MerIA: Artificial Intelligence and Bioinformatics to Accelerate the Identification of Mercury-Resistance Genes in Latin America

MerIA: Artificial Intelligence and Bioinformatics to Accelerate the Identification of Mercury-Resistance Genes in Latin America

Mercury pollution is one of the most persistent environmental and health problems in Latin America, with significant impacts on ecosystems and communities exposed to this highly toxic metal, particularly those involved in artisanal mining. Kidney problems, neurological disorders and foetal malformations are among its most severe manifestations. Faced with this challenge, the Ecuadorian MerIA (Mercury Identification through Artificial Intelligence) project, which benefited from the capabilities of BELLA II’s Bioinformatics and Artificial Intelligence testbed through the Early Adopters BIO+IA programme launched in March 2026, utilised this testbed to investigate how certain microorganisms manage to resist and detoxify mercury. Indeed, the solution lies in nature itself, but the research data supporting this is so scattered that analysing it is almost a feat in itself.

 

Uncontrollable tremors of the hands and eyelids, memory loss, difficulty concentrating, emotional changes, problems with motor coordination, peripheral vision and balance, permanent kidney damage, intestinal irritation, erosive gastritis, severe diarrhoea, dermatitis and skin rashes, delayed brain and nervous system development in foetuses and children, cognitive problems, and irreversible neurological damage that can cause paralysis or extreme weakness, known as Minamata disease. These are some of the effects that mercury pollution has on people, a cumulative process caused by breathing in vapours (common in Latin American mining areas), absorption through the skin and even by frequently eating contaminated fish.

There is no annual register of people contaminated by mercury; in fact, the United Nations (UN) and prestigious scientific journals have highlighted an alarming lack of consolidated data and standardised medical records in the region. Artisanal – which accounts for the highest number of contaminated people – and illegal mining present staggering figures across Latin America. And that is not all: in August 2025, in a report by Deutsche Welle, journalist Judit Alonso stated that “from April 2019 to June 2025, approximately 200 tonnes of mercury were trafficked in Latin America, according to the US Environmental Investigation Agency”. Clearly, the number of people affected will not fall rapidly, even though that was the goal set on 10 October 2013 in Kumamoto and Minamata, Japan, by the 128 countries and the European Union that signed the Minamata Convention on mercury.

Understanding the biological mechanisms that enable certain microorganisms to resist and detoxify this metal represents a significant opportunity for the development of bioremediation strategies on the continent.

In this context, the MerIA team sought to conduct an in-depth study of the mer operon. ‘The operon what?’, you might ask. The mer operon is a set of genes found in bacteria that confers resistance to mercury toxicity. Its main function is to transform highly toxic mercury compounds into less harmful or volatile forms that the organism can expel. Explained in this way, the subject does not seem so complex; however, the information required to carry out this analysis was scattered across scientific literature, biomolecular databases, specialised repositories and logical knowledge bases, making its integration and validation difficult.

The central hypothesis of the project was to determine whether the combination of bioinformatics and artificial intelligence would enable the automatic identification and validation of the function of genes in the mer operon, based on scientific evidence available in the specialist literature.

In just eight weeks, a team from the Chimborazo Higher Polytechnic School (ESPOCH), led by Celso Recalde and comprising three researchers and five research assistants, succeeded in significantly optimising the analysis of scientific information relating to the biological detoxification of mercury. Using advanced bioinformatics and artificial intelligence tools, the group filtered more than 10,000 initial records to build a reliable knowledge base, from which 12 scientific hypotheses were formulated with a view to strengthening future bioremediation strategies for ecosystems affected by this pollutant.

This work was made possible by access to specialised computing infrastructure provided through CEDIA and RedCLARA, as part of the BELLA II project and the Early Adopters BIO+IA programme. This support enabled the team to focus their efforts on scientific research, utilising advanced technological capabilities without the need to deploy and maintain their own complex computing environments.

The MerIA experience demonstrates the potential of collaboration between human talent and artificial intelligence to tackle complex challenges affecting Latin America. Whilst AI tools contributed to the processing and analysis of over 6,700 scientific publications, the validation of results, the definition of methodological criteria and the interpretation of findings remained in the hands of the researchers. All of this was carried out within an Open Science framework, promoting transparency, reproducibility and the sharing of knowledge and methodologies with the regional scientific community.

For the ESPOCH team, access to the BIO+IA testbed proved crucial, as it enabled them to tackle a computational challenge that would have been difficult to implement independently.

“We applied to the Early Adopters programme because the problem we were facing—identifying and validating the functionality of the mer operon genes amongst thousands of noisy NCBI sequences—requires exactly what the testbed offers: a computational environment capable of supporting bioinformatics pipelines combined with automated reading of scientific literature using AI. As a research group, this presented an implementation challenge that involved setting up large-scale computational tools ourselves. The value of the proposal lay in the fact that it allowed us to move from manual validation—which was slow and dependent on expert judgement—to a reproducible workflow: analysing the mechanism and functionality of the mer operon, processing the full text of the articles using AI, and structuring the evidence (gene function, organism, type of evidence, level of confidence). The testbed tied in directly with our objective of understanding and building a dataset on the functionality and behaviour of mercury-resistance genes with implications for bioremediation in the region’s mining areas,” explains Celso Guillermo Recalde Moreno, project coordinator.

The solution: An AI-driven knowledge pipeline

The project developed a workflow (pipeline) capable of combining multiple sources of information, including knowledge bases in Prolog, annotations from PubTator, articles indexed in PubMed and data from specialised bioinformatics repositories.

The methodology integrated the following processes:

  • Collection and enrichment of expert knowledge.
  • Curation and alignment of biological entities.
  • Processing of scientific literature using artificial intelligence.
  • Construction of knowledge bases and inference of biological pathways.
  • Automated generation of reports and scientific hypotheses.
  • Validation and refinement of regulatory relationships associated with the mer operon.

Through the testbed, the team was able to structure the work into successive phases of logical extraction, narrative convergence, contextualisation using PubTator, hypothesis generation and strengthening of its knowledge base. This approach made it possible to combine expert knowledge with advanced automated processing capabilities.

Key results

In just eight weeks of work within the BIO+IA Early Adopters Programme, the project achieved significant results from both a scientific and methodological perspective.

Key achievements include:

  • Identification and alignment of 58 expert entities relevant to the analysis.
  • Processing of over 10,000 events contained in the original knowledge base.
  • Analysis of 6,781 PubMed identifiers.
  • Confirmation of 473 regulatory pathways associated with the mer operon.
  • Reconstruction of regulatory events supported by scientific evidence.
  • Generation of 12 testable hypotheses relating to the function and regulation of mercury resistance genes.
  • Prioritisation of factors according to available levels of scientific evidence.

One of the most significant results was the refinement of an initial database containing over 10,000 events to consolidate a core of around 1,000 events supported by robust evidence, thereby improving the quality and traceability of the information used by the project.

Beyond the quantitative indicators, the main contribution of the Early Adopters programme was to enable a profound methodological transformation. According to the research team, the experience made it possible to move from a manual, slow process that was highly dependent on expert judgement to a reproducible, documented and scalable workflow. This evolution enabled them to work with large volumes of biological data and scientific literature without compromising consistency or the ability to validate findings.

Do you think that having taken part in Early Adopters BIO+IA and used the testbed was beneficial for you? The question put to the research team was answered by its coordinator, Celso Recalde: “Yes, it was undoubtedly beneficial; we managed to move from having a preliminary hypothesis scattered across unstructured literature to having concrete and verifiable results: the reconstruction of confirmed regulatory events in the mer operon, 58 expert entities aligned with PubTator, and 12 testable hypotheses, including the effects of mer gene deficiency and also the factors that are prioritised according to their level of evidence. We also defined clear criteria for influencing gene function within the mer operon mechanism, in order to refine our knowledge base (from over 10,000 events down to a core of ~1,000 supported by evidence). For the next phase, we have identified the remaining gaps in the literature: SIRT1, ergothioneine and ferroptosis. Beyond the technical results, the main difference was methodological: we moved from a manual, ad-hoc process to a reproducible and documented pipeline, aligned with the principles of open science, which we can now share with the regional scientific community, as well as potentially exploring this methodology in other heavy metal resistance genes in the future.”

Furthermore, participation in the Early Adopters BIO+IA programme facilitated exchange with a regional community specialising in bioinformatics and applied artificial intelligence, strengthening the validation of the scientific approach and providing new perspectives for the project’s development.

Impact on open science and the region

MerIA demonstrates how the combination of artificial intelligence, bioinformatics and regional digital infrastructure can accelerate the generation of scientific knowledge with potential environmental and social impact.

The results obtained open up the possibility of improving future bioremediation strategies in areas affected by mining pollution and lay the foundations for studying other mechanisms of heavy metal resistance.

Following the completion of the Early Adopters BIO+IA Programme, the team plans to:

  • Openly publish the dataset and the methodology developed.
  • Experimentally validating the hypotheses with the strongest scientific support.
  • Extending the approach to other detoxification mechanisms, including operons associated with arsenic and zinc.
  • Exploring applications related to biosensors, gene editing and new biotechnological strategies for the treatment of environmental pollutants.

MerIA demonstrates the potential of the testbeds driven by BELLA II to accelerate scientific research in Latin America and the Caribbean. By providing access to specialised infrastructure, expert knowledge and opportunities for regional collaboration, the BIO+IA Early Adopters Programme enabled a complex bioinformatics challenge to be transformed into a reproducible platform for knowledge generation.

This experience confirms that regional cooperation, promoted by RedCLARA and its partners, can act as a catalyst for scientific innovation, strengthening local capacities and helping to address challenges relevant to the region through the strategic use of advanced technologies.

ACKNOWLEDGEMENTS

BELLA II receives funding from the European Union through the Neighbourhood, Development and International Cooperation Instrument (NDICI), under agreement number 438-964 with DG-INTPA, signed in December 2022. The implementation period of BELLA II is 48 months.

Contact

For more information about BELLA II please contact:

redclara_comunica@redclara.net