Edit

Benchmarking: the foundation of reliable AI in biomedicine

Establishing shared evaluation standards to ensure AI models in biology are rigorous, reliable, and reproducible

Credit: Karen Arnott, EMBL-EBI

As AI becomes embedded in scientific research, the life sciences community is facing a challenge: how can we ensure AI models are delivering accurate and reproducible results? 

AI systems are increasingly used for a range of biomedical applications: to interpret genomes, analyse gene expression, predict protein structure, and accelerate drug discovery and literature research, among many others. But as these models grow more powerful and complex, evaluating them becomes harder. Without standardised assessments known as benchmarks, we can’t compare models effectively, reproduce findings, or determine whether their predictions reflect genuine biological insight.

In a recently published Perspective titled “Benchmarking biomedical foundation models”,  Julio Saez-Rodriguez, Head of Research at EMBL-EBI; Philipp Sven Lars Schäfer, PhD student in the Saez-Rodriguez Group; Nikolas Kalavros and Gustavo Stolovitzky at NYU discuss the role of benchmarking, providing a set of recommendations for scientists to perform rigorous comparisons and ensure reproducibility. The perspective explores the benchmarking of foundation models and its limitations. Foundation models are large-scale machine learning systems trained on broad, heterogeneous datasets. 

Below, Saez-Rodriguez explains why benchmarking for AI models is more important than ever. He is also a co-founder and member of the governance committee of BEACON (Benchmarking, Evaluation, and Assessment Consortium for Science), a new global initiative focused on AI benchmarking. 

What is benchmarking in the context of AI models?

Benchmarking is the structured and systematic assessment of the performance of methods used to predict, detect and characterise systems. In our context, it’s about determining how accurately an algorithm answers a specific biological question. While bioinformatics has long relied on benchmarking to evaluate and compare methods, the rise of generative and foundation-model based AI requires us to rethink this strategy as these models can answer all sorts of questions.

Why is benchmarking important for AI development?

AI models today are more complex and less interpretable than the algorithms life scientists used in the past. Those models were built to answer single, well-defined questions. AI foundation models are different because they are designed for multi-tasking across an entire domain. 

In single-cell biology, for example, a single AI model might be expected to handle everything from annotation of cell types to predicting the effects of genetic perturbations. Because these models can perform many tasks, benchmarking becomes more challenging. Our testing must evolve into dynamic, cross-task frameworks and include tasks not originally anticipated. 

The benchmarking triangle

To achieve reproducible AI models, scientists rely on three interconnected pillars:

  • Data quality: High-quality, curated, and openly accessible data used to train and assess AI models. Without reliable data, evaluation is impossible. 
  • Model repositories: Centralised hubs that host validated AI models to ensure they are accessible to the global community. 
  • Benchmarking frameworks: Standardised shared frameworks for evaluating model performance, such as those coordinated by BEACON. 

How do we ensure rigorous and unbiased evaluation?

To ensure that AI evaluations are fair and unbiased, we need to isolate performance assessment from method development. Developers should train and refine their methods without access to a withheld dataset reserved exclusively for final evaluation. Although developers must assess their methods during development, relying solely on them to determine final performance creates a conflict of interest. It’s effectively asking the same party to develop the method, define the test, and judge the result. Independent evaluation on previously unseen data provides a more credible measure of performance and generalisability.

Supporting this essential practice, community-driven benchmark initiatives such as CASP and DREAM act as independent referees. Scientists submit their AI models, and organisers test them on previously unseen data under standardised conditions. 

Carefully choosing the performance metrics is key for benchmarking. For example, first-generation single-cell foundation models have been shown to not outperform linear models in predicting the effect of perturbations on single-cells on one commonly used measure, whereas emerging studies argued that such conclusions are metric-dependent, and deep-learning models may show improved performance if other metrics are used. This illustrates how conclusions about model performance depend not only on the task, but also on the evaluation protocol. We have thus built a taxonomy of metrics with a corresponding toolbox to facilitate exploring the impact of metrics and to help align the choice of metrics with the intended downstream application of perturbation prediction models.

By testing models under identical, unbiased conditions and metrics, benchmarking gives scientists rigorous information to confidently choose the most reliable AI tools for their research.

What is BEACON’s role in this landscape? 

BEACON provides an umbrella structure that connects existing initiatives like CASP, DREAM Challenges, Sage Bionetworks, OpenADMET, and Conscience. 

For example, CASP, which stands for Critical Assessment of Structure Prediction, is a biennial community initiative where researchers blind-test protein structure prediction algorithms against newly experimentally solved protein structures. CASP served as a catalyst for advances in AI-driven modelling, with Google DeepMind’s AlphaFold representing an example of success. AlphaFold predictions are now openly available and widely used by the scientific community through the AlphaFold Database, co-developed with EMBL-EBI. In practice, these initiatives work like standardised open competitions. The organisers set a specific scientific challenge, provide clean datasets for scientists to train their AI models, and then evaluate the models on unseen data to test their accuracy. 

Rigorous benchmarking helps define the best approaches to analyse data, in particular publicly available data, such as the resources available at EMBL-EBI. By strengthening how we evaluate predictive models, we ensure that AI delivers reliable, real-world impact for science and medicine. BEACON will thus accelerate research across biomedical domains.

How does benchmarking link to EMBL’s wider AI strategy? 

The EMBL AI Strategy is based on innovative methodological development, practical application of AI to existing data resources, and the use of AI in the scientific process, including in lab-based experiments. 

The BioAIrepo is a core component of this strategy. It acts as a curated library for the life sciences community, aiming to aggregate information from other AI repositories, such as EMBL’s BioImage Model Zoo, in a single location. By using the insights gained from benchmarking, we can identify high-performing models and make them easily accessible to the global scientific community. 

Some EMBL research groups focus strongly on method development. For them, this approach provides a standardised environment where they can develop and assess their AI models with rigour. Other groups may deploy such methods to answer biological questions. For them, it is important to understand the limits of existing models to be able to make informed decisions when choosing and using models.

Looking ahead, how do you see the future of benchmarking unfolding?

We need rigorous assessment to understand both the power and the limitations of AI models in the life sciences. 

The success of AlphaFold, facilitated by data from the Protein Data Bank (PDBe) and the framework of CASP, showed us a path forward. While image analysis and multi-omics already have active modelling and benchmarking communities, we aim to amplify existing efforts to strengthen and standardise benchmarking practices. To achieve this, we encourage researchers to continue actively contributing to community benchmarking efforts by submitting their AI models to catalogues such as BioAIrepo.

Another key area will be the benchmark of large language model-based agentic frameworks, tools that are increasingly used to perform a wide range of analyses autonomously. Here, benchmarking is even more complex as we need to evaluate not just the output of the model, but how it got there. This is something we are starting to address with the Karenina framework together with OpenTargets.

By establishing these benchmarks, we ensure that as AI reshapes biological research, the results remain as rigorous, transparent, and reproducible.


Tags: alphafold, benchmarking, bioinformatics, embl-ebi, saez-rodriguez

News archive

EMBLetc archive

News archive

For press

Contact the Press Office
Edit