Combining Machine Learning and Iterative Experiments to keep pace with Emerging Viral Variants of Concern

Thomas Sheffield, Ryan C. Bruneau, Stephen Won, Kenneth L. Sale, Brooke Harmon, Le Thanh Mai Pham

Abstract

Modeling and predicting viral mutations before they emerge plays a crucial role in pandemic preparedness, enabling the early identification of emerging variants of concern (VOCs) and guiding timely updates to vaccines, diagnostic tests, and therapeutic strategies. However, existing machine learning models and large-scale experiments lose their predictive power as viral variants evolve further from the original strains in sequence space. Here, we present a scalable framework that integrates random forest and neural network machine learning models with targeted high-throughput experimentation to anticipate and evaluate emerging SARS-CoV-2 receptor-binding domain (RBD) variants.

Introduction

The COVID-19 pandemic, caused by the SARS-CoV-2 virus, profoundly affected global health and the economy, driving unprecedented efforts to mitigate its impact. One critical area of focus has been the discovery and development of neutralizing antibodies, which play a key role in preventing and treating infection. These antibodies target the SARS-CoV-2 spike protein, blocking its ability to bind the ACE2 receptor on human cells and replicate. However, the rapid evolution of SARS-CoV-2, with the emergence of variants carrying mutations in the spike protein, has posed significant challenges to ensuring the sustained efficacy of these antibodies.

Materials and Methods:

There were six different datasets that were used to build machine learning models: three came from public datasets (“PEx”, “PACE”, “PAnti”) and four came from in-house experiments (“I1”, “I2”, “I3”, “IACE”). PEx and PACE were drawn from Starr et al. (11) and contained RBD expression and ACE2 binding data, respectively. PAnti was drawn from Greaney et al. (10) and contained antibody binding data for various antibodies.

Discussion:

We successfully developed the PEx_NN, PACE_NN, and PAnti_NN ML models for RBD expression, ACE2 binding, and antibody escape using publicly available datasets and employed them to guide targeted experiments, and data from these experiments was in turn used to refine the models. This iterative framework enabled adaptation to emerging SARS-CoV-2 variants, including those outside the sequence space of previously characterized datasets. The cycle can be repeated indefinitely to maintain alignment with ongoing viral evolution. We implemented the global epistasis model in Starr et al. on a combined PACE and I3 dataset and used fivefold cross-validation to estimate its accuracy.

Acknowledgments:

We sincerely thank Umakant Mishra for his thoughtful review and constructive feedback, which significantly improved the clarity and quality of this manuscript.

Citation: Sheffield T, Bruneau RC, Won S, Sale KL, Harmon B, Pham LTM (2026) Combining machine learning and iterative experiments to keep pace with emerging viral variants of concern. PLoS Comput Biol 22(6): e1014394. https://doi.org/10.1371/journal.pcbi.1014394

Editor: Eric C. Dykeman, University of York, UNITED KINGDOM OF GREAT BRITAIN AND NORTHERN IRELAND

Received: August 4, 2025; Accepted: June 2, 2026; Published: June 17, 2026.

Copyright: This is an open access article, free of all copyright, and may be freely reproduced, distributed, transmitted, modified, built upon, or otherwise used by anyone for any lawful purpose. The work is made available under the Creative Commons CC0 public domain dedication.

Data Availability: The data used in the submission is all included in the manuscript, supporting Information and references in the manuscript. All code used for model training, evaluation, and figure generation, along with preprocessed datasets and trained model weights, is available at https://github.com/sandialabs/IterML-for-VOCs.

Funding: This work was supported by the Laboratory Directed Research and Development Program at Sandia National Laboratories: Project 225922 and Project 233120. L.T.M.P., T.S., and K.L.S. were funded by project 225922 and Project 233120. B.H., R.C.B., and S.W. were partly funded by project 225922. This paper describes objective technical results and analysis. Any subjective views or opinions that might be expressed in the paper do not necessarily represent the views of the U.S. Department of Energy or the United States Government. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Competing interests: The authors have declared that no competing interests exist.