Beyond Human Intuition

AI-Driven platforms and the future of Molecular Binding Design

Hannah Sanford-Crane, Founder, MAAD Scientist Technologies

AI-driven platforms are beginning to predict novel molecular binding interactions that traditional approaches would miss entirely, redefining what is possible in medicinal chemistry and binder development. Realising this potential demands disciplined computational-to-experimental workflows, robust wet-lab validation, and vigilance against training bias. Built responsibly, these tools could reshape pharmaceutical R&D accessibility worldwide.

novel molecular binding interactions

A New Lens on Molecular Interaction

In late 2024, Nabla Bio demonstrated the first computationally designed antibodies capable of binding G protein-coupled receptors (GPCRs)—a class of membrane proteins targeted by roughly a third of all approved small-molecule drugs, but notoriously resistant to antibody-based approaches due to their minimal extracellular exposure. Conventional methods for generating antibodies against GPCRs rely on immunisation campaigns or library screening, both constrained by throughput and by the biases of the biological systems they depend on. The company’s generative AI platform, JAM, took a different approach, computationally designing binders against multipass membrane protein targets including CXCR7 and Claudin-4, with hit rates exceeding those of conventional campaigns. In 2025, the company scaled this approach with JAM-2, reporting hits across sixteen previously unseen soluble protein targets in initial assays using a four-person team operating in parallel over less than a month.

Whether these computational hits translate to clinical-grade leads remains to be demonstrated through further validation, and independent analyses have noted gaps in the published methodology, including missing controls and unverified epitope targeting claims [19]. But as a proof of principle, the results were striking: a generative model had navigated molecular spaces too vast for manual analysis and returned candidates that conventional pipelines, constrained by time and throughput, were unlikely to have reached.

The implications extend well beyond speed. These platforms can specify binding epitopes upfront, enabling researchers to direct antibodies to precise locations on a target protein rather than accepting whatever the immune system or a random library happens to produce. This level of intentional design opens access to targets previously considered undruggable and allows systematic exploration of epitope biology that was simply not feasible with earlier methods. The 2024 Nobel Prize in Chemistry, awarded for breakthroughs in protein structure prediction and AI-designed proteins, cemented the scientific legitimacy of the field.

Meanwhile, AI-driven small molecule discovery platforms are compressing lead generation timelines and identifying novel scaffolds optimised for binding affinity, metabolic stability, and synthetic accessibility simultaneously. Several AI-designed drug candidates have entered clinical trials, with the first fully AI-discovered molecule reaching Phase IIa results in 2025. Multi-billion-dollar partnerships between AI startups and major pharmaceutical companies — including Isomorphic Labs’ collaborations with Eli Lilly and Novartis, collectively valued at approximately $3 billion in upfront and milestone payments — signal serious industry commitment to the approach.

Discipline at the Interface of Computation and Experiment

Discipline at the Interface of Computation and Experiment

The platforms that have demonstrated genuine results share a common architecture: tight integration between generative AI and experimental validation, with clearly defined handoff points between computational prediction and wet-lab confirmation. This integration is not a nice-to-have. It is structural.

Generative models trained on protein data learn statistical patterns, not physical laws. They can produce molecular structures that are statistically plausible but biologically incorrect — confidently predicting ordered structure in regions that are experimentally disordered, or assigning high-resolution features to flexible loops that adopt no fixed conformation. A 2025 preprint evaluating AlphaFold 3’s performance on intrinsically disordered proteins found that 32 per cent of residues were misaligned with experimental annotations, with 22 per cent representing outright hallucinations. Eighteen per cent of residues associated with biological processes showed hallucinations, raising direct concerns for drug discovery and target identification. These findings are consistent with the original AlphaFold 3 publication, which acknowledged that the model’s diffusion-based architecture can produce spurious structural order in disordered regions — hallucinations that may form ordered secondary structures rather than expected disordered conformations, making them harder to identify by visual inspection alone [1, Extended Data Fig. 1]. Independent benchmarking has further confirmed that AlphaFold 2 remains the more reliable tool for identifying intrinsically disordered regions, precisely because AlphaFold 3’s architectural changes introduce new categories of structural hallucination (Figure 1).

A binding prediction with a strong confidence score looks indistinguishable from a validated result — until someone synthesises the molecule and tests it. Without rigorous wet-lab validation loops built into the workflow from the outset, research teams risk spending months pursuing computationally generated leads that were never viable. The companies demonstrating the strongest results have built for this. Leading platforms generate thousands of protein designs in mammalian systems, measure binding, developability, and cellular function in contexts that closely reflect human biology, and feed experimental data back into model training — a continuous cycle where the AI improves based on real-world outcomes rather than purely computational benchmarks.

But this necessary discipline carries a cost that the field has not yet fully reckoned with. The entire promise of AI in molecular design is that it can venture beyond what human intuition and established chemistry would predict, surfacing binding interactions in unexplored regions of proteomic space and proposing scaffolds that no medicinal chemist would conceive. Yet conservative confidence thresholds filter out not only hallucinated predictions but also genuinely novel ones that score low precisely because they are unlike anything in the training data. Anchoring to established datasets biases models toward well-characterised target families. Validation frameworks preferentially confirm predictions that resemble what conventional methods already produce. A model trained predominantly on kinase inhibitor data will excel at proposing kinase inhibitor variants and struggle with anything that looks different — not because the different prediction is wrong, but because the model has no basis for confidence in territory it has never mapped. The very constraints that keep AI from being wrong also keep it from being new (Figure 2).

In early 2025, two prominent AI-native biotechs faced public scrutiny over antibody candidates described as “de novo” — designed from scratch by AI. Independent critics argued that one company’s Phase I candidate was only a few mutations from an established antibody, and another’s preprint overstated the novelty of its designs. Whether these cases represent genuine failures of novelty or legitimate AI-assisted optimisation being held to an impossibly strict standard remains actively debated. The definition of “de novo” itself is contested: some researchers require that a computationally designed antibody share no sequence or structural homology with known binders, while others accept that starting from learned representations of antibody space and generating functional variants constitutes genuine design. But beneath the definitional dispute lies a harder problem the field has not resolved: there are no agreed-upon criteria for what counts as a genuinely novel computationally designed molecule, and existing validation frameworks are not built to make that distinction. Until they are, every claim of AI-driven novelty will invite the same question — and there will be no principled way to answer it.

 AI-driven novelty

What the Models Inherit

AI models for drug discovery learn from published academic research, curated databases such as ChEMBL and the Protein Data Bank, gene expression profiles, and compound-target relationship datasets. These sources are treated as ground truth.

They are not.

The biomedical research literature carries a well-documented reproducibility problem. In a widely cited internal review, Amgen attempted to independently reproduce 53 high-profile preclinical oncology studies and succeeded with only six. Reproducibility rates vary across therapeutic areas and methodologies, and preclinical oncology represents a particularly challenging subfield. But the underlying issues — statistical practices, selective reporting, methodological variability — pervade published research more broadly. Much of the historical data in compound-target relationship databases was generated under unstandardised conditions, with incomplete metadata and limited provenance tracking. In human research networks, non-reproducible work gets filtered informally: a postdoc mentions at a conference that a key finding didn’t hold up in their hands; a laboratory head steers students away from a published protocol that quietly failed replication; institutional memory accumulates around which results are trustworthy and which are not. An AI model trained on the published literature has no access to any of this. It absorbs everything with equal weight, building predictions on foundations that may include the very studies the scientific community has learned to treat with caution.

The problem compounds. No established mechanism currently exists to flag computationally generated entries in major molecular databases or to track their provenance through downstream model training. When generative models produce plausible but incorrect molecular structures — binding predictions that look convincing but do not survive experimental testing — those outputs can enter the broader data ecosystem. If hallucinated results are published, deposited in databases, or used to train subsequent models, each generation of AI tools could train on data that includes the confident errors of the previous generation. This feedback loop is not yet widely documented, but the infrastructure to prevent it does not yet exist either (Figure 1).

Training data bias adds another dimension. AI models trained predominantly on data from specific populations or well-studied protein families systematically underperform on underrepresented groups and novel chemical spaces. Multiple studies have documented significant accuracy differentials across ethnic populations in AI diagnostic models, and similar patterns emerge in drug discovery: tools optimised for well-characterised target classes such as kinases or nuclear receptors perform well within those families but degrade when applied to less-studied areas of the proteome. The irony is precise — the targets where AI-driven design could add the most value, by definition, sit in the regions of biological space where training data is thinnest and model confidence is lowest. And fewer than one-third of AI research papers are themselves reproducible, with issues ranging from unpublished code to sensitivity to training conditions. A 2024 replication study found that 86 per cent of AI papers sharing both code and data could be reproduced, compared to only 33 per cent of those sharing data alone — underscoring how much of the reproducibility gap stems from incomplete methodological transparency rather than fundamental scientific failure. When AI tools that are difficult to reproduce are applied to biomedical data carrying its own reproducibility problems, the uncertainty does not cancel. It multiplies.

Building from Honest Ground

None of these problems are reasons to abandon AI-driven molecular design. They are reasons to build it differently. Adding more validation checkpoints will not resolve the tension between preventing hallucination and enabling discovery. The field needs a fundamentally different approach to confidence assessment — frameworks that can distinguish a genuinely novel binding prediction from a confident hallucination without defaulting to the assumption that anything unfamiliar is wrong. Independent metrics beyond standard confidence scores. Interpretability built into generative architectures so researchers can interrogate why a prediction was made — which training examples influenced it, which structural features drove the score — not just how confident the model was. Diverse, well-characterised experimental datasets that allow models to be tested against biology they have never seen, including targets from underrepresented protein families and understudied therapeutic areas.

Early work points toward what this might look like. Ensemble disagreement methods — running multiple models and examining where their predictions diverge — can flag regions of genuine uncertainty rather than collapsing everything into a single confidence score. Out-of-distribution detection techniques can distinguish between a prediction the model finds unfamiliar because it is wrong and one it finds unfamiliar because nothing in the training data resembles it. Experimental designs that deliberately probe model behaviour at the boundaries of training distributions, rather than validating only within comfortable territory, would begin to map where AI-driven discovery actually extends the frontier and where it merely retraces known ground. None of these approaches are mature. All of them are buildable.

The provenance of training data deserves equal scrutiny. Bias auditing— for demographic representation, chemical diversity, and the reproducibility status of source studies—is not a compliance exercise. It is a prerequisite for trusting what these models produce. Acknowledging how difficult this is, honestly, remains more productive than pretending current approaches have solved it.

Access matters too, and not only as an equity argument. Synthetic binders designed through computational platforms could replace antibody-based reagents that currently cost hundreds of dollars per vial, offering comparable specificity at a fraction of the price with room-temperature stability that eliminates cold-chain requirements. But cheaper, more stable reagents also change what is experimentally feasible. The diverse validation panels this field needs—testing AI predictions against underrepresented protein families, novel target classes, biology the models have never seen—become economically realistic only when the per-target cost of running those assays drops dramatically. Public benefit research organisations and open-access foundries have a role to play here: developing and validating proof-of-concept binding technologies not for proprietary advantage but for broad licensing and integration into the scientific commons, so that the tools required to stress-test AI predictions are not locked behind the same cost barriers that limit who can use AI-driven discovery in the first place.

The same logic applies to how the field organises itself. Open-source models such as MIT’s BoltzGen are making generative protein design accessible without proprietary licences. BoltzGen’s all-atom generative diffusion framework—experimentally validated across nanobodies, minibinders, peptides, and cyclic peptides against diverse and novel targets through distributed academic and industry collaboration—demonstrates that frontier-grade binder design tools can be developed, shared, and improved openly rather than locked behind commercial platforms. Partnership-driven approaches that cross institutional and geographic boundaries expand not just who can use these tools but the diversity of the experimental data that flows back into them. AI-driven molecular design will not mature inside silos. The models need broader biology, and broader biology requires open infrastructure.

Regulatory frameworks are beginning to respond, though their scope reveals the same tension the field itself faces. The United States Food and Drug Administration issued draft guidance in January 2025 on AI in drug development, establishing a risk-based credibility assessment framework but explicitly excluding AI used in early discovery. The European Medicines Agency’s 2024 Reflection Paper takes broader lifecycle coverage, highlighting data bias and human oversight. The early-discovery exclusion is telling: regulators, like the models themselves, lack established frameworks for assessing novelty. Design decisions made at that stage shape everything downstream, yet the tools for evaluating whether an AI-generated molecule represents genuine innovation or confident pattern-matching do not yet exist at any level not in the models, not in the validation pipelines, and not in the regulatory infrastructure.

AI-driven platforms offer a fundamentally new way to navigate the gap between academic innovation and clinical application, processing complexity at scales no human team could match, identifying connections that would otherwise remain invisible, making the tools of molecular discovery available to researchers who have historically been priced out. The field has not yet built what it needs most: frameworks that can tell the difference between a prediction that is wrong and a prediction that is merely new. The ambiguity in every confidence score is real, and no amount of computational power alone will resolve it. But it is an ambiguity the field can learn to work with, through better tools, more honest benchmarks, open infrastructure, and the willingness to build for what we do not yet know rather than only for what we do.

References

1. Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, 630, 493–500 (2024).
2. Gopalan, S. and Narayanan, S. Hallucinations in AlphaFold 3 for Intrinsically Disordered Proteins with disorder in Biological Process Residues. arXiv:2510.15939 [preprint] (2025).
3. Nabla Bio. De novo design of epitope-specific antibodies against soluble and multipass membrane proteins with high specificity, developability, and function. Technical Report (2024).
4. Begley, C.G. and Ellis, L.M. Raise standards for preclinical cancer research. Nature, 483, 531–533 (2012).
5. Lin, F. Scratch That? De Novo Antibody Design Enters the AI Drug Discovery Toolbox. GEN Biotechnology (2025).
6. Niazi, S.K. and Mariam, Z. Artificial intelligence in drug development: reshaping the therapeutic landscape. Therapeutic Innovation & Regulatory Science (2025).
7. Hutson, M. Artificial intelligence faces reproducibility crisis. Science, 359, 725–726 (2018).
8. Fan, X., Wu, J. and Wang, L. Exploring the ethical issues posed by AI and big data technologies in drug development. Frontiers in Pharmacology (2025).
9. Drug Target Review. AI in drug discovery: 2025 in review (2025).
10. Stärk, H. et al. BoltzGen: Toward Universal Binder Design. MIT Jameel Clinic (2025). Available at: https://github.com/HannesStark/boltzgen
11. Isomorphic Labs. Isomorphic Labs kicks off 2024 with two pharmaceutical collaborations [press release]. 7 January 2024. Available at: https://www.isomorphiclabs.com/articles/isomorphic-labs-kicks-off-2024-with-two-pharmaceutical-collaborations
12. Semmelrock, L. et al. Reproducibility in machine-learning-based research: Overview, barriers, and drivers. AI Magazine (2025).
13. US Food and Drug Administration. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products. Draft Guidance (January 2025).
14. European Medicines Agency. Reflection Paper on the Use of Artificial Intelligence in the Medicinal Product Lifecycle (2024).
15. Krokidis, M.G. et al. AlphaFold3: An Overview of Applications and Performance Insights. International Journal of Molecular Sciences, 26(8), 3671 (2025).
16. Fang, Z. et al. AlphaFold 3: an unprecedented opportunity for fundamental research and drug development. Precision Clinical Medicine, 8(3) (2025).
17. Semmelrock, L. et al. The Unreasonable Effectiveness of Open Science in AI: A Replication Study. arXiv:2412.17859 (2024).
18. Nabla Bio. JAM-2: Drug-like antibodies straight from the computer. Technical Report (2025).
19. Yapici, E. I Stress-Tested the Claims in Nabla’s JAM-2 Paper. These Are the Gaps. Medium (December 2025).

--PFA Issue 63--

Author Bio

Hannah Sanford-Crane

Dr. Hannah Sanford-Crane is the founder of MAAD Scientist Technologies, a Canadian public benefit research foundry specialising in surface chemistry, bioconjugation, and bioorthogonal click chemistry. Her work focuses on developing affordable, accessible nanotech-based research tools through proof-of-concept development and IP licensing partnerships with academic and industry collaborators.