AI & Machine Learning

How AI Is Discovering New Drugs: The Real Pipeline from Lab to FDA

Drug discovery has a productivity problem. Bringing a single new drug to market takes roughly a decade and costs well over a billion dollars, and the vast majority of candidates fail. The attrition is brutal, and much of it happens in early phases when the biology turns out to be wrong or the molecule turns out to be toxic or ineffective. Artificial intelligence promises to compress that timeline and reduce that waste. The question is whether the promise is real, and after several years of clinical data, the answer is becoming more nuanced but more encouraging.

AlphaFold and the Protein Revolution

In 2020, DeepMind’s AlphaFold 2 solved a problem that had stumped biologists for fifty years: predicting a protein’s three-dimensional structure from its amino acid sequence. The Protein Data Bank had been built painstakingly over decades, but predicting structure computationally had been intractable. AlphaFold 2 achieved accuracy that rivalled experimental methods.

DeepMind then released predicted structures for essentially all known proteins — over 200 million — through the AlphaFold Protein Structure Database, a resource made freely available. In 2023, DeepMind and Isomorphic Labs released AlphaFold 3, extending the model beyond proteins to handle nucleic acids, ligands, ions, and chemical modifications — richer detail that matters enormously for drug design, because most drugs work by binding to proteins.

The impact was recognised by the 2024 Nobel Prize in Chemistry, awarded to Demis Hassabis and John Jumper of DeepMind and David Baker of the University of Washington for their work on protein structure prediction and design.

AlphaFold alone has not created a drug. But it has given researchers a powerful tool — a structural starting point for rational drug design, and a means to prioritise which proteins to target.

AI-Designed Molecules in Clinical Trials

The more consequential milestone is AI-designed molecules reaching clinical trials. The leading example is Insilico Medicine, a company with roots in Hong Kong and offices worldwide, which used its generative AI platform to design a molecule for idiopathic pulmonary fibrosis. The drug, ISM001-055 (rentosertib), targets TNIK, a kinase involved in fibrosis. It was designed using AI that proposed novel molecular structures targeting a novel target.

The significance lies in two things: the target was discovered by AI (a “first-in-class” mechanism), and the molecule was designed by AI. ISM001-055 advanced into Phase 2a trials, and positive topline results were reported in 2025 — patients showed improvement in lung function, and the drug appeared well tolerated. This is among the most advanced real-world validations of AI-driven, end-to-end drug discovery. It is not proof that AI makes better drugs, and the trial was small, but it is a genuine milestone: a molecule designed by AI, targeting a target identified by AI, showing clinical activity in humans.

Other Players and Approaches

The field includes dozens of companies. Recursion Pharmaceuticals combines AI with high-throughput cell imaging and wet-lab experiments, creating a data flywheel; it has multiple candidates in clinical trials and acquired competitor Exscientia in 2024. Isomorphic Labs, spun out of DeepMind, has struck partnerships with Eli Lilly and Novartis worth billions in potential milestone payments to apply AlphaFold-based design to real drug programs.

BenevolentAI (UK) has had a mixed record — some programs advanced, others failed — illustrating that AI reduces but does not eliminate discovery risk. Exscientia was among the first to put an AI-designed molecule into clinical trials (for obsessive-compulsive disorder). Schrödinger, a computational chemistry company, uses physics-based simulation combined with machine learning and has proprietary programs and partnerships.

Big pharma has jumped in. Companies including AstraZeneca, Sanofi, and Bayer have built AI-discovery capabilities or partnered with AI startups.

Where AI Actually Helps

It is worth being precise about what AI does well in the drug pipeline:

  • Target identification: Analysing genomic, proteomic, and literature data to find proteins associated with disease.
  • Molecule design: Generating novel chemical structures predicted to bind a target with desirable properties (potency, selectivity, solubility, metabolic stability).
  • Virtual screening: Rapidly evaluating millions of compounds against a target computationally rather than experimentally.
  • ADMET prediction: Predicting absorption, distribution, metabolism, excretion, and toxicity to filter out bad candidates early.
  • Clinical trial optimisation: Identifying suitable patients, predicting response, and optimising trial design.
  • Chemistry synthesis planning: Retrosynthesis tools such as those pioneered by AiZynthFinder and others help chemists plan how to make molecules.

The Limits

AI cannot solve the fundamental problem that biology is messy and clinical trials are expensive. Predicting a molecule’s effects from structure is not the same as knowing whether it will work in a human body. Toxicity and efficacy surprises remain common. And the most expensive part of drug development — Phase 3 trials with thousands of patients — is not something AI can compress much.

Data quality is also a bottleneck. AI models are only as good as the data they are trained on, and much biological data is noisy, inconsistent, or proprietary. Sharing data across companies and institutions remains difficult.

The Canadian Connection

Canada is well-positioned in this field. Toronto and Montreal are hubs of machine learning research with deep connections to the health sciences. The University of Toronto, the Vector Institute, Mila, and Amii are sources of AI talent, and Canadian hospitals generate rich clinical data. Several Canadian startups are applying AI to drug discovery, and international companies have established research presences in the country.

The Wet-Lab Bottleneck

AI accelerates the design phase, but the physical validation remains stubbornly slow. A predicted molecule must be synthesised in a lab, tested for activity, screened for toxicity, and then advanced through animal studies before human trials. Each step takes weeks or months and consumes reagents and skilled labour. The bottleneck has shifted from ideation to experimentation: AI can generate thousands of candidate molecules, but a lab can only test so many. The companies that pair AI with their own high-throughput wet labs — using automation and robotics to run experiments at scale — are betting that closing this loop is the key to faster discovery.

The Data Moat

AI models need data, and much of the most valuable biological data — proprietary assay results, clinical outcomes, patient records — is locked inside companies and hospitals. Public datasets such as the Protein Data Bank and ChEMBL are foundational but incomplete. This has created a competitive dynamic in which the firms with proprietary experimental data have an edge, and in which data-sharing consortia and industry partnerships have emerged to pool resources. Privacy regulations, commercial secrecy, and the sheer difficulty of standardising biological data all limit what can be shared. Data, more than algorithms, may be the limiting factor.

Clinical Trials and the Cost Wall

Even a perfectly designed molecule must pass through clinical trials, and it is here that the economics dominate. Phase 1 (safety) enrols dozens; Phase 2 (efficacy) hundreds; Phase 3 (confirmation) thousands, at a cost that can reach hundreds of millions of dollars per drug. AI can help design smarter trials — better patient selection, adaptive protocols, predictive biomarkers — but it cannot shorten the biological time it takes to observe outcomes or eliminate the cost of running large studies. The most valuable AI contributions may therefore be in choosing the right targets and molecules so that fewer expensive failures occur downstream.

Failures and the Realistic Picture

It is worth naming the failures. BenevolentAI’s lead asset for a neurological condition failed in trials, and the company restructured. Exscientia’s early clinical programmes did not all succeed and the company was ultimately acquired. Recursion has had pipeline setbacks. These are not indictments of the technology — most drug candidates fail regardless of how they were designed — but they are correctives to the narrative that AI guarantees success. AI can raise the probability of success and reduce the cost of the search; it cannot repeal the brutal base rates of drug development.

The Regulatory Question

Regulators are adapting. The U.S. FDA has issued guidance on the use of AI in drug development and has begun to accept AI-derived evidence in certain contexts. The European Medicines Agency is developing its own framework. The key questions — how to validate AI-generated predictions, how to document model provenance, and how to assess the reliability of AI-informed decisions — are being worked out. Clear regulatory pathways will accelerate adoption; ambiguity will slow it. Getting this right matters, because a model that helps design a drug must also be defensible to the regulators who decide whether it reaches patients.

The Reproducibility Problem

A final technical caveat applies to all AI-driven science, including drug discovery: reproducibility. Machine-learning models are sensitive to their training data, hyperparameters, random seeds, and implementation details. Two labs using the same architecture can produce different results, and published findings are not always reproducible. In drug discovery, where a model’s prediction guides experiments costing real money and time, non-reproducibility is a serious problem. Efforts to standardise reporting, publish models and data, and share evaluation protocols are improving the situation, but the field is still grappling with how to make AI-driven results as reliable as traditional experimental science. This, too, will take years to resolve.

Conclusion

AI is beginning to deliver on its drug-discovery promise — not by replacing the lengthy clinical trial process, but by making the early, expensive, high-attrition phase faster and more rational. The ISM001-055 results show a molecule that might not have existed otherwise, active in humans. That is a meaningful proof point. The next few years will reveal whether AI-discovered drugs translate into higher overall success rates — the metric that ultimately matters. The early evidence is encouraging, but the honest assessment is that we are at the beginning, not the end, of this transformation.

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button