Independent field guide · current 4 August 2026

AI in Drug Discovery: What Is Settled, What Is Sold, What Is Still Unknown

A read of the public record by category, evidence tier, and economics — written to be falsifiable.

Observed verifiable public record · Sponsor company-reported, not independently confirmed · Inference my reasoning · Forecast probabilistic claim

1 · Joint thesis

ObservedHigh confidence — a platform transition has occurred. Structure prediction, generative chemistry, ML-guided property optimization and automated retrosynthesis are now default infrastructure across large pharma and biotech, not novelty. This is not seriously contested.

InferenceLow confidence — that this infrastructure produces clinically superior drugs. The transition is real at the level of tooling and cycle time. Evidence that it changes the probability a molecule helps a patient remains thin, and the strongest published numbers sit at the trial stage least predictive of that.

ObservedBounded finding. Across the records reviewed for this guide, no approved drug was found for which AI both selected the target and generated the molecule. This is a statement about the records reviewed — it is a bounded search result, not a proof of absence, and any counterexample with contemporaneous documentation would retire it.

2 · Timeline, 2012–2026

  1. 2012ObservedAlexNet and the Merck molecular-activity contest. Deep networks beat hand-curated QSAR descriptors; the field's attention shifts.
  2. 2015ObservedAtomNet: structure-based deep convolutional bioactivity prediction, the first widely cited application of the paradigm to binding.
  3. 2018ObservedNeural retrosynthesis with Monte Carlo tree search (Segler et al., Nature). Synthesis planning becomes tractable — the least glamorous and most durable win.
  4. 2019ObservedInferenceDDR1 inhibitor in 21 days (Insilico, Nature Biotechnology). Widely read as de-novo discovery; the record shows a known target with extensive prior chemistry, six compounds synthesized, and a speed claim rather than a novelty claim. The conflation of those two claims still distorts coverage.
  5. 2021 / 2024ObservedAlphaFold2, then AlphaFold3. Structure supply is largely solved; binding affinity, conformational dynamics and induced fit are not.
  6. Mar 2026ObservedAF2BIND reaches peer review: folding networks repurposed to infer ligand-binding sites, extending structure models toward the druggability question.
  7. Jul 2026SponsorRentosertib (TNIK, idiopathic pulmonary fibrosis) enters a 320-participant, 52-week Phase III — the closest asset to the strict definition, and the single most informative pending readout in the field.
  8. Dec 2025SponsorInferenceGB-0895, a generatively designed anti-TSLP antibody, advances into two Phase III programs on a known, de-risked target with no Phase II efficacy readout. Biology risk was partly de-risked by tezepelumab-class precedent; what is being tested is molecule engineering, not target discovery.
  9. 2026ObservedZasocitinib (TYK2), a physics-first design lineage, posts positive Phase III results with an NDA planned. Not approved as of this writing.

3 · Five-category taxonomy

Most public disagreement dissolves once claims are sorted into these five buckets. Use arrow keys to move between tabs.

Strict: AI selects the target and generates the molecule

The only category that would settle the headline claim. Membership is small — rentosertib is the canonical member. ObservedNo approval was identified by the cut-off. InferenceBecause target selection is the step with the highest attrition and the weakest ground truth, this is also the category where AI's claim is boldest and its evidence thinnest.

4 · Clinical ledger

Status as of 4 August 2026. Sponsor-reported items are tagged as such.
Asset / programCategoryStatusRead
Rentosertib (TNIK, IPF)StrictSponsorPhase III, 320 participants, 52 weeksThe decisive test. Phase IIa was encouraging and small.
GB-0895 (anti-TSLP)Protein engineeringSponsorTwo Phase III programsSkipped Phase II efficacy on a validated target — speed bought with borrowed biology.
Zasocitinib (TYK2)Physics-firstObservedPositive Phase III; NDA planned; not approvedNearest-term approval in the guide; weakest claim to being “AI-discovered.”
REC-3964 (C. difficile)Workflow / phenotypicSponsorMid-stageSurviving asset of a heavily pruned pipeline.
REC-994Phenotypic screeningObservedDiscontinued after longer follow-upEarly exploratory trends did not hold.
DSP-1181, EXS-21546Generative chemistryObservedDiscontinuedThe field's first “AI-designed” clinical entries; neither established patient benefit.
Isomorphic LabsStrict (stated intent)ObservedZero publicly disclosed clinical-stage candidatesThe best-capitalized pure-AI platform has not yet placed an asset in public clinical trials. Included explicitly because its absence is data.

5 · Evidence ladder and benchmark leakage

Ladder. T0 in-silico benchmark score · T1 retrospective enrichment on known actives · T2 prospective wet-lab hit · T3 IND and Phase I safety · T4 Phase II efficacy against control · T5 Phase III · T6 approval with confirmatory evidence. InferenceThe overwhelming majority of public “AI drug discovery” claims sit at T0–T3. Nearly all economic value attributed to the field is priced off tiers that historically predict T4 poorly.

Leakage. ObservedStandard docking and affinity benchmarks carry documented artifacts: decoy sets separable on physicochemical properties alone, scaffold overlap between train and test splits, and near-absent temporal splitting. InferenceConsequence: rank-order improvements on these benchmarks can be uninformative about prospective enrichment. The corrective is not a better leaderboard but a CASP-style protocol — timestamped predictions registered before wet-lab results exist, scored against a properly prepared physics baseline rather than a weak ML strawman.

6 · Economics

ObservedTrial-stage rates (Jayatunga et al.). Assets disclosed by AI-native companies: 21 of 24 Phase I successes (~88%) but 4 of 10 in Phase II (~40%). This is not a strict all-AI-origin cohort.

InferenceThe Phase I number is the most-quoted statistic in the field and the least load-bearing. Phase I chiefly tests tolerability, and it is exactly what careful molecule optimization on well-characterized targets should improve. The Phase II collapse toward baseline is consistent with the plain explanation: AI-native companies choose de-risked targets and design clean molecules, which fixes chemistry risk and leaves biology risk untouched. Sample sizes are small and survivorship is unaccounted for; the honest reading is that the question is still open, not that it has been answered favorably.

ObservedCompany economics. Recursion reported roughly $74.7m 2025 revenue, $475.3m in R&D expense, and a $648.1m operating loss. Schrödinger reported $199.5m in software-segment revenue and a $103.3m net loss. InferenceThe revenue that exists is substantially tooling and partnership revenue. Selling computational chemistry to pharma is a demonstrated business; owning a profitable AI-derived pipeline is not yet one. Picks-and-shovels and pipeline option value should be valued separately — they have different failure modes and should not share a discount rate.

7 · Forecast

ForecastState of the field at end-2028 — 58 / 22 / 20.

ForecastStrict-definition approval — AI selected the target and generated the molecule — by 31 December 2030: ~30%. The estimate is dominated by rentosertib-class timelines: even a clean Phase III leaves review and confirmatory work inside a narrow window, and the category currently has too few shots on goal for redundancy.

8 · What to do, by background

9 · Five falsifiers

Each would move me materially. Listed in descending order of impact.

  1. A strict-definition approval. An approved new molecular entity where the sponsor documents, with contemporaneous records, both AI target selection and AI molecule generation. Voids the bounded finding and upgrades the low-confidence claim.
  2. A prospective head-to-head. A pre-registered Phase II in which AI-derived assets beat matched conventional controls on efficacy, not merely on timeline or cost.
  3. Phase II rates that hold. AI-native-company Phase II success stabilizing above ~55% across n ≥ 30 assets. That would break the selection-effect explanation in §6.
  4. A leakage-audited prospective benchmark. Blind, temporally split evaluation showing generative models achieving ≥5× prospective enrichment over a well-prepared physics baseline.
  5. Pure-platform clinical translation. An Isomorphic-class platform disclosing multiple clinical candidates that clear Phase II — falsifying the current read that no pure-AI platform has yet translated to durable clinical output.

10 · Sources

Primary papers, registries, company records, and filings supporting the central facts. Sponsor figures should be checked against the current registry entry before use.

Methods and foundations
Programs and clinical records
Economics, regulation, and reliability