Independent field guide · current 4 August 2026
AI in Drug Discovery: What Is Settled, What Is Sold, What Is Still Unknown
A read of the public record by category, evidence tier, and economics — written to be falsifiable.
Observed verifiable public record · Sponsor company-reported, not independently confirmed · Inference my reasoning · Forecast probabilistic claim
1 · Joint thesis
ObservedHigh confidence — a platform transition has occurred. Structure prediction, generative chemistry, ML-guided property optimization and automated retrosynthesis are now default infrastructure across large pharma and biotech, not novelty. This is not seriously contested.
InferenceLow confidence — that this infrastructure produces clinically superior drugs. The transition is real at the level of tooling and cycle time. Evidence that it changes the probability a molecule helps a patient remains thin, and the strongest published numbers sit at the trial stage least predictive of that.
ObservedBounded finding. Across the records reviewed for this guide, no approved drug was found for which AI both selected the target and generated the molecule. This is a statement about the records reviewed — it is a bounded search result, not a proof of absence, and any counterexample with contemporaneous documentation would retire it.
2 · Timeline, 2012–2026
- 2012 — ObservedAlexNet and the Merck molecular-activity contest. Deep networks beat hand-curated QSAR descriptors; the field's attention shifts.
- 2015 — ObservedAtomNet: structure-based deep convolutional bioactivity prediction, the first widely cited application of the paradigm to binding.
- 2018 — ObservedNeural retrosynthesis with Monte Carlo tree search (Segler et al., Nature). Synthesis planning becomes tractable — the least glamorous and most durable win.
- 2019 — ObservedInferenceDDR1 inhibitor in 21 days (Insilico, Nature Biotechnology). Widely read as de-novo discovery; the record shows a known target with extensive prior chemistry, six compounds synthesized, and a speed claim rather than a novelty claim. The conflation of those two claims still distorts coverage.
- 2021 / 2024 — ObservedAlphaFold2, then AlphaFold3. Structure supply is largely solved; binding affinity, conformational dynamics and induced fit are not.
- Mar 2026 — ObservedAF2BIND reaches peer review: folding networks repurposed to infer ligand-binding sites, extending structure models toward the druggability question.
- Jul 2026 — SponsorRentosertib (TNIK, idiopathic pulmonary fibrosis) enters a 320-participant, 52-week Phase III — the closest asset to the strict definition, and the single most informative pending readout in the field.
- Dec 2025 — SponsorInferenceGB-0895, a generatively designed anti-TSLP antibody, advances into two Phase III programs on a known, de-risked target with no Phase II efficacy readout. Biology risk was partly de-risked by tezepelumab-class precedent; what is being tested is molecule engineering, not target discovery.
- 2026 — ObservedZasocitinib (TYK2), a physics-first design lineage, posts positive Phase III results with an NDA planned. Not approved as of this writing.
3 · Five-category taxonomy
Most public disagreement dissolves once claims are sorted into these five buckets. Use arrow keys to move between tabs.
Strict: AI selects the target and generates the molecule
The only category that would settle the headline claim. Membership is small — rentosertib is the canonical member. ObservedNo approval was identified by the cut-off. InferenceBecause target selection is the step with the highest attrition and the weakest ground truth, this is also the category where AI's claim is boldest and its evidence thinnest.
Known-target protein engineering
Generative design of antibodies, binders and enzymes against a target chosen by conventional biology. GB-0895 is representative. ObservedNo approval was identified by the cut-off. InferenceThis is the most likely category to produce the first genuine success, precisely because it removes target risk — which is also why such a success would not validate the strict thesis.
Physics-first with ML acceleration
Free-energy perturbation and molecular dynamics leading, machine learning accelerating sampling and triage. Zasocitinib sits here. ObservedNo approval was identified; one NDA was planned. InferenceThe strongest track record in the guide — and notably the least dependent on generative models. Deserves credit as a separate lineage rather than absorption into the “AI drug” label.
Repurposing
ML over existing approved molecules. Produced real pandemic-era label expansions, and also produced conspicuous false signals that consumed trial capacity. InferenceOutputs are new indications, not new molecular entities; counting them toward “AI-discovered drugs” is the most common category error in circulation.
Workflow AI
Assay QC, retrosynthetic route planning, literature and patent triage, trial-site selection, protocol and submission drafting. InferenceThe largest category by realized value and the least discussed. If the platform transition is real anywhere, it is here — and its value shows up as cost and cycle time, not as approval rate.
4 · Clinical ledger
| Asset / program | Category | Status | Read |
|---|---|---|---|
| Rentosertib (TNIK, IPF) | Strict | SponsorPhase III, 320 participants, 52 weeks | The decisive test. Phase IIa was encouraging and small. |
| GB-0895 (anti-TSLP) | Protein engineering | SponsorTwo Phase III programs | Skipped Phase II efficacy on a validated target — speed bought with borrowed biology. |
| Zasocitinib (TYK2) | Physics-first | ObservedPositive Phase III; NDA planned; not approved | Nearest-term approval in the guide; weakest claim to being “AI-discovered.” |
| REC-3964 (C. difficile) | Workflow / phenotypic | SponsorMid-stage | Surviving asset of a heavily pruned pipeline. |
| REC-994 | Phenotypic screening | ObservedDiscontinued after longer follow-up | Early exploratory trends did not hold. |
| DSP-1181, EXS-21546 | Generative chemistry | ObservedDiscontinued | The field's first “AI-designed” clinical entries; neither established patient benefit. |
| Isomorphic Labs | Strict (stated intent) | ObservedZero publicly disclosed clinical-stage candidates | The best-capitalized pure-AI platform has not yet placed an asset in public clinical trials. Included explicitly because its absence is data. |
5 · Evidence ladder and benchmark leakage
Ladder. T0 in-silico benchmark score · T1 retrospective enrichment on known actives · T2 prospective wet-lab hit · T3 IND and Phase I safety · T4 Phase II efficacy against control · T5 Phase III · T6 approval with confirmatory evidence. InferenceThe overwhelming majority of public “AI drug discovery” claims sit at T0–T3. Nearly all economic value attributed to the field is priced off tiers that historically predict T4 poorly.
Leakage. ObservedStandard docking and affinity benchmarks carry documented artifacts: decoy sets separable on physicochemical properties alone, scaffold overlap between train and test splits, and near-absent temporal splitting. InferenceConsequence: rank-order improvements on these benchmarks can be uninformative about prospective enrichment. The corrective is not a better leaderboard but a CASP-style protocol — timestamped predictions registered before wet-lab results exist, scored against a properly prepared physics baseline rather than a weak ML strawman.
6 · Economics
ObservedTrial-stage rates (Jayatunga et al.). Assets disclosed by AI-native companies: 21 of 24 Phase I successes (~88%) but 4 of 10 in Phase II (~40%). This is not a strict all-AI-origin cohort.
InferenceThe Phase I number is the most-quoted statistic in the field and the least load-bearing. Phase I chiefly tests tolerability, and it is exactly what careful molecule optimization on well-characterized targets should improve. The Phase II collapse toward baseline is consistent with the plain explanation: AI-native companies choose de-risked targets and design clean molecules, which fixes chemistry risk and leaves biology risk untouched. Sample sizes are small and survivorship is unaccounted for; the honest reading is that the question is still open, not that it has been answered favorably.
ObservedCompany economics. Recursion reported roughly $74.7m 2025 revenue, $475.3m in R&D expense, and a $648.1m operating loss. Schrödinger reported $199.5m in software-segment revenue and a $103.3m net loss. InferenceThe revenue that exists is substantially tooling and partnership revenue. Selling computational chemistry to pharma is a demonstrated business; owning a profitable AI-derived pipeline is not yet one. Picks-and-shovels and pipeline option value should be valued separately — they have different failure modes and should not share a discount rate.
7 · Forecast
ForecastState of the field at end-2028 — 58 / 22 / 20.
- 58% Infrastructure consolidation. AI is universal and unremarkable; no demonstrated approval-rate advantage; value accrues to tooling, cycle time and cost, and the term “AI drug” quietly stops being used as a differentiator.
- 22% Clinical validation. At least one strict-definition asset posts positive confirmatory Phase III data with a defensible superiority narrative, and capital re-rates the category.
- 20% Retrenchment. High-profile Phase II failures trigger capital withdrawal; surviving platforms are absorbed into pharma as internal functions.
ForecastStrict-definition approval — AI selected the target and generated the molecule — by 31 December 2030: ~30%. The estimate is dominated by rentosertib-class timelines: even a clean Phase III leaves review and confirmatory work inside a narrow window, and the category currently has too few shots on goal for redundancy.
8 · What to do, by background
9 · Five falsifiers
Each would move me materially. Listed in descending order of impact.
- A strict-definition approval. An approved new molecular entity where the sponsor documents, with contemporaneous records, both AI target selection and AI molecule generation. Voids the bounded finding and upgrades the low-confidence claim.
- A prospective head-to-head. A pre-registered Phase II in which AI-derived assets beat matched conventional controls on efficacy, not merely on timeline or cost.
- Phase II rates that hold. AI-native-company Phase II success stabilizing above ~55% across n ≥ 30 assets. That would break the selection-effect explanation in §6.
- A leakage-audited prospective benchmark. Blind, temporally split evaluation showing generative models achieving ≥5× prospective enrichment over a well-prepared physics baseline.
- Pure-platform clinical translation. An Isomorphic-class platform disclosing multiple clinical candidates that clear Phase II — falsifying the current read that no pure-AI platform has yet translated to durable clinical output.
10 · Sources
Primary papers, registries, company records, and filings supporting the central facts. Sponsor figures should be checked against the current registry entry before use.