Steven Sun

July 15, 2026

Where the money actually goes in a clinical trial

The number everyone quotes is around two billion dollars per approved drug, and the usual gloss is that discovery is hard. Discovery is hard, but that isn’t where the two billion comes from. It comes from everything that fails on the way to the one thing that doesn’t.

The model below makes that concrete. It follows a cohort of 100 single-asset biotech companies from pre-clinical work through FDA review, drawn as two Sankey diagrams on a calendar timeline. Most of them fail, and it tracks exactly where and when: some run out of money, most fail the science, and only a handful reach approval. Press Run both and watch the cohort resolve.

100 companies per cohort · same cohort, same luck, two sets of assumptions

Two futures for the same hundred companies.

Both diagrams run the identical cohort with identical random draws — the same molecule gets the same luck in each — so any difference between them is caused by the switches, not by chance. Left to right is years.

Compare Scenario A Scenario BB vs A
Cost per approved drug$977M$529M-46%
Approvals per cohort2.14.2+95%
Capital consumed$2.10B$2.21B+5%
Viable drugs lost to cash
would have been approved with money
1.83.9+123%
Pipeline length13.7 yrs13.7 yrs0%
Put another way: of every drug that would have been approved with money behind it, 45% died of financing instead of science in Scenario A (48% in B), at the normal market. That share falls toward 20% in a good market and climbs past 60% in a bad one — when capital is scarce, a viable drug's biggest enemy is the market, not the biology.
Capital deployed, and what it is doingBOTH BARS ON ONE SCALE · MAX $2.19B
SCENARIO A$0M deployed
Still at work$0M
Written off$0M
SCENARIO B$0M deployed
Still at work$0M
Written off$0M
IN PROGRESSPRE-CLINICALINDPHASE IPHASE IIPHASE IIIREVIEWAPPROVED
WRITTEN OFFPRE-CLINICALINDPHASE IPHASE IIPHASE IIIREVIEWOUT OF MONEY

Both bars are cut by phase, on the same colours as the diagrams below: the live ramp for money still in play, the rust ramp for money lost at that phase. Segments run pre-clinical to review, left to right, so the written-off bar reddens toward the right as the losses move later and get more expensive.

Measure inYEAR 0.0 / 14.6 · All indications · normal market
SCENARIO APRE-CLINICAL2.5yIND0.7yPHASE I2.3yPHASE II3.6yPHASE III3.3yREVIEW1.3y100entered the pipeRan out of cash0 / 19Failed at pre-clinical0 / 36Failed at IND0 / 7Failed at Phase I0 / 18Failed at Phase II0 / 14 +1Failed at Phase III0 / 2 +2Approved0 / 4 3RUNS TO 2041202620332040202620332041CALENDAR YEARS FROM A 2026 START — 1 YEAR = 24PX THROUGHOUT →
SCENARIO BPRE-CLINICAL2.5yIND0.7yPHASE I2.3yPHASE II3.6yPHASE III3.3yREVIEW1.3y100entered the pipeRan out of cash0 / 19Failed at pre-clinical0 / 36Failed at IND0 / 7Failed at Phase I0 / 18Failed at Phase II0 / 13 1Approved0 / 7 +3RUNS TO 2041202620332040202620332041CALENDAR YEARS FROM A 2026 START — 1 YEAR = 24PX THROUGHOUT →

Both diagrams share the same pixels-per-year and the same vertical scale, so lengths and thicknesses are directly comparable — a scenario that finishes sooner is drawn shorter. The green and rust figures beside each lane are that lane's difference against the other scenario. Blue is still in play — pale inside a stage, solid once past it. Rust is failed, black is out of money, green is approved.

Shared setup
Therapeutic area
Financing market

Long-run average — specialists, selective. Sets the odds of closing each round — and so the size of the black lane and how many viable drugs die of money rather than science.

Financing gate

The full gate: a market-driven raise curve, a penalty on programmes that drag, and easier raises for programmes that have compressed their time and cost.

Comparisons worth running
Scenario A
Saves time & money0/5 ON
PRE-CLINICAL 2.51.5yPRE-CLINICAL cost $3.0M$1.5MIND 0.70.5y
PHASE I 2.32.0yPHASE II 3.63.2yPHASE III 3.32.9y
PHASE II 3.62.9yPHASE III 3.32.6yPHASE II cost $13.0M$12.0M
REVIEW 1.31.0y
PHASE I 2.31.3yPHASE I cost $4.0M$1.5M
Improves the odds0/2 ON
PHASE I pass 47%85%
PHASE II pass 28%41%PHASE III pass 55%80%
Scenario B
Saves time & money0/5 ON
PRE-CLINICAL 2.51.5yPRE-CLINICAL cost $3.0M$1.5MIND 0.70.5y
PHASE I 2.32.0yPHASE II 3.63.2yPHASE III 3.32.9y
PHASE II 3.62.9yPHASE III 3.32.6yPHASE II cost $13.0M$12.0M
REVIEW 1.31.0y
PHASE I 2.31.3yPHASE I cost $4.0M$1.5M
Improves the odds1/2 ON
PHASE I pass 47%85%
PHASE II pass 28%41%PHASE III pass 55%80%
Sensitivity — each lever on its ownAll indications · normal market · from an all-off baseline

Each switch flipped on by itself, measured two ways: its effect on cost per approved drug and on PTS — the probability of technical success, i.e. the odds a programme clears every scientific gate. The split is the whole point: time-and-money levers move cost and leave PTS untouched, while odds levers move both.

COST / APPROVED DRUG
PTS (LIKELIHOOD OF APPROVAL)
Genetically supported target selection
-46%
+110%
In-silico ADMET & toxicity prediction
-28%
+81%
Run Phase I in China
-20%
Generative chemistry & in-silico screening
-18%
AI patient matching & enrollment
-11%
AI site selection & activation
-8%
Automated submission drafting
-0%

Baseline at these settings: $977M per approved drug, 4.1% PTS. Bars are scaled within each column. The ordering by cost impact holds across markets, but the magnitudes move — odds levers pull further ahead of time levers as the market worsens, because a higher PTS also rescues more drugs from dying of cash.

What each switch does
Saves time & money
These shorten bars. None of them changes whether the drug works.

Generative chemistry & in-silico screeningModel-designed candidates and computational triage before the bench. Cheaper, faster candidates — who then face exactly the same clinical gauntlet.

AI site selection & activationRanking investigator sites on predicted enrollment instead of prior relationships. Vendor-reported, unaudited.

AI patient matching & enrollmentScreening records against protocol criteria to fill cohorts faster. The biggest cycle-time claim in the vendor literature, and the least audited.

Automated submission draftingGenerating CSRs and module text from trial data. Real savings, on the shortest and cheapest stage in the pipeline — watch how little it moves.

Run Phase I in ChinaFirst-in-human in China: direct costs about 60% lower and enrollment two to three times faster, on a large pool of treatment-naive patients. Phase II onward stays in the US, so nothing here turns on the single-country data question.

Improves the odds
These change where the spine narrows, and so how many reach the end.

In-silico ADMET & toxicity predictionPredicting absorption, metabolism and tox before dosing humans. Jayatunga et al. 2024 found AI-discovered molecules clear Phase I at 80–90% — real, and the cheapest phase to win.

Genetically supported target selectionChoosing targets with human genetic causal evidence, which ML is what makes tractable at genome scale. Nelson et al. 2015 estimated a doubling of clinical success; Minikel/Nelson 2024 refined it to 2.6x; King et al. replicated it and located the effect in Phases II and III. Only about 15% of programmes currently have that support — the best-evidenced lever here by some distance.

Why the areas differ

The area you pick changes every transition rate, duration and cost in the model. Here is what is behind those numbers.

All indicationsLOA 4.1% · 13.7 yrsSELECTED
Not a disease area so much as the industry's weighted average, and it is dominated by whatever fills the pipeline — which for two decades has been oncology. That drags the pooled likelihood of approval down toward 6–7% and hides an enormous spread: the best and worst areas differ by roughly fourfold. Useful as a benchmark, misleading as a plan.
OncologyLOA 3.2% · 13.2 yrs
The largest share of the industry pipeline and the lowest odds in it. The structural problem is that oncology's early signals are weak predictors: a single-arm Phase II reporting response rate can look convincing and still fail to produce a survival benefit against an active comparator in Phase III. So attrition concentrates late, after the expensive trial has been run — which is why the money leaves this pipeline at the far right. Biomarker-selected populations and accelerated approval have improved the odds in targeted subsets, and FDA's Project Optimus has pushed dose-finding earlier, but the pooled numbers still look like this.
HematologyLOA 14.9% · 13.0 yrs
The most successful major area, for reasons that compound. Targets are often unusually well characterised at the molecular level — BCR-ABL, BTK, CD19, BCMA — so mechanism and disease are tightly linked. Endpoints are objective and read out early: response, remission, minimal residual disease, rather than survival years later. Populations are small and well defined, which brings orphan designation, smaller trials and faster paths. It is the closest thing in the model to what the rest of drug development would look like if target biology were actually understood.
CardiovascularLOA 4.7% · 15.0 yrs
A study in mismatched phases. Phase II is comparatively cheap because surrogate endpoints — LDL cholesterol, blood pressure — move fast and read out in modest samples. Phase III then demands cardiovascular outcome trials with tens of thousands of patients followed for years, which is where the budget goes. The field emptied out through the 2000s after a run of high-profile outcome-trial failures made that bet look unaffordable, and refilled only when PCSK9 inhibitors and later the GLP-1 agents showed hard outcome benefit.
NeurologyLOA 3.7% · 15.6 yrs
The longest timelines and among the least forgiving biology. Efficacy usually rests on subjective rating scales with large placebo responses, so trials must run long with many patients to separate signal from noise. Animal models translate poorly and the blood–brain barrier constrains what can be dosed at all. Alzheimer's is the extreme case: decades of amyloid-targeting failures preceded the recent anti-amyloid approvals, whose effect sizes remain contested. The parameters here are the most punishing in the model and they are not an artefact.
Benchmarks & sources
Transitions Citeline/BIO phase-transition data 2014–2023: Phase I 47%, II 28%, III 55%, review 92%; likelihood of approval 6.7%, down from 10.4% in 2014. Area splits from BIO's 14-category cut and Hay et al.
Durations BIO: 2.3 / 3.6 / 3.3 years plus 1.3 to approval; 10.5 years average, 9.2–12.2 by area. These set the column widths.
Costs Sertkaya et al.: Phase I $1.4–6.6M, II $7.0–19.6M, III $11.5–52.9M. Out-of-pocket and uncapitalised — add a discount rate and the time switches bite considerably harder.
Burn $100–250k/month early stage; a step-up of 50%+ moving into Phase II; a six-month delay costs eight to nine months of runway.
China DIA/GlobalData on cost and enrollment; FDA/ODAC 2022 on single-country data.
AI Jayatunga et al., Drug Discovery Today, 2024. As of early 2026 no AI-discovered drug has been approved anywhere.
Genetics Nelson et al., Nature Genetics 2015 (roughly a doubling); King et al., PLOS Genetics 2019 (replication, effect located in Phases II and III); Minikel & Nelson, Nature 2024 (2.6x, rising with confidence in the causal gene). Modelled here as 1.45x on each of the Phase II and Phase III transitions, compounding to about 2.1x overall.
Stated assumptions

One asset per company. The unit here is a company whose fate is one molecule's fate — it lives or dies with its lead program. That is deliberately not the whole industry: in the US, platform companies outnumber single-asset ones by roughly three to one among venture-backed biotechs, because larger funds prefer bets where one failure doesn't end the company. But single-asset and single-lead-asset companies remain common, especially at the clinical-stage IPOs that dominate the funnel — and even nominal platforms are often, in practice, valued and funded on one lead program. The single-asset frame is chosen because a molecule's journey maps cleanly onto a company's; it overstates binary company death relative to a diversified platform, which reallocates rather than dies when a program fails.

The financing gate, with its actual numbers. A round is required to begin Phase I, II and III. The chance it closes is 0.62 + 0.36 × climate, where the market sets climate to 0.30 (bad), 0.60 (normal) or 0.85 (good) — so a raise succeeds about 73% of the time in a bad market and about 93% in a good one, per attempt, compounded across three gates. The three settings bracket the observed range, from the 2022–23 winter to a 2021-style peak. Phase III is scaled slightly harder than Phase I (×1.02 vs ×0.95) because a pivotal round is the largest. Two adjustments then apply: a program already past 55% of its benchmark clock loses a further 15% (×0.85), reflecting investor fatigue with a program that is dragging; and a program that has compressed its time and cost needs a smaller, shorter round, which scales its chance of failing to raise by 0.15 + 0.85 × need (need = this round's time-and-cost against benchmark, so exactly neutral at benchmark and easier below it). These coefficients are calibrated to land the pooled likelihood of approval and the share dying of financing in a plausible range — they are not drawn from a specific dataset, and this gate is the weakest structural assumption in the model.

Modelled vs data-anchored. The financing-gate switch replaces all of that with a curve tied to real numbers. Biotech follow-on runs about 50% per round in normal markets; in the 2022–23 winter SVB counted 356 Series A biotechs producing only 102 Series B (~29%), and EY put 55% of emerging biotechs at under two years of cash. The data-anchored curve is calibrated to that spread (~4% fail-to-raise in a good market, ~23% in a bad one) and drops the drag and round-size add-ons. Honest limits: observed graduation conflates money with science and M&A, so the viable-but-unfunded share can't be isolated cleanly; funding rounds don't map one-to-one onto clinical phases; the winter figure is one lender's book. The two modes land close, which is some reassurance — but data-anchored removes the "faster programmes raise more easily" effect, because that isn't visible in the data.

The rest. Costs are out-of-pocket and include company overhead and the capital burned by companies that died of financing rather than science. The China route charges no extra time or money for a US bridging study, only the Phase I speed and cost gains, so it reads a little favourably. The time switches use vendor-reported reductions and stack multiplicatively; nothing audits them. Failure-mode splits come from older, coarse literature. The calendar years start from the financing climate you pick, but that climate is then held fixed for the whole run — a 2021 cohort does not live through 2022's crash here, which would mean treating the start year as a vintage rather than a setting.

Each diagram animates one cohort of 100; the comparison table pools 60 cohorts per scenario (6,000 programmes each). Both scenarios draw from the same fixed random strips, so the table's differences are attributable to the switches — but the levels themselves still carry sampling error of roughly ±8%.

Failures are cheap early and ruinous late

The reason a surviving drug costs what it does is visible the moment you switch the model from measuring companies to measuring capital. Killing a program in pre-clinical costs almost nothing. Killing one in Phase III, after years of trials and hundreds of patients, is where the money actually goes — and that cost of failure, spread across the whole cohort, is what the survivor has to carry. The written-off bar reddens toward the right precisely because the losses that land late are the expensive ones.

Testing what actually changes the math

The model is also a tool for asking how AI and other changes might move those numbers. You can run two scenarios side by side against the identical cohort — same companies, same luck — and toggle interventions grouped into two kinds: those that save time and money (generative chemistry, AI-run trial operations, running Phase I in China) and those that improve the actual odds of success (in-silico toxicity prediction, genetically supported target selection). Because both scenarios share their random draws, any difference between them is caused by the switches rather than by chance.

The recurring lesson is that levers which merely go faster help far less than levers that change whether a drug works, because cost-per-approval is driven by the probability of success. The single best-evidenced intervention in the model — choosing targets with human genetic support — outperforms every speed-and-cost improvement combined. Every number is sourced from published benchmarks (Citeline/BIO, Tufts CSDD, DiMasi, Nelson/Minikel), with the assumptions and their weaknesses stated openly at the bottom of the model.