Every stopped clinical trial on ClinicalTrials.gov carries a sponsor’s note saying why. A model read all 38,617 of them, then read the eligibility criteria of 23,253 trials to test one question: can you see a recruitment failure coming in the protocol itself? Mostly yes, about two years early, and the attempts that didn’t work taught as much as the one that did.
A trial that stops early wastes the money spent on it, and the time of every patient who enrolled. The registry records why each one stopped, but only as a free-text note that nobody aggregates. And the question sponsors care about most, whether a protocol will recruit, is decided when the eligibility criteria are written, long before anyone knows the answer.
Source: ClinicalTrials.gov API, pulled September 28, 2026.
By the time low accrual is visible in enrollment numbers, the protocol, the sites and the budget are locked. The cheapest moment to fix recruitment is before the first patient is screened, and that is exactly when nobody has data.
I build clinical trial software. Recruitment is the problem every sponsor names first, and protocol design is where it is decided. The registry makes the claim testable in public, with outcomes anyone can check.
The ideaRead every stop reason and every eligibility section with a small decision model, match each failed trial to finished trials like it, and keep only the findings that survive controls and a test on future trials.
Sources: ClinicalTrials.gov (278,757 ended interventional trials, 453,567 for competition counts); TypeSafe jev-1.13.0 answers, run September 28, 2026.
Each dot is one stopped trial, colored by the main reason its sponsor gave. Red and orange are the scientific reasons: the treatment wasn’t safe or didn’t work. Everything else is operational. Tap a dot to read the sponsor’s note.
Tap any dot to read that trial’s reason.
Share of stopped trials giving each main reason, by the year the trial started. Business decisions have doubled since 2015, and COVID-19 left a mark on everything started just before it.
Trials started after 2022 have had less time to stop, so recent years are thinner. Years with fewer than 200 stopped trials are left out.
Each attempt to predict low-accrual failure, scored on trials that started after the model’s training data (AUC: 0.5 is a coin flip, 1.0 is perfect). Tap an attempt to read what I expected and what happened.
Protocols with six or more lab-value cutoffs looked twice as likely to fail. Then I split out cancer trials, which are lab-heavy and fail more on their own.
Failure rises once a protocol passes five exclusions, then flattens. The sixth exclusion matters; the twentieth barely does.
Share of trials in the matched sample that terminated for low accrual. One in three by design.
Seven features of an eligibility section, each tested with controls for oncology, phase, sponsor type and start year. Switch them on to build a protocol and see how its odds of dying from low accrual compare with a protocol that has none.
The final model learned only from trials that started in 2008 to 2014. Here it scores 12,305 trials that started in 2015 to 2021, sorted into tenths from lowest to highest risk. Slide to choose how many new protocols a design team would review.
Paste eligibility criteria, or open a real one. Each criterion gets an estimate of how much of the patient population it screens out, and the red flags are counted. Examples use the study’s saved answers; your own text is read live by Claude on your account.
The same model applied to the 2,241 interventional trials in Sorting Sea, all recruiting now. Each tick is a trial, placed by how its risk compares with trials that finished. Search or tap one to see its flags.
Say a sponsor or CRO wants this analysis: every stop reason coded and every eligibility section rated. Here is the cost by hand against the cost with the model and a person checking what it’s unsure about.
In practice nobody does this by hand; feasibility teams read a handful of comparable protocols. The hand cost is what the full answer would take.
Feasibility today means a few experienced people comparing a new protocol with the handful of trials they remember. Here is how the work shifts.
Reading thousands of comparable trials disappears. Every design decision stays with people: a narrow criterion is sometimes exactly right for safety or science, and the model can only say what it tends to cost in recruitment. It reads the registry, not the full protocol, the sites or the budget.
The registry shows each trial’s latest eligibility criteria, not the version it launched with, and the version history isn’t open to automated access, so I couldn’t test original against amended criteria or planned against actual enrollment. Stop reasons are written by sponsors and can be vague on purpose. Odds come from a matched sample, one failure to two finishers, so they compare protocols rather than give a trial’s absolute chance. Everything here is association, not proof that a criterion caused a failure. About 0.70 AUC seems to be the ceiling for public text: the rest is site performance and screening execution, which only operational systems see. Results are pinned to jev-1.13.0.