Product Field Work All builds
BuildLegible

The protocol predicted it

Every stopped clinical trial on ClinicalTrials.gov carries a sponsor’s note saying why. A model read all 38,617 of them, then read the eligibility criteria of 23,253 trials to test one question: can you see a recruitment failure coming in the protocol itself? Mostly yes, about two years early, and the attempts that didn’t work taught as much as the one that did.

The problem

A trial that stops early wastes the money spent on it, and the time of every patient who enrolled. The registry records why each one stopped, but only as a free-text note that nobody aggregates. And the question sponsors care about most, whether a protocol will recruit, is decided when the eligibility criteria are written, long before anyone knows the answer.

–interventional trials started since 2005 that have ended
–of them stopped early: terminated, withdrawn or suspended
2.2 yrsmedian time a trial ran before stopping for low accrual

Source: ClinicalTrials.gov API, pulled September 28, 2026.

By the time low accrual is visible in enrollment numbers, the protocol, the sites and the budget are locked. The cheapest moment to fix recruitment is before the first patient is screened, and that is exactly when nobody has data.

How I measure the problem
  1. Why trials really stop, trial by trial, from every sponsor’s written reason
  2. Whether the eligibility criteria predict low-accrual deaths, failed trials against matched ones that finished
  3. Whether the signal holds on trials the model never saw, trained on 2008 to 2014 starts, tested on 2015 to 2021
Why I chose it

I build clinical trial software. Recruitment is the problem every sponsor names first, and protocol design is where it is decided. The registry makes the claim testable in public, with outcomes anyone can check.

The ideaRead every stop reason and every eligibility section with a small decision model, match each failed trial to finished trials like it, and keep only the findings that survive controls and a test on future trials.

Sources: ClinicalTrials.gov (278,757 ended interventional trials, 453,567 for competition counts); TypeSafe jev-1.13.0 answers, run September 28, 2026.

38,617 stopped trials, sorted by why

Each dot is one stopped trial, colored by the main reason its sponsor gave. Red and orange are the scientific reasons: the treatment wasn’t safe or didn’t work. Everything else is operational. Tap a dot to read the sponsor’s note.

Tap any dot to read that trial’s reason.

The reasons are changing

Share of stopped trials giving each main reason, by the year the trial started. Business decisions have doubled since 2015, and COVID-19 left a mark on everything started just before it.

Trials started after 2022 have had less time to stop, so recent years are thinner. Years with fewer than 200 stopped trials are left out.

What I tried, in order

Each attempt to predict low-accrual failure, scored on trials that started after the model’s training data (AUC: 0.5 is a coin flip, 1.0 is perfect). Tap an attempt to read what I expected and what happened.

0.50 coin flip0.75

Claim that didn’t survive: lab cutoffs double the risk

Protocols with six or more lab-value cutoffs looked twice as likely to fail. Then I split out cancer trials, which are lab-heavy and fail more on their own.

Claim that half survived: more exclusions, more failure

Failure rises once a protocol passes five exclusions, then flattens. The sixth exclusion matters; the twentieth barely does.

Share of trials in the matched sample that terminated for low accrual. One in three by design.

Red flags stack up

Seven features of an eligibility section, each tested with controls for oncology, phase, sponsor type and start year. Switch them on to build a protocol and see how its odds of dying from low accrual compare with a protocol that has none.

1.0×

Two years early, on trials it never saw

The final model learned only from trials that started in 2008 to 2014. Here it scores 12,305 trials that started in 2015 to 2021, sorted into tenths from lowest to highest risk. Slide to choose how many new protocols a design team would review.

Lowest risk tenthHighest risk tenth
–
–

Check a protocol

Paste eligibility criteria, or open a real one. Each criterion gets an estimate of how much of the patient population it screens out, and the red flags are counted. Examples use the study’s saved answers; your own text is read live by Claude on your account.

Sorting Sea’s live trials, scored

The same model applied to the 2,241 interventional trials in Sorting Sea, all recruiting now. Each tick is a trial, placed by how its risk compares with trials that finished. Search or tap one to see its flags.

Safer than finished trialsRiskier than 90% of them

Pick a trial

Its risk percentile, red flags and estimated eligible share appear here.

What a complete answer costs

Say a sponsor or CRO wants this analysis: every stop reason coded and every eligibility section rated. Here is the cost by hand against the cost with the model and a person checking what it’s unsure about.

–saved per full pass
–analyst hours by hand
–analyst hours with the model

    In practice nobody does this by hand; feasibility teams read a handful of comparable protocols. The hand cost is what the full answer would take.

    From public data

    Median pay, BLS, May 2024.

    Private industry, BLS ECEC.

    Measured: 38,617 stop reasons and 23,253 eligibility sections; about $6 of model time for everything on this page.

    Assumptions you can change

    No public benchmark exists for these; they are starting points, not findings.

    What it means for the people

    Feasibility today means a few experienced people comparing a new protocol with the handful of trials they remember. Here is how the work shifts.

    Today
    Reading and rating by hand
    With the model
    Reviewing unsure answers
    Freed for protocol design

    Protocol authorsSee which criteria screen out the most patients while the protocol is still a draft, and argue for each one with evidence instead of habit.
    Feasibility and CRO bid teamsFlag risky protocols before quoting timelines, and price recruitment risk from thousands of comparable trials rather than a few memories.
    Sites and investigatorsDecide which studies to take on knowing how often protocols like this one fill.
    PatientsFewer trials that enroll a handful of people and then stop, leaving their participation without an answer.
    What disappears, and what doesn’t

    Reading thousands of comparable trials disappears. Every design decision stays with people: a narrow criterion is sometimes exactly right for safety or science, and the model can only say what it tends to cost in recruitment. It reads the registry, not the full protocol, the sites or the budget.

    How it works

    Two public pulls278,757 ended interventional trials for the analysis, 453,567 of every status for competition counts.
    Three reading passesWhy each trial stopped; seven judgments about each eligibility section; a screened-out estimate for every single criterion.
    Matched, then tested forwardEach failed trial is paired with two finished trials of the same phase, start year and sponsor type, and the final model is judged on later trials only.
    Every number opens a trialTap a dot, an example or a live trial to see the text behind it.

    Evidence

    –accuracy (AUC) on 12,305 trials that started after the training data, against 0.64 for registry fields alone
    0.77correlation between this project’s enrollment-difficulty reading and the one Sorting Sea made independently, on the same 2,241 trials
    ~$6of model time for all three reading passes, about 80,000 requests, with no failed calls
    What I haven’t proven yet

    The registry shows each trial’s latest eligibility criteria, not the version it launched with, and the version history isn’t open to automated access, so I couldn’t test original against amended criteria or planned against actual enrollment. Stop reasons are written by sponsors and can be vague on purpose. Odds come from a matched sample, one failure to two finishers, so they compare protocols rather than give a trial’s absolute chance. Everything here is association, not proof that a criterion caused a failure. About 0.70 AUC seems to be the ceiling for public text: the rest is site performance and screening execution, which only operational systems see. Results are pinned to jev-1.13.0.

    Choices

    Test on the future, not a random splitA random split lets the model peek at trials from the same years. Training on earlier starts and testing on later ones is what “predicts it years ahead” actually requires.
    Report the findings that diedThe lab-cutoff doubling and the funnel’s accuracy gain didn’t survive testing. Showing them is how a reader knows the ones that did were tested the same way.
    Model for judgment, code for countingExclusions, lab cutoffs and washouts are counted in code. Whether a window is acute or a biomarker is rare needs reading, so the model does that, with its uncertainty kept.