Product Field Work All builds
BuildLegible

Recall Radar

FDA publishes a written notice for every medical device recall, then files more than a quarter of them under “Other” or “Under investigation.” A model read all 22,383 recall notices since 2003 in about five minutes for about a dollar. The pile FDA left unexplained turns out to be readable, and a trend everyone cites turns out to be a labeling artifact.

The problem

Device makers, hospital risk managers and investors all trend public recall data: which failures are rising, in which device types, at which firms. The notices themselves are free text. The only structure FDA adds is a root-cause label, and for a large share of recalls that label says nothing. So people either read notices one at a time or search them by keyword, and keyword search misses most of what it is looking for.

–product recalls in FDA’s public device recall database, each with a written reason
–of recall events filed as Other, Under investigation, Unknown or Pending
–new recall events a year on average, 2004 to 2025, each one a notice someone has to read

The cost isn’t only reading time. A question that takes weeks of manual coding usually doesn’t get asked, and one answered by keyword search gets a confident, incomplete answer.

How I measure the problem
  1. How much of FDA’s unexplained pile can be named from the notice text, counted recall by recall
  2. Whether the model finds software recalls FDA’s labels find, and ones they miss, checked against FDA’s own labels and a keyword baseline
  3. What one new question of the whole archive costs, by hand and with the model, with assumptions you can change
Why I chose it

I’ve spent my career in regulated software, where the difference between what a record says and how it was filed matters. This dataset is public, large, messy in a specific way, and has a built-in answer key for most of it, so the model’s reading can be checked instead of trusted.

The ideaAsk a small decision model five typed questions about every notice in one pass, keep FDA’s label beside its answers, and let code, not the model, decide what counts.

Sources: openFDA device recall export (59,273 product recalls), downloaded September 26, 2026; TypeSafe jev-1.13.0 answers, run September 27, 2026.

Every recall, sorted

Each dot is one recall event, colored by what the model says failed. Switch the grouping to watch FDA’s unexplained pile sort itself. Tap a dot to read the notice.

Tap any dot to read that recall.

Recalls where the model is less sure than this fade out. At 0.60, of recalls go to a person.

The model’s read of the notice, from no harm (0) to life-threatening (3). Not FDA’s recall class.

Software recalls didn’t start in 2007. The label did.

FDA’s software root-cause labels barely appear before 2007, so labeled software recalls seem to jump from near zero. Read from the notices, software is behind roughly one recall in six from the first year.

Model: software is the causeFDA: software root cause

The threshold lives in code. Moving it recounts every recall without asking the model again.

What one new question costs

Say a regulatory team wants to ask one new question of the whole archive, like “which recalls involve a battery?” Here is the cost by hand against the cost with the model and a person reviewing what it is unsure about.

–saved a year at the question rate you set
–analyst hours per question, by hand
–analyst hours per question, with the model

    In practice nobody codes 22,000 notices by hand. They search by keyword, which found of FDA-labeled software recalls here. The hand cost is what a complete answer would take.

    From public data

    Median pay, BLS, May 2024.

    Private industry, BLS ECEC.

    Assumptions you can change

    No public benchmark exists for these. They are starting points, not findings.

    What it means for the people

    Today, regulatory affairs and post-market surveillance analysts read notices one by one or settle for keyword counts. Here is how the hours for one complete question shift.

    Today
    Reading and coding every notice
    With the model
    Spot checks
    Unsure cases
    Freed for analysis

    Analysts ask more questionsWhen a question of the whole archive costs a dollar and a few hours of review, the questions nobody had time for get asked.
    Unsure cases arrive flaggedEvery answer carries a probability, so the one recall in ten the model can’t call goes to a person instead of into the count.
    Every count opens its noticeEach dot links to the notice it came from, with FDA’s label beside the model’s, so a reviewer can check any number in a minute.
    What disappears, and what doesn’t

    First-pass tagging disappears. Deciding whether a trend is real, what to do about it, and anything filed with FDA stays with people. So does the real root cause: FDA’s label comes from the firm’s investigation, and most notices don’t say why a failure happened. The model reads what was written, nothing more.

    Try it: look up a device

    Search the notices for a device, a firm or a word. You’ll see what failed across the matching recalls, and the chart above highlights them.

    How it works

    One public export59,273 product recalls from openFDA, reduced to 23,094 unique notice texts so nothing is read twice.
    Five typed questions, one requestWhat failed, where it started, whether software caused it, worst plausible harm, and whether the notice says why.
    Code decides what countsThresholds, grouping and every total run in the page, so changing a rule never calls the model again.
    Every dot opens its noticeThe notice text, FDA’s label and the model’s answers, side by side.

    Evidence

    –of FDA-labeled software recalls the model also flagged, against for keyword search, at similar precision
    –of FDA’s no-root-cause recalls where the model named what failed; it marked as “notice doesn’t say”
    ~$1.10for the whole archive: 23,094 requests, about 22 million input tokens, a few minutes of wall time, no failed calls
    What I haven’t proven yet

    FDA’s labels are the only answer key, and they aren’t ground truth: many of the model’s “false positives” I read by hand were software problems FDA filed under device design. I haven’t built an independent hand-labeled set, so accuracy on the unexplained pile rests on spot checks. Harm is the model’s reading of the text, not FDA’s recall class. Only 29% of notices say why a failure happened, so the model can’t recover most root causes, and FDA’s investigation still knows more than the text. Results are pinned to jev-1.13.0; a new version needs the checks rerun.

    Choices

    A reading model, not a classifier trained on FDA’s labelsTraining on the labels would learn their blind spots, including the missing pre-2007 software category and the 28% labeled nothing.
    A small decision model over a chat modelJev returns typed answers with probabilities and bills input tokens only, at $0.042 per million. Five questions ride in one request, and there is no free text to parse.
    FDA’s label beside the model’s, not replacedFDA’s root cause comes from an investigation the notice text doesn’t carry. Showing both keeps the model honest about what it can’t see.