FDA publishes a written notice for every medical device recall, then files more than a quarter of them under “Other” or “Under investigation.” A model read all 22,383 recall notices since 2003 in about five minutes for about a dollar. The pile FDA left unexplained turns out to be readable, and a trend everyone cites turns out to be a labeling artifact.
Device makers, hospital risk managers and investors all trend public recall data: which failures are rising, in which device types, at which firms. The notices themselves are free text. The only structure FDA adds is a root-cause label, and for a large share of recalls that label says nothing. So people either read notices one at a time or search them by keyword, and keyword search misses most of what it is looking for.
The cost isn’t only reading time. A question that takes weeks of manual coding usually doesn’t get asked, and one answered by keyword search gets a confident, incomplete answer.
I’ve spent my career in regulated software, where the difference between what a record says and how it was filed matters. This dataset is public, large, messy in a specific way, and has a built-in answer key for most of it, so the model’s reading can be checked instead of trusted.
The ideaAsk a small decision model five typed questions about every notice in one pass, keep FDA’s label beside its answers, and let code, not the model, decide what counts.
Sources: openFDA device recall export (59,273 product recalls), downloaded September 26, 2026; TypeSafe jev-1.13.0 answers, run September 27, 2026.
Each dot is one recall event, colored by what the model says failed. Switch the grouping to watch FDA’s unexplained pile sort itself. Tap a dot to read the notice.
Tap any dot to read that recall.
FDA’s software root-cause labels barely appear before 2007, so labeled software recalls seem to jump from near zero. Read from the notices, software is behind roughly one recall in six from the first year.
The threshold lives in code. Moving it recounts every recall without asking the model again.
Say a regulatory team wants to ask one new question of the whole archive, like “which recalls involve a battery?” Here is the cost by hand against the cost with the model and a person reviewing what it is unsure about.
In practice nobody codes 22,000 notices by hand. They search by keyword, which found of FDA-labeled software recalls here. The hand cost is what a complete answer would take.
Today, regulatory affairs and post-market surveillance analysts read notices one by one or settle for keyword counts. Here is how the hours for one complete question shift.
First-pass tagging disappears. Deciding whether a trend is real, what to do about it, and anything filed with FDA stays with people. So does the real root cause: FDA’s label comes from the firm’s investigation, and most notices don’t say why a failure happened. The model reads what was written, nothing more.
Search the notices for a device, a firm or a word. You’ll see what failed across the matching recalls, and the chart above highlights them.
FDA’s labels are the only answer key, and they aren’t ground truth: many of the model’s “false positives” I read by hand were software problems FDA filed under device design. I haven’t built an independent hand-labeled set, so accuracy on the unexplained pile rests on spot checks. Harm is the model’s reading of the text, not FDA’s recall class. Only 29% of notices say why a failure happened, so the model can’t recover most root causes, and FDA’s investigation still knows more than the text. Results are pinned to jev-1.13.0; a new version needs the checks rerun.