Product Field Work All builds
BuildAccountable

The Handoff Line

Every day, about 210 people tell the U.S. government their car might be dangerous. Someone at NHTSA reads every one. This is an experiment in letting AI take the clear-cut complaints, and proving when it shouldn’t.

The problem

NHTSA screens every vehicle safety complaint it receives and forwards the ones with safety significance for investigation. The work is repetitive and noisy: owners describe symptoms, not causes, and often name the wrong part. When volume outruns attention, a real defect can hide in plain sight.

–complaints a year, counted from NHTSA’s public file
50–75%of complaints misidentify the affected part, per a 2015 federal audit
9,266complaints about GM vehicles reached NHTSA before it acted on the ignition switch defect

That defect was later tied to 124 deaths and 266 injuries, and a recall of 2.6 million cars. Missing a signal is the expensive failure. Reading everything slowly is the everyday one.

How I measure the problem
  1. Hours spent reading routine complaints. –
  2. Mistakes the AI makes when it acts alone, measured on complaints it never saw
  3. Share of complaints a person still reads, and whether those are the hard ones
Why I chose it

It is the cleanest version of a question every AI product eventually faces: when should software decide on its own, and who checks? The data is public, the stakes are real, and the answer can be measured instead of asserted.

The ideaLet a model route the complaints it’s sure about, define “sure” with a measured guarantee, and send everything else to a person.

Sources: DOT Inspector General testimony on ODI’s screening process; The Regulatory Review on the 2015 audit; AP on the GM compensation fund.

900 real complaints, sorted by how sure the model is

Each dot is a complaint the model never saw in training. Higher means more confident.

AI routes it, correctly AI routes it, wrongly Sent to a person
1%10%
AI handles alone–
Its actual mistake rate–
People review–

What the line is worth

Every number below moves with the slider above. Public figures are linked; assumptions are yours to change.

–saved a year
–people’s time freed
–complaints a year start in the wrong queue

    From public data

    Median pay for compliance officers, BLS, May 2025.

    Private industry, BLS ECEC, June 2026. Benefits make up the rest.

    About 600 tokens in and 120 out at Claude Haiku 4.5’s $1 and $5 per million.

    Assumptions you can change

    No public benchmark exists for either. These are conservative starting points, not findings.

    What it means for the people

    Today
    Reading and routing every complaint
    With the line
    The hard cases
    Fixing misroutes
    Freed for finding defect trends

    They read the hard ones, not all of them
    They own the lineThe mistake budget is a policy the team sets, not a vendor default. Tightening it sends more to people, and the page shows exactly what that costs.
    Their calls make it betterEvery decision on a handed-off complaint becomes a trusted label for the next time the line is set. The owner’s own label often isn’t one.
    Why the freed time matters

    The job was never reading complaints. It was spotting a defect trend before it hurt someone. A 2015 federal audit found that half to three quarters of complaints misidentify the affected part, and that NHTSA received 9,266 complaints about GM vehicles later recalled for the ignition switch before it acted. Volume buried the signal. The line hands the routine back to software so screeners have time to look for patterns.

    Source: The Regulatory Review, summarizing the DOT Inspector General’s 2015 audit. What disappears: routing clear-cut complaints like a squealing brake. What stays human: every uncertain call, and every decision to investigate.

    Try your own

    This runs entirely in your browser. What you type is never sent anywhere.

    How it works

    ComplaintFree text an owner filed with NHTSA.
    ClassifyA small linear model scores all ten components.
    The lineAbove it, the AI routes alone. Below it, a person decides.
    QueuesEach engineering team gets its complaints; reviewers get the rest.

    Evidence

    –of held-out complaints routed to the right component when the AI must always answer
    –error targets where the measured mistake rate stayed under what you asked for
    –complaints used to train, set the line, and test it, kept strictly apart
    Where it gets it wrong

    Choices

    A 1.5 MB linear model over an LLM classifierIt answers instantly, costs nothing to run, and every decision points to the words that drove it.
    A measured line over self-reported confidenceThe line is set on one held-out set with a 95% upper bound on error, then checked on another the model never saw.
    Your browser over a serverThe model ships as a file, so the demo is private by design and free to host.