Public companies put out about 16,900 earnings releases a year, and the headline numbers mix official figures with adjusted ones. This build checks every headline claim against what the company filed with the SEC, and shows the receipt.
An earnings release is written to be read, not audited. The headline might be GAAP net income, or it might be an adjusted number that leaves out charges the company would rather you skip. Analysts and data teams have to know which is which, for every figure, for every company, every quarter.
Adjusted numbers aren’t wrong; they’re a choice. The risk is not knowing which one you’re looking at. A model built on an adjusted headline, mistaken for GAAP, can overstate earnings by a lot: in this sample, one company’s adjusted EPS was $0.77 while it filed a loss.
It tests the most common promise in AI products, “grounded answers,” in a domain where grounded has a precise meaning: a number either matches the filing or it doesn’t. Every source is public, and every verdict can be checked by anyone with a browser.
The ideaPull each headline number out of the release, find the exact figure the company filed in XBRL, and let code, not a model, decide whether they match.
Sources: counts from EDGAR full-text search for 8-K Item 2.02; Calcbench and Suffolk University, July 2026.
Each tile is one headline claim. Pick one to see the sentence, the filed number and the verdict.
Every adjusted EPS claim in the sample, next to the diluted EPS the company filed, largest gap first.
Measured rates come from the board above. Public figures are linked; assumptions are yours to change.
Retyping and cross-checking headline figures that already match the filing goes away. Deciding whether an adjustment is reasonable, asking management about it on the call, and deciding what it means for the stock stay human.
The rules take one number per sentence, so a headline that packs GAAP and adjusted net income together keeps only the first. Segment figures, like one division’s revenue, are sometimes caught only by a size check and sent to a person rather than recognized. And the two unlabeled differences on the board are almost certainly definitions, not errors: net sales versus total revenue, and a bank’s “managed” revenue. The tool can’t tell a definition from a mistake. That’s why it flags instead of accusing.