Product Field Work All builds
BuildLegible

The Engine

Every team in a software company runs the same loop by hand: notice a signal, decide what it means, do the work, check it. This build takes a typical SaaS company apart department by department and shows that loop running on its own, around the clock, with people moving from doing the work to deciding what the engine is for.

Watch it in two minutes. The same idea as the page below, from the way work runs today to where people stay once the engine runs. Written as code, drawn frame by frame, and narrated with an ElevenLabs voice.

The problem

A signal shows up: a drop in a funnel, a spike in tickets, a failing build, a late invoice. Today it waits for the next standup, planning meeting or month-end close. Someone writes it up, someone prioritizes it, someone does the work, someone checks it. Most of the elapsed time is waiting between people, not work.

The pieces to run that loop continuously now exist. What’s missing is the design: which decisions go to which kind of model, who checks the work, and how a company hands over control without losing it.

12AI agents the average company runs in 2026, half of them fully on their own, per Salesforce’s benchmark report
70–500 msper typed decision from TypeSafe’s Jev, a decision model launched September 15 at $0.042 per million input tokens
90%of Uber’s roughly 65,000 weekly code changes are reviewed by its AI review system, uReview
How I measure the problem
  1. What each department takes in and produces, reduced to signals, decisions, work and checks
  2. Time from signal to shipped change, today versus each stage of handing over control
  3. Cost per decision, for a person, a frontier model and a decision model, at published prices
Why I chose it

In an earlier build, The Governor, I simulated the control layer that always-on agents will need: how much they may spend in a year, which model handles each task, when they wake up, and what data may never leave the company. That build was about the limits. This one is about the work itself: what the engine actually does all day, team by team. Product management is, at bottom, the job of turning signals into decisions into shipped work, so it’s the job I most want to see clearly before an engine takes it over.

The ideaTreat the company as one loop. Fast decision models make the thousands of small calls; large models do the few pieces of real work; separate agents check everything; and autonomy is earned one decision type at a time, with numbers.

What exists today

Decision modelsJev and CLM-8BJev returns typed decisions with calibrated probabilities instead of text. A week later, an open model, CLM-8B, matched it on several tasks at up to 9× lower latency.
Agents that work in the backgroundCoding agents in parallelCursor’s background agents plan, change, test and open a pull request unattended; Claude Code’s agent teams run a lead agent coordinating several others.
Agents for everyoneMeta’s MuseA personal agent that books, emails and shops across a person’s accounts, with about 730,000 downloads in its first five days.
Engine parts in productionReview, triage, routingUber’s uReview reviews most code changes; support, sales and finance tools now ship task agents that run alone.
AdoptionAgents inside the appsGartner expected 40% of enterprise applications to embed task-specific agents by the end of 2026, up from under 5% in 2025.
What’s missingThe loop across departments, and a way to hand it overParts run in isolation. Nothing yet connects signals in one team to work in another, or measures when a decision type has earned autonomy.

The same company, before and after the engine

Each colored dot is one piece of work from the team at the edge where it starts. Start with how work has moved for the last twenty years, then switch to the engine and watch the same work move through it.

    Signals anyone looks at–
    Decisions made a day–
    Pieces of work finished a day–
    Checks and tests run a day–
    Waiting right now0
    How the waiting line moves–
    Median time, signal to finished work–
    Hours a week the loop runs–
    • One piece of work, in its team’s color
    • A signal nobody noticed, fading away
    • Sparks: fast decisions by a decision model
    • Larger dot: real work being generated
    • Outlined: high-risk, always goes to a person
    Live: following one piece of work every couple of seconds

      Then and now, step by step

      StepFor the last twenty yearsIn the engine
      1 · SenseSomeone notices it in a weekly report, an email, a spreadsheet or a customer complaint. Most signals are never looked at.Watchers read every signal as it arrives, from every tool, day and night.
      2 · DecideIt waits for the next triage or planning meeting. Often someone has to ask an analyst or an engineer to pull the data first.A decision model answers in about a tenth of a second, with the data already attached and a confidence score.
      3 · ActIt goes into a sprint, a campaign or a queue, and gets done in business hours when someone is free.An agent starts immediately and works nights and weekends. Simple calls skip this step entirely.
      4 · VerifyManual QA, a review meeting, one A/B test at a time.A separate tester writes tests for every change; many experiments run in parallel; failures go straight back.
      PeopleDo every step, and pass work between them.Approve everything at first, then a sample, then only goals and exceptions.
      TimeDays to months per itemMinutes to hours per item

      Team by team

      The crew inside the loop

      SenseWatchersSubscribe to every signal source and turn raw events into a state a model can judge: what changed, how much, compared to what.
      DecideDecision modelsAnswer thousands of bounded questions a minute with a probability attached. Low confidence routes up to a large model or a person.
      ActBuildersLarge models do the open-ended work: code, copy, specs, forecasts. Few calls, each expensive, each checked.
      VerifyTestersA separate agent that didn’t see the builder’s reasoning writes the tests and tries to break the work.
      OverseeOverseerWatches the loop itself: guardrail metrics, drift, cost, and whether the engine is gaming its own goals.
      ReportReporterWrites the daily account for leadership: what changed, why, what it cost, and what it wants permission to do next.

      What a decision costs

      Person: two minutes of judgment at a $50 loaded hourly cost (an assumption). Frontier model: Claude Sonnet 5 at $2 in and $10 out per million tokens, 2,000 tokens in and 300 out. Decision model: Jev at $0.042 per million input tokens with output free, 2,000 tokens in. Bars are on a log scale; the gap is roughly 20,000× from person to decision model.

      That ratio is the whole economic argument for the engine. It’s also why cost discipline matters: once a decision costs almost nothing, the temptation is to make a million of them, and to send too many of them to the expensive models. Keeping that in check is its own layer, which I explore in The Governor: a policy gate in front of every agent call that holds the year to a budget, routes each task to the cheapest model that can do it, and keeps sensitive data on models the company controls.

      Handing it over

      A person approving every change will soon look absurd. An engine that has run a thousand checked experiments before lunch doesn’t need someone reading each one. But the handover shouldn’t happen by feel. It should happen one decision type at a time, when the numbers say so.

      Measured, not trustedError rate on an audited sample beats the human baseline for that decision type, for long enough to count.
      ReversibleA bad call can be undone in minutes: a rollback, a paused ad, a refund reversed.
      Small blast radiusOne bad decision can’t cost more than a set amount or reach more than a set number of customers.
      Not adversarialNobody is actively trying to fool it, or the decision is screened first.

      Where I’d keep a person, even when the engine is better

      An engine optimizes whatever it’s pointed at, very well, including the wrong thing. These are the places I’d keep a named person, even after the numbers say the engine is more accurate.

      • Choosing the goal. Conversion, retention, margin and trust pull against each other. Deciding the trade is a human job.
      • Irreversible or legally loaded calls. Money out, contracts, filings, hiring and firing. Someone has to be accountable by name.
      • Novel situations. An incident, a lawsuit, a market shock. The engine’s history doesn’t cover it.
      • Watching for gaming. Metrics that go up while customers get worse are the classic failure of optimizers. People have to look.

      What it means for the people

      The org chart becomes a control panel. The engine reports up to leadership directly; middle layers that existed to move information between teams mostly disappear.

      Department heads become loop ownersThey set the goals, guardrails and budget for their part of the engine and answer for its results.
      Individual contributors become verifiers and specialistsThe remaining roles check high-risk work, handle exceptions and bring the judgment the engine borrows.
      Leadership reads the engine’s reportDaily: what changed, why, what it cost, and what it wants permission to do next.

      Evidence

      ~20,000×cheaper per decision for a decision model than a person, at published prices
      ~5×more finished work than people can review in Stage 1, so the approval queue grows without limit; that is the pressure that forces the handover
      7 teamsbroken into signals, decisions, work, checks and outputs, each with a line where autonomy stops
      What I haven’t proven yet

      This is an analysis and a simulation, not a running engine. Decision-model speeds and prices are the vendors’ own claims, weeks old; independent tests so far show them strong on some tasks and weaker on others. Volumes and timings in the loop are illustrative. And the hardest part in practice, clean signals flowing between tools that were never designed to talk to each other, is assumed rather than solved here.

      Choices

      Decision models for the many, large models for the fewMost of a company’s work is bounded judgment. Typed answers with probabilities are faster, cheaper and easier to check than prose.
      The tester never sees the builder’s reasoningAgents that check their own work agree with themselves. Independence is what makes verification worth anything.
      Autonomy earned per decision typeNot “trust the engine” but “trust this kind of decision, at this error rate, this reversible.” That’s how the handover survives a bad week.

      Sources

      1. TypeSafe AI, “Introducing System One Models & Jev,” September 2026, and Flowtivity on its published pricing and 70–500 ms latency.
      2. explainx.ai on CLM-8B, September 24, 2026: an open System One model at up to 9× lower latency than Jev.
      3. Towards Data Science: an independent test of Jev, strong on some tasks and weaker on others.
      4. Uber Engineering: uReview analyzing about 90% of roughly 65,000 weekly changes.
      5. Belitsoft, citing Salesforce’s 2026 Connectivity Benchmark: 12 agents on average, half fully on their own.
      6. Gartner forecast, via Wikipedia: 40% of enterprise applications embedding task-specific agents by the end of 2026.
      7. David Daniel Research, April 2026: Cursor background agents and Claude Code agent teams.
      8. TechCrunch and AI Agent Store: Meta’s Muse, launched September 8, about 730,000 downloads in five days.
      9. Claude API pricing, September 2026: Sonnet 5 at $2/$10 per million tokens.