Every team in a software company runs the same loop by hand: notice a signal, decide what it means, do the work, check it. This build takes a typical SaaS company apart department by department and shows that loop running on its own, around the clock, with people moving from doing the work to deciding what the engine is for.
Watch it in two minutes. The same idea as the page below, from the way work runs today to where people stay once the engine runs. Written as code, drawn frame by frame, and narrated with an ElevenLabs voice.
A signal shows up: a drop in a funnel, a spike in tickets, a failing build, a late invoice. Today it waits for the next standup, planning meeting or month-end close. Someone writes it up, someone prioritizes it, someone does the work, someone checks it. Most of the elapsed time is waiting between people, not work.
The pieces to run that loop continuously now exist. What’s missing is the design: which decisions go to which kind of model, who checks the work, and how a company hands over control without losing it.
In an earlier build, The Governor, I simulated the control layer that always-on agents will need: how much they may spend in a year, which model handles each task, when they wake up, and what data may never leave the company. That build was about the limits. This one is about the work itself: what the engine actually does all day, team by team. Product management is, at bottom, the job of turning signals into decisions into shipped work, so it’s the job I most want to see clearly before an engine takes it over.
The ideaTreat the company as one loop. Fast decision models make the thousands of small calls; large models do the few pieces of real work; separate agents check everything; and autonomy is earned one decision type at a time, with numbers.
Each colored dot is one piece of work from the team at the edge where it starts. Start with how work has moved for the last twenty years, then switch to the engine and watch the same work move through it.
| Step | For the last twenty years | In the engine |
|---|---|---|
| 1 · Sense | Someone notices it in a weekly report, an email, a spreadsheet or a customer complaint. Most signals are never looked at. | Watchers read every signal as it arrives, from every tool, day and night. |
| 2 · Decide | It waits for the next triage or planning meeting. Often someone has to ask an analyst or an engineer to pull the data first. | A decision model answers in about a tenth of a second, with the data already attached and a confidence score. |
| 3 · Act | It goes into a sprint, a campaign or a queue, and gets done in business hours when someone is free. | An agent starts immediately and works nights and weekends. Simple calls skip this step entirely. |
| 4 · Verify | Manual QA, a review meeting, one A/B test at a time. | A separate tester writes tests for every change; many experiments run in parallel; failures go straight back. |
| People | Do every step, and pass work between them. | Approve everything at first, then a sample, then only goals and exceptions. |
| Time | Days to months per item | Minutes to hours per item |
Person: two minutes of judgment at a $50 loaded hourly cost (an assumption). Frontier model: Claude Sonnet 5 at $2 in and $10 out per million tokens, 2,000 tokens in and 300 out. Decision model: Jev at $0.042 per million input tokens with output free, 2,000 tokens in. Bars are on a log scale; the gap is roughly 20,000× from person to decision model.
That ratio is the whole economic argument for the engine. It’s also why cost discipline matters: once a decision costs almost nothing, the temptation is to make a million of them, and to send too many of them to the expensive models. Keeping that in check is its own layer, which I explore in The Governor: a policy gate in front of every agent call that holds the year to a budget, routes each task to the cheapest model that can do it, and keeps sensitive data on models the company controls.
A person approving every change will soon look absurd. An engine that has run a thousand checked experiments before lunch doesn’t need someone reading each one. But the handover shouldn’t happen by feel. It should happen one decision type at a time, when the numbers say so.
An engine optimizes whatever it’s pointed at, very well, including the wrong thing. These are the places I’d keep a named person, even after the numbers say the engine is more accurate.
The org chart becomes a control panel. The engine reports up to leadership directly; middle layers that existed to move information between teams mostly disappear.
This is an analysis and a simulation, not a running engine. Decision-model speeds and prices are the vendors’ own claims, weeks old; independent tests so far show them strong on some tasks and weaker on others. Volumes and timings in the loop are illustrative. And the hardest part in practice, clean signals flowing between tools that were never designed to talk to each other, is assumed rather than solved here.