Every token is getting cheaper, and every AI bill is getting bigger. When agents run around the clock in every department, something has to decide what they may spend, where data may go, and when they should sleep. This build is that layer, running live.
Today, the throttle on AI spend is a person at a terminal deciding to start a task. That’s about to change. Agents that watch usage, logs and feedback and improve the product on their own don’t stop at 5 p.m., and the same pattern is spreading to support, marketing, sales and finance.
Unit prices are falling fast. Total bills are rising faster, because volume and the cost of frontier work are growing even quicker.
Uber’s response was a per-engineer cap of $1,500 a month. A cap limits spending, but it can’t decide which work deserves the budget, which model should do it, or which data must never leave the building. That takes policy.
It’s where I think the next year goes. Once agents run on their own, the product question stops being “what can AI do” and becomes “what should it be allowed to do, and at what price.” That’s a governance question, and governance is the part of software I know best.
The ideaPut one policy gate in front of every agent call. It routes each task to the cheapest model that can do it, keeps sensitive data on models the company controls, wakes agents only when a signal is worth acting on, and keeps the year inside its budget.
Sources are listed at the end of the page.
A simulated week at a 1,000-person software company, one day per minute. Each dot is about thirty model calls, colored by department. Change the policy and every decision after that follows it. Prices are real; volumes are stated assumptions.
R0With routing off, everything goes to the most capable model.
R1Personal data is redacted by a local model first; confidential work runs only on a private endpoint or locally.
R2Low difficulty goes to the cheapest capable model.
R3Medium difficulty goes to a fast mid-tier model.
R4Hard work goes to Sonnet; long-horizon coding escalates to Opus.
R5Anything that changes something real waits for a named approver.
R6Work that can wait runs in the overnight batch at half price.
R7Over the day’s budget, background work is held until the next window.
Software is the biggest driver, but agents are showing up in every department. These are the always-on and everyday workloads in the simulation, and where the current policy sends each one.
| Department | Work | Per day | Tokens per call | Data | Goes to |
|---|
Volumes and tokens per call are assumptions for a 1,000-person software company with 250 engineers, set so engineering spend lands inside the $500 to $2,000 per engineer per month Uber reported.
A full year for the simulated company, calculated from the same rules and prices as the live view. It follows whatever policy is set above.
Today, engineering leaders set seat caps and finance finds out at the end of the month. With a governor, both manage the rules instead of the bill.
Per-seat caps, month-end surprises and hand-checking which tool saw which data go away. Deciding what the company is willing to spend, what agents may change without asking, and approving the changes that matter stay with people.
Routers already exist. What’s missing is the layer above them that a CFO and a CISO would both sign: a year’s budget, the data rules and the wake policy, enforced on every call and explained in a log.
That’s a gap a new company could fill, and the cloud platforms that host many labs’ models are well placed to try. None is fully neutral yet, and neutrality is the product.
# The policy the simulation runs, as a company would write it budget: { annual: 1_600_000, daily_share: true, over_budget: hold_background } wake: { product_optimizer: on_signal, security_triage: on_signal, content: on_signal } data: { personal: redact_locally_first, confidential: private_endpoint_or_local } route: { low: cheapest_open, medium: haiku, high: sonnet, long_horizon_code: opus } approve: { ships_code: true, sends_to_customer: true, moves_money: true } batch: { if_can_wait: overnight_batch }
This is a simulation. Prices are published list prices; call volumes, tokens per call and difficulty labels are my assumptions, calibrated to Uber’s reported per-engineer spend. Routing savings assume the cheaper model does the job as well. Published routers support that on benchmarks, but a real deployment has to measure cost per completed task, including retries. And the difficulty of a request is given here; in production a small classifier has to guess it, and sometimes guesses wrong.