Product Field Work All builds
BuildAccountable

The Governor

Every token is getting cheaper, and every AI bill is getting bigger. When agents run around the clock in every department, something has to decide what they may spend, where data may go, and when they should sleep. This build is that layer, running live.

The problem

Today, the throttle on AI spend is a person at a terminal deciding to start a task. That’s about to change. Agents that watch usage, logs and feedback and improve the product on their own don’t stop at 5 p.m., and the same pattern is spreading to support, marketing, sales and finance.

Unit prices are falling fast. Total bills are rising faster, because volume and the cost of frontier work are growing even quicker.

4 monthsfor Uber to use up its entire 2026 AI budget, as adoption of coding agents outran its forecast
13× a yearhow fast the price of a given level of model capability has fallen, per Epoch AI
~18× a yearhow fast the cost of running frontier-level models has risen, per a 2026 study of benchmark prices

Uber’s response was a per-engineer cap of $1,500 a month. A cap limits spending, but it can’t decide which work deserves the budget, which model should do it, or which data must never leave the building. That takes policy.

How I measure the problem
  1. A year of AI spend for a realistic company, with and without a policy layer, at published list prices
  2. How much each lever saves: routing, waking on signals, batching, measured one at a time
  3. Calls carrying customer data that leave the company, which cost nothing extra and matter most
Why I chose it

It’s where I think the next year goes. Once agents run on their own, the product question stops being “what can AI do” and becomes “what should it be allowed to do, and at what price.” That’s a governance question, and governance is the part of software I know best.

The ideaPut one policy gate in front of every agent call. It routes each task to the cheapest model that can do it, keeps sensitive data on models the company controls, wakes agents only when a signal is worth acting on, and keeps the year inside its budget.

Sources are listed at the end of the page.

A company’s AI, running live

A simulated week at a 1,000-person software company, one day per minute. Each dot is about thirty model calls, colored by department. Change the policy and every decision after that follows it. Prices are real; volumes are stated assumptions.

When agents wake
Rules
Simulated timeMon 07:00
Spent today, against a daily share of the budget (–)$0
Calls on local models–
Sent for approval0
Held or batched0
Customer data sent outside0
    The rules, in order

    R0With routing off, everything goes to the most capable model.

    R1Personal data is redacted by a local model first; confidential work runs only on a private endpoint or locally.

    R2Low difficulty goes to the cheapest capable model.

    R3Medium difficulty goes to a fast mid-tier model.

    R4Hard work goes to Sonnet; long-horizon coding escalates to Opus.

    R5Anything that changes something real waits for a named approver.

    R6Work that can wait runs in the overnight batch at half price.

    R7Over the day’s budget, background work is held until the next window.

    Where the tokens go

    Software is the biggest driver, but agents are showing up in every department. These are the always-on and everyday workloads in the simulation, and where the current policy sends each one.

    DepartmentWorkPer dayTokens per callDataGoes to

    Volumes and tokens per call are assumptions for a 1,000-person software company with 250 engineers, set so engineering spend lands inside the $500 to $2,000 per engineer per month Uber reported.

    What the policy is worth

    A full year for the simulated company, calculated from the same rules and prices as the live view. It follows whatever policy is set above.

    –saved a year against no governor
    –a year with this policy, down from
    –per engineer per month, down from
    –calls with customer data kept inside the company each year

    What it means for the people

    Today, engineering leaders set seat caps and finance finds out at the end of the month. With a governor, both manage the rules instead of the bill.

    Engineers stop being the throttleNobody has to decide, task by task, whether a run is worth it. The policy decides, and the log shows why.
    Finance plans a year, not a monthThe annual budget becomes a daily allowance the engine respects, with background work waiting when a day runs hot.
    Security writes the data rules onceWhat may leave the company, and what must stay on a controlled model, is set in one place and enforced on every call.
    What disappears, and what doesn’t

    Per-seat caps, month-end surprises and hand-checking which tool saw which data go away. Deciding what the company is willing to spend, what agents may change without asking, and approving the changes that matter stay with people.

    Who should own this?

    Build it in-houseLike Uber’s own gatewayFull control, including PII redaction and a security review before any team gets access. But every company rebuilds the same plumbing, and prices and models change monthly.
    Let a model vendor routeA router from the people selling tokensConvenient, but the router’s owner earns more when the expensive model wins. Hard to call that a neutral judge.
    My positionThe company owns the policy; a neutral layer enforces itBudget, data rules and approval rules belong to the company. Enforcing them across every vendor is a product, and it should be neutral.

    Routers already exist. What’s missing is the layer above them that a CFO and a CISO would both sign: a year’s budget, the data rules and the wake policy, enforced on every call and explained in a log.

    That’s a gap a new company could fill, and the cloud platforms that host many labs’ models are well placed to try. None is fully neutral yet, and neutrality is the product.

    How it works

    Every call, taggedDepartment, task, difficulty, data sensitivity, and whether it changes anything.
    One policy gateOrdered rules pick a model, a timing, and whether a person must approve.
    People approve changesAnything that ships, sends or pays waits for a named approver.
    A log that explainsEach decision records the rules that made it and what it cost.
    # The policy the simulation runs, as a company would write it
    budget:  { annual: 1_600_000, daily_share: true, over_budget: hold_background }
    wake:    { product_optimizer: on_signal, security_triage: on_signal, content: on_signal }
    data:    { personal: redact_locally_first, confidential: private_endpoint_or_local }
    route:   { low: cheapest_open, medium: haiku, high: sonnet, long_horizon_code: opus }
    approve: { ships_code: true, sends_to_customer: true, moves_money: true }
    batch:   { if_can_wait: overnight_batch }

    Evidence

    –of the savings come from routing alone, consistent with published router results
    –a year saved by waking the always-on agents on signals instead of every hour
    –of calls run on a local model, mostly for data reasons rather than cost
    What I haven’t proven yet

    This is a simulation. Prices are published list prices; call volumes, tokens per call and difficulty labels are my assumptions, calibrated to Uber’s reported per-engineer spend. Routing savings assume the cheaper model does the job as well. Published routers support that on benchmarks, but a real deployment has to measure cost per completed task, including retries. And the difficulty of a request is given here; in production a small classifier has to guess it, and sometimes guesses wrong.

    Choices

    Rules before a learned routerOrdered, readable rules can be audited and argued with. A learned router can come later, inside the same gate.
    Local models for data, not just costMoving sensitive work local barely moves the bill. It changes where customer data goes, which is the bigger risk.
    A budget as a daily allowanceCaps stop spending after the damage. A daily share holds background work when a day runs hot, so the year lands on plan.

    Sources

    1. Fortune, May 2026: Uber’s COO on AI token spend, after the company used its full-year AI coding budget in four months.
    2. Outlook Business, citing Bloomberg, June 2026: Uber’s $1,500 monthly cap on AI coding tool spend.
    3. Uber Engineering, “GenAI Gateway,” 2024: one gateway for all LLM use, with PII redaction and a security review before access.
    4. Epoch AI, “The plunging price of thought,” 2026: price of a fixed capability falling about 13× a year.
    5. “The Price of Progress,” arXiv, 2026: running frontier-level models getting about 18× more expensive a year.
    6. RouteLLM, ICLR 2025: learned routers cutting cost substantially while keeping most of the strong model’s quality.
    7. Anthropic, Claude Opus 5.5 and Claude API pricing, Sept 2026: Opus 5.5 $4/$20, Sonnet 5 $2/$10, Haiku 4.5 $1/$5 per million tokens.
    8. gpt-oss-120b pricing, Aug 2026 and gpt-oss-20b pricing: open-weight rates used for the hosted and local lanes (local is priced at the hosted rate as a stand-in for hardware cost).