Resources
The off switch is not governance: what token pricing does to enterprise AI budgets
Banning a tool or capping spend controls only cost, the one AI-coding risk that self-reports. Code already merged, data sent in prompts, and the missing audit trail stay exactly where they were. Govern the work, not the tool.
News broke in May 2026: Microsoft had ordered thousands of engineers in its Experiences and Devices division to stop using Claude Code and move to the company's own Copilot CLI by June 30 (TheStreet, May 2026, and Forbes, June 2026). The tool had not failed. It had spread. Engineers building Windows, Office, and Teams kept picking a rival's coding agent over the product their employer sells, and every one of those choices was metered by the token. Microsoft called the move toolchain consolidation, and the strategic motive is easy to see. Harder to see is what the order actually controlled.
A line item with no line to the product
Uber ran the same experiment in the open. It put Claude Code in front of roughly 5,000 engineers. Typical bills ran $150 to $250 per engineer per month, and the heaviest users generated $500 to $2,000 (Forbes, May 2026). Four months into 2026, the year's AI budget was spent (Fortune, May 2026).
That alone is a procurement story. What makes it an economics story is what Uber's COO, Andrew Macdonald, said on the record: the link between token spend and product "is not there yet," and the trade "becomes harder to justify" (Fortune, May 2026). Uber did what the market advised. Pick a strong tool, give it to everyone. Adoption climbed. The line from spend to shipped product has not appeared.
The heavy user changed sides
For thirty years, developer tools were priced per seat, like a gym membership. Gyms charge flat and pray you never show up. Per-seat software charged flat and did not care if you came: usage cost the vendor nothing, so the heavy user was the vendor's best customer. Same fee, more value taken out, easier renewal. Rollout plans still assume that shape.
Token pricing moves the heavy user to the other side of the ledger. The better the tool works, the more engineers reach for it, and the more context it reloads on every attempt. The reward for successful adoption is a bigger invoice. Nor is this one vendor's experiment: GitHub moved Copilot to usage-based billing in June 2026 (GitHub Blog, June 2026), and whatever mix of seats and meters vendors land on, the direction is metered, because metered is where their own costs sit.
The tools are fine. The budgeting model that carried them into the building is not. Unmanaged adoption is now a cost that grows with success.
Cost is the only risk that reports itself
Cost has one property no other AI coding risk shares: it self-reports. It arrives monthly, itemized, in a currency the CFO already reads. Every other liability of ungoverned AI coding is silent.
Output quality has no meter. In a small controlled study by DryRun Security, a code-security vendor, 26 of 30 AI-agent pull requests (87%) introduced at least one vulnerability (DryRun Security, March 2026). Thirty pull requests with no human baseline make that an illustration, not an industry rate, and the illustration is enough: no invoice announces it. Veracode's 2025 tests, which ran code generation against a fixed task set across more than 100 models rather than live agent work, found a security flaw in 45% of tasks (Veracode, 2025). Data exposure has no meter either. Nothing bills you the month source code or customer records leave in a prompt. Auditability is the quietest of all: no vendor sends a monthly figure for whether this year's AI-written changes will stand up as change-management evidence under DORA, or meet the development-security measures NIS2 names.
These costs do not arrive monthly. They arrive at once, on the day something breaks. In March 2026, a configuration change went around change management at Amazon, and a roughly six-hour outage cost an estimated 6.3 million orders (Business Insider, March 2026, from internal documents whose interpretation Amazon disputes). The reporting tied the incident to AI-assisted code changes. Amazon rejects the link and calls it user error. Either way, the response is the telling part: a 90-day code-safety reset with mandatory senior sign-off on AI-assisted changes (Digital Trends, 2026). The company that took the hit did not reach for an off switch. It put governance on the work.
An off switch is a fuse
A fuse trips when current crosses a limit, instantly and without judgment. It also knows nothing about what the current was doing. Microsoft's order works the same way. It closes one vendor's line on the invoice, and that is all it does. It does not review the code the agents already merged, or recall the data that left in prompts, or produce the record an auditor will ask for. And it did not stop AI-written code from shipping at Microsoft. The same engineers now generate it under a different meter. Substitution is not governance either.
Agent code does not ship uncontrolled today. In a regulated shop it passes the same pipeline as everything else: mandatory review, static analysis, CI gates, change management. That stack judges the diff. It cannot say which model wrote the change, from what prompt, under which policy, or whether the agent behaved inside the rules on the way to the merge. Provenance is the gap, and no dashboard you already own displays it.
An off switch trips on cost because cost is the only signal wired to it. The rest of the circuit stays dark.
Govern the work, not the tool
One tempting conclusion from May 2026 is to pick a cheaper agent or set a harder cap. A cap is a fuse with a dial. Amazon's reset points somewhere more durable: put the policy on the work itself, above whichever tool touched it, and make every AI-assisted change carry its evidence.
That is the shape of the factory. It is deliberately slow, hours and sometimes days, because the audit trail is built during delivery instead of reconstructed for the auditor afterwards. You do not prompt it. It prompts you, interviewing stakeholders and reading your policies until it is confident enough to build. A run ends in the same standard set of nine artifacts, specification through DPIA, held to the published bar of 22 checks. The subscription is priced on the outcomes delivered rather than tokens burned, so the bill and the shipped software sit on the same line.
The factory does not replace the assistants your engineers use. Those make developers faster inside the pipeline you already run. The factory is a separate route for the work that has to arrive with the record attached. Complement, not competitor.
A spending cap controls one number. Governance accounts for all of them. Microsoft's off switch handled the visible problem. The invisible ones are still in production, at Microsoft and everywhere else. The invoice was the cheap surprise. It was the only one that announced itself.
Sources
- TheStreet, May 2026 (Microsoft orders engineers off Claude Code)
- Forbes, June 2026 (Microsoft ends Claude Code licenses, pushes Copilot CLI)
- The Next Web, 2026 (scope of Microsoft's staged Claude Code retreat)
- Forbes, May 2026 (Uber engineer count and per-engineer cost ranges)
- Fortune, May 2026 (Uber COO Andrew Macdonald on AI token spend)
- GitHub Blog, June 2026 (Copilot moves to usage-based billing)
- DryRun Security, March 2026, via Help Net Security (26 of 30 agent PRs, in Taiga receipts registry)
- Veracode, 2025 GenAI Code Security Report (in Taiga receipts registry)
- Business Insider, March 2026, via Digital Trends (Amazon outage estimate and code-safety reset, in Taiga receipts registry)
Frequently asked questions
Why did Microsoft tell its engineers to stop using Claude Code?+
Reporting in May and June 2026 revealed that Microsoft had ordered thousands of engineers in its Experiences and Devices division to move off Claude Code and onto its own Copilot CLI by June 30 (TheStreet, May 2026, and Forbes, June 2026). The official reason was toolchain consolidation, and token-metered usage at that scale had become a visible recurring cost. The switch changed which meter runs, and did not answer what the agents had already shipped.
How much does AI coding cost per engineer per month?+
At Uber, typical bills ran $150 to $250 per engineer per month, and the heaviest users generated $500 to $2,000, across roughly 5,000 engineers (Forbes, May 2026). Because pricing is metered by the token, cost scales with usage rather than headcount. Uber's 2026 AI budget was spent in four months (Fortune, May 2026).
Is token-based pricing replacing per-seat pricing for AI coding tools?+
The direction is visible. GitHub moved Copilot to usage-based billing in June 2026 (GitHub Blog, June 2026), and agentic coding tools are metered by the token. Under per-seat pricing heavy use was free at the margin, and under metered pricing adoption itself drives the bill.
How do you govern AI coding costs in an enterprise?+
Spending caps and tool bans control the invoice, and the invoice is the only AI coding risk that reports itself. Governance means policy enforced on the work before it ships, above whichever agent produced it. In the EU that has concrete names, such as change-management evidence under DORA, development-security measures under NIS2, and the logging the AI Act expects.
Does banning an AI coding tool reduce risk?+
It caps future spend and stops new exposure. It does nothing about code already merged, data already sent in prompts, or the audit trail that was never created. A ban is a fuse. It trips on cost, and the other liabilities stay where they are.
See how Taiga puts governance on the work
Bring the project this article made you think about.