All articles

Sep 1, 2026

AI Weekly: The Invoice Is the Easy Half

AI spend got harder to forecast, and per-seat pricing is the first casualty

The Invoice Is the Easy Half

Sundeep Goel stood in a room of IT professionals in Atlanta a few months ago and told them token prices were going up. The room did not take it calmly. "Half the room violently agreed with me. Half of the room violently disagreed with me."

Goel runs Mavvrik, which sells AI cost governance, so he has a commercial interest in the anxiety. He also has the number. In research his company publishes with BenchmarkIT, the share of companies able to forecast AI spend within ten percent fell from 15% last year to 11% this year while aggregate spend climbed. Forecasting got worse as the stakes got higher. His explanation is structural, not a maturity problem. "AI is fundamentally non-deterministic." The same task costs a different amount every run, the model roster keeps changing, and today's prices are partly underwritten by venture subsidy.

The damage lands on the price list before the P&L. "the traditional software pricing models of per seat or percentage of spend models, no longer are viable because those models assume relatively flat cost of goods sold." Nearly half of companies with an AI-bearing product have already re-priced.

Keith Townsend priced the other side of it in public. He built a retrieval pipeline over thirteen years of conference recordings, roughly 237,000 segments, split between managed cloud primitives and an NVIDIA workstation on his desk, then ran the on-premise argument at its strongest. Peg the box at 100% for three years and it still loses: "Call it $2.90 per million, fully utilized. The same model on bedrock? About 40 cents per million." Seven to nine times slower, too. "Utilization is the wrong lever." And: "You can't idle your way past a through put ceiling."

That is the cheap finding. The expensive one is what the convenient path charges. "The managed data plane is cheap and fast, and the cost it hides is the chunking judgment it takes from you." The managed wrapper re-chunked his data on its own strategy and switched on parsing he never asked for, which would have broken filtered retrieval without ever failing. On the part that decides whether an answer is true: "Generating LLM findings is free and easy. Generating true ones is not." Most of his own promising discoveries did not survive the gate, and what worked was not the priciest model. "A cheap, strict judge, plus one frontier judge trusted where they agree, did the work."

Nick Heinzmann has the buyer's version. Zip's research, which he runs, finds "only 17% of organizations from respondents actually said they could demonstrate a clear ROI from AI." The separator is not the tool, it is intensity: "84% of those people use AI multiple times a day versus just 40% of the lager group," the ones he calls bystanders. Occasional use produces neither returns nor the pressure to redesign the work. "the dabbling is really the actual issue." What that group hires for says where the cost moved: "the number one thing people are looking at when they're taking on new people is the ability to evaluate AI output. It's slop detection, basically." Generation is solved. "The bottleneck actually becomes how do we verify what's useful and what's good and then actually apply it."

Goel's instruction to finance is to stop treating the bill as the whole question: "I would encourage finance to not view AI cost governance as just what's my token invoice from anthropic or open AI." Ben Horowitz has the failure mode in one sentence: "They can burn a lot of tokens and spend a lot of money and get nothing productive done." The invoice arrives itemized. Nothing on it says which line bought anything.

Sources: Interviews from Metrics that Measure Up (Aug 26, Sundeep Goel of Mavvrik), Art of Procurement (Aug 24, Nick Heinzmann of Zip), and a16z (Aug 28, Ben Horowitz), and an essay by Keith Townsend of Layer2C Labs published on The CTO Advisor (Aug 24).

Finance and procurement leaders report AI cost forecasting getting worse as spend rises, per-seat pricing breaking on variable cost of goods sold, and verification replacing generation as the real bottleneck.