MarketCatalyst
← Back to Research
Research Desk

The AI Bill Nobody Can Predict: Cheap Models Cost More Nearly 1 in 3 Times

Only 11% of 396 organizations forecast their AI spending within 10%, and a test of 6,800+ tasks found cheaper models cost more than premium ones 32% of the time, as token use swings by model, task and even repeat runs of the same prompt.

5 min read

Published Oct 4, 2026 · 11:31 PM ET

Companies are spending heavily on AI, but few can say in advance what the bill will be. A new study of 396 organizations found that only 11% forecast their AI spending to within 10% of the actual figure, and a separate test across more than 6,800 tasks found that lower-priced models ended up costing more than premium alternatives in 32% of cases. The cause is simple: AI is billed by the token, and the number of tokens a model burns on a job can swing widely by model, by task and even between repeated runs of the same prompt.

11%
of 396 organizations forecast AI spend within 10%
32%
of 6,800+ tasks where the cheaper model cost more
$14 vs $1
cost of one task on Gemini 3 Flash vs. Gemini 3.1 Pro

The forecasting gap

The survey results show how widely budgets missed the mark:

How far off the forecast wasShare of organizations
Within 10%11%
Missed by 11% to 25%39%
Missed by 26% to 50%35%
Missed by more than 50%13%

In other words, nearly nine in ten organizations were off by more than 10%, and about one in eight were off by more than half. The finding is in line with other surveys of enterprise AI spending. A separate benchmark of 372 companies, cited by the cost-management firm PointFive, put the share forecasting within 10% at just 15%.

When "cheap" costs more

A model's price per million tokens says little about what a task will cost, because the cost depends on how many tokens the model uses to finish it. Models that "think" at length, or take many back-and-forth steps with tools, can rack up a bill that dwarfs their low sticker price. The extreme example in the research: Google's cheaper Gemini 3 Flash spent about $14 and still failed after nearly 1,000 steps on one task, while Gemini 3.1 Pro completed the same task in 85 steps for about $1.

Independent testing points the same way. Artificial Analysis found that Google's Gemini 3.5 Flash cost about 5.5 times as much to run in benchmark testing as Gemini 3 Flash, and nearly twice as much as the Pro-tier Gemini 3.1, even though its per-token price is lower than Pro's. The reason was volume: it averaged 49 interaction turns per task, against 23 for Gemini 3.1 Pro, and the extra turns multiplied the input tokens billed.

Same prompt, different bill

The study also ran identical prompts five times each on models from Anthropic, Google and OpenAI. Costs differed widely from run to run, and failed runs still consumed paid tokens. Academic work on AI coding agents has found the same pattern, with runs of the same task differing by as much as 30 times in total tokens. For a finance team, that means even a well-tested workload does not have a single, stable cost per task.

What it means for companies

Because AI providers have shifted to pricing that scales with activity, costs increasingly look like a utility bill rather than a fixed software license. Three practical points follow:

  • Measure cost per completed task, not cost per token. A model that fails and retries can be far more expensive than its price list implies.
  • Cap steps and spending per job. Without limits, a stuck agent can burn through budget while producing nothing.
  • Match the model to the task. Accenture analyzed 9,368 occupational tasks and found fewer than 10% needed a frontier model, yet many employees default to the most expensive one because they cannot see the cost difference.

What it means for investors

Rising token use is the engine behind demand for AI computing. Goldman Sachs has forecast token consumption growing roughly 24-fold by 2030, according to PointFive. For cloud and model providers such as Alphabet, that variable, usage-based revenue is a strength when adoption is climbing. The same unpredictability, though, can push customers toward tighter controls, routing work to cheaper models and slower rollouts, which would temper the growth.

It also complicates the idea that falling per-token prices make AI steadily cheaper. Google's Gemini 3.5 Flash was priced at $1.50 per million input tokens and $9 per million output tokens, three times the $0.50 and $3 charged for Gemini 3 Flash. Things to watch: how cloud providers talk about customer AI consumption on their earnings calls, any further price changes at the "cheap" tier, and whether companies start to report AI cost controls as a priority.

The bottom line. The price list is no longer a reliable guide to the AI bill. With only about one in ten organizations forecasting spend accurately and cheaper models costing more nearly a third of the time, the winners in AI are likely to include not just the best models but the tools and habits that help customers control what they spend.

Readers also read

MarketCatalyst LLC is not a registered investment advisor and does not manage client assets. Content on this platform is provided for informational and educational purposes only. It is not investment advice, and MarketCatalyst is not a stock-picking or trade-alert service. Trading stocks and options involves risk, including the possible loss of principal. Consider your own goals, time horizon, and risk tolerance, and consult a qualified financial advisor before making any investment decision.