Spending on AI Is Becoming Almost Impossible for Businesses to Budget

A recent study found that only 11% of nearly 400 businesses surveyed were able to accurately forecast AI spending.


First, there was #tokenmaxxing, whereby American businesses encouraged their workers to use as much AI as possible. Then came the bill.

A recent study found that only 11% of nearly 400 businesses surveyed were able to accurately forecast AI spending.

Companies started to realize they need to be more careful about counting their tokens, the units that measure AI use. And that’s hard to do: A recent study found that only 11% of nearly 400 businesses surveyed were able to accurately forecast AI spending.

Unlike traditional software, AI behaves more like a human worker: It takes action, makes decisions, sometimes even makes mistakes—all on the clock.

While more-advanced models tend to cost more per token, they can sometimes perform tasks more efficiently, leading to lower overall costs. Likewise, asking a “cheap” model to do something it isn’t suited to handle could cause a token run-up.

Researchers from Stanford University, Carnegie Mellon University, and the University of California, Berkeley, as well as Microsoft Research put this to the test earlier this year, running models through more than 6,800 tasks spanning math, programming, science and other areas. In 32% of cases, lower-priced models actually cost more than higher-priced models.

“The practical takeaway is clear,” said Lingjiao Chen, one of the researchers. “Price alone should not be used to infer which model is actually cheaper.”

When cheap gets expensive

Here’s an example of the same prompt presented to two Google models, the Gemini 3.1 Pro—generally used for tasks that require reasoning and judgment—and the less expensive, lighter, speedier Gemini 3 Flash.

The pricier Pro model finished in 85 steps, whereas the less expensive Flash model went through nearly 1,000 steps…then failed.

The less expensive model failed after running up $14 worth of token use, while the pricier one succeeded for just $1. While this result could have been an anomaly, it is something that can happen in regular model use.

Same model, different results

Even the same model might use a different number of tokens each time it completes a task.

Suppose you pay a lawn service $20 an hour to mow. The first week, it takes two hours and costs $40. A week later, the job takes eight hours and costs $160. On week three, the job takes five hours, but only half your lawn is trimmed.

Here’s an example of that, from the researchers’ data. We selected these two prompts to demonstrate the variability of results.

Here you see two prompts—computer-programming requests—given to models from leading AI labs at Anthropic, Google and OpenAI. From each lab, there’s both a fast model and a deeper reasoning model.Scroll to continue ↓

__________________

Each model received the same prompt five times—and it cost a different amount each time.

__________________

In one instance, a quick-response model ran up an inconsistently high cost.

__________________

Even zoomed into the lower-cost end of the pricing spectrum, the results varied widely.

__________________

The full picture shows pricing variability, not only among models but within a given model’s output.

__________________

And even these unsuccessful runs—when the model didn’t return the anticipated correct solution—used up tokens and cost money.

__________________

Google, Anthropic and OpenAI—whose models researchers used for this testing—have since released new models that perform better in industry benchmarks.

“Some prompt-level fluctuation is inherent to AI, and our testing shows this averages out across a high volume of real-world, diverse workloads,” said a Google spokeswoman in a statement. “Total costs depend on many factors for a given task,” she added, “which can make it hard to forecast new and evolving technology with precision.” She said the company offers customers more control via spending caps and flexible pricing.

As more companies add AI to their daily workflows, they will also need strategies to manage their use, such as training employees on which model to use for a particular task. Otherwise, they could end up with IT bills that come out of a black box.

News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.

Write to Stephanie Stamm at stephanie.stamm@wsj.com



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *