What a token costs and what it burns are unrelated numbers
Vals AI's audit finds a model fifteen times cheaper per token with the same footprint, days before its chief executive said token spend may start to eclipse salary spend.
Vals AI published an audit on 3 September of what open-weight models actually consume, and the headline finding is that the price you pay has almost nothing to do with the resources you use. DeepSeek V4 Pro is nearly fifteen times cheaper per output token than Kimi K3 and carries roughly the same environmental footprint. Per completed task, it is about six times cheaper. Same work, same physical cost, wildly different invoice.
The method matters, so here it is briefly. Vals took token usage from its own index of 2,157 long-horizon agentic tasks across finance, coding, healthcare and law, then ran it through EcoLogits, an open-source estimator that needs a model's parameter counts. Which is why only open-weight models appear: fourteen Chinese, one from Thinking Machines Lab, one from Mistral. Nobody can do this to GPT or Claude (you need the parameter counts, and those are the whole secret).
Now the arithmetic nobody has done. Kimi K3 needs about 1.7m watt-hours for the full index, which is roughly 788 watt-hours per task. Sam Altman and Google both put an average text query between 0.24 and 0.34 watt-hours last year. So one agentic task on the most accurate open model costs something like 2,600 average chatbot queries. Ling 3.0 Flash, at the bottom of the table, comes in near 11 watt-hours a task — about seventy times less than Kimi K3, for a model built to answer simple things.
That is the shape of the problem. Capability is being bought in orders of magnitude, and Bloomberg's write-up of the same report puts the gap between a quick question and building a web app at up to 10,000 times.
So why does the price not track any of this? Because it is set by market share. Vals says so plainly: providers offer tokens at steep discounts while trying to capture share, which distorts the relationship between price and impact. Cheap tokens can also be less efficient tokens, since a model that thinks longer bills you less per token and more per job.
Six days later the company's chief executive said the quiet part in a podcast.

Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival:
"The labs side of it is very clear. If you're raising lots of money, investing heavily in building models, it's essential for you to show why your model is getting better and why the customer should pay a premium for them."
"But what I think is still underappreciated is, on the enterprise side, this is turning out to be existential as well."
"We're in this world where it is still very unclear what ROI looks like and how to value this intelligence that's being used... Token spend may start to eclipse salary spend."
"If this is such a meaningful line item in your costs, you have to justify the ROI much more cleanly... Over time, I think a firm really is just its evals."
"The ability for a company to make its evals legible in order to solve this ROI calculus is going to be the reason why that company wins out over its competitors in the long term."
@RayanKrishnan @eriktorenberg
Rayan Krishnan's claim that "token spend may start to eclipse salary spend" is the sentence that turns an environmental report into a finance one. If inference is heading toward the size of payroll, then a per-token price is a payroll system that bills by the keystroke. You would not run a headcount plan that way. (Imagine a payroll line item denominated in emails sent.)
Our read: the unit of AI procurement is going to move from tokens to completed tasks, and the vendors who benefit are the ones whose models finish quickly rather than the ones with the lowest sticker. We would expect at least one major enterprise AI contract to be reported on a per-task basis within twelve months. If the market is still quoting dollars per million tokens in late 2027 with no efficiency term attached, we were early or wrong.
The honest caveat is that these are estimates on assumed hardware, not meter readings, and Vals counts several of the labs it measured as customers. It says so. Take the ratios seriously and the absolute numbers loosely.
One question worth putting to your own vendor, then. How many watt-hours did my last job take, and if you cannot answer, what exactly are you charging me for?