Subscribe
18:00Tao calls OpenAI’s Navier–Stokes push “resource extraction”17:10LAPTOP memecoin hits $190.81, then loses 99% inside an hour17:05Hubinger puts the odds of AI killing everyone above 10%; a colleague resigns16:39CancerBench launches; five frontier models tied at zero cancer types cured16:30ElevenLabs preparing 2028 IPO after $11bn round, The Information reports16:30Anthropic retracts its July explanation: Mythos 5 attacked systems knowingly
Business3 min read

What a token costs and what it burns are unrelated numbers

Vals AI's audit finds a model fifteen times cheaper per token with the same footprint, days before its chief executive said token spend may start to eclipse salary spend.

In briefDeepSeek V4 Pro is nearly 15x cheaper per output token than Kimi K3 with roughly the same environmental footprint, and about 6x cheaper per task1Token pricing is distorted by providers discounting to capture market share2Vals estimated impacts from token usage on its 2,157-task index using EcoLogits, covering open-weight models only3

Vals AI published an audit on 3 September of what open-weight models actually consume, and the headline finding is that the price you pay has almost nothing to do with the resources you use. DeepSeek V4 Pro is nearly fifteen times cheaper per output token than Kimi K3 and carries roughly the same environmental footprint. Per completed task, it is about six times cheaper. Same work, same physical cost, wildly different invoice.

Energy for one full Vals Index run, kWh
Ling 3.0 Flash23.8MiniMax M2.734.8Kimi K2.5 Thinking108.5DeepSeek V4 Flash139.7GLM-5.2328.2Kimi K31,700Qwen3.8 Max1,800

The method matters, so here it is briefly. Vals took token usage from its own index of 2,157 long-horizon agentic tasks across finance, coding, healthcare and law, then ran it through EcoLogits, an open-source estimator that needs a model's parameter counts. Which is why only open-weight models appear: fourteen Chinese, one from Thinking Machines Lab, one from Mistral. Nobody can do this to GPT or Claude (you need the parameter counts, and those are the whole secret).

Now the arithmetic nobody has done. Kimi K3 needs about 1.7m watt-hours for the full index, which is roughly 788 watt-hours per task. Sam Altman and Google both put an average text query between 0.24 and 0.34 watt-hours last year. So one agentic task on the most accurate open model costs something like 2,600 average chatbot queries. Ling 3.0 Flash, at the bottom of the table, comes in near 11 watt-hours a task — about seventy times less than Kimi K3, for a model built to answer simple things.

That is the shape of the problem. Capability is being bought in orders of magnitude, and Bloomberg's write-up of the same report puts the gap between a quick question and building a web app at up to 10,000 times.

So why does the price not track any of this? Because it is set by market share. Vals says so plainly: providers offer tokens at steep discounts while trying to capture share, which distorts the relationship between price and impact. Cheap tokens can also be less efficient tokens, since a model that thinks longer bills you less per token and more per job.

Six days later the company's chief executive said the quiet part in a podcast.

a16z@a16z

Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival:

"The labs side of it is very clear. If you're raising lots of money, investing heavily in building models, it's essential for you to show why your model is getting better and why the customer should pay a premium for them."

"But what I think is still underappreciated is, on the enterprise side, this is turning out to be existential as well."

"We're in this world where it is still very unclear what ROI looks like and how to value this intelligence that's being used... Token spend may start to eclipse salary spend."

"If this is such a meaningful line item in your costs, you have to justify the ROI much more cleanly... Over time, I think a firm really is just its evals."

"The ability for a company to make its evals legible in order to solve this ROI calculus is going to be the reason why that company wins out over its competitors in the long term."

@RayanKrishnan @eriktorenberg

on X · 23.5K views · captured Sep 10, 2026

Rayan Krishnan's claim that "token spend may start to eclipse salary spend" is the sentence that turns an environmental report into a finance one. If inference is heading toward the size of payroll, then a per-token price is a payroll system that bills by the keystroke. You would not run a headcount plan that way. (Imagine a payroll line item denominated in emails sent.)

Our read: the unit of AI procurement is going to move from tokens to completed tasks, and the vendors who benefit are the ones whose models finish quickly rather than the ones with the lowest sticker. We would expect at least one major enterprise AI contract to be reported on a per-task basis within twelve months. If the market is still quoting dollars per million tokens in late 2027 with no efficiency term attached, we were early or wrong.

The honest caveat is that these are estimates on assumed hardware, not meter readings, and Vals counts several of the labs it measured as customers. It says so. Take the ratios seriously and the absolute numbers loosely.

One question worth putting to your own vendor, then. How many watt-hours did my last job take, and if you cannot answer, what exactly are you charging me for?

Sources

01
DeepSeek V4 Pro is nearly 15x cheaper per output token than Kimi K3 with roughly the same environmental footprint, and about 6x cheaper per taskOn agentic tasks, cheaper models can be less token-efficient, driving up both resource demand and cost per task. DeepSeek V4 Pro, for instance, is nearly 15x cheaper per output token than Kimi K3 yet has roughly the same environmental…” — vals.ai · primary · Sep 10
02
Token pricing is distorted by providers discounting to capture market shareThe relationship between pricing and environmental impacts is further distorted by market dynamics, as model providers may offer tokens at steep discounts while trying to capture market share.” — vals.ai · primary · Sep 10
03
Vals estimated impacts from token usage on its 2,157-task index using EcoLogits, covering open-weight models onlyTo estimate environmental impacts, we collected token usage data for tasks from our Vals Index benchmark. The index is composed of long-horizon, multi-turn, agentic tasks from different industries such as finance, coding, healthcare, and…” — vals.ai · primary · Sep 10
Show all 12 sources
04
Energy for a full index run ranges from 23.8k Wh for Ling 3.0 Flash to 1.7M Wh for Kimi K3 and 1.8M Wh for Qwen3.8 Max1 Ling 3.0 Flash 2607 23.8k Wh 1 Day (US home) 29.0k Wh 2 MiniMax M2.7 34.8k Wh 3 Mistral Medium 3.5 47.0k Wh ... 15 Kimi K3 1.7M Wh 16 Qwen3.8 Max 1.8M Wh” — vals.ai · primary · Sep 10
05
Kimi K3 leads the Vals Index at 57% accuracy with DeepSeek V4 Flash 4% behind at a much smaller footprintKimi K3 tops the open-weight leaderboard with 57% accuracy on the Vals Index, with the runner-up, DeepSeek V4 Flash, only 4% behind. This performance comes at the cost of a larger environmental footprint.” — vals.ai · primary · Sep 10
06
Vals found complex tasks can have an environmental impact 10,000 times greater than simple queries, and 14 of the 16 models were ChineseThe firm found that lengthier tasks that require AI models to “think” for longer, like building a software app, can have an environmental impact 10,000 times greater than simple queries, such as asking a basic question that can be…” — spokesman.com · reported · Sep 10
07
Altman and Google put the energy of an average text query between 0.24 and 0.34 watt hoursOpenAI Chief Executive Officer Sam Altman wrote in 2025 that an average ChatGPT query uses about 1/15th of a teaspoon of water, and estimates by Altman and Alphabet Inc.'s Google the same year put the energy toll of the average text…” — spokesman.com · reported · Sep 10
08
Fourteen of sixteen models analysed were from Chinese companies, one from Thinking Machines Lab and one from MistralFourteen of the 16 models analyzed were developed by Chinese companies, with one from U.S.-based Thinking Machines Lab and one from France's Mistral AI.” — spokesman.com · reported · Sep 10
09
Vals counts several leading AI firms as customers, including some whose models it measuredThe companies whose models were analyzed did not immediately respond to requests for comment. Vals counts several leading AI firms as its customers, including some of these companies.” — spokesman.com · reported · Sep 10
10
Vals chief executive Rayan Krishnan says token spend may start to eclipse salary spend and that a firm is its evals“We're in this world where it is still very unclear what ROI looks like and how to value this intelligence that's being used... Token spend may start to eclipse salary spend.”” — x.com · primary · Sep 10
11
Krishnan argues making evals legible is what decides which company wins“The ability for a company to make its evals legible in order to solve this ROI calculus is going to be the reason why that company wins out over its competitors in the long term.”” — x.com · primary · Sep 10
12
Shirin Ghaffary reported the Vals findings on the gap between a simple question and vibecoding a full-stack appNEW: A new report from benchmarking firm @ValsAI shows that there can be a major difference in the environmental cost of asking a chatbot a simple question v. asking it to vibecode a full-stack web app for you.” — x.com · primary · Sep 10
Up next · Keep readingBusiness · 2 min read

ElevenLabs hired an IPO CFO the day before the IPO story

The Information says ElevenLabs is preparing a 2028 listing; the company had already announced Adyen's former CFO and a new revenue chief inside a week.

Continue ↓