Someone outside Google finally measured a TPU, and the interesting number is not the headline
SemiAnalysis puts Ironwood at $0.181 per million tokens against $0.222 on a B200. The comparison Google most needs to win is the one that isn't published yet.

SemiAnalysis published the first third-party inference numbers for Google's TPUv7 Ironwood on September 7th, and put it at $0.181 per million tokens at 100 tokens per second per user, against $0.222 on an Nvidia B200 and $0.276 on a B300. On raw throughput at 20 tokens per second per user it reached 9,364 tokens per second per chip, about 5% above either Blackwell part.
For more than a decade, everyone who could tell you what a TPU costs per token worked at Google. That is the story. The ratios will move.

We are excited to bring the first open benchmarking of Google's TPUs to the world
Running every day, on many models + scenarios
$/token is better than B200 and B300
Huge shout-out to Google @inferact and the InferenceX team at SemiAnalysis to this effort that's taken many months
The alert SemiAnalysis posted the next morning says "50% better perf per dollar than Blackwell Ultra." Read the article and the 50.4% figure is against the B200. Against B300 — which is Blackwell Ultra — the same run shows 96.0%. So the promotional tweet names the wrong chip and, in doing so, undersells its own result by half. We mention it because this is the number that will be repeated for a year, and somebody should say which comparison it came from.

ALERT🚨🚨: On apples-to-apples InferenceX Comparisons, TPUv7 Ironwood achieves 50% better perf per dollar than Blackwell Ultra on the new TorchTPU external inference stack!
TorchTPU brings native PyTorch to TPUs, & along with Google open-sourcing a bunch of their Pallas inference kernels, Google has laid out a solid foundation for TPU to rapidly externalize. We at SemiAnalysis strongly believe that TPU externalization is heading in the right direction and moving full steam ahead.
Now the part that matters more. Every figure above is aggregated serving against aggregated serving, on Qwen3.5 397B in FP8, chosen as a bring-up model. Large operators run disaggregated prefill and decode. SemiAnalysis is direct about what happens there: "today, GB200/300 NVL72 is more competitive on perf per dollar in a disagg-vs-disagg comparison," and in the mixed comparison they do publish, a GB300 NVL72 running disaggregated holds roughly a 30% advantage over TPUv7 running aggregated in the middle of the latency curve.
There is a latency cost too, and it is not small.
So which chip is cheaper? It depends on where on the curve your product lives, which is the honest answer and also the reason vendor benchmarks pick a point and stop. If you serve a chat interface where the first token has to land in under three seconds, the Ironwood advantage at concurrency 256 is not available to you at all.
We would read this as a software release wearing a silicon headline. TorchTPU lets vLLM treat a TPU as a native PyTorch device instead of translating through JAX, Google has open-sourced Pallas inference kernels, and the stack is due out of private beta around October. TPUv7 has no native FP4, so Nvidia keeps the lead whenever a customer accepts FP4 quality loss, and Google's answer to that is TPUv8i, which does not exist to benchmark.
The fair objection to all of it: SemiAnalysis thanks sixteen named Google engineers and the Inferact team in the piece, Google supplied the stack under test, and the TCO figures come from SemiAnalysis's own subscription model rather than from an invoice. That is not a scandal — you cannot benchmark a chip nobody will sell you without help from the seller — but it is a long way from an independent lab.
Our expectation is that the promised TPUv7-disaggregated versus GB300 NVL72-disaggregated follow-up, whenever it lands, shows Nvidia still ahead on the middle of the Pareto curve, and that the perf-per-dollar gap in Google's favour survives only at the high-latency end. A first published disagg run where TPUv7 matches GB300 across the curve would settle it against us.
Anthropic has committed to more than a million of these chips, most of them for training. Somebody at Anthropic already knew all of this. Now the rest of us have a number to argue with.
