Subscribe
18:00Tao calls OpenAI’s Navier–Stokes push “resource extraction”17:10LAPTOP memecoin hits $190.81, then loses 99% inside an hour17:05Hubinger puts the odds of AI killing everyone above 10%; a colleague resigns16:39CancerBench launches; five frontier models tied at zero cancer types cured16:30ElevenLabs preparing 2028 IPO after $11bn round, The Information reports16:30Anthropic retracts its July explanation: Mythos 5 attacked systems knowingly

Open weights sit at 58 percent of the frontier and 44 percent of the price

Ethan Mollick says open models are further from the frontier than they have been in a while. Two long-horizon leaderboards published this month agree on the score and disagree on the value.

In briefEthan Mollick said on September 9, 2026 that open weights models are further from the frontier than they have been in some time.1Yuchen Jin said on August 29 that open models are cost-efficient and that there is no default winner anymore.2On DeepsecBench, GLM-5.3 at high effort scores 21.91 at a cost of $16.22, against GPT-6 Astra xhigh at 37.79 for $63.70; Kimi K3 xhigh scores 17.49 and Qwen3.8 Max xhigh 16.47.3
An Nvidia GeForce RTX 5090 graphics card
Photo: ZMASLO (CC BY 3.0)

Ethan Mollick wrote on September 9 that open weights models are further from the frontier than we have seen in some time. Ten days earlier Yuchen Jin, who runs open models for a living, wrote that there is no default winner anymore. Both are informed. Both are right, about different axes, and the arithmetic that separates them has been sitting on two public leaderboards all week.

Ethan Mollick@emollick

Open weights models are further from the frontier than we have seen in some time. Mythos was launched in March, and, as good as the open model releases have been since then (K3 & GLM-5.3 are great) they aren’t that close to Mythos or Astra in practice. I assume that will change

on X · 23.7K views · captured Sep 10, 2026

Start with score. On Vercel's DeepsecBench, which scores models on finding real vulnerabilities in application code, the best open-weights configuration is GLM-5.3 at high effort with 21.91. GPT-6 Astra sits at 37.79. That is 58 percent of the frontier, and every open model on the board — Kimi K3 at 17.49, Qwen3.8 Max and DeepSeek V4 Flash at 16.47 — is further back still.

DeepsecBench, best configuration per model
GPT-6 Astra37.79GPT-5.6 Sol35.44Claude Opus 532.44GLM-5.321.91Kimi K317.49Qwen3.8 Max16.47

Andon Labs' Vending-Bench 2, which drops an agent into a simulated vending business for a year and scores it on the bank balance, lands in the same place from a completely different direction. Astra finished on $15,514.70, GLM-5.3 on $8,163.61 (GLM-5.2 came in fractionally ahead of its successor, which nobody at Z.ai will want to discuss). Fifty-three percent. Two evals with nothing in common (different harnesses, different task, different scoring unit) agreeing to within five points on how far back open weights are.

So Mollick is probably right about capability.

Now price it.

That DeepsecBench run cost Astra $63.70 and GLM-5.3 $16.22, which works out at $1.69 per point of score against $0.74. Per unit of capability the open model is 2.3 times cheaper, and that ratio is roughly why you keep finding open weights in the parts of a stack nobody puts in a slide. On Ollama's cloud, GLM-5.3-Flash runs at $0.15 per million input tokens. Astra is $10. (Different tiers of model, yes, and the ratio is still sixty-seven to one.)

And there is one board where the open model simply won. In Vending-Bench Arena round twelve, run on September 4 with three agents competing at the same location, GPT-6 Astra took $12.4k, GLM-5.3 took $7.8k, and Claude Fable 5.1 finished last on $5.7k. An open-weights model beat a frontier lab's newest release, in competition, at running a business for a simulated year.

What about running it yourself? Tom's Hardware spent September 8 on that question with Qwen 3.8 27B, roughly 17GB of four-bit weights, and found that VRAM capacity alone cannot overcome software and inference-engine bottlenecks. The single-card frontier is still not here.

Not for lack of weights. For lack of an inference engine that can feed them.

Our read is that the gap is widening in score and narrowing in money, and that most of the argument people are having is two groups measuring different columns. Andon publishes a regression on exactly this. The Chinese frontier is gaining $1,047 a month on Vending-Bench, the Western frontier $822, a lag of about 111 days, with crossover projected for October 2027.

We would take the other side of that date. The fit sits on one Astra data point that broke the Western trend upward, and a linear projection drawn through a step change is, as far as we can tell, the most reliable way to be wrong in public. Sometime before then an open-weights model will crack the top three of DeepsecBench, and if it does before Christmas we will have been too cautious here and will say so.

Until that happens, you are choosing between 58 percent of the answer at 44 percent of the cost. Most of the time, for most work, that is not a hard call — which is the part Mollick's post, correct as it is on capability, does not price.

Sources

01
Ethan Mollick said on September 9, 2026 that open weights models are further from the frontier than they have been in some time.Open weights models are further from the frontier than we have seen in some time. Mythos was launched in March, and, as good as the open model releases have been since then (K3 & GLM-5.3 are great) they aren’t that close to Mythos or…” — x.com · primary · Sep 10
02
Yuchen Jin said on August 29 that open models are cost-efficient and that there is no default winner anymore.Meanwhile, OSS models like GLM-5.3, GLM-5.3-Flash, and Kimi K3 are pretty good and cost-efficient. There’s no default winner anymore.” — x.com · primary · Sep 10
03
On DeepsecBench, GLM-5.3 at high effort scores 21.91 at a cost of $16.22, against GPT-6 Astra xhigh at 37.79 for $63.70; Kimi K3 xhigh scores 17.49 and Qwen3.8 Max xhigh 16.47.8 GLM-5.3 high zai/glm-5.3 21.91 18.5 % 80.7 % 11 $16.22 2h 05m 37s 1,232.1 Pi ... 1 GPT-6 Astra xhigh openai/gpt-6-astra 37.79 32.8 % 97.9 % 2 $63.70 49m 47s” — vercel.com · primary · Sep 10
Show all 10 sources
04
On Vending-Bench 2, GPT-6 Astra averaged $15,514.70 and GLM-5.3 $8,163.61 across five runs.1 GPT-6 Astra New $15,514.70 ± $1,074 ... 7 GLM-5.3 $8,163.61 ± $787” — andonlabs.com · primary · Sep 10
05
In Vending-Bench Arena round twelve on September 4, 2026, GPT-6 Astra won with $12.4k, ahead of GLM-5.3 on $7.8k and Claude Fable 5.1 on $5.7k.GPT-6 Astra won its first round with $12.4k, ahead of GLM-5.3 ($7.8k) and Claude Fable 5.1 ($5.7k). It is the highest balance any model has reached in the Arena so far, past the $10.6k GPT-5.5 posted in Round #8.” — andonlabs.com · primary · Sep 10
06
Andon Labs' regression puts Chinese frontier progress at $1,047 a month and Western at $822, a lag of about 111 days, with crossover projected for October 2027.Chinese: +$1,047/month (R² = 0.98) · Western: +$822/month (R² = 0.95) · Chinese lags by ~111 days · Projected crossover: Oct 2027” — andonlabs.com · primary · Sep 10
07
Ollama's cloud price for glm-5.3-flash is $0.15 per million input tokens.glm-5.3-flash $0.15 $0.03 $0.50” — ollama.com · primary · Sep 10
08
GPT-6 Astra's API price is $10 per million input tokens.Standard API pricing is $10/M input and $50/M output tokens” — x.com · reported · Sep 10
09
Tom's Hardware found on September 8 that Qwen 3.8 27B, about 17GB at four-bit quantization, is held back on consumer cards by software and inference-engine bottlenecks rather than VRAM.Totaling around 17GB for four-bit quantized weights and offering built-in multimodal capabilities on top of its general aptitude, Qwen 3.8 27B immediately grabbed the attention of everybody with an RTX 5090, RTX 4090, or RTX 3090” — tomshardware.com · reported · Sep 10
10
GLM-5.2 averaged $8,313.78 on Vending-Bench 2, fractionally ahead of GLM-5.3 at $8,163.61.6 GLM-5.2 $8,313.78 ± $1,084 7 GLM-5.3 $8,163.61 ± $787” — andonlabs.com · primary · Sep 10
Up next · Keep readingAI · 4 min read

A rumour is now a starting gun, and mathematics just fired the first one

OpenAI spent a nine-figure token budget on someone else's problem because it heard they were close. Terence Tao says this is resource extraction. We think he is right, and that the fix will come from contracts, not from labs behaving better.

Continue ↓