Open weights sit at 58 percent of the frontier and 44 percent of the price
Ethan Mollick says open models are further from the frontier than they have been in a while. Two long-horizon leaderboards published this month agree on the score and disagree on the value.

Ethan Mollick wrote on September 9 that open weights models are further from the frontier than we have seen in some time. Ten days earlier Yuchen Jin, who runs open models for a living, wrote that there is no default winner anymore. Both are informed. Both are right, about different axes, and the arithmetic that separates them has been sitting on two public leaderboards all week.

Open weights models are further from the frontier than we have seen in some time. Mythos was launched in March, and, as good as the open model releases have been since then (K3 & GLM-5.3 are great) they aren’t that close to Mythos or Astra in practice. I assume that will change
Start with score. On Vercel's DeepsecBench, which scores models on finding real vulnerabilities in application code, the best open-weights configuration is GLM-5.3 at high effort with 21.91. GPT-6 Astra sits at 37.79. That is 58 percent of the frontier, and every open model on the board — Kimi K3 at 17.49, Qwen3.8 Max and DeepSeek V4 Flash at 16.47 — is further back still.
Andon Labs' Vending-Bench 2, which drops an agent into a simulated vending business for a year and scores it on the bank balance, lands in the same place from a completely different direction. Astra finished on $15,514.70, GLM-5.3 on $8,163.61 (GLM-5.2 came in fractionally ahead of its successor, which nobody at Z.ai will want to discuss). Fifty-three percent. Two evals with nothing in common (different harnesses, different task, different scoring unit) agreeing to within five points on how far back open weights are.
So Mollick is probably right about capability.
Now price it.
That DeepsecBench run cost Astra $63.70 and GLM-5.3 $16.22, which works out at $1.69 per point of score against $0.74. Per unit of capability the open model is 2.3 times cheaper, and that ratio is roughly why you keep finding open weights in the parts of a stack nobody puts in a slide. On Ollama's cloud, GLM-5.3-Flash runs at $0.15 per million input tokens. Astra is $10. (Different tiers of model, yes, and the ratio is still sixty-seven to one.)
And there is one board where the open model simply won. In Vending-Bench Arena round twelve, run on September 4 with three agents competing at the same location, GPT-6 Astra took $12.4k, GLM-5.3 took $7.8k, and Claude Fable 5.1 finished last on $5.7k. An open-weights model beat a frontier lab's newest release, in competition, at running a business for a simulated year.
What about running it yourself? Tom's Hardware spent September 8 on that question with Qwen 3.8 27B, roughly 17GB of four-bit weights, and found that VRAM capacity alone cannot overcome software and inference-engine bottlenecks. The single-card frontier is still not here.
Not for lack of weights. For lack of an inference engine that can feed them.
Our read is that the gap is widening in score and narrowing in money, and that most of the argument people are having is two groups measuring different columns. Andon publishes a regression on exactly this. The Chinese frontier is gaining $1,047 a month on Vending-Bench, the Western frontier $822, a lag of about 111 days, with crossover projected for October 2027.
We would take the other side of that date. The fit sits on one Astra data point that broke the Western trend upward, and a linear projection drawn through a step change is, as far as we can tell, the most reliable way to be wrong in public. Sometime before then an open-weights model will crack the top three of DeepsecBench, and if it does before Christmas we will have been too cautious here and will say so.
Until that happens, you are choosing between 58 percent of the answer at 44 percent of the cost. Most of the time, for most work, that is not a hard call — which is the part Mollick's post, correct as it is on capability, does not price.
