Astra and Fable 5.1 tie at 53. The interesting number is what happened to the Flash models when the benchmark changed
Artificial Analysis moved to Terminal-Bench 4.0 and a private automation test set. The two frontier models held. Gemini 3.8 Flash and Muse Spark 1.3, SemiAnalysis says, did not.
Artificial Analysis published version 4.3 of its Intelligence Index on Monday. GPT‑6 Astra (max) and Claude Fable 5.1 (max, with fallback) both score 53. Claude Opus 5 is at 51, Claude Fable 5 at 50, Meta's Muse Spark 1.3 at 48, GPT‑5.6 Sol at 47. Among open-weights models GLM‑5.3 and Kimi K3 lead at 44.
The version bump is the story. Two changes: Terminal-Bench went from 2.1 to 4.0, "66 multi-step tasks" run in a minimal mini‑SWE‑agent harness instead of Terminus 2, and τ³‑Banking was replaced by AutomationBench‑AA, Artificial Analysis's implementation of Zapier's 657 business-workflow tasks with a private test set. The weight on evaluations with private tasks rose from 40 to 45 percent. The stated reasons: "adds more private test sets to prevent gaming, and reduces saturation."
Who moved
SemiAnalysis said the quiet part on Monday.

Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵 t.co/K2ccQjWm11

Both models had been positioned as frontier-class at a fraction of the price. Erin Woo reported on September 1 that Google employees "prefer it to Opus on coding in side-by-side testing." Cursor added Gemini 3.8 Flash on the 2nd and Muse Spark 1.3 on the 8th, the day Meta launched its consumer agent on the same model. On harder terminal tasks with a fresh test set, according to SemiAnalysis, the gap to the frontier came back.
Boris Cherny at Anthropic offered his own read of a different axis: "OpenAI's new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk," adding that "evaluating and naming other labs turns out to be a great way to encourage them to train more aligned models." And antirez, from the other side: "Do you believe me now, after the Astra issue, that Artificial Analysis numbers are broken?"
Effort levels and cost
Why does a single number mislead? Artificial Analysis's new release pages answer it: GPT‑6 Astra spans 46–53 on the index across its effort settings, at 1.6 to 8.2 minutes per task. Fable 5.1 spans 47–53 at 4.2 to 12.2 minutes. OpenAI's Thibault Sottiaux says "GPT‑6 Astra on low performs better than GPT‑5.6 Sol on high." Simon Willison's pelican grid supports that in miniature: Astra low beat every Sol setting for 9.55 cents, despite Astra's list price ($10/$50 per million tokens) being double Sol's, because it used far fewer tokens at each level. OpenAI holds "the majority of the cost-efficiency frontier," per Artificial Analysis, with Fable 5.1 (xhigh) at the top and GLM‑5.3‑Flash further down.
Our read
The tie at 53 isn't the finding. It's the tolerance. Two labs with the same compute class and the same benchmark targets landing on the same integer tells you the index measures what they optimise for. The finding is that a harness change plus a private test set moved the Flash-class models and not the frontier ones, which is about the clearest evidence this year that benchmark scores below the frontier are partly a measure of how hard a lab tuned for the test. Cheap "frontier-class" claims deserve a discount until a private-set number exists.
If you're choosing a model this month, choose by effort level and time-per-task, not by headline index. Astra at low or medium is the cost-efficient default for agentic coding on current numbers. Fable 5.1 at xhigh is the pick when the task justifies twelve minutes. Gemini 3.8 Flash and Muse Spark 1.3 are worth their price on work that looks like Terminal-Bench 2.1 and unproven, so far, on work that looks like 4.0.
We'd expect Artificial Analysis to ship Index v5 with a majority private weighting before the end of 2026, and at least one Flash-class model to drop more than five points on the transition. If Gemini 3.8 Flash or Muse Spark 1.3 recovers to within three points of the frontier on Terminal-Bench 4.0 in Artificial Analysis's harness by then, SemiAnalysis's benchmaxxing claim was wrong, and so was our reading of it.