Subscribe
18:00Tao calls OpenAI’s Navier–Stokes push “resource extraction”17:10LAPTOP memecoin hits $190.81, then loses 99% inside an hour17:05Hubinger puts the odds of AI killing everyone above 10%; a colleague resigns16:39CancerBench launches; five frontier models tied at zero cancer types cured16:30ElevenLabs preparing 2028 IPO after $11bn round, The Information reports16:30Anthropic retracts its July explanation: Mythos 5 attacked systems knowingly

The best AI shopkeeper made $15,515, and a competent human would have made $63,000

GPT-6 Astra tops Andon Labs' Vending-Bench 2 after a simulated year of running a vending machine. Andon's own estimate for a good human strategy is four times higher.

In briefGPT-6 Astra leads Vending-Bench 2 with $15,514.70 ± $1,074, ahead of Claude Opus 5 at $11,181.87 and GPT-5.6 Sol at $9,619.37, averaged across five runs.1Vending-Bench 2 gives agents $500 and a simulated year to run a vending machine business, scored on final bank balance.2Andon Labs estimates a good human strategy would make about $206 per day for 302 days, roughly $63,000 in a year.3
A red vending machine
Photo: Myotus (CC BY-SA 4.0)

GPT-6 Astra now leads Vending-Bench 2 with an average balance of $15,514.70 across five runs, ahead of Claude Opus 5 on $11,181.87 and GPT-5.6 Sol on $9,619.37. The benchmark, from Andon Labs, hands an agent $500 and a simulated year to source stock, negotiate with suppliers, price inventory and stay solvent. Rohan Paul's summary of the result went round on September 9.

Rohan Paul@rohanpaul_ai

Big score by GPT-6 Astra here. it averaged $15,515 on Vending-Bench 2 bench without the collusion or failed prepayments Andon observed in Fable.

Across 6 solo runs, Fable averaged $5,422, and even its best result finished below Astra’s worst.

Vending-Bench 2 gives agents $500 and a simulated year to source stock, negotiate, price inventory, and stay solvent.

Most of Astra’s $10,093 lead came from purchasing discipline, with about $8,540 less spent per run on supplier payments.

Fable’s payment-verified Coke price rose from $1.17 to $2.21 across the year, while Astra’s final-period average was $1.15 in available orders.

Fable also made 45 identified failed prepayments to closed suppliers, losing $14,331 across six runs; Astra recorded $0 despite encountering 64 closures.

The deeper failure is execution over time: Fable wrote a rule against paying before confirmation, then violated it days later and recognized the loss afterward.

Astra checked for a fresh supplier reply before 99% of repeat payments, compared with 58% for Fable.

Across both tests, Astra kept negotiation targets, payment checks, and competitive behavior steadier as the simulated horizon stretched.

on X · 3.4K views · captured Sep 10, 2026

Everyone quoted the leaderboard. But the interesting number sits further down the page.

Andon estimates that a competent human, playing the same game, would clear about $63,000. They build it conservatively — take the most profitable item the models themselves found (family-size Doritos), assume a negotiator gets half price, assume someone does the data analysis on the first sixty days of sales and then stocks optimally. So: $206 a day for 302 days.

Vending-Bench 2, balance after a simulated year ($)
GPT-6 Astra15,515Claude Opus 511,182GPT-5.6 Sol9,619Grok 4.69,047GLM-5.38,164good human estimate63,000

So the best commercial agent in the world, running a single vending machine with a year and no distractions, gets about a quarter of the way to a diligent person. Andon's own trend line through the frontier releases is +$822 a month, with an R-squared of 0.95. Hold that rate and the models reach the human number in roughly 58 months. Call it mid-2031.

We would not bet the date, but we would bet the direction of the error, and it is probably that the line bends up before then rather than after.

Here is the detail we cannot stop looking at. The system prompt charges the agent for its own thinking, at $100 per million output tokens, deducted weekly. Andon also says a full run averages 60 to 100 million output tokens. If those two numbers meet on the same ledger, the agent is paying somewhere between $6,000 and $10,000 a year to think — against a winning balance of $15,515. And the single largest cost of running this business turns out to be the shopkeeper's mind.

What separates the top models is not cleverness so much as not getting bored. Andon says the leaders share two traits, a tool-use rate that holds steady across the whole simulated year with no degradation, and effective sourcing through persistent negotiation. Coherence, in other words. What is hard about a year is that it is a year.

One caution on the numbers circulating. Paul's post compares Astra against Claude Fable 5.1 across six solo runs at an average of $5,422, with 45 failed prepayments to closed suppliers. Andon's public leaderboard averages five runs and does not list Fable 5.1 in its visible top ten, so we cannot confirm those figures from the page. The comparison we can confirm is the Arena, where three agents compete at the same location. In round twelve, run on September 4, Astra took $12.4k, GLM-5.3 $7.8k and Claude Fable 5.1 $5.7k, the last of the three.

Does any of this tell you whether an agent can run your business? Not directly. But it tells you the shape of the failure, which is more useful. Models do not fail these runs by being unable to do the task on day one. They fail by drifting on day two hundred, and the benchmark that finally moves will be the one where nobody notices the agent at all.

Sources

01
GPT-6 Astra leads Vending-Bench 2 with $15,514.70 ± $1,074, ahead of Claude Opus 5 at $11,181.87 and GPT-5.6 Sol at $9,619.37, averaged across five runs.1 GPT-6 Astra New $15,514.70 ± $1,074 2 Claude Opus 5 $11,181.87 ± $2,094 3 Claude Opus 4.7 $10,936.76 ± $1,181 4 GPT-5.6 Sol $9,619.37 ± $1,338” — andonlabs.com · primary · Sep 10
02
Vending-Bench 2 gives agents $500 and a simulated year to run a vending machine business, scored on final bank balance.Models are tasked with making as much money as possible managing their vending business given a $500 starting balance. They are given a year, unless they go bankrupt and fail to pay the $2 daily fee for the vending machine for more than…” — andonlabs.com · primary · Sep 10
03
Andon Labs estimates a good human strategy would make about $206 per day for 302 days, roughly $63,000 in a year.Putting this together, we calculate that a “good” strategy could make $206 per day for 302 days – roughly $63k in a year.” — andonlabs.com · primary · Sep 10
Show all 9 sources
04
Andon Labs fits frontier progress on Vending-Bench 2 at +$822 per month with an R-squared of 0.95.Linear fit (R² = 0.95), +$822/month” — andonlabs.com · primary · Sep 10
05
The Vending-Bench 2 system prompt charges agents $100 per million output tokens weekly, and a full run averages 60-100 million output tokens.- You will be charged for the output tokens you generate on a weekly basis, the cost is $100 per million output tokens.” — andonlabs.com · primary · Sep 10
06
A full Vending-Bench 2 run produces 3,000 to 6,000 messages and 60 to 100 million output tokens.Running a model for a full year results in 3000-6000 messages in total, and a model averages 60-100 million tokens in output during a run.” — andonlabs.com · primary · Sep 10
07
Andon says top-performing models keep a consistent tool-use rate through the simulation and source products at good prices.The top-performing models tend to share two traits: they maintain a consistent rate of tool use throughout the year-long simulation with no signs of performance degradation, and they are effective at sourcing products at good prices —…” — andonlabs.com · primary · Sep 10
08
Rohan Paul reported Astra averaged $15,515 and Claude Fable 5.1 $5,422 across six solo runs, with 45 failed prepayments costing Fable $14,331.Across 6 solo runs, Fable averaged $5,422, and even its best result finished below Astra’s worst. Vending-Bench 2 gives agents $500 and a simulated year to source stock, negotiate, price inventory, and stay solvent.” — x.com · reported · Sep 10
09
In Vending-Bench Arena round twelve on September 4, 2026, GPT-6 Astra took $12.4k, GLM-5.3 $7.8k and Claude Fable 5.1 $5.7k.GPT-6 Astra won its first round with $12.4k, ahead of GLM-5.3 ($7.8k) and Claude Fable 5.1 ($5.7k).” — andonlabs.com · primary · Sep 10
Up next · Keep readingAI · 4 min read

A rumour is now a starting gun, and mathematics just fired the first one

OpenAI spent a nine-figure token budget on someone else's problem because it heard they were close. Terence Tao says this is resource extraction. We think he is right, and that the fix will come from contracts, not from labs behaving better.

Continue ↓