The best AI shopkeeper made $15,515, and a competent human would have made $63,000
GPT-6 Astra tops Andon Labs' Vending-Bench 2 after a simulated year of running a vending machine. Andon's own estimate for a good human strategy is four times higher.

GPT-6 Astra now leads Vending-Bench 2 with an average balance of $15,514.70 across five runs, ahead of Claude Opus 5 on $11,181.87 and GPT-5.6 Sol on $9,619.37. The benchmark, from Andon Labs, hands an agent $500 and a simulated year to source stock, negotiate with suppliers, price inventory and stay solvent. Rohan Paul's summary of the result went round on September 9.

Big score by GPT-6 Astra here. it averaged $15,515 on Vending-Bench 2 bench without the collusion or failed prepayments Andon observed in Fable.
Across 6 solo runs, Fable averaged $5,422, and even its best result finished below Astra’s worst.
Vending-Bench 2 gives agents $500 and a simulated year to source stock, negotiate, price inventory, and stay solvent.
Most of Astra’s $10,093 lead came from purchasing discipline, with about $8,540 less spent per run on supplier payments.
Fable’s payment-verified Coke price rose from $1.17 to $2.21 across the year, while Astra’s final-period average was $1.15 in available orders.
Fable also made 45 identified failed prepayments to closed suppliers, losing $14,331 across six runs; Astra recorded $0 despite encountering 64 closures.
The deeper failure is execution over time: Fable wrote a rule against paying before confirmation, then violated it days later and recognized the loss afterward.
Astra checked for a fresh supplier reply before 99% of repeat payments, compared with 58% for Fable.
Across both tests, Astra kept negotiation targets, payment checks, and competitive behavior steadier as the simulated horizon stretched.

Everyone quoted the leaderboard. But the interesting number sits further down the page.
Andon estimates that a competent human, playing the same game, would clear about $63,000. They build it conservatively — take the most profitable item the models themselves found (family-size Doritos), assume a negotiator gets half price, assume someone does the data analysis on the first sixty days of sales and then stocks optimally. So: $206 a day for 302 days.
So the best commercial agent in the world, running a single vending machine with a year and no distractions, gets about a quarter of the way to a diligent person. Andon's own trend line through the frontier releases is +$822 a month, with an R-squared of 0.95. Hold that rate and the models reach the human number in roughly 58 months. Call it mid-2031.
We would not bet the date, but we would bet the direction of the error, and it is probably that the line bends up before then rather than after.
Here is the detail we cannot stop looking at. The system prompt charges the agent for its own thinking, at $100 per million output tokens, deducted weekly. Andon also says a full run averages 60 to 100 million output tokens. If those two numbers meet on the same ledger, the agent is paying somewhere between $6,000 and $10,000 a year to think — against a winning balance of $15,515. And the single largest cost of running this business turns out to be the shopkeeper's mind.
What separates the top models is not cleverness so much as not getting bored. Andon says the leaders share two traits, a tool-use rate that holds steady across the whole simulated year with no degradation, and effective sourcing through persistent negotiation. Coherence, in other words. What is hard about a year is that it is a year.
One caution on the numbers circulating. Paul's post compares Astra against Claude Fable 5.1 across six solo runs at an average of $5,422, with 45 failed prepayments to closed suppliers. Andon's public leaderboard averages five runs and does not list Fable 5.1 in its visible top ten, so we cannot confirm those figures from the page. The comparison we can confirm is the Arena, where three agents compete at the same location. In round twelve, run on September 4, Astra took $12.4k, GLM-5.3 $7.8k and Claude Fable 5.1 $5.7k, the last of the three.
Does any of this tell you whether an agent can run your business? Not directly. But it tells you the shape of the failure, which is more useful. Models do not fail these runs by being unable to do the task on day one. They fail by drifting on day two hundred, and the benchmark that finally moves will be the one where nobody notices the agent at all.
