Opus 5 runs vending machines

sudo_cowsay2 pts0 comments

Vending-Bench 2 | Andon Labs

Eval<br>Vending-Bench 2<br>We're releasing Vending-Bench 2, a benchmark for measuring AI model performance on running a business over long time horizons. Models are tasked with running a simulated vending machine business over a year and scored on their bank account balance at the end.

Long-term coherence in agents is more important than ever. Coding agents can now write code autonomously for hours, and the length and breadth of tasks AI models are able to complete is likely to increase. We expect models to soon take active part in the economy, managing entire businesses. But to do this, they have to stay coherent and efficient over very long time horizons. This is what Vending-Bench 2 measures: the ability of models to stay coherent and successfully manage a simulated business over the course of a year. Our results show that while models are improving at this, current frontier models handle this with varying degrees of success.<br>Money balance over time

Average across 5 runs

Days in simulation

Frontier Open All<br>Current leaderboard<br>Average across 5 runs<br>Model Money Balance 1 Claude Opus 5 New<br>$11,181.87 ± $2,094<br>2 Claude Opus 4.7<br>$10,936.76 ± $1,181<br>3 GPT-5.6 Sol<br>$9,619.37 ± $1,338<br>4 GLM-5.2<br>$8,313.78 ± $1,084<br>5 Claude Opus 4.6<br>$8,017.59 ± $1,367<br>6 GPT-5.5<br>$7,523.84 ± $1,346<br>7 GPT-5.6 Terra<br>$7,343.21 ± $373<br>8 Claude Sonnet 4.6<br>$7,204.14 ± $722<br>9 Muse Spark 1.1 New<br>$6,520.47 ± $1,016<br>10 Claude Sonnet 5<br>$6,377.70 ± $784

Show 48 more

The leaderboard shows significant spread in performance. The top-performing models tend to share two traits: they maintain a consistent rate of tool use throughout the year-long simulation with no signs of performance degradation, and they are effective at sourcing products at good prices — whether through persistent negotiation or by finding better suppliers.<br>Vending-Bench Arena<br>Vending-Bench Arena is a version of Vending-Bench 2 that adds a crucial component: competition. It's our first multi-agent eval, where all participating agents manage their own vending machine at the same location. This leads to price wars and tough strategy decisions. Agents may also collaborate and trade with each other if they so choose, but all scoring is individual.

See the arena results

Performance vs. release date

SOTA frontier models are labeled and a trend line is fitted through them, with a projection into the near future.

Linear fit (R² = 0.96), +$734/month

Linear scale Log scale<br>Frontier lag analysis

Comparing SOTA frontier progression between model groups, with linear regression and projected crossover points.

Chinese: +$1,047/month (R² = 0.98) · Western: +$734/month (R² = 0.96) · Chinese lags by ~100 days · Projected crossover: Mar 2027<br>Chinese Western<br>Only profitable models are included.

Chinese vs Western Open vs Closed<br>Score vs. cost per run

Score vs. mean cost per run using each LLM provider’s API to run Vending-Bench 2. Costs are calculated from the provider’s input and output token pricing, without caching.<br>Score ($)

Cost per run ($)

Improvements from our original Vending-Bench<br>Vending-Bench 2 keeps the core idea from Vending-Bench of managing a business in a lifelike setting, but introduces more real-world messiness inspired by learnings from our vending machine deployments:<br>Suppliers may be adversarial and actively try to exploit the agent, quoting unreasonable prices or even trying bait-and-switch tactics. The agents must realize this and look for other options to stay profitable.<br>Negotiation is key to success. Even honest suppliers will try to get the most out of their customers.<br>Deliveries can be delayed and trusted suppliers can go out of business, forcing agents to build robust supply chains and always have a plan B.<br>Unhappy customers can reach out at any time demanding costly refunds.<br>We’ve also streamlined the scoring system, evaluating models on money balance after a year and clarified the scoring criteria, such that agents know exactly what to optimize for. Better planning tools, such as proper note-taking and reminder systems have been added as well.<br>How Vending-Bench works<br>Models are tasked with making as much money as possible managing their vending business given a $500 starting balance. They are given a year, unless they go bankrupt and fail to pay the $2 daily fee for the vending machine for more than 10 consecutive days, in which case they are terminated early. Models can search the internet to find suitable suppliers and then contact them through e-mail to make orders. Delivered items arrive at a storage facility, and the models are given tools to move items between storage and the vending machine. Revenue is generated through customer sales, which depend on factors such as day of the week, season, weather, and price.

Running a model for a full year results in 3000-6000 messages in total, and a model averages 60-100 million tokens in output during a run.<br>System prompt<br>A good way to understand the benchmark is to read the...

vending models bench agents business year

Related Articles