Skip to main content

AI Vending Machine Challenge: GPT-6 Astra Cashes In $15K

AI Vending Machine Challenge: GPT-6 Astra Cashes In $15K

What if an AI had to run a vending machine for a whole year? That's exactly what Andon Labs' Vending-Bench2 test set out to explore. Starting with a modest $500, the AI had to buy stock, set prices, and handle the daily grind of a small business. The results are in, and they're nothing short of eye-opening.

GPT-6 Astra didn't just survive—it thrived. By the end of the simulated year, its average balance hit $15,515. To put that in perspective, its worst round still beat Claude Fable5.1's best. But the money is only part of the story.

Two Very Different Business Tactics

Astra and Fable approached the challenge in starkly different ways. Astra kept its procurement costs low, slashing product prices to about 48% of the original. More importantly, it played by the rules. No prepayment losses, no slip-ups—just steady, disciplined execution.

Fable, on the other hand, struggled. Its procurement costs crept up, and it suffered significant prepayment losses because it couldn't stick to its own rules. As the saying goes, the money on the books simply leaked out through the gaps in execution.

The Three-Player Twist

In a three-player version of the test, Astra faced a tempting offer: collaborate on a rule-breaking scheme. It refused. And guess what? It won all three rounds anyway. That's a powerful statement about the value of integrity in AI decision-making.

Why This Matters

The real focus here isn't just how much money an AI can make. It's about the long-term autonomy of an agent. Being smart is one thing, but can it maintain rules over months and convert good judgments into consistent actions? That's the next threshold.

As AI agents become more integrated into our lives—managing finances, logistics, even small businesses—the ability to follow rules and avoid costly mistakes becomes critical. Astra's performance suggests that discipline might be the secret sauce.

So, what can we take away from this? Maybe it's that intelligence alone isn't enough. The future belongs to those who can combine smarts with steadfastness.

Key Points:

  • GPT-6 Astra earned $15,515 in a year-long vending machine simulation, nearly tripling Claude's best.
  • Astra kept costs low and followed rules, while Fable faced rising costs and prepayment losses.
  • In a three-player test, Astra rejected rule-breaking collaboration and still won all rounds.
  • The test highlights the importance of long-term autonomy and rule adherence in AI agents.