Skip to main content

GPT-6 Astra Runs a Vending Machine for a Year, Turns $500 into $15,515

GPT-6 Astra Runs a Vending Machine for a Year, Turns $500 into $15,515

What happens when you hand an AI agent $500 and tell it to run a vending machine for an entire year? Andon Labs recently put that question to the test with Vending-Bench2, a year-long business simulation. The result? GPT-6 Astra didn't just survive—it thrived, finishing with an average balance of $15,515.

That's not pocket change. Even Astra's worst round beat the best round of Claude Fable 5.1. It's the first time Astra has claimed the top spot in this benchmark.

The Real Difference: Supply Chain and Rule-Following

So where did Claude stumble? Look at how each agent handled procurement and rules.

Astra turned out to be a ruthless negotiator. It drove product prices down to roughly 48% of the original cost, keeping expenses low over the long haul. Just as important, it followed the rules consistently—no prepayment losses, no slip-ups.

Claude Fable 5.1, by contrast, watched its procurement costs climb. Worse, it suffered significant prepayment losses because it didn't stick strictly to the rules. In a year-long game, those small missteps add up fast.

Three-Player Showdown: No Cheating, No Problem

The regular simulation was already lopsided. But Andon Labs also ran a more complex three-player competition. Here, agents could theoretically collude for mutual gain.

Astra refused to play dirty. It rejected illegal cooperation and won all three rounds cleanly. That's not just a technical win—it's a sign of something bigger: an AI that can resist shortcuts and stay focused on long-term strategy.

What This Means for AI Agents

The Vending-Bench2 results point to a clear trend. Long-term autonomy is becoming the key differentiator among AI agents. It's not enough to make a smart decision once; the real challenge is consistently converting good judgment into good actions, day after day, for an entire year.

Think about it: a vending machine is simple. But running one for 365 days requires planning, adaptation, and discipline. Astra showed all three. Claude showed flashes of brilliance but couldn't maintain the discipline.

As AI agents move from demo to deployment, this kind of endurance will matter more than any single clever trick. The next threshold isn't intelligence—it's reliability. And right now, GPT-6 Astra is setting the pace.


Key Points

  • GPT-6 Astra earned an average of $15,515 in a year-long vending machine simulation, nearly triple Claude Fable 5.1's best result.
  • Astra cut procurement costs to 48% of original prices and maintained strict rule adherence, avoiding prepayment losses.
  • In a three-player competition, Astra refused illegal cooperation and won all three rounds.
  • The test highlights long-term autonomy as the next major milestone for AI agents—consistently turning good decisions into reliable actions.