AI and Humans Just Cracked a 2017 Math Puzzle Together
AI and Humans Just Cracked a 2017 Math Puzzle Together
Remember when AI could only beat humans at chess or Go? Those days feel almost quaint now. On FrontierMath, a brutal benchmark that makes most humans weep, a team of three researchers and an AI called GPT-6 Astra just pulled off something remarkable. They solved an open problem that had been sitting there since 2017, waiting for someone—or something—to crack it.
The problem itself sounds like a political science nightmare: in a committee election based on approval voting, is the "core" always empty? For the uninitiated, the core is that elusive group of voters who can't be outvoted by any other coalition. The original goal was to find a counterexample where the core is empty—a scenario where no stable, fair committee exists. But Astra didn't just find a counterexample. It proved the opposite: such a counterexample doesn't exist. In other words, an absolutely fair committee must always exist, no matter what.
That's not just solving a problem. That's rewriting the rules of the game.
But wait, there's more. The model didn't stop at the proof. It went on to invent a new voting rule based on something called "harmonic entropy," then provided a polynomial-time algorithm and proved that a local optimal solution already satisfies the core requirement. That's like not only finding a needle in a haystack but also building a magnet that makes needles easier to find in the future.
How did this collaboration work? The humans set the direction and logical framework. Astra handled the knowledge base, computational drills, and those sudden sparks of insight. Their division of labor looked less like a scientist using a tool and more like a well-oiled research team, with each member playing to their strengths. You might even say the AI was the enthusiastic postdoc who never sleeps.
This is the first time AI has solved a "major breakthrough" level math problem, and the significance ripples far beyond one theorem. It pushes large models one step further from being just "problem-solving machines" to becoming collaborators capable of discovering mathematical patterns. Even Epoch AI, which maintains the FrontierMath leaderboard, has added a new "Human + AI" status tag, specifically reserving a place for such joint achievements. It's a small change with big implications: the scientific community is starting to recognize that the best results might come from teams, not solo geniuses.
As mathematicians begin to play the role of "prompt engineers," the way scientific research is conducted is quietly being rewritten. Instead of spending months on tedious calculations, researchers can focus on asking the right questions and guiding the AI's intuition. It's a shift from being the person who does the math to being the person who directs the math.
So what does this mean for the rest of us? It means the line between human and machine intelligence is blurring in ways that are both exciting and a little unsettling. But if this collaboration is any indication, the future of discovery might not be about humans versus AI—it might be about humans and AI, working together, solving problems we didn't even know how to approach.
Key Points:
- GPT-6 Astra, working with three human researchers, solved a 2017 open problem in approval-based committee elections.
- The AI proved that a fair committee always exists, contrary to the original search for a counterexample.
- It also invented a new voting rule based on harmonic entropy with a polynomial-time algorithm.
- This marks the first time AI has solved a "major breakthrough" level math problem, signaling a shift from problem-solving to pattern discovery.
- Epoch AI added a "Human + AI" tag to its leaderboard, acknowledging the growing role of collaboration in research.