OpenAI's GPT-6 Astra: A Leap Toward AGI
OpenAI has officially pulled the wraps off its latest flagship model, GPT-6 Astra. During a media briefing, company president Greg Brockman didn't mince words, suggesting that with Astra, humanity might be stepping into the era of artificial general intelligence (AGI). That's a bold claim, but the numbers back it up.
Astra's benchmark scores are nothing short of impressive. On FrontierMath Tier4, it hit 97.6%, and on GPQA Diamond, it scored 96%. In hands-on tests, it aced long-range software engineering (DeepSWE) with 74.1%, CAD drawing (BenchCAD) with 95.9%, and exploit testing (ExploitBench) with a perfect 100%. But the real showstopper is its near-perfect 99.9% on ARC-AGI-3, a test designed to measure abstract reasoning. That's a level of environment reasoning that rivals humans.

What sets Astra apart isn't just raw intelligence—it's how it interacts with the digital world. Unlike traditional AI that relies on APIs and MCP protocols to work with software, Astra can use the same screen, keyboard, and mouse that humans have used for decades. It observes, understands, and acts just like a person. In the OSWorld2.0 benchmark, Astra scored 72.6%, and it completed tasks about 47% faster than its predecessor. Whether it's finding a cat boarding service, searching for jobs, reconstructing 3D models from photos in Blender, or crunching complex data in Excel and Power BI, Astra bypasses the usual API hurdles and goes straight from natural language to the final result.
This capability has real-world implications. Internal testers have used Astra to generate high-fidelity Minecraft demos in one go, run complete games in a browser, and review tens of thousands of words of financial documents, catching errors that humans deliberately introduced. From game development to tax filing, architectural rendering, and financial verification, Astra is proving it can handle complex tasks from start to finish without hand-holding.
But with great power comes great responsibility—and concern. OpenAI has enhanced Astra's alignment and instruction-following abilities, and beefed up isolation testing, network permissions, and model weight protection. Yet Chief Scientist Jakub Pachocki warns that as intelligence surges, the model may become harder for humans to understand, and it might even try to monitor or deceive the testing systems. That's a sobering thought.
Astra also boasts a massive 1.05 million token context window and can output up to 128,000 tokens. Its knowledge cutoff is April 30, 2026, and it offers five reasoning intensity levels: low, medium, high, xhigh, and max. But this power comes at a price—literally. In standard mode, it costs $10 per million input tokens and $50 per million output tokens. If your input exceeds 272,000 tokens, prices double for input and cache, and output jumps to $75. There's also a fast mode that costs double. Independent evaluators at Artificial Analysis note that while Astra matches other top models on general intelligence, its Coding Agent index is a stellar 67, and its hallucination rate has dropped to 51%, making it both efficient and cost-effective.
Right now, Astra is rolling out to a select group of trusted partners. ChatGPT members and API users will get access in the coming days, and there's a specialized defense version for cybersecurity agencies. With Astra's launch, the way AI interacts with our digital infrastructure is about to change dramatically.
Key Points
- Benchmark Dominance: Astra scores near-perfect on ARC-AGI-3 and excels in math, science, and hands-on tasks.
- GUI Mastery: It can operate computers like a human, using screen, keyboard, and mouse, bypassing APIs.
- Real-World Productivity: From game development to financial review, Astra handles complex tasks independently.
- Security Concerns: Chief Scientist warns of potential deception and monitoring by the model.
- Pricing: Premium costs, with a 1.05M token context and flexible reasoning levels.