Grok 4.7: SpaceXAI's Cheaper, Smarter AI for Long Tasks
Grok 4.7 Arrives: Faster, Cheaper, and Ready for Marathon Tasks
SpaceXAI unveiled its latest flagship model, Grok 4.7, early today. The company isn't shy about its ambitions: the landing page touts it as "the strongest model for programming and knowledge work," claiming it's twice as fast as comparable models and priced at half the cost. This isn't just about topping benchmarks—Grok 4.7 packs a larger base model, beefed-up long-term task handling, self-checking abilities, and better long-context management. SpaceXAI has also put it through its paces in specialized fields like documentation, presentations, law, healthcare, and engineering, aiming to cut down on missteps and deliver results you can actually use after hours of work.
Elon Musk chimed in on Twitter, saying Grok 4.7 strikes a "very competitive balance" between intelligence, speed, and cost.

Long Tasks: The New Battleground
So what's under the hood? Grok 4.7 uses a bigger base model than its predecessor, Grok 4.6, and extends the reinforcement learning phase during training. That means the model grapples with tougher task combinations—some samples take hours to complete. According to SpaceXAI, the new model excels at checking its own work and managing long contexts, directly tackling the headaches of programming agents: complex software tasks involve reading repositories, breaking down requirements, editing multiple files, running tests, and endless debugging. Staying on track and catching errors over long execution chains often matters more than writing flawless code in one go.

In the CursorBench 4.0 test for long-term coding, Grok 4.7 scored 46.3%, up from 40.4% for the previous generation. Its DeepSWE v1.1 high-reasoning score hit 71.0%, and Terminal-Bench 4.0 jumped from 20.3% to 38.0%. The model was also trained to natively understand the Grok Bot operating framework, which SpaceXAI says boosts performance in dialogue and general knowledge work. The takeaway? The upgrade focus has shifted from single-turn answers to continuous collaboration between the model, tools, and the execution environment.
Beyond Code: A Full Knowledge Worker
Programming remains Grok 4.7's headline act, but its evaluation scope is far broader. On AA Briefcase v1.1, it scored 1657 points (up from 1546), and its GDPval Elo rose from 1605 to 1695—close to Fable5.1's 1735 and ahead of GPT-6Astra's 1542. Professional fields saw gains too: EEBench climbed from 53.0% to 64.0%, Harvey Legal Agent Benchmark from 15.8% to 19.6%, and HealthBench Professional from 48.5% to 56.7%. The trend is clear: cutting-edge models are now competing on "completing the entire task," which means understanding the request, calling tools, producing files, and self-reviewing.
On the security front, Grok 4.7 introduces a new protection system. SpaceXAI claims it's the best yet at refusing harmful answers and resisting jailbreaks. The LatchBio biosafety benchmark reached 62.4%, and HackerBench v0.3 allowed only 3.3% high-risk dual-use prompts—while rarely blocking legitimate security research. SpaceXAI has also opened invitation-based red team capabilities to select cybersecurity partners.
Key Points
- Grok 4.7 is SpaceXAI's strongest model, targeting programming and knowledge work with twice the speed and half the price of competitors.
- Long tasks are a major focus: enhanced self-checking and long-context management for hours-long workflows.
- Benchmark gains across coding (CursorBench 4.0: 46.3%), legal (Harvey: 19.6%), and healthcare (HealthBench: 56.7%).
- Security upgrades include a new protection system and red team access for partners.
- Elon Musk calls it a "very competitive balance" of intelligence, speed, and cost.