Skip to main content

Japanese AI Startup Sakana Unveils Fugu Cyber, Outshining GPT and Claude in Security Tests

The cybersecurity landscape just got a new contender. On July 21, Japanese AI startup Sakana AI released Fugu Cyber, a multi-agent system designed specifically for modern network defense. And it's already making waves: in the industry's toughest security benchmarks, Fugu Cyber outperformed leading closed-source models from OpenAI and Anthropic.

Image

What Makes Fugu Cyber Special?

Fugu Cyber isn't just another large language model. It's a multi-agent system, meaning it coordinates multiple AI agents to tackle complex tasks. This approach allows it to handle the messy, real-world scenarios that cybersecurity professionals face daily. Instead of relying on a single monolithic model, Fugu Cyber uses specialized agents that can plan, execute, and adapt—much like a human team.

Benchmark Results That Turn Heads

In the CyberGym test, which simulates realistic cyber attack and defense scenarios, Fugu Cyber achieved an 86.9% success rate. That's a significant leap over OpenAI's GPT-5.5-Cyber and Anthropic's Claude Mythos-Preview. On the CTI-REALM test, which evaluates cyber threat intelligence capabilities, it scored 72.1%, again surpassing its competitors.

These numbers aren't just academic. They suggest that a specialized, multi-agent approach can outperform general-purpose models in specific domains. For cybersecurity, where precision and adaptability are critical, this could be a game-changer.

Implications for the Industry

Sakana AI's success with Fugu Cyber highlights a growing trend: the move from monolithic AI models to more modular, specialized systems. Instead of trying to build one model that does everything, companies are increasingly creating ecosystems of models that work together. This approach can be more efficient, more adaptable, and potentially more secure.

For cybersecurity teams, Fugu Cyber offers a glimpse of a future where AI doesn't just assist but actively leads in threat detection and response. The model's ability to outperform closed-source giants suggests that open-source or startup-driven innovation can still compete—and win—in high-stakes fields.

Key Points

  • Fugu Cyber is a multi-agent system for cybersecurity, not a single model.
  • It scored 86.9% on CyberGym and 72.1% on CTI-REALM, beating GPT-5.5-Cyber and Claude.
  • The system uses multiple specialized agents to handle complex tasks.
  • This marks a shift toward modular, specialized AI over monolithic models.
  • Sakana AI, a Japanese startup, shows that smaller players can outperform tech giants in niche areas.