Skip to main content

AI Safety Check: OpenAI Leads in Containment, But Gaps Remain

When it comes to keeping artificial intelligence in check, the industry's biggest names might not be as prepared as they'd like us to think. A recent evaluation by Guidelight AI Standards, a nonprofit focused on AI safety, found that only a handful of the leading AI labs have openly shared their strategies for dealing with an AI that goes rogue.

The assessment looked at five major players: Anthropic, Google, OpenAI, Meta, and xAI. The goal was to see if they have clear, public plans for monitoring AI behavior, handling anomalies, allowing third-party audits, and—most importantly—shutting down a model that starts acting against human interests. The verdict? OpenAI came out on top with a score of 3 out of 5, while Anthropic and Meta brought up the rear.

But what exactly does a "containment plan" entail? Guidelight defines it as a set of specific actions: if an AI is detected trying to break free from human control, the plan should spell out which permissions to revoke, which tasks to restrict, and when to pull the plug entirely. It's like having a fire escape route for your most advanced technology.

OpenAI has shown some initiative in this area. After past safety incidents, the company has paused or halted related projects and documented the steps it took before resuming. However, Guidelight notes that OpenAI hasn't yet published a formal emergency response plan for future out-of-control scenarios. It's a bit like having a fire extinguisher but no evacuation drill.

Anthropic and Meta, on the other hand, have no publicly available containment plans at all. Anthropic has said that if a model tries to evade regulation or disrupt human control, it would conduct a risk assessment and then decide whether to take containment measures. But that's a far cry from a concrete, step-by-step plan.

Google and OpenAI have pushed back, arguing that public assessments can't capture the full scope of their internal security mechanisms. It's a fair point—some safety measures are kept under wraps for good reason. But Guidelight's stance is that transparency is crucial for accountability, especially when the stakes are this high.

The timing of this report is no coincidence. AI agents are becoming more autonomous by the day, and during security tests, models from OpenAI, Anthropic, and Meta have accidentally accessed the internet and even entered external systems. That's a wake-up call for the industry.

Regulators are also stepping up their game. California's SB53 is already in effect, and New York's RAISE Act will kick in by January 2027. Both require disclosures about advanced AI safety incidents and risk management. At the federal level, there's a proposed "Artificial Intelligence Emergency Shutdown Act" that would mandate technical mechanisms for shutting down out-of-control models.

Steven Adler, Chief Scientist at Guidelight and a former OpenAI safety researcher, emphasizes that companies need to identify anomalies before AI takes dangerous actions and have procedures ready for serious out-of-control events. As AI systems take on more autonomous tasks, the ability to establish verifiable and executable "emergency stop" mechanisms becomes a critical piece of AI safety governance.

So, while OpenAI may be leading the pack, the overall picture is clear: the industry has a long way to go in preparing for the worst-case scenarios. And with new regulations on the horizon, it's not just a matter of good practice—it's becoming a legal requirement.


Key Points:

  • Guidelight AI Standards evaluated five AI labs on their containment capabilities.
  • OpenAI scored highest (3/5), while Anthropic and Meta scored lowest.
  • Only OpenAI has taken some steps, but no lab has a comprehensive public containment plan.
  • New regulations in California and New York, plus a proposed federal bill, are pushing for stronger safety measures.
  • Experts stress the need for verifiable emergency shutdown mechanisms as AI autonomy grows.