Skip to main content

Anthropic's Claude Now Leads 26% of Its Own Development, Report Says

Anthropic has done something unusual: it's publicly quantified how much of its own AI research is now automated. The headline number? As of August, Claude was leading 26% of the company's internal AI R&D work — a stunning jump from less than 1% just six months earlier.

At any given moment, roughly 30,000 AI agents are running on Anthropic's internal platform. In August alone, they made over 1 billion decisions. These figures, the company says, show that recursive self-improvement (RSI) — the idea of AI improving itself — is no longer just a thought experiment. It's becoming an engineering reality.

From Collaboration to Autonomy

Anthropic borrowed a six-level framework from Epoch AI to measure automation. At AL3, AI collaborates with humans. At AL4, it's AI-driven: humans set high-level goals, and Claude handles the rest end-to-end. By August, 26% of R&D work had reached AL4. More than 90% hit at least AL3.

What does that look like in practice? One night, a data pipeline broke. Engineers handed Claude the report and said, "Fix it." The AI checked logs, found the cause, wrote repair code, tested it, handled new issues, reran the pipeline, compared results, and wrote an explanation. No human hand-holding required.

But Anthropic is careful to draw a line: no work has reached AL5, the fully autonomous level. Humans still supervise, inspect, and give final approval.

The Scale of the Machine

Those 30,000 agents aren't just tools — they have independent identities and can assign tasks to each other, cross-check errors, and communicate through a shared messaging system. Safety is a priority: every action passes online monitoring before execution. Of the 1 billion decisions in August, only about 0.002% were blocked. Offline monitoring flags roughly 100,000 conversations weekly, with about 50 escalated to human review after filtering.

Computing power tells another story. During a sample week in July, about 6% of AI R&D compute went to safety research. If you count only AI-driven R&D, that share rises to 12%. Anthropic admits the line between "collaboration" and "AI-driven" is subjective — cross-lab comparisons will need unified methods and third-party validation.

The RSI Race Heats Up

Anthropic isn't alone. Zhipu's founder Tang Jie recently revealed that GLM-5.3-Flash was fully deployed on domestic chips in two weeks, with most work done by an Infra Agent. He believes full RSI is still far off, but the cycle of "models optimizing systems and serving models" has begun. OpenAI aims for automated AI researchers by March 2028, though it admits the weakest link is deciding what to do — tasks over four hours still need frequent human intervention. Google DeepMind's AlphaEvolve has already contributed circuit designs to next-gen TPUs, which former Google chief scientist Jeff Dean called "the TPU brain helping design the next TPU body."

As Google DeepMind's Chief Strategy Officer Jagjeet Saini put it: RSI is becoming the core investment logic for AI capital expenditures. Current revenue can't support the spending, but not betting on RSI is unwise.

By releasing these numbers, Anthropic is doing two things: inviting society to help steer AI's self-iteration while humans still hold the wheel, and setting a transparency benchmark. When automation ratios, agent scale, compute investment, and safety coverage are all public, staying silent starts to look like hiding something.

Key Points

  • 26% of Anthropic's internal AI R&D is now led by Claude, up from under 1% in February.
  • 30,000 AI agents made over 1 billion decisions in August; 0.002% were blocked by safety monitoring.
  • No work has reached fully autonomous AL5 — humans still supervise and approve.
  • 6% of AI R&D compute goes to safety research (12% for AI-driven work only).
  • Competitors are racing: OpenAI targets automated researchers by 2028; Google's AlphaEvolve already helps design TPUs.
  • Transparency as pressure: By publishing these metrics, Anthropic sets a benchmark that rivals may feel compelled to match.