In February, an AI led essentially none of the research and engineering that goes into building Anthropic's models. By August, it led 26% of it.

That is the number Anthropic put on the table this week, in a disclosure the Associated Press picked up on September 17 (AP via Local10). The company says its own model, Claude, now "leads" a quarter of its AI R&D work โ€” meaning an engineer hands over a high-level goal and the model does most of the job end to end, with a human setting direction rather than doing the typing.

To run that work, Anthropic says it has roughly 30,000 AI agents operating at the same time on research and engineering tasks (Digital Today). The tool is now a meaningful chunk of the workforce that builds the next version of the tool.

The thesis is simple and a little vertiginous: the loop where AI helps make better AI has stopped being a thought experiment and started showing up in a percentage.

๐Ÿง  Why This Matters

For years, "recursive self-improvement" was the part of the AI conversation you could safely file under speculation. A model that helps design its successor, which helps design its successor, and so on โ€” clean on a whiteboard, hard to point at in the real world.

Anthropic just pointed at it. The jump from under 1% in February to 26% in August is the headline (36Kr), and the slope matters more than the level. Seven months, a 25-point climb. Nobody is claiming the model designs itself. But the share of the work it can carry with light human oversight is growing fast enough that Anthropic felt the need to publish a metric for it.

The company framed the disclosure as a transparency move. In the AP's account, Anthropic argued the industry "should do everything possible to minimize the gap between what frontier labs know and what the public knows" โ€” which is a polite way of saying the people building this would rather you hear the numbers from them than find out later.

"We should do everything possible to minimize the gap between what frontier labs know and what the public knows." โ€” Anthropic, in its disclosure (via the Associated Press)

๐Ÿ“Š Deep Dive

Anthropic scores the automation on an internal ladder it calls AL0 through AL5. AL3 is collaboration โ€” the model does big chunks under close human direction. AL4 is leading โ€” you give a goal, it does most of the end-to-end work. AL5 is fully autonomous, no human in the loop. The 26% figure is the share of work that has reached AL4.

Here is how the picture breaks down:

  • 26% of AI R&D work is now at AL4 ("leading") โ€” up from under 1% in February 2026 (36Kr).
  • Over 90% of the work has reached at least AL3 ("collaboration").
  • 0% has reached AL5 โ€” no fully autonomous, human-free work, by Anthropic's own account.
  • ~30,000 agents run simultaneously on internal research and engineering.
  • To build the number, Anthropic sampled 15,000 fine-grained R&D tasks and sorted them into a 542-node task tree; human and model ratings fully agreed 59% of the time and landed within one level 97% of the time (Digital Today).

That last line is the part worth slowing down on. The 26% is not a hard sensor reading โ€” it is a rating exercise, part model self-assessment, part human check, and the two sides disagree by a full level roughly 3% of the time and by some amount 41% of the time. Anthropic admits there is real "subjective judgment space" between "collaborating" and "leading." The headline number is a considered estimate, not a thermometer.

โš ๏ธ The Catch

Thirty thousand agents making decisions is a lot of decisions to watch. Anthropic says it monitored more than a billion agent decisions in August, with humans reviewing about 50 high-priority cases a week (Digital Today). A billion to fifty is a supervision ratio that leans hard on automated filtering to decide what a human ever sees.

On safety spend, Anthropic reports putting 6% of its total AI R&D compute toward safety research, rising to 12% of the compute specifically used for AI-conducted R&D. Useful figures โ€” and also a reminder that the large majority of the compute is pointed at making the system more capable, not at checking it.

The company itself flagged the direction of travel. As models take on more, the AP noted, it becomes "more challenging for humans to understand or control these systems." That is the vendor talking, not a critic.

As AI leads more of the work, it gets "more challenging for humans to understand or control these systems." โ€” Anthropic (via the Associated Press)

๐ŸŽฏ What Happens Next

Anthropic says it will keep publishing these metrics and will accept independent third-party evaluations of them (Digital Today). If outside auditors get to grade the grader, the 26% becomes a lot more useful as a public benchmark.

The rest of the field is on similar clocks. OpenAI has said it is aiming for an "automated AI researcher" by March 2028, and reports that for every day its humans work, its agents already run roughly 24.8 hours in parallel (36Kr). Google DeepMind is pushing a framework it calls Dream-RSI, focused on automated verification. Three labs, three routes, one destination.

๐Ÿงฉ Bigger Picture

Strip away the ladder and the task trees and you get a straightforward shift in what a frontier lab is. The scarce input used to be brilliant researchers. Increasingly, one of the largest contributors to building the model is the previous model โ€” spun up 30,000 times, working through the night, graded by a mix of itself and a thin layer of humans.

None of this proves a runaway. Progress at AL5 is flat at zero, the numbers are estimates, and a percentage that climbed fast can also plateau. What changed this week is that a company at the frontier stopped describing the feedback loop as a hypothetical and started reporting it as a quarterly figure you can argue about.

The gap between the model that writes the code and the model that is the code just got a lot narrower โ€” and the company doing the narrowing decided you should see the receipts.


Sources