Anthropic Puts AI Self-Improvement on the Dashboard


R&D Automation Index
Anthropic’s internal metric for estimating how much model research and development work is performed by AI systems at different levels of autonomy.
AL4, or “leads”
A level on the automation scale where AI completes most of a task end-to-end from a high-level prompt, while a human remains in a supervisory role.
Recursive self-improvement
The idea that an AI system helps improve or build the next, more capable version of itself, potentially accelerating future progress.
Embedded evaluator
An outside evaluator given access inside an AI lab to examine models, safeguards, alignment work or internal processes.
26% AI-led
Anthropic says Claude now leads 26% of its model R&D work, up from under 1% in February 2026.
Human supervision
“Leads” means Claude can complete most of a task from a high-level prompt while humans still supervise and approve key steps.
30,000 agents
Anthropic says about 30,000 internal agents were active at once on its main research and engineering platform in August.
Anthropic’s claim that Claude now “leads” 26% of its model research and development marks a shift in how frontier AI labs describe self-improvement: less as a theoretical risk than as an internal production metric. The company says that share rose from under 1% in February 2026 to 26% in August, with more than 90% of model R&D now involving Claude at a collaboration level or higher.12
The key caveat: “leads” does not mean Claude is autonomously designing and releasing its successor. Anthropic defines the level as AI completing most of a task end-to-end from a high-level prompt while a human supervises. No measured category has reached full autonomy.26 In practice, Claude may inspect logs, diagnose a failed data pipeline, test a fix and compare outputs, but a human engineer still decides whether to deploy the change.1
That distinction matters for AI engineering and policy readers because Anthropic is trying to turn an unsettled question — how much of the next model is being built by the current one — into something that can be tracked, audited and potentially compared across labs. The company’s R&D Automation Index rates internal work on a scale from no AI involvement to full autonomy, using task categories weighted by staff time.36
Anthropic’s scale draws a line between AI as a helper and AI as the primary executor. At AL3, described as “collaborates,” Claude performs large chunks of work under close human direction. At AL4, described as “leads,” it can take a higher-level prompt and complete most of the task while a human remains in the loop. AL5 would mean fully autonomous work, and Anthropic says none of the measured R&D has reached that level.6
That puts the 26% figure somewhere between ordinary AI-assisted coding and independent machine research. It suggests Claude is no longer merely accelerating isolated subtasks, but is being trusted to carry larger workflows across debugging, experimentation, analysis and engineering.
It does not imply that Claude chooses the overall research agenda, controls deployment or removes human approval from consequential steps.12
The operational boundary is supervision, not involvement. Anthropic’s reported system still depends on humans to approve important decisions, review blocked actions and retain accountability for model behavior. The company’s own framing emphasizes that Claude is not operating fully autonomously in any measured subset of AI R&D work.1
Anthropic built the index by sampling 20% of staff in its model R&D departments each week in July 2026. A Claude research agent reviewed internal materials such as Slack messages and documentation, identified about 15,000 granular tasks, and organized them into 542 categories.17
A separate Claude judge then rated the automation level of each category. That makes the index partly self-assessed: Claude helped evaluate how much Claude was doing.
Anthropic checked those ratings against employees familiar with the work and reported exact agreement between model and human raters 59% of the time, with scores within one level 97% of the time. The same reporting noted that pairs of human employees matched exactly 35% of the time, underscoring both the usefulness and ambiguity of the measurement exercise.17
For policymakers, that methodology is as important as the headline number. The index is not an external audit, and the lab’s own model is part of the measuring apparatus. But the disclosure also creates a concrete target for scrutiny: task definitions, rating rubrics, inter-rater agreement, sampling procedures and the share of work that requires human approval can all be examined more directly than broad claims about AI capability.16
The R&D metric arrived alongside new details about Anthropic’s internal agent oversight. The company says roughly 30,000 agents were working at once on its most-used research and engineering platform in August 2026, with every action passing through an online monitor before execution.13
Out of more than one billion agent decisions that month, the monitor blocked about one in 47,000, or roughly 0.002%. Anthropic also described offline monitoring that flags large numbers of transcripts for later review, with a smaller number escalated to humans each week.16
Those figures show the scale problem facing frontier labs. Human supervision is still central, but humans are not manually approving every low-level action. Instead, oversight is increasingly mediated through automated monitors, sampling, escalation queues and after-the-fact review.
That makes the quality of monitors — and the incentives around what they flag — part of the safety case for AI-accelerated R&D.
The immediate policy significance is that Anthropic is proposing a vocabulary for recursive self-improvement that can be measured over time. If the index is repeated, outside observers could track whether the share of AI-led R&D rises gradually, jumps after a model release, or approaches categories where humans mostly supervise rather than perform the work.
Anthropic also disclosed compute-allocation metrics, saying about 6% of AI R&D compute in a July snapshot went to safety work, rising to about 12% when narrowed to AI-driven AI R&D. The company described compute allocation as a potentially verifiable input that could matter if labs ever coordinate on pacing or safety investment.16
The broader debate is already entering mainstream policy discussion. CNN’s coverage connected Anthropic’s disclosure to concerns that models accelerating their own development could make AI systems harder for humans to understand or control, and to calls for greater transparency about both technical progress and resource allocation.5
Anthropic is explicitly positioning the index as a framework other frontier developers could adopt. Quartz reported that the company framed the R&D Automation Index, agent oversight metrics and compute-allocation disclosures as replicable measurements that could become meaningful across organizations if reported consistently.3
The next step is verification. Anthropic says it plans to embed outside evaluators with access comparable to internal risk teams, and TechCrunch reported that Accenture staff will begin working inside the company to scrutinize models, alignment and safeguards. Anthropic also said more evaluators are expected, while acknowledging that standards for evaluator access and communications do not yet exist.8
That is the path from disclosure to norm: common definitions, repeated reporting, independent access and comparability across labs. Without those pieces, the 26% figure remains a company-reported snapshot built with significant help from the model being measured. With them, it could become an early template for auditing how quickly AI systems are taking over the work of building more capable AI systems.
The central conclusion is not that Claude is independently building its successor. It is that a frontier lab is now quantifying how much of that successor’s development pipeline is AI-led, publishing the number and inviting comparison.
Recursive self-improvement is moving from speculation into an engineering dashboard. Whether that dashboard becomes trustworthy will depend on how much of it outsiders are allowed to inspect.
Comments