Updated 9 hours ago
Claude leads 26% of Anthropic's AI research. What does that actually measure?

Claude leads 26% of Anthropic's AI research. What does that actually measure?

Anthropic says Claude leads 26% of its measured model‑development work. Its task weighting, human review, monitoring boundaries and safety‑compute denominators explain what that figure means.

Anthropic says Claude could lead 26% of the work involved in developing its future AI models as of August. The figure, [published on September 17](https://www.anthropic.com/institute/measuring‑pace‑of‑ai‑development), sounds like a measure of a lab increasingly building itself. It is narrower, and more revealing, than that. “Leads” means that Claude can do most of a defined task from a high‑level human instruction while a person supervises the result. Anthropic reported no measured category in which Claude operated fully autonomously. The number describes how one company classified its own work, not how much of an entire model was designed without people. That distinction matters because the same paper proposes a way for outsiders to track AI development inside frontier labs. It gives a view of the work basket, the monitoring of internal agents and the computing capacity spent on safety. Those measurements could become useful public accountability tools, but only if readers can tell what each denominator contains, where human judgment enters, and which parts another organization could verify. [Reuters' initial report](https://www.marketscreener.com/news/anthropic‑says‑claude‑now‑leads‑a‑quarter‑of‑work‑building‑its‑next‑ai‑models‑ce785bd3d18efe2c) and [Associated Press coverage](https://apnews.com/article/anthropic‑claude‑ai‑model‑self‑improvement‑4d3a7430f57cbc7c39e1c5f2b7d7e132) established the news; the methods explain what the headline can support.

A task can be led by AI while a human still decides whether it ships

Anthropic uses a six‑level scale, adapted from [Epoch AI's automation framework](https://epochai.substack.com/p/toward‑an‑onet‑for‑ai‑r‑and‑d), to describe individual kinds of research work. At AL3, Claude collaborates: it handles substantial parts of a task under close human direction. At AL4, it leads: a person states the problem, then the agent carries out most of the task and brings the result back for review. AL5 would mean the agent notices, scopes, completes and deploys the work without needing a human in the loop. Anthropic reports Claude at AL4 for 26% of its measured R&D basket, at AL3 or higher for more than 90%, and at AL5 for none of the measured categories. The paper's own example is a failed overnight data pipeline. At AL4, Claude could inspect logs, find the broken stage, write and test a fix, handle surprises and document the result. An engineer would still decide whether to release it. That is a meaningful change in who performs the investigation, but it is not a claim that Claude chooses the lab's research direction or deploys changes on its own. [Anthropic's example and footnotes](https://www.anthropic.com/institute/measuring‑pace‑of‑ai‑development) make the human handoff part of the definition, not an incidental caveat. The change over time also needs careful wording. Anthropic's figure rose from less than 1% in February to 26% in August. That is a rise of more than 25 percentage points in its index. Calling it a 26‑fold increase would imply a known February baseline when the company only supplies an upper bound; a zero starting value would make a ratio meaningless. Nor does the increase alone tell us whether model R&D became 26% faster. An automation rating concerns who does a type of task, while speed depends on the time saved, review burden, failed experiments and other bottlenecks.

The 26% comes from a constructed basket of work, not a task counter

In the [methodological appendix](https://www.anthropic.com/institute/measuring‑pace‑of‑ai‑development), Anthropic says it sampled 20% of staff each week in July across departments involved in model development. A Claude research agent used work records, including internal messages and documents, to list roughly 15,000 granular tasks. Another Claude‑assisted process organized them into a tree of 542 nodes, including 378 leaf categories. A model then investigated how each kind of work was done, and an independent Claude judge assigned an automation level. The 26% is an aggregate across this categorized and weighted basket; it is not the fraction of 15,000 task tickets closed by an agent. Anthropic weighted categories using sampled person‑time. Each person contributed one unit per week, divided evenly across the tasks recorded for that person. Someone associated with four tasks would contribute 0.25 to each; someone associated with ten would contribute 0.10. This avoids treating a tiny routine activity as equal to a large workstream just because each has one label, but it is still a proxy. It assumes the sampled task list and the equal division of a person's time approximate the actual importance and effort of the work. The work tree is also frozen for comparison, so a rising index describes automation of that basket and may miss a shift toward new kinds of human work. The company did ask staff to rate relevant work areas without seeing the model's evidence. It reports 59% exact agreement between its judge and human raters, compared with 35% exact agreement between human raters, and 97% model‑human agreement within one automation level. Those figures show why the AL3‑to‑AL4 boundary is debatable; they do not independently validate every work record or the resulting company‑wide percentage. Anthropic itself identifies the problem of using its own models to judge its own systems and proposes outside verification. Until that exists, the 26% should be read as a disclosed internal estimate with an unusually explicit method, not an audited universal measure.

Thirty thousand agents are monitored, but the coverage has a boundary

Anthropic separately says that roughly 30,000 agents were doing research and engineering work at a given time in August on its most‑used internal platform. The [published monitoring description](https://www.anthropic.com/institute/measuring‑pace‑of‑ai‑development) says every action on that platform passes through an online monitor before execution, and every action is later ingested by an offline monitor. That is 100% coverage of the specified platform's actions; the disclosure does not establish that every agent system throughout the company is covered in the same way. The online monitor blocked about 0.002% of more than a billion decisions in August, which Anthropic expresses as roughly one in 47,000. Its offline process flags about 100,000 transcripts a week for additional classifier review, with about 50 high‑priority cases reaching people. Dividing 50 by 100,000 gives about 0.05% of flagged transcripts escalated to human review. It would be a mistake to compare that directly with the online block rate: one denominator is individual decisions in a month, the other is transcripts flagged in a week. Neither figure by itself is a measured rate of harmful agent behavior, since a block may prevent a problem and a flag may be a false positive. For a future cross‑lab comparison, the useful questions are more specific than “does the company monitor its agents?” Which actions can the monitor stop before they run? How quickly are consequential flags reviewed? What proportion of agent activity lies outside the reported platform? Can independent evaluators test whether known misbehavior is caught? Anthropic says it plans to embed outside evaluators and that its current measurements would need greater completeness and verification. Those are concrete tests for a proposed accountability system, not proof that the system already catches every important failure.

Safety‑compute figures are a one‑week allocation, not a safety score

Anthropic's third measurement concerns the computing capacity it used between July 13 and July 20. It classed about 6% of AI R&D compute as safety work and about 12% of AI‑driven AI R&D compute as safety work. The second group is a subset with a different denominator; neither percentage is a share of all compute used by Anthropic. The company says work that advanced capabilities and safety equally was counted outside the safety category, and that safeguards classifiers were excluded from these figures. These choices make the reported safety share deliberately conservative by its definition.

The appendix also describes a limitation that matters to anyone comparing labs: workload tags are often best‑effort, model classifiers helped sort the runs, and only one week was measured. Even a consistent compute share would not establish the quality or effectiveness of safety research. A more efficient safety method could use fewer tokens while doing more useful work. Anthropic argues for published category definitions and independent checks so future figures can be compared on a like‑for‑like basis. Until then, 6% and 12% are a transparent snapshot of one classification exercise, not a ranking of how safe the lab is.

OpenAI's September 6 account of research acceleration supplies a useful comparison of what cannot yet be compared: it reports 3.1 agent‑workdays of runtime for every human workday in its research organization as of mid‑August. Runtime measures the amount of agent activity; Anthropic's 26% measures the person‑time‑weighted share of task categories judged to be AI‑led. A long‑running agent can generate many runtime hours without leading a larger share of important work. Both companies also note that other bottlenecks limit what these indicators say about overall research progress.

The September disclosure therefore establishes a sharp change in Anthropic's own account of its internal workflow, alongside a proposed way to inspect that change. It does not put a date on fully self‑improving AI. The next informative update would keep the same work basket or explain a revision, show whether any category reaches AL5, disclose monitoring coverage beyond the main platform, and let independent reviewers test the classifications. That would tell readers whether the numbers are becoming a dependable series rather than just a striking first point.

Share this article

PostShare

Related News