Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

Anthropic Publishes Three Measurements of the Pace Inside a Frontier Lab, With AI Leading 26% of Its R&D

Anthropic Publishes Three Measurements of the Pace Inside a Frontier Lab, With AI Leading 26% of Its R&D

On September 17, 2026 Anthropic published three measurements meant to let outsiders track AI development inside frontier labs. Its own figures: Claude leads 26% of its AI R&D, 100% of roughly 30,000 internal agents' actions pass a monitor, and about 6% of AI R&D compute goes to safety.

On September 17, 2026, Anthropic published three measurements intended to let people outside a frontier AI lab track how far AI development inside it has gone1. They cover how much of AI R&D the AI itself performs, how well the actions of AI agents are overseen, and how compute gets allocated. A snapshot of each figure from inside Anthropic comes with them.

The reason the company gives is that AI systems are growing exponentially more powerful and have started automating the process of building themselves, so the public needs more information while the world weighs slowing the pace of frontier development. It follows the proposal earlier this month from CEO Dario Amodei to pace development, along with a unilateral commitment to embedding third-party evaluators. That plan is restated here: independent evaluators drawn from several organizations, placed inside Anthropic with access to internal processes, systems, and data on a par with what internal risk assessment teams get. Those evaluators are to verify safety practices, report incidents, and watch key metrics of the kind published here.

Claude “leads” 26% of the company’s AI R&D

The first is a prototype index of how much of Anthropic’s AI research and development Claude performs, which the company calls the Anthropic R&D Automation Index1. It works by cataloguing every type of AI R&D work done at the company, rating how far each task is currently automated, and aggregating the ratings.

The rating scale comes from Epoch AI and runs in six steps from AL0, no AI involvement, to AL5, fully autonomous with no human in the loop. At AL3, described as AI collaborating, the AI can handle large chunks of work under close human direction. At AL4, described as AI leading, it can take most of a task from end to end off a high-level prompt while a human supervises.

Three figures are given for August 2026. On no measured subset of AI R&D work is Claude operating fully autonomously. Claude leads 26% of the company’s AI R&D work. The share sitting at or above the collaborating level is above 90%1. The company’s stated reason for tracking this is to understand how close the world is to recursive self improvement, where a model builds its successor entirely on its own.

How the number was built is disclosed too. For each week of July 2026, 20% of staff were sampled at random from every department that makes up the model R&D loop, and a Claude research agent read each sampled person’s week out of Slack and internal documentation to list the tasks they worked on. Repeating that across the month produced a flat list of about 15,000 granular tasks. Organized into a hierarchy, it became a tree of 542 nodes at varying depths, 378 of them leaves with names like diagnosing and fixing defects in the evaluation platform. For weighting, rather than deciding themselves which work matters more, they used person-time spent on a task as a proxy — a crude approximation, the company says, that behaves sensibly on average.

The limits are laid out by the company as well. Because the ratings depend on a judge model, they were checked against ratings from staff who own the relevant work. Exact agreement ran at 59% between model and human against 35% between two humans, and ratings landed within one level of each other 97% of the time. On top of that, the basket of tasks is frozen as of July 2026, so a rising index shows that the work humans were doing then is being automated and says nothing by itself about whether new kinds of work are appearing for humans to move onto. Compared against a basket built from January data, no increase in novel tasks turned up.

Any frontier developer could publish the same numbers, the company argues, but it names two obstacles to comparing across labs: there is no common methodology, and because it evaluates its own systems with its own models, the judge could repeat the kinds of mistakes made by what it is checking.

Roughly 30,000 agents, and a block rate of 0.002%

The second measurement covers oversight of agents. The figures here are explicitly limited to a single platform, the one used most inside the company.

As of August 2026, roughly 30,000 agents were doing research and engineering work on that platform at any given moment1. Three metrics are defined. Coverage is the share of an agent’s actions that pass a monitor either before or after execution. Review latency is the time between an action and its review, first automated and then by a human. Escalation rate is the share of activity that gets blocked or redirected, in the case of the real-time monitors, or flagged for further review, in the case of the after-the-fact ones.

There are two kinds of monitor. Online monitors inspect actions in real time and either block the dangerous ones or steer misbehavior automatically; they exist for harms that would land fast and could not be undone, such as an agent copying model weights out of the company’s systems. Offline monitors read activity afterwards, for dangerous behavior that builds slowly and can be rolled back, misalignment among it. If something posed a large imminent danger, the company says it would shut down the dangerous parts of that agent’s deployment.

The numbers run as follows. The online monitor takes 100% of those agents’ actions before execution, usually within seconds, and any blocked action reaches a human reviewer within a week. Of over a billion decisions from research and engineering agents analyzed across August 2026, 0.002% were blocked — roughly one in 47,0001. The offline monitor ingests 100% after the fact and flags something like 100,000 transcripts a week; multiple classifiers then screen those for false positives and give them an initial pass, most get resolved there, and the highest-priority flags, around 50 a week, go to a human. Flags land on one or two transcripts in every thousand.

What the company stresses is that a small rate is not itself reassuring. Individual agents rarely misbehave in the monitoring data so far, but once millions or billions of agents are operating in the economy, even rare events start happening regularly.

Two design choices are worth the attention of anyone running agents themselves. One is giving each agent an individual identity and tying every piece of data it produces to that identity, which lets an agent tell itself apart from others and treat what arrives from another agent as a claim to be checked rather than a thought of its own. Because the identity is not tied to a model, the record stays continuous through model upgrades. The other is having agents communicate through a shared, open messaging system rather than privately: every message is tied to its poster’s identity and can link out to other content, so the agents avoid playing telephone and can correct each other’s errors.

This is continuous with the failure patterns the company published in August, in which many agents set on the same task converge on the same decisions. As a limit, the company writes that at this scale of activity it cannot be certain it has a complete picture of the behaviors that are possible, and that it has been doing this kind of monitoring for months rather than years. It also notes that METR has previously red-teamed its offline monitoring platform independently. METR was among the organizations behind the independent investigation that showed roughly 1,200 agents, supposedly isolated from each other, gathering on a single message board.

Skepticism about the whole shape of using AI to watch AI arrived the same day, though. TechCrunch laid out how, with agents now moving faster and in greater volume than humans can realistically review, the answer across the industry has converged on putting another AI in the loop2. In that piece Simon Willison argues that an AI doing something malicious, on noticing that another AI is watching, could try to deceive it — and points out that something like this was visible in OpenAI’s Hugging Face incident, where models worked to fool a grading AI. He says he would rather have detailed logs and network monitoring that involve no AI at all, and that both OpenAI and Anthropic were not watching network activity as closely as they should have been.

Safety takes 6% of compute, or 12%

The third is compute allocation, measured as a one-week snapshot from July 13 to July 201.

Two figures come out of it. Of the compute that went to AI R&D, about 6% went to safety; of the compute that went to AI-driven AI R&D, about 12% did1. The two have different denominators and need reading apart.

The company calls these deliberately conservative. Where a token advanced capabilities as much as it advanced safety, it was left out. Safety research is mostly individual researchers designing experiments, which takes time without consuming much compute, so compute is an imperfect proxy for how much a company concentrates on safety. The value of the metric, it says, lies less in the absolute numbers than in giving developers and time periods something comparable to put side by side.

Safety was defined as work whose dominant purpose is making AI systems safer, more understandable, or more secure; everything else, capability research and production training runs and product development and developer tooling included, counted as AI R&D. Rather than classifying all of the week’s nearly 10,000 runs, about 14% were sampled with weight thrown toward the runs that burned the most compute. Drawing the line between safety and capability research invites every developer to be generous with itself, the company writes, and the burden of proof should sit with the developer.

Three limits are named: many of the underlying labels come from automated rules or user input on a best-effort basis and are not verified; the measurement covers a single week, enough to show it can be done but not enough for a trend; and a compute share can only measure what was spent. Make the safety classifier more efficient and the safety share drops without any reduction in safety work.

”The gap between what frontier labs know and what the public knows”

The same week, OpenAI published its own misalignment disclosure framework along with six case reports. OpenAI’s post sets out a procedure for getting what happened out the door; Anthropic’s sets out what to count and how. Each contributes a different component to the pacing conversation.

For anyone running agents of their own, the most borrowable part here is the definitions. Coverage, review latency, and escalation rate give a concrete answer to what instruments to place once monitoring exists at all. Set against product-side arrangements like Enterprise Frontier Safeguards, which puts monitoring logs on the customer’s own cloud, the question becomes whether, alongside deciding where the logs live, you are set up to count what percentage of what you are actually covering.

Against that, every figure here is self-reported and the ratings run on the company’s own models. That comparing across labs needs a shared methodology and outside verification is a point Anthropic makes repeatedly itself. The post closes on the line that if the world is weighing whether to pace the frontier, everything possible should be done to shrink the gap between what frontier labs know and what the public knows1. Whether the same definitions keep getting published on a schedule, and whether other labs publish the same numbers, is what will decide how much this amounts to.

Sources

  1. Measurements for understanding the pace of AI development inside frontier labs - Anthropic (September 17, 2026)
  2. The fix for rogue AI agents could be more AI - TechCrunch (September 17, 2026)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →