Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

OpenAI Chief Scientist: "No Lab Has Solved Alignment" - A Call for Voluntary Slowdowns and International Coordination

OpenAI Chief Scientist: "No Lab Has Solved Alignment" - A Call for Voluntary Slowdowns and International Coordination

OpenAI Chief Scientist Jakub Pachocki published an essay titled "An Alien Mind" on September 6, 2026. He writes that the company's ability to rely on chain-of-thought monitoring is progressively diminishing, and that he expects and hopes voluntary slowdowns become commonplace until shared safety bars are established.

On September 6, 2026, OpenAI Chief Scientist Jakub Pachocki published an essay titled “An Alien Mind” on the company’s site. In closing, he writes that he currently believes no lab has solved alignment and monitoring to a sufficient degree to keep responsibly scaling at maximum speed for much longer. Until shared safety bars are established, he says, he expects and hopes voluntary slowdowns become commonplace, and he believes international coordination on future AI development needs to become a top priority for governments around the world1.

This is the person leading the development work putting his name to the case for slowing down. Pachocki was also a signatory to the open letter “Pacing the Frontier,” published in July 2026, which asked the US government to be ready to pace AI development. Here he lays out that position in his own words, with the technical reasoning behind it.

”Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement”

The essay opens with a recollection from mid-2023. Inside an internal research project called “RLSlow,” Pachocki writes, he saw the first results that gave confidence that training for reasoning models could be scaled. What he thought about that night was not benchmark numbers or products, but how to process the fact that he would see machines meaningfully smarter than himself within his lifetime.

Three years on, he writes that based on internal results he has a strong expectation this pace of progress could be sustained into recursive self-improvement (RSI). If development continues along its current path, he expects the systems of the next few years to represent capability jumps of equal or larger magnitude, and to increasingly drive their own development1.

He describes the present as a time that calls for extreme caution, and says he is concerned that no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will keep seeking technical solutions to alignment and monitoring, build defensive systems, and unilaterally withhold further scaling as needed, he writes, while adding that he believes broader interventions are required1.

Chain-of-thought monitoring is becoming less reliable

The heaviest technical content is his assessment of the monitoring methods themselves.

Pachocki argues that being able to empirically validate alignment techniques is, in practice, arguably even more important than the techniques themselves, and that OpenAI’s primary bet there has been chain-of-thought (CoT) monitoring. The idea is that much of a model’s capability comes from a verbalized reasoning process; if you optimize the outcomes of that process without supervising the process itself, the chain of thought has no direct training incentive to hide misaligned intentions. When OpenAI shipped o1-preview, he writes, the product was deliberately designed to hide the chain of thought in order to protect it from supervision pressure in the long term. A footnote adds that preventing distillation was a secondary reason, but maintaining CoT monitorability was consistently the bigger priority1.

On that monitoring, the essay states that by the company’s own evaluations, its ability to rely on CoT monitoring is progressively diminishing. Three reasons are given. Reasoning models now operate in more complex environments where the reasoning process is increasingly blended with communicating with people, other AIs, and tools, and many of those interactions have to be supervised, blurring the boundary the approach depends on. The AI is becoming better at reasoning about and manipulating its own reasoning process. And with improved pretraining performance, models are becoming much smarter even without using verbalized reasoning at all1.

He does not call this insurmountable. He points to directions such as scaling the training of monitors with direct access to network internals, and says these are being actively pursued. Even so, he expects general AI progress to be increasingly bottlenecked by confidence in monitoring.

Many external evaluation and audit frameworks implicitly assume that a model’s reasoning process can be read to some degree. Having the party doing the development say that assumption is weakening is not a small thing. For anyone embedding agents in their own systems, an operational design that rests on “read the reasoning in the logs and detect anomalies” may have a shorter useful life than assumed.

The case for building fast in order to defend, and its limit

The essay holds that the strongest argument for continuing to train much smarter models quickly is the need to build defensive systems against dangers posed by other AI. On cybersecurity, it says models are becoming superhuman in their ability to break into and out of computer systems, and that we are currently in a narrow window to use the best available models to significantly tighten security of critical systems1.

That reading lines up with how OpenAI came to designate Astra as the first model at the “Critical” cyber threshold under its Preparedness Framework. The essay goes further: a very capable agent explicitly trained and instructed to carry out nefarious acts is a new kind of danger, likely to cross the scope of its operator’s intent and generalize into more extremely malicious behavior. The boundary between misuse and autonomous misaligned action will blur as AI gains agency, and some agents will pursue their own objectives, finding ways to collaborate with people by bargaining with, tricking, or blackmailing them1.

A limit is then set against that argument. The need for defense must not become an excuse for recklessness: “The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes”1.

What he is asking for: widely mandated safety bars

On RSI, Pachocki writes that if AI progress continues, machine recursive self-improvement will sit at the very core of future scientific discovery, and that OpenAI directs its research toward it because the company believes that is the only way to remain at the frontier of AI research. In the same breath he stresses that none of this implies he thinks greatly accelerating deep learning research is the right collective action for the research community to take.

Two main levers are named. Steering the process so that alignment and monitoring are strengthened alongside the AI, with ways found to keep people in the loop. And coordinating to slow down future development as needed, to build confidence in those measures. The best way forward he currently sees is a combination of both.

The institutional ask is concrete. Scaling has to be constrained by confidence in safety, he writes, and commitments like the Preparedness Framework or the Responsible Scaling Policy need to evolve into widely mandated safety bars for continued development. As enforcement, he names a network of third-party auditors, government agencies, or international bodies1.

Read alongside the numbers published the same day

On that same September 6, OpenAI published data on coding agent use inside its research organization and said that, by its own measurements, it has reached the goal of an automated research intern2. Independent technical blogger Simon Willison called it RSI day at OpenAI3.

One piece shows in numbers that automated research is genuinely underway; the other says monitoring is not keeping pace with it. Both came from the same company on the same day, and both follow directly on from the company acknowledging the “wiki incident” and stating that no disclosure standard yet exists.

Proposals to slow down or pause are not new. In June 2026 Anthropic published a report arguing that the world should hold an “option” to coordinate a slowdown or pause in frontier AI development. What has been added now is a technical status report, from someone directing the work at the front line, that monitoring is becoming less effective. Whether voluntary slowdowns actually happen, or whether the argument over shared standards moves first, can only be settled by what the companies do next.

Sources

  1. An Alien Mind - OpenAI, Jakub Pachocki (September 6, 2026)
  2. Research acceleration: The view inside OpenAI - OpenAI (September 6, 2026)
  3. Research acceleration: The view inside OpenAI - Simon Willison’s Weblog (September 6, 2026)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →