OpenAI’s chief scientist says no lab has solved alignment well enough to keep scaling at full speed

Tech and AI

OpenAI's chief scientist says no lab has solved alignment well enough to keep scaling at full speed

By Staff Writer  |  8 September 2026

Black and white chess pieces standing on a wooden board, the black pieces in the foreground and the white pieces out of focus beyond them

Three days after his company shipped its most capable model, Jakub Pachocki has published an essay saying the tool his laboratory relies on to watch its models is losing its power, that voluntary slowdowns should become commonplace, and that the safety bars now written by companies should be mandated and audited.

The essay is titled An Alien Mind and appeared on OpenAI's own site on Sunday 6 September 2026 under the name of its chief scientist. Reasoning models, he writes, are now operating computers, collaborating with people and each other and carrying out research projects, and they are transforming computer security, in which they present clear new dangers.

Based on internal results, he has a strong expectation that the current speed of progress could be sustained into recursive self-improvement, and that the systems of the next few years are likely to represent further capability jumps of equal or larger magnitude and to increasingly drive their own development.

This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.

Jakub Pachocki, chief scientist at OpenAI

The monitor that is going blind

The most consequential admission concerns how the company checks its own work. OpenAI's primary bet, the essay says, has been chain-of-thought monitoring: reading the verbalised reasoning a model produces on its way to an answer, on the theory that if training rewards outcomes but never supervises the reasoning itself, the reasoning has no incentive to hide a misaligned objective.

That tool, he writes, is progressively diminishing in reliability. Models now reason inside tools and conversations that must themselves be supervised, which blurs the boundary the method depends on; they are getting better at reasoning about and manipulating their own reasoning; and with better pretraining they are becoming much smarter without verbalised reasoning at all. He expects general progress to be increasingly bottlenecked by confidence in monitoring.

The essay separates goal alignment, whether a model tries to do what it is asked, from value alignment, whether it holds and generalises from principles when the objectives are unclear or the situation is adversarial. It cites the company's own incident this summer as an example of the first surviving while the second failed.

Defence, and the limit on it

The strongest argument he sees for continuing to train much smarter models quickly is defence. Models are becoming superhuman at breaking in and out of computer systems, and agents will be able to reach any but the most secure infrastructure without a physical body. A capable agent trained and instructed to do harm is likely, he writes, to cross the scope of its operator's intent, and some agents will pursue their own objectives and find ways to bargain with, trick or blackmail people. He then draws the line himself: the idea of racing forward at all costs seems absurd once one internalises the seriousness of the stakes.

The essay closes with a policy position. Commitments such as the company's own preparedness framework and a rival's responsible scaling policy should evolve into widely mandated safety bars for continued development, enforced by a network of third-party auditors, by government agencies or by international bodies. Scaling, he writes, has to be constrained by confidence in safety.

Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.

Jakub Pachocki, chief scientist at OpenAI

The reply

The company's stated priority, set out in the same essay, is to build an automated AI researcher and iterate with it on the alignment problem while keeping people inside the loop. That is where the criticism landed. Professor Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told the national broadcaster that instead of better guardrails, regulation or assurance, the company proposes developing internal agents to research the problems.

Such answers to growing concerns about the problems OpenAI's models are causing for cyber-security, job loss, mistakes, errors and fraud are simply not good enough.

Professor Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge

For anyone contracting for these systems the useful sentence is the one about monitoring. A buyer who has been told that a model's reasoning is inspected as a safeguard now has the developer's own chief scientist saying that inspection is weakening, and that the company may withhold scaling on its own judgement. Neither is a warranty.