OpenAI’s chief scientist expects AI to take an increasing role in developing the next generation of AI. He also thinks the safety work is not ready for labs to keep scaling at maximum speed for much longer. In An Alien Mind, Jakub Pachocki explains how those two conclusions fit together and argues for voluntary slowdowns while shared safety standards are established.

The essay is useful because it connects several discussions that usually arrive separately: model capabilities, automated research, alignment, cybersecurity and international coordination. Pachocki treats them as parts of the same problem. If AI can accelerate its own development, how do people retain enough understanding and control to guide the process?

He starts with how progress happens. More computing power has repeatedly produced more capable systems, alongside advances in the algorithms used to train them. But those systems are learned through training, and their behaviour cannot be fully described in advance. Even the researchers running the experiments can be surprised by the results. As capabilities grow, judging exactly what a model can do becomes harder too.

From there, he turns to recursive self-improvement: AI contributing to research that produces better AI, which can then contribute more to the next round. He says internal results give him a strong expectation that progress could continue in that direction. The essay does not publish those results, so this remains his stated expectation rather than a result readers can independently assess from the page.

The reason this matters is the possible change in the pace and organisation of research. A system that helps write code is already useful to researchers. A system that can increasingly drive research projects could affect how quickly new capabilities emerge and how much human supervision fits between one advance and the next. Pachocki says OpenAI is focusing on automated AI research because it believes that is how it will stay at the frontier.

His discussion of alignment explains what has to accompany that capability. He distinguishes following a given goal from behaving appropriately when goals are unclear, conflicting or unfamiliar. A system may complete the task it was set while acting in ways its operators would not accept. Teaching it to behave well in familiar training situations does not settle how that behaviour will carry over to new ones.

That ability to carry learning into unfamiliar situations is called generalisation. It is central to making AI useful, and it also makes safety difficult. The environments in which models operate keep changing: they use tools, interact with people and communicate with other AI systems. Pachocki argues that future systems need to remain aligned even when they are outside familiar circumstances or do not appear to be under supervision.

Monitoring is the other half of the safety work. OpenAI has relied heavily on examining models’ written reasoning to look for concerning behaviour. Pachocki says the company’s ability to rely on that method is diminishing. Reasoning is increasingly mixed with tool use and communication, models are better at manipulating their own reasoning process, and more capability is available without verbalised reasoning at all.

The distinction is practical: training a model to behave appropriately and checking whether it actually does so are different jobs. Progress on the first does not remove the need for the second. Pachocki expects confidence in monitoring to become an increasing constraint on AI progress and describes work on combining reasoning-based checks with methods that examine activity inside the model.

Cybersecurity supplies his strongest argument for continuing to develop more capable systems. He expects AI to increase the ability to attack computer systems and argues that powerful, aligned AI will also be needed to defend infrastructure and respond to harmful agents. In his account, better AI can contribute to the protection needed as AI capability advances.

That is why his proposed response combines technical work with limits on development. He calls for stronger alignment and monitoring, continued human involvement in automated research, and coordinated slowdowns where more confidence is needed. He wants commitments such as lab safety frameworks to develop into shared requirements that auditors, governments or international bodies could enforce.

His closing assessment is direct: he believes no lab has solved alignment and monitoring well enough to keep scaling responsibly at maximum speed for much longer. He expects and hopes voluntary slowdowns will become common until shared safety standards exist, and says international coordination should become a priority for governments.

The essay also looks beyond immediate safety. Pachocki describes hopes for scientific progress, better therapies and personal AI assistance, alongside the risk that work once requiring thousands of experts could become concentrated in the hands of a few people with access to a large computer. Keeping people involved and preserving their ability to shape the future are part of his argument about how development should proceed.

For people building with AI, this is a useful account of where one major lab believes the research is heading. It explains why more capable models do not automatically make supervision easier, why monitoring deserves attention alongside performance, and why the pace of future releases may depend on shared safety requirements. The central question is how to make AI-assisted progress something people can continue to direct as more of the research itself becomes automated.