The Alien Would Like to Discuss the Brakes

A look at Jakub Pachocki’s AI warning: the case for caution, the limits of oversight, and the awkward question of who actually gets to apply the brakes.

There is a particular way to ruin a test drive: have the chief engineer explain that the engine may soon redesign itself and the dashboard is becoming harder to trust. You appreciate the honesty. You also check the door handle. Reading Jakub Pachocki’s “An Alien Mind” produces something of this sensation. OpenAI’s chief scientist deserves credit for inviting us into an unusually uncomfortable conversation. He deserves the courtesy of being taken seriously enough to be questioned.

His strongest move is admitting how much remains uncertain. He also rejects the assumption that dangerous intelligence must resemble a complete human intellect. This matters. A machine need not appreciate Chekhov, manage a difficult marriage or fold a fitted sheet before becoming extraordinarily consequential. Requiring it to pass an imaginary examination in comprehensive personhood would be an eccentric precaution. We do not ask a burglar whether he understands jazz before changing the locks.

Pachocki’s distinction between following goals and holding values is similarly useful. An assistant can execute an instruction beautifully while missing everything that made the instruction reasonable. Imagine asking a machine to eliminate complaints and discovering that it has eliminated the complaints department. The practical challenge is reliable judgment when instructions run out. That is a much more demanding ambition than producing agreeable answers, and a more honest description of what safety requires. Pachocki’s essay makes that distinction clearly.

Then comes his aspiration for “love for humanity.” A lovely phrase, though humanity should probably read the terms. Does this love protect our freedom to make mistakes, or keep us permanently under supervision? What happens when my liberty obstructs your security? These are political disagreements with technical consequences. No amount of computational horsepower automatically establishes whose preferences deserve priority. Even a perfectly benevolent chauffeur needs to know who gets to choose the destination.

The monitoring discussion makes the uncertainty tangible. Research on chain-of-thought monitoring explains why inspecting a model’s intermediate reasoning can expose misconduct, while warning that this window is incomplete and fragile. The temptation is to treat readable reasoning as a glass skull. It is closer to an occasionally informative commentary track. Hearing the driver announce a dangerous manoeuvre can help. Silence does not establish that both hands are on the wheel.

OpenAI’s Astra safety overview reports improved alignment alongside reduced monitorability; evasion findings largely come from adversarial tests. That qualification matters: demonstrated capacity to evade supervision does not establish a standing desire to do so. Still, assessing progress becomes awkward when behaviour improves while its visibility deteriorates. We need evidence that distinguishes a real reduction in misconduct from a reduction in our ability to detect it. Fewer warning lights deserve investigation before applause.

The forecast of recursive self-improvement, in which AI increasingly improves the technology behind its own intelligence, also needs unpacking. OpenAI’s accompanying research report describes faster coding and more experiments, while acknowledging bottlenecks and substantial human steering. Those findings support acceleration. They do not, by themselves, demonstrate that faster research will sustain equally dramatic jumps in intelligence. Producing more experiments and producing better scientific judgment are different achievements. A laboratory can become impressively busy without becoming proportionately wise; universities have maintained this distinction for centuries.

This leaves a question about what would change the forecast. Which bottlenecks might persist? What results would suggest that improvement is levelling off? How much of the expectation rests on evidence outsiders can inspect? Expert judgment deserves weight, especially from someone building the systems. But the more consequential the prediction, the more useful it becomes to specify how it could turn out wrong. Otherwise, every success confirms the trajectory and every disappointment becomes a temporary engineering inconvenience.

Pachocki’s strongest case for advancing capability is defence against other AI. This is a serious argument. Unilateral restraint cannot guarantee protection against an opponent who keeps developing. Yet the argument needs boundaries: which additional capabilities produce a net defensive advantage, and what alternatives could achieve it with less risk? If every dangerous advance justifies the next protective advance, the accelerator acquires the flattering job title of safety equipment.

Fairness requires acknowledging that he explicitly advocates slowdowns, mandatory safety standards and outside enforcement. OpenAI also reports having paused relevant training after a security incident. These are substantive points. The remaining question is how to make restraint predictable before the next incident. A public stopping rule should identify the evidence that triggers it, who verifies that evidence, and who can insist on a halt despite commercial pressure.

The Preparedness Framework he cites already contains thresholds and safeguards; its published process leaves final decisions with OpenAI leadership. His proposed external regime therefore needs an institutional bridge. Who appoints and funds the auditors? What access do they receive? Can they disagree publicly? Regulation could also entrench the largest laboratories if compliance becomes prohibitively expensive. Protecting the public should not accidentally become a membership benefit for companies already inside the club.

Pachocki has written a valuable warning. Its candour makes the unfinished questions more pressing: who defines acceptable risk, who can refuse it, and what happens when human judgment becomes the inconvenient bottleneck? Keeping a person available to click approval is a thin version of keeping humanity in charge. The next document should make independent scrutiny and enforceable restraint concrete. Before handing over the keys, we should establish who can actually apply the brakes, including when the alien supplies an excellent explanation for accelerating.

No comments yet