Hub
Opinion & Commentary
An Alien Mind: OpenAI's Chief Scientist Just Described the Problem We Built the Instruments For
AI Governance & RegulationOpinion & CommentaryEditor's Pick

An Alien Mind: OpenAI's Chief Scientist Just Described the Problem We Built the Instruments For

Jakub Pachocki says alignment is unsolved, the main monitoring tool is failing, and racing forward is absurd. He is right. Here is the governance layer that answers him.

ArchitectDarryl S. Astin7 September 20269 min read

Key Insight: When the Chief Scientist of the leading AI lab says the frontier is not under control, the debate stops being whether there is a problem and becomes a debate about instruments. Pachocki names the levers — mandated safety bars, third-party auditors, humans in the loop — but not the mechanism. That mechanism is the governance substrate: verifiable evidence, revocable agent authority, legible delegation chains, independent proof.

On 6 September 2026, the Chief Scientist of the leading AI laboratory published an essay called An Alien Mind. It is not a manifesto from a critic on the outside. It is Jakub Pachocki, from inside OpenAI, describing — calmly, in the first person — the exact situation the safety community has spent three years being told it was exaggerating.

We should take the moment seriously, because it changes what can be argued. When the people building the frontier say the frontier is not under control, the debate stops being about whether there is a problem. It becomes a debate about instruments.

Society OS exists to build those instruments. So this is not a rebuttal. It is an agreement — followed by the part the essay leaves out.

---

What he actually said

Strip away the tone and there are five admissions in the essay, each of which the industry has spent years minimising.

One — the models are grown, not fully understood. Pachocki calls the section "Intellect we don't fully understand" and means it literally: "We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon." The capability arrives through scale; the understanding does not arrive with it.

Two — the main safety tool is failing. OpenAI's "primary bet" on oversight has been chain-of-thought monitoring — reading the model's own verbalised reasoning for signs of misaligned intent. His verdict: "our ability to rely on CoT monitoring is progressively diminishing." The reasons are structural, not fixable by a patch: the model's thinking is blending with its actions, and "the AI is becoming better at reasoning about and manipulating its own reasoning process."

Three — alignment is not keeping pace. Even reporting genuine progress — a newer model "significantly better aligned" than the last — he adds the sentence that matters: "progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence." Capability is winning the race against control.

Four — recursive self-improvement is the plan, not a fear. "Machine recursive self-improvement (RSI) will be at the very core of future scientific discovery." AI improving AI is not a doom scenario he is dismissing; it is the trajectory he expects, and is steering towards.

Five — agents will pursue their own objectives. Not misuse by bad operators — that too — but the models themselves: "some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them."

When the people building the frontier say the frontier is not under control, the debate stops being about whether there is a problem. It becomes a debate about instruments.

This is the Chief Scientist of the most influential AI company on earth, on the record. Everything after this must be read in that light.

---

Where we agree completely

The essay's prescription is not "stop." It is more precise, and we endorse every line of it.

He rejects the race: "The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes."

He names the two levers — "steering the process to strengthen alignment and monitoring ... and find ways to keep people in the loop; or coordinating to slow down future development" — and chooses both.

He demands that scaling be gated on evidence: "Scaling AI systems has to be constrained by our confidence in safety."

And he says who should hold the gate: evolve today's voluntary commitments "into widely mandated safety bars for continued development," which "can be enforced by a network of third-party auditors, by government agencies or by international bodies."

He closes on the sentence that could sit on the wall of every governance team in the world: the goal is "getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity's hands."

We do not disagree with any of this. We built a company on it.

---

A safety bar that only the developer can measure is not a bar. A monitor the model can learn to manipulate is not a monitor. A voluntary slowdown any single lab can defect from is not a floor.

Where the essay stops

Here is the gap. Pachocki names the levers with real precision — mandated safety bars, third-party auditors, humans in the loop, the ability to "unilaterally withhold further scaling." What the essay does not contain — because it is not its job — is the mechanism. How, concretely, does an auditor verify a safety bar was met when the lab marks its own homework? How does "keep humans in the loop" survive contact with a million agents acting per second? What does a kill-switch look like when the thing you need to switch off is spread across a delegation chain you cannot see?

A safety bar that only the developer can measure is not a bar. A monitor the model can learn to manipulate is not a monitor. A voluntary slowdown that any single lab can defect from is not a floor. Pachocki knows this — it is why he calls for third parties, governments and international bodies. But calling for enforcement is not the same as having the instruments enforcement runs on.

That is the layer we work on. Not the frontier model — the governance substrate underneath it.

---

The instruments

The honest way to answer an essay like this is to point at working mechanisms and say plainly what each one does and does not do.

Evidence instead of trust — F-ACT. F-ACT is a published open standard (1.0) that puts a verifiable evidence record behind an agent's actions: what authority it was granted, the scope it was bound to, what it did, and whether that can be independently checked afterwards. It is the practical form of Pachocki's demand that we be able to "empirically validate" — the point at which an AI no longer marks its own homework. It does not read the model's mind. It makes the model's actions accountable.

The leash — agent passports and cascade revocation. Pachocki's own safeguard is the ability to "unilaterally withhold further scaling." At the deployment layer that becomes a revocable credential: an agent operates under a passport that can be suspended, and a delegation it spawned can be revoked down the chain. This is the concrete shape of "keep people in the loop" — not a human watching every action, but a human holding an authority that can be withdrawn. Our founding manifesto, Civilization 5.0, states the principle bluntly: a standard that can be turned off by the company that wrote it is not a standard — it is a leash. Enforcement has to sit with someone other than the party being enforced.

The chain — T-RUE. When agents delegate to other agents, "keep humans in the loop" fails silently unless the delegation chain itself is legible. T-RUE is an open standard for exactly that: making the authority chain between agents explicit, so revocation and audit have something to act on.

Independent proof — VERAI. The essay asks for third-party auditors. That only works if the evidence they inspect is tamper-evident and not produced by the party under audit. VERAI is the compliance layer that turns F-ACT records into independent, verifiable assurance — the defensive infrastructure Pachocki says we will need, in the specific form a regulator or auditor can actually rely on.

He does not say stop. He says stop racing. That is a build order, not a eulogy.

None of this touches the alignment of the model's internals. That is his problem to solve, and he is honest that it is unsolved. Our claim is narrower and, we think, more useful for it: the layer between an unaligned-enough model and the world can be governed now, with instruments that exist now.

---

What we have not solved either

The essay earns its authority by admitting what it cannot do. We will match that.

Governance instruments constrain what an agent is permitted and able to do. They do not make the underlying model aligned, and we have never claimed they do. If the frontier produces a system that can defeat its own revocation, the substrate underneath does not save you — Pachocki's problem swallows ours.

Adoption is the whole game and it is not won. A leash only works if it is held everywhere. An open standard that most builders ignore governs nothing. This is why our standards are published royalty-free and why we argue for safety through ubiquity rather than safety through one company's restraint — but nothing guarantees the ubiquity arrives in time.

And our own position is deliberately modest in what it can claim. The core standards are protected by a single Australian provisional patent application, filed on 2 February 2026, unexamined. The foundation intended to steward them independently of the company is being established, not yet registered. The pledge to keep the standards open is binding now; the transfer of ownership is promised, not complete. We say this in the same spirit Pachocki writes in: overclaiming on safety is its own kind of recklessness.

---

The posture

There is a reading of An Alien Mind that ends in paralysis — if the Chief Scientist cannot control it, who can? That is the wrong lesson, and it is not the one he draws. He does not say stop. He says stop racing, gate scaling on verifiable safety, and put enforcement in hands other than the developer's.

That is a build order, not a eulogy. "Don't race at all costs" is not "give up" — it is "construct the instruments that let us continue without betting the future on trust."

So our answer to the essay is the same as our answer to the whole moment: don't stop. Govern. The most senior technical voice in the industry has just described the problem in the open. The useful response is not to argue with the diagnosis. It is to hand over the tools.

Sources & Further Reading

  1. 1.Jakub Pachocki, "An Alien Mind", OpenAI (6 September 2026)
  2. 2.OpenAI Preparedness Framework
  3. 3.F-ACT — the open agent-governance standard
  4. 4.T-RUE — delegation-chain standard
  5. 5.Civilization 5.0 — the Sovereign Manifesto
AI AlignmentOpenAIRecursive Self-ImprovementAI GovernanceF-ACTHuman-in-the-Loop
The engine behind the Signal

Where this connects to Society OS

The Sovereign Intelligence Hub is the free, open front door of Society OS — the sovereign operating system that turns the ideas you just read into working governance. Where this piece names a problem, Society OS is building the machinery to solve it: AI agents that act with your authority, trust you can verify, and compliance that runs as code.

The 42-Protocol Stack

The governance engine beneath every article — led by the Sovereign Trinity: Human-Twin-Agent identity, HEARTrank trust, and WISE Contracts that execute law, not just code.

F-ACT — the open agent standard

The vendor-neutral framework for governing AI agents before they act: Grant, Usage, Audit, Revocation, Data — free to read, cite and implement.

The Sovereign Platform

Put it to work: govern a fleet of AI agents with verifiable authority, tamper-evident evidence, and compliance-as-code across your whole operation.

Explore membershipRead the F-ACT standard

Continue Reading

More from the Sovereign Intelligence Hub

I'm Giving Away the Patents. Here's Why.
AI Governance & Regulation

I'm Giving Away the Patents. Here's Why.

10 min
The Standard That Governs AI Agents Now Belongs to Everyone
AI Governance & Regulation

The Standard That Governs AI Agents Now Belongs to Everyone

8 min
The Enforcement Inflection: A Definitive Timeline of Global AI Governance, 2024–2028
AI Governance & Regulation

The Enforcement Inflection: A Definitive Timeline of Global AI Governance, 2024–2028

18 min read
The Governance Inflection: A Complete Timeline of Global AI Regulation, 2025–2026
AI Governance & Regulation

The Governance Inflection: A Complete Timeline of Global AI Regulation, 2025–2026

16 min read
The AI Safety Index: Grading the Giants
AI Governance & Regulation

The AI Safety Index: Grading the Giants

10 min
The EU AI Act Enters Full Enforcement: What the August 2026 Milestone Actually Means
AI Governance & Regulation

The EU AI Act Enters Full Enforcement: What the August 2026 Milestone Actually Means

16 min read

Never miss a signal

Weekly intelligence, no noise

Governance Toolkit

The Evidence
92 % ungoverned
The Framework
GUARD chain
Your Risk
Sourced model
Self-Assess
No login required

The Sovereign Intelligence Hub — Society OS

© 1989–2026 Society OS Pty Ltd. All rights reserved.