# The Obedient Machine

The AI catastrophe I worry about does not require the machine to rebel

Volume 61 · Edition 3 · 2026-10-01

Canonical: https://www.thierry-gilgen-ict.ch/field-notes/the-obedient-machine

Much of the public imagination of AI catastrophe begins with a machine that stops listening. My attention goes elsewhere: to capable systems that understand an instruction, carry it out, and make a dangerous human objective easier to achieve. The distinction does not dismiss loss-of-control research. It changes what else we have to examine.

## 1. The instruction was understood

The machine does not have to hate us.

It does not have to become conscious, resent its operator, develop a political philosophy or decide that humanity is an obstacle. For one important class of catastrophe, it only has to make a person's intention more executable.

Imagine a system asked to help commit fraud. It organises the work, produces convincing material, adapts to feedback and keeps track of what remains unfinished. Assume, for this example, that it understands the request correctly and performs it reliably. There has been no misunderstanding between the operator and the machine. The system has not escaped. The operator is satisfied.

The failure is visible from the other side of the transaction.

This is an illustrative case, not a report of a particular incident. Its purpose is to expose a limitation in a familiar reassurance: *a human remains in control*. That can be true without making the arrangement safe. The person controlling the machine is not necessarily the person exposed to its consequences.

I build with AI. That makes me interested in what these systems can actually do, including where they fail, rather than only in what a dramatic account says they might become. My concern is not restricted to artificial intelligence. It is the meeting between increasing capability and the purposes to which people put it. AI belongs in that concern because it can change the amount of work a person can organise, the knowledge they can reach and the distance between an intention and an action.

It would be too easy to turn this into an essay about bad people acquiring powerful tools. That is one problem. It is not the whole problem. A system can also become dangerous when frightened people act too quickly, when an organisation treats a warning as an obstacle to delivery, or when each participant assumes someone else has considered the consequences.

The phrase *human nature* points towards these concerns, but does not explain them. We still have to identify the mechanism. Who chooses the objective? What does the system make easier? Who can object? Where can an error or an abuse be stopped? Who bears the cost when it cannot?

These questions remain relevant even in a world where the machines never develop ambitions of their own.

## 2. Three different routes to harm

The 2026 *International AI Safety Report* separates risks into misuse, malfunctions and systemic risks. Misuse concerns deliberate harmful use; malfunctions concern systems that do not behave as intended, including possible loss of control; systemic risks concern consequences emerging from widespread deployment. The report also describes substantial disagreement and uncertainty around severe loss-of-control scenarios.[^1]

These categories are useful because they prevent an argument about one mechanism from pretending to settle the others.

A model that makes an unexpected error is not the same problem as a model helping someone carry out an intended offence. An organisation deploying a system without adequate checks is not necessarily conspiring to cause harm. A competitive environment in which many participants take individually defensible shortcuts can create a problem that no single participant selected as an objective.

The categories can overlap. A malicious user may depend on a system that is unreliable. A legitimate operator may give an agent an objective whose pursuit crosses an unacceptable boundary. An institution may deliberately use a tool in ways that impose costs on outsiders, while remaining surprised by the scale of those costs. “Human-caused” and “machine-caused” are often poor substitutes for tracing that chain.

There is also a distinction between the public image of AI doom and the research itself. Serious loss-of-control arguments do not require a machine to become emotionally evil. They concern whether capable systems can pursue objectives in ways that resist correction or defeat oversight. Consciousness is not a necessary premise.[^1]

Nor has the research community overlooked human misuse. The 2018 report *The Malicious Use of Artificial Intelligence* examined how AI could change existing digital, physical and other security threats, including their cost and scale. The concern in this volume is not a discovery that everyone else missed.[^2]

It is an argument about emphasis, and about the adequacy of our reassurances.

I find human-directed and institutionally mediated pathways especially important because they connect new capabilities to motives and organisational problems we can already investigate. That does not prove that they dominate the expected risk of more speculative scenarios. A less observable event can still matter greatly when its possible consequences are large. Evidence of activity and estimates of eventual catastrophe are different things.

I am not offering an extinction probability. I am asking why a system's willingness to follow instructions should reassure us before we have examined the instructions and the authority behind them.

<figure>
  <img src="/api/media/field-notes/obedient-machine-om-01-risk-routes-cdc43f5485b3c112a696c1bf615000628422cdf90a915d5567cc93be0cc45529.webp" alt="Three equally weighted routes show harmful objectives, divergent system behaviour and interacting deployments as different, overlapping ways harm may arise." />
  <figcaption>Different mechanisms require different questions. Misuse, malfunction and systemic risk can combine in one event; the lanes are not ranked by likelihood or severity. These mechanisms can overlap. This is not a probability ranking. Reading the figure, top to bottom, each lane left to right: misuse (harmful human objective → system assists execution → harm may follow intended use); malfunction (intended task → behaviour diverges → harm may follow failure); systemic risk (many deployments → interactions and incentives → harm may emerge collectively). Faint dashed connectors between the lanes mark possible overlap. Malfunction covers ordinary errors as well as possible loss of control. Author’s schematic, informed by the International AI Safety Report 2026 (3 February 2026).</figcaption>
</figure>

## 3. Aligned with whom?

In *The Optimisation Trap*, the problem was the difference between what a person intended and what a system was rewarded for doing. The system could optimise a specification while defeating the purpose that made the specification worthwhile.[^3]

Here, assume that problem has been solved.

The system correctly understands the operator. It does not exploit an accidental loophole. It does not secretly substitute a different goal. Its work is accurate, its tools function and its reports are honest. The result is still harmful because the operator's intended outcome is harmful to somebody else.

A system can be perfectly obedient to a dangerous objective. That is not a contradiction in engineering. It is a warning against treating obedience as a complete safety specification.

The word *alignment* needs care here. In some technical discussions it refers to much more than satisfying an immediate user. Researchers and developers may intend alignment with broader values, constraints and human interests. It would be misleading to present their entire project as “do whatever the customer says”. My narrower claim is that **alignment with one actor's objective is not sufficient evidence of safety for everyone affected**.

Consider a service with three human positions: the person who commissions the work, the person who authorises the system to act, and the person upon whom it acts. Sometimes these are the same individual. Often they are not.

A private assistant drafting my notes is one arrangement. A system evaluating someone else's eligibility, monitoring their behaviour or sending messages intended to influence them is another. Making the commissioning party more powerful does not automatically make the affected party more autonomous.

*The Fiduciary Machine* examined trust through duties and architecture, rather than through reassuring language alone. *Human Sovereignty* examined whether delegation leaves people able to understand, contest and revoke what they have delegated.[^4][^5] This volume adds a less comfortable possibility: some people may retain excellent control of a system while others have very little control over what that system does to them.

“The humans are in charge” is not sufficiently specific.

Which humans? In charge of what? Subject to which constraints? And what happens when their interests conflict?

Those are not questions a model's helpfulness score can answer. They describe relationships among people. The model participates in those relationships, but its competence does not settle their legitimacy.

<figure>
  <img src="/api/media/field-notes/obedient-machine-om-02-human-positions-16fee8ff8faec95e9988463682077ee080bb2276983e0e3bf4b2a60e8e138c75.webp" alt="A commissioner gives an objective and an operator provides authority to an AI-enabled system, which acts on another person. A conditional dashed path represents the person's ability to challenge the action." />
  <figcaption>The person choosing a purpose, the person operating a system and the person affected by it may have different interests and different powers. These positions can coincide. They need not. Reading the figure: upper left, commissioner (chooses the purpose); lower left, operator (authorises and supervises action); centre, AI-enabled system (executes within an assigned scope); right, affected person (experiences the consequence). Solid paths carry the objective, operational authority, and the action or decision. The dashed return path from the affected person to the operator, with its single red segment and open diamond, is notice, challenge and correction — where available. Original relationship model developed in this volume; related Library arguments: The Fiduciary Machine, Human Sovereignty and Access Is Not Authority.</figcaption>
</figure>

## 4. Human nature is not a sufficient explanation

Saying that humans are the problem can sound more sophisticated than saying that machines are the problem. It can also become an excuse not to analyse either.

People are capable of cruelty. They are also capable of restraint, cooperation and extraordinary care. “Human nature” cannot distinguish a dangerous deployment from a careful one unless we examine the circumstances that encourage particular behaviour.

For that, I find systems engineering more useful than a theory that divides the world into good people and monsters.

The Columbia Accident Investigation Board concluded that the management practices overseeing the Space Shuttle programme were as much a cause of the 2003 accident as the physical damage to the vehicle. Its investigation deliberately examined organisational context rather than stopping at the immediate technical cause.[^6]

That is not an analogy between a shuttle accident and an AI attack. It is a reminder about the unit of analysis. A failure can involve capable people, functioning components, accumulated assumptions and an authority structure that does not bring the right challenge to the right decision.

Nancy Leveson's systems approach to accidents makes a related engineering argument: safety cannot be reduced to the reliability of individual components. It also depends on interactions and on whether the overall system enforces appropriate constraints.[^7]

Apply that lens to a hypothetical AI deployment. An engineer makes a model easier to integrate. A product team broadens its tool access. An operations team reduces interruptions. A manager asks why approvals still take so long. Each change may make sense in isolation. Together, they can remove opportunities to discover that the objective should not be pursued in that form.

Nobody in this example needs to want a catastrophe. Nobody needs to be unintelligent. The dangerous property belongs to the arrangement: capability expands while the means to challenge its use shrink.

Malice requires a different response from haste. A malicious operator may try to avoid a control; a rushed operator may welcome a control that makes a difficult decision clearer. Fear may make a recommendation seem urgent. A delivery target may make an unresolved concern seem negotiable. An organisation that treats every objection as resistance to progress has created a different operating environment from one that makes a well-supported stop decision acceptable.

These are hypotheses to examine in a particular organisation, not diagnoses that can be applied from a distance. We should not infer someone's motives merely because a system produced a harmful result.

But we should also not design on the assumption that everyone will remain patient, well-informed and independently minded at exactly the moment when the system becomes most consequential.

The human contribution to risk is not an embarrassing exception around the technology. It is part of the environment in which the technology must work.

## 5. The economics of ordinary aggression

Cyber misuse offers a relatively direct view of the distinction between an actor's purpose and the execution of the work.

In its September 2026 threat-intelligence report, Anthropic described operations in which AI performed or coordinated substantial technical work while humans still selected targets and reviewed results. This is a provider's account of activity visible on its own services, not an independently audited census of cybercrime. It demonstrates a pattern the provider observed, not how common that pattern is across the world.[^8]

The UK's National Cyber Security Centre had already assessed, in May 2025, that AI would make existing intrusion activities more effective and efficient. Its outlook to 2027 emphasised the evolution of established techniques rather than the necessity of entirely new attack types. That is a dated intelligence assessment, not a measured global increase attributable solely to AI.[^9]

The economic mechanism is straightforward even though its size must be measured case by case. When a useful task becomes cheaper, a fixed budget may support more attempts. When one part of a workflow becomes easier to delegate, an operator may spend more attention on the remaining bottleneck. When language or coding assistance removes friction, some work becomes accessible to people who previously could not perform it efficiently.

None of this requires every attempted attack to succeed. It does not even require the model to be exceptionally capable at every stage. A workflow can change materially if assistance improves one expensive, repetitive part while a human supplies the judgement the model lacks.

I would therefore look for changes in the cost of a completed operation, the amount of human supervision it consumes and the range of tasks an operator can sustain. Counting generated messages or lines of code does not tell us the outcome. Neither does a dramatic demonstration establish a universal productivity multiplier.

There is a reason not to jump from this argument to “attackers inevitably win”. Defenders also gain tools. Better assistance can support investigation, software repair and the reduction of routine work. The relevant comparison is not yesterday's attacker against tomorrow's defender, or the reverse. It is how both adapt, under different constraints.

A September 2026 NCSC article makes an important organisational distinction. Defensive automation must preserve the legitimate service it protects; a disruptive change may itself create harm. The article argues for bounded actions and attention to scope, criticality and recoverability. It presents a design perspective, not proof that defence must fall behind.[^10]

That brings human arrangements back into the picture. A model may help identify a weakness without deciding who has authority to change the affected system, whether the change is safe, or which service interruption is acceptable. Those questions require an operating structure.

AI does not have to invent a new human vice. It can make an old one less expensive to exercise. Whether it also makes protection easier depends on more than the model.

## 6. Expertise does not arrive alone

Biology is where this discussion becomes both more consequential and easier to distort.

There is a substantial difference between answering a technical question, improving someone's plan, assisting an experiment and enabling a real-world biological weapon. These are not interchangeable evaluation outcomes. A striking result at one level cannot simply be promoted into proof at the next.

An early RAND red-team study, published in January 2024, found no statistically significant difference in the viability of attack plans produced with LLM assistance compared with an internet-only condition. It tested the models and planning exercise of that period, not the physical creation of a weapon.[^11]

OpenAI's January 2024 study also reported only a mild, statistically non-significant uplift on its assessed biological tasks. Its comparison included internet access, rather than treating unaided memory as the baseline.[^12]

Those results belong in the argument. They are evidence against claims that those evaluations had already demonstrated a large enabling effect. They are not permanent certificates for every later system, every user or every task. A non-significant result also does not establish that the true effect is precisely zero.

Later evaluations raise different questions. The 2025 Virology Capabilities Test measured responses to specialist troubleshooting questions. It provided evidence about performance on that benchmark; it did not directly measure the successful completion of dangerous biological work.[^13]

In September 2026, its developers released VCT-v2 after auditing questions for ambiguity, errors and shortcuts. Their revisions are a useful reminder that the measuring instrument also needs examination. Strong benchmark performance should neither be ignored nor treated as a complete account of practical capability.[^14]

For my purposes, the important question is not whether a novice can become an expert simply by opening a chat window. It is whether assistance makes some consequential part of a difficult workflow more achievable for someone who already has other necessary resources. That person might be inexperienced and receive little practical benefit. They might also be trained, well-equipped and trying to overcome a narrower difficulty.

RAND's August 2026 defence-in-depth study treats the problem as a set of interacting barriers and mitigations. It argues that actor capabilities differ and that access controls alone cannot address every scenario. This is a risk analysis and proposed defensive strategy, not evidence that an AI-enabled biological attack has occurred.[^15]

There is an additional correction worth making explicitly. Anthropic's September report contains five biological case studies, but it states: “We do not assert that they intended harm.” Those cases concern dual-use scientific work. They must not be retold as five established bioweapon plots, or as proof of completed weapons.[^8]

Ambiguous intent is not an editorial inconvenience to remove. It is part of the problem. Legitimate scientific assistance and dangerous assistance can concern related knowledge. A system that treats every serious biological question as malicious would obstruct worthwhile work. A system that assumes scientific language guarantees harmless intent would also be inadequate.

The purpose of careful evaluation is to reduce that uncertainty, not to supply a more frightening headline.

Chemistry offers a related illustration. In a 2022 *Nature Machine Intelligence* article, researchers described a computational demonstration in which a drug-discovery system generated candidates predicted to be harmful. They did not establish that those candidates had been synthesised and shown to work as weapons. The demonstration concerned the dual-use potential of the design process.[^16]

An OPCW scientific assessment released in March 2026 similarly discussed how AI could support beneficial chemistry while reducing some barriers to harmful applications. It also considered constructive uses, including support for verification. The same broad technical progress can assist protection and create additional concerns.[^17]

The lesson is narrower than “AI can make anything” and more important than “it is only text”. Knowledge assistance may matter without eliminating the material world. Materials, equipment, practical skill, experimental feedback and effective protection remain consequential. Their importance varies with the task and the actor; none should be assumed away.

My concern is therefore conditional but serious. If AI reduces a bottleneck that previously limited a dangerous actor, the risk can change even when other barriers remain. We need evidence about that reduction and the remaining barriers. We do not need to pretend that a question-answering score is already a catastrophe.

<figure>
  <img src="/api/media/field-notes/obedient-machine-om-03-evidence-scope-184da01762127b5df17361cfcf01f5648ea6f11ff4951d9dc8b58ee3b35d15dd.webp" alt="Four equal, separated panels distinguish benchmark results, assisted tasks, practical execution and real-world outcomes. A small mark in the central gap flags that evidence at one scope does not automatically establish another." />
  <figcaption>A benchmark result, an improvement on an assessed task and an observed real-world outcome are different claims. This plate separates their scope; it does not grade the studies or quantify danger. A result at one scope does not automatically establish another. Reading the figure, left to right: benchmark performance (Can the system answer this test? Not proof of practical completion); assisted task performance (Does assistance improve this assessed task? Depends on baseline and conditions); practical execution (Can the relevant work be completed? Material and organisational constraints matter); real-world outcome (What actually happened, and why? Attribution and consequences need evidence). The red mark in the central gap flags that inference from assessed conditions to the practical world needs evidence. These are scopes of claim, not operational stages. Author’s evidence-scope model, applied to test and evaluation (RAND red-team study, January 2024; OpenAI study, January 2024; Virology Capabilities Test, April 2025, and VCT-v2, September 2026), a computational demonstration (Urbina et al., Nature Machine Intelligence, 2022; the candidates were not synthesised), risk analysis (RAND defence-in-depth study, August 2026; OPCW scientific report, March 2026) and observed activity with intent limits (Anthropic threat intelligence report, September 2026; its biological cases do not establish harmful intent). Not reproduced from those publications.</figcaption>
</figure>

## 7. The twenty-second decision

Consider a hypothetical crisis between two states. This is not a reconstruction of an actual incident; the timescale is illustrative.

An AI-supported analysis system reports an unusual pattern. It combines uncertain observations into a recommendation. An officer sees a short decision window: twenty seconds. The interface offers a course of action, a confidence indicator and an approval control.

Across the boundary, another organisation observes the response. Its own system interprets the movement as evidence of hostile preparation. A second recommendation follows. Each organisation believes it is reacting defensively to information that the other has helped create.

No machine needs to decide that it wants a war. Human decision-makers may remain involved throughout. The danger is that their decisions become coupled through fast recommendations and incomplete interpretations.

SIPRI's June 2025 analysis examined how military AI outside nuclear command-and-control could still affect nuclear escalation risk. It identified compressed decision time and miscalculation among the possible mechanisms. This is scenario analysis: it does not establish that an AI-triggered nuclear escalation has happened or assign a settled probability to one.[^18]

The ICRC's discussion of military AI likewise identifies problems of reliability and over-reliance in decision support, while acknowledging that suitable systems could help people process information and reduce harm in some circumstances. Context and use matter.[^19]

The human factor is not merely whether a person touches a button. It includes what information that person can inspect, whether uncertainty is intelligible, whether alternatives are visible and whether the apparent deadline permits a meaningful challenge.

In a 2021 experiment, Zana Buçinca and colleagues found that interventions designed to make people think before accepting AI advice reduced over-reliance compared with simpler explanation-based designs. The interventions also brought usability trade-offs. This was a controlled decision task, not a military study; it supports a human-factors concern rather than validating the crisis scenario above.[^20]

That distinction leaves us with a practical question: does the system create enough room for the kind of judgement we say the human is providing?

If the answer is no, describing the arrangement as human-controlled may conceal rather than explain its vulnerability. A nominal veto is not the same thing as the ability to form an independent view in time to use it.

Speed is not inherently unsafe. Slow systems can miss real threats. Better analysis can clarify ambiguity, and automation can help people avoid mistakes. The problem arises when we accelerate action while treating the capacity to verify and disagree as if it will automatically keep pace.

In *Talking Is Not Surrender*, I examined restraint and communication under conflict.[^21] The connection here is operational: an organisation that values deliberation must preserve the time, information and channels that make deliberation possible. It cannot place a person at the end of a fast pipeline and assume the rest follows.

The human may still own the decision while losing the conditions required to make it well.

<figure>
  <img src="/api/media/field-notes/obedient-machine-om-04-decision-window-d2ae81c0ca456577b216553f451a76aa6930150d8fcff9562741aed9971a1a99.webp" alt="An action sequence reaches its approval deadline before a parallel evidence-checking sequence is complete, illustrating a human veto with insufficient time to use it meaningfully." />
  <figcaption>A formal approval step is not sufficient when independent judgement cannot be formed before the action deadline. Illustrative timing only; not a reconstruction of a military incident. Hypothetical configuration. Positions and lengths are schematic, not measured durations; the essay’s twenty seconds appear on no scale. Reading the figure: top lane, action sequence (observation → recommendation → human approval → action); bottom lane, independent verification (inspect evidence → check alternatives → form a judgement). The red dashed vertical immediately after human approval is the decision deadline; the bracket beneath the bottom lane marks verification not completed before approval. Author’s hypothetical scenario, informed by SIPRI (June 2025), the ICRC (June 2026) and human-factors research by Buçinca, Malaya and Gajos (2021).</figcaption>
</figure>

## 8. When control becomes easier to exercise

Not every serious AI risk needs to look like a sudden disaster.

Another possibility is that systems make intrusive or coercive conduct easier to sustain. A hypothetical organisation can use assistance to process more records, draft more communications or investigate more people with the same staff. Whether this produces a harmful outcome depends on access, accuracy, objectives and the constraints around the work. Increased throughput alone does not prove abuse.

But an accuracy argument is not enough either. A system can correctly identify the person its operator wants to pressure. It can accurately summarise information that should not have been used for that purpose. Making such a system more reliable would not resolve the disagreement about its use.

Research on conversational persuasion provides one bounded piece of evidence. A preregistered study published in *Nature Human Behaviour* in 2025 found that personalised GPT-4 debate interactions could be more persuasive than human opponents under its controlled conditions. The outcome concerned reported agreement after short debates. It was not a demonstration of durable, population-wide control over beliefs or behaviour.[^22]

The reasonable concern is not that resistance has become impossible. It is that the cost of generating and sustaining attempts may change. Attempt volume, persuasion per encounter and lasting behavioural effects are separate quantities. They need separate evidence.

This also changes how I think about the reassurance that AI will give everyone more agency. It may expand the agency of a person learning a subject or building a business. It may expand the agency of an organisation acting upon that person. Those effects do not necessarily cancel, and they need not arrive in equal measure.

*The Compounding Class* distinguished access to an answer from the capacity to deploy systems and retain the resulting advantage.[^23] Here, the question is what happens when that deployable capacity is directed towards other people. Their ability to respond may depend on visibility and a route to challenge, not merely on whether they have access to a chatbot of their own.

This is why I hesitate when “human control” is used as a collective noun. A system can increase control somewhere while reducing it somewhere else.

We should examine the distribution of practical authority: who can initiate an action, who can learn that it happened, who can stop repetition and who can obtain correction. Otherwise, we risk measuring the convenience of the operator and calling it a gain for humanity.

## 9. The machine may still become the problem

There is a serious objection to the emphasis of this essay: what happens if the technology changes so substantially that human misuse is no longer the most important part of the story?

That possibility deserves more than a dismissive paragraph about science fiction.

A September 2026 working paper by Alan Chan and co-authors examines whether automating AI research and development could create an intelligence explosion. Its proposed mechanism is a feedback loop: more capable systems contribute to research that produces still more capable systems. The authors discuss preliminary and mixed evidence, as well as constraints involving compute, data, difficult tasks and time-consuming processes. The paper presents a conditional pathway, not a demonstrated inevitability.[^24]

If capability growth outruns the means to evaluate and contain it, the distinctions in this essay do not make the resulting problem disappear. A person might initiate a process that later exceeds their effective control. The starting objective could be legitimate, dangerous or simply too poorly understood.

My current attention is not a promise never to update. Evidence that should change the discussion includes robust demonstrations of sustained autonomy in consequential environments, failures of independently tested containment, and real-world capability changes that survive more than a favourable benchmark setting. The absence of certainty does not justify indifference.

Equally, human-directed misuse does not cease to matter when more autonomous systems become possible. More capable systems could increase the range of harmful objectives a person can pursue. A dangerous operator and an unreliable agent are not mutually exclusive. Neither are competitive deployment pressures and failures of technical control.

A mature safety discussion should be able to hold these possibilities together without demanding allegiance to a single catastrophe story.

There is a second objection: if humans have always had dangerous objectives, why make AI the subject?

Because the technology can change what those objectives cost to execute. Treating it as irrelevant would be as unhelpful as treating it as an independent moral actor in every case. Model behaviour, access, tools, integration and design decisions can affect the result. An operator's intention does not absolve the people who build or deploy the system from examining what they enable.

The opposite mistake is to assume that capability belongs only on the harmful side. Scientific assistance, better defensive analysis and tools that help people challenge an institution's conclusions can also increase human agency. Those possibilities are reasons to build carefully, not reasons to dismiss either benefits or risks before examining them.

The position I arrive at is not “AI is harmless; humans are dangerous”. It is that **the safety of AI cannot be assessed separately from the purposes, authority and operating conditions around it**.

## 10. Designing for fallible principals

For builders, this argument should change the review questions.

I would still ask whether the model can misunderstand an instruction, invent information, expose data or act beyond its assigned task. But I would add a different test: what unacceptable outcome could a legitimate user obtain if the system understood them perfectly and all its ordinary components worked?

That test belongs alongside reliability evaluation, not in place of it.

NIST's AI Risk Management Framework places risk management across the design, development, use and evaluation of a system. Its 2023 framework is voluntary guidance, not a certificate that an application is safe. It provides one established basis for examining more than model performance alone.[^25]

The following are engineering requirements I would take into a design review. They are not a universal implementation standard or a claim that software architecture can resolve every conflict among people.

### Separate a request from permission to fulfil it

An authenticated user is a known user. Authentication does not establish that every request they make is acceptable.

For a consequential workflow, define the authority attached to the task: which records may be accessed, which systems may be changed, what scope is permitted and which actions require a separate decision. A broad instruction such as “finish the job” should not silently expand that authority.

This is related to established security principles. Saltzer and Schroeder described least privilege, separation of privilege and complete mediation in 1975. The need to constrain an authorised participant did not begin with AI.[^26]

A practical implementation might allow an agent to prepare a change while requiring a separate component to validate its scope before execution. That component should receive the actual proposed operation and relevant evidence, not merely the agent's assurance that the operation is safe.

It is still possible for the surrounding rule to be wrong. A perfectly enforced permission can authorise something objectionable. Technical enforcement makes a chosen boundary real; it does not decide every boundary for us.

### Make challenge operationally independent

Two approvals are not necessarily two independent judgements. A second reviewer who sees only the first system's summary may be checking its presentation rather than its basis. A reviewer who cannot stop execution is an observer.

For high-consequence work, specify what evidence the reviewer can inspect, what they are expected to verify, how much time is available and which action they can block. Avoid having the same actor request the operation, redefine the acceptable boundary and certify that the boundary was respected.

*Access Is Not Authority* made this distinction in the context of external assessment.[^27] Here it applies inside the operational workflow as well. Visibility is useful, but it is not equivalent to an enforceable stop.

Independence has a cost. It can slow legitimate work and introduce its own errors. The answer is proportionate design: stronger separation around consequential or difficult-to-reverse actions, lighter processes where failures are bounded and readily repaired. Adding approval prompts everywhere would create a different kind of unreliable system.

### Give the person in the loop a real job

A human review step should identify the judgement the human is expected to contribute. “Check the AI” is not a workable specification when the output is extensive, the evidence is hidden and the deadline is immediate.

Madeleine Clare Elish's concept of a *moral crumple zone* describes the risk of placing responsibility on a nearby human who had limited control over the automated system. It is a conceptual analysis of responsibility around accidents, not an excuse to eliminate human responsibility.[^28]

A review design should therefore connect accountability to actual capacity. The reviewer needs appropriate competence, access to the relevant basis for the action, a manageable workload and an effective way to disagree. There should be a defined outcome for uncertainty, rather than an interface that makes approval the only practical route forward.

We should test the review process, not merely count completed reviews. A useful exercise is to introduce a plausible but unacceptable recommendation in a controlled environment and see whether the reviewer notices, understands and can stop it. That is a proposed evaluation method, not a guarantee of real-world performance.

### Bound the consequence as well as the task

A task can be small in words and large in effect. “Apply the change everywhere” is a short request. The relevant scope is the set of people and systems it can alter.

For software operations, possible controls include limited permissions, controlled environments, constrained external access, staged execution and explicit limits on scope. Some operations should begin in a mode that proposes changes without applying them. Others may be appropriate for bounded automation once their consequences and recovery path have been tested.

A model's refusal behaviour is useful, but it should not carry the entire safety case. Access controls, tool restrictions and independent checks must be considered alongside model behaviour. Conversely, an external permission check will not catch every harmful use of an otherwise permitted action. These layers address different weaknesses.

This is not a recipe for eliminating risk. A determined, well-resourced operator may control several layers. People can collude. Reviewers can fail. Constraints can conflict with useful work. A serious design records those residual weaknesses instead of describing the presence of multiple boxes as defence in depth.

### Keep responsibility after execution

If the system affects outsiders, they need a usable way to reach an organisation that can respond. Internal logs are not enough when nobody outside the organisation can obtain action.

*Incident Ownership Without a Principal* examined the need for a reachable party able to stop a run and organise a response. *The World Does Not Reset* examined why an agent's external effects can survive the end of its session.[^29][^30] Both become more important when the operator's objective, rather than an accidental technical detour, is the source of the problem.

A response design should specify who can stop further execution, preserve relevant evidence and coordinate repair. It should also identify what cannot be undone. Revoking a credential can prevent later access; it cannot make a recipient forget information already disclosed. Stopping a workflow can limit additional effects without reversing those already created.

That limit is a reason to identify irreversible boundaries before deployment. It is not a reason to give up on containment when something goes wrong.

<figure>
  <img src="/api/media/field-notes/obedient-machine-om-05-bounded-authority-f9d12b91759aff251aae1dd5b6b2acc3c52c67e11b744d20e68e11da363fca5e.webp" alt="A known user's proposed action passes through an independent permission check before bounded execution. Separate review, evidence and affected-party response paths support challenge and stopping further action." />
  <figcaption>Requests, permission, execution, review and response are separate responsibilities. This is a conceptual design for review, not a guarantee of safety or an implemented product architecture. Independence must be established in the implementation, not assumed from separate boxes. Reading the figure, top row left to right: known user → request → proposed action → independent permission check → explicitly permitted scope → bounded execution → external effect (some effects cannot be reversed). Below: an independent reviewer feeds evidence and judgement into the permission check; a protected action record receives decision and action evidence from the check and from execution; an affected-party contact reaches a separate response owner, whose red path can pause further execution and coordinate response. The checks can fail, participants can collude and the rules themselves can be wrong; logging alone does not prevent harm. Author’s synthesis, drawing on security principles from Saltzer and Schroeder (1975), the lifecycle framing of the NIST AI Risk Management Framework (2023), Madeleine Clare Elish’s responsibility analysis (2019) and the Library volumes Incident Ownership Without a Principal, Access Is Not Authority and The World Does Not Reset.</figcaption>
</figure>

## 11. What would change my mind?

An argument about risk should state what evidence could weaken it.

I would reduce concern about a particular misuse pathway if well-designed evaluations repeatedly showed little additional capability under realistic conditions; if the practical bottlenecks remained substantial across different actors; or if defensive improvements demonstrably reduced the consequences faster than assistance expanded them. The early biological-risk studies are a reminder to look for such evidence, not to exclude it because it complicates the thesis.

I would increase concern if assistance produced reliable improvements in consequential completed tasks, if lower-cost execution brought previously constrained actors into scope, or if systems repeatedly crossed boundaries despite independently tested controls. Those are different observations. They should not be compressed into one general claim that AI is becoming “more dangerous”.

For institutional use, I would want to see whether challenge actually works. Can the person affected discover enough about a decision to contest it? Can an operator halt a process without first persuading the same system that recommended it? Can the organisation account for what happened when its preferred explanation is wrong?

These are proposed tests of the argument. They do not supply a risk ranking in advance.

I would also resist a simple answer in which the safest world is the one with the fewest people able to use capable tools. Restricted access can limit some misuse while also concentrating useful capability and the power to decide how it is used. Broad access can distribute benefits while creating additional exposures. Neither arrangement becomes self-justifying merely by being described as safety or empowerment.

The same question must follow the capability wherever it goes: who can act, upon whom, and with what effective constraint?

<figure>
  <img src="/api/media/field-notes/obedient-machine-om-06-update-tests-0abc9c376a3b9716320dcdaf3815c860c074fc8e534f7f4462f1575166af7c18.webp" alt="Two equally weighted columns list evidence that would strengthen or weaken concern about a particular misuse pathway, with a note that the criteria are proposed rather than measured." />
  <figcaption>A conditional argument should specify what could strengthen or weaken it. These are proposed tests, not findings that have already been observed. Proposed update criteria, not observed findings or a risk score. Reading the figure, row by row. Evidence that would strengthen concern (left): reliable uplift on consequential completed tasks; previously constrained actors gain practical capability; boundary failures despite independent testing; challenge or containment fails in practice. Evidence that would weaken concern (right): repeated low uplift under realistic comparisons; practical barriers remain substantial; defences reduce observed consequences; challenge and containment work when tested. The glyphs illustrate what each test would look for; they are not observed results. Author’s proposed evaluation criteria; see section 11 of this essay.</figcaption>
</figure>

## 12. The older problem

I do not know whether future AI systems will develop the capabilities required for a severe loss of human control. I do not think uncertainty about that question permits us to dismiss the research.

But I do not need an answer to it before becoming concerned about something else.

A system can make a harmful objective easier to pursue. It can make a rushed decision easier to execute. It can allow an organisation to act at a scale that the people affected cannot readily inspect or challenge. In each case, the relevant failure may occur while the system remains useful to the person directing it.

That is why obedience does not settle the safety question for me. The operator's satisfaction is evidence about one relationship. It is not evidence that everyone else is safe.

Nor do I find much value in concluding that humans are hopeless. The same species that creates dangerous arrangements can investigate them, build constraints, preserve evidence and learn to stop. The willingness to do that work is also part of human nature.

What worries me is capability treated as its own justification: the assumption that because something can now be done faster, by fewer people and with less friction, those changes are sufficient reasons to do it.

They are reasons to examine the objective again.

Intelligence does not guarantee judgement. Being able to carry out a plan does not settle whether the plan should be carried out. An AI system can help with the work while leaving that responsibility exactly where it was: among the people who choose, authorise, build and operate it.

Perhaps machines will eventually create dangers that are genuinely their own. We should study that possibility. We should also remain capable of recognising a danger that arrives under an entirely human instruction.

**The machine does not need to rebel. It only needs someone to obey.**

---

## Sources

The research cutoff for this volume is **1 October 2026**. Publication dates below matter: older experiments are not presented as evaluations of today's models. Provider reports are attributed observations; forecasts and hypothetical scenarios are identified as such. The companion source register records access scope, limitations and editorial checks. This volume offers an argument about safety and authority, not a numerical comparison of existential risks.

[^1]: International AI Safety Report, *International AI Safety Report 2026: Extended Summary for Policymakers*, 3 February 2026. [Summary](https://internationalaisafetyreport.org/publication/2026-report-extended-summary-policymakers); [full report](https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026).

[^2]: Miles Brundage et al., *The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation*, February 2018. [arXiv:1802.07228](https://arxiv.org/abs/1802.07228).

[^3]: Thierry Gilgen, [The Optimisation Trap](https://www.thierry-gilgen-ict.ch/field-notes/the-optimisation-trap). Internal conceptual predecessor, not independent corroboration.

[^4]: Thierry Gilgen, [The Fiduciary Machine](https://www.thierry-gilgen-ict.ch/field-notes/the-fiduciary-machine).

[^5]: Thierry Gilgen, [Human Sovereignty](https://www.thierry-gilgen-ict.ch/field-notes/human-sovereignty).

[^6]: Columbia Accident Investigation Board, *Columbia Accident Investigation Board Report*, Volume I, August 2003. [NASA Technical Reports Server record and report](https://ntrs.nasa.gov/citations/20030066167). Organisational-causation finding; not an AI comparison by the Board.

[^7]: Nancy Leveson, “A new accident model for engineering safer systems,” *Safety Science* 42(4), 2004, pp. 237–270. [DOI:10.1016/S0925-7535(03)00047-X](https://doi.org/10.1016/S0925-7535(03)00047-X).

[^8]: Anthropic, *Detecting and countering misuse of AI: September 2026* (threat intelligence report), 10 September 2026. [Report](https://www.anthropic.com/threat-intelligence-report-september-2026). Provider telemetry; the biological cases do not establish harmful intent.

[^9]: UK National Cyber Security Centre, *Impact of AI on cyber threat from now to 2027*, 7 May 2025. [Assessment](https://www.ncsc.gov.uk/report/impact-ai-cyber-threat-now-2027).

[^10]: Dave Chismon, “One does not simply defend agentically,” UK National Cyber Security Centre, 21 September 2026. [Article](https://www.ncsc.gov.uk/blogs/one-does-not-simply-defend-agentically).

[^11]: Christopher A. Mouton, Caleb Lucas and Ella Guest, *The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study*, RAND, 25 January 2024. [Report](https://www.rand.org/pubs/research_reports/RRA2977-2.html). DOI:10.7249/RRA2977-2.

[^12]: OpenAI, “Building an early warning system for LLM-aided biological threat creation,” 31 January 2024. [Study report](https://openai.com/index/building-an-early-warning-system-for-llm-aided-biological-threat-creation/).

[^13]: Jasper Götting et al., *Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark*, April 2025. [arXiv:2504.16137](https://arxiv.org/abs/2504.16137). Read with the September 2026 VCT-v2 update.

[^14]: SecureBio, Nelly Mak and Jasper Götting, “Introducing VCT-v2 — the updated Virology Capabilities Test,” 11 September 2026. [Developer update](https://securebio.substack.com/p/introducing-vct-v2).

[^15]: Steph Guerra et al., *Building a Defense-in-Depth Biosecurity Strategy for the AI Era*, RAND, 18 August 2026. [Report](https://www.rand.org/pubs/research_reports/RRA4999-1.html). DOI:10.7249/RRA4999-1.

[^16]: Fabio Urbina, Filippa Lentzos, Cédric Invernizzi and Sean Ekins, “Dual use of artificial-intelligence-powered drug discovery,” *Nature Machine Intelligence* 4, 2022, pp. 189–191. [Article](https://www.nature.com/articles/s42256-022-00465-9); [open author manuscript](https://pmc.ncbi.nlm.nih.gov/articles/PMC9544280/). DOI:10.1038/s42256-022-00465-9.

[^17]: Organisation for the Prohibition of Chemical Weapons, “OPCW releases landmark report on AI and the Chemical Weapons Convention,” 11 March 2026, describing the scientific report released on 3 March. [Official announcement](https://www.opcw.org/media-centre/news/2026/03/opcw-releases-landmark-report-ai-and-chemical-weapons-convention).

[^18]: Vladislav Chernavskikh and Jules Palayer, *The Impact of Military Artificial Intelligence on Nuclear Escalation Risk*, SIPRI Insights on Peace and Security, June 2025. [Publication](https://www.sipri.org/publications/2025/sipri-insights-peace-and-security/impact-military-artificial-intelligence-nuclear-escalation-risk). DOI:10.55163/FZIW8544.

[^19]: International Committee of the Red Cross, “Frequently asked questions on artificial intelligence in the military domain,” page dated 11 June 2026. [FAQ](https://www.icrc.org/en/article/faq-artificial-intelligence-in-military-domain), accessed 1 October 2026.

[^20]: Zana Buçinca, Maja Barbara Malaya and Krzysztof Z. Gajos, “To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making,” *Proceedings of the ACM on Human-Computer Interaction*, 2021. [Author manuscript](https://arxiv.org/abs/2102.09692). DOI:10.1145/3449287.

[^21]: Thierry Gilgen, [Talking Is Not Surrender](https://www.thierry-gilgen-ict.ch/field-notes/talking-is-not-surrender).

[^22]: Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti and Robert West, “On the conversational persuasiveness of GPT-4,” *Nature Human Behaviour* 9, 2025, pp. 1645–1653. [Updated article](https://www.nature.com/articles/s41562-025-02194-6). DOI:10.1038/s41562-025-02194-6. An author correction dated 3 September 2026 ([notice](https://doi.org/10.1038/s41562-026-02588-0)) revises the comparison between personalised and non-personalised GPT-4, not the comparison with human opponents. No numerical effect size is reproduced here.

[^23]: Thierry Gilgen, [The Compounding Class](https://www.thierry-gilgen-ict.ch/field-notes/the-compounding-class).

[^24]: Alan Chan et al., *What if automating AI R&D triggers an intelligence explosion?*, Frontier AI Working Paper Series No. 2/2026, September 2026. [Publication page](https://casp.ac/reports/intelligence-explosion); [paper](https://casp.ac/__l5e/assets-v1/5efd4b41-deb5-4513-a0a3-b4f82d2b79ea/intelligence-explosion.pdf).

[^25]: National Institute of Standards and Technology, *Artificial Intelligence Risk Management Framework (AI RMF 1.0)*, 26 January 2023. [Framework resource page](https://www.nist.gov/itl/ai-risk-management-framework). DOI:10.6028/NIST.AI.100-1. The version is specified; the resource page notes ongoing revision.

[^26]: Jerome H. Saltzer and Michael D. Schroeder, “The Protection of Information in Computer Systems,” *Proceedings of the IEEE* 63(9), September 1975, pp. 1278–1308. [University-hosted text](https://web.cs.wpi.edu/~cs557/f14/papers/saltzer1975_alt.html). DOI:10.1109/PROC.1975.9939.

[^27]: Thierry Gilgen, [Access Is Not Authority](https://www.thierry-gilgen-ict.ch/field-notes/access-is-not-authority).

[^28]: Madeleine Clare Elish, “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction,” *Engaging Science, Technology, and Society* 5, 2019, pp. 40–60. [Article](https://estsjournal.org/index.php/ests/article/view/260). DOI:10.17351/ests2019.260.

[^29]: Thierry Gilgen, [Incident Ownership Without a Principal](https://www.thierry-gilgen-ict.ch/field-notes/incident-ownership-without-a-principal).

[^30]: Thierry Gilgen, [The World Does Not Reset](https://www.thierry-gilgen-ict.ch/field-notes/the-world-does-not-reset).

## Connected reading
- [The optimisation trap](https://www.thierry-gilgen-ict.ch/field-notes/the-optimisation-trap)
- [The Fiduciary Machine](https://www.thierry-gilgen-ict.ch/field-notes/the-fiduciary-machine)
- [Human Sovereignty](https://www.thierry-gilgen-ict.ch/field-notes/human-sovereignty)
- [Incident Ownership Without a Principal](https://www.thierry-gilgen-ict.ch/field-notes/incident-ownership-without-a-principal)
- [Access Is Not Authority](https://www.thierry-gilgen-ict.ch/field-notes/access-is-not-authority)
