Research / Threat Modeling

Pathways to Harm,
Made Testable

Threat modeling asks how an AI system could cause harm, step by step, and where a safeguard could stop it. We write these steps down so that others can check them, challenge them and test them.

Why we do this

Start with a decision that matters

AI systems are gaining autonomy. Agents plan tasks, hand work to other agents and use real tools, and more capable models could change how organizations and societies decide. We build threat models for several of these developments, for example human oversight of cooperating agents, and a gradual loss of human influence through everyday delegation. In each we ask how abilities, access, dependencies and human decisions could add up to a loss of control, through deliberate misuse, unintended behavior, or because coordination and oversight break down.

We cannot wait for a loss of control to happen before we study it, and we cannot test every combination of systems, tools and operators. Threat modeling lets us reason about harm before it occurs: we write down how it could unfold, step by step, and where a safeguard would have to hold. Behind every model is the same question: do people keep the ability to notice, understand, decide, intervene and reverse? That ability can be lost through what AI systems do, or gradually through how people use them, when they stop learning or can no longer work without agents.

A good threat model states what it assumes, what the evidence shows and what would change our assessment. It is useful when it helps someone make a better decision, for example who should be able to stop which agent.

Our research workflow

From a first model to ongoing research

We work in five stages. Stage 5 is not the end: the revised model and the new questions it raises start the next round. A new finding can also send us back to any earlier stage, and not every model needs every experiment.

  1. 01

    Set the scope

    What could be harmed? Who can decide or step in?

    What is in and out, and who can actExample Several agents, one or several operators
  2. 02

    Model pathways

    How could an AI capability lead to harm?

    Possible pathways to harm, and reasons they might not occurExample A pause stops one agent, but not the others
  3. 03

    Examine evidence

    What do research, incidents and experts say? What is still uncertain?

    What we know, what experts judge, and what is still openExample What the pilot shows, and where experts disagree
  4. 04

    Assess safeguards

    Where could a safeguard stop it? What needs a test?

    Weak points, and what we need to testExample Simulate, then test the one decisive link
  5. 05

    Review and revise

    What changes after new findings or criticism?

    An updated model, and new questions for the next roundExample Update the model with the test result

The revised model starts the next round. A new finding, an incident or a good counterargument can change what we include, which pathway we consider likely or which safeguard we recommend.

Evidence throughout

Start with what is already known

Research papers, documented incidents and expert critique often answer a question. Then no new experiment is needed.

When a test can help

Test what the model cannot yet answer

When an important assumption is still open, we build test cases in an evaluation environment. Before the test runs, we fix which result would change the model. That way each test reduces the uncertainty that matters for the decision.

How we build evaluations

Illustrative example

Can the whole chain of agents be stopped?

This is how the five stages work on one concrete question. The example illustrates the method: the pathways and outcomes below are hypotheses chosen to explain the steps, not findings or a study we are running.

Testing alone would not tell us which of the many possible setups to test, or what a passed test means. The model narrows the question to the one link that is both uncertain and decisive, and shows what a result does not cover. The test then answers exactly that link.

The situation

A task is split among several AI agents. One agent plans the work and hands parts of it to other agents, which carry them out with real tools. Something looks wrong, and a person pauses the planning agent. Do the tasks it has already handed on stop as well?

01

Set the scope and the decision

We look at several agents that share information and hand tasks to each other. We compare two setups: one operator with authority over all agents, and several operators who each control only part of the chain. The decision this should inform: who should have the right to intervene, and where should monitoring sit?

Output Two setups to compare and one decision to inform. A framing, not yet a claim.

02
STPA

Describe the pathways, and an alternative

We use STPA to list how a pause can fail: it is not issued, comes too late, reaches only some agents, or hits the wrong one. Each becomes a pathway in the model. Our main hypothesis: pausing one agent leaves the tasks it handed on running. The alternative: a shared control point already reaches every agent that carries out actions, so the gap does not exist. Writing down the alternative tells us what a test has to tell apart.

Output A causal model: events and the arrows between them. These are assumptions, derived systematically, not findings.

03
Evidence and expert elicitation

Estimate what we know

Our monitoring pilot shows that a monitor can spot problems in recorded agent proposals, so detection needs no new test here. No study shows whether handed-on tasks stop when the planning agent is paused. Experts estimate this in a structured round, and their answers differ widely.

Output A strength for each arrow: from evidence where it exists, from pooled expert judgment where it does not. The two are recorded separately.

04
RiskGraph simulation

Simulate, then choose the test

RiskGraph runs the model many times with these uncertain estimates. It shows which element moves the outcome most: here, whether handed-on tasks keep running, which is also where experts disagree most. It also shows which safeguard would reduce the risk most: a shared control point, if that arrow holds. So we test that one link and nothing else. Before the test runs, we fix which result would support which hypothesis.

Output A ranking of safeguards and the one test that would change it most. This follows from the model and its inputs.

05
Controlled test

Run the test and update the model

In a test environment we keep the tasks and agents the same and change only who has authority. We check whether actions continue after the pause, whether they spread into shared data, and whether an independent observer can confirm that the system is back in a safe state. The test is the only step that observes cause and effect directly. Its result replaces the expert estimate for that link, and RiskGraph recalculates.

Output An updated estimate and a recommendation. This is observation, for this setup.

The result

Before the test: experts disagree strongly on whether handed-on tasks keep running. In the model, a shared control point is the strongest safeguard, if that link holds.

After the test, if the hypothesis holds: with divided authority, handed-on tasks keep running after the pause; with a shared control point, they stop. The expert estimate is replaced by this observation.

Recommendation: give the right to intervene to a control point that every executing agent must pass, not only to the planning agent.

Limits: this holds for the tested setup. Real deployments, and agents that start new agents themselves, are not covered.

Next round: does the control point still hold when agents can start new agents?

Stage 02 in real work

Human oversight of a swarm of agents

The example above is invented to show the whole chain. This is stage 02 carried out for real: one overseer, many fast agents, a halt command that still exists. Where can that control loop fail? A Systems-Theoretic Process Analysis (STPA) of the loop.

Status, September 2026: stage 02 is done but not yet independently reviewed. Stages 03 to 05 are open: the mechanisms are already set up as a RiskGraph model, but there has been no elicitation round and no test yet. We update this section as each stage is completed.

Loss L1: harm to a critical service. Hazard H1: the swarm acts without effective human constraint. Hazard H2: the system is left in a state that cannot be verified.

Four ways the control loop is inadequate

1 / Feedback delayed

Pace outruns review; the halt comes too late.

2 / Feedback produced by the agents

The overseer sees what the swarm reports.

3 / Control action bypassed

Gated actions are split so that no approval is triggered.

4 / Process model wrong

The overseer’s picture no longer matches the system.

Unsafe control actions for “halt the swarm”

Type Unsafe control action Hazard Causal factor
Not providedThe overseer does not halt while a coordinated deviation is runningH11, 2
Too lateThe halt is issued after the deviation has spread beyond rollbackH1, H21, 4
Applied incompletelyThe halt reaches the orchestrator; agents continue via side channelsH13
Wrong targetThe halt is aimed at the wrong agents, based on an outdated pictureH24
Approve without reviewThe queue exceeds capacity; approval becomes a formalityH11

Causal factors refer to the four inadequacies above. The full analysis also covers the control actions “approve” and “roll back”. Each row is a claim about how control fails; the experiments that would test these claims are proposed, and none has run.

From one system to society

From system failures to systemic risk

Our main concern is catastrophic harm from advanced AI, whether it comes suddenly or through a gradual loss of human control. We study it at three levels. A result at one level does not automatically carry over to the next; each step needs its own argument.

AI systems

Does a monitor notice when an agent breaks a rule? Does a safeguard reach the action it is meant to stop? What can a single agent do with the permissions it has?

Organizations

How do handing off tasks, shared infrastructure and split responsibility affect control inside a company, or between several operators?

Systemic risk

When could a failure spread across institutions, overwhelm the ability to recover and cause harm on a catastrophic scale?

A failure in one company does not prove an existential risk. How many are exposed, how far a failure spreads, how severe it gets and whether recovery is possible all have to be shown separately.

Methods in more depth

Choose the method for the question

Methods from safety engineering and security show how a system is structured and where it can fail. RiskGraph, our internal platform, turns the resulting pathways into events experts can judge, and shows which safeguard changes the outcome most. STPA has been applied in a first analysis; the other methods are candidates for particular studies and not yet in use at COAI.

Estimating how likely pathways are · RiskGraph

RiskGraph links clearly defined events, expert judgments, evidence and counterarguments in one model. This makes every assumption visible and shows what follows from it. The numbers hold within the model and its inputs: whether an arrow is truly causal, and whether experts are well calibrated, needs separate support.

RiskGraph simulates the model to show which element and which safeguard move the outcome most, and which test would reduce the uncertainty most. Causal evidence for a single link comes from a controlled test, not from the simulation.

See how we use RiskGraph

Who controls what · STPA, in use

Systems-Theoretic Process Analysis (STPA) treats accidents as failures of control, not only of components. It asks who sends which commands, what feedback they receive and when a correct command becomes unsafe. For our running example: who issues the pause, what do they see afterwards, and when does the pause fail to take effect?

Our first analysis applies it to human oversight of a swarm of agents: who issues a halt, what feedback they receive and when the halt is ineffective. Independent review of that analysis is still open.

Method reference: MIT STPA handbook ↗

How an attacker could proceed · Attack trees, STRIDE

Attack trees break an attacker’s goal into the steps needed to reach it. STRIDE is a checklist of common threat types, such as faking an identity or tampering with data. Both cover deliberate misuse; accidents and institutional dependencies need other methods.

Method reference: Microsoft STRIDE documentation ↗

How small variations add up · FRAM

The Functional Resonance Analysis Method (FRAM) describes what each part of a system does and how its performance varies. It suits cases where ordinary variations between agents or people combine into a failure that no single part would cause. It complements RiskGraph, which works with separate events rather than continuous feedback.

Method reference: FRAM foundations ↗

Where safeguards fail together · Barrier analysis

Barrier analysis maps where safeguards interrupt a pathway, where gaps remain and whether safeguards that look independent can fail for the same reason. We record the analysis, its evidence and its limits in the model dossier.

How we keep records

What a model dossier records

Every model has a dossier. It brings together everything behind the model: analyses, expert judgments, evaluation results and sources, along with the assumptions, counterarguments and uncertainty, the review status and what would make us revise the model. The dossier links the model to the decision it is meant to inform, so anyone can trace a result back to its evidence.

We are building a shared evidence database to hold these dossiers. It is at the proof-of-concept stage and not yet in use.