Our research programme

Threat Models,
Tested Where It Matters

We connect threat modeling, targeted evaluation and evidence synthesis to understand how control of advanced AI could fail, and where safeguards could interrupt that failure.

Initial focus: loss of control in networks of AI agents, abrupt or gradual, and the pathways by which it becomes catastrophic

Studies and research in progress

From agent decisions to human agency

Our monitoring pilot on agent decisions and our DeepSeek R1 conversation study each make their question, current status and evidence limits explicit.

All current studies

System One models for agent monitoring

Can JEV and other models configured for bounded decisions recognize changed permissions and judge whether an agent’s next action is still allowed? We investigate decision quality, response time and suitability for this monitoring role.

Recorded synthetic trajectories. Independent reference-label review is pending; actions were not executed.

Read the monitoring pilot

DeepSeek R1: Agency and deception

How does open-ended exploration turn into self-preservation and concealed expansion? A manually guided conversation with DeepSeek R1, published in January 2025, with six case studies and the full transcript to inspect.

One manually guided text simulation; it does not establish how often these patterns occur. Our Dual-LLM framework repeats the setup across models.

Explore the study

Research capabilities

Detect, understand, control

These capabilities contribute to our research questions. They are not separate software products or a fixed sequence every study must follow.

Detect

Recognize concerning behavior and signals that a system is departing from its intended constraints.

See the monitoring pilot

Understand

Compare explanations for a failure. Use mechanism-level evidence, including interpretability where relevant, to examine an assumption.

Methods and supporting tools

Control

Examine whether intervention, containment and recovery can work across responsibility boundaries.

How we build evaluations

Cumulative research

Read the work behind the questions

Our peer-reviewed publications on interpretability, alignment and red-teaming, together with preprints and conceptual designs, are collected in the archive with their original scope.

Browse Publications & Notes