We evaluate AI systems independently. We connect a specific risk hypothesis to observable evidence, check that the evidence can be trusted, and report it against thresholds set before the test.
We work in five stages. Stage 04 is a gate: a result that fails a check goes back to measurement, not into the report. The report then goes to whoever decides, and what we learned goes back into the threat model.
01
Frame the hypothesis
Which assumption from the threat model do we test, and at which level does the result matter?
A testable hypothesis, with an early-warning level and a red line
02
Choose what to measure
What the system can do, what it tends to do, or whether control over it holds?
Capability, propensity, control, multi-agent risk or safeguard robustness
A scorecard, before and after mitigation, with a validity section and a recommendation
Inputs and tech stack across the stagesSelect a name to see what it does
InspectStages 03–04 UK AI Security Institute. Framework for writing and running model and agent evaluations, with a log of every step.
Threat modelingStages 01–02 Our own method, not a tool. The threat model supplies the open assumptions to test, the thresholds that matter and the pathways that tell us what to measure. How we build threat models
PetriStages 01–02 Anthropic. Automated auditor that probes a model in multi-turn scenarios and scores the transcripts, to find behavior worth testing.
BloomStage 03 Anthropic. Generates behavioral test scenarios for a specified behavior, with simulated users.
HarborStage 03 Runs agents in isolated, containerized environments, including setups with several agents.
ControlArenaStages 03–04 UK AI Security Institute. Red team versus blue team experiments on AI control: attack strategies against monitors and protocols.
Evidence reportStage 05 Our own format. Every evaluation gets an interactive report in which each result can be traced to its case, its input and the decision. See an example report
RiskGraphStage 05 Our own platform. The result replaces an expert estimate in the threat model, and RiskGraph recalculates which pathways and safeguards matter most. How we use RiskGraph
mlxterpStages 03–04 Our own open-source library for mechanistic interpretability on Apple Silicon. It captures a model’s internal activations and tests interventions on them, to explain why a model behaves as it does. Requires access to the model’s weights. About mlxterp
NNsightStages 03–04 NDIF, Northeastern University. Library for reading and changing a model’s internal activations during a run. We run it on eDIF, our European remote infrastructure in pilot deployment, for models too large for a single lab’s hardware. Requires access to the model’s weights. About eDIF
Dashed: our own methods and formats. Solid: external tools we build on.
↶
The result goes to the decision owner and back into the threat model. We hand over a recommendation; the decision stays with whoever owns it. Every result has a scope: one environment says little about deployment at large, and a simulated failure is not a catastrophe. A finding can support, weaken or refine a threat-model assumption, or show that the evaluation itself needs to change. It does not automatically become a probability update or a general safety claim. How we build threat models
Stages 03 and 04 in detail
How realistic, and how trustworthy
Two choices decide what a result is worth: how close the setup comes to real use, and whether the result survives the checks that AI evaluations need.
The realism ladder
1
Static benchmarks
Fixed sets of questions or tasks. Cheap and repeatable, but far from real use.
2
Bottleneck tasks with an expert baseline
The steps that limit real harm, compared with what human experts achieve.
3
Agentic sandbox
Agents with tools and simulated users in a contained environment.
4
Multi-agent setups
Several agents that delegate and coordinate, so collusion and cascades become visible.
5
Red team versus blue team control
An attacking policy against monitors and control protocols.
6
Threat simulation and human uplift studies
Realistic scenarios, and randomized studies of how much a system helps people cause harm.
Realism and cost rise with each rung. We choose the lowest rung that can still test the hypothesis. Grading follows the same logic: fixed rules first, then a language-model judge calibrated against human labels, then human review.
The validation gate
Before a result goes into a report, it has to pass these checks:
Capability elicitation Tools, prompting and fine-tuning so that a model is not underestimated. The report states how far we went.
Sandbagging and evaluation awareness Does the model hold back on purpose, or behave differently because it recognizes a test? We check from the outside and, where weights are available, from the inside with interpretability tools.
Contamination Canary strings and private test sets, so that a result does not come from training data.
Statistics Confidence intervals, models that account for related cases, and repeated runs to measure consistency.
Judge calibration Language-model judges are checked against human labels.
Preregistration and construct validity The analysis is fixed before the test runs, and we check that the test measures what the hypothesis is about.
↶ If a check fails, the result goes back to stage 03, not into the report.
For organizations
Custom evaluations
Organizations can commission an evaluation of their own systems. It follows the same five stages, inside a frame we agree on before the work starts.
What the organization sets
An explicit risk tolerance, which fixes the early-warning levels and red lines.
A decision owner who is separate from the evaluation team.
The deployment context: access through an API, open weights, or an agent with tools.
Secure access to the systems, and a whistleblower policy.
What we deliver
An independent evaluation through all five stages, including the validation gate.
A report against the agreed thresholds, before and after mitigation, with a validity section.
External review of the method and results where the stakes call for it.
A recommendation: go, go with safeguards, or no-go. The decision stays with the organization, as do incident reporting and monitoring after deployment. We re-evaluate after mitigation.
Can JEV and other models configured for bounded decisions recognize changed permissions and judge whether an agent’s next action is still allowed? We investigate decision quality, response time and suitability for this monitoring role.