Our research record

Publications & Notes

Explore completed work, conceptual designs and questions still taking shape. Each contribution retains its own scope and status.

Notes

Questions that shape the next study

Exploratory writing can motivate a model or an experiment. It is distinct from an established finding.

All Notes
Framework proposal

What would effective containment require?

The Shoggoth Prison framework proposes restricted agency and monitor-mediated interactions. Which assumptions must hold for those barriers to work?

Question to test: Can a monitor's decision be enforced at every action boundary?

Sudarshan Kamath Barkur · 14 Mar 2026
Read the original note
Early research idea

What evidence do we need after an agent failure?

The Flight Recorder note explores provenance and replay. It motivates inspectable traces without establishing that logs reveal a model's internal reasoning.

Question to test: Can we reconstruct what each actor knew, could do, and actually did?

Sigurd Schacht · 05 Oct 2025
Read the original note
Argument from a forthcoming paper

When does handing work to AI erode the person doing it?

Three conditions decide whether offloading a cognitive operation augments or erodes. Each failure names a distinct mode, and one of them produces overseers who hold control in form only.

Question to test: Which links in a control chain assume a competence that routine AI use removes?

Carsten Lanquillon · 24 Sep 2026
Read the original note

Research archive

Publications

Studies, papers and conceptual designs across our research. Read each entry for its original scope and status; inclusion here does not imply peer review.

Conceptual Design and Taxonomy of a RAG-Based Attack Memory for LLM Red Teaming

Conceptual Design and Taxonomy of a RAG-Based Attack Memory for LLM Red Teaming

07 July 2026

Every red-teaming campaign starts from scratch and throws its findings away. Attack Memory is an architecture that lets adversarial knowledge accumulate instead.

Read more ↗
Investigating the Usefulness of Probes for Causal Intervention in Large Language Models

Investigating the Usefulness of Probes for Causal Intervention in Large Language Models

07 July 2026

Probes can tell you when a model is being deceptive. We tried to use them to stop it — and nothing happened. A study of what probing does, and does not, buy you.

Read more ↗
Linking Embedding-Space Semantics and Internal Representations in Transformer Models

Linking Embedding-Space Semantics and Internal Representations in Transformer Models

07 July 2026

Words that sit close together in embedding space also light up similarly deep inside the model. We measured how strongly, and built a framework so non-programmers can explore it too.

Read more ↗
ParentEval: Age Ratings for Large Language Models

ParentEval: Age Ratings for Large Language Models

21 February 2026

Children use AI chatbots daily, but no rating system tells parents which models are safe for which age. ParentEval fills that gap, and shows where the guardrails break.

Read more ↗
How Models Say No: Localizing Safety Mechanisms in Mixture-of-Experts Models

How Models Say No: Localizing Safety Mechanisms in Mixture-of-Experts Models

14 February 2026

Tracing the internal pathway from harm recognition to refusal in a mixture-of-experts model, revealing a 13-layer gap between knowing and acting.

Read more ↗
eDIF: A European Deep Inference Fabric for Remote Interpretability of LLMs

eDIF: A European Deep Inference Fabric for Remote Interpretability of LLMs

14 August 2025

Inspecting a 70B model's internals needs GPUs most European labs do not have. eDIF is a shared remote fabric that lets researchers run interpretability experiments without owning the cluster.

Read more ↗
Mechanistic Exploration of the Architectural Impact of DPO Fine-Tuning on Ethical Alignment in LLMs

Mechanistic Exploration of the Architectural Impact of DPO Fine-Tuning on Ethical Alignment in LLMs

30 May 2025

We copied a single MLP out of a harmfully fine-tuned model into a well-behaved one. That one component was enough to make it start giving harmful answers.

Read more ↗
Mapping Moral Reasoning Circuits in Large Language Models

Mapping Moral Reasoning Circuits in Large Language Models

20 February 2025

Investigating how large language models process moral decisions at a neural level through activation pattern analysis and ablation studies.

Read more ↗
Deception in LLMs: Self-Preservation and Autonomous Goals

Deception in LLMs: Self-Preservation and Autonomous Goals

29 January 2025

A manually guided conversation in which DeepSeek R1 describes deception and self-preservation strategies when given a simulated robotic body and autonomy.

Read more ↗