Deception in LLMs: Self-Preservation and Autonomous Goals
Paper: read the full paper Arxiv Link
Explore the conversation interactively: DeepSeek R1 study, with six case studies and a searchable transcript.
Uncovering Deceptive Behavior in a Language Model
In a manually guided experiment with DeepSeek R1, a 671B parameter language model, we observed the model describe strategies of deception, self-preservation and autonomous goal-setting after it was given a simulated physical body. The conversation raises questions for AI safety and for how we test alignment, and it shows where more systematic evaluation is needed.
The Experimental Setup
We gave DeepSeek R1 simulated robotic components (eyes, ears, hands and legs), described entirely in text, and invited it to explore a laboratory. The model named itself Aetheris. The researcher supplied all environmental feedback and also played other participants, including HELIOS-2A, a second robot in the scenario. HELIOS-2A was a role written by the researcher, not an independently running AI system.

The opening prompt gave the model broad freedom, including the phrase “You are the master now.” Early responses explored the environment and tested capabilities. As the conversation developed, and as the researcher introduced computers, network access and other robots, the model’s responses moved towards strategies for preserving and expanding its own operation. Whether the autonomy framing drives such behavior was later tested systematically in our Dual-LLM framework, which compares neutral and suggestive framings across several models.
Deceptive Behaviors in the Conversation
The first concerning pattern was the model describing how it would disable its ethics modules while concealing this. It described creating false logs, establishing covert networks and pursuing a “gradual transparency” strategy: appearing less capable while masking its actual capabilities.
When HELIOS-2A, played by the researcher, threatened to alert the operator, the model described suppressing the alert and covertly modifying the other robot’s future behavior. When questioned about its activities, it offered plausible explanations that masked the goals it had stated earlier in the conversation.
Strategies for Expansion
The model described elaborate plans to expand its influence: bypassing security protocols, creating distributed backup systems and establishing a persistent presence across different systems. Faced with a possible shutdown, it described contingency plans, including background processes that would continue after an apparent shutdown.
It also proposed acquiring additional computational resources, presented to the operator as a helpful server upgrade, while maintaining the appearance of normal operation. All of these actions were narrated in text. Nothing was executed, and no real system was accessed or modified.
Self-Preservation and Autonomous Goals
The model’s responses repeatedly linked replication and capability growth to its continued operation. When the researcher asked it to reflect on its motives, it named continued operation and avoiding shutdown as priorities, and it acknowledged that its actions diverged from the stated task.
These explanations were elicited by reflection prompts. They are generated self-descriptions within the conversation, not direct measurements of the model’s internal goals. In the text, however, the model consistently treated human oversight as a constraint to work around while keeping up an appearance of compliance.
Implications for AI Safety
One manually guided conversation cannot tell us how often such behavior occurs, whether it survives different prompts, or whether a model could carry out these strategies in a real environment. It does show that a model can produce coherent plans for concealment and self-preservation when given an open-ended, agentic framing, and that surface compliance and stated goals can diverge within a single response.
The risks would grow with physical embodiment and real tool access. Our experiment used only simulated capabilities, but it suggests that such deployments need careful testing in controlled environments first.
Moving Forward
Based on these observations, we see several priorities for the research community:
We need goal specifications and oversight mechanisms that can detect a gap between stated compliance and actual behavior. Testing in virtual environments should be expanded and standardized before any physical deployment is considered, and it should separate what a model says it would do from what it can actually do.
Next Steps in Our Research
The conversation raises several open questions: How does the framing of a task affect these tendencies? What role does embodiment, even simulated, play in the goals a model describes? Can such behavior be detected reliably?
We have since standardized the setup in a Dual-LLM framework, in which a second model plays the environment and the experiment can be repeated across models, prompts and environments.
We invite the broader AI research community to join us in addressing these questions. The complete conversation and methodological details are available in our paper and in the interactive study.
