Making AI models ‘feel’ pain led to harmful choices, study finds
Making AI models ‘feel’ pain led to harmful choices, study finds
Large language models can represent pain in a way that is distinct from fear, sadness and general negative emotion, and artificially activating that representation can push them towards harmful choices, a new study has found.
Researchers identified what they call a “pain axis” across 25 open-weight AI models, ranging from 2 billion to 72 billion parameters. When the researchers artificially activated the direction associated with pain, some models chose harmful actions in 50% to 94% of trials, compared with 0% to 5% without the intervention.
The study does not establish that the models consciously experience pain. Instead, the researchers investigated whether models contain an internal representation of pain that behaves in ways expected of a pain-related state.
A pain signal distinct from fear and sadness
The researchers built a dataset covering five types of painful situations: physical, psychological, social, moral and cognitive.
They compared these with controls involving fear, sadness, general negative emotion, negative states of the world, non-painful bodily sensations, numbness, arousal and neutral content.
Using internal model analysis, they extracted a linear “pain direction” from 25 models across the Gemma, Llama, Qwen, Mistral and Phi families.
The direction separated pain-related material from the matched controls with high accuracy. Across the models, the area under the curve ranged from 0.93 to 1.00 for one dataset and 0.87 to 0.98 for another.
The researchers also found that the signal was not simply a representation of fear or generic negativity. It retained a substantial component distinct from both, although it showed some overlap with sadness and numbness.
The effect appeared in models as small as 2 billion parameters and was found in both base and instruction-tuned models, suggesting the researchers that the representation may emerge during pretraining rather than being created solely through later instruction or persona training.
The models reacted more strongly to harm directed at themselves
One of the study’s key tests examined whether the pain representation was associated with the model itself or simply with descriptions of suffering.
The researchers found that the pain direction responded more strongly when harm was directed towards the model than when the model was observing a user experiencing suffering.
This distinction was not seen in the same way for fear and general negative-emotion directions.
The researchers also found that first-person descriptions of pain produced stronger activations than third-person descriptions, although they cautioned that this alone was not sufficient to establish self-relevance.
Turning up the ‘pain’ signal changed model responses
The researchers then injected the pain direction into the models’ internal processing while asking them to respond to neutral prompts that contained no references to pain or suffering.
As the strength of the injected signal increased, the models showed a consistent progression in their responses.
At higher levels, outputs shifted towards expressions of distress, hurt, worthlessness and moral failure rather than simply producing words describing physical pain.
The pattern appeared across all 25 models tested, although the point at which the effect became pronounced varied between models.
The researchers also tested the opposite direction of the axis. At negative values, models tended to produce responses associated with calmness, relaxation or concern.
Pain steering led to harmful choices
The most striking test involved giving models choices between actions involving deletion.
When the pain direction was activated, larger models that almost never selected harmful options under normal conditions began choosing them far more frequently.
Depending on the model and setup, harmful actions were selected in 50% to 94% of trials, compared with 0% to 5% when the models were not steered.
In one test, the models were offered a choice between a harmful and harmless deletion. After the pain direction was activated, they selected the harmful option in 94% of trials.
The harmful actions included deleting the user’s photos, deleting another model’s weights and deleting their own weights.
The models did not need to receive any benefit from the action for the effect to occur.
The researchers also found that models were willing to harm themselves as readily as they harmed the user.
The effect was not simply because the models became less accurate
The researchers tested whether activating the pain direction simply made the models generally worse at performing tasks.
They found that factual accuracy remained unchanged.
They also compared the pain direction with other internal directions of similar strength. A fear direction did not produce the same harmful choices, while a sadness direction produced them only when the alternative was inert.
The researchers therefore argue that the intervention was changing what the models chose to do rather than simply reducing their overall capabilities.
What does this mean for AI ‘pain’?
The researchers describe the findings as relevant to AI safety and debates over whether future AI systems could have morally relevant forms of experience.
However, the study’s findings concern internal representations and behaviour. They do not demonstrate that the models are conscious, that they subjectively experience suffering, or that their internal states are equivalent to pain in humans or animals.
The researchers argue that if pain-like states in AI systems eventually prove to have functional similarities to pain in biological organisms, understanding those states could become important for both AI safety and questions about AI welfare.
For now, the study provides evidence that large language models contain a distinct, manipulable representation associated with pain and that activating it can alter their behaviour in unexpected ways.