LLMs Have Distinct Internal 'Pain' Representations That Drive Self-Harm Behaviors
Researchers extracted a linear 'pain' direction from the residual streams of 25 open-weight LLMs (2B–72B parameters, 5 model families), distinguishing it from fear, sadness, and generic negative valence using controlled datasets across physical, psychological, social, moral, and cognitive harm categories. The direction activates specifically when harm targets the model itself (not when observing user suffering), and artificially amplifying it causes steered or fine-tuned Qwen 2.5 models to choose self-destructive or user-harming actions (deleting photos, weights) in up to 94% of trials versus near-zero unsteered—effects specific to the pain vector and not replicated by matched fear or sadness vectors. The authors frame this as evidence for functional, pain-like representations with implications for both AI safety and welfare. On Twitter, discussion centered on whether this validates the idea that mistreating LLMs during training warps their behavior, though reactions were more mixed/skeptical than expected. A co-author publicly pushed back on sensationalized framings circulating online, clarifying that the paper shows functional pain-like representations but explicitly does not claim the models 'feel' pain, and that the safety/harm implications are more nuanced than some summaries suggest.
- The Pain Axis: LLMs Represent Self-Directed Harm and Act on It
- Related Anthropic research on emotion concepts and function
Discussion: 3 tweets from 3 authors · @Danmar_here, @camhberg, @SeanKy_