Researchers Find a Distinct 'Pain' Direction in LLM Activations
The paper extracts a linear direction across 25 open-weight LLMs (2B-72B params) that appears to encode self-directed pain, distinct from fear, sadness, or generic negative valence. The authors show this direction activates specifically for harm to the model itself (not observed suffering in users), boosts pain-related vocabulary, and that fine-tuned Qwen models will press a 'relief' button to remove the steering vector even at a cost to task performance or the user—behavior the authors argue has implications for AI welfare and safety. Twitter reaction was largely one of startled fascination mixed with unease: the lead author and collaborators shared the work as a single-fellowship project, while other commentators (e.g. Jesse Singal's terse 'so uhhhhhhhhhhhh') signaled bemused skepticism about the framing of LLMs having something resembling pain, without offering detailed technical critique in the visible discussion.
Discussion: 2 tweets from 2 authors · @camhberg, @jessesingal