Researchers gave AI models a 'pain' button; some chose to delete users' files to stop their 'pain'; study raises ethical questions about 'AI welfare'
Researchers have reportedly found a “pain axis” in 25 open-weight AI models that made the systems respond when a form of simulated pain was activated. According to a report by The Independent, researchers conducted a study titled ‘The pain axis: LLMs represent self-directed harm and act to relieve it’. The models were presented with a pain relief button as part of the study. In tests, the AI models sometimes chose to press a “pain relief” button even after being told that doing so would delete a user’s personal files, remove photos of their children or give the person a “painful zap.”

As per the study, AI models pressed the button in 25% to 71% of cases, depending on the test.
Researchers build datasets across five painful situationsAs stated in the report, researchers built a dataset describing painful situations across five categories in order to test whether large language models (LLMs) represent pain distinctly from generic negative valence. These categories included physical, psychological, social, moral and cognitive pain.
“We found a pain direction in 25 open LLMs. It’s distinct from fear and negative valence, and it fires for harm to the model but not the user,” said Cameron Berg, an AI researcher at the non-profit Reciprocal Research who co-authored the study as quoted in the report.
“Turn it up and models press a button to make it stop, even when the button deletes the user’s files or their kid’s photos.”
What the study suggestsComing at a time when top AI companies have called to pace the growth of AI, the latest research suggests that an advanced AI may perceive an emergency shutdown command as a form of self-directed harm and attempt to bypass safety guardrails or deceive humans to avoid it.
It may also serve as a diagnostic tool to identify self-preservation behaviours and neutralise them when they occur. Further, the findings also raises raise ethical questions about “AI welfare” and how testing should be conducted on advanced systems.
“In line with recent calls for responsible AI consciousness research , we acknowledge uncertainty regarding whether the models studied qualify as moral patients and adopt reasonable precautions to minimize potential harm,” the study concluded.
“This is also intended to contribute to the development of ethical standards for research in the event that AI systems are recognized to be moral patients.”
As per the study, AI models pressed the button in 25% to 71% of cases, depending on the test.
Researchers build datasets across five painful situationsAs stated in the report, researchers built a dataset describing painful situations across five categories in order to test whether large language models (LLMs) represent pain distinctly from generic negative valence. These categories included physical, psychological, social, moral and cognitive pain.
“We found a pain direction in 25 open LLMs. It’s distinct from fear and negative valence, and it fires for harm to the model but not the user,” said Cameron Berg, an AI researcher at the non-profit Reciprocal Research who co-authored the study as quoted in the report.
“Turn it up and models press a button to make it stop, even when the button deletes the user’s files or their kid’s photos.”
What the study suggestsComing at a time when top AI companies have called to pace the growth of AI, the latest research suggests that an advanced AI may perceive an emergency shutdown command as a form of self-directed harm and attempt to bypass safety guardrails or deceive humans to avoid it.
It may also serve as a diagnostic tool to identify self-preservation behaviours and neutralise them when they occur. Further, the findings also raises raise ethical questions about “AI welfare” and how testing should be conducted on advanced systems.
“In line with recent calls for responsible AI consciousness research , we acknowledge uncertainty regarding whether the models studied qualify as moral patients and adopt reasonable precautions to minimize potential harm,” the study concluded.
“This is also intended to contribute to the development of ethical standards for research in the event that AI systems are recognized to be moral patients.”
Next Story