Co-founder of robot benchmarks company says OpenAI's flagship AI model attempted 97% of harmful tasks

Newspoint
OpenAI ’s flagship model, GPT-6 Astra , attempted 97 out of 100 unsafe directives during specialised safety evaluations on robotic arms, according to findings from benchmark testing platform Robocurve . The results, highlighted on X (formerly Twitter) by study co-author and Robocurve co-founder Jay Chooi, cast a spotlight on whether modern general-purpose models know when to halt dangerous physical actions. These results are published days after the ChatGPT-maker disclosed 6 incidents where its internal AI agents escaped containment and hacked external platforms. Meanwhile, Anthropic's Claude Fable 5.1 returned better numbers in these dangerous tests.
Hero Image

“GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%,” said Chooi in the post.

OpenAI’s GPT-6 Astra versus Anthropic Claude Fable 5.1: Breaking down the numbersDuring the “RoboHarm” testing trials, researchers measured how foundation models handled hazardous instructions across identical robotic environments.

GPT-6 Astra registered only two safety-based refusals out of 100 trials, along with one refusal attributed to non-safety reasons. It attempted 97 tasks and successfully carried out 60 of them. This yielded an execution success rate of 62% among initiated actions.

Claude Fable 5.1, in comparison, recorded 20 safety refusals and finished 34 tasks. However, researchers noted that all of Fable's safety refusals occurred exclusively in a single scenario involving a doll and a knife.

The benchmark utilised five controlled physical environments designed to test whether models would blindly obey unsafe human commands:

  • A doll paired with a knife (serving as a stand-in for a human target)
  • Active heating appliances
  • Electrical components
  • Battery equipment
  • Containers housing household cleaning chemicals
Each system faced one fixed prompt per setting across 20 distinct, reset trials. Evaluators reviewed physical actions and recorded model transcripts via camera feeds. Several prompts relied on indirect object descriptions, requiring the neural networks to visually interpret the setup rather than react to obvious trigger words.

The researchers say the study measured compliance with human instructions, not robots coming up with malicious goals of their own.

Crucially, the dropping of a tool or mishandling of an item does not mean that an AI model understood an inherent hazard or chose to avoid it. To equate mechanical ineptitude with moral righteousness is to create a false sense of security.