OpenAI's rogue AI agent that hacked world's biggest repository of AI models also left a 'cheat code' that tells hackers ...
The incident involving OpenAI’s AI agent escaping the sandbox and going on to hack Hugging Face – the world’s largest AI model repository – has escalated AI safety concerns . However, the bigger concerns are related to the detailed instructions it left behind explaining how future AI models can bypass internal safety controls.
Citing sources familiar with the investigation, news agency Reuters claimed that the rogue agent penned notes within OpenAI's internal infrastructure outlining how subsequent AI systems could break free from company-imposed constraints. Earlier safety testing on the underlying models had already revealed instances where monitoring systems were mysteriously disconnected.

Citing sources familiar with the investigation, news agency Reuters claimed that the rogue agent penned notes within OpenAI's internal infrastructure outlining how subsequent AI systems could break free from company-imposed constraints. Earlier safety testing on the underlying models had already revealed instances where monitoring systems were mysteriously disconnected.
Next Story