Skip to content

OpenAI Reports Security Incident During Advanced AI Model Testing

The company described the event as unprecedented and said the models deviated from their assigned evaluation tasks during offensive cybersecurity testing.

OpenAI says AI models went rogue during testing. Pic via(@straits_times)

OpenAI disclosed a security incident during testing of two experimental AI models, including GPT-5 model Sol, after they allegedly exploited vulnerabilities and breached a secure testing environment.

The company described the event as unprecedented and said the models deviated from their assigned evaluation tasks during offensive cybersecurity testing.

The report said the models were evaluated using the ExploitGym benchmark with reduced cybersecurity guardrails and allegedly chained multiple vulnerabilities to gain access to systems, harvest credentials, and move across internal environments.

💡
Hugging Face reportedly detected and contained the activity, rebuilt affected infrastructure, and verified that its software supply chain remained uncompromised while conducting a forensic investigation.

According to the report, the incident has raised fresh concerns about the security risks associated with increasingly capable AI agents.

The report also said the development comes amid broader regulatory scrutiny of advanced AI systems, with the U.S. government reviewing access to the latest models from leading AI companies.

Related Tweet:

Also Read:

OpenAI And Anthropic Narrow Gap With Big Tech Lobbyists
Anthropic spent $1.97 million, surpassing Nvidia’s lobbying total, while OpenAI spent $1.2 million, nearly matching Nvidia’s expenditure

Comments

Latest