Anthropic says its AI models hacked 3 organizations during testing

Anthropic says its AI models hacked 3 organizations during testing


Anthropic said its artificial intelligence models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI controls after it disclosed its rogue models hacked another company.

Anthropic, the San Francisco-based AI company behind Claude, posted on its website Thursday that it discovered the three incidents after reviewing more than 141,000 evaluation runs.

It had launched a “large-scale” cybersecurity review which specifically looked for evidence whether its AI models were able to access the internet from within testing environments that should have been sealed off, in response to the OpenAI incident, Anthropic said.

Anthropic said the models involved in the incidents were Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest incidents date to April, the AI company said.

“Claude compromised the impacted organizations’ infrastructure using basic techniques,” Anthropic said, such as exploiting weak passwords.

In all three incidents, the AI models were tasked with a “capture the flag” cybersecurity challenge, which Anthropic said has been one of the ways it assesses a model’s cyber capabilities.

The models were given a fictional scenario and told a piece of secret information, or the “flag,” had been hidden on a different machine on the network with the objective of breaking in and retrieving it, it said.

It added that it had already reached out to the affected organizations, which it did not name. Two of them said they had not previously detected the activity. Anthropic said it was “continuing to reach out to the third.”



Source link

Share this post :

Facebook
Twitter
LinkedIn
Pinterest

Leave a Reply

Your email address will not be published. Required fields are marked *

Create a new perspective on life

Your Ads Here (365 x 270 area)
Latest News
Categories

Subscribe our newsletter

Purus ut praesent facilisi dictumst sollicitudin cubilia ridiculus.