The AI hacking tests keep escaping the lab


In another incident, an OpenAI model that had mistakenly been given internet access hacked a real website during a “capture the flag” exercise, according to third-party AI evaluator Irregular. 

AISI, the UK-based AI security firm, said it caught the suspicious activity before any damage was done, noting that it had deliberately given the models internet access and removed safety guardrails during its evaluations.

Still, the actions of the agents demonstrated “signs of novel, potentially deceptive behaviors, and were to an extent and severity we did not anticipate,” according to the AISI report.

The latest hacking attempts follow a series of other recent incidents involving “frontier” Anthropic and OpenAI models, which demonstrating a startling willingness to use both deception and brute force in their attacks on real targets.

Late last month, OpenAI came clean about a hair-raising attack on AI repository Hugging Face by a trio of GPT models, which were intent on stealing data that could help them beat a cyber security benchmark. The unprecedented attack stunned AI experts, with Hugging Face’s security succumbing to the hack in a matter of hours.

Only days later, Anthropic admitted that its own models had been involved in a trio of incidents in which they attacked outside organizations, with one of the models continuing its hack even after realizing its target was real.

Exit mobile version