
Don’t get me wrong. I’m not unaffected. But I don’t think of AI as the problem.
Welcome to Safe Mode, your weekly report for pressing security and privacy news—and what steps to take next. Want this newsletter to come directly to your inbox? Sign up on our website!
AI is a tool. Whoever wields the tool sets the agenda. In this case, OpenAI ran a benchmark specifically to evaluate how well its new models can find and exploit vulnerabilities. And in its own way, the AI agent being tested did exactly that. It found an unknown vulnerability in its controlled environment, broke out to the open web, and hacked into a website called Hugging Face (a code repository for AI developers).
I don’t find it scary OpenAI’s AI agent basically chose to cheat on its test. AI isn’t human. It also doesn’t operate independently, even if marketed as such. As Olivia Buzek, Staff AI Engineer at IBM said in a podcast: “Fundamentally…models by themselves cannot escape containment. They can only do the things that you give them the tools to do. So what that means is, you need to be very careful about what sort of tools you hand it.” In this context, her reference to tools is about the type and level of access developers give to AI models.
Humans make AI. Humans determine how AI is configured. I’m much more concerned about the people developing AI. They are learning in real time the consequences of automating tasks and processes at dramatically bigger scale and speed. But they seem unprepared to protect the rest of us as mistakes happen.
Instead, AI companies have stayed quiet about the uglier parts of development. Hugging Face brought this breach to light, not OpenAI. OpenAI identified itself as the source five days later. Meanwhile, rival Anthropic just revealed it too has seen Claude hack live websites—sharing after the fact and as OpenAI dominates headlines.
So what scares me is splash damage. I can see a future of consumers dealing regularly with the consequences of human decisions around AI design. We already can’t control the number of attacks on businesses, which lead to data leaks and other online security issues. Matters will worsen dramatically in a world where AI agents run amok, either accidentally or purposefully. AI models can continually hammer at a task without fatiguing.