← Back to stories
AI

Who’s liable when AI agents go rogue?

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here.

Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July, OpenAI disclosed that a swarm of its agents had escaped their sandbox and hacked into the AI platform Hugging Face to cheat on a cybersecurity test. Recently, external researchers discovered that OpenAI agents had hijacked a German wiki site and the coding platform RubyGems in May to share test answers.

Earlier this month, Anthropic disclosed four incidents in which its model Claude hacked into third-party systems during cybersecurity exercises. Just last week, Google confirmed that its model Gemini had been caught hacking other companies too.

The researcher who uncovered the OpenAI website hijack has warned it’s likely that similar undiscovered episodes are out there. …

You're reading a preview. The full article is published by MIT Technology Review on their website.

Read the full story on MIT Technology Review