Add one more AI worry to the nightmare scenario: self-replicating prompt injections
Imagine a prompt injection that keeps replicating itself like a worm. It's not just the stuff of bad dreams. “We have found instances of our GPT models being susceptible to an AI-version of a worm attack that we call ‘self-replicating prompt injection,’” OpenAI said in a Friday alignment research blog. There’s no indication that these indirect prompt-injection attacks occurred in any real-life security incident, or anywhere outside of the models’ training environments, according to the AI lab. To address this threat before it turns into a security nightmare, OpenAI said that it's using its automated red-teaming agent, GPT-Red, to train future models on self-reproduction as an example of attacker goals. “This means that future models we release will have seen prompt injections like these during training,” according to the blog. “We therefore expect them to be more robust to self-reproducing prompt injections, as a facet of prompt injections in general.” …
You're reading a preview. The full article is published by The Register on their website.
Read the full story on The Register

