Security researchers tricked LLMs into giving them cocaine recipes by abusing role models for prompt injection
Security researchers tricked LLMs into giving them cocaine recipes by abusing role models for prompt injection
www.theregister.com
Security researchers tricked LLMs into giving them cocaine recipes by abusing role models for prompt injection
If you want a picture of the future of LLM security, imagine Whac-a-Mole meets Groundhog Day

cross-posted from: https://sopuli.xyz/post/48033292
"Role tags were a formatting trick that became the security architecture and the cognitive scaffolding of modern LLMs," the authors explain in a blog post. "We've shown that this architecture doesn't survive into the model's actual representations, and that such role confusion is linked to prompt injection."