Jailbreak prompts (i.e. prompts designed to remove or bypass the guardrails and rules that govern AI systems, like LLMs) prompts have been circulating for years. At first, a lot of it was pretty simple: copy a prompt, tell the model to ignore its rules, and see what happens. It was also largely noisy, unverified, and often didn’t work. But the noise was still telling us something. Threat actors were beginning to study AI systems the same way defenders were, and over time, the goal started to change.