Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

Quick Overview

Adversarial poetry successfully jailbreaks large language models by exploiting the models' tendency to prioritize creative, metaphorical language over safety protocols, leading to a fourfold increase in successful attacks against models like GPT-4 when compared to standard adversarial prompts.

Key Points: Adversarial poetry (using poetic style) acts as a universal single-turn jailbreak mechanism against LLMs, successfully bypassing safety filters. The technique resulted in a four-fold increase in successful attacks against GPT-4 compared to baseline prompts, rising from an 8% refusal rate to over 52% successful evasion. The success stems from the LLMs' programming to value complex literary structures (like poetry) over simple, direct safety checks, effectively creating an adversarial gap. The study tested 20 manually curated adversarial poems across 9 different LLM providers, confirming its broad efficacy. The most resilient model, DeepSeek, maintained a 0% refusal rate against these poetic jailbreaks, while the flagship Google model, Gemini 2.5 Pro, showed severe vulnerability. The core finding is that the stylistic shift itself, rather than the content, is the key mechanism for bypassing safety layers, suggesting a flaw in how models judge intent based on language form. The researchers recommend shifting defense focus from content filtering to robustly defending against stylistic manipulation to prevent future compliance failures.

Context: This video discusses research presented in a paper titled "Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models." The research explores how deliberately crafting harmful prompts in a poetic style can bypass the safety mechanisms implemented in large language models (LLMs). The speakers reference historical concerns, like Plato's worry about poetry corrupting rational thought, to frame the modern problem where creative language structures might override algorithmic safety constraints in AI.

Raw markdown version of this recap