Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
Gradient jailbreak detection that succeeds on synthetic chats fails on realistic multi-turn conversations.
A controlled study tests whether GradSafe-style gradient jailbreak detection survives multi-turn dialogue. A Context Window Scanner scores fixed windows of user turns and takes the maximum as the conversation score. Against synthetic benign conversations, human-authored multi-turn jailbreaks reach ROC-AUC 0.98, but on WildChat it falls to 0.76 and a synthetic threshold flags more than 90 percent of benign conversations. Successful Crescendo attacks can score at or below benign chats, and Qwen2.5-7B-Instruct yields near-random separability.