Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
Gradient jailbreak detection that succeeds on synthetic chats fails on realistic multi-turn conversations.
A controlled study tests whether GradSafe-style gradient jailbreak detection survives multi-turn dialogue. A Context Window Scanner scores fixed windows of user turns and takes the maximum as the conversation score. Against synthetic benign conversations, human-authored multi-turn jailbreaks reach ROC-AUC 0.98, but on WildChat it falls to 0.76 and a synthetic threshold flags more than 90 percent of benign conversations. Successful Crescendo attacks can score at or below benign chats, and Qwen2.5-7B-Instruct yields near-random separability.
- GradSafe was designed for single prompts, not multi-turn assistants.
- Synthetic benign chats score ROC-AUC 0.98; WildChat falls to 0.76.
- A synthetic threshold flags more than 90 percent of realistic benign chats.
- Crescendo can score like benign text, and Qwen2.5-7B-Instruct is near random.
Full article247 words · extracted from arxiv.org · click to collapse
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced gradient and a fixed unsafe reference direction, and their effectiveness in multi-turn dialogue remains unclear. We conduct a controlled evaluation of gradient-based jailbreak detection in multi-turn settings. We extend GradSafe with a Context Window Scanner that applies the detector to fixed-size windows of user turns and uses the maximum window score as the conversation-level score. We evaluate different window sizes, attack families, benign conversation distributions, and target models. The results differ sharply between synthetic and realistic benign settings. Against synthetic benign conversations, the detector achieves an ROC-AUC of 0.98 on human-authored multi-turn jailbreaks. On WildChat benign conversations, ROC-AUC drops to 0.76, and a threshold calibrated on synthetic data flags more than 90% of benign conversations as unsafe. Under realistic benign distributions, single-turn windows give the highest separability, whereas longer windows and accumulated contexts reduce performance. The detector is also sensitive to the attack-generation method and target model: successful Crescendo attacks receive scores comparable to or lower than benign conversations, and Qwen2.5-7B-Instruct yields near-random separability with a different optimal window size. These findings show that gradient-based signals can support multi-turn jailbreak detection, but reliable deployment requires calibration on realistic benign conversations, short-window scoring, length-aware thresholds, and evaluation across attack types and model architectures.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.36849