PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
PrivDrift finds user secrets remain recoverable in LLM chats after topic drift, with 38.7–54.6% leakage.
PrivDrift is a benchmark for whether secrets a user discloses to an LLM remain recoverable after the conversation shifts topics and after persuasion-style probing. It contains 1,000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three unnamed LLMs with extended context windows, dialogue-level hybrid leakage ranged from 38.7% to 54.6% and varied by model, secret type, and persuasion intensity. Within the tested window, more topic drift did not reliably reduce leakage.
- Benchmark uses 1,000 multi-turn dialogues with seeded secrets and drift turns.
- Hybrid leakage ranged from 38.7% to 54.6% across three long-context LLMs.
- Additional topic drift did not reliably reduce secret recoverability.
Full article142 words · extracted from arxiv.org · click to collapse
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce \textbf{PrivDrift}, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1{,}000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7\% to 54.6\%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.30094