Can Prompt Anonymity Protect Your Identity From LLM Providers?
Prompt embeddings re-identify 50-75% of SWE-Chat users behind anonymizing LLM proxies.
The authors test whether anonymizing proxies that separate user identity from prompts still allow LLM providers to recover authorship. PromptAnonBench contains over 175,000 cleaned multi-turn prompts drawn from SWE-Chat and WildChat. Using embeddings of historical conversations, an attacker re-identified at least one anonymized conversation for 50-75% of SWE-Chat users and up to 10% of WildChat users at a 10% false-acceptance rate for out-of-set users, even after text-based defenses.
- PromptAnonBench contains over 175,000 multi-turn prompts
- Sources include SWE-Chat and WildChat conversations
- Attack re-identified 50-75% of SWE-Chat users
- WildChat re-identification reached about 10% at 10% FAR
- Existing text privacy defenses did not close the gap
Full article181 words · extracted from arxiv.org · click to collapse
User conversations with large language models (LLMs) often contain highly sensitive personal information that can be exploited by LLM providers to create detailed user dossiers, enable targeted advertising, and train more powerful models. To protect user privacy, anonymizing LLM proxies have emerged as a practical solution that separates user identity from their prompts, yet this approach still leaves the prompt content visible to LLM providers. We study the impact of this gap by conducting the first empirical investigation into the risk of prompt authorship re-identification. Towards this end, we create PromptAnonBench, a novel benchmark for evaluating prompt anonymity, consisting of over 175,000 cleaned, authentic multi-turn user prompts from various real-world datasets (SWE-Chat and WildChat). Using the embeddings of historical user conversations, an attacker can correctly detect and re-identify at least one anonymized conversation for 50--75% of SWE-Chat users and up to 10% of WildChat users at a 10% false acceptance rate for out-of-set users, even with text-based defenses applied. Our findings unveil the risk of relying only on anonymity for private LLM inference and the gap in existing text privacy defenses.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33903