The Privacy Fallacy of Crowdsourced Fine-Tuning: Extracting Proprietary Data via Topic-Based Poisoning
Fifty poisoned crowdsourced examples can more than triple near-verbatim extraction of other users' fine-tuning data.
The paper shows a malicious contributor can poison a small share of crowdsourced supervised fine-tuning data to amplify extraction of other users' unseen instructions. With only black-box, output-only access, 50 poisoned examples raised near-verbatim extraction to 3.71 times the baseline for Qwen2.5-14B on OpenMathInstruct and 3.08 times for Llama-3.1-8B on AceReason. Experiments covered four models and two datasets. Data filtering was largely ineffective: the best method scored 0.378 F1 and left most poisoned samples undetected.
- Fifty poisoned samples raise near-verbatim extraction by up to 3.71 times.
- Qwen2.5-14B on OpenMathInstruct and Llama-3.1-8B on AceReason were tested.
- The best data filter scored only 0.378 F1 and missed most poison.
- The attack needs only black-box, output-only access to the deployed model.
Full article181 words · extracted from arxiv.org · click to collapse
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also allows untrusted users to contribute data to the fine-tuning pipeline. We investigate an underexplored privacy risk arising from this setting: can a malicious user poison a small fraction of the crowdsourced data to amplify extraction of previously unseen instructions contributed by other users? We show that this is possible using only black-box, output-only access to the deployed model. Experiments across four models and two datasets demonstrate substantial increases in training-data extraction: with only 50 poisoned examples, near-verbatim extraction reaches $3.71\times$ the rate without poisoning for Qwen2.5-14B on OpenMathInstruct and $3.08\times$ for Llama-3.1-8B on AceReason. Data filtering also proves largely ineffective in detecting poisoned samples: even the best-performing method achieves only 0.378 in F-1 score, leaving the majority of poisoned samples undetected. These findings demonstrate that seemingly benign crowdsourced contributions can amplify leakage of other records while remaining difficult to identify through data filtering.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33985