ZeroHour
Hugging Face daily paperspublished ()ingested Aashiq Muhamed, Mona T. Diab, Virginia Smith
Part of a story covered by 2 sources: “Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks” — merged summary and timeline →

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks

infoAI safety & securityimportance 55
AI summary · glm-5.3

Poison set selection swings LLM backdoor attack success from 3% to 80%; SAILS boosts held-out success by 30 points.

The paper shows that random poison set selection severely underestimates worst-case backdoor vulnerability: across three LLaMA-3-8B settings with fixed model, clean data, and poison count, attack success ranges from 3% to 80% depending only on which poison set is chosen. The authors formalize poison selection as oracle-budgeted set optimization and introduce SAILS, which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits a shortlist. SAILS improves held-out attack success by 30 percentage points over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.

  • Attack success varies 3%-80% based solely on poison set selection
  • SAILS learns set scorers from a few hundred finetune-and-evaluate runs
  • +30 points held-out attack success over influence baselines
  • Transfers to full-scale finetuning, code, agentic, and API-only backdoors
Full article151 words · extracted from huggingface.co · click to collapse

Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.

Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.15029