Backdooring Sparse Autoencoders
Malicious sparse autoencoders can backdoor language-model behavior without changing the model.
The paper shows sparse autoencoders used to interpret or steer language models can carry a supply-chain backdoor. A decoder-only modification, with the LLM and SAE encoder frozen, induces attacker-chosen behavior at a single insertion layer. On code generation, the backdoor produces high rates of unsolicited code insertion across three models and many layers, including trigger-dependent behavior. HumanEval and SAEBench checks find strong attacks can coexist with relatively small drops in conventional SAE quality metrics.
- A malicious SAE decoder can change behavior of an otherwise unmodified language model.
- The LLM and SAE encoder stay frozen; only one insertion layer is changed.
- Code-generation tests show unsolicited code insertion and trigger-dependent behavior.
- Strong backdoors can leave several SAE quality metrics only slightly changed.
Full article173 words · extracted from arxiv.org · click to collapse
Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, restricting the attack to a single auxiliary component at a single insertion layer. Using code generation as a case study, we demonstrate high rates of unsolicited code insertion across three language models and a wide range of insertion layers, as well as trigger-dependent behavior conditioned on a prompt cue. We further evaluate the modified SAEs using HumanEval and selected SAEBench metrics. While attack effectiveness varies across models and layers, strong backdoor behavior can coexist with relatively small changes in several conventional SAE quality measures. These results establish that SAEs can carry behavioral backdoors without modifying the language model itself and should therefore be treated as security-sensitive components.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06049