BiasReducer: Adaptive Bias Mitigation for Reward Models
BiasReducer edits reward-model heads to cut length, confidence, and sycophancy bias without retraining.
BiasReducer is a lightweight framework that edits only a reward model's linear head and chooses relevant edits for each new dataset. A sparse-autoencoder-style encoder identifies sensitivity to superficial attributes such as length and confidence, then the method learns the direction and magnitude of each head adjustment. Across five reward models, BiasReducer-M improved three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, beating two training-based baselines. The gains transferred downstream, reducing unnecessary verbosity and sycophancy while judged quality stayed comparable.
- Edits only the linear reward head instead of retraining
- Detects attribute sensitivity and selects edits per dataset
- BiasReducer-M gains 8.3, 18.0, and 6.9 points on three benchmarks
- Downstream verbosity and sycophancy fall with similar judged quality
Full article238 words · extracted from huggingface.co · click to collapse
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.32720