Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
Circuits that match GPT-2's correct answers often fail to reproduce most of its errors.
The paper shows that circuit explanations validated by ablation can closely reproduce a model's successful decisions while failing to account for most of its errors. On indirect object identification with GPT-2 small under mean ablation, manual and automated circuits agree on 97.3-99.5% of prompts the model answers correctly but only 11.4-41.7% of its errors. Restoring omitted attention heads raises error reproduction from 14.2% to 75.1% on a held-out set, with a 0.41 percentage point drop in correct agreement. The authors argue exact error reproduction should be a necessary test of circuit explanations, based on IOI, Docstring, and the Mechanistic Interpretability Benchmark.
- Ablation-validated circuits can match successes while missing most errors.
- GPT-2 small IOI circuits agree on 97.3-99.5% of correct answers.
- Those circuits reproduce only 11.4-41.7% of the model's errors.
- Restoring omitted attention heads lifts error agreement to 75.1%.
Full article263 words · extracted from arxiv.org · click to collapse
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35686