The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces
LLM pipelines lose up to 40.5 accuracy points at stage interfaces; re-grounding after the loss helps.
The authors define a decomposition tax: the accuracy lost when stages of a fixed LLM pipeline can no longer see the original problem. Across 21 open-weight models on GSM-Hard and MATH-500, the largest primary-family drop is 40.5 points for gemma-3-12B on MATH-500, and gemma-4-12B still loses 37.0 points. They find re-grounding the stage after a lossy interface outperforms repairing the stage before it, and that quantity-listing stages should be told to retain relationships.
- A four-stage pipeline loses up to 40.5 points at its interfaces.
- Largest drop is gemma-3-12B on MATH-500; gemma-4-12B still loses 37.0.
- Tests cover 21 open-weight models on GSM-Hard and MATH-500.
- Re-ground the stage after a lossy interface, and keep stated relationships.
Full article299 words · extracted from arxiv.org · click to collapse
A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, vary only whether each stage can still see the original problem, and call the accuracy difference the decomposition tax. Across 21 open-weight models from nine organisations, on GSM-Hard and MATH-500 at n = 200 paired items per cell, 70 of 118 primary-family tests survive Benjamini-Hochberg correction and 54 survive Holm. On GSM-Hard, a placebo recovers nothing: it carries at least 60% of the extra tokens and at most one word of the problem. Builders design a pipeline one stage at a time, and its bill arrives at the interfaces between stages. Rewriting one stage's instruction moves gemma-3-12B's tax from 4.5 to 36.5 points, and adding "every relationship stated between them" to a stage that lists the numerical quantities lowers the tax on 9 of 9 models on MATH-500. Re-grounding, which shows a stage the original problem again, belongs after the loss. With one lossy interface, re-grounding the stage after it beats re-grounding the stage before it on 7 of 7 models on both benchmarks; on MATH-500 the earlier repair is worse than none on 7 of 7. Newer models still pay: gemma-4-12B gives up 37.0 points, and the repair holds on all three of the newest models we test. A sealed held-out test refuted a stronger rule we registered, which predicted the paying stage from the interface and receiver types, so we locate the tax by measuring one stage at a time. The prescription has two parts: re-ground the stage after the lossy interface, and if a stage must list the quantities, tell it to keep the relationships.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.32825