ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Zheyuan Wang

Signed Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference

infoAI researchimportance 22
AI summary · glm-5.3-flash

Signed Rescue Routing improves LLM cascade efficiency by predicting when a larger model actually corrects a smaller one rather than uncertainty.

Signed Rescue Routing (SRR) is a budgeted cascade method that separately predicts rescues and regressions when escalating from a small to a large model, ranking requests by their signed difference. The authors prove this signed conditional gain is Bayes-optimal under a fixed escalation budget and add only a lightweight two-head router needing small-model output statistics at deployment. Evaluation with Qwen3-4B and Qwen3-8B on MMLU, HellaSwag, and ARC-Challenge shows better accuracy-compute tradeoffs than entropy routing and learned error predictors.

  • Predicts rescue vs. harm separately instead of small-model uncertainty
  • Signed conditional gain shown Bayes-optimal under fixed escalation budget
  • Deployment requires only small-model output statistics plus a two-head router
  • Tested with Qwen3-4B and Qwen3-8B on MMLU, HellaSwag, ARC-Challenge
Full article185 words · extracted from arxiv.org · click to collapse

Large language model (LLM) cascades answer easy requests with a small model and escalate selected requests to a larger model. Most routers prioritize examples on which the small model appears uncertain or likely to be wrong. This proxy ignores a decisive fact: escalation is useful only when the large model corrects the small model, and it is harmful when the large model replaces a correct answer with an incorrect one. We introduce Signed Rescue Routing (SRR), a budgeted routing method that predicts these two events separately and ranks requests by their difference. We show that this signed conditional gain is the Bayes-optimal routing score under a fixed escalation budget. SRR requires only the small model's output statistics at deployment and adds a lightweight two-head router. We evaluate SRR with Qwen3-4B and Qwen3-8B on TBD examples from MMLU, HellaSwag, and ARC-Challenge. Across the accuracy-compute curve, SRR reaches an area of TBD, compared with TBD for a learned small-model error predictor and TBD for entropy routing. These results show that predicting incremental value, rather than model uncertainty, is a simple and effective objective for efficient LLM cascades.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.07786