Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Study shows LLM test-time scaling cost depends on generation schedule: eight serial A100 calls use ~4.7x more GPU energy than one batched call.
The paper shows that candidate count N alone poorly describes the systems cost of sampling-based test-time scaling. Fixing N=8 with Phi-3-mini and Qwen2.5-1.5B on 500 GSM8K prompts, going from 1 to 8 candidates improved accuracy by 8.4 and 18.4 percentage points respectively, but eight serial generation calls on A100 GPUs consumed 4.64-4.86x more gross GPU-device energy and had 5.77-6.12x higher P95 latency than one batched 8-candidate call. The authors argue evaluations should report generation schedules and GPU-level metrics, and recommend fewer, larger batches when candidates are independent and memory allows.