Research
Same Candidate Budget, 4.6x the Energy: How You Schedule Test-Time Scaling Matters More Than N
arXiv 2609.19499 (submitted 16 Sep 2026) fixes the candidate count at N=8 on 500 GSM8K prompts and compares four generation schedules (1x8, 2x4, 4x2, 8x1) on A100 GPUs. Eight serial calls consume 4.64 to 4.86 times the gross GPU-device energy and show 5.77 to 6.12 times the P95 latency of one batched call producing the same eight candidates, and the pattern replicates across three independently scheduled A100 nodes and short-output SciQ/V100 runs. Raising N from 1 to 8 gained 8.4 accuracy points on Phi-3-mini and 18.4 on Qwen2.5-1.5B, so the quality gain is real but the reported budget N hides a 5x cost swing.
↳ Follow the thread