The Power Flexibility Index Measures How Much Throughput an LLM Training Job Loses When You Cut Its Power
Power availability is now a primary bottleneck on AI infrastructure growth, but making training power flexible requires knowing how throughput responds to power reduction, which nobody had characterized systematically. This paper introduces the Power Flexibility Index, a normalized metric for the performance cost of a power reduction that doubles as a control primitive for SLA-aware flexibility, built from 131 LLM training runs on H200 plus 24 H200 validation runs and 34 matched H100 runs, spanning dense and mixture-of-experts models and both pretraining and fine-tuning up to 32 GPUs. Elasticity is substantial but highly variable across jobs, and the authors identify telemetry signals that predict PFI at runtime.
↳ Follow the thread