Together AI sells preemptible GPU capacity at 50% of on-demand with a five-minute drain
Together AI Blog·medium signal
Together GPU Clusters added preemptible nodes at a flat 50% of the on-demand rate, billed every one to two minutes. On reclaim, the node is cordoned and workloads get SIGTERM with up to 300 seconds to checkpoint. The cluster refills toward its preemptible target on its own. Together warns that a 321B model checkpoint runs about 3.5 TB and can take up to four minutes to write, so this suits chunked batch, experiments and inference bursts, not uncheckpointed multi-day runs.