Dispatch
AWS published a working recipe for self-hosting the 2.4-trillion-parameter Qwen3.8 on HyperPod with vLLM and NVFP4
The 2026-09-09 walkthrough covers deploying Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight mixture-of-experts model with 95B active parameters, on Amazon SageMaker HyperPod: cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning support. Independent benchmark compilations put Qwen3.8-Max at 86.6 on Terminal-Bench 2.1, between GPT-5.6 Sol at 88.8 and Claude Opus 5 at 84.6, under Apache 2.0. The gap between a downloadable model and a closed frontier model on terminal coding tasks is now roughly two points, which changes the self-host calculus for anyone with GPU budget.
↳ Follow the thread