Tools
Transformers 5.17.0 ships a 780B-parameter MoE with a 1M context and a per-channel-gated linear attention model
Released 2026-09-09, transformers 5.17.0 adds Hy4-Preview: 780B total parameters activating 49B per token, 256 routed experts plus one always-active shared expert, top-8 routing and a 1M-token context. It combines Multi-head Latent Attention, DeepSeek Sparse Attention with shared indexer layers, gated MLA with learnable attention sinks, and Independent Hyper-Connections replacing the plain residual path. The release also lands Moonshot's KimiLinear with per-channel forget gates, Microsoft's VibeVoice long-form multi-speaker TTS, and Alibaba's 800M Fun-ASR-Nano.
Source
↳ Follow the thread