Fetching from the wire…
Research2026-06-20 · source-backed
A systematic negative rounding error in non-uniform FP4 (E2M1) formats compounds across layers, degrading training (arXiv:2606.20381). The proposed UFP4 recipe trains on uniform E1M2/INT4 grids with a Random Hadamard Transform on all three GEMMs, and shows lower loss degradation than E2M1 baselines while staying stable on Dense 1.5B, MoE 7.9B, and MoE 124B models. If you track the economics of low-precision pretraining, this is a concrete reason the naive 4-bit path was leaving accuracy on the table.
Each link below shares sources, entities, or timing with this story.
This work pairs E2M1 payloads with unsigned E5M3 block scales whose wider range permits periodic tensor scaling, applies selective stochastic rounding only to backward gradients, drops the Hadamard transform entirely, and uses FP4 in every eligible internal linear. Pretraining...
The authors extract a steering direction from the model's existing tool-use preference signal and apply it at inference, producing monotonic control over how often the agent reaches for a tool while keeping invocations valid (arXiv 2608.25198). Open-domain QA accuracy with liv...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
AGENTQ is the first study of this attack against agents rather than free-text generation, where the payload is a structured function call nobody reads (arXiv 2609.14060). Naive adaptation of prior backdoor methods wrecks benign utility; AGENTQ combines layer-banded LoRA inject...
Power availability is now a primary limit on AI infrastructure growth, but making training power-flexible requires knowing how throughput responds to reduction, which nobody had characterized. The index is a normalized metric for the performance cost of a power cut that double...
Most looped-transformer results compare at fixed model size, conflating architecture with extra compute. SMELT matches per-token FLOPs, non-embedding parameters and KV cache against an unlooped baseline, scaling to 54B non-embedding parameters with a separate Chinchilla-style...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.