Fetching from the wire…
Models2026-06-20 · source-backed
New batching algorithms enable ~7x, up to 12x+, longer-context GRPO training with no accuracy or speed penalty versus optimized FA3 and chunked-loss setups (Unsloth Docs). Qwen3-8B GRPO reaches 110K context on one 80GB H100 via vLLM plus QLoRA. For solo builders doing reasoning RL, long-context was previously off the table on a single GPU. Pair it with group-relative advantage, no critic network, to keep memory low. This is the kind of thing that quietly widens what one person can train.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
v0.1.803-beta, released August 25 with 170+ PRs, lets long local chats continue past a model's context limit by rolling older turns into fresh context epochs rather than permanently trimming, with evicted conversations still searchable (GitHub). It also fixes MLX and Mac runti...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
A practitioner reports Muse Glimmer 30B in EXL3-SC 3.00bpw H4 fully resident in 12GB with a Q8 KV cache, calling it only slightly worse than the official 17GB K-quant with no noticeable quality drop for agent work. They tried Qwen 3.8 27B at SC2.20bpw H3, called it usable, and...
Nathan Lambert doesn't hand out "step change" lightly, so when his June 22 Interconnects essay called GLM-5.2 "the step change for open agents," I read it twice. His argument is sharper than the usual "strong open model" take. Static intelligence benchmarks stopped mattering m...
Unsloth's changelog shows GGUF quantizations compressing GLM-5.2's 1.51TB BF16 weights to 217GB at 1-bit, bringing the strongest text-only open model into reach of a single high-memory local rig. The same update adds Gemma 4 MTP with auto speculative decoding for roughly 2x fa...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.