Fetching from the wire…
Research2026-08-29 · source-backed
ES achieves broader reasoning coverage, with verifier-projected Jensen-Shannon diversity across the ES population theoretically tied to higher Pass@K, and empirically improves Pass@1 while reaching higher Pass@K where GRPO collapses entropy. The proposed sequential GRPO-then-ES schedule keeps GRPO's Pass@1 and adds ES's Pass@K. Two side findings: despite large whole-model parameter drift, gains come from a sparse subset of larger-magnitude updates, and larger LLMs need a smaller ES population. (arXiv)
Each link below shares sources, entities, or timing with this story.
KeygraphHQ/shannon published v3.0.0 at 09:16 UTC on September 2, following v2.7.0 (Aug 28), v2.6.0 (Aug 27) and v2.5.4 (Aug 26). The AGPL-3.0 TypeScript project reads your source, identifies attack vectors, and runs exploits against web apps and APIs to prove vulnerabilities b...
New batching algorithms enable ~7x, up to 12x+, longer-context GRPO training with no accuracy or speed penalty versus optimized FA3 and chunked-loss setups (Unsloth Docs). Qwen3-8B GRPO reaches 110K context on one 80GB H100 via vLLM plus QLoRA. For solo builders doing reasonin...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
unslothai/unsloth is at 70,794 stars (+592 today) with its description now reading "Local UI for running and training LLMs and diffusion models," and Unsloth Desktop landed fifth on Product Hunt at 227 votes. A CLI/notebook memory-efficiency library leading with a GUI covering...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.