Fetching from the wire…
Public story · 2026-02-28 · source-backed
Hybrid on/off-policy RL giving agents non-parametric memory for exploration. 128.6% improvement over GRPO on ScienceWorld, 11.3% on WebShop. Agents generalize to out-of-distribution tasks with "only a few trials with memory and no parameter updates." (arXiv)
Each link below shares sources, entities, or timing with this story.
SkillPyramid extends Voyager-style libraries with a hierarchical topology plus a self-evolution loop, raising average reward 38% and cutting execution steps 27.7% across ALFWorld, WebShop, and ScienceWorld. The takeaway: don't append every successful trajectory as a new flat s...
Two specialized RL agents — Memory Manager (ADD/UPDATE/DELETE operations) and Answer Agent — fine-tuned with PPO and GRPO. With only 152 training QA pairs, outperforms baselines across three benchmarks. Directly applicable to persistent agent memory systems. arXiv 2508.19828
New batching algorithms enable ~7x, up to 12x+, longer-context GRPO training with no accuracy or speed penalty versus optimized FA3 and chunked-loss setups (Unsloth Docs). Qwen3-8B GRPO reaches 110K context on one 80GB H100 via vLLM plus QLoRA. For solo builders doing reasonin...
LLM as post-hoc critic for step-level Q-values. +7.7% WebShop, +13.8% ALFWorld over GRPO. Third paper in the online RL-for-agents cluster this week. arXiv:2603.08754
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
It turns a course brief into finished slides or a self-contained interactive HTML page in one pass (arXiv 2608.30968). Across 220k production requests the median slide takes 17 seconds and an interactive page 59. A hybrid rule-plus-VLM reward drives GRPO, hardened after the te...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.