Fetching from the wire…
Public story · 2026-03-03 · source-backed
Two specialized RL agents — Memory Manager (ADD/UPDATE/DELETE operations) and Answer Agent — fine-tuned with PPO and GRPO. With only 152 training QA pairs, outperforms baselines across three benchmarks. Directly applicable to persistent agent memory systems. arXiv 2508.19828
Each link below shares sources, entities, or timing with this story.
Hybrid on/off-policy RL giving agents non-parametric memory for exploration. 128.6% improvement over GRPO on ScienceWorld, 11.3% on WebShop. Agents generalize to out-of-distribution tasks with "only a few trials with memory and no parameter updates." (arXiv)
New batching algorithms enable ~7x, up to 12x+, longer-context GRPO training with no accuracy or speed penalty versus optimized FA3 and chunked-loss setups (Unsloth Docs). Qwen3-8B GRPO reaches 110K context on one 80GB H100 via vLLM plus QLoRA. For solo builders doing reasonin...
arXiv 2607.28457 is an oracle-free multi-turn RL framework where the model emits a solution plus a discrete correctness verdict and a confidence score each turn, keeping its answer only when the verdict is Correct and confidence clears threshold. Ground-truth correctness shape...
It turns a course brief into finished slides or a self-contained interactive HTML page in one pass (arXiv 2608.30968). Across 220k production requests the median slide takes 17 seconds and an interactive page 59. A hybrid rule-plus-VLM reward drives GRPO, hardened after the te...
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% pea...
RobertGolds1/Gradient, created August 23 and at 423 stars in about a day, Apache-2.0, built on OpenPipe ART. It ships a Research Environment containing a reproducible company workspace of emails, contracts, policies, meeting notes and customer records with evidence deliberatel...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.