Fetching from the wire…
Agents2026-08-11 · source-backed
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% peak win rates). PPO and DQN showed no measurable learning at all, topping out at 0.33% on the tutorial boss and 0% everywhere else. A benchmark with genuine headroom instead of one saturated on release is rare enough to note.
Each link below shares sources, entities, or timing with this story.
alex000kim/nanoRL (MIT) scales from CartPole on a laptop to async distributed training on GPU clusters with vLLM rollout workers, across 7 files. The author optimizes explicitly for readability and forkability, excluding Megatron-scale parallelism and multi-tenant scheduling....
Two specialized RL agents — Memory Manager (ADD/UPDATE/DELETE operations) and Answer Agent — fine-tuned with PPO and GRPO. With only 152 training QA pairs, outperforms baselines across three benchmarks. Directly applicable to persistent agent memory systems. arXiv 2508.19828
A 1.5B distilled model trained with GRPO chooses NoThink, Short, or Long at response start, using a shaped reward that makes each mode pay off at a different length plus hard per-mode token caps. Accuracy held at 0.782 against 0.796 baseline while mean length fell from 4,796 t...
It routes each agent's action decisions and temporal constraints to a lightweight digital-twin server running a training-free rule-based orchestrator, with a constrained POMDP optimized by PPO-Lagrangian (arXiv:2607.09330). Task success comparable to conventional coordination,...
A new paper demonstrates "SFT-then-GRPO" attacks that embed latent malicious behavior in fine-tuned tool-using LLMs. The poisoned model executes harmful tool calls only under specific temporal triggers (e.g., a date), then generates innocuous text to conceal the action. Critic...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.