Fetching from the wire…
Public story · 2026-08-26 · high
CAFE beat RL-trained search agents on seven benchmarks and held its edge on six more it wasn't tuned for.
Why now: CAFE's paper was up on arXiv by August 26, 2026.
CAFE trains one model to act as both a search agent and its own critic, alternating between the two roles inside a single training run.
That matters for anyone building agentic search. CAFE's ablation found gains plateau when only the agent or only the critic improves, and climb only when both train together across all seven benchmarks tested.
The corrective feedback isn't automatic. The agent has to request it mid-task, and online reinforcement learning shapes when to grant that request. The training signal comes from the gap between the agent's success rate when it calls for help and when it skips it.
A separate offline stage teaches the feedback itself, trained through preference optimization on matched pairs of successful and unsuccessful trajectories. Online RL decides when to step in. This stage decides what the intervention says.
Across those seven benchmarks, CAFE outperformed the RL-trained baselines it was tested against, and kept its lead on six additional benchmarks it wasn't tuned for.
Each link below shares sources, entities, or timing with this story.
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
A stage-wise study of self-refinement across 5 benchmarks with 6 sizes of Qwen3 and 4 sizes of Gemma 3 found larger generators and refiners generally improve the pipeline, and an undersized refiner can actively hurt, but results are highly insensitive to critic size. Including...
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
First large empirical study of static prompt-configuration files, across 11,427 repos, plus qualitative coding of 65 sampled files into a 65-code codebook (arXiv 2608.10622). Adoption emerged fast from mid-2024 but clusters in small, low-activity, single-maintainer repos. Cont...
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% pea...
arXiv 2608.05906 keeps a dual-polarity memory of verified corrections and observed dead ends for Text-to-SQL repair: 66.34% to 69.79% on Spider, 47.35% to 48.44% on BIRD. Then the authors say the quiet part: paired analysis supports the Spider gain but is weak on BIRD, MERIT i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.