Fetching from the wire…
Research2026-07-29 · source-backed
arXiv 2607.25886 isolates data-centric research capability by fixing the entire post-training stack so only the agent's data strategy varies. Four frontier agents across six benchmarks. Among searches that continued past the best observed score, 78.26% ended on a lower-scoring final attempt and the rest merely recovered the peak. The practical read: your self-improvement loop needs explicit checkpoint preservation, not more iterations. Code at github.com/evolvent-ai/RSIBench-Data.
Each link below shares sources, entities, or timing with this story.
Engram, a bi-temporal memory engine, scored 83.6% versus 73.2% for a full-context baseline on the 500-question LongMemEval_S benchmark, a statistically significant +10.4 points, while using ~9.6k tokens instead of 79k (arXiv 2606.09900). Roughly 8x fewer tokens and more accura...
arXiv 2608.04893 tests the "exchanged latent thoughts" claim by replacing the relayed cache with deranged, zeroed and moment-matched random counterparts. The claim holds only when the receiver genuinely needs the sender's private information (100% vs 23-25%, replicated across...
Four stories about things going wrong. Here's one about something working, with actual numbers attached. In an August 7 disclosure covered by TechCrunch, Airbnb said AI now writes 60% of its new code, that concept-to-launch time on key initiatives has dropped by as much as 60%...
CausalMix tries to attribute downstream capability gains to specific data sources instead of grid-searching the mix as a hyperparameter. Data-mix selection is one of the highest-leverage and least-transparent training decisions, so a principled attribution method here is pract...
Here's a stat that should change how you plan your engineering workflow: OpenAI Codex has generated over 400,000 pull requests in two months. Code review agents are no longer experimental. They're routine gatekeepers in development workflows at scale. So researchers did what n...
A code model and test model trained adversarially via RL — each model's failures become the other's training signal. This self-improving loop could break the benchmark dependency bottleneck for training coding agents. Source
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.