Fetching from the wire…
Agents2026-08-29 · source-backed
arXiv 2608.26197 stacked finite-state control, forced tool selection, output validation and bounded retries on two open-weight models, and got mixed results across all four model-task cells. Adding structured planning, where the plan is checked against a fixed schema before any tool fires, pushed three of four cells to a Reproducibility Rate and Determinism Index of 1.000 at N=100, with task success at 100%. Token cost fell in every cell. Latency split by model, one faster and one markedly slower, so that trade needs measuring rather than assuming. The retry and validation layers everyone reaches for first are not where determinism comes from. (arXiv)
Each link below shares sources, entities, or timing with this story.
Nearly all cache-compaction research assumes a static context where future queries are known offline, which agents never have. Comparing token eviction against attention matching across proxy-query sources on BrowseComp-Plus and WideSearch, compacting a turn immediately often...
"Memory in the Loop" (arXiv:2607.05690) moves memory read/write inside the agent's per-step loop, viable only with an in-process store answering in ~100µs. The behavioral number is the story: redundant actions were 0.0 of 12 at in-process speed but 7.2 of 12 at a 110ms cloud r...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
This is the paper of the week. arXiv 2607.28871 introduces BSG-VA, which replays every validation command an agent runs across three code states: the original buggy code (B), the candidate patch (S), and the gold developer fix (G). If a test passes in all three states, it neve...
"Towards a Science of AI Agent Reliability" (arXiv 2602.16666) — 12 concrete metrics decomposing reliability along consistency, robustness, predictability, and safety. Key finding: stronger performance on benchmarks does NOT correlate with reliable real-world operation. Intera...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.