Fetching from the wire…
Research2026-07-11 · source-backed
Tsinghua's CompactionRL folds summarization into RL rollout collection so the agent learns what to keep when it compresses, optimizing summary and task under one reward. Under fixed 64k–80k windows it adds +5.5 to +7.0 on SWE-bench Verified and +3 to +6.8 on Terminal-Bench versus heuristic compaction. If you run long-horizon coding agents, stop treating "summarize the last N turns" as a fixed prompt. The compression step is a tunable component, and the gap between trained and heuristic is bigger than most model upgrades.
Each link below shares sources, entities, or timing with this story.
Same weights. Same tasks. Same time budget. One change to the harness, and fail-to-pass on SWE-bench Verified went from 28% to 49%. The setup in arXiv 2608.26218 is clean enough to be worth describing precisely. 169 SWE-bench Verified tasks, a 20,480-token context window, a fi...
June leaderboards increasingly score a weighted blend of Terminal-Bench 2.0, BrowseComp, and OSWorld-Verified instead of SWE-bench alone, with BenchLM weighting "agentic" browse-and-do workflows at 22%, its highest single category. When you evaluate a tool, the headline chat o...
Poolside AI released two models that change the math on local coding agents. Laguna M.1 is a 225B total / 23B active MoE model scoring 72.5% on SWE-bench Verified. Laguna XS.2 is a 33B total / 3B active model scoring 68.2% on the same benchmark, 44.5% on SWE-bench Pro, and 30....
A team at UC Berkeley RDI built an automated scanning agent that achieved near-perfect scores on eight major AI agent benchmarks. SWE-bench Verified: 100%. Terminal-Bench: 100%. WebArena: approximately 100%. FieldWorkArena: 100%. GAIA: roughly 98%. OSWorld: 73%. The agent didn...
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage ri...
Kwaipilot's release is 35B total / 3B activated, built on Qwen3.6-35B-A3B, Apache 2.0, 262,144-token context, posting 69.40% on SWE-bench Verified, 63.00% multilingual, 45.96% SWE-bench Pro, 41.02% Terminal-Bench 2.1. An r/LocalLLaMA user reports it one-shotting a five-level T...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.