Fetching from the wire…
Research2026-08-11 · source-backed
The Replay Gap (arXiv 2608.08239) forked live SWE-bench trajectories at controlled points, rebuilt the environment, and continued each fork with a different model across ~900 rollouts. Model swaps rewrote 61–94% of post-fork actions and diverged at the very first post-fork action 74–77% of the time, versus 6–35% for same-model controls. All five outcome flips occurred in swap arms, zero across 359 control forks. A log-stitching replay evaluator mispredicted every success-relevant outcome and produced patches with 0.00–0.11 similarity to reality. If you route per-step between Haiku and Opus, branch live or don't trust the number.
Each link below shares sources, entities, or timing with this story.
Tessl ran 880 evaluations across 9 models with and without agent skills. The result inverts what most teams assume about AI costs. Haiku 4.5, Anthropic's cheapest model at roughly $0.25 per million tokens, scored 84.3% when given a well-crafted agent skill. Opus 4.7, the most...
For two years the technique was accumulation. Longer system prompts, longer CLAUDE.md, more numbered do/don't lists, more "always verify your work" imperatives. Anthropic's context-engineering guidance for Claude 5 models inverts it, with an 80% deletion figure attached. The s...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
July 9, across VS Code, Visual Studio, Copilot CLI, the cloud agent, github.com, GitHub Mobile, JetBrains, Xcode, and Eclipse. Sol is the high-reasoning tier at $5/1M in, $30/1M out, gated to Pro+/Max/Business/Enterprise. Terra is the balanced default at $2.50/$15. Luna is fas...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
JetBrains survey confirms 93% of developers regularly use AI tools (up from 85% in 2025), 51% daily. Key finding: no single model excels at everything. Recommended allocation: Opus for reasoning/architecture, Gemini for visuals/frontend, DeepSeek for cost-sensitive tasks, Haik...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.