Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.04804 sends a 7B searcher into the repo first, sandbox-verifies its reproduction claims and strips false ones, then routes to one of four frontier fixers. On the full 266-task Python slice under the official capped budget it solves 159 vs 158 for the best single model. The honest ablation is the finding: always using the cheapest fixer with the handoff ties the routed system. The verified context handoff carries the result, not the routing. Calibration suggests it redistributes solving ability upward for cheap models while slightly hurting the strongest.
Each link below shares sources, entities, or timing with this story.
119 repository-level tasks from 98 GitHub repos across 20 scientific domains, split into issue-driven, expert-exploratory, and engineering-integration paradigms. Claude Code with Opus-5 (max) lands below 50%. arXiv The ablation is the better finding: stripping explicit scienti...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
I've spent real hours tuning the CLAUDE.md in my own repos. Rewriting architecture notes. Adding conventions. Trimming when it got long. So this one stung. arXiv 2607.27250 ran a two-agent ablation across Claude Code and Codex: 17 real tasks from 3 repositories, 288 gold-test-...
arXiv 2608.09802, accepted at COLM 2026, audits the benchmark everyone quotes in funding decks and finds a large chunk of the reported headroom is measurement error, not model failure. Their replacement is 170 expert-curated multilingual refactoring instances across Python, Ja...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
Paritok-4B (arXiv 2608.24188) is a LoRA on Qwen3-4B distilled from a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories. It's extractive rather than paraphrasing, with 96.0% of emitted identifiers, paths and numbers already present in its input, and intent-conditione...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.