Fetching from the wire…
Research2026-08-28 · source-backed
SWE-Prime's premise is that a successful trajectory still contains ineffective, redundant and risky steps, so SFT on all resolved runs teaches bad habits (arXiv 2608.27449). It filters at trajectory level on process quality, result quality and representativeness, then at segment level by grouping consecutive steps and scoring each on contribution, learnability and risk. All segments stay in the sequence for context, but only selected ones contribute to the loss. The 10% subset beats the full resolved dataset by up to 12.2% on SWE-Bench Pro and 24.2% on Verified.
Each link below shares sources, entities, or timing with this story.
SWE-Touch mines task-critical regions from repair trajectories, builds plausible Counter-Edits that conflict with task completion, and injects them with contextual user messages when the agent reaches that code. Across nine models on SWE-bench Verified, average resolve rate dr...
arXiv 2607.29658 attacks the fact that repair agents treat every issue independently and throw away procedural knowledge. STAIR converts historical trajectories into multi-level trees spanning fine-grained diagnostic actions up through high-level strategies, then tailors plan...
arXiv 2608.05144 runs Manager, Planner, Engineer and Reviewer roles over persistent project state with *fixed* model weights, self-evolving through runtime state and control policies rather than training. 76.8% on AARRI-Bench, and mature waves use 21% fewer solve-input tokens...
The Agent Lightning result (arXiv 2608.17528) is a 9B model gaining 14.6 points from 3,500 lines of training code and modest compute. The point isn't that a 9B beats anything, it's the cost curve: 6K examples is a dataset a small team can build. Task-specific agentic RL on an...
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the buil...
This is the most directly usable paper of the day and it does something rare: it bolts onto an existing agent without touching it. Ledger is a deterministic runtime wrapper that distills an agent's completed interactions into explicit state. What has been observed, what has be...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.