Fetching from the wire…
Public story · 2026-02-27 · source-backed
Each link below shares sources, entities, or timing with this story.
Every number you use to pick a harness comes from public repositories the models may have trained on. Specific Labs built the version that doesn't: tasks drawn from licensed private company codebases, including a 200K-user event app and a fintech processing over 100K bank stat...
Most coding agents run on JavaScript runtimes and eat 300MB of RAM just sitting idle. Zerostack is built in pure Rust, uses ~8MB on idle, ~12MB while working, and posts 0.0% idle CPU. It hit 1.0 on crates.io this week with 474 points and 252 comments on Hacker News. The archit...
Tackles cascading errors where one agent's bad output poisons downstream agents. A "rectify-or-reject" pruning framework acts as an active firewall between handoffs without retraining. Practical pattern: add quality gates between agent handoffs. arXiv 2602.23258
On the May 15 – July 1 window (111 problems from 65 repositories, all post-dating training cutoffs), per the announcement thread and the leaderboard itself: Fable 5 at 64.5%, Grok 4.5 at 63.8%, Opus 5 at 63.4%, GLM-5.2 at 62.9%, GPT-5.6 Sol at 62.3%. Compare that to the 95.0 F...
26. arXiv 2601.10338 — Agent Skills in the Wild 27. arXiv 2602.19555 — Agentic AI Attack Surface 28. arXiv 2602.23163 — Steganographic LLM Monitoring 29. arXiv 2602.16666 — Agent Reliability 30. arXiv 2602.23047 — CL4SE Context Learning 31. arXiv 2602.22675 — Search More Think...
Microsoft researchers curated 101 tasks from two production repositories for AL, the DSL behind Dynamics 365 Business Central, adapting SWE-bench's method to an ecosystem with scarce public training data. In bug-fixing, differences between frontier models exceeded differences...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.