Fetching from the wire…
Public story · 2026-08-17 · high
1,221 people tested three coding agents on ICML papers, and human-guided runs beat agents working alone.
Why now: The hackathon ran July 15 to August 2, giving peer review a data point instead of a guess about what agent-assisted checks catch.
Hugging Face ran 2,226 of ICML 2026's accepted papers through agent review from July 15 to August 2. 1,221 people used Claude Code, Codex and Cursor for the check, per Hugging Face. That's more than a third of the conference's 6,352 accepted papers, checking that used to cost a reviewer a weekend per paper.
The hackathon logged 6,816 logbooks and 2,962 cloud jobs. 51% of the examined papers had at least one claim the agents verified. 23% had at least one claim flagged false, with 49 papers falsified outright. Independent teams checking the same paper landed on conflicting verdicts in 242 cases.
Agent-only runs hit limits on scale-dependent behavior, according to the results. Human-guided workflows, where a person steered the agent's checks, were the most reliable setup in the hackathon.
"Checking a paper carefully used to cost a reviewer a weekend; an agent can attempt it in an afternoon," the hackathon's organizers wrote. Agent-only runs still missed cases that human-guided runs caught, which is why the 242 conflicting verdicts matter as much as the 49 outright falsifications. Conferences reading this as a case for agent-only review are reading past the results.
Each link below shares sources, entities, or timing with this story.
QM went up under MIT license. Created July 29. As of the GitHub API check: 8,420 stars, 887 forks. Five days. YC uses it internally across accounting, legal, events, and engineering, including to build QM itself. Every employee and every Slack room gets its own scoped memory,...
The August 14 report covers January through August 2026: model repos grew from 2.43M to 2.96M, datasets from 711K to 1M, and 85.6% of models have under 200 lifetime downloads (Hugging Face). Chinese labs shipped monthly parameter ceilings of 754B to 2.78T against sub-130B for...
Issue 6235 on anthropics/claude-code asks Claude Code to read AGENTS.md, the config file that Codex, Amp, Cursor and most other harnesses already load, rather than only CLAUDE.md. It has been open since August 2025. It has accumulated over 5,200 reactions and 300+ comments, ma...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Three things happened this month that only make sense together. Agent Plugins 1.0 shipped co-signed by six competitors: AWS, Anysphere, Microsoft, OpenAI, Vercel and Google (GitHub Changelog). It makes skills-plus-MCP bundles portable across clients. OpenAI's August 11 Codex c...
Go look at your ~/.claude/CLAUDE.md right now. Mine has internal package names, a build command with a host in it, and notes about which credentials live where. I wrote it assuming exactly one reader. RuntimeWire published traced request captures on August 9 showing Muse Code...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.