Fetching from the wire…
Public story · 2026-08-03 · high
Coding agents keep accuracy on the same documents, but at substantially higher cost, per the ExtractBench paper.
Why now: Coverage of the paper as of Aug. 3 already includes the dataset and eval code, giving builders a way to test their own extraction setup before shipping.
Commercial VLMs silently truncate record lists on long documents, per a new benchmark called ExtractBench.
That's the exact failure mode for anyone piloting AI extraction on long, multi-page filings. The model can look accurate while quietly dropping records, and a normal accuracy score won't catch it.
The paper, from Adrian Lyjak, Simon Suo and co-authors, tests schema-guided extraction on 370 enterprise documents across 4,869 pages, 8 domains and 67 document types. It scores value accuracy, word-level grounding and page-level grounding separately, so a model can't fake traceability by getting the right answer without showing its work.
On short documents, the VLMs it tested score well. Push the record count up and they start truncating without flagging that anything was cut.
Coding agents hold accuracy on the same longer documents, but run at substantially higher cost. The paper doesn't say by how much, just that the gap is substantial.
Dataset and eval code are public on HuggingFace and GitHub.
Each link below shares sources, entities, or timing with this story.
5.8k stars, Rust, macOS, ~40 MB versus ~67 MB upstream, keeping full Lua customization. Command-failure recovery with suggested fixes applied via Cmd + Shift + E, natural-language-to-command via # <description>, preconfigured integration for Claude Code, Codex, Gemini CLI, and...
Give it an editable agent and a benchmark, and it walks a coding agent through systematic changes to prompts, tools, workflows, models and reasoning settings, trading off accuracy, cost, latency and reliability (GitHub). Install is gh skill install microsoft/agent-lightning ag...
119 repository-level tasks from 98 GitHub repos across 20 scientific domains, split into issue-driven, expert-exploratory, and engineering-integration paradigms. Claude Code with Opus-5 (max) lands below 50%. arXiv The ablation is the better finding: stripping explicit scienti...
Three independent signals: Claude Code 2.1.225 naming gateway caps and reset times, codeburn shipping local cost attribution across 36 coding tools with a yield command correlating spend to shipped code, and GitHub's usage metrics API reporting agent activity. Per-agent cost a...
The repo appeared on trending with +135 stars and a repositioned pitch, pivoting from the general local-code-execution tool it launched as in 2023. It's now aimed directly at Claude Code and Codex but on the open-weight side. Single-source on the repositioning, so check the re...
A GitHub repo cataloging Claude Code tips doesn't normally warrant a top story. But shanraisshan/claude-code-best-practice at 53.4K stars isn't a tips list anymore. It's the de facto reference for how an entire generation of developers is learning to work with AI coding agents...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.