Fetching from the wire…
Research2026-08-18 · source-backed
The authors formalize Dense Same-Class Attribute Misbinding and built InstaBind-Lite to measure it: 524 images, 529 groups of 3-6 same-class entities, 9,580 deterministically evaluated questions with source-instance annotations that separate copying from hallucination (arXiv 2608.16805). Five open-source models average 19.84% misbinding versus 7.55% for two commercial API systems, and ~81% of identifiable transfers come from adjacent instances. If you're running extraction over dense scenes, dashboards, or tables of similar items, standard accuracy metrics hide this entirely.
Each link below shares sources, entities, or timing with this story.
The study extracted 130 clean atomic state transitions from 707 real issues in SWE-bench Lite and Verified. Plain RAG scored 0.57-0.59 answer accuracy; an LLM reranker didn't help and added latency, about 18 seconds against 2.1. A (subject, relation, object) supersession memor...
arXiv 2607.29677, from a team including Adrian Lyjak and Simon Suo, evaluates schema-guided extraction across 370 enterprise documents, 4,869 pages, 8 domains and 67 document types, scoring order-insensitive value F1, word-level grounding F1 and page-level grounding F1 separat...
Testing five VLMs across two benchmarks and five visual-token budgets, native-resolution table images match text on accuracy and efficiency, but downscaling makes models compensate for lost readability with longer, weaker reasoning traces that cancel the token savings. The exp...
Ockhamareto (arXiv 2608.24473) reinforces a unit-test rollout only when it's non-dominated on both mutation-killing and test count, then ties each test's killing power back to specific source tokens. Against MIST-RL that's a 3.4x better per-test trade-off, plus 30 to 35 percen...
A July 21 paper pairs two near-identical agents: an Explore Agent that inspects untrusted input but holds no tools, and a Safe Agent that takes privileged actions using its own context plus length-constrained hints from the explorer (arXiv 2607.19595). Borrowing from residual...
Niclas Lietzow, Danielle Bitterman, and Carsten Eickhoff probe what happens when a vision-language model's eyes disagree with its memorized world knowledge, identifying a "vision-default, prior-override" causal mechanism. This is directly useful for debugging the maddening cla...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.