Fetching from the wire…
Research2026-06-29 · source-backed
Niclas Lietzow, Danielle Bitterman, and Carsten Eickhoff probe what happens when a vision-language model's eyes disagree with its memorized world knowledge, identifying a "vision-default, prior-override" causal mechanism. This is directly useful for debugging the maddening class of VLM hallucination where the model describes what it expects instead of what's actually in the image. It gives you a mechanistic handle on when to trust perception versus baked-in memory.
Each link below shares sources, entities, or timing with this story.
The system improves frozen VLM agents with zero parameter updates. It collects experiences in verifiable environments, distills lessons through verifier-guided reflection, attaches a Transfer Reliability Score to each, and retrieves only relevant and reliable lessons at infere...
arXiv 2607.29677, from a team including Adrian Lyjak and Simon Suo, evaluates schema-guided extraction across 370 enterprise documents, 4,869 pages, 8 domains and 67 document types, scoring order-insensitive value F1, word-level grounding F1 and page-level grounding F1 separat...
arXiv 2607.27180 decouples decision-making from execution: an off-the-shelf VLM issues atomic skill commands, a controller translates them into sub-second chunks of physically simulated full-body motion, so balance and motor failures are factored out. On 1,218 long-horizon ego...
OSReward builds human-verified ground truth for computer-use trajectory judgments and finds even state-of-the-art models fall short with a consistent bias toward misclassifying failures as successes. The authors release OS-Shepherd at 9B and 35B, trained on a 100K corpus, clai...
Testing five VLMs across two benchmarks and five visual-token budgets, native-resolution table images match text on accuracy and efficiency, but downscaling makes models compensate for lost readability with longer, weaker reasoning traces that cancel the token savings. The exp...
OpenMOSS (Xipeng Qiu's group, 32 authors) released MOSS-VL on Aug 15, built on gated cross-attention so it can ingest incoming video frames during generation, with visual tokens kept outside the decoded sequence. 66.0 on OmniMMI Proactive Alerting against a 37.5 baseline, time...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.