Fetching from the wire…
Public story · 2026-07-30 · high
The best of nine top VLMs scored 16.8% on 1,218 indoor tasks, failing not at recognition but at tracking their own position.
Why now: This is new arXiv research, flagged in the July 30 briefing, with no comparable benchmark cited for these tasks.
Meta built a benchmark called HumanCLAW that strips motor control out of VLM navigation tests, and the best of nine models still scored 16.8%, per the new arXiv paper.
The design is the point. A separate, physically simulated controller handles balance, movement, and collision avoidance, translating a model's high-level commands into sub-second chunks of full-body motion. Whatever's left over after that is pure decision-making, tested across 1,218 long-horizon find-navigate-interact episodes in 41 indoor scenes. All nine state-of-the-art vision-language models tested failed to clear even a fifth of them.
The failure mode isn't the obvious one. The paper's diagnosis: recognizing the target object isn't what trips these models up. They lose track of their own body in space, and can't reliably tell whether they've arrived at a target or just collided with something near it.
That reframes where embodied AI's problem actually sits. It isn't perception, and with motor control factored out, it isn't motor skill either. It's self-localization, a model's sense of where it is and what it's touching. Scaling up the VLM doesn't obviously fix that. Watch whether follow-up work bolts on an explicit spatial-state model instead of a bigger one.
Each link below shares sources, entities, or timing with this story.
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Launched August 8 as the new Quality Mode at grok.com/imagine and in the Grok mobile apps, pitching precision editing, crisp text rendering, and improved factuality, with API access promised but not shipped (The Decoder). On the August 7 Arena leaderboards the faster "low" var...
Muse Image, from Alexandr Wang's lab, generates and edits from text, writes and runs code to produce accurate plots and QR codes, and searches the web to ground outputs. It launched free in Meta AI, WhatsApp, and Instagram Stories, and immediately drew pushback for training on...
Reads, edits, and executes actions on local files and applications directly on the user's machine. Positions Meta against OpenClaw and Claude Cowork in local agent runtime. No prior announcement preceded the release. OneNewsPage ---
Designed with Broadcom, built by TSMC, going into manufacturing after a six-week test run found no major issues (Reuters). Iris supplements rather than replaces Meta's Nvidia/AMD GPUs and supports a buildout targeting 7GW by end-2026 and 14GW in 2027 against up to $145B in 202...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.