Fetching from the wire…
Public story · 2026-08-07 · high
ByteDance Seed's GST-Bench, from 6,790 minutes of video, shows models track local space fine but lose the global picture.
Why now: The GST-Bench paper posted to arXiv and made the August 7, 2026 briefing.
GST-Bench catches vision-language models losing track of space over long video, per ByteDance Seed's paper on arXiv.
The gap is stark: top models average 42.68 on the benchmark's spatial consistency tests, against 79.08 for humans, a roughly 36-point spread. That matters most for agents that operate over long video, since a wrong mental map compounds with every step that follows it.
The benchmark draws on 6,790 minutes of synthetic video with human-verified questions, built to isolate one specific failure. Models handle local spatial relations, judging what sits near what in a single view, without much trouble.
The failure shows up when models have to stitch separate observations into one consistent scene as an episode stretches on.
ByteDance Seed also released GST-Train, a training set built to target that consolidation problem directly.
Whether GST-Train closes the 36-point gap is the thing to watch next. If it doesn't, the failure sits deeper than a data problem, in how these models represent space at all.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
arXiv 2607.26998 flips the pentest agent's observation-action loop against it, replacing static honeytokens with a trajectory-adaptive policy that constructs new decoy artifacts conditioned on the agent's interaction history, folding validated ones into a factually consistent...
arXiv 2607.05458 formalizes agent execution as a finite-horizon MDP where a lightweight controller picks structural actions (when to verify, retry, branch) while the LLM executor stays frozen, trained from offline rollouts via advantage-weighted regression using only terminal...
ADeptS-Bench tests seven models on paired benign and malicious GUI tasks and finds none clearing 80% task success while staying under 30% attack success. The ablation is the usable finding: removing the refusal tool raises attack success 21-23pp for tool-dependent models and 1...
It treats harness improvement as offline learning: diagnose failure traces, generate structured patches that edit the harness as source code, then select updates by validation over mini-batches of failures (arXiv 2608.23041). Gains: 9.0 on GAIA2, 9.6 on SWE-Bench Pro, 10.0 on...
This one falsifies an assumption a lot of this year's agent tooling is built on, mine included. The paper is WER (Write, Execute, Refine), arXiv 2608.17587, published 2026-08-18. It opens with a measurement rather than a method: skills that an agent authors for itself perform...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.