Show-Harness gets frontier VLMs controlling robots zero-shot through discrete semantic action units
arXiv 2609.10522 (2026-09-09) exposes a compact semantic interface of discrete action units that a vision-language model can reason over directly, with embodiment-specific interpreters deterministically grounding each unit into local robot actions so the VLM stays responsible for fine-grained physical decisions. Through the same interface the authors unlock closed-source frontier VLMs for zero-shot robot control and adapt small open-source VLMs for low-cost deployment with only a few GPU-hours of fine-tuning, and their GUMI GUI Manipulation Interface extends the action space to demonstration collection without teleoperation hardware. Agents generalize across tasks, embodiments and environments, outperforming representative agentic and VLA baselines — the argument being that interface design, not model capacity or embodiment-specific pretraining, is the binding constraint.
Source
↳ Follow the thread