EG
IMBench: New Benchmark Reveals Systematic Gap in Robots’ Intuitive Physical Reasoning—VLMs Reason but Can’t Act, VLAs Act but Can’t Reason
ResearchJuly 20, 2026Stax

IMBench: New Benchmark Reveals Systematic Gap in Robots’ Intuitive Physical Reasoning—VLMs Reason but Can’t Act, VLAs Act but Can’t Reason

IMBench, a new benchmark for intuitive robot manipulation, reveals a systematic gap: VLMs can reason about physics but can’t act, while VLAs can act but fail to respect task constraints. The 35-task benchmark challenges the assumption that current VLA performance on controlled demos transfers to novel scenes.

#IMBench#benchmark#robot manipulation#intuitive physics#VLA#VLM#embodied AI#arXiv
Reading in English

A new research benchmark called IMBench (Intuitive Manipulation Benchmark) has been published on arXiv, designed to evaluate what researchers call “intuitive manipulation”—the ability of robots to combine physical reasoning with motor execution in an integrated way.

Published on July 15, 2026, IMBench proposes a suite of 35 tasks covering contact-rich manipulation, tool use, and multi-step dependencies, accompanied by 14,000 filtered trajectories and tools for generating new scenarios at scale. Unlike existing benchmarks, IMBench requires models to first identify the relevant physical structure of a scene—weight, friction, geometric constraints—before producing an executable action sequence under explicit constraints.

The researchers tested both state-of-the-art Vision-Language Models (VLMs) and Vision-Language-Action (VLA) models on the task set. The results reveal a systematic gap in current systems: VLMs show partial physical reasoning capability but fail to translate that reasoning into executable action plans, while state-of-the-art VLAs fail to respect task constraints and generalize poorly across scenarios.

For the robotics industry, these findings confirm a concern already widespread among integrators: the performance displayed by generative policies on controlled demonstrations does not guarantee transferable physical understanding to novel scenes. IMBench thus provides a more rigorous measure to distinguish systems that truly “understand” a scene from those that reproduce training patterns.

This work joins a recent wave of benchmarks aiming to bridge the gap between simulation evaluation and real-world deployment, as models like GR00T N2, Pi-0, and Helix pursue generalist manipulation across varied environments. By explicitly isolating “intuitive reasoning” as a missing axis in current robotic policies, the authors position IMBench as a diagnostic tool to guide the next generation of robot foundation models, rather than a simple raw performance ranking.

Language: English- Showing content in English