LLMs & Models3 min read
3D Reasoning
SceneActBench: Why even the best VLMs fail at 3D action
SceneActBench tests eleven VLM configurations on five 3D tasks in a unified agent loop. Overall scores range from 38.6 to 50.2, with no model performing consistently. The benchmark exposes a blind spot in vision-language agents: acting on full scenes, not just describing them.
2026-07-31