3D reasoning
2 published articles
LLMs & Models3 min read
3D Reasoning
SceneActBench: Why even the best VLMs fail at 3D action
SceneActBench tests eleven VLM configurations on five 3D tasks in a unified agent loop. Overall scores range from 38.6 to 50.2, with no model performing consistently. The benchmark exposes a blind spot in vision-language agents: acting on full scenes, not just describing them.
2026-07-31
AI4 min read
Benchmark
Oxford's GauntletBench puts AI agents through 100 real tasks. They failed 81% of them.
Oxford's GauntletBench puts AI agents through 100 challenging real-world tasks. Frontier systems cap out at 19.1% success, far below human performance. The benchmark targets overlooked capabilities in temporal perception, graphical understanding, and 3D reasoning across five professional applications.
2026-07-01