VLM
2 published articles
LLMs & Models3 min read
3D Reasoning
SceneActBench: Why even the best VLMs fail at 3D action
SceneActBench tests eleven VLM configurations on five 3D tasks in a unified agent loop. Overall scores range from 38.6 to 50.2, with no model performing consistently. The benchmark exposes a blind spot in vision-language agents: acting on full scenes, not just describing them.
2026-07-31
Labs & ResearchFeatured3 min read
Open-endedness & VLMs
Sakana AI's Picbreeder reboot shows what creativity metrics miss
A Sakana AI-led team replaced human Picbreeder users with VLMs and found the synthetic archives lack the boldness and diversity of the human originals. Their experiments with noise, memory, and thousands of prompted personalities reveal both the promise and the limits of using large models for open-ended discovery.
2026-07-20