Cost vs. capability
The $1.40 model that crushed its $2.82 sibling on physics
Four AI models were asked to build interactive destruction scenes with real physics. Opus 5 passed all three at the lowest cost among top performers, while Fable 5 failed each one at twice the price. A reminder that price tags don't predict physics.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-25 · 3 min read

Atomic.chat, a platform known for practical AI benchmarks, dropped a focused test on Thursday. The assignment: code three HTML files that simulate destruction physics, a tornado lifting objects, a wrecking ball striking a building, and a truck snapping a truss bridge. Simple enough for any model that claims to understand how things move in space.
The four contestants: Opus 5 (Anthropic), Fable 5 (Anthropic), Kimi K3 (Moonshot AI), and GPT 5.6 (OpenAI). Opus 5 used 55.9K tokens and cost $1.40. Fable 5 used 55.1K tokens but cost $2.82, nearly double. Kimi K3 used 35.7K tokens for $0.55. GPT 5.6 used 20.1K tokens for $0.31. Cheaper isn't always better, but Opus 5's performance at that price point is hard to ignore, especially given its known cost advantage across other benchmarks.
What Opus 5 got right
Opus 5 delivered three working scenes. Its tornado had houses spiraling up the funnel and shooting out the top, the kind of behavior a real physics simulation should show. The wrecking ball scene displayed the wall cracking at impact and rubble piling on the ground. The bridge collapsed in stages, dropping the truck into the river below. Atomic.chat gave it a pass on all three. It's the kind of performance you'd expect from a model that, as much smaller models have shown, doesn't need to be the biggest to be the best.
Where Fable 5 stumbled
Fable 5's scenes lacked basic physical coherence. Its tornado had almost nothing on the ground to pick up, which defeats the point. The building collapsed on its own before the wrecking ball reached it. The bridge disintegrated all at once, no progressive failure, no intermediate state. For a model marketed toward engineering and coding, these are embarrassing failures.
GPT 5.6 and Kimi K3 lag behind
GPT 5.6 was the cheapest at 31 cents, but its wrecking ball never reached the building, and its bridge broke in a way nothing breaks in real life. Kimi K3, a newer entrant from China, ended up with similar results. Despite Kimi K3's growing popularity, it could not match Opus 5 on physical reasoning.
Why it matters
The test is limited but telling. Many AI applications in gaming, robotics, and education need models that reason about how objects behave in space. This is a known weak spot for most large language models, which are trained primarily on text. LLMs can describe data but struggle to reason through it, and this benchmark underlines the gap. That Opus 5 handled it well while Fable 5 did not suggests that architectural choices or training data can make a big difference.
Cost also emerges as an unexpected twist. Opus 5 outperformed its more expensive sibling by a wide margin. Developers shopping for a simulation model would do well to test across the family rather than assume pricier is better, a lesson that applies beyond this benchmark, as real-world deployment often reveals.
The test confirms something researchers have noted before: language models excel at syntax but often struggle with the semantics of the physical world. The size of the gap between Opus 5 and Fable 5, both from the same lab, is a reminder that specialization comes with trade-offs. If you're building a robot that needs to understand collisions, you might want to see how the simulation stack is evolving before picking your model.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.