LiveBench leaderboard: cost-performance divergence at the top
The LiveBench top four are separated by 2.2 points. The cost difference is brutal.
GPT-5.6 Sol takes the overall crown on LiveBench with an 82.4 average, but Claude Fable 5 trails by just 1.6 points at nearly three times the cost. The real story is how the pack below has thinned out, and where the dollar smarts stop.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-22 · Last updated: 2026-07-30 · 4 min read

LiveBench dropped its latest update, and what it shows is increasingly clear. The top is dominated by two labs, OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5, separated by a rounding error on the composite score but worlds apart on pricing. The broader trend is that the field is flattening out.
GPT-5.6 Sol Max Effort posts an 82.4 overall average. Claude Fable 5 Max Effort sits at 80.8. That 1.6-point gap costs nearly 2.7 times more per token if you run Anthropic's top config. GPT-5.6 Terra, the cheaper variant in the same family, lands at 79.8, closer to the top-of-stack Claude than the gap between Sol and Fable. The marginal returns on spending up are tiny and they compress fast. This pattern mirrors Microsoft's 89% cost reduction with in-house models, where the same economics of diminishing returns reshaped inference decisions.
What each model does well
The headline average hides real category-level divergence. Claude Fable 5 dominates writing at 90.7, tied with GPT-5.5 Thinking and Claude 4.8 Opus Thinking, and a full 3 points ahead of GPT-5.6 Sol at 87.7. If your primary load is prose, Claude Fable is the clear pick, and the closest GPT variant (5.5 Thinking) is $0.53 per token versus Fable's $1.57. Anyone running high-volume text generation should factor that budget difference into their choice. The landscape of model alternatives is vast, as seen in Macaron Venti's four-specialist design that beats GPT-5.5 and Opus 4.8.

Reasoning is a rougher story. GPT-5.6 Sol hits 96.2 on reasoning benchmarks, slightly ahead of Claude Fable's 96.0 and GPT-5.5 Thinking's 95.9. These are all within noise. The real drop-off happens after the top four; Gemini 3.1 Pro Preview scores 91.0 and Claude 4.7 Opus Thinking lands at 92.9. Beyond that, numbers slip into the 80s fast. This compression at the top happened just as Chinese labs reached parity on benchmarks, as four Chinese AI labs rewrote the competitive landscape shows.
Math remains a stronghold for the GPT line. GPT-5.6 Sol leads at 91.7, followed by GPT-5.6 Terra at 90.6 and GPT-5.5 Thinking at 89.7. Claude Fable 5 matches the GPT-5.5 score at 89.7, but Claude 4.8 Opus Thinking also trails at 89.7, identical on paper. The variation is small enough that a single bad seed could flip the order.
Coding benchmarks spread further. GPT-5.6 Sol scores 83.9 on coding. Claude Fable 5 scores 86.0. Several Claude line variants outperform GPT on coding, Claude Sonnet 5 at 80.7 and Claude 4.6 Opus Thinking at 78.2, but none match the top two. The coding leaderboard is the one category where a model from outside the duopoly, Gemini 3.1 Pro Preview at 76.5, lands within striking distance of the next tier down.
The budget zone: open models and outsiders
Below 75 overall, the field gets crowded with interesting options. DeepSeek V4 Pro, open weight, posts a 71.6 average at $0.05 per token, roughly 30 times cheaper than Claude Fable 5 for about 89% of the composite performance. Kimi K3, also open, scores 78.5 at $0.379, making it the best value in the high-70s band. Qwen 3.7 Max at 73.1 and $0.182 is competitive with GPT-5.4 Nano at 69.6 and $0.091 in the same price neighborhood. The open-weight surge shows value models can compete in specific tasks, as seen in Kimi K3's position as the largest open model yet still not the best.
Grok 4.5 at 76.3 and $0.128 is a newcomer worth watching. It beats several higher-priced Claude and GPT variants on the composite while charging a fraction of the cost. Its weakness is writing (68.6) and coding (70.8), but reasoning at 90.8 is genuinely top-tier. If xAI iterates on those weaknesses, this line could climb the leaderboard fast.
The contraction at the top
The most striking feature of this update is how the top four models cluster within 2.2 points of each other while everything else drops off smoothly. This is not the exponential progress curve AI labs like to project. It is a long tail of incremental gains at exponential cost. The difference between GPT-5.6 Sol and GPT-5.4 Nano is 12.8 points on the composite, but the price difference is about 6.5x. Pay
- Source : LiveBench
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.