SevenTnewS

Model race

The LiveBench top four are separated by 2.2 points. The cost difference is brutal.

GPT-5.6 Sol takes the overall crown on LiveBench with an 82.4 average, but Claude Fable 5 trails by just 1.6 points at nearly three times the cost. The real story is how the pack below has thinned out, and where the dollar smarts stop.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-07-22 · 4 min read

The LiveBench top four are separated by 2.2 points. The cost difference is brutal.
Sources : LiveBench

LiveBench dropped its latest update, and the shape of the leaderboard is telling. The top is dominated by two labs, OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5, separated by a rounding error on the composite score but worlds apart on pricing. The broader trend: the field is flattening out.

GPT-5.6 Sol Max Effort posts an 82.4 overall average. Claude Fable 5 Max Effort sits at 80.8. That 1.6-point gap costs nearly 2.7 times more per token if you run Anthropic's top config. GPT-5.6 Terra, the cheaper variant in the same family, lands at 79.8, closer to the top-of-stack Claude than the gap between Sol and Fable. The marginal returns on spending up are tiny, and they compress fast. This pattern echoes the $314 billion assumption that broke, where diminishing returns at scale rewrote the economics of AI inference.

What each model does well

The headline average hides real category-level divergence. Claude Fable 5 dominates writing at 90.7, tied with GPT-5.5 Thinking and Claude 4.8 Opus Thinking, and a full 3 points ahead of GPT-5.6 Sol at 87.7. If your primary load is prose, Claude Fable is the clear pick, and the closest GPT variant (5.5 Thinking) is $0.53 per token versus Fable's $1.57. Anyone running high-volume text generation should factor that into their budget planning, mirroring the cost-efficiency calculus seen in Microsoft's $600 million bet on Kimi K3.

Graphique : Top model composite scores · Category scores: GPT vs Claude · Price-performance curve
The top three models on LiveBench composite score cluster within 2.6 points, per the article. Category-level divergence between the top two models, as reported by LiveBench data in the article. Composite scores across price tiers, showing diminishing returns near the top per the article's data.

Reasoning is a rougher story. GPT-5.6 Sol hits 96.2 on reasoning benchmarks, slightly ahead of Claude Fable's 96.0 and GPT-5.5 Thinking's 95.9. These are all within noise. The real drop-off happens after the top four, Gemini 3.1 Pro Preview scores 91.0 and Claude 4.7 Opus Thinking lands at 92.9. Beyond that, numbers slip into the 80s pretty fast. The compression at the top recalls Grok 4.5's narrow lead on the SWE Marathon, where tiny margins define the frontier.

Math remains a stronghold for the GPT line. GPT-5.6 Sol leads at 91.7, followed by GPT-5.6 Terra at 90.6 and GPT-5.5 Thinking at 89.7. Claude Fable 5 matches the GPT-5.5 score at 89.7, but Claude 4.8 Opus Thinking trails at 89.7 as well, identical on paper. The variation is small enough that a single bad seed could flip the order.

The coding results show more spread. GPT-5.6 Sol scores 83.9 on coding. Claude Fable 5 scores 86.0. Several Claude line variants outperform GPT on coding, Claude Sonnet 5 at 80.7 and Claude 4.6 Opus Thinking at 78.2, but none match the top two. The coding leaderboard is the one category where a model from outside the duopoly, Gemini 3.1 Pro Preview at 76.5, lands within striking distance of the next tier down.

The budget zone: open models and outsiders

Below 75 overall, the field gets crowded with interesting options. DeepSeek V4 Pro, open weight, posts a 71.6 average at $0.05 per token, roughly 30 times cheaper than Claude Fable 5 for about 89% of the composite performance. Kimi K3, also open, scores 78.5 at $0.379, making it the best value in the high-70s band. Qwen 3.7 Max at 73.1 and $0.182 is competitive with GPT-5.4 Nano at 69.6 and $0.091 in the same price neighborhood. The open-weight surge here echoes Kimi K3's benchmark win over Claude and GPT, showing value models can compete in specific tasks.

Grok 4.5 at 76.3 and $0.128 is a newcomer worth watching, it beats several higher-priced Claude and GPT variants on the composite while charging a fraction of the cost. Its weakness is writing (68.6) and coding (70.8), but reasoning at 90.8 is genuinely top-tier. If xAI iterates on the weaknesses, this line could climb the leaderboard fast.

The contraction at the top

The most striking feature of this update is how the top four models cluster within 2.2 points of each other while everything else drops off smoothly. This is not the exponential progress curve AI labs like to project. It is a long tail of incremental gains at exponential cost. The difference between GPT-5.6 Sol and GPT-5.4 Nano is 12.8 points on the composite, but the price difference is about 6.5x. Paying more gets you more, but the curve is flattening.

For anyone deploying at scale, the sweet spot is probably the $0.10, 0.20 range where models like Grok 4.5, DeepSeek V4 Pro, and Qwen 3.7 Max live. They are good enough for most production loads, and the cost delta vs the top tier funds a lot of evaluation runs before you cross the break-even point on accuracy gains. The pattern of cost-driven re-evaluation is consistent with the 95% gap benchmark, where throwing more compute at a problem doesn't solve structural limitations.

The LiveBench data does not say which model is best, that depends on your task and your budget. It does say that the gap between the best and the very good is shrinking, and that the best is getting expensive faster than it is getting better.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.