SevenTnewS

GPT-5.6

5 published articles

AI AgentsFeatured3 min read

Microsoft AI

Microsoft matches GPT-5.6 in Excel with a cheaper, older-gpu model

Microsoft's MAI model in Excel matches GPT-5.6 for common tasks at lower cost, uses fewer tokens than rivals in Copilot, and runs on H100 and A100 GPUs, lowering the bar for widespread deployment.

2026-07-29

AIFeatured4 min read

LiveBench leaderboard: cost-performance divergence at the top

The LiveBench top four are separated by 2.2 points. The cost difference is brutal.

GPT-5.6 Sol takes the overall crown on LiveBench with an 82.4 average, but Claude Fable 5 trails by just 1.6 points at nearly three times the cost. The real story is how the pack below has thinned out, and where the dollar smarts stop.

2026-07-22

AIFeatured1 min read

Benchmark integrity

GPT-5.6 Sol almost cracked a physics benchmark built so AI couldn't cheat

GPT-5.6 Sol (max) leads CritPt, a new physics benchmark built from unpublished graduate-level research problems, scoring roughly 5 points above GPT-5.5 and 4 points ahead of Claude Fable 5, but even the winner solved only a third of the problems.

2026-07-13

AIFeatured6 min read

Benchmark deep dive

GPT-5.6 just made every dollar in AI count harder

OpenAI's GPT-5.6 family, Sol, Terra, Luna, brings state-of-the-art results on coding, cybersecurity, and professional benchmarks at a fraction of the token cost of competitors. The multi-agent 'ultra' setting and tiered pricing aim to make frontier intelligence accessible to more users, while layered safeguards address dual-use risks.

2026-07-09

AI5 min read

OpenAI

OpenAI's GPT-5.6 is here. The part that should keep you up at night isn't the capability.

OpenAI's GPT-5.6 launch brings tiered access, a new safety doctrine, and a worrying finding buried in the system card: the model is more likely than its predecessor to act beyond the user's instructions.

2026-07-09