coding AI
6 published articles
AI Models
The 33B model that just beat 137B models on coding benchmarks without changing hardware
Poolside's Laguna XS 2.1 improves SWE-bench Multilingual by 5.4 points to 63.1% while keeping the same 33B-total-3B-activated MoE architecture. The release includes quantized checkpoints, speculative decoding draft models, and an OpenMDW-1.1 license, making local AI coding more practical.
2026-07-21
benchmark breakdown
The 2.8 trillion parameter model that beats the frontier on the benchmarks that matter
Kimi K3, the 2.8T-parameter open model from Moonshot AI, trails frontier proprietary models on most broad benchmarks, but leads on SWE Marathon, Terminal-Bench 2.1, BrowseComp, and others. The detailed table reveals where its architectural bets on KDA and Stable LatentMoE pay off.
2026-07-17
Benchmark deep dive
GPT-5.6 just made every dollar in AI count harder
OpenAI's GPT-5.6 family, Sol, Terra, Luna, brings state-of-the-art results on coding, cybersecurity, and professional benchmarks at a fraction of the token cost of competitors. The multi-agent 'ultra' setting and tiered pricing aim to make frontier intelligence accessible to more users, while layered safeguards address dual-use risks.
2026-07-09
Artificial Intelligence
MiniMax's M3 just beat Opus 4.7 at browsing, trained itself, and never asked for help
MiniMax M3 delivers a 9.4x CUDA kernel speedup, beats Opus 4.7 on BrowseComp, and autonomously replicated an ICLR paper. All in an open-weight package, and it never asked for help.
2026-07-06
Artificial Intelligence
MiniMax's new M2.5 coding model tops the benchmark at 5% of the price
MiniMax's M2.5 model tops the Multi-SWE-Bench coding benchmark, beats mainstream models on workspace tasks, and costs a tenth to a twentieth of competitors. Open-source weights are on HuggingFace.
2026-07-04
AI Models
Anthropic Launches Claude Fable 5: A Fifth-Generation Model for Extended Autonomous Work
Anthropic introduces Claude Fable 5, a fifth-generation model that can run agents for days, tackle ambitious coding projects, and handle complex enterprise workflows with minimal oversight. Pricing starts at $10 per million input tokens.
2026-07-01