Benchmarks
20 published articles
local ai
The uncensored model paradox: one knob removes both the annoying refusals and the safety guardrail
Benchmarking five locally run uncensored LLMs shows abliteration cuts over-refusal from 44% to near zero with no hit to reasoning, but the same edit collapses safety refusals from 41.5% to 9.5%, because both ride on the same internal direction. The real reason to run uncensored may not be what you think.
2026-07-19
benchmark breakdown
The 2.8 trillion parameter model that beats the frontier on the benchmarks that matter
Kimi K3, the 2.8T-parameter open model from Moonshot AI, trails frontier proprietary models on most broad benchmarks, but leads on SWE Marathon, Terminal-Bench 2.1, BrowseComp, and others. The detailed table reveals where its architectural bets on KDA and Stable LatentMoE pay off.
2026-07-17
Artificial Intelligence
Inkling is open-source AI's $53 billion reality check
Inkling is the first openly available near-1T parameter model with native audio, image, and text input alongside a 1M-token context window. The raw benchmark scores are strong. The real story is how the open-source ecosystem has moved from playing catch-up to competing at the frontier, and where Inkling fits on that new map.
2026-07-15
Artificial intelligence
Sonnet 4.6 just made Opus look expensive. That changes everything.
Anthropic's Claude Sonnet 4.6 matches or beats Opus-class performance on key benchmarks while keeping the same pricing as Sonnet 4.5. Early developer preference data and customer testimonials suggest the model is already displacing costlier alternatives for real-world agentic and coding tasks.
2026-07-14
Generative AI
An AI that can see its own mistakes and undo them mid-generation
Google introduces CO2Jump, a training-free sampler for joint text and image generation. It uses a self-correcting Markov jump process where each modality's confidence scores guide the other's updates in real time, catching cross-modal errors mid-generation.
2026-07-14
Model strategy shift
Claude Sonnet 5 just made the Opus price gap harder to justify
Anthropic releases Claude Sonnet 5, a model that closes the gap with Opus-class models on agentic tasks like coding, tool use, and reasoning. Priced at $2 per million input tokens through August 2026, it launches today across plans and the Claude API.
2026-07-10
Artificial Intelligence
Kimi K2.7 Code is faster and cheaper. But open-source coding just hit a wall called GPT-5.5.
Moonshot AI's Kimi K2.7 Code makes big gains on long-horizon coding tasks with 30% less token waste. Yet GPT-5.5 and Claude Opus 4.8 still lead on key benchmarks, highlighting the real-world trade-offs of open-source decisions.
2026-07-09
Paper 2606.23050 decoded
A 33-page preprint just landed. Here's what it reveals about where AI is heading
A 33-page preprint, paper 2606.23050, released only five days ago, marks a serious contribution to AI research. This article unpacks its methodology, key findings, and what it signals for the trajectory of machine learning and natural language processing.
2026-07-07