SevenTnewS

Benchmarks

20 published articles

AIFeatured5 min read

local ai

The uncensored model paradox: one knob removes both the annoying refusals and the safety guardrail

Benchmarking five locally run uncensored LLMs shows abliteration cuts over-refusal from 44% to near zero with no hit to reasoning, but the same edit collapses safety refusals from 41.5% to 9.5%, because both ride on the same internal direction. The real reason to run uncensored may not be what you think.

2026-07-19

LLMs & ModelsFeatured6 min read

benchmark breakdown

The 2.8 trillion parameter model that beats the frontier on the benchmarks that matter

Kimi K3, the 2.8T-parameter open model from Moonshot AI, trails frontier proprietary models on most broad benchmarks, but leads on SWE Marathon, Terminal-Bench 2.1, BrowseComp, and others. The detailed table reveals where its architectural bets on KDA and Stable LatentMoE pay off.

2026-07-17

LLMs & Models7 min read

Artificial Intelligence

Inkling is open-source AI's $53 billion reality check

Inkling is the first openly available near-1T parameter model with native audio, image, and text input alongside a 1M-token context window. The raw benchmark scores are strong. The real story is how the open-source ecosystem has moved from playing catch-up to competing at the frontier, and where Inkling fits on that new map.

2026-07-15

AI5 min read

Artificial intelligence

Sonnet 4.6 just made Opus look expensive. That changes everything.

Anthropic's Claude Sonnet 4.6 matches or beats Opus-class performance on key benchmarks while keeping the same pricing as Sonnet 4.5. Early developer preference data and customer testimonials suggest the model is already displacing costlier alternatives for real-world agentic and coding tasks.

2026-07-14

AIFeatured4 min read

Generative AI

An AI that can see its own mistakes and undo them mid-generation

Google introduces CO2Jump, a training-free sampler for joint text and image generation. It uses a self-correcting Markov jump process where each modality's confidence scores guide the other's updates in real time, catching cross-modal errors mid-generation.

2026-07-14

Anthropic / ClaudeFeatured4 min read

Model strategy shift

Claude Sonnet 5 just made the Opus price gap harder to justify

Anthropic releases Claude Sonnet 5, a model that closes the gap with Opus-class models on agentic tasks like coding, tool use, and reasoning. Priced at $2 per million input tokens through August 2026, it launches today across plans and the Claude API.

2026-07-10

AI3 min read

Artificial Intelligence

Kimi K2.7 Code is faster and cheaper. But open-source coding just hit a wall called GPT-5.5.

Moonshot AI's Kimi K2.7 Code makes big gains on long-horizon coding tasks with 30% less token waste. Yet GPT-5.5 and Claude Opus 4.8 still lead on key benchmarks, highlighting the real-world trade-offs of open-source decisions.

2026-07-09

Labs & Research4 min read

Paper 2606.23050 decoded

A 33-page preprint just landed. Here's what it reveals about where AI is heading

A 33-page preprint, paper 2606.23050, released only five days ago, marks a serious contribution to AI research. This article unpacks its methodology, key findings, and what it signals for the trajectory of machine learning and natural language processing.

2026-07-07

← PreviousPage 9 / 2 · 20 articlesNext →