SevenTnewS

AI benchmarks

19 published articles

Labs & Research1 min read

Analysis

Benchmarks, not bargains: China's AI labs reach parity

Chinese AI labs Qwen, DeepSeek, MiniMax, and Kimi have matched or beaten US frontier models in reasoning, coding, and document parsing benchmarks. With aggressive pricing and open-weight strategies, they are no longer just cheap alternatives, they are real competitors.

2026-07-27

LLMs & ModelsFeatured3 min read

Open Source AI

Kimi K3 is the biggest open model ever. It's still not the best.

Kimi K3 is the largest open model at 2.8T parameters, but it fails to beat the best proprietary models on overall benchmarks. While it excels on specific coding and agentic tasks, the gap to frontier leaders like Claude Fable 5 and GPT 5.6 Sol exposes the limits of scaling without architectural and data breakthroughs.

2026-07-27

AI4 min read

Artificial Intelligence

The AI that taught itself to revise: four agents, one puzzle, no consensus

Researchers propose ARCANA, a multi-agent system that cracks ARC-AGI-2 puzzles by splitting reasoning into four steps and closing the loop with reflective feedback. Each failure teaches the next attempt. The architecture improves reasoning efficiency, but questions about generalization and compute cost remain.

2026-07-19

AI1 min read

Orchestrator swap

Perplexity swapped its orchestrator model and the cost-performance chart is brutal

Perplexity swapped Grok 4.5 into Computer as its orchestrator model, claiming top accuracy on its internal benchmark at roughly half the cost of Claude Opus 4.8. The move signals a broader shift: which model you pair with an agent matters as much as the model itself.

2026-07-14

AIFeatured1 min read

Benchmark integrity

GPT-5.6 Sol almost cracked a physics benchmark built so AI couldn't cheat

GPT-5.6 Sol (max) leads CritPt, a new physics benchmark built from unpublished graduate-level research problems, scoring roughly 5 points above GPT-5.5 and 4 points ahead of Claude Fable 5, but even the winner solved only a third of the problems.

2026-07-13

VibeCoding7 min read

Frontier AI Deployment

Anthropic split its most powerful model in two, and that changes how AI gets deployed

Anthropic's Claude Fable 5 brings Mythos-class capability to public users, while Mythos 5 remains trusted-access only. The deployment model, capability routing, fallback classifiers, and tiered access, represents a fundamental shift in how frontier AI is released and used.

2026-07-10

AI4 min read

Artificial Intelligence

Ifbench reveals the instruction-following gap that other benchmarks miss

IFBench measures language models' ability to follow precise natural-language instructions. xAI's Grok models lead while Claude models lag, showing instruction following is a distinct capability from general intelligence.

2026-07-07

← PreviousPage 9 / 2 · 19 articlesNext →