Benchmarks
20 published articles
Google's workhorse model, bigger in the 3.7 update
Gemini 3.7 Flash halves its price, then makes the case for it
Google's Gemini 3.7 Flash ships at half the intro price of 3.6 Flash while claiming gains on coding, web development, and knowledge-work benchmarks. The release lands three weeks after the previous Flash, and comes with a new price-performance argument for agent builders.
2026-08-16
Artificial Intelligence
Grok 4.6 ties GPT-5.6 Sol at 61, then the component scores split
xAI's Grok 4.6 ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The breakdown is lopsided: wins on knowledge work and most coding tests, losses on DeepSWE and Terminal-Bench. Available now in Cursor and Grok Build, from $2 per million input tokens.
2026-08-13
On-device AI agents
LFM2.5-2.6B: the tiny agent that outruns models 4x its size
Liquid AI's LFM2.5-2.6B fits an agentic model into 2.6B parameters and under 2.5 GB of memory, topping every instruction-following benchmark it was tested on. It runs 220 tokens/s on a laptop; coding is the one clear gap.
2026-08-12
Frontier model access
Anthropic's smartest Claude is the one you can't use
Anthropic launched Claude Fable 5 for everyone and kept Claude Mythos 5 for vetted partners. Same model class, two access doors. The benchmarks matter less than the routing, and the routing is how frontier AI ships from here.
2026-08-03
Messier Corpus
957,253 records can't hide the gap: AI agents surge in coding but stall where enterprises need them
The Messier corpus, with 957,253 records across 30 benchmarks, reveals uneven AI agent progress: function calling is saturated, programming improves fastest, and enterprise workflows lag. It also shows that strict all-pass aggregation can obscure gains and flip leaderboard rankings.
2026-08-02
Open Source
OpenForgeRL trains AI agents without touching their inference harnesses
OpenForgeRL uses a proxy and Kubernetes to train AI agents without modifying their inference harnesses. In tests, OpenForgeGUI matched models several times larger on web navigation and desktop benchmarks.
2026-08-01
GUI Agents
Alibaba's Qwen-UI-Agent scores higher on real phones than in sandboxes
Alibaba Tongyi Lab's Qwen-UI-Agent claims state-of-the-art mobile scores (97.5% AndroidDaily) and competitive computer-use results against Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. Its real-device benchmark score beats its sandbox score, while 40% partial progress on OSWorld-v2 is the honest limit.
2026-07-31
Multi-Agent Engineering
Qoder's multi-agent experiment: 60% fewer mistakes, but at what cost?
An analysis of Alibaba Cloud's Qoder Experts Mode, which deploys specialized AI agents as a coordinated team. Claims of significant error reduction are scrutinized against real-world use cases and the trade-offs of multi-agent complexity.
2026-07-31
AI Models
Macaron-V1-Venti's 748B four-specialist design beats GPT-5.5 and Opus 4.8
Macaron-V1-Venti from MindLab Research leads benchmarks in personal intelligence, coding, terminal use, and Generative UI, using a novel Mixture of LoRA architecture that keeps the model compact while specialized.
2026-07-27
KlingTeam
Video generation's dirty secret is finally on the record
KlingTeam introduces MultiRef-Compass and KeyFrame-Compass, two benchmarks that push video generation models beyond text-to-video and into multi-reference and keyframe-conditioned tasks. Early tests on eight and nine systems respectively show consistent failures in binding entities, preserving temporal order, and handling dense constraints.
2026-07-26
Cyber Defense
Sakana's new cyber agent matches frontier models but tells you not to trust it alone
Sakana AI released Fugu-Cyber, a multi-agent orchestration model matching frontier cyber models on benchmarks. The company argues raw model access alone cannot fix enterprise security without human expertise and verification workflows.
2026-07-21
Foundations of AI
DeepMind's new framework shows why AI is brilliant at some things and terrible at others
A Google DeepMind paper introduces a three-part framework for thinking about rigor in AI: conceptual, epistemic, and operational. It argues that deep learning's rapid progress has come from prioritizing performance-driven iteration over scientific understanding, and that closing the gaps will require more than better benchmarks.
2026-07-20