SevenTnewS

Benchmarks

20 published articles

Google DeepMindFeatured3 min read

Google's workhorse model, bigger in the 3.7 update

Gemini 3.7 Flash halves its price, then makes the case for it

Google's Gemini 3.7 Flash ships at half the intro price of 3.6 Flash while claiming gains on coding, web development, and knowledge-work benchmarks. The release lands three weeks after the previous Flash, and comes with a new price-performance argument for agent builders.

2026-08-16

xAI / GrokFeatured4 min read

Artificial Intelligence

Grok 4.6 ties GPT-5.6 Sol at 61, then the component scores split

xAI's Grok 4.6 ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The breakdown is lopsided: wins on knowledge work and most coding tests, losses on DeepSWE and Terminal-Bench. Available now in Cursor and Grok Build, from $2 per million input tokens.

2026-08-13

LLMs & Models4 min read

On-device AI agents

LFM2.5-2.6B: the tiny agent that outruns models 4x its size

Liquid AI's LFM2.5-2.6B fits an agentic model into 2.6B parameters and under 2.5 GB of memory, topping every instruction-following benchmark it was tested on. It runs 220 tokens/s on a laptop; coding is the one clear gap.

2026-08-12

Anthropic / Claude5 min read

Frontier model access

Anthropic's smartest Claude is the one you can't use

Anthropic launched Claude Fable 5 for everyone and kept Claude Mythos 5 for vetted partners. Same model class, two access doors. The benchmarks matter less than the routing, and the routing is how frontier AI ships from here.

2026-08-03

AI5 min read

Messier Corpus

957,253 records can't hide the gap: AI agents surge in coding but stall where enterprises need them

The Messier corpus, with 957,253 records across 30 benchmarks, reveals uneven AI agent progress: function calling is saturated, programming improves fastest, and enterprise workflows lag. It also shows that strict all-pass aggregation can obscure gains and flip leaderboard rankings.

2026-08-02

Open Source2 min read

Open Source

OpenForgeRL trains AI agents without touching their inference harnesses

OpenForgeRL uses a proxy and Kubernetes to train AI agents without modifying their inference harnesses. In tests, OpenForgeGUI matched models several times larger on web navigation and desktop benchmarks.

2026-08-01

AI Agents3 min read

GUI Agents

Alibaba's Qwen-UI-Agent scores higher on real phones than in sandboxes

Alibaba Tongyi Lab's Qwen-UI-Agent claims state-of-the-art mobile scores (97.5% AndroidDaily) and competitive computer-use results against Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. Its real-device benchmark score beats its sandbox score, while 40% partial progress on OSWorld-v2 is the honest limit.

2026-07-31

AI AgentsFeatured2 min read

Multi-Agent Engineering

Qoder's multi-agent experiment: 60% fewer mistakes, but at what cost?

An analysis of Alibaba Cloud's Qoder Experts Mode, which deploys specialized AI agents as a coordinated team. Claims of significant error reduction are scrutinized against real-world use cases and the trade-offs of multi-agent complexity.

2026-07-31

LLMs & ModelsFeatured3 min read

AI Models

Macaron-V1-Venti's 748B four-specialist design beats GPT-5.5 and Opus 4.8

Macaron-V1-Venti from MindLab Research leads benchmarks in personal intelligence, coding, terminal use, and Generative UI, using a novel Mixture of LoRA architecture that keeps the model compact while specialized.

2026-07-27

LabFeatured3 min read

KlingTeam

Video generation's dirty secret is finally on the record

KlingTeam introduces MultiRef-Compass and KeyFrame-Compass, two benchmarks that push video generation models beyond text-to-video and into multi-reference and keyframe-conditioned tasks. Early tests on eight and nine systems respectively show consistent failures in binding entities, preserving temporal order, and handling dense constraints.

2026-07-26

AIFeatured5 min read

Cyber Defense

Sakana's new cyber agent matches frontier models but tells you not to trust it alone

Sakana AI released Fugu-Cyber, a multi-agent orchestration model matching frontier cyber models on benchmarks. The company argues raw model access alone cannot fix enterprise security without human expertise and verification workflows.

2026-07-21

AIFeatured5 min read

Foundations of AI

DeepMind's new framework shows why AI is brilliant at some things and terrible at others

A Google DeepMind paper introduces a three-part framework for thinking about rigor in AI: conceptual, epistemic, and operational. It argues that deep learning's rapid progress has come from prioritizing performance-driven iteration over scientific understanding, and that closing the gaps will require more than better benchmarks.

2026-07-20

← PreviousPage 1 / 2 · 20 articlesNext →