SevenTnewS

LLM benchmarks

5 published articles

Benchmarks & Tests5 min read

AI in Education

AI tutors over-help by design. TutorMoments quantifies it

A benchmark built on 462 real tutoring sessions finds AI tutors default to doing the work for students. Explicit prompting helps, but models still struggle with the core judgment call: when to push, when to scaffold, when to step back.

2026-08-10

Benchmarks & Tests5 min read

Voice Agents

Why Grok 4.1 Fast beats smarter models on a phone call

Overall leaderboards pick the wrong LLMs for voice calls. BenchLM's 2026 ranking puts latency first: Grok 4.1 Fast answers in 0.54s while Claude Opus 4.6, the top scorer, needs 1.78s. Fast models take the conversation; reasoning models stay on background tool calls.

2026-08-05

Benchmarks & TestsFeatured5 min read

AI Evaluation

The old AI benchmarks broke. Here's what replaced them.

Static benchmarks like MMLU and GSM8K are saturated and contaminated. The industry has moved to dynamic regeneration, expert-level exams, and agentic task environments to get a real measure of AI capability.

2026-08-04

AIFeatured4 min read

LiveBench leaderboard: cost-performance divergence at the top

The LiveBench top four are separated by 2.2 points. The cost difference is brutal.

GPT-5.6 Sol takes the overall crown on LiveBench with an 82.4 average, but Claude Fable 5 trails by just 1.6 points at nearly three times the cost. The real story is how the pack below has thinned out, and where the dollar smarts stop.

2026-07-22

Tools & Frameworks4 min read

Frameworks & Tools

Microsoft's Flint hides the chart boilerplate so AI agents stop drawing wrong axes

Microsoft Research introduced Flint, a visualization intermediate language that helps LLMs and AI agents create polished charts without hand-coding low-level parameters like scales and axis formatting. In a study across three models, Flint outperformed direct Vega-Lite generation, and is already used internally in Data Formulator.

2026-07-08