LLM benchmarks
5 published articles
AI in Education
AI tutors over-help by design. TutorMoments quantifies it
A benchmark built on 462 real tutoring sessions finds AI tutors default to doing the work for students. Explicit prompting helps, but models still struggle with the core judgment call: when to push, when to scaffold, when to step back.
2026-08-10
Voice Agents
Why Grok 4.1 Fast beats smarter models on a phone call
Overall leaderboards pick the wrong LLMs for voice calls. BenchLM's 2026 ranking puts latency first: Grok 4.1 Fast answers in 0.54s while Claude Opus 4.6, the top scorer, needs 1.78s. Fast models take the conversation; reasoning models stay on background tool calls.
2026-08-05
AI Evaluation
The old AI benchmarks broke. Here's what replaced them.
Static benchmarks like MMLU and GSM8K are saturated and contaminated. The industry has moved to dynamic regeneration, expert-level exams, and agentic task environments to get a real measure of AI capability.
2026-08-04
LiveBench leaderboard: cost-performance divergence at the top
The LiveBench top four are separated by 2.2 points. The cost difference is brutal.
GPT-5.6 Sol takes the overall crown on LiveBench with an 82.4 average, but Claude Fable 5 trails by just 1.6 points at nearly three times the cost. The real story is how the pack below has thinned out, and where the dollar smarts stop.
2026-07-22
Frameworks & Tools
Microsoft's Flint hides the chart boilerplate so AI agents stop drawing wrong axes
Microsoft Research introduced Flint, a visualization intermediate language that helps LLMs and AI agents create polished charts without hand-coding low-level parameters like scales and axis formatting. In a study across three models, Flint outperformed direct Vega-Lite generation, and is already used internally in Data Formulator.
2026-07-08