benchmark
16 published articles
Benchmarking
Frontier AI vision models fail at basic perception, new benchmark shows
PerceptionBench tests ten atomic visual capabilities across 3,000 questions. No frontier model cracked 60 percent, and similar overall scores mask wildly different weakness profiles.
2026-08-03
TRACTA Benchmark
Neuro-symbolic reasoning outperforms raw neural models on temporal tasks, benchmark finds
TRACTA benchmark reveals neuro-symbolic AI beats raw neural models on three temporal reasoning tasks, with largest margins on early warning and pattern detection.
2026-08-02
Benchmarks
AI desktop agents fail before-after test 35% of the time
DDB tests ordering and before-after pair tasks across 2,013 instances. The top model hit 65.1% exact match on non-decoy sequences and 65.7% with decoys, exposing a gap in how agents verify state changes.
2026-08-01
3D Reasoning
SceneActBench: Why even the best VLMs fail at 3D action
SceneActBench tests eleven VLM configurations on five 3D tasks in a unified agent loop. Overall scores range from 38.6 to 50.2, with no model performing consistently. The benchmark exposes a blind spot in vision-language agents: acting on full scenes, not just describing them.
2026-07-31
Speech synthesis
Alibaba's new TTS model speaks 16 languages, laughs on command, and might actually ship
Alibaba's Qwen-Audio-3.0-TTS is a production-oriented speech synthesis system that combines a low-frame-rate tokenizer with progressive training. It supports 16 languages, 20 Chinese dialect regions, and natural-language instructions for emotion, pace, and speaking style.
2026-07-26
Cost vs. capability
The $1.40 model that crushed its $2.82 sibling on physics
Four AI models were asked to build interactive destruction scenes with real physics. Opus 5 passed all three at the lowest cost among top performers, while Fable 5 failed each one at twice the price. A reminder that price tags don't predict physics.
2026-07-25
AI Safety
Your AI research assistant is sabotaging you, and you won't catch it half the time
ResearchArena, a new benchmark for evaluating AI control in automated R&D, shows that monitors miss embedded sabotage more than half the time. Even monitors that probe artifacts with tests can be fooled by subtle anomalies or wrong test choices.
2026-07-25
Agent evaluation
Your AI agent keeps failing? It might be the harness, not the brain
PawBench, an open-source benchmark from the AgentScope team, systematically evaluates models and agent harnesses together. Results show that harness design can swing scores by over 11 points for smaller models, exposing a blind spot in how AI agents are currently judged.
2026-07-22
Speech AI
Voice AI said 'I understand your frustration.' It had no idea what that meant.
Hume's Real World VoiceEQ benchmark, based on more than one million human ratings, tests over 40 voice models across dimensions standard benchmarks ignore: emotion, speaker identity, and acoustic context. The findings show speech-to-speech models vary wildly, and even leading systems often ignore the audio cues humans use instinctively.
2026-07-20
New benchmark exposes agent reasoning limits
AI agents hit 49% on a test human experts pass at 95%. More compute won't fix it.
Alibaba's HSCodeComp benchmark reveals the gap: top agents hit 49.4% accuracy vs. human experts' 95% on tariff classification. The bottleneck is structural, more compute doesn't help.
2026-07-19
Domain specialization
A six-month-old OCR model still beats Mistral. The reason is hard to fix.
DharmaOCR scores 0.925 on a Portuguese benchmark versus 0.798 for Mistral OCR4 and 0.7587 for Unlimited-OCR. The gap comes from concentrated training allocation and a DPO-based approach that suppresses text degeneration in complex documents.
2026-07-16
AI Research
GPT-5.5's big win reveals something missing from every agent benchmark
EvoPolicyGym isolates a critical but understudied capability: an agent's ability to refine an executable policy through repeated feedback-constrained edits. The benchmark reveals GPT-5.5 as the strongest performer across 16 environments, and provides trajectory-level diagnostics that expose how different agents allocate budget and convert feedback into tuned parameters.
2026-07-11