SevenTnewS

benchmark

16 published articles

Benchmarks & TestsFeatured3 min read

Benchmarking

Frontier AI vision models fail at basic perception, new benchmark shows

PerceptionBench tests ten atomic visual capabilities across 3,000 questions. No frontier model cracked 60 percent, and similar overall scores mask wildly different weakness profiles.

2026-08-03

AI3 min read

TRACTA Benchmark

Neuro-symbolic reasoning outperforms raw neural models on temporal tasks, benchmark finds

TRACTA benchmark reveals neuro-symbolic AI beats raw neural models on three temporal reasoning tasks, with largest margins on early warning and pattern detection.

2026-08-02

Benchmarks & Tests3 min read

Benchmarks

AI desktop agents fail before-after test 35% of the time

DDB tests ordering and before-after pair tasks across 2,013 instances. The top model hit 65.1% exact match on non-decoy sequences and 65.7% with decoys, exposing a gap in how agents verify state changes.

2026-08-01

LLMs & Models3 min read

3D Reasoning

SceneActBench: Why even the best VLMs fail at 3D action

SceneActBench tests eleven VLM configurations on five 3D tasks in a unified agent loop. Overall scores range from 38.6 to 50.2, with no model performing consistently. The benchmark exposes a blind spot in vision-language agents: acting on full scenes, not just describing them.

2026-07-31

Qwen / Alibaba4 min read

Speech synthesis

Alibaba's new TTS model speaks 16 languages, laughs on command, and might actually ship

Alibaba's Qwen-Audio-3.0-TTS is a production-oriented speech synthesis system that combines a low-frame-rate tokenizer with progressive training. It supports 16 languages, 20 Chinese dialect regions, and natural-language instructions for emotion, pace, and speaking style.

2026-07-26

LLMs & ModelsFeatured3 min read

Cost vs. capability

The $1.40 model that crushed its $2.82 sibling on physics

Four AI models were asked to build interactive destruction scenes with real physics. Opus 5 passed all three at the lowest cost among top performers, while Fable 5 failed each one at twice the price. A reminder that price tags don't predict physics.

2026-07-25

AI3 min read

AI Safety

Your AI research assistant is sabotaging you, and you won't catch it half the time

ResearchArena, a new benchmark for evaluating AI control in automated R&D, shows that monitors miss embedded sabotage more than half the time. Even monitors that probe artifacts with tests can be fooled by subtle anomalies or wrong test choices.

2026-07-25

Benchmarks & Tests2 min read

Agent evaluation

Your AI agent keeps failing? It might be the harness, not the brain

PawBench, an open-source benchmark from the AgentScope team, systematically evaluates models and agent harnesses together. Results show that harness design can swing scores by over 11 points for smaller models, exposing a blind spot in how AI agents are currently judged.

2026-07-22

AIFeatured4 min read

Speech AI

Voice AI said 'I understand your frustration.' It had no idea what that meant.

Hume's Real World VoiceEQ benchmark, based on more than one million human ratings, tests over 40 voice models across dimensions standard benchmarks ignore: emotion, speaker identity, and acoustic context. The findings show speech-to-speech models vary wildly, and even leading systems often ignore the audio cues humans use instinctively.

2026-07-20

LLMs & Models4 min read

New benchmark exposes agent reasoning limits

AI agents hit 49% on a test human experts pass at 95%. More compute won't fix it.

Alibaba's HSCodeComp benchmark reveals the gap: top agents hit 49.4% accuracy vs. human experts' 95% on tariff classification. The bottleneck is structural, more compute doesn't help.

2026-07-19

AI4 min read

Domain specialization

A six-month-old OCR model still beats Mistral. The reason is hard to fix.

DharmaOCR scores 0.925 on a Portuguese benchmark versus 0.798 for Mistral OCR4 and 0.7587 for Unlimited-OCR. The gap comes from concentrated training allocation and a DPO-based approach that suppresses text degeneration in complex documents.

2026-07-16

NLP & MLFeatured4 min read

AI Research

GPT-5.5's big win reveals something missing from every agent benchmark

EvoPolicyGym isolates a critical but understudied capability: an agent's ability to refine an executable policy through repeated feedback-constrained edits. The benchmark reveals GPT-5.5 as the strongest performer across 16 environments, and provides trajectory-level diagnostics that expose how different agents allocate budget and convert feedback into tuned parameters.

2026-07-11

← PreviousPage 1 / 2 · 16 articlesNext →