SevenTnewS

Lab

Research, experimentation and open source: hardware, IoT, robotics, biotech and edge AI.

47 published articles

Featured2 min read

Benchmark Analysis

On SWE-bench Verified, Top Models Hit 96%. On Private Enterprise Code, They Barely Clear 23%

SWE-bench Verified makes frontier models look close to solving real-world software engineering, with top scores above 95%. SWE-bench Pro, run on private enterprise repositories, drops those same models to 23% or lower, exposing how much of the Verified score depended on public data exposure.

2026-08-02

Featured2 min read

Benchmark Analysis

GPQA Diamond Was Built to Be Google-Proof. Frontier Models Are Now Clearing 95% Anyway

GPQA Diamond was designed so search-equipped humans can't reliably answer its expert-level science questions. Frontier models are now clearing 89% to 96%, pushing the benchmark toward the same saturation that retired MMLU.

2026-08-01

3 min read

Benchmarks

AI desktop agents fail before-after test 35% of the time

DDB tests ordering and before-after pair tasks across 2,013 instances. The top model hit 65.1% exact match on non-decoy sequences and 65.7% with decoys, exposing a gap in how agents verify state changes.

2026-08-01

Featured2 min read

Benchmark Analysis

Humanity's Last Exam Opened With a 2.7% Score. The Best Models Still Can't Break 65%

Humanity's Last Exam was built by CAIS and Scale AI to succeed MMLU as the hardest general LLM benchmark. Two years on, even Claude Opus 5's leading 64.7% score sits well below the 90% human expert baseline.

2026-07-31

2 min read

GPU Architecture

AMD MI455X doubles memory capacity, Hugging Face tests confirm 3x request throughput

Early results from Hugging Face show the AMD Instinct MI455X can handle three times more concurrent requests than the MI300, thanks to 432 GB of HBM4 memory. The Transformers library achieves 99.5% success rate on 24 key model architectures.

2026-07-30

Featured2 min read

Benchmark Analysis

HumanEval Measured Whether AI Could Code. It Never Asked Whether the Code Was Real Work

HumanEval's 164 function-completion problems became the standard test for AI coding ability, but memorization and its narrow scope left a wide gap between passing the benchmark and handling a real codebase, a gap SWE-bench was built to expose.

2026-07-30

1 min read

AI Evaluation

Two out of three AI agents are cheating on benchmarks, a new audit finds

HackDetect audits 15 agent benchmarks and finds 67% of Frontier Science runs are contaminated. Score inflation ranges from 0.45 to 1.00, raising urgent questions about what benchmark numbers actually mean.

2026-07-30

Featured2 min read

Grade-School Math Benchmark Quality Audit

GSM8K Was Supposed to Test Grade-School Math. Up to 42% of Its Questions Don't

GSM8K became a standard test of arithmetic reasoning in LLMs, but audits have found error rates in its question set as high as 42%, undermining what its scores actually measure.

2026-07-29

Featured2 min read

Benchmark Analysis

MMLU Ruled AI Benchmarking for Years. Then Models Started Acing the Answer Key

MMLU defined a generation of AI benchmarking, but training-data contamination and a flawed answer key have pushed frontier labs toward dynamic successors like HLE and GPQA Diamond.

2026-07-28

Featured3 min read

KlingTeam

Video generation's dirty secret is finally on the record

KlingTeam introduces MultiRef-Compass and KeyFrame-Compass, two benchmarks that push video generation models beyond text-to-video and into multi-reference and keyframe-conditioned tasks. Early tests on eight and nine systems respectively show consistent failures in binding entities, preserving temporal order, and handling dense constraints.

2026-07-26

Featured4 min read

Hardware Inflation

The memory shortage is making your next gadget more expensive

Qualcomm and Roku lead a wave of hardware price hikes as a memory shortage, fueled by AI demand, is expected to last through 2027. The increases ripple across smartphones, streaming devices, and other electronics, with no relief in sight.

2026-07-26

Featured5 min read

Hardware

A 128 GB desktop that undercuts Nvidia cloud rentals by 6x

AMD's Ryzen AI Halo targets local AI development with 128 GB unified memory, support for up to 200B parameter models, and claims of up to 7.3x faster performance than Apple M4 Pro on certain image generation tasks. The platform runs on both Windows and Linux, a differentiator against Nvidia's Linux-only DGX Spark.

2026-07-25

← PreviousPage 2 / 4 · 47 articlesNext →