SevenTnewS

Lab

Research, experimentation and open source: hardware, IoT, robotics, biotech and edge AI.

47 published articles

5 min read

Consumer Hardware

CMF, Elektron and OhSnap show budget gear beats feature bloat

CMF's $79 Clip Pro earbuds, Elektron's $349 Model grooveboxes and OhSnap's $50 Snap Grip Stand all won the Verge's reviews by owning their trade-offs. The same pattern shows up in AR glasses. Buyers should take note.

2026-08-23

4 min read

Financial reasoning benchmark

LLMs know accounting formulas. FinIndices shows they can't apply them

LLMs can recite accounting formulas, but FinIndices, a benchmark built on real Chinese financial statements, shows they struggle to apply them. Gemini-3.1-Pro's accuracy fell from 70.70% to 38.22% when formula hints were removed, and generating tables actively degraded models' reasoning.

2026-08-20

4 min read

HDR10+ ADVANCED: Samsung and Prime Video ship it first

Can scene-by-scene metadata fix motion smoothing's bad reputation?

Prime Video is the first streaming service to adopt HDR10+ ADVANCED on Samsung's 2026 TVs from August 2026. The standard adds Enhanced Overall Brightness and scene-aware Intelligent Motion Smoothing, the setting with the worst reputation in home theater, now backed by smarter metadata.

2026-08-11

5 min read

AI in Education

AI tutors over-help by design. TutorMoments quantifies it

A benchmark built on 462 real tutoring sessions finds AI tutors default to doing the work for students. Explicit prompting helps, but models still struggle with the core judgment call: when to push, when to scaffold, when to step back.

2026-08-10

5 min read

Voice Agents

Why Grok 4.1 Fast beats smarter models on a phone call

Overall leaderboards pick the wrong LLMs for voice calls. BenchLM's 2026 ranking puts latency first: Grok 4.1 Fast answers in 0.54s while Claude Opus 4.6, the top scorer, needs 1.78s. Fast models take the conversation; reasoning models stay on background tool calls.

2026-08-05

Featured5 min read

AI Evaluation

The old AI benchmarks broke. Here's what replaced them.

Static benchmarks like MMLU and GSM8K are saturated and contaminated. The industry has moved to dynamic regeneration, expert-level exams, and agentic task environments to get a real measure of AI capability.

2026-08-04

Featured3 min read

Architecture & Cloud

Why your IoT dashboard shows yesterday's data, and how to fix it

Most IoT monitoring platforms fail not at data collection but at reliable delivery. A layered architecture using Alibaba Cloud DataWorks automates scheduled synchronization, eliminates risks like duplication and latency, and keeps your dashboards current without custom middleware.

2026-08-03

Featured8 min read

Special Report: AI Evaluation

The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them

A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.

2026-08-03

Featured3 min read

Benchmarking

Frontier AI vision models fail at basic perception, new benchmark shows

PerceptionBench tests ten atomic visual capabilities across 3,000 questions. No frontier model cracked 60 percent, and similar overall scores mask wildly different weakness profiles.

2026-08-03

Featured2 min read

Benchmark Analysis

LiveBench Refuses to Sit Still. That's the Whole Point

LiveBench replaces a fixed answer key with a continuously refreshed pool of coding problems, repositories, and prediction questions, sidestepping the contamination that undermined MMLU and GSM8K. It's part of a broader shift toward dynamic evaluation alongside LiveCodeBench, ForecastBench, and LLMEval-Fair.

2026-08-03

4 min read

Model Review

The 7B model that just made GPT-4o-mini look expensive

A deep dive into Alibaba's open-source omni-model: processes text, images, audio, and video simultaneously while generating streaming speech; outperforms GPT-4o-mini and Gemini on multiple benchmarks; fits on a single consumer GPU. The catch? Text-only reasoning takes a hit.

2026-08-03

Featured2 min read

Benchmark Analysis

5 Million Votes, 25 Elo Points: Inside the Chatbot Arena Leaderboard Nobody Can Shake

LMSYS Chatbot Arena ranks models by blind human preference on an Elo scale, with nearly five million votes now packing the top labs into a 25-point band. The LLM-as-a-Judge methods that scale this kind of evaluation carry their own documented biases toward verbosity and position.

2026-08-02

← PreviousPage 1 / 4 · 47 articlesNext →