SevenTnewS

Features

In-depth special reports combining feature, analysis, interview and profile reporting.

31 published features

6 min read

Frontier video models: special report

Perceptron Mk1 bets video AI's future on watching, not making

Perceptron Mk1 is a closed-source vision-language model priced for continuous use, aimed at timestamped video understanding and embodied reasoning. The real AI video race is not about making prettier clips, and benchmarks that blend both categories mislead.

2026-08-06

Featured4 min read

AI Safety

Shieldstral, the 3B classifier that outguns models seven times its size

A 3B safety classifier matches text models nearly seven times its size and sets a new multimodal moderation state of the art, per a July 2026 arXiv paper. The trick: moderation reframed as binary question answering, trained on roughly 54.1 million samples.

2026-08-05

Featured5 min read

AI Evaluation

The old AI benchmarks broke. Here's what replaced them.

Static benchmarks like MMLU and GSM8K are saturated and contaminated. The industry has moved to dynamic regeneration, expert-level exams, and agentic task environments to get a real measure of AI capability.

2026-08-04

5 min read

AI agents

Treating SOPs as code: why compilation alone lifts strong agents by 16 points

New research from Hong Kong and mainland China demonstrates that compiling SOPs into executable pseudo-code and running them on a stack-paged virtual machine cleanly separates capable agents from brittle ones. The work yields a precise deployment rule: compile first, page only after a model-level discipline check.

2026-08-04

Featured8 min read

Special Report: AI Evaluation

The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them

A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.

2026-08-03

5 min read

Artificial Intelligence

The AI that remembers every failure it fixes, and gets better with each one

OpsMem couples a short-term memory (STM) for the evolving diagnostic state with a long-term memory (LTM) for reusable operational experience. Through a mechanism called cross-memory resonance, the system activates state-relevant experience from LTM to guide multi-agent diagnosis. On a Huawei microservice dataset, it outperformed agentic-reasoning and knowledge-augmented baselines by significant margins.

2026-08-03

5 min read

AI interpretability

Your neural network's black-box decisions just got a discoverable memory

Researchers show that neural network action scores can be expressed as exact weighted sums of training-case returns, using Gram geometry. This allows post-training audit signals that identify influential cases, measure action coherence, and flag weak support, without retraining or accessing the original optimization trajectory.

2026-08-03

3 min read

AI Research

The monitor that goes silent when AI reasoning fails

Novelis Research shows token log-probability fails as a decoder monitor for quantized reasoning models, being blind to confident loops. They introduce a calibrated e-CUSUM controller that combines uncertainty and repetition signals, achieving selectivity on GSM8K with DeepSeek-R1-Distill-Qwen-1.5B.

2026-08-03

4 min read

AI research

Over half your AI's reasoning is froth, and nobody noticed until now

LLMs often generate reasoning chains that are correct but padded with unnecessary steps. A new diagnostic benchmark shows current evaluators miss this inefficiency entirely, and half of human-written reasoning steps may be compressible.

2026-07-31

3 min read

Formal mathematics

The bottleneck no one saw in AI math proofs: decomposition, not compute

Nanjing University's ToMap framework achieves state-of-the-art results in full-proof autoformalization by identifying the decomposition step as the critical bottleneck. Using iterative, Pareto-guided evolution of proof decompositions, ToMap lifts joint syntactic-semantic accuracy by 19% on the ProofFlowBench benchmark while reducing test-time costs.

2026-07-30

5 min read

Research

Six words that could break the AI agent safety ceiling: disrupt, validate, broker

A novel heterogeneous agent cohort architecture separates divergent exploration, runtime safety gating, and cross-domain knowledge retrieval into specialized roles. The Disrupter generates high-entropy proposals, the Validator enforces hard tool-call checks, and the Broker imports out-of-domain analogies via contrastive novelty retrieval. Execution failures are compiled into signed constraint patches called Scars, cached for future generations. In evaluations, the cohort achieved 95% remote target discovery, zero executed breaches, and 15.1% token savings from Scars, with a 55.9% cost reduction under resource constraints via credit-based bandwidth allocation.

2026-07-30

4 min read

Enterprise analytics

Alibaba built a data agent that finally understands what 'valid users' means

Alibaba's QwenPaw-Data introduces a three-subsystem architecture, DataBridge, Skill-Hub, and Host, to turn fragmented enterprise data into reusable analytical assets. Early evaluations show substantial gains over existing systems on both public benchmarks and industrial BI workloads.

2026-07-30

← PreviousPage 1 / 3 · 31 featuresNext →