SevenTnewS

Features

In-depth special reports combining feature, analysis, interview and profile reporting.

31 published features

4 min read

Artificial Intelligence

The hardest lesson for AI reasoning engines: when to shut up

MIT researchers propose OS-Pruner, a plug-in that dynamically stops chain-of-thought reasoning when further computation isn't worth the token cost. Tests show 20-60% length reduction with minimal accuracy sacrifice.

2026-07-29

6 min read

Agentic Finance

Why returns alone lie: NextFund opens the black box on AI trading agents

NextFund records every decision an AI trading agent makes in live markets, letting users compare models, inspect rationales, and diagnose failures. An eight-LLM test across US, China, and Hong Kong equities shows that similar intermediate signals can diverge into starkly different portfolio behaviors, and that quarterly rankings flip depending on which metric matters most.

2026-07-27

Featured5 min read

Cyber Defense

Sakana's new cyber agent matches frontier models but tells you not to trust it alone

Sakana AI released Fugu-Cyber, a multi-agent orchestration model matching frontier cyber models on benchmarks. The company argues raw model access alone cannot fix enterprise security without human expertise and verification workflows.

2026-07-21

7 min read

Biologically plausible AI's hidden warning

The CIFAR-10 ablation that should make every AI researcher rethink their benchmarks

Sakana AI's Error Diffusion scales brain-like, backprop-free learning to CIFAR-10 and RL. The catch is in the ablation tables: which design choices matter reverses entirely between benchmarks, and the cost of Dale's principle grows with task difficulty.

2026-07-19

5 min read

Governed agentic research

Nvidia's AI ran a hospital study on 286,000 patients, and never touched their data

Nvidia's AI Technology Center unveils NAIS, a governed agentic research system that orchestrates end-to-end biomedical workflows on protected hospital data. In a real-world hypertension GWAS deployment involving 286,422 individuals, the system produced results comparable to expert-led analyses while preserving privacy and enabling human oversight.

2026-07-19

Featured5 min read

local ai

The uncensored model paradox: one knob removes both the annoying refusals and the safety guardrail

Benchmarking five locally run uncensored LLMs shows abliteration cuts over-refusal from 44% to near zero with no hit to reasoning, but the same edit collapses safety refusals from 41.5% to 9.5%, because both ride on the same internal direction. The real reason to run uncensored may not be what you think.

2026-07-19

6 min read

Model Orchestration

The model choice is now the router's problem, not yours

Sakana AI's Fugu routes each task to a specialized model behind a single API, promising less integration overhead and better cost-performance. It lands in a fast-forming category where the model itself is becoming an implementation detail.

2026-07-18

Featured5 min read

special report / edge ai

Nvidia just cracked open DeepStream. Your edge AI project will never be the same.

The full source code of Nvidia's DeepStream video analytics SDK is now on GitHub under Apache 2.0 and CC-BY-4.0, opening edge AI development to a wider audience. Version 9.1 brings LLM-based coding agents, Triton Inference Server integration, and consolidated repositories for end-to-end pipelines.

2026-07-16

5 min read

Machine Learning Theory

Google just proved why diffusion models invent, not just copy

Google researchers reveal that the creativity of diffusion models stems from a 'score smoothing' effect caused by neural network regularization. This theoretical framework explains why models interpolate between training data points rather than merely memorizing them, opening the path for controlled novelty in generative AI.

2026-07-16

5 min read

Portfolio strategy autopsy

Mistral killed half its model family. What survived tells you everything.

A mid-2026 audit of Mistral AI's model portfolio reveals 36 active models across frontier, specialist, and legacy tiers, with 19 models slated for deprecation. The company's strategy emphasizes small specialist models for agents and coding, while retiring experimental variants like Magistrate and earlier Devstral versions.

2026-07-13

Featured4 min read

Special Report

OpenAI's bet on shared agents is the quietest shift in enterprise AI this year

OpenAI launches workspace agents: persistent, cloud-based AI workers that run across ChatGPT and Slack, handle multi-step workflows, and share context across teams. Free until May 6, 2026, then credit-based. A structural shift from GPTs to organizational AI.

2026-07-11

Featured3 min read

Grok 4.5

Cursor's Grok 4.5 was built by AI agents, not humans. That's the real story.

Cursor's Grok 4.5 is a Mixture-of-Experts model built using reinforcement learning in environments created by earlier AI agents, not humans. It handles complex, long-duration tasks across software engineering, data science, finance, and law, and it's available now.

2026-07-10

← PreviousPage 2 / 3 · 31 featuresNext →