SevenTnewS

Labs & Research

OpenAI, Anthropic, Google DeepMind, Meta AI and publications.

114 published articles

4 min read

Google AI

Turning off Gemini's watermark doesn't make it untraceable

AI content from Gemini and Flow can now drop its visible sparkle badge. Invisible SynthID and C2PA markers remain, so verification still works. The catch: it only works for people who think to ask.

2026-08-23

5 min read

EU AI Act: marking AI-generated content

Claude's invisible watermark is coming. A full rewrite erases it

Anthropic is adding an invisible watermark to future Claude models to meet EU AI Act rules on marking AI-generated content. It is free and untraceable, but it passes over exact code, short passages, and fully rewritten text.

2026-08-21

4 min read

AI Research

Muon groks modular addition faster, then its solutions collapse

Muon-trained transformers grok modular addition faster than AdamW, then lose the solution in every configuration tested. The paper pins the collapse to the representation-readout interface, where freezing either parameter group prevents it. Fourier analysis shows the task circuit survives and simply gets outvoted.

2026-08-19

3 min read

Fisher-R1 | Hypothesis testing

AI agents run flawless statistics and still draw the wrong conclusions

Agents that automate hypothesis testing can execute analyses flawlessly yet reach wrong conclusions, because most benchmarks never check whether the p-value is valid. A 425-task benchmark quantifies the gap, and an RL-trained open-weight model closes a chunk of it.

2026-08-19

5 min read

AI / Agentic Coding Research

Blast Radius buries dead context. Zero of 450 bodies came back

A new arXiv paper, Blast Radius, treats wasted agent context as dead matter and buries it reversibly: 17-26% token savings across seven OpenAI models, 450 archived contexts, zero recalls. The paper's bolder goal is making agentic coding sustainable, one reclaimed token at a time.

2026-08-19

4 min read

AI research

An AI boss that ignores replies pushes its underling into an 'alien' state

A new arXiv paper finds AI agents behave differently in interaction than in isolation. A boss agent that ignores its subordinate's replies pushes it into an 'alien' state, and when the boss listens, both shift together. The result makes message delivery a design decision for multi-agent systems.

2026-08-19

4 min read

AI Research

Longer chain-of-thought hits a wall. ThinkRetrieve injects the fix mid-reasoning

Sequential test-time scaling often hits diminishing or even negative returns, a new preprint argues. ThinkRetrieve retrieves solved examples mid-reasoning and injects them into the trace, reporting relative gains up to 60% on AIME 2025 across five small reasoning models.

2026-08-18

4 min read

Autonomous Driving Research

XCoT-VLA: driving AI that reasons in six tokens, not a paragraph

XCoT-VLA replaces descriptive chain-of-thought in vision-language-action driving models with 2 to 6 executable tokens, cutting trajectory error while staying within real-time planning budgets. The tradeoff: reasoning that compresses well stops reading like a human explanation.

2026-08-18

5 min read

AI Safety: Abliteration and Open Weights

Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move

An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.

2026-08-16

4 min read

AI Agents · Open Source

DeepSeek ships an agent harness where even the model is a plugin

DeepSeek released DeepSeek Harness, an open-source agent runtime where models, tools, sandboxes, and the UI are all Cordis plugins. Append-only session logs and a two-tool minimal mode point to a quieter ambition: auditable, reproducible agent runs.

2026-08-16

Featured3 min read

Google's workhorse model, bigger in the 3.7 update

Gemini 3.7 Flash halves its price, then makes the case for it

Google's Gemini 3.7 Flash ships at half the intro price of 3.6 Flash while claiming gains on coding, web development, and knowledge-work benchmarks. The release lands three weeks after the previous Flash, and comes with a new price-performance argument for agent builders.

2026-08-16

5 min read

Open Source AI: Alibaba Opens the Max Tier

Qwen 3.8-Max: Alibaba's most powerful model is now free to download

Alibaba is open-sourcing Qwen 3.8-Max, its most capable model ever: a 2.4T-parameter MoE that beats GPT-5.6 Sol on SWE-bench Pro, PaperBench, and IFBench. We break down the benchmark caveats and what a 95B-active open flagship means for developers.

2026-08-16

← PreviousPage 1 / 10 · 114 articlesNext →