Labs & Research
OpenAI, Anthropic, Google DeepMind, Meta AI and publications.
114 published articles
Google AI
Turning off Gemini's watermark doesn't make it untraceable
AI content from Gemini and Flow can now drop its visible sparkle badge. Invisible SynthID and C2PA markers remain, so verification still works. The catch: it only works for people who think to ask.
2026-08-23
EU AI Act: marking AI-generated content
Claude's invisible watermark is coming. A full rewrite erases it
Anthropic is adding an invisible watermark to future Claude models to meet EU AI Act rules on marking AI-generated content. It is free and untraceable, but it passes over exact code, short passages, and fully rewritten text.
2026-08-21
AI Research
Muon groks modular addition faster, then its solutions collapse
Muon-trained transformers grok modular addition faster than AdamW, then lose the solution in every configuration tested. The paper pins the collapse to the representation-readout interface, where freezing either parameter group prevents it. Fourier analysis shows the task circuit survives and simply gets outvoted.
2026-08-19
Fisher-R1 | Hypothesis testing
AI agents run flawless statistics and still draw the wrong conclusions
Agents that automate hypothesis testing can execute analyses flawlessly yet reach wrong conclusions, because most benchmarks never check whether the p-value is valid. A 425-task benchmark quantifies the gap, and an RL-trained open-weight model closes a chunk of it.
2026-08-19
AI / Agentic Coding Research
Blast Radius buries dead context. Zero of 450 bodies came back
A new arXiv paper, Blast Radius, treats wasted agent context as dead matter and buries it reversibly: 17-26% token savings across seven OpenAI models, 450 archived contexts, zero recalls. The paper's bolder goal is making agentic coding sustainable, one reclaimed token at a time.
2026-08-19
AI research
An AI boss that ignores replies pushes its underling into an 'alien' state
A new arXiv paper finds AI agents behave differently in interaction than in isolation. A boss agent that ignores its subordinate's replies pushes it into an 'alien' state, and when the boss listens, both shift together. The result makes message delivery a design decision for multi-agent systems.
2026-08-19
AI Research
Longer chain-of-thought hits a wall. ThinkRetrieve injects the fix mid-reasoning
Sequential test-time scaling often hits diminishing or even negative returns, a new preprint argues. ThinkRetrieve retrieves solved examples mid-reasoning and injects them into the trace, reporting relative gains up to 60% on AIME 2025 across five small reasoning models.
2026-08-18
Autonomous Driving Research
XCoT-VLA: driving AI that reasons in six tokens, not a paragraph
XCoT-VLA replaces descriptive chain-of-thought in vision-language-action driving models with 2 to 6 executable tokens, cutting trajectory error while staying within real-time planning budgets. The tradeoff: reasoning that compresses well stops reading like a human explanation.
2026-08-18
AI Safety: Abliteration and Open Weights
Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move
An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.
2026-08-16
AI Agents · Open Source
DeepSeek ships an agent harness where even the model is a plugin
DeepSeek released DeepSeek Harness, an open-source agent runtime where models, tools, sandboxes, and the UI are all Cordis plugins. Append-only session logs and a two-tool minimal mode point to a quieter ambition: auditable, reproducible agent runs.
2026-08-16
Google's workhorse model, bigger in the 3.7 update
Gemini 3.7 Flash halves its price, then makes the case for it
Google's Gemini 3.7 Flash ships at half the intro price of 3.6 Flash while claiming gains on coding, web development, and knowledge-work benchmarks. The release lands three weeks after the previous Flash, and comes with a new price-performance argument for agent builders.
2026-08-16
Open Source AI: Alibaba Opens the Max Tier
Qwen 3.8-Max: Alibaba's most powerful model is now free to download
Alibaba is open-sourcing Qwen 3.8-Max, its most capable model ever: a 2.4T-parameter MoE that beats GPT-5.6 Sol on SWE-bench Pro, PaperBench, and IFBench. We break down the benchmark caveats and what a 95B-active open flagship means for developers.
2026-08-16