AI
Artificial intelligence: LLMs, agents, diffusion, vision, NLP and the latest from top labs.
565 published articles
AI research
An AI boss that ignores replies pushes its underling into an 'alien' state
A new arXiv paper finds AI agents behave differently in interaction than in isolation. A boss agent that ignores its subordinate's replies pushes it into an 'alien' state, and when the boss listens, both shift together. The result makes message delivery a design decision for multi-agent systems.
2026-08-19
AI slop
Nobody can define AI slop. SlopFinder is averaging the answer
SlopFinder collects one-click anonymous votes on AI-generated text and exports the averaged results to Hugging Face. The project's premise is that 'slop' is measurable even if it is undefinable, which makes it easier to build on.
2026-08-18
Federated Learning
When users can't be shared, FedCGR shares a language instead
FedCGR treats federated cross-domain recommendation as generation over a stable semantic item language. Items become discrete semantic IDs drawn from public metadata, so clients align without exchanging private interactions. The trade-off: a semantic-only bottleneck that local collaborative filtering evidence must fill.
2026-08-18
AI Research
Longer chain-of-thought hits a wall. ThinkRetrieve injects the fix mid-reasoning
Sequential test-time scaling often hits diminishing or even negative returns, a new preprint argues. ThinkRetrieve retrieves solved examples mid-reasoning and injects them into the trace, reporting relative gains up to 60% on AIME 2025 across five small reasoning models.
2026-08-18
Autonomous Driving Research
XCoT-VLA: driving AI that reasons in six tokens, not a paragraph
XCoT-VLA replaces descriptive chain-of-thought in vision-language-action driving models with 2 to 6 executable tokens, cutting trajectory error while staying within real-time planning budgets. The tradeoff: reasoning that compresses well stops reading like a human explanation.
2026-08-18
Knowledge Graphs & Urban AI
Stations aren't islands: how RTSKG rewires urban transit data
City-scale transit models often ignore how stations relate to roads and businesses. RTSKG, a knowledge graph dataset posted to arXiv, models those interactions explicitly and reports gains on store recommendation and ridership prediction.
2026-08-18
Open Source AI
Boris-2 feeds its 125M model 90B tokens, its 250M just 60B
Boris-2 pre-announces three small models with lopsided token budgets: the 125M gets 90B tokens, the 250M just 60B. Training runs from August 12 to September 10, with no benchmarks shared, only the stated ambition to reach SmolLM2-135M strength.
2026-08-16
AI Video Generation
Alibaba's Wan3.0 sells video by the second: $6 for a 30-second clip
Wan3.0 can turn a PDF or a brand deck into a 30-second video in one generation, with per-second API pricing that tops out at $0.20 for 1080P. We break down the per-clip math and the gaps Alibaba admits in its own testing.
2026-08-16
AI Safety: Abliteration and Open Weights
Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move
An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.
2026-08-16
Open-weight speech for production voice agents
Magpie TTS spends 32ms of your voice agent's latency budget
Magpie TTS reports 32ms time-to-first-audio on an NVIDIA B200 and adds Arabic, Korean and Brazilian Portuguese, bringing its roster to 12 languages. The open-weights pitch: self-hosted speech synthesis no longer loses the latency argument.
2026-08-16
Red-Teaming AI Swarms
IO Factory's 100,000-agent sandbox gives influence campaigns nowhere to hide
IO Factory simulates coordinated AI influence campaigns as swarms of up to 100,000 agents that adapt to platform feedback and hide in ordinary social interaction. The paper argues detection must track whole operations, not isolated messages.
2026-08-16
CARE-X Medical AI
Mild aortic dilation: 12% caught on first reads, 93% with a measuring VLM
Aortic dilation is rarely quantified on chest X-rays, and mild cases get missed: 5 of 43 on initial reads. A tool-augmented VLM caught 40. New CARE-X research explains why radiology AI needs rulers, not just eyes.
2026-08-16