Labs & Research
OpenAI, Anthropic, Google DeepMind, Meta AI and publications.
114 published articles
AI Agents
Grok Bot goes after the last 10% of work most AI leaves undone
Grok Bot, xAI's new agent product, runs on its own cloud computer, signs into the tools you already use, learns workflows by watching you do them once, and finishes jobs end to end. The pitch behind it: most AI stops at 90% done.
2026-08-14
AI Economics
Alibaba's $18 plan runs Qwen and DeepSeek inside Claude Code
Alibaba Cloud's Token Plan Individual bundles Qwen and third-party models such as DeepSeek into a single credit pool, from $6 a month, usable inside Claude Code and Cursor. The company claims roughly 40% savings over pay-as-you-go, but the plan only works from Singapore and only inside approved tools.
2026-08-14
AI Models | NVIDIA Nemotron 3.5 Lightning
Why Nemotron 3.5 Lightning bets most agent steps don't need a big model
Nvidia's new open model runs locally with 3B active parameters and a 1M-token context, built for agents that stay running. Nvidia claims up to 4x throughput and 30% faster task completion than comparable open models.
2026-08-14
Sovereign AI in Europe
Mistral promises AI sovereignty, and its 1 GW compute bet shows the catch
Mistral's sovereignty push now has concrete products: region-locked endpoints, an SLA-backed tier, and third-party model hosting starting with GLM-5.2. The hard part is compute, which rests on a coalition of commitments and a 1 GW target for 2030.
2026-08-13
Artificial Intelligence
Grok 4.6 ties GPT-5.6 Sol at 61, then the component scores split
xAI's Grok 4.6 ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The breakdown is lopsided: wins on knowledge work and most coding tests, losses on DeepSWE and Terminal-Bench. Available now in Cursor and Grok Build, from $2 per million input tokens.
2026-08-13
Open Source AI
Meta's 30B Muse Glimmer hits Apple Silicon first, NVIDIA support follows
Meta released Muse Glimmer, a 30B open-weights multimodal model built for agent workloads, and Ollama shipped it for Apple Silicon the same day. DFlash makes it 1.5x to 1.8x faster on Mac hardware. NVIDIA and AMD support follow in the coming days.
2026-08-13
Agentic AI's growing cyber-capability problem
OpenAI paused Astra on fears it can hack hardened systems unaided
OpenAI paused internal work on Astra, its in-development model, after evaluations concluded the company cannot rule out 'critical cyber capabilities' under its Preparedness Framework. The full threshold describes a model that finds zero-day exploits in hardened systems without human intervention.
2026-08-13
Anthropic / Claude
Anthropic's London founder house is watch-only now
Applications for Anthropic's Claude Founder House London are closed, and the livestream on Sep 23 is the only door left. The agenda shows a lab courting the UK's top AI founders with its internal playbook, office hours, and a pitch that startups and Anthropic win together.
2026-08-11
Open source AI
Muse Glimmer: Meta's 30B agent fits under 20GB, cloud optional
Meta open-sourced Muse Glimmer, a 30B agentic model that runs offline on a single consumer GPU. Quantization keeps it under 20 GB; a DFlash drafter delivers up to 3.1x faster decoding on an RTX 5090, and the weights are on Hugging Face under Apache 2.0.
2026-08-10
Open Source
Meta's 30B Muse Glimmer lands on Apple Silicon today via Ollama's MLX engine
Meta opened the weights for Muse Glimmer, a 30B dense model, and Ollama ships it the same day on Apple Silicon. Local coding agents gain a native backend, with Muse Spark 1.2 teased for later.
2026-08-10
AI Safety
Claude Fable 5's biology fallbacks drop 85% as Anthropic eases safeguards
Anthropic cut Claude Fable 5's biology-related fallbacks by about 85% after retraining its safety classifier. Ordinary health and education queries now reach the full model more often, but virology, toxicology, and molecular design still route to Opus 5.
2026-08-10
Open Source AI
Alibaba's biggest model coded alone for 16 days. Next week, its weights go public.
Qwen3.8-Max activates just 95 billion of its 2.4 trillion parameters, ranks second in Vision Arena, and built an agent framework in a 16-day autonomous run. Alibaba publishes the weights next week in its first open-weight flagship release.
2026-08-10