AI
Artificial intelligence: LLMs, agents, diffusion, vision, NLP and the latest from top labs.
565 published articles
Sovereign AI in Europe
Mistral promises AI sovereignty, and its 1 GW compute bet shows the catch
Mistral's sovereignty push now has concrete products: region-locked endpoints, an SLA-backed tier, and third-party model hosting starting with GLM-5.2. The hard part is compute, which rests on a coalition of commitments and a 1 GW target for 2030.
2026-08-13
Artificial Intelligence
Grok 4.6 ties GPT-5.6 Sol at 61, then the component scores split
xAI's Grok 4.6 ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The breakdown is lopsided: wins on knowledge work and most coding tests, losses on DeepSWE and Terminal-Bench. Available now in Cursor and Grok Build, from $2 per million input tokens.
2026-08-13
Open Source AI
Meta's 30B Muse Glimmer hits Apple Silicon first, NVIDIA support follows
Meta released Muse Glimmer, a 30B open-weights multimodal model built for agent workloads, and Ollama shipped it for Apple Silicon the same day. DFlash makes it 1.5x to 1.8x faster on Mac hardware. NVIDIA and AMD support follow in the coming days.
2026-08-13
Terminal Security
The '$HOME' trap: AI coding agents need sandboxes, not 'allow?' prompts
Qoder's terminal sandbox blocks close to a hundred destructive agent commands every day. The cases behind those blocks explain why 'allow?' prompts fail: a project folder named $HOME, a cleanup that targeted /root, and a click that nearly cost an entire disk.
2026-08-13
Agentic AI's growing cyber-capability problem
OpenAI paused Astra on fears it can hack hardened systems unaided
OpenAI paused internal work on Astra, its in-development model, after evaluations concluded the company cannot rule out 'critical cyber capabilities' under its Preparedness Framework. The full threshold describes a model that finds zero-day exploits in hardened systems without human intervention.
2026-08-13
Clinical AI: heart-failure phenotyping preprint
nMAS automates heart-failure EHR features, tested only on 500 dummy patients
nMAS, a multi-agent pipeline, generated 132 structured and 70 rubric-scored features from 500 dummy patient records, lifting held-out HFrEF phenotyping AUROC from 0.895 to 0.963. The preprint has not been validated on real patient data.
2026-08-13
Qoder Computer Use
One engineer shipped a macOS agent without knowing Swift
An engineer who could not read Swift shipped production-grade macOS software with Qoder's Computer Use. His approach: judge code by behavior, make the Agent generate its own tests, and keep every lesson in the file system so no round starts from zero.
2026-08-12
Video Understanding
Gemini 3.6 Flash counts state changes but misses blinks
Video language models fail at simple event bookkeeping, a new arXiv study shows. Gemini 3.6 Flash counts persistent state changes up to 12 events but has no reliable region for transient blinks; extra frames inflate accuracy without faithful recovery, with only 0.2% of high-count final counts correct.
2026-08-12
On-device AI agents
LFM2.5-2.6B: the tiny agent that outruns models 4x its size
Liquid AI's LFM2.5-2.6B fits an agentic model into 2.6B parameters and under 2.5 GB of memory, topping every instruction-following benchmark it was tested on. It runs 220 tokens/s on a laptop; coding is the one clear gap.
2026-08-12
Anthropic / Claude
Anthropic's London founder house is watch-only now
Applications for Anthropic's Claude Founder House London are closed, and the livestream on Sep 23 is the only door left. The agenda shows a lab courting the UK's top AI founders with its internal playbook, office hours, and a pitch that startups and Anthropic win together.
2026-08-11
China's AI pricing war, receipts pending
Zhipu's viral $0.07 GLM-5.2 price already 10x'd, one reply claims
Zhipu's GLM-5.2 went viral at $0.07 per million tokens, a 95% cut posters called "almost free." One reply says the price rebounded 10x as a stunt for eyeballs. Practitioners add that GLM-5.2 never tested as frontier-grade.
2026-08-10
Open source AI
Muse Glimmer: Meta's 30B agent fits under 20GB, cloud optional
Meta open-sourced Muse Glimmer, a 30B agentic model that runs offline on a single consumer GPU. Quantization keeps it under 20 GB; a DFlash drafter delivers up to 3.1x faster decoding on an RTX 5090, and the weights are on Hugging Face under Apache 2.0.
2026-08-10