LLMs & Models
Large language models: GPT, Claude, Gemini, Mistral and open weights.
93 published articles
AI Music Generation: Qwen's answer to prompt drift
Music 3.0 swaps the one-tag prompt for a timeline that keeps AI songs on track
AI tracks tend to drift from the prompt as they unfold: instruments drop out, emotion flattens, the vocal style comes and goes. Music 3.0 swaps the one-tag description for a time-sequential Structured Caption, backed by an 8B/0.6B Hybrid-LM that splits structure from detail.
2026-08-21
Open-weight AI
Xiaohongshu's dots3-note: a 280B open MoE that only activates 16B
Xiaohongshu's dots studio has released dots3-note preview, its first open-weight model: a multimodal MoE with 280B parameters, 16B active, and a 512K context window. The sparse design targets low serving cost on one 8-GPU node, but the card has not published benchmark numbers yet.
2026-08-20
Open Source AI
Boris-2 feeds its 125M model 90B tokens, its 250M just 60B
Boris-2 pre-announces three small models with lopsided token budgets: the 125M gets 90B tokens, the 250M just 60B. Training runs from August 12 to September 10, with no benchmarks shared, only the stated ambition to reach SmolLM2-135M strength.
2026-08-16
AI Video Generation
Alibaba's Wan3.0 sells video by the second: $6 for a 30-second clip
Wan3.0 can turn a PDF or a brand deck into a 30-second video in one generation, with per-second API pricing that tops out at $0.20 for 1080P. We break down the per-clip math and the gaps Alibaba admits in its own testing.
2026-08-16
Open-weight speech for production voice agents
Magpie TTS spends 32ms of your voice agent's latency budget
Magpie TTS reports 32ms time-to-first-audio on an NVIDIA B200 and adds Arabic, Korean and Brazilian Portuguese, bringing its roster to 12 languages. The open-weights pitch: self-hosted speech synthesis no longer loses the latency argument.
2026-08-16
On-device AI agents
LFM2.5-2.6B: the tiny agent that outruns models 4x its size
Liquid AI's LFM2.5-2.6B fits an agentic model into 2.6B parameters and under 2.5 GB of memory, topping every instruction-following benchmark it was tested on. It runs 220 tokens/s on a laptop; coding is the one clear gap.
2026-08-12
China's AI pricing war, receipts pending
Zhipu's viral $0.07 GLM-5.2 price already 10x'd, one reply claims
Zhipu's GLM-5.2 went viral at $0.07 per million tokens, a 95% cut posters called "almost free." One reply says the price rebounded 10x as a stunt for eyeballs. Practitioners add that GLM-5.2 never tested as frontier-grade.
2026-08-10
Research
Beacon: your tool-using AI model is making easy questions harder
A new paper from KlingTeam measures when multimodal models actually need tools and when tools hurt. The proposed Beacon model, trained with necessity-aware rewards, improves both accuracy and tool discipline.
2026-08-09
Video generation
MiniMax scrapped its proven architecture to make H3 do everything
MiniMax says H3 unifies text, image, video and audio generation in one model, prices 2K output below a third of mainstream models, and plans to open the weights within days. The small print is the real story: the company abandoned the architecture that gave it an edge to get there.
2026-08-07
Small-model race
BananaMind 2 Micro: 2.9M parameters, 75B tokens, and an extreme overtraining bet
BananaMind 2 Micro will pack 2.9M parameters and train on 75B tokens, roughly 25,900 tokens per parameter. Training starts August 3, release is estimated August 4 to 6, and a BananaMind 2 Pro public preview arrives the same day. No benchmarks have been shared.
2026-08-06
AI Safety
Shieldstral, the 3B classifier that outguns models seven times its size
A 3B safety classifier matches text models nearly seven times its size and sets a new multimodal moderation state of the art, per a July 2026 arXiv paper. The trick: moderation reframed as binary question answering, trained on roughly 54.1 million samples.
2026-08-05
Artificial Intelligence
LFM2.5-Encoders make the small-model case: 3.7× faster than ModernBERT on CPU
Liquid AI's open-weight LFM2.5-Encoders make the case that production NLP belongs on small models. A 230M encoder beats ModernBERT-base on benchmarks and scans full documents in about 28 seconds on a laptop CPU, roughly 3.7× faster.
2026-08-05