Apache 2.0
9 published articles
Open-weight AI
Xiaohongshu's dots3-note: a 280B open MoE that only activates 16B
Xiaohongshu's dots studio has released dots3-note preview, its first open-weight model: a multimodal MoE with 280B parameters, 16B active, and a 512K context window. The sparse design targets low serving cost on one 8-GPU node, but the card has not published benchmark numbers yet.
2026-08-20
Open source AI
Muse Glimmer: Meta's 30B agent fits under 20GB, cloud optional
Meta open-sourced Muse Glimmer, a 30B agentic model that runs offline on a single consumer GPU. Quantization keeps it under 20 GB; a DFlash drafter delivers up to 3.1x faster decoding on an RTX 5090, and the weights are on Hugging Face under Apache 2.0.
2026-08-10
AI Safety: 3B Classifier, Apache 2.0 Weights
Mistral's Shieldstral puts your moderation policy in the prompt, not the weights
Shieldstral frames moderation as a binary question: an instruction, a yes/no query, and the content to judge. Mistral says the 3B model matches open guardrails up to seven times its size on text safety, with Apache 2.0 weights that run on one 16GB GPU.
2026-08-05
Open Source TTS
Audio8's new TTS model clones voices in 11 languages for free
Audio8 releases Audio8-TTS Preview under Apache 2.0, a 0.6B model supporting 11 languages, zero-shot voice cloning, and a bundled 44.1 kHz codec. The release targets edge speech applications that need speed and privacy.
2026-07-30
Edge AI
Mistral Nano runs on $80 hardware and matches 85% of its bigger sibling. That gap decides everything.
Mistral AI's new Mistral Nano targets edge devices with under 1GB of RAM, achieving 85% of the reasoning performance of its larger Mistral Small model. That 15% gap is either irrelevant or decisive, depending on what you need it to do.
2026-07-25
Open Source AI
480 ms and open source: a streaming model that finally matches Whisper's quality
Voxtral Realtime matches offline transcription quality at sub-second latency and is open source. The 13-language model uses a novel causal audio encoder and is trained end-to-end for streaming rather than adapted from offline systems.
2026-07-17
Multimodal AI
Mistral's 12B model just embarrassed a 90B one. The scaling orthodoxy has a problem.
Pixtral-12B matches or beats models seven times its size on multimodal benchmarks, without sacrificing language performance. Mistral also releases a new open benchmark for practical vision-language evaluation, challenging the idea that bigger is always better.
2026-07-17
LLMs & Models
Mistral's cascade recipe shrinks LLMs without killing reasoning
Mistral's cascade distillation shrinks large models into small ones while preserving reasoning and vision. The 3B variant packs capabilities that used to require ten times the parameters, and it's all Apache 2.0.
2026-07-17
Google DeepMind
Gemma 4 just made every other open-weight model look 10x too big
Google DeepMind's Gemma 4 natively multimodal open-weight family introduces thinking mode, encoder-free architecture, and MoE options. The 2.3B model matches Gemma 3's 27B performance. The 31B model tops open-weight leaderboards.
2026-07-13