SevenTnewS

Apache 2.0

9 published articles

LLMs & Models5 min read

Open-weight AI

Xiaohongshu's dots3-note: a 280B open MoE that only activates 16B

Xiaohongshu's dots studio has released dots3-note preview, its first open-weight model: a multimodal MoE with 280B parameters, 16B active, and a 512K context window. The sparse design targets low serving cost on one 8-GPU node, but the card has not published benchmark numbers yet.

2026-08-20

Meta AIFeatured4 min read

Open source AI

Muse Glimmer: Meta's 30B agent fits under 20GB, cloud optional

Meta open-sourced Muse Glimmer, a 30B agentic model that runs offline on a single consumer GPU. Quantization keeps it under 20 GB; a DFlash drafter delivers up to 3.1x faster decoding on an RTX 5090, and the weights are on Hugging Face under Apache 2.0.

2026-08-10

Mistral AIFeatured4 min read

AI Safety: 3B Classifier, Apache 2.0 Weights

Mistral's Shieldstral puts your moderation policy in the prompt, not the weights

Shieldstral frames moderation as a binary question: an instruction, a yes/no query, and the content to judge. Mistral says the 3B model matches open guardrails up to seven times its size on text safety, with Apache 2.0 weights that run on one 16GB GPU.

2026-08-05

AIFeatured1 min read

Open Source TTS

Audio8's new TTS model clones voices in 11 languages for free

Audio8 releases Audio8-TTS Preview under Apache 2.0, a 0.6B model supporting 11 languages, zero-shot voice cloning, and a bundled 44.1 kHz codec. The release targets edge speech applications that need speed and privacy.

2026-07-30

AI4 min read

Edge AI

Mistral Nano runs on $80 hardware and matches 85% of its bigger sibling. That gap decides everything.

Mistral AI's new Mistral Nano targets edge devices with under 1GB of RAM, achieving 85% of the reasoning performance of its larger Mistral Small model. That 15% gap is either irrelevant or decisive, depending on what you need it to do.

2026-07-25

Open SourceFeatured3 min read

Open Source AI

480 ms and open source: a streaming model that finally matches Whisper's quality

Voxtral Realtime matches offline transcription quality at sub-second latency and is open source. The 13-language model uses a novel causal audio encoder and is trained end-to-end for streaming rather than adapted from offline systems.

2026-07-17

LLMs & Models2 min read

Multimodal AI

Mistral's 12B model just embarrassed a 90B one. The scaling orthodoxy has a problem.

Pixtral-12B matches or beats models seven times its size on multimodal benchmarks, without sacrificing language performance. Mistral also releases a new open benchmark for practical vision-language evaluation, challenging the idea that bigger is always better.

2026-07-17

LLMs & ModelsFeatured5 min read

LLMs & Models

Mistral's cascade recipe shrinks LLMs without killing reasoning

Mistral's cascade distillation shrinks large models into small ones while preserving reasoning and vision. The 3B variant packs capabilities that used to require ten times the parameters, and it's all Apache 2.0.

2026-07-17

LLMs & ModelsFeatured4 min read

Google DeepMind

Gemma 4 just made every other open-weight model look 10x too big

Google DeepMind's Gemma 4 natively multimodal open-weight family introduces thinking mode, encoder-free architecture, and MoE options. The 2.3B model matches Gemma 3's 27B performance. The 31B model tops open-weight leaderboards.

2026-07-13