small language models
4 published articles
Open Source
Sub-200M models are booming while the frontier spends billions
Hugging Face posted a thank-you to the tinkerers behind a finetuning and pretraining explosion of sub-200M parameter models. No benchmarks, no launches, just a signal that open source AI's center of gravity is shifting.
2026-08-18
Open Source AI
Boris-2 feeds its 125M model 90B tokens, its 250M just 60B
Boris-2 pre-announces three small models with lopsided token budgets: the 125M gets 90B tokens, the 250M just 60B. Training runs from August 12 to September 10, with no benchmarks shared, only the stated ambition to reach SmolLM2-135M strength.
2026-08-16
Small-model race
BananaMind 2 Micro: 2.9M parameters, 75B tokens, and an extreme overtraining bet
BananaMind 2 Micro will pack 2.9M parameters and train on 75B tokens, roughly 25,900 tokens per parameter. Training starts August 3, release is estimated August 4 to 6, and a BananaMind 2 Pro public preview arrives the same day. No benchmarks have been shared.
2026-08-06
LLMs & Models
Mistral's cascade recipe shrinks LLMs without killing reasoning
Mistral's cascade distillation shrinks large models into small ones while preserving reasoning and vision. The 3B variant packs capabilities that used to require ten times the parameters, and it's all Apache 2.0.
2026-07-17