SevenTnewS

Open Source AI

Boris-2 feeds its 125M model 90B tokens, its 250M just 60B

Boris-2 pre-announces three small models with lopsided token budgets: the 125M gets 90B tokens, the 250M just 60B. Training runs from August 12 to September 10, with no benchmarks shared, only the stated ambition to reach SmolLM2-135M strength.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-08-16 · 3 min read

Boris-2 feeds its 125M model 90B tokens, its 250M just 60B

Boris-2 does not exist yet, and the announcement post is fine with that. The project has laid out plans for three small models, the 75M, the 125M and the 250M, with training windows that run from August 12 to September 10 and token budgets that look like someone stopped sorting them by size.

The number that jumps out is the middle one. The 125M model is slated to train on 90B tokens, half again as much as the 60B planned for the 250M. The post answers the obvious question before it gets asked. "The straight answer is time," it says. Training a bigger model consumes more clock, and the project expects the 250M to pull ahead anyway on raw capacity.

The token math behind the 90B budget for a 125M model

Convert those budgets into tokens per parameter and the allocation looks different. The 75M works through about 347 tokens per parameter at 26B total, the 125M roughly 720, and the 250M just 240. That is deliberate. "It saves time, while still allowing the 250M model to exceed the 125M model," the post explains.

ModelParametersTraining tokensTokens per paramEstimated start
Boris-2-75M75M26B~347Aug 12
Boris-2-125M125M90B720Aug 20
Boris-2-250M250M60B240Sep 10
BananaMind 2 Micro (reference)2.9M75B~25,900Aug 3

Another small-model project, BananaMind, is pushing the ratio far harder in the same window. BananaMind 2 Micro packs 2.9M parameters and a 75B token budget, roughly 25,900 tokens per parameter, in the team's words, "to get the maximum intelligence per parameter." Boris-2's heaviest ratio, 720, does not come close to that intensity, so the project's own pitch rests on the "unique architecture and layering scheme" it says it is attempting. We unpacked that bet in our BananaMind 2 Micro coverage.

A target set by hope rather than benchmarks

What Boris-2 offers as a yardstick is the 135M class of open models. The project says it hopes to land near SmolLM2-135M strength and "in the ballpark of AxiomicLabs/GPT-X2.5-135M or BananaMind/BananaMind-2-Pro-Preview." "Fingers crossed" is the operative phrase. Training has not started, no benchmarks have been shared, and the goal is framed as an aspiration rather than a result.

The dates are the only hard commitments on the table. The 75M is estimated to start training August 12, the 125M August 20, the 250M September 10. None of these are release dates, and the post does not say when the models might ship.

The overtraining wave hits the smallest models

Boris-2 is part of a wave of small open models being fed outsized data diets on purpose. BananaMind placed the same kind of bet on its 2 Micro, with training set to start August 3 and a release estimated for August 4 to 6, alongside a public preview of BananaMind 2 Pro. No benchmarks have been shared there either. The shared wager is that data density can substitute for scale, and neither project has evidence yet to cash it. Other small-model pushes have harder numbers to show, like Liquid AI's CPU-speed encoder results.

What is checkable is the arithmetic, and it shows two different bets. BananaMind is pushing tokens per parameter to an extreme. Boris-2 spends its token budget unevenly, favoring the middle model, while hoping the architecture lets all three outperform their parameter counts. Small models have outrun their size before, like a 2.6B agent that outruns models four times its size.

Variants, and the long wait until August 12

The roadmap does not end at the three base sizes. Boris-2 says it will follow with Pro, Instruct and Pro-Instruct variants. "More info will be coming soon," the post closes, which is the only timeline offered apart from the training dates.

Until the first run starts, Boris-2 is a specification with dates attached. The numbers are public, the reasoning is public, and the outcome is entirely unproven. For small-model watchers, that is the point: a bet laid out in the open before a single training run, in a class that has already produced upsets like a 3B classifier that outguns models seven times its size.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.