SevenTnewS

Small-model race

BananaMind 2 Micro: 2.9M parameters, 75B tokens, and an extreme overtraining bet

BananaMind 2 Micro will pack 2.9M parameters and train on 75B tokens, roughly 25,900 tokens per parameter. Training starts August 3, release is estimated August 4 to 6, and a BananaMind 2 Pro public preview arrives the same day. No benchmarks have been shared.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-08-06 · 4 min read

BananaMind 2 Micro: 2.9M parameters, 75B tokens, and an extreme overtraining bet

BananaMind has announced BananaMind 2 Micro, the smallest model in its BananaMind 2 family, with an unusual caveat: the model does not exist yet. Training has not started, the announcement post says, and nothing has been released. What exists is a specification, and the specification is the story. BananaMind 2 Micro will use 2.9M parameters while being overtrained on 75B tokens "to get the maximum intelligence per parameter," in the team's own words.

The schedule is just as specific. Training starts August 3, and the release date is estimated for August 4 to 6. The same day training begins, BananaMind says it will release a public preview of BananaMind 2 Pro, which the post mentions only by name.

The numbers at a glance:

ModelParametersTraining tokensTokens per parameter
BananaMind 2 Micro2.9M75B~25,900
Gemma 4, smallest config2.3Bn/an/a
Gemma 3, previous flagship27Bn/an/a

25,900 tokens per parameter, and a deliberate imbalance

The two headline numbers deserve to be read together. 75B tokens across 2.9M parameters works out to roughly 25,900 tokens for every parameter in the model. For a model that small, the ratio is deliberately lopsided. The word the team chose, "overtrained," signals that the imbalance is the strategy rather than a side effect: feed a tiny model an enormous volume of data, then see how much capability survives per parameter.

The post offers no benchmark scores, no evaluation results, and no comparisons. The promise of "maximum intelligence per parameter" is, at this point, an intention, backed by nothing public that can be checked. Whether 25,900 tokens per parameter buys real capability is a question only the August training run can answer. Overtraining is not the only lever that moves small-model scores: an external monitoring controller added 4.5 points on MATH-500 to a quantized model.

Muon instead of AdamW, plus a gate from the TX4

Three technical changes come with the announcement. BananaMind is dropping AdamW in favor of the Muon optimizer, which it says offers up to 2x faster convergence and tolerates a higher learning rate. The learning rate moves to 2.2e-2. The model also gains what the team calls the XSA refresh gate from the TX4 architecture, integrated into BananaMind's own design. AdamW has been the default optimizer for much of the field's recent large-scale training, so a public move away from it is the kind of detail that signals a team optimizing every variable at once.

That is the full extent of the disclosed detail. The post does not explain how the 2x convergence figure was measured, what the refresh gate does, or where the TX4 architecture comes from. Every claim rests on the team's own description, with no outside validation attached, and the dates carry the same caveat: the release window is an estimate, not a commitment.

A bet built for the small-model race

The timing gives the announcement its context. Google DeepMind says the smallest model in its Gemma 4 generation, at 2.3B parameters, matches the performance of Gemma 3's 27B model, a 10x parameter-efficiency gain in one generation, and Gemma 4 brought thinking mode to open weights, previously the domain of closed models. Independent work has started mapping what Gemma 4's internal representations actually contain. BananaMind 2 Micro sits roughly 800 times below even that 2.3B scale. The other end of the open-weight spectrum is just as extreme: Kimi K3, the largest open model yet, carries 2.8T parameters. This is not an incremental step down. It is a different weight class, and the announcement reads like a deliberate positioning statement for it.

The gaps are as visible as the ambition. No benchmarks, no demonstration, and nothing said about training hardware or licensing terms. The window is short: if training starts August 3 and release lands anywhere in the estimated August 4 to 6 slot, the team is implying either a 75B-token run that fits in days or a deliberately flexible timeline. The post does not say which. What it does say is that BananaMind is comfortable betting on the extreme end of overtraining. If a 2.9M-parameter model with that much data behind it works, the economics of small models move again; compact models have already won on narrow ground, like the 0.8B OvisOCR2 taking state-of-the-art scores on document parsing. If it does not, the field learns where the limit sits. Either way, August 3 settles part of the argument.

BananaMind 2 Pro appears in the announcement only as a name and a date. Its public preview arrives August 3. No parameter count, no capabilities, no further detail.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.