SevenTnewS

AI Research

MiniMax's VTP learned to understand before it learned to generate

MiniMax's VTP framework rethinks how visual tokenizers are pre-trained, replacing pure reconstruction objectives with a mix that rewards semantic understanding. The result: a tokenizer that actually scales with compute, and a 65.8% FID improvement on downstream generation.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-07-22 · Last updated: 2026-07-30 · 1 min read

MiniMax's VTP learned to understand before it learned to generate
Sources : Towards Scalabl…·VTP GitHub repo…

For years, the standard story for visual tokenizers was simple: make the pixels look as close to the original as possible, and the generator downstream will pick up the slack. That story turns out to be wrong.

A new paper from MiniMax, in collaboration with HUST Vision Lab, documents what they call the "pre-training scaling problem": pouring ever more compute into reconstruction-only training of a visual autoencoder can make downstream generation worse. Their proposed framework, VTP (Visual Tokenizer Pre-training), flips the logic entirely: the tokenizer becomes a semantic learner first, a pixel reconstructor second. The results, still at the research stage, suggest this might be the path forward for scalable generative vision, a finding that resonates with work from other Chinese AI labs that have reached parity with US frontiers.

The paradox at the heart of VAE-based generation

Variational autoenc

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.