SevenTnewS

Video generation

MiniMax scrapped its proven architecture to make H3 do everything

MiniMax says H3 unifies text, image, video and audio generation in one model, prices 2K output below a third of mainstream models, and plans to open the weights within days. The small print is the real story: the company abandoned the architecture that gave it an edge to get there.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-08-07 · 5 min read

MiniMax scrapped its proven architecture to make H3 do everything

Launching a model is routine. Announcing in the same post that you abandoned the architecture behind your previous generation is not. MiniMax's H3 launch does both, and the confession is the most revealing part of the announcement: the company says it dropped the Hailuo-02 architecture, which by its own account gave it a significant advantage, because it added complexity to a model built around one first principle: task unification.

According to the launch post, H3 is an omni-modal generation model. It reads a multimodal context made of text, images, video and audio, understands it as one unit, and generates audio-video with native dual-channel sound, up to 15 seconds at 2K resolution. MiniMax's pricing pitch is blunt: at 2K, the per-second cost is less than a third of mainstream models, and 768P output runs at half the rate of mainstream 720P, with 2K set as the default. The company also plans to open the weights within days.

The architecture MiniMax gave up

MiniMax's own history puts the sacrifice in context. Hailuo 01 built the generation system from scratch; Hailuo 02 focused on architecture efficiency, data quality and scale. The H3 team says it set that design aside anyway, because task generalization is an irreversible trend and architecture tricks should yield to the model definition. The replacement, the H3-Omni Transformer, splits understanding and generation into a heterogeneous training setup. Multimodal context tripled the variance of sequence lengths, and the two workloads now differ sharply in compute. MiniMax says the setup, tuned per workload with load balancing across samples, gains nearly 30% in end-to-end training throughput.

One pipeline instead of a stack of experts

The unification runs through pretraining. Text-to-image, text-to-video with joint audio, and text-to-audio all live in one model. Voice, sound effects and music are modeled together, not as separate domains, a different route from the melody-first planning behind Alibaba's Qwen-Music. Generalized reference and editing covers image-to-image, image-to-video, audio-to-audio and audio-video-to-audio-video, with the relationship stated in natural language, not picked from a fixed menu. MiniMax's example prompt is one sentence that folds three references together: the Hitchcock camera movement from a video clip, the person in an image, and a voice from an audio file. The model resolves the mapping itself. Language, in MiniMax's design, is the bridge that generalizes across any task and modality.

Two technical decisions carry the economics. The H3-VAE tokenizer was rebuilt for this generation, and its higher compression delivers a 4x gain in sequence length, cutting training and inference cost and making native 2K output feasible. Instead of a dedicated upscaler, H3 regenerates its own low-resolution results in context, reusing the base model's generation ability and the original multimodal context. MiniMax says this preserves small text and details that conventional upscalers would have to guess.

The captioning pipeline is where the unification gets expensive. Describing the target video, the relationship between context and target, and the relations between elements inside the context consumes about 100K tokens of inference per piece of material, collapsing to an average of about 4K tokens. MiniMax calls this Contextual Omni Representation, and it points to early invite-testing feedback as evidence: strong instruction following, solid rendering of text and brand information, and video-to-video motion transfer, for advertising, e-commerce, product design, UI/UX and games.

The price attack and the open-weights plan

The pricing claim is the most checkable part of the announcement and the least auditable, because the mainstream models it compares against are unnamed. Below a third at 2K, half at 768P, plus an assertion of the industry's best price-performance. Direction is clear; receipts are not.

ClaimDetails from MiniMax
OutputNative dual-channel audio-video, up to 15 seconds at 2K, 2K by default
UnderstandingText, images, video and sound handled as one multimodal context
Pricing2K: under 1/3 of mainstream per-second cost; 768P: half of mainstream 720P
AudioVoice, effects and music modeled together, all output dual-channel
Open weightsPlanned within days, subject to applicable law and regulations
Known limitsContext understanding, model scale, fine detail at high resolution

The open-weights plan is a three-part wager: grow the open-source community, speed up adaptation of domestic Chinese chips, several of which H3 was designed to work with, and let users customize their own builds. It is also an answer to the closed-source dominance that MiniMax says has held video generation back, a case it makes by comparing with the LLM ecosystem's faster iteration. Openness alone is not a proven cure: Kimi K3, the largest open model released to date, still trails the best proprietary systems on overall benchmarks.

What MiniMax admits it cannot do yet

The announcement is candid about gaps. Multimodal context understanding is the foundation of generation, MiniMax says, and it has the most room to improve. Long-form video is where most models fall apart, a problem Nvidia has tackled separately with its open AV-Flamingo model, built for long, complex video understanding. The next H-series version will merge with the M-series model line. Model scale is a constraint; MiniMax names scaling as the explicit direction. Fine detail in some scenes still needs work; higher resolution and finer quality are on the roadmap.

The wager is coherent, which is not the same as safe. Specialized pipelines are predictable: pick a task, pick a tool, tune it. A general model trades that predictability for flexibility, and MiniMax is betting that flexibility is worth more than the architecture it gave up, at prices rivals will have to answer. The bet cuts against the current leaderboard logic, where Macaron-V1-Venti beats GPT-5.5 and Opus 4.8 by splitting into four specialists. Whether one model can truly replace the expert stack is the open question. MiniMax has chosen to answer it in public.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.