multimodal AI
13 published articles
Open-weight AI
Xiaohongshu's dots3-note: a 280B open MoE that only activates 16B
Xiaohongshu's dots studio has released dots3-note preview, its first open-weight model: a multimodal MoE with 280B parameters, 16B active, and a 512K context window. The sparse design targets low serving cost on one 8-GPU node, but the card has not published benchmark numbers yet.
2026-08-20
Video generation
MiniMax scrapped its proven architecture to make H3 do everything
MiniMax says H3 unifies text, image, video and audio generation in one model, prices 2K output below a third of mainstream models, and plans to open the weights within days. The small print is the real story: the company abandoned the architecture that gave it an edge to get there.
2026-08-07
Qwen / Alibaba Cloud
Qwen2.5-Omni outperforms Gemini-1.5-Pro on OmniBench, fits under 12GB
Alibaba's open-source Qwen2.5-Omni outscored Gemini-1.5-Pro on OmniBench and topped the MMAU audio reasoning leaderboard. Quantized builds cut VRAM below 12GB and MNN support brings real-time voice chat to phones.
2026-08-07
Model Review
The 7B model that just made GPT-4o-mini look expensive
A deep dive into Alibaba's open-source omni-model: processes text, images, audio, and video simultaneously while generating streaming speech; outperforms GPT-4o-mini and Gemini on multiple benchmarks; fits on a single consumer GPU. The catch? Text-only reasoning takes a hit.
2026-08-03
Graph Foundation Models
Zero-shot transfer on graphs has a multimodal blind spot. CHARM offers a fix
CHARM replaces isolated node features with structured graph contexts that capture multimodal semantics and cross-modal relations. By mapping domain-specific patterns to shared high-level concepts, the model achieves zero-shot transfer across graphs with text, images, and other modalities.
2026-08-01
Video generation
MiniMax H3 undercuts video rivals and plans open weights
MiniMax packs image, video and audio generation into one model, prices output at under a third of mainstream per-second rates, and plans to open the weights within days. How well the outputs hold up outside its own demos is the open question.
2026-07-30
Artificial Intelligence
The 7B model that beats bigger ones, runs on a phone, and changes the AI cost equation
Alibaba Cloud's Qwen2.5-Omni is a 7B model that handles text, images, audio, and video, generating speech and text in real time. Its Thinker-Talker architecture and TMRoPE position embedding enable streaming interaction, and the 3B variant runs on mobile SoCs like Snapdragon 8 Gen 1 at usable speeds.
2026-07-23
Multimodal AI
Alibaba just shipped a model that hears, sees, and speaks, and it runs on a phone
Alibaba's Qwen2.5-Omni processes text, images, audio, and video end-to-end, generating text and natural speech in real time. Benchmarks show it matches specialized models in vision and audio, and a 4-bit quantized version runs on an RTX 3080 or a Snapdragon phone.
2026-07-20
Multimodal AI
Mistral's 12B model just embarrassed a 90B one. The scaling orthodoxy has a problem.
Pixtral-12B matches or beats models seven times its size on multimodal benchmarks, without sacrificing language performance. Mistral also releases a new open benchmark for practical vision-language evaluation, challenging the idea that bigger is always better.
2026-07-17
Physical AI
Alibaba's Qwen is now the brain inside 150,000 robots, cars, glasses, and drones
Alibaba's Qwen AI family now powers over 150,000 hardware devices, from humanoid robots to children's cameras, as the company pivots from chatbots to Physical AI, integrating multimodal models into robots, cars, glasses, and drones.
2026-07-14
Google DeepMind
Gemma 4 just made every other open-weight model look 10x too big
Google DeepMind's Gemma 4 natively multimodal open-weight family introduces thinking mode, encoder-free architecture, and MoE options. The 2.3B model matches Gemma 3's 27B performance. The 31B model tops open-weight leaderboards.
2026-07-13
Alibaba Cloud
Alibaba Cloud's EMR Serverless Spark now processes images and video in plain SQL, no Python needed
Alibaba Cloud’s EMR Serverless Spark now supports images and video frames directly in SQL, letting data engineers skip Python overhead. A case study on autonomous driving data preprocessing shows automated ETL pipelines powered by Qwen vision models replacing manual annotation.
2026-07-09