SevenTnewS

multimodal AI

13 published articles

LLMs & Models5 min read

Open-weight AI

Xiaohongshu's dots3-note: a 280B open MoE that only activates 16B

Xiaohongshu's dots studio has released dots3-note preview, its first open-weight model: a multimodal MoE with 280B parameters, 16B active, and a 512K context window. The sparse design targets low serving cost on one 8-GPU node, but the card has not published benchmark numbers yet.

2026-08-20

LLMs & Models5 min read

Video generation

MiniMax scrapped its proven architecture to make H3 do everything

MiniMax says H3 unifies text, image, video and audio generation in one model, prices 2K output below a third of mainstream models, and plans to open the weights within days. The small print is the real story: the company abandoned the architecture that gave it an edge to get there.

2026-08-07

Qwen / Alibaba4 min read

Qwen / Alibaba Cloud

Qwen2.5-Omni outperforms Gemini-1.5-Pro on OmniBench, fits under 12GB

Alibaba's open-source Qwen2.5-Omni outscored Gemini-1.5-Pro on OmniBench and topped the MMAU audio reasoning leaderboard. Quantized builds cut VRAM below 12GB and MNN support brings real-time voice chat to phones.

2026-08-07

IoT & Sensors4 min read

Model Review

The 7B model that just made GPT-4o-mini look expensive

A deep dive into Alibaba's open-source omni-model: processes text, images, audio, and video simultaneously while generating streaming speech; outperforms GPT-4o-mini and Gemini on multiple benchmarks; fits on a single consumer GPU. The catch? Text-only reasoning takes a hit.

2026-08-03

Labs & Research3 min read

Graph Foundation Models

Zero-shot transfer on graphs has a multimodal blind spot. CHARM offers a fix

CHARM replaces isolated node features with structured graph contexts that capture multimodal semantics and cross-modal relations. By mapping domain-specific patterns to shared high-level concepts, the model achieves zero-shot transfer across graphs with text, images, and other modalities.

2026-08-01

AI4 min read

Video generation

MiniMax H3 undercuts video rivals and plans open weights

MiniMax packs image, video and audio generation into one model, prices output at under a third of mainstream per-second rates, and plans to open the weights within days. How well the outputs hold up outside its own demos is the open question.

2026-07-30

Qwen / Alibaba2 min read

Artificial Intelligence

The 7B model that beats bigger ones, runs on a phone, and changes the AI cost equation

Alibaba Cloud's Qwen2.5-Omni is a 7B model that handles text, images, audio, and video, generating speech and text in real time. Its Thinker-Talker architecture and TMRoPE position embedding enable streaming interaction, and the 3B variant runs on mobile SoCs like Snapdragon 8 Gen 1 at usable speeds.

2026-07-23

AI5 min read

Multimodal AI

Alibaba just shipped a model that hears, sees, and speaks, and it runs on a phone

Alibaba's Qwen2.5-Omni processes text, images, audio, and video end-to-end, generating text and natural speech in real time. Benchmarks show it matches specialized models in vision and audio, and a 4-bit quantized version runs on an RTX 3080 or a Snapdragon phone.

2026-07-20

LLMs & Models2 min read

Multimodal AI

Mistral's 12B model just embarrassed a 90B one. The scaling orthodoxy has a problem.

Pixtral-12B matches or beats models seven times its size on multimodal benchmarks, without sacrificing language performance. Mistral also releases a new open benchmark for practical vision-language evaluation, challenging the idea that bigger is always better.

2026-07-17

NLP & MLFeatured4 min read

Physical AI

Alibaba's Qwen is now the brain inside 150,000 robots, cars, glasses, and drones

Alibaba's Qwen AI family now powers over 150,000 hardware devices, from humanoid robots to children's cameras, as the company pivots from chatbots to Physical AI, integrating multimodal models into robots, cars, glasses, and drones.

2026-07-14

LLMs & ModelsFeatured4 min read

Google DeepMind

Gemma 4 just made every other open-weight model look 10x too big

Google DeepMind's Gemma 4 natively multimodal open-weight family introduces thinking mode, encoder-free architecture, and MoE options. The 2.3B model matches Gemma 3's 27B performance. The 31B model tops open-weight leaderboards.

2026-07-13

LLMs & Models2 min read

Alibaba Cloud

Alibaba Cloud's EMR Serverless Spark now processes images and video in plain SQL, no Python needed

Alibaba Cloud’s EMR Serverless Spark now supports images and video frames directly in SQL, letting data engineers skip Python overhead. A case study on autonomous driving data preprocessing shows automated ETL pipelines powered by Qwen vision models replacing manual annotation.

2026-07-09

← PreviousPage 1 / 2 · 13 articlesNext →