Vision & Diffusion
Image and video generation and multimodal models.
8 published articles
Video AI research
How Context-Matched Distillation stops video teachers from seeing the future
Video distillation has long trained causal students against teachers that score whole clips with future knowledge. CMD replaces that scoring with a causal teacher, adds prefix-scored targets, and reports state-of-the-art results among autoregressive methods.
2026-08-14
Product-image AI tools, live at Reve.com
Swap one element, and Reve's Templates redraw the rest
Reve's Templates turn one product shot into a full design: swap the product or text and the rest of the template updates automatically. The feature is live at Reve.com.
2026-08-08
Computer Vision
CABiNet stays within 2 points of YOLO26x at an eighth of the compute
The VDD Semantic Segmentation Model Zoo brings YOLO26 and CABiNet models trained on varied drone footage to Hugging Face. YOLO26x-sem leads at 78.83% mIoU, but CABiNet-Large's 77.76% at 54.8 GFLOPs makes the efficiency case.
2026-08-07
AI Video Generation
Wan3.0-Video charges by the second: a 30-second clip runs to $6
Alibaba's Wan3.0-Video bills per second on DashScope, with a 30-second 1080P clip costing $6 per generation. We break down the pricing tiers, the multi-input workflow, and the questions the listing leaves unanswered.
2026-08-07
AI image generation
Reve 2.1 hits second on Arena with a tenth of the compute
Reve 2.1 claims second place overall on the Arena leaderboards while trained on less than a tenth of the compute of the labs around it. The company also details 4K output, addressable regions, and multilingual text rendering.
2026-08-05
Computer Vision, World Models & AI Research
PhiZero: teaching video AI to think in physics before it renders
PhiZero, a CASIA world model, learns a compact discrete "physical language" from raw video and uses it to reason about how a scene will evolve before rendering frames. The authors argue this reason-then-render design produces more physically coherent video than direct pixel prediction.
2026-07-31
AI Research
VideoCoCo fixes AI video's broken physics by thinking in Blender code
VideoCoCo treats executable Blender code as a chain of thought: a coding agent scripts a scene, a simulator plays it out, and a video engine makes the result photorealistic. The split targets text-to-video's physics problem and posts best average scores on PhyGenBench and VBench-2.0.
2026-07-30
AI Image Generation
The one thing Midjourney still can't do: render a Chinese poster correctly
Zhipu AI's GLM-Image fixes a glaring blind spot in AI image generation: accurate Chinese text rendering, for free, with no limits or watermarks. A practical tool for Chinese-language creators.
2026-07-17