video understanding
3 published articles
Multimodal / Open Source
Qwen plugin pack gives your coding agent eyes, hands, and video memory
Alibaba's Qwen team released Qwen-MM-Plugins, an Apache-2.0 toolkit that gives coding agents native multimodal skills: dynamic-resolution image and video reading, OCR, grounding, ASR, long-video memory, video generation, and thin-client control of Blender and FreeCAD. One script installs it across Claude Code, Codex, Qoder, OpenClaw, Qwen Code, and Gemini CLI.
2026-08-15
Video Understanding
Gemini 3.6 Flash counts state changes but misses blinks
Video language models fail at simple event bookkeeping, a new arXiv study shows. Gemini 3.6 Flash counts persistent state changes up to 12 events but has no reliable region for transient blinks; extra frames inflate accuracy without faithful recovery, with only 0.2% of high-count final counts correct.
2026-08-12
Open Source AI
Nvidia's AV-Flamingo solves the one thing every other video AI gets wrong: time
Nvidia releases AV-Flamingo, an open audio-visual large language model designed for long, complex video understanding. It uses a three-stage curriculum and a timestamped chain-of-thought framework to handle temporal and cross-modal reasoning that trips up most models.
2026-07-28