Open Source AI
Nvidia's av-flamingo solves the one thing every other video AI gets wrong: time
Nvidia releases AV-Flamingo, an open audio-visual large language model designed for long, complex video understanding. It uses a three-stage curriculum and a timestamped chain-of-thought framework to handle temporal and cross-modal reasoning that trips up most models.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-28 · Last updated: 2026-07-30 · 1 min read

Video understanding is a solved problem only if by video you mean a five-second clip of a cat on a skateboard. The moment you ask a model to track a conversation across a ten-minute meeting recording or piece together what happened in a surveillance feed spanning several minutes, most current systems stop being useful. As OpenAI's recent long-context update acknowledges, even frontier labs are still working on handling extended content.
Nvidia, in collaboration with researchers from multiple institutions, has released AV-Flamingo, a fully open audio-visual large language model that takes a different bet. Instead of optimizing for the short clips that dominate most benchmarks, the model is trained from
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.