temporal reasoning
3 published articles
Video Understanding
Gemini 3.6 Flash counts state changes but misses blinks
Video language models fail at simple event bookkeeping, a new arXiv study shows. Gemini 3.6 Flash counts persistent state changes up to 12 events but has no reliable region for transient blinks; extra frames inflate accuracy without faithful recovery, with only 0.2% of high-count final counts correct.
2026-08-12
TRACTA Benchmark
Neuro-symbolic reasoning outperforms raw neural models on temporal tasks, benchmark finds
TRACTA benchmark reveals neuro-symbolic AI beats raw neural models on three temporal reasoning tasks, with largest margins on early warning and pattern detection.
2026-08-02
Benchmark
Oxford's GauntletBench puts AI agents through 100 real tasks. They failed 81% of them.
Oxford's GauntletBench puts AI agents through 100 challenging real-world tasks. Frontier systems cap out at 19.1% success, far below human performance. The benchmark targets overlooked capabilities in temporal perception, graphical understanding, and 3D reasoning across five professional applications.
2026-07-01