vision-language models
2 published articles
A new study reveals the blind judge
Why your AI model's self-review is structurally blind to visual defects
New research from Pine AI and the University of Washington introduces 'grounding' as the key variable governing a third axis of test-time compute: interaction scaling. The findings show that a deterministic instrument measuring actual layout outperforms VLM-on-screenshot evaluation, fixing 40-74% of defects on visual modalities while the standard metric sees nothing.
2026-07-31
Multimodal AI
The three-stage rhythm that stops AI from seeing things that aren't there
New research reveals a stable three-stage redistribution of multimodal attention in VLMs, operationalized as the Visual Relay Window (VRW). The TRACE framework uses lightweight trained modules to schedule this window per task, improving grounding-sensitive benchmarks by 4.33 points on average and up to 6.6 points.
2026-07-23