test-time compute
2 published articles
LLMs & Models3 min read
AI Research
Penelope hides its reasoning in a single decoder layer to cut inference costs
Penelope, a latent-reasoning framework for decoder-only Transformers, localizes recurrent computation to a narrow decoder interval, cutting inference latency without a major accuracy hit. The paper describes a curriculum that shifts reasoning from visible tokens to an internal GRU loop.
2026-08-04
AI4 min read
A new study reveals the blind judge
Why your AI model's self-review is structurally blind to visual defects
New research from Pine AI and the University of Washington introduces 'grounding' as the key variable governing a third axis of test-time compute: interaction scaling. The findings show that a deterministic instrument measuring actual layout outperforms VLM-on-screenshot evaluation, fixing 40-74% of defects on visual modalities while the standard metric sees nothing.
2026-07-31