AIME 2025
2 published articles
Labs & Research4 min read
AI Research
Longer chain-of-thought hits a wall. ThinkRetrieve injects the fix mid-reasoning
Sequential test-time scaling often hits diminishing or even negative returns, a new preprint argues. ThinkRetrieve retrieves solved examples mid-reasoning and injects them into the trace, reporting relative gains up to 60% on AIME 2025 across five small reasoning models.
2026-08-18
AI4 min read
AI Benchmarks
GPT-5.4 Leads the 2026 Math Benchmark Pack as Frontier Scores Saturate
GPT-5.4 edges out rivals with a sweep of the top competition-math rows, but GPT-5.2 Pro ties every cell and GPT-5.3 Codex offers nearly identical scores at lower cost. The real differentiator: how models perform on next-generation benchmarks like FrontierMath and HMMT Feb 2026.
2026-07-01