token efficiency
4 published articles
Test-Time Scaling
CoBa routing matches best-of-16 voting with 58.9% fewer tokens
CoBa, a compute-balanced routing policy from a new arXiv paper, matches best-of-16 majority voting within 0.01 points while cutting parameter-weighted tokens by 58.9%. On 3,129 evaluations across MATH-500, AIME, and AMC, it also beats single-sample decoding outright, though a small best-of-16 edge remains when budget is no object.
2026-08-19
AI Coding
OpenCode bets token efficiency wins, adding Ling 3.0 Flash for free
OpenCode now offers inclusionAI's Ling 3.0 Flash for free, days after release. The move extends a cost-conscious strategy betting that token efficiency wins the AI coding race.
2026-08-07
AI Coding
OpenCode adds Ling 3.0 Flash for free, betting token efficiency wins the AI coding war
OpenCode offers Ling 3.0 Flash for free, a token-efficient model from inclusionAI that could make AI coding cheaper. The addition comes days after the model's release and continues OpenCode's cost-conscious strategy.
2026-07-24
Gemini triple model launch lands July 21
Google ships Gemini models: cheaper Flash, faster Lite, and a vetted cyber variant
Google's three-model Gemini drop targets production AI with better token efficiency and lower latency. 3.6 Flash cuts token usage by 17% versus its predecessor. 3.5 Flash-Lite runs at 350 output tokens per second. A cyber-focused variant ships exclusively to vetted partners.
2026-07-22