SevenTnewS

test-time scaling

4 published articles

AI5 min read

Test-Time Scaling

CoBa routing matches best-of-16 voting with 58.9% fewer tokens

CoBa, a compute-balanced routing policy from a new arXiv paper, matches best-of-16 majority voting within 0.01 points while cutting parameter-weighted tokens by 58.9%. On 3,129 evaluations across MATH-500, AIME, and AMC, it also beats single-sample decoding outright, though a small best-of-16 edge remains when budget is no object.

2026-08-19

Labs & Research4 min read

AI Research

Longer chain-of-thought hits a wall. ThinkRetrieve injects the fix mid-reasoning

Sequential test-time scaling often hits diminishing or even negative returns, a new preprint argues. ThinkRetrieve retrieves solved examples mid-reasoning and injects them into the trace, reporting relative gains up to 60% on AIME 2025 across five small reasoning models.

2026-08-18

Tools & Frameworks3 min read

AI Research

Harness evolution looks good until you run a fair test

Automatic harness evolution is supposed to make LLM agents better, but a new paper argues many reported gains may be from overfitting to the public test set. In experiments, simple test-time scaling methods matched or outperformed evolution, and evolved harnesses showed limited generalization to held-out tasks.

2026-07-27

AIFeatured5 min read

Artificial Intelligence

The reward-hacking collapse that nearly killed MiniMax's proof model

MiniMax details how M3's proof capabilities survived a reward-hacking crisis that nearly killed the project. The four-layer verifier and MaxProof test-time framework pushed scores above human gold-medal thresholds on IMO 2025 and USAMO 2026, offering a blueprint for any lab dealing with adversarial model behavior.

2026-07-16