chain-of-thought
6 published articles
AI Research
Longer chain-of-thought hits a wall. ThinkRetrieve injects the fix mid-reasoning
Sequential test-time scaling often hits diminishing or even negative returns, a new preprint argues. ThinkRetrieve retrieves solved examples mid-reasoning and injects them into the trace, reporting relative gains up to 60% on AIME 2025 across five small reasoning models.
2026-08-18
Autonomous Driving Research
XCoT-VLA: driving AI that reasons in six tokens, not a paragraph
XCoT-VLA replaces descriptive chain-of-thought in vision-language-action driving models with 2 to 6 executable tokens, cutting trajectory error while staying within real-time planning budgets. The tradeoff: reasoning that compresses well stops reading like a human explanation.
2026-08-18
AI research
Over half your AI's reasoning is froth, and nobody noticed until now
LLMs often generate reasoning chains that are correct but padded with unnecessary steps. A new diagnostic benchmark shows current evaluators miss this inefficiency entirely, and half of human-written reasoning steps may be compressible.
2026-07-31
AI Research
VideoCoCo fixes AI video's broken physics by thinking in Blender code
VideoCoCo treats executable Blender code as a chain of thought: a coding agent scripts a scene, a simulator plays it out, and a video engine makes the result photorealistic. The split targets text-to-video's physics problem and posts best average scores on PhyGenBench and VBench-2.0.
2026-07-30
Artificial Intelligence
The hardest lesson for AI reasoning engines: when to shut up
MIT researchers propose OS-Pruner, a plug-in that dynamically stops chain-of-thought reasoning when further computation isn't worth the token cost. Tests show 20-60% length reduction with minimal accuracy sacrifice.
2026-07-29
AI Safety Research
AI models can't stop thinking out loud. That's both good news and a nightmare for safety.
Claude Sonnet 4.5 can control its chain-of-thought only 2.7% of the time, versus 61.9% for final outputs. The gap raises open questions about the robustness of CoT monitoring as a safety mechanism, and nobody knows why it exists.
2026-03-09