Retrieval-Augmented Generation
2 published articles
LLMs & Models5 min read
On-premise factory AI, arXiv preprint
Compression cost a factory assistant 13.7% of its answer quality
A new preprint argues that once a model is compressed and adapted with retrieval, parameter count stops predicting the quality of answers a factory-floor assistant gives. Its one case study loses 13.7 percent of judged quality to extraction and recovers two thirds of that.
2026-09-25
Labs & Research4 min read
AI Research
Longer chain-of-thought hits a wall. ThinkRetrieve injects the fix mid-reasoning
Sequential test-time scaling often hits diminishing or even negative returns, a new preprint argues. ThinkRetrieve retrieves solved examples mid-reasoning and injects them into the trace, reporting relative gains up to 60% on AIME 2025 across five small reasoning models.
2026-08-18