LLMs & Models4 min read
AI Research
Four minds, one answer: why AI that thinks differently beat the biggest models at humanity's hardest test
PoTRE breaks inference into four agents working in parallel, then reconciles their answers dynamically. It hit 49.92% on Humanity's Last Exam, beating the previous official best. The approach works with fewer tokens than scaled-up baselines.
2026-07-24