Tools & Frameworks3 min read
AI Research
Harness evolution looks good until you run a fair test
Automatic harness evolution is supposed to make LLM agents better, but a new paper argues many reported gains may be from overfitting to the public test set. In experiments, simple test-time scaling methods matched or outperformed evolution, and evolved harnesses showed limited generalization to held-out tasks.
2026-07-27