agent evaluation
2 published articles
AI AgentsFeatured2 min read
AI Agents
Nobody Can Say What Their AI Coding Agents Actually Did Today
Enterprises spending heavily on AI coding agents can't see what those agents actually did, the benchmarks measuring their improvement may be overfit to public test sets, and the code they write isn't reviewed by default. Three separate problems with the same root cause: adoption outran the tooling to observe it.
2026-07-30
Tools & Frameworks3 min read
AI Research
Harness evolution looks good until you run a fair test
Automatic harness evolution is supposed to make LLM agents better, but a new paper argues many reported gains may be from overfitting to the public test set. In experiments, simple test-time scaling methods matched or outperformed evolution, and evolved harnesses showed limited generalization to held-out tasks.
2026-07-27