NLP & MLFeatured4 min read
AI Research
GPT-5.5's big win reveals something missing from every agent benchmark
EvoPolicyGym isolates a critical but understudied capability: an agent's ability to refine an executable policy through repeated feedback-constrained edits. The benchmark reveals GPT-5.5 as the strongest performer across 16 environments, and provides trajectory-level diagnostics that expose how different agents allocate budget and convert feedback into tuned parameters.
2026-07-11