AI Agents3 min read
Artificial intelligence · Agent safety
SafeEvolve cuts agent attack success 3x by evolving harness and model together
SafeEvolve, an arXiv paper submitted on 2 September, pairs runtime harness updates with policy training so an agent's own trajectories drive both. It reports a threefold cut in attack success on AgentDojo and a modest utility gain.
2026-09-25