Benchmarks & Tests3 min read
Benchmarks
AI desktop agents fail before-after test 35% of the time
DDB tests ordering and before-after pair tasks across 2,013 instances. The top model hit 65.1% exact match on non-decoy sequences and 65.7% with decoys, exposing a gap in how agents verify state changes.
2026-08-01