coding benchmarks
2 published articles
Closed-loop coding agents
Alibaba built a coding model that learns from its own users. The numbers are hard to ignore.
Alibaba's new coding model Qwen-Coder-Qoder beats Cursor Composer-1 on the Qoder Bench benchmark. Production metrics show a 3.85% code retention increase, a 61.5% drop in tool errors, and a 14.5% reduction in token use. The model trains on real agent traces through a rewarder-attacker framework against reward hacking.
2026-07-23
agent reliability
AI agents can't tell when a Java migration is actually done
IBM Research introduces ScarfBench, an open benchmark for evaluating AI agents on enterprise Java framework migration. Early tests reveal that frontier agents are systematically overconfident about their own results, that configuration layers dominate effort, and that environment issues like Docker caches regularly derail migrations even when code transformations succeed.
2026-07-06