benchmarking
2 published articles
AIFeatured5 min read
AI Research
The AI agent bottleneck isn't exploration. It's knowing what good looks like.
Two new Hugging Face papers tackle the same core problem from opposite directions: how to make AI agents reliably evaluate their own actions. AJ-Bench builds a benchmark for environment-aware judge agents, while HeavySkill argues the best judge lives inside the model's parameters.
2026-07-17
Benchmarks & Tests4 min read
Agent evaluation
Sandbox benchmarks are hiding how agents really fail. HKU just built the fix.
UniClawBench evaluates proactive agents across five fundamental capabilities in 400 bilingual real-world tasks, using live Docker containers and a three-agent closed-loop evaluation. It disentangles base model abilities from framework choices, revealing where agents truly break.
2026-07-13