LLMs & Models4 min read
Bottleneck
AI agents hit 49% on a test human experts pass at 95%. More compute won't fix it.
Alibaba's HSCodeComp benchmark reveals the gap: top agents hit 49.4% accuracy vs. human experts' 95% on tariff classification. The bottleneck is structural, more compute doesn't help.
2026-07-19