agentic systems
3 published articles
Artificial intelligence
What breaks when AI agents leave the lab: a production reality check
A new tutorial on agentic systems highlights the gap between benchmark success and production reliability. Researchers share concrete patterns for handling failures, from verification pipelines to human-in-the-loop design, drawing on case studies in drug discovery and finance.
2026-07-24
Systems Design / LLM Infrastructure
The real cost of model routing has nothing to do with sticker price
Across 417 agentic tasks, GPT-4.1 cost nearly double Claude Sonnet despite lower sticker pricing, because caching rewrote the math. A deep-dive on why model routing fails as classification, and how optimization-based routers beat difficulty-based ones.
2026-07-15
Research analysis
AI document corruption in delegated workflows: what a new stress test reveals
The DELEGATE-52 benchmark evaluates AI systems on long-horizon delegated document editing tasks, finding that frontier models accumulate semantic fidelity loss of 19–34% over 20 iterations. Python workflows showed less than 1% degradation on average, but the study underscores that reliable long-horizon delegation remains an open challenge.
2026-07-04