autonomous agents
5 published articles
Open-source AI
Qwen3.8-Max beat 458 human teams in 24 hours, working alone
Qwen3.8-Max beat 458 of 526 human teams in a 24-hour contest while working alone, and Alibaba will open-source its weights next week. Every number is self-reported so far, which is exactly why the autonomy claims deserve scrutiny.
2026-08-03
Productivity metrics
The one metric AI coding agents can't fake: human hours
Cognition Labs built the first automated system to estimate how many human engineering hours each Devin session saves. The model reaches R² 0.70 but deliberately underestimates. The deeper challenge: converting session traces into defensible ROI numbers remains unsolved.
2026-07-15
Special Report
OpenAI's bet on shared agents is the quietest shift in enterprise AI this year
OpenAI launches workspace agents: persistent, cloud-based AI workers that run across ChatGPT and Slack, handle multi-step workflows, and share context across teams. Free until May 6, 2026, then credit-based. A structural shift from GPTs to organizational AI.
2026-07-11
AI Research
GPT-5.5's big win reveals something missing from every agent benchmark
EvoPolicyGym isolates a critical but understudied capability: an agent's ability to refine an executable policy through repeated feedback-constrained edits. The benchmark reveals GPT-5.5 as the strongest performer across 16 environments, and provides trajectory-level diagnostics that expose how different agents allocate budget and convert feedback into tuned parameters.
2026-07-11
AI Models
Anthropic Launches Claude Fable 5: A Fifth-Generation Model for Extended Autonomous Work
Anthropic introduces Claude Fable 5, a fifth-generation model that can run agents for days, tackle ambitious coding projects, and handle complex enterprise workflows with minimal oversight. Pricing starts at $10 per million input tokens.
2026-07-01