long-horizon tasks
2 published articles
Qwen / AlibabaFeatured5 min read
Open-source AI
Qwen3.8-Max beat 458 human teams in 24 hours, working alone
Qwen3.8-Max beat 458 of 526 human teams in a 24-hour contest while working alone, and Alibaba will open-source its weights next week. Every number is self-reported so far, which is exactly why the autonomy claims deserve scrutiny.
2026-08-03
AI Agents5 min read
Artificial Intelligence
StructAgent lifts AI agent success rates from 27% to 79% without bigger models
An academic paper introduces StructAgent, a state-centered framework that restructures how digital agents track task progress. It achieves state-of-the-art results on OSWorld-Verified with open models, and generalizes to Minecraft. The work identifies that raw interaction history is a bottleneck for long-horizon tasks, and proposes a structured state plus verifier-backed workflow as the fix.
2026-07-22