agentic coding
9 published articles
AI / Agentic Coding Research
Blast Radius buries dead context. Zero of 450 bodies came back
A new arXiv paper, Blast Radius, treats wasted agent context as dead matter and buries it reversibly: 17-26% token savings across seven OpenAI models, 450 archived contexts, zero recalls. The paper's bolder goal is making agentic coding sustainable, one reclaimed token at a time.
2026-08-19
Agentic coding & prompt bloat
Catastrophic remembering: why CLAUDE.md files grow 226% and never shrink
An arXiv study of 1,867 repositories finds agentic coding instruction files like CLAUDE.md triple in size over their lifetime, because deleting a line risks regressions once its rationale is lost. The authors name it catastrophic remembering and show documenting the reasoning cuts excess instructions by 99.3%.
2026-08-15
Artificial Intelligence
Grok 4.6 ties GPT-5.6 Sol at 61, then the component scores split
xAI's Grok 4.6 ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The breakdown is lopsided: wins on knowledge work and most coding tests, losses on DeepSWE and Terminal-Bench. Available now in Cursor and Grok Build, from $2 per million input tokens.
2026-08-13
AI IDEs
Cursor attacks agent costs with an India plan at ₹649 and a model router
Cursor's July update adds a ₹649-a-month India plan and an Auto mode that routes requests to cheaper models, alongside full PR review on iPhone and iPad and agents that work in Slack. The release is really about controlling agentic coding costs.
2026-08-05
Open-source AI
Qwen3.8-Max beat 458 human teams in 24 hours, working alone
Qwen3.8-Max beat 458 of 526 human teams in a 24-hour contest while working alone, and Alibaba will open-source its weights next week. Every number is self-reported so far, which is exactly why the autonomy claims deserve scrutiny.
2026-08-03
Benchmark Analysis
On SWE-bench Verified, Top Models Hit 96%. On Private Enterprise Code, They Barely Clear 23%
SWE-bench Verified makes frontier models look close to solving real-world software engineering, with top scores above 95%. SWE-bench Pro, run on private enterprise repositories, drops those same models to 23% or lower, exposing how much of the Verified score depended on public data exposure.
2026-08-02
Product Launch
Cursor's India plan: model router picks cheapest AI, 649 INR/month
Cursor Start gives Indian developers local pricing with UPI, while Cursor Router lets teams control the cost-intelligence tradeoff per request. Two features that lower the barrier and optimize spending.
2026-08-01
Alibaba Qoder
AI code is twice as likely to have bugs. Qoder fixes that during the session
Qoder Security embeds a dedicated security engineer into every AI coding session, catching vulnerabilities at the moment code is written, not days later at deployment.
2026-07-30
AI Coding Security
The blind spot in AI-generated code that Alibaba fixes mid-sentence
Alibaba's in-session code review catches vulnerabilities before they reach the repo, but the real test is whether developers will let it run.
2026-07-26