SevenTnewS

agentic coding

9 published articles

Labs & Research5 min read

AI / Agentic Coding Research

Blast Radius buries dead context. Zero of 450 bodies came back

A new arXiv paper, Blast Radius, treats wasted agent context as dead matter and buries it reversibly: 17-26% token savings across seven OpenAI models, 450 archived contexts, zero recalls. The paper's bolder goal is making agentic coding sustainable, one reclaimed token at a time.

2026-08-19

AI Agents4 min read

Agentic coding & prompt bloat

Catastrophic remembering: why CLAUDE.md files grow 226% and never shrink

An arXiv study of 1,867 repositories finds agentic coding instruction files like CLAUDE.md triple in size over their lifetime, because deleting a line risks regressions once its rationale is lost. The authors name it catastrophic remembering and show documenting the reasoning cuts excess instructions by 99.3%.

2026-08-15

xAI / GrokFeatured4 min read

Artificial Intelligence

Grok 4.6 ties GPT-5.6 Sol at 61, then the component scores split

xAI's Grok 4.6 ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The breakdown is lopsided: wins on knowledge work and most coding tests, losses on DeepSWE and Terminal-Bench. Available now in Cursor and Grok Build, from $2 per million input tokens.

2026-08-13

AI IDEs5 min read

AI IDEs

Cursor attacks agent costs with an India plan at ₹649 and a model router

Cursor's July update adds a ₹649-a-month India plan and an Auto mode that routes requests to cheaper models, alongside full PR review on iPhone and iPad and agents that work in Slack. The release is really about controlling agentic coding costs.

2026-08-05

Qwen / AlibabaFeatured5 min read

Open-source AI

Qwen3.8-Max beat 458 human teams in 24 hours, working alone

Qwen3.8-Max beat 458 of 526 human teams in a 24-hour contest while working alone, and Alibaba will open-source its weights next week. Every number is self-reported so far, which is exactly why the autonomy claims deserve scrutiny.

2026-08-03

Benchmarks & TestsFeatured2 min read

Benchmark Analysis

On SWE-bench Verified, Top Models Hit 96%. On Private Enterprise Code, They Barely Clear 23%

SWE-bench Verified makes frontier models look close to solving real-world software engineering, with top scores above 95%. SWE-bench Pro, run on private enterprise repositories, drops those same models to 23% or lower, exposing how much of the Verified score depended on public data exposure.

2026-08-02

AI IDEs2 min read

Product Launch

Cursor's India plan: model router picks cheapest AI, 649 INR/month

Cursor Start gives Indian developers local pricing with UPI, while Cursor Router lets teams control the cost-intelligence tradeoff per request. Two features that lower the barrier and optimize spending.

2026-08-01

AI5 min read

Alibaba Qoder

AI code is twice as likely to have bugs. Qoder fixes that during the session

Qoder Security embeds a dedicated security engineer into every AI coding session, catching vulnerabilities at the moment code is written, not days later at deployment.

2026-07-30

CybersecurityFeatured4 min read

AI Coding Security

The blind spot in AI-generated code that Alibaba fixes mid-sentence

Alibaba's in-session code review catches vulnerabilities before they reach the repo, but the real test is whether developers will let it run.

2026-07-26