SevenTnewS

LLMs

7 published articles

Google DeepMindFeatured3 min read

Google's workhorse model, bigger in the 3.7 update

Gemini 3.7 Flash halves its price, then makes the case for it

Google's Gemini 3.7 Flash ships at half the intro price of 3.6 Flash while claiming gains on coding, web development, and knowledge-work benchmarks. The release lands three weeks after the previous Flash, and comes with a new price-performance argument for agent builders.

2026-08-16

Qwen / Alibaba4 min read

AI Economics

Alibaba's $18 plan runs Qwen and DeepSeek inside Claude Code

Alibaba Cloud's Token Plan Individual bundles Qwen and third-party models such as DeepSeek into a single credit pool, from $6 a month, usable inside Claude Code and Cursor. The company claims roughly 40% savings over pay-as-you-go, but the plan only works from Singapore and only inside approved tools.

2026-08-14

Qwen / Alibaba2 min read

Qwen / Alibaba

Alibaba splits its flagship AI model in two because one size fits nobody

Alibaba's Qwen team released Qwen3-2507, splitting its model line into dedicated instruct and thinking variants. The update brings substantial gains in reasoning, instruction following, and 256K context support, extensible to 1M tokens.

2026-07-29

AI1 min read

AI Models

OpenAI shipped a long context fix. It just won't tell you how much better it got.

OpenAI shipped a long context fix without sharing any benchmarks. That silence might tell developers more than the update itself.

2026-07-28

LLMs & ModelsFeatured3 min read

AI Models

Claude Opus 5 nearly matches Fable 5 across benchmarks, with cybersecurity as a deliberate blind spot

Claude Opus 5 launches with near-Fable-level scores on coding and knowledge benchmarks at half the price. But its deliberately weakened cybersecurity capabilities create a clear trade-off for developers.

2026-07-24

AI4 min read

Nature Neuroscience paper

Microsoft's new method turns black-box brain AI into readable theories

GCT translates uninterpretable LLM-based brain models into short phrases like 'food preparation' or 'location names,' then uses an LLM to write stories that causally test those explanations in real subjects. The method promises to bridge predictive AI and human-readable scientific theory.

2026-07-05

LLMs & Models3 min read

effective context, output ceilings, and the hidden tax of long windows

Your AI model says it can read 1 million tokens. It's lying. Here's the real math.

All four frontier LLMs advertise 1M+ token contexts, but effective recall, output limits, and real-world cost differ sharply. DeepSeek V4 Pro leads in output ceiling and cost, Gemini excels under 200K tokens, and Claude Opus wins on caching for interactive code review. This analysis breaks down the numbers from April 2026.

2026-06-30