LLMs
7 published articles
Google's workhorse model, bigger in the 3.7 update
Gemini 3.7 Flash halves its price, then makes the case for it
Google's Gemini 3.7 Flash ships at half the intro price of 3.6 Flash while claiming gains on coding, web development, and knowledge-work benchmarks. The release lands three weeks after the previous Flash, and comes with a new price-performance argument for agent builders.
2026-08-16
AI Economics
Alibaba's $18 plan runs Qwen and DeepSeek inside Claude Code
Alibaba Cloud's Token Plan Individual bundles Qwen and third-party models such as DeepSeek into a single credit pool, from $6 a month, usable inside Claude Code and Cursor. The company claims roughly 40% savings over pay-as-you-go, but the plan only works from Singapore and only inside approved tools.
2026-08-14
Qwen / Alibaba
Alibaba splits its flagship AI model in two because one size fits nobody
Alibaba's Qwen team released Qwen3-2507, splitting its model line into dedicated instruct and thinking variants. The update brings substantial gains in reasoning, instruction following, and 256K context support, extensible to 1M tokens.
2026-07-29
AI Models
OpenAI shipped a long context fix. It just won't tell you how much better it got.
OpenAI shipped a long context fix without sharing any benchmarks. That silence might tell developers more than the update itself.
2026-07-28
AI Models
Claude Opus 5 nearly matches Fable 5 across benchmarks, with cybersecurity as a deliberate blind spot
Claude Opus 5 launches with near-Fable-level scores on coding and knowledge benchmarks at half the price. But its deliberately weakened cybersecurity capabilities create a clear trade-off for developers.
2026-07-24
Nature Neuroscience paper
Microsoft's new method turns black-box brain AI into readable theories
GCT translates uninterpretable LLM-based brain models into short phrases like 'food preparation' or 'location names,' then uses an LLM to write stories that causally test those explanations in real subjects. The method promises to bridge predictive AI and human-readable scientific theory.
2026-07-05
effective context, output ceilings, and the hidden tax of long windows
Your AI model says it can read 1 million tokens. It's lying. Here's the real math.
All four frontier LLMs advertise 1M+ token contexts, but effective recall, output limits, and real-world cost differ sharply. DeepSeek V4 Pro leads in output ceiling and cost, Gemini excels under 200K tokens, and Claude Opus wins on caching for interactive code review. This analysis breaks down the numbers from April 2026.
2026-06-30