efficiency
3 published articles
Multimodal AI
Mistral's 12B model just embarrassed a 90B one. The scaling orthodoxy has a problem.
Pixtral-12B matches or beats models seven times its size on multimodal benchmarks, without sacrificing language performance. Mistral also releases a new open benchmark for practical vision-language evaluation, challenging the idea that bigger is always better.
2026-07-17
Benchmark deep dive
GPT-5.6 just made every dollar in AI count harder
OpenAI's GPT-5.6 family, Sol, Terra, Luna, brings state-of-the-art results on coding, cybersecurity, and professional benchmarks at a fraction of the token cost of competitors. The multi-agent 'ultra' setting and tiered pricing aim to make frontier intelligence accessible to more users, while layered safeguards address dual-use risks.
2026-07-09
Open-source AI research
Million-token context on a tenth of the KV cache: DeepSeek-V4's efficiency bet
DeepSeek's V4 preview pairs a 1.6T-parameter Pro model with a 284B Flash variant, both at one million tokens of context. The paper claims 27% of the inference FLOPs and 10% of the KV cache of V3.2, the numbers that make the context cost-effective to serve.
2026-06-22