DeepSeek4 min read
Open-source AI research
Million-token context on a tenth of the KV cache: DeepSeek-V4's efficiency bet
DeepSeek's V4 preview pairs a 1.6T-parameter Pro model with a 284B Flash variant, both at one million tokens of context. The paper claims 27% of the inference FLOPs and 10% of the KV cache of V3.2, the numbers that make the context cost-effective to serve.
2026-06-22