Cost engineering
Microsoft's $600 million bet on a Chinese model might reshape AI costs
Microsoft is preparing to test Moonshot AI's Kimi K3 inside Copilot in a bid to cut AI inference costs by up to $600 million. The company is shopping for cheaper alternatives to OpenAI and Anthropic models, but the Chinese open-weight model still needs to pass quality and latency checks.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-21 · 2 min read

Microsoft is preparing to run Moonshot AI's Kimi K3 model inside Copilot, a move that could slash its AI inference bill by as much as $600 million, The Information reports. The figure is not confirmed by Microsoft, but it gives some idea of what inference costs look like at Copilot's scale, see K3's recent benchmark performance.
Engineers working on Copilot plan to evaluate whether the 2.8 trillion-parameter open-weight model can power features that currently rely on systems from OpenAI and Anthropic. The tests are not a done deal: Microsoft's engineers must first determine whether Kimi K3 meets the product's requirements for response quality, reliability, safety, and latency, with Kimi K3's benchmark record showing it trails frontier models on most broad evaluations.

The $600 million estimate, which Microsoft has not publicly confirmed, speaks to the sheer cost of running inference at Copilot's scale. Inference is the computing used each time a model receives a prompt and generates a response. For a service serving millions of users and processing large volumes of tokens, those costs add up fast. This is exactly the kind of economic pressure that has pushed companies to shop around, as token-level cost visibility has become a priority for enterprise AI.
Microsoft has previously tested DeepSeek models and earlier versions of Kimi for Copilot, according to a person familiar with its plans cited by the report. Those evaluations suggest the company is serious about mixing and matching models rather than handing every workload to a single provider, a strategy detailed in Sakana AI's Fugu routing approach.
Kimi K3, released July 16, is Moonshot's flagship model. The startup describes it as natively multimodal with a one million token context window, designed for coding, knowledge work, and complex reasoning. On some benchmarks it leads the pack: SWE Marathon, BrowseComp, Terminal-Bench 2.1. But it trails frontier proprietary models on most broad evaluations. Its open-weight structure would give Microsoft more control over how the model is hosted and optimized on Azure. The sudden demand for K3 has already caused capacity issues, as Kimi.ai paused new K3 subscriptions after a demand spike.
Microsoft has been pushing this model-agnostic approach through Foundry, its platform that encourages developers to pick models based on capability, safety, latency, and cost rather than defaulting to the most powerful option. The platform's pricing page already lists earlier Moonshot models, including Kimi K2 Thinking, K2.5 Thinking, and K2.6 Thinking. Kimi K3 has not appeared there yet, which tracks with the report that Azure integration is still in progress.
If the tests succeed, it would mean a Chinese open-weight model running inside one of Microsoft's flagship AI products, deepening the company's ties to Moonshot and adding more pricing pressure on the US labs currently powering Copilot. At $600 million in potential savings, the pressure is clearly there.
- Source : Crypto Briefing post on X
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.