SevenTnewS

The cost engineering behind the model swap

Microsoft's Kimi K3 Copilot trial targets $600M in inference savings

Microsoft is preparing to test Moonshot AI's Kimi K3 inside Copilot in a bid to cut AI inference costs by up to $600 million. The company is shopping for cheaper alternatives to OpenAI and Anthropic models, but the Chinese open-weight model still needs to pass quality and latency checks.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-07-21 · Last updated: 2026-08-03 · 2 min read

Microsoft's Kimi K3 Copilot trial targets $600M in inference savings
Sources : Crypto Briefing…

Microsoft is preparing to run Moonshot AI's Kimi K3 model inside Copilot, a move that could cut its AI inference bill by as much as $600 million, according to The Information. Microsoft has not confirmed the figure, but it gives a rare glimpse of what inference costs look like at Copilot's scale.

Engineers working on Copilot plan to evaluate whether the 2.8 trillion-parameter open-weight model can power features that currently rely on systems from OpenAI and Anthropic. This is not a rubber-stamp review. Kimi K3 has to clear Microsoft's requirements for response quality, reliability, safety, and latency, and on most broad public benchmarks it still trails the proprietary frontier models. Public leaderboards are a weak proxy for real enterprise code, as the gap between SWE-bench Verified and private code shows.

Diagram: Microsoft's cost-cutting evaluation for Copilot
Microsoft's decision path to potentially replace existing AI providers with Moonshot's Kimi K3 model to cut Copilot's inference costs, as reported by The Information.

That estimate also shows how fast inference costs add up. Each time a model receives a prompt and generates a response, it consumes compute, and a service with millions of users and large token volumes racks that up quickly. The pressure is pushing companies to compare models on price as well as capability. Some are already building routers that pick the cheapest model for each task.

Microsoft has previously tested DeepSeek models and earlier versions of Kimi for Copilot, according to a person familiar with its plans cited in The Information's report. Those evaluations suggest the company is serious about mixing and matching models rather than giving every workload to a single provider, the same model-swapping logic behind Sakana AI's orchestration approach.

Kimi K3, released July 16, is Moonshot's flagship model. The startup describes it as natively multimodal, with a one million token context window designed for coding, knowledge work, and complex reasoning. It comes out ahead on SWE Marathon, BrowseComp, and Terminal-Bench 2.1. On most broad evaluations it still trails proprietary frontier models. Its open-weight structure would give Microsoft more control over how the model is hosted and tuned on Azure. Demand spiked fast enough that Kimi.ai paused new K3 subscriptions.

Microsoft has been pushing this model-agnostic approach through Foundry, a platform that lets developers pick models by capability, safety, latency, and cost instead of defaulting to the most powerful option. Its pricing page already lists earlier Moonshot models, including Kimi K2 Thinking, K2.5 Thinking, and K2.6 Thinking. Kimi K3 has not appeared there yet, which matches the report that Azure integration is still in progress. Cost pressure is also pushing model selection down to individual calls, where per-call routing frameworks operate.

If the tests succeed, a Chinese open-weight model would be running inside one of Microsoft's flagship AI products, deepening the company's ties to Moonshot and adding pricing pressure on the US labs that currently power Copilot. That prospect is already splitting opinions in Silicon Valley, as coverage of China's open-weight push shows. At $600 million in potential savings, Microsoft has a clear reason to keep testing.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.