inference cost
3 published articles
Open-weight reasoning
Nvidia's Nemotron-4 runs 30% cheaper than GPT-4. The catch is buried in the benchmark.
Nvidia's Nemotron-4 line challenges Llama 3.3 and Mistral Large with competitive MMLU scores and a claimed 30% inference savings. The open license and dual-size strategy position it as a practical alternative for teams running agentic workloads at scale, pending third-party verification.
2026-07-25
Cost engineering
Microsoft's $600 million bet on a Chinese model might reshape AI costs
Microsoft is preparing to test Moonshot AI's Kimi K3 inside Copilot in a bid to cut AI inference costs by up to $600 million. The company is shopping for cheaper alternatives to OpenAI and Anthropic models, but the Chinese open-weight model still needs to pass quality and latency checks.
2026-07-21
AI Infrastructure
Nvidia's 55-billion-token trick just rewrote the math on agentic AI costs
Nemotron 3 Ultra, Nvidia's sparse 550B model with 55B active parameters, promises to cut inference costs by up to 30% versus peer open models while keeping frontier reasoning. With a 1M-token context window and native NVFP4 quantization, it targets the cost of multi-step agentic workflows.
2026-06-04