inference cost
4 published articles
Open-weight reasoning
Nvidia's Nemotron-4 runs 30% cheaper than GPT-4. The catch is buried in the benchmark.
Nvidia's Nemotron-4 line challenges Llama 3.3 and Mistral Large with competitive MMLU scores and a claimed 30% inference savings. The open license and dual-size strategy position it as a practical alternative for teams running agentic workloads at scale, pending third-party verification.
2026-07-25
The cost engineering behind the model swap
Microsoft's Kimi K3 Copilot trial targets $600M in inference savings
Microsoft is preparing to test Moonshot AI's Kimi K3 inside Copilot in a bid to cut AI inference costs by up to $600 million. The company is shopping for cheaper alternatives to OpenAI and Anthropic models, but the Chinese open-weight model still needs to pass quality and latency checks.
2026-07-21
AI Infrastructure
Nvidia's 55-billion-token trick just rewrote the math on agentic AI costs
Nemotron 3 Ultra, Nvidia's sparse 550B model with 55B active parameters, promises to cut inference costs by up to 30% versus peer open models while keeping frontier reasoning. With a 1M-token context window and native NVFP4 quantization, it targets the cost of multi-step agentic workflows.
2026-06-04
Qwen3.7
The timezone loophole that cuts AI inference costs by 80%
Model Studio automatically slashes Qwen3.7-Max and Qwen3.7-Plus API costs by up to 80% during a nightly window that coincides with working hours across the US and Europe. No signup, no code changes, just cheaper inference for anyone who times their calls right.
2026-05-20