SevenTnewS

inference cost

4 published articles

AI3 min read

Open-weight reasoning

Nvidia's Nemotron-4 runs 30% cheaper than GPT-4. The catch is buried in the benchmark.

Nvidia's Nemotron-4 line challenges Llama 3.3 and Mistral Large with competitive MMLU scores and a claimed 30% inference savings. The open license and dual-size strategy position it as a practical alternative for teams running agentic workloads at scale, pending third-party verification.

2026-07-25

AIFeatured2 min read

The cost engineering behind the model swap

Microsoft's Kimi K3 Copilot trial targets $600M in inference savings

Microsoft is preparing to test Moonshot AI's Kimi K3 inside Copilot in a bid to cut AI inference costs by up to $600 million. The company is shopping for cheaper alternatives to OpenAI and Anthropic models, but the Chinese open-weight model still needs to pass quality and latency checks.

2026-07-21

AIFeatured5 min read

AI Infrastructure

Nvidia's 55-billion-token trick just rewrote the math on agentic AI costs

Nemotron 3 Ultra, Nvidia's sparse 550B model with 55B active parameters, promises to cut inference costs by up to 30% versus peer open models while keeping frontier reasoning. With a 1M-token context window and native NVFP4 quantization, it targets the cost of multi-step agentic workflows.

2026-06-04

AIFeatured2 min read

Qwen3.7

The timezone loophole that cuts AI inference costs by 80%

Model Studio automatically slashes Qwen3.7-Max and Qwen3.7-Plus API costs by up to 80% during a nightly window that coincides with working hours across the US and Europe. No signup, no code changes, just cheaper inference for anyone who times their calls right.

2026-05-20