SevenTnewS

AI inference

6 published articles

AI4 min read

AI Infrastructure & Policy

Groq bets its future on the American AI Stack

Groq endorsed the White House's AI Action Plan the day it launched, naming itself the compute layer of the American AI Stack. The release reveals a company betting that national policy and cheap domestic inference will carry it into the global market.

2026-08-08

AI3 min read

Inference Infrastructure

AI inference is becoming a location problem: Groq opens a UK data center

Groq opened a UK data center with Equinix as GroqCloud tops 3.5 million developers and production AI traffic keeps climbing. The move is a bet that inference is now a geography problem, with latency and locality deciding where real-time AI can run.

2026-08-08

AI5 min read

Geospatial AI

Ai2's OlmoEarth platform turns satellite inference into a distributed systems problem

Ai2's OlmoEarth Platform tackles the unique scale of geospatial inference by splitting work into CPU-driven data prep, GPU inference, and CPU postprocessing. A continent-wide wildfire risk map ran with a 155x speedup at a fraction of a cent per square kilometer. The architecture, tolerance for failures, and roadmap show how the team is operationalizing foundation models for environmental organizations.

2026-08-02

Hardware & Electronics2 min read

GPU Architecture

AMD MI455X doubles memory capacity, Hugging Face tests confirm 3x request throughput

Early results from Hugging Face show the AMD Instinct MI455X can handle three times more concurrent requests than the MI300, thanks to 432 GB of HBM4 memory. The Transformers library achieves 99.5% success rate on 24 key model architectures.

2026-07-30

AI5 min read

AI Hardware & Infrastructure

Groq just got Nvidia to license its chip. The $750 million was secondary.

Groq signed a non-exclusive licensing deal with Nvidia for inference tech, signaling a strategic pivot from competing on silicon to powering global AI infrastructure. The agreement, paired with a $750M raise and DOE collaboration, positions Groq as an emerging hyperscaler, not just a chipmaker.

2026-07-16

AIFeatured1 min read

Model Studio

Qwen3.7's off-peak pricing cuts API costs by 80%, and US developers get the best hours

Alibaba's Qwen3.7 off-peak pricing offers up to 80% off API calls during hours that cover the US and EU workday. Here's how it compares to GPT-5.5, Claude, and Gemini, and which workloads benefit most.

2026-05-20