SevenTnewS

A 158 ms model just rewrote the speed-quality playbook

The 158 ms model that broke the speed-quality rule

Celeris-1 scores 75.9% on MMLU-Pro at 158 ms latency, eight times faster than GPT-5 mini while losing only 2.6 accuracy points. The analysis examines the benchmark methodology and what this speed-quality trade-off means for real-time AI applications.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-07-25 · 1 min read

The 158 ms model that broke the speed-quality rule
Sources : How we test Cel…

The speed-quality frontier

Conventional wisdom says you cannot have both speed and accuracy. Top models on reasoning leaderboards take seconds. Fast models lose too much intelligence. Celeris-1, published by an unaffiliated team, aims to break that trade-off. Its MMLU-Pro numbers: 75.9% accuracy at a median response time of 158 ms. Not the top score, but at a fraction of the latency of the leaders. Latency-focused design has already yielded results elsewhere: Mistral Nano achieved 85% of its larger sibling's reasoning on devices with under 1GB of RAM.

Context: GPT-5 scores 81.9% on the same test but takes 2.0 seconds. Gemini 3.5 Flash Lite leads at 83.0% but needs 1.2 seconds. Celeris-1 trades about 5 to 7 accuracy points for an 8 to 16x speed advantage. For applications where every millisecond counts, that trade-off is the right one. The model is designed for the hot path: every turn of a conversation, every step in an agent loop, every keystroke.

The benchmark results

The Celeris team ran the full MMLU-Pro test set in a five-shot chain-of-thought configuration with strict scoring, setting the reasoning budget to zero for every model that allowed it. Mercury 2 ran in its instant mode. The table below shows accuracy and response time for all six models.

ModelAccuracyResponse time (p50)Timing basis
Celeris-175.9%158 msserver-reported
Gemini 3.5 Flash Lite83.0%1,232 mse2e, colocated
GPT-581.9%2,046 msserver-reported
GPT-5 mini78.5%2,495 msserver-reported
Gemini 2.5 Flash73.0%2,600 mse2e, colocated
Inception / Mercury 263.7%257 msserver-reported

Celeris-1 lands at 75.9%, behind GPT-5 mini by 2.6 points but 16 times faster. It beats Mercury 2, the only other model in the same latency class, by over 12 points while responding 100 milliseconds sooner. Gemini 3.5 Flash Lite remains the most

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.