Qwen3.7
The timezone loophole that cuts AI inference costs by 80%
Model Studio automatically slashes Qwen3.7-Max and Qwen3.7-Plus API costs by up to 80% during a nightly window that coincides with working hours across the US and Europe. No signup, no code changes, just cheaper inference for anyone who times their calls right.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-05-20 · Last updated: 2026-07-30 · 2 min read

Alibaba Cloud's Model Studio has introduced an automatic pricing discount that can reduce API costs for its Qwen3.7 models, the latest in the line that Alibaba recently split into thinking and instruct variants, according to the Qwen3-2507 announcement. It only applies between 22:00 and 08:00 Beijing time (UTC+8). For developers in North America and Europe, that window covers most of the working day, making the discount effectively the default price for anyone paying attention.
The Deal: Up to 80% Off Frontier AI
No signup, no coupon codes, no code changes required. Any API call to Qwen3.7-Max or Qwen3.7-Plus that lands in the off-peak window receives the reduced rate automatically at the billing layer. The discount is determined by the timestamp when the request is received, not when the response is sent. As long as the call arrives before 08:00 Beijing time, the lower price applies.
The savings are substantial. For Qwen3.7-Max, standard input costs $1.25 per million tokens and output $3.75. Off-peak, those drop to $0.50 and $1.50 respectively, a 60% reduction on both, and up to 80% cheaper than some peak rates depending on the comparison. Qwen3.7-Plus goes from $0.32 input and $1.28 output to $0.128 and $0.512, also a 60% cut.
| Model | Standard Input | Standard Output | Off-Peak Input | Off-Peak Output | Savings |
|---|---|---|---|---|---|
| Qwen3.7-Max | $1.25 | $3.75 | $0.50 | $1.50 | Up to 80% |
| Qwen3.7-Plus | $0.32 | $1.28 | $0.128 | $0.512 | Up to 60% |
The promotion runs from June 23, 2026 to July 30, 2026, per Alibaba Cloud's pricing page. There are no tiers to unlock and no volume commitments. The same latency and SLA apply during off-peak hours as during peak, only the bill changes. Similar cost-reduction strategies have been reported by Microsoft.
A convenient overlap for Western timezones
The off-peak window, defined by Beijing time, maps to convenient hours across North America and Europe. For the US West Coast (PDT, UTC-7), off-peak runs from 07:00 to 17:00, the entire business day. For the East Coast (EDT, UTC-4), it's 10:00 to 20:00. Central Europe (CEST, UTC+2) sees 16:00 to 02:00, and the UK (BST, UTC+1) from 15:00 to 01:00. A full timezone table is available in Alibaba Cloud's announcement.
| Region | Timezone | Off-Peak Local Time |
|---|---|---|
| US West Coast | PDT (UTC-7) | 07:00 |
- Source : The timezone loophole that cuts AI inference costs by 80% — 2026-05-20
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.