Qwen releases Qwen3.8 Max Prime

24-09-2026

Alibaba’s Qwen team released Qwen3.8 Max Prime on 23 September 2026 as a new premium model in the Qwen lineup.

Written by:

Senne Doets

Online Marketeer at DataNorth | Next-Gen AI & Tech Apprentice

qwen releases qwen3.8 max
Sign up for our Newsletter

Published: 24 September 2026

Alibaba’s Qwen team released Qwen3.8 Max Prime on 23 September 2026, sold as a higher-throughput version of Qwen3.8 Max at $4 per million input tokens and $12 per million output tokens. That is exactly twice the price of the standard model. On the only public measurements available, it is not measurably faster.

What is Qwen3.8 Max Prime?

It is the same model on different serving capacity, offered as a separate product. Qwen3.8 Max Prime takes text, images and video as input and returns text. Reasoning is on by default, and you can set how much effort it puts in. Tool calling and structured outputs work as they do on the standard model.

The context window is 1 million tokens with up to 131,072 tokens of output, unchanged from Qwen3.8 Max. Alibaba Cloud International is the only host. There is no announcement post, and Alibaba Cloud’s Model Studio documentation still lists only qwen3.8-max and its 2 September snapshot.

Qwen3.8 Max Prime against Qwen3.8 Max

Take this from the table: the price doubles, and the two speed figures do not move in the direction that would justify it.

What is measuredQwen3.8 Max PrimeQwen3.8 Max
Input, per million tokens$4.00$2.00
Output, per million tokens$12.00$6.00
Cache read, per million tokens$0.50$0.25
Measured speed40 tokens per second37 tokens per second
Measured response time2.03 seconds1.42 seconds
Context window1,000,000 tokens1,000,000 tokens

All of these figures come from OpenRouter, which is the only place Qwen3.8 Max Prime is documented at all. The speed and response-time numbers are OpenRouter’s own measurements of live traffic, taken shortly after launch, and a new endpoint can be slow while capacity settles. Treat them as a first reading rather than a settled verdict. Qwen has published no throughput claim of its own to compare them against.

What Qwen is not saying

Everything, essentially. There is no blog post, no model card, no documentation entry and no benchmark. A model named Prime, priced at double, has to be better at something, and the vendor has not said what. On the evidence available it is the same model with a bigger reservation on a shared machine.

That may well be the honest answer, and for some buyers it is a real product. Guaranteed capacity is worth paying for if your traffic is spiky and you cannot afford to be queued behind everyone else. But that is a service-level promise, and Qwen has not made one. There is no published rate limit, no availability target and no statement of what you get for the extra $2 and $6.

What this means

Safe to ignore unless you are already hitting rate limits on Qwen3.8 Max. This is the clearest case of the pattern several labs have adopted this month: take a working model, stand it up on more capacity, charge a premium and publish nothing. Z.ai did the same thing the same day with GLM 5.3 Prime, and at least Z.ai put a speed claim behind it.

If you do run into capacity problems on Qwen3.8 Max, run a week of your own traffic through both and compare the tail. The median speed is not the reason to buy this; the worst five percent of requests might be. That is the number worth measuring, and it is the one nobody has published. Until Qwen documents what Prime guarantees, paying double for three extra tokens per second is not a decision you can defend to a finance team.

For more information, visit the Qwen3.8 Max Prime model page on OpenRouter.

Add DataNorth AI to your Google favorites