Guide

Qwen3.8-Max, explained: five models named some version of "Qwen3 Max" since 2025, and which one you're actually reading about

Type "qwen3 max" into Google today and you'll get results for a model Alibaba discontinued a year ago, alongside results for the one it shipped this month, with no obvious signal which is which. One independent tracker still benchmarks the old one. We built the actual timeline, checked which versions are open-weight and which aren't (the answer is more complicated than a yes/no), and pulled the current price directly from a live API listing.

Bijgewerkt 8 Sept 202610 min read
Quick answer

The current flagship is Qwen3.8-Max, specifically the "-0902" refresh Alibaba shipped on 2 September 2026: 2.4 trillion total parameters, a sparse mixture-of-experts design, a 1-million-token context window, and $2.00 per million input tokens / $6.00 per million output tokens on OpenRouter.

It is closed, API-only. Alibaba did not release Qwen3.8-Max's weights. What it did release under the same "3.8" naming, separately, was a reduced text-only checkpoint (under a restrictive custom license, not Apache) and an unrelated 27B dense model (free, Apache 2.0): three different products, one confusing shared name.

"Qwen3-Max" (no ".8", no dot-version) is a different, older, discontinued model from September 2025, over 1 trillion parameters, that Alibaba's own successor obsoletes. At least one independent benchmark tracker still lists it as a live comparison point, which is exactly how the confusion perpetuates.

The one active-parameter figure you'll see everywhere ("~95B active") is not an Alibaba-confirmed number. It's a widely repeated secondary-source estimate. Alibaba has disclosed the 2.4T total but not the active count.

The timeline, because nobody else has laid it out this way

Alibaba's Qwen team has shipped five distinctly-named things carrying "Qwen3" and "Max" in some combination over the past year. Read left to right and the naming looks incremental. It isn't: these are different models with different capabilities, different licenses, and in two cases, different companies' pricing pages disagreeing about what the same model costs.

Date              Release                        What it is
------------------------------------------------------------------------
Sept 2025         Qwen3-Max-Preview, then         >1T params, MoE, closed,
                   Qwen3-Max (stable, Sept 23-24)   API/Qwen Chat only.
                                                     Now discontinued.
Jan 2026           Qwen3-Max-Thinking              Reasoning-mode variant
                                                     of the same Max line,
                                                     not a separate model
                                                     family. Closed.
Jul 19, 2026       Qwen3.8-Max (preview)           2.4T params announced
                                                     at WAIC Shanghai.
Aug 3, 2026        Qwen3.8-Max (GA)                Full release: 1M
                                                     context, $2/$6 per
                                                     1M tokens. Closed.
Aug 12-13, 2026    Qwen3.8-2.4T-A95B               Same base architecture,
                                                     text-only, reduced.
                                                     Open weights, but
                                                     restrictive custom
                                                     license (not Apache).
Aug 13-14, 2026    Qwen3.8-27B                     Unrelated dense model,
                                                     27B params, genuinely
                                                     open, Apache 2.0,
                                                     runs on one GPU.
Sept 2, 2026       Qwen3.8-Max-0902                Post-training refresh
                                                     of the Aug 3 build.
                                                     Same price/specs,
                                                     stronger coding and
                                                     agent benchmarks.

The model you almost certainly mean when you search "qwen3 max" today is the last row: Qwen3.8-Max-0902, live on OpenRouter as of 3 September 2026. Everything above it in the table is either discontinued, a mode rather than a separate model, or a different product entirely that happens to share a name.

Why the confusion is real, not just careless writing

We checked Artificial Analysis, one of the more widely cited independent benchmark trackers, and its "Qwen3 Max" model page tracks the original, September 2025 model, not Qwen3.8-Max. That page states Alibaba never disclosed a parameter count for that model, marks it deprecated with benchmarking discontinued except at a 10,000-token input length, and lists pricing at $1.20 per million input tokens and $6.00 per million output tokens. OpenRouter's own listing for the same nominal model shows $0.78 input and $3.90 output instead, a real, unresolved conflict between two live sources about the price of the model everyone agrees is now obsolete. As of this writing, Artificial Analysis had not yet published an independent score for Qwen3.8-Max itself, meaning most of the benchmark numbers circulating for the current model trace back to Alibaba's own published tables, not third-party verification.

That gap showed up in public within hours of one aggregator's ranking update. A Hacker News thread discussing Qwen3.8-Max being ranked "best overall model by agentic index" caught the underlying benchmark provider swapping its grading model mid-cycle, watching Qwen's score shift from 55.4 to 58.4 and a rival model (Claude Opus) overtake it again within the same news cycle. Commenters called out the volatility directly: a leaderboard position changing meaningfully "in a few seconds" despite an identical benchmark description underneath it. If you're choosing a model based on a single leaderboard screenshot, that thread is a useful reminder that the screenshot can be stale before you finish reading it.

What Qwen3.8-Max actually is

The current flagship is a sparse mixture-of-experts model with 2.4 trillion total parameters. The commonly cited "~95 billion active parameters" figure appears across nearly every secondary source covering this model, but it is not a number Alibaba has confirmed in its own announcement; treat it as a widely repeated estimate rather than a specification. It's multimodal, accepting text, image, and video input, and Alibaba positions it around coding, full-stack development, data analysis, and general office work, claiming improvements over the prior Qwen3.7-Max generation on those specific tasks.

The context window is up to 1 million tokens (roughly 991,000 usable for input once you account for output budget), with as much as 131,072 tokens available for a single completion. On OpenRouter's live listing, it's priced at $2.00 per million input tokens and $6.00 per million output tokens, with cache pricing at $0.25 per million for a cache read and $2.50 per million for a 5-minute cache write. It's available through Alibaba's own Qwen Chat interface, the DashScope API, OpenRouter, and the Qoder IDE integration. There's a separate independent claim, found in one source only and not corroborated elsewhere, that the model supports a native 262,144-token context extensible to 1,010,000; we can't confirm that specific framing and would treat the simpler "up to 1M" figure as the safer one to repeat.

Open-weight or not? It depends which "Qwen3.8" you mean

This is the single most consequential point of confusion, because the honest answer splits three ways depending on exactly which product you're looking at:

  • Qwen3.8-Max itself, the full multimodal flagship with the 1M context window, the one you'd reach through Qwen Chat, the API, or OpenRouter, is closed. Alibaba has not released its weights, and coverage from its July 2026 preview announcement described it plainly as proprietary, API-access-only, with no downloadable release planned.
  • Qwen3.8-2.4T-A95B, released separately around 12-13 August 2026, shares the same base architecture but is text-only, drops the native 1M context and vision input, and Alibaba did release its weights, but under a custom license, not Apache 2.0 or another OSI-approved license. That license reportedly requires a paid commercial agreement once a deployment crosses revenue or user thresholds (cited figures include $50M/year in MaaS revenue, or 100 million monthly active users / $20M in monthly revenue). A rumor circulated after release that the license also carried territorial download restrictions in the US, EU, UK, and South Korea; the published license text on Hugging Face contains no such territorial clause, so treat that specific claim as debunked rather than repeat it as fact. This is not the same product as Qwen3.8-Max, even though the name overlaps almost completely.
  • Qwen3.8-27B, released around the same week (13-14 August 2026), is a much smaller 27-billion-parameter dense model, genuinely open under Apache 2.0, and small enough to run on a single consumer GPU. It's an entirely separate model line from the Max family; the "3.8" in its name is the only thing it shares with the flagship. Our own Qwen3.8-27B guide covers what it can and can't do if local, self-hosted inference is what you're actually after.

Put plainly: if you searched "is Qwen3.8-Max open source" hoping to run the flagship locally, the answer is no. If a headline told you Qwen released the weights for "Qwen3.8," it was very likely referring to the smaller 2.4T-A95B checkpoint or the unrelated 27B model, not the full Max product people mean when they say "Qwen3.8-Max." Our Qwen3.8-Flash-Next coverage covers a third, separate open-weight release from the same family with its own 1M-context, non-parametric embedding architecture, a genuinely different design from either the Max line or the 27B model despite the overlapping generational name.

How it stacks up, with the caveat that most of this is self-reported

Independent, third-party benchmarking of Qwen3.8-Max specifically is thin right now. Most of the comparison numbers circulating, including Alibaba's own claims of beating GPT-5.6 Sol Max and Claude Fable 5 on agentic computer-use tasks, trace back to Alibaba's own published tables rather than a neutral third party running the same test on every model. One useful, if narrow, data point: a Hacker News discussion cited independent cost-per-task figures putting Qwen at roughly $1.13 per task against Claude Opus at roughly $1.80 for comparable agentic work, a real efficiency gap if it replicates, but drawn from one source's methodology rather than a broad consensus. Where Gemini's benchmark numbers get described by at least one tracker as independently audited, Qwen's are described in the same source as self-reported. That asymmetry matters if you're making a purchasing decision based on a comparison chart: check whether the chart's Qwen numbers came from Alibaba or from the same lab that ran every other model in the row.

Where this fits if you're not deep in model benchmarks

If your actual use case is summarizing or chatting with documents you already have rather than picking a frontier coding model, Qwen3.8-Max isn't really competing in that lane: it has no document-grounding product built around it the way Gemini Notebook or ChatPDF do, and its 1M-token context is a raw capacity number, not a citation-and-retrieval system. If you're evaluating it because you're choosing between hosted frontier models for coding or agentic work, the practical checklist is: confirm you're pricing the -0902 refresh rather than the deprecated original Qwen3-Max (the two have different prices on at least one tracker), confirm whether you need the closed API flagship or would actually be served by the open-weight 27B model instead, and re-verify current pricing directly against Alibaba's own DashScope page or a live OpenRouter listing rather than a static blog post, since this specific model family has changed name and price twice in under twelve months.

Access itself splits along the same lines as the licensing. Reaching Qwen3.8-Max through Alibaba's Qwen Chat app or through the DashScope API requires an Alibaba Cloud account and, for the API, a billing setup against the per-token pricing above. OpenRouter is the simpler route if you already use it for other models: it lists Qwen3.8-Max as a single-provider model (Alibaba Cloud International only, no fallback routing), so an outage or rate limit on Alibaba's side has nowhere else to fail over to, unlike some multi-provider listings on the same platform. If you actually want to run something yourself rather than call an API, remember that Qwen3.8-Max is not an option: the open-weight paths in this family are the restrictively-licensed Qwen3.8-2.4T-A95B checkpoint or the genuinely open Qwen3.8-27B, and only the 27B model realistically fits on hardware most individuals or small teams own.

People also ask

What is Qwen3.8-Max?

Alibaba's current flagship AI model: a 2.4-trillion-parameter sparse mixture-of-experts model with a 1-million-token context window, released 3 August 2026 and refreshed as "Qwen3.8-Max-0902" on 2 September 2026. It's closed and API-only, priced at $2.00 per million input tokens and $6.00 per million output tokens on OpenRouter.

Is Qwen3.8-Max the same as Qwen3-Max?

No. Qwen3-Max (no ".8") is an older, discontinued model from September 2025 with over 1 trillion parameters. Qwen3.8-Max is its 2026 successor with more than double the parameter count and a much larger context window. At least one benchmark tracker still lists the old model separately, which is a common source of confusion.

Is Qwen3.8-Max open source?

The full Qwen3.8-Max flagship is not open-weight; it's closed and available only through Alibaba's API, Qwen Chat, and third-party platforms like OpenRouter. A related but different, reduced text-only checkpoint (Qwen3.8-2.4T-A95B) was released with open weights under a restrictive custom license, and a separate, smaller 27B model (Qwen3.8-27B) was released under a genuinely open Apache 2.0 license. These are three different products despite the shared "3.8" naming.

How much does Qwen3.8-Max cost?

$2.00 per million input tokens and $6.00 per million output tokens, per OpenRouter's live listing for the current "-0902" version, as of September 2026. The older, discontinued Qwen3-Max had different, and inconsistently reported, pricing, so don't reuse a price you find for that model.

What is Qwen3-Max-Thinking?

A reasoning-focused variant of the original Qwen3-Max line, released around January 2026. Evidence suggests it's a mode or checkpoint within the same Max lineage (enabled via an API parameter) rather than a fully separate model architecture, though coverage isn't fully consistent on this point. It remains closed, API and Qwen Chat only.

notebooklm-to-pdf.comAlle handleidingen

Exporteer je NotebookLM in één klik

Gratis Chrome-extensie. PDF, Word en Markdown. Gerenderd op je eigen computer — er wordt niets geüpload.

Verder lezen