Qwen4's architecture is available ahead of time, featuring 6 billion parameters out of a total of 125 billion.
Alibaba's Qwen team has introduced Qwen3.8-Flash-Next, an open-weight preview of the architecture intended for Qwen4, featuring 125B parameters but activating only 6B for every token. Its license may not meet the EU AI Act's criteria for the open-source exemption.
The Qwen team has unveiled the architecture planned for Qwen4, with Qwen3.8-Flash-Next containing 125B parameters and engaging only 6B of them for each produced token.
The focus of their claim centers on cost rather than capability. The team expresses concern regarding how architectural decisions impact inference expenses as agentic tasks with very long contexts become the standard workload.
They compare it to their previous model. Qwen3.7-Plus has 397B parameters and activates 17B, meaning this new model operates on about one-third of the active computing power.
Three out of the four modifications are fairly conventional. A new sparse attention method operates on micro-blocks instead of individual tokens, a gated residual mechanism regulates the flow between layers, and the training approach eliminates batch-size warmup entirely.
The fourth change is more unusual. Instead of incorporating experts, the model adds 51B parameters as a separate embedding indexed by two- and three-character fragments, which the team claims reduces computational costs and simplifies offloading to memory-limited accelerators.
That last point deserves a second look. Designing for memory-constrained accelerators is typically necessary when top-tier chips are inaccessible, a situation resulting from export controls affecting Chinese labs.
All the data presented comes from Alibaba itself. TNW noted earlier this month that the company labeled Qwen3.8 as the second-best model globally without providing evidence, and these ratings again originate from the team's own assessments.
Some of the decisions made are quite transparent. The model card indicates that Humanity’s Last Exam score was evaluated by GPT-4o rather than the benchmark's official grader, reflecting openness rather than neutrality.
Europe raises questions regarding the license. The weights are available on Hugging Face under a qwen-community license, and TNW reported on August 7 that Alibaba plans to charge its largest commercial users.
The AI Act considers this differentiation significant. Article 53(2) removes two documentation requirements for models under a free and open-source license, and Recital 103 states that components offered for a fee or monetized in any way should not qualify for the exemption.
European companies are not waiting for clarification. Thomson Reuters has developed its model using Qwen, meaning anyone who engages with it will adopt the eventual license outcome.
Published August 26, 2026 - 7:55 pm UTC
Back to top
Other articles
Qwen4's architecture is available ahead of time, featuring 6 billion parameters out of a total of 125 billion.
Qwen3.8-Flash-Next utilizes 6 billion of its 125 billion parameters for each token. Under the EU AI Act, the aspect that is significant is its licensing rather than its architecture.
