Alibaba Previews Qwen4 Architecture with 125B Parameters and Reduced Inference Costs
The Qwen3.8-Flash-Next model activates 6 billion parameters per token while raising licensing questions under the EU AI Act.

Alibaba’s Qwen artificial intelligence research division has released an architectural preview of Qwen3.8-Flash-Next, offering an early look at the technical foundation planned for its next-generation Qwen4 model family. The open-weight release features a total parameter count of 125 billion, though it activates only 6 billion parameters per generated token during processing.
The model design prioritizes lowering operational expenditure over achieving raw performance capability gains. Technical leadership for the Qwen project indicated that architectural decisions are increasingly driven by the ongoing cost of running inference, particularly as long-context agentic workloads become standard across enterprise deployments.
To illustrate compute efficiency, the development team compared the preview to its predecessor, Qwen3.7-Plus. That model contains 397 billion total parameters and activates 17 billion per token, meaning Qwen3.8-Flash-Next operates on roughly one-third of the active computing power required by the previous generation.
The new structural design incorporates four main modifications. Three adjustments follow standard deep learning practices: a sparse attention mechanism that operates on micro-blocks rather than individual tokens, a gated residual system regulating signal flow between layers, and a modified training regimen that eliminates batch-size warmup completely.
The fourth architectural change departs from typical mixture-of-experts designs by adding 51 billion parameters into a separate embedding layer indexed by two- and three-character fragments. According to the development team, this mechanism reduces computational requirements and simplifies memory offloading for hardware accelerators with restricted memory capacity—a design consideration relevant to Chinese research labs operating under semiconductor export controls.
All benchmark metrics associated with the preview originate from Alibaba's internal testing environment. First reported by The Next Web, previous model releases from the group, including Qwen3.8, were launched with self-reported evaluation scores. In the model card published alongside Qwen3.8-Flash-Next, the team explicitly noted that performance on the "Humanity's Last Exam" benchmark was evaluated using OpenAI's GPT-4o model rather than the benchmark's native grading system.
The architecture preview also carries regulatory implications regarding open-source exemptions under European technology regulations. The model weights are hosted on the Hugging Face platform under Alibaba's proprietary "qwen-community" license. The Next Web previously reported on Aug. 7 that Alibaba intends to require licensing fees from its largest commercial users.
This monetization framework intersects directly with rules established under the European Union’s Artificial Intelligence Act. While Article 53(2) waives two specific technical documentation requirements for models distributed under free and open-source licenses, Recital 103 specifies that software components provided against payment or otherwise monetized do not qualify for the exemption. Consequently, enterprise developers utilizing Qwen foundational models—including corporate deployments such as Thomson Reuters—may inherit regulatory documentation liabilities depending on final licensing enforcement.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



