Qwen3.8-Flash-Next previews the Qwen4 architecture with sparse attention and offloadable N-gram embeddings
On August 27, 2026, Alibaba published the open weights of Qwen3.8-Flash-Next, a 125-billion-parameter model that activates only 6 billion per token. It is the public preview of the Qwen4 architecture, built on sparse attention and an N-gram embedding table that can live in RAM or on SSD.