Alibaba to Release Qwen 3.8-Flash-Next as a Preview of What Qwen 4 Will Offer

Summary

Alibaba is releasing Qwen 3.8-Flash-Next, a multimodal model framed as a preview of the upcoming Qwen 4 architecture rather than a final flagship. It is described as a 125-billion-parameter model with only 6 billion active parameters per token, likely using a mixture-of-experts design so only relevant submodels run for each task. That would give it large overall capacity with much lower compute cost. Alibaba and Hugging Face say it is an early Qwen 4 preview intended to help developers prepare for the full release. No official benchmarks are available yet, so the claimed size and efficiency are unverified in practice. The release fits China’s fast-moving open-weight trend, where powerful downloadable models are increasingly challenging closed APIs and making near-frontier capability cheaper to deploy.