Alibaba Qwen released Qwen3.8-Flash, an open-weights 125B MoE model (6B active) previewing the Qwen4 architecture with 1/9 training cost, 262K native context, and ultra-cheap API pricing.
Key Takeaways
- ✓New core architecture combining GDN gated residuals, sparse attention, N-gram embeddings, and Muon optimizer;
- ✓Coding & SWE benchmarks: 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, and 73.9 on CoWorkBench;
- ✓Day-0 support across QwenCloud, OpenRouter, vLLM, SGLang, and Unsloth (runnable locally with ~75GB RAM via GGUF).