DFlash 2: Keep Drafting Parallel
DFlash 2 is the successor to our widely deployed parallel drafter: close to 3× the speed of autoregressive decoding, with the same output. Drafters for Qwen3.8-27B and Meta's Muse Glimmer are out today.
Qwen
Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.
DFlash 2 is the successor to our widely deployed parallel drafter: close to 3× the speed of autoregressive decoding, with the same output. Drafters for Qwen3.8-27B and Meta's Muse Glimmer are out today.
Friday’s big release was Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba’s Qwen research lab. I’ve been looking forward to this one: 27B is an …
Qwen 27B 3.8 was given the same CRUDbench task in non-thinking mode and XHigh thinking mode.
Qwen
An open-weights multimodal MoE model that doubles as an early preview of the Qwen4 architecture, the same role Qwen3-Next played for Qwen3.5. It pairs Gated DeltaNet with Qwen Sparse Attention, widens the residual stream into four gated branches, and adds 51B N-gram embedding parameters that cost almost nothing per token. Natively 262K context, extensible to 1M with YaRN.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.