Qwen3.8 125B (6B active) Flash-Next
on M1 Max 64GB · ds4 · 4,096 ctx
Oct 3, 2026
Use cases
agenticcodinglong-context
Summary
User reports Qwen3.8-Flash-Next Q2_0 at 44.0 tok/s decode and 328 tok/s prefill at 4K context on a 2021 M1 Max with 64 GB unified memory.
Setup is antirez's ds4 engine (forked for Apple Silicon) with Q2_0 GGUF weights, MTP speculative decoding on, single request, model fully in unified memory. Decode holds 38.8 tok/s at 256K and 35.4 tok/s at 398K; prefill falls from 328 to 292 tok/s over the same range. IQ3_XXS (44 GiB) gives 35.2 tok/s at 4K and 31.7 at 259K.
Sustained five-prompt loop for 5 minutes: 43.2 tok/s first quarter, 43.1 last, 73.1 C, 0.70 J per token. In OpenCode sessions (162 requests, context up to 371K) median decode is 44.7 tok/s under 64K and 38.1 at 128-192K. User compares against a Splash port of Qwen3.8-27B (14 tok/s at 128K+ vs 38) and a DGX Spark running vLLM NVFP4 + MTP (36.8 tok/s prose, 45.8 code). DeepSeek V4 Flash (81 GiB Q2, experts streamed from SSD) runs at 11.8 tok/s at 4K on the same Mac.