Qwen3.8 Flash-Next
M3 Max 48GB · oMLX · 65,536 ctx
- reported speed:
- 38.0 tokens/s generation
- quant:
- T5 (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports a custom oMLX fork with a Metal kernel for ternary experts running at 30.7 t/s on a 65,536-token prompt plus 256 output tokens, with prefill at 210-380 tok/s. Setup is a ternary routed expert gate/up (Bonsai T5 packing, ~1.875 bpw), Q3 expert down projections, and a 53 GB n-gram table left on SSD as Q8 and mmapped per token, all low-bit tensors fitted with Unsloth's imatrix. Stock oMLX/mlx-lm will not load it. 64K context is confirmed; 96K trips the prefill guard. Physical peak is 42.3 GiB at 64K and ~41.5 GiB at 8K, with ~35.6 GiB resident and ~2 GiB swapped once at load. The oMLX memory guard is on the 'safe' profile with a 48 GB limit, one model and one request at a time. Against Unsloth UD-Q4_K_XL, which does not fit in 48 GB, the user reports KLD vs Q8_0 of 0.49 vs 0.036, MMLU 83.0% vs 89.7%, GSM8K 90.0% vs 92.0%, and HumanEval 92.7% vs 95.7%. The model sometimes ignores 'answer with just the letter' in Chinese.