llamaperf
Oct 7, 2026
Throughput
57.4 t/s gen · 586.0 t/s pp
Quant
oQ4e (MLX)
System RAM
128 GB

Use cases

codingagentic

Summary

User reports Qwen3.8-Flash-Next at 57.4 tok/s decode and 586 tok/s prefill on an M4 Max 128GB, where the model fits fully in memory. Setup is oMLX with the Jundot oQ4e+MTP MLX checkpoint; the 36GB paging configuration uses oMLX 0.7.0 expert offload with resident fraction 0.28 (143/512 experts per layer), PLE on SSD and MTP off. The same model on an M4 Max Studio 36GB measures 15.55 tok/s steady (n>=1024) and 13.0 tok/s short-run (n=128); exact replay on repeated prompts adds about 40% (14.3 to 20.1 tok/s). Other memory tiers are simulated: 48GB at 192 experts ~15.9 tok/s, 48GB at 128 experts ~11.8 tok/s, M4 Pro 32GB ~9.7 tok/s, M4 24GB ~4.9 tok/s.