llamaperf
Sep 21, 2026
Throughput
7.1 t/s gen
Quant
Q2 (GGUF)
System RAM
128 GB
VRAM reported
124 GB

Use cases

codingmultilinguallong-context

Summary

User reports DeepSeek V4.1 Flash Q2 (~340 GiB MoE) running on a single AMD Strix Halo 128GB box with 124 GiB unified memory, reaching a highest completion-token rate of 7.12 tok/s including reasoning. Setup is the kyuz0 ds4 ROCm runtime with SSD expert streaming (--ssd-streaming-cache-experts), ROCm 7.2.4, 8k context, and a host-side patch. Small coding and Czech tests completed at request-wall medians of 29.2 s (Python), 26.4 s (bugfix), and 20.4 s (Czech); 64k prompts pass but 128k never produced a final answer in a 905 s bounded calibration. User stopped the experiment, calling it not a practical winner. A Three.js coding test produced 20,575 output tokens over ~83 minutes at ~3.9-4.2 end-to-end tok/s with zero accepted games. The same box measured Qwen3.8 Flash-Next at 44.66 tok/s short, 41.37 @ ~32k, 40.08 @ ~64k, and 38.03 @ ~126k with MTP decode, and GLM-5.3-Flash at 14.63 tok/s on ROCm versus 8.57 tok/s on Vulkan.