- throughput:
- 50.0 t/s gen
- quant:
- Q8_XL
User emphasizes energy efficiency and versatility of Strix Halo compared to Nvidia cards. Mentions running Qwen 3.6 35B at 50 tps on Q8_XL quant.
AMD · 128GB unified memory · 4 reports
User emphasizes energy efficiency and versatility of Strix Halo compared to Nvidia cards. Mentions running Qwen 3.6 35B at 50 tps on Q8_XL quant.
AMD Strix Halo 128GB · llama.cpp
MTP enabled with --spec-type draft-mtp --spec-draft-n-max 3. Baseline without MTP: 11.7 tok/s. Also tested Q8_0: 7.4 → 18.1 tok/s (2.44×).
AMD Strix Halo 128GB · llama.cpp · 128,000 ctx
Benchmark compares MTP vs non-MTP for 27B and 35B-A3B models. 27B-MTP shows significant speedup in generation and overall wall time for long-context chat; 35B-MTP shows mixed results with faster generation but slower end-to-end due to prefill overhead.
AMD Strix Halo 128GB · llama.cpp
Strix Halo 128GB, ROCm backend, Q4_K_M quant, chat workload. Also tested RTX 3090 and RTX 5070. Multiple models and quants reported.