llamaperf

AMD Strix Halo 128GB

AMD · 128GB unified memory · 35 reports

See what fits on this GPU →

Use the calculator to check which models fit in 128 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
Tone: positive
reported speed:
52.0 tokens/s generation · 1300.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next at 52 t/s decode and 1300 t/s prefill on Strix Halo. Setup is forked llama.cpp builds and halogen-flash-server, described as roughly double decode and 5-6x prefill performance over prior Strix Halo results. The post is a community reflection on local LLM optimization rather than a detailed benchmark report.

Sep 13, 2026
Tone: positive
reported speed:
1200.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash Next reaching 1.2k t/s prefill on Strix Halo 128GB. Setup is a custom llama.cpp branch with a custom HIP runtime, built to match the closed-source Halogen server's prefill numbers; the community fork had only reached 400 t/s. The user plans to clean up the work and submit PRs to mainline and the community fork, noting the approach may also help the GLM 5.3 Flash architecture.

Sep 13, 2026
Tone: mixed
reported speed:
33.2 tokens/s generation · 1192.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User benchmarks five Strix Halo llama.cpp forks with Qwen3.8 Flash-Next on a Flow Z13 (Ryzen AI Max+ 395, 128 GB, gfx1151). halogen 0.5.4 at 70W reaches prefill 1192/1350/1315 t/s at 8k/32k/131k and decode 33.2 t/s serial and 43.1 t/s drafted. Nathan's toolbox v0.7.4.1 on Vulkan reaches 26 t/s serial and 34.3 t/s drafted, with prefill 539/433 t/s at 8k/32k. myhacsint reaches 25.7 t/s serial and 35.6 t/s drafted, with prefill 532/584/538/427 t/s. strix-llama was not run; the thread reports 14 t/s serial and prefill 532/397 t/s. Speculative decoding fails the byte-identity gate on every GGUF build tested, with 3/5 to 9/10 prompts differing at temp 0. Serial reruns are byte-identical 10/10. The halogen 0.5.3 changelog claims 1246/1424/1358 t/s prefill at 8k/32k/131k. The model is Qwen3.8 Flash-Next; parameter count is not stated.

Sep 11, 2026
Tone: mixed
quant:
UD-Q5_K_XL (GGUF)
kv:
Q8
mtp (multi-token prediction):
on

User reports Qwen3.8-Flash-Next-UD-Q5_K_XL on Strix Halo 128GB with mmproj, MTP, a 64000 Q8 KV cache and 2GB cache-ram, leaving only 1-2GB free memory. Q4_K_XL (77GB) and Q5_K_XL (108GB) were also tested. MTP adds 5.5GB even when the file is 2.8GB. F16 KV cache uses 1GB per 10000 tokens; Q8 uses 1GB per 20000 tokens. Perplexity results: Q5_K_XL Top-1% 94.795, Mean KLD 0.018175, 99.9% KLD 0.591240; Q4_K_XL+50GB-NGRAM Top-1% 93.786, Mean KLD 0.027853, 99.9% KLD 0.948222; Q4_K_XL+25GB-NGRAM Top-1% 93.285, Mean KLD 0.033378, 99.9% KLD 1.090875. The F16 base-kld could not be run due to insufficient memory.

Sep 11, 2026
Tone: positive
reported speed:
12.5 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)
mtp (multi-token prediction):
on
rating:
5/5

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User compares Qwen 3.8 27B at Q6_K and about 31 GiB resident against Qwen 3.8 Flash Next at UD-Q4_K_XL and about 86 GiB on an ASUS ROG Flow Z13 with Ryzen AI Max+ 395, 128 GB unified memory, running Arch Linux. Decode is 10-15 tok/s. Setup is the llama.cpp backend with the LlamaStash tool and the Pi coding harness. Cold prefill runs 31k tokens in 3 min (about 172 t/s) and a 128k window in about 18 min (about 118 t/s). Warm follow-up turns take about 45s. MTP gives 7.3 to 22.4 tok/s on an empty window and 1.15x at the full 256k. Flash Next completes the same 5/5 coding tasks with 45% fewer tokens, 76.5s against 289.8s for 27B. On the Artificial Analysis index Flash Next scores 40 against Opus 4.8 at 42, and 27B xhigh scores 34 against Opus 4.6 at 32. Thinking is 90-95% of generated tokens. The setup costs $0/month and runs fully offline.

Sep 11, 2026
Tone: positive
reported speed:
21.5 tokens/s generation · 199.0 tokens/s prompt processing
quant:
Q5_K_M (GGUF)
kv:
f16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionlong-context

User benchmarks Qwen3.8-Flash-Next-Uncensored on a Bosgame M5 (128GB/2TB) running Fedora 44 with Vulkan driver 26.1.7, reaching 198.99 t/s prompt processing and 21.53 t/s generation at 200,000 context. Setup is a llama.cpp fork (strix-halo-qwen4exp-b10685) with the Q5_K_M quant and a shared-Q8_0 MTP draft model. llama-benchy results across context lengths: pp2048 460.38 t/s, tg512 34.86 t/s; pp4096 465.96 t/s, tg512 33.04 t/s; pp8192 437.14 t/s, tg512 31.18 t/s; pp16384 440.89 t/s, tg512 31.21 t/s; pp32768 407.48 t/s, tg512 29.65 t/s; pp65536 351.55 t/s, tg512 27.40 t/s; pp131072 265.54 t/s, tg512 23.87 t/s; pp200000 198.99 t/s, tg512 21.53 t/s. MTP acceptance rate stays at 80% even at context above 200K.

Sep 10, 2026

Qwen3-Coder-Next

AMD Strix Halo 128GB · llama.cpp · 262,144 ctx

Tone: mixed
reported speed:
36.8 tokens/s generation · 545.8 tokens/s prompt processing
quant:
UD-Q6_K_XL (GGUF)
kv:
f16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User reports Qwen3-Coder-Next UD-Q6_K_XL at a peak of 37.94 t/s for tg32. Setup is llama-benchy 0.4.1 in API latency mode at -c 262144. A 27B 8-bit XL model was also mentioned but was too slow to be workable.

Sep 10, 2026
Tone: positive
reported speed:
76.9 tokens/s generation
quant:
ROCmFP4 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Nex-N2.5-mini in ROCmFP4 GGUF on AMD Strix Halo against Qwen3.8-27B, decoding at 76.9 t/s versus 14-34 t/s. Setup is the HaloFPX engine with weights from julianmb/Nex-N2.5-mini-ROCmFP4-GGUF. Nex-N2.5-mini scores 73.4 on Terminal-Bench 2.1 against 73.0, 63.4 on WebArena against 64.8, and 43.8 on SWE-Bench against 61.7.

Sep 10, 2026
Tone: positive
reported speed:
40.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running 3 local sessions of GLM 4.7 at about 40 tok/s each on an AMD Strix Halo mini PC with 128 GB unified RAM. The model family is inferred as GLM-4.7 from the post title and text. The user is positive about the setup for budget local AI.

Sep 9, 2026
Tone: positive
reported speed:
10.0 tokens/s generation
quant:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen 3.8 Uncensored Q8 at 256k context on the AI PC (Strix Halo) at ~10 t/s. Setup is a custom AI framework running across multiple devices. A Qwen3.6 35B MOE runs on an RTX 5090 at ~200 t/s, and a Gemma 4 model runs on a MacBook Air for security.

Sep 9, 2026
reported speed:
32.0 tokens/s generation · 245.0 tokens/s prompt processing
quant:
ROCmFP2

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash on Strix Halo at 32.0 tok/s decode with DSpark speculative draft and 25.31 tok/s autoregressive. Setup is ROCmFP2 mixed precision at roughly 2.88 bits per parameter, with a Q4RMFP4 draft model. Prefill runs at 245 tok/s sparse. The user compares this to LocalMaxxing entries of 18.99 tok/s for HipFire and 15.6 tok/s for DwarfStar.

Sep 7, 2026
Tone: mixed
reported speed:
35.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks whether anyone has run DeepSeek-V4-Flash-Strix-Halo-GGUF on a Strix Halo. The user notes it fits with ~64K context and is reported at 35 t/s. The user is cautiously optimistic about the result.

Sep 7, 2026
Tone: positive
reported speed:
26.8 tokens/s generation · 236.0 tokens/s prompt processing
quant:
IQ3_XXS (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash 0731 at 26.76 t/s sustained decode and 236 t/s prefill on a Strix Halo (Ryzen AI MAX+ 395, Radeon 8060S). Setup is llama.cpp Vulkan v0.6.1 with DSpark speculative decoding, a bf16 draft model, and a q8_0 KV cache. The user compares this with the DGX Spark and notes that q8_0 KV beats f16, the draft must be bf16, Vulkan beats ROCm, ubatch tops out at 2048, and a v0.6 regression is fixed in v0.6.1.

Sep 7, 2026
reported speed:
24.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 24 t/s on a Ryzen AI Max+ 395 and 51 t/s on a Radeon AI PRO R9700. The 51 t/s figure is attributed to the R9700, not the Strix Halo. User asks about t/s for various quants and settings.

Sep 7, 2026

Qwen3.8 27B

AMD Strix Halo 128GB · llama.cpp · 142,000 ctx

Tone: positive
reported speed:
17.5 tokens/s generation
quant:
Q8_0 (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Qwen3.8-27B Q8_0 on a Ryzen AI Max+ 395 in a ROG Flow Z13 with 128 GB unified memory, generating at 9-19 t/s and sustaining 16-19 t/s. Setup is Lemonade Server with llama.cpp ROCm, a Q8 KV cache, native MTP speculative decoding, a 64 GB VRAM and 64 GB RAM split, and roughly 142k context. MTP acceptance runs 97-99%. The model generated a flight simulator through an agent with file and bash tools.

Sep 7, 2026
Tone: positive
reported speed:
28.5 tokens/s generation · 210.0 tokens/s prompt processing
quant:
UD-IQ3_XXS (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a speculative decoding sweep with DSpark drafters, from a 20.48 t/s baseline to a peak of 28.5 t/s at n_max=3, a 1.39x gain. Q2_K_S and Q8_0 drafters land within 1-3% of each other. The user estimates 22-28 t/s for realistic mixed use, with prefill at 200-220 t/s at 64k context.

Sep 7, 2026
Tone: mixed
reported speed:
45.0 tokens/s generation
quant:
Q4_K_M (gguf)
kv:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Q4_K_M at 1M context on an AMD Strix Halo 128GB with an eGPU RTX 3080 Ti 12GB over Oculink, reaching about 45 t/s. Setup splits layers: 2GB of weights and FastMTP on the 3080, the rest on the iGPU. At 262K context, Q6 gives 10 t/s on the Strix alone, 24 t/s with MTP n=4, and about 28 t/s with FastMTP offloaded to the 3080. Q4_K_M at 262K context with 5GB of weights and FastMTP on the 3080 gives about 53 t/s. The user notes 60-70GB free on the AMD, enough to load an embedding and reranker, and compares to Qwen3.6-35B on 4x3090 at 450 t/s.

Sep 7, 2026

Qwen3.8 27B

AMD Strix Halo 128GB · llama.cpp · 65,536 ctx

Tone: positive
reported speed:
31.4 tokens/s generation · 300.0 tokens/s prompt processing
quant:
Q5_K_XL (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a dense 27B model on a Strix Halo APU, with a DFlash2 drafter giving +40% decode over MTP. Setup is Nathan's Vulkan fork of llama.cpp, with Q5_K_XL decoding faster than Q4_K_XL due to higher draft acceptance. Prefill is ~300 t/s at 3k context. FP4 builds are slower and worse quality.

Sep 7, 2026

Qwen3.8 27B

AMD Strix Halo 128GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
153.3 tokens/s generation
quant:
Q4_K_M (gguf)
kv:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User reports 153.32 t/s at 32K context and 87.74 t/s at 200K context. Setup uses a Q4 KV cache, speculative decoding with MTP and n-gram, and the vision encoder on the iGPU. HumanEval scored 159/164 in 29.7 min.

Sep 7, 2026
Tone: positive
reported speed:
23.0 tokens/s generation · 390.3 tokens/s prompt processing
quant:
UD-IQ4_XS (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports a hybrid attention model (Gated DeltaNet plus sparse QSA) with a 51B n-gram embedding table, which llama-bench reports as 176.94B all-in. Setup is a 3-part GGUF, 87 GiB on disk, built from llama.cpp PR #27742 with a crash fix, on the Vulkan (RADV) backend. The model uses about 91 GB resident at 131k context and loads in about 45s. The KV cache must stay F16, as quantized KV asserts.

Sep 7, 2026
reported speed:
10.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B dense at 10 t/s on Strix Halo 128GB. User considers eGPU options with 32GB and 48GB of VRAM for faster dense model inference and larger models. User mentions 0731 (DeepSeek V4 Flash) and dflash speculation.

Sep 7, 2026
Tone: positive
reported speed:
49.0 tokens/s generation · 682.0 tokens/s prompt processing
quant:
Q4_K_XL (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a hybrid setup with Strix Halo 128GB and an R9700 32GB over x4 PCIe. Dense layers, KV cache, and the MTP drafter run on the R9700, while routed experts run on the Strix, using a custom llama.cpp fork. The user compares this to stock at 24 tok/s.

Sep 7, 2026
Tone: positive
reported speed:
26.7 tokens/s generation · 138.6 tokens/s prompt processing
quant:
Q5_K_M (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

creative-writing

User reports average prefill and token generation throughput with an MTP draft model. User finds quality superior to Qwen3.8-27B for architectural tasks.

Sep 7, 2026
Tone: mixed
reported speed:
15.0 tokens/s generation
quant:
Q8_0 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8-Flash-Next at Q8_0 generating for over 45 minutes because of excessive thinking even at low reasoning effort. Setup is llama.cpp RPC on 2x Strix Halo 128GB. User compares the run with Qwen3.8-27B and DeepSeek V4 Flash on the same hardware.

Sep 7, 2026
Tone: positive
reported speed:
84.0 tokens/s generation · 408.0 tokens/s prompt processing
quant:
Q4_K_XL (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 84 t/s with 4 streams at 196K context on a model split across an iGPU (Vulkan) and an RTX 3090 Ti (CUDA). The model has 512 experts per layer, 36 layers of gated DeltaNet, 12 layers of top-k sparse attention, a 26.8 GiB n-gram table, and a built-in MTP draft head, at 103.69 GiB. Baseline on the iGPU alone is 22.2 tok/s. HumanEval+ scores 155/164, against 156/164 for Qwen3.8-27B on dual 3090 vLLM.

Sep 7, 2026

Qwen3.8 27B

AMD Strix Halo 128GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
26.0 tokens/s generation · 200.0 tokens/s prompt processing
quant:
Q4_K_XL (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports running the pi coding agent with llama.cpp on Strix Halo, decoding at 26+ t/s and prefilling at ~200 t/s in the fast band. Setup uses Qwen3.8-27B with the UD-Q4_K_XL quant, a DFlash2 drafter, f16 KV cache and 256k context. A Flash-Next model with IQ4_XS quant and MTP is also mentioned. The KV cache hit rate is ~94%.

Sep 7, 2026
reported speed:
237.9 tokens/s prompt processing
quant:
UD-Q8_K_XL (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6 35B A3B prompt processing at 237.9 t/s at depth 0 on Strix Halo 128GB.

Sep 3, 2026

LFM2.5 2.6B

AMD Strix Halo 128GB · llama.cpp · 128,000 ctx

Tone: positive
reported speed:
113.0 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-useagentic

User reports a 2.69B parameter model at 128K context with tool calling, benchmarked on a Ryzen AI Max+ 395, a phone at 30 t/s, and an M5 Max at 220 t/s. Benchmarks are ToolSandbox 77.83, IFBench 59.17, BFCLv4 56.88, and LiveCodeBench 59.41. The user does not recommend the model for agentic coding.

Aug 28, 2026
Tone: positive
reported speed:
50.0 tokens/s generation
quant:
Q8_XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen 3.6 35B at 50 t/s on Strix Halo. Setup is the Q8_XL quant. User emphasizes energy efficiency and versatility of Strix Halo compared to Nvidia cards.

Jul 11, 2026
reported speed:
21.2 tokens/s generation
quant:
Q4_K_M (gguf)
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MTP enabled with --spec-type draft-mtp --spec-draft-n-max 3. Baseline without MTP is 11.7 tok/s. Q8_0 was also tested: 7.4 → 18.1 tok/s (2.44×).

May 19, 2026
reported speed:
20.6 tokens/s generation
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.6-27B Dense with MTP on Strix Halo under Windows, at 18.7 t/s for a poem, 19.8 t/s for editing HTML and 23.2 t/s for creating HTML. Setup is llama.cpp with spec-draft-n-max 3.

May 17, 2026
reported speed:
15.5 tokens/s generation
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.6-27B Dense with MTP at 14.5 t/s on a poem task, 14.2 t/s on an edit HTML task, and 17.9 t/s on a create HTML task, on Strix Halo. Setup is llama.cpp on Windows with spec-draft-n-max 6.

May 17, 2026
reported speed:
12.6 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.6-27B Dense without MTP at 12.5 t/s on a poem task, 12.6 t/s on an edit HTML task, and 12.6 t/s on a create HTML task on Strix Halo. Setup is llama.cpp on Windows.

May 17, 2026

Qwen3.6 27B

AMD Strix Halo 128GB · llama.cpp · 128,000 ctx

Tone: mixed
quant:
Q8_0 (gguf)
flash attention:
on
long-context

User benchmarks MTP against non-MTP for 27B and 35B-A3B models. The 27B-MTP run shows significant speedup in generation and overall wall time for long-context chat. The 35B-MTP run shows mixed results: faster generation but slower end-to-end due to prefill overhead.

May 17, 2026
reported speed:
21.1 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports multiple models and quants on Strix Halo 128GB with the ROCm backend, plus tests on an RTX 3090 and an RTX 5070. Setup is the Q4_K_M quant for a chat workload.

May 17, 2026