gpt-oss 117B (5.1B active)
AMD Strix Halo 128GB · llama.cpp
- generation:
- 54.1 tokens/s
- prompt processing (prefill):
- 645.1 tokens/s
- quant:
- MXFP4 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.
User reports gpt-oss-120b at 54.11 t/s generation and 645.07 t/s prompt processing on a GMKtec EVO-X2 with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB unified memory. Setup is llama.cpp build 11456 with MXFP4 GGUF weights, Vulkan (Mesa RADV) backend, flash attention on, whole model on the iGPU, no speculative decoding, one request at a time. Three runs after a warm-up gave 54.11, 54.64 and 54.50 t/s generation with 645.07, 656.81 and 657.14 t/s prompt processing on an 839-token prompt. Two small embedding and reranker models were loaded but idle. The user runs this model for agentic tool use and research.