llamaperf
Oct 9, 2026
Throughput
54.1 t/s gen · 645.1 tokens/s prompt processing (prefill)Prompt processing measures input prefill speed; it does not tell you the time to first token.
Quant
MXFP4 (GGUF)
System RAM
128 GB

Use cases

agentictool-use

Summary

User reports gpt-oss-120b at 54.11 t/s generation and 645.07 t/s prompt processing on a GMKtec EVO-X2 with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB unified memory. Setup is llama.cpp build 11456 with MXFP4 GGUF weights, Vulkan (Mesa RADV) backend, flash attention on, whole model on the iGPU, no speculative decoding, one request at a time. Three runs after a warm-up gave 54.11, 54.64 and 54.50 t/s generation with 645.07, 656.81 and 657.14 t/s prompt processing on an 839-token prompt. Two small embedding and reranker models were loaded but idle. The user runs this model for agentic tool use and research.