llamaperf
Oct 8, 2026
Throughput
112.0 t/s gen
Quant
MXFP4 (GGUF)
VRAM reported
32 GB

Use cases

coding

Summary

User reports 112 tok/s decode on coding prompts with WHIRL, the engine they wrote, for Swift-1.5 27B MXFP4 on a single Radeon AI PRO R9700, against 61 tok/s for llama.cpp b11214 on the same GGUF. Setup is one Radeon AI PRO R9700 32GB over USB4 on Windows 11, MXFP4 weights, WHIRL v0.1.3 with speculative decoding, single request. Prefill reaches 1,757 tok/s at 128K and 970 at 256K, and 4 concurrent users get 182 tok/s in total. On Ornith-1.5-35B-A3B MXFP4 WHIRL decodes at 258 tok/s and prefills over 11,000 tok/s at 8K.