Qwen3.8 27B
AMD RX 7600 XT 16GB · llama.cpp · 163,840 ctx
- reported speed:
- 18.0 tokens/s generation · 141.0 tokens/s prompt processing
- quant:
- IQ3_XXS (GGUF)
- kv:
- q8_0
- mtp (multi-token prediction):
- off
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at 18 t/s decode and 141 t/s prefill on a 1.8k-token prompt, running on an AMD Radeon RX 7600 XT 16 GB at 163,840 context. Setup is llama.cpp llama-server build 10480 with the unsloth Qwen3.8-27B-UD-IQ3_XXS GGUF (10.2 GiB), q8_0 KV cache, flash attention on, and Vulkan (RADV) backend. The model is a hybrid architecture where only 16 of 64 layers keep a full KV cache, so KV is about 5.3 GiB at this context. With MTP on the same model runs about 24 t/s at 98k context and 39 t/s at 124k. The user notes 160k is a VRAM-math ceiling, not a quality claim, and that decode drops to about 6 t/s if the GPU spills to system RAM.