Qwen3.8 27B
NVIDIA RTX 4060 · llama.cpp · 16,384 ctx
- generation:
- 6.0 tokens/s
- quant:
- UD-Q4_K_XL (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.
User reports Qwen3.8-27B at ~5.99 tok/s with MTP speculative decoding on an RTX 4060 8GB with 32GB system RAM. Setup is llama.cpp with UD-Q4_K_XL GGUF at 16K context, one parallel slot, partial CPU offload since the 17.56GB model exceeds 8GB VRAM. MTP accepted 99 of 126 draft tokens (~79%), improving decode by roughly 57% over the ~3.82 tok/s without MTP. The IQ4_XS quant measured ~4.17 tok/s. User also ran Qwen3-VL 4B Q4_K_M invoice extraction: 5 of 6 image-only invoices completed in 8-38 sec, and 4 of 6 with OCR transcripts in 25-40 sec. In a 2-invoice comparison Qwen matched 22 of 24 top-level fields versus Gemma's 5 of 24. User notes these are application-level results, not standardized benchmarks.