Qwen3.8 27B
on NVIDIA RTX 4060 · llama.cpp · 16,384 ctx
Oct 11, 2026
Use cases
codingagenticvision
Summary
User reports Qwen3.8-27B at ~5.99 tok/s with MTP speculative decoding on an RTX 4060 8GB with 32GB system RAM.
Setup is llama.cpp with UD-Q4_K_XL GGUF at 16K context, one parallel slot, partial CPU offload since the 17.56GB model exceeds 8GB VRAM. MTP accepted 99 of 126 draft tokens (~79%), improving decode by roughly 57% over the ~3.82 tok/s without MTP. The IQ4_XS quant measured ~4.17 tok/s.
User also ran Qwen3-VL 4B Q4_K_M invoice extraction: 5 of 6 image-only invoices completed in 8-38 sec, and 4 of 6 with OCR transcripts in 25-40 sec. In a 2-invoice comparison Qwen matched 22 of 24 top-level fields versus Gemma's 5 of 24. User notes these are application-level results, not standardized benchmarks.