llamaperf

Qwen3.8 27B

on NVIDIA RTX 4070 Ti · llama.cpp · 4,096 ctx

Sep 28, 2026
Throughput
43.6 t/s gen · 943.3 t/s pp
Quant
UD-IQ2_XXS (GGUF)
KV cache
Q8
System RAM
64 GB
VRAM reported
12 GB

Summary

User benchmarks Qwen3.8-27B at 43.643 generation tok/s and 943.278 prompt tok/s on an RTX 4070 Ti 12GB. Setup is llama.cpp b10448 with UD-IQ2_XXS GGUF, Q8 KV cache, 4096-token context, all 66 layers on GPU, one slot, 256-token outputs. A second quant UD-Q2_K_XL decoded 38.030 tok/s and 995.095 prompt tok/s with 1,583 MiB more peak VRAM. An IQ2 context ladder gave 41.124 tok/s at 4K, 40.522 at 8K and 39.201 at 16K (12,831 prompt tokens, 11.119 s TTFT). An isolated IQ2 draft-mtp experiment raised throughput 47.284% on prose and 92.651% on Python code, but prose output diverged at token 16, so MTP stays off by default. In a 24-task quality pass, Q2 scored 10/24 and IQ2 9/24 (McNemar p = 1.0).