llamaperf

Qwen3.8 27B

on NVIDIA RTX 4070 Super · llama.cpp · 64,000 ctx

Tone: positive
Oct 7, 2026
Throughput
18.0 t/s gen · 500-600 t/s pp
Quant
IQ3_S (GGUF)
KV cache
q4_0
System RAM
32 GB
VRAM reported
12 GB

Summary

User reports Qwen3.8-27B at ~18 tok/s decode and ~500-600 tok/s prefill at 64k context on an RTX 4070 Super 12GB with 32GB system RAM. Setup is stock llama.cpp with an ISTA-DASLab GSQ-RCO IQ3_S quant (~11GB), q4_0 KV cache, -ngl 58 and token_embd offloaded to CPU. Only 16 of 64 layers need KV cache, so 64k context takes about 1.1GB instead of 4GB at f16. The same recipe works with 0bserverx' Qwen3.8-27B-Heretic-GSQ-RCO IQ3_S quant at about a 7% speed loss. User calls GSQ-RCO the best Qwen3.8-27B quant tested and says it performs close to full precision.