Qwen3.8 27B
NVIDIA RTX 4070 Super · llama.cpp · 64,000 ctx
- reported speed:
- 18.0 tokens/s generation · 500-600 tokens/s prompt processing
- quant:
- IQ3_S (GGUF)
- kv:
- q4_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-27B at ~18 tok/s decode and ~500-600 tok/s prefill at 64k context on an RTX 4070 Super 12GB with 32GB system RAM. Setup is stock llama.cpp with an ISTA-DASLab GSQ-RCO IQ3_S quant (~11GB), q4_0 KV cache, -ngl 58 and token_embd offloaded to CPU. Only 16 of 64 layers need KV cache, so 64k context takes about 1.1GB instead of 4GB at f16. The same recipe works with 0bserverx' Qwen3.8-27B-Heretic-GSQ-RCO IQ3_S quant at about a 7% speed loss. User calls GSQ-RCO the best Qwen3.8-27B quant tested and says it performs close to full precision.