Qwen3.8 27B
RTX 5070 Ti Laptop 12GB · Unsloth Studio · 8,192 ctx
- reported speed:
- 4.5 tokens/s generation · 23.5 tokens/s prompt processing
- quant:
- UD-Q4_K_XL
- kv:
- Q8
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
MTP acceptance 78-83%. Longer response 3.26 t/s, short factual 4.42 t/s, coding 4.53 t/s. Prompt processing 20-27 t/s. Model size ~17.9GB, offloading to CPU.