Qwen3.8 27B
3× V100 16GB · 256,000 ctx
- reported speed:
- 35.0 tokens/s generation · 650.0 tokens/s prompt processing
- quant:
- Q8
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at 30-40 t/s decode and 600-700 t/s prefill on 3x V100 16GB. Setup uses Q8 quantization with speculative decoding, MTP, and prefix caching, tuned by an automated agent. On heavy agentic work at 256k context, decode dropped to around 20 t/s. Qwen3.8 Flash Next at Q4 ran at 20 t/s decode and 90 t/s prefill. The build cost around $1500 and uses a 3D printed cooling block; GPU temperatures stay under 55C.