Qwen3.8 27B
on 2× NVIDIA Tesla P100 16GB · llama.cpp · 262,144 ctx
Sep 27, 2026
Use cases
codingcreative-writing
Summary
User reports Qwen3.8 27B at 50-60 t/s generation and 350 t/s prefill on 2x Tesla P100 16GB with a custom llama.cpp fork.
Setup is llama.cpp with Q6_K quant, 262144 context, batch 32768, ubatch 1024, and MTP speculative decoding.
At 260k context, generation is 30-35 t/s and prefill is 110 t/s. GPUs are capped at 175W/250W each and run at 79C with minor thermal throttling; user estimates 5-10% higher numbers with better cooling.