Qwen3.8 27B
on 2× NVIDIA Tesla P100 16GB · llama.cpp
Sep 20, 2026
Summary
User reports Qwen3.8 27B at 40+ t/s on two Tesla P100s, with a roughly 30 t/s average across long contexts and up to 55 t/s at 0 context.
Setup is a custom llama.cpp fork with P100 kernel optimizations, Q6_K quant, both cards capped at 175W and communicating over PCIe gen 3.
Speeds vary by about ±2 t/s; the user notes fp16 math saves 40% or more on prefill with negligible accuracy loss.