llamaperf

Qwen3.8 27B

on 2× NVIDIA Tesla P100 16GB · llama.cpp

Tone: positive
Sep 20, 2026
Throughput
40.0 t/s gen
Quant
Q6_K (GGUF)

Summary

User reports Qwen3.8 27B at 40+ t/s on two Tesla P100s, with a roughly 30 t/s average across long contexts and up to 55 t/s at 0 context. Setup is a custom llama.cpp fork with P100 kernel optimizations, Q6_K quant, both cards capped at 175W and communicating over PCIe gen 3. Speeds vary by about ±2 t/s; the user notes fp16 math saves 40% or more on prefill with negligible accuracy loss.