llamaperf

Qwen3.8 27B

on 2× NVIDIA Tesla P100 16GB · llama.cpp · 262,144 ctx

Tone: positive
Sep 27, 2026
Throughput
50-60 t/s gen · 350.0 t/s pp
Quant
Q6_K (GGUF)

Use cases

codingcreative-writing

Summary

User reports Qwen3.8 27B at 50-60 t/s generation and 350 t/s prefill on 2x Tesla P100 16GB with a custom llama.cpp fork. Setup is llama.cpp with Q6_K quant, 262144 context, batch 32768, ubatch 1024, and MTP speculative decoding. At 260k context, generation is 30-35 t/s and prefill is 110 t/s. GPUs are capped at 175W/250W each and run at 79C with minor thermal throttling; user estimates 5-10% higher numbers with better cooling.