llamaperf

Qwen3.8 27B

on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 52,000 ctx

Tone: mixed
Oct 6, 2026
Throughput
22-23 t/s gen
Quant
UD-Q6_K_XL (GGUF)
KV cache
f16

Use cases

codinglong-context

Summary

User reports Qwen3.8-27B GGUF quants fail to converge on long thinking tasks on 2x RTX 5060 Ti 16GB, while NVFP4 mixed-precision quants converge in ~17K thinking tokens. Setup is llama.cpp and vLLM with UD-Q6_K_XL GGUF (23.6 GB) and f16 KV cache, tensor-parallel across 2 GPUs. Throughput was healthy at 23-41 t/s on llama.cpp and 22-23 t/s on vLLM. The failure appears after 30K-38K thinking tokens with no closing token, tail-looping, or engine crash. The same model in NVFP4 (FP8 attention/GDN, FP4 MLP) converges in ~17K tokens on SGLang at 47-50 t/s. User hypothesizes recurrent state error accumulation under uniform quantization.