Qwen3.8 27B
on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 52,000 ctx
Oct 6, 2026
Use cases
codinglong-context
Summary
User reports Qwen3.8-27B GGUF quants fail to converge on long thinking tasks on 2x RTX 5060 Ti 16GB, while NVFP4 mixed-precision quants converge in ~17K thinking tokens.
Setup is llama.cpp and vLLM with UD-Q6_K_XL GGUF (23.6 GB) and f16 KV cache, tensor-parallel across 2 GPUs. Throughput was healthy at 23-41 t/s on llama.cpp and 22-23 t/s on vLLM.
The failure appears after 30K-38K thinking tokens with no closing token, tail-looping, or engine crash. The same model in NVFP4 (FP8 attention/GDN, FP4 MLP) converges in ~17K tokens on SGLang at 47-50 t/s. User hypothesizes recurrent state error accumulation under uniform quantization.