Qwen3.8 27B
on 2× AMD RX 6700 XT · llama.cpp · 144,000 ctx
Sep 28, 2026
Use cases
agenticlong-context
Summary
User reports Qwen3.8 27B at 18 t/s generation and 150 t/s prompt processing on a 2007 Dell Precision T5400 with dual RX 6700 XT / RX 6700 (22GB total VRAM).
Setup is llama.cpp with UD-IQ3_XXS quant at 144k context, using MTP q4_0 speculative decoding, on 24GB DDR2 and dual Xeon X5460.
User compares five systems and argues older dual-Xeon dual-GPU setups beat a 2025 HP Omen with RTX 5070 in context length and speed, concluding DDR5 is not worth the cost for agentic tasks.