llamaperf

Qwen3.8 27B

on 2× AMD RX 6700 XT · llama.cpp · 144,000 ctx

Tone: positive
Sep 28, 2026
Throughput
18.0 t/s gen · 150.0 t/s pp
Quant
UD-IQ3_XXS (GGUF)
System RAM
46 GB
VRAM reported
22 GB

Use cases

agenticlong-context

Summary

User reports Qwen3.8 27B at 18 t/s generation and 150 t/s prompt processing on a 2007 Dell Precision T5400 with dual RX 6700 XT / RX 6700 (22GB total VRAM). Setup is llama.cpp with UD-IQ3_XXS quant at 144k context, using MTP q4_0 speculative decoding, on 24GB DDR2 and dual Xeon X5460. User compares five systems and argues older dual-Xeon dual-GPU setups beat a 2025 HP Omen with RTX 5070 in context length and speed, concluding DDR5 is not worth the cost for agentic tasks.