Qwen3.8 27B
2× RX 6800 16GB · llama.cpp · 131,072 ctx
- reported speed:
- 45.0 tokens/s generation
- quant:
- IQ4_XS
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User currently runs llama.cpp on two RTX 2060 12GB cards (24GB total) with Qwen3.8 27B IQ4_XS at 131k context, getting ~45 tok/s. Considering upgrade to RX 6800 16GB + RX 6800 XT 16GB (32GB total) and asks about performance and ROCm/Vulkan support.