Qwen3.8 27B
AMD RX 7600 XT 16GB · llama.cpp · 114,688 ctx
- quant:
- GSQ-RCO-IQ3_XXS (GGUF)
- kv:
- Q4_0
- mtp (multi-token prediction):
- on
User benchmarks Qwen3.8 27B on an RX 7600 XT 16GB, completing 15/15 tasks with 1.000 correctness in 348.5 seconds. Setup is llama.cpp HIP ROCm with the GSQ-RCO-IQ3_XXS GGUF and a Q4_0 KV cache at 114,688 context, fully in VRAM with no CPU offload. The same benchmark also ran Ornith-1.5-9B (15/15, 0.983, 184.9s) and K2-Horizon-7B (11/15, 0.909, 498.1s); the user calls Qwen3.8 27B the undisputed winner for agentic coding.