llamaperf

Qwen3.8 27B

on AMD RX 7600 XT 16GB · llama.cpp · 114,688 ctx

Tone: positive
Sep 23, 2026
Quant
GSQ-RCO-IQ3_XXS (GGUF)
KV cache
Q4_0
MTP (Multi-Token Prediction)
on
VRAM reported
16 GB

Use cases

codingagentic

Summary

User benchmarks Qwen3.8 27B on an RX 7600 XT 16GB, completing 15/15 tasks with 1.000 correctness in 348.5 seconds. Setup is llama.cpp HIP ROCm with the GSQ-RCO-IQ3_XXS GGUF and a Q4_0 KV cache at 114,688 context, fully in VRAM with no CPU offload. The same benchmark also ran Ornith-1.5-9B (15/15, 0.983, 184.9s) and K2-Horizon-7B (11/15, 0.909, 498.1s); the user calls Qwen3.8 27B the undisputed winner for agentic coding.