Qwen3.8 125B (6B active) Flash-Next
on AMD RX 7600 8GB · Strata · 32,768 ctx
Oct 6, 2026
Summary
User reports Qwen3.8-Flash-Next at 24.0 tok/s decode and 67.5 tok/s prompt on an RX 7600 8GB.
Setup is Strata 0.1.39 with IQ2_XS and a 32768-token context, splitting MoE experts across GPU, CPU and RAM with about 33 GB of pinned RAM.
User also reports Qwen3.6-35B-A3B Q6_K_XL on llama.cpp at 23.7 tok/s decode without MTP and 32.0 tok/s with MTP (98% draft acceptance), 161.7 tok/s prompt. Strata runs one model and one request at a time; prompt speed is limited by the 8 GB card reading prompts in 256-token chunks.