llamaperf

Qwen3.8 125B (6B active) Flash-Next

on AMD RX 7600 8GB · Strata · 32,768 ctx

Tone: mixed
Oct 6, 2026
Throughput
24.0 t/s gen · 67.5 t/s pp
Quant
IQ2_XS (GGUF)
System RAM
64 GB
VRAM reported
8 GB

Summary

User reports Qwen3.8-Flash-Next at 24.0 tok/s decode and 67.5 tok/s prompt on an RX 7600 8GB. Setup is Strata 0.1.39 with IQ2_XS and a 32768-token context, splitting MoE experts across GPU, CPU and RAM with about 33 GB of pinned RAM. User also reports Qwen3.6-35B-A3B Q6_K_XL on llama.cpp at 23.7 tok/s decode without MTP and 32.0 tok/s with MTP (98% draft acceptance), 161.7 tok/s prompt. Strata runs one model and one request at a time; prompt speed is limited by the 8 GB card reading prompts in 256-token chunks.