llamaperf

Qwen3.8 125B (6B active) Flash-Next

on AMD Strix Halo 128GB · Gufo · 256,000 ctx

Tone: positive
Sep 29, 2026
Throughput
40-50 t/s gen
Quant
Q4 (GGUF)
System RAM
128 GB

Use cases

agentictool-use

Summary

User reports Qwen3.8 Flash-Next running at approximately 40-50 tok/s on a Ryzen Strix Halo PC with 128GB unified RAM. Setup is gufo with Q4 weights and 256k context, supporting 2 concurrent sessions. The user notes the model impressed them with its context understanding and tooling, and that it provides capability comparable to good API models.