llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 4090 · Strata · 65,536 ctx

Tone: mixed
Oct 5, 2026
Throughput
166.4 t/s gen
Quant
IQ2_XS (GGUF)
KV cache
INT8
System RAM
64 GB
VRAM reported
24 GB

Use cases

coding

Summary

User reports Qwen3.8-Flash-Next at 166.36 output tokens/s on a rented RTX 4090 24 GB. Setup is Strata with IQ2_XS, INT8 KV cache, 65,536-token context and 32,768 GPU-resident tokens, MTP draft 4 at 0.5 minimum probability. The rate is a weighted generation figure (sum of completion tokens over predicted time), not whole-request throughput; MTP acceptance was 84.9% and peak sampled GPU memory 23,918 MiB. The generated Hill Climb Racing-style game was rejected for poor vehicle physics and game feel.