llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 3060 12GB · ik_llama.cpp · 262,144 ctx

Tone: positive
Sep 23, 2026
Throughput
22.1 t/s gen · 66.1 t/s pp
Quant
Q4_K_XL
KV cache
Q4
Flash Attention
on
MTP (Multi-Token Prediction)
on
Rating
4/5
System RAM
96 GB
VRAM reported
12 GB

Use cases

codingtool-useagenticlong-context

Summary

User reports Qwen3.8 Flash Next Q4_K_XL at 66.14 t/s prompt processing and 22.09 t/s generation on an RTX 3060 12GB, Intel i7-14700 and 96GB DDR5. Setup is ik_llama / llama-server on Windows 11 at 262,144 tokens of context with MTP3 and NCMOE48. A fresh OpenCode agent session produced a 10,890-token initial prompt, and MTP draft acceptance for that request was 89.5%. After the initial ~11K-token prefill, a second short prompt completed in 1.6s with short-generation throughput around 22–26 t/s. At approximately 71K active tokens, generation was around 10 t/s, with an OpenCode session average around 12 t/s.