llamaperf
Sep 30, 2026
Throughput
51.0 t/s gen · 1500.0 t/s pp
Quant
IQ3_XXS (GGUF)
KV cache
8bit
System RAM
64 GB
VRAM reported
12 GB

Summary

User reports Qwen3.8 Flash-Next at 51 t/s generation and 1500 t/s prompt processing on an RTX 5070 Ti Laptop 12GB with 64GB RAM. Setup is the Strata engine with an IQ3_XXS GGUF quant and 8-bit KV cache, using 131k context. The model was partially offloaded, using 11GB VRAM and 56GB system RAM. The generation speed was measured at 43k context depth. The user compares this to stock llama.cpp, which reached 23 t/s generation and 100 t/s prompt processing with the same quant.