llamaperf
Oct 7, 2026
Throughput
97.0 t/s gen
Quant
UD-Q4_K_XL (GGUF)
System RAM
121 GB

Use cases

codingagentic

Summary

User reports Qwen3.8-Flash-Next at 97 t/s decode on 2x RTX 3090 with Strata, with about half the experts in system RAM. Setup is Strata v0.1.40.1 with unsloth UD-Q4_K_XL GGUF, a layer split, speculative decoding, and 121 GB DDR4 on a Ryzen 9 3950X. The 97 t/s is at a 95% expert hit rate; shrinking the cache to 90% and 81% hit rates gave 84 t/s and 68 t/s. The user estimates +34% decode from eliminating misses on their box and about +75% headroom at a single-24GB-card share, and reports +16% decode from --pcie-frac 0.2 --spec-min-p 0.7 over about 280 runs. The prefetch feature is not implemented yet.