Qwen3.8 125B (6B active) Flash-Next
2× A100 80GB · TensorSharp
- reported speed:
- 56.8 tokens/s generation · 1019.1 tokens/s prompt processing
- quant:
- UD-Q2_K_XL (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks TensorSharp against llama.cpp on Qwen 3.8 Flash Next, reporting prompt processing and generation speeds for 1 GPU and 2 GPU layer split configurations. The promptTps and generationTps fields capture the 1 GPU pp512 and tg64 values respectively.