- reported speed:
- 280.0 tokens/s generation
- quant:
- NVFP4 (NVFP4)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Victoria, a fine-tune of Qwen3.8-Flash-Next with 44% of experts cut via REAP, at 280 tok/s single stream on one B300 with the draft head, versus 135 without it. Setup is NVFP4 weights retrained at 4-bit, 48.0 GiB of weights including the draft head, with a separate 95.4 GiB n-gram table not counted in that number. Terminal-Bench 2.1 scored 70.0% averaged over 3 runs with an 8h per-task timeout, versus 62.5% for the previous NVFP4 build; HumanEval 159/164. The GGUF Q4_K_M build is 49.17 GiB and scored 75.3% on Terminal-Bench in a single noisy run and 93.2% on HumanEval averaged over 5 runs. Uses 35% fewer output tokens than the previous build. A second fine-tune, Maple, is a Canada-first model; its figures are not reported here.