- reported speed:
- 43.0 tokens/s generation
- quant:
- MXFP4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 34 t/s at the start of a run, rising to 43 t/s by the end. The cached tokens were the default chat prompt and the query was 13k.