This page is thin (2 of 3 reports needed for indexing).
Help fill it in.
- reported speed:
- 16.0 tokens/s generation
- quant:
- IQ3_XXS
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports a patched engine improved speed from 5-6 to 15-16 tok/s.
Setup uses the Unsloth UD-IQ3_XXS quant.
The user notes a wired limit at 120GB.
- reported speed:
- 31.2 tokens/s generation · 530.2 tokens/s prompt processing
- quant:
- IQ4_NL (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks a custom llama.cpp branch with faster Metal inference and n-gram SSD offload at pp512 530.18 t/s and tg128 31.17 t/s with resident n-gram.
The model is 68.37 GiB and 125.74 B params.
With SSD read mode, pp512 is 197.92 t/s and tg128 is 28.47 t/s.