Qwen3.8 27B
M2 Ultra 192GB · llama.cpp · 131,072 ctx
- reported speed:
- 22.4 tokens/s generation · 360.2 tokens/s prompt processing
- quant:
- Q6_K_XL (gguf)
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports llama-bench results and a serving setup, with about 16 t/s on WebUI with a specific prompt.