Qwen3.8 27B
M5 16GB · llama.cpp · 8,192 ctx
- reported speed:
- 9.0 tokens/s generation
- quant:
- Q3_xxs
- kv:
- 8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
summarization
User is new to local LLMs and asks if ~9 t/s is normal for this setup. They also ask for model recommendations for PDF summaries on 16GB RAM.