Thin page (2 of 3 reports needed for indexing).
Add yours.
- reported speed:
- 64.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User fixed the MTP head on Ornith1.5 35B A3B, achieving 64 t/s (up from 60) and 33% faster wall clock. Mentions running on a PC with a hackRF receiver and a 5 watt quansheng portable. Model is a splice of a trained MTP head onto an APEX requant. Link to Ollama model and testing results provided.
- reported speed:
- 36.0 tokens/s generation · 950.0 tokens/s prompt processing
- quant:
- Q6_K (gguf)
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingagentic
User tested Ornith-1.5-9B on AMD RX 9060 XT 16GB with llama.cpp. Reports ~950 tok/s prompt eval and ~36 tok/s generation, dropping to ~500/25 tok/s during long tasks. Ran continuously for ~3.5 hours on a coding agent task. Contrasts with Qwen 3.8 27B which was frustratingly slow on same hardware.