llamaperf

Ornith1.5

2 reports

Thin page (2 of 3 reports needed for indexing). Add yours.

By engine

EngineAvg t/sRangeN
Ollama64.064–641
llama.cpp36.036–361

Ornith1.5 35B (3B active)

Unknown GPU · Ollama

Tone: positive
reported speed:
64.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User fixed the MTP head on Ornith1.5 35B A3B, achieving 64 t/s (up from 60) and 33% faster wall clock. Mentions running on a PC with a hackRF receiver and a 5 watt quansheng portable. Model is a splice of a trained MTP head onto an APEX requant. Link to Ollama model and testing results provided.

Ornith1.5 9B

RX 9060 XT 16GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
36.0 tokens/s generation · 950.0 tokens/s prompt processing
quant:
Q6_K (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User tested Ornith-1.5-9B on AMD RX 9060 XT 16GB with llama.cpp. Reports ~950 tok/s prompt eval and ~36 tok/s generation, dropping to ~500/25 tok/s during long tasks. Ran continuously for ~3.5 hours on a coding agent task. Contrasts with Qwen 3.8 27B which was frustratingly slow on same hardware.