Ornith1.5 35B (3B active)
on NVIDIA RTX 5060 Laptop 8GB · llama.cpp
Sep 26, 2026
Use cases
codingagentic
Summary
User reports Ornith 1.5 35B-A3B at 30-40 t/s generation and 300-400 t/s prompt processing on an RTX 5060 Laptop 8GB with 32GB DDR5 RAM.
Setup uses llama.cpp via OpenCode, with generation at 30 t/s for 100k-140k context and up to 40 t/s below 100k context.
User also ran Qwen 3.8 27B on the same hardware and compares the two models for coding tasks, noting Ornith handles contained tasks well but struggles with complex multi-file changes.