llamaperf

Occamy 1.0

on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 175,000 ctx

Tone: positive
Sep 24, 2026
Throughput
100.0 t/s gen

Use cases

agenticlong-context

Summary

User reports Occamy 1.0 at around 100 tok/s on two RTX 5060 Ti cards, with enough memory for two independent ~175K context pools for concurrent subagents. Setup uses llama.cpp on the desktop worker box; the cards cost about $400 each. A 96 GB M5 Ultra Mac Studio is planned as the primary inference appliance running Qwen 3.8 Next Flash through oMLX, with DeepSeek V4.1 Flash via API filling in until it arrives.