Occamy 1.0
on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 175,000 ctx
Sep 24, 2026
Use cases
agenticlong-context
Summary
User reports Occamy 1.0 at around 100 tok/s on two RTX 5060 Ti cards, with enough memory for two independent ~175K context pools for concurrent subagents.
Setup uses llama.cpp on the desktop worker box; the cards cost about $400 each.
A 96 GB M5 Ultra Mac Studio is planned as the primary inference appliance running Qwen 3.8 Next Flash through oMLX, with DeepSeek V4.1 Flash via API filling in until it arrives.