Bonsai 2 27B Ternary
on NVIDIA RTX 4090 · NInfer · 262,144 ctx
Sep 26, 2026
Use cases
codingtool-uselong-contextmathmultilingual
Summary
User reports Ternary Bonsai 2 27B at 188 tok/s decode averaged over six mixed prompts on a single RTX 4090 under native Windows.
Setup is NInfer, a from-scratch C++/CUDA engine, with Q4/Q5 weights, E8 4-bit KV cache, MTP plus n-gram speculative decoding, and 262K context per request.
On edit-style prompts the model reaches 250 tok/s with MTP and 532 tok/s with MTP + n-gram; through the server on 7K-11K-token file edits with thinking on it goes 215 to 378 tok/s. Three concurrent requests give 360 tok/s aggregate. Against Prism's llama.cpp fork on the same card: 101 vs 77 tok/s decode and 3,061 vs 1,363 tok/s pp512. N-gram gains depend on repetition and are zero on free prose (167 tok/s both ways).