LFM2.5 2.6B
Unknown GPU · custom engine · 128,000 ctx
- reported speed:
- 17.0 tokens/s generation
- quant:
- Q4_K_M (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Running on OnePlus 13, pure CPU. Custom inference engine built from scratch, 450kb, supports other model architectures.