LFM2.5 1.2B Instruct
RTX 4050 6GB · LM Studio · 256,000 ctx
- reported speed:
- 129.0 tokens/s generation · 8500.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
tool-usesummarization
User benchmarked 20 small models on RTX 4050 6GB. Selected LFM2.5-1.2B-Instruct as cheap always-on model. Also tested Granite 4.1 3B, Gemma-4-agentic-e2b, Nemotron-3-Nano-4B, LFM2.5-8B-A1B. Full results table included.