llamaperf

LFM2.5

Liquid AI · 3 reports

By engine

EngineAvg t/sRangeN
LM Studio129.0129–1291
llama.cpp113.0113–1131

LFM2.5 2.6B

Unknown GPU · custom engine · 128,000 ctx

reported speed:
17.0 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

Running on OnePlus 13, pure CPU. Custom inference engine built from scratch, 450kb, supports other model architectures.

LFM2.5 1.2B Instruct

RTX 4050 6GB · LM Studio · 256,000 ctx

Tone: positive
reported speed:
129.0 tokens/s generation · 8500.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-usesummarization

User benchmarked 20 small models on RTX 4050 6GB. Selected LFM2.5-1.2B-Instruct as cheap always-on model. Also tested Granite 4.1 3B, Gemma-4-agentic-e2b, Nemotron-3-Nano-4B, LFM2.5-8B-A1B. Full results table included.

LFM2.5 2.6B

AMD Strix Halo 128GB · llama.cpp · 128,000 ctx

Tone: positive
reported speed:
113.0 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-useagentic

Vendor benchmark on Ryzen AI Max+ 395. Also reports 30 tok/s on phone and 220 tok/s on M5 Max. Model is 2.69B params, 128K context, tool calling. Benchmarks: ToolSandbox 77.83, IFBench 59.17, BFCLv4 56.88, LiveCodeBench 59.41. Not recommended for agentic coding.