llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX 3090
Compare setup details (1 active)

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
43.9 tokens/s generation · 1817.0 tokens/s prompt processing
quant:
Q6_K_XL (GGUF)
kv:
int8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-Flash-Next at 43.9 tok/s generation and 1817 tok/s prompt processing on an RTX 3090 + RTX A4000 at 32k context. Setup is Strata with UD-Q6_K_XL GGUF and int8 KV cache, 262144 native context, speculative decoding with MTP, experts offloaded to CPU. At 200k context Strata Q6 does 48.8 tok/s versus llama.cpp's 10.6 tok/s, and TTFT drops from 595 s to 118 s. Q6 is 12-31% faster than Q8 and 19 GB smaller.

Oct 7, 2026
Tone: mixed
reported speed:
97.0 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Qwen3.8-Flash-Next at 97 t/s decode on 2x RTX 3090 with Strata, with about half the experts in system RAM. Setup is Strata v0.1.40.1 with unsloth UD-Q4_K_XL GGUF, a layer split, speculative decoding, and 121 GB DDR4 on a Ryzen 9 3950X. The 97 t/s is at a 95% expert hit rate; shrinking the cache to 90% and 81% hit rates gave 84 t/s and 68 t/s. The user estimates +34% decode from eliminating misses on their box and about +75% headroom at a single-24GB-card share, and reports +16% decode from --pcie-frac 0.2 --spec-min-p 0.7 over about 280 runs. The prefetch feature is not implemented yet.

Oct 7, 2026
reported speed:
19.8 tokens/s generation · 437.8 tokens/s prompt processing
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports GLM-5.3-Flash at 19.8 tok/s decode and 437.8 tok/s prefill on one RTX 3090 24 GB with 251 GB RAM. Setup is the Strata engine with a UD-Q4_K_XL GGUF pack, --chunk 4096, 16K-token prompt, tiered expert cache streaming from RAM and NVMe. Decode splits into 12.4 ms expert copies, 0.2 ms expert kernels and 37.2 ms dense/sync per 49.8 ms token. A 64K prompt prefills at 420.7 tok/s; a 1K prompt at 132 tok/s prefill and 19.8 tok/s decode; disk-only tier decodes at 1.15 tok/s. --chunk 8192 OOMs on 24 GB.

Oct 6, 2026
Tone: positive
reported speed:
116.0 tokens/s generation · 2908.0 tokens/s prompt processing
quant:
IQ3_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticvisionlong-context

User reports Qwen3.8-Flash-Next at 116 t/s decode and 2,908 t/s prompt read on two RTX 3090s with NVLink. Setup is a modified Strata build (v0.1.38 plus 29 commits) with IQ3_S GGUF at 262K context, greedy, median of 3, on a Ryzen 9 3950X with 121 GB RAM. The second card acts as a peer expert tier read over NVLink; the pair holds about 19,500 of 24,576 IQ3_S experts with a 0.98 hit rate on real agent chats. Decode gains from the second card are modest (+10%) because the verify window is latency-bound on the primary. Sampled with official settings it sits around 102-104 t/s, a 1.5K prompt reads at ~1,670 t/s, and switching chats takes 1 s. One stream at a time; answers are word-for-word identical to the one-card run.

Oct 6, 2026
Tone: positive
reported speed:
37-40 tokens/s generation · 1700.0 tokens/s prompt processing
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running DeepSeek V4 Flash with Strata on an RTX 3090 alone at 1700 t/s prefill and 37-40 t/s generation, and on both an RTX 3090 and RTX 5070 Ti at 1850 t/s prefill and 45-50 t/s generation, at 131K context. Setup is Strata with UD-Q4_K_XL quant, 96GB DDR4 system RAM, RTX 5070 Ti on PCIe x16 and RTX 3090 on PCIe x4. User compares against Unsloth Studio on the same model at 131K context, which gave 100 t/s prefill and 14-16 t/s generation, with a 100K conversation taking over 10 minutes to process. User notes Strata initially lacked prefix caching but now supports it, and that the same generation speed matches Qwen3.8-27B-UD-Q8_K_L on both GPUs.

Oct 5, 2026
Tone: positive
reported speed:
34.0 tokens/s generation · 520.0 tokens/s prompt processing
quant:
Q8 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next at Q8 running at 34 t/s generation and up to 520 t/s prefill on 2x RTX 3090 with 256GB of system RAM. Setup uses the Strata engine with MTP set to 5; most of the model sits in system RAM rather than GPU memory. For comparison, the user says llama.cpp without MTP gives 13 t/s generation and 115 t/s prefill on the same setup, and notes MTP usually does not help when most of the model is in system RAM.

Oct 5, 2026
Tone: positive
reported speed:
167.0 tokens/s generation · 2466.0 tokens/s prompt processing
quant:
IQ3_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next at 2466 t/s prompt processing and 167 t/s generation on an RTX 3090 plus RTX 5070 Ti with the Strata engine and IQ3_S quant. Setup is Strata with IQ3_S; the user also added UD-Q4_K_XL support, which reached 2341 t/s prompt and 126 t/s generation. The user doubled Strata throughput over a week of profiling and reached 5x llama.cpp on UD-Q4_K_XL, where the starting point was 6 t/s. Earlier steps were 21 t/s, 27 t/s and 51 t/s on IQ3_XXS.

Oct 3, 2026
Tone: positive
reported speed:
38-61 tokens/s generation · 1650.0 tokens/s prompt processing
quant:
UD-Q3_K_XL (GGUF)
kv:
fp16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingtool-uselong-context

User reports Qwen3.8 Flash-Next running on an RTX 3090 24GB with 128GB system RAM, reaching about 1650 t/s prompt processing and 38 to 61 t/s generation depending on context under Strata. Setup is Strata with Unsloth UD-Q3_K_XL quant, fp16 KV cache, 256k context, speculative decoding with 4 draft tokens, and an expert cache of 6517 slots (~14GB). The same model on llama.cpp master gave up to 700 t/s PP and 23 t/s TG. Generation is a range: 38 t/s at 182k context and about 61 t/s at short context. The user notes the quant Strata recommends by default was faster but produced minor errors and lower quality, so Unsloth's was chosen for accuracy.

Oct 3, 2026
Showing 1–8 of 8
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090111AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1385 total
Qwen3.8771
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other212

Records by engine

1063 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other225

Use cases

coding 444agentic 282long-context 204tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding444agentic282long-context204tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active278DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q436Q8_0354-bit24Q6_K23