llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX 5060 8GB
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Qwen3.8 27B

NVIDIA RTX 5060 8GB · llama.cpp · 8,192 ctx

reported speed:
30.1 tokens/s generation
quant:
UD-IQ2_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at 30.10 tok/s median decode on an RTX 5060 8GB, over five runs at an occupied 8192-token prompt. Setup is llama.cpp with UD-IQ2_XXS weights and q4_0 KV cache, 65/65 layers on CUDA, peak VRAM 7767 MiB. A short-context FULL_GPU decode of 31.39 tok/s and a ctx512 control of 18.11 tok/s are also given; quality gate and uncensored checkpoint are not done.

Oct 7, 2026
reported speed:
2.4 tokens/s generation
quant:
FP8 (safetensors)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4.1-Flash (552B backbone + 196B Engram, 384 experts top-6, FP8 + FP4) running on a single RTX 5060 8GB at ~1.6 tokens/s from disk only and ~2.4 tokens/s with a 16 GB RAM cache. Setup streams experts and Engram rows from NVMe using DeepSeek's reference code with a modified storage layer; only the dense part lives on the GPU (cap --vram_gb 7.3), with 240 experts (4.2 GiB) read per token. Batch size 1, text only, context limit 8192, prefill max 700 tokens. Live chat via the OpenAI-compatible server with a 16 GB RAM cache reaches 2.43 tokens/s; peak VRAM 7.06 GiB. A 311-token prompt takes 17.9 s. The user notes the speed-ups changed nothing in output (136/136 tokens identical) and that DeepSeek's untouched reference could not be run side by side because it cannot load the model in 8 GB.

Oct 5, 2026
reported speed:
40.0 tokens/s generation · 500.0 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks whether speculative decoding (MTP) is still worth enabling for a 35B A3B MoE model offloaded across an 8GB RTX 5060 and 32GB of system RAM. Current llama-server setup uses a Q4_K_XL GGUF with a Q8 KV cache, flash attention on, 4096 batch and ubatch, 16 CPU cores, 40 MoE layers on CPU and 99 GPU layers, yielding about 40 t/s generation and 500 t/s prompt processing without MTP. User recalls earlier reports that MTP hurt prompt processing and wants to know if that is still the case and how others configure llama-server.

Sep 18, 2026
Showing 1–3 of 3
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409037NVIDIA RTX 5070 Ti30

Records by model

1397 total
Qwen3.8781
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other213

Records by engine

1074 total
llama.cpp570
vLLM153
Strata47
NInfer40
Ollama34
other230

Use cases

coding 453agentic 287long-context 208tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding453agentic287long-context208tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B468Qwen3.8 125B · 6B active283DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22GLM-5.3 320B · 18B active18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24Q6_K23