llamaperf

NVIDIA RTX 5060 8GB

NVIDIA · 8GB · 3 reports

As of 7 Oct 2026, the models most run on the NVIDIA RTX 5060 8GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 2 · ds4 1

Run models on your NVIDIA RTX 5060 8GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 5060 8GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 8 GB of VRAM.

Qwen3.8 27B

NVIDIA RTX 5060 8GB · llama.cpp · 8,192 ctx

reported speed:
30.1 tokens/s generation
quant:
UD-IQ2_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at 30.10 tok/s median decode on an RTX 5060 8GB, over five runs at an occupied 8192-token prompt. Setup is llama.cpp with UD-IQ2_XXS weights and q4_0 KV cache, 65/65 layers on CUDA, peak VRAM 7767 MiB. A short-context FULL_GPU decode of 31.39 tok/s and a ctx512 control of 18.11 tok/s are also given; quality gate and uncensored checkpoint are not done.

Oct 7, 2026
reported speed:
2.4 tokens/s generation
quant:
FP8 (safetensors)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4.1-Flash (552B backbone + 196B Engram, 384 experts top-6, FP8 + FP4) running on a single RTX 5060 8GB at ~1.6 tokens/s from disk only and ~2.4 tokens/s with a 16 GB RAM cache. Setup streams experts and Engram rows from NVMe using DeepSeek's reference code with a modified storage layer; only the dense part lives on the GPU (cap --vram_gb 7.3), with 240 experts (4.2 GiB) read per token. Batch size 1, text only, context limit 8192, prefill max 700 tokens. Live chat via the OpenAI-compatible server with a 16 GB RAM cache reaches 2.43 tokens/s; peak VRAM 7.06 GiB. A 311-token prompt takes 17.9 s. The user notes the speed-ups changed nothing in output (136/136 tokens identical) and that DeepSeek's untouched reference could not be run side by side because it cannot load the model in 8 GB.

Oct 5, 2026
reported speed:
40.0 tokens/s generation · 500.0 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks whether speculative decoding (MTP) is still worth enabling for a 35B A3B MoE model offloaded across an 8GB RTX 5060 and 32GB of system RAM. Current llama-server setup uses a Q4_K_XL GGUF with a Q8 KV cache, flash attention on, 4096 batch and ubatch, 16 CPU cores, 40 MoE layers on CPU and 99 GPU layers, yielding about 40 t/s generation and 500 t/s prompt processing without MTP. User recalls earlier reports that MTP hurt prompt processing and wants to know if that is still the case and how others configure llama-server.

Sep 18, 2026

Get a weekly email of new NVIDIA RTX 5060 8GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B
NVIDIA RTX 5060 8GB
UD-IQ2_XXS
llama.cpp
8,19230.1 tokens/s
DeepSeek V4.1 Flash 552B (16B active)
NVIDIA RTX 5060 8GB
FP8
ds4
8,1922.4 tokens/s
Qwen3.6 35B (3B active)
NVIDIA RTX 5060 8GB
Q4_K_XL
llama.cpp
Not reported40.0 tokens/s