llamaperf

NVIDIA GTX 1080 Ti

NVIDIA · 11GB · 2 reports

Engines people use on it: llama.cpp 2

Run models on your NVIDIA GTX 1080 Ti? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA GTX 1080 Ti

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 11 GB of VRAM.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: positive
reported speed:
52.0 tokens/s generation · 340.9 tokens/s prompt processing
quant:
Q6_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma 4 26B-A4B at 52.03 t/s generation and 340.86 t/s prompt processing on a dual-GPU setup of GTX 1080 Ti and Radeon MI50 16GB, totaling 27GB VRAM. Setup is llama.cpp Vulkan pre-built binary build 851cb34f2 (11055) with the UD-Q6_K_XL GGUF, flash attention on, and 99 GPU layers offloaded. Single-GPU GTX 1080 Ti runs of the same model reached 11.76 t/s generation and 162.12 t/s prompt processing. The user also benchmarked Qwen3.6-35B-A3B MXFP4 MoE, Nemotron 31B-A3.5B Q5_K_M, Qwen3.8 27B Q6_K, and medgemma 27B Q6_K_XL, noting dense models benefited most from the second GPU.

Sep 19, 2026
reported speed:
19.7 tokens/s generation · 332.6 tokens/s prompt processing
quant:
MXFP4 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.6 35B-A3B at 19.7 t/s generation and 332.57 t/s prompt processing on a mixed GTX 1080 Ti and Radeon VII Vulkan setup. Setup is llama.cpp Vulkan build b29c606e2 with the MXFP4 MoE GGUF and Flash Attention enabled, across two GPUs. User also reports results for twelve other models including llama 7B, bailingmoe2 16B-A1B, gpt-oss 20B, Gemma 4 26B-A4B, Qwen3.5 27B, Qwen3-Coder 30B-A3B, granite 4.0, and Phi-3.5-MoE.

Sep 18, 2026

Get a weekly email of new NVIDIA GTX 1080 Ti reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Gemma 4 26B (4B active)
2× NVIDIA GTX 1080 Ti
Q6_K_XL
llama.cpp
Not reported52.0 tokens/s
Qwen3.6 35B (3B active)
2× NVIDIA GTX 1080 Ti
MXFP4
llama.cpp
Not reported19.7 tokens/s