llamaperf

RTX 3090 Ti

NVIDIA · 24GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 24 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.6 27B

RTX 3090 Ti · llama.cpp · 196,608 ctx

Tone: positive
reported speed:
100.0 tokens/s generation
quant:
Q8_0 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports tensor split-mode raising throughput from 70+ t/s to 100+ t/s, with a peak of 130 t/s. Power draw is 750W+.

Jun 23, 2026