llamaperf

T4 16GB

NVIDIA · 16GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 16 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 125B (6B active) Flash-Next

T4 16GB · ik_llama.cpp · 262,144 ctx

Tone: positive
reported speed:
17.6 tokens/s generation · 159.6 tokens/s prompt processing
quant:
UD-Q4_K_XL (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8-Flash-Next, 180B total with 6B active across 512 experts, generating 17.6 t/s on short prompts and 16.1 t/s at 12.5K context. Non-expert weights run on a T4 using 4606 MiB, with the experts held in host RAM. Prompt processing reaches 159.6 t/s on a cold 12.5K prompt. User reports solid refactor and coding performance, and finds it less verbose than Opus.

Sep 4, 2026