llamaperf

D700 12GB

AMD · 12GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 12 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.5 9B

D700 12GB · llama.cpp · 70,000 ctx

Tone: positive
reported speed:
11.0 tokens/s generation
quant:
Q4 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

Also tested Qwen 2.5 coder q4 at 22 t/s. User compares Qwen 3.5 favorably to Claude Sonnet 4.6 for planning tasks.