Llama 3 8B
2× NVIDIA RTX 3090 · llama.cpp · 65,536 ctx
- reported speed:
- 129.1 tokens/s generation · 5124.4 tokens/s prompt processing
- quant:
- Q4_0 (GGUF)
- kv:
- f16
- flash attention:
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks the new tensor parallel (split mode tensor) implementation in llama.cpp against graph parallel in ik_llama.cpp on a 2x RTX 3090 system with 48 GiB of VRAM. Setup is llama.cpp with Q4_0 quantized Llama 3 8B, f16 KV cache, 65536 context, 100 GPU layers, and flash attention enabled. The run crashed with CUDA out of memory at 34816 tokens of context. User reports 129.08 t/s generation and 5124.38 t/s prompt processing at zero context, degrading to 26.89 t/s generation and 3092.68 t/s prompt processing at 34816 tokens. User calls the PR a gimmick not ready for prime time, noting memory is not released and performance lags well behind ik_llama.cpp graph parallel.