llamaperf

Llama 3 8B

on 2× NVIDIA RTX 3090 · llama.cpp · 65,536 ctx

Tone: negative
Oct 8, 2026
Throughput
129.1 t/s gen · 5124.4 t/s pp
Quant
Q4_0 (GGUF)
KV cache
f16
Flash Attention
on
VRAM reported
48 GB

Summary

User benchmarks the new tensor parallel (split mode tensor) implementation in llama.cpp against graph parallel in ik_llama.cpp on a 2x RTX 3090 system with 48 GiB of VRAM. Setup is llama.cpp with Q4_0 quantized Llama 3 8B, f16 KV cache, 65536 context, 100 GPU layers, and flash attention enabled. The run crashed with CUDA out of memory at 34816 tokens of context. User reports 129.08 t/s generation and 5124.38 t/s prompt processing at zero context, degrading to 26.89 t/s generation and 3092.68 t/s prompt processing at 34816 tokens. User calls the PR a gimmick not ready for prime time, noting memory is not released and performance lags well behind ik_llama.cpp graph parallel.