Llama 3 8B
on 2× NVIDIA RTX 3090 · llama.cpp · 65,536 ctx
Oct 8, 2026
Summary
User benchmarks the new tensor parallel (split mode tensor) implementation in llama.cpp against graph parallel in ik_llama.cpp on a 2x RTX 3090 system with 48 GiB of VRAM.
Setup is llama.cpp with Q4_0 quantized Llama 3 8B, f16 KV cache, 65536 context, 100 GPU layers, and flash attention enabled. The run crashed with CUDA out of memory at 34816 tokens of context.
User reports 129.08 t/s generation and 5124.38 t/s prompt processing at zero context, degrading to 26.89 t/s generation and 3092.68 t/s prompt processing at 34816 tokens. User calls the PR a gimmick not ready for prime time, noting memory is not released and performance lags well behind ik_llama.cpp graph parallel.