Gemma 4
2× NVIDIA RTX 2000 Ada · TensorSharp
- reported speed:
- 51.7 tokens/s generation · 2488.0 tokens/s prompt processing
- quant:
- Q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks TensorSharp multi-GPU tensor parallelism on 2x RTX 2000 Ada 16GB. Gemma 4 E4B Q8_0 is the primary model among several tested. TP=2 raises decode speed from 37.3 to 51.7 tok/s.