llamaperf

H200

NVIDIA · 141GB · 3 reports

See what fits on this GPU →

Use the calculator to check which models fit in 141 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
Tone: positive
kv:
fp8_e4m3

Deployment of DeepSeek V4 Flash with DSpark speculative decoding on HGX-H200 (4 GPUs, TP=4) via SGLang. Compares marlin and flashinfer MoE backends. DSpark is faster than EAGLE: 3.2x at bs=1, +46% throughput at bs=24. Reports TTFT and accept length benchmarks.

Tone: positive
reported speed:
4.8 tokens/s generation
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

vision

Trained a 40.1M connector to give DeepSeek V4 Flash basic vision. Model loaded across four B200s in a custom SGLang stack. Training on 5x H200. Also trained a Laguna XS 2.1 version (33B total / 3B active) with 30.7M connector. Inspired by Baseten's GLM-5.2 Vision NVFP4.

User pretrained a 500M parameter LLM and 330M image generator from scratch using 8xH200 from modal.com. Total cost $800. Includes GGUF weights. Also mentions pretraining a 1B model next.