llamaperf

NVIDIA B300 288GB

NVIDIA · 288GB · 3 reports

Engines people use on it: vLLM 1

Run models on your NVIDIA B300 288GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA B300 288GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 288 GB of VRAM.

reported speed:
280.0 tokens/s generation
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Victoria, a fine-tune of Qwen3.8-Flash-Next with 44% of experts cut via REAP, at 280 tok/s single stream on one B300 with the draft head, versus 135 without it. Setup is NVFP4 weights retrained at 4-bit, 48.0 GiB of weights including the draft head, with a separate 95.4 GiB n-gram table not counted in that number. Terminal-Bench 2.1 scored 70.0% averaged over 3 runs with an 8h per-task timeout, versus 62.5% for the previous NVFP4 build; HumanEval 159/164. The GGUF Q4_K_M build is 49.17 GiB and scored 75.3% on Terminal-Bench in a single noisy run and 93.2% on HumanEval averaged over 5 runs. Uses 35% fewer output tokens than the previous build. A second fine-tune, Maple, is a Canada-first model; its figures are not reported here.

Sep 29, 2026
reported speed:
92.0 tokens/s generation
quant:
MXFP4 (MXFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Kimi K3 (2.8T parameters) at 92 tok/s decode on 8x B300 via Modal. Setup is vLLM with MXFP4 weights, tensor parallel 8, cold boot about 27 min for a 1.56 TB load. TTFT is 0.92 to 1.02 s and average decode over 4 prompts is 83 tok/s. Cost is $56.79 per hour, $190 per million output tokens, about $36 per run, or $1,363 a day left warm. User also ran Unsloth's Dynamic GGUF 1-bit UD-IQ1_S (594 GB) on 8x A100-80GB via llama.cpp at about 9 tok/s with TTFT 7 to 60 s, $19.99 per hour and about $620 per million tokens, 3.3x more expensive per token. Quality at 1-bit was fine.

Sep 27, 2026
kv:
fp8
summarization

User reports a single B300 GPU running vLLM 0.25.0 at batch 256 with about 300 output tokens per item. Setup is tensor_parallel_size=1, block_size=256, max_num_seqs=256, enable_prefix_caching=true, moe_backend=flashinfer_trtllm, reasoning_parser=deepseek_v4, attention_config use_fp4_indexer_cache=true, and compilation_config cudagraph_mode=FULL_AND_PIECEWISE. The user reports deep_gemm_mega_moe requires expert parallel on a single GPU, and that disabling DSpark speculative decoding doubled throughput.

Sep 7, 2026

Get a weekly email of new NVIDIA B300 288GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 Victoria
NVIDIA B300 288GB
NVFP4
Engine not reported
Not reported280.0 tokens/s
Kimi K3 2800B (104B active)
8× NVIDIA B300 288GB
MXFP4
vLLM
Not reported92.0 tokens/s