Ling-3.0 flash
DGX Spark · vLLM
- reported speed:
- 40.9 tokens/s generation
- quant:
- INT4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Benchmark of Ling-3.0-flash on DGX Spark with MTP n=1. Reported 40.9 tok/s on short coding task with CUDA graphs and MTP n=1. Also reports prose throughput 38.7 tok/s at 512-token output and 37.3 at 2048-token output. MTP n=2 and n=3 slower for prose. Acceptance lengths reported: n=1: 1.87, n=2: 2.39, n=3: 2.77. Baseline without MTP: 22.9 tok/s with CUDA graphs, 20.8 eager.