llamaperf

Ling-3.0

Ant Group · 3 reports

reported speed:
40.9 tokens/s generation
quant:
INT4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

Benchmark of Ling-3.0-flash on DGX Spark with MTP n=1. Reported 40.9 tok/s on short coding task with CUDA graphs and MTP n=1. Also reports prose throughput 38.7 tok/s at 512-token output and 37.3 at 2048-token output. MTP n=2 and n=3 slower for prose. Acceptance lengths reported: n=1: 1.87, n=2: 2.39, n=3: 2.77. Baseline without MTP: 22.9 tok/s with CUDA graphs, 20.8 eager.

Tone: positive
reported speed:
40.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmarked on DGX Spark by sudoingX. Q5_K_M fastest at 40.2 tok/s, Q4_K_M 38.2, Q6_K 32.0. DeepSeek V4 Flash measured 16.5 tok/s on same box.

Tone: positive
reported speed:
38.7 tokens/s generation
quant:
INT4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Community GGUF achieved 35.2 tok/s; DeepSeek V4 Flash comparison at 2.4x speedup. Official INT4 quants initially reported as not running on single Spark, later corrected to be fastest path.