Skip to content

Instantly share code, notes, and snippets.

@jodoherty
Last active May 16, 2026 03:26
Show Gist options
  • Select an option

  • Save jodoherty/1c92906d1604d864d7cc3aab3d341168 to your computer and use it in GitHub Desktop.

Select an option

Save jodoherty/1c92906d1604d864d7cc3aab3d341168 to your computer and use it in GitHub Desktop.
RTX Pro 6000 Blackwell benchmarks (vLLM)

RTX Pro 6000 Blackwell vLLM Benchmarks

Hardware:

01:00.0 VGA compatible controller: NVIDIA Corporation GB202GL [RTX PRO 6000 Blackwell Workstation Edition] (rev a1)

Summary

Configuration Concurrency Prefill (t/s) TTFT (ms) Gen Speed (tg512) Gen Speed (tg8192)
BF16 c1 20,645 413.7 139.4 t/s 137.9 t/s
c5 6,821 1,425.4 70.4 t/s 74.4 t/s
NVFP4 (Winner) c1 24,584 360.5 166.0 t/s 163.5 t/s
c5 7,813 1,185.8 110.9 t/s 115.5 t/s
MTP Enabled c1 20,651 402.8 138.0 t/s 136.5 t/s
c5 6,047 1,478.7 98.7 t/s 100.5 t/s

google/gemma-4-26B-A4B-it (MTP)

docker run --restart=unless-stopped -d --gpus all \
  --privileged --ipc=host -p 8000:8000 \
  -v hfcache:/root/.cache/huggingface \
  vllm/vllm-openai:gemma4-0505-cu130 google/gemma-4-26B-A4B-it \
  --tensor-parallel-size 1 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.90 \
  --max-model-len auto \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4 \
  --chat-template examples/tool_chat_template_gemma4.jinja \
  --reasoning-parser gemma4 \
  --language-model-only \
  --speculative-config '{"model":"google/gemma-4-26B-A4B-it-assistant","num_speculative_tokens":4}'
model test t/s (total) t/s (req) peak t/s peak t/s (req) ttfr (ms) est_ppt (ms) e2e_ttft (ms)
google/gemma-4-26B-A4B-it pp8192 (c1) 20651.05 ± 1490.71 20651.05 ± 1490.71 402.16 ± 31.01 398.97 ± 31.01 402.82 ± 30.67
google/gemma-4-26B-A4B-it tg512 (c1) 137.97 ± 17.48 137.97 ± 17.48 157.50 ± 25.10 157.50 ± 25.10
google/gemma-4-26B-A4B-it pp8192 (c2) 21388.76 ± 764.84 11341.00 ± 2315.66 743.21 ± 87.71 740.01 ± 87.71 743.32 ± 87.77
google/gemma-4-26B-A4B-it tg512 (c2) 228.50 ± 18.63 125.09 ± 11.92 272.30 ± 18.12 141.90 ± 13.49
google/gemma-4-26B-A4B-it pp8192 (c3) 21553.78 ± 521.25 9058.30 ± 3390.59 987.20 ± 225.38 984.00 ± 225.38 987.20 ± 225.38
google/gemma-4-26B-A4B-it tg512 (c3) 254.15 ± 24.14 98.21 ± 8.69 323.60 ± 25.43 121.47 ± 10.45
google/gemma-4-26B-A4B-it pp8192 (c4) 21364.69 ± 400.81 7201.75 ± 2511.86 1246.79 ± 318.84 1243.59 ± 318.84 1246.79 ± 318.84
google/gemma-4-26B-A4B-it tg512 (c4) 343.30 ± 31.86 106.46 ± 11.13 480.10 ± 13.94 129.12 ± 10.83
google/gemma-4-26B-A4B-it pp8192 (c5) 20922.07 ± 350.58 6466.23 ± 3267.18 1478.73 ± 466.32 1475.54 ± 466.32 1478.73 ± 466.32
google/gemma-4-26B-A4B-it tg512 (c5) 377.84 ± 33.09 98.66 ± 15.14 567.80 ± 33.89 125.60 ± 11.58
google/gemma-4-26B-A4B-it pp8192 (c1) 18976.19 ± 1645.58 18976.19 ± 1645.58 438.39 ± 40.08 435.20 ± 40.08 438.39 ± 40.08
google/gemma-4-26B-A4B-it tg8192 (c1) 136.52 ± 13.71 136.52 ± 13.71 155.10 ± 17.32 155.10 ± 17.32
google/gemma-4-26B-A4B-it pp8192 (c2) 20181.92 ± 854.40 10720.02 ± 2185.10 786.22 ± 93.99 783.03 ± 93.99 786.46 ± 94.14
google/gemma-4-26B-A4B-it tg8192 (c2) 238.64 ± 20.62 128.03 ± 9.55 289.50 ± 14.52 150.60 ± 13.63
google/gemma-4-26B-A4B-it pp8192 (c3) 19700.95 ± 2384.44 8325.95 ± 3021.38 1085.53 ± 308.63 1082.33 ± 308.63 1085.83 ± 308.72
google/gemma-4-26B-A4B-it tg8192 (c3) 250.77 ± 22.39 99.35 ± 12.88 354.10 ± 52.09 137.37 ± 42.51
google/gemma-4-26B-A4B-it pp8192 (c4) 20526.53 ± 183.70 7981.86 ± 4513.84 1243.66 ± 405.69 1240.46 ± 405.69 1243.72 ± 405.68
google/gemma-4-26B-A4B-it tg8192 (c4) 311.20 ± 35.81 106.08 ± 10.07 493.20 ± 26.63 136.93 ± 13.66
google/gemma-4-26B-A4B-it pp8192 (c5) 20288.39 ± 456.89 6047.00 ± 2212.57 1521.55 ± 461.21 1518.36 ± 461.21 1521.55 ± 461.21
google/gemma-4-26B-A4B-it tg8192 (c5) 374.25 ± 17.80 100.47 ± 9.88 593.20 ± 29.22 130.42 ± 9.84

google/gemma-4-26B-A4B-it

docker run --restart=unless-stopped -d --gpus all \
  --privileged --ipc=host -p 8000:8000 \
  -v hfcache:/root/.cache/huggingface \
  vllm/vllm-openai:gemma4-0505-cu130 google/gemma-4-26B-A4B-it \
  --tensor-parallel-size 1 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.90 \
  --max-model-len auto \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4 \
  --chat-template examples/tool_chat_template_gemma4.jinja \
  --reasoning-parser gemma4 \
  --language-model-only
model test t/s (total) t/s (req) peak t/s peak t/s (req) ttfr (ms) est_ppt (ms) e2e_ttft (ms)
google/gemma-4-26B-A4B-it pp8192 (c1) 20645.11 ± 2399.57 20645.11 ± 2399.57 413.72 ± 49.34 402.56 ± 49.34 413.72 ± 49.34
google/gemma-4-26B-A4B-it tg512 (c1) 139.43 ± 1.99 139.43 ± 1.99 148.20 ± 6.90 148.20 ± 6.90
google/gemma-4-26B-A4B-it pp8192 (c2) 22245.63 ± 936.78 11765.39 ± 2078.39 721.08 ± 77.97 709.91 ± 77.97 721.08 ± 77.97
google/gemma-4-26B-A4B-it tg512 (c2) 210.97 ± 11.22 110.06 ± 1.76 230.30 ± 9.27 118.65 ± 8.03
google/gemma-4-26B-A4B-it pp8192 (c3) 22782.29 ± 535.52 9539.13 ± 3708.30 945.28 ± 205.66 934.12 ± 205.66 945.32 ± 205.69
google/gemma-4-26B-A4B-it tg512 (c3) 233.56 ± 8.00 83.51 ± 3.30 268.00 ± 12.89 95.23 ± 11.84
google/gemma-4-26B-A4B-it pp8192 (c4) 22190.73 ± 544.42 7682.54 ± 3236.00 1198.83 ± 314.74 1187.67 ± 314.74 1198.83 ± 314.74
google/gemma-4-26B-A4B-it tg512 (c4) 297.11 ± 12.03 82.33 ± 3.87 364.50 ± 14.73 92.38 ± 4.82
google/gemma-4-26B-A4B-it pp8192 (c5) 21807.74 ± 489.31 6820.89 ± 3605.16 1425.41 ± 453.21 1414.25 ± 453.21 1425.41 ± 453.21
google/gemma-4-26B-A4B-it tg512 (c5) 310.71 ± 10.43 70.40 ± 4.72 390.20 ± 17.75 81.54 ± 4.98
google/gemma-4-26B-A4B-it pp8192 (c1) 19850.77 ± 1940.72 19850.77 ± 1940.72 427.70 ± 39.08 416.53 ± 39.08 427.70 ± 39.08
google/gemma-4-26B-A4B-it tg8192 (c1) 137.90 ± 2.23 137.90 ± 2.23 147.11 ± 5.26 147.11 ± 5.26
google/gemma-4-26B-A4B-it pp8192 (c2) 21657.73 ± 1030.81 11854.73 ± 2735.89 726.03 ± 105.59 714.86 ± 105.59 726.03 ± 105.59
google/gemma-4-26B-A4B-it tg8192 (c2) 196.18 ± 23.62 111.12 ± 5.93 226.00 ± 9.10 119.30 ± 10.79
google/gemma-4-26B-A4B-it pp8192 (c3) 21658.94 ± 555.97 8759.12 ± 2743.27 1003.00 ± 194.03 991.83 ± 194.03 1003.04 ± 194.05
google/gemma-4-26B-A4B-it tg8192 (c3) 217.63 ± 5.51 87.24 ± 6.06 260.90 ± 6.07 108.10 ± 20.11
google/gemma-4-26B-A4B-it pp8192 (c4) 22194.42 ± 390.27 8401.15 ± 4547.35 1169.67 ± 361.38 1158.50 ± 361.38 1169.69 ± 361.40
google/gemma-4-26B-A4B-it tg8192 (c4) 254.20 ± 37.51 84.56 ± 8.01 354.00 ± 2.65 98.67 ± 16.05
google/gemma-4-26B-A4B-it pp8192 (c5) 21860.70 ± 409.96 6941.26 ± 3681.21 1407.52 ± 456.03 1396.35 ± 456.03 1407.52 ± 456.03
google/gemma-4-26B-A4B-it tg8192 (c5) 275.75 ± 44.40 74.41 ± 10.08 380.50 ± 10.28 93.48 ± 18.09

nvidia/Gemma-4-26B-A4B-NVFP4

docker run --restart=unless-stopped -d --gpus all \
  --privileged --ipc=host -p 8000:8000 \
  -v hfcache:/root/.cache/huggingface \
  vllm/vllm-openai:gemma4-0505-cu130 nvidia/Gemma-4-26B-A4B-NVFP4 \
  --tensor-parallel-size 1 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.90 \
  --max-model-len auto \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4 \
  --chat-template examples/tool_chat_template_gemma4.jinja \
  --reasoning-parser gemma4 \
  --language-model-only
model test t/s (total) t/s (req) peak t/s peak t/s (req) ttfr (ms) est_ppt (ms) e2e_ttft (ms)
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c1) 23936.64 ± 2957.70 23936.64 ± 2957.70 360.49 ± 43.29 347.60 ± 43.29 360.49 ± 43.29
nvidia/Gemma-4-26B-A4B-NVFP4 tg512 (c1) 166.01 ± 1.99 166.01 ± 1.99 175.70 ± 7.32 175.70 ± 7.32
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c2) 25871.79 ± 1424.07 14032.42 ± 2395.88 609.20 ± 74.49 596.32 ± 74.49 609.20 ± 74.49
nvidia/Gemma-4-26B-A4B-NVFP4 tg512 (c2) 275.87 ± 10.63 143.25 ± 3.32 303.90 ± 11.94 153.25 ± 5.86
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c3) 27186.01 ± 1070.15 10783.24 ± 2999.72 809.88 ± 143.18 796.99 ± 143.18 809.88 ± 143.18
nvidia/Gemma-4-26B-A4B-NVFP4 tg512 (c3) 332.00 ± 24.37 125.33 ± 6.49 396.70 ± 18.50 137.43 ± 9.79
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c4) 27628.53 ± 912.11 9904.41 ± 4701.95 956.41 ± 267.23 943.53 ± 267.23 956.41 ± 267.23
nvidia/Gemma-4-26B-A4B-NVFP4 tg512 (c4) 426.41 ± 17.52 121.50 ± 7.38 545.10 ± 26.41 136.57 ± 6.62
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c5) 26588.54 ± 456.06 7873.37 ± 3372.24 1185.76 ± 343.01 1172.88 ± 343.01 1185.76 ± 343.01
nvidia/Gemma-4-26B-A4B-NVFP4 tg512 (c5) 474.54 ± 41.77 110.91 ± 9.76 625.70 ± 24.87 127.44 ± 5.55
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c1) 24584.12 ± 2168.29 24584.12 ± 2168.29 349.25 ± 35.30 336.37 ± 35.30 349.25 ± 35.30
nvidia/Gemma-4-26B-A4B-NVFP4 tg8192 (c1) 163.54 ± 2.57 163.54 ± 2.57 173.90 ± 8.13 173.90 ± 8.13
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c2) 24397.78 ± 1610.53 12514.53 ± 896.71 670.93 ± 47.25 658.05 ± 47.25 671.07 ± 47.36
nvidia/Gemma-4-26B-A4B-NVFP4 tg8192 (c2) 248.77 ± 41.98 145.81 ± 4.53 304.60 ± 12.60 156.45 ± 9.28
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c3) 25814.52 ± 1448.95 10186.65 ± 2278.61 851.34 ± 155.92 838.46 ± 155.92 851.34 ± 155.92
nvidia/Gemma-4-26B-A4B-NVFP4 tg8192 (c3) 311.80 ± 21.87 126.85 ± 6.35 388.30 ± 10.89 139.50 ± 10.37
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c4) 25683.49 ± 844.64 9259.50 ± 4382.21 1021.18 ± 289.65 1008.30 ± 289.65 1021.18 ± 289.65
nvidia/Gemma-4-26B-A4B-NVFP4 tg8192 (c4) 370.62 ± 31.54 123.99 ± 7.95 529.30 ± 18.09 138.78 ± 10.88
nvidia/Gemma-4-26B-A4B-NVFP4 pp8192 (c5) 25287.12 ± 752.49 7812.57 ± 3601.66 1222.64 ± 386.49 1209.76 ± 386.49 1222.71 ± 386.53
nvidia/Gemma-4-26B-A4B-NVFP4 tg8192 (c5) 413.20 ± 40.57 115.50 ± 10.45 628.10 ± 22.28 133.46 ± 10.19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment