Skip to content

Instantly share code, notes, and snippets.

@vikhyat
Last active August 26, 2026 20:06
Show Gist options
  • Select an option

  • Save vikhyat/793e094a10c5b61e8278205475f6fcd7 to your computer and use it in GitHub Desktop.

Select an option

Save vikhyat/793e094a10c5b61e8278205475f6fcd7 to your computer and use it in GitHub Desktop.
Photon vs vLLM: NVIDIA B200 ChartQA reasoning benchmark

Photon vs vLLM on NVIDIA B200 (final shipping board)

Status: Final 28-cell board against vLLM 0.27.1. The board retains 26 validated, unchanged cells from 7f0dddc2, retains the frozen Gemma E4B C8 measurement from e1d6c1dd, and replaces Gemma E4B C4 with the final fast/correct release measurement from 5c06c090. Every Photon/vLLM pair is internally matched on driver, model revision, harness, ordered request stream, request count, and B200 hardware class; exact engine commits are recorded per cell.

ChartQA multimodal inference with reasoning enabled on NVIDIA B200. The primary metric is natural-output tokens completed per active end-to-end service second; accuracy and output-length differences are reported alongside it.

At a glance

  • Photon leads in 28/28 end-to-end output-token service-throughput cells.
  • Photon leads in 28/28 completed-request-rate cells.
  • Photon leads in 28/28 timing-matched model-prefill cells.
  • Photon leads in 27/28 timing-matched decode cells.
  • The 1 timing-matched decode non-win is Gemma 4 E4B C8 (-2.4%).
  • Pair checks: request stream 28/28, input tokens 28/28, model revision 28/28, engine-side model fingerprint 26/28, driver 28/28, exact GPU UUID 27/28, matching harness hash 28/28.
  • 6 comparable cells cross the report warning threshold of 2 accuracy points or 10% mean-output-length difference; 6 retained Gemma rows use the superseded parser and are excluded from accuracy comparisons.
  • 24 cells use 128 requests per arm; Qwen 0.8B/2B/4B/9B C8 use 256 to satisfy the declared-concurrency floor.
  • The Kestrel 0.6.1 release tree passed its full CPU suite (839 passed, 46 hardware-dependent skips); focused query, streaming, and scheduler regressions also passed.

Output-token service throughput

Output-token service throughput

Photon service-throughput delta

Model C Photon output tok/s vLLM output tok/s Service delta Req/s delta Decode delta Accuracy P / V
Moondream 3 1 300.6 117.6 +155.6% +140.6% +59.0% 85.2% / 83.6%
Moondream 3 2 453.6 214.1 +111.8% +108.4% +61.8% 85.2% / 83.6%
Moondream 3 4 662.5 353.1 +87.6% +82.1% +57.3% 84.4% / 83.6%
Moondream 3 8 906.4 458.8 +97.6% +92.4% +60.3% 84.4% / 84.4%
Qwen3.5 0.8B 1 1,189.4 622.5 +91.1% +118.6% +83.2% 75.0% / 71.9%
Qwen3.5 0.8B 2 2,016.9 1,188.5 +69.7% +71.7% +64.0% 71.9% / 72.7%
Qwen3.5 0.8B 4 3,230.1 2,167.1 +49.0% +56.4% +39.7% 76.6% / 73.4%
Qwen3.5 0.8B 8 4,875.7 3,997.4 +22.0% +20.6% +15.6% 73.0% / 73.0%
Qwen3.5 2B 1 822.7 520.0 +58.2% +56.4% +50.3% 75.0% / 76.6%
Qwen3.5 2B 2 1,462.4 953.4 +53.4% +52.1% +49.6% 74.2% / 75.0%
Qwen3.5 2B 4 2,430.3 1,769.4 +37.3% +37.3% +29.2% 74.2% / 75.8%
Qwen3.5 2B 8 3,859.4 3,284.8 +17.5% +25.6% +14.1% 75.8% / 75.4%
Qwen3.5 4B 1 458.2 314.3 +45.8% +44.2% +40.6% 72.7% / 75.8%
Qwen3.5 4B 2 800.9 621.4 +28.9% +27.8% +25.8% 75.0% / 75.0%
Qwen3.5 4B 4 1,371.8 1,256.4 +9.2% +12.2% +6.3% 75.0% / 75.0%
Qwen3.5 4B 8 2,346.1 2,153.3 +8.9% +1.5% +5.7% 74.2% / 78.1%
Qwen3.5 9B 1 301.8 225.2 +34.0% +39.4% +30.0% 79.7% / 77.3%
Qwen3.5 9B 2 564.0 441.8 +27.7% +29.3% +25.2% 80.5% / 78.9%
Qwen3.5 9B 4 1,008.4 822.5 +22.6% +21.2% +18.4% 77.3% / 78.9%
Qwen3.5 9B 8 1,775.8 1,600.7 +10.9% +16.3% +9.9% 78.5% / 77.0%
Gemma 4 E2B 1 431.2 329.9 +30.7% +33.3% +25.9% n/a (legacy parser)
Gemma 4 E2B 2 796.6 599.5 +32.9% +34.6% +29.0% n/a (legacy parser)
Gemma 4 E2B 4 1,403.6 1,111.9 +26.2% +24.3% +22.8% n/a (legacy parser)
Gemma 4 E2B 8 2,260.1 2,018.2 +12.0% +11.9% +6.3% n/a (legacy parser)
Gemma 4 E4B 1 290.2 227.2 +27.7% +28.2% +22.8% n/a (legacy parser)
Gemma 4 E4B 2 498.2 411.4 +21.1% +21.4% +17.6% n/a (legacy parser)
Gemma 4 E4B 4 905.1 868.1 +4.3% +6.5% +0.9% 82.8% / 79.7%
Gemma 4 E4B 8 1,590.4 1,560.3 +1.9% +4.6% -2.4% 82.8% / 83.6%

Model-prefill throughput

Model-prefill tok/s is the timing-matched model prefill rate recorded by the service run. Prefix caching is disabled.

Model-prefill throughput

Model-prefill speedup

Model C Photon prefill tok/s vLLM prefill tok/s Photon vs vLLM
Moondream 3 1 61,406.9 19,597.6 +213.3%
Moondream 3 2 46,835.7 22,150.2 +111.4%
Moondream 3 4 39,525.6 19,838.4 +99.2%
Moondream 3 8 30,384.0 14,310.9 +112.3%
Qwen3.5 0.8B 1 28,463.4 11,714.8 +143.0%
Qwen3.5 0.8B 2 28,176.1 19,057.9 +47.8%
Qwen3.5 0.8B 4 25,433.4 18,375.3 +38.4%
Qwen3.5 0.8B 8 28,665.2 19,884.5 +44.2%
Qwen3.5 2B 1 24,443.3 10,526.0 +132.2%
Qwen3.5 2B 2 23,639.6 17,126.0 +38.0%
Qwen3.5 2B 4 23,502.8 15,089.8 +55.8%
Qwen3.5 2B 8 25,487.6 18,308.3 +39.2%
Qwen3.5 4B 1 27,866.2 7,162.8 +289.0%
Qwen3.5 4B 2 26,453.5 11,787.5 +124.4%
Qwen3.5 4B 4 25,668.6 12,061.6 +112.8%
Qwen3.5 4B 8 28,420.4 14,623.6 +94.3%
Qwen3.5 9B 1 23,637.0 6,387.9 +270.0%
Qwen3.5 9B 2 21,709.0 9,937.8 +118.4%
Qwen3.5 9B 4 20,212.0 9,425.1 +114.4%
Qwen3.5 9B 8 22,242.1 11,875.0 +87.3%
Gemma 4 E2B 1 19,265.4 9,171.4 +110.1%
Gemma 4 E2B 2 17,538.0 9,620.5 +82.3%
Gemma 4 E2B 4 17,142.2 9,435.6 +81.7%
Gemma 4 E2B 8 15,915.6 7,852.4 +102.7%
Gemma 4 E4B 1 17,974.6 7,946.8 +126.2%
Gemma 4 E4B 2 14,962.9 8,604.8 +73.9%
Gemma 4 E4B 4 14,530.8 8,538.6 +70.2%
Gemma 4 E4B 8 13,082.4 7,563.7 +73.0%

Serving trade-off

The horizontal axis is aggregate end-to-end output tok/s and includes prefill. The vertical axis is decode tok/s per concurrent user and excludes prefill. Higher and farther right are better, but the axes are not commensurable.

Serving trade-off

Latency and jitter

Model C Photon p50 / p95 / p99 ms vLLM p50 / p95 / p99 ms Photon p95-p50 vLLM p95-p50
Moondream 3 1 51.4 / 89.2 / 124.0 132.6 / 187.8 / 228.8 37.8 55.2
Moondream 3 2 66.9 / 113.0 / 135.0 141.5 / 232.4 / 287.3 46.0 90.9
Moondream 3 4 92.0 / 168.0 / 226.6 172.3 / 270.7 / 399.0 76.1 98.4
Moondream 3 8 133.5 / 246.5 / 279.6 238.6 / 696.3 / 876.6 113.0 457.7
Qwen3.5 0.8B 1 205.5 / 1247.0 / 1267.1 413.9 / 2304.8 / 2323.5 1041.5 1890.9
Qwen3.5 0.8B 2 247.2 / 1485.2 / 1516.8 440.7 / 2425.1 / 2469.1 1238.0 1984.4
Qwen3.5 0.8B 4 287.7 / 1818.6 / 1938.5 469.5 / 2563.7 / 2671.2 1530.9 2094.2
Qwen3.5 0.8B 8 422.3 / 2420.2 / 2478.9 523.8 / 2837.8 / 2924.7 1997.9 2314.0
Qwen3.5 2B 1 265.7 / 1820.5 / 1840.2 490.1 / 2704.8 / 2794.6 1554.8 2214.7
Qwen3.5 2B 2 314.4 / 2047.2 / 2106.2 555.6 / 3042.3 / 3094.9 1732.8 2486.6
Qwen3.5 2B 4 367.0 / 2457.0 / 2568.1 575.2 / 3187.0 / 3289.6 2090.0 2611.7
Qwen3.5 2B 8 449.2 / 3018.8 / 3200.5 587.7 / 3514.8 / 3613.7 2569.6 2927.1
Qwen3.5 4B 1 540.4 / 3352.0 / 3357.3 799.6 / 4715.6 / 4720.8 2811.6 3916.0
Qwen3.5 4B 2 617.0 / 3821.3 / 3849.4 847.5 / 4811.1 / 4857.2 3204.3 3963.6
Qwen3.5 4B 4 662.1 / 4365.0 / 4433.0 809.6 / 4670.8 / 4721.4 3702.9 3861.1
Qwen3.5 4B 8 814.9 / 5202.3 / 5246.1 899.1 / 5459.3 / 5531.9 4387.5 4560.2
Qwen3.5 9B 1 809.1 / 5066.7 / 5100.3 1203.1 / 6604.3 / 6631.2 4257.7 5401.2
Qwen3.5 9B 2 847.0 / 5368.7 / 5459.6 1160.6 / 6770.4 / 6865.8 4521.7 5609.8
Qwen3.5 9B 4 939.9 / 5959.6 / 6008.5 1174.8 / 7028.2 / 7166.4 5019.7 5853.4
Qwen3.5 9B 8 1130.4 / 6709.7 / 6863.3 1267.6 / 7359.0 / 7531.6 5579.2 6091.4
Gemma 4 E2B 1 736.1 / 1479.3 / 2250.9 982.7 / 2418.9 / 3476.2 743.2 1436.2
Gemma 4 E2B 2 784.0 / 1692.0 / 2487.6 1085.5 / 2261.4 / 2901.6 908.0 1175.9
Gemma 4 E2B 4 925.7 / 2120.9 / 2767.1 1198.2 / 2551.8 / 3403.0 1195.2 1353.6
Gemma 4 E2B 8 1093.7 / 2610.9 / 3373.0 1239.6 / 2717.3 / 3566.8 1517.2 1477.6
Gemma 4 E4B 1 903.2 / 1862.5 / 3144.3 1141.3 / 2321.2 / 5178.8 959.3 1179.9
Gemma 4 E4B 2 1052.7 / 1960.1 / 2549.0 1290.1 / 2371.7 / 4416.5 907.4 1081.6
Gemma 4 E4B 4 1145.1 / 2420.6 / 4194.0 1228.9 / 2477.2 / 4336.5 1275.5 1248.3
Gemma 4 E4B 8 1301.5 / 2573.6 / 3244.7 1322.1 / 2629.2 / 4310.0 1272.1 1307.1

Quality audit

Accuracy and generation length are measured outcomes. Six retained Gemma r40 rows used the superseded answer parser and are excluded instead of publishing their invalid scores. A warning means an absolute accuracy gap of at least 2 percentage points or a mean-output-length gap of at least 10%.

Model C Photon correct vLLM correct Accuracy delta Mean output P / V Assessment
Moondream 3 1 109/128 107/128 +1.56 pp 17.1 / 16.1 within threshold
Moondream 3 2 109/128 107/128 +1.56 pp 16.5 / 16.3 within threshold
Moondream 3 4 108/128 107/128 +0.78 pp 16.9 / 16.4 within threshold
Moondream 3 8 108/128 108/128 +0.00 pp 16.8 / 16.4 within threshold
Qwen3.5 0.8B 1 96/128 92/128 +3.12 pp 378.1 / 432.5 warning
Qwen3.5 0.8B 2 92/128 93/128 -0.78 pp 429.6 / 434.7 within threshold
Qwen3.5 0.8B 4 98/128 94/128 +3.12 pp 412.6 / 432.8 warning
Qwen3.5 0.8B 8 187/256 187/256 +0.00 pp 429.7 / 424.8 within threshold
Qwen3.5 2B 1 96/128 98/128 -1.56 pp 419.4 / 414.7 within threshold
Qwen3.5 2B 2 95/128 96/128 -0.78 pp 423.3 / 419.8 within threshold
Qwen3.5 2B 4 95/128 97/128 -1.56 pp 425.5 / 425.2 within threshold
Qwen3.5 2B 8 194/256 193/256 +0.39 pp 413.6 / 442.1 within threshold
Qwen3.5 4B 1 93/128 97/128 -3.12 pp 493.6 / 488.3 warning
Qwen3.5 4B 2 96/128 96/128 +0.00 pp 492.1 / 488.1 within threshold
Qwen3.5 4B 4 96/128 96/128 +0.00 pp 480.7 / 494.0 within threshold
Qwen3.5 4B 8 190/256 200/256 -3.91 pp 494.6 / 460.7 warning
Qwen3.5 9B 1 102/128 99/128 +2.34 pp 417.6 / 434.4 warning
Qwen3.5 9B 2 103/128 101/128 +1.56 pp 420.4 / 425.8 within threshold
Qwen3.5 9B 4 99/128 101/128 -1.56 pp 429.1 / 424.1 within threshold
Qwen3.5 9B 8 201/256 197/256 +1.56 pp 434.8 / 455.9 within threshold
Gemma 4 E4B 4 106/128 102/128 +3.12 pp 277.0 / 282.8 warning
Gemma 4 E4B 8 106/128 107/128 -0.78 pp 269.8 / 277.0 within threshold

These are measured output differences, not token-normalized quality comparisons. The artifact ledger preserves the exact natural outputs and reports every threshold crossing.

GPU operating characteristics at C1

Model Backend Peak process VRAM Power mean / p95 Temperature p95 Incremental J/output token
Moondream 3 Photon 18.90 GiB 451.6 / 496.6 W 43 C 0.813
Moondream 3 vLLM 161.46 GiB 345.2 / 359.7 W 39 C 1.145
Qwen3.5 0.8B Photon 6.64 GiB 500.1 / 518.0 W 45 C 0.249
Qwen3.5 0.8B vLLM 159.92 GiB 341.6 / 354.8 W 41 C 0.221
Qwen3.5 2B Photon 15.09 GiB 665.8 / 687.4 W 53 C 0.568
Qwen3.5 2B vLLM 159.66 GiB 441.2 / 473.1 W 45 C 0.468
Qwen3.5 4B Photon 25.05 GiB 807.4 / 822.4 W 57 C 1.283
Qwen3.5 4B vLLM 159.69 GiB 559.7 / 583.5 W 52 C 1.096
Qwen3.5 9B Photon 40.08 GiB 929.6 / 948.4 W 60 C 2.370
Qwen3.5 9B vLLM 159.57 GiB 665.5 / 691.6 W 56 C 2.025
Gemma 4 E2B Photon 16.69 GiB 500.3 / 506.6 W 46 C 0.707
Gemma 4 E2B vLLM 158.54 GiB 387.9 / 400.0 W 42 C 0.593
Gemma 4 E4B Photon 28.87 GiB 618.5 / 627.9 W 50 C 1.434
Gemma 4 E4B vLLM 158.62 GiB 481.9 / 503.1 W 44 C 1.209

Methodology

  • NVIDIA B200, tensor parallelism 1, ChartQA test split, greedy decoding, reasoning enabled, no prefix cache.
  • 24 cells contain 128 requests per arm. Qwen 0.8B/2B/4B/9B C8 contain 256 requests per arm to remove finite-window boundary underfill.
  • The final board uses 26 unchanged validated rows from runtime 7f0dddc232ea9a3406108bbaea96edfda62ec0be plus public runtime f49bd3812197132fe9bd60f27363f4c4c140adcb, Gemma E4B C8 from runtime e1d6c1dd7fd094303c6cf7d18fd867b03f53016b plus parser 3a85db8fc74c87701f0b9edc97621338fc896efc, and Gemma E4B C4 from runtime 5c06c09047d830d8008c16e9eb275a9f55bf5714 plus parser b4c20e2cb04b4839a0cc063e636d009be704d272. The baseline is vLLM 0.27.1 at e3acdf4951dabbe2369fa03dc5b90e1da1408411.
  • Gemma E4B C8 is the predeclared median of three exact repeats; the selected run is the median, not the best run. Its Photon repeat and retained vLLM baseline use separate B200s on the same host; all other pairs match exact GPU UUID.
  • Every pair uses a byte-identical ordered request stream and the same model revision, driver, harness, and B200 hardware class.
  • Photon is in-process; vLLM is served out-of-process over HTTP. The chart reports each engine in its native deployment shape.
  • Service throughput is completed natural-output tokens divided by the closed-loop active service wall. Request rate uses completed requests over the same wall. Latency is client-observed end-to-end time.
  • Model-prefill throughput is the timing-matched model prefill rate from the natural-output service run; prefix caching is disabled.
  • Decode tok/s counts tokens after the first token over first-token-to-completion time. The first token is produced by prefill.

Validation findings and release status

  1. Final board: 26 unchanged validated rows and the frozen Gemma E4B C8 row are retained; Gemma E4B C4 is replaced by the final fast/correct release measurement.
  2. Pair fairness: the report rejects any pair whose B200 hardware class, driver, model revision, harness, ordered request stream, or request count differs; exact GPU UUID is also reported.
  3. Natural outputs: accuracy, token counts, and output-length gaps are reported directly for both engines; warning cells are not silently normalized away.

C4 release source trees: runtime ce9b25ce831a1e88da87bf13a464716da5245b51, public 30e07fd1ccdd4c9d4dbcf400084bf94dd0ad6434; B200-installed CPython 3.12 x86_64 kernels wheel SHA-256 38ecc42ad852ae9d356ba4680d3777bc04810faa1d73866fbf2f632b7f63c5ff; companion A/B bundle SHA-256 396d2698564f29b65385a0bafaf141df6c87c73fe60e144c5a0e6e219e66cafd / 6f7f3ee9e83507b1b6f75c6f4ecbed0a929de4836f16da4b0c5c8e6b9f079a59. Retained C8 source tree: aa8e173b26a1988d52f1bdc5bdfc5be2e1ee652c.

Artifact inventory

  • report-data.json: normalized values and pair-level provenance checks used by this report.
  • SHA256SUMS: hashes for the report, data ledger, generator, and charts.
  • Private raw manifests, requests, telemetry, runtime-batch traces, and cold-start records remain in the retained artifact tree.
#!/usr/bin/env python3
"""Generate the final B200 presales report from retained paired summaries."""
from __future__ import annotations
import hashlib
import json
import os
from collections import Counter
from pathlib import Path
SCRIPT_DIR = Path(__file__).resolve().parent
MATPLOTLIB_CACHE = SCRIPT_DIR.parent / ".matplotlib"
MATPLOTLIB_CACHE.mkdir(parents=True, exist_ok=True)
os.environ.setdefault("MPLCONFIGDIR", str(MATPLOTLIB_CACHE))
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import numpy as np
OUTPUT = SCRIPT_DIR
INPUT = OUTPUT.parent / "collected"
BOARD_MAP = OUTPUT.parent / "28-row-recompute.json"
CONCURRENCIES = (1, 2, 4, 8)
BACKENDS = ("photon", "vllm")
MODELS = {
"md3": "Moondream 3",
"qwen08": "Qwen3.5 0.8B",
"qwen2": "Qwen3.5 2B",
"qwen4": "Qwen3.5 4B",
"qwen9": "Qwen3.5 9B",
"gemmae2": "Gemma 4 E2B",
"gemmae4": "Gemma 4 E4B",
}
SHORT_MODELS = {
"md3": "MD3",
"qwen08": "Qwen 0.8B",
"qwen2": "Qwen 2B",
"qwen4": "Qwen 4B",
"qwen9": "Qwen 9B",
"gemmae2": "Gemma E2B",
"gemmae4": "Gemma E4B",
}
PHOTON = "#007C83"
VLLM = "#C84A5A"
GRID = "#D7DDE2"
TEXT = "#1F2933"
C8_SOURCE_COMMIT = "e1d6c1dd7fd094303c6cf7d18fd867b03f53016b"
C8_SOURCE_TREE = "aa8e173b26a1988d52f1bdc5bdfc5be2e1ee652c"
C8_PUBLIC_COMMIT = "3a85db8fc74c87701f0b9edc97621338fc896efc"
C4_SOURCE_COMMIT = "5c06c09047d830d8008c16e9eb275a9f55bf5714"
C4_SOURCE_TREE = "ce9b25ce831a1e88da87bf13a464716da5245b51"
C4_PUBLIC_COMMIT = "b4c20e2cb04b4839a0cc063e636d009be704d272"
C4_PUBLIC_TREE = "30e07fd1ccdd4c9d4dbcf400084bf94dd0ad6434"
VALIDATED_SOURCE_COMMIT = "7f0dddc232ea9a3406108bbaea96edfda62ec0be"
VALIDATED_PUBLIC_COMMIT = "f49bd3812197132fe9bd60f27363f4c4c140adcb"
BOARD_MAP_SHA256 = "d18425b0b21fdf96d3a1a041d14504a8694e7aa1cbcf3a956b2c85cc20aca1ee"
WHEEL_SHA256 = "38ecc42ad852ae9d356ba4680d3777bc04810faa1d73866fbf2f632b7f63c5ff"
BUNDLE_A_SHA256 = "396d2698564f29b65385a0bafaf141df6c87c73fe60e144c5a0e6e219e66cafd"
BUNDLE_B_SHA256 = "6f7f3ee9e83507b1b6f75c6f4ecbed0a929de4836f16da4b0c5c8e6b9f079a59"
VLLM_VERSION = "0.27.1"
VLLM_COMMIT = "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
BOARD_STATUS = "final_2026_08_26_shipping_board"
GIST_RAW = "https://gist.githubusercontent.com/vikhyat/793e094a10c5b61e8278205475f6fcd7/raw"
REQUEST_COUNTS = {
(tag, 8): 256 for tag in ("qwen08", "qwen2", "qwen4", "qwen9")
}
EXPECTED_SOURCE_COHORTS = {
(VALIDATED_SOURCE_COMMIT, VALIDATED_PUBLIC_COMMIT): 26,
(C8_SOURCE_COMMIT, C8_PUBLIC_COMMIT): 1,
(C4_SOURCE_COMMIT, C4_PUBLIC_COMMIT): 1,
}
EXPECTED_CELL_SOURCES = {
("gemmae4", 4): (C4_SOURCE_COMMIT, C4_PUBLIC_COMMIT),
("gemmae4", 8): (C8_SOURCE_COMMIT, C8_PUBLIC_COMMIT),
}
def read_json(path: Path) -> dict:
with path.open() as handle:
return json.load(handle)
def percent(delta: float, *, bold_win: bool = True) -> str:
text = f"{delta:+.1f}%"
return f"**{text}**" if bold_win and delta > 0 else text
def rate_delta(photon: float, vllm: float) -> float:
return (photon / vllm - 1.0) * 100.0
def gib(value: float) -> float:
return value / (1024.0**3)
def table(headers: list[str], rows: list[list[str]]) -> str:
lines = [
"| " + " | ".join(headers) + " |",
"|" + "|".join("---:" if index else "---" for index in range(len(headers))) + "|",
]
lines.extend("| " + " | ".join(row) + " |" for row in rows)
return "\n".join(lines)
def load_cells() -> list[dict]:
cells: list[dict] = []
for tag, model in MODELS.items():
for concurrency in CONCURRENCIES:
input_root = INPUT / f"{tag}-c{concurrency}"
arms = {}
for backend in BACKENDS:
arm = input_root / "cells" / tag / f"c{concurrency}" / backend
summary = read_json(arm / "summary.json")
manifest = read_json(arm / "manifest.json")
placement_path = arm / "placement.json"
arms[backend] = {
"summary": summary,
"manifest": manifest,
"placement": read_json(placement_path) if placement_path.exists() else None,
}
p_summary = arms["photon"]["summary"]
v_summary = arms["vllm"]["summary"]
p_manifest = arms["photon"]["manifest"]
v_manifest = arms["vllm"]["manifest"]
p_git = p_manifest["source"]["git"]
v_git = v_manifest["source"]["git"]
p_commits = tuple(entry["commit"] for entry in p_git)
v_commits = tuple(entry["commit"] for entry in v_git)
p_req = p_summary["requests"]
v_req = v_summary["requests"]
p_prefill = p_summary["model_prefill"]
v_prefill = v_summary["model_prefill"]
cells.append(
{
"tag": tag,
"model": model,
"concurrency": concurrency,
"photon": arms["photon"],
"vllm": arms["vllm"],
"request_delta_pct": rate_delta(
p_req["requests_per_second"], v_req["requests_per_second"]
),
"output_delta_pct": rate_delta(
p_req["output_tokens_per_service_second"],
v_req["output_tokens_per_service_second"],
),
"decode_delta_pct": rate_delta(
p_req["decode_tokens_per_second"],
v_req["decode_tokens_per_second"],
),
"prefill_delta_pct": rate_delta(
p_prefill["model_prefill_tokens_per_second"],
v_prefill["model_prefill_tokens_per_second"],
),
"accuracy_delta_pp": (p_req["accuracy"] - v_req["accuracy"]) * 100.0,
"accuracy_comparable": not (
tag in {"gemmae2", "gemmae4"}
and p_commits[1]
not in {C8_PUBLIC_COMMIT, C4_PUBLIC_COMMIT}
),
"output_mean_gap_pct": abs(
p_req["output_tokens"]["mean"] - v_req["output_tokens"]["mean"]
)
/ max(
p_req["output_tokens"]["mean"],
v_req["output_tokens"]["mean"],
1.0,
)
* 100.0,
"same_gpu_uuid": p_manifest["gpu"]["gpu_uuid"]
== v_manifest["gpu"]["gpu_uuid"],
"same_gpu_model": (
p_manifest["gpu"]["gpu_name"],
p_manifest["gpu"]["memory_total_bytes"],
)
== (
v_manifest["gpu"]["gpu_name"],
v_manifest["gpu"]["memory_total_bytes"],
),
"same_driver": p_manifest["gpu"]["driver_version"]
== v_manifest["gpu"]["driver_version"],
"same_request_stream": p_manifest["workload"]["request_stream"]["sha256"]
== v_manifest["workload"]["request_stream"]["sha256"],
"same_harness": p_manifest["source"]["harness_sha256"]
== v_manifest["source"]["harness_sha256"],
"same_model_revision": p_manifest["model"]["revision"]
== v_manifest["model"]["revision"],
"same_model_fingerprint": p_manifest["model"]["fingerprint"]
== v_manifest["model"]["fingerprint"],
"same_input_tokens": p_req["total_input_tokens"]
== v_req["total_input_tokens"],
"request_count": p_req["request_count"],
"same_request_count": p_req["request_count"] == v_req["request_count"],
"source_commit": p_commits[0],
"public_commit": p_commits[1],
"vllm_commit": v_commits[0],
"source_commit_exact": all(
entry["tracked_dirty"] is False for entry in p_git + v_git
),
}
)
return cells
def arm_metric(cell: dict, backend: str, section: str, key: str) -> float:
return cell[backend]["summary"][section][key]
def style_axes(axis, title: str) -> None:
axis.set_title(title, fontsize=11, color=TEXT, pad=8)
axis.grid(True, color=GRID, linewidth=0.7, alpha=0.8)
axis.spines[["top", "right"]].set_visible(False)
axis.tick_params(colors=TEXT, labelsize=8)
def save_figure(figure, name: str) -> None:
figure.tight_layout(pad=2.0)
figure.savefig(OUTPUT / name, dpi=180, bbox_inches="tight", facecolor="white")
plt.close(figure)
def line_chart(cells: list[dict], section: str, key: str, ylabel: str, output: str) -> None:
figure, axes = plt.subplots(2, 4, figsize=(15, 7.7))
for axis, (tag, name) in zip(axes.flat, MODELS.items(), strict=False):
selected = [cell for cell in cells if cell["tag"] == tag]
for backend, color, label in (
("photon", PHOTON, "Photon"),
("vllm", VLLM, "vLLM"),
):
values = [arm_metric(cell, backend, section, key) for cell in selected]
axis.plot(
CONCURRENCIES,
values,
marker="o",
linewidth=2.2,
markersize=5,
color=color,
label=label,
)
style_axes(axis, name)
axis.set_xticks(CONCURRENCIES)
axis.set_xlabel("Concurrency", fontsize=8)
axis.set_ylabel(ylabel, fontsize=8)
axes.flat[-1].axis("off")
handles, labels = axes.flat[0].get_legend_handles_labels()
figure.legend(handles, labels, loc="lower right", bbox_to_anchor=(0.91, 0.11), frameon=False)
save_figure(figure, output)
def speedup_chart(cells: list[dict], field: str, title: str, output: str) -> None:
values = np.array(
[[next(cell[field] for cell in cells if cell["tag"] == tag and cell["concurrency"] == c) for c in CONCURRENCIES] for tag in MODELS]
)
bound = max(20.0, float(np.ceil(np.max(np.abs(values)) / 10.0) * 10.0))
figure, axis = plt.subplots(figsize=(9.2, 5.3))
image = axis.imshow(values, cmap="RdYlGn", vmin=-bound, vmax=bound, aspect="auto")
axis.set_xticks(range(len(CONCURRENCIES)), [f"C{c}" for c in CONCURRENCIES])
axis.set_yticks(range(len(MODELS)), [SHORT_MODELS[tag] for tag in MODELS])
axis.set_title(title, fontsize=13, color=TEXT, pad=12)
for row in range(values.shape[0]):
for column in range(values.shape[1]):
value = values[row, column]
axis.text(column, row, f"{value:+.1f}%", ha="center", va="center", fontsize=9, color="#111111")
colorbar = figure.colorbar(image, ax=axis, shrink=0.82)
colorbar.set_label("Photon vs vLLM", color=TEXT)
save_figure(figure, output)
def tradeoff_chart(cells: list[dict]) -> None:
figure, axes = plt.subplots(2, 4, figsize=(15, 7.7))
for axis, (tag, name) in zip(axes.flat, MODELS.items(), strict=False):
selected = [cell for cell in cells if cell["tag"] == tag]
for backend, color, label in (
("photon", PHOTON, "Photon"),
("vllm", VLLM, "vLLM"),
):
label_offset = (4, 5) if backend == "photon" else (4, -5)
label_va = "bottom" if backend == "photon" else "top"
x_values = [
arm_metric(cell, backend, "requests", "output_tokens_per_service_second")
for cell in selected
]
y_values = [
arm_metric(cell, backend, "requests", "decode_tokens_per_second")
/ cell["concurrency"]
for cell in selected
]
axis.plot(x_values, y_values, marker="o", linewidth=2.0, markersize=5, color=color, label=label)
for x_value, y_value, concurrency in zip(x_values, y_values, CONCURRENCIES, strict=True):
axis.annotate(
f"C{concurrency}",
(x_value, y_value),
xytext=label_offset,
textcoords="offset points",
fontsize=7,
color=color,
va=label_va,
)
style_axes(axis, name)
axis.set_xlabel("Aggregate output tok/s", fontsize=8)
axis.set_ylabel("Decode tok/s/user", fontsize=8)
axes.flat[-1].axis("off")
handles, labels = axes.flat[0].get_legend_handles_labels()
figure.legend(handles, labels, loc="lower right", bbox_to_anchor=(0.91, 0.11), frameon=False)
save_figure(figure, "decode-scaling.png")
def serializable_cell(cell: dict) -> dict:
result = {
key: value
for key, value in cell.items()
if key not in {"photon", "vllm"}
}
for backend in BACKENDS:
summary = cell[backend]["summary"]
manifest = cell[backend]["manifest"]
requests = summary["requests"]
telemetry = summary["telemetry"]
result[backend] = {
"requests_per_second": requests["requests_per_second"],
"output_tokens_per_service_second": requests["output_tokens_per_service_second"],
"decode_tokens_per_second": requests["decode_tokens_per_second"],
"accuracy": requests["accuracy"],
"correct": requests["correct"],
"request_count": requests["request_count"],
"mean_input_tokens": requests["input_tokens"]["mean"],
"mean_output_tokens": requests["output_tokens"]["mean"],
"total_input_tokens": requests["total_input_tokens"],
"total_output_tokens": requests["total_output_tokens"],
"latency_ms": {
key: requests["client_wall_ms"][key]
for key in ("p50", "p95", "p99", "p95_minus_p50")
},
"aggregate_prefill_tokens_per_second": summary["prefill_service"]["aggregate_prompt_tokens_per_second"],
"per_user_model_prefill_tokens_per_second": summary["model_prefill"]["model_prefill_tokens_per_second"],
"process_vram_peak_bytes": telemetry["process_vram_bytes"]["max"],
"active_power_mean_w": telemetry["active_power_w"]["mean"],
"active_power_p95_w": telemetry["active_power_w"]["p95"],
"temperature_p95_c": telemetry["active_temperature_c"]["p95"],
"idle_subtracted_joules_per_output_token": telemetry["idle_subtracted_joules_per_output_token"],
"gpu_uuid": manifest["gpu"]["gpu_uuid"],
"driver_version": manifest["gpu"]["driver_version"],
"model_revision": manifest["model"]["revision"],
"request_stream_sha256": manifest["workload"]["request_stream"]["sha256"],
"harness_sha256": manifest["source"]["harness_sha256"],
"source_fingerprint": manifest["source"]["fingerprint"],
}
return result
def generate_readme(cells: list[dict]) -> str:
service_wins = sum(cell["output_delta_pct"] > 0 for cell in cells)
request_wins = sum(cell["request_delta_pct"] > 0 for cell in cells)
prefill_wins = sum(cell["prefill_delta_pct"] > 0 for cell in cells)
decode_wins = sum(cell["decode_delta_pct"] > 0 for cell in cells)
nonwins = [cell for cell in cells if cell["decode_delta_pct"] <= 0]
nonwin_text = ", ".join(
f'{cell["model"]} C{cell["concurrency"]} ({cell["decode_delta_pct"]:+.1f}%)'
for cell in nonwins
)
if nonwins:
nonwin_line = (
f"- The {len(nonwins)} timing-matched decode non-win"
f"{'s are' if len(nonwins) != 1 else ' is'} {nonwin_text}."
)
else:
nonwin_line = "- Photon leads in every timing-matched decode cell."
same_gpu = sum(cell["same_gpu_uuid"] for cell in cells)
same_harness = sum(cell["same_harness"] for cell in cells)
same_stream = sum(cell["same_request_stream"] for cell in cells)
same_driver = sum(cell["same_driver"] for cell in cells)
same_revision = sum(cell["same_model_revision"] for cell in cells)
same_model_fingerprint = sum(cell["same_model_fingerprint"] for cell in cells)
same_input_tokens = sum(cell["same_input_tokens"] for cell in cells)
quality_warnings = [
cell
for cell in cells
if cell["accuracy_comparable"]
and (
abs(cell["accuracy_delta_pp"]) >= 2.0
or cell["output_mean_gap_pct"] >= 10.0
)
]
incomparable_accuracy = sum(not cell["accuracy_comparable"] for cell in cells)
service_rows = []
prefill_rows = []
latency_rows = []
quality_rows = []
c1_rows = []
for cell in cells:
p_req = cell["photon"]["summary"]["requests"]
v_req = cell["vllm"]["summary"]["requests"]
p_tel = cell["photon"]["summary"]["telemetry"]
v_tel = cell["vllm"]["summary"]["telemetry"]
service_rows.append(
[
cell["model"],
str(cell["concurrency"]),
f'{p_req["output_tokens_per_service_second"]:,.1f}',
f'{v_req["output_tokens_per_service_second"]:,.1f}',
percent(cell["output_delta_pct"]),
percent(cell["request_delta_pct"]),
percent(cell["decode_delta_pct"]),
(
f'{p_req["accuracy"] * 100:.1f}% / {v_req["accuracy"] * 100:.1f}%'
if cell["accuracy_comparable"]
else "n/a (legacy parser)"
),
]
)
prefill_rows.append(
[
cell["model"],
str(cell["concurrency"]),
f'{cell["photon"]["summary"]["model_prefill"]["model_prefill_tokens_per_second"]:,.1f}',
f'{cell["vllm"]["summary"]["model_prefill"]["model_prefill_tokens_per_second"]:,.1f}',
percent(cell["prefill_delta_pct"]),
]
)
latency_rows.append(
[
cell["model"],
str(cell["concurrency"]),
" / ".join(f'{p_req["client_wall_ms"][key]:.1f}' for key in ("p50", "p95", "p99")),
" / ".join(f'{v_req["client_wall_ms"][key]:.1f}' for key in ("p50", "p95", "p99")),
f'{p_req["client_wall_ms"]["p95_minus_p50"]:.1f}',
f'{v_req["client_wall_ms"]["p95_minus_p50"]:.1f}',
]
)
if cell["accuracy_comparable"]:
quality_rows.append(
[
cell["model"],
str(cell["concurrency"]),
f'{p_req["correct"]}/{p_req["request_count"]}',
f'{v_req["correct"]}/{v_req["request_count"]}',
f'{cell["accuracy_delta_pp"]:+.2f} pp',
f'{p_req["output_tokens"]["mean"]:.1f} / {v_req["output_tokens"]["mean"]:.1f}',
"warning" if cell in quality_warnings else "within threshold",
]
)
if cell["concurrency"] == 1:
for backend, summary, telemetry in (
("Photon", p_req, p_tel),
("vLLM", v_req, v_tel),
):
c1_rows.append(
[
cell["model"],
backend,
f'{gib(telemetry["process_vram_bytes"]["max"]):.2f} GiB',
f'{telemetry["active_power_w"]["mean"]:.1f} / {telemetry["active_power_w"]["p95"]:.1f} W',
f'{telemetry["active_temperature_c"]["p95"]:.0f} C',
f'{telemetry["idle_subtracted_joules_per_output_token"]:.3f}',
]
)
lines = [
"# Photon vs vLLM on NVIDIA B200 (final shipping board)",
"",
f"> **Status:** Final 28-cell board against vLLM {VLLM_VERSION}. The board retains 26 validated, unchanged cells from `{VALIDATED_SOURCE_COMMIT[:8]}`, retains the frozen Gemma E4B C8 measurement from `{C8_SOURCE_COMMIT[:8]}`, and replaces Gemma E4B C4 with the final fast/correct release measurement from `{C4_SOURCE_COMMIT[:8]}`. Every Photon/vLLM pair is internally matched on driver, model revision, harness, ordered request stream, request count, and B200 hardware class; exact engine commits are recorded per cell.",
"",
"ChartQA multimodal inference with reasoning enabled on NVIDIA B200. The primary metric is natural-output tokens completed per active end-to-end service second; accuracy and output-length differences are reported alongside it.",
"",
"## At a glance",
"",
f"- Photon leads in **{service_wins}/28** end-to-end output-token service-throughput cells.",
f"- Photon leads in **{request_wins}/28** completed-request-rate cells.",
f"- Photon leads in **{prefill_wins}/28** timing-matched model-prefill cells.",
f"- Photon leads in **{decode_wins}/28** timing-matched decode cells.",
nonwin_line,
f"- Pair checks: request stream {same_stream}/28, input tokens {same_input_tokens}/28, model revision {same_revision}/28, engine-side model fingerprint {same_model_fingerprint}/28, driver {same_driver}/28, exact GPU UUID {same_gpu}/28, matching harness hash {same_harness}/28.",
f"- {len(quality_warnings)} comparable cells cross the report warning threshold of 2 accuracy points or 10% mean-output-length difference; {incomparable_accuracy} retained Gemma rows use the superseded parser and are excluded from accuracy comparisons.",
"- 24 cells use 128 requests per arm; Qwen 0.8B/2B/4B/9B C8 use 256 to satisfy the declared-concurrency floor.",
"- The Kestrel 0.6.1 release tree passed its full CPU suite (839 passed, 46 hardware-dependent skips); focused query, streaming, and scheduler regressions also passed.",
"",
"## Output-token service throughput",
"",
f"![Output-token service throughput]({GIST_RAW}/throughput-by-concurrency.png)",
"",
f"![Photon service-throughput delta]({GIST_RAW}/speedup-vs-vllm.png)",
"",
table(
["Model", "C", "Photon output tok/s", "vLLM output tok/s", "Service delta", "Req/s delta", "Decode delta", "Accuracy P / V"],
service_rows,
),
"",
"## Model-prefill throughput",
"",
"Model-prefill tok/s is the timing-matched model prefill rate recorded by the service run. Prefix caching is disabled.",
"",
f"![Model-prefill throughput]({GIST_RAW}/model-prefill-throughput.png)",
"",
f"![Model-prefill speedup]({GIST_RAW}/model-prefill-speedup.png)",
"",
table(
["Model", "C", "Photon prefill tok/s", "vLLM prefill tok/s", "Photon vs vLLM"],
prefill_rows,
),
"",
"## Serving trade-off",
"",
"The horizontal axis is aggregate end-to-end output tok/s and includes prefill. The vertical axis is decode tok/s per concurrent user and excludes prefill. Higher and farther right are better, but the axes are not commensurable.",
"",
f"![Serving trade-off]({GIST_RAW}/decode-scaling.png)",
"",
"## Latency and jitter",
"",
table(
["Model", "C", "Photon p50 / p95 / p99 ms", "vLLM p50 / p95 / p99 ms", "Photon p95-p50", "vLLM p95-p50"],
latency_rows,
),
"",
"## Quality audit",
"",
"Accuracy and generation length are measured outcomes. Six retained Gemma r40 rows used the superseded answer parser and are excluded instead of publishing their invalid scores. A warning means an absolute accuracy gap of at least 2 percentage points or a mean-output-length gap of at least 10%.",
"",
table(
["Model", "C", "Photon correct", "vLLM correct", "Accuracy delta", "Mean output P / V", "Assessment"],
quality_rows,
),
"",
"These are measured output differences, not token-normalized quality comparisons. The artifact ledger preserves the exact natural outputs and reports every threshold crossing.",
"",
"## GPU operating characteristics at C1",
"",
table(
["Model", "Backend", "Peak process VRAM", "Power mean / p95", "Temperature p95", "Incremental J/output token"],
c1_rows,
),
"",
"## Methodology",
"",
"- NVIDIA B200, tensor parallelism 1, ChartQA test split, greedy decoding, reasoning enabled, no prefix cache.",
"- 24 cells contain 128 requests per arm. Qwen 0.8B/2B/4B/9B C8 contain 256 requests per arm to remove finite-window boundary underfill.",
f"- The final board uses 26 unchanged validated rows from runtime `{VALIDATED_SOURCE_COMMIT}` plus public runtime `{VALIDATED_PUBLIC_COMMIT}`, Gemma E4B C8 from runtime `{C8_SOURCE_COMMIT}` plus parser `{C8_PUBLIC_COMMIT}`, and Gemma E4B C4 from runtime `{C4_SOURCE_COMMIT}` plus parser `{C4_PUBLIC_COMMIT}`. The baseline is vLLM {VLLM_VERSION} at `{VLLM_COMMIT}`.",
"- Gemma E4B C8 is the predeclared median of three exact repeats; the selected run is the median, not the best run. Its Photon repeat and retained vLLM baseline use separate B200s on the same host; all other pairs match exact GPU UUID.",
"- Every pair uses a byte-identical ordered request stream and the same model revision, driver, harness, and B200 hardware class.",
"- Photon is in-process; vLLM is served out-of-process over HTTP. The chart reports each engine in its native deployment shape.",
"- Service throughput is completed natural-output tokens divided by the closed-loop active service wall. Request rate uses completed requests over the same wall. Latency is client-observed end-to-end time.",
"- Model-prefill throughput is the timing-matched model prefill rate from the natural-output service run; prefix caching is disabled.",
"- Decode tok/s counts tokens after the first token over first-token-to-completion time. The first token is produced by prefill.",
"",
"## Validation findings and release status",
"",
"1. **Final board:** 26 unchanged validated rows and the frozen Gemma E4B C8 row are retained; Gemma E4B C4 is replaced by the final fast/correct release measurement.",
"2. **Pair fairness:** the report rejects any pair whose B200 hardware class, driver, model revision, harness, ordered request stream, or request count differs; exact GPU UUID is also reported.",
"3. **Natural outputs:** accuracy, token counts, and output-length gaps are reported directly for both engines; warning cells are not silently normalized away.",
"",
f"C4 release source trees: runtime `{C4_SOURCE_TREE}`, public `{C4_PUBLIC_TREE}`; B200-installed CPython 3.12 x86_64 kernels wheel SHA-256 `{WHEEL_SHA256}`; companion A/B bundle SHA-256 `{BUNDLE_A_SHA256}` / `{BUNDLE_B_SHA256}`. Retained C8 source tree: `{C8_SOURCE_TREE}`.",
"",
"## Artifact inventory",
"",
"- `report-data.json`: normalized values and pair-level provenance checks used by this report.",
"- `SHA256SUMS`: hashes for the report, data ledger, generator, and charts.",
"- Private raw manifests, requests, telemetry, runtime-batch traces, and cold-start records remain in the retained artifact tree.",
"",
]
return "\n".join(lines)
def main() -> None:
unresolved = {
name: value
for name, value in {
"C4_SOURCE_COMMIT": C4_SOURCE_COMMIT,
"C4_SOURCE_TREE": C4_SOURCE_TREE,
"C4_PUBLIC_COMMIT": C4_PUBLIC_COMMIT,
"C4_PUBLIC_TREE": C4_PUBLIC_TREE,
"BOARD_MAP_SHA256": BOARD_MAP_SHA256,
"WHEEL_SHA256": WHEEL_SHA256,
"BUNDLE_A_SHA256": BUNDLE_A_SHA256,
"BUNDLE_B_SHA256": BUNDLE_B_SHA256,
}.items()
if value.startswith("__FINAL_")
}
if unresolved:
raise SystemExit(
"refusing to generate with unresolved release identities: "
+ ", ".join(sorted(unresolved))
)
if not INPUT.is_dir():
raise SystemExit(f"missing input directory: {INPUT}")
if not BOARD_MAP.is_file():
raise SystemExit(f"missing board map: {BOARD_MAP}")
board_map_sha256 = hashlib.sha256(BOARD_MAP.read_bytes()).hexdigest()
if board_map_sha256 != BOARD_MAP_SHA256:
raise AssertionError(
f"board map hash mismatch: expected {BOARD_MAP_SHA256}, got {board_map_sha256}"
)
OUTPUT.mkdir(parents=True, exist_ok=True)
cells = load_cells()
if len(cells) != 28:
raise AssertionError(f"expected 28 paired cells, found {len(cells)}")
assert all(
cell["request_count"] == REQUEST_COUNTS.get((cell["tag"], cell["concurrency"]), 128)
for cell in cells
)
assert all(cell["same_request_stream"] for cell in cells)
assert all(cell["same_driver"] for cell in cells)
assert all(cell["same_model_revision"] for cell in cells)
assert all(cell["same_input_tokens"] for cell in cells)
assert all(cell["same_gpu_model"] for cell in cells)
assert all(cell["same_harness"] for cell in cells)
assert all(cell["same_request_count"] for cell in cells)
assert all(cell["source_commit_exact"] for cell in cells)
assert all(cell["vllm_commit"] == VLLM_COMMIT for cell in cells)
source_cohorts = Counter(
(cell["source_commit"], cell["public_commit"]) for cell in cells
)
assert dict(source_cohorts) == EXPECTED_SOURCE_COHORTS
assert all(
(cell["source_commit"], cell["public_commit"])
== EXPECTED_CELL_SOURCES.get(
(cell["tag"], cell["concurrency"]),
(VALIDATED_SOURCE_COMMIT, VALIDATED_PUBLIC_COMMIT),
)
for cell in cells
)
line_chart(
cells,
"requests",
"output_tokens_per_service_second",
"Output tokens/s",
"throughput-by-concurrency.png",
)
speedup_chart(
cells,
"output_delta_pct",
"Photon output-token service-throughput delta",
"speedup-vs-vllm.png",
)
line_chart(
cells,
"model_prefill",
"model_prefill_tokens_per_second",
"Prompt tokens/s",
"model-prefill-throughput.png",
)
speedup_chart(cells, "prefill_delta_pct", "Photon model-prefill delta", "model-prefill-speedup.png")
tradeoff_chart(cells)
provenance = {
"board_map_sha256": BOARD_MAP_SHA256,
"c4_source_commit": C4_SOURCE_COMMIT,
"c4_source_tree": C4_SOURCE_TREE,
"c4_public_commit": C4_PUBLIC_COMMIT,
"c4_public_tree": C4_PUBLIC_TREE,
"c8_source_commit": C8_SOURCE_COMMIT,
"c8_source_tree": C8_SOURCE_TREE,
"c8_public_commit": C8_PUBLIC_COMMIT,
"wheel_sha256": WHEEL_SHA256,
"bundle_a_sha256": BUNDLE_A_SHA256,
"bundle_b_sha256": BUNDLE_B_SHA256,
"vllm_version": VLLM_VERSION,
"vllm_commit": VLLM_COMMIT,
"source_cohorts": [
{
"runtime_commit": runtime_commit,
"public_commit": public_commit,
"cells": count,
}
for (runtime_commit, public_commit), count in sorted(source_cohorts.items())
],
"measured_cells": len(cells),
"request_count_per_arm": {
"default": 128,
"qwen08-c8": 256,
"qwen2-c8": 256,
"qwen4-c8": 256,
"qwen9-c8": 256,
},
"service_throughput_wins": sum(cell["output_delta_pct"] > 0 for cell in cells),
"request_throughput_wins": sum(cell["request_delta_pct"] > 0 for cell in cells),
"decode_throughput_wins": sum(cell["decode_delta_pct"] > 0 for cell in cells),
"model_prefill_wins": sum(cell["prefill_delta_pct"] > 0 for cell in cells),
"same_gpu_pairs": sum(cell["same_gpu_uuid"] for cell in cells),
"same_harness_pairs": sum(cell["same_harness"] for cell in cells),
"same_request_stream_pairs": sum(cell["same_request_stream"] for cell in cells),
"same_driver_pairs": sum(cell["same_driver"] for cell in cells),
"same_model_fingerprint_pairs": sum(
cell["same_model_fingerprint"] for cell in cells
),
"same_input_token_pairs": sum(cell["same_input_tokens"] for cell in cells),
"quality_warning_cells": sum(
cell["accuracy_comparable"]
and (
abs(cell["accuracy_delta_pp"]) >= 2.0
or cell["output_mean_gap_pct"] >= 10.0
)
for cell in cells
),
"accuracy_comparable_cells": sum(cell["accuracy_comparable"] for cell in cells),
}
validation_findings = [
{
"type": "service_throughput_nonwin",
"cell": f'{cell["tag"]}-c{cell["concurrency"]}',
"model": cell["model"],
"concurrency": cell["concurrency"],
"output_delta_pct": cell["output_delta_pct"],
}
for cell in cells
if cell["output_delta_pct"] <= 0
]
validation_findings.extend(
{
"type": "decode_throughput_nonwin",
"cell": f'{cell["tag"]}-c{cell["concurrency"]}',
"model": cell["model"],
"concurrency": cell["concurrency"],
"decode_delta_pct": cell["decode_delta_pct"],
}
for cell in cells
if cell["decode_delta_pct"] <= 0
)
validation_findings.extend(
{
"type": "quality_warning",
"cell": f'{cell["tag"]}-c{cell["concurrency"]}',
"model": cell["model"],
"concurrency": cell["concurrency"],
"accuracy_delta_pp": cell["accuracy_delta_pp"],
"output_mean_gap_pct": cell["output_mean_gap_pct"],
}
for cell in cells
if cell["accuracy_comparable"]
and (
abs(cell["accuracy_delta_pp"]) >= 2.0
or cell["output_mean_gap_pct"] >= 10.0
)
)
ledger = {
"schema_version": 2,
"status": BOARD_STATUS,
"provenance": provenance,
"cells": [serializable_cell(cell) for cell in cells],
"validation_findings": validation_findings,
}
(OUTPUT / "report-data.json").write_text(
json.dumps(ledger, indent=2, sort_keys=True) + "\n", encoding="utf-8"
)
(OUTPUT / "README.md").write_text(generate_readme(cells), encoding="utf-8")
output_names = [
"README.md",
"report-data.json",
"generate.py",
"throughput-by-concurrency.png",
"speedup-vs-vllm.png",
"model-prefill-throughput.png",
"model-prefill-speedup.png",
"decode-scaling.png",
]
sums = []
for name in output_names:
digest = hashlib.sha256((OUTPUT / name).read_bytes()).hexdigest()
sums.append(f"{digest} {name}")
(OUTPUT / "SHA256SUMS").write_text("\n".join(sums) + "\n", encoding="ascii")
print(json.dumps(provenance, indent=2, sort_keys=True))
if __name__ == "__main__":
main()
{
"cells": [
{
"accuracy_comparable": true,
"accuracy_delta_pp": 1.5625,
"concurrency": 1,
"decode_delta_pct": 59.0125094056811,
"model": "Moondream 3",
"output_delta_pct": 155.63125665277383,
"output_mean_gap_pct": 5.863490609253321,
"photon": {
"accuracy": 0.8515625,
"active_power_mean_w": 451.5742916666667,
"active_power_p95_w": 496.64745,
"aggregate_prefill_tokens_per_second": 45572.903512705176,
"correct": 109,
"decode_tokens_per_second": 412.6919481919769,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-be7f328f-bb91-2065-b8c1-4f8040cc1ac0",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.8129130980091451,
"latency_ms": {
"p50": 51.41467892099172,
"p95": 89.24412068445237,
"p95_minus_p50": 37.82944176346065,
"p99": 124.0118882013486
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 17.0546875,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 300.6393741306282,
"per_user_model_prefill_tokens_per_second": 61406.902089487965,
"process_vram_peak_bytes": 20289945600.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 17.62796146986734,
"source_fingerprint": "8f68fcb33c41bd1edff6440d4973d899e9b0b5ee2d4dee14ecb9bf4fd5df652e",
"temperature_p95_c": 43.0,
"total_input_tokens": 96200,
"total_output_tokens": 2183
},
"prefill_delta_pct": 213.33813912384142,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 140.64234192462214,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "md3",
"vllm": {
"accuracy": 0.8359375,
"active_power_mean_w": 345.1964571428571,
"active_power_p95_w": 359.6708,
"aggregate_prefill_tokens_per_second": 9809.76475481446,
"correct": 107,
"decode_tokens_per_second": 259.5342654074438,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-be7f328f-bb91-2065-b8c1-4f8040cc1ac0",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.1452794350081097,
"latency_ms": {
"p50": 132.56514602107927,
"p95": 187.7853721671272,
"p95_minus_p50": 55.220226146047935,
"p99": 228.77975672367032
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 16.0546875,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 117.60665658307556,
"per_user_model_prefill_tokens_per_second": 19597.64689392566,
"process_vram_peak_bytes": 173365264384.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 7.325378122936093,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 39.0,
"total_input_tokens": 96200,
"total_output_tokens": 2055
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 1.5625,
"concurrency": 2,
"decode_delta_pct": 61.82205510224104,
"model": "Moondream 3",
"output_delta_pct": 111.80415875899689,
"output_mean_gap_pct": 1.6052880075542966,
"photon": {
"accuracy": 0.8515625,
"active_power_mean_w": 511.20504255319145,
"active_power_p95_w": 601.5657000000001,
"aggregate_prefill_tokens_per_second": 51553.461890509265,
"correct": 109,
"decode_tokens_per_second": 334.60583246545775,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-17fee275-053d-0826-5778-5977365b2e92",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.6718221351567335,
"latency_ms": {
"p50": 66.91287900321186,
"p95": 112.96138102188706,
"p95_minus_p50": 46.04850201867521,
"p99": 135.04406418418515
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 16.546875,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 453.5655539025692,
"per_user_model_prefill_tokens_per_second": 46835.720929553776,
"process_vram_peak_bytes": 20564672512.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 27.410949433205317,
"source_fingerprint": "8f68fcb33c41bd1edff6440d4973d899e9b0b5ee2d4dee14ecb9bf4fd5df652e",
"temperature_p95_c": 50.7,
"total_input_tokens": 96200,
"total_output_tokens": 2118
},
"prefill_delta_pct": 111.44589049431298,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 108.40409199893753,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "md3",
"vllm": {
"accuracy": 0.8359375,
"active_power_mean_w": 420.8397422680413,
"active_power_p95_w": 474.10499999999996,
"aggregate_prefill_tokens_per_second": 19251.897747363142,
"correct": 107,
"decode_tokens_per_second": 206.77393588534636,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-17fee275-053d-0826-5778-5977365b2e92",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.0167185612042766,
"latency_ms": {
"p50": 141.52954402379692,
"p95": 232.4307778966614,
"p95_minus_p50": 90.90123387286448,
"p99": 287.31910773902206
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 16.28125,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 214.1438376659367,
"per_user_model_prefill_tokens_per_second": 22150.215745533,
"process_vram_peak_bytes": 173086343168.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 13.15278849387711,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 47.0,
"total_input_tokens": 96200,
"total_output_tokens": 2084
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 0.78125,
"concurrency": 4,
"decode_delta_pct": 57.313026494608145,
"model": "Moondream 3",
"output_delta_pct": 87.60558749003853,
"output_mean_gap_pct": 2.954755309325946,
"photon": {
"accuracy": 0.84375,
"active_power_mean_w": 555.1217575757576,
"active_power_p95_w": 682.0768,
"aggregate_prefill_tokens_per_second": 62681.18624977897,
"correct": 108,
"decode_tokens_per_second": 238.42246640614354,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-27fd3c5e-69b4-4789-d473-b977761394b1",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.5205343888253626,
"latency_ms": {
"p50": 91.95472206920385,
"p95": 168.01079894648865,
"p95_minus_p50": 76.0560768772848,
"p99": 226.61869481438788
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 16.921875,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 662.5153062064276,
"per_user_model_prefill_tokens_per_second": 39525.57464323883,
"process_vram_peak_bytes": 21099446272.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 39.151412370462936,
"source_fingerprint": "8f68fcb33c41bd1edff6440d4973d899e9b0b5ee2d4dee14ecb9bf4fd5df652e",
"temperature_p95_c": 46.0,
"total_input_tokens": 96200,
"total_output_tokens": 2166
},
"prefill_delta_pct": 99.23799370122224,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 82.06230143308446,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "md3",
"vllm": {
"accuracy": 0.8359375,
"active_power_mean_w": 505.5893220338983,
"active_power_p95_w": 592.9142,
"aggregate_prefill_tokens_per_second": 32405.886946141087,
"correct": 107,
"decode_tokens_per_second": 151.55926481035274,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-27fd3c5e-69b4-4789-d473-b977761394b1",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.852472182672398,
"latency_ms": {
"p50": 172.31369146611542,
"p95": 270.67310936399736,
"p95_minus_p50": 98.35941789788194,
"p99": 398.9698972937187
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 16.421875,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 353.1426302756605,
"per_user_model_prefill_tokens_per_second": 19838.372144276596,
"process_vram_peak_bytes": 173059080192.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 21.504403746567338,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 44.0,
"total_input_tokens": 96200,
"total_output_tokens": 2102
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 0.0,
"concurrency": 8,
"decode_delta_pct": 60.3163147028154,
"model": "Moondream 3",
"output_delta_pct": 97.5500216209584,
"output_mean_gap_pct": 2.604651162790698,
"photon": {
"accuracy": 0.84375,
"active_power_mean_w": 615.7250833333334,
"active_power_p95_w": 778.2176,
"aggregate_prefill_tokens_per_second": 65743.61838220243,
"correct": 108,
"decode_tokens_per_second": 159.23294462848028,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-f68ad80c-9250-2667-5ed1-dec1d7a0349f",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.4377784608085209,
"latency_ms": {
"p50": 133.46018560696393,
"p95": 246.47215622244403,
"p95_minus_p50": 113.0119706154801,
"p99": 279.58005774533376
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 16.796875,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 906.4046701361575,
"per_user_model_prefill_tokens_per_second": 30383.96709268066,
"process_vram_peak_bytes": 21252538368.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 53.96269664066426,
"source_fingerprint": "8f68fcb33c41bd1edff6440d4973d899e9b0b5ee2d4dee14ecb9bf4fd5df652e",
"temperature_p95_c": 53.0,
"total_input_tokens": 96200,
"total_output_tokens": 2150
},
"prefill_delta_pct": 112.31408494058721,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 92.40453268571484,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "md3",
"vllm": {
"accuracy": 0.84375,
"active_power_mean_w": 533.8851086956522,
"active_power_p95_w": 688.502,
"aggregate_prefill_tokens_per_second": 41285.71590227106,
"correct": 108,
"decode_tokens_per_second": 99.32422967909199,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-f68ad80c-9250-2667-5ed1-dec1d7a0349f",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.6982190222459065,
"latency_ms": {
"p50": 238.6042860453017,
"p95": 696.3062826893292,
"p95_minus_p50": 457.70199664402753,
"p99": 876.6237004997678
},
"mean_input_tokens": 751.5625,
"mean_output_tokens": 16.359375,
"model_revision": "5112966d1a723413b1c9a1e8bea272b72e647b35",
"output_tokens_per_service_second": 458.82286557036525,
"per_user_model_prefill_tokens_per_second": 14310.857944814703,
"process_vram_peak_bytes": 173331709952.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 28.046478888732928,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 51.0,
"total_input_tokens": 96200,
"total_output_tokens": 2094
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 3.125,
"concurrency": 1,
"decode_delta_pct": 83.23949814495784,
"model": "Qwen3.5 0.8B",
"output_delta_pct": 91.06111494781888,
"output_mean_gap_pct": 12.592793280953671,
"photon": {
"accuracy": 0.75,
"active_power_mean_w": 500.10672972972975,
"active_power_p95_w": 517.9579,
"aggregate_prefill_tokens_per_second": 28368.640149833736,
"correct": 96,
"decode_tokens_per_second": 1250.4207672242721,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-73f7107c-066d-e32e-2714-cba521de3564",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.24852967494797232,
"latency_ms": {
"p50": 205.4614459630102,
"p95": 1246.9797140685841,
"p95_minus_p50": 1041.518268105574,
"p99": 1267.0868718391284
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 378.0703125,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 1189.4108745245032,
"per_user_model_prefill_tokens_per_second": 28463.39320004205,
"process_vram_peak_bytes": 7134511104.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 3.146004420869473,
"source_fingerprint": "8f68fcb33c41bd1edff6440d4973d899e9b0b5ee2d4dee14ecb9bf4fd5df652e",
"temperature_p95_c": 45.0,
"total_input_tokens": 44927,
"total_output_tokens": 48393
},
"prefill_delta_pct": 142.9690026357147,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 118.58737067522145,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen08",
"vllm": {
"accuracy": 0.71875,
"active_power_mean_w": 341.56402587176603,
"active_power_p95_w": 354.8396,
"aggregate_prefill_tokens_per_second": 7669.358131418871,
"correct": 92,
"decode_tokens_per_second": 682.3969612900186,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-73f7107c-066d-e32e-2714-cba521de3564",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.2209887996624275,
"latency_ms": {
"p50": 413.89824147336185,
"p95": 2304.7922479338013,
"p95_minus_p50": 1890.8940064604394,
"p99": 2323.454610792687
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 432.5390625,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 622.5290137395806,
"per_user_model_prefill_tokens_per_second": 11714.824891765076,
"process_vram_peak_bytes": 171712708608.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.4392434526987503,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 41.0,
"total_input_tokens": 44927,
"total_output_tokens": 55365
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -0.78125,
"concurrency": 2,
"decode_delta_pct": 63.99078293846032,
"model": "Qwen3.5 0.8B",
"output_delta_pct": 69.70573168901275,
"output_mean_gap_pct": 1.1896632282643855,
"photon": {
"accuracy": 0.71875,
"active_power_mean_w": 492.44331250000005,
"active_power_p95_w": 507.2835,
"aggregate_prefill_tokens_per_second": 32502.3043278718,
"correct": 92,
"decode_tokens_per_second": 1052.6918301416506,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-0e95d6ef-9a19-e489-0142-b24d69f0894b",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.14219970155141448,
"latency_ms": {
"p50": 247.22146603744477,
"p95": 1485.191235539969,
"p95_minus_p50": 1237.9697695025243,
"p99": 1516.830346379429
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 429.5625,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 2016.9414332431918,
"per_user_model_prefill_tokens_per_second": 28176.12753436353,
"process_vram_peak_bytes": 7134511104.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 4.695338706807954,
"source_fingerprint": "8f68fcb33c41bd1edff6440d4973d899e9b0b5ee2d4dee14ecb9bf4fd5df652e",
"temperature_p95_c": 52.0,
"total_input_tokens": 44927,
"total_output_tokens": 54984
},
"prefill_delta_pct": 47.844859202869095,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 71.74896598222762,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen08",
"vllm": {
"accuracy": 0.7265625,
"active_power_mean_w": 347.7979871794872,
"active_power_p95_w": 355.7618,
"aggregate_prefill_tokens_per_second": 13525.4566420125,
"correct": 93,
"decode_tokens_per_second": 641.9213392844688,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-0e95d6ef-9a19-e489-0142-b24d69f0894b",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.12092145941340907,
"latency_ms": {
"p50": 440.6808084459044,
"p95": 2425.0949896813836,
"p95_minus_p50": 1984.4141812354792,
"p99": 2469.108702467056
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 434.734375,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 1188.4934074821083,
"per_user_model_prefill_tokens_per_second": 19057.90142875441,
"process_vram_peak_bytes": 171704320000.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 2.73383812237555,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 48.0,
"total_input_tokens": 44927,
"total_output_tokens": 55646
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 3.125,
"concurrency": 4,
"decode_delta_pct": 39.731436021168484,
"model": "Qwen3.5 0.8B",
"output_delta_pct": 49.049846541665445,
"output_mean_gap_pct": 4.683777344596058,
"photon": {
"accuracy": 0.765625,
"active_power_mean_w": 442.38224539877297,
"active_power_p95_w": 463.8906,
"aggregate_prefill_tokens_per_second": 47190.20368691797,
"correct": 98,
"decode_tokens_per_second": 844.7799049906781,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-00c0cafc-2251-3f2d-eff7-2d6ac0967a4c",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.07532424815790607,
"latency_ms": {
"p50": 287.6521560829133,
"p95": 1818.570830929093,
"p95_minus_p50": 1530.9186748461798,
"p99": 1938.5403614933605
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 412.5703125,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 3230.079324853146,
"per_user_model_prefill_tokens_per_second": 25433.375770911854,
"process_vram_peak_bytes": 7186939904.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 7.829160816928983,
"source_fingerprint": "35b645ff0b5f667b907a1452ed6fad91d5d78b528bfdf8a73d217677bbf9b45a",
"temperature_p95_c": 43.0,
"total_input_tokens": 44927,
"total_output_tokens": 52809
},
"prefill_delta_pct": 38.410370287747455,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 56.37405930417982,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen08",
"vllm": {
"accuracy": 0.734375,
"active_power_mean_w": 347.072703125,
"active_power_p95_w": 357.19075,
"aggregate_prefill_tokens_per_second": 29177.886594245425,
"correct": 94,
"decode_tokens_per_second": 604.5739806629475,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-00c0cafc-2251-3f2d-eff7-2d6ac0967a4c",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.06610391484029267,
"latency_ms": {
"p50": 469.4851109525189,
"p95": 2563.663464272395,
"p95_minus_p50": 2094.178353319876,
"p99": 2671.1555312480778
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 432.84375,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 2167.113485722515,
"per_user_model_prefill_tokens_per_second": 18375.339736493213,
"process_vram_peak_bytes": 171819663360.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 5.006687715191718,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 40.0,
"total_input_tokens": 44927,
"total_output_tokens": 55404
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 0.0,
"concurrency": 8,
"decode_delta_pct": 15.612380499668221,
"model": "Qwen3.5 0.8B",
"output_delta_pct": 21.97128263536794,
"output_mean_gap_pct": 1.1381611243431937,
"photon": {
"accuracy": 0.73046875,
"active_power_mean_w": 456.4589777777778,
"active_power_p95_w": 476.4208,
"aggregate_prefill_tokens_per_second": 78508.56317800854,
"correct": 187,
"decode_tokens_per_second": 639.4976115448139,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-8c48cd96-8229-b187-74c5-eb6787035398",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.04918443668158653,
"latency_ms": {
"p50": 422.30539256706834,
"p95": 2420.1855122228153,
"p95_minus_p50": 1997.880119655747,
"p99": 2478.947615204379
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 429.6953125,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 4875.6812164906605,
"per_user_model_prefill_tokens_per_second": 28665.170359004285,
"process_vram_peak_bytes": 7361003520.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 11.346833615948883,
"source_fingerprint": "35b645ff0b5f667b907a1452ed6fad91d5d78b528bfdf8a73d217677bbf9b45a",
"temperature_p95_c": 49.0,
"total_input_tokens": 105702,
"total_output_tokens": 110002
},
"prefill_delta_pct": 44.15858323693567,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 256,
"request_delta_pct": 20.583052913549405,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen08",
"vllm": {
"accuracy": 0.73046875,
"active_power_mean_w": 366.7497389705882,
"active_power_p95_w": 385.92015,
"aggregate_prefill_tokens_per_second": 37669.663333781114,
"correct": 187,
"decode_tokens_per_second": 553.1393859212587,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-8c48cd96-8229-b187-74c5-eb6787035398",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.03783933390446137,
"latency_ms": {
"p50": 523.8134315004572,
"p95": 2837.8001830133144,
"p95_minus_p50": 2313.986751512857,
"p99": 2924.670825555222
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 424.8046875,
"model_revision": "2fc06364715b967f1860aea9cf38778875588b17",
"output_tokens_per_service_second": 3997.400954670999,
"per_user_model_prefill_tokens_per_second": 19884.47008520532,
"process_vram_peak_bytes": 171930812416.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 9.409973741570353,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 47.0,
"total_input_tokens": 105702,
"total_output_tokens": 108750
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -1.5625,
"concurrency": 1,
"decode_delta_pct": 50.30910544027505,
"model": "Qwen3.5 2B",
"output_delta_pct": 58.19868208839911,
"output_mean_gap_pct": 1.1231350928495596,
"photon": {
"accuracy": 0.75,
"active_power_mean_w": 665.7675858895706,
"active_power_p95_w": 687.39775,
"aggregate_prefill_tokens_per_second": 24864.535930230413,
"correct": 96,
"decode_tokens_per_second": 852.1248508275826,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-ccbbb0a9-59e7-ebf9-8ffe-8a1b5c9fe1d4",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.5676134607461217,
"latency_ms": {
"p50": 265.66908054519445,
"p95": 1820.4996104119346,
"p95_minus_p50": 1554.83052986674,
"p99": 1840.1636486244388
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 419.4453125,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 822.6774201219528,
"per_user_model_prefill_tokens_per_second": 24443.30716722956,
"process_vram_peak_bytes": 16198402048.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.9613460816109436,
"source_fingerprint": "35b645ff0b5f667b907a1452ed6fad91d5d78b528bfdf8a73d217677bbf9b45a",
"temperature_p95_c": 53.0,
"total_input_tokens": 44927,
"total_output_tokens": 53689
},
"prefill_delta_pct": 132.2173110374916,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 56.42189717343879,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen2",
"vllm": {
"accuracy": 0.765625,
"active_power_mean_w": 441.16591185112634,
"active_power_p95_w": 473.071,
"aggregate_prefill_tokens_per_second": 8308.175422801345,
"correct": 98,
"decode_tokens_per_second": 566.9149905001412,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-ccbbb0a9-59e7-ebf9-8ffe-8a1b5c9fe1d4",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.46836113214334835,
"latency_ms": {
"p50": 490.0882520596497,
"p95": 2704.7611344605693,
"p95_minus_p50": 2214.6728824009197,
"p99": 2794.5674144593067
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 414.734375,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 520.0279858603706,
"per_user_model_prefill_tokens_per_second": 10526.048664512859,
"process_vram_peak_bytes": 171433787392.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.253882044044144,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 45.0,
"total_input_tokens": 44927,
"total_output_tokens": 53086
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -0.78125,
"concurrency": 2,
"decode_delta_pct": 49.59597503397482,
"model": "Qwen3.5 2B",
"output_delta_pct": 53.38818415445672,
"output_mean_gap_pct": 0.8452367770272764,
"photon": {
"accuracy": 0.7421875,
"active_power_mean_w": 646.1771297297297,
"active_power_p95_w": 667.2296,
"aggregate_prefill_tokens_per_second": 28566.93257165784,
"correct": 95,
"decode_tokens_per_second": 760.2668733894786,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-4b9a4813-a57b-4323-c377-10e272fc0ffd",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.3013397282245118,
"latency_ms": {
"p50": 314.4272605422884,
"p95": 2047.2393695497885,
"p95_minus_p50": 1732.8121090075,
"p99": 2106.225579569582
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 423.328125,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 1462.4381247807225,
"per_user_model_prefill_tokens_per_second": 23639.627673007744,
"process_vram_peak_bytes": 16273899520.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 3.454620750229441,
"source_fingerprint": "35b645ff0b5f667b907a1452ed6fad91d5d78b528bfdf8a73d217677bbf9b45a",
"temperature_p95_c": 57.0,
"total_input_tokens": 44927,
"total_output_tokens": 54186
},
"prefill_delta_pct": 38.03365260974834,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 52.09169081036891,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen2",
"vllm": {
"accuracy": 0.75,
"active_power_mean_w": 459.90626287744226,
"active_power_p95_w": 474.29810000000003,
"aggregate_prefill_tokens_per_second": 12606.352774391267,
"correct": 96,
"decode_tokens_per_second": 508.21345508581635,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-4b9a4813-a57b-4323-c377-10e272fc0ffd",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.2683342641809532,
"latency_ms": {
"p50": 555.612844938878,
"p95": 3042.2582867497113,
"p95_minus_p50": 2486.6454418108333,
"p99": 3094.8929824738298
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 419.75,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 953.4229333519568,
"per_user_model_prefill_tokens_per_second": 17125.988645567613,
"process_vram_peak_bytes": 171431690240.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 2.271406630975478,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 52.0,
"total_input_tokens": 44927,
"total_output_tokens": 53728
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -1.5625,
"concurrency": 4,
"decode_delta_pct": 29.172170863403824,
"model": "Qwen3.5 2B",
"output_delta_pct": 37.34960771919169,
"output_mean_gap_pct": 0.05325106961200169,
"photon": {
"accuracy": 0.7421875,
"active_power_mean_w": 590.6654821428572,
"active_power_p95_w": 615.24785,
"aggregate_prefill_tokens_per_second": 42014.32517954294,
"correct": 95,
"decode_tokens_per_second": 634.5040421305171,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-d88fb48d-2dd9-972b-8ad5-a88a88a5e06d",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.15971559612598196,
"latency_ms": {
"p50": 366.99557187967,
"p95": 2457.027105835732,
"p95_minus_p50": 2090.031533956062,
"p99": 2568.0806689825845
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 425.4609375,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 2430.2588290259137,
"per_user_model_prefill_tokens_per_second": 23502.81209531701,
"process_vram_peak_bytes": 16284385280.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 5.712061002135862,
"source_fingerprint": "35b645ff0b5f667b907a1452ed6fad91d5d78b528bfdf8a73d217677bbf9b45a",
"temperature_p95_c": 44.0,
"total_input_tokens": 44927,
"total_output_tokens": 54459
},
"prefill_delta_pct": 55.752994977169394,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 37.276467583973314,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen2",
"vllm": {
"accuracy": 0.7578125,
"active_power_mean_w": 449.04954397394135,
"active_power_p95_w": 465.2278,
"aggregate_prefill_tokens_per_second": 22273.755257156507,
"correct": 97,
"decode_tokens_per_second": 491.2080039294907,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-d88fb48d-2dd9-972b-8ad5-a88a88a5e06d",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.14084637478161244,
"latency_ms": {
"p50": 575.2284920308739,
"p95": 3186.958805186441,
"p95_minus_p50": 2611.7303131555673,
"p99": 3289.5827712200116
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 425.234375,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 1769.3962650366832,
"per_user_model_prefill_tokens_per_second": 15089.797855098775,
"process_vram_peak_bytes": 171482021888.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 4.160990665528118,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 41.0,
"total_input_tokens": 44927,
"total_output_tokens": 54430
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 0.390625,
"concurrency": 8,
"decode_delta_pct": 14.056303202211762,
"model": "Qwen3.5 2B",
"output_delta_pct": 17.495048088350696,
"output_mean_gap_pct": 6.453722111773802,
"photon": {
"accuracy": 0.7578125,
"active_power_mean_w": 585.0007883211679,
"active_power_p95_w": 615.57785,
"aggregate_prefill_tokens_per_second": 60111.37718391565,
"correct": 194,
"decode_tokens_per_second": 509.30474587912863,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-2e989258-545a-9591-4d90-e15ce6929cfa",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.09528177605963654,
"latency_ms": {
"p50": 449.20701009687036,
"p95": 3018.7786474707536,
"p95_minus_p50": 2569.5716373738833,
"p99": 3200.4631873685867
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 413.55859375,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 3859.4412165411877,
"per_user_model_prefill_tokens_per_second": 25487.622801056867,
"process_vram_peak_bytes": 16494100480.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 9.332271834917439,
"source_fingerprint": "35b645ff0b5f667b907a1452ed6fad91d5d78b528bfdf8a73d217677bbf9b45a",
"temperature_p95_c": 54.0,
"total_input_tokens": 105702,
"total_output_tokens": 105871
},
"prefill_delta_pct": 39.21330937301764,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 256,
"request_delta_pct": 25.600986742347676,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen2",
"vllm": {
"accuracy": 0.75390625,
"active_power_mean_w": 464.9696162790698,
"active_power_p95_w": 480.91549999999995,
"aggregate_prefill_tokens_per_second": 37846.17668179058,
"correct": 193,
"decode_tokens_per_second": 446.53800936908874,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-2e989258-545a-9591-4d90-e15ce6929cfa",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.07831252469456326,
"latency_ms": {
"p50": 587.7441025222652,
"p95": 3514.8060717328917,
"p95_minus_p50": 2927.0619692106266,
"p99": 3613.7167446431704
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 442.08984375,
"model_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
"output_tokens_per_service_second": 3284.769255670308,
"per_user_model_prefill_tokens_per_second": 18308.323331904703,
"process_vram_peak_bytes": 171626725376.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 7.430094362284947,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 52.0,
"total_input_tokens": 105702,
"total_output_tokens": 113175
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -3.125,
"concurrency": 1,
"decode_delta_pct": 40.62514443859675,
"model": "Qwen3.5 4B",
"output_delta_pct": 45.774761537172395,
"output_mean_gap_pct": 1.06842680089273,
"photon": {
"accuracy": 0.7265625,
"active_power_mean_w": 807.420353879623,
"active_power_p95_w": 822.4288,
"aggregate_prefill_tokens_per_second": 28709.152889660207,
"correct": 93,
"decode_tokens_per_second": 464.9723494719227,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-4a0e08bf-75e9-6590-fca2-8c0d11a4af89",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.2828834666355193,
"latency_ms": {
"p50": 540.3806593967602,
"p95": 3352.0046980003826,
"p95_minus_p50": 2811.6240386036225,
"p99": 3357.271444920916
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 493.5703125,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 458.23552760361184,
"per_user_model_prefill_tokens_per_second": 27866.242430698258,
"process_vram_peak_bytes": 26902265856.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 0.9284098253044988,
"source_fingerprint": "5d09767acf28428f21ad08776c430e6922b927c8964548eb87e16772cc7d59a0",
"temperature_p95_c": 57.0,
"total_input_tokens": 44927,
"total_output_tokens": 63177
},
"prefill_delta_pct": 289.0428593457143,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 44.217264915971796,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen4",
"vllm": {
"accuracy": 0.7578125,
"active_power_mean_w": 559.6798772635815,
"active_power_p95_w": 583.4734500000001,
"aggregate_prefill_tokens_per_second": 7953.688555057904,
"correct": 97,
"decode_tokens_per_second": 330.64666445548113,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-4a0e08bf-75e9-6590-fca2-8c0d11a4af89",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.0963909766809237,
"latency_ms": {
"p50": 799.5882540126331,
"p95": 4715.580011269776,
"p95_minus_p50": 3915.991757257143,
"p99": 4720.807805925142
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 488.296875,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 314.3448994679112,
"per_user_model_prefill_tokens_per_second": 7162.76928397124,
"process_vram_peak_bytes": 171463147520.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 0.6437577538621585,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 52.0,
"total_input_tokens": 44927,
"total_output_tokens": 62502
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 0.0,
"concurrency": 2,
"decode_delta_pct": 25.781158287888275,
"model": "Qwen3.5 4B",
"output_delta_pct": 28.884172583806,
"output_mean_gap_pct": 0.8286635022939057,
"photon": {
"accuracy": 0.75,
"active_power_mean_w": 763.6189592875318,
"active_power_p95_w": 777.1569999999999,
"aggregate_prefill_tokens_per_second": 34784.27225901024,
"correct": 96,
"decode_tokens_per_second": 407.4617872967088,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-d0c2119b-1596-d0fc-4fe7-c2d51b964acd",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.6766568018982873,
"latency_ms": {
"p50": 617.0120724709705,
"p95": 3821.319415536709,
"p95_minus_p50": 3204.3073430657387,
"p99": 3849.448163774796
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 492.1328125,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 800.8544084761036,
"per_user_model_prefill_tokens_per_second": 26453.51574070803,
"process_vram_peak_bytes": 26908557312.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.6273135790475333,
"source_fingerprint": "5d09767acf28428f21ad08776c430e6922b927c8964548eb87e16772cc7d59a0",
"temperature_p95_c": 55.0,
"total_input_tokens": 44927,
"total_output_tokens": 62993
},
"prefill_delta_pct": 124.41913327099452,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 27.81615648537055,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen4",
"vllm": {
"accuracy": 0.75,
"active_power_mean_w": 572.4340527363184,
"active_power_p95_w": 584.1172,
"aggregate_prefill_tokens_per_second": 11642.791938653492,
"correct": 96,
"decode_tokens_per_second": 323.9450111948477,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-d0c2119b-1596-d0fc-4fe7-c2d51b964acd",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.576354610853281,
"latency_ms": {
"p50": 847.4994894932024,
"p95": 4811.053738661576,
"p95_minus_p50": 3963.5542491683736,
"p99": 4857.22626772942
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 488.0546875,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 621.375295679036,
"per_user_model_prefill_tokens_per_second": 11787.549196513657,
"process_vram_peak_bytes": 171456856064.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.27316735520348,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 52.0,
"total_input_tokens": 44927,
"total_output_tokens": 62471
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 0.0,
"concurrency": 4,
"decode_delta_pct": 6.250030010317853,
"model": "Qwen3.5 4B",
"output_delta_pct": 9.180790155727593,
"output_mean_gap_pct": 2.699539797257761,
"photon": {
"accuracy": 0.75,
"active_power_mean_w": 692.9157723214286,
"active_power_p95_w": 713.4110999999999,
"aggregate_prefill_tokens_per_second": 45888.3252314031,
"correct": 96,
"decode_tokens_per_second": 355.5682334159427,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3b683bea-442e-2758-038a-a8c289f74656",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.35557037088206667,
"latency_ms": {
"p50": 662.1374459937215,
"p95": 4365.025370393414,
"p95_minus_p50": 3702.8879243996926,
"p99": 4433.039798294194
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 480.671875,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 1371.7950249980165,
"per_user_model_prefill_tokens_per_second": 25668.600782971705,
"process_vram_peak_bytes": 26977763328.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 2.8539115690886145,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 48.0,
"total_input_tokens": 44927,
"total_output_tokens": 61526
},
"prefill_delta_pct": 112.81221615520849,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 12.209942201949131,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": false,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen4",
"vllm": {
"accuracy": 0.75,
"active_power_mean_w": 585.1145666003976,
"active_power_p95_w": 604.5881,
"aggregate_prefill_tokens_per_second": 24099.77826112465,
"correct": 96,
"decode_tokens_per_second": 334.6523604571347,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3b683bea-442e-2758-038a-a8c289f74656",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.297341194444575,
"latency_ms": {
"p50": 809.6372344880365,
"p95": 4670.765752292937,
"p95_minus_p50": 3861.1285178049,
"p99": 4721.370619542431
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 494.0078125,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 1256.443576787993,
"per_user_model_prefill_tokens_per_second": 12061.619979677787,
"process_vram_peak_bytes": 171523964928.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 2.54336782738227,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 46.0,
"total_input_tokens": 44927,
"total_output_tokens": 63233
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -3.90625,
"concurrency": 8,
"decode_delta_pct": 5.706620338520807,
"model": "Qwen3.5 4B",
"output_delta_pct": 8.949371248392278,
"output_mean_gap_pct": 6.869620153202242,
"photon": {
"accuracy": 0.7421875,
"active_power_mean_w": 680.4242648148148,
"active_power_p95_w": 703.0849999999999,
"aggregate_prefill_tokens_per_second": 55693.288605428716,
"correct": 190,
"decode_tokens_per_second": 302.517401744762,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-caa4aca3-4a4e-c925-1dee-68cc55d12061",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.20161713810737814,
"latency_ms": {
"p50": 814.8620079737157,
"p95": 5202.317798801232,
"p95_minus_p50": 4387.4557908275165,
"p99": 5246.125819627196
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 494.6484375,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 2346.054595814292,
"per_user_model_prefill_tokens_per_second": 28420.359535844305,
"process_vram_peak_bytes": 27191672832.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 4.742872751547491,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 59.0,
"total_input_tokens": 105702,
"total_output_tokens": 126630
},
"prefill_delta_pct": 94.34581430631977,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 256,
"request_delta_pct": 1.4649632843255933,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": false,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen4",
"vllm": {
"accuracy": 0.78125,
"active_power_mean_w": 571.6950784671533,
"active_power_p95_w": 600.0525,
"aggregate_prefill_tokens_per_second": 33944.11710023019,
"correct": 200,
"decode_tokens_per_second": 286.1858611844399,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-caa4aca3-4a4e-c925-1dee-68cc55d12061",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.16850686850012345,
"latency_ms": {
"p50": 899.1244694916531,
"p95": 5459.292831568746,
"p95_minus_p50": 4560.168362077093,
"p99": 5531.91402227385
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 460.66796875,
"model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"output_tokens_per_service_second": 2153.3438595671673,
"per_user_model_prefill_tokens_per_second": 14623.60259071457,
"process_vram_peak_bytes": 171404427264.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 4.674394587082233,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 53.0,
"total_input_tokens": 105702,
"total_output_tokens": 117931
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 2.34375,
"concurrency": 1,
"decode_delta_pct": 30.010040323156062,
"model": "Qwen3.5 9B",
"output_delta_pct": 33.99852625598705,
"output_mean_gap_pct": 3.857705519486359,
"photon": {
"accuracy": 0.796875,
"active_power_mean_w": 929.5731507622812,
"active_power_p95_w": 948.3855,
"aggregate_prefill_tokens_per_second": 25860.35777236419,
"correct": 102,
"decode_tokens_per_second": 305.73389404782284,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-a44201cb-1fef-6391-c186-cb0d047c1c60",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 2.3699737207840115,
"latency_ms": {
"p50": 809.0583495795727,
"p95": 5066.741403250489,
"p95_minus_p50": 4257.683053670917,
"p99": 5100.348583662417
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 417.640625,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 301.78001505791167,
"per_user_model_prefill_tokens_per_second": 23637.03239527488,
"process_vram_peak_bytes": 43031461888.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 0.7225829983802742,
"source_fingerprint": "5d09767acf28428f21ad08776c430e6922b927c8964548eb87e16772cc7d59a0",
"temperature_p95_c": 60.0,
"total_input_tokens": 44927,
"total_output_tokens": 53458
},
"prefill_delta_pct": 270.02983367660664,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 39.37521148212897,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen9",
"vllm": {
"accuracy": 0.7734375,
"active_power_mean_w": 665.5142810854597,
"active_power_p95_w": 691.6486,
"aggregate_prefill_tokens_per_second": 8354.099866205877,
"correct": 99,
"decode_tokens_per_second": 235.16175619043221,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-a44201cb-1fef-6391-c186-cb0d047c1c60",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 2.0246463361136873,
"latency_ms": {
"p50": 1203.099783451762,
"p95": 6604.3188379495405,
"p95_minus_p50": 5401.219054497778,
"p99": 6631.1978773912415
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 434.3984375,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 225.2114433567721,
"per_user_model_prefill_tokens_per_second": 6387.872069778253,
"process_vram_peak_bytes": 171341512704.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 0.5184444139644772,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 56.0,
"total_input_tokens": 44927,
"total_output_tokens": 55603
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 1.5625,
"concurrency": 2,
"decode_delta_pct": 25.194093703703736,
"model": "Qwen3.5 9B",
"output_delta_pct": 27.653264185786732,
"output_mean_gap_pct": 1.269678176947635,
"photon": {
"accuracy": 0.8046875,
"active_power_mean_w": 950.309389937107,
"active_power_p95_w": 965.6444,
"aggregate_prefill_tokens_per_second": 28806.824215437147,
"correct": 103,
"decode_tokens_per_second": 287.42436558520865,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-9e70266c-c50d-7938-4156-a2c7db4c3c43",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.2524737903530085,
"latency_ms": {
"p50": 846.9928574049845,
"p95": 5368.673236342147,
"p95_minus_p50": 4521.6803789371625,
"p99": 5459.580619460903
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 420.390625,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 564.011086973766,
"per_user_model_prefill_tokens_per_second": 21708.95244891651,
"process_vram_peak_bytes": 43031461888.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.3416357393168938,
"source_fingerprint": "5d09767acf28428f21ad08776c430e6922b927c8964548eb87e16772cc7d59a0",
"temperature_p95_c": 66.0,
"total_input_tokens": 44927,
"total_output_tokens": 53810
},
"prefill_delta_pct": 118.44721684284112,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 29.294893229023412,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen9",
"vllm": {
"accuracy": 0.7890625,
"active_power_mean_w": 722.8411605839416,
"active_power_p95_w": 737.6824,
"aggregate_prefill_tokens_per_second": 11844.288371092493,
"correct": 101,
"decode_tokens_per_second": 229.58300753824258,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-9e70266c-c50d-7938-4156-a2c7db4c3c43",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.1102058856979378,
"latency_ms": {
"p50": 1160.6392259709537,
"p95": 6770.406584383454,
"p95_minus_p50": 5609.7673584125005,
"p99": 6865.778888191562
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 425.796875,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 441.8305247196057,
"per_user_model_prefill_tokens_per_second": 9937.848036093186,
"process_vram_peak_bytes": 171301666816.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.0376556303275022,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 61.0,
"total_input_tokens": 44927,
"total_output_tokens": 54502
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -1.5625,
"concurrency": 4,
"decode_delta_pct": 18.35224640641888,
"model": "Qwen3.5 9B",
"output_delta_pct": 22.594918210932203,
"output_mean_gap_pct": 1.1616897305171159,
"photon": {
"accuracy": 0.7734375,
"active_power_mean_w": 861.105924632353,
"active_power_p95_w": 883.2842999999999,
"aggregate_prefill_tokens_per_second": 36258.394110856294,
"correct": 99,
"decode_tokens_per_second": 261.24155136661807,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3bda5f05-a365-0190-c229-bfc74673be59",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.6359303489197351,
"latency_ms": {
"p50": 939.8579314583912,
"p95": 5959.567129320931,
"p95_minus_p50": 5019.709197862539,
"p99": 6008.489653277211
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 429.0625,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 1008.3852246557135,
"per_user_model_prefill_tokens_per_second": 20211.959862131826,
"process_vram_peak_bytes": 43008393216.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 2.350205913254394,
"source_fingerprint": "5d09767acf28428f21ad08776c430e6922b927c8964548eb87e16772cc7d59a0",
"temperature_p95_c": 54.0,
"total_input_tokens": 44927,
"total_output_tokens": 54920
},
"prefill_delta_pct": 114.44769610454553,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 21.170745635939948,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen9",
"vllm": {
"accuracy": 0.7890625,
"active_power_mean_w": 674.7200439393939,
"active_power_p95_w": 694.5207,
"aggregate_prefill_tokens_per_second": 20465.94234261492,
"correct": 101,
"decode_tokens_per_second": 220.7322288328357,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3bda5f05-a365-0190-c229-bfc74673be59",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.5604360464493299,
"latency_ms": {
"p50": 1174.801565532107,
"p95": 7028.190784453182,
"p95_minus_p50": 5853.389218921075,
"p99": 7166.408103093272
},
"mean_input_tokens": 350.9921875,
"mean_output_tokens": 424.078125,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 822.5342774165597,
"per_user_model_prefill_tokens_per_second": 9425.123342093766,
"process_vram_peak_bytes": 171362484224.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.9395819518315398,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 51.0,
"total_input_tokens": 44927,
"total_output_tokens": 54282
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 1.5625,
"concurrency": 8,
"decode_delta_pct": 9.870228302303929,
"model": "Qwen3.5 9B",
"output_delta_pct": 10.9425911062897,
"output_mean_gap_pct": 4.6429336417769775,
"photon": {
"accuracy": 0.78515625,
"active_power_mean_w": 870.7769681020734,
"active_power_p95_w": 897.0839,
"aggregate_prefill_tokens_per_second": 41969.73944857664,
"correct": 201,
"decode_tokens_per_second": 232.1369820266797,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3dbab81a-f32f-2a5d-34db-18e733ee51e0",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.36235908109867904,
"latency_ms": {
"p50": 1130.4141470463946,
"p95": 6709.659235551953,
"p95_minus_p50": 5579.245088505559,
"p99": 6863.346347783226
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 434.75,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 1775.8296285406736,
"per_user_model_prefill_tokens_per_second": 22242.09734668247,
"process_vram_peak_bytes": 43394269184.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 4.084714499230992,
"source_fingerprint": "5d09767acf28428f21ad08776c430e6922b927c8964548eb87e16772cc7d59a0",
"temperature_p95_c": 61.0,
"total_input_tokens": 105702,
"total_output_tokens": 111296
},
"prefill_delta_pct": 87.3018088860438,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 256,
"request_delta_pct": 16.344383634367855,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "qwen9",
"vllm": {
"accuracy": 0.76953125,
"active_power_mean_w": 711.9061152263374,
"active_power_p95_w": 737.795,
"aggregate_prefill_tokens_per_second": 33954.55029901592,
"correct": 197,
"decode_tokens_per_second": 211.2828794602695,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3dbab81a-f32f-2a5d-34db-18e733ee51e0",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.30483702077066976,
"latency_ms": {
"p50": 1267.5742120482028,
"p95": 7358.99924847763,
"p95_minus_p50": 6091.425036429428,
"p99": 7531.591436831513
},
"mean_input_tokens": 412.8984375,
"mean_output_tokens": 455.91796875,
"model_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"output_tokens_per_service_second": 1600.6743765695192,
"per_user_model_prefill_tokens_per_second": 11875.004026370494,
"process_vram_peak_bytes": 171211489280.0,
"request_count": 256,
"request_stream_sha256": "f9516855a032cdf73caedce4f135b684652cdcd470ba0c39ce4d8c8c043530d4",
"requests_per_second": 3.5108824093029765,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 60.0,
"total_input_tokens": 105702,
"total_output_tokens": 116715
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": false,
"accuracy_delta_pp": -71.875,
"concurrency": 1,
"decode_delta_pct": 25.923858951114333,
"model": "Gemma 4 E2B",
"output_delta_pct": 30.724006337702182,
"output_mean_gap_pct": 1.9063032009920342,
"photon": {
"accuracy": 0.0,
"active_power_mean_w": 500.25492791127544,
"active_power_p95_w": 506.55245,
"aggregate_prefill_tokens_per_second": 14810.10301647187,
"correct": 0,
"decode_tokens_per_second": 442.20265913028715,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-785504e5-03fd-7665-e8df-6f381f0ce90f",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.7067175566968952,
"latency_ms": {
"p50": 736.1076704692096,
"p95": 1479.3495318852365,
"p95_minus_p50": 743.241861416027,
"p99": 2250.9498284664005
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 364.625,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 431.1984528436384,
"per_user_model_prefill_tokens_per_second": 19265.386660623706,
"process_vram_peak_bytes": 17920163840.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.1825806043020595,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 46.0,
"total_input_tokens": 38997,
"total_output_tokens": 46672
},
"prefill_delta_pct": 110.05887220402482,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 33.26443044098244,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "gemmae2",
"vllm": {
"accuracy": 0.71875,
"active_power_mean_w": 387.85746393897364,
"active_power_p95_w": 400.01085,
"aggregate_prefill_tokens_per_second": 4837.7963839264585,
"correct": 92,
"decode_tokens_per_second": 351.16669931625694,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-785504e5-03fd-7665-e8df-6f381f0ce90f",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.5929590574045147,
"latency_ms": {
"p50": 982.650906953495,
"p95": 2418.874065962155,
"p95_minus_p50": 1436.2231590086599,
"p99": 3476.209186662458
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 371.7109375,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 329.85406806590214,
"per_user_model_prefill_tokens_per_second": 9171.422496218931,
"process_vram_peak_bytes": 170230022144.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 0.8873940333431866,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 42.0,
"total_input_tokens": 38997,
"total_output_tokens": 47579
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": false,
"accuracy_delta_pp": -71.875,
"concurrency": 2,
"decode_delta_pct": 29.046654595640533,
"model": "Gemma 4 E2B",
"output_delta_pct": 32.889181360025475,
"output_mean_gap_pct": 1.2909905341446923,
"photon": {
"accuracy": 0.0,
"active_power_mean_w": 500.2739112627986,
"active_power_p95_w": 511.11975,
"aggregate_prefill_tokens_per_second": 16192.153781798159,
"correct": 0,
"decode_tokens_per_second": 411.8402555453764,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-23915d64-a606-0b33-a84f-49cf54eef191",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.37441996956370777,
"latency_ms": {
"p50": 783.9619250735268,
"p95": 1692.0025395578707,
"p95_minus_p50": 908.0406144843439,
"p99": 2487.5582748861057
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 364.9765625,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 796.6183419760393,
"per_user_model_prefill_tokens_per_second": 17537.96238601708,
"process_vram_peak_bytes": 17930649600.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 2.182656158848664,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 49.0,
"total_input_tokens": 38997,
"total_output_tokens": 46717
},
"prefill_delta_pct": 82.29783916136641,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 34.627205843853105,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "gemmae2",
"vllm": {
"accuracy": 0.71875,
"active_power_mean_w": 393.615969581749,
"active_power_p95_w": 406.269,
"aggregate_prefill_tokens_per_second": 7003.28992714474,
"correct": 92,
"decode_tokens_per_second": 319.14059053747,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-23915d64-a606-0b33-a84f-49cf54eef191",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.321350513141332,
"latency_ms": {
"p50": 1085.5219829827547,
"p95": 2261.3768783048727,
"p95_minus_p50": 1175.854895322118,
"p99": 2901.6276456264322
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 369.75,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 599.4606436906465,
"per_user_model_prefill_tokens_per_second": 9620.499325004519,
"process_vram_peak_bytes": 166027329536.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.621259347371593,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 45.0,
"total_input_tokens": 38997,
"total_output_tokens": 47328
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": false,
"accuracy_delta_pp": -72.65625,
"concurrency": 4,
"decode_delta_pct": 22.77426913709477,
"model": "Gemma 4 E2B",
"output_delta_pct": 26.241442052524675,
"output_mean_gap_pct": 1.5214155779725484,
"photon": {
"accuracy": 0.0,
"active_power_mean_w": 500.6832587209302,
"active_power_p95_w": 511.60040000000004,
"aggregate_prefill_tokens_per_second": 29563.7594684603,
"correct": 0,
"decode_tokens_per_second": 364.44622133971535,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-8b73f859-fad5-e91d-3880-7e8c8a1a898a",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.2089828146542606,
"latency_ms": {
"p50": 925.6545396056026,
"p95": 2120.883117616177,
"p95_minus_p50": 1195.2285780105744,
"p99": 2767.0983714703475
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 377.9375,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 1403.6365792989266,
"per_user_model_prefill_tokens_per_second": 17142.242873019117,
"process_vram_peak_bytes": 18035507200.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 3.713938360969543,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 44.0,
"total_input_tokens": 38997,
"total_output_tokens": 48376
},
"prefill_delta_pct": 81.67557850635878,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 24.32078508728037,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "gemmae2",
"vllm": {
"accuracy": 0.7265625,
"active_power_mean_w": 385.93309112149535,
"active_power_p95_w": 398.10595,
"aggregate_prefill_tokens_per_second": 11466.98743521101,
"correct": 93,
"decode_tokens_per_second": 296.8425093475896,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-8b73f859-fad5-e91d-3880-7e8c8a1a898a",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.17044133987086924,
"latency_ms": {
"p50": 1198.2174179865979,
"p95": 2551.836267922773,
"p95_minus_p50": 1353.6188499361751,
"p99": 3402.9667292209365
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 372.1875,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 1111.8667186286752,
"per_user_model_prefill_tokens_per_second": 9435.634119871056,
"process_vram_peak_bytes": 163579953152.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 2.987383291025828,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 37.0,
"total_input_tokens": 38997,
"total_output_tokens": 47640
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": false,
"accuracy_delta_pp": -71.875,
"concurrency": 8,
"decode_delta_pct": 6.278553770717954,
"model": "Gemma 4 E2B",
"output_delta_pct": 11.983106271199716,
"output_mean_gap_pct": 0.11004359419308418,
"photon": {
"accuracy": 0.0,
"active_power_mean_w": 495.3612392344497,
"active_power_p95_w": 511.2754,
"aggregate_prefill_tokens_per_second": 43460.90073589207,
"correct": 0,
"decode_tokens_per_second": 301.5562777248356,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-6819489f-ffbc-9caa-ce6c-f7bc63fbf267",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.12610554000391808,
"latency_ms": {
"p50": 1093.6856720363721,
"p95": 2610.8703801757656,
"p95_minus_p50": 1517.1847081393935,
"p99": 3372.952606694307
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 369.171875,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 2260.056953393711,
"per_user_model_prefill_tokens_per_second": 15915.557295741604,
"process_vram_peak_bytes": 18383634432.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 6.121964067261924,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 51.0,
"total_input_tokens": 38997,
"total_output_tokens": 47254
},
"prefill_delta_pct": 102.68335631662491,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 11.859876036169803,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "gemmae2",
"vllm": {
"accuracy": 0.71875,
"active_power_mean_w": 419.7131111111111,
"active_power_p95_w": 428.50155,
"aggregate_prefill_tokens_per_second": 14988.444433695497,
"correct": 92,
"decode_tokens_per_second": 283.74142009440936,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-6819489f-ffbc-9caa-ce6c-f7bc63fbf267",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.10071675204169293,
"latency_ms": {
"p50": 1239.609285024926,
"p95": 2717.2569016343928,
"p95_minus_p50": 1477.6476166094667,
"p99": 3566.8139905750313
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 368.765625,
"model_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
"output_tokens_per_service_second": 2018.2124149336637,
"per_user_model_prefill_tokens_per_second": 7852.424384999266,
"process_vram_peak_bytes": 165989580800.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 5.472886511408605,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 45.0,
"total_input_tokens": 38997,
"total_output_tokens": 47202
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": false,
"accuracy_delta_pp": -75.78125,
"concurrency": 1,
"decode_delta_pct": 22.842488996050236,
"model": "Gemma 4 E4B",
"output_delta_pct": 27.747350528712733,
"output_mean_gap_pct": 0.383514437373204,
"photon": {
"accuracy": 0.0546875,
"active_power_mean_w": 618.4702113360324,
"active_power_p95_w": 627.85,
"aggregate_prefill_tokens_per_second": 14865.943189187692,
"correct": 7,
"decode_tokens_per_second": 297.04589394069865,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3b683bea-442e-2758-038a-a8c289f74656",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.4340331499991568,
"latency_ms": {
"p50": 903.1670185504481,
"p95": 1862.4719202169222,
"p95_minus_p50": 959.304901666474,
"p99": 3144.2583910631993
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 280.0390625,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 290.22756763765466,
"per_user_model_prefill_tokens_per_second": 17974.57183736164,
"process_vram_peak_bytes": 30993809408.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.0363824426731705,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 50.0,
"total_input_tokens": 38997,
"total_output_tokens": 35845
},
"prefill_delta_pct": 126.18655674708354,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 28.239166245631743,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "gemmae4",
"vllm": {
"accuracy": 0.8125,
"active_power_mean_w": 481.91670580808085,
"active_power_p95_w": 503.0721,
"aggregate_prefill_tokens_per_second": 5026.591088977804,
"correct": 104,
"decode_tokens_per_second": 241.8103836615111,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-3b683bea-442e-2758-038a-a8c289f74656",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 1.2092755064084646,
"latency_ms": {
"p50": 1141.285497986246,
"p95": 2321.2325231346776,
"p95_minus_p50": 1179.9470251484317,
"p99": 5178.760194020585
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 281.1171875,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 227.1887177592952,
"per_user_model_prefill_tokens_per_second": 7946.790514814008,
"process_vram_peak_bytes": 170322296832.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 0.8081637404660474,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 44.0,
"total_input_tokens": 38997,
"total_output_tokens": 35983
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": false,
"accuracy_delta_pp": -76.5625,
"concurrency": 2,
"decode_delta_pct": 17.63539509553631,
"model": "Gemma 4 E4B",
"output_delta_pct": 21.104272510590594,
"output_mean_gap_pct": 0.22127031857252277,
"photon": {
"accuracy": 0.0546875,
"active_power_mean_w": 586.6381203966006,
"active_power_p95_w": 600.0285,
"aggregate_prefill_tokens_per_second": 15447.27365813537,
"correct": 7,
"decode_tokens_per_second": 257.2302184977964,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-caa4aca3-4a4e-c925-1dee-68cc55d12061",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.7682504272836969,
"latency_ms": {
"p50": 1052.7166540268809,
"p95": 1960.137265850791,
"p95_minus_p50": 907.4206118239101,
"p99": 2548.967027941255
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 274.7890625,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 498.22561293812686,
"per_user_model_prefill_tokens_per_second": 14962.898097761927,
"process_vram_peak_bytes": 30998003712.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.8131202472373764,
"source_fingerprint": "5c5eda395c7652445fe69fe25d15218be6a33492191643ca6d053e63066a95de",
"temperature_p95_c": 59.0,
"total_input_tokens": 38997,
"total_output_tokens": 35173
},
"prefill_delta_pct": 73.88962233026398,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"request_count": 128,
"request_delta_pct": 21.37283456830039,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be",
"source_commit_exact": true,
"tag": "gemmae4",
"vllm": {
"accuracy": 0.8203125,
"active_power_mean_w": 481.16680163360564,
"active_power_p95_w": 491.0008,
"aggregate_prefill_tokens_per_second": 6738.588698946801,
"correct": 105,
"decode_tokens_per_second": 218.66736477476837,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-caa4aca3-4a4e-c925-1dee-68cc55d12061",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.6672179966838432,
"latency_ms": {
"p50": 1290.1016450487077,
"p95": 2371.736789983698,
"p95_minus_p50": 1081.6351449349904,
"p99": 4416.473485321044
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 275.3984375,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 411.4021764959258,
"per_user_model_prefill_tokens_per_second": 8604.825231803246,
"process_vram_peak_bytes": 166125895680.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 1.4938435389486397,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 51.0,
"total_input_tokens": 38997,
"total_output_tokens": 35251
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": 3.125,
"concurrency": 4,
"decode_delta_pct": 0.9182623683115176,
"model": "Gemma 4 E4B",
"output_delta_pct": 4.264614563809133,
"output_mean_gap_pct": 2.0610012156039343,
"photon": {
"accuracy": 0.828125,
"active_power_mean_w": 556.7003887468031,
"active_power_p95_w": 571.1725,
"aggregate_prefill_tokens_per_second": 26745.715410367073,
"correct": 106,
"decode_tokens_per_second": 235.0663460013063,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-785504e5-03fd-7665-e8df-6f381f0ce90f",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.3958987520287473,
"latency_ms": {
"p50": 1145.0920845381916,
"p95": 2420.5717907985677,
"p95_minus_p50": 1275.4797062603761,
"p99": 4194.041937044829
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 276.953125,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 905.1398381686105,
"per_user_model_prefill_tokens_per_second": 14530.760299358903,
"process_vram_peak_bytes": 31000100864.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 3.2682059036835582,
"source_fingerprint": "192f568a6b97fb7fc61c57163d90211c2d13b56d5a5ef1fe3d8adbf1740bfd73",
"temperature_p95_c": 49.0,
"total_input_tokens": 38997,
"total_output_tokens": 35450
},
"prefill_delta_pct": 70.17737970646544,
"public_commit": "b4c20e2cb04b4839a0cc063e636d009be704d272",
"request_count": 128,
"request_delta_pct": 6.4587302891857545,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": true,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "5c06c09047d830d8008c16e9eb275a9f55bf5714",
"source_commit_exact": true,
"tag": "gemmae4",
"vllm": {
"accuracy": 0.796875,
"active_power_mean_w": 485.937757793765,
"active_power_p95_w": 495.8428,
"aggregate_prefill_tokens_per_second": 10930.404104999921,
"correct": 102,
"decode_tokens_per_second": 232.92746078346815,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-785504e5-03fd-7665-e8df-6f381f0ce90f",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.3359568772951413,
"latency_ms": {
"p50": 1228.8938495330513,
"p95": 2477.1565746807037,
"p95_minus_p50": 1248.2627251476524,
"p99": 4336.4511709089875
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 282.78125,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 868.11795349291,
"per_user_model_prefill_tokens_per_second": 8538.596800833715,
"process_vram_peak_bytes": 163582050304.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 3.0699275623575115,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 43.0,
"total_input_tokens": 38997,
"total_output_tokens": 36196
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
},
{
"accuracy_comparable": true,
"accuracy_delta_pp": -0.78125,
"concurrency": 8,
"decode_delta_pct": -2.3713008644001166,
"model": "Gemma 4 E4B",
"output_delta_pct": 1.9293518598028747,
"output_mean_gap_pct": 2.594618985842405,
"photon": {
"accuracy": 0.828125,
"active_power_mean_w": 604.2621520737326,
"active_power_p95_w": 622.6712,
"aggregate_prefill_tokens_per_second": 34775.819928323916,
"correct": 106,
"decode_tokens_per_second": 209.25814196657547,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-6819489f-ffbc-9caa-ce6c-f7bc63fbf267",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.24085726397848506,
"latency_ms": {
"p50": 1301.4992920216173,
"p95": 2573.6308254767214,
"p95_minus_p50": 1272.131533455104,
"p99": 3244.6758230170235
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 269.828125,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 1590.3778357287285,
"per_user_model_prefill_tokens_per_second": 13082.39819229523,
"process_vram_peak_bytes": 31329353728.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 5.894040273706563,
"source_fingerprint": "6024600b0b129af2f956030c6a2f2df862b09c1f99680202d4b80664ae22c83b",
"temperature_p95_c": 54.0,
"total_input_tokens": 38997,
"total_output_tokens": 34538
},
"prefill_delta_pct": 72.96303062628122,
"public_commit": "3a85db8fc74c87701f0b9edc97621338fc896efc",
"request_count": 128,
"request_delta_pct": 4.644477336408914,
"same_driver": true,
"same_gpu_model": true,
"same_gpu_uuid": false,
"same_harness": true,
"same_input_tokens": true,
"same_model_fingerprint": true,
"same_model_revision": true,
"same_request_count": true,
"same_request_stream": true,
"source_commit": "e1d6c1dd7fd094303c6cf7d18fd867b03f53016b",
"source_commit_exact": true,
"tag": "gemmae4",
"vllm": {
"accuracy": 0.8359375,
"active_power_mean_w": 493.106436123348,
"active_power_p95_w": 507.5383,
"aggregate_prefill_tokens_per_second": 14844.865809409215,
"correct": 107,
"decode_tokens_per_second": 214.3408073848547,
"driver_version": "580.95.05",
"gpu_uuid": "GPU-23915d64-a606-0b33-a84f-49cf54eef191",
"harness_sha256": "499a39f597dfc64cfb96b41eef82c3e28cf491418908e63a73563d47112aff92",
"idle_subtracted_joules_per_output_token": 0.18781059509428108,
"latency_ms": {
"p50": 1322.069188056048,
"p95": 2629.169589933009,
"p95_minus_p50": 1307.1004018769609,
"p99": 4310.0025493337325
},
"mean_input_tokens": 304.6640625,
"mean_output_tokens": 277.015625,
"model_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
"output_tokens_per_service_second": 1560.274647793492,
"per_user_model_prefill_tokens_per_second": 7563.696210065944,
"process_vram_peak_bytes": 165987483648.0,
"request_count": 128,
"request_stream_sha256": "799b5eef05ff402c06c7878a21aa9d434b1e2da2ba52331c55ee761f2752f1dd",
"requests_per_second": 5.632442746843222,
"source_fingerprint": "42df5872580ae3e742159ca7465b1a375ff289ad11ffab8b4f18d5ba938e7252",
"temperature_p95_c": 45.0,
"total_input_tokens": 38997,
"total_output_tokens": 35458
},
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411"
}
],
"provenance": {
"accuracy_comparable_cells": 22,
"board_map_sha256": "d18425b0b21fdf96d3a1a041d14504a8694e7aa1cbcf3a956b2c85cc20aca1ee",
"bundle_a_sha256": "396d2698564f29b65385a0bafaf141df6c87c73fe60e144c5a0e6e219e66cafd",
"bundle_b_sha256": "6f7f3ee9e83507b1b6f75c6f4ecbed0a929de4836f16da4b0c5c8e6b9f079a59",
"c4_public_commit": "b4c20e2cb04b4839a0cc063e636d009be704d272",
"c4_public_tree": "30e07fd1ccdd4c9d4dbcf400084bf94dd0ad6434",
"c4_source_commit": "5c06c09047d830d8008c16e9eb275a9f55bf5714",
"c4_source_tree": "ce9b25ce831a1e88da87bf13a464716da5245b51",
"c8_public_commit": "3a85db8fc74c87701f0b9edc97621338fc896efc",
"c8_source_commit": "e1d6c1dd7fd094303c6cf7d18fd867b03f53016b",
"c8_source_tree": "aa8e173b26a1988d52f1bdc5bdfc5be2e1ee652c",
"decode_throughput_wins": 27,
"measured_cells": 28,
"model_prefill_wins": 28,
"quality_warning_cells": 6,
"request_count_per_arm": {
"default": 128,
"qwen08-c8": 256,
"qwen2-c8": 256,
"qwen4-c8": 256,
"qwen9-c8": 256
},
"request_throughput_wins": 28,
"same_driver_pairs": 28,
"same_gpu_pairs": 27,
"same_harness_pairs": 28,
"same_input_token_pairs": 28,
"same_model_fingerprint_pairs": 26,
"same_request_stream_pairs": 28,
"service_throughput_wins": 28,
"source_cohorts": [
{
"cells": 1,
"public_commit": "b4c20e2cb04b4839a0cc063e636d009be704d272",
"runtime_commit": "5c06c09047d830d8008c16e9eb275a9f55bf5714"
},
{
"cells": 26,
"public_commit": "f49bd3812197132fe9bd60f27363f4c4c140adcb",
"runtime_commit": "7f0dddc232ea9a3406108bbaea96edfda62ec0be"
},
{
"cells": 1,
"public_commit": "3a85db8fc74c87701f0b9edc97621338fc896efc",
"runtime_commit": "e1d6c1dd7fd094303c6cf7d18fd867b03f53016b"
}
],
"vllm_commit": "e3acdf4951dabbe2369fa03dc5b90e1da1408411",
"vllm_version": "0.27.1",
"wheel_sha256": "38ecc42ad852ae9d356ba4680d3777bc04810faa1d73866fbf2f632b7f63c5ff"
},
"schema_version": 2,
"status": "final_2026_08_26_shipping_board",
"validation_findings": [
{
"cell": "gemmae4-c8",
"concurrency": 8,
"decode_delta_pct": -2.3713008644001166,
"model": "Gemma 4 E4B",
"type": "decode_throughput_nonwin"
},
{
"accuracy_delta_pp": 3.125,
"cell": "qwen08-c1",
"concurrency": 1,
"model": "Qwen3.5 0.8B",
"output_mean_gap_pct": 12.592793280953671,
"type": "quality_warning"
},
{
"accuracy_delta_pp": 3.125,
"cell": "qwen08-c4",
"concurrency": 4,
"model": "Qwen3.5 0.8B",
"output_mean_gap_pct": 4.683777344596058,
"type": "quality_warning"
},
{
"accuracy_delta_pp": -3.125,
"cell": "qwen4-c1",
"concurrency": 1,
"model": "Qwen3.5 4B",
"output_mean_gap_pct": 1.06842680089273,
"type": "quality_warning"
},
{
"accuracy_delta_pp": -3.90625,
"cell": "qwen4-c8",
"concurrency": 8,
"model": "Qwen3.5 4B",
"output_mean_gap_pct": 6.869620153202242,
"type": "quality_warning"
},
{
"accuracy_delta_pp": 2.34375,
"cell": "qwen9-c1",
"concurrency": 1,
"model": "Qwen3.5 9B",
"output_mean_gap_pct": 3.857705519486359,
"type": "quality_warning"
},
{
"accuracy_delta_pp": 3.125,
"cell": "gemmae4-c4",
"concurrency": 4,
"model": "Gemma 4 E4B",
"output_mean_gap_pct": 2.0610012156039343,
"type": "quality_warning"
}
]
}
7df3695fe7cd4ec7d9c4964f7aace36da8e993584c515c040d1b63c22a484da5 README.md
03214b37a46128144c5bd875f2a01c41fd8c0fcbce4e45b7d6aee570fbd8c21a report-data.json
7ddccb5f27f1960f100287149ec68b308302c77348f49e0c5af106d7480fe618 generate.py
05d53b70b24df746e8e4ebce56a81ac55233c2563c46da0fb95015addd078c83 throughput-by-concurrency.png
e16965c7848f5b7b07889779dbb609a520a069cdfde7f99a6aa25de264120e2a speedup-vs-vllm.png
7ffcb15a6e691b575790a55e33d3b13d1d59f4d3895ae514d5889301e67b42b4 model-prefill-throughput.png
ecd673958e17395c0bef6b8ea58a62598708bf17d8ffa09e9d186aa6ce46828a model-prefill-speedup.png
5d0823fa6ba2365410067b14e8f3cb99ba0f2f8cc4def5395ebdb9a51f50e089 decode-scaling.png
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment