Status: Final 28-cell board against vLLM 0.27.1. The board retains 26 validated, unchanged cells from
7f0dddc2, retains the frozen Gemma E4B C8 measurement frome1d6c1dd, and replaces Gemma E4B C4 with the final fast/correct release measurement from5c06c090. Every Photon/vLLM pair is internally matched on driver, model revision, harness, ordered request stream, request count, and B200 hardware class; exact engine commits are recorded per cell.
ChartQA multimodal inference with reasoning enabled on NVIDIA B200. The primary metric is natural-output tokens completed per active end-to-end service second; accuracy and output-length differences are reported alongside it.
- Photon leads in 28/28 end-to-end output-token service-throughput cells.
- Photon leads in 28/28 completed-request-rate cells.
- Photon leads in 28/28 timing-matched model-prefill cells.
- Photon leads in 27/28 timing-matched decode cells.
- The 1 timing-matched decode non-win is Gemma 4 E4B C8 (-2.4%).
- Pair checks: request stream 28/28, input tokens 28/28, model revision 28/28, engine-side model fingerprint 26/28, driver 28/28, exact GPU UUID 27/28, matching harness hash 28/28.
- 6 comparable cells cross the report warning threshold of 2 accuracy points or 10% mean-output-length difference; 6 retained Gemma rows use the superseded parser and are excluded from accuracy comparisons.
- 24 cells use 128 requests per arm; Qwen 0.8B/2B/4B/9B C8 use 256 to satisfy the declared-concurrency floor.
- The Kestrel 0.6.1 release tree passed its full CPU suite (839 passed, 46 hardware-dependent skips); focused query, streaming, and scheduler regressions also passed.
| Model | C | Photon output tok/s | vLLM output tok/s | Service delta | Req/s delta | Decode delta | Accuracy P / V |
|---|---|---|---|---|---|---|---|
| Moondream 3 | 1 | 300.6 | 117.6 | +155.6% | +140.6% | +59.0% | 85.2% / 83.6% |
| Moondream 3 | 2 | 453.6 | 214.1 | +111.8% | +108.4% | +61.8% | 85.2% / 83.6% |
| Moondream 3 | 4 | 662.5 | 353.1 | +87.6% | +82.1% | +57.3% | 84.4% / 83.6% |
| Moondream 3 | 8 | 906.4 | 458.8 | +97.6% | +92.4% | +60.3% | 84.4% / 84.4% |
| Qwen3.5 0.8B | 1 | 1,189.4 | 622.5 | +91.1% | +118.6% | +83.2% | 75.0% / 71.9% |
| Qwen3.5 0.8B | 2 | 2,016.9 | 1,188.5 | +69.7% | +71.7% | +64.0% | 71.9% / 72.7% |
| Qwen3.5 0.8B | 4 | 3,230.1 | 2,167.1 | +49.0% | +56.4% | +39.7% | 76.6% / 73.4% |
| Qwen3.5 0.8B | 8 | 4,875.7 | 3,997.4 | +22.0% | +20.6% | +15.6% | 73.0% / 73.0% |
| Qwen3.5 2B | 1 | 822.7 | 520.0 | +58.2% | +56.4% | +50.3% | 75.0% / 76.6% |
| Qwen3.5 2B | 2 | 1,462.4 | 953.4 | +53.4% | +52.1% | +49.6% | 74.2% / 75.0% |
| Qwen3.5 2B | 4 | 2,430.3 | 1,769.4 | +37.3% | +37.3% | +29.2% | 74.2% / 75.8% |
| Qwen3.5 2B | 8 | 3,859.4 | 3,284.8 | +17.5% | +25.6% | +14.1% | 75.8% / 75.4% |
| Qwen3.5 4B | 1 | 458.2 | 314.3 | +45.8% | +44.2% | +40.6% | 72.7% / 75.8% |
| Qwen3.5 4B | 2 | 800.9 | 621.4 | +28.9% | +27.8% | +25.8% | 75.0% / 75.0% |
| Qwen3.5 4B | 4 | 1,371.8 | 1,256.4 | +9.2% | +12.2% | +6.3% | 75.0% / 75.0% |
| Qwen3.5 4B | 8 | 2,346.1 | 2,153.3 | +8.9% | +1.5% | +5.7% | 74.2% / 78.1% |
| Qwen3.5 9B | 1 | 301.8 | 225.2 | +34.0% | +39.4% | +30.0% | 79.7% / 77.3% |
| Qwen3.5 9B | 2 | 564.0 | 441.8 | +27.7% | +29.3% | +25.2% | 80.5% / 78.9% |
| Qwen3.5 9B | 4 | 1,008.4 | 822.5 | +22.6% | +21.2% | +18.4% | 77.3% / 78.9% |
| Qwen3.5 9B | 8 | 1,775.8 | 1,600.7 | +10.9% | +16.3% | +9.9% | 78.5% / 77.0% |
| Gemma 4 E2B | 1 | 431.2 | 329.9 | +30.7% | +33.3% | +25.9% | n/a (legacy parser) |
| Gemma 4 E2B | 2 | 796.6 | 599.5 | +32.9% | +34.6% | +29.0% | n/a (legacy parser) |
| Gemma 4 E2B | 4 | 1,403.6 | 1,111.9 | +26.2% | +24.3% | +22.8% | n/a (legacy parser) |
| Gemma 4 E2B | 8 | 2,260.1 | 2,018.2 | +12.0% | +11.9% | +6.3% | n/a (legacy parser) |
| Gemma 4 E4B | 1 | 290.2 | 227.2 | +27.7% | +28.2% | +22.8% | n/a (legacy parser) |
| Gemma 4 E4B | 2 | 498.2 | 411.4 | +21.1% | +21.4% | +17.6% | n/a (legacy parser) |
| Gemma 4 E4B | 4 | 905.1 | 868.1 | +4.3% | +6.5% | +0.9% | 82.8% / 79.7% |
| Gemma 4 E4B | 8 | 1,590.4 | 1,560.3 | +1.9% | +4.6% | -2.4% | 82.8% / 83.6% |
Model-prefill tok/s is the timing-matched model prefill rate recorded by the service run. Prefix caching is disabled.
| Model | C | Photon prefill tok/s | vLLM prefill tok/s | Photon vs vLLM |
|---|---|---|---|---|
| Moondream 3 | 1 | 61,406.9 | 19,597.6 | +213.3% |
| Moondream 3 | 2 | 46,835.7 | 22,150.2 | +111.4% |
| Moondream 3 | 4 | 39,525.6 | 19,838.4 | +99.2% |
| Moondream 3 | 8 | 30,384.0 | 14,310.9 | +112.3% |
| Qwen3.5 0.8B | 1 | 28,463.4 | 11,714.8 | +143.0% |
| Qwen3.5 0.8B | 2 | 28,176.1 | 19,057.9 | +47.8% |
| Qwen3.5 0.8B | 4 | 25,433.4 | 18,375.3 | +38.4% |
| Qwen3.5 0.8B | 8 | 28,665.2 | 19,884.5 | +44.2% |
| Qwen3.5 2B | 1 | 24,443.3 | 10,526.0 | +132.2% |
| Qwen3.5 2B | 2 | 23,639.6 | 17,126.0 | +38.0% |
| Qwen3.5 2B | 4 | 23,502.8 | 15,089.8 | +55.8% |
| Qwen3.5 2B | 8 | 25,487.6 | 18,308.3 | +39.2% |
| Qwen3.5 4B | 1 | 27,866.2 | 7,162.8 | +289.0% |
| Qwen3.5 4B | 2 | 26,453.5 | 11,787.5 | +124.4% |
| Qwen3.5 4B | 4 | 25,668.6 | 12,061.6 | +112.8% |
| Qwen3.5 4B | 8 | 28,420.4 | 14,623.6 | +94.3% |
| Qwen3.5 9B | 1 | 23,637.0 | 6,387.9 | +270.0% |
| Qwen3.5 9B | 2 | 21,709.0 | 9,937.8 | +118.4% |
| Qwen3.5 9B | 4 | 20,212.0 | 9,425.1 | +114.4% |
| Qwen3.5 9B | 8 | 22,242.1 | 11,875.0 | +87.3% |
| Gemma 4 E2B | 1 | 19,265.4 | 9,171.4 | +110.1% |
| Gemma 4 E2B | 2 | 17,538.0 | 9,620.5 | +82.3% |
| Gemma 4 E2B | 4 | 17,142.2 | 9,435.6 | +81.7% |
| Gemma 4 E2B | 8 | 15,915.6 | 7,852.4 | +102.7% |
| Gemma 4 E4B | 1 | 17,974.6 | 7,946.8 | +126.2% |
| Gemma 4 E4B | 2 | 14,962.9 | 8,604.8 | +73.9% |
| Gemma 4 E4B | 4 | 14,530.8 | 8,538.6 | +70.2% |
| Gemma 4 E4B | 8 | 13,082.4 | 7,563.7 | +73.0% |
The horizontal axis is aggregate end-to-end output tok/s and includes prefill. The vertical axis is decode tok/s per concurrent user and excludes prefill. Higher and farther right are better, but the axes are not commensurable.
| Model | C | Photon p50 / p95 / p99 ms | vLLM p50 / p95 / p99 ms | Photon p95-p50 | vLLM p95-p50 |
|---|---|---|---|---|---|
| Moondream 3 | 1 | 51.4 / 89.2 / 124.0 | 132.6 / 187.8 / 228.8 | 37.8 | 55.2 |
| Moondream 3 | 2 | 66.9 / 113.0 / 135.0 | 141.5 / 232.4 / 287.3 | 46.0 | 90.9 |
| Moondream 3 | 4 | 92.0 / 168.0 / 226.6 | 172.3 / 270.7 / 399.0 | 76.1 | 98.4 |
| Moondream 3 | 8 | 133.5 / 246.5 / 279.6 | 238.6 / 696.3 / 876.6 | 113.0 | 457.7 |
| Qwen3.5 0.8B | 1 | 205.5 / 1247.0 / 1267.1 | 413.9 / 2304.8 / 2323.5 | 1041.5 | 1890.9 |
| Qwen3.5 0.8B | 2 | 247.2 / 1485.2 / 1516.8 | 440.7 / 2425.1 / 2469.1 | 1238.0 | 1984.4 |
| Qwen3.5 0.8B | 4 | 287.7 / 1818.6 / 1938.5 | 469.5 / 2563.7 / 2671.2 | 1530.9 | 2094.2 |
| Qwen3.5 0.8B | 8 | 422.3 / 2420.2 / 2478.9 | 523.8 / 2837.8 / 2924.7 | 1997.9 | 2314.0 |
| Qwen3.5 2B | 1 | 265.7 / 1820.5 / 1840.2 | 490.1 / 2704.8 / 2794.6 | 1554.8 | 2214.7 |
| Qwen3.5 2B | 2 | 314.4 / 2047.2 / 2106.2 | 555.6 / 3042.3 / 3094.9 | 1732.8 | 2486.6 |
| Qwen3.5 2B | 4 | 367.0 / 2457.0 / 2568.1 | 575.2 / 3187.0 / 3289.6 | 2090.0 | 2611.7 |
| Qwen3.5 2B | 8 | 449.2 / 3018.8 / 3200.5 | 587.7 / 3514.8 / 3613.7 | 2569.6 | 2927.1 |
| Qwen3.5 4B | 1 | 540.4 / 3352.0 / 3357.3 | 799.6 / 4715.6 / 4720.8 | 2811.6 | 3916.0 |
| Qwen3.5 4B | 2 | 617.0 / 3821.3 / 3849.4 | 847.5 / 4811.1 / 4857.2 | 3204.3 | 3963.6 |
| Qwen3.5 4B | 4 | 662.1 / 4365.0 / 4433.0 | 809.6 / 4670.8 / 4721.4 | 3702.9 | 3861.1 |
| Qwen3.5 4B | 8 | 814.9 / 5202.3 / 5246.1 | 899.1 / 5459.3 / 5531.9 | 4387.5 | 4560.2 |
| Qwen3.5 9B | 1 | 809.1 / 5066.7 / 5100.3 | 1203.1 / 6604.3 / 6631.2 | 4257.7 | 5401.2 |
| Qwen3.5 9B | 2 | 847.0 / 5368.7 / 5459.6 | 1160.6 / 6770.4 / 6865.8 | 4521.7 | 5609.8 |
| Qwen3.5 9B | 4 | 939.9 / 5959.6 / 6008.5 | 1174.8 / 7028.2 / 7166.4 | 5019.7 | 5853.4 |
| Qwen3.5 9B | 8 | 1130.4 / 6709.7 / 6863.3 | 1267.6 / 7359.0 / 7531.6 | 5579.2 | 6091.4 |
| Gemma 4 E2B | 1 | 736.1 / 1479.3 / 2250.9 | 982.7 / 2418.9 / 3476.2 | 743.2 | 1436.2 |
| Gemma 4 E2B | 2 | 784.0 / 1692.0 / 2487.6 | 1085.5 / 2261.4 / 2901.6 | 908.0 | 1175.9 |
| Gemma 4 E2B | 4 | 925.7 / 2120.9 / 2767.1 | 1198.2 / 2551.8 / 3403.0 | 1195.2 | 1353.6 |
| Gemma 4 E2B | 8 | 1093.7 / 2610.9 / 3373.0 | 1239.6 / 2717.3 / 3566.8 | 1517.2 | 1477.6 |
| Gemma 4 E4B | 1 | 903.2 / 1862.5 / 3144.3 | 1141.3 / 2321.2 / 5178.8 | 959.3 | 1179.9 |
| Gemma 4 E4B | 2 | 1052.7 / 1960.1 / 2549.0 | 1290.1 / 2371.7 / 4416.5 | 907.4 | 1081.6 |
| Gemma 4 E4B | 4 | 1145.1 / 2420.6 / 4194.0 | 1228.9 / 2477.2 / 4336.5 | 1275.5 | 1248.3 |
| Gemma 4 E4B | 8 | 1301.5 / 2573.6 / 3244.7 | 1322.1 / 2629.2 / 4310.0 | 1272.1 | 1307.1 |
Accuracy and generation length are measured outcomes. Six retained Gemma r40 rows used the superseded answer parser and are excluded instead of publishing their invalid scores. A warning means an absolute accuracy gap of at least 2 percentage points or a mean-output-length gap of at least 10%.
| Model | C | Photon correct | vLLM correct | Accuracy delta | Mean output P / V | Assessment |
|---|---|---|---|---|---|---|
| Moondream 3 | 1 | 109/128 | 107/128 | +1.56 pp | 17.1 / 16.1 | within threshold |
| Moondream 3 | 2 | 109/128 | 107/128 | +1.56 pp | 16.5 / 16.3 | within threshold |
| Moondream 3 | 4 | 108/128 | 107/128 | +0.78 pp | 16.9 / 16.4 | within threshold |
| Moondream 3 | 8 | 108/128 | 108/128 | +0.00 pp | 16.8 / 16.4 | within threshold |
| Qwen3.5 0.8B | 1 | 96/128 | 92/128 | +3.12 pp | 378.1 / 432.5 | warning |
| Qwen3.5 0.8B | 2 | 92/128 | 93/128 | -0.78 pp | 429.6 / 434.7 | within threshold |
| Qwen3.5 0.8B | 4 | 98/128 | 94/128 | +3.12 pp | 412.6 / 432.8 | warning |
| Qwen3.5 0.8B | 8 | 187/256 | 187/256 | +0.00 pp | 429.7 / 424.8 | within threshold |
| Qwen3.5 2B | 1 | 96/128 | 98/128 | -1.56 pp | 419.4 / 414.7 | within threshold |
| Qwen3.5 2B | 2 | 95/128 | 96/128 | -0.78 pp | 423.3 / 419.8 | within threshold |
| Qwen3.5 2B | 4 | 95/128 | 97/128 | -1.56 pp | 425.5 / 425.2 | within threshold |
| Qwen3.5 2B | 8 | 194/256 | 193/256 | +0.39 pp | 413.6 / 442.1 | within threshold |
| Qwen3.5 4B | 1 | 93/128 | 97/128 | -3.12 pp | 493.6 / 488.3 | warning |
| Qwen3.5 4B | 2 | 96/128 | 96/128 | +0.00 pp | 492.1 / 488.1 | within threshold |
| Qwen3.5 4B | 4 | 96/128 | 96/128 | +0.00 pp | 480.7 / 494.0 | within threshold |
| Qwen3.5 4B | 8 | 190/256 | 200/256 | -3.91 pp | 494.6 / 460.7 | warning |
| Qwen3.5 9B | 1 | 102/128 | 99/128 | +2.34 pp | 417.6 / 434.4 | warning |
| Qwen3.5 9B | 2 | 103/128 | 101/128 | +1.56 pp | 420.4 / 425.8 | within threshold |
| Qwen3.5 9B | 4 | 99/128 | 101/128 | -1.56 pp | 429.1 / 424.1 | within threshold |
| Qwen3.5 9B | 8 | 201/256 | 197/256 | +1.56 pp | 434.8 / 455.9 | within threshold |
| Gemma 4 E4B | 4 | 106/128 | 102/128 | +3.12 pp | 277.0 / 282.8 | warning |
| Gemma 4 E4B | 8 | 106/128 | 107/128 | -0.78 pp | 269.8 / 277.0 | within threshold |
These are measured output differences, not token-normalized quality comparisons. The artifact ledger preserves the exact natural outputs and reports every threshold crossing.
| Model | Backend | Peak process VRAM | Power mean / p95 | Temperature p95 | Incremental J/output token |
|---|---|---|---|---|---|
| Moondream 3 | Photon | 18.90 GiB | 451.6 / 496.6 W | 43 C | 0.813 |
| Moondream 3 | vLLM | 161.46 GiB | 345.2 / 359.7 W | 39 C | 1.145 |
| Qwen3.5 0.8B | Photon | 6.64 GiB | 500.1 / 518.0 W | 45 C | 0.249 |
| Qwen3.5 0.8B | vLLM | 159.92 GiB | 341.6 / 354.8 W | 41 C | 0.221 |
| Qwen3.5 2B | Photon | 15.09 GiB | 665.8 / 687.4 W | 53 C | 0.568 |
| Qwen3.5 2B | vLLM | 159.66 GiB | 441.2 / 473.1 W | 45 C | 0.468 |
| Qwen3.5 4B | Photon | 25.05 GiB | 807.4 / 822.4 W | 57 C | 1.283 |
| Qwen3.5 4B | vLLM | 159.69 GiB | 559.7 / 583.5 W | 52 C | 1.096 |
| Qwen3.5 9B | Photon | 40.08 GiB | 929.6 / 948.4 W | 60 C | 2.370 |
| Qwen3.5 9B | vLLM | 159.57 GiB | 665.5 / 691.6 W | 56 C | 2.025 |
| Gemma 4 E2B | Photon | 16.69 GiB | 500.3 / 506.6 W | 46 C | 0.707 |
| Gemma 4 E2B | vLLM | 158.54 GiB | 387.9 / 400.0 W | 42 C | 0.593 |
| Gemma 4 E4B | Photon | 28.87 GiB | 618.5 / 627.9 W | 50 C | 1.434 |
| Gemma 4 E4B | vLLM | 158.62 GiB | 481.9 / 503.1 W | 44 C | 1.209 |
- NVIDIA B200, tensor parallelism 1, ChartQA test split, greedy decoding, reasoning enabled, no prefix cache.
- 24 cells contain 128 requests per arm. Qwen 0.8B/2B/4B/9B C8 contain 256 requests per arm to remove finite-window boundary underfill.
- The final board uses 26 unchanged validated rows from runtime
7f0dddc232ea9a3406108bbaea96edfda62ec0beplus public runtimef49bd3812197132fe9bd60f27363f4c4c140adcb, Gemma E4B C8 from runtimee1d6c1dd7fd094303c6cf7d18fd867b03f53016bplus parser3a85db8fc74c87701f0b9edc97621338fc896efc, and Gemma E4B C4 from runtime5c06c09047d830d8008c16e9eb275a9f55bf5714plus parserb4c20e2cb04b4839a0cc063e636d009be704d272. The baseline is vLLM 0.27.1 ate3acdf4951dabbe2369fa03dc5b90e1da1408411. - Gemma E4B C8 is the predeclared median of three exact repeats; the selected run is the median, not the best run. Its Photon repeat and retained vLLM baseline use separate B200s on the same host; all other pairs match exact GPU UUID.
- Every pair uses a byte-identical ordered request stream and the same model revision, driver, harness, and B200 hardware class.
- Photon is in-process; vLLM is served out-of-process over HTTP. The chart reports each engine in its native deployment shape.
- Service throughput is completed natural-output tokens divided by the closed-loop active service wall. Request rate uses completed requests over the same wall. Latency is client-observed end-to-end time.
- Model-prefill throughput is the timing-matched model prefill rate from the natural-output service run; prefix caching is disabled.
- Decode tok/s counts tokens after the first token over first-token-to-completion time. The first token is produced by prefill.
- Final board: 26 unchanged validated rows and the frozen Gemma E4B C8 row are retained; Gemma E4B C4 is replaced by the final fast/correct release measurement.
- Pair fairness: the report rejects any pair whose B200 hardware class, driver, model revision, harness, ordered request stream, or request count differs; exact GPU UUID is also reported.
- Natural outputs: accuracy, token counts, and output-length gaps are reported directly for both engines; warning cells are not silently normalized away.
C4 release source trees: runtime ce9b25ce831a1e88da87bf13a464716da5245b51, public 30e07fd1ccdd4c9d4dbcf400084bf94dd0ad6434; B200-installed CPython 3.12 x86_64 kernels wheel SHA-256 38ecc42ad852ae9d356ba4680d3777bc04810faa1d73866fbf2f632b7f63c5ff; companion A/B bundle SHA-256 396d2698564f29b65385a0bafaf141df6c87c73fe60e144c5a0e6e219e66cafd / 6f7f3ee9e83507b1b6f75c6f4ecbed0a929de4836f16da4b0c5c8e6b9f079a59. Retained C8 source tree: aa8e173b26a1988d52f1bdc5bdfc5be2e1ee652c.
report-data.json: normalized values and pair-level provenance checks used by this report.SHA256SUMS: hashes for the report, data ledger, generator, and charts.- Private raw manifests, requests, telemetry, runtime-batch traces, and cold-start records remain in the retained artifact tree.









