Status: Final 28-cell board against vLLM 0.27.1. The board retains 26 validated, unchanged cells from
7f0dddc2, retains the frozen Gemma E4B C8 measurement frome1d6c1dd, and replaces Gemma E4B C4 with the final fast/correct release measurement from5c06c090. Every Photon/vLLM pair is internally matched on driver, model revision, harness, ordered request stream, request count, and B200 hardware class; exact engine commits are recorded per cell.
ChartQA multimodal inference with reasoning enabled on NVIDIA B200. The primary metric is natural-output tokens completed per active end-to-end service second; accuracy and output-length differences are reported alongside it.
- Photon leads in 28/28 end-to-end output-token service-throughput cells.
- Photon leads in 28/28 completed-request-rate cells.