Last updated: 2026-08-17
Hardware under test: Windows 11 + WSL2 Ubuntu 24.04 · Intel Core i9-12900K · RTX 5090 32 GB (sm_120) · 32 GB system RAM (WSL capped ~20 GB)
Stack: vLLM 0.27.1 · FlashInfer 0.6.16.post3 · CUDA 13.2 · Unsloth Qwen3.8-27B-NVFP4 (~22 GB weights)
Goal: Document real A/B results so others can reproduce. Prefer stable single-stream decode (interactive agent/chat) over theoretical multi-stream throughput.
TL;DR
- MTP speculative decode ×2 is the big win (~+73% tok/s vs no MTP).
- A popular “performance” trio (
max-num-batched-tokens 8192+ explicit--attention-backend flashinfer+gpu-memory-utilization 0.93) regressed short-stream decode on this box — twice.