Skip to content

Instantly share code, notes, and snippets.

vLLM on RTX 5090 (32 GB) — Qwen3.8-27B NVFP4 startup flags: what we tested

Last updated: 2026-08-17
Hardware under test: Windows 11 + WSL2 Ubuntu 24.04 · Intel Core i9-12900K · RTX 5090 32 GB (sm_120) · 32 GB system RAM (WSL capped ~20 GB)
Stack: vLLM 0.27.1 · FlashInfer 0.6.16.post3 · CUDA 13.2 · Unsloth Qwen3.8-27B-NVFP4 (~22 GB weights)
Goal: Document real A/B results so others can reproduce. Prefer stable single-stream decode (interactive agent/chat) over theoretical multi-stream throughput.

TL;DR

  1. MTP speculative decode ×2 is the big win (~+73% tok/s vs no MTP).
  2. A popular “performance” trio (max-num-batched-tokens 8192 + explicit --attention-backend flashinfer + gpu-memory-utilization 0.93) regressed short-stream decode on this box — twice.
@cya9nide
cya9nide / ec2cleanup.rb
Last active January 11, 2018 23:10
ec2 cleanup with cloudwatch metrics, similar to the ebscleanup script
require 'aws-sdk'
require 'logger'
require 'json'
# Logging options
@log = Logger.new('ec2cleanup.log','weekly')
@log_candidate = Logger.new('ec2candidates.log','weekly')
# shared profile account
@profile = "profilename"