Skip to content

Instantly share code, notes, and snippets.

@AshyIsMe
Created June 26, 2026 09:03
Show Gist options
  • Select an option

  • Save AshyIsMe/ef9dc0e57d7af5626f3935f55c38543a to your computer and use it in GitHub Desktop.

Select an option

Save AshyIsMe/ef9dc0e57d7af5626f3935f55c38543a to your computer and use it in GitHub Desktop.
Qwen3.6-35B-A3B-MTP on Radeon RX 7900 XTX (24 GB) — llama.cpp ROCm benchmarks

Qwen3.6-35B-A3B-MTP on Radeon RX 7900 XTX (24 GB) — llama.cpp ROCm benchmarks

Goal: run unsloth/Qwen3.6-35B-A3B-MTP-GGUF (Q4_K_XL) entirely on the GPU with the largest possible context, and find the fastest generation settings — including the model's built-in MTP (Multi-Token Prediction) self-speculation.

This is the MoE sibling of the dense 27B (see Benchmarks.md): 35.5 B total params but only ~3 B active per token, so it generates ~2.4× faster while still fitting the full 256k context on 24 GB.

Environment

GPU AMD Radeon RX 7900 XTX, gfx1100, 24560 MiB VRAM
CPU AMD Ryzen 5 3600 (6c/12t)
RAM 128 GiB
Build llama.cpp ROCm (libggml 0.15.2, build be4a6a6), ./llama-bench / ./llama-server
Model Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf, 21.27 GiB, 35.51 B params (~3 B active)

Model architecture (from GGUF metadata + llama.cpp source)

Arch qwen35moe is a hybrid SSM + sparse-attention Mixture-of-Experts model. Two things make large contexts feasible on 24 GB: only ~1 in 4 layers keeps a KV cache, and the active-param count is tiny (MoE), so decode is fast.

key value meaning
block_count 41 total layers
expert_count / expert_used_count 256 / 8 MoE — 256 experts, 8 routed per token (~3 B active)
expert_feed_forward_length 512 small per-expert FFN (+ 512 shared)
full_attention_interval 4 only ~1 in 4 layers keeps a KV cache (~10 layers); the rest are SSM/linear (fixed-size state, no growth with context)
head_count / head_count_kv 16 / 2 GQA — very small KV (2 KV heads)
key_length/value_length 256
context_length 262144 native 256k, no rope scaling needed (rope.freq_base 1e7)
nextn_predict_layers 1 one trained MTP head → used as the speculative draft
ssm.* conv 4, state 128, groups 16, inner 4096 Mamba-style layers

Per-token KV (f16) ≈ 10 attn-layers × 2 KV-heads × (256+256) × 2 B ≈ 20 KiB/token — about 3× smaller than the dense 27B (64 KiB/token), thanks to fewer attention layers and head_count_kv=2. That extra KV headroom is what lets a 21 GiB model still reach 256k. Quantizing the KV cache is still required to fit.


1. Fit & raw throughput — llama-bench (synthetic depth -d)

All layers on GPU (-ngl 999), flash attention on. "OK" = the whole model + that much KV fits in VRAM; "OOM" = it doesn't.

KV type ctx pp t/s tg t/s fits?
f16 2k 2724 87.2 ✅
f16 144k — — ❌ OOM
f16 256k — — ❌ OOM
q8_0 2k 2717 85.1 ✅
q8_0 144k 572 52.9 ✅
q8_0 256k — — ❌ OOM
q4_0 2k 2692 82.9 ✅
q4_0 144k 572 50.0 ✅
q4_0 256k 353 36.9 ✅ only KV quant that fits full 256k

Batch/ubatch tuning at depth 0: ubatch 1024 lifts prompt-processing to ~3640 t/s (vs ~2700 at ubatch 512) with no change to tg — worth setting for prompt-heavy use.

Takeaway: full 256k context fits only with q4_0 KV. q8_0 is viable up to ~144k at slightly higher KV quality. (Same frontier as the 27B — but every throughput number here is ~2.5–3× higher because only ~3 B params are active.)


2. Real high-context speed + MTP — llama-server

Measured by filling the KV cache with actual text (concatenated llama.cpp source), then generating 128 tokens (ignore_eos, greedy). The server reports prompt-eval and generation throughput separately. Config: -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1; MTP adds --spec-type draft-mtp --spec-draft-n-max 4 -ctkd q4_0 -ctvd q4_0.

--parallel 1 is essential: the server otherwise auto-selects n_parallel=4, which multiplies the compute/output buffers and OOMs MTP.

pp = prompt processing (context ingest), tg = generation; all t/s. MTP speeds up generation 1.35–1.70× but slightly lowers pp (extra draft pass):

context filled pp plain pp MTP tg plain tg MTP tg speedup MTP acceptance (rate / mean len)
~16k 1918 1817 79.6 135.2 1.70× 0.82 / 4.23
~64k 1368 1289 66.0 108.2 1.64× 0.82 / 4.23
~144k 916 848 50.1 68.5 1.37× 0.65 / 3.53
~196k 753 694 43.4 58.6 1.35× 0.64 / 3.53
~245k 641 OOM 38.3 OOM — MTP draft ctx + compute overflows VRAM

Prompt-processing (pp) throughput falls with depth as expected, but stays ~2.4× faster than the dense 27B at every depth — pp is compute-bound on active params (~3 B vs 27.3 B), so the MoE wins by roughly the active-param ratio, flat across the whole context range (plain decoding; MTP shaves a few % off pp for the extra draft pass):

context filled 27B pp t/s 35B-A3B pp t/s speedup
~16k 815 1918 2.35×
~64k 573 1368 2.39×
~144k 382 916 2.40×
~196k 314 753 2.40×
~245k 267 641 2.40×

(Both decay ~3× from 16k→245k — that part is the attention/KV cost, which scales the same for both. These are server real-text figures; the llama-bench synthetic-depth pp in §1 is measured differently and is not comparable.)

Acceptance is content-dependent (more predictable text accepts more). Expect ~1.35–1.70× in practice — best at low/mid context, tapering as the context fills.


3. The frontier: max context OR MTP, not both

The MTP head runs as a separate "draft context" on the same model, with its own KV cache sized to n_ctx (quantized by -ctkd/-ctvd) plus its own compute buffers. Combined with the main q4_0 KV and compute buffers, MTP fits up to n_ctx ≈ 204800 (~200k) but OOMs by ~217k (failed to create MTP context). Plain decoding reaches the full 256k.

  • Full ~256k context → plain decoding, ~38 t/s at depth.
  • Up to ~200k context → MTP, ~1.35–1.70× faster (~59–135 t/s depending on depth).

vs. the dense 27B

27B (dense) 35B-A3B (MoE)
File size (Q4_K_XL) 16.67 GiB 21.27 GiB
Active params/token 27.3 B (all) ~3 B of 35.5 B
KV per token (f16) ~64 KiB ~20 KiB
tg @16k (plain / MTP) 29 / 50 t/s 80 / 135 t/s
tg @256k (plain) ~14 t/s ~38 t/s
Max ctx (q4_0 KV) 256k 256k
MTP context ceiling ~208k ~200k

The MoE model is bigger on disk but much faster and has a smaller KV, so it wins on every throughput axis; its only regression is a slightly lower MTP context ceiling (the 21 GiB weights leave less room for the MTP draft buffers).


Recommended configurations (see qwen-35b.sh)

Maximum context (256k), plain decoding:

./llama-server -m <model> -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1 -c 262144

Fast (MTP speculative), up to ~200k context:

./llama-server -m <model> -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1 -c 204800 \
  --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 -ctkd q4_0 -ctvd q4_0

Notes / further tuning:

  • --spec-draft-n-max 4 matches the observed mean accepted length (~4 at low depth). Try 5–6 if your workload is very predictable (code), lower if poor.
  • For ≤144k you can use q8_0 KV (-ctk q8_0 -ctv q8_0, and -ctkd q8_0 -ctvd q8_0) for slightly higher KV quality; MTP still fits.
  • Add -ub 1024 for ~35% faster prompt processing if your prompts are large.
  • --parallel 1 is required for the MTP configs to fit at high context.

Reproduce

  • bench-fit-35b.sh — llama-bench fit + raw-throughput sweep (KV quant × depth).
  • bench-context-35b.sh — llama-server real-context depth sweep, plain vs MTP (DEPTHS="..." SPECS="..." ./bench-context-35b.sh to target specific points).
  • Raw results: bench-fit-*.csv, bench-context-*.csv, server logs in bench-ctx/.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment