Goal: run unsloth/Qwen3.6-35B-A3B-MTP-GGUF (Q4_K_XL) entirely on the GPU with the
largest possible context, and find the fastest generation settings — including
the model's built-in MTP (Multi-Token Prediction) self-speculation.
This is the MoE sibling of the dense 27B (see Benchmarks.md): 35.5 B total
params but only ~3 B active per token, so it generates ~2.4× faster while
still fitting the full 256k context on 24 GB.
| GPU | AMD Radeon RX 7900 XTX, gfx1100, 24560 MiB VRAM |
| CPU | AMD Ryzen 5 3600 (6c/12t) |
| RAM | 128 GiB |
| Build | llama.cpp ROCm (libggml 0.15.2, build be4a6a6), ./llama-bench / ./llama-server |
| Model | Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf, 21.27 GiB, 35.51 B params (~3 B active) |
Arch qwen35moe is a hybrid SSM + sparse-attention Mixture-of-Experts model.
Two things make large contexts feasible on 24 GB: only ~1 in 4 layers keeps a KV
cache, and the active-param count is tiny (MoE), so decode is fast.
| key | value | meaning |
|---|---|---|
block_count |
41 | total layers |
expert_count / expert_used_count |
256 / 8 | MoE — 256 experts, 8 routed per token (~3 B active) |
expert_feed_forward_length |
512 | small per-expert FFN (+ 512 shared) |
full_attention_interval |
4 | only ~1 in 4 layers keeps a KV cache (~10 layers); the rest are SSM/linear (fixed-size state, no growth with context) |
head_count / head_count_kv |
16 / 2 | GQA — very small KV (2 KV heads) |
key_length/value_length |
256 | |
context_length |
262144 | native 256k, no rope scaling needed (rope.freq_base 1e7) |
nextn_predict_layers |
1 | one trained MTP head → used as the speculative draft |
ssm.* |
conv 4, state 128, groups 16, inner 4096 | Mamba-style layers |
Per-token KV (f16) ≈ 10 attn-layers × 2 KV-heads × (256+256) × 2 B ≈ 20 KiB/token
— about 3× smaller than the dense 27B (64 KiB/token), thanks to fewer
attention layers and head_count_kv=2. That extra KV headroom is what lets a
21 GiB model still reach 256k. Quantizing the KV cache is still required to fit.
All layers on GPU (-ngl 999), flash attention on. "OK" = the whole model +
that much KV fits in VRAM; "OOM" = it doesn't.
| KV type | ctx | pp t/s | tg t/s | fits? |
|---|---|---|---|---|
| f16 | 2k | 2724 | 87.2 | ✅ |
| f16 | 144k | — | — | ❌ OOM |
| f16 | 256k | — | — | ❌ OOM |
| q8_0 | 2k | 2717 | 85.1 | ✅ |
| q8_0 | 144k | 572 | 52.9 | ✅ |
| q8_0 | 256k | — | — | ❌ OOM |
| q4_0 | 2k | 2692 | 82.9 | ✅ |
| q4_0 | 144k | 572 | 50.0 | ✅ |
| q4_0 | 256k | 353 | 36.9 | ✅ only KV quant that fits full 256k |
Batch/ubatch tuning at depth 0: ubatch 1024 lifts prompt-processing to ~3640 t/s (vs ~2700 at ubatch 512) with no change to tg — worth setting for prompt-heavy use.
Takeaway: full 256k context fits only with q4_0 KV. q8_0 is viable up
to ~144k at slightly higher KV quality. (Same frontier as the 27B — but every
throughput number here is ~2.5–3× higher because only ~3 B params are active.)
Measured by filling the KV cache with actual text (concatenated llama.cpp
source), then generating 128 tokens (ignore_eos, greedy). The server reports
prompt-eval and generation throughput separately. Config: -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1; MTP adds --spec-type draft-mtp --spec-draft-n-max 4 -ctkd q4_0 -ctvd q4_0.
--parallel 1is essential: the server otherwise auto-selectsn_parallel=4, which multiplies the compute/output buffers and OOMs MTP.
pp = prompt processing (context ingest), tg = generation; all t/s. MTP speeds up generation 1.35–1.70× but slightly lowers pp (extra draft pass):
| context filled | pp plain | pp MTP | tg plain | tg MTP | tg speedup | MTP acceptance (rate / mean len) |
|---|---|---|---|---|---|---|
| ~16k | 1918 | 1817 | 79.6 | 135.2 | 1.70× | 0.82 / 4.23 |
| ~64k | 1368 | 1289 | 66.0 | 108.2 | 1.64× | 0.82 / 4.23 |
| ~144k | 916 | 848 | 50.1 | 68.5 | 1.37× | 0.65 / 3.53 |
| ~196k | 753 | 694 | 43.4 | 58.6 | 1.35× | 0.64 / 3.53 |
| ~245k | 641 | OOM | 38.3 | OOM | — | MTP draft ctx + compute overflows VRAM |
Prompt-processing (pp) throughput falls with depth as expected, but stays ~2.4× faster than the dense 27B at every depth — pp is compute-bound on active params (~3 B vs 27.3 B), so the MoE wins by roughly the active-param ratio, flat across the whole context range (plain decoding; MTP shaves a few % off pp for the extra draft pass):
| context filled | 27B pp t/s | 35B-A3B pp t/s | speedup |
|---|---|---|---|
| ~16k | 815 | 1918 | 2.35× |
| ~64k | 573 | 1368 | 2.39× |
| ~144k | 382 | 916 | 2.40× |
| ~196k | 314 | 753 | 2.40× |
| ~245k | 267 | 641 | 2.40× |
(Both decay ~3× from 16k→245k — that part is the attention/KV cost, which scales
the same for both. These are server real-text figures; the llama-bench
synthetic-depth pp in §1 is measured differently and is not comparable.)
Acceptance is content-dependent (more predictable text accepts more). Expect ~1.35–1.70× in practice — best at low/mid context, tapering as the context fills.
The MTP head runs as a separate "draft context" on the same model, with its own
KV cache sized to n_ctx (quantized by -ctkd/-ctvd) plus its own compute
buffers. Combined with the main q4_0 KV and compute buffers, MTP fits up to
n_ctx ≈ 204800 (~200k) but OOMs by ~217k (failed to create MTP context).
Plain decoding reaches the full 256k.
- Full ~256k context → plain decoding, ~38 t/s at depth.
- Up to ~200k context → MTP, ~1.35–1.70× faster (~59–135 t/s depending on depth).
| 27B (dense) | 35B-A3B (MoE) | |
|---|---|---|
| File size (Q4_K_XL) | 16.67 GiB | 21.27 GiB |
| Active params/token | 27.3 B (all) | ~3 B of 35.5 B |
| KV per token (f16) | ~64 KiB | ~20 KiB |
| tg @16k (plain / MTP) | 29 / 50 t/s | 80 / 135 t/s |
| tg @256k (plain) | ~14 t/s | ~38 t/s |
| Max ctx (q4_0 KV) | 256k | 256k |
| MTP context ceiling | ~208k | ~200k |
The MoE model is bigger on disk but much faster and has a smaller KV, so it wins on every throughput axis; its only regression is a slightly lower MTP context ceiling (the 21 GiB weights leave less room for the MTP draft buffers).
Maximum context (256k), plain decoding:
./llama-server -m <model> -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1 -c 262144
Fast (MTP speculative), up to ~200k context:
./llama-server -m <model> -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1 -c 204800 \
--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 -ctkd q4_0 -ctvd q4_0
Notes / further tuning:
--spec-draft-n-max 4matches the observed mean accepted length (~4 at low depth). Try 5–6 if your workload is very predictable (code), lower if poor.- For ≤144k you can use
q8_0KV (-ctk q8_0 -ctv q8_0, and-ctkd q8_0 -ctvd q8_0) for slightly higher KV quality; MTP still fits. - Add
-ub 1024for ~35% faster prompt processing if your prompts are large. --parallel 1is required for the MTP configs to fit at high context.
bench-fit-35b.sh— llama-bench fit + raw-throughput sweep (KV quant × depth).bench-context-35b.sh— llama-server real-context depth sweep, plain vs MTP (DEPTHS="..." SPECS="..." ./bench-context-35b.shto target specific points).- Raw results:
bench-fit-*.csv,bench-context-*.csv, server logs inbench-ctx/.