Card: AMD Instinct MI50 32 GB (gfx906 / Vega 20 / GCN5, ~1 TB/s HBM2, no matrix cores, passive-cooled) · replaced an Intel Arc B580 (permanent swap)
Host: Nobara/Fedora, Ryzen 7 8845HS (8c/16t), 30 GB RAM · SELinux disabled · Docker
Workload: n8n LLM node — "big data in → small data out" (prefill-dominated), single-user (--parallel 1)
Compiled: 2026-07-03 — consolidates every test from 2026-06-20 → 2026-07-03.
Source reports: MI50-report.md, MI50-MTP-speculative-decoding-analysis.md, MI50-Q5-Q6-quant-test.md, DFlash-setup.md, dflash-bench-2026-06-29.md, ornith-mtp-optimization-report.md, ornith-model-docker-logs-2026-07-01.md, llama-stats-2026-06-23.md, memory mi50-*.
Current production (2026-07-03):
llama-hipgraphscontainer, imagellama-hipgraphs:upstream-rocm-7.2.4(build commit0eca4d4, 2026-06-30), model Ornith-1.0-35B Q4_K_M + embedded MTP,--spec-type draft-mtp --spec-draft-n-max 2, 262 k ctx, q8_0 KV,--parallel 1, port 8089. ~70 t/s.