llm-bench sweep · aggregate decode tok/s at concurrency 1 · Apple M5 Max, 128 GB unified memory, 40 GPU cores · 30s per cell, greedy, ignore_eos, max_tokens 8192
Full serve and benchmark commands for every series on this chart, extracted from the canonical benchmark records (the source of truth). One block per engine/variant; the swept draft-depth parameter is shown as N with its range.
engine: MTPLX 2.11.3 (brew install youssofal/mtplx/mtplx; native process, no container)