Quality-first setup (current 2026-08-28, matches docker-compose.qwen3.8-27b-xl.yml), measured on:
- CPU: Intel Core Ultra 7 270K Plus (24C/24T, no HT)
- RAM: 96GB DDR5-5600 (2x48GB dual channel, ≈89.6 GB/s theoretical bandwidth)
- GPU: RTX 4090 24GB
Use the BeeLlama fork image (
server-cuda13-v0.4.3) — mainline llama.cpp silently falls back to CPU for non-q4 KV caches on Qwen3.x hybrid architecture (no error is reported).