This guide sets up Qwen3.8 27B with pi through llama.cpp on a 32 GB machine: an Apple Silicon Mac or a Linux box with a 24 GB GPU. Thinking is kept in the conversation on every turn, so the KV cache survives between messages. No API key, no cloud, works offline.
| download | 17.56 GB (UD-Q4_K_XL) |
| resident | about 19.4 GB at 64k context |
