Multi-Cluster Inference, Active-Active Placement, Cloud Burst, Token Governance, and Closed-Loop Optimization
Status: architecture direction, grounded in the current burst-routing-v1 branches
Scope: Grid + Praxis AI + Praxis core + llm-d/EPP + self-hosted and cloud inference providers + shared token state
Primary design rule: Grid decides asynchronously; Praxis executes locally from an immutable snapshot
LoRA / multi-LoRA: deferred