Status: draft for review. Background reading: agentEvaluationReference.md.
A reliable, low-setup way to answer "is the agent working properly?" — and, once a prompt or model changes, "did that make it better or worse?"
Non-goals for the first release: production traffic splitting, statistical A/B