I benchmarked every model in OpenRouter's Discounted Models collection β plus a full-price frontier model and a local 27B β to answer one question: which of these should drive an agent that calls tools and verifies its work, rather than guessing?
Short answer: the model with the smallest discount in the collection won, and a $15/M frontier model couldn't beat it by more than noise.
Scope, stated up front. Everything here is deterministic-verifier work β did it call get_contacts, did the cron job land in place, is the arithmetic right. That excludes everything where the answer is a judgment call: prose quality, voice, taste. A low score here means "don't put it in an agent loop." It says nothing about whether a model writes well. See What this doesn't measure.
Harness. BenchLocal 0.3.0, temperature 0, 1 run per test, driven through its Agent API. Three valid