Cached input price is everything. A coding agent's token mix is 95.6% cache reads (measured: 8.04B tokens, 16 May 2026). Headline input price is nearly irrelevant.
Formula: Effective $/1M = cache_read × 0.9564 + cache_write × 0.0273 + input × 0.0134 + output × 0.0029
Important 25 June 2026 fix: the Anthropic-style cache_write = 1.25 × input default is not universal. For OpenAI, Gemini, DeepSeek, and most non-Anthropic providers, use cache_write = input/cache-miss unless the provider publishes a separate write price.
| # | Model | Cache Read | Input | Output | Cache Write | Effective $/1M | vs cheapest | AA Intelligence Index | Verification |
|---|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash | $0.0028 | $0.14 | $0.28 | $0.14 | $0.0092 | 1.0x | 40 | DeepSeek official |
| 2 | MiMo V2.5 Pro | $0.0036 | $0.435 | $0.87 | Free promo | $0.0118 | 1.3x | 42 | Xiaomi/OpenRouter |
| 3 | DeepSeek V4 Pro | $0.003625 | $0.435 | $0.87 | $0.435 | $0.0237 | 2.6x | 44 | DeepSeek official |
| 4 | MiniMax-M3 | $0.06 | $0.3 | $1.2 | $0.375 | $0.0751 | 8.2x | 44 | AA/OpenRouter; MiniMax official page blocked |
| 5 | Qwen3.7 Plus | $0.064 | $0.32 | $1.28 | $0.4 | $0.0801 | 8.7x | 39 | OpenRouter + Alibaba latest model page |
| 6 | Claude Haiku 4.5 | $0.1 | $1 | $5 | $1.25 | $0.1577 | 17.2x | n/a | Anthropic official |
| 7 | Gemini 3.5 Flash | $0.15 | $1.5 | $9 | $1.5 | $0.2306 | 25.1x | 50 | Google official |
| 8 | GPT-5.3 Codex | $0.175 | $1.75 | $14 | $1.75 | $0.2792 | 30.4x | 44 | OpenAI official/OpenRouter |
| 9 | GLM-5.2 | $0.26 | $1.4 | $4.4 | Free promo | $0.2802 | 30.5x | 51 | Z.AI official |
| 10 | Gemini 3.1 Pro Preview | $0.2 | $2 | $12 | $2 | $0.3075 | 33.5x | 46 | Google official |
| 11 | Qwen3.7 Max | $0.25 | $1.25 | $3.75 | $1.562 | $0.3094 | 33.7x | 46 | OpenRouter + Alibaba latest model page |
| 12 | Claude Sonnet 4.6 | $0.3 | $3 | $15 | $3.75 | $0.4730 | 51.5x | 47 | Anthropic official |
| 13 | GPT-5.5 | $0.5 | $5 | $30 | $5 | $0.7687 | 83.7x | 55 | OpenAI official |
| 14 | Claude Opus 4.8 | $0.5 | $5 | $25 | $6.25 | $0.7883 | 85.8x | 56 | Anthropic official |
| 15 | Claude Fable 5 | $1 | $10 | $50 | $12.5 | $1.5766 | 171.6x | 60 | Anthropic official/AA |
All prices: USD per 1M tokens. AA Intelligence Index: higher = smarter; n/a means not re-verified in the current public leaderboard scrape.
- Cheapest effective cost: DeepSeek V4 Flash at $0.0092/M, followed by MiMo V2.5 Pro at $0.0118/M.
- New cheap strong middle: MiniMax-M3 at $0.0751/M effective, with AA Intelligence Index 44.
- GLM update: replace GLM-5.1 with GLM-5.2. Its official pricing is unchanged vs GLM-5.1 in this table: $1.40 input / $0.26 cached / $4.40 output, effective $0.2802/M.
- Qwen update: Qwen3.7 Plus is much cheaper than Qwen3.7 Max under this model ($0.0801/M vs $0.3094/M) because its cache-read price is lower.
- Frontier tax: GPT-5.5 and Claude Opus 4.8 are ~84–86× DeepSeek V4 Flash effective cost; Claude Fable 5 is ~172× but currently leads AA's intelligence leaderboard.
- DeepSeek V4 Flash / V4 Pro pricing — https://api-docs.deepseek.com/quick_start/pricing
- Z.AI GLM-5.2 pricing — https://docs.z.ai/guides/overview/pricing
- MiniMax-M3 pricing/benchmark — https://artificialanalysis.ai/models/minimax-m3 and https://openrouter.ai/minimax/minimax-m3/api
- Qwen3.7 pricing/version aliases — https://openrouter.ai/api/v1/models and https://help.aliyun.com/zh/model-studio/model-pricing
- Anthropic Claude pricing — https://docs.anthropic.com/en/docs/about-claude/pricing
- OpenAI pricing — https://platform.openai.com/docs/pricing
- Gemini pricing/caching — https://ai.google.dev/gemini-api/docs/pricing and https://ai.google.dev/gemini-api/docs/caching
- Artificial Analysis Intelligence Index — https://artificialanalysis.ai/leaderboards/models