This matrix compares the seven model providers behind every model in the Model Selection guide, provider by provider, so the vendor choice is as visible as the model choice.
Provider choice is a bundle decision, list price, cache discount, batch policy, context flatness, weights, and where your subscription does and does not transfer, and the challengers win that bundle on every axis except subscriptions.
Legend: ✓ supported, ✗ not supported, ~ partial or conditional, ? not verified. Each column links to the provider’s canonical pricing page; every cell traces to the guide or to a source cited in its references.
The matrix #
| Provider | |||||||
|---|---|---|---|---|---|---|---|
| Flagship model | Qwen3.8-Max, $2 / $6 | Claude Fable 5.1, $10 / $50 | DeepSeek V4 Pro, $1.32 / $3.96 peak, the Flash routing cancelled, rates continue | Gemini 3.1 Pro Preview, $2 / $12 | Kimi K3, $3 / $15 list (OpenRouter hosts floor at $0.88 / $4.90) | GPT-6 Astra, $10 / $50 | GLM-5.3, $1.40 / $4.40 |
| Workhorse model | Qwen3.8-Max doubles as the workhorse | Claude Sonnet 5, $2 / $10 | deepseek-flash (V4.1-Flash), $0.30 / $1.20 peak | Gemini 3.1 Pro Preview, $2 / $12 | kimi-k2.7-code, $0.95 / $4.00 list (host floor $0.66 / $3.30) | gpt-6-sol, $2 / $10 (gpt-5.6-terra at $2 / $12 is off OpenAI’s pricing page) | GLM-5.3 covers the class at $1.40 / $4.40 |
| Cheap model | Qwen3.8-Flash, $0.15 / $0.47 | Claude Haiku 4.5, $1 / $5 | deepseek-flash doubles as the cheap tier, $0.15 / $0.60 off-peak | Flash-Lite $0.25 / $1.50, 3.6 Flash $0.75 intro | k2.7-code doubles as the cheap tier | gpt-6-luna, $0.10 / $0.50 | GLM-5.3-Flash, $0.15 / $0.50, GLM-4.7-Flash free |
| Coding-specialized model | ? none in this guide (the open Coder family is separate) | ✗ none, the general lineup carries coding | ✗ none in the guide | ✗ none in the guide | ✓ kimi-k2.7-code | ✓ gpt-5.3-codex, $1.75 / $14 | ✗ none in the guide |
| Cached input discount | ~ discount stated, rate unpublished | ✓ one tenth, Opus 5.5 at one twentieth, the Fable pair at one fortieth | ✓ about one fiftieth on Flash, the deepest | ✓ about one tenth | ✓ one tenth on K3, one fifth on k2.7-code | ✓ about one tenth | ✓ about one fifth |
| Long-context policy | ✓ 1M flat | ✓ 1M flat from Claude 4.6 onward | ✓ 1M flat, 384K max output | ~ doubles past thresholds, Pro beyond 200K | ~ K3 flat 1M, k2.7-code 256K | ~ doubles past thresholds | ✓ 1M flat |
| Batch or time discounts | ~ batch half price on Max | ✓ batch half price | ~ off-peak half price instead of batch | ✓ batch half price | ? not stated | ✓ batch half price | ? not stated |
| Live pricing windows | ~ 1M-token free quota for new accounts, 90 days | ~ Sonnet 5’s planned increase was cancelled | ~ off-peak is standing policy; the 2026-09-14 Flash routing was cancelled, V4 Pro rates continue | ~ Flash intro rate ends 2026-12-31 | ~ HighSpeed variant at $1.90 / $8.00 | ~ Sol promo through 2026-11-21 | ✓ none; the 50% promo ended 2026-09-09 into the $0.15 / $0.50 list |
| Weights | ✓ open weights (27B and Flash-Next) | ✗ closed | ✓ open weights | ✗ closed | ✓ open weights (K3, K2.7) | ✗ closed | ✓ open weights |
| Subscription includes an agent | ? none verified | ✓ Claude Code from Pro $20 to Max 20x $200 | ✗ API and BYOK only | ~ the Antigravity platform is free | ~ Kimi OAuth reuse via Kimi Code | ✓ Codex from ChatGPT Free tier up | ~ coding plans from $18/month |
| BYOK harness fit | ~ OpenAI-compatible, rides OpenRouter or any compatible endpoint | ~ subscription does not transfer to third-party harnesses | ✓ native OpenCode provider, plus an Anthropic-format endpoint | ? not verified here | ✓ native OpenCode provider | ~ API key works everywhere, subscription locked to Codex | ✓ native OpenCode provider |
Reading the matrix #
I read this table as a bundle scorecard rather than a price list, because the rows move together in ways the model-level table hides. The big three monetize closed frontier weights plus subscriptions, and the four challengers monetize cheap tokens plus open weights, with almost no overlap in which rows they win. On the price rows the challengers sweep: every challenger flagship costs less out than the big-three workhorses, and Zhipu’s free GLM-4.7-Flash has no big-three analogue at any price. On the policy rows the field is closer than list prices suggest, and this is the part I check before switching: DeepSeek’s cache discount is the deepest at about one fiftieth on Flash, Alibaba leaves its cache rate unpublished, and batch is half price at the big three and Qwen but unverified at Moonshot and Zhipu. Long context splits cleanly by flatness: Anthropic, DeepSeek, Alibaba, Moonshot K3, and Zhipu keep 1M pricing flat, while OpenAI and Google double past thresholds, which decides whole-repository prompting before any model-quality question. The weights row is the quiet differentiator: open weights from Alibaba, DeepSeek, Moonshot, and Zhipu mean a private deployment path the closed big three do not offer at all. The subscription row runs the other way: only OpenAI and Anthropic ship a first-party agent on subscription, and Anthropic’s is the one that explicitly does not transfer to third-party harnesses.
Choosing between providers #
- Default workhorse loop: OpenAI, Anthropic, and Google price identically at the tier, so pick by harness fit and let the challenger prices below pressure that default yearly.
- Price-floor BYOK loops: DeepSeek and Zhipu, with Alibaba close behind at flagship quality for $6 out.
- Flat 1M context: Anthropic, DeepSeek, Alibaba, Moonshot K3, and Zhipu, the five vendors that keep whole-repository reads predictable.
- Coding-specialized spend: Moonshot’s k2.7-code or OpenAI’s gpt-5.3-codex, the only two columns with a model built for the job.
- Weights you can host: Alibaba, DeepSeek, Moonshot, or Zhipu, and none of the big three at any price.
- Subscription-first teams: OpenAI or Anthropic, and read the transfer row before assuming the subscription follows your harness.
Changes #
- 2026-09-08 - Created on owner request with seven provider columns sorted alphabetically as a companion to the model selection guide.
- 2026-09-08 - Moved Zhipu’s long-context cell to 1M flat on models.dev evidence, widened the flat-1M club to five, and added the models.dev reference.
- 2026-09-09 - Updated the Moonshot workhorse and pricing-window cells and fixed the Zhipu cached-price prose contradiction.
- 2026-09-10 - Moved the DeepSeek workhorse cell to deepseek-flash and updated the Zhipu cells as the promo resolved to list.
- 2026-09-12 - Updated the DeepSeek cells for the routing cancellation and gave Moonshot’s flagship cell the OpenRouter price with the platform lag.
- 2026-09-16 - Updated the Moonshot workhorse cell for OpenRouter’s kimi-k2.7-code cut to $3.21 and moved its cache discount to one quarter.
- 2026-09-18 - Updated the Moonshot flagship cell for OpenRouter’s kimi-k3 cut to $2.10 / $10.95.
- 2026-09-20 - Updated the Moonshot flagship cell again for OpenRouter’s kimi-k3 cut to $1.70 / $8.50 with cache hits at $0.17, and moved its cache-discount cell from about one ninth to one tenth.
- 2026-09-24 - Removed the verification preamble line on owner request.
- 2026-09-25 - Re-grounded the Moonshot flagship and workhorse cells on Moonshot’s official list prices ($3/$15 and $0.95/$4.00) after the endpoints API showed the earlier “OpenRouter price with platform lag” readings were third-party host floors, and moved the OpenAI workhorse and cheap cells to the 2026-09-22 GPT-6 Sol and Luna releases.
- 2026-09-27 - Refreshed against all seven pricing pages: noted the new Claude Opus 5.5 tier in the cache-discount cell, moved the Kimi K3 host floor to $0.88 / $4.90, and flagged gpt-5.6-terra’s removal from OpenAI’s pricing page in the workhorse cell.
See also #
- Model Selection for Coding Tasks - the model-level pricing table every cell here traces to
- Harness Feature Matrix - the client side of the same purchase, capability by capability
- Agentic Coding Tools Landscape - the four-layer map that puts model vendors in context
- Kimi Code - a vendor building its own harness to sell its own tokens
- OpenCode - the lean BYOK harness the challenger provider columns ride
References #
https://platform.openai.com/docs/pricing - GPT-5.6 and GPT-6 family prices including the 2026-09-22 GPT-6 Sol and Luna, batch discount, long-context doubling, gpt-5.6-terra and gpt-5.6-luna no longer listed, gpt-5.6-cyber added under Daybreak (verified 2026-09-27)
https://docs.claude.com/en/docs/about-claude/pricing - Claude lineup prices including Opus 5.5 ($4/$20, 5% cache), cache multipliers, 1M-context policy (verified 2026-09-27)
https://cloud.google.com/vertex-ai/generative-ai/pricing - Gemini 3 family prices, intro windows, long-context and batch rates, 3.8 Flash Cyber at $1.50 / $7.50 (verified 2026-09-27)
https://help.aliyun.com/zh/model-studio/billing-for-model-studio - Qwen3.8-Max and Qwen3.8-Flash official prices, batch half price, cache discount, free quota, qwen3.8-max-prime speed variant (fetched 2026-09-27)
https://api-docs.deepseek.com/news/news260910 - the 2026-09-10 V4.1-Flash release and V4 Pro routing announcement behind the DeepSeek column’s workhorse and windows cells (fetched 2026-09-10, routing cancelled per the DeepSeek pricing page fetched 2026-09-27)
https://openrouter.ai/api/v1/models - aggregator listings for the Qwen pair and the GPT-6 pair (fetched 2026-09-27)
https://openrouter.ai/api/v1/models/moonshotai/kimi-k3/endpoints - per-provider endpoint prices separating Moonshot’s official list prices from third-party host floors, now $0.88 / $4.90 on Morph (fetched 2026-09-27)
https://models.dev - the community model list reference: model ids, release dates, and context windows behind the lineup, including the GPT-6 pair’s and Opus 5.5’s 2026-09-22 release dates (fetched 2026-09-27)