On this page
- Measure the workload first#
- What I would trial first#
- Normalize the subscription costs#
- Start with the API baseline#
- Compare with OpenRouter too#
- Convert the allowance before comparing it#
- Native coding subscriptions belong in the trial too#
- Temporary offers, kept separate#
- The next analysis should measure accepted work#
- Reproduce and update the comparison#
- Changelog#
- 2026-09-07 — Published as a data artifact#
- 2026-09-07 — Display the analysis date#
- 2026-09-07 — Index by subscription and model#
- 2026-09-07 — OpenRouter route comparison#
- 2026-09-07 — Normalize subscription costs#
- 2026-09-07 — Measured usage, supported models and capacity#
- 2026-09-07 — Initial research draft#
curl -s https://jedarden.com/data/inference-plans.md Once coding workers can keep themselves busy, inference becomes a purchasing decision. A subscription buys an allowance that expires. An API bills for the work requested. Both can be economical, and both can be wasted.
In The unit economics of running cattle, I argued for measuring cost per completed outcome. This artifact applies that discipline to buying inference: which offers deserve a trial, what their advertised units mean, and where the apparent savings disappear.
Analysis date: September 7, 2026.
This is a maintained research artifact. Prices are in USD before tax; individual source checks and usage captures retain their own dates in the dataset. The calculations model billing; they do not establish model quality, sustained throughput, or the cost of a shipped application. The changelog records substantive revisions. Publication dates, editorial updates, and source-verification dates serve different purposes.
Measure the workload first#
The baseline now comes from an actual local usage scan captured September 7, 2026 at 14:55 UTC, using Tokscale 4.15.1. It covers the locally discovered Claude Code, Codex and ZCode histories in one scan, across 871,701 recorded messages. This is an accumulated history snapshot, not one month of usage or a deduplicated total across every machine. The aggregate measurement contains the exact counters and derivation; no prompts or session records are included.
Fresh input2,293,232,314
- Share of all tokens
- 2.794%
- Per 1M output
- 5.985M
Cache reads78,848,276,612
- Share of all tokens
- 96.082%
- Per 1M output
- 205.776M
Cache writes538,768,695
- Share of all tokens
- 0.657%
- Per 1M output
- 1.406M
Output, including billed reasoning383,174,954
- Share of all tokens
- 0.467%
- Per 1M output
- 1.000M
82.06B measured tokens; 213.2:1 total input/output; 96.53% of input served from cache. All input includes fresh input, cache reads and cache writes.
Tokscale’s model report keeps separately reported reasoning outside its output field. I add that bucket once to obtain billable output; reasoning already included by a client stays inside output. Cache reads and writes are separate from fresh input. This normalization follows the versioned scanner implementation.
These are measured usage proportions, transferred into hypothetical purchases below. Another model, tokenizer, provider cache or session pattern can change them. In particular, the high cache-hit rate is not guaranteed on another service. Most processed tokens here are repeated context; maximizing that total alone would reward repetition rather than useful applications.
What I would trial first#
My shortlist from the published economics is OpenCode Go for routine coding, Z.ai for sustained GLM-5.3 work, and Synthetic when its model selection or additional concurrency justifies a pack. Keep a capped API balance available for overflow and difficult tasks. These are candidates to measure in the actual harness, not a ranking of coding quality.
The key eligibility question is whether the allowance covers the workload. A developer running a supported coding tool, an unattended coding worker, and a finished application serving customers can fall under different rules. An API-shaped endpoint does not establish permission for all three.
Each expanded plan reports total processed tokens, with output in parentheses, at the measured mix above. Each example spends the entire allowance on its named model; examples within one plan are alternatives. Weekly quotas are prorated to 30 days and assume full utilization within shorter limits. B means billion and M means million tokens. This is quota capacity, not a measured throughput promise or a model’s per-request context window.
OpenCode Go$1027 models / variants · 3 calculations
Supported models / catalog
- MiMo V2.5
- MiniMax M3
- Kimi K3
- DeepSeek V4 Flash
- GLM-5.3
- GPT-5.6 Luna
- MiMo V2.5 Pro
- GLM-5.3-Flash
- DeepSeek V4 Pro
- Grok 4.6
- GLM-5.2
- GLM-5.1
- Kimi K2.7 Code
- Kimi K2.6
- MiniMax M2.7
- Qwen 3.8 Max
- Qwen 3.8 Flash
- Qwen 3.7 Max
- Qwen 3.7 Plus
- Qwen 3.6 Plus
- LongCat 2.0
- Muse Spark 1.3 Contributor
- Muse Spark 1.2 Contributor
- DeepSeek V4 Flash Vision Exp
- Hy4
- Hy3
- Omen Alpha
Estimated processed tokens / 30 days
- MiMo V2.56.80B total (31.73M output)
- MiniMax M3815.16M total (3.81M output)
- Kimi K332.48M total (0.15M output)
Allowance and practical limit
$60 API allowance for MiMo V2.5 and MiniMax M3; $30 for DeepSeek V4 Flash; $15 for K3, GLM-5.3 and DeepSeek V4 Pro. One shared pool, with shorter limits. Coding-agent traffic only.
Z.ai Lite / Pro / Max$18 / $80 / $1682 models / variants · 3 calculations
Estimated processed tokens / 30 days
- Lite, GLM-5.3 off-peak432.12M total (2.02M output)
- Pro, GLM-5.3 off-peak2.59B total (12.11M output)
- Max, GLM-5.3 off-peak6.05B total (28.25M output)
Allowance and practical limit
10k / 60k / 140k credits weekly; model and cache weights apply. Half credit consumption off-peak. Supported coding tools; dynamic concurrency.
Synthetic$30 per pack11 models / variants · 2 calculations
Supported models / catalog
- syn:large:text
- syn:small:text
- syn:large:vision
- syn:small:vision
- openai/gpt-oss-120b
- zai-org/GLM-5.3-Flash
- zai-org/GLM-5.2
- moonshotai/Kimi-K3
- Qwen/Qwen3.8-27B
- zai-org/GLM-4.7-Flash
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
Estimated processed tokens / 30 days
- GLM-5.3-Flash2.24B total (10.45M output)
- Kimi-K3169.75M total (0.79M output)
Allowance and practical limit
$24 of its own API credits weekly; 500 weighted requests per five hours; one concurrent request per model per pack. More packs increase each limit.
camelStream$5 per stream5 models / variants · 1 calculation
Supported models / catalog
- DeepSeek V4 Flash
- Gemini 3.7 Flash
- GLM-5.3-Flash
- GPT-5.6 Luna
- Muse Spark 1.2
Estimated processed tokens / 30 days
- Unmetered; throughput unknown
Allowance and practical limit
Unmetered tokens; one generation per stream. Provider selects the model. No throughput guarantee; standard terms permit training on submitted content.
MiniMax Token Plan$22 / $55 / $1322 models / variants · 1 calculation
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
Five-hour and weekly limits; text, image and speech share quota. Current documentation lacks sufficient numerical quota detail for a defensible API-equivalent calculation.
Kimi membership / Code$19 / $39 / $99 / $1993 models / variants · 1 calculation
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
Membership and coding benefits share resources. No fixed comparable token entitlement established here; personal Code benefits and production API access differ.
Alibaba Cloud Coding Plan$5010 models / variants · 1 calculation
Supported models / catalog
- qwen3.7-plus
- qwen3.6-plus
- kimi-k2.5
- glm-5
- MiniMax-M2.5
- qwen3.5-plus
- qwen3-max-2026-01-23
- qwen3-coder-next
- qwen3-coder-plus
- glm-4.7
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
90k requests/month, 45k/week, 6k/five hours. Explicitly excludes automated scripts, application backends and noninteractive use.
deepseekv4pro.com$19.90 / $49.90; Agent $29.906 models / variants · 1 calculation
Supported models / catalog
- DeepSeek V4 Flash
- DeepSeek V4 Pro
- GLM-5.3
- GLM-5.2
- MiniMax M3
- Kimi K3
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
Independent reseller. Recurring promotional prices observed. Pricing cards and integration pages disagree on some quotas; model weighting and upstream plan rules apply.
MiMo Token Plan$6 / $16 / $50 / $1002 models / variants · 4 calculations
Estimated processed tokens / 30 days
- Lite, V2.5 daytime650.12M total (3.04M output)
- Standard, V2.5 daytime1.74B total (8.14M output)
- Pro, V2.5 daytime6.03B total (28.13M output)
- Max, V2.5 daytime13.00B total (60.71M output)
Allowance and practical limit
4.1B / 11B / 38B / 82B credits, with model-specific deductions. These are not raw tokens. Programming-tool restrictions apply.
Chutes$10 / $2014 models / variants · 2 calculations
Supported models / catalog
- Qwen/Qwen3-32B-TEE
- Qwen/Qwen3.5-397B-A17B-TEE
- google/gemma-4-31B-turbo-TEE
- zai-org/GLM-5.1-TEE
- deepseek-ai/DeepSeek-V3.2-TEE
- Qwen/Qwen3.6-27B-TEE
- moonshotai/Kimi-K2.6-TEE
- Qwen/Qwen3.8-27B-TEE
- deepseek-ai/DeepSeek-V4-Flash-0731-TEE
- zai-org/GLM-5.2-TEE
- moonshotai/Kimi-K3-TEE
- unsloth/Mistral-Nemo-Instruct-2407-TEE
- Qwen/Qwen3-235B-A22B-Thinking-2507-TEE
- Nemotron-3-Nano-Omni-30B-TEE
Estimated processed tokens / 30 days
- Plus, Flash conditional ceiling785.87M total (3.67M output)
- Pro, Flash conditional ceiling1.57B total (7.34M output)
Allowance and practical limit
Daily quotas and discounted overage. Published February policy caps subscription value at 5× its PAYG rates; confirm current rolling limits in the account.
NanoGPT$12 advertised296 models / variants · 2 calculations
Supported models / catalog
- GLM 5.3 Flash
- MiniMax M3
- MiMo V2.5
- MiMo V2.5 Pro
- GLM 5.3
- DeepSeek V4 Flash 0731
- DeepSeek V4 Pro 0813
Showing 7 compared routes from 296 recorded models / variants. Browse the complete catalog.
Estimated processed tokens / 30 days
- 1× input weight258.35M total (1.21M output)
- 2× input weight129.17M total (0.60M output)
Allowance and practical limit
Help lists 60M input-token units weekly, including cached input, with model multipliers. New subscription terms exclude commercial products; signup availability requires checking.
Kilo Pass$19 / $49 / $199371 models / variants · 3 calculations
Supported models / catalog
- Anthropic: Claude Sonnet 5
- Z.ai: GLM 5.3
- MoonshotAI: Kimi K3
- MiniMax: MiniMax M3
- Z.ai: GLM 5.3 Flash
- DeepSeek: DeepSeek V4 Pro 0813
- DeepSeek: DeepSeek V4 Flash 0731
- Xiaomi: MiMo-V2.5-Pro
- Xiaomi: MiMo-V2.5
Showing 9 compared routes from 371 recorded models / variants. Browse the complete catalog.
Estimated processed tokens / 30 days
- Starter, Sonnet 5 annual91.59M total (0.43M output)
- Pro, Sonnet 5 annual236.21M total (1.10M output)
- Expert, Sonnet 5 annual959.32M total (4.48M output)
Allowance and practical limit
Credits plus welcome/loyalty bonuses; annual billing gives 50% extra monthly credits. Full use yields a 33.3% effective discount. Bonus credits expire monthly; Gateway supported.
Several qualifications matter more than another decimal place in the price:
- Z.ai supports specific coding tools and uses dynamic concurrency limits. Its published examples include automated development tasks, but that does not establish unrestricted custom-backend access. Usage policy
- camelStream’s standard terms permit retention and training on prompts and outputs, without an opt-out. Model choice, queue time and throughput are not guaranteed. It merits a small experiment with suitable non-sensitive work. Terms
- Kimi Code’s personal benefits differ from production API access. The membership page alone does not establish a fixed Code token allowance or identical context limits across products. Code benefits
- NanoGPT’s new subscription terms explicitly exclude building commercial products. They apply to new users now and existing users from September 14. PAYG permits commercial use. Terms, quota accounting
- The DeepSeek reseller needs further clarification. Its pricing cards and DeepSeek integration page give different allowances. Its separate Agent Plan documentation describes an upstream Volcengine product. I have not treated either as an official DeepSeek subscription or assigned it a reliable savings multiplier.
- Chutes’ old unlimited-value anecdotes are obsolete. Its published revision caps monthly subscription benefit at five times its own PAYG value and permits shorter rolling limits.
The capacity CSV includes additional models, peak/off-peak variants and all four token buckets. The catalog snapshot preserves the large public model lists. Model availability, subscription inclusion and a selectable API alias are distinct; unresolved tier or alias details are marked in the comparison.
Normalize the subscription costs#
The same measured mix makes every quantifiable allowance comparable in token units: 1M billable output tokens requires approximately 213.167M input tokens, for 214.167M total processed tokens. The input is divided among the fresh, cache-read and cache-write buckets above before applying each provider’s prices or credit weights. Dividing a subscription’s fee by its advertised credit count would skip that conversion.
output_M = allowance_for_period / deductions_per_1M_output_and_its_input
processed_M = output_M × (1 + measured_input_output_ratio)
cost_per_1M_processed = fee_for_period / processed_M
cost_per_1M_output_with_input = fee_for_period / output_M
These are fully utilized, steady-state 30-day estimates from the dated September 7 price snapshot. Output costs include the associated input and cache processing; they are not the provider’s output-only API rate. Processed-token costs include cache reads, which dominate this workload. Both measures describe the same workload and produce the same cost ordering. Differences in model quality still need an accepted-work benchmark.
OpenCode Go: GoMiMo V2.5$0.32 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 6.80B
- Output / 30 days
- 31.73M
- $ / 1M processed
- $0.00147
- $ / 1M output + its input
- $0.32
- Saving vs direct API
- 83.3% less
OpenCode Go: GoMiniMax M3$2.63 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 815.16M
- Output / 30 days
- 3.81M
- $ / 1M processed
- $0.01227
- $ / 1M output + its input
- $2.63
- Saving vs direct API
- 83.3% less
OpenCode Go: GoKimi K3$65.94 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 32.48M
- Output / 30 days
- 0.15M
- $ / 1M processed
- $0.30788
- $ / 1M output + its input
- $65.94
- Saving vs direct API
- 33.3% less
Synthetic: One packGLM-5.3-Flash$2.87 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 2.24B
- Output / 30 days
- 10.45M
- $ / 1M processed
- $0.01340
- $ / 1M output + its input
- $2.87
- Saving vs direct API
- 63.1% less
Synthetic: One packKimi-K3$37.85 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 169.75M
- Output / 30 days
- 0.79M
- $ / 1M processed
- $0.17673
- $ / 1M output + its input
- $37.85
- Saving vs direct API
- 61.7% less
MiMo Token Plan: LiteMiMo V2.5 daytime$1.98 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 650.12M
- Output / 30 days
- 3.04M
- $ / 1M processed
- $0.00923
- $ / 1M output + its input
- $1.98
- Saving vs direct API
- 4.5% more expensive
MiMo Token Plan: StandardMiMo V2.5 daytime$1.96 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 1.74B
- Output / 30 days
- 8.14M
- $ / 1M processed
- $0.00917
- $ / 1M output + its input
- $1.96
- Saving vs direct API
- 3.9% more expensive
MiMo Token Plan: ProMiMo V2.5 daytime$1.78 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 6.03B
- Output / 30 days
- 28.13M
- $ / 1M processed
- $0.00830
- $ / 1M output + its input
- $1.78
- Saving vs direct API
- 6.0% less
MiMo Token Plan: MaxMiMo V2.5 daytime$1.65 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 13.00B
- Output / 30 days
- 60.71M
- $ / 1M processed
- $0.00769
- $ / 1M output + its input
- $1.65
- Saving vs direct API
- 12.9% less
Chutes: PlusDeepSeek V4 Flash · conditional ceiling$2.73 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 785.87M
- Output / 30 days
- 3.67M
- $ / 1M processed
- $0.01272
- $ / 1M output + its input
- $2.73
- Saving vs direct API
- 26.9% less
Chutes: ProDeepSeek V4 Flash · conditional ceiling$2.73 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 1.57B
- Output / 30 days
- 7.34M
- $ / 1M processed
- $0.01272
- $ / 1M output + its input
- $2.73
- Saving vs direct API
- 26.9% less
NanoGPT: ProIncluded model at 1x input weight$9.95 / 1M output
- Fee / month
- $12.00
- Processed / 30 days
- 258.35M
- Output / 30 days
- 1.21M
- $ / 1M processed
- $0.04645
- $ / 1M output + its input
- $9.95
- Saving vs direct API
- Not established
NanoGPT: ProIncluded model at 2x input weight$19.90 / 1M output
- Fee / month
- $12.00
- Processed / 30 days
- 129.17M
- Output / 30 days
- 0.60M
- $ / 1M processed
- $0.09290
- $ / 1M output + its input
- $19.90
- Saving vs direct API
- Not established
Kilo Pass: Starter annual billingAnthropic: Claude Sonnet 5$44.43 / 1M output
- Fee / month
- $19.00
- Processed / 30 days
- 91.59M
- Output / 30 days
- 0.43M
- $ / 1M processed
- $0.20744
- $ / 1M output + its input
- $44.43
- Saving vs direct API
- Not established
Kilo Pass: Pro annual billingAnthropic: Claude Sonnet 5$44.43 / 1M output
- Fee / month
- $49.00
- Processed / 30 days
- 236.21M
- Output / 30 days
- 1.10M
- $ / 1M processed
- $0.20744
- $ / 1M output + its input
- $44.43
- Saving vs direct API
- Not established
Kilo Pass: Expert annual billingAnthropic: Claude Sonnet 5$44.43 / 1M output
- Fee / month
- $199.00
- Processed / 30 days
- 959.32M
- Output / 30 days
- 4.48M
- $ / 1M processed
- $0.20744
- $ / 1M output + its input
- $44.43
- Saving vs direct API
- Not established
Z.ai Lite / Pro / Max: LiteGLM-5.3 (off-peak)$8.92 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 432.12M
- Output / 30 days
- 2.02M
- $ / 1M processed
- $0.04166
- $ / 1M output + its input
- $8.92
- Saving vs direct API
- 86.9% less
Z.ai Lite / Pro / Max: ProGLM-5.3 (off-peak)$6.61 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 2.59B
- Output / 30 days
- 12.11M
- $ / 1M processed
- $0.03086
- $ / 1M output + its input
- $6.61
- Saving vs direct API
- 90.3% less
Z.ai Lite / Pro / Max: MaxGLM-5.3 (off-peak)$5.95 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 6.05B
- Output / 30 days
- 28.25M
- $ / 1M processed
- $0.02777
- $ / 1M output + its input
- $5.95
- Saving vs direct API
- 91.3% less
The comparison covers the quantifiable examples from the plan overview. Other model allocations and time windows use the same calculation:
Show the remaining model and time-of-day calculations
OpenCode Go: GoDeepSeek V4 Flash off-peak$1.24 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 1.72B
- Output / 30 days
- 8.05M
- $ / 1M processed
- $0.00580
- $ / 1M output + its input
- $1.24
- Saving vs direct API
- 66.7% less
OpenCode Go: GoGLM-5.3$45.50 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 47.07M
- Output / 30 days
- 0.22M
- $ / 1M processed
- $0.21245
- $ / 1M output + its input
- $45.50
- Saving vs direct API
- 33.3% less
OpenCode Go: GoGPT-5.6 Luna$4.58 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 468.02M
- Output / 30 days
- 2.19M
- $ / 1M processed
- $0.02137
- $ / 1M output + its input
- $4.58
- Saving vs direct API
- Not established
Synthetic: One packgpt-oss-120b$1.45 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 4.45B
- Output / 30 days
- 20.76M
- $ / 1M processed
- $0.00675
- $ / 1M output + its input
- $1.45
- Saving vs direct API
- Not established
Synthetic: One packGLM-5.2$12.63 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 508.57M
- Output / 30 days
- 2.37M
- $ / 1M processed
- $0.05899
- $ / 1M output + its input
- $12.63
- Saving vs direct API
- Not established
Synthetic: One packQwen3.8-27B$7.01 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 916.11M
- Output / 30 days
- 4.28M
- $ / 1M processed
- $0.03275
- $ / 1M output + its input
- $7.01
- Saving vs direct API
- Not established
Synthetic: One packGLM-4.7-Flash$1.56 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 4.11B
- Output / 30 days
- 19.21M
- $ / 1M processed
- $0.00729
- $ / 1M output + its input
- $1.56
- Saving vs direct API
- Not established
Synthetic: One packNVIDIA-Nemotron-3-Super-120B-A12B-NVFP4$4.54 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 1.42B
- Output / 30 days
- 6.61M
- $ / 1M processed
- $0.02120
- $ / 1M output + its input
- $4.54
- Saving vs direct API
- Not established
MiMo Token Plan: LiteMiMo V2.5 off-peak$1.58 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 812.66M
- Output / 30 days
- 3.79M
- $ / 1M processed
- $0.00738
- $ / 1M output + its input
- $1.58
- Saving vs direct API
- 16.4% less
MiMo Token Plan: LiteMiMo V2.5 Pro daytime$4.88 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 263.55M
- Output / 30 days
- 1.23M
- $ / 1M processed
- $0.02277
- $ / 1M output + its input
- $4.88
- Saving vs direct API
- 1.0% more expensive
MiMo Token Plan: LiteMiMo V2.5 Pro off-peak$3.90 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 329.44M
- Output / 30 days
- 1.54M
- $ / 1M processed
- $0.01821
- $ / 1M output + its input
- $3.90
- Saving vs direct API
- 19.2% less
MiMo Token Plan: StandardMiMo V2.5 off-peak$1.57 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 2.18B
- Output / 30 days
- 10.18M
- $ / 1M processed
- $0.00734
- $ / 1M output + its input
- $1.57
- Saving vs direct API
- 16.9% less
MiMo Token Plan: StandardMiMo V2.5 Pro daytime$4.85 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 707.10M
- Output / 30 days
- 3.30M
- $ / 1M processed
- $0.02263
- $ / 1M output + its input
- $4.85
- Saving vs direct API
- 0.4% more expensive
MiMo Token Plan: StandardMiMo V2.5 Pro off-peak$3.88 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 883.87M
- Output / 30 days
- 4.13M
- $ / 1M processed
- $0.01810
- $ / 1M output + its input
- $3.88
- Saving vs direct API
- 19.7% less
MiMo Token Plan: ProMiMo V2.5 off-peak$1.42 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 7.53B
- Output / 30 days
- 35.17M
- $ / 1M processed
- $0.00664
- $ / 1M output + its input
- $1.42
- Saving vs direct API
- 24.8% less
MiMo Token Plan: ProMiMo V2.5 Pro daytime$4.38 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 2.44B
- Output / 30 days
- 11.41M
- $ / 1M processed
- $0.02047
- $ / 1M output + its input
- $4.38
- Saving vs direct API
- 9.2% less
MiMo Token Plan: ProMiMo V2.5 Pro off-peak$3.51 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 3.05B
- Output / 30 days
- 14.26M
- $ / 1M processed
- $0.01638
- $ / 1M output + its input
- $3.51
- Saving vs direct API
- 27.3% less
MiMo Token Plan: MaxMiMo V2.5 off-peak$1.32 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 16.25B
- Output / 30 days
- 75.89M
- $ / 1M processed
- $0.00615
- $ / 1M output + its input
- $1.32
- Saving vs direct API
- 30.3% less
MiMo Token Plan: MaxMiMo V2.5 Pro daytime$4.06 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 5.27B
- Output / 30 days
- 24.61M
- $ / 1M processed
- $0.01897
- $ / 1M output + its input
- $4.06
- Saving vs direct API
- 15.8% less
MiMo Token Plan: MaxMiMo V2.5 Pro off-peak$3.25 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 6.59B
- Output / 30 days
- 30.77M
- $ / 1M processed
- $0.01518
- $ / 1M output + its input
- $3.25
- Saving vs direct API
- 32.6% less
Kilo Pass: Starter annual billingMiniMax: MiniMax M3$10.51 / 1M output
- Fee / month
- $19.00
- Processed / 30 days
- 387.20M
- Output / 30 days
- 1.81M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
Kilo Pass: Pro annual billingMiniMax: MiniMax M3$10.51 / 1M output
- Fee / month
- $49.00
- Processed / 30 days
- 998.57M
- Output / 30 days
- 4.66M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
Kilo Pass: Expert annual billingMiniMax: MiniMax M3$10.51 / 1M output
- Fee / month
- $199.00
- Processed / 30 days
- 4.06B
- Output / 30 days
- 18.94M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
Z.ai Lite / Pro / Max: LiteGLM-5.3 (peak)$17.84 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 216.06M
- Output / 30 days
- 1.01M
- $ / 1M processed
- $0.08331
- $ / 1M output + its input
- $17.84
- Saving vs direct API
- 73.9% less
Z.ai Lite / Pro / Max: LiteGLM-5.3-Flash list (off-peak)$2.94 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 1.31B
- Output / 30 days
- 6.11M
- $ / 1M processed
- $0.01375
- $ / 1M output + its input
- $2.94
- Saving vs direct API
- 62.2% less
Z.ai Lite / Pro / Max: LiteGLM-5.3-Flash list (peak)$5.89 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 654.52M
- Output / 30 days
- 3.06M
- $ / 1M processed
- $0.02750
- $ / 1M output + its input
- $5.89
- Saving vs direct API
- 24.3% less
Z.ai Lite / Pro / Max: ProGLM-5.3 (peak)$13.22 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 1.30B
- Output / 30 days
- 6.05M
- $ / 1M processed
- $0.06171
- $ / 1M output + its input
- $13.22
- Saving vs direct API
- 80.6% less
Z.ai Lite / Pro / Max: ProGLM-5.3-Flash list (off-peak)$2.18 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 7.85B
- Output / 30 days
- 36.67M
- $ / 1M processed
- $0.01019
- $ / 1M output + its input
- $2.18
- Saving vs direct API
- 72.0% less
Z.ai Lite / Pro / Max: ProGLM-5.3-Flash list (peak)$4.36 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 3.93B
- Output / 30 days
- 18.34M
- $ / 1M processed
- $0.02037
- $ / 1M output + its input
- $4.36
- Saving vs direct API
- 43.9% less
Z.ai Lite / Pro / Max: MaxGLM-5.3 (peak)$11.89 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 3.02B
- Output / 30 days
- 14.12M
- $ / 1M processed
- $0.05554
- $ / 1M output + its input
- $11.89
- Saving vs direct API
- 82.6% less
Z.ai Lite / Pro / Max: MaxGLM-5.3-Flash list (off-peak)$1.96 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 18.33B
- Output / 30 days
- 85.57M
- $ / 1M processed
- $0.00917
- $ / 1M output + its input
- $1.96
- Saving vs direct API
- 74.8% less
Z.ai Lite / Pro / Max: MaxGLM-5.3-Flash list (peak)$3.93 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 9.16B
- Output / 30 days
- 42.79M
- $ / 1M processed
- $0.01833
- $ / 1M output + its input
- $3.93
- Saving vs direct API
- 49.5% less
“Saving vs direct API” compares the same named model against the direct API row below, using ordinary list prices for GLM Flash and off-peak prices for DeepSeek. Serving configurations and cache behavior can differ. “Not established” means this dataset lacks a defensible same-model baseline; it does not mean zero savings. Chutes estimates remain conditional on its published ceiling and model eligibility. Kilo examples assume annual billing with the full monthly bonus consumed before expiry. Alternative model allocations within a subscription cannot be added together.
The capacity CSV contains all four token buckets, both normalized costs, direct API savings and break-even utilization. It also separates a discount against the serving provider’s own PAYG rates from a discount against the original model API. Missing or conflicting quotas remain blank; camelStream’s unmetered service has no defensible fixed token capacity or normalized cost here. MiniMax, Kimi and the DeepSeek reseller need quota clarification; Alibaba needs tokens per billable request.
At 50% utilization, each subscription’s effective cost per token doubles. Break-even utilization is the fraction of modeled allowance that must replace direct API spending to recover the fee. A figure above 100% means even full utilization is more expensive than that API baseline. Real concurrency, request and reset limits can prevent reaching the modeled allowance.
Start with the API baseline#
The priced workload is 1M billable output tokens plus the measured proportions of fresh input, cache reads and cache writes shown above. Tokenizers differ, so equal token counts do not mean identical amounts of source code. These are estimates at current prices, not the historical cash cost of the measured sessions.
For prices stated per million tokens:
cost = fresh_input_M × input_price
+ cache_read_M × cache_read_price
+ cache_write_M × cache_creation_price
+ output_M × output_price
+ additional_charges
The comparison excludes paid tools, search, additional retries, infrastructure and long-context premiums. Cache creation uses the ordinary input price unless the serving provider specifies a separate full creation rate. Free cache storage or a waived write surcharge does not make the initial input processing free. Kilo’s Sonnet 5 example, for instance, uses its explicit cache-creation rate in the capacity calculator.
MiMo V2.5$1.89 / 1M output
- Input / M
- $0.14
- Cached / M
- $0.0028
- Output / M
- $0.28
- Cost per 1M output + its input
- $1.89
- $100 processed capacity
- 11.33B total (52.88M output)
DeepSeek V4 Flash off-peak$3.73 / 1M output
- Input / M
- $0.22
- Cached / M
- $0.007
- Output / M
- $0.66
- Cost per 1M output + its input
- $3.73
- $100 processed capacity
- 5.75B total (26.84M output)
GLM-5.3-Flash promotional$3.89 / 1M output
- Input / M
- $0.075
- Cached / M
- $0.015
- Output / M
- $0.25
- Cost per 1M output + its input
- $3.89
- $100 processed capacity
- 5.50B total (25.70M output)
MiMo V2.5 Pro$4.83 / 1M output
- Input / M
- $0.435
- Cached / M
- $0.0036
- Output / M
- $0.87
- Cost per 1M output + its input
- $4.83
- $100 processed capacity
- 4.44B total (20.72M output)
GLM-5.3-Flash list$7.78 / 1M output
- Input / M
- $0.15
- Cached / M
- $0.03
- Output / M
- $0.50
- Cost per 1M output + its input
- $7.78
- $100 processed capacity
- 2.75B total (12.85M output)
DeepSeek V4 Pro off-peak$11.39 / 1M output
- Input / M
- $0.66
- Cached / M
- $0.022
- Output / M
- $1.98
- Cost per 1M output + its input
- $11.39
- $100 processed capacity
- 1.88B total (8.78M output)
MiniMax M3$15.76 / 1M output
- Input / M
- $0.3
- Cached / M
- $0.06
- Output / M
- $1.20
- Cost per 1M output + its input
- $15.76
- $100 processed capacity
- 1.36B total (6.34M output)
GLM-5.3$68.25 / 1M output
- Input / M
- $1.4
- Cached / M
- $0.26
- Output / M
- $4.40
- Cost per 1M output + its input
- $68.25
- $100 processed capacity
- 313.80M total (1.47M output)
Kimi K3$98.91 / 1M output
- Input / M
- $3
- Cached / M
- $0.3
- Output / M
- $15.00
- Cost per 1M output + its input
- $98.91
- $100 processed capacity
- 216.54M total (1.01M output)
Without cache hits, the same workload costs $30.12 on MiMo V2.5, $47.56 on off-peak DeepSeek Flash, and $654.50 on Kimi K3. A plan that counts cached input at full token weight can lose much of its advantage against an API with inexpensive cache reads. Repeated context also inflates token totals without representing newly generated work.
DeepSeek’s listed prices apply outside weekday peak windows of 01:00–04:00 and 06:00–10:00 UTC; peak rates double. MiniMax M3’s listed rate applies through 512K input per request. These timing and context conditions belong in the comparison, not in a footnote to the invoice. DeepSeek pricing, MiniMax pricing
Compare with OpenRouter too#
OpenRouter narrows some subscription discounts because its alternative providers can undercut the model author’s API. The result depends heavily on cached-input pricing. The following snapshot was captured on September 7, 2026, through 16:11 UTC from OpenRouter’s public model-endpoint API. Each quote keeps one provider’s input, cache and output rates together.
The comparison is grouped by subscription, then model, then tier or time window. All twelve researched coding subscriptions appear, including those with unknown costs. Models from the saved catalogs remain visible in the index; large unpriced catalogs expand underneath their subscription. Catalog presence, subscription eligibility and model pinning are separate questions, so each group retains those conditions. Native workflow subscriptions are discussed separately below.
The unit-cost fields report cash cost per 1M output tokens plus 213.167M associated input tokens, using the measured cache mix. Subscription summaries show the monthly fee, and capacity details show total processed tokens with output in parentheses. Each calculated scenario spends the full allowance on its named model; scenarios within one subscription are alternative allocations. “Not established” means the snapshot lacks a defensible entitlement or calculation; “Not quoted” means no matching OpenRouter route was priced.
OpenRouter figures include its standard card-purchase fee: 5.5%, with a $0.80 minimum. The reference purchase is $100 of inference credits plus $5.50 in fees, so quoted inference cost is multiplied by 1.055 once. A $10 credit purchase instead incurs an 8% effective fee. Taxes and special billing arrangements are excluded. OpenRouter fee policy
Jump to subscription: OpenCode Go · Z.ai Lite / Pro / Max · Synthetic · camelStream · MiniMax Token Plan · Kimi membership / Code · Alibaba Cloud Coding Plan · deepseekv4pro.com · MiMo Token Plan · Chutes · NanoGPT · Kilo Pass
OpenCode Go27 recorded models / aliases$10 / month
One shared allowance. Regional and experimental conditions apply; uncalculated models retain unknown capacity.
MiMo V2.5$0.32 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 6.80B (31.73M output)
- Subscription $ / 1M output + input
- $0.32
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 84.2% less
MiniMax M3$2.63 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 815.16M (3.81M output)
- Subscription $ / 1M output + input
- $2.63
- OpenRouter $ / 1M output + input
- $15.50DeepInfra, FP8
- Subscription saving
- 83.1% less
Kimi K3$65.94 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 32.48M (0.15M output)
- Subscription $ / 1M output + input
- $65.94
- OpenRouter $ / 1M output + input
- $96.95Sail Research, FP4
- Subscription saving
- 32.0% less
DeepSeek V4 Flash$1.24 subscription
- Tier / conditions
- Go; off-peak
- Processed / 30 days (output)
- 1.72B (8.05M output)
- Subscription $ / 1M output + input
- $1.24
- OpenRouter $ / 1M output + input
- $2.36StreamLake, FP8
- Subscription saving
- 47.3% less
GLM-5.3$45.50 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 47.07M (0.22M output)
- Subscription $ / 1M output + input
- $45.50
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- 1.0% less
GPT-5.6 Luna$4.58 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 468.02M (2.19M output)
- Subscription $ / 1M output + input
- $4.58
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiMo V2.5 Pro$21.21 OpenRouter
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- Not established
GLM-5.3-Flash$3.90 OpenRouter
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- Not established
DeepSeek V4 Pro$12.01 OpenRouter
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $12.01DeepSeek
- Subscription saving
- Not established
Grok 4.6Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GLM-5.2Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GLM-5.1Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Kimi K2.7 CodePricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Kimi K2.6Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax M2.7Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.8 MaxPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.8 FlashPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.7 MaxPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.7 PlusPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.6 PlusPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
LongCat 2.0Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Muse Spark 1.3 ContributorPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Muse Spark 1.2 ContributorPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
DeepSeek V4 Flash Vision ExpPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Hy4Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Hy3Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Omen AlphaPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Z.ai Lite / Pro / Max2 recorded models / aliases$18 / $80 / $168 / month
Older aliases route to these models. Each row allocates the entire tier allowance to one model and time window.
GLM-5.36 scenarios
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 432.12M (2.02M output)
- Subscription $ / 1M output + input
- $8.92
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- 80.6% less
- Tier / conditions
- Lite; peak
- Processed / 30 days (output)
- 216.06M (1.01M output)
- Subscription $ / 1M output + input
- $17.84
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- 61.2% less
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 2.59B (12.11M output)
- Subscription $ / 1M output + input
- $6.61
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- 85.6% less
- Tier / conditions
- Pro; peak
- Processed / 30 days (output)
- 1.30B (6.05M output)
- Subscription $ / 1M output + input
- $13.22
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- 71.2% less
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 6.05B (28.25M output)
- Subscription $ / 1M output + input
- $5.95
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- 87.1% less
- Tier / conditions
- Max; peak
- Processed / 30 days (output)
- 3.02B (14.12M output)
- Subscription $ / 1M output + input
- $11.89
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- 74.1% less
GLM-5.3-Flash6 scenarios
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 1.31B (6.11M output)
- Subscription $ / 1M output + input
- $2.94
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- 24.5% less
- Tier / conditions
- Lite; peak
- Processed / 30 days (output)
- 654.52M (3.06M output)
- Subscription $ / 1M output + input
- $5.89
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- 51.0% more expensive
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 7.85B (36.67M output)
- Subscription $ / 1M output + input
- $2.18
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- 44.1% less
- Tier / conditions
- Pro; peak
- Processed / 30 days (output)
- 3.93B (18.34M output)
- Subscription $ / 1M output + input
- $4.36
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- 11.9% more expensive
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 18.33B (85.57M output)
- Subscription $ / 1M output + input
- $1.96
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- 49.7% less
- Tier / conditions
- Max; peak
- Processed / 30 days (output)
- 9.16B (42.79M output)
- Subscription $ / 1M output + input
- $3.93
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- 0.7% more expensive
Synthetic11 recorded models / aliases$30 per pack / month
The saved served catalog includes four automatic aliases. Catalog availability alone does not establish subscription eligibility.
syn:large:textPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
syn:small:textPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
syn:large:visionPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
syn:small:visionPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
openai/gpt-oss-120b$1.45 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 4.45B (20.76M output)
- Subscription $ / 1M output + input
- $1.45
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
zai-org/GLM-5.3-Flash$2.87 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 2.24B (10.45M output)
- Subscription $ / 1M output + input
- $2.87
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- 26.4% less
zai-org/GLM-5.2$12.63 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 508.57M (2.37M output)
- Subscription $ / 1M output + input
- $12.63
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
moonshotai/Kimi-K3$37.85 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 169.75M (0.79M output)
- Subscription $ / 1M output + input
- $37.85
- OpenRouter $ / 1M output + input
- $96.95Sail Research, FP4
- Subscription saving
- 61.0% less
Qwen/Qwen3.8-27B$7.01 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 916.11M (4.28M output)
- Subscription $ / 1M output + input
- $7.01
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
zai-org/GLM-4.7-Flash$1.56 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 4.11B (19.21M output)
- Subscription $ / 1M output + input
- $1.56
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4$4.54 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 1.42B (6.61M output)
- Subscription $ / 1M output + input
- $4.54
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
camelStream5 recorded models / aliases$5 per stream / month
Unmetered service has no fixed token entitlement. Named routes are possible automatic choices, not selectable allocations.
DeepSeek V4 Flash$2.36 OpenRouter
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.36StreamLake, FP8
- Subscription saving
- Not established
Gemini 3.7 FlashPricing not established
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GLM-5.3-Flash$3.90 OpenRouter
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- Not established
GPT-5.6 LunaPricing not established
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Muse Spark 1.2Pricing not established
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax Token Plan2 recorded models / aliases$22 / $55 / $132 / month
Text models only. Image and speech share quota but are outside this token comparison; published numerical allowance is insufficient.
MiniMax M3$15.50 OpenRouter
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $15.50DeepInfra, FP8
- Subscription saving
- Not established
MiniMax M2.7Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Kimi membership / Code3 recorded models / aliases$19 / $39 / $99 / $199 / month
No fixed comparable entitlement established. Membership K3 and Code aliases are distinguished; alias costs are not assumed to equal K3.
Kimi K3$96.95 OpenRouter
- Tier / conditions
- Membership; K3 advertised across Kimi
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $96.95Sail Research, FP4
- Subscription saving
- Not established
kimi-for-codingPricing not established
- Tier / conditions
- Code alias; backing version not pinned
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
kimi-for-coding-highspeedPricing not established
- Tier / conditions
- Code HighSpeed alias; Allegretto+; backing version not pinned
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Alibaba Cloud Coding Plan10 recorded models / aliases$50 / month
Request quotas cannot be converted without tokens per billable request. Automated scripts and application backends are excluded.
qwen3.7-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3.6-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
kimi-k2.5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
glm-5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax-M2.5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3.5-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3-max-2026-01-23Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3-coder-nextPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3-coder-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
glm-4.7Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
deepseekv4pro.com6 recorded models / aliases$19.90 / $49.90; Agent $29.90 / month
Independent reseller; published quotas conflict. Models added by the Agent plan are marked separately.
DeepSeek V4 Flash$2.36 OpenRouter
- Tier / conditions
- DeepSeek plans / Agent plan
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.36StreamLake, FP8
- Subscription saving
- Not established
DeepSeek V4 Pro$12.01 OpenRouter
- Tier / conditions
- DeepSeek plans / Agent plan
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $12.01DeepSeek
- Subscription saving
- Not established
GLM-5.3$45.95 OpenRouter
- Tier / conditions
- Agent plan only
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- Not established
GLM-5.2Pricing not established
- Tier / conditions
- Agent plan only
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax M3$15.50 OpenRouter
- Tier / conditions
- Agent plan only
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $15.50DeepInfra, FP8
- Subscription saving
- Not established
Kimi K3$96.95 OpenRouter
- Tier / conditions
- Agent plan only
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $96.95Sail Research, FP4
- Subscription saving
- Not established
MiMo Token Plan2 recorded models / aliases$6 / $16 / $50 / $100 / month
Text models only. Credit pools are shared across models; daytime and off-peak rows are alternative allocations.
MiMo V2.58 scenarios
- Tier / conditions
- Lite; daytime
- Processed / 30 days (output)
- 650.12M (3.04M output)
- Subscription $ / 1M output + input
- $1.98
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 0.9% less
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 812.66M (3.79M output)
- Subscription $ / 1M output + input
- $1.58
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 20.7% less
- Tier / conditions
- Standard; daytime
- Processed / 30 days (output)
- 1.74B (8.14M output)
- Subscription $ / 1M output + input
- $1.96
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 1.5% less
- Tier / conditions
- Standard; off-peak
- Processed / 30 days (output)
- 2.18B (10.18M output)
- Subscription $ / 1M output + input
- $1.57
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 21.2% less
- Tier / conditions
- Pro; daytime
- Processed / 30 days (output)
- 6.03B (28.13M output)
- Subscription $ / 1M output + input
- $1.78
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 10.9% less
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 7.53B (35.17M output)
- Subscription $ / 1M output + input
- $1.42
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 28.7% less
- Tier / conditions
- Max; daytime
- Processed / 30 days (output)
- 13.00B (60.71M output)
- Subscription $ / 1M output + input
- $1.65
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 17.4% less
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 16.25B (75.89M output)
- Subscription $ / 1M output + input
- $1.32
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 33.9% less
MiMo V2.5 Pro8 scenarios
- Tier / conditions
- Lite; daytime
- Processed / 30 days (output)
- 263.55M (1.23M output)
- Subscription $ / 1M output + input
- $4.88
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 77.0% less
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 329.44M (1.54M output)
- Subscription $ / 1M output + input
- $3.90
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 81.6% less
- Tier / conditions
- Standard; daytime
- Processed / 30 days (output)
- 707.10M (3.30M output)
- Subscription $ / 1M output + input
- $4.85
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 77.2% less
- Tier / conditions
- Standard; off-peak
- Processed / 30 days (output)
- 883.87M (4.13M output)
- Subscription $ / 1M output + input
- $3.88
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 81.7% less
- Tier / conditions
- Pro; daytime
- Processed / 30 days (output)
- 2.44B (11.41M output)
- Subscription $ / 1M output + input
- $4.38
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 79.3% less
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 3.05B (14.26M output)
- Subscription $ / 1M output + input
- $3.51
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 83.5% less
- Tier / conditions
- Max; daytime
- Processed / 30 days (output)
- 5.27B (24.61M output)
- Subscription $ / 1M output + input
- $4.06
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 80.8% less
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 6.59B (30.77M output)
- Subscription $ / 1M output + input
- $3.25
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- 84.7% less
Chutes14 recorded models / aliases$10 / $20 / month
Served catalog; exact subscription tier eligibility needs account confirmation. Flash capacity remains a conditional ceiling.
Qwen/Qwen3-32B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3.5-397B-A17B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
google/gemma-4-31B-turbo-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
zai-org/GLM-5.1-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
deepseek-ai/DeepSeek-V3.2-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3.6-27B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
moonshotai/Kimi-K2.6-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3.8-27B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
deepseek-ai/DeepSeek-V4-Flash-0731-TEE2 scenarios
- Tier / conditions
- Plus; conditional ceiling
- Processed / 30 days (output)
- 785.87M (3.67M output)
- Subscription $ / 1M output + input
- $2.73
- OpenRouter $ / 1M output + input
- $2.36StreamLake, FP8
- Subscription saving
- 15.5% more expensive
- Tier / conditions
- Pro; conditional ceiling
- Processed / 30 days (output)
- 1.57B (7.34M output)
- Subscription $ / 1M output + input
- $2.73
- OpenRouter $ / 1M output + input
- $2.36StreamLake, FP8
- Subscription saving
- 15.5% more expensive
zai-org/GLM-5.2-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
moonshotai/Kimi-K3-TEE$96.95 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $96.95Sail Research, FP4
- Subscription saving
- Not established
unsloth/Mistral-Nemo-Instruct-2407-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3-235B-A22B-Thinking-2507-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Nemotron-3-Nano-Omni-30B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
NanoGPT296 recorded models / aliases$12 advertised / month
Subscription-included model/variant catalog. Per-model multipliers were not captured, so the 1x/2x quota examples cannot be assigned to individual IDs.
Unassigned quota examples
- Included model at 1x input weight$9.95 / 1M output + input
- Included model at 2x input weight$19.90 / 1M output + input
Per-model weights remain unknown.
GLM 5.3 Flash$3.90 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- Not established
MiniMax M3$15.50 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $15.50DeepInfra, FP8
- Subscription saving
- Not established
MiMo V2.5$1.99 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- Not established
MiMo V2.5 Pro$21.21 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- Not established
GLM 5.3$45.95 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- Not established
DeepSeek V4 Flash 0731$2.36 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.36StreamLake, FP8
- Subscription saving
- Not established
DeepSeek V4 Pro 0813$12.01 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $12.01DeepSeek
- Subscription saving
- Not established
Show 289 additional NanoGPT catalog models / variants
These IDs have no matched allowance or OpenRouter quote in this snapshot. Catalog coverage: Subscription-included endpoint.
- Muse Spark 1.3 Contributor
- Mercury 2.5 Preview
- Granite 4.2 8B
- TheDrummer/Artemis v1.1
- DeepSeek V4 Flash Vision Exp Uncensored
- GLM 5.3 Flash Uncensored
- Gemma 4 26B A4B Uncensored
- Gemma 4 26B A4B Uncensored Thinking
- Qwen 3.8 27B Uncensored
- Qwen 3.8 27B Fable
- Qwen 3.8 27B Uncensored Thinking
- Qwen 3.8 27B Obliterated
- Qwen 3.8 27B Obliterated Thinking
- Qwen 3.6 35B A3B Uncensored
- Qwen 3.6 35B A3B Uncensored Thinking
- Ornith 1.5 9B
- Ornith 1.5 9B Thinking
- Gemma 4 31B MeroMero v2
- Gemma 4 31B MeroMero v2 Thinking
- Gemma 4 26B A4B MeroMero
- Gemma 4 26B A4B MeroMero Thinking
- Gemma 4 26B A4B Musica
- Gemma 4 26B A4B Shadow Siren
- Gemma 4 26B A4B Chimera X
- Gemma 4 26B A4B Luminous Mirror
- Gemma 4 26B A4B Dark Soul
- Gemma 4 26B A4B Moonlight Dusk
- Gemma 4 26B A4B Opus Distill
- Gemma 4 31B Fabled
- Gemma 4 31B DarkIdol
- Gemma 4 31B Garnet
- Gemma 4 31B Novelist
- Gemma 4 31B Isometry
- Gemma 4 31B Gembrain
- Gemma 4 31B Gemsicle
- Ornith 1.5 35B
- Ornith 1.5 35B Thinking
- Qwen3.8 27B
- Qwen3.8 27B Thinking
- Gemma 4 12B Instruct
- Gemma 4 E4B Instruct
- Gemma 4 E2B Instruct
- Laguna S 2.1
- Laguna S 2.1 Thinking
- Nvidia Nemotron 3.5 Lightning
- Nvidia Nemotron 3.5 Lightning Thinking
- Nvidia Nemotron 3 Ultra 550B
- Nvidia Nemotron 3 Ultra 550B Thinking
- LFM2.5 2.6B
- Muse Glimmer 30B
- Muse Spark 1.2 Contributor
- LongCat 2.0
- LongCat 2.0 Thinking
- Step 3.7 Flash Thinking
- Doubao Seed Character
- NanoGPT Help
- Auto model
- Auto model (Basic)
- Auto model (Standard)
- Auto model (Premium)
- Claw High
- Claw Medium
- Claw Low
- Hermes High
- Hermes Medium
- Hermes Low
- GPT OSS 120B
- GPT OSS 20B
- Amoral Gemma3 27B v2
- Mistral Devstral Small 2505
- Veiled Calla 12B
- Qwen: QvQ Max
- Step 3.5 Flash 2603
- Step 3.5 Flash
- Nex N2 Mini
- Qwen 3 Coder 480B
- Llama 4 Maverick
- Llama 4 Scout
- DeepSeek R1 0528
- Kimi K2 Thinking
- Kimi K2.5
- Kimi K2.5 Thinking
- Kimi K2.6
- Kimi K2.7 Code
- Kimi K2.6 Thinking
- Ministral 3 14B
- Mistral Small 4 119B
- Mistral Small 4 119B Thinking
- Mistral Large 3 675B
- Devstral 2 123B
- Hermes 4 Large (Thinking)
- OpenReasoning Nemotron 32B
- DeepSeek R1
- DeepSeek V3/Deepseek Chat
- MiniMax 01
- MiniMax M2.5
- Qwen 3 235b A22B
- Qwen3.5 9B
- Qwen 3 32b
- Qwen 3 14b
- Qwen3 30B A3B
- Qwen3 Coder 30B A3B Instruct
- Qwen 3 235b A22B 2507
- Qwen 3 235b A22B 2507 Thinking
- Qwen3 Next 80B A3B (Instruct)
- Qwen3 Next 80B A3B (Thinking)
- MiniMax M2
- MiniMax M3 Thinking
- MiniMax M2.7
- MiniMax Latest
- MiMo V2.5 Thinking
- MiMo V2.5 Pro Thinking
- MiniMax M2.1
- GLM 4.6
- GLM 4.6 Thinking
- GLM 5
- GLM 5 Thinking
- GLM 5.1
- GLM Latest
- GLM 5.1 Thinking
- GLM 5.2
- GLM 5.2 Thinking
- GLM 5.3 Thinking
- GLM 4.7 Flash Original
- GLM 4.7 Flash Original Thinking
- GLM 4.7 Flash
- GLM 4.7 Flash Thinking
- GLM 4.7
- GLM 4.7 Thinking
- GLM 4.6V
- Qwen3 30B A3B Instruct 2507
- Llama 3.3 70b Instruct
- Nvidia Nemotron 70b
- Sao10K Stheno 8b
- Grayline Qwen3 8B
- Hermes 4 Large
- Hermes 3 70B
- Qwen3.8 Flash
- Qwen3.5 122B A10B
- Qwen3.5 122B A10B Thinking
- Qwen3.5 27B
- Qwen3.5 27B Thinking
- Qwen3.5 35B A3B
- Qwen3.5 35B A3B Thinking
- Qwen3.6 35B A3B
- Qwen3.6 35B A3B Thinking
- Qwen3.6 27B
- Qwen3.6 27B Thinking
- DeepSeek V3.2 Exp
- DeepSeek V3.2 Exp Thinking
- DeepSeek V3.2
- DeepSeek V3.2 Thinking
- DeepSeek V4 Flash
- DeepSeek V4 Flash Vision Exp
- DeepSeek V4 Flash Latest
- DeepSeek V4 Flash 0731 (Thinking)
- DeepSeek V4 Flash (Thinking)
- DeepSeek V4 Pro 0813 Thinking
- DeepSeek V4 Pro
- DeepSeek Latest
- DeepSeek V4 Pro (Thinking)
- Qwen3.5 397B A17B
- Qwen3.5 397B A17B Thinking
- Qwen 2.5 Coder 32b
- Phi 4 Multimodal
- Phi 4 Mini
- The Drummer Cydonia 24B v2
- The Drummer Cydonia 24B v4
- The Drummer Cydonia 24B v4.1
- The Drummer Cydonia 24B v4.3
- The Drummer Magidonia 24B v4.3
- MS3.2 24B Magnum Diamond
- Omega Directive 24B Unslop v2.0
- EVA Llama 3.33 70B
- Steelskull Nevoria 70b
- Steelskull Nevoria R1 70b
- Steelskull Electra R1 70b
- Qwen2 72B Dracarys
- Lumimaid v0.2
- DeepSeek V3/Chat Cheaper
- Llama 3.3 70B Instruct abliterated
- MythoMax 13B
- Qwen2.5 72B
- EVA-Qwen2.5-32B-v0.2
- TheDrummer Skyfall 36B V2
- Qwen 3 8B
- K2-Think
- DeepSeek V3.1
- DeepSeek V3.1 Thinking
- DeepSeek V3.1 Terminus
- DeepSeek V3.1 Terminus (Thinking)
- DeepSeek Chat 0324
- GLM 4.5 (Thinking)
- GLM 4.5
- GLM 4.5 Air
- GLM 4.5 Air (Thinking)
- MN-LooseCannon-12B-v1
- EVA-Qwen2.5-72B-v0.2
- EVA-LLaMA-3.33-70B-v0.1
- Llama 3.1 8b Instruct
- ReMM SLERP 13B
- Mistral Saba
- Neural Daredevil 8B abliterated
- Llama 3 70B abliterated
- Magnum V2 72B
- Mistral Nemo
- DeepSeek Reasoner
- Llama 3.05 Storybreaker Ministral 70b
- Nemotron Tenyxchat Storybreaker 70b
- Mag Mell R1
- Qwerky 72B
- Anubis 70B v1
- Anubis 70B v1.1
- Llama 3.2 3b Instruct
- Llama 3.1 8B (decentralized)
- Llama 3.1 70B Hanami
- Rocinante 12b
- Llama 3.3 70B Euryale
- Llama 3.1 70B Euryale
- Llama 3.3 70B Cu Mai
- UnslopNemo 12b v4
- NemoMix 12B Unleashed
- Mistral Nemo Starcannon 12b v1
- Llama 3.1 70B Celeste v0.1
- DeepSeek R1 Qwen Abliterated
- DeepSeek R1 Llama 70B Abliterated
- Qwen 2.5 32B Abliterated
- Deepseek R1 Cheaper
- Llama 3.3 70B Wayfarer
- Gemma 3 27B IT
- Gemma 3 12B IT
- Gemma 3 4B IT
- Qwen25 VL 72b
- Holo3-35B-A3B
- Holo3-35B-A3B Thinking
- Cogito v1 Preview Qwen 32B
- Llama-xLAM-2 70B fc-r
- Mistral Small 31 24b Instruct
- Mistral Small 3.2 24b Instruct
- Nvidia Nemotron Super 49B
- Shisa V2 Llama 3.3 70B
- Shisa V2.1 Llama 3.3 70B
- GLM 4 9B 0414
- GLM 4 32B 0414
- Qwen3.5 27B Blossom V6.4 Derestricted
- Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted
- Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted Lite
- Gemma 4 31B Agares v1
- Gemma 4 31B Animus V14.1
- Gemma 4 31B AssGuard
- Gemma 4 31B Dark Gemistry
- Gemma 4 31B Gembrain Uncensored Heretic
- Gemma 4 31B Gembrain X Core
- Gemma 4 31B Isometry RP
- Gemma 4 31B Mero Artemis v0.3.1
- Gemma 4 31B Novelist (ArliAI)
- Gemma 4 31B SDFT Heretic RP
- Gemma 4 31B StyleTune
- Qwen3.5 27B BlueStar v3 Derestricted
- Qwen3.5 27B Queen Derestricted
- Gemma 4 31B Claude 4.6 Opus Reasoning Distilled
- Gemma 4 31B Cognitive Unshackled
- Gemma 4 31B DarkIdol (ArliAI)
- Gemma 4 31B Fabled (ArliAI)
- Gemma 4 31B Garnet V2
- Gemma 4 31B K1 v5
- Gemma 4 31B MeroMero
- Gemma 4 31B Queen
- GLM 4.6 Derestricted v5
- Venice Uncensored
- Gemma 4 26B A4B
- Gemma 4 26B A4B Thinking
- Tencent Hy3
- Qwen3 Coder Next
- Ling 3.0 Flash
- Ling 3.0 Flash Thinking
- Gemma 4 31B
- Gemma 4 31B Thinking
- Nvidia Nemotron 3 Nano 30B
- Nvidia Nemotron 3 Nano Omni
- Nvidia Nemotron 3 Super 120B
- Nvidia Nemotron 3 Super 120B Thinking
- Nex N2 Pro
- Manta Mini 1.0
- MiMo V2.5 Pro (Crof)
- MiMo V2.5 Pro Thinking (Crof)
- Mistral Code Agent Latest
- Synth 2.5 Flash Preview
- Synth 2.5 Pro Preview
Kilo Pass371 recorded models / aliases$19 / $49 / $199 / month
Gateway catalog includes free and paid variants. Annual credit bonus is modeled only where provider rates were captured; catalog presence alone does not establish eligibility.
Anthropic: Claude Sonnet 53 scenarios
- Tier / conditions
- Starter annual billing
- Processed / 30 days (output)
- 91.59M (0.43M output)
- Subscription $ / 1M output + input
- $44.43
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 36.8% less
- Tier / conditions
- Pro annual billing
- Processed / 30 days (output)
- 236.21M (1.10M output)
- Subscription $ / 1M output + input
- $44.43
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 36.8% less
- Tier / conditions
- Expert annual billing
- Processed / 30 days (output)
- 959.32M (4.48M output)
- Subscription $ / 1M output + input
- $44.43
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 36.8% less
Z.ai: GLM 5.3$45.95 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $45.95BaseTen, FP4
- Subscription saving
- Not established
MoonshotAI: Kimi K3$96.95 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $96.95Sail Research, FP4
- Subscription saving
- Not established
MiniMax: MiniMax M33 scenarios
- Tier / conditions
- Starter annual billing
- Processed / 30 days (output)
- 387.20M (1.81M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $15.50DeepInfra, FP8
- Subscription saving
- 32.2% less
- Tier / conditions
- Pro annual billing
- Processed / 30 days (output)
- 998.57M (4.66M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $15.50DeepInfra, FP8
- Subscription saving
- 32.2% less
- Tier / conditions
- Expert annual billing
- Processed / 30 days (output)
- 4.06B (18.94M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $15.50DeepInfra, FP8
- Subscription saving
- 32.2% less
Z.ai: GLM 5.3 Flash$3.90 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $3.90Relace, FP4
- Subscription saving
- Not established
DeepSeek: DeepSeek V4 Pro 0813$12.01 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $12.01DeepSeek
- Subscription saving
- Not established
DeepSeek: DeepSeek V4 Flash 0731$2.36 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.36StreamLake, FP8
- Subscription saving
- Not established
Xiaomi: MiMo-V2.5-Pro$21.21 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $21.21DeepInfra, FP8
- Subscription saving
- Not established
Xiaomi: MiMo-V2.5$1.99 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- Not established
Show 362 additional Kilo Pass catalog models / variants
These IDs have no matched allowance or OpenRouter quote in this snapshot. Catalog coverage: Served catalog; consult plan terms for eligibility.
- Auto Frontier
- Auto Balanced
- Auto Efficient
- Auto Free
- StepFun: Step 3.7 Flash (free)
- Poolside: Laguna S 2.1 (free)
- MiniMax: MiniMax M3 (free)
- MiniMax: MiniMax M2.7 (free)
- Dots Studio: Dots3-Note Preview (free)
- Claude Opus 5
- OpenAI: GPT-5.6 Sol
- OpenAI: GPT-5.6 Sol (50% off)
- OpenAI: GPT-6 Astra ($$$$)
- OpenAI: GPT-6 Astra Pro ($$$$)
- inclusionAI: Ling 3.0 Flash Sante (free)
- Qwen: Qwen3.8 Max (0902)
- Meta: Muse Spark 1.3 Contributor
- Meta: Muse Spark 1.3
- Google: Gemini 3.8 Flash
- Anthropic: Claude Fable 5.1 ($$$$)
- Inception: Mercury 2.5 Preview
- IBM: Granite 4.2 8B
- Tencent: Hy4 preview
- inclusionAI: Ling 3.0 Flash Fin
- inclusionAI: Ling 3.0 Flash Fin (free)
- Z.ai: GLM Flash Latest
- Qwen: Qwen3.8 Flash
- Meta: Muse Spark 1.2 Contributor
- DeepSeek: DeepSeek V4 Flash Vision Exp
- Tencent: Hy-MT2-1.8B
- Tencent: Hy-MT2-30B-A3B
- Z.ai: GLM Latest
- Tencent: Hy-MT2-7B
- Qwen: Qwen3.8 27B
- Google: Gemini 3.7 Flash
- ByteDance Seed: Seed 2.1 Turbo
- Qwen: Qwen3.8 2.4T A95B
- ByteDance Seed: Seed-2.0-Code
- SpaceXAI: Grok 4.6
- LiquidAI: LFM2.5-2.6B (free)
- NVIDIA: Nemotron 3.5 Lightning
- NVIDIA: Nemotron 3.5 Lightning (free)
- Sakana: Sakana Namazu
- Upstage: Solar Pro 4
- Meta: Muse Glimmer 30B
- Meta: Muse Spark 1.2
- DeepSeek V4 Flash Latest
- Thinking Machines: Inkling Small
- Thinking Machines: Inkling Small (free)
- Qwen: Qwen3.7 Flash
- inclusionAI: Ling 3.0 Flash
- Poolside: Laguna S 2.1
- Google: Gemini 3.6 Flash
- Google: Gemini 3.5 Flash Lite
- Meituan: LongCat 2.0
- Thinking Machines: Inkling
- Thinking Machines: Inkling (free)
- OpenRouter Auto Router (Beta)
- Meta: Muse Spark 1.1
- Kwaipilot: KAT-Coder-Pro V2.5
- OpenAI: GPT-5.6 Luna Pro
- OpenAI: GPT-5.6 Luna
- OpenAI: GPT-5.6 Terra Pro
- OpenAI: GPT-5.6 Terra
- OpenAI: GPT-5.6 Sol Pro
- SpaceXAI: Grok 4.5
- xAI: Grok Latest
- AionLabs: Aion-3.0-Mini
- AionLabs: Aion-3.0
- Tencent: Hy3
- Poolside: Laguna XS 2.1
- Poolside: Laguna XS 2.1 (free)
- Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
- Nex AGI: Nex-N2-Mini (retires Sep 8)
- Sakana: Fugu Ultra
- Google: Nano Banana 2 (Gemini 3.1 Flash Image)
- Google: Nano Banana Pro (Gemini 3 Pro Image)
- Cohere: North Mini Code (free)
- Z.ai: GLM 5.2
- OpenRouter: Fusion
- MoonshotAI: Kimi K2.7 Code
- Anthropic: Claude Fable Latest ($$$$)
- Anthropic: Claude Fable 5 ($$$$)
- Nex AGI: Nex-N2-Pro (retires Sep 8)
- NVIDIA: Nemotron 3.5 Content Safety
- NVIDIA: Nemotron 3.5 Content Safety (free)
- NVIDIA: Nemotron 3 Ultra
- NVIDIA: Nemotron 3 Ultra (free)
- Qwen: Qwen3.7 Plus (20% off)
- StepFun: Step 3.7 Flash
- Anthropic: Claude Opus 4.8
- Qwen: Qwen3.7 Max (50% off)
- SpaceXAI: Grok Build 0.1
- Google: Gemini 3.5 Flash
- Perceptron: Perceptron Mk1
- Google: Gemini 3.1 Flash Lite
- OpenAI: GPT Chat Latest
- SpaceXAI: Grok 4.3
- Mistral: Mistral Medium 3.5
- NVIDIA: Nemotron 3 Nano Omni (free)
- Anthropic Claude Haiku Latest
- OpenAI GPT Mini Latest
- Google Gemini Pro Latest
- MoonshotAI Kimi Latest
- Google Gemini Flash Latest
- Anthropic Claude Sonnet Latest
- OpenAI GPT Latest
- Qwen: Qwen3.5 Plus 2026-04-20
- Qwen: Qwen3.6 Flash
- Qwen: Qwen3.6 35B A3B
- Qwen: Qwen3.6 Max Preview
- Qwen: Qwen3.6 27B
- OpenAI: GPT-5.5 Pro ($$$$)
- OpenAI: GPT-5.5
- DeepSeek: DeepSeek V4 Pro 0423
- DeepSeek: DeepSeek V4 Flash 0423
- Tencent: Hy3 preview
- OpenAI: GPT-5.4 Image 2
- Anthropic: Claude Opus Latest
- OpenRouter Pareto Code Router
- MoonshotAI: Kimi K2.6
- Anthropic: Claude Opus 4.7
- Z.ai: GLM 5.1
- Google: Gemma 4 26B A4B
- Google: Gemma 4 31B
- Qwen: Qwen3.6 Plus
- Z.ai: GLM 5V Turbo
- Arcee AI: Trinity Large Thinking
- SpaceXAI: Grok 4.20 Multi-Agent
- SpaceXAI: Grok 4.20
- Google: Lyria 3 Pro Preview
- Google: Lyria 3 Clip Preview
- Kwaipilot: KAT-Coder-Pro V2
- Reka Edge
- MiniMax: MiniMax M2.7
- OpenAI: GPT-5.4 Nano
- OpenAI: GPT-5.4 Mini
- Mistral: Mistral Small 4
- Z.ai: GLM 5 Turbo
- NVIDIA: Nemotron 3 Super
- NVIDIA: Nemotron 3 Super (free)
- ByteDance Seed: Seed-2.0-Lite
- Qwen: Qwen3.5-9B
- OpenAI: GPT-5.4 Pro ($$$$)
- OpenAI: GPT-5.4
- Inception: Mercury 2
- Google: Gemini 3.1 Flash Lite Preview
- ByteDance Seed: Seed-2.0-Mini
- Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
- Qwen: Qwen3.5-35B-A3B
- Qwen: Qwen3.5-27B
- Qwen: Qwen3.5-122B-A10B
- Qwen: Qwen3.5-Flash
- Google: Gemini 3.1 Pro Preview Custom Tools
- OpenAI: GPT-5.3-Codex
- AionLabs: Aion-2.0
- Google: Gemini 3.1 Pro Preview
- Anthropic: Claude Sonnet 4.6
- Qwen: Qwen3.5 Plus 2026-02-15
- Qwen: Qwen3.5 397B A17B
- MiniMax: MiniMax M2.5
- Z.ai: GLM 5
- Qwen: Qwen3 Max Thinking
- Anthropic: Claude Opus 4.6
- Qwen: Qwen3 Coder Next
- OpenRouter Free Models Router
- StepFun: Step 3.5 Flash
- MoonshotAI: Kimi K2.5
- Upstage: Solar Pro 3
- MiniMax: MiniMax M2-her
- Writer: Palmyra X5
- OpenAI: GPT Audio
- OpenAI: GPT Audio Mini
- Z.ai: GLM 4.7 Flash (retires Sep 10)
- OpenAI: GPT-5.2-Codex
- ByteDance Seed: Seed 1.6 Flash
- ByteDance Seed: Seed 1.6
- MiniMax: MiniMax M2.1
- Z.ai: GLM 4.7
- Google: Gemini 3 Flash Preview
- NVIDIA: Nemotron 3 Nano 30B A3B
- OpenAI: GPT-5.2 Chat
- OpenAI: GPT-5.2 Pro ($$$$)
- OpenAI: GPT-5.2
- Mistral: Devstral 2 2512
- Relace: Relace Search
- Z.ai: GLM 4.6V
- OpenRouter Body Builder (beta)
- OpenAI: GPT-5.1-Codex-Max
- Amazon: Nova 2 Lite
- Mistral: Ministral 3 14B 2512
- Mistral: Ministral 3 8B 2512
- Mistral: Ministral 3 3B 2512
- Mistral: Mistral Large 3 2512
- DeepSeek: DeepSeek V3.2
- Anthropic: Claude Opus 4.5
- Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
- OpenAI: GPT-5.1
- OpenAI: GPT-5.1-Codex
- OpenAI: GPT-5.1-Codex-Mini
- MoonshotAI: Kimi K2 Thinking
- Amazon: Nova Premier 1.0
- Perplexity: Sonar Pro Search
- Mistral: Voxtral Small 24B 2507
- OpenAI: gpt-oss-safeguard-20b
- MiniMax: MiniMax M2
- Qwen: Qwen3 VL 32B Instruct
- IBM: Granite 4.0 Micro
- OpenAI: GPT-5 Image Mini
- Anthropic: Claude Haiku 4.5
- Qwen: Qwen3 VL 8B Thinking
- Qwen: Qwen3 VL 8B Instruct
- OpenAI: GPT-5 Image ($$$$)
- Google: Nano Banana (Gemini 2.5 Flash Image)
- Qwen: Qwen3 VL 30B A3B Thinking
- Qwen: Qwen3 VL 30B A3B Instruct
- OpenAI: GPT-5 Pro ($$$$)
- Z.ai: GLM 4.6
- Anthropic: Claude Sonnet 4.5
- DeepSeek: DeepSeek V3.2 Exp
- TheDrummer: Cydonia 24B V4.1
- Relace: Relace Apply 3
- Qwen: Qwen3 VL 235B A22B Thinking
- Qwen: Qwen3 VL 235B A22B Instruct
- Qwen: Qwen3 Max
- Qwen: Qwen3 Coder Plus
- DeepSeek: DeepSeek V3.1 Terminus
- Qwen: Qwen3 Coder Flash
- Qwen: Qwen3 Next 80B A3B Thinking
- Qwen: Qwen3 Next 80B A3B Instruct
- Qwen: Qwen Plus 0728
- MoonshotAI: Kimi K2 0905
- Qwen: Qwen3 30B A3B Thinking 2507
- Nous: Hermes 4 70B
- Nous: Hermes 4 405B
- DeepSeek: DeepSeek V3.1
- Mistral: Mistral Medium 3.1
- Z.ai: GLM 4.5V
- OpenAI: GPT-5
- OpenAI: GPT-5 Mini
- OpenAI: GPT-5 Nano
- OpenAI: gpt-oss-120b
- OpenAI: gpt-oss-20b
- Anthropic: Claude Opus 4.1 ($$$$)
- Mistral: Codestral 2508
- Qwen: Qwen3 Coder 30B A3B Instruct
- Qwen: Qwen3 30B A3B Instruct 2507
- Z.ai: GLM 4.5
- Z.ai: GLM 4.5 Air
- Qwen: Qwen3 235B A22B Thinking 2507
- Qwen: Qwen3 Coder 480B A35B
- ByteDance: UI-TARS 7B
- Google: Gemini 2.5 Flash Lite
- Qwen: Qwen3 235B A22B Instruct 2507
- MoonshotAI: Kimi K2 0711
- Venice: Uncensored
- Tencent: Hunyuan A13B Instruct
- Morph: Morph V3 Large
- Morph: Morph V3 Fast
- Baidu: ERNIE 4.5 VL 424B A47B
- Mistral: Mistral Small 3.2 24B
- MiniMax: MiniMax M1
- Google: Gemini 2.5 Flash
- Google: Gemini 2.5 Pro
- OpenAI: o3 Pro ($$$$)
- Google: Gemini 2.5 Pro Preview 06-05
- DeepSeek: R1 0528
- Anthropic: Claude Opus 4 ($$$$)
- Anthropic: Claude Sonnet 4
- Mistral: Mistral Medium 3
- Google: Gemini 2.5 Pro Preview 05-06
- Meta: Llama Guard 4 12B
- Qwen: Qwen3 30B A3B
- Qwen: Qwen3 8B
- Qwen: Qwen3 14B
- Qwen: Qwen3 32B
- Qwen: Qwen3 235B A22B
- OpenAI: o4 Mini High
- OpenAI: o3
- OpenAI: o4 Mini
- OpenAI: GPT-4.1
- OpenAI: GPT-4.1 Mini
- OpenAI: GPT-4.1 Nano
- Meta: Llama 4 Maverick
- Meta: Llama 4 Scout
- DeepSeek: DeepSeek V3 0324
- OpenAI: o1-pro ($$$$)
- Mistral: Mistral Small 3.1 24B
- Google: Gemma 3 4B
- Google: Gemma 3 12B
- Cohere: Command A
- Reka Flash 3
- Google: Gemma 3 27B
- TheDrummer: Skyfall 36B V2
- Perplexity: Sonar Reasoning Pro
- Perplexity: Sonar Pro
- Perplexity: Sonar Deep Research
- Mistral: Saba
- OpenAI: o3 Mini High
- AionLabs: Aion-RP 1.0 (8B)
- Qwen: Qwen2.5 VL 72B Instruct
- Qwen: Qwen-Plus
- OpenAI: o3 Mini
- Mistral: Mistral Small 3
- Perplexity: Sonar
- DeepSeek: R1 Distill Llama 70B
- DeepSeek: R1
- MiniMax: MiniMax-01
- Microsoft: Phi 4
- DeepSeek: DeepSeek V3
- Sao10K: Llama 3.3 Euryale 70B
- OpenAI: o1 ($$$$)
- Cohere: Command R7B (12-2024)
- Meta: Llama 3.3 70B Instruct
- Amazon: Nova Lite 1.0
- Amazon: Nova Micro 1.0
- Amazon: Nova Pro 1.0
- OpenAI: GPT-4o (2024-11-20)
- Mistral Large 2407
- Qwen2.5 Coder 32B Instruct
- TheDrummer: UnslopNemo 12B
- Magnum v4 72B
- Qwen: Qwen2.5 7B Instruct
- Meta: Llama 3.2 1B Instruct
- Meta: Llama 3.2 3B Instruct
- Qwen2.5 72B Instruct
- Cohere: Command R (08-2024)
- Cohere: Command R+ (08-2024)
- Sao10K: Llama 3.1 Euryale 70B v2.2
- Nous: Hermes 3 70B Instruct
- Nous: Hermes 3 405B Instruct
- Sao10K: Llama 3 8B Lunaris
- OpenAI: GPT-4o (2024-08-06)
- Meta: Llama 3.1 70B Instruct
- Meta: Llama 3.1 8B Instruct
- Mistral: Mistral Nemo
- OpenAI: GPT-4o-mini
- OpenAI: GPT-4o-mini (2024-07-18)
- Google: Gemma 2 27B
- OpenAI: GPT-4o
- OpenAI: GPT-4o (2024-05-13)
- Mistral: Mixtral 8x22B Instruct
- WizardLM-2 8x22B
- OpenAI: GPT-4 Turbo ($$$$)
- Anthropic: Claude 3 Haiku
- Mistral Large
- OpenAI: GPT-3.5 Turbo (older v0613)
- OpenAI: GPT-4 Turbo Preview ($$$$)
- OpenRouter Auto Router
- OpenAI: GPT-3.5 Turbo Instruct
- OpenAI: GPT-3.5 Turbo 16k
- Mancer: Weaver (alpha)
- ReMM SLERP 13B
- MythoMax 13B
- OpenAI: GPT-3.5 Turbo
- OpenAI: GPT-4 ($$$$)
- Stealth: Qwen3.6 Plus (50% off)
- Stealth: Claude Opus 4.8 (20% off)
- Stealth: Claude Opus 4.7 (20% off)
- Stealth: Claude Sonnet 4.6 (20% off)
- Stealth: Claude Opus 4.6 (20% off)
- Auto Small
The selected route is the lowest modeled cost in the captured endpoints with status 0, advertised tool support and explicit cache-read pricing. Nonzero-status routes are excluded from selection; the Xiaomi Pro author quote is retained only in the dataset and quote CSV for comparison. MiniMax M3’s lower GMICloud quote is excluded because its endpoint metadata does not advertise tool support. These are dated quotes, not a promise of availability or default routing. The snapshot records provider tags, context limits, quantization, rate overrides and source URLs in the price dataset. Reported discounts are already reflected in endpoint prices and are not applied again. Endpoint API
Headline input prices can pick the wrong provider. MiMo V2.5’s DeepInfra route quotes $0.13/M fresh input, below Xiaomi’s $0.14/M. But its cache-read price is $0.026/M versus Xiaomi’s $0.0028/M. Including their output rates and the purchase fee, this measured workload costs $7.34 through DeepInfra versus $1.99 through Xiaomi. The API quote, provider tag and context limits for that example are retained in the dataset. MiMo provider quotes
Reproducing the measured 96.53% input cache-hit rate is an assumption. OpenRouter uses sticky routing to help preserve caches, but fallback can move requests to a different provider. Verify actual billed cache reads in the coding harness. Pinning a provider or quantization changes the available routes; an FP4 quote is not evidence of equal application quality to an FP8 or author-hosted route. Each request must also fit that endpoint’s context and pricing tier. Caching, provider routing
GLM Flash’s current OpenRouter quotes include promotional pricing; the author’s offer ends September 9 at 16:00 UTC. Standing subscription estimates above exclude temporary quota boosts. DeepSeek comparisons pin OpenRouter’s Flash 0731 and Pro 0813 versions; a subscription’s generic alias may not guarantee the same checkpoint. The author-route DeepSeek quotes are off-peak. Sonnet uses standard service and five-minute cache creation. Recheck these conditions before extending the estimates beyond this snapshot.
At full modeled utilization, OpenCode Go’s MiMo allocation buys about 6.33× the tokens per dollar of the selected OpenRouter route, and Z.ai Max’s off-peak GLM-5.3 allocation buys about 7.73×. The advantage is smaller for Synthetic GLM Flash against current promotional OpenRouter quotes. Chutes Pro’s conditional DeepSeek Flash ceiling is actually more expensive than the selected StreamLake quote. These are different serving configurations, so accepted-work testing still decides whether the cheaper tokens help.
OpenRouter is therefore a useful baseline and overflow source. A subscription only beats it when enough relevant allowance is consumed: Go’s MiMo allocation breaks even at 15.8% utilization, while Synthetic’s Flash example needs about 73.6% against the current selected quote. The OpenRouter comparison CSV includes cash-budget token capacity; the subscription comparison CSV includes savings and break-even utilization for every mapped, quantifiable offer.
Convert the allowance before comparing it#
Z.ai illustrates why a flat subscription cannot be reduced to one universal token price. Credits depend on the selected model, uncached input, cached input, output and time of day. Off-peak model usage consumes half the normal credits.
Applying its credit formula to the example workload gives these fully utilized, off-peak, steady-state 30-day equivalents:
Lite7.7× API value
- Monthly fee
- $18
- GLM-5.3 output + input
- 2.0M + 430M
- Direct API equivalent
- $138
- Value multiple
- 7.7×
Pro10.3× API value
- Monthly fee
- $80
- GLM-5.3 output + input
- 12.1M + 2581M
- Direct API equivalent
- $826
- Value multiple
- 10.3×
Max11.5× API value
- Monthly fee
- $168
- GLM-5.3 output + input
- 28.2M + 6021M
- Direct API equivalent
- $1,928
- Value multiple
- 11.5×
The weekly allowance is prorated for comparison; it is not a guaranteed calendar-month entitlement. Five-hour limits, concurrency, tool consumption and uneven demand can reduce realized value. All-peak operation halves the token allowances. The same subscriptions produce approximately 2.6–4.0× ordinary GLM Flash API value in this scenario, because Flash’s direct API is already much cheaper.
OpenCode Go has a different mechanism. Some models receive six times the subscription fee in API allowance; others receive three or one-and-a-half times. Allocating its entire qualifying allowance to MiMo V2.5 supports approximately 31.7M output plus 6.76B input tokens for $10 in this scenario. Allocating it to MiniMax M3 supports approximately 3.8M output plus 811M input. Those are alternative uses of one pool, still subject to shorter limits. Allowance rules
MiMo’s own plan shows the danger of reading credits as tokens. Each uncached V2.5 input token costs 100 credits; each output token costs 200. Its 4.1B-credit Lite plan therefore represents approximately $5.74 of daytime direct API usage for a $6 fee, or $7.18 with the off-peak credit discount. Large numbers need units. Conversion rules, API prices
Synthetic’s $24 weekly allowance represents about $103 per 30 days at Synthetic’s rates. Always compare the serving provider’s cache prices and model configuration with the original API before calling that a direct-provider discount. Limits, model metadata and prices
NanoGPT illustrates the effect of counting cache hits at full input weight. At the measured ratio, its 60M weekly input-unit allowance yields approximately 258M total processed tokens, including 1.21M output, per 30 days on a 1× model. A 2× model halves those estimates. A request-count plan needs another measurement—tokens per billable request—before it can be converted. The local message count is not assumed to equal provider API calls. NanoGPT quota definitions
Utilization is the other half of the arithmetic. A $10 plan containing $60 of relevant API usage breaks even once it replaces more than $10 of API spending. The unused $50 has no value. More packs can buy parallelism, but idle parallelism is still a bill.
Native coding subscriptions belong in the trial too#
Codex through ChatGPT, Claude Code, Cursor and Copilot can be economical when their native workflows produce useful results. Their included allowances are not interchangeable general API credits, and published limits do not establish a universal cost per token.
- Codex through ChatGPT: $20/$100/$200 options, with model-dependent five-hour and weekly allowances. API billing is separate. Pricing
- Claude Code: Pro $20, Max from $100. Personal native use and product integrations have different access rules. Pricing, usage rules
- Cursor: $20/$60/$200, with distinct usage pools for Cursor models and other models. Pricing documentation
- Copilot: $10/$39/$100 plans advertise $15/$70/$200 total feature credits, including variable flex credits. Plans
- Gemini CLI: the free Google-account tier advertises up to 1,000 requests daily, subject to model availability and shared limits. It is a useful free baseline. Quotas
Temporary offers, kept separate#
As checked September 7, Z.ai’s GLM-5.3-Flash direct API is half price until September 9 at 16:00 UTC. That promotional row is explicitly separated from ordinary pricing above. API offer
Through September 20, Z.ai also advertises zero-quota Flash use through ZCode daily 15:00–01:00 UTC, and doubled quota through other supported agents during that window. OpenCode separately advertises a temporary Flash allowance increase. These offers can improve a trial; the standing subscription calculations exclude them. After expiry, treat this paragraph as a dated observation until it is revised. Z.ai campaign, OpenCode campaign
The next analysis should measure accepted work#
Run 20–50 representative tasks through the same harness, tests and retry budget. Record input, cache hits, billed output, subscription consumption, cash spend, retries, queue time, review time and whether the change was accepted. Include abandoned attempts in the bill.
The metric is total inference and retry cost divided by changes that pass tests and review. Report elapsed time alongside it: a cheap queue can become expensive when it delays everything else. Allocate the full subscription fee across the measured period, including unused quota, rather than claiming the theoretical maximum discount.
That experiment may favor a cheap default with occasional stronger calls. It may favor one reliable model that finishes sooner. This initial research does not establish which outcome the fleet will produce.
Reproduce and update the comparison#
Download the dated price dataset, measured usage profile, model catalogs, calculator, API comparison CSV, plan capacity CSV, Z.ai comparison CSV, OpenRouter quote CSV, and OpenRouter versus subscriptions CSV. The dataset records source URLs and check dates; the artifact’s comparisons and CSVs are generated from it.
Put calculate.mjs and prices.json in the same directory. With Node 22 or newer:
node calculate.mjs --format api > api.csv
node calculate.mjs --format plans > capacity.csv
node calculate.mjs --format zai > zai.csv
node calculate.mjs --format openrouter > openrouter.csv
node calculate.mjs --format openrouter-plans > openrouter-plans.csv
# Sensitivity scenario; defaults above use the measured profile.
node calculate.mjs --ratio 50 --cache-hit 0.95 --cache-write-share 0.2 --format plans
--cache-write-share is the fraction of non-cache-read input spent creating caches. It preserves separate write pricing while --cache-hit controls reads as a fraction of all input. The calculator does not convert a message count into requests.
For OpenRouter, --topup-credits 10 models the fee on a $10 inference-credit purchase; --budget remains a cash budget. The default amortizes a $100 credit purchase. Ratio overrides reprice the captured provider routes; they do not select a new cheapest provider. Use the recorded endpoint API URLs to refresh route selection for a different workload.
Future revisions can add measured task outcomes, new providers, different cache-hit scenarios and throughput observations without moving this artifact. The page-level verification date advances only after the comparison has been checked as a whole; partial refreshes retain individual source dates. Corrections and changed recommendations belong in the changelog, newest first.
Changelog#
2026-09-07 — Published as a data artifact#
- Retired the private Notes draft and published the maintained comparison at
/data/inference-plans/. - Kept the report, Markdown twin, calculator, source data and generated CSVs together under one public artifact path.
- This changes the publication surface only; the analysis, measurement and source-verification dates are unchanged.
2026-09-07 — Display the analysis date#
- Added a prominent analysis date near the top of the artifact, separate from publication, measurement and source-verification dates.
2026-09-07 — Index by subscription and model#
- Reorganized the OpenRouter comparison into all twelve researched coding subscriptions, then their models, with tier and time-window rows underneath.
- Included the saved model catalogs; large unpriced catalogs expand within the relevant subscription instead of disappearing from the comparison.
- Added token capacity beside each calculated model/tier cost and retained explicit unknowns for missing allowances, rates and model multipliers.
- Preserved the existing measured mix and dated prices; this is a coverage and presentation update, not a source refresh.
2026-09-07 — OpenRouter route comparison#
- Added nine model comparisons using coherent provider quotes, measured cache composition and the standard credit-purchase fee.
- Added selected provider, quantization, context and status metadata; separated flagged author quotes and current promotional rates.
- Added OpenRouter cash-budget capacity and subscription savings/break-even CSVs, plus a configurable top-up size.
- Found that provider cache pricing can reverse headline-price rankings, and that the selected DeepSeek Flash route undercuts Chutes Pro’s modeled ceiling.
- Retained the earlier usage measurement and subscription snapshot; this revision refreshes OpenRouter quotes only.
2026-09-07 — Normalize subscription costs#
- Converted the measured mix into effective costs per million processed tokens and per million output tokens including associated input for every quantifiable offer.
- Added a main normalization table, expandable model/time-window calculations and matching CSV fields.
- Added explicit same-model direct API savings and break-even utilization, preserving negative savings and unknown comparisons.
- Kept the existing measurement and dated price snapshot; this revision adds arithmetic rather than claiming a new measurement or source refresh.
2026-09-07 — Measured usage, supported models and capacity#
- Replaced the illustrative 20:1 input/output ratio and 80% cache-hit assumption with a fresh local measurement: 213.2:1 and 96.53% of input from cache.
- Normalized separately reported reasoning into billable output and retained distinct cache-read and cache-write buckets.
- Added model coverage for every plan in the main table, dated catalog snapshots and reproducible token-capacity estimates using serving-provider rates.
- Kept unknown quotas, unmetered service and conditional ceilings explicit; added downloadable usage and capacity data.
- Recalculated API examples. The shortlist remains a set of trial candidates; no accepted-work benchmark has yet been run.
2026-09-07 — Initial research draft#
- Compared twelve subscription offerings with direct API costs and native coding alternatives.
- Added reproducible API and Z.ai calculations, source-check dates, and explicit workload assumptions.
- Recorded automation restrictions, conflicting reseller documentation and expiring promotions.
- Established the cost-per-accepted-change experiment as the next analysis; no paid performance benchmark has been run for this note.