Model Catalog
AIOrouter is an All-In-One AI Router — providing a single OpenAI-compatible API endpoint that routes to Western models (Google Gemini, Anthropic Claude), Chinese models (DeepSeek, Alibaba Qwen, Moonshot Kimi, Zhipu GLM), and more — all with built-in privacy protection (bidirectional PII pseudonymization, technical secret redaction, AI Firewall) and Canada-resident infrastructure.
LLM Public Token Pricing
Prices below are public provider rates in USD per 1 million tokens before CAD conversion, taxes, and any account-specific retail presentation. The Dashboard and API model response are the customer-facing source for your exact billable rate.
Pricing basis. Model pricing is based on the rates published on this page. AIOrouter reserves the right to adjust published rates from time to time; adjustments take effect once announced (see the Terms of Service).
| Model | Input (USD/1M tokens) | Output (USD/1M tokens) | Cache pricing |
|---|---|---|---|
| claude-fable-5 | $10.00 | $50.00 | Cache write: $12.50; Cache read: $1.00 |
| claude-haiku-4.5 | $1.00 | $5.00 | Cache write: $1.25; Cache read: $0.100 |
| claude-opus-5 | $5.00 | $25.00 | Cache write: $6.25; Cache read: $0.500 |
| claude-sonnet-5 | $2.00 | $10.00 | Cache write: $2.50; Cache read: $0.200 |
| deepseek-v4-flash (version 0731) | $0.138 | $0.275 | Implicit cache: $0.028 |
| deepseek-v4-pro → deepseek-v4-pro-0813 | $0.636 | $1.91 | Cache read: $0.064; Implicit cache: $0.064; explicit creation $0.795 |
| gemini-2.5-flash | $0.300 | $2.50 | Implicit cache: $0.030 |
| gemini-2.5-pro | $1.25 up to 200K / $2.50 above | $10.00 | Implicit cache: $0.125 |
| glm-5.1 | $0.825 | $3.30 | Cache read: $0.083; Implicit cache: $0.165; explicit creation $1.03 |
| glm-5.2 | $1.40 | $4.40 | Cache read: $0.140; Implicit cache: $0.260; explicit creation $1.75 |
| grok-4.6 | $2.00 | $6.00 | Implicit cache: $0.500 |
| kimi-k2.6 | $0.950 | $4.00 | Implicit cache: $0.160 |
| kimi-k2.7-code | $0.950 | $4.00 | Implicit cache: $0.190 |
| kimi-k3 | $3.00 | $15.00 | Implicit cache: $0.300 |
| qwen3.6-flash | $0.165 | $0.990 | Cache read: $0.017; Implicit cache: $0.033; explicit creation $0.206 |
| qwen3.6-plus → qwen3.7-max-2026-05-20 | $0.276 | $1.65 | Cache read: $0.028; Implicit cache: $0.055; explicit creation $0.345 |
| qwen3.7-max | $0.825 | $2.48 | Cache read: $0.083; Implicit cache: $0.165; explicit creation $1.03 |
| qwen3.7-plus | $0.221 | $0.881 | Cache read: $0.022; Implicit cache: $0.045; explicit creation $0.275 |
| qwen3.8-max | $1.65 | $4.95 | Cache read: $0.137; Implicit cache: $0.206; explicit creation $2.06 |
Time-of-day pricing (peak/off-peak): requests during the off-peak hours (22:00–08:00 (Asia/Shanghai, UTC+8)) are billed at the prices in the table above; during peak hours the gateway automatically applies the 2× peak rate below (
X-Billing-Modelresponse header shows the resolved variant).
| Model | Off-peak (USD/1M in/out) | Peak (USD/1M in/out) |
|---|---|---|
| deepseek-v4-flash (version 0731) | $0.220 / $0.660 | $0.440 / $1.32 |
| deepseek-v4-pro → deepseek-v4-pro-0813 | $0.636 / $1.91 | $1.27 / $3.82 |
Current Exchange Rate
| Date | Source | USD → CAD Rate | FX Buffer Applied |
|---|---|---|---|
| 2026-08-18 | Bank of Canada — Bank of Canada VALET API | 1.4142 | Yes |
Token costs in CAD are computed as: USD_price × CAD_USD_rate. The rate above is refreshed daily
at 09:00 EST from the Bank of Canada VALET API with the FX buffer (5% execution fee covering currency
conversion and payment processing) applied per pricing policy.
Current Exchange Rate
| Date | Source | USD → CAD Rate | 2% Buffer Applied |
|---|---|---|---|
| 2026-08-20 | Bank of Canada — Bank of Canada VALET API | 1.4515 | Yes |
Token costs in CAD are computed as: USD_price × CAD_USD_rate. The rate above is refreshed daily
at 09:00 EST from the Bank of Canada VALET API with a 5% buffer applied per pricing policy.
Official Pricing Sources
- Anthropic Claude pricing: Anthropic API pricing
- Google Gemini pricing: Google Gemini API pricing documentation
- DeepSeek pricing: DeepSeek API pricing documentation
- Alibaba Qwen and GLM pricing: Alibaba Model Studio Console
- Kimi pricing: Moonshot public API documentation
Data freshness: This BETA catalog was updated on 2026-08-20. Token rates are public provider rates.