THE COST OF BUILDING WITH AI
Cheapest LLM API: compare the actual workload
Compare published LLM input, output and cache-read prices, with a consistent 1M-input / 250K-output example and exact-model OpenRouter references.
THE SHORT ANSWER
The cheapest LLM API depends on the model and token mix. The table below orders enabled Tokely text models by an uncached workload of one million input tokens and 250,000 output tokens. Exact-model OpenRouter rates appear where a matching public offer is available.
Input-heavy and output-heavy apps have different winners
For classification and extraction, input can dominate. For code generation and long answers, output can dominate. A low input rate alone is not a useful ranking: price both parts, then use the calculator to change the mix.
Same model first. Different model second.
An exact-model comparison helps evaluate a gateway switch. Comparing different models is a separate quality decision: test representative tasks and constraints before replacing a model. The table orders prices; it does not claim equal intelligence or interchangeability.
Include the costs outside the token table
Some gateways charge a platform or credit-purchase fee. OpenRouter's published Standard plan lists a 5.5% platform fee; its other plans have different conditions. Our token comparisons exclude that fee, minimum fees, tax, special units and negotiated plans, so the quoted reference is not a complete checkout bill.
Compare a matched text workload
| Model | Tokely input / output | Tokely example | OpenRouter example | Difference vs OpenRouter |
|---|---|---|---|---|
| GPT-5.4 MiniOpenAI | $0.07425 / $0.4455Cache read: $0.007425 | $0.185625 | $1.875$0.75 / $4.5 per 1M | 90.1% lower |
| Gemini 3.1 Flash LiteGoogle | $0.050552 / $0.303312Cache read: $0.005056 | $0.12638 | $0.625$0.25 / $1.5 per 1M | 79.8% lower |
| Gemini 2.5 FlashGoogle | $0.060663 / $0.505924Cache read: $0.006067 | $0.187144 | $0.925$0.3 / $2.5 per 1M | 79.8% lower |
| GPT-6 LunaOpenAI | $0.09 / $0.45Cache read: $0.009 | $0.2025 | $0.225$0.1 / $0.5 per 1M | 10.0% lower |
| DeepSeek V4.1 FlashDeepSeek | $0.135 / $0.54Cache read: $0.0027 | $0.27 | $0.6$0.3 / $1.2 per 1M | 55.0% lower |
| Gemini 3.7 FlashGoogle | $0.151656 / $0.758279Cache read: $0.015166 | $0.341226 | $1.6875$0.75 / $3.75 per 1M | 79.8% lower |
| Gemini 3.8 FlashGoogle | $0.151656 / $0.758279Cache read: $0.015166 | $0.341226 | $1.6875$0.75 / $3.75 per 1M | 79.8% lower |
| GPT-5 CodexOpenAI | $0.126397 / $1.011175Cache read: $0.01264 | $0.379191 | Not matched | — |
| GPT-5.6 LunaOpenAI | $0.18 / $1.08Cache read: $0.018 | $0.45 | $0.5$0.2 / $1.2 per 1M | 10.0% lower |
| Gemini 3.5 Flash LiteGoogle | $0.218386 / $1.819882Cache read: $0.021839 | $0.673357 | $0.925$0.3 / $2.5 per 1M | 27.2% lower |
| Grok 4.3Grok | $0.505519 / $1.011038Cache read: $0.18 | $0.758279 | $1.875$1.25 / $2.5 per 1M | 59.6% lower |
| Gemini 3.5 FlashGoogle | $0.303312 / $1.819868Cache read: $0.030332 | $0.758279 | $3.75$1.5 / $9 per 1M | 79.8% lower |
| GPT-6 SolOpenAI | $0.505588 / $1.516763Cache read: $0.040447 | $0.884779 | $4.5$2 / $10 per 1M | 80.3% lower |
| GPT-6.1 SolOpenAI | $0.505588 / $1.516763Cache read: $0.020224 | $0.884779 | $4.5$2 / $10 per 1M | 80.3% lower |
| GPT-5.6 TerraOpenAI | $0.505588 / $1.820115Cache read: $0.040447 | $0.960617 | $5$2 / $12 per 1M | 80.8% lower |
| Claude Sonnet 5Anthropic | $0.48532 / $2.4266Cache read: $0.048532 | $1.09197 | $4.5$2 / $10 per 1M | 75.7% lower |
| Gemini 2.5 ProGoogle | $0.505519 / $3.033113Cache read: $0.050552 | $1.263797 | $3.75$1.25 / $10 per 1M | 66.3% lower |
| Grok 4.5Grok | $0.80883 / $2.42649Cache read: $0.202208 | $1.415453 | $3.5$2 / $6 per 1M | 59.6% lower |
| Grok 4.6Grok | $0.80883 / $2.42649Cache read: $0.202208 | $1.415453 | $3.5$2 / $6 per 1M | 59.6% lower |
| Grok 4.7Grok | $0.80883 / $2.42649Cache read: $0.202208 | $1.415453 | $3.5$2 / $6 per 1M | 59.6% lower |
| Claude Sonnet 4.6Anthropic | $0.72798 / $3.6399Cache read: $0.072798 | $1.637955 | $6.75$3 / $15 per 1M | 75.7% lower |
| GPT-5.2 CodexOpenAI | $0.641667 / $5.133334Cache read: $0.064167 | $1.925001 | $5.25$1.75 / $14 per 1M | 63.3% lower |
| Gemini 3 ProGoogle | $0.75 / $5.25Cache read: $0.75 | $2.0625 | Not matched | — |
| GPT-5.5OpenAI | $1.011175 / $4.550288Cache read: $0.101118 | $2.148747 | $12.5$5 / $30 per 1M | 82.8% lower |
| Claude Sonnet 5.5Anthropic | $0.970585 / $4.852925Cache read: $0.097059 | $2.183816 | $4.5$2 / $10 per 1M | 51.5% lower |
| Claude Opus 5.5Anthropic | $0.97064 / $4.8532Cache read: $0.048532 | $2.18394 | $9$4 / $20 per 1M | 75.7% lower |
| Gemini 3.1 ProGoogle | $0.833334 / $5.833334Cache read: $0.18 | $2.291668 | Not matched | — |
| GPT-5.6 SolOpenAI | $1.263969 / $4.550288Cache read: $0.101118 | $2.401541 | $4.5$2 / $10 per 1M | 46.6% lower |
| GPT-5.1 CodexOpenAI | $1.008334 / $8.066667Cache read: $0.100834 | $3.025001 | $3.75$1.25 / $10 per 1M | 19.3% lower |
| Kimi K3Moonshot | $0.45 / $10.8Cache read: $0.27 | $3.15 | $3.5$0.5 / $12 per 1M | 10.0% lower |
| GPT-5.3 CodexOpenAI | $1.05875 / $8.47Cache read: $0.105875 | $3.17625 | $5.25$1.75 / $14 per 1M | 39.5% lower |
| GPT-6 AstraOpenAI | $2.527938 / $7.583813Cache read: $0.202235 | $4.423891 | $22.5$10 / $50 per 1M | 80.3% lower |
| GPT-5.2OpenAI | $1.575 / $12.6Cache read: $0.1575 | $4.725 | $5.25$1.75 / $14 per 1M | 10.0% lower |
| Claude Opus 4.6Anthropic | $2.426463 / $12.132313Cache read: $0.242647 | $5.459541 | $11.25$5 / $25 per 1M | 51.5% lower |
| Claude Opus 4.8Anthropic | $2.426463 / $12.132313Cache read: $0.242647 | $5.459541 | $11.25$5 / $25 per 1M | 51.5% lower |
| Claude Opus 5Anthropic | $2.426463 / $12.132313Cache read: $0.242647 | $5.459541 | $11.25$5 / $25 per 1M | 51.5% lower |
| GPT-5.4OpenAI | $2.25 / $13.5Cache read: $0.225 | $5.625 | $6.25$2.5 / $15 per 1M | 10.0% lower |
| Claude Opus 4.7Anthropic | $3.033079 / $12.132313Cache read: $0.242647 | $6.066157 | $11.25$5 / $25 per 1M | 46.1% lower |
| Claude Fable 5Anthropic | $4.852925 / $24.264625Cache read: $0.485293 | $10.919081 | $22.5$10 / $50 per 1M | 51.5% lower |
Questions before you switch
Are the cheapest models always the best value?
No. Compare price per accepted result on your own tasks, including retries, latency and limits. The rate table contains no quality or performance ranking.
Do these comparisons include prompt caching?
The example table uses uncached input. The calculator allows cached tokens where both the published rate and request usage support them. Missing cache-read prices are shown as unavailable.
Keep exploring
AI API pricing. Know what you pay. ↗
Published Tokely rates for text, images, video and audio. Pay from one USD balance, compare exact models and estimate your workload before you switch.
A cost-focused OpenRouter alternative ↗
Compare exact-model rates, supported endpoints and migration requirements before moving to Tokely. One balance for supported text, image, video and audio models.
AI API price index ↗
An exportable first-party dataset of enabled Tokely models, exact-model OpenRouter token references and supported media variants. Sources, units and dates are included.
START WITH YOUR OWN WORKLOAD
Know the price.
Then make the call.
Check the model, compare the rate, and connect with a Tokely key.