THE COST OF BUILDING WITH AI

Cheapest LLM API: compare the actual workload

Compare published LLM input, output and cache-read prices, with a consistent 1M-input / 250K-output example and exact-model OpenRouter references.

THE SHORT ANSWER

The cheapest LLM API depends on the model and token mix. The table below orders enabled Tokely text models by an uncached workload of one million input tokens and 250,000 output tokens. Exact-model OpenRouter rates appear where a matching public offer is available.

Input-heavy and output-heavy apps have different winners

For classification and extraction, input can dominate. For code generation and long answers, output can dominate. A low input rate alone is not a useful ranking: price both parts, then use the calculator to change the mix.

Same model first. Different model second.

An exact-model comparison helps evaluate a gateway switch. Comparing different models is a separate quality decision: test representative tasks and constraints before replacing a model. The table orders prices; it does not claim equal intelligence or interchangeability.

Include the costs outside the token table

Some gateways charge a platform or credit-purchase fee. OpenRouter's published Standard plan lists a 5.5% platform fee; its other plans have different conditions. Our token comparisons exclude that fee, minimum fees, tax, special units and negotiated plans, so the quoted reference is not a complete checkout bill.

Compare a matched text workload

USD per 1M tokens. Example: 1M uncached input + 250K output. Tokely rates dated 2026-10-09; OpenRouter references checked 2026-10-09. Platform/checkout fees, taxes, special units and context tiers excluded. Different models are not quality equivalents. Methodology.
ModelTokely input / outputTokely exampleOpenRouter exampleDifference vs OpenRouter
GPT-5.4 MiniOpenAI$0.07425 / $0.4455Cache read: $0.007425$0.185625$1.875$0.75 / $4.5 per 1M90.1% lower
Gemini 3.1 Flash LiteGoogle$0.050552 / $0.303312Cache read: $0.005056$0.12638$0.625$0.25 / $1.5 per 1M79.8% lower
Gemini 2.5 FlashGoogle$0.060663 / $0.505924Cache read: $0.006067$0.187144$0.925$0.3 / $2.5 per 1M79.8% lower
GPT-6 LunaOpenAI$0.09 / $0.45Cache read: $0.009$0.2025$0.225$0.1 / $0.5 per 1M10.0% lower
DeepSeek V4.1 FlashDeepSeek$0.135 / $0.54Cache read: $0.0027$0.27$0.6$0.3 / $1.2 per 1M55.0% lower
Gemini 3.7 FlashGoogle$0.151656 / $0.758279Cache read: $0.015166$0.341226$1.6875$0.75 / $3.75 per 1M79.8% lower
Gemini 3.8 FlashGoogle$0.151656 / $0.758279Cache read: $0.015166$0.341226$1.6875$0.75 / $3.75 per 1M79.8% lower
GPT-5 CodexOpenAI$0.126397 / $1.011175Cache read: $0.01264$0.379191Not matched—
GPT-5.6 LunaOpenAI$0.18 / $1.08Cache read: $0.018$0.45$0.5$0.2 / $1.2 per 1M10.0% lower
Gemini 3.5 Flash LiteGoogle$0.218386 / $1.819882Cache read: $0.021839$0.673357$0.925$0.3 / $2.5 per 1M27.2% lower
Grok 4.3Grok$0.505519 / $1.011038Cache read: $0.18$0.758279$1.875$1.25 / $2.5 per 1M59.6% lower
Gemini 3.5 FlashGoogle$0.303312 / $1.819868Cache read: $0.030332$0.758279$3.75$1.5 / $9 per 1M79.8% lower
GPT-6 SolOpenAI$0.505588 / $1.516763Cache read: $0.040447$0.884779$4.5$2 / $10 per 1M80.3% lower
GPT-6.1 SolOpenAI$0.505588 / $1.516763Cache read: $0.020224$0.884779$4.5$2 / $10 per 1M80.3% lower
GPT-5.6 TerraOpenAI$0.505588 / $1.820115Cache read: $0.040447$0.960617$5$2 / $12 per 1M80.8% lower
Claude Sonnet 5Anthropic$0.48532 / $2.4266Cache read: $0.048532$1.09197$4.5$2 / $10 per 1M75.7% lower
Gemini 2.5 ProGoogle$0.505519 / $3.033113Cache read: $0.050552$1.263797$3.75$1.25 / $10 per 1M66.3% lower
Grok 4.5Grok$0.80883 / $2.42649Cache read: $0.202208$1.415453$3.5$2 / $6 per 1M59.6% lower
Grok 4.6Grok$0.80883 / $2.42649Cache read: $0.202208$1.415453$3.5$2 / $6 per 1M59.6% lower
Grok 4.7Grok$0.80883 / $2.42649Cache read: $0.202208$1.415453$3.5$2 / $6 per 1M59.6% lower
Claude Sonnet 4.6Anthropic$0.72798 / $3.6399Cache read: $0.072798$1.637955$6.75$3 / $15 per 1M75.7% lower
GPT-5.2 CodexOpenAI$0.641667 / $5.133334Cache read: $0.064167$1.925001$5.25$1.75 / $14 per 1M63.3% lower
Gemini 3 ProGoogle$0.75 / $5.25Cache read: $0.75$2.0625Not matched—
GPT-5.5OpenAI$1.011175 / $4.550288Cache read: $0.101118$2.148747$12.5$5 / $30 per 1M82.8% lower
Claude Sonnet 5.5Anthropic$0.970585 / $4.852925Cache read: $0.097059$2.183816$4.5$2 / $10 per 1M51.5% lower
Claude Opus 5.5Anthropic$0.97064 / $4.8532Cache read: $0.048532$2.18394$9$4 / $20 per 1M75.7% lower
Gemini 3.1 ProGoogle$0.833334 / $5.833334Cache read: $0.18$2.291668Not matched—
GPT-5.6 SolOpenAI$1.263969 / $4.550288Cache read: $0.101118$2.401541$4.5$2 / $10 per 1M46.6% lower
GPT-5.1 CodexOpenAI$1.008334 / $8.066667Cache read: $0.100834$3.025001$3.75$1.25 / $10 per 1M19.3% lower
Kimi K3Moonshot$0.45 / $10.8Cache read: $0.27$3.15$3.5$0.5 / $12 per 1M10.0% lower
GPT-5.3 CodexOpenAI$1.05875 / $8.47Cache read: $0.105875$3.17625$5.25$1.75 / $14 per 1M39.5% lower
GPT-6 AstraOpenAI$2.527938 / $7.583813Cache read: $0.202235$4.423891$22.5$10 / $50 per 1M80.3% lower
GPT-5.2OpenAI$1.575 / $12.6Cache read: $0.1575$4.725$5.25$1.75 / $14 per 1M10.0% lower
Claude Opus 4.6Anthropic$2.426463 / $12.132313Cache read: $0.242647$5.459541$11.25$5 / $25 per 1M51.5% lower
Claude Opus 4.8Anthropic$2.426463 / $12.132313Cache read: $0.242647$5.459541$11.25$5 / $25 per 1M51.5% lower
Claude Opus 5Anthropic$2.426463 / $12.132313Cache read: $0.242647$5.459541$11.25$5 / $25 per 1M51.5% lower
GPT-5.4OpenAI$2.25 / $13.5Cache read: $0.225$5.625$6.25$2.5 / $15 per 1M10.0% lower
Claude Opus 4.7Anthropic$3.033079 / $12.132313Cache read: $0.242647$6.066157$11.25$5 / $25 per 1M46.1% lower
Claude Fable 5Anthropic$4.852925 / $24.264625Cache read: $0.485293$10.919081$22.5$10 / $50 per 1M51.5% lower

Change the token mix in the calculator →

Questions before you switch

Are the cheapest models always the best value?

No. Compare price per accepted result on your own tasks, including retries, latency and limits. The rate table contains no quality or performance ranking.

Do these comparisons include prompt caching?

The example table uses uncached input. The calculator allows cached tokens where both the published rate and request usage support them. Missing cache-read prices are shown as unavailable.

Keep exploring

OpenRouter migration checklist · Billing documentation

START WITH YOUR OWN WORKLOAD

Know the price.
Then make the call.

Check the model, compare the rate, and connect with a Tokely key.

Read the quickstart ↗Explore all models →