GUIDES
Understand your usage.
Model prices, account balance and output reservations.
Prices per model
Model pages publish input, output and, where available, cached input prices per million tokens in USD. Check the model’s current rates before calling it. Reference comparisons are labeled separately.
Charges use reported token usage and applicable rates. Reasoning can consume the output budget. Cache pricing applies when the model reports eligible cached input.
Balance and reservations
View your available USD balance in the dashboard. Requests use this balance across your API keys.
Before generation, the service reserves estimated input plus maximum output. It then settles actual usage and releases the unused reservation. Your balance and key budget must cover the reservation; a high output limit can cause a 402 even if the eventual answer would be short.
Inspect usage
Read usage in a completed response. Cache and reasoning fields depend on the model and endpoint. Monetary rounding can affect very small requests.
X-Tokely-Pricing-Version identifies the active pricing version on successful generation responses. It does not replace published model rates.