API REFERENCE
Errors & limits
Diagnose failed requests and keep your integration within service limits.
HTTP status codes
Errors use an OpenAI-style error object. Save the X-Request-ID response header in your logs to identify the request.
| Status | What to check |
|---|---|
| 400 | JSON, parameters, model availability, endpoint or required streaming. |
| 401 | Missing, invalid, expired or revoked key. |
| 402 | Insufficient balance or key budget. |
| 403 | Key or account permission. |
| 404 / 405 | Unsupported endpoint or HTTP method. |
| 413 / 415 | Body over 256 KB or Content-Type is not application/json. |
| 429 | Rate limit. Honor Retry-After; gateway errors use 60 seconds. |
| 502 / 503 / 504 | Model service error, unavailability or timeout. |
{
"error": {
"message": "Your Tokely balance or API key budget is insufficient for this request.",
"type": "invalid_request_error",
"param": null,
"code": "insufficient_quota"
}
}Fix your first request
Start with the smallest request from Quickstart. Check the HTTP status and error.code before changing model settings. If your SDK fails before sending HTTP, check the local setup first.
| Symptom | Next step |
|---|---|
| TOKELY_API_KEY is missing | Set the variable in the terminal that runs your script; restart your server after changing environment variables. |
| Module not found / import error | Install openai in the same Python environment or JavaScript project. Run JavaScript examples as .mjs files. |
| invalid_api_key / 401 | Use a Tokely key, copy its full value, and check that it is active and unexpired. |
| insufficient_quota / 402 | Check your account balance and key budget. Lower an unnecessarily high output limit. |
| model_not_available | Copy an enabled ID exactly and check that it supports the endpoint you selected. |
| stream_required | Add stream=true and read the request as a stream. |
| unsupported_parameter | Remove the named field. Tokely supports a defined subset of model API options. |
| No text / finish_reason=length | Increase the output budget; reasoning models may use it for reasoning before visible text. |
| HTML or endpoint_not_supported | Use https://api.tokely.me/v1 as the SDK base URL. Append /chat/completions only for a direct HTTP request. |
Service limits
Model-specific context, output and feature limits also apply.
| Limit | Value |
|---|---|
| Request rate | 60 requests/minute per account, across keys. |
| JSON body | 256 KB maximum. |
| Messages / input items | 1–128 per request. |
| Function tools | 32 maximum. |
| Completions | One per request (n=1). |
| Default output | 256 tokens, including reasoning. |
| Maximum output | Model limit capped at 65,536; 4,096 if unspecified. |
| Gateway timeout | 120 seconds. |
Retry deliberately
Fix authentication, balance and validation errors before retrying. For transient errors, use bounded backoff and honor Retry-After. Avoid immediately replaying after a timeout: generation may already have begun.
Tokely does not automatically retry inference. SDK retry settings are separate. Every new generation attempt can consume usage.