Docs/Errors & limits

API REFERENCE

Errors & limits

Diagnose failed requests and keep your integration within service limits.

HTTP status codes

Errors use an OpenAI-style error object. Save the X-Request-ID response header in your logs to identify the request.

StatusWhat to check
400JSON, parameters, model availability, endpoint or required streaming.
401Missing, invalid, expired or revoked key.
402Insufficient balance or key budget.
403Key or account permission.
404 / 405Unsupported endpoint or HTTP method.
413 / 415Body over 256 KB or Content-Type is not application/json.
429Rate limit. Honor Retry-After; gateway errors use 60 seconds.
502 / 503 / 504Model service error, unavailability or timeout.
{
  "error": {
    "message": "Your Tokely balance or API key budget is insufficient for this request.",
    "type": "invalid_request_error",
    "param": null,
    "code": "insufficient_quota"
  }
}

Fix your first request

Start with the smallest request from Quickstart. Check the HTTP status and error.code before changing model settings. If your SDK fails before sending HTTP, check the local setup first.

SymptomNext step
TOKELY_API_KEY is missingSet the variable in the terminal that runs your script; restart your server after changing environment variables.
Module not found / import errorInstall openai in the same Python environment or JavaScript project. Run JavaScript examples as .mjs files.
invalid_api_key / 401Use a Tokely key, copy its full value, and check that it is active and unexpired.
insufficient_quota / 402Check your account balance and key budget. Lower an unnecessarily high output limit.
model_not_availableCopy an enabled ID exactly and check that it supports the endpoint you selected.
stream_requiredAdd stream=true and read the request as a stream.
unsupported_parameterRemove the named field. Tokely supports a defined subset of model API options.
No text / finish_reason=lengthIncrease the output budget; reasoning models may use it for reasoning before visible text.
HTML or endpoint_not_supportedUse https://api.tokely.me/v1 as the SDK base URL. Append /chat/completions only for a direct HTTP request.
Back to Quickstart

Service limits

Model-specific context, output and feature limits also apply.

LimitValue
Request rate60 requests/minute per account, across keys.
JSON body256 KB maximum.
Messages / input items1–128 per request.
Function tools32 maximum.
CompletionsOne per request (n=1).
Default output256 tokens, including reasoning.
Maximum outputModel limit capped at 65,536; 4,096 if unspecified.
Gateway timeout120 seconds.

Retry deliberately

Fix authentication, balance and validation errors before retrying. For transient errors, use bounded backoff and honor Retry-After. Avoid immediately replaying after a timeout: generation may already have begun.

Tokely does not automatically retry inference. SDK retry settings are separate. Every new generation attempt can consume usage.

Get your API key Browse models