All models

OpenAI

gpt-5-nano-2025-08-07

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

textAPI availablestreamingreasoningfunction tools
gpt-5-nano-2025-08-07Version date 2025-08-07

Tokely price / 1M tokens

$0.091913input
$0.009192cached input
$0.7353output

USD · Published rate · Updated 2026-10-06

OpenRouter reference$0.05 input · $0.4 output / 1M tokensStandard listed rates · 2026-10-06 · family reference; snapshot rates may differ

Context window

400,000tokens · model specification

Model max output

128,000tokens

First request

Available
curl https://tokely.me/v1/chat/completions \
  -H "Authorization: Bearer $TOKELY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-nano-2025-08-07",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 256
  }'

Set TOKELY_API_KEY to your dashboard key. Requests require a funded Tokely balance. Base URL: https://tokely.me/v1.

Compare nearby prices

USD per million tokens. Compare price and context; these are not quality benchmarks.
ModelInputOutputContext
gpt-5-nano-2025-08-07OpenAI · this model$0.091913$0.7353400,000
gpt-5-nanoOpenAI$0.091913$0.7353400,000
gpt-4.1-nanoOpenAI$0.183825$0.73531,047,576
gpt-4.1-nano-2025-04-14OpenAI$0.183825$0.73531,047,576
qwen-turboQwen$0.1375$0.55Unconfirmed

Capabilities & limits

Provider
OpenAI
API model ID
gpt-5-nano-2025-08-07
Category
Text
Endpoints
POST /v1/chat/completionsPOST /v1/responses
Tokely max output
65,536 tokens / request, including reasoning tokens
Default output
256 tokens. Set an explicit limit for longer responses.
Request limits
256 KB JSON body · 128 messages · 32 function tools
Rate limit
60 requests / minute / account, across its keys
Supported input
Text messages and function tool results
Model modalities
text, image, file. Tokely currently accepts text input.

The context includes your instructions, conversation history, tool definitions and generated output. Keep the input and requested output within the model’s context window.

OpenRouter model reference

API reference

Authenticate with Authorization: Bearer YOUR_TOKELY_API_KEY. Use an OpenAI SDK with the Tokely base URL.

model
Required string: gpt-5-nano-2025-08-07
messages
An array of system, developer, user, assistant and tool messages.
max_tokens
Optional integer, 1–65,536. Defaults to 256.
stream
Set true for Server-Sent Events (SSE).
reasoning_effort
Optional reasoning effort. Supported values depend on the model; omit to use its default.
tools
Function definitions executed by your application. Provider-hosted tools are not enabled.

The response contains choices with message or tool_calls, finish_reason, and usage (prompt_tokens, completion_tokens, total_tokens). Streaming sends chat completion chunks.

400 / 415
Invalid request, parameter, endpoint for this model, or content type.
401 / 403
Invalid, revoked or restricted API key.
402
Insufficient account balance or API key budget.
429
Request limit reached. Retry after the indicated delay.
502 / 503 / 504
Upstream failure, temporary service unavailability or request timeout.

Frequently asked questions

How much does gpt-5-nano-2025-08-07 cost?

$0.091913 per million input tokens, $0.7353 per million output tokens and $0.009192 per million cached input tokens. Usage is charged to your Tokely balance.

Do I need a separate OpenAI account?

Use your Tokely API key and balance. A separate provider key is not required.

How do I switch to this model?

Set model to gpt-5-nano-2025-08-07 and use /v1/chat/completions. Keep your Tokely base URL and API key.

What is the context window?

400,000 tokens, including input and output. The Tokely request body and output limits listed above also apply.

Related versions