DeepSeek
deepseek-v4-flash-0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Tokely price / 1M tokens
USD · Published rate · Updated 2026-10-06
Context window
1,048,576tokens · model specificationModel max output
943,718tokensFirst request
Availablecurl https://tokely.me/v1/chat/completions \
-H "Authorization: Bearer $TOKELY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash-0731",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 256
}'Set TOKELY_API_KEY to your dashboard key. Requests require a funded Tokely balance. Base URL: https://tokely.me/v1.
Compare nearby prices
| Model | Input | Output | Context |
|---|---|---|---|
| deepseek-v4-flash-0731DeepSeek · this model | $1.21 | $3.63 | 1,048,576 |
| glm-4.5-airChatGLM | $0.825 | $4.033354 | 131,072 |
| qwen3-14bQwen | $0.9625 | $3.85 | 131,072 |
| qwen3-vl-flashQwen | $0.55 | $4.4 | Unconfirmed |
| grok-4Grok | $0.82725 | $4.13625 | Unconfirmed |
Capabilities & limits
- Provider
- DeepSeek
- API model ID
deepseek-v4-flash-0731- Category
- Text
- Endpoints
POST /v1/chat/completions- Tokely max output
- 65,536 tokens / request, including reasoning tokens
- Default output
- 256 tokens. Set an explicit limit for longer responses.
- Request limits
- 256 KB JSON body · 128 messages · 32 function tools
- Rate limit
- 60 requests / minute / account, across its keys
- Supported input
- Text messages and function tool results
- Model modalities
- text. Tokely currently accepts text input.
The context includes your instructions, conversation history, tool definitions and generated output. Keep the input and requested output within the model’s context window.
OpenRouter model referenceAPI reference
Authenticate with Authorization: Bearer YOUR_TOKELY_API_KEY. Use an OpenAI SDK with the Tokely base URL.
- model
- Required string:
deepseek-v4-flash-0731 - messages
- An array of system, developer, user, assistant and tool messages.
- max_tokens
- Optional integer, 1–65,536. Defaults to 256.
- stream
- Set true for Server-Sent Events (SSE).
- reasoning_effort
- Optional reasoning effort. Supported values depend on the model; omit to use its default.
- tools
- Function definitions executed by your application. Provider-hosted tools are not enabled.
The response contains choices with message or tool_calls, finish_reason, and usage (prompt_tokens, completion_tokens, total_tokens). Streaming sends chat completion chunks.
- 400 / 415
- Invalid request, parameter, endpoint for this model, or content type.
- 401 / 403
- Invalid, revoked or restricted API key.
- 402
- Insufficient account balance or API key budget.
- 429
- Request limit reached. Retry after the indicated delay.
- 502 / 503 / 504
- Upstream failure, temporary service unavailability or request timeout.
Frequently asked questions
How much does deepseek-v4-flash-0731 cost?
$1.21 per million input tokens, $3.63 per million output tokens and $0.0385 per million cached input tokens. Usage is charged to your Tokely balance.
Do I need a separate DeepSeek account?
Use your Tokely API key and balance. A separate provider key is not required.
How do I switch to this model?
Set model to deepseek-v4-flash-0731 and use /v1/chat/completions. Keep your Tokely base URL and API key.
What is the context window?
1,048,576 tokens, including input and output. The Tokely request body and output limits listed above also apply.