Qwen
qwen3.8-max
240 trillion parameter MoE flagship, comprehensive improvement in programming and office capabilities, capable of independently programming and delivering complete projects in just over ten days. Competent in hundreds of professional tasks such as law, finance, design, delivering production-level results end-to-end in a single conversation. Native visual understanding throughout the planning, execution, and verification process, supporting deep semantic analysis of long documents and lengthy videos. Autonomous planning and closed-loop iteration in long-range tasks, continuously evolving.
Tokely price / 1M tokens
USD · Published rate · Updated 2026-10-06
Context window
UnconfirmedNo verified context specificationFirst request
Availablecurl https://tokely.me/v1/chat/completions \
-H "Authorization: Bearer $TOKELY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-max",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 256
}'Set TOKELY_API_KEY to your dashboard key. Requests require a funded Tokely balance. Base URL: https://tokely.me/v1.
Compare nearby prices
| Model | Input | Output | Context |
|---|---|---|---|
| qwen3.8-maxQwen · this model | $5.5 | $16.5 | Unconfirmed |
| qwen-maxQwen | $4.4 | $17.6 | Unconfirmed |
| qwen3.8-max-0902Qwen | $5.5 | $16.5 | 1,000,000 |
| qwq-32bQwen | $5.5 | $16.5 | Unconfirmed |
| glm-4.7ChatGLM | $3.85 | $17.966667 | 204,800 |
Capabilities & limits
- Provider
- Qwen
- API model ID
qwen3.8-max- Category
- Text
- Endpoints
POST /v1/chat/completions- Tokely max output
- 4,096 tokens / request, including reasoning tokens
- Default output
- 256 tokens. Set an explicit limit for longer responses.
- Request limits
- 256 KB JSON body · 128 messages · 32 function tools
- Rate limit
- 60 requests / minute / account, across its keys
- Supported input
- Text messages and function tool results
API reference
Authenticate with Authorization: Bearer YOUR_TOKELY_API_KEY. Use an OpenAI SDK with the Tokely base URL.
- model
- Required string:
qwen3.8-max - messages
- An array of system, developer, user, assistant and tool messages.
- max_tokens
- Optional integer, 1–4,096. Defaults to 256.
- stream
- Set true for Server-Sent Events (SSE).
The response contains choices with message or tool_calls, finish_reason, and usage (prompt_tokens, completion_tokens, total_tokens). Streaming sends chat completion chunks.
- 400 / 415
- Invalid request, parameter, endpoint for this model, or content type.
- 401 / 403
- Invalid, revoked or restricted API key.
- 402
- Insufficient account balance or API key budget.
- 429
- Request limit reached. Retry after the indicated delay.
- 502 / 503 / 504
- Upstream failure, temporary service unavailability or request timeout.
Frequently asked questions
How much does qwen3.8-max cost?
$5.5 per million input tokens, $16.5 per million output tokens and $0.6875 per million cached input tokens. Usage is charged to your Tokely balance.
Do I need a separate Qwen account?
Use your Tokely API key and balance. A separate provider key is not required.
How do I switch to this model?
Set model to qwen3.8-max and use /v1/chat/completions. Keep your Tokely base URL and API key.