Qwen
qwen3-vl-32b-thinking
The reasoning version of the largest Dense model in the Qwen3-VL series. Multimodal reasoning ranks second only to Qwen3-VL-235B-Thinking, with outstanding STEM and math problem-solving, general image and video understanding, and SOTA multimodal Agent capabilities — ideal for complex multimodal reasoning tasks.
Tokely price / 1M tokens
USD · Published rate · Updated 2026-10-06
Context window
UnconfirmedNo verified context specificationFirst request
Availablecurl https://tokely.me/v1/chat/completions \
-H "Authorization: Bearer $TOKELY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-32b-thinking",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 256
}'Set TOKELY_API_KEY to your dashboard key. Requests require a funded Tokely balance. Base URL: https://tokely.me/v1.
Compare nearby prices
| Model | Input | Output | Context |
|---|---|---|---|
| qwen3-vl-32b-thinkingQwen · this model | $0.44 | $1.76 | Unconfirmed |
| qwen3-32bQwen | $0.44 | $1.76 | 131,072 |
| qwen3-vl-32b-instructQwen | $0.44 | $1.76 | 131,072 |
| gpt-6-lunaOpenAI | $0.5625 | $1.6875 | 1,050,000 |
| qwen3-vl-8b-instructQwen | $0.495 | $1.925001 | 262,144 |
Capabilities & limits
- Provider
- Qwen
- API model ID
qwen3-vl-32b-thinking- Category
- Text
- Endpoints
POST /v1/chat/completions- Tokely max output
- 4,096 tokens / request, including reasoning tokens
- Default output
- 256 tokens. Set an explicit limit for longer responses.
- Request limits
- 256 KB JSON body · 128 messages · 32 function tools
- Rate limit
- 60 requests / minute / account, across its keys
- Supported input
- Text messages and function tool results
API reference
Authenticate with Authorization: Bearer YOUR_TOKELY_API_KEY. Use an OpenAI SDK with the Tokely base URL.
- model
- Required string:
qwen3-vl-32b-thinking - messages
- An array of system, developer, user, assistant and tool messages.
- max_tokens
- Optional integer, 1–4,096. Defaults to 256.
- stream
- Set true for Server-Sent Events (SSE).
The response contains choices with message or tool_calls, finish_reason, and usage (prompt_tokens, completion_tokens, total_tokens). Streaming sends chat completion chunks.
- 400 / 415
- Invalid request, parameter, endpoint for this model, or content type.
- 401 / 403
- Invalid, revoked or restricted API key.
- 402
- Insufficient account balance or API key budget.
- 429
- Request limit reached. Retry after the indicated delay.
- 502 / 503 / 504
- Upstream failure, temporary service unavailability or request timeout.
Frequently asked questions
How much does qwen3-vl-32b-thinking cost?
$0.44 per million input tokens, $1.76 per million output tokens and $0.44 per million cached input tokens. Usage is charged to your Tokely balance.
Do I need a separate Qwen account?
Use your Tokely API key and balance. A separate provider key is not required.
How do I switch to this model?
Set model to qwen3-vl-32b-thinking and use /v1/chat/completions. Keep your Tokely base URL and API key.