API REFERENCE
Chat Completions
Generate a response from a sequence of text messages.
POST /v1/chat/completions
Send a model ID and an array of messages. Supported roles are system, developer, user, assistant and tool; role support also depends on the model. Message content must be text, with null allowed for an assistant function call.
curl https://api.tokely.me/v1/chat/completions \
-H "Authorization: Bearer $TOKELY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello, Tokely!"}],
"max_tokens": 256
}'Core parameters
Optional generation settings must be supported by the selected model.
| Parameter | Description |
|---|---|
| model | Required. Exact enabled model ID. |
| messages | Required. Between 1 and 128 text messages. |
| max_tokens / max_completion_tokens | Output budget. Use one, never both. Default: 256. |
| stream | Set true for SSE. Default: false. |
| temperature / top_p | Sampling settings, when supported. |
| tools / tool_choice | Up to 32 function tools; model support required. |
| response_format | text, json_object or json_schema; model support required. |
| n | Only 1 is supported. |
Continue a conversation
Send the prior messages along with the new user message. The following example makes two requests and includes the first answer in the second request.
Each request includes and bills its submitted input history. Trim old turns when the conversation approaches the model context limit. For assistant tool calls, preserve the tool_calls fields and corresponding tool results as described in Function calling.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TOKELY_API_KEY"],
base_url="https://api.tokely.me/v1",
)
messages = [
{"role": "system", "content": "Answer briefly."},
{"role": "user", "content": "What is an API?"},
]
first = client.chat.completions.create(
model="gpt-4o-mini", messages=messages, max_tokens=256,
)
print(first.choices[0].message.content)
messages.append({
"role": "assistant", "content": first.choices[0].message.content,
})
messages.append({"role": "user", "content": "Give me an example."})
second = client.chat.completions.create(
model="gpt-4o-mini", messages=messages, max_tokens=256,
)
print(second.choices[0].message.content)Read the response
Read choices[0].message.content for text, choices[0].message.tool_calls for function calls, and usage for token counts. A finish reason of length indicates that the output budget was reached.
Carry conversation history forward in messages on the next request. Tokely does not automatically retain your conversation.