Docs/Chat Completions

API REFERENCE

Chat Completions

Generate a response from a sequence of text messages.

POST /v1/chat/completions

Send a model ID and an array of messages. Supported roles are system, developer, user, assistant and tool; role support also depends on the model. Message content must be text, with null allowed for an assistant function call.

curl https://api.tokely.me/v1/chat/completions \
  -H "Authorization: Bearer $TOKELY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Hello, Tokely!"}],
    "max_tokens": 256
  }'

Core parameters

Optional generation settings must be supported by the selected model.

ParameterDescription
modelRequired. Exact enabled model ID.
messagesRequired. Between 1 and 128 text messages.
max_tokens / max_completion_tokensOutput budget. Use one, never both. Default: 256.
streamSet true for SSE. Default: false.
temperature / top_pSampling settings, when supported.
tools / tool_choiceUp to 32 function tools; model support required.
response_formattext, json_object or json_schema; model support required.
nOnly 1 is supported.

Continue a conversation

Send the prior messages along with the new user message. The following example makes two requests and includes the first answer in the second request.

Each request includes and bills its submitted input history. Trim old turns when the conversation approaches the model context limit. For assistant tool calls, preserve the tool_calls fields and corresponding tool results as described in Function calling.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["TOKELY_API_KEY"],
    base_url="https://api.tokely.me/v1",
)
messages = [
    {"role": "system", "content": "Answer briefly."},
    {"role": "user", "content": "What is an API?"},
]
first = client.chat.completions.create(
    model="gpt-4o-mini", messages=messages, max_tokens=256,
)
print(first.choices[0].message.content)
messages.append({
    "role": "assistant", "content": first.choices[0].message.content,
})
messages.append({"role": "user", "content": "Give me an example."})
second = client.chat.completions.create(
    model="gpt-4o-mini", messages=messages, max_tokens=256,
)
print(second.choices[0].message.content)

Read the response

Read choices[0].message.content for text, choices[0].message.tool_calls for function calls, and usage for token counts. A finish reason of length indicates that the output budget was reached.

Carry conversation history forward in messages on the next request. Tokely does not automatically retain your conversation.

Get your API key Browse models