GUIDES
Let the response flow.
Display generated text as it arrives with server-sent events.
Enable streaming
Set stream=true to receive text/event-stream. For Chat Completions, read text from choices[0].delta.content. Some models require streaming and return stream_required for a non-streaming request.
# Use the client configured in Quickstart.
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Tell me a short story."}],
max_tokens=512,
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)Two endpoints, two event formats
Chat Completions sends data events containing completion chunks and ends with [DONE]. Responses uses typed events such as response.output_text.delta and response.completed. Use an SSE parser or the SDK; network chunks are not always complete events.
Handle interruptions
A connection can close after partial output. Preserve that text and show an interrupted state instead of treating it as complete. Cancel the request when your user stops generation.
Tokely does not automatically retry inference. A retry is a new request and can incur additional usage. SDKs may retry independently; configure that behavior in your client.