integration · 2026-09-28

Streaming, timeouts and retries in an API client

Streaming helps users see output sooner, while connection, reading and retry failures need separate handling.

RelayAI · 中文

Enable streaming

Set stream: true on a compatible Chat Completions request to receive SSE chunks. Render chunks as they arrive and save the complete answer when the stream ends. Distinguish a failure before the first chunk from an interruption after partial output.

Choose timeouts by phase

A connection timeout detects an unreachable path; a read timeout must allow ongoing model output. Measure your own model, prompt sizes and user expectations before choosing SDK and proxy thresholds.

Retry with a clear boundary

Use limited backoff for 429, 502 and 503, honoring Retry-After where available. Replaying an already-partial stream can produce a different answer and another billable request. Let the application decide when a user should start over and keep the request_id for support.

Sources

Read API docs

Related articles

Connect the OpenAI Python SDK to RelayAI