Streaming
Stream model output token by token via SSE.
Set stream: true on POST /v1/chat/completions to receive tokens via Server-Sent Events (SSE).
Request
curl https://api.caedral.com/v1/chat/completions \ -H "Authorization: Bearer $CAEDRAL_API_KEY" \ -H "Content-Type: application/json" \ -N \ -d '{ "model": "deepseek/deepseek-v4-flash", "messages": [{ "role": "user", "content": "Count to five." }], "stream": true }'Event format
Content-Type is text/event-stream. Each event is a data: line with a chat.completion.chunk JSON object. The stream ends with data: [DONE]. Chunks stream through Caedral unmodified — deltas arrive exactly as the upstream model emits them.
Usage in the final chunk
When the model reports token usage, it arrives on the final chunk's usage object rather than a separate event. Billing settles against that usage once the stream completes; if usage cannot be parsed, the reservation made at request start is billed instead.
Timeouts and aborts
| Guard | Value | What happens |
|---|---|---|
| Idle watchdog | 60s without bytes from the model | Caedral aborts the upstream call, emits a terminal SSE error event, and releases the billing reservation — you are not charged |
| Non-streaming total timeout | 120s hard cap | Applies only to non-streaming calls; streams have no total cap while tokens keep flowing |
| Mid-stream interruption | Any upstream failure | A terminal data: line carrying the error object is written before the stream closes |
Because HTTP 200 is already sent when streaming starts, later failures surface inside the stream: expect a final data: event shaped like {"error":{"type":"upstream_error","code":502,"message":"…"}} instead of a silent end. Clients should treat any terminal error event as a retryable failure and check whether partial output arrived.