API reference
All gateway endpoints — chat, models, usage, images, embeddings, audio, rerank.
On this page
- POST /v1/chat/completions
- Request parameters
- Example request
- Example response
- Response fields
- POST /v1/embeddings
- Request parameters
- Response fields
- POST /v1/rerank
- Request parameters
- Response fields
- GET /v1/models
- Response fields
- GET /v1/models/:modelId
- GET /v1/usage
- GET /health
- GET /v1/status
- GET /v1/status/models
This page documents the production gateway at https://api.caedral.com. Authenticate with Authorization: Bearer <api_key> unless noted. Paid routes draw from included usage pools; exhausted pools return HTTP 402 unless on-demand is enabled. Every successful debit is recorded in an append-only billing ledger tied to the request.
POST /v1/chat/completions
Create a chat completion. Schema is OpenAI-compatible so existing SDKs work when baseURL is set to https://api.caedral.com/v1.
Request parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| model | string | Yes | — | Any live catalog model ID from GET /v1/models (example: deepseek/deepseek-v4-flash) |
| messages | array | Yes | — | Ordered chat messages; each item has role and content |
| messages[].role | string | Yes | — | system | user | assistant | tool |
| messages[].content | string | array | Yes* | — | Message text or multimodal parts where the model allows. *Assistant messages with tool_calls may omit content |
| temperature | number | No | model default | Sampling temperature, typically 0–2 |
| top_p | number | No | model default | Nucleus sampling |
| max_tokens | number | No | — | Maximum completion tokens to generate |
| stream | boolean | No | false | If true, respond with text/event-stream SSE |
| stop | string | string[] | No | — | Stop sequences |
| presence_penalty | number | No | 0 | Presence penalty |
| frequency_penalty | number | No | 0 | Frequency penalty |
| user | string | No | — | End-user identifier for your own abuse tracking |
| tools | array | No | — | Tool/function definitions — passed through for models that support tool calling |
| tool_choice | string | object | No | — | none | auto | required | specific tool |
| response_format | object | No | — | e.g. { "type": "json_object" } for structured output |
| notre | object | No | — | Optional Notre public contract v1: { mode: "off"|"auto", telemetry?: boolean }. Stripped before upstream. See /docs/notre. |
Example request
{ "model": "deepseek/deepseek-v4-flash", "messages": [ { "role": "system", "content": "You are a helpful assistant." }, { "role": "user", "content": "Explain Caedral subscription billing in one paragraph." } ], "temperature": 0.3, "max_tokens": 500}Example response
{ "id": "cd_req_01HXYZ", "object": "chat.completion", "created": 1722268800, "model": "deepseek/deepseek-v4-flash", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Usage draws from included monthly pools..." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 42, "completion_tokens": 96, "total_tokens": 138 }}Response fields
| Field | Type | Description |
|---|---|---|
| id | string | Request id (cd_req_… when not supplied upstream) for support and logs |
| object | string | Always chat.completion for non-streaming |
| created | number | Unix timestamp |
| model | string | Model id that served the request — passed through as-is |
| choices | array | Completion choices (usually length 1) |
| choices[].message.role | string | assistant |
| choices[].message.content | string | Assistant text |
| choices[].message.tool_calls | array | Present when the model requests a tool call |
| choices[].finish_reason | string | stop | length | tool_calls | content_filter |
| usage.prompt_tokens | number | Input tokens billed |
| usage.completion_tokens | number | Output tokens billed |
| usage.total_tokens | number | prompt + completion |
POST /v1/embeddings
Create vector embeddings with caedral-embed (Caedral E1 Small, model caedral-embed-e1-small-v1). Returns 384-dimensional L2-normalized vectors (512-token context). Free through 28 September 2026 (130 RPM); afterward billed from the Caedral Models pool at published rates.
Request parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| model | string | No | caedral-embed | Embedding model id (caedral-embed-e1-small-v1) |
| input | string | string[] | Yes | — | Text or batch of texts to embed |
| encoding_format | string | No | float | float | base64 when supported |
| user | string | No | — | Optional end-user id |
{ "model": "caedral-embed", "input": [ "Included pools stop spend when exhausted (402) unless on-demand is enabled.", "n8n HTTP Request node with Bearer auth." ]}{ "id": "cd_req_...", "object": "list", "model": "caedral-embed-e1-small-v1", "data": [ { "object": "embedding", "index": 0, "embedding": [0.0123, -0.0441, 0.0087] }, { "object": "embedding", "index": 1, "embedding": [0.0011, 0.0330, -0.0194] } ], "usage": { "prompt_tokens": 18, "total_tokens": 18 }}Response fields
| Field | Type | Description |
|---|---|---|
| data | array | One embedding object per input |
| data[].index | number | Position matching the input array |
| data[].embedding | number[] | 384-dimensional L2-normalized dense vector |
| usage.prompt_tokens | number | Tokens processed for billing |
| usage.total_tokens | number | Same as prompt_tokens for embeddings |
POST /v1/rerank
Reorder documents by relevance to a query using caedral-rerank (BAAI/bge-reranker-v2-m3). Hosted on Caedral infrastructure. Free through 28 September 2026 (130 RPM); afterward $0.0005 per search from the External Models pool. Use after vector retrieval to improve grounded generation.
Request parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| model | string | No | caedral-rerank | Rerank model id |
| query | string | Yes | — | Search / user query |
| documents | string[] | Yes | — | Candidate documents (maximum 100) |
| top_n | number | No | all | Return only the top N results |
| return_documents | boolean | No | false | Include document text in results when supported |
{ "model": "caedral-rerank", "query": "How does Caedral billing work?", "documents": [ "API usage draws from subscription pools (or on-demand when enabled).", "Semantic cache hits bill at half price; opt out with X-Caedral-Cache: off.", "Free-tier catalog models cost $0 with an active subscription." ], "top_n": 2}{ "id": "cd_req_...", "model": "caedral-rerank", "results": [ { "index": 0, "relevance_score": 0.91 }, { "index": 2, "relevance_score": 0.74 } ]}Response fields
| Field | Type | Description |
|---|---|---|
| results | array | Ranked hits, highest score first |
| results[].index | number | Index into the request documents array |
| results[].relevance_score | number | Higher means more relevant |
GET /v1/models
The live public catalog: real lab model IDs plus the Caedral-hosted products. Authentication is not required. Unknown IDs return 404.
curl https://api.caedral.com/v1/modelsResponse fields
| Field | Type | Description |
|---|---|---|
| object | string | list |
| data | array | Model objects |
| data[].id | string | Model id to pass as model in requests |
| data[].object | string | model |
| data[].owned_by | string | caedral (or the hosting lab for hosted products) |
| data[].name | string | Human-readable name |
| data[].context_length | number | Maximum context window when known |
| data[].pricing | object | Published pricing for the model |
| data[].architecture | object | Modality and tokenizer metadata when known |
| data[].supported_parameters | array | Parameters this model accepts |
| data[].is_free | boolean | True for free-tier catalog models ($0, active-subscription gate) |
| data[].is_caedral_hosted | boolean | True for Caedral-hosted products (Embed and Voice debit Caedral Models; hosted rerank debit External Models) |
| data[].recommended_endpoint | string | Endpoint Caedral recommends for this model |
GET /v1/models/:modelId
Metadata for a single model. Returns 404 for unknown ids.
curl https://api.caedral.com/v1/models/deepseek/deepseek-v4-flashGET /v1/usage
Returns subscription plan, usage pools, billing period, and on-demand state for the API key owner. Bearer auth required.
| Field | Type | Description |
|---|---|---|
| accountStatus | string | active | payment_pending |
| plan | object | id, name, interval, status |
| billingPeriod | object | start / end ISO timestamps of the current cycle |
| pools.caedral | object | usedMilli, limitMilli, usedFormatted, limitFormatted, percentUsed for Caedral-hosted products |
| pools.external | object | External models pool (Pro or higher); available flags whether the plan includes external access |
| onDemand | object | mode (disabled | fixed | unlimited), allowed, blocked, accrued/spent totals |
GET /health
Gateway liveness. Public. Response: { "status": "ok" }.
GET /v1/status
Public status index. Points to health and per-model uptime endpoints. No authentication.
GET /v1/status/models
Per-model uptime for public catalog model IDs. Probe uptime is measured every minute from Caedral gateway catalog health (embed/rerank also require model-inference). Request success rate uses Caedral execution logs. Public — no API key. Also on https://caedral.com/status and GET /api/status.