Skip to content

OpenAI-compatible AI API

One key. Subscription pools. Notre on eligible chat.

Caedral is an OpenAI-compatible API: chat, Caedral-operated Embed, Rerank, and Voice, plus frontier labs, all on one key. Every plan includes monthly usage pools. Notre Runtime is included on every plan and applied on eligible POST /v1/chat/completions when the platform runs Notre — not on embeddings, rerank, voice, or images.

POST/v1/chat/completions

Request

import { Caedral } from "caedral"

const client = new Caedral({
  apiKey: process.env.CAEDRAL_API_KEY,
})

const completion = await client.chat.completions.create({
  model: "anthropic/claude-haiku-4.5",
  messages: [{ role: "user", content: "Hello" }],
})

Result · illustrative Notre workload

Status
200 OK
Model
anthropic/claude-haiku-4.5
Input sent
0
Output
0
Notre
0 saved

Notre Runtime applies on chat completions only. Embed, Rerank, and Voice are different request shapes.

Live platform

Models in catalog
0
Caedral-operated products
0

Embed, Rerank, Voice, Notre Runtime

Up to, verified Notre
0.0%

More than 60% fewer tokens in validated Notre workloads.

How Caedral works

One path from your app to the model ecosystem.

Your client talks to Caedral. Caedral authenticates, meters, optionally compiles chat context, and routes to a Caedral-operated product or an external lab. You do not hold a stack of provider keys.

  1. 01

    Your application

    SDK, OpenAI-compatible client, n8n, or REST.

  2. 02 · Caedral

    • Unified API
    • Notre Runtime
    • Routing / model access
    • Usage & billing
  3. 03

    Caedral products + external labs

    Embed, Voice, hosted Rerank, and frontier chat models — billed on the pool that model belongs to.

Workloads

One API. Different request shapes.

Chat, retrieval, speech, and automation share authentication and billing. They do not share a compiler.

POST /v1/chat/completions

Model inference on the OpenAI-compatible chat path. Notre Runtime may compile context on this endpoint when the deployment allows.

Notre Runtime is native here — included, not a SKU.

Open this path

Built by Caedral

Caedral-operated products, on the same key.

Frontier labs are in the catalog. These four are developed or hosted by Caedral. Classification is explicit — we do not hide our own work behind someone else’s name.

Caedral Embed

caedral-embed-e1-small-v1

384-dimensional embeddings. Developed and operated by Caedral. POST /v1/embeddings.

Caedral Models

Caedral Rerank

BAAI/bge-reranker-v2-m3

Hosted by Caedral over an external foundation model. POST /v1/rerank.

External Models

Caedral Voice

caedral-voice-1

English WAV speech. Developed and operated by Caedral. Notre Runtime does not apply. POST /v1/audio/speech.

Caedral Models

Notre Runtime

eligible chat

Native context compiler on POST /v1/chat/completions. Not a SKU. MCP is an integration, not the product name.

Included

Catalog

Live model IDs. Published rates.

551 models across 86 providers, plus the Caedral-operated products above.

ModelIDContext
Anthropic: Claude Fable Latest~anthropic/claude-fable-latest1,000,000
Anthropic: Claude Haiku Latest~anthropic/claude-haiku-latest200,000
Anthropic: Claude Opus Latest~anthropic/claude-opus-latest1,000,000
Anthropic: Claude Sonnet Latest~anthropic/claude-sonnet-latest1,000,000
DeepSeek: DeepSeek Flash Latest~deepseek/deepseek-flash-latest1,048,576
DeepSeek: DeepSeek Pro Latest~deepseek/deepseek-pro-latest1,048,576
DeepSeek: DeepSeek V4 Flash Latest~deepseek/deepseek-v4-flash-latest1,310,720
Full catalog

Notre Runtime

Same plan. More useful model work.

Notre Runtime is built into Caedral. No add-on. No separate bill. Designed to reduce unnecessary context without degrading the model response. More than 60% fewer tokens in validated Notre workloads.

Verified benchmark — M1.4 development corpus — Notre Runtime
  1. 01Large request context
  2. 02Notre Runtime
  3. 03Reduced model input
  4. 04Model inference
  5. 05Result
Input before
0
Input sent
0
Saved
0

0.0%reduction

TRUE AGGREGATE on the recorded corpus (n=4,962). Output tokens are generated by the model and are not claimed as Notre savings.

Notre is Caedral's research and development program.Notre Runtime is the production layer on eligible chat. Embed, Voice, and Rerank are different request shapes — we do not pretend they share this compiler.

Notre economics

Verified reduction. Projected at Caedral prices.

Inspect recorded compiler benchmarks, then scale the same ratio to a realistic input budget. The model control sets the current Caedral input tariff — Notre was not A/B tested on each catalog model.

Verified Notre benchmark

Same plan. Same price. More useful AI work.

Notre Runtime is built into Caedral. No add-on. No separate bill.Reduction is from the compiler corpus — not a live A/B of the catalog model you pick. Dollars use that model's current Caedral input tariff.

Up to
0.0%

fewer input tokens, strongest verified

Median
0.0%

across 2 verified workloads

Notre Runtime
Included

No add-on fee. No separate Notre bill.

Benchmark

Model (Caedral input tariff)

Economic scale

Projected from verified benchmark reduction. Same ratio, normalized workload.

Same budget

0.00×

equivalent input workload

  1. 01Input before Notre
  2. 02Notre Runtime compiles
  3. 03Input sent to the model
  • Input before0
  • Input sent0
  • Input saved0
Reduction
0.0%
Quality
Required-set recall 91.9%

Projected from verified benchmark

Input tokens avoided

0

of 10,000,000 baseline input

Usage value preserved · External Models pool

$0.00

Current Caedral input tariff only. Output is not counted as savings.

More use from the same plan

0→0

equivalent requests on the same input budget

Compiler time (not model round-trip)

p50 12.0 ms · p95 133.7 ms

Pro+ External Models pool · $70/mo

Without Notre-equivalent: 20,289,855 input tokens.

With this reduction: 55,340,826 equivalent input tokens.

Same included allowance, projected from the verified ratio and this model's current input price.

At this benchmark's 63.3% reduction rate, a 10,000,000-token input allowance can process roughly 2.73× the equivalent workload before exhausting the same input budget.

Benchmark details
Workload
M1.4 development corpus
Date
2026-08-30
Runs
4,962
Policy
V1
Input before (recorded)
567,755,348
Input after (recorded)
208,158,685
Quality methodology
4,470 of 4,962 required tool-sets preserved (492 failures). Design constraint is to reduce unnecessary context without degrading the model response. This is not a claim of zero quality loss.
Source
caedral-notre-mcp/benchmarks/results/M1.4-GUARDED-V3.json#v1_shadow_baseline.m14_development
Confidence
TRUE AGGREGATE SUM/SUM. Holdout not run. Decision at test time: shadow only. Not a production SLA.
Fallback rate
36.4%

Full verified table

Model economics

Stop collecting provider accounts.

Direct integration means a key, a bill, and an SDK per lab. Caedral is one integration, one key, one catalog, and one billing layer — Notre included on eligible chat.

Direct integration

  • Provider A key
  • Provider B key
  • Provider C key
  • Separate invoices
  • Separate SDK behavior
  • Separate usage dashboards

Caedral

  • One API · OpenAI-compatible
  • One key
  • One catalog of live model IDs
  • Two usage pools, hard stop, optional on-demand
  • Notre Runtime included
  • One console for usage, keys, teams
Platform overview

Long-running work

Agents, RAG, and tools that keep going.

Coding assistants, retrieval pipelines, and n8n automations send large context. Notre exists so eligible chat can keep more of that work inside the included pool.

  • Agents and coding tools. One key across models. Chat completions can carry Notre when the deployment allows.
  • RAG. Caedral Embed and Caedral Rerank on the same account as chat. Compiler does not apply to those endpoints.
  • Automation. n8n checks usage, then calls chat. HTTP 402 is the hard stop unless on-demand is on.

Developers

Keep the client you already have.

Point an OpenAI-compatible SDK at https://api.caedral.com/v1, or use the official caedral packages. Automation goes through n8n. Agents can speak MCP locally.

Quickstart
  1. 01

    REST API

    Chat, embeddings, rerank, images, audio, video on api.caedral.com.

  2. 02

    TypeScript SDK

    npm install caedral. Official client. OpenAI-compatible clients also work with a base URL change.

  3. 03

    Python SDK

    pip install caedral. Same methods as the TypeScript client.

  4. 04

    Go, Java, C

    First-party clients in the monorepo for systems languages. Prefer TS/Python when a registry install is enough.

  5. 05

    n8n

    Official community node n8n-nodes-caedral, or HTTP Request with Bearer auth. GET /v1/usage before spend.

  6. 06

    CLI

    The caedral CLI talks to Notre Runtime through the gateway. There is no separate MCP product and no public MCP hostname.

Console

Usage, Notre, keys, and billing in one account.

The dashboard is the same identity as your API key. Pools, Notre analytics, teams, and service health live here — not in a third-party vendor console.

Open dashboard

Dashboard preview

Caedral Models

Included pool

External Models

Paid plans

Notre Runtime

Input saved · reduction · value preserved

Control

API keys · team usage · billing portal · status

Pricing

Two pools. A hard stop. Optional on-demand.

Caedral Models covers hosted embed and voice. External Models covers frontier APIs and hosted rerank. Annual billing saves 20%. Notre Runtime is built into Caedral. No add-on. No separate bill.

PlanPriceIncluded
HobbyFree$200 Caedral Models / mo
Pro$20/mo$500 Caedral + $20 external
TeamsFrom $40/seatPooled usage, owner billing
Buy a plan

Research

Notre is a program, not a checkbox.

Runtime is the customer-facing compiler on chat. Engine is local CPU inference, pre-alpha, and not a hosted API. Notre DB is research and is not a product.

  • Notre Runtime

    Production/customer output of Notre. Experimental. Native to eligible chat.

  • Notre Engine

    PRE-ALPHA local CPU MoE inference. Not a hosted Caedral API.

Create a key and send a request.

Sign up, generate a cd_live_ key, then POST /v1/chat/completions to https://api.caedral.com. Pricing covers pools; the docs cover curl, SDKs, and HTTP 402.