caedral-embed-e1-small-v1
384-dimensional embeddings. Developed and operated by Caedral. POST /v1/embeddings.
Caedral Models
OpenAI-compatible AI API
Caedral is an OpenAI-compatible API: chat, Caedral-operated Embed, Rerank, and Voice, plus frontier labs, all on one key. Every plan includes monthly usage pools. Notre Runtime is included on every plan and applied on eligible POST /v1/chat/completions when the platform runs Notre — not on embeddings, rerank, voice, or images.
POST/v1/chat/completions
Request
import { Caedral } from "caedral"
const client = new Caedral({
apiKey: process.env.CAEDRAL_API_KEY,
})
const completion = await client.chat.completions.create({
model: "anthropic/claude-haiku-4.5",
messages: [{ role: "user", content: "Hello" }],
})Result · illustrative Notre workload
Notre Runtime applies on chat completions only. Embed, Rerank, and Voice are different request shapes.
Live platform
Embed, Rerank, Voice, Notre Runtime
More than 60% fewer tokens in validated Notre workloads.
How Caedral works
Your client talks to Caedral. Caedral authenticates, meters, optionally compiles chat context, and routes to a Caedral-operated product or an external lab. You do not hold a stack of provider keys.
01
Your application
SDK, OpenAI-compatible client, n8n, or REST.
02 · Caedral
03
Caedral products + external labs
Embed, Voice, hosted Rerank, and frontier chat models — billed on the pool that model belongs to.
Workloads
Chat, retrieval, speech, and automation share authentication and billing. They do not share a compiler.
POST /v1/chat/completions
Model inference on the OpenAI-compatible chat path. Notre Runtime may compile context on this endpoint when the deployment allows.
Notre Runtime is native here — included, not a SKU.
Open this pathBuilt by Caedral
Frontier labs are in the catalog. These four are developed or hosted by Caedral. Classification is explicit — we do not hide our own work behind someone else’s name.
caedral-embed-e1-small-v1
384-dimensional embeddings. Developed and operated by Caedral. POST /v1/embeddings.
Caedral Models
BAAI/bge-reranker-v2-m3
Hosted by Caedral over an external foundation model. POST /v1/rerank.
External Models
caedral-voice-1
English WAV speech. Developed and operated by Caedral. Notre Runtime does not apply. POST /v1/audio/speech.
Caedral Models
eligible chat
Native context compiler on POST /v1/chat/completions. Not a SKU. MCP is an integration, not the product name.
Included
Catalog
551 models across 86 providers, plus the Caedral-operated products above.
| Model | ID | Context |
|---|---|---|
| Anthropic: Claude Fable Latest | ~anthropic/claude-fable-latest | 1,000,000 |
| Anthropic: Claude Haiku Latest | ~anthropic/claude-haiku-latest | 200,000 |
| Anthropic: Claude Opus Latest | ~anthropic/claude-opus-latest | 1,000,000 |
| Anthropic: Claude Sonnet Latest | ~anthropic/claude-sonnet-latest | 1,000,000 |
| DeepSeek: DeepSeek Flash Latest | ~deepseek/deepseek-flash-latest | 1,048,576 |
| DeepSeek: DeepSeek Pro Latest | ~deepseek/deepseek-pro-latest | 1,048,576 |
| DeepSeek: DeepSeek V4 Flash Latest | ~deepseek/deepseek-v4-flash-latest | 1,310,720 |
Notre Runtime
Notre Runtime is built into Caedral. No add-on. No separate bill. Designed to reduce unnecessary context without degrading the model response. More than 60% fewer tokens in validated Notre workloads.
0.0%reduction
TRUE AGGREGATE on the recorded corpus (n=4,962). Output tokens are generated by the model and are not claimed as Notre savings.
Notre is Caedral's research and development program.Notre Runtime is the production layer on eligible chat. Embed, Voice, and Rerank are different request shapes — we do not pretend they share this compiler.
Notre economics
Inspect recorded compiler benchmarks, then scale the same ratio to a realistic input budget. The model control sets the current Caedral input tariff — Notre was not A/B tested on each catalog model.
Verified Notre benchmark
Same plan. Same price. More useful AI work.
Notre Runtime is built into Caedral. No add-on. No separate bill.Reduction is from the compiler corpus — not a live A/B of the catalog model you pick. Dollars use that model's current Caedral input tariff.
fewer input tokens, strongest verified
across 2 verified workloads
No add-on fee. No separate Notre bill.
Model (Caedral input tariff)
Same budget
0.00×
equivalent input workload
Projected from verified benchmark
Input tokens avoided
0
of 10,000,000 baseline input
Usage value preserved · External Models pool
$0.00
Current Caedral input tariff only. Output is not counted as savings.
More use from the same plan
0→0
equivalent requests on the same input budget
Compiler time (not model round-trip)
p50 12.0 ms · p95 133.7 ms
Pro+ External Models pool · $70/mo
Without Notre-equivalent: 20,289,855 input tokens.
With this reduction: 55,340,826 equivalent input tokens.
Same included allowance, projected from the verified ratio and this model's current input price.
At this benchmark's 63.3% reduction rate, a 10,000,000-token input allowance can process roughly 2.73× the equivalent workload before exhausting the same input budget.
Model economics
Direct integration means a key, a bill, and an SDK per lab. Caedral is one integration, one key, one catalog, and one billing layer — Notre included on eligible chat.
Direct integration
Caedral
Long-running work
Coding assistants, retrieval pipelines, and n8n automations send large context. Notre exists so eligible chat can keep more of that work inside the included pool.
Developers
Point an OpenAI-compatible SDK at https://api.caedral.com/v1, or use the official caedral packages. Automation goes through n8n. Agents can speak MCP locally.
01
Chat, embeddings, rerank, images, audio, video on api.caedral.com.
02
npm install caedral. Official client. OpenAI-compatible clients also work with a base URL change.
03
pip install caedral. Same methods as the TypeScript client.
04
First-party clients in the monorepo for systems languages. Prefer TS/Python when a registry install is enough.
05
Official community node n8n-nodes-caedral, or HTTP Request with Bearer auth. GET /v1/usage before spend.
06
The caedral CLI talks to Notre Runtime through the gateway. There is no separate MCP product and no public MCP hostname.
Console
The dashboard is the same identity as your API key. Pools, Notre analytics, teams, and service health live here — not in a third-party vendor console.
Dashboard preview
Caedral Models
Included pool
External Models
Paid plans
Notre Runtime
Input saved · reduction · value preserved
Control
API keys · team usage · billing portal · status
Pricing
Caedral Models covers hosted embed and voice. External Models covers frontier APIs and hosted rerank. Annual billing saves 20%. Notre Runtime is built into Caedral. No add-on. No separate bill.
| Plan | Price | Included |
|---|---|---|
| Hobby | Free | $200 Caedral Models / mo |
| Pro | $20/mo | $500 Caedral + $20 external |
| Teams | From $40/seat | Pooled usage, owner billing |
Research
Runtime is the customer-facing compiler on chat. Engine is local CPU inference, pre-alpha, and not a hosted API. Notre DB is research and is not a product.
Production/customer output of Notre. Experimental. Native to eligible chat.
PRE-ALPHA local CPU MoE inference. Not a hosted Caedral API.
Sign up, generate a cd_live_ key, then POST /v1/chat/completions to https://api.caedral.com. Pricing covers pools; the docs cover curl, SDKs, and HTTP 402.