Notre Runtime
ExperimentalCompile the smallest useful context before the model spends tokens.
Notre Runtimeis Caedral's context compiler and capability runtime. It indexes a workspace, retrieves what a task needs, compiles a context package, tracks delivery, and can reach coding agents over MCP, CLI, or the Caedral chat API.
Formerly marketed as Notre MCP. MCP remains an integration method. The product is the runtime — local SQLite state, symbol index, hybrid search, context VM, and estimated input-token accounting — not the transport.
Start here
Hosted chat field first. Local compiler tools need the experimental CLI and MCP binaries from Notre Runtime source.
1. Native chat path
Notre Runtime is included. Production currently accepts the notre field for compatibility; apply is environment-gated.
2. CLI
From Notre Runtime source (package caedral-cli): caedral init, index, compile-context.
3. MCP
Point the host at the caedral-mcp binary (@caedral/mcp). MCP is the transport, not the product name.
Problem
IDE agents and long chat sessions dump repository maps, stale history, and unused tools into the model. You pay for tokens that do not change the answer. A protocol adapter does not solve that; a compiler does.
Product
Notre Runtime plans a capability surface, retrieves from a local index, compiles context, and optionally records whether it intervened. When the Caedral gateway can apply it, the same notre field opts a chat completion into that path. Fail-open: errors never drop the completion.
How it works (customer level)
Local runtime first. Hosted apply is additive and environment-gated.
Index
SQLite-backed project state, tree-sitter symbols, hybrid search.
Compile
Context VM builds the minimum useful package for the task.
Deliver
CLI, MCP tools, or the chat body Caedral sends upstream.
Account
Estimated input tokens before/sent. Public JSON stays small.
Customer value
Verified Notre benchmark
Same plan. Same price. More useful AI work.
Notre Runtime is built into Caedral. No add-on. No separate bill.Reduction is from the compiler corpus — not a live A/B of the catalog model you pick. Dollars use that model's current Caedral input tariff.
- Up to
- 0.0%
- Median
- 0.0%
- Notre Runtime
- Included
fewer input tokens, strongest verified
across 2 verified workloads
No add-on fee. No separate Notre bill.
Model (Caedral input tariff)
Same budget
0.00×
equivalent input workload
- 01Input before Notre
- 02Notre Runtime compiles
- 03Input sent to the model
- Input before0
- Input sent0
- Input saved0
- Reduction
- 0.0%
- Quality
- Required-set recall 91.9%
Projected from verified benchmark
Input tokens avoided
0
of 10,000,000 baseline input
Usage value preserved · External Models pool
$0.00
Current Caedral input tariff only. Output is not counted as savings.
More use from the same plan
0→0
equivalent requests on the same input budget
Compiler time (not model round-trip)
p50 12.0 ms · p95 133.7 ms
Pro+ External Models pool · $70/mo
Without Notre-equivalent: 20,289,855 input tokens.
With this reduction: 55,340,826 equivalent input tokens.
Same included allowance, projected from the verified ratio and this model's current input price.
At this benchmark's 63.3% reduction rate, a 10,000,000-token input allowance can process roughly 2.73× the equivalent workload before exhausting the same input budget.
Benchmark details
- Workload
- M1.4 development corpus
- Date
- 2026-08-30
- Runs
- 4,962
- Policy
- V1
- Input before (recorded)
- 567,755,348
- Input after (recorded)
- 208,158,685
- Quality methodology
- 4,470 of 4,962 required tool-sets preserved (492 failures). Design constraint is to reduce unnecessary context without degrading the model response. This is not a claim of zero quality loss.
- Source
- caedral-notre-mcp/benchmarks/results/M1.4-GUARDED-V3.json#v1_shadow_baseline.m14_development
- Confidence
- TRUE AGGREGATE SUM/SUM. Holdout not run. Decision at test time: shadow only. Not a production SLA.
- Fallback rate
- 36.4%
Use cases
- IDE assistants
- Connect via stdio MCP. Local tools compile repo context; chat still goes through Caedral when the agent calls the API.
- CLI / terminal agents
- caedral index and compile-context produce a package without pasting the tree into every prompt.
- Tool-heavy chat
- Gateway notre.mode=auto may shrink the provider-bound body when the deployment applies the runtime.
- Hosted MCP (local/dev)
- Streamable HTTP exposes caedral_chat, list_models, and notre_contract — not filesystem tools. No public mcp.caedral.com.
Integrations
MCP is one way in
CLI, IDE/agent hosts, Remote MCP (local/dev), and the Caedral chat field all reach the same runtime family.
CLI
caedral CLI
# Experimental CLI from Notre Runtime source (package caedral-cli)
caedral init
caedral index
caedral compile-context --task "fix the auth middleware"Chat API
POST /v1/chat/completions
{
"model": "anthropic/claude-haiku-4.5",
"messages": [{"role": "user", "content": "…"}],
"notre": { "mode": "auto", "telemetry": true }
}MCP (stdio)
Local tools: get_context_for_task, caedral_search, caedral_symbol, caedral_expand, caedral_diff, caedral_repo_map.
.cursor/mcp.json
{
"mcpServers": {
"caedral": {
"type": "stdio",
"command": "caedral-mcp",
"env": {
"CAEDRAL_REPO_ROOT": "/absolute/path/to/your-repo"
}
}
}
}Compatibility
- Chat completions (stream and non-stream, tools where the route supports them).
- Not embeddings, rerank, Voice, images, or video.
- SDKs, n8n, and Discord pass the notre field through.
- Public telemetry: enabled, mode, intervened, fallback_used.
Billing, limits, maturity
- Billed as ordinary chat usage. No Notre line item.
- Input before / sent / saved are Notre estimates, not billed provider tokens.
- Experimental. Production currently strips the field.
- Related: Notre Engine and models.