Notre
Runtime for chat. Engine for local inference.
Notre is Caedral's research and development program: how requests are prepared, optimized, executed, given context, served efficiently, and observed. Notre Runtime is the production, customer-facing output of that program.
Notre Runtime is built into Caedral. No add-on. No separate bill. Customers do not buy Notre separately and do not manage it as a dashboard switch.
Production chat currently accepts the notre field for compatibility and records telemetry. Optimizer apply is environment-gated (canary/local). Fail-open is the contract. Embeddings, rerank, Voice, images, and video never enter this compiler.
Why it exists
Agents, coding tools, and long conversations accumulate history, tools, and retrieval on every turn. Treating every request identically wastes context, money, and model attention.
What it changes
Notre Runtime compiles the smallest useful context before inference on eligible chat. Same Caedral plan. More useful work.Designed to reduce unnecessary context without degrading the model response.
Request path
Large context in. Reduced input to the model.
- 01Large request context
- 02Notre Runtime
- 03Reduced model input
- 04Model inference
- 05Result
- Input before
- 0
- Input sent
- 0
- Saved
- 0
0.0%reduction
TRUE AGGREGATE on the recorded corpus (n=4,962). Output tokens are generated by the model and are not claimed as Notre savings.
Economics
Same plan. More useful model work.
Verified compiler reduction, projected at current Caedral input tariffs. Notre is not a second bill.
Verified Notre benchmark
Same plan. Same price. More useful AI work.
Notre Runtime is built into Caedral. No add-on. No separate bill.Reduction is from the compiler corpus — not a live A/B of the catalog model you pick. Dollars use that model's current Caedral input tariff.
- Up to
- 0.0%
- Median
- 0.0%
- Notre Runtime
- Included
fewer input tokens, strongest verified
across 2 verified workloads
No add-on fee. No separate Notre bill.
Model (Caedral input tariff)
Same budget
0.00×
equivalent input workload
- 01Input before Notre
- 02Notre Runtime compiles
- 03Input sent to the model
- Input before0
- Input sent0
- Input saved0
- Reduction
- 0.0%
- Quality
- Required-set recall 91.9%
Projected from verified benchmark
Input tokens avoided
0
of 10,000,000 baseline input
Usage value preserved · External Models pool
$0.00
Current Caedral input tariff only. Output is not counted as savings.
More use from the same plan
0→0
equivalent requests on the same input budget
Compiler time (not model round-trip)
p50 12.0 ms · p95 133.7 ms
Pro+ External Models pool · $70/mo
Without Notre-equivalent: 20,289,855 input tokens.
With this reduction: 55,340,826 equivalent input tokens.
Same included allowance, projected from the verified ratio and this model's current input price.
At this benchmark's 63.3% reduction rate, a 10,000,000-token input allowance can process roughly 2.73× the equivalent workload before exhausting the same input budget.
Benchmark details
- Workload
- M1.4 development corpus
- Date
- 2026-08-30
- Runs
- 4,962
- Policy
- V1
- Input before (recorded)
- 567,755,348
- Input after (recorded)
- 208,158,685
- Quality methodology
- 4,470 of 4,962 required tool-sets preserved (492 failures). Design constraint is to reduce unnecessary context without degrading the model response. This is not a claim of zero quality loss.
- Source
- caedral-notre-mcp/benchmarks/results/M1.4-GUARDED-V3.json#v1_shadow_baseline.m14_development
- Confidence
- TRUE AGGREGATE SUM/SUM. Holdout not run. Decision at test time: shadow only. Not a production SLA.
- Fallback rate
- 36.4%
Quality
Operating constraint, not a universal proof
Designed to reduce unnecessary context without degrading the model response.
Notre Engine documents a lossless-by-default contract versus llama.cpp under named conditions — that is Engine research, not a Runtime SLA. Runtime quality is an operating objective. Do not read one experimental test as a mathematical guarantee on every request.
On the M1.4 development corpus (V1 policy, 4,962 rows), compiled schema/request context used 63.3% fewer tokens than the uncompiled baseline. Observed recall on that set was 91.9%. That is a measured workload result, not a guarantee for every request. Methodology.
Native integration
How Notre fits Caedral
- Eligible:
POST /v1/chat/completions(stream, non-stream, tools). SDKs, n8n, and Discord chat pass the same path. - Not this compiler: embeddings, rerank, audio/speech, images, video. Different modalities.
- Compatibility:
notre.moderemainsoff | autoon the public contract. That is not a customer product toggle. - More than 60% fewer tokens in validated Notre workloads.
Products
What is public
Runtime is the customer product on eligible chat. Engine is local and pre-alpha. Notre DB stays in research: no server, no API, no plan.
Context compiler and capability runtime for chat, IDE, CLI, and agent workflows. Native to eligible Caedral chat. MCP is one integration, not the product name.
Local CPU Mixture-of-Experts inference for low-memory machines. Pre-alpha. Not a hosted Caedral API. Notre DB is research only and is not offered.