Skip to content

Notre

Runtime for chat. Engine for local inference.

Notre is Caedral's research and development program: how requests are prepared, optimized, executed, given context, served efficiently, and observed. Notre Runtime is the production, customer-facing output of that program.

Notre Runtime is built into Caedral. No add-on. No separate bill. Customers do not buy Notre separately and do not manage it as a dashboard switch.

Production chat currently accepts the notre field for compatibility and records telemetry. Optimizer apply is environment-gated (canary/local). Fail-open is the contract. Embeddings, rerank, Voice, images, and video never enter this compiler.

Why it exists

Agents, coding tools, and long conversations accumulate history, tools, and retrieval on every turn. Treating every request identically wastes context, money, and model attention.

What it changes

Notre Runtime compiles the smallest useful context before inference on eligible chat. Same Caedral plan. More useful work.Designed to reduce unnecessary context without degrading the model response.

Request path

Large context in. Reduced input to the model.

Verified benchmark — M1.4 development corpus — Notre Runtime
  1. 01Large request context
  2. 02Notre Runtime
  3. 03Reduced model input
  4. 04Model inference
  5. 05Result
Input before
0
Input sent
0
Saved
0

0.0%reduction

TRUE AGGREGATE on the recorded corpus (n=4,962). Output tokens are generated by the model and are not claimed as Notre savings.

Economics

Same plan. More useful model work.

Verified compiler reduction, projected at current Caedral input tariffs. Notre is not a second bill.

Verified Notre benchmark

Same plan. Same price. More useful AI work.

Notre Runtime is built into Caedral. No add-on. No separate bill.Reduction is from the compiler corpus — not a live A/B of the catalog model you pick. Dollars use that model's current Caedral input tariff.

Up to
0.0%

fewer input tokens, strongest verified

Median
0.0%

across 2 verified workloads

Notre Runtime
Included

No add-on fee. No separate Notre bill.

Benchmark

Model (Caedral input tariff)

Economic scale

Projected from verified benchmark reduction. Same ratio, normalized workload.

Same budget

0.00×

equivalent input workload

  1. 01Input before Notre
  2. 02Notre Runtime compiles
  3. 03Input sent to the model
  • Input before0
  • Input sent0
  • Input saved0
Reduction
0.0%
Quality
Required-set recall 91.9%

Projected from verified benchmark

Input tokens avoided

0

of 10,000,000 baseline input

Usage value preserved · External Models pool

$0.00

Current Caedral input tariff only. Output is not counted as savings.

More use from the same plan

0→0

equivalent requests on the same input budget

Compiler time (not model round-trip)

p50 12.0 ms · p95 133.7 ms

Pro+ External Models pool · $70/mo

Without Notre-equivalent: 20,289,855 input tokens.

With this reduction: 55,340,826 equivalent input tokens.

Same included allowance, projected from the verified ratio and this model's current input price.

At this benchmark's 63.3% reduction rate, a 10,000,000-token input allowance can process roughly 2.73× the equivalent workload before exhausting the same input budget.

Benchmark details
Workload
M1.4 development corpus
Date
2026-08-30
Runs
4,962
Policy
V1
Input before (recorded)
567,755,348
Input after (recorded)
208,158,685
Quality methodology
4,470 of 4,962 required tool-sets preserved (492 failures). Design constraint is to reduce unnecessary context without degrading the model response. This is not a claim of zero quality loss.
Source
caedral-notre-mcp/benchmarks/results/M1.4-GUARDED-V3.json#v1_shadow_baseline.m14_development
Confidence
TRUE AGGREGATE SUM/SUM. Holdout not run. Decision at test time: shadow only. Not a production SLA.
Fallback rate
36.4%

Full verified table

Quality

Operating constraint, not a universal proof

Designed to reduce unnecessary context without degrading the model response.

Notre Engine documents a lossless-by-default contract versus llama.cpp under named conditions — that is Engine research, not a Runtime SLA. Runtime quality is an operating objective. Do not read one experimental test as a mathematical guarantee on every request.

On the M1.4 development corpus (V1 policy, 4,962 rows), compiled schema/request context used 63.3% fewer tokens than the uncompiled baseline. Observed recall on that set was 91.9%. That is a measured workload result, not a guarantee for every request. Methodology.

Native integration

How Notre fits Caedral

  • Eligible: POST /v1/chat/completions (stream, non-stream, tools). SDKs, n8n, and Discord chat pass the same path.
  • Not this compiler: embeddings, rerank, audio/speech, images, video. Different modalities.
  • Compatibility: notre.mode remains off | auto on the public contract. That is not a customer product toggle.
  • More than 60% fewer tokens in validated Notre workloads.

Products

What is public

Runtime is the customer product on eligible chat. Engine is local and pre-alpha. Notre DB stays in research: no server, no API, no plan.

  • Notre Runtime
    Experimental

    Customer-facing runtime

    Context compiler and capability runtime for chat, IDE, CLI, and agent workflows. Native to eligible Caedral chat. MCP is one integration, not the product name.

  • Notre Engine
    Experimental

    PRE-ALPHA

    Local CPU Mixture-of-Experts inference for low-memory machines. Pre-alpha. Not a hosted Caedral API. Notre DB is research only and is not offered.