Skip to content

PRE-ALPHA

Experimental

Notre Engine

Notre Engine (cne) is a PRE-ALPHA local Mixture-of-Experts inference engine for low-RAM, CPU-only machines. It is built so large sparse models can run with explicit memory budgets, predictable resource behavior, and lossless output by default — not so Caedral can host another chat SKU.

The research objective: close the gap between a huge MoE artifact and ordinary RAM without changing what the model computes. A regime classifier inspects hardware and the GGUF file, then selects streaming, expert caching, speculation, and residency. llama.cpp remains the inference kernel; memory regimes live outside it.

It is MIT-licensed experimental software. It is not on the production POST /v1/chat/completions path, is not a hosted Caedral API, is not billed through subscription pools, and is not generally available. There is no production SLA, latency target, or capacity commitment. Hardware trivia and unpublished token-per-second figures are not product claims.

Notre Runtime prepares context. Notre Engine is a separate local runtime for the weights themselves. They share a research family, not a request path today.

Documentation: /docs/notre-engine.

Integration docs