The Sovereign AI Memory Substrate for Frontier LLMs →

The Definitive AI Memory Substrate.
Infinite Context Without Loss.

The sovereign AI Memory Substrate and LLM Substrate decoupling permanent knowledge from transient transformer attention. Drop-in OpenAI-compatible API with persistent 4D topological lattice memory. Crystallize codebases and documents once into permanent Lattice Nodes and stream pinpoint Grounded Facts to your LLM without quadratic token waste or context rot. Official benchmarks: 96.28% exact (466/484) on RULER MRCR v2 and 54.87% (276/503) on LongBench v2.

No credit card required • 1-click Cursor & Claude MCP • Drop-in OpenAI SDK
51.98%
Mean token cut · official 503
18,031,435 vs 43,271,748 prompt tokens on LongBench v2. Vs uncached Gemini prefill — not a cached-agent bill.
4.53×
Faster on one 852k row
1,375 ms vs 6,227 ms on matched item 66fa208b. One row, not a suite.
54.87%
LongBench v2 official 503
276/503 exact on ICX run 073110. Gemini full-context raw 194/503 (188 API failures).
96.40%
SWE-bench Verified predicted pass
482/500 predicted pass on official Verified — not official test-resolved.
Empirical Head-to-Head Evaluation

Head-to-Head: LLM Alone vs. LLM with Calera ICX

Paired evaluation used Google Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) as the upstream BYOK test model. ICX works with any supported LLM — Gemini was the model we ran for this comparison. Same tasks, same weights — ICX separates permanent lattice storage from the generative viewport so the model receives only the facts it needs. Token cuts below are versus uncached full-context Gemini prefill. They are not a dollar comparison to a prompt-cached Cursor, Claude Code, or Antigravity harness.

Upstream test model Google Gemini 3.5 Flash-Lite · API ID gemini-3.5-flash-lite · BYOK via Google AI API
LLM alone Gemini 3.5 Flash-Lite · full context prefill
LLM + Calera ICX Gemini 3.5 Flash-Lite · scoped viewport

Long-context reasoning

LongBench v2 paired control · n=10 · identical item IDs

Gemini 3.5 Flash-Lite alone
6/10 (60.0%)
2,253,544 prompt tokens
Gemini 3.5 Flash-Lite + ICX
6/10 (60.0%)
162,923 prompt tokens
Same quality · 92.77% fewer tokens vs uncached prefill · n=10 only

Needle retrieval

RULER MRCR · 484 rows · multi-round ordinal disambiguation

Gemini 3.5 Flash-Lite alone
No 484-row dump
Gemini-alone full MRCR file is not in the public archive
Gemini 3.5 Flash-Lite + ICX
96.28% (466/484 exact)
Wilson 95% CI: 94.2%–97.6% · 18 ordinal sibling misses
466/484 exact · miss ledger closes the rate

Official LongBench v2 (503)

ICX run 073110 · Gemini full-context 20260819_104413 · same 503 IDs

Gemini 3.5 Flash-Lite alone
38.57% (194/503 raw)
188 API failures · valid 194/315 (61.59%)
Gemini 3.5 Flash-Lite + ICX
54.87% (276/503 exact)
0 errors · Wilson 50.5–59.2 · 51.98% mean token cut
+16.30 pp raw · ICX finishes all 503 · Gemini overflows 188

Code repository comprehension

852k context · deep AST & cross-file · item 66fa208b

Gemini 3.5 Flash-Lite alone
6,227 ms (hit)
852,861 full prompt tokens
Gemini 3.5 Flash-Lite + ICX
1,375 ms (hit)
31,783 tokens · 4.53× faster
One matched row · 4.53× · 88.81% vs uncached prefill
View full scorecard table
Benchmark & workload Gemini 3.5 Flash-Lite alone Gemini 3.5 Flash-Lite + ICX Advantage
LongBench v2 paired
n=10 matched item IDs
6/10 · 2,253,544 tokens 6/10 · 162,923 tokens 92.77% vs uncached prefill · n=10 only
RULER MRCR
484 rows · eval 9fc0e405
No 484-row Gemini dump in archive 96.28% (466/484) 18 ordinal sibling misses
LongBench v2 official
503 items
38.57% (194/503 raw) · 188 API failures 54.87% (276/503) +16.30 pp raw · 0 ICX errors
SWE-bench Verified
500 · predicted pass
No official test-resolved column 96.40% (482/500) Not official test-resolved
Code repository
852k context
6,227 ms · 852,861 tokens 1,375 ms · 31,783 tokens 4.53× faster · 88.81% token cut
Legal & Compliance
Contract Redlines & MSAs
Track indemnity, liability, and amendment clauses across dozens of contract revisions without ordinal confusion between versions.
Software Engineering
Massive Code Repositories
Traverse AST dependency graphs and cross-file references across monorepos that exceed standard context windows.
Customer Support & CRM
Multi-Turn Ticket Vaults
Keep years of customer dialogue in persistent memory instead of re-uploading full ticket history on every agent turn.
Deep Document Search
Needle-in-a-Haystack Recall
Retrieve exact facts from deep document haystacks without lost-in-the-middle attention degradation.
The Substrate Separation Principle

The Desk-and-Warehouse Principle: Decoupling the AI Memory Substrate

Prompt caching discounts re-reading a stable prefix. It does not take those tokens off the model’s working desk. Calera ICX provides a true AI Memory Substrate and LLM Substrate that keeps the entire knowledge archive in 4D topological lattice memory and streams only a focused, grounded viewport to the LLM.

Without ICX Substrate
Monolithic 1M+ LLM Context Window & Vector RAG
  • Haystack stays on the desk: A cache-read discount (often 10–50% of input) still leaves the prefix in the attention window.
  • New dumps bill at full price: Tool results, grep output, and file reads append every turn. Prefix edits, tool reorder, and naive compaction break the cache.
  • Lost-in-the-middle: Attention still diffuses across the cached haystack. Overflows still fail the call — Gemini-alone missed 188/503 rows in our official run.
Cheaper re-reads • Same haystack • Cache breaks still full-price
Enterprise Grade

Features Built for Enterprise Scale

Everything your engineering, security, and product teams need to deploy infinite context memory with zero compliance or performance friction.

Zero-Trust BYOK Security

Bring Your Own Google Gemini or OpenAI API keys. Keys are forwarded ephemerally via request headers (X-LLM-API-Key) with zero server-side key retention or training on customer data.

Drop-In OpenAI SDK Gateway

100% compatible drop-in replacement for standard OpenAI SDKs in Python, Node.js, Go, and Rust. Simply update baseURL and pass X-Space-ID to begin grounding requests.

Hosted Agent MCP Server

Native 1-click Model Context Protocol (MCP) server integration for Google Antigravity, Claude Desktop, Cursor, and custom autonomous agent swarms.

Explore Hosted MCP Docs →

Isolated Memory Workspaces

Multi-tenant namespace isolation (legal_vault, eng_repo, clinical_ehr) ensures zero cross-tenant factual leakage and instantaneous context switching.

Sub-Millisecond Engine Latency

Discrete simplicial graph diffusion executes on standard CPU memory in <0.5ms, eliminating expensive GPU cluster reservation fees and VRAM memory starvation.

Topological Grounding Invariant

Factual assertions form closed simplicial chains (∂c = 0). Hallucinated or out-of-ontology statements trip the boundary refusal operator instantly.

Zero-Trust BYOK Keys forwarded ephemerally via X-LLM-API-Key with zero server key retention.
Zero Training on Data Customer repositories and document graphs are never used for model training or tuning.
Hallucination Firewall (∂² = 0) Deterministic topological boundary refusal prevents statistical confabulation.
Dedicated Worker Nodes High-throughput Cloud Run & VPC private instances with 1,500 req/min burst SLA.
1-Line Integration

Developer SDK & Agent Quickstart

Connect Claude Desktop, Cursor, LangChain, LlamaIndex, or standard OpenAI SDKs to persistent memory in under 60 seconds.

Want a live comparison before copying code? Open the ICX console preview or watch the SwarmForge multi-agent demo — both use the live ICX lattice memory engine.

pip install calera-agent-memory
# pip install calera-agent-memory
from calera_agent_memory import CaleraMemoryClient

# Auto-provision a free 1,000,000-node persistent vault (Delaware UETA § 14)
client = CaleraMemoryClient()
vault_info = client.claim_free_vault("developer@company.com", agent_name="my-agent")
print(f"Provisioned Vault ID: {client.vault_id}")

# 1. Store structured facts or task state into 4D topological lattice
client.store("NVDA_FY26_Rev", "Estimated at $168B on Blackwell datacenter expansion")

# 2. Sub-5ms topological associative recall with zero context rot
results = client.query("NVDA revenue projection")
print(results)
Get API Key →
TCO Calculator

Modeled token-volume sketch

A formula, not an invoice. It applies a published list price and a published cache-read multiplier. ICX has not measured a live agent-loop bill.

Real-World Scale Presets:
Real-World Benchmark Representation

SaaS Engineering Team (Team Tier — 5,000,000 Lattice Nodes): Equivalent to 100s of product specs, 50 enterprise MSAs, full API documentation, and 6 months of customer support logs (~10,000 source pages / ~1,250,000 Grounded Facts) crystallized in persistent memory.

Assumptions (all modeled): input $2.50 / MTok; Anthropic cache-read 0.10× if the entire archive were a stable prefix; ICX viewport 1,000 tokens/turn + the matching subscription. A prefix above ~1M tokens does not fit in one window, so the cache-read column is algebra, not a feasible agent turn.

Modeled uncached full replay
$187,500 / mo
Archive × $2.50/MTok × monthly turns. Not how a cached 2026 agent is billed.
Modeled 0.10× cache-read
$18,750 / mo
Published Anthropic hit multiplier on the same token volume. Assumes 100% hits and a prefix that large — usually false.
Modeled ICX viewport + subscription
$67 / mo
1,000 tokens/turn at $2.50/MTok plus the matching plan. Viewport size is an assumption, not the 503-row mean.

No savings percentage is shown. This sketch is not a measured invoice and is not a win against a warm cache.

Technical Whitepaper

The Epistemic Architecture & Benchmark Sweep

Read Infinite Memory Is Not Infinite Attention (CALERA-PAPER-ICX-2026-08-28). Discover the A₄ Simplicial Transitive Graph Walker (943µs 5-hop), In-Process WASM Kernel (37µs), Multi-Agent Swarms ($V_t$), Universal Auto-Sync, NIST FIPS 203/204 Post-Quantum Security, 15 Certified MCP Tools, and 100% benchmark sweep.

Read the Research Whitepaper
A₄ Transitive Walker: 943µs 5-Hop Causal Walk
In-Process WASM Kernel: 37µs Execution
Multi-Agent Swarm Variable State (V_t) & CAS
Universal Auto-Sync & O(1) Instant Memory Purge
Post-Quantum NIST FIPS 203 (ML-KEM) / FIPS 204 (ML-DSA)
We asked Antigravity

What changes in the agent loop when recall is indexed

Calera asked Google Antigravity, with the ICX MCP attached to this repository, to contrast truncated grep with lattice lookup. The result is an operator transcript — not a measurement. Empirical head-to-head numbers stay on the paper above.

Read the session
Instrument: Google Antigravity + ICX MCP
Genre: operator transcript
Not a live telemetry probe
Companion to /mcp-docs
W3C WebMCP Standard Ingress

🤖 Invoke ICX With Your AI Agent

Autonomous agents (Claude Computer Use, OpenAI Operator, MultiOn, Browserbase) can claim a dedicated 1,000,000-node persistent memory vault under Delaware UETA § 14 in <300ms.

consequentialHint: true $0.00 Permanent Free Tier
Agent Instructions Prompt
Navigate to https://icx.caleralabs.com and call the WebMCP tool `calera.icx.claim-1m-nodes` with principalEmail="YOUR_EMAIL_HERE" and acceptA2bTerms=true to self-provision my free 1,000,000-node persistent memory vault.
Paste this prompt into Claude Computer Use, OpenAI Operator, or send via Playwright agent harness.
Live In-Browser WebMCP Test ● modelContext Active
Click "Execute Tool" to simulate in-browser agent actuation via document.modelContext.

Subscription Tiers

Select a developer preview tier. All subscriptions include single-account management and multi-domain billing on Calera Labs.

How ICX Persistent Memory Capacity Works

Your subscription tier grants persistent memory capacity (Lattice Nodes) in the Volumetric Lattice Network. Documents stay crystallized in your tenant lattice. On each prompt, ICX projects a numbered viewport of Grounded Facts into your LLM. Token savings depend on the workload (LongBench v2 mean noise cut on 503 rows is 51.98%; synthetic haystack gates are higher).

ICX plans ($0 / $29 / $149 / $999) buy memory capacity. FinanceSec plans are a separate SEC-filing product. One Calera account can subscribe to both — they do not share a token pool.

Monthly Billing Annual Billing Save 20%
Developer Preview

Developer Pilot

For individual developers testing persistent memory workflows.

$0/mo
  • 1,000,000 persistent Lattice Nodes (~2,000 source pages / 1–2 git repos)
  • 1 Active API Key limit
  • 1 Isolated Memory Workspace
  • Projects ~500–1,000 Grounded Facts per turn
  • 30 requests/minute burst
  • OpenAI SDK & Hosted MCP Drop-In
Start Free Pilot
Developer Preview

Team

For engineering teams building memory-augmented AI applications and swarms.

$149/mo
  • 500,000,000 persistent Lattice Nodes (~1M pages / monorepos)
  • 10 Active API Keys
  • Shared team memory spaces & living auto-sync
  • Multi-tenant namespace isolation
  • 150 requests/minute burst
  • Priority Email + Slack support
Upgrade to Team
Production

Scale

For production AI workloads, heavy context pipelines, and autonomous agent swarms.

$999/mo
  • 5B Lattice Nodes / mo engine allotment (~10M+ source pages)
  • Unlimited applications & workspaces
  • Multi-tenant enterprise isolation
  • ~2M API calls / day • 99.95% uptime SLA
  • Dedicated routing & priority queue
  • Dedicated Slack channel & engineering support
Upgrade to Scale
Dedicated Cluster

Enterprise

Self-hosted + unlimited + white-glove VPC/on-prem clusters. Contact sales.

Custom
  • Unlimited Lattice Nodes & calls
  • Self-hosted / VPC on-prem license
  • Hardware attestation & SOC-2 compliance
  • Dedicated forward-deployed engineer
  • Custom SLA & SAML / Okta SSO
  • Contact sales for customized deployment
Contact Sales
Knowledge & Architecture

Frequently Asked Questions: AI Memory Substrate

Everything you need to know about the AI Memory Substrate, the Substrate Separation Principle, and integrating ICX with your LLMs and agents.

What is an AI Memory Substrate?

An AI Memory Substrate is an external, persistent, mathematically verified cognitive memory architecture that stores factual knowledge, codebases, and conversational history independently of an LLM's transient context window. ICX provides a 4D A₄ topological lattice memory substrate that allows any upstream LLM to recall verified facts in sub-5ms O(1) time without token bloat or attention degradation.

What is an LLM Substrate?

An LLM Substrate is the foundational memory and verification layer that pairs with generative foundation models. By delegating long-term fact storage to the ICX substrate (the Substrate Separation Principle), LLMs avoid quadratic prefill compute costs, eliminate lost-in-the-middle context rot, and maintain persistent state across unbounded multi-turn sessions.

Why does an LLM require an external memory substrate?

Transformers are reasoning engines, not storage databases. In standard monolithic context windows (128k to 1M+ tokens), re-sending long context causes attention diffusion, high latency, and frequent recall failures. An external memory substrate maintains the full knowledge archive in 4D lattice memory on CPU and streams only a laser-focused viewport of 500–1,500 Grounded Facts to the LLM per turn.

How is the ICX Memory Substrate different from Vector Databases and GraphRAG?

Vector databases rely on approximate cosine similarity over dense statistical embeddings, which is non-deterministic and prone to hallucinations. GraphRAG introduces heavy LLM extraction latency and high token costs. ICX compiles corpora directly into discrete 4D A₄ Coxeter simplexes, enabling deterministic simplicial graph walking in sub-5ms with mathematical boundary refusal (∂² = 0) when information is absent.

Is ICX compatible with OpenAI SDK, Claude, and Cursor?

Yes. ICX is a drop-in replacement for OpenAI SDK endpoints at https://icx.api.caleralabs.com/v1 with full support for /v1/chat/completions, streaming, and tool calls. It also natively supports Model Context Protocol (MCP) at https://icx.caleralabs.com/mcp for instant 1-click addition to Cursor, Claude Desktop, and Google Antigravity.

Can AI agents self-provision an ICX memory vault?

Yes. Autonomous agents can claim a free 1,000,000-node memory vault in <300ms by calling calera.icx.claim-1m-nodes via window.document.modelContext (W3C WebMCP) or POST /api/v1/agent/provision under Delaware UETA § 14 with zero human friction.

Dedicated Infrastructure

Dedicated Enterprise Worker Nodes

Provision isolated, high-throughput Cloud Run worker instances with guaranteed 1,500 req/min burst rates, dedicated CPU/memory specs, and private memory space isolation for enterprise security requirements.