Empirical head-to-head benchmark results

Context Without Limits.
Memory Without Loss.

Drop-in OpenAI-compatible API with persistent lattice memory. Crystallize documents once into permanent Lattice Nodes and recall Grounded Facts through a compact scoped viewport. Official-scale results: 96.28% exact (466/484) on RULER MRCR v2 and 54.87% (276/503) on LongBench v2.

51.98%
Mean token cut · official 503
18,031,435 vs 43,271,748 prompt tokens on LongBench v2. Vs uncached Gemini prefill — not a cached-agent bill.
4.53×
Faster on one 852k row
1,375 ms vs 6,227 ms on matched item 66fa208b. One row, not a suite.
54.87%
LongBench v2 official 503
276/503 exact on ICX run 073110. Gemini full-context raw 194/503 (188 API failures).
96.40%
SWE-bench Verified predicted pass
482/500 predicted pass on official Verified — not official test-resolved.
Empirical Head-to-Head Evaluation

Head-to-Head: LLM Alone vs. LLM with Calera ICX

Paired evaluation used Google Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) as the upstream BYOK test model. ICX works with any supported LLM — Gemini was the model we ran for this comparison. Same tasks, same weights — ICX separates permanent lattice storage from the generative viewport so the model receives only the facts it needs. Token cuts below are versus uncached full-context Gemini prefill. They are not a dollar comparison to a prompt-cached Cursor, Claude Code, or Antigravity harness.

Upstream test model Google Gemini 3.5 Flash-Lite · API ID gemini-3.5-flash-lite · BYOK via Google AI API
LLM alone Gemini 3.5 Flash-Lite · full context prefill
LLM + Calera ICX Gemini 3.5 Flash-Lite · scoped viewport

Long-context reasoning

LongBench v2 paired control · n=10 · identical item IDs

Gemini 3.5 Flash-Lite alone
6/10 (60.0%)
2,253,544 prompt tokens
Gemini 3.5 Flash-Lite + ICX
6/10 (60.0%)
162,923 prompt tokens
Same quality · 92.77% fewer tokens vs uncached prefill · n=10 only

Needle retrieval

RULER MRCR · 484 rows · multi-round ordinal disambiguation

Gemini 3.5 Flash-Lite alone
No 484-row dump
Gemini-alone full MRCR file is not in the public archive
Gemini 3.5 Flash-Lite + ICX
96.28% (466/484 exact)
Wilson 95% CI: 94.2%–97.6% · 18 ordinal sibling misses
466/484 exact · miss ledger closes the rate

Official LongBench v2 (503)

ICX run 073110 · Gemini full-context 20260819_104413 · same 503 IDs

Gemini 3.5 Flash-Lite alone
38.57% (194/503 raw)
188 API failures · valid 194/315 (61.59%)
Gemini 3.5 Flash-Lite + ICX
54.87% (276/503 exact)
0 errors · Wilson 50.5–59.2 · 51.98% mean token cut
+16.30 pp raw · ICX finishes all 503 · Gemini overflows 188

Code repository comprehension

852k context · deep AST & cross-file · item 66fa208b

Gemini 3.5 Flash-Lite alone
6,227 ms (hit)
852,861 full prompt tokens
Gemini 3.5 Flash-Lite + ICX
1,375 ms (hit)
31,783 tokens · 4.53× faster
One matched row · 4.53× · 88.81% vs uncached prefill
View full scorecard table
Benchmark & workload Gemini 3.5 Flash-Lite alone Gemini 3.5 Flash-Lite + ICX Advantage
LongBench v2 paired
n=10 matched item IDs
6/10 · 2,253,544 tokens 6/10 · 162,923 tokens 92.77% vs uncached prefill · n=10 only
RULER MRCR
484 rows · eval 9fc0e405
No 484-row Gemini dump in archive 96.28% (466/484) 18 ordinal sibling misses
LongBench v2 official
503 items
38.57% (194/503 raw) · 188 API failures 54.87% (276/503) +16.30 pp raw · 0 ICX errors
SWE-bench Verified
500 · predicted pass
No official test-resolved column 96.40% (482/500) Not official test-resolved
Code repository
852k context
6,227 ms · 852,861 tokens 1,375 ms · 31,783 tokens 4.53× faster · 88.81% token cut
Legal & Compliance
Contract Redlines & MSAs
Track indemnity, liability, and amendment clauses across dozens of contract revisions without ordinal confusion between versions.
Software Engineering
Massive Code Repositories
Traverse AST dependency graphs and cross-file references across monorepos that exceed standard context windows.
Customer Support & CRM
Multi-Turn Ticket Vaults
Keep years of customer dialogue in persistent memory instead of re-uploading full ticket history on every agent turn.
Deep Document Search
Needle-in-a-Haystack Recall
Retrieve exact facts from deep document haystacks without lost-in-the-middle attention degradation.
Architectural Foundation

The Desk-and-Warehouse Principle

Prompt caching discounts re-reading a stable prefix. It does not take those tokens off the model’s working desk. ICX keeps the archive in lattice memory and streams a scoped viewport.

Without ICX
Standard 1M+ monolithic LLM context window
  • Haystack stays on the desk: A cache-read discount (often 10–50% of input) still leaves the prefix in the attention window.
  • New dumps bill at full price: Tool results, grep output, and file reads append every turn. Prefix edits, tool reorder, and naive compaction break the cache.
  • Lost-in-the-middle: Attention still diffuses across the cached haystack. Overflows still fail the call — Gemini-alone missed 188/503 rows in our official run.
Cheaper re-reads • Same haystack • Cache breaks still full-price
Enterprise Grade

Features Built for Enterprise Scale

Everything your engineering, security, and product teams need to deploy infinite context memory with zero compliance or performance friction.

Zero-Trust BYOK Security

Bring Your Own Google Gemini or OpenAI API keys. Keys are forwarded ephemerally via request headers (X-LLM-API-Key) with zero server-side key retention or training on customer data.

Drop-In OpenAI SDK Gateway

100% compatible drop-in replacement for standard OpenAI SDKs in Python, Node.js, Go, and Rust. Simply update baseURL and pass X-Space-ID to begin grounding requests.

Hosted Agent MCP Server

Native 1-click Model Context Protocol (MCP) server integration for Google Antigravity, Claude Desktop, Cursor, and custom autonomous agent swarms.

Explore Hosted MCP Docs →

Isolated Memory Workspaces

Multi-tenant namespace isolation (legal_vault, eng_repo, clinical_ehr) ensures zero cross-tenant factual leakage and instantaneous context switching.

Sub-Millisecond Engine Latency

Discrete simplicial graph diffusion executes on standard CPU memory in <0.5ms, eliminating expensive GPU cluster reservation fees and VRAM memory starvation.

Topological Grounding Invariant

Factual assertions form closed simplicial chains (∂c = 0). Hallucinated or out-of-ontology statements trip the boundary refusal operator instantly.

1-Line Integration

Developer SDK Quickstart

Drop Calera ICX into your existing OpenAI, LangChain, LlamaIndex, or REST workflows in less than 60 seconds.

Want a live comparison before copying code? Open the ICX console preview or run a FinanceSec query — both use the same Calera account, billed per product.

baseURL: "https://icx.api.caleralabs.com/v1"
import requests
from openai import OpenAI

API_KEY = "icx_live_your_api_key"
BASE_URL = "https://icx.api.caleralabs.com/v1"
SPACE_ID = "legal_vault"

# 1. Ingest document into isolated memory space via REST API
with open("contracts/vendor_agreement.pdf", "rb") as f:
    requests.post(
        f"{BASE_URL}/ingest/file",
        headers={"Authorization": f"Bearer {API_KEY}", "X-Space-ID": SPACE_ID},
        files={"file": f}
    )

# 2. Query associative memory via standard OpenAI client drop-in
client = OpenAI(
    api_key=API_KEY,
    base_url=BASE_URL,
    default_headers={"X-Space-ID": SPACE_ID}
)

response = client.chat.completions.create(
    model="calera-icx-v1",
    messages=[
        {"role": "user", "content": "Recall exact contractual indemnity cap across all MSAs."}
    ]
)

print(response.choices[0].message.content)
Get API Key →
TCO Calculator

Modeled token-volume sketch

A formula, not an invoice. It applies a published list price and a published cache-read multiplier. ICX has not measured a live agent-loop bill.

Real-World Scale Presets:
Real-World Benchmark Representation

SaaS Engineering Team (Team Tier — 5,000,000 Lattice Nodes): Equivalent to 100s of product specs, 50 enterprise MSAs, full API documentation, and 6 months of customer support logs (~10,000 source pages / ~1,250,000 Grounded Facts) crystallized in persistent memory.

Assumptions (all modeled): input $2.50 / MTok; Anthropic cache-read 0.10× if the entire archive were a stable prefix; ICX viewport 1,000 tokens/turn + the matching subscription. A prefix above ~1M tokens does not fit in one window, so the cache-read column is algebra, not a feasible agent turn.

Modeled uncached full replay
$187,500 / mo
Archive × $2.50/MTok × monthly turns. Not how a cached 2026 agent is billed.
Modeled 0.10× cache-read
$18,750 / mo
Published Anthropic hit multiplier on the same token volume. Assumes 100% hits and a prefix that large — usually false.
Modeled ICX viewport + subscription
$67 / mo
1,000 tokens/turn at $2.50/MTok plus the matching plan. Viewport size is an assumption, not the 503-row mean.

No savings percentage is shown. This sketch is not a measured invoice and is not a win against a warm cache.

Technical note

The public scorecard

Read Infinite Memory Is Not Infinite Attention (CALERA-PAPER-ICX-2026-08-23). Named-run scores: RULER 466/484, LongBench 276/503 vs Gemini 194/503 raw, SWE-bench Verified 482/500 predicted pass. Row dumps are not published on that page.

Read the technical note
Official suites: n = 484 / 503 / 500
Named run IDs; no unverifiable hashes
SWE-bench labeled predicted pass
Gemini 503 reported with 188 API failures
Dollar TCO kept as a model, not a measurement
We asked Antigravity

What changes in the agent loop when recall is indexed

Calera asked Google Antigravity, with the ICX MCP attached to this repository, to contrast truncated grep with lattice lookup. The result is an operator transcript — not a measurement. Empirical head-to-head numbers stay on the paper above.

Read the session
Instrument: Google Antigravity + ICX MCP
Genre: operator transcript
Not a live telemetry probe
Companion to /mcp-docs

Subscription Tiers

Select a developer preview tier. All subscriptions include single-account management and multi-domain billing on Calera Labs.

How ICX Persistent Memory Capacity Works

Your subscription tier grants persistent memory capacity (Lattice Nodes) in the Volumetric Lattice Network. Documents stay crystallized in your tenant lattice. On each prompt, ICX projects a numbered viewport of Grounded Facts into your LLM. Token savings depend on the workload (LongBench v2 mean noise cut on 503 rows is 51.98%; synthetic haystack gates are higher).

ICX plans ($0 / $29 / $149 / $999) buy memory capacity. FinanceSec plans are a separate SEC-filing product. One Calera account can subscribe to both — they do not share a token pool.

Monthly Billing Annual Billing Save 20%
Developer Preview

Developer Pilot

For individual developers testing persistent memory workflows.

$0/mo
  • 1,000,000 persistent Lattice Nodes (~2,000 source pages)
  • 1 Active API Key limit
  • 1 Isolated Memory Workspace
  • Projects ~500–1,000 Grounded Facts per turn
  • 30 requests/minute burst
  • OpenAI SDK Drop-In
Start Free Pilot
Developer Preview

Team

For engineering teams building memory-augmented AI applications.

$149/mo
  • 500,000,000 persistent Lattice Nodes
  • 10 Active API Keys
  • Shared team memory spaces
  • Multi-tenant namespaces
  • 150 requests/minute burst
  • Email + Slack support
Upgrade to Team
Production

Scale

For production AI workloads and heavy context applications.

$999/mo
  • 5B Lattice Nodes / mo engine allotment
  • Unlimited applications
  • Multi-tenant isolation
  • ~2M API calls / day
  • 99.95% uptime SLA
  • Priority support
Upgrade to Scale
Dedicated Cluster

Enterprise

Self-hosted + unlimited + white-glove. Contact sales.

Custom
  • Unlimited Lattice Nodes & calls
  • Self-hosted license
  • Hardware attestation
  • Dedicated engineering
  • SSO / custom SLA
  • Contact sales — no self-serve Checkout
Contact Sales
Dedicated Infrastructure

Dedicated Enterprise Worker Nodes

Provision isolated, high-throughput Cloud Run worker instances with guaranteed 1,500 req/min burst rates, dedicated CPU/memory specs, and private memory space isolation for enterprise security requirements.

Contact Sales & Engineering View Enterprise Architecture Docs