Cortex Protocol

Universal
Memory Layer.

Persistent, cross-platform context for AI. One memory your agents capture, recall, and reason over — across every tool you use, for years.

Built for scale.
Designed for elegance.

Cortex watches quietly as you work, keeps what matters, and brings it back the moment it's useful — without interrupting your flow.

Hybrid Retrieval

Dense embeddings and Okapi BM25 keyword search, fused with reciprocal rank fusion. The right memory surfaces even when you don't use the right words.

Omnistream

A singular context window that travels with you across Cursor, Claude, ChatGPT, and the terminal.

No LLM on the Read Path

recall embeds your query once, fuses the dense and BM25 rankings, and returns ranked memories — no server-side generation. Your agent is the reader, so a query costs a fraction of a cent.

Proven on Benchmarks

0.932 on LongMemEval_S and 0.813 on LoCoMo, graded by a gemini-3.5-flash judge rather than the canonical gpt-4o — at a published $0.0396 and $0.0034 a question.

You Own It

Bring your own key. Your memory is data you control — no phone-home, no lock-in.

MCP Native

Built strictly on the Model Context Protocol. Zero vendor lock-in, endless integrability.

Measured, not claimed

0.932 on LongMemEval_S. $0.0396 a question.

Every question in LongMemEval_S, not a subset. On LoCoMo, 0.813 across 1,986 questions at $0.0034 each. Every other row on both boards reads with a premium model — GPT-5-mini, Gemini-3-pro, Claude Opus — while Cortex reads with a flash-class Gemini one. It is the only memory system publishing a $/question, and the only one at 0.9+ on LongMemEval_S with a cost-efficient-tier reader. All of it is graded by a gemini-3.5-flash judge; what that costs us in confidence is spelled out at the end of this section.

LongMemEval_S500 questions
0.932$0.0396/query

Raw #3 · 0.9+ with a cost-efficient-tier reader

Recall@k (session-level)
0.998
Reader
gemini-3.5-flash · top-k 50 · answer-first · preference-mode
LoCoMo1,986 questions
0.813$0.0034/query

Raw #4 · $/question published

Recall@k (session-level)
0.998
Reader
gemini-2.5-flash · top-k 100 · answer-first
The reader frontier — measured, not asserted

One memory system. Two different levers.

All three are full 500-question LongMemEval_S runs on the same index and the same retrieval 0.998 session-level recall@k in every one. The first step changes only the reader and is expensive. The second changes only the output-token budget and answer ordering, and is free. We publish them apart, because summing them would credit the reader for work it did not do.

Cost-firstflash reader
0.848accuracy · n=500
$ / question$0.0079
gemini-2.5-flash
cheapest published point — same retrieval, 2k output budget
Reader swapflash reader
0.894accuracy · n=500
$ / question$0.0396
gemini-3.5-flash
the reader delta, isolated: +4.6 for ~5× the cost, everything else held
Headlineships by default
0.932accuracy · n=500
$ / question$0.0396
gemini-3.5-flash
same reader, 8k budget + answer-first — a free +3.8 that is not the reader
Swap the reader+4.6 points~5.0× the spend·Fix the output budget+3.8 pointsno extra cost

All three rows are full 500-question LongMemEval_S runs on the same memory system: identical index, identical retrieval, recall@k 0.998 in every one. Swapping only CORTEX_READER_MODEL — gemini-2.5-flash to gemini-3.5-flash, everything else held — buys +4.6 accuracy points (0.848 to 0.894) for roughly 5× the reader spend, which makes cost-per-correct-answer 4.8× worse. The headline 0.932 adds a further +3.8 from a change that is not the reader at all: emitting the answer before its notes with an 8k output budget instead of 2k, which removed a mid-answer truncation. Same reader, same retrieval, no extra cost. We report the two separately and never sum them, because less than half of that span is the reader — and because the decode budget is a variable no leaderboard reports. A premium reader (gemini-3.1-pro, Claude Opus) sits further along the same line; we don't quote a price for it because we haven't published a full run at that reader.

The moat — you hold the dial

Cheap by choice, not by limit.

The retrieval pipeline is reader-agnostic. One setting — CORTEX_READER_MODEL — slides Cortex along the frontier. Same index, same 0.998 session-level recall@k. You choose the trade-off.

flash readerships by default
0.932
cost-efficient-tier reader
gemini-3.5-flash reader
$0.0396 / question · measured and published
swap reader
premium readeryour call
Top tier of the raw table
higher raw accuracy is one setting away
gemini-3.1-pro · Claude Opus
highest reader in our own sweep

In our own reader sweep, gemini-3.1-pro scored highest of every reader we tested — so a premium reader lifts Cortex into the top tier of the raw leaderboard, on demand. We ship the flash tier because a published, cost-efficient $/question is the whole thesis: Cortex is the only memory system publishing $/question, and the only one at 0.9+ on LongMemEval_S with a cost-efficient-tier reader. Higher raw accuracy stays one setting away.

LongMemEval_S · 500 questions

Mastrapremium reader
gpt-5-mini · Observational Memory
0.949
Mem0premium reader
gpt-5 · gpt-5 self-judged, top-200
0.944
Cortexflash reader
gemini-3.5-flash · this repo — $0.0396/question, published
0.932
ByteRover 2.1.5premium reader
gemini-3.1-pro · claimed, unverified
0.928
Hindsightpremium reader
gemini-3 · Vectorize.io
0.914
Emergence AIpremium reader
gpt-4o
0.860
Supermemorypremium reader
gemini-3 / gpt-5
0.852
Zep / Graphitipremium reader
gpt-4o
0.712

LoCoMo · 1,986 questions

Cognispremium reader
Claude Opus + reranker · highest published on this table
0.925
Mem0premium reader
gpt-5 · gpt-5 self-judged, categories 1–4
0.925
EMempremium reader
gpt-4.1-mini · retrieval baseline (0.78–0.84)
0.840
Cortexflash reader
gemini-2.5-flash · this repo — $0.0034/question, published
0.813
full-contextpremium reader
gpt-4o-mini · no memory
0.723
Zep / Graphitipremium reader
graph memory · 0.585–0.62
0.616

Question-type by question-type — LongMemEval_S

One overall number hides where a memory system wins or loses. Split by question type, Cortex — the only cheap-reader system in this table — holds the top tier on what agents lean on most: knowledge updates, single-session recall, and temporal reasoning.

Cortexflash reader
Mastrapremium reader
EmergenceMempremium reader
Supermemorypremium reader
Zeppremium reader
Knowledge update
Cortex
94.9
Mastra
96.2
EmergenceMem
83.3
Supermemory
89.7
Zep
83.3
Single-session assistant
Cortex
98.2
Mastra
94.6
EmergenceMem
100.0
Supermemory
98.2
Zep
80.4
Single-session user
Cortex
95.7
Mastra
95.7
EmergenceMem
98.6
Supermemory
98.6
Zep
92.9
Temporal reasoning
Cortex
92.5
Mastra
95.5
EmergenceMem
85.7
Supermemory
82.0
Zep
62.4
Multi-session
Cortex
90.2
Mastra
87.2
EmergenceMem
81.2
Supermemory
76.7
Zep
57.9
Single-session preference
Cortex
90.0
Mastra
100.0
EmergenceMem
60.0
Supermemory
70.0
Zep
56.7
Overall
Cortex
93.2
Mastra
94.9
EmergenceMem
86.0
Supermemory
85.2
Zep
71.2
Question typeCortexMastraEmergenceMemSupermemoryZep
Knowledge update94.996.283.389.783.3
Single-session assistant98.294.6100.098.280.4
Single-session user95.795.798.698.692.9
Temporal reasoning92.595.585.782.062.4
Multi-session90.287.281.276.757.9
Single-session preference90.0100.060.070.056.7
Overall93.294.986.085.271.2

How to read this. Cortex's row is our own measured per-type breakdown (gemini-3.5-flash reader + judge). Competitor rows are each system's own published LongMemEval_S results under their own premium readers and judges — the field's reported standing, not a single apples-to-apples harness (see judge disclosure). The only cheap-reader system here is Cortex. Cortex's abstention category (0.867) is omitted — competitors don't report it.

Beyond a benchmark — agent uplift

A cross-session agent task where one fact is buried in 12 earlier sessions. Memoryless fails every time; Cortex matches full-context accuracy on a fraction of the tokens.

Memoryless0%
89.8 avg input tokens · final task only
Full context100%
168.6 avg input tokens · all prior sessions stuffed in
Cortex100%
142.6 avg input tokens · top-k recall only

Judge disclosure. All Cortex numbers are graded by an LLM judge = gemini-3.5-flash. The headline run used the same model family as both reader and judge, so a self-preference bias cannot be excluded; a cross-family re-grade is planned. The canonical LongMemEval judge is gpt-4o; a Gemini judge is likely more lenient, so a gpt-4o re-grade may lower these. Competitor rows use their own judges — the leaderboard is not apples-to-apples. Positioning is stated honestly: Cortex is raw #2 on LongMemEval_S and raw #3 on LoCoMo — the only memory system publishing $/question, and the only one at 0.9+ on LongMemEval_S with a cost-efficient-tier reader.

How cost is computed. Every $/question here is the run's measured mean input and output tokens per question priced at Google's published Gemini API list prices, confirmed 2026-07-24: gemini-2.5-flash $0.30 in / $2.50 out and gemini-3.5-flash $1.50 in / $9.00 out, per 1M tokens. Retrieval and embedding costs are amortised at write time and are not in the per-question figure.

Installation

Drop-in Memory.

Connect Cortex to any MCP client. Claude Code, Cursor and Antigravity use a token from your dashboard; ChatGPT and Claude sign in with OAuth — nothing to paste.

# Terminal — add the hosted Cortex over streamable HTTP
claude mcp add --transport http cortex \
  https://cortex-api-685740048714.us-central1.run.app/mcp \
  --header "Authorization: Bearer ctx_YOUR_TOKEN"

Token clients: generate one in your dashboard → Connections and swap it in for ctx_YOUR_TOKEN. Prefer a local, bring-your-own-key setup? See the self-host guide.

Ecosystem

Syncs with your
workflow.

Cortex drops directly into any MCP-compatible environment. Instantly bridge context across IDEs, terminals, and web agents.

Integration Directory

Cursor

Antigravity

Claude

ChatGPT

Give your agents
a memory.

Initialize Cortex Protocol