Engineering notes / 04 October 2026

Memory retrieval.
What the tests show.

We ran the full memory retrieval sets. The engine stayed small. The results also show where we need to improve.

What we measured

instantKV found 95.13% of labelled source sessions in the first ten results on LongMemEval-S. We queried all 500 questions. On LoCoMo, we queried all 1,986 questions across ten histories. Source-turn Recall@10 was 57.66% across the 1,533 questions with positive evidence labels.

These are evidence retrieval scores. They are not final answer accuracy. All five full retrieval files were independently checked with pytrec_eval against official source labels. They had zero retrieval failures and zero truncated queries.

Evidence retrieval · higher is better

Full memory retrieval results

Recall@10: the share of labelled source evidence found in the first ten results.

instantKVSQLite FTS5Supermemory local
LongMemEval-SSource sessions · 500 scored questions
instantKV95.13%
SQLite FTS595.43%
Supermemory localNot complete

All 500 questions. Full local Supermemory run is incomplete. Raw data ↗

LoCoMoSource turns · 1,533 scored questions
instantKV57.66%
SQLite FTS557.19%
Supermemory local57.97%

All ten histories; 1,986 questions queried. Recall scores 1,533 positive-label questions. Raw data ↗

Why the engine stays small

The core stores content, topic, tags, event time and custom metadata. Ordered indexes filter structured memory. BM25 ranks document content. English stemming handles word forms. WAND skips postings that cannot improve the top results.

Search has fixed limits for terms, candidates, index reads and bytes. Long questions can select fewer terms and report query_reduced. A work-limit hit reports truncated. That flag matters: a partial search has no complete top-k guarantee. Records and indexes change in one local redb transaction.

The memory service makes no model calls and needs no embedding model. Your agent decides what to save. It adds retrieved facts to its model context.

What it costs

The measured macOS ARM64 binary is 8.31 MiB. The largest sampled Rust server RSS was 14.34 MiB on LongMemEval-S and 11.25 MiB on LoCoMo. These samples do not measure peak RAM.

A separate three-run workload stored 10,000 records per database. Topic-query p95 was 0.152–0.173 ms. Durable-save p95 was 6.665–6.923 ms. We verified all 30,000 records after abrupt process restarts. This is a Mac test, not a phone, battery or power-loss test.

Where the baselines do better

SQLite FTS5 reached 95.43% LongMemEval-S recall. Local Supermemory reached 57.97% LoCoMo recall. instantKV had the highest observed LoCoMo nDCG@10 at 44.42%. These small differences need repeated runs and confidence intervals before we claim a significant lead.

Secondary BEIR tests show wider lexical wins: SciFact Recall@10 was 81.43% against the recorded local control's 74.80%; ArguAna was 76.96% against 56.40%. NFCorpus remains a loss: 15.31% against 17.08%. ArguAna rejected zero questions, but 688 hit work limits and 1,149 used reduced terms.

Supermemory is the local v0.0.8 direct embedding path. Extraction, query rewriting and reranking are disabled. This comparison does not measure its full hosted product. BEIR controls were recorded separately on 3 October. Parallel timings and different process scopes do not support speed or RAM ratios.

What we do next

Improve paraphrase and zero-overlap recall on separate development data. Test multi-hop links, event order, stale facts, updates, contradictions and abstention. Reduce work-limit hits through general index changes. Keep every change accountable for CPU, RAM, storage and correctness.

Retrieval quality is one part of useful memory. We are also testing a simple path from install to MCP connection, save, restart and recall. The local MCP process must open its data before it offers tools. One process owns each database; concurrent harnesses need one shared HTTP server. The harness guide gives setup steps, and the operations guide covers backup and failure checks. Packaged releases and real harness tests remain planned.

Full native GPT-6 Luna QA is complete: 426/500 correct, or 85.20%, with zero API failures and a 95% interval of 82.0–88.2%. The first parallel runs hit API rate limits; their failures remain recorded. The completed pass uses saved native contexts and a shared limiter. Its official judge rubric uses a different model from the leaderboard. Competitor QA is incomplete.QA settings and raw outputs →

LongMemEval-V2, AMA-Bench, BEAM, more baselines and 100K–10M+ record tests remain incomplete. Native phone bindings and real device tests remain planned. Quality comes first; local operation stays a requirement.

All benchmark results → · Raw verification receipts ↗ · Evaluation policy →