Measured results · 04 October 2026

Benchmarks

Inspect retrieval quality, model answers and storage measurements. Each result keeps its workload and comparison limits.

Apple M4 Pro24 GiB RAMmacOS 27 · ARM64Run conditions ↗

instantKV at a glance

Five complete retrieval sets. Select a benchmark to compare engines.

instantKV retrieval quality, sampled RAM, query p95 and work limits for each completed benchmark
BenchmarkRecall@10 Higher is betterSampled RAM Lower is betterQuery p95 Lower is betterQuery limits
LongMemEval-S500 scored · source sessions95.13%14.34MiB5.65msNo limits hit0 failures
LoCoMo1,533 scored · source turns57.66%11.25MiB0.79msNo limits hit0 failures
SciFact300 scored · beir documents81.43%23.14MiB1.55msNo limits hit0 failures
ArguAna1,406 scored · beir documents76.96%24.78MiB46.83ms688 truncated0 failures · 1,149 reduced
NFCorpus323 scored · beir documents15.31%22.92MiB0.93msNo limits hit0 failures

Recall@10 is the share of labelled source evidence found in the first ten results. Query time excludes model inference. RAM is sampled process RSS.

Compare engines

One benchmark at a time. The same source data and questions.

LongMemEval-S

Find the source conversation session behind a question.

Source evidence found in the first ten results. Higher is better.

instantKV95.13%
SQLite FTS595.43%
Supermemory localNo complete run
LongMemEval-S: all published providers with recall, sampled RAM, p95 response time and failures
EngineRecall@10Sampled RAMQuery p95Failed / truncated
instantKVRust server · BM2595.13%14.34 MiB5.65 ms0 / 0
SQLite FTS5In-process · lexical95.43%Highest observed237.69 MiB1.73 ms0 / 0
Supermemory localIncompleteN/AN/AN/AN/A

Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.

All metrics, scope and raw data

All 500 questions. Full local Supermemory run is incomplete.

instantKV

Ranking · nDCG@10
89.43%
Query p50 / p95 / p99
4.40 / 5.65 / 6.64 ms
Write p95
13.87 ms
Individual record write request
Startup p95
73.40 ms
Database size
8.04 MiB
Largest individual corpus database
Write throughput
102.92 records/s
Binary size
8.31 MiB
Query samples / passes
1,500 / 3

RAM scope: Native Rust server; fresh database per corpus. Loopback HTTP · shared host load.

Raw results for instantKV ↗

SQLite FTS5

Ranking · nDCG@10
90.20%
Query p50 / p95 / p99
0.95 / 1.73 / 2.29 ms
Write p95
0.77 ms
Individual record write request
Startup p95
0.77 ms
Database size
4.94 MiB
Largest individual corpus database
Write throughput
3969.22 records/s
Binary size
Not recorded
Query samples / passes
500 / 1

RAM scope: In-process Python runner + SQLite; fresh database per corpus. In-process · shared host load.

Raw results for SQLite FTS5 ↗

Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.

LoCoMo

Find the source turns across ten long conversation histories.

Source evidence found in the first ten results. Higher is better.

instantKV57.66%
SQLite FTS557.19%
Supermemory local57.97%
LoCoMo: all published providers with recall, sampled RAM, p95 response time and failures
EngineRecall@10Sampled RAMQuery p95Failed / truncated
instantKVRust server · BM2557.66%11.25 MiB0.79 ms0 / 0
SQLite FTS5In-process · lexical57.19%174.63 MiB1.71 ms0 / 0
Supermemory localLocal v0.0.8 · embeddings57.97%Highest observed2547.42 MiB25.12 ms0 / 0

Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.

All metrics, scope and raw data

All ten histories; 1,986 questions queried. Recall scores 1,533 positive-label questions.

instantKV

Ranking · nDCG@10
44.42%
Query p50 / p95 / p99
0.68 / 0.79 / 0.95 ms
Write p95
23.19 ms
Individual record write request
Startup p95
158.52 ms
Database size
3.75 MiB
Largest individual corpus database
Write throughput
88.35 records/s
Binary size
8.31 MiB
Query samples / passes
1,986 / 1

RAM scope: Native Rust server; fresh database per corpus. Loopback HTTP · shared host load.

Raw results for instantKV ↗

SQLite FTS5

Ranking · nDCG@10
43.93%
Query p50 / p95 / p99
1.28 / 1.71 / 2.18 ms
Write p95
8.65 ms
Individual record write request
Startup p95
21.06 ms
Database size
4.29 MiB
Largest individual corpus database
Write throughput
437.42 records/s
Binary size
Not recorded
Query samples / passes
1,986 / 1

RAM scope: In-process Python runner + SQLite; fresh database per corpus. In-process · shared host load.

Raw results for SQLite FTS5 ↗

Supermemory local

Ranking · nDCG@10
44.08%
Query p50 / p95 / p99
21.16 / 25.12 / 28.49 ms
Write p95
57.00 ms
Individual record write request
Startup p95
4109.55 ms
Database size
392.02 MiB
Cumulative suite database
Write throughput
25.60 records/s
Binary size
Not recorded
Query samples / passes
1,986 / 1

RAM scope: Server + descendants; cumulative corpus scopes in one suite server. Loopback HTTP · shared host load.

Raw results for Supermemory local ↗

Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.

SciFact

Find scientific documents that support a claim.

Source evidence found in the first ten results. Higher is better.

instantKV81.43%
SQLite FTS5No complete run
Supermemory local74.80%
SciFact: all published providers with recall, sampled RAM, p95 response time and failures
EngineRecall@10Sampled RAMQuery p95Failed / truncated
instantKVRust server · BM2581.43%Highest observed23.14 MiB1.55 ms0 / 0
SQLite FTS5Not runN/AN/AN/AN/A
Supermemory localLocal v0.0.8 · embeddings74.80%2606.55 MiB53.82 ms0 / 0

Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.

All metrics, scope and raw data

All 300 queries. Zero native truncations. 5,183 source documents.

instantKV

Ranking · nDCG@10
68.52%
Query p50 / p95 / p99
0.86 / 1.55 / 1.75 ms
Write p95
12.57 ms
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
81.00 MiB
One full corpus database
Write throughput
Not recorded
Binary size
8.31 MiB
Query samples / passes
300 / 1

RAM scope: Native Rust server. Current native run · loopback HTTP.

Raw results for instantKV ↗

Supermemory local

Ranking · nDCG@10
63.24%
Query p50 / p95 / p99
43.89 / 53.82 / 56.94 ms
Write p95
354.49 ms
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
Not recorded
One full corpus database
Write throughput
Not recorded
Binary size
258.20 MiB
Query samples / passes
300 / 1

RAM scope: Server + descendants, embedding runtime included. Recorded 3 October control · loopback HTTP.

Raw results for Supermemory local ↗

Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.

ArguAna

Find a relevant counterargument to a long argument.

Source evidence found in the first ten results. Higher is better.

instantKV76.96%
SQLite FTS5No complete run
Supermemory local56.40%
ArguAna: all published providers with recall, sampled RAM, p95 response time and failures
EngineRecall@10Sampled RAMQuery p95Failed / truncated
instantKVRust server · BM2576.96%Highest observed24.78 MiB46.83 ms0 / 6881,149 reduced
SQLite FTS5Not runN/AN/AN/AN/A
Supermemory localLocal v0.0.8 · embeddings56.40%3874.09 MiB332.66 ms0 / 0

Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.

All metrics, scope and raw data

688 of 1,406 native queries truncated; 1,149 reduced. Zero rejections. 8,674 source documents.

instantKV

Ranking · nDCG@10
47.52%
Query p50 / p95 / p99
21.17 / 46.83 / 49.30 ms
Write p95
Not recorded
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
128.50 MiB
One full corpus database
Write throughput
101.76 records/s
Binary size
8.31 MiB
Query samples / passes
1,406 / 1

RAM scope: Native Rust server. Current native run · loopback HTTP.

Raw results for instantKV ↗

Supermemory local

Ranking · nDCG@10
41.05%
Query p50 / p95 / p99
167.27 / 332.66 / 403.07 ms
Write p95
Not recorded
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
Not recorded
One full corpus database
Write throughput
4.51 records/s
Binary size
258.20 MiB
Query samples / passes
1,406 / 1

RAM scope: Server + descendants, embedding runtime included. Recorded 3 October control · loopback HTTP.

Raw results for Supermemory local ↗

Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.

NFCorpus

Find medical documents for natural-language questions.

Source evidence found in the first ten results. Higher is better.

instantKV15.31%
SQLite FTS5No complete run
Supermemory local17.08%
NFCorpus: all published providers with recall, sampled RAM, p95 response time and failures
EngineRecall@10Sampled RAMQuery p95Failed / truncated
instantKVRust server · BM2515.31%22.92 MiB0.93 ms0 / 0
SQLite FTS5Not runN/AN/AN/AN/A
Supermemory localLocal v0.0.8 · embeddings17.08%Highest observed3161.84 MiB39.76 ms0 / 0

Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.

All metrics, scope and raw data

All 323 queries. Zero native truncations. 3,633 source documents.

instantKV

Ranking · nDCG@10
32.34%
Query p50 / p95 / p99
0.54 / 0.93 / 1.19 ms
Write p95
Not recorded
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
56.63 MiB
One full corpus database
Write throughput
104.39 records/s
Binary size
8.31 MiB
Query samples / passes
323 / 1

RAM scope: Native Rust server. Current native run · loopback HTTP.

Raw results for instantKV ↗

Supermemory local

Ranking · nDCG@10
35.97%
Query p50 / p95 / p99
33.55 / 39.76 / 43.40 ms
Write p95
Not recorded
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
Not recorded
One full corpus database
Write throughput
2.77 records/s
Binary size
258.20 MiB
Query samples / passes
323 / 1

RAM scope: Server + descendants, embedding runtime included. Recorded 3 October control · loopback HTTP.

Raw results for Supermemory local ↗

Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.

Answer quality

A separate test: can a model answer correctly from the retrieved memory?

LongMemEval-S · instantKV + GPT-6 Luna85.20%

Answer accuracy · 426 / 500 correct

95% confidence interval
82.00–88.20%
API failures
0
Evaluation coverage
500 questions · one full pass

This score includes model inference. Reader and judge: GPT-6 Luna, medium reasoning, 16,384-token context limit. Official LongMemEval rubric with a different model; not official leaderboard parity. Competitor QA is incomplete. The model API is used for evaluation; instantKV retrieval stays local.

Accuracy by task and model response time
All questions, including abstention
TaskQuestionsAccuracy
User facts6495.31%
Time reasoning12782.68%
Assistant facts56100.00%
Updated facts7291.67%
Preferences3066.67%
Multiple sessions12175.21%
Abstention3090.00%

Cloud reader response: p50 2.55 s · p95 6.01 s · p99 8.52 s. Excludes retrieval, rate-limit waiting and judge calls. This is not local-model performance. Confidence interval: source-history-cluster bootstrap, 2,000 resamples.

Storage & speed

Separate structured-memory test. Three databases with 10,000 records each.

Recorded Rust binary
8.31 MiB
Durable save · p95
6.90 ms
Topic recall · p95
0.161 ms
Recovered after restart
30,000 / 30,000

512-byte content plus metadata. Warm loopback HTTP, one request at a time, immediate durable commits. No embedding model or cloud API is required by the memory engine.

All operation timings and recovery notes
Milliseconds · median of three run percentiles
Operationp50p95p99p95 range
RememberNew memory + indexes; immediate commit5.6976.9007.8236.665–6.923
Recall by topicUp to 10 memories; 16 KiB response budget0.1280.1610.1970.152–0.173
Recall by tagUp to 10 memories; 16 KiB response budget0.1280.1560.2000.152–0.157
Recall by event timeUp to 10 memories; 16 KiB response budget0.1270.1550.2160.154–0.167
Topic + tag + keywordsUp to 10 memories; 16 KiB response budget0.1290.1580.1710.147–0.159
BrowseUp to 10 memories; 16 KiB response budget0.1660.3840.5670.360–0.419
Sparse keyword / first page1,000 scanned candidates; no match on this page1.4031.6941.8261.660–1.714
Read exact keyOne complete memory and revision0.0870.1120.1500.108–0.118
ForgetRevision-checked delete + index removal5.8127.4207.8736.844–7.741

Idle RSS 6.34–6.36 MiB. All 30,000 records recovered after abrupt process restarts. This was not a device power-loss test.

Raw structured-memory runs ↗

Method & limits

What was measured, what differs, and what is still pending.

Hardware & build
Apple M4 Pro, 24 GiB RAM, macOS 27.0 ARM64. Retrieval engine frozen at fafa202. No engine tuning during the full runs.
Quality scores
LongMemEval-S scores source sessions; LoCoMo scores source turns. BEIR scores source documents. Recall is not answer accuracy. Small score gaps do not establish a significant win.
RAM boundaries
instantKV: Rust server only. SQLite: Python runner and in-process adapter. Supermemory: server and descendants, including embeddings; memory-suite corpus scopes accumulate. Sampled RSS is not peak RAM.
Response speed
Memory suites ran under shared host load. BEIR controls are recorded same-host runs. Timings do not support isolated speed ratios. CPU utilization, CPU time and energy were not recorded.
Supermemory control
Local v0.0.8, direct embedding retrieval. Model extraction, query rewriting and reranking are disabled. The hosted Supermemory product is not measured.
Reproduction
Same source data, lossless chunks and questions. SQLite and instantKV use matching response pagination bounds. Published full-memory rankings were checked with pytrec_eval. Read the full run notes ↗
Pending benchmarks and device tests

LongMemEval-V2 Small / Medium, AMA-Bench, BEAM and 100K / 1M / 10M+ record tests are incomplete. Mem0, Zep and independent dense-vector controls are not complete. Competitor end-to-end QA, phone latency, battery and native bindings remain pending. No score or resource estimate is assigned to an unfinished run.