Recall@10 is the share of labelled source evidence found in the first ten results. Query time excludes model inference. RAM is sampled process RSS.
Compare engines
One benchmark at a time. The same source data and questions.
LongMemEval-S
Find the source conversation session behind a question.
Source evidence found in the first ten results. Higher is better.
instantKV
95.13%
SQLite FTS5
95.43%
Supermemory local
No complete run
0100%
LongMemEval-S: all published providers with recall, sampled RAM, p95 response time and failures
Engine
Recall@10
Sampled RAM
Query p95
Failed / truncated
instantKVRust server · BM25
95.13%
14.34 MiB
5.65 ms
0 / 0
SQLite FTS5In-process · lexical
95.43%Highest observed
237.69 MiB
1.73 ms
0 / 0
Supermemory localIncomplete
N/A
N/A
N/A
N/A
Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.
All metrics, scope and raw data
All 500 questions. Full local Supermemory run is incomplete.
Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.
LoCoMo
Find the source turns across ten long conversation histories.
Source evidence found in the first ten results. Higher is better.
instantKV
57.66%
SQLite FTS5
57.19%
Supermemory local
57.97%
0100%
LoCoMo: all published providers with recall, sampled RAM, p95 response time and failures
Engine
Recall@10
Sampled RAM
Query p95
Failed / truncated
instantKVRust server · BM25
57.66%
11.25 MiB
0.79 ms
0 / 0
SQLite FTS5In-process · lexical
57.19%
174.63 MiB
1.71 ms
0 / 0
Supermemory localLocal v0.0.8 · embeddings
57.97%Highest observed
2547.42 MiB
25.12 ms
0 / 0
Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.
All metrics, scope and raw data
All ten histories; 1,986 questions queried. Recall scores 1,533 positive-label questions.
Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.
SciFact
Find scientific documents that support a claim.
Source evidence found in the first ten results. Higher is better.
instantKV
81.43%
SQLite FTS5
No complete run
Supermemory local
74.80%
0100%
SciFact: all published providers with recall, sampled RAM, p95 response time and failures
Engine
Recall@10
Sampled RAM
Query p95
Failed / truncated
instantKVRust server · BM25
81.43%Highest observed
23.14 MiB
1.55 ms
0 / 0
SQLite FTS5Not run
N/A
N/A
N/A
N/A
Supermemory localLocal v0.0.8 · embeddings
74.80%
2606.55 MiB
53.82 ms
0 / 0
Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.
All metrics, scope and raw data
All 300 queries. Zero native truncations. 5,183 source documents.
instantKV
Ranking · nDCG@10
68.52%
Query p50 / p95 / p99
0.86 / 1.55 / 1.75 ms
Write p95
12.57 ms
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
81.00 MiB
One full corpus database
Write throughput
Not recorded
Binary size
8.31 MiB
Query samples / passes
300 / 1
RAM scope: Native Rust server. Current native run · loopback HTTP.
Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.
ArguAna
Find a relevant counterargument to a long argument.
Source evidence found in the first ten results. Higher is better.
instantKV
76.96%
SQLite FTS5
No complete run
Supermemory local
56.40%
0100%
ArguAna: all published providers with recall, sampled RAM, p95 response time and failures
Engine
Recall@10
Sampled RAM
Query p95
Failed / truncated
instantKVRust server · BM25
76.96%Highest observed
24.78 MiB
46.83 ms
0 / 6881,149 reduced
SQLite FTS5Not run
N/A
N/A
N/A
N/A
Supermemory localLocal v0.0.8 · embeddings
56.40%
3874.09 MiB
332.66 ms
0 / 0
Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.
All metrics, scope and raw data
688 of 1,406 native queries truncated; 1,149 reduced. Zero rejections. 8,674 source documents.
instantKV
Ranking · nDCG@10
47.52%
Query p50 / p95 / p99
21.17 / 46.83 / 49.30 ms
Write p95
Not recorded
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
128.50 MiB
One full corpus database
Write throughput
101.76 records/s
Binary size
8.31 MiB
Query samples / passes
1,406 / 1
RAM scope: Native Rust server. Current native run · loopback HTTP.
Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.
NFCorpus
Find medical documents for natural-language questions.
Source evidence found in the first ten results. Higher is better.
instantKV
15.31%
SQLite FTS5
No complete run
Supermemory local
17.08%
0100%
NFCorpus: all published providers with recall, sampled RAM, p95 response time and failures
Engine
Recall@10
Sampled RAM
Query p95
Failed / truncated
instantKVRust server · BM25
15.31%
22.92 MiB
0.93 ms
0 / 0
SQLite FTS5Not run
N/A
N/A
N/A
N/A
Supermemory localLocal v0.0.8 · embeddings
17.08%Highest observed
3161.84 MiB
39.76 ms
0 / 0
Different process scopes and transports. These resource figures do not establish matched speed or RAM ratios. Highest observed scores are not significance claims.
All metrics, scope and raw data
All 323 queries. Zero native truncations. 3,633 source documents.
instantKV
Ranking · nDCG@10
32.34%
Query p50 / p95 / p99
0.54 / 0.93 / 1.19 ms
Write p95
Not recorded
Record request; provider commit guarantees differ
Startup p95
Not recorded
Database size
56.63 MiB
One full corpus database
Write throughput
104.39 records/s
Binary size
8.31 MiB
Query samples / passes
323 / 1
RAM scope: Native Rust server. Current native run · loopback HTTP.
Write guarantees differ. Write throughput excludes query work. CPU utilization, energy, separate index bytes and concurrent query throughput were not recorded.
Answer quality
A separate test: can a model answer correctly from the retrieved memory?
LongMemEval-S · instantKV + GPT-6 Luna85.20%
Answer accuracy · 426 / 500 correct
95% confidence interval
82.00–88.20%
API failures
0
Evaluation coverage
500 questions · one full pass
This score includes model inference. Reader and judge: GPT-6 Luna, medium reasoning, 16,384-token context limit. Official LongMemEval rubric with a different model; not official leaderboard parity. Competitor QA is incomplete. The model API is used for evaluation; instantKV retrieval stays local.
Accuracy by task and model response time
All questions, including abstention
Task
Questions
Accuracy
User facts
64
95.31%
Time reasoning
127
82.68%
Assistant facts
56
100.00%
Updated facts
72
91.67%
Preferences
30
66.67%
Multiple sessions
121
75.21%
Abstention
30
90.00%
Cloud reader response: p50 2.55 s · p95 6.01 s · p99 8.52 s. Excludes retrieval, rate-limit waiting and judge calls. This is not local-model performance. Confidence interval: source-history-cluster bootstrap, 2,000 resamples.
Separate structured-memory test. Three databases with 10,000 records each.
Recorded Rust binary
8.31 MiB
Durable save · p95
6.90 ms
Topic recall · p95
0.161 ms
Recovered after restart
30,000 / 30,000
512-byte content plus metadata. Warm loopback HTTP, one request at a time, immediate durable commits. No embedding model or cloud API is required by the memory engine.
All operation timings and recovery notes
Milliseconds · median of three run percentiles
Operation
p50
p95
p99
p95 range
RememberNew memory + indexes; immediate commit
5.697
6.900
7.823
6.665–6.923
Recall by topicUp to 10 memories; 16 KiB response budget
0.128
0.161
0.197
0.152–0.173
Recall by tagUp to 10 memories; 16 KiB response budget
0.128
0.156
0.200
0.152–0.157
Recall by event timeUp to 10 memories; 16 KiB response budget
0.127
0.155
0.216
0.154–0.167
Topic + tag + keywordsUp to 10 memories; 16 KiB response budget
0.129
0.158
0.171
0.147–0.159
BrowseUp to 10 memories; 16 KiB response budget
0.166
0.384
0.567
0.360–0.419
Sparse keyword / first page1,000 scanned candidates; no match on this page
1.403
1.694
1.826
1.660–1.714
Read exact keyOne complete memory and revision
0.087
0.112
0.150
0.108–0.118
ForgetRevision-checked delete + index removal
5.812
7.420
7.873
6.844–7.741
Idle RSS 6.34–6.36 MiB. All 30,000 records recovered after abrupt process restarts. This was not a device power-loss test.
What was measured, what differs, and what is still pending.
Hardware & build
Apple M4 Pro, 24 GiB RAM, macOS 27.0 ARM64. Retrieval engine frozen at fafa202. No engine tuning during the full runs.
Quality scores
LongMemEval-S scores source sessions; LoCoMo scores source turns. BEIR scores source documents. Recall is not answer accuracy. Small score gaps do not establish a significant win.
RAM boundaries
instantKV: Rust server only. SQLite: Python runner and in-process adapter. Supermemory: server and descendants, including embeddings; memory-suite corpus scopes accumulate. Sampled RSS is not peak RAM.
Response speed
Memory suites ran under shared host load. BEIR controls are recorded same-host runs. Timings do not support isolated speed ratios. CPU utilization, CPU time and energy were not recorded.
Supermemory control
Local v0.0.8, direct embedding retrieval. Model extraction, query rewriting and reranking are disabled. The hosted Supermemory product is not measured.
Reproduction
Same source data, lossless chunks and questions. SQLite and instantKV use matching response pagination bounds. Published full-memory rankings were checked with pytrec_eval. Read the full run notes ↗
Pending benchmarks and device tests
LongMemEval-V2 Small / Medium, AMA-Bench, BEAM and 100K / 1M / 10M+ record tests are incomplete. Mem0, Zep and independent dense-vector controls are not complete. Competitor end-to-end QA, phone latency, battery and native bindings remain pending. No score or resource estimate is assigned to an unfinished run.