Documentation menu / searchSearch documentation →

Start here

Quick startConnect through MCPOpenCode native memoryOpenCode memory controllerMemory API & local models

Use the service

Memory & compactionMemory lifecycle controllerLocal agents & swarmsCLI referenceHTTP reference

Run a node

ConfigurationOperations & backupsLocal memory & ARM

Evidence

Performance & device targetsFull retrieval reportBenchmark methodologyEvaluation policyMemory benchmark notes

Build with us

Architecture & schemaTechnology & learning mapRepository maintenanceContributingSecurityWebsite & deploymentSearch & agent discoveryPrivate product measurementEngineering references

Project

Cleanup & release planRoadmapLocal AI memory: when instantKV fitsFeaturesChangelog

History

Verification history

Proposals

Distributed memory proposal

Project / single-node · source MVP

Local AI memory: when instantKV fits

Who it helps, what makes it useful and when another memory tool fits better.

instantKV is an open-source local AI memory service. It stores facts and task state outside the model’s context window. After a context reset or restart, your agent can retrieve that state and continue.

Five tools save, recall, search, browse and delete structured memories. Checkpoints preserve the goal and next action before compaction. Your runtime selects what to keep. Storage and retrieval need no model call.

Who should use it

The service fits local LLM apps, coding agents and tasks that continue across sessions. Multiple workers can share project facts while keeping private notes. Developers can add it to an existing runtime through HTTP, CLI, MCP or the Rust core.

The native binary runs beside your model. Docker is optional. Checkpoints contain the essential task state; detailed facts remain in separate records.

What makes it worth using

  • One process to run. The database lives inside the Rust service. No separate vector or graph database is needed.
  • Exact recall. Read the saved record by key. Use revisions to spot changes and avoid overwriting someone else’s update.
  • Simple discovery. The source MVP adds indexed topic/tag/time retrieval and bounded content keywords, with custom metadata. No embeddings are required.
  • Checkpoints built in. The context note and the session’s latest checkpoint commit together. A restore reports changed, missing or inaccessible references.
  • Shared facts, private work. Give workers read-only project knowledge and separate namespaces (storage sections with their own permissions) for notes. Permissions apply to every request.
  • No model calls for memory. Storing and reading records require no LLM or embedding calls. The local-first software has no per-agent or per-request fee; you still use your own hardware and handle backups.

The storage, permissions and checkpoint tools work together in one local process. The host runtime still controls memory selection and model context.

How the alternatives compare

Reviewed against official documentation on 2026-10-07. The last column is our judgment about fit, not a result from testing those products.

ToolWhat it gives youWhen to choose it
Mem0A memory layer with configurable models, embeddings and vector storage; its default flow extracts facts and searches memoriesYou want it to choose facts from conversations and find related memories by meaning
SupermemoryA memory API with automatic learning, semantic retrieval, user profiles, MCP and agent plugins; its local server can use local modelsYou want automatic memory updates and ready agent integrations
Zep / GraphitiMemory built around changing facts and relationships; Graphiti’s setup uses a graph backend and model providersRelationships and time matter to retrieval, and you want graph-based search
LettaA stateful agent SDK with persistent memory, git-backed files and background memory updatesYou want the agent runtime and memory system together
Upstash RedisManaged Redis over HTTP or TCP, durable storage and multi-region replicationYou want someone else to operate the database, or your application needs Redis features
ValkeyA general-purpose key/value server with transactions, persistence and clusteringYou need its broader database features and want to build the agent handoff logic yourself
instantKVBM25 document ranking, exact records, durable checkpoints, shared/private permissions, HTTP/CLI/MCP in one local-first Rust processYou know what the agent should save and want a direct way to store it and resume work

Several alternatives support self-hosting. Mem0 can use local models, depending on its configuration. instantKV combines direct record access with checkpoint tools; it does not replace your agent runtime. Supermemory discontinued Company Brain and Nova in September 2026. Its memory API, MCP and plugins continue.

The evidence so far

The source MVP’s structured-memory benchmark used three fresh 10,000-memory databases on an M4 Pro. Topic query p95 was 0.152–0.173 ms. The largest sampled server RSS was 20.88 MiB. All 30,000 memories were verified after abrupt restarts. These warm local-HTTP results exclude the model and phones.

The full retrieval report uses all 500 LongMemEval-S questions and all ten LoCoMo histories. Full native QA scored 85.20% (426/500), using the GPT-6 Luna model variant. Competitor QA remains incomplete.

These alternatives have not been compared in this suite. Supermemory local is measured separately in the retrieval notes. The results do not establish a speed or cost ranking. The evidence does not establish SOTA or a statistically significant retrieval win. Semantic search, automatic runtime hooks and replicas remain planned.

Run your own node and read the benchmark limits. Agent knowledge and task checkpoints are separate from a model’s inference KV cache.

A separate SciFact comparison measured instantKV BM25 against the recorded Supermemory local control: 81.43% versus 74.80% Recall@10. Full ArguAna scored 76.96% versus 56.40%, with zero query rejections. 688 searches hit work limits; 1,149 used reduced queries. Supermemory was not rerun here. That result applies to the tested corpus and configurations. It does not rank the other products above or measure real-agent memory quality.