Documentation menu / search
Search documentation →Start here
Quick startConnect through MCPOpenCode native memoryOpenCode memory controllerMemory API & local modelsUse the service
Memory & compactionMemory lifecycle controllerLocal agents & swarmsCLI referenceHTTP referenceRun a node
ConfigurationOperations & backupsLocal memory & ARMEvidence
Performance & device targetsFull retrieval reportBenchmark methodologyEvaluation policyMemory benchmark notesBuild with us
Architecture & schemaTechnology & learning mapRepository maintenanceContributingSecurityWebsite & deploymentSearch & agent discoveryPrivate product measurementEngineering referencesProject
Cleanup & release planRoadmapLocal AI memory: when instantKV fitsFeaturesChangelogHistory
Verification historyProposals
Distributed memory proposalStart here / single-node · source MVP
Memory API & local models
Remember, recall, search, browse and forget. Find memories by topic, tag, time or keywords. Try the local-model example.
Status: source MVP, unreleased. Updated: 2026-10-07.
Build from this checkout to use these commands. Earlier release archives do not contain them.
instantKV stores facts, preferences, decisions and observations beside your model. Queries can rank content by relevance or select a topic, tag, event time and literal keywords. The memory engine requires no model or embedding calls. Your runtime selects what to save and adds retrieved facts to the model context.
Five everyday tools
| CLI / MCP tool | Use |
|---|---|
remember | Save content, topic, tags, optional event time and custom metadata |
recall | Find memories using any combination of supported filters |
search | Rank content with BM25, English stemming and optional label/time filters |
browse | Page through permitted memories, newest event time first |
forget | Explicitly delete one memory and its indexes |
The seven original KV and checkpoint MCP tools remain available.
Structured memory uses a durable JSON namespace, such as knowledge.
RAM scratch remains a separate raw KV cache.
The get and list operations still provide access to ordinary records.
Ordinary records do not become structured memories automatically.
Try it locally
Run these commands from the repository. You need Rust 1.98 or later.
Start the server, then run the memory commands from its directory in another terminal.
cargo install --path crates/instantkv --locked
instantkv start --dir my-local-memory
In another terminal:
cd my-local-memory
instantkv remember "Prefer Rust for local tools" \
--key preferences/language --topic preferences --tag local \
--metadata '{"source":"user","app":{"confidence":0.9}}'
instantkv recall --topic preferences --query Rust
instantkv search "preferred language for local tooling"
instantkv browse --limit 10
remember returns a key and revision. If you omit --key, the server generates a UUID key.
A stable key lets you inspect the result after an uncertain save.
A repeated creation request fails if that key already exists.
Before you retry or update, read the saved value and revision:
instantkv get knowledge preferences/language
instantkv remember "Prefer Rust for embedded tools" \
--key preferences/language --topic preferences --if-revision 1
instantkv forget preferences/language --if-revision 2
The revisions above are examples. Use the revisions from your responses. Other writes in the same namespace can advance its revision counter.
Topics, tags and time
Topics and tags are exact labels. The server removes outer spaces and converts the labels to lowercase. Each memory can have one topic and up to eight tags. The server removes duplicate normalized tags.
| Field | Limit |
|---|---|
| Topic or tag | 1–64 UTF-8 bytes; no control characters |
| Structured-memory key | 256 bytes, even if the raw KV key limit is larger |
| Content | 16 KiB |
| Custom metadata | 8 KiB JSON object |
The namespace value limit applies to the complete stored JSON envelope.
Each save uses the current server time as its default event time.
written_at_ms records when the server saved that version.
occurred_at_ms records when the event happened.
An update without an explicit event time uses the current server time.
For an older event, supply --occurred-at-ms:
instantkv remember "Chose redb for the local app" \
--key decisions/storage --topic decisions --tag local \
--occurred-at-ms 1790985600000
instantkv recall --topic decisions --tag local \
--since-ms 1790985600000 --until-ms 1791072000000
Times use Unix milliseconds. Both time-range endpoints are inclusive. The host runtime must convert phrases such as “last week” with the user’s timezone. Results have descending event-time order. Equal times have descending key order.
Ranked retrieval
Use search when you need relevant documents, rather than newest events:
instantkv search "preferred language for local tooling" --limit 10
instantkv search "storage decisions" --topic decisions --tag local
The engine splits content into alphanumeric words, removes common English stop words, and applies English Snowball stemming. BM25 combines term rarity, term frequency and document length. Query words use OR matching. More relevant records come first; equal scores use ascending key order. Custom metadata is returned but is not searched. Scores compare results within one query. They are not confidence scores.
This is lexical retrieval. English stemming is fixed. Other languages need evaluation.
Queries accept 1–16,384 UTF-8 bytes. Search uses at most 64 original indexed terms.
Long questions select terms from the whole input by corpus rarity and square-root
query frequency. query_reduced reports this selection. Scores then describe the
selected terms, not all words in the question.
The default index-read budget is 20,000 per request, set with
memory.max_search_postings. Short searches score all postings. Larger searches use
WAND score bounds to skip documents. index_reads counts iterator reads, including
seeks and exhausted lists. scored_candidates counts documents scored. Record,
scan-byte and response limits also apply. Check truncated: a work limit removes
the complete top-k guarantee. Scores of returned candidates use all selected terms.
Expansion is optional and off by default. Set expand: true for a small English
synonym list, or supply up to eight single words in expansion_terms. Search adds
at most eight indexed related terms at one-quarter weight. This is not embedding
similarity or argument stance detection. Check expansion_terms in the response.
Topic/tag/time filters run after scoring. They can require further pages.
Continue with next_cursor and unchanged filters. Ranked pages rescore the query.
A managed memory write, delete or expiry cleanup invalidates the namespace’s ranked
cursor; start a new search after such a change. Unchanged data permits continuation
after restart with the same runtime contract. This update invalidates old ranked
cursors; restart the search after upgrade. The database format is unchanged.
Expired records are hidden; corpus statistics include them until cleanup.
Literal retrieval and pagination
--query "Rust local" requires both literal substrings in the content.
Matching ignores letter case. A query accepts at most eight terms and 256 bytes.
Whitespace separates the terms.
The MVP returns metadata but does not search its fields.
Use search for relevance ranking. Neither tool provides semantic similarity or automatic fact extraction.
Time, topic and tag queries use ordered indexes. A combined topic/tag query uses the topic index, then tests the tag. Keyword filters test the candidate content. A keyword query without a topic or tag uses the time index. It can require several pages.
| Query setting | Default or limit |
|---|---|
| Results per page | Default 20; range 1–100 |
| Complete response size | Default 16 KiB; configured maximum 64 KiB |
| Candidate records examined | At most 1,000 per page by default |
| Encoded records examined | At most 4 MiB per page by default |
Set the scan and response limits in [memory].
These limits do not cap the process RSS.
instantkv recall --topic decisions --limit 10 --max-bytes 16384
# Copy the returned next_cursor, then keep exactly the same filters:
instantkv recall --topic decisions --limit 10 --max-bytes 16384 \
--cursor 'RETURNED_CURSOR'
- Keep the namespace and filters unchanged.
- Supply the returned
next_cursor, even after an empty page. - Continue until
next_cursoris null.
You can change the page size or response budget.
If one matching memory exceeds the budget, increase max_bytes.
The API returns an error instead of a partial memory.
Pages do not form a stable snapshot across writes.
Updates and new events can move records between pages.
HTTP and MCP
Routes under /v1/namespaces/{namespace}:
| Method / route | Body / result | Required grants |
|---|---|---|
POST /memories | {key?, memory: {content, topic?, tags?, metadata?, occurred_at_ms?}, ttl_seconds?, if_revision?} → saved memory | put |
GET /memories | Query filters → {items, next_cursor, scanned, scanned_bytes} | list + get |
POST /search | {query, topic?, tag?, since_ms?, until_ms?, limit?, max_bytes?, cursor?} → scored items and work counters | list + get |
GET /memories/{key} | Exact structured memory, revision and timestamps | get |
DELETE /memories/{key} | Forget; optional If-Match: "revision" | delete |
If you omit if_revision, the key must not exist.
To update a memory, supply its key and the observed positive revision.
TTL follows the namespace policy.
The forget operation rejects raw KV records.
The MCP remember tool accepts memory fields directly. It does not use the HTTP memory wrapper.
The recall, search and browse tools accept query fields directly.
Credentials, request limits and namespace permissions apply to each call.
MCP setup.
The command instantkv schema --kind memory generates the HTTP schemas.
Connect an actual local model
Run this isolated storage test without a model:
cargo build --release --locked -p instantkv
python3 scripts/memory-bench.py --records 100 --runs 1 --queries 10 \
--output /tmp/instantkv-memory-demo.json
The script starts a temporary server and saves structured memories. It tests topic, tag and time queries. Then it terminates and restarts the server. It verifies every saved field and calls the actual MCP bridge. It removes its temporary state.
This test verifies storage and transport. The model example below tests inference behavior.
The Ollama example reads five tool schemas from the instantKV MCP adapter. It sends tool calls and results between the model and adapter. It uses the Python standard library.
- Start instantKV.
- Start Ollama with an installed model that supports tools.
- Run the example from your memory directory:
python3 /absolute/path/to/instantKV/examples/local-llm.py \
--model YOUR_INSTALLED_MODEL --trace \
--prompt "Remember that I prefer Rust for local tools. Use topic preferences."
# Start a fresh invocation, with no previous chat messages:
python3 /absolute/path/to/instantKV/examples/local-llm.py \
--model YOUR_INSTALLED_MODEL --trace \
--prompt "What language do I prefer for local tools? Check memory first."
Trace output contains only tool names. Tool-call behavior depends on the model. The HTTP/MCP tools and the example’s protocol contract have automated tests. The example’s MCP bridge also passed a test with a real local server.
No installed Ollama model was available during implementation. Model recall quality and phone support remain unverified. The example follows Ollama’s tool-calling API.
Embed and extend
The Rust example saves a memory and closes the engine. A new engine instance retrieves that memory. The example requires no HTTP server or asynchronous runtime.
cargo run -p instantkv-core --example memory -- /tmp/instantkv-embedded-memory
The app selects namespaces, limits and custom metadata. The app controls authorization, cleanup scheduling and the data path.
Run synchronous storage calls outside the UI thread.
Swift/Kotlin bindings and device lifecycle tests are planned. Agents can add metadata fields and connect through typed tools. The service does not execute uploaded plugins.
Storage and upgrade boundary
Structured records use the reserved _instantkv_memory: 1 JSON envelope in records_v1.
The additional memory_index_v1 table contains time, topic and tag entries.
search_postings_v1, search_terms_v1 and search_stats_v1 hold the sparse search index.
Each entry points to a structured record.
Writes, updates, deletes and expiry cleanup change records and indexes in the same immediate transaction.
A raw KV write that replaces a structured record also removes the old index entries.
On first open, this build backfills the search index in one transaction. Stored record bytes and revisions do not change. Large existing corpora can take longer to open. Subsequent opens reuse the index. Ordinary records need no conversion.
- Back up existing data offline before you use this unreleased build.
- Use
rememberto create indexed memories when needed.
Older binaries do not maintain the new indexes. Do not use an older writer with an MVP database. To downgrade, restore a backup from before the MVP upgrade.
Logical quotas count the full JSON envelope. They exclude physical index and database overhead.
Evidence and next steps
The performance plan separates Mac measurements from device targets. The roadmap defines local-model evaluation, native mobile integration and portable exports. Local embeddings and real-agent retrieval evaluation remain future work.