Documentation menu / searchSearch documentation →

Start here

Quick startConnect through MCPOpenCode native memoryOpenCode memory controllerMemory API & local models

Use the service

Memory & compactionMemory lifecycle controllerLocal agents & swarmsCLI referenceHTTP reference

Run a node

ConfigurationOperations & backupsLocal memory & ARM

Evidence

Performance & device targetsFull retrieval reportBenchmark methodologyEvaluation policyMemory benchmark notes

Build with us

Architecture & schemaTechnology & learning mapRepository maintenanceContributingSecurityWebsite & deploymentSearch & agent discoveryPrivate product measurementEngineering references

Project

Cleanup & release planRoadmapLocal AI memory: when instantKV fitsFeaturesChangelog

History

Verification history

Proposals

Distributed memory proposal

Build with us / single-node · source MVP

Architecture & schema

Rust, the bounded engine, redb transactions and storage schema.

View the diagram
instantKV architecture and stored schemaHTTP, CLI and MCP clients authenticate with a scoped token. Grants and bounded policy route durable knowledge and checkpoints to an embedded redb database and optional scratch to RAM with TTL and FIFO eviction. The stored record carries a namespace, key, revision, value, TTL and checksum; a checkpoint commits a capsule and the session latest pointer in one transaction.02 / ARCHITECTURERUST · TOKIO · AXUM · REDBAGENT CLIENTSHTTPbearer token, conditional writesCLIrecords, checkpoints, benchMCP · 7 TOOLSstore, list, get, delete,checkpoint, restore, healthBOUNDED POLICYGRANTStoken → namespacesper-operation rightsLIMITSrecord, key and value sizenamespace quotas, restore budgetadmission controlCHECKSUMread verified against the writeevery read is accountedSTORAGEDURABLE · REDBshared + private knowledgeimmutable capsulessession latest pointerno TTL by defaultSCRATCH · RAMoptional namespaceTTL and FIFO evictioncleared on restartdisposable notesauthdurableramSCHEMA 1recordnamespacekeyrevisionvaluettl · checksumcapsulegoal · constraintsunfinished · nextreferences → keyssession latestONE TRANSACTIONCAPSULE + SESSION LATEST POINTERa checkpoint write commits the immutable capsule and moves the session pointer togetherREAD PATHEXACT KEY OR PREFIX, NO MODEL CALLSrestore reports which referenced record changed or went missingStack: Rust with Tokio and Axum; redb supplies durable transactions without a separate database service.one binary · one config · one data directory
authenticated request durable writedashed frame: stored schema, version 1

Status: single-node implementation with an unreleased memory MVP. Updated: 2026-10-03. Verification history records completed checks. Future work appears below.

Clients connect through HTTP, CLI or MCP. Permissions and limits apply before storage access. Durable records use redb; scratch uses RAM. Rust apps can also embed the core directly.

Stack decision

LayerChoiceWhy / constraint
CoreRust 2024, minimum Rust 1.98Ownership and explicit memory/concurrency bounds
HTTPTokio + AxumAsync networking, middleware, graceful shutdown
Durable storageredbPure Rust, embedded ACID transactions; one writer
ScratchPer-namespace Mutex + ordered mapsAtomic quotas; ordered prefix, deadline and FIFO indexes
ConfigurationSerde + TOMLStrict fields and startup validation
Agent toolsOfficial rmcp SDK, stdio → HTTPSame auth and policy as every client
DeploymentBinary / non-root Docker + volumeNo database service or orchestration dependency
VerificationFake-clock invariant tests + HTTP/MCP tests + CLI load clientDeterministic lifecycle checks and measured request paths

Rust provides explicit control over memory and concurrency. redb handles transactions and persistence inside the process, which permits one binary. Go could also implement the service. Valkey could provide storage when a separate database is acceptable. Language choice alone does not prove a speed advantage.

Sources: Rust ownership, Axum, redb concurrency.

Boundaries and request paths

instantkv-core owns configuration policy, input validation, clocks, quotas, indexes, storage and checkpoint transactions. instantkv provides HTTP, authentication, CLI, MCP, demos and benchmarks. HTTP, CLI and MCP use the same permission layer. An app that embeds the core must provide its own authorization. Local integration.

Write flow: authorization → body limit → format/TTL/size validation → revision and quota validation → commit → success response. A failed write preserves the old live value. Input validation can reclaim expired data.

Read flow: authorization → lookup → expiry test → value and revision. Raw record lists return metadata only, with at most 1,000 items and bounded scans. An extra empty page is possible. Continue until next_cursor is null. Pages do not form a stable snapshot across writes.

A semaphore limits active requests and cleanup work. Synchronous storage runs on spawn_blocking. A submitted task keeps its permit even after the HTTP deadline. There is no unbounded writer queue. A timed-out write can still commit. Tokio blocking-task behavior explains this boundary.

After a timeout, inspect the revision or retry the same checkpoint ID and payload.

Stored schema: format 1

All tables use instantkv.redb. Record and counter integers use little-endian u64 encoding. Namespace names and ordinary keys cannot contain NUL. The composite record key is therefore unambiguous.

TableKeyValue
records_v1namespace + NUL + key32-byte header + raw value bytes
usage_v1namespaceentry count, key/value bytes, revision high-water mark
expiry_v1padded UTC deadline + revision + composite keycomposite record key
metadata_v1format / namespace identityformat version / storage mode + purpose
memory_index_v1namespace + index kind + optional normalized label + event time + keycomposite structured-memory record key

The record header contains revision, write time, expiry time and insertion order. An expiry of zero means no expiry. Ordinary JSON records accept your own fields without a fixed wrapper. Namespaces can also accept raw bytes or UTF-8 text.

Structured-memory APIs use a _instantkv_memory: 1 envelope in durable JSON records. Time, topic and tag indexes change in the same transaction as records, quotas and expiry entries. Deletion, replacement and expiry cleanup remove old index entries. Recall reads indexes and values in one snapshot, with candidate, scan-byte and response limits. Literal recall filters content. Ranked search uses sparse BM25 postings and English stemming in the same redb transaction. Neither requires a vector index or model. Ranked search accepts 16 KiB input, selects up to 64 original indexed terms, and uses exhaustive sparse scoring or bounded WAND pruning. Optional expansion adds at most eight related terms at low weight. Query reduction and work limits are visible in responses. Ranked cursor fingerprints bind the query and expansion settings. Memory shape and cursor contract.

Checkpoint namespace reserved records:

__checkpoint/<id>            -> JSON CheckpointRequest (immutable capsule + references)
__latest/<agent>/<session>    -> checkpoint ID bytes

The bundle, latest pointer and usage counters share one redb write transaction. expected_latest_revision = null requires no prior pointer. Later saves use the last returned pointer revision. Identical ID/payload retries do not rewind the pointer. Old bundles can be deleted, but the current latest bundle is protected. Deletion removes retry history; checkpoint IDs must remain unique.

Writes retain redb’s immediate durability. Success follows commit. The filesystem and storage device must honor synchronization. Restore reads the bundle and pointer in one snapshot. Reference tests run separately and report current observations, not a snapshot of all referenced keys.

Retention and accounting

RAM expiry uses a monotonic deadline. Durable expiry uses UTC and can change with the wall clock. Reads hide expired records immediately. Cleanup has a batch limit. Expired records count toward quota until removal. RAM writes also run a bounded cleanup batch. Durable input validation does not scan all expired records.

Scratch eviction follows insertion order. Reads and live overwrites do not change that order. Recreating an expired entry gives it a new insertion order. Durable storage rejects over-quota writes and does not auto-evict. Revisions do not reset after deletion, so stale conditions cannot match a recreated key.

Quotas count stored key/value bytes and entries, including checkpoint bundles and pointers. They exclude allocator overhead, process RSS, indexes and physical database size. Deleted disk pages can be reused. Deletion does not provide secure erasure.

Access and operations

Namespaces define API access boundaries. Agent and session labels are metadata. Private workers need separate namespace grants. Grants also apply to referenced records. The server loads credentials at startup and compares fixed-size token hashes. The swarm profile gives workers read-only shared access and private record/checkpoint access. The operator can access all swarm namespaces.

Legacy namespace/operation shorthand remains supported. It cannot be combined with explicit grants. Checkpoint saves require readable reference namespaces. Restore uses current GET grants. Swarm setup.

Disabled authentication requires loopback. Remote access uses SSH tunneling or an HTTPS proxy. Logs exclude request bodies and tokens. Metrics report aggregate counters.

Health reports a running listener after successful startup validation and database opening. It does not periodically test disk writes. Framework parsing errors can be plain text; service errors use a JSON object. One process must own each data file. Backup and upgrades.

Optional memory controller

The source-only Python controller uses the existing HTTP API. It owns fact interpretation, provenance, conflicts, profiles and forgetting. Each slot commits its current content and lifecycle metadata in one native structured-memory transaction. Local-model extraction proposes assertions for explicit review; the Rust node makes no model calls.

The controller design defines retries, temporal ordering, tombstones and finite bounds. It requires a durable namespace without native TTL. It does not add multi-fact transactions or consistent read snapshots.

Future proposals

  1. Connect a real runtime’s compaction hooks and evaluate task continuation.
  2. Add session retirement, online backup and deeper fault tests.
  3. Measure writer contention, expiry delay and memory growth before adding storage concurrency features.
  4. Implement the distributed proposal in stages. Start with export/import, immutable baselines and private overlays. Add reviewed publication before synchronization.
  5. Evaluate BM25 ranking with more corpora and real tasks before adding local embeddings.

Topic/tag/time indexes, literal filters and bounded BM25 ranking work in the source MVP. Semantic search, automatic failover, a custom write-ahead log and Redis compatibility remain separate design choices.