memo-mcp operator guide
This guide is for the people who install, configure, tune and run memo-mcp. It is written for readers comfortable with a terminal, SQLite, environment variables and Prometheus.
memo-mcp is a single, pure-Go binary. It exposes a knowledge base to AI agents over the Model Context Protocol (MCP) on stdio. It stores everything in one SQLite file per knowledge base, and it ships with a CLI and a read-only web UI over the same store. There is no daemon to deploy, no external database and no network dependency after a one-time model download.
What you will find here
| Part | Read it when you need to… |
|---|---|
| I. How it works | Understand the processes, the data model, the search pipeline and the security boundaries |
| II. Install and deploy | Install, connect MCP clients, lay out several knowledge bases, upgrade, run offline |
| III. Configuration reference | Look up an environment variable, a command, a flag, the overrides file or a path on disk |
| IV. Retrieval tuning | Change ranking, switch embedding models, try the reranker, and measure whether a change helped |
| V. Knowledge lifecycle | Administer ingest, facts, trust, the graph index, compaction and deletion |
| VI. Monitoring and observability | Scrape metrics, write alerts, read logs, use the call log |
| VII. Operations | Back up, verify, size, export, troubleshoot, release |
Conventions
- Commands are shown as
memo-mcp <command>. Every command reads the same environment (MEMO_HOME,MEMO_KB, …), so set it the same way the MCP client does. - “Knowledge base” (KB) means one SQLite file. A “namespace” is a label inside one file.
- Addresses such as
memo://chunk/42identify records; every CLI command and the UI accept them. - Defaults quoted here are the shipped values for v1.4.
memo-mcp profiles showandmemo-mcp --helpprint the live values for your build.
Deeper references
This guide states what to do and what it costs. The design documents explain why:
- Architecture: the 1.x contract, formulas and defaults.
- Schema: every table and column, trust transitions, time rules.
- Eval reports: measured quality and cost per release.
Looking for the non-technical introduction? See the user guide.
Part I: How it works
Four short chapters give you the mental model the rest of the guide assumes.
- System overview: one binary, three ways to run it, one file per knowledge base.
- Data model: layers, addresses, trust and the two clocks.
- The search pipeline: arms, fusion, cutoff, abstention and explain.
- Security model: what is exposed, what leaves the machine, who may change what.
System overview
One binary, three processes
memo-mcp is one statically linked Go binary (CGO_ENABLED=0, about 30 MB). The same
binary runs in three roles, which are separate processes that share only the database file:
| Role | Started by | Talks over | Lifetime |
|---|---|---|---|
MCP server (memo-mcp or memo-mcp serve) | The MCP client (Claude Code, Claude Desktop, …) | stdio: JSON-RPC on stdin/stdout, logs on stderr | As long as the client session |
Web UI (memo-mcp ui) | An operator | HTTP on a loopback port, GET only | Until Ctrl-C |
CLI (memo-mcp <command>) | An operator or a script | Terminal | One command |
Each MCP client session usually starts its own server process. Several processes can hold the same file open at once; SQLite WAL mode and a 5-second busy timeout serialise writers.
MCP client ── stdio ──▶ memo-mcp serve ─┐
operator ── browser ──▶ memo-mcp ui ────┼──▶ $MEMO_HOME/kb/<name>.db (SQLite, WAL)
operator ── shell ────▶ memo-mcp <cmd> ─┘
└─▶ ~/.cache/memo-mcp/models (embedding models)
Inside the process
| Package | Responsibility |
|---|---|
internal/server | The ten MCP tools and the memo:// resources; request middleware (metrics, logs, call log) |
internal/cli | Commands, configuration resolution, serve and UI wiring |
internal/retrieve | The search pipeline: arms, fusion, recency, cutoff, budget, explain |
internal/kb | Schema and migrations, ingest, facts, trust, time, graph tables, pages, export |
internal/embedding | The model registry, download, in-process inference (hugot / GoMLX, pure-Go ONNX) |
internal/graph, internal/compact | Entity extraction and resolution; compaction work items and checks |
internal/ui | The read-only web UI (html/template, embedded assets) |
internal/obs | Metrics registry, Prometheus exposition, logging setup |
The ten MCP tools
| Tool | Writes? | Purpose |
|---|---|---|
ingest | yes | Store a markdown document with provenance; same content is a no-op, changed content a new revision |
search | no* | Hybrid search with optional per-result explanations |
read | no | Dereference a memo:// address as a passage, section or document |
remember | yes | Record one fact with evidence and validity dates; supersedes corrects an old one |
forget | yes | Retire a document or fact with a reason (agent-written records only) |
promote | via a human | Ask a human to raise trust, through an elicitation dialog or a CLI command |
explore | no | Walk the entity graph from one name |
compact | work items | Propose pages, refreshes, conflict and merge decisions, duplicates |
submit | yes | Hand back a page or a decision for a work item; checked before it is stored |
status | no | Namespaces, model, pending vectors, jobs, graph and page counts |
* search writes only to the opt-in query log.
Resources mirror the addresses (memo://doc/{id}, memo://chunk/{id}, …), and
memo://index returns a one-line-per-document index under 8 KB.
The one network call
The first time a process needs an embedding model that is not cached, it downloads the model
files from Hugging Face into ~/.cache/memo-mcp/models. Nothing else in the default
configuration opens an outbound connection. The optional Ollama executor and the
--metrics-addr listener are explicit opt-ins, and both are loopback-only by default. See
Air-gapped installs to avoid even the download.
Design reference: architecture §1–§4.
Data model
One knowledge base is one SQLite file. Inside it, knowledge is stored in layers. Each layer points down to the one it was derived from, so every answer can be traced to its source.
L5 pages markdown an agent wrote from passages; cites them; marked stale when a source changes
L4 graph entities, aliases, mentions (entity ↔ passage), merge candidates, typed edges
L3 facts one-sentence claims with a validity window and an evidence passage
L2 chunks passages of ~200 estimated tokens, indexed three ways (two FTS5 + vectors)
L1 documents the text as ingested, one row per revision
L0 sources where it came from: URI or file, library, version, hash, trust, origin
bookkeeping namespaces, models, jobs, audit, query_log, call_log
| Layer | Rebuildable? | Notes |
|---|---|---|
| L0–L2 | Yes, deterministically from the sources | Chunking and indexing are pure functions of the text |
| L3 | No: facts are additive | Corrections supersede or invalidate; nothing is deleted in place |
| L4 | Yes: memo-mcp graph rebuild | Merge decisions made by a human are kept |
| L5 | No: written by agents | Stored as is_inference, with the passages each page cites |
Addresses
| Address | Resolves to |
|---|---|
memo://source/<id> | The latest live document of a source |
memo://doc/<id> | One document revision |
memo://chunk/<n> | One passage (read with section adds its neighbours) |
memo://fact/<id> | One fact, with its evidence address and its two clocks |
memo://entity/<id> | One entity, its aliases, passages and neighbours |
memo://page/<id> | One curated page with its sources and stale flag |
memo://index, memo://ns/<namespace>/index | One line per document and fact, under 8 KB |
Document and fact ids are UUIDv7, which sort by time. Chunk ids are integers. A forgotten record’s address still resolves, to “forgotten on … because …”.
Namespaces and knowledge bases
- A namespace is a label on sources inside one file. A search spans all namespaces unless the request scopes it. Namespaces organise; they do not isolate.
- A knowledge base is a file selected by
MEMO_KB. Files are the isolation and privacy boundary: no command or tool reads across files.
Trust and origin
Every source, and therefore every passage and fact, carries two labels:
| Label | Values | Set by |
|---|---|---|
| trust | agent < user < curated | The channel the write came through, never the request. MCP tool writes are agent. CLI writes are user by default, curated only with --trust curated |
| origin | web, user-said, agent-derived | Declared by the writer: fetched from the web, said by a person, or inferred by an agent |
Only a human can raise trust: memo-mcp trust promote, or an elicitation dialog the client
shows for the promote tool. Tool calls may forget or supersede only agent records. Every
transition is written to the audit table. See Trust administration.
Two clocks
- Recorded time: when the knowledge base learned something.
- Valid time: when a fact is true in the world (
valid_from,valid_to).
as_of queries (“what did we believe on 1 March?”) filter on recorded time, and also on valid
time for facts that have a window. Superseded revisions and replaced facts stay readable under
as_of. Forgotten records are excluded in both views. See Facts and time.
Full column-by-column reference: schema.md.
The search pipeline
Every search, from the MCP tool, the CLI or the UI, runs the same pipeline in
internal/retrieve:
scope filter ─▶ arms (parallel ranked lists) ─▶ fusion ─▶ recency ─▶ abstention
─▶ cutoff ─▶ limit ─▶ token budget ─▶ results + why + trace
1. Scope filter
Namespaces, kinds, sources, library, version, tags, dates and minimum trust are applied
inside each arm’s SQL, before top-k. A filtered search therefore never loses a result to a
candidate that was filtered out later. The live filter excludes superseded and forgotten
records. Under as_of, it is replaced by “recorded on or before T and not superseded before T”.
2. Arms
| Arm | Index | Runs when |
|---|---|---|
keyword | FTS5, Porter stemmer, BM25, over text and its section header | Always, except mode: semantic |
exact | FTS5 that keeps identifiers whole (net/http, ERR_CONN_RESET) | mode: exact, or auto when the query looks like an identifier |
semantic | Cosine similarity over unit vectors in a plain SQLite table | A model is loaded and passages have vectors for it |
fact | Facts matched by keyword or meaning; a hit votes for its evidence passage | Always |
entity | Passages mentioning entities the query names | Routed: the query names 2+ known entities, or 1 with relational phrasing |
graph | Personalised PageRank over the mention graph, one walk per named entity | Same routing; reports only passages not every seed mentions directly |
Each arm returns up to fetch_depth candidates (100 by default).
3. Fusion
Weighted reciprocal rank fusion: a passage’s score is the sum over arms of
weight / (k + rank), with k = 60. Default weights are semantic 0.5, keyword 0.5, exact 0.3,
fact 0.4, entity 0.4 and graph 0.5. The minmax profile uses score fusion instead.
4. Recency
For kinds that age (note, conversation by default), the score is multiplied by
floor + (1 − floor) · 0.5^(age_days / half_life), with floor 0.8 and half-life 90 days.
Versioned documents and code do not age.
5. Abstention and relevance bands
A candidate whose only evidence is a semantic similarity below semantic_floor is dropped. If
nothing survives, the response has zero results, a reason and a hint. That is a deliberate
“not in the knowledge base”, not an error.
relevance is the best raw cosine for a result, interpreted against the model’s bands. For
the default model, granite-small-r2, strong is 0.88, moderate 0.80 and weak 0.72, because
unrelated text already scores about 0.63 under it. score orders one list and is not
comparable across queries; relevance is.
6. Cutoff, limit and budget
The list ends at the first result, past min_results (3), whose score drops by more than
cutoff_gap (half) from the previous one. It is then capped at limit (default 10, maximum 100),
and then packed into max_tokens. The response always reports truncated, the count of ranked
results that did not fit, and narrow_hint, the scope field that would shorten the list.
7. Explain and trace
With response_format: explain (MCP) or --explain (CLI), every result carries why: each
arm’s rank, raw score and contribution, the fused and final score, relevance band, provenance
and freshness. Every response carries a trace: arms run and why, candidates and latency per
arm, filter counts, cutoff kind, budget use and degraded state. The UI renders the same
structures, and a test asserts that the numbers are identical.
Degraded mode
If the embedding model is not available, because it is downloading, failed to load or has no
vectors yet, search runs without the semantic arm. It sets degraded{flag, reason}, and the
memo_search_degraded_total metric counts it. Writes made in that state queue their vectors.
memo-mcp backfill or the next server start embeds them.
Tuning these constants: Profiles and constants. Full formulas: architecture §5.
Security model
memo-mcp is designed to run on one person’s machine with that person’s privileges. Its security properties come from being local, small and explicit.
Exposure
| Surface | Exposure | Guard |
|---|---|---|
| MCP server | stdio of the client process only | No network listener. stdout is reserved for JSON-RPC; all diagnostics go to stderr |
| Web UI | Loopback (127.0.0.1:0 by default) | GET only, Host header allowlist against DNS rebinding, CSP, no mutating routes. A non-loopback bind needs --allow-remote |
| Metrics endpoint | Off by default; loopback only, with no override | GET /metrics only, Host check, 404 for any other path |
| Ollama executor | Off by default; loopback unless --allow-remote | CLI only; no MCP code path reaches it |
Outbound traffic
- Model download from Hugging Face on first use of a model, into
~/.cache/memo-mcp/models. Avoidable; see Air-gapped installs. - Nothing else. The server never fetches URLs. Agents fetch content and pass the text to
ingest. There is no telemetry: metrics are pull-only on loopback, logs go to stderr, and the call log is a table in your own file.
Data at rest
- One SQLite file per knowledge base, in plaintext. Directories are created
0700and files0600. A pre-existing directory with wider permissions is not tightened; check it. - Use full-disk encryption if the file may hold sensitive material.
forget --redacterases a record’s text. Plainforgethides the record but keeps the text in history.
Trust gates
- Trust is set by channel. An agent cannot write a record above
agent, and cannot forget or supersedeuserorcuratedrecords. - Raising trust requires a human:
memo-mcp trust promote, or an MCP elicitation dialog that shows the excerpt, the source and the target level. A client configured to auto-accept elicitation removes this protection. In that case, treatuserandcuratedas no stronger thanagent. - Every write and every trust change is recorded in the
audittable with its actor and channel.
Prompt injection
Ingested content is data, but agents read it. memo-mcp limits the blast radius: content
cannot raise its own trust, cannot delete human records and cannot trigger network calls.
Results carry provenance and trust, so an agent, or a human reviewing its answer, can discount
agent-trust web content.
The MCP host
Everything an agent writes or searches passes through the MCP host as tool input and output.
The knowledge base is as private as that host. memo-mcp does not advertise the MCP logging
capability; logs stay on stderr, which the host may display or store.
Part II: Install and deploy
memo-mcp has no installer and no service to register. Installing it means putting one binary on disk and telling an MCP client how to start it.
- Release archives: download, verify and place the binary.
- Build from source: reproducible pure-Go builds.
- Connecting MCP clients: Claude Code, Claude Desktop and any stdio client.
- Knowledge bases and MEMO_HOME: one file per KB, per-project setups.
- Upgrading and migrations: schema versions and what to run after an upgrade.
- Air-gapped installs: no network at runtime, not even the model download.
Release archives
Every tag vX.Y.Z publishes five archives and a checksum file on the
releases page:
| Platform | Archive |
|---|---|
| macOS, Apple silicon | memo-mcp_X.Y.Z_darwin_arm64.tar.gz |
| macOS, Intel | memo-mcp_X.Y.Z_darwin_amd64.tar.gz |
| Linux, x86-64 | memo-mcp_X.Y.Z_linux_amd64.tar.gz |
| Linux, ARM64 | memo-mcp_X.Y.Z_linux_arm64.tar.gz |
| Windows, x86-64 | memo-mcp_X.Y.Z_windows_amd64.zip |
| All | checksums.txt (SHA-256) |
Each archive contains the binary, LICENSE, README.md and SKILL.md.
Install
VERSION=1.4.0
ARCH=darwin_arm64 # see the table above
curl -LO https://github.com/kKEo/memory-find/releases/download/v${VERSION}/memo-mcp_${VERSION}_${ARCH}.tar.gz
curl -LO https://github.com/kKEo/memory-find/releases/download/v${VERSION}/checksums.txt
shasum -a 256 -c --ignore-missing checksums.txt # Linux: sha256sum -c --ignore-missing
tar xzf memo-mcp_${VERSION}_${ARCH}.tar.gz
install -m 0755 memo-mcp ~/.local/bin/memo-mcp # any directory; note the absolute path
memo-mcp version
memo-mcp version prints the build version, the MCP protocol version (2026-07-28), the Go
version and the model directory.
macOS Gatekeeper. A binary downloaded by a browser carries a quarantine attribute and is blocked on first run. Clear it with
xattr -d com.apple.quarantine ~/.local/bin/memo-mcp, or download withcurl, which does not set it.
MCP registry
The server is listed in the MCP registry as io.github.kKEo/memory-find. Its server.json
points at the release archives with their checksums, so registry-aware clients can install it
directly.
What the binary needs at runtime
- No shared libraries, no CGo, no interpreter.
- A writable
MEMO_HOME, by default~/.memo-mcp, and a writable~/.cache/memo-mcp/models. - About 200 MB of disk for the default model, plus the knowledge-base files. See Capacity and performance.
- Network access to
huggingface.coonce per model, unless you pre-seed the cache.
Build from source
Requirements: Go 1.26 or newer. No C toolchain is needed.
git clone https://github.com/kKEo/memory-find
cd memory-find
make build # CGO_ENABLED=0, stripped, version from `git describe`
./memo-mcp version
make build runs:
CGO_ENABLED=0 go build -ldflags "-s -w -X main.version=$(VERSION)" -o memo-mcp ./cmd/memo-mcp/
Cross-compiling is plain Go, for example
GOOS=linux GOARCH=arm64 CGO_ENABLED=0 go build ./cmd/memo-mcp.
Reproducing a release
make snapshot runs the same GoReleaser configuration as the release workflow without
publishing. It writes the five archives and checksums.txt under dist/. Release builds use
-trimpath and the commit timestamp, so rebuilding a tag gives the same binaries.
Development targets
| Target | What it does |
|---|---|
make check | gofmt, vet, golangci-lint, then all tests under the race detector |
make test | All tests, no race detector |
make eval | The retrieval benchmark against the recorded baseline |
make golden-update | Re-record the tools/list golden after a deliberate tool-surface change |
make baseline-update | Re-record the eval baseline after a deliberate retrieval change |
make docs | Build these guides with mdBook into site/ |
Tests use a deterministic hash embedder and need no network or model download.
Connecting MCP clients
memo-mcp speaks MCP over stdio. A client starts the binary, writes JSON-RPC to its stdin and reads responses from its stdout. Configuration is passed through environment variables in the client’s server entry, not the shell. The client starts the process, so your shell profile does not apply.
Claude Code
# In a project directory: this project only, private to you (the default "local" scope)
claude mcp add memo --env MEMO_KB=my-project -- /absolute/path/to/memo-mcp
# Shared with the team through .mcp.json in the repository
claude mcp add memo --scope project --env MEMO_KB=my-project -- memo-mcp
# Available in every project
claude mcp add memo --scope user --env MEMO_KB=personal -- /absolute/path/to/memo-mcp
Add more variables with additional --env flags, for example
--env MEMO_QUERY_LOG=1 --env MEMO_LOG_LEVEL=warn. Check the connection with /mcp inside
Claude Code. Claude Code shows a stdio server’s stderr, so memo-mcp’s logs appear there.
Claude Desktop
Edit claude_desktop_config.json. On macOS it is in ~/Library/Application Support/Claude/;
on Windows, in %APPDATA%\Claude\. Use absolute paths; ~ is not expanded.
{
"mcpServers": {
"memo": {
"command": "/Users/you/.local/bin/memo-mcp",
"env": {
"MEMO_KB": "my-project",
"MEMO_QUERY_LOG": "1"
}
}
}
}
Restart Claude Desktop completely after editing.
Any other stdio client
The generic shape is the same everywhere: a command, optional args and an environment map.
| Field | Value |
|---|---|
| command | Absolute path to memo-mcp |
| args | none, or ["serve", "--metrics-addr", "127.0.0.1:9469"] |
| env | MEMO_KB, and optionally MEMO_HOME, MEMO_MODEL, MEMO_PROFILE, MEMO_QUERY_LOG, MEMO_LOG_FORMAT, MEMO_LOG_LEVEL |
The server implements protocol version 2026-07-28: stateless, with server/discover in
place of initialize. It advertises tools and resources. It does not advertise logging or
prompts. The promote tool uses elicitation when the client supports it, and otherwise
returns the CLI command a human can run.
Teaching the agent
Ship SKILL.md from the release archive with the project, for example as a Claude Code
skill, or paste it into the system prompt. It describes the search-then-read loop, scoping,
writing facts and how to treat trust. For a static orientation, add the output of
memo-mcp export --index to CLAUDE.md or AGENTS.md.
Verifying a connection without a client
MEMO_KB=my-project memo-mcp status # read-only; fails if the file does not exist yet
A server that starts and immediately exits usually has an invalid MEMO_KB name or an
unwritable MEMO_HOME. The reason is on stderr. See the
troubleshooting runbook.
Knowledge bases and MEMO_HOME
Where files go
| Path | Contents | Created with |
|---|---|---|
$MEMO_HOME (default ~/.memo-mcp) | Base directory | 0700 |
$MEMO_HOME/kb/<name>.db | One knowledge base (plus -wal and -shm while open) | 0600 |
$MEMO_HOME/profiles.json | Optional ranking-profile overrides | You |
~/.cache/memo-mcp/models/ | Downloaded embedding and reranker models, one directory each plus a .ok marker | 0755 (public model files) |
<name> comes from MEMO_KB, default default. It must match ^[A-Za-z0-9._-]{1,64}$ and must
not be . or ... The model cache location is fixed and shared by every knowledge base and
process of the same user.
One KB per project
Files are the isolation boundary, so the usual pattern is one MEMO_KB per project or client:
cd ~/src/payments && claude mcp add memo --env MEMO_KB=payments -- ~/.local/bin/memo-mcp
cd ~/src/analytics && claude mcp add memo --env MEMO_KB=analytics -- ~/.local/bin/memo-mcp
Claude Code’s default local scope ties each entry to its project directory, so the same server
name memo opens a different file in each project.
Namespaces within a KB
Use namespaces for topics that may be searched together: docs, decisions, grpc. Writes
default to the namespace default. Searches span all namespaces unless scoped. Namespace
counts are visible in memo-mcp status and as the memo_kb_namespace_documents metric.
Separate homes
MEMO_HOME moves everything except the model cache. Use it to keep test data apart, run CI
against a throwaway directory, or put knowledge bases on an encrypted volume:
MEMO_HOME=/Volumes/secure/memo MEMO_KB=client-a memo-mcp status
Sharing between processes
Several processes may open the same file. WAL mode lets readers proceed during writes, and writers wait up to 5 seconds for each other. Each process serialises its own writes over one connection. Do not put knowledge bases on network file systems such as NFS or SMB: SQLite’s locking is not reliable there.
Legacy variables
JOURNAL_TOKEN and JOURNAL_PATH are accepted for one more release as aliases of MEMO_KB
and MEMO_HOME, with a deprecation warning. Files from the pre-1.0 journal are not opened or
migrated.
Upgrading and migrations
Replacing the binary
Stop or let finish any running memo-mcp processes, replace the binary, and start the client
again. Each MCP session starts its own process, so a new binary is picked up on the next session.
Schema versions
The schema version is SQLite’s user_version. Migrations are additive and run forward only.
| Version | Release | Adds |
|---|---|---|
| 1 | 0.4 to 1.0 | Sources, documents, chunks and their indexes, facts, bookkeeping |
| 2 | 1.1 | Graph layer: entities, aliases, mentions, merge candidates, edges |
| 3 | 1.2 | Pages layer: pages, page sources, page vectors, work items |
| 4 | 1.4 | call_log table |
memo-mcp status prints the version in its first line, and the metric
memo_kb_schema_version exposes it.
What migrates and when
| Situation | Behaviour |
|---|---|
A writing process opens an older file: serve, ingest, remember, forget, search, verify, backfill, migrate, … | Migrates forward automatically, one transaction per migration |
A read-only command opens an older file: status, ui, ls, read, export, metrics, facts, explore, lint, pages, trust ls, graph merges | Refuses with file is at schema vN, this binary expects vM; run memo-mcp migrate. Read-only opens never write |
| Any process opens a newer file than it knows | Refuses with database schema is at version N, but this binary only supports up to version M; upgrade memo-mcp |
To migrate explicitly before anything else touches the file:
MEMO_KB=my-project memo-mcp migrate
Downgrades are not supported. Back up the file before upgrading across a schema version if you may need to roll back. See Backup and restore.
After upgrading to 1.1 or later from 1.0
Files written before the graph layer have no entity mentions. Populate them once:
memo-mcp graph rebuild # all namespaces; or --ns <name>
Run it again after any release whose notes mention an extraction change.
After changing the embedding model
A model change is not a schema change. Vectors are stored per model, so the server embeds the
missing ones in the background at start, and search reports degraded until it finishes. See
Embedding models.
Release notes
Read the changelog and the
eval report for the tag, docs/eval/vX.Y.Z.md, before upgrading. The 1.x contract keeps tool
names and parameters, memo:// addresses, explain field names and the export format stable.
Air-gapped installs
After the binary is in place, the only network access memo-mcp ever makes is downloading an embedding model the first time it is needed. To run with no network at all, pre-seed the model cache.
Pre-seed the model cache
On a connected machine with the same memo-mcp version:
memo-mcp model pull granite-small-r2 # or the model you will set in MEMO_MODEL
ls ~/.cache/memo-mcp/models/
# onnx-community_granite-embedding-small-english-r2-ONNX/
# onnx-community_granite-embedding-small-english-r2-ONNX.ok
Copy both the directory and its .ok marker to the same path on the target machine,
~/.cache/memo-mcp/models/, for the user who runs memo-mcp. The marker records a completed,
verified download. Without it, the files are treated as an interrupted download and fetched
again.
Then check that the model loads offline:
memo-mcp model smoke granite-small-r2
For the optional reranker, also copy cross-encoder_ms-marco-MiniLM-L6-v2 and its marker.
Choosing a smaller model
| Model id | Download | Notes |
|---|---|---|
granite-small-r2 (default) | about 195 MB | Best measured quality; about 0.2 s per query on this backend |
potion | about 130 MB | Static lookup table, no neural network; about 6 ms per query; weaker on paraphrase |
minilm | about 87 MB | The previous default |
See Embedding models for the full registry.
Running with no model at all
If no model is cached and the download fails, memo-mcp keeps working. Search runs on the
keyword, exact and fact arms and reports degraded. New passages queue their vectors. When a
model becomes available, memo-mcp backfill or the next server start embeds the backlog.
Part III: Configuration reference
memo-mcp is configured through environment variables, command flags and one optional file. There is no configuration file for the server itself.
- Environment variables: every variable, its default and its effect.
- Commands and flags: every command and flag.
- profiles.json: ranking-profile overrides.
- Files on disk: what memo-mcp creates and where.
Precedence. A flag beats its environment variable, which beats the built-in default.
Examples: serve --metrics-addr over MEMO_METRICS_ADDR, search --profile over
MEMO_PROFILE.
Environment variables
Set these in the MCP client’s server entry for the server, and in your shell for CLI commands.
Both must agree on MEMO_HOME and MEMO_KB to see the same file.
Storage and selection
| Variable | Default | Effect |
|---|---|---|
MEMO_HOME | ~/.memo-mcp | Base directory. Knowledge bases live in $MEMO_HOME/kb/; overrides in $MEMO_HOME/profiles.json |
MEMO_KB | default | Knowledge-base name; the file is $MEMO_HOME/kb/<name>.db. Must match ^[A-Za-z0-9._-]{1,64}$ |
Retrieval
| Variable | Default | Effect |
|---|---|---|
MEMO_MODEL | granite-small-r2 | Embedding model id from memo-mcp model ls. Queries use this model’s vectors; missing vectors are embedded in the background at server start |
MEMO_PROFILE | default | Ranking profile for the server and the UI (memo-mcp profiles show lists them). The CLI search --profile flag defaults to it |
MEMO_RERANK | unset | 1 loads the cross-encoder reranker. Only profiles with rerank on (precise) use it. Measured slower and worse than the default; experiments only |
MEMO_RERANKER | ms-marco-minilm | Reranker id to load when MEMO_RERANK=1 |
Logging and observability
| Variable | Default | Effect |
|---|---|---|
MEMO_LOG_FORMAT | text | text or json. Logs always go to stderr |
MEMO_LOG_LEVEL | info | debug, info, warn or error. Unknown values fall back to the default with one warning |
MEMO_METRICS_ADDR | empty (off) | Same as serve --metrics-addr: serve Prometheus metrics at http://<addr>/metrics. Loopback addresses only |
MEMO_QUERY_LOG | unset | 1 records every search in query_log and every tool call in call_log, in the same file. See Call and query logs |
Compaction executor (CLI only)
| Variable | Default | Effect |
|---|---|---|
MEMO_OLLAMA_URL | http://127.0.0.1:11434 | Ollama endpoint for memo-mcp compact --executor ollama. Loopback only unless --allow-remote |
MEMO_OLLAMA_MODEL | none (required) | Ollama model name for the executor. Use a non-thinking instruction model |
Deprecated
| Variable | Replacement | Behaviour |
|---|---|---|
JOURNAL_TOKEN | MEMO_KB | Used as the name when MEMO_KB is unset, with a warning. Old journal files are not opened |
JOURNAL_PATH | MEMO_HOME | Used as the base directory when MEMO_HOME is unset, with a warning |
Not configurable
- The model cache,
~/.cache/memo-mcp/models, is fixed per user. - SQLite pragmas are fixed: WAL,
synchronous=NORMAL,busy_timeout=5000, foreign keys on.
Commands and flags
Running memo-mcp with no arguments is the same as memo-mcp serve. memo-mcp --help prints
the summary. Flags may come before or after positional arguments.
Opens shows how each command opens the knowledge base:
- rw: creates and migrates the file.
- rw, no create: migrates, but refuses a missing file.
- ro: never writes, and refuses a missing or out-of-date file.
Server and UI
| Command | Opens | Purpose and flags |
|---|---|---|
memo-mcp serve | rw | MCP server on stdio. --metrics-addr 127.0.0.1:PORT exposes /metrics on loopback (default $MEMO_METRICS_ADDR, off when empty). Starts a background backfill and reindex for the current model |
memo-mcp ui | ro | Read-only web UI. --addr (default 127.0.0.1:0, a random port; the URL is printed), --allow-remote permits a non-loopback address, --no-model gives keyword-only search |
Writing
| Command | Opens | Purpose and flags |
|---|---|---|
memo-mcp ingest <file|dir|-> | rw | Add markdown documents. --ns (default default), --kind doc|note|code|conversation, --uri, --title, --library, --version, --trust user|curated (default user), --origin web|user-said|agent-derived, --context "<sentence>", --embed=false queues vectors instead of embedding now |
memo-mcp remember "<fact>" | rw | Record a fact. --ns, --about a,b, --valid-from, --valid-to (YYYY-MM-DD), --supersedes memo://fact/.., --evidence memo://chunk/n, --trust user|curated, --origin, --no-model |
memo-mcp forget <memo://doc/..|memo://fact/..> | rw | Retire a record. --reason "<why>" (required), --redact also erases the text |
memo-mcp trust ls | ro | List records above agent trust |
memo-mcp trust promote <uri> --to user|curated | rw | Raise trust (human channel; audited) |
memo-mcp trust demote <uri> --to agent|user | rw | Lower trust (audited) |
Reading and search
| Command | Opens | Purpose and flags |
|---|---|---|
memo-mcp search "<q>" | rw, no create | Hybrid search. --mode auto|hybrid|keyword|exact|semantic, --ns a,b, --library, --version, --kind a,b, --min-trust agent|user|curated, --limit (10), --granularity chunk|document, --format table|json|md, --explain, --max-tokens (8000), --profile (default $MEMO_PROFILE), --rerank (default $MEMO_RERANK), --no-model |
memo-mcp explain "<q>" [<memo://...>] | rw, no create | search --explain; with an address, the full explanation for that one result. Same flags as search |
memo-mcp read <memo://...> | ro | Print any record with its provenance. --history prints the revision or supersession chain |
memo-mcp ls | ro | Live documents, newest first. --ns, --kind, --since YYYY-MM-DD, --limit (50), --json |
memo-mcp facts ls | ro | Facts. --ns, --as-of YYYY-MM-DD, --history includes replaced facts, --json |
memo-mcp explore <name> | ro | Walk the graph from one entity. --ns, --hops 1|2, --as-of, --json |
memo-mcp export --md <dir> | ro | Markdown files with front matter, plus an _index.md per namespace. --ns |
memo-mcp export --index | ro | A compact index for AGENTS.md and CLAUDE.md. --ns, --library name[@version], --max-bytes (8192) |
Graph and compaction
| Command | Opens | Purpose and flags |
|---|---|---|
memo-mcp graph merges | ro | Merge-candidate review queue. --state open|merged|rejected|all (default open) |
memo-mcp graph merge <id> / memo-mcp graph reject <id> | rw | Decide a merge candidate |
memo-mcp graph rebuild | rw | Re-extract entity mentions. --ns |
memo-mcp compact | rw | Create or list work items. --ns, --kinds page,stale,conflict,merge,duplicate, --lint, --json, --executor ollama, --apply (with an executor; dry run otherwise), --allow-remote |
memo-mcp submit <item-id> | rw | Complete a work item. Pages: --content-file f.md|-, --title. Conflicts: --keep memo://fact/... Merges: --accept or --reject. Any item: --skip, --reason, --dry-run |
memo-mcp lint | ro | Contradictions, orphan entities, missing and stale pages, expired facts. --ns, --json |
memo-mcp pages ls | ro | Curated pages. --ns, --stale, --json |
Maintenance
| Command | Opens | Purpose and flags |
|---|---|---|
memo-mcp status | ro | File, schema version, counts, model, jobs, namespaces |
memo-mcp migrate | rw, no create | Bring the file to this binary’s schema version |
memo-mcp verify | rw, no create | Check documents without chunks, orphan vectors and facts, missing vectors for the current model, and FTS integrity. --repair queues missing vectors and rebuilds broken indexes |
memo-mcp backfill | rw, no create | Embed passages whose vectors are pending |
memo-mcp reindex | rw, no create | Embed every passage lacking a vector for a model. --model <id> (default $MEMO_MODEL) |
Observability
| Command | Opens | Purpose and flags |
|---|---|---|
memo-mcp metrics | ro | Snapshot from the file: gauges, plus per-tool and search statistics from the opt-in logs. --json, --since (24h) |
memo-mcp log tail | rw, no create | Recent searches, then recent tool calls. --n (20) |
memo-mcp log calls | rw, no create | Recent tool calls with latency and outcome. --n (50) |
memo-mcp log show <id> | rw, no create | One logged search with its full trace |
memo-mcp log replay | rw, no create | Logged searches as JSON lines, ready to label as eval queries. --n (200) |
memo-mcp log prune | rw, no create | Keep at most 10,000 rows and 30 days in both logs |
Models, profiles and evaluation
| Command | Opens | Purpose and flags |
|---|---|---|
memo-mcp model ls | ro | Registry: id, dimension, licence, stored vectors, notes; marks the selected model |
memo-mcp model smoke <id>|--all | none | Prove a model loads under the pure-Go backend, and time it |
memo-mcp model pull <id> | none | Download a model and check that it loads |
memo-mcp model use <id> | rw, no create | Record the knowledge base’s default model. Also set MEMO_MODEL in the client config |
memo-mcp model redownload [<id>] | none | Discard the cached copy and download again. Without an id it re-downloads minilm, so pass the id you use |
memo-mcp profiles show [<name>] | none | Every ranking constant with its derivation |
memo-mcp eval | temporary | Benchmark on built-in labelled corpora. --models hash,minilm,…, --profiles default,…|all, --corpus notes|kb|all, --format table|md|json, --explain-failures, --rerank, --agent-proxy |
memo-mcp version | none | Build version, MCP protocol version, Go version, model directory |
Exit codes and output
0on success.1on a runtime error, printed asfatal: …on stderr.2on a usage error.- Results go to stdout, and diagnostics and logs to stderr.
--jsonoutput is stable enough to script against. - The legacy flags
--statsand--redownload-modelstill work for one more release, with a deprecation warning.
profiles.json
$MEMO_HOME/profiles.json changes ranking constants without rebuilding the binary. It is read
at start-up by memo-mcp serve and memo-mcp ui. A missing file is fine. A malformed
file stops the process with the path and the JSON error.
Note. In v1.4, the CLI
search,explainandevalcommands do not readprofiles.json. They only see the built-in profiles. Test an override through the UI, which uses the same pipeline as the server, or by running the server.
Shape
The file is a JSON object mapping a profile name to the fields to change. Only fields that are present are applied.
- An existing name, such as
default, is modified in place. - A new name creates a profile. It starts from
base, or fromdefaultwhenbaseis absent.
{
"default": {
"weights": { "exact": 0.4 }
},
"team-docs": {
"base": "precise",
"weights": { "semantic": 0.6, "keyword": 0.4 },
"cutoff_gap": 0.4,
"recency_kinds": [],
"derivation": "Docs-heavy KB: favour meaning over wording; no recency."
}
}
Select it with MEMO_PROFILE=team-docs in the server’s environment.
Fields
| Field | Type | Meaning | Shipped default |
|---|---|---|---|
base | string | Profile to copy for a new name | default |
weights | map of arm to number | Arm weights; arms: semantic, keyword, exact, fact, entity, graph. Merged into the base’s weights; 0 disables an arm | 0.5 / 0.5 / 0.3 / 0.4 / 0.4 / 0.5 |
rrf_k | integer | RRF damping constant | 60 |
fusion | "rrf" or "minmax" | Rank fusion, or min-max score fusion | rrf |
fetch_depth | integer | Candidates per arm before fusion | 100 |
half_life_days | number | Recency half-life | 90 |
recency_floor | number 0–1 | Fraction of score the oldest item keeps | 0.8 |
recency_kinds | list of kinds | Kinds that age; [] turns recency off | ["note","conversation"] |
cutoff_gap | number 0–1 | Stop before a drop larger than this fraction | 0.5 |
min_results | integer | Never gap-cut below this many results | 3 |
semantic_floor | number | Drop semantic-only candidates below this cosine | 0.30 |
bands | [strong, moderate, weak] | Relevance bands for the generic profile; exactly three numbers | [0.60, 0.45, 0.30] |
derivation | string | Your reason, shown by profiles show | — |
The reranker switch, its depth and graph routing are not overridable. Use the precise profile
as a base to get reranking. A model’s own relevance bands, set in the model registry, take
precedence over bands.
Checking an override
memo-mcp profiles show team-docs # built-in profiles only in v1.4; see the note above
memo-mcp ui # search there with MEMO_PROFILE=team-docs set; the trace shows the profile
Before adopting an override, measure it. See Measuring a change.
Files on disk
| Path | Created by | Mode | Contents |
|---|---|---|---|
$MEMO_HOME/ | First writing command | 0700 | Base directory (default ~/.memo-mcp) |
$MEMO_HOME/kb/ | First writing command | 0700 | Knowledge-base files |
$MEMO_HOME/kb/<name>.db | First writing command | 0600 | The SQLite database: all layers, logs and audit |
$MEMO_HOME/kb/<name>.db-wal | SQLite, while open | 0600 | Write-ahead log, checkpointed into the main file |
$MEMO_HOME/kb/<name>.db-shm | SQLite, while open | 0600 | Shared-memory index for the WAL |
$MEMO_HOME/profiles.json | You | yours | Optional profile overrides |
~/.cache/memo-mcp/models/<org>_<model>/ | First use of a model | 0755 | Model files from Hugging Face |
~/.cache/memo-mcp/models/<org>_<model>.ok | Completed download | — | Marker; without it the download is redone |
A directory that already existed with wider permissions is not tightened.
What is inside the database
| Group | Tables |
|---|---|
| L0–L2 | sources, documents, chunks, chunks_fts and chunks_fts_exact (FTS5), chunk_vecs (one row per passage per model) |
| L3 | facts, facts_fts, fact_vecs |
| L4 | entities, entity_aliases, mentions, merge_candidates, edges |
| L5 | pages, pages_fts, page_sources, page_vecs, work_items |
| Bookkeeping | namespaces, models, jobs, audit, query_log, call_log |
Every column is described in
schema.md. The file is plain
SQLite. You can inspect it with the sqlite3 shell, for example
sqlite3 ~/.memo-mcp/kb/default.db '.tables'. Do not write to it by hand: triggers keep the
indexes consistent only for writes made through memo-mcp.
Nothing else
memo-mcp writes no PID files, no lock files of its own, no logs to disk and no temporary files outside the model cache. Logs go to stderr. Redirect them yourself if you want them on disk.
Part IV: Retrieval tuning
The defaults are measured, not guessed. Each constant has a written derivation, and every release carries an eval report. Tune only when you have a reason, and measure before and after.
- Profiles and constants: the shipped profiles and what each constant does.
- Embedding models: the model registry, switching models, and what it costs.
- The reranker: an opt-in cross-encoder, and why it is off.
- Measuring a change: the eval harness, the query log, and reading
explain.
| Symptom | First lever |
|---|---|
| Identifier or code queries miss | code profile, or raise weights.exact |
| Paraphrased questions miss | Check degraded and the model; try a stronger model |
| Old notes crowd out new ones | recency profile, or a shorter half_life_days |
| Lists are too long or too short | cutoff_gap, min_results, or the request’s limit |
| “No results” when there should be some | semantic_floor, and the model’s bands; look at filtered in the trace |
Profiles and constants
A profile is a named set of ranking constants. The server and the UI use MEMO_PROFILE,
default default. The CLI takes search --profile <name>. MCP clients cannot pick a profile per
request in 1.x.
memo-mcp profiles show # every profile, every constant, with its derivation
memo-mcp profiles show precise
Shipped profiles
| Profile | Weights: sem / kw / exact / fact / entity / graph | Other differences | Use for |
|---|---|---|---|
default | 0.5 / 0.5 / 0.3 / 0.4 / 0.4 / 0.5 | — | General use |
precise | same as default | fetch_depth 200, cutoff_gap 0.6, rerank on when a reranker is attached | Rarely worded answers; shorter, surer lists |
recency | same as default | Every kind ages; half-life 30 days; floor 0.6 | “What did I write lately” |
code | 0.4 / 0.4 / 0.6 / 0.4 / 0.5 / 0.5 | No recency | Identifier-heavy questions |
minmax | same as default | Min-max score fusion instead of RRF | Experiment |
text-only | same as default | No entity or graph arm (the 1.0 arms) | Ablation |
no-graph | same as default | Entity arm without the graph walk | Ablation |
keyword-only | keyword 1.0 only | — | Ablation; the baseline to beat |
semantic-only | semantic 1.0 only | — | Ablation |
The constants
| Constant | Default | Raise it to… | Lower it to… |
|---|---|---|---|
weights.<arm> | see above | Let that arm’s ranking count for more | Mute an arm (0 disables it) |
rrf_k | 60 | Flatten the difference between rank 1 and rank 10 | Reward top ranks more |
fetch_depth | 100 | Keep rarely worded hits in the fused list, at some latency cost | Speed up large KBs |
half_life_days | 90 | Age more slowly | Favour recent items more strongly |
recency_floor | 0.8 | Make age matter less | Make age matter more (floor is the minimum kept) |
recency_kinds | note, conversation | Age more kinds | [] turns recency off |
cutoff_gap | 0.5 | Cut later, giving longer lists | Cut sooner at a score cliff |
min_results | 3 | Always show more before a gap cut | Allow single-answer lists |
semantic_floor | 0.30 | Abstain more readily on vague matches | Keep weaker meaning-only matches |
bands | 0.60 / 0.45 / 0.30 | Generic relevance bands; per-model bands override them | — |
Worked example: fusion
A passage at keyword rank 1 and semantic rank 2 under default:
keyword 0.5 / (60 + 1) = 0.008197
semantic 0.5 / (60 + 2) = 0.008065
fused = 0.016262
A keyword-only hit at rank 1 scores 0.008197, the same as a semantic-only hit at rank 1. Equal weights therefore let either arm put a result on the first page. That is the main reason the default is balanced.
Graph routing
The entity and graph arms run only when the query names two or more known entities, or one
entity with relational phrasing (“depends on”, “related to”, …). The trace’s routing_reason
says why they ran. Ablations show that routing them on every query hurts single-entity
lookups. That is why routing is not a tunable constant.
Customising
Overrides go in profiles.json. The derivations behind each
default are in
architecture §5–§6
and the eval reports.
Embedding models
The semantic arm needs an embedding model. Models run in-process through a pure-Go ONNX
backend (hugot / GoMLX), so no Python, no ONNX Runtime and no GPU are involved. potion is a
static lookup table with no neural network at all.
memo-mcp model ls
| Id | Dim | Licence | Size | Query latency* | Notes |
|---|---|---|---|---|---|
granite-small-r2 (default) | 384 | Apache-2.0 | ~195 MB | ~0.2 s | IBM granite-embedding-small-english-r2; best measured quality; 8k context |
minilm | 384 | Apache-2.0 | ~87 MB | ~0.12 s | The previous default; no paraphrase recall in the bake-off |
potion | 512 | MIT | ~130 MB | ~6 ms | model2vec table; instant; weaker on paraphrase; good fallback |
granite-r2 | 768 | Apache-2.0 | ~600 MB | slow | Runs, but about 12× slower than MiniLM on this backend |
arctic-m-v2 | 768 | Apache-2.0 | ~1.2 GB | slow | Snowflake; multilingual |
gemma-256 | 256 | Gemma Terms of Use | ~310 MB | slow | Opt-in only: not an Apache or MIT licence |
* p50 query latency including query embedding, measured in docs/eval/v0.7.0.md. Writes cost
more: about 0.6 s per passage with granite-small-r2 on this backend.
Each model has its own relevance bands in the registry. For granite-small-r2 they are
0.88 / 0.80 / 0.72, because unrelated text already scores about 0.63 under it.
How vectors are stored
Vectors are stored per passage per model in chunk_vecs(chunk_id, model_id, embedding).
Several models can coexist in one file, so switching back is free once both are embedded.
Each 384-dimension vector adds about 1.5 KB per passage.
Switching models
- Pre-download and check, optionally:
memo-mcp model pull potion memo-mcp model smoke potion - Embed the existing passages. This is optional, because the server does it in the
background; doing it ahead avoids a degraded period:
MEMO_KB=my-project memo-mcp reindex --model potionreindexis resumable. It only embeds passages that lack a vector for that model. - Record the default for the file, and set
MEMO_MODELin the MCP client’s server entry:MEMO_KB=my-project memo-mcp model use potion - Restart the client session. At start, the server runs
backfill, thenreindexfor the current model, in the background. Until that finishes, search reportsdegradedandmemo_kb_pending_embeddings{model="potion"}is above zero.
Downloads
- Source: Hugging Face, over HTTPS, into
~/.cache/memo-mcp/models/. - A completed download writes a
.okmarker. An interrupted one is detected and retried. memo-mcp model redownload <id>discards the cache for one model and fetches it again. Always pass the id: without one it re-downloadsminilm.- Metrics:
memo_embed_model_downloads_total{model,outcome},memo_embed_model_load_seconds,memo_embed_model_loaded.
Offline installs: Air-gapped installs.
The reranker
A cross-encoder can re-score the top results by reading the query and each passage together. memo-mcp ships one as an opt-in experiment:
| Setting | Value |
|---|---|
| Model | cross-encoder/ms-marco-MiniLM-L6-v2 (Apache-2.0, ~91 MB), id ms-marco-minilm |
| Enable | MEMO_RERANK=1 (server and UI), or search --rerank (CLI) |
| Choose another | MEMO_RERANKER=<id> |
| Used by | Profiles with rerank on: precise, which re-scores the top 30 |
Why it is off by default
It failed its promotion gate. In the bake-off it lowered nDCG@10 by about 0.2 and added 0.7 to 4 seconds per query:
| Embedder | Corpus | default nDCG@10 | precise + rerank nDCG@10 | p50 ms |
|---|---|---|---|---|
| minilm | notes | 0.977 | 0.732 | 806 |
| minilm | kb | 0.882 | 0.724 | 2,333 |
| granite-small-r2 | notes | 0.978 | 0.676 | 3,987 |
The pure-Go pipeline cannot pass sentence-pair segment ids, so the model’s scores are compressed near zero. Details: spike S8.
If you try it
MEMO_RERANK=1 MEMO_PROFILE=precise memo-mcp ui
Watch memo_search_rerank_duration_seconds and the trace’s rerank{model, top_n, latency_ms}.
If the reranker fails to load, the server continues without it. Search then reports
degraded with reason reranker_failed.
Measuring a change
Every retrieval change in memo-mcp’s own history was measured against a labelled baseline. Use the same tools before you change a constant, a profile or a model.
1. The built-in benchmark
memo-mcp eval --models hash,granite-small-r2 --profiles default,precise,code --corpus all
memo-mcp eval --models potion --profiles all --format md > potion.md
memo-mcp eval --explain-failures # one line per missed query, and why
It loads two labelled corpora into a throwaway knowledge base, runs every query under each model and profile combination, and prints quality next to cost:
| Column | Meaning |
|---|---|
| recall@1/5/10 | Share of relevant items found in the top k |
| MRR | Mean reciprocal rank of the first relevant item |
| nDCG@10 | Ranking quality, rewarding relevant items near the top |
| abstention | Share of no-answer queries that correctly returned nothing |
| p50 ms, tokens p50 | Median latency and response size |
hash is the deterministic test embedder and needs no download. Other model ids download real
models. The corpora are the project’s own fixtures, so they tell you about relative effects
(this profile against that one), not about your data.
2. Your own queries
- Turn on the opt-in log in the server’s environment:
MEMO_QUERY_LOG=1. - Use the system normally for a while.
- Export what was asked:
Each line is an unlabelled eval query, with the query and the addresses that were returned.memo-mcp log replay --n 200 > candidates.jsonl - Label the relevant addresses, then compare runs under different profiles in the UI or with
memo-mcp search --profile <p> --format json.
3. Reading one result
memo-mcp explain "how do we rotate credentials" # ranking table for every hit
memo-mcp explain "how do we rotate credentials" memo://chunk/812 # the full why for one hit
What to look for:
| Field | Tells you |
|---|---|
arms[] | Which arms saw the result, at what rank, and what each contributed |
recency_factor | Whether age pushed it down |
relevance, band | Whether the model thinks it is about the question at all |
trace filtered | How many candidates scope, revocation, minimum trust or the semantic floor removed |
trace cutoff | Whether the list ended on a gap, the limit or the budget |
trace degraded | Whether a capability was missing |
The UI’s search page shows the same table, with links.
4. In production
After changing a profile or a model, watch these in Prometheus:
memo_search_abstentions_totalandmemo_search_results: is the system answering less?memo_search_duration_seconds: did latency move?memo_search_cutoff_total{kind}: are lists ending differently?
See Prometheus: scrape, rules, alerts.
Part V: Knowledge lifecycle
How records enter, change, gain trust, get connected and summarised, and leave. Agents do most of this through MCP tools; an operator does the parts that need a human through the CLI.
- Ingest and revisions
- Facts and time
- Trust administration
- The graph index
- Compaction and pages
- Forget and redact
Every write in every chapter lands in the audit table: who (actor), through which channel
(tool, cli, elicitation, worker), what operation, on which record. The same events
increment memo_store_writes_total{op,channel}.
Ingest and revisions
The write path
- Normalise line endings and whitespace, and hash the content (SHA-256).
- Deduplicate. If the source’s current revision has the same hash, nothing happens and the
result says
dedup: true. If the content differs, a new revision is created. The previous revision is marked superseded and stays readable withas_of. - Chunk. The text is split by headings, then paragraphs, then sentences, then a hard
split. The target is about 200 estimated tokens and the cap 400, clamped to the model’s
limit, with about 40 tokens of overlap. Fenced code blocks are never split. Each passage gets
a context header
title > section paththat the keyword index also sees. - Index. Both FTS5 indexes are filled by triggers inside the transaction. Entity mentions are extracted and linked.
- Embed outside the transaction, in batches of 16. If no model is available, the
passages are queued as a job, and
backfillor the next server start embeds them. A write never reports success it did not get. - Audit: one row per write.
A source is identified by its URI, or by its file path for CLI ingests. Re-ingesting the
same URI with changed content, or with a new version, creates a new revision of that source.
From the CLI
memo-mcp ingest ./docs --ns handbook --kind doc
memo-mcp ingest page.md --uri https://pkg.go.dev/google.golang.org/grpc \
--library grpc/grpc-go --version v1.64.0 --context "gRPC-Go API docs for deadlines"
cat notes.md | memo-mcp ingest - --kind note --title "Standup 2026-10-03"
memo-mcp ingest ./big-folder --embed=false && memo-mcp backfill # fast load, embed later
CLI writes are trust user, or curated with --trust curated. Origin defaults to web
when --uri is set, and to user-said otherwise.
Kinds
| Kind | Ages under recency? | Typical content |
|---|---|---|
doc | No | Documentation, specs, handbooks; usually versioned |
code | No | Code explanations and snippets |
note | Yes | Short notes, decisions, observations |
conversation | Yes | Conversation summaries |
Costs
| Step | Cost |
|---|---|
| Parse, chunk, index | About 60 documents per second without embedding |
Embedding with granite-small-r2 | About 0.6 s per passage on the pure-Go backend |
Embedding with potion | Near instant |
| Disk | About 4.4 KB per passage before vectors, plus about 1.5 KB per passage per 384-dimension model |
For bulk loads, ingest with --embed=false, then run memo-mcp backfill, or let the server
embed in the background. See Capacity and performance.
Observing it
memo_store_ingests_total{outcome=new|revision|dedup|error},
memo_store_ingest_duration_seconds, memo_store_chunks_written_total,
memo_store_embed_batches_total{outcome}, and memo_kb_pending_embeddings{model} for the
backlog.
Facts and time
A fact is one sentence with an optional subject list, an evidence passage and a validity window. Facts are searched by their own arm. A matching fact votes for its evidence passage.
memo-mcp remember "Production runs Postgres 16" --about postgres,production \
--evidence memo://chunk/812 --valid-from 2026-06-01
memo-mcp facts ls --ns default
memo-mcp facts ls --as-of 2026-05-15 --history
Correcting facts
- Supersede:
remember "<new>" --supersedes memo://fact/<old>. The old fact is invalidated, not deleted, and remains visible underas_ofand--history. Agents may supersede onlyagent-trust facts. - Expire: set
--valid-to.memo-mcp lintlists facts past their window. - Retire:
memo-mcp forget memo://fact/<id> --reason "…".
Two clocks
| Clock | Columns | Question it answers |
|---|---|---|
| Recorded time | recorded_at, invalidated_at | When did the knowledge base learn or drop this? |
| Valid time | valid_from, valid_to | When is this true in the world? |
as_of=T, available on the search and explore tools and on facts ls --as-of and
explore --as-of in the CLI, shows what was recorded on or before T and not superseded before
T. For facts with a window, it also applies valid time.
Conflicts
Two live facts about the same subject, with overlapping windows and differing statements, form
a conflict. compact turns each conflict into a work item, together with the rule the
server would apply: higher trust wins, then the newer fact. A human or agent settles it with
submit <item> --keep memo://fact/<id>. When conflicting facts both match a search and score
within 10% of each other, trust breaks the tie.
Trust administration
| From → To | MCP tool call | CLI (human) | Elicitation dialog (human) |
|---|---|---|---|
create agent | yes | — | — |
create user / curated | no | yes (--trust) | — |
agent → user / curated | no; promote asks | trust promote | yes, when accepted |
| lower trust | no | trust demote | — |
forget or supersede agent records | yes | yes | — |
forget or supersede user / curated records | no | yes | — |
Commands
memo-mcp trust ls # everything above agent trust
memo-mcp trust promote memo://doc/0199… --to curated
memo-mcp trust demote memo://fact/0199… --to agent
The promote tool
When an agent calls promote:
- If the client supports elicitation, it shows the human a dialog with the excerpt, the source
and the target level. Accepting it applies the change through the
elicitationchannel. Declining or cancelling changes nothing. - If the client does not support elicitation, the tool returns the exact
memo-mcp trust promote …command for a human to run.
Outcomes are counted in
memo_mcp_elicitations_total{tool="promote",outcome=asked|accept|decline|cancel|unsupported}.
Warning. Hooks or settings that auto-accept elicitation dialogs defeat this gate. If you run such a client, do not rely on trust levels.
Auditing trust
Every change is an audit row, with op promote or demote and the channel it came
through. Inspect it with sqlite3:
sqlite3 ~/.memo-mcp/kb/default.db \
"SELECT datetime(ts/1000,'unixepoch'), actor, channel, op, target_uri FROM audit WHERE op IN ('promote','demote') ORDER BY ts DESC LIMIT 20"
The graph index
The graph layer is an index, not a source of truth. It records which entities each passage
mentions. Search uses it to find passages that connect things the query names, and explore
uses it to walk from one name.
Extraction and resolution
- Extraction runs on every ingest, using heuristics only: backticked spans, dotted, camel
and snake identifiers, headings and capitalised names. Clients may also declare
entities[]andrelations[]when they ingest. - Resolution is deterministic. The same normalised key means the same entity. A near key, with 3-gram Jaccard of at least 0.8, not too short and not differing only in a number, becomes a merge candidate for a human. Anything else is a new entity.
Commands
memo-mcp explore "Quorum Replication" --hops 2
memo-mcp graph merges # open candidates with their similarity
memo-mcp graph merge 17 # same thing: merge the names
memo-mcp graph reject 18 # different things: keep apart
memo-mcp graph rebuild --ns default # re-extract mentions
Nothing is merged without a decision. Agents can propose decisions through compact and
submit. Merge precision is gated at 0.95 or better in CI.
When to rebuild
- Once, after upgrading a file written before 1.1.
- After a release whose notes mention an extraction change.
- If
memo-mcp lintreports many orphan entities after large deletions.
Rebuild keeps merge decisions. It re-reads every live passage, so on large knowledge bases run it when the system is quiet.
Runtime behaviour
The server keeps an in-memory mention graph per namespace. It is built on first use and
reused while the namespace’s count of live mentions is unchanged; any ingest, revision or
forget that changes that count triggers a rebuild on the next routed query. as_of queries
always build a graph for that point in time and do not cache it.
memo_graph_cache_total{event=hit|build|build_asof} and memo_graph_build_duration_seconds
show how often this happens and what it costs.
Compaction and pages
The server never writes prose. Compaction is a queue of work items that an agent, or a person, completes. The server checks the result before storing it.
Work items
compact, through the tool or memo-mcp compact, scans a namespace and creates items:
| Kind | Created when |
|---|---|
page | An entity has 3 or more live passages and no page |
stale | A page’s sources changed or were forgotten, or the entity gained 2 or more passages since the page was built |
conflict | Two live facts about one subject overlap in time and disagree |
merge | An open merge candidate exists |
duplicate | Two passages from different documents are near duplicates (word 3-gram Jaccard of at least 0.75) |
Each item’s payload carries everything needed to do it: the passages, the facts and the previous page. No second call is needed.
Submitting
memo-mcp compact --ns default --kinds page,stale --json
memo-mcp submit 42 --content-file page.md --title "Quorum replication" --dry-run
memo-mcp submit 42 --content-file page.md --title "Quorum replication"
memo-mcp submit 43 --keep memo://fact/0199…
memo-mcp submit 44 --accept # merge item
memo-mcp submit 45 --skip --reason "not worth a page"
Every page submission reports:
- diff: a line diff against the previous page.
- omission check: recorded facts about the subject whose content words are mostly absent from the page. Less than 75% coverage is flagged.
- corruption check: page sentences that share fewer than half of their content words with any source.
Pages are stored as is_inference = 1, with the passages they cite. They are searchable with
granularity: page, and marked stale when a source changes. Raw rows are never modified.
Lint
memo-mcp lint --json
Reports contradictions, orphan entities, missing pages, stale pages and expired facts, with addresses. Lint changes nothing.
The optional Ollama executor
For unattended page writing from the terminal:
MEMO_OLLAMA_MODEL=qwen2.5:7b-instruct memo-mcp compact --kinds page --executor ollama # dry run
MEMO_OLLAMA_MODEL=qwen2.5:7b-instruct memo-mcp compact --kinds page --executor ollama --apply # submit
- The endpoint is
MEMO_OLLAMA_URL, defaulthttp://127.0.0.1:11434. It must be loopback unless--allow-remoteis passed. - Use a non-thinking instruction model. Every page still goes through the same checks.
- No MCP code path can reach the executor.
Observing it
memo_kb_work_items_open, memo_kb_pages, memo_kb_pages_stale,
memo_store_work_items_total{kind,event} and memo_store_pages_marked_stale_total.
Forget and redact
memo-mcp forget memo://doc/0199… --reason "superseded by the 2026 policy"
memo-mcp forget memo://fact/0199… --reason "wrong owner" --redact
forget | forget --redact | |
|---|---|---|
| Leaves every index (keyword, exact, vectors, graph) | yes | yes |
Excluded from search, including under as_of | yes | yes |
| Address resolves to “forgotten on … because …” | yes | yes |
| Text kept in the file | yes | no: cleared |
| Pages built from it | marked stale | marked stale |
| Audit row | yes | yes |
--reasonis required and is shown to anyone who dereferences the address later.- MCP tool calls may forget only
agent-trust records. The CLI may forget anything.
Reclaiming disk space
SQLite does not shrink the file when rows are cleared. After large redactions, compact it while no process has the file open:
sqlite3 ~/.memo-mcp/kb/my-project.db 'VACUUM'
VACUUM also removes freed pages that might still hold redacted bytes. Run it after redacting
secrets.
Part VI: Monitoring and observability
All observability in memo-mcp is local and pull-based. Metrics are served on loopback, logs go to stderr, and the optional call log is a table in your own file. Nothing is pushed anywhere, and there is no telemetry.
- Signals overview: what exists, where, and the per-process caveat.
- The metrics endpoint: enabling
/metricson the server and the UI. - Metric catalogue: every series, its type, labels and meaning.
- Prometheus: scrape, rules, alerts: ready-to-adapt configuration.
- Logs: format, level, fields, and where they end up.
- Call and query logs: per-call history inside the knowledge base.
Signals overview
| Signal | Where | On by default | Scope |
|---|---|---|---|
| Metrics (Prometheus text 0.0.4) | serve --metrics-addr and UI GET /metrics | In memory, always; exposed only when asked | The process that serves them |
Logs (log/slog, text or JSON) | stderr | Yes, at info | The process |
| Call log and query log | call_log and query_log tables | No (MEMO_QUERY_LOG=1) | The file: every process writing to it |
Snapshot (memo-mcp metrics) | stdout | On demand | The file: gauges plus log statistics |
| Audit | audit table | Always | The file: every write |
The per-process caveat
The MCP server, the UI and each CLI command are separate processes. Counters such as
memo_mcp_tool_calls_total live in the memory of the process that served the calls:
- A server process exists only while its client session is open. Its counters start at
zero with each session. Scrape it while it runs; Prometheus
rate()handles the resets. - The UI has its own counters, mostly
memo_ui_*. Like every process, it also serves thememo_kb_*gauges, which are read from the file. - History across sessions comes from the file:
memo-mcp metrics(withMEMO_QUERY_LOG=1) and thememo_kb_*gauges.
Recommended setups
| Goal | Setup |
|---|---|
| Glance at health now | memo-mcp metrics, or memo-mcp status |
| Trends of knowledge-base size, backlog and pages | Run memo-mcp ui --addr 127.0.0.1:9470 as a long-lived process and scrape its /metrics for memo_kb_* |
| Live tool and search latency, errors, degraded mode | Add --metrics-addr to one server entry and scrape it while sessions run |
| Per-call history and slow-call forensics | MEMO_QUERY_LOG=1, then memo-mcp log calls and memo-mcp metrics --since 7d |
Naming
- Prefix
memo_; counters end in_total; units are suffixes (_seconds,_bytes), and seconds are never milliseconds. - Labels are bounded enums: tool names, outcomes, arms, models, namespaces. Never query text, ids or paths.
- OpenTelemetry mapping: replace
_with.and drop_total. For example,memo_mcp_tool_calls_totalmaps tomemo.mcp.tool_calls.
The metrics endpoint
On the MCP server
memo-mcp serve --metrics-addr 127.0.0.1:9469
# or, in the client's server entry: "env": { "MEMO_METRICS_ADDR": "127.0.0.1:9469" }
Claude Code example:
claude mcp add memo --env MEMO_KB=my-project --env MEMO_METRICS_ADDR=127.0.0.1:9469 -- /path/to/memo-mcp
At start, the server logs metrics listening url=http://127.0.0.1:9469/metrics.
| Property | Behaviour |
|---|---|
| Bind | Loopback only (127.0.0.1, ::1, localhost). Any other address is refused at start; there is no override |
| Methods | GET and HEAD; anything else returns 405 |
| Paths | /metrics only; anything else returns 404 |
| Host header | Must match the bound address, which defends against DNS rebinding; otherwise 403 |
| Content type | text/plain; version=0.0.4; charset=utf-8 |
| stdout | Untouched: the MCP stream stays clean |
Warning: one port, one process. The address is bound when the server starts. If a second session starts with the same entry while the first is running, for example two Claude Code windows on one project, the second fails with
fatal: listen tcp 127.0.0.1:9469: bind: address already in use. Until this is relaxed:
- set
--metrics-addron one entry you use for a single long session;- or give each knowledge base’s entry its own port;
- or leave it off and use the UI’s endpoint for file-level gauges plus the call log for per-call data.
On the UI
memo-mcp ui always serves GET /metrics, under the same guards as the rest of the UI:
loopback, GET only, and the Host check. Pin the port to scrape it:
MEMO_KB=my-project memo-mcp ui --addr 127.0.0.1:9470 --no-model
curl -s http://127.0.0.1:9470/metrics | grep memo_kb_documents_live
--no-model keeps a long-running UI light. Its search is then keyword-only, which is fine for a
metrics sidecar.
Checking it
curl -s http://127.0.0.1:9469/metrics | head -20
curl -s -o /dev/null -w '%{http_code}\n' -X POST http://127.0.0.1:9469/metrics # 405
curl -s -o /dev/null -w '%{http_code}\n' -H 'Host: example.com' http://127.0.0.1:9469/metrics # 403
memo_kb_* gauges are read from the database at scrape time and cached for 5 seconds, so
frequent scrapes do not load the file. Read failures increment
memo_store_status_scrape_errors_total.
Metric catalogue
Every series memo-mcp exposes, as of v1.4. Types: C counter, G gauge, H histogram.
A histogram exposes _bucket{le}, _sum and _count. A family with no observations yet
shows only its # HELP and # TYPE lines.
Bucket sets:
| Set | Boundaries |
|---|---|
| latency (s) | 0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10 |
| tokens | 50, 100, 200, 500, 1000, 2000, 4000, 8000, 16000 |
| count | 0, 1, 2, 5, 10, 20, 50, 100, 200, 500 |
| bytes | 1 KiB, 4 KiB, 16 KiB, 64 KiB, 256 KiB, 1 MiB, 4 MiB |
MCP server
| Metric | Type | Labels | Meaning |
|---|---|---|---|
memo_mcp_requests_total | C | method, outcome = ok, error | Every MCP request received (tools/call, resources/read, server/discover, …) |
memo_mcp_requests_in_flight | G | — | Requests being served now |
memo_mcp_tool_calls_total | C | tool, outcome = ok, tool_error, input_required, error | Tool calls. input_required is promote waiting for a human |
memo_mcp_tool_call_duration_seconds | H latency | tool | Tool call latency |
memo_mcp_tool_result_tokens | H tokens | tool | Estimated tokens in each tool result: the cost the agent pays |
memo_mcp_tool_errors_total | C | tool, class = not_found, forgotten, needs_human, canceled, invalid_args, internal | Why calls failed |
memo_mcp_elicitations_total | C | tool, outcome = asked, accept, decline, cancel, unsupported | Human approval dialogs |
memo_mcp_resource_reads_total | C | kind = doc, chunk, source, fact, entity, page, index, ns_index, unknown; outcome = ok, error | memo:// resource reads |
memo_mcp_sessions_total | C | client | Sessions started, by client name |
Search
| Metric | Type | Labels | Meaning |
|---|---|---|---|
memo_search_total | C | mode_requested, mode_resolved, granularity, outcome = results, abstain | Searches that completed (failed searches show up as tool errors) |
memo_search_duration_seconds | H latency | granularity | End-to-end search latency |
memo_search_arm_duration_seconds | H latency | arm | Time in each retrieval arm |
memo_search_arm_candidates | H count | arm | Candidates each arm returned before fusion |
memo_search_results | H count | granularity | Results returned per search |
memo_search_truncated_results_total | C | — | Ranked results left out by the token budget |
memo_search_cutoff_total | C | kind = gap, budget, limit, none | Why lists ended |
memo_search_degraded_total | C | reason = no_embedder, query_embed_failed, reranker_failed, other | Searches served without a capability |
memo_search_abstentions_total | C | reason | Searches that returned nothing on purpose |
memo_search_entities_matched | H count | — | Known entities named per query |
memo_search_rerank_duration_seconds | H latency | model | Cross-encoder time, only with the reranker |
memo_graph_cache_total | C | event = hit, build, build_asof | In-memory mention-graph cache |
memo_graph_build_duration_seconds | H latency | — | Time to build a namespace graph |
Store (write path)
| Metric | Type | Labels | Meaning |
|---|---|---|---|
memo_store_writes_total | C | op, channel = tool, cli, elicitation, worker | Every audited write |
memo_store_ingests_total | C | outcome = new, revision, dedup, error | Ingest calls |
memo_store_ingest_duration_seconds | H latency | — | Ingest latency, embedding included |
memo_store_chunks_written_total | C | — | Passages written |
memo_store_document_bytes | H bytes | — | Size of ingested documents |
memo_store_vectors_stored_total | C | model | Vectors stored |
memo_store_embed_batches_total | C | outcome = ok, error, queued | Embedding batches on the write path |
memo_store_jobs_total | C | kind, event = queued, done, failed | Background jobs, for example the embed backlog |
memo_store_mentions_linked_total | C | — | Entity mentions linked to passages |
memo_store_pages_marked_stale_total | C | — | Pages marked stale by a source change or forget |
memo_store_work_items_total | C | kind, event = created, done, skipped | Compaction work items |
memo_store_status_scrape_errors_total | C | — | Failures reading the gauges below |
Knowledge base (gauges read from the file, cached for 5 s)
| Metric | Labels | Meaning |
|---|---|---|
memo_kb_sources | — | Sources |
memo_kb_documents_live | — | Live documents (latest revision, not forgotten) |
memo_kb_revisions | — | Document revisions stored |
memo_kb_chunks | — | Live passages |
memo_kb_facts | — | Live facts |
memo_kb_entities | — | Entities |
memo_kb_mentions | — | Entity mentions |
memo_kb_edges | — | Typed edges |
memo_kb_merge_review_open | — | Open merge candidates awaiting a human |
memo_kb_pages | — | Curated pages |
memo_kb_pages_stale | — | Stale pages |
memo_kb_work_items_open | — | Open compaction work items |
memo_kb_jobs_queued | — | Jobs queued or running |
memo_kb_jobs_failed | — | Jobs failed |
memo_kb_pending_embeddings | model | Passages without a vector for that model |
memo_kb_schema_version | — | Schema version of the file |
memo_kb_db_size_bytes | — | Size of the main SQLite file, excluding the WAL |
memo_kb_last_write_timestamp_seconds | — | Unix time of the last audited write |
memo_kb_namespace_documents | namespace | Live documents per namespace |
memo_kb_namespace_chunks | namespace | Live passages per namespace |
All are gauges.
Embedding
| Metric | Type | Labels | Meaning |
|---|---|---|---|
memo_embed_duration_seconds | H latency | model, role = query, doc | Embedding call latency |
memo_embed_texts_total | C | model, role | Texts embedded |
memo_embed_errors_total | C | model, role | Embedding calls that failed |
memo_embed_batch_size | H count | model | Texts per call |
memo_embed_model_downloads_total | C | model, outcome = ok, error, corrupt_retry | Model downloads |
memo_embed_model_load_seconds | G | model | Time the last model load took |
memo_embed_model_loaded | G | model, backend = hugot, static | 1 while a model is loaded in this process |
Web UI
| Metric | Type | Labels | Meaning |
|---|---|---|---|
memo_ui_requests_total | C | route (for example GET /search), status | UI requests |
memo_ui_request_duration_seconds | H latency | route | UI render time |
Runtime and build
| Metric | Type | Labels | Meaning |
|---|---|---|---|
memo_build_info | G (always 1) | version, go_version, mcp_protocol, goos, goarch | What is running |
go_info | G | version | Go version |
go_goroutines | G | — | Goroutines |
go_memstats_heap_alloc_bytes | G | — | Live heap |
go_memstats_sys_bytes | G | — | Memory obtained from the OS |
go_gc_cycles_total | C | — | Completed GC cycles |
go_gc_pauses_seconds | H | — | GC pause distribution |
process_start_time_seconds | G | — | Process start time; use it to spot session restarts |
Prometheus: scrape, rules, alerts
These are starting points written against the metric catalogue. Adapt the
thresholds to your usage. The scrape configuration and both rule groups pass
promtool check config and promtool check rules (Prometheus 3.15). Re-check them after you
edit them.
Scrape configuration
scrape_configs:
# Long-lived UI sidecar: knowledge-base gauges, always up.
- job_name: memo-kb
scrape_interval: 60s
static_configs:
- targets: ["127.0.0.1:9470"]
labels: { kb: my-project }
# MCP server: live tool and search metrics; only up while a client session runs.
- job_name: memo-mcp
scrape_interval: 15s
static_configs:
- targets: ["127.0.0.1:9469"]
labels: { kb: my-project }
Prometheus must run on the same machine, because both endpoints are loopback-only. A Grafana Alloy or OpenTelemetry Collector agent on the machine can scrape them and forward the samples. That forwarding is your choice; memo-mcp itself never pushes.
Do not alert on
up{job="memo-mcp"} == 0. The server exists only while a client session is open, so being down is normal. Alert onup{job="memo-kb"}instead, if you rely on the sidecar.
Recording rules
groups:
- name: memo-recording
interval: 1m
rules:
- record: memo:tool_calls:rate5m
expr: sum by (kb, tool, outcome) (rate(memo_mcp_tool_calls_total[5m]))
- record: memo:tool_error_ratio:rate5m
expr: |
sum by (kb) (rate(memo_mcp_tool_calls_total{outcome="error"}[5m]))
/
sum by (kb) (rate(memo_mcp_tool_calls_total[5m]))
- record: memo:tool_latency_seconds:p95_5m
expr: histogram_quantile(0.95, sum by (kb, tool, le) (rate(memo_mcp_tool_call_duration_seconds_bucket[5m])))
- record: memo:search_latency_seconds:p95_5m
expr: histogram_quantile(0.95, sum by (kb, le) (rate(memo_search_duration_seconds_bucket[5m])))
- record: memo:search_abstain_ratio:rate15m
expr: |
sum by (kb) (rate(memo_search_total{outcome="abstain"}[15m]))
/
sum by (kb) (rate(memo_search_total[15m]))
- record: memo:arm_latency_seconds:p95_5m
expr: histogram_quantile(0.95, sum by (kb, arm, le) (rate(memo_search_arm_duration_seconds_bucket[5m])))
Alerts
groups:
- name: memo-alerts
rules:
- alert: MemoToolErrorsHigh
expr: memo:tool_error_ratio:rate5m > 0.05
for: 10m
labels: { severity: warning }
annotations:
summary: "memo-mcp {{ $labels.kb }}: more than 5% of tool calls fail"
description: "Check memo_mcp_tool_errors_total by class and the server's stderr."
- alert: MemoSearchSlow
expr: memo:search_latency_seconds:p95_5m > 2
for: 15m
labels: { severity: warning }
annotations:
summary: "memo-mcp {{ $labels.kb }}: p95 search latency above 2 s"
description: "Look at memo:arm_latency_seconds:p95_5m to find the slow arm."
- alert: MemoSearchDegraded
expr: sum by (kb, reason) (increase(memo_search_degraded_total[30m])) > 0
for: 30m
labels: { severity: warning }
annotations:
summary: "memo-mcp {{ $labels.kb }}: searches degraded ({{ $labels.reason }})"
description: "The embedding model is missing or failing; search is keyword-only."
- alert: MemoEmbeddingBacklog
expr: max by (kb, model) (memo_kb_pending_embeddings) > 0
for: 2h
labels: { severity: info }
annotations:
summary: "{{ $value }} passages lack vectors for {{ $labels.model }}"
description: "Run memo-mcp backfill, or start a server session to drain the backlog."
- alert: MemoJobsFailed
expr: max by (kb) (memo_kb_jobs_failed) > 0
labels: { severity: warning }
annotations:
summary: "memo-mcp {{ $labels.kb }}: background jobs failed"
- alert: MemoDatabaseGrowth
expr: delta(memo_kb_db_size_bytes[1d]) > 500e6
labels: { severity: info }
annotations:
summary: "memo-mcp {{ $labels.kb }}: knowledge base grew more than 500 MB in a day"
- alert: MemoMergeQueueBacklog
expr: max by (kb) (memo_kb_merge_review_open) > 50
for: 1d
labels: { severity: info }
annotations:
summary: "{{ $value }} merge candidates waiting for review (memo-mcp graph merges)"
- alert: MemoKBSidecarDown
expr: up{job="memo-kb"} == 0
for: 10m
labels: { severity: info }
annotations:
summary: "memo-mcp UI sidecar for {{ $labels.kb }} is not running"
Dashboard queries
| Panel | PromQL |
|---|---|
| Tool calls by tool | sum by (tool) (rate(memo_mcp_tool_calls_total[5m])) |
| Tool errors by class | sum by (class) (rate(memo_mcp_tool_errors_total[15m])) |
| p50 and p95 tool latency | histogram_quantile(0.5, sum by (le) (rate(memo_mcp_tool_call_duration_seconds_bucket[5m]))) (and 0.95) |
| Result size the agent pays for | histogram_quantile(0.9, sum by (tool, le) (rate(memo_mcp_tool_result_tokens_bucket[15m]))) |
| Search outcomes | sum by (outcome) (rate(memo_search_total[15m])) |
| Mode resolution | sum by (mode_resolved) (rate(memo_search_total[1h])) |
| Where time goes, per arm | histogram_quantile(0.95, sum by (arm, le) (rate(memo_search_arm_duration_seconds_bucket[5m]))) |
| Why lists end | sum by (kind) (rate(memo_search_cutoff_total[1h])) |
| Budget truncation | rate(memo_search_truncated_results_total[1h]) |
| Knowledge-base size | memo_kb_documents_live, memo_kb_chunks, memo_kb_facts |
| Per namespace | memo_kb_namespace_documents |
| Backlog | memo_kb_pending_embeddings, memo_kb_jobs_queued, memo_kb_work_items_open, memo_kb_pages_stale |
| File size | memo_kb_db_size_bytes |
| Writes by channel | sum by (channel) (rate(memo_store_writes_total[1h])) |
| Human approvals | sum by (outcome) (increase(memo_mcp_elicitations_total[1d])) |
| Embedding cost | histogram_quantile(0.95, sum by (model, role, le) (rate(memo_embed_duration_seconds_bucket[5m]))) |
| Session restarts | changes(process_start_time_seconds{job="memo-mcp"}[1d]) |
| Version running | memo_build_info |
Logs
memo-mcp logs with Go’s log/slog, to stderr only. In server mode, stdout carries the MCP
JSON-RPC stream and is never written to by anything else; a test enforces this.
| Variable | Values | Default |
|---|---|---|
MEMO_LOG_FORMAT | text, json | text |
MEMO_LOG_LEVEL | debug, info, warn, error | info |
An unknown value falls back to the default and prints one warning.
What is logged
| Level | Event | Fields |
|---|---|---|
| INFO | tool call: one per MCP tool call | tool, client, latency_ms, outcome, n_results, tokens_out, and error_class and err on failure |
| INFO | search: one per search, including CLI and UI searches | mode, resolved, granularity, outcome, n_results, truncated, cutoff, latency_ms, profile, and degraded and entities when present |
| DEBUG | search detail | routing, latency_ms_per_arm, candidates_per_arm, scope |
| INFO | Start-up and background work | metrics listening, backfilled chunk vectors, embedded passages for the current model, reranker loaded |
| WARN | Degraded or deprecated | embedding unavailable…, backfill failed, reindex failed, reranker unavailable…, download retries, licence notes, deprecated variables and flags |
| ERROR | Unexpected failures | — |
Logs never contain document text, fact statements, or forget reasons. Search log lines do not include the query text.
Example, with MEMO_LOG_FORMAT=json:
{"time":"2026-10-03T16:23:45.136+02:00","level":"INFO","msg":"search","mode":"auto","resolved":"semantic+keyword+fact","granularity":"chunk","outcome":"results","n_results":1,"truncated":0,"cutoff":"none","latency_ms":"461.8","profile":"default"}
latency_ms is a string with one decimal place.
Where logs end up
-
Claude Code shows a stdio server’s stderr. Run
claude --debugto see it live. -
Claude Desktop writes each server’s stderr to its own log file. On macOS this is
~/Library/Logs/Claude/mcp-server-memo.log. -
CLI commands print logs to your terminal’s stderr. Every
memo-mcp searchprints its INFOsearchline. UseMEMO_LOG_LEVEL=warnfor quiet interactive use, or2>/dev/null. -
To capture a server’s logs yourself, wrap the binary in a small script:
#!/bin/sh exec /path/to/memo-mcp "$@" 2>>"$HOME/.memo-mcp/serve.log"Point the client’s
commandat the script. Rotate the file with your usual tool.
MCP logging capability
memo-mcp does not advertise the MCP logging capability. It is deprecated in protocol
2026-07-28, and stdio hosts already surface stderr. Do not expect log notifications over the
protocol.
Call and query logs
With MEMO_QUERY_LOG=1 in a process’s environment, that process records:
- every search in
query_log: arguments, the addresses returned, scores and the trace; - every MCP tool call in
call_log: client, tool, latency, outcome, error class, result count, tokens out and an argument summary.
Both tables live in the knowledge-base file, so they persist across sessions and processes. They are off by default.
What is stored, and what never is
| Stored | Never stored |
|---|---|
Search query text and its scope, mode, granularity and as_of | Passage, document or page text returned |
| Returned addresses and scores | content of ingest (only its length, as content_len) |
| Tool name, client name, latency, outcome, error class | statement of remember |
Allowlisted arguments: ids, namespaces, flags, dates, and for ingest the source object (URI, title, kind, library, version) | reason of forget and submit, and context |
The argument summary is an allowlist per tool, so a new argument is not logged until it is added explicitly. Search queries are stored verbatim. Treat the logs as sensitive as the questions people ask.
Reading the logs
memo-mcp log tail --n 20 # recent searches, then recent tool calls
memo-mcp log calls --n 50 # tool calls: time, client, tool, latency, results, tokens, args
memo-mcp log show 128 # one search with its full trace
memo-mcp log replay --n 200 # searches as JSON lines, for building eval queries
memo-mcp metrics --since 7d # per-tool calls, errors, p50/p95/max latency; search aggregates
memo-mcp metrics --json | jq .
The UI’s /log page shows both tables.
Retention
The logs are not pruned automatically. Prune them on a schedule:
memo-mcp log prune # both logs: keep at most 10,000 rows and nothing older than 30 days
For example, with cron:
0 3 * * * MEMO_KB=my-project /path/to/memo-mcp log prune
Querying with SQL
sqlite3 -header -column ~/.memo-mcp/kb/my-project.db "
SELECT tool, COUNT(*) AS calls, SUM(ok = 0) AS errors, ROUND(AVG(latency_ms)) AS avg_ms
FROM call_log WHERE ts > (strftime('%s','now','-1 day') * 1000)
GROUP BY tool ORDER BY calls DESC"
ts is in Unix milliseconds. Column reference:
schema §5.8.
Part VII: Operations
Day-two work: keeping data safe, checking integrity, planning capacity, getting data out, and fixing things.
- Backup and restore
- Integrity and repair
- Capacity and performance
- Exporting
- Troubleshooting runbook
- Release engineering
Backup and restore
A knowledge base is one SQLite file in WAL mode. Back it up with SQLite’s own tools, not with a
plain file copy while processes have it open: a copy of the main file alone can miss committed
transactions that are still in the -wal file.
While memo-mcp is running (safe online backup)
DB=~/.memo-mcp/kb/my-project.db
sqlite3 "$DB" ".backup '/backups/my-project-$(date +%F).db'"
# or, which also compacts the copy:
sqlite3 "$DB" "VACUUM INTO '/backups/my-project-$(date +%F).db'"
Both take a consistent snapshot while readers and writers continue. Writers wait at most
memo-mcp’s 5-second busy timeout. sqlite3 creates the copy with your umask, typically 0644.
Restrict it with chmod 600, because the backup holds everything the knowledge base holds.
While nothing is running
When no memo-mcp process has the file open, -wal and -shm are empty or absent, and copying
<name>.db is enough. If a -wal file with content remains, copy it alongside, or run
sqlite3 <name>.db 'PRAGMA wal_checkpoint(TRUNCATE)' first.
What to back up
| Path | Back up? |
|---|---|
$MEMO_HOME/kb/*.db | Yes: everything lives here, including logs and audit |
$MEMO_HOME/profiles.json | Yes, if you use it |
~/.cache/memo-mcp/models/ | No: it is re-downloadable. Keep a copy only for air-gapped machines |
Restore
- Stop every process using the knowledge base: close the client sessions, the UI and running commands.
- Remove the stale
-waland-shmfiles next to the target, if present. - Copy the backup to
$MEMO_HOME/kb/<name>.db, thenchmod 600it. - Check it:
MEMO_KB=<name> memo-mcp verify MEMO_KB=<name> memo-mcp status
A backup taken by an older binary is migrated forward on first write. A backup from a newer binary is refused by an older one.
A human-readable copy
memo-mcp export --md <dir> writes every live document as markdown with full provenance front
matter. It is not a complete backup: there are no facts, history, graph or logs. But it is
readable without memo-mcp, and re-ingesting it creates no new revisions. See
Exporting.
Integrity and repair
verify
memo-mcp verify
memo-mcp verify --repair
| Check | Problem reported | --repair does |
|---|---|---|
| Every live document has passages | N live document(s) have no chunks | Reports only; re-ingest the source |
| Vectors point at existing passages | N vector(s) point at missing chunks | Reports only |
| Facts point at existing evidence | N fact(s) point at missing evidence chunks | Reports only |
| Every live passage has a vector for the current model | N live chunk(s) have no vector for model X | Queues an embed job; run memo-mcp backfill |
| FTS5 index integrity | <index> failed integrity-check | Rebuilds the index |
The command prints ok: no problems found, or one problem: line per finding and one
repaired: line per fix.
backfill and reindex
| Command | Embeds | When |
|---|---|---|
memo-mcp backfill | Passages queued for vectors, from --embed=false, a degraded write or verify --repair | After bulk loads, or after the model was unavailable |
memo-mcp reindex [--model id] | Every live passage lacking a vector for the model | After switching models |
Both are resumable and safe to interrupt. The server runs both in the background at start.
SQLite-level checks
sqlite3 ~/.memo-mcp/kb/my-project.db 'PRAGMA integrity_check' # expect: ok
sqlite3 ~/.memo-mcp/kb/my-project.db 'PRAGMA user_version' # schema version
When to run what
| After… | Run |
|---|---|
| A crash or power loss | verify, then PRAGMA integrity_check |
| Restoring a backup | verify |
Bulk ingest with --embed=false | backfill |
Changing MEMO_MODEL | reindex --model <id> |
| Upgrading across 1.0 to 1.1 | graph rebuild |
| Large redactions | VACUUM with no process attached |
Capacity and performance
Numbers measured on an Apple-silicon laptop with the shipped defaults. They are a guide to proportions, not a benchmark of your hardware.
Disk
| Item | Size |
|---|---|
| Fixed overhead of an empty knowledge base | ~0.3 MB |
| Per passage, before vectors (text, two FTS5 indexes, graph rows) | ~4.4 KB |
Per passage per 384-dimension model (granite-small-r2, minilm) | ~1.5 KB |
| Per passage per 768-dimension model | ~3 KB |
| Typical passages per KB of markdown | ~1 per 0.9 KB |
| Default model in the cache | ~195 MB |
Example: this project’s docs/ and articles/ folders (756 KB of markdown) became 101
documents and 818 passages, in a 3.6 MB file before vectors. Vectors for one 384-dimension model
add about 1.2 MB more.
Time
| Operation | Cost |
|---|---|
| Ingest without embedding | ~60 documents per second |
Embedding a passage, granite-small-r2 | ~0.6 s |
Embedding a passage, minilm | ~0.25 s |
Embedding, potion | near instant |
Search p50, granite-small-r2 (includes embedding the query) | ~0.2 s |
Search p50, potion | ~6 ms |
| Search p50, keyword-only | ~1–3 ms |
Reranker (precise + MEMO_RERANK=1) | +0.7–4 s per query |
Query embedding dominates search latency. Retrieval itself is milliseconds at these sizes.
memo_search_arm_duration_seconds{arm} shows the split on your data.
Memory
A server process holds the embedding model in memory: a few hundred MB for granite-small-r2,
less for potion. It also holds one in-memory mention graph per namespace it has searched with
graph routing. go_memstats_sys_bytes shows the total.
Scaling guidance
- Everything is single-file SQLite with exact (brute-force) vector search. Tens of thousands of
passages per knowledge base are comfortable. Beyond about 100,000 passages, watch
memo_search_arm_duration_seconds{arm="semantic"}, and split knowledge bases by project. - Bulk loads: ingest with
--embed=false, thenbackfill. Or choosepotionfor fast, slightly weaker semantics. - Several sessions on one file are fine. Writes are serialised and short; embedding happens outside the write lock.
Exporting
Markdown with provenance
memo-mcp export --md ./kb-export # all namespaces
memo-mcp export --md ./kb-export --ns handbook
The layout is <namespace>/<kind>/<slug>-<shortid>.md, plus an _index.md per namespace. Each
file has YAML front matter:
---
memo_uri: memo://doc/01a1017e-ec2c-777a-b64a-4b67d1ac8f97
title: Deploy checklist
kind: doc
namespace: default
revision: 1
content_hash: sha256:cf966c40e2f66a5d595ab316004e2875ff5f70aa51679c31669852fc121d7326
fetched_at: 2026-10-03T11:20:57Z
trust: user
origin: user-said
---
Empty fields are left out. source_uri, library, version and context appear when they
are set.
The export opens as an Obsidian vault. Re-ingesting it produces zero new revisions, because the content hashes match. Facts, history, graph and logs are not included.
An index for agents
memo-mcp export --index > kb-index.md # ≤ 8 KB, one line per document
memo-mcp export --index --library grpc/grpc-go@v1.64.0 --max-bytes 4096
Each line has the title, address, kind, version and trust, and facts are summarised. Lines that do
not fit the budget are counted in a footer. Paste the output into CLAUDE.md or AGENTS.md so
an agent knows what exists before it searches. The same content is the MCP resource
memo://index.
Raw data
The file is standard SQLite, so any tool can read it. For example, every live fact as CSV:
sqlite3 -csv -header ~/.memo-mcp/kb/my-project.db \
"SELECT id, namespace, statement, trust, valid_from, valid_to FROM facts WHERE invalidated_at IS NULL AND deleted_at IS NULL"
Column names: schema.md.
Troubleshooting runbook
Start with these three commands. They answer most questions:
memo-mcp version # which build, protocol, model dir
MEMO_KB=<name> memo-mcp status # file, schema, counts, model, jobs
MEMO_KB=<name> memo-mcp verify # integrity
The client does not show memo’s tools
| Check | Fix |
|---|---|
/mcp in Claude Code, or the developer settings in Claude Desktop, shows an error | Read the server’s stderr; the reason is the last fatal: line |
command is relative or uses ~ | Use an absolute path |
fatal: invalid name … | MEMO_KB must match ^[A-Za-z0-9._-]{1,64}$ |
fatal: … address already in use | Another session holds the same --metrics-addr. See the metrics endpoint |
fatal: database schema is at version N, but this binary only supports up to version M | The file was written by a newer memo-mcp; upgrade the binary |
| macOS refuses to run the binary | xattr -d com.apple.quarantine /path/to/memo-mcp |
Search returns nothing, or too little
| Check | Meaning, and fix |
|---|---|
Response has a reason and zero results | Deliberate abstention: nothing passed the semantic floor. Check with memo-mcp search "<q>" --mode keyword |
status shows 0 documents, or a different count than you expect | Wrong MEMO_KB or MEMO_HOME: the CLI and the client are not looking at the same file |
trace filtered.by_scope is high | The request’s scope (namespace, version, library, minimum trust) excludes the answer |
degraded is set | No model: see the next section |
truncated is above 0 | The token budget cut results; raise max_tokens or narrow with narrow_hint |
| Lists stop after 3 results | Gap cutoff (cutoff.kind = gap); see cutoff_gap in Profiles |
Search is degraded
| Check | Fix |
|---|---|
stderr shows embedding unavailable… | Model download or load failed; the error follows. Retry with memo-mcp model pull <id> |
| Download interrupted or corrupted | memo-mcp model redownload <id> (pass the id) |
memo_kb_pending_embeddings above 0 | Backlog after a model switch or a degraded write: memo-mcp backfill, or reindex --model <id> |
| No network by design | Pre-seed the cache: Air-gapped installs |
Slow searches
memo-mcp explain "<q>", or the UI: seelatency_ms_per_armin the trace.semanticis slow: query embedding dominates. Considerpotion, or accept about 0.2 s.graphis slow on its first call: the namespace graph is built once. Checkmemo_graph_cache_total{event="build"}. Repeated builds mean frequent writes.- Reranker on?
MEMO_RERANK=1adds 0.7–4 s. Turn it off.
Slow writes
Embedding dominates: about 0.6 s per passage with the default model. For bulk loads, use
ingest --embed=false, then backfill.
“database is locked”
Another process held the write lock for more than 5 seconds. This is rare, because writes are
short and embedding runs outside the lock. Look for a long graph rebuild, VACUUM or a
manual sqlite3 session holding a transaction. Never put the file on a network file system.
Read-only commands fail with “run memo-mcp migrate”
The file is older than the binary. Run memo-mcp migrate, or any writing command, once.
The UI is unreachable
- Use the exact URL it printed. The Host header must match the bound address.
- It binds loopback only. From another machine, use an SSH tunnel
(
ssh -L 9470:127.0.0.1:9470 host) rather than--allow-remote. The UI has no authentication and shows everything in the knowledge base.
Promote never shows a dialog
The client does not support elicitation, so the tool result contains a memo-mcp trust promote …
command for a human to run. memo_mcp_elicitations_total{outcome="unsupported"} counts these.
Reporting a bug
Include memo-mcp version, the status output, the stderr lines from
MEMO_LOG_LEVEL=debug, and, for ranking issues, memo-mcp explain "<query>". Do not include
document text you would not share.
Release engineering
For maintainers and for operators who build their own releases.
Pipelines
| Workflow | Trigger | Does |
|---|---|---|
ci.yml | Push, pull request | gofmt, go vet, tests under the race detector (these include the retrieval eval gate and the tools/list golden), a pure-Go build and smoke run on Linux and macOS (Windows is currently disabled in the test matrix), golangci-lint, and a GoReleaser snapshot that must produce five archives |
nightly.yml | Daily at 03:17 UTC | The eval with the hash embedder and granite-small-r2; uploads the report as an artifact |
release.yml | Tag v* | Tests, then GoReleaser (five archives, checksums.txt), then publishes server.json to the MCP registry with GitHub OIDC (no stored secret) |
docs.yml | Push to master touching docs/guide/**, pull requests (build only), manual | Builds these guides with mdBook, runs scripts/docs-check.sh, deploys to GitHub Pages |
Gates a change must pass
- Eval gate: no single query may drop a rank band, and no category mean may fall by more
than 0.02 against
internal/eval/testdata/baseline.json. - Tool-surface golden:
internal/server/testdata/tools.golden.jsonmust match. Tool names, parameters and descriptions are the 1.x contract. - Docs check: every
MEMO_*variable, command and metric in the code must appear in this guide.
Cutting a release
make check # lint + race tests
make eval # retrieval against the baseline
make snapshot # local GoReleaser dry run into dist/
# write docs/eval/vX.Y.Z.md and the CHANGELOG entry
git tag -a vX.Y.Z -m "vX.Y.Z" && git push origin vX.Y.Z
The release notes footer points at docs/eval/vX.Y.Z.md. Every tag should have one.
Versioning
Semantic versioning over the 1.x contract: tool names and parameters, memo:// addresses,
explain field names and the export front matter. Schema migrations are additive and are
listed in Upgrading and migrations.
Appendix
- Glossary: the technical terms used in this guide.
- Further reading: design documents, research and history.
Glossary
| Term | Meaning |
|---|---|
| Abstention | Returning zero results on purpose, with a reason, when nothing passes the semantic floor |
| Address | A memo://<kind>/<id> URI identifying one record; it resolves even after the record is forgotten |
| Arm | One retrieval method producing a ranked list: keyword, exact, semantic, fact, entity, graph |
as_of | A time-travel view: what the knowledge base had recorded, and believed valid, at a given time |
| Audit | The table recording every write: actor, channel, operation, target |
| Band | A relevance class (strong, moderate, weak) from the best raw cosine, using thresholds per model |
| BM25 | The term-frequency ranking function FTS5 uses for the keyword and exact arms |
| Channel | How a write arrived: tool, cli, elicitation, worker. It determines trust |
| Chunk / passage | A section of a document of about 200 tokens: the unit of indexing and retrieval |
| Compaction | Agent-performed maintenance: writing pages, settling conflicts, deciding merges |
| Cutoff | Where a result list ends: a score gap, the limit, the token budget, or none |
| Degraded | A search served without a capability, usually the embedding model |
| Elicitation | The MCP mechanism by which a server asks the human a question through the client |
| Embedding | A vector representing a text’s meaning; compared with cosine similarity |
| Entity | A named thing (system, person, identifier) extracted from passages into the graph layer |
| Fact | One-sentence claim with evidence, a subject list and a validity window |
| FTS5 | SQLite’s full-text search extension; memo-mcp keeps two indexes (stemmed, and identifier-preserving) |
| Fusion | Combining arm lists into one ranking: RRF (rank-based) or min-max (score-based) |
| Granularity | What a search returns: chunk, document, fact or page |
| Knowledge base (KB) | One SQLite file, selected by MEMO_KB; the isolation boundary |
| MCP | Model Context Protocol: the JSON-RPC protocol between AI clients and tool servers |
| Merge candidate | Two entity names similar enough to possibly be the same thing; queued for a human |
| Namespace | A label grouping sources inside one knowledge base; searches span all by default |
| Origin | Declared provenance class: web, user-said, agent-derived |
| Page | Agent-written markdown about an entity, stored as an inference with its cited passages |
| PPR | Personalised PageRank: the random walk the graph arm runs from each named entity |
| Profile | A named set of ranking constants (MEMO_PROFILE) |
| Recency | A multiplicative score factor that decays with age for kinds that age |
| Revision | One version of a source’s document; changed content creates a new revision |
| RRF | Reciprocal rank fusion: sum of weight / (k + rank) over arms |
| Source | Where content came from (URI or file), with library, version and trust |
| Stale | A page whose sources changed or were forgotten since it was built |
| Supersede | Replace a revision or fact with a newer one, keeping the old one as history |
| Trace | The per-query explanation: arms run, candidates, latency, filters, cutoff, budget |
| Trust | Who vouched for a record: agent < user < curated; only humans raise it |
| WAL | SQLite’s write-ahead log mode, which lets readers and a writer work concurrently |
| Why | The per-result explanation: each arm’s rank and contribution, fused and final scores |
| Work item | One compaction task with everything needed to do it in its payload |
Further reading
All links point to the repository on GitHub.
Design
- Architecture: the 1.x contract, formulas, explain fields and observability.
- Schema: every table and column, trust transitions, time rules and migrations.
- Roadmap: phases, decisions, verification and the glossary.
- Research: the state of the art behind the design.
Measurements
- Eval reports: one per release.
- Spikes: the vector store, FTS5, embedding models, the graph, elicitation and the reranker.
Articles
One article per development phase, written for newcomers: articles/. Two of them matter most for operators:
For agents
- SKILL.md: how an agent should use the tools.