Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

memo-mcp operator guide

This guide is for the people who install, configure, tune and run memo-mcp. It is written for readers comfortable with a terminal, SQLite, environment variables and Prometheus.

memo-mcp is a single, pure-Go binary. It exposes a knowledge base to AI agents over the Model Context Protocol (MCP) on stdio. It stores everything in one SQLite file per knowledge base, and it ships with a CLI and a read-only web UI over the same store. There is no daemon to deploy, no external database and no network dependency after a one-time model download.

What you will find here

PartRead it when you need to…
I. How it worksUnderstand the processes, the data model, the search pipeline and the security boundaries
II. Install and deployInstall, connect MCP clients, lay out several knowledge bases, upgrade, run offline
III. Configuration referenceLook up an environment variable, a command, a flag, the overrides file or a path on disk
IV. Retrieval tuningChange ranking, switch embedding models, try the reranker, and measure whether a change helped
V. Knowledge lifecycleAdminister ingest, facts, trust, the graph index, compaction and deletion
VI. Monitoring and observabilityScrape metrics, write alerts, read logs, use the call log
VII. OperationsBack up, verify, size, export, troubleshoot, release

Conventions

  • Commands are shown as memo-mcp <command>. Every command reads the same environment (MEMO_HOME, MEMO_KB, …), so set it the same way the MCP client does.
  • “Knowledge base” (KB) means one SQLite file. A “namespace” is a label inside one file.
  • Addresses such as memo://chunk/42 identify records; every CLI command and the UI accept them.
  • Defaults quoted here are the shipped values for v1.4. memo-mcp profiles show and memo-mcp --help print the live values for your build.

Deeper references

This guide states what to do and what it costs. The design documents explain why:

  • Architecture: the 1.x contract, formulas and defaults.
  • Schema: every table and column, trust transitions, time rules.
  • Eval reports: measured quality and cost per release.

Looking for the non-technical introduction? See the user guide.

Part I: How it works

Four short chapters give you the mental model the rest of the guide assumes.

System overview

One binary, three processes

memo-mcp is one statically linked Go binary (CGO_ENABLED=0, about 30 MB). The same binary runs in three roles, which are separate processes that share only the database file:

RoleStarted byTalks overLifetime
MCP server (memo-mcp or memo-mcp serve)The MCP client (Claude Code, Claude Desktop, …)stdio: JSON-RPC on stdin/stdout, logs on stderrAs long as the client session
Web UI (memo-mcp ui)An operatorHTTP on a loopback port, GET onlyUntil Ctrl-C
CLI (memo-mcp <command>)An operator or a scriptTerminalOne command

Each MCP client session usually starts its own server process. Several processes can hold the same file open at once; SQLite WAL mode and a 5-second busy timeout serialise writers.

 MCP client ── stdio ──▶ memo-mcp serve ─┐
 operator ── browser ──▶ memo-mcp ui ────┼──▶ $MEMO_HOME/kb/<name>.db  (SQLite, WAL)
 operator ── shell ────▶ memo-mcp <cmd> ─┘
                                          └─▶ ~/.cache/memo-mcp/models  (embedding models)

Inside the process

PackageResponsibility
internal/serverThe ten MCP tools and the memo:// resources; request middleware (metrics, logs, call log)
internal/cliCommands, configuration resolution, serve and UI wiring
internal/retrieveThe search pipeline: arms, fusion, recency, cutoff, budget, explain
internal/kbSchema and migrations, ingest, facts, trust, time, graph tables, pages, export
internal/embeddingThe model registry, download, in-process inference (hugot / GoMLX, pure-Go ONNX)
internal/graph, internal/compactEntity extraction and resolution; compaction work items and checks
internal/uiThe read-only web UI (html/template, embedded assets)
internal/obsMetrics registry, Prometheus exposition, logging setup

The ten MCP tools

ToolWrites?Purpose
ingestyesStore a markdown document with provenance; same content is a no-op, changed content a new revision
searchno*Hybrid search with optional per-result explanations
readnoDereference a memo:// address as a passage, section or document
rememberyesRecord one fact with evidence and validity dates; supersedes corrects an old one
forgetyesRetire a document or fact with a reason (agent-written records only)
promotevia a humanAsk a human to raise trust, through an elicitation dialog or a CLI command
explorenoWalk the entity graph from one name
compactwork itemsPropose pages, refreshes, conflict and merge decisions, duplicates
submityesHand back a page or a decision for a work item; checked before it is stored
statusnoNamespaces, model, pending vectors, jobs, graph and page counts

* search writes only to the opt-in query log.

Resources mirror the addresses (memo://doc/{id}, memo://chunk/{id}, …), and memo://index returns a one-line-per-document index under 8 KB.

The one network call

The first time a process needs an embedding model that is not cached, it downloads the model files from Hugging Face into ~/.cache/memo-mcp/models. Nothing else in the default configuration opens an outbound connection. The optional Ollama executor and the --metrics-addr listener are explicit opt-ins, and both are loopback-only by default. See Air-gapped installs to avoid even the download.

Design reference: architecture §1–§4.

Data model

One knowledge base is one SQLite file. Inside it, knowledge is stored in layers. Each layer points down to the one it was derived from, so every answer can be traced to its source.

L5  pages       markdown an agent wrote from passages; cites them; marked stale when a source changes
L4  graph       entities, aliases, mentions (entity ↔ passage), merge candidates, typed edges
L3  facts       one-sentence claims with a validity window and an evidence passage
L2  chunks      passages of ~200 estimated tokens, indexed three ways (two FTS5 + vectors)
L1  documents   the text as ingested, one row per revision
L0  sources     where it came from: URI or file, library, version, hash, trust, origin
    bookkeeping namespaces, models, jobs, audit, query_log, call_log
LayerRebuildable?Notes
L0–L2Yes, deterministically from the sourcesChunking and indexing are pure functions of the text
L3No: facts are additiveCorrections supersede or invalidate; nothing is deleted in place
L4Yes: memo-mcp graph rebuildMerge decisions made by a human are kept
L5No: written by agentsStored as is_inference, with the passages each page cites

Addresses

AddressResolves to
memo://source/<id>The latest live document of a source
memo://doc/<id>One document revision
memo://chunk/<n>One passage (read with section adds its neighbours)
memo://fact/<id>One fact, with its evidence address and its two clocks
memo://entity/<id>One entity, its aliases, passages and neighbours
memo://page/<id>One curated page with its sources and stale flag
memo://index, memo://ns/<namespace>/indexOne line per document and fact, under 8 KB

Document and fact ids are UUIDv7, which sort by time. Chunk ids are integers. A forgotten record’s address still resolves, to “forgotten on … because …”.

Namespaces and knowledge bases

  • A namespace is a label on sources inside one file. A search spans all namespaces unless the request scopes it. Namespaces organise; they do not isolate.
  • A knowledge base is a file selected by MEMO_KB. Files are the isolation and privacy boundary: no command or tool reads across files.

Trust and origin

Every source, and therefore every passage and fact, carries two labels:

LabelValuesSet by
trustagent < user < curatedThe channel the write came through, never the request. MCP tool writes are agent. CLI writes are user by default, curated only with --trust curated
originweb, user-said, agent-derivedDeclared by the writer: fetched from the web, said by a person, or inferred by an agent

Only a human can raise trust: memo-mcp trust promote, or an elicitation dialog the client shows for the promote tool. Tool calls may forget or supersede only agent records. Every transition is written to the audit table. See Trust administration.

Two clocks

  • Recorded time: when the knowledge base learned something.
  • Valid time: when a fact is true in the world (valid_from, valid_to).

as_of queries (“what did we believe on 1 March?”) filter on recorded time, and also on valid time for facts that have a window. Superseded revisions and replaced facts stay readable under as_of. Forgotten records are excluded in both views. See Facts and time.

Full column-by-column reference: schema.md.

The search pipeline

Every search, from the MCP tool, the CLI or the UI, runs the same pipeline in internal/retrieve:

scope filter ─▶ arms (parallel ranked lists) ─▶ fusion ─▶ recency ─▶ abstention
            ─▶ cutoff ─▶ limit ─▶ token budget ─▶ results + why + trace

1. Scope filter

Namespaces, kinds, sources, library, version, tags, dates and minimum trust are applied inside each arm’s SQL, before top-k. A filtered search therefore never loses a result to a candidate that was filtered out later. The live filter excludes superseded and forgotten records. Under as_of, it is replaced by “recorded on or before T and not superseded before T”.

2. Arms

ArmIndexRuns when
keywordFTS5, Porter stemmer, BM25, over text and its section headerAlways, except mode: semantic
exactFTS5 that keeps identifiers whole (net/http, ERR_CONN_RESET)mode: exact, or auto when the query looks like an identifier
semanticCosine similarity over unit vectors in a plain SQLite tableA model is loaded and passages have vectors for it
factFacts matched by keyword or meaning; a hit votes for its evidence passageAlways
entityPassages mentioning entities the query namesRouted: the query names 2+ known entities, or 1 with relational phrasing
graphPersonalised PageRank over the mention graph, one walk per named entitySame routing; reports only passages not every seed mentions directly

Each arm returns up to fetch_depth candidates (100 by default).

3. Fusion

Weighted reciprocal rank fusion: a passage’s score is the sum over arms of weight / (k + rank), with k = 60. Default weights are semantic 0.5, keyword 0.5, exact 0.3, fact 0.4, entity 0.4 and graph 0.5. The minmax profile uses score fusion instead.

4. Recency

For kinds that age (note, conversation by default), the score is multiplied by floor + (1 − floor) · 0.5^(age_days / half_life), with floor 0.8 and half-life 90 days. Versioned documents and code do not age.

5. Abstention and relevance bands

A candidate whose only evidence is a semantic similarity below semantic_floor is dropped. If nothing survives, the response has zero results, a reason and a hint. That is a deliberate “not in the knowledge base”, not an error.

relevance is the best raw cosine for a result, interpreted against the model’s bands. For the default model, granite-small-r2, strong is 0.88, moderate 0.80 and weak 0.72, because unrelated text already scores about 0.63 under it. score orders one list and is not comparable across queries; relevance is.

6. Cutoff, limit and budget

The list ends at the first result, past min_results (3), whose score drops by more than cutoff_gap (half) from the previous one. It is then capped at limit (default 10, maximum 100), and then packed into max_tokens. The response always reports truncated, the count of ranked results that did not fit, and narrow_hint, the scope field that would shorten the list.

7. Explain and trace

With response_format: explain (MCP) or --explain (CLI), every result carries why: each arm’s rank, raw score and contribution, the fused and final score, relevance band, provenance and freshness. Every response carries a trace: arms run and why, candidates and latency per arm, filter counts, cutoff kind, budget use and degraded state. The UI renders the same structures, and a test asserts that the numbers are identical.

Degraded mode

If the embedding model is not available, because it is downloading, failed to load or has no vectors yet, search runs without the semantic arm. It sets degraded{flag, reason}, and the memo_search_degraded_total metric counts it. Writes made in that state queue their vectors. memo-mcp backfill or the next server start embeds them.

Tuning these constants: Profiles and constants. Full formulas: architecture §5.

Security model

memo-mcp is designed to run on one person’s machine with that person’s privileges. Its security properties come from being local, small and explicit.

Exposure

SurfaceExposureGuard
MCP serverstdio of the client process onlyNo network listener. stdout is reserved for JSON-RPC; all diagnostics go to stderr
Web UILoopback (127.0.0.1:0 by default)GET only, Host header allowlist against DNS rebinding, CSP, no mutating routes. A non-loopback bind needs --allow-remote
Metrics endpointOff by default; loopback only, with no overrideGET /metrics only, Host check, 404 for any other path
Ollama executorOff by default; loopback unless --allow-remoteCLI only; no MCP code path reaches it

Outbound traffic

  • Model download from Hugging Face on first use of a model, into ~/.cache/memo-mcp/models. Avoidable; see Air-gapped installs.
  • Nothing else. The server never fetches URLs. Agents fetch content and pass the text to ingest. There is no telemetry: metrics are pull-only on loopback, logs go to stderr, and the call log is a table in your own file.

Data at rest

  • One SQLite file per knowledge base, in plaintext. Directories are created 0700 and files 0600. A pre-existing directory with wider permissions is not tightened; check it.
  • Use full-disk encryption if the file may hold sensitive material.
  • forget --redact erases a record’s text. Plain forget hides the record but keeps the text in history.

Trust gates

  • Trust is set by channel. An agent cannot write a record above agent, and cannot forget or supersede user or curated records.
  • Raising trust requires a human: memo-mcp trust promote, or an MCP elicitation dialog that shows the excerpt, the source and the target level. A client configured to auto-accept elicitation removes this protection. In that case, treat user and curated as no stronger than agent.
  • Every write and every trust change is recorded in the audit table with its actor and channel.

Prompt injection

Ingested content is data, but agents read it. memo-mcp limits the blast radius: content cannot raise its own trust, cannot delete human records and cannot trigger network calls. Results carry provenance and trust, so an agent, or a human reviewing its answer, can discount agent-trust web content.

The MCP host

Everything an agent writes or searches passes through the MCP host as tool input and output. The knowledge base is as private as that host. memo-mcp does not advertise the MCP logging capability; logs stay on stderr, which the host may display or store.

Part II: Install and deploy

memo-mcp has no installer and no service to register. Installing it means putting one binary on disk and telling an MCP client how to start it.

Release archives

Every tag vX.Y.Z publishes five archives and a checksum file on the releases page:

PlatformArchive
macOS, Apple siliconmemo-mcp_X.Y.Z_darwin_arm64.tar.gz
macOS, Intelmemo-mcp_X.Y.Z_darwin_amd64.tar.gz
Linux, x86-64memo-mcp_X.Y.Z_linux_amd64.tar.gz
Linux, ARM64memo-mcp_X.Y.Z_linux_arm64.tar.gz
Windows, x86-64memo-mcp_X.Y.Z_windows_amd64.zip
Allchecksums.txt (SHA-256)

Each archive contains the binary, LICENSE, README.md and SKILL.md.

Install

VERSION=1.4.0
ARCH=darwin_arm64            # see the table above
curl -LO https://github.com/kKEo/memory-find/releases/download/v${VERSION}/memo-mcp_${VERSION}_${ARCH}.tar.gz
curl -LO https://github.com/kKEo/memory-find/releases/download/v${VERSION}/checksums.txt
shasum -a 256 -c --ignore-missing checksums.txt      # Linux: sha256sum -c --ignore-missing
tar xzf memo-mcp_${VERSION}_${ARCH}.tar.gz
install -m 0755 memo-mcp ~/.local/bin/memo-mcp       # any directory; note the absolute path
memo-mcp version

memo-mcp version prints the build version, the MCP protocol version (2026-07-28), the Go version and the model directory.

macOS Gatekeeper. A binary downloaded by a browser carries a quarantine attribute and is blocked on first run. Clear it with xattr -d com.apple.quarantine ~/.local/bin/memo-mcp, or download with curl, which does not set it.

MCP registry

The server is listed in the MCP registry as io.github.kKEo/memory-find. Its server.json points at the release archives with their checksums, so registry-aware clients can install it directly.

What the binary needs at runtime

  • No shared libraries, no CGo, no interpreter.
  • A writable MEMO_HOME, by default ~/.memo-mcp, and a writable ~/.cache/memo-mcp/models.
  • About 200 MB of disk for the default model, plus the knowledge-base files. See Capacity and performance.
  • Network access to huggingface.co once per model, unless you pre-seed the cache.

Build from source

Requirements: Go 1.26 or newer. No C toolchain is needed.

git clone https://github.com/kKEo/memory-find
cd memory-find
make build            # CGO_ENABLED=0, stripped, version from `git describe`
./memo-mcp version

make build runs:

CGO_ENABLED=0 go build -ldflags "-s -w -X main.version=$(VERSION)" -o memo-mcp ./cmd/memo-mcp/

Cross-compiling is plain Go, for example GOOS=linux GOARCH=arm64 CGO_ENABLED=0 go build ./cmd/memo-mcp.

Reproducing a release

make snapshot runs the same GoReleaser configuration as the release workflow without publishing. It writes the five archives and checksums.txt under dist/. Release builds use -trimpath and the commit timestamp, so rebuilding a tag gives the same binaries.

Development targets

TargetWhat it does
make checkgofmt, vet, golangci-lint, then all tests under the race detector
make testAll tests, no race detector
make evalThe retrieval benchmark against the recorded baseline
make golden-updateRe-record the tools/list golden after a deliberate tool-surface change
make baseline-updateRe-record the eval baseline after a deliberate retrieval change
make docsBuild these guides with mdBook into site/

Tests use a deterministic hash embedder and need no network or model download.

Connecting MCP clients

memo-mcp speaks MCP over stdio. A client starts the binary, writes JSON-RPC to its stdin and reads responses from its stdout. Configuration is passed through environment variables in the client’s server entry, not the shell. The client starts the process, so your shell profile does not apply.

Claude Code

# In a project directory: this project only, private to you (the default "local" scope)
claude mcp add memo --env MEMO_KB=my-project -- /absolute/path/to/memo-mcp

# Shared with the team through .mcp.json in the repository
claude mcp add memo --scope project --env MEMO_KB=my-project -- memo-mcp

# Available in every project
claude mcp add memo --scope user --env MEMO_KB=personal -- /absolute/path/to/memo-mcp

Add more variables with additional --env flags, for example --env MEMO_QUERY_LOG=1 --env MEMO_LOG_LEVEL=warn. Check the connection with /mcp inside Claude Code. Claude Code shows a stdio server’s stderr, so memo-mcp’s logs appear there.

Claude Desktop

Edit claude_desktop_config.json. On macOS it is in ~/Library/Application Support/Claude/; on Windows, in %APPDATA%\Claude\. Use absolute paths; ~ is not expanded.

{
  "mcpServers": {
    "memo": {
      "command": "/Users/you/.local/bin/memo-mcp",
      "env": {
        "MEMO_KB": "my-project",
        "MEMO_QUERY_LOG": "1"
      }
    }
  }
}

Restart Claude Desktop completely after editing.

Any other stdio client

The generic shape is the same everywhere: a command, optional args and an environment map.

FieldValue
commandAbsolute path to memo-mcp
argsnone, or ["serve", "--metrics-addr", "127.0.0.1:9469"]
envMEMO_KB, and optionally MEMO_HOME, MEMO_MODEL, MEMO_PROFILE, MEMO_QUERY_LOG, MEMO_LOG_FORMAT, MEMO_LOG_LEVEL

The server implements protocol version 2026-07-28: stateless, with server/discover in place of initialize. It advertises tools and resources. It does not advertise logging or prompts. The promote tool uses elicitation when the client supports it, and otherwise returns the CLI command a human can run.

Teaching the agent

Ship SKILL.md from the release archive with the project, for example as a Claude Code skill, or paste it into the system prompt. It describes the search-then-read loop, scoping, writing facts and how to treat trust. For a static orientation, add the output of memo-mcp export --index to CLAUDE.md or AGENTS.md.

Verifying a connection without a client

MEMO_KB=my-project memo-mcp status      # read-only; fails if the file does not exist yet

A server that starts and immediately exits usually has an invalid MEMO_KB name or an unwritable MEMO_HOME. The reason is on stderr. See the troubleshooting runbook.

Knowledge bases and MEMO_HOME

Where files go

PathContentsCreated with
$MEMO_HOME (default ~/.memo-mcp)Base directory0700
$MEMO_HOME/kb/<name>.dbOne knowledge base (plus -wal and -shm while open)0600
$MEMO_HOME/profiles.jsonOptional ranking-profile overridesYou
~/.cache/memo-mcp/models/Downloaded embedding and reranker models, one directory each plus a .ok marker0755 (public model files)

<name> comes from MEMO_KB, default default. It must match ^[A-Za-z0-9._-]{1,64}$ and must not be . or ... The model cache location is fixed and shared by every knowledge base and process of the same user.

One KB per project

Files are the isolation boundary, so the usual pattern is one MEMO_KB per project or client:

cd ~/src/payments  && claude mcp add memo --env MEMO_KB=payments  -- ~/.local/bin/memo-mcp
cd ~/src/analytics && claude mcp add memo --env MEMO_KB=analytics -- ~/.local/bin/memo-mcp

Claude Code’s default local scope ties each entry to its project directory, so the same server name memo opens a different file in each project.

Namespaces within a KB

Use namespaces for topics that may be searched together: docs, decisions, grpc. Writes default to the namespace default. Searches span all namespaces unless scoped. Namespace counts are visible in memo-mcp status and as the memo_kb_namespace_documents metric.

Separate homes

MEMO_HOME moves everything except the model cache. Use it to keep test data apart, run CI against a throwaway directory, or put knowledge bases on an encrypted volume:

MEMO_HOME=/Volumes/secure/memo MEMO_KB=client-a memo-mcp status

Sharing between processes

Several processes may open the same file. WAL mode lets readers proceed during writes, and writers wait up to 5 seconds for each other. Each process serialises its own writes over one connection. Do not put knowledge bases on network file systems such as NFS or SMB: SQLite’s locking is not reliable there.

Legacy variables

JOURNAL_TOKEN and JOURNAL_PATH are accepted for one more release as aliases of MEMO_KB and MEMO_HOME, with a deprecation warning. Files from the pre-1.0 journal are not opened or migrated.

Upgrading and migrations

Replacing the binary

Stop or let finish any running memo-mcp processes, replace the binary, and start the client again. Each MCP session starts its own process, so a new binary is picked up on the next session.

Schema versions

The schema version is SQLite’s user_version. Migrations are additive and run forward only.

VersionReleaseAdds
10.4 to 1.0Sources, documents, chunks and their indexes, facts, bookkeeping
21.1Graph layer: entities, aliases, mentions, merge candidates, edges
31.2Pages layer: pages, page sources, page vectors, work items
41.4call_log table

memo-mcp status prints the version in its first line, and the metric memo_kb_schema_version exposes it.

What migrates and when

SituationBehaviour
A writing process opens an older file: serve, ingest, remember, forget, search, verify, backfill, migrate, …Migrates forward automatically, one transaction per migration
A read-only command opens an older file: status, ui, ls, read, export, metrics, facts, explore, lint, pages, trust ls, graph mergesRefuses with file is at schema vN, this binary expects vM; run memo-mcp migrate. Read-only opens never write
Any process opens a newer file than it knowsRefuses with database schema is at version N, but this binary only supports up to version M; upgrade memo-mcp

To migrate explicitly before anything else touches the file:

MEMO_KB=my-project memo-mcp migrate

Downgrades are not supported. Back up the file before upgrading across a schema version if you may need to roll back. See Backup and restore.

After upgrading to 1.1 or later from 1.0

Files written before the graph layer have no entity mentions. Populate them once:

memo-mcp graph rebuild            # all namespaces; or --ns <name>

Run it again after any release whose notes mention an extraction change.

After changing the embedding model

A model change is not a schema change. Vectors are stored per model, so the server embeds the missing ones in the background at start, and search reports degraded until it finishes. See Embedding models.

Release notes

Read the changelog and the eval report for the tag, docs/eval/vX.Y.Z.md, before upgrading. The 1.x contract keeps tool names and parameters, memo:// addresses, explain field names and the export format stable.

Air-gapped installs

After the binary is in place, the only network access memo-mcp ever makes is downloading an embedding model the first time it is needed. To run with no network at all, pre-seed the model cache.

Pre-seed the model cache

On a connected machine with the same memo-mcp version:

memo-mcp model pull granite-small-r2     # or the model you will set in MEMO_MODEL
ls ~/.cache/memo-mcp/models/
#   onnx-community_granite-embedding-small-english-r2-ONNX/
#   onnx-community_granite-embedding-small-english-r2-ONNX.ok

Copy both the directory and its .ok marker to the same path on the target machine, ~/.cache/memo-mcp/models/, for the user who runs memo-mcp. The marker records a completed, verified download. Without it, the files are treated as an interrupted download and fetched again.

Then check that the model loads offline:

memo-mcp model smoke granite-small-r2

For the optional reranker, also copy cross-encoder_ms-marco-MiniLM-L6-v2 and its marker.

Choosing a smaller model

Model idDownloadNotes
granite-small-r2 (default)about 195 MBBest measured quality; about 0.2 s per query on this backend
potionabout 130 MBStatic lookup table, no neural network; about 6 ms per query; weaker on paraphrase
minilmabout 87 MBThe previous default

See Embedding models for the full registry.

Running with no model at all

If no model is cached and the download fails, memo-mcp keeps working. Search runs on the keyword, exact and fact arms and reports degraded. New passages queue their vectors. When a model becomes available, memo-mcp backfill or the next server start embeds the backlog.

Part III: Configuration reference

memo-mcp is configured through environment variables, command flags and one optional file. There is no configuration file for the server itself.

Precedence. A flag beats its environment variable, which beats the built-in default. Examples: serve --metrics-addr over MEMO_METRICS_ADDR, search --profile over MEMO_PROFILE.

Environment variables

Set these in the MCP client’s server entry for the server, and in your shell for CLI commands. Both must agree on MEMO_HOME and MEMO_KB to see the same file.

Storage and selection

VariableDefaultEffect
MEMO_HOME~/.memo-mcpBase directory. Knowledge bases live in $MEMO_HOME/kb/; overrides in $MEMO_HOME/profiles.json
MEMO_KBdefaultKnowledge-base name; the file is $MEMO_HOME/kb/<name>.db. Must match ^[A-Za-z0-9._-]{1,64}$

Retrieval

VariableDefaultEffect
MEMO_MODELgranite-small-r2Embedding model id from memo-mcp model ls. Queries use this model’s vectors; missing vectors are embedded in the background at server start
MEMO_PROFILEdefaultRanking profile for the server and the UI (memo-mcp profiles show lists them). The CLI search --profile flag defaults to it
MEMO_RERANKunset1 loads the cross-encoder reranker. Only profiles with rerank on (precise) use it. Measured slower and worse than the default; experiments only
MEMO_RERANKERms-marco-minilmReranker id to load when MEMO_RERANK=1

Logging and observability

VariableDefaultEffect
MEMO_LOG_FORMATtexttext or json. Logs always go to stderr
MEMO_LOG_LEVELinfodebug, info, warn or error. Unknown values fall back to the default with one warning
MEMO_METRICS_ADDRempty (off)Same as serve --metrics-addr: serve Prometheus metrics at http://<addr>/metrics. Loopback addresses only
MEMO_QUERY_LOGunset1 records every search in query_log and every tool call in call_log, in the same file. See Call and query logs

Compaction executor (CLI only)

VariableDefaultEffect
MEMO_OLLAMA_URLhttp://127.0.0.1:11434Ollama endpoint for memo-mcp compact --executor ollama. Loopback only unless --allow-remote
MEMO_OLLAMA_MODELnone (required)Ollama model name for the executor. Use a non-thinking instruction model

Deprecated

VariableReplacementBehaviour
JOURNAL_TOKENMEMO_KBUsed as the name when MEMO_KB is unset, with a warning. Old journal files are not opened
JOURNAL_PATHMEMO_HOMEUsed as the base directory when MEMO_HOME is unset, with a warning

Not configurable

  • The model cache, ~/.cache/memo-mcp/models, is fixed per user.
  • SQLite pragmas are fixed: WAL, synchronous=NORMAL, busy_timeout=5000, foreign keys on.

Commands and flags

Running memo-mcp with no arguments is the same as memo-mcp serve. memo-mcp --help prints the summary. Flags may come before or after positional arguments.

Opens shows how each command opens the knowledge base:

  • rw: creates and migrates the file.
  • rw, no create: migrates, but refuses a missing file.
  • ro: never writes, and refuses a missing or out-of-date file.

Server and UI

CommandOpensPurpose and flags
memo-mcp serverwMCP server on stdio. --metrics-addr 127.0.0.1:PORT exposes /metrics on loopback (default $MEMO_METRICS_ADDR, off when empty). Starts a background backfill and reindex for the current model
memo-mcp uiroRead-only web UI. --addr (default 127.0.0.1:0, a random port; the URL is printed), --allow-remote permits a non-loopback address, --no-model gives keyword-only search

Writing

CommandOpensPurpose and flags
memo-mcp ingest <file|dir|->rwAdd markdown documents. --ns (default default), --kind doc|note|code|conversation, --uri, --title, --library, --version, --trust user|curated (default user), --origin web|user-said|agent-derived, --context "<sentence>", --embed=false queues vectors instead of embedding now
memo-mcp remember "<fact>"rwRecord a fact. --ns, --about a,b, --valid-from, --valid-to (YYYY-MM-DD), --supersedes memo://fact/.., --evidence memo://chunk/n, --trust user|curated, --origin, --no-model
memo-mcp forget <memo://doc/..|memo://fact/..>rwRetire a record. --reason "<why>" (required), --redact also erases the text
memo-mcp trust lsroList records above agent trust
memo-mcp trust promote <uri> --to user|curatedrwRaise trust (human channel; audited)
memo-mcp trust demote <uri> --to agent|userrwLower trust (audited)
CommandOpensPurpose and flags
memo-mcp search "<q>"rw, no createHybrid search. --mode auto|hybrid|keyword|exact|semantic, --ns a,b, --library, --version, --kind a,b, --min-trust agent|user|curated, --limit (10), --granularity chunk|document, --format table|json|md, --explain, --max-tokens (8000), --profile (default $MEMO_PROFILE), --rerank (default $MEMO_RERANK), --no-model
memo-mcp explain "<q>" [<memo://...>]rw, no createsearch --explain; with an address, the full explanation for that one result. Same flags as search
memo-mcp read <memo://...>roPrint any record with its provenance. --history prints the revision or supersession chain
memo-mcp lsroLive documents, newest first. --ns, --kind, --since YYYY-MM-DD, --limit (50), --json
memo-mcp facts lsroFacts. --ns, --as-of YYYY-MM-DD, --history includes replaced facts, --json
memo-mcp explore <name>roWalk the graph from one entity. --ns, --hops 1|2, --as-of, --json
memo-mcp export --md <dir>roMarkdown files with front matter, plus an _index.md per namespace. --ns
memo-mcp export --indexroA compact index for AGENTS.md and CLAUDE.md. --ns, --library name[@version], --max-bytes (8192)

Graph and compaction

CommandOpensPurpose and flags
memo-mcp graph mergesroMerge-candidate review queue. --state open|merged|rejected|all (default open)
memo-mcp graph merge <id> / memo-mcp graph reject <id>rwDecide a merge candidate
memo-mcp graph rebuildrwRe-extract entity mentions. --ns
memo-mcp compactrwCreate or list work items. --ns, --kinds page,stale,conflict,merge,duplicate, --lint, --json, --executor ollama, --apply (with an executor; dry run otherwise), --allow-remote
memo-mcp submit <item-id>rwComplete a work item. Pages: --content-file f.md|-, --title. Conflicts: --keep memo://fact/... Merges: --accept or --reject. Any item: --skip, --reason, --dry-run
memo-mcp lintroContradictions, orphan entities, missing and stale pages, expired facts. --ns, --json
memo-mcp pages lsroCurated pages. --ns, --stale, --json

Maintenance

CommandOpensPurpose and flags
memo-mcp statusroFile, schema version, counts, model, jobs, namespaces
memo-mcp migraterw, no createBring the file to this binary’s schema version
memo-mcp verifyrw, no createCheck documents without chunks, orphan vectors and facts, missing vectors for the current model, and FTS integrity. --repair queues missing vectors and rebuilds broken indexes
memo-mcp backfillrw, no createEmbed passages whose vectors are pending
memo-mcp reindexrw, no createEmbed every passage lacking a vector for a model. --model <id> (default $MEMO_MODEL)

Observability

CommandOpensPurpose and flags
memo-mcp metricsroSnapshot from the file: gauges, plus per-tool and search statistics from the opt-in logs. --json, --since (24h)
memo-mcp log tailrw, no createRecent searches, then recent tool calls. --n (20)
memo-mcp log callsrw, no createRecent tool calls with latency and outcome. --n (50)
memo-mcp log show <id>rw, no createOne logged search with its full trace
memo-mcp log replayrw, no createLogged searches as JSON lines, ready to label as eval queries. --n (200)
memo-mcp log prunerw, no createKeep at most 10,000 rows and 30 days in both logs

Models, profiles and evaluation

CommandOpensPurpose and flags
memo-mcp model lsroRegistry: id, dimension, licence, stored vectors, notes; marks the selected model
memo-mcp model smoke <id>|--allnoneProve a model loads under the pure-Go backend, and time it
memo-mcp model pull <id>noneDownload a model and check that it loads
memo-mcp model use <id>rw, no createRecord the knowledge base’s default model. Also set MEMO_MODEL in the client config
memo-mcp model redownload [<id>]noneDiscard the cached copy and download again. Without an id it re-downloads minilm, so pass the id you use
memo-mcp profiles show [<name>]noneEvery ranking constant with its derivation
memo-mcp evaltemporaryBenchmark on built-in labelled corpora. --models hash,minilm,…, --profiles default,…|all, --corpus notes|kb|all, --format table|md|json, --explain-failures, --rerank, --agent-proxy
memo-mcp versionnoneBuild version, MCP protocol version, Go version, model directory

Exit codes and output

  • 0 on success. 1 on a runtime error, printed as fatal: … on stderr. 2 on a usage error.
  • Results go to stdout, and diagnostics and logs to stderr. --json output is stable enough to script against.
  • The legacy flags --stats and --redownload-model still work for one more release, with a deprecation warning.

profiles.json

$MEMO_HOME/profiles.json changes ranking constants without rebuilding the binary. It is read at start-up by memo-mcp serve and memo-mcp ui. A missing file is fine. A malformed file stops the process with the path and the JSON error.

Note. In v1.4, the CLI search, explain and eval commands do not read profiles.json. They only see the built-in profiles. Test an override through the UI, which uses the same pipeline as the server, or by running the server.

Shape

The file is a JSON object mapping a profile name to the fields to change. Only fields that are present are applied.

  • An existing name, such as default, is modified in place.
  • A new name creates a profile. It starts from base, or from default when base is absent.
{
  "default": {
    "weights": { "exact": 0.4 }
  },
  "team-docs": {
    "base": "precise",
    "weights": { "semantic": 0.6, "keyword": 0.4 },
    "cutoff_gap": 0.4,
    "recency_kinds": [],
    "derivation": "Docs-heavy KB: favour meaning over wording; no recency."
  }
}

Select it with MEMO_PROFILE=team-docs in the server’s environment.

Fields

FieldTypeMeaningShipped default
basestringProfile to copy for a new namedefault
weightsmap of arm to numberArm weights; arms: semantic, keyword, exact, fact, entity, graph. Merged into the base’s weights; 0 disables an arm0.5 / 0.5 / 0.3 / 0.4 / 0.4 / 0.5
rrf_kintegerRRF damping constant60
fusion"rrf" or "minmax"Rank fusion, or min-max score fusionrrf
fetch_depthintegerCandidates per arm before fusion100
half_life_daysnumberRecency half-life90
recency_floornumber 0–1Fraction of score the oldest item keeps0.8
recency_kindslist of kindsKinds that age; [] turns recency off["note","conversation"]
cutoff_gapnumber 0–1Stop before a drop larger than this fraction0.5
min_resultsintegerNever gap-cut below this many results3
semantic_floornumberDrop semantic-only candidates below this cosine0.30
bands[strong, moderate, weak]Relevance bands for the generic profile; exactly three numbers[0.60, 0.45, 0.30]
derivationstringYour reason, shown by profiles show—

The reranker switch, its depth and graph routing are not overridable. Use the precise profile as a base to get reranking. A model’s own relevance bands, set in the model registry, take precedence over bands.

Checking an override

memo-mcp profiles show team-docs   # built-in profiles only in v1.4; see the note above
memo-mcp ui                        # search there with MEMO_PROFILE=team-docs set; the trace shows the profile

Before adopting an override, measure it. See Measuring a change.

Files on disk

PathCreated byModeContents
$MEMO_HOME/First writing command0700Base directory (default ~/.memo-mcp)
$MEMO_HOME/kb/First writing command0700Knowledge-base files
$MEMO_HOME/kb/<name>.dbFirst writing command0600The SQLite database: all layers, logs and audit
$MEMO_HOME/kb/<name>.db-walSQLite, while open0600Write-ahead log, checkpointed into the main file
$MEMO_HOME/kb/<name>.db-shmSQLite, while open0600Shared-memory index for the WAL
$MEMO_HOME/profiles.jsonYouyoursOptional profile overrides
~/.cache/memo-mcp/models/<org>_<model>/First use of a model0755Model files from Hugging Face
~/.cache/memo-mcp/models/<org>_<model>.okCompleted download—Marker; without it the download is redone

A directory that already existed with wider permissions is not tightened.

What is inside the database

GroupTables
L0–L2sources, documents, chunks, chunks_fts and chunks_fts_exact (FTS5), chunk_vecs (one row per passage per model)
L3facts, facts_fts, fact_vecs
L4entities, entity_aliases, mentions, merge_candidates, edges
L5pages, pages_fts, page_sources, page_vecs, work_items
Bookkeepingnamespaces, models, jobs, audit, query_log, call_log

Every column is described in schema.md. The file is plain SQLite. You can inspect it with the sqlite3 shell, for example sqlite3 ~/.memo-mcp/kb/default.db '.tables'. Do not write to it by hand: triggers keep the indexes consistent only for writes made through memo-mcp.

Nothing else

memo-mcp writes no PID files, no lock files of its own, no logs to disk and no temporary files outside the model cache. Logs go to stderr. Redirect them yourself if you want them on disk.

Part IV: Retrieval tuning

The defaults are measured, not guessed. Each constant has a written derivation, and every release carries an eval report. Tune only when you have a reason, and measure before and after.

SymptomFirst lever
Identifier or code queries misscode profile, or raise weights.exact
Paraphrased questions missCheck degraded and the model; try a stronger model
Old notes crowd out new onesrecency profile, or a shorter half_life_days
Lists are too long or too shortcutoff_gap, min_results, or the request’s limit
“No results” when there should be somesemantic_floor, and the model’s bands; look at filtered in the trace

Profiles and constants

A profile is a named set of ranking constants. The server and the UI use MEMO_PROFILE, default default. The CLI takes search --profile <name>. MCP clients cannot pick a profile per request in 1.x.

memo-mcp profiles show            # every profile, every constant, with its derivation
memo-mcp profiles show precise

Shipped profiles

ProfileWeights: sem / kw / exact / fact / entity / graphOther differencesUse for
default0.5 / 0.5 / 0.3 / 0.4 / 0.4 / 0.5—General use
precisesame as defaultfetch_depth 200, cutoff_gap 0.6, rerank on when a reranker is attachedRarely worded answers; shorter, surer lists
recencysame as defaultEvery kind ages; half-life 30 days; floor 0.6“What did I write lately”
code0.4 / 0.4 / 0.6 / 0.4 / 0.5 / 0.5No recencyIdentifier-heavy questions
minmaxsame as defaultMin-max score fusion instead of RRFExperiment
text-onlysame as defaultNo entity or graph arm (the 1.0 arms)Ablation
no-graphsame as defaultEntity arm without the graph walkAblation
keyword-onlykeyword 1.0 only—Ablation; the baseline to beat
semantic-onlysemantic 1.0 only—Ablation

The constants

ConstantDefaultRaise it to…Lower it to…
weights.<arm>see aboveLet that arm’s ranking count for moreMute an arm (0 disables it)
rrf_k60Flatten the difference between rank 1 and rank 10Reward top ranks more
fetch_depth100Keep rarely worded hits in the fused list, at some latency costSpeed up large KBs
half_life_days90Age more slowlyFavour recent items more strongly
recency_floor0.8Make age matter lessMake age matter more (floor is the minimum kept)
recency_kindsnote, conversationAge more kinds[] turns recency off
cutoff_gap0.5Cut later, giving longer listsCut sooner at a score cliff
min_results3Always show more before a gap cutAllow single-answer lists
semantic_floor0.30Abstain more readily on vague matchesKeep weaker meaning-only matches
bands0.60 / 0.45 / 0.30Generic relevance bands; per-model bands override them—

Worked example: fusion

A passage at keyword rank 1 and semantic rank 2 under default:

keyword   0.5 / (60 + 1) = 0.008197
semantic  0.5 / (60 + 2) = 0.008065
fused                    = 0.016262

A keyword-only hit at rank 1 scores 0.008197, the same as a semantic-only hit at rank 1. Equal weights therefore let either arm put a result on the first page. That is the main reason the default is balanced.

Graph routing

The entity and graph arms run only when the query names two or more known entities, or one entity with relational phrasing (“depends on”, “related to”, …). The trace’s routing_reason says why they ran. Ablations show that routing them on every query hurts single-entity lookups. That is why routing is not a tunable constant.

Customising

Overrides go in profiles.json. The derivations behind each default are in architecture §5–§6 and the eval reports.

Embedding models

The semantic arm needs an embedding model. Models run in-process through a pure-Go ONNX backend (hugot / GoMLX), so no Python, no ONNX Runtime and no GPU are involved. potion is a static lookup table with no neural network at all.

memo-mcp model ls
IdDimLicenceSizeQuery latency*Notes
granite-small-r2 (default)384Apache-2.0~195 MB~0.2 sIBM granite-embedding-small-english-r2; best measured quality; 8k context
minilm384Apache-2.0~87 MB~0.12 sThe previous default; no paraphrase recall in the bake-off
potion512MIT~130 MB~6 msmodel2vec table; instant; weaker on paraphrase; good fallback
granite-r2768Apache-2.0~600 MBslowRuns, but about 12× slower than MiniLM on this backend
arctic-m-v2768Apache-2.0~1.2 GBslowSnowflake; multilingual
gemma-256256Gemma Terms of Use~310 MBslowOpt-in only: not an Apache or MIT licence

* p50 query latency including query embedding, measured in docs/eval/v0.7.0.md. Writes cost more: about 0.6 s per passage with granite-small-r2 on this backend.

Each model has its own relevance bands in the registry. For granite-small-r2 they are 0.88 / 0.80 / 0.72, because unrelated text already scores about 0.63 under it.

How vectors are stored

Vectors are stored per passage per model in chunk_vecs(chunk_id, model_id, embedding). Several models can coexist in one file, so switching back is free once both are embedded. Each 384-dimension vector adds about 1.5 KB per passage.

Switching models

  1. Pre-download and check, optionally:
    memo-mcp model pull potion
    memo-mcp model smoke potion
    
  2. Embed the existing passages. This is optional, because the server does it in the background; doing it ahead avoids a degraded period:
    MEMO_KB=my-project memo-mcp reindex --model potion
    
    reindex is resumable. It only embeds passages that lack a vector for that model.
  3. Record the default for the file, and set MEMO_MODEL in the MCP client’s server entry:
    MEMO_KB=my-project memo-mcp model use potion
    
  4. Restart the client session. At start, the server runs backfill, then reindex for the current model, in the background. Until that finishes, search reports degraded and memo_kb_pending_embeddings{model="potion"} is above zero.

Downloads

  • Source: Hugging Face, over HTTPS, into ~/.cache/memo-mcp/models/.
  • A completed download writes a .ok marker. An interrupted one is detected and retried.
  • memo-mcp model redownload <id> discards the cache for one model and fetches it again. Always pass the id: without one it re-downloads minilm.
  • Metrics: memo_embed_model_downloads_total{model,outcome}, memo_embed_model_load_seconds, memo_embed_model_loaded.

Offline installs: Air-gapped installs.

The reranker

A cross-encoder can re-score the top results by reading the query and each passage together. memo-mcp ships one as an opt-in experiment:

SettingValue
Modelcross-encoder/ms-marco-MiniLM-L6-v2 (Apache-2.0, ~91 MB), id ms-marco-minilm
EnableMEMO_RERANK=1 (server and UI), or search --rerank (CLI)
Choose anotherMEMO_RERANKER=<id>
Used byProfiles with rerank on: precise, which re-scores the top 30

Why it is off by default

It failed its promotion gate. In the bake-off it lowered nDCG@10 by about 0.2 and added 0.7 to 4 seconds per query:

EmbedderCorpusdefault nDCG@10precise + rerank nDCG@10p50 ms
minilmnotes0.9770.732806
minilmkb0.8820.7242,333
granite-small-r2notes0.9780.6763,987

The pure-Go pipeline cannot pass sentence-pair segment ids, so the model’s scores are compressed near zero. Details: spike S8.

If you try it

MEMO_RERANK=1 MEMO_PROFILE=precise memo-mcp ui

Watch memo_search_rerank_duration_seconds and the trace’s rerank{model, top_n, latency_ms}. If the reranker fails to load, the server continues without it. Search then reports degraded with reason reranker_failed.

Measuring a change

Every retrieval change in memo-mcp’s own history was measured against a labelled baseline. Use the same tools before you change a constant, a profile or a model.

1. The built-in benchmark

memo-mcp eval --models hash,granite-small-r2 --profiles default,precise,code --corpus all
memo-mcp eval --models potion --profiles all --format md > potion.md
memo-mcp eval --explain-failures            # one line per missed query, and why

It loads two labelled corpora into a throwaway knowledge base, runs every query under each model and profile combination, and prints quality next to cost:

ColumnMeaning
recall@1/5/10Share of relevant items found in the top k
MRRMean reciprocal rank of the first relevant item
nDCG@10Ranking quality, rewarding relevant items near the top
abstentionShare of no-answer queries that correctly returned nothing
p50 ms, tokens p50Median latency and response size

hash is the deterministic test embedder and needs no download. Other model ids download real models. The corpora are the project’s own fixtures, so they tell you about relative effects (this profile against that one), not about your data.

2. Your own queries

  1. Turn on the opt-in log in the server’s environment: MEMO_QUERY_LOG=1.
  2. Use the system normally for a while.
  3. Export what was asked:
    memo-mcp log replay --n 200 > candidates.jsonl
    
    Each line is an unlabelled eval query, with the query and the addresses that were returned.
  4. Label the relevant addresses, then compare runs under different profiles in the UI or with memo-mcp search --profile <p> --format json.

3. Reading one result

memo-mcp explain "how do we rotate credentials"                  # ranking table for every hit
memo-mcp explain "how do we rotate credentials" memo://chunk/812  # the full why for one hit

What to look for:

FieldTells you
arms[]Which arms saw the result, at what rank, and what each contributed
recency_factorWhether age pushed it down
relevance, bandWhether the model thinks it is about the question at all
trace filteredHow many candidates scope, revocation, minimum trust or the semantic floor removed
trace cutoffWhether the list ended on a gap, the limit or the budget
trace degradedWhether a capability was missing

The UI’s search page shows the same table, with links.

4. In production

After changing a profile or a model, watch these in Prometheus:

  • memo_search_abstentions_total and memo_search_results: is the system answering less?
  • memo_search_duration_seconds: did latency move?
  • memo_search_cutoff_total{kind}: are lists ending differently?

See Prometheus: scrape, rules, alerts.

Part V: Knowledge lifecycle

How records enter, change, gain trust, get connected and summarised, and leave. Agents do most of this through MCP tools; an operator does the parts that need a human through the CLI.

Every write in every chapter lands in the audit table: who (actor), through which channel (tool, cli, elicitation, worker), what operation, on which record. The same events increment memo_store_writes_total{op,channel}.

Ingest and revisions

The write path

  1. Normalise line endings and whitespace, and hash the content (SHA-256).
  2. Deduplicate. If the source’s current revision has the same hash, nothing happens and the result says dedup: true. If the content differs, a new revision is created. The previous revision is marked superseded and stays readable with as_of.
  3. Chunk. The text is split by headings, then paragraphs, then sentences, then a hard split. The target is about 200 estimated tokens and the cap 400, clamped to the model’s limit, with about 40 tokens of overlap. Fenced code blocks are never split. Each passage gets a context header title > section path that the keyword index also sees.
  4. Index. Both FTS5 indexes are filled by triggers inside the transaction. Entity mentions are extracted and linked.
  5. Embed outside the transaction, in batches of 16. If no model is available, the passages are queued as a job, and backfill or the next server start embeds them. A write never reports success it did not get.
  6. Audit: one row per write.

A source is identified by its URI, or by its file path for CLI ingests. Re-ingesting the same URI with changed content, or with a new version, creates a new revision of that source.

From the CLI

memo-mcp ingest ./docs --ns handbook --kind doc
memo-mcp ingest page.md --uri https://pkg.go.dev/google.golang.org/grpc \
  --library grpc/grpc-go --version v1.64.0 --context "gRPC-Go API docs for deadlines"
cat notes.md | memo-mcp ingest - --kind note --title "Standup 2026-10-03"
memo-mcp ingest ./big-folder --embed=false && memo-mcp backfill   # fast load, embed later

CLI writes are trust user, or curated with --trust curated. Origin defaults to web when --uri is set, and to user-said otherwise.

Kinds

KindAges under recency?Typical content
docNoDocumentation, specs, handbooks; usually versioned
codeNoCode explanations and snippets
noteYesShort notes, decisions, observations
conversationYesConversation summaries

Costs

StepCost
Parse, chunk, indexAbout 60 documents per second without embedding
Embedding with granite-small-r2About 0.6 s per passage on the pure-Go backend
Embedding with potionNear instant
DiskAbout 4.4 KB per passage before vectors, plus about 1.5 KB per passage per 384-dimension model

For bulk loads, ingest with --embed=false, then run memo-mcp backfill, or let the server embed in the background. See Capacity and performance.

Observing it

memo_store_ingests_total{outcome=new|revision|dedup|error}, memo_store_ingest_duration_seconds, memo_store_chunks_written_total, memo_store_embed_batches_total{outcome}, and memo_kb_pending_embeddings{model} for the backlog.

Facts and time

A fact is one sentence with an optional subject list, an evidence passage and a validity window. Facts are searched by their own arm. A matching fact votes for its evidence passage.

memo-mcp remember "Production runs Postgres 16" --about postgres,production \
  --evidence memo://chunk/812 --valid-from 2026-06-01
memo-mcp facts ls --ns default
memo-mcp facts ls --as-of 2026-05-15 --history

Correcting facts

  • Supersede: remember "<new>" --supersedes memo://fact/<old>. The old fact is invalidated, not deleted, and remains visible under as_of and --history. Agents may supersede only agent-trust facts.
  • Expire: set --valid-to. memo-mcp lint lists facts past their window.
  • Retire: memo-mcp forget memo://fact/<id> --reason "…".

Two clocks

ClockColumnsQuestion it answers
Recorded timerecorded_at, invalidated_atWhen did the knowledge base learn or drop this?
Valid timevalid_from, valid_toWhen is this true in the world?

as_of=T, available on the search and explore tools and on facts ls --as-of and explore --as-of in the CLI, shows what was recorded on or before T and not superseded before T. For facts with a window, it also applies valid time.

Conflicts

Two live facts about the same subject, with overlapping windows and differing statements, form a conflict. compact turns each conflict into a work item, together with the rule the server would apply: higher trust wins, then the newer fact. A human or agent settles it with submit <item> --keep memo://fact/<id>. When conflicting facts both match a search and score within 10% of each other, trust breaks the tie.

Trust administration

From → ToMCP tool callCLI (human)Elicitation dialog (human)
create agentyes——
create user / curatednoyes (--trust)—
agent → user / curatedno; promote askstrust promoteyes, when accepted
lower trustnotrust demote—
forget or supersede agent recordsyesyes—
forget or supersede user / curated recordsnoyes—

Commands

memo-mcp trust ls                                      # everything above agent trust
memo-mcp trust promote memo://doc/0199… --to curated
memo-mcp trust demote  memo://fact/0199… --to agent

The promote tool

When an agent calls promote:

  1. If the client supports elicitation, it shows the human a dialog with the excerpt, the source and the target level. Accepting it applies the change through the elicitation channel. Declining or cancelling changes nothing.
  2. If the client does not support elicitation, the tool returns the exact memo-mcp trust promote … command for a human to run.

Outcomes are counted in memo_mcp_elicitations_total{tool="promote",outcome=asked|accept|decline|cancel|unsupported}.

Warning. Hooks or settings that auto-accept elicitation dialogs defeat this gate. If you run such a client, do not rely on trust levels.

Auditing trust

Every change is an audit row, with op promote or demote and the channel it came through. Inspect it with sqlite3:

sqlite3 ~/.memo-mcp/kb/default.db \
  "SELECT datetime(ts/1000,'unixepoch'), actor, channel, op, target_uri FROM audit WHERE op IN ('promote','demote') ORDER BY ts DESC LIMIT 20"

The graph index

The graph layer is an index, not a source of truth. It records which entities each passage mentions. Search uses it to find passages that connect things the query names, and explore uses it to walk from one name.

Extraction and resolution

  • Extraction runs on every ingest, using heuristics only: backticked spans, dotted, camel and snake identifiers, headings and capitalised names. Clients may also declare entities[] and relations[] when they ingest.
  • Resolution is deterministic. The same normalised key means the same entity. A near key, with 3-gram Jaccard of at least 0.8, not too short and not differing only in a number, becomes a merge candidate for a human. Anything else is a new entity.

Commands

memo-mcp explore "Quorum Replication" --hops 2
memo-mcp graph merges                 # open candidates with their similarity
memo-mcp graph merge 17               # same thing: merge the names
memo-mcp graph reject 18              # different things: keep apart
memo-mcp graph rebuild --ns default   # re-extract mentions

Nothing is merged without a decision. Agents can propose decisions through compact and submit. Merge precision is gated at 0.95 or better in CI.

When to rebuild

  • Once, after upgrading a file written before 1.1.
  • After a release whose notes mention an extraction change.
  • If memo-mcp lint reports many orphan entities after large deletions.

Rebuild keeps merge decisions. It re-reads every live passage, so on large knowledge bases run it when the system is quiet.

Runtime behaviour

The server keeps an in-memory mention graph per namespace. It is built on first use and reused while the namespace’s count of live mentions is unchanged; any ingest, revision or forget that changes that count triggers a rebuild on the next routed query. as_of queries always build a graph for that point in time and do not cache it. memo_graph_cache_total{event=hit|build|build_asof} and memo_graph_build_duration_seconds show how often this happens and what it costs.

Compaction and pages

The server never writes prose. Compaction is a queue of work items that an agent, or a person, completes. The server checks the result before storing it.

Work items

compact, through the tool or memo-mcp compact, scans a namespace and creates items:

KindCreated when
pageAn entity has 3 or more live passages and no page
staleA page’s sources changed or were forgotten, or the entity gained 2 or more passages since the page was built
conflictTwo live facts about one subject overlap in time and disagree
mergeAn open merge candidate exists
duplicateTwo passages from different documents are near duplicates (word 3-gram Jaccard of at least 0.75)

Each item’s payload carries everything needed to do it: the passages, the facts and the previous page. No second call is needed.

Submitting

memo-mcp compact --ns default --kinds page,stale --json
memo-mcp submit 42 --content-file page.md --title "Quorum replication" --dry-run
memo-mcp submit 42 --content-file page.md --title "Quorum replication"
memo-mcp submit 43 --keep memo://fact/0199…
memo-mcp submit 44 --accept            # merge item
memo-mcp submit 45 --skip --reason "not worth a page"

Every page submission reports:

  • diff: a line diff against the previous page.
  • omission check: recorded facts about the subject whose content words are mostly absent from the page. Less than 75% coverage is flagged.
  • corruption check: page sentences that share fewer than half of their content words with any source.

Pages are stored as is_inference = 1, with the passages they cite. They are searchable with granularity: page, and marked stale when a source changes. Raw rows are never modified.

Lint

memo-mcp lint --json

Reports contradictions, orphan entities, missing pages, stale pages and expired facts, with addresses. Lint changes nothing.

The optional Ollama executor

For unattended page writing from the terminal:

MEMO_OLLAMA_MODEL=qwen2.5:7b-instruct memo-mcp compact --kinds page --executor ollama          # dry run
MEMO_OLLAMA_MODEL=qwen2.5:7b-instruct memo-mcp compact --kinds page --executor ollama --apply  # submit
  • The endpoint is MEMO_OLLAMA_URL, default http://127.0.0.1:11434. It must be loopback unless --allow-remote is passed.
  • Use a non-thinking instruction model. Every page still goes through the same checks.
  • No MCP code path can reach the executor.

Observing it

memo_kb_work_items_open, memo_kb_pages, memo_kb_pages_stale, memo_store_work_items_total{kind,event} and memo_store_pages_marked_stale_total.

Forget and redact

memo-mcp forget memo://doc/0199… --reason "superseded by the 2026 policy"
memo-mcp forget memo://fact/0199… --reason "wrong owner" --redact
forgetforget --redact
Leaves every index (keyword, exact, vectors, graph)yesyes
Excluded from search, including under as_ofyesyes
Address resolves to “forgotten on … because …”yesyes
Text kept in the fileyesno: cleared
Pages built from itmarked stalemarked stale
Audit rowyesyes
  • --reason is required and is shown to anyone who dereferences the address later.
  • MCP tool calls may forget only agent-trust records. The CLI may forget anything.

Reclaiming disk space

SQLite does not shrink the file when rows are cleared. After large redactions, compact it while no process has the file open:

sqlite3 ~/.memo-mcp/kb/my-project.db 'VACUUM'

VACUUM also removes freed pages that might still hold redacted bytes. Run it after redacting secrets.

Part VI: Monitoring and observability

All observability in memo-mcp is local and pull-based. Metrics are served on loopback, logs go to stderr, and the optional call log is a table in your own file. Nothing is pushed anywhere, and there is no telemetry.

Signals overview

SignalWhereOn by defaultScope
Metrics (Prometheus text 0.0.4)serve --metrics-addr and UI GET /metricsIn memory, always; exposed only when askedThe process that serves them
Logs (log/slog, text or JSON)stderrYes, at infoThe process
Call log and query logcall_log and query_log tablesNo (MEMO_QUERY_LOG=1)The file: every process writing to it
Snapshot (memo-mcp metrics)stdoutOn demandThe file: gauges plus log statistics
Auditaudit tableAlwaysThe file: every write

The per-process caveat

The MCP server, the UI and each CLI command are separate processes. Counters such as memo_mcp_tool_calls_total live in the memory of the process that served the calls:

  • A server process exists only while its client session is open. Its counters start at zero with each session. Scrape it while it runs; Prometheus rate() handles the resets.
  • The UI has its own counters, mostly memo_ui_*. Like every process, it also serves the memo_kb_* gauges, which are read from the file.
  • History across sessions comes from the file: memo-mcp metrics (with MEMO_QUERY_LOG=1) and the memo_kb_* gauges.
GoalSetup
Glance at health nowmemo-mcp metrics, or memo-mcp status
Trends of knowledge-base size, backlog and pagesRun memo-mcp ui --addr 127.0.0.1:9470 as a long-lived process and scrape its /metrics for memo_kb_*
Live tool and search latency, errors, degraded modeAdd --metrics-addr to one server entry and scrape it while sessions run
Per-call history and slow-call forensicsMEMO_QUERY_LOG=1, then memo-mcp log calls and memo-mcp metrics --since 7d

Naming

  • Prefix memo_; counters end in _total; units are suffixes (_seconds, _bytes), and seconds are never milliseconds.
  • Labels are bounded enums: tool names, outcomes, arms, models, namespaces. Never query text, ids or paths.
  • OpenTelemetry mapping: replace _ with . and drop _total. For example, memo_mcp_tool_calls_total maps to memo.mcp.tool_calls.

The metrics endpoint

On the MCP server

memo-mcp serve --metrics-addr 127.0.0.1:9469
# or, in the client's server entry:  "env": { "MEMO_METRICS_ADDR": "127.0.0.1:9469" }

Claude Code example:

claude mcp add memo --env MEMO_KB=my-project --env MEMO_METRICS_ADDR=127.0.0.1:9469 -- /path/to/memo-mcp

At start, the server logs metrics listening url=http://127.0.0.1:9469/metrics.

PropertyBehaviour
BindLoopback only (127.0.0.1, ::1, localhost). Any other address is refused at start; there is no override
MethodsGET and HEAD; anything else returns 405
Paths/metrics only; anything else returns 404
Host headerMust match the bound address, which defends against DNS rebinding; otherwise 403
Content typetext/plain; version=0.0.4; charset=utf-8
stdoutUntouched: the MCP stream stays clean

Warning: one port, one process. The address is bound when the server starts. If a second session starts with the same entry while the first is running, for example two Claude Code windows on one project, the second fails with fatal: listen tcp 127.0.0.1:9469: bind: address already in use. Until this is relaxed:

  • set --metrics-addr on one entry you use for a single long session;
  • or give each knowledge base’s entry its own port;
  • or leave it off and use the UI’s endpoint for file-level gauges plus the call log for per-call data.

On the UI

memo-mcp ui always serves GET /metrics, under the same guards as the rest of the UI: loopback, GET only, and the Host check. Pin the port to scrape it:

MEMO_KB=my-project memo-mcp ui --addr 127.0.0.1:9470 --no-model
curl -s http://127.0.0.1:9470/metrics | grep memo_kb_documents_live

--no-model keeps a long-running UI light. Its search is then keyword-only, which is fine for a metrics sidecar.

Checking it

curl -s http://127.0.0.1:9469/metrics | head -20
curl -s -o /dev/null -w '%{http_code}\n' -X POST http://127.0.0.1:9469/metrics   # 405
curl -s -o /dev/null -w '%{http_code}\n' -H 'Host: example.com' http://127.0.0.1:9469/metrics  # 403

memo_kb_* gauges are read from the database at scrape time and cached for 5 seconds, so frequent scrapes do not load the file. Read failures increment memo_store_status_scrape_errors_total.

Metric catalogue

Every series memo-mcp exposes, as of v1.4. Types: C counter, G gauge, H histogram. A histogram exposes _bucket{le}, _sum and _count. A family with no observations yet shows only its # HELP and # TYPE lines.

Bucket sets:

SetBoundaries
latency (s)0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10
tokens50, 100, 200, 500, 1000, 2000, 4000, 8000, 16000
count0, 1, 2, 5, 10, 20, 50, 100, 200, 500
bytes1 KiB, 4 KiB, 16 KiB, 64 KiB, 256 KiB, 1 MiB, 4 MiB

MCP server

MetricTypeLabelsMeaning
memo_mcp_requests_totalCmethod, outcome = ok, errorEvery MCP request received (tools/call, resources/read, server/discover, …)
memo_mcp_requests_in_flightG—Requests being served now
memo_mcp_tool_calls_totalCtool, outcome = ok, tool_error, input_required, errorTool calls. input_required is promote waiting for a human
memo_mcp_tool_call_duration_secondsH latencytoolTool call latency
memo_mcp_tool_result_tokensH tokenstoolEstimated tokens in each tool result: the cost the agent pays
memo_mcp_tool_errors_totalCtool, class = not_found, forgotten, needs_human, canceled, invalid_args, internalWhy calls failed
memo_mcp_elicitations_totalCtool, outcome = asked, accept, decline, cancel, unsupportedHuman approval dialogs
memo_mcp_resource_reads_totalCkind = doc, chunk, source, fact, entity, page, index, ns_index, unknown; outcome = ok, errormemo:// resource reads
memo_mcp_sessions_totalCclientSessions started, by client name
MetricTypeLabelsMeaning
memo_search_totalCmode_requested, mode_resolved, granularity, outcome = results, abstainSearches that completed (failed searches show up as tool errors)
memo_search_duration_secondsH latencygranularityEnd-to-end search latency
memo_search_arm_duration_secondsH latencyarmTime in each retrieval arm
memo_search_arm_candidatesH countarmCandidates each arm returned before fusion
memo_search_resultsH countgranularityResults returned per search
memo_search_truncated_results_totalC—Ranked results left out by the token budget
memo_search_cutoff_totalCkind = gap, budget, limit, noneWhy lists ended
memo_search_degraded_totalCreason = no_embedder, query_embed_failed, reranker_failed, otherSearches served without a capability
memo_search_abstentions_totalCreasonSearches that returned nothing on purpose
memo_search_entities_matchedH count—Known entities named per query
memo_search_rerank_duration_secondsH latencymodelCross-encoder time, only with the reranker
memo_graph_cache_totalCevent = hit, build, build_asofIn-memory mention-graph cache
memo_graph_build_duration_secondsH latency—Time to build a namespace graph

Store (write path)

MetricTypeLabelsMeaning
memo_store_writes_totalCop, channel = tool, cli, elicitation, workerEvery audited write
memo_store_ingests_totalCoutcome = new, revision, dedup, errorIngest calls
memo_store_ingest_duration_secondsH latency—Ingest latency, embedding included
memo_store_chunks_written_totalC—Passages written
memo_store_document_bytesH bytes—Size of ingested documents
memo_store_vectors_stored_totalCmodelVectors stored
memo_store_embed_batches_totalCoutcome = ok, error, queuedEmbedding batches on the write path
memo_store_jobs_totalCkind, event = queued, done, failedBackground jobs, for example the embed backlog
memo_store_mentions_linked_totalC—Entity mentions linked to passages
memo_store_pages_marked_stale_totalC—Pages marked stale by a source change or forget
memo_store_work_items_totalCkind, event = created, done, skippedCompaction work items
memo_store_status_scrape_errors_totalC—Failures reading the gauges below

Knowledge base (gauges read from the file, cached for 5 s)

MetricLabelsMeaning
memo_kb_sources—Sources
memo_kb_documents_live—Live documents (latest revision, not forgotten)
memo_kb_revisions—Document revisions stored
memo_kb_chunks—Live passages
memo_kb_facts—Live facts
memo_kb_entities—Entities
memo_kb_mentions—Entity mentions
memo_kb_edges—Typed edges
memo_kb_merge_review_open—Open merge candidates awaiting a human
memo_kb_pages—Curated pages
memo_kb_pages_stale—Stale pages
memo_kb_work_items_open—Open compaction work items
memo_kb_jobs_queued—Jobs queued or running
memo_kb_jobs_failed—Jobs failed
memo_kb_pending_embeddingsmodelPassages without a vector for that model
memo_kb_schema_version—Schema version of the file
memo_kb_db_size_bytes—Size of the main SQLite file, excluding the WAL
memo_kb_last_write_timestamp_seconds—Unix time of the last audited write
memo_kb_namespace_documentsnamespaceLive documents per namespace
memo_kb_namespace_chunksnamespaceLive passages per namespace

All are gauges.

Embedding

MetricTypeLabelsMeaning
memo_embed_duration_secondsH latencymodel, role = query, docEmbedding call latency
memo_embed_texts_totalCmodel, roleTexts embedded
memo_embed_errors_totalCmodel, roleEmbedding calls that failed
memo_embed_batch_sizeH countmodelTexts per call
memo_embed_model_downloads_totalCmodel, outcome = ok, error, corrupt_retryModel downloads
memo_embed_model_load_secondsGmodelTime the last model load took
memo_embed_model_loadedGmodel, backend = hugot, static1 while a model is loaded in this process

Web UI

MetricTypeLabelsMeaning
memo_ui_requests_totalCroute (for example GET /search), statusUI requests
memo_ui_request_duration_secondsH latencyrouteUI render time

Runtime and build

MetricTypeLabelsMeaning
memo_build_infoG (always 1)version, go_version, mcp_protocol, goos, goarchWhat is running
go_infoGversionGo version
go_goroutinesG—Goroutines
go_memstats_heap_alloc_bytesG—Live heap
go_memstats_sys_bytesG—Memory obtained from the OS
go_gc_cycles_totalC—Completed GC cycles
go_gc_pauses_secondsH—GC pause distribution
process_start_time_secondsG—Process start time; use it to spot session restarts

Prometheus: scrape, rules, alerts

These are starting points written against the metric catalogue. Adapt the thresholds to your usage. The scrape configuration and both rule groups pass promtool check config and promtool check rules (Prometheus 3.15). Re-check them after you edit them.

Scrape configuration

scrape_configs:
  # Long-lived UI sidecar: knowledge-base gauges, always up.
  - job_name: memo-kb
    scrape_interval: 60s
    static_configs:
      - targets: ["127.0.0.1:9470"]
        labels: { kb: my-project }

  # MCP server: live tool and search metrics; only up while a client session runs.
  - job_name: memo-mcp
    scrape_interval: 15s
    static_configs:
      - targets: ["127.0.0.1:9469"]
        labels: { kb: my-project }

Prometheus must run on the same machine, because both endpoints are loopback-only. A Grafana Alloy or OpenTelemetry Collector agent on the machine can scrape them and forward the samples. That forwarding is your choice; memo-mcp itself never pushes.

Do not alert on up{job="memo-mcp"} == 0. The server exists only while a client session is open, so being down is normal. Alert on up{job="memo-kb"} instead, if you rely on the sidecar.

Recording rules

groups:
  - name: memo-recording
    interval: 1m
    rules:
      - record: memo:tool_calls:rate5m
        expr: sum by (kb, tool, outcome) (rate(memo_mcp_tool_calls_total[5m]))
      - record: memo:tool_error_ratio:rate5m
        expr: |
          sum by (kb) (rate(memo_mcp_tool_calls_total{outcome="error"}[5m]))
          /
          sum by (kb) (rate(memo_mcp_tool_calls_total[5m]))
      - record: memo:tool_latency_seconds:p95_5m
        expr: histogram_quantile(0.95, sum by (kb, tool, le) (rate(memo_mcp_tool_call_duration_seconds_bucket[5m])))
      - record: memo:search_latency_seconds:p95_5m
        expr: histogram_quantile(0.95, sum by (kb, le) (rate(memo_search_duration_seconds_bucket[5m])))
      - record: memo:search_abstain_ratio:rate15m
        expr: |
          sum by (kb) (rate(memo_search_total{outcome="abstain"}[15m]))
          /
          sum by (kb) (rate(memo_search_total[15m]))
      - record: memo:arm_latency_seconds:p95_5m
        expr: histogram_quantile(0.95, sum by (kb, arm, le) (rate(memo_search_arm_duration_seconds_bucket[5m])))

Alerts

groups:
  - name: memo-alerts
    rules:
      - alert: MemoToolErrorsHigh
        expr: memo:tool_error_ratio:rate5m > 0.05
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "memo-mcp {{ $labels.kb }}: more than 5% of tool calls fail"
          description: "Check memo_mcp_tool_errors_total by class and the server's stderr."

      - alert: MemoSearchSlow
        expr: memo:search_latency_seconds:p95_5m > 2
        for: 15m
        labels: { severity: warning }
        annotations:
          summary: "memo-mcp {{ $labels.kb }}: p95 search latency above 2 s"
          description: "Look at memo:arm_latency_seconds:p95_5m to find the slow arm."

      - alert: MemoSearchDegraded
        expr: sum by (kb, reason) (increase(memo_search_degraded_total[30m])) > 0
        for: 30m
        labels: { severity: warning }
        annotations:
          summary: "memo-mcp {{ $labels.kb }}: searches degraded ({{ $labels.reason }})"
          description: "The embedding model is missing or failing; search is keyword-only."

      - alert: MemoEmbeddingBacklog
        expr: max by (kb, model) (memo_kb_pending_embeddings) > 0
        for: 2h
        labels: { severity: info }
        annotations:
          summary: "{{ $value }} passages lack vectors for {{ $labels.model }}"
          description: "Run memo-mcp backfill, or start a server session to drain the backlog."

      - alert: MemoJobsFailed
        expr: max by (kb) (memo_kb_jobs_failed) > 0
        labels: { severity: warning }
        annotations:
          summary: "memo-mcp {{ $labels.kb }}: background jobs failed"

      - alert: MemoDatabaseGrowth
        expr: delta(memo_kb_db_size_bytes[1d]) > 500e6
        labels: { severity: info }
        annotations:
          summary: "memo-mcp {{ $labels.kb }}: knowledge base grew more than 500 MB in a day"

      - alert: MemoMergeQueueBacklog
        expr: max by (kb) (memo_kb_merge_review_open) > 50
        for: 1d
        labels: { severity: info }
        annotations:
          summary: "{{ $value }} merge candidates waiting for review (memo-mcp graph merges)"

      - alert: MemoKBSidecarDown
        expr: up{job="memo-kb"} == 0
        for: 10m
        labels: { severity: info }
        annotations:
          summary: "memo-mcp UI sidecar for {{ $labels.kb }} is not running"

Dashboard queries

PanelPromQL
Tool calls by toolsum by (tool) (rate(memo_mcp_tool_calls_total[5m]))
Tool errors by classsum by (class) (rate(memo_mcp_tool_errors_total[15m]))
p50 and p95 tool latencyhistogram_quantile(0.5, sum by (le) (rate(memo_mcp_tool_call_duration_seconds_bucket[5m]))) (and 0.95)
Result size the agent pays forhistogram_quantile(0.9, sum by (tool, le) (rate(memo_mcp_tool_result_tokens_bucket[15m])))
Search outcomessum by (outcome) (rate(memo_search_total[15m]))
Mode resolutionsum by (mode_resolved) (rate(memo_search_total[1h]))
Where time goes, per armhistogram_quantile(0.95, sum by (arm, le) (rate(memo_search_arm_duration_seconds_bucket[5m])))
Why lists endsum by (kind) (rate(memo_search_cutoff_total[1h]))
Budget truncationrate(memo_search_truncated_results_total[1h])
Knowledge-base sizememo_kb_documents_live, memo_kb_chunks, memo_kb_facts
Per namespacememo_kb_namespace_documents
Backlogmemo_kb_pending_embeddings, memo_kb_jobs_queued, memo_kb_work_items_open, memo_kb_pages_stale
File sizememo_kb_db_size_bytes
Writes by channelsum by (channel) (rate(memo_store_writes_total[1h]))
Human approvalssum by (outcome) (increase(memo_mcp_elicitations_total[1d]))
Embedding costhistogram_quantile(0.95, sum by (model, role, le) (rate(memo_embed_duration_seconds_bucket[5m])))
Session restartschanges(process_start_time_seconds{job="memo-mcp"}[1d])
Version runningmemo_build_info

Logs

memo-mcp logs with Go’s log/slog, to stderr only. In server mode, stdout carries the MCP JSON-RPC stream and is never written to by anything else; a test enforces this.

VariableValuesDefault
MEMO_LOG_FORMATtext, jsontext
MEMO_LOG_LEVELdebug, info, warn, errorinfo

An unknown value falls back to the default and prints one warning.

What is logged

LevelEventFields
INFOtool call: one per MCP tool calltool, client, latency_ms, outcome, n_results, tokens_out, and error_class and err on failure
INFOsearch: one per search, including CLI and UI searchesmode, resolved, granularity, outcome, n_results, truncated, cutoff, latency_ms, profile, and degraded and entities when present
DEBUGsearch detailrouting, latency_ms_per_arm, candidates_per_arm, scope
INFOStart-up and background workmetrics listening, backfilled chunk vectors, embedded passages for the current model, reranker loaded
WARNDegraded or deprecatedembedding unavailable…, backfill failed, reindex failed, reranker unavailable…, download retries, licence notes, deprecated variables and flags
ERRORUnexpected failures—

Logs never contain document text, fact statements, or forget reasons. Search log lines do not include the query text.

Example, with MEMO_LOG_FORMAT=json:

{"time":"2026-10-03T16:23:45.136+02:00","level":"INFO","msg":"search","mode":"auto","resolved":"semantic+keyword+fact","granularity":"chunk","outcome":"results","n_results":1,"truncated":0,"cutoff":"none","latency_ms":"461.8","profile":"default"}

latency_ms is a string with one decimal place.

Where logs end up

  • Claude Code shows a stdio server’s stderr. Run claude --debug to see it live.

  • Claude Desktop writes each server’s stderr to its own log file. On macOS this is ~/Library/Logs/Claude/mcp-server-memo.log.

  • CLI commands print logs to your terminal’s stderr. Every memo-mcp search prints its INFO search line. Use MEMO_LOG_LEVEL=warn for quiet interactive use, or 2>/dev/null.

  • To capture a server’s logs yourself, wrap the binary in a small script:

    #!/bin/sh
    exec /path/to/memo-mcp "$@" 2>>"$HOME/.memo-mcp/serve.log"
    

    Point the client’s command at the script. Rotate the file with your usual tool.

MCP logging capability

memo-mcp does not advertise the MCP logging capability. It is deprecated in protocol 2026-07-28, and stdio hosts already surface stderr. Do not expect log notifications over the protocol.

Call and query logs

With MEMO_QUERY_LOG=1 in a process’s environment, that process records:

  • every search in query_log: arguments, the addresses returned, scores and the trace;
  • every MCP tool call in call_log: client, tool, latency, outcome, error class, result count, tokens out and an argument summary.

Both tables live in the knowledge-base file, so they persist across sessions and processes. They are off by default.

What is stored, and what never is

StoredNever stored
Search query text and its scope, mode, granularity and as_ofPassage, document or page text returned
Returned addresses and scorescontent of ingest (only its length, as content_len)
Tool name, client name, latency, outcome, error classstatement of remember
Allowlisted arguments: ids, namespaces, flags, dates, and for ingest the source object (URI, title, kind, library, version)reason of forget and submit, and context

The argument summary is an allowlist per tool, so a new argument is not logged until it is added explicitly. Search queries are stored verbatim. Treat the logs as sensitive as the questions people ask.

Reading the logs

memo-mcp log tail --n 20            # recent searches, then recent tool calls
memo-mcp log calls --n 50           # tool calls: time, client, tool, latency, results, tokens, args
memo-mcp log show 128               # one search with its full trace
memo-mcp log replay --n 200         # searches as JSON lines, for building eval queries
memo-mcp metrics --since 7d         # per-tool calls, errors, p50/p95/max latency; search aggregates
memo-mcp metrics --json | jq .

The UI’s /log page shows both tables.

Retention

The logs are not pruned automatically. Prune them on a schedule:

memo-mcp log prune                  # both logs: keep at most 10,000 rows and nothing older than 30 days

For example, with cron:

0 3 * * *  MEMO_KB=my-project /path/to/memo-mcp log prune

Querying with SQL

sqlite3 -header -column ~/.memo-mcp/kb/my-project.db "
  SELECT tool, COUNT(*) AS calls, SUM(ok = 0) AS errors, ROUND(AVG(latency_ms)) AS avg_ms
  FROM call_log WHERE ts > (strftime('%s','now','-1 day') * 1000)
  GROUP BY tool ORDER BY calls DESC"

ts is in Unix milliseconds. Column reference: schema §5.8.

Part VII: Operations

Day-two work: keeping data safe, checking integrity, planning capacity, getting data out, and fixing things.

Backup and restore

A knowledge base is one SQLite file in WAL mode. Back it up with SQLite’s own tools, not with a plain file copy while processes have it open: a copy of the main file alone can miss committed transactions that are still in the -wal file.

While memo-mcp is running (safe online backup)

DB=~/.memo-mcp/kb/my-project.db
sqlite3 "$DB" ".backup '/backups/my-project-$(date +%F).db'"
# or, which also compacts the copy:
sqlite3 "$DB" "VACUUM INTO '/backups/my-project-$(date +%F).db'"

Both take a consistent snapshot while readers and writers continue. Writers wait at most memo-mcp’s 5-second busy timeout. sqlite3 creates the copy with your umask, typically 0644. Restrict it with chmod 600, because the backup holds everything the knowledge base holds.

While nothing is running

When no memo-mcp process has the file open, -wal and -shm are empty or absent, and copying <name>.db is enough. If a -wal file with content remains, copy it alongside, or run sqlite3 <name>.db 'PRAGMA wal_checkpoint(TRUNCATE)' first.

What to back up

PathBack up?
$MEMO_HOME/kb/*.dbYes: everything lives here, including logs and audit
$MEMO_HOME/profiles.jsonYes, if you use it
~/.cache/memo-mcp/models/No: it is re-downloadable. Keep a copy only for air-gapped machines

Restore

  1. Stop every process using the knowledge base: close the client sessions, the UI and running commands.
  2. Remove the stale -wal and -shm files next to the target, if present.
  3. Copy the backup to $MEMO_HOME/kb/<name>.db, then chmod 600 it.
  4. Check it:
    MEMO_KB=<name> memo-mcp verify
    MEMO_KB=<name> memo-mcp status
    

A backup taken by an older binary is migrated forward on first write. A backup from a newer binary is refused by an older one.

A human-readable copy

memo-mcp export --md <dir> writes every live document as markdown with full provenance front matter. It is not a complete backup: there are no facts, history, graph or logs. But it is readable without memo-mcp, and re-ingesting it creates no new revisions. See Exporting.

Integrity and repair

verify

memo-mcp verify
memo-mcp verify --repair
CheckProblem reported--repair does
Every live document has passagesN live document(s) have no chunksReports only; re-ingest the source
Vectors point at existing passagesN vector(s) point at missing chunksReports only
Facts point at existing evidenceN fact(s) point at missing evidence chunksReports only
Every live passage has a vector for the current modelN live chunk(s) have no vector for model XQueues an embed job; run memo-mcp backfill
FTS5 index integrity<index> failed integrity-checkRebuilds the index

The command prints ok: no problems found, or one problem: line per finding and one repaired: line per fix.

backfill and reindex

CommandEmbedsWhen
memo-mcp backfillPassages queued for vectors, from --embed=false, a degraded write or verify --repairAfter bulk loads, or after the model was unavailable
memo-mcp reindex [--model id]Every live passage lacking a vector for the modelAfter switching models

Both are resumable and safe to interrupt. The server runs both in the background at start.

SQLite-level checks

sqlite3 ~/.memo-mcp/kb/my-project.db 'PRAGMA integrity_check'   # expect: ok
sqlite3 ~/.memo-mcp/kb/my-project.db 'PRAGMA user_version'      # schema version

When to run what

After…Run
A crash or power lossverify, then PRAGMA integrity_check
Restoring a backupverify
Bulk ingest with --embed=falsebackfill
Changing MEMO_MODELreindex --model <id>
Upgrading across 1.0 to 1.1graph rebuild
Large redactionsVACUUM with no process attached

Capacity and performance

Numbers measured on an Apple-silicon laptop with the shipped defaults. They are a guide to proportions, not a benchmark of your hardware.

Disk

ItemSize
Fixed overhead of an empty knowledge base~0.3 MB
Per passage, before vectors (text, two FTS5 indexes, graph rows)~4.4 KB
Per passage per 384-dimension model (granite-small-r2, minilm)~1.5 KB
Per passage per 768-dimension model~3 KB
Typical passages per KB of markdown~1 per 0.9 KB
Default model in the cache~195 MB

Example: this project’s docs/ and articles/ folders (756 KB of markdown) became 101 documents and 818 passages, in a 3.6 MB file before vectors. Vectors for one 384-dimension model add about 1.2 MB more.

Time

OperationCost
Ingest without embedding~60 documents per second
Embedding a passage, granite-small-r2~0.6 s
Embedding a passage, minilm~0.25 s
Embedding, potionnear instant
Search p50, granite-small-r2 (includes embedding the query)~0.2 s
Search p50, potion~6 ms
Search p50, keyword-only~1–3 ms
Reranker (precise + MEMO_RERANK=1)+0.7–4 s per query

Query embedding dominates search latency. Retrieval itself is milliseconds at these sizes. memo_search_arm_duration_seconds{arm} shows the split on your data.

Memory

A server process holds the embedding model in memory: a few hundred MB for granite-small-r2, less for potion. It also holds one in-memory mention graph per namespace it has searched with graph routing. go_memstats_sys_bytes shows the total.

Scaling guidance

  • Everything is single-file SQLite with exact (brute-force) vector search. Tens of thousands of passages per knowledge base are comfortable. Beyond about 100,000 passages, watch memo_search_arm_duration_seconds{arm="semantic"}, and split knowledge bases by project.
  • Bulk loads: ingest with --embed=false, then backfill. Or choose potion for fast, slightly weaker semantics.
  • Several sessions on one file are fine. Writes are serialised and short; embedding happens outside the write lock.

Exporting

Markdown with provenance

memo-mcp export --md ./kb-export            # all namespaces
memo-mcp export --md ./kb-export --ns handbook

The layout is <namespace>/<kind>/<slug>-<shortid>.md, plus an _index.md per namespace. Each file has YAML front matter:

---
memo_uri: memo://doc/01a1017e-ec2c-777a-b64a-4b67d1ac8f97
title: Deploy checklist
kind: doc
namespace: default
revision: 1
content_hash: sha256:cf966c40e2f66a5d595ab316004e2875ff5f70aa51679c31669852fc121d7326
fetched_at: 2026-10-03T11:20:57Z
trust: user
origin: user-said
---

Empty fields are left out. source_uri, library, version and context appear when they are set.

The export opens as an Obsidian vault. Re-ingesting it produces zero new revisions, because the content hashes match. Facts, history, graph and logs are not included.

An index for agents

memo-mcp export --index > kb-index.md                        # ≤ 8 KB, one line per document
memo-mcp export --index --library grpc/grpc-go@v1.64.0 --max-bytes 4096

Each line has the title, address, kind, version and trust, and facts are summarised. Lines that do not fit the budget are counted in a footer. Paste the output into CLAUDE.md or AGENTS.md so an agent knows what exists before it searches. The same content is the MCP resource memo://index.

Raw data

The file is standard SQLite, so any tool can read it. For example, every live fact as CSV:

sqlite3 -csv -header ~/.memo-mcp/kb/my-project.db \
  "SELECT id, namespace, statement, trust, valid_from, valid_to FROM facts WHERE invalidated_at IS NULL AND deleted_at IS NULL"

Column names: schema.md.

Troubleshooting runbook

Start with these three commands. They answer most questions:

memo-mcp version                       # which build, protocol, model dir
MEMO_KB=<name> memo-mcp status         # file, schema, counts, model, jobs
MEMO_KB=<name> memo-mcp verify         # integrity

The client does not show memo’s tools

CheckFix
/mcp in Claude Code, or the developer settings in Claude Desktop, shows an errorRead the server’s stderr; the reason is the last fatal: line
command is relative or uses ~Use an absolute path
fatal: invalid name …MEMO_KB must match ^[A-Za-z0-9._-]{1,64}$
fatal: … address already in useAnother session holds the same --metrics-addr. See the metrics endpoint
fatal: database schema is at version N, but this binary only supports up to version MThe file was written by a newer memo-mcp; upgrade the binary
macOS refuses to run the binaryxattr -d com.apple.quarantine /path/to/memo-mcp

Search returns nothing, or too little

CheckMeaning, and fix
Response has a reason and zero resultsDeliberate abstention: nothing passed the semantic floor. Check with memo-mcp search "<q>" --mode keyword
status shows 0 documents, or a different count than you expectWrong MEMO_KB or MEMO_HOME: the CLI and the client are not looking at the same file
trace filtered.by_scope is highThe request’s scope (namespace, version, library, minimum trust) excludes the answer
degraded is setNo model: see the next section
truncated is above 0The token budget cut results; raise max_tokens or narrow with narrow_hint
Lists stop after 3 resultsGap cutoff (cutoff.kind = gap); see cutoff_gap in Profiles

Search is degraded

CheckFix
stderr shows embedding unavailable…Model download or load failed; the error follows. Retry with memo-mcp model pull <id>
Download interrupted or corruptedmemo-mcp model redownload <id> (pass the id)
memo_kb_pending_embeddings above 0Backlog after a model switch or a degraded write: memo-mcp backfill, or reindex --model <id>
No network by designPre-seed the cache: Air-gapped installs

Slow searches

  1. memo-mcp explain "<q>", or the UI: see latency_ms_per_arm in the trace.
  2. semantic is slow: query embedding dominates. Consider potion, or accept about 0.2 s.
  3. graph is slow on its first call: the namespace graph is built once. Check memo_graph_cache_total{event="build"}. Repeated builds mean frequent writes.
  4. Reranker on? MEMO_RERANK=1 adds 0.7–4 s. Turn it off.

Slow writes

Embedding dominates: about 0.6 s per passage with the default model. For bulk loads, use ingest --embed=false, then backfill.

“database is locked”

Another process held the write lock for more than 5 seconds. This is rare, because writes are short and embedding runs outside the lock. Look for a long graph rebuild, VACUUM or a manual sqlite3 session holding a transaction. Never put the file on a network file system.

Read-only commands fail with “run memo-mcp migrate”

The file is older than the binary. Run memo-mcp migrate, or any writing command, once.

The UI is unreachable

  • Use the exact URL it printed. The Host header must match the bound address.
  • It binds loopback only. From another machine, use an SSH tunnel (ssh -L 9470:127.0.0.1:9470 host) rather than --allow-remote. The UI has no authentication and shows everything in the knowledge base.

Promote never shows a dialog

The client does not support elicitation, so the tool result contains a memo-mcp trust promote … command for a human to run. memo_mcp_elicitations_total{outcome="unsupported"} counts these.

Reporting a bug

Include memo-mcp version, the status output, the stderr lines from MEMO_LOG_LEVEL=debug, and, for ranking issues, memo-mcp explain "<query>". Do not include document text you would not share.

Release engineering

For maintainers and for operators who build their own releases.

Pipelines

WorkflowTriggerDoes
ci.ymlPush, pull requestgofmt, go vet, tests under the race detector (these include the retrieval eval gate and the tools/list golden), a pure-Go build and smoke run on Linux and macOS (Windows is currently disabled in the test matrix), golangci-lint, and a GoReleaser snapshot that must produce five archives
nightly.ymlDaily at 03:17 UTCThe eval with the hash embedder and granite-small-r2; uploads the report as an artifact
release.ymlTag v*Tests, then GoReleaser (five archives, checksums.txt), then publishes server.json to the MCP registry with GitHub OIDC (no stored secret)
docs.ymlPush to master touching docs/guide/**, pull requests (build only), manualBuilds these guides with mdBook, runs scripts/docs-check.sh, deploys to GitHub Pages

Gates a change must pass

  • Eval gate: no single query may drop a rank band, and no category mean may fall by more than 0.02 against internal/eval/testdata/baseline.json.
  • Tool-surface golden: internal/server/testdata/tools.golden.json must match. Tool names, parameters and descriptions are the 1.x contract.
  • Docs check: every MEMO_* variable, command and metric in the code must appear in this guide.

Cutting a release

make check                       # lint + race tests
make eval                        # retrieval against the baseline
make snapshot                    # local GoReleaser dry run into dist/
# write docs/eval/vX.Y.Z.md and the CHANGELOG entry
git tag -a vX.Y.Z -m "vX.Y.Z" && git push origin vX.Y.Z

The release notes footer points at docs/eval/vX.Y.Z.md. Every tag should have one.

Versioning

Semantic versioning over the 1.x contract: tool names and parameters, memo:// addresses, explain field names and the export front matter. Schema migrations are additive and are listed in Upgrading and migrations.

Appendix

Glossary

TermMeaning
AbstentionReturning zero results on purpose, with a reason, when nothing passes the semantic floor
AddressA memo://<kind>/<id> URI identifying one record; it resolves even after the record is forgotten
ArmOne retrieval method producing a ranked list: keyword, exact, semantic, fact, entity, graph
as_ofA time-travel view: what the knowledge base had recorded, and believed valid, at a given time
AuditThe table recording every write: actor, channel, operation, target
BandA relevance class (strong, moderate, weak) from the best raw cosine, using thresholds per model
BM25The term-frequency ranking function FTS5 uses for the keyword and exact arms
ChannelHow a write arrived: tool, cli, elicitation, worker. It determines trust
Chunk / passageA section of a document of about 200 tokens: the unit of indexing and retrieval
CompactionAgent-performed maintenance: writing pages, settling conflicts, deciding merges
CutoffWhere a result list ends: a score gap, the limit, the token budget, or none
DegradedA search served without a capability, usually the embedding model
ElicitationThe MCP mechanism by which a server asks the human a question through the client
EmbeddingA vector representing a text’s meaning; compared with cosine similarity
EntityA named thing (system, person, identifier) extracted from passages into the graph layer
FactOne-sentence claim with evidence, a subject list and a validity window
FTS5SQLite’s full-text search extension; memo-mcp keeps two indexes (stemmed, and identifier-preserving)
FusionCombining arm lists into one ranking: RRF (rank-based) or min-max (score-based)
GranularityWhat a search returns: chunk, document, fact or page
Knowledge base (KB)One SQLite file, selected by MEMO_KB; the isolation boundary
MCPModel Context Protocol: the JSON-RPC protocol between AI clients and tool servers
Merge candidateTwo entity names similar enough to possibly be the same thing; queued for a human
NamespaceA label grouping sources inside one knowledge base; searches span all by default
OriginDeclared provenance class: web, user-said, agent-derived
PageAgent-written markdown about an entity, stored as an inference with its cited passages
PPRPersonalised PageRank: the random walk the graph arm runs from each named entity
ProfileA named set of ranking constants (MEMO_PROFILE)
RecencyA multiplicative score factor that decays with age for kinds that age
RevisionOne version of a source’s document; changed content creates a new revision
RRFReciprocal rank fusion: sum of weight / (k + rank) over arms
SourceWhere content came from (URI or file), with library, version and trust
StaleA page whose sources changed or were forgotten since it was built
SupersedeReplace a revision or fact with a newer one, keeping the old one as history
TraceThe per-query explanation: arms run, candidates, latency, filters, cutoff, budget
TrustWho vouched for a record: agent < user < curated; only humans raise it
WALSQLite’s write-ahead log mode, which lets readers and a writer work concurrently
WhyThe per-result explanation: each arm’s rank and contribution, fused and final scores
Work itemOne compaction task with everything needed to do it in its payload

Further reading

All links point to the repository on GitHub.

Design

  • Architecture: the 1.x contract, formulas, explain fields and observability.
  • Schema: every table and column, trust transitions, time rules and migrations.
  • Roadmap: phases, decisions, verification and the glossary.
  • Research: the state of the art behind the design.

Measurements

  • Eval reports: one per release.
  • Spikes: the vector store, FTS5, embedding models, the graph, elicitation and the reranker.

Articles

One article per development phase, written for newcomers: articles/. Two of them matter most for operators:

For agents

  • SKILL.md: how an agent should use the tools.

Project