Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Air-gapped installs

After the binary is in place, the only network access memo-mcp ever makes is downloading an embedding model the first time it is needed. To run with no network at all, pre-seed the model cache.

Pre-seed the model cache

On a connected machine with the same memo-mcp version:

memo-mcp model pull granite-small-r2     # or the model you will set in MEMO_MODEL
ls ~/.cache/memo-mcp/models/
#   onnx-community_granite-embedding-small-english-r2-ONNX/
#   onnx-community_granite-embedding-small-english-r2-ONNX.ok

Copy both the directory and its .ok marker to the same path on the target machine, ~/.cache/memo-mcp/models/, for the user who runs memo-mcp. The marker records a completed, verified download. Without it, the files are treated as an interrupted download and fetched again.

Then check that the model loads offline:

memo-mcp model smoke granite-small-r2

For the optional reranker, also copy cross-encoder_ms-marco-MiniLM-L6-v2 and its marker.

Choosing a smaller model

Model idDownloadNotes
granite-small-r2 (default)about 195 MBBest measured quality; about 0.2 s per query on this backend
potionabout 130 MBStatic lookup table, no neural network; about 6 ms per query; weaker on paraphrase
minilmabout 87 MBThe previous default

See Embedding models for the full registry.

Running with no model at all

If no model is cached and the download fails, memo-mcp keeps working. Search runs on the keyword, exact and fact arms and reports degraded. New passages queue their vectors. When a model becomes available, memo-mcp backfill or the next server start embeds the backlog.