Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Ingest and revisions

The write path

  1. Normalise line endings and whitespace, and hash the content (SHA-256).
  2. Deduplicate. If the source’s current revision has the same hash, nothing happens and the result says dedup: true. If the content differs, a new revision is created. The previous revision is marked superseded and stays readable with as_of.
  3. Chunk. The text is split by headings, then paragraphs, then sentences, then a hard split. The target is about 200 estimated tokens and the cap 400, clamped to the model’s limit, with about 40 tokens of overlap. Fenced code blocks are never split. Each passage gets a context header title > section path that the keyword index also sees.
  4. Index. Both FTS5 indexes are filled by triggers inside the transaction. Entity mentions are extracted and linked.
  5. Embed outside the transaction, in batches of 16. If no model is available, the passages are queued as a job, and backfill or the next server start embeds them. A write never reports success it did not get.
  6. Audit: one row per write.

A source is identified by its URI, or by its file path for CLI ingests. Re-ingesting the same URI with changed content, or with a new version, creates a new revision of that source.

From the CLI

memo-mcp ingest ./docs --ns handbook --kind doc
memo-mcp ingest page.md --uri https://pkg.go.dev/google.golang.org/grpc \
  --library grpc/grpc-go --version v1.64.0 --context "gRPC-Go API docs for deadlines"
cat notes.md | memo-mcp ingest - --kind note --title "Standup 2026-10-03"
memo-mcp ingest ./big-folder --embed=false && memo-mcp backfill   # fast load, embed later

CLI writes are trust user, or curated with --trust curated. Origin defaults to web when --uri is set, and to user-said otherwise.

Kinds

KindAges under recency?Typical content
docNoDocumentation, specs, handbooks; usually versioned
codeNoCode explanations and snippets
noteYesShort notes, decisions, observations
conversationYesConversation summaries

Costs

StepCost
Parse, chunk, indexAbout 60 documents per second without embedding
Embedding with granite-small-r2About 0.6 s per passage on the pure-Go backend
Embedding with potionNear instant
DiskAbout 4.4 KB per passage before vectors, plus about 1.5 KB per passage per 384-dimension model

For bulk loads, ingest with --embed=false, then run memo-mcp backfill, or let the server embed in the background. See Capacity and performance.

Observing it

memo_store_ingests_total{outcome=new|revision|dedup|error}, memo_store_ingest_duration_seconds, memo_store_chunks_written_total, memo_store_embed_batches_total{outcome}, and memo_kb_pending_embeddings{model} for the backlog.