Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The search pipeline

Every search, from the MCP tool, the CLI or the UI, runs the same pipeline in internal/retrieve:

scope filter ─▶ arms (parallel ranked lists) ─▶ fusion ─▶ recency ─▶ abstention
            ─▶ cutoff ─▶ limit ─▶ token budget ─▶ results + why + trace

1. Scope filter

Namespaces, kinds, sources, library, version, tags, dates and minimum trust are applied inside each arm’s SQL, before top-k. A filtered search therefore never loses a result to a candidate that was filtered out later. The live filter excludes superseded and forgotten records. Under as_of, it is replaced by “recorded on or before T and not superseded before T”.

2. Arms

ArmIndexRuns when
keywordFTS5, Porter stemmer, BM25, over text and its section headerAlways, except mode: semantic
exactFTS5 that keeps identifiers whole (net/http, ERR_CONN_RESET)mode: exact, or auto when the query looks like an identifier
semanticCosine similarity over unit vectors in a plain SQLite tableA model is loaded and passages have vectors for it
factFacts matched by keyword or meaning; a hit votes for its evidence passageAlways
entityPassages mentioning entities the query namesRouted: the query names 2+ known entities, or 1 with relational phrasing
graphPersonalised PageRank over the mention graph, one walk per named entitySame routing; reports only passages not every seed mentions directly

Each arm returns up to fetch_depth candidates (100 by default).

3. Fusion

Weighted reciprocal rank fusion: a passage’s score is the sum over arms of weight / (k + rank), with k = 60. Default weights are semantic 0.5, keyword 0.5, exact 0.3, fact 0.4, entity 0.4 and graph 0.5. The minmax profile uses score fusion instead.

4. Recency

For kinds that age (note, conversation by default), the score is multiplied by floor + (1 − floor) · 0.5^(age_days / half_life), with floor 0.8 and half-life 90 days. Versioned documents and code do not age.

5. Abstention and relevance bands

A candidate whose only evidence is a semantic similarity below semantic_floor is dropped. If nothing survives, the response has zero results, a reason and a hint. That is a deliberate “not in the knowledge base”, not an error.

relevance is the best raw cosine for a result, interpreted against the model’s bands. For the default model, granite-small-r2, strong is 0.88, moderate 0.80 and weak 0.72, because unrelated text already scores about 0.63 under it. score orders one list and is not comparable across queries; relevance is.

6. Cutoff, limit and budget

The list ends at the first result, past min_results (3), whose score drops by more than cutoff_gap (half) from the previous one. It is then capped at limit (default 10, maximum 100), and then packed into max_tokens. The response always reports truncated, the count of ranked results that did not fit, and narrow_hint, the scope field that would shorten the list.

7. Explain and trace

With response_format: explain (MCP) or --explain (CLI), every result carries why: each arm’s rank, raw score and contribution, the fused and final score, relevance band, provenance and freshness. Every response carries a trace: arms run and why, candidates and latency per arm, filter counts, cutoff kind, budget use and degraded state. The UI renders the same structures, and a test asserts that the numbers are identical.

Degraded mode

If the embedding model is not available, because it is downloading, failed to load or has no vectors yet, search runs without the semantic arm. It sets degraded{flag, reason}, and the memo_search_degraded_total metric counts it. Writes made in that state queue their vectors. memo-mcp backfill or the next server start embeds them.

Tuning these constants: Profiles and constants. Full formulas: architecture §5.