akm docs

Search Architecture

Search uses a multi-signal scoring pipeline combining lexical matching, semantic similarity, and relevance boosts to find the most useful assets for a query.

Pipeline Overview

Query
  │
  ├─ FTS5 (lexical) ──────────────┐
  │   Multi-column BM25 with      │
  │   field weighting              │   Normalize + combine
  │                                ├──────────────────────── Score + Boost ── Sort ── Return
  └─ Vector (semantic) ───────────┘        0.7 FTS          │
      Cosine similarity, exact             0.3 Vec          │
      scan of stored vectors                                │
                                                            │
                                              Boosts applied:
                                              • exact name match
                                              • type relevance
                                              • tag/alias/hint match
                                              • description match
                                              • metadata quality
                                              • usage history (M-2)

Indexed Search (primary)

When an index exists (~/.local/share/akm/index.db), local search uses two ranking signals:

1. FTS5 (lexical)

SQLite full-text search with Porter stemming, using multi-column field weighting. The FTS5 table has separate columns for different metadata fields, each weighted differently in the BM25 scoring:

Column BM25 Weight Contents
name 10.0 Asset name
description 5.0 Description text
tags 3.0 Tags and aliases
hints 2.0 Search hints and examples
content 1.0 bounded body projection, TOC headings, parameters

A name match is weighted 10x higher than a content match. This ensures that searching "docker" ranks the docker-homelab skill above a knowledge doc that merely mentions Docker in its table of contents.

Progressive fallback: Search runs strict AND, then prefix-AND, then one OR/prefix-OR recovery only if both strict forms return no candidates. Unicode letter/number tokens are deduplicated and capped; callers do not strip stopwords or maintain separate result collections.

matchStage: Each hit that has an FTS component reports which rung of the ladder produced it via matchStage: "exact" (strict AND), "prefix" (prefix AND), or "relaxed" (OR/prefix-OR recovery). The field is omitted for hits with no FTS contribution (e.g. a pure-semantic hybrid match) and for hits that never go through the lexical ladder (registry/browse hits). Because the ladder stops at the first non-empty stage, all FTS-only hits in a single response share one stage; in hybrid ranking mode a response can legitimately mix hits that carry matchStage with hits that omit it.

2. Semantic (vector)

Cosine similarity between query embedding and stored entry embeddings. Requires an embedding model — either the declared external @huggingface/transformers dependency (default model bge-small-en-v1.5) or a remote OpenAI-compatible endpoint. AKM imports the package directly and does not copy, wrap, alias, or hash a second runtime under src/ or dist/. Dependency installation and platform support remain the upstream package manager's responsibility. Users who need a different runtime or native/GPU-class throughput can configure a remote embedding endpoint.

An LRU cache (100 entries) avoids redundant embedding computation for repeated queries.

Query-time semantic failure is not conflated with configured keyword search. FTS still serves the request, while searchMode: "fts-fallback" and one sanitized warning name the normalized endpoint without URL userinfo, query parameters, fragments, API keys, or raw runtime error text. Curate deduplicates that warning across its internal fallback queries. searchMode: "keyword" means no ready semantic attempt failed for this request.

Score Normalization

FTS5 BM25 scores are negative (lower = better match). They are normalized to a 0.3–1.0 range using min-max normalization across the result set:

When vector search is also active, scores are combined with weighted addition:

base_score = 0.7 × normalized_bm25 + 0.3 × cosine_similarity

Design decision — why not RRF? An earlier version used Reciprocal Rank Fusion (RRF) to merge FTS and vector results. RRF uses rank positions instead of raw scores, which avoids scale mismatch but destroys score differentiation. With RRF (K=60), the best result scores 1/61 = 0.0164 and the 5th result scores 1/65 = 0.0154 — a 6% difference that makes all results look equally relevant. Normalized BM25 + weighted combination preserves the actual relevance signal: the best match might score 1.0 while the 5th scores 0.45 — a 55% difference that enables meaningful ranking. See the Scoring Design section below for the full rationale.

Scoring Design

After normalization, multiplicative boosts are applied in a single pass:

final_score = base_score × (1 + sum_of_boosts)

Boost signals (ordered by strength)

Signal Boost When it fires
Exact name match +2.0 Query exactly equals the asset's base name
Alias exact match +1.5 Query exactly matches a defined alias
Near-exact name +1.0 Query is a substring of the name or vice versa
Name token overlap +0.3/token (max 0.9) Individual query tokens found in name
Type: skill +0.4 Asset is a skill
Type: command / workflow +0.35 Asset is a command or workflow
Type: agent +0.3 Asset is an agent
Type: knowledge / fact / instruction +0.22 Asset is a knowledge doc, fact, or project instruction file
Type: script +0.2 Asset is a script
Type: memory -0.02 Asset is a memory (slight demotion)
Tag exact match +0.15/tag (max 0.3) Query token exactly matches a tag
All-token description +0.25 Every query token appears in description
Partial description +0.1 Some query tokens in description
Search hint match +0.12/hint (max 0.24) Query token found in a search hint
Alias token match +0.3 Query token found in any alias
Curated metadata +0.05 Non-generated metadata (quality signal)
Confidence up to +0.05 Based on metadata source reliability
Usage history (M-2) up to +0.5 (capped 1.5×) Utility score from usage telemetry

Design principles

Exact name match is the strongest signal. If a user types "docker-homelab", the asset named docker-homelab must rank first with a decisive score gap. This is the single most predictable and important ranking behavior.

Actionable assets rank above reference material. Skills, commands, and agents are things you can execute or dispatch. Knowledge docs are reference material — valuable but secondary when a user is searching for something to use. The type boost encodes this hierarchy.

Author-curated signals are stronger than auto-generated ones. Tags, aliases, and search hints are explicitly chosen by the asset author. They carry more intent than terms that happen to appear in file content.

Boosts are multiplicative on the base score. This means a high base score (strong FTS match) amplified by boosts produces a much larger gap than a low base score with the same boosts. Highly relevant assets separate clearly from marginally relevant ones.

Worked example

Query: "docker homelab" → asset: skills/docker-homelab

Base BM25 (normalized):  1.0    (best FTS match in result set)
+ Name token overlap:    +0.6   (2 tokens match: "docker", "homelab")
+ Type boost (skill):    +0.4
+ Tag exact match:       +0.3   (tags "docker" and "homelab" both match)
+ Alias token match:     +0.3   (alias "docker-compose" contains "docker")
+ Search hint match:     +0.24  (hints contain "docker" and "homelab")
+ All-token description: +0.25  (all query tokens in description)
+ Curated quality:       +0.05
─────────────────────────────
BoostSum:                2.14
Final score:             1.0 × (1 + 2.14) = 3.14

Compare to a sub-reference knowledge doc for the same query:

Base BM25 (normalized):  0.85   (good FTS match but not the best)
+ Name token overlap:    +0.3   (1 token "homelab" in path-derived name)
+ Type boost:            +0.0   (knowledge docs get no type boost)
+ Tag match:             +0.15  (1 tag matches)
─────────────────────────────
BoostSum:                0.45
Final score:             0.85 × (1 + 0.45) = 1.23

The skill scores 3.14 vs the sub-reference at 1.23 — a 2.6× gap that clearly communicates which result is more useful.

Result Merging

There is one scoring pipeline because there is one data store. All sources — filesystem, git, website, npm — are materialised to local directories and indexed together. The indexer holds every searchable asset, so source results do not need cross-merging.

Source + registry results (--from all)

When registry results (static-index, skills.sh) are included:

Substring Fallback

When no index is available, search falls back to scanning the primary bundle and installed bundle roots, filtering by substring match. This ensures search always works, even before akm index has been run.

Output Modes

Standard output

By default (--format json, --detail brief), search emits minimal fields: type, name, ref, action, and estimatedTokens.

--detail normal emits type, name, description, action, score, and estimatedTokens, plus optional warning, quality, or env-key fields. --detail full adds whyMatched, origin, path, and timing data. --shape summary returns metadata only (no content), under 200 tokens.

Agent-optimized output

For each materialized local hit, --shape agent includes name, canonical ref, type, absolute path, current-policy editable, description, action, score, and optional estimatedTokens/env keys. editHint appears only when editable is false. The hint is supplemental: action remains the normal show, run, or use action. Registry-only hits do not acquire local access fields.

--format jsonl outputs one JSON object per line for streaming consumption.

Explainability

whyMatched explains which signals contributed to a hit's ranking. Examples:

Visible in --detail full output.