Search Architecture
Search uses a multi-signal scoring pipeline combining lexical matching, semantic similarity, and relevance boosts to find the most useful assets for a query.
Pipeline Overview
Query
│
├─ FTS5 (lexical) ──────────────┐
│ Multi-column BM25 with │
│ field weighting │ Normalize + combine
│ ├──────────────────────── Score + Boost ── Sort ── Return
└─ Vector (semantic) ───────────┘ 0.7 FTS │
Cosine similarity, exact 0.3 Vec │
scan of stored vectors │
│
Boosts applied:
• exact name match
• type relevance
• tag/alias/hint match
• description match
• metadata quality
• usage history (M-2)
Indexed Search (primary)
When an index exists (~/.local/share/akm/index.db), local search uses two
ranking signals:
1. FTS5 (lexical)
SQLite full-text search with Porter stemming, using multi-column field weighting. The FTS5 table has separate columns for different metadata fields, each weighted differently in the BM25 scoring:
| Column | BM25 Weight | Contents |
|---|---|---|
name |
10.0 | Asset name |
description |
5.0 | Description text |
tags |
3.0 | Tags and aliases |
hints |
2.0 | Search hints and examples |
content |
1.0 | bounded body projection, TOC headings, parameters |
A name match is weighted 10x higher than a content match. This ensures that
searching "docker" ranks the docker-homelab skill above a knowledge doc that
merely mentions Docker in its table of contents.
Progressive fallback: Search runs strict AND, then prefix-AND, then one OR/prefix-OR recovery only if both strict forms return no candidates. Unicode letter/number tokens are deduplicated and capped; callers do not strip stopwords or maintain separate result collections.
matchStage: Each hit that has an FTS component reports which rung of
the ladder produced it via matchStage: "exact" (strict AND), "prefix"
(prefix AND), or "relaxed" (OR/prefix-OR recovery). The field is omitted
for hits with no FTS contribution (e.g. a pure-semantic hybrid match) and for
hits that never go through the lexical ladder (registry/browse hits). Because
the ladder stops at the first non-empty stage, all FTS-only hits in a single
response share one stage; in hybrid ranking mode a response can legitimately
mix hits that carry matchStage with hits that omit it.
2. Semantic (vector)
Cosine similarity between query embedding and stored entry embeddings.
Requires an embedding model — either the declared external
@huggingface/transformers dependency (default model
bge-small-en-v1.5) or a remote OpenAI-compatible endpoint. AKM imports the
package directly and does not copy, wrap, alias, or hash a second runtime under
src/ or dist/. Dependency installation and platform support remain the
upstream package manager's responsibility. Users who need a different runtime
or native/GPU-class throughput can configure a remote embedding endpoint.
An LRU cache (100 entries) avoids redundant embedding computation for repeated queries.
Query-time semantic failure is not conflated with configured keyword search.
FTS still serves the request, while searchMode: "fts-fallback" and one
sanitized warning name the normalized endpoint without URL userinfo, query
parameters, fragments, API keys, or raw runtime error text. Curate deduplicates
that warning across its internal fallback queries. searchMode: "keyword"
means no ready semantic attempt failed for this request.
Score Normalization
FTS5 BM25 scores are negative (lower = better match). They are normalized to a 0.3–1.0 range using min-max normalization across the result set:
- Best FTS match → 1.0
- Worst FTS match → 0.3 (floor prevents zero-score entries)
When vector search is also active, scores are combined with weighted addition:
base_score = 0.7 × normalized_bm25 + 0.3 × cosine_similarity
Design decision — why not RRF? An earlier version used Reciprocal Rank Fusion (RRF) to merge FTS and vector results. RRF uses rank positions instead of raw scores, which avoids scale mismatch but destroys score differentiation. With RRF (K=60), the best result scores 1/61 = 0.0164 and the 5th result scores 1/65 = 0.0154 — a 6% difference that makes all results look equally relevant. Normalized BM25 + weighted combination preserves the actual relevance signal: the best match might score 1.0 while the 5th scores 0.45 — a 55% difference that enables meaningful ranking. See the Scoring Design section below for the full rationale.
Scoring Design
After normalization, multiplicative boosts are applied in a single pass:
final_score = base_score × (1 + sum_of_boosts)
Boost signals (ordered by strength)
| Signal | Boost | When it fires |
|---|---|---|
| Exact name match | +2.0 | Query exactly equals the asset's base name |
| Alias exact match | +1.5 | Query exactly matches a defined alias |
| Near-exact name | +1.0 | Query is a substring of the name or vice versa |
| Name token overlap | +0.3/token (max 0.9) | Individual query tokens found in name |
| Type: skill | +0.4 | Asset is a skill |
| Type: command / workflow | +0.35 | Asset is a command or workflow |
| Type: agent | +0.3 | Asset is an agent |
| Type: knowledge / fact / instruction | +0.22 | Asset is a knowledge doc, fact, or project instruction file |
| Type: script | +0.2 | Asset is a script |
| Type: memory | -0.02 | Asset is a memory (slight demotion) |
| Tag exact match | +0.15/tag (max 0.3) | Query token exactly matches a tag |
| All-token description | +0.25 | Every query token appears in description |
| Partial description | +0.1 | Some query tokens in description |
| Search hint match | +0.12/hint (max 0.24) | Query token found in a search hint |
| Alias token match | +0.3 | Query token found in any alias |
| Curated metadata | +0.05 | Non-generated metadata (quality signal) |
| Confidence | up to +0.05 | Based on metadata source reliability |
| Usage history (M-2) | up to +0.5 (capped 1.5×) | Utility score from usage telemetry |
Design principles
Exact name match is the strongest signal. If a user types "docker-homelab",
the asset named docker-homelab must rank first with a decisive score gap.
This is the single most predictable and important ranking behavior.
Actionable assets rank above reference material. Skills, commands, and agents are things you can execute or dispatch. Knowledge docs are reference material — valuable but secondary when a user is searching for something to use. The type boost encodes this hierarchy.
Author-curated signals are stronger than auto-generated ones. Tags, aliases, and search hints are explicitly chosen by the asset author. They carry more intent than terms that happen to appear in file content.
Boosts are multiplicative on the base score. This means a high base score (strong FTS match) amplified by boosts produces a much larger gap than a low base score with the same boosts. Highly relevant assets separate clearly from marginally relevant ones.
Worked example
Query: "docker homelab" → asset: skills/docker-homelab
Base BM25 (normalized): 1.0 (best FTS match in result set)
+ Name token overlap: +0.6 (2 tokens match: "docker", "homelab")
+ Type boost (skill): +0.4
+ Tag exact match: +0.3 (tags "docker" and "homelab" both match)
+ Alias token match: +0.3 (alias "docker-compose" contains "docker")
+ Search hint match: +0.24 (hints contain "docker" and "homelab")
+ All-token description: +0.25 (all query tokens in description)
+ Curated quality: +0.05
─────────────────────────────
BoostSum: 2.14
Final score: 1.0 × (1 + 2.14) = 3.14
Compare to a sub-reference knowledge doc for the same query:
Base BM25 (normalized): 0.85 (good FTS match but not the best)
+ Name token overlap: +0.3 (1 token "homelab" in path-derived name)
+ Type boost: +0.0 (knowledge docs get no type boost)
+ Tag match: +0.15 (1 tag matches)
─────────────────────────────
BoostSum: 0.45
Final score: 0.85 × (1 + 0.45) = 1.23
The skill scores 3.14 vs the sub-reference at 1.23 — a 2.6× gap that clearly communicates which result is more useful.
Result Merging
There is one scoring pipeline because there is one data store. All sources — filesystem, git, website, npm — are materialised to local directories and indexed together. The indexer holds every searchable asset, so source results do not need cross-merging.
Source + registry results (--from all)
When registry results (static-index, skills.sh) are included:
- Source hits and registry hits stay in separate response fields:
hitsandregistryHits - Registry results are never rank-merged into
hits - Registry raw scores (which may be on a 0-100 scale) do not leak through
Substring Fallback
When no index is available, search falls back to scanning the primary bundle
and installed bundle roots, filtering by substring match. This ensures search
always works, even before akm index has been run.
Output Modes
Standard output
By default (--format json, --detail brief), search emits minimal fields:
type, name, ref, action, and estimatedTokens.
--detail normal emits type, name, description, action, score, and
estimatedTokens, plus optional warning, quality, or env-key fields.
--detail full adds whyMatched, origin, path, and timing data.
--shape summary returns metadata only (no content), under 200 tokens.
Agent-optimized output
For each materialized local hit, --shape agent includes name, canonical
ref, type, absolute path, current-policy editable, description,
action, score, and optional estimatedTokens/env keys. editHint appears
only when editable is false. The hint is supplemental: action remains the
normal show, run, or use action. Registry-only hits do not acquire local access
fields.
--format jsonl outputs one JSON object per line for streaming consumption.
Explainability
whyMatched explains which signals contributed to a hit's ranking. Examples:
"exact name match"— query matched the asset name exactly"skill type boost"— asset is a skill (actionable)"matched tags"— query token matched a tag"near-exact name match"— query is a substring of the name"usage history boost"— utility score from M-2 telemetry
Visible in --detail full output.