akm docs

Design: bundle organization & back-linking conventions

Status: accepted (conventions shipped in the bundle skeleton) Author: akm Date: 2026-07-11

Amended by the 0.9.0 surface decisions. Two mechanisms this document records as closed are superseded; the sections below that narrate the 2026-07-12 amendments are kept as a historical record, not current guidance:

  • A tooled rename (akm mv, SPEC-7) — the command still ships, but only as an Experimental surface outside the stability contract: its inbound-ref rewriting is inverted relative to the body-ref grammar. A rename is delete plus create: the destination is a new identity and learned state does not follow it. The recommended forced-rename procedure is the manual one (move, fix inbound refs, akm index, akm lint). See 0.9.0-decisions.md D3.
  • The ref-prefix filter (SPEC-4) — the <type>:<prefix>/ spelling is retired in favour of conceptId prefixes (memories/projecta/, bundle//skills/). See 0.9.0-decisions.md D4.

The no-rename default this document argues for is therefore stronger under 0.9.0, not weaker: no tool you may rely on makes a rename cheap.

Problem

The bundle skeleton already ships per-type authoring conventions (facts/conventions/assets/<type>.md) that tell an agent how to write each asset type. Nothing told an agent where to place an asset (which subdirectory) or how to cross-link it. Yet placement and linking are exactly what determine whether the same knowledge is retrievable two sessions later, or re-derived from scratch — and whether cross-project reuse happens or duplicates accumulate.

The request that drove this: define conventions that instruct agents how to organize assets under their type directories (subdirectories for projects, domains, etc.) and how to back-link them to improve retrieval during agentic tasks. This document records the analysis, a structured debate, and the resulting conventions.

What AKM's mechanics dictate (the ground truth)

Any organization scheme has to fit how AKM actually stores and retrieves. The load-bearing facts, each verified in code:

  1. Path becomes the ref. A file's subpath under its type dir is part of its ref name — memories/projectA/auth-tip.md → memory:projectA/auth-tip. Works for every type. A ref-prefix query syntax exists since SPEC-4 landed (amended 2026-07-11): akm search "memory:projectA/" enumerates exactly that subtree (typed, recursive, /-boundary exact) and a bare memory: enumerates the whole type. At review time the idiom was false — sanitizeFtsQuery strips : and / and entry_type is not an FTS column, so the query degenerated to token noise (empirically tested; see the Review round). akm search "projectA" --type memory remains the token-match alternative. The path is therefore the one facet you cannot express twice and cannot change without breaking the ref (and every inbound reference to it).
  2. Retrieval is search, not browse — and folders are highly visible to the ranker. There is no folder walk at query time — agents akm search/ akm curate, then akm show <ref>. A subdirectory's real payoff: its tokens join the FTS name column at the highest bm25 weight (10.0), always merge into tags (since SPEC-2 landed; see fact 6), and match the cwd project-context boost (projectContextRankingContributor, +0.2/token, cap 0.5). No "indexing-confidence bump" for subdirectories exists in code.
  3. The FTS surface is name, description, tags, hints, content (entries_fts, src/storage/repositories/index-schema.ts). There is no project field — a bare project: frontmatter value is invisible to search. Off-axis facets must be tags to be retrievable.
  4. xrefs: fold into the FTS hints field for all types (src/indexer/search/search-fields.ts:58). Back-links are a retrieval signal, not decoration — which means both too few and too many degrade ranking.
  5. Project relevance is recovered at query time. Ranking blends a per-project usage signal — scopedUtility * 0.7 + globalUtility * 0.3 (SCOPED_UTILITY_BLEND_SCOPED, src/indexer/search/ranking-contributors.ts), keyed by a cwd-anchor (a hash of the querying project root), not the asset's path. It is NOT rename-proof for the asset: entry_key includes the full name, so renaming a file mints a new entry row and orphans its accumulated global + scoped utility history. As of 0.9.0 that is the defined semantics — a rename is delete plus create, and no command preserves learned state across one (see 0.9.0-decisions.md D3). A reusable asset does not need the project baked into its path to rank well inside a project — and a manual rename costs learned ranking.
  6. Directory (scope/domain) tokens always merge into tags — since SPEC-2 landed (extractDirTagsFromName, src/indexer/passes/metadata.ts) they are derived from the canonical ref subpath, so explicit tags no longer suppress the scope token, both indexing walks agree, and the exact-tag ranking boost fires for the path token without restating it. Filename tokens are still auto-derived only when tags is empty (they already live in the FTS name column and in aliases). The old explicit-tags footgun is gone; existing installs pick the merged tags up on the next reindex.
  7. The entity/relation graph boost extracts from body content, not the path (graph-extraction.ts). Canonical entity naming must live in the prose.
  8. category: convention|meta facts are surfaced to every non-wiki authoring flow (resolveStashStandards); facts/conventions/assets/<type>.md is surfaced type-scoped (resolveTypeConventions). This is the only channel that reaches a browse-blind agent mid-task — so conventions must ship as facts.
  9. Non-wiki xref breakage is caught by akm lint, not at write time. The deterministic missing-ref check covers body refs, the refs: frontmatter array, and (since SPEC-1 landed) the xrefs:/supersededBy:/ contradictedBy: frontmatter channels. A rename dangles inbound links silently — nothing catches them until the next akm lint run flags them, so run akm lint as the last step of any rename. (akm mv automates this, but its rewriting is inverted relative to the body-ref grammar, so it ships Experimental in 0.9.0 and does not substitute for the lint run — see 0.9.0-decisions.md D3.) Wikis are excluded from akm lint's directory sweep; the equivalent orphan/broken-xref/broken-source/stale-index/uncited-raw structural checks are implemented in the LLM Wiki adapter (core/adapter/adapters/llm-wiki-adapter.ts, ported from the pre-0.9.0 akm wiki lint). Amendment: the akm wiki command family, including akm wiki lint, was removed in the 0.9.0 bundle-adapter cutover (see wikis.md) — there is currently no CLI surface that invokes these checks.

Research inputs

Three parallel research briefs (Karpathy's LLM wiki; agent memory / RAG / GraphRAG; PKM taxonomies — Zettelkasten, PARA, Johnny.Decimal) converged on a few transferable principles:

The debate

Four positions were argued and adversarially critiqued:

Position Thesis Verdict
Domain-first taxonomy Path = stable subject domain; lowest rename rate, highest lexical signal. reject-keep-ideas
Scope-first namespacing Path = project/client/team; scope is the dominant retrieval filter. reject-keep-ideas
Flat + faceted + dense backlinks Minimal folders; organizing signal lives in frontmatter + a dense link graph the index can read. adopt-with-changes
Hybrid "narrowest stable scope" ladder One folder = the narrowest stable boundary (client > project > domain > root), rest in facets. adopt-with-changes

The critiques were decisive:

Decision

Resolve the project-vs-domain tension by asset TYPE, not per-asset judgment. This is the one rule that is both deterministic mid-task and mechanically justified:

Supporting rules, all shipped as convention facts:

Rejected over-builds: mandatory dense/bidirectional xrefs, per-namespace hub wikis, the per-asset scope ladder, and any hand-maintained catalog.

What shipped

Convention facts in the bundle skeleton (src/assets/stash-skeleton/facts/conventions/):

Because these are category: convention facts, resolveStashStandards injects them into the authoring context automatically — no code change was required to enforce the surfacing. They are soft guidance; a bundle owner edits them to match how their bundle is queried.

The code changes the finalized conventions imply are specced separately in stash-conventions-code-spec.md (8 specs, prioritized and sequenced; nothing implemented in this change).

Review round

A two-reviewer adversarial review (retrieval mechanics / IR + KM methodology / agent usability), with cross-rebuttals, tuned the shipped conventions. Decisive corrections, each code-verified or empirically tested:

Open questions / future work

Items with an implementation path were specced in stash-conventions-code-spec.md, and all eight specs have since landed (amended 2026-07-12): the lint frontmatter channels (SPEC-1), the tag merge (SPEC-2), --xref (SPEC-3), the ref-prefix filter (SPEC-4), --supersedes demotion (SPEC-5), category capture (SPEC-6 — shipped capture-only; the rank-time demotion was dropped after measurement showed no crowding), akm mv (SPEC-7), and bounded native body indexing (SPEC-8) — see the closed bullets below. What remains genuinely open: vocabulary governance, scheduled consolidation (argued down in the spec — the improve pipeline already injects the amended conventions), and the typed provenance channel.