akm docs

Configuration

AKM reads one user configuration file: $XDG_CONFIG_HOME/akm/config.json (normally ~/.config/akm/config.json on Linux and macOS, or %APPDATA%\akm\config.json on Windows). Set AKM_CONFIG_DIR to override the directory. Project .akm/config.json files are not merged. A config file may optionally extend one other config via extends (see "Sharing configuration across installs" below) — this is a single, explicit, user-opted-in key, not automatic project-config discovery.

Version 0.9

configVersion is "0.9.0", the only value akm has ever shipped. It is read, never gated on: a file without the field loads silently, and a file declaring any other value is named once on stderr (config.json declares configVersion "X"; this release reads it as 0.9.0.) and read as the current shape anyway — nothing is rewritten on disk. The next config write (akm config set, etc.) and akm migrate apply's config step both persist "0.9.0", which silences the note. When a newer akm wrote the shared config, akm health's binary-config-skew advisory is what says so. Pre-0.9 config and database layouts are not runtime inputs. Historical task sources are handled by the standalone akm-migrate executable, also invoked by akm migrate / akm upgrade; ordinary runtime code reads only the current shape.

{
  "configVersion": "0.9.0",
  "$schema": "https://itlackey.github.io/akm/schemas/akm-config.json",
  "engines": {
    "fast": {
      "kind": "llm",
      "endpoint": "http://localhost:11434/v1/chat/completions",
      "model": "qwen3",
      "apiKey": "${LOCAL_LLM_API_KEY}"
    },
    "reviewer": {
      "kind": "agent",
      "platform": "opencode",
      "model": "anthropic/claude-sonnet-4-6"
    }
  },
  "defaults": {
    "engine": "reviewer",
    "llmEngine": "fast",
    "improveStrategy": "default"
  },
  "workflow": {
    "maxConcurrency": 8,
    "judgeEngine": "reviewer"
  },
  "execution": {
    "allowedTools": ["read_file", "search"]
  },
  "improve": {
    "strategies": {
      "nightly": {
        "engine": "fast",
        "processes": {
          "reflect": {},
          "memoryInference": { "model": "qwen3-small", "llm": { "temperature": 0.1 } }
        }
      }
    }
  }
}

Scheduler activation

scheduler.enabled is this host's list of scheduled refs, such as stash//tasks/nightly. A ref that is not listed is disabled. On disk each entry is still written as the {kind, ref, sourceId} object 0.9.16 reads, so that release keeps working against a config this one wrote; in memory it is the ref. A config without the list (written before 0.9.17) means "keep what is installed": the first akm task sync fills it from the akm-written native scheduler rows. The 0.9.17-alpha {kind, ref, sourceId} entries are read as their ref. Authored task/workflow files may describe schedules but cannot put themselves on the list.

This key is deliberately local: if a config uses extends, any scheduler section in the base is ignored with a warning. Only the top-level local config can activate schedules. Prefer akm task enable <ref> and akm task disable <ref> over editing the JSON by hand; both update the allow-list and sync the affected bundle. An unscoped akm task sync reconciles enabled refs across all enabled configured bundles.

Engines

engines is the only public execution map. An engine name is lowercase kebab-case, at most 63 characters, and cannot start with akm-.

Kind Required fields Use
llm endpoint, model OpenAI-compatible chat completions
agent platform A registered dispatch-capable harness

LLM endpoints must be complete http:// or https:// chat-completions URLs ending in /chat/completions, without userinfo, query, or fragment. API keys are symbolic only: $VAR or ${VAR}. AKM resolves them only at dispatch.

An LLM engine may set enableThinking: false to turn thinking off and reasoningEffort to a value such as "none", "low", or "high". AKM sends both wire forms — chat_template_kwargs.enable_thinking and top-level enable_thinking — whenever enableThinking resolves, from engine config or a calling process (improve's consolidate/reflect always request enableThinking: false for a machine-readable payload, and the reflect and distill quality-gate judges do too unless the judge's engine sets enableThinking: true); reasoningEffort is always sent as top-level reasoning_effort when set. Backend support: llama.cpp direct honors both forms (reasoning_effort from build ≥ b10644); vLLM honors chat_template_kwargs; Bifrost drops chat_template_kwargs and passes reasoning_effort through, so also set reasoningEffort: "none" behind it. A strict hosted API that rejects the two fields (OpenAI answers 400 Unknown parameter: 'chat_template_kwargs') gets one retry without both, and AKM stops sending them to that endpoint and model for the rest of the process, as it does for response_format below. Both fields are AKM-owned, not settable via extraParams. A response with reasoning tokens despite enableThinking: false triggers a runtime warning and the akm health thinking-control advisory.

When a call asks for JSON that matches a schema, an LLM engine sends the schema as response_format (json_schema, strict) unless the engine sets supportsJsonSchema: false. If the endpoint rejects it with a 4xx other than 429, AKM retries once without response_format and stops sending it to that endpoint and model for the rest of the process. An agent engine receives the schema as one instruction at the end of its prompt (Respond with ONLY a JSON value matching this JSON Schema (no prose, no code fences): followed by the schema), plus the harness's own schema channel where it has one (codex --output-schema).

An agent engine may set bin, args, workspace, model, and timeoutMs; it takes no inference of its own (see Inference on an agent engine). Only platform: "opencode-sdk" may set llmEngine; it names the LLM engine used as that SDK engine's fallback connection. With no llmEngine, an SDK engine has no fallback connection and opencode resolves provider, model and auth from its own configuration. defaults.llmEngine is not a substitute.

On an engine without timeoutMs, model work (an improve process, a quality or triage judge, or an index pass) stops after 600 seconds, whatever the engine's kind; other work on an agent engine runs until it finishes. An improve stage's reply that the stage cannot read gets one corrective retry. Reflect holds its reply to its own contract the same way on every engine kind: a reply that is not the JSON object of reflect's schema gets one repair turn, and one still invalid fails with parse_error.

Executable assets may request tools, but the request is not authority. Configure the host-local execution.allowedTools list to define the ceiling; "*" is an explicit allow-all. The default is an empty list. Asset frontmatter cannot set workspace, environment, or opaque runtime values; those belong to local engine configuration or workflow environment bindings.

platform: "opencode-sdk" needs the opencode binary on PATH (or a bin pointing at it). akm bundles @opencode-ai/sdk, but that package is an HTTP client with no dependencies — it spawns opencode serve and talks to it — so the npm dependency alone does not make the platform usable. Install the binary with npm i -g opencode-ai or opencode's own installer.

Inference on an agent engine

An agent engine sets no inference of its own. For opencode, set it on the model in your own opencode config, which akm leaves as it is. akm carries inference (temperature, reasoningEffort, enableThinking, maxTokens, contextLength) into opencode only where it writes opencode's config itself:

Inference from an asset, a workflow's llm: or a models.json alias reaches an LLM engine. On an agent engine it is reported as an untranslated-field notice and dispatch continues; model work's agent carries temperature, reasoningEffort and enableThinking without one.

Reasoning effort has one word in a request, reasoningEffort. effort, as a models.json alias or an asset's effort: frontmatter spells it, is read as reasoningEffort wherever layers are merged, so an LLM engine sends it as reasoning_effort.

Engines for unattended model work

Unattended model work is the work akm hands a model with no one watching:

It runs on any engine kind under one tool policy. The model may read, edit files only inside a scratch working directory that akm creates for the dispatch and removes after it, and run akm search and akm show. The stash stays read-only to it: what model work changes reaches the stash only as a proposal, through the review queue.

Each engine enforces as much of the policy as it can, and grants nothing it cannot enforce:

Engine What the model gets
LLM No tools.
claude Read and Edit inside the working directory, and Bash for akm search and akm show only. akm runs it with --restricted, so your user, project and local settings cannot widen that.
opencode, opencode-sdk Read, grep and glob in the working directory and the stash, edit in the working directory only, and akm_search and akm_show, through an injected akm-model-work agent. No bash, because opencode cannot stop a redirect such as akm show x > file from writing elsewhere.
codex, copilot, pi, gemini, aider, amazonq, openhands Cannot run model work: akm refuses the request before it starts.

The one rule. Every key model work reads its engine from must name an LLM engine or a claude, opencode or opencode-sdk agent engine. Those keys are:

A config that breaks the rule fails to load, with an error that names the key, the engine and its platform. Other engine keys, defaults.engine and workflow.judgeEngine, may name any configured engine.

Further details:

Model-map files

AKM ships an immutable models.json package asset with three intent aliases: fast, balanced, and reasoning. The installed starter has separate columns for Claude Code, OpenCode, and OpenCode SDK only. A provider-specific identifier is not pretended to work on unrelated engines. Add mappings for Gemini, Codex, named direct LLM engines, or other harnesses in your user file. Config-root and per-engine modelAliases are rejected; this file is the only alias definition surface.

An optional operator-owned file lives beside config.json at $XDG_CONFIG_HOME/akm/models.json (or <AKM_CONFIG_DIR>/models.json). It uses the same version-1 schema as the installed file:

{
  "version": 1,
  "aliases": {
    "fast": {
      "gemini": "gemini-2.5-flash"
    },
    "reasoning": {
      "claude": {
        "inference": {
          "effort": "medium"
        }
      },
      "local-reasoner": {
        "model": "qwen3:30b",
        "inference": {
          "effort": "high"
        }
      }
    }
  }
}

Each engine mapping is either a non-empty exact model string or a structured profile with the documented fields model, inference, and engine. A user profile may omit model when the installed layer already supplies it, as the partial Claude override above does. After overlay, every alias/engine entry must have a usable model. Unknown profile fields are rejected; JSON-safe fields inside inference are preserved for engine adapters to lower optimistically. An inference.effort is read as reasoningEffort (see Inference on an agent engine).

A profile's engine field (0.9.15, #946) borrows a column's model (and, for an llm-kind engine, its inference defaults) from a configured engines.<name> connection instead of hand-typing a literal model a second time:

{
  "version": 1,
  "aliases": {
    "fast": {
      "opencode": { "engine": "local-fast" }
    }
  }
}

With engines.local-fast configured (agent-kind or llm-kind), this column resolves to that engine's own model string. model and engine are mutually exclusive on the same profile — engine is an indirection for the model value, never an engine-selection override; which engine akm agent dispatches to is still decided entirely by --engine/defaults.engine (see Engine selection). The referenced engine's model must itself be literal, not another alias, and akm copies it verbatim: it does not translate between an engine's connection and an agent platform's own provider registry, so the value must already be meaningful for the column's platform (e.g. a kind: "agent", platform: "opencode" engine's model should already be a string opencode itself understands, such as krang/qwen3.5-9b). Run akm models list to see, for every alias/column, the resolved model and whether it came from the installed defaults, the user overlay, and a literal value or an engine reference.

The user file overlays the installed file by alias, engine, and nested object field. Objects merge recursively. Arrays, scalars, and explicit null replace the lower value; omitted fields preserve it. A layer setting a literal model clears any engine inherited from a farther layer, and vice versa — the nearer layer's choice of literal-vs-engine always wins outright rather than merging. Alias and engine keys are case-normalized, and case-colliding definitions are rejected. Unknown model inputs still pass through byte-for-byte as exact identifiers. Once a name is a known merged alias, selecting an engine with no mapping is an actionable configuration error rather than silently sending the alias as a model ID.

The common execution cascade reads these files for current direct command and non-interactive agent calls, task source v4 runs, and improve/proposal/index model work routed through that resolver. A structured alias expands as defaults at the layer that selected it; explicit sibling fields and nearer layers still win. The resulting request carries the exact model ID and merged inference object. Engine lowerers consume that exact selection and never run alias resolution again. New workflow starts persist the exact request and symbolic runner selection in the durable plan v4 family's executable irVersion: 5; resume consumes that frozen material without resolving aliases again.

Copy the complete installed starter into the user configuration directory when you want to customize all fields:

akm models copy-defaults
akm models copy-defaults --overwrite  # explicit replacement confirmation

The command validates the installed asset, creates the config directory, and writes a fully synced sibling before publication. Without --overwrite, a hard-link/no-replace operation makes publication atomic: a racing creator wins without losing its bytes. A filesystem that cannot provide that operation fails safely instead of falling back to a clobbering rename.

With --overwrite, the portable guarantee is an atomic pathname replacement that never follows the target when it is a symlink. AKM verifies the observed regular-file identity again immediately before rename, but the portable filesystem APIs do not provide a conditional compare-and-swap rename. Another process can still change the directory entry after that check; AKM replaces the entry at the pathname without dereferencing it. Consequently, overwritten: true means overwrite was requested for an entry AKM observed, not that an inode identity was transactionally locked. Symlinks and other non-regular targets observed at either check are refused.

AKM does not auto-create or sync this file, and authoritative defaults never live in the cache. npm/Node and normal Bun installs read the packaged dist/assets/models.json lazily, so akm health can report a missing or malformed package asset as a model-map-files failure. A standalone binary has the same authoritative bytes embedded at compile time and therefore has no external model-map asset that can later disappear; its health check validates the embedded copy, and release tests pin copied bytes to src/assets/models.json. The health check passes when the optional user file is absent and warns with its path and JSON location when the user file is unreadable or invalid.

defaults.engine names an LLM or agent engine. defaults.llmEngine names the default engine for unattended model work, so it follows the one rule. There is no first-engine fallback: an unset defaults.engine never resolves to some arbitrary entry in engines. It resolves instead to a synthesized, config-free opencode-sdk engine when the opencode binary is on PATH — announced once per run, and preempted by any opencode-sdk engine you configure yourself. Naming an engine that is not configured is always an error and is never rescued by that fallback.

defaults.llmEngine is not an opencode-sdk engine's fallback connection. An SDK engine gets an LLM fallback only from its own llmEngine, so the synthesized engine, which sets none, runs on opencode's own provider, model and auth.

Index passes select engines through index.defaults.engine or index.<pass>.engine, which follow the one rule. Per-pass model, timeoutMs, and llm fields are invocation overrides; enabled: false disables that pass. Connection fields such as endpoint, provider, apiKey, and apiKeyFile belong only on named engines.

workflow.maxConcurrency is the native workflow engine ceiling. An explicit value is clamped to 1..64. When absent, AKM derives the cap once from the CPU count (min(16, max(1, cores - 2))) and freezes it into the run plan, so resume does not change policy on a different host or after config edits.

workflow.defaultMapConcurrency is the width a map step freezes when it declares no concurrency: of its own. Unset means 4 — map steps are parallel by default as of 0.9.1. An explicit value is clamped to 1..64; set it to 1 to restore the pre-0.9.1 serial-by-default fan-out for every workflow on this machine. It is only a default: an authored map.concurrency always wins, and it never raises a step past workflow.maxConcurrency, the selected engine's concurrency, or the host CPU cap. An LLM engine that declares no engines.<name>.concurrency gets 1 on a loopback endpoint (a local model server holds one loaded model) and 4 on a remote one.

workflow.judgeEngine names the LLM or agent engine used to verify every non-empty workflow ### gate rubric. It is required when a workflow declares completion criteria and is frozen into each new run, so later config edits do not change an in-flight run's verifier. Missing, failed, or malformed verifier results reject the gate; criteria are never silently bypassed.

Strategies

Improve presets live under improve.strategies; invoke one with akm improve --strategy <name>. The selection order is --strategy, defaults.improveStrategy, then built-in default. A strategy and each process can select engine, model, timeoutMs, and LLM request overrides:

{
  "improve": {
    "strategies": {
      "nightly": {
        "engine": "fast",
        "processes": {
          "reflect": { "llm": { "temperature": 0.2 } },
          "memoryInference": { "model": "qwen3-small" }
        }
      }
    }
  }
}

An improve process's engine follows the one rule; an explicit invalid or incompatible engine never falls back to another engine. Built-in strategies are complete presets. User-defined strategies inherit omitted fields from the built-in default strategy before applying their own overrides.

processes.triage.judgment explicitly controls the optional judgment tier. Use true to enable it, false to disable it, or an object with enabled, engine, model, timeoutMs, and/or llm overrides. Existing object values such as {} and { "engine": "reviewer" } remain enabled by default. Unknown object keys are rejected so misspellings cannot silently change execution; the retired mode and profile keys continue to report their engine migration guidance. When enabled, engine selection is judgment → triage → strategy → defaults.llmEngine, and resolution fails closed if none is available.

{
  "improve": {
    "strategies": {
      "nightly": {
        "processes": {
          "triage": {
            "enabled": true,
            "judgment": { "enabled": true, "engine": "reviewer" }
          }
        }
      }
    }
  }
}

processes.reflect.qualityGate and processes.distill.qualityGate control each process's LLM-as-judge quality gate. Each is on unless it sets enabled: false, and each follows only its own switch. A reflect revision the judge passes is staged for the triage drain to accept; a distill lesson it passes is deferred for a person (reason distill-review), which the drain and its judgment tier leave alone. With the distill gate off nothing is judged, and the drain decides. The judge is the process's own engine when that is an LLM engine, or the defaults.llmEngine engine when an agent generates. engine, model, timeoutMs and llm give the gate a judge of its own, resolved over the process's settings the way triage.judgment resolves over triage's. The judge may be any engine that follows the one rule. A gate whose settings resolve to no engine fails before anything is generated; it never falls back to another engine. The judge runs at temperature 0 with thinking off unless its engine sets enableThinking: true. Thinking is slow: on a 27B llama.cpp server, a thinking judgment took a median of 30–67 s and up to about 3 minutes, against about 5 s without. Behind a gateway that drops chat_template_kwargs (Bifrost), point a thinking judge's engine at the server directly.

{
  "engines": {
    "judge": {
      "kind": "llm",
      "endpoint": "http://127.0.0.1:8080/v1/chat/completions",
      "model": "qwen3-27b",
      "enableThinking": true
    }
  },
  "improve": {
    "strategies": {
      "nightly": {
        "processes": {
          "reflect": { "qualityGate": { "engine": "judge" } }
        }
      }
    }
  }
}

processes.reflect.defectFilter sets the wording of the checks reflect runs before the judge. With the quality gate on or off, reflect refuses a revision that adds placeholder text, talks about its own edit, or copies frontmatter into its body: no proposal, and no judge call. Each rule counts only what the revision adds to its source, so wording the asset already had and kept is not held against it. Each of the three lists is optional. A list you set replaces that rule's default list, and [] turns the rule off.

List Rule Default
placeholders placeholder_added please confirm, please verify, to be confirmed, to be determined, to be verified
metaCommentary meta_commentary_added feedback signal, feedback signals, feedback indicate, feedback indicates, feedback suggest, feedback suggests, feedback ask, feedback asks, feedback says, feedback report, feedback reports, feedback request, feedback requests, this revision, the source asset, the source note, the source memory, the original asset, the original note, the original memory, the original version of this, quality gate rejected, proposal rejected
frontmatterKeys frontmatter_copied_into_body sources, updated, inferenceProcessed, captureMode, beliefState, xrefs, contradictedBy, outcomeData, orderedActions, generated, verified, description, when_to_use, tags, searchHints, quality, salience, salienceInputs, lint_skip, type

placeholders and metaCommentary entries are plain phrases, not patterns: whole words, in any case, with any run of whitespace between words. frontmatterKeys entries are exact key names: a line outside a code fence that starts with key: counts. The rule also refuses a sources, xrefs or contradictedBy value copied into the body; frontmatterKeys: [] turns that off too. Every entry must be a non-empty string, or the config does not load.

{
  "improve": {
    "strategies": {
      "nightly": {
        "processes": {
          "reflect": {
            "defectFilter": {
              "placeholders": ["please confirm", "to be confirmed", "[draft]"],
              "frontmatterKeys": []
            }
          }
        }
      }
    }
  }
}

No shipped strategy turns improve-stage session extraction on. proactiveMaintenance is on only in the proactive-maintenance preset; run akm improve --strategy proactive-maintenance to use that opt-in preset. Because strategies inherit from default, a preset that omits either process also inherits the off value. User strategy overrides are applied last, so an explicit enabled: true still opts the selected strategy in.

These improve-stage defaults do not gate explicit standalone extraction through akm proposal extract --type <harness> or akm proposal extract --auto. The interactive scheduled-task step also continues to offer the bundled core/extract template as an unselected opt-in; it is not installed merely because the template is bundled.

Indexing

AKM-native Markdown contributes a normalized body projection to the lowest-weight content search field. The projection is capped at 16,384 characters, removes frontmatter, comments, fenced code, and link destinations, and is never produced for secret, env, session, or session-checkpoint assets. Embedding input is separately capped at 8,192 characters with structured metadata placed before body content.

semanticSearchMode (top-level, "off" | "auto", default "off") gates embedding-based search. "auto" lets AKM set up embeddings (which downloads a local model unless you point embedding at a remote provider) and falls back to keyword-only FTS if the embedding runtime is unavailable; "off" disables semantic search outright and search is always keyword-only FTS. If a backend marked ready cannot serve a query, search still returns the FTS results but reports searchMode: "fts-fallback" and one sanitized warning. This is distinct from searchMode: "keyword", which is the normal result when semantic search is disabled or has not been built. A read-only sandbox that cannot record best-effort usage telemetry does not by itself mark search as degraded. The npm/Bun package declares @huggingface/transformers as a normal dependency. AKM imports that external package directly; it does not carry a copied runtime under src/ or dist/. If the dependency is unavailable, reinstall akm-cli or configure a remote embedding.endpoint. Setup does not mutate a global installation to add runtime packages. The default is "off" so a bare or headless install (akm bundle create, --yes, --config) never silently downloads the local embedding model on first index. The interactive akm setup wizard pre-selects semantic search on regardless of this default, and warns that choosing it downloads the model unless a remote embedding config is provided.

{ "semanticSearchMode": "off" }

embedding configures the connection used for semantic search and akm improve's memory-inference/consolidate passes when they call an embedding model: provider, endpoint, model, apiKey (symbolic reference, same rules as engine apiKey), dimension, localModel, maxInputTokens, maxTokens, batchSize, contextLength, timeoutMs, queryTimeoutMs, queryTemplate, documentTemplate, concurrency, and ollamaOptions.num_ctx.

Retrieval models expect a prompt around queries and documents. akm picks it by model name (src/llm/embedders/profile.ts): Qwen3-Embedding gets Instruct: Given a question or task, retrieve the knowledge asset that helps with it\nQuery:{text} on queries; nomic-embed search_query: / search_document: ; the BGE English, mxbai and arctic models Represent this sentence for searching relevant passages: on queries; E5 query: / passage: ; any other model none. embedding.queryTemplate and embedding.documentTemplate override the preset ({text} marks where the text goes, a template without it is a prefix, and "" means none). The document template is part of the embedding fingerprint, so changing it re-embeds the index; the query template applies at search time only. embedding.queryTimeoutMs (default 3000) bounds how long a search waits for its query embedding before falling back to keyword ranking with a warning.

The knobs that bound request/document size and rate, all optional (defaults apply when unset), for a remote endpoint (src/llm/embedders/remote.ts):

Key Default Bounds
embedding.maxInputTokens 512 Per-DOCUMENT cap, applied before batching (#956). A document's embedded text is truncated to its head (unicode-safe) at this many estimated tokens instead of ever being skipped for size alone — a document is skipped only when its truncated head is empty.
embedding.maxTokens 6000 (DEFAULT_TOKEN_BUDGET) Per-REQUEST token budget: how many (already-capped) documents' estimated tokens fit in one HTTP request. With the 512-token default document cap, a request carries about 11 documents by default. Lowered from 8000 to 6000 (#954): the 4-chars-per-token estimator undercounts dense technical text by 7-55%, so 8000 regularly overshot an 8192-token endpoint's real context window.
embedding.batchSize 100 Per-REQUEST document-COUNT safety cap, independent of the token budget — guards against many tiny documents packing an oversized request.
embedding.contextLength unset Ollama's num_ctx ONLY, forwarded verbatim as options.num_ctx on the native /api/embed request. Does not feed the request token budget above (#956) — the two used to share this one field, so setting it for the server's context window silently changed request batching too.
embedding.timeoutMs 120000 (120s) Per-request wall timeout — see below.
embedding.concurrency 1 loopback / 2 remote In-flight request window — see below.

Which knob fixed the field's 8k-context overflow, worked examples. A 0.9.15-beta field report described documents estimated under the request budget that still tokenized to 8.5k-12.4k real tokens against an 8192-token endpoint, because the 4-chars-per-token estimator undercounts dense technical text. Three knobs changed shape between beta and this release; only one of them makes that overflow structurally unreachable:

A field config of contextLength: 8192 + maxTokens: 8000 (the exact 0.9.15-beta values from the original report) produces no 400s on 0.9.15: maxInputTokens (512, new this release) caps every document before it is counted, so the original 8.5k-12.4k-token documents that overflowed the 8192-token endpoint can never reach the request budget in the first place — independent of whatever maxTokens or contextLength are set to.

embedding.timeoutMs (positive integer, default 120000 — 120s) is the budget for a request at the FULL token budget (embedding.maxTokens); a local model server on a large, token-budget-bounded batch legitimately takes longer than the prior fixed 30s cut off. A smaller request gets a proportionally smaller timeout — clamp(timeoutMs × requestTokens / tokenBudget, 30000, timeoutMs) — so a dead endpoint is still detected in seconds on the common case of small documents. Set embedding.timeoutMs lower to fail fast against a known-fast endpoint, or higher for a slow local server on large batches.

embedding.maxTokens (or its default) is also a run-scoped adaptive starting point, not a hard ceiling (#954): on the FIRST rejection of an akm index run for exceeding the endpoint's context window, akm shrinks the request budget to three quarters of its current value — floored at twice embedding.maxInputTokens — for every request not yet sent, and prints one line naming the new value. This never changes the rejected request's own split-and-retry (below), never shrinks a second time in the same run, and never grows the budget back up. Users who set embedding.maxTokens explicitly are unaffected by the LOWERED DEFAULT above but still benefit from this same-run recovery if their own value turns out to be too high for the endpoint.

A request TIMEOUT (not a rejection for exceeding the context window) never drops its batch immediately: field confirmation showed that once akm abandons a timed-out request, the endpoint (e.g. llama-server) keeps computing it anyway, so dropping it right away just grows the provider's queue while every following batch dies the same way. Instead akm backs off (5s, doubling, capped at 60s) and retries the same request once; a second timeout splits it in half and retries each half the same way, down to individual documents, and a single document that still times out is finally skipped (logged at the default warn level). After 3 consecutive failures at single-document size (timeout or network error), or 3 consecutive network errors at any size, the embedding phase stops dispatching further requests and reports failure — batches already committed are kept; rerun akm index once the endpoint is healthy.

akm index keeps a small number of /v1/embeddings requests in flight at once (a remote endpoint only; the local transformer path is unaffected): 1 for a loopback endpoint (localhost, 127.0.0.0/8, etc. — a local model server serves one inference at a time, and parallel requests thrash it) and 2 for a remote one, unless embedding.concurrency (positive integer, 1-16) overrides it. This default holds for the overwhelming majority of setups; set the override only for an endpoint that genuinely serves parallel requests — a local server started with a multi-slot flag (llama.cpp's --parallel N, vLLM) — not to "speed up" an ordinary single-slot model server, which the default already protects from reload-thrash. Request SIZE remains the first throughput lever regardless: embedding.batchSize (a document-count cap, default 100) together with embedding.maxTokens (an estimated token budget per request, default 6000 — NOT embedding.contextLength, see the table above) control how many documents land in one request — with the default 512-token embedding.maxInputTokens document cap, that is about 11 documents, taking about the same wall time as a single one against a healthy endpoint.

Search tuning

search sets which types search leaves out by default, and the optional curate reranker:

Key Purpose
search.defaultExcludeTypes Asset types excluded from results by default

Curate rerank (#951)

An optional cross-encoder rerank pass over the top fused search candidates akm curate fetches, via a standalone /rerank-style HTTP endpoint (NOT one of the engines.* "llm"/"agent" kinds). Each candidate is sent as its name, description and the start of its indexed content (2,000 characters in all). Disabled by default; a misconfigured endpoint, network failure, timeout, or malformed response keeps the fused order.

Key Purpose
search.curateRerank.enabled Turn the rerank pass on (default false)
search.curateRerank.endpoint Full URL of the reranker's rerank endpoint
search.curateRerank.model Model name sent to the endpoint (optional)
search.curateRerank.apiKey $VAR/secret://<name> credential reference (optional)
search.curateRerank.timeoutMs Request timeout (default 10000)
search.curateRerank.topN How many of the top fused candidates to rerank (default 30, max 50)

Feedback

feedback configures akm feedback:

Key Purpose
feedback.requireReason Whether akm feedback --negative without --reason is a hard error. Defaults to true when unset — set false to downgrade the check to a warning instead

Bundles and write target

bundles (replacing the retired stashDir/sources[]/installed[] trio) and defaultBundle are the 0.9 source configuration shape — see Concepts and the CLI reference for the full bundle model (path, git, website, npm, writable, registryId, components). defaultBundle must name a key in bundles when set. A bundle's components.<id>.adapter key pins it to a specific format adapter instead of relying on auto-detection — see Bundle Types for the full adapter list and what each one reads/writes.

Each physical content root has one bundle id. Duplicate paths and symbolic-link aliases are rejected because source ownership, scheduler authority, and default selection must not depend on which spelling a caller used. If an older config contains aliases, choose the id whose durable refs should survive and remove the other entry before running ordinary commands.

defaultWriteTarget

defaultWriteTarget names the bundle that write commands (akm remember, akm env/secret create, akm improve, etc.) fall back to when no explicit destination flag is given and the command isn't already scoped to a specific source. It must name a configured bundle; setting it with no bundles configured, or naming an unconfigured bundle, is rejected at config set (or config load) time. The full write-target resolution order is the command's destination flag (--bundle on remember/clone/ improve, --target on env/secret create) -> defaultWriteTarget -> working bundle (defaultBundle) -> ConfigError.

Memory scope

akm remember's scope flags (--user, --agent, --run, --channel) write four canonical top-level frontmatter keys on the memory file: scope_user, scope_agent, scope_run, scope_channel (one key per non-empty scope value; string values). This is not a config-file setting — it is documented here because it is the multi-tenant/multi-agent contract that akm search --filter and akm show --filter read back: --filter user=<id> / --filter agent=<id> / --filter run=<id> / --filter channel=<name> (repeatable) narrow results/resolution to assets whose frontmatter scope matches, without changing ranking. A memory with only scope flags and no tags is valid — the tag-required check is independent of scope. --scope was removed in 0.9.0 with no alias; use --filter.

archiveRetentionDays (default 90 when unset) controls how long a pending proposal is kept before akm improve's maintenance pass archives it (status rejected, reason "expired: no action within retention window"; counted from the last akm proposal reopen, if there was one) — akm proposal itself has no archive/expire verb, though akm proposal reopen puts an expired proposal back. Setting it to 0 or less disables expiry entirely.

Registries

registries (top-level array, distinct from bundles) lists remote package registries akm registry/akm bundle add can search and install from. Each entry is { url, name?, enabled?, provider?, options? }; provider defaults to "static-index". See Registries for the full field reference and provider list.

Registry url values must not contain username/password userinfo. The built-in providers do not currently support authenticated registry requests; options does not add an authentication mechanism. Use a credential-free HTTPS endpoint.

Output defaults

output.format (one of json|yaml|text|jsonl|md|html, default json) and output.detail (brief|normal|full, default brief) set the CLI's default --format/--detail when the flags are omitted. Per-command flags always override these.

Setup-derived recommendations

setup is reserved for configuration derived by akm setup. It currently holds no keys — the setup.taskSchedules sub-key was removed in 0.9.0 after nothing in the setup flow or the tasks subsystem was found to read or write it. Scheduling lives in the tasks subsystem (akm task).

Experimental opt-ins

experimental holds explicit opt-ins for behavior outside the 0.9 stability contract (see STABILITY.md for full classification). Every key defaults to off; an absent experimental section, an absent key, and an explicit false all read identically as off.

{
  "experimental": {
    "improveAutonomy": false
  }
}

Managing Config

akm config list
akm config get engines.fast
akm config set engines.fast '{"kind":"llm","endpoint":"http://localhost:11434/v1/chat/completions","model":"qwen3"}'
akm config set engines.fast.apiKey '$LOCAL_LLM_API_KEY'
akm config unset engines.old

Object values passed to config set deep-merge with their current value. Arrays replace, null is only valid for nullable fields, and config unset is the only deletion operation. configVersion cannot be set or unset with the generic walker.

config get <key> --show-source wraps the (redacted) value as { value, source }, where source is "local" when the local file's own JSON sets the key, "extends:<ref>" for the nearest extends chain member that sets it, or "default" when neither does. It is opt-in — plain config get keeps its Stable, script-safe bare-value shape.

Sharing configuration across installs

Five hosts running the same fleet often carry an identical engines map and improve.strategies block, differing only in credential delivery (apiKey vs apiKeyFile), bundle paths, and cron offsets. Hand-syncing that block across hosts drifts silently. extends fixes this: put the shared block in one file, and have each host's local config extend it.

// bundles/fleet/config/shared.json — versioned with the bundle, shared by every host
{
  "configVersion": "0.9.0",
  "engines": {
    "fast": { "kind": "llm", "endpoint": "https://api.example.test/v1/chat/completions", "model": "qwen3" }
  },
  "improve": { "strategies": { "nightly": { "engine": "fast" } } }
}
// ~/.config/akm/config.json — this host's local file, under 20 lines
{
  "configVersion": "0.9.0",
  "extends": "fleet//config/shared.json",
  "bundles": {
    "fleet": { "git": "https://github.com/example/fleet-bundle.git" },
    "stash": { "path": "~/akm-stash", "writable": true }
  },
  "defaultBundle": "stash",
  "engines": { "fast": { "apiKeyFile": "/run/secrets/fast-api-key" } }
}

extends accepts either form:

Shared layers carry portable policy, not host authority. bundles, source and write defaults, registries, embedding connections, scheduler activation, execution, experimental, and setup state are ignored when inherited. Engine definitions may be shared, but credentials and executable authority (apiKey, apiKeyFile, bin, args, and workspace) must be supplied by the local file. Improve publication (strategies.*.sync) and reranker network configuration are local as well. Bundle-relative chains stay physically inside the bundle root for every hop; lexical .. paths and symlink escapes are both rejected before a referenced file is read.

There is no extends: <url> form: config load is synchronous and runs on every invocation, and akm deliberately does not fetch network resources at load time (the same reason registries is never fetched until a registry-touching command runs). A URL-backed shared config should be synced as a git/website bundle and referenced as extends: bundle//<path> once materialized, reusing the sync machinery akm already has instead of a second one inside config load.

The base config runs through the exact same load pipeline as the local file — its own version shim, its own legacy-shape shim — so it can carry an older configVersion independently, and it may itself set extends (chained). Cycle detection (ConfigError, "extends cycle detected") stops A extends B extends A instead of recursing forever. Merge order is DEFAULT_CONFIG (outermost) → the resolved extends chain → the local file's own keys (local always wins) — the same deepMergeConfig "override wins" semantics config set already uses. A referenced file/bundle that does not already exist locally is a load-time ConfigError naming the ref — akm never fetches or syncs one on your behalf.

akm config diff <path|bundle//path> compares this host's EFFECTIVE config (its own extends already applied) against another config file or bundle-relative file (loaded through the same loader, so ITS extends is honoured too), printing sorted { path, local, other } rows for every leaf that differs. Both sides are redacted the same way config get/list are before comparison, so a differing secret never round-trips into the diff output. Cross-host comparison (ssh host2 akm config diff ... in a loop) is left to the operator; akm has no concept of a networked fleet to compare against directly.

akm config diff ~/other-host/config.json
akm config diff fleet//config/shared.json

Environment

Variable Purpose
AKM_CONFIG_DIR Override the user config directory (or set XDG_CONFIG_HOME)
AKM_ENGINE_<NAME>_API_KEY Fallback credential for LLM engine <name>
AKM_LLM_API_KEY Fallback only for the selected defaults.llmEngine
AKM_EMBED_API_KEY Embedding credential
AKM_BUNDLE_DIR Override the bundle directory
AKM_DATA_DIR Override the data directory — index.db, durable state.db, and akm.lock (or set XDG_DATA_HOME)
AKM_CACHE_DIR Override the cache directory — regenerable caches (or set XDG_CACHE_HOME)
AKM_STATE_DIR Override the state directory — task-scheduler invocation state, and (per stash) akm improve's machine-local writers and whole-run lock (or set XDG_STATE_HOME)
AKM_SQLITE_JOURNAL_MODE SQLite journal mode: WAL (default), DELETE, or TRUNCATE
AKM_VERBOSE Truthy value enables the same diagnostics as --verbose
AKM_DEBUG 1 prints a stack trace on unexpected internal errors

For an engine named fast, its fallback variable is AKM_ENGINE_FAST_API_KEY. An explicit apiKey symbolic reference is authoritative and does not fall through to another variable.

engines.<name>.apiKeyFile is a file-backed alternative to apiKey, for a host that refuses to put secrets in the process environment (a container runtime's mounted secret, for example). It is a plain filesystem path — ~ expands to the home directory — read at dispatch time and trimmed of one trailing newline; the raw path is safe to keep in config.json since it is not itself a secret. Setting both apiKey and apiKeyFile on the same engine is rejected. A missing, unreadable, or empty file fails the call closed, naming the engine and path but never the file's content.

engines.<name>.apiKey also accepts secret://<name>, a reference into AKM's own secret store (akm secret set <name> --from-file <file>), for a launch context where the credential's environment variable is deliberately not sourced into the process — a scheduled task's crontab preamble, or a container entrypoint that keeps the user's env out on purpose — and a file-backed credential is not an option. Like apiKeyFile, only the reference is kept in config.json; the store lookup happens at dispatch time, and an unresolved reference fails the call closed, naming the reference but never the value. akm improve, workflow LLM steps, and akm health's engine probes all resolve secret:// the same way direct LLM and embedding calls have since 0.9.13 (#917); resolution order for a single apiKey field is: an env reference ($VAR/${VAR}) first, then apiKeyFile, then secret://<name> — though in practice a config sets only one of the three per engine.

embedding.apiKey accepts the same three forms and resolves secret:// the same way, on every path that sends an embedding request: akm index (including the reindex akm bundle update runs and the targeted re-embed a write command like akm remember triggers), akm improve's consolidate pass (memory dedup and similarity clustering). All of them build the provider request through the same RemoteEmbedder/resolveSecret boundary, so a secret:// reference resolves identically regardless of which command triggered the request (#953).

Use AKM_SQLITE_JOURNAL_MODE=DELETE or TRUNCATE when WAL is unavailable, such as on some NFS/SMB mounts. With the default WAL setting, AKM detects a network filesystem for the data directory and falls back to DELETE.

Retired Configuration

profiles, llm, agent, features, stashes, defaults.llm, defaults.agent, and defaults.improve are rejected in 0.9. Recreate the configuration using engines, defaults.engine, defaults.llmEngine, and improve.strategies; AKM deliberately does not infer or rename ambiguous profile identities.

embedding.chunkSize was never read by anything under src/ (#954), so a config that still sets it is simply ignored — it still loads, unvalidated and without warning.

index.graph.* and every strategy's processes.graphExtraction.* are retired in 0.9.17-alpha.9: the LLM entity graph they configured is gone — akm show's links come from declared links instead (see ## Strategies above). A config that still sets them loads; each key is named once as unknown, and akm migrate apply removes it. The built-in graph-refresh strategy is retired too, but not the same way as an ordinary unknown name: naming it via --strategy or a task always fails with a message pointing at the retirement, even when improve.strategies["graph-refresh"] still has a leftover override from customizing the built-in (the message names it; akm migrate apply drops it — a leftover override is never resolved as a new custom strategy, which would silently run a full, unplanned improve pass). defaults.improveStrategy: "graph-refresh" still loads config successfully; the refusal happens lazily, when the strategy is actually resolved.

improve.strategies.<name>.processes.consolidate.incrementalSince and .neighborsPerChanged are retired in 0.9.17-alpha.9: the consolidate pair pass is now the candidate generator, narrowing per initiator through the improve ledger rather than a global time window. A config that still sets either key loads; each is named once as unknown, and akm migrate apply removes it.