Configuration
AKM reads one user configuration file: $XDG_CONFIG_HOME/akm/config.json
(normally ~/.config/akm/config.json on Linux and macOS, or
%APPDATA%\akm\config.json on Windows). Set AKM_CONFIG_DIR to override the
directory. Project .akm/config.json files are not merged. A config file may
optionally extend one other config via extends (see "Sharing configuration
across installs" below) — this is a single, explicit, user-opted-in key, not
automatic project-config discovery.
Version 0.9
configVersion is "0.9.0", the only value akm has ever shipped. It is
read, never gated on: a file without the field loads silently, and a file
declaring any other value is named once on stderr (config.json declares configVersion "X"; this release reads it as 0.9.0.) and read as the
current shape anyway — nothing is rewritten on disk. The next config write
(akm config set, etc.) and akm migrate apply's config step both persist
"0.9.0", which silences the note. When a newer akm wrote the shared
config, akm health's binary-config-skew advisory is what says so. Pre-0.9
config and database layouts are not runtime inputs. Historical task sources
are handled by the standalone akm-migrate executable, also invoked by akm migrate / akm upgrade; ordinary runtime code reads only the current shape.
{
"configVersion": "0.9.0",
"$schema": "https://itlackey.github.io/akm/schemas/akm-config.json",
"engines": {
"fast": {
"kind": "llm",
"endpoint": "http://localhost:11434/v1/chat/completions",
"model": "qwen3",
"apiKey": "${LOCAL_LLM_API_KEY}"
},
"reviewer": {
"kind": "agent",
"platform": "opencode",
"model": "anthropic/claude-sonnet-4-6"
}
},
"defaults": {
"engine": "reviewer",
"llmEngine": "fast",
"improveStrategy": "default"
},
"workflow": {
"maxConcurrency": 8,
"judgeEngine": "reviewer"
},
"execution": {
"allowedTools": ["read_file", "search"]
},
"improve": {
"strategies": {
"nightly": {
"engine": "fast",
"processes": {
"reflect": {},
"memoryInference": { "model": "qwen3-small", "llm": { "temperature": 0.1 } }
}
}
}
}
}
Scheduler activation
scheduler.enabled is this host's list of scheduled refs, such as
stash//tasks/nightly. A ref that is not listed is disabled. On disk each
entry is still written as the {kind, ref, sourceId} object 0.9.16 reads,
so that release keeps working against a config this one wrote; in memory it
is the ref.
A config without the list (written before 0.9.17) means "keep what is
installed": the first akm task sync fills it from the akm-written native
scheduler rows. The 0.9.17-alpha {kind, ref, sourceId} entries are read as
their ref. Authored task/workflow files may describe schedules but cannot
put themselves on the list.
This key is deliberately local: if a config uses extends, any scheduler
section in the base is ignored with a warning. Only the top-level local config
can activate schedules. Prefer akm task enable <ref> and akm task disable <ref> over editing the JSON by hand; both update the allow-list and sync the
affected bundle. An unscoped akm task sync reconciles enabled refs across all
enabled configured bundles.
Engines
engines is the only public execution map. An engine name is lowercase
kebab-case, at most 63 characters, and cannot start with akm-.
| Kind | Required fields | Use |
|---|---|---|
llm |
endpoint, model |
OpenAI-compatible chat completions |
agent |
platform |
A registered dispatch-capable harness |
LLM endpoints must be complete http:// or https:// chat-completions URLs
ending in /chat/completions, without userinfo, query, or fragment. API keys
are symbolic only: $VAR or ${VAR}. AKM resolves them only at dispatch.
An LLM engine may set enableThinking: false to turn thinking off and
reasoningEffort to a value such as "none", "low", or "high". AKM sends
both wire forms — chat_template_kwargs.enable_thinking and top-level
enable_thinking — whenever enableThinking resolves, from engine config
or a calling process (improve's consolidate/reflect always request
enableThinking: false for a machine-readable payload, and the reflect and
distill quality-gate judges do too unless the judge's engine sets
enableThinking: true); reasoningEffort is always sent as top-level
reasoning_effort when set. Backend support: llama.cpp direct honors both forms
(reasoning_effort from build ≥ b10644); vLLM honors
chat_template_kwargs; Bifrost drops chat_template_kwargs and passes
reasoning_effort through, so also set reasoningEffort: "none" behind it. A
strict hosted API that rejects the two fields (OpenAI answers 400 Unknown parameter: 'chat_template_kwargs') gets one retry without both, and AKM stops
sending them to that endpoint and model for the rest of the process, as it
does for response_format below. Both fields are AKM-owned, not
settable via extraParams. A response with reasoning tokens despite
enableThinking: false triggers a runtime warning and the akm health
thinking-control advisory.
When a call asks for JSON that matches a schema, an LLM engine sends the schema
as response_format (json_schema, strict) unless the engine sets
supportsJsonSchema: false. If the endpoint rejects it with a 4xx other than
429, AKM retries once without response_format and stops sending it to that
endpoint and model for the rest of the process. An agent engine receives the
schema as one instruction at the end of its prompt (Respond with ONLY a JSON value matching this JSON Schema (no prose, no code fences): followed by the
schema), plus the harness's own schema channel where it has one (codex
--output-schema).
An agent engine may set bin, args, workspace, model, and timeoutMs;
it takes no inference of its own (see
Inference on an agent engine). Only platform: "opencode-sdk" may set llmEngine; it names
the LLM engine used as that SDK engine's fallback connection. With no
llmEngine, an SDK engine has no fallback connection and opencode resolves
provider, model and auth from its own configuration. defaults.llmEngine is
not a substitute.
On an engine without timeoutMs, model work (an improve process, a quality or
triage judge, or an index pass) stops after 600 seconds, whatever the engine's
kind; other work on an agent engine runs until it finishes. An improve stage's
reply that the stage cannot read gets one corrective retry. Reflect holds its
reply to its own contract the same way on every engine kind: a reply that is not
the JSON object of reflect's schema gets one repair turn, and one still invalid
fails with parse_error.
Executable assets may request tools, but the request is not authority. Configure
the host-local execution.allowedTools list to define the ceiling; "*" is an
explicit allow-all. The default is an empty list. Asset frontmatter cannot set
workspace, environment, or opaque runtime values; those belong to local
engine configuration or workflow environment bindings.
platform: "opencode-sdk" needs the opencode binary on PATH (or a bin
pointing at it). akm bundles @opencode-ai/sdk, but that package is an HTTP
client with no dependencies — it spawns opencode serve and talks to it — so
the npm dependency alone does not make the platform usable. Install the binary
with npm i -g opencode-ai or opencode's own installer.
Inference on an agent engine
An agent engine sets no inference of its own. For opencode, set it on the
model in your own opencode config, which akm leaves as it is. akm carries
inference (temperature, reasoningEffort, enableThinking, maxTokens,
contextLength) into opencode only where it writes opencode's config itself:
- Model work on
opencodeandopencode-sdk: an improve process'sllmoverlay becomes the options of theakm-model-workagent that runs the dispatch, so opencode's own title call on the same model keeps its defaults. - An
opencode-sdkengine'sllmEnginefallback: the model akm declares for it underakm-customcarries the fallback's inference and the request's.maxTokensandcontextLengthbecomelimit.outputandlimit.context, and only together: opencode refuses half alimit.
Inference from an asset, a workflow's llm: or a models.json alias reaches an
LLM engine. On an agent engine it is reported as an untranslated-field
notice and dispatch continues; model work's agent carries temperature,
reasoningEffort and enableThinking without one.
Reasoning effort has one word in a request, reasoningEffort. effort, as a
models.json alias or an asset's effort: frontmatter spells it, is read as
reasoningEffort wherever layers are merged, so an LLM engine sends it as
reasoning_effort.
Engines for unattended model work
Unattended model work is the work akm hands a model with no one watching:
- the improve processes (reflect, distill, consolidate, memory inference, extract, and validation's repair);
- the quality, triage and retrieval-gate judges;
- index passes;
akm remember --enrich.
It runs on any engine kind under one tool policy. The model may read, edit
files only inside a scratch working directory that akm creates for the
dispatch and removes after it, and run akm search and akm show. The
stash stays read-only to it: what model work changes reaches the stash only
as a proposal, through the review queue.
Each engine enforces as much of the policy as it can, and grants nothing it cannot enforce:
| Engine | What the model gets |
|---|---|
| LLM | No tools. |
claude |
Read and Edit inside the working directory, and Bash for akm search and akm show only. akm runs it with --restricted, so your user, project and local settings cannot widen that. |
opencode, opencode-sdk |
Read, grep and glob in the working directory and the stash, edit in the working directory only, and akm_search and akm_show, through an injected akm-model-work agent. No bash, because opencode cannot stop a redirect such as akm show x > file from writing elsewhere. |
codex, copilot, pi, gemini, aider, amazonq, openhands |
Cannot run model work: akm refuses the request before it starts. |
The one rule. Every key model work reads its engine from must name an
LLM engine or a claude, opencode or opencode-sdk agent engine. Those
keys are:
defaults.llmEngine;index.defaults.engineandindex.<pass>.engine;improve.strategies.<name>.engine;improve.strategies.<name>.processes.<process>.engine;- an enabled
processes.triage.judgment.engine; processes.<process>.qualityGate.engine.
A config that breaks the rule fails to load, with an error that names the
key, the engine and its platform. Other engine keys, defaults.engine and
workflow.judgeEngine, may name any configured engine.
Further details:
- A model-work dispatch builds its own command, so the engine's
argsdo not apply to it, except a--modelthey name, and neither does itsworkspace. - An improve process's
llmoverlay reaches theakm-model-workagent onopencodeandopencode-sdk(see Inference on an agent engine); on any other agent engine it is reported asuntranslated-fieldnotices, not errors. - opencode model work may read the stash and write only its working directory.
akm_searchandakm_showcome from the akm-opencode plugin (0.9.21 or later), which akm does not load: put it in your opencode config ("plugin": ["akm-opencode"]). akm turns off the plugin's curation, learning and write gate for these dispatches and keeps its state in akm's state directory. The stash is protected from edits only while the temporary directory is outside a git repository.
Model-map files
AKM ships an immutable models.json package asset with three intent aliases:
fast, balanced, and reasoning. The installed starter has separate
columns for Claude Code, OpenCode, and OpenCode SDK only. A provider-specific
identifier is not pretended to work on unrelated engines. Add mappings for
Gemini, Codex, named direct LLM engines, or other harnesses in your user file.
Config-root and per-engine modelAliases are rejected; this file is the only
alias definition surface.
An optional operator-owned file lives beside config.json at
$XDG_CONFIG_HOME/akm/models.json (or <AKM_CONFIG_DIR>/models.json). It uses
the same version-1 schema as the installed file:
{
"version": 1,
"aliases": {
"fast": {
"gemini": "gemini-2.5-flash"
},
"reasoning": {
"claude": {
"inference": {
"effort": "medium"
}
},
"local-reasoner": {
"model": "qwen3:30b",
"inference": {
"effort": "high"
}
}
}
}
}
Each engine mapping is either a non-empty exact model string or a structured
profile with the documented fields model, inference, and engine. A user
profile may omit model when the installed layer already supplies it, as the
partial Claude override above does. After overlay, every alias/engine entry
must have a usable model. Unknown profile fields are rejected; JSON-safe
fields inside inference are preserved for engine adapters to lower
optimistically. An inference.effort is read as reasoningEffort
(see Inference on an agent engine).
A profile's engine field (0.9.15, #946) borrows a column's model (and, for
an llm-kind engine, its inference defaults) from a configured
engines.<name> connection instead of hand-typing a literal model a second
time:
{
"version": 1,
"aliases": {
"fast": {
"opencode": { "engine": "local-fast" }
}
}
}
With engines.local-fast configured (agent-kind or llm-kind), this column
resolves to that engine's own model string. model and engine are
mutually exclusive on the same profile — engine is an indirection for the
model value, never an engine-selection override; which engine akm agent
dispatches to is still decided entirely by --engine/defaults.engine (see
Engine selection). The referenced engine's model must itself be
literal, not another alias, and akm copies it verbatim: it does not translate
between an engine's connection and an agent platform's own provider registry,
so the value must already be meaningful for the column's platform (e.g. a
kind: "agent", platform: "opencode" engine's model should already be a
string opencode itself understands, such as krang/qwen3.5-9b). Run
akm models list to see, for every alias/column, the resolved model and
whether it came from the installed defaults, the user overlay, and a literal
value or an engine reference.
The user file overlays the installed file by alias, engine, and nested object
field. Objects merge recursively. Arrays, scalars, and explicit null replace
the lower value; omitted fields preserve it. A layer setting a literal model
clears any engine inherited from a farther layer, and vice versa — the
nearer layer's choice of literal-vs-engine always wins outright rather than
merging. Alias and engine keys are case-normalized, and case-colliding
definitions are rejected. Unknown model inputs still pass through
byte-for-byte as exact identifiers. Once a name is a known merged alias,
selecting an engine with no mapping is an actionable configuration error
rather than silently sending the alias as a model ID.
The common execution cascade reads these files for current direct command and
non-interactive agent calls, task source v4 runs, and improve/proposal/index
model work routed through that resolver. A structured alias expands as
defaults at the layer that selected it; explicit sibling fields and nearer
layers still win. The resulting request carries the exact model ID and merged
inference object. Engine lowerers consume that exact selection and never run
alias resolution again. New workflow starts persist the exact request and
symbolic runner selection in the durable plan v4 family's executable
irVersion: 5; resume consumes that frozen material without resolving aliases
again.
Copy the complete installed starter into the user configuration directory when you want to customize all fields:
akm models copy-defaults
akm models copy-defaults --overwrite # explicit replacement confirmation
The command validates the installed asset, creates the config directory, and
writes a fully synced sibling before publication. Without --overwrite, a
hard-link/no-replace operation makes publication atomic: a racing creator wins
without losing its bytes. A filesystem that cannot provide that operation
fails safely instead of falling back to a clobbering rename.
With --overwrite, the portable guarantee is an atomic pathname replacement
that never follows the target when it is a symlink. AKM verifies the observed
regular-file identity again immediately before rename, but the portable
filesystem APIs do not provide a conditional compare-and-swap rename. Another
process can still change the directory entry after that check; AKM replaces
the entry at the pathname without dereferencing it. Consequently,
overwritten: true means overwrite was requested for an entry AKM observed,
not that an inode identity was transactionally locked. Symlinks and other
non-regular targets observed at either check are refused.
AKM does not auto-create or sync this file, and authoritative defaults never
live in the cache. npm/Node and normal Bun installs read the packaged
dist/assets/models.json lazily, so akm health can report a missing or
malformed package asset as a model-map-files failure. A standalone binary has
the same authoritative bytes embedded at compile time and therefore has no
external model-map asset that can later disappear; its health check validates
the embedded copy, and release tests pin copied bytes to src/assets/models.json.
The health check passes when the optional user file is absent and warns with
its path and JSON location when the user file is unreadable or invalid.
defaults.engine names an LLM or agent engine. defaults.llmEngine names the
default engine for unattended model work, so it follows
the one rule. There is no first-engine
fallback: an unset defaults.engine
never resolves to some arbitrary entry in engines. It resolves instead to a
synthesized, config-free opencode-sdk engine when the opencode binary is on
PATH — announced once per run, and preempted by any opencode-sdk engine you
configure yourself. Naming an engine that is not configured is always an error
and is never rescued by that fallback.
defaults.llmEngine is not an opencode-sdk engine's fallback connection. An
SDK engine gets an LLM fallback only from its own llmEngine, so the
synthesized engine, which sets none, runs on opencode's own provider, model and
auth.
Index passes select engines through index.defaults.engine or
index.<pass>.engine, which follow
the one rule. Per-pass model, timeoutMs, and llm fields are
invocation overrides; enabled: false disables that pass. Connection fields
such as endpoint, provider, apiKey, and apiKeyFile belong only on
named engines.
workflow.maxConcurrency is the native workflow engine ceiling. An explicit
value is clamped to 1..64. When absent, AKM derives the cap once from the CPU
count (min(16, max(1, cores - 2))) and freezes it into the run plan, so resume
does not change policy on a different host or after config edits.
workflow.defaultMapConcurrency is the width a map step freezes when it
declares no concurrency: of its own. Unset means 4 — map steps are
parallel by default as of 0.9.1. An explicit value is clamped to 1..64; set
it to 1 to restore the pre-0.9.1 serial-by-default fan-out for every workflow
on this machine. It is only a default: an authored map.concurrency always
wins, and it never raises a step past workflow.maxConcurrency, the selected
engine's concurrency, or the host CPU cap. An LLM engine that declares no
engines.<name>.concurrency gets 1 on a loopback endpoint (a local model
server holds one loaded model) and 4 on a remote one.
workflow.judgeEngine names the LLM or agent engine used to verify every
non-empty workflow ### gate rubric. It is required when a workflow declares
completion criteria and is frozen into each new run, so later config edits do
not change an in-flight run's verifier. Missing, failed, or malformed verifier
results reject the gate; criteria are never silently bypassed.
Strategies
Improve presets live under improve.strategies; invoke one with
akm improve --strategy <name>. The selection order is --strategy,
defaults.improveStrategy, then built-in default. A strategy and each process
can select engine, model, timeoutMs, and LLM request overrides:
{
"improve": {
"strategies": {
"nightly": {
"engine": "fast",
"processes": {
"reflect": { "llm": { "temperature": 0.2 } },
"memoryInference": { "model": "qwen3-small" }
}
}
}
}
}
An improve process's engine follows
the one rule; an explicit invalid or
incompatible engine never falls back to another engine. Built-in strategies
are complete presets. User-defined strategies inherit omitted fields from the
built-in default strategy before applying their own overrides.
processes.triage.judgment explicitly controls the optional judgment tier.
Use true to enable it, false to disable it, or an object with enabled,
engine, model, timeoutMs, and/or llm overrides. Existing object values
such as {} and { "engine": "reviewer" } remain enabled by default. Unknown
object keys are rejected so misspellings cannot silently change execution;
the retired mode and profile keys continue to report their engine migration
guidance. When enabled, engine selection is judgment → triage → strategy →
defaults.llmEngine, and resolution fails closed if none is available.
{
"improve": {
"strategies": {
"nightly": {
"processes": {
"triage": {
"enabled": true,
"judgment": { "enabled": true, "engine": "reviewer" }
}
}
}
}
}
}
processes.reflect.qualityGate and processes.distill.qualityGate control
each process's LLM-as-judge quality gate. Each is on unless it sets
enabled: false, and each follows only its own switch. A reflect revision the
judge passes is staged for the triage drain to accept; a distill lesson it
passes is deferred for a person (reason distill-review), which the drain and
its judgment tier leave alone. With the distill gate off nothing is judged, and
the drain decides. The judge is the
process's own engine when that is an LLM engine, or the defaults.llmEngine
engine when an agent generates. engine, model, timeoutMs and llm give the gate a judge of its own,
resolved over the process's settings the way triage.judgment resolves over
triage's. The judge may be any engine that follows
the one rule. A gate whose settings
resolve to no engine fails before anything is generated; it never falls back
to another engine. The judge runs at
temperature 0 with thinking off unless its engine sets enableThinking: true.
Thinking is slow: on a 27B llama.cpp server, a thinking judgment took a median
of 30–67 s and up to about 3 minutes, against about 5 s without.
Behind a gateway that drops chat_template_kwargs (Bifrost), point a thinking
judge's engine at the server directly.
{
"engines": {
"judge": {
"kind": "llm",
"endpoint": "http://127.0.0.1:8080/v1/chat/completions",
"model": "qwen3-27b",
"enableThinking": true
}
},
"improve": {
"strategies": {
"nightly": {
"processes": {
"reflect": { "qualityGate": { "engine": "judge" } }
}
}
}
}
}
processes.reflect.defectFilter sets the wording of the checks reflect runs
before the judge. With the quality gate on or off, reflect refuses a revision
that adds placeholder text, talks about its own edit, or copies frontmatter into
its body: no proposal, and no judge call. Each rule counts only what the
revision adds to its source, so wording the asset already had and kept is not
held against it. Each of the three lists is optional. A list you set replaces
that rule's default list, and [] turns the rule off.
| List | Rule | Default |
|---|---|---|
placeholders |
placeholder_added |
please confirm, please verify, to be confirmed, to be determined, to be verified |
metaCommentary |
meta_commentary_added |
feedback signal, feedback signals, feedback indicate, feedback indicates, feedback suggest, feedback suggests, feedback ask, feedback asks, feedback says, feedback report, feedback reports, feedback request, feedback requests, this revision, the source asset, the source note, the source memory, the original asset, the original note, the original memory, the original version of this, quality gate rejected, proposal rejected |
frontmatterKeys |
frontmatter_copied_into_body |
sources, updated, inferenceProcessed, captureMode, beliefState, xrefs, contradictedBy, outcomeData, orderedActions, generated, verified, description, when_to_use, tags, searchHints, quality, salience, salienceInputs, lint_skip, type |
placeholders and metaCommentary entries are plain phrases, not patterns:
whole words, in any case, with any run of whitespace between words.
frontmatterKeys entries are exact key names: a line outside a code fence that
starts with key: counts. The rule also refuses a sources, xrefs or
contradictedBy value copied into the body; frontmatterKeys: [] turns that
off too. Every entry must be a non-empty string, or the config does not load.
{
"improve": {
"strategies": {
"nightly": {
"processes": {
"reflect": {
"defectFilter": {
"placeholders": ["please confirm", "to be confirmed", "[draft]"],
"frontmatterKeys": []
}
}
}
}
}
}
}
No shipped strategy turns improve-stage session extraction on.
proactiveMaintenance is on only in the proactive-maintenance preset; run
akm improve --strategy proactive-maintenance to use that opt-in preset.
Because strategies inherit from default, a preset that omits either process
also inherits the off value. User strategy overrides
are applied last, so an explicit enabled: true still opts the selected
strategy in.
These improve-stage defaults do not gate explicit standalone extraction through
akm proposal extract --type <harness> or akm proposal extract --auto. The interactive
scheduled-task step also continues to offer the bundled core/extract template
as an unselected opt-in; it is not installed merely because the template is
bundled.
Indexing
AKM-native Markdown contributes a normalized body projection to the
lowest-weight content search field. The projection is capped at 16,384
characters, removes frontmatter, comments, fenced code, and link destinations,
and is never produced for secret, env, session, or session-checkpoint assets.
Embedding input is separately capped at 8,192 characters with structured
metadata placed before body content.
Semantic search
semanticSearchMode (top-level, "off" | "auto", default "off") gates
embedding-based search. "auto" lets AKM set up embeddings (which downloads
a local model unless you point embedding at a remote provider) and falls
back to keyword-only FTS if the embedding runtime is unavailable; "off"
disables semantic search outright and search is always keyword-only FTS.
If a backend marked ready cannot serve a query, search still returns the FTS
results but reports searchMode: "fts-fallback" and one sanitized warning.
This is distinct from searchMode: "keyword", which is the normal result when
semantic search is disabled or has not been built. A read-only sandbox that
cannot record best-effort usage telemetry does not by itself mark search as
degraded.
The npm/Bun package declares @huggingface/transformers as a normal dependency.
AKM imports that external package directly; it does not carry a copied runtime
under src/ or dist/. If the dependency is unavailable, reinstall akm-cli
or configure a remote embedding.endpoint. Setup does not mutate a global
installation to add runtime packages.
The default is "off" so a bare or headless install (akm bundle create, --yes,
--config) never silently downloads the local embedding model on first
index.
The interactive akm setup wizard pre-selects semantic search on
regardless of this default, and warns that choosing it downloads the model
unless a remote embedding config is provided.
{ "semanticSearchMode": "off" }
embedding configures the connection used for semantic search and
akm improve's memory-inference/consolidate passes when they call an
embedding model: provider, endpoint, model, apiKey (symbolic
reference, same rules as engine apiKey), dimension, localModel,
maxInputTokens, maxTokens, batchSize, contextLength, timeoutMs,
queryTimeoutMs, queryTemplate, documentTemplate, concurrency, and
ollamaOptions.num_ctx.
Retrieval models expect a prompt around queries and documents. akm picks it by
model name (src/llm/embedders/profile.ts): Qwen3-Embedding gets
Instruct: Given a question or task, retrieve the knowledge asset that helps with it\nQuery:{text}
on queries; nomic-embed search_query: / search_document: ; the BGE
English, mxbai and arctic models Represent this sentence for searching relevant passages: on queries; E5 query: / passage: ; any other model
none. embedding.queryTemplate and embedding.documentTemplate override the
preset ({text} marks where the text goes, a template without it is a prefix,
and "" means none). The document template is part of the embedding
fingerprint, so changing it re-embeds the index; the query template applies
at search time only. embedding.queryTimeoutMs (default 3000) bounds how
long a search waits for its query embedding before falling back to keyword
ranking with a warning.
The knobs that bound request/document size and rate, all optional (defaults
apply when unset), for a remote endpoint (src/llm/embedders/remote.ts):
| Key | Default | Bounds |
|---|---|---|
embedding.maxInputTokens |
512 |
Per-DOCUMENT cap, applied before batching (#956). A document's embedded text is truncated to its head (unicode-safe) at this many estimated tokens instead of ever being skipped for size alone — a document is skipped only when its truncated head is empty. |
embedding.maxTokens |
6000 (DEFAULT_TOKEN_BUDGET) |
Per-REQUEST token budget: how many (already-capped) documents' estimated tokens fit in one HTTP request. With the 512-token default document cap, a request carries about 11 documents by default. Lowered from 8000 to 6000 (#954): the 4-chars-per-token estimator undercounts dense technical text by 7-55%, so 8000 regularly overshot an 8192-token endpoint's real context window. |
embedding.batchSize |
100 |
Per-REQUEST document-COUNT safety cap, independent of the token budget — guards against many tiny documents packing an oversized request. |
embedding.contextLength |
unset | Ollama's num_ctx ONLY, forwarded verbatim as options.num_ctx on the native /api/embed request. Does not feed the request token budget above (#956) — the two used to share this one field, so setting it for the server's context window silently changed request batching too. |
embedding.timeoutMs |
120000 (120s) |
Per-request wall timeout — see below. |
embedding.concurrency |
1 loopback / 2 remote |
In-flight request window — see below. |
Which knob fixed the field's 8k-context overflow, worked examples. A 0.9.15-beta field report described documents estimated under the request budget that still tokenized to 8.5k-12.4k real tokens against an 8192-token endpoint, because the 4-chars-per-token estimator undercounts dense technical text. Three knobs changed shape between beta and this release; only one of them makes that overflow structurally unreachable:
embedding.maxInputTokens: 512— per-DOCUMENT cap, applied before batching. Example: a 6,000-character API reference page is truncated to its first ~2,000 characters (512 estimated tokens) before it is ever counted toward a request. This is the fix for the original overflow: no single document can contribute more than 512 estimated tokens to a request, no matter howmaxTokensorcontextLengthare set.embedding.maxTokens: 6000— per-REQUEST budget: how many already-capped documents' estimated tokens fit in one HTTP request. Example: with the default 512-token document cap, a request packs about 11 documents before this budget is reached and the request is sent; if the run's first request is still rejected for exceeding the endpoint's real context window, akm shrinks this budget to three quarters of its value (floored at twicemaxInputTokens) for every later request in the same run. A request-level budget alone cannot stop one oversized document from overflowing a request — only the per-document cap above does that.embedding.contextLength: 8192— Ollama'snum_ctxonly, forwarded verbatim on a native/api/embedrequest. It has no effect on request or document sizing, and no effect at all against a non-Ollama endpoint — see below for why that used not to be true.
A field config of contextLength: 8192 + maxTokens: 8000 (the exact
0.9.15-beta values from the original report) produces no 400s on 0.9.15:
maxInputTokens (512, new this release) caps every document before it is
counted, so the original 8.5k-12.4k-token documents that overflowed the
8192-token endpoint can never reach the request budget in the first place —
independent of whatever maxTokens or contextLength are set to.
embedding.timeoutMs (positive integer, default 120000 — 120s) is the
budget for a request at the FULL token budget (embedding.maxTokens); a
local model server on a large, token-budget-bounded batch legitimately takes
longer than the prior fixed 30s cut off. A smaller request gets a
proportionally smaller timeout —
clamp(timeoutMs × requestTokens / tokenBudget, 30000, timeoutMs) — so a
dead endpoint is still detected in seconds on the common case of small
documents. Set embedding.timeoutMs lower to fail fast against a
known-fast endpoint, or higher for a slow local server on large batches.
embedding.maxTokens (or its default) is also a run-scoped adaptive
starting point, not a hard ceiling (#954): on the FIRST rejection of an
akm index run for exceeding the endpoint's context window, akm shrinks
the request budget to three quarters of its current value — floored at
twice embedding.maxInputTokens — for every request not yet sent, and
prints one line naming the new value. This never changes the rejected
request's own split-and-retry (below), never shrinks a second time in the
same run, and never grows the budget back up. Users who set
embedding.maxTokens explicitly are unaffected by the LOWERED DEFAULT
above but still benefit from this same-run recovery if their own value
turns out to be too high for the endpoint.
A request TIMEOUT (not a rejection for exceeding the context window) never
drops its batch immediately: field confirmation showed that once akm
abandons a timed-out request, the endpoint (e.g. llama-server) keeps
computing it anyway, so dropping it right away just grows the provider's
queue while every following batch dies the same way. Instead akm backs off
(5s, doubling, capped at 60s) and retries the same request once; a second
timeout splits it in half and retries each half the same way, down to
individual documents, and a single document that still times out is finally
skipped (logged at the default warn level). After 3 consecutive failures
at single-document size (timeout or network error), or 3 consecutive
network errors at any size, the embedding phase stops dispatching further
requests and reports failure — batches already committed are kept; rerun
akm index once the endpoint is healthy.
akm index keeps a small number of /v1/embeddings requests in flight at
once (a remote endpoint only; the local transformer path is unaffected):
1 for a loopback endpoint (localhost, 127.0.0.0/8, etc. — a local
model server serves one inference at a time, and parallel requests thrash
it) and 2 for a remote one, unless embedding.concurrency (positive
integer, 1-16) overrides it. This default holds for the overwhelming
majority of setups; set the override only for an endpoint that genuinely
serves parallel requests — a local server started with a multi-slot flag
(llama.cpp's --parallel N, vLLM) — not to "speed up" an ordinary
single-slot model server, which the default already protects from
reload-thrash. Request SIZE remains the first throughput lever regardless:
embedding.batchSize (a document-count cap, default 100) together with
embedding.maxTokens (an estimated token budget per request, default 6000
— NOT embedding.contextLength, see the table above) control how many
documents land in one request — with the default 512-token
embedding.maxInputTokens document cap, that is about 11 documents,
taking about the same wall time as a single one against a healthy endpoint.
Search tuning
search sets which types search leaves out by default, and the optional curate reranker:
| Key | Purpose |
|---|---|
search.defaultExcludeTypes |
Asset types excluded from results by default |
Curate rerank (#951)
An optional cross-encoder rerank pass over the top fused search candidates
akm curate fetches, via a standalone /rerank-style HTTP endpoint (NOT one
of the engines.* "llm"/"agent" kinds). Each candidate is sent as its
name, description and the start of its indexed content (2,000 characters in
all). Disabled by default; a misconfigured endpoint, network failure, timeout,
or malformed response keeps the fused order.
| Key | Purpose |
|---|---|
search.curateRerank.enabled |
Turn the rerank pass on (default false) |
search.curateRerank.endpoint |
Full URL of the reranker's rerank endpoint |
search.curateRerank.model |
Model name sent to the endpoint (optional) |
search.curateRerank.apiKey |
$VAR/secret://<name> credential reference (optional) |
search.curateRerank.timeoutMs |
Request timeout (default 10000) |
search.curateRerank.topN |
How many of the top fused candidates to rerank (default 30, max 50) |
Feedback
feedback configures akm feedback:
| Key | Purpose |
|---|---|
feedback.requireReason |
Whether akm feedback --negative without --reason is a hard error. Defaults to true when unset — set false to downgrade the check to a warning instead |
Bundles and write target
bundles (replacing the retired stashDir/sources[]/installed[] trio)
and defaultBundle are the 0.9 source configuration shape — see
Concepts and the CLI reference for the
full bundle model (path, git, website, npm, writable, registryId,
components). defaultBundle must name a key in bundles when set. A
bundle's components.<id>.adapter key pins it to a specific format adapter
instead of relying on auto-detection — see Bundle Types
for the full adapter list and what each one reads/writes.
Each physical content root has one bundle id. Duplicate paths and symbolic-link aliases are rejected because source ownership, scheduler authority, and default selection must not depend on which spelling a caller used. If an older config contains aliases, choose the id whose durable refs should survive and remove the other entry before running ordinary commands.
defaultWriteTarget
defaultWriteTarget names the bundle that write commands (akm remember,
akm env/secret create, akm improve, etc.) fall back to when no
explicit destination flag is given and the command isn't already scoped to a
specific source. It must name a configured bundle; setting it with no
bundles configured, or naming an unconfigured bundle, is rejected at
config set (or config load) time. The full write-target resolution order
is the command's destination flag (--bundle on remember/clone/
improve, --target on env/secret create) -> defaultWriteTarget ->
working bundle (defaultBundle) -> ConfigError.
Memory scope
akm remember's scope flags (--user, --agent, --run, --channel)
write four canonical top-level frontmatter keys on the memory file:
scope_user, scope_agent, scope_run, scope_channel (one key per
non-empty scope value; string values). This is not a config-file setting —
it is documented here because it is the multi-tenant/multi-agent contract
that akm search --filter and akm show --filter read back:
--filter user=<id> / --filter agent=<id> / --filter run=<id> /
--filter channel=<name> (repeatable) narrow results/resolution to assets
whose frontmatter scope matches, without changing ranking. A memory with
only scope flags and no tags is valid — the tag-required check is
independent of scope. --scope was removed in 0.9.0 with no alias; use
--filter.
archiveRetentionDays (default 90 when unset) controls how long a pending
proposal is kept before akm improve's maintenance pass archives it (status
rejected, reason "expired: no action within retention window"; counted
from the last akm proposal reopen, if there was one) — akm proposal itself
has no archive/expire verb, though akm proposal reopen puts an expired
proposal back. Setting it to 0 or less disables expiry entirely.
Registries
registries (top-level array, distinct from bundles) lists remote package
registries akm registry/akm bundle add can search and install from.
Each entry is { url, name?, enabled?, provider?, options? }; provider
defaults to "static-index". See Registries for the full
field reference and provider list.
Registry url values must not contain username/password userinfo. The built-in
providers do not currently support authenticated registry requests; options
does not add an authentication mechanism. Use a credential-free HTTPS endpoint.
Output defaults
output.format (one of json|yaml|text|jsonl|md|html,
default json) and output.detail (brief|normal|full, default
brief) set the CLI's default --format/--detail when the flags are
omitted. Per-command flags always override these.
Setup-derived recommendations
setup is reserved for configuration derived by akm setup. It currently
holds no keys — the setup.taskSchedules sub-key was removed in 0.9.0 after
nothing in the setup flow or the tasks subsystem was found to read or write
it. Scheduling lives in the tasks subsystem (akm task).
Experimental opt-ins
experimental holds explicit opt-ins for behavior outside the 0.9
stability contract (see STABILITY.md for full
classification). Every key defaults to off; an absent experimental
section, an absent key, and an explicit false all read identically as off.
{
"experimental": {
"improveAutonomy": false
}
}
experimental.improveAutonomy— gates only the autonomousmemoryInference,triagePromote, andmemoryCleanuplanes.akm improveitself always runs; this only gates mutations without a human in the loop. Consolidation is not gated: it remains advisory and emits reviewable proposals.sync.pushis deliberately not gated by this key.
Managing Config
akm config list
akm config get engines.fast
akm config set engines.fast '{"kind":"llm","endpoint":"http://localhost:11434/v1/chat/completions","model":"qwen3"}'
akm config set engines.fast.apiKey '$LOCAL_LLM_API_KEY'
akm config unset engines.old
Object values passed to config set deep-merge with their current value.
Arrays replace, null is only valid for nullable fields, and config unset is
the only deletion operation. configVersion cannot be set or unset with the
generic walker.
config get <key> --show-source wraps the (redacted) value as
{ value, source }, where source is "local" when the local file's own
JSON sets the key, "extends:<ref>" for the nearest extends chain member
that sets it, or "default" when neither does. It is opt-in — plain
config get keeps its Stable, script-safe bare-value shape.
Sharing configuration across installs
Five hosts running the same fleet often carry an identical engines map and
improve.strategies block, differing only in credential delivery (apiKey
vs apiKeyFile), bundle paths, and cron offsets. Hand-syncing that block
across hosts drifts silently. extends fixes this: put the shared block in
one file, and have each host's local config extend it.
// bundles/fleet/config/shared.json — versioned with the bundle, shared by every host
{
"configVersion": "0.9.0",
"engines": {
"fast": { "kind": "llm", "endpoint": "https://api.example.test/v1/chat/completions", "model": "qwen3" }
},
"improve": { "strategies": { "nightly": { "engine": "fast" } } }
}
// ~/.config/akm/config.json — this host's local file, under 20 lines
{
"configVersion": "0.9.0",
"extends": "fleet//config/shared.json",
"bundles": {
"fleet": { "git": "https://github.com/example/fleet-bundle.git" },
"stash": { "path": "~/akm-stash", "writable": true }
},
"defaultBundle": "stash",
"engines": { "fast": { "apiKeyFile": "/run/secrets/fast-api-key" } }
}
extends accepts either form:
- A filesystem path — relative paths resolve against the directory of the
config file that declares them; a leading
~expands. - A
bundle//<path>ref — a plain file path relative to that bundle's content root (e.g.config/shared.json), resolved through the bundle's configuredpath, not the search index — so it never needsakm indexto have run. This is not an asset ref: the path after//needs no asset type (scripts/,knowledge/, …) and the shared file is never indexed; it can live anywhere under the bundle. An empty, absolute, or content-root-escaping path is rejected. Only a filesystem bundle (bundles.<id>.path) can host anextendssource; sync agit/websitebundle withakm bundle add/akm syncfirst so the file is materialized locally, then pointextendsat it.
Shared layers carry portable policy, not host authority. bundles, source and
write defaults, registries, embedding connections, scheduler activation,
execution, experimental, and setup state are ignored when inherited.
Engine definitions may be shared, but credentials and executable authority
(apiKey, apiKeyFile, bin, args, and workspace) must be supplied by
the local file. Improve publication (strategies.*.sync) and reranker network
configuration are local as well. Bundle-relative chains stay physically inside
the bundle root for every hop; lexical .. paths and symlink escapes are both
rejected before a referenced file is read.
There is no extends: <url> form: config load is synchronous and runs on
every invocation, and akm deliberately does not fetch network resources at
load time (the same reason registries is never fetched until a
registry-touching command runs). A URL-backed shared config should be synced
as a git/website bundle and referenced as extends: bundle//<path>
once materialized, reusing the sync machinery akm already has instead of a
second one inside config load.
The base config runs through the exact same load pipeline as the local
file — its own version shim, its own legacy-shape shim — so it can carry an
older configVersion independently, and it may itself set extends
(chained). Cycle detection (ConfigError, "extends cycle detected") stops A
extends B extends A instead of recursing forever. Merge order is
DEFAULT_CONFIG (outermost) → the resolved extends chain → the local
file's own keys (local always wins) — the same deepMergeConfig "override
wins" semantics config set already uses. A referenced file/bundle that does
not already exist locally is a load-time ConfigError naming the ref — akm
never fetches or syncs one on your behalf.
akm config diff <path|bundle//path> compares this host's EFFECTIVE
config (its own extends already applied) against another config file or
bundle-relative file (loaded through the same loader, so ITS extends is
honoured too), printing sorted { path, local, other } rows for every leaf
that differs. Both sides are redacted the same way config get/list are
before comparison, so a differing secret never round-trips into the diff
output. Cross-host comparison (ssh host2 akm config diff ... in a loop) is
left to the operator; akm has no concept of a networked fleet to compare
against directly.
akm config diff ~/other-host/config.json
akm config diff fleet//config/shared.json
Environment
| Variable | Purpose |
|---|---|
AKM_CONFIG_DIR |
Override the user config directory (or set XDG_CONFIG_HOME) |
AKM_ENGINE_<NAME>_API_KEY |
Fallback credential for LLM engine <name> |
AKM_LLM_API_KEY |
Fallback only for the selected defaults.llmEngine |
AKM_EMBED_API_KEY |
Embedding credential |
AKM_BUNDLE_DIR |
Override the bundle directory |
AKM_DATA_DIR |
Override the data directory — index.db, durable state.db, and akm.lock (or set XDG_DATA_HOME) |
AKM_CACHE_DIR |
Override the cache directory — regenerable caches (or set XDG_CACHE_HOME) |
AKM_STATE_DIR |
Override the state directory — task-scheduler invocation state, and (per stash) akm improve's machine-local writers and whole-run lock (or set XDG_STATE_HOME) |
AKM_SQLITE_JOURNAL_MODE |
SQLite journal mode: WAL (default), DELETE, or TRUNCATE |
AKM_VERBOSE |
Truthy value enables the same diagnostics as --verbose |
AKM_DEBUG |
1 prints a stack trace on unexpected internal errors |
For an engine named fast, its fallback variable is
AKM_ENGINE_FAST_API_KEY. An explicit apiKey symbolic reference is
authoritative and does not fall through to another variable.
engines.<name>.apiKeyFile is a file-backed alternative to apiKey, for a
host that refuses to put secrets in the process environment (a container
runtime's mounted secret, for example). It is a plain filesystem path — ~
expands to the home directory — read at dispatch time and trimmed of one
trailing newline; the raw path is safe to keep in config.json since it is
not itself a secret. Setting both apiKey and apiKeyFile on the same
engine is rejected. A missing, unreadable, or empty file fails the call
closed, naming the engine and path but never the file's content.
engines.<name>.apiKey also accepts secret://<name>, a reference into
AKM's own secret store (akm secret set <name> --from-file <file>), for a
launch context where the credential's environment variable is deliberately
not sourced into the process — a scheduled task's crontab preamble, or a
container entrypoint that keeps the user's env out on purpose — and a
file-backed credential is not an option. Like apiKeyFile, only the
reference is kept in config.json; the store lookup happens at dispatch
time, and an unresolved reference fails the call closed, naming the
reference but never the value. akm improve, workflow LLM steps, and akm health's engine probes all resolve secret:// the same way direct LLM and
embedding calls have since 0.9.13 (#917); resolution order for a single
apiKey field is: an env reference ($VAR/${VAR}) first, then
apiKeyFile, then secret://<name> — though in practice a config sets only
one of the three per engine.
embedding.apiKey accepts the same three forms and resolves secret:// the
same way, on every path that sends an embedding request: akm index
(including the reindex akm bundle update runs and the targeted re-embed a
write command like akm remember triggers), akm improve's
consolidate pass (memory dedup and similarity clustering). All of them build the
provider request through the same RemoteEmbedder/resolveSecret boundary,
so a secret:// reference resolves identically regardless of which command
triggered the request (#953).
Use AKM_SQLITE_JOURNAL_MODE=DELETE or TRUNCATE when WAL is unavailable,
such as on some NFS/SMB mounts. With the default WAL setting, AKM detects a
network filesystem for the data directory and falls back to DELETE.
Retired Configuration
profiles, llm, agent, features, stashes, defaults.llm,
defaults.agent, and defaults.improve are rejected in 0.9. Recreate the
configuration using engines, defaults.engine, defaults.llmEngine, and
improve.strategies; AKM deliberately does not infer or rename ambiguous
profile identities.
embedding.chunkSize was never read by anything under src/ (#954), so a
config that still sets it is simply ignored — it still loads, unvalidated
and without warning.
index.graph.* and every strategy's processes.graphExtraction.* are retired
in 0.9.17-alpha.9: the LLM entity graph they configured is gone —
akm show's links come from declared links instead (see ## Strategies
above). A config that still sets them loads; each key is named once as
unknown, and akm migrate apply removes it. The built-in graph-refresh
strategy is retired too, but not the same way as an ordinary unknown name:
naming it via --strategy or a task always fails with a message pointing at
the retirement, even when improve.strategies["graph-refresh"] still has a
leftover override from customizing the built-in (the message names it;
akm migrate apply drops it — a leftover override is never resolved as a new
custom strategy, which would silently run a full, unplanned improve pass).
defaults.improveStrategy: "graph-refresh" still loads config successfully;
the refusal happens lazily, when the strategy is actually resolved.
improve.strategies.<name>.processes.consolidate.incrementalSince and
.neighborsPerChanged are retired in 0.9.17-alpha.9: the consolidate pair
pass is now the candidate generator, narrowing per initiator through the
improve ledger rather than a global time window. A config that still sets
either key loads; each is named once as unknown, and akm migrate apply
removes it.