akm docs

akm 0.9.15 — release plan

Status: in progress 2026-09-09. Branch: release/0.9.15 (work lands via claude/0-9-15-milestone-h9huwi). Scope: the sixteen issues on the 0.9.15 milestone (#942–#957), all filed from a 2026-09-08/09 review of five akm instances. Implementation is done by Sonnet agents, one batch per git worktree; every batch diff is reviewed by the coordinator before it reaches the release branch. Nothing is pushed unreviewed.

1. What the review found, from the top

Every issue in the milestone is one of three shapes:

  1. akm did the wrong thing silently. A credential form the schema accepts crashes engine dispatch (#953); an improve run with every LLM process skipped exits 0 and looks normal (#957); state.db contention surfaces as a raw driver string and exit 70 (#948); two scopes each believe a workflow ref has no active run (#942); an index rebuild that dies mid-way starts from zero because the fingerprint is only written on success (#955/#956).
  2. akm already had the data and threw it away. Per-process engine/model routing is computed on every run and never shown (#947); per-call LLM usage is persisted and never cross-tabbed per run (#944); health knows which process an engine is bound to but not whether it was ever used (#950).
  3. A performance or ergonomics gap with an existing precedent to copy. Embeddings are batched per request but committed once at the very end of the whole pass (#954); the thinking switch is lowered by provider name (#949); task list, show env, scheduler log growth, path discovery (#951); reflect treats feedback as fact and leaks its truncation marker (#952); a timed-out command task should be a counted failure (#943); fleet config is copied by hand (#945/#946).

The default is subtraction and reuse: every batch brief names the existing helper to extend, and every guard added has to pass AGENTS.md's three tests.

2. Decisions the coordinator made (and why)

# Decision Rationale
954 Commit per provider batch; a fixed in-flight window (1 loopback / 2 otherwise, no knob — embedding.batchSize and the token budget are the tuning surface); split a batch on a context-size rejection; throughput line. Defaults unchanged. Request batching already shipped (#874). The single end-of-run transaction is the real defect. Lab measurement: a 32-input batch costs one input's wall time; 2 in flight saturates a 4-slot server.
955 Write the new fingerprint at purge time (resume, not re-purge); store the vector identity the server reports (response model id + observed dimension) and canary-verify (median cosine of 8 samples ≥ 0.999) before purging; akm index --reembed. No /models hashing, no indexKey. Field-review addendum: a sample whose canary re-embed failed is excluded from the median rather than scored as zero — a partial provider failure is not evidence of a different model — and when half or fewer of the sampled entries verify, the outcome is unverifiable rather than trusting a too-thin sample either way. Five stashes re-embedded identical vectors after a transport change; the config string cannot tell that apart, the server's model id can.
956 PID-liveness rebuild lock, never blocking: akm index --skip-if-locked skips; without the flag warn and proceed; a dead holder is reclaimed silently. Interactive writers never wait on a rebuild: the write-path index upsert (remember, import, proposal accept, source clone, extract session assets) skips as success, not a wait, whenever the rebuild lock is held. Field-review correction: the planned 30 s/10-minute interactive-vs-batch lease-wait split is dropped — remember/import write the asset file directly and never took the (pre-existing, unrelated) synchronous asset-mutation lease's wait path in the first place, so the premise for a shorter interactive bound didn't hold. That pre-existing lease (withAssetMutationLeaseSync) keeps its unchanged 10-minute bound but now logs a wait notice every 15 s naming the holder, instead of giving no progress feedback at all. #872 removed the blocking lease for good reasons; a 24k-entry rebuild was restarted from zero by four legitimate neighbours and remember blocked for over ten minutes.
948 One SQLite-contention classifier; exhaustion → STATE_DB_CONTENDED, a new TransientError kind exiting 75 (sysexits EX_TEMPFAIL); RUN_LEASE_HELD moves to the same kind; akm workflow run --skip-if-locked. No per-run "waiting for" bookkeeping. Schedulers treat exit 2 as "fix the command line"; contention is "try again later". SQLite cannot name its writer.
953 Full secret:// parity with apiKeyFile (#905) across resolution, runner specs, redaction, health; warn once when $VAR resolves empty. 0.9.13 already announced the feature; a partial rollout caused this issue.
957 Credential-aware improve plan; skippedProcesses on the result; ok/exit unchanged (#912 precedent); health active-improve-strategy fails only when the run would be a no-op; opt-in --require-engines that aborts before any index work and names the unresolved reference (never a value), added to the shipped scheduled task templates. Three scheduled scopes would otherwise have skipped every LLM process, spent the budget on an index pass and exited 0.
947 Extend --dry-run's plan with per-process routing (--plan is a pure alias); no network probe in dry-run. Dry-run stays offline and fast; health already probes reachability.
944 usageReport (process×engine×model cross-tab + no-call reasons) on every run, shared projection helper, akm improve report [--run <id>|--since <window>] via scope interception. citty cannot take a subcommand under improve without breaking akm improve <scope>.
949 Send both chat_template_kwargs.enable_thinking and enable_thinking only when enableThinking resolves to a value (from the engine's own config or a calling process's per-call override), nothing otherwise; reasoning_effort sent whenever it is set on the engine; delete the provider branch; passive thinking-control health advisory from usage telemetry. No thinking/thinkingWire fields. The tri-state already exists; a strict hosted API rejects unknown keys, so the conditional matters more than the wire; gateways drop chat_template_kwargs but pass reasoning_effort; health must not wake a cold model.
950 cli-version advisory (GitHub releases via checkForUpdate, --probe-gated), env-asset credential hint (never names the variable), engine-last-used advisory (30-day window, quiet on fresh installs). Fleet-pinned version deferred until a field exists. Reuses akm upgrade's source of truth and the plugin-staleness shape.
942 Keep per-scope partitioning of the active-run guard; warn on cross-scope collisions naming id and scope; --all-scopes on list/status; envelope names the scope searched; errors name id and scope. Per-project runs of one workflow are a documented design; visibility fixes the incident without changing the contract in a patch.
943 Pin timeout → failed for command tasks (test first); health task-fail-rate gains a reason breakdown and a "timeout-dominant" suffix. Keep the rich AgentFailureReason vocabulary. Static reading finds the propagation already correct; the missing piece is data in health.
951 Do: show env key names in text, cron log truncates per run, akm info exposes data/config/cache/state dirs, akm task list as a pure alias, remove the false --rerank docs. Defer: a rerank engine kind (own issue). Out of scope: the operator's Discord script. An alias reintroduces no logic; rerank is a feature, not a consistency nit.
952 Feedback framed as a signal; truncation marker detected at creation (deferred to review) and rejected at promote; context-aware content budget on the LLM path only. Both failures reproduced on two quants; a truncated body replacing a full asset is data loss, which passes the defensive-code carve-out. The argv cap stays where argv exists.
945 extends: <path|bundle//conceptId> (no URL), chain with cycle detection, akm config diff <other>, opt-in config get --show-source. Config load is synchronous and offline; Stable output shapes stay unchanged.
946 { "engine": "<name>" } as a per-column indirection in models.json; akm models list. Engine selection is unchanged. Respects the engine-selection vs model-mapping separation in the approved design.

3. Batches

Batch Worktree branch Issues (in order) Why together
A index wt/0.9.15/index 954 → 955 → 956 Same materializer/embedder files; resumability builds on per-batch commits.
B improve wt/0.9.15/improve 953 → 957 → 947 → 944 Credential availability feeds the plan; routing projection feeds the report.
C engines-health wt/0.9.15/engines-health 949 → 950 Both add health advisories over usage telemetry.
D state-workflow wt/0.9.15/state-workflow 948 → 942 state.db and workflow-run repositories.
E tasks-misc wt/0.9.15/tasks-misc 943 → 951 Task runner, health task metrics, small CLI fixes.
F config-models wt/0.9.15/config-models 945 → 946 Config load and model-map resolution.
G reflect wt/0.9.15/reflect 952 Prompt assets and reflect sanitiser.

Priority for merging: B, A, D, E first (correctness and reliability), then C, G, F.

4. Working agreement

5. Outcome (2026-09-09)

Batches (all merged)

Batch Issues Merge commit
A index 954 → 955 → 956 e828a95 merge: 0.9.15 index batch (#954, #955, #956)
B improve 953 → 957 → 947 → 944 60e4054 merge: 0.9.15 improve batch (#944, #947, #953, #957)
C engines-health 949 → 950 711be77 merge: 0.9.15 engines-health batch (#949, #950)
D state-workflow 948 → 942 a997503 merge: 0.9.15 state-workflow batch (#942, #948)
E tasks-misc 943 → 951 d90c1d4 merge: 0.9.15 tasks-misc batch (#943, #951)
F config-models 945 → 946 8ed13b5 merge: 0.9.15 config-models batch (#945, #946)
G reflect 952 e68c4c1 merge: 0.9.15 reflect batch (#952)

Deferred / open