akm 0.9.15 — release plan
Status: in progress 2026-09-09. Branch: release/0.9.15 (work lands via
claude/0-9-15-milestone-h9huwi). Scope: the sixteen issues on the 0.9.15
milestone (#942–#957), all filed from a 2026-09-08/09 review of five akm
instances. Implementation is done by Sonnet agents, one batch per git
worktree; every batch diff is reviewed by the coordinator before it reaches
the release branch. Nothing is pushed unreviewed.
1. What the review found, from the top
Every issue in the milestone is one of three shapes:
- akm did the wrong thing silently. A credential form the schema accepts crashes engine dispatch (#953); an improve run with every LLM process skipped exits 0 and looks normal (#957); state.db contention surfaces as a raw driver string and exit 70 (#948); two scopes each believe a workflow ref has no active run (#942); an index rebuild that dies mid-way starts from zero because the fingerprint is only written on success (#955/#956).
- akm already had the data and threw it away. Per-process engine/model routing is computed on every run and never shown (#947); per-call LLM usage is persisted and never cross-tabbed per run (#944); health knows which process an engine is bound to but not whether it was ever used (#950).
- A performance or ergonomics gap with an existing precedent to copy.
Embeddings are batched per request but committed once at the very end of
the whole pass (#954); the thinking switch is lowered by provider name
(#949);
task list,show env, scheduler log growth, path discovery (#951); reflect treats feedback as fact and leaks its truncation marker (#952); a timed-out command task should be a counted failure (#943); fleet config is copied by hand (#945/#946).
The default is subtraction and reuse: every batch brief names the existing helper to extend, and every guard added has to pass AGENTS.md's three tests.
2. Decisions the coordinator made (and why)
| # | Decision | Rationale |
|---|---|---|
| 954 | Commit per provider batch; a fixed in-flight window (1 loopback / 2 otherwise, no knob — embedding.batchSize and the token budget are the tuning surface); split a batch on a context-size rejection; throughput line. Defaults unchanged. |
Request batching already shipped (#874). The single end-of-run transaction is the real defect. Lab measurement: a 32-input batch costs one input's wall time; 2 in flight saturates a 4-slot server. |
| 955 | Write the new fingerprint at purge time (resume, not re-purge); store the vector identity the server reports (response model id + observed dimension) and canary-verify (median cosine of 8 samples ≥ 0.999) before purging; akm index --reembed. No /models hashing, no indexKey. Field-review addendum: a sample whose canary re-embed failed is excluded from the median rather than scored as zero — a partial provider failure is not evidence of a different model — and when half or fewer of the sampled entries verify, the outcome is unverifiable rather than trusting a too-thin sample either way. |
Five stashes re-embedded identical vectors after a transport change; the config string cannot tell that apart, the server's model id can. |
| 956 | PID-liveness rebuild lock, never blocking: akm index --skip-if-locked skips; without the flag warn and proceed; a dead holder is reclaimed silently. Interactive writers never wait on a rebuild: the write-path index upsert (remember, import, proposal accept, source clone, extract session assets) skips as success, not a wait, whenever the rebuild lock is held. Field-review correction: the planned 30 s/10-minute interactive-vs-batch lease-wait split is dropped — remember/import write the asset file directly and never took the (pre-existing, unrelated) synchronous asset-mutation lease's wait path in the first place, so the premise for a shorter interactive bound didn't hold. That pre-existing lease (withAssetMutationLeaseSync) keeps its unchanged 10-minute bound but now logs a wait notice every 15 s naming the holder, instead of giving no progress feedback at all. |
#872 removed the blocking lease for good reasons; a 24k-entry rebuild was restarted from zero by four legitimate neighbours and remember blocked for over ten minutes. |
| 948 | One SQLite-contention classifier; exhaustion → STATE_DB_CONTENDED, a new TransientError kind exiting 75 (sysexits EX_TEMPFAIL); RUN_LEASE_HELD moves to the same kind; akm workflow run --skip-if-locked. No per-run "waiting for" bookkeeping. |
Schedulers treat exit 2 as "fix the command line"; contention is "try again later". SQLite cannot name its writer. |
| 953 | Full secret:// parity with apiKeyFile (#905) across resolution, runner specs, redaction, health; warn once when $VAR resolves empty. |
0.9.13 already announced the feature; a partial rollout caused this issue. |
| 957 | Credential-aware improve plan; skippedProcesses on the result; ok/exit unchanged (#912 precedent); health active-improve-strategy fails only when the run would be a no-op; opt-in --require-engines that aborts before any index work and names the unresolved reference (never a value), added to the shipped scheduled task templates. |
Three scheduled scopes would otherwise have skipped every LLM process, spent the budget on an index pass and exited 0. |
| 947 | Extend --dry-run's plan with per-process routing (--plan is a pure alias); no network probe in dry-run. |
Dry-run stays offline and fast; health already probes reachability. |
| 944 | usageReport (process×engine×model cross-tab + no-call reasons) on every run, shared projection helper, akm improve report [--run <id>|--since <window>] via scope interception. |
citty cannot take a subcommand under improve without breaking akm improve <scope>. |
| 949 | Send both chat_template_kwargs.enable_thinking and enable_thinking only when enableThinking resolves to a value (from the engine's own config or a calling process's per-call override), nothing otherwise; reasoning_effort sent whenever it is set on the engine; delete the provider branch; passive thinking-control health advisory from usage telemetry. No thinking/thinkingWire fields. |
The tri-state already exists; a strict hosted API rejects unknown keys, so the conditional matters more than the wire; gateways drop chat_template_kwargs but pass reasoning_effort; health must not wake a cold model. |
| 950 | cli-version advisory (GitHub releases via checkForUpdate, --probe-gated), env-asset credential hint (never names the variable), engine-last-used advisory (30-day window, quiet on fresh installs). Fleet-pinned version deferred until a field exists. |
Reuses akm upgrade's source of truth and the plugin-staleness shape. |
| 942 | Keep per-scope partitioning of the active-run guard; warn on cross-scope collisions naming id and scope; --all-scopes on list/status; envelope names the scope searched; errors name id and scope. |
Per-project runs of one workflow are a documented design; visibility fixes the incident without changing the contract in a patch. |
| 943 | Pin timeout → failed for command tasks (test first); health task-fail-rate gains a reason breakdown and a "timeout-dominant" suffix. Keep the rich AgentFailureReason vocabulary. |
Static reading finds the propagation already correct; the missing piece is data in health. |
| 951 | Do: show env key names in text, cron log truncates per run, akm info exposes data/config/cache/state dirs, akm task list as a pure alias, remove the false --rerank docs. Defer: a rerank engine kind (own issue). Out of scope: the operator's Discord script. |
An alias reintroduces no logic; rerank is a feature, not a consistency nit. |
| 952 | Feedback framed as a signal; truncation marker detected at creation (deferred to review) and rejected at promote; context-aware content budget on the LLM path only. | Both failures reproduced on two quants; a truncated body replacing a full asset is data loss, which passes the defensive-code carve-out. The argv cap stays where argv exists. |
| 945 | extends: <path|bundle//conceptId> (no URL), chain with cycle detection, akm config diff <other>, opt-in config get --show-source. |
Config load is synchronous and offline; Stable output shapes stay unchanged. |
| 946 | { "engine": "<name>" } as a per-column indirection in models.json; akm models list. Engine selection is unchanged. |
Respects the engine-selection vs model-mapping separation in the approved design. |
3. Batches
| Batch | Worktree branch | Issues (in order) | Why together |
|---|---|---|---|
| A index | wt/0.9.15/index |
954 → 955 → 956 | Same materializer/embedder files; resumability builds on per-batch commits. |
| B improve | wt/0.9.15/improve |
953 → 957 → 947 → 944 | Credential availability feeds the plan; routing projection feeds the report. |
| C engines-health | wt/0.9.15/engines-health |
949 → 950 | Both add health advisories over usage telemetry. |
| D state-workflow | wt/0.9.15/state-workflow |
948 → 942 | state.db and workflow-run repositories. |
| E tasks-misc | wt/0.9.15/tasks-misc |
943 → 951 | Task runner, health task metrics, small CLI fixes. |
| F config-models | wt/0.9.15/config-models |
945 → 946 | Config load and model-map resolution. |
| G reflect | wt/0.9.15/reflect |
952 | Prompt assets and reflect sanitiser. |
Priority for merging: B, A, D, E first (correctness and reliability), then C, G, F.
4. Working agreement
- One Sonnet agent per issue, in the batch's worktree, with the issue text, the reader's code map and the coordinator's brief. Tests first. Commits are frequent and small. No agent pushes.
- Each issue diff is reviewed adversarially by a second Sonnet agent (defects, complexity, SOLID/DRY, tests that a null implementation would pass, docs); confirmed findings go back for a fix, at most two rounds.
- Each batch ends with a gate: biome, tsc,
bun run lint,bun run test:unit,bun run test:integrationin the worktree. - The coordinator reads every batch diff before it is merged into the release
branch, resolves conflicts between batches, and pushes. CHANGELOG entries
are written as fragments under
docs/plans/0.9.15/changelog/and assembled at closeout, then the fragment directory is deleted. - Closeout: CHANGELOG section,
docs/migration/release-notes/0.9.15.md(+ README index line),package.jsonbump,./tests/release-check.sh --skip-docker, push.
5. Outcome (2026-09-09)
- #954 — the originally proposed configurable
embedding.concurrencywas superseded by the dev-team field review before it shipped: the in-flight window shipped fixed at 1 request (loopback) / 2 requests (remote), not a config knob, with request size —embedding.batchSizeandembedding.maxTokens/contextLength— as the throughput lever instead. The 0.9.15-beta.1 field test then surfaced two further problems: the per-batch-commit claim did not hold forakm bundle update(its coordinator ran the embedding phase inside its own outer transaction, so per-batch commits nested as unobservable SAVEPOINTs), and a multi-slot local server (llama.cpp--parallel N, vLLM) was left idle by the fixed concurrency. #954 (before the stable 0.9.15) fixes both: a drift guard makes the ambient-transaction case a loud failure instead of a silent one,akm bundle updatenow runs the shared embedding pass on its own connection after its own commit, andembedding.concurrency(bounded 1-16) reverses the "no knob" call above, opt-in only for a server that genuinely serves parallel slots. #954 also makesembedding.timeoutMsconfigurable (default 120s, replacing a fixed 30s that cut off a slow local batch), adds a 3-consecutive-failure circuit breaker so a dead endpoint fails fast instead of grinding for hours, and surfaces provider/progress failures at the default log level instead of--verbose-only. - #956 — the originally planned 30 s/10-minute interactive-vs-batch lease
split was dropped:
remember/importnever take the asset-mutation lease in the first place (they write the asset file directly), so the premise for a shorter interactive-writer bound was wrong. The pre-existing sync lease wait (withAssetMutationLeaseSync) keeps its unchanged 10-minute bound but now prints a notice every 15 s naming the holder. Separately, the write-path index upsert skips — and reports as success — rather than waiting, whenever the new rebuild lock is held during a live rebuild. - #957 — a workflow agent working #947 hit a pre-existing #957
regression (
--dry-run's all-disabled guard threw when every process was credential-unavailable, even though a dry run never dispatches), filed GitHub issue #958 against it, and left the characterization test.skip-ped per AGENTS.md's explicit-skip-with-linked-issue rule so the #947 batch could land green. The coordinator reversed that deferral: the regression is fixed in this release (--dry-run/--plannever abort on an unavailable credential; the preview reads the resolved engine/model/ context-length metadata a credential-unavailable process still carries viaEngineUnavailableProcess'sengine/model/contextLengthfields), the characterization test is restored with updated assertions, and #958 can be closed by its owner as fixed. - #948 —
RUN_LEASE_HELDand the newSTATE_DB_CONTENDEDboth exit 75 (EX_TEMPFAIL), not 2, per the field review — schedulers read exit 2 as "fix the command line," not "retry shortly." - #949 — both thinking-control wire forms (
chat_template_kwargs .enable_thinkingand bareenable_thinking) are sent only whenenableThinkingresolves to a value, whether from the engine's own config or a calling process's per-call override — not gated onproviderany more.reasoning_effortis sent whenever it is set on the engine, unconditionally. - #955 — vector identity is keyed on the endpoint's reported response
model plus the observed vector width (not a config string or
/modelshash), and the fingerprint-rename canary re-embeds a cosine-similarity sample before any rebuild. A field-review addendum refined the verdict rule: a sample whose canary re-embed failed is excluded from the median rather than scored as zero, and when half or fewer of the sampled entries verify, the outcome isunverifiablerather than trusting a too-thin sample either way. - #955 — pulled forward from a later milestone on
2026-09-09, alongside #954 (both land via the same "embed" worktree/
batch, #954 first). Embedding salvage across full rebuilds and index-
generation bumps: vectors about to be discarded wholesale are copied into
a transient
embedding_salvagetable (keyed bysha256(search_text)plus the fingerprint they were generated under) in the same transaction as the discard, and the next embedding pass hands them back to unchanged entries before any provider call — so the 0.9.14 v22→v23 bump's multi-hour re-embed does not recur onakm index --fullor a future generation bump.--reembedand a canary "rebuild" verdict still purge everything; a canary "keep" verdict relabels leftover salvage to the new fingerprint string instead of discarding it.
Batches (all merged)
| Batch | Issues | Merge commit |
|---|---|---|
| A index | 954 → 955 → 956 | e828a95 merge: 0.9.15 index batch (#954, #955, #956) |
| B improve | 953 → 957 → 947 → 944 | 60e4054 merge: 0.9.15 improve batch (#944, #947, #953, #957) |
| C engines-health | 949 → 950 | 711be77 merge: 0.9.15 engines-health batch (#949, #950) |
| D state-workflow | 948 → 942 | a997503 merge: 0.9.15 state-workflow batch (#942, #948) |
| E tasks-misc | 943 → 951 | d90c1d4 merge: 0.9.15 tasks-misc batch (#943, #951) |
| F config-models | 945 → 946 | 8ed13b5 merge: 0.9.15 config-models batch (#945, #946) |
| G reflect | 952 | e68c4c1 merge: 0.9.15 reflect batch (#952) |
Deferred / open
- #951 — a
rerankengine kind is deferred to its own issue; this release only removed the false--rerankdocumentation claims. - #952 — the issue's confidence-gate claim does not match the code that shipped; documented in the reflect batch's coordinator brief rather than re-litigated here.
- #943 — the issue's original premise (timeout/killed command tasks not
counted as failures) was not reproduced; static reading found the
underlying failure propagation was already correct. The fix that shipped
is the
task-fail-ratehealth evidence/keys line (agentFailureReasonCountsand the "timeout-dominant" suffix), pinning that contract with a regression test instead of leaving it implicit.