akm docs

Benchmark findings: what the first real A/B says about tuning akm

Status: live findings, updated as runs land. Source runs: akm-bench branch claude/harbor-akm-agent-p0, results/harbor/2026-08-24/. Companions: benchmark-harness-consolidation.md (the plan), benchmark-harness-decisions.md (decision register).

1. The headline, and what it is not

19-task eval slice, baseline (opencode) vs akm-static (akm-opencode), k=3, 114 trials, opencode/qwen3.5-plus, zero errors and zero timeouts.

metric akm arm baseline
pass@1 75.4% [0.579, 0.895] 52.6% [0.316, 0.737]
paired delta +0.228 [0.070, 0.404] —
akm tool engagement 28.1% 0.0%

The CI excludes zero. Split by what each task family is built to measure:

family tasks baseline akm delta
retrieval (drillbit--*, inkwell--*) 12 0.33 0.67 +0.33
compliance (workflow-compliance--*) 4 0.75 0.83 +0.08
other (opencode--*) 3 1.00 1.00 0.00

drillbit and inkwell are fictional tools whose syntax exists ONLY in the per-task stash, so the baseline must invent it. The compliance family tests the opposite reflex (a noisy stash of plausible-but-wrong material, or a minimal stash with no answer at all) where NOT reaching for akm is correct, so it dilutes the aggregate by design.

What this is not: one model, n=19 tasks, k=3. The CI is ±0.17 wide. It is evidence that akm helps where retrieval is the only path, not a general capability claim.

2. The biggest finding: akm is invisible on "edit an existing file"

Engagement is not uniformly low — it is bimodal by task shape. Counting every akm-arm trial in the eval slice, split by whether the task's workspace ships a pre-existing artifact to edit:

task shape trials engaging akm trials rate
create-new (workspace has no artifact) 16 33 48%
edit-existing (workspace ships service.yaml) 0 24 0%

Zero out of twenty-four. Not one akm call across every edit-shaped task.

Replicated on the disjoint train slice (27 tasks, k=3, 162 trials, zero errors). Same model, different task instances, same direction:

slice shape tasks engagement baseline -> akm
eval create-new 11 48% 0.45 -> 0.88 (+0.42)
train create-new 8 38% 0.58 -> 0.71 (+0.12)
eval edit-existing 8 0% 0.62 -> 0.58 (-0.04)
train edit-existing 19 5% 0.77 -> 0.81 (+0.04)

The train slice's aggregate paired delta is +0.062 [-0.012, 0.148] — the CI includes zero. That is not a contradiction of eval's +0.228: train is 19 of 27 edit-shaped, i.e. dominated by the shape where akm never fires.

Consequence for the tuning protocol: train is a poor surface for tuning the DELTA (its aggregate is diluted by inert tasks and its CI spans zero) but a good surface for tuning ENGAGEMENT, which is the mechanism actually under change. Tune against engagement on train; spend an eval measurement only to confirm the delta moved. The 5% on train (vs 0% on eval) also shows nothing structurally PREVENTS engagement on edit tasks — the model simply rarely thinks to.

The plugin's own framing predicts this. AKM_HINTS_PREFIX (akm-plugins/opencode/index.ts) opens with:

"Before writing anything from scratch, call akm_curate ..."

Editing an existing file is not "from scratch". The trigger condition, read literally, excludes the exact case where engagement collapses.

The failure is legible in the trajectories. On inkwell--configure-scaling (0.00 on both arms) the akm arm did:

read  /app/service.yaml
edit  /app/service.yaml   (invented a `scaling:` block)
text  "Done. Added autoscaling configuration with min: 2, max: 20, ..."

The task spells out every value (min: 2, max: 20, metric: rps, target: 100); the only unknown is the key name and nesting, which lives in the stash. A visible file makes the task look self-sufficient, so the model never suspects there is a convention to look up. The model does not know what it does not know, and nothing in the current framing tells it that an unfamiliar format is as much a lookup trigger as an unfamiliar tool.

Three of the six edit-shaped tasks scored 0.00 on both arms (inkwell--configure-scaling, --cpu-scaling, --workflow-configure-scaling) — i.e. retrieval would plausibly have flipped them. The other three were guessable and passed on both arms.

2b. CORRECTION: shape was confounded with tool familiarity

The §2 finding ("akm is invisible on edit-shaped tasks") is half right, and the half that is wrong matters. Shape is real, but it was measured on a corpus where every edit-shaped task also happened to involve a tool the model already knows. Pooling both post-fix runs and cross-tabbing (workflow-compliance excluded — it tests the opposite reflex):

create-shaped edit-shaped
fictional tools (drillbit, inkwell — syntax exists ONLY in the stash) 96% (27/28) 20% (4/20)
real/known tools (az-cli, docker-homelab) 29% (5/17) 0% (0/35)

Shape survives as an effect (96% vs 20% within fictional tools), but 35 of the trials dragging the old "edit-shaped" aggregate down were az-cli and docker-homelab, where declining to consult a stash is correct — the model knows docker compose. Those were never a miss.

The sharper, and more useful, statement:

akm is under-consulted specifically when editing an existing file for a tool the model does not actually know. That is the 20% cell, and it is the only cell where non-engagement is a genuine miss.

This is exactly the confound akm-bench#6 called out, and it only became visible because that issue's fix added an edit-shaped drillbit task — which engaged 3/3. Without it the effect would still read as pure shape.

Per-family edit-shaped engagement on the post-fix train slice, for the record:

drillbit             3/3  = 100%   (the new decoupling task)
inkwell              1/3  =  33%
workflow-compliance  3/14 =  21%
az-cli               0/17 =   0%
docker-homelab       0/15 =   0%

2c. Engagement is the mechanism, not a proxy metric

Pooling both post-fix runs (jobs/akm-corpus-eval-ab and jobs/akm-corpus-train-ab-postfix, 48 tasks, errored trials excluded) and splitting tasks by whether the akm arm EVER called an akm_* tool, then pairing each group against the baseline arm on the same tasks:

group tasks mean paired delta (akm - baseline)
tasks where akm was NEVER called 29 -0.011
tasks where akm WAS called 19 +0.561

Trial-level, for the record: akm-arm trials WITH an akm_* call mean reward 0.857 (n=49); WITHOUT, 0.800 (n=90).

Injected context alone contributes nothing. Where the model does not call a tool, the treatment arm is statistically indistinguishable from baseline (-0.011). The hints block, the curated-file pointer and the rest of the injected context net to zero by themselves.

Therefore engagement is the mechanism, not a proxy metric: the entire measured +0.246 eval delta is carried by trials that actually invoked akm_*.

Caveat that has to be stated. The +0.561 is confounded by task selection — the model engages precisely on the tasks where retrieval is needed and the baseline fails, so it is NOT a clean causal estimate of what a call is worth. The -0.011 is the clean half: same tasks, both arms, paired, and the treatment's extra context changes nothing.

The within-task comparison does not settle the other half either. Across the 6 tasks that had both engaged and non-engaged akm-arm trials, the mean engaged-minus-non-engaged delta is +0.250, but it splits 2 higher / 1 lower / 3 tied — not decisive on its own.

Consequence for prioritisation. A fix for akm-plugins#99 must produce actual TOOL CALLS; more or better injected context is already known to be worth ~0. And akm#819 (retrieval ceiling) is a multiplier that only pays off once engagement rises — the calls that do happen already convert well, so retrieval quality is not the current binding constraint.

2d. The memory half: akm cannot retrieve conversational content at all

Two independent memory packs, three arms each, smoke scale. baseline is a LONG-CONTEXT REFERENCE ARM — it receives the whole conversation and never retrieves — so it is a ceiling, not a control. The only fair comparison is akm-memory vs raw-vector, two retrieval arms on equal footing.

LongMemEval (5 questions, qwen3.5-plus):

variant judgedPass R@k zero-hit
baseline (long-context) 1.00 — —
raw-vector 0.20 0.20 0/5
akm-memory 0.00 0 5/5 (100%)

LoCoMo (5 questions, same model):

variant tokenF1 zero-hit
baseline (long-context, 16k) 0.567 —
raw-vector 0.233 0/5
akm-memory 0.200 2/5 (40%)

akm scored 0 on LongMemEval because it returned no documents at all — on every question. The run emits this itself, unprompted:

5/5 retrieval queries returned zero hits (>=50%). The aggregate score for this run is dominated by prompts with no retrieved context at all, not by answer quality on retrieved context.

This is a measurement of the RETRIEVAL CEILING (akm#819), not of memory quality, and must not be published as the latter. On LoCoMo, where akm did retrieve, the answer was correct.

The mechanism is structural, not a tuning problem. akm indexes synthesized frontmatter — name, description, tags, searchHints, headings — and never body prose. A LongMemEval haystack is a chat transcript whose answer lives entirely in the body, so it is unsearchable by construction. LoCoMo reaches 40% rather than 100% only because some of its questions happen to match a synthesized field.

The sharpest contrast: raw-vector is a naive in-memory cosine store and had 0/5 zero-hits on both packs. It always returns something. The gap is not ranking quality — it is COVERAGE.

What this means for the project's priorities. The +0.439 benchmark result came from a corpus of DOCUMENTED CONVENTIONS — content that lives in frontmatter, exactly where akm looks. Conversational recall is the opposite shape. So the benchmark win does not transfer to memory, and akm#819 is not an optimisation: it is the precondition for akm working as memory at all rather than as a convention library.

Caveats. n=5 questions, one sample, one model per pack — these numbers separate the arms on COVERAGE, which is a structural property, but they are far too small to rank answer quality. Also note LongMemEval reports exactMatch/tokenF1/containsExpected as a hardcoded 0 because the pack scores on the official LLM judge (akm-eval#6); only judgedPass carries signal there.

2e. RESOLVED: akm#819 lifted the ceiling — measured on 0.9.2-alpha.2

§2d's diagnosis was correct and the fix confirms it. akm-cli 0.9.2-alpha.2 (npm next) carries #819's retrieval fix. Re-measuring with a RETRIEVAL-ONLY probe — each pack adapter's exact ingest and query, no LLM in the loop, so nothing here is confounded by model or judge — on identical corpora, identical code paths, only the CLI version differing:

pack metric 0.9.1 0.9.2-alpha.2
LoCoMo (conv-26, 419 docs, 40 questions) zero-hit rate 75.0% 0.0%
evidence recall@5 0.154 0.590
LongMemEval (20 questions, ~61 sessions each) zero-hit rate 100% 0.0%
evidence recall@5 0.000 1.000

Zero-hit going to 0% on its own would be ambiguous — a retriever that returns five arbitrary documents for every query also scores 0% zero-hit. Evidence recall is what rules that out: at topK=5 over a 419-document haystack, chance recall is ~1%. 0.590 is retrieval, not noise.

The isolated unit-level proof. akm-eval's live-CLI integration test had encoded the ceiling as an assertion — a body-only term MUST return zero hits. Against 0.9.2 it returns exactly one hit, the correct document. The assertion has been inverted to guard the fix instead (akm-eval#9); it was, until then, a tripwire that fired on an improvement.

What this does and does not settle. It settles COVERAGE, which is what §2d identified as the gap. It does not yet re-rank answer quality: the end-to-end judged numbers (LongMemEval judgedPass, LoCoMo tokenF1) still need a rerun against 0.9.2 before the memory-half result can be restated. The published 0.00 / 0.200 remain the last measured end-to-end values and remain floored by 0.9.1 retrieval — they should be treated as superseded-pending- rerun, not as current.

A defect surfaced by the fix — since fixed and released. Under 0.9.2-alpha.2, 4 of 20 LongMemEval questions ABORTED on akm-eval's contamination guard. Root cause was a latent akm bug that 0.9.1's near-total zero-hit rate had hidden: a document under memories/ whose body contains an ordinary dollar amount ($1,200, $2,000) matched COMMAND_PLACEHOLDER_RE = /\$ARGUMENTS|\$[123]\b/ — \b sits between the 2 and the comma — and was indexed as a command, its ref becoming commands/memories/<slug>. That content heuristic (specificity 18) outranked the explicit parent-directory classification (15), so body sniffing beat the directory the file was sitting in. Clean separation on the corpus: 3 of 51 documents affected, and those 3 exactly the 3 matching the regex. Present in 0.9.1 identically — newly VISIBLE, not new. Fixed in akm#824 / #826 and shipped in 0.9.2-alpha.3, drawn along ambiguity rather than directories: $ARGUMENTS keeps its precedence over a directory hint, while the numeric placeholders exclude currency and defer to a directory that declares a type.

Verified against the published 0.9.2-alpha.3 artifact:

pack metric 0.9.1 0.9.2-alpha.2 0.9.2-alpha.3
LoCoMo zero-hit 75.0% 0.0% 0.0%
evidence recall@5 0.154 0.590 0.590
LongMemEval zero-hit 100% 0.0% 0.0%
evidence recall@5 0.000 1.000 (of 16) 0.950 (of 20)
questions aborted — 4 of 20 0 of 20

LongMemEval recall@5 reads LOWER on alpha.3 only because all 20 questions now complete: alpha.2's 1.000 was over the 16 that did not abort. The alpha.3 number is the honest one.

Ordered by expected value. Every one changes the TREATMENT (the product), not the measurement — see §4 for why that distinction is what keeps the number honest.

akm-plugins (highest value, smallest change)

  1. Rewrite the AKM_HINTS_PREFIX trigger so it covers editing. "Before writing anything from scratch" is the single highest-leverage string in the plugin, and it currently gates out 24 of 57 trials. It should name the case that actually fails: writing or editing a config file, manifest, or command for a tool whose exact syntax you are not certain of. Cheap to change, directly falsifiable on the train slice.

  2. Make the akm_curate description compete with the built-ins. On edit-shaped tasks the model reaches for read/glob/edit because those obviously act on the visible file. akm_curate's description ("PRIMARY discovery entry point for the stash ... describe the task in natural language") reads as project-discovery, not as "check this format's conventions before you write it". Paid models went further and preferred opencode's built-in skill tool.

  3. Consider a targeted nudge when an unfamiliar format is in play. The plugin already runs experimental.chat.system.transform on every request and already tracks a curated file per session. A line that names the concrete asset types available for the current stash would convert "there is a stash" into "there is a documented schema for the file you are about to edit".

akm (CLI / retrieval)

  1. Query shaping is a known ceiling. akm's FTS is a conjunctive AND with no stopword removal, so sentence-shaped queries zero-hit while keyword queries rank 1. Anything that makes the plugin's generated query keyword- shaped (or that makes akm tolerant of sentence-shaped input) raises the payoff of every engagement that does happen.

  2. Body prose is never indexed — only name/description/tags/hints/ headings. Question-form searchHints recovered most of this in the treatment library (0/8 -> 7/8). Worth making that authoring rule explicit in the docs, since it decides retrieval outcomes more than ranking does.

akm-bench (corpus quality, not tuning)

  1. drillbit--backup-policy's verifier is too lenient. Both arms wrote drillbit backup configure --cluster ... against a gold of drillbit backup --cluster ... and both scored 1.0. A task that accepts an invented form cannot discriminate; it under-reports akm's effect whenever a model guesses close.

  2. The edit-shaped inkwell tasks are the corpus's most valuable assets — they are the only ones that isolate the blind spot in §2. Keep them, and consider adding edit-shaped drillbit equivalents so the shape is not confounded with the tool family.

4. How to tune without invalidating the benchmark

The rule: change the treatment, never the measurement.

Legitimate (ships to real users): tool descriptions, injected framing, retrieval quality, indexing, query shaping. A model that consults akm more because the tools are better described is a product improvement.

Invalidating: putting akm instructions in task prompts (contaminates the baseline or becomes a treatment-only confound), touching verifiers or timeouts, giving the treatment arm anything unrelated to akm.

The real risk is iterating against the eval slice. Tune until eval improves and the plugin is fitted to those 19 tasks; the number keeps looking rigorous while it stops generalizing. The corpus already solves this:

akm-tasks-train  27 tasks
akm-tasks-eval   19 tasks
overlap: NONE          (same families, disjoint instances)

Protocol:

  1. Iterate on train (harbor/jobs/.corpus-train-ab-local.yaml), freely.
  2. Decide the change and predict its effect before touching eval.
  3. Measure on eval rarely — ideally once per meaningful plugin change. Today's +0.228 [0.070, 0.404] at 28.1% engagement is the pre-registered baseline. SUPERSEDED as of akm-bench 0578025 — see the note below.

The eval baseline was invalidated on purpose, once. Fixing akm-bench#6 tightened the verifiers of 11 of the 19 eval-slice tasks (5 drillbit + 6 inkwell). The task LIST is unchanged, but the scoring is not, so the +0.228 [0.070, 0.404] figure is no longer comparable to anything measured after that commit. That was the right trade — a correct verifier beats a comparable-but-wrong one, and the old one scored an invented drillbit backup configure ... as a pass — but it means the NEXT eval run re-establishes the baseline rather than testing against it. Since both arms are always re-run together, the new run is internally valid on its own; only the cross-run comparison is lost.

  1. Always re-run both arms together. The baseline moves too when opencode, the corpus, or the harness changes; comparing a new treatment against a stored baseline silently confounds.
  2. At n=19/k=3 the CI is ±0.17. Small engagement gains will not separate from noise on eval — another reason to do the tuning where looking is free, and to raise k when you need to resolve a smaller effect.

5. Environment notes that affect any rerun