Benchmark findings: what the first real A/B says about tuning akm
Status: live findings, updated as runs land.
Source runs: akm-bench branch claude/harbor-akm-agent-p0,
results/harbor/2026-08-24/.
Companions: benchmark-harness-consolidation.md (the plan),
benchmark-harness-decisions.md (decision register).
1. The headline, and what it is not
19-task eval slice, baseline (opencode) vs akm-static (akm-opencode),
k=3, 114 trials, opencode/qwen3.5-plus, zero errors and zero timeouts.
| metric | akm arm | baseline |
|---|---|---|
| pass@1 | 75.4% [0.579, 0.895] | 52.6% [0.316, 0.737] |
| paired delta | +0.228 [0.070, 0.404] | — |
| akm tool engagement | 28.1% | 0.0% |
The CI excludes zero. Split by what each task family is built to measure:
| family | tasks | baseline | akm | delta |
|---|---|---|---|---|
retrieval (drillbit--*, inkwell--*) |
12 | 0.33 | 0.67 | +0.33 |
compliance (workflow-compliance--*) |
4 | 0.75 | 0.83 | +0.08 |
other (opencode--*) |
3 | 1.00 | 1.00 | 0.00 |
drillbit and inkwell are fictional tools whose syntax exists ONLY in the
per-task stash, so the baseline must invent it. The compliance family tests
the opposite reflex (a noisy stash of plausible-but-wrong material, or a
minimal stash with no answer at all) where NOT reaching for akm is correct,
so it dilutes the aggregate by design.
What this is not: one model, n=19 tasks, k=3. The CI is ±0.17 wide. It is evidence that akm helps where retrieval is the only path, not a general capability claim.
2. The biggest finding: akm is invisible on "edit an existing file"
Engagement is not uniformly low — it is bimodal by task shape. Counting every akm-arm trial in the eval slice, split by whether the task's workspace ships a pre-existing artifact to edit:
| task shape | trials engaging akm | trials | rate |
|---|---|---|---|
| create-new (workspace has no artifact) | 16 | 33 | 48% |
edit-existing (workspace ships service.yaml) |
0 | 24 | 0% |
Zero out of twenty-four. Not one akm call across every edit-shaped task.
Replicated on the disjoint train slice (27 tasks, k=3, 162 trials, zero errors). Same model, different task instances, same direction:
| slice | shape | tasks | engagement | baseline -> akm |
|---|---|---|---|---|
| eval | create-new | 11 | 48% | 0.45 -> 0.88 (+0.42) |
| train | create-new | 8 | 38% | 0.58 -> 0.71 (+0.12) |
| eval | edit-existing | 8 | 0% | 0.62 -> 0.58 (-0.04) |
| train | edit-existing | 19 | 5% | 0.77 -> 0.81 (+0.04) |
The train slice's aggregate paired delta is +0.062 [-0.012, 0.148] — the CI includes zero. That is not a contradiction of eval's +0.228: train is 19 of 27 edit-shaped, i.e. dominated by the shape where akm never fires.
Consequence for the tuning protocol: train is a poor surface for tuning the DELTA (its aggregate is diluted by inert tasks and its CI spans zero) but a good surface for tuning ENGAGEMENT, which is the mechanism actually under change. Tune against engagement on train; spend an eval measurement only to confirm the delta moved. The 5% on train (vs 0% on eval) also shows nothing structurally PREVENTS engagement on edit tasks — the model simply rarely thinks to.
The plugin's own framing predicts this. AKM_HINTS_PREFIX
(akm-plugins/opencode/index.ts) opens with:
"Before writing anything from scratch, call
akm_curate..."
Editing an existing file is not "from scratch". The trigger condition, read literally, excludes the exact case where engagement collapses.
The failure is legible in the trajectories. On inkwell--configure-scaling
(0.00 on both arms) the akm arm did:
read /app/service.yaml
edit /app/service.yaml (invented a `scaling:` block)
text "Done. Added autoscaling configuration with min: 2, max: 20, ..."
The task spells out every value (min: 2, max: 20, metric: rps,
target: 100); the only unknown is the key name and nesting, which lives in
the stash. A visible file makes the task look self-sufficient, so the model
never suspects there is a convention to look up. The model does not know
what it does not know, and nothing in the current framing tells it that an
unfamiliar format is as much a lookup trigger as an unfamiliar tool.
Three of the six edit-shaped tasks scored 0.00 on both arms
(inkwell--configure-scaling, --cpu-scaling, --workflow-configure-scaling)
— i.e. retrieval would plausibly have flipped them. The other three were
guessable and passed on both arms.
2b. CORRECTION: shape was confounded with tool familiarity
The §2 finding ("akm is invisible on edit-shaped tasks") is half right, and the half that is wrong matters. Shape is real, but it was measured on a corpus where every edit-shaped task also happened to involve a tool the model already knows. Pooling both post-fix runs and cross-tabbing (workflow-compliance excluded — it tests the opposite reflex):
| create-shaped | edit-shaped | |
|---|---|---|
fictional tools (drillbit, inkwell — syntax exists ONLY in the stash) |
96% (27/28) | 20% (4/20) |
real/known tools (az-cli, docker-homelab) |
29% (5/17) | 0% (0/35) |
Shape survives as an effect (96% vs 20% within fictional tools), but 35 of
the trials dragging the old "edit-shaped" aggregate down were az-cli and
docker-homelab, where declining to consult a stash is correct — the model
knows docker compose. Those were never a miss.
The sharper, and more useful, statement:
akm is under-consulted specifically when editing an existing file for a tool the model does not actually know. That is the 20% cell, and it is the only cell where non-engagement is a genuine miss.
This is exactly the confound akm-bench#6 called out, and it only became visible
because that issue's fix added an edit-shaped drillbit task — which engaged
3/3. Without it the effect would still read as pure shape.
Per-family edit-shaped engagement on the post-fix train slice, for the record:
drillbit 3/3 = 100% (the new decoupling task)
inkwell 1/3 = 33%
workflow-compliance 3/14 = 21%
az-cli 0/17 = 0%
docker-homelab 0/15 = 0%
2c. Engagement is the mechanism, not a proxy metric
Pooling both post-fix runs (jobs/akm-corpus-eval-ab and
jobs/akm-corpus-train-ab-postfix, 48 tasks, errored trials excluded) and
splitting tasks by whether the akm arm EVER called an akm_* tool, then
pairing each group against the baseline arm on the same tasks:
| group | tasks | mean paired delta (akm - baseline) |
|---|---|---|
| tasks where akm was NEVER called | 29 | -0.011 |
| tasks where akm WAS called | 19 | +0.561 |
Trial-level, for the record: akm-arm trials WITH an akm_* call mean reward
0.857 (n=49); WITHOUT, 0.800 (n=90).
Injected context alone contributes nothing. Where the model does not call a tool, the treatment arm is statistically indistinguishable from baseline (-0.011). The hints block, the curated-file pointer and the rest of the injected context net to zero by themselves.
Therefore engagement is the mechanism, not a proxy metric: the entire
measured +0.246 eval delta is carried by trials that actually invoked akm_*.
Caveat that has to be stated. The +0.561 is confounded by task selection — the model engages precisely on the tasks where retrieval is needed and the baseline fails, so it is NOT a clean causal estimate of what a call is worth. The -0.011 is the clean half: same tasks, both arms, paired, and the treatment's extra context changes nothing.
The within-task comparison does not settle the other half either. Across the 6 tasks that had both engaged and non-engaged akm-arm trials, the mean engaged-minus-non-engaged delta is +0.250, but it splits 2 higher / 1 lower / 3 tied — not decisive on its own.
Consequence for prioritisation. A fix for akm-plugins#99 must produce actual TOOL CALLS; more or better injected context is already known to be worth ~0. And akm#819 (retrieval ceiling) is a multiplier that only pays off once engagement rises — the calls that do happen already convert well, so retrieval quality is not the current binding constraint.
2d. The memory half: akm cannot retrieve conversational content at all
Two independent memory packs, three arms each, smoke scale. baseline is a
LONG-CONTEXT REFERENCE ARM — it receives the whole conversation and never
retrieves — so it is a ceiling, not a control. The only fair comparison is
akm-memory vs raw-vector, two retrieval arms on equal footing.
LongMemEval (5 questions, qwen3.5-plus):
| variant | judgedPass | R@k | zero-hit |
|---|---|---|---|
| baseline (long-context) | 1.00 | — | — |
| raw-vector | 0.20 | 0.20 | 0/5 |
| akm-memory | 0.00 | 0 | 5/5 (100%) |
LoCoMo (5 questions, same model):
| variant | tokenF1 | zero-hit |
|---|---|---|
| baseline (long-context, 16k) | 0.567 | — |
| raw-vector | 0.233 | 0/5 |
| akm-memory | 0.200 | 2/5 (40%) |
akm scored 0 on LongMemEval because it returned no documents at all — on every question. The run emits this itself, unprompted:
5/5 retrieval queries returned zero hits (>=50%). The aggregate score for this run is dominated by prompts with no retrieved context at all, not by answer quality on retrieved context.
This is a measurement of the RETRIEVAL CEILING (akm#819), not of memory quality, and must not be published as the latter. On LoCoMo, where akm did retrieve, the answer was correct.
The mechanism is structural, not a tuning problem. akm indexes synthesized frontmatter — name, description, tags, searchHints, headings — and never body prose. A LongMemEval haystack is a chat transcript whose answer lives entirely in the body, so it is unsearchable by construction. LoCoMo reaches 40% rather than 100% only because some of its questions happen to match a synthesized field.
The sharpest contrast: raw-vector is a naive in-memory cosine store and had
0/5 zero-hits on both packs. It always returns something. The gap is not
ranking quality — it is COVERAGE.
What this means for the project's priorities. The +0.439 benchmark result came from a corpus of DOCUMENTED CONVENTIONS — content that lives in frontmatter, exactly where akm looks. Conversational recall is the opposite shape. So the benchmark win does not transfer to memory, and akm#819 is not an optimisation: it is the precondition for akm working as memory at all rather than as a convention library.
Caveats. n=5 questions, one sample, one model per pack — these numbers
separate the arms on COVERAGE, which is a structural property, but they are
far too small to rank answer quality. Also note LongMemEval reports
exactMatch/tokenF1/containsExpected as a hardcoded 0 because the pack
scores on the official LLM judge (akm-eval#6); only judgedPass carries
signal there.
2e. RESOLVED: akm#819 lifted the ceiling — measured on 0.9.2-alpha.2
§2d's diagnosis was correct and the fix confirms it. akm-cli 0.9.2-alpha.2
(npm next) carries #819's retrieval fix. Re-measuring with a RETRIEVAL-ONLY
probe — each pack adapter's exact ingest and query, no LLM in the loop, so
nothing here is confounded by model or judge — on identical corpora, identical
code paths, only the CLI version differing:
| pack | metric | 0.9.1 | 0.9.2-alpha.2 |
|---|---|---|---|
LoCoMo (conv-26, 419 docs, 40 questions) |
zero-hit rate | 75.0% | 0.0% |
| evidence recall@5 | 0.154 | 0.590 | |
| LongMemEval (20 questions, ~61 sessions each) | zero-hit rate | 100% | 0.0% |
| evidence recall@5 | 0.000 | 1.000 |
Zero-hit going to 0% on its own would be ambiguous — a retriever that returns five arbitrary documents for every query also scores 0% zero-hit. Evidence recall is what rules that out: at topK=5 over a 419-document haystack, chance recall is ~1%. 0.590 is retrieval, not noise.
The isolated unit-level proof. akm-eval's live-CLI integration test had encoded the ceiling as an assertion — a body-only term MUST return zero hits. Against 0.9.2 it returns exactly one hit, the correct document. The assertion has been inverted to guard the fix instead (akm-eval#9); it was, until then, a tripwire that fired on an improvement.
What this does and does not settle. It settles COVERAGE, which is what
§2d identified as the gap. It does not yet re-rank answer quality: the
end-to-end judged numbers (LongMemEval judgedPass, LoCoMo tokenF1) still
need a rerun against 0.9.2 before the memory-half result can be restated. The
published 0.00 / 0.200 remain the last measured end-to-end values and remain
floored by 0.9.1 retrieval — they should be treated as superseded-pending-
rerun, not as current.
A defect surfaced by the fix — since fixed and released. Under
0.9.2-alpha.2, 4 of 20 LongMemEval questions ABORTED on akm-eval's
contamination guard. Root cause was a latent akm bug that 0.9.1's near-total
zero-hit rate had hidden: a document under memories/ whose body contains an
ordinary dollar amount ($1,200, $2,000) matched
COMMAND_PLACEHOLDER_RE = /\$ARGUMENTS|\$[123]\b/ — \b sits between the 2
and the comma — and was indexed as a command, its ref becoming
commands/memories/<slug>. That content heuristic (specificity 18) outranked
the explicit parent-directory classification (15), so body sniffing beat the
directory the file was sitting in. Clean separation on the corpus: 3 of 51
documents affected, and those 3 exactly the 3 matching the regex. Present in
0.9.1 identically — newly VISIBLE, not new. Fixed in akm#824 / #826 and
shipped in 0.9.2-alpha.3, drawn along ambiguity rather than directories:
$ARGUMENTS keeps its precedence over a directory hint, while the numeric
placeholders exclude currency and defer to a directory that declares a type.
Verified against the published 0.9.2-alpha.3 artifact:
| pack | metric | 0.9.1 | 0.9.2-alpha.2 | 0.9.2-alpha.3 |
|---|---|---|---|---|
| LoCoMo | zero-hit | 75.0% | 0.0% | 0.0% |
| evidence recall@5 | 0.154 | 0.590 | 0.590 | |
| LongMemEval | zero-hit | 100% | 0.0% | 0.0% |
| evidence recall@5 | 0.000 | 1.000 (of 16) | 0.950 (of 20) | |
| questions aborted | — | 4 of 20 | 0 of 20 |
LongMemEval recall@5 reads LOWER on alpha.3 only because all 20 questions now complete: alpha.2's 1.000 was over the 16 that did not abort. The alpha.3 number is the honest one.
3. Recommended changes
Ordered by expected value. Every one changes the TREATMENT (the product), not the measurement — see §4 for why that distinction is what keeps the number honest.
akm-plugins (highest value, smallest change)
-
Rewrite the
AKM_HINTS_PREFIXtrigger so it covers editing. "Before writing anything from scratch" is the single highest-leverage string in the plugin, and it currently gates out 24 of 57 trials. It should name the case that actually fails: writing or editing a config file, manifest, or command for a tool whose exact syntax you are not certain of. Cheap to change, directly falsifiable on the train slice. -
Make the
akm_curatedescription compete with the built-ins. On edit-shaped tasks the model reaches forread/glob/editbecause those obviously act on the visible file.akm_curate's description ("PRIMARY discovery entry point for the stash ... describe the task in natural language") reads as project-discovery, not as "check this format's conventions before you write it". Paid models went further and preferred opencode's built-inskilltool. -
Consider a targeted nudge when an unfamiliar format is in play. The plugin already runs
experimental.chat.system.transformon every request and already tracks a curated file per session. A line that names the concrete asset types available for the current stash would convert "there is a stash" into "there is a documented schema for the file you are about to edit".
akm (CLI / retrieval)
-
Query shaping is a known ceiling. akm's FTS is a conjunctive AND with no stopword removal, so sentence-shaped queries zero-hit while keyword queries rank 1. Anything that makes the plugin's generated query keyword- shaped (or that makes akm tolerant of sentence-shaped input) raises the payoff of every engagement that does happen.
-
Body prose is never indexed — only name/description/tags/hints/ headings. Question-form
searchHintsrecovered most of this in the treatment library (0/8 -> 7/8). Worth making that authoring rule explicit in the docs, since it decides retrieval outcomes more than ranking does.
akm-bench (corpus quality, not tuning)
-
drillbit--backup-policy's verifier is too lenient. Both arms wrotedrillbit backup configure --cluster ...against a gold ofdrillbit backup --cluster ...and both scored 1.0. A task that accepts an invented form cannot discriminate; it under-reports akm's effect whenever a model guesses close. -
The edit-shaped inkwell tasks are the corpus's most valuable assets — they are the only ones that isolate the blind spot in §2. Keep them, and consider adding edit-shaped
drillbitequivalents so the shape is not confounded with the tool family.
4. How to tune without invalidating the benchmark
The rule: change the treatment, never the measurement.
Legitimate (ships to real users): tool descriptions, injected framing, retrieval quality, indexing, query shaping. A model that consults akm more because the tools are better described is a product improvement.
Invalidating: putting akm instructions in task prompts (contaminates the baseline or becomes a treatment-only confound), touching verifiers or timeouts, giving the treatment arm anything unrelated to akm.
The real risk is iterating against the eval slice. Tune until eval improves and the plugin is fitted to those 19 tasks; the number keeps looking rigorous while it stops generalizing. The corpus already solves this:
akm-tasks-train 27 tasks
akm-tasks-eval 19 tasks
overlap: NONE (same families, disjoint instances)
Protocol:
- Iterate on train (
harbor/jobs/.corpus-train-ab-local.yaml), freely. - Decide the change and predict its effect before touching eval.
- Measure on eval rarely — ideally once per meaningful plugin change.
Today's
+0.228 [0.070, 0.404]at 28.1% engagement is the pre-registered baseline. SUPERSEDED as of akm-bench0578025— see the note below.
The eval baseline was invalidated on purpose, once. Fixing akm-bench#6 tightened the verifiers of 11 of the 19 eval-slice tasks (5 drillbit + 6 inkwell). The task LIST is unchanged, but the scoring is not, so the
+0.228 [0.070, 0.404]figure is no longer comparable to anything measured after that commit. That was the right trade — a correct verifier beats a comparable-but-wrong one, and the old one scored an inventeddrillbit backup configure ...as a pass — but it means the NEXT eval run re-establishes the baseline rather than testing against it. Since both arms are always re-run together, the new run is internally valid on its own; only the cross-run comparison is lost.
- Always re-run both arms together. The baseline moves too when opencode, the corpus, or the harness changes; comparing a new treatment against a stored baseline silently confounds.
- At n=19/k=3 the CI is ±0.17. Small engagement gains will not separate from noise on eval — another reason to do the tuning where looking is free, and to raise k when you need to resolve a smaller effect.
5. Environment notes that affect any rerun
- Model choice is not neutral. The plugin injects its context as
ADDITIONAL entries in opencode's
systemarray, so the treatment arm sends more than one system message. Chat templates that require a single leading system message (qwen3.6-35b-a3b,devstral-small-2) return HTTP 500 on the treatment arm ONLY — an arm-asymmetric hard failure, not a score. - Engagement varies by model far more than reward does. On the same
akm-relevant task:
qwen3.5-plusandqwen3-30b-a3b-2507called akm tools;kimi-k2.7-code,deepseek-v4-pro,glm-5.2,hy3-freeandnemotron-3-ultra-freedid not, despite the plugin curating successfully (prompt_recall: ok, reason=explicit-akm) in every case. Any engagement number must name its model. - Agent-phase timeouts are benchmark-defined. Each task's own
[agent] timeout_secgoverns; the harness no longer overrides it.