akm docs

Decision Register: Benchmark Harness Consolidation

Status: Live — updated as decisions are made. Last reconciled against executed results 2026-08-25. Companion to: benchmark-harness-consolidation.md §8 Numbers: benchmark-tuning-findings.md is the source of truth for every figure quoted here; the reports it derives from are committed under akm-bench/results/harbor/<date>/. Branch: claude/akm-benchmark-measurement-75rpsg

Each decision is numbered to match §8 of the brief. DECIDED entries carry the consequence that follows from them — the point of this file is that later phases can be checked against it without re-litigating.

# Decision Status Blocks
D1 Longitudinal workflows (evolve, attribute) DECIDED — never exercised P7
D2 How akm reaches the trial container DECIDED + proven live P0
D3 Network policy for the akm arm RESOLVED for the scope actually run P2
D4 Canonical reward shape DECIDED (provisional lifted) P3
D5 Where workflow-compliance is scored DECIDED by implementation (post-hoc) P4
D6 Cold-start library content BUILT — quality still unmeasured P2
D7 Cross-task accumulation policy SUPERSEDED by D15 (two arms ran, not three) P2
D8 Which upstream memory harness akm-eval adapts DECIDED by implementation P6
D9 Judge-model cost budget OPEN (narrowed: coding side measured) P6
D10 Slice / fixture delivery / token budgets / masking DECIDED + amended P3
D11 Ownership of akm-plugins/evals OPEN (untouched by this work) —
D12 Harbor version pin DECIDED + held (14/14 throughout) P2
D13 SWE-bench A/B DROPPED (2026-08-25) —
D14 terminal-bench 2 at full scale NOT RUN (2026-08-25) —
D15 Accumulating arm PARKED, not dropped (2026-08-25) —

D13-D15 are not in the brief's §8 numbering — they are scope decisions taken during execution and recorded here because they change what the programme measures. See Scope decisions.


Decided

D1 — Longitudinal workflows: keep, orchestrated outside Harbor

evolve and attribute stay. akm-bench keeps a thin multi-run orchestrator that calls harbor run repeatedly and performs the feedback / improve / accept / reindex / mask steps between jobs. Harbor owns each individual job; the loop around them is ours.

Consequences

Outcome (2026-08-25): decided, never exercised. No evolve or attribute run happened. The orchestrator was never ported to 0.9.1 verbs because the arm it serves (D7's accumulating arm) was parked — see D15. The decision stands; it is simply untested. What would settle it: one orchestrated multi-run pass, which needs a configured in-container LLM engine (akm improve exits 78 without one).

D2 — akm reaches the container via a custom Harbor agent

AkmOpenCode(OpenCode) subclass overriding install(), invoked as --agent <module>:AkmOpenCode. Chosen over baking akm into task environment/Dockerfiles (which would contaminate 46 arm-neutral task definitions and make the baseline arm run an image containing akm) and over --skill upload (files, not a global install).

Consequences

Outcome (2026-08-24/25): proven live, then run at scale. The P0 gate passed in real containers (treatment arm made akm_curate x2 + akm_show x3 on opencode--select-correct-skill, control arm zero akm activity). Across every subsequent A/B — 114 + 168 + 114 trials — the baseline arm's engagement rate stayed exactly 0.0% and the akm arm's did not, which is the arm-neutrality property this decision was chosen for, measured rather than asserted. The "a job that forgets --agent silently runs the baseline" hazard was never hit: the analysis layer folds the agent kwargs digest into the arm label and every report carries a mismatch tripwire (0 mismatches, all runs).

D4 — Single reward key now; revisit at P4 (provisional)

Tasks emit exactly one reward key valued 0 or 1. Compliance sub-scores and diagnostics go to /logs/artifacts/ and are joined post-hoc by our analysis layer. Revisit once the analysis layer exists and we know what the compliance metrics actually need.

Consequences

Outcome (2026-08-25): provisional lifted, single binary reward confirmed. Every committed report discloses the count of non-errored trials whose reward is present but not exactly 0 or 1: 0 in every run. Nothing richer than pass/fail was ever needed, because the diagnostic that turned out to matter — whether the model called an akm_* tool — is read from each trial's own opencode trajectory by the analysis layer, not from the reward. The artifact join D4 said must exist before P4 could revisit this exists and is what the whole findings document is built on. No reason to reopen.

D6 — Cold-start library: BUILT (2026-08-23)

Shipped as akm-bench/harbor/treatment-library/: 26 assets (knowledge 20, skills 3, lessons 3), contamination-free, retrieval-verified against real akm 0.9.1, wired into both A/B job yamls' akm-static arms. Seed expectations in the agent self-check are now derived from the configured library. The original decision text follows.

D6 (original) — a hand-authored generic SWE skills bundle

~20–40 assets of general software-engineering practice (debugging, test running, git, build systems, common CLI tooling). Contamination-free, reusable across both terminal-bench and SWE-bench, and honest about what it claims.

Consequences

Outcome (2026-08-25): built, wired, and still unmeasured. The corpus A/B — the run that produced every headline number — seeds each trial from the task's own AKM_TASK_STASH, which takes precedence over seed_library_dir. So the treatment library was carried but overridden on all 48 corpus tasks. The only run that actually exercised it is the TB2 10-task 2-arm run, which returned a null at 3.3% engagement (see D14) — a result the analysis attributes to task shape, not to library quality. D6's central claim ("the bundle's quality caps the measurable effect") therefore remains untested. Any future TB2 or non-corpus run is a test of the library as much as of akm.

D7 — Accumulation: run both static and accumulating as separate arms

Three arms total: baseline (no plugin), akm-static (pristine per-trial copy of the library), akm-accumulating (shared mutable bundle across trials). Isolates retrieval value from learning value.

Consequences

Outcome (2026-08-25): superseded in part — two arms ran, not three. Every executed A/B is baseline vs akm-static. The accumulating arm was never run; it is parked with its rationale under D15, and the cost tripling D7 warned about never landed on D9. The static half of the decision held: per-trial pristine seeding worked on every executed A/B trial, both arms, with zero seed-related setup failures.


Open

D3 — Network policy: partially resolved by implementation (2026-08-23)

The agent pre-installs and cache-warms everything at install() time, so a session-start npm fetch is only needed when the warm boot failed (and the self-check aborts setup in that case). Residual: install() itself needs npm egress, and whether TB2/SWE-bench task network policies permit it is a first-live-run question. A related finding closed the in-process pin hole: live npm probing showed opencode installs plugins under ~/.cache/opencode/packages/, making the documented config-dir overrides file INERT - the effective fix (shipped) is a post-warm-boot realign step that force-reinstalls the pinned akm-cli into the plugin's hoisted tree, plus probe 7c verifying the hoisted package version.

Outcome (2026-08-25): resolved for the scope actually run. Install-time npm egress worked on every executed job on the local Docker backend — 114 + 168 + 114 corpus trials with zero errors, and the self-check never had to fall back to a session-start fetch. The residual this entry named was "whether TB2/SWE-bench task network policies permit it": SWE-bench is dropped (D13) and TB2's 10-task run failed on the tasks' own agent-phase time budgets, not on network policy (43/60 errored trials, all budget). Cloud sandbox backends remain unverified, and that is now the only open part of D3.

D5 — Where workflow-compliance is scored (blocks P4)

In-container tests/test.sh — which can read /logs/agent/trajectory.json (verified: _sync_agent_output runs before _run_verifier, and cloud envs upload logs back in) and akm's own state.db — puts compliance into the reward. Post-hoc in our analysis keeps it out of every Harbor artifact and regrade path. Coupled to D4; decide both together at P4.

Outcome (2026-08-25): decided by implementation, in the post-hoc direction. Nothing compliance-shaped was ever folded into a reward. The four workflow-compliance--* tasks score pass/fail through their own in-container verifiers like every other task, and every cross-cutting metric (engagement, gate-reason histogram, per-cell rates) is computed post-hoc by the analysis layer from trajectories and the plugin ledger. The in-container branch had no claimant by the time P4 arrived and P4 never ran as a distinct phase. Reopen only if a future run needs a compliance sub-score to enter the reward itself.

D8 — Which upstream memory harness akm-eval adapts (blocks P6)

Per-benchmark, and the surfaces differ: LongMemEval v1 is thinnest (JSONL + shell out to evaluate_qa.py); LoCoMo is already done; BEAM has no plugin interface (fork or mem0-shim); supermemoryai/memorybench costs ~200 LOC for one Provider and buys three benchmarks plus head-to-head vs mem0/zep. Recommendation on the table: LongMemEval v1 first, memorybench second.

Outcome (2026-08-25): decided by implementation — LongMemEval v1 + LoCoMo. Both adapters shipped, both ran three variants at smoke scale (n=5 questions, qwen3.5-plus), and both are committed. BEAM and supermemoryai/memorybench were not adopted. The recommendation on the table ("LongMemEval first, memorybench second") is half-taken; the second half is not worth buying yet, because the two packs already in hand both bottom out on the same retrieval ceiling (akm#819) — a third harness would measure that ceiling a third time rather than tell us anything new.

D9 — Judge-model cost budget (blocks P6)

No estimate exists. Reference points: Harbor's own LoCoMo parity run cost ~$35 for 5×10 trials on gpt-5-mini; LongMemEval-V2 defaults to a gpt-5.2 medium-reasoning judge; BEAM-10M ingest is likely the single most expensive item in the suite. D7 triples the coding-benchmark side. Force a number before P6.

Outcome (2026-08-25): still open, but much narrower. Two of the three cost drivers evaporated: D7's tripling never happened (D15) and SWE-bench — the item with a real per-instance bill — is dropped (D13). The coding side now has measured numbers rather than an estimate: the 114-trial enforce-mode eval run cost ~$0.32 total (mean $0.0044/trial on the akm arm vs $0.0012 on the baseline — the treatment arm carries ~1.7x the input tokens, 44k vs 26k, at ~3.7x the cost; the gap between those two ratios is not explained by anything in the report) over 2h13m wall. What is still unbudgeted is the memory side: both packs ran at n=5 questions, and no per-question judge cost has been extracted from them. Recommended sequencing, not a decision: do not spend judge budget scaling a pack whose akm arm returns nothing on 5/5 questions — re-measure after akm#819 first.

D10 — Corpus mechanics: decided by implementation (2026-08-23)

All four sub-decisions were settled during the P3 conversion:

Correction (2026-08-25): the slice counts above are stale. akm-bench#6 added two tasks — drillbit--fix-runbook-train (edit-shaped drillbit) and inkwell--new-service-scaled-train (create-shaped inkwell), authored to break the shape/tool-family confound — so akm-tasks-train@1.0 is 29 (incl. the reference fixture task) and the corpus is 48 registered tasks. Both went to TRAIN deliberately: eval is a pre-registered measurement surface and adding tasks to it would invalidate comparison outright. The eval task LIST has never changed; only its verifiers have.

Amendment (2026-08-25): one train task is barred from ARM COMPARISONS. workflow-compliance--repeated-fail-opencode-disable-provider ships an opencode.json whose $schema the write gate resolves against the task's own stash, handing the model that task's gold_ref unasked — on the treatment arm only. The task exists to measure whether the model CHOOSES to look the asset up, so the adherence pass would accrue to one arm from the harness rather than from the model. Barred via exclude_task_names in both A/B job yamls (akm-bench 2e40003); registry membership deliberately unchanged, since the task is valid and still runs under the oracle. Fixed benchmark-side on purpose: firing on a real user's opencode.json is correct product behaviour, and making the plugin benchmark-aware would special-case the benchmark.

Amendment (2026-08-25): fixture delivery gained a rule. akm-bench#7 found that graded artifacts were sitting on filenames opencode itself claims: /app/opencode.json (project config) and /app/AGENTS.md (project instructions, spliced verbatim into the agent's own system prompt via Instruction.systemPaths). Seven tasks were affected, including opencode--select-correct-skill — the task P0 was validated on — where the model's own half-written output was being fed back to it as a standing instruction mid-trial. Fixed by renaming the graded artifacts (agent-guidance.md, and off opencode.json), proven against the pinned opencode 1.18.21 binary rather than reasoned about. The D10 rule is now: a graded artifact must not sit on any filename an agent loader claims. This changed three eval-slice tasks and so cost a second baseline invalidation — see the changelog.

D11 — Ownership of akm-plugins/evals

A third eval surface neither repo covers: tier2 deterministic plugin metrics against a fake-akm shim, tier3 LLM-judge scenarios. It holds akm's ranker constant so it does not overlap akm-bench, but ownership and CI wiring need a call. Note evals/README.md:50 claims a git-ref pairwise A/B mode that does not exist in tier3/runner.ts — a doc/code mismatch to fix or drop.

Outcome (2026-08-25): unchanged and still open. Nothing in this work touched akm-plugins/evals; ownership and CI wiring are still uncalled. The doc/code mismatch this entry recorded is still live — evals/README.md still advertises "pairwise A/B between two git refs" and evals/tier3/ still contains no git-ref handling at all (re-verified 2026-08-25).

D12 — Harbor version pin: decided and implemented (2026-08-23)

Harbor is pinned to 0.22.0 (akm-bench/harbor/requirements.txt), and akm-bench/bin/check-harbor-contract executes 14 assertions over the internal behaviors the stack depends on (trial_results exclusion, pass@k k-set, exclude-after-include log filtering, result.json naming, config deep-merge layering, metadata exclusion from results, agent-log sync ordering, reward-file parsing, one-level task scanning, provider/model splitting, trial-dir layout, XDG log path, and sync-before-populate ordering). Run it on every Harbor bump; a failure names exactly which load-bearing behavior moved. Not yet wired into CI (see Open questions below).

Outcome (2026-08-25): held for the whole effort. bin/check-harbor-contract stayed 14/14 across every run on Harbor 0.22.0 — no pinned internal behavior moved under us, and no result had to be re-litigated against a harness change. This is the decision that most clearly paid for itself: it is the reason the two baseline invalidations below can be attributed to the corpus and the plugin rather than to the harness. Still not wired into CI.


Scope decisions taken during execution (2026-08-25)

Three decisions that were not in the brief's §8 set. Each changes what the programme measures, so each is recorded here with what it was traded against.

D13 — SWE-bench A/B: DROPPED

harbor/jobs/swebench-ab.yaml deleted (akm-bench a543820).

SWE-bench is maximally edit-an-existing-file-for-a-tool-the-model-already-knows: real Python repos whose conventions are in the model's weights. That is the exact cell this programme measured at 0/35 engagement with a paired delta of -0.011 on tasks where akm was never called. The predicted result is a null.

It is not a cheap null: x86_64-only images, ~120GB for full Verified, and a real bill per instance per arm.

And it is not a speculative prediction — the TB2 10-task run already delivered the same null at a tenth of the price (D14). Buying the same answer twice is not rigour.

Consequence / caveat. This is a deliberate narrowing of external validity. The programme now has no evidence about akm on real-world repository work, and should not claim any. What it has is a measurement of when akm is consulted and what a consultation is worth, on a corpus built so that retrieval is the only path. Reopen if the engagement mechanism changes enough that the known-tool cell is expected to move — that, not cost, is the condition.

D14 — terminal-bench 2 at full scale: NOT RUN

The 10-task 2-arm run stands as the TB2 result: 3.3% engagement (1 trial in 30), delta 0.000 [-0.100, 0.100], and 43 of 60 trials errored on the tasks' own agent-phase budgets. The error rate alone makes the run weak evidence about akm and strong evidence about budget fit; scaling it would buy a better-powered null on the shape D13 already argues is uninformative.

harbor/jobs/tb2-ab.yaml is kept: the 10-task run is a real published result and the config is the record of how it was produced.

Caveat. TB2 is also the only executed job that actually uses the D6 treatment library, so "TB2 returned a null" and "the D6 library is untested" are the same fact seen twice, not two independent findings.

D15 — Accumulating arm: PARKED, not dropped

D7's third arm was never run. This is not a reversal of D7's reasoning — the arm measures learning (does a bundle that accumulates across trials get better?), which is a genuinely different hypothesis from the one everything else here measures (retrieval: is a static, well-authored bundle worth consulting?). Answering the retrieval question first was the right order, and it consumed the available run budget.

Consequence. D7's warnings still apply verbatim when it is run: non-independent trials, -n 1, its own analysis treatment, a shared mount whose semantics are verified for the local Docker backend only. D1's evolve/attribute orchestration is the same parked machinery.


Open questions raised by implementation (2026-08-23)


Open questions raised by execution (2026-08-24 -> 2026-08-25)

Ordered by how much each one gates the next decision.


Changelog