akm docs

Testing Workflow

This project is a CLI with three risk-heavy areas:

The safest way to test it is in layers: fast local checks first, then end-to-end CLI coverage, then Docker-based deployment and upgrade validation. This document owns automated test workflow and test-authoring guidance only. The authoritative manual setup, command matrix, expected behavior, release gates, evidence, and cleanup live in AKM Manual Testing Runbook.

What To Validate

Test Layers In This Repo

1. Fast local correctness

Run these before any release candidate or merge:

bun run test:unit
bunx biome check --write src/ tests/
bunx tsc --noEmit

Use this when you want the shortest full-project signal:

bun run check

Relevant coverage:

Writing deterministic, isolated tests

Both scripts/test-unit.sh (bun run test:unit) and scripts/test-integration.sh (bun run test:integration) shard their target across up to min(nproc, 8) concurrent bun test processes, splitting the file list round-robin — neither runner is a single shared process, and neither is "the unit suite" alone. Within one shard, every file assigned to that process still runs in one shared process: one process.env, one module-singleton namespace, and (under fake timers) one global clock for every file in that shard. A test that mutates shared state without restoring it, or that asserts on a wall-clock measurement, can pass on one scheduling and fail on another — the two release/0.8.0 flakes (scoring-pipeline Issue #14 reading the wrong index DB after a sibling mutated XDG_DATA_HOME; the llm-client timeout test racing real timers) were both this class.

Rules (conventions, not statically enforced — the lint script that policed them was deleted in 0.9.8; the runtime guard in src/core/paths.ts throws TEST_ISOLATION_MISSING and the preload tripwire catches leaks):

  1. Never mutate AKM_* / XDG_* / HOME on process.env directly. Use the sanctioned helpers in tests/_helpers/sandbox.ts:

    • sandboxStashDir, sandboxXdgConfigHome, sandboxXdgDataHome, sandboxXdgCacheHome, sandboxXdgStateHome, sandboxHome — set the env var to an isolated temp dir and return a cleanup that restores the prior value. Chain them and call cleanup() in afterEach.
    • withEnv({ AKM_BUNDLE_DIR }, async () => …) — scoped override that always restores in a finally, even on throw. Use this for per-call overrides around an in-process CLI invocation.
    • makeStashDir / makeSandboxDir — temp dirs that are NOT wired into env (pass them to withEnv or a subprocess env object yourself).

    The DB path resolves from XDG_DATA_HOME; if you akmIndex then akmSearch, both must see the same sandboxed XDG_DATA_HOME or the search reads a different (empty/stale) DB. Sandbox it in beforeEach.

  2. Do not assert on a measured wall-clock delta. expect(Date.now() - start) .toBeLessThan(N) races the scheduler. Assert the observable result instead (e.g. result.reason === "timeout"), or drive time deterministically with jest.useFakeTimers() + jest.advanceTimersByTime(...) and a fetch/spawn stub. Asserting on a durationMs field computed from injected fixture timestamps is fine (it is deterministic).

  3. Sort/compare on the value the user sees. When a test asserts ordering, make the production sort key the same quantized value that is displayed (db-search.ts sorts on the clamped+rounded score, then breaks ties by name), so an invisible sub-display-precision epsilon can never reorder visible ties.

  4. Avoid stateful feedback within one test. Re-running akmSearch without skipLogging: true writes utility/recency rows that perturb the next search's ranking. Pass skipLogging: true when you need repeatable ranking across calls in a single test.

If a file legitimately needs a literal env value (e.g. pure path-resolution unit tests) and restores via its own save/restore wrapper, add it to the linter's ENV_ASSIGN_ALLOWED set with a one-line justification — the list may only shrink.

2. End-to-end CLI validation

Run the full integration suite when changing CLI behavior, indexing, search, config, source management, or output shaping:

bun run test:integration

Collectively, these suites (tests/integration/) exercise real flows, including:

Semantic Search States

Semantic search does not behave as a simple on/off feature at runtime. Testing should distinguish between saved-config state and actual readiness.

Config intent

Semantic search intent is saved independently as semanticSearchMode:

Setup should not flip intent from auto back to off because preparation or verification fails transiently.

Runtime readiness

Actual semantic readiness is tracked separately from config intent. Runtime state can be:

These cases are covered by tests/integration/setup-run.test.ts and the focused semantic/config suites.

The real-model gate (AKM_SEMANTIC_TESTS=1) honors an explicit HF_HOME and otherwise uses the ignored, repo-local .ci-cache/huggingface directory. That path is outside the preload's disposable test HOME, so a second local run reuses the model without exposing the developer's normal AKM cache. Gated CI sets the same path and restores it from a model/source-identified Actions cache. Candidate tags and manual runs are restore-only; only the scheduled default-branch run may save a cache entry.

The same gate runs real semantic index/search round-trips against the declared external dependency. Package acceptance separately verifies the ordinary npm package surface; there is no installed-consumer compatibility architecture or copied semantic runtime to audit.

What to test explicitly

3. Docker deployment validation

Run the Docker matrix when changing install, packaging, startup, runtime dependencies, or platform behavior:

AKM_DOCKER_TESTS=1 bun test tests/integration/docker-install.test.ts

Or run the shell orchestrator directly:

./tests/docker/run-docker-tests.sh

This validates two deployment methods across four Linux families:

Binary validation currently excludes Alpine. The compiled Linux binary used in this repo's Docker tests is not packaged for Alpine/musl compatibility, so the binary deployment gate focuses on the glibc-based targets we currently support.

The Docker smoke test in tests/docker/smoke-test.sh verifies:

4. Benchmark (agent utility)

The LLM-provider-driven, multi-seed agent-utility benchmark (akm-bench) now lives in the standalone repo itlackey/akm-bench. Run it from there after any change to src/output/, src/commands/read/show.ts, APPLY directives, or other content that affects what agents see.

For curate/search ranking quality, run the retrieval eval of itlackey/akm-eval (evals/retrieval/run --corpus public, with AKM_BIN naming the build under test). It scores akm search and akm curate against graded judgments and calls no model service. The in-repo curate benchmark that used to do this, akm-eval-curate-bench, left with the rest of scripts/akm-eval/; see docs/maintainers/eval.md.

Normal change

Use this for most code changes:

bun run test:unit
bun run test:integration
bunx biome check --write src/ tests/
bunx tsc --noEmit

Use this when touching src/cli.ts, src/commands/sources/self-update.ts, install flows, source management, or Docker assets:

bun run test:unit
bun test tests/integration/self-update.test.ts tests/integration/setup-run.test.ts tests/integration/install-script.test.ts
./tests/docker/run-docker-tests.sh
bunx biome check --write src/ tests/
bunx tsc --noEmit

Release gate

Use this before publishing a release:

bun run release:check

If Docker is available, prefer ./tests/docker/run-docker-tests.sh over a single-variant container run so both bun and binary installs are covered.

For a local release gate without Docker, use:

./tests/release-check.sh --skip-docker

That script now runs a dedicated install/setup regression suite before the full test run so first-run, installer, and wizard failures surface early.

The local script is only one half of release validation. Follow the maintainer release checklist to run Gated CI from a gated-ci/candidate-* tag targeting the exact candidate commit, or manually dispatch it with the full candidate SHA after the workflow reaches the default branch. Its real-embedding, Docker, and Linux/macOS/Windows scheduler jobs must all succeed, and the release PR must link the resulting Actions run. The weekly scheduled run detects default-branch drift; it cannot attest a different release-candidate commit.

Manual QA Authority

Use AKM Manual Testing Runbook for every manual flow. Its full sandbox is mandatory; the smaller historical XDG-only setup is unsafe because AKM_CONFIG_DIR can bypass it. The runbook also owns expected exits/channels, fixtures, release evidence, platform gates, and cleanup.

Docker Deployment Validation

Automated matrix

The repo already contains Dockerfiles for:

Dockerfile.alpine-binary does not exist and is deliberately not part of this matrix: binary validation excludes Alpine (see above) because the compiled binary targets glibc and Alpine is musl-based.

Run one variant if you need a focused repro:

./tests/docker/run-docker-tests.sh ubuntu-binary
./tests/docker/run-docker-tests.sh --bun-only
./tests/docker/run-docker-tests.sh --binary-only

What the Docker matrix proves

Treat this matrix as a release gate for shell-level regressions too. The Docker smoke path exercises real entrypoint scripts and container command execution, which can catch failures that unit and subprocess tests miss.

What it does not prove by itself

Those should be validated in disposable containers as described next.

Published Artifact Acceptance

Published binary, installer, checksum, native-platform, and self-upgrade procedures are release-facing manual gates. Run sections 21-22 of the manual runbook; do not validate release bytes with an installer fetched from raw main.

Upgrade Regression Coverage

Managed-source updates and CLI self-upgrade are separate implementations. Provider/registry suites cover source add/update/remove/cache behavior; tests/integration/self-update.test.ts covers install-method detection, orchestration, checksum, and failure paths. Real managed-source and self-upgrade acceptance is manual and lives in section 21 of the manual runbook.

Upgrade rehearsal gate

tests/integration/upgrade-rehearsal/ installs the PREVIOUS published stable akm-cli release as a real global npm package (real npm pack / npm install --global into a throwaway prefix — see scripts/package-install.ts) under the label live, drives it to build a realistic home — a filesystem bundle, a git bundle (with an in-bundle symlink at its root), a website bundle, an npm bundle, three scheduled tasks (two granted, one left ungranted) plus a manual task, and a synced native (fake) crontab — then installs the CANDIDATE build OVER live, in the SAME prefix: an in-place swap, the same thing a real npm i -g akm-cli@…/bun add -g upgrade does, not a side-by-side install. It then runs the CANDIDATE against that home (migrate status/apply, bundle list with every bundle confirmed enabled, search/show, task sync dry-run and real — plain, no --rebind, as an upgrading user actually runs it — executing the generated cron command and confirming it ran the candidate, task run, health, improve --plan), and finally installs a separate untouched copy of the PREVIOUS release and runs it back against the candidate-written home to prove read-back still works. It exists because no other suite drives a real prior release against a real candidate build: the symlink-abort regression fixed in 0.9.17-alpha.3 (d76af0a7b) went undetected by the full unit/integration suite, tests/release-check.sh, and the Docker matrix, because no fixture combined a git bundle with an in-bundle symlink, a website bundle, an npm bundle, an ungranted task, and a real crontab row in one home. Plain task sync proves clean specifically because the candidate lands inside the SAME prefix the previous release's scheduler rows already point at — no launcher path moves, so nothing needs rebinding; a side-by-side install (a separate prefix) would need --rebind instead, which is not what an in-place package-manager upgrade does. tests/integration/previous-release-corpus.test.ts covers persisted-data shapes read in isolation; this gate covers the whole CLI surface driven end to end across two real installed releases.

Run it locally with:

bun run build
AKM_UPGRADE_REHEARSAL=1 TMPDIR=/tmp bun test --timeout=900000 tests/integration/upgrade-rehearsal/

Unset (the default), the suite logs one line and skips — it never fails "inconclusively" (#795): every missing capability (no npm on PATH, no network reachable to resolve/fetch the previous release, no candidate build) throws naming the fix instead. Env overrides:

Fetched previous-release tarballs are cached under ${TMPDIR:-/tmp}/akm-upgrade-rehearsal/previous-release/<version>/ so repeat runs do not re-hit the network. It is wired into CI as the upgrade-rehearsal job (.github/workflows/ci.yml) and into tests/release-check.sh right after packing the release candidate.

Coverage Gap Guide

This repo now has broad coverage across the major CLI, indexing, registry, and semantic-search paths. Use this file as a current gap guide, not a greenfield test plan.

Areas With Strong Existing Coverage

Highest-Value Remaining Gaps

  1. corrupt or version-mismatched DB fallback behavior during search
  2. local model download and ONNX startup failures
  3. partial embedding-generation failures during indexing
  4. concurrent search/embedder behavior under load
  5. semantic readiness reporting parity between setup, search, and info
  6. broader platform CI for Alpine/musl, ARM, and Windows edge cases

When Adding Tests

Useful Existing Suites To Extend