Testing Workflow
This project is a CLI with three risk-heavy areas:
- command behavior across real files, config, and cache directories
- managed-source install/update flows
- binary deployment and self-upgrade behavior
The safest way to test it is in layers: fast local checks first, then end-to-end CLI coverage, then Docker-based deployment and upgrade validation. This document owns automated test workflow and test-authoring guidance only. The authoritative manual setup, command matrix, expected behavior, release gates, evidence, and cleanup live in AKM Manual Testing Runbook.
What To Validate
- core CLI flows:
setup,index,search,show,info,bundle list,config - asset lifecycle: add assets, re-index, search, show, and incremental refresh
- managed-source lifecycle:
akm bundle add,akm bundle list,akm bundle update,akm bundle remove - binary lifecycle: install, run,
akm upgrade --check,akm upgrade - cross-environment behavior on Ubuntu, Debian, Alpine, and Fedora containers
Test Layers In This Repo
1. Fast local correctness
Run these before any release candidate or merge:
bun run test:unit
bunx biome check --write src/ tests/
bunx tsc --noEmit
Use this when you want the shortest full-project signal:
bun run check
Relevant coverage:
tests/integration/self-update.test.ts- self-upgrade detection and checksum enforcementtests/integration/registry-*.test.ts,tests/provider-registry.test.ts-list,remove,update, cache cleanup, install resolution, tar safety, local/git/npm pathstests/integration/setup-run.test.ts- full setup wizard orchestration and failure handlingtests/integration/install-script.test.ts- repeatableinstall.shedge cases and permission paths
Writing deterministic, isolated tests
Both scripts/test-unit.sh (bun run test:unit) and scripts/test-integration.sh
(bun run test:integration) shard their target across up to min(nproc, 8)
concurrent bun test processes, splitting the file list round-robin — neither
runner is a single shared process, and neither is "the unit suite" alone.
Within one shard, every file assigned to that process still runs in one
shared process: one process.env, one module-singleton namespace, and
(under fake timers) one global clock for every file in that shard. A test that
mutates shared state without restoring it, or that asserts on a wall-clock
measurement, can pass on one scheduling and
fail on another — the two release/0.8.0 flakes (scoring-pipeline Issue #14
reading the wrong index DB after a sibling mutated XDG_DATA_HOME; the
llm-client timeout test racing real timers) were both this class.
Rules (conventions, not statically enforced — the lint script that policed
them was deleted in 0.9.8; the runtime guard in src/core/paths.ts throws
TEST_ISOLATION_MISSING and the preload tripwire catches leaks):
-
Never mutate
AKM_*/XDG_*/HOMEonprocess.envdirectly. Use the sanctioned helpers intests/_helpers/sandbox.ts:sandboxStashDir,sandboxXdgConfigHome,sandboxXdgDataHome,sandboxXdgCacheHome,sandboxXdgStateHome,sandboxHome— set the env var to an isolated temp dir and return acleanupthat restores the prior value. Chain them and callcleanup()inafterEach.withEnv({ AKM_BUNDLE_DIR }, async () => …)— scoped override that always restores in afinally, even on throw. Use this for per-call overrides around an in-process CLI invocation.makeStashDir/makeSandboxDir— temp dirs that are NOT wired into env (pass them towithEnvor a subprocess env object yourself).
The DB path resolves from
XDG_DATA_HOME; if youakmIndexthenakmSearch, both must see the same sandboxedXDG_DATA_HOMEor the search reads a different (empty/stale) DB. Sandbox it inbeforeEach. -
Do not assert on a measured wall-clock delta.
expect(Date.now() - start) .toBeLessThan(N)races the scheduler. Assert the observable result instead (e.g.result.reason === "timeout"), or drive time deterministically withjest.useFakeTimers()+jest.advanceTimersByTime(...)and a fetch/spawn stub. Asserting on adurationMsfield computed from injected fixture timestamps is fine (it is deterministic). -
Sort/compare on the value the user sees. When a test asserts ordering, make the production sort key the same quantized value that is displayed (
db-search.tssorts on the clamped+rounded score, then breaks ties by name), so an invisible sub-display-precision epsilon can never reorder visible ties. -
Avoid stateful feedback within one test. Re-running
akmSearchwithoutskipLogging: truewrites utility/recency rows that perturb the next search's ranking. PassskipLogging: truewhen you need repeatable ranking across calls in a single test.
If a file legitimately needs a literal env value (e.g. pure path-resolution unit
tests) and restores via its own save/restore wrapper, add it to the linter's
ENV_ASSIGN_ALLOWED set with a one-line justification — the list may only shrink.
2. End-to-end CLI validation
Run the full integration suite when changing CLI behavior, indexing, search, config, source management, or output shaping:
bun run test:integration
Collectively, these suites (tests/integration/) exercise real flows,
including:
- fallback search without an index
index -> search -> show- CLI subprocess execution through
src/cli.ts - config read/write behavior
- registry-source compatibility
- progressive indexing and re-indexing
- update and upgrade command error paths
- knowledge
#fragmentselection and mixed asset discovery
Semantic Search States
Semantic search does not behave as a simple on/off feature at runtime. Testing should distinguish between saved-config state and actual readiness.
Config intent
Semantic search intent is saved independently as semanticSearchMode:
offwhen the user opts out explicitlyautowhen the user wants semantic search enabled
Setup should not flip intent from auto back to off because preparation or
verification fails transiently.
Runtime readiness
Actual semantic readiness is tracked separately from config intent. Runtime state can be:
pendingwhen semantic search is enabled but not yet verifiedready-jswhen every entry has a vector (the name is historical; there is no other vector path)blockedwhen semantic search cannot run with the current provider/setup
These cases are covered by tests/integration/setup-run.test.ts and the focused
semantic/config suites.
The real-model gate (AKM_SEMANTIC_TESTS=1) honors an explicit HF_HOME and
otherwise uses the ignored, repo-local .ci-cache/huggingface directory. That
path is outside the preload's disposable test HOME, so a second local run
reuses the model without exposing the developer's normal AKM cache. Gated CI
sets the same path and restores it from a model/source-identified Actions
cache. Candidate tags and manual runs are restore-only; only the scheduled
default-branch run may save a cache entry.
The same gate runs real semantic index/search round-trips against the declared external dependency. Package acceptance separately verifies the ordinary npm package surface; there is no installed-consumer compatibility architecture or copied semantic runtime to audit.
What to test explicitly
- config stays
offonly when the user disables semantic search intentionally - config stays
autowhen preparation is skipped intentionally - config stays
autowhen preparation fails but runtime status becomesblocked - runtime status becomes
pending,ready-js, orblockedas appropriate - index and info output report readiness state instead of only config intent
3. Docker deployment validation
Run the Docker matrix when changing install, packaging, startup, runtime dependencies, or platform behavior:
AKM_DOCKER_TESTS=1 bun test tests/integration/docker-install.test.ts
Or run the shell orchestrator directly:
./tests/docker/run-docker-tests.sh
This validates two deployment methods across four Linux families:
- bun-based install: Ubuntu, Debian, Alpine, Fedora
- compiled binary install: Ubuntu, Debian, Fedora
Binary validation currently excludes Alpine. The compiled Linux binary used in this repo's Docker tests is not packaged for Alpine/musl compatibility, so the binary deployment gate focuses on the glibc-based targets we currently support.
The Docker smoke test in tests/docker/smoke-test.sh verifies:
akm --helpakm bundle create- bundle directory creation
akm indexakm searchakm showakm infoakm bundle list- incremental re-index after adding a new asset
4. Benchmark (agent utility)
The LLM-provider-driven, multi-seed agent-utility benchmark (akm-bench) now
lives in the standalone repo itlackey/akm-bench.
Run it from there after any change to src/output/, src/commands/read/show.ts,
APPLY directives, or other content that affects what agents see.
For curate/search ranking quality, run the retrieval eval of
itlackey/akm-eval
(evals/retrieval/run --corpus public, with AKM_BIN naming the build under
test). It scores akm search and akm curate against graded judgments and
calls no model service. The in-repo curate benchmark that used to do this,
akm-eval-curate-bench, left with the rest of scripts/akm-eval/; see
docs/maintainers/eval.md.
Recommended Workflow
Normal change
Use this for most code changes:
bun run test:unit
bun run test:integration
bunx biome check --write src/ tests/
bunx tsc --noEmit
Install, packaging, or release-related change
Use this when touching src/cli.ts, src/commands/sources/self-update.ts, install flows,
source management, or Docker assets:
bun run test:unit
bun test tests/integration/self-update.test.ts tests/integration/setup-run.test.ts tests/integration/install-script.test.ts
./tests/docker/run-docker-tests.sh
bunx biome check --write src/ tests/
bunx tsc --noEmit
Release gate
Use this before publishing a release:
bun run release:check
If Docker is available, prefer ./tests/docker/run-docker-tests.sh over a
single-variant container run so both bun and binary installs are covered.
For a local release gate without Docker, use:
./tests/release-check.sh --skip-docker
That script now runs a dedicated install/setup regression suite before the full test run so first-run, installer, and wizard failures surface early.
The local script is only one half of release validation. Follow the
maintainer release checklist to
run Gated CI from a gated-ci/candidate-* tag targeting the exact candidate
commit, or manually dispatch it with the full candidate
SHA after the workflow reaches the default branch. Its real-embedding, Docker,
and Linux/macOS/Windows scheduler jobs must all succeed, and the release PR
must link the resulting Actions run. The weekly scheduled run detects
default-branch drift; it cannot attest a different release-candidate commit.
Manual QA Authority
Use AKM Manual Testing Runbook for every manual
flow. Its full sandbox is mandatory; the smaller historical XDG-only setup is
unsafe because AKM_CONFIG_DIR can bypass it. The runbook also owns expected
exits/channels, fixtures, release evidence, platform gates, and cleanup.
Docker Deployment Validation
Automated matrix
The repo already contains Dockerfiles for:
tests/docker/Dockerfile.ubuntu-buntests/docker/Dockerfile.debian-buntests/docker/Dockerfile.alpine-buntests/docker/Dockerfile.fedora-buntests/docker/Dockerfile.ubuntu-binarytests/docker/Dockerfile.debian-binarytests/docker/Dockerfile.fedora-binary
Dockerfile.alpine-binary does not exist and is deliberately not part of this
matrix: binary validation excludes Alpine (see above) because the compiled
binary targets glibc and Alpine is musl-based.
Run one variant if you need a focused repro:
./tests/docker/run-docker-tests.sh ubuntu-binary
./tests/docker/run-docker-tests.sh --bun-only
./tests/docker/run-docker-tests.sh --binary-only
What the Docker matrix proves
- the CLI starts in minimal Linux images
- runtime dependencies are sufficient for
bundle create,index,search, andshow - bun-linked installs work after building from source
- compiled Linux binaries run correctly when copied into the image
- the CLI can create a fresh bundle, build an index, and discover new assets
Treat this matrix as a release gate for shell-level regressions too. The Docker smoke path exercises real entrypoint scripts and container command execution, which can catch failures that unit and subprocess tests miss.
What it does not prove by itself
- that the published release artifact matches the local compiled binary
- that
install.shworks against a real GitHub release - that self-upgrade can replace the running binary in-place
Those should be validated in disposable containers as described next.
Published Artifact Acceptance
Published binary, installer, checksum, native-platform, and self-upgrade
procedures are release-facing manual gates. Run sections 21-22 of the
manual runbook; do not validate release bytes
with an installer fetched from raw main.
Upgrade Regression Coverage
Managed-source updates and CLI self-upgrade are separate implementations.
Provider/registry suites cover source add/update/remove/cache behavior;
tests/integration/self-update.test.ts covers install-method detection,
orchestration, checksum, and failure paths. Real managed-source and self-upgrade
acceptance is manual and lives in section 21 of the
manual runbook.
Upgrade rehearsal gate
tests/integration/upgrade-rehearsal/ installs the PREVIOUS published stable
akm-cli release as a real global npm package (real npm pack / npm install --global into a throwaway prefix — see scripts/package-install.ts)
under the label live, drives it to build a realistic home — a filesystem
bundle, a git bundle (with an in-bundle symlink at its root), a website
bundle, an npm bundle, three scheduled tasks (two granted, one left
ungranted) plus a manual task, and a synced native (fake) crontab — then
installs the CANDIDATE build OVER live, in the SAME prefix: an in-place
swap, the same thing a real npm i -g akm-cli@…/bun add -g upgrade does,
not a side-by-side install. It then runs the CANDIDATE against that home
(migrate status/apply, bundle list with every bundle confirmed enabled,
search/show, task sync dry-run and real — plain, no --rebind, as an
upgrading user actually runs it — executing the generated cron command and
confirming it ran the candidate, task run, health, improve --plan), and
finally installs a separate untouched copy of the PREVIOUS release and runs
it back against the candidate-written home to prove read-back still works.
It exists because no other suite drives a real prior release against a real
candidate build: the symlink-abort regression fixed in 0.9.17-alpha.3
(d76af0a7b) went undetected by the full unit/integration suite,
tests/release-check.sh, and the Docker matrix, because no fixture combined
a git bundle with an in-bundle symlink, a website bundle, an npm bundle, an
ungranted task, and a real crontab row in one home. Plain task sync proves
clean specifically because the candidate lands inside the SAME prefix the
previous release's scheduler rows already point at — no launcher path moves,
so nothing needs rebinding; a side-by-side install (a separate prefix) would
need --rebind instead, which is not what an in-place package-manager
upgrade does. tests/integration/previous-release-corpus.test.ts covers
persisted-data shapes read in isolation; this gate covers the whole CLI
surface driven end to end across two real installed releases.
Run it locally with:
bun run build
AKM_UPGRADE_REHEARSAL=1 TMPDIR=/tmp bun test --timeout=900000 tests/integration/upgrade-rehearsal/
Unset (the default), the suite logs one line and skips — it never fails
"inconclusively" (#795): every missing capability (no npm on PATH, no
network reachable to resolve/fetch the previous release, no candidate build)
throws naming the fix instead. Env overrides:
AKM_UPGRADE_FROM=<version>— pin the previous release instead of resolving the highest published stable version below the candidate'spackage.jsonversion.AKM_CANDIDATE_TARBALL=<path>— use an already-packed candidate tarball (tests/release-check.shpasses$PACKAGE_CANDIDATE) instead of runningnpm packon the working tree (which requiresdist/cli.js— runbun run buildfirst).
Fetched previous-release tarballs are cached under
${TMPDIR:-/tmp}/akm-upgrade-rehearsal/previous-release/<version>/ so repeat
runs do not re-hit the network. It is wired into CI as the upgrade-rehearsal
job (.github/workflows/ci.yml) and into tests/release-check.sh right
after packing the release candidate.
Coverage Gap Guide
This repo now has broad coverage across the major CLI, indexing, registry, and semantic-search paths. Use this file as a current gap guide, not a greenfield test plan.
Areas With Strong Existing Coverage
- database and scoring (
tests/integration/db.test.ts,tests/integration/db-scoring.test.ts,tests/integration/fts-field-weighting.test.ts) - search/show CLI surfaces (
tests/integration/commands/search-cli-envelope.test.ts,tests/integration/commands/show.test.ts, and the othersearch-*/show-*suites undertests/integration/) - registry install/search/update/list flows
- workflow, vault, and wiki behavior
- semantic status, vector search, and embedding config behavior
- CLI error handling and output shaping
- Docker install validation
Highest-Value Remaining Gaps
- corrupt or version-mismatched DB fallback behavior during search
- local model download and ONNX startup failures
- partial embedding-generation failures during indexing
- concurrent search/embedder behavior under load
- semantic readiness reporting parity between
setup,search, andinfo - broader platform CI for Alpine/musl, ARM, and Windows edge cases
When Adding Tests
- prefer focused
bun:testfiles undertests/ - use isolated temp config/cache/stash dirs
- cover the user-visible CLI behavior when the risk is output shaping or command routing
- cover internal units directly when the risk is scoring, metadata extraction, or persistence
Useful Existing Suites To Extend
tests/integration/commands/search-cli-envelope.test.tstests/integration/commands/show.test.tstests/integration/vector-search.test.tstests/integration/semantic-status.test.tstests/setup-wizard.test.ts,tests/setup-scheduled-tasks.test.tstests/integration/info-command.test.tstests/integration/docker-install.test.ts