AKM Data & Telemetry
AKM stores your data locally on your machine and has no telemetry: it does not send usage data, analytics, or crash reports to Anthropic or to the AKM project, and it has no servers of its own that receive your data. It does, however, make network requests to the endpoints you configure — your LLM/embedding provider, the registries and source hosts you install from, and GitHub for upgrades — and those endpoints necessarily receive whatever those requests contain. This document describes exactly what AKM reads, writes, and sends.
No Telemetry
AKM does not:
- Send usage data, events, or crash reports to Anthropic or the AKM project
- Contact any AKM-operated analytics or telemetry endpoint at runtime
- Include any analytics SDK or beacon
- Collect your email, name, or any personally-identifying information for the project's benefit
AKM adds no network destinations of its own. The requests it does make all go to endpoints you chose or invoked, and those third parties receive whatever the request contains:
- Your configured LLM/embedding provider (e.g. Anthropic, OpenAI, a local Ollama, or any OpenAI-compatible endpoint) receives the prompts and asset content sent for reflect/propose/distill/embedding when you enable those features. If you point AKM at Anthropic, Anthropic receives those requests.
- Registry metadata and bundle packages from sources you explicitly configure (GitHub, npm, git remotes, websites) — those hosts receive the fetch/clone/crawl requests, and website sources receive requests for the pages you crawl.
akm upgrade— fetches the latest release from GitHub releases (GitHub sees the request).akm setup— a single DNS lookup forgithub.comto decide whether to skip network-dependent steps (Ollama detection, remote embedding probes) when offline. No HTTP request is made by this probe; if it succeeds, akm proceeds with the network-dependent steps you already configured.akm improvedead-link checks — a full-scope improve run (the default for a bareakm improve) sends best-effortHEADrequests (following redirects, with a short per-request timeout, checked at a bounded concurrency rather than all at once) to every URL found in the bodies of the knowledge assets it is improving, to flag dead links. The hosts of those URLs see aHEADrequest; no asset content is sent. Keep URLs you don't want probed out of knowledge-asset bodies, or run improve with an explicit narrower scope.
In every case the receiving endpoint is one you configured or invoked; the data leaving your machine is the data you directed AKM to send there.
Dry runs and diagnostic output
Command dry-run is an intentionally zero-write diagnostic. A command dry-run does not mutate or write authored source. A command dry-run does not mutate or write durable state. A command dry-run records no usage. A command dry-run emits no events. A command dry-run performs no accounting.
akm command run --dry-run still reads the selected source and configuration,
performs authorization, and lowers a request. It does not dispatch or
materialize credentials. Its output contains only safe field provenance and
fixed lowering notices; resolved prompt, command, environment, endpoint,
model, and credential values are excluded. Live --verbose writes the same
safe diagnostic metadata to stderr while preserving normal stdout.
Local On-Disk Surface
AKM writes to these locations on your machine. All paths follow XDG Base Directory conventions on Linux/macOS and Windows conventions on Windows.
Config Directory ($XDG_CONFIG_HOME/akm or ~/.config/akm/)
| Path | Contents | Safe to delete? |
|---|---|---|
config.json |
Your AKM configuration: engines, improve strategies, bundles (bundle sources), and experimental opt-ins — see Configuration | No — deleting resets all settings |
Override: set AKM_CONFIG_DIR or XDG_CONFIG_HOME.
Data Directory ($XDG_DATA_HOME/akm or ~/.local/share/akm/)
| Path | Contents | Safe to delete? |
|---|---|---|
index.db |
Search index for all your bundle assets (FTS5 + metadata) | Yes — rebuilds via akm index --full |
state.db |
Events, local usage telemetry, proposals, task history, improve run results, and workflow run state/history (the former workflow.db was folded in during the 0.9.0 cutover) |
No — deletes event/usage logs, proposal queue, improve history, and workflow run history |
logs.db |
Structured, high-volume task/run log lines ({ts, task_id, run_id, stream, level, line}), joined to state.db's task_history rows by task_id@started_at. Kept separate from state.db because log lines are append-only and freely purgeable, unlike durable state |
Yes — log lines are regenerable per run; deleting loses historical run output only |
akm.lock |
Inter-process write lock | Yes — recreated automatically |
backups/tasks/ |
Copies of task files taken by akm migrate apply before it rewrites them, one timestamped directory per run that rewrote a file |
Yes — once the migrated tasks are verified |
akm.lock.lck |
Lock write sentinel | Yes — recreated automatically |
Override: set AKM_DATA_DIR or XDG_DATA_HOME.
These files take your process umask — akm does not set or change their
permissions. They hold task history, captured command output, and indexed
content, so on a shared machine you probably do not want them world-readable;
set a tighter umask, or chmod the directory yourself. akm will not do it for
you, and akm health will not nag about it either — 0644 under a default
022 umask is simply the expected state.
If akm cannot read this directory — a uid/ownership mismatch, for instance
when two accounts share one $XDG_DATA_HOME — commands fail loudly with a
DATA_DIR_UNREADABLE config error (exit 78) naming the path, the errno, the
mode and owner, and the uid you are running as. They do not report an empty
index. akm health stays runnable in that state and reports it as a failing
state-db-readable check, so it remains the command to reach for.
0.9.1 note. A pre-release build briefly chmodded this directory to
0700and the databases to0600on every open. That was reverted: it silently changed the permissions of directories akm did not create, which broke installs sharing$XDG_DATA_HOMEbetween two uids. If a 0.9.1 pre-release tightened your data directory and you need it shared again,chmodit back.
Cache Directory ($XDG_CACHE_HOME/akm or ~/.cache/akm/)
Everything in the cache is regenerable. It is safe to delete the entire cache directory; AKM will recreate what it needs on next use.
| Path | Contents | Safe to delete? |
|---|---|---|
config-backups/config-<timestamp>.json |
Pre-save config snapshots (5 retained; owner-only permissions — file 0600, dir 0700, since 08-F4) |
Yes |
config-backups/config.latest.json |
Latest backup alias (owner-only 0600) |
Yes |
registry/ |
Downloaded registry tarballs (bundle packages from npm, GitHub, etc.) | Yes — re-downloaded on next akm bundle add or akm bundle update |
registry-index/ |
Legacy per-URL JSON cache (v0.7 artifact) | Yes — fully replaced by index.db in 0.8.0 |
semantic-status.json |
Semantic index build status marker | Yes |
bin/ |
Downloaded AKM binary cache (used by akm upgrade) |
Yes |
tasks/logs/ |
Scheduled task log files. Written at your umask; they hold captured command/agent output, so tighten the directory yourself if the machine is shared | Yes — ephemeral logs |
tasks/history/ |
Legacy task history JSONL (v0.7 migration artifact) | Yes |
Override: set AKM_CACHE_DIR or XDG_CACHE_HOME.
Bundle Directory (~/akm/ by default, or user-configured)
| Path | Contents | Safe to delete? |
|---|---|---|
<stash>/ |
All your asset files: agents, skills, commands, knowledge, instructions, workflows, scripts, memories, env files, secrets, lessons, tasks, sessions, facts, plus any bundle-adapter-owned content (e.g. llm-wiki bundle roots — not an AKM PLACEMENT_SPECS type) |
No — this is YOUR data |
<stash>/.akm/ |
Hidden AKM metadata (v0.7 proposals, legacy runs) | Caution — check for pending proposals first |
Override: set AKM_BUNDLE_DIR, or configure bundles/defaultBundle in config.json (the top-level stashDir key from 0.8 is retired and rejected in 0.9 — see Configuration).
What Is Stored in state.db
state.db holds four categories of non-regenerable data:
1. Events Table
An append-only log of every mutating action you perform with AKM. Events are stored locally for self-improvement (the improve loop uses them to surface usage patterns) and for inspection via akm log.
What is recorded:
event_type— what action was taken (see full list below)ts— ISO-8601 UTC timestampref— the asset ref affected (e.g.skills/code-review), if applicablemetadata— structured payload specific to the event type (e.g. query text forsearch, score forfeedback)
What is NOT recorded:
- File contents
- LLM prompts or responses
- API keys or secrets (config is not stored in events)
- Personal information
Retention: Events older than 90 days are purged automatically when akm improve runs its maintenance pass. The window is improve.eventRetentionDays (default 90; set 0 to disable), enforced by purgeOldEvents().
Full event type list. EventType (src/core/events.ts) is an open
string union — new types can be added without a schema bump — so this is
the set of types the code actually emits at HEAD (verified against every
appendEvent(...) and insertEventOnce(...) call site, 2026-09-30), grouped by area:
Asset lifecycle
| Event type | When emitted | Key metadata fields |
|---|---|---|
add |
akm bundle add <source> |
target, name, writable; provider when given |
remove |
akm bundle remove <source> |
ref |
update |
akm bundle update [source] |
target, all, processed |
remember |
akm remember <text> |
ref, path, force |
import |
akm import <file> |
ref, source, path, force |
rekey |
scripts/rekey-asset-ref.ts moved at least one row onto a renamed asset's new ref — nothing is emitted on a no-op re-run |
ref (the new ref); metadata {from, to, changed} (row counts only) |
Search, retrieval, sync
| Event type | When emitted | Key metadata fields |
|---|---|---|
search |
akm search <query> |
query, hitCount, resultRefs, mode |
curate |
akm curate <prompt> |
query, itemCount, itemRefs |
show |
akm show <ref> |
ref, type, name |
select |
akm show after a search returning the same ref |
ref, query, searchTs, rankPosition |
feedback |
akm feedback <ref> |
signal (positive/negative), reason, tags, fix (source, the number of replacements and, for --superseded-by or --outdated, the beliefState the proposal leaves and the supersededBy ref, when a fix was attached), contentHash (sha256 of the asset's body, without its frontmatter, as it stood when the feedback was given: it lets reflect mark feedback given on an earlier version of the text, and the loop's distill pass tell that a memory flagged wrong still has it; left out for an env or secret file and when the file cannot be read) |
sync |
akm sync |
name, message, ok |
index_db_vacuumed |
akm index VACUUMed index.db, after an index layout migration or because more than half its pages were free |
pagesBefore, pagesAfter, freelistRatioBefore |
index_completed |
akm index, once when a run finishes |
mode, totalMs, walkMs, llmMs, embedMs, ftsMs, finalizeMs (phase timings in milliseconds) |
stash_synced |
akm improve's internal auto-sync pass (the sync.push feature), distinct from the akm sync command above |
committed, pushed, skipped, reason, attributed (paths the run wrote and staged), unattributed (in-scope paths that went dirty during the run without the run writing them — left for their author) |
env_access |
akm env run <name> -- <command> (audit trail: key names only, values never recorded) |
ref, keys |
secret_access |
akm secret run <ref> <VAR> -- <command> (audit trail: var name only, value never recorded) |
ref, var |
Proposals
| Event type | When emitted | Key metadata fields |
|---|---|---|
promoted |
akm proposal accept <id> |
ref |
rejected |
akm proposal reject <id> |
ref |
proposal_reopened |
akm proposal reopen <id> (a rejected proposal goes back to pending) |
ref, proposalId, source, reason (when given) |
proposal_reverted |
akm proposal revert <id> (undoes a previously-accepted proposal, restores prior content) |
ref |
proposal_expired |
A pending proposal aged past the retention window and was auto-expired | ref |
proposal_expiration_pass |
Summary emitted once per akm improve maintenance run after per-proposal proposal_expired events |
expiry counts |
proposal_orphan_purge |
Stale proposals whose target asset no longer exists on disk, pruned by improve maintenance | checked, rejected |
proposal_creation_rejected |
createProposal() validation failed before write |
ref, reason, source |
triage_drained |
akm proposal drain run summary |
promoted, rejected, deferredByReason, skippedByCap, applyMode |
triage_deferred |
akm proposal drain left items unresolved after the (optional) judgment tier |
deferred, deferredByReason, reason |
akm improve pipeline
| Event type | When emitted | Key metadata fields |
|---|---|---|
improve_invoked |
Start of an akm improve run |
ref (scope); strategy, scope, dryRun, eligibleCount |
improve_completed |
akm improve run finished |
run stats |
improve_failed |
akm improve run errored |
error |
improve_skipped |
akm improve left a ref, a lane, or a group of refs out |
reason (no_new_signal, not_retrieved, distill_no_new_signal, budget_exhausted, budget_exhausted_batch, asset_missing_on_disk, strategy_filtered_all_passes, autonomy_gated, engine_unavailable, pool_below_min_size, consolidation_no_memory_updates, below_min_new_sessions, derived_memory_reflect_skipped, memory_distill_requires_feedback, distill_flagged_wrong, distill_positive_without_reason, distill_deprecated_or_superseded); count, remaining, strategy, lane or configKey where they apply |
improve_lock_recovered |
Stale improve lock cleared at startup | |
improve_review_needed |
akm feedback pushed a high-utility asset's utility below the review threshold — a review-needed escalation is recorded (not a proposal, so it can't accidentally overwrite the asset) |
ref, previousUtility, nextUtility |
reflect_invoked |
Start of reflect phase in akm improve |
ref, engine |
reflect_completed |
Reflect phase produced a proposal | ref |
improve_reflect_outcome |
Per-asset reflect result | ref, ok, durationMs, reason |
propose_invoked |
akm proposal new |
ref |
distill_invoked |
Distill phase inside the akm improve/akm proposal new pipeline. akm distill is not a CLI command — there is no standalone verb by that name |
ref, outcome (queued, skipped with a skipReason such as lesson_exists, nothing_reusable or conflict_noop, llm_failed, validation_failed, quality_rejected, review_needed) |
extract_invoked |
akm proposal extract --type <harness> / --auto, or improve-stage session extraction |
outcome, sessionId, harness |
extract_triaged |
The pre-LLM extract triage gate evaluated at least one session | evaluated, passed, triagedOut, sourceRun (aggregated) |
schema_repair_invoked |
The schema-repair pass inside akm improve (runSchemaRepairPass) attempts to patch missing frontmatter on an asset that failed schema validation. There is no akm lint --repair flag — lint has --fix/--auto-fix, unrelated to this event |
ref, outcome |
proactive_selected |
The proactive-maintenance selector runs (once per akm improve run) |
count, dueTotal, neverReflected (aggregated) |
events_purged |
Old events deleted by improve maintenance (90-day default retention) | purgedCount, retentionDays |
improve_runs_purged |
Old improve_runs rows deleted by improve maintenance (same retention window as events) |
purgedCount, retentionDays |
asset_state_gc |
Improve maintenance found asset_salience/asset_outcome rows that no longer resolve against the index (pending) or deleted them (improve.stateGc.collect); a run with neither emits nothing |
pending, collected, byTable |
state_db_vacuumed |
state.db was VACUUMed after the retention purge because more than half its pages were free | pagesBefore, pagesAfter, freelistRatioBefore |
task_logs_purged |
Old scheduled-task log files purged by improve maintenance |
Workflows
| Event type | When emitted | Key metadata fields |
|---|---|---|
workflow_started |
akm workflow run <ref> creates a run (including native workflow task execution) |
ref, runId |
workflow_step_completed |
The run completion path records a genuine completed transition |
ref, runId, stepId, status |
workflow_step_updated |
The run completion path records a non-completed transition (failed/skipped/blocked) |
ref, runId, stepId, status |
workflow_finished |
A run transition makes the run terminal |
ref, runId |
workflow_abandoned |
akm workflow abandon |
runId only — never the workflow title |
workflow_unit_started |
A unit begins through akm workflow run |
ids/status only — never unit instructions or results |
workflow_unit_finished |
A workflow unit terminates | ids/status/tokens only — never unit instructions or results |
LLM usage and health
| Event type | When emitted | Key metadata fields |
|---|---|---|
llm_usage |
Per-attempt LLM call usage telemetry (#576), written by every akm command that makes an LLM call (index, curate, workflows, agent dispatch, command run, improve, proposal drain) |
model provenance, terminal outcome, duration, optional token usage |
llm_usage_summary |
The owning LLM telemetry sink's terminal-record count marker. The process-wide sink that covers commands other than improve writes none when it saw no call |
expectedTerminalRecords |
health_probe |
akm health's state.db round-trip write/read probe. Not durably retained: the row is inserted then deleted within the same connection once the round trip is confirmed, so the net effect on the events table is always zero rows |
n/a (ephemeral) |
llm_usage rows also carry process/engine/stage (each optional; a call
made outside any attributed scope carries none of them). akm improve
(#944) builds a process x engine x model cross-tab from the LLM call records
the run's own usage sink collects — summarizeLlmUsageRecordsCrossTab in
src/commands/health/llm-usage.ts — persisted on the run result as
usageReport.byProcessEngineModel and queryable per-run or aggregated with
akm improve report; see docs/reference/cli.md's #### improve report
section. summarizeLlmUsageCrossTab is the events form of the same
aggregation, used by improve report to recompute the cross-tab from stored
llm_usage events when a run has no persisted usageReport.
2. Usage Events Table
usage_events is the local analytical record behind utility scores,
retrieval-demand counts, GRR, and real-query eval generation (0.9.0: its CLI
read surface, akm history, was removed — the table itself and everything
below still applies). It stores
search/curate queries, per-entry search impressions, explicit show/curate
engagement, feedback signals, stable refs, and timestamps. It never leaves the
machine unless you explicitly copy the database or send derived content to a
configured endpoint.
Successful search, curate, and show commands always record usage. Machine
reads are stamped by source (below), so they never skew ranking or eval.
The search summary row (the one with no entry_ref) carries resultCount,
stashHitCount, registryHitCount, resolvedCount and mode in its metadata,
plus latency for akm metrics: totalMs for the whole search, and rankMs and
embedMs for the ranking and query-embedding phases when the local search
reported them (a registry-only search has totalMs alone).
Every runtime writer stamps provenance as user, improve, task, audit, or
unknown. Direct interactive CLI traffic defaults to user; internal improve,
scheduled-task, and eval subprocesses preserve their stamp across nested
search/curate/show/remember/agent reads. Omitted or invalid writer provenance is
unknown, and pre-provenance rows rescued at the 0.9 cutover are also
unknown. Only exact source='user' rows contribute demand, utility, GRR, or
real-query labels.
Per-entry search, curate, and show rows carry a local-only
metadata.downstreamAttribution object. Version 1 uses control: true for
current traffic where memory inference does not apply; rows without the
version marker are historical/unattributed. Attributed rows use
control: false and may contain:
memoryInference:directwhen the emitted ref is an inferred child, orsurfacewhen derived description/tags were actually present in the emitted search or selected curate output. Brief output and internally replaced descriptions are controls, not surface attribution.graphExtraction: written only by releases that boosted search with the graph — the graph contribution applied to the hit, plusbodyHashandextractionRunIdwhen available. Current releases do not rank by the graph and never write it.
Attribution metadata contains fully-qualified refs and graph identifiers, never
asset bodies or provenance content. It is not added to search, curate, or
show result payloads; there is no CLI surface that reads it back (0.9.0:
akm history was removed). The full index still applies its existing
higher-priority-wins (type, entry.name) dedup across sources: attribution
source-qualifies every indexed row but does not invent a lower-priority row for
an identity that production indexing omitted.
Retention: usage events older than 90 days are purged on every akm index.
created_at is YYYY-MM-DD HH:MM:SS in UTC (SQLite datetime('now')), not
ISO 8601, so time bounds compare datetime(created_at) with the bound rather
than the raw text.
akm metrics is the read surface for this table: it aggregates the rows of a
window into per-asset usage, queries (including searches that returned nothing),
feedback, and daily series, and --format html carries the window's rows in the
page so they can be filtered and exported in the browser. It reads
usage_events, the select events, utility_scores, asset_outcome,
proposals, the workflow tables, and llm_usage events, and writes nothing.
A window longer than a store's retention is reported in the result's notes.
3. Proposals Table
The proposal queue: pending, accepted, rejected, and reverted improvement proposals for your bundle assets. Generated by akm improve, akm proposal new, and related proposal-producing flows.
Contents:
- Proposal UUID (primary key)
- Target asset ref
- Status (pending/accepted/rejected/reverted)
- Source (which process generated it — e.g.
reflect,distill) - Full proposal content (Markdown text)
- Created/updated timestamps
Beside it, the improve_ledger table records what each improve stage last did
with each asset — one row per bundle, asset ref and stage: the outcome
(proposed, accepted, rejected, quality_rejected, review_needed,
expired, unchanged, failed, judged_no_action), when it was attempted,
and the earliest time the stage may try that asset again. It holds refs,
timestamps, a proposal id and a short reason — never asset content. A row whose
next attempt depends on the asset changing rather than on a clock (the
consolidate pair pass, and a consolidate promotion once accepted or rejected)
also holds a hash of the asset's body, and no earliest-retry time.
4. Task History Table
A record of scheduled task runs (from akm task):
- Task ID, status, start/end times
- Log file path (the log content stays in
$CACHE/tasks/logs/)
How to Inspect and Clear Local Data
Report on usage
# What was searched, shown and rated in the last 30 days
akm metrics
# A dashboard you can open from disk
akm metrics --format html --output metrics.html
Inspect events
# List recent events
akm log
# Filter by type
akm log --type search --limit 20
# Filter by asset ref
akm log --ref skills/code-review
Inspect proposals
# List pending proposals
akm proposal list
# Show a specific proposal
akm proposal show <id>
Clear specific data
# Delete the search index (safe — rebuilds with akm index --full)
rm ~/.local/share/akm/index.db
# Delete all cached registry downloads
rm -rf ~/.cache/akm/registry/
# Delete config backups
rm -rf ~/.cache/akm/config-backups/
# Delete the events log from state.db (non-reversible)
# There is no akm CLI command to do this directly (`akm log` only exposes
# `list`/`tail`, no delete/purge verb). Use SQLite directly.
# Stop akm first (no `akm` process or scheduled task running): an older
# `sqlite3` (< 3.51) opened read-write alongside a running akm can corrupt
# the database.
sqlite3 ~/.local/share/akm/state.db "DELETE FROM events;"
# Delete all proposals (same precondition: stop akm first)
sqlite3 ~/.local/share/akm/state.db "DELETE FROM proposals;"
Start completely fresh (nuclear reset)
rm -f ~/.config/akm/config.json
rm -rf ~/.local/share/akm/
rm -rf ~/.cache/akm/
# Your stash files in ~/akm/ are NOT touched by the above.
Environment Variable Overrides
You can redirect any AKM directory to a custom path:
| Variable | Overrides |
|---|---|
AKM_CONFIG_DIR |
Config directory (~/.config/akm/) |
AKM_DATA_DIR |
Data directory (~/.local/share/akm/) |
AKM_SQLITE_JOURNAL_MODE |
SQLite journal mode: WAL (default), DELETE, or TRUNCATE. Use DELETE/TRUNCATE on network filesystems (NFS/SMB) where WAL is impossible. When left at the WAL default, akm auto-detects a network FS for the data dir and falls back to DELETE. |
AKM_STATE_DIR |
State directory (~/.local/state/akm/) |
AKM_CACHE_DIR |
Cache directory (~/.cache/akm/) |
AKM_BUNDLE_DIR |
Default bundle directory (~/akm/) |
XDG_CONFIG_HOME |
XDG base — akm appends /akm |
XDG_DATA_HOME |
XDG base — akm appends /akm |
XDG_STATE_HOME |
XDG base — akm appends /akm |
XDG_CACHE_HOME |
XDG base — akm appends /akm |