Current 0.9 upgrade troubleshooting
This page covers the supported upgrade boundaries in the current 0.9 line. There is no general config/storage compatibility coordinator.
akm upgrade succeeded but AKM rejects config
akm upgrade updates executable code; it does not translate an old config.
Move the incompatible config aside and run akm setup, then configure current
bundles, defaultBundle, engines, and credentials deliberately.
Do not restore an old config over the generated current file. Keep the archive with the executable version that created it.
An old database does not open
Managed current databases apply exact-prefix additive schema migrations
automatically. Released migration 002's task_history table rebuild is also
automatic and preserves every existing row. Released migration 018 is the
exception: it removes retired dead-lane state, so an ordinary command stops
before that migration and tells you to run either of:
akm upgrade # applies pending state migrations first, then installs if a release is newer
akm migrate apply # applies them alongside the task-source and config migrations
Both run on an already-current install and never need an install to reach
the migration, so either is safe as a container entrypoint: on a current
database the step is a no-op. akm upgrade installs a newer release first
when there is one, then runs akm-migrate apply — the migrator that shipped
with whatever is now installed; akm migrate apply is that step without the
release check. Immediately before
the migration, AKM takes a SQLite writer-exclusion lock, rechecks the exact
ledger, and writes a consistent SQLite snapshot beside the database as
state.db.pre-018-drop-dead-lane-schema.<UTC-digits>.<UUID>.bak. The randomized
path is reserved with exclusive creation, kept owner-only while it is written,
checked against its held file descriptor to reject symlink/inode replacement,
fsynced, and required to pass PRAGMA quick_check. Its final permissions are
never broader than the source database. Migration 018 and its ledger row commit
before that writer-exclusion window is released, so a concurrent WAL writer
cannot land between the recovery snapshot and the destructive DDL. AKM never
substitutes a raw copy of the main database file, which could omit committed
WAL content. Keep that safety copy until the upgraded installation has
completed its checks. This is a one-purpose historical ledger guard, not a
general storage migrator.
An existing database with no applied migration IDs is never treated as a fresh
install, whether its schema_migrations table is absent or empty. Ordinary
commands reject it without writing. The explicit akm upgrade / akm migrate apply path
takes and verifies a descriptor-bound snapshot named
state.db.pre-001-initial-schema.<UTC-digits>.<UUID>.bak before it creates the
ledger or applies any migration from 001 through 022. A truly new database is
distinguished only by the inode this invocation created atomically. The
pre-001 snapshot, ledger initialization, migration 001, and migration 002 all
run within one writer-exclusion transaction. A concurrent process therefore
cannot commit an old-schema row after the snapshot but before migration 002
rebuilds task_history.
Both the source reader and safety-copy target stay bound to held file descriptors while SQLite snapshots and verifies them; a pathname replacement cannot substitute a different database. AKM never unlinks a reserved backup pathname after reservation. If creation or verification fails, the error names that retained path for inspection; do not treat it as verified recovery media.
Unknown or divergent ledgers still fail closed.
Do not delete state.db as a generic repair for a current installation. First
stop writers and preserve the database plus its -wal and -shm companions.
If it came from a 0.8 installation, keep it in the 0.8 archive and initialize
fresh current state. It is not a supported input to the current runtime.
index.db is regenerable and can be rebuilt with akm index after its live
file is moved aside. state.db is durable history and is not regenerable.
If migration 018 fails after the safety copy was reported, stop every AKM
writer and preserve the failed database plus its sidecars. Verify the sibling
copy with sqlite3 /path/to/safety-copy.bak 'PRAGMA quick_check;' and use it
only with the matching older executable in an isolated data directory. Do not
overwrite the failed database while diagnosing it.
Task migration is blocked
Preview the task-only plan:
akm migrate status
akm migrate apply --dry-run
The migrator runs two generations in one pass — task-v2 to task-v3, then
task-v3 to task source v4 against the resulting files — and either
generation can block a file. A blocked file is intentionally unchanged.
Common causes at the v2-to-v3 stage are argv arrays, shell-sensitive command
forms, invalid YAML, unsupported fields, or a source that cannot be proven
writable; rewrite that file manually as task v3 and preview again. Common
causes at the v3-to-v4 stage are a GitHub Action uses: locator (no v4
equivalent) or a with: block on a non-command target (v4 wants declared
inputs: instead); rewrite that file manually as task source v4 and preview
again.
The task migrator does not repair config or databases.
Task migration was interrupted
The migrator validates and backs up each changed task immediately before its atomic replacement. Re-run the preview. Already-current task source v4 files are skipped; remaining v2 and v3 files are planned again from their current bytes. A changed input generation fails closed instead of applying a stale plan.
Use the per-file backup only to reverse that file deliberately. Do not copy a backup over a file while a task sync or scheduler process is running.
A workflow will not resume
Only the durable plan v4 family's irVersion: 5 executes. Pre-irVersion-5
stored plans are rejected rather than decoded by a compatibility runtime.
Start a new run from the current Markdown or YAML workflow source.
For an irVersion: 5 run, a missing or changed authored source is not a resume
blocker: the run uses its frozen plan. A plan-hash or schema failure is
durable-state corruption and must fail closed.
A stale transaction journal is reported
Current atomic source mutations may leave a transaction journal after a crash. Stop AKM writers and inspect the named journal and its target files. Reconcile or restore those current-source files from their recorded backups, then remove the journal only after verifying the result.
This is source-transaction recovery, not the removed 0.8 config/storage cutover protocol.
Scheduler entries point at an old installation
Run:
akm task doctor
akm task sync
Inspect ownership and the planned command. Use akm task sync --rebind only
for a scheduler entry you deliberately want the current installation to adopt.
Unknown or foreign entries must remain fail-closed.
Roll back executable code
Executable rollback and data rollback are separate:
- stop tasks and running AKM processes;
- preserve the current config/data/state/cache and authored bundles;
- restore the previous executable;
- restore the matching archived data set, if one is required.
Never combine a restored old executable with current state, or a current executable with an archived old state layout.
See Moving from 0.8 to current 0.9 and Migrating from 0.9.1 to 0.9.2.