akm docs

Implementation Plan: Porting Inform Crawler Capabilities into akm

Status: Draft — awaiting review Source repo: fwdslsh/inform (Bun web crawler / content ingester, CC-BY-4.0) Target repo: itlackey/akm (MPL-2.0) Branch: claude/inform-akm-porting-nk81wa


1. Goal and scope

Port the highest-value capabilities from the inform codebase into akm's existing website-source architecture, rewritten (not copy-pasted) to akm's TypeScript conventions, security posture, and test-isolation harness.

In scope (in priority order)

# Capability Inform source akm destination
P1 robots.txt compliance (Disallow / Crawl-delay, per-domain cache) src/RobotsParser.js new src/sources/snapshot-fetchers/robots.ts, wired into website-ingest.ts crawlWebsite()
P2 Main-content extraction before HTML→Markdown conversion (strip nav/header/footer/aside, prefer <main>/<article>/content selectors) src/WebCrawler.js (extractContentWithHTMLRewriter) new src/sources/snapshot-fetchers/content-extract.ts, called from htmlToMarkdown() in website-ingest.ts
P3 RSS / Atom / RDF feed ingestion as a WikiSnapshotFetcher src/sources/rss.js new src/sources/snapshot-fetchers/rss.ts, registered in registry.ts BUILTIN_FETCHERS
P4 Bluesky profile ingestion as a WikiSnapshotFetcher src/sources/bluesky.js new src/sources/snapshot-fetchers/bluesky.ts, registered in registry.ts
P5 X/Twitter ingestion (API v2 with bearer token; RSS-template fallback) as a WikiSnapshotFetcher src/sources/x.js new src/sources/snapshot-fetchers/x.ts, registered in registry.ts

Explicitly out of scope

Approved dependencies

akm has been dependency-averse on this path, but the owner has approved adding real parsers rather than extending the hand-rolled regex converters (see §1.1). The following are approved; anything beyond them needs a fresh decision:

Dependency Phase Purpose
turndown P2 HTML→Markdown conversion (the library inform uses; ships its own DOM shim for Node)
a DOM selector lib (node-html-parser proposed) P2 Selecting the main-content region before conversion
fast-xml-parser P3 RSS 2.0 / Atom 1.0 / RDF feed parsing (also inform's choice)

The P2 spec must confirm the exact selector library and verify that bun run build still produces working standalone binaries — akm compiles with bun build --compile, so a dependency with native bindings or dynamic require would break the release artifacts. ./tests/release-check.sh must pass before P2 closes.

1.1 Decisions on record

Settled by the repo owner on 2026-08-01, before implementation started. These are inputs to the phase specs, not open questions:

  1. robots.txt is honored by default. respectRobots defaults to true. This is an intentional behavior change — existing website sources may return fewer pages after upgrade. Requires a CHANGELOG.md callout and a documented opt-out.
  2. P2 uses a real DOM parser plus turndown, not extended regex scanning. One code path on both Bun and Node; no HTMLRewriter runtime branching.
  3. P3 uses fast-xml-parser, matching inform, rather than hand-rolled feed scanning.
  4. X bearer token resolves from X_BEARER_TOKEN and akm's secret store via src/core/env-secret-ref.ts. P5 stays in scope.

Licensing note

inform is CC-BY-4.0; akm is MPL-2.0. All ported code is a clean rewrite in TypeScript against akm's contracts, using inform as a behavioral reference. Attribution to inform goes in the PR description and this document, not in per-file headers. Every new file carries akm's standard MPL-2.0 header.


2. Architectural constraints (from AGENTS.md + docs/architecture/architecture.md)

Implementers and reviewers must treat these as hard requirements:

  1. MPL-2.0 header on every new src/ and tests/ file.
  2. No barrel exports, no public API. New modules are internal; import them directly by path.
  3. SSRF guards are non-negotiable. Any new fetch path (robots.txt, feed URLs, Bluesky XRPC, X API, RSS-template URLs) must go through the same validation as website-ingest.ts: assertWebsiteRequestUrl + assertResolvedHostAllowed (resolve-then-validate on every redirect hop), fetchWithRetry from src/core/common.ts with explicit timeouts, readBodyWithByteCap with a byte cap, and redirect caps. No raw fetch. 3a. The allowPrivateHosts test escape hatch must thread through any new fetch path exactly as it does today, or the test suite cannot exercise loopback fixtures.
  4. Fetcher contract is WikiSnapshotFetcher (snapshot-fetchers/types.ts): { name, matches(url, context), fetch(url, context) } returning WikiSnapshotResult | null. New ingesters are built-in fetchers appended to BUILTIN_FETCHERS in registry.ts. Order matters: more specific matchers (youtube, bluesky, x, rss) run before the generic website fallback.
  5. Output shape: fetchers produce markdown snapshots that flow through buildMarkdownSnapshot frontmatter (name/description/sourceUrl/title/ updated/tags, lint_skip: [stale-path]). Reserved basenames index.md / log.md must be remapped (avoidReservedBasename).
  6. Style: Biome-formatted (bunx biome check --write src/ tests/), tsc --noEmit clean. Long prose/templates go in external .md/.xml asset files imported with { type: "text" }, not inline template literals.
  7. Errors: user-facing failures use UsageError / ConfigError from src/core/errors.ts (exit codes 2 / 78); never bare throw new Error on user-input paths. Diagnostics via warn() from src/core/warn.ts, not console.*.
  8. Config: any new tunables (e.g. respectRobots, feed limit) ride on the existing SourceConfigEntry.options bag, validated by the schema in src/core/config/schema/sources-bundles.ts if a schema change is needed. Default behavior must not change for existing users except where the change is the feature (robots.txt compliance — see §4.1).

Test conventions


3. Process: test-first with independent review

Work proceeds in five phases (P1–P5), strictly sequential merges into the feature branch, each phase following the same six-step cycle:

┌─────────────────────────────────────────────────────────────────┐
│ Step 1  SPEC        Opus 5 agent writes a behavior spec          │
│ Step 2  TESTS       Sonnet 5 agent writes failing tests          │
│ Step 3  TEST REVIEW Sonnet 5 agent (independent) reviews tests   │
│ Step 4  IMPLEMENT   Sonnet 5 agent makes the tests pass          │
│ Step 5  CODE REVIEW Opus 5 agent (independent) adversarial review│
│ Step 6  GATE        bun run check green + review findings closed │
└─────────────────────────────────────────────────────────────────┘

This cycle is mechanized in .claude/workflows/port-inform-phase.js. Run one phase with Workflow({name: "port-inform-phase", args: {phase: "P1"}}). The script owns sequencing, the review gates, and the bounded fix loops, so the process cannot be shortcut by an agent deciding it is done. Agents receive pointers into this plan (section headings to read) rather than pasted plan text, which keeps each subagent's context to the sections it actually needs.

Agent roles and models

Role Model Independence rule
Spec author Opus 5 (claude-opus-5) Reads inform source + akm architecture; produces the phase spec (behavior table, edge cases, security requirements, acceptance criteria).
Test author Sonnet 5 (claude-sonnet-5) Works only from the spec + akm test conventions. Writes tests that fail against current main-of-branch.
Test reviewer Sonnet 5 (claude-sonnet-5) Fresh context; has NOT seen the test author's reasoning. Checks: tests actually pin the spec, cover the edge cases, are hermetic, would catch a null implementation, don't over-fit an anticipated implementation.
Implementer Sonnet 5 Works from spec + reviewed tests. May not modify tests except with a written justification the code reviewer must countersign.
Code reviewer (final gate) Opus 5 Fresh context; adversarial. The heavyweight review of each phase. Reviews the diff for: convention violations (§2), SSRF gaps, silent behavior changes, missing error paths, style drift. Verdict per finding: CONFIRMED (must fix) or ADVISORY.
Orchestrator session model Runs the cycle, resolves disputes, commits at gates.

Independence is enforced by context: reviewer agents are launched as fresh subagents given only the spec, the diff, and this plan — never the transcript of the agent whose work they review. A reviewer who authored any artifact in a phase cannot review that phase.

Step details

Step 1 — Spec. One markdown spec per phase, committed to docs/plans/specs/pN-<name>.md. Contains: behavior table (input → expected output), inform-parity notes ("inform does X; we deliberately do Y because…"), security requirements, config surface, list of files to be created/modified, acceptance criteria checklist.

Step 2 — Tests. New tests/**.test.ts files (unit) and, where the phase touches akm add behavior, tests/integration/**. Committed with the message test(pN): failing tests for <capability>. CI/gate expectation at this commit: new tests fail, everything pre-existing passes (bun test tests/<new-file>.test.ts shows red; bun run test:unit on untouched files stays green).

Step 3 — Test review. Sonnet reviewer returns a findings list. CONFIRMED findings are fixed by the test author before implementation starts. The reviewer explicitly answers: "Would a trivially wrong implementation (returns empty, ignores robots, echoes input) pass these tests?" If yes, tests are insufficient — back to Step 2.

Step 4 — Implementation. Sonnet implementer writes src/ code until the phase's tests pass plus bun run lint, bunx tsc --noEmit, and the full bun run test:unit && bun run test:integration are green. Commit: feat(pN): <capability>.

Step 5 — Code review (final gate for the phase). Opus reviewer gets the full phase diff (git diff <phase-start>..HEAD). Findings are triaged: CONFIRMED → implementer fixes and re-runs the gate; ADVISORY → recorded in the phase spec's "review log" section. A phase needs a clean CONFIRMED-free review to close. If a fix commit changes more than ~30 lines, one more review round on the fix diff.

Step 6 — Gate. Orchestrator verifies: all tests green, Biome clean, typecheck clean, review log closed, spec acceptance boxes checked. Then a single squash-tidy commit boundary; next phase begins.

Dispute rule

If implementer and reviewer disagree after one round-trip, the orchestrator decides, recording the rationale in the phase spec. Reviewers cannot demand scope beyond the spec; scope changes require a spec amendment (Step 1 redo, cheap by design).


4. Phase breakdown

P1 — robots.txt compliance (highest value, most isolated)

New: src/sources/snapshot-fetchers/robots.ts Modified: website-ingest.ts (crawlWebsite, fetchWebsitePage call path)

Port of inform's RobotsParser semantics, adapted:

Key tests: parser table-tests (groups, wildcards, $ anchors, comments, crlf, case-insensitivity of directives), fail-open on 404/timeout, fail-open vs 5xx decision per spec, crawl skips disallowed paths (loopback integration test with a robots.txt-serving fixture server), crawl-delay clamping, respectRobots: false bypass, SSRF guard still applied to the robots.txt URL.

P2 — main-content extraction

New: src/sources/snapshot-fetchers/content-extract.ts Modified: website-ingest.ts (htmlToMarkdown gains a pre-pass)

Per decision §1.1.2, this phase replaces the hand-rolled converter with a real DOM parse plus turndown, giving one code path on both Bun and Node (no HTMLRewriter branching).

P3 — RSS/Atom/RDF fetcher

New: src/sources/snapshot-fetchers/rss.ts (+ registry entry)

P4 — Bluesky fetcher

New: src/sources/snapshot-fetchers/bluesky.ts (+ registry entry)

P5 — X/Twitter fetcher

New: src/sources/snapshot-fetchers/x.ts (+ registry entry)

Cross-cutting final pass (after P5)


5. Deliverables & sequencing summary

Order Deliverable Est. new files Risk
1 This plan (committed) 1 —
2 P1 robots.txt 2 src, 2–3 test Low — isolated
3 P2 content extraction 1 src, 1–2 test + fixtures Medium — golden churn
4 P3 RSS/Atom 1 src, 1–2 test + fixtures Medium — parser breadth
5 P4 Bluesky 1 src, 1 test + fixtures Low
6 P5 X/Twitter 1 src, 1 test + fixtures Low–medium — secret handling
7 Docs + changelog + final review sweep — Low

Each phase lands as its own reviewed commit series on claude/inform-akm-porting-nk81wa; the branch is pushed after every gate so progress is always recoverable.

6. Acceptance criteria (branch-level)