akm docs

Website Snapshot Extraction Fix

Status: Implemented and verified Date: 2026-08-04 Reference: fwdslsh/inform at f708313 (CC-BY-4.0)

Goal

Make akm import <url> produce focused, readable Markdown for repository pages, X posts, X Articles linked from posts, and ordinary article/documentation pages. Keep the existing website provider, snapshot-fetcher registry, SSRF guards, and secret-resolution boundary.

Findings

The current behavior is not caused by an obsolete importer. The active import path sends every HTTP URL through fetchWebsiteMarkdownSnapshot().

Inform avoids the reported GitHub failure because its CLI classifies GitHub URLs before generic crawling and sends them to GitCrawler. Its generic crawler also removes navigation, menus, sidebars, ads, sharing controls, comments, related content, breadcrumbs, cookie notices, popups, modals, and overlays from a selected content region.

AKM currently differs in four material ways:

  1. GitHub repository roots reach the generic HTML converter, so GitHub's whole application <main> can win over the much narrower README.
  2. The X fetcher matches profiles only. /user/status/<id> therefore reaches the generic converter and captures X's loading shell.
  3. Content-region cleanup runs only for the <body> fallback, not inside a matched <main> or <article>.
  4. Redirected short links are classified only before the redirect. A URL that resolves to a supported site never gets a second specialized dispatch.

Design

Keep one extension point

Extend WikiSnapshotFetcher; do not add provider kinds or URL-specific command branches. Built-in fetchers remain ordered before the generic website fallback, and stash-local fetchers retain precedence over built-ins.

GitHub repository roots

Add a github-repository fetcher for github.com/<owner>/<repo> roots. Fetch the preferred README through GitHub's documented GET /repos/{owner}/{repo}/readme endpoint using the rendered-HTML media type, then pass that narrow HTML through AKM's existing safe HTML-to-Markdown converter.

X resources

Replace profile-only URL parsing with a discriminated parser for:

Profile behavior remains unchanged. For a post:

  1. Fetch its public X HTML through fetchGuardedResponse() without credentials.
  2. If the serialized page data contains an ArticleEntity, extract its title and plain_text body using a bounded JavaScript-string scanner plus JSON.parse. The scanner executes no page code and emits plain text only.
  3. Otherwise, when a bearer token exists, use documented GET /2/tweets/{id} with note_tweet and author fields.
  4. Without API content, use the target page's Open Graph description as the public fallback for an ordinary post.

This supports the reported X Article URL because it is an Article's enclosing status URL and X includes the complete Article body in that public response. A bare /i/article/<id> is recognized and the same public-data extraction is attempted, but remains best-effort: X currently omits the body from that page and documents create/publish Article endpoints only, not Article lookup. Do not bind AKM to private GraphQL operation IDs or third-party reader services.

All post and Article text is escaped with the existing escapeMarkdownStructure() helper. Bearer tokens stay confined to the fixed api.x.com request and never enter page requests, warnings, or snapshots.

Redirect re-dispatch

After the guarded generic fetch follows redirects, compare the final URL with the input URL. If it changed, offer the final URL to the same registry before using the already-produced generic Markdown. This enables short links that resolve to GitHub, X, feeds, or stash-local fetchers without weakening per-hop SSRF checks.

Only the single-URL import path gets this second dispatch. Multi-page crawls keep their existing same-origin and robots behavior.

Generic extraction

Improve the shared extractor rather than adding per-site selector tables:

Files

Verification

Tests must prove:

Run:

bunx biome check --write src/ tests/
bun test tests/website-content-extract.test.ts
bun test tests/website-feed-fetchers.test.ts
bun test tests/website-github-fetcher.test.ts
bunx tsc --noEmit
bun run check:changed

Run bun run check if focused verification exposes shared website-provider or CLI-contract changes.