P1a — behavior spec: fail-closed with rejection + target-ref classifier
Status: ready for test authoring
Phase: P1a of the akm task/workflow refactor
Branch: claude/breaking-changes-0-9-2-3cfyvp
Predecessor: docs/plans/specs/p0-invariants.md (P0 pinned every row named here)
This document is the single source of truth for the P1a implementation and
test lanes. Test authors and implementers read this spec, not the parent
plan. Every file:line below was verified at the current head of
claude/breaking-changes-0-9-2-3cfyvp. Do not re-derive the design decisions —
they are binding and reproduced verbatim in §1.
0. What P1a is (and is not)
P1a makes exactly one behavior change and one structural change:
- The fail-closed correction (Lane A). A workflow step that authors
with:on auses: tasks/<ref>target is rejected at freeze instead of having its authored mapping silently dropped. This flips P0 row R-01(c). - The classification seam (Lane B). Workflow
usesclassification stops delegating to the task-v3 grammar. A newclassifyTargetRefowns canonical asset refs; workflow classification imports nothing fromsrc/tasks/source-v3.ts. This is behavior-preserving — see the parity table in §4.3.
Everything else is diagnostics plumbing (§1, D7): five new UsageError codes
declared, two wired.
P1a is NOT: task-input support (P2b), the typed preparer (P1b), child
workflows (P3), grammar removal (P4). Nothing in this phase implements with
bindings; the rejection message must say so.
Exit-code contract is untouched. Every code introduced here is a
UsageError, so classifyExitCode still maps it to exit 2. No test that
asserts an exit code changes in P1a.
1. Binding design decisions (verbatim)
These were decided before this spec was written. They are reproduced exactly; do not re-derive or re-litigate them.
D7 diagnostics: add UsageError codes
COMPOSITION_INVALID,TASK_SOURCE_INVALID,TARGET_REF_INVALID,WORKFLOW_SOURCE_INVALID,INPUT_BINDING_INVALIDtosrc/core/errors.tswithUSAGE_HINTSentries. In P1a, only two are WIRED:COMPOSITION_INVALID(the new with-rejection) andTASK_SOURCE_INVALID(re-code thesourceErrorfunnel atsrc/tasks/source-v3.ts:210-226— message text unchanged, code changes fromINVALID_FLAG_VALUE). The other three are declared now, wired in later phases. Exit-code contract untouched (all are UsageError -> exit 2).
Lane A (with-rejection): at the head of
taskDispatch(src/workflows/ir/source-freeze-v4.ts:211), ifsource.with !== undefinedthrow UsageError codeCOMPOSITION_INVALID, message naming the step id and stating task-call inputs are not yet supported (implemented in a later phase; today they were silently ignored — this is the fail-closed correction). The DECODER still acceptswithon task steps (schema.ts:144unchanged); rejection is at freeze. builtin-command with-consumption (source-freeze-v4.ts:145-151) must be untouched.
Lane B (classifier): new
src/execution/target-ref.tsexportingclassifyTargetRef(value): {kind:"command"|"script"|"task"|"workflow", ref}for canonical asset refs, throwing UsageErrorTARGET_REF_INVALIDon anything else (fragments, malformed, empty). NO GitHub grammar, NOakm/commandspecial case (callers layer builtin detection). Rewiresrc/workflows/source-ir/semantics.ts(classifyWorkflowStepUsesat:111-148) andsrc/workflows/source-ir/uses.ts(currently a delegator toclassifyTaskV3Usesat:39-41) so workflow classification imports NOTHING fromsrc/tasks/source-v3.ts. PARITY REQUIREMENT (P1a is behavior-preserving except the with-rejection): every currently-classifiable workflowusesvalue classifies identically;akm/commandstill routes to builtin; a GitHub-locator-SHAPED value (slash-segmentedowner/repo[/path]@revshape) in a workflow step must STILL throwWorkflowSourceSemanticErrorcoderemote-action-acquisition-out-of-scope—semantics.tsimplements its own minimal locator-shape detection (shape only, no full grammar) to preserve that error until P4 removes it. Nested-workflow rejection atsemantics.ts:141-146stays. Task documents keep usingclassifyTaskV3Uses(untouched until P4).
Ratchet: new
tests/architecture/diagnostic-codes.test.tscounting the literal stringINVALID_FLAG_VALUEinsrc/tasks/**+src/workflows/**and asserting count<=the post-P1a baseline (measure it during implementation; hardcode the measured number with a comment explaining the ratchet-only-declines rule, mirroringtests/architecture/src-fn-size-ratchet.test.tsstyle).
Docs ride with code (docs-only commits skip CI):
docs/reference/workflow-schema.mddocuments thatwith:on a task-step ref errors (COMPOSITION_INVALID) pending task-input support;CHANGELOG.mdgets an[Unreleased]"Breaking changes & migration" entry for the with-rejection AND a note that task-source validation errors now use codeTASK_SOURCE_INVALID(scripts consuming JSON envelopes must update);docs/migration/v0.9.1-to-v0.9.2.mdgains the with-rejection note.
2. Diagnostics (D7) — exact edits
2.1 src/core/errors.ts
Append five members to the UsageErrorCode union (src/core/errors.ts:66-93).
The union is a closed string-literal type; USAGE_HINTS
(src/core/errors.ts:132) is Partial<Record<UsageErrorCode, string>>, so an
entry per code is optional at the type level and required by this spec.
| New code | Wired in P1a | Thrown from | USAGE_HINTS entry (exact) |
|---|---|---|---|
COMPOSITION_INVALID |
yes | source-freeze-v4.ts taskDispatch head |
Remove the step's with: block; task-call inputs arrive in a later 0.9.x release. |
TASK_SOURCE_INVALID |
yes | source-v3.ts sourceError funnel |
Fix the task source at the reported path and line, then re-run. |
TARGET_REF_INVALID |
yes (new module only) | src/execution/target-ref.ts |
Targets are canonical asset refs: `commands/review`, `scripts/build.sh`, `tasks/nightly`, `workflows/release`. |
WORKFLOW_SOURCE_INVALID |
no — declared only | (P1b+) | Run `akm workflow validate <ref>` to see the failing source location. |
INPUT_BINDING_INVALID |
no — declared only | (P2b) | Check the step's with: keys against the target's declared inputs. |
TARGET_REF_INVALID is "wired" in the narrow sense that the new module throws
it; whether it reaches a user surface in P1a depends on the wrapping described
in §4.4 (workflow callers convert it to a WorkflowSourceSemanticError, so the
user-visible workflow codes do not change — that is the parity requirement).
2.2 The TASK_SOURCE_INVALID re-code
sourceError (src/tasks/source-v3.ts:209-226) is the single funnel for 62
task-v3 source validation call sites in that module. Change only its thrown
code:
throw new UsageError(`Invalid task v3 source at ${location}: ${dotted} ${detail}`, "TASK_SOURCE_INVALID");
Message text, field-path rendering ($ for the empty path), and the
file:line location string are unchanged, byte for byte.
Scope boundary — read this before flipping any test. The re-code covers the
sourceError funnel and nothing else. In particular it does not cover:
classifyTaskV3Uses's ownUsageErrorthrows (src/tasks/source-v3.ts:523-593, including the trailingTask v3 uses must be akm/command, …message). These are directnew UsageError(…, "INVALID_FLAG_VALUE")constructions, not funnel calls; they keepINVALID_FLAG_VALUEin P1a and are deleted in P4.taskV2UnsupportedError(src/tasks/source-v3.ts:49-57), which already carries its own codeTASK_SCHEMA_VERSION_UNSUPPORTEDand its own hint (TASK_V2_MIGRATION_HINT). It is unaffected by P1a.- Any
INVALID_FLAG_VALUEoutsidesrc/tasks/source-v3.ts(runtime-v3, schedule, scheduler-binding, task-id, workflow modules).
3. Lane A — the fail-closed with rejection
3.1 The change
At the head of taskDispatch (src/workflows/ir/source-freeze-v4.ts:211-216),
before resolveOwnedAsset at :217 — so the rejection does not depend on the
task asset resolving:
if (source.with !== undefined) {
throw new UsageError(
`Workflow step ${source.id} cannot pass with: to task target ${refInput}; task-call inputs are not supported yet.`,
"COMPOSITION_INVALID",
);
}
Pinned message contract (test authors assert this verbatim, with <id> and
<ref> substituted):
Workflow step <id> cannot pass with: to task target <ref>; task-call inputs are not supported yet.
<id>issource.id— the authored step id, which every decodedWorkflowSourceStepcarries.<ref>isrefInput— the classified task ref as authored (e.g.tasks/nightly,team//tasks/nightly), not the resolved owned ref.- Rejection fires on
source.with !== undefined, i.e. an authoredwith:block of any shape that survived decode — including an empty mappingwith: {}. An absentwith:isundefinedand freezes exactly as today.
3.2 What Lane A must NOT touch
| Site | Why it stays |
|---|---|
src/workflows/source-ir/schema.ts:144 (with?: Record<string, WorkflowSourceScalar>) |
The DECODER still accepts with on task steps. Rejection is at freeze, not decode. R-01(a) stays green. |
schema.ts:389 (scalarRecord(step.with, …, true)) and :393 (step <id> with is legal only with uses) |
R-01(b) guardrails fire on the shape of with, before and independent of the new rejection. |
source-freeze-v4.ts:145-151 (resolveStep's builtin-command branch, action = source.with) |
R-01(d): with: on uses: akm/command is consumed and target-validated. The new rejection lives inside taskDispatch only, which the target.kind === "task" branch at :143 reaches — the builtin branch at :145-151 never enters it. |
source-freeze-v4.ts:220-222 (nested-workflow UsageError) |
R-03 site 2 keeps its INVALID_FLAG_VALUE code and message. It is now unreachable for a step that also authors with: (the new guard fires first) — fixtures for R-03 must not author with:. |
3.3 Ordering consequence (test authors: read this)
The new guard is the first statement in taskDispatch. For a task step that
authors with: and targets a nested workflow, COMPOSITION_INVALID now
wins over R-03's A workflow task step cannot compose a nested workflow target.
The P0 R-03 fixtures author no with:, so R-03 is unaffected — but a new
fixture must not combine the two.
4. Lane B — the target-ref classifier seam
4.1 New module: src/execution/target-ref.ts
export type TargetRefKind = "command" | "script" | "task" | "workflow";
export interface ClassifiedTargetRef {
readonly kind: TargetRefKind;
readonly ref: string;
}
export function classifyTargetRef(value: string): ClassifiedTargetRef;
Accepts exactly a canonical asset ref, using the repo's one ref parser
(parseBundleRef / bundleRefToString from src/core/asset/asset-ref.ts).
All five conditions must hold:
parseBundleRef(value)does not throw;parsed.fragment === undefined(no#fragment);bundleRefToString(parsed) === value(round-trips — rejects non-canonical spellings such asakm:commands/reviewandbad.bundle//commands/review);parsed.conceptIdcontains a/, and the family before the first/is one ofcommands,scripts,tasks,workflows;- the name after the first
/is non-empty.
Kind mapping: commands → command, scripts → script, tasks →
task, workflows → workflow. ref is the input string unchanged. The
returned object is frozen (Object.freeze), matching classifyTaskV3Uses's
existing contract.
Rejects everything else — empty string, whitespace/inner whitespace,
${{ … }} expressions, fragments, non-canonical refs, other families
(agents/, knowledge/, …), bare words, local paths, docker://, and every
GitHub locator — with:
throw new UsageError(
`Target ref ${JSON.stringify(value)} must be a canonical commands/, scripts/, tasks/, or workflows/ asset ref.`,
"TARGET_REF_INVALID",
);
Explicit non-goals (binding): no GitHub locator grammar, no akm/command
special case, no resolution, no filesystem access, no guessing. Callers layer
builtin detection.
4.2 Rewire: src/workflows/source-ir/uses.ts
Today uses.ts:12 imports classifyTaskV3Uses + TaskV3UsesTarget from
../../tasks/source-v3, and :39-41 is a one-line delegator. After P1a the
file imports nothing from src/tasks/**:
WorkflowSourceUsesTargetis declared locally as the union of:{ kind: "command" | "script" | "task" | "workflow"; ref: string },{ kind: "builtin-command"; ref: "akm/command" }, and a structural{ kind: "github-action"; ref: string; owner: string; repository: string; path?: string; revision: string }member. Thegithub-actionmember is retained as a type only — it is what an externally injected classifier may still return (see §4.4 step 8) and what keepsGithubWorkflowSourceOptions.classifyUsestype-compatible. P1a introduces no code that produces it.classifyWorkflowSourceUsesbecomes the builtin-layering wrapper:
export function classifyWorkflowSourceUses(value: string): WorkflowSourceUsesTarget {
if (value === "akm/command") return Object.freeze({ kind: "builtin-command" as const, ref: "akm/command" as const });
return classifyTargetRef(value); // throws UsageError TARGET_REF_INVALID
}
- The module docstring at
uses.ts:5-10(which describes delegation "to WP6's canonical task-v3 classifier") must be rewritten to describe the new seam.
4.3 Rewire: src/workflows/source-ir/semantics.ts
classifyWorkflowStepUses (:111-148) keeps its signature — including the
injected-classifier parameter, which github-yaml.ts:587,608 and
compile.ts:41 depend on and which the P0 delegation tests exercise:
export function classifyWorkflowStepUses(
value: string,
classifier: WorkflowSourceUsesClassifier = classifyWorkflowSourceUses,
): WorkflowSourceUsesTarget
Required evaluation order (deviating from it breaks parity):
value.includes("${{")→WorkflowSourceSemanticError("unsupported-github-expression", "GitHub expressions are unsupported in uses.")— unchanged (:115-120).- empty / untrimmed / inner-whitespace →
WorkflowSourceSemanticError("unsupported-uses-target", "uses must be one exact, non-empty executable ref")— unchanged (:121-126). canonicalTaskTarget(value)→{ kind: "task", ref: value }— unchanged (:127-128, helper at:150-163). Still first, still without calling the classifier.- Call the injected
classifier(value)insidetry— unchanged (:129-134). - Returned
kind === "workflow"→WorkflowSourceSemanticError("nested-workflow-unsupported", \Nested workflow target ${JSON.stringify(value)} is unsupported in a workflow step.`)— unchanged (:141-146`). - Returned
kind === "github-action"→WorkflowSourceSemanticError("remote-action-acquisition-out-of-scope", …)— unchanged (:134-139), retained for injected classifiers. - Otherwise return the target (
command/script/task/builtin-command). - On a classifier throw (the
catchat:131-133): compute the codeusesFailure(:229-239) would assign. If — and only if — that code isunsupported-uses-targetandisGithubLocatorShape(value)is true, throwWorkflowSourceSemanticError("remote-action-acquisition-out-of-scope", \Remote action acquisition is out of scope for ${JSON.stringify(value)}.`). Otherwisethrow usesFailure(value, cause)` exactly as today.
Step 8's ordering is load-bearing: usesFailure's prefix classifications
(docker:// → docker-action-unsupported, ./ ../ / →
local-action-path-unsupported, agents/ → non-executable-asset-ref) must
keep winning over locator-shape detection, because today's full grammar rejects
those values before reaching the locator branch.
isGithubLocatorShape(value) — minimal shape detection, local to
semantics.ts, shape only:
- exactly one
@, at index > 0 (at = value.lastIndexOf("@"); at > 0 && at === value.indexOf("@")); - the locator
value.slice(0, at)splits on/into ≥ 2 non-empty segments; the first segment (owner) matches/^[A-Za-z0-9][A-Za-z0-9._-]*$/— looser thansrc/tasks/source-v3.ts'sGITHUB_OWNER(no 39-char cap,./_allowed), which is the one-directional slack Accepted deviation A-1 covers; the second segment (repository) matches/^[A-Za-z0-9_.-]+$/and is neither.nor..— this mirrorsGITHUB_REPOSITORY(src/tasks/source-v3.ts:178) exactly, not the owner regex (see the Review log: an initial draft applied the owner regex to both segments, which rejected old-valid repository shapes); - the revision
value.slice(at + 1)is non-empty, contains none of~ ^ : ? * [ \or a C0/DEL control character, does not start or end with/, does not contain..or@{or a literal@, and no/-separated revision segment is./.., starts with., or ends with./.lock— this mirrorsvalidGithubRevision(src/tasks/source-v3.ts:497-520) by forbidden-character set, not a strict charset allowlist (see the Review log: an initial draft used a strict/^[A-Za-z0-9._/-]+$/allowlist, which rejected old-valid revisions such asv1.0+meta/%40).
This is deliberately not the full grammar
(src/tasks/source-v3.ts:562-588) and must not import it.
Accepted deviation A-1 (recorded, not a defect): a value that is
locator-shaped under the rule above but that the old full grammar rejected for
a reason the shape rule does not encode now yields
remote-action-acquisition-out-of-scope where it previously yielded
unsupported-uses-target. Both are WorkflowSourceSemanticError rejections of
the same value at the same boundary, both keep the source non-compiling, and P4
deletes the row entirely. No existing test pins such a value — the pinned table
in §4.5 is the complete set, and its one locator near-miss
(actions/checkout@bad:ref, revision contains :) is excluded by the charset
rule and keeps unsupported-uses-target.
4.4 Rewire: src/workflows/source-ir/compile.ts (required — do not skip)
compile.ts:41 injects options.classifyUses ?? classifyTaskV3Uses into
parseGithubWorkflowSource, which overrides the semantics.ts default for the
entire GitHub-YAML entrypoint (github-yaml.ts:587 → :608). Rewiring
semantics.ts and uses.ts alone would leave the GitHub path still classifying
through the task-v3 grammar — and tests/workflows/source-ir-contract.test.ts
(the parity gate, §4.5) runs through exactly that path.
Change the default only:
classifyUses: options.classifyUses ?? classifyWorkflowSourceUses,
classifyTriggers: options.classifyTriggers ?? classifyTaskV3Triggers
(compile.ts:42) stays — trigger classification is not target-ref
classification and is owned by P2a. compile.ts therefore keeps exactly one
import from src/tasks/source-v3.ts: classifyTaskV3Triggers. The
"imports NOTHING from src/tasks/source-v3.ts" invariant is scoped to
uses classification: semantics.ts and uses.ts import nothing from it
at all, and compile.ts retains only the trigger import.
The two classification entrypoints after P1a:
| Entrypoint | Path | Default uses classifier after P1a |
|---|---|---|
| GitHub YAML | compileGithubWorkflowSource → parseGithubWorkflowSource → github-yaml.ts:608 → classifyWorkflowStepUses(uses, injected) |
classifyWorkflowSourceUses (via compile.ts:41) |
| Strict decode / Markdown | decodeWorkflowSourceIrV1 (schema.ts:358) — no injection |
classifyWorkflowSourceUses (the semantics.ts:113 parameter default) |
Both must resolve to the same function.
4.5 PARITY TABLE — the binding acceptance for Lane B
This is tests/workflows/source-ir-contract.test.ts:429-457, reproduced as the
authoritative input→expected table. It runs through
compileGithubWorkflowSource with no injected classifier. That test file
must stay green UNCHANGED — it is the parity gate, not an authorized flip.
Accepted (result.ok === true):
uses: value |
Classified kind |
|---|---|
akm/command (with with: { content: … }) |
builtin-command |
commands/review |
command |
team//commands/review |
command |
tasks/review |
task (via canonicalTaskTarget, classifier never called) |
team//tasks/review |
task (via canonicalTaskTarget) |
scripts/build.sh |
script |
Rejected (result.ok === false, errors[].code asserted; message not asserted):
uses: value |
Expected code |
Which rule produces it after P1a |
|---|---|---|
actions/checkout@v4 |
remote-action-acquisition-out-of-scope |
§4.3 step 8, locator shape |
./actions/review |
local-action-path-unsupported |
usesFailure prefix (./) |
docker://alpine:latest |
docker-action-unsupported |
usesFailure prefix (docker://) |
agents/reviewer |
non-executable-asset-ref |
usesFailure regex (agents/) |
workflows/child |
nested-workflow-unsupported |
§4.3 step 5 |
akm:commands/review |
unsupported-uses-target |
classifyTargetRef round-trip check → usesFailure |
bad.bundle//commands/review |
unsupported-uses-target |
classifyTargetRef round-trip check → usesFailure |
commands/review#fragment |
unsupported-uses-target |
classifyTargetRef fragment check → usesFailure |
actions/checkout@bad:ref |
unsupported-uses-target |
not locator-shaped (: in revision) → usesFailure |
review |
unsupported-uses-target |
classifyTargetRef reject → usesFailure |
usesFailure derives its message from the thrown cause, so the wrapped
message text changes (from the task-v3 trailing message to the
TARGET_REF_INVALID message). No test asserts that message — verified: the
contract table asserts code only, and the only test asserting the task-v3
trailing message asserts it against classifyTaskV3Uses directly (which is
untouched). Message drift here is authorized; code drift is not.
5. Behavior table (input → expected)
| # | Input | Expected after P1a | Lane |
|---|---|---|---|
| B-01 | Workflow step uses: tasks/nightly with with: {a: 1}, decoded |
Decodes with no error (unchanged, schema.ts:144) |
A |
| B-02 | Same step, frozen | UsageError, code COMPOSITION_INVALID, message Workflow step <id> cannot pass with: to task target tasks/nightly; task-call inputs are not supported yet.; exit 2 |
A |
| B-03 | Same step with with: {} (empty mapping), frozen |
Same rejection as B-02 (!== undefined, not "non-empty") |
A |
| B-04 | Same step without with:, frozen |
Freezes exactly as today — resolved dispatch unchanged from P0's pinned value | A |
| B-05 | Workflow step uses: akm/command with with: {content: …}, frozen |
Unchanged: consumed into the command target, content present in the frozen plan | A |
| B-06 | Workflow step with with: but no uses: |
Unchanged decode failure step <id> with is legal only with uses (schema.ts:393) |
A |
| B-07 | Workflow step uses: tasks/x with a non-scalar with value |
Unchanged decode failure from scalarRecord (schema.ts:389) — fires at decode, before freeze |
A |
| B-08 | Task document (tasks/*.yml) with with: on a command / script / workflow uses |
Unchanged: P-01 / P-02 / P-03 behavior from runtime-v3.ts, code INVALID_FLAG_VALUE |
— |
| B-09 | classifyTargetRef("commands/review") |
{ kind: "command", ref: "commands/review" }, frozen |
B |
| B-10 | classifyTargetRef("team//scripts/build.sh") |
{ kind: "script", ref: "team//scripts/build.sh" }, frozen |
B |
| B-11 | classifyTargetRef("tasks/nightly") / ("workflows/release") |
{ kind: "task", … } / { kind: "workflow", … }, frozen |
B |
| B-12 | classifyTargetRef("akm/command") |
throws UsageError TARGET_REF_INVALID (no builtin special case) |
B |
| B-13 | classifyTargetRef("commands/review#fragment"), ("akm:commands/review"), ("bad.bundle//commands/review"), ("agents/reviewer"), ("review"), (""), (" commands/review "), ("owner/repo@v1"), ("docker://alpine:latest"), ("./x") |
throws UsageError TARGET_REF_INVALID, message Target ref "<value>" must be a canonical commands/, scripts/, tasks/, or workflows/ asset ref. |
B |
| B-14 | Every row of the §4.5 parity table, through compileGithubWorkflowSource |
Byte-identical ok / code outcomes to pre-P1a |
B |
| B-15 | Every row of the §4.5 parity table, through decodeWorkflowSourceIrV1 (strict decode / Markdown path) |
Same classification outcome as B-14 (both defaults are now one function) | B |
| B-16 | classifyWorkflowStepUses("tasks/build", spy) |
{ kind: "task", ref: "tasks/build" }; spy not called |
B |
| B-17 | classifyWorkflowStepUses("commands/review", spy) where the spy returns {kind:"command", ref} |
Spy called once with "commands/review"; its return passed through |
B |
| B-18 | Task v3 source with neither akm.schedule nor on: |
UsageError, code TASK_SOURCE_INVALID, message unchanged: Invalid task v3 source at <path>:1: $ must declare exactly one scheduling source: akm.schedule or on.; exit 2 |
— |
| B-19 | Task v3 source with both scheduling sources | Same as B-18 (byte-identical message, same new code) | — |
| B-20 | classifyTaskV3Uses("review") (direct call) |
Unchanged: UsageError INVALID_FLAG_VALUE with the trailing task-v3 message — not re-coded in P1a |
— |
| B-21 | Task document uses: actions/checkout@v4, prepared |
Unchanged R-04(b): UsageError INVALID_FLAG_VALUE, GitHub action "actions/checkout@v4" is recognized but remote action acquisition is unsupported in 0.9.2. |
— |
| B-22 | Task v3 source with version: 2 |
Unchanged: TASK_SCHEMA_VERSION_UNSUPPORTED + TASK_V2_MIGRATION_HINT |
— |
| B-23 | Any of the five new codes reaching the CLI | {ok:false, error, code} on stderr, exit 2 (all are UsageError) |
— |
| B-24 | new UsageError("x", "<new code>").hint() for each of the five |
Returns the §2.1 hint string | — |
6. Per-lane file lists
Lane 0 — diagnostics (lands first; both lanes rebase on it)
| File | Change |
|---|---|
src/core/errors.ts |
Five codes appended to UsageErrorCode (:66-93); five USAGE_HINTS entries (:132-147) |
src/tasks/source-v3.ts |
sourceError (:209-226) code INVALID_FLAG_VALUE → TASK_SOURCE_INVALID. Nothing else in this file. |
tests/integration/tasks-scheduling-characterization.test.ts |
Authorized flip, §7 row F-02 |
tests/integration/cli-errors.test.ts |
Add hint coverage for the five new codes (B-24) — additive only; existing assertions untouched |
Lane A — with-rejection
| File | Change |
|---|---|
src/workflows/ir/source-freeze-v4.ts |
Guard at the head of taskDispatch (:211-217). No other edit; :145-151 and :220-222 untouched. |
tests/workflows/characterization-with-drop.test.ts |
Authorized flip, §7 rows F-01a/F-01b; new B-03 (empty with: {}) coverage |
docs/reference/workflow-schema.md |
§8 |
CHANGELOG.md |
§8 |
docs/migration/v0.9.1-to-v0.9.2.md |
§8 (file does not exist yet — create it) |
Lane B — classifier seam
| File | Change |
|---|---|
src/execution/target-ref.ts |
new — classifyTargetRef, TargetRefKind, ClassifiedTargetRef (§4.1) |
src/workflows/source-ir/uses.ts |
Local WorkflowSourceUsesTarget union; classifyWorkflowSourceUses layers builtin over classifyTargetRef; docstring rewritten; no src/tasks/** import (§4.2) |
src/workflows/source-ir/semantics.ts |
Evaluation order §4.3; new local isGithubLocatorShape; usesFailure (:229-239) unchanged; no src/tasks/** import |
src/workflows/source-ir/compile.ts |
:41 default → classifyWorkflowSourceUses; :42 unchanged (§4.4) |
tests/execution/target-ref.test.ts |
new — B-09…B-13 unit coverage |
tests/architecture/diagnostic-codes.test.ts |
new — §9 ratchet + import-seam assertion |
Untouched by both lanes (assert this in review): src/tasks/runtime-v3.ts,
src/tasks/runner.ts, src/workflows/source-ir/github-yaml.ts,
src/workflows/source-ir/schema.ts, classifyTaskV3Uses itself.
7. Authorized test flips
Only the rows marked FLIP may change. Every other test in the repo must stay green UNCHANGED. A test that goes red and is not in this table is a regression, not a decision — stop and re-read §4.3/§4.4 before editing it.
| # | Test file:line | What it pins today | P1a action |
|---|---|---|---|
| F-01a | tests/workflows/characterization-with-drop.test.ts:161 — R-01(c) a task-composed step freezes byte-identically whether or not with: is authored… |
Freeze-equality of the with: and no-with: halves |
FLIP. The with: half must now assert the COMPOSITION_INVALID rejection (type + code + exact message, §3.1). |
| F-01b | same test, no-with: half |
The with-free fixture freezes |
KEEP GREEN inside the flipped test — the without-block still freezes, unchanged (B-04). Assert both halves in the rewritten test. |
| F-02 | tests/integration/tasks-scheduling-characterization.test.ts:49 and :64 (R-06 neither-case and both-case) |
UsageError + INVALID_FLAG_VALUE + EXACTLY_ONE_SCHEDULING_SOURCE |
FLIP the CODE assertion only → TASK_SOURCE_INVALID. The EXACTLY_ONE_SCHEDULING_SOURCE message constant and the instanceof UsageError assertion are unchanged. |
| F-03 | tests/workflows/characterization-classification.test.ts:149 and :162 — the two classifyWorkflowStepUses delegation tests (spy-based) |
Task-ref priority (spy not called) and delegation of non-task refs | KEEP UNCHANGED. The injected-classifier parameter is retained (§4.3), so both stay green as written. Authorized (P0-recorded advisory) to be rewritten as observable-result assertions through the new seam if the implementer finds them brittle — but rewriting is not required and should be avoided. |
Examined and not flipped (verified against the §2.2 scope boundary)
| Test file:line | Pins | Why it does not flip |
|---|---|---|
tests/workflows/characterization-with-drop.test.ts:64,:72,:97,:103 |
R-01(a) decode acceptance, R-01(b) scalarRecord + with is legal only with uses |
Rejection moved to freeze; decode is unchanged (B-01/B-06/B-07) |
tests/workflows/characterization-with-drop.test.ts:116,:252 |
R-01(d) builtin-command with consumption |
source-freeze-v4.ts:145-151 untouched (B-05) |
tests/workflows/characterization-classification.test.ts:91-114 (asserts at :94, :108) |
R-04(a) trailing classifyTaskV3Uses message + INVALID_FLAG_VALUE |
Not funnel-produced — a direct new UsageError(…) at source-v3.ts:587-590, outside the §2.2 re-code scope. Stays INVALID_FLAG_VALUE; P4 deletes it. (This site was listed as a flip candidate; inspection shows no flip is needed.) |
tests/workflows/characterization-classification.test.ts:116 (assert at :134) |
R-04(b) prepare-time rejection (runtime-v3.ts:366-371) |
Runtime module, not the source funnel |
tests/workflows/characterization-classification.test.ts:191 |
R-04(c) remote-action-acquisition-out-of-scope via the default classifier |
Parity requirement — §4.3 step 8 keeps it green |
tests/workflows/characterization-classification.test.ts:180 |
R-03 site 1 nested-workflow-unsupported |
semantics.ts:141-146 unchanged |
tests/workflows/characterization-classification.test.ts:217 (assert at :243) |
R-03 site 2 INVALID_FLAG_VALUE at source-freeze-v4.ts:220-222 |
Untouched; fixture authors no with: (§3.3) |
tests/workflows/characterization-classification.test.ts:527 (assert at :532) |
R-05(b) multi-job INVALID_FLAG_VALUE |
source-freeze-v4.ts:105-110 untouched; flips in P4 |
tests/workflows/source-ir-contract.test.ts:429-457 |
The §4.5 parity table (10 rejection codes + 6 acceptances) | Parity gate. Must stay green unchanged — this is Lane B's primary acceptance evidence. |
tests/integration/tasks-with-classification-characterization.test.ts:77,:98,:155 |
P-01 / P-02 / P-04, runtime-v3.ts |
Task layer, not the source funnel (B-08) |
tests/integration/commands/tasks-cli-envelope.test.ts:113,:136 |
{ok:false, code:"INVALID_FLAG_VALUE"} for task run commands/nightly and a non-task adapter claiming a task-shaped ref |
Ref-projection / adapter-boundary rejections, not parseTaskV3Yaml. Implementer must re-verify by running the file after the Lane 0 commit; if either turns red, the error came through the funnel and the row becomes an authorized code-only flip to TASK_SOURCE_INVALID (record it in the Review log). |
tests/integration/commands/tasks-lifecycle.test.ts:217 |
akmTasksAdd workflow+engine rejection |
Add-path flag validation. Same re-verify instruction as the row above. |
tests/tasks-task-id.test.ts:36,:54,:68, tests/tasks-parse-ref.test.ts:50, tests/integration/tasks-run-attempt-observability.test.ts:168,:172, tests/integration/commands/tasks-bundle-target.test.ts:226, tests/integration/okf-conformance.test.ts:717 |
INVALID_FLAG_VALUE on task ids / refs / run args |
task-id.ts, parse-ref, scheduler-binding — none reach sourceError |
tests/integration/cli-errors.test.ts:202,:217, tests/completions.test.ts:227 |
INVALID_FLAG_VALUE hint and flag-value envelopes |
Generic CLI surface; the INVALID_FLAG_VALUE hint is unchanged |
tests/tasks/migrate-v2-to-v3.test.ts, tests/integration/migrate-format.test.ts, tests/migrate/task-v2-to-v3-files.test.ts |
task-v2 rejection + migration hint | taskV2UnsupportedError already carries TASK_SCHEMA_VERSION_UNSUPPORTED (source-v3.ts:49-57) — never INVALID_FLAG_VALUE, so the funnel re-code cannot reach it |
8. Docs that ride with the code
Docs-only commits skip CI (.github/workflows/ci.yml ignores docs/**,
CHANGELOG.md), so these edits must land in the same commits as their code.
| File | Required content |
|---|---|
docs/reference/workflow-schema.md |
In the step-uses/with reference (near the with: example at :58 and the grammar section at :278): with: on a task-step ref is an error — code COMPOSITION_INVALID, exit 2 — pending task-input support in a later 0.9.x release. with: on uses: akm/command is unchanged and still required for the builtin action. |
CHANGELOG.md |
Under ## [Unreleased], a ### Breaking changes & migration section with two entries: (1) a workflow step that passes with: to a tasks/<ref> target is now rejected (COMPOSITION_INVALID) instead of having the mapping silently dropped — remove the block, or wait for task-call inputs; (2) task-source validation errors now report code TASK_SOURCE_INVALID instead of INVALID_FLAG_VALUE — scripts that branch on the code field of the JSON error envelope must be updated; messages and exit code 2 are unchanged. |
docs/migration/v0.9.1-to-v0.9.2.md |
Create (it does not exist). Carries the with-rejection note: what breaks, the exact new message, and the fix (delete the with: block from task steps). |
9. Ratchet — tests/architecture/diagnostic-codes.test.ts (new)
Mirrors tests/architecture/src-fn-size-ratchet.test.ts in style: a hardcoded
baseline number with a comment stating the rule, and a failure message that
tells the next author what to do.
Assertion 1 — the code ratchet. Count occurrences of the literal string
INVALID_FLAG_VALUE across all files under src/tasks/** and
src/workflows/**; assert count <= INVALID_FLAG_VALUE_BASELINE.
- Pre-P1a measurement (
grep -rn "INVALID_FLAG_VALUE" src/tasks/ src/workflows/ | wc -l, measured 2026-08-26 at branch head): 83 —runner.ts1,runtime-v3.ts11,schedule.ts12,scheduler-binding.ts10,scheduler-sync.ts3,source-v3.ts12,task-id.ts7,ir/environment-v4.ts3,ir/freeze-v4.ts2,ir/params.ts2,ir/source-freeze-v4.ts9,runtime/runs.ts2,runtime/workflow-asset-loader.ts4,source-files.ts5. - Expected post-P1a: 82 (the one funnel literal at
source-v3.ts:225becomesTASK_SOURCE_INVALID; Lane A addsCOMPOSITION_INVALID, notINVALID_FLAG_VALUE). Re-measure during implementation and hardcode the measured number — do not copy 82 on faith. - Required comment on the constant: the baseline only ever declines. A later
phase that re-codes more sites lowers it; nothing may raise it. Raising it
means new
INVALID_FLAG_VALUEthrows were added to task/workflow code, which is exactly what this ratchet exists to prevent.
Assertion 2 — the classification import seam. Assert that the source text of
src/workflows/source-ir/semantics.ts and src/workflows/source-ir/uses.ts
contains no import from tasks/source-v3, and that
src/workflows/source-ir/compile.ts's only tasks/source-v3 import binding is
classifyTaskV3Triggers. This is what keeps §4.2/§4.4 from silently regressing.
The new test file carries the MPL-2.0 header
(scripts/lint-license-headers.ts) and uses no env/cwd/fetch mutation
(scripts/lint-tests-isolation.ts).
10. Acceptance criteria
- [ ]
src/core/errors.tsdeclares all five newUsageErrorCodemembers and aUSAGE_HINTSentry for each, with the §2.1 strings. - [ ]
sourceError(src/tasks/source-v3.ts:209-226) throwsTASK_SOURCE_INVALID; its rendered message is byte-identical to pre-P1a. - [ ] No other
INVALID_FLAG_VALUEinsrc/tasks/source-v3.tswas re-coded (classifyTaskV3UsesandtaskV2UnsupportedErroruntouched). - [ ]
taskDispatch(src/workflows/ir/source-freeze-v4.ts:211) rejectssource.with !== undefinedwithCOMPOSITION_INVALIDand the exact §3.1 message, beforeresolveOwnedAsset. - [ ]
schema.ts:144and thescalarRecord/with is legal only with usesguardrails are unchanged; R-01(a) and R-01(b) are green unchanged. - [ ]
source-freeze-v4.ts:145-151(builtin-commandwithconsumption) is unchanged; R-01(d) is green unchanged. - [ ]
src/execution/target-ref.tsexists, exportsclassifyTargetRef, contains no GitHub-locator grammar and noakm/commandbranch, and returns frozen objects. - [ ]
src/workflows/source-ir/semantics.tsandsrc/workflows/source-ir/uses.tsimport nothing fromsrc/tasks/source-v3.ts;src/workflows/source-ir/compile.tsimports onlyclassifyTaskV3Triggersfrom it. - [ ]
compile.ts:41's default classifier isclassifyWorkflowSourceUses; both classification entrypoints (§4.4) resolve to the same function. - [ ] Every row of the §4.5 parity table produces its listed outcome, through both entrypoints (B-14, B-15).
- [ ]
tests/workflows/source-ir-contract.test.tsis green unchanged. - [ ] Exactly the flips in §7 (F-01a/F-01b, F-02) are present in the test diff.
git diff --statovertests/shows no other pre-existing test file modified, except any row the implementer re-verified into the table per thetasks-cli-envelope/tasks-lifecycleinstruction (recorded in the Review log). - [ ]
tests/execution/target-ref.test.tscovers B-09…B-13, including the frozen-result assertion and each rejection shape. - [ ]
tests/architecture/diagnostic-codes.test.tsexists, hardcodes the measured post-P1a baseline with the only-ever-declines comment, and carries the §9 import-seam assertion. - [ ]
docs/reference/workflow-schema.md,CHANGELOG.md, and the newly createddocs/migration/v0.9.1-to-v0.9.2.mdcarry the §8 content, committed with their code (never as a docs-only commit). - [ ] No exit-code test changed;
COMPOSITION_INVALIDandTASK_SOURCE_INVALIDboth exit 2. - [ ] Every new test file carries the MPL-2.0 header;
bun scripts/lint-license-headers.tsandbun scripts/lint-tests-isolation.tspass. - [ ]
bunx biome check --write src/ tests/produces no further changes;bunx tsc --noEmitis clean. - [ ]
bun run checkpasses (lint + typecheck +test:unit+test:integration). - [ ] Any behavior difference discovered during implementation that is not authorized by §4.5 (Accepted deviation A-1) or §7 is recorded in the Review log and not silently absorbed.
Review log
-
2026-08-26 —
isGithubLocatorShaperepository/revision charsets regressed 5 values in the direction A-1 does NOT authorize; fixed in code, not merely documented. A code-review pass on Lane B diffed pre-P1a classification (classifyTaskV3Uses+ the oldsemantics.tsdelegator) against post-P1aisGithubLocatorShapeover a 50-value corpus. The function as first implemented matched §4.3's prescribed regexes verbatim (GITHUB_LOCATOR_OWNER_REPO_SEGMENT_REapplied to both the owner and repository segments; a strict/^[A-Za-z0-9._/-]+$/revision allowlist) — but that prescribed rule itself was narrower than the grammar it replaces for two of its three components:owner/.github@v1,owner/_repo@v1,owner/-repo@v1(oldGITHUB_REPOSITORY = /^[A-Za-z0-9_.-]+$/allows a leading./_/-in the repository segment; the shared segment regex required a leading alphanumeric) andowner/repo@v1.0+meta,owner/repo@%40(oldvalidGithubRevisiononly forbids~ ^ : ? * [ \and../@; the strict allowlist rejects+and%). All five wereremote-action-acquisition-out-of-scopebefore P1a and had regressed tounsupported-uses-target— the opposite of what Accepted deviation A-1 authorizes (A-1 covers only previously-unsupported-uses-targetvalues becomingremote-action-acquisition-out-of-scope, never the reverse). Fixed by giving the repository segment its own/^[A-Za-z0-9_.-]+$/check (mirroringGITHUB_REPOSITORYexactly, distinct from the intentionally-looser owner check) and replacing the revision allowlist with a forbidden-character check mirroringvalidGithubRevision(src/tasks/source-v3.ts:497-520) — seeisGithubLocatorShapeandisGithubLocatorRevisionShapeinsrc/workflows/source-ir/semantics.ts, and the corrected §4.3 rule text above. Verified against the same corpus post-fix: all five now match pre-P1a exactly (remote-action-acquisition-out-of-scope); the five values that move in A-1's authorized direction (owner/repo/../x@v1,owner/repo/.@v1,o.wner/repo@v1,own_er/repo@v1, a >39-char owner — all reachable only through the intentionally-loose owner segment or the intentionally-shape-only path-segment handling) are unaffected and remain covered by A-1 as originally scoped. No pinned test in §4.5 ortests/execution/target-ref.test.tsexercises any of these ten values, so this was a silent, undetected divergence, not a test failure —bun run checkwas green both before and after this fix. -
2026-08-26 — CHANGELOG's
TASK_SOURCE_INVALIDentry overclaimed its scope and omitted ahint/detailbehavior change; corrected in docs, not code. A code-review pass on the Lane 0 diagnostics + CHANGELOG diff found the[Unreleased]"Breaking changes & migration" entry forTASK_SOURCE_INVALIDinaccurate in two ways. First, it read as if every task-source validation error now reports the new code, but per §2.2's scope boundary only thesourceErrorfunnel re-codes; the seven directUsageErrorthrows inparseTaskV3Yaml/yamlAstError(src/tasks/source-v3.ts:810,896,901,911,919,926,936— YAML syntax, size, structure, and expansion failures) render the identicalInvalid task v3 source at …message prefix but keepINVALID_FLAG_VALUEin 0.9.2, so a script branching only on the new code would miss the most common failure shape. Second, the entry's "messages … are unchanged" claim is true only of the JSON envelope'serrorfield:taskV3SourceErrorDetail(source-v3.ts:884) appendshint()to the message, so for a re-coded (sourceError-funneled) failure the envelope's separatehintfield (src/cli/shared.ts:113) and the composeddetailstring thatakm lint(src/core/adapter/adapters/akm-lint.ts:328), the akm-task adapter (src/core/adapter/adapters/akm-task-adapter.ts:104), and scheduler-sync (src/tasks/scheduler-sync.ts:711) surface for every invalid task YAML all changed — from… Run \akm--help` to see accepted values. to… Fix the task source at the reported path and line, then re-run.— a delta no test or golden pins, so it was invisible inbun run check. Neither finding required a source change: §2.2's scope boundary ("thesourceErrorfunnel and nothing else") is binding for P1a, so the seven direct-throw sites stayINVALID_FLAG_VALUE`. The CHANGELOG entry was rewritten to state both qualifications explicitly instead of silently absorbing the gap, per §10's final acceptance criterion.
2026-08-26 — phase close-out (orchestrator). Test review: clean after 2 rounds. Code review ran its full 3-round budget (3 → 1 → 2 CONFIRMED) without a confirming fourth round; the orchestrator adjudicated per the dispute rule. Round 1 found real code issues (type widening at the classifier seam, a five-value locator parity regression, a falsified workflow-schema.md sentence) — fixed in round 1's fix commit, with the parity fix independently differentially fuzzed by the round-2 reviewer over 353k locator-shaped values (sole divergence class: the A-1-authorized widening). Round 2's single finding (ratchet file describing itself as unmeasured/RED) and round 3's two findings (CHANGELOG overclaiming the TASK_SOURCE_INVALID scope; the unrecorded hint/detail drift) were each fixed in their round's fix commit — verified directly against the reviewers' required changes. Severity converged monotonically (code → stale comments → changelog wording), so the budget expiry is a bookkeeping artifact, not an unresolved defect.
Applied at close-out (round-2 advisory): the ten locator parity values are now pinned in
tests/execution/target-ref.test.ts's rejection table (five regression values + five A-1
widenings, all remote-action-acquisition-out-of-scope, live-verified before pinning).
Advisories carried to later phase specs: import-seam ratchet hardening against namespace-import /
re-export / dynamic-import evasion (P4a spec — the seam must hold through the grammar deletion);
CLI envelope + exit-2 coverage for COMPOSITION_INVALID / TASK_SOURCE_INVALID (P1b spec);
lint-stage diagnostic for with: on task steps — the freeze-stage rejection is invisible to
akm lint, workflow authoring, and scheduler-sync (P2b spec, where bindings land);
WORKFLOW_SOURCE_INVALID's hint names akm workflow validate, which does not exist — the phase that
wires the code retargets the hint at akm lint or ships the verb (P3 specs);
SAFE_TASK_ATTEMPT_ERROR_CODES allowlist tracks the new codes when P1b splits the runner;
with: on uses: commands/<ref>/scripts/<ref> is still accepted-then-ignored (P0-unpinned,
outside R-01) — P2b's binding work must close it and the docs note it until then.
2026-08-26 — phase gate green. lint green, tsc green, unit 3908 pass / 0 fail (291 files), integration 5622 pass / 57 skip / 0 fail (419 files), with the close-out parity pins included. P1a is complete.
2026-08-28 — P4 deletion note (row F-02 / §6 Lane 0 file list): the canary this row flips no
longer exists. tests/integration/tasks-scheduling-characterization.test.ts — named above as the
FLIP target for F-02 (the INVALID_FLAG_VALUE → TASK_SOURCE_INVALID code-only flip on R-06's
neither-case and both-case) and listed in §6's Lane 0 file table — was deleted by commit
0969162 ("refactor(p4): remove task source v3 acceptance from src", spec
docs/plans/specs/p4-deletions-closeout.md §3.2, its own row F-A2.7: "DELETE — all three tests
are R-06, resolved by deletion (§5.5)"). This is a downstream consequence of R-06 itself, not a
regression of F-02: R-06 ("task v3 requires exactly one scheduling source") is a v3 rule, and P4
§3.2 removed task v3 acceptance from src entirely, so the rule the deleted file's three tests
pinned has no document shape left to apply to — docs/plans/specs/p0-invariants.md's final
disposition table records it as "RESOLVED by deletion." F-02's own flip is not falsified by this:
it landed as specified and stayed byte-unchanged through P2a's own close-out sweep
(docs/plans/specs/p2a-task-source-v4.md:1318-1326, "Every canary named in §7 —
… tests/integration/tasks-scheduling-characterization.test.ts … — is byte-unchanged from P1b's
head … to this commit; no test file under tests/ was deleted anywhere in the phase") before P4
deleted the file three phases later. Where the coverage lives now: nowhere, by design — P4's
F-A2.7 disposition is deletion, not relocation; no replacement test exists for the neither-case,
both-case, or akm.schedule success-shape assertions, because task source v4 (the only schema src
still accepts) makes schedule: optional rather than exactly-one-required (P2a §1.5 D2-N6), so the
behavior itself has no live subject to characterize. This entry is prose-only; the row's F-02 text
above is not edited — it is historically accurate for P1a's own scope and is superseded, not wrong.
2026-08-28 — P2b supersession note (§3.1's pinned message, rows B-02/B-03). §3.1's code block and
"pinned message contract" (Workflow step <id> cannot pass with: to task target <ref>; task-call inputs are not supported yet.) and the behavior table's B-02/B-03 rows above describe an
unconditional rejection: any authored with: on a uses: tasks/<ref> step throws
COMPOSITION_INVALID, regardless of what the target declares. That is no longer what the shipped
code does. docs/plans/specs/p2b-input-bindings.md (§1.7 A-N5/A-N6, §3) turned the unconditional
rejection into real typed bindings: src/workflows/freeze/targets/task.ts's taskDispatch now
parses the composed target's own inputs: contract and only raises COMPOSITION_INVALID (via the
renamed noDeclaredInputsError) when the target declares no inputs at all — if (source.with !== undefined && contract === undefined). The pinned message text changed with it: `Workflow step ${stepId} cannot pass with: to task target ${ref}; ${ref} declares no inputs.` (verified at HEAD,
src/workflows/freeze/targets/task.ts:32-37,117-119 — note the message now also names the target
ref a second time in place of the retired "task-call inputs are not supported yet." clause). When
the target DOES declare inputs:, an authored with: binds through freezeTaskInputBindings
instead of being rejected at all. This entry is prose-only; §3.1's code block and the B-02/B-03 rows
above are not edited — they are historically accurate for P1a's own scope (the fail-closed gate that
existed before task input bindings shipped) and are superseded, not wrong. The corresponding test
(tests/workflows/characterization-with-drop.test.ts:169, R-01(c) — this spec's own F-01a flip
target) was authorized to flip again, message-bytes-only, under P2b's own §7 flips table
(docs/plans/specs/p2b-input-bindings.md:1174); see that spec for the current pinned assertions.