Files
notarius/docs/roadmap/implementation.md

32 KiB

Domain-Typed Pipeline Implementation Roadmap

Status

The domain-typed pipeline and the first five remediation stages were completed on 2026-07-17. A subsequent review identified three additional implementation issues. Stages 1 through 5 are retained as the completed implementation record; Stages 6 through 8 are the decision-complete plan for the remaining work.

Implement the pending stages in order. Keep each stage independently reviewable and leave the repository passing its full validation suite before beginning the next stage. Unless a stage explicitly says otherwise, preserve public CLI, configuration, module-key, validator-key, durable output, and checkpoint file contracts.

Current architectural policy is defined by Architecture. The design context remains in the feature roadmap and ADR-0002, ADR-0003, and ADR-0004.

Delivered Baseline

The existing implementation already provides:

  • domain-first production extensions and package-family registrars;
  • engine-owned source provenance using source.SourceRef and workspace schema v2;
  • one domain-owned Go type through extraction, merge, normalization, and typed validation;
  • schema-aware artifact codecs at checkpoint, debug, and output boundaries;
  • pre-execution typed resolution, option validation, and preparation;
  • one shared scheduled LLM client; and
  • bounded chunk-first, lane-second extraction with deterministic public outcomes.

Stages 1 through 5 corrected checkpoint identity, completed merge and normalize retry debugging, removed kind-ambiguous registry lookup, strengthened the initial architectural guard, and removed the obsolete sequential extraction path. Stages 6 through 8 complete candidate-artifact handling, make attempt debugging comprehensive across all stages, and close the remaining import-guard loopholes.

Implementation Rules

Apply these rules to every stage:

  1. Read the task-specific documents listed in Development before changing the affected subsystem.

  2. Add focused regression tests that fail against the pre-stage code and pass after the change.

  3. Preserve deterministic behavior. Do not base digests, public ordering, reference resolution, or reported errors on Go map iteration or goroutine completion order.

  4. Keep framework-owned type erasure private. Do not reintroduce module-facing any, raw JSON handoffs, or a generic processing interface.

  5. Update current-behavior documentation in the same stage when externally observable or documented internal behavior changes. Follow Documentation Policy; do not describe a later stage as already implemented.

  6. Run focused tests while iterating. At the end of every stage, run:

    go test ./...
    go test -race ./...
    go vet ./...
    go build ./cmd/notarius
    git diff --check
    

Completed Remediation Record

The following five stages have been implemented and are retained to document the decisions and completion criteria that produced the current baseline.

Stage 1: Include Resolved Validator Policy In Pipeline Identity

Goal

Ensure that every effective validator-chain change alters the resolved pipeline digest and therefore prevents reuse of checkpoints created under a different validation policy.

Required Changes

  1. Extend resolvedPipelineDigest in internal/framework/pipeline/profile.go to include the complete resolved validator-chain collection. Include, for every chain:

    • stage, lane ID, and owning module key;
    • validator order;
    • each complete resolved binding, including module key, LLM profile, retries, options, and references;
    • validator execution class;
    • validator target; and
    • artifact kind.
  2. Hash the already canonical ResolvedPipeline.ValidatorChains order produced by resolution. Do not independently sort validators or otherwise weaken configured order. Continue excluding only the digest field itself.

  3. Use the same deterministic JSON-and-SHA-256 mechanism as the existing pipeline digest. Go map keys encoded inside bindings must retain the deterministic ordering supplied by encoding/json.

  4. Do not add validator fingerprints separately to individual stage dependencies. Pipeline/workspace identity is the authoritative invalidation boundary for validation-policy changes.

  5. Do not advance the workspace schema version. The corrected digest naturally selects a different checkpoint directory; existing checkpoints remain intact and become reuse misses for the changed pipeline identity.

Tests

Add focused tests under internal/framework/pipeline and, where useful, internal/core/workspace proving that:

  • registering a different default chain changes the resolved pipeline digest;
  • adding, removing, or reordering a default validator changes the digest;
  • changing a resolved validator binding option or LLM profile changes the digest;
  • changing execution class, target, or artifact kind changes the digest;
  • resolving the same pipeline and catalog repeatedly produces the same digest;
  • an explicit empty override produces an empty resolved chain and differs from a non-empty inherited default; and
  • a changed validator-chain digest produces a different checkpoint identity, while an unchanged chain preserves it.

Retain the existing tests proving that the digest field itself is excluded and that artifact schema identity participates in the digest.

Documentation

Update the checkpoint invalidation description in docs/operations.md and the identity description in docs/internal/pipeline.md to state concisely that the effective resolved validator policy participates in pipeline identity.

Completion Gate

A checkpoint accepted under one resolved default or explicit validator chain must not be reusable after that chain changes.

Stage 2: Complete Attempt-Scoped Debugging For Merge And Normalize

Goal

Give merge and normalize retries the same attempt-level observability and nested LLM-call association already provided for chunk and extract attempts.

Required Changes

  1. In internal/framework/pipeline/runner_typed.go, create an attempt-specific debug context before invoking each merge or normalize module. Pass that context to both the module operation and its validation chain.

  2. Use these stable paths:

    merge/<lane-id>/attempt-<NN>.json
    merge/<lane-id>/attempt-<NN>/prompt-<NNNN>.json
    merge/<lane-id>/attempt-<NN>/response-<NNNN>.json
    merge/<lane-id>/attempt-<NN>/response-content-<NNNN>.<ext>
    
    normalize/<lane-id>/attempt-<NN>.json
    normalize/<lane-id>/attempt-<NN>/prompt-<NNNN>.json
    normalize/<lane-id>/attempt-<NN>/response-<NNNN>.json
    normalize/<lane-id>/attempt-<NN>/response-content-<NNNN>.<ext>
    

    Continue using two-digit retry attempt numbers and the existing debug path sanitization and LLM call numbering behavior.

  3. Write one module-attempt envelope for every attempted merge and normalize operation:

    • on success, record the codec-backed candidate artifact and warnings that will be promoted if the attempt is accepted;
    • on validator rejection, record the rejection without promoting discarded warnings;
    • on module, validation, serialization, or debug failure, record the error; and
    • in all cases, attach the LLM calls recorded in the module attempt scope.
  4. Keep validator-specific debug scopes under the existing validate/... hierarchy. A validator's LLM calls remain linked to its validator attempt; the merge or normalize module envelope links only calls made by that module attempt.

  5. Preserve the existing stage-level input.json and output.json artifacts. Checkpoint reuse should continue to emit stage-level artifacts but should not synthesize retry attempts that did not execute.

  6. Factor common attempt-envelope behavior into a small private helper where it prevents chunk, extract, merge, and normalize instrumentation from drifting. Do not introduce a new public runner abstraction solely for debugging.

  7. Treat any debug write failure as a framework error, consistent with current debug policy.

Tests

Add tests with instrumented LLM-backed fake mergers and normalizers proving that:

  • first-attempt success writes the expected attempt and nested LLM artifacts;
  • a failed first attempt followed by success writes two distinct attempt envelopes and associates each call with the correct attempt;
  • module errors, validator errors, and final rejection are represented in the corresponding attempt envelope;
  • validator LLM calls remain under validator paths rather than being attributed to the module scope;
  • discarded-attempt warnings are not promoted;
  • checkpoint reuse produces no module-attempt files; and
  • no merge or normalize call falls back to an unscoped stage-name debug path.

Extend the CLI debug integration test only as needed to verify the public debug directory layout. Keep most behavioral coverage in the pipeline package.

Documentation

Update docs/operations.md if necessary to show the stable merge and normalize attempt paths. Update docs/internal/pipeline.md only where its implementation description needs clarification; its existing attempt-scoping guarantee should become fully true rather than be weakened.

Completion Gate

Every executed merge and normalize retry must have a distinct debug envelope, and every LLM call made by that module attempt must be nested under and linked from that attempt.

Stage 3: Make Typed Variant Spec Lookup Kind-Specific

Goal

Ensure that CLI reference-target discovery and other behavior-sensitive lookup select the merger or normalizer specification for the lane's actual artifact kind, never an arbitrary map entry.

Required Changes

  1. Add explicit kind-specific lookup methods to MergerRegistry and NormalizerRegistry, named:

    SpecForArtifact(key string, kind contracts.ArtifactKind) (ModuleSpec, bool)
    

    Normalize the key and artifact kind in the same way as typed registration and return a cloned spec.

  2. Change CLI reference-target discovery in internal/cli/run.go to retain the selected extractor's declared artifact kind and use it for merger and normalizer spec lookup in that lane.

  3. A missing typed variant must produce a deterministic error naming the pipeline, lane, stage, module key, requested artifact kind, and sorted registered kinds, matching the quality of full pipeline resolution errors.

  4. Keep Spec(key) for kind-neutral catalog inspection and compatibility with existing callers, but remove its map-order dependence. Select the first registered artifact kind in sorted order before cloning its spec. Add a comment making clear that behavior-sensitive code must use SpecForArtifact.

  5. Do not require typed variants under one reusable key to expose identical reference slots or capabilities. Their kind-specific specs are allowed to differ, and resolution must consistently choose the matching variant.

  6. Audit all merger and normalizer Spec callers. Convert any caller making a lane-specific decision to SpecForArtifact; leave only catalog or display callers on the kind-neutral method.

Tests

Register at least two artifact-kind variants under the same merger key and the same normalizer key with intentionally different reference slots. Prove that:

  • SpecForArtifact returns the correct cloned variant;
  • lookup is stable regardless of registration order;
  • CLI qualified and unqualified reference discovery uses the selected lane's variant;
  • a reference accepted by one variant is not incorrectly accepted for another;
  • a missing variant reports sorted available kinds; and
  • repeated Spec(key) calls return the same deterministic catalog result.

Retain coverage for the existing single-variant D&D catalog behavior.

Documentation

Update docs/internal/pipeline.md or docs/internal/modules.md only if either currently describes registry lookup mechanics. No CLI or configuration syntax change is intended.

Completion Gate

No behavior-sensitive merger or normalizer lookup may depend on map iteration, and CLI reference resolution must use the artifact kind selected by the lane's extractor.

Stage 4: Enforce Domain Import Boundaries Generically

Goal

Make the repository guard enforce ADR-0004 for present and future module families without a hardcoded list of domain names or domain pairs.

Required Changes

  1. Refactor internal/modules/import_boundaries_test.go so that a module family is derived from the first path segment under internal/modules/, rather than recognized by a fixed isDomain list.

  2. Treat generic as the domain-neutral reusable extension family. Treat every other production family, including dnd, seriatim, and future families, as concrete for import-boundary purposes.

  3. Enforce these rules:

    • a family root package must not import its own child implementation or registrar packages;
    • child packages within the same concrete family may import the family root, shared helpers, or sibling implementations when needed;
    • a concrete family must not import another concrete family;
    • generic packages must not import concrete families;
    • concrete implementation packages must not import generic implementation packages directly;
    • internal/modules/<family>/register may import its own family packages and generic packages to compose typed strategies;
    • the generic registrar may import generic child packages; and
    • the application composition root and designated black-box integration tests may compose multiple families.
  4. Exclude internal/modules/integration from production-family discovery. Preserve its exemption only for _test.go black-box composition files; do not create a blanket exemption for production Go files.

  5. Apply production import rules to white-box tests located in concrete module packages. If an existing test composes concrete and generic implementations, move that cross-family coverage to internal/modules/integration or replace the foreign implementation with a package-local test double. Do not exempt arbitrary _test.go files merely because they are tests.

  6. Keep the test based on parsed Go imports. Do not add a new build tool or external dependency for this guard.

Tests

Expand the table-driven import tests to cover:

  • a hypothetical future concrete family, demonstrating that no code change is needed to enforce its boundaries;
  • generic-to-concrete rejection for both current and hypothetical families;
  • concrete-to-peer-concrete rejection;
  • concrete implementation-to-generic rejection;
  • concrete registrar-to-generic acceptance;
  • family-root-to-child rejection;
  • child-to-family-root and same-family sibling acceptance;
  • application composition-root acceptance; and
  • black-box integration-test acceptance without exempting non-test files.

Run the guard against the complete current repository and confirm that no production package must be moved to satisfy it.

Documentation

No ADR change is required. Update docs/internal/modules.md only if its package boundary description needs to name the registrar-only generic composition rule more clearly.

Completion Gate

Adding a new directory under internal/modules/<new-family> must automatically receive the same cross-family and registrar enforcement as current production families.

Stage 5: Remove The Obsolete Sequential Extraction Path

Goal

Make the runner's structure match the implemented concurrent architecture: the coordinator owns extraction, and a lane continuation begins from finalized extract results and performs only merge and normalize work.

Required Changes

  1. Replace runTypedLane with a continuation-oriented private function whose input explicitly contains the finalized extraction state needed by merge and normalize:

    • accepted typed extract artifacts;
    • serialized checkpoint artifacts;
    • accepted warnings;
    • rejected outputs; and
    • the original extract checkpoint decision.

    Use a private struct if it keeps the call boundary clear and avoids a long positional parameter list.

  2. Change continueLane in runner_concurrent.go to pass the finalized state directly. Do not make completed extraction look like checkpoint reuse.

  3. Remove:

    • completedExtractLoader;
    • the RunInput.extractDecision private override;
    • the sequential extraction branch formerly contained in runTypedLane; and
    • duplicate extraction retry, validation, checkpoint, and debug logic made unreachable by concurrent coordination.
  4. Keep extraction checkpoint loading and validation in prepareLaneExtract, extract execution in runExtractJob, and deterministic final aggregation in finalizeLaneExtract.

  5. Preserve existing behavior exactly:

    • extract checkpoint events report the real loader decision;
    • reused and freshly computed extract results enter continuation through the same typed state;
    • accepted artifacts reach merge in chunk-index order;
    • warnings and rejections retain stable ordering and promotion semantics;
    • a lane with no accepted extracts follows the current merge behavior;
    • stage-level extract input/output debug artifacts remain unchanged; and
    • failure classification and cancellation retain their deterministic scope.
  6. Keep merge and normalize serial within a lane and keep cross-lane continuations bounded by the existing worker policy. Do not introduce a goroutine per lane or chunk.

Tests

Add or adjust focused runner tests proving parity for:

  • fresh extraction, fully reused extraction, and a mixture of accepted and rejected chunk results;
  • deterministic extract ordering under reverse completion;
  • warning promotion across retries;
  • real checkpoint decision reporting for reused and recomputed extracts;
  • extract input/output and attempt debug paths;
  • framework cancellation and deterministic error selection; and
  • unchanged bounded worker and global LLM concurrency behavior.

Use source search or a package-local compile-time assertion where practical to confirm there is only one extraction execution path and no remaining completedExtractLoader or extractDecision compatibility shim.

Documentation

Update docs/internal/pipeline.md if its execution-flow description names the old continuation mechanism. This stage is an internal refactor and must not change operator or integration contracts.

Completion Gate

The concurrent coordinator must be the only code path that executes extraction, and lane continuation must consume finalized extraction state without routing it through a synthetic checkpoint loader.

Pending Remediation Plan

Implement Stages 6 through 8 in order. Do not perform final roadmap closeout until all three stages and their documentation changes have landed.

Stage 6: Validate Merge And Normalize Candidates Before Final Encoding

Goal

Preserve the two-zone artifact model at merge and normalize boundaries: a module result remains a candidate artifact until typed validation accepts it, and only an accepted artifact is encoded with the final artifact codec and eligible for checkpointing or downstream use.

Required Changes

  1. Refactor the merge and normalize attempt paths in internal/framework/pipeline/runner_typed.go so each attempt proceeds in this order:

    1. execute the typed module and collect its warnings;
    2. serialize the returned value with ArtifactCodec.EncodeCandidate for attempt-debug capture;
    3. execute the resolved typed validator chain against the in-memory value;
    4. on validator rejection, record the rejected candidate and finish the attempt without invoking ArtifactCodec.Encode or creating a checkpoint artifact; and
    5. only after validator acceptance, invoke ArtifactCodec.Encode through the normal checkpoint-artifact path.
  2. Use one private candidate-serialization helper for merge, normalize, and any equivalent pre-validation debug path. The helper must preserve the codec's media type and schema digest in the debug artifact, but its result must never be used as a final checkpoint or downstream artifact.

  3. Treat candidate-encoding and final-encoding failures as framework errors, not validator rejections. Record the applicable attempt envelope before returning the error using the existing attempt-debug behavior; Stage 7 will consolidate that behavior across every retrying stage.

  4. Preserve retry behavior:

    • a validator rejection may be retried according to the stage retry policy;
    • warnings from rejected attempts remain attempt-local unless existing warning-promotion policy says otherwise;
    • a final accepted value and its serialized checkpoint artifact are published only after the attempt has completed successfully; and
    • checkpoint identity, checkpoint event reporting, and durable artifact formats remain unchanged.
  5. Do not weaken a domain codec so that its final Encode accepts invalid domain values. In particular, keep the D&D spell codec's semantic checks at the final-artifact boundary and use EncodeCandidate for its intentionally incomplete pre-validation representation.

Tests

Add focused merge and normalize tests using an instrumented codec for which EncodeCandidate accepts an invalid value but Encode rejects it. Prove that:

  • the validator, rather than final encoding, determines that the candidate is rejected;
  • rejected candidates never call final Encode and never produce checkpoint artifacts;
  • an accepted candidate calls final Encode exactly once and produces the existing checkpoint representation;
  • candidate-encoding and post-acceptance final-encoding failures are reported as framework errors and have attempt debug records; and
  • retry warnings, rejection promotion, and checkpoint event ordering remain deterministic.

Retain or add a D&D spell regression test demonstrating that an incomplete candidate can reach validation without making the final spell codec permissive.

Documentation

Update docs/internal/pipeline.md to state explicitly that merge and normalize debug payloads may contain candidate encodings, while checkpoints and stage outputs always contain validator-approved final encodings. Update docs/internal/artifacts.md if it currently implies that every serialized debug artifact is a final Domain Artifact Zone representation.

Completion Gate

No merge or normalize code path may invoke final artifact encoding before typed validation acceptance, and a validator-rejected candidate must remain observable as a rejection even when the codec would refuse to encode it as a final artifact.

Stage 7: Record Every Executed Attempt And Its Terminal Outcome

Goal

Make attempt debugging complete and uniform across chunk, extract, merge, and normalize so every executed retry has one terminal attempt envelope, including validator rejection, module error, validator error, and serialization error.

Required Changes

  1. Introduce or consolidate a private attempt-terminal recorder used by all four retrying stages. It must write exactly one envelope per executed attempt and support these terminal outcomes:

    • accepted success;
    • validator rejection;
    • module execution error;
    • validator execution error;
    • candidate-serialization error; and
    • final-serialization error where final encoding occurs within the attempt.
  2. Preserve the existing debug directory layout and envelope schema. Populate the envelope consistently with the attempt number, warnings accumulated by that attempt, any available candidate payload or rejection details, and the terminal error text when an error occurred. A validator rejection without an execution error remains a rejection and must not acquire a synthetic error string.

  3. Complete the extract path in internal/framework/pipeline/runner_concurrent.go:

    • write an attempt envelope before returning a validator rejection or validator error;
    • write an attempt envelope before returning a final-serialization error; and
    • stop discarding errors returned while writing module-error attempt data.
  4. Complete the chunk path in internal/framework/pipeline/runner.go by recording validator execution errors in the envelope's error field and by propagating attempt-write failures on every terminal path.

  5. Route the existing merge and normalize attempt handling through the same terminal-recording behavior without changing their retry or validation semantics. Coordinate this work with Stage 6 so rejected candidates use candidate encoding and accepted values use final encoding.

  6. Keep LLM-call ownership scoped to the operation that made the call:

    • module calls belong to the module attempt envelope;
    • validator calls remain in their validator-specific debug scope; and
    • attaching a call to an attempt must not duplicate it in another module attempt or orphan it from the attempt that initiated it.
  7. Never silently discard a debug write failure. If another error already exists, return an errors.Join result that preserves both the primary error and the contextualized debug error. If no primary error exists, return the contextualized debug error. Do not replace a validator rejection with a framework error unless persisting its required debug record fails.

  8. Keep debug failures subject to the retry boundary already surrounding the relevant attempt; do not add a second retry loop specifically for debug persistence.

Tests

Add table-driven attempt-debug tests covering each terminal outcome for chunk and extract, and retain equivalent merge and normalize coverage. At minimum, prove that:

  • a first-attempt extract rejection followed by success writes both envelopes;
  • a terminal extract rejection, validator error, module error, candidate-codec error, and final-codec error each write an envelope with the correct fields;
  • a chunk validator execution error appears in its attempt envelope;
  • debug write failures are returned, and are joined with the primary error when both occur;
  • module LLM calls are attached to the correct attempt while validator LLM calls remain isolated; and
  • success, rejection, warning, and retry ordering remains deterministic.

Use a shared assertion helper to verify the invariant that the number and indices of attempt envelopes equal the attempts actually executed.

Documentation

Update docs/internal/pipeline.md and docs/operations.md so their attempt debug guarantees name all terminal outcomes and explain the separation between module-attempt and validator-call scopes. Do not promise that a payload exists when failure occurred before a candidate value was available.

Completion Gate

For chunk, extract, merge, and normalize, every entered attempt must leave exactly one terminal envelope or return an error that explicitly reports why that envelope could not be persisted. No attempt-related debug write error may be ignored.

Stage 8: Close Production Import-Guard Loopholes

Goal

Make the automated import guard enforce the complete domain-first dependency policy for production code, including the framework, the application composition root, and the integration-test-only package.

Required Changes

  1. Extend internal/modules/import_boundaries_test.go so validation considers both the importing file's repository-relative path and whether it is a test file. Do not return early merely because the importer is outside internal/modules.

  2. Treat internal/modules/integration as test infrastructure, not as a module family or a production dependency target:

    • reject every import of internal/modules/integration from a non-test Go file; and
    • continue allowing designated black-box tests under internal/modules/integration to import concrete families for composition coverage.
  3. Reject imports of any internal/modules/** package from non-test files under internal/framework/** or internal/core/**. These packages define inward framework and core layers and must remain independent of all module implementations, including generic.

  4. For non-test files under the production CLI composition root internal/cli/**, allow module imports only when the target is the exact family registrar package internal/modules/<family>/register. Reject direct imports of family roots, domain leaves, adapters, and generic implementation leaves. Apply the same rule automatically to newly added families.

  5. Preserve test-only assembly allowances needed for black-box and compatibility coverage. In particular, _test.go files in the CLI, framework, and core trees may import module packages, and designated integration black-box tests may compose concrete families. These allowances must be based on the importing file being a test file, not on a directory exemption that production files could inherit.

  6. Retain all Stage 4 within-module rules: concrete peer isolation, registrar-only concrete-to-generic composition, family-root direction, and same-family child relationships. Keep path parsing structural so new module families receive the rules without a hard-coded family list.

  7. Improve failure messages to name the importing file, import target, and the applicable boundary rule.

Tests

Add table-driven fixtures proving that the guard:

  • rejects production imports of internal/modules/integration from module and non-module packages;
  • rejects production framework and core imports of concrete and generic module packages;
  • accepts an exact registrar import from production CLI code;
  • rejects production CLI imports of a family root, concrete leaf, generic leaf, or other non-registrar module package;
  • accepts equivalent direct imports from _test.go compatibility tests;
  • accepts designated black-box integration-test composition while rejecting a production .go file in the same directory; and
  • retains every cross-family, registrar, and family-root case covered by Stage 4.

Run the guard against the full repository and ensure all current production imports comply without adding path-specific exemptions.

Documentation

Update docs/internal/modules.md to distinguish production composition through family registrars from test-only direct composition, and to state that internal/modules/integration is not a production dependency target. Update other current-behavior documentation only if it describes broader composition permissions.

Completion Gate

A new production file cannot bypass the domain-first dependency rules by being placed outside internal/modules, under internal/modules/integration, or in the CLI composition tree. Test-only allowances must remain explicit and file-scoped.

Final Verification And Closeout

After Stages 6 through 8:

  1. Run the full validation commands from this document on a clean worktree.
  2. Exercise the maintained D&D production example with both a fresh workspace and checkpoint resume.
  3. Verify that a merge or normalize candidate rejected by typed validation is not passed to final encoding or checkpoint creation.
  4. Verify that every executed chunk, extract, merge, and normalize retry has a terminal attempt envelope with correctly scoped LLM-call data.
  5. Confirm current production imports satisfy the strengthened framework, registrar-only CLI, integration-target, and module-family boundary rules.
  6. Re-read current-behavior documentation for statements made true or obsolete by these stages.
  7. Replace this roadmap's status with a concise completion record only after all stages and documentation updates have landed.

Open Questions

None. The implementation choices required for Stages 6 through 8 are specified above. Final roadmap closeout is intentionally deferred until those stages are complete.