Files
notarius/docs/roadmap/implementation.md

36 KiB

Domain-Typed Pipeline Implementation Plan

Purpose

This document is the executable implementation plan for the target state in the domain-typed pipeline feature roadmap. It assumes the decisions in ADR-0002, ADR-0003, and ADR-0004.

The intended operator is an LLM coding agent working through one stage per implementation prompt. Complete the stages in order. Each stage must leave the repository buildable and tested; do not defer a broken intermediate state to a later stage.

Implementation Rules

For every stage:

  1. Read docs/development.md and follow its task-specific reading guide. Read the current implementation and focused tests for every touched subsystem.
  2. Treat the feature roadmap as the canonical owner of desired behavior and this document as the canonical owner of task sequencing. Do not restate future behavior in current-behavior documentation before it exists.
  3. Preserve unrelated user changes. Use mechanical moves where possible so file history and test intent remain legible.
  4. Add focused tests with the change. Run those tests while iterating, then run go test ./..., go vet ./..., and go build ./cmd/notarius before ending the stage.
  5. Run go test -race ./... in stages that introduce or change concurrency and in the final stage.
  6. Update the canonical current-behavior documents in the same stage in which behavior changes. At minimum, reconsider docs/policy/architecture.md, docs/internal/overview.md, docs/internal/pipeline.md, docs/internal/modules.md, docs/internal/llm.md, docs/config.md, docs/operations.md, and docs/integrations/ according to the documentation policy; edit only the documents whose owned facts changed.
  7. Do not change user-visible module keys, validator keys, output paths, durable D&D JSON, prompt/schema identities, default chains, or rejection semantics unless this plan explicitly requires it.
  8. Stop after a stage if an exit criterion cannot be met. Record the concrete blocker rather than implementing a second architecture alongside this one.

Fixed Technical Decisions

The following choices are inputs to implementation, not questions to reopen in individual stages.

Source types

Keep the existing names source.SourceDocument, source.SourceUnit, and source.SourceRef. Add Ref source.SourceRef to SourceUnit. Move contracts.SourceChunk to internal/core/source as source.Chunk, with this logical shape:

type Chunk struct {
	ID        string
	SourceID  string
	Index     int
	Ref       SourceRef
	Content   []byte
	MediaType string
	Units     []SourceUnit
	Metadata  map[string]any
}

Remove StartUnitID and EndUnitID; Ref is the only chunk-boundary representation. A Seriatim unit's self-reference is {SourceID: document ID, StartUnitID: unit ID, EndUnitID: unit ID}. A chunk reference spans its first and last included units. Advance persisted workspace state from notarius.workspace.v1 to notarius.workspace.v2; v1 state is incompatible and must be recomputed, but never deleted automatically.

Artifact contracts

Place engine-owned artifact primitives with the other universal contracts under internal/framework/contracts:

type ArtifactKind string

type ArtifactSchema struct {
	ID         string
	Name       string
	Version    string
	JSONSchema []byte
}

type SerializedArtifact struct {
	Kind      ArtifactKind
	Schema    ArtifactSchema
	MediaType string
	Content   []byte
	Metadata  map[string]any
}

type ArtifactCodec[T any] interface {
	Kind() ArtifactKind
	Schema() ArtifactSchema
	MediaType() string
	Encode(T) ([]byte, error)
	Decode([]byte) (T, error)
}

Use strict JSON decoding for JSON codecs: reject unknown fields and trailing tokens. Encoding must be deterministic for equal canonical values. Clone byte slices and maps at framework ownership boundaries. Validate non-empty artifact kind, schema ID, schema name, schema version, media type, and JSON Schema during registration. Compute and retain a SHA-256 digest of the JSON Schema bytes.

Use generic Extractor[T], Merger[T], Normalizer[T], and TypedValidator[T] contracts. The exact request/result structs may retain existing names where that reduces churn, but they must satisfy these rules:

  • the extractor returns T, warnings, and framework-owned chunk provenance;
  • merge receives accepted per-chunk typed values carrying lane ID, source ID, chunk ID, chunk index, and chunk reference, already sorted by chunk index;
  • normalize receives and returns T;
  • typed validators receive T plus the relevant universal source, chunk, reference, lane, stage, profile, and metadata context;
  • the LLM client and decoded module options are held by constructed implementations, not passed in operation requests; and
  • no module-facing Zone-B request or result contains RawPayload, any, or serialized JSON as its artifact value.

Retain separate framework wrappers around T for extract, merge, and normalize provenance. Do not put lane IDs, module keys, or framework warnings into the D&D domain value itself.

Support a second SerializedValidator contract for representation-level generic validators. Its request contains immutable bytes, media type, and optional schema metadata. For a Zone-B value, the framework produces that request with the lane codec. Keep a separate non-generic ChunkValidator contract for semantic validation of immutable []source.Chunk; when a representation validator is selected at chunk, the framework instead supplies its canonical JSON chunk encoding. valid_json and valid_json_schema use the serialized path, domain validators use TypedValidator[T], and generic approve/reject validators register explicit chunk and typed-artifact variants.

Typed registry model

Because Go methods cannot introduce type parameters, expose free generic registration functions in internal/framework/pipeline, backed by private non-generic registry entries. Use reflect.TypeFor[T]() only inside registration and framework assembly to prove exact type equality.

  • Codec registry key: artifact kind. Exactly one codec may be registered per kind.
  • Extractor registry key: existing module key. Each entry declares one artifact kind and exact Go type.
  • Merger and normalizer registry key: (existing module key, artifact kind).
  • Typed validator registry key: (existing validator key, artifact kind).
  • Chunk-validator registry key: existing validator key in the distinct chunk target namespace.
  • Serialized validators retain their existing validator key and declare whether they support chunk values, artifact values, or both.
  • Duplicate keys/variants, missing codecs, Go-type mismatches, and incompatible selected variants are errors.

The extractor selected for a lane establishes the lane artifact kind. During resolution, look up merger, normalizer, and validators against that kind. Record artifact kind, schema ID, schema version, and schema digest on the resolved lane and in the resolved-pipeline digest. A resolved pipeline with an incompatible lane must fail before preparation or source execution.

The private erased lane entry owns closures for construction and execution of one concrete T. It may store a value as any internally, but it must verify the exact registered reflect.Type at every erased boundary and return a descriptive framework error rather than panic.

Construction and preparation

Use one uniform construction context:

type ModuleDependencies struct {
	LLM contracts.StructuredLLMClient
}

type BuildRequest struct {
	Dependencies ModuleDependencies
	Options      map[string]any
}

Each registry entry stores both an option-validation closure and a constructor. Implementations own concrete option structs and one decoder used by both closures. Resolution/configuration validation calls the decoder and discards the value; preparation calls it once and supplies the decoded value to the constructor. Reject unknown option fields. Empty options produce the implementation's explicit defaults.

Add:

func Prepare(
	resolved ResolvedPipeline,
	registries Registries,
	deps ModuleDependencies,
) (*PreparedPipeline, error)

PreparedPipeline retains explicit resolved input, chunk, artifact-lane, and output fields and all constructed validators. Construct in stable pipeline order: input; chunk and its validators; each resolved lane in order with extract, merge, normalize, and their validator chains in stage order; then output. On the first failure, return an error identifying stage, lane if any, module or validator key, and cause. No operation method may have run.

Create the one scheduled production LLM client first, inject it into preparation, and then run only the prepared pipeline. Test-only deterministic modules may accept a nil LLM dependency; any implementation that declares or uses LLM-backed execution must reject a nil client at preparation. Remove LLM clients and raw option maps from operation requests after every production implementation has migrated.

Constructed modules and validators are reused for a run. Anything callable from parallel extract workers must be concurrency-safe; production implementations should be immutable after construction.

Package registration

Each package-family registrar exposes:

func Register(registries pipeline.Registries, assets *llm.AssetRegistry) error

The generic and Seriatim registrars ignore the asset argument until they need it. Registrars validate the registry pointers they use and return contextual errors. The CLI creates one complete registry set and one asset registry, then calls registrars in this order: generic, Seriatim, D&D. The D&D registrar owns D&D codecs, implementations, typed generic specializations, prompt/schema assets, and default validator chains.

internal/modules/dnd owns D&D shared types. Its register sibling may import children and generic strategies; the root package must not import its children. internal/modules/generic never imports D&D. internal/modules/seriatim does not import D&D. Move domain-neutral embedded prompt-filesystem helpers to internal/framework/promptfs.

D&D typed model

Define dnd.SpellList, dnd.SpellCast, and evidence/source-reference fields at the D&D package root. Use source.SourceRef; do not create another D&D unit-ref type for artifact provenance. The stable artifact kind is dnd/spell-list.

The spell extractor owns a private LLM DTO and the existing dnd_spells_llm.v1.json response schema. It canonicalizes and maps that DTO to dnd.SpellList. The codec package owns dnd_spells.v1.json and the durable encoding. Keep those schemas separate. Remove the validator-only spellpayload model after all three D&D validators consume dnd.SpellList.

Make appendorder a generic strategy that accepts a typed combine function at registration/construction. The D&D registrar supplies a function that appends spell casts in already-sorted source-chunk order. Make noop a generic typed strategy. Neither generic package imports D&D.

Serialized boundaries

After normalize, encode T once to SerializedArtifact for final output. Output remains domain-neutral and receives serialized artifacts. Preserve the existing output bundle and index contract.

Extract, merge, and normalize checkpoints encode and decode through the same lane codec. Checkpoint identity includes artifact kind, schema ID, schema version, and schema digest. A missing codec or any mismatch invalidates reuse and recomputes the stage; it is not a fatal run error by itself. Codec decode failure also invalidates that checkpoint and records the reason. Never pass serialized checkpoint content directly to the next typed stage.

Debug recording uses the codec for typed artifact values and preserves existing opt-in/sensitive-content rules. Artifact/checkpoint digests use the stable codec bytes. Aggregate manifests, checkpoint indexes, warning slices, and rejection slices have one coordinator writer.

Concurrency

Version-2 configuration gains:

concurrency:
  total_llm: 4
  stage_workers:
    extract: 4

Represent effective stage worker limits as map[string]int. Initially accept only extract; reject unknown keys. Missing extract defaults to total_llm and its valid range is 1..total_llm. Add the environment override NOTARIUS_STAGE_WORKERS_EXTRACT. Preserve precedence rules and keep the file configuration version at 2.

Use one run-wide fixed extract worker pool. Dispatch jobs in round-robin order with source chunk as the outer loop and resolved lane as the inner loop. Use one job channel whose capacity equals the effective extract worker count, so the dispatcher applies bounded backpressure. Do not create one goroutine per job. A job includes extract retries and extract-stage validators. Store results by lane index and chunk index.

When all extract jobs for a lane are terminal, run that lane's merge and then normalize serially. Lane continuations may overlap. All LLM calls at all stages, including retries and validators, use the single injected scheduled client, so total_llm remains the authoritative process-wide provider-call ceiling.

Rejections are terminal results and do not cancel other work. A framework error cancels a derived run context, stops dispatching jobs not yet started, and waits for started tasks to finish or observe cancellation. Choose the returned error as follows:

  1. if the parent context is canceled, return its error;
  2. otherwise discard internal context.Canceled/DeadlineExceeded errors when at least one non-context framework error exists; and
  3. choose the earliest remaining error by stage order (extract, merge, normalize), resolved lane index, chunk index for chunk-scoped work, and configured validator/operation index.

Use a sentinel chunk index after all real chunks for lane-scoped merge and normalize errors. Retain other started-task errors only in opt-in diagnostics. Sort accepted artifacts, warnings, and rejections by resolved lane index, source chunk index where applicable, stage order, validator order, and original within-result order. Completion timing must not affect public output. Run output only if every lane has a successful or rejection-only terminal outcome and no framework error occurred.

Staged Implementation

Stage 1: Compatibility Baselines and ADR Acceptance

Goal: lock down behavior that subsequent internal migrations must preserve and record the architectural decisions as accepted.

Tasks:

  • Add semantic or golden compatibility tests for the maintained Seriatim-to-D&D path: durable output files and JSON, output index, manifest provenance, warnings, rejections, and stable lane/chunk ordering.
  • Snapshot production module keys, validator keys, default validator chains, prompt/schema identities, and maintained example/profile resolution in tests.
  • Strengthen runner tests for fixed topology, validator rejection as a nonfatal outcome, framework-error abort, retries, parent cancellation, checkpoint reuse and invalidation, diagnostics, and opt-in debug behavior.
  • Add an instrumented scheduled-client test showing that all existing production LLM callers share concurrency.total_llm. It need not demonstrate parallel lanes yet.
  • Review the three ADRs against this decision-complete plan, set their status to Accepted, and update their dates only if ADR policy requires an acceptance date. Do not rewrite accepted decision text after this stage.

Exit criteria:

  • compatibility tests fail on an unintended durable-output, key, chain, provenance, or outcome-semantics change; and
  • ADR-0002, ADR-0003, and ADR-0004 are accepted.

Stage 2: Registrar Composition Without Package Moves

Goal: replace CLI leaf-by-leaf registration with package-family composition before changing imports.

Tasks:

  • Add internal/modules/generic/register, internal/modules/seriatim/register, and internal/modules/dnd/register using the fixed registrar signature.
  • Initially let those registrars import the existing stage-oriented packages. Move ownership of production validators, default chains, and prompt assets out of internal/cli/catalog.go and into the appropriate registrar.
  • Have the CLI allocate complete registries and the asset registry once, invoke generic, Seriatim, then D&D registration, and retain its existing test injection paths.
  • Test nil registry handling, duplicate registration errors, stable registered keys, default chains, and asset identities.

Exit criteria:

  • the CLI composition root names only the three registrar packages, framework registry types, and asset registry; and
  • no production key, chain, prompt, schema, or runtime behavior changes.

Stage 3: Mechanical Generic and Seriatim Package Moves

Goal: establish the domain-first generic and source-format trees without changing contracts.

Tasks:

  • Move the Seriatim adapter and tests to internal/modules/seriatim/input/transcript.
  • Move the generic chunker to internal/modules/generic/chunk/units, retaining the configured key generic.
  • Move append-order merge, no-op normalize, JSON output, and all generic validators to their target paths under internal/modules/generic.
  • Update only registrar imports and affected black-box tests. Preserve package behavior and all public registry keys.
  • Remove the emptied old directories.

Exit criteria:

  • generic and Seriatim production implementations exist only under their target trees; and
  • compatibility baselines remain green.

Stage 4: Mechanical D&D Package Move and Import Guard

Goal: establish the D&D tree and enforce ADR-0004 dependency direction while legacy contracts are still intact.

Tasks:

  • Move the D&D scenes chunker, spell extractor, validators, schemas, prompt assets, and D&D shared helpers into the target D&D tree. Do not create the typed root model or codec yet.
  • Move domain-neutral prompt filesystem helpers from internal/modules/sharedassets to internal/framework/promptfs.
  • Move tests with their owning implementation. Relocate tests that intentionally compose domains to a black-box integration-test package rather than creating peer-domain production imports.
  • Add a Go-parser-based import-boundary test. It must reject concrete D&D-to-Seriatim and Seriatim-to-D&D imports, all generic-to-D&D imports, and root-domain imports of child implementations. Allow domain registrars, the CLI composition root, and designated external integration tests to compose packages.
  • Remove old empty stage-oriented and validator directories.

Exit criteria:

  • all production extensions use the target domain-first package layout except the not-yet-created typed codec/model pieces;
  • the import guard detects a deliberate fixture violation; and
  • behavior and keys remain unchanged.

Stage 5: Source-Unit Provenance

Goal: add canonical provenance to engine-owned source units without yet changing the chunk type.

Tasks:

  • Add Ref source.SourceRef to source.SourceUnit, including clone and debug representations.
  • Make the Seriatim adapter assign the fixed self-reference for every unit.
  • Extend source.ValidateDocument to require the unit reference's source ID to match the document, require start and end IDs to equal the unit ID, reject missing/invalid/reversed references, and preserve the existing unit-order and uniqueness checks.
  • Make source digests and source checkpoint serialization include the new reference deterministically.
  • Add focused tests for valid refs and missing, foreign, non-self, and reversed refs, plus source checkpoint/debug round trips.

Exit criteria:

  • every produced source unit has a validated self-reference; and
  • source state preserves it through clone, debug, digest, and checkpoint paths.

Stage 6: Engine-Owned Chunks and Workspace v2

Goal: finish the Zone-A source model and make its persisted compatibility break explicit.

Tasks:

  • Add source.Chunk with the fixed shape and update universal chunk contracts, modules, validators, runner code, checkpoints, debug envelopes, and tests to use it.
  • Derive Chunk.Ref from the first and last included unit references. Validate source identity, non-empty ordered units, contiguous boundary agreement, and exact correspondence between the chunk ref and first/last unit refs.
  • Remove contracts.SourceChunk, StartUnitID, and EndUnitID after all consumers migrate. Do not keep aliases.
  • Advance workspace.WorkspaceSchemaVersion to notarius.workspace.v2. Make loader behavior explicitly classify v1 as incompatible and recompute while leaving files untouched.
  • Update checkpoint identity/digest tests and operations documentation for the one-time v1 resume miss.

Exit criteria:

  • no production code imports a framework-owned chunk type;
  • chunk provenance round-trips exactly; and
  • v1 workspaces are safely ignored while v2 workspaces reuse successfully.

Stage 7: Artifact and Codec Foundation

Goal: add the typed primitives and prove strict serialization independently of production lanes.

Tasks:

  • Add the fixed artifact types, codec interface, schema digest helper, and clone helpers under internal/framework/contracts.
  • Add ArtifactCodecRegistry to pipeline.Registries and ModuleCatalog.
  • Implement generic codec registration and private erasure/type tracking.
  • Validate registration metadata and duplicates. Ensure erased encode/decode returns typed errors, never reflection panics.
  • Use two small test artifact types to cover registration, exact type identity, deterministic encoding, strict decoding, cloning, duplicate kind rejection, and schema metadata/digest behavior.
  • Wire the new empty registry through CLI/test registry constructors without changing production lane execution.

Exit criteria:

  • codecs for heterogeneous test types can coexist and safely round-trip through erased framework storage; and
  • current production behavior remains on the legacy raw path and unchanged.

Stage 8: Typed Contracts, Variants, and Resolution

Goal: resolve a complete type-compatible lane before executing it.

Tasks:

  • Add the typed stage, typed validator, chunk validator, serialized validator, and provenance wrapper contracts from the fixed decisions.
  • Extend extractor specs with artifact kind/type. Convert merger, normalizer, and typed validator registries to artifact-kind variants while retaining serialized-validator registration by key.
  • Implement free generic registration helpers and private erased entries.
  • Extend lane resolution to derive kind from extractor, require its codec, select exact merger/normalizer/validator variants, and include artifact/schema identity in resolved lanes and pipeline digest.
  • Keep legacy registration helpers only as explicitly named transitional APIs; do not let a raw registration satisfy a typed lane.
  • Add composition tests with two artifact types and heterogeneous lanes. Cover missing codec, missing variant, Go-type mismatch, duplicate variant, wrong validator kind, stable resolution order, and digest changes on schema identity or schema digest changes.

Exit criteria:

  • heterogeneous typed test lanes resolve without module-facing erasure;
  • every incompatible selection fails before execution; and
  • existing raw production lanes continue to resolve only through their visible transitional path.

Stage 9: Preparation and Construction Foundation

Goal: construct and validate an entire resolved pipeline before source work.

Tasks:

  • Add ModuleDependencies, BuildRequest, PreparedPipeline, and Prepare as specified.
  • Extend registry entries/specs with option validation and construction functions. Supply adapters for legacy zero-argument constructors during the migration.
  • Call option validation for every selected module and validator during resolution/config validation. Reject unknown fields and contextualize errors.
  • Have preparation construct all selected components in fixed order and retain immutable prepared lane executors.
  • Update the runner API so Run receives a prepared pipeline. At the CLI, create the shared scheduled LLM client, prepare, and only then invoke the runner.
  • Prove with fakes that malformed options, missing required LLM dependencies, and late-component construction failures occur before the input adapter's Parse method.

Exit criteria:

  • all components are constructed before source work;
  • preparation errors identify exact scope and perform no operations; and
  • legacy production modules still run through temporary construction adapters.

Stage 10: Migrate Universal Modules to Construction

Goal: remove legacy option/dependency handling from input, chunk, and output.

Tasks:

  • Give Seriatim input, generic units chunking, D&D scenes chunking, generic JSON output, and all applicable chunk validators implementation-owned option structs and strict decoders.
  • Build those implementations with decoded options and injected dependencies. Require the LLM client for D&D scenes; keep deterministic implementations independent of it.
  • Remove Options and LLMClient from the corresponding operation requests. Retain per-run source, reference, profile, session, and metadata fields.
  • Update registrars and focused tests. Verify options are decoded once during preparation and operation methods do not inspect raw maps.

Exit criteria:

  • no universal production module parses raw options during execution; and
  • every universal LLM call uses the injected shared client.

Stage 11: Canonical D&D Model, Codec, and Typed Extractor

Goal: establish the first production T and its extraction boundary.

Tasks:

  • Add canonical spell types at the D&D root using engine-owned source refs.
  • Add internal/modules/dnd/codec/spells, register kind dnd/spell-list, and make it own the existing durable dnd_spells.v1.json schema and strict stable encoding.
  • Keep the private spell-extraction LLM DTO and dnd_spells_llm.v1.json in the extractor package. Map canonicalized DTO values to dnd.SpellList and do not expose the DTO to validators or the codec.
  • Convert the extractor to Extractor[dnd.SpellList], construction-time options and dependency injection. Preserve prompt assets, retries, warnings, evidence, and LLM response validation.
  • Register the codec and typed extractor from the D&D registrar.
  • Add codec compatibility tests against existing durable fixtures and tests proving LLM schema ownership is separate from durable schema ownership.

Exit criteria:

  • the typed extractor returns the canonical domain model;
  • codec output is semantically identical to the maintained durable spell JSON; and
  • no downstream production consumer is switched until the next stages.

Stage 12: Typed Validators and Generic Typed Strategies

Goal: complete all typed components required by the D&D lane.

Tasks:

  • Convert D&D spell shape, source-reference, and source-relatedness validators to TypedValidator[dnd.SpellList] with construction-time options/dependencies.
  • Remove JSON reparsing from those validators. Remove the duplicate spellpayload package after its final consumer migrates.
  • Convert valid_json and valid_json_schema to serialized validators. Register them so the framework uses the D&D codec when they occur in the spell chain.
  • Implement generic typed append-order merge and no-op normalize. In the D&D registrar, register D&D variants using a spell-list append function and noop[dnd.SpellList].
  • Convert always-accept/reject into explicit chunk and typed variants and register the D&D variants without changing their keys. Verify serialized validators operate on the framework encoding at chunk and the codec encoding at artifact stages.
  • Test typed validator requests, source refs, relatedness LLM injection, generic strategy reuse with a second test type, default chain order, and rejection behavior.

Exit criteria:

  • every selected spell-lane component has a compatible D&D typed variant;
  • generic packages import no D&D code; and
  • the duplicate validator payload model and inter-validator JSON parsing are gone.

Stage 13: Typed D&D Runner Vertical Slice

Goal: execute one complete production lane as dnd.SpellList while retaining temporary raw output/checkpoint adapters.

Tasks:

  • Implement the erased typed lane executor and runner path for extract, stage retries, typed/serialized validation, merge, and normalize.
  • Keep accepted extract values indexed by chunk and pass them to merge in source order. Preserve warnings and rejections in stable scope order.
  • Add narrow transitional adapters from typed stage outputs to the existing raw checkpoint/debug/output envelopes. These adapters must use the registered codec and be named/commented as migration-only.
  • Route the production D&D lane through the typed path; leave no production raw D&D stage module registered in parallel.
  • Add end-to-end and checkpoint-disabled tests proving the maintained D&D output and outcome semantics are unchanged. Add incompatible-lane tests proving failure occurs before input.

Exit criteria:

  • D&D values remain typed from extractor through normalize and typed validation;
  • the runner's only D&D erasure is its private lane adapter and explicit codec boundary; and
  • durable output remains unchanged through the transitional adapter.

Stage 14: Typed Checkpoints, Debugging, and Output

Goal: move every serialization side effect and the final Zone-C boundary to the codec model.

Tasks:

  • Change runner/output contracts so final normalized results are SerializedArtifact values. Update generic JSON output without importing D&D.
  • Preserve logical file names, index fields, media types, manifest contents, and durable spell JSON.
  • Change extract, merge, and normalize checkpoints to store codec bytes plus artifact kind, schema ID, version, and schema digest. Decode reused values back to T before the next stage.
  • Implement safe invalidation for missing/mismatched codecs and decode failure, with explicit checkpoint-event reasons.
  • Change typed debug envelopes to serialize through the codec, preserving opt-in and redaction behavior. Use stable codec bytes for artifact digests.
  • Remove the transitional raw output/checkpoint/debug adapters introduced in Stage 13 after all tests use the typed boundaries.

Exit criteria:

  • typed values round-trip at every stage checkpoint;
  • incompatible checkpoint artifacts recompute safely;
  • output and debug code are domain-neutral; and
  • no D&D typed lane depends on a raw-boundary adapter.

Stage 15: Finish Construction Migration and Remove Raw Contracts

Goal: leave one production extension system rather than parallel raw and typed models.

Tasks:

  • Migrate any remaining merge, normalize, extract, and validator implementations to construction-time option decoding and dependency injection.
  • Remove LLM clients and raw option maps from all remaining operation requests.
  • Remove legacy raw extractor/merger/normalizer/validator contracts, constructors, registry entries, RawPayload, ResponseSchema if superseded, raw clone helpers, raw checkpoint envelopes, and migration-only adapters.
  • Remove dead duplicate models and compatibility helpers. Search for production references to old stage-oriented paths and raw Zone-B types.
  • Keep serialized artifacts only at codec, checkpoint/debug, and output boundaries.

Exit criteria:

  • no production Zone-B handoff uses JSON bytes, RawPayload, or any;
  • all production configuration options are validated before execution;
  • all production LLM users receive the one injected client; and
  • there is no legacy production registration path.

Stage 16: Stage-Worker Configuration

Goal: add the decided scheduling control without changing runner execution yet.

Tasks:

  • Add StageWorkers map[string]int to effective concurrency configuration and stage_workers to the version-2 YAML shape. Deep-clone the map.
  • Accept only extract, reject unknown or empty keys, default missing extract to effective total_llm, and validate the inclusive range 1..total_llm after file/environment precedence resolves.
  • Add NOTARIUS_STAGE_WORKERS_EXTRACT with the existing environment precedence and integer error style.
  • Preserve redaction/effective-config diagnostics and version 2.
  • Update docs/config.md and maintained examples that intentionally demonstrate concurrency. Do not add the field to every example when the default conveys the intended behavior.
  • Add file, default, merge/precedence, environment, unknown-key, boundary, and redacted-effective-config tests.

Exit criteria:

  • every run has a validated effective extract worker count;
  • omitted configuration preserves current effective behavior at the default total_llm: 1; and
  • current configuration documentation owns the implemented contract.

Stage 17: Concurrent Lane and Extract Scheduling

Goal: implement bounded concurrent lanes while preserving deterministic public behavior and the global LLM invariant.

Tasks:

  • Add the fixed-size run-wide worker pool and central round-robin dispatcher. Bound queued work so dispatch applies backpressure; do not enqueue the whole run into unbounded memory.
  • Treat extract, retries, and extract validators as one job. Publish immutable results to a coordinator indexed by lane and chunk.
  • Start each lane's serial merge/normalize continuation only after all its extract jobs are terminal. Permit different lane continuations to overlap.
  • Make the coordinator the sole writer of aggregate output, manifest, checkpoint-event collection, warnings, and rejections. Use attempt-specific debug paths and synchronize any recorder state that remains shared.
  • Implement rejection, cancellation, stable sorting, deterministic primary error selection, and output gating exactly as specified under Concurrency.
  • Audit every concurrently reused extractor, validator, codec, LLM/debug wrapper, checkpoint loader/recorder, and manifest metadata provider. Make production implementations immutable or add narrowly scoped synchronization.
  • Add deterministic barrier-controlled tests that force reverse completion order, simultaneous failures, parent cancellation, rejections mixed with successes, lane continuation overlap, and undispatched-job cancellation.
  • Add instrumented integration tests issuing calls from multiple lanes, retries, and LLM-backed validators. Assert provider calls never exceed total_llm, extract jobs never exceed the effective extract worker count, and both limits are exercised independently.
  • Run go test -race ./... and eliminate races rather than weakening tests.

Exit criteria:

  • lanes and extract jobs actually overlap when configured above one;
  • job and provider-call limits are independently enforced;
  • public results and primary errors are identical across forced completion orders; and
  • full race testing passes.

Stage 18: Documentation, Cleanup, and Final Verification

Goal: make the implemented repository and its canonical documentation agree, then close the focused roadmap.

Tasks:

  • Review every current-behavior document routed by docs/development.md and update only its owned facts: architecture and dependency direction, package inventory, pipeline resolution/preparation/execution, typed module contracts, LLM scheduling, configuration, checkpoint compatibility, diagnostics, operations, and durable integration contracts.

  • Verify maintained examples and copyable files against the implementation.

  • Remove stale old package paths, raw-contract terminology, and superseded future-work entries. Validate all changed documentation links.

  • Update ADR consequences only in ways permitted for accepted ADR metadata; do not edit accepted decision text. Record a new superseding ADR if final code required an architectural change.

  • Mark the feature roadmap implemented and reduce this implementation plan to a concise completion record, or move it to the repository's established completed-roadmap location if one exists. Do not let this file become a second current-behavior reference.

  • Run final checks:

    go test ./...
    go test -race ./...
    go vet ./...
    go build ./cmd/notarius
    

Exit criteria:

  • all feature-roadmap completion criteria are met;
  • no stale production paths or legacy typed/raw bridge remain;
  • current documentation and examples describe only implemented behavior; and
  • all final validation commands pass.

Open Questions

None. The feature-policy decisions and the implementation choices required to begin each stage are resolved above. If implementation evidence contradicts one of the accepted architectural decisions, stop and handle that as an ADR change or supersession rather than treating it as an implicit implementation choice.