13 KiB
Pipeline Internals
The implemented resolver and runner live in internal/framework/pipeline.
Their fixed workflow and ownership boundaries are defined by
Architecture. Configuration fields,
defaults, and selectable keys are defined in
Configuration.
Pipeline execution is serial. Resolution fixes the selected lanes and all stage bindings; preparation constructs every selected implementation before the runner begins source work.
Resolution
internal/core/config.Config.Resolve validates the loaded configuration,
selects the named profile, applies the runtime inputs supplied by the CLI, and
calls pipeline.ResolvePipeline.
ResolvePipeline:
- selects and sorts artifact lanes;
- completes omitted bindings using the documented configuration defaults;
- looks up each module and validator spec without constructing it;
- for a typed extractor, derives its artifact kind, requires the codec, and selects exact-type merger, normalizer, and validator variants;
- checks required and provided capabilities in workflow order;
- resolves target-aware reference bindings and validator chains;
- validates each selected module and validator option set through its registry entry; and
- calculates a digest over the resolved structure, including typed artifact kind and schema identity.
Resolution returns a ResolvedPipeline containing ordered lanes, concrete
bindings, validator chains, reference targets, and the digest. It does not read
reference bytes or construct runtime modules. CLI lane and reference selector
syntax is defined in the CLI reference.
Reference Materialization
The CLI calls MaterializeReferences after resolution and before constructing
the LLM client or running the pipeline. The materializer checks each binding
against its resolved target declaration, reads and validates the file, and
builds both a contracts.ReferenceSet and provenance-only metadata on the
corresponding ResolvedReferenceTarget.
The runner clones the resulting set into the chunk, extract, merge, or normalize request that owns the target. LLM-backed extensions may convert those items into named prompt inputs. Reference content remains separate from source evidence and source digests.
Binding precedence, path resolution, accepted content, and media-type behavior are configuration contracts; see Configuration. Durable provenance is defined in the JSON output contract, while runtime sensitive-data handling belongs in Operations.
Registries And Specs
pipeline.Registries holds option validators and run-local builders used during
resolution and preparation.
pipeline.ModuleCatalog exposes their specs during configuration validation and
resolution. Separate registries exist for every stage and for validators;
ValidatorChainRegistry stores production default-chain mappings. Both
containers also carry an ArtifactCodecRegistry. Generic registration records
one codec per stable artifact kind, validates its schema metadata and JSON
Schema, retains the exact schema digest and Go type, and safely encodes or
decodes framework-erased values with typed errors on incompatibility.
Typed extractor entries are keyed by module key and declare one artifact kind. Merger, normalizer, and typed-validator variants are keyed by module or validator key plus artifact kind. Chunk and serialized validators occupy separate target namespaces; serialized registrations declare whether they support chunks, artifacts, or both. Duplicate variants and exact Go-type mismatches are rejected deterministically.
Production composition registers the D&D spell-list codec and typed extractor, matching typed merge, normalize, and semantic-validator variants, and serialized JSON validators. The production D&D lane has no parallel raw stage registration. A standalone raw registration cannot satisfy a typed lane.
A ModuleSpec declares its stage plus required and provided capabilities.
Chunk, extract, merge, and normalize specs may also declare reference slots.
Registry implementations defensively copy spec metadata, reject duplicate keys,
and verify that a constructed implementation reports the registered key.
Builder registrations accept ModuleDependencies and cloned raw options through
one BuildRequest. Production input, chunk, and output builders decode those
options and retain typed values or injected dependencies in the constructed
implementation. A typed extractor registration may explicitly supply a raw
adapter builder for a still-raw downstream lane; the adapter is selected as one
unit and does not expose the typed value to raw consumers. Remaining production
raw-stage registrations are adapted from their zero-argument constructors.
A ValidatorSpec declares a validator key and execution class. Resolution uses
the execution class to reject incompatible profile bindings before execution.
The current production catalog and default chain are listed only in
Configuration.
Preparation And Runner Boundary
pipeline.Prepare receives a resolved pipeline, the registries, and shared
module dependencies. It constructs input; chunk and its validators; each lane's
extract, merge, and normalize modules and validator chains in resolved order;
then output. It stops at the first error with pipeline, stage, lane, module, and
validator context as applicable. It never invokes an operation method.
PreparedPipeline keeps private constructed executors and exposes cloned
resolved input, chunk, lane, and output identities. pipeline.RunInput carries
that prepared pipeline, raw source input, run identity and timing, optional
session and profile metadata, and checkpoint/debug collaborators. The runner
parses source bytes through the already constructed input adapter. Later stage
requests receive the generic source model; extract requests receive
chunk-scoped input material, while chunk, merge, and normalize requests retain
access to the original source material. Input, chunk, and output operation
requests do not carry raw module options. The chunk request also does not carry
an LLM client; an LLM-backed chunker receives the shared client during
preparation. Their operation requests retain run-specific source, reference,
profile, session, and metadata context as applicable.
Prepared typed lanes retain exact-type-checked erased operation closures. The runner uses those closures to keep each value typed through extraction, validation, merge, and normalization. Legacy raw lanes continue through their existing executor while they migrate independently.
Source validation requires every unit to carry a canonical self-reference to its containing document and its own unit ID. Explicit clone, checkpoint, and debug boundaries retain that reference, and the canonical source digest covers it deterministically. Chunks use the same source model and carry one canonical reference spanning the first selected unit through the last.
pipeline.RunOutput carries the run manifest, accepted normalized serialized
artifacts with lane and normalizer provenance,
rejected results, warnings, checkpoint events, and logical files returned by the
output encoder. The CLI owns diagnostics and durable filesystem writes after the
runner returns.
Execution Flow
The runner:
- validates its prepared input;
- parses the raw input with the prepared adapter and validates the generic source document;
- obtains or executes the chunk result;
- validates and canonicalizes chunks;
- executes each resolved artifact lane in order;
- invokes the prepared output encoder and validates its logical file results;
- returns the assembled manifest, outcomes, warnings, and files.
Within each artifact lane, it reuses the prepared extractor, merger, normalizer, and validators while performing these transitions:
- extract once per accepted chunk and add runner-owned lane, source, and chunk provenance;
- validate each extract result and omit rejected results from merge input;
- skip the rest of the lane when no extract result is accepted;
- merge accepted extract results in their existing order;
- validate the merge result and skip normalization on rejection;
- normalize the accepted merge result;
- validate and append the accepted normalized result.
Module-provided warnings and payload warnings are promoted only from attempts whose results are accepted and used.
Chunk Canonicalization
Before lane execution, generic validation requires unique chunk IDs, matching source identity, indexes matching returned order, a valid canonical reference, non-empty content and media type, and at least one valid source unit per chunk. Units may not repeat inside a chunk and must form a contiguous range in source-document order. The chunk reference must exactly match the source and the first and last unit references.
The runner then rebuilds each chunk's unit slice from the source document by unit ID. It preserves the canonical reference, content, media type, and cloned metadata. The framework permits gaps and overlap between separate chunks; stricter coverage policy belongs to the chunk implementation.
Validation And Retries
Chunk, extract, merge, and normalize results pass through the resolved validator chain for their stage and module. Chunk validators receive canonical chunks; typed validators receive the domain value; and serialized validators receive canonical chunk JSON or artifact codec bytes. Validators execute in resolved order and stop at the first error or rejection. An empty chain approves the result.
runWithRetry applies the effective retry policy around module execution and
its complete validation chain. A module or validator error becomes a framework
error when attempts are exhausted. A rejection becomes a recorded
RejectedOutput when attempts are exhausted. Cancellation stops retry
processing immediately.
Rejected output is a non-fatal pipeline outcome and does not advance. Warnings from discarded attempts are not promoted. Configuration owns retry counts and validator overrides; see Module Bindings.
Checkpoint And Debug Hooks
The runner depends on recorder and loader interfaces, using no-op implementations when collaborators are absent. Each checkpointed workflow boundary records a running, succeeded, or failed transition. Reuse decisions are consulted in workflow order and accepted payloads are cloned before entering the normal handoff path. Typed extract, merge, and normalize checkpoints store codec bytes with artifact kind, schema ID and version, exact schema digest, and media type. Reuse compares that identity with the prepared codec and decodes through the codec; missing identity, mismatches, corrupt bytes, and decode failures become explicit reuse misses and execute the step normally. Dependency fingerprints and debug content digests use the same stable codec bytes that cross those boundaries.
Debug instrumentation wraps run, stage, attempt, validator, and structured LLM boundaries. Context scopes associate nested LLM calls with the module or validator attempt that made them. Debug-write failures are framework errors; debug data is never used as a checkpoint source. Typed artifact debug envelopes are domain-neutral, redact sensitive metadata and bytes through the common debug policy, and record codec identity plus schema and content digests.
Checkpoint identity, physical layout, reuse behavior, and debug artifact handling are operator contracts in Operations. Serialization and recorder implementation are inventoried in Internal Overview.
Results And Failures
The runner owns manifest assembly and handoff summaries but not the durable JSON schema. It records resolved module and lane provenance, validator chains, source/reference identities, selected LLM profiles, normalized and rejected summaries, status, and timing. Raw payload bytes remain outside the manifest. Module metadata providers may add non-secret singleton or lane-scoped metadata.
Execution errors include stage, module, lane, or validator context. Once a manifest exists, a failing run returns it with failed status and completion time. Successful status reflects whether any result was rejected. The durable manifest and logical file schemas are defined in the JSON output contract.
Tests To Inspect
internal/core/config/effective_config_test.go: config-to-resolution boundary.internal/framework/pipeline/profile_test.go: selection, defaults, capabilities, validator chains, and digest behavior.internal/framework/pipeline/artifact_codec_registry_test.go: typed codec metadata, registration, erasure safety, strict decoding, and cloning.internal/framework/pipeline/typed_resolution_test.go: heterogeneous typed lane resolution and preparation, target-specific validators, incompatibilities, ordering, and schema-sensitive pipeline identity.internal/framework/pipeline/preparation_test.go: option validation, construction order, dependency failures, and the before-source-work boundary.internal/framework/pipeline/references_test.go: target resolution and materialization.internal/framework/pipeline/runner_test.go: stage transitions, retries, rejections, warnings, checkpoints, debug hooks, and manifests.internal/framework/pipeline/walking_skeleton_test.go: fake-backed complete workflow composition.internal/framework/checkpoint/*_test.go: checkpoint serialization and reuse collaborators.