8.5 KiB
Pipeline Internals
The implemented pipeline runner lives in internal/framework/pipeline. It
executes the fixed workflow defined by the architecture policy:
input -> chunk -> extract -> merge -> normalize -> output
Pipeline execution is serial. The runner executes the resolved lanes one after another in the fixed workflow order.
Profile Resolution
Config loading produces pipeline.PipelineProfile values. Resolution happens
before execution:
internal/core/config.Config.Resolvevalidates config and finds the named pipeline.- The optional lane selection is passed to
pipeline.ResolvePipeline. - Module bindings are defaulted:
- chunk:
generic - merge:
appendorder - normalize:
noop - output:
json - LLM profile:
default
- chunk:
- The module catalog is checked for each bound module key.
- Module capabilities are checked in workflow order.
- A digest is calculated from the resolved pipeline without the digest field.
The CLI writes the resolved pipeline and digest to diagnostics.
Pipeline profiles and artifact lanes may include reference binding maps keyed by
reference slot name. During resolution, pipeline-level bindings act as defaults
for selected chunk, extractor, and normalizer targets that declare the slot;
target-local bindings override or add bindings for that target. Runtime
--reference requests override target config bindings, and runtime unbinds
remove optional target bindings. Flat runtime slot names are resolved only when
exactly one selected target declares the slot; otherwise the CLI requires a more
specific selector such as chunk.slot, lane.extract.slot, or
lane.normalize.slot. Resolution validates bindings against the declaring
target specs and stores the bindings in target-aware resolved reference holders.
It does not read reference files or include reference bytes in source digests.
During run preparation, resolved file references for chunk, extractor, and
normalizer targets are materialized before any LLM-backed pipeline work. Config
bindings resolve relative to the config file, and CLI bindings resolve relative
to the current working directory. Materialization accepts UTF-8 text files,
computes sha256: content digests, records file origins, infers canonical base
media types from file extensions, enforces declared byte limits, and warns for
empty bound files. Media-type acceptance is checked only when a slot declares
AcceptedMediaTypes; unknown extensions are recorded as
application/octet-stream. Reference content is omitted from diagnostics and
manifests. The CLI writes provenance-only resolved reference diagnostics, and
the run manifest records target-stage reference provenance separately from
source digests. Runtime reference content is passed to the matching chunker,
extractor, or normalizer request.
Prompt bundles can declare reference slots and use reference and
hasreference template functions. Bundle loading validates string-literal slot
names against the declaration. Rendering receives a lane reference set from the
caller; unbound optional slots render as empty strings, and hasreference
returns true only when at least one bound item has content.
Registries And Module Specs
pipeline.Registries holds concrete constructors for execution. A
pipeline.ModuleCatalog exposes module specs for config validation and
resolution.
Every production module registers a ModuleSpec with:
Key: module key used in config;Stage: module kind such as input, chunk, extract, merge, normalize, validate, or output;Provides: capabilities added after that module runs;Requires: capabilities that must already be available.
Chunk, extract, and normalize specs may also declare reference slots. Slot declarations are available from registry metadata without constructing module instances. Input, merge, validate, and output specs must not declare reference slots.
Capability checks prevent incompatible pipeline composition before a run starts.
Runner Input And Output
pipeline.RunInput carries:
- a
ResolvedPipeline; - optional source ID, input path, and raw input bytes;
- a structured LLM client;
- run ID, start time, LLM profile manifest metadata, and CLI metadata.
pipeline.RunOutput carries:
- run manifest;
- approved artifacts;
- rejected artifacts;
- warnings;
- logical output files returned by the output encoder.
The CLI owns durable file writes and diagnostics writes after the runner returns.
Execution
The runner:
- validates run input and registries;
- builds the input adapter and parses the raw input into a source document;
- validates the source document;
- builds the chunker and produces source chunks;
- validates source chunks against framework invariants;
- runs each selected artifact lane in sorted resolved order;
- builds the output encoder and validates logical output file names.
Chunk Results
Chunkers implement contracts.Chunker and receive a contracts.ChunkRequest
with the validated source document, reference set, structured LLM client, the
configured LLM profile, module options, and run metadata. Deterministic and
LLM-backed chunkers use the same contract; provider construction stays outside
chunk modules.
After Chunk returns, the runner appends chunker warnings before returning any
chunker error. When chunking succeeds, the runner validates generic chunk
invariants before running extractors:
- chunk IDs must be non-empty and unique in the chunk result;
- each chunk
SourceIDmust match the source document ID; - each chunk
Indexmust match its zero-based returned order; - each chunk must contain at least one source unit;
- a chunk must not repeat a source unit;
- every chunk source unit must exist in the source document;
- source units inside each chunk must appear in source-document order.
After validation, the runner rebuilds each chunk from source-document units by
ID, preserving the chunk boundary order and cloning chunk metadata. Extractors
and downstream stages therefore see canonical source units, while
SourceChunk.Metadata remains the supported place for chunker-owned context.
The framework does not require complete source-unit coverage and does not reject overlap between different chunks. Stricter policies, such as full coverage or non-overlap, belong to individual chunk modules when they are part of that module's contract.
Within an artifact lane, the runner:
- builds the extractor, merger, and normalizer;
- records module manifest metadata when modules provide it;
- extracts candidates from each chunk;
- normalizes candidate envelope fields such as index, extractor key, artifact type, and schema version;
- merges candidates;
- normalizes merged candidates;
- validates candidate envelope consistency;
- runs validators;
- converts approved candidates to artifacts.
Validators
If a lane declares validators in config, the runner builds those validators from the validator registry. Otherwise it uses validators returned by the extractor.
Each validator must return exactly one decision for each eligible candidate. The
runner enforces decision cardinality with internal/framework/validate.
Rejected candidates are removed before the next validator runs. Approved
candidates continue through the chain.
The production CLI currently registers no standalone validator modules. The current D&D spell extractor supplies deterministic shape and source-reference validators.
Warnings And Failures
Warnings from chunking, extraction, merging, normalization, validation, and
output encoding are accumulated in RunOutput.Warnings.
Errors wrap the operation and module key or lane context. If execution fails
after a manifest exists, the returned manifest is marked failed and receives a
completion timestamp.
On successful execution, the manifest validation status is:
approvedwhen no candidates were rejected;rejectedwhen at least one candidate was rejected.
Manifest Population
The manifest records run ID, pipeline ID, pipeline digest, module keys, top-level module metadata, artifact lanes, LLM profile metadata, source digest, reference provenance, validation status, and timing.
Singleton pipeline modules may add non-secret metadata by implementing
contracts.ManifestMetadataProvider. The runner records that metadata under
module_metadata with stable keys for input, chunker, and output.
Lane-owned modules may add non-secret metadata through
artifact_lanes[].metadata. The runner records extractor, merger, and
normalizer metadata there. The D&D spell extractor uses lane metadata for
prompt and response-schema provenance.