892 lines
27 KiB
Markdown
892 lines
27 KiB
Markdown
# Implementation Plan: Checkpoint 3 Pipeline Stages
|
|
|
|
## Status
|
|
|
|
This is a staged implementation plan for
|
|
[`3-pipeline-stages-chunking-merge-normalize.md`](3-pipeline-stages-chunking-merge-normalize.md).
|
|
It is intended for an LLM coding agent to follow stage by stage.
|
|
|
|
This plan implements only checkpoint 3. Do not add real Seriatim parsing, real
|
|
D&D extraction, production config file loading, CLI command behavior, LLM
|
|
provider clients, prompt assets, response schema assets, diagnostics run
|
|
directories, or durable output writing.
|
|
|
|
## Policy Context
|
|
|
|
Follow:
|
|
|
|
- [`docs/policy/architecture.md`](../policy/architecture.md)
|
|
- [`docs/policy/documentation.md`](../policy/documentation.md)
|
|
|
|
Required boundaries:
|
|
|
|
- framework code must stay source-agnostic and domain-agnostic;
|
|
- concrete business logic remains under `internal/modules/<stage>/...`, but this
|
|
checkpoint should use fake modules only;
|
|
- framework plumbing should remain consolidated under `internal/framework/pipeline`
|
|
unless a boundary proves itself;
|
|
- capabilities are flat strings, not a type system;
|
|
- pipeline profiles are fixed-shape templates, not arbitrary step lists;
|
|
- no structural module wiring through ad hoc CLI flags.
|
|
|
|
## Global Implementation Decisions
|
|
|
|
- Add no third-party dependencies.
|
|
- Keep new checkpoint 3 framework code in `internal/framework/pipeline`.
|
|
- Keep shared request/result contracts in `internal/framework/contracts` when
|
|
stage modules will implement or consume them.
|
|
- Do not create new framework packages for chunking, merge, normalize, output,
|
|
warnings, response schemas, or structured output.
|
|
- Use in-memory pipeline profile structs only. Do not implement YAML/TOML
|
|
parsing in this checkpoint.
|
|
- Keep one shared chunker per resolved pipeline. Per-lane chunker overrides are
|
|
deferred.
|
|
- Execute chunks serially in checkpoint 3. Contracts must not prevent later
|
|
parallel execution.
|
|
- Sort artifact lane IDs during resolution for deterministic output and digest
|
|
behavior.
|
|
- Use `sha256:<hex>` for resolved pipeline digests.
|
|
- Preserve existing validator decision-cardinality behavior.
|
|
- Run `gofmt` on all touched Go files before validation.
|
|
|
|
## Stage 1: Stage Contracts And Source Chunks
|
|
|
|
### Goal
|
|
|
|
Define the source chunk model and contracts for chunk, merge, normalize, and
|
|
output stages.
|
|
|
|
### Files To Update
|
|
|
|
- `internal/framework/contracts/contracts.go`
|
|
- `internal/framework/contracts/contracts_test.go`
|
|
- `internal/framework/contracts/composition_test.go`
|
|
|
|
### Required API
|
|
|
|
Add:
|
|
|
|
```go
|
|
type SourceChunk struct {
|
|
ID string `json:"id"`
|
|
SourceID string `json:"source_id"`
|
|
Index int `json:"index"`
|
|
Units []source.SourceUnit `json:"units"`
|
|
Metadata map[string]any `json:"metadata,omitempty"`
|
|
}
|
|
|
|
type ChunkRequest struct {
|
|
Source *source.SourceDocument `json:"-"`
|
|
Metadata map[string]any `json:"metadata,omitempty"`
|
|
}
|
|
|
|
type ChunkResult struct {
|
|
Chunks []SourceChunk `json:"chunks"`
|
|
Warnings []Warning `json:"warnings,omitempty"`
|
|
}
|
|
|
|
type Chunker interface {
|
|
Key() string
|
|
Chunk(ctx context.Context, req ChunkRequest) (ChunkResult, error)
|
|
}
|
|
```
|
|
|
|
Extend `ExtractionRequest`:
|
|
|
|
```go
|
|
type ExtractionRequest struct {
|
|
Source *source.SourceDocument `json:"-"`
|
|
Chunk *SourceChunk `json:"chunk,omitempty"`
|
|
AmbientContext map[string]any `json:"ambient_context,omitempty"`
|
|
LLMClient StructuredLLMClient `json:"-"`
|
|
Metadata map[string]any `json:"metadata,omitempty"`
|
|
}
|
|
```
|
|
|
|
Add:
|
|
|
|
```go
|
|
type ChunkArtifacts struct {
|
|
Chunk SourceChunk `json:"chunk"`
|
|
Candidates []artifacts.ArtifactCandidate `json:"candidates"`
|
|
}
|
|
|
|
type MergeRequest struct {
|
|
Source *source.SourceDocument `json:"-"`
|
|
LaneID string `json:"lane_id"`
|
|
ChunkArtifacts []ChunkArtifacts `json:"chunk_artifacts"`
|
|
Metadata map[string]any `json:"metadata,omitempty"`
|
|
}
|
|
|
|
type MergeResult struct {
|
|
Candidates []artifacts.ArtifactCandidate `json:"candidates"`
|
|
Warnings []Warning `json:"warnings,omitempty"`
|
|
}
|
|
|
|
type Merger interface {
|
|
Key() string
|
|
Merge(ctx context.Context, req MergeRequest) (MergeResult, error)
|
|
}
|
|
|
|
type NormalizeRequest struct {
|
|
Source *source.SourceDocument `json:"-"`
|
|
LaneID string `json:"lane_id"`
|
|
Candidates []artifacts.ArtifactCandidate `json:"candidates"`
|
|
Metadata map[string]any `json:"metadata,omitempty"`
|
|
}
|
|
|
|
type NormalizeResult struct {
|
|
Candidates []artifacts.ArtifactCandidate `json:"candidates"`
|
|
Warnings []Warning `json:"warnings,omitempty"`
|
|
}
|
|
|
|
type Normalizer interface {
|
|
Key() string
|
|
Normalize(ctx context.Context, req NormalizeRequest) (NormalizeResult, error)
|
|
}
|
|
|
|
type OutputRequest struct {
|
|
Manifest artifacts.RunManifest `json:"manifest"`
|
|
Approved []artifacts.Artifact `json:"approved,omitempty"`
|
|
Rejected []artifacts.RejectedArtifact `json:"rejected,omitempty"`
|
|
Warnings []Warning `json:"warnings,omitempty"`
|
|
Metadata map[string]any `json:"metadata,omitempty"`
|
|
}
|
|
|
|
type OutputResult struct {
|
|
Bytes []byte `json:"-"`
|
|
ContentType string `json:"content_type,omitempty"`
|
|
Warnings []Warning `json:"warnings,omitempty"`
|
|
}
|
|
|
|
type OutputEncoder interface {
|
|
Key() string
|
|
Encode(ctx context.Context, req OutputRequest) (OutputResult, error)
|
|
}
|
|
```
|
|
|
|
If `gofmt` aligns fields differently, keep the formatted output.
|
|
|
|
### Required Behavior
|
|
|
|
- Existing fake extractor tests must still compile with `ExtractionRequest{Source: doc}`.
|
|
- `SourceChunk` is a framework contract type, not a core source type.
|
|
- `SourceChunk.Units` should be copied by implementations where mutation risk
|
|
exists; this stage only defines the contract.
|
|
|
|
### Required Tests
|
|
|
|
Update fake implementations in existing contract tests to prove:
|
|
|
|
- fake chunker satisfies `contracts.Chunker`;
|
|
- fake merger satisfies `contracts.Merger`;
|
|
- fake normalizer satisfies `contracts.Normalizer`;
|
|
- fake output encoder satisfies `contracts.OutputEncoder`;
|
|
- fake extractor can receive a `SourceChunk` and ambient context.
|
|
|
|
### Validation
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w internal/framework/contracts
|
|
go test ./internal/framework/contracts
|
|
go test ./...
|
|
```
|
|
|
|
## Stage 2: Module Metadata And Stage Registries
|
|
|
|
### Goal
|
|
|
|
Teach pipeline registries about module metadata and add registries for chunk,
|
|
merge, normalize, configured validators, and output modules.
|
|
|
|
### Files To Update
|
|
|
|
- `internal/framework/pipeline/input_registry.go`
|
|
- `internal/framework/pipeline/input_registry_test.go`
|
|
- `internal/framework/pipeline/extractor_registry.go`
|
|
- `internal/framework/pipeline/extractor_registry_test.go`
|
|
|
|
### Files To Add
|
|
|
|
- `internal/framework/pipeline/module.go`
|
|
- `internal/framework/pipeline/chunker_registry.go`
|
|
- `internal/framework/pipeline/chunker_registry_test.go`
|
|
- `internal/framework/pipeline/merger_registry.go`
|
|
- `internal/framework/pipeline/merger_registry_test.go`
|
|
- `internal/framework/pipeline/normalizer_registry.go`
|
|
- `internal/framework/pipeline/normalizer_registry_test.go`
|
|
- `internal/framework/pipeline/validator_registry.go`
|
|
- `internal/framework/pipeline/validator_registry_test.go`
|
|
- `internal/framework/pipeline/output_registry.go`
|
|
- `internal/framework/pipeline/output_registry_test.go`
|
|
|
|
### Required API
|
|
|
|
Add:
|
|
|
|
```go
|
|
type ModuleStage string
|
|
|
|
const (
|
|
StageInput ModuleStage = "input"
|
|
StageChunk ModuleStage = "chunk"
|
|
StageExtract ModuleStage = "extract"
|
|
StageMerge ModuleStage = "merge"
|
|
StageNormalize ModuleStage = "normalize"
|
|
StageValidate ModuleStage = "validate"
|
|
StageOutput ModuleStage = "output"
|
|
)
|
|
|
|
type ModuleSpec struct {
|
|
Key string
|
|
Stage ModuleStage
|
|
Provides []string
|
|
Requires []string
|
|
}
|
|
```
|
|
|
|
Add metadata registration methods to existing registries:
|
|
|
|
```go
|
|
func (r *InputAdapterRegistry) RegisterWithSpec(spec ModuleSpec, constructor InputAdapterConstructor) error
|
|
func (r *InputAdapterRegistry) Spec(key string) (ModuleSpec, bool)
|
|
|
|
func (r *ExtractorRegistry) RegisterWithSpec(spec ModuleSpec, constructor ExtractorConstructor) error
|
|
func (r *ExtractorRegistry) Spec(key string) (ModuleSpec, bool)
|
|
```
|
|
|
|
Keep existing `Register` methods. They should call `RegisterWithSpec` with a
|
|
default spec using the provided key and correct stage.
|
|
|
|
Add typed registries for other stages:
|
|
|
|
```go
|
|
type ChunkerConstructor func() (contracts.Chunker, error)
|
|
type ChunkerRegistry struct { /* unexported fields */ }
|
|
func NewChunkerRegistry() *ChunkerRegistry
|
|
func (r *ChunkerRegistry) Register(key string, constructor ChunkerConstructor) error
|
|
func (r *ChunkerRegistry) RegisterWithSpec(spec ModuleSpec, constructor ChunkerConstructor) error
|
|
func (r *ChunkerRegistry) Build(key string) (contracts.Chunker, error)
|
|
func (r *ChunkerRegistry) Spec(key string) (ModuleSpec, bool)
|
|
func (r *ChunkerRegistry) RegisteredKeys() []string
|
|
```
|
|
|
|
Repeat the same pattern for `MergerRegistry`, `NormalizerRegistry`,
|
|
`ValidatorRegistry`, and `OutputEncoderRegistry`.
|
|
|
|
`ValidatorRegistry` should use:
|
|
|
|
```go
|
|
type ValidatorConstructor func() (contracts.Validator, error)
|
|
```
|
|
|
|
Because `contracts.Validator` uses `Name()` rather than `Key()`,
|
|
`ValidatorRegistry.Build` must require `validator.Name()` to match the
|
|
registered key.
|
|
|
|
### Required Behavior
|
|
|
|
- Registry keys are trimmed.
|
|
- Specs are normalized by trimming `Key`, `Provides`, and `Requires`.
|
|
- Empty capability strings are ignored.
|
|
- `Provides` and `Requires` are deduplicated and sorted for deterministic
|
|
behavior.
|
|
- `RegisterWithSpec` rejects an empty key, wrong stage, nil constructor, and
|
|
duplicate key.
|
|
- `Spec` returns false for unknown keys and nil registries.
|
|
- `Build` rejects unknown keys, constructor errors, nil modules, and key
|
|
mismatches.
|
|
- Existing registry tests continue to pass.
|
|
|
|
### Required Tests
|
|
|
|
For each new registry, cover the same behavior as existing input/extractor
|
|
registries:
|
|
|
|
- successful registration and build;
|
|
- metadata registration and lookup;
|
|
- capability normalization;
|
|
- wrong stage rejection;
|
|
- key trimming;
|
|
- empty key rejection;
|
|
- duplicate key rejection;
|
|
- nil constructor rejection;
|
|
- unknown key build error;
|
|
- constructor error wrapping;
|
|
- nil module rejection;
|
|
- key mismatch rejection;
|
|
- sorted `RegisteredKeys`;
|
|
- nil registry behavior.
|
|
|
|
For `ValidatorRegistry`, key mismatch checks should compare the registry key to
|
|
`Validator.Name()`.
|
|
|
|
For existing input/extractor registries, add tests for metadata registration,
|
|
default specs from `Register`, capability normalization, and wrong stage
|
|
rejection.
|
|
|
|
### Validation
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w internal/framework/pipeline
|
|
go test ./internal/framework/pipeline
|
|
go test ./...
|
|
```
|
|
|
|
## Stage 3: Resolved Pipeline Model
|
|
|
|
### Goal
|
|
|
|
Add fixed-shape pipeline profile structs, default resolution, lane selection,
|
|
capability validation, and resolved pipeline digesting.
|
|
|
|
### Files To Add
|
|
|
|
- `internal/framework/pipeline/profile.go`
|
|
- `internal/framework/pipeline/profile_test.go`
|
|
|
|
### Files To Update
|
|
|
|
- `internal/core/artifacts/artifacts.go`
|
|
- `internal/core/artifacts/artifacts_test.go`
|
|
|
|
### Required API
|
|
|
|
Update `artifacts.RunManifest` to include:
|
|
|
|
```go
|
|
PipelineID string `json:"pipeline_id,omitempty"`
|
|
PipelineDigest string `json:"pipeline_digest,omitempty"`
|
|
```
|
|
|
|
Add lane-level manifest details rather than relying on the existing shared
|
|
`Extractors`, `Merger`, or `Normalizer` fields for multi-lane runs:
|
|
|
|
```go
|
|
type ArtifactLaneManifest struct {
|
|
ID string `json:"id"`
|
|
Extractor string `json:"extractor"`
|
|
Merger string `json:"merger"`
|
|
Normalizer string `json:"normalizer"`
|
|
Validators []string `json:"validators,omitempty"`
|
|
}
|
|
```
|
|
|
|
Then add:
|
|
|
|
```go
|
|
ArtifactLanes []ArtifactLaneManifest `json:"artifact_lanes,omitempty"`
|
|
```
|
|
|
|
Do not remove the existing singular fields in this checkpoint unless doing so is
|
|
required to make the package compile. Treat `ArtifactLanes` as the authoritative
|
|
manifest representation for checkpoint 3's multi-lane pipeline model.
|
|
|
|
Add pipeline model types:
|
|
|
|
```go
|
|
type ModuleBinding struct {
|
|
Module string `json:"module"`
|
|
LLMProfile string `json:"llm_profile,omitempty"`
|
|
Options map[string]any `json:"options,omitempty"`
|
|
}
|
|
|
|
type ArtifactLaneProfile struct {
|
|
Extract ModuleBinding `json:"extract"`
|
|
Merge ModuleBinding `json:"merge,omitempty"`
|
|
Normalize ModuleBinding `json:"normalize,omitempty"`
|
|
Validators []ModuleBinding `json:"validators,omitempty"`
|
|
}
|
|
|
|
type PipelineProfile struct {
|
|
ID string `json:"id"`
|
|
Input ModuleBinding `json:"input"`
|
|
Chunk ModuleBinding `json:"chunk,omitempty"`
|
|
Artifacts map[string]ArtifactLaneProfile `json:"artifacts"`
|
|
Output ModuleBinding `json:"output,omitempty"`
|
|
}
|
|
|
|
type ResolveOptions struct {
|
|
Only []string
|
|
}
|
|
|
|
type ResolvedArtifactLane struct {
|
|
ID string
|
|
Extract ModuleBinding
|
|
Merge ModuleBinding
|
|
Normalize ModuleBinding
|
|
Validators []ModuleBinding
|
|
}
|
|
|
|
type ResolvedPipeline struct {
|
|
ID string
|
|
Digest string
|
|
Input ModuleBinding
|
|
Chunk ModuleBinding
|
|
ArtifactLanes []ResolvedArtifactLane
|
|
Output ModuleBinding
|
|
}
|
|
```
|
|
|
|
Add defaults:
|
|
|
|
```go
|
|
const (
|
|
DefaultChunkModule = "generic"
|
|
DefaultMergeModule = "appendorder"
|
|
DefaultNormalizeModule = "noop"
|
|
DefaultOutputModule = "json"
|
|
DefaultLLMProfile = "default"
|
|
)
|
|
```
|
|
|
|
Add:
|
|
|
|
```go
|
|
type ModuleCatalog struct {
|
|
Inputs *InputAdapterRegistry
|
|
Chunkers *ChunkerRegistry
|
|
Extractors *ExtractorRegistry
|
|
Mergers *MergerRegistry
|
|
Normalizers *NormalizerRegistry
|
|
Validators *ValidatorRegistry
|
|
Outputs *OutputEncoderRegistry
|
|
}
|
|
|
|
func Binding(module string) ModuleBinding
|
|
func ResolvePipeline(profile PipelineProfile, options ResolveOptions, catalog ModuleCatalog) (ResolvedPipeline, error)
|
|
```
|
|
|
|
### Required Behavior
|
|
|
|
`ResolvePipeline` must:
|
|
|
|
- trim `PipelineProfile.ID`;
|
|
- reject empty pipeline IDs;
|
|
- require non-empty input module;
|
|
- require at least one artifact lane before and after `Only` filtering;
|
|
- apply defaults for missing chunk, merge, normalize, output, and binding
|
|
`llm_profile` values;
|
|
- trim all module keys;
|
|
- trim lane names in `ResolveOptions.Only`;
|
|
- reject empty or unknown lane names in `ResolveOptions.Only`;
|
|
- deduplicate repeated `Only` lane names;
|
|
- sort selected lane IDs alphabetically;
|
|
- reject unknown module keys using registry metadata, not constructors;
|
|
- validate flat capability requirements before execution;
|
|
- compute `ResolvedPipeline.Digest` from canonical JSON of the resolved
|
|
pipeline excluding the digest field itself.
|
|
|
|
Capability validation algorithm:
|
|
|
|
1. Start with capabilities provided by the input module.
|
|
2. Validate chunker requirements against current capabilities, then add chunker
|
|
provided capabilities.
|
|
3. For each selected lane, start with input plus chunk capabilities.
|
|
4. Validate extractor requirements, then add extractor provided capabilities.
|
|
5. Validate merger requirements, then add merger provided capabilities.
|
|
6. Validate normalizer requirements, then add normalizer provided capabilities.
|
|
7. Validate configured validator requirements in lane order, then add validator
|
|
provided capabilities. If no validators are configured for a lane, do not
|
|
attempt capability validation for extractor-owned default validators.
|
|
8. Validate output requirements against the union of pipeline-level capabilities
|
|
and all selected lane capabilities.
|
|
|
|
Error messages must name the pipeline ID, lane ID when applicable, stage, module
|
|
key, and missing capability when applicable.
|
|
|
|
`Binding(module)` should trim the module key and leave `LLMProfile` empty and
|
|
`Options` nil before resolution. Resolution should fill empty `LLMProfile`
|
|
values with `DefaultLLMProfile` on all resolved bindings.
|
|
|
|
### Required Tests
|
|
|
|
Cover:
|
|
|
|
- successful resolution with explicit modules;
|
|
- default chunk, merge, normalize, output, and `llm_profile`;
|
|
- lane selection with `Only`;
|
|
- `Only` trimming and deduplication;
|
|
- unknown `Only` lane rejection;
|
|
- empty `Only` lane rejection;
|
|
- empty artifact set rejection;
|
|
- empty pipeline ID rejection;
|
|
- missing input rejection;
|
|
- unknown module key rejection for each stage;
|
|
- missing capability rejection for input/chunk/extract/merge/normalize/validate/output
|
|
interactions;
|
|
- deterministic lane ordering;
|
|
- deterministic digest for equivalent input maps;
|
|
- digest changes when a module binding changes;
|
|
- manifest JSON still omits empty optional fields.
|
|
|
|
Use fake registries and fake specs only. Do not add config file parsing.
|
|
|
|
### Validation
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w internal/core/artifacts internal/framework/pipeline
|
|
go test ./internal/core/artifacts
|
|
go test ./internal/framework/pipeline
|
|
go test ./...
|
|
```
|
|
|
|
## Stage 4: Generic Merge And Normalize Helpers
|
|
|
|
### Goal
|
|
|
|
Add generic append-in-chunk-order merge and no-op normalization behavior inside
|
|
`internal/framework/pipeline`.
|
|
|
|
### Files To Add
|
|
|
|
- `internal/framework/pipeline/generic_stages.go`
|
|
- `internal/framework/pipeline/generic_stages_test.go`
|
|
|
|
### Required API
|
|
|
|
Add:
|
|
|
|
```go
|
|
type AppendOrderMerger struct{}
|
|
func (m AppendOrderMerger) Key() string
|
|
func (m AppendOrderMerger) Merge(ctx context.Context, req contracts.MergeRequest) (contracts.MergeResult, error)
|
|
|
|
type NoopNormalizer struct{}
|
|
func (n NoopNormalizer) Key() string
|
|
func (n NoopNormalizer) Normalize(ctx context.Context, req contracts.NormalizeRequest) (contracts.NormalizeResult, error)
|
|
```
|
|
|
|
### Required Behavior
|
|
|
|
`AppendOrderMerger` must:
|
|
|
|
- return key `appendorder`;
|
|
- concatenate candidates by `ChunkArtifacts` slice order and candidate order;
|
|
- copy candidate slices so caller mutation does not affect the result;
|
|
- preserve source refs and metadata exactly;
|
|
- append no warnings.
|
|
|
|
`NoopNormalizer` must:
|
|
|
|
- return key `noop`;
|
|
- copy candidates so caller mutation does not affect the result;
|
|
- preserve ordering and source refs exactly;
|
|
- append no warnings.
|
|
|
|
### Required Tests
|
|
|
|
Cover:
|
|
|
|
- key values;
|
|
- append order across multiple chunks;
|
|
- stable candidate order within chunks;
|
|
- no-op normalize preserves order;
|
|
- result slices are caller-mutation safe;
|
|
- nil/empty input returns empty candidates and no error.
|
|
|
|
### Validation
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w internal/framework/pipeline
|
|
go test ./internal/framework/pipeline
|
|
go test ./...
|
|
```
|
|
|
|
## Stage 5: Pipeline Runner Refactor
|
|
|
|
### Goal
|
|
|
|
Refactor the existing runner from extractor-only execution to resolved pipeline
|
|
execution across input, chunk, extract, merge, normalize, validate, and output.
|
|
|
|
### Files To Update
|
|
|
|
- `internal/framework/pipeline/runner.go`
|
|
- `internal/framework/pipeline/runner_test.go`
|
|
- `internal/framework/pipeline/registry_integration_test.go`
|
|
|
|
### Required API
|
|
|
|
Replace the extractor-only runner constructor with:
|
|
|
|
```go
|
|
type Registries struct {
|
|
Inputs *InputAdapterRegistry
|
|
Chunkers *ChunkerRegistry
|
|
Extractors *ExtractorRegistry
|
|
Mergers *MergerRegistry
|
|
Normalizers *NormalizerRegistry
|
|
Validators *ValidatorRegistry
|
|
Outputs *OutputEncoderRegistry
|
|
}
|
|
|
|
type Runner struct {
|
|
// unexported fields
|
|
}
|
|
|
|
func New(registries Registries) *Runner
|
|
```
|
|
|
|
Replace `RunInput` with:
|
|
|
|
```go
|
|
type RunInput struct {
|
|
Pipeline ResolvedPipeline
|
|
SourceID string
|
|
Path string
|
|
RawInput []byte
|
|
LLMClient contracts.StructuredLLMClient
|
|
Metadata map[string]any
|
|
}
|
|
```
|
|
|
|
Replace or extend `RunOutput` with:
|
|
|
|
```go
|
|
type RunOutput struct {
|
|
Manifest artifacts.RunManifest `json:"manifest"`
|
|
Approved []artifacts.Artifact `json:"approved,omitempty"`
|
|
Rejected []artifacts.RejectedArtifact `json:"rejected,omitempty"`
|
|
Warnings []contracts.Warning `json:"warnings,omitempty"`
|
|
EncodedOutput []byte `json:"-"`
|
|
ContentType string `json:"content_type,omitempty"`
|
|
}
|
|
```
|
|
|
|
Keep the existing candidate metadata normalization behavior, but it should now
|
|
run per extractor and across all selected lanes.
|
|
|
|
### Required Behavior
|
|
|
|
`Run` must:
|
|
|
|
1. reject nil runner;
|
|
2. reject missing registries for stages used by the pipeline;
|
|
`Validators` is required only when at least one selected lane declares
|
|
validator bindings;
|
|
3. reject empty resolved pipeline ID or digest;
|
|
4. build the input adapter and parse `RawInput`;
|
|
5. validate the resulting `SourceDocument`;
|
|
6. build the chunker and produce chunks;
|
|
7. reject empty chunk results;
|
|
8. for each selected artifact lane in resolved order:
|
|
- build extractor, merger, and normalizer;
|
|
- extract from every chunk serially;
|
|
- pass `Source`, active `Chunk`, `LLMClient`, and metadata to the extractor;
|
|
- normalize candidate metadata with global monotonically increasing indices;
|
|
- merge per-chunk candidates;
|
|
- normalize merged candidates;
|
|
- run configured lane validators against normalized candidates when the lane
|
|
declares validators;
|
|
- otherwise run the extractor-owned default validator chain;
|
|
- append approved artifacts and rejected artifacts to output;
|
|
9. build the output encoder and encode the final output bundle;
|
|
10. populate `RunManifest` with pipeline ID, pipeline digest, input module,
|
|
chunker, lane-level extractor/merger/normalizer/validator keys, output
|
|
encoder, source digest, and validation status;
|
|
11. return partial output plus wrapped errors on failure.
|
|
|
|
Validation status:
|
|
|
|
- `approved` when all candidates from all lanes survive validation and no lane
|
|
returns rejected artifacts;
|
|
- `rejected` when at least one candidate is rejected and execution otherwise
|
|
succeeds;
|
|
- `failed` when returning an error after a manifest has been initialized.
|
|
|
|
Warnings:
|
|
|
|
- collect warnings from input, chunk, extract, merge, normalize, validator, and
|
|
output stages where those stages expose warnings;
|
|
- existing input adapters do not expose warnings, so no input warnings are
|
|
required in this checkpoint.
|
|
|
|
Do not validate source references inside the runner in checkpoint 3. Source
|
|
reference validation remains a later deterministic validator.
|
|
|
|
### Required Tests
|
|
|
|
Refactor existing runner tests and add coverage for:
|
|
|
|
- nil runner error;
|
|
- missing registry errors;
|
|
- invalid resolved pipeline error;
|
|
- input adapter build and parse errors;
|
|
- invalid source document error;
|
|
- chunker build and chunk errors;
|
|
- empty chunk result error;
|
|
- extractor runs once per chunk;
|
|
- extractor receives active chunk and LLM client;
|
|
- merge receives per-chunk candidates;
|
|
- normalize receives merged candidates;
|
|
- validator approvals and rejections still work;
|
|
- configured validators are built and run in lane order;
|
|
- extractor-owned default validators run when a lane does not configure
|
|
validators;
|
|
- configured validators replace extractor-owned default validators when present;
|
|
- global candidate indices across lanes and chunks;
|
|
- warnings from chunk, extract, merge, normalize, validator, and output stages;
|
|
- output encoder receives manifest and artifacts;
|
|
- manifest pipeline ID, digest, and artifact lane details are populated;
|
|
- partial output is returned when a later lane fails.
|
|
|
|
Use fake modules only.
|
|
|
|
### Validation
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w internal/framework/pipeline
|
|
go test ./internal/framework/pipeline
|
|
go test ./...
|
|
```
|
|
|
|
## Stage 6: Fixture-Driven Walking Skeleton
|
|
|
|
### Goal
|
|
|
|
Add an end-to-end fixture-driven integration test that exercises the full staged
|
|
path with fake modules and a fake LLM client.
|
|
|
|
### Files To Add
|
|
|
|
- `internal/framework/pipeline/testdata/walking_skeleton_input.json`
|
|
- `internal/framework/pipeline/testdata/walking_skeleton_output.json`
|
|
- `internal/framework/pipeline/walking_skeleton_test.go`
|
|
|
|
### Required Fixture Shape
|
|
|
|
Use a small input fixture with two or three source units in a simple fake JSON
|
|
format. Do not use Seriatim.
|
|
|
|
Example input shape:
|
|
|
|
```json
|
|
{
|
|
"id": "fixture-source",
|
|
"units": [
|
|
{"id": "u1", "text": "First event."},
|
|
{"id": "u2", "text": "Second event."},
|
|
{"id": "u3", "text": "Third event."}
|
|
]
|
|
}
|
|
```
|
|
|
|
The expected output fixture should contain:
|
|
|
|
- manifest with `pipeline_id` and `pipeline_digest`;
|
|
- approved artifacts emitted from multiple chunks;
|
|
- no rejected artifacts for the happy path;
|
|
- warnings only if fake stages deliberately emit them.
|
|
|
|
Compare expected output by unmarshaled structural equality, not raw bytes, so
|
|
formatting differences do not fail the test.
|
|
|
|
### Required Test Setup
|
|
|
|
Register fake modules with metadata:
|
|
|
|
- fake input adapter key `fake/input`, provides `plain_text`;
|
|
- fake chunker key `fake/chunk`, requires `plain_text`, provides `chunks`;
|
|
- fake extractor key `fake/extract`, requires `chunks`, provides
|
|
`fake_artifacts`;
|
|
- append-order merger key `appendorder`, requires `fake_artifacts`;
|
|
- no-op normalizer key `noop`;
|
|
- fake output encoder key `json`.
|
|
|
|
Use an in-memory `PipelineProfile`:
|
|
|
|
```go
|
|
PipelineProfile{
|
|
ID: "walking-skeleton",
|
|
Input: Binding("fake/input"),
|
|
Chunk: Binding("fake/chunk"),
|
|
Artifacts: map[string]ArtifactLaneProfile{
|
|
"events": {Extract: Binding("fake/extract")},
|
|
},
|
|
Output: Binding("json"),
|
|
}
|
|
```
|
|
|
|
Resolve the profile, run the runner with fixture bytes and fake LLM client, then
|
|
assert the encoded output matches the expected fixture structurally.
|
|
|
|
### Required Behavior
|
|
|
|
- The fake input adapter parses raw fixture bytes into `SourceDocument`.
|
|
- The fake chunker produces at least two chunks.
|
|
- The fake extractor calls the fake structured LLM client once per chunk.
|
|
- The fake extractor returns one artifact candidate per chunk.
|
|
- The append-order merger preserves chunk order.
|
|
- The no-op normalizer preserves candidates.
|
|
- The fake output encoder produces deterministic JSON bytes.
|
|
- The test fails before execution if required capabilities are missing.
|
|
|
|
### Additional Tests
|
|
|
|
Add focused tests for the walking skeleton failure path:
|
|
|
|
- missing required capability fails during resolution;
|
|
- `Only` with an unknown lane fails during resolution;
|
|
- fake LLM call count equals chunk count in the happy path.
|
|
|
|
### Validation
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w internal/framework/pipeline
|
|
go test ./internal/framework/pipeline
|
|
go test ./...
|
|
```
|
|
|
|
## Stage 7: Final Checkpoint 3 Review Pass
|
|
|
|
### Goal
|
|
|
|
Clean up naming, formatting, and accidental scope creep before checkpoint 3 is
|
|
considered complete.
|
|
|
|
### Required Review
|
|
|
|
Check:
|
|
|
|
- no Seriatim parser exists;
|
|
- no D&D module exists;
|
|
- no production config file loader exists;
|
|
- no CLI behavior was added;
|
|
- no LLM provider code exists;
|
|
- no prompt, response schema, diagnostics, or durable output package exists;
|
|
- no new framework package was created beyond `contracts`, `pipeline`, and
|
|
`validate`;
|
|
- stage contracts remain source/domain agnostic;
|
|
- capabilities remain flat strings;
|
|
- resolved pipeline profiles remain fixed-shape templates;
|
|
- walking skeleton uses fake modules only.
|
|
|
|
### Required Validation
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w internal/core internal/framework
|
|
go test ./...
|
|
go build ./cmd/notarius
|
|
```
|
|
|
|
## Open Questions
|
|
|
|
None. The deferred choices identified while drafting this plan are intentionally
|
|
out of scope for checkpoint 3 and should remain deferred until later roadmap
|
|
items.
|