Document raw pipeline completion

This commit is contained in:
2026-07-07 19:32:15 +00:00
parent 7c95791e94
commit a9d8505cdb
7 changed files with 91 additions and 822 deletions

View File

@@ -248,9 +248,9 @@ merge, and normalize binding.
| chunk | `generic` | Splits source units into ordered chunks. |
| chunk | `dnd/scenes` | Uses an LLM to split transcript source units into D&D scenes. |
| extract | `dnd/spells` | Extracts D&D spell raw outputs. |
| merge | `appendorder` | Keeps raw extract outputs in append order. |
| merge | `appendorder` | Merges JSON raw extract outputs in chunk order. |
| normalize | `noop` | Passes merged raw outputs through unchanged. |
| output | `json` | Produces JSON output files. |
| output | `json` | Produces JSON output files for normalized `application/json` lanes. |
The `generic` chunker accepts:
@@ -308,6 +308,6 @@ Pipeline resolution additionally checks:
- required module keys are present;
- module keys are registered for the expected slot;
- module capability requirements are satisfied;
- bound reference slots are declared by selected chunk, extractor, or
- bound reference slots are declared by selected chunk, extractor, merger, or
normalizer targets;
- required reference slots are bound for selected targets.

View File

@@ -15,7 +15,9 @@ CompleteStructured(ctx, request, out) (response, error)
The request contains prompt ID/version, profile ID, session ID, prompt input
materials, and variables. The caller supplies a pointer target for decoded
structured output.
structured output. The response also carries the raw structured output bytes
returned by the runtime so modules can preserve raw payloads in pipeline stage
outputs.
Modules that call the LLM own their prompts, schemas, prompt IDs, validators,
and domain-specific interpretation. Provider adapters should not contain
@@ -59,6 +61,7 @@ Notarius prompt requests into Scriptorium `RunRequest` values. It:
- lets Scriptorium render prompts, call the configured provider, and validate
structured output;
- unmarshals successful JSON into the caller-provided target;
- returns the validated raw structured output bytes to the caller;
- maps token usage and selected profile/model metadata into the Notarius
response and manifest profile recorder.

View File

@@ -178,7 +178,7 @@ Response schema identity:
The extractor adds prompt and response-schema provenance to lane manifest
metadata under `artifact_lanes[].metadata.extractor`. Durable raw output
details belong in the
[D&D spell artifact contract](../integrations/dnd-spell-artifacts.md).
[D&D spell raw output contract](../integrations/dnd-spell-artifacts.md).
The `dnd/scenes` chunker and `dnd/spells` extractor declare optional `players`,
`party`, and `glossary` reference slots accepting UTF-8 plain text, Markdown,

View File

@@ -121,6 +121,8 @@ The runner:
chunk validators;
6. runs each selected artifact lane in sorted resolved order;
7. builds the output encoder and validates logical output file names.
8. passes accepted normalized raw outputs, rejected output records, warnings,
and the manifest to the output encoder.
## Chunk Results
@@ -179,8 +181,8 @@ Within an artifact lane, the runner:
## Validators
The current runner handoff is raw-output based. Extractors, mergers, and
normalizers do not advertise candidate validator chains through their module
interfaces. Runner-side raw validation chains receive the raw module output plus
normalizers do not advertise validator chains through their module interfaces.
Runner-side raw validation chains receive the raw module output plus
stage, lane, module, source, and chunk provenance. Empty raw validation chains
approve output by default.
@@ -208,7 +210,12 @@ On successful execution, the manifest validation status is:
The manifest records run ID, pipeline ID, pipeline digest, module keys, top-level
module metadata, artifact lanes, LLM profile metadata, source digest,
reference provenance, validation status, and timing.
reference provenance, normalized raw output summaries, rejected output
summaries, validation status, and timing. Raw output summaries include lane ID,
normalizer module key, media type, source ID, and response-schema provenance
when present. Rejected output summaries include stage, lane, module, chunk,
validator or reason, message, attempt count, and optional diagnostic artifact
path. The manifest does not include raw output payload bytes.
Singleton pipeline modules may add non-secret metadata by implementing
`contracts.ManifestMetadataProvider`. The runner records that metadata under
@@ -218,3 +225,12 @@ Lane-owned modules may add non-secret metadata through
`artifact_lanes[].metadata`. The runner records extractor, merger, and
normalizer metadata there. The D&D spell extractor uses lane metadata for
prompt and response-schema provenance.
## JSON Output
The production JSON output encoder writes `manifest.json`, `index.json`,
`warnings.json`, `rejected.json`, and one pretty-printed JSON file per accepted
normalized lane output under `lanes/`. It accepts only normalized outputs with
valid `application/json` payloads. Unsupported media types, invalid JSON, unsafe
logical paths, and duplicate sanitized lane file names fail the run before
durable output files are written.

View File

@@ -1,419 +1,18 @@
# Raw Pipeline Data Model Implementation Plan
# Raw Pipeline Implementation
This plan implements the target state in
[`pipeline.md`](pipeline.md). It is intentionally staged for multiple coding
passes. Do not skip stages: later work assumes the contracts and tests from
earlier stages are already in place.
The raw pipeline migration described here has been implemented.
The desired end state is a fixed-shape pipeline:
Current behavior is documented in:
```text
input -> chunk -> extract -> merge -> normalize -> output
```
- [Pipeline Internals](../internal/pipeline.md)
- [Modules](../internal/modules.md)
- [LLM Runtime](../internal/llm.md)
- [Configuration](../config.md)
- [CLI Reference](../cli.md)
- [Operations](../operations.md)
- [Troubleshooting](../troubleshooting.md)
- [JSON Output](../integrations/json-output.md)
- [D&D Spell Raw Output](../integrations/dnd-spell-artifacts.md)
`chunk` produces typed chunk envelopes. `extract`, `merge`, and `normalize`
produce raw byte payload envelopes with media type and provenance. Rejected
outputs do not pass to the next stage.
## Stage 1: Integer Source Units And Chunk Payloads
Update the core source and chunk contracts before changing downstream stages.
Implementation tasks:
- Change `internal/core/source.SourceUnit.ID` from `string` to `int`.
- Change `internal/core/source.SourceRef.StartUnitID` and `EndUnitID` from
`string` to `int`.
- Treat source-unit IDs as positive, stable, input-adapter-owned integers.
`0` and negative IDs should be invalid.
- Update source validation helpers, including unit lookup and source-ref
ordering checks, to use integer IDs.
- Update `contracts.SourceChunk` to include:
- `StartUnitID int`;
- `EndUnitID int`;
- `Content []byte`;
- `MediaType string`.
- Keep `SourceChunk.Units []source.SourceUnit` for framework scheduling,
diagnostics, and modules that need unit metadata.
- Update chunk canonicalization in `internal/framework/pipeline` to enforce
minimal scheduling invariants:
- at least one chunk;
- non-empty chunk ID;
- source ID matches the source document;
- indexes are deterministic and contiguous from zero;
- start/end unit IDs exist and are ordered;
- `Units` are canonical copies from the source document;
- content is non-empty;
- media type is non-empty.
- Update the Seriatim input adapter to parse segment IDs as integers.
Accept JSON numbers and numeric strings only when they are positive integers.
Reject non-numeric IDs such as `seg-001`.
- Update Seriatim examples, fixtures, tests, and integration docs to use
integer segment IDs.
- Update D&D source-reference helper code under
`internal/modules/sharedassets/dnd` so source refs remain integer-valued all
the way into `source.SourceRef`.
- Remove fallback behavior that converts integer LLM unit refs back into string
source-unit IDs.
- Update the generic chunker and D&D scenes chunker to populate
`StartUnitID`, `EndUnitID`, `Content`, and `MediaType`.
- For current transcript chunkers, use `application/json` content containing a
canonical JSON encoding of the chunk's source units. Keep the original
`Units` field populated as framework-owned typed metadata.
Tests to update or add:
- `go test ./internal/core/source`
- `go test ./internal/modules/input/seriatim`
- `go test ./internal/modules/chunk/generic`
- `go test ./internal/modules/chunk/dnd/scenes`
- `go test ./internal/framework/contracts`
- `go test ./internal/framework/pipeline`
Completion criteria:
- All source refs and source-unit IDs in core/framework types are integers.
- Chunkers return extraction-ready content bytes plus media type.
- Framework chunk validation still allows partial coverage and overlap unless a
validator rejects them.
## Stage 2: Raw Module Output Contracts
Replace artifact-candidate stage contracts with raw payload envelopes.
Implementation tasks:
- Add contract types in `internal/framework/contracts` for raw stage payloads.
The exact names may differ, but the contracts must preserve:
- raw bytes;
- media type;
- lane ID;
- stage module key;
- source ID;
- chunk ID and chunk index where applicable;
- response schema ID, name, and version when applicable;
- metadata;
- warnings.
- Recommended shape:
- `ExtractOutput`
- `MergeOutput`
- `NormalizeOutput`
- a small shared payload/provenance helper if it reduces duplication.
- Update `Extractor` so `Extract` returns one raw `ExtractOutput` for one input
chunk instead of `[]artifacts.ArtifactCandidate`.
- Remove `Extractor.ArtifactType`, `Extractor.SchemaVersion`, and
`Extractor.Validators` from the generic extractor contract. Schema and prompt
provenance should come from the returned output and manifest metadata, not
artifact-specific methods.
- Update `Merger` so `Merge` receives ordered accepted `[]ExtractOutput` and
returns one raw `MergeOutput`.
- Update `Normalizer` so `Normalize` receives one accepted `MergeOutput` and
returns one raw `NormalizeOutput`.
- Update `OutputRequest` so output encoders receive the validated normalized
lane outputs rather than approved `Artifact` values.
- Add a raw rejected-output record type for manifests/output files. It should
identify stage, lane, module, chunk when relevant, validator or error reason,
message, attempt count, and diagnostic artifact path if present. It must not
include large raw content by default.
- Leave old artifact/candidate types in `internal/core/artifacts` only if
temporary compatibility code still needs them during the migration. They
should not remain in the final stage contracts.
Tests to update or add:
- Contract composition tests proving fake chunk, extract, merge, normalize, and
output modules compose using raw envelopes.
- Defensive-copy tests for raw content and metadata maps.
- Registry integration tests proving the new interfaces build and run.
Completion criteria:
- Framework contracts no longer require extractors, mergers, normalizers, or
output encoders to use `ArtifactCandidate` or `Artifact`.
- Stage outputs carry raw bytes and media type.
## Stage 3: Config, References, Profiles, And Retries
Make `merge` a first-class LLM/profile/reference stage and add retry policy.
Implementation tasks:
- Update `pipeline.referenceSlotStage` so `merge` modules may declare reference
slots.
- Add `MergeReferences` to `ResolvedArtifactLane`.
- Resolve and materialize merge references using the same declared-slot model
used by `chunk`, `extract`, and `normalize`.
- Update reference provenance so merge-stage references appear in manifests and
resolved reference diagnostics.
- Update config parsing, config validation, redaction, cloning, and examples so
`artifacts.<lane>.merge.references` is valid.
- Update CLI `--reference` selector parsing and ambiguity checks to support:
- `merge.<slot>=<path>` where unambiguous;
- `<lane>.merge.<slot>=<path>`;
- lane-qualified shorthand only when it resolves unambiguously.
- Update `--without-reference` behavior to support merge-stage slots.
- Update `ModuleStage` LLM-capable helpers so `chunk`, `extract`, `merge`, and
`normalize` are all considered LLM-capable.
- Update `--llm-profile` override handling so it applies to `chunk`, every lane
`extract`, every lane `merge`, and every lane `normalize`.
- Update explicit Scriptorium profile validation so it inspects only those four
LLM-capable stages.
- Add `retries` to `pipeline.ModuleBinding` as a non-negative integer count of
extra attempts after the first attempt. Default is `0`.
- Apply `retries` only to `chunk`, `extract`, `merge`, and `normalize` runtime
execution. Keep input and output retry behavior out of scope.
- Validate `retries >= 0` in config validation.
- Ensure redacted config preserves retry counts.
Tests to update or add:
- Config accepts merge references and rejects unknown/undeclared merge slots.
- CLI reference selectors work for merge and remain strict for ambiguous slots.
- `--llm-profile` overrides merge bindings as well as chunk/extract/normalize.
- Explicit profile validation includes merge and ignores non-LLM stages.
- Negative retries are rejected.
Completion criteria:
- Merge has the same runtime plumbing class as chunk/extract/normalize.
- Retry policy is representable in pipeline config and resolved bindings.
## Stage 4: Runner Orchestration And Retry Semantics
Rewrite pipeline execution around raw envelopes.
Implementation tasks:
- Introduce a small retry helper in `internal/framework/pipeline` that reruns
the same module with the same input when:
- the module returns a framework-level error;
- the module returns output that is rejected by that stage's validator chain.
- Preserve context cancellation: never retry after `ctx.Err()` is non-nil.
- Record attempt count and compact attempt diagnostics for manifests/output.
- Do not log or manifest raw payload bytes by default.
- Run chunk once per source input, with retries from the chunk binding.
- After successful chunk execution, run chunk validation when validator mapping
support exists. A rejected chunk result after final retry should stop
downstream execution and produce a rejected run record.
- Run extract once per accepted chunk, with per-chunk retries from the extract
binding.
- Omit rejected extract outputs from merge input.
- If a lane has no accepted extract outputs after retries, skip merge and
normalize for that lane, record rejected output details, and continue to the
next lane.
- Run merge once per lane with accepted extract outputs, with retries from the
merge binding.
- If merge output is rejected after final retry, omit that lane from normalize
and output.
- Run normalize once per accepted merge output, with retries from the normalize
binding.
- If normalize output is rejected after final retry, omit that lane from output.
- Treat validator rejection as non-fatal run outcome. Treat unrecovered
framework-level execution errors as run failures.
- Preserve deterministic ordering:
- chunks by chunk index;
- extract outputs by chunk index;
- lane outputs by resolved lane order.
- Continue to collect warnings from every successful module attempt whose output
is used. For failed attempts, record compact attempt diagnostics rather than
promoting warnings as final-stage warnings unless the implementation already
has a clear warning policy.
- Set manifest validation status to:
- `approved` when all produced normalized lane outputs pass;
- `rejected` when one or more module outputs are rejected;
- failed-run status through existing failure handling for unrecovered runtime
errors.
Validation seam scope:
- Add the runner-side raw validation seam needed by this data-model refactor.
- Represent the validation boundary as raw module output plus stage provenance,
consistent with [`validation.md`](validation.md).
- Support empty validator chains as approval.
- Add fake validator/test-only coverage for approval, rejection, and
retry-on-rejection behavior.
- Do not implement the full production validator package migration, default
validator mappings, or concrete generic validators in this plan. Those remain
owned by the validation roadmap.
Tests to update or add:
- Runner passes chunk content and media type to extractors.
- Runner passes ordered accepted extract outputs to merge.
- Runner passes accepted merge output to normalize.
- Rejected extract output is omitted from merge input.
- A lane with no accepted extracts is omitted and recorded.
- Rejected merge output prevents normalize for that lane.
- Rejected normalize output prevents output for that lane.
- Retry reruns the same module input after framework error.
- Retry reruns the same module input after validator rejection.
- Retry stops after configured attempts and records attempt count.
- Context cancellation stops retries.
Completion criteria:
- The runner no longer depends on candidate materialization between extract,
merge, normalize, and output.
- Rejected outputs never pass to the next stage.
## Stage 5: Production Module Migration
Update concrete modules to the new contracts.
Implementation tasks:
- Update `internal/modules/extract/dnd/spells`:
- call Scriptorium with the existing prompt/schema;
- capture the returned raw JSON response bytes;
- return one `ExtractOutput` with `application/json` media type and response
schema provenance;
- do not convert spell casts into `ArtifactCandidate`;
- do not run shape/source-ref/source-relatedness validation inside the module.
- Keep D&D spell prompt and schema asset metadata in manifest metadata.
- Update `internal/modules/merge/appendorder` to merge raw JSON extract outputs
deterministically:
- preserve input order by chunk index;
- if all extract outputs are JSON objects with one common top-level array
field, concatenate those arrays into one JSON object under that field;
- otherwise produce a JSON array containing each extract output's decoded JSON
value in order;
- reject non-JSON media types with a clear error unless a future option
deliberately supports them.
- Update `internal/modules/normalize/noop` to pass accepted merge output through
unchanged as normalized output.
- Update chunk modules from Stage 1 as needed after interface changes.
- Remove or quarantine old candidate validators from D&D module packages if
they no longer compile. Their replacement belongs under the validation
roadmap in `internal/validators`.
- Update fake/test modules throughout the repository to the raw contracts.
Tests to update or add:
- D&D spells extractor returns raw JSON with schema provenance and no candidate
materialization.
- Append-order merger concatenates common top-level JSON arrays.
- Append-order merger falls back to ordered JSON value arrays when shapes differ.
- Append-order merger rejects invalid JSON and non-JSON media types clearly.
- No-op normalizer returns defensive copies of merge output bytes and metadata.
- Production catalog still registers all modules successfully.
Completion criteria:
- Production modules compile against raw contracts.
- D&D spell extraction no longer rejects spell candidates inside the extractor.
## Stage 6: Output, Manifest, Diagnostics, And Files
Update durable output around normalized lane outputs.
Implementation tasks:
- Update `RunOutput` to carry normalized lane outputs and rejected module-output
records instead of approved/rejected artifacts.
- Update `artifacts.RunManifest` or introduce a more accurately named manifest
package if the old artifact naming becomes misleading. Prefer the smallest
rename that keeps manifest output clear and avoids a broad unrelated refactor.
- Preserve existing manifest fields that remain meaningful:
- pipeline ID and digest;
- module keys;
- lane IDs;
- source digest;
- LLM profile provenance;
- reference provenance;
- module metadata;
- validation status.
- Add manifest/reporting fields needed for raw outputs:
- normalized lane output media type;
- output schema provenance where present;
- rejected stage/lane/chunk/module information;
- retry attempt counts.
- Update `internal/modules/output/json`:
- write `manifest.json`;
- write `index.json`;
- write `warnings.json`;
- write `rejected.json`;
- write one file per normalized lane output.
- For the JSON output encoder, accept `application/json` normalized outputs and
write them as raw JSON files. Recommended file path:
`lanes/<safe-lane-id>.json`.
- Reject non-JSON normalized output media types in the JSON encoder with a clear
error. Future text/Markdown/binary encoders can support other media types.
- Keep output file name validation strict.
- Ensure diagnostics redaction still prevents raw source, prompt, reference,
schema, and payload bytes from appearing in ordinary errors or manifests.
Tests to update or add:
- JSON output writes lane output files and index entries deterministically.
- JSON output rejects invalid JSON bytes and unsupported media types.
- Rejected output records are written without raw payload bytes.
- Manifest includes raw-output provenance and retry attempt counts.
- Diagnostics tests still prove sensitive/large content is not emitted.
Completion criteria:
- Durable output no longer assumes artifact candidates.
- JSON output remains deterministic and safe.
## Stage 7: Documentation, Examples, And Full Validation
After behavior is implemented, update canonical current-behavior docs.
Implementation tasks:
- Update `docs/internal/pipeline.md` for the raw data model.
- Update `docs/internal/modules.md` for new module contracts.
- Update `docs/internal/llm.md` for merge-stage LLM/profile plumbing.
- Update `docs/config.md` for:
- integer source-unit expectations where config examples include source refs;
- merge references;
- module `retries`;
- LLM profile override scope.
- Update `docs/cli.md` for merge reference selectors and profile behavior.
- Update `docs/integrations/json-output.md` for lane output files.
- Update `docs/integrations/dnd-spell-artifacts.md`; if the artifact-specific
contract is no longer durable, rename or replace it with a D&D spell raw-output
integration doc.
- Update `docs/troubleshooting.md` for common raw-output, media-type, retry, and
validation failures.
- Update examples so Seriatim segment IDs are positive integers and expected
output shape matches lane-normalized output.
- Replace `docs/roadmap/implementation.md` with a completed-note document only
after this implementation is finished and reviewed.
- Leave `docs/roadmap/pipeline.md` as target-state roadmap context until the
feature is implemented, then move implemented behavior into canonical docs and
remove stale roadmap language.
Validation commands:
```sh
go test ./internal/core/source
go test ./internal/framework/contracts
go test ./internal/framework/pipeline
go test ./internal/core/config
go test ./internal/cli
go test ./internal/modules/input/seriatim
go test ./internal/modules/chunk/generic
go test ./internal/modules/chunk/dnd/scenes
go test ./internal/modules/extract/dnd/spells
go test ./internal/modules/merge/appendorder
go test ./internal/modules/normalize/noop
go test ./internal/modules/output/json
go test ./...
go vet ./...
go build ./cmd/notarius
```
Documentation inspection:
```sh
rg -n "ArtifactCandidate|approved artifacts|spell artifact|start_unit_id\": \"|end_unit_id\": \"" docs examples internal
rg -n "chunk`, `extract`, and `normalize|deterministic-only stage|json.RawMessage" docs/roadmap docs/internal docs/config.md docs/cli.md
```
Expected result:
- Remaining `ArtifactCandidate` mentions are only in deliberately retained
compatibility code or historical roadmap context.
- Current-behavior docs describe raw lane outputs, integer source-unit IDs,
merge LLM/reference support, and retry behavior.
Follow-up validator work remains tracked separately in
[Validation Roadmap](validation.md).

View File

@@ -1,403 +1,12 @@
# Raw Pipeline Data Model
This roadmap defines the target data model for the pipeline stages after input.
It is paired with the validation system roadmap in
[`validation.md`](validation.md): validators should evaluate module outputs, but
the pipeline first needs a raw-output handoff model that does not require every
module response shape to be represented by bespoke Go structs.
The raw pipeline data model has been implemented.
## Goals
Current behavior is documented in:
- Make typed chunk envelopes and raw module-output envelopes first-class
pipeline handoffs.
- Keep framework-owned provenance around raw payloads so runs remain
deterministic, ordered, and auditable.
- Avoid requiring each extractor, merger, or normalizer response schema to have
matching Go structs.
- Let extract modules produce one raw output per input chunk.
- Let merge modules receive accepted extract outputs and merge them into one raw
payload.
- Let normalize modules optionally post-process accepted merge output.
- Let output modules write normalized output bytes directly, along with
manifests, warnings, rejected outputs, and diagnostics as appropriate.
- Preserve the fixed workflow shape:
- [Pipeline Internals](../internal/pipeline.md)
- [Modules](../internal/modules.md)
- [JSON Output](../integrations/json-output.md)
```text
input -> chunk -> extract -> merge -> normalize -> output
```
## Non-Goals
- Do not turn the pipeline into an arbitrary DAG or general workflow language.
- Do not make validators mutate, materialize, or rewrite module outputs.
- Do not require a generic source-reference convention for every raw payload in
this pass.
- Do not eliminate typed framework data where the framework genuinely needs it,
such as input source documents and chunk boundaries.
## Cross-Stage Runtime Plumbing
LLM/profile/reference plumbing should be available to every stage that may need
LLM-backed or reference-aware module behavior: `chunk`, `extract`, `merge`, and
`normalize`.
For all four stages, the framework should provide the same categories of runtime
support where the concrete module contract needs them:
- configured LLM profile and profile override handling;
- Scriptorium client access through framework-owned LLM contracts;
- session ID propagation;
- declared reference slots and resolved reference bindings;
- module options and metadata;
- prompt, schema, profile, and reference provenance for manifests and
diagnostics.
`merge` must not be treated as a deterministic-only stage. A merge module may be
simple and deterministic, but it may also be LLM-backed, reference-aware, and
validator-gated in the same way as `chunk`, `extract`, and `normalize`.
## Chunk Stage Target Model
The chunk stage receives the canonical `SourceDocument` from the input stage,
plus any original source input material needed for prompt construction. The
`SourceDocument` remains the source of truth for source identity, source-unit
ordering, source-unit IDs, and source provenance. Original raw input material
may be supplied to LLM-backed chunkers, but it should not replace the
`SourceDocument` as the framework handoff.
Input adapters are responsible for assigning stable integer source-unit IDs. How
those IDs are assigned depends on the input format. A numbered JSON transcript
may map source units directly to transcript segment numbers; a PDF input adapter
may assign page or extracted-text unit numbers; an adapter for unordered source
material may assign deterministic IDs as part of input parsing. Downstream
framework code should treat those IDs as opaque integers owned by the input
adapter.
The chunk module returns one or more ordered `SourceChunk` envelopes. A single
chunk representing the whole source is valid and should be supported.
Each chunk should include framework-readable provenance and ordering metadata,
plus content suitable for extraction. The content may be JSON, plain text,
Markdown, PDF page text, or another module/input-specific representation, as
long as the framework can still associate the chunk with the source and preserve
deterministic order.
Conceptually:
```go
type SourceChunk struct {
ID string
SourceID string
Index int
StartUnitID int
EndUnitID int
Content []byte
MediaType string
Units []source.SourceUnit
Metadata map[string]any
}
```
The exact implementation shape can differ, but it should preserve:
- chunk identity;
- source identity;
- deterministic chunk order;
- source locator or range, normally integer start/end unit IDs;
- chunk content and media type;
- optional source-unit projection when useful;
- metadata needed by extractors, validators, manifests, diagnostics, and output
encoders.
The framework owns the minimal invariants required to schedule extract work:
- the chunker returns at least one chunk;
- chunk IDs and indexes are stable and non-empty;
- chunk source IDs match the source document;
- chunk ordering is deterministic;
- chunk provenance is sufficient to trace the chunk back to the input source.
Domain-specific chunk acceptability belongs in validators. For example, full
coverage, no gaps, no overlap, scene metadata quality, expected media type, and
D&D scene-boundary policy should be explicit validator concerns rather than
hidden framework rules, except where a minimal invariant is required for
extraction to run safely.
## Extract Stage Target Model
The extract stage receives one input chunk from the chunk stage and produces one
raw extract output for that chunk.
The extractor owns:
- prompt selection;
- input material assembly;
- Scriptorium request construction;
- response-schema selection;
- LLM profile and session usage;
- returned raw payload metadata.
The framework owns:
- chunk iteration;
- deterministic ordering;
- association between each extract output and its input chunk;
- execution errors when an extractor does not return output;
- handoff of returned output to validation and later stages.
An extract output should be an envelope, not just raw bytes. Conceptually:
```go
type ExtractOutput struct {
LaneID string
ExtractorKey string
ChunkID string
ChunkIndex int
SourceID string
RawContent []byte
MediaType string
SchemaID string
SchemaName string
SchemaVersion string
Metadata map[string]any
Warnings []contracts.Warning
}
```
The exact implementation shape can differ, but it should preserve:
- raw returned content;
- chunk provenance;
- chunk order;
- source identity;
- module identity;
- response schema provenance;
- metadata needed by validators, mergers, manifests, diagnostics, and output
encoders.
Extractor modules should not be required to convert raw LLM output into
`ArtifactCandidate` values. Domain-specific Go projection may still exist for
specific modules when it is useful, but it should not be the generic pipeline
contract.
Rejected extract outputs are not passed to merge. This keeps downstream
contracts simple and makes rejection behavior explicit. Support for passing
rejected outputs forward as marked data is deferred future work.
## Merge Stage Target Model
The merge stage receives the accepted extract outputs for a lane. Each extract
output corresponds to one input chunk and carries enough metadata to recover the
original chunk order.
The merger owns:
- merge/reconciliation strategy;
- deterministic or LLM-backed merge logic;
- prompt and schema usage when LLM-backed;
- the shape of merged raw output.
The framework owns:
- passing the ordered accepted extract-output set to the merger;
- lane identity;
- source context;
- references;
- runtime LLM plumbing;
- validation handoff for merged output;
- manifest provenance.
Simple merge modules may concatenate raw extract outputs in chunk order. Other
merge modules may deterministically merge JSON documents, reconcile duplicates,
or use an LLM to produce a more coherent merged result.
Conceptually:
```go
type MergeRequest struct {
LaneID string
Source *source.SourceDocument
ExtractOutputs []ExtractOutput
SourceInput contracts.LLMInputMaterial
SessionID string
References contracts.ReferenceSet
LLMClient contracts.StructuredLLMClient
LLMProfile string
Options map[string]any
Metadata map[string]any
}
type MergeOutput struct {
LaneID string
MergerKey string
RawContent []byte
MediaType string
SchemaID string
SchemaName string
SchemaVersion string
Metadata map[string]any
Warnings []contracts.Warning
}
```
The exact implementation shape can differ, but the key contract is that
merge consumes ordered accepted extract outputs and produces one merged raw
output for the lane.
If no accepted extract outputs remain for a lane, the framework should not pass
rejected outputs to the merger. The lane should produce no merge output and
should be reported as rejected or omitted according to run reporting policy.
## Normalize Stage Target Model
The normalize stage receives accepted merge output for a lane and optionally
post-processes it into the final normalized raw payload for that lane.
Normalization may be a no-op, deterministic cleanup, schema conversion, or an
LLM-backed post-processing pass.
The normalizer owns:
- post-merge processing strategy;
- deterministic or LLM-backed normalization logic;
- prompt and schema usage when LLM-backed;
- the shape of normalized raw output.
The framework owns:
- passing accepted merge output to the normalizer;
- lane identity;
- source context;
- references;
- runtime LLM plumbing;
- validation handoff for normalized output;
- manifest provenance.
Conceptually:
```go
type NormalizeRequest struct {
LaneID string
Source *source.SourceDocument
MergeOutput MergeOutput
SourceInput contracts.LLMInputMaterial
SessionID string
References contracts.ReferenceSet
LLMClient contracts.StructuredLLMClient
LLMProfile string
Options map[string]any
Metadata map[string]any
}
type NormalizeOutput struct {
LaneID string
NormalizerKey string
RawContent []byte
MediaType string
SchemaID string
SchemaName string
SchemaVersion string
Metadata map[string]any
Warnings []contracts.Warning
}
```
The exact implementation shape can differ, but the key contract is that
normalize consumes one accepted merged output and produces one normalized raw
output for the lane.
## Output Stage Target Model
The output stage receives validated normalized output bytes and writes them to
its configured destination.
Output modules should not require normalized output to be converted into
`Artifact` values. Output modules should use the normalized output media type to
decide how to serialize or wrap the payload. A JSON output module can write raw
normalized JSON directly, while a text or Markdown output module can write text
payloads directly. Output modules may also write run manifests, warnings,
rejected outputs, and indexes.
Output modules may still choose to provide convenience layouts, grouping, or file
naming conventions, but those should be output concerns rather than constraints
on extractor or normalizer response schemas.
## Validation Relationship
Validation chains should attach to returned module outputs, not to hidden
module-internal conversions.
For chunk:
```text
chunk(input) -> SourceChunk set -> chunk validators -> extract input
```
For extract:
```text
extract(chunk) -> raw ExtractOutput -> extract validators -> merge input
```
For merge:
```text
merge(extract outputs) -> raw MergeOutput -> merge validators -> normalize input
```
For normalize:
```text
normalize(merge output) -> raw NormalizeOutput -> normalize validators -> output
```
Validators must be read-only. They inspect raw output and metadata, return
accept/reject decisions and warnings, and do not rewrite output.
An empty validator chain approves returned output for that validation point.
Framework/runtime errors remain separate from validation rejections: if a module
or Scriptorium call fails before output is returned, the pipeline reports an
execution error rather than asking validators to evaluate nonexistent output.
Rejected outputs do not pass to the next stage. An empty validator chain still
approves returned output for that validation point.
## Retry Policy
A retry means re-running the same module with the same input after that module
fails to produce valid output. Failure to produce valid output includes both:
- framework-level execution errors, such as module errors, Scriptorium errors,
provider errors, or missing returned output;
- validator rejection of returned output.
Retries should be configurable at the pipeline or lane level. Chunk retries are
per source input. Extract retries are per chunk. Merge and normalize retries,
when configured, are per lane. Retry attempts should preserve deterministic
reporting: the final accepted or rejected output should record attempt count and
enough diagnostics/provenance to understand prior failures without leaking
secrets or large payloads by default.
## Relationship To Validation Roadmap
This roadmap defines the data model that validation should evaluate. The
validation system roadmap in [`validation.md`](validation.md) defines validator
registration, mapping, execution classes, and concrete validator behavior.
The shared boundary between the roadmaps is a returned module output: validators
inspect typed chunk output or raw module-output envelopes and decide whether
that output may continue through the pipeline.
Validator mapping and concrete validator behavior remain tracked separately in
[Validation Roadmap](validation.md).

View File

@@ -200,7 +200,8 @@ Fix:
- Or use an existing Scriptorium profile ID with `--llm-profile`.
Use `--llm-profile <id>` when one run should force every LLM-backed binding to
the same Scriptorium profile.
the same Scriptorium profile. The override applies to effective chunk, extract,
merge, and normalize bindings.
## Missing API Key Environment Variable
@@ -285,17 +286,58 @@ Symptoms include:
- `create output directory`
- `write output file`
- `output file name must`
- `unsupported media type`
- `invalid JSON`
Fix:
- Ensure `--output-dir` points to a directory path or a path that can be
created.
- Check filesystem permissions and available disk space.
- The production JSON output encoder writes lane payloads under `lanes/` and
accepts only valid `application/json` normalized outputs. If an error names an
unsupported media type or invalid JSON, inspect the lane's merge and normalize
module output.
- If diagnostics were retained, inspect `run-report.json`, `run-manifest.json`,
and `error.log`.
The CLI rejects unsafe logical output paths before writing files.
## Raw Output Rejection
Symptoms include a successful run with:
- `validation_status` set to `rejected`;
- non-empty `rejected.json`;
- `rejected_outputs` entries in `manifest.json`.
Explanation and fixes:
- Validator rejection is a non-fatal run outcome. Rejected module outputs do not
pass to the next pipeline stage.
- Check `rejected.json` for the stage, lane, module, chunk, validator, reason,
message, and attempt count.
- Increase a module binding's `retries` only when re-running the same module
input can reasonably produce an acceptable output.
- If rejection is deterministic, fix the source input, module configuration, or
validator configuration rather than adding retries.
## Retry Exhaustion
Symptoms include:
- errors containing `failed after ... attempt(s)`;
- rejected output records with `attempt_count` greater than `1`.
Fix:
- `retries` is the number of extra attempts after the first attempt for chunk,
extract, merge, and normalize bindings.
- Framework-level errors after the last attempt fail the run.
- Validator rejections after the last attempt are recorded as rejected outputs.
- Check retained `error.log`, `run-manifest.json`, and `rejected.json` for the
operation, module key, lane, chunk, and attempt count.
## Diagnostics Directory Surprise
Symptom: the diagnostics directory is missing after a successful run.