From a9d8505cdbc54fbab7c5112cdb6a0bd1af38281b Mon Sep 17 00:00:00 2001 From: Eric Rakestraw Date: Tue, 7 Jul 2026 19:32:15 +0000 Subject: [PATCH] Document raw pipeline completion --- docs/config.md | 6 +- docs/internal/llm.md | 5 +- docs/internal/modules.md | 2 +- docs/internal/pipeline.md | 22 +- docs/roadmap/implementation.md | 429 ++------------------------------- docs/roadmap/pipeline.md | 405 +------------------------------ docs/troubleshooting.md | 44 +++- 7 files changed, 91 insertions(+), 822 deletions(-) diff --git a/docs/config.md b/docs/config.md index f5b922e..608dc28 100644 --- a/docs/config.md +++ b/docs/config.md @@ -248,9 +248,9 @@ merge, and normalize binding. | chunk | `generic` | Splits source units into ordered chunks. | | chunk | `dnd/scenes` | Uses an LLM to split transcript source units into D&D scenes. | | extract | `dnd/spells` | Extracts D&D spell raw outputs. | -| merge | `appendorder` | Keeps raw extract outputs in append order. | +| merge | `appendorder` | Merges JSON raw extract outputs in chunk order. | | normalize | `noop` | Passes merged raw outputs through unchanged. | -| output | `json` | Produces JSON output files. | +| output | `json` | Produces JSON output files for normalized `application/json` lanes. | The `generic` chunker accepts: @@ -308,6 +308,6 @@ Pipeline resolution additionally checks: - required module keys are present; - module keys are registered for the expected slot; - module capability requirements are satisfied; -- bound reference slots are declared by selected chunk, extractor, or +- bound reference slots are declared by selected chunk, extractor, merger, or normalizer targets; - required reference slots are bound for selected targets. diff --git a/docs/internal/llm.md b/docs/internal/llm.md index 16d90c4..76bb709 100644 --- a/docs/internal/llm.md +++ b/docs/internal/llm.md @@ -15,7 +15,9 @@ CompleteStructured(ctx, request, out) (response, error) The request contains prompt ID/version, profile ID, session ID, prompt input materials, and variables. The caller supplies a pointer target for decoded -structured output. +structured output. The response also carries the raw structured output bytes +returned by the runtime so modules can preserve raw payloads in pipeline stage +outputs. Modules that call the LLM own their prompts, schemas, prompt IDs, validators, and domain-specific interpretation. Provider adapters should not contain @@ -59,6 +61,7 @@ Notarius prompt requests into Scriptorium `RunRequest` values. It: - lets Scriptorium render prompts, call the configured provider, and validate structured output; - unmarshals successful JSON into the caller-provided target; +- returns the validated raw structured output bytes to the caller; - maps token usage and selected profile/model metadata into the Notarius response and manifest profile recorder. diff --git a/docs/internal/modules.md b/docs/internal/modules.md index 6741eb8..292789c 100644 --- a/docs/internal/modules.md +++ b/docs/internal/modules.md @@ -178,7 +178,7 @@ Response schema identity: The extractor adds prompt and response-schema provenance to lane manifest metadata under `artifact_lanes[].metadata.extractor`. Durable raw output details belong in the -[D&D spell artifact contract](../integrations/dnd-spell-artifacts.md). +[D&D spell raw output contract](../integrations/dnd-spell-artifacts.md). The `dnd/scenes` chunker and `dnd/spells` extractor declare optional `players`, `party`, and `glossary` reference slots accepting UTF-8 plain text, Markdown, diff --git a/docs/internal/pipeline.md b/docs/internal/pipeline.md index 2c68306..2da2277 100644 --- a/docs/internal/pipeline.md +++ b/docs/internal/pipeline.md @@ -121,6 +121,8 @@ The runner: chunk validators; 6. runs each selected artifact lane in sorted resolved order; 7. builds the output encoder and validates logical output file names. +8. passes accepted normalized raw outputs, rejected output records, warnings, + and the manifest to the output encoder. ## Chunk Results @@ -179,8 +181,8 @@ Within an artifact lane, the runner: ## Validators The current runner handoff is raw-output based. Extractors, mergers, and -normalizers do not advertise candidate validator chains through their module -interfaces. Runner-side raw validation chains receive the raw module output plus +normalizers do not advertise validator chains through their module interfaces. +Runner-side raw validation chains receive the raw module output plus stage, lane, module, source, and chunk provenance. Empty raw validation chains approve output by default. @@ -208,7 +210,12 @@ On successful execution, the manifest validation status is: The manifest records run ID, pipeline ID, pipeline digest, module keys, top-level module metadata, artifact lanes, LLM profile metadata, source digest, -reference provenance, validation status, and timing. +reference provenance, normalized raw output summaries, rejected output +summaries, validation status, and timing. Raw output summaries include lane ID, +normalizer module key, media type, source ID, and response-schema provenance +when present. Rejected output summaries include stage, lane, module, chunk, +validator or reason, message, attempt count, and optional diagnostic artifact +path. The manifest does not include raw output payload bytes. Singleton pipeline modules may add non-secret metadata by implementing `contracts.ManifestMetadataProvider`. The runner records that metadata under @@ -218,3 +225,12 @@ Lane-owned modules may add non-secret metadata through `artifact_lanes[].metadata`. The runner records extractor, merger, and normalizer metadata there. The D&D spell extractor uses lane metadata for prompt and response-schema provenance. + +## JSON Output + +The production JSON output encoder writes `manifest.json`, `index.json`, +`warnings.json`, `rejected.json`, and one pretty-printed JSON file per accepted +normalized lane output under `lanes/`. It accepts only normalized outputs with +valid `application/json` payloads. Unsupported media types, invalid JSON, unsafe +logical paths, and duplicate sanitized lane file names fail the run before +durable output files are written. diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md index 886497d..975d441 100644 --- a/docs/roadmap/implementation.md +++ b/docs/roadmap/implementation.md @@ -1,419 +1,18 @@ -# Raw Pipeline Data Model Implementation Plan +# Raw Pipeline Implementation -This plan implements the target state in -[`pipeline.md`](pipeline.md). It is intentionally staged for multiple coding -passes. Do not skip stages: later work assumes the contracts and tests from -earlier stages are already in place. +The raw pipeline migration described here has been implemented. -The desired end state is a fixed-shape pipeline: +Current behavior is documented in: -```text -input -> chunk -> extract -> merge -> normalize -> output -``` +- [Pipeline Internals](../internal/pipeline.md) +- [Modules](../internal/modules.md) +- [LLM Runtime](../internal/llm.md) +- [Configuration](../config.md) +- [CLI Reference](../cli.md) +- [Operations](../operations.md) +- [Troubleshooting](../troubleshooting.md) +- [JSON Output](../integrations/json-output.md) +- [D&D Spell Raw Output](../integrations/dnd-spell-artifacts.md) -`chunk` produces typed chunk envelopes. `extract`, `merge`, and `normalize` -produce raw byte payload envelopes with media type and provenance. Rejected -outputs do not pass to the next stage. - -## Stage 1: Integer Source Units And Chunk Payloads - -Update the core source and chunk contracts before changing downstream stages. - -Implementation tasks: - -- Change `internal/core/source.SourceUnit.ID` from `string` to `int`. -- Change `internal/core/source.SourceRef.StartUnitID` and `EndUnitID` from - `string` to `int`. -- Treat source-unit IDs as positive, stable, input-adapter-owned integers. - `0` and negative IDs should be invalid. -- Update source validation helpers, including unit lookup and source-ref - ordering checks, to use integer IDs. -- Update `contracts.SourceChunk` to include: - - `StartUnitID int`; - - `EndUnitID int`; - - `Content []byte`; - - `MediaType string`. -- Keep `SourceChunk.Units []source.SourceUnit` for framework scheduling, - diagnostics, and modules that need unit metadata. -- Update chunk canonicalization in `internal/framework/pipeline` to enforce - minimal scheduling invariants: - - at least one chunk; - - non-empty chunk ID; - - source ID matches the source document; - - indexes are deterministic and contiguous from zero; - - start/end unit IDs exist and are ordered; - - `Units` are canonical copies from the source document; - - content is non-empty; - - media type is non-empty. -- Update the Seriatim input adapter to parse segment IDs as integers. - Accept JSON numbers and numeric strings only when they are positive integers. - Reject non-numeric IDs such as `seg-001`. -- Update Seriatim examples, fixtures, tests, and integration docs to use - integer segment IDs. -- Update D&D source-reference helper code under - `internal/modules/sharedassets/dnd` so source refs remain integer-valued all - the way into `source.SourceRef`. -- Remove fallback behavior that converts integer LLM unit refs back into string - source-unit IDs. -- Update the generic chunker and D&D scenes chunker to populate - `StartUnitID`, `EndUnitID`, `Content`, and `MediaType`. -- For current transcript chunkers, use `application/json` content containing a - canonical JSON encoding of the chunk's source units. Keep the original - `Units` field populated as framework-owned typed metadata. - -Tests to update or add: - -- `go test ./internal/core/source` -- `go test ./internal/modules/input/seriatim` -- `go test ./internal/modules/chunk/generic` -- `go test ./internal/modules/chunk/dnd/scenes` -- `go test ./internal/framework/contracts` -- `go test ./internal/framework/pipeline` - -Completion criteria: - -- All source refs and source-unit IDs in core/framework types are integers. -- Chunkers return extraction-ready content bytes plus media type. -- Framework chunk validation still allows partial coverage and overlap unless a - validator rejects them. - -## Stage 2: Raw Module Output Contracts - -Replace artifact-candidate stage contracts with raw payload envelopes. - -Implementation tasks: - -- Add contract types in `internal/framework/contracts` for raw stage payloads. - The exact names may differ, but the contracts must preserve: - - raw bytes; - - media type; - - lane ID; - - stage module key; - - source ID; - - chunk ID and chunk index where applicable; - - response schema ID, name, and version when applicable; - - metadata; - - warnings. -- Recommended shape: - - `ExtractOutput` - - `MergeOutput` - - `NormalizeOutput` - - a small shared payload/provenance helper if it reduces duplication. -- Update `Extractor` so `Extract` returns one raw `ExtractOutput` for one input - chunk instead of `[]artifacts.ArtifactCandidate`. -- Remove `Extractor.ArtifactType`, `Extractor.SchemaVersion`, and - `Extractor.Validators` from the generic extractor contract. Schema and prompt - provenance should come from the returned output and manifest metadata, not - artifact-specific methods. -- Update `Merger` so `Merge` receives ordered accepted `[]ExtractOutput` and - returns one raw `MergeOutput`. -- Update `Normalizer` so `Normalize` receives one accepted `MergeOutput` and - returns one raw `NormalizeOutput`. -- Update `OutputRequest` so output encoders receive the validated normalized - lane outputs rather than approved `Artifact` values. -- Add a raw rejected-output record type for manifests/output files. It should - identify stage, lane, module, chunk when relevant, validator or error reason, - message, attempt count, and diagnostic artifact path if present. It must not - include large raw content by default. -- Leave old artifact/candidate types in `internal/core/artifacts` only if - temporary compatibility code still needs them during the migration. They - should not remain in the final stage contracts. - -Tests to update or add: - -- Contract composition tests proving fake chunk, extract, merge, normalize, and - output modules compose using raw envelopes. -- Defensive-copy tests for raw content and metadata maps. -- Registry integration tests proving the new interfaces build and run. - -Completion criteria: - -- Framework contracts no longer require extractors, mergers, normalizers, or - output encoders to use `ArtifactCandidate` or `Artifact`. -- Stage outputs carry raw bytes and media type. - -## Stage 3: Config, References, Profiles, And Retries - -Make `merge` a first-class LLM/profile/reference stage and add retry policy. - -Implementation tasks: - -- Update `pipeline.referenceSlotStage` so `merge` modules may declare reference - slots. -- Add `MergeReferences` to `ResolvedArtifactLane`. -- Resolve and materialize merge references using the same declared-slot model - used by `chunk`, `extract`, and `normalize`. -- Update reference provenance so merge-stage references appear in manifests and - resolved reference diagnostics. -- Update config parsing, config validation, redaction, cloning, and examples so - `artifacts..merge.references` is valid. -- Update CLI `--reference` selector parsing and ambiguity checks to support: - - `merge.=` where unambiguous; - - `.merge.=`; - - lane-qualified shorthand only when it resolves unambiguously. -- Update `--without-reference` behavior to support merge-stage slots. -- Update `ModuleStage` LLM-capable helpers so `chunk`, `extract`, `merge`, and - `normalize` are all considered LLM-capable. -- Update `--llm-profile` override handling so it applies to `chunk`, every lane - `extract`, every lane `merge`, and every lane `normalize`. -- Update explicit Scriptorium profile validation so it inspects only those four - LLM-capable stages. -- Add `retries` to `pipeline.ModuleBinding` as a non-negative integer count of - extra attempts after the first attempt. Default is `0`. -- Apply `retries` only to `chunk`, `extract`, `merge`, and `normalize` runtime - execution. Keep input and output retry behavior out of scope. -- Validate `retries >= 0` in config validation. -- Ensure redacted config preserves retry counts. - -Tests to update or add: - -- Config accepts merge references and rejects unknown/undeclared merge slots. -- CLI reference selectors work for merge and remain strict for ambiguous slots. -- `--llm-profile` overrides merge bindings as well as chunk/extract/normalize. -- Explicit profile validation includes merge and ignores non-LLM stages. -- Negative retries are rejected. - -Completion criteria: - -- Merge has the same runtime plumbing class as chunk/extract/normalize. -- Retry policy is representable in pipeline config and resolved bindings. - -## Stage 4: Runner Orchestration And Retry Semantics - -Rewrite pipeline execution around raw envelopes. - -Implementation tasks: - -- Introduce a small retry helper in `internal/framework/pipeline` that reruns - the same module with the same input when: - - the module returns a framework-level error; - - the module returns output that is rejected by that stage's validator chain. -- Preserve context cancellation: never retry after `ctx.Err()` is non-nil. -- Record attempt count and compact attempt diagnostics for manifests/output. -- Do not log or manifest raw payload bytes by default. -- Run chunk once per source input, with retries from the chunk binding. -- After successful chunk execution, run chunk validation when validator mapping - support exists. A rejected chunk result after final retry should stop - downstream execution and produce a rejected run record. -- Run extract once per accepted chunk, with per-chunk retries from the extract - binding. -- Omit rejected extract outputs from merge input. -- If a lane has no accepted extract outputs after retries, skip merge and - normalize for that lane, record rejected output details, and continue to the - next lane. -- Run merge once per lane with accepted extract outputs, with retries from the - merge binding. -- If merge output is rejected after final retry, omit that lane from normalize - and output. -- Run normalize once per accepted merge output, with retries from the normalize - binding. -- If normalize output is rejected after final retry, omit that lane from output. -- Treat validator rejection as non-fatal run outcome. Treat unrecovered - framework-level execution errors as run failures. -- Preserve deterministic ordering: - - chunks by chunk index; - - extract outputs by chunk index; - - lane outputs by resolved lane order. -- Continue to collect warnings from every successful module attempt whose output - is used. For failed attempts, record compact attempt diagnostics rather than - promoting warnings as final-stage warnings unless the implementation already - has a clear warning policy. -- Set manifest validation status to: - - `approved` when all produced normalized lane outputs pass; - - `rejected` when one or more module outputs are rejected; - - failed-run status through existing failure handling for unrecovered runtime - errors. - -Validation seam scope: - -- Add the runner-side raw validation seam needed by this data-model refactor. -- Represent the validation boundary as raw module output plus stage provenance, - consistent with [`validation.md`](validation.md). -- Support empty validator chains as approval. -- Add fake validator/test-only coverage for approval, rejection, and - retry-on-rejection behavior. -- Do not implement the full production validator package migration, default - validator mappings, or concrete generic validators in this plan. Those remain - owned by the validation roadmap. - -Tests to update or add: - -- Runner passes chunk content and media type to extractors. -- Runner passes ordered accepted extract outputs to merge. -- Runner passes accepted merge output to normalize. -- Rejected extract output is omitted from merge input. -- A lane with no accepted extracts is omitted and recorded. -- Rejected merge output prevents normalize for that lane. -- Rejected normalize output prevents output for that lane. -- Retry reruns the same module input after framework error. -- Retry reruns the same module input after validator rejection. -- Retry stops after configured attempts and records attempt count. -- Context cancellation stops retries. - -Completion criteria: - -- The runner no longer depends on candidate materialization between extract, - merge, normalize, and output. -- Rejected outputs never pass to the next stage. - -## Stage 5: Production Module Migration - -Update concrete modules to the new contracts. - -Implementation tasks: - -- Update `internal/modules/extract/dnd/spells`: - - call Scriptorium with the existing prompt/schema; - - capture the returned raw JSON response bytes; - - return one `ExtractOutput` with `application/json` media type and response - schema provenance; - - do not convert spell casts into `ArtifactCandidate`; - - do not run shape/source-ref/source-relatedness validation inside the module. -- Keep D&D spell prompt and schema asset metadata in manifest metadata. -- Update `internal/modules/merge/appendorder` to merge raw JSON extract outputs - deterministically: - - preserve input order by chunk index; - - if all extract outputs are JSON objects with one common top-level array - field, concatenate those arrays into one JSON object under that field; - - otherwise produce a JSON array containing each extract output's decoded JSON - value in order; - - reject non-JSON media types with a clear error unless a future option - deliberately supports them. -- Update `internal/modules/normalize/noop` to pass accepted merge output through - unchanged as normalized output. -- Update chunk modules from Stage 1 as needed after interface changes. -- Remove or quarantine old candidate validators from D&D module packages if - they no longer compile. Their replacement belongs under the validation - roadmap in `internal/validators`. -- Update fake/test modules throughout the repository to the raw contracts. - -Tests to update or add: - -- D&D spells extractor returns raw JSON with schema provenance and no candidate - materialization. -- Append-order merger concatenates common top-level JSON arrays. -- Append-order merger falls back to ordered JSON value arrays when shapes differ. -- Append-order merger rejects invalid JSON and non-JSON media types clearly. -- No-op normalizer returns defensive copies of merge output bytes and metadata. -- Production catalog still registers all modules successfully. - -Completion criteria: - -- Production modules compile against raw contracts. -- D&D spell extraction no longer rejects spell candidates inside the extractor. - -## Stage 6: Output, Manifest, Diagnostics, And Files - -Update durable output around normalized lane outputs. - -Implementation tasks: - -- Update `RunOutput` to carry normalized lane outputs and rejected module-output - records instead of approved/rejected artifacts. -- Update `artifacts.RunManifest` or introduce a more accurately named manifest - package if the old artifact naming becomes misleading. Prefer the smallest - rename that keeps manifest output clear and avoids a broad unrelated refactor. -- Preserve existing manifest fields that remain meaningful: - - pipeline ID and digest; - - module keys; - - lane IDs; - - source digest; - - LLM profile provenance; - - reference provenance; - - module metadata; - - validation status. -- Add manifest/reporting fields needed for raw outputs: - - normalized lane output media type; - - output schema provenance where present; - - rejected stage/lane/chunk/module information; - - retry attempt counts. -- Update `internal/modules/output/json`: - - write `manifest.json`; - - write `index.json`; - - write `warnings.json`; - - write `rejected.json`; - - write one file per normalized lane output. -- For the JSON output encoder, accept `application/json` normalized outputs and - write them as raw JSON files. Recommended file path: - `lanes/.json`. -- Reject non-JSON normalized output media types in the JSON encoder with a clear - error. Future text/Markdown/binary encoders can support other media types. -- Keep output file name validation strict. -- Ensure diagnostics redaction still prevents raw source, prompt, reference, - schema, and payload bytes from appearing in ordinary errors or manifests. - -Tests to update or add: - -- JSON output writes lane output files and index entries deterministically. -- JSON output rejects invalid JSON bytes and unsupported media types. -- Rejected output records are written without raw payload bytes. -- Manifest includes raw-output provenance and retry attempt counts. -- Diagnostics tests still prove sensitive/large content is not emitted. - -Completion criteria: - -- Durable output no longer assumes artifact candidates. -- JSON output remains deterministic and safe. - -## Stage 7: Documentation, Examples, And Full Validation - -After behavior is implemented, update canonical current-behavior docs. - -Implementation tasks: - -- Update `docs/internal/pipeline.md` for the raw data model. -- Update `docs/internal/modules.md` for new module contracts. -- Update `docs/internal/llm.md` for merge-stage LLM/profile plumbing. -- Update `docs/config.md` for: - - integer source-unit expectations where config examples include source refs; - - merge references; - - module `retries`; - - LLM profile override scope. -- Update `docs/cli.md` for merge reference selectors and profile behavior. -- Update `docs/integrations/json-output.md` for lane output files. -- Update `docs/integrations/dnd-spell-artifacts.md`; if the artifact-specific - contract is no longer durable, rename or replace it with a D&D spell raw-output - integration doc. -- Update `docs/troubleshooting.md` for common raw-output, media-type, retry, and - validation failures. -- Update examples so Seriatim segment IDs are positive integers and expected - output shape matches lane-normalized output. -- Replace `docs/roadmap/implementation.md` with a completed-note document only - after this implementation is finished and reviewed. -- Leave `docs/roadmap/pipeline.md` as target-state roadmap context until the - feature is implemented, then move implemented behavior into canonical docs and - remove stale roadmap language. - -Validation commands: - -```sh -go test ./internal/core/source -go test ./internal/framework/contracts -go test ./internal/framework/pipeline -go test ./internal/core/config -go test ./internal/cli -go test ./internal/modules/input/seriatim -go test ./internal/modules/chunk/generic -go test ./internal/modules/chunk/dnd/scenes -go test ./internal/modules/extract/dnd/spells -go test ./internal/modules/merge/appendorder -go test ./internal/modules/normalize/noop -go test ./internal/modules/output/json -go test ./... -go vet ./... -go build ./cmd/notarius -``` - -Documentation inspection: - -```sh -rg -n "ArtifactCandidate|approved artifacts|spell artifact|start_unit_id\": \"|end_unit_id\": \"" docs examples internal -rg -n "chunk`, `extract`, and `normalize|deterministic-only stage|json.RawMessage" docs/roadmap docs/internal docs/config.md docs/cli.md -``` - -Expected result: - -- Remaining `ArtifactCandidate` mentions are only in deliberately retained - compatibility code or historical roadmap context. -- Current-behavior docs describe raw lane outputs, integer source-unit IDs, - merge LLM/reference support, and retry behavior. +Follow-up validator work remains tracked separately in +[Validation Roadmap](validation.md). diff --git a/docs/roadmap/pipeline.md b/docs/roadmap/pipeline.md index 2b8aac2..5e21b3d 100644 --- a/docs/roadmap/pipeline.md +++ b/docs/roadmap/pipeline.md @@ -1,403 +1,12 @@ # Raw Pipeline Data Model -This roadmap defines the target data model for the pipeline stages after input. -It is paired with the validation system roadmap in -[`validation.md`](validation.md): validators should evaluate module outputs, but -the pipeline first needs a raw-output handoff model that does not require every -module response shape to be represented by bespoke Go structs. +The raw pipeline data model has been implemented. -## Goals +Current behavior is documented in: -- Make typed chunk envelopes and raw module-output envelopes first-class - pipeline handoffs. -- Keep framework-owned provenance around raw payloads so runs remain - deterministic, ordered, and auditable. -- Avoid requiring each extractor, merger, or normalizer response schema to have - matching Go structs. -- Let extract modules produce one raw output per input chunk. -- Let merge modules receive accepted extract outputs and merge them into one raw - payload. -- Let normalize modules optionally post-process accepted merge output. -- Let output modules write normalized output bytes directly, along with - manifests, warnings, rejected outputs, and diagnostics as appropriate. -- Preserve the fixed workflow shape: +- [Pipeline Internals](../internal/pipeline.md) +- [Modules](../internal/modules.md) +- [JSON Output](../integrations/json-output.md) -```text -input -> chunk -> extract -> merge -> normalize -> output -``` - -## Non-Goals - -- Do not turn the pipeline into an arbitrary DAG or general workflow language. -- Do not make validators mutate, materialize, or rewrite module outputs. -- Do not require a generic source-reference convention for every raw payload in - this pass. -- Do not eliminate typed framework data where the framework genuinely needs it, - such as input source documents and chunk boundaries. - -## Cross-Stage Runtime Plumbing - -LLM/profile/reference plumbing should be available to every stage that may need -LLM-backed or reference-aware module behavior: `chunk`, `extract`, `merge`, and -`normalize`. - -For all four stages, the framework should provide the same categories of runtime -support where the concrete module contract needs them: - -- configured LLM profile and profile override handling; -- Scriptorium client access through framework-owned LLM contracts; -- session ID propagation; -- declared reference slots and resolved reference bindings; -- module options and metadata; -- prompt, schema, profile, and reference provenance for manifests and - diagnostics. - -`merge` must not be treated as a deterministic-only stage. A merge module may be -simple and deterministic, but it may also be LLM-backed, reference-aware, and -validator-gated in the same way as `chunk`, `extract`, and `normalize`. - -## Chunk Stage Target Model - -The chunk stage receives the canonical `SourceDocument` from the input stage, -plus any original source input material needed for prompt construction. The -`SourceDocument` remains the source of truth for source identity, source-unit -ordering, source-unit IDs, and source provenance. Original raw input material -may be supplied to LLM-backed chunkers, but it should not replace the -`SourceDocument` as the framework handoff. - -Input adapters are responsible for assigning stable integer source-unit IDs. How -those IDs are assigned depends on the input format. A numbered JSON transcript -may map source units directly to transcript segment numbers; a PDF input adapter -may assign page or extracted-text unit numbers; an adapter for unordered source -material may assign deterministic IDs as part of input parsing. Downstream -framework code should treat those IDs as opaque integers owned by the input -adapter. - -The chunk module returns one or more ordered `SourceChunk` envelopes. A single -chunk representing the whole source is valid and should be supported. - -Each chunk should include framework-readable provenance and ordering metadata, -plus content suitable for extraction. The content may be JSON, plain text, -Markdown, PDF page text, or another module/input-specific representation, as -long as the framework can still associate the chunk with the source and preserve -deterministic order. - -Conceptually: - -```go -type SourceChunk struct { - ID string - SourceID string - Index int - - StartUnitID int - EndUnitID int - - Content []byte - MediaType string - - Units []source.SourceUnit - Metadata map[string]any -} -``` - -The exact implementation shape can differ, but it should preserve: - -- chunk identity; -- source identity; -- deterministic chunk order; -- source locator or range, normally integer start/end unit IDs; -- chunk content and media type; -- optional source-unit projection when useful; -- metadata needed by extractors, validators, manifests, diagnostics, and output - encoders. - -The framework owns the minimal invariants required to schedule extract work: - -- the chunker returns at least one chunk; -- chunk IDs and indexes are stable and non-empty; -- chunk source IDs match the source document; -- chunk ordering is deterministic; -- chunk provenance is sufficient to trace the chunk back to the input source. - -Domain-specific chunk acceptability belongs in validators. For example, full -coverage, no gaps, no overlap, scene metadata quality, expected media type, and -D&D scene-boundary policy should be explicit validator concerns rather than -hidden framework rules, except where a minimal invariant is required for -extraction to run safely. - -## Extract Stage Target Model - -The extract stage receives one input chunk from the chunk stage and produces one -raw extract output for that chunk. - -The extractor owns: - -- prompt selection; -- input material assembly; -- Scriptorium request construction; -- response-schema selection; -- LLM profile and session usage; -- returned raw payload metadata. - -The framework owns: - -- chunk iteration; -- deterministic ordering; -- association between each extract output and its input chunk; -- execution errors when an extractor does not return output; -- handoff of returned output to validation and later stages. - -An extract output should be an envelope, not just raw bytes. Conceptually: - -```go -type ExtractOutput struct { - LaneID string - ExtractorKey string - - ChunkID string - ChunkIndex int - SourceID string - - RawContent []byte - MediaType string - - SchemaID string - SchemaName string - SchemaVersion string - - Metadata map[string]any - Warnings []contracts.Warning -} -``` - -The exact implementation shape can differ, but it should preserve: - -- raw returned content; -- chunk provenance; -- chunk order; -- source identity; -- module identity; -- response schema provenance; -- metadata needed by validators, mergers, manifests, diagnostics, and output - encoders. - -Extractor modules should not be required to convert raw LLM output into -`ArtifactCandidate` values. Domain-specific Go projection may still exist for -specific modules when it is useful, but it should not be the generic pipeline -contract. - -Rejected extract outputs are not passed to merge. This keeps downstream -contracts simple and makes rejection behavior explicit. Support for passing -rejected outputs forward as marked data is deferred future work. - -## Merge Stage Target Model - -The merge stage receives the accepted extract outputs for a lane. Each extract -output corresponds to one input chunk and carries enough metadata to recover the -original chunk order. - -The merger owns: - -- merge/reconciliation strategy; -- deterministic or LLM-backed merge logic; -- prompt and schema usage when LLM-backed; -- the shape of merged raw output. - -The framework owns: - -- passing the ordered accepted extract-output set to the merger; -- lane identity; -- source context; -- references; -- runtime LLM plumbing; -- validation handoff for merged output; -- manifest provenance. - -Simple merge modules may concatenate raw extract outputs in chunk order. Other -merge modules may deterministically merge JSON documents, reconcile duplicates, -or use an LLM to produce a more coherent merged result. - -Conceptually: - -```go -type MergeRequest struct { - LaneID string - Source *source.SourceDocument - ExtractOutputs []ExtractOutput - - SourceInput contracts.LLMInputMaterial - SessionID string - References contracts.ReferenceSet - LLMClient contracts.StructuredLLMClient - LLMProfile string - Options map[string]any - Metadata map[string]any -} - -type MergeOutput struct { - LaneID string - MergerKey string - - RawContent []byte - MediaType string - - SchemaID string - SchemaName string - SchemaVersion string - - Metadata map[string]any - Warnings []contracts.Warning -} -``` - -The exact implementation shape can differ, but the key contract is that -merge consumes ordered accepted extract outputs and produces one merged raw -output for the lane. - -If no accepted extract outputs remain for a lane, the framework should not pass -rejected outputs to the merger. The lane should produce no merge output and -should be reported as rejected or omitted according to run reporting policy. - -## Normalize Stage Target Model - -The normalize stage receives accepted merge output for a lane and optionally -post-processes it into the final normalized raw payload for that lane. -Normalization may be a no-op, deterministic cleanup, schema conversion, or an -LLM-backed post-processing pass. - -The normalizer owns: - -- post-merge processing strategy; -- deterministic or LLM-backed normalization logic; -- prompt and schema usage when LLM-backed; -- the shape of normalized raw output. - -The framework owns: - -- passing accepted merge output to the normalizer; -- lane identity; -- source context; -- references; -- runtime LLM plumbing; -- validation handoff for normalized output; -- manifest provenance. - -Conceptually: - -```go -type NormalizeRequest struct { - LaneID string - Source *source.SourceDocument - MergeOutput MergeOutput - - SourceInput contracts.LLMInputMaterial - SessionID string - References contracts.ReferenceSet - LLMClient contracts.StructuredLLMClient - LLMProfile string - Options map[string]any - Metadata map[string]any -} - -type NormalizeOutput struct { - LaneID string - NormalizerKey string - - RawContent []byte - MediaType string - - SchemaID string - SchemaName string - SchemaVersion string - - Metadata map[string]any - Warnings []contracts.Warning -} -``` - -The exact implementation shape can differ, but the key contract is that -normalize consumes one accepted merged output and produces one normalized raw -output for the lane. - -## Output Stage Target Model - -The output stage receives validated normalized output bytes and writes them to -its configured destination. - -Output modules should not require normalized output to be converted into -`Artifact` values. Output modules should use the normalized output media type to -decide how to serialize or wrap the payload. A JSON output module can write raw -normalized JSON directly, while a text or Markdown output module can write text -payloads directly. Output modules may also write run manifests, warnings, -rejected outputs, and indexes. - -Output modules may still choose to provide convenience layouts, grouping, or file -naming conventions, but those should be output concerns rather than constraints -on extractor or normalizer response schemas. - -## Validation Relationship - -Validation chains should attach to returned module outputs, not to hidden -module-internal conversions. - -For chunk: - -```text -chunk(input) -> SourceChunk set -> chunk validators -> extract input -``` - -For extract: - -```text -extract(chunk) -> raw ExtractOutput -> extract validators -> merge input -``` - -For merge: - -```text -merge(extract outputs) -> raw MergeOutput -> merge validators -> normalize input -``` - -For normalize: - -```text -normalize(merge output) -> raw NormalizeOutput -> normalize validators -> output -``` - -Validators must be read-only. They inspect raw output and metadata, return -accept/reject decisions and warnings, and do not rewrite output. - -An empty validator chain approves returned output for that validation point. -Framework/runtime errors remain separate from validation rejections: if a module -or Scriptorium call fails before output is returned, the pipeline reports an -execution error rather than asking validators to evaluate nonexistent output. - -Rejected outputs do not pass to the next stage. An empty validator chain still -approves returned output for that validation point. - -## Retry Policy - -A retry means re-running the same module with the same input after that module -fails to produce valid output. Failure to produce valid output includes both: - -- framework-level execution errors, such as module errors, Scriptorium errors, - provider errors, or missing returned output; -- validator rejection of returned output. - -Retries should be configurable at the pipeline or lane level. Chunk retries are -per source input. Extract retries are per chunk. Merge and normalize retries, -when configured, are per lane. Retry attempts should preserve deterministic -reporting: the final accepted or rejected output should record attempt count and -enough diagnostics/provenance to understand prior failures without leaking -secrets or large payloads by default. - -## Relationship To Validation Roadmap - -This roadmap defines the data model that validation should evaluate. The -validation system roadmap in [`validation.md`](validation.md) defines validator -registration, mapping, execution classes, and concrete validator behavior. - -The shared boundary between the roadmaps is a returned module output: validators -inspect typed chunk output or raw module-output envelopes and decide whether -that output may continue through the pipeline. +Validator mapping and concrete validator behavior remain tracked separately in +[Validation Roadmap](validation.md). diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index ded07ae..6198554 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -200,7 +200,8 @@ Fix: - Or use an existing Scriptorium profile ID with `--llm-profile`. Use `--llm-profile ` when one run should force every LLM-backed binding to -the same Scriptorium profile. +the same Scriptorium profile. The override applies to effective chunk, extract, +merge, and normalize bindings. ## Missing API Key Environment Variable @@ -285,17 +286,58 @@ Symptoms include: - `create output directory` - `write output file` - `output file name must` +- `unsupported media type` +- `invalid JSON` Fix: - Ensure `--output-dir` points to a directory path or a path that can be created. - Check filesystem permissions and available disk space. +- The production JSON output encoder writes lane payloads under `lanes/` and + accepts only valid `application/json` normalized outputs. If an error names an + unsupported media type or invalid JSON, inspect the lane's merge and normalize + module output. - If diagnostics were retained, inspect `run-report.json`, `run-manifest.json`, and `error.log`. The CLI rejects unsafe logical output paths before writing files. +## Raw Output Rejection + +Symptoms include a successful run with: + +- `validation_status` set to `rejected`; +- non-empty `rejected.json`; +- `rejected_outputs` entries in `manifest.json`. + +Explanation and fixes: + +- Validator rejection is a non-fatal run outcome. Rejected module outputs do not + pass to the next pipeline stage. +- Check `rejected.json` for the stage, lane, module, chunk, validator, reason, + message, and attempt count. +- Increase a module binding's `retries` only when re-running the same module + input can reasonably produce an acceptable output. +- If rejection is deterministic, fix the source input, module configuration, or + validator configuration rather than adding retries. + +## Retry Exhaustion + +Symptoms include: + +- errors containing `failed after ... attempt(s)`; +- rejected output records with `attempt_count` greater than `1`. + +Fix: + +- `retries` is the number of extra attempts after the first attempt for chunk, + extract, merge, and normalize bindings. +- Framework-level errors after the last attempt fail the run. +- Validator rejections after the last attempt are recorded as rejected outputs. +- Check retained `error.log`, `run-manifest.json`, and `rejected.json` for the + operation, module key, lane, chunk, and attempt count. + ## Diagnostics Directory Surprise Symptom: the diagnostics directory is missing after a successful run.