# Validation System Refactor This roadmap defines the target state for making validation a first-class, composable pipeline concern. Current pipeline behavior is raw-output based. The runner can execute `contracts.RawValidator` chains from `pipeline.RawValidationRegistry` for `chunk`, `extract`, `merge`, and `normalize` outputs. Empty chains approve by default, validator rejection records a rejected raw output, and rejected output does not pass to the next stage. Production currently registers no raw validators, and non-empty pipeline-configured validator lists are rejected so they cannot appear in manifests without executing. Legacy candidate validator contracts and D&D spell validators still exist under `internal/modules/extract/dnd/spells`, but they are not part of the current runner path. The desired end state is that validator implementations, validator registration, and default module-to-validator mappings are explicit, reviewable, and independent of concrete module packages. ## Goals - Move artifact and module-output validation behavior out of `internal/modules` and into `internal/validators`. - Retire or replace the legacy candidate-oriented `contracts.Validator`, `ValidationRequest`, and `ValidationResult` path after equivalent raw-output validators exist. - Keep each validator in its own package. - Mirror the stage and domain shape of `internal/modules` where a validator is module-specific. - Support deterministic and LLM-backed validators through the same framework contract. - Allow validators to be mapped to modules at any pipeline stage that returns module output for validation: `chunk`, `extract`, `merge`, or `normalize`. - Make default production module-to-validator mappings centralized and human-readable. - Allow pipeline configuration to override default mappings for advanced use. - Preserve the distinction between an unset validator override and an explicit empty validator override. - Treat an empty validator set as valid and equivalent to approval. - Preserve the rule that module output passes forward unless a validator rejects it. - Make successfully returned module output the explicit validation boundary: questions about output syntax, media type, schema conformance, and domain acceptability should be answered by validators. ## Non-Goals - Do not create a general workflow engine or arbitrary validation DAG. - Do not revive the legacy artifact-candidate validation model as the primary runner path. - Do not enforce validator compatibility with a module or stage in this pass. - Do not move ordinary runtime invariant checks into validator packages. - Do not require every module to have validators. - Do not require LLM-backed validators for modules that can be checked deterministically. - Do not silently reorder configured validator chains unless that behavior is introduced deliberately and documented as part of the validator contract. ## Validation Boundary Validation packages should own approve/reject/warning evaluation of successfully returned module outputs. This means logic that decides whether a chunk result or raw extract, merge, or normalize payload should continue through the pipeline belongs in `internal/validators`. The boundary is: - no module output was returned: execution failed, and the pipeline should report a module or runtime error; - module output was returned: the validator chain decides whether that output is acceptable, and an empty validator chain approves it. Scriptorium and provider errors are execution failures rather than validator rejections. This includes provider timeouts, authentication failures, transport/runtime failures, Scriptorium structured-output retry exhaustion, and malformed responses that Scriptorium rejects before returning module output. Other validation-like checks should remain with their owning packages: - input parsing and source-format validation stay in input modules; - source document and source reference invariants stay in `internal/core/source`; - generic framework chunk invariants that make extraction possible stay in the runner, such as non-empty chunk content, valid unit ranges, and canonical source-unit ordering; - config validation stays in `internal/core/config`; - registry, profile, and pipeline consistency checks stay in framework and CLI code; - response schema loading stays in module asset code; - Scriptorium runtime errors stay in LLM runtime code. Domain validators may call reusable core helpers such as `source.ValidateRef`, but the module-output approval or rejection decision should be made by a validator. Validators should answer module-output questions such as: - is returned content syntactically valid JSON; - does returned JSON conform to the module's declared schema; - does the returned media type match the module or pipeline policy; - are required domain fields present and non-empty; - are source references valid and appropriately grounded; - does domain-specific output satisfy the configured policy. ## Audita Patterns To Adapt The validator architecture should adapt useful patterns from [`audita`](https://gitea.maximumdirect.net/eric/audita) without copying its narrower transcript-correction shape directly. Useful patterns: - concrete validators live under `internal/validators`; - shared validator runtime mechanics live under a framework package; - built-in validator keys are stable and centrally registered; - built-in chains are centrally reviewable; - validators carry execution-class metadata; - deterministic and LLM-backed validators implement one contract; - LLM-backed validator runtime can share batching, diagnostics, structured response handling, and malformed-response policy; - reports and manifests can classify validator decisions by execution class. Important Notarius differences: - mappings must be keyed by stage and module key, not module key alone; - mappings should be owned by the central production catalog, not resolved inside concrete module constructors; - configured mapping order should be authoritative unless the config explicitly opts into a different ordering policy; - validators must support chunk, extract, merge, and normalize outputs rather than only one proposal shape. ## Validator Package Layout Concrete validators should live under `internal/validators`. Module-specific validators should mirror the module tree and use one package per validator: ```text internal/validators/extract/dnd/spells/shape internal/validators/extract/dnd/spells/source_refs internal/validators/extract/dnd/spells/source_relatedness ``` D&D spell validation policy belongs under `internal/validators/extract/dnd/spells`, split by concern rather than bundled inside the extractor module. Generic validators may live under stage-specific generic paths when they operate on a particular stage output shape: ```text internal/validators/chunk/generic/... internal/validators/extract/generic/... internal/validators/merge/generic/... internal/validators/normalize/generic/... ``` Truly stage-independent validators may live under `internal/validators/generic` once there is a real shared validator that justifies that location. Generic JSON syntax and JSON schema validators are likely candidates for `internal/validators/generic/valid_json` and `internal/validators/generic/valid_json_schema`. Each validator package should expose: - a stable validator key; - execution-class metadata; - a constructor; - a validator spec suitable for registration; - a `Register` function; - focused tests for decisions, warnings, errors, and diagnostics behavior. Reusable validator runtime mechanics should live in framework code, such as `internal/framework/validators`, not in concrete validator packages. This package can own shared helpers for decision cardinality, approval/rejection construction, LLM validator batching, validator diagnostics, and Scriptorium request plumbing. The concrete validator packages should own policy: what they inspect, what they approve or reject, what warning reason codes they emit, and how they interpret domain-specific data. ## Validator Design Policy Validators should follow a small-tool model: each validator should do one thing well. If a validator both rejects output and emits unrelated warnings, split those concerns into separate validators so production mappings can include, exclude, and order them independently. Validators are read-only. A validator must not mutate pipeline state, rewrite module output, materialize raw output into typed stage output, or enrich the `ModuleOutput` passed to later validators. A validator returns an accept/reject verdict for the output it evaluates, plus any warnings or diagnostic references. Any conversion from raw module output into a downstream representation is a separate materialization concern and must not be hidden inside a validator. The initial generic validator set should include: - `generic/always_accept`: accepts returned module output unchanged. This is functionally equivalent to a no-op validator and is primarily useful for tests, demonstrations, and explicit pass-through configurations. - `generic/always_reject`: rejects returned module output without inspecting it. This is primarily useful for tests and for proving rejection plumbing, manifests, and diagnostics. - `generic/valid_json`: inspects raw returned module output and accepts only syntactically valid JSON. - `generic/valid_json_schema`: compares raw returned JSON with the module's configured response schema and accepts only schema-conformant output. The exact keys may be adjusted during implementation to match local naming conventions, but the validator set should preserve these four behaviors. For the current D&D spell behavior, the target split is: - `generic/valid_json`: rejects returned module output that is not syntactically valid JSON. - `generic/valid_json_schema`: rejects returned JSON that does not conform to the configured response schema. - `extract/dnd/spells/shape`: rejects malformed spell-cast payloads and missing required spell fields. - `extract/dnd/spells/source_refs`: rejects missing or invalid source references. - `extract/dnd/spells/source_relatedness`: warning-only validator that reports when a spell name is not found in the cited source text. Production default mappings should generally list deterministic validators before LLM-backed validators. This keeps cheap structural failures from consuming model calls and keeps diagnostics easier to interpret. Pipeline-configured order should still be authoritative; if a user explicitly lists an LLM-backed validator before a deterministic validator, the framework should honor that order rather than silently reshuffling it. ## Execution Classes Validator specs should declare an execution class: ```go type ExecutionClass string const ( ExecutionClassDeterministic ExecutionClass = "deterministic" ExecutionClassLLMBacked ExecutionClass = "llm_backed" ) ``` Execution class should be metadata on the validator spec or registered definition, not an ad hoc convention inferred from package paths. It should be used for: - human-readable catalog and manifest reporting; - diagnostics and timing summaries; - operational policy such as concurrency budgeting for LLM-backed validators; - default mapping review, where deterministic validators should usually appear before LLM-backed validators. Execution class should not by itself imply compatibility with a stage or module. ## Validator Contract The validator framework should support validation of outputs from `chunk`, `extract`, `merge`, and `normalize` stages. The contract should be generalized enough for stage-specific validators to inspect the output they care about while ignoring irrelevant fields. The current `contracts.RawValidationRequest` is the right starting point. It already carries stage, lane, module, source, source and chunk provenance, response schema metadata, raw payload, and run metadata. The final contract should evolve from that raw-output shape rather than from the legacy artifact-candidate `ValidationRequest`. Additional fields needed for the full validator system include: - source input material when a validator needs to compare module output to the original source payload; - session ID; - resolved references for the validated target; - LLM client and profile for LLM-backed validators; - validator options; - chunk output collections when validating a chunk module; - ordered upstream output envelopes when validating merge or normalize behavior. A shared module-output envelope should represent the validation boundary. Validators may inspect raw returned content and any already-existing typed stage envelope, such as `SourceChunk` values for chunk validation, but they must not modify it. Conceptually: ```go type ModuleOutput struct { Stage pipeline.ModuleStage ModuleKey string LaneID string RawContent []byte MediaType string ResponseSchema contracts.ResponseSchema SourceID string ChunkID string ChunkIndex int Chunks []contracts.SourceChunk Warnings []contracts.Warning } ``` The final implementation does not need to use this exact shape, but it should preserve the boundary: returned raw module output enters validation as immutable module output. Validators may parse raw bytes internally to decide approve, reject, or warn, but parsing inside a validator must not create or replace the payload passed to later stages. The result should continue to express validator identity, warnings, and explicit decisions. For output collections, the implementation should define an explicit decision shape rather than silently mutating lists. Validator decisions reject or approve output; validators do not rewrite output. An empty validator list is always valid. With no validators, the framework should pass module output forward unchanged and treat the output as approved for that validation point. If the approved output cannot be consumed by a later stage because its media type or envelope shape is unsuitable, that failure should be reported at the downstream boundary that requires a different shape, not as an implicit pre-validation rejection. ## Module Development Workflow The validation system should make iterative module development easier. Once pipeline overrides are implemented, a module author should be able to start with an explicit empty validator mapping and inspect returned raw LLM output without first satisfying JSON syntax, schema, media-type, or domain validators. A typical development path should be: 1. Configure an empty validator set for the module and inspect raw returned output. 2. Add `generic/valid_json` and adjust prompts until the model reliably returns syntactically valid JSON. 3. Add `generic/valid_json_schema` and iterate on prompt/schema alignment. 4. Add media-type or schema validators appropriate to the module's intended output format. 5. Add domain-specific validators one at a time until production policy is represented explicitly in the chain. This workflow is a central reason for making output validation explicit and composable rather than hiding schema, shape, or domain checks inside module implementation code. ## Central Production Mappings Production defaults should be defined in a central, human-readable location near the production module and validator registries. The mapping should be keyed by stage and module key, not only by module key, so future modules can share keys only when stage context makes their ownership unambiguous. Conceptually: ```go { Stage: pipeline.StageExtract, Module: "dnd/spells", Validators: []pipeline.ModuleBinding{ pipeline.Binding("generic/valid_json"), pipeline.Binding("generic/valid_json_schema"), pipeline.Binding("extract/dnd/spells/shape"), pipeline.Binding("extract/dnd/spells/source_refs"), pipeline.Binding("extract/dnd/spells/source_relatedness"), }, } ``` The production catalog should expose three related surfaces together: - available modules; - available validators; - default module-to-validator mappings. This makes production validation policy reviewable without constructing concrete modules or searching inside module implementation packages. The validator registry should expose registered validator specs without building validators, including key and execution class. Building a validator should still be available for runtime execution. The mapping surface should preserve the current useful behavior of stage/module lookup and empty-chain approval while adding validator specs, execution-class metadata, production registration, config override integration, and manifest reporting of the resolved chain. ## Pipeline Overrides Pipeline configuration should be able to override the central default mapping for a module binding. Current configuration validation rejects non-empty validator lists, and the current config/profile structs do not preserve whether an empty list was explicitly configured or simply omitted. The target config model must preserve that distinction. Override semantics should distinguish three states: - unset validators: use the central production/default mapping; - explicit empty validators: run no validators and pass output forward; - explicit non-empty validators: run exactly the configured validators in the configured order. This keeps the happy path concise while preserving advanced control for experimentation, debugging, and custom deployments. The run manifest should record the resolved validator chain for each validation point so completed runs remain auditable after defaults or configuration change. Each manifest entry should include at least validator key and execution class, and should preserve the resolved order actually used for the run. ## LLM-Backed Validator Runtime LLM-backed validators should use the same validator contract as deterministic validators. Shared framework runtime should provide common support for: - Scriptorium request construction; - validator prompt and schema provenance; - diagnostics redaction; - optional batching or context-window controls when validator inputs are large; - mapping successful LLM validator responses into validator decisions and warnings; - consistent handling of Scriptorium/runtime errors. Scriptorium errors raised during validator execution should be treated as validator execution errors unless a specific validator deliberately converts a successful response into reject/warn decisions. This keeps provider/runtime failure distinct from a validator's semantic rejection of module output. ## Stage Coverage Validators should be composable across all LLM-eligible stages: - `chunk`: validators can evaluate chunk boundaries, coverage, overlap, metadata, or module-specific chunk quality. - `extract`: validators can evaluate raw extracted output, source references, payload shape, evidence quality, media type, or domain constraints. - `merge`: validators can evaluate merged output, cross-chunk consistency, deduplication results, media type, or domain-specific reconciliation. - `normalize`: validators can evaluate normalized output, final shape, post-processing results, media type, or domain-specific policy. The framework should not require compatibility declarations in this pass. A validator mapped to an unsuitable output shape should return a clear error, or approve unchanged only when that is explicitly the validator's documented behavior. ## Documentation Impact Current-behavior docs and policy should describe the implemented validation system once the refactor is complete: - `docs/policy/architecture.md` should describe centralized validator mappings rather than module-owned validator chains. - `docs/internal/modules.md` should remove claims that concrete modules own validator defaults. - Internal validation docs should describe validator package ownership, mapping precedence, empty-chain approval behavior, and LLM-backed validator support. - User/config docs should replace the current "configured validators are reserved and rejected" language with the implemented pipeline override contract. Roadmap docs should not remain the canonical description of implemented validation behavior after the refactor is complete.