diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md new file mode 100644 index 0000000..3b4cbe3 --- /dev/null +++ b/docs/roadmap/implementation.md @@ -0,0 +1,305 @@ +# Extraction References Implementation Plan + +## Purpose + +Implement the extraction-reference feature described in +[references.md](references.md). This plan is decision-complete for an LLM coding +agent: implement each stage in order, keep the repository compiling after each +stage, and do not move planned behavior into non-roadmap docs until the relevant +behavior exists. + +Core decisions to preserve: + +- references are opaque framework inputs and domain semantics stay in extract + modules; +- references are not evidence and must not be addressable through `SourceRef`; +- reference binding is lane-scoped; +- extractors expose `ReferenceSlots()` directly on the first-class extractor + contract; +- token budgeting is deferred; enforce only UTF-8 text handling, empty-file + warnings, and declared `MaxBytes`; +- run manifests record references in a dedicated section, separate from + `source_digests`; +- CLI unbinding uses `--without-reference`. + +## Stage 1: Contracts and Mechanical Adoption + +Add the framework contracts needed to describe references without changing +runtime behavior. Slot declarations must be available without constructing +extractor modules, because pipeline/config validation should use registry +metadata rather than runtime module instances. + +Implementation steps: + +- In `internal/framework/contracts`, add reference model types: + `ReferenceSlot`, `ReferenceOrigin`, `ReferenceItem`, + `ResolvedReferenceSlot`, `ReferenceSet`, and a binding-source enum or string + constants for `config` and `cli`. +- Include `Name`, `Description`, `Required`, `AcceptedMediaTypes`, `Multiple`, + and `MaxBytes` on `ReferenceSlot`. +- Include slot name, media type, content bytes, digest, origin, size bytes, and + binding source on `ReferenceItem`. +- Add `References ReferenceSet` to `contracts.ExtractionRequest`. +- Add `ReferenceSlots() []ReferenceSlot` to `contracts.Extractor`. +- Extend extractor registration metadata so reference slots are also declared + through the extractor's registry spec. Prefer the smallest idiomatic change to + the existing registry model, such as adding `ReferenceSlots` to `ModuleSpec` + with validation that non-extractor modules leave it empty, unless the codebase + shape clearly supports a narrower extractor-specific spec. +- Update every concrete extractor and all extractor fakes/test doubles to + implement `ReferenceSlots()`. Existing extractors without references should + return `nil`. +- Add tests that compare a production extractor's runtime `ReferenceSlots()` + with its registered spec slots so the two declarations cannot drift. +- Add contract tests for empty reference sets, slot copying expectations if + helpers are introduced, and compile-time coverage through existing fakes. + +Verification: + +- `go test ./internal/framework/contracts` +- `go test ./internal/framework/pipeline` +- `go test ./...` + +## Stage 2: Config Shape and Pipeline-Level Resolution + +Add unresolved reference bindings to config and resolved lane bindings to the +pipeline model. Do not read reference files in this stage. + +Implementation steps: + +- Add `references` maps to file config parsing at both pipeline and artifact + lane level. +- Add corresponding fields to `pipeline.PipelineProfile` and + `pipeline.ArtifactLaneProfile`. +- Preserve deterministic map handling and duplicate-after-trim validation. +- Extend config cloning, effective config, redaction, validation, and tests for + the new fields. +- Add resolved reference binding structures to `internal/framework/pipeline`. + They should represent lane ID, slot name, source URI/path, and binding source, + but not file bytes. +- During `pipeline.ResolvePipeline`, collect selected lanes, read each lane + extractor's declared slots from registry metadata, and validate without + building extractor instances: + - every bound slot is declared by the lane extractor; + - required slots are bound after applying pipeline-level and lane-level config; + - required slots remain bound after any CLI unbinds supplied to resolution; + - selected lanes under `--only` are the only lanes considered. +- Apply pipeline-level bindings as defaults only to lanes whose extractor + declares the matching slot. +- Apply lane-level bindings as overrides or additions for that lane. +- Keep reference bindings out of source digests and artifact source references. +- Add tests proving reference-slot validation works through registry specs even + when extractor constructors would fail if called. + +Verification: + +- `go test ./internal/core/config` +- `go test ./internal/framework/pipeline` +- `go test ./...` + +## Stage 3: CLI Reference Overrides and Unbinds + +Add run-time CLI syntax for reference binding overrides and optional unbinding. + +Implementation steps: + +- Add repeatable `--reference` flags to `notarius run`. + Accepted forms: + - `slot=path` for unambiguous slot names across selected lanes; + - `lane.slot=path` for explicit lane-scoped binding. +- Add repeatable `--without-reference` flags to `notarius run`. + Accepted forms: + - `slot`; + - `lane.slot`. +- Reject empty paths for `--reference`; use `--without-reference` for unbinding. +- Reject malformed values with concise CLI errors before expensive work. +- Pass parsed override/unbind requests into config/pipeline resolution. +- Resolve flat CLI names only when exactly one selected lane declares the slot. + If multiple selected lanes declare the same slot, fail and instruct the user + to use `lane.slot`. +- Let CLI bindings override config bindings for the same lane and slot. +- Let CLI unbinds remove config-bound optional slots for the same lane and slot. +- Fail if unbinding leaves a required slot unbound. +- Add CLI tests for flat binding, lane-qualified binding, ambiguous flat + binding, malformed syntax, optional unbind, and required-slot unbind failure. + +Verification: + +- `go test ./internal/cli` +- `go test ./internal/core/config` +- `go test ./internal/framework/pipeline` +- `go test ./...` + +## Stage 4: Run Preparation and Reference Materialization + +Read, validate, digest, and materialize resolved file references before any LLM +call. + +Implementation steps: + +- Add a reference resolver/materializer near pipeline run preparation. Keep file + I/O out of pure config parsing. +- Ensure run preparation receives the loaded config path or config directory so + config-relative reference paths can be resolved after pure config parsing. +- Resolve config-relative paths relative to the config file path and + CLI-relative paths relative to the current working directory. +- For MVP, accept only UTF-8 text files. Reject non-UTF-8 content with an error + naming pipeline, lane, slot, and path. +- Compute `sha256:` content digests over the raw reference bytes. +- Populate `ReferenceItem` values with content bytes, media type, digest, + origin type `file`, normalized origin URI/path, size bytes, and binding source. +- Enforce declared `MaxBytes` when greater than zero. The error should name the + pipeline, lane, slot, actual size, limit, and path. +- Emit a warning for empty bound files, but do not fail. +- Add `ReferenceSet` values to the runner input or resolved pipeline path in a + way that keeps lane-scoped references available when calling each extractor. +- Pass the correct lane-specific `ReferenceSet` into + `contracts.ExtractionRequest`. +- Ensure no reference content is written to ordinary diagnostics, logs, errors, + or manifests. + +Verification: + +- Focused resolver/materializer tests for path resolution, digest stability, + UTF-8 rejection, empty-file warning, `MaxBytes`, and binding source. +- `go test ./internal/cli` +- `go test ./internal/framework/pipeline` +- `go test ./...` + +## Stage 5: Prompt Template Reference Functions + +Make references available to module-owned prompt templates. + +Implementation steps: + +- Extend `internal/framework/prompt` so prompt bundles can be compiled with + declared reference slots. +- Add `reference` and `hasreference` template functions. +- Validate at bundle build time, or the earliest feasible equivalent, that + templates reference only declared slots. +- Render a declared but unbound optional slot as an empty string. +- Ensure `hasreference` returns true only when the slot has at least one bound + item with content. +- Render multiple items deterministically if future `Multiple` support is + enabled; for MVP, reject multiple bindings unless the slot declares + `Multiple`. +- Keep prompt metadata hashes based on template source. Do not include rendered + reference content in prompt identity. +- Add deterministic rendering tests proving byte-identical output across runs + with the same reference bytes and config. + +Verification: + +- `go test ./internal/framework/prompt` +- `go test ./internal/modules/extract/dnd/spells` +- `go test ./...` + +## Stage 6: Manifest and Diagnostics Provenance + +Record reference provenance separately from source provenance. + +Implementation steps: + +- Add a dedicated references section to `artifacts.RunManifest`. + The shape should be lane-scoped and include lane ID, slot name, origin type, + origin URI/path, digest, media type, size bytes, and binding source. +- Do not add reference digests to `source_digests`. +- Include reference digests in any cache/idempotency key if such a key exists. + If no cache/idempotency key exists, add a test or comment documenting that no + additional key needs updating yet. +- Write a diagnostics artifact for resolved references that contains provenance + only, not full content, consistent with redacted effective config behavior. +- Ensure durable JSON output manifests include the new manifest section. +- Add manifest round-trip tests and a CLI/run test where two runs that differ + only in reference bytes produce distinguishable manifests. + +Verification: + +- `go test ./internal/core/artifacts` +- `go test ./internal/core/diagnostics` +- `go test ./internal/modules/output/json` +- `go test ./internal/cli` +- `go test ./...` + +## Stage 7: D&D Spells Consumer + +Use the new reference feature in the first production extractor. + +Implementation steps: + +- Declare optional `roster` and `glossary` slots on `dnd/spells`. +- Set accepted media type to text/UTF-8. Add conservative `MaxBytes` limits only + if a clear module-owned limit is chosen; otherwise leave `MaxBytes` unset. +- Update the D&D spells prompt bundle to include conditional reference sections + using `hasreference` and `reference`. +- Frame references as supporting material only. The prompt must instruct the + model to extract only spell-cast events present in the source transcript and + use references only for disambiguation. +- Update prompt metadata tests as needed while preserving template-hash + semantics. +- Add fixture coverage with no references, with roster/glossary references, and + with a roster that mentions a spell never cast in the transcript. The last + case must assert no spell-cast artifact is produced for the uncast spell. +- If existing deterministic source-reference validation can be extended + cleanly, add warning-level relatedness checks for spell names or close + variants near cited source text. If this becomes large, defer that validator + enhancement to a separate roadmap item and keep the prompt/regression fixture + guard in this stage. + +Verification: + +- `go test ./internal/modules/extract/dnd/spells` +- `go test ./internal/framework/pipeline` +- `go test ./internal/cli` +- `go test ./...` + +## Stage 8: Canonical Documentation and Examples + +Move implemented behavior out of roadmap-only status once code exists. + +Implementation steps: + +- Update `docs/cli.md` with `--reference` and `--without-reference` syntax, + precedence, ambiguity behavior, and examples. +- Update `docs/config.md` with pipeline-level and lane-level `references` + blocks. +- Update `docs/internal/modules.md` or the most appropriate internal docs with + module-author guidance for `ReferenceSlots()`, reference request delivery, + prompt functions, evidence exclusion, and provenance. +- Update `docs/internal/pipeline.md` with reference resolution lifecycle and + lane-scoped delivery. +- Update `docs/integrations/json-output.md` with the manifest reference + provenance shape. +- Update `docs/operations.md` or `docs/troubleshooting.md` for common reference + errors such as unknown slot, ambiguous flat override, missing required slot, + unreadable file, non-UTF-8 content, and `MaxBytes` failures. +- Add maintained example reference files and update `examples/dnd-spells.config.yml` + only after the CLI/config behavior is implemented and covered by tests. +- Keep future-only material in `docs/roadmap/references.md`; do not duplicate + canonical current behavior there after implementation. + +Verification: + +- `rg -n "references:|--reference|--without-reference|ReferenceSlots|reference \"|hasreference" docs examples` +- `go test ./...` +- `go vet ./...` +- `go build ./cmd/notarius` + +## Final Acceptance Criteria + +The feature is complete when: + +- extractor modules can declare reference slots through the first-class + extractor contract; +- config and CLI can bind and unbind lane-scoped file references; +- selected-pipeline validation catches unknown, ambiguous, or missing required + references before any LLM call; +- run preparation materializes UTF-8 text references with digests, size checks, + and empty-file warnings; +- extractors receive lane-scoped resolved references; +- prompt templates can render `reference` and `hasreference` deterministically; +- D&D spell extraction uses optional roster and glossary references; +- run manifests and diagnostics record reference provenance without recording + full content or treating references as source evidence; +- canonical docs and maintained examples describe only implemented behavior; +- `go test ./...`, `go vet ./...`, and `go build ./cmd/notarius` pass. diff --git a/docs/roadmap/references.md b/docs/roadmap/references.md index 1376d49..aa9544a 100644 --- a/docs/roadmap/references.md +++ b/docs/roadmap/references.md @@ -1,150 +1,117 @@ -# Feature Roadmap Proposal: Extraction Reference +# Feature Roadmap: Extraction References ## Status -This document captures proposed design and implementation sequencing for the -extraction-reference feature in Notarius. It describes planned work, not -implemented behavior. Go snippets are conceptual sketches; the implementing -agent should adapt names and shapes to the existing contracts, package -boundaries, and conventions in this repository. +This document defines the target state and policy choices for the planned +extraction-reference feature in Notarius. It describes planned behavior, not +implemented behavior. The staged implementation plan lives in +[implementation.md](implementation.md). ## Goal -Extraction quality improves significantly when the LLM receives reference -material alongside the source input. For the initial D&D spell extractor, -useful reference material includes a party roster (mapping players to player -characters), a player list, and a campaign glossary. +Extraction quality improves when an LLM-backed extractor receives stable +reference material alongside the source input. For the initial D&D spell +extractor, useful reference material includes a party roster, a player list, and +a campaign glossary. Notarius should support passing this material to extractors as **named reference -items** without introducing any domain-specific concepts into core or framework +items** without introducing domain-specific concepts into core or framework packages. The framework should know only that: - extractors declare named reference slots they accept; -- pipeline config and CLI flags bind content (initially files) to those slots; +- pipeline config and CLI flags bind content, initially files, to those slots; - bound content is rendered into module-owned prompt templates; - bound content is digested and recorded as run provenance. -Only extract modules should know what a "roster" or "glossary" means. All -domain semantics live in module-owned slot declarations and prompt templates. +Only extract modules should know what a "roster" or "glossary" means. Domain +semantics live in module-owned slot declarations and prompt templates. ## Definitions - **Reference slot**: a named, typed-by-convention input declared by an - extractor, with a human-readable description and a required/optional flag. - Example: extractor `dnd/spells` declares an optional slot named `roster`. -- **Reference item**: resolved content bound to a slot for a given run: name, - content bytes, media type, content digest, and origin (initially a file - path). + extractor, with a human-readable description, required/optional status, and + optional guardrails such as accepted media types and maximum bytes. +- **Reference item**: resolved content bound to a slot for a given run: slot + name, content bytes, media type, content digest, size, origin, and binding + source. - **Reference binding**: the association of a slot name to a content source, defined in pipeline config and overridable per run via CLI. +- **Reference set**: the lane-scoped collection of resolved reference items + delivered to an extractor. ## Architectural Principles -- Reference is opaque to the framework. Core and framework packages must not +- References are opaque to the framework. Core and framework packages must not interpret reference content or recognize domain slot names. -- Reference is an input. Anything that changes extraction output must be - digested into the run manifest and participate in any cache key. -- A Reference is not evidence. `SourceRef` values must only ever reference source - units. Reference items must not receive unit IDs and must not be addressable - by source references. -- Slots are declared, not ad hoc. Binding an undeclared slot name, or omitting - a required slot, should fail at config-load time, before any LLM call. +- References are inputs. Anything that can change extraction output must be + digested into the run manifest and participate in any cache or idempotency key. +- References are not evidence. `SourceRef` values must only ever reference + source units. Reference items must not receive unit IDs and must not be + addressable by source references. +- Slots are declared, not ad hoc. Binding an undeclared slot name, or omitting a + required slot, should fail before any LLM call. +- Reference delivery is lane-scoped. Pipeline-level bindings may apply to + multiple lanes, but each lane receives only the references declared by its + extractor after pipeline, lane, CLI override, and CLI unbind rules are + resolved. - Optional slots degrade gracefully. Prompt templates should render cleanly whether or not an optional slot is bound. -- Determinism. Identical input, config, prompts, and reference bytes should - produce byte-identical rendered prompts. Reference slots should render in a - stable, documented order (declaration order). +- Rendering must be deterministic. Identical source input, config, prompts, and + reference bytes should produce byte-identical rendered prompts. Reference + slots should render in declaration order, with stable binding order within a + slot. -## Proposed Contracts +## Target Contracts -### Slot declaration (extractor contract extension) +Extractors should declare accepted reference slots directly on the extractor +contract. This is a first-class feature, so mechanical updates to existing +extractors and test fakes are acceptable. -Extractors should declare the reference slots they accept: +The target slot declaration includes: -```go -type ReferenceSlot struct { - Name string - Description string - Required bool +- `Name`; +- `Description`; +- `Required`; +- `AcceptedMediaTypes`; +- `Multiple`; +- `MaxBytes`. - // MVP can leave these empty/defaulted, but having the fields now - // makes validation and future docs easier. - AcceptedMediaTypes []string - Multiple bool - MaxBytes int64 -} -``` +Extractors with no reference needs return an empty slot list. -The extractor interface should gain a method such as: +Resolved reference items should be content-bearing values, not unresolved file +paths. The MVP producer is "read this file," but the item shape should permit +future producers such as prior-run artifacts, derived summaries, or entity +registries without changing extractor-facing contracts. -```go -ReferenceSlots() []ReferenceSlot -``` +The extraction request should carry the lane-scoped resolved reference set. +Framework and core code should treat the set as opaque bytes plus metadata. -Extractors with no reference needs return an empty slice. Existing extractors -should require no other changes. +## Binding Lifecycle -### Resolved reference item +Reference handling should be split across existing lifecycle boundaries: -```go -type ReferenceItem struct { - SlotName string - MediaType string - Content []byte - Digest string - Origin ReferenceOrigin +1. Config parsing records pipeline-level and lane-level reference bindings + without reading files. +2. Pipeline resolution validates selected lanes, declared extractor slots, + missing required slots, unknown bindings, and ambiguous flat CLI bindings. +3. Run preparation resolves paths, reads files, validates media type and size, + computes digests, and materializes reference items. +4. Extraction receives the lane-specific resolved reference set. - SizeBytes int64 - TokenEstimate int -} - -type ReferenceOrigin struct { - Type string // "file" for MVP - URI string // path or future artifact URI -} - -type ReferenceSet struct { - // Stable declaration order, then stable binding order within a slot. - Slots []ResolvedReferenceSlot -} - -type ResolvedReferenceSlot struct { - Name string - Items []ReferenceItem -} -``` - -`ReferenceItem` is a resolved-content type, not a file path. The only MVP -producer is "read this file," but the shape should permit future producers -(prior-run artifacts, derived summaries, entity registries) without contract -changes. - -### Binding resolution - -A resolver should, at config-load time: - -1. Collect declared slots from every extractor selected by the active - pipeline (respecting lane selection, e.g. `--only`). -2. Collect bindings from pipeline config (pipeline-level and lane-level) and - CLI overrides, applying the standard layering: config file, then CLI. -3. Fail with a clear error if a required slot is unbound, or if a binding - references a slot no extractor within the selected pipeline declares. - Errors should name the pipeline, lane, slot, and the slot description. -4. Read, digest, and materialize each bound source into a `ReferenceItem`. -5. Enforce size guardrails (see Validation and Guardrails). +Config-relative paths resolve relative to the config file. CLI-relative paths +resolve relative to the current working directory. ## Configuration and CLI -### Pipeline config +Reference bindings should live in pipeline config because the initial use cases +are campaign-invariant more often than run-variant. Bindings should be supported +at two levels: -Reference bindings should live in pipeline config, because the initial use cases -(roster, glossary) are campaign-invariant rather than run-variant. Bindings -should be supported at two levels: +- pipeline level: defaults shared by artifact lanes whose extractors declare + matching slots; +- lane level: additions or overrides for a single artifact lane. -- pipeline level: shared by all artifact lanes; -- lane level: additions or overrides for a single lane. - -Illustrative shape (adapt to the existing config format): +Illustrative config shape: ```yaml pipelines: @@ -158,22 +125,29 @@ pipelines: extract: dnd/spells npcs: extract: dnd/npcs - reference: + references: npc_registry: ./campaign/npcs.md ``` -### CLI - -Per-run override flag, repeatable: +Per-run CLI binding overrides should be repeatable: ```text notarius run dnd-session --input session-014.json --reference roster=./alt_roster.md ``` -CLI bindings override config bindings for the same slot name. The existing -pipeline-describe/config-validate commands (or their nearest equivalents) -should surface declared slots, descriptions, required flags, and current -bindings so users can discover what a pipeline accepts. +Flat CLI slot names are allowed when unambiguous across selected lanes. Lane +qualified names, such as `spells.roster=./alt_roster.md`, disambiguate or target +a specific lane. CLI bindings override config bindings for the same lane and +slot. + +Users should also be able to unbind a config-bound optional slot for a run with +an explicit repeatable flag: + +```text +notarius run dnd-session --input session-014.json --without-reference roster +``` + +Unbinding a required slot should fail during pipeline/reference resolution. ## Prompt Template Integration @@ -185,30 +159,29 @@ Prompt templates are module-owned. Template rendering should expose: Rules: -- Referencing an **undeclared** slot from a template is a module bug and - should fail at prompt registration/build time (or earliest feasible point), - not silently at render time. -- Referencing a declared but unbound **optional** slot should render as - empty; templates should use `hasreference` to avoid dangling section headers. +- Referencing an undeclared slot from a template is a module bug and should fail + at prompt registration/build time or the earliest feasible equivalent. +- Referencing a declared but unbound optional slot should render as empty; + templates should use `hasreference` to avoid dangling section headers. - Rendering must be deterministic and independent of map iteration order. -- Prompt identity (registry hash) should be computed over the **template**, - not the rendered prompt. Reference digests are recorded separately in the - manifest, so a reference edit is visible as a reference change, not a prompt - change. +- Prompt identity should be computed over the template, not the rendered prompt. + Reference digests are recorded separately in the manifest so a reference edit + is visible as a reference change, not a prompt change. -Note: reference content is repeated in every per-chunk prompt. Diagnostics -should record per-slot token or byte counts so reference cost is observable. -Per-slot inclusion policies (e.g., roster in every chunk, glossary on demand) -are explicitly out of scope until cost data justifies them. +Reference content is repeated in every per-chunk prompt in the MVP. Per-slot or +per-chunk inclusion policies are deferred until cost data justifies them. ## Provenance -The run manifest must record, for every bound slot: +The run manifest must record resolved references separately from source +digests. For every bound lane and slot, it should record: +- lane ID; - slot name; -- origin (path); +- origin type and URI; - content digest; - media type; +- size in bytes; - whether the binding came from config or CLI override. Reference digests must participate in any idempotency/cache key alongside source @@ -216,125 +189,50 @@ digests, prompt hashes, schema versions, model, and parameters. Two runs that differ only in reference content must be distinguishable from the manifest alone. -Diagnostics for a run should include the resolved binding set (with digests, -not necessarily full content) in the run directory, consistent with the -existing redacted-effective-config pattern. - -## Path Resolution - -- Config-relative paths resolve relative to the pipeline config file. -- CLI-relative paths resolve relative to the current working directory. -- Manifest records the normalized absolute path or a redacted/display path -according to existing diagnostics policy. +Diagnostics for a run should include the resolved binding set with digests, not +full reference content, consistent with the existing redacted-effective-config +pattern. ## Validation and Guardrails -### References are not evidence +### References Are Not Evidence -The primary new failure mode: the model extracts facts from references rather -than from the source input. Example: the roster lists a PC's known spells, and -the model emits a `SpellCast` for a spell that was never cast in the session, -with a fabricated or misattributed source reference. +The primary new failure mode is the model extracting facts from references +rather than from the source input. For example, a roster may list a player +character's known spells, and the model might emit a spell-cast artifact for a +spell that was never cast in the session. Defenses, in priority order: -1. **Structural.** `SourceRef` remains the only grounding mechanism and can - only reference source units. No contract change should make references +1. **Structural.** `SourceRef` remains the only grounding mechanism and can only + reference source units. No contract change should make references addressable as evidence. 2. **Prompt discipline.** Module templates should frame references explicitly as - reference material, e.g. "use the roster to resolve speakers to - characters; extract only events that occur in the transcript." This - guidance belongs in the module prompt guidelines, not framework code. -3. **Validator support.** The source-reference validator (or a sibling - deterministic validator) should support checking that referenced source - text plausibly relates to the extracted fact (e.g., spell name or a close - variant appears in or near the referenced range). Severity should be - `warn`, not `fail`, given paraphrase and nickname casting. -4. **Regression fixtures.** Golden-file tests must include a fixture in which - the bound roster mentions a spell that is never cast in the transcript, - asserting no artifact record is produced for it. This regression is likely - to be reintroduced by future prompt edits; the fixture is the guard. + reference material, such as "use the roster to resolve speakers to + characters; extract only events that occur in the transcript." +3. **Validator support.** The source-reference validator, or a sibling + deterministic validator, should warn when referenced source text does not + plausibly relate to the extracted fact. Severity should be `warn`, not + `fail`, because transcripts can use paraphrase, nicknames, and abbreviations. +4. **Regression fixtures.** Tests should include a fixture in which a bound + roster mentions a spell that is never cast in the transcript, asserting no + artifact record is produced for it. -### Size and sanity guardrails +### Size and Sanity Guardrails -- Fail fast, before any LLM call, if bound references plus template plus largest - chunk exceeds the configured model context budget, with an error that names - the offending slot(s) and sizes. -- Empty bound files should produce a warning (probable user error). -- MVP accepts text content only (`utf-8`); other media - types should be rejected with a clear error. +- MVP accepts UTF-8 text content only. Other media types should be rejected with + a clear error. +- A slot-level `MaxBytes` value should be enforced when declared. +- Empty bound files should produce a warning because they are likely user error. -## Out of Scope (MVP) +## Out of Scope -- Non-file reference producers (prior-run artifacts, derived summaries, entity - registries). The `ReferenceItem` shape should permit them later. -- Per-chunk or per-slot inclusion policies and context budgeting beyond the - fail-fast guardrail. -- Structured/parsed references (e.g., typed roster schemas). References are opaque - text handed to prompts. -- Reference caching or preprocessing (summarization, embedding, retrieval). -- Making reference addressable as evidence, in any form. +- Token budgeting and model context-window management for references. +- Non-file reference producers, including prior-run artifacts, derived + summaries, and entity registries. +- Per-chunk or per-slot inclusion policies. +- Structured or parsed references such as typed roster schemas. References are + opaque text handed to prompts. +- Reference caching, preprocessing, summarization, embedding, or retrieval. +- Making references addressable as evidence in any form. -## Checkpoint Sequencing - -Each checkpoint should leave the repository compiling, with targeted tests -covering newly introduced contracts or behavior. - -1. **Contracts and resolution.** Add `ReferenceSlot`, `ReferenceItem`, and the - extractor `ReferenceSlots()` method (empty default for existing extractors). - Implement config parsing for pipeline- and lane-level bindings, CLI - override flag, layering, and load-time validation (unknown slot, missing - required slot, unreadable file, empty file warning). Unit tests for - resolution and error cases. -2. **Prompt rendering.** Add `reference`/`hasreference` template functions, - declaration-order rendering, undeclared-slot failure at registration, and - deterministic-render tests (byte-identical output across runs). -3. **Provenance.** Record bindings (name, origin, digest, media type, - binding source) in the run manifest and diagnostics; include reference - digests in the cache/idempotency key if one exists. Tests: manifest - round-trip; two runs differing only in reference content produce differing - manifests. -4. **Guardrails and validation.** Context-window fail-fast check; - relatedness `warn` validator (or extension of the source-reference - validator); media-type rejection. -5. **First consumer.** Declare `roster` (optional) and `glossary` (optional) - slots on the D&D spells extractor; update its prompt template with - conditional reference sections and reference-material framing; add golden - fixtures with and without references bound, including the - roster-mentions-uncast-spell fixture. This checkpoint is the acceptance - test for the feature: spell extraction quality with a roster bound should - visibly improve speaker-to-character attribution in fixtures. - -## Open Design Questions - -The implementing agent should resolve these against existing code and record -decisions in the implementation plan: - -- Should slot names be namespaced per lane in config and CLI (e.g., - `spells.roster=...`) or flat with lane-level config as the only - disambiguator? (Recommended default: flat names; lane-level config for - overrides; revisit if two extractors in one pipeline want the same slot - name with different content.) -- Where does binding resolution live relative to the existing config and - pipeline packages? It must run at load time, alongside existing pipeline - validation. -- Does the existing prompt registry hash templates or rendered prompts? If - rendered, this feature requires moving to template hashing as described in - Provenance. -- Should CLI overrides be permitted to bind slots that config leaves unbound - (yes, presumably), and to *unbind* a config-bound optional slot (e.g., - `--reference roster=` to clear)? Decide and test both directions. - -## Documentation Tasks - -Once implemented, move contracts out of this roadmap into canonical docs: - -- `docs/cli.md`: `--reference` flag syntax, layering, and examples; -- `docs/config.md`: pipeline- and lane-level `references` blocks; -- `docs/internal/`: slot/item contracts, resolution flow, evidence - exclusion rule, and template function reference for module authors; -- module-author guidance: how to declare slots, write conditional reference - sections, and frame reference material in prompts; -- `examples/`: a maintained example pipeline with a roster and glossary - bound, plus matching fixture files. -``` \ No newline at end of file