Add a staged implementation plan for background context references

This commit is contained in:
2026-07-05 09:08:31 -05:00
parent 11d8187052
commit f9999a73df
2 changed files with 439 additions and 236 deletions

View File

@@ -0,0 +1,305 @@
# Extraction References Implementation Plan
## Purpose
Implement the extraction-reference feature described in
[references.md](references.md). This plan is decision-complete for an LLM coding
agent: implement each stage in order, keep the repository compiling after each
stage, and do not move planned behavior into non-roadmap docs until the relevant
behavior exists.
Core decisions to preserve:
- references are opaque framework inputs and domain semantics stay in extract
modules;
- references are not evidence and must not be addressable through `SourceRef`;
- reference binding is lane-scoped;
- extractors expose `ReferenceSlots()` directly on the first-class extractor
contract;
- token budgeting is deferred; enforce only UTF-8 text handling, empty-file
warnings, and declared `MaxBytes`;
- run manifests record references in a dedicated section, separate from
`source_digests`;
- CLI unbinding uses `--without-reference`.
## Stage 1: Contracts and Mechanical Adoption
Add the framework contracts needed to describe references without changing
runtime behavior. Slot declarations must be available without constructing
extractor modules, because pipeline/config validation should use registry
metadata rather than runtime module instances.
Implementation steps:
- In `internal/framework/contracts`, add reference model types:
`ReferenceSlot`, `ReferenceOrigin`, `ReferenceItem`,
`ResolvedReferenceSlot`, `ReferenceSet`, and a binding-source enum or string
constants for `config` and `cli`.
- Include `Name`, `Description`, `Required`, `AcceptedMediaTypes`, `Multiple`,
and `MaxBytes` on `ReferenceSlot`.
- Include slot name, media type, content bytes, digest, origin, size bytes, and
binding source on `ReferenceItem`.
- Add `References ReferenceSet` to `contracts.ExtractionRequest`.
- Add `ReferenceSlots() []ReferenceSlot` to `contracts.Extractor`.
- Extend extractor registration metadata so reference slots are also declared
through the extractor's registry spec. Prefer the smallest idiomatic change to
the existing registry model, such as adding `ReferenceSlots` to `ModuleSpec`
with validation that non-extractor modules leave it empty, unless the codebase
shape clearly supports a narrower extractor-specific spec.
- Update every concrete extractor and all extractor fakes/test doubles to
implement `ReferenceSlots()`. Existing extractors without references should
return `nil`.
- Add tests that compare a production extractor's runtime `ReferenceSlots()`
with its registered spec slots so the two declarations cannot drift.
- Add contract tests for empty reference sets, slot copying expectations if
helpers are introduced, and compile-time coverage through existing fakes.
Verification:
- `go test ./internal/framework/contracts`
- `go test ./internal/framework/pipeline`
- `go test ./...`
## Stage 2: Config Shape and Pipeline-Level Resolution
Add unresolved reference bindings to config and resolved lane bindings to the
pipeline model. Do not read reference files in this stage.
Implementation steps:
- Add `references` maps to file config parsing at both pipeline and artifact
lane level.
- Add corresponding fields to `pipeline.PipelineProfile` and
`pipeline.ArtifactLaneProfile`.
- Preserve deterministic map handling and duplicate-after-trim validation.
- Extend config cloning, effective config, redaction, validation, and tests for
the new fields.
- Add resolved reference binding structures to `internal/framework/pipeline`.
They should represent lane ID, slot name, source URI/path, and binding source,
but not file bytes.
- During `pipeline.ResolvePipeline`, collect selected lanes, read each lane
extractor's declared slots from registry metadata, and validate without
building extractor instances:
- every bound slot is declared by the lane extractor;
- required slots are bound after applying pipeline-level and lane-level config;
- required slots remain bound after any CLI unbinds supplied to resolution;
- selected lanes under `--only` are the only lanes considered.
- Apply pipeline-level bindings as defaults only to lanes whose extractor
declares the matching slot.
- Apply lane-level bindings as overrides or additions for that lane.
- Keep reference bindings out of source digests and artifact source references.
- Add tests proving reference-slot validation works through registry specs even
when extractor constructors would fail if called.
Verification:
- `go test ./internal/core/config`
- `go test ./internal/framework/pipeline`
- `go test ./...`
## Stage 3: CLI Reference Overrides and Unbinds
Add run-time CLI syntax for reference binding overrides and optional unbinding.
Implementation steps:
- Add repeatable `--reference` flags to `notarius run`.
Accepted forms:
- `slot=path` for unambiguous slot names across selected lanes;
- `lane.slot=path` for explicit lane-scoped binding.
- Add repeatable `--without-reference` flags to `notarius run`.
Accepted forms:
- `slot`;
- `lane.slot`.
- Reject empty paths for `--reference`; use `--without-reference` for unbinding.
- Reject malformed values with concise CLI errors before expensive work.
- Pass parsed override/unbind requests into config/pipeline resolution.
- Resolve flat CLI names only when exactly one selected lane declares the slot.
If multiple selected lanes declare the same slot, fail and instruct the user
to use `lane.slot`.
- Let CLI bindings override config bindings for the same lane and slot.
- Let CLI unbinds remove config-bound optional slots for the same lane and slot.
- Fail if unbinding leaves a required slot unbound.
- Add CLI tests for flat binding, lane-qualified binding, ambiguous flat
binding, malformed syntax, optional unbind, and required-slot unbind failure.
Verification:
- `go test ./internal/cli`
- `go test ./internal/core/config`
- `go test ./internal/framework/pipeline`
- `go test ./...`
## Stage 4: Run Preparation and Reference Materialization
Read, validate, digest, and materialize resolved file references before any LLM
call.
Implementation steps:
- Add a reference resolver/materializer near pipeline run preparation. Keep file
I/O out of pure config parsing.
- Ensure run preparation receives the loaded config path or config directory so
config-relative reference paths can be resolved after pure config parsing.
- Resolve config-relative paths relative to the config file path and
CLI-relative paths relative to the current working directory.
- For MVP, accept only UTF-8 text files. Reject non-UTF-8 content with an error
naming pipeline, lane, slot, and path.
- Compute `sha256:` content digests over the raw reference bytes.
- Populate `ReferenceItem` values with content bytes, media type, digest,
origin type `file`, normalized origin URI/path, size bytes, and binding source.
- Enforce declared `MaxBytes` when greater than zero. The error should name the
pipeline, lane, slot, actual size, limit, and path.
- Emit a warning for empty bound files, but do not fail.
- Add `ReferenceSet` values to the runner input or resolved pipeline path in a
way that keeps lane-scoped references available when calling each extractor.
- Pass the correct lane-specific `ReferenceSet` into
`contracts.ExtractionRequest`.
- Ensure no reference content is written to ordinary diagnostics, logs, errors,
or manifests.
Verification:
- Focused resolver/materializer tests for path resolution, digest stability,
UTF-8 rejection, empty-file warning, `MaxBytes`, and binding source.
- `go test ./internal/cli`
- `go test ./internal/framework/pipeline`
- `go test ./...`
## Stage 5: Prompt Template Reference Functions
Make references available to module-owned prompt templates.
Implementation steps:
- Extend `internal/framework/prompt` so prompt bundles can be compiled with
declared reference slots.
- Add `reference` and `hasreference` template functions.
- Validate at bundle build time, or the earliest feasible equivalent, that
templates reference only declared slots.
- Render a declared but unbound optional slot as an empty string.
- Ensure `hasreference` returns true only when the slot has at least one bound
item with content.
- Render multiple items deterministically if future `Multiple` support is
enabled; for MVP, reject multiple bindings unless the slot declares
`Multiple`.
- Keep prompt metadata hashes based on template source. Do not include rendered
reference content in prompt identity.
- Add deterministic rendering tests proving byte-identical output across runs
with the same reference bytes and config.
Verification:
- `go test ./internal/framework/prompt`
- `go test ./internal/modules/extract/dnd/spells`
- `go test ./...`
## Stage 6: Manifest and Diagnostics Provenance
Record reference provenance separately from source provenance.
Implementation steps:
- Add a dedicated references section to `artifacts.RunManifest`.
The shape should be lane-scoped and include lane ID, slot name, origin type,
origin URI/path, digest, media type, size bytes, and binding source.
- Do not add reference digests to `source_digests`.
- Include reference digests in any cache/idempotency key if such a key exists.
If no cache/idempotency key exists, add a test or comment documenting that no
additional key needs updating yet.
- Write a diagnostics artifact for resolved references that contains provenance
only, not full content, consistent with redacted effective config behavior.
- Ensure durable JSON output manifests include the new manifest section.
- Add manifest round-trip tests and a CLI/run test where two runs that differ
only in reference bytes produce distinguishable manifests.
Verification:
- `go test ./internal/core/artifacts`
- `go test ./internal/core/diagnostics`
- `go test ./internal/modules/output/json`
- `go test ./internal/cli`
- `go test ./...`
## Stage 7: D&D Spells Consumer
Use the new reference feature in the first production extractor.
Implementation steps:
- Declare optional `roster` and `glossary` slots on `dnd/spells`.
- Set accepted media type to text/UTF-8. Add conservative `MaxBytes` limits only
if a clear module-owned limit is chosen; otherwise leave `MaxBytes` unset.
- Update the D&D spells prompt bundle to include conditional reference sections
using `hasreference` and `reference`.
- Frame references as supporting material only. The prompt must instruct the
model to extract only spell-cast events present in the source transcript and
use references only for disambiguation.
- Update prompt metadata tests as needed while preserving template-hash
semantics.
- Add fixture coverage with no references, with roster/glossary references, and
with a roster that mentions a spell never cast in the transcript. The last
case must assert no spell-cast artifact is produced for the uncast spell.
- If existing deterministic source-reference validation can be extended
cleanly, add warning-level relatedness checks for spell names or close
variants near cited source text. If this becomes large, defer that validator
enhancement to a separate roadmap item and keep the prompt/regression fixture
guard in this stage.
Verification:
- `go test ./internal/modules/extract/dnd/spells`
- `go test ./internal/framework/pipeline`
- `go test ./internal/cli`
- `go test ./...`
## Stage 8: Canonical Documentation and Examples
Move implemented behavior out of roadmap-only status once code exists.
Implementation steps:
- Update `docs/cli.md` with `--reference` and `--without-reference` syntax,
precedence, ambiguity behavior, and examples.
- Update `docs/config.md` with pipeline-level and lane-level `references`
blocks.
- Update `docs/internal/modules.md` or the most appropriate internal docs with
module-author guidance for `ReferenceSlots()`, reference request delivery,
prompt functions, evidence exclusion, and provenance.
- Update `docs/internal/pipeline.md` with reference resolution lifecycle and
lane-scoped delivery.
- Update `docs/integrations/json-output.md` with the manifest reference
provenance shape.
- Update `docs/operations.md` or `docs/troubleshooting.md` for common reference
errors such as unknown slot, ambiguous flat override, missing required slot,
unreadable file, non-UTF-8 content, and `MaxBytes` failures.
- Add maintained example reference files and update `examples/dnd-spells.config.yml`
only after the CLI/config behavior is implemented and covered by tests.
- Keep future-only material in `docs/roadmap/references.md`; do not duplicate
canonical current behavior there after implementation.
Verification:
- `rg -n "references:|--reference|--without-reference|ReferenceSlots|reference \"|hasreference" docs examples`
- `go test ./...`
- `go vet ./...`
- `go build ./cmd/notarius`
## Final Acceptance Criteria
The feature is complete when:
- extractor modules can declare reference slots through the first-class
extractor contract;
- config and CLI can bind and unbind lane-scoped file references;
- selected-pipeline validation catches unknown, ambiguous, or missing required
references before any LLM call;
- run preparation materializes UTF-8 text references with digests, size checks,
and empty-file warnings;
- extractors receive lane-scoped resolved references;
- prompt templates can render `reference` and `hasreference` deterministically;
- D&D spell extraction uses optional roster and glossary references;
- run manifests and diagnostics record reference provenance without recording
full content or treating references as source evidence;
- canonical docs and maintained examples describe only implemented behavior;
- `go test ./...`, `go vet ./...`, and `go build ./cmd/notarius` pass.

View File

@@ -1,150 +1,117 @@
# Feature Roadmap Proposal: Extraction Reference # Feature Roadmap: Extraction References
## Status ## Status
This document captures proposed design and implementation sequencing for the This document defines the target state and policy choices for the planned
extraction-reference feature in Notarius. It describes planned work, not extraction-reference feature in Notarius. It describes planned behavior, not
implemented behavior. Go snippets are conceptual sketches; the implementing implemented behavior. The staged implementation plan lives in
agent should adapt names and shapes to the existing contracts, package [implementation.md](implementation.md).
boundaries, and conventions in this repository.
## Goal ## Goal
Extraction quality improves significantly when the LLM receives reference Extraction quality improves when an LLM-backed extractor receives stable
material alongside the source input. For the initial D&D spell extractor, reference material alongside the source input. For the initial D&D spell
useful reference material includes a party roster (mapping players to player extractor, useful reference material includes a party roster, a player list, and
characters), a player list, and a campaign glossary. a campaign glossary.
Notarius should support passing this material to extractors as **named reference Notarius should support passing this material to extractors as **named reference
items** without introducing any domain-specific concepts into core or framework items** without introducing domain-specific concepts into core or framework
packages. The framework should know only that: packages. The framework should know only that:
- extractors declare named reference slots they accept; - extractors declare named reference slots they accept;
- pipeline config and CLI flags bind content (initially files) to those slots; - pipeline config and CLI flags bind content, initially files, to those slots;
- bound content is rendered into module-owned prompt templates; - bound content is rendered into module-owned prompt templates;
- bound content is digested and recorded as run provenance. - bound content is digested and recorded as run provenance.
Only extract modules should know what a "roster" or "glossary" means. All Only extract modules should know what a "roster" or "glossary" means. Domain
domain semantics live in module-owned slot declarations and prompt templates. semantics live in module-owned slot declarations and prompt templates.
## Definitions ## Definitions
- **Reference slot**: a named, typed-by-convention input declared by an - **Reference slot**: a named, typed-by-convention input declared by an
extractor, with a human-readable description and a required/optional flag. extractor, with a human-readable description, required/optional status, and
Example: extractor `dnd/spells` declares an optional slot named `roster`. optional guardrails such as accepted media types and maximum bytes.
- **Reference item**: resolved content bound to a slot for a given run: name, - **Reference item**: resolved content bound to a slot for a given run: slot
content bytes, media type, content digest, and origin (initially a file name, content bytes, media type, content digest, size, origin, and binding
path). source.
- **Reference binding**: the association of a slot name to a content source, - **Reference binding**: the association of a slot name to a content source,
defined in pipeline config and overridable per run via CLI. defined in pipeline config and overridable per run via CLI.
- **Reference set**: the lane-scoped collection of resolved reference items
delivered to an extractor.
## Architectural Principles ## Architectural Principles
- Reference is opaque to the framework. Core and framework packages must not - References are opaque to the framework. Core and framework packages must not
interpret reference content or recognize domain slot names. interpret reference content or recognize domain slot names.
- Reference is an input. Anything that changes extraction output must be - References are inputs. Anything that can change extraction output must be
digested into the run manifest and participate in any cache key. digested into the run manifest and participate in any cache or idempotency key.
- A Reference is not evidence. `SourceRef` values must only ever reference source - References are not evidence. `SourceRef` values must only ever reference
units. Reference items must not receive unit IDs and must not be addressable source units. Reference items must not receive unit IDs and must not be
by source references. addressable by source references.
- Slots are declared, not ad hoc. Binding an undeclared slot name, or omitting - Slots are declared, not ad hoc. Binding an undeclared slot name, or omitting a
a required slot, should fail at config-load time, before any LLM call. required slot, should fail before any LLM call.
- Reference delivery is lane-scoped. Pipeline-level bindings may apply to
multiple lanes, but each lane receives only the references declared by its
extractor after pipeline, lane, CLI override, and CLI unbind rules are
resolved.
- Optional slots degrade gracefully. Prompt templates should render cleanly - Optional slots degrade gracefully. Prompt templates should render cleanly
whether or not an optional slot is bound. whether or not an optional slot is bound.
- Determinism. Identical input, config, prompts, and reference bytes should - Rendering must be deterministic. Identical source input, config, prompts, and
produce byte-identical rendered prompts. Reference slots should render in a reference bytes should produce byte-identical rendered prompts. Reference
stable, documented order (declaration order). slots should render in declaration order, with stable binding order within a
slot.
## Proposed Contracts ## Target Contracts
### Slot declaration (extractor contract extension) Extractors should declare accepted reference slots directly on the extractor
contract. This is a first-class feature, so mechanical updates to existing
extractors and test fakes are acceptable.
Extractors should declare the reference slots they accept: The target slot declaration includes:
```go - `Name`;
type ReferenceSlot struct { - `Description`;
Name string - `Required`;
Description string - `AcceptedMediaTypes`;
Required bool - `Multiple`;
- `MaxBytes`.
// MVP can leave these empty/defaulted, but having the fields now Extractors with no reference needs return an empty slot list.
// makes validation and future docs easier.
AcceptedMediaTypes []string
Multiple bool
MaxBytes int64
}
```
The extractor interface should gain a method such as: Resolved reference items should be content-bearing values, not unresolved file
paths. The MVP producer is "read this file," but the item shape should permit
future producers such as prior-run artifacts, derived summaries, or entity
registries without changing extractor-facing contracts.
```go The extraction request should carry the lane-scoped resolved reference set.
ReferenceSlots() []ReferenceSlot Framework and core code should treat the set as opaque bytes plus metadata.
```
Extractors with no reference needs return an empty slice. Existing extractors ## Binding Lifecycle
should require no other changes.
### Resolved reference item Reference handling should be split across existing lifecycle boundaries:
```go 1. Config parsing records pipeline-level and lane-level reference bindings
type ReferenceItem struct { without reading files.
SlotName string 2. Pipeline resolution validates selected lanes, declared extractor slots,
MediaType string missing required slots, unknown bindings, and ambiguous flat CLI bindings.
Content []byte 3. Run preparation resolves paths, reads files, validates media type and size,
Digest string computes digests, and materializes reference items.
Origin ReferenceOrigin 4. Extraction receives the lane-specific resolved reference set.
SizeBytes int64 Config-relative paths resolve relative to the config file. CLI-relative paths
TokenEstimate int resolve relative to the current working directory.
}
type ReferenceOrigin struct {
Type string // "file" for MVP
URI string // path or future artifact URI
}
type ReferenceSet struct {
// Stable declaration order, then stable binding order within a slot.
Slots []ResolvedReferenceSlot
}
type ResolvedReferenceSlot struct {
Name string
Items []ReferenceItem
}
```
`ReferenceItem` is a resolved-content type, not a file path. The only MVP
producer is "read this file," but the shape should permit future producers
(prior-run artifacts, derived summaries, entity registries) without contract
changes.
### Binding resolution
A resolver should, at config-load time:
1. Collect declared slots from every extractor selected by the active
pipeline (respecting lane selection, e.g. `--only`).
2. Collect bindings from pipeline config (pipeline-level and lane-level) and
CLI overrides, applying the standard layering: config file, then CLI.
3. Fail with a clear error if a required slot is unbound, or if a binding
references a slot no extractor within the selected pipeline declares.
Errors should name the pipeline, lane, slot, and the slot description.
4. Read, digest, and materialize each bound source into a `ReferenceItem`.
5. Enforce size guardrails (see Validation and Guardrails).
## Configuration and CLI ## Configuration and CLI
### Pipeline config Reference bindings should live in pipeline config because the initial use cases
are campaign-invariant more often than run-variant. Bindings should be supported
at two levels:
Reference bindings should live in pipeline config, because the initial use cases - pipeline level: defaults shared by artifact lanes whose extractors declare
(roster, glossary) are campaign-invariant rather than run-variant. Bindings matching slots;
should be supported at two levels: - lane level: additions or overrides for a single artifact lane.
- pipeline level: shared by all artifact lanes; Illustrative config shape:
- lane level: additions or overrides for a single lane.
Illustrative shape (adapt to the existing config format):
```yaml ```yaml
pipelines: pipelines:
@@ -158,22 +125,29 @@ pipelines:
extract: dnd/spells extract: dnd/spells
npcs: npcs:
extract: dnd/npcs extract: dnd/npcs
reference: references:
npc_registry: ./campaign/npcs.md npc_registry: ./campaign/npcs.md
``` ```
### CLI Per-run CLI binding overrides should be repeatable:
Per-run override flag, repeatable:
```text ```text
notarius run dnd-session --input session-014.json --reference roster=./alt_roster.md notarius run dnd-session --input session-014.json --reference roster=./alt_roster.md
``` ```
CLI bindings override config bindings for the same slot name. The existing Flat CLI slot names are allowed when unambiguous across selected lanes. Lane
pipeline-describe/config-validate commands (or their nearest equivalents) qualified names, such as `spells.roster=./alt_roster.md`, disambiguate or target
should surface declared slots, descriptions, required flags, and current a specific lane. CLI bindings override config bindings for the same lane and
bindings so users can discover what a pipeline accepts. slot.
Users should also be able to unbind a config-bound optional slot for a run with
an explicit repeatable flag:
```text
notarius run dnd-session --input session-014.json --without-reference roster
```
Unbinding a required slot should fail during pipeline/reference resolution.
## Prompt Template Integration ## Prompt Template Integration
@@ -185,30 +159,29 @@ Prompt templates are module-owned. Template rendering should expose:
Rules: Rules:
- Referencing an **undeclared** slot from a template is a module bug and - Referencing an undeclared slot from a template is a module bug and should fail
should fail at prompt registration/build time (or earliest feasible point), at prompt registration/build time or the earliest feasible equivalent.
not silently at render time. - Referencing a declared but unbound optional slot should render as empty;
- Referencing a declared but unbound **optional** slot should render as templates should use `hasreference` to avoid dangling section headers.
empty; templates should use `hasreference` to avoid dangling section headers.
- Rendering must be deterministic and independent of map iteration order. - Rendering must be deterministic and independent of map iteration order.
- Prompt identity (registry hash) should be computed over the **template**, - Prompt identity should be computed over the template, not the rendered prompt.
not the rendered prompt. Reference digests are recorded separately in the Reference digests are recorded separately in the manifest so a reference edit
manifest, so a reference edit is visible as a reference change, not a prompt is visible as a reference change, not a prompt change.
change.
Note: reference content is repeated in every per-chunk prompt. Diagnostics Reference content is repeated in every per-chunk prompt in the MVP. Per-slot or
should record per-slot token or byte counts so reference cost is observable. per-chunk inclusion policies are deferred until cost data justifies them.
Per-slot inclusion policies (e.g., roster in every chunk, glossary on demand)
are explicitly out of scope until cost data justifies them.
## Provenance ## Provenance
The run manifest must record, for every bound slot: The run manifest must record resolved references separately from source
digests. For every bound lane and slot, it should record:
- lane ID;
- slot name; - slot name;
- origin (path); - origin type and URI;
- content digest; - content digest;
- media type; - media type;
- size in bytes;
- whether the binding came from config or CLI override. - whether the binding came from config or CLI override.
Reference digests must participate in any idempotency/cache key alongside source Reference digests must participate in any idempotency/cache key alongside source
@@ -216,125 +189,50 @@ digests, prompt hashes, schema versions, model, and parameters. Two runs that
differ only in reference content must be distinguishable from the manifest differ only in reference content must be distinguishable from the manifest
alone. alone.
Diagnostics for a run should include the resolved binding set (with digests, Diagnostics for a run should include the resolved binding set with digests, not
not necessarily full content) in the run directory, consistent with the full reference content, consistent with the existing redacted-effective-config
existing redacted-effective-config pattern. pattern.
## Path Resolution
- Config-relative paths resolve relative to the pipeline config file.
- CLI-relative paths resolve relative to the current working directory.
- Manifest records the normalized absolute path or a redacted/display path
according to existing diagnostics policy.
## Validation and Guardrails ## Validation and Guardrails
### References are not evidence ### References Are Not Evidence
The primary new failure mode: the model extracts facts from references rather The primary new failure mode is the model extracting facts from references
than from the source input. Example: the roster lists a PC's known spells, and rather than from the source input. For example, a roster may list a player
the model emits a `SpellCast` for a spell that was never cast in the session, character's known spells, and the model might emit a spell-cast artifact for a
with a fabricated or misattributed source reference. spell that was never cast in the session.
Defenses, in priority order: Defenses, in priority order:
1. **Structural.** `SourceRef` remains the only grounding mechanism and can 1. **Structural.** `SourceRef` remains the only grounding mechanism and can only
only reference source units. No contract change should make references reference source units. No contract change should make references
addressable as evidence. addressable as evidence.
2. **Prompt discipline.** Module templates should frame references explicitly as 2. **Prompt discipline.** Module templates should frame references explicitly as
reference material, e.g. "use the roster to resolve speakers to reference material, such as "use the roster to resolve speakers to
characters; extract only events that occur in the transcript." This characters; extract only events that occur in the transcript."
guidance belongs in the module prompt guidelines, not framework code. 3. **Validator support.** The source-reference validator, or a sibling
3. **Validator support.** The source-reference validator (or a sibling deterministic validator, should warn when referenced source text does not
deterministic validator) should support checking that referenced source plausibly relate to the extracted fact. Severity should be `warn`, not
text plausibly relates to the extracted fact (e.g., spell name or a close `fail`, because transcripts can use paraphrase, nicknames, and abbreviations.
variant appears in or near the referenced range). Severity should be 4. **Regression fixtures.** Tests should include a fixture in which a bound
`warn`, not `fail`, given paraphrase and nickname casting. roster mentions a spell that is never cast in the transcript, asserting no
4. **Regression fixtures.** Golden-file tests must include a fixture in which artifact record is produced for it.
the bound roster mentions a spell that is never cast in the transcript,
asserting no artifact record is produced for it. This regression is likely
to be reintroduced by future prompt edits; the fixture is the guard.
### Size and sanity guardrails ### Size and Sanity Guardrails
- Fail fast, before any LLM call, if bound references plus template plus largest - MVP accepts UTF-8 text content only. Other media types should be rejected with
chunk exceeds the configured model context budget, with an error that names a clear error.
the offending slot(s) and sizes. - A slot-level `MaxBytes` value should be enforced when declared.
- Empty bound files should produce a warning (probable user error). - Empty bound files should produce a warning because they are likely user error.
- MVP accepts text content only (`utf-8`); other media
types should be rejected with a clear error.
## Out of Scope (MVP) ## Out of Scope
- Non-file reference producers (prior-run artifacts, derived summaries, entity - Token budgeting and model context-window management for references.
registries). The `ReferenceItem` shape should permit them later. - Non-file reference producers, including prior-run artifacts, derived
- Per-chunk or per-slot inclusion policies and context budgeting beyond the summaries, and entity registries.
fail-fast guardrail. - Per-chunk or per-slot inclusion policies.
- Structured/parsed references (e.g., typed roster schemas). References are opaque - Structured or parsed references such as typed roster schemas. References are
text handed to prompts. opaque text handed to prompts.
- Reference caching or preprocessing (summarization, embedding, retrieval). - Reference caching, preprocessing, summarization, embedding, or retrieval.
- Making reference addressable as evidence, in any form. - Making references addressable as evidence in any form.
## Checkpoint Sequencing
Each checkpoint should leave the repository compiling, with targeted tests
covering newly introduced contracts or behavior.
1. **Contracts and resolution.** Add `ReferenceSlot`, `ReferenceItem`, and the
extractor `ReferenceSlots()` method (empty default for existing extractors).
Implement config parsing for pipeline- and lane-level bindings, CLI
override flag, layering, and load-time validation (unknown slot, missing
required slot, unreadable file, empty file warning). Unit tests for
resolution and error cases.
2. **Prompt rendering.** Add `reference`/`hasreference` template functions,
declaration-order rendering, undeclared-slot failure at registration, and
deterministic-render tests (byte-identical output across runs).
3. **Provenance.** Record bindings (name, origin, digest, media type,
binding source) in the run manifest and diagnostics; include reference
digests in the cache/idempotency key if one exists. Tests: manifest
round-trip; two runs differing only in reference content produce differing
manifests.
4. **Guardrails and validation.** Context-window fail-fast check;
relatedness `warn` validator (or extension of the source-reference
validator); media-type rejection.
5. **First consumer.** Declare `roster` (optional) and `glossary` (optional)
slots on the D&D spells extractor; update its prompt template with
conditional reference sections and reference-material framing; add golden
fixtures with and without references bound, including the
roster-mentions-uncast-spell fixture. This checkpoint is the acceptance
test for the feature: spell extraction quality with a roster bound should
visibly improve speaker-to-character attribution in fixtures.
## Open Design Questions
The implementing agent should resolve these against existing code and record
decisions in the implementation plan:
- Should slot names be namespaced per lane in config and CLI (e.g.,
`spells.roster=...`) or flat with lane-level config as the only
disambiguator? (Recommended default: flat names; lane-level config for
overrides; revisit if two extractors in one pipeline want the same slot
name with different content.)
- Where does binding resolution live relative to the existing config and
pipeline packages? It must run at load time, alongside existing pipeline
validation.
- Does the existing prompt registry hash templates or rendered prompts? If
rendered, this feature requires moving to template hashing as described in
Provenance.
- Should CLI overrides be permitted to bind slots that config leaves unbound
(yes, presumably), and to *unbind* a config-bound optional slot (e.g.,
`--reference roster=` to clear)? Decide and test both directions.
## Documentation Tasks
Once implemented, move contracts out of this roadmap into canonical docs:
- `docs/cli.md`: `--reference` flag syntax, layering, and examples;
- `docs/config.md`: pipeline- and lane-level `references` blocks;
- `docs/internal/`: slot/item contracts, resolution flow, evidence
exclusion rule, and template function reference for module authors;
- module-author guidance: how to declare slots, write conditional reference
sections, and frame reference material in prompts;
- `examples/`: a maintained example pipeline with a roster and glossary
bound, plus matching fixture files.
```