Files
notarius/docs/roadmap/references.md

340 lines
14 KiB
Markdown

# Feature Roadmap Proposal: Extraction Reference
## Status
This document captures proposed design and implementation sequencing for the
extraction-reference feature in Notarius. It describes planned work, not
implemented behavior. Go snippets are conceptual sketches; the implementing
agent should adapt names and shapes to the existing contracts, package
boundaries, and conventions in this repository.
## Goal
Extraction quality improves significantly when the LLM receives reference
material alongside the source input. For the initial D&D spell extractor,
useful reference material includes a party roster (mapping players to player
characters), a player list, and a campaign glossary.
Notarius should support passing this material to extractors as **named reference
items** without introducing any domain-specific concepts into core or framework
packages. The framework should know only that:
- extractors declare named reference slots they accept;
- pipeline config and CLI flags bind content (initially files) to those slots;
- bound content is rendered into module-owned prompt templates;
- bound content is digested and recorded as run provenance.
Only extract modules should know what a "roster" or "glossary" means. All
domain semantics live in module-owned slot declarations and prompt templates.
## Definitions
- **Reference slot**: a named, typed-by-convention input declared by an
extractor, with a human-readable description and a required/optional flag.
Example: extractor `dnd/spells` declares an optional slot named `roster`.
- **Reference item**: resolved content bound to a slot for a given run: name,
content bytes, media type, content digest, and origin (initially a file
path).
- **Reference binding**: the association of a slot name to a content source,
defined in pipeline config and overridable per run via CLI.
## Architectural Principles
- Reference is opaque to the framework. Core and framework packages must not
interpret reference content or recognize domain slot names.
- Reference is an input. Anything that changes extraction output must be
digested into the run manifest and participate in any cache key.
- A Reference is not evidence. `SourceRef` values must only ever reference source
units. Reference items must not receive unit IDs and must not be addressable
by source references.
- Slots are declared, not ad hoc. Binding an undeclared slot name, or omitting
a required slot, should fail at config-load time, before any LLM call.
- Optional slots degrade gracefully. Prompt templates should render cleanly
whether or not an optional slot is bound.
- Determinism. Identical input, config, prompts, and reference bytes should
produce byte-identical rendered prompts. Reference slots should render in a
stable, documented order (declaration order).
## Proposed Contracts
### Slot declaration (extractor contract extension)
Extractors should declare the reference slots they accept:
```go
type ReferenceSlot struct {
Name string
Description string
Required bool
// MVP can leave these empty/defaulted, but having the fields now
// makes validation and future docs easier.
AcceptedMediaTypes []string
Multiple bool
MaxBytes int64
}
```
The extractor interface should gain a method such as:
```go
ReferenceSlots() []ReferenceSlot
```
Extractors with no reference needs return an empty slice. Existing extractors
should require no other changes.
### Resolved reference item
```go
type ReferenceItem struct {
SlotName string
MediaType string
Content []byte
Digest string
Origin ReferenceOrigin
SizeBytes int64
TokenEstimate int
}
type ReferenceOrigin struct {
Type string // "file" for MVP
URI string // path or future artifact URI
}
type ReferenceSet struct {
// Stable declaration order, then stable binding order within a slot.
Slots []ResolvedReferenceSlot
}
type ResolvedReferenceSlot struct {
Name string
Items []ReferenceItem
}
```
`ReferenceItem` is a resolved-content type, not a file path. The only MVP
producer is "read this file," but the shape should permit future producers
(prior-run artifacts, derived summaries, entity registries) without contract
changes.
### Binding resolution
A resolver should, at config-load time:
1. Collect declared slots from every extractor selected by the active
pipeline (respecting lane selection, e.g. `--only`).
2. Collect bindings from pipeline config (pipeline-level and lane-level) and
CLI overrides, applying the standard layering: config file, then CLI.
3. Fail with a clear error if a required slot is unbound, or if a binding
references a slot no extractor within the selected pipeline declares.
Errors should name the pipeline, lane, slot, and the slot description.
4. Read, digest, and materialize each bound source into a `ReferenceItem`.
5. Enforce size guardrails (see Validation and Guardrails).
## Configuration and CLI
### Pipeline config
Reference bindings should live in pipeline config, because the initial use cases
(roster, glossary) are campaign-invariant rather than run-variant. Bindings
should be supported at two levels:
- pipeline level: shared by all artifact lanes;
- lane level: additions or overrides for a single lane.
Illustrative shape (adapt to the existing config format):
```yaml
pipelines:
dnd-session:
input: seriatim
references:
roster: ./campaign/party_roster.md
glossary: ./campaign/glossary.md
artifacts:
spells:
extract: dnd/spells
npcs:
extract: dnd/npcs
reference:
npc_registry: ./campaign/npcs.md
```
### CLI
Per-run override flag, repeatable:
```text
notarius run dnd-session --input session-014.json --reference roster=./alt_roster.md
```
CLI bindings override config bindings for the same slot name. The existing
pipeline-describe/config-validate commands (or their nearest equivalents)
should surface declared slots, descriptions, required flags, and current
bindings so users can discover what a pipeline accepts.
## Prompt Template Integration
Prompt templates are module-owned. Template rendering should expose:
- `{{ reference "roster" }}`: renders the content of the bound item;
- `{{ hasreference "glossary" }}`: predicate for conditional sections, so
optional slots can be included only when bound.
Rules:
- Referencing an **undeclared** slot from a template is a module bug and
should fail at prompt registration/build time (or earliest feasible point),
not silently at render time.
- Referencing a declared but unbound **optional** slot should render as
empty; templates should use `hasreference` to avoid dangling section headers.
- Rendering must be deterministic and independent of map iteration order.
- Prompt identity (registry hash) should be computed over the **template**,
not the rendered prompt. Reference digests are recorded separately in the
manifest, so a reference edit is visible as a reference change, not a prompt
change.
Note: reference content is repeated in every per-chunk prompt. Diagnostics
should record per-slot token or byte counts so reference cost is observable.
Per-slot inclusion policies (e.g., roster in every chunk, glossary on demand)
are explicitly out of scope until cost data justifies them.
## Provenance
The run manifest must record, for every bound slot:
- slot name;
- origin (path);
- content digest;
- media type;
- whether the binding came from config or CLI override.
Reference digests must participate in any idempotency/cache key alongside source
digests, prompt hashes, schema versions, model, and parameters. Two runs that
differ only in reference content must be distinguishable from the manifest
alone.
Diagnostics for a run should include the resolved binding set (with digests,
not necessarily full content) in the run directory, consistent with the
existing redacted-effective-config pattern.
## Path Resolution
- Config-relative paths resolve relative to the pipeline config file.
- CLI-relative paths resolve relative to the current working directory.
- Manifest records the normalized absolute path or a redacted/display path
according to existing diagnostics policy.
## Validation and Guardrails
### References are not evidence
The primary new failure mode: the model extracts facts from references rather
than from the source input. Example: the roster lists a PC's known spells, and
the model emits a `SpellCast` for a spell that was never cast in the session,
with a fabricated or misattributed source reference.
Defenses, in priority order:
1. **Structural.** `SourceRef` remains the only grounding mechanism and can
only reference source units. No contract change should make references
addressable as evidence.
2. **Prompt discipline.** Module templates should frame references explicitly as
reference material, e.g. "use the roster to resolve speakers to
characters; extract only events that occur in the transcript." This
guidance belongs in the module prompt guidelines, not framework code.
3. **Validator support.** The source-reference validator (or a sibling
deterministic validator) should support checking that referenced source
text plausibly relates to the extracted fact (e.g., spell name or a close
variant appears in or near the referenced range). Severity should be
`warn`, not `fail`, given paraphrase and nickname casting.
4. **Regression fixtures.** Golden-file tests must include a fixture in which
the bound roster mentions a spell that is never cast in the transcript,
asserting no artifact record is produced for it. This regression is likely
to be reintroduced by future prompt edits; the fixture is the guard.
### Size and sanity guardrails
- Fail fast, before any LLM call, if bound references plus template plus largest
chunk exceeds the configured model context budget, with an error that names
the offending slot(s) and sizes.
- Empty bound files should produce a warning (probable user error).
- MVP accepts text content only (`utf-8`); other media
types should be rejected with a clear error.
## Out of Scope (MVP)
- Non-file reference producers (prior-run artifacts, derived summaries, entity
registries). The `ReferenceItem` shape should permit them later.
- Per-chunk or per-slot inclusion policies and context budgeting beyond the
fail-fast guardrail.
- Structured/parsed references (e.g., typed roster schemas). References are opaque
text handed to prompts.
- Reference caching or preprocessing (summarization, embedding, retrieval).
- Making reference addressable as evidence, in any form.
## Checkpoint Sequencing
Each checkpoint should leave the repository compiling, with targeted tests
covering newly introduced contracts or behavior.
1. **Contracts and resolution.** Add `ReferenceSlot`, `ReferenceItem`, and the
extractor `ReferenceSlots()` method (empty default for existing extractors).
Implement config parsing for pipeline- and lane-level bindings, CLI
override flag, layering, and load-time validation (unknown slot, missing
required slot, unreadable file, empty file warning). Unit tests for
resolution and error cases.
2. **Prompt rendering.** Add `reference`/`hasreference` template functions,
declaration-order rendering, undeclared-slot failure at registration, and
deterministic-render tests (byte-identical output across runs).
3. **Provenance.** Record bindings (name, origin, digest, media type,
binding source) in the run manifest and diagnostics; include reference
digests in the cache/idempotency key if one exists. Tests: manifest
round-trip; two runs differing only in reference content produce differing
manifests.
4. **Guardrails and validation.** Context-window fail-fast check;
relatedness `warn` validator (or extension of the source-reference
validator); media-type rejection.
5. **First consumer.** Declare `roster` (optional) and `glossary` (optional)
slots on the D&D spells extractor; update its prompt template with
conditional reference sections and reference-material framing; add golden
fixtures with and without references bound, including the
roster-mentions-uncast-spell fixture. This checkpoint is the acceptance
test for the feature: spell extraction quality with a roster bound should
visibly improve speaker-to-character attribution in fixtures.
## Open Design Questions
The implementing agent should resolve these against existing code and record
decisions in the implementation plan:
- Should slot names be namespaced per lane in config and CLI (e.g.,
`spells.roster=...`) or flat with lane-level config as the only
disambiguator? (Recommended default: flat names; lane-level config for
overrides; revisit if two extractors in one pipeline want the same slot
name with different content.)
- Where does binding resolution live relative to the existing config and
pipeline packages? It must run at load time, alongside existing pipeline
validation.
- Does the existing prompt registry hash templates or rendered prompts? If
rendered, this feature requires moving to template hashing as described in
Provenance.
- Should CLI overrides be permitted to bind slots that config leaves unbound
(yes, presumably), and to *unbind* a config-bound optional slot (e.g.,
`--reference roster=` to clear)? Decide and test both directions.
## Documentation Tasks
Once implemented, move contracts out of this roadmap into canonical docs:
- `docs/cli.md`: `--reference` flag syntax, layering, and examples;
- `docs/config.md`: pipeline- and lane-level `references` blocks;
- `docs/internal/`: slot/item contracts, resolution flow, evidence
exclusion rule, and template function reference for module authors;
- module-author guidance: how to declare slots, write conditional reference
sections, and frame reference material in prompts;
- `examples/`: a maintained example pipeline with a roster and glossary
bound, plus matching fixture files.
```