14 KiB
Feature Roadmap Proposal: Extraction Reference
Status
This document captures proposed design and implementation sequencing for the extraction-reference feature in Notarius. It describes planned work, not implemented behavior. Go snippets are conceptual sketches; the implementing agent should adapt names and shapes to the existing contracts, package boundaries, and conventions in this repository.
Goal
Extraction quality improves significantly when the LLM receives reference material alongside the source input. For the initial D&D spell extractor, useful reference material includes a party roster (mapping players to player characters), a player list, and a campaign glossary.
Notarius should support passing this material to extractors as named reference items without introducing any domain-specific concepts into core or framework packages. The framework should know only that:
- extractors declare named reference slots they accept;
- pipeline config and CLI flags bind content (initially files) to those slots;
- bound content is rendered into module-owned prompt templates;
- bound content is digested and recorded as run provenance.
Only extract modules should know what a "roster" or "glossary" means. All domain semantics live in module-owned slot declarations and prompt templates.
Definitions
- Reference slot: a named, typed-by-convention input declared by an
extractor, with a human-readable description and a required/optional flag.
Example: extractor
dnd/spellsdeclares an optional slot namedroster. - Reference item: resolved content bound to a slot for a given run: name, content bytes, media type, content digest, and origin (initially a file path).
- Reference binding: the association of a slot name to a content source, defined in pipeline config and overridable per run via CLI.
Architectural Principles
- Reference is opaque to the framework. Core and framework packages must not interpret reference content or recognize domain slot names.
- Reference is an input. Anything that changes extraction output must be digested into the run manifest and participate in any cache key.
- A Reference is not evidence.
SourceRefvalues must only ever reference source units. Reference items must not receive unit IDs and must not be addressable by source references. - Slots are declared, not ad hoc. Binding an undeclared slot name, or omitting a required slot, should fail at config-load time, before any LLM call.
- Optional slots degrade gracefully. Prompt templates should render cleanly whether or not an optional slot is bound.
- Determinism. Identical input, config, prompts, and reference bytes should produce byte-identical rendered prompts. Reference slots should render in a stable, documented order (declaration order).
Proposed Contracts
Slot declaration (extractor contract extension)
Extractors should declare the reference slots they accept:
type ReferenceSlot struct {
Name string
Description string
Required bool
// MVP can leave these empty/defaulted, but having the fields now
// makes validation and future docs easier.
AcceptedMediaTypes []string
Multiple bool
MaxBytes int64
}
The extractor interface should gain a method such as:
ReferenceSlots() []ReferenceSlot
Extractors with no reference needs return an empty slice. Existing extractors should require no other changes.
Resolved reference item
type ReferenceItem struct {
SlotName string
MediaType string
Content []byte
Digest string
Origin ReferenceOrigin
SizeBytes int64
TokenEstimate int
}
type ReferenceOrigin struct {
Type string // "file" for MVP
URI string // path or future artifact URI
}
type ReferenceSet struct {
// Stable declaration order, then stable binding order within a slot.
Slots []ResolvedReferenceSlot
}
type ResolvedReferenceSlot struct {
Name string
Items []ReferenceItem
}
ReferenceItem is a resolved-content type, not a file path. The only MVP
producer is "read this file," but the shape should permit future producers
(prior-run artifacts, derived summaries, entity registries) without contract
changes.
Binding resolution
A resolver should, at config-load time:
- Collect declared slots from every extractor selected by the active
pipeline (respecting lane selection, e.g.
--only). - Collect bindings from pipeline config (pipeline-level and lane-level) and CLI overrides, applying the standard layering: config file, then CLI.
- Fail with a clear error if a required slot is unbound, or if a binding references a slot no extractor within the selected pipeline declares. Errors should name the pipeline, lane, slot, and the slot description.
- Read, digest, and materialize each bound source into a
ReferenceItem. - Enforce size guardrails (see Validation and Guardrails).
Configuration and CLI
Pipeline config
Reference bindings should live in pipeline config, because the initial use cases (roster, glossary) are campaign-invariant rather than run-variant. Bindings should be supported at two levels:
- pipeline level: shared by all artifact lanes;
- lane level: additions or overrides for a single lane.
Illustrative shape (adapt to the existing config format):
pipelines:
dnd-session:
input: seriatim
references:
roster: ./campaign/party_roster.md
glossary: ./campaign/glossary.md
artifacts:
spells:
extract: dnd/spells
npcs:
extract: dnd/npcs
reference:
npc_registry: ./campaign/npcs.md
CLI
Per-run override flag, repeatable:
notarius run dnd-session --input session-014.json --reference roster=./alt_roster.md
CLI bindings override config bindings for the same slot name. The existing pipeline-describe/config-validate commands (or their nearest equivalents) should surface declared slots, descriptions, required flags, and current bindings so users can discover what a pipeline accepts.
Prompt Template Integration
Prompt templates are module-owned. Template rendering should expose:
{{ reference "roster" }}: renders the content of the bound item;{{ hasreference "glossary" }}: predicate for conditional sections, so optional slots can be included only when bound.
Rules:
- Referencing an undeclared slot from a template is a module bug and should fail at prompt registration/build time (or earliest feasible point), not silently at render time.
- Referencing a declared but unbound optional slot should render as
empty; templates should use
hasreferenceto avoid dangling section headers. - Rendering must be deterministic and independent of map iteration order.
- Prompt identity (registry hash) should be computed over the template, not the rendered prompt. Reference digests are recorded separately in the manifest, so a reference edit is visible as a reference change, not a prompt change.
Note: reference content is repeated in every per-chunk prompt. Diagnostics should record per-slot token or byte counts so reference cost is observable. Per-slot inclusion policies (e.g., roster in every chunk, glossary on demand) are explicitly out of scope until cost data justifies them.
Provenance
The run manifest must record, for every bound slot:
- slot name;
- origin (path);
- content digest;
- media type;
- whether the binding came from config or CLI override.
Reference digests must participate in any idempotency/cache key alongside source digests, prompt hashes, schema versions, model, and parameters. Two runs that differ only in reference content must be distinguishable from the manifest alone.
Diagnostics for a run should include the resolved binding set (with digests, not necessarily full content) in the run directory, consistent with the existing redacted-effective-config pattern.
Path Resolution
- Config-relative paths resolve relative to the pipeline config file.
- CLI-relative paths resolve relative to the current working directory.
- Manifest records the normalized absolute path or a redacted/display path according to existing diagnostics policy.
Validation and Guardrails
References are not evidence
The primary new failure mode: the model extracts facts from references rather
than from the source input. Example: the roster lists a PC's known spells, and
the model emits a SpellCast for a spell that was never cast in the session,
with a fabricated or misattributed source reference.
Defenses, in priority order:
- Structural.
SourceRefremains the only grounding mechanism and can only reference source units. No contract change should make references addressable as evidence. - Prompt discipline. Module templates should frame references explicitly as reference material, e.g. "use the roster to resolve speakers to characters; extract only events that occur in the transcript." This guidance belongs in the module prompt guidelines, not framework code.
- Validator support. The source-reference validator (or a sibling
deterministic validator) should support checking that referenced source
text plausibly relates to the extracted fact (e.g., spell name or a close
variant appears in or near the referenced range). Severity should be
warn, notfail, given paraphrase and nickname casting. - Regression fixtures. Golden-file tests must include a fixture in which the bound roster mentions a spell that is never cast in the transcript, asserting no artifact record is produced for it. This regression is likely to be reintroduced by future prompt edits; the fixture is the guard.
Size and sanity guardrails
- Fail fast, before any LLM call, if bound references plus template plus largest chunk exceeds the configured model context budget, with an error that names the offending slot(s) and sizes.
- Empty bound files should produce a warning (probable user error).
- MVP accepts text content only (
utf-8); other media types should be rejected with a clear error.
Out of Scope (MVP)
- Non-file reference producers (prior-run artifacts, derived summaries, entity
registries). The
ReferenceItemshape should permit them later. - Per-chunk or per-slot inclusion policies and context budgeting beyond the fail-fast guardrail.
- Structured/parsed references (e.g., typed roster schemas). References are opaque text handed to prompts.
- Reference caching or preprocessing (summarization, embedding, retrieval).
- Making reference addressable as evidence, in any form.
Checkpoint Sequencing
Each checkpoint should leave the repository compiling, with targeted tests covering newly introduced contracts or behavior.
- Contracts and resolution. Add
ReferenceSlot,ReferenceItem, and the extractorReferenceSlots()method (empty default for existing extractors). Implement config parsing for pipeline- and lane-level bindings, CLI override flag, layering, and load-time validation (unknown slot, missing required slot, unreadable file, empty file warning). Unit tests for resolution and error cases. - Prompt rendering. Add
reference/hasreferencetemplate functions, declaration-order rendering, undeclared-slot failure at registration, and deterministic-render tests (byte-identical output across runs). - Provenance. Record bindings (name, origin, digest, media type, binding source) in the run manifest and diagnostics; include reference digests in the cache/idempotency key if one exists. Tests: manifest round-trip; two runs differing only in reference content produce differing manifests.
- Guardrails and validation. Context-window fail-fast check;
relatedness
warnvalidator (or extension of the source-reference validator); media-type rejection. - First consumer. Declare
roster(optional) andglossary(optional) slots on the D&D spells extractor; update its prompt template with conditional reference sections and reference-material framing; add golden fixtures with and without references bound, including the roster-mentions-uncast-spell fixture. This checkpoint is the acceptance test for the feature: spell extraction quality with a roster bound should visibly improve speaker-to-character attribution in fixtures.
Open Design Questions
The implementing agent should resolve these against existing code and record decisions in the implementation plan:
- Should slot names be namespaced per lane in config and CLI (e.g.,
spells.roster=...) or flat with lane-level config as the only disambiguator? (Recommended default: flat names; lane-level config for overrides; revisit if two extractors in one pipeline want the same slot name with different content.) - Where does binding resolution live relative to the existing config and pipeline packages? It must run at load time, alongside existing pipeline validation.
- Does the existing prompt registry hash templates or rendered prompts? If rendered, this feature requires moving to template hashing as described in Provenance.
- Should CLI overrides be permitted to bind slots that config leaves unbound
(yes, presumably), and to unbind a config-bound optional slot (e.g.,
--reference roster=to clear)? Decide and test both directions.
Documentation Tasks
Once implemented, move contracts out of this roadmap into canonical docs:
docs/cli.md:--referenceflag syntax, layering, and examples;docs/config.md: pipeline- and lane-levelreferencesblocks;docs/internal/: slot/item contracts, resolution flow, evidence exclusion rule, and template function reference for module authors;- module-author guidance: how to declare slots, write conditional reference sections, and frame reference material in prompts;
examples/: a maintained example pipeline with a roster and glossary bound, plus matching fixture files.