Files
notarius/docs/roadmap/references.md

14 KiB

Feature Roadmap Proposal: Extraction Reference

Status

This document captures proposed design and implementation sequencing for the extraction-reference feature in Notarius. It describes planned work, not implemented behavior. Go snippets are conceptual sketches; the implementing agent should adapt names and shapes to the existing contracts, package boundaries, and conventions in this repository.

Goal

Extraction quality improves significantly when the LLM receives reference material alongside the source input. For the initial D&D spell extractor, useful reference material includes a party roster (mapping players to player characters), a player list, and a campaign glossary.

Notarius should support passing this material to extractors as named reference items without introducing any domain-specific concepts into core or framework packages. The framework should know only that:

  • extractors declare named reference slots they accept;
  • pipeline config and CLI flags bind content (initially files) to those slots;
  • bound content is rendered into module-owned prompt templates;
  • bound content is digested and recorded as run provenance.

Only extract modules should know what a "roster" or "glossary" means. All domain semantics live in module-owned slot declarations and prompt templates.

Definitions

  • Reference slot: a named, typed-by-convention input declared by an extractor, with a human-readable description and a required/optional flag. Example: extractor dnd/spells declares an optional slot named roster.
  • Reference item: resolved content bound to a slot for a given run: name, content bytes, media type, content digest, and origin (initially a file path).
  • Reference binding: the association of a slot name to a content source, defined in pipeline config and overridable per run via CLI.

Architectural Principles

  • Reference is opaque to the framework. Core and framework packages must not interpret reference content or recognize domain slot names.
  • Reference is an input. Anything that changes extraction output must be digested into the run manifest and participate in any cache key.
  • A Reference is not evidence. SourceRef values must only ever reference source units. Reference items must not receive unit IDs and must not be addressable by source references.
  • Slots are declared, not ad hoc. Binding an undeclared slot name, or omitting a required slot, should fail at config-load time, before any LLM call.
  • Optional slots degrade gracefully. Prompt templates should render cleanly whether or not an optional slot is bound.
  • Determinism. Identical input, config, prompts, and reference bytes should produce byte-identical rendered prompts. Reference slots should render in a stable, documented order (declaration order).

Proposed Contracts

Slot declaration (extractor contract extension)

Extractors should declare the reference slots they accept:

type ReferenceSlot struct {
    Name        string
    Description string
    Required    bool

    // MVP can leave these empty/defaulted, but having the fields now
    // makes validation and future docs easier.
    AcceptedMediaTypes []string
    Multiple           bool
    MaxBytes           int64
}

The extractor interface should gain a method such as:

ReferenceSlots() []ReferenceSlot

Extractors with no reference needs return an empty slice. Existing extractors should require no other changes.

Resolved reference item

type ReferenceItem struct {
    SlotName  string
    MediaType string
    Content   []byte
    Digest    string
    Origin    ReferenceOrigin

    SizeBytes     int64
    TokenEstimate int
}

type ReferenceOrigin struct {
    Type string // "file" for MVP
    URI  string // path or future artifact URI
}

type ReferenceSet struct {
    // Stable declaration order, then stable binding order within a slot.
    Slots []ResolvedReferenceSlot
}

type ResolvedReferenceSlot struct {
    Name  string
    Items []ReferenceItem
}

ReferenceItem is a resolved-content type, not a file path. The only MVP producer is "read this file," but the shape should permit future producers (prior-run artifacts, derived summaries, entity registries) without contract changes.

Binding resolution

A resolver should, at config-load time:

  1. Collect declared slots from every extractor selected by the active pipeline (respecting lane selection, e.g. --only).
  2. Collect bindings from pipeline config (pipeline-level and lane-level) and CLI overrides, applying the standard layering: config file, then CLI.
  3. Fail with a clear error if a required slot is unbound, or if a binding references a slot no extractor within the selected pipeline declares. Errors should name the pipeline, lane, slot, and the slot description.
  4. Read, digest, and materialize each bound source into a ReferenceItem.
  5. Enforce size guardrails (see Validation and Guardrails).

Configuration and CLI

Pipeline config

Reference bindings should live in pipeline config, because the initial use cases (roster, glossary) are campaign-invariant rather than run-variant. Bindings should be supported at two levels:

  • pipeline level: shared by all artifact lanes;
  • lane level: additions or overrides for a single lane.

Illustrative shape (adapt to the existing config format):

pipelines:
  dnd-session:
    input: seriatim
    references:
      roster: ./campaign/party_roster.md
      glossary: ./campaign/glossary.md
    artifacts:
      spells:
        extract: dnd/spells
      npcs:
        extract: dnd/npcs
        reference:
          npc_registry: ./campaign/npcs.md

CLI

Per-run override flag, repeatable:

notarius run dnd-session --input session-014.json --reference roster=./alt_roster.md

CLI bindings override config bindings for the same slot name. The existing pipeline-describe/config-validate commands (or their nearest equivalents) should surface declared slots, descriptions, required flags, and current bindings so users can discover what a pipeline accepts.

Prompt Template Integration

Prompt templates are module-owned. Template rendering should expose:

  • {{ reference "roster" }}: renders the content of the bound item;
  • {{ hasreference "glossary" }}: predicate for conditional sections, so optional slots can be included only when bound.

Rules:

  • Referencing an undeclared slot from a template is a module bug and should fail at prompt registration/build time (or earliest feasible point), not silently at render time.
  • Referencing a declared but unbound optional slot should render as empty; templates should use hasreference to avoid dangling section headers.
  • Rendering must be deterministic and independent of map iteration order.
  • Prompt identity (registry hash) should be computed over the template, not the rendered prompt. Reference digests are recorded separately in the manifest, so a reference edit is visible as a reference change, not a prompt change.

Note: reference content is repeated in every per-chunk prompt. Diagnostics should record per-slot token or byte counts so reference cost is observable. Per-slot inclusion policies (e.g., roster in every chunk, glossary on demand) are explicitly out of scope until cost data justifies them.

Provenance

The run manifest must record, for every bound slot:

  • slot name;
  • origin (path);
  • content digest;
  • media type;
  • whether the binding came from config or CLI override.

Reference digests must participate in any idempotency/cache key alongside source digests, prompt hashes, schema versions, model, and parameters. Two runs that differ only in reference content must be distinguishable from the manifest alone.

Diagnostics for a run should include the resolved binding set (with digests, not necessarily full content) in the run directory, consistent with the existing redacted-effective-config pattern.

Path Resolution

  • Config-relative paths resolve relative to the pipeline config file.
  • CLI-relative paths resolve relative to the current working directory.
  • Manifest records the normalized absolute path or a redacted/display path according to existing diagnostics policy.

Validation and Guardrails

References are not evidence

The primary new failure mode: the model extracts facts from references rather than from the source input. Example: the roster lists a PC's known spells, and the model emits a SpellCast for a spell that was never cast in the session, with a fabricated or misattributed source reference.

Defenses, in priority order:

  1. Structural. SourceRef remains the only grounding mechanism and can only reference source units. No contract change should make references addressable as evidence.
  2. Prompt discipline. Module templates should frame references explicitly as reference material, e.g. "use the roster to resolve speakers to characters; extract only events that occur in the transcript." This guidance belongs in the module prompt guidelines, not framework code.
  3. Validator support. The source-reference validator (or a sibling deterministic validator) should support checking that referenced source text plausibly relates to the extracted fact (e.g., spell name or a close variant appears in or near the referenced range). Severity should be warn, not fail, given paraphrase and nickname casting.
  4. Regression fixtures. Golden-file tests must include a fixture in which the bound roster mentions a spell that is never cast in the transcript, asserting no artifact record is produced for it. This regression is likely to be reintroduced by future prompt edits; the fixture is the guard.

Size and sanity guardrails

  • Fail fast, before any LLM call, if bound references plus template plus largest chunk exceeds the configured model context budget, with an error that names the offending slot(s) and sizes.
  • Empty bound files should produce a warning (probable user error).
  • MVP accepts text content only (utf-8); other media types should be rejected with a clear error.

Out of Scope (MVP)

  • Non-file reference producers (prior-run artifacts, derived summaries, entity registries). The ReferenceItem shape should permit them later.
  • Per-chunk or per-slot inclusion policies and context budgeting beyond the fail-fast guardrail.
  • Structured/parsed references (e.g., typed roster schemas). References are opaque text handed to prompts.
  • Reference caching or preprocessing (summarization, embedding, retrieval).
  • Making reference addressable as evidence, in any form.

Checkpoint Sequencing

Each checkpoint should leave the repository compiling, with targeted tests covering newly introduced contracts or behavior.

  1. Contracts and resolution. Add ReferenceSlot, ReferenceItem, and the extractor ReferenceSlots() method (empty default for existing extractors). Implement config parsing for pipeline- and lane-level bindings, CLI override flag, layering, and load-time validation (unknown slot, missing required slot, unreadable file, empty file warning). Unit tests for resolution and error cases.
  2. Prompt rendering. Add reference/hasreference template functions, declaration-order rendering, undeclared-slot failure at registration, and deterministic-render tests (byte-identical output across runs).
  3. Provenance. Record bindings (name, origin, digest, media type, binding source) in the run manifest and diagnostics; include reference digests in the cache/idempotency key if one exists. Tests: manifest round-trip; two runs differing only in reference content produce differing manifests.
  4. Guardrails and validation. Context-window fail-fast check; relatedness warn validator (or extension of the source-reference validator); media-type rejection.
  5. First consumer. Declare roster (optional) and glossary (optional) slots on the D&D spells extractor; update its prompt template with conditional reference sections and reference-material framing; add golden fixtures with and without references bound, including the roster-mentions-uncast-spell fixture. This checkpoint is the acceptance test for the feature: spell extraction quality with a roster bound should visibly improve speaker-to-character attribution in fixtures.

Open Design Questions

The implementing agent should resolve these against existing code and record decisions in the implementation plan:

  • Should slot names be namespaced per lane in config and CLI (e.g., spells.roster=...) or flat with lane-level config as the only disambiguator? (Recommended default: flat names; lane-level config for overrides; revisit if two extractors in one pipeline want the same slot name with different content.)
  • Where does binding resolution live relative to the existing config and pipeline packages? It must run at load time, alongside existing pipeline validation.
  • Does the existing prompt registry hash templates or rendered prompts? If rendered, this feature requires moving to template hashing as described in Provenance.
  • Should CLI overrides be permitted to bind slots that config leaves unbound (yes, presumably), and to unbind a config-bound optional slot (e.g., --reference roster= to clear)? Decide and test both directions.

Documentation Tasks

Once implemented, move contracts out of this roadmap into canonical docs:

  • docs/cli.md: --reference flag syntax, layering, and examples;
  • docs/config.md: pipeline- and lane-level references blocks;
  • docs/internal/: slot/item contracts, resolution flow, evidence exclusion rule, and template function reference for module authors;
  • module-author guidance: how to declare slots, write conditional reference sections, and frame reference material in prompts;
  • examples/: a maintained example pipeline with a roster and glossary bound, plus matching fixture files.