Files
notarius/docs/roadmap/prompts.md

464 lines
22 KiB
Markdown

# D&D Prompt Asset Cleanup
## Purpose
This roadmap defines the target state for the embedded D&D prompt and private
schema assets. The work is intended to make prompts easier to inspect, compare,
and revise; eliminate unnecessary repetition; preserve useful prompt-cache
boundaries; and establish conventions that future D&D modules can follow.
The central design principle is that each instruction should have one clear
owner. Shared behavior belongs in a shared prompt asset, module-specific
behavior belongs in the module's instruction asset, variable grounding belongs
in a dedicated input asset, and purely structural requirements belong in the
JSON schema.
## Current State
The root `assets` package correctly centralizes embedded prompt and schema
content without owning D&D business logic. The D&D prompt assets nevertheless
retain several conventions inherited from their former module-local layouts:
- prompt definitions use prompt-ID-derived filenames rather than a consistent
local filename;
- most prompt directories divide fixed module instructions between `task.md`
and `instructions.md`, although that division generally does not represent a
meaningful cache, trust, or input boundary;
- module-specific instructions repeat some rules already supplied by shared D&D
messages or enforced by their response schemas;
- one generic transcript fragment currently describes full-session input,
chunk-scoped input, and selected normalization windows even though those
materials have different completeness and inference boundaries;
- variable context assets use several naming styles, including generic names
such as `catalog.md` and `grounding.md`;
- the private scene and spell response-schema identifiers do not follow the
`_llm`/`.llm` convention used by the other private D&D schemas.
These are internal authoring and maintenance inconsistencies. The durable
artifact contracts and the separation between content ownership in `assets`
and behavior ownership in `internal/modules` remain sound.
## Goals
The completed cleanup should:
- give every D&D prompt directory the same predictable entry-point names;
- consolidate fixed module-specific instructions into one message unless a
distinct message boundary has a concrete purpose;
- remove module-specific restatements of rules already owned by shared assets;
- retain separate messages for shared policy, stable references, variable
transcript content, and other inputs where the boundary communicates trust,
provenance, reuse, or cache behavior;
- distinguish full-session transcripts, extraction chunks, and normalization
evidence windows through separate shared assets with explicit scope rules;
- keep all shared D&D filenames visibly identifiable in prompt definitions;
- make variable context filenames self-describing;
- normalize private LLM response-schema identities without changing durable
output schemas;
- document the conventions future prompt authors must follow; and
- preserve the rule that the root `assets` Go package contains no business
logic.
## Decisions
### Canonical prompt directory layout
Every D&D prompt directory should use the following local layout:
```text
prompts/
prompt.yaml
instructions.md
[optional variable or context fragments]
```
The prompt definition filename is always `prompt.yaml`. The prompt's `id` field,
not its filename, remains the canonical global identifier. This is safe because
each module registers a distinct asset subtree with PromptKit; the local
`prompt.yaml` filenames do not share a virtual directory.
Use the `.yaml` extension consistently. It matches the existing embedded prompt
and profile assets.
For prompts that currently have both `task.md` and `instructions.md`, combine
their non-duplicative content into `instructions.md`. For prompts that currently
have only `task.md`, rename that asset to `instructions.md`. Do not preserve an
otherwise artificial two-message split merely to distinguish a short task
statement from its rules.
### Retain the `common-dnd-` prefix
Shared D&D assets must retain the `common-dnd-` filename prefix, including:
- `common-dnd-system.md`
- `common-dnd-identity.md`
- `common-dnd-references.md`
- `common-dnd-transcript-full.md`
- `common-dnd-transcript-chunk.md`
- `common-dnd-transcript-windows.md`
- `common-dnd-extraction-evidence.md`
- `common-dnd-npcs.md`
- `common-dnd-entity-reconciliation.md`
The prefix is a deliberate visual provenance marker. When reviewing a prompt
definition, an author should be able to distinguish shared messages from
module-owned messages immediately. Because asset references are explicit and
machine-resolved, retaining the prefix has no meaningful mechanical cost.
### One owner per instruction
Prompt prose should follow this ownership hierarchy:
1. Universal D&D prompting behavior belongs in a shared D&D asset.
2. Rules shared by a prompt family belong in the narrowest applicable shared
asset.
3. Semantics of a variable input belong in the fragment that renders that
input.
4. Lane-specific extraction, classification, or normalization policy belongs in
that lane's `instructions.md`.
5. Object shape, required fields, enum values, and rejection of unknown fields
belong in the JSON schema unless prose is needed to explain their domain
meaning.
Module-specific prompt assets should omit instructions already provided by the
shared prompt set. In particular, they should not routinely repeat generic rules
about transcript-only evidence, references not being evidence, omission of
`source_id`, narrow transcript ranges, schema-only output, or explanatory text.
This cleanup should begin with the minimal instruction set that is both correct
and non-repetitive. Repetition may be reintroduced only when evaluation produces
repeatable observational evidence that a target model needs reinforcement. Any
such repetition should be narrow, placed close to the rule's canonical owner,
and documented as a deliberate model-reliability tradeoff rather than allowed to
become accidental drift.
This rule does not authorize removal of domain semantics. Similar wording is not
duplication when it defines a distinct lane contract, such as event precedence,
holder transitions, scene-kind definitions, identity inclusion criteria, or the
relationship between a combat actor and an extracted event.
### System-prompt eligibility
`common-dnd-system.md` should own every fixed instruction that genuinely applies
to every maintained D&D LLM prompt. Consolidating those universal rules into the
system message gives them one authoritative owner, applies the appropriate
instruction priority, and preserves one byte-identical initial message across
chunking, extraction, and normalization prompts.
An instruction belongs in the shared system message only when it is:
- applicable to scene chunking, every extraction lane, and both entity
normalizers;
- independent of prompt-specific or variable inputs;
- static across calls; and
- appropriate to treat as a high-priority behavioral constraint.
The system message should therefore own the D&D transcript-processing role,
reliance only on supplied inputs, treatment of transcription uncertainty, and
the requirement to return exactly one schema-conforming JSON object without
explanatory prose.
Rules shared by only a subset of prompts must remain in the narrowest applicable
shared user-message fragment. Transcript-evidence and citation rules, identity
rules, campaign-reference semantics, registry grounding, and reconciliation
policy are not universal system instructions. No rendered transcript,
reference, registry, catalog, candidate, or other variable material belongs in
the system message.
### Shared prompt ownership
The shared D&D files should have clear, non-overlapping responsibilities:
| Asset | Canonical responsibility |
| --- | --- |
| `common-dnd-system.md` | D&D domain role, reliance on supplied inputs, treatment of transcription uncertainty, and the universal requirement to emit exactly one schema-conforming JSON object without commentary. |
| `common-dnd-identity.md` | Rules for resolving supported in-world identities and the distinction between disambiguation and participation evidence. |
| `common-dnd-references.md` | Rendering and interpretation of campaign references, including that references provide context rather than establishing transcript events. |
| `common-dnd-transcript-full.md` | Rendering the complete ordered session transcript and identifying it as complete session input. |
| `common-dnd-transcript-chunk.md` | Rendering one extraction chunk, limiting the task to that chunk, and explaining that its unit IDs retain their source-wide meaning. |
| `common-dnd-transcript-windows.md` | Rendering selected reconciliation evidence windows and explaining that they may be incomplete, non-contiguous, or overlapping. |
| `common-dnd-extraction-evidence.md` | Transcript-evidence policy, source-range construction, omission of `source_id`, and preference for narrow supporting ranges. |
| `common-dnd-npcs.md` | Rendering and interpreting an upstream NPC registry as grounding rather than evidence. |
| `common-dnd-entity-reconciliation.md` | Safe duplicate-group proposal mechanics shared by entity normalizers. |
As part of normalization, universal output-format prose should move to
`common-dnd-system.md`, while evidence-specific prose should remain in
`common-dnd-extraction-evidence.md`. The system message should refer generally
to supplied inputs rather than only transcripts and references, because
normalizers and grounded extractors receive additional input types.
### Transcript-scope ownership
The current `common-dnd-transcript.md` should be replaced rather than retained
as an alias. All three replacements may render the same prompt input named
`transcript`, but each must accurately describe the material supplied by its
caller:
- `common-dnd-transcript-full.md` is used by scene chunking. It identifies the
material as the complete ordered session and permits whole-session coverage
decisions without implying that the entire session is one scene.
- `common-dnd-transcript-chunk.md` is used by every extraction lane. It states
that only the current chunk is in scope, that source-wide unit IDs remain
authoritative, and that the model must not infer or report events outside the
supplied chunk.
- `common-dnd-transcript-windows.md` is used by NPC and location normalization.
It states that the material consists of selected evidence windows that may be
incomplete, non-contiguous, or overlapping. The model may use those windows
to evaluate candidate identity but must not treat absence outside them as
evidence.
The scope fragments should describe input completeness and permissible use;
they should not absorb scene policy, extraction evidence mechanics, or entity
reconciliation rules owned by other assets. Future prompt families must select
the fragment matching the material actually passed by code rather than choosing
one based merely on an input variable named `transcript`.
### Module-specific ownership
After shared duplication is removed, each module's `instructions.md` should
retain only its substantive lane policy:
| Prompt | Module-owned policy |
| --- | --- |
| Scene chunking | Scene-boundary rules, ordered coverage, and valid treatment of gaps or overlaps. |
| Scene descriptions | `kind` classification definitions and precedence, plus title and summary semantics. |
| NPC extraction | Inclusion and exclusion rules for individually identifiable NPCs. |
| NPC interactions | Interaction kinds, precedence, occurrence splitting, and NPC-specific event semantics. |
| Location extraction | Inclusion and exclusion rules for identifiable in-world places. |
| Location occurrences | Occurrence kinds, precedence, splitting, and location-specific event semantics. |
| Spell extraction | What constitutes a cast, catalog-name policy, and caster attribution. |
| Combat turns | Combat event boundaries, event kinds, actor semantics, and combat-specific inclusion rules. |
| Item events | Event kinds, possession transitions, holder semantics, quantities, and currency treatment. |
| Enemy events | Enemy-event kinds, outcome semantics, encounter identity, and combat-grounding rules. |
| NPC normalization | NPC-specific duplicate criteria and canonical-name selection. |
| Location normalization | Location-specific duplicate criteria and canonical-name selection. |
Schema field lists should not be copied into prose solely to restate the schema.
Prose may still explain how a field is interpreted or how two fields relate when
the schema cannot express that semantic rule.
### Message boundaries and cache ordering
Combining `task.md` and `instructions.md` must not lead to indiscriminate message
consolidation. Separate messages remain appropriate when they preserve a shared
byte-identical prefix, distinguish fixed instructions from variable material,
or identify the provenance and trust level of an input.
Extraction prompts should retain this general order:
```text
system
identity
campaign references
chunk transcript
evidence policy
optional grounding or variable context
module instructions
```
The scene chunker should retain this general order:
```text
system
campaign references
module instructions
full transcript
```
Entity normalizers should retain this general order:
```text
system
module instructions
shared entity-reconciliation policy
candidates
transcript windows
```
Existing cache controls should remain attached to the same semantic boundaries
unless PromptKit's contract requires a mechanical adjustment. The cleanup must
not move variable transcript or candidate content into an earlier stable prefix
merely to reduce message count.
### Variable context asset names
Optional non-shared fragments should describe the input they render rather than
use generic names. Normalize the current assets as follows where applicable:
- `catalog.md` to `spell-catalog.md`;
- `grounding.md` to `combat-grounding.md`;
- `locations.md` to `location-registry.md`.
The name `candidates.md` may remain in normalizer-local directories because its
meaning is clear in that context. Future context assets should use similarly
specific nouns and should remain separate from fixed `instructions.md` content.
### Private response-schema identity
Private schemas used only to constrain LLM responses should consistently use
the `_llm` filename/key/name convention and the `.llm` schema-ID suffix already
used by most D&D modules.
Normalize the scene schema to:
- file: `dnd_scenes_llm.v1.json`;
- schema ID: `notarius.dnd.scenes.llm`;
- registry key: `dnd_scenes_llm`;
- response-schema name: `notarius_dnd_scenes_llm_v1`.
Normalize the spell schema metadata to:
- file: `dnd_spells_llm.v1.json` (already named correctly);
- schema ID: `notarius.dnd.spells.llm`;
- registry key: `dnd_spells_llm`;
- response-schema name: `notarius_dnd_spells_llm_v1`.
Private schema filenames must remain globally unique. They should not be renamed
to a generic local name because the schema registry exposes a shared namespace.
Durable artifact schema identifiers and shapes are out of scope and must not
change.
### Explicit prompt definitions remain explicit
The repeated `version`, `default_profile`, input, message, output, and schema
declarations in `prompt.yaml` files should remain explicit. They are declarative
configuration, not prompt-prose duplication. Introducing a generator, template
layer, or Go builder for these small definitions would make inspection harder
and create another synchronization boundary.
Similarly, module manifests should continue to declare the assets their prompt
uses. Shared asset selection is module behavior and does not belong in the root
`assets` package.
## Compatibility and Cache Effects
This is an intentional internal prompt revision. Renaming prompt definition
files, consolidating messages, editing shared prose, and normalizing private
schema identifiers will change prompt and component fingerprints. Existing
checkpoints may no longer be reusable, and backend prompt caches will experience
a one-time cold start. Those fingerprint, checkpoint, and cache effects are
accepted consequences of reaching the target prompt layout.
The following contracts must remain stable:
- public prompt IDs;
- module IDs and configuration fields;
- default profile selection;
- required and optional prompt inputs;
- durable artifact schema IDs and artifact shapes;
- extraction and normalization semantics, except for removing unintended prompt
repetition; and
- pipeline dependency behavior.
No backward-compatible copies of old private prompt assets or schema aliases are
required while the application remains pre-release.
## Documentation and Maintenance Rules
The implemented convention should be documented in the canonical internal D&D
and LLM documentation rather than in an assets-package README that could drift.
The documentation should explain:
- the canonical `prompt.yaml` and `instructions.md` names;
- why shared assets retain the `common-dnd-` prefix;
- the instruction-ownership hierarchy;
- the eligibility test for moving a rule into `common-dnd-system.md`;
- the full, chunk, and evidence-window transcript contracts and which prompt
families select each one;
- when a separate message boundary is warranted;
- the cache-oriented message ordering;
- the prohibition on business logic in `assets`; and
- the evidence threshold for deliberately repeating an instruction.
The Go file in the root `assets` package should remain a minimal embedding and
filesystem-access layer. Prompt composition, schema selection, module policy,
and validation logic must remain in their existing owning packages.
`docs/internal/dnd.md` should be the canonical owner of these D&D prompt
authoring conventions and the concrete message-order patterns. Its existing
prompt-construction section should be revised when the assets change rather
than creating an `assets` README or duplicating the rules in contributor
orientation. `docs/internal/llm.md` should receive only any short update needed
to describe provider-neutral PromptKit preparation or cache-control mechanics;
it should link to the D&D guide instead of repeating domain-specific asset names
or ordering.
No new ADR is warranted for this cleanup. ADR-0011 already records the
architecturally significant decisions: physical centralization, content-only
package ownership, and retention of prompt semantics and ordering by modules.
The filename, instruction-ownership, transcript-scope, and system-message rules
are internal authoring conventions that should be allowed to evolve with model
evaluation. They do not create a new dependency boundary or reverse an accepted
architectural decision. If a later change promotes these conventions into a
cross-domain runtime-enforced prompt composition contract, that broader
decision should receive its own ADR at that time; ADR-0011 must not be edited to
retrofit the new rationale.
## Validation and Evaluation Policy
Automated validation should establish that the refactor preserves executable
contracts without freezing prompt prose. It should cover:
- preparation of every embedded prompt through PromptKit;
- prompt input declarations and required-input behavior;
- private response-schema registration and compatibility with the owning Go
response types;
- expected cache controls and meaningful stable-prefix boundaries;
- selection of the transcript-scope fragment appropriate to each prompt
family, verified through prompt preparation and owning-module behavior rather
than exact prose snapshots;
- manifest references after asset renames;
- absence of references to retired filenames and private schema identifiers;
and
- the full repository test suite.
Tests must not assert exact prompt wording, exact message counts, shared-prefix
lengths, or asset inventories solely as change detectors. Such tests would make
intentional prompt iteration unnecessarily expensive and conflict with the
project's testing policy.
Model quality should be evaluated by running representative transcripts through
all affected lanes, with particular attention to cost-effective smaller models.
Review should compare extraction correctness, unsupported claims, missing
events, source-range quality, and schema reliability. The initial evaluation
should use the minimal non-repetitive prompts. A repeated instruction should be
restored only if failures recur and there is a credible causal reason that
reinforcement will help.
## Non-Goals
This roadmap does not include:
- changes to durable D&D artifact schemas;
- new extraction lanes or pipeline behavior;
- changes to profile selection or model settings;
- a prompt-definition generator or templating framework;
- moving business logic into `assets`;
- collapsing all prompt material into one message;
- placing family-specific rules or rendered variable content in the system
message;
- preserving old prompt fingerprints or warm backend cache entries; or
- introducing brittle prompt snapshots or change-detector tests.
## Target End State
Every D&D prompt is discoverable through a local `prompt.yaml`, has one fixed
module-owned `instructions.md`, and uses separately named context fragments only
where inputs or cache boundaries justify them. Shared D&D assets remain visibly
marked with `common-dnd-` and have documented, non-overlapping responsibilities.
Module instructions contain no routine repetition of shared or schema-enforced
rules, while retaining all lane-specific domain semantics.
The shared system message contains all and only universal, static D&D behavior.
Scene chunking renders the explicitly full transcript fragment, extraction
lanes render the explicitly chunk-scoped fragment, and entity normalizers render
the explicitly incomplete evidence-window fragment. No prompt presents one
scope as another merely because all three inputs use the name `transcript`.
Private scene and spell response schemas follow the established LLM-only naming
convention, durable contracts are unchanged, all manifests and PromptKit
definitions prepare successfully, and human evaluation begins from a minimal
non-repetitive prompt set. Future modules can follow the same conventions
without copying an existing extractor and inheriting accidental prose or naming
drift.