Hardened D&D spell extraction prompts

This commit is contained in:
2026-07-20 20:25:23 -05:00
parent fa5076f5f1
commit c4c907d421
7 changed files with 130 additions and 32 deletions

View File

@@ -50,11 +50,21 @@ Each spell cast contains exactly these required fields:
- `source_refs`: transcript source references with extractor-assigned source
IDs and evidence unit ranges. It must contain at least one entry.
All four string fields must be non-empty. `caster` is the in-world caster, not
the transcript speaker. The `spell` value must resolve through the effective
SRD-plus-overlay catalog as either a canonical name or alias. Catalog
validation accepts aliases but does not rewrite them; unknown fields are
rejected.
All four string fields must be non-empty. `caster` is the canonical in-world
caster, not the human player, transcript speaker, or GM when the associated
character or creature can be identified. Player and party references may
disambiguate that identity, but do not independently establish that a cast
occurred. The `spell` value must resolve through the effective SRD-plus-overlay
catalog as either a canonical name or alias. Catalog validation accepts aliases
but does not rewrite them; unknown fields are rejected.
`effect` and `narrative_description` record what the transcript establishes
about the session as played. They do not correct the table from published D&D
rules or supplement the transcript with model knowledge. When the transcript
contains a nonstandard or disputed ruling, the artifact may preserve the
observed resolution and attribute relevant reasoning to the GM or table; it
must not present that reasoning as a universal game rule. The spell catalog is
name-recognition policy, not evidence for spell mechanics or outcomes.
## Source References
@@ -63,6 +73,17 @@ Each source reference contains exactly three required fields: `source_id`,
The unit IDs must be positive integers present in the input, and the start unit
must not appear after the end unit. Unknown fields are rejected.
For each cast, the complete `source_refs` collection identifies the transcript
evidence for every factual claim in `caster`, `spell`, `effect`, and
`narrative_description`. A cast declaration and a later resolution may be
cited with separate narrow ranges when intervening units are unrelated. A
reported target, roll, amount, condition, interruption, or outcome must be
supported by the cited units; otherwise the artifact describes only the
supported attempt or declaration. The deterministic validators establish that
ranges are structurally valid and that the spell name is related to cited
text. Semantic claim completeness is an extraction policy and remains subject
to evaluation rather than deterministic proof.
Reference slot keys and accepted file types are defined in
[Configuration](../config.md#implemented-production-modules). References are
supporting disambiguation material, not source evidence, and are not
@@ -145,3 +166,9 @@ included in manifest metadata. The `normalizer` metadata uses the same catalog
identity fields when that module is selected. Overlay origin, media type, byte
size, and raw digest are recorded separately in the manifest's reference
provenance; see the [JSON output contract](json-output.md#manifestjson).
The extractor's prompt hash, private response-schema hash, and effective catalog
digest also contribute independently scoped semantic checkpoint fingerprints.
Changing any of those prepared contracts intentionally produces a cold
checkpoint miss. Fingerprints contain only digests, never prompt, schema,
catalog, or reference content.

View File

@@ -158,11 +158,25 @@ reference into an immutable effective catalog and adds a generated
canonical-name-only JSON input to every structured completion request. Overlay
failures therefore stop construction before source parsing or an LLM call;
campaign references remain separate disambiguation inputs and never become
source evidence. Both the extractor and deterministic catalog validator expose
source evidence.
The prompt requires each cast's source references collectively to support all
reported facts, using separate narrow ranges when a declaration and resolution
are non-contiguous. It grounds caster names in in-world identities, using the
player and party references to disambiguate transcript speakers without
treating those references as event evidence. Effects describe the session as
played: model rules knowledge cannot supplement or correct the transcript, and
nonstandard adjudication is attributed to the GM or table rather than stated as
a universal rule. Structural source validation remains deterministic; semantic
claim completeness is enforced through extraction policy and evaluation.
Both the extractor and deterministic catalog validator expose
the effective base-plus-overlay semantic digest as scoped prepared-component
checkpoint identity. Raw overlay provenance independently covers file-byte
changes, while the semantic digest also invalidates reuse when the embedded
catalog or catalog composition changes. The separate
catalog or catalog composition changes. The extractor additionally fingerprints
its complete prompt assets and private response schema, so either semantic
contract changing invalidates previously recorded extraction checkpoints. The separate
`internal/modules/dnd/codec/spells` package
owns the durable schema and stable JSON representation for artifact kind
`dnd/spell-list`. The runner keeps the result typed through validators and later

View File

@@ -2,9 +2,27 @@ Source references must use integer source-unit IDs from the transcript. Provide
start_unit_id and end_unit_id for each source reference; the extractor assigns
source_id automatically.
For each artifact, its source references must collectively support every
factual claim in caster, spell, effect, and narrative_description. Include the
casting declaration or action and every later transcript unit needed to support
reported targets, saving throws, damage, healing, conditions, interruptions,
and outcomes. When supporting evidence is non-contiguous, provide multiple
narrow source references. Do not use one broad range merely to bridge unrelated
table conversation. If a detail is not supported by the cited transcript
units, omit that detail or describe only the supported attempt or declaration.
Return only D&D spell-cast artifacts. For each spell cast, identify the in-world
caster, spell name, effect, narrative description, and source references.
The caster must be the canonical in-world character or creature, not the human
player, transcript speaker, or GM. Use the player and party references together
with transcript context to map first-person player speech to the associated
player character and use the canonical character name from the references.
Likewise, attribute a spell narrated by the GM to the in-world creature that
casts it. Never return a player or GM name when the associated in-world caster
can be identified. If the caster cannot be resolved, use only the most specific
in-world identity supported by the transcript; do not invent a name.
Use the canonical spell-name catalog to select spell names. Do not return a
spell name absent from that catalog, even when it is suggested by general D&D
knowledge or reference material.
@@ -13,4 +31,11 @@ Use player, party, and glossary reference material only to clarify source text;
references are not source evidence for a spell cast. Do not return spells,
casters, or effects that are mentioned only in reference material.
Effects and narrative descriptions are session records, not rules summaries.
Report only mechanics, explanations, and outcomes established by the cited
transcript units. Preserve the table's observed resolution without silently
correcting it from general D&D knowledge. If the transcript gives a possibly
nonstandard rationale, use wording such as "the GM rules" or "the table
resolves" rather than asserting that rationale as a universal rule.
Return exactly one JSON object and no explanatory text.

View File

@@ -4,6 +4,13 @@ Extract only spell casts that are supported by the transcript. Do not infer
spells from general D&D knowledge or from table chatter that does not identify a
spell being cast.
Describe the session as it was played and adjudicated. The transcript is
authoritative for what happened in this session, even when a table ruling may
differ from published D&D rules. Do not correct the transcript or fill in
unstated mechanics from general D&D knowledge. When a ruling or mechanical
explanation matters, attribute it to the GM or table instead of presenting it
as a universal game rule.
Use the provided canonical spell-name catalog when naming each extracted spell.
Return the canonical catalog spelling exactly. The catalog is a recognition
aid; it does not establish that a spell was cast.
aid; it does not establish that a spell was cast or how the spell works.

View File

@@ -20,23 +20,28 @@
"properties": {
"caster": {
"type": "string",
"minLength": 1
"minLength": 1,
"description": "Canonical in-world character or creature that casts the spell, never the human player, transcript speaker, or GM when the in-world caster can be identified."
},
"spell": {
"type": "string",
"minLength": 1
"minLength": 1,
"description": "Canonical spell name from the provided spell-name catalog."
},
"effect": {
"type": "string",
"minLength": 1
"minLength": 1,
"description": "Concise effect or resolution established by the cited transcript units; do not infer mechanics from general D&D rules knowledge."
},
"narrative_description": {
"type": "string",
"minLength": 1
"minLength": 1,
"description": "Short session-grounded description containing only details established by the cited transcript units."
},
"source_refs": {
"type": "array",
"minItems": 1,
"description": "One or more narrow transcript ranges that collectively support every factual claim in this spell-cast object, including any reported resolution or outcome.",
"items": {
"type": "object",
"additionalProperties": false,

View File

@@ -51,6 +51,8 @@ type Extractor struct {
llm contracts.StructuredLLMClient
effectiveCatalog spellcatalog.EffectiveCatalog
catalogPromptInput contracts.LLMInputMaterial
promptSHA string
responseSchemaSHA string
}
func New(llmClient contracts.StructuredLLMClient, _ Options, references ...contracts.ReferenceSet) (*Extractor, error) {
@@ -72,10 +74,20 @@ func New(llmClient contracts.StructuredLLMClient, _ Options, references ...contr
if err != nil {
return nil, extractorErrorf("prepare spell catalog prompt input: %w", err)
}
promptSHA, err := scriptoriumPromptMetadata()
if err != nil {
return nil, extractorErrorf("load prompt metadata: %w", err)
}
responseSchema, err := loadResponseSchema()
if err != nil {
return nil, extractorErrorf("load response schema: %w", err)
}
return &Extractor{
llm: llmClient,
effectiveCatalog: effectiveCatalog,
catalogPromptInput: catalogPromptInput,
promptSHA: promptSHA,
responseSchemaSHA: responseSchema.SHA256,
}, nil
}
@@ -88,24 +100,18 @@ func (e *Extractor) ReferenceSlots() []contracts.ReferenceSlot {
}
func (e *Extractor) ManifestMetadata() map[string]any {
promptSHA, err := scriptoriumPromptMetadata()
if err != nil {
promptSHA = ""
}
metadata := map[string]any{
"prompt_id": PromptID,
"prompt_version": SchemaVersion,
"prompt_sha256": promptSHA,
"catalog_base_id": e.effectiveCatalog.BaseID(),
"catalog_digest": e.effectiveCatalog.Digest(),
"catalog_overlay_ids": e.effectiveCatalog.OverlayIDs(),
"response_schema_key": string(ResponseSchemaKey),
"response_schema_id": ResponseSchemaID,
"response_schema_name": ResponseSchemaName,
}
if schema, err := loadResponseSchema(); err == nil {
metadata["response_schema_version"] = schema.Version
metadata["response_schema_sha256"] = schema.SHA256
"prompt_id": PromptID,
"prompt_version": SchemaVersion,
"prompt_sha256": e.promptSHA,
"catalog_base_id": e.effectiveCatalog.BaseID(),
"catalog_digest": e.effectiveCatalog.Digest(),
"catalog_overlay_ids": e.effectiveCatalog.OverlayIDs(),
"response_schema_key": string(ResponseSchemaKey),
"response_schema_id": ResponseSchemaID,
"response_schema_name": ResponseSchemaName,
"response_schema_version": SchemaVersion,
"response_schema_sha256": e.responseSchemaSHA,
}
return metadata
}
@@ -114,7 +120,11 @@ func (e *Extractor) CheckpointFingerprints() []pipeline.CheckpointFingerprint {
if e == nil {
return nil
}
return []pipeline.CheckpointFingerprint{{Name: "effective_catalog", Value: e.effectiveCatalog.Digest()}}
return []pipeline.CheckpointFingerprint{
{Name: "effective_catalog", Value: e.effectiveCatalog.Digest()},
{Name: "prompt", Value: e.promptSHA},
{Name: "response_schema", Value: e.responseSchemaSHA},
}
}
func (e *Extractor) Extract(ctx context.Context, req contracts.TypedExtractionRequest) (contracts.TypedExtractionResult[dnd.SpellList], error) {

View File

@@ -118,8 +118,18 @@ func TestExtractPromptUsesCanonicalOverlayNamesWithoutAliasesOrMetadata(t *testi
t.Fatalf("catalog overlay metadata = %#v", metadata["catalog_overlay_ids"])
}
fingerprints := newExtractor(t, &fakeSpellsLLMClient{}, overlaySpellCatalogReference()).CheckpointFingerprints()
if len(fingerprints) != 1 || fingerprints[0].Name != "effective_catalog" || fingerprints[0].Value != metadata["catalog_digest"] {
t.Fatalf("checkpoint fingerprints = %#v, want manifest catalog digest %#v", fingerprints, metadata["catalog_digest"])
wantFingerprints := map[string]any{
"effective_catalog": metadata["catalog_digest"],
"prompt": metadata["prompt_sha256"],
"response_schema": metadata["response_schema_sha256"],
}
if len(fingerprints) != len(wantFingerprints) {
t.Fatalf("checkpoint fingerprints = %#v, want prompt, response schema, and catalog identities", fingerprints)
}
for _, fingerprint := range fingerprints {
if want, ok := wantFingerprints[fingerprint.Name]; !ok || fingerprint.Value != want {
t.Fatalf("checkpoint fingerprint %q = %q, want %#v", fingerprint.Name, fingerprint.Value, want)
}
}
encoded, err := json.Marshal(metadata)
if err != nil {