139 lines
5.5 KiB
Markdown
139 lines
5.5 KiB
Markdown
# D&D NPC Artifact
|
|
|
|
This document defines the durable D&D NPC-list artifact, its JSON codec, and
|
|
the selectable production NPC pipeline. The normalized JSON payload can be
|
|
passed explicitly to the spell extractor as an optional caster-name registry
|
|
or to the combat extractor and normalizer as an actor/target registry. It
|
|
remains a reference, not spell or combat evidence.
|
|
|
|
## Identity
|
|
|
|
- Artifact kind: `dnd/npc-list`
|
|
- Durable schema ID: `notarius.dnd.npcs`
|
|
- Durable schema name: `notarius_dnd_npcs_v1`
|
|
- Durable schema version: `v1`
|
|
- Media type: `application/json`
|
|
- Identity policy: `dnd.npcs.identity.v1`
|
|
|
|
The durable JSON Schema is owned by the D&D NPC codec. NPC IDs are derived from
|
|
the Unicode-normalized, case-folded canonical name using the identity policy.
|
|
The durable codec enforces the artifact shape and ID syntax; registry identity
|
|
validation remains a separate deterministic concern.
|
|
|
|
The extractor's private LLM response schema is a separate structural transport
|
|
contract. It omits framework-assigned NPC and source IDs and admits semantic
|
|
candidates for the deterministic shape and source-reference validators; it is
|
|
not part of this durable contract.
|
|
|
|
## Output Shape
|
|
|
|
The payload is one object with a required top-level `npcs` array:
|
|
|
|
```json
|
|
{"npcs": []}
|
|
```
|
|
|
|
The array may be empty. Every object and nested object rejects unknown fields.
|
|
|
|
## NPC Fields
|
|
|
|
Each NPC contains exactly these required fields:
|
|
|
|
- `id`: `npc:sha256:` followed by 64 lowercase hexadecimal characters;
|
|
- `name`: the canonical display name;
|
|
- `source_refs`: at least one source reference supporting the NPC record.
|
|
|
|
Each source reference contains required `source_id`, `start_unit_id`, and
|
|
`end_unit_id`; unit IDs are positive integers. Source document identity, unit
|
|
existence, and range ordering are validated by the source-reference validator
|
|
when the artifact is used by a pipeline.
|
|
|
|
## Codec Boundary
|
|
|
|
`EncodeCandidate` and `DecodeCandidate` provide strict single-value JSON
|
|
serialization while preserving typed values that still need semantic
|
|
validation. `Encode` and `Decode` are the approved-artifact boundary and
|
|
require all durable structural fields, non-empty required strings, valid source
|
|
reference shapes, and the NPC ID pattern.
|
|
|
|
Codec metadata contains only `npc_count`. Schema bytes and returned metadata
|
|
are independent values so callers cannot mutate codec-owned state.
|
|
|
|
## Production Pipeline
|
|
|
|
The production identities are:
|
|
|
|
- extractor: `dnd/npcs`;
|
|
- artifact kind: `dnd/npc-list`;
|
|
- normalizer: `dnd/npcs`; and
|
|
- durable schema: `notarius.dnd.npcs`, version `v1`, media type
|
|
`application/json`.
|
|
|
|
The extractor maps private model records to the current source identity and
|
|
assigns deterministic IDs. Extraction validation checks shape, source
|
|
references, and source relatedness. The normalizer then consolidates records
|
|
only when their normalized canonical names match, preserves the first record's
|
|
display and output position, unions exact evidence, and validates the retained
|
|
registry's identity. No LLM is used for consolidation.
|
|
|
|
The extraction prompt asks only for individually identifiable NPC names backed
|
|
by source evidence. Groups, generic roles, invented labels, and descriptive or
|
|
relationship enrichment are outside the contract.
|
|
|
|
The default extraction chain is `generic/valid_json`,
|
|
`generic/valid_json_schema`, `extract/dnd/npcs/shape`,
|
|
`extract/dnd/npcs/source_refs`, and
|
|
`extract/dnd/npcs/source_relatedness`. The normalize chain adds
|
|
`normalize/dnd/npcs/identity` before the source-reference and relatedness
|
|
checks. Relatedness emits bounded warnings when an NPC canonical name is not
|
|
present near its cited transcript text; opaque campaign
|
|
references may explain such a warning but do not become evidence.
|
|
|
|
## Manifest And Artifact Handoff
|
|
|
|
The NPC extractor records prompt and response-schema identities. The durable
|
|
codec records only `npc_count`; raw names, source references, and payload bytes
|
|
stay in the lane file rather than manifest
|
|
metadata. The normalized lane can be consumed by a later ordered step through
|
|
the registered canonical codec:
|
|
|
|
```yaml
|
|
steps:
|
|
- id: identify-npcs
|
|
artifacts:
|
|
npcs:
|
|
extract: dnd/npcs
|
|
normalize: dnd/npcs
|
|
- id: grounded-events
|
|
references:
|
|
npcs:
|
|
artifact:
|
|
step: identify-npcs
|
|
lane: npcs
|
|
artifacts:
|
|
spells:
|
|
extract: dnd/spells
|
|
normalize: dnd/spells
|
|
combat:
|
|
extract: dnd/combat-turns
|
|
normalize: dnd/combat-turns
|
|
```
|
|
|
|
The framework hands only an accepted normalized artifact across the barrier. It
|
|
validates the canonical bytes against each consumer slot and clones the
|
|
operation-time reference for the spell and combat consumers. Generated
|
|
provenance records the artifact kind, schema identity, media type, canonical
|
|
digest, size, and producer step/lane/module, but not names, source
|
|
ranges, or payload bytes. External normalized files remain supported as
|
|
explicit references and retain their file provenance.
|
|
|
|
NPC source references are registry provenance and are never accepted as spell
|
|
or combat evidence. Current transcript units remain the only event evidence.
|
|
|
|
Consumers receive a separate names-only projection in normalized registry
|
|
order, for example `{"npcs":[{"name":"Mira Thorn"}]}`. The projection omits
|
|
IDs and evidence. Its digest covers the exact projected bytes and is used for
|
|
consumer-local checkpoint identity, while the full durable artifact digest
|
|
remains the manifest and generated-reference provenance identity. The unbound
|
|
projection is exactly `{"npcs":[]}` and also has a projection digest.
|