Plan D&D spells extractor implementation

This commit is contained in:
2026-07-03 23:25:32 +00:00
parent 39fcfba605
commit a4d64f6f16
2 changed files with 482 additions and 339 deletions

View File

@@ -23,7 +23,7 @@ In scope:
- module metadata/capability requirements for pipeline validation;
- source-reference and schema validators in the extractor chain;
- fake LLM tests;
- CLI-level integration test if the CLI path is ready.
- runner-level integration tests with fake infrastructure.
Out of scope:
@@ -31,89 +31,71 @@ Out of scope:
- NPC extraction;
- combat extraction;
- cross-slice deduplication beyond simple deterministic merging;
- broad D&D rules validation.
- broad D&D rules validation;
- a `notarius run` command.
## Proposed Stages
## Target End State
### Stage 1: Spell Artifact Schema
The repository should contain a real D&D spells extract-stage module at
`internal/modules/extract/dnd/spells`.
Define the D&D spell artifact model.
The spells module should be registered under the stable extractor key
`dnd/spells`. It should be selectable as a named artifact lane in pipeline
configuration, for example a lane named `spells` whose extractor module is
`dnd/spells`.
Initial shape:
The module should translate generic source chunks into spell-cast artifact
candidates with this spell payload:
```go
type SpellCast struct {
Player string `json:"player"`
Spell string `json:"spell"`
Effect string `json:"effect"`
NarrativeDescription string `json:"narrative_description"`
SourceRefs []SourceRef `json:"source_refs"`
}
```
- `caster`;
- `spell`;
- `effect`;
- `narrative_description`.
Keep this schema inside the D&D spells extract module or a D&D artifact package,
not inside core framework packages.
The LLM structured response should also include source references for each
spell cast. The extractor should copy those references into the generic
artifact envelope rather than duplicating source references inside the durable
spell payload.
### Stage 2: Structured Response Schema
Add a structured response schema asset for spell extraction.
The schema should require:
- spell-cast array;
- non-empty player, spell, effect, and narrative description fields;
- at least one source reference per spell cast.
### Stage 3: Prompt Assets
Add embedded prompt assets for D&D spell extraction.
The caster is the in-world character or creature casting the spell, not the
table speaker. Speaker metadata from transcript source units may be used as
optional prompt context when present, but it should not be a required or durable
field in the spell payload.
Prompts should:
- describe the generic source-unit input format;
- explain that source references must use source-unit IDs;
- explain that source references must use source-unit IDs exactly;
- avoid relying on transcript-specific fields except as optional metadata;
- request only spell-cast artifacts.
- request only D&D spell-cast artifacts.
### Stage 4: Process Module Implementation
The extractor should attach deterministic validators by default. Validation
should cover:
Implement `internal/modules/extract/dnd/spells`.
- spell payload shape;
- non-empty required spell fields;
- at least one source reference per spell cast;
- source references that validate against the source document.
The extractor should:
The module should declare flat capabilities for pipeline validation. Initial
capabilities should require chunked transcript source material and provide a
D&D spell-cast artifact capability.
- satisfy the framework `Extractor` contract;
- declare module metadata for pipeline-profile validation;
- build LLM messages from a source document or source chunk;
- call the structured LLM client;
- return artifact candidates with source references;
- attach its validator chain.
Implementation staging belongs in
[`implementation.md`](implementation.md).
### Stage 5: Validators And Tests
## Fixtures And Tests
Wire deterministic validators:
The checkpoint should add synthetic fixtures and focused tests for:
- schema/shape validation;
- source-reference validation;
- required-field validation if not covered by schema handling.
Add tests using a fake structured LLM client:
- successful spell extraction;
- empty result;
- invalid source reference rejection;
- successful spell extraction with a fake structured LLM client;
- empty spell-cast results;
- malformed structured output handling;
- stable output ordering.
### Stage 6: CLI Integration
If the CLI path is ready, add an end-to-end test using:
```sh
notarius run dnd-session --input ./transcript.json --only spells
```
The test should use fake LLM wiring, fixture input, and a named pipeline profile
with a `spells` artifact lane.
- invalid source-reference rejection;
- stable output ordering;
- pipeline-profile selection of a `spells` artifact lane;
- runner integration from Seriatim input through the spells extractor using
fake chunk and output modules.
## Done Criteria
@@ -123,8 +105,8 @@ with a `spells` artifact lane.
- D&D concepts are contained in extract module/artifact packages and docs.
- The spells module can be selected as a named artifact lane in pipeline
configuration.
- The first meaningful vertical slice is available through tests, and through
CLI if the CLI path is ready.
- The first meaningful vertical slice is available through tests.
- No CLI `run` behavior is documented or implemented until the CLI path exists.
## Review Questions

View File

@@ -1,15 +1,14 @@
# Implementation Plan: Checkpoint 5 Seriatim Input Module
# Implementation Plan: Checkpoint 6 D&D Spells Extractor
## Status
This is a staged implementation plan for
[`5-seriatim-input-module.md`](5-seriatim-input-module.md). It is intended for
an LLM coding agent to follow stage by stage.
[`6-dnd-spells-extractor.md`](6-dnd-spells-extractor.md). It is intended for an
LLM coding agent to follow stage by stage.
This plan implements only checkpoint 5. Do not add D&D extraction, real domain
prompts or schemas, LLM extraction calls, a `notarius run` command, broad
Seriatim schema support, or transcript-specific behavior in core framework
packages in this checkpoint.
This plan implements only checkpoint 6. Do not add D&D item, NPC, combat, or
encounter extraction; do not add broad D&D rules validation; do not add a
`notarius run` command; and do not make core framework packages D&D-specific.
## Policy Context
@@ -20,404 +19,564 @@ Follow:
Required boundaries:
- keep Seriatim JSON schema details inside `internal/modules/input/seriatim`;
- keep core source, runner, pipeline, extractor, validator, LLM, and config
- keep D&D-specific artifact semantics, prompt data shaping, and response
interpretation inside `internal/modules/extract/dnd/spells`;
- keep source-format details inside input modules and do not depend on concrete
Seriatim package types from the spells extractor;
- keep core source, runner, pipeline, LLM, prompt, config, and validator
packages source-agnostic and domain-agnostic;
- do not add transcript-specific typed fields to `SourceDocument`,
`SourceUnit`, runner contracts, or pipeline contracts;
- preserve transcript-specific values only as source metadata conventions;
- register the input module through the existing input adapter registry instead
of adding ad hoc conditionals;
- use embedded prompt and schema assets instead of inline prompt/schema strings;
- register the extractor through the existing extractor registry;
- use flat capability strings in module metadata;
- keep future or planned behavior in `docs/roadmap/` until implemented.
## Global Implementation Decisions
- Add no new third-party dependency. Use `encoding/json` with
`Decoder.UseNumber` for Seriatim JSON parsing.
- Use `seriatim` as the stable input adapter key.
- Put all concrete Seriatim input code under
`internal/modules/input/seriatim`.
- Expose a small module API:
- Add no new third-party dependency.
- Use `dnd/spells` as the stable extractor module key.
- Put concrete extractor code under `internal/modules/extract/dnd/spells`.
- Use `dnd.spell_cast` as the artifact type and `v1` as the schema version.
- Use `dnd.spell_casts` as the extractor-provided capability.
- The extractor module spec must require `chunks` and `source.transcript`.
It must not require `transcript.speaker` or `transcript.timestamps`; those
metadata values may be included in prompts when present, but extraction must
not depend on Seriatim-specific helper APIs.
- Keep durable artifact payloads focused on spell data:
```go
const Key = "seriatim"
func New() *Adapter
func ModuleSpec() pipeline.ModuleSpec
func Register(registry *pipeline.InputAdapterRegistry) error
type SpellCast struct {
Caster string `json:"caster"`
Spell string `json:"spell"`
Effect string `json:"effect"`
NarrativeDescription string `json:"narrative_description"`
}
```
- `ModuleSpec()` must return stage `pipeline.StageInput`, no required
capabilities, and these provided capabilities:
`source.transcript`, `transcript.speaker`, and `transcript.timestamps`.
- The Seriatim adapter should satisfy `contracts.InputAdapter`.
- Use `source.SourceDocument.Kind = "transcript"`.
- Use `source.SourceDocument.Format =
"application/vnd.seriatim.minimal+json"`.
- Use `source.SourceUnit.Kind = "transcript_segment"`.
- Compute `SourceDocument.Digest` from the exact raw input bytes as
`sha256:<hex>`.
- Resolve `SourceDocument.ID` in this order:
1. trimmed `contracts.ParseRequest.SourceID`, if non-empty;
2. trimmed string `metadata.id`, if present and non-empty;
3. trimmed string `metadata.source_id`, if present and non-empty;
4. deterministic fallback `seriatim:<first-16-hex-chars-of-raw-sha256>`.
- Segment IDs become source unit IDs exactly after validation. Reject segment
IDs with leading or trailing whitespace rather than silently rewriting them.
- Copy top-level Seriatim `metadata` into `SourceDocument.Metadata`.
- Store segment `speaker`, `start`, and `end` in `SourceUnit.Metadata` under
keys with those exact names.
- Store `start` and `end` as `json.Number` values so JSON serialization remains
numeric and the original decimal representation is preserved.
- Require top-level `metadata` to be present and be an object, but do not
require any specific metadata keys in checkpoint 5.
- Require top-level `segments` to be present and contain at least one segment.
- Reject unknown or extra JSON fields only if they prevent parsing the minimal
shape. Otherwise ignore them so the module can tolerate compatible Seriatim
additions.
- Return module-specific errors prefixed with useful Seriatim context, for
example `seriatim input: segment "s1" text must not be empty`.
- Keep examples free of private transcript content. Use synthetic fixture text.
- The LLM response model must include source references so the extractor can
populate `artifacts.ArtifactCandidate.SourceRefs`, but source references must
live in the generic artifact envelope for durable pipeline output rather than
being duplicated inside the `SpellCast` payload.
- `caster` means the in-world character or creature casting the spell. Do not
use table speaker/player as the durable caster field. Source-unit `speaker`
metadata may be rendered as optional prompt context when present.
- Use `source.SourceRef` JSON field names exactly as the core type defines
them: `source_id`, `start_unit_id`, and `end_unit_id`.
- Trim text fields before marshaling candidate payloads. Do not silently trim
source-reference IDs; invalid source refs must be rejected by validation.
- `Extractor.Validators()` must return module-owned deterministic validators by
default. Do not require pipeline profiles to configure validators explicitly
for the default checkpoint behavior.
- Treat an empty `spell_casts` response array as a successful empty extraction.
Treat a missing or null `spell_casts` field as malformed structured output.
- Use synthetic D&D transcript fixture text only. Do not include private
campaign transcript content.
## Stage 1: Seriatim Package Skeleton And External Model
## Stage 1: Domain Model, Module Skeleton, And Registry Metadata
### Goal
Create the Seriatim input module package, define the module-local JSON model,
and add registry-facing module metadata without changing framework contracts.
Create the D&D spells extractor package, define the artifact and response
models, and make the extractor discoverable through the existing extractor
registry without calling an LLM yet.
### Files To Add
### Files To Add Or Update
- `internal/modules/input/seriatim/adapter.go`
- `internal/modules/input/seriatim/model.go`
- `internal/modules/input/seriatim/metadata.go`
- `internal/modules/input/seriatim/registry_test.go`
- `internal/modules/extract/dnd/spells/extractor.go`
- `internal/modules/extract/dnd/spells/model.go`
- `internal/modules/extract/dnd/spells/registry_test.go`
### Required API
Add:
```go
package seriatim
package spells
const Key = "seriatim"
const Key = "dnd/spells"
const ArtifactType = "dnd.spell_cast"
const SchemaVersion = "v1"
const (
DocumentKind = "transcript"
UnitKind = "transcript_segment"
Format = "application/vnd.seriatim.minimal+json"
)
type SpellCast struct {
Caster string `json:"caster"`
Spell string `json:"spell"`
Effect string `json:"effect"`
NarrativeDescription string `json:"narrative_description"`
}
const (
MetadataSpeaker = "speaker"
MetadataStart = "start"
MetadataEnd = "end"
)
type Extractor struct{}
type Adapter struct{}
func New() *Adapter
func (a *Adapter) Key() string
func (a *Adapter) Parse(ctx context.Context, req contracts.ParseRequest) (*source.SourceDocument, error)
func New() *Extractor
func (e *Extractor) Key() string
func (e *Extractor) ArtifactType() string
func (e *Extractor) SchemaVersion() string
func (e *Extractor) Validators() []contracts.Validator
func (e *Extractor) Extract(ctx context.Context, req contracts.ExtractionRequest) (contracts.ExtractionResult, error)
func ModuleSpec() pipeline.ModuleSpec
func Register(registry *pipeline.InputAdapterRegistry) error
func Register(registry *pipeline.ExtractorRegistry) error
```
Add typed metadata helpers:
Add unexported response structs in the same package:
```go
func Speaker(unit source.SourceUnit) (string, bool)
func Start(unit source.SourceUnit) (json.Number, bool)
func End(unit source.SourceUnit) (json.Number, bool)
```
### Seriatim JSON Shape
Define module-local structs for the minimal external shape:
```go
type transcript struct {
Metadata map[string]any `json:"metadata"`
Segments []segment `json:"segments"`
type extractionResponse struct {
SpellCasts []spellCastResponse `json:"spell_casts"`
}
type segment struct {
ID string `json:"id"`
Start json.Number `json:"start"`
End json.Number `json:"end"`
Speaker string `json:"speaker"`
Text string `json:"text"`
type spellCastResponse struct {
Caster string `json:"caster"`
Spell string `json:"spell"`
Effect string `json:"effect"`
NarrativeDescription string `json:"narrative_description"`
SourceRefs []source.SourceRef `json:"source_refs"`
}
```
Use an internal decode helper based on `json.NewDecoder(bytes.NewReader(raw))`
and `UseNumber`.
### Required Behavior
- `New()` returns a non-nil adapter.
- `Adapter.Key()` returns `Key`.
- `ModuleSpec()` returns defensive slices and the capability set listed in the
global decisions.
- `Register()` calls `InputAdapterRegistry.RegisterWithSpec(ModuleSpec(), ...)`.
- `New()` returns a non-nil extractor.
- `ModuleSpec()` returns defensive slices with stage `pipeline.StageExtract`,
key `dnd/spells`, requires `chunks` and `source.transcript`, and provides
`dnd.spell_casts`.
- `Register()` calls `ExtractorRegistry.RegisterWithSpec(ModuleSpec(), ...)`.
- `Register(nil)` returns an error from the registry path rather than panicking.
- Keep external JSON structs unexported.
- `Extract()` may return a clear not-yet-implemented error in this stage only.
- `Validators()` returns nil in this stage only; later stages must replace it
with the real validator chain.
### Required Tests
- `New()` returns an adapter whose key is `seriatim`.
- `ModuleSpec()` uses input stage and declares the required provided
capabilities.
- `Register()` makes the adapter buildable from an `InputAdapterRegistry`.
- Registry lookup returns the Seriatim module spec.
- `New()` returns an extractor whose key, artifact type, and schema version
match the constants.
- `ModuleSpec()` uses extract stage and declares the required capabilities.
- Caller mutation of `ModuleSpec().Provides` or `ModuleSpec().Requires` does not
affect later calls.
- `Register()` makes the extractor buildable from an `ExtractorRegistry`.
- Registry lookup returns the spells module spec.
- `Register(nil)` returns an error containing extractor registry context.
### Validation
Run:
```sh
gofmt -w internal/modules/input/seriatim
go test ./internal/modules/input/seriatim
gofmt -w internal/modules/extract/dnd/spells
go test ./internal/modules/extract/dnd/spells
go test ./...
```
## Stage 2: Parse, Validate, And Map To SourceDocument
## Stage 2: Structured Response Schema Asset
### Goal
Implement the Seriatim parser and mapper from minimal Seriatim JSON into the
generic source model.
Add the structured response schema used by the D&D spells extractor and register
it through the existing LLM schema registry.
### Files To Add Or Update
- `internal/modules/input/seriatim/adapter.go`
- `internal/modules/input/seriatim/model.go`
- `internal/modules/input/seriatim/adapter_test.go`
- `internal/modules/input/seriatim/testdata/valid_minimal.json`
- `internal/modules/input/seriatim/testdata/duplicate_segment_id.json`
- `internal/framework/llm/assets/schemas/dnd_spells.v1.json`
- `internal/framework/llm/schema_registry.go`
- `internal/framework/llm/schema_registry_test.go`
- `internal/modules/extract/dnd/spells/schema_test.go`
### Required Validation
### Schema Decisions
Reject:
Register:
- nil or canceled context before parsing;
- empty raw input;
- malformed JSON;
- valid JSON with trailing non-whitespace data;
- missing, null, or non-object top-level `metadata`;
- missing, null, empty, or non-array top-level `segments`;
- segment IDs that are empty after trimming;
- segment IDs with leading or trailing whitespace;
- duplicate segment IDs;
- missing or empty `speaker`;
- missing, empty, non-numeric, negative, or non-finite `start`;
- missing, empty, non-numeric, negative, or non-finite `end`;
- segments where `end < start`;
- missing or empty `text`.
- response schema key: `dnd_spells`
- schema ID: `notarius.dnd.spells`
- schema version: `v1`
- response schema name: `notarius_dnd_spells_v1`
- asset path: `assets/schemas/dnd_spells.v1.json`
The parser may preserve leading and trailing whitespace in segment text as long
as the text is not empty after trimming.
The JSON schema must require a top-level object:
### Mapping Rules
```json
{
"spell_casts": [
{
"caster": "Aria",
"spell": "Cure Wounds",
"effect": "Heals an injured ally.",
"narrative_description": "Aria casts Cure Wounds after the fight.",
"source_refs": [
{
"source_id": "session-alpha",
"start_unit_id": "seg-001",
"end_unit_id": "seg-002"
}
]
}
]
}
```
- `segment.id` becomes `SourceUnit.ID`.
- `segment.text` becomes `SourceUnit.Text`.
- `speaker`, `start`, and `end` become unit metadata under the exact keys
defined in stage 1.
- The document metadata is a shallow copy of top-level Seriatim metadata.
- The document digest is based on raw input bytes, not normalized JSON.
- Call `source.ValidateDocument` before returning the document and wrap any
validation failure with Seriatim context.
Required schema constraints:
- `additionalProperties: false` at every object level.
- top-level `spell_casts` is required and must be an array.
- `spell_casts` may be empty.
- each spell cast requires non-empty `caster`, `spell`, `effect`, and
`narrative_description`.
- each spell cast requires `source_refs` with `minItems: 1`.
- each source ref requires non-empty `source_id`, `start_unit_id`, and
`end_unit_id`.
### Required Tests
Add tests for:
- valid minimal transcript parses to a source document with expected ID, kind,
format, digest, units, and metadata;
- `ParseRequest.SourceID` overrides metadata-derived IDs;
- fallback document ID is deterministic and has prefix `seriatim:`;
- malformed JSON returns an actionable Seriatim parse error;
- missing metadata is rejected;
- missing or empty segments is rejected;
- duplicate segment IDs are rejected;
- empty segment text is rejected;
- missing speaker is rejected;
- invalid timestamp values are rejected;
- `end < start` is rejected;
- typed metadata helpers return the expected speaker and timestamp values;
- a `source.SourceRef` using the first and last generated unit IDs validates
with `source.ValidateRef`.
- `llm.LookupResponseSchema(llm.DNDSpellsSchemaKey)` succeeds.
- Registered schema list remains sorted and now includes the D&D spells schema.
- The schema content is valid JSON and mutation-safe through registry lookups.
- The spells package can look up the schema key it will use during extraction.
- The schema diagnostics map omits raw schema content.
### Validation
Run:
```sh
gofmt -w internal/modules/input/seriatim
go test ./internal/modules/input/seriatim
gofmt -w internal/framework/llm internal/modules/extract/dnd/spells
go test ./internal/framework/llm
go test ./internal/modules/extract/dnd/spells
go test ./...
```
## Stage 3: Pipeline Resolution And Config Compatibility
## Stage 3: Prompt Assets And Prompt Rendering
### Goal
Prove the Seriatim input module participates in pipeline-profile resolution and
capability validation through existing registries and config loading.
Add embedded prompt assets for D&D spell extraction and render deterministic
system/user messages from generic source chunks.
### Files To Add Or Update
- `internal/modules/input/seriatim/config_test.go`
- `internal/modules/input/seriatim/testdata/pipeline.yml`
- `internal/framework/prompt/assets/dnd/spells/system.md`
- `internal/framework/prompt/assets/dnd/spells/user.md`
- `internal/framework/prompt/registry.go`
- `internal/framework/prompt/registry_test.go`
- `internal/framework/prompt/render_test.go`
- `internal/modules/extract/dnd/spells/prompt.go`
- `internal/modules/extract/dnd/spells/prompt_test.go`
### Prompt Decisions
Register:
- prompt ID: `dnd.spells`
- prompt constant: `DNDSpellsPromptID`
- prompt version: `v1`
- embedded directory: `assets/dnd/spells`
The system prompt must:
- include the shared prompt hardening text via `{{ hardening }}`;
- identify the task as extracting D&D spell casts only;
- state that the model must not infer spells not supported by the provided
source text;
- state that source references must use source-unit IDs exactly as provided.
The user prompt must render:
- source document ID;
- chunk ID and chunk index when a chunk is present;
- source units in their existing order;
- each source unit's ID and text;
- selected metadata only as optional context, using generic labels. Include
`speaker`, `start`, and `end` when present without importing Seriatim helper
APIs.
Add a module-local prompt data builder that accepts
`contracts.ExtractionRequest` and returns template data. It must not mutate the
request, source document, chunk, units, or metadata maps.
### Required Tests
- Prompt metadata lookup succeeds for `dnd.spells`.
- Registered prompt metadata remains sorted.
- Rendering includes hardening text.
- Rendering includes source ID, chunk ID, unit IDs, unit text, and optional
speaker/timestamp metadata when present.
- Rendering fails clearly if required template data is missing.
- Module-local prompt data construction works with a generic source chunk and
does not depend on concrete Seriatim package helpers.
### Validation
Run:
```sh
gofmt -w internal/framework/prompt internal/modules/extract/dnd/spells
go test ./internal/framework/prompt
go test ./internal/modules/extract/dnd/spells
go test ./...
```
## Stage 4: Extractor LLM Call And Candidate Mapping
### Goal
Implement `Extractor.Extract` so a source chunk can flow through prompt
rendering, structured LLM completion, and candidate creation.
### Files To Add Or Update
- `internal/modules/extract/dnd/spells/extractor.go`
- `internal/modules/extract/dnd/spells/model.go`
- `internal/modules/extract/dnd/spells/extractor_test.go`
### Required Behavior
`Extract()` must:
- reject nil extractor, nil context, canceled context, nil source, nil chunk,
empty chunk units, and nil LLM client with clear `dnd spells` error context;
- render the `dnd.spells` prompt using the stage 3 prompt data builder;
- look up the stage 2 response schema;
- call `req.LLMClient.CompleteStructured` with:
- `StageName: Key`;
- two messages, system then user;
- `ResponseSchemaName: schema.Name`;
- `ResponseSchema: schema.JSONSchema`;
- decode into `extractionResponse`;
- reject a nil `SpellCasts` slice as malformed structured output;
- return no candidates for an empty `SpellCasts` slice;
- preserve response order when creating candidates;
- trim `caster`, `spell`, `effect`, and `narrative_description` before
marshaling the `SpellCast` payload;
- copy response source refs into `ArtifactCandidate.SourceRefs`;
- leave `ExtractorKey`, `ArtifactType`, `SchemaVersion`, and `Index` empty or
zero so the runner's existing candidate normalization remains authoritative;
- wrap LLM, prompt, schema lookup, response validation, and JSON marshal errors
with useful `dnd spells` context.
Do not call validators from `Extract()`. Validation belongs to the runner's
validator phase.
### Required Tests
Use a fake `contracts.StructuredLLMClient`.
- Successful extraction returns one candidate with expected payload fields,
source refs, response schema name, schema JSON, prompt messages, and stage
name.
- Empty `spell_casts: []` returns no candidates and no error.
- Missing or null `spell_casts` is rejected as malformed structured output.
- LLM client errors are wrapped with D&D spells context.
- Nil source, nil chunk, empty chunk units, nil LLM client, nil context, and
canceled context are rejected.
- Multiple response spell casts produce candidates in response order.
- Candidate source refs are copied so later mutation of the fake response does
not mutate returned candidates.
### Validation
Run:
```sh
gofmt -w internal/modules/extract/dnd/spells
go test ./internal/modules/extract/dnd/spells
go test ./...
```
## Stage 5: Deterministic Validator Chain
### Goal
Add module-owned deterministic validators for spell payload shape, required
fields, and source-reference grounding.
### Files To Add Or Update
- `internal/modules/extract/dnd/spells/validator.go`
- `internal/modules/extract/dnd/spells/validator_test.go`
- `internal/modules/extract/dnd/spells/extractor.go`
### Validator Decisions
Add two validators:
1. `ShapeValidator`
- `Name()` returns `dnd/spells/shape`.
- Rejects malformed JSON payloads with reason code `invalid_payload`.
- Rejects blank `caster`, `spell`, `effect`, or `narrative_description`
with reason code `missing_required_field`.
- Approves candidates with valid payload shape and required fields.
2. `SourceRefValidator`
- `Name()` returns `dnd/spells/source_refs`.
- Rejects candidates with no source refs using reason code
`missing_source_ref`.
- Rejects any invalid source ref using reason code `invalid_source_ref` and
the `source.ValidateRef` error message.
- Approves candidates whose source refs all validate against the request
source document.
`Extractor.Validators()` must return `ShapeValidator{}` followed by
`SourceRefValidator{}`. Return a fresh slice each time.
Both validators must:
- satisfy `contracts.Validator`;
- return one decision for every candidate;
- preserve candidate indexes in decisions;
- return an error, not rejection decisions, when called with nil source only if
source-reference validation cannot run. Shape validation does not require a
source document.
### Required Tests
- `Extractor.Validators()` returns the two validators in the required order and
is mutation-safe.
- Each validator approves a valid candidate.
- Shape validator rejects malformed payload JSON.
- Shape validator rejects each blank required field.
- Source ref validator rejects missing refs.
- Source ref validator rejects unknown source IDs, unknown unit IDs, and
reversed unit ranges.
- Both validators return one decision per candidate and preserve indexes.
- Approved decisions use `validate.Approved`; rejected decisions use the reason
codes listed above.
### Validation
Run:
```sh
gofmt -w internal/modules/extract/dnd/spells
go test ./internal/modules/extract/dnd/spells
go test ./...
```
## Stage 6: Pipeline Resolution And Runner Integration
### Goal
Prove Seriatim input can flow through the existing runner into the real D&D
spells extractor, fake downstream infrastructure, and the extractor-owned
validators.
### Files To Add Or Update
- `internal/modules/extract/dnd/spells/config_test.go`
- `internal/modules/extract/dnd/spells/runner_test.go`
- `internal/modules/extract/dnd/spells/testdata/seriatim_spell_session.json`
- `internal/modules/extract/dnd/spells/testdata/pipeline.yml`
### Required Test Catalog
Build a test-only module catalog with:
Build test-only catalogs and registries with:
- Seriatim input registered through `seriatim.Register`;
- a fake chunker requiring `source.transcript` and providing `chunks`;
- a fake extractor requiring `chunks`, `transcript.speaker`, and
`transcript.timestamps`, and providing `fake.artifacts`;
- a fake chunker registered as `fake/chunk`, requiring `source.transcript` and
providing `chunks`;
- D&D spells extractor registered through `spells.Register`;
- `pipeline.AppendOrderMerger` registered as `appendorder`, requiring
`fake.artifacts`;
`dnd.spell_casts`;
- `pipeline.NoopNormalizer` registered as `noop`;
- a fake `json` output encoder registered as output stage.
Do not add real extract, chunk, normalize, or output modules for this checkpoint.
Do not add real chunk or output modules for this checkpoint.
### YAML Fixture
Use a synthetic pipeline fixture shaped like:
Use a synthetic pipeline fixture:
```yaml
version: 1
pipelines:
seriatim-fixture:
dnd-spells-fixture:
input: seriatim
chunk: fake/chunk
artifacts:
events:
extract: fake/extract
spells:
extract: dnd/spells
merge: appendorder
normalize: noop
output: json
```
The default LLM profile supplied by `config.Default()` is sufficient. Do not
add real provider settings to this fixture.
### Required Tests
- `config.ParseFileConfigYAML` and `Config.ApplyFileConfig` load the fixture.
- `Config.Resolve` succeeds with pipeline ID `seriatim-fixture` and the
- `Config.Resolve` succeeds with pipeline ID `dnd-spells-fixture` and the
test-only catalog.
- The resolved pipeline input module is `seriatim`.
- The resolved artifact lane ID is `spells` and extractor module is
`dnd/spells`.
- The resolved pipeline digest is non-empty and stable across repeated
resolution.
- Removing `transcript.timestamps` from the Seriatim module spec in the
test-only catalog causes resolution to fail with a missing capability error.
- Removing `source.transcript` from the Seriatim module spec causes resolution
to fail with a missing capability error for `dnd/spells`.
- Removing `dnd.spell_casts` from the extractor spec causes resolution to fail
with a missing capability error for `appendorder`.
- Selecting an unknown `--only` lane still fails through existing resolution
behavior.
- Runner success path:
- parse the Seriatim fixture;
- use a fake LLM that returns at least two valid spell casts;
- assert approved artifacts preserve response order;
- assert approved payloads contain spell data;
- assert approved artifact source refs validate with `source.ValidateRef`;
- assert the manifest records input module `seriatim`, extractor
`dnd/spells`, lane `spells`, and validation status `approved`.
- Runner rejection path:
- fake LLM returns a spell cast with an invalid source ref;
- runner completes with rejected artifacts and validation status `rejected`;
- rejection contains validator `dnd/spells/source_refs` and reason code
`invalid_source_ref`.
- Runner malformed-output path:
- fake LLM returns or reports malformed structured output;
- runner returns an extraction error with D&D spells context and failed
validation status.
### Validation
Run:
```sh
gofmt -w internal/modules/input/seriatim
go test ./internal/modules/input/seriatim
go test ./internal/core/config
go test ./...
```
## Stage 4: Runner Integration With Fake Downstream Stages
### Goal
Prove real Seriatim input can flow through the existing runner into fake
downstream stages while preserving source-unit IDs and metadata.
### Files To Add Or Update
- `internal/modules/input/seriatim/runner_test.go`
### Required Behavior
Use the same Seriatim fixture from stage 2 and a resolved pipeline from stage 3.
Register fake downstream stages only inside the test.
The fake extractor should:
- inspect the received `SourceDocument` and `SourceChunk`;
- assert that unit IDs match Seriatim segment IDs;
- assert that speaker and timestamp metadata are present;
- return one generic artifact candidate with a source reference pointing at
existing Seriatim-derived unit IDs.
The test should then assert:
- `Runner.Run` succeeds;
- the manifest records input module `seriatim`;
- the manifest source digest equals the parsed document digest;
- approved artifacts preserve valid source references;
- no transcript-specific type has been added outside the module.
### Required Tests
- successful runner execution from Seriatim JSON through fake chunk, extract,
merge, normalize, and output stages;
- runner failure when the Seriatim adapter returns an invalid source document,
using a malformed fixture or test input;
- validation of the fake candidate's source reference with
`source.ValidateRef`.
### Validation
Run:
```sh
gofmt -w internal/modules/input/seriatim
gofmt -w internal/modules/extract/dnd/spells
go test ./internal/modules/extract/dnd/spells
go test ./internal/modules/input/seriatim
go test ./internal/framework/pipeline
go test ./...
```
## Stage 5: Documentation And Final Verification
## Stage 7: Documentation And Final Verification
### Goal
Document the implemented Seriatim integration contract once the module exists,
without describing unimplemented D&D extraction or run-command behavior.
Document the implemented D&D spells extraction contract without describing
unimplemented extractors or a `notarius run` command.
### Files To Add Or Update
- `docs/integrations/seriatim.md`
- `docs/roadmap/5-seriatim-input-module.md`
- `docs/integrations/dnd-spells.md`
- `docs/roadmap/6-dnd-spells-extractor.md`
### Required Documentation
Create `docs/integrations/seriatim.md` as implemented-behavior documentation
Create `docs/integrations/dnd-spells.md` as implemented-behavior documentation
with:
- accepted minimal JSON shape;
- required fields and validation rules;
- mapping from Seriatim fields to `SourceDocument` and `SourceUnit`;
- metadata key conventions for `speaker`, `start`, and `end`;
- capability strings declared by the module;
- note that broader Seriatim schema variants are not yet supported.
- module key, artifact type, schema version, prompt ID, and response schema key;
- accepted source expectations: generic source document/chunk with transcript
capability supplied by pipeline resolution;
- spell payload fields;
- source-reference behavior: LLM response includes refs, durable artifact
output carries refs in the generic artifact envelope;
- default deterministic validators and rejection reason codes;
- declared required/provided capabilities;
- note that item, NPC, combat, encounter, broad D&D rules, and CLI `run`
workflows are not implemented by this checkpoint.
Update `docs/roadmap/5-seriatim-input-module.md` only if implementation
reveals a real scope or policy correction. Keep future D&D extraction behavior
out of the integration doc.
Update `docs/roadmap/6-dnd-spells-extractor.md` only if implementation reveals
a real scope or policy correction. Keep implementation staging in this file,
not in the feature roadmap.
### Final Validation
Run:
```sh
gofmt -w internal/modules/input/seriatim
gofmt -w internal/modules/extract/dnd/spells internal/framework/llm internal/framework/prompt
go test ./...
go build ./cmd/notarius
rm -f ./notarius
@@ -427,17 +586,19 @@ rm -f ./notarius
- `go test ./...` passes.
- `go build ./cmd/notarius` passes.
- Seriatim minimal transcript JSON maps into `SourceDocument`.
- Unit IDs are stable and validate in source references.
- Transcript fields do not appear in core runner contracts.
- The input module is selectable through the input registry and pipeline-profile
resolution.
- The input module declares transcript-oriented flat capabilities for pipeline
validation.
- Tests prove transcript-specific assumptions are isolated to
`internal/modules/input/seriatim`.
- The D&D spells extractor is registered through the extractor registry.
- Pipeline-profile resolution can select a `spells` artifact lane using
extractor module `dnd/spells`.
- Seriatim minimal transcript JSON can flow through the runner into the D&D
spells extractor in tests.
- Spell artifacts include valid generic source references in the artifact
envelope.
- Invalid source references are rejected by the extractor-owned validator chain.
- D&D concepts do not appear in core runner, source, pipeline, config, prompt,
LLM, or validator contracts.
- No CLI `run` behavior is documented or implemented in this checkpoint.
## Open Questions
None. This plan chooses the checkpoint-5 behavior needed to implement the
None. This plan chooses the checkpoint-6 behavior needed to implement the
feature without requiring additional product decisions.