Files
notarius/docs/roadmap/dnd-combat-turn-extraction.md

9.3 KiB

D&D Combat-Turn Extraction

Status: Accepted.

Purpose

Add a production D&D combat-turn pipeline that converts session transcripts into an ordered, evidence-backed account of combat activity. This is the next recommended increment because spell and NPC extraction now provide the two most important grounding vocabularies, while combat turns are the next useful artifact explicitly identified by the sequential-pipeline strategy.

The feature uses the existing fixed pipeline and scene chunks for a more structurally demanding artifact without requiring automatic workflow composition, richer scene routing, or generic semantic deduplication.

Target Outcome

An operator can:

  1. run the NPC pipeline against a transcript;
  2. run a combat-turn pipeline over the same transcript, explicitly binding the normalized NPC artifact as a reference;
  3. receive an ordered JSON list of validated combat turns and interrupting combat events; and
  4. trace every reported declaration and immediate resolution back to the transcript units that support it.

The runs remain independent CLI invocations. Notarius does not discover prior outputs, schedule dependent pipelines, or reconcile spell and combat artifacts automatically.

Combat-Turn Artifact

The new typed artifact represents an ordered list of combat-turn records. Each record contains:

  • the canonical in-world actor;
  • a turn kind distinguishing an ordinary turn from a reaction, legendary action, lair action, or other interrupting combat event;
  • the combat round as a positive integer when it is explicit or unambiguous, and null otherwise;
  • one or more ordered actions;
  • a concise turn-level summary; and
  • one or more transcript source references that collectively support every reported field.

Each action contains a conservative category, a concise declaration, zero or more targets, and its immediate observed resolution. Resolution is null when the cited passage establishes the declaration but no immediate resolution. Categories cover attacks, spells, movement, items, ability checks, saving throws, condition or state changes, and an other fallback without requiring the transcript to use formal rules terminology.

Reactions and similar out-of-turn events appear at the point where they occur in transcript chronology rather than being moved to the reacting creature's later turn. Output order is derived from cited source position; numeric source unit IDs are identifiers, not chronology.

Inclusion And Evidence Policy

Include a record when the transcript establishes that an in-world participant takes a combat turn or performs a discrete interrupting combat event. Report only declarations and their immediate resolutions, including directly associated rolls, damage, healing, movement, conditions, or target outcomes. For every detail reported, cite all supporting transcript units.

Exclude:

  • initiative setup that contains no turn or combat event;
  • tactical planning, table talk, rules lookup, and hypothetical actions;
  • corrected or abandoned declarations that never become an attempted action, except where the correction is necessary to describe the final declaration;
  • recap of combat that occurred outside the current source passage; and
  • downstream consequences that occur on later turns or elsewhere in the scene.

Do not infer a round number, action-economy classification, target, roll, amount, condition, or outcome merely from D&D rules knowledge. Preserve the session as played, and attribute nonstandard rulings to the GM or table when that detail is relevant to the immediate resolution.

Identity And Reference Grounding

The extractor uses the existing D&D transcript, player, party, and glossary prompt inputs. It also accepts the normalized NPC artifact through an optional structured npcs reference slot with the same validation, size, provenance, content-safety, and semantic-checkpoint rules used by spell extraction.

The NPC registry helps select canonical actors and targets and recognize aliases. It does not establish that combat occurred and never becomes source evidence. Unmatched actors and targets remain permitted because a session may introduce combatants that were omitted from an earlier NPC run.

The deterministic normalizer also accepts the registry. Exact canonical-name or alias matches are rewritten to the registry's canonical display name for actors and targets; ambiguous aliases and unmatched values remain unchanged for validation and human review. Opaque player and party references continue to guide the LLM but are not parsed into a new roster contract in this scope.

Extraction, Validation, And Normalization

The extractor uses one structured LLM call per supplied chunk and returns typed combat-turn candidates. It must preserve malformed candidates for the normal validation and retry boundary rather than silently repairing unsupported content in mapping code.

The production default validator chain is deterministic and covers:

  • required fields, arrays, nullable-round shape, and supported enum values;
  • a required non-empty evidence collection, source identity, unit existence, and range order;
  • actor and declared-action relatedness to cited transcript text, expressed as bounded warnings where deterministic substring checks are only advisory; and
  • normalized identity and duplicate invariants.

Normalization is deterministic and conservative. It normalizes display whitespace, canonicalizes exact NPC identity matches, orders and deduplicates exact source references, and collapses only exact duplicate records with the same normalized actor, turn kind, round value, and complete valid evidence set. The first record is retained without synthesizing or merging prose. Every mutation or collapse emits a scoped warning.

An LLM-backed validator and semantic reconciliation normalizer are outside the production chain. Human evaluation owns judgments such as whether the model grouped a long turn correctly or omitted a subtle reaction.

Scene Strategy

The combat extractor processes every chunk delivered by the configured chunker. Existing D&D scene annotations remain useful context, but a primary_mode value does not suppress an LLM call. Avoiding a call based on an imperfect non-combat classification could silently lose the very turns this artifact is intended to recover.

Scene-classification and routing improvements remain separate future work. The combat pipeline stays compatible with generic chunks that carry no D&D annotation.

Provenance And Checkpoints

The extractor, normalizer, and validators report stable semantic identities through the existing manifest and prepared-component fingerprint contracts. Prompt, private response schema, artifact policy, normalization policy, and a bound NPC registry's semantic digest must invalidate incompatible checkpoints.

Metadata and fingerprints contain identities, counts, and digests only. They must not contain transcript text, combat records, NPC names, reference paths, or raw reference content. Preparation failures remain bounded and content-safe and occur before checkpoint handlers or pipeline execution are constructed.

Evaluation

Use human-reviewed development runs rather than exact model-output goldens. Evaluate at least the existing transcripts used for spell and NPC development, with separate attention to:

  • combat-turn detection precision and recall;
  • actor and target identity;
  • turn boundaries and chronological order;
  • reactions and other interrupting events;
  • declaration and immediate-resolution fidelity;
  • round-number restraint;
  • completeness and precision of evidence; and
  • duplicate behavior at chunk or scene boundaries.

Frontier and inexpensive development models may differ substantially in semantic quality. Deterministic tests should protect structure, provenance, identity, ordering, normalization, and orchestration rather than require exact combat prose.

Out Of Scope

  • Initiative trackers, current hit points, complete encounter-state replay, or rules-engine validation.
  • Automatic comparison or reconciliation with spell artifacts.
  • NPC discovery, campaign-wide entity persistence, or deterministic PC-roster parsing.
  • LLM-backed validation or generic LLM-assisted deduplication.
  • Automatic pipeline scheduling, prior-output discovery, or DAG execution.
  • Scene-classification changes or skipping provider calls for non-combat chunks.
  • Narrative summaries outside the immediate combat-turn scope.

Acceptance Criteria

  • A selectable D&D combat lane produces a typed, durable JSON artifact with the turn, action, chronology, and evidence semantics above.
  • The default deterministic validator chain rejects malformed or invalidly sourced records and emits bounded advisory relatedness warnings.
  • Deterministic normalization canonicalizes exact NPC identities and collapses only safely identical combat records while preserving chronology and evidence.
  • A normalized NPC artifact can be bound explicitly to both extraction and normalization without becoming combat evidence.
  • Semantic contracts and structured references participate in checkpoint identity without leaking application content.
  • Maintained configuration and operational examples demonstrate independent NPC and combat invocations over the same transcript.
  • Human review on representative sessions demonstrates useful turn extraction without requiring scene-based call suppression or semantic reconciliation.