diff --git a/docs/roadmap/dnd-npc-semantic-normalization.md b/docs/roadmap/dnd-npc-semantic-normalization.md deleted file mode 100644 index e7e4899..0000000 --- a/docs/roadmap/dnd-npc-semantic-normalization.md +++ /dev/null @@ -1,328 +0,0 @@ -# D&D NPC Semantic Normalization - -## Status - -Implemented. - -## Purpose - -Reconcile NPC records that extraction produced under different display names -when the complete transcript establishes that they represent the same -individual. This work addresses identities split across scene or chunk -boundaries while preserving the existing rule that extraction records narrow -source evidence and normalization owns document-wide reconciliation. - -The model should make only the semantic identity judgment. Deterministic code -must continue to own identity derivation, proposal validation, artifact -mutation, evidence preservation, ordering, diagnostics, and final validation. - -## Target Behavior - -The target `dnd/npcs` normalization contract combines: - -- the current deterministic display-name normalization, stable-ID derivation, - source-reference canonicalization, and equal-comparison-key consolidation; -- one LLM determination of whether remaining, distinctly named NPC records - represent the same individual and which existing display name is canonical; - and -- deterministic proposal validation and application followed by the configured - normalize validator chain. - -Each configured normalization attempt should make at most one semantic pass -over the merged document-level NPC list, not one pass per extraction chunk or -candidate pair. The pass should be skipped when fewer than two distinct -candidate identities remain after deterministic preprocessing. - -False consolidation is more damaging than a missed consolidation. Prompt -policy should therefore require affirmative contextual evidence that names -identify the same individual and should prefer no group when identity remains -ambiguous. - -Canonical selection should favor the most complete stable proper name supported -by the transcript. A complete proper name is preferable to an abbreviation, -while an unadorned proper name is preferable to the same name plus a contextual -class, role, title, or relationship descriptor unless the transcript establishes -that descriptor as part of the character's name. The model must still select -one supplied display name rather than synthesize a better one. - -## Model Proposal Contract - -The private structured response should contain only proposed duplicate groups: - -```json -{ - "duplicate_groups": [ - { - "members": [ - "Billy", - "Billy the druid" - ], - "canonical_name": "Billy" - } - ] -} -``` - -`members` identifies candidates by their supplied display names. -`canonical_name` selects one existing member as the retained display name. An -empty `duplicate_groups` array is a valid determination that no sufficiently -supported duplicates exist. - -The model must not receive or return deterministic `npc:sha256:` IDs. Those IDs -are long, non-semantic implementation identities and remain exclusively owned -by deterministic code. The model must not invent a replacement name, rewrite -an NPC record, propose new source references, or return a complete replacement -artifact. - -Before the call, deterministic preprocessing ensures that each supplied display -name has a unique NPC identity comparison key. Returned names may resolve using -the existing comparison-key equivalences for case, whitespace, Unicode -compatibility, and supported apostrophes. Resolution must not use fuzzy, -substring, edit-distance, embedding, or other approximate matching. - -An applicable group must: - -- contain at least two distinct, known members; -- resolve every member uniquely against the supplied candidate set; -- select a known `canonical_name` that belongs to the group; -- contain no repeated member; and -- share no resolved member with any other proposed group. - -Unknown or ambiguous names, singleton groups, invalid canonical selections, and -other unsafe groups must not mutate the artifact. Groups are otherwise -independent: a locally valid group may be applied when none of its resolved -members appears in any other proposed group. - -Overlap is a proposal conflict even when one participating group is already -locally invalid. Deterministic code must discard every group in the connected -conflict set rather than selecting a winner by response order. Locally valid, -non-conflicting groups remain safe to apply. This permits useful partial -reconciliation without allowing an unsafe group to influence an NPC identity -that another group would mutate. - -## Retry And Safe Fallback - -An invalid or partially unsafe private proposal is a retryable normalization -attempt. This includes: - -- structurally invalid model output classified by the LLM boundary as an - invalid structured completion; and -- a structurally valid response containing any locally invalid or conflicting - group. - -The framework owns retry counting and attempt diagnostics. The normalizer -returns a safe candidate together with a bounded, content-safe retry -diagnostic. For structurally invalid output, that candidate is the -deterministic pre-LLM result. For a decoded proposal with unsafe groups, it -also includes every independently valid, non-conflicting group from that -attempt. When normalize retries remain, the framework invokes the normalizer -again from the same merged input; safe groups are not accumulated across -attempts. A later completely safe proposal is applied normally. - -When the configured attempt budget is exhausted, the framework validates and -accepts the final attempt's safe fallback instead of failing or rejecting the -NPC lane, provided that fallback passes the configured normalize validators. -It promotes one durable warning stating that one or more proposed groups were -omitted. The fallback may therefore contain a safe partial reconciliation, or -only the deterministic pre-LLM result when no group could be applied. Earlier -retry diagnostics, candidate values, and warnings remain attempt-local in -debug artifacts. - -The normal module-binding default remains `retries: 0`, meaning one total -normalization attempt and immediate fallback after its invalid proposal. -Retries occur only when the user configures a positive normalize retry count. -For example, `retries: 2` permits the initial proposal plus two additional -attempts before fallback. - -Transport, authentication, prompt-preparation, cancellation, input-encoding, -and other operational errors are not safe-proposal failures. They retain the -existing module/framework error behavior rather than being converted into an -accepted fallback. - -## Transcript Context Policy - -The model should receive every candidate's display name and source references, -together with transcript windows derived from those references. It should not -receive the complete transcript by default. - -Each window must contain: - -- the complete inclusive source range cited by the NPC record; -- up to two source units immediately before the cited range; and -- up to two source units immediately after the cited range. - -The surrounding-unit count is defined once as a named module policy constant -with value `2`; it is not a user-accessible configuration field. Window -construction consumes that value through one clear boundary so a later -configuration option can replace the fixed value without changing prompt or -reconciliation contracts. The context-radius policy participates in module -metadata and checkpoint identity so changing it invalidates incompatible -normalize checkpoints. - -The value was selected by reviewing the July 19 evaluation transcript. Its NPC -citations are generally self-contained; where additional context is useful, -two surrounding units capture the relevant question-and-answer exchange, -speaker transition, or short anaphora chain. A third unit frequently begins a -separate joke or table exchange and adds distraction without improving the -identity evidence. - -Window expansion must use source-document positions rather than arithmetic on -unit IDs. It must clamp at document boundaries, preserve document order, and -coalesce overlapping or adjacent expanded windows without duplicating units. -The prepared model input must distinguish originally cited units from -surrounding context. - -Only records with a non-empty NPC identity comparison key and a non-empty -source-reference collection that is wholly valid against the current source -document are eligible for semantic reconciliation. Ineligible records remain -unchanged so the configured deterministic validators retain ownership of their -rejection. They are not supplied to the model and cannot participate in a -proposed group. - -Surrounding units inform the semantic decision but do not automatically become -durable NPC evidence. An applied group unions only the source references -already present on its member records. The LLM cannot add references from the -context window. - -Full-transcript mode and a configurable context radius may be evaluated later. -They are not part of this scope. Provider prompt-cache behavior should be -measured rather than assumed before expanding context solely for cache -economics. - -## Deterministic Application - -For each approved group, deterministic code should: - -- retain the model-selected existing display name; -- union source references from every group member; -- canonicalize and exact-deduplicate those references in source-document order; -- derive the resulting stable NPC ID from the retained display name under the - existing NPC identity policy; and -- emit one bounded duplicate-collapse warning describing the applied group. - -The consolidated record should occupy the earliest original member position so -model response ordering cannot reorder the artifact. Unrelated records must -remain present and retain their relative order. Every input record must be -represented by exactly one output record, either unchanged or through one -approved consolidation. - -The existing NPC shape, identity, source-reference, schema, and relatedness -validators remain the final artifact boundary. No LLM-backed validator is -needed: model judgment occurs in the normalizer, while deterministic validators -continue to enforce the durable artifact contract. - -## Prompt, Provenance, And Diagnostics - -NPC semantic-normalization instructions and the private response schema should -be module-owned prompt assets. Stable D&D-wide identity or transcript guidance -may be reused through the existing shared prompt-asset mechanism where its -meaning is genuinely common. - -The normalizer should use the configured normalize-stage LLM profile and the -application-wide scheduled LLM client. Invalid structured output and unsafe -semantic proposals use the framework-owned retryable-fallback contract; -operational failures retain existing error semantics. - -Manifest metadata and component checkpoint fingerprints should identify the -semantic-normalization policy, prompt identity, private response-schema -identity, deterministic NPC identity policy, and evidence-context policy. -They must not contain transcript text, NPC names, source paths, raw model -responses, or other source content. - -Warnings and preparation or runtime errors must follow the established bounded -diagnostic and content-safety policies. Debug artifacts may retain the normal -attempt-local model request, response, proposal decisions, and warnings under -the existing debug sensitivity contract. - -## Configuration And Documentation - -The maintained complete D&D configuration should demonstrate a normalize-stage -LLM profile and an explicit positive retry count for the NPC lane. The -application-wide default remains zero additional retries. No new configuration -field is introduced for the context radius in this scope. - -Canonical documentation ownership is: - -- Configuration owns the LLM-backed `dnd/npcs` selection and its use of the - existing normalize binding's profile and retry fields. -- Operations owns the document-level semantic reconciliation call and its - checkpoint behavior. -- Pipeline internals own the provider-neutral invalid-structured-output - classification and normalize retryable-fallback contract. -- The NPC integration contract owns externally observable consolidation, - evidence, ordering, and warning behavior. -- Internal LLM and module documentation own the proposal boundary, - context-window construction, deterministic application, metadata, and - fingerprints. -- The broader generic LLM-assisted deduplication item in - [future.md](future.md) remains future work until another artifact demonstrates - that extracting a shared generic facility is worthwhile. - -## Quality Expectations - -Tests should protect the semantic and safety boundaries through deterministic -LLM fakes rather than live-provider calls. Coverage should demonstrate: - -- distinct display variants can be consolidated when the model proposes a - valid group; -- model-facing requests and responses use display names rather than stable - hash IDs; -- the selected existing canonical name controls ID derivation while all member - evidence is preserved; -- window construction uses source-document position, handles document edges, - coalesces overlap, and distinguishes cited evidence from context; -- no surrounding context is promoted into durable source references; -- empty proposals and fewer-than-two-candidate inputs preserve deterministic - normalization behavior; -- unknown, ambiguous, repeated, overlapping, and otherwise malformed semantic - groups cannot corrupt or reorder the artifact; -- locally valid groups are applied independently, while every group that - participates in a resolved-member conflict is discarded; -- an invalid proposal consumes only configured retry budget, a later valid - proposal can succeed, and exhaustion accepts the final attempt's safe - fallback with one durable warning; -- the default zero-retry binding makes exactly one proposal attempt before - fallback; -- warnings and failures remain bounded and do not expose transcript content; -- prompt, schema, policy, or context-policy changes invalidate relevant - checkpoint reuse; and -- an assembled ordered D&D pipeline supplies the reconciled NPC registry to - downstream consumers. - -Prompt tests should assert prepared message structure, supplied materials, and -cache-boundary behavior. They must not act as change detectors for particular -words or phrases in natural-language prompt text. - -Human review of representative transcripts should compare missed and false -consolidations, latency, input-token cost, and provider cache use. Probabilistic -model quality is an evaluation activity, not a deterministic CI assertion. - -## Non-Goals - -This scope does not: - -- change the durable NPC artifact schema; -- add a durable alias collection or identity history; -- expose the context radius or full-transcript selection as configuration; -- ask an LLM to read, reproduce, or derive stable NPC IDs; -- permit model-authored source references or arbitrary replacement records; -- add fuzzy deterministic name matching; -- add an LLM-backed normalize validator; -- change the global default retry count; -- reconcile NPCs concurrently across lanes or runs; -- add canonical NPC IDs to downstream artifact schemas; -- implement a general DAG or implicit cross-lane dependency; or -- implement the generic cross-artifact deduplication normalizer described in - `future.md`. - -## Completion Criteria - -The scope is complete when `dnd/npcs` can use one document-level LLM proposal -to reconcile differently named records conservatively, deterministic code -validates and independently applies only safe non-conflicting name-based -groups, the resulting NPC preserves all member evidence under a newly derived -stable ID, the two-unit context policy is centralized and fingerprinted, -invalid or partially unsafe proposals use configured framework retries and -then an accepted safe fallback, downstream ordered steps receive the -reconciled registry, and the canonical current-behavior documentation reflects -the implemented contract. diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md deleted file mode 100644 index f55c906..0000000 --- a/docs/roadmap/implementation.md +++ /dev/null @@ -1,365 +0,0 @@ -# D&D NPC Semantic Normalization Follow-up Plan - -## Status - -Ready for implementation. - -## Objective - -Close the four findings from the post-implementation review of -[D&D NPC Semantic Normalization](dnd-npc-semantic-normalization.md) without -changing its durable artifact schema, deterministic identity rules, -conflict-aware partial-application policy, retry accounting, or generated -reference handoff. - -Complete the four follow-up stages below in order. Each stage should leave its -affected packages passing before the next stage begins. - -## Completed Work Summary - -The original five-stage implementation is complete: - -1. **Framework retryable normalize fallback.** The framework now classifies - invalid structured output separately from operational LLM failures. - Normalizers can return a safe candidate and retry directive; the runner owns - retry accounting, attempt debug artifacts, final fallback validation, - accepted-only warning promotion, and checkpoint recording. -2. **NPC prompt and context material.** `dnd/npcs` owns a private normalization - prompt and structural response schema. It sends display-name candidates and - coalesced transcript windows with a fixed two-unit radius, without exposing - stable NPC IDs or using the full source input. -3. **Semantic proposal validation and application.** The normalizer performs - deterministic preprocessing, resolves returned names through the existing - comparison-key policy, applies independent safe groups, discards locally - invalid and conflicting groups, preserves evidence and ordering, and emits - bounded D&D diagnostics. -4. **Production integration and checkpoint identity.** Production registration - includes the normalize prompt and schema. Prompt, schema, identity, - normalization, and context policies contribute checkpoint fingerprints. - Assembled tests prove retry behavior and the ordered normalized-NPC handoff - to downstream consumers. -5. **Configuration and documentation.** The maintained complete D&D example - demonstrates an explicit normalize profile and retry count. Configuration, - operations, integration, LLM, module, pipeline, and component documentation - describe the implemented feature. - -The completed implementation passed the full test suite, vet, build, diff -checks, and the targeted race suite. The stages below are limited follow-up -corrections and do not reopen the original feature design. - -## Decisions Applying To Every Follow-up Stage - -- Retain module key `dnd/npcs`, artifact kind `dnd/npc-list`, durable schema - `notarius.dnd.npcs` v1, prompt ID `dnd.npcs.normalize`, response-schema - identity, empty reference-slot contract, strict empty options, and default - normalize validator chain. -- Retain the context radius of two and the existing candidate and transcript - input shapes. -- Retain conflict-aware partial application: independently safe groups may be - applied, every group participating in a resolved-member conflict is - discarded, retries restart from the original merge input, and final - exhaustion accepts only the final attempt's safe candidate. -- Retain `retries: 0` as the application-wide default. Changing that default - remains separate work. -- Retain accepted-only warning promotion, checkpoint format, public - configuration shape, and durable output JSON. -- Do not add an LLM-backed validator, normalizer-local retry loop, generic - deduplication framework, prompt-language change-detector test, or live - provider test. - -## Stage 1: Make Context-Material Failures Content-Safe - -### Goal - -Prevent arbitrary source metadata keys or other source-derived details from -appearing in errors returned by the NPC normalization context builder. - -### Implementation - -Keep `source.CloneMetadata` unchanged. Its detailed errors remain useful to -trusted direct callers and its own package tests. Redaction belongs at the -external source-material consumption boundary in -`internal/modules/dnd/normalize/npcs/context_material.go`. - -At that boundary: - -- When copying source-unit metadata fails, return a fixed category such as - `build NPC normalization context: invalid source metadata`. -- Do not wrap the underlying clone error with `%w`. -- The returned error must not include metadata keys, values, traversal - locations, transcript text, NPC names, source IDs, origin paths, encoded - material, concrete Go types derived from source metadata, or the underlying - error text. -- Apply the same fixed-error treatment to candidate and transcript JSON - encoding failures. Those paths are effectively defensive after metadata - validation, but future DTO changes must not create a content-exposure path. -- Keep the negative-radius error detailed: the radius is code-owned policy, - not source content. -- Keep failure timing unchanged. Context construction still fails before the - normalization LLM call, retry directive creation, artifact validation, or - checkpoint recording. - -The outer `Normalize` error may retain its existing operation context because -the nested context-builder error is fixed and content-safe. - -### Tests - -Add narrow boundary coverage in the NPC normalization package: - -- Use source metadata with a recognizable sensitive key and an unsupported or - non-finite value. -- Assert that normalization fails before the LLM fake is called. -- Assert that the error contains a stable context-material failure category - but none of the recognizable key, value, transcript text, NPC names, source - ID, or origin path. -- Retain direct `source.CloneMetadata` tests as the owner of detailed cloning - behavior; do not duplicate its type matrix at the normalizer boundary. - -Do not assert the complete error sentence. Protect the fixed category and -absence of source-derived sentinels. - -### Completion Gate - -Every context-material construction failure exposed by `dnd/npcs` is -content-safe, while trusted source-package callers retain detailed clone -diagnostics. - -## Stage 2: Remove Conflicting Shared Identity Prompt Guidance - -### Goal - -Give the NPC normalizer one unambiguous canonical-name policy and avoid telling -it about campaign references that it does not receive. - -### Implementation - -The existing `common-dnd-identity.md` fragment serves extraction prompts: it -asks for the most specific supported in-world identity and discusses campaign -and registry references. It is not wholly applicable to reconciliation of an -already extracted NPC list. - -Remove `common-dnd-identity.md` from the NPC normalization prompt only: - -- Remove its message from - `assets/prompts/dnd.npcs.normalize.yaml`. -- Remove it from the normalization package's `PromptAssetManifest.SharedFiles`. -- Do not change the shared fragment or any extraction prompt that currently - consumes it. -- Do not replace it with a copied or newly shared identity fragment. NPC - extraction has already established that the candidates are in-world NPC - records; normalization needs only the package-owned duplicate and canonical - selection instructions. -- Retain `common-dnd-system.md` and `common-dnd-transcript.md`. - -After removing the identity message, preserve two intentional prompt-cache -tiers: - -1. Mark the common D&D system message as the end of the stable shared tier. -2. Retain the existing boundary after the package-owned normalization - instructions. - -The package-owned instructions remain authoritative: - -- consolidate only when transcript context clearly identifies one individual; -- select a supplied canonical name; -- prefer an unadorned proper name over the same name plus a contextual class, - role, title, or relationship descriptor unless established as part of the - name; and -- prefer no consolidation when identity is ambiguous. - -Do not add references, stable IDs, new prompt inputs, new natural-language -repair material, or schema constraints. The prompt digest will change and -prior NPC normalize checkpoints will intentionally become cold misses. Keep -the existing prompt ID and v1 version because checkpoint compatibility is -already content-fingerprinted and no external prompt protocol is versioned by -this private asset. - -Update the canonical internal LLM documentation to describe the normalization -prompt's actual shared tier and to stop claiming that it uses shared identity -guidance. - -### Tests - -Update the existing offline prepared-prompt test to prove: - -- the prompt prepares with five messages in the intended order; -- the first cache boundary follows the common D&D system message; -- the second follows the package-owned instructions; -- candidate and transcript messages remain in the variable tail without cache - boundaries; -- only the two declared dynamic inputs are rendered in their corresponding - messages; and -- prompt metadata and fingerprint construction remain valid. - -Tests may assert asset/message identity, ordering, inputs, and cache-control -structure. They must not require the presence or absence of particular words -or phrases in natural-language prompt content. - -### Completion Gate - -The normalization prompt contains no extraction-specific identity or reference -guidance, preserves its intended cache structure, and still carries one clear -package-owned canonical-name policy. - -## Stage 3: Enforce Generic Normalize-Retry Diagnostic Bounds - -### Goal - -Make the framework-level `NormalizeRetry` contract safely reusable by enforcing -the mechanical parts of its diagnostic contract at the runner boundary. - -### Contract - -Add provider-neutral limits in `internal/framework/contracts`: - -- maximum retry reason-code length: 128 bytes; -- maximum retry message length: 4,096 bytes; and -- both values must be valid UTF-8 and nonblank after trimming surrounding - whitespace for the blank check. - -Expose the limits as named constants alongside `NormalizeRetry` so normalizer -authors can construct compliant diagnostics without importing a D&D package. -Do not import `internal/modules/dnd/shared/diagnostics` into the framework. - -These are encoded-byte limits, not rune limits. The framework validates but -does not truncate or rewrite caller-supplied diagnostics. Silent truncation -could collapse stable reason identities or conceal a module defect. - -Content safety remains the module's semantic responsibility: the framework -cannot determine whether otherwise valid text contains transcript content, -credentials, paths, names, or other sensitive values. - -### Runner Implementation - -Extract the current inline blank check into a small framework-owned validator -for `*contracts.NormalizeRetry` and call it before placing the directive in a -debug payload. - -The validator must reject: - -- a blank reason code; -- a blank message; -- invalid UTF-8 in either field; -- a reason code exceeding 128 bytes; and -- a message exceeding 4,096 bytes. - -Return a fixed module-contract error that identifies the invalid field and -failure category without echoing either supplied value. Preserve existing -retry-loop behavior for module contract errors; do not create a new error -class, rejection, or durable warning. - -Do not impose a generic count or text-size policy on `FallbackWarnings` in this -stage. Warning limits remain owned by their existing domain and validator -contracts; broad warning-policy unification is separate work. - -Update `docs/internal/pipeline.md` to document the normalize-retry limits and -the division of responsibility between mechanical framework validation and -module-owned content safety. - -### Tests - -Extend the generic runner retry tests with table-driven boundary cases: - -- a reason code and message exactly at their byte limits are accepted; -- one byte over either limit is rejected; -- invalid UTF-8 in either field is rejected; -- blank fields remain rejected; and -- errors contain only fixed field/category context and do not echo supplied - diagnostic sentinels. - -Use test-controlled strings and the exported contract constants. Keep the -existing NPC diagnostic tests as the owner of D&D aggregation and omission -behavior. - -### Completion Gate - -Any typed normalizer receives the same reusable retry mechanism, and the runner -cannot write a blank, invalid-UTF-8, or unbounded retry diagnostic into debug -state. - -## Stage 4: Correct Safe-Fallback Documentation - -### Goal - -Make canonical pipeline documentation accurately describe the generic -framework contract and the NPC module's model-derived safe partial fallback. - -### Documentation - -Update `docs/internal/pipeline.md`: - -- replace `deterministic safe candidate` with `module-supplied safe candidate` - or equivalent wording; -- state that the framework treats the candidate as opaque and relies on the - normalizer plus the configured validator chain for its safety; -- retain the rules that intermediate candidates are attempt-local, only the - final fallback is validated for acceptance, and accepted-only warnings and - checkpoints remain unchanged; and -- keep the generic contract separate from the NPC-specific conflict and - partial-application rules. - -Review the affected current-behavior documentation for the same inaccurate -determinism claim. The NPC integration and operations documents may describe -deterministic validation and mutation, but must not imply that every safe -partial fallback is independent of the LLM proposal. - -Do not modify the completed feature roadmap except if a factual contradiction -with the implemented target state is discovered. Do not add tests for prose -wording. - -### Verification - -- Validate local documentation links. -- Confirm configuration defaults, module keys, prompt IDs, and schema versions - against current code. -- Run `git diff --check`. - -### Completion Gate - -Canonical documentation distinguishes the generic opaque safe-candidate -contract from deterministic NPC validation and application, with no duplicated -or contradictory retry semantics. - -## Final Verification - -From the repository root, run: - -```sh -git diff --check -go test ./... -go vet ./... -go build ./cmd/notarius -go test -race ./internal/framework/contracts ./internal/framework/llm ./internal/framework/pipeline ./internal/modules/dnd/normalize/npcs ./internal/modules/dnd/register ./internal/modules/integration -``` - -Review the final diff for: - -- no durable artifact, checkpoint-format, or public configuration change; -- no model-facing stable NPC IDs or full-transcript normalization input; -- no source-derived metadata keys, values, paths, names, or transcript content - in context-material errors; -- no extraction-specific shared identity fragment in the normalize prompt; -- no prompt-wording change-detector or live-provider test; -- no framework dependency on D&D diagnostic packages; -- no silent retry-diagnostic truncation; -- no change to conflict-aware partial application, retry accounting, warning - promotion, validator ordering, or downstream handoff; and -- documentation that describes only implemented behavior in its canonical - owner. - -## Assumptions - -- Detailed `source.CloneMetadata` errors remain appropriate inside the trusted - source package; only the NPC context-material boundary requires redaction. -- The normalizer does not need general extraction identity guidance because its - candidates are already accepted NPC artifacts. -- Mechanical UTF-8 and byte limits are appropriate framework invariants; - semantic content safety remains a module obligation. -- Prompt-content and cache-boundary changes intentionally invalidate existing - NPC normalize checkpoints through the current fingerprint mechanism. -- Backward compatibility for development checkpoints is not required. - -## Open Questions - -None. The review findings and decisions above are sufficient to implement the -follow-up work without additional product or architecture choices. diff --git a/docs/roadmap/scene-aware-combat-extraction.md b/docs/roadmap/scene-aware-combat-extraction.md deleted file mode 100644 index 59c8051..0000000 --- a/docs/roadmap/scene-aware-combat-extraction.md +++ /dev/null @@ -1,190 +0,0 @@ -# Scene-Aware Combat Extraction - -## Status - -Implemented. This document preserves the feature contract and rationale; the -current configuration, operational, integration, and internal behavior is -maintained in their canonical documentation. - -## Purpose - -Use an accepted D&D scene-description artifact as the eligibility boundary for -combat-turn extraction. The combat LLM should run only for a chunk whose exact -scene record has `kind: combat`; every other chunk should be ignored by the -combat extractor. This work should connect the existing scene-description lane, -ordered generated-reference handoff, and combat-turn lane without expanding the -minimal combat-turn artifact contract or introducing D&D policy into generic -pipeline code. - -## Target Workflow - -The complete D&D pipeline has two ordered steps: - -1. The first step extracts and normalizes NPCs and scene descriptions for the - accepted scene chunks. -2. At the step barrier, the accepted NPC and scene-description artifacts become - generated references. -3. The second step supplies the NPC artifact to its existing consumers and the - scene-description artifact to combat-turn extraction. -4. For each chunk, combat-turn extraction performs its existing LLM extraction - only for an exact `combat` classification. Every other chunk produces a - deterministic empty artifact without an LLM call. - -The scene-description dependency must remain explicit in pipeline -configuration. The combat extractor must not discover another lane implicitly, -read output files directly, inspect chunker-private state, or require generic -chunk materialization to interpret D&D scene classifications. - -## Scene-Description Reference Contract - -The combat extractor should declare a required structured reference slot named -`scene_descriptions`. The slot accepts one approved -`dnd/scene-description-list` artifact using the existing durable scene -description schema. - -The reference may be supplied as: - -- a generated artifact from an earlier ordered step; or -- an external artifact through the existing reference-materialization - boundary. - -External artifacts must be decoded and validated before source parsing or LLM -execution. Generated artifacts must cross the existing typed step-handoff -boundary and be validated before use. A bound artifact that is malformed, -incompatible, or internally inconsistent is an error; it must not be treated as -though the slot were unbound. - -Pipeline resolution must reject combat extraction when this slot is not bound. -In an ordered same-run workflow, failure of the configured scene-description -producer to yield an accepted normalized artifact must fail the run before the -combat consumer step starts, consistent with existing required generated -reference semantics. - -The prepared reference view should be immutable and safe for concurrent -extract jobs. Its metadata and checkpoint identity should be content-safe and -must not expose scene titles, summaries, paths, or source text. - -## Chunk Matching And Gating Policy - -A scene record authorizes combat extraction for the current chunk only when all -of the following are true: - -- the scene ID exactly equals the current accepted chunk ID; -- the scene source ID exactly equals the chunk source ID; -- the scene start and end unit IDs exactly equal the chunk's inclusive source - range; and -- the approved artifact contains exactly one such record. - -The normalized scene-description contract already rejects conflicting IDs and -ranges. The combat extractor must nevertheless require the exact match above at -its own decision boundary rather than relying on array position, range overlap, -title, summary, or inferred chronology. - -An exact match with `kind: combat` performs the existing combat-turn LLM -extraction. An exact match with `narrative`, `recap`, or `meta` returns a typed -`dnd/combat-turn-list` containing an empty `combat_turns` array without making -an LLM call. - -A valid artifact with no exact match, incomplete coverage, or a chunk identity -or range mismatch also returns the deterministic empty artifact. It must not -infer combat eligibility from an overlapping or adjacent scene. Emit a bounded, -content-safe warning for missing or mismatched coverage so operators can -distinguish an intentional non-combat classification from an unavailable exact -classification. The warning must not include scene prose or transcript text. - -This policy relies on the scene contract's mixed-scene precedence: any chunk in -which combat is a substantive central activity is classified as `combat`. -Scene-aware gating must not add a second classification policy. - -## Extraction, Validation, And Provenance - -The deterministic empty result follows the same typed extractor and validator -boundaries as an LLM-produced empty result. It is not a rejection, does not -consume retry budget, and continues through merge and normalization normally. -The durable combat-turn schema remains unchanged. - -Existing combat extraction behavior—including prompt assets, NPC grounding, -candidate mapping, validators, retries, warnings, and normalization—remains -unchanged for chunks classified as combat. The scene-description artifact is -control context only: - -- it must not be added to the combat prompt; -- its title or summary must not become combat evidence; -- its source references must not be copied into combat turns; and -- it must not create, repair, or classify a combat turn. - -Generated-reference provenance and dependency fingerprints should cover the -scene artifact through the existing ordered-handoff machinery. External -references should contribute their existing materialization provenance plus a -component-local semantic fingerprint sufficient to invalidate combat extract -checkpoints when a classification or chunk identity changes. Checkpoint reuse -must never preserve a skipped result after the effective scene classification -changes. - -Run manifests may report bounded module metadata such as the number of approved -scene records. They must not contain scene prose or duplicate the referenced -artifact payload. - -## Configuration And Documentation - -The maintained complete D&D example should bind the normalized -`scene-descriptions` lane from the first step to the `scene_descriptions` slot -in the second step. The minimal example should remain unchanged. - -Current-behavior documentation is maintained in these canonical locations: - -- Configuration owns the new selectable reference slot and binding example. -- Operations owns the ordered scene-aware workflow. -- The combat-turn integration contract owns externally observable extraction, - empty-result, and provenance behavior. -- Internal pipeline and module documentation own preparation, handoff, matching, - and checkpoint mechanics. -- `future.md` should remove the completed scene-aware combat item. - -## Quality Expectations - -Tests should protect behavior and architectural boundaries rather than internal -helper shape. Coverage should demonstrate: - -- exact non-combat matches produce accepted empty combat artifacts without an - LLM call; -- exact combat matches retain the existing LLM path; -- an unbound required reference is rejected during pipeline resolution; -- missing and mismatched chunk coverage produces an accepted empty combat - artifact without an LLM call and emits a bounded warning; -- malformed external and generated artifacts fail at the appropriate - preparation or handoff boundary; -- retries are neither consumed nor attempted for deterministic skips; -- scene artifact changes invalidate relevant checkpoint reuse; -- no scene prose is exposed through combat prompts, warnings, metadata, or - manifests; and -- the complete maintained configuration resolves and materializes the intended - ordered dependency. - -Model-output fixtures should remain deterministic test doubles. This feature -does not require live-provider tests or assertions over probabilistic model -quality. - -## Non-Goals - -This scope does not: - -- change the durable scene-description or combat-turn schemas; -- add scene fields to combat-turn artifacts; -- annotate generic chunks with D&D classifications; -- make the combat lane depend implicitly on the scene-description lane; -- skip spell, NPC, NPC-interaction, or scene-description extraction; -- infer combat from scene titles, summaries, overlap, or campaign references; -- introduce arbitrary DAG scheduling or concurrent cross-step execution; -- add prior-run artifact discovery or new reference syntax; or -- add an LLM-backed validator, repair pass, or generic deduplication stage. - -## Completion Criteria - -The scope is complete when an explicitly configured ordered D&D pipeline can -hand an approved scene-description artifact to combat extraction, exact -`combat` scene matches are the only chunks that invoke the combat LLM, every -other chunk produces a deterministic empty result, provenance and checkpoint -identity remain correct, the complete example demonstrates the workflow, and -the canonical current-behavior documentation reflects the implemented -contract. diff --git a/internal/framework/contracts/typed_pipeline.go b/internal/framework/contracts/typed_pipeline.go index 9aa8b16..dcc5fae 100644 --- a/internal/framework/contracts/typed_pipeline.go +++ b/internal/framework/contracts/typed_pipeline.go @@ -93,13 +93,15 @@ type TypedNormalizeResult[T any] struct { Retry *NormalizeRetry } -// NormalizeRetry asks the framework to retry normalization while retaining a -// safe candidate for acceptance if the retry budget is exhausted. +// Normalize retry diagnostic limits bound module-provided values before the +// framework persists them in debug artifacts. const ( MaxNormalizeRetryReasonCodeBytes = 128 MaxNormalizeRetryMessageBytes = 4096 ) +// NormalizeRetry asks the framework to retry normalization while retaining a +// safe candidate for acceptance if the retry budget is exhausted. type NormalizeRetry struct { ReasonCode string Message string