From 26bd59a5a213235f50125855ffe734107d2885eb Mon Sep 17 00:00:00 2001 From: Eric Rakestraw Date: Sat, 25 Jul 2026 20:13:12 -0500 Subject: [PATCH] Plan LLM-assisted NPC normalization --- .../roadmap/dnd-npc-semantic-normalization.md | 328 +++++++ docs/roadmap/implementation.md | 883 ++++++++++++++++++ 2 files changed, 1211 insertions(+) create mode 100644 docs/roadmap/dnd-npc-semantic-normalization.md create mode 100644 docs/roadmap/implementation.md diff --git a/docs/roadmap/dnd-npc-semantic-normalization.md b/docs/roadmap/dnd-npc-semantic-normalization.md new file mode 100644 index 0000000..41f108c --- /dev/null +++ b/docs/roadmap/dnd-npc-semantic-normalization.md @@ -0,0 +1,328 @@ +# D&D NPC Semantic Normalization + +## Status + +Accepted scope; not implemented. + +## Purpose + +Reconcile NPC records that extraction produced under different display names +when the complete transcript establishes that they represent the same +individual. This work addresses identities split across scene or chunk +boundaries while preserving the existing rule that extraction records narrow +source evidence and normalization owns document-wide reconciliation. + +The model should make only the semantic identity judgment. Deterministic code +must continue to own identity derivation, proposal validation, artifact +mutation, evidence preservation, ordering, diagnostics, and final validation. + +## Target Behavior + +The target `dnd/npcs` normalization contract combines: + +- the current deterministic display-name normalization, stable-ID derivation, + source-reference canonicalization, and equal-comparison-key consolidation; +- one LLM determination of whether remaining, distinctly named NPC records + represent the same individual and which existing display name is canonical; + and +- deterministic proposal validation and application followed by the configured + normalize validator chain. + +Each configured normalization attempt should make at most one semantic pass +over the merged document-level NPC list, not one pass per extraction chunk or +candidate pair. The pass should be skipped when fewer than two distinct +candidate identities remain after deterministic preprocessing. + +False consolidation is more damaging than a missed consolidation. Prompt +policy should therefore require affirmative contextual evidence that names +identify the same individual and should prefer no group when identity remains +ambiguous. + +Canonical selection should favor the most complete stable proper name supported +by the transcript. A complete proper name is preferable to an abbreviation, +while an unadorned proper name is preferable to the same name plus a contextual +class, role, title, or relationship descriptor unless the transcript establishes +that descriptor as part of the character's name. The model must still select +one supplied display name rather than synthesize a better one. + +## Model Proposal Contract + +The private structured response should contain only proposed duplicate groups: + +```json +{ + "duplicate_groups": [ + { + "members": [ + "Billy", + "Billy the druid" + ], + "canonical_name": "Billy" + } + ] +} +``` + +`members` identifies candidates by their supplied display names. +`canonical_name` selects one existing member as the retained display name. An +empty `duplicate_groups` array is a valid determination that no sufficiently +supported duplicates exist. + +The model must not receive or return deterministic `npc:sha256:` IDs. Those IDs +are long, non-semantic implementation identities and remain exclusively owned +by deterministic code. The model must not invent a replacement name, rewrite +an NPC record, propose new source references, or return a complete replacement +artifact. + +Before the call, deterministic preprocessing ensures that each supplied display +name has a unique NPC identity comparison key. Returned names may resolve using +the existing comparison-key equivalences for case, whitespace, Unicode +compatibility, and supported apostrophes. Resolution must not use fuzzy, +substring, edit-distance, embedding, or other approximate matching. + +An applicable group must: + +- contain at least two distinct, known members; +- resolve every member uniquely against the supplied candidate set; +- select a known `canonical_name` that belongs to the group; +- contain no repeated member; and +- share no resolved member with any other proposed group. + +Unknown or ambiguous names, singleton groups, invalid canonical selections, and +other unsafe groups must not mutate the artifact. Groups are otherwise +independent: a locally valid group may be applied when none of its resolved +members appears in any other proposed group. + +Overlap is a proposal conflict even when one participating group is already +locally invalid. Deterministic code must discard every group in the connected +conflict set rather than selecting a winner by response order. Locally valid, +non-conflicting groups remain safe to apply. This permits useful partial +reconciliation without allowing an unsafe group to influence an NPC identity +that another group would mutate. + +## Retry And Safe Fallback + +An invalid or partially unsafe private proposal is a retryable normalization +attempt. This includes: + +- structurally invalid model output classified by the LLM boundary as an + invalid structured completion; and +- a structurally valid response containing any locally invalid or conflicting + group. + +The framework owns retry counting and attempt diagnostics. The normalizer +returns a safe candidate together with a bounded, content-safe retry +diagnostic. For structurally invalid output, that candidate is the +deterministic pre-LLM result. For a decoded proposal with unsafe groups, it +also includes every independently valid, non-conflicting group from that +attempt. When normalize retries remain, the framework invokes the normalizer +again from the same merged input; safe groups are not accumulated across +attempts. A later completely safe proposal is applied normally. + +When the configured attempt budget is exhausted, the framework validates and +accepts the final attempt's safe fallback instead of failing or rejecting the +NPC lane, provided that fallback passes the configured normalize validators. +It promotes one durable warning stating that one or more proposed groups were +omitted. The fallback may therefore contain a safe partial reconciliation, or +only the deterministic pre-LLM result when no group could be applied. Earlier +retry diagnostics, candidate values, and warnings remain attempt-local in +debug artifacts. + +The normal module-binding default remains `retries: 0`, meaning one total +normalization attempt and immediate fallback after its invalid proposal. +Retries occur only when the user configures a positive normalize retry count. +For example, `retries: 2` permits the initial proposal plus two additional +attempts before fallback. + +Transport, authentication, prompt-preparation, cancellation, input-encoding, +and other operational errors are not safe-proposal failures. They retain the +existing module/framework error behavior rather than being converted into an +accepted fallback. + +## Transcript Context Policy + +The model should receive every candidate's display name and source references, +together with transcript windows derived from those references. It should not +receive the complete transcript by default. + +Each window must contain: + +- the complete inclusive source range cited by the NPC record; +- up to two source units immediately before the cited range; and +- up to two source units immediately after the cited range. + +The surrounding-unit count is defined once as a named module policy constant +with value `2`; it is not a user-accessible configuration field. Window +construction consumes that value through one clear boundary so a later +configuration option can replace the fixed value without changing prompt or +reconciliation contracts. The context-radius policy participates in module +metadata and checkpoint identity so changing it invalidates incompatible +normalize checkpoints. + +The value was selected by reviewing the July 19 evaluation transcript. Its NPC +citations are generally self-contained; where additional context is useful, +two surrounding units capture the relevant question-and-answer exchange, +speaker transition, or short anaphora chain. A third unit frequently begins a +separate joke or table exchange and adds distraction without improving the +identity evidence. + +Window expansion must use source-document positions rather than arithmetic on +unit IDs. It must clamp at document boundaries, preserve document order, and +coalesce overlapping or adjacent expanded windows without duplicating units. +The prepared model input must distinguish originally cited units from +surrounding context. + +Only records with a non-empty NPC identity comparison key and a non-empty +source-reference collection that is wholly valid against the current source +document are eligible for semantic reconciliation. Ineligible records remain +unchanged so the configured deterministic validators retain ownership of their +rejection. They are not supplied to the model and cannot participate in a +proposed group. + +Surrounding units inform the semantic decision but do not automatically become +durable NPC evidence. An applied group unions only the source references +already present on its member records. The LLM cannot add references from the +context window. + +Full-transcript mode and a configurable context radius may be evaluated later. +They are not part of this scope. Provider prompt-cache behavior should be +measured rather than assumed before expanding context solely for cache +economics. + +## Deterministic Application + +For each approved group, deterministic code should: + +- retain the model-selected existing display name; +- union source references from every group member; +- canonicalize and exact-deduplicate those references in source-document order; +- derive the resulting stable NPC ID from the retained display name under the + existing NPC identity policy; and +- emit one bounded duplicate-collapse warning describing the applied group. + +The consolidated record should occupy the earliest original member position so +model response ordering cannot reorder the artifact. Unrelated records must +remain present and retain their relative order. Every input record must be +represented by exactly one output record, either unchanged or through one +approved consolidation. + +The existing NPC shape, identity, source-reference, schema, and relatedness +validators remain the final artifact boundary. No LLM-backed validator is +needed: model judgment occurs in the normalizer, while deterministic validators +continue to enforce the durable artifact contract. + +## Prompt, Provenance, And Diagnostics + +NPC semantic-normalization instructions and the private response schema should +be module-owned prompt assets. Stable D&D-wide identity or transcript guidance +may be reused through the existing shared prompt-asset mechanism where its +meaning is genuinely common. + +The normalizer should use the configured normalize-stage LLM profile and the +application-wide scheduled LLM client. Invalid structured output and unsafe +semantic proposals use the framework-owned retryable-fallback contract; +operational failures retain existing error semantics. + +Manifest metadata and component checkpoint fingerprints should identify the +semantic-normalization policy, prompt identity, private response-schema +identity, deterministic NPC identity policy, and evidence-context policy. +They must not contain transcript text, NPC names, source paths, raw model +responses, or other source content. + +Warnings and preparation or runtime errors must follow the established bounded +diagnostic and content-safety policies. Debug artifacts may retain the normal +attempt-local model request, response, proposal decisions, and warnings under +the existing debug sensitivity contract. + +## Configuration And Documentation + +The maintained complete D&D configuration should demonstrate a normalize-stage +LLM profile and an explicit positive retry count for the NPC lane. The +application-wide default remains zero additional retries. No new configuration +field is introduced for the context radius in this scope. + +Canonical documentation ownership is: + +- Configuration owns the LLM-backed `dnd/npcs` selection and its use of the + existing normalize binding's profile and retry fields. +- Operations owns the document-level semantic reconciliation call and its + checkpoint behavior. +- Pipeline internals own the provider-neutral invalid-structured-output + classification and normalize retryable-fallback contract. +- The NPC integration contract owns externally observable consolidation, + evidence, ordering, and warning behavior. +- Internal LLM and module documentation own the proposal boundary, + context-window construction, deterministic application, metadata, and + fingerprints. +- The broader generic LLM-assisted deduplication item in + [future.md](future.md) remains future work until another artifact demonstrates + that extracting a shared generic facility is worthwhile. + +## Quality Expectations + +Tests should protect the semantic and safety boundaries through deterministic +LLM fakes rather than live-provider calls. Coverage should demonstrate: + +- distinct display variants can be consolidated when the model proposes a + valid group; +- model-facing requests and responses use display names rather than stable + hash IDs; +- the selected existing canonical name controls ID derivation while all member + evidence is preserved; +- window construction uses source-document position, handles document edges, + coalesces overlap, and distinguishes cited evidence from context; +- no surrounding context is promoted into durable source references; +- empty proposals and fewer-than-two-candidate inputs preserve deterministic + normalization behavior; +- unknown, ambiguous, repeated, overlapping, and otherwise malformed semantic + groups cannot corrupt or reorder the artifact; +- locally valid groups are applied independently, while every group that + participates in a resolved-member conflict is discarded; +- an invalid proposal consumes only configured retry budget, a later valid + proposal can succeed, and exhaustion accepts the final attempt's safe + fallback with one durable warning; +- the default zero-retry binding makes exactly one proposal attempt before + fallback; +- warnings and failures remain bounded and do not expose transcript content; +- prompt, schema, policy, or context-policy changes invalidate relevant + checkpoint reuse; and +- an assembled ordered D&D pipeline supplies the reconciled NPC registry to + downstream consumers. + +Prompt tests should assert prepared message structure, supplied materials, and +cache-boundary behavior. They must not act as change detectors for particular +words or phrases in natural-language prompt text. + +Human review of representative transcripts should compare missed and false +consolidations, latency, input-token cost, and provider cache use. Probabilistic +model quality is an evaluation activity, not a deterministic CI assertion. + +## Non-Goals + +This scope does not: + +- change the durable NPC artifact schema; +- add a durable alias collection or identity history; +- expose the context radius or full-transcript selection as configuration; +- ask an LLM to read, reproduce, or derive stable NPC IDs; +- permit model-authored source references or arbitrary replacement records; +- add fuzzy deterministic name matching; +- add an LLM-backed normalize validator; +- change the global default retry count; +- reconcile NPCs concurrently across lanes or runs; +- add canonical NPC IDs to downstream artifact schemas; +- implement a general DAG or implicit cross-lane dependency; or +- implement the generic cross-artifact deduplication normalizer described in + `future.md`. + +## Completion Criteria + +The scope is complete when `dnd/npcs` can use one document-level LLM proposal +to reconcile differently named records conservatively, deterministic code +validates and independently applies only safe non-conflicting name-based +groups, the resulting NPC preserves all member evidence under a newly derived +stable ID, the two-unit context policy is centralized and fingerprinted, +invalid or partially unsafe proposals use configured framework retries and +then an accepted safe fallback, downstream ordered steps receive the +reconciled registry, and the canonical current-behavior documentation reflects +the implemented contract. diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md new file mode 100644 index 0000000..5b92b7e --- /dev/null +++ b/docs/roadmap/implementation.md @@ -0,0 +1,883 @@ +# D&D NPC Semantic Normalization Implementation Plan + +## Status + +Ready for implementation. + +## Objective + +Implement the accepted +[D&D NPC Semantic Normalization](dnd-npc-semantic-normalization.md) contract. +Enhance the existing `dnd/npcs` normalizer with one document-level LLM identity +proposal per configured normalization attempt, but keep name resolution, +proposal safety, evidence handling, ordering, stable-ID derivation, mutation, +diagnostics, and final artifact validation deterministic. + +Complete the stages below in order. Each stage should leave its affected +packages passing before the next stage begins. Do not introduce a generic +deduplication framework during this work. + +## Decisions Applying To Every Stage + +- Retain the production module key `dnd/npcs`, artifact kind + `dnd/npc-list`, durable schema `notarius.dnd.npcs` v1, codec, default + normalize validator chain, and empty reference-slot contract. +- Give normalization its own private prompt identity. Use: + - prompt ID and prompt filesystem directory: `dnd.npcs.normalize`; + - prompt version: `v1`; + - response-schema key: `dnd_npcs_normalize_llm`; + - response-schema ID: `notarius.dnd.npcs.normalize.llm`; + - response-schema name: `notarius_dnd_npcs_normalize_llm_v1`; and + - response-schema version: `v1`. +- Use `gemini-2-flash` as the normalize prompt's default profile, matching the + current NPC extraction prompt. An explicit normalize binding + `llm_profile` continues to override that prompt default through the existing + framework contract. +- Retain empty, strict `Options`. This scope adds no user-configurable context + mode, radius, batching, or prompt option. +- Define one package-owned context-radius constant with value `2`. Pass that + constant into a context builder that accepts a radius argument, so future + configuration can replace the caller-supplied value without redesigning the + builder. +- Use two model inputs: + - `candidates`, `application/json`, containing display names and integer + start/end source ranges; and + - `transcript`, `application/json`, containing the coalesced context windows. + Compute a `sha256:` digest over each exact encoded input, leave its origin URI + empty, and do not pass the original full-source `SourceInput` to the prompt. +- Never place `npc:sha256:` IDs in either model input, the private response + schema, or natural-language prompt material. Stable IDs remain internal to + deterministic application and durable artifacts. +- Keep the private response schema structural: require the top-level + `duplicate_groups` array and each group's `members` array and + `canonical_name` string; reject unknown fields and incompatible JSON types. + Do not encode semantic constraints such as non-empty strings, minimum or + unique array sizes, known names, canonical membership, or non-overlap in JSON + Schema. +- Resolve model-returned names only through `identity.ComparisonKey`. Do not add + fuzzy, substring, edit-distance, embedding, or heuristic name resolution. +- Classify structurally invalid model output separately from operational LLM + failures at the provider-neutral structured-completion boundary. Invalid + output and partially unsafe semantic proposals request a framework-owned + normalize retry with a safe candidate fallback. +- Apply locally valid proposal groups independently only when none of their + resolved members appears in another group. Discard all groups participating + in a resolved-member conflict, including locally valid groups that conflict + with an invalid group. Do not invent repairs, choose a conflict winner by + response order, produce a durable rejection, or add a normalizer-local retry + loop. +- Preserve accepted-only warning promotion. If retry budget remains, all + fallback warnings remain attempt-local. On the last configured attempt, the + runner validates and accepts the supplied safe fallback with one durable + exhaustion warning. +- Preserve the binding default of `retries: 0`; this feature does not change + global configuration defaults. Zero retries means one proposal call followed + immediately by fallback when that proposal is invalid. A configured value of + `2` permits three total proposal attempts. +- Do not add an LLM-backed validator. The existing deterministic validators + remain the artifact-acceptance boundary. +- Do not add a new stage, reference type, durable alias field, schema migration, + or compatibility adapter. The only framework contract expansion is the + provider-neutral invalid-output classification and normalize retryable + fallback described below. + +## Stage 1: Framework Retryable Normalize Fallback + +### Goal + +Add the smallest domain-neutral result contract needed for a normalizer to ask +the runner for another attempt while supplying a safe candidate to accept if +the configured retry budget is exhausted. + +### Invalid Structured-Output Classification + +Add a provider-neutral `ErrInvalidStructuredOutput` sentinel in +`internal/framework/contracts`. It identifies failures in which a provider +returned no usable value for the caller's declared structured-output contract. + +Update `internal/framework/llm/scriptorium_client.go` to wrap that sentinel for: + +- Scriptorium structured-output validation failure; +- an empty structured completion; and +- failure to decode returned structured content into the caller's output + target. + +Preserve the current contextual and redacted error text. Do not apply the +sentinel to nil clients or output targets, prompt preparation, authentication, +transport/provider execution, context cancellation, or other operational +failures. Existing modules that do not inspect the sentinel retain their +current behavior. + +### Normalize Result Contract + +Add a normalize-only, domain-neutral retry directive to +`contracts.TypedNormalizeResult[T]`. Use a named structure rather than Boolean +fields: + +```go +type NormalizeRetry struct { + ReasonCode string + Message string + FallbackWarnings []Warning +} + +type TypedNormalizeResult[T any] struct { + Value T + Warnings []Warning + Retry *NormalizeRetry +} +``` + +`Value` is always the safe fallback candidate when `Retry` is non-nil. +`Warnings` are ordinary attempt warnings. `FallbackWarnings` are promoted only +if the runner exhausts retry budget and accepts that fallback. `ReasonCode` and +`Message` are attempt-local retry diagnostics for debug artifacts and must be +non-empty, bounded, and content-safe when supplied by a module. + +Update normalizer type erasure and cloning so the directive and both warning +collections cross the registry boundary without mutable shared backing +storage. Normalizers returning `Retry: nil` remain behaviorally unchanged. + +### Runner Semantics + +In the normalize attempt path, serialize the returned candidate before acting +on a retry directive so debug can inspect the safe fallback through the +existing candidate envelope. + +When `Retry` is non-nil: + +- Reject a blank reason code or message as a framework/module contract error. +- Add a structured `retry` object to the normalize attempt debug payload with + its reason code, message, whether another attempt remains, and whether the + fallback is being accepted. Do not encode this as a + `contracts.RejectedOutput`. +- If `attempt <= lane.Normalize.Retries`, record the terminal attempt envelope + and return an unaccepted result with no rejection or error to + `runWithRetry`. The loop then performs the next configured attempt. +- On the final attempt, append defensively copied `FallbackWarnings` to the + attempt warnings and continue through the ordinary typed and serialized + normalize validator chain using `Value`. + +If final fallback validation approves, store and publish it exactly like any +other accepted normalized result, including its accepted warnings and +checkpoint. If validation rejects it, preserve the ordinary terminal normalize +rejection with the correct attempt count. Do not record intermediate fallbacks +as checkpoints or durable warnings. + +A later attempt with `Retry: nil` follows the existing candidate validation and +acceptance path; no earlier retry diagnostic or fallback warning becomes +durable. Debug persistence failure, cancellation, and ordinary module or +validator errors retain their current non-fallback behavior. + +This contract is normalize-specific because a safe reconciled fallback is the +demonstrated requirement. Do not add speculative retry fields to extraction, +merge, chunking, or validation results. + +### Stage 1 Tests + +Add generic framework tests at stable contracts: + +- `errors.Is` recognizes invalid structured output for validation, empty-body, + and decode failures, but not operational or cancellation failures. +- Normalizer registry erasure preserves a retry directive and returns defensive + warning copies. +- With `retries: 0`, one retryable result validates and accepts its fallback, + promotes only fallback and final-attempt warnings, and records one LLM/module + attempt. +- With positive retries, an invalid result followed by a regular valid result + accepts the later value and promotes no earlier fallback warnings. +- Exhausting positive retries validates and accepts only the last safe fallback + and promotes its fallback warning. +- A final fallback rejected by normalize validators remains a normal rejection + with the total attempt count. +- Intermediate fallback attempts produce debug retry diagnostics but no + checkpoint or durable warning. +- Blank retry reason fields, debug persistence failures, errors, and + cancellation retain the correct failure semantics. + +Keep these tests domain-neutral and use test-controlled retry values. Do not +assert the global default retry count here; configuration contract tests already +own that default. + +Run: + +```sh +go test ./internal/framework/contracts +go test ./internal/framework/llm +go test ./internal/framework/pipeline +``` + +### Stage 1 Completion Gate + +Any typed normalizer can request a framework-managed retry with a safe fallback; +zero and positive retry budgets behave deterministically; invalid structured +output is distinguishable from operational failure; and existing normalizers +remain unchanged. + +## Stage 2: Private Prompt Contract And Context Material + +### Goal + +Add the package-owned prompt, schema, input models, and deterministic +context-window builder without changing normalizer runtime behavior yet. + +### Prompt And Schema Assets + +Under `internal/modules/dnd/normalize/npcs`, add the same package-owned asset +structure used by LLM-backed D&D extractors: + +- `assets.go` embedding only this package's prompt YAML, Markdown, and private + response schema; +- `schema.go` owning the identities listed above and loading + `assets/schemas/dnd_npcs_normalize_llm.v1.json`; +- `scriptorium_assets.go` owning the ordered prompt manifest, registration, + and prompt-content digest; +- `assets/prompts/dnd.npcs.normalize.yaml`; +- package-local `task.md`, `instructions.md`, and a small candidate-input + rendering template; and +- `assets/schemas/dnd_npcs_normalize_llm.v1.json`. + +The prompt manifest should reuse only semantically applicable shared assets: + +1. `common-dnd-system.md`; +2. `common-dnd-identity.md`; +3. package-owned task and normalization instructions; +4. the package-owned `candidates` input message; and +5. `common-dnd-transcript.md` rendering the windowed `transcript` input. + +Place cache boundaries after the stable shared identity tier and after the +stable package instructions. Candidate data and transcript windows are +document-specific and should follow the final stable boundary. Do not reuse +the extraction-evidence or campaign-reference prompt fragments: normalization +does not extract new events, return citations, or consume campaign references. +Set Scriptorium `repair_attempts` to `0`, consistent with the current +module-owned structured-completion boundary. + +Prompt policy must tell the model to: + +- return only duplicate groups that the supplied transcript context clearly + establishes as one individual; +- prefer no group when identity is ambiguous; +- copy supplied display names into `members`; +- select `canonical_name` from the same group's supplied members; +- prefer a complete stable proper name over an abbreviation, but prefer an + unadorned proper name over that name plus a contextual class, role, title, or + relationship descriptor unless the descriptor is established as part of the + name; and +- avoid inventing names, source references, replacement records, or + explanatory output. + +### Model Input Types + +Add private DTOs rather than exposing durable NPC types directly: + +- Candidate input: + - a required `npcs` array in deterministic current-record order; + - each item contains only `name` and `source_refs`; and + - each source range contains only `start_unit_id` and `end_unit_id`. +- Transcript input: + - a required `windows` array in source-document order; + - each window contains its ordered `units`; and + - each unit contains generic source-unit `id`, `kind`, `text`, a defensively + cloned `metadata` value when present, and a Boolean `cited` marker. +- Proposal response: + - `duplicate_groups`; + - each group has `members []string` and `canonical_name string`. + +The candidate DTO deliberately omits stable NPC IDs and source-document IDs. +The transcript DTO deliberately preserves generic source-unit metadata because +the Seriatim adapter records speaker and timing context there; it must not +import or interpret Seriatim-specific types or metadata keys. + +Encode candidates and windows with `encoding/json` into independently owned +bytes. Return fixed contextual errors without including transcript text, NPC +names, encoded payloads, source paths, or model response content. + +### Context Builder + +Implement context selection inside `internal/modules/dnd/normalize/npcs`; do +not add it to generic framework or shared D&D packages for this first concrete +use. + +Use `source.NewDocumentIndex` and source-document slice positions. For each +preprocessed record: + +1. Require a non-empty `identity.ComparisonKey` for its display name. +2. Require a non-empty source-reference list. +3. Require every reference to validate against the current document. A record + with any invalid reference is ineligible for the semantic call, remains + unchanged, and is left to the configured normalize validators. +4. Convert each valid inclusive range to source slice positions. +5. Expand its start and end positions by the supplied radius, clamping to the + document bounds. +6. Sort intervals by document position and coalesce intervals that overlap or + are directly adjacent. +7. Emit each source unit at most once and preserve document order. +8. Mark a unit `cited: true` only when its position belongs to at least one + original unexpanded reference; surrounding units remain `false`. + +Only eligible records appear in the model's candidate list. Skip semantic +normalization when fewer than two eligible, comparison-distinct candidates +remain. Do not emit a new warning merely because a record is ineligible; shape +and source-reference validators already own that diagnostic. + +No artificial candidate-count, byte, token, or window-count limit is added in +this scope. If the selected material exceeds a provider's context limit, the +normal structured-completion failure remains a normalizer error. Batching and +token budgeting stay deferred pending observed need. + +### Stage 2 Tests + +Use package-level behavioral tests and the real prompt asset registry with +offline Scriptorium preparation: + +- The private schema accepts structurally valid empty and non-empty group + arrays, including semantically invalid values reserved for deterministic + checking, while rejecting missing fields, unknown fields, and wrong JSON + types. +- Prepared prompt messages contain exactly the two declared dynamic inputs in + the intended order and at the intended cache boundaries. +- Candidate and transcript input JSON contain display names, ranges, selected + source units, and citation markers but no stable NPC hash IDs or full-source + origin path. +- Context construction is based on source slice order for non-monotonic unit + IDs, clamps document edges, expands complete multi-unit citations, coalesces + overlap and adjacency, and does not duplicate units. +- Records with empty, missing, foreign-source, reversed, or otherwise invalid + references are excluded from semantic candidates without being mutated. +- Returned materials and copied metadata do not alias caller-owned values. + +Test the window mechanism with a caller-supplied test radius rather than +duplicating the production constant throughout the suite. Prompt tests must +assert message structure, inputs, schema identity, and cache behavior; they +must not require particular words or phrases in natural-language prompt files. + +Run: + +```sh +go test ./internal/modules/dnd/normalize/npcs +go test ./internal/modules/dnd/register +``` + +### Stage 2 Completion Gate + +The new prompt and schema can be registered and prepared offline; context +material is deterministic, content-safe, document-position-aware, and contains +no stable NPC IDs; existing normalization behavior remains unchanged. + +## Stage 3: Deterministic Proposal Validation And LLM-Assisted Normalization + +### Goal + +Wire the private structured completion into `dnd/npcs` and apply only safe +name-based groups while preserving all existing deterministic behavior. + +### Normalizer Construction And Identity + +Change the normalizer to retain: + +- the injected `contracts.StructuredLLMClient`; +- the prompt asset digest; and +- the private response-schema digest. + +Change `New` to accept the LLM client and return `(*Normalizer, error)`. Reject +a nil client and fail construction if prompt or schema metadata cannot load. +Update the normalizer registry builder to pass +`request.Dependencies.LLM`. Keep `Options` strict and empty and +`ReferenceSlots()` empty. + +Bump the complete normalization policy from `dnd.npcs.normalize.v2` to +`dnd.npcs.normalize.v3`. Add a stable semantic-context policy identity such as +`dnd.npcs.semantic_context.v1` and keep the radius as a separately inspectable +metadata value. + +Extend manifest metadata with: + +- `prompt_id`, `prompt_version`, and `prompt_sha256`; +- `response_schema_key`, `response_schema_id`, `response_schema_name`, + `response_schema_version`, and `response_schema_sha256`; +- `identity_policy`; +- `normalization_policy`; +- `semantic_context_policy`; and +- `semantic_context_radius`. + +Extend component fingerprints with local names: + +- `prompt`; +- `response_schema`; +- `identity_policy`; +- `normalization_policy`; and +- `semantic_context_policy`. + +The semantic-context fingerprint value must cover both its policy identity and +the radius value. Metadata and fingerprints must not contain names, source +text, model output, paths, timestamps, or invocation-specific data. Existing +pipeline scoping will prefix these local names and incorporate them into +checkpoint identity. The new fingerprints intentionally make prior NPC +normalize checkpoints cold misses. + +### Deterministic Preprocessing + +Refactor the current `normalizeList` implementation into a form that preserves +all existing behavior before the semantic call: + +- display whitespace normalization; +- stable-ID recomputation; +- source-reference canonicalization and exact deduplication; +- equal-`identity.ComparisonKey` consolidation; +- first-occurrence output anchoring; +- source-reference union; and +- existing field, ID, evidence, and duplicate warnings. + +Each preprocessed record must retain the sorted original input indexes it +represents and its earliest input index. This provenance is internal only and +supports deterministic application, ordering, scopes, and warnings. + +Return the deterministic result immediately, without an LLM call, when fewer +than two semantically eligible records remain. Preserve `nil` versus empty list +behavior and do not mutate or alias the merge input. + +### Structured Completion + +For two or more eligible records on each framework attempt: + +1. Build the candidate and transcript-window inputs from Stage 2 using the + production radius constant. +2. Call the injected scheduled client exactly once with: + - `StageName: Key`; + - the normalization prompt ID and v1 version; + - `ProfileID: req.LLMProfile`; + - `SessionID: req.SessionID`; and + - only the `candidates` and `transcript` input materials. +3. Decode into the private proposal DTO. +4. If completion returns `contracts.ErrInvalidStructuredOutput`, return the + deterministic pre-LLM artifact as `Value` with a normalize retry directive. +5. Wrap every other completion failure with normalizer and completion context, + relying on the existing Scriptorium redaction boundary and never appending + raw response content. + +Do not perform an LLM call per candidate pair or group. Do not make a second +module-local repair call. A framework retry invokes the normalizer again and +therefore makes one new proposal call with the same deterministic inputs. + +### Proposal Validation + +Validate the complete decoded proposal and isolate independently safe groups +before applying them: + +1. Build the unique map from `identity.ComparisonKey(displayName)` to eligible + preprocessed record position. +2. For each response group, resolve every member and `canonical_name`. +3. Mark a group locally invalid when it has: + - fewer than two distinct members; + - a blank, unknown, or ambiguous member; + - a repeated resolved member; + - a blank, unknown, or ambiguous canonical name; or + - a canonical record that is not a group member. +4. For every group, including locally invalid groups, retain each member that + resolves uniquely and count ownership of those resolved record positions + across the complete proposal. +5. Mark every group containing a record position owned by more than one + proposal group as conflicting. Discard every participant rather than + choosing a winner. Applying this rule to every group also discards all + groups in a chained conflict. +6. Classify a group as safe only when it is locally valid and non-conflicting. + Apply every safe group independently; discard every other group. +7. If any group is discarded, return the artifact containing the safe groups + as a retryable fallback. If none is discarded, return the applied artifact + through the ordinary successful result path. Response order may order + attempt diagnostics but must never choose artifact order or conflict + winners. + +Name resolution may accept only the equivalences already implemented by +`identity.ComparisonKey`. The validator must not infer which supplied name the +model intended. + +Add: + +- `npc_semantic_proposal_invalid` as the attempt-local retry reason for + structurally or semantically invalid proposals; +- `npc_semantic_reconciliation_exhausted` as the durable fallback-warning + reason; and +- `npc_normalization_warnings_omitted` for total accepted-warning truncation. + +For a partially unsafe semantic proposal, build one bounded retry message with +`shared/diagnostics.Aggregate`. It may identify proposal-group indexes and +stable issue categories, including every group participating in an overlap, +but must not echo model-returned names, transcript content, source paths, or +raw completion text. + +Apply the safe groups to the deterministic pre-LLM artifact and return that +partial result as the retry directive's safe `Value`. If structured output +could not be decoded or no group is safe, `Value` remains the deterministic +pre-LLM artifact. Ordinary deterministic preprocessing warnings and warnings +for safely applied groups remain in `Warnings`. + +Put one bounded warning with reason +`npc_semantic_reconciliation_exhausted` in `FallbackWarnings`; it should state +the exact number of groups omitted from the final proposal after the configured +attempt budget. For structurally invalid output, where no decoded group count +exists, state that the semantic proposal could not be applied. That warning +becomes durable only when the runner accepts final fallback. + +Each framework attempt starts from the same merge output and performs +deterministic preprocessing again. Do not retain or accumulate safe groups +from earlier attempts. If retry budget is exhausted, the final attempt's safe +candidate is authoritative. + +### Deterministic Group Application + +Apply approved groups against the preprocessed record positions: + +- Iterate records in existing order. +- Emit one consolidated record at the earliest member position and suppress + later members. +- Take the display name from the model-selected existing canonical member, + regardless of which member supplies the output anchor. +- Union every member's existing source references and canonicalize them with + the current source-document order. +- Derive a new stable ID from the selected display name. +- Combine the represented original input-index sets without loss. +- Leave unrelated records present and in their prior relative order. + +Use the existing `duplicate_npc_collapsed` reason code for both deterministic +equal-name and approved semantic consolidation. The semantic warning scope is +the earliest represented input index. Its bounded message should identify +input indexes and, when different, the input index that supplied the canonical +display name; it should not include raw names or transcript text. + +Build duplicate details with `shared/diagnostics.Aggregate`, and pass the +complete NPC normalization warning list through +`diagnostics.LimitWarnings(..., "npcs", +ReasonCodeNPCNormalizationWarningsOmitted)` before returning. This total cap +applies to old and new normalizer warnings and must preserve deterministic +warning order. When returning a retry directive, reserve one position within +`diagnostics.MaxWarnings` for the exhaustion fallback warning. If deterministic +or safely applied-group warnings require truncation, include their +omission-summary warning before the reserved exhaustion warning so final +accepted fallback still contains at most `MaxWarnings` warnings and reports +the exact omitted-warning count. + +For a proposal whose groups are all safe, return the applied artifact with +`Retry: nil`. +The configured normalize validator chain receives it through the ordinary +framework path. For a final fallback, Stage 1 sends the safe candidate through +that same validator chain. No validator, codec, artifact schema, or warning +JSON shape changes are needed. + +### Stage 3 Tests + +Use a small recording LLM fake at the normalizer boundary and real +deterministic collaborators: + +- A nil client fails construction before execution, while normal construction + exposes valid prompt and schema metadata. +- Nil receiver, nil context, canceled context, operational completion failure, + and JSON input-encoding failure follow content-safe error paths. +- Invalid structured output returns a retryable deterministic fallback rather + than an ordinary error. +- Zero or one eligible candidate performs no LLM call and retains existing + deterministic normalization. +- A valid proposal consolidates differently named records, selects an existing + canonical display name, recomputes its ID, unions and orders all member + evidence, occupies the earliest member position, preserves unrelated order, + and does not alias input. +- The completion request carries the normalize profile and session ID and + contains names but no stable hash IDs. +- Empty proposals leave comparison-distinct records unchanged. +- Case, whitespace, Unicode compatibility, and supported apostrophe variants + resolve through the existing comparison policy. +- Unknown, blank, repeated, singleton, non-member-canonical, and overlapping + groups are discarded and return a retryable safe fallback. +- A valid independent group is applied even when a disjoint group is invalid. +- A locally valid group sharing a resolved member with an invalid group is + discarded, and chained overlaps discard every group in the connected + conflict set without response-order dependence. +- A partially safe final attempt preserves its applied groups in the accepted + fallback and reports the exact omitted-group count; a later retry starts + from the original merge output rather than accumulating earlier groups. +- Invalid or ineligible source-reference records cannot participate in semantic + groups and remain available for final validator rejection. +- Surrounding context never becomes output evidence. +- Retry diagnostics and fallback warnings remain valid UTF-8 and individually + bounded. Accepted deterministic, applied-group, and exhaustion warnings are + subject to the total warning cap with an accurate omission summary. +- Metadata and fingerprints contain all stable policy and asset identities but + no names, source text, raw model responses, or paths. + +Do not add tests that assert exact natural-language prompt wording. Do not +repeat generic checkpoint-identity tests: package tests should prove that the +normalizer contributes the required stable fingerprints, while existing +framework tests continue to own the rule that fingerprint changes alter +identity. + +Run: + +```sh +go test ./internal/modules/dnd/normalize/npcs +go test ./internal/modules/dnd/validate/npcs/... +go test ./internal/framework/pipeline +``` + +### Stage 3 Completion Gate + +Each normalizer invocation performs at most one scheduled LLM call over +eligible document-level candidates. It independently applies safe, +non-conflicting name-based groups, requests a framework retry when any group +is discarded, supplies the resulting safe candidate as fallback, preserves +artifact integrity and evidence, emits bounded diagnostics, and contributes +complete checkpoint identity. + +## Stage 4: Production Registration, Pipeline Integration, And Checkpoint Coverage + +### Goal + +Make the new normalize prompt available in production composition and prove +that reconciled NPCs cross an ordered step barrier correctly. + +### Registration + +Update `internal/modules/dnd/register/modules.go` to register NPC normalize +prompt and schema assets in addition to the existing NPC extraction assets. +Use a distinct registration label and call the normalize package's +`RegisterPromptAssets`. + +Update registrar tests to cover the new prompt filesystem and schema +registration under `dnd.npcs.normalize`. Preserve existing module key, +artifact-kind registration, normalizer spec, reference slots, options, and +default normalize validator chain. + +Update every production test constructor or direct normalizer call for the new +LLM-injected constructor. Preparation tests should continue using the existing +production fake client; no real provider is permitted. + +### Existing Fake And Fixture Adaptation + +Several integration fakes currently assume that every `dnd/npcs` pipeline LLM +request is extraction. Change them to dispatch by `PromptID`: + +- extraction prompt requests continue returning the existing private NPC + extraction response; +- normalization prompt requests return a private + `duplicate_groups` response; and +- unrelated prompt IDs retain their existing behavior. + +For tests unrelated to semantic reconciliation, return +`{"duplicate_groups":[]}` and update only call-count assertions whose observable +contract now includes the one document-level normalize call. Do not duplicate +proposal edge cases already owned by normalize-package tests. + +### Assembled Behavior + +Add one focused assembled D&D workflow case with two extraction chunks that +produce a short proper name and a role-qualified variant for the same NPC. +Have the normalize fake propose one group using those display names. Prove: + +- extraction still produces and merge still preserves both candidates; +- normalization emits one canonical NPC at the earliest member position; +- its evidence is the canonical union of both chunks' source references; +- its ID is derived from the selected canonical display name; +- the accepted normalized artifact crosses the existing ordered step barrier; +- a downstream NPC consumer receives a names-only registry projection + containing the canonical name once; and +- no stable NPC hash ID or transcript content leaks into manifest component + metadata or reference provenance. + +Keep this integration test focused on collaboration among extraction, merge, +normalization, validation, and generated handoff. Proposal-malformation, +context-window edge cases, and schema parsing remain at their narrower owners. + +### Checkpoint And Retry Coverage + +Update the existing NPC preparation and grounded-pipeline assertions to include +the scoped normalize fingerprints for prompt, response schema, identity +policy, normalization policy, and semantic-context policy. Confirm the +normalize prompt/schema metadata appears under the normalizer manifest +component. + +Rely on existing generic checkpoint tests for ordering and identity sensitivity; +do not add another generic fingerprint mutation test. The assembled NPC test +should demonstrate that the new fingerprints are present before execution. +Existing checkpoint construction will then cause a one-time cold miss for old +NPC normalize checkpoints. + +Stage 1 framework tests own the generic retryable-fallback mechanism. Add only +the narrow NPC integration coverage needed to prove feature wiring: + +- with normalize retries omitted, one partially unsafe NPC proposal makes one + normalize completion call, accepts its safe partial fallback, and promotes + the exhaustion warning; and +- with `retries: 1`, a partially unsafe NPC proposal followed by a completely + safe proposal makes two normalize completion calls, applies the later + proposal, and does not promote the earlier fallback warning. + +Do not repeat every proposal-invalidity category or generic final-validator +case at integration scope. Existing NPC fakes should return a fresh +normalization response for each actual attempt. + +### Stage 4 Tests + +Run: + +```sh +go test ./internal/modules/dnd/register +go test ./internal/modules/integration +go test ./internal/cli +go test ./internal/framework/pipeline +``` + +### Stage 4 Completion Gate + +Production preparation registers and constructs the LLM-assisted normalizer, +representative pipelines route both NPC prompt identities correctly, prepared +identity contains all new fingerprints, and downstream generated references +observe the reconciled registry. + +## Stage 5: Maintained Configuration And Canonical Documentation + +### Goal + +Update each canonical owner to describe the implemented behavior without +duplicating volatile contracts. + +### Configuration And Example + +Update `examples/dnd-complete.config.yml` so the NPC normalize binding uses +object form and demonstrates the normalize-stage profile and retry knobs: + +```yaml +normalize: + module: dnd/npcs + llm_profile: gemini-2-flash + retries: 2 +``` + +Do not add a radius or context-mode option. Keep the minimal example unchanged +unless its current NPC composition requires object form for correctness. +Retain the maintained-example resolution tests as the owner of example +validity. + +Update `docs/config.md` to: + +- identify `dnd/npcs` normalization as LLM-assisted; +- explain that `llm_profile` and `retries` use the ordinary normalize binding + contract; +- retain and clearly state the global `retries: 0` default, including that it + allows one initial normalization attempt and no additional attempts; +- retain the absence of normalizer references and options; +- state that context selection is currently module policy rather than + configuration; and +- accurately list the affected prompt default without duplicating the complete + example. + +### Operations, Integration, And Internals + +Update: + +- `docs/operations.md` with the document-level NPC semantic-normalization call, + its place before the ordered generated-reference barrier, configured retry + behavior, safe partial fallback after invalid-proposal exhaustion, and + checkpoint reuse; +- `docs/internal/pipeline.md` with the provider-neutral invalid structured + output classification, normalize retry directive, attempt accounting, + accepted-only warning promotion, fallback validation, debug behavior, and + checkpoint rules; +- `docs/integrations/dnd-npc-artifacts.md` with externally observable + name-based consolidation, canonical selection, evidence union, ordering, + warnings, and unchanged durable v1 artifact shape; +- `docs/internal/modules.md` with deterministic preprocessing, candidate + eligibility, prompt inputs, private response contract, proposal validation, + application, diagnostics, metadata, and fingerprints; +- `docs/internal/llm.md` with the normalization prompt's stable and variable + message tiers, cache boundaries, windowed transcript input, and distinct + prompt/schema identity; and +- `docs/internal/overview.md` with a concise current package responsibility. + +Do not put implementation details in configuration or operations documentation. +Do not describe the private proposal schema as part of the durable NPC +integration schema. Keep the generic LLM-assisted deduplication item in +`future.md`; this feature is intentionally the first D&D-specific use and does +not complete that generic work. + +After all implementation and documentation checks pass: + +- change the feature roadmap status to `Implemented` and preserve it as the + feature contract and rationale; and +- change this implementation plan's status to `Completed`. + +### Documentation And Test Policy + +Update existing behavioral and contract tests where their owned behavior +changes. Do not add prompt-language change detectors, broad golden snapshots, +live-provider tests, redundant schema cases at higher layers, or assertions +against private helper shape. + +Run: + +```sh +go test ./internal/cli +go test ./internal/modules/integration +``` + +### Stage 5 Completion Gate + +The maintained complete example resolves, canonical current-behavior +documentation describes the shipped feature in its proper owners, future work +still distinguishes generic deduplication, and both roadmap statuses agree with +implementation state. + +## Final Verification + +From the repository root, run: + +```sh +git diff --check +go test ./... +go vet ./... +go build ./cmd/notarius +go test -race ./internal/modules/dnd/normalize/npcs ./internal/modules/dnd/register ./internal/framework/pipeline ./internal/cli ./internal/modules/integration +``` + +Review the final diff for: + +- no durable artifact or public configuration schema change; +- no model-facing stable NPC IDs; +- no full-transcript normalize input; +- no transcript, name, path, or raw response content in metadata, + fingerprints, provenance, or content-safe errors; +- no fuzzy matching or response-order-dependent application; +- no retryable-result expansion beyond the demonstrated normalize fallback or + domain behavior in the framework contract; +- no real-provider or prompt-wording change-detector tests; and +- no documentation describing future behavior as already implemented before + the final implementation stage is complete. + +## Assumptions + +- The current source document supplied by the runner remains the authoritative + source-unit order; unit IDs are identifiers and may be non-monotonic. +- Default production extraction validators normally provide valid NPC + references before merge. The normalizer still excludes invalid records from + semantic reconciliation so custom validator chains cannot make them unsafe + model operands. +- A deterministically validated partial reconciliation after retry exhaustion + is preferable to failing the complete NPC lane or discarding independent + safe groups. Operators can inspect its durable bounded warning and + attempt-local debug diagnostics. +- Invalid structured output is distinguishable from operational LLM failure. + The former uses retryable fallback; the latter retains existing error + behavior. +- The configuration-wide retry default remains zero. Any future decision to + change it to one or two is separate work with application-wide cost, + latency, and failure-semantics consequences. +- One document-level request is sufficient for the current workload. Batching, + token budgeting, retrieval, full-transcript mode, and context-radius + configuration require separate evidence and scope. +- Checkpoint compatibility is subordinate to semantic correctness. No + migration or cleanup is required for normalize checkpoints made cold by the + new component fingerprints. + +## Open Questions + +None. The feature roadmap and the decisions above are sufficient to implement +the work without additional product or architecture choices.