Files
notarius/docs/roadmap/implementation.md

884 lines
39 KiB
Markdown

# D&D NPC Semantic Normalization Implementation Plan
## Status
Completed.
## Objective
Implement the accepted
[D&D NPC Semantic Normalization](dnd-npc-semantic-normalization.md) contract.
Enhance the existing `dnd/npcs` normalizer with one document-level LLM identity
proposal per configured normalization attempt, but keep name resolution,
proposal safety, evidence handling, ordering, stable-ID derivation, mutation,
diagnostics, and final artifact validation deterministic.
Complete the stages below in order. Each stage should leave its affected
packages passing before the next stage begins. Do not introduce a generic
deduplication framework during this work.
## Decisions Applying To Every Stage
- Retain the production module key `dnd/npcs`, artifact kind
`dnd/npc-list`, durable schema `notarius.dnd.npcs` v1, codec, default
normalize validator chain, and empty reference-slot contract.
- Give normalization its own private prompt identity. Use:
- prompt ID and prompt filesystem directory: `dnd.npcs.normalize`;
- prompt version: `v1`;
- response-schema key: `dnd_npcs_normalize_llm`;
- response-schema ID: `notarius.dnd.npcs.normalize.llm`;
- response-schema name: `notarius_dnd_npcs_normalize_llm_v1`; and
- response-schema version: `v1`.
- Use `gemini-2-flash` as the normalize prompt's default profile, matching the
current NPC extraction prompt. An explicit normalize binding
`llm_profile` continues to override that prompt default through the existing
framework contract.
- Retain empty, strict `Options`. This scope adds no user-configurable context
mode, radius, batching, or prompt option.
- Define one package-owned context-radius constant with value `2`. Pass that
constant into a context builder that accepts a radius argument, so future
configuration can replace the caller-supplied value without redesigning the
builder.
- Use two model inputs:
- `candidates`, `application/json`, containing display names and integer
start/end source ranges; and
- `transcript`, `application/json`, containing the coalesced context windows.
Compute a `sha256:` digest over each exact encoded input, leave its origin URI
empty, and do not pass the original full-source `SourceInput` to the prompt.
- Never place `npc:sha256:` IDs in either model input, the private response
schema, or natural-language prompt material. Stable IDs remain internal to
deterministic application and durable artifacts.
- Keep the private response schema structural: require the top-level
`duplicate_groups` array and each group's `members` array and
`canonical_name` string; reject unknown fields and incompatible JSON types.
Do not encode semantic constraints such as non-empty strings, minimum or
unique array sizes, known names, canonical membership, or non-overlap in JSON
Schema.
- Resolve model-returned names only through `identity.ComparisonKey`. Do not add
fuzzy, substring, edit-distance, embedding, or heuristic name resolution.
- Classify structurally invalid model output separately from operational LLM
failures at the provider-neutral structured-completion boundary. Invalid
output and partially unsafe semantic proposals request a framework-owned
normalize retry with a safe candidate fallback.
- Apply locally valid proposal groups independently only when none of their
resolved members appears in another group. Discard all groups participating
in a resolved-member conflict, including locally valid groups that conflict
with an invalid group. Do not invent repairs, choose a conflict winner by
response order, produce a durable rejection, or add a normalizer-local retry
loop.
- Preserve accepted-only warning promotion. If retry budget remains, all
fallback warnings remain attempt-local. On the last configured attempt, the
runner validates and accepts the supplied safe fallback with one durable
exhaustion warning.
- Preserve the binding default of `retries: 0`; this feature does not change
global configuration defaults. Zero retries means one proposal call followed
immediately by fallback when that proposal is invalid. A configured value of
`2` permits three total proposal attempts.
- Do not add an LLM-backed validator. The existing deterministic validators
remain the artifact-acceptance boundary.
- Do not add a new stage, reference type, durable alias field, schema migration,
or compatibility adapter. The only framework contract expansion is the
provider-neutral invalid-output classification and normalize retryable
fallback described below.
## Stage 1: Framework Retryable Normalize Fallback
### Goal
Add the smallest domain-neutral result contract needed for a normalizer to ask
the runner for another attempt while supplying a safe candidate to accept if
the configured retry budget is exhausted.
### Invalid Structured-Output Classification
Add a provider-neutral `ErrInvalidStructuredOutput` sentinel in
`internal/framework/contracts`. It identifies failures in which a provider
returned no usable value for the caller's declared structured-output contract.
Update `internal/framework/llm/scriptorium_client.go` to wrap that sentinel for:
- Scriptorium structured-output validation failure;
- an empty structured completion; and
- failure to decode returned structured content into the caller's output
target.
Preserve the current contextual and redacted error text. Do not apply the
sentinel to nil clients or output targets, prompt preparation, authentication,
transport/provider execution, context cancellation, or other operational
failures. Existing modules that do not inspect the sentinel retain their
current behavior.
### Normalize Result Contract
Add a normalize-only, domain-neutral retry directive to
`contracts.TypedNormalizeResult[T]`. Use a named structure rather than Boolean
fields:
```go
type NormalizeRetry struct {
ReasonCode string
Message string
FallbackWarnings []Warning
}
type TypedNormalizeResult[T any] struct {
Value T
Warnings []Warning
Retry *NormalizeRetry
}
```
`Value` is always the safe fallback candidate when `Retry` is non-nil.
`Warnings` are ordinary attempt warnings. `FallbackWarnings` are promoted only
if the runner exhausts retry budget and accepts that fallback. `ReasonCode` and
`Message` are attempt-local retry diagnostics for debug artifacts and must be
non-empty, bounded, and content-safe when supplied by a module.
Update normalizer type erasure and cloning so the directive and both warning
collections cross the registry boundary without mutable shared backing
storage. Normalizers returning `Retry: nil` remain behaviorally unchanged.
### Runner Semantics
In the normalize attempt path, serialize the returned candidate before acting
on a retry directive so debug can inspect the safe fallback through the
existing candidate envelope.
When `Retry` is non-nil:
- Reject a blank reason code or message as a framework/module contract error.
- Add a structured `retry` object to the normalize attempt debug payload with
its reason code, message, whether another attempt remains, and whether the
fallback is being accepted. Do not encode this as a
`contracts.RejectedOutput`.
- If `attempt <= lane.Normalize.Retries`, record the terminal attempt envelope
and return an unaccepted result with no rejection or error to
`runWithRetry`. The loop then performs the next configured attempt.
- On the final attempt, append defensively copied `FallbackWarnings` to the
attempt warnings and continue through the ordinary typed and serialized
normalize validator chain using `Value`.
If final fallback validation approves, store and publish it exactly like any
other accepted normalized result, including its accepted warnings and
checkpoint. If validation rejects it, preserve the ordinary terminal normalize
rejection with the correct attempt count. Do not record intermediate fallbacks
as checkpoints or durable warnings.
A later attempt with `Retry: nil` follows the existing candidate validation and
acceptance path; no earlier retry diagnostic or fallback warning becomes
durable. Debug persistence failure, cancellation, and ordinary module or
validator errors retain their current non-fallback behavior.
This contract is normalize-specific because a safe reconciled fallback is the
demonstrated requirement. Do not add speculative retry fields to extraction,
merge, chunking, or validation results.
### Stage 1 Tests
Add generic framework tests at stable contracts:
- `errors.Is` recognizes invalid structured output for validation, empty-body,
and decode failures, but not operational or cancellation failures.
- Normalizer registry erasure preserves a retry directive and returns defensive
warning copies.
- With `retries: 0`, one retryable result validates and accepts its fallback,
promotes only fallback and final-attempt warnings, and records one LLM/module
attempt.
- With positive retries, an invalid result followed by a regular valid result
accepts the later value and promotes no earlier fallback warnings.
- Exhausting positive retries validates and accepts only the last safe fallback
and promotes its fallback warning.
- A final fallback rejected by normalize validators remains a normal rejection
with the total attempt count.
- Intermediate fallback attempts produce debug retry diagnostics but no
checkpoint or durable warning.
- Blank retry reason fields, debug persistence failures, errors, and
cancellation retain the correct failure semantics.
Keep these tests domain-neutral and use test-controlled retry values. Do not
assert the global default retry count here; configuration contract tests already
own that default.
Run:
```sh
go test ./internal/framework/contracts
go test ./internal/framework/llm
go test ./internal/framework/pipeline
```
### Stage 1 Completion Gate
Any typed normalizer can request a framework-managed retry with a safe fallback;
zero and positive retry budgets behave deterministically; invalid structured
output is distinguishable from operational failure; and existing normalizers
remain unchanged.
## Stage 2: Private Prompt Contract And Context Material
### Goal
Add the package-owned prompt, schema, input models, and deterministic
context-window builder without changing normalizer runtime behavior yet.
### Prompt And Schema Assets
Under `internal/modules/dnd/normalize/npcs`, add the same package-owned asset
structure used by LLM-backed D&D extractors:
- `assets.go` embedding only this package's prompt YAML, Markdown, and private
response schema;
- `schema.go` owning the identities listed above and loading
`assets/schemas/dnd_npcs_normalize_llm.v1.json`;
- `scriptorium_assets.go` owning the ordered prompt manifest, registration,
and prompt-content digest;
- `assets/prompts/dnd.npcs.normalize.yaml`;
- package-local `task.md`, `instructions.md`, and a small candidate-input
rendering template; and
- `assets/schemas/dnd_npcs_normalize_llm.v1.json`.
The prompt manifest should reuse only semantically applicable shared assets:
1. `common-dnd-system.md`;
2. `common-dnd-identity.md`;
3. package-owned task and normalization instructions;
4. the package-owned `candidates` input message; and
5. `common-dnd-transcript.md` rendering the windowed `transcript` input.
Place cache boundaries after the stable shared identity tier and after the
stable package instructions. Candidate data and transcript windows are
document-specific and should follow the final stable boundary. Do not reuse
the extraction-evidence or campaign-reference prompt fragments: normalization
does not extract new events, return citations, or consume campaign references.
Set Scriptorium `repair_attempts` to `0`, consistent with the current
module-owned structured-completion boundary.
Prompt policy must tell the model to:
- return only duplicate groups that the supplied transcript context clearly
establishes as one individual;
- prefer no group when identity is ambiguous;
- copy supplied display names into `members`;
- select `canonical_name` from the same group's supplied members;
- prefer a complete stable proper name over an abbreviation, but prefer an
unadorned proper name over that name plus a contextual class, role, title, or
relationship descriptor unless the descriptor is established as part of the
name; and
- avoid inventing names, source references, replacement records, or
explanatory output.
### Model Input Types
Add private DTOs rather than exposing durable NPC types directly:
- Candidate input:
- a required `npcs` array in deterministic current-record order;
- each item contains only `name` and `source_refs`; and
- each source range contains only `start_unit_id` and `end_unit_id`.
- Transcript input:
- a required `windows` array in source-document order;
- each window contains its ordered `units`; and
- each unit contains generic source-unit `id`, `kind`, `text`, a defensively
cloned `metadata` value when present, and a Boolean `cited` marker.
- Proposal response:
- `duplicate_groups`;
- each group has `members []string` and `canonical_name string`.
The candidate DTO deliberately omits stable NPC IDs and source-document IDs.
The transcript DTO deliberately preserves generic source-unit metadata because
the Seriatim adapter records speaker and timing context there; it must not
import or interpret Seriatim-specific types or metadata keys.
Encode candidates and windows with `encoding/json` into independently owned
bytes. Return fixed contextual errors without including transcript text, NPC
names, encoded payloads, source paths, or model response content.
### Context Builder
Implement context selection inside `internal/modules/dnd/normalize/npcs`; do
not add it to generic framework or shared D&D packages for this first concrete
use.
Use `source.NewDocumentIndex` and source-document slice positions. For each
preprocessed record:
1. Require a non-empty `identity.ComparisonKey` for its display name.
2. Require a non-empty source-reference list.
3. Require every reference to validate against the current document. A record
with any invalid reference is ineligible for the semantic call, remains
unchanged, and is left to the configured normalize validators.
4. Convert each valid inclusive range to source slice positions.
5. Expand its start and end positions by the supplied radius, clamping to the
document bounds.
6. Sort intervals by document position and coalesce intervals that overlap or
are directly adjacent.
7. Emit each source unit at most once and preserve document order.
8. Mark a unit `cited: true` only when its position belongs to at least one
original unexpanded reference; surrounding units remain `false`.
Only eligible records appear in the model's candidate list. Skip semantic
normalization when fewer than two eligible, comparison-distinct candidates
remain. Do not emit a new warning merely because a record is ineligible; shape
and source-reference validators already own that diagnostic.
No artificial candidate-count, byte, token, or window-count limit is added in
this scope. If the selected material exceeds a provider's context limit, the
normal structured-completion failure remains a normalizer error. Batching and
token budgeting stay deferred pending observed need.
### Stage 2 Tests
Use package-level behavioral tests and the real prompt asset registry with
offline Scriptorium preparation:
- The private schema accepts structurally valid empty and non-empty group
arrays, including semantically invalid values reserved for deterministic
checking, while rejecting missing fields, unknown fields, and wrong JSON
types.
- Prepared prompt messages contain exactly the two declared dynamic inputs in
the intended order and at the intended cache boundaries.
- Candidate and transcript input JSON contain display names, ranges, selected
source units, and citation markers but no stable NPC hash IDs or full-source
origin path.
- Context construction is based on source slice order for non-monotonic unit
IDs, clamps document edges, expands complete multi-unit citations, coalesces
overlap and adjacency, and does not duplicate units.
- Records with empty, missing, foreign-source, reversed, or otherwise invalid
references are excluded from semantic candidates without being mutated.
- Returned materials and copied metadata do not alias caller-owned values.
Test the window mechanism with a caller-supplied test radius rather than
duplicating the production constant throughout the suite. Prompt tests must
assert message structure, inputs, schema identity, and cache behavior; they
must not require particular words or phrases in natural-language prompt files.
Run:
```sh
go test ./internal/modules/dnd/normalize/npcs
go test ./internal/modules/dnd/register
```
### Stage 2 Completion Gate
The new prompt and schema can be registered and prepared offline; context
material is deterministic, content-safe, document-position-aware, and contains
no stable NPC IDs; existing normalization behavior remains unchanged.
## Stage 3: Deterministic Proposal Validation And LLM-Assisted Normalization
### Goal
Wire the private structured completion into `dnd/npcs` and apply only safe
name-based groups while preserving all existing deterministic behavior.
### Normalizer Construction And Identity
Change the normalizer to retain:
- the injected `contracts.StructuredLLMClient`;
- the prompt asset digest; and
- the private response-schema digest.
Change `New` to accept the LLM client and return `(*Normalizer, error)`. Reject
a nil client and fail construction if prompt or schema metadata cannot load.
Update the normalizer registry builder to pass
`request.Dependencies.LLM`. Keep `Options` strict and empty and
`ReferenceSlots()` empty.
Bump the complete normalization policy from `dnd.npcs.normalize.v2` to
`dnd.npcs.normalize.v3`. Add a stable semantic-context policy identity such as
`dnd.npcs.semantic_context.v1` and keep the radius as a separately inspectable
metadata value.
Extend manifest metadata with:
- `prompt_id`, `prompt_version`, and `prompt_sha256`;
- `response_schema_key`, `response_schema_id`, `response_schema_name`,
`response_schema_version`, and `response_schema_sha256`;
- `identity_policy`;
- `normalization_policy`;
- `semantic_context_policy`; and
- `semantic_context_radius`.
Extend component fingerprints with local names:
- `prompt`;
- `response_schema`;
- `identity_policy`;
- `normalization_policy`; and
- `semantic_context_policy`.
The semantic-context fingerprint value must cover both its policy identity and
the radius value. Metadata and fingerprints must not contain names, source
text, model output, paths, timestamps, or invocation-specific data. Existing
pipeline scoping will prefix these local names and incorporate them into
checkpoint identity. The new fingerprints intentionally make prior NPC
normalize checkpoints cold misses.
### Deterministic Preprocessing
Refactor the current `normalizeList` implementation into a form that preserves
all existing behavior before the semantic call:
- display whitespace normalization;
- stable-ID recomputation;
- source-reference canonicalization and exact deduplication;
- equal-`identity.ComparisonKey` consolidation;
- first-occurrence output anchoring;
- source-reference union; and
- existing field, ID, evidence, and duplicate warnings.
Each preprocessed record must retain the sorted original input indexes it
represents and its earliest input index. This provenance is internal only and
supports deterministic application, ordering, scopes, and warnings.
Return the deterministic result immediately, without an LLM call, when fewer
than two semantically eligible records remain. Preserve `nil` versus empty list
behavior and do not mutate or alias the merge input.
### Structured Completion
For two or more eligible records on each framework attempt:
1. Build the candidate and transcript-window inputs from Stage 2 using the
production radius constant.
2. Call the injected scheduled client exactly once with:
- `StageName: Key`;
- the normalization prompt ID and v1 version;
- `ProfileID: req.LLMProfile`;
- `SessionID: req.SessionID`; and
- only the `candidates` and `transcript` input materials.
3. Decode into the private proposal DTO.
4. If completion returns `contracts.ErrInvalidStructuredOutput`, return the
deterministic pre-LLM artifact as `Value` with a normalize retry directive.
5. Wrap every other completion failure with normalizer and completion context,
relying on the existing Scriptorium redaction boundary and never appending
raw response content.
Do not perform an LLM call per candidate pair or group. Do not make a second
module-local repair call. A framework retry invokes the normalizer again and
therefore makes one new proposal call with the same deterministic inputs.
### Proposal Validation
Validate the complete decoded proposal and isolate independently safe groups
before applying them:
1. Build the unique map from `identity.ComparisonKey(displayName)` to eligible
preprocessed record position.
2. For each response group, resolve every member and `canonical_name`.
3. Mark a group locally invalid when it has:
- fewer than two distinct members;
- a blank, unknown, or ambiguous member;
- a repeated resolved member;
- a blank, unknown, or ambiguous canonical name; or
- a canonical record that is not a group member.
4. For every group, including locally invalid groups, retain each member that
resolves uniquely and count ownership of those resolved record positions
across the complete proposal.
5. Mark every group containing a record position owned by more than one
proposal group as conflicting. Discard every participant rather than
choosing a winner. Applying this rule to every group also discards all
groups in a chained conflict.
6. Classify a group as safe only when it is locally valid and non-conflicting.
Apply every safe group independently; discard every other group.
7. If any group is discarded, return the artifact containing the safe groups
as a retryable fallback. If none is discarded, return the applied artifact
through the ordinary successful result path. Response order may order
attempt diagnostics but must never choose artifact order or conflict
winners.
Name resolution may accept only the equivalences already implemented by
`identity.ComparisonKey`. The validator must not infer which supplied name the
model intended.
Add:
- `npc_semantic_proposal_invalid` as the attempt-local retry reason for
structurally or semantically invalid proposals;
- `npc_semantic_reconciliation_exhausted` as the durable fallback-warning
reason; and
- `npc_normalization_warnings_omitted` for total accepted-warning truncation.
For a partially unsafe semantic proposal, build one bounded retry message with
`shared/diagnostics.Aggregate`. It may identify proposal-group indexes and
stable issue categories, including every group participating in an overlap,
but must not echo model-returned names, transcript content, source paths, or
raw completion text.
Apply the safe groups to the deterministic pre-LLM artifact and return that
partial result as the retry directive's safe `Value`. If structured output
could not be decoded or no group is safe, `Value` remains the deterministic
pre-LLM artifact. Ordinary deterministic preprocessing warnings and warnings
for safely applied groups remain in `Warnings`.
Put one bounded warning with reason
`npc_semantic_reconciliation_exhausted` in `FallbackWarnings`; it should state
the exact number of groups omitted from the final proposal after the configured
attempt budget. For structurally invalid output, where no decoded group count
exists, state that the semantic proposal could not be applied. That warning
becomes durable only when the runner accepts final fallback.
Each framework attempt starts from the same merge output and performs
deterministic preprocessing again. Do not retain or accumulate safe groups
from earlier attempts. If retry budget is exhausted, the final attempt's safe
candidate is authoritative.
### Deterministic Group Application
Apply approved groups against the preprocessed record positions:
- Iterate records in existing order.
- Emit one consolidated record at the earliest member position and suppress
later members.
- Take the display name from the model-selected existing canonical member,
regardless of which member supplies the output anchor.
- Union every member's existing source references and canonicalize them with
the current source-document order.
- Derive a new stable ID from the selected display name.
- Combine the represented original input-index sets without loss.
- Leave unrelated records present and in their prior relative order.
Use the existing `duplicate_npc_collapsed` reason code for both deterministic
equal-name and approved semantic consolidation. The semantic warning scope is
the earliest represented input index. Its bounded message should identify
input indexes and, when different, the input index that supplied the canonical
display name; it should not include raw names or transcript text.
Build duplicate details with `shared/diagnostics.Aggregate`, and pass the
complete NPC normalization warning list through
`diagnostics.LimitWarnings(..., "npcs",
ReasonCodeNPCNormalizationWarningsOmitted)` before returning. This total cap
applies to old and new normalizer warnings and must preserve deterministic
warning order. When returning a retry directive, reserve one position within
`diagnostics.MaxWarnings` for the exhaustion fallback warning. If deterministic
or safely applied-group warnings require truncation, include their
omission-summary warning before the reserved exhaustion warning so final
accepted fallback still contains at most `MaxWarnings` warnings and reports
the exact omitted-warning count.
For a proposal whose groups are all safe, return the applied artifact with
`Retry: nil`.
The configured normalize validator chain receives it through the ordinary
framework path. For a final fallback, Stage 1 sends the safe candidate through
that same validator chain. No validator, codec, artifact schema, or warning
JSON shape changes are needed.
### Stage 3 Tests
Use a small recording LLM fake at the normalizer boundary and real
deterministic collaborators:
- A nil client fails construction before execution, while normal construction
exposes valid prompt and schema metadata.
- Nil receiver, nil context, canceled context, operational completion failure,
and JSON input-encoding failure follow content-safe error paths.
- Invalid structured output returns a retryable deterministic fallback rather
than an ordinary error.
- Zero or one eligible candidate performs no LLM call and retains existing
deterministic normalization.
- A valid proposal consolidates differently named records, selects an existing
canonical display name, recomputes its ID, unions and orders all member
evidence, occupies the earliest member position, preserves unrelated order,
and does not alias input.
- The completion request carries the normalize profile and session ID and
contains names but no stable hash IDs.
- Empty proposals leave comparison-distinct records unchanged.
- Case, whitespace, Unicode compatibility, and supported apostrophe variants
resolve through the existing comparison policy.
- Unknown, blank, repeated, singleton, non-member-canonical, and overlapping
groups are discarded and return a retryable safe fallback.
- A valid independent group is applied even when a disjoint group is invalid.
- A locally valid group sharing a resolved member with an invalid group is
discarded, and chained overlaps discard every group in the connected
conflict set without response-order dependence.
- A partially safe final attempt preserves its applied groups in the accepted
fallback and reports the exact omitted-group count; a later retry starts
from the original merge output rather than accumulating earlier groups.
- Invalid or ineligible source-reference records cannot participate in semantic
groups and remain available for final validator rejection.
- Surrounding context never becomes output evidence.
- Retry diagnostics and fallback warnings remain valid UTF-8 and individually
bounded. Accepted deterministic, applied-group, and exhaustion warnings are
subject to the total warning cap with an accurate omission summary.
- Metadata and fingerprints contain all stable policy and asset identities but
no names, source text, raw model responses, or paths.
Do not add tests that assert exact natural-language prompt wording. Do not
repeat generic checkpoint-identity tests: package tests should prove that the
normalizer contributes the required stable fingerprints, while existing
framework tests continue to own the rule that fingerprint changes alter
identity.
Run:
```sh
go test ./internal/modules/dnd/normalize/npcs
go test ./internal/modules/dnd/validate/npcs/...
go test ./internal/framework/pipeline
```
### Stage 3 Completion Gate
Each normalizer invocation performs at most one scheduled LLM call over
eligible document-level candidates. It independently applies safe,
non-conflicting name-based groups, requests a framework retry when any group
is discarded, supplies the resulting safe candidate as fallback, preserves
artifact integrity and evidence, emits bounded diagnostics, and contributes
complete checkpoint identity.
## Stage 4: Production Registration, Pipeline Integration, And Checkpoint Coverage
### Goal
Make the new normalize prompt available in production composition and prove
that reconciled NPCs cross an ordered step barrier correctly.
### Registration
Update `internal/modules/dnd/register/modules.go` to register NPC normalize
prompt and schema assets in addition to the existing NPC extraction assets.
Use a distinct registration label and call the normalize package's
`RegisterPromptAssets`.
Update registrar tests to cover the new prompt filesystem and schema
registration under `dnd.npcs.normalize`. Preserve existing module key,
artifact-kind registration, normalizer spec, reference slots, options, and
default normalize validator chain.
Update every production test constructor or direct normalizer call for the new
LLM-injected constructor. Preparation tests should continue using the existing
production fake client; no real provider is permitted.
### Existing Fake And Fixture Adaptation
Several integration fakes currently assume that every `dnd/npcs` pipeline LLM
request is extraction. Change them to dispatch by `PromptID`:
- extraction prompt requests continue returning the existing private NPC
extraction response;
- normalization prompt requests return a private
`duplicate_groups` response; and
- unrelated prompt IDs retain their existing behavior.
For tests unrelated to semantic reconciliation, return
`{"duplicate_groups":[]}` and update only call-count assertions whose observable
contract now includes the one document-level normalize call. Do not duplicate
proposal edge cases already owned by normalize-package tests.
### Assembled Behavior
Add one focused assembled D&D workflow case with two extraction chunks that
produce a short proper name and a role-qualified variant for the same NPC.
Have the normalize fake propose one group using those display names. Prove:
- extraction still produces and merge still preserves both candidates;
- normalization emits one canonical NPC at the earliest member position;
- its evidence is the canonical union of both chunks' source references;
- its ID is derived from the selected canonical display name;
- the accepted normalized artifact crosses the existing ordered step barrier;
- a downstream NPC consumer receives a names-only registry projection
containing the canonical name once; and
- no stable NPC hash ID or transcript content leaks into manifest component
metadata or reference provenance.
Keep this integration test focused on collaboration among extraction, merge,
normalization, validation, and generated handoff. Proposal-malformation,
context-window edge cases, and schema parsing remain at their narrower owners.
### Checkpoint And Retry Coverage
Update the existing NPC preparation and grounded-pipeline assertions to include
the scoped normalize fingerprints for prompt, response schema, identity
policy, normalization policy, and semantic-context policy. Confirm the
normalize prompt/schema metadata appears under the normalizer manifest
component.
Rely on existing generic checkpoint tests for ordering and identity sensitivity;
do not add another generic fingerprint mutation test. The assembled NPC test
should demonstrate that the new fingerprints are present before execution.
Existing checkpoint construction will then cause a one-time cold miss for old
NPC normalize checkpoints.
Stage 1 framework tests own the generic retryable-fallback mechanism. Add only
the narrow NPC integration coverage needed to prove feature wiring:
- with normalize retries omitted, one partially unsafe NPC proposal makes one
normalize completion call, accepts its safe partial fallback, and promotes
the exhaustion warning; and
- with `retries: 1`, a partially unsafe NPC proposal followed by a completely
safe proposal makes two normalize completion calls, applies the later
proposal, and does not promote the earlier fallback warning.
Do not repeat every proposal-invalidity category or generic final-validator
case at integration scope. Existing NPC fakes should return a fresh
normalization response for each actual attempt.
### Stage 4 Tests
Run:
```sh
go test ./internal/modules/dnd/register
go test ./internal/modules/integration
go test ./internal/cli
go test ./internal/framework/pipeline
```
### Stage 4 Completion Gate
Production preparation registers and constructs the LLM-assisted normalizer,
representative pipelines route both NPC prompt identities correctly, prepared
identity contains all new fingerprints, and downstream generated references
observe the reconciled registry.
## Stage 5: Maintained Configuration And Canonical Documentation
### Goal
Update each canonical owner to describe the implemented behavior without
duplicating volatile contracts.
### Configuration And Example
Update `examples/dnd-complete.config.yml` so the NPC normalize binding uses
object form and demonstrates the normalize-stage profile and retry knobs:
```yaml
normalize:
module: dnd/npcs
llm_profile: gemini-2-flash
retries: 2
```
Do not add a radius or context-mode option. Keep the minimal example unchanged
unless its current NPC composition requires object form for correctness.
Retain the maintained-example resolution tests as the owner of example
validity.
Update `docs/config.md` to:
- identify `dnd/npcs` normalization as LLM-assisted;
- explain that `llm_profile` and `retries` use the ordinary normalize binding
contract;
- retain and clearly state the global `retries: 0` default, including that it
allows one initial normalization attempt and no additional attempts;
- retain the absence of normalizer references and options;
- state that context selection is currently module policy rather than
configuration; and
- accurately list the affected prompt default without duplicating the complete
example.
### Operations, Integration, And Internals
Update:
- `docs/operations.md` with the document-level NPC semantic-normalization call,
its place before the ordered generated-reference barrier, configured retry
behavior, safe partial fallback after invalid-proposal exhaustion, and
checkpoint reuse;
- `docs/internal/pipeline.md` with the provider-neutral invalid structured
output classification, normalize retry directive, attempt accounting,
accepted-only warning promotion, fallback validation, debug behavior, and
checkpoint rules;
- `docs/integrations/dnd-npc-artifacts.md` with externally observable
name-based consolidation, canonical selection, evidence union, ordering,
warnings, and unchanged durable v1 artifact shape;
- `docs/internal/modules.md` with deterministic preprocessing, candidate
eligibility, prompt inputs, private response contract, proposal validation,
application, diagnostics, metadata, and fingerprints;
- `docs/internal/llm.md` with the normalization prompt's stable and variable
message tiers, cache boundaries, windowed transcript input, and distinct
prompt/schema identity; and
- `docs/internal/overview.md` with a concise current package responsibility.
Do not put implementation details in configuration or operations documentation.
Do not describe the private proposal schema as part of the durable NPC
integration schema. Keep the generic LLM-assisted deduplication item in
`future.md`; this feature is intentionally the first D&D-specific use and does
not complete that generic work.
After all implementation and documentation checks pass:
- change the feature roadmap status to `Implemented` and preserve it as the
feature contract and rationale; and
- change this implementation plan's status to `Completed`.
### Documentation And Test Policy
Update existing behavioral and contract tests where their owned behavior
changes. Do not add prompt-language change detectors, broad golden snapshots,
live-provider tests, redundant schema cases at higher layers, or assertions
against private helper shape.
Run:
```sh
go test ./internal/cli
go test ./internal/modules/integration
```
### Stage 5 Completion Gate
The maintained complete example resolves, canonical current-behavior
documentation describes the shipped feature in its proper owners, future work
still distinguishes generic deduplication, and both roadmap statuses agree with
implementation state.
## Final Verification
From the repository root, run:
```sh
git diff --check
go test ./...
go vet ./...
go build ./cmd/notarius
go test -race ./internal/modules/dnd/normalize/npcs ./internal/modules/dnd/register ./internal/framework/pipeline ./internal/cli ./internal/modules/integration
```
Review the final diff for:
- no durable artifact or public configuration schema change;
- no model-facing stable NPC IDs;
- no full-transcript normalize input;
- no transcript, name, path, or raw response content in metadata,
fingerprints, provenance, or content-safe errors;
- no fuzzy matching or response-order-dependent application;
- no retryable-result expansion beyond the demonstrated normalize fallback or
domain behavior in the framework contract;
- no real-provider or prompt-wording change-detector tests; and
- no documentation describing future behavior as already implemented before
the final implementation stage is complete.
## Assumptions
- The current source document supplied by the runner remains the authoritative
source-unit order; unit IDs are identifiers and may be non-monotonic.
- Default production extraction validators normally provide valid NPC
references before merge. The normalizer still excludes invalid records from
semantic reconciliation so custom validator chains cannot make them unsafe
model operands.
- A deterministically validated partial reconciliation after retry exhaustion
is preferable to failing the complete NPC lane or discarding independent
safe groups. Operators can inspect its durable bounded warning and
attempt-local debug diagnostics.
- Invalid structured output is distinguishable from operational LLM failure.
The former uses retryable fallback; the latter retains existing error
behavior.
- The configuration-wide retry default remains zero. Any future decision to
change it to one or two is separate work with application-wide cost,
latency, and failure-semantics consequences.
- One document-level request is sufficient for the current workload. Batching,
token budgeting, retrieval, full-transcript mode, and context-radius
configuration require separate evidence and scope.
- Checkpoint compatibility is subordinate to semantic correctness. No
migration or cleanup is required for normalize checkpoints made cold by the
new component fingerprints.
## Open Questions
None. The feature roadmap and the decisions above are sufficient to implement
the work without additional product or architecture choices.