Plan LLM-assisted NPC normalization
This commit is contained in:
328
docs/roadmap/dnd-npc-semantic-normalization.md
Normal file
328
docs/roadmap/dnd-npc-semantic-normalization.md
Normal file
@@ -0,0 +1,328 @@
|
||||
# D&D NPC Semantic Normalization
|
||||
|
||||
## Status
|
||||
|
||||
Accepted scope; not implemented.
|
||||
|
||||
## Purpose
|
||||
|
||||
Reconcile NPC records that extraction produced under different display names
|
||||
when the complete transcript establishes that they represent the same
|
||||
individual. This work addresses identities split across scene or chunk
|
||||
boundaries while preserving the existing rule that extraction records narrow
|
||||
source evidence and normalization owns document-wide reconciliation.
|
||||
|
||||
The model should make only the semantic identity judgment. Deterministic code
|
||||
must continue to own identity derivation, proposal validation, artifact
|
||||
mutation, evidence preservation, ordering, diagnostics, and final validation.
|
||||
|
||||
## Target Behavior
|
||||
|
||||
The target `dnd/npcs` normalization contract combines:
|
||||
|
||||
- the current deterministic display-name normalization, stable-ID derivation,
|
||||
source-reference canonicalization, and equal-comparison-key consolidation;
|
||||
- one LLM determination of whether remaining, distinctly named NPC records
|
||||
represent the same individual and which existing display name is canonical;
|
||||
and
|
||||
- deterministic proposal validation and application followed by the configured
|
||||
normalize validator chain.
|
||||
|
||||
Each configured normalization attempt should make at most one semantic pass
|
||||
over the merged document-level NPC list, not one pass per extraction chunk or
|
||||
candidate pair. The pass should be skipped when fewer than two distinct
|
||||
candidate identities remain after deterministic preprocessing.
|
||||
|
||||
False consolidation is more damaging than a missed consolidation. Prompt
|
||||
policy should therefore require affirmative contextual evidence that names
|
||||
identify the same individual and should prefer no group when identity remains
|
||||
ambiguous.
|
||||
|
||||
Canonical selection should favor the most complete stable proper name supported
|
||||
by the transcript. A complete proper name is preferable to an abbreviation,
|
||||
while an unadorned proper name is preferable to the same name plus a contextual
|
||||
class, role, title, or relationship descriptor unless the transcript establishes
|
||||
that descriptor as part of the character's name. The model must still select
|
||||
one supplied display name rather than synthesize a better one.
|
||||
|
||||
## Model Proposal Contract
|
||||
|
||||
The private structured response should contain only proposed duplicate groups:
|
||||
|
||||
```json
|
||||
{
|
||||
"duplicate_groups": [
|
||||
{
|
||||
"members": [
|
||||
"Billy",
|
||||
"Billy the druid"
|
||||
],
|
||||
"canonical_name": "Billy"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
`members` identifies candidates by their supplied display names.
|
||||
`canonical_name` selects one existing member as the retained display name. An
|
||||
empty `duplicate_groups` array is a valid determination that no sufficiently
|
||||
supported duplicates exist.
|
||||
|
||||
The model must not receive or return deterministic `npc:sha256:` IDs. Those IDs
|
||||
are long, non-semantic implementation identities and remain exclusively owned
|
||||
by deterministic code. The model must not invent a replacement name, rewrite
|
||||
an NPC record, propose new source references, or return a complete replacement
|
||||
artifact.
|
||||
|
||||
Before the call, deterministic preprocessing ensures that each supplied display
|
||||
name has a unique NPC identity comparison key. Returned names may resolve using
|
||||
the existing comparison-key equivalences for case, whitespace, Unicode
|
||||
compatibility, and supported apostrophes. Resolution must not use fuzzy,
|
||||
substring, edit-distance, embedding, or other approximate matching.
|
||||
|
||||
An applicable group must:
|
||||
|
||||
- contain at least two distinct, known members;
|
||||
- resolve every member uniquely against the supplied candidate set;
|
||||
- select a known `canonical_name` that belongs to the group;
|
||||
- contain no repeated member; and
|
||||
- share no resolved member with any other proposed group.
|
||||
|
||||
Unknown or ambiguous names, singleton groups, invalid canonical selections, and
|
||||
other unsafe groups must not mutate the artifact. Groups are otherwise
|
||||
independent: a locally valid group may be applied when none of its resolved
|
||||
members appears in any other proposed group.
|
||||
|
||||
Overlap is a proposal conflict even when one participating group is already
|
||||
locally invalid. Deterministic code must discard every group in the connected
|
||||
conflict set rather than selecting a winner by response order. Locally valid,
|
||||
non-conflicting groups remain safe to apply. This permits useful partial
|
||||
reconciliation without allowing an unsafe group to influence an NPC identity
|
||||
that another group would mutate.
|
||||
|
||||
## Retry And Safe Fallback
|
||||
|
||||
An invalid or partially unsafe private proposal is a retryable normalization
|
||||
attempt. This includes:
|
||||
|
||||
- structurally invalid model output classified by the LLM boundary as an
|
||||
invalid structured completion; and
|
||||
- a structurally valid response containing any locally invalid or conflicting
|
||||
group.
|
||||
|
||||
The framework owns retry counting and attempt diagnostics. The normalizer
|
||||
returns a safe candidate together with a bounded, content-safe retry
|
||||
diagnostic. For structurally invalid output, that candidate is the
|
||||
deterministic pre-LLM result. For a decoded proposal with unsafe groups, it
|
||||
also includes every independently valid, non-conflicting group from that
|
||||
attempt. When normalize retries remain, the framework invokes the normalizer
|
||||
again from the same merged input; safe groups are not accumulated across
|
||||
attempts. A later completely safe proposal is applied normally.
|
||||
|
||||
When the configured attempt budget is exhausted, the framework validates and
|
||||
accepts the final attempt's safe fallback instead of failing or rejecting the
|
||||
NPC lane, provided that fallback passes the configured normalize validators.
|
||||
It promotes one durable warning stating that one or more proposed groups were
|
||||
omitted. The fallback may therefore contain a safe partial reconciliation, or
|
||||
only the deterministic pre-LLM result when no group could be applied. Earlier
|
||||
retry diagnostics, candidate values, and warnings remain attempt-local in
|
||||
debug artifacts.
|
||||
|
||||
The normal module-binding default remains `retries: 0`, meaning one total
|
||||
normalization attempt and immediate fallback after its invalid proposal.
|
||||
Retries occur only when the user configures a positive normalize retry count.
|
||||
For example, `retries: 2` permits the initial proposal plus two additional
|
||||
attempts before fallback.
|
||||
|
||||
Transport, authentication, prompt-preparation, cancellation, input-encoding,
|
||||
and other operational errors are not safe-proposal failures. They retain the
|
||||
existing module/framework error behavior rather than being converted into an
|
||||
accepted fallback.
|
||||
|
||||
## Transcript Context Policy
|
||||
|
||||
The model should receive every candidate's display name and source references,
|
||||
together with transcript windows derived from those references. It should not
|
||||
receive the complete transcript by default.
|
||||
|
||||
Each window must contain:
|
||||
|
||||
- the complete inclusive source range cited by the NPC record;
|
||||
- up to two source units immediately before the cited range; and
|
||||
- up to two source units immediately after the cited range.
|
||||
|
||||
The surrounding-unit count is defined once as a named module policy constant
|
||||
with value `2`; it is not a user-accessible configuration field. Window
|
||||
construction consumes that value through one clear boundary so a later
|
||||
configuration option can replace the fixed value without changing prompt or
|
||||
reconciliation contracts. The context-radius policy participates in module
|
||||
metadata and checkpoint identity so changing it invalidates incompatible
|
||||
normalize checkpoints.
|
||||
|
||||
The value was selected by reviewing the July 19 evaluation transcript. Its NPC
|
||||
citations are generally self-contained; where additional context is useful,
|
||||
two surrounding units capture the relevant question-and-answer exchange,
|
||||
speaker transition, or short anaphora chain. A third unit frequently begins a
|
||||
separate joke or table exchange and adds distraction without improving the
|
||||
identity evidence.
|
||||
|
||||
Window expansion must use source-document positions rather than arithmetic on
|
||||
unit IDs. It must clamp at document boundaries, preserve document order, and
|
||||
coalesce overlapping or adjacent expanded windows without duplicating units.
|
||||
The prepared model input must distinguish originally cited units from
|
||||
surrounding context.
|
||||
|
||||
Only records with a non-empty NPC identity comparison key and a non-empty
|
||||
source-reference collection that is wholly valid against the current source
|
||||
document are eligible for semantic reconciliation. Ineligible records remain
|
||||
unchanged so the configured deterministic validators retain ownership of their
|
||||
rejection. They are not supplied to the model and cannot participate in a
|
||||
proposed group.
|
||||
|
||||
Surrounding units inform the semantic decision but do not automatically become
|
||||
durable NPC evidence. An applied group unions only the source references
|
||||
already present on its member records. The LLM cannot add references from the
|
||||
context window.
|
||||
|
||||
Full-transcript mode and a configurable context radius may be evaluated later.
|
||||
They are not part of this scope. Provider prompt-cache behavior should be
|
||||
measured rather than assumed before expanding context solely for cache
|
||||
economics.
|
||||
|
||||
## Deterministic Application
|
||||
|
||||
For each approved group, deterministic code should:
|
||||
|
||||
- retain the model-selected existing display name;
|
||||
- union source references from every group member;
|
||||
- canonicalize and exact-deduplicate those references in source-document order;
|
||||
- derive the resulting stable NPC ID from the retained display name under the
|
||||
existing NPC identity policy; and
|
||||
- emit one bounded duplicate-collapse warning describing the applied group.
|
||||
|
||||
The consolidated record should occupy the earliest original member position so
|
||||
model response ordering cannot reorder the artifact. Unrelated records must
|
||||
remain present and retain their relative order. Every input record must be
|
||||
represented by exactly one output record, either unchanged or through one
|
||||
approved consolidation.
|
||||
|
||||
The existing NPC shape, identity, source-reference, schema, and relatedness
|
||||
validators remain the final artifact boundary. No LLM-backed validator is
|
||||
needed: model judgment occurs in the normalizer, while deterministic validators
|
||||
continue to enforce the durable artifact contract.
|
||||
|
||||
## Prompt, Provenance, And Diagnostics
|
||||
|
||||
NPC semantic-normalization instructions and the private response schema should
|
||||
be module-owned prompt assets. Stable D&D-wide identity or transcript guidance
|
||||
may be reused through the existing shared prompt-asset mechanism where its
|
||||
meaning is genuinely common.
|
||||
|
||||
The normalizer should use the configured normalize-stage LLM profile and the
|
||||
application-wide scheduled LLM client. Invalid structured output and unsafe
|
||||
semantic proposals use the framework-owned retryable-fallback contract;
|
||||
operational failures retain existing error semantics.
|
||||
|
||||
Manifest metadata and component checkpoint fingerprints should identify the
|
||||
semantic-normalization policy, prompt identity, private response-schema
|
||||
identity, deterministic NPC identity policy, and evidence-context policy.
|
||||
They must not contain transcript text, NPC names, source paths, raw model
|
||||
responses, or other source content.
|
||||
|
||||
Warnings and preparation or runtime errors must follow the established bounded
|
||||
diagnostic and content-safety policies. Debug artifacts may retain the normal
|
||||
attempt-local model request, response, proposal decisions, and warnings under
|
||||
the existing debug sensitivity contract.
|
||||
|
||||
## Configuration And Documentation
|
||||
|
||||
The maintained complete D&D configuration should demonstrate a normalize-stage
|
||||
LLM profile and an explicit positive retry count for the NPC lane. The
|
||||
application-wide default remains zero additional retries. No new configuration
|
||||
field is introduced for the context radius in this scope.
|
||||
|
||||
Canonical documentation ownership is:
|
||||
|
||||
- Configuration owns the LLM-backed `dnd/npcs` selection and its use of the
|
||||
existing normalize binding's profile and retry fields.
|
||||
- Operations owns the document-level semantic reconciliation call and its
|
||||
checkpoint behavior.
|
||||
- Pipeline internals own the provider-neutral invalid-structured-output
|
||||
classification and normalize retryable-fallback contract.
|
||||
- The NPC integration contract owns externally observable consolidation,
|
||||
evidence, ordering, and warning behavior.
|
||||
- Internal LLM and module documentation own the proposal boundary,
|
||||
context-window construction, deterministic application, metadata, and
|
||||
fingerprints.
|
||||
- The broader generic LLM-assisted deduplication item in
|
||||
[future.md](future.md) remains future work until another artifact demonstrates
|
||||
that extracting a shared generic facility is worthwhile.
|
||||
|
||||
## Quality Expectations
|
||||
|
||||
Tests should protect the semantic and safety boundaries through deterministic
|
||||
LLM fakes rather than live-provider calls. Coverage should demonstrate:
|
||||
|
||||
- distinct display variants can be consolidated when the model proposes a
|
||||
valid group;
|
||||
- model-facing requests and responses use display names rather than stable
|
||||
hash IDs;
|
||||
- the selected existing canonical name controls ID derivation while all member
|
||||
evidence is preserved;
|
||||
- window construction uses source-document position, handles document edges,
|
||||
coalesces overlap, and distinguishes cited evidence from context;
|
||||
- no surrounding context is promoted into durable source references;
|
||||
- empty proposals and fewer-than-two-candidate inputs preserve deterministic
|
||||
normalization behavior;
|
||||
- unknown, ambiguous, repeated, overlapping, and otherwise malformed semantic
|
||||
groups cannot corrupt or reorder the artifact;
|
||||
- locally valid groups are applied independently, while every group that
|
||||
participates in a resolved-member conflict is discarded;
|
||||
- an invalid proposal consumes only configured retry budget, a later valid
|
||||
proposal can succeed, and exhaustion accepts the final attempt's safe
|
||||
fallback with one durable warning;
|
||||
- the default zero-retry binding makes exactly one proposal attempt before
|
||||
fallback;
|
||||
- warnings and failures remain bounded and do not expose transcript content;
|
||||
- prompt, schema, policy, or context-policy changes invalidate relevant
|
||||
checkpoint reuse; and
|
||||
- an assembled ordered D&D pipeline supplies the reconciled NPC registry to
|
||||
downstream consumers.
|
||||
|
||||
Prompt tests should assert prepared message structure, supplied materials, and
|
||||
cache-boundary behavior. They must not act as change detectors for particular
|
||||
words or phrases in natural-language prompt text.
|
||||
|
||||
Human review of representative transcripts should compare missed and false
|
||||
consolidations, latency, input-token cost, and provider cache use. Probabilistic
|
||||
model quality is an evaluation activity, not a deterministic CI assertion.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
This scope does not:
|
||||
|
||||
- change the durable NPC artifact schema;
|
||||
- add a durable alias collection or identity history;
|
||||
- expose the context radius or full-transcript selection as configuration;
|
||||
- ask an LLM to read, reproduce, or derive stable NPC IDs;
|
||||
- permit model-authored source references or arbitrary replacement records;
|
||||
- add fuzzy deterministic name matching;
|
||||
- add an LLM-backed normalize validator;
|
||||
- change the global default retry count;
|
||||
- reconcile NPCs concurrently across lanes or runs;
|
||||
- add canonical NPC IDs to downstream artifact schemas;
|
||||
- implement a general DAG or implicit cross-lane dependency; or
|
||||
- implement the generic cross-artifact deduplication normalizer described in
|
||||
`future.md`.
|
||||
|
||||
## Completion Criteria
|
||||
|
||||
The scope is complete when `dnd/npcs` can use one document-level LLM proposal
|
||||
to reconcile differently named records conservatively, deterministic code
|
||||
validates and independently applies only safe non-conflicting name-based
|
||||
groups, the resulting NPC preserves all member evidence under a newly derived
|
||||
stable ID, the two-unit context policy is centralized and fingerprinted,
|
||||
invalid or partially unsafe proposals use configured framework retries and
|
||||
then an accepted safe fallback, downstream ordered steps receive the
|
||||
reconciled registry, and the canonical current-behavior documentation reflects
|
||||
the implemented contract.
|
||||
883
docs/roadmap/implementation.md
Normal file
883
docs/roadmap/implementation.md
Normal file
@@ -0,0 +1,883 @@
|
||||
# D&D NPC Semantic Normalization Implementation Plan
|
||||
|
||||
## Status
|
||||
|
||||
Ready for implementation.
|
||||
|
||||
## Objective
|
||||
|
||||
Implement the accepted
|
||||
[D&D NPC Semantic Normalization](dnd-npc-semantic-normalization.md) contract.
|
||||
Enhance the existing `dnd/npcs` normalizer with one document-level LLM identity
|
||||
proposal per configured normalization attempt, but keep name resolution,
|
||||
proposal safety, evidence handling, ordering, stable-ID derivation, mutation,
|
||||
diagnostics, and final artifact validation deterministic.
|
||||
|
||||
Complete the stages below in order. Each stage should leave its affected
|
||||
packages passing before the next stage begins. Do not introduce a generic
|
||||
deduplication framework during this work.
|
||||
|
||||
## Decisions Applying To Every Stage
|
||||
|
||||
- Retain the production module key `dnd/npcs`, artifact kind
|
||||
`dnd/npc-list`, durable schema `notarius.dnd.npcs` v1, codec, default
|
||||
normalize validator chain, and empty reference-slot contract.
|
||||
- Give normalization its own private prompt identity. Use:
|
||||
- prompt ID and prompt filesystem directory: `dnd.npcs.normalize`;
|
||||
- prompt version: `v1`;
|
||||
- response-schema key: `dnd_npcs_normalize_llm`;
|
||||
- response-schema ID: `notarius.dnd.npcs.normalize.llm`;
|
||||
- response-schema name: `notarius_dnd_npcs_normalize_llm_v1`; and
|
||||
- response-schema version: `v1`.
|
||||
- Use `gemini-2-flash` as the normalize prompt's default profile, matching the
|
||||
current NPC extraction prompt. An explicit normalize binding
|
||||
`llm_profile` continues to override that prompt default through the existing
|
||||
framework contract.
|
||||
- Retain empty, strict `Options`. This scope adds no user-configurable context
|
||||
mode, radius, batching, or prompt option.
|
||||
- Define one package-owned context-radius constant with value `2`. Pass that
|
||||
constant into a context builder that accepts a radius argument, so future
|
||||
configuration can replace the caller-supplied value without redesigning the
|
||||
builder.
|
||||
- Use two model inputs:
|
||||
- `candidates`, `application/json`, containing display names and integer
|
||||
start/end source ranges; and
|
||||
- `transcript`, `application/json`, containing the coalesced context windows.
|
||||
Compute a `sha256:` digest over each exact encoded input, leave its origin URI
|
||||
empty, and do not pass the original full-source `SourceInput` to the prompt.
|
||||
- Never place `npc:sha256:` IDs in either model input, the private response
|
||||
schema, or natural-language prompt material. Stable IDs remain internal to
|
||||
deterministic application and durable artifacts.
|
||||
- Keep the private response schema structural: require the top-level
|
||||
`duplicate_groups` array and each group's `members` array and
|
||||
`canonical_name` string; reject unknown fields and incompatible JSON types.
|
||||
Do not encode semantic constraints such as non-empty strings, minimum or
|
||||
unique array sizes, known names, canonical membership, or non-overlap in JSON
|
||||
Schema.
|
||||
- Resolve model-returned names only through `identity.ComparisonKey`. Do not add
|
||||
fuzzy, substring, edit-distance, embedding, or heuristic name resolution.
|
||||
- Classify structurally invalid model output separately from operational LLM
|
||||
failures at the provider-neutral structured-completion boundary. Invalid
|
||||
output and partially unsafe semantic proposals request a framework-owned
|
||||
normalize retry with a safe candidate fallback.
|
||||
- Apply locally valid proposal groups independently only when none of their
|
||||
resolved members appears in another group. Discard all groups participating
|
||||
in a resolved-member conflict, including locally valid groups that conflict
|
||||
with an invalid group. Do not invent repairs, choose a conflict winner by
|
||||
response order, produce a durable rejection, or add a normalizer-local retry
|
||||
loop.
|
||||
- Preserve accepted-only warning promotion. If retry budget remains, all
|
||||
fallback warnings remain attempt-local. On the last configured attempt, the
|
||||
runner validates and accepts the supplied safe fallback with one durable
|
||||
exhaustion warning.
|
||||
- Preserve the binding default of `retries: 0`; this feature does not change
|
||||
global configuration defaults. Zero retries means one proposal call followed
|
||||
immediately by fallback when that proposal is invalid. A configured value of
|
||||
`2` permits three total proposal attempts.
|
||||
- Do not add an LLM-backed validator. The existing deterministic validators
|
||||
remain the artifact-acceptance boundary.
|
||||
- Do not add a new stage, reference type, durable alias field, schema migration,
|
||||
or compatibility adapter. The only framework contract expansion is the
|
||||
provider-neutral invalid-output classification and normalize retryable
|
||||
fallback described below.
|
||||
|
||||
## Stage 1: Framework Retryable Normalize Fallback
|
||||
|
||||
### Goal
|
||||
|
||||
Add the smallest domain-neutral result contract needed for a normalizer to ask
|
||||
the runner for another attempt while supplying a safe candidate to accept if
|
||||
the configured retry budget is exhausted.
|
||||
|
||||
### Invalid Structured-Output Classification
|
||||
|
||||
Add a provider-neutral `ErrInvalidStructuredOutput` sentinel in
|
||||
`internal/framework/contracts`. It identifies failures in which a provider
|
||||
returned no usable value for the caller's declared structured-output contract.
|
||||
|
||||
Update `internal/framework/llm/scriptorium_client.go` to wrap that sentinel for:
|
||||
|
||||
- Scriptorium structured-output validation failure;
|
||||
- an empty structured completion; and
|
||||
- failure to decode returned structured content into the caller's output
|
||||
target.
|
||||
|
||||
Preserve the current contextual and redacted error text. Do not apply the
|
||||
sentinel to nil clients or output targets, prompt preparation, authentication,
|
||||
transport/provider execution, context cancellation, or other operational
|
||||
failures. Existing modules that do not inspect the sentinel retain their
|
||||
current behavior.
|
||||
|
||||
### Normalize Result Contract
|
||||
|
||||
Add a normalize-only, domain-neutral retry directive to
|
||||
`contracts.TypedNormalizeResult[T]`. Use a named structure rather than Boolean
|
||||
fields:
|
||||
|
||||
```go
|
||||
type NormalizeRetry struct {
|
||||
ReasonCode string
|
||||
Message string
|
||||
FallbackWarnings []Warning
|
||||
}
|
||||
|
||||
type TypedNormalizeResult[T any] struct {
|
||||
Value T
|
||||
Warnings []Warning
|
||||
Retry *NormalizeRetry
|
||||
}
|
||||
```
|
||||
|
||||
`Value` is always the safe fallback candidate when `Retry` is non-nil.
|
||||
`Warnings` are ordinary attempt warnings. `FallbackWarnings` are promoted only
|
||||
if the runner exhausts retry budget and accepts that fallback. `ReasonCode` and
|
||||
`Message` are attempt-local retry diagnostics for debug artifacts and must be
|
||||
non-empty, bounded, and content-safe when supplied by a module.
|
||||
|
||||
Update normalizer type erasure and cloning so the directive and both warning
|
||||
collections cross the registry boundary without mutable shared backing
|
||||
storage. Normalizers returning `Retry: nil` remain behaviorally unchanged.
|
||||
|
||||
### Runner Semantics
|
||||
|
||||
In the normalize attempt path, serialize the returned candidate before acting
|
||||
on a retry directive so debug can inspect the safe fallback through the
|
||||
existing candidate envelope.
|
||||
|
||||
When `Retry` is non-nil:
|
||||
|
||||
- Reject a blank reason code or message as a framework/module contract error.
|
||||
- Add a structured `retry` object to the normalize attempt debug payload with
|
||||
its reason code, message, whether another attempt remains, and whether the
|
||||
fallback is being accepted. Do not encode this as a
|
||||
`contracts.RejectedOutput`.
|
||||
- If `attempt <= lane.Normalize.Retries`, record the terminal attempt envelope
|
||||
and return an unaccepted result with no rejection or error to
|
||||
`runWithRetry`. The loop then performs the next configured attempt.
|
||||
- On the final attempt, append defensively copied `FallbackWarnings` to the
|
||||
attempt warnings and continue through the ordinary typed and serialized
|
||||
normalize validator chain using `Value`.
|
||||
|
||||
If final fallback validation approves, store and publish it exactly like any
|
||||
other accepted normalized result, including its accepted warnings and
|
||||
checkpoint. If validation rejects it, preserve the ordinary terminal normalize
|
||||
rejection with the correct attempt count. Do not record intermediate fallbacks
|
||||
as checkpoints or durable warnings.
|
||||
|
||||
A later attempt with `Retry: nil` follows the existing candidate validation and
|
||||
acceptance path; no earlier retry diagnostic or fallback warning becomes
|
||||
durable. Debug persistence failure, cancellation, and ordinary module or
|
||||
validator errors retain their current non-fallback behavior.
|
||||
|
||||
This contract is normalize-specific because a safe reconciled fallback is the
|
||||
demonstrated requirement. Do not add speculative retry fields to extraction,
|
||||
merge, chunking, or validation results.
|
||||
|
||||
### Stage 1 Tests
|
||||
|
||||
Add generic framework tests at stable contracts:
|
||||
|
||||
- `errors.Is` recognizes invalid structured output for validation, empty-body,
|
||||
and decode failures, but not operational or cancellation failures.
|
||||
- Normalizer registry erasure preserves a retry directive and returns defensive
|
||||
warning copies.
|
||||
- With `retries: 0`, one retryable result validates and accepts its fallback,
|
||||
promotes only fallback and final-attempt warnings, and records one LLM/module
|
||||
attempt.
|
||||
- With positive retries, an invalid result followed by a regular valid result
|
||||
accepts the later value and promotes no earlier fallback warnings.
|
||||
- Exhausting positive retries validates and accepts only the last safe fallback
|
||||
and promotes its fallback warning.
|
||||
- A final fallback rejected by normalize validators remains a normal rejection
|
||||
with the total attempt count.
|
||||
- Intermediate fallback attempts produce debug retry diagnostics but no
|
||||
checkpoint or durable warning.
|
||||
- Blank retry reason fields, debug persistence failures, errors, and
|
||||
cancellation retain the correct failure semantics.
|
||||
|
||||
Keep these tests domain-neutral and use test-controlled retry values. Do not
|
||||
assert the global default retry count here; configuration contract tests already
|
||||
own that default.
|
||||
|
||||
Run:
|
||||
|
||||
```sh
|
||||
go test ./internal/framework/contracts
|
||||
go test ./internal/framework/llm
|
||||
go test ./internal/framework/pipeline
|
||||
```
|
||||
|
||||
### Stage 1 Completion Gate
|
||||
|
||||
Any typed normalizer can request a framework-managed retry with a safe fallback;
|
||||
zero and positive retry budgets behave deterministically; invalid structured
|
||||
output is distinguishable from operational failure; and existing normalizers
|
||||
remain unchanged.
|
||||
|
||||
## Stage 2: Private Prompt Contract And Context Material
|
||||
|
||||
### Goal
|
||||
|
||||
Add the package-owned prompt, schema, input models, and deterministic
|
||||
context-window builder without changing normalizer runtime behavior yet.
|
||||
|
||||
### Prompt And Schema Assets
|
||||
|
||||
Under `internal/modules/dnd/normalize/npcs`, add the same package-owned asset
|
||||
structure used by LLM-backed D&D extractors:
|
||||
|
||||
- `assets.go` embedding only this package's prompt YAML, Markdown, and private
|
||||
response schema;
|
||||
- `schema.go` owning the identities listed above and loading
|
||||
`assets/schemas/dnd_npcs_normalize_llm.v1.json`;
|
||||
- `scriptorium_assets.go` owning the ordered prompt manifest, registration,
|
||||
and prompt-content digest;
|
||||
- `assets/prompts/dnd.npcs.normalize.yaml`;
|
||||
- package-local `task.md`, `instructions.md`, and a small candidate-input
|
||||
rendering template; and
|
||||
- `assets/schemas/dnd_npcs_normalize_llm.v1.json`.
|
||||
|
||||
The prompt manifest should reuse only semantically applicable shared assets:
|
||||
|
||||
1. `common-dnd-system.md`;
|
||||
2. `common-dnd-identity.md`;
|
||||
3. package-owned task and normalization instructions;
|
||||
4. the package-owned `candidates` input message; and
|
||||
5. `common-dnd-transcript.md` rendering the windowed `transcript` input.
|
||||
|
||||
Place cache boundaries after the stable shared identity tier and after the
|
||||
stable package instructions. Candidate data and transcript windows are
|
||||
document-specific and should follow the final stable boundary. Do not reuse
|
||||
the extraction-evidence or campaign-reference prompt fragments: normalization
|
||||
does not extract new events, return citations, or consume campaign references.
|
||||
Set Scriptorium `repair_attempts` to `0`, consistent with the current
|
||||
module-owned structured-completion boundary.
|
||||
|
||||
Prompt policy must tell the model to:
|
||||
|
||||
- return only duplicate groups that the supplied transcript context clearly
|
||||
establishes as one individual;
|
||||
- prefer no group when identity is ambiguous;
|
||||
- copy supplied display names into `members`;
|
||||
- select `canonical_name` from the same group's supplied members;
|
||||
- prefer a complete stable proper name over an abbreviation, but prefer an
|
||||
unadorned proper name over that name plus a contextual class, role, title, or
|
||||
relationship descriptor unless the descriptor is established as part of the
|
||||
name; and
|
||||
- avoid inventing names, source references, replacement records, or
|
||||
explanatory output.
|
||||
|
||||
### Model Input Types
|
||||
|
||||
Add private DTOs rather than exposing durable NPC types directly:
|
||||
|
||||
- Candidate input:
|
||||
- a required `npcs` array in deterministic current-record order;
|
||||
- each item contains only `name` and `source_refs`; and
|
||||
- each source range contains only `start_unit_id` and `end_unit_id`.
|
||||
- Transcript input:
|
||||
- a required `windows` array in source-document order;
|
||||
- each window contains its ordered `units`; and
|
||||
- each unit contains generic source-unit `id`, `kind`, `text`, a defensively
|
||||
cloned `metadata` value when present, and a Boolean `cited` marker.
|
||||
- Proposal response:
|
||||
- `duplicate_groups`;
|
||||
- each group has `members []string` and `canonical_name string`.
|
||||
|
||||
The candidate DTO deliberately omits stable NPC IDs and source-document IDs.
|
||||
The transcript DTO deliberately preserves generic source-unit metadata because
|
||||
the Seriatim adapter records speaker and timing context there; it must not
|
||||
import or interpret Seriatim-specific types or metadata keys.
|
||||
|
||||
Encode candidates and windows with `encoding/json` into independently owned
|
||||
bytes. Return fixed contextual errors without including transcript text, NPC
|
||||
names, encoded payloads, source paths, or model response content.
|
||||
|
||||
### Context Builder
|
||||
|
||||
Implement context selection inside `internal/modules/dnd/normalize/npcs`; do
|
||||
not add it to generic framework or shared D&D packages for this first concrete
|
||||
use.
|
||||
|
||||
Use `source.NewDocumentIndex` and source-document slice positions. For each
|
||||
preprocessed record:
|
||||
|
||||
1. Require a non-empty `identity.ComparisonKey` for its display name.
|
||||
2. Require a non-empty source-reference list.
|
||||
3. Require every reference to validate against the current document. A record
|
||||
with any invalid reference is ineligible for the semantic call, remains
|
||||
unchanged, and is left to the configured normalize validators.
|
||||
4. Convert each valid inclusive range to source slice positions.
|
||||
5. Expand its start and end positions by the supplied radius, clamping to the
|
||||
document bounds.
|
||||
6. Sort intervals by document position and coalesce intervals that overlap or
|
||||
are directly adjacent.
|
||||
7. Emit each source unit at most once and preserve document order.
|
||||
8. Mark a unit `cited: true` only when its position belongs to at least one
|
||||
original unexpanded reference; surrounding units remain `false`.
|
||||
|
||||
Only eligible records appear in the model's candidate list. Skip semantic
|
||||
normalization when fewer than two eligible, comparison-distinct candidates
|
||||
remain. Do not emit a new warning merely because a record is ineligible; shape
|
||||
and source-reference validators already own that diagnostic.
|
||||
|
||||
No artificial candidate-count, byte, token, or window-count limit is added in
|
||||
this scope. If the selected material exceeds a provider's context limit, the
|
||||
normal structured-completion failure remains a normalizer error. Batching and
|
||||
token budgeting stay deferred pending observed need.
|
||||
|
||||
### Stage 2 Tests
|
||||
|
||||
Use package-level behavioral tests and the real prompt asset registry with
|
||||
offline Scriptorium preparation:
|
||||
|
||||
- The private schema accepts structurally valid empty and non-empty group
|
||||
arrays, including semantically invalid values reserved for deterministic
|
||||
checking, while rejecting missing fields, unknown fields, and wrong JSON
|
||||
types.
|
||||
- Prepared prompt messages contain exactly the two declared dynamic inputs in
|
||||
the intended order and at the intended cache boundaries.
|
||||
- Candidate and transcript input JSON contain display names, ranges, selected
|
||||
source units, and citation markers but no stable NPC hash IDs or full-source
|
||||
origin path.
|
||||
- Context construction is based on source slice order for non-monotonic unit
|
||||
IDs, clamps document edges, expands complete multi-unit citations, coalesces
|
||||
overlap and adjacency, and does not duplicate units.
|
||||
- Records with empty, missing, foreign-source, reversed, or otherwise invalid
|
||||
references are excluded from semantic candidates without being mutated.
|
||||
- Returned materials and copied metadata do not alias caller-owned values.
|
||||
|
||||
Test the window mechanism with a caller-supplied test radius rather than
|
||||
duplicating the production constant throughout the suite. Prompt tests must
|
||||
assert message structure, inputs, schema identity, and cache behavior; they
|
||||
must not require particular words or phrases in natural-language prompt files.
|
||||
|
||||
Run:
|
||||
|
||||
```sh
|
||||
go test ./internal/modules/dnd/normalize/npcs
|
||||
go test ./internal/modules/dnd/register
|
||||
```
|
||||
|
||||
### Stage 2 Completion Gate
|
||||
|
||||
The new prompt and schema can be registered and prepared offline; context
|
||||
material is deterministic, content-safe, document-position-aware, and contains
|
||||
no stable NPC IDs; existing normalization behavior remains unchanged.
|
||||
|
||||
## Stage 3: Deterministic Proposal Validation And LLM-Assisted Normalization
|
||||
|
||||
### Goal
|
||||
|
||||
Wire the private structured completion into `dnd/npcs` and apply only safe
|
||||
name-based groups while preserving all existing deterministic behavior.
|
||||
|
||||
### Normalizer Construction And Identity
|
||||
|
||||
Change the normalizer to retain:
|
||||
|
||||
- the injected `contracts.StructuredLLMClient`;
|
||||
- the prompt asset digest; and
|
||||
- the private response-schema digest.
|
||||
|
||||
Change `New` to accept the LLM client and return `(*Normalizer, error)`. Reject
|
||||
a nil client and fail construction if prompt or schema metadata cannot load.
|
||||
Update the normalizer registry builder to pass
|
||||
`request.Dependencies.LLM`. Keep `Options` strict and empty and
|
||||
`ReferenceSlots()` empty.
|
||||
|
||||
Bump the complete normalization policy from `dnd.npcs.normalize.v2` to
|
||||
`dnd.npcs.normalize.v3`. Add a stable semantic-context policy identity such as
|
||||
`dnd.npcs.semantic_context.v1` and keep the radius as a separately inspectable
|
||||
metadata value.
|
||||
|
||||
Extend manifest metadata with:
|
||||
|
||||
- `prompt_id`, `prompt_version`, and `prompt_sha256`;
|
||||
- `response_schema_key`, `response_schema_id`, `response_schema_name`,
|
||||
`response_schema_version`, and `response_schema_sha256`;
|
||||
- `identity_policy`;
|
||||
- `normalization_policy`;
|
||||
- `semantic_context_policy`; and
|
||||
- `semantic_context_radius`.
|
||||
|
||||
Extend component fingerprints with local names:
|
||||
|
||||
- `prompt`;
|
||||
- `response_schema`;
|
||||
- `identity_policy`;
|
||||
- `normalization_policy`; and
|
||||
- `semantic_context_policy`.
|
||||
|
||||
The semantic-context fingerprint value must cover both its policy identity and
|
||||
the radius value. Metadata and fingerprints must not contain names, source
|
||||
text, model output, paths, timestamps, or invocation-specific data. Existing
|
||||
pipeline scoping will prefix these local names and incorporate them into
|
||||
checkpoint identity. The new fingerprints intentionally make prior NPC
|
||||
normalize checkpoints cold misses.
|
||||
|
||||
### Deterministic Preprocessing
|
||||
|
||||
Refactor the current `normalizeList` implementation into a form that preserves
|
||||
all existing behavior before the semantic call:
|
||||
|
||||
- display whitespace normalization;
|
||||
- stable-ID recomputation;
|
||||
- source-reference canonicalization and exact deduplication;
|
||||
- equal-`identity.ComparisonKey` consolidation;
|
||||
- first-occurrence output anchoring;
|
||||
- source-reference union; and
|
||||
- existing field, ID, evidence, and duplicate warnings.
|
||||
|
||||
Each preprocessed record must retain the sorted original input indexes it
|
||||
represents and its earliest input index. This provenance is internal only and
|
||||
supports deterministic application, ordering, scopes, and warnings.
|
||||
|
||||
Return the deterministic result immediately, without an LLM call, when fewer
|
||||
than two semantically eligible records remain. Preserve `nil` versus empty list
|
||||
behavior and do not mutate or alias the merge input.
|
||||
|
||||
### Structured Completion
|
||||
|
||||
For two or more eligible records on each framework attempt:
|
||||
|
||||
1. Build the candidate and transcript-window inputs from Stage 2 using the
|
||||
production radius constant.
|
||||
2. Call the injected scheduled client exactly once with:
|
||||
- `StageName: Key`;
|
||||
- the normalization prompt ID and v1 version;
|
||||
- `ProfileID: req.LLMProfile`;
|
||||
- `SessionID: req.SessionID`; and
|
||||
- only the `candidates` and `transcript` input materials.
|
||||
3. Decode into the private proposal DTO.
|
||||
4. If completion returns `contracts.ErrInvalidStructuredOutput`, return the
|
||||
deterministic pre-LLM artifact as `Value` with a normalize retry directive.
|
||||
5. Wrap every other completion failure with normalizer and completion context,
|
||||
relying on the existing Scriptorium redaction boundary and never appending
|
||||
raw response content.
|
||||
|
||||
Do not perform an LLM call per candidate pair or group. Do not make a second
|
||||
module-local repair call. A framework retry invokes the normalizer again and
|
||||
therefore makes one new proposal call with the same deterministic inputs.
|
||||
|
||||
### Proposal Validation
|
||||
|
||||
Validate the complete decoded proposal and isolate independently safe groups
|
||||
before applying them:
|
||||
|
||||
1. Build the unique map from `identity.ComparisonKey(displayName)` to eligible
|
||||
preprocessed record position.
|
||||
2. For each response group, resolve every member and `canonical_name`.
|
||||
3. Mark a group locally invalid when it has:
|
||||
- fewer than two distinct members;
|
||||
- a blank, unknown, or ambiguous member;
|
||||
- a repeated resolved member;
|
||||
- a blank, unknown, or ambiguous canonical name; or
|
||||
- a canonical record that is not a group member.
|
||||
4. For every group, including locally invalid groups, retain each member that
|
||||
resolves uniquely and count ownership of those resolved record positions
|
||||
across the complete proposal.
|
||||
5. Mark every group containing a record position owned by more than one
|
||||
proposal group as conflicting. Discard every participant rather than
|
||||
choosing a winner. Applying this rule to every group also discards all
|
||||
groups in a chained conflict.
|
||||
6. Classify a group as safe only when it is locally valid and non-conflicting.
|
||||
Apply every safe group independently; discard every other group.
|
||||
7. If any group is discarded, return the artifact containing the safe groups
|
||||
as a retryable fallback. If none is discarded, return the applied artifact
|
||||
through the ordinary successful result path. Response order may order
|
||||
attempt diagnostics but must never choose artifact order or conflict
|
||||
winners.
|
||||
|
||||
Name resolution may accept only the equivalences already implemented by
|
||||
`identity.ComparisonKey`. The validator must not infer which supplied name the
|
||||
model intended.
|
||||
|
||||
Add:
|
||||
|
||||
- `npc_semantic_proposal_invalid` as the attempt-local retry reason for
|
||||
structurally or semantically invalid proposals;
|
||||
- `npc_semantic_reconciliation_exhausted` as the durable fallback-warning
|
||||
reason; and
|
||||
- `npc_normalization_warnings_omitted` for total accepted-warning truncation.
|
||||
|
||||
For a partially unsafe semantic proposal, build one bounded retry message with
|
||||
`shared/diagnostics.Aggregate`. It may identify proposal-group indexes and
|
||||
stable issue categories, including every group participating in an overlap,
|
||||
but must not echo model-returned names, transcript content, source paths, or
|
||||
raw completion text.
|
||||
|
||||
Apply the safe groups to the deterministic pre-LLM artifact and return that
|
||||
partial result as the retry directive's safe `Value`. If structured output
|
||||
could not be decoded or no group is safe, `Value` remains the deterministic
|
||||
pre-LLM artifact. Ordinary deterministic preprocessing warnings and warnings
|
||||
for safely applied groups remain in `Warnings`.
|
||||
|
||||
Put one bounded warning with reason
|
||||
`npc_semantic_reconciliation_exhausted` in `FallbackWarnings`; it should state
|
||||
the exact number of groups omitted from the final proposal after the configured
|
||||
attempt budget. For structurally invalid output, where no decoded group count
|
||||
exists, state that the semantic proposal could not be applied. That warning
|
||||
becomes durable only when the runner accepts final fallback.
|
||||
|
||||
Each framework attempt starts from the same merge output and performs
|
||||
deterministic preprocessing again. Do not retain or accumulate safe groups
|
||||
from earlier attempts. If retry budget is exhausted, the final attempt's safe
|
||||
candidate is authoritative.
|
||||
|
||||
### Deterministic Group Application
|
||||
|
||||
Apply approved groups against the preprocessed record positions:
|
||||
|
||||
- Iterate records in existing order.
|
||||
- Emit one consolidated record at the earliest member position and suppress
|
||||
later members.
|
||||
- Take the display name from the model-selected existing canonical member,
|
||||
regardless of which member supplies the output anchor.
|
||||
- Union every member's existing source references and canonicalize them with
|
||||
the current source-document order.
|
||||
- Derive a new stable ID from the selected display name.
|
||||
- Combine the represented original input-index sets without loss.
|
||||
- Leave unrelated records present and in their prior relative order.
|
||||
|
||||
Use the existing `duplicate_npc_collapsed` reason code for both deterministic
|
||||
equal-name and approved semantic consolidation. The semantic warning scope is
|
||||
the earliest represented input index. Its bounded message should identify
|
||||
input indexes and, when different, the input index that supplied the canonical
|
||||
display name; it should not include raw names or transcript text.
|
||||
|
||||
Build duplicate details with `shared/diagnostics.Aggregate`, and pass the
|
||||
complete NPC normalization warning list through
|
||||
`diagnostics.LimitWarnings(..., "npcs",
|
||||
ReasonCodeNPCNormalizationWarningsOmitted)` before returning. This total cap
|
||||
applies to old and new normalizer warnings and must preserve deterministic
|
||||
warning order. When returning a retry directive, reserve one position within
|
||||
`diagnostics.MaxWarnings` for the exhaustion fallback warning. If deterministic
|
||||
or safely applied-group warnings require truncation, include their
|
||||
omission-summary warning before the reserved exhaustion warning so final
|
||||
accepted fallback still contains at most `MaxWarnings` warnings and reports
|
||||
the exact omitted-warning count.
|
||||
|
||||
For a proposal whose groups are all safe, return the applied artifact with
|
||||
`Retry: nil`.
|
||||
The configured normalize validator chain receives it through the ordinary
|
||||
framework path. For a final fallback, Stage 1 sends the safe candidate through
|
||||
that same validator chain. No validator, codec, artifact schema, or warning
|
||||
JSON shape changes are needed.
|
||||
|
||||
### Stage 3 Tests
|
||||
|
||||
Use a small recording LLM fake at the normalizer boundary and real
|
||||
deterministic collaborators:
|
||||
|
||||
- A nil client fails construction before execution, while normal construction
|
||||
exposes valid prompt and schema metadata.
|
||||
- Nil receiver, nil context, canceled context, operational completion failure,
|
||||
and JSON input-encoding failure follow content-safe error paths.
|
||||
- Invalid structured output returns a retryable deterministic fallback rather
|
||||
than an ordinary error.
|
||||
- Zero or one eligible candidate performs no LLM call and retains existing
|
||||
deterministic normalization.
|
||||
- A valid proposal consolidates differently named records, selects an existing
|
||||
canonical display name, recomputes its ID, unions and orders all member
|
||||
evidence, occupies the earliest member position, preserves unrelated order,
|
||||
and does not alias input.
|
||||
- The completion request carries the normalize profile and session ID and
|
||||
contains names but no stable hash IDs.
|
||||
- Empty proposals leave comparison-distinct records unchanged.
|
||||
- Case, whitespace, Unicode compatibility, and supported apostrophe variants
|
||||
resolve through the existing comparison policy.
|
||||
- Unknown, blank, repeated, singleton, non-member-canonical, and overlapping
|
||||
groups are discarded and return a retryable safe fallback.
|
||||
- A valid independent group is applied even when a disjoint group is invalid.
|
||||
- A locally valid group sharing a resolved member with an invalid group is
|
||||
discarded, and chained overlaps discard every group in the connected
|
||||
conflict set without response-order dependence.
|
||||
- A partially safe final attempt preserves its applied groups in the accepted
|
||||
fallback and reports the exact omitted-group count; a later retry starts
|
||||
from the original merge output rather than accumulating earlier groups.
|
||||
- Invalid or ineligible source-reference records cannot participate in semantic
|
||||
groups and remain available for final validator rejection.
|
||||
- Surrounding context never becomes output evidence.
|
||||
- Retry diagnostics and fallback warnings remain valid UTF-8 and individually
|
||||
bounded. Accepted deterministic, applied-group, and exhaustion warnings are
|
||||
subject to the total warning cap with an accurate omission summary.
|
||||
- Metadata and fingerprints contain all stable policy and asset identities but
|
||||
no names, source text, raw model responses, or paths.
|
||||
|
||||
Do not add tests that assert exact natural-language prompt wording. Do not
|
||||
repeat generic checkpoint-identity tests: package tests should prove that the
|
||||
normalizer contributes the required stable fingerprints, while existing
|
||||
framework tests continue to own the rule that fingerprint changes alter
|
||||
identity.
|
||||
|
||||
Run:
|
||||
|
||||
```sh
|
||||
go test ./internal/modules/dnd/normalize/npcs
|
||||
go test ./internal/modules/dnd/validate/npcs/...
|
||||
go test ./internal/framework/pipeline
|
||||
```
|
||||
|
||||
### Stage 3 Completion Gate
|
||||
|
||||
Each normalizer invocation performs at most one scheduled LLM call over
|
||||
eligible document-level candidates. It independently applies safe,
|
||||
non-conflicting name-based groups, requests a framework retry when any group
|
||||
is discarded, supplies the resulting safe candidate as fallback, preserves
|
||||
artifact integrity and evidence, emits bounded diagnostics, and contributes
|
||||
complete checkpoint identity.
|
||||
|
||||
## Stage 4: Production Registration, Pipeline Integration, And Checkpoint Coverage
|
||||
|
||||
### Goal
|
||||
|
||||
Make the new normalize prompt available in production composition and prove
|
||||
that reconciled NPCs cross an ordered step barrier correctly.
|
||||
|
||||
### Registration
|
||||
|
||||
Update `internal/modules/dnd/register/modules.go` to register NPC normalize
|
||||
prompt and schema assets in addition to the existing NPC extraction assets.
|
||||
Use a distinct registration label and call the normalize package's
|
||||
`RegisterPromptAssets`.
|
||||
|
||||
Update registrar tests to cover the new prompt filesystem and schema
|
||||
registration under `dnd.npcs.normalize`. Preserve existing module key,
|
||||
artifact-kind registration, normalizer spec, reference slots, options, and
|
||||
default normalize validator chain.
|
||||
|
||||
Update every production test constructor or direct normalizer call for the new
|
||||
LLM-injected constructor. Preparation tests should continue using the existing
|
||||
production fake client; no real provider is permitted.
|
||||
|
||||
### Existing Fake And Fixture Adaptation
|
||||
|
||||
Several integration fakes currently assume that every `dnd/npcs` pipeline LLM
|
||||
request is extraction. Change them to dispatch by `PromptID`:
|
||||
|
||||
- extraction prompt requests continue returning the existing private NPC
|
||||
extraction response;
|
||||
- normalization prompt requests return a private
|
||||
`duplicate_groups` response; and
|
||||
- unrelated prompt IDs retain their existing behavior.
|
||||
|
||||
For tests unrelated to semantic reconciliation, return
|
||||
`{"duplicate_groups":[]}` and update only call-count assertions whose observable
|
||||
contract now includes the one document-level normalize call. Do not duplicate
|
||||
proposal edge cases already owned by normalize-package tests.
|
||||
|
||||
### Assembled Behavior
|
||||
|
||||
Add one focused assembled D&D workflow case with two extraction chunks that
|
||||
produce a short proper name and a role-qualified variant for the same NPC.
|
||||
Have the normalize fake propose one group using those display names. Prove:
|
||||
|
||||
- extraction still produces and merge still preserves both candidates;
|
||||
- normalization emits one canonical NPC at the earliest member position;
|
||||
- its evidence is the canonical union of both chunks' source references;
|
||||
- its ID is derived from the selected canonical display name;
|
||||
- the accepted normalized artifact crosses the existing ordered step barrier;
|
||||
- a downstream NPC consumer receives a names-only registry projection
|
||||
containing the canonical name once; and
|
||||
- no stable NPC hash ID or transcript content leaks into manifest component
|
||||
metadata or reference provenance.
|
||||
|
||||
Keep this integration test focused on collaboration among extraction, merge,
|
||||
normalization, validation, and generated handoff. Proposal-malformation,
|
||||
context-window edge cases, and schema parsing remain at their narrower owners.
|
||||
|
||||
### Checkpoint And Retry Coverage
|
||||
|
||||
Update the existing NPC preparation and grounded-pipeline assertions to include
|
||||
the scoped normalize fingerprints for prompt, response schema, identity
|
||||
policy, normalization policy, and semantic-context policy. Confirm the
|
||||
normalize prompt/schema metadata appears under the normalizer manifest
|
||||
component.
|
||||
|
||||
Rely on existing generic checkpoint tests for ordering and identity sensitivity;
|
||||
do not add another generic fingerprint mutation test. The assembled NPC test
|
||||
should demonstrate that the new fingerprints are present before execution.
|
||||
Existing checkpoint construction will then cause a one-time cold miss for old
|
||||
NPC normalize checkpoints.
|
||||
|
||||
Stage 1 framework tests own the generic retryable-fallback mechanism. Add only
|
||||
the narrow NPC integration coverage needed to prove feature wiring:
|
||||
|
||||
- with normalize retries omitted, one partially unsafe NPC proposal makes one
|
||||
normalize completion call, accepts its safe partial fallback, and promotes
|
||||
the exhaustion warning; and
|
||||
- with `retries: 1`, a partially unsafe NPC proposal followed by a completely
|
||||
safe proposal makes two normalize completion calls, applies the later
|
||||
proposal, and does not promote the earlier fallback warning.
|
||||
|
||||
Do not repeat every proposal-invalidity category or generic final-validator
|
||||
case at integration scope. Existing NPC fakes should return a fresh
|
||||
normalization response for each actual attempt.
|
||||
|
||||
### Stage 4 Tests
|
||||
|
||||
Run:
|
||||
|
||||
```sh
|
||||
go test ./internal/modules/dnd/register
|
||||
go test ./internal/modules/integration
|
||||
go test ./internal/cli
|
||||
go test ./internal/framework/pipeline
|
||||
```
|
||||
|
||||
### Stage 4 Completion Gate
|
||||
|
||||
Production preparation registers and constructs the LLM-assisted normalizer,
|
||||
representative pipelines route both NPC prompt identities correctly, prepared
|
||||
identity contains all new fingerprints, and downstream generated references
|
||||
observe the reconciled registry.
|
||||
|
||||
## Stage 5: Maintained Configuration And Canonical Documentation
|
||||
|
||||
### Goal
|
||||
|
||||
Update each canonical owner to describe the implemented behavior without
|
||||
duplicating volatile contracts.
|
||||
|
||||
### Configuration And Example
|
||||
|
||||
Update `examples/dnd-complete.config.yml` so the NPC normalize binding uses
|
||||
object form and demonstrates the normalize-stage profile and retry knobs:
|
||||
|
||||
```yaml
|
||||
normalize:
|
||||
module: dnd/npcs
|
||||
llm_profile: gemini-2-flash
|
||||
retries: 2
|
||||
```
|
||||
|
||||
Do not add a radius or context-mode option. Keep the minimal example unchanged
|
||||
unless its current NPC composition requires object form for correctness.
|
||||
Retain the maintained-example resolution tests as the owner of example
|
||||
validity.
|
||||
|
||||
Update `docs/config.md` to:
|
||||
|
||||
- identify `dnd/npcs` normalization as LLM-assisted;
|
||||
- explain that `llm_profile` and `retries` use the ordinary normalize binding
|
||||
contract;
|
||||
- retain and clearly state the global `retries: 0` default, including that it
|
||||
allows one initial normalization attempt and no additional attempts;
|
||||
- retain the absence of normalizer references and options;
|
||||
- state that context selection is currently module policy rather than
|
||||
configuration; and
|
||||
- accurately list the affected prompt default without duplicating the complete
|
||||
example.
|
||||
|
||||
### Operations, Integration, And Internals
|
||||
|
||||
Update:
|
||||
|
||||
- `docs/operations.md` with the document-level NPC semantic-normalization call,
|
||||
its place before the ordered generated-reference barrier, configured retry
|
||||
behavior, safe partial fallback after invalid-proposal exhaustion, and
|
||||
checkpoint reuse;
|
||||
- `docs/internal/pipeline.md` with the provider-neutral invalid structured
|
||||
output classification, normalize retry directive, attempt accounting,
|
||||
accepted-only warning promotion, fallback validation, debug behavior, and
|
||||
checkpoint rules;
|
||||
- `docs/integrations/dnd-npc-artifacts.md` with externally observable
|
||||
name-based consolidation, canonical selection, evidence union, ordering,
|
||||
warnings, and unchanged durable v1 artifact shape;
|
||||
- `docs/internal/modules.md` with deterministic preprocessing, candidate
|
||||
eligibility, prompt inputs, private response contract, proposal validation,
|
||||
application, diagnostics, metadata, and fingerprints;
|
||||
- `docs/internal/llm.md` with the normalization prompt's stable and variable
|
||||
message tiers, cache boundaries, windowed transcript input, and distinct
|
||||
prompt/schema identity; and
|
||||
- `docs/internal/overview.md` with a concise current package responsibility.
|
||||
|
||||
Do not put implementation details in configuration or operations documentation.
|
||||
Do not describe the private proposal schema as part of the durable NPC
|
||||
integration schema. Keep the generic LLM-assisted deduplication item in
|
||||
`future.md`; this feature is intentionally the first D&D-specific use and does
|
||||
not complete that generic work.
|
||||
|
||||
After all implementation and documentation checks pass:
|
||||
|
||||
- change the feature roadmap status to `Implemented` and preserve it as the
|
||||
feature contract and rationale; and
|
||||
- change this implementation plan's status to `Completed`.
|
||||
|
||||
### Documentation And Test Policy
|
||||
|
||||
Update existing behavioral and contract tests where their owned behavior
|
||||
changes. Do not add prompt-language change detectors, broad golden snapshots,
|
||||
live-provider tests, redundant schema cases at higher layers, or assertions
|
||||
against private helper shape.
|
||||
|
||||
Run:
|
||||
|
||||
```sh
|
||||
go test ./internal/cli
|
||||
go test ./internal/modules/integration
|
||||
```
|
||||
|
||||
### Stage 5 Completion Gate
|
||||
|
||||
The maintained complete example resolves, canonical current-behavior
|
||||
documentation describes the shipped feature in its proper owners, future work
|
||||
still distinguishes generic deduplication, and both roadmap statuses agree with
|
||||
implementation state.
|
||||
|
||||
## Final Verification
|
||||
|
||||
From the repository root, run:
|
||||
|
||||
```sh
|
||||
git diff --check
|
||||
go test ./...
|
||||
go vet ./...
|
||||
go build ./cmd/notarius
|
||||
go test -race ./internal/modules/dnd/normalize/npcs ./internal/modules/dnd/register ./internal/framework/pipeline ./internal/cli ./internal/modules/integration
|
||||
```
|
||||
|
||||
Review the final diff for:
|
||||
|
||||
- no durable artifact or public configuration schema change;
|
||||
- no model-facing stable NPC IDs;
|
||||
- no full-transcript normalize input;
|
||||
- no transcript, name, path, or raw response content in metadata,
|
||||
fingerprints, provenance, or content-safe errors;
|
||||
- no fuzzy matching or response-order-dependent application;
|
||||
- no retryable-result expansion beyond the demonstrated normalize fallback or
|
||||
domain behavior in the framework contract;
|
||||
- no real-provider or prompt-wording change-detector tests; and
|
||||
- no documentation describing future behavior as already implemented before
|
||||
the final implementation stage is complete.
|
||||
|
||||
## Assumptions
|
||||
|
||||
- The current source document supplied by the runner remains the authoritative
|
||||
source-unit order; unit IDs are identifiers and may be non-monotonic.
|
||||
- Default production extraction validators normally provide valid NPC
|
||||
references before merge. The normalizer still excludes invalid records from
|
||||
semantic reconciliation so custom validator chains cannot make them unsafe
|
||||
model operands.
|
||||
- A deterministically validated partial reconciliation after retry exhaustion
|
||||
is preferable to failing the complete NPC lane or discarding independent
|
||||
safe groups. Operators can inspect its durable bounded warning and
|
||||
attempt-local debug diagnostics.
|
||||
- Invalid structured output is distinguishable from operational LLM failure.
|
||||
The former uses retryable fallback; the latter retains existing error
|
||||
behavior.
|
||||
- The configuration-wide retry default remains zero. Any future decision to
|
||||
change it to one or two is separate work with application-wide cost,
|
||||
latency, and failure-semantics consequences.
|
||||
- One document-level request is sufficient for the current workload. Batching,
|
||||
token budgeting, retrieval, full-transcript mode, and context-radius
|
||||
configuration require separate evidence and scope.
|
||||
- Checkpoint compatibility is subordinate to semantic correctness. No
|
||||
migration or cleanup is required for normalize checkpoints made cold by the
|
||||
new component fingerprints.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None. The feature roadmap and the decisions above are sufficient to implement
|
||||
the work without additional product or architecture choices.
|
||||
Reference in New Issue
Block a user