Clean up completed roadmap work

This commit is contained in:
2026-07-25 21:50:39 -05:00
parent 7b077c269d
commit 5a968b64eb
4 changed files with 4 additions and 885 deletions

View File

@@ -1,328 +0,0 @@
# D&D NPC Semantic Normalization
## Status
Implemented.
## Purpose
Reconcile NPC records that extraction produced under different display names
when the complete transcript establishes that they represent the same
individual. This work addresses identities split across scene or chunk
boundaries while preserving the existing rule that extraction records narrow
source evidence and normalization owns document-wide reconciliation.
The model should make only the semantic identity judgment. Deterministic code
must continue to own identity derivation, proposal validation, artifact
mutation, evidence preservation, ordering, diagnostics, and final validation.
## Target Behavior
The target `dnd/npcs` normalization contract combines:
- the current deterministic display-name normalization, stable-ID derivation,
source-reference canonicalization, and equal-comparison-key consolidation;
- one LLM determination of whether remaining, distinctly named NPC records
represent the same individual and which existing display name is canonical;
and
- deterministic proposal validation and application followed by the configured
normalize validator chain.
Each configured normalization attempt should make at most one semantic pass
over the merged document-level NPC list, not one pass per extraction chunk or
candidate pair. The pass should be skipped when fewer than two distinct
candidate identities remain after deterministic preprocessing.
False consolidation is more damaging than a missed consolidation. Prompt
policy should therefore require affirmative contextual evidence that names
identify the same individual and should prefer no group when identity remains
ambiguous.
Canonical selection should favor the most complete stable proper name supported
by the transcript. A complete proper name is preferable to an abbreviation,
while an unadorned proper name is preferable to the same name plus a contextual
class, role, title, or relationship descriptor unless the transcript establishes
that descriptor as part of the character's name. The model must still select
one supplied display name rather than synthesize a better one.
## Model Proposal Contract
The private structured response should contain only proposed duplicate groups:
```json
{
"duplicate_groups": [
{
"members": [
"Billy",
"Billy the druid"
],
"canonical_name": "Billy"
}
]
}
```
`members` identifies candidates by their supplied display names.
`canonical_name` selects one existing member as the retained display name. An
empty `duplicate_groups` array is a valid determination that no sufficiently
supported duplicates exist.
The model must not receive or return deterministic `npc:sha256:` IDs. Those IDs
are long, non-semantic implementation identities and remain exclusively owned
by deterministic code. The model must not invent a replacement name, rewrite
an NPC record, propose new source references, or return a complete replacement
artifact.
Before the call, deterministic preprocessing ensures that each supplied display
name has a unique NPC identity comparison key. Returned names may resolve using
the existing comparison-key equivalences for case, whitespace, Unicode
compatibility, and supported apostrophes. Resolution must not use fuzzy,
substring, edit-distance, embedding, or other approximate matching.
An applicable group must:
- contain at least two distinct, known members;
- resolve every member uniquely against the supplied candidate set;
- select a known `canonical_name` that belongs to the group;
- contain no repeated member; and
- share no resolved member with any other proposed group.
Unknown or ambiguous names, singleton groups, invalid canonical selections, and
other unsafe groups must not mutate the artifact. Groups are otherwise
independent: a locally valid group may be applied when none of its resolved
members appears in any other proposed group.
Overlap is a proposal conflict even when one participating group is already
locally invalid. Deterministic code must discard every group in the connected
conflict set rather than selecting a winner by response order. Locally valid,
non-conflicting groups remain safe to apply. This permits useful partial
reconciliation without allowing an unsafe group to influence an NPC identity
that another group would mutate.
## Retry And Safe Fallback
An invalid or partially unsafe private proposal is a retryable normalization
attempt. This includes:
- structurally invalid model output classified by the LLM boundary as an
invalid structured completion; and
- a structurally valid response containing any locally invalid or conflicting
group.
The framework owns retry counting and attempt diagnostics. The normalizer
returns a safe candidate together with a bounded, content-safe retry
diagnostic. For structurally invalid output, that candidate is the
deterministic pre-LLM result. For a decoded proposal with unsafe groups, it
also includes every independently valid, non-conflicting group from that
attempt. When normalize retries remain, the framework invokes the normalizer
again from the same merged input; safe groups are not accumulated across
attempts. A later completely safe proposal is applied normally.
When the configured attempt budget is exhausted, the framework validates and
accepts the final attempt's safe fallback instead of failing or rejecting the
NPC lane, provided that fallback passes the configured normalize validators.
It promotes one durable warning stating that one or more proposed groups were
omitted. The fallback may therefore contain a safe partial reconciliation, or
only the deterministic pre-LLM result when no group could be applied. Earlier
retry diagnostics, candidate values, and warnings remain attempt-local in
debug artifacts.
The normal module-binding default remains `retries: 0`, meaning one total
normalization attempt and immediate fallback after its invalid proposal.
Retries occur only when the user configures a positive normalize retry count.
For example, `retries: 2` permits the initial proposal plus two additional
attempts before fallback.
Transport, authentication, prompt-preparation, cancellation, input-encoding,
and other operational errors are not safe-proposal failures. They retain the
existing module/framework error behavior rather than being converted into an
accepted fallback.
## Transcript Context Policy
The model should receive every candidate's display name and source references,
together with transcript windows derived from those references. It should not
receive the complete transcript by default.
Each window must contain:
- the complete inclusive source range cited by the NPC record;
- up to two source units immediately before the cited range; and
- up to two source units immediately after the cited range.
The surrounding-unit count is defined once as a named module policy constant
with value `2`; it is not a user-accessible configuration field. Window
construction consumes that value through one clear boundary so a later
configuration option can replace the fixed value without changing prompt or
reconciliation contracts. The context-radius policy participates in module
metadata and checkpoint identity so changing it invalidates incompatible
normalize checkpoints.
The value was selected by reviewing the July 19 evaluation transcript. Its NPC
citations are generally self-contained; where additional context is useful,
two surrounding units capture the relevant question-and-answer exchange,
speaker transition, or short anaphora chain. A third unit frequently begins a
separate joke or table exchange and adds distraction without improving the
identity evidence.
Window expansion must use source-document positions rather than arithmetic on
unit IDs. It must clamp at document boundaries, preserve document order, and
coalesce overlapping or adjacent expanded windows without duplicating units.
The prepared model input must distinguish originally cited units from
surrounding context.
Only records with a non-empty NPC identity comparison key and a non-empty
source-reference collection that is wholly valid against the current source
document are eligible for semantic reconciliation. Ineligible records remain
unchanged so the configured deterministic validators retain ownership of their
rejection. They are not supplied to the model and cannot participate in a
proposed group.
Surrounding units inform the semantic decision but do not automatically become
durable NPC evidence. An applied group unions only the source references
already present on its member records. The LLM cannot add references from the
context window.
Full-transcript mode and a configurable context radius may be evaluated later.
They are not part of this scope. Provider prompt-cache behavior should be
measured rather than assumed before expanding context solely for cache
economics.
## Deterministic Application
For each approved group, deterministic code should:
- retain the model-selected existing display name;
- union source references from every group member;
- canonicalize and exact-deduplicate those references in source-document order;
- derive the resulting stable NPC ID from the retained display name under the
existing NPC identity policy; and
- emit one bounded duplicate-collapse warning describing the applied group.
The consolidated record should occupy the earliest original member position so
model response ordering cannot reorder the artifact. Unrelated records must
remain present and retain their relative order. Every input record must be
represented by exactly one output record, either unchanged or through one
approved consolidation.
The existing NPC shape, identity, source-reference, schema, and relatedness
validators remain the final artifact boundary. No LLM-backed validator is
needed: model judgment occurs in the normalizer, while deterministic validators
continue to enforce the durable artifact contract.
## Prompt, Provenance, And Diagnostics
NPC semantic-normalization instructions and the private response schema should
be module-owned prompt assets. Stable D&D-wide identity or transcript guidance
may be reused through the existing shared prompt-asset mechanism where its
meaning is genuinely common.
The normalizer should use the configured normalize-stage LLM profile and the
application-wide scheduled LLM client. Invalid structured output and unsafe
semantic proposals use the framework-owned retryable-fallback contract;
operational failures retain existing error semantics.
Manifest metadata and component checkpoint fingerprints should identify the
semantic-normalization policy, prompt identity, private response-schema
identity, deterministic NPC identity policy, and evidence-context policy.
They must not contain transcript text, NPC names, source paths, raw model
responses, or other source content.
Warnings and preparation or runtime errors must follow the established bounded
diagnostic and content-safety policies. Debug artifacts may retain the normal
attempt-local model request, response, proposal decisions, and warnings under
the existing debug sensitivity contract.
## Configuration And Documentation
The maintained complete D&D configuration should demonstrate a normalize-stage
LLM profile and an explicit positive retry count for the NPC lane. The
application-wide default remains zero additional retries. No new configuration
field is introduced for the context radius in this scope.
Canonical documentation ownership is:
- Configuration owns the LLM-backed `dnd/npcs` selection and its use of the
existing normalize binding's profile and retry fields.
- Operations owns the document-level semantic reconciliation call and its
checkpoint behavior.
- Pipeline internals own the provider-neutral invalid-structured-output
classification and normalize retryable-fallback contract.
- The NPC integration contract owns externally observable consolidation,
evidence, ordering, and warning behavior.
- Internal LLM and module documentation own the proposal boundary,
context-window construction, deterministic application, metadata, and
fingerprints.
- The broader generic LLM-assisted deduplication item in
[future.md](future.md) remains future work until another artifact demonstrates
that extracting a shared generic facility is worthwhile.
## Quality Expectations
Tests should protect the semantic and safety boundaries through deterministic
LLM fakes rather than live-provider calls. Coverage should demonstrate:
- distinct display variants can be consolidated when the model proposes a
valid group;
- model-facing requests and responses use display names rather than stable
hash IDs;
- the selected existing canonical name controls ID derivation while all member
evidence is preserved;
- window construction uses source-document position, handles document edges,
coalesces overlap, and distinguishes cited evidence from context;
- no surrounding context is promoted into durable source references;
- empty proposals and fewer-than-two-candidate inputs preserve deterministic
normalization behavior;
- unknown, ambiguous, repeated, overlapping, and otherwise malformed semantic
groups cannot corrupt or reorder the artifact;
- locally valid groups are applied independently, while every group that
participates in a resolved-member conflict is discarded;
- an invalid proposal consumes only configured retry budget, a later valid
proposal can succeed, and exhaustion accepts the final attempt's safe
fallback with one durable warning;
- the default zero-retry binding makes exactly one proposal attempt before
fallback;
- warnings and failures remain bounded and do not expose transcript content;
- prompt, schema, policy, or context-policy changes invalidate relevant
checkpoint reuse; and
- an assembled ordered D&D pipeline supplies the reconciled NPC registry to
downstream consumers.
Prompt tests should assert prepared message structure, supplied materials, and
cache-boundary behavior. They must not act as change detectors for particular
words or phrases in natural-language prompt text.
Human review of representative transcripts should compare missed and false
consolidations, latency, input-token cost, and provider cache use. Probabilistic
model quality is an evaluation activity, not a deterministic CI assertion.
## Non-Goals
This scope does not:
- change the durable NPC artifact schema;
- add a durable alias collection or identity history;
- expose the context radius or full-transcript selection as configuration;
- ask an LLM to read, reproduce, or derive stable NPC IDs;
- permit model-authored source references or arbitrary replacement records;
- add fuzzy deterministic name matching;
- add an LLM-backed normalize validator;
- change the global default retry count;
- reconcile NPCs concurrently across lanes or runs;
- add canonical NPC IDs to downstream artifact schemas;
- implement a general DAG or implicit cross-lane dependency; or
- implement the generic cross-artifact deduplication normalizer described in
`future.md`.
## Completion Criteria
The scope is complete when `dnd/npcs` can use one document-level LLM proposal
to reconcile differently named records conservatively, deterministic code
validates and independently applies only safe non-conflicting name-based
groups, the resulting NPC preserves all member evidence under a newly derived
stable ID, the two-unit context policy is centralized and fingerprinted,
invalid or partially unsafe proposals use configured framework retries and
then an accepted safe fallback, downstream ordered steps receive the
reconciled registry, and the canonical current-behavior documentation reflects
the implemented contract.

View File

@@ -1,365 +0,0 @@
# D&D NPC Semantic Normalization Follow-up Plan
## Status
Ready for implementation.
## Objective
Close the four findings from the post-implementation review of
[D&D NPC Semantic Normalization](dnd-npc-semantic-normalization.md) without
changing its durable artifact schema, deterministic identity rules,
conflict-aware partial-application policy, retry accounting, or generated
reference handoff.
Complete the four follow-up stages below in order. Each stage should leave its
affected packages passing before the next stage begins.
## Completed Work Summary
The original five-stage implementation is complete:
1. **Framework retryable normalize fallback.** The framework now classifies
invalid structured output separately from operational LLM failures.
Normalizers can return a safe candidate and retry directive; the runner owns
retry accounting, attempt debug artifacts, final fallback validation,
accepted-only warning promotion, and checkpoint recording.
2. **NPC prompt and context material.** `dnd/npcs` owns a private normalization
prompt and structural response schema. It sends display-name candidates and
coalesced transcript windows with a fixed two-unit radius, without exposing
stable NPC IDs or using the full source input.
3. **Semantic proposal validation and application.** The normalizer performs
deterministic preprocessing, resolves returned names through the existing
comparison-key policy, applies independent safe groups, discards locally
invalid and conflicting groups, preserves evidence and ordering, and emits
bounded D&D diagnostics.
4. **Production integration and checkpoint identity.** Production registration
includes the normalize prompt and schema. Prompt, schema, identity,
normalization, and context policies contribute checkpoint fingerprints.
Assembled tests prove retry behavior and the ordered normalized-NPC handoff
to downstream consumers.
5. **Configuration and documentation.** The maintained complete D&D example
demonstrates an explicit normalize profile and retry count. Configuration,
operations, integration, LLM, module, pipeline, and component documentation
describe the implemented feature.
The completed implementation passed the full test suite, vet, build, diff
checks, and the targeted race suite. The stages below are limited follow-up
corrections and do not reopen the original feature design.
## Decisions Applying To Every Follow-up Stage
- Retain module key `dnd/npcs`, artifact kind `dnd/npc-list`, durable schema
`notarius.dnd.npcs` v1, prompt ID `dnd.npcs.normalize`, response-schema
identity, empty reference-slot contract, strict empty options, and default
normalize validator chain.
- Retain the context radius of two and the existing candidate and transcript
input shapes.
- Retain conflict-aware partial application: independently safe groups may be
applied, every group participating in a resolved-member conflict is
discarded, retries restart from the original merge input, and final
exhaustion accepts only the final attempt's safe candidate.
- Retain `retries: 0` as the application-wide default. Changing that default
remains separate work.
- Retain accepted-only warning promotion, checkpoint format, public
configuration shape, and durable output JSON.
- Do not add an LLM-backed validator, normalizer-local retry loop, generic
deduplication framework, prompt-language change-detector test, or live
provider test.
## Stage 1: Make Context-Material Failures Content-Safe
### Goal
Prevent arbitrary source metadata keys or other source-derived details from
appearing in errors returned by the NPC normalization context builder.
### Implementation
Keep `source.CloneMetadata` unchanged. Its detailed errors remain useful to
trusted direct callers and its own package tests. Redaction belongs at the
external source-material consumption boundary in
`internal/modules/dnd/normalize/npcs/context_material.go`.
At that boundary:
- When copying source-unit metadata fails, return a fixed category such as
`build NPC normalization context: invalid source metadata`.
- Do not wrap the underlying clone error with `%w`.
- The returned error must not include metadata keys, values, traversal
locations, transcript text, NPC names, source IDs, origin paths, encoded
material, concrete Go types derived from source metadata, or the underlying
error text.
- Apply the same fixed-error treatment to candidate and transcript JSON
encoding failures. Those paths are effectively defensive after metadata
validation, but future DTO changes must not create a content-exposure path.
- Keep the negative-radius error detailed: the radius is code-owned policy,
not source content.
- Keep failure timing unchanged. Context construction still fails before the
normalization LLM call, retry directive creation, artifact validation, or
checkpoint recording.
The outer `Normalize` error may retain its existing operation context because
the nested context-builder error is fixed and content-safe.
### Tests
Add narrow boundary coverage in the NPC normalization package:
- Use source metadata with a recognizable sensitive key and an unsupported or
non-finite value.
- Assert that normalization fails before the LLM fake is called.
- Assert that the error contains a stable context-material failure category
but none of the recognizable key, value, transcript text, NPC names, source
ID, or origin path.
- Retain direct `source.CloneMetadata` tests as the owner of detailed cloning
behavior; do not duplicate its type matrix at the normalizer boundary.
Do not assert the complete error sentence. Protect the fixed category and
absence of source-derived sentinels.
### Completion Gate
Every context-material construction failure exposed by `dnd/npcs` is
content-safe, while trusted source-package callers retain detailed clone
diagnostics.
## Stage 2: Remove Conflicting Shared Identity Prompt Guidance
### Goal
Give the NPC normalizer one unambiguous canonical-name policy and avoid telling
it about campaign references that it does not receive.
### Implementation
The existing `common-dnd-identity.md` fragment serves extraction prompts: it
asks for the most specific supported in-world identity and discusses campaign
and registry references. It is not wholly applicable to reconciliation of an
already extracted NPC list.
Remove `common-dnd-identity.md` from the NPC normalization prompt only:
- Remove its message from
`assets/prompts/dnd.npcs.normalize.yaml`.
- Remove it from the normalization package's `PromptAssetManifest.SharedFiles`.
- Do not change the shared fragment or any extraction prompt that currently
consumes it.
- Do not replace it with a copied or newly shared identity fragment. NPC
extraction has already established that the candidates are in-world NPC
records; normalization needs only the package-owned duplicate and canonical
selection instructions.
- Retain `common-dnd-system.md` and `common-dnd-transcript.md`.
After removing the identity message, preserve two intentional prompt-cache
tiers:
1. Mark the common D&D system message as the end of the stable shared tier.
2. Retain the existing boundary after the package-owned normalization
instructions.
The package-owned instructions remain authoritative:
- consolidate only when transcript context clearly identifies one individual;
- select a supplied canonical name;
- prefer an unadorned proper name over the same name plus a contextual class,
role, title, or relationship descriptor unless established as part of the
name; and
- prefer no consolidation when identity is ambiguous.
Do not add references, stable IDs, new prompt inputs, new natural-language
repair material, or schema constraints. The prompt digest will change and
prior NPC normalize checkpoints will intentionally become cold misses. Keep
the existing prompt ID and v1 version because checkpoint compatibility is
already content-fingerprinted and no external prompt protocol is versioned by
this private asset.
Update the canonical internal LLM documentation to describe the normalization
prompt's actual shared tier and to stop claiming that it uses shared identity
guidance.
### Tests
Update the existing offline prepared-prompt test to prove:
- the prompt prepares with five messages in the intended order;
- the first cache boundary follows the common D&D system message;
- the second follows the package-owned instructions;
- candidate and transcript messages remain in the variable tail without cache
boundaries;
- only the two declared dynamic inputs are rendered in their corresponding
messages; and
- prompt metadata and fingerprint construction remain valid.
Tests may assert asset/message identity, ordering, inputs, and cache-control
structure. They must not require the presence or absence of particular words
or phrases in natural-language prompt content.
### Completion Gate
The normalization prompt contains no extraction-specific identity or reference
guidance, preserves its intended cache structure, and still carries one clear
package-owned canonical-name policy.
## Stage 3: Enforce Generic Normalize-Retry Diagnostic Bounds
### Goal
Make the framework-level `NormalizeRetry` contract safely reusable by enforcing
the mechanical parts of its diagnostic contract at the runner boundary.
### Contract
Add provider-neutral limits in `internal/framework/contracts`:
- maximum retry reason-code length: 128 bytes;
- maximum retry message length: 4,096 bytes; and
- both values must be valid UTF-8 and nonblank after trimming surrounding
whitespace for the blank check.
Expose the limits as named constants alongside `NormalizeRetry` so normalizer
authors can construct compliant diagnostics without importing a D&D package.
Do not import `internal/modules/dnd/shared/diagnostics` into the framework.
These are encoded-byte limits, not rune limits. The framework validates but
does not truncate or rewrite caller-supplied diagnostics. Silent truncation
could collapse stable reason identities or conceal a module defect.
Content safety remains the module's semantic responsibility: the framework
cannot determine whether otherwise valid text contains transcript content,
credentials, paths, names, or other sensitive values.
### Runner Implementation
Extract the current inline blank check into a small framework-owned validator
for `*contracts.NormalizeRetry` and call it before placing the directive in a
debug payload.
The validator must reject:
- a blank reason code;
- a blank message;
- invalid UTF-8 in either field;
- a reason code exceeding 128 bytes; and
- a message exceeding 4,096 bytes.
Return a fixed module-contract error that identifies the invalid field and
failure category without echoing either supplied value. Preserve existing
retry-loop behavior for module contract errors; do not create a new error
class, rejection, or durable warning.
Do not impose a generic count or text-size policy on `FallbackWarnings` in this
stage. Warning limits remain owned by their existing domain and validator
contracts; broad warning-policy unification is separate work.
Update `docs/internal/pipeline.md` to document the normalize-retry limits and
the division of responsibility between mechanical framework validation and
module-owned content safety.
### Tests
Extend the generic runner retry tests with table-driven boundary cases:
- a reason code and message exactly at their byte limits are accepted;
- one byte over either limit is rejected;
- invalid UTF-8 in either field is rejected;
- blank fields remain rejected; and
- errors contain only fixed field/category context and do not echo supplied
diagnostic sentinels.
Use test-controlled strings and the exported contract constants. Keep the
existing NPC diagnostic tests as the owner of D&D aggregation and omission
behavior.
### Completion Gate
Any typed normalizer receives the same reusable retry mechanism, and the runner
cannot write a blank, invalid-UTF-8, or unbounded retry diagnostic into debug
state.
## Stage 4: Correct Safe-Fallback Documentation
### Goal
Make canonical pipeline documentation accurately describe the generic
framework contract and the NPC module's model-derived safe partial fallback.
### Documentation
Update `docs/internal/pipeline.md`:
- replace `deterministic safe candidate` with `module-supplied safe candidate`
or equivalent wording;
- state that the framework treats the candidate as opaque and relies on the
normalizer plus the configured validator chain for its safety;
- retain the rules that intermediate candidates are attempt-local, only the
final fallback is validated for acceptance, and accepted-only warnings and
checkpoints remain unchanged; and
- keep the generic contract separate from the NPC-specific conflict and
partial-application rules.
Review the affected current-behavior documentation for the same inaccurate
determinism claim. The NPC integration and operations documents may describe
deterministic validation and mutation, but must not imply that every safe
partial fallback is independent of the LLM proposal.
Do not modify the completed feature roadmap except if a factual contradiction
with the implemented target state is discovered. Do not add tests for prose
wording.
### Verification
- Validate local documentation links.
- Confirm configuration defaults, module keys, prompt IDs, and schema versions
against current code.
- Run `git diff --check`.
### Completion Gate
Canonical documentation distinguishes the generic opaque safe-candidate
contract from deterministic NPC validation and application, with no duplicated
or contradictory retry semantics.
## Final Verification
From the repository root, run:
```sh
git diff --check
go test ./...
go vet ./...
go build ./cmd/notarius
go test -race ./internal/framework/contracts ./internal/framework/llm ./internal/framework/pipeline ./internal/modules/dnd/normalize/npcs ./internal/modules/dnd/register ./internal/modules/integration
```
Review the final diff for:
- no durable artifact, checkpoint-format, or public configuration change;
- no model-facing stable NPC IDs or full-transcript normalization input;
- no source-derived metadata keys, values, paths, names, or transcript content
in context-material errors;
- no extraction-specific shared identity fragment in the normalize prompt;
- no prompt-wording change-detector or live-provider test;
- no framework dependency on D&D diagnostic packages;
- no silent retry-diagnostic truncation;
- no change to conflict-aware partial application, retry accounting, warning
promotion, validator ordering, or downstream handoff; and
- documentation that describes only implemented behavior in its canonical
owner.
## Assumptions
- Detailed `source.CloneMetadata` errors remain appropriate inside the trusted
source package; only the NPC context-material boundary requires redaction.
- The normalizer does not need general extraction identity guidance because its
candidates are already accepted NPC artifacts.
- Mechanical UTF-8 and byte limits are appropriate framework invariants;
semantic content safety remains a module obligation.
- Prompt-content and cache-boundary changes intentionally invalidate existing
NPC normalize checkpoints through the current fingerprint mechanism.
- Backward compatibility for development checkpoints is not required.
## Open Questions
None. The review findings and decisions above are sufficient to implement the
follow-up work without additional product or architecture choices.

View File

@@ -1,190 +0,0 @@
# Scene-Aware Combat Extraction
## Status
Implemented. This document preserves the feature contract and rationale; the
current configuration, operational, integration, and internal behavior is
maintained in their canonical documentation.
## Purpose
Use an accepted D&D scene-description artifact as the eligibility boundary for
combat-turn extraction. The combat LLM should run only for a chunk whose exact
scene record has `kind: combat`; every other chunk should be ignored by the
combat extractor. This work should connect the existing scene-description lane,
ordered generated-reference handoff, and combat-turn lane without expanding the
minimal combat-turn artifact contract or introducing D&D policy into generic
pipeline code.
## Target Workflow
The complete D&D pipeline has two ordered steps:
1. The first step extracts and normalizes NPCs and scene descriptions for the
accepted scene chunks.
2. At the step barrier, the accepted NPC and scene-description artifacts become
generated references.
3. The second step supplies the NPC artifact to its existing consumers and the
scene-description artifact to combat-turn extraction.
4. For each chunk, combat-turn extraction performs its existing LLM extraction
only for an exact `combat` classification. Every other chunk produces a
deterministic empty artifact without an LLM call.
The scene-description dependency must remain explicit in pipeline
configuration. The combat extractor must not discover another lane implicitly,
read output files directly, inspect chunker-private state, or require generic
chunk materialization to interpret D&D scene classifications.
## Scene-Description Reference Contract
The combat extractor should declare a required structured reference slot named
`scene_descriptions`. The slot accepts one approved
`dnd/scene-description-list` artifact using the existing durable scene
description schema.
The reference may be supplied as:
- a generated artifact from an earlier ordered step; or
- an external artifact through the existing reference-materialization
boundary.
External artifacts must be decoded and validated before source parsing or LLM
execution. Generated artifacts must cross the existing typed step-handoff
boundary and be validated before use. A bound artifact that is malformed,
incompatible, or internally inconsistent is an error; it must not be treated as
though the slot were unbound.
Pipeline resolution must reject combat extraction when this slot is not bound.
In an ordered same-run workflow, failure of the configured scene-description
producer to yield an accepted normalized artifact must fail the run before the
combat consumer step starts, consistent with existing required generated
reference semantics.
The prepared reference view should be immutable and safe for concurrent
extract jobs. Its metadata and checkpoint identity should be content-safe and
must not expose scene titles, summaries, paths, or source text.
## Chunk Matching And Gating Policy
A scene record authorizes combat extraction for the current chunk only when all
of the following are true:
- the scene ID exactly equals the current accepted chunk ID;
- the scene source ID exactly equals the chunk source ID;
- the scene start and end unit IDs exactly equal the chunk's inclusive source
range; and
- the approved artifact contains exactly one such record.
The normalized scene-description contract already rejects conflicting IDs and
ranges. The combat extractor must nevertheless require the exact match above at
its own decision boundary rather than relying on array position, range overlap,
title, summary, or inferred chronology.
An exact match with `kind: combat` performs the existing combat-turn LLM
extraction. An exact match with `narrative`, `recap`, or `meta` returns a typed
`dnd/combat-turn-list` containing an empty `combat_turns` array without making
an LLM call.
A valid artifact with no exact match, incomplete coverage, or a chunk identity
or range mismatch also returns the deterministic empty artifact. It must not
infer combat eligibility from an overlapping or adjacent scene. Emit a bounded,
content-safe warning for missing or mismatched coverage so operators can
distinguish an intentional non-combat classification from an unavailable exact
classification. The warning must not include scene prose or transcript text.
This policy relies on the scene contract's mixed-scene precedence: any chunk in
which combat is a substantive central activity is classified as `combat`.
Scene-aware gating must not add a second classification policy.
## Extraction, Validation, And Provenance
The deterministic empty result follows the same typed extractor and validator
boundaries as an LLM-produced empty result. It is not a rejection, does not
consume retry budget, and continues through merge and normalization normally.
The durable combat-turn schema remains unchanged.
Existing combat extraction behavior—including prompt assets, NPC grounding,
candidate mapping, validators, retries, warnings, and normalization—remains
unchanged for chunks classified as combat. The scene-description artifact is
control context only:
- it must not be added to the combat prompt;
- its title or summary must not become combat evidence;
- its source references must not be copied into combat turns; and
- it must not create, repair, or classify a combat turn.
Generated-reference provenance and dependency fingerprints should cover the
scene artifact through the existing ordered-handoff machinery. External
references should contribute their existing materialization provenance plus a
component-local semantic fingerprint sufficient to invalidate combat extract
checkpoints when a classification or chunk identity changes. Checkpoint reuse
must never preserve a skipped result after the effective scene classification
changes.
Run manifests may report bounded module metadata such as the number of approved
scene records. They must not contain scene prose or duplicate the referenced
artifact payload.
## Configuration And Documentation
The maintained complete D&D example should bind the normalized
`scene-descriptions` lane from the first step to the `scene_descriptions` slot
in the second step. The minimal example should remain unchanged.
Current-behavior documentation is maintained in these canonical locations:
- Configuration owns the new selectable reference slot and binding example.
- Operations owns the ordered scene-aware workflow.
- The combat-turn integration contract owns externally observable extraction,
empty-result, and provenance behavior.
- Internal pipeline and module documentation own preparation, handoff, matching,
and checkpoint mechanics.
- `future.md` should remove the completed scene-aware combat item.
## Quality Expectations
Tests should protect behavior and architectural boundaries rather than internal
helper shape. Coverage should demonstrate:
- exact non-combat matches produce accepted empty combat artifacts without an
LLM call;
- exact combat matches retain the existing LLM path;
- an unbound required reference is rejected during pipeline resolution;
- missing and mismatched chunk coverage produces an accepted empty combat
artifact without an LLM call and emits a bounded warning;
- malformed external and generated artifacts fail at the appropriate
preparation or handoff boundary;
- retries are neither consumed nor attempted for deterministic skips;
- scene artifact changes invalidate relevant checkpoint reuse;
- no scene prose is exposed through combat prompts, warnings, metadata, or
manifests; and
- the complete maintained configuration resolves and materializes the intended
ordered dependency.
Model-output fixtures should remain deterministic test doubles. This feature
does not require live-provider tests or assertions over probabilistic model
quality.
## Non-Goals
This scope does not:
- change the durable scene-description or combat-turn schemas;
- add scene fields to combat-turn artifacts;
- annotate generic chunks with D&D classifications;
- make the combat lane depend implicitly on the scene-description lane;
- skip spell, NPC, NPC-interaction, or scene-description extraction;
- infer combat from scene titles, summaries, overlap, or campaign references;
- introduce arbitrary DAG scheduling or concurrent cross-step execution;
- add prior-run artifact discovery or new reference syntax; or
- add an LLM-backed validator, repair pass, or generic deduplication stage.
## Completion Criteria
The scope is complete when an explicitly configured ordered D&D pipeline can
hand an approved scene-description artifact to combat extraction, exact
`combat` scene matches are the only chunks that invoke the combat LLM, every
other chunk produces a deterministic empty result, provenance and checkpoint
identity remain correct, the complete example demonstrates the workflow, and
the canonical current-behavior documentation reflects the implemented
contract.

View File

@@ -93,13 +93,15 @@ type TypedNormalizeResult[T any] struct {
Retry *NormalizeRetry
}
// NormalizeRetry asks the framework to retry normalization while retaining a
// safe candidate for acceptance if the retry budget is exhausted.
// Normalize retry diagnostic limits bound module-provided values before the
// framework persists them in debug artifacts.
const (
MaxNormalizeRetryReasonCodeBytes = 128
MaxNormalizeRetryMessageBytes = 4096
)
// NormalizeRetry asks the framework to retry normalization while retaining a
// safe candidate for acceptance if the retry budget is exhausted.
type NormalizeRetry struct {
ReasonCode string
Message string