Files
notarius/docs/roadmap/implementation.md

413 lines
19 KiB
Markdown

# Scene-Aware Combat Extraction Implementation Plan
## Status
Ready for implementation. Execute each stage in order, with one implementation
prompt per stage.
## Feature Contract
The target behavior and policy choices are defined in
[Scene-Aware Combat Extraction](scene-aware-combat-extraction.md). That document
is authoritative when this plan and the feature contract appear to overlap.
The central rule is strict opt-in: `dnd/combat-turns` may call the LLM only for
an accepted chunk with one exact scene-description record whose `kind` is
`combat`. Exact non-combat matches and missing or mismatched coverage produce an
accepted empty combat-turn list without an LLM call.
## Required Reading
Before Stage 1, read:
- `docs/development.md`;
- every document under `docs/policy/`;
- `docs/roadmap/scene-aware-combat-extraction.md`;
- `docs/internal/pipeline.md`;
- `docs/internal/modules.md`;
- `docs/integrations/dnd-scene-description-artifacts.md`;
- `docs/integrations/dnd-combat-turn-artifacts.md`; and
- the focused code and tests named in the applicable stage.
Before changing a subsystem in a later stage, reread its focused current
documentation and tests if they have changed since the earlier stage.
## Global Implementation Constraints
- Keep scene interpretation under `internal/modules/dnd`; generic framework and
source packages must not import D&D types or interpret scene kinds.
- Use the existing ordered generated-reference mechanism. Do not add implicit
lane discovery, a new dependency syntax, a general DAG, or direct output-file
reads.
- Add `scene_descriptions` only to the combat extractor. The combat normalizer
continues to accept only the optional `npcs` registry.
- Keep both durable v1 schemas unchanged.
- Do not add the scene artifact or a derived classification to the combat
prompt. Do not change combat prompt assets, prompt ordering, prompt inputs,
response schema, or prompt fingerprint.
- Preserve existing combat extraction behavior for eligible combat chunks,
including NPC grounding, retries, validators, candidate mapping, and
normalization.
- Treat the scene artifact as control context, never as combat evidence. Never
copy its references, title, or summary into combat turns.
- Use deterministic, offline tests and existing LLM fakes. Do not call a live
provider or turn probabilistic model output into a correctness gate.
- Test stable behavior at the narrowest owning boundary. Do not add tests for
private helper shape, cache implementation details, or exact diagnostic prose.
- Do not add a framework abstraction unless implementation proves the existing
reference contracts cannot express this feature. No such framework gap is
currently expected.
## Stage 1 — Add the Domain Scene-Eligibility Reference Boundary
### Goal
Create one domain-owned, immutable boundary that decodes an approved
scene-description reference, retains only gating data, and supports both
construction-time external references and operation-time generated references.
This stage must not change combat extraction behavior.
### Implementation
1. Add `internal/modules/dnd/scenedescriptions/registry`, following the useful
external/generated resolution pattern in `internal/modules/dnd/npcs/registry`
without copying NPC-specific projection or identity policy.
2. Define these package contracts:
- reference slot: `scene_descriptions`;
- maximum item size: 1 MiB (`1048576` bytes);
- accepted representation: exactly one `application/json` item decoded by
`codec/scenedescriptions`;
- eligibility projection: scene ID, exact `source_ref`, and `kind` only.
3. Decode through the approved scene-description codec. Reject:
- more than one item;
- invalid or non-JSON media types;
- oversized content;
- invalid approved scene JSON; and
- duplicate scene IDs, even if duplicate records are otherwise identical.
4. Do not retain titles, summaries, original bytes, paths, or prompt material in
the resolved eligibility view. Framework reference provenance remains the
owner of raw reference identity.
5. Build an immutable lookup keyed by exact scene ID. Expose a chunk-match result
with three stable states:
- exact: ID, source ID, start unit ID, and end unit ID all match;
- missing: no record has the current chunk ID;
- mismatched: the ID exists but any source-range field differs.
Only the exact state exposes the scene kind.
6. Compute a semantic eligibility digest from a canonical projection sorted by
scene ID. The digest must:
- ignore title, summary, and input array order; and
- change when an ID, source-range field, or kind changes.
Represent an unbound view with the canonical empty projection
`{"scenes":[]}` and its `sha256:` digest so generated-reference construction
has a stable sentinel without pretending the required reference was bound.
7. Provide a resolver that:
- seeds and validates an external materialized item during construction;
- permits an unbound seed because a generated required reference is not
populated until the operation-time handoff;
- resolves an operation-time generated item when present;
- otherwise returns the seeded view; and
- safely reuses resolved immutable views across concurrent extract jobs.
8. Expose only the accessors required by combat extraction: bound state, record
count, eligibility digest, and exact chunk matching. Return defensive values
rather than mutable maps or scene records.
### Tests
Add focused package tests that own:
- approved decoding and exact match classification;
- missing and mismatched match states;
- strict item count, media type, size, codec, and duplicate-ID failures;
- semantic digest stability across prose and ordering changes;
- semantic digest changes for each gating field;
- construction-time seed versus operation-time generated override; and
- immutability and concurrent-safe reuse at the public package boundary where
practical.
Do not duplicate the scene codec's exhaustive JSON-shape tests.
### Validation
Run:
```sh
go test ./internal/modules/dnd/scenedescriptions/registry
go test ./internal/modules/dnd/codec/scenedescriptions
git diff --check
```
Stage 1 is complete when the new package is independently usable and tested but
no production module selects it yet.
## Stage 2 — Enforce the Required Combat Scene Gate
### Goal
Make `dnd/combat-turns` require the scene-description artifact and invoke its
existing LLM path only for exact combat scenes. Update directly affected
configuration fixtures and contract tests in the same stage so the repository
does not retain knowingly invalid combat configurations.
### Module Contract
1. In `internal/modules/dnd/extract/combatturns`, add aliases for the
scene-registry slot and size limit and declare a reference slot with:
- name `scene_descriptions`;
- `Required: true`;
- media type `application/json`;
- accepted artifact kind `dnd/scene-description-list`;
- maximum size 1 MiB; and
- a description stating that it gates combat eligibility and is not combat
evidence.
2. Preserve sorted reference-slot ordering and the existing agreement between
`ReferenceSlots()` and `ModuleSpec().ReferenceSlots`.
3. Add a scene resolver to the immutable extractor. Constructor behavior must:
- validate a materialized external scene artifact before execution;
- allow the construction reference set to lack an item for a configured
generated binding; and
- retain no scene prose.
4. Define a stable gating policy identity
`dnd.combat_turns.scene_gate.v1`.
5. Extend manifest metadata with `scene_gate_policy`. For a seeded external
reference only, also report `scene_eligibility_digest` and
`scene_description_count`. Do not publish those operation-time values for a
generated reference in singleton component metadata; framework provenance
owns that handoff.
6. Extend prepared-component checkpoint fingerprints with:
- `scene_gate_policy`; and
- `scene_eligibility`, using the seeded external eligibility digest or the
registry's stable unbound sentinel for a generated binding.
Keep the existing prompt, response-schema, mapping-policy, and NPC-registry
fingerprints unchanged.
### Extraction Behavior
In `Extractor.Extract`, retain common request validation through
`shared.PrepareChunkExtraction`, then resolve and apply the scene gate before
building prompt inputs or resolving NPC prompt grounding:
1. If the operation-time scene reference is absent, return a contextual module
error stating that the required reference is missing. Normal configured runs
should have failed earlier during resolution or required-producer handoff;
this check protects direct module use.
2. For an exact `combat` match, continue through the existing NPC resolution,
prompt construction, LLM request, response mapping, and validation path
without semantic changes.
3. For an exact `narrative`, `recap`, or `meta` match, return
`dnd.CombatTurnList{CombatTurns: []dnd.CombatTurn{}}`. Make no LLM call and
emit no warning.
4. For a missing or mismatched match, return the same non-nil empty list without
an LLM call and attach exactly one content-safe warning:
- scope: `scene_descriptions`;
- reason code: `scene_classification_unavailable`;
- message: explain that no exact classification was available and combat
extraction was skipped, without scene prose or transcript content.
5. The deterministic result is an accepted extract result, not an error or
rejection. It must pass the existing combat shape, source-reference, schema,
merge, and normalization boundaries and must not consume retry attempts.
6. Do not pass `scene_descriptions` into `shared.PromptInputs` or add it to the
structured completion request. The current shared helper already selects
only campaign and transcript inputs; preserve that boundary explicitly.
### Required Test And Fixture Migration
1. Update combat extractor test helpers so tests of the existing LLM behavior
supply an exact combat scene artifact for their current chunk. Keep
NPC-specific test intent unchanged.
2. Add a compact table-driven extractor test covering:
- exact combat;
- exact narrative, recap, and meta;
- missing scene ID;
- matching ID with each mismatched range identity; and
- an operation-time missing required reference.
Assert typed output, warning reason where applicable, and whether the fake LLM
was called. Do not assert private helper calls.
3. Update module specification, registration, metadata, and fingerprint tests
for the required slot and new policy identities. Assert that changing only
scene prose does not change the eligibility fingerprint, while changing
kind or range does.
4. Update every in-repository pipeline or programmatic test configuration that
selects `dnd/combat-turns`:
- the maintained complete example must bind the earlier
`scene-descriptions` lane to the later `scene_descriptions` step reference;
- CLI combat contract configurations must bind a compatible external scene
artifact when testing a standalone one-step profile;
- the ordered D&D integration fixture must produce scene descriptions in its
first step and bind them in its consumer step; and
- fake clients used by those integration tests must return an exact combat
scene where the test expects the combat LLM path.
5. Add or update a configuration-resolution assertion proving that a selected
combat extractor without `scene_descriptions` is rejected as an unbound
required slot.
6. Update existing expected lane counts, schemas, request sets, fingerprints,
and provenance counts only where the new scene producer or consumer
materially changes the contract. Preserve the original NPC-grounding
assertions.
### Validation
Run:
```sh
go run ./cmd/notarius config validate \
--config examples/dnd-complete.config.yml \
--pipeline dnd-session
go test ./internal/modules/dnd/...
go test ./internal/modules/integration
go test ./internal/cli
git diff --check
```
Stage 2 is complete when required-slot resolution is enforced, every combat
extract path obeys the eligibility rule, all affected configurations are valid,
and focused suites pass.
## Stage 3 — Prove Ordered Handoff, Provenance, And Retry Behavior
### Goal
Protect the cross-component behavior that cannot be established by the
scene-registry and extractor package tests alone. Keep this stage lean: reuse
the real codecs, resolver, pipeline, and validators, with a deterministic fake
only at the LLM boundary.
### Integration Coverage
1. Extend or refactor the focused ordered D&D integration test so it proves, in
one representative assembled workflow:
- scene descriptions and NPCs normalize in the first step;
- their accepted artifacts are supplied to the intended second-step targets;
- an exact combat scene causes exactly one combat completion request;
- the combat request contains transcript, campaign, and NPC inputs as
configured but no `scene_descriptions` input;
- the combat output and existing NPC canonicalization remain correct;
- the manifest records one generated scene-description reference for the
combat extract target with the expected kind, schema, media type, digest,
and producer identity; and
- manifest and component metadata contain no scene title, summary, transcript
text, or generated artifact payload.
2. Add one focused non-combat ordered-run case using an exact `narrative`
classification. Configure combat retries to a value greater than zero, then
prove:
- no combat completion request occurs;
- one accepted empty combat lane reaches normalized output;
- no retry or rejection is recorded for the deterministic skip; and
- no unavailable-classification warning is emitted for an exact non-combat
match.
3. Add one focused missing-or-mismatched case at the narrowest stable boundary
not already covered by Stage 2. Prove the accepted empty result and
`scene_classification_unavailable` warning reach the runner's durable warning
collection without a combat LLM call. Do not repeat every mismatch variant
at integration scope.
### Checkpoint And Identity Coverage
1. Confirm the existing framework-generated-reference dependency fingerprint
already includes the canonical generated scene artifact. Reuse that
mechanism; do not add a second framework fingerprint path.
2. Add focused D&D coverage proving:
- a seeded external reference's component fingerprint changes when eligibility
kind or exact range changes;
- title-only or summary-only changes leave that semantic component
fingerprint unchanged; and
- the generated reference appears in the combat extract checkpoint
dependencies, so changing the generated artifact prevents reuse.
3. Prefer an existing checkpoint identity or prepared-pipeline test boundary.
Do not add a full CLI resume test if the same dependency invalidation is
already credibly protected by the framework checkpoint tests plus the D&D
prepared-pipeline assertion.
### Validation
Run:
```sh
go test ./internal/modules/integration
go test ./internal/framework/pipeline
go test ./internal/cli
go test ./internal/modules/dnd/extract/combatturns
go test ./internal/modules/dnd/scenedescriptions/registry
git diff --check
```
Stage 3 is complete when the ordered generated handoff, no-call behavior,
warning propagation, provenance privacy, and checkpoint invalidation are
protected at their stable owning boundaries without redundant end-to-end cases.
## Stage 4 — Update Canonical Documentation And Finish Validation
### Goal
Move the implemented behavior out of future-only documentation and into each
canonical current-behavior owner, then perform repository-wide verification.
### Documentation
1. Update `docs/config.md` to own:
- the combat extractor's required `scene_descriptions` slot;
- accepted artifact kind, JSON media type, single-item and 1 MiB limits;
- configuration-time rejection when the slot is unbound; and
- the explicit generated binding used by the complete example.
Update illustrative combat configuration fragments so none show an invalid
unbound combat extractor.
2. Update `docs/operations.md` to explain that the first step produces NPC and
scene-description artifacts and the second step runs combat extraction only
for exact combat scenes.
3. Update `docs/integrations/dnd-combat-turn-artifacts.md` to own the externally
observable eligibility, empty-result, warning, retry, prompt-exclusion, and
provenance behavior.
4. Add a concise downstream-consumer link or summary to
`docs/integrations/dnd-scene-description-artifacts.md` without duplicating
the combat contract.
5. Update `docs/internal/modules.md` with the scene eligibility registry,
construction-time versus operation-time resolution, exact matching,
fingerprint, and content-retention boundaries. Keep generic ordered-handoff
mechanics canonical in `docs/internal/pipeline.md`; update that document only
if implementation changed or clarified a generic mechanism.
6. Update `docs/internal/overview.md` only if its component inventory requires
the new domain package.
7. Remove the completed combat-gating bullet from
`docs/roadmap/future.md`. Preserve the remaining scene-chunking evaluation
and ordered-pipeline guidance.
8. Review `docs/roadmap/scene-aware-combat-extraction.md` for consistency with
the final implementation. Keep it as feature history until a later explicit
retirement request; do not duplicate its implementation sequence there.
9. Verify that README and CLI documentation need no changes. Update them only
if an existing command or linked example became inaccurate.
### Final Verification
Run all repository checks:
```sh
go run ./cmd/notarius config validate \
--config examples/dnd-minimal.config.yml \
--pipeline dnd-session
go run ./cmd/notarius config validate \
--config examples/dnd-complete.config.yml \
--pipeline dnd-session
go test ./...
go vet ./...
go build ./cmd/notarius
git diff --check
```
Also verify:
- only the two maintained D&D configuration examples remain;
- every current documentation link points to an existing file;
- no current-behavior document describes combat extraction without the required
scene gate;
- no prompt asset or combat response schema changed; and
- no generated build artifact is left in the working tree.
Stage 4 is complete when documentation ownership is correct, future work no
longer lists the implemented item, both examples validate, and all repository
checks pass.
## Open Questions
None. The feature contract and this plan resolve the required behavior,
reference ownership, matching semantics, failure policy, warning policy,
checkpoint identity, documentation ownership, and test boundaries.