241 lines
12 KiB
Markdown
241 lines
12 KiB
Markdown
# D&D Validation Reliability Hardening
|
|
|
|
## Purpose
|
|
|
|
Improve the correctness and efficiency of D&D validation retries by giving the
|
|
producer complete, contextual, semantically useful correction guidance; moving
|
|
simple closed-value checks into compatible LLM-facing schemas; and removing
|
|
duplicated source-range containment logic. The work should make one retry more
|
|
likely to repair every detectable defect without weakening deterministic
|
|
validation or exposing internal identifiers to a model.
|
|
|
|
## Current State And Findings
|
|
|
|
The validation architecture is coherent and the major safety boundaries are in
|
|
place. Validators run through the framework-managed retry protocol, model-facing
|
|
guidance is distinct from reason codes and operator diagnostics, schemas use
|
|
provider-compatible required fields, and D&D extraction adapters canonically
|
|
order safely resolvable source-reference endpoints before validation. The
|
|
review nevertheless found four opportunities to make that architecture more
|
|
consistent and effective.
|
|
|
|
### Generic correction guidance
|
|
|
|
Most rejecting D&D validators produce detailed operator-facing messages but
|
|
pair them with a fixed, general correction sentence. The framework correctly
|
|
constructs retry text from `CorrectionGuidance` alone, so the model does not see
|
|
the internal diagnostic—and therefore often does not learn which records were
|
|
wrong or how each record must change. Item-occurrence shape validation and the
|
|
combat-semantics validator already demonstrate the intended contextual pattern.
|
|
|
|
Passing the operator message through unchanged is not an acceptable fix.
|
|
Operator messages may contain Go paths, array indexes that do not reliably map
|
|
back to canonicalized model output, reason codes, validator terminology, or
|
|
opaque application identities. Those details are useful for diagnosis but are
|
|
not appropriate model context.
|
|
|
|
### Missing closed-value schema constraints
|
|
|
|
Four private LLM response schemas describe categorical fields as unconstrained
|
|
strings even though their prompts, durable artifact schemas, and deterministic
|
|
validators define closed value sets:
|
|
|
|
- `combat-turns`: `turn_kind`;
|
|
- `enemy-events`: `kind`;
|
|
- `npc-occurrences`: `kind`; and
|
|
- `scene-descriptions`: `kind`.
|
|
|
|
This defers an inexpensive structural check until the outer semantic-validation
|
|
loop. Item and location occurrences already use the preferable private-schema
|
|
enum pattern.
|
|
|
|
### First-defect validation
|
|
|
|
Spell shape validation returns after the first invalid cast field, and enemy
|
|
engagement validation returns after the first duplicated engagement subject.
|
|
When multiple defects exist, each retry can therefore reveal only one of them.
|
|
That needlessly consumes the bounded producer-attempt budget and makes
|
|
correction less reliable for smaller models.
|
|
|
|
### Duplicated source-range containment
|
|
|
|
Nine extraction source-reference validators independently determine whether a
|
|
citation belongs to the current chunk. Five validate every source unit spanned
|
|
by a reference, while four check only the endpoints. Those behaviors are
|
|
equivalent for today's contiguous materialized chunks, but the duplicated
|
|
implementations and different apparent contracts create a drift risk.
|
|
Scene-description validation has a separate exact-one-scene range contract and
|
|
should remain specialized.
|
|
|
|
## Target End State
|
|
|
|
- Every producer-correctable D&D rejection supplies bounded contextual guidance
|
|
that identifies all affected records using transcript-grounded names, source
|
|
ranges, and response fields as appropriate, explains each semantic defect,
|
|
and requests one complete corrected replacement.
|
|
- Model-facing guidance never contains reason codes, validator keys, Go field
|
|
paths, opaque or hash-derived entity IDs, or raw internal error text.
|
|
Operator diagnostics remain detailed and separately available.
|
|
- Every detectable issue in a candidate is collected in one validator pass,
|
|
subject to the repository's established diagnostic bounds. A retry is not
|
|
spent merely to reveal the next defect.
|
|
- Closed categorical fields are constrained by compatible enums in the private
|
|
LLM schemas. Deterministic typed validators retain the same rules as
|
|
defense-in-depth and as protection for non-LLM producers and later stages.
|
|
- One D&D-owned helper implements full-span chunk containment for extraction
|
|
evidence. All applicable source-reference validators use it and no
|
|
module-local endpoint-only or coverage implementation remains.
|
|
- Validator results remain immutable, deterministic, bounded, and suitable for
|
|
the existing feedback-aware replacement-request protocol. Retry budgets,
|
|
warning policy, durable artifact schemas, and accepted output semantics do
|
|
not change.
|
|
|
|
## Required Work
|
|
|
|
### Contextual correction guidance
|
|
|
|
Audit each rejecting D&D validator according to the stage that can actually
|
|
correct its result:
|
|
|
|
- Extraction validators should build model-facing issue descriptions from the
|
|
candidate and its transcript context. Shape, registry-membership, identity,
|
|
catalog, engagement, and source-reference validators should name the
|
|
contextual artifact and cited range where useful, state the invalid value or
|
|
relationship in plain language, and state the required replacement shape.
|
|
- LLM-backed NPC, item, and location registry reconciliation should continue to
|
|
translate its typed proposal issues through the shared category-to-prose
|
|
renderer, including any domain-specific supplement.
|
|
- Normalization-only deterministic invariant validators should retain useful
|
|
operator diagnostics. They should not claim that a producer can directly
|
|
repair deterministically derived ordering, identity, or normalization state.
|
|
If such a rejection can reach a feedback-capable producer, guidance must be
|
|
expressed only in terms of the source candidate that producer controls.
|
|
- LLM-backed semantic validators may continue using their model explanation
|
|
when it is bounded and semantically meaningful.
|
|
|
|
Introduce a small D&D-shared diagnostic utility only for demonstrated common
|
|
mechanics such as stable grouping, de-duplication, contextual source-range
|
|
rendering, and bounded correction aggregation. Domain validators must continue
|
|
to own the meaning and prose of their rules. The generic pipeline must not gain
|
|
D&D knowledge, and operator `Message` values must never be mechanically copied
|
|
into `CorrectionGuidance`.
|
|
|
|
Update validator policy fingerprints wherever correction behavior changes so
|
|
checkpoints created under generic feedback are not reused as though the policy
|
|
were identical.
|
|
|
|
### Private-schema enum guardrails
|
|
|
|
Add the existing supported value sets to the four private LLM schemas:
|
|
|
|
- `assets/dnd/combat-turns/schemas/dnd_combat_turns_llm.v1.json`;
|
|
- `assets/dnd/enemy-events/schemas/dnd_enemy_events_llm.v1.json`;
|
|
- `assets/dnd/npc-occurrences/schemas/dnd_npc_occurrences_llm.v1.json`; and
|
|
- `assets/dnd/scene-descriptions/schemas/dnd_scene_descriptions_llm.v1.json`.
|
|
|
|
Use the exact values owned by the corresponding durable contract and typed
|
|
domain constants. Keep all object properties required, retain current nullable
|
|
types where present, and do not introduce `uniqueItems`, optional properties,
|
|
or conditional schema logic. Revise schema tests that currently accept unknown
|
|
values, and rely on computed asset fingerprints to invalidate incompatible
|
|
LLM-output checkpoints. Keep deterministic enum validation in place.
|
|
|
|
### Complete per-attempt issue collection
|
|
|
|
Refactor spell shape and enemy engagement validation to inspect the complete
|
|
candidate and collect every detectable violation before returning. Build the
|
|
operator diagnostic and contextual correction request from the same evaluated
|
|
issue set while preserving their different audiences. De-duplicate repeated
|
|
semantic instructions, retain enough contextual identification for every
|
|
affected record, and use the established bounded diagnostic behavior rather
|
|
than an unbounded error string.
|
|
|
|
Review the surrounding D&D validators while applying the contextual-guidance
|
|
change. Remove any additional accidental early exits that prevent independent
|
|
candidate defects from being reported together, but retain immediate returns
|
|
for request-level prerequisites whose absence makes further inspection unsafe
|
|
or meaningless.
|
|
|
|
### Shared full-span chunk containment
|
|
|
|
Add one helper under `internal/modules/dnd/shared` that determines whether a
|
|
source reference's complete document-ordered span is contained by a chunk. The
|
|
helper must:
|
|
|
|
- validate source identity;
|
|
- resolve endpoints through `source.DocumentIndex` rather than numeric-ID
|
|
assumptions;
|
|
- require every unit in the inclusive span to be present in the chunk;
|
|
- handle nil or unresolved inputs without panic;
|
|
- avoid mutating the source, chunk, or reference; and
|
|
- leave source-reference validity and error wording to the consuming validator.
|
|
|
|
Use the helper from the spell, NPC-registry, NPC-occurrence, item-registry,
|
|
item-occurrence, location-registry, location-occurrence, combat-turn, and
|
|
enemy-event source-reference validators. Keep the scene-description exact-range
|
|
validator separate because it enforces a materially different contract.
|
|
|
|
## Testing And Documentation
|
|
|
|
Add lean offline behavioral coverage at the narrowest stable boundary:
|
|
|
|
- correction tests should prove that multiple contextual defects produce one
|
|
actionable, bounded request and that internal identifiers and diagnostic
|
|
syntax are absent; they should not snapshot exact prose or message length;
|
|
- schema tests should prove rejection of representative unsupported categorical
|
|
values and acceptance of the supported sets without duplicating every
|
|
provider behavior;
|
|
- spell-shape and engagement tests should prove that independent defects are
|
|
reported together;
|
|
- the shared containment helper should own the full case matrix, including
|
|
partial spans, non-monotonic unit IDs, wrong sources, unresolved endpoints,
|
|
and nil inputs; consuming validators need only enough coverage to prove they
|
|
use the common policy; and
|
|
- existing validator tests should be simplified when the shared helper makes
|
|
module-local cases redundant.
|
|
|
|
Update `docs/internal/dnd.md` to document the durable conventions for
|
|
contextual correction guidance, complete issue collection, private-schema enum
|
|
guardrails, and shared full-span extraction containment. Update
|
|
`docs/internal/pipeline.md` only if framework behavior changes; the intended
|
|
work applies its existing contract and should normally require only a link or
|
|
no change. No new ADR is required because ADR-0014 and the architecture policy
|
|
already decide the separation between semantic guidance and internal
|
|
diagnostics. Create an ADR only if implementation requires changing that
|
|
framework-level decision.
|
|
|
|
## Non-Goals
|
|
|
|
- Passing operator diagnostics or internal errors directly to an LLM.
|
|
- Exposing durable IDs, hashes, reason codes, validator keys, or raw provider
|
|
responses in correction text.
|
|
- Changing stage retry counts, PromptKit structural-repair budgets, validator
|
|
failure policy, or warning classification.
|
|
- Silently repairing domain-semantic defects or weakening deterministic
|
|
validators after adding schema guardrails.
|
|
- Adding provider-sensitive schema constructs beyond simple enums.
|
|
- Moving D&D-specific behavior into the generic pipeline framework.
|
|
- Generalizing the scene-description exact-range contract into the shared
|
|
extraction containment helper.
|
|
|
|
## Acceptance Criteria
|
|
|
|
- Every producer-correctable D&D rejection reviewed in this work gives the next
|
|
attempt all currently detectable, actionable corrections in contextual prose.
|
|
- No model-facing correction request contains an opaque entity ID, hash,
|
|
validator key, reason code, Go-style field path, or unfiltered operator error.
|
|
- Spell shape and enemy engagement validation aggregate independent defects in
|
|
one pass and keep outputs immutable.
|
|
- The four private LLM schemas reject unsupported categorical values, retain
|
|
provider-compatible required-only shapes, and leave durable v1 contracts
|
|
unchanged.
|
|
- All nine applicable extraction source-reference validators use one D&D-shared
|
|
full-span containment policy; scene descriptions retain their specialized
|
|
exact-range check.
|
|
- Validator and asset fingerprints change wherever their effective policy
|
|
changes, preventing reuse of stale checkpoints.
|
|
- Canonical internal documentation records the conventions future D&D
|
|
validators must follow.
|
|
- Focused tests, `go test ./...`, `go vet ./...`, and
|
|
`go build ./cmd/notarius` pass in a supported development environment.
|