Document future validation and retry work

This commit is contained in:
2026-08-25 14:29:00 +00:00
parent b4a81f8b09
commit 5e5c69bf9d

View File

@@ -5,13 +5,230 @@ configuration, operations, internal, and integration docs. This roadmap records
future work only. Items are ordered roughly by current value and specificity, future work only. Items are ordered roughly by current value and specificity,
not as committed release dates. not as committed release dates.
## Near-Term Validation And LLM Reliability
The following work forms one related program but should be promoted into
separate feature roadmaps and implemented in dependency order. PromptKit owns
structural output repair within one completion. Notarius owns stage candidates,
validator chains, semantic rejection policy, and whether another stage attempt
is warranted.
### 1. Upgrade To PromptKit v0.8.0
- Upgrade the PromptKit dependency and follow the upstream v0.8.0 release
guide, including its now-active `repair_attempts` behavior and compatibility
checks for prompt definitions, prepared execution, validation results,
capacity accounting, error adaptation, debug data, and cumulative token
usage.
- Use PromptKit's bounded repair loop for eligible structural output failures.
A repair must resend the immutable original prompt followed by the latest
defective assistant response and one bounded user correction message. It
must not accumulate the history of every defective candidate.
- Review every maintained LLM prompt's structural repair budget. Prefer a
small positive budget where a corrected structured response can be useful,
while retaining zero where a second call would be inappropriate or where
observational evidence does not justify it.
- Keep PromptKit repair attempts distinct from Notarius stage retries.
PromptKit repairs a response that fails its deterministic `basic`, `json`, or
`json_schema` contract within one stage attempt; Notarius may later retry a
complete stage candidate after application validation. Document and test the
multiplicative maximum provider-call cost created by both configured
budgets, including PromptKit repair calls made by any LLM-backed validators.
Attribute generation count, token usage, latency, and failure provenance to
the producer, validator, PromptKit repair, and Notarius correction layers so
operators can explain the actual cost of a run.
- Treat exhaustion of PromptKit structural repair as the absence of a usable
structured candidate. The default terminal policy is to fail the run after
the applicable Notarius stage-attempt budget is also exhausted. A future
configurable alternative may reject the affected stage or lane where the
pipeline can represent that outcome, but it must never accept undecodable or
structurally invalid output merely with a warning.
### 2. Feedback-Aware Stage Validation Retries
- Model Notarius's corrective stage-retry conversation explicitly after
PromptKit v0.8.0. The first attempt sends the ordinary complete initial
prompt. If application validation rejects the resulting LLM-produced
candidate and another stage attempt is available, reconstruct that complete
initial prompt byte-for-byte and append exactly two messages: an assistant
message containing the defective response and an application-owned user
message detailing every applicable semantic validation error and requesting
one corrected, complete replacement response. This is a freshly constructed
correction request, not continuation of an accumulating conversation.
- Use the configured stage `retries` value as the one outer retry budget for
this loop. `retries: N` continues to mean at most `N` additional complete
chunk, extract, merge, or normalize attempts after the initial attempt,
whether an attempt is needed because of a producer error or semantic
rejection. Do not add a second semantic-correction count. PromptKit's
prompt-level `repair_attempts` budget is independent and internal to each
individual LLM completion, and does not consume or replenish the Notarius
stage budget.
- Extend the framework-managed validation boundary for chunk, extract, merge,
and normalize stages so a rejected LLM-produced candidate and its exact raw
model response remain available to construct the next stage attempt.
Deterministic producers cannot improve by repeating the same inputs; a
rejection from a deterministic stage is therefore terminal under the
configured rejection policy rather than consuming retries mechanically.
- Preserve the original session ID, selected profile, structured-output
contract, prompt inputs, and reusable prompt prefix. Carry only the latest
candidate and latest aggregate feedback; do not build an unbounded retry
conversation. Keep model-facing corrective guidance separate from
operator-facing diagnostics, and apply explicit size, redaction, and debug
disclosure rules to both.
- Run every applicable validator in the configured chain before deciding
whether to retry. Do not short-circuit merely because an earlier validator
rejected the candidate. Aggregate all semantic rejection reason codes and
corrective guidance into the retry message so one retry can address the
whole candidate. A validator is applicable only when its declared target and
prerequisites can be satisfied; record a deterministic skipped diagnostic
rather than invoking a validator on an input it cannot interpret. Initially
execute the chain sequentially in configured order so results, diagnostics,
costs, and feedback ordering remain deterministic; consider validator
concurrency only in response to measured latency.
- Continue running independent applicable validators after one validator
execution failure so the attempt retains as much useful diagnostic
information as practical. Do not present validator operational failures as
defects in the producer candidate and do not include them in corrective
feedback.
- Distinguish three terminal conditions and make their policies configurable
at a coherent pipeline or binding scope:
- **producer structural failure:** PromptKit could not return a usable
structured candidate after its repair budget. Default to `fail_run`; an
allowed alternative may record a terminal stage or lane rejection where
execution can safely continue, but may not accept the invalid output;
- **semantic rejection:** one or more validators completed and rejected the
candidate. Default to `fail_run` after corrective stage retries are
exhausted; allow an explicit alternative that records the existing
rejected-output outcome without advancing that output;
- **validator execution failure:** a validator could not produce a valid
decision because of generation, structural-output, transport, or internal
failure. Default to a genuine warning and an explicitly recorded
`validation_incomplete` or equivalent degraded state while allowing the
candidate to continue; allow strict configuration to fail the run instead.
- An LLM-backed validator uses the same scheduled PromptKit boundary as every
other LLM-backed module. Its own response may use PromptKit's bounded
structural repair. Distinguish its possible output states:
- output rejected by PromptKit's structural contract should consume only the
validator prompt's configured PromptKit repair budget;
- output that is structurally valid but violates a deterministically
checkable validator-result invariant should be classified as a validator
execution failure;
- output that satisfies the complete validator-result contract is the
validator's decision, even though an LLM judgment may remain imperfect.
Automatically judging that judgment would require another semantic
validator and is outside this feature.
If the validator cannot return a contract-valid decision, do not recursively
create another Notarius semantic-validation loop around it. Apply the
configured validator-failure policy. The default warning must identify the
validator and affected stage without exposing sensitive content.
- Separate validator execution retry from producer correction. A transient
validator operational failure must not automatically discard and regenerate
an otherwise usable producer candidate. Any bounded retry of the validator
itself should reuse that same immutable candidate and remain subordinate to
PromptKit and provider retry behavior.
- Preserve attempt-level provenance, cumulative token usage, validator
outcomes, aggregated correction feedback, and terminal policy decisions in
the debug and manifest models without copying raw source material into
ordinary errors or durable summaries.
- Define terminal-outcome precedence. A semantic rejection dominates a
validator execution failure for the same candidate: use the completed
rejections to correct the producer while separately recording incomplete
validation. If a later candidate has no semantic rejection but one validator
still fails, apply the configured validator-failure policy to that candidate.
Never allow a known semantic rejection to become accepted through a
warn-and-continue setting, and never accept a structurally invalid producer
response. Permissive policy may preserve a rejected-output outcome or accept
a structurally valid candidate with explicitly incomplete validation; it may
not relabel known-invalid output as approved.
Before implementation, record the generic validation and retry state machine
in an ADR. The ADR should own the separation between PromptKit repair and
Notarius correction, use of the existing stage-retry budget, reconstruction of
correction conversations, all-applicable-validator aggregation, deterministic
validator ordering, non-recursive validator failure handling, outcome
precedence, default fail-open/fail-closed choices, configurable terminal
policies, and provenance and sensitive-data constraints. A dependency-upgrade
ADR is not needed for PromptKit v0.8.0 itself. Current behavior remains
authoritative until the validation ADR is implemented and the canonical
architecture, configuration, operations, and internal documentation are
updated.
### 3. D&D Combat Scene Semantic Validation
- Add an optional production LLM-backed D&D validator that determines whether
proposed scene boundaries and classifications represent substantive active
combat correctly. Its central quality goal is that active combat is kept in
coherent scenes classified as `combat`, rather than split incorrectly or
hidden inside scenes classified as `narrative`, `recap`, or `meta`.
- Resolve the validator's exact target before implementation. The current
`dnd/scenes` chunker owns only complete, gap-free source ranges, while the
per-chunk `dnd/scene-descriptions` extractor owns the `combat`, `narrative`,
`recap`, and `meta` classification. The preferred initial placement is
therefore an extract-stage validator for `dnd/scene-descriptions`, where it
can compare one proposed kind with the corresponding transcript chunk.
- Consider a chunk-stage LLM validator only for a distinct boundary-coherence
question that can be answered from the complete transcript and proposed
range map, such as whether one continuous combat was fragmented across
inappropriate scene boundaries. Do not duplicate the same classification
judgment at both stages. Moving classification into chunk-plan annotations
would change the deliberately minimal, annotation-free chunk contract and
requires an explicit architecture review before it is selected.
- Validate both false negatives and false positives: a non-combat kind must not
omit substantive active combat, and a combat kind must be supported by such
combat. Keep the existing deterministic downstream rule that combat-turn
extraction runs only for an exact `combat` scene classification; semantic
review improves the upstream classification but does not replace that gate.
- Run the semantic validator through PromptKit, use a minimal required-field
structured response schema, and let PromptKit repair structural validator
output within its bounded budget. A contract-invalid final validator response
is a validator execution failure, not a semantic rejection and not a reason
to recursively validate the validator.
- Evaluate the prompt and decision policy against a small human-reviewed set
containing combat setup, active turns, interruptions, multi-phase encounters,
brief rules discussion, aftermath, recalled combat, and false-positive
hostile dialogue. Measure false acceptance, false rejection, retry success,
added calls, latency, and token cost before placing it in the production
default chain.
- An ADR is not required if classification remains owned by
`dnd/scene-descriptions` and the validator follows the generic validation ADR.
Create or supersede an ADR if the work transfers scene classification into
the chunker or otherwise changes stage ownership or the durable chunk-plan
contract.
### 4. Warning Signal And Presentation Reform
- Audit every warning producer and representative successful runs. Ordinary
success producing dozens of warnings is a failed operator experience: the
volume obscures actionable problems and trains operators to ignore the
warning channel.
- Define a small warning taxonomy that distinguishes actionable degradation,
incomplete validation, lossy fallback, and data-quality risk from routine
normalization observations or informational diagnostics. Preserve detailed
traceability in debug or manifest data without promoting every observation
to a top-level CLI warning.
- Consider stable deduplication and aggregation by scope and reason code,
bounded samples plus omitted counts, and a concise CLI summary with a path to
detailed diagnostics. Do not suppress genuine validator execution failures
merely to reduce the count.
- Decide which warnings affect process status, rejection summaries, durable run
receipts, or only debug output. Ensure warning ordering and aggregation are
deterministic across concurrent execution.
- Establish a representative warning-volume acceptance target and human review
workflow before changing individual producers piecemeal. The intended result
is not zero warnings; it is a small set in which every surfaced warning merits
operator attention.
- This work does not require an ADR unless it changes validation acceptance,
failure, or durable contract semantics. CLI presentation and diagnostic
taxonomy otherwise belong in a feature roadmap followed by updates to their
canonical configuration, operations, integration, and internal documents.
## Near-Term D&D Pipeline ## Near-Term D&D Pipeline
### Evaluate Spell Extraction And Normalization ### Evaluate Spell Extraction And Normalization
- Evaluate ordinary extraction retries and the completed normalization path - Evaluate ordinary extraction retries and the completed normalization path
against a human-reviewed transcript set before adding repair-aware retries or against a human-reviewed transcript set before and after adopting the shared
an LLM-backed semantic validator. PromptKit repair and Notarius validation-retry policies above.
- Maintain a small set of human-reviewed transcripts and outputs for prompt, - Maintain a small set of human-reviewed transcripts and outputs for prompt,
validator, and normalizer development. Treat model-quality review as an validator, and normalizer development. Treat model-quality review as an
iterative human evaluation aid, not a deterministic correctness gate. iterative human evaluation aid, not a deterministic correctness gate.