Files
notarius/docs/roadmap/validation-retries.md

610 lines
31 KiB
Markdown

# Feedback-Aware Stage Validation Retries
## Status
Proposed. This is the active feature roadmap for the next Notarius work set.
Its design decisions are settled. Current behavior remains authoritative until
this roadmap is implemented and the corresponding ADR and canonical
documentation are updated.
## Purpose
Make validation an effective corrective boundary around LLM-produced stage
candidates. When deterministic or LLM-backed validators reject a structurally
valid candidate, Notarius should give the producing model the complete,
ordered validation feedback and use the stage's existing retry budget to ask
for a corrected replacement. The feature must distinguish semantic rejection
from producer failure and validator execution failure, preserve the boundary
between PromptKit repair and Notarius stage retries, and remain safe under
concurrency, cancellation, caching, checkpoints, and sensitive input.
This work is domain-neutral. It establishes the framework behavior required by
future LLM-backed validators such as D&D combat-scene review, but it does not
add that validator.
## User Intent
- A stage candidate should be evaluated by every applicable configured
validator before Notarius decides whether to retry or terminate.
- A semantic retry should be materially more useful than repeating the same
request. The producing model should see its latest defective response and
all actionable semantic feedback.
- PromptKit's bounded structural repair and Notarius's stage retry loop are
separate. Each stage attempt receives its own complete PromptKit repair
budget; PromptKit repair never consumes or replenishes the stage budget.
- Deterministic rejection, semantic rejection, producer structural failure,
and validator execution failure are different outcomes and must not be
collapsed into one generic error path.
- The default posture is strict for known-invalid producer output and tolerant
but visible when a validator itself cannot make a decision.
- Corrective prompts must not expose opaque application identifiers, secrets,
or unbounded diagnostic content merely because those values exist in an
internal artifact or operator-facing error.
## Current State
The current runner already provides useful foundations:
- chunk, extract, merge, and normalize producer bindings have one `retries`
value interpreted as additional stage attempts;
- `runWithRetry` retries producer errors and semantic rejections within that
budget;
- PromptKit performs bounded structural repair inside each structured
completion;
- validator targets, execution classes, profile selection, repair policy,
attempts, debug scopes, checkpoint identity, and deterministic public
ordering are already explicit; and
- the structured-completion response retains the model's validated raw bytes
and PromptKit repair metadata.
The current behavior is not yet the desired corrective workflow:
- the runner repeats the ordinary producer request after rejection and does
not pass the previous model response or validator feedback;
- validation stops at the first rejection or execution failure, so later
applicable validators do not contribute findings;
- validator execution failure is immediately a framework error rather than a
configurable incomplete-validation outcome;
- `ValidationResult.Message` currently serves operator diagnostics and does
not define separately bounded model-facing guidance;
- typed stage results do not carry the exact model response needed for the
next correction attempt;
- validator-binding `retries` values participate in resolved configuration but
are not used to retry a failed validator against the same candidate; and
- Notarius still pins PromptKit v0.8, while PromptKit v0.9.0 now provides the
append-only request-message API needed for application-owned correction
attempts.
## Target End State
For chunk, extract, merge, and normalize stages, Notarius owns one explicit
candidate-attempt state machine:
1. The producer creates one candidate using the ordinary request. An
LLM-backed producer may use PromptKit structural repair internally.
2. The framework establishes one immutable validation candidate and runs every
applicable validator sequentially in configured order.
3. The framework aggregates approvals, warnings, semantic rejections,
execution failures, and skipped-validator diagnostics without allowing one
validator to mutate the candidate seen by another.
4. A candidate with one or more semantic rejections is never accepted. If the
LLM-backed producer has another stage attempt available, Notarius rebuilds
the complete original prompt and appends the latest defective assistant
response followed by one application-owned correction message containing
every actionable rejection. It then requests one complete replacement
candidate.
5. A producer error consumes the same stage attempt budget under the existing
retry rules, but semantic correction material is used only when a
structurally valid candidate was actually rejected.
6. A validator execution failure is retried, when configured, against the same
immutable candidate. It never regenerates the producer candidate by itself.
7. When budgets are exhausted, the configured terminal policies decide
whether the run fails, a rejected output is recorded, or a structurally
valid candidate advances with explicitly incomplete validation.
The first attempt remains byte-for-byte the ordinary prompt rendered from the
selected prompt definition. Every correction attempt starts from that same
ordinary prompt rather than from the prior correction conversation. It appends
exactly two messages:
- an `assistant` message containing the producer-supplied exact defective
response for the latest candidate; and
- a `user` message containing deterministic, bounded, application-owned
correction guidance and asking for one complete replacement response.
The session ID, prompt ID and version, selected profile, reasoning settings,
structured-output contract, repair budget, named inputs, variables, references,
and reusable prompt prefix remain unchanged across stage attempts.
## Architectural Ownership
### PromptKit
PromptKit continues to own prompt loading and rendering, profile resolution,
backend admission, provider generation, structural validation, and bounded
structural repair within one completion. A PromptKit repair conversation is
private to that completion and is not exposed as a Notarius stage attempt.
PromptKit v0.9.0 owns the mechanical operation of appending explicitly supplied
messages to a normally rendered prompt before creating the immutable prepared
execution. `RunRequest.AppendedMessages` preserves the original rendered
messages as an exact prefix, validates and defensively copies additions,
includes the complete sequence in prepared details and the rendered-prompt
hash, and runs it through the ordinary generation and structural-repair path.
PromptKit does not impose message-count, byte-size, token, or context-window
limits and permits empty content, so Notarius retains its stricter
application-owned correction validation and bounds.
### Notarius Framework
The framework owns stage budgets, immutable candidate preparation, complete
validator-chain execution, result aggregation, outcome precedence, correction
message construction, terminal policy, public ordering, checkpoint effects,
manifest summaries, warnings, and debug lifecycle.
The framework must remain domain-neutral. It may format stable reason codes and
validator-supplied corrective guidance, but it must not infer D&D or other
domain rules from artifact JSON.
### Producers And Artifact Families
The producing module owns prompt selection, prompt inputs, typed decoding, and
the model-facing representation that corresponds to its candidate. An
LLM-backed producer that supports feedback-aware correction must return the
exact response material that the model should see as its prior assistant turn.
It must not substitute a normalized artifact containing deterministically
attached UUIDs or other opaque application identity.
Artifact-family validators own semantic decisions and domain-specific
corrective guidance. Operator-facing explanation and model-facing correction
are separate contract fields even when their concise text happens to match.
## Validation Outcome Model
Each validator invocation produces one of four framework outcomes:
| Outcome | Meaning | Effect |
| --- | --- | --- |
| Approved | The validator completed and accepted the whole candidate. | Retain its warnings and continue the chain. |
| Rejected | The validator completed and found a semantic defect in the candidate. | Record the finding, continue the chain, and make the candidate ineligible for acceptance. |
| Failed | The validator could not return a usable decision because of an internal, transport, generation, structural-output, or result-invariant failure. | Retry that validator when eligible, then record incomplete validation and continue the chain unless cancellation or framework integrity prevents it. |
| Skipped | Runtime prerequisites for an otherwise selected validator cannot be satisfied. | Record a deterministic incomplete-validation diagnostic and continue; do not invent a semantic decision. |
Configured validator order controls invocation order and aggregate feedback
order. Execution remains sequential initially. The framework must continue
after a rejection and after an isolated validator failure when it can safely
prepare the remaining validator requests. Cancellation, inability to preserve
an immutable candidate, debug persistence failure, or another framework
integrity failure remains immediately terminal.
### Outcome Precedence
For one candidate, apply this precedence:
1. A producer structural failure means no acceptable candidate exists and
cannot be converted into validator approval.
2. Any completed semantic rejection makes the candidate rejected, even when
another validator failed or was skipped.
3. With no semantic rejection, a validator failure or skip makes validation
incomplete and invokes the validator-failure policy.
4. Only a structurally valid candidate with no rejection and either complete
validation or an explicit `warn_continue` decision may advance.
Do not turn a known rejection into acceptance through a permissive
validator-failure policy. Do not turn a structurally invalid response into a
rejected-but-usable artifact.
## Corrective Feedback Contract
`ValidationResult` should gain a separately bounded, optional model-facing
correction field. A rejecting production validator should provide:
- a stable reason code suitable for aggregation and provenance;
- an operator-facing message suitable for ordinary diagnostics; and
- concise corrective guidance that explains the violated rule without asking
the model to reproduce opaque identity or leaking unrelated source data.
The framework constructs one deterministic correction message from all
rejections in validator order. Each entry identifies the stable reason code
and corrective guidance. Duplicate identical entries may be collapsed while
preserving first occurrence; distinct findings must not be discarded merely
to shorten the message. If a validator rejects without model-facing guidance,
the framework uses a generic reason-code-based correction rather than copying
the operator message automatically.
Warnings, validator failures, skipped diagnostics, provider messages, stack
traces, debug paths, and sensitive values are not corrective guidance. They may
be recorded through their proper diagnostic channels but must not be presented
to the producer as candidate defects.
The framework must validate UTF-8, role, non-empty content, and
application-owned size limits before constructing the correction request. Oversized or
invalid correction material is a framework-owned inability to perform a
feedback retry; it must never be silently truncated into a misleading or
syntactically defective assistant response.
## Producer Correction Contracts
Introduce application-owned, defensively copied correction contracts at the
framework boundary:
- chunk, typed extraction, typed merge, and typed normalize results can carry
optional model-facing candidate material associated with their returned
value;
- the corresponding requests can carry an optional correction containing the
latest assistant material and aggregated guidance;
- `StructuredCompletionRequest` can carry the two bounded appended messages
without importing PromptKit types into module or pipeline contracts; and
- the PromptKit adapter translates those application-owned messages into
`RunRequest.AppendedMessages` using `promptkit.RoleAssistant` and
`promptkit.RoleUser` before preparation.
A semantic correction always supplies exactly two appended messages: the
latest defective response as `assistant`, followed by the aggregate correction
request as `user`. The framework does not expose the other PromptKit-supported
roles through this contract and does not accumulate messages from earlier
stage attempts. PromptKit preserves message content exactly, but Notarius must
reject empty content and enforce its own per-message and aggregate byte limits
before the adapter is called.
Correction material is attempt-local sensitive data. It is not part of the
artifact schema, checkpoint value, cache key, durable output bundle, ordinary
error, or configuration summary. The policy and capability that affect
execution do participate in resolved pipeline and checkpoint identity.
LLM-backed modules selected with both `retries > 0` and a non-empty validator
chain must declare whether they can produce and consume correction material.
Preparation must reject a pipeline that could request feedback-aware semantic
retries from an LLM-backed producer without that capability. An LLM-backed
producer with no validators may continue to use its retry budget for
operational failures without declaring semantic-correction capability.
A correction-capable producer must supply the exact single LLM response that
directly controlled the candidate being validated. Direct D&D chunk and
extraction producers expose their exact structured response. The shared
semantic-reconciliation path exposes its exact proposal response through its
typed normalizers without turning request-local batch handles into durable
identity. Deterministic transformations after that response are permitted only
when the validated candidate remains directly traceable to it.
A producer whose candidate combines multiple LLM responses is not
correction-capable under this initial protocol. It may continue to use ordinary
operational retries when no semantic correction can occur, but configuration
must reject a validator-backed retry workflow for it. Supporting compound
producers later requires a separately reviewed multi-response protocol; the
framework must not synthesize an assistant message by serializing the final
typed artifact.
Deterministic producers do not receive correction material. A deterministic
candidate rejected by validation immediately applies the terminal semantic
rejection policy without consuming retries that cannot change the result.
## Retry Budgets
### Producer Stage Budget
The existing producer binding `retries` field remains the sole outer stage
budget. `retries: N` means at most `N` additional complete producer attempts
after the initial attempt. Producer operational errors, producer structural
failures, module-requested normalize retries, and semantic corrections all
draw from this same budget. Do not add a separate semantic retry counter.
Every LLM-backed producer attempt receives the configured PromptKit
`structured_output_repair_attempts` value independently. Notarius does not
decrement that value across stage attempts.
### Validator Budget
Use the existing `retries` field on an LLM-backed validator binding for
additional attempts to obtain a usable decision about the same immutable
candidate. A completed approval or rejection is terminal for that validator
and does not consume another validator attempt. A validator retry reconstructs
the same ordinary validator prompt; it does not append semantic feedback about
the validator's prior failed judgment and does not create a recursive
Notarius correction loop.
Reject a positive validator `retries` value on a deterministic validator at
configuration resolution because repeating the same pure decision cannot
improve it. Validator retries do not consume the producer stage budget.
## PromptKit v0.9.0 Adoption
The target end state pins PromptKit v0.9.0 for correction requests. The
resolved dependency graph includes its independently versioned
OpenRouter and Rakestrawhome catalog modules through ordinary Go module
resolution; Notarius must not import or register those catalogs directly.
PromptKit continues to own their built-in backend and profile IDs, source
precedence, credentials, and capacity behavior.
Notarius's PromptKit compatibility documentation and built-in-profile
checkpoint marker identify v0.9.0 rather than v0.8.0. The PromptKit release
identity remains the conservative checkpoint identity for the exact catalog
versions selected by that release; Notarius should not duplicate upstream
catalog module versions in a second hand-maintained marker.
PromptKit v0.9.0 restricts text-chat roles to `developer`, `system`, `user`, and
`assistant`. Maintained Notarius prompt definitions already use only `system`
and `user`; correction requests add only `assistant` and `user`. PromptKit
`RunRequest` literals remain keyed. These compatibility conditions must remain
covered by the ordinary production-asset and adapter checks without adding a
brittle inventory test that merely counts prompt messages or literals.
## Terminal Policy Configuration
Add an optional `validation_policy` object at pipeline scope and on chunk,
extract, merge, and normalize producer bindings:
```yaml
validation_policy:
producer_structural_failure: fail_run
semantic_rejection: fail_run
validator_failure: warn_continue
```
The binding object overrides individual pipeline values; resolution is
field-by-field in binding, pipeline, application-default order. Omitted values
inherit rather than replacing the complete object. Explicit null, unknown
fields, and unknown enum values are invalid. The effective policy is resolved
and detached before execution, appears in redacted effective configuration and
run provenance, and participates in the resolved pipeline digest and checkpoint
identity.
The initial enum values and defaults are:
- `producer_structural_failure`: `fail_run` by default; `reject_output` may
retain a terminal rejection and final raw candidate for debug, but may not
advance or publish an invalid artifact;
- `semantic_rejection`: `fail_run` by default after stage attempts are
exhausted; `reject_output` records the aggregate rejection and allows
unrelated work to complete without advancing that candidate; and
- `validator_failure`: `warn_continue` by default, which advances a
structurally valid and otherwise unrejected candidate with explicit
incomplete-validation provenance and one genuine warning; `fail_run`
terminates the run.
Producer structural policy applies only to LLM-backed producers. Semantic and
validator-failure policies apply to any validated producer. Input and output
bindings do not accept `validation_policy`, and validator bindings do not own
terminal policy; they own only their decision and their own operational retry
budget. Candidate disposition belongs to the chunk, extract, merge, or
normalize producer binding after its complete validator chain has run.
Keep the current file-configuration version. The syntax is strictly
decodable without a version change. The project is pre-v1, but the behavior
and output changes should still be called out in the next release note and
downstream documentation.
## Stage-Specific Behavior
### Chunk
A generated chunk plan is structurally validated and materialized before the
validator chain runs. Semantic feedback applies to the exact raw chunker
response associated with that plan.
When an automatically reused chunk-plan record is rejected by the current
validator chain, treat the record as unusable for this invocation and enter
ordinary generation at attempt one. A cache hit is not a new model attempt and
does not supply model-facing assistant material. Do not overwrite the cached
record until a newly generated plan is accepted. Refresh and bypass modes
retain their existing publication rules.
### Extract
Each chunk-scoped extraction job owns its own attempt state and correction
conversation. One rejected chunk candidate does not cancel unrelated chunks or
lanes unless terminal policy converts it into a framework error. Deterministic
public ordering remains chunk-first and lane-second regardless of concurrent
completion.
### Merge And Normalize
Merge and normalize remain serial within a lane. A correction attempt receives
the same accepted upstream artifacts and references as the initial attempt.
The existing safe-fallback `NormalizeRetry` mechanism must be reconciled with
the shared attempt state rather than layered into a second retry loop: it uses
the same stage budget, retains its documented fallback behavior, and cannot
override a known validator rejection.
The initial feature supports one exact producer-supplied assistant response per
candidate attempt. Future multi-request normalization or batching must define
which response directly represents the candidate, or supply a new explicitly
reviewed correction protocol, before it can claim feedback-aware correction.
## LLM-Backed Validators
An LLM-backed validator uses the same scheduled PromptKit client, selected
profile, session, timeout, and structural-repair policy as other LLM-backed
modules. PromptKit may structurally repair its response inside one validator
attempt.
- A contract-valid validator response is its decision; Notarius does not ask a
second LLM to judge that judgment.
- A structurally invalid final validator response, transport failure, or
deterministic violation of the validator-result contract is a validator
execution failure.
- Validator execution retries reuse the immutable producer candidate and do
not regenerate it.
- Exhaustion invokes `validator_failure` policy and emits a genuine warning
under `warn_continue`.
This feature supplies the generic execution model only. It does not register a
production LLM-backed validator or change a D&D default validator chain.
## Provenance, Diagnostics, And Sensitive Data
Attempt debug output should make the state machine auditable. When debug is
enabled, record:
- producer attempt number and whether it was initial, error retry, module
retry, or semantic correction;
- PromptKit repair count and cumulative usage for every completion;
- each validator's configured-order outcome and validator attempt count;
- aggregate rejection codes and the bounded correction message;
- effective terminal policy and the decision it produced; and
- whether validation was complete, rejected, or incomplete.
Raw assistant responses and correction messages belong only in explicitly
requested detailed debug traces, following existing allowlisted content-file,
redaction, permission, and retention rules. Ordinary errors, CLI output,
warnings, manifests, checkpoints, caches, and run receipts contain identities,
counts, bounded safe summaries, and reason codes—not raw source or model
content.
The durable run manifest and rejection summaries should record enough
structured information to distinguish:
- the number and kinds of producer attempts;
- completed semantic rejection and all rejecting validator identities;
- incomplete validation and failed or skipped validator identities;
- the effective terminal policy and terminal result; and
- successful use of a correction attempt without treating it as a warning.
Warnings from abandoned producer attempts must not be promoted. Warnings from
the accepted attempt remain eligible. A warn-and-continue validator failure
produces one bounded, deterministically ordered warning per affected validator
after its retry budget is exhausted; detailed repeated failures stay in debug
provenance.
## Checkpoints, Caches, Concurrency, And Cancellation
- Effective validation policy, producer correction capability/protocol
version, validator chain, validator retry budgets, and prompt assets must all
affect checkpoint identity.
- Only accepted, completely validated stage outputs may be checkpointed or
reused. Rejected, structurally invalid, and validation-incomplete outputs
accepted under a permissive policy must not be written as reusable stage
checkpoints. This conservative rule avoids treating a transient validator
outage as durable validation success; a future checkpoint-status contract may
revisit it explicitly.
- Correction attempts use the same run-wide scheduler and worker bounds as
initial completions. No retry path may bypass provider admission.
- A scheduled permit covers the complete PromptKit operation, including its
internal structural repair, and is reacquired normally for a later Notarius
stage attempt.
- Parent cancellation dominates producer, validator, retry, debug, cache, and
checkpoint work. Cancellation never becomes a rejection, warning, or
incomplete-validation acceptance.
- Framework errors retain deterministic selection and cancellation behavior
across concurrently executing chunks and lanes.
## Architecture Record And Canonical Documentation
The target end state includes an accepted ADR that records:
- the separation between PromptKit structural repair, producer stage attempts,
and validator execution retries;
- the complete validator-chain aggregation rule and outcome precedence;
- the fresh reconstruction plus two-message correction protocol;
- module ownership of model-facing candidate material;
- deterministic producer and non-recursive validator behavior;
- default fail-closed and fail-open terminal policies; and
- provenance, cache, identity, and sensitive-data constraints.
The canonical owners describe the implemented behavior without duplicating
one another:
- `docs/policy/architecture.md` for durable validation and retry invariants;
- `docs/config.md` for fields, values, precedence, defaults, and validation;
- `docs/operations.md` for costs, failure behavior, warnings, debug handling,
and recovery;
- `docs/internal/pipeline.md` for the attempt state machine, aggregation,
checkpoint behavior, and concurrency;
- `docs/internal/llm.md` for appended correction messages and the distinction
from PromptKit repair;
- `docs/internal/modules.md` for producer and validator contracts;
- `docs/integrations/pkg-promptkit.md` for PromptKit v0.9.0,
`RunRequest.AppendedMessages`, supported message roles, application-owned
bounds, and the independently versioned upstream catalog boundary; and
- affected output and subprocess integration documents for durable validation
status and rejection summaries.
Until this target state is implemented, canonical current-state documentation
continues to describe the existing behavior.
## Testing Strategy
Tests should protect observable state-machine behavior rather than private
helper layout or exact prose. The target test suite includes:
- contract tests proving the first request is unchanged and a correction
request contains the same initial messages plus exactly one assistant and one
user message;
- behavioral runner tests for all-approved, multiple-rejection,
rejection-plus-failure, failure-only, skipped, retry-success, and each
terminal policy outcome;
- one representative path for chunk, extract, merge, and normalize, without
duplicating the complete state matrix at every stage;
- proof that all validators see immutable equivalent candidates and execute in
configured order after an earlier rejection or isolated failure;
- proof that validator retries reuse the candidate and do not consume producer
retries;
- proof that deterministic rejection does not repeat the producer;
- focused PromptKit-adapter tests proving that application-owned correction
messages map to the two intended PromptKit roles without content leakage;
- config parsing, precedence, invalid-placement, round-trip, redaction, and
digest tests for effective policy;
- checkpoint and chunk-cache tests for rejected, corrected, incomplete, and
accepted outcomes;
- warning, manifest, receipt, debug, and sensitive-content tests at their
canonical boundaries; and
- a small assembled D&D pipeline test proving a rejected direct extraction can
be corrected without a live provider.
Tests remain offline and deterministic. Do not reproduce PromptKit's internal
message-copying, rendered-hash, prepared-execution, capacity, or repair suite.
One representative adapter or assembled-run test should prove that PromptKit
structural repair remains usable after Notarius appends semantic-correction
messages. Do not snapshot full prompts or error prose, assert private
constants, or multiply equivalent tests across every D&D artifact family.
## Acceptance Criteria
- Every applicable validator runs in configured order and contributes one
explicit outcome before candidate disposition.
- Multiple semantic rejections produce one bounded, deterministic correction
request containing all actionable findings.
- Correction attempts reconstruct the exact ordinary prompt and append only
the latest defective assistant response and one correction message.
- Notarius pins PromptKit v0.9.0 and routes correction messages through
`RunRequest.AppendedMessages`; it does not maintain paired correction prompt
manifests or bypass PromptKit's normal execution path.
- Every correction-capable LLM producer exposes the exact single response that
directly controlled its candidate. Configuration rejects semantic retries
for compound producers that cannot satisfy that contract.
- The existing producer `retries` value is the only producer-stage budget;
PromptKit structural repair and validator execution retries remain separate.
- Deterministic producers are not repeated after semantic rejection.
- Validator execution failure is never described to the producer as a
candidate defect and never creates recursive semantic validation.
- Default terminal behavior is `fail_run` for structural failure and semantic
rejection, and `warn_continue` with explicit incomplete validation for
validator failure.
- Terminal policy resolves field by field from producer-binding override to
pipeline default to application default; individual validators do not own
candidate disposition.
- Permissive policy never advances known rejected or structurally invalid
output.
- Raw model responses and correction content are confined to model requests and
explicitly requested debug traces.
- Checkpoint, cache, manifest, warning, concurrency, cancellation, and
deterministic-ordering invariants remain intact.
- Canonical architecture, configuration, operations, internal, integration,
and ADR documentation accurately describe the implemented behavior.
- Focused, full, and race-enabled Go tests; vet; builds; example validation;
formatting; link checks; and repository hygiene checks pass.
## Non-Goals
- Adding the D&D combat-scene semantic validator.
- Redesigning the warning taxonomy beyond the warnings required for validator
failure and retry outcomes.
- Concurrent validator execution.
- Unbounded or accumulating conversational history.
- A second semantic retry counter.
- Recursive LLM judgment of LLM-validator decisions.
- Provider-specific retry policy or bypassing PromptKit.
- General workflow graphs or new pipeline stages.
- Large-collection reconciliation batching or a generic multi-response
correction protocol.