Plan feedback-aware validation retries

This commit is contained in:
2026-08-26 18:51:36 +00:00
parent 916d9210fd
commit c8a29a5fa2
7 changed files with 1264 additions and 1843 deletions

View File

@@ -83,9 +83,7 @@ entity identity, checkpoint contracts, or cross-request correlation. Changes
to shared protocol and policy assets must participate in the normal prompt, to shared protocol and policy assets must participate in the normal prompt,
schema, and checkpoint fingerprint mechanisms. schema, and checkpoint fingerprint mechanisms.
Acceptance of this decision does not imply that the shared mechanism or its The shared mechanism and its initial D&D registry consumers are now
consumer migrations are implemented. The implemented. Current behavior is documented in
[feature roadmap](../roadmap/semantic-reconciliation.md) owns target behavior [Module Internals](../internal/modules.md#semantic-reconciliation) and
and status, and the [D&D Module Internals](../internal/dnd.md#semantic-registry-reconciliation).
[implementation plan](../roadmap/implementation.md) owns delivery sequence
until the work is complete.

View File

@@ -1,224 +0,0 @@
# D&D Subprocess Consumer Documentation
## Status
Completed. The target guide is `docs/consumers/dnd-pipeline.md`.
## Purpose
Provide one task-oriented guide for applications that run Notarius as a
subprocess to execute the maintained complete D&D pipeline and consume its
published artifacts. The initial concrete consumer is Narratio, but the guide
must describe the public Notarius workflow rather than depend on Narratio
internals.
The guide should make the safe integration path obvious without duplicating
the CLI, input, receipt, output-bundle, or individual artifact contracts that
already have canonical documentation.
## Current State
The public integration surface is documented accurately but is distributed
across several documents:
- `docs/consumers/subprocess.md` defines the generic subprocess workflow;
- `docs/cli.md` owns commands, flags, stream behavior, and exit statuses;
- `docs/integrations/seriatim.md` owns the accepted transcript input shape;
- `docs/integrations/run-result.md` owns the machine-readable successful-run
receipt;
- `docs/integrations/json-output.md` owns bundle discovery and logical files;
- the D&D integration documents own the individual lane payload contracts;
- `examples/dnd-complete.config.yml` is the maintained complete pipeline.
A consumer can reconstruct the full workflow from those documents, but there
is no D&D-focused guide that connects the maintained example to its input,
invocation, complete artifact inventory, discovery procedure, and downstream
acceptance decisions.
## Target Documentation Set
### Create `docs/consumers/dnd-pipeline.md`
This document should own the end-to-end consumer workflow for the maintained
complete D&D configuration. It should be useful to Narratio and to another
subprocess orchestrator with the same needs.
The guide should contain the following sections.
#### Prerequisites And Deployment Configuration
- Link to `examples/dnd-complete.config.yml` rather than embedding a second
complete configuration.
- Explain that a deployment must provide the configured PromptKit profile and
campaign reference files.
- Recommend absolute paths for a service or orchestrator deployment.
- Call out the path-resolution distinction explicitly: YAML reference paths
are relative to the Notarius configuration file, while
`promptkit.profile_file` is relative to the Notarius process working
directory.
- Recommend validating the selected configuration and `dnd-session` pipeline
before processing sessions.
#### Transcript Input
- State that the complete pipeline consumes a Seriatim JSON document.
- Link to the canonical Seriatim contract for required fields and validation.
- Recommend the caller's final trimmed transcript when the caller maintains
transcript tiers. For Narratio, identify the implemented source as
`narratio.transcript.final_trimmed`, normally stored at
`transcripts/final.trimmed.json`.
- Explain that segment IDs must remain stable because D&D source references
cite those units.
- Explain that Notarius derives its default prompt session from the input
module and exact input bytes and that ordinary callers should not supply
`--session-id`.
#### Subprocess Invocation
- Show one concise invocation using `notarius run dnd-session`, explicit
absolute `--config`, `--input`, and `--output-dir` paths, and `--json`.
- Direct callers to capture stdout and stderr separately, propagate
cancellation, impose an operator-appropriate timeout, and wait for process
completion before parsing stdout.
- State that only exit status zero permits receipt decoding and link to the CLI
contract for the complete exit-status definition.
- Recommend retaining stderr and the invocation context for diagnosis without
logging secrets or transcript content.
#### Receipt And Bundle Discovery
- Require callers to accept only supported run-result schema versions while
tolerating unknown fields allowed by that version.
- Direct callers to obtain the exact run-specific bundle from the receipt's
absolute `output_directory`; they must not scan for the newest run directory
or construct a run ID.
- Require a confinement check when resolving `index_file` beneath the reported
bundle root.
- Direct callers to discover lane payloads by `lane_id` in `index.json`, then
verify descriptor media type and schema identity before decoding them.
- Explain that descriptor paths are untrusted relative paths and require the
same confinement discipline.
#### Complete D&D Artifact Inventory
Include a compact table for the ten lane IDs selected by the maintained
complete configuration:
- `item-registry`;
- `npc-registry`;
- `location-registry`;
- `scene-descriptions`;
- `item-occurrences`;
- `spells`;
- `combat-turns`;
- `npc-occurrences`;
- `location-occurrences`;
- `enemy-events`.
For each row, give a one-line purpose and link to the corresponding canonical
D&D artifact contract. Do not copy its fields or schema rules into the
consumer guide.
Document the four always-published bundle files—`index.json`, `manifest.json`,
`rejected.json`, and `warnings.json`—and the complete example's configured
`chunk-map.json` and `evidence-context.json` pipeline-wide artifacts. Link to
their canonical contracts and distinguish pipeline-wide artifacts from lane
outputs.
The inventory must say that a file is available only when its corresponding
artifact was accepted and published. It must not imply that process success
guarantees every configured lane.
#### Downstream Acceptance And Retention
- Explain that exit status zero can coexist with rejected outputs, warnings,
or absent lane descriptors.
- Require the consumer to define its required lane set explicitly. Recommend
treating all ten lanes as required when the caller claims to consume the
complete D&D workflow, while allowing another consumer to adopt a narrower
documented policy.
- Recommend retaining the receipt, the complete published bundle, and captured
diagnostic streams long enough to support provenance and failure analysis.
- Explain that `evidence-context.json` is a reading excerpt; authoritative
citations remain in lane payloads.
- Treat transcripts, lane artifacts, evidence context, manifests, and logs as
sensitive campaign data.
#### Compatibility Checklist
End with a concise checklist covering process exit, receipt schema, path
confinement, pipeline identity, index decoding, required descriptors,
descriptor schema/media compatibility, warnings and rejections, checksums or
retention, and secure handling. Compatibility should be based on published
receipt and artifact contracts rather than parsing a human version string.
### Update Existing Navigation
- Add a short link from `docs/consumers/subprocess.md` to the D&D-specific
workflow. Keep generic subprocess policy in the existing document.
- Add the guide to the documentation links in `README.md`.
- Extend the subprocess-consumer row in `docs/development.md` so maintainers
working on the D&D workflow are routed to the new guide and the canonical
contracts.
### Verify Canonical Contract Documents
Review the linked integration documents and the complete example while writing
the guide. Correct an integration document only if repository inspection finds
an actual stale contract. Do not move schema definitions, field tables, CLI
flags, or configuration semantics into the new guide.
## Narratio Alignment
The guide may name Narratio as the motivating consumer and identify its current
final-trimmed transcript source. It must not claim that Narratio already has a
Notarius adapter or extraction stage. Until that feature is implemented,
Narratio-specific architecture, configuration, stage behavior, manifest
records, and artifact source IDs belong in Narratio's roadmap.
Once Narratio implements the integration, its own integration documentation
should link to this guide and the durable Notarius contracts instead of
repeating them.
## Validation
Documentation implementation should include:
```sh
go run ./cmd/notarius config validate \
--config examples/dnd-complete.config.yml \
--pipeline dnd-session
go test ./...
```
Also verify all new and changed relative Markdown links, compare the artifact
inventory directly with the maintained complete configuration, and confirm
that commands and path semantics match the CLI and configuration references.
If the repository still has no automated link checker, record that fact and
perform a focused manual link review.
## Acceptance Criteria
- A subprocess integrator can follow one D&D-focused guide from a Seriatim
transcript through safe discovery of every artifact configured by the
complete example.
- The guide makes stdout, stderr, exit-status, receipt, and path-confinement
responsibilities unambiguous.
- The ten configured D&D lanes and both configured pipeline-wide artifacts are
listed and linked to their canonical contracts.
- The guide distinguishes process success from the caller's required-artifact
policy.
- The profile-path and reference-path resolution rules are clearly stated.
- Existing navigation makes the guide discoverable.
- No volatile contract is defined in two places, and no unimplemented Narratio
behavior is presented as current.
## Non-Goals
- Implementing or documenting Narratio's future adapter or stage as current
Notarius behavior.
- Adding a new Notarius command, receipt version, output format, or artifact
schema.
- Duplicating the complete configuration or individual D&D payload schemas in
prose.
- Defining a universal partial-result policy for every Notarius consumer.

View File

@@ -13,124 +13,15 @@ structural output repair within one completion. Notarius owns stage candidates,
validator chains, semantic rejection policy, and whether another stage attempt validator chains, semantic rejection policy, and whether another stage attempt
is warranted. is warranted.
### 1. Upgrade To PromptKit v0.8.0 ### Feedback-Aware Stage Validation Retries
This item has been promoted to the standalone This item has been promoted to the standalone
[PromptKit v0.8.0 Upgrade](promptkit-v0.8.md) roadmap. That document owns the [Feedback-Aware Stage Validation Retries](validation-retries.md) roadmap. That
release-by-release compatibility review, adopted features, structured-repair document owns the target validation state machine, correction protocol, retry
policy, target integration boundary, acceptance criteria, and settled design budgets, terminal policies, provenance requirements, settled producer
decisions. contracts, and PromptKit v0.9.0 adoption.
### 2. Feedback-Aware Stage Validation Retries ### D&D Combat Scene Semantic Validation
- Model Notarius's corrective stage-retry conversation explicitly after
PromptKit v0.8.0. The first attempt sends the ordinary complete initial
prompt. If application validation rejects the resulting LLM-produced
candidate and another stage attempt is available, reconstruct that complete
initial prompt byte-for-byte and append exactly two messages: an assistant
message containing the defective response and an application-owned user
message detailing every applicable semantic validation error and requesting
one corrected, complete replacement response. This is a freshly constructed
correction request, not continuation of an accumulating conversation.
- Use the configured stage `retries` value as the one outer retry budget for
this loop. `retries: N` continues to mean at most `N` additional complete
chunk, extract, merge, or normalize attempts after the initial attempt,
whether an attempt is needed because of a producer error or semantic
rejection. Do not add a second semantic-correction count. PromptKit's
prompt-level `repair_attempts` budget is independent and internal to each
individual LLM completion, and does not consume or replenish the Notarius
stage budget.
- Extend the framework-managed validation boundary for chunk, extract, merge,
and normalize stages so a rejected LLM-produced candidate and its exact raw
model response remain available to construct the next stage attempt.
Deterministic producers cannot improve by repeating the same inputs; a
rejection from a deterministic stage is therefore terminal under the
configured rejection policy rather than consuming retries mechanically.
- Preserve the original session ID, selected profile, structured-output
contract, prompt inputs, and reusable prompt prefix. Carry only the latest
candidate and latest aggregate feedback; do not build an unbounded retry
conversation. Keep model-facing corrective guidance separate from
operator-facing diagnostics, and apply explicit size, redaction, and debug
disclosure rules to both.
- Run every applicable validator in the configured chain before deciding
whether to retry. Do not short-circuit merely because an earlier validator
rejected the candidate. Aggregate all semantic rejection reason codes and
corrective guidance into the retry message so one retry can address the
whole candidate. A validator is applicable only when its declared target and
prerequisites can be satisfied; record a deterministic skipped diagnostic
rather than invoking a validator on an input it cannot interpret. Initially
execute the chain sequentially in configured order so results, diagnostics,
costs, and feedback ordering remain deterministic; consider validator
concurrency only in response to measured latency.
- Continue running independent applicable validators after one validator
execution failure so the attempt retains as much useful diagnostic
information as practical. Do not present validator operational failures as
defects in the producer candidate and do not include them in corrective
feedback.
- Distinguish three terminal conditions and make their policies configurable
at a coherent pipeline or binding scope:
- **producer structural failure:** PromptKit could not return a usable
structured candidate after its repair budget. Default to `fail_run`; an
allowed alternative may record a terminal stage or lane rejection where
execution can safely continue, but may not accept the invalid output;
- **semantic rejection:** one or more validators completed and rejected the
candidate. Default to `fail_run` after corrective stage retries are
exhausted; allow an explicit alternative that records the existing
rejected-output outcome without advancing that output;
- **validator execution failure:** a validator could not produce a valid
decision because of generation, structural-output, transport, or internal
failure. Default to a genuine warning and an explicitly recorded
`validation_incomplete` or equivalent degraded state while allowing the
candidate to continue; allow strict configuration to fail the run instead.
- An LLM-backed validator uses the same scheduled PromptKit boundary as every
other LLM-backed module. Its own response may use PromptKit's bounded
structural repair. Distinguish its possible output states:
- output rejected by PromptKit's structural contract should consume only the
validator prompt's configured PromptKit repair budget;
- output that is structurally valid but violates a deterministically
checkable validator-result invariant should be classified as a validator
execution failure;
- output that satisfies the complete validator-result contract is the
validator's decision, even though an LLM judgment may remain imperfect.
Automatically judging that judgment would require another semantic
validator and is outside this feature.
If the validator cannot return a contract-valid decision, do not recursively
create another Notarius semantic-validation loop around it. Apply the
configured validator-failure policy. The default warning must identify the
validator and affected stage without exposing sensitive content.
- Separate validator execution retry from producer correction. A transient
validator operational failure must not automatically discard and regenerate
an otherwise usable producer candidate. Any bounded retry of the validator
itself should reuse that same immutable candidate and remain subordinate to
PromptKit and provider retry behavior.
- Preserve attempt-level provenance, cumulative token usage, validator
outcomes, aggregated correction feedback, and terminal policy decisions in
the debug and manifest models without copying raw source material into
ordinary errors or durable summaries.
- Define terminal-outcome precedence. A semantic rejection dominates a
validator execution failure for the same candidate: use the completed
rejections to correct the producer while separately recording incomplete
validation. If a later candidate has no semantic rejection but one validator
still fails, apply the configured validator-failure policy to that candidate.
Never allow a known semantic rejection to become accepted through a
warn-and-continue setting, and never accept a structurally invalid producer
response. Permissive policy may preserve a rejected-output outcome or accept
a structurally valid candidate with explicitly incomplete validation; it may
not relabel known-invalid output as approved.
Before implementation, record the generic validation and retry state machine
in an ADR. The ADR should own the separation between PromptKit repair and
Notarius correction, use of the existing stage-retry budget, reconstruction of
correction conversations, all-applicable-validator aggregation, deterministic
validator ordering, non-recursive validator failure handling, outcome
precedence, default fail-open/fail-closed choices, configurable terminal
policies, and provenance and sensitive-data constraints. A dependency-upgrade
ADR is not needed for PromptKit v0.8.0 itself. Current behavior remains
authoritative until the validation ADR is implemented and the canonical
architecture, configuration, operations, and internal documentation are
updated.
### 3. D&D Combat Scene Semantic Validation
- Add an optional production LLM-backed D&D validator that determines whether - Add an optional production LLM-backed D&D validator that determines whether
proposed scene boundaries and classifications represent substantive active proposed scene boundaries and classifications represent substantive active
@@ -172,7 +63,7 @@ updated.
the chunker or otherwise changes stage ownership or the durable chunk-plan the chunker or otherwise changes stage ownership or the durable chunk-plan
contract. contract.
### 4. Warning Signal And Presentation Reform ### Warning Signal And Presentation Reform
- Audit every warning producer and representative successful runs. Ordinary - Audit every warning producer and representative successful runs. Ordinary
success producing dozens of warnings is a failed operator experience: the success producing dozens of warnings is a failed operator experience: the

File diff suppressed because it is too large Load Diff

View File

@@ -1,520 +0,0 @@
# PromptKit v0.8.0 Upgrade
## Status
Proposed.
## Purpose
Upgrade Notarius from PromptKit v0.5.0 to v0.8.0 and deliberately adopt the
useful correctness, profile-composition, provider-diagnostic, backend, and
structured-output-repair capabilities introduced in PromptKit v0.6.0, v0.7.0,
and v0.8.0.
The upgrade should improve structured-output reliability without confusing
PromptKit's bounded deterministic repair with Notarius's existing stage retry
budget or the future feedback-aware semantic-validation loop. PromptKit types
and provider behavior must remain behind Notarius's transport-neutral LLM
boundary.
## Current State
Notarius currently pins PromptKit v0.5.0. Its production adapter prepares one
frozen execution, records credential-redacted details, and runs that same
prepared value. It maps PromptKit capacity failures to an application-owned
error, maps failed structured validation to `ErrInvalidStructuredOutput`, and
returns PromptKit's raw validated bytes and usage metadata.
Every maintained production prompt uses JSON Schema validation and currently
declares `repair_attempts: 0`. Notarius stage bindings separately expose
`retries`, which reruns a complete stage operation after an error or rejected
candidate. The two mechanisms have different ownership and must remain
independent.
Notarius also maintains:
- embedded prompt, schema, and fallback-profile filesystems;
- operator profile-file and profile-directory sources;
- one optional conventional `local` backend registration;
- explicit profile preflight through PromptKit inspection;
- one application-wide scheduled LLM client around the PromptKit adapter;
- PromptKit profile-source fingerprints for checkpoint safety; and
- redacted debug and manifest provenance at application-owned boundaries.
The upgrade must preserve those established responsibilities while revising
the pinned integration contract and any behavior affected by the three
intervening releases.
This roadmap is based on PromptKit's pinned release guides for
[v0.6.0](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.8.0/docs/releases/v0.6.0.md),
[v0.7.0](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.8.0/docs/releases/v0.7.0.md),
and
[v0.8.0](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.8.0/docs/releases/v0.8.0.md),
plus the public API and format documentation at the v0.8.0 tag.
## Target End State
- `go.mod` and `go.sum` pin PromptKit v0.8.0 without a local replacement or
vendored copy.
- Every maintained PromptKit prompt and profile prepares successfully under
v0.8.0's stricter validation and source-loading rules.
- Eligible Notarius structured completions use one PromptKit corrective call by
default after a structurally invalid response. Operators can explicitly set
a value from zero through three for a configured pipeline, with a more local
LLM-backed binding override where needed.
- PromptKit repair remains an inner operation within one Notarius stage
attempt. It never consumes or replenishes the binding's `retries` budget.
- A successful repaired result exposes cumulative usage and the actual repair
count to Notarius's application-owned response and debug models. A repaired
success is not itself a warning.
- Exhausted PromptKit validation remains an invalid structured-output result,
preserving the final candidate and diagnostics for debug and for any
applicable outer Notarius stage policy. Invalid structured output is never
accepted merely because the repair budget was exhausted.
- Profile inheritance, the built-in Rakestrawhome backend/profile, optional
credential behavior, and structured generation errors work through the
existing Notarius PromptKit boundary and are accurately documented.
- Provider-specific PromptKit types do not escape `internal/framework/llm`.
- Checkpoint identity, effective configuration, redacted summaries, and debug
provenance reflect every execution-affecting repair or profile change.
- Current documentation pins and describes v0.8.0; future Notarius semantic
validation retries remain roadmap behavior rather than being conflated with
this dependency upgrade.
## Release-by-Release Adoption
### PromptKit v0.6.0: Correctness, Safety, And Efficiency
PromptKit v0.6.0 adds no public declarations, but intentionally rejects several
formerly permissive or ambiguous inputs. The upgrade must audit Notarius's
embedded and operator-facing integration against these rules:
- YAML `id` and `version` metadata, rather than filenames, define prompt and
profile identity.
- Prompt `content_file` paths are exact, relative, contained paths; built-in
file artifacts must resolve to regular files.
- execution controls, output contracts, and repair budgets must be finite and
within their documented ranges;
- provider endpoints must be absolute HTTP or HTTPS URLs with a host and no
user information, query, or fragment;
- JSON documents and successful provider responses contain exactly one value;
- successful provider responses are bounded to 16 MiB; and
- JSON-compatible values are bounded for depth and expansion.
Notarius should rely on PromptKit for these rules rather than duplicate its
parsers or internal limits. Existing Notarius validation may retain a narrower
application rule where it has independent value, but overlapping validation
must agree with PromptKit and must not accept a value PromptKit will reject
later.
The upgrade automatically receives operation-local schema-plan reuse,
artifact-text memoization, improved cancellation checks, and transport error
identity preservation. Notarius should verify these changes through its real
adapter boundary and avoid adding a second cache or response-body layer that
would duplicate PromptKit's ownership.
### PromptKit v0.7.0: Profiles, Backend Access, And Generation Errors
#### Profile Inheritance
Operator profiles may use `base_profile` to alias or selectively refine a
built-in, fallback, or higher-precedence operator profile. Notarius must pass
profile sources through unchanged and let PromptKit own parent lookup, merge
rules, source precedence, cycle detection, and fully resolved prepared targets.
Preflight inspection must resolve inherited profiles through the same source
and backend composition used at execution. The selected leaf profile ID remains
the public profile identity, while effective backend, endpoint, model, and
reasoning provenance reflect the resolved chain. Notarius must not implement a
second inheritance parser.
The existing complete `dnd-extraction` fallback remains a standalone profile:
PromptKit v0.8.0 does not provide a built-in `openai/gpt-5.6-luna` profile that
would be an appropriate parent. Documentation should nevertheless explain how
operators can use inheritance for environment-specific workload profiles and
should link to PromptKit's pinned format contract rather than duplicate its
field-by-field merge algorithm.
Checkpoint safety must cover inherited behavior. Operator file/directory
digests already cover changes to definitions in those sources, fallback asset
digests cover application parents, and the PromptKit built-in catalog marker
must change from its v0.5.0 identity to v0.8.0 so a changed built-in parent
cannot reuse an incompatible checkpoint.
#### Rakestrawhome Backend And Profile
PromptKit's reserved `rakestrawhome` backend and
`rakestrawhome-gemma-4-31b` profile become available without Notarius-specific
registration. Notarius must not register or shadow the reserved backend ID.
Profile preflight, backend-capacity reporting, scheduling, generation, and
provenance should work for it through the same generic paths used by OpenRouter
and `local`.
The D&D default remains `dnd-extraction`; this upgrade does not silently move a
production workload to Rakestrawhome. Operator documentation should identify
the built-in profile as an available selection and link to PromptKit for its
endpoint, credential environment, model, and capacity defaults.
#### Optional Credentials
An absent or blank optional `APIKeyEnv` now causes PromptKit to omit the
`Authorization` header and send the request. Notarius must not restore the old
failure behavior by pre-reading provider credential environment variables or
by adding provider-specific authentication logic.
Profile inspection may report an explicit `APIKeyRequired` policy without
reading the credential, and execution remains the boundary at which that
requirement is enforced. For optional profiles, an authentication-requiring
provider may instead return a structured 401 or 403 generation failure. The
configuration and operations documentation must explain this distinction.
Notarius does not currently expose PromptKit's in-memory profile-registration
API to operators, and PromptKit's filesystem profile format does not expose
`APIKeyRequired`; therefore Notarius must not promise that an operator profile
can force local credential preflight. Operators should provision the named
environment variable, while Notarius should preserve the provider's structured
authentication failure when it is absent.
Notarius must continue to document mechanisms and environment-variable names,
never secret values.
#### Structured Generation Errors
The adapter should recognize `*promptkit.GenerationError` with `errors.As` and
translate useful information into an immutable, provider-neutral Notarius
error classification. At minimum, retain the HTTP status code so callers and
future retry policy can distinguish transport success with provider rejection
from other generation failures.
PromptKit's provider code, type, and message accessors are bounded but remain
untrusted and potentially sensitive. They must never appear automatically in
ordinary CLI output, warnings, manifests, checkpoint identity, or cache data.
If retained for an explicitly requested debug trace, they must pass through
Notarius's known-secret and bearer redaction and remain clearly identified as
untrusted provider diagnostics. Default error formatting should continue to
use a bounded, redacted application-owned message.
Capacity and cancellation retain their current more specific classifications
and precedence. This upgrade does not add automatic provider-error retry
classification; it only preserves safe structured data needed for diagnosis
and later policy.
### PromptKit v0.8.0: Bounded Structured-Output Repair
#### Default Policy
Every maintained production prompt whose output is consumed as structured data
should declare one repair attempt. All current production prompts use eligible
JSON Schema validation, so no current prompt needs a zero default merely
because of its output mode.
One repair means at most one corrective generation after the initial
candidate. PromptKit reconstructs the immutable original conversation and
appends only the latest invalid assistant candidate and latest deterministic
validation diagnostics. It preserves the selected target, direct session ID,
provider-native structured-output contract, and backend capacity policy. This
shape preserves the original cacheable prompt prefix and avoids accumulating
unbounded failed history.
The default is deliberately small. A single repair captures the common case in
which a capable model can correct malformed JSON or a schema violation after
receiving an exact diagnostic, while bounding the extra latency and cost of a
single structured completion.
#### Configuration Contract
The public configuration is an optional, presence-aware
`structured_output_repair_attempts` integer at pipeline scope and at each
LLM-backed module or validator binding. Its effective precedence is:
1. the binding value, when present;
2. the pipeline value, when present; and
3. the selected prompt's declared `repair_attempts` value.
The value must be from zero through three. Explicit zero disables PromptKit
repair at that scope. A deterministic binding must reject the field because it
cannot perform structured LLM repair. Validator bindings may use it only when
the selected validator is LLM-backed. Shorthand module bindings continue to
inherit the pipeline or prompt default.
The long, provider-neutral name is intentional: it distinguishes PromptKit's
inner structural repair from the existing binding `retries` field, which owns
complete stage attempts, without exposing a dependency name in generic
pipeline contracts.
The effective value must survive file parsing, cloning, redacted summaries,
pipeline resolution, and pipeline digest construction without pointer aliasing
or loss of presence. It must affect checkpoint identity because it can change
the selected result, latency, token usage, and provider cost.
#### Adapter Contract
The transport-neutral structured-completion request should carry an optional
application-owned structural-repair budget. No `promptkit.OutputContract` or
other PromptKit type may cross the adapter boundary.
PromptKit v0.8.0 request validation replaces the complete prompt output
contract rather than merging one field. When Notarius has a configured
override, the adapter must therefore inspect the selected prompt, copy its
normalized declared format, validation mode, and schema path, change only the
repair count, and supply that complete contract on the prepared request. A nil
override continues to use the prompt declaration directly. Inspection and
preparation must use the same immutable engine sources; a small adapter-local
cache keyed by normalized prompt ID and version is acceptable but not required
without measured need.
This approach prevents configuration from accidentally dropping JSON Schema
validation, avoids duplicating schema paths in pipeline YAML, and keeps prompt
assets authoritative for every output-contract field other than the explicit
operator override.
The transport-neutral structured-completion response should report the actual
number of PromptKit repair calls. PromptKit's returned token usage is already
cumulative and must be passed through without re-summing it. Debug records
should distinguish the configured budget from the actual count. Ordinary run
manifests need not gain raw prompt or response data merely to report repairs;
any durable aggregate should be added only if it has a clear consumer contract.
#### Result And Failure Semantics
- A valid initial candidate returns normally with zero actual repairs.
- A valid corrected candidate returns normally with cumulative usage and its
positive actual repair count. It does not emit a warning solely because a
repair occurred.
- Exhausting the repair budget returns PromptKit's final candidate and failed
validation result. The adapter maps this to
`ErrInvalidStructuredOutput`, preserves the response and debug material, and
does not decode or accept the candidate.
- An explicitly empty or whitespace-only candidate participates in the
declared structural validation and repair flow. Missing, `null`, or
non-string provider content remains a malformed provider response.
- A generation failure during a corrective call is an operational generation
failure and uses the same safe structured-error adaptation as an initial
generation failure.
- Context cancellation remains authoritative throughout the initial and
corrective calls.
PromptKit repair happens inside one scheduled `CompleteStructured` operation.
The Notarius scheduler holds one permit for that logical operation while
PromptKit performs its initial and serial corrective calls; PromptKit
reacquires its own selected-backend capacity for each corrective generation.
Because corrective calls are serial, this cannot expand actual concurrent
provider work beyond the number of admitted Notarius operations, but
documentation must stop describing the Notarius permit as a separate admission
event for every internal repair call.
One `CompleteStructured` invocation with effective PromptKit repair budget `R`
may make at most `R + 1` provider calls. If one stage attempt makes `C`
structured-completion invocations, a binding with `retries: N` has an upper
bound of `(N + 1) * C * (R + 1)` provider calls; `C` may itself be a bounded,
data-dependent module property, as it is for batched semantic reconciliation.
LLM-backed validators have their own corresponding invocation counts, budgets,
and costs. These formulas are upper bounds, not promises that every failure is
retryable or that every attempt reaches the provider.
## Profile And Prompt Source Compatibility
The upgrade must preserve Notarius's source precedence: an operator source,
then registered application fallback profiles, then PromptKit built-ins. A
selected malformed definition remains authoritative and fails rather than
falling through. Parent resolution introduced by profile inheritance observes
that same precedence.
All embedded prompt manifests, shared content fragments, response schemas, and
fallback profiles must be prepared or inspected offline under v0.8.0. The
review should specifically catch:
- IDs inferred accidentally from filenames;
- stale or escaping `content_file` paths;
- missing or non-regular embedded artifacts;
- repair values outside zero through three or paired with ineligible
validation;
- schemas or examples that are not exact single JSON documents;
- unsupported endpoint forms; and
- JSON-compatible variables or profile extras that exceed upstream bounds.
No prompt prose, schema shape, durable D&D artifact contract, or default D&D
model should change merely to exercise the dependency. Prompt manifests should
change only as needed to enable the adopted repair default and satisfy v0.8.0
contracts.
## Provenance, Debugging, And Security
- Update the opaque PromptKit built-in profile-catalog identity from v0.5.0 to
v0.8.0. Do not hash or publish PromptKit's internal catalog bytes.
- Ensure a prompt's repair default remains covered by its existing prompt asset
fingerprint and a configured effective override remains covered by the
resolved pipeline digest.
- Preserve selected leaf profile identity while recording the inherited
effective target already exposed by PromptKit inspection and prepared
details.
- Add actual structural-repair count and, when useful, the configured budget to
application-owned debug material. Token totals remain PromptKit's cumulative
values.
- Do not generate a warning for a successful repair. Repair exhaustion is an
invalid-output failure, while provider rejection is a generation failure.
- Never expose raw provider diagnostic fields without explicit debug capture
and application redaction. Do not place them in normal errors or durable
summaries.
- Preserve context and transport error identity sufficiently for
`errors.Is`-based cancellation and deadline handling after adapting the
external error.
## Documentation And Examples
Implementation should update current-state documentation only when the new
behavior lands:
- `docs/integrations/pkg-promptkit.md` must pin v0.8.0 and define the revised
prepared-execution, repair, profile-inheritance, backend, credential, and
error-adaptation boundary.
- `docs/config.md` must own the repair configuration fields, precedence,
allowed range, explicit-zero behavior, profile inheritance availability, and
optional credential semantics.
- `docs/operations.md` must explain structural repair cost, timeout and
concurrency effects, credential failures, and its distinction from stage
retries.
- `docs/internal/llm.md` must describe adapter contract replacement, actual
repair metadata, error adaptation, source compatibility, and scheduling.
- `docs/internal/pipeline.md` must describe how effective repair configuration
is resolved and how inner repair differs from outer stage attempts.
- `docs/policy/architecture.md` should receive only the durable ownership rule:
PromptKit owns bounded deterministic structural repair within one completion,
while Notarius owns stage attempts and semantic validation policy. Detailed
fields and retry formulas belong in their canonical configuration and
operations documents.
Update maintained configuration examples only if the public Notarius
configuration contract changes. A short inheritance illustration may remain in
the configuration reference; do not create a complete example solely to copy
PromptKit's upstream profile catalog. All upstream links must point to the
v0.8.0 tag. Historical release or archived roadmap references should remain
historical.
No ADR is required solely to pin a newer dependency. The durable separation
between PromptKit structural repair and Notarius semantic stage retries should
be stated in architecture documentation now; the more extensive future
validation state machine still warrants the separate ADR already identified in
`future.md` when that work is promoted.
## Validation And Acceptance Criteria
The implementation is complete when:
- the repository builds and tests against PromptKit v0.8.0 with no replacement
directive, workspace dependency, or vendored source;
- every maintained prompt and profile prepares or inspects successfully under
the v0.8.0 source, path, endpoint, output-contract, and JSON-value rules;
- an invalid first JSON Schema candidate followed by a valid correction returns
the valid raw output, cumulative usage, and actual repair count through the
Notarius adapter;
- repair exhaustion returns the final raw candidate and debug material with an
error matching `ErrInvalidStructuredOutput`;
- a corrective generation failure retains safe generation classification and
provider status without leaking untrusted provider detail;
- explicit empty content follows structural validation rather than being
misclassified by Notarius;
- repair configuration is presence-aware, range checked, rejected on
deterministic bindings, resolved with documented precedence, and included in
effective pipeline identity;
- inherited profiles resolve consistently during preflight and execution, and
changes to any relevant operator, fallback, or built-in parent invalidate
checkpoint reuse;
- the Rakestrawhome built-in profile reaches generic preflight, scheduling, and
provenance paths without application-specific registration;
- optional missing credentials and explicitly required credentials behave as
documented without contacting real providers in tests;
- cancellation, timeout, backend capacity, prepared-execution snapshot,
session ID, raw-output, debug-redaction, and existing profile provenance
behavior remain intact;
- maintained examples validate successfully; and
- canonical documentation contains no active v0.5.0 pin or claim that PromptKit
is always single-pass.
Tests should follow `docs/policy/testing.md`: exercise observable Notarius
contracts with offline fake clients or `httptest` boundaries, and do not copy
PromptKit's entire internal repair test suite or assert its exact correction
message prose. The dependency's internal wording is not a Notarius contract.
## Non-Goals
- Implementing Notarius's future feedback-aware semantic stage-retry loop.
- Adding the D&D combat-scene semantic validator.
- Redesigning warning policy or treating successful structural repair as a
warning.
- Adding provider transport retries or deciding which HTTP statuses should
consume a stage retry.
- Exposing PromptKit request, response, profile, validation, capacity, or error
types outside the LLM adapter.
- Changing durable artifact schemas, D&D prompt semantics, the D&D default
model, or the fixed pipeline shape.
- Reimplementing PromptKit profile inheritance, schema validation, response
bounds, repair conversations, backend admission, or provider parsing inside
Notarius.
## Decisions
### 1. Default Structured-Output Repair Budget
**Decision: default to one repair attempt.** Set every maintained
eligible production prompt to `repair_attempts: 1`. One corrective call is a
strong fit for Notarius because every current production LLM response has a
strict JSON Schema contract, smaller cost-effective models are a deliberate
deployment target, and a precise structural diagnostic often makes one retry
materially more successful. The budget is paid only after a structurally
invalid candidate and remains tightly bounded.
**Alternative considered: retain zero by default.** This preserves single-pass
cost and latency and requires operators to opt in. It is preferable for an
environment where every additional request is expensive or where upstream
provider-native schema enforcement already produces negligible invalid output.
It is less suitable as the Notarius default because one malformed response can
otherwise discard substantial completed pipeline work.
**Alternative considered: default to two.** This may improve recovery for
weak models, but it doubles the worst-case corrective cost relative to the
selected default and compounds with outer stage retries. It should be an
operator choice supported by configuration, not the initial default, unless
observational evidence shows that the second correction has a worthwhile
marginal success rate.
### 2. Repair Override Scope
**Decision: support both pipeline and LLM-backed binding overrides.** Use
the presence-aware `structured_output_repair_attempts` field and precedence
defined above. A pipeline value provides the convenient one-line control the
operator requested, while a binding value permits an expensive normalizer or
future LLM-backed validator to use a deliberately different budget. This
mirrors Notarius's established pipeline/binding profile inheritance and scales
without editing embedded prompts.
**Alternative considered: support only a pipeline override.** This is smaller to
implement and document and still permits global enablement or disablement for
one pipeline. Its drawback is that one exceptional prompt cannot opt out or
request a larger budget without changing an embedded asset for every pipeline.
**Alternative considered: expose one global value under the top-level
`promptkit` configuration.** This makes client construction simple, but applies
the same budget to unrelated pipelines and leaks an execution policy into the
dependency configuration block. It is less compositional than pipeline-owned
policy and therefore not recommended.
### 3. Retention Of Provider-Supplied Generation Details
**Decision: retain status in the application-owned error contract and
retain redacted provider code, type, and message only in explicitly requested
debug traces.** Status is useful for diagnosis and future retry policy without
usually containing sensitive data. The other fields can materially explain a
400 response but may echo request or schema content, so they belong only in the
already-sensitive debug surface after Notarius redaction.
**Alternative considered: retain only HTTP status and discard all provider fields.**
This is the safest and smallest policy and still improves typed failure
handling. It sacrifices potentially decisive provider diagnostics, leaving an
operator with less information when a provider returns a terse status and the
problem cannot be reproduced easily.
**Alternative considered: include bounded provider code and type in normal
errors while keeping message debug-only.** Codes and types are often stable and
less sensitive than messages, but PromptKit explicitly classifies every
provider field as untrusted. Promoting them to ordinary output creates a
disclosure and compatibility burden that is not currently justified.

View File

@@ -1,282 +0,0 @@
# Source-Only Releases
## Status
Implemented. Creating the first release under this procedure remains a
separate maintainer operation.
## Purpose
Define a repeatable, guarded release process for Notarius without taking on a
binary-distribution system that its current operator audience does not need.
The process should make an exact source revision, its compatibility impact,
and its validation status easy to identify while keeping installation in the
hands of technically capable operators and deployment automation.
The model is adapted from Weatherreporter's release procedure, but its target
is deliberately narrower: an immutable source tag and checked-in release note
are the release. Notarius does not publish executable archives or support
Windows as part of this work.
## Release Model
Notarius releases come from commits on `main` and use stable semantic-version
tags in the form `vMAJOR.MINOR.PATCH`. Prerelease tags are not part of the
initial process.
Every release has one nonempty, version-matched note at
`docs/releases/<tag>.md`. The note and every affected current-state document
must be present in the tagged commit. The Git tag and checked-in note together
are the durable release record; no separately editable release page is
required.
Published tags are immutable. A maintainer must never move, reuse, or delete a
published tag. If a published candidate is defective, the correction is made
on `main` and released under a new patch version. An unpublished local tag may
be deleted when candidate inspection finds a problem before any remote push.
Before `v1.0.0`, a minor release may intentionally change a documented CLI,
configuration, durable artifact, integration, or operating contract when its
release note explains the impact and required operator action. A patch release
must not intentionally break those documented contracts within its minor
line.
The existing `v0.1.0`, `v0.2.0`, and `v0.3.0` tags remain unchanged. They
predate this procedure and do not need retrospective release notes. The first
release made under this process establishes the release-note series.
## Source-Only Distribution
Notarius does not publish release binaries, archives, installers, container
images, package-manager entries, checksum files, or signatures. A release tag
is suitable for Go-native installation and for an operator-controlled build
from an exact checkout.
The primary installation form is:
```sh
GOWORK=off go install \
gitea.maximumdirect.net/eric/notarius/cmd/notarius@vMAJOR.MINOR.PATCH
```
Operator documentation should also describe cloning the repository, checking
out the tag in detached-head state, and building `./cmd/notarius` with the Go
version declared by `go.mod`. Private-module authentication and `GOPRIVATE`
configuration belong to the operator environment and must be documented by
mechanism rather than with real credentials.
Consumers such as Narratio should pin the desired Notarius tag in provisioning
or deployment configuration. They must continue to decide runtime
compatibility from Notarius's published receipt and artifact schema contracts,
not merely from the executable's product version.
Packaged binaries may be reconsidered if distribution demand, installation
friction, or a broader user audience justifies their build, signing, retention,
and platform-support costs. They are not a prerequisite for a disciplined
release process.
## Platform Policy
Linux is the supported deployment platform. Release validation must run the
test suite and the release build on Linux and must confirm that the command
builds with `CGO_ENABLED=0` for Linux `amd64` and `arm64`.
macOS is a best-effort development and testing platform. Release validation
should confirm that the command cross-compiles with `CGO_ENABLED=0` for Darwin
`amd64` and `arm64`, but the project does not promise packaged artifacts or a
separate runtime test environment for those targets.
Windows is unsupported. The release process must not require Windows builds,
Windows-specific compatibility work, or Windows documentation. Platform-
specific implementation may intentionally use Unix facilities when they are
important to Notarius's filesystem safety and operational model. Any later
decision to support Windows requires its own feature scope and validation
policy.
## Version Reporting
Add a root `notarius --version` interface for deployment diagnostics. It
prints exactly one line:
```text
notarius vMAJOR.MINOR.PATCH
```
when the build has a valid release version, and:
```text
notarius development
```
when no release version is available.
The implementation must obtain the main-module version from Go build
information so `go install ...@vMAJOR.MINOR.PATCH` reports the selected tag. It
must also accept an optional link-time version override so controlled builds
and release CI can identify an exact tag from a checkout. The override must be
validated and must not silently turn arbitrary text into a release version.
Ordinary unversioned checkout builds remain `development`; the release process
must not modify a tracked source constant for each release.
Version reporting is an informational product interface. It does not replace
receipt, configuration, prompt, or artifact schema versioning, and it must not
be used as the sole downstream compatibility check.
## Release Notes
Each new `docs/releases/<tag>.md` document has this minimum structure:
```markdown
# Notarius vMAJOR.MINOR.PATCH
This release ...
## Summary
## Compatibility
## Upgrade
## Changes
```
The note should concisely explain the release's purpose, compatibility with the
preceding release, operator actions, and material user-visible, operational,
integration, and maintainer-visible changes. It should link to canonical
current-state documentation for exact contracts rather than duplicating those
contracts.
Release notes are durable historical summaries. They must not contain
credentials, private infrastructure detail, sensitive campaign material, or
claims that are not true of the tagged candidate. A release note does not
excuse stale current-state documentation; affected canonical documents are
updated in the same candidate.
## Candidate Validation
The release procedure must provide copyable POSIX-shell guards that validate
the release version, release-note filename and heading, required note sections,
repository state, and module hygiene. Validation must be run from the Notarius
repository root with Go workspace behavior disabled.
At minimum, a candidate must pass:
- no tracked `go.work` or `go.work.sum`, no vendored tree, and no `replace`
directive in `go.mod`;
- `GOWORK=off go test -count=1 ./...`;
- `GOWORK=off go test -race -count=1 ./...`;
- `GOWORK=off go vet ./...`;
- `GOWORK=off go build ./...`;
- `GOWORK=off go mod tidy -diff`;
- `gofmt` verification for every tracked Go file;
- `git diff --check` and `git diff --cached --check`;
- validation of both maintained D&D configuration examples with their selected
pipeline;
- Linux `amd64` and `arm64` static command builds;
- best-effort Darwin `amd64` and `arm64` static command builds; and
- a focused manual or automated check that every added or changed local
Markdown link resolves.
The candidate review also checks for generated binaries, test output,
credentials, temporary files, module replacements, vendored dependencies, and
other unintended source-control content. Tests remain offline and do not call
an LLM provider or require live credentials.
## Candidate Publication
The release procedure must guard the exact commit immediately before tagging.
It requires:
- the current branch is `main`;
- the worktree and index are clean;
- the candidate commit has been pushed and exactly matches `origin/main`;
- the matching release note exists in that commit;
- no local or remote tag already uses the selected version; and
- the substantive release checks have passed for that exact candidate.
The maintainer records the exact candidate commit, creates a lightweight tag
bound explicitly to that commit, verifies the local tag target, and pushes only
that tag ref. The procedure must not recommend `git push --tags`.
After publication, the maintainer verifies that the remote tag resolves to the
guarded commit and that the release note can be read from the tagged tree. A
fresh temporary checkout or `go install ...@<tag>` must build successfully, and
the resulting command must report the expected version through `--version`.
## Validation-Only Release Automation
Add a tag-triggered Woodpecker pipeline that validates source releases without
publishing artifacts. It should:
- accept only stable semantic-version tags;
- require the version-matched release note;
- run the same substantive module, test, race, vet, build, formatting, and
whitespace checks as the documented local procedure;
- validate the maintained configuration examples;
- perform the supported and best-effort cross-build checks; and
- verify a release-version build's `notarius --version` output on the CI host.
The pipeline must not upload binaries, create archives or checksums, create or
edit a Gitea release object, or require a release API token. Local guards remain
authoritative before tag publication because CI begins only after the tag is
already remote.
If tag validation fails, preserve the published tag, fix the cause on `main`,
select a new patch version, and repeat the full process. Do not weaken tag
immutability merely because the release contains source rather than binaries.
## Documentation Ownership
In the target state:
- `docs/release.md` owns the maintainer release procedure, commands, ordering,
publication checks, and failure recovery;
- `docs/releases/` owns one historical summary per release made under the new
process;
- `docs/cli.md` owns the `--version` contract;
- `README.md` owns the shortest source-installation example and links to the
release procedure where useful;
- `docs/development.md` routes release preparation, tagging, and verification
work to `docs/release.md`;
- `docs/policy/documentation.md` assigns canonical ownership to the release
procedure and release notes;
- `docs/policy/architecture.md` records Linux support, best-effort macOS
development, unsupported Windows, and source-only distribution only if those
are judged durable development invariants rather than release mechanics; and
- `docs/operations.md` describes only installation or deployment consequences
relevant to operators and links to canonical CLI and release contracts.
Current-state documentation must not describe the new release process,
`--version`, or automated validation until the corresponding behavior exists.
## Acceptance Criteria
- A maintainer can prepare, validate, tag, publish, and verify a source release
by following `docs/release.md` without relying on undocumented knowledge.
- Every new release has an immutable semantic-version tag and matching
checked-in release note in the tagged commit.
- The guarded candidate is clean, synchronized with `origin/main`, and passes
the documented substantive checks before tagging.
- Tag-triggered CI independently validates the published source and never
publishes binary artifacts.
- `go install` of a tagged version succeeds and `notarius --version` reports
that version; ordinary unversioned builds report `development`.
- Linux is the documented supported deployment platform, macOS has a
best-effort development build check, and Windows is explicitly unsupported.
- Downstream compatibility remains based on durable Notarius contracts rather
than the product version alone.
- Existing pre-procedure tags remain untouched and require no invented release
history.
## Non-Goals
- Publishing executable archives, installers, container images, checksums,
signatures, or package-manager entries.
- Supporting or cross-compiling for Windows.
- Creating or maintaining a mutable Gitea release page.
- Supporting prerelease tag syntax in the initial procedure.
- Automating version selection, release-note authorship, commits, or tag
creation.
- Retrospectively creating release notes for `v0.1.0` through `v0.3.0`.
- Treating a product version as a substitute for receipt, configuration,
prompt, or artifact schema compatibility.

View File

@@ -0,0 +1,609 @@
# Feedback-Aware Stage Validation Retries
## Status
Proposed. This is the active feature roadmap for the next Notarius work set.
Its design decisions are settled. Current behavior remains authoritative until
this roadmap is implemented and the corresponding ADR and canonical
documentation are updated.
## Purpose
Make validation an effective corrective boundary around LLM-produced stage
candidates. When deterministic or LLM-backed validators reject a structurally
valid candidate, Notarius should give the producing model the complete,
ordered validation feedback and use the stage's existing retry budget to ask
for a corrected replacement. The feature must distinguish semantic rejection
from producer failure and validator execution failure, preserve the boundary
between PromptKit repair and Notarius stage retries, and remain safe under
concurrency, cancellation, caching, checkpoints, and sensitive input.
This work is domain-neutral. It establishes the framework behavior required by
future LLM-backed validators such as D&D combat-scene review, but it does not
add that validator.
## User Intent
- A stage candidate should be evaluated by every applicable configured
validator before Notarius decides whether to retry or terminate.
- A semantic retry should be materially more useful than repeating the same
request. The producing model should see its latest defective response and
all actionable semantic feedback.
- PromptKit's bounded structural repair and Notarius's stage retry loop are
separate. Each stage attempt receives its own complete PromptKit repair
budget; PromptKit repair never consumes or replenishes the stage budget.
- Deterministic rejection, semantic rejection, producer structural failure,
and validator execution failure are different outcomes and must not be
collapsed into one generic error path.
- The default posture is strict for known-invalid producer output and tolerant
but visible when a validator itself cannot make a decision.
- Corrective prompts must not expose opaque application identifiers, secrets,
or unbounded diagnostic content merely because those values exist in an
internal artifact or operator-facing error.
## Current State
The current runner already provides useful foundations:
- chunk, extract, merge, and normalize producer bindings have one `retries`
value interpreted as additional stage attempts;
- `runWithRetry` retries producer errors and semantic rejections within that
budget;
- PromptKit performs bounded structural repair inside each structured
completion;
- validator targets, execution classes, profile selection, repair policy,
attempts, debug scopes, checkpoint identity, and deterministic public
ordering are already explicit; and
- the structured-completion response retains the model's validated raw bytes
and PromptKit repair metadata.
The current behavior is not yet the desired corrective workflow:
- the runner repeats the ordinary producer request after rejection and does
not pass the previous model response or validator feedback;
- validation stops at the first rejection or execution failure, so later
applicable validators do not contribute findings;
- validator execution failure is immediately a framework error rather than a
configurable incomplete-validation outcome;
- `ValidationResult.Message` currently serves operator diagnostics and does
not define separately bounded model-facing guidance;
- typed stage results do not carry the exact model response needed for the
next correction attempt;
- validator-binding `retries` values participate in resolved configuration but
are not used to retry a failed validator against the same candidate; and
- Notarius still pins PromptKit v0.8, while PromptKit v0.9.0 now provides the
append-only request-message API needed for application-owned correction
attempts.
## Target End State
For chunk, extract, merge, and normalize stages, Notarius owns one explicit
candidate-attempt state machine:
1. The producer creates one candidate using the ordinary request. An
LLM-backed producer may use PromptKit structural repair internally.
2. The framework establishes one immutable validation candidate and runs every
applicable validator sequentially in configured order.
3. The framework aggregates approvals, warnings, semantic rejections,
execution failures, and skipped-validator diagnostics without allowing one
validator to mutate the candidate seen by another.
4. A candidate with one or more semantic rejections is never accepted. If the
LLM-backed producer has another stage attempt available, Notarius rebuilds
the complete original prompt and appends the latest defective assistant
response followed by one application-owned correction message containing
every actionable rejection. It then requests one complete replacement
candidate.
5. A producer error consumes the same stage attempt budget under the existing
retry rules, but semantic correction material is used only when a
structurally valid candidate was actually rejected.
6. A validator execution failure is retried, when configured, against the same
immutable candidate. It never regenerates the producer candidate by itself.
7. When budgets are exhausted, the configured terminal policies decide
whether the run fails, a rejected output is recorded, or a structurally
valid candidate advances with explicitly incomplete validation.
The first attempt remains byte-for-byte the ordinary prompt rendered from the
selected prompt definition. Every correction attempt starts from that same
ordinary prompt rather than from the prior correction conversation. It appends
exactly two messages:
- an `assistant` message containing the producer-supplied exact defective
response for the latest candidate; and
- a `user` message containing deterministic, bounded, application-owned
correction guidance and asking for one complete replacement response.
The session ID, prompt ID and version, selected profile, reasoning settings,
structured-output contract, repair budget, named inputs, variables, references,
and reusable prompt prefix remain unchanged across stage attempts.
## Architectural Ownership
### PromptKit
PromptKit continues to own prompt loading and rendering, profile resolution,
backend admission, provider generation, structural validation, and bounded
structural repair within one completion. A PromptKit repair conversation is
private to that completion and is not exposed as a Notarius stage attempt.
PromptKit v0.9.0 owns the mechanical operation of appending explicitly supplied
messages to a normally rendered prompt before creating the immutable prepared
execution. `RunRequest.AppendedMessages` preserves the original rendered
messages as an exact prefix, validates and defensively copies additions,
includes the complete sequence in prepared details and the rendered-prompt
hash, and runs it through the ordinary generation and structural-repair path.
PromptKit does not impose message-count, byte-size, token, or context-window
limits and permits empty content, so Notarius retains its stricter
application-owned correction validation and bounds.
### Notarius Framework
The framework owns stage budgets, immutable candidate preparation, complete
validator-chain execution, result aggregation, outcome precedence, correction
message construction, terminal policy, public ordering, checkpoint effects,
manifest summaries, warnings, and debug lifecycle.
The framework must remain domain-neutral. It may format stable reason codes and
validator-supplied corrective guidance, but it must not infer D&D or other
domain rules from artifact JSON.
### Producers And Artifact Families
The producing module owns prompt selection, prompt inputs, typed decoding, and
the model-facing representation that corresponds to its candidate. An
LLM-backed producer that supports feedback-aware correction must return the
exact response material that the model should see as its prior assistant turn.
It must not substitute a normalized artifact containing deterministically
attached UUIDs or other opaque application identity.
Artifact-family validators own semantic decisions and domain-specific
corrective guidance. Operator-facing explanation and model-facing correction
are separate contract fields even when their concise text happens to match.
## Validation Outcome Model
Each validator invocation produces one of four framework outcomes:
| Outcome | Meaning | Effect |
| --- | --- | --- |
| Approved | The validator completed and accepted the whole candidate. | Retain its warnings and continue the chain. |
| Rejected | The validator completed and found a semantic defect in the candidate. | Record the finding, continue the chain, and make the candidate ineligible for acceptance. |
| Failed | The validator could not return a usable decision because of an internal, transport, generation, structural-output, or result-invariant failure. | Retry that validator when eligible, then record incomplete validation and continue the chain unless cancellation or framework integrity prevents it. |
| Skipped | Runtime prerequisites for an otherwise selected validator cannot be satisfied. | Record a deterministic incomplete-validation diagnostic and continue; do not invent a semantic decision. |
Configured validator order controls invocation order and aggregate feedback
order. Execution remains sequential initially. The framework must continue
after a rejection and after an isolated validator failure when it can safely
prepare the remaining validator requests. Cancellation, inability to preserve
an immutable candidate, debug persistence failure, or another framework
integrity failure remains immediately terminal.
### Outcome Precedence
For one candidate, apply this precedence:
1. A producer structural failure means no acceptable candidate exists and
cannot be converted into validator approval.
2. Any completed semantic rejection makes the candidate rejected, even when
another validator failed or was skipped.
3. With no semantic rejection, a validator failure or skip makes validation
incomplete and invokes the validator-failure policy.
4. Only a structurally valid candidate with no rejection and either complete
validation or an explicit `warn_continue` decision may advance.
Do not turn a known rejection into acceptance through a permissive
validator-failure policy. Do not turn a structurally invalid response into a
rejected-but-usable artifact.
## Corrective Feedback Contract
`ValidationResult` should gain a separately bounded, optional model-facing
correction field. A rejecting production validator should provide:
- a stable reason code suitable for aggregation and provenance;
- an operator-facing message suitable for ordinary diagnostics; and
- concise corrective guidance that explains the violated rule without asking
the model to reproduce opaque identity or leaking unrelated source data.
The framework constructs one deterministic correction message from all
rejections in validator order. Each entry identifies the stable reason code
and corrective guidance. Duplicate identical entries may be collapsed while
preserving first occurrence; distinct findings must not be discarded merely
to shorten the message. If a validator rejects without model-facing guidance,
the framework uses a generic reason-code-based correction rather than copying
the operator message automatically.
Warnings, validator failures, skipped diagnostics, provider messages, stack
traces, debug paths, and sensitive values are not corrective guidance. They may
be recorded through their proper diagnostic channels but must not be presented
to the producer as candidate defects.
The framework must validate UTF-8, role, non-empty content, and
application-owned size limits before constructing the correction request. Oversized or
invalid correction material is a framework-owned inability to perform a
feedback retry; it must never be silently truncated into a misleading or
syntactically defective assistant response.
## Producer Correction Contracts
Introduce application-owned, defensively copied correction contracts at the
framework boundary:
- chunk, typed extraction, typed merge, and typed normalize results can carry
optional model-facing candidate material associated with their returned
value;
- the corresponding requests can carry an optional correction containing the
latest assistant material and aggregated guidance;
- `StructuredCompletionRequest` can carry the two bounded appended messages
without importing PromptKit types into module or pipeline contracts; and
- the PromptKit adapter translates those application-owned messages into
`RunRequest.AppendedMessages` using `promptkit.RoleAssistant` and
`promptkit.RoleUser` before preparation.
A semantic correction always supplies exactly two appended messages: the
latest defective response as `assistant`, followed by the aggregate correction
request as `user`. The framework does not expose the other PromptKit-supported
roles through this contract and does not accumulate messages from earlier
stage attempts. PromptKit preserves message content exactly, but Notarius must
reject empty content and enforce its own per-message and aggregate byte limits
before the adapter is called.
Correction material is attempt-local sensitive data. It is not part of the
artifact schema, checkpoint value, cache key, durable output bundle, ordinary
error, or configuration summary. The policy and capability that affect
execution do participate in resolved pipeline and checkpoint identity.
LLM-backed modules selected with both `retries > 0` and a non-empty validator
chain must declare whether they can produce and consume correction material.
Preparation must reject a pipeline that could request feedback-aware semantic
retries from an LLM-backed producer without that capability. An LLM-backed
producer with no validators may continue to use its retry budget for
operational failures without declaring semantic-correction capability.
A correction-capable producer must supply the exact single LLM response that
directly controlled the candidate being validated. Direct D&D chunk and
extraction producers expose their exact structured response. The shared
semantic-reconciliation path exposes its exact proposal response through its
typed normalizers without turning request-local batch handles into durable
identity. Deterministic transformations after that response are permitted only
when the validated candidate remains directly traceable to it.
A producer whose candidate combines multiple LLM responses is not
correction-capable under this initial protocol. It may continue to use ordinary
operational retries when no semantic correction can occur, but configuration
must reject a validator-backed retry workflow for it. Supporting compound
producers later requires a separately reviewed multi-response protocol; the
framework must not synthesize an assistant message by serializing the final
typed artifact.
Deterministic producers do not receive correction material. A deterministic
candidate rejected by validation immediately applies the terminal semantic
rejection policy without consuming retries that cannot change the result.
## Retry Budgets
### Producer Stage Budget
The existing producer binding `retries` field remains the sole outer stage
budget. `retries: N` means at most `N` additional complete producer attempts
after the initial attempt. Producer operational errors, producer structural
failures, module-requested normalize retries, and semantic corrections all
draw from this same budget. Do not add a separate semantic retry counter.
Every LLM-backed producer attempt receives the configured PromptKit
`structured_output_repair_attempts` value independently. Notarius does not
decrement that value across stage attempts.
### Validator Budget
Use the existing `retries` field on an LLM-backed validator binding for
additional attempts to obtain a usable decision about the same immutable
candidate. A completed approval or rejection is terminal for that validator
and does not consume another validator attempt. A validator retry reconstructs
the same ordinary validator prompt; it does not append semantic feedback about
the validator's prior failed judgment and does not create a recursive
Notarius correction loop.
Reject a positive validator `retries` value on a deterministic validator at
configuration resolution because repeating the same pure decision cannot
improve it. Validator retries do not consume the producer stage budget.
## PromptKit v0.9.0 Adoption
The target end state pins PromptKit v0.9.0 for correction requests. The
resolved dependency graph includes its independently versioned
OpenRouter and Rakestrawhome catalog modules through ordinary Go module
resolution; Notarius must not import or register those catalogs directly.
PromptKit continues to own their built-in backend and profile IDs, source
precedence, credentials, and capacity behavior.
Notarius's PromptKit compatibility documentation and built-in-profile
checkpoint marker identify v0.9.0 rather than v0.8.0. The PromptKit release
identity remains the conservative checkpoint identity for the exact catalog
versions selected by that release; Notarius should not duplicate upstream
catalog module versions in a second hand-maintained marker.
PromptKit v0.9.0 restricts text-chat roles to `developer`, `system`, `user`, and
`assistant`. Maintained Notarius prompt definitions already use only `system`
and `user`; correction requests add only `assistant` and `user`. PromptKit
`RunRequest` literals remain keyed. These compatibility conditions must remain
covered by the ordinary production-asset and adapter checks without adding a
brittle inventory test that merely counts prompt messages or literals.
## Terminal Policy Configuration
Add an optional `validation_policy` object at pipeline scope and on chunk,
extract, merge, and normalize producer bindings:
```yaml
validation_policy:
producer_structural_failure: fail_run
semantic_rejection: fail_run
validator_failure: warn_continue
```
The binding object overrides individual pipeline values; resolution is
field-by-field in binding, pipeline, application-default order. Omitted values
inherit rather than replacing the complete object. Explicit null, unknown
fields, and unknown enum values are invalid. The effective policy is resolved
and detached before execution, appears in redacted effective configuration and
run provenance, and participates in the resolved pipeline digest and checkpoint
identity.
The initial enum values and defaults are:
- `producer_structural_failure`: `fail_run` by default; `reject_output` may
retain a terminal rejection and final raw candidate for debug, but may not
advance or publish an invalid artifact;
- `semantic_rejection`: `fail_run` by default after stage attempts are
exhausted; `reject_output` records the aggregate rejection and allows
unrelated work to complete without advancing that candidate; and
- `validator_failure`: `warn_continue` by default, which advances a
structurally valid and otherwise unrejected candidate with explicit
incomplete-validation provenance and one genuine warning; `fail_run`
terminates the run.
Producer structural policy applies only to LLM-backed producers. Semantic and
validator-failure policies apply to any validated producer. Input and output
bindings do not accept `validation_policy`, and validator bindings do not own
terminal policy; they own only their decision and their own operational retry
budget. Candidate disposition belongs to the chunk, extract, merge, or
normalize producer binding after its complete validator chain has run.
Keep the current file-configuration version. The syntax is strictly
decodable without a version change. The project is pre-v1, but the behavior
and output changes should still be called out in the next release note and
downstream documentation.
## Stage-Specific Behavior
### Chunk
A generated chunk plan is structurally validated and materialized before the
validator chain runs. Semantic feedback applies to the exact raw chunker
response associated with that plan.
When an automatically reused chunk-plan record is rejected by the current
validator chain, treat the record as unusable for this invocation and enter
ordinary generation at attempt one. A cache hit is not a new model attempt and
does not supply model-facing assistant material. Do not overwrite the cached
record until a newly generated plan is accepted. Refresh and bypass modes
retain their existing publication rules.
### Extract
Each chunk-scoped extraction job owns its own attempt state and correction
conversation. One rejected chunk candidate does not cancel unrelated chunks or
lanes unless terminal policy converts it into a framework error. Deterministic
public ordering remains chunk-first and lane-second regardless of concurrent
completion.
### Merge And Normalize
Merge and normalize remain serial within a lane. A correction attempt receives
the same accepted upstream artifacts and references as the initial attempt.
The existing safe-fallback `NormalizeRetry` mechanism must be reconciled with
the shared attempt state rather than layered into a second retry loop: it uses
the same stage budget, retains its documented fallback behavior, and cannot
override a known validator rejection.
The initial feature supports one exact producer-supplied assistant response per
candidate attempt. Future multi-request normalization or batching must define
which response directly represents the candidate, or supply a new explicitly
reviewed correction protocol, before it can claim feedback-aware correction.
## LLM-Backed Validators
An LLM-backed validator uses the same scheduled PromptKit client, selected
profile, session, timeout, and structural-repair policy as other LLM-backed
modules. PromptKit may structurally repair its response inside one validator
attempt.
- A contract-valid validator response is its decision; Notarius does not ask a
second LLM to judge that judgment.
- A structurally invalid final validator response, transport failure, or
deterministic violation of the validator-result contract is a validator
execution failure.
- Validator execution retries reuse the immutable producer candidate and do
not regenerate it.
- Exhaustion invokes `validator_failure` policy and emits a genuine warning
under `warn_continue`.
This feature supplies the generic execution model only. It does not register a
production LLM-backed validator or change a D&D default validator chain.
## Provenance, Diagnostics, And Sensitive Data
Attempt debug output should make the state machine auditable. When debug is
enabled, record:
- producer attempt number and whether it was initial, error retry, module
retry, or semantic correction;
- PromptKit repair count and cumulative usage for every completion;
- each validator's configured-order outcome and validator attempt count;
- aggregate rejection codes and the bounded correction message;
- effective terminal policy and the decision it produced; and
- whether validation was complete, rejected, or incomplete.
Raw assistant responses and correction messages belong only in explicitly
requested detailed debug traces, following existing allowlisted content-file,
redaction, permission, and retention rules. Ordinary errors, CLI output,
warnings, manifests, checkpoints, caches, and run receipts contain identities,
counts, bounded safe summaries, and reason codes—not raw source or model
content.
The durable run manifest and rejection summaries should record enough
structured information to distinguish:
- the number and kinds of producer attempts;
- completed semantic rejection and all rejecting validator identities;
- incomplete validation and failed or skipped validator identities;
- the effective terminal policy and terminal result; and
- successful use of a correction attempt without treating it as a warning.
Warnings from abandoned producer attempts must not be promoted. Warnings from
the accepted attempt remain eligible. A warn-and-continue validator failure
produces one bounded, deterministically ordered warning per affected validator
after its retry budget is exhausted; detailed repeated failures stay in debug
provenance.
## Checkpoints, Caches, Concurrency, And Cancellation
- Effective validation policy, producer correction capability/protocol
version, validator chain, validator retry budgets, and prompt assets must all
affect checkpoint identity.
- Only accepted, completely validated stage outputs may be checkpointed or
reused. Rejected, structurally invalid, and validation-incomplete outputs
accepted under a permissive policy must not be written as reusable stage
checkpoints. This conservative rule avoids treating a transient validator
outage as durable validation success; a future checkpoint-status contract may
revisit it explicitly.
- Correction attempts use the same run-wide scheduler and worker bounds as
initial completions. No retry path may bypass provider admission.
- A scheduled permit covers the complete PromptKit operation, including its
internal structural repair, and is reacquired normally for a later Notarius
stage attempt.
- Parent cancellation dominates producer, validator, retry, debug, cache, and
checkpoint work. Cancellation never becomes a rejection, warning, or
incomplete-validation acceptance.
- Framework errors retain deterministic selection and cancellation behavior
across concurrently executing chunks and lanes.
## Architecture Record And Canonical Documentation
The target end state includes an accepted ADR that records:
- the separation between PromptKit structural repair, producer stage attempts,
and validator execution retries;
- the complete validator-chain aggregation rule and outcome precedence;
- the fresh reconstruction plus two-message correction protocol;
- module ownership of model-facing candidate material;
- deterministic producer and non-recursive validator behavior;
- default fail-closed and fail-open terminal policies; and
- provenance, cache, identity, and sensitive-data constraints.
The canonical owners describe the implemented behavior without duplicating
one another:
- `docs/policy/architecture.md` for durable validation and retry invariants;
- `docs/config.md` for fields, values, precedence, defaults, and validation;
- `docs/operations.md` for costs, failure behavior, warnings, debug handling,
and recovery;
- `docs/internal/pipeline.md` for the attempt state machine, aggregation,
checkpoint behavior, and concurrency;
- `docs/internal/llm.md` for appended correction messages and the distinction
from PromptKit repair;
- `docs/internal/modules.md` for producer and validator contracts;
- `docs/integrations/pkg-promptkit.md` for PromptKit v0.9.0,
`RunRequest.AppendedMessages`, supported message roles, application-owned
bounds, and the independently versioned upstream catalog boundary; and
- affected output and subprocess integration documents for durable validation
status and rejection summaries.
Until this target state is implemented, canonical current-state documentation
continues to describe the existing behavior.
## Testing Strategy
Tests should protect observable state-machine behavior rather than private
helper layout or exact prose. The target test suite includes:
- contract tests proving the first request is unchanged and a correction
request contains the same initial messages plus exactly one assistant and one
user message;
- behavioral runner tests for all-approved, multiple-rejection,
rejection-plus-failure, failure-only, skipped, retry-success, and each
terminal policy outcome;
- one representative path for chunk, extract, merge, and normalize, without
duplicating the complete state matrix at every stage;
- proof that all validators see immutable equivalent candidates and execute in
configured order after an earlier rejection or isolated failure;
- proof that validator retries reuse the candidate and do not consume producer
retries;
- proof that deterministic rejection does not repeat the producer;
- focused PromptKit-adapter tests proving that application-owned correction
messages map to the two intended PromptKit roles without content leakage;
- config parsing, precedence, invalid-placement, round-trip, redaction, and
digest tests for effective policy;
- checkpoint and chunk-cache tests for rejected, corrected, incomplete, and
accepted outcomes;
- warning, manifest, receipt, debug, and sensitive-content tests at their
canonical boundaries; and
- a small assembled D&D pipeline test proving a rejected direct extraction can
be corrected without a live provider.
Tests remain offline and deterministic. Do not reproduce PromptKit's internal
message-copying, rendered-hash, prepared-execution, capacity, or repair suite.
One representative adapter or assembled-run test should prove that PromptKit
structural repair remains usable after Notarius appends semantic-correction
messages. Do not snapshot full prompts or error prose, assert private
constants, or multiply equivalent tests across every D&D artifact family.
## Acceptance Criteria
- Every applicable validator runs in configured order and contributes one
explicit outcome before candidate disposition.
- Multiple semantic rejections produce one bounded, deterministic correction
request containing all actionable findings.
- Correction attempts reconstruct the exact ordinary prompt and append only
the latest defective assistant response and one correction message.
- Notarius pins PromptKit v0.9.0 and routes correction messages through
`RunRequest.AppendedMessages`; it does not maintain paired correction prompt
manifests or bypass PromptKit's normal execution path.
- Every correction-capable LLM producer exposes the exact single response that
directly controlled its candidate. Configuration rejects semantic retries
for compound producers that cannot satisfy that contract.
- The existing producer `retries` value is the only producer-stage budget;
PromptKit structural repair and validator execution retries remain separate.
- Deterministic producers are not repeated after semantic rejection.
- Validator execution failure is never described to the producer as a
candidate defect and never creates recursive semantic validation.
- Default terminal behavior is `fail_run` for structural failure and semantic
rejection, and `warn_continue` with explicit incomplete validation for
validator failure.
- Terminal policy resolves field by field from producer-binding override to
pipeline default to application default; individual validators do not own
candidate disposition.
- Permissive policy never advances known rejected or structurally invalid
output.
- Raw model responses and correction content are confined to model requests and
explicitly requested debug traces.
- Checkpoint, cache, manifest, warning, concurrency, cancellation, and
deterministic-ordering invariants remain intact.
- Canonical architecture, configuration, operations, internal, integration,
and ADR documentation accurately describe the implemented behavior.
- Focused, full, and race-enabled Go tests; vet; builds; example validation;
formatting; link checks; and repository hygiene checks pass.
## Non-Goals
- Adding the D&D combat-scene semantic validator.
- Redesigning the warning taxonomy beyond the warnings required for validator
failure and retry outcomes.
- Concurrent validator execution.
- Unbounded or accumulating conversational history.
- A second semantic retry counter.
- Recursive LLM judgment of LLM-validator decisions.
- Provider-specific retry policy or bypassing PromptKit.
- General workflow graphs or new pipeline stages.
- Large-collection reconciliation batching or a generic multi-response
correction protocol.