From 61436d7c186af57b06004e090ceac898e83cdb71 Mon Sep 17 00:00:00 2001 From: Eric Rakestraw Date: Thu, 27 Aug 2026 19:05:13 +0000 Subject: [PATCH] Prepare the v0.5.0 release --- docs/internal/state.md | 5 +- docs/operations.md | 9 +- docs/releases/v0.5.0.md | 85 ++ .../warning-signal-and-presentation-audit.md} | 0 docs/roadmap/audit-plan.md | 399 ---------- docs/roadmap/future.md | 32 +- docs/roadmap/implementation.md | 742 ------------------ 7 files changed, 96 insertions(+), 1176 deletions(-) create mode 100644 docs/releases/v0.5.0.md rename docs/roadmap/{audit.md => archive/warning-signal-and-presentation-audit.md} (100%) delete mode 100644 docs/roadmap/audit-plan.md delete mode 100644 docs/roadmap/implementation.md diff --git a/docs/internal/state.md b/docs/internal/state.md index 92bf7037..6be86bad 100644 --- a/docs/internal/state.md +++ b/docs/internal/state.md @@ -68,8 +68,9 @@ workspace schema v4 plus an exact non-empty invocation identity, and deliberatel does not require extract or merge checkpoint files or dependency fingerprints. The runner performs canonical codec and producer-provenance validation before cloning the artifact into normal step output. Success restores only stored -normalize warnings and emits one normalize decision; failure retains the files, -records the decision, and stops without executing the producer or consumer. +normalize diagnostics and emits one normalize decision; failure retains the +files, records the decision, and stops without executing the producer or +consumer. The loader assigns a typed category and reason code at each validation site; diagnostic prose is not classified after the fact. The runner then applies diff --git a/docs/operations.md b/docs/operations.md index 2f212431..e143e80f 100644 --- a/docs/operations.md +++ b/docs/operations.md @@ -129,13 +129,14 @@ validator execution failure normally uses `warn_continue`, which keeps an otherwise accepted result in the current run with `incomplete` validation provenance. It emits one bounded warning for every validator whose execution budget was exhausted. A corrected result that later passes validation does not -retain abandoned-attempt warnings. +retain diagnostics from abandoned attempts. Treat a successful process exit as a completed run, not as proof that every candidate was fully validated. Inspect the receipt's `validation_status`, -`validation_summaries`, rejection count, and warning count when an orchestrator -requires complete validation. The durable fields and their meanings are owned -by the [run-result receipt](integrations/run-result.md) and +`validation_summaries`, rejection count, and warning group and occurrence +counts when an orchestrator requires complete validation. The durable fields +and their meanings are owned by the +[run-result receipt](integrations/run-result.md) and [published JSON output contract](integrations/json-output.md). ## Output Bundles diff --git a/docs/releases/v0.5.0.md b/docs/releases/v0.5.0.md new file mode 100644 index 00000000..0174ec48 --- /dev/null +++ b/docs/releases/v0.5.0.md @@ -0,0 +1,85 @@ +# Notarius v0.5.0 + +This release separates actionable process warnings from extraction-quality +advisories and routine normalization observations, giving operators a quiet +warning channel without discarding durable diagnostic detail. + +## Summary + +Notarius now carries one validated, origin-aware diagnostic contract from +producers and validators through retries, reusable state, output publication, +debug summaries, run receipts, and CLI presentation. Warnings are reserved for +process degradation or incomplete configured work. Data-quality findings are +advisories, and successful deterministic cleanup is recorded as observations. +An ordinary successful run therefore reports zero warnings while retaining +bounded diagnostic provenance for later review. + +The framework aggregates findings deterministically by their stable identity +and complete pipeline origin, preserves exact occurrence counts, and retains +bounded representative samples. Warning groups fail rather than truncate; +advisory and observation representation is bounded with explicit truncation +metadata and exact unrepresented-occurrence counts. + +## Compatibility + +- `warnings.json` now uses the incompatible grouped + `notarius.warnings.v2` envelope and contains process warnings only. Consumers + of the former flat warning payload must migrate to the current + [JSON output contract](../integrations/json-output.md). +- The new `diagnostics.json` file uses `notarius.diagnostics.v1` and contains + advisory and observation groups. Production `index.json` files always expose + both `warnings_file` and `diagnostics_file`. +- The machine-readable run receipt is now `notarius.run-result.v2`. It replaces + `warning_count` with exact warning group and occurrence counts and adds + advisory/observation group, occurrence, and truncation fields. See the + current [run-result receipt](../integrations/run-result.md). +- Custom output modules must return their complete logical file set or an + error. The former `OutputResult.Warnings` field has been removed; an output + module cannot report a warning after serializing its output. +- Reusable state now uses `notarius.workspace.v4` and chunk-plan records use + `notarius.chunk-plan.v3` so they can preserve structured diagnostics. Older + pre-release reusable state is not reused under these contracts; start with + clean state when deterministic continuity with an older workspace is not + required. +- Validation acceptance, semantic retry budgets, rejection policy, and D&D + artifact schema identities are unchanged by this release. + +## Upgrade + +1. Update subprocess consumers to require `notarius.run-result.v2` and read + `warning_group_count`, `warning_occurrence_count`, + `diagnostic_group_count`, `diagnostic_occurrence_count`, and + `diagnostics_truncated`. +2. Update output-bundle consumers to decode `notarius.warnings.v2`, discover + `diagnostics.json` through `index.json`, and treat diagnostics as review + information rather than process warnings. +3. Update any custom output module for the removal of + `OutputResult.Warnings`; return an error when encoding cannot complete. +4. Clear pre-release reusable state before the first upgraded production run + when deterministic continuity with an older workspace is not required. +5. Run `notarius config validate --config --pipeline ` and perform + one representative run before promoting the release in an automated + pipeline. + +## Changes + +- Added validated diagnostic dispositions, categories, origins, stable reason + codes, exact occurrence counts, and bounded representative samples. +- Added deterministic run-level aggregation with separate limits for + actionable warning groups and advisory/observation groups. +- Reclassified D&D source-relatedness and unresolved-identity findings as + data-quality advisories and routine normalization changes as observations. +- Preserved structured diagnostics across producer retries, validation, + generated-reference handoff, checkpoints, chunk-plan reuse, and debug + summaries while discarding superseded-attempt findings. +- Added grouped `warnings.json`, a new grouped `diagnostics.json`, and the + corresponding production index entries. +- Upgraded the machine-readable run receipt and human CLI summary to report + exact warning and diagnostic counts without allowing advisory volume to + create warning output. +- Removed post-encoding output warnings and hardened diagnostic validation, + overflow handling, aggregate memory bounds, and warning-file path + presentation. +- Documented diagnostic ownership, classification, operator interpretation, + durable contracts, and architectural invariants in ADR-0015 and the + canonical CLI, operations, integration, and internal documentation. diff --git a/docs/roadmap/audit.md b/docs/roadmap/archive/warning-signal-and-presentation-audit.md similarity index 100% rename from docs/roadmap/audit.md rename to docs/roadmap/archive/warning-signal-and-presentation-audit.md diff --git a/docs/roadmap/audit-plan.md b/docs/roadmap/audit-plan.md deleted file mode 100644 index bc609e6b..00000000 --- a/docs/roadmap/audit-plan.md +++ /dev/null @@ -1,399 +0,0 @@ -# Warning Signal And Presentation Audit Plan - -## Purpose - -This document defines the audit required before redesigning Notarius warning -semantics and presentation. The audit must explain why an ordinary successful -D&D run can produce dozens of warnings, distinguish actionable degradation -from routine diagnostic observations, and recommend a bounded, deterministic -warning contract that operators will actually review. - -This is an audit plan, not a feature design or implementation plan. The audit -must first establish the current producers, propagation behavior, empirical -volume, public contracts, and compatibility constraints. The completed audit -selects the target behavior from those findings; `implementation.md` owns the -subsequent implementation sequence. - -## Context - -Notarius currently uses one `contracts.Warning` shape—optional `scope`, required -`reason_code`, and required `message`—for findings from modules, validators, -normalizers, reference preparation, retries, and framework policy. Accepted -warnings are collected by the pipeline, counted by the CLI and machine-readable -run receipt, and published as a flat array in `warnings.json`. - -The architecture already guarantees deterministic public ordering and bounded -warning production at several local boundaries. Recent validation work also -gave incomplete validation explicit provenance. Those guarantees must be -preserved. The observed problem is signal quality and aggregate volume: a -successful run can report many warnings even when most describe routine, -expected cleanup or weak heuristic observations rather than conditions that -require operator action. - -Initial repository inspection identifies several likely warning families: - -- D&D source-relatedness validators can produce one warning per artifact - record when a contextual name is not found in its cited text; -- deterministic normalizers report canonicalization, whitespace cleanup, - ordering changes, source-reference cleanup, and duplicate consolidation; -- semantic registry reconciliation reports discarded proposals, exhausted - fallback, and duplicate consolidation; -- optional reference materialization can report unavailable or skipped - reference inputs; -- the producer-attempt state machine reports validators that exhausted their - execution budgets under `warn_continue`; and -- producer modules can return their own warnings, which are promoted only from - the terminal accepted or rejected attempt. - -These are audit hypotheses. The audit must verify exact behavior and frequency -from current code and representative runs rather than assume that every family -is noisy or incorrectly classified. - -## Policy And Architecture Constraints - -The audit and its recommendations must follow: - -- `docs/policy/architecture.md` for module/framework ownership, deterministic - ordering, validation semantics, durable provenance, and sensitive-data - boundaries; -- `docs/policy/documentation.md` for canonical documentation ownership and the - distinction between current and future behavior; -- `docs/policy/testing.md` for risk-based, behavior-oriented test - recommendations; and -- ADR-0014 for the distinction between semantic rejection, validator failure, - correction guidance, and durable diagnostic provenance. - -In particular: - -- modules and validators may report findings, but the framework and output - boundaries own aggregation and presentation; -- reducing warning volume must not silently change artifact acceptance, - validation policy, retry behavior, rejection behavior, or process exit - status; -- genuine validator execution failures and validation-incomplete outcomes must - remain visible and auditable; -- public ordering must remain deterministic despite concurrent lane execution; -- aggregation must be bounded and must not leak raw model responses, - correction guidance, credentials, private reference content, or unnecessary - transcript text; and -- detailed diagnostics may move to a more appropriate durable or debug surface, - but must not be discarded when they are needed to understand a lossy or - degraded result. - -## Audit Objectives - -The audit must answer five questions. - -1. Which code paths produce warnings, under what conditions, and with what - expected multiplicity? -2. How are warnings cloned, bounded, replayed, ordered, promoted, persisted, - counted, and presented from their producer through the final CLI and output - bundle? -3. Which warnings identify actionable degradation or data-quality risk, and - which merely describe routine successful transformations or advisory - heuristics? -4. What aggregation and presentation model would materially reduce operator - noise without hiding genuine failures, lossy fallback, or incomplete - validation? -5. Which proposed changes affect only presentation, and which would alter a - public receipt, durable output contract, validation invariant, or other - compatibility boundary? - -## Scope - -### Warning Producers - -Inventory every production warning constructor and direct warning literal -under `internal/`. Include warnings originating from: - -- input, chunk, extract, merge, normalize, and output modules; -- typed, chunk, and serialized validators; -- semantic reconciliation and normalizer fallback; -- external and generated reference materialization; -- chunk-plan and checkpoint reuse or fallback; -- producer-attempt exhaustion and validation-incomplete continuation; -- pipeline orchestration and cancellation handling; and -- CLI or publication logic, if it creates warnings rather than only presenting - them. - -Do not infer completeness from one textual search. Inspect shared constructors, -returned result types, reason-code constants, registration paths, and tests that -exercise warning behavior. - -### Warning Propagation - -Trace each warning family through: - -- stage result and validation result contracts; -- producer attempts, including abandoned attempts, semantic retries, structural - retries, terminal rejection, and `warn_continue`; -- concurrent lane collection and stable public ordering; -- merge and normalize continuation; -- chunk-plan caching and checkpoint recording or hydration; -- ordered step output merging and generated-reference handoff; -- `RunOutput`, run manifests, validation summaries, debug records, and CLI - result construction; -- `warnings.json`, `manifest.json`, the run-result receipt, human-readable - standard error, and debug bundles. - -For every boundary, determine whether warnings are copied, filtered, -deduplicated, bounded, summarized, replayed from reusable state, or dropped. -Pay particular attention to amplification across chunks, lanes, validators, -retries, and resumed runs. - -### D&D Warning Semantics - -Review every implemented D&D artifact family. At minimum, distinguish: - -- evidence/source-relatedness advisories; -- deterministic canonicalization or whitespace changes; -- order and source-reference normalization; -- exact and semantic duplicate consolidation; -- unresolved catalog or registry membership; -- semantic-reconciliation proposal failure or fallback; and -- validator execution incompleteness. - -Determine whether the reason-code vocabulary is consistent across artifact -families, whether scopes are sufficiently contextual, and whether equivalent -conditions produce near-duplicate warnings with different codes or prose. - -### Public And Operator Surfaces - -Review the implemented contracts and documentation for: - -- CLI success output and warning-count output; -- `notarius.run-result.v1` and its `warning_count` and validation fields; -- `warnings.json` and the published JSON index; -- manifest validation summaries and rejection summaries; -- debug summary and detailed debug artifacts; and -- downstream subprocess guidance, especially the complete D&D consumer - workflow. - -Identify which surfaces are intended for immediate operator attention, durable -machine consumption, forensic detail, or debugging. Record any places where -the same flat count or warning list is being asked to serve incompatible -audiences. - -### Tests And Documentation - -Inventory tests that protect warning production, bounds, ordering, retry -promotion, checkpoint replay, JSON publication, receipt counts, and CLI stream -behavior. Identify meaningful gaps, redundant exact-prose assertions, and tests -that would unnecessarily obstruct a taxonomy or aggregation redesign. - -Review `docs/cli.md`, `docs/config.md`, `docs/operations.md`, -`docs/integrations/json-output.md`, `docs/integrations/run-result.md`, -`docs/consumers/`, and relevant internal documentation for current warning -claims. Record the canonical document that would own each future contract -change; do not rewrite those documents during the audit. - -## Out Of Scope - -The audit must not: - -- implement warning filtering, severity levels, aggregation, or new CLI flags; -- change validator decisions, default chains, retry budgets, or terminal - validation policy; -- suppress warnings merely to meet a numerical target; -- redesign rejections, errors, logs, metrics, or debug bundles except where - their boundary with warnings must be clarified; -- add the planned D&D combat-scene semantic validator; -- create deterministic tests that assert one exact global warning count for - all future runs; or -- use live paid LLM calls as part of the default automated test suite. - -## Inventory Method - -Create a warning inventory with one row per semantically distinct production -condition. Each row should record: - -| Field | Required analysis | -| --- | --- | -| Producer | Package, function, module or validator key, stage, and artifact family. | -| Trigger | Exact condition that emits the warning and whether it follows successful mutation, heuristic doubt, fallback, or failure. | -| Identity | Reason code, scope format, and whether either is stable enough for aggregation or machine use. | -| Multiplicity | Maximum per record, chunk, lane, validator, attempt, step, and run. | -| Lifecycle | Whether abandoned attempts are discarded, terminal warnings promoted, and cached or checkpointed warnings replayed. | -| Consequence | Whether data changed, evidence is questionable, output is incomplete, fallback occurred, or no externally meaningful consequence exists. | -| Actionability | What an operator can reasonably do in response, if anything. | -| Surfaces | CLI, receipt count, `warnings.json`, manifest, rejection, checkpoint, or debug presence. | -| Bounds | Existing local caps, message limits, omission records, and any missing aggregate bound. | -| Sensitivity | Whether scope or message can contain source-derived or otherwise sensitive content. | -| Coverage | Existing tests and the meaningful regression risk they protect. | - -Treat different reason codes that represent the same operator condition as -potential consolidation candidates, but do not merge them in the audit -document without explaining lost diagnostic information. - -## Propagation Analysis - -Produce a compact propagation map from warning creation to each terminal -surface. The analysis must explicitly verify: - -- only the terminal candidate's warnings are promoted after retries; -- whether rejected candidates retain warnings and where; -- whether `warn_continue` creates one warning per exhausted validator and how - that relates to validation summaries; -- whether cached chunk plans or hydrated checkpoints replay historical warnings - into the current run; -- whether the same warning can be appended at more than one handoff boundary; -- how concurrent completion is reordered before publication; -- whether local per-module caps compose into an unbounded or excessively large - run-level result; and -- whether warning counts on stderr, receipts, manifests, and `warnings.json` - refer to exactly the same collection. - -Any suspected duplicate append or unstable ordering is a correctness finding, -not merely a presentation concern. - -## Empirical D&D Run Analysis - -Static inspection must be supplemented with representative run evidence. Use -at least: - -- one ordinary successful complete D&D run known to produce high warning - volume; -- one smaller maintained example or synthetic run; -- one run with semantic registry reconciliation activity; -- one run with a validator execution failure allowed through - `warn_continue`, using an offline test double where practical; and -- one retrying producer case to verify abandoned-attempt warning treatment. - -For real campaign runs, analyze only bounded metadata unless the operator -explicitly provides source content for review. Record counts grouped by stage, -lane, module or validator, reason code, and scope family. Also record unique -reason-code count, repeated-message count, maximum group size, validation -status, rejected-output count, and whether each warning led to a plausible -operator action. - -Compare the same logical run under ordinary execution and checkpoint resume -when practical. The audit must distinguish warning volume caused by actual data -conditions from volume caused by orchestration or replay. - -Do not make a model-quality judgment solely from warning frequency. Manually -inspect a bounded sample from each high-volume reason code to estimate false -positive rate and operational value. - -## Classification Rubric - -Classify each warning condition along independent dimensions rather than force -an immediate single severity enum: - -- **result impact:** none, routine mutation, lossy mutation, uncertain data - quality, fallback, or incomplete validation; -- **operator action:** none, informational review, configuration or reference - correction, source/model review, or rerun required; -- **scope:** record, chunk, lane, step, pipeline, or infrastructure; -- **persistence need:** top-level attention, durable detail, debug-only detail, - or metric/trace candidate; and -- **confidence:** deterministic fact, heuristic advisory, or execution failure. - -The audit should then test whether a small durable taxonomy can represent the -meaningful combinations. A promising starting hypothesis is that top-level -operator warnings should be limited to actionable degradation, incomplete -validation, lossy fallback, and material data-quality risk, while routine -successful normalization observations remain available as lower-level durable -diagnostics. The audit must validate or revise that hypothesis from evidence. - -## Design Questions The Audit Must Resolve - -The findings must give a recommendation, with at least one viable alternative -and tradeoffs, for each of these questions: - -1. Should warning severity or disposition become an explicit contract field, - or should stable reason-code metadata drive presentation policy outside the - warning payload? -2. Should `warnings.json` remain the complete durable detail while the CLI and - receipt expose an aggregated actionable summary, or should durable warnings - themselves be separated from routine observations? -3. Where should global deduplication and aggregation live so modules retain - semantic ownership but concurrent pipeline results remain deterministic? -4. What is the stable aggregation key: reason code, stage/lane/module identity, - normalized scope, message template, or an explicit structured grouping key? -5. How should bounded samples and omitted counts be represented without - converting a summary record into another warning that inflates the count? -6. Should `warning_count` continue to mean the length of `warnings.json`, or - should a new receipt or schema field distinguish actionable warning groups - from detailed observations? -7. Which normalization changes are sufficiently lossy or surprising to remain - warnings, and which are ordinary provenance that belongs in manifest or - debug data? -8. Should heuristic source-relatedness findings remain warnings, become - grouped data-quality advisories, or be strengthened into configurable - validation decisions only after demonstrated precision? -9. How should warnings loaded from checkpoints be identified or aggregated - relative to newly produced warnings? -10. Does the chosen target alter durable or validation semantics enough to - require a new ADR or a versioned run-result/output contract? - -## Required Audit Deliverable - -Write the completed findings to `docs/roadmap/audit.md`. It should contain: - -1. an executive assessment of current warning quality and risk; -2. the complete warning-producer inventory; -3. the warning propagation and surface map; -4. empirical measurements and bounded representative samples; -5. findings ranked by operator impact, correctness risk, and implementation - leverage; -6. a recommended target taxonomy and presentation model; -7. compatibility, documentation, ADR, and migration implications; -8. implementation implications and dependencies sufficient to support a - separate implementation plan; and -9. open decisions only where repository evidence cannot support a responsible - recommendation. - -Each finding should identify the supporting code paths, tests, documentation, -and run evidence. Separate observed facts from recommendations and avoid -changelog or development-history framing. - -## High-Level Audit Sequence - -The detailed execution sequence should be written separately if needed. At a -high level, perform the audit in this order: - -1. **Static inventory:** enumerate warning producers, reason codes, scopes, - bounds, and existing tests. -2. **Propagation audit:** trace promotion, ordering, replay, persistence, - counting, and presentation across the framework and CLI. -3. **Empirical analysis:** measure representative D&D runs and inspect bounded - samples from high-volume warning groups. -4. **Classification:** apply the rubric, identify duplicate concepts and - misplaced routine diagnostics, and evaluate public-contract options. -5. **Synthesis:** rank findings and recommend a decision-complete target for a - subsequent feature roadmap. - -Static inventory and propagation may be performed as separate focused agent -prompts. Empirical analysis should be isolated because it may require operator -artifacts or opt-in provider execution. Classification and synthesis should -consume the earlier written evidence rather than rediscover the repository. - -## Validation Of The Audit - -Before considering the audit complete, verify that: - -- every production warning literal or constructor is represented in the - inventory; -- every reason code observed in representative `warnings.json` files maps to a - known producer or is recorded as an unexplained finding; -- counts agree across the runner result, CLI receipt, stderr summary, and - published warning collection for each sampled run; -- retry, rejection, incomplete-validation, cache, checkpoint, concurrency, and - ordered-step paths are covered; -- recommended aggregation preserves deterministic ordering and bounded memory; -- recommendations distinguish warnings from errors, rejections, validation - summaries, logs, and debug diagnostics; -- no recommendation hides a condition that changes output completeness or - correctness; -- public compatibility and schema-version consequences are explicit; and -- proposed tests protect meaningful behavior without asserting incidental - prose or one permanently fixed global warning count. - -## Completion Criteria - -The audit is ready to become a feature roadmap when it can explain the current -high warning count quantitatively, identify the dominant producers and any -amplification defects, classify every warning family by consequence and -actionability, and recommend where each class should appear. The findings must -be specific enough that a later roadmap can define the target contract without -repeating the discovery work. diff --git a/docs/roadmap/future.md b/docs/roadmap/future.md index d3d67ced..f987b450 100644 --- a/docs/roadmap/future.md +++ b/docs/roadmap/future.md @@ -10,8 +10,9 @@ not as committed release dates. PromptKit now owns structural output repair within one completion. Notarius owns stage candidates, validator chains, semantic rejection policy, bounded feedback-aware stage retries, validation provenance, and reusable-state -eligibility. The remaining near-term work applies those completed foundations -to domain review and operator-facing diagnostics. +eligibility, and the separation of actionable process warnings from quality +diagnostics. The remaining near-term work applies those completed foundations +to domain review and empirical evaluation. ### D&D Combat Scene Semantic Validation @@ -55,33 +56,6 @@ to domain review and operator-facing diagnostics. the chunker or otherwise changes stage ownership or the durable chunk-plan contract. -### Warning Signal And Presentation Reform - -- Audit every warning producer and representative successful runs. Ordinary - success producing dozens of warnings is a failed operator experience: the - volume obscures actionable problems and trains operators to ignore the - warning channel. -- Define a small warning taxonomy that distinguishes actionable degradation, - incomplete validation, lossy fallback, and data-quality risk from routine - normalization observations or informational diagnostics. Preserve detailed - traceability in debug or manifest data without promoting every observation - to a top-level CLI warning. -- Consider stable deduplication and aggregation by scope and reason code, - bounded samples plus omitted counts, and a concise CLI summary with a path to - detailed diagnostics. Do not suppress genuine validator execution failures - merely to reduce the count. -- Decide which warnings affect process status, rejection summaries, durable run - receipts, or only debug output. Ensure warning ordering and aggregation are - deterministic across concurrent execution. -- Establish a representative warning-volume acceptance target and human review - workflow before changing individual producers piecemeal. The intended result - is not zero warnings; it is a small set in which every surfaced warning merits - operator attention. -- This work does not require an ADR unless it changes validation acceptance, - failure, or durable contract semantics. CLI presentation and diagnostic - taxonomy otherwise belong in a feature roadmap followed by updates to their - canonical configuration, operations, integration, and internal documents. - ## Near-Term D&D Pipeline ### Evaluate Spell Extraction And Normalization diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md deleted file mode 100644 index 3c9a203a..00000000 --- a/docs/roadmap/implementation.md +++ /dev/null @@ -1,742 +0,0 @@ -# Warning Signal And Presentation Implementation Plan - -## Purpose - -This document is the ordered implementation plan for the warning and diagnostic -reform defined by [the completed audit](audit.md). It is intended to be executed -stage by stage by a `gpt-5.6-terra` coding agent. Each numbered stage is one -implementation prompt and must leave the repository compiling, internally -coherent, and covered at the narrowest durable test boundaries relevant to that -stage. - -The audit is the canonical source for findings, evidence, and target rationale. -This document is the canonical source for implementation order and task -breakdown. - -## Governing Decisions - -The following decisions are final for this work set. - -1. **Warnings are process signals.** A warning means that the run completed - under policy despite process-level degradation or incomplete configured - work. Examples are an empty configured reference, an applicable validator - that could not complete under `warn_continue`, an unavailable required - upstream classification, or exhausted semantic-reconciliation fallback. -2. **Extraction-quality signals are not warnings.** LLM-judged uncertainty, - lexical source-relatedness findings, unresolved entity grounding, and other - accepted-artifact quality signals are advisories. They must never be - promoted to warnings merely because a model or heuristic expressed doubt. -3. **Routine successful transformations are observations.** Canonicalization, - sorting, source-reference cleanup, ID repair, and accepted duplicate - consolidation remain inspectable but do not require operator action. -4. **Ordinary successful runs have zero warnings.** A successful-but-degraded - run may have warnings when policy permits continuation. Advisory or - observation volume alone must not produce warning stderr or a nonzero - warning count. -5. **Errors and rejections remain separate.** This work must not change - validation approval, rejection, retry budgets, exit status, or error policy. - A corrected superseded attempt leaves no final warning. An exhausted - rejection remains a rejection or run failure according to existing policy. -6. **Modules own meaning; the framework owns context and presentation.** A - producer chooses disposition, category, reason code, scope, and safe message. - The framework attaches stage and pipeline origin, aggregates deterministically, - enforces bounds, and supplies the final collections to output and CLI code. -7. **Durable contracts are versioned.** The incompatible grouped warning file - is `notarius.warnings.v2`, the new diagnostic file is - `notarius.diagnostics.v1`, and the machine-readable run receipt becomes - `notarius.run-result.v2`. Do not silently redefine the v1 receipt or warning - payload. -8. **There is no first-release configuration surface for presentation policy.** - Classification, sample bounds, and aggregation are fixed application - policy. Do not add suppression, escalation, verbosity, or per-reason config - in this work set. - -## Target Contract - -Use the following model unless existing Go naming requires a narrowly scoped -variation. Any naming variation must preserve the specified fields and -semantics. - -### Producer Diagnostic - -Add a framework contract representing a locally grouped producer diagnostic: - -- `disposition`: `warning`, `advisory`, or `observation`; -- `category`: one of `configuration`, `degradation`, - `validation_incomplete`, `fallback`, `data_quality`, or `normalization`; -- `reason_code`: stable, nonblank producer-owned identity; -- `occurrence_count`: exact positive number of represented occurrences; -- `samples`: deterministic bounded samples containing safe `scope` and - `message`; and -- `omitted_sample_count`: exactly `occurrence_count - len(samples)`. - -Validate these combinations: - -| Disposition | Allowed categories | -| --- | --- | -| `warning` | `configuration`, `degradation`, `validation_incomplete`, `fallback` | -| `advisory` | `data_quality` | -| `observation` | `normalization` | - -This strict matrix is intentional. It makes the process-only warning rule a -contract invariant rather than a convention inferred from reason codes. - -Use these fixed limits: - -- reason code: 128 UTF-8 bytes; -- scope: 512 UTF-8 bytes; -- sample message: 4 KiB of valid UTF-8; -- retained distinct samples per group: 3; and -- groups returned by one producer or validator result: 64. - -Blank or invalid required fields, invalid disposition/category combinations, -invalid UTF-8, inconsistent counts, excessive samples, or excessive local -groups make the producer result invalid. They must return a bounded contextual -error rather than be silently repaired. The three-sample and 64-group bounds -are public operational safeguards and may be asserted directly in their owning -contract tests. - -The shared collector must count every occurrence and retain the first three -distinct samples in producer order. Repeated identical samples still increase -`occurrence_count`. The omission count is numeric metadata, never another -diagnostic record. - -### Framework Origin And Aggregation - -The framework adds this origin before final aggregation: - -- stage: `references`, `chunk`, `extract`, `merge`, or `normalize`; -- step ID, when applicable; -- lane ID, when applicable; -- module key, when applicable; and -- validator key, for validator-produced diagnostics. - -Chunk ID and zero-based chunk index belong on samples, not group origin, -because otherwise identical per-chunk findings cannot aggregate. Use an -optional integer representation that preserves chunk index zero. - -The stable aggregation key is: - -```text -disposition + category + reason_code -+ stage + step_id + lane_id + module_key + validator_key -``` - -Scope, message, chunk ID, and chunk index are excluded from the key. Groups and -samples retain first-occurrence order from the runner's existing canonical -ordering; completion timing must never affect the result. When groups merge, -sum exact occurrence counts and retain the first three distinct samples. - -Use separate global bounds: - -- at most 128 actionable warning groups; exceeding this limit is a framework - error because Notarius must not hide process degradation; and -- at most 256 advisory/observation groups. Additional non-warning groups are - omitted from representation while their occurrences contribute to an exact - `unrepresented_occurrence_count` and set `truncated: true`. - -The non-warning envelope's group count means represented groups. Its occurrence -count includes represented and unrepresented occurrences. The warning -collection is never truncated, so both warning counts are exact. - -### Classification Matrix - -Migrate existing producers according to this matrix. - -| Disposition and category | Existing conditions | -| --- | --- | -| Warning / `configuration` | `empty_reference` | -| Warning / `validation_incomplete` | `validator_execution_incomplete`, covering both exhausted failure and skip when `warn_continue` advances the candidate | -| Warning / `degradation` | `scene_classification_unavailable` from combat-turn and enemy-event extraction gates | -| Warning / `fallback` | `npc_semantic_reconciliation_exhausted`, `item_semantic_reconciliation_exhausted`, `location_semantic_reconciliation_exhausted` | -| Advisory / `data_quality` | All ten source-relatedness families; `spell_name_unresolved`; `item_occurrence_unknown_item_id`; `location_occurrence_unknown_location_id`; guarded invalid semantic-consolidation proposals | -| Observation / `normalization` | Field and whitespace normalization; durable ID recomputation; source-reference sorting/deduplication; canonical record ordering; exact or approved semantic duplicate consolidation | - -Rename `item_occurrence_source_unrelated` to -`item_occurrence_not_near_source` when it moves into the new contract. Keep -other existing reason codes unless this plan explicitly changes them. Split the -item registry's internal retry reason from its accepted advisory: use -`item_semantic_retry_proposal_invalid` for the retry directive and retain -`item_semantic_proposal_invalid` only for the accepted data-quality advisory. - -Local `*_warnings_omitted` reason codes disappear. Omission is represented by -group counts. Rejections, structural failures, provider failures, cancellation, -and successful corrected retries do not receive a diagnostic solely to mirror -their existing error, rejection, validation, or attempt-debug records. - -### Durable And CLI Contracts - -Always publish both companion files, including for empty collections: - -- `warnings.json` with schema version `notarius.warnings.v2`, actionable groups, - exact `group_count`, and exact `occurrence_count`; -- `diagnostics.json` with schema version `notarius.diagnostics.v1`, advisory and - observation groups, represented `group_count`, total `occurrence_count`, - `truncated`, and `unrepresented_occurrence_count`; and -- `index.json` with both `warnings_file` and `diagnostics_file`. - -The run receipt becomes `notarius.run-result.v2` and replaces ambiguous -`warning_count` with: - -- `warning_group_count`; -- `warning_occurrence_count`; -- `diagnostic_group_count`; -- `diagnostic_occurrence_count`; and -- `diagnostics_truncated`. - -The human CLI writes a warning summary to stderr only when -`warning_group_count > 0`. It reports group and occurrence counts plus the -durable warning path. Advisory and observation counts do not create a warning -line. Preserve the existing successful stdout summary and process exit rules. - -Remove `OutputResult.Warnings`. An output encoder either returns its complete -logical files or returns an error. It cannot discover a warning after the -warning file has already been serialized. - -## Stage 1 ✅ — Record The Decision And Add Core Diagnostic Primitives - -### Goal - -Create the durable decision record and the generic, independently tested -diagnostic model without changing current runtime behavior. - -### Work - -- Add `docs/adr/0015-separate-process-warnings-from-quality-diagnostics.md` - using the repository ADR format. Record the governing decisions, origin and - aggregation ownership, bounded samples, fresh/resume equivalence, versioned - public contracts, and removal of output-result warnings. Mark it accepted; - do not claim implementation is complete. -- Add the producer diagnostic, sample, disposition, category, origin, group, - and collection types under `internal/framework/contracts/`. -- Add contract validation with the exact allowed combinations and limits in - this plan. -- Add a small generic collector under `internal/framework/diagnostics/` that - groups producer-local occurrences by disposition, category, and reason code, - preserves exact counts, and retains three distinct samples. -- Keep `contracts.Warning` and existing result fields temporarily. Do not alter - current warning output in this stage. - -### Tests - -Use table-driven contract tests for valid classifications, invalid UTF-8, -blank/oversized fields, invalid counts, the disposition/category matrix, sample -bounds, exact repeated-occurrence counting, and deterministic distinct-sample -selection. Test behavior through exported package contracts rather than private -collector structure. - -### Acceptance Criteria - -- Existing application behavior and public JSON remain unchanged. -- The new primitives cannot represent an LLM-quality warning because - `warning` plus `data_quality` is rejected. -- `go test ./internal/framework/contracts ./internal/framework/diagnostics` - passes. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 2 ✅ — Add Framework Diagnostic Transport And Transitional Projection - -### Goal - -Carry structured diagnostics through every framework result boundary and add -origin-aware aggregation while preserving the current public warning contract -temporarily. - -### Work - -- Add structured diagnostic fields alongside legacy warning fields in chunk - plan, typed extract, typed merge, typed normalize, normalize-retry fallback, - validation, run input/output, and output request contracts. -- Extend erased typed results, producer-attempt state, lane result collection, - ordered-step handoff, and run finalization to transport diagnostics. -- Preserve validator identity in validation reports instead of flattening new - diagnostics through a warning-only helper. -- Attach current stage, step, lane, module, validator, and sample-level chunk - context at promotion time. -- Add a run-level aggregator implementing the stable key, canonical - first-occurrence ordering, count merging, and bounds from this plan. -- Preserve terminal-attempt semantics: diagnostics from superseded attempts - remain debug-only; only terminal accepted or terminal rejected candidate - diagnostics are promoted. -- Add a private transitional projection from each structured diagnostic sample - to the existing legacy `Warning` collection so public surfaces remain - unchanged until Stage 8. Mark it explicitly temporary and ensure it does not - duplicate diagnostics already supplied through the legacy path. - -### Tests - -Extend focused framework tests for terminal attempt promotion, terminal -rejection, reversed concurrent completion, stage/lane/chunk origin, stable -grouping across chunks, and global bounds. Use controlled test dispositions and -origins; do not assert incidental diagnostic prose. - -### Acceptance Criteria - -- A producer may return either legacy warnings or new structured diagnostics - during migration, but not produce duplicate public records through both. -- Structured groups are deterministic under concurrent completion. -- Superseded attempt diagnostics never reach final groups. -- Existing public warning tests still pass through the transitional projection. -- Focused framework tests and `go test ./internal/framework/pipeline` pass. - -This stage is appropriately sized for one high-reasoning -`gpt-5.6-terra` prompt. Do not combine it with D&D migration. - -## Stage 3 ✅ — Migrate Framework Process Signals And Close The Output Boundary - -### Goal - -Move framework-owned warning conditions into the new process-only contract and -remove the output-encoder consistency defect. - -### Work - -- Convert `empty_reference` to warning/configuration with reference origin and - bounded, non-sensitive sample content. -- Convert incomplete validation to warning/validation-incomplete whenever - `warn_continue` advances a candidate after an applicable validator either - exhausts with failure or skips. Retain validator key and typed failure/skip - detail in validation summaries; do not copy provider errors or arbitrary - validator prose into diagnostic samples. -- Aggregate all incomplete-validator occurrences rather than generating - omission warning records. -- Validate module and validator diagnostic results before promotion. An invalid - diagnostic contract is a bounded contextual framework error. -- Remove `Warnings` from `contracts.OutputResult`, the output adapter, test - doubles, and runner append logic. Keep `OutputRequest` diagnostic input; an - encoder failure remains an error. -- Add a regression test proving no post-encoding result can make receipt/debug - warning state disagree with the already encoded logical files. - -### Acceptance Criteria - -- Failed and skipped validators are both visible when incomplete validation is - allowed to continue. -- A fail-run validation policy still fails rather than converting the failure - into a warning. -- Output encoders have no successful post-encoding warning capability. -- Existing validation decisions, retry budgets, and exit behavior are - unchanged. -- `go test ./internal/framework/... ./internal/modules/generic/output/json` - passes. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 4 ✅ — Version Diagnostic-Bearing Cache And Checkpoint State - -### Goal - -Persist and replay structured producer diagnostics safely without allowing old -warning-only state to be interpreted as complete new state. - -### Work - -- Store producer-local structured diagnostics in chunk-plan and extract, - merge, and normalize checkpoint envelopes. Do not persist current pipeline - origin in source-keyed chunk-plan state. -- On reuse, attach the current run's origin at the same logical promotion point - used by fresh execution, then aggregate once. -- Bump the checkpoint workspace schema from v3 to - `notarius.workspace.v4` and the chunk-plan schema from v2 to - `notarius.chunk-plan.v3`. -- Treat earlier workspace and chunk-plan versions as incompatible cache misses, - never as corrupt fatal state and never as reusable diagnostic-complete state. -- Preserve the rule that validation-incomplete results and their dependents are - not reusable. - -### Tests - -Add or adapt behavioral tests for old-version invalidation, structured -diagnostic round trips, fresh/resume equality of groups and counts, exact-once -replay, and current-validator diagnostics on a reused chunk plan. - -### Acceptance Criteria - -- Fresh and resumed logical runs produce deeply equal diagnostic groups and - counts. -- Reuse provenance remains in checkpoint events and does not alter diagnostic - identity. -- Old cache data is safely bypassed. -- `go test ./internal/framework/checkpoint ./internal/framework/pipeline` - passes. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 5 ✅ — Migrate All D&D Source-Relatedness Validators - -### Goal - -Move all ten heuristic relatedness families from warnings to bounded -data-quality advisories. - -### Work - -- Migrate combat turns, enemy events, item occurrences, item registry, - location occurrences, location registry, NPC occurrences, NPC registry, - scene descriptions, and spells source-relatedness validators. -- Use the shared generic collector so every family returns exact occurrence - counts and at most three distinct samples. Remove each relatedness omission - reason code. -- Bring NPC and spell validators under the same contract bounds as every other - family. -- Rename `item_occurrence_source_unrelated` to - `item_occurrence_not_near_source`. -- Preserve current approval behavior, validator chain registration and order, - lexical algorithms, scope selection, and messages except for changes needed - to satisfy generic safety bounds. - -### Tests - -Consolidate repetitive limiter tests where a shared collector contract already -owns the behavior. Retain family tests for realistic triggering and -non-triggering cases. Prove that a relatedness finding is advisory/data-quality -and cannot increment process warning groups. - -### Acceptance Criteria - -- All ten relatedness validators use the same structured advisory convention. -- No relatedness validator returns a legacy warning or an omission record. -- High-cardinality NPC and spell results remain bounded while reporting exact - occurrence counts. -- `go test ./internal/modules/dnd/validate/...` passes. - -This repetitive but cohesive migration is appropriately sized for one -`gpt-5.6-terra` prompt. Do not combine it with normalizer migration. - -## Stage 6 — Migrate D&D Registry Normalizers And Semantic Reconciliation ✅ - -### Goal - -Apply the new taxonomy to NPC, item, and location registry normalization and -semantic-reconciliation outcomes. - -### Work - -- Migrate deterministic field cleanup, ID recomputation, source-reference - normalization, ordering, and accepted duplicate consolidation to - observation/normalization groups. -- Migrate exhausted NPC, item, and location semantic reconciliation to - warning/fallback groups. -- Migrate guarded invalid item semantic proposals to advisory/data-quality. -- Split the item retry directive reason to - `item_semantic_retry_proposal_invalid`; keep correction control data separate - from the accepted advisory reason and from model-facing text. -- Replace local warning limiters and omission records with the shared - diagnostic collector while preserving artifact values, safe fallback, and - retry budgets. - -### Tests - -Protect exact artifact outcomes, classification, retry exhaustion, safe -currency behavior, occurrence counts, and bounded samples. Do not preserve -old warning slice lengths or omission prose. - -### Acceptance Criteria - -- Successful registry cleanup and consolidation produce observations only. -- A guarded model proposal produces an advisory, not a warning. -- Exhausted semantic reconciliation is the only registry-normalizer process - warning family. -- Semantic retry behavior and artifacts are unchanged. -- Registry normalizer and semantic-reconciliation tests pass. - -This stage is appropriately sized for one high-reasoning -`gpt-5.6-terra` prompt. - -## Stage 7 — Migrate Remaining D&D Producers ✅ - -### Goal - -Complete D&D diagnostic classification across spells, occurrences, combat, -enemy events, scenes, and extraction gates. - -### Work - -- Migrate spell, combat-turn, enemy-event, item-occurrence, - location-occurrence, NPC-occurrence, and scene-description normalizers. -- Classify deterministic cleanup, ordering, source-reference normalization, - and duplicate consolidation as observation/normalization. -- Classify unresolved spell, item, or location membership as - advisory/data-quality, never warning. -- Convert combat-turn and enemy-event `scene_classification_unavailable` - extraction gates to warning/degradation with module and chunk provenance. -- Remove all remaining D&D local warning omission reason codes and all uses of - `LimitWarnings`; retain text-safety helpers that still have value. - -### Tests - -Adapt family tests to assert artifact invariants and classification. Add one -assembled D&D test demonstrating that quality advisories and normalization -observations can be present while the warning group count remains zero. - -### Acceptance Criteria - -- No D&D producer returns a legacy `contracts.Warning`. -- No accepted-artifact uncertainty or unresolved grounding appears as a - process warning. -- Missing required scene classification remains an actionable process warning. -- All D&D package tests pass. - -This stage is appropriately sized for one high-reasoning -`gpt-5.6-terra` prompt. - -## Stage 8 ✅ — Remove The Legacy Warning Path And Finalize Aggregation - -### Goal - -Make the structured diagnostic collection the sole in-memory signal path. - -### Work - -- Migrate any remaining generic test modules and framework fixtures from - `contracts.Warning` to structured diagnostics. -- Remove `contracts.Warning`, legacy `Warnings` fields, clone helpers, warning - append helpers, transitional projections, and obsolete D&D warning limiters. -- Make `RunOutput` and `OutputRequest` carry the finalized warning and - non-warning grouped collections derived from one aggregator result. -- Enforce the 128-warning-group error bound and 256-non-warning-group truncation - behavior. Ensure occurrence counts remain exact and truncation metadata is - deterministic. -- Verify all producers are validated at their stable framework boundary. - -### Tests - -Add focused behavioral coverage for warning overflow failure, non-warning -truncation, exact occurrence totals, distinct sample retention, stable group -order, and no legacy double-counting. Remove tests whose only purpose was to -assert old flat-list caps or omission prose. - -### Acceptance Criteria - -- Repository search finds no production `contracts.Warning`, legacy warning - result field, or `*_warnings_omitted` reason code. -- One finalized collection supplies all later surfaces. -- An advisory-only successful run has zero warning groups. -- `go test ./internal/framework/... ./internal/modules/dnd/...` passes. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 9 ✅ — Publish Versioned Warning And Diagnostic Files - -### Goal - -Replace the flat durable warning payload with the two target grouped contracts. - -### Work - -- Update the production JSON encoder to always emit grouped `warnings.json` - (`notarius.warnings.v2`) and `diagnostics.json` - (`notarius.diagnostics.v1`). -- Include exact counts and diagnostic truncation metadata specified above. -- Add `diagnostics_file` to `index.json`; retain `warnings_file`. -- Ensure warning groups appear only in `warnings.json` and advisory/observation - groups appear only in `diagnostics.json`. -- Update `docs/integrations/json-output.md` in the same stage. It owns file - names, envelope schemas, count meanings, group/sample fields, bounds, - truncation, and index discovery. Link rather than duplicate CLI behavior. -- Update maintained output examples or test fixtures only where they encode - implemented output contracts. - -### Tests - -Use encoder contract tests for empty and populated files, schema versions, -partitioning, counts, truncation, index paths, deterministic JSON ordering, and -newline/valid-JSON conventions. Update assembled pipeline tests to compare both -durable files with the finalized in-memory collections. - -### Acceptance Criteria - -- Every production JSON bundle contains both companion files and index paths. -- No diagnostic appears in both files. -- Published counts agree with the in-memory collection. -- Output encoder and maintained example contract tests pass. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 10 ✅ — Introduce Run Result V2 And Quiet CLI Presentation - -### Goal - -Give subprocess consumers unambiguous counts and make an ordinary successful -run quiet on the warning stream. - -### Work - -- Change the emitted receipt to `notarius.run-result.v2` with the five fields - specified in the target contract. Remove v1 `warning_count`; do not emit two - competing count models. -- Derive receipt counts from the same finalized collection used by the JSON - encoder. -- Print a stderr warning summary only when actionable warning groups exist. It - must include group count, occurrence count, and the output-relative or - absolute details path consistent with current CLI path conventions. -- Do not mention advisory/observation counts as warnings. Preserve ordinary - stdout completion output and all exit classifications. -- Update `docs/cli.md`, `docs/integrations/run-result.md`, - `docs/consumers/subprocess.md`, and `docs/consumers/dnd-pipeline.md` in the - same stage. Each document must keep to its canonical scope and link to the - grouped JSON contract rather than duplicating schemas. - -### Tests - -Cover zero-warning approved success, advisory-only success, degraded success -with warnings, rejected output, JSON receipt delivery, and non-production -output modules. Assert semantic fields and stream choice, not complete prose. - -### Acceptance Criteria - -- Advisory-only and observation-only successful runs write no warning stderr. -- Degraded success writes one concise actionable warning summary. -- Receipt counts and durable files agree. -- Receipt, CLI command, and consumer contract tests pass. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 11 ✅ — Align Debug, Manifest, And Resume Surfaces - -### Goal - -Ensure forensic and provenance surfaces reflect the new model without creating -another competing warning contract. - -### Work - -- Replace final flat warning debug summaries with final grouped warning and - non-warning diagnostic projections, using separate clearly named files or - one versioned diagnostic summary envelope consistently with existing debug - layout. -- Keep attempt-local candidate diagnostics in detailed debug records with - stage/attempt provenance. Preserve redaction and the rule that raw model - responses and correction messages appear only in detailed debug. -- Keep manifest validation and rejection summaries authoritative for validation - status; do not duplicate full diagnostic groups into the manifest. -- Ensure checkpoint events, not diagnostic origin, identify reuse. -- Update `docs/operations.md`, `docs/internal/state.md`, and the debug portions - of `docs/internal/pipeline.md` in the same stage. - -### Tests - -Cover successful debug capture, partial failure debug capture, -terminal-attempt-only final grouping, redaction, and fresh/resume equality. -Retain existing validation summary tests rather than duplicating every outcome -in diagnostic tests. - -### Acceptance Criteria - -- Debug consumers can distinguish process warnings, quality advisories, and - observations. -- Debug final counts match receipt and durable output when output is published. -- Attempt detail remains bounded and sensitive-data rules are preserved. -- Debug, state, and resume-focused tests pass. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 12 ✅ — Update Architecture And Internal Producer Documentation - -### Goal - -Make the implemented ownership and classification rules durable for future -modules without duplicating public schemas. - -### Work - -- Update `docs/policy/architecture.md` to state the process-only warning - invariant, module semantic ownership, framework origin/aggregation ownership, - and ordinary-success-zero-warnings expectation. Link to ADR-0015 for - rationale. -- Update `docs/internal/pipeline.md` for terminal promotion, origin enrichment, - aggregation, bounds, concurrency ordering, checkpoint replay, and output - handoff. -- Update `docs/internal/modules.md` with the generic producer contract and the - prohibition on using warnings for extraction quality. -- Update `docs/internal/dnd.md` with the implemented classification matrix and - shared collector convention. Link to integration contracts for durable - schemas rather than restating them. -- Update `docs/internal/overview.md` only if the new generic diagnostics package - changes the component inventory. -- Remove stale terminology from current docs outside `docs/roadmap/`, but do - not write release history or document unimplemented configuration. - -### Acceptance Criteria - -- Each volatile fact has one canonical owner under the documentation policy. -- Future module authors can determine which disposition/category to use and - where origin and aggregation are attached. -- Current documentation contains no claim that heuristic quality doubt is a - warning. -- Relative Markdown links resolve. - -This stage is appropriately sized for one `gpt-5.6-terra` prompt. - -## Stage 13 ✅ — Final Behavioral Verification And Cleanup - -### Goal - -Verify the complete migration, remove transitional residue, and establish that -the audit findings are fully addressed. - -### Work - -- Run repository searches for legacy warning types, warning omission reason - codes, old run-result v1 fields, old flat warning JSON assumptions, and - `OutputResult.Warnings`. -- Review every producer inventoried in `audit.md` against the final - classification matrix. -- Run the maintained minimal and complete D&D workflows with offline fake LLM - clients. Do not require provider credentials or add live calls to the default - suite. -- Verify these relationships end to end: - - ordinary approved success and advisory-only success have zero warnings; - - successful process degradation has a warning; - - corrected retries retain only terminal diagnostics; - - rejected and failed runs retain their established semantics; - - warning and diagnostic group order is deterministic under concurrency; - - fresh and resumed runs are equivalent; - - global group and sample bounds hold; and - - receipt, stderr, durable files, index, and debug counts agree. -- Remove obsolete helpers and redundant tests made unnecessary by stronger - package-level collector or end-to-end contract tests. - -### Validation Commands - -Run at minimum: - -```sh -go test ./... -go vet ./... -go build ./cmd/notarius -``` - -Also run any repository documentation or link checker discovered during the -stage. If none exists, perform a focused relative-link review for the documents -changed by Stages 9–12. - -### Acceptance Criteria - -- All three repository-wide Go commands pass offline. -- No live API key is required. -- No legacy flat-warning code path or public v1 warning-count claim remains in - current-behavior documentation. -- All seven audit findings are addressed without changing artifact values, - validator chain order, retry budgets, rejection policy, or exit semantics. -- The ordinary maintained successful workflow emits zero actionable warnings; - non-warning diagnostics remain available in `diagnostics.json`. - -This final verification is appropriately sized for one `gpt-5.6-terra` -prompt. - -## Deferred Work - -The following are deliberately outside this implementation plan: - -- provider-backed measurement of production advisory precision or frequency; -- configurable warning suppression, escalation, verbosity, or reason filters; -- converting heuristic relatedness checks into rejection rules; -- the planned LLM-backed D&D combat-scene validator; -- metrics or telemetry export beyond the specified durable and debug files; -- a two-phase output encoder protocol; and -- release preparation or release-note creation. - -Provider-backed production data may inform later advisory tuning, but it is not -required to implement or validate this architecture.