Files
notarius/docs/roadmap/audit-sequence.md

42 KiB
Raw Blame History

Staged Codebase Audit Sequence

Purpose

This document turns the scope in Codebase Audit Plan into an ordered, execution-ready audit sequence. Each numbered stage is intended to be run as one prompt by an LLM coding agent. Stages must be completed in order because later stages rely on the target snapshot, terminology, candidate ledger, and architectural conclusions established earlier.

This is an audit, not an implementation plan. The only audit deliverable is docs/roadmap/audit.md. No stage may modify production code, tests, assets, examples, policies, ADRs, integration documentation, or other current-behavior documentation.

Execution Rules For Every Stage

Before beginning a stage:

  1. Read AGENTS.md, docs/development.md, docs/roadmap/audit-plan.md, this sequence, and the current docs/roadmap/audit.md.
  2. Follow the task-specific reading guide in docs/development.md and the architecture and testing policies.
  3. Read the focused implementation and tests before drawing conclusions from metrics, naming, or textual similarity.
  4. Prefer the current code knowledge graph for symbol discovery, call tracing, dependency inspection, and similarity candidates. Refresh the graph if its indexed head is stale. Use textual search for literals, configuration, diagnostics, asset content, and non-code files.
  5. Verify that production code, assets, tests, examples, policies, ADRs, and current-behavior documentation have not changed since the audit target was recorded. Roadmap-only commits are permitted. If the audited implementation changed, stop and report that the baseline must be re-established.
  6. Treat every command as read-only except edits to docs/roadmap/audit.md. Do not fix a defect, refactor a helper, add a test, or update current documentation during the audit.
  7. Add evidence to the audit document as the stage proceeds. Do not rely on the final response to preserve findings.
  8. Report both findings and significant inspected areas with no findings.

Do not report a similarity or complexity score as a finding. Trace callers, read the code, compare behavior, and identify an actual risk or concrete simplification first. Do not recommend an abstraction that weakens static typing, crosses an ownership boundary, or has more policy parameters than the code it replaces.

Audit Document Contract

Stage 1 creates docs/roadmap/audit.md. All later stages maintain it. The document must use this structure:

# Codebase Audit

## Audit Metadata
## Executive Summary
## Finding Index
## Findings
## Intentional Complexity And Duplication To Preserve
## Areas Reviewed Without Findings
## Validation Record
## Coverage Matrix

Until the final stage, Executive Summary and Finding Index may state that they are pending synthesis. Findings belong under area-specific third-level headings within ## Findings.

Use stable IDs by area:

  • ARCH-### for architecture and dependency boundaries;
  • CFGCLI-### for configuration and CLI composition;
  • PIPE-### for resolution, preparation, and typed registries;
  • REF-### for references and ordered handoffs;
  • RUN-### for execution, validation, retry, and concurrency;
  • STATE-### for filesystem state, checkpoints, and debug behavior;
  • LLM-### for the LLM runtime, PromptKit, prompts, and assets;
  • MOD-### for generic and Seriatim modules;
  • DND-CORE-### for shared D&D mechanics and codecs;
  • DND-REG-### for NPC, item, and location registries;
  • DND-OCC-### for NPC, item, and location occurrences;
  • DND-SCENE-### for spells, scene chunking, and scene descriptions;
  • DND-COMBAT-### for combat turns and enemy events; and
  • TEST-### or COMMENT-### for cross-cutting test or comment findings found during final synthesis.

Every finding must contain:

### AREA-001 — Concise title

- **Severity:** High, Medium, or Low
- **Category:** Correctness, Efficiency, Duplication, Simplicity, Test Quality,
  or Documentation/Comments
- **Evidence:** Exact files, symbols, call paths, and observed behavior
- **Impact:** The concrete risk or cost
- **Recommendation:** A bounded implementation direction
- **Preserve:** Invariants and contracts that remediation must retain
- **Validation:** Focused checks that would demonstrate success
- **Grouping:** Independent, or the IDs with which this should be implemented

If later evidence disproves a finding, remove it and record the examined design under Areas Reviewed Without Findings when that negative result is useful. If two stages identify the same root cause, retain one finding and add the second stage's evidence to it. Do not preserve duplicate symptoms merely to show that each stage produced output.

The Intentional Complexity And Duplication To Preserve section should record only notable false-positive candidates: exact symbols, why explicit code is preferable, and the invariant or ownership boundary it protects. It is not a list of every repeated one-line method.

The coverage matrix must have one row per stage area with status Pending, Reviewed, or Revisit, the packages/documents inspected, validation run, and finding IDs. Each stage updates its own row.

Stage 1: Establish The Baseline And Architecture Boundaries

Goal

Freeze the implementation snapshot, establish the audit document and evidence format, verify the baseline, and audit high-level dependency direction before subsystem review begins.

Required Reading

  • docs/policy/architecture.md
  • docs/policy/testing.md
  • docs/policy/documentation.md
  • docs/internal/overview.md
  • all accepted ADRs under docs/adr/
  • internal/cli/catalog.go
  • production registrars under internal/modules/*/register
  • root assets package Go source

Work

  1. Record git rev-parse HEAD, branch, Go version, PromptKit version, and audit date in Audit Metadata. State that later roadmap-only commits do not change the production target.
  2. Confirm the worktree contains no unexplained production changes. Record any pre-existing roadmap changes without treating them as audited code.
  3. Refresh the repository knowledge graph and record the returned current project identity. Do not reuse a graph whose indexed head predates the audit target.
  4. Create the complete audit document structure, finding template, coverage matrix, and initial validation record.
  5. Run the baseline commands below. Record exact pass/fail results; do not edit code to repair a failure.
  6. Map production package imports and trace graph-reported calls from framework or core into modules and from modules into CLI. Distinguish test-only edges, interface implementation edges, and real production imports.
  7. Verify the composition root, assets leaf, domain/framework direction, fixed pipeline shape, typed artifact boundary, and physical-state ownership.
  8. Record architecture findings, intentional boundaries that initially resemble duplication, and areas examined without findings.

Validation

go test ./...
go vet ./...
go build ./cmd/notarius
git diff --check

Acceptance Criteria

  • docs/roadmap/audit.md exists with the required structure.
  • The exact production target and baseline results are recorded.
  • The code graph is current.
  • Production dependency direction has been inspected, not inferred from folder names.
  • The architecture coverage row is Reviewed or explains a specific Revisit.
  • No file other than docs/roadmap/audit.md was changed by this stage.

This stage is suitable for one audit prompt.

Stage 2: Audit Configuration And CLI Composition

Goal

Audit configuration loading and validation, effective resolution, process composition, session/profile selection, run publication, and terminal error handling.

Required Reading

  • docs/config.md, docs/cli.md, and docs/operations.md
  • docs/internal/configuration.md and docs/internal/cli.md
  • internal/core/config/
  • internal/cli/run.go, catalog.go, session.go, promptkit_profiles.go, run_result.go, and run_terminal.go
  • configuration, session, state, production, example, and command contract tests

Work

  1. Trace RunWithOptions through configuration discovery, effective config, catalog construction, pipeline selection, session resolution, state construction, framework invocation, output publication, and terminalization.
  2. Review the high-complexity paths around runPipelineCommand, validatePipelineProfiles, selected reference targets, recomputation policy, and option normalization. Separate necessary orchestration from policy or transformation that has a clearer existing owner.
  3. Verify file/environment/CLI/profile precedence, strict parsing, unknown option handling, redaction, and equality between validation-time and runtime profile sources.
  4. Verify that default session identity depends on the intended input identity and input-module key but not references, pipeline-local changes, or unstable paths.
  5. Trace every terminal outcome: success, validation rejection, framework error, cancellation, output failure, and requested debug/report failure. Confirm the primary error and publication guarantees are stable.
  6. Compare repetitive CLI state setup and cleanup paths. Recommend a helper only if it would reduce policy duplication without hiding when physical state is allocated.
  7. Review tests for duplicated lower-layer policy or missing consequential CLI behavior.
  8. Add CFGCLI-* findings and update the validation and coverage sections.

Validation

go test ./internal/core/config ./internal/cli
go vet ./internal/core/config ./internal/cli

Acceptance Criteria

  • Configuration precedence, session derivation, profile setup, output publication, and terminalization have each been traced end to end.
  • Every CLI/config finding names its correct owner rather than pushing process policy into the framework.
  • Dense functions examined without a justified refactor are recorded as such when that conclusion will prevent repeat work.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 3: Audit Pipeline Resolution, Preparation, And Typed Registries

Goal

Audit static composition, option validation, typed registry mechanics, and the construction boundary that must fail before source parsing.

Required Reading

  • docs/internal/pipeline.md and docs/internal/modules.md
  • internal/framework/contracts/
  • internal/framework/pipeline/profile.go, options.go, module.go, construction.go, prepare.go, and prepared_fingerprints.go
  • all stage, validator, validator-chain, codec, and evidence-projector registry implementations and focused tests

Work

  1. Trace configuration-owned profiles into ResolvePipeline, resolved digest construction, Prepare, and the private prepared pipeline.
  2. Review capability checks, selected-lane filtering, validator-chain resolution, effective LLM profile application, option validation, and complete registry-set validation.
  3. Inspect high-complexity resolution functions including ResolvePipeline, resolveArtifactLane, applyEffectiveLLMProfiles, and generated-binding validation only to the point where Stage 4 assumes ownership of reference semantics.
  4. Compare every typed registry implementation. Identify genuinely shared registration mechanics, drift, redundant wrapper layers, dead compatibility paths, and opportunities to reduce code while retaining exact Go types and stage-specific diagnostics.
  5. Verify private type erasure reports incompatibility rather than panicking and that preparation clones all retained mutable values.
  6. Verify checkpoint fingerprints are collected once from the complete prepared implementation set and distinguish semantic execution identity from scheduling or diagnostics.
  7. Review focused tests for registry mechanism duplication and gaps at typed package boundaries.
  8. Add PIPE-* findings and update coverage and validation.

Validation

go test ./internal/framework/contracts ./internal/framework/pipeline
go vet ./internal/framework/contracts ./internal/framework/pipeline

Acceptance Criteria

  • Resolution, preparation, registry typing, and fingerprint collection are all covered.
  • Reference-specific questions are handed to Stage 4 rather than partially decided here.
  • Proposed helpers preserve compile-time type guarantees and correct ownership.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 4: Audit References And Ordered Handoffs

Goal

Audit external reference materialization, target resolution, generated artifact handoffs, provenance, ordered steps, and dependency invalidation as one coherent correctness boundary.

Required Reading

  • reference and ordered-step sections of docs/config.md, docs/internal/pipeline.md, and docs/internal/state.md
  • internal/framework/pipeline/references.go, handoff.go, relevant portions of profile.go, checkpoint.go, and runner handoff code
  • internal/cli reference-selector and recomputation-policy code
  • reference, handoff, typed-resolution, recomputation, and assembled pipeline tests

Work

  1. Trace external reference configuration through target resolution, materialization, preparation, and operation request cloning.
  2. Trace a generated normalized artifact through codec canonicalization, producer provenance, operation reference construction, consumer fingerprints, and later-step execution.
  3. Verify required/unbound/default/local reference behavior for selected and unselected lanes and stages.
  4. Verify rejection of unknown slots, forward references, cycles, ambiguity, incompatible artifact kinds, schema/media mismatches, rejected producers, and missing accepted normalized state.
  5. Review validateGeneratedBindings, resolveReferenceTargetBindings, buildStepReferenceSets, materializeReferenceTarget, and generatedReferenceItem for repeated scans, mixed policy, repeated encoding, or opportunities for clearer data structures.
  6. Verify references cannot become source evidence and that registry/reference content is not leaked into errors, manifests, or state metadata.
  7. Reconcile any overlap with PIPE-* findings rather than reporting duplicate symptoms.
  8. Add REF-* findings and update coverage and validation.

Validation

go test ./internal/framework/pipeline ./internal/cli
go test ./internal/modules/integration/...

Acceptance Criteria

  • At least one external and one generated reference path have been traced end to end.
  • Ordinary execution, resume dependencies, and selective recomputation are distinguished.
  • Findings preserve canonical codec and provenance checks.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 5: Audit Execution, Validation, Retry, And Concurrency

Goal

Audit the runtime state machine for deterministic ordering, bounded work, correct cancellation, whole-output validation, retries, and typed handoff.

Required Reading

  • execution and validation sections of docs/policy/architecture.md and docs/internal/pipeline.md
  • runner.go, runner_chunk_plan.go, runner_concurrent.go, typed_execution.go, runner_typed.go, normalize_retry.go, chunk_validation.go, and synchronized_collaborators.go
  • concurrency, retry, rejection, session, typed-checkpoint, debug, manifest, and handoff tests

Work

  1. Trace success and failure from source parsing through chunk validation, extraction, lane continuation, merge, normalize, and output request.
  2. Build a concise runtime state diagram for audit use. Verify bounded worker groups, chunk-first/lane-second dispatch, serial lane continuation, and overlap only where documented.
  3. Trace cancellation before dispatch, while queued, in an active provider or module call, during validation, and after the first framework error. Check that started work is awaited and undispatched work cannot begin.
  4. Verify stable public ordering and selected error are independent of goroutine completion order. Check warnings and rejections across retry attempts.
  5. Verify rejection versus framework-error semantics and output suppression.
  6. Inspect repeated extract/merge/normalize orchestration, candidate encoding, hydration, validation calls, and synchronization. Look for double work, aliasing, double recording, permit leaks, lock-order hazards, or a smaller state representation.
  7. Identify exact symbols where invariant-level comments would materially aid future changes.
  8. Add RUN-* findings and update coverage and validation.

Validation

go test -race ./internal/framework/pipeline
go test -count=1 ./internal/framework/pipeline

Run shuffle testing only if ordinary tests pass:

go test -shuffle=on ./internal/framework/pipeline

Acceptance Criteria

  • All terminal states and cancellation positions listed above were examined.
  • Concurrency findings cite an actual interleaving or ownership risk, not only a complex function.
  • Comment findings state the invariant to explain.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 6: Audit State, Checkpoints, Debugging, And File Safety

Goal

Audit physical-state collaborators, cache identities, canonical persistence, resume/recompute decisions, debug isolation, redaction, and safe publication.

Required Reading

  • docs/internal/state.md, docs/operations.md, and state/security portions of docs/policy/architecture.md
  • internal/core/fileio, debugbundle, and source digest/clone helpers
  • internal/framework/checkpoint, chunkplan, chunkmap, debug, and evidencecontext
  • CLI cache, recompute, state-hardening, run-contract, and terminalization tests

Work

  1. Trace output, cache, and debug root construction and prove they remain independently optional and application-owned.
  2. Trace chunk-plan keying, validation, materialization, and atomic publication.
  3. Trace checkpoint recording and loading for cold execution, ordinary resume, invalidated dependencies, accepted normalize hydration, forced recompute, and a required predecessor failure.
  4. Verify reason categories/codes are assigned at validation sites rather than inferred from prose, and that bounded details cannot leak caller content.
  5. Inspect path-component validation, safe joins, permissions, atomic writes, symlink handling, overwrite/move/delete scope, and recoverability.
  6. Inspect defensive clones and encode/decode/hash sequences for aliasing or redundant canonicalization. Preserve copies that establish an ownership boundary.
  7. Compare repeated stage recorder/loader methods, path validators, codecs, and clone helpers. Classify intentional interface adapters separately from extractable primitives.
  8. Verify debug persistence cannot affect reuse and terminal persistence errors cannot replace the primary failure.
  9. Add STATE-* findings and update coverage and validation.

Validation

go test -race ./internal/core/fileio ./internal/core/debugbundle \
  ./internal/framework/checkpoint ./internal/framework/chunkplan \
  ./internal/framework/chunkmap ./internal/framework/debug \
  ./internal/framework/evidencecontext
go test ./internal/cli

Acceptance Criteria

  • Each physical-state family and its lifecycle was reviewed separately.
  • Both ordinary resume and selective recomputation were traced.
  • File-safety conclusions are grounded in path and write implementations, not only tests or documentation.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 7: Audit The LLM Runtime, Prompt Filesystems, And Assets

Goal

Audit the transport-neutral LLM boundary, PromptKit integration, scheduling, profiles, sessions, prompt/schema identity, asset filesystems, caching order, and secret handling.

Required Reading

  • docs/internal/llm.md, the PromptKit integration document, relevant ADRs, and LLM sections of architecture/configuration/operations docs
  • internal/framework/llm/ and internal/framework/promptfs/
  • internal/cli/promptkit_profiles.go and session/profile construction paths
  • root assets Go package, shared prompt fragments, all prompt manifests, and representative module prompt/schema loaders
  • LLM, scheduler, promptfs, profile, prompt-asset, schema, and secret tests

Work

  1. Trace one structured completion from a module request through the scheduled client, PromptKit preparation and provider execution, structured validation, response decoding, profile recording, and debug capture.
  2. Verify every production LLM call shares the one Notarius scheduler and that cancellation, FIFO admission, backend limits, and permit release compose correctly.
  3. Compare profile inspection, preflight, runtime, fallback assets, and local backend source construction. Verify semantic profile changes invalidate checkpoints without storing secrets or paths.
  4. Verify session propagation and backend-caching goals: stable session identity, exactly reusable prefix bytes, prompt ordering, cache-control metadata, and separation of stable references from changing transcript material.
  5. Audit prompt and schema filesystem flattening, duplicate detection, scoping, defensive reads, and content hashing. Compare llm asset registry and promptfs adapters for truly shared filesystem mechanics.
  6. Search model-visible assets and private schemas for opaque IDs, digests, UUID-copy tasks, provider-specific values, or duplicate instructions.
  7. Verify the root assets Go package remains a minimal content-only leaf.
  8. Inspect error adaptation and redaction for credentials, profile content, request/response bytes, provider-specific error types, and capacity errors.
  9. Add LLM-* findings and update coverage and validation.

Validation

go test -race ./internal/framework/llm ./internal/framework/promptfs
go test ./internal/cli ./internal/modules/dnd/register

Use textual searches for model-visible IDs and sensitive fields, but inspect every match semantically; durable schemas and non-model runtime metadata are not violations.

Acceptance Criteria

  • Scheduling, profiles, sessions, prompts, schemas, filesystem adapters, and redaction were each reviewed.
  • Prompt-cache recommendations preserve exact-byte identity and correct message ordering.
  • No recommendation moves domain prompt ownership into the generic runtime.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 8: Audit Generic And Seriatim Modules

Goal

Audit the non-D&D production extensions for strict options, source-boundary correctness, chunk-plan behavior, validators, output publication, and useful shared module mechanics.

Required Reading

  • docs/internal/modules.md
  • Seriatim, JSON output, chunk-map, and evidence-context integration contracts
  • all code and tests under internal/modules/generic and internal/modules/seriatim
  • relevant output preparation and example contract tests

Work

  1. Trace Seriatim bytes through parsing, validation, generic source document construction, metadata, and source-unit self-references.
  2. Trace generic unit chunking from options through source-addressed plan and materialization. Review integer option parsing for strictness and avoidable complexity.
  3. Review generic validators for correct target ownership and redundant parsing or semantic duplication.
  4. Trace JSON output options, lane allowlists, evidence-context preparation, logical file construction, deterministic names, and chunk-map publication.
  5. Inspect clone/metadata helpers, option decoders, module specs, registration, diagnostics, and tests for drift or reusable domain-neutral mechanics.
  6. Ensure no Seriatim fields leak beyond the input adapter and no D&D concepts enter generic modules.
  7. Add MOD-* findings and update coverage and validation.

Validation

go test ./internal/modules/generic/... ./internal/modules/seriatim/...
go test ./internal/framework/chunkmap ./internal/framework/evidencecontext

Acceptance Criteria

  • Input, chunking, validation, output, and registration paths were all traced.
  • Strict option parsing was evaluated as behavior, not merely line count.
  • Shared-helper recommendations remain domain-neutral and have demonstrated callers.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 9: Audit Shared D&D Types, Codecs, And Family Mechanics

Goal

Establish a current convention matrix for all D&D artifact families and audit shared types, codecs, source-reference mechanics, prompt/schema registration, mergers, diagnostics, and family registration before lane-specific review.

Required Reading

  • docs/internal/dnd.md, docs/internal/modules.md, and all D&D integration contracts
  • root files and shared packages under internal/modules/dnd
  • every package under internal/modules/dnd/codec
  • D&D merger packages and internal/modules/dnd/register
  • shared D&D assets and representative module asset declarations
  • corresponding focused tests

Work

  1. Add a D&D convention matrix to the audit document with one row for each of the ten durable artifact families. Include artifact kind/type, extractor, merger, normalizer, validators, codec, reference dependencies, execution class, prompt/schema ownership, and documented exceptions.
  2. Compare codec construction, Encode, Decode, candidate decoding, validation, schema loading, metadata, and defensive-copy behavior. Determine whether exact similarities are safe typed adapters or a useful codec primitive with one natural owner.
  3. Compare source-reference canonicalization, ordering, equality, exact identity, nil/empty handling, diagnostics, and clone helpers.
  4. Review artifact registration, default chains, evidence projectors, prompt assets, fallback profiles, and registrar failure behavior.
  5. Review shared registry resolver, entity reconciliation, diagnostics, comparison policies, and candidate JSON without deciding domain-specific registry behavior reserved for Stage 10.
  6. Classify repeated ManifestMetadata, CheckpointFingerprints, Register, New, and DecodeOptions methods. Report only helpers that reduce drift without obscuring per-module semantics.
  7. Record convention divergence that later lane stages must confirm or refute. Mark the D&D coverage row Revisit until Stages 1013 complete.
  8. Add DND-CORE-* findings and update validation.

Validation

go test ./internal/modules/dnd/codec/... ./internal/modules/dnd/shared/... \
  ./internal/modules/dnd/register

Acceptance Criteria

  • The convention matrix covers every current D&D artifact family.
  • Exact similarities are classified by semantics and ownership.
  • Lane-specific questions are explicitly handed to Stages 1013.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 10: Audit NPC, Item, And Location Registries

Goal

Audit the three canonical-noun registry families end to end and compare their identity, extraction, reconciliation, normalization, validation, and immutable reference behavior.

Required Reading

  • NPC, item, and location registry integration contracts and relevant ADRs
  • registry extractor packages for NPCs, items, and locations
  • npcs, items, and locations identity and registry packages
  • NPC, item, and location registry normalizers and validators
  • shared entity reconciliation and registry resolver code
  • corresponding prompt assets, private schemas, and focused tests

Work

  1. Trace each registry from model response through source-reference canonicalization, deterministic identity, merge, normalization, reconciliation/fallback, validation, codec, and generated reference.
  2. Compare all three extractors for duplicated or drifted construction, canonicalization, manifest metadata, fingerprints, prompt inputs, and response mapping.
  3. Compare identity derivation, comparison keys, exact-ID validation, registry lookup indexes, immutable projections, resolvers, digests, and defensive copies. Preserve same-name location and currency-specific behavior.
  4. Trace entity reconciliation candidate construction, descriptor mapping, eligibility, collision handling, group assessment, diagnostics, retry, and safe fallback for all three domains.
  5. Compare normalizers and validators for shared mechanics and domain policy that should remain local.
  6. Inspect runtime complexity relative to registry records and evidence ranges: nested scans, repeated canonicalization, repeated JSON projections, hashing, and cloning.
  7. Verify proper-name scope for NPCs and locations and item/currency registry rules as implemented; do not treat future-roadmap behavior as current.
  8. Resolve or refine Stage 9 registry-related candidates. Add DND-REG-* findings and update the convention matrix, coverage, and validation.

Validation

go test ./internal/modules/dnd/extract/npcregistry \
  ./internal/modules/dnd/extract/itemregistry \
  ./internal/modules/dnd/extract/locationregistry \
  ./internal/modules/dnd/npcs/... ./internal/modules/dnd/items/... \
  ./internal/modules/dnd/locations/... \
  ./internal/modules/dnd/normalize/npcregistry \
  ./internal/modules/dnd/normalize/itemregistry \
  ./internal/modules/dnd/normalize/locationregistry \
  ./internal/modules/dnd/validate/npcregistry/... \
  ./internal/modules/dnd/validate/itemregistry/... \
  ./internal/modules/dnd/validate/locationregistry/...

Acceptance Criteria

  • All three registry families were traced to durable validated output.
  • Shared mechanics and domain-specific exceptions are explicitly separated.
  • Identity, reconciliation, retry/fallback, and generated-reference behavior were each examined.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 11: Audit NPC, Item, And Location Occurrences

Goal

Audit the three registry-consuming occurrence families for contextual model selection, deterministic identity attachment, evidence separation, canonicalization efficiency, and consistent normalization/validation.

Required Reading

  • NPC-, item-, and location-occurrence integration contracts and grounding ADR
  • the three occurrence extractor, domain helper, codec, normalizer, and validator families
  • NPC, item, and location registry prompt projections/resolvers used by the occurrence extractors
  • occurrence prompt assets, private schemas, and focused tests
  • assembled generated-reference tests in internal/cli

Work

  1. Trace each occurrence response through contextual registry resolution, current-source evidence mapping, durable ID attachment, ordering, deduplication, merge, normalize, validation, and encoding.
  2. Verify model inputs/private responses contain no opaque entity IDs and that unknown or ambiguous selection uses the correct all-or-nothing policy.
  3. Verify registry evidence and context cannot become occurrence evidence. Confirm source IDs are attached only by deterministic code for the current document.
  4. Compare the three canonicalization paths for repeated lookup, conversion, sort, deduplication, clone, or intermediate response structures.
  5. Compare occurrence normalizers, registry validators, invariant validators, source-reference validators, and source-relatedness validators. Identify helpers only where semantics and diagnostics are actually identical.
  6. Verify nil versus present-empty results, stable kind ordering, duplicate collapse, holder/quantity behavior, same-name location selectors, and exact registry identity checks.
  7. Trace generated NPC/item/location registry handoffs into each consumer and checkpoint identity.
  8. Resolve or refine Stage 9 occurrence candidates. Add DND-OCC-* findings and update the convention matrix, coverage, and validation.

Validation

go test ./internal/modules/dnd/extract/npcoccurrences \
  ./internal/modules/dnd/extract/itemoccurrences \
  ./internal/modules/dnd/extract/locationoccurrences \
  ./internal/modules/dnd/normalize/npcoccurrences \
  ./internal/modules/dnd/normalize/itemoccurrences \
  ./internal/modules/dnd/normalize/locationoccurrences \
  ./internal/modules/dnd/validate/npcoccurrences/... \
  ./internal/modules/dnd/validate/itemoccurrences/... \
  ./internal/modules/dnd/validate/locationoccurrences/...
go test ./internal/cli

Acceptance Criteria

  • Every registry-consuming occurrence path was traced end to end.
  • Model semantics, deterministic identity, and evidence ownership were reviewed as separate concerns.
  • Repeated mechanics were evaluated against all three domain exceptions.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 12: Audit Spells, Scene Chunking, And Scene Descriptions

Goal

Audit spell extraction/catalog behavior and the scene planning/description family, including chunk-map metadata, eligibility, normalization, and reference use.

Required Reading

  • spell, scene-description, and accepted chunk-map integration contracts
  • D&D scene chunker and scene-description extractor/registry/normalizer/ validator packages
  • spell extractor, catalog, normalizer, and validators
  • related prompt assets, schemas, codecs, examples, and CLI integration tests

Work

  1. Trace scene chunking from whole-source prompt input through private response, plan canonicalization, validation, materialized chunks, accepted chunk-map annotations, and output publication.
  2. Trace scene descriptions through combat/non-combat eligibility, extraction, normalization, deterministic scene identity/ranges, validation, and codec.
  3. Trace spell extraction through effective catalog overlays, NPC registry grounding, canonicalization, normalization, validation, checkpoint identity, and retry behavior.
  4. Review the high-complexity scene planFromResponse and spell catalog construction/composition functions for repeated scans, intermediate maps, multiple decoding passes, and missing invariant comments.
  5. Compare these lanes with D&D conventions from Stage 9 and classify justified exceptions.
  6. Verify chunk metadata and references remain context, eligibility, or output annotations according to their owners and do not become fabricated evidence.
  7. Check prompt ordering/cache boundaries and model-visible contextual identity.
  8. Add DND-SCENE-* findings and update the convention matrix, coverage, and validation.

Validation

go test ./internal/modules/dnd/chunk/scenes \
  ./internal/modules/dnd/extract/scenedescriptions \
  ./internal/modules/dnd/normalize/scenedescriptions \
  ./internal/modules/dnd/validate/scenedescriptions/... \
  ./internal/modules/dnd/scenedescriptions/... \
  ./internal/modules/dnd/extract/spells \
  ./internal/modules/dnd/normalize/spells \
  ./internal/modules/dnd/validate/spells/... \
  ./internal/modules/dnd/spells/...

Acceptance Criteria

  • Scene planning, scene description, and spell paths were traced end to end.
  • Catalog, eligibility, chunk-map, and checkpoint semantics were examined.
  • Convention differences are classified rather than normalized mechanically.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 13: Audit Combat Turns And Enemy Events

Goal

Audit the combat family, including scene eligibility, NPC grounding, actor and enemy identity, ordering, normalization, and validation.

Required Reading

  • combat-turn and enemy-event integration contracts and relevant D&D internal documentation
  • combat-turn and enemy-event extractors, domain helpers, codecs, normalizers, validators, assets, and tests
  • scene-description and NPC-registry inputs consumed by these lanes
  • CLI combat integration tests and complete maintained pipeline example

Work

  1. Trace combat-turn extraction from combat-scene eligibility and NPC registry grounding through canonicalization, normalization, validation, codec, and output.
  2. Trace enemy events through contextual grounding, event mapping, engagement validation, ordering, duplicate collapse, and output.
  3. Verify non-combat scenes cannot enter combat extraction, while missing, rejected, or malformed scene metadata fails according to the intended boundary.
  4. Verify actor/enemy names remain contextual and no model-visible opaque ID or registry evidence becomes direct occurrence evidence.
  5. Compare actor grounding, source-reference conversion, ordering, duplicate collapse, invariants, and metadata with each other and with Stage 9 conventions.
  6. Inspect loops and lookup structures relative to turns, events, NPCs, and evidence ranges. Identify repeated work or clearer one-pass mappings.
  7. Verify generated reference provenance, step ordering, checkpoint identity, retry, and rejection behavior in the complete pipeline.
  8. Add DND-COMBAT-* findings and finalize all rows of the D&D convention matrix. Update coverage and validation.

Validation

go test ./internal/modules/dnd/extract/combatturns \
  ./internal/modules/dnd/normalize/combatturns \
  ./internal/modules/dnd/validate/combatturns/... \
  ./internal/modules/dnd/extract/enemyevents \
  ./internal/modules/dnd/normalize/enemyevents \
  ./internal/modules/dnd/validate/enemyevents/... \
  ./internal/modules/dnd/enemyevents
go test ./internal/cli

Acceptance Criteria

  • Both combat lanes were traced from eligibility/reference inputs to durable output.
  • Scene gating, contextual identity, ordering, and engagement rules were explicitly checked.
  • The D&D convention matrix no longer contains unexplained placeholders.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Stage 14: Audit Test Ownership And Comments, Then Synthesize

Goal

Perform the cross-cutting test/comment review, reconcile all stage findings, run final verification, and turn the accumulated audit into a concise, decision-useful final report.

Required Reading

  • the complete accumulated docs/roadmap/audit.md
  • docs/policy/testing.md
  • focused test inventories in all internal documentation
  • production symbols and tests cited by every unresolved finding
  • coverage-matrix rows marked Revisit

Work

  1. Review each finding against the audit-plan quality bar. Remove findings based only on taste, line count, metrics, hypothetical reuse, or an unverified assumption.
  2. Re-read every cited symbol and its callers. Consolidate duplicate symptoms under one root cause, update grouping, and resolve contradictions between stages.
  3. Review test ownership across layers. Add TEST-* findings only for a meaningful unprotected risk, harmful duplication, brittle implementation coupling, or disproportionately elaborate test infrastructure.
  4. Review necessarily complex production symbols identified by prior stages. Add COMMENT-* findings only when the exact invariant and intended comment location can be named.
  5. Review intentional duplication entries. Retain only decisions valuable to a future remediation agent; remove trivial false positives.
  6. Run the final validation suite below. Record failures exactly and determine whether they confirm a finding or are pre-existing/unrelated.
  7. Finalize the finding index ordered by severity, then correctness risk, dependency order, and expected remediation value. Do not renumber stable area IDs.
  8. Write the executive summary with:
    • target snapshot and baseline health;
    • number of High, Medium, and Low findings;
    • principal correctness risks;
    • highest-value refactoring themes;
    • areas where intentional complexity should remain; and
    • recommended independent remediation work sets, without writing an implementation plan.
  9. Complete every coverage row as Reviewed or explain the exact unresolved Revisit. Ensure areas without findings are visible.
  10. Verify the final document contains no implementation-history narrative, secrets, unsupported claims, or instructions to implement unbuilt behavior outside the roadmap.

Validation

go test -count=1 ./...
go test -race ./...
go vet ./...
go build ./cmd/notarius
git diff --check

If the ordinary suite passes, also run:

go test -shuffle=on ./...

Acceptance Criteria

  • Every finding satisfies the evidence contract and is unique at the root-cause level.
  • Test and comment recommendations are specific, risk-based, and non-duplicative.
  • The finding index, executive summary, intentional-complexity section, no-findings section, validation record, and coverage matrix are complete.
  • Final validation results are recorded accurately.
  • The report proposes bounded remediation groups but does not implement them or turn itself into an implementation plan.
  • Only docs/roadmap/audit.md changed.

This stage is suitable for one audit prompt.

Completion And Handoff

After Stage 14, the audit is complete when:

  • docs/roadmap/audit.md is decision-useful without requiring access to stage commentary;
  • the audited production snapshot is unambiguous;
  • every repository area in audit-plan.md has a completed coverage record;
  • findings are evidence-backed, deduplicated, severity-ranked, and grouped;
  • intentional explicit code is distinguished from refactoring candidates;
  • validation results and limitations are visible; and
  • no production or current-behavior file was modified by the audit.

A later planning pass may convert accepted findings into a staged remediation plan. That later pass should choose work by dependency and risk rather than blindly following finding-ID order.

Open Questions

None. The sequence is ready to execute against the production snapshot recorded by Stage 1.