diff --git a/docs/roadmap/audit-findings.md b/docs/roadmap/audit-findings.md new file mode 100644 index 0000000..cb8e08e --- /dev/null +++ b/docs/roadmap/audit-findings.md @@ -0,0 +1,385 @@ +# Codebase Audit Findings + +Status: in progress + +This document is the working ledger and final report for the audit defined by +the [audit plan](audit-plan.md) and [audit sequence](audit-sequence.md). The +audit is investigative: candidate findings below are not remediation changes. + +## Audit Identity And Baseline + +| Field | Value | +| --- | --- | +| Audited revision | `74e2d21de5fb2ada0be5ef3fe9333e0d48ac7fb3` (`Close the completed roadmap documents`) | +| Branch | `main`, attached worktree | +| Initial worktree state | Untracked `docs/roadmap/audit-plan.md` and `docs/roadmap/audit-sequence.md`; no production or test changes | +| Audit date | 2026-08-10 UTC | +| Toolchain | `go version go1.26.5 linux/amd64` | +| Platform | `GOOS=linux`, `GOARCH=amd64` | +| Repository root | `/home/eric/Workspace/narratio` | + +The two initial untracked files are the audit specification supplied for this +run. Adding this ledger and tracking those documents changes documentation +only; all implementation and test evidence remains pinned to the revision +above. If implementation or tests change, affected audit stages must be rerun +and this section must record the new revision. + +### Baseline Commands + +| Command | Result | Wall time | Evidence or limitation | +| --- | --- | --- | --- | +| `go test -count=1 ./...` | pass | 3.34 s | All 23 packages passed; `cmd/narratio` has no test files. | +| `go test -race -count=1 ./...` | fail | 55.65 s | Race in `internal/adapters/whisperx.(*FakeClient).Transcribe` at `fake.go:45`, reached concurrently by `TestTranscribeStageTranscribesPreparedAudio`; candidate `TST-001`. All packages reported before `internal/stage` passed. | +| `go vet ./...` | pass | 0.47 s | No diagnostics. | +| `go build -o "$audit_build_dir/narratio" ./cmd/narratio` | pass | 1.03 s | Built outside the repository in `/tmp/tmp.x11pJL7014`. | +| `go test -coverprofile="$audit_build_dir/coverage.out" ./...` | pass | 11.05 s | Diagnostic coverage only; no percentage is treated as a gate. | + +Coverage ranged from 69.8% (`internal/manifest`) to 100% (`internal/logging`) +among tested packages. `cmd/narratio` reported 0% because it has no tests. The +remaining package results ranged from 70.1% to 88.1%. Stage 12 owns the +risk-based interpretation; these numbers are inventory signals only. + +### Code Graph Freshness And Structural Inventory + +The `narratio` graph was rebuilt in `moderate` mode after the revision was +pinned. Its branch record reports the exact audited HEAD, `main`, and the +repository root above. The index contains 2,407 nodes and 13,181 edges across +224 modeled files: 1,494 functions, 136 methods, 226 structs, 12 interfaces, +and 20 modeled package nodes. The moderate filter excluded documentation, +examples, `.git`, `.codex`, and `cmd/narratio`; the executable entry point was +therefore verified through `go list` and direct inspection instead of graph +evidence. Internal production code is represented at the pinned revision. + +Repository inventory at that revision: + +- 23 Go packages, including `cmd/narratio`; +- 221 tracked Go files and 95 tracked `_test.go` files; +- 278 tracked files total; +- one process entry point, `cmd/narratio/main.go`, delegating to + `internal/app.Execute`; +- 11 canonical stages returned by `internal/stage.All`; and +- 12 modeled interfaces, of which 11 are Narratio boundaries and one is the + private AWS S3 client seam. + +Graph call tracing from `internal/app.Execute` confirms command dispatch into +run, single-stage, clean, and session-helper paths, followed by configuration, +artifact/path, manifest, stage, storage, restore, and cleanup owners. The +production import inventory shows no lower-level package importing +`internal/app`; apparent graph rollups such as `stage -> app`, `adapters -> app`, +and `config -> app` came from test relationships or graph classification and +are rejected as production dependency reversals at this mapping stage. + +### Metric Signals For Later Review + +These are prioritization signals, not findings: + +| Signal | Evidence | Assigned review | +| --- | --- | --- | +| High fan-in | `app.Error` (207), stage `Run` symbols (151), `app.Execute` (108), `stage.sessionPathsForEnv` (105), `manifest.New` (72), `manifest.MarkStageSucceeded` (56), `app.executeStages` (48), `artifacts.S3SessionPrefix` (41), and `artifacts.SessionWorkDirForCampaign` (37) | Owning behavior stages, then Stage 11 | +| High complexity | `app.executeStages` cyclomatic 54/cognitive 96; `previouscache.BuildPlan` 22/38; `analyzeStage.Run` 20/27; `audita.NewSubprocessRunner` 17/25; `app.SessionInit` 20/21 | Stages 2, 5, 7, 10, then 11 | +| Exact similarity | `app.Analyze`/`app.Publish`, `manifest.Load`/`LoadRun`, `manifest.Create`/`CreateRun`, adapter constructors, and Seriatim fake methods | Owning behavior stages, then Stage 11 | +| Test-heavy hotspot noise | Several test functions and fakes rank highly in transitive-depth and fan-in results | Stage 12; do not infer production risk from the metric | + +### Automation And Fixture Inventory + +- `.woodpecker/release.yml` is tag-only release automation. It cross-builds + Linux, macOS, and Windows binaries with Go 1.25, then publishes release + assets. It does not run tests, race tests, vet, or example validation. +- `examples/` contains 19 maintained files: pipeline, campaign, session, + template, stable-input, and placeholder-audio fixtures. Configuration tests + are documented as their validation owner. +- No fuzz tests, golden files, golden-update switches, opt-in/live test tags, or + `go:generate` test mechanisms were found. +- Platform build constraints exist for the native no-replace directory tests + and unsupported-platform fallback in `internal/fileops`. + +## Execution Coverage Ledger + +| Stage | Status | Evidence and result | +| --- | --- | --- | +| 0: baseline | complete | Revision/environment pinned; graph refreshed; inventories and every prescribed baseline command recorded. `TST-001` owns the non-blocking race limitation. | +| 1: contract and boundary map | complete | Canonical contracts and focused internal docs read; ownership, stage-contract, lifecycle, scenario, area, and preliminary risk-to-test matrices recorded below. | +| 2: runner and manifest | not_started | Assigned lifecycle and dual-ledger questions below. | +| 3: paths and filesystem | not_started | Assigned path, lock, artifact, and mutation questions below. | +| 4: publish and cleanup | not_started | Assigned remote commit and cleanup scenarios below. | +| 5: restore and previous state | not_started | Assigned restore and previous-cache scenarios below. | +| 6: configuration and composition | not_started | Assigned configuration, CLI composition, and process-boundary areas below. | +| 7: adapters and shared support | not_started | Assigned external-boundary and cancellation areas below. | +| 8: ordinary stages | not_started | Assigned prepare/transcript behavior and disabled-outcome questions below. | +| 9: extraction | not_started | Assigned extraction promotion, provenance, and resume scenario below. | +| 10: analyze and dependencies | not_started | Assigned artifact dependency/source and selection scenario below. | +| 11: maintainability | not_started | Seeded by graph complexity, similarity, and fan-in signals only. | +| 12: test policy | not_started | Seeded by intended owners and baseline execution observations. | +| 13: synthesis | not_started | No final ranking or accepted-risk decisions yet. | + +## Area Coverage And Ownership + +Every area in the audit plan has a primary execution owner. `assigned` means it +has been mapped but not behaviorally audited. + +| Inspection area | Canonical implementation owner | Primary audit stage | Status | +| --- | --- | --- | --- | +| Process and application boundary | `cmd/narratio`, `internal/app` | 6 (runner lifecycle portions in 2; publish/restore portions in 4-5) | assigned | +| Stage registry and runner | `internal/stage`, `internal/app` | 2 | assigned | +| Configuration | `internal/config` | 6 | assigned | +| Prepare and audio | `internal/stage`, `internal/audio`, `internal/previouscache` | 8 | assigned | +| Transcript stages | `internal/stage` plus tool adapters | 8 | assigned | +| Extraction | `internal/stage`, Notarius adapter, `internal/fileops` | 9 | assigned | +| Analyze and artifact dependencies | `internal/stage`, `internal/artifacts`, `internal/artifactpolicy` | 10 | assigned | +| Publish and cleanup | `internal/stage`, `internal/app` | 4 | assigned | +| Manifest state | `internal/manifest`, transition policy in `internal/app` | 2 | assigned | +| Artifacts, paths, and policy | `internal/artifacts`, `internal/artifactpolicy`, `internal/pathsafe` | 3 (resolution consumption revisited in 10) | assigned | +| Restore | `internal/app`, `internal/artifacts`, `internal/previouscache`, `internal/audio` | 5 | assigned | +| File operations | `internal/fileops`, `internal/pathsafe`, local artifact store | 3 (promotion vertical slice in 9) | assigned | +| External adapters and storage | `internal/adapters`, `internal/audio` | 7 | assigned | +| Shared models and diagnostics | `internal/artifactmodel`, `internal/contracts`, `internal/logging` | 7 (maintainability revisited in 11) | assigned | +| Tests, examples, and automation | package test owners, `examples/`, `.woodpecker/` | 12 | assigned | + +## Package And Interface Ownership Map + +| Package | Owned contract or policy | Important boundaries | Audit owner | +| --- | --- | --- | --- | +| `cmd/narratio` | Process entry and exit; CLI delegates behavior to app | `main -> app.Execute` | 6 | +| `internal/app` | Command dispatch, composition, locking, planning, lifecycle, restore, cleanup, reporting | `Execute`, `executeStages`; consumes stage/artifact/manifest/adapter contracts | 2, 4-6 | +| `internal/config` | Strict discovery, defaults, resolve, template, and validation rules | Config models and load/resolve/validate functions | 6 | +| `internal/stage` | Canonical order and stage behavior | `Stage`, `ResumeValidator`, `Env`; adapter interfaces are injected | 2, 4, 8-10 | +| `internal/manifest` | Session/run models, transitions, validation, atomic persistence | `Store`; transition methods record but do not choose policy | 2 | +| `internal/artifacts` | Artifact identity/resolution, paths/keys, local store, remote current-state mechanics | `Store`; consumes explicit storage keys | 3, 5, 10 | +| `internal/artifactpolicy` | Configured source/destination identity and safety policy | Narrow validators used by config, artifacts, app, and stages | 3, 10 | +| `internal/artifactmodel` | Shared serialized artifact, contract, and provenance models | Data contract only | 3, 7 | +| `internal/pathsafe` | Confined relative path and destination mechanics | Narrow validation helpers; no stage policy | 3 | +| `internal/fileops` | Atomic files, copies, hashing, no-replace directory promotion | Filesystem mechanics receive explicit paths | 3, 9 | +| `internal/previouscache` | Deterministic previous-session requirement planning/materialization | Uses explicit object-store and artifact contracts | 5, 8, 10 | +| `internal/audio` | S3 audio spool/cache materialization | Uses `storage.ObjectStore`; no stage ordering | 5, 8 | +| `internal/contracts` | Bounds and shared JSON validation models | Data contract only | 7, 8 | +| `internal/logging` | Shared logger construction | `slog` composition | 7, 11 | +| `internal/adapters/whisperx` | WhisperX HTTP protocol | `Client` | 7 | +| `internal/adapters/seriatim` | Merge/normalize/trim/render subprocess protocol | `Runner` | 7 | +| `internal/adapters/audita` | Audita subprocess protocol | `Runner` | 7 | +| `internal/adapters/scriptorium` | Scriptorium run/render subprocess protocol | `Runner` | 7 | +| `internal/adapters/notarius` | Notarius invocation and receipt boundary | `Runner` | 7 (vertical behavior in 9) | +| `internal/adapters/notify` | Notification transport | `Sender` | 7 | +| `internal/adapters/storage` | Explicit bucket-relative object-store operations and S3 mechanics | `ObjectStore`; private `s3API` test seam | 7 | +| `internal/adapters/subprocess` | Shared bounded subprocess/config/log mechanics | Concrete helper package, not stage policy | 7 | + +The graph reported no inbound production callers of `Stage.Declares`; text +search found definitions and test stubs but no production invocation. This +reduces the current impact of `ARC-001` but makes the interface's intended owner +and future use an explicit question rather than resolving the mismatch. + +## Stage Contract Matrix + +The table separates declared/static contracts from dynamic behavior. All +executed stages use the runner's session/run transitions. Unless noted, a +successful result records returned outputs, diagnostics, generated +configuration, and metadata; a different effective executed outcome can stale +succeeded downstream work, while force pre-stales succeeded downstream work. + +| Order and stage | Inputs and outputs | Configuration and adapters | Skip/resume behavior | Materialization and manifest effects | +| --- | --- | --- | --- | --- | +| 1 `prepare` | Config, stable inputs, one audio mode, optional previous requirements -> canonical `inputs/**`, `audio/**`, optional `previous/**`, `manifest.inputs` | All resolved config; storage for S3/current previous state; audio/artifact/previous-cache services | No stage-specific resume validator or explicit self-skip | Writes canonical session inputs and deterministic input records; unlike processing stages, `Declares` labels produced canonical files as inputs. | +| 2 `transcribe` | Prepared FLAC files -> raw per-speaker JSON | WhisperX URL/language/retry/timeout/concurrency; `whisperx.Client` | Ordinary succeeded-record skip; no validator/self-skip | Bounded concurrent run-local writes, validation, then canonical transcript materialization. | +| 3 `merge` | Raw transcripts, speakers, autocorrect -> base transcript, optional report | Seriatim merge fields; `seriatim.Runner` | Ordinary succeeded-record skip | Normalized scratch inputs and run-local results validate before canonical transcript/report materialization. | +| 4 `polish` | Base transcript, glossary -> polished transcript, optional report | Audita fields/credential reference; `audita.Runner` | Ordinary succeeded-record skip | Run-local output, report, logs, and generated config; validates before canonical materialization. | +| 5 `normalize` | Polished transcript -> final transcript, optional report | Normalize plus Seriatim fields; `seriatim.Runner` | Ordinary succeeded-record skip | Manifest-first input; run-local validation then configurable canonical output/report. | +| 6 `trim` | Final transcript -> final-trimmed transcript and, when enabled, bounds | Trim, bounds, Scriptorium, and Seriatim fields; both runners when enabled | Disabled trim copies input and still succeeds; no explicit self-skip or resume validator | Run-local bounds/trim result validates then materializes; debug render is diagnostic, not output. | +| 7 `extract` | Final-trimmed source -> immutable index and configured lane outputs | Notarius executable/config/pipeline/timeout/output contracts; `notarius.Runner` | Disabled is explicit `notarius_disabled` self-skip; only current `ResumeValidator`; obsolete reruns, unsafe validation errors | Validates run-local receipt/bundle completely, promotes to unique immutable bundle, records checksums/contracts/provenance; identical repeated self-skip is stable. | +| 8 `render` | Final and final-trimmed JSON -> two Markdown transcripts | Render and Seriatim fields; `seriatim.Runner` | Disabled returns a zero-disposition result with skip metadata, therefore runner-level success rather than explicit self-skip; no validator | Run-local text validates non-empty before canonical materialization when enabled. | +| 9 `analyze` | Dynamic built-in, prepared, extraction, configured, and previous sources -> selected configured artifact outputs | Scriptorium artifact graph/selection; `scriptorium.Runner` | Missing config or no executable artifacts returns success with skip metadata; no validator | Topological run-local generation/reuse, validation, canonical outputs, deterministic metadata; static `Declares` omits dynamic outputs and several input families. | +| 10 `publish` | Session/run state, selected output rules, locks, previous cache -> remote run/output/current objects | Publish/storage/selection fields; `storage.ObjectStore` | Disabled publish or run upload returns success with skip metadata; force cannot bypass locks; no validator | Deterministic uploads; `current/manifest.json` before `current/run_id.txt`; commit metadata gates cleanup. Static prerequisites omit extract because disabled extraction is valid and lane resolution enforces required extraction state when selected. | +| 11 `notify` | No implemented persisted pipeline input/output | Optional `notify.Sender`; default no-op | Ordinary succeeded-record skip; no explicit self-skip or validator | Placeholder metadata and optional notification call; no returned output. `Declares` nevertheless advertises placeholder input/output paths. | + +Configuration, adapters, skip policy, and dynamic outputs are not represented +by `IODecl`; their current canonical owners are the focused stage, +configuration, and integration contracts. Whether `IODecl` should remain a +partial display type or become an enforceable declaration is deferred as +`ARC-001`. + +## Lifecycle Matrix + +This is the intended contract map to be walked through both durable ledgers in +Stage 2. Cells marked `unknown` are not treated as implementation conclusions. + +| Outcome | Session manifest intent | Invocation manifest intent | Downstream and next-invocation intent | +| --- | --- | --- | --- | +| First run | Pending/non-succeeded stage becomes running, then succeeded/failed/skipped; executing clears older result payload first | New run record; action `run`; terminal status records this invocation | Success enables later stages; failure stops current execution and an effective outcome change may stale succeeded downstream records. | +| Already-succeeded skip | Existing succeeded session record and payload remain unchanged, subject to resume validation | Action/status record a skip and stable reason for this invocation | Reusable result remains authoritative; pipeline continues. | +| Explicit self-skip | Session stage becomes skipped, clears older result payload, and may record bounded current skip details | Action was `run`, outcome is skipped with reason | Reconsidered later; a changed effective upstream outcome stales succeeded downstream work; identical extraction disabled skip is stable. | +| Failure | Current stage becomes failed with error; current output/log/config/metadata payload is cleared | Action `run`, failed outcome and overall failed run | Current execution stops; affected succeeded downstream work is intended to stale; later invocation reruns non-succeeded stages. | +| Interruption | Model admits `interrupted`, but no production transition reference has yet been found; process death can leave persisted `running` state | A process death can leave the run non-terminal; exact recovery semantics are unknown | CLI promises continuation of interrupted/partial sessions because non-succeeded stages run; explicit status normalization is `RSK-001` for Stage 2. | +| Forced replacement | Target execution starts fresh; succeeded downstream records are pre-marked stale; current target payload clears on running | Force flag and `run` action recorded | Replacement result determines later execution; locks and safety policy remain authoritative. | +| Non-resumable success | Prior success becomes stale while retaining details long enough for diagnosis/validation, then running clears them | Current invocation records execution after validation rejects skip | Obsolete result reruns; unsafe inability to decide stops without silently replacing current success. | +| Successful rerun | Target becomes succeeded with only new outputs/diagnostics/config/metadata | Current invocation records its own new success; earlier run manifests remain immutable | Changed effective outcome stales succeeded downstream work; identical effective outcome should avoid unnecessary invalidation. | + +Stage 2 must separately verify save failures before and after each session/run +transition, first- and last-stage behavior, and which fields are retained in +historical invocation records. + +## Cross-Boundary Scenario Assignments + +| Scenario | Primary audit stage | Supporting packages and focused tests | +| --- | --- | --- | +| 1. Success becomes non-resumable, rerun fails, later reuse decision | 2 | `internal/app`, `internal/manifest`, `stage.ResumeValidator`; `runner_test.go`, `extract_lifecycle_test.go`, manifest transition tests | +| 2. Forced/changed upstream outcome with succeeded, self-skipped, disabled downstream | 2 | `internal/app`, `internal/stage`, `internal/manifest`; run-control, runner, extraction-lifecycle tests; disabled-stage detail revisited in 8 | +| 3. Extraction bundle followed by configuration/transitive-input change | 9 | Extract/resume, Notarius adapter, artifacts/fileops tests; downstream resolution revisited in 10 | +| 4. Published/restored/prepared previous state consumed locally by analyze | 5 | `internal/app`, `internal/previouscache`, `internal/audio`, `internal/artifacts`; restore/prepare tests; analyze consumption revisited in 10 | +| 5. Publish failure at every upload boundary, then status/restore/retry | 4 | Publish stage, storage fake/adapter, app status/restore; publish and operator-helper tests; restore interpretation revisited in 5 | +| 6. Restore identical/conflict/unsafe/cache/pre-manifest-install cases | 5 | Restore discovery/plan/execute/report, artifacts, previouscache, audio; restore test suite | +| 7. Cleanup after skipped/failed/locked/partial/committed publish | 4 | Publish metadata, post-publish cleanup, cleanup targets, pathsafe; publish/cleanup tests | +| 8. Cancellation through workers, HTTP, subprocess, storage, manifests | 7 | Adapter and subprocess tests; transcribe/stage tests in 8; runner reporting in 2 | +| 9. Disabled/unselected/reused/generated/extraction/previous source then publish filtering | 10 | Analyze, artifact catalog/resolver/policy, publish tests; config ownership in 6 and publish result in 4 | +| 10. Concurrent same-session invocation and lock cleanup failures | 3 | Runner lock lifetime in 2; local artifact store, path/file cleanup and lock tests in 3 | + +## Preliminary Risk-To-Test Matrix + +This matrix identifies intended owners only. It makes no sufficiency judgment. + +| Architectural invariant or risk | Implementation owner | Intended test owner | +| --- | --- | --- | +| One deterministic canonical stage order | `internal/stage`, planner in `internal/app` | `internal/app/planner_test.go`, narrow registry tests | +| Session manifest is cross-invocation authority; run manifest is immutable invocation audit | `internal/app`, `internal/manifest` | Manifest transition tests plus assembled runner/run-stage tests | +| First run, skip, self-skip, failure, force, invalidation, and rerun transitions | `internal/app`, `internal/manifest` | App lifecycle tests as primary; manifest helpers own field mutation | +| Obsolete versus unsafe resume validation | Stage-specific `ResumeValidator`, runner | Extract resume tests plus runner integration tests | +| Run-local validation before canonical materialization | Individual stages and `run_local.go` | Focused stage package behavioral tests; fileops owns atomic mechanism | +| Strict config, defaults, identity, and cross-field validation | `internal/config` | Config package tests; example load/validation test samples assembly | +| Canonical path/key ownership and traversal confinement | Artifacts, artifactpolicy, pathsafe | Owning package tests; app/stage tests only for assembled policy | +| Immutable extraction promotion and provenance/checksum validation | Extract, fileops, artifacts, Notarius adapter | Fileops mechanism, extract behavior, artifact hydration, adapter contract tests | +| Deterministic artifact dependency and source resolution | Artifacts, artifactpolicy, analyze | Artifact/package tests and analyze package behavior tests | +| Previous-session consumption remains local in analyze | Previouscache/prepare/artifacts/analyze | Previouscache and prepare tests; one analyze boundary test for no remote call | +| Remote current pointer is publish's final commit point | Publish stage | Publish tests with stateful object-store fake; storage tests own transport only | +| Restore is confined, deterministic, conflict-safe, and installs manifest last | Restore app modules, artifacts/previouscache/audio | Restore plan/execution/workflow tests plus low-level path/file tests | +| Cleanup requires explicit scope and committed publish metadata | App cleanup modules, pathsafe | Cleanup-target and post-publish integration tests | +| Session single-writer lock and safe release | Local artifact store, app lifetime | Artifact local-store tests plus assembled concurrent runner tests | +| Adapter cancellation, error adaptation, and resource closure | Each adapter and shared subprocess package | Focused adapter boundary tests; stage tests sample propagation | +| Bounded deterministic transcription concurrency | Transcribe stage and WhisperX client | Stage concurrency/result-order tests; HTTP adapter retry/cancel tests | +| Secrets never persist or appear in diagnostics | Config/app composition and each adapter/logging boundary | Owning config/adapter tests plus selected assembled redaction checks | +| Default suite remains deterministic, offline, and credential-free | Every package; automation | Stage 12 repository-wide execution and test-policy audit | + +## Candidate Register + +No candidate is confirmed by Stage 1. Later owning stages must inspect the +implementation, focused tests, canonical contract, realistic scenario, and +callers before promoting or rejecting it. + +### `ARC-001`: `IODecl` is not a complete or consistently classified stage contract + +- Category: architectural boundary/ownership candidate. +- Evidence: `prepare.Declares` lists files it produces under `Inputs`; + `analyze.Declares` omits dynamic input families and has no outputs; + `publish.Declares` exposes only the manifest; and `notify.Declares` advertises + placeholder paths although its result has no persisted output. No production + caller of `Declares` was found. +- Contract tension: architecture says every stage declares required inputs, + produced output state, configuration, adapters, lifecycle, and failure + behavior; the Go interface declares only partial static artifacts. +- Realistic risk: a future planner, validator, or operator view could treat the + interface as authoritative and make incorrect dependency or readiness + decisions. Current likelihood appears low because the method has no + production caller. +- Confirmation owners: Stages 8-10 for dynamic contracts, then Stage 11 for + interface purpose/simplification. Smallest plausible outcome may be clearer + naming/documentation, a complete contract, or removal; do not choose yet. + +### `ARC-002`: disabled-stage “skip” terminology spans two different durable outcomes + +- Category: architectural/lifecycle ownership candidate. +- Evidence: production use of `StageDispositionSkipped` was found only in + extraction. Disabled render, absent/no-op analyze, and disabled publish return + zero-disposition results with skip metadata, which the runner treats as + success. Focused and operator docs use “skip” for several of these cases, + while manifest docs reserve self-skip for a durable skipped state. +- Realistic risk: maintainers or operator features may assume all disabled + outcomes clear state, are reconsidered, and invalidate downstream work in the + same way. Conversely, changing them to explicit self-skip could break valid + pipeline continuation or cleanup semantics. +- Confirmation owners: Stage 2 for runner truth table, Stage 4 for publish and + cleanup, Stage 8 for ordinary disabled stages. Treat wording and behavior as + unresolved until those flows are traced. + +### `RSK-001`: interruption has a documented state but no mapped production transition + +- Category: correctness/operational risk candidate. +- Evidence: `manifest.StatusInterrupted` is admitted and external docs promise + continuation of interrupted sessions, but graph-augmented code search found + the constant only in its declaration and an artifact rejection test. No + production transition to it was found during mapping. +- Realistic risk: a killed process may leave session/run records as `running`, + producing confusing status or dual-ledger interpretation even though the + planner reruns all non-succeeded stages. +- Confirmation owner: Stage 2 must trace load normalization, process failure + boundaries, status reporting, and next-invocation behavior before deciding + whether this is a defect, compatibility state, or unused model value. + +### `TST-001`: full race baseline fails in the concurrent transcribe test + +- Category: test-suite execution candidate. +- Evidence: the race detector reported concurrent slice access in + `internal/adapters/whisperx/fake.go:45` from transcribe workers in + `TestTranscribeStageTranscribesPreparedAudio`. +- Observed impact: the canonical full race command exits nonzero, weakening its + signal for other packages. The report currently points to a test fake, not a + production data race. +- Confirmation owners: Stage 8 should inspect the worker/fake contract; Stage + 12 should classify suite impact and the smallest durable fix. Do not change + the fake during this investigative stage. + +## Candidate Classification Log + +| Candidate signal | Classification | Reason | +| --- | --- | --- | +| Graph rollups `stage -> app`, `adapters -> app`, `config -> app` | rejected as a production reversal at Stage 1 | `go list` production imports contain no lower-level import of `internal/app`; graph connections include tests and ambiguous package grouping. Reopen only with a concrete production edge. | +| Similar wrapper/manifest/adapter functions | deferred metric signals, not findings | Similarity alone does not establish duplicated policy; owning behavior stages must first establish contracts. | +| Coverage percentages | deferred diagnostic signals, not findings | Stage 12 must reason from risk and test ownership, not a numeric target. | + +## Unresolved Questions And Follow-Up + +- Does loading a persisted `running` stage or run normalize it to + `interrupted`, or is `interrupted` only a compatibility value? +- What exact session/run disagreement states are possible when either save + fails at each transition boundary? +- Are disabled render/analyze/publish outcomes intentionally successful so + pipeline continuation and optional outputs work, and do all operator views + describe that distinction accurately? +- Is `IODecl` intended only for display/tests, or should it own enforceable + dependency declarations? +- Does publish's omission of extract from its static prerequisite list combine + safely with every configured extraction output rule and disabled extraction? +- Which native CI runner limitations explain the absence of validation jobs in + the tag-only release workflow? Stage 12 owns the automation conclusion. + +No accepted risks or final audit conclusions are recorded yet. + +## Completed-Stage Evidence + +### Stage 0 + +- Contracts and records: development guide, audit plan and sequence, all policy + documents, repository/branch/toolchain state. +- Graph evidence: refreshed moderate index at exact HEAD; architecture, + interface, complexity, similarity, fan-in, and `Execute` call trace queries. +- Commands: every baseline command listed above; Go/package/file/test and + automation inventories. +- Candidates: `TST-001`; metric signals assigned to later owners. +- Explicit no-finding conclusion: no production dependency reversal into + `internal/app` was found in the package import inventory. +- Limitation disposition: the graph excludes the executable entry point, which + was verified directly; the race failure is owned by Stages 8 and 12 and does + not prevent read-only audit work. + +### Stage 1 + +- Contracts reviewed: architecture, testing and documentation policy; internal + overview and every focused internal document; CLI, configuration, + operations, and every integration contract. +- Code/evidence reviewed: canonical registry and stage declarations; all + modeled interfaces; production import graph; application dispatch trace; + explicit self-skip usages; interrupted-state usages; focused test ownership + references. +- Outputs: package/interface ownership, area coverage, stage contract, + lifecycle, cross-boundary scenario, and preliminary risk-to-test matrices. +- Candidates: `ARC-001`, `ARC-002`, `RSK-001`; no candidate was confirmed from + mapping evidence alone. +- Explicit no-finding conclusion: the canonical stage order agrees across the + registry, internal overview, CLI, and operations contract. +- Follow-up: all unresolved behavior has a named owner in Stages 2-12; every + area and invariant has an implementation owner and intended test owner. diff --git a/docs/roadmap/audit-plan.md b/docs/roadmap/audit-plan.md new file mode 100644 index 0000000..71010b5 --- /dev/null +++ b/docs/roadmap/audit-plan.md @@ -0,0 +1,332 @@ +# Codebase Audit Plan + +Status: proposed + +## Purpose + +This audit will evaluate Narratio for correctness, efficiency, maintainability, +and test-suite value. It will identify defects and credible risks, duplicated or +near-duplicated behavior, code that can be made smaller or more idiomatic, and +complex code whose remaining invariants need focused explanation. + +The audit is investigative. It should produce evidence-backed findings and a +prioritized remediation backlog, not make opportunistic production changes as +it proceeds. The [Audit Sequence](audit-sequence.md) assigns this scope to +concrete execution stages. + +## Authoritative Baseline + +Review implemented behavior against its canonical owner rather than treating +the current implementation or tests as the specification: + +- [Architecture](../policy/architecture.md) for system boundaries, dependency + direction, state and path ownership, safety properties, and pipeline + invariants; +- [Internal Overview](../internal/overview.md) and its focused internal + documents for implemented ownership and mechanics; +- [Testing Policy](../policy/testing.md) for risk-based sufficiency, durable + boundaries, test-double guidance, and test lifecycle decisions; +- the [CLI](../cli.md), [Configuration](../config.md), + [Operations](../operations.md), and [integration contracts](../integrations/) + for externally observable behavior; and +- the [Documentation Policy](../policy/documentation.md) for canonical ownership + and the distinction between current and proposed behavior. + +Where code, tests, and documentation disagree, record the disagreement. Do not +assume which one is wrong until the canonical contract and caller expectations +have been traced. + +## Audit Principles + +1. Review correctness before cleanup. A shorter implementation is not an + improvement if it weakens a state transition, safety check, or external + contract. +2. Trace behavior across boundaries. Narratio's most important properties often + emerge from the interaction of application orchestration, stages, manifests, + artifact resolution, filesystem operations, and adapters. +3. Distinguish repeated syntax from repeated policy. Extract a helper only when + the behavior has one stable owner and the shared abstraction makes that + ownership clearer. Similar stage code may be intentionally explicit. +4. Prefer narrow, idiomatic Go over generic frameworks. In particular, proposed + refactors must preserve the explicit canonical stage sequence and must not + turn Narratio into a workflow engine or a second configuration system for + downstream tools. +5. Optimize credible work. Flag repeated I/O, hashing, serialization, remote + calls, subprocess work, allocation, or poor asymptotic behavior when the + relevant path can matter. Require a benchmark or workload argument for + performance changes whose benefit is not evident. +6. Treat comments as explanations of intent. Recommend comments for invariants, + ordering constraints, non-obvious failure policy, or security reasoning—not + as narration of ordinary Go or a substitute for simplifying code. +7. Judge tests as a suite. A test can be locally reasonable and still add no + marginal protection, while a compact test can be inadequate for a + consequential cross-component failure. + +## Evidence And Finding Standard + +Begin from a cleanly identified revision and record toolchain and platform +assumptions. Use the code knowledge graph to find ownership, callers, callees, +similarity candidates, high-complexity functions, and weakly protected +boundaries. Confirm every candidate by reading the implementation, its focused +tests, and the applicable contract. Text search and static analysis supplement +the graph for literals, configuration, generated files, and patterns that are +not modeled reliably. + +Each finding should record: + +- category: correctness defect, correctness risk, duplication, simplification, + efficiency, architectural boundary, comment/clarity, or test-suite issue; +- source locations and the affected contract or invariant; +- concrete evidence and a realistic failure or maintenance scenario; +- impact, likelihood, confidence, and estimated remediation scope separately; +- the smallest plausible improvement and its intended owner; +- tests that already protect the behavior, tests that should change or be + added, and tests that may become redundant; and +- dependencies on, or conflicts with, other findings. + +Do not report a metric alone as a finding. Complexity, similarity, coverage, +fan-in, file size, and test count are prioritization signals that require manual +confirmation. Consolidate findings that share one root cause. + +## Cross-Cutting Review Lenses + +### Correctness And Pipeline Semantics + +Construct an explicit lifecycle matrix for every stage outcome: first run, +already-succeeded skip, self-skip, failure, interruption, forced replacement, +non-resumable result, and successful rerun. Trace how each outcome changes the +session manifest, invocation manifest, downstream stage state, artifacts, +diagnostics, and cleanup eligibility. + +Across the pipeline, verify: + +- the registry exposes one deterministic canonical order; +- each stage's declared inputs, outputs, configuration, adapters, and manifest + effects agree with its implementation and focused documentation; +- inputs are resolved through manifest and artifact contracts rather than + incidental directory contents; +- run-local outputs are fully validated before canonical materialization; +- failure, cancellation, or process interruption cannot advertise partial work + as successful; +- force and changed outcomes invalidate exactly the intended succeeded + downstream work; +- repeated execution is idempotent where promised, and ordering is stable + wherever maps, directory reads, remote listings, or dependency graphs are + involved; +- session, campaign, run, source, checksum, contract, and external provenance + identities cannot be confused across runs; and +- errors preserve useful causes and do not expose secrets or private content. + +Use fault-oriented reasoning at durability boundaries: fail immediately before +and after manifest saves, canonical renames, external process completion, +uploads, current-manifest publication, the current-run commit marker, restore +manifest installation, and cleanup. Determine which state is authoritative and +whether the next invocation recovers safely. + +### Duplication And Helper Ownership + +Search for exact and semantic duplication in production and tests, including: + +- repeated stage setup, input resolution, output validation, run-local + materialization, metadata construction, and error adaptation; +- repeated manifest create/load/save and session/run transition handling; +- repeated adapter construction, timeout parsing, command execution, generated + configuration, log handling, and output checks; +- repeated source-ID, destination, remote-key, and path validation policy; +- repeated sorting, deduplication, checksum, copy, and atomic-write mechanics; + and +- repeated test fixtures and assertions that encode the same policy at several + layers. + +For each candidate, decide whether it is coincidental similarity, a repeated +mechanism, or duplicated policy. Recommend extraction only when the helper can +have a clear package owner, a narrow contract, and callers that become easier +to understand. Prefer an unexported local helper when sharing is package-local. +Do not create a broad utility package, force unlike stage results into one data +model, or move policy into storage/file-operation helpers. + +Initial similarity and complexity signals should seed, but not predetermine, +inspection of the single-stage command wrappers, session/run manifest +persistence pairs, adapter constructors, Scriptorium operations, stage fakes, +and common stage materialization paths. + +### Simplification, Go Idioms, And Efficiency + +Review long or branch-heavy functions for separable decisions, state +transitions, or data transformations. Pay particular attention to orchestration, +configuration validation, artifact dependency resolution, resume verification, +restore/previous-cache planning, and analyze/publish selection logic. A useful +refactor should reduce cognitive load while leaving the important ordering +visible. + +Check for: + +- unnecessary nesting, defensive branches made unreachable by earlier + validation, repeated normalization, and overly wide parameter lists; +- interfaces defined for hypothetical extensibility rather than a demonstrated + consumer boundary; +- manual slice, map, string, error, and filesystem logic with a clearer standard + library form; +- incorrect or inconsistent `errors.Is`/`errors.As`, wrapping, context + propagation, deferred cleanup, response-body closure, process waiting, and + goroutine/channel ownership; +- redundant filesystem scans, `stat`/checksum passes, whole-file buffering, + copying, YAML/JSON round trips, sorting, remote listings, downloads, uploads, + or adapter initialization; +- linear searches nested in loops and repeated dependency or artifact lookup + that should use an indexed map or a single planning pass; +- unbounded concurrency, leaked work after cancellation, serialized independent + work, and nondeterministic result collection; and +- obsolete dependencies, portability assumptions, and platform-sensitive path + or atomic-rename behavior. + +Keep correctness and diagnosability ahead of micro-optimization. When a simpler +algorithm changes performance characteristics, specify the representative +input size and validation method. + +### Comments And Local Explanation + +Review high fan-in, high-complexity, security-sensitive, and commit-boundary +code after likely simplifications have been identified. Add a comment +recommendation when a maintainer needs to know why: + +- state transitions or persistence operations occur in a specific order; +- a stale record intentionally retains data while another transition clears it; +- a path is checked more than once to resist traversal, symlink replacement, or + time-of-check/time-of-use hazards; +- an artifact is accepted only with particular manifest, checksum, contract, or + provenance evidence; +- a partial operation is intentionally not rolled back; +- a remote pointer or local manifest must be installed last; or +- concurrency, cancellation, compatibility, or downstream-tool behavior makes + an apparently simpler approach unsafe. + +Prefer a named helper, typed state, or smaller control flow when that removes the +need for explanation. Check existing comments for stale claims as well as +missing rationale. + +### Test Suite Against The Canonical Policy + +Build a risk-to-test matrix rather than auditing tests file by file in +isolation. For each important behavior, identify its proper owner—parser, +validator, domain package, adapter, orchestrator, CLI, integration, or end to +end—and identify all tests that claim to protect it. + +Evaluate: + +- protection of data integrity, destructive operations, compatibility, + security, concurrency, idempotency, recovery, and partial failure; +- manifest transitions, force/invalidation, resume validation, atomic + materialization, publish commit order, restore install order, and cleanup + gates as assembled behaviors; +- realistic HTTP, subprocess, filesystem, and object-store boundary behavior, + including cancellation and malformed responses; +- whether higher-level tests intentionally sample lower-level behavior or + redundantly reproduce its full policy; +- whether tests assert durable outcomes or private constants, exact error text, + incidental paths, call choreography, or oversized snapshots; +- whether real fast collaborators could replace elaborate doubles, and whether + stateful fakes are realistic enough for the risk they protect; +- fixture/helper duplication, oversized test cases, and setup that obscures the + behavior under test without introducing a heavyweight test framework; +- deterministic, offline, credential-free, order-independent execution and + safe handling of environment and process-global state; +- focused fuzz candidates in parsing, normalization, source IDs, remote/local + path mapping, manifest decoding, and configuration boundaries; and +- the presence and value of a small number of representative assembled + workflows. + +Use coverage only to locate unexpectedly weak consequential branches. Also +inspect packages with extensive coverage for redundant tests and refactoring +friction. For every proposed addition, deletion, or consolidation, state the +realistic defect and marginal confidence involved. + +The audit baseline should include the repository's canonical commands plus +targeted diagnostic runs where supported: + +```sh +go test ./... +go test -race ./... +go vet ./... +go build ./cmd/narratio +``` + +Use focused repeated or shuffled runs to investigate state leakage and +flakiness, and collect package/branch coverage for diagnosis. Review continuous +integration to determine whether the appropriate offline validation is enforced; +do not turn coverage percentage into a gate merely for this audit. + +## Area-By-Area Inspection Map + +| Area | Primary locations | What to inspect | +| --- | --- | --- | +| Process and application boundary | `cmd/narratio`, `internal/app` | Command dispatch, configuration selection, production composition, secret loading, lock lifetime, object-store initialization, context/error propagation, and separation of CLI reporting from orchestration policy. Review operator commands for consistent current-state authority and shared read-only mechanics. | +| Stage registry and runner | `internal/stage/placeholders.go`, `internal/stage/stage.go`, `internal/app/planner.go`, `internal/app/runner.go`, `internal/app/run_stage.go` | Canonical order, action decisions, resume/force/self-skip/failure transitions, downstream invalidation, session/run manifest consistency, resource lifecycle, cleanup triggering, and opportunities to decompose the runner without hiding its state machine. | +| Configuration | `internal/config` | Strict decoding, discovery and precedence, centralized defaults, normalization, templating, validation order, unknown fields, empty-value behavior, secret references, cross-field constraints, path confinement, deterministic errors, duplicated validator policy, and compatibility with maintained examples. | +| Prepare and audio | `internal/stage/prepare.go`, `internal/audio`, `internal/previouscache` | Local/S3 exclusivity, cache and spool identity, partial downloads, checksum/reuse policy, previous-session required/optional planning, deterministic input records, clearing semantics, traversal safety, and avoiding repeated remote or filesystem work. | +| Transcript stages | `internal/stage/transcribe.go`, `merge.go`, `polish.go`, `normalize.go`, `trim.go`, `render.go` | Contract parity across similar stages, bounded concurrency and cancellation, deterministic speaker/input ordering, run-local validation and canonical promotion, report/diagnostic classification, disabled behavior, and narrow opportunities for shared mechanics. | +| Extraction | `internal/stage/extract.go`, `extract_resume.go`, `internal/adapters/notarius`, `internal/fileops/directory.go` | External receipt and lane validation, configuration fingerprint limits, immutable promotion, symlink/root replacement defenses, provenance and checksum checks, immediate and cross-invocation reuse, obsolete versus unsafe outcomes, failure residue, and whether dense verification logic can be clarified without weakening it. | +| Analyze and artifact dependencies | `internal/stage/analyze.go`, `internal/artifacts`, `internal/artifactpolicy` | Source-family validation, runtime catalog state, enabled/selected/reused distinctions, topological ordering and cycle handling, required/optional inputs, local-only previous sources, deterministic metadata, repeated lookup/scanning, and ownership shared with config and publish. | +| Publish and cleanup | `internal/stage/publish.go`, `internal/app/post_publish_cleanup.go`, `internal/app/cleanup_targets.go` | Prerequisite success, output selection, locks, required/optional behavior, exclusion rules, deterministic upload set, retry/idempotency implications, current-manifest then commit-marker ordering, metadata gates, and destructive path confinement. | +| Manifest state | `internal/manifest` | Validation and backward compatibility, atomic persistence, timestamps, session/run identity, transition truth table, clearing versus retaining payload, create/load/save duplication, failure during dual-manifest updates, and whether state mutation has a single owner. | +| Artifacts, paths, and policy | `internal/artifacts`, `internal/artifactpolicy`, `internal/pathsafe` | Canonical helper coverage, ad hoc reconstruction by callers, source-ID ownership, manifest-first resolution, extraction/current-state identity, destination normalization, stable ordering, typed missing-state errors, symlink/traversal defenses, and duplicate policy across config/stages/app. | +| Restore | `internal/app/restore*.go`, `internal/previouscache`, `internal/audio` | Remote authority, confined mapping, deterministic plan actions, local conflict and force behavior, dry-run purity, temp-file installation, manifest-last ordering, partial failure/retry behavior, report accuracy, cache reuse, and shared current-state mechanics. | +| File operations | `internal/fileops`, `internal/pathsafe`, local-store code in `internal/artifacts` | Atomic-write and promotion guarantees, permissions, close/sync/rename error handling, temp cleanup, same-filesystem assumptions, replacement policy, regular-file-only traversal, symlink and root-swap resistance, lock cleanup, and portability. | +| External adapters and storage | `internal/adapters`, `internal/audio` | Transport isolation, shared subprocess mechanics versus adapter-specific policy, command/config duplication, quoting and working directories, timeouts/cancellation, stdout/stderr separation, HTTP body and retry behavior, S3 pagination/streaming/not-found mapping, credential independence, and external error adaptation. | +| Shared models and diagnostics | `internal/artifactmodel`, `internal/contracts`, `internal/logging` | Serialization and validation invariants, unnecessary conversions, ownership of shared types, stable diagnostic structure, redaction, and whether small shared packages remain cohesive. | +| Tests, examples, and automation | all `*_test.go`, `examples/`, `.woodpecker/` | Risk ownership, semantic duplication, fixture cost, policy-coupled assertions, realistic boundary tests, end-to-end sufficiency, default-suite isolation, example validation, diagnostic coverage, flakiness, runtime cost, and enforcement of canonical validation. | + +## Narratio-Specific Cross-Boundary Scenarios + +In addition to package-local review, trace these complete scenarios because a +modular pipeline can look correct within every package while violating an +end-to-end invariant: + +1. A stage succeeds, its result becomes non-resumable, the rerun fails, and a + later invocation decides what remains usable. +2. An upstream forced or changed outcome interacts with already-succeeded, + self-skipped, and disabled downstream stages. +3. Extraction produces a valid immutable bundle, then configuration or + transitive Notarius inputs change before analyze or publish. +4. Previous-session state is published, restored or prepared into the local + cache, and consumed by analyze without an unintended remote read. +5. Publish fails at each upload boundary, especially between current manifest + and current-run pointer, followed by status, restore, and retry. +6. Restore encounters identical files, conflicting files, unsafe remote keys, + cache hits, and a failure immediately before manifest installation. +7. Automatic or manual cleanup is requested after skipped, failed, locked, + partially uploaded, and fully committed publish outcomes. +8. Cancellation reaches bounded transcription work, HTTP requests, + subprocesses, object storage, and manifest reporting without leaks or false + success. +9. A configured artifact is disabled, unselected, reused, generated from + another artifact, sourced from extraction, or sourced from a previous + session, then filtered for publish. +10. The same session is invoked concurrently, including lock contention and + cleanup/release failures. + +## Completion Criteria + +The audit is complete when: + +- every area in the inspection map has been reviewed against its canonical + contracts and focused tests; +- the stage lifecycle matrix and cross-boundary scenarios have explicit + conclusions; +- duplication candidates have been classified rather than merely counted; +- simplification and performance recommendations explain their correctness + constraints and expected benefit; +- comment recommendations identify the non-obvious rationale to preserve; +- the test suite has a risk-based sufficiency assessment, including gaps, + redundancy, durability, execution properties, and automation; +- findings are deduplicated, evidence-backed, and ranked by risk and dependency; + and +- unresolved questions and intentionally accepted risks are recorded rather + than silently omitted. + +## Execution + +The [Audit Sequence](audit-sequence.md) is the canonical owner of execution +order, stage boundaries, checkpoints, validation, and audit deliverables. This +document remains the canonical owner of audit scope, review criteria, and the +finding standard. diff --git a/docs/roadmap/audit-sequence.md b/docs/roadmap/audit-sequence.md new file mode 100644 index 0000000..8b2b74e --- /dev/null +++ b/docs/roadmap/audit-sequence.md @@ -0,0 +1,734 @@ +# Codebase Audit Sequence + +Status: proposed + +## Purpose And Relationship To The Audit Plan + +This document turns the [Codebase Audit Plan](audit-plan.md) into a bounded, +execution-ready sequence. The plan owns scope, review criteria, and the finding +standard. This document owns ordering, dependencies, working records, +validation, and exit gates. + +The sequence is for investigation only. Do not mix production refactors or bug +fixes into the audit. A confirmed urgent defect may justify stopping to request +a separate remediation change, but its fix is not part of this sequence. + +## Audit Run Records + +Create `docs/roadmap/audit-findings.md` when the audit begins. It is the single +working ledger and final audit report. Initialize it with: + +- the audited revision, branch/worktree state, Go version, platform, and audit + date; +- baseline command results and timings; +- an area coverage ledger; +- the stage lifecycle matrix; +- the cross-boundary scenario matrix from the audit plan; +- a risk-to-test matrix; +- candidate and confirmed finding registers; and +- unresolved questions, accepted risks, and final conclusions. + +Track each execution stage in the coverage ledger with one of `not_started`, +`in_progress`, `complete`, or `blocked`. For a completed stage, record: + +- contracts, packages, files, and important symbols reviewed; +- graph traces, commands, tests, or other evidence used; +- finding and candidate IDs produced; +- explicit no-finding conclusions for reviewed high-risk behavior; and +- follow-up questions assigned to later stages. + +Use stable finding IDs with these prefixes: + +| Prefix | Category | +| --- | --- | +| `COR` | Confirmed correctness defect | +| `RSK` | Correctness or operational risk | +| `ARC` | Ownership or architectural-boundary issue | +| `DUP` | Duplicated mechanism or policy | +| `SIM` | Simplification or idiomatic-Go opportunity | +| `EFF` | Efficiency or resource-use issue | +| `COM` | Missing, misleading, or stale explanatory comment | +| `TST` | Test-suite gap, redundancy, brittleness, or execution issue | + +Candidate IDs remain candidates until manual inspection confirms the behavior, +contract, realistic scenario, and affected callers. Rejected candidates remain +in a short classification log so later stages do not reopen them without new +evidence. + +## Execution Rules + +1. Pin the audit to the revision recorded in Stage 0. If the worktree or HEAD + changes, record the change and rerun every affected stage; do not silently + combine evidence from different implementations. +2. Use codebase graph search and call/data-flow traces before broad source + search. Read the exact implementation, focused tests, and canonical contract + before confirming a finding. +3. Record test-policy observations during every behavior pass. Stage 12 owns the + suite-wide conclusion but must not rediscover the suite from scratch. +4. Record cross-area observations as candidates for the stage that owns the + conclusion. Avoid producing duplicate findings from several review passes. +5. Treat baseline failures as evidence, not automatic blockers. Continue when + read-only inspection remains sound, and state the limitation. Stop only when + the repository cannot be identified, required sources are unavailable, or a + failure makes later evidence unreliable. +6. Do not exercise a suspected destructive, credentialed, paid, or live-service + path merely to prove a defect. Use source reasoning, existing safe fakes, or + a narrowly controlled offline reproduction. +7. Escalate a credible active data-loss, secret-exposure, or unsafe-cleanup + defect immediately. Preserve the evidence and do not wait for final + synthesis before reporting it. +8. A stage is complete only when its exit gate is met. A package test passing is + evidence, not proof that the review is complete. + +## Sequence Overview + +| Stage | Focus | Depends on | Primary result | +| --- | --- | --- | --- | +| 0 | Pin revision and establish baseline | None | Reproducible audit record | +| 1 | Contract, boundary, and lifecycle map | 0 | Review matrices and ownership map | +| 2 | Runner and manifest state machine | 1 | Lifecycle and dual-ledger conclusions | +| 3 | Paths, artifacts, and filesystem safety | 1-2 | State/path authority and mutation conclusions | +| 4 | Publish, remote commit, and cleanup | 2-3 | Commit-boundary and destructive-operation conclusions | +| 5 | Restore and remote/previous state | 2-4 | Restore authority and recovery conclusions | +| 6 | Configuration and application composition | 1-5 | Validation and wiring conclusions | +| 7 | External adapters and shared support | 3, 6 | Boundary, cancellation, and resource conclusions | +| 8 | Prepare and transcript-processing stages | 2-3, 6-7 | Ordinary stage-contract conclusions | +| 9 | Extraction vertical slice | 2-3, 6-7 | Promotion, provenance, and resume conclusions | +| 10 | Analyze and artifact dependency slice | 3, 6, 8-9 | Dependency and source-resolution conclusions | +| 11 | Cross-codebase duplication, simplicity, efficiency, and comments | 2-10 | Classified maintainability candidates | +| 12 | Test-suite policy audit | 2-11 | Risk-based suite sufficiency assessment | +| 13 | Synthesis and audit closeout | 0-12 | Final deduplicated audit report | + +Stages are intentionally ordered. Later stages may resolve candidates raised by +earlier ones, but they must not invalidate an earlier stage silently. Return to +the owning stage, update its coverage record, and note the new evidence. + +## Stage 0: Pin Revision And Establish Baseline + +### Entry + +- Repository root and `docs/development.md` are available. +- The audit plan and canonical policy documents can be read. + +### Execute + +1. Record `git rev-parse HEAD`, branch/detached state, `git status --short`, + `go version`, `go env GOOS GOARCH`, and the current date. +2. Confirm that the code knowledge graph represents the recorded repository and + revision; refresh the index if it is missing or stale. +3. Capture the package/file/test inventory, entry points, architecture + boundaries, high fan-in symbols, complexity signals, and similarity signals. +4. Run the default offline baseline and record wall time and failures: + + ```sh + go test -count=1 ./... + go test -race -count=1 ./... + go vet ./... + ``` + +5. Build into an external temporary directory so validation does not add a + workspace binary: + + ```sh + audit_build_dir="$(mktemp -d)" + go build -o "$audit_build_dir/narratio" ./cmd/narratio + go test -coverprofile="$audit_build_dir/coverage.out" ./... + ``` + +6. Inventory the repository's CI/release validation, maintained examples, fuzz + tests, golden data, opt-in tests, and generated-test update mechanisms. + +### Output + +- Baseline and inventory sections in `audit-findings.md`. +- Initial coverage ledger containing Stages 0-13. +- Unconfirmed metric-driven candidates, clearly labeled as such. + +### Exit Gate + +- Revision and environment are reproducible. +- Every baseline command has a recorded result. +- Graph freshness is known. +- Any limitation that affects later stages has an owner and disposition. + +## Stage 1: Build The Contract, Boundary, And Lifecycle Map + +### Entry + +- Stage 0 is complete. + +### Execute + +1. Read the architecture, internal overview, testing policy, focused internal + documents, and the relevant CLI/configuration/operations/integration + contracts using the development guide's routing rules. +2. Map each package and important interface to its owned policy. Mark every + cross-package dependency that appears to reverse or blur the intended + direction for later confirmation. +3. Build a stage-contract matrix with canonical order, declared inputs, + outputs, configuration, adapters, skip behavior, resume validation, + materialization boundary, manifest effects, and downstream invalidation. +4. Build the lifecycle matrix required by the audit plan: first run, + already-succeeded skip, self-skip, failure, interruption, forced replacement, + non-resumable result, and successful rerun. +5. Assign each of the ten cross-boundary scenarios in the audit plan to its + primary execution stage and list supporting packages/tests. +6. Seed the risk-to-test matrix with the intended test owner for each + architectural invariant. Do not judge sufficiency yet. + +### Output + +- Package ownership, stage-contract, lifecycle, scenario, and preliminary + risk-to-test matrices. +- `ARC` and `RSK` candidates for apparent disagreements, without deciding from + documentation alone which artifact is wrong. + +### Exit Gate + +- Every area in the audit plan's inspection map has an assigned stage. +- Every architectural invariant has an implementation owner and intended test + owner. +- Unknown or contradictory contracts are explicitly recorded. + +## Stage 2: Audit The Runner And Manifest State Machine + +### Entry + +- Stage 1 matrices are complete. + +### Execute + +1. Trace the entry paths into full-run and single-stage execution through + `internal/app/planner.go`, `runner.go`, `run_stage.go`, and related helpers. +2. Inspect `internal/manifest` models, validation, session/run creation, + loading, normalization, atomic saves, and all transition methods. +3. Walk every lifecycle-matrix cell through both manifests. Verify clearing + versus retention of outputs, diagnostics, generated configuration, metadata, + errors, actions, timestamps, and downstream state. +4. Reason about failures before and after each session-manifest and run-manifest + save. Determine which disagreement states are possible and how a later + invocation interprets them. +5. Review force, changed-result, self-skip, failed-result, and non-resumable + invalidation separately. Confirm behavior at the first and last canonical + stage. +6. Review session lock acquisition/release and concurrent invocation behavior, + while leaving path implementation details to Stage 3. +7. Classify the runner's complexity and repeated session/run persistence paths: + state-machine clarity, justified explicitness, candidate local helpers, and + comments that preserve ordering rationale. +8. Review focused app/manifest tests against the matrix and add observations to + the risk-to-test ledger. + +### Validation + +```sh +go test -count=1 ./internal/app ./internal/manifest +go test -race -count=1 ./internal/app ./internal/manifest +``` + +### Exit Gate + +- Every lifecycle cell has a source-backed conclusion for both manifests. +- Cross-boundary scenarios 1, 2, and the lock portion of 10 are resolved or + carry explicit questions. +- All runner/manifest candidates are confirmed, rejected, or assigned to a + named later stage. + +## Stage 3: Audit Paths, Artifacts, And Filesystem Safety + +### Entry + +- Stages 1-2 are complete. + +### Execute + +1. Review `internal/artifacts`, `internal/artifactpolicy`, `internal/pathsafe`, + `internal/fileops`, and local-store filesystem code. +2. Inventory canonical path and key helpers, then search callers for ad hoc + reconstruction, double normalization, mixed slash/filesystem semantics, or + policy implemented outside its owner. +3. Trace built-in, configured, extraction, previous-session, and current-state + artifact resolution. Verify identity, checksum, contract, provenance, + deterministic ordering, and typed missing-state behavior. +4. Review atomic file writes, copies, directory promotion, temp cleanup, + permission preservation, close/sync/rename errors, existing-destination + behavior, same-filesystem assumptions, and platform sensitivity. +5. Walk traversal, absolute path, broad root, symlink component, inspected-root + replacement, non-regular file, and time-of-check/time-of-use scenarios. +6. Confirm that low-level file/storage helpers receive explicit destinations + and do not infer stage, campaign, session, run, or publish policy. +7. Inspect lock-file implementation and cleanup errors to finish scenario 10. +8. Record focused test ownership and gaps without duplicating Stage 2's state + conclusions. + +### Validation + +```sh +go test -count=1 ./internal/artifacts ./internal/artifactpolicy ./internal/pathsafe ./internal/fileops +go test -race -count=1 ./internal/artifacts ./internal/fileops +``` + +### Exit Gate + +- Every canonical path/key family has one identified owner. +- Every material filesystem mutation has documented confinement and atomicity + conclusions. +- Scenario 10 is resolved. +- Safety checks that appear repetitive are classified before any simplification + recommendation is made. + +## Stage 4: Audit Publish, Remote Commit, And Cleanup + +### Entry + +- Stages 2-3 are complete. + +### Execute + +1. Trace publish from stage selection through object-store calls, manifest + metadata, commit-marker publication, run completion, and post-publish + cleanup. +2. Verify prerequisite stage-state checks, selected/configured/extraction + output resolution, required versus optional outputs, static and remote + locks, run-file exclusions, previous-cache inclusion, and deterministic + upload order. +3. Enumerate failures before and after every upload. Prove that + `current/run_id.txt` is written last and is the only remote-current commit + point. +4. Review retry/idempotency behavior, existing remote objects, partial uploads, + pointer/manifest disagreement, and status/restore interpretation after each + partial outcome. +5. Trace automatic and manual cleanup gates. Confirm publish execution, + `uploaded`, `current_pointer_written`, explicit policy, and confined targets + are all required at the correct boundary. +6. Confirm that `--force` cannot override publish locks or cleanup safety. +7. Review duplication between publish planning, artifact destination policy, + operator views, and cleanup metadata only after ownership is established. + +### Validation + +```sh +go test -count=1 ./internal/stage ./internal/app ./internal/artifacts ./internal/adapters/storage +``` + +### Exit Gate + +- Cross-boundary scenarios 5 and 7 are resolved for every relevant failure + boundary. +- Remote-current authority and local-cleanup eligibility have explicit truth + tables. +- Publish findings distinguish stage policy from storage mechanics. + +## Stage 5: Audit Restore And Remote/Previous State + +### Entry + +- Stages 2-4 are complete. + +### Execute + +1. Trace restore discovery, planning, execution, reporting, audio + materialization, and previous-cache planning through `internal/app`, + `internal/artifacts`, `internal/previouscache`, `internal/audio`, and storage. +2. Confirm remote pointer/manifest identity and campaign/session/run authority, + including missing and inconsistent current state. +3. Verify remote-to-local confinement, deterministic action ordering, + `download`/`skip_same`/`conflict` decisions, force semantics, and dry-run + purity. +4. Walk failures during download, checksum or manifest validation, atomic + install, report persistence, and the manifest-last boundary. Record the + intentional lack of rollback and retry consequences. +5. Review audio spool/cache identity, cache-hit verification, partial download + behavior, and duplicate remote/filesystem work. +6. Review previous-session requirement planning, required/optional behavior, + identity checks, published-path fallback, and deterministic local mapping. +7. Confirm which mechanics are shared with status/validate/operator commands + and which caller-specific missing-state policies must remain separate. + +### Validation + +```sh +go test -count=1 ./internal/app ./internal/previouscache ./internal/audio ./internal/artifacts ./internal/adapters/storage +``` + +### Exit Gate + +- Cross-boundary scenarios 4 and 6 are resolved through retry/recovery. +- Restore authority, manifest-last installation, and partial-write behavior are + explicit. +- Previous-cache conclusions are ready for the prepare and analyze passes. + +## Stage 6: Audit Configuration And Application Composition + +### Entry + +- Stage 1 is complete and Stages 2-5 have identified the policies that + configuration and composition must supply. + +### Execute + +1. Review `internal/config`, `cmd/narratio`, application command dispatch, + configuration selection, secret-file environment loading, and production + collaborator construction. +2. Trace discovery, precedence, strict YAML decoding, defaults, empty values, + normalization, session templating, and validation order across pipeline, + campaign, and session configuration. +3. Verify cross-field constraints for stage enablement, paths, timeouts, + concurrency, artifacts, Notarius, Scriptorium, publish, storage, cleanup, + audio, and previous-session behavior. +4. Compare validation logic with maintained examples and the public + configuration contract. Record contract drift rather than silently choosing + code or docs. +5. Check that filesystem secrets are loaded before the boundary that consumes + them and are excluded from logs, manifests, reports, generated files, and + errors. +6. Review conditional construction of expensive/external collaborators and + cleanup of anything with a lifecycle. Confirm test injection cannot create a + behavior different from production composition. +7. Classify repeated validators, path checks, timeout parsing, constructor + wrappers, and single-stage command wrappers by policy owner. + +### Validation + +```sh +go test -count=1 ./internal/config ./internal/app ./cmd/narratio +go vet ./... +``` + +### Exit Gate + +- Every operator-visible field used by audited behavior has a traced default, + normalization, validation, and consumer. +- Composition conclusions cover enabled and disabled stages without requiring + live services or credentials. +- Maintained examples have an explicit validity conclusion. + +## Stage 7: Audit External Adapters And Shared Support + +### Entry + +- Stages 3 and 6 are complete. + +### Execute + +1. Review `internal/adapters`, `internal/audio`, `internal/logging`, + `internal/contracts`, and `internal/artifactmodel` at their public package + boundaries. +2. For each HTTP, subprocess, notification, and object-storage adapter, compare + implementation with its integration contract and trace all production + callers. +3. Verify context cancellation, timeout ownership, process termination and + waiting, goroutine/channel closure, HTTP response-body closure, retries, + malformed responses, streaming, pagination, not-found mapping, and local + file cleanup. +4. Confirm command argument construction, working directory, environment, + generated configuration, stdout/stderr separation, output validation, and + external error adaptation stay inside the owning adapter. +5. Compare subprocess implementations to the shared subprocess package. Classify + repeated constructor/config/log/output mechanics separately from + adapter-specific protocol policy. +6. Review fakes for realistic state and concurrency behavior, but defer their + suite-wide value judgment to Stage 12. +7. Check shared models for avoidable conversions, stable serialization, + validation ownership, and redaction-sensitive diagnostic fields. + +### Validation + +```sh +go test -count=1 ./internal/adapters/... ./internal/audio ./internal/logging ./internal/contracts ./internal/artifactmodel +go test -race -count=1 ./internal/adapters/... ./internal/audio +``` + +### Exit Gate + +- Every external resource has an explicit acquisition, cancellation, and + release conclusion. +- Transport types and protocol policy have not leaked into stages. +- Adapter duplication candidates identify the correct shared or specific + owner. + +## Stage 8: Audit Prepare And Transcript-Processing Stages + +### Entry + +- Stages 2-3 and 6-7 are complete. + +### Execute + +1. Review `prepare`, `transcribe`, `merge`, `polish`, `normalize`, `trim`, and + `render` as vertical slices from resolved configuration and manifest input + through adapter call, run-local output, validation, canonical + materialization, and recorded result. +2. Verify each implementation against the Stage 1 contract matrix and focused + internal document. Record any undeclared input, output, diagnostic, config, + adapter, or skip/failure behavior. +3. For prepare, confirm local/S3 exclusivity, stable input copying, + previous-cache clearing/hydration, and deterministic manifest input records. +4. For transcribe, confirm unique speaker identities, bounded runtime + concurrency, cancellation, deterministic result ordering, adapter-returned + path identity, and partial failure behavior. +5. For transformation/render stages, confirm manifest-first resolution, + run-local paths, schema/report validation, disabled/default behavior, + canonical promotion, and diagnostic-versus-artifact classification. +6. Compare similar stage implementations for shared mechanisms only after + listing meaningful differences. Avoid a generic stage framework. +7. Add stage-focused test ownership, gaps, and redundancy candidates to the + risk-to-test matrix. + +### Validation + +```sh +go test -count=1 ./internal/stage ./internal/audio ./internal/previouscache ./internal/adapters/whisperx ./internal/adapters/seriatim ./internal/adapters/audita ./internal/adapters/scriptorium +go test -race -count=1 ./internal/stage ./internal/audio +``` + +### Exit Gate + +- Every reviewed stage has a completed contract-matrix row. +- Cross-boundary scenario 8 is resolved for transcription and subprocess-backed + transformation stages. +- Similarity candidates are classified as intentional explicitness, local + helper candidates, or shared-owner findings. + +## Stage 9: Audit The Extraction Vertical Slice + +### Entry + +- Stages 2-3 and 6-7 are complete. + +### Execute + +1. Trace extraction from configuration validation and composition through + transcript resolution, invocation fingerprint, Notarius execution, receipt + and lane validation, directory promotion, manifest recording, catalog + hydration, resume validation, analyze, and publish consumers. +2. Verify run-local isolation, regular-file and confined-index requirements, + required-lane policy, contract/provenance construction, checksum timing, + immutable destination identity, and no-replacement promotion. +3. Enumerate failures before and after subprocess completion, receipt parsing, + payload inspection, promotion, and manifest persistence. Determine what + remains diagnostic, durable, advertised, and reusable. +4. Walk every resume validation branch. Distinguish obsolete/missing outcomes + that trigger rerun from unsafe conditions that must stop execution. +5. Evaluate the fingerprint's intentionally observable and unobservable inputs + against documentation and force guidance. +6. Review the dense validation code for named sub-decisions and comments while + preserving the visible security proof and check ordering. +7. Confirm focused tests cover immediate reuse, cross-invocation reuse, + configuration change, payload tampering, provenance mismatch, symlinks/root + replacement, failure residue, and downstream invalidation at the correct + layers. + +### Validation + +```sh +go test -count=1 ./internal/stage ./internal/artifacts ./internal/fileops ./internal/adapters/notarius ./internal/app +``` + +### Exit Gate + +- Cross-boundary scenario 3 is resolved, including transitive-input limits. +- Promotion, advertisement, and resume each have a distinct authority and + failure conclusion. +- Every proposed simplification states which security or compatibility checks + it preserves. + +## Stage 10: Audit Analyze And Artifact Dependencies + +### Entry + +- Stages 3, 6, 8, and 9 are complete. + +### Execute + +1. Trace all analyze source families from configuration validation through + runtime catalog registration, availability, resolution, Scriptorium + execution/reuse, materialization, metadata, and publish selection. +2. Verify enabled, selected, executable, reused, generated, and unavailable + states are distinct and deterministic. +3. Review configured-artifact dependency validation and runtime topological + ordering for cycles, missing dependencies, stable ordering, and consistency + between configuration and execution. +4. Confirm required/optional behavior and guidance for built-in transcripts, + prepared stable inputs, configured artifacts, extraction sources, and + previous-session sources. +5. Prove previous-session resolution is local-only during analyze and that + disabled artifacts are reused only under the documented conditions. +6. Inspect repeated resolution branches, parameter width, nested lookup, and + ordering work for a smaller representation or indexed plan without merging + distinct source policies. +7. Review tests for each state transition and source family at the narrowest + stable owner, noting semantic duplication across config, artifacts, stage, + publish, and assembled runner tests. + +### Validation + +```sh +go test -count=1 ./internal/stage ./internal/artifacts ./internal/artifactpolicy ./internal/config ./internal/adapters/scriptorium ./internal/app +``` + +### Exit Gate + +- Cross-boundary scenario 9 is resolved for every source family and selection + state. +- Dependency ordering and source availability have explicit determinism and + complexity conclusions. +- Config, artifact-policy, catalog, stage, and publish ownership is unambiguous + or represented by an `ARC` finding. + +## Stage 11: Audit Duplication, Simplicity, Efficiency, And Comments + +### Entry + +- Behavior stages 2-10 are complete, so structural candidates can be judged + against known contracts. + +### Execute + +1. Rerun graph similarity, complexity, fan-in/fan-out, call-path, loop-depth, + scan-in-loop, and change-coupling analyses on production code. Add targeted + text/static searches for patterns the graph cannot represent. +2. Revisit all `DUP`, `SIM`, `EFF`, and `COM` candidates collected earlier. + Search for additional occurrences and trace all callers before assigning an + owner. +3. For duplication, classify coincidental syntax, shared mechanism, duplicated + policy, or deliberately explicit security/state logic. Propose only the + narrowest helper that improves ownership and comprehension. +4. For complexity, sketch the smaller control flow or data model and verify it + leaves state transitions, validation order, and commit boundaries visible. +5. For efficiency, state the input scale or call frequency, current and proposed + complexity/I/O behavior, expected benefit, and benchmark or measurement + needed. Reject micro-optimizations without a credible workload. +6. Review standard-library usage, errors, slices/maps, allocations, copying, + sorting, serialization, filesystem passes, adapter initialization, remote + calls, goroutines/channels, and interface breadth across the complete codebase. +7. Review comments only after simplification decisions. Recommend why-comments + for remaining invariants, compatibility limits, safety checks, partial + failure, and ordering; flag comments that restate code or no longer match it. +8. Check dependencies and platform assumptions for clear correctness, + portability, complexity, or maintenance consequences. + +### Validation + +- Run focused package tests for any behavior used to disprove or confirm a + candidate. +- Run existing benchmarks where relevant. Propose a benchmark rather than + inventing performance claims when representative measurement is absent. + +### Exit Gate + +- Every structural candidate is confirmed, rejected with a reason, or merged + into a stronger root-cause finding. +- No helper recommendation creates a generic workflow abstraction or moves + policy into a low-level utility. +- Every efficiency finding has a credible workload and validation method. +- Every comment finding states the non-obvious rationale that should be + preserved. + +## Stage 12: Audit The Test Suite Against Policy + +### Entry + +- Stages 2-11 have populated the risk-to-test matrix and test observations. + +### Execute + +1. Complete the risk-to-test matrix. For every consequential invariant, list + the current tests, proper owner, protected defect, missing failure modes, and + overlap with other layers. +2. Review tests by behavior cluster rather than filename: parsing/validation, + domain/state, filesystem, adapters, orchestration, CLI, integration, and + representative assembled workflows. +3. Classify gaps for data integrity, destructive operations, compatibility, + security, concurrency, idempotency, recovery, cancellation, and partial + success. Confirm the gap is not credibly protected elsewhere. +4. Classify redundancy and brittleness: private constants/defaults, exact error + wording, incidental formatting/paths, mock choreography, oversized + snapshots, helper-level duplication, and the same policy repeated across + layers. +5. Review doubles using the policy order: real deterministic collaborator, + stateful fake, stub, then mock when interaction is contractual. Check that + fakes model the failure and state semantics used by the tests. +6. Inspect test helpers and large test functions for simplification and + meaningful table-driven boundaries without creating a fixture framework + whose maintenance cost exceeds its value. +7. Review determinism and isolation: credentials, network access, paid APIs, + environment, working directory, clocks, randomness, ports, temp paths, + process-global state, ordering, cleanup, and parallel execution. +8. Use coverage to investigate consequential weak branches, not as a score. + Review heavily covered behavior for marginal-value duplication as well. +9. Identify focused fuzz opportunities for parsers, YAML/JSON normalization, + source IDs, confined paths, remote/local mapping, and manifest decoding. +10. Compare local requirements with `.woodpecker/` and other automation. Record + missing enforcement as a risk/cost decision, not an assumption that every + diagnostic command belongs in CI. +11. Investigate order dependence and flakiness with bounded runs, recording + runtime and any reproducible seed: + + ```sh + go test -shuffle=on -count=3 ./... + go test -race -shuffle=on -count=1 ./... + ``` + +### Exit Gate + +- Every important risk has a sufficiency conclusion and one intended test + owner. +- Every proposed test addition names the realistic defect and marginal value. +- Every deletion/consolidation names the stronger remaining protection. +- Default-suite determinism, offline behavior, runtime, flakiness, and CI + enforcement have explicit conclusions. + +## Stage 13: Synthesize And Close The Audit + +### Entry + +- Stages 0-12 meet their exit gates or have explicitly accepted limitations. + +### Execute + +1. Reconcile candidates and findings across stages. Merge shared root causes and + remove repeated symptoms while retaining all affected locations and + contracts. +2. Recheck every confirmed finding against current source, callers, tests, and + canonical documentation. Downgrade or reject anything supported only by a + metric or hypothetical preference. +3. Rank impact, likelihood, confidence, and remediation scope separately. Order + the recommended backlog by dependency: correctness/data safety first, + architectural ownership next, then simplification/duplication, tests, + efficiency, and comments where they remain necessary. +4. Record positive conclusions for high-risk areas where the current design and + tests are sufficient. The report should not imply that only defective areas + were reviewed. +5. Reconcile the area coverage ledger, lifecycle matrix, cross-boundary scenario + matrix, and risk-to-test matrix with the audit plan's completion criteria. +6. Record any accepted risks, ambiguous contracts, environmental limitations, + and deferred investigations with an explicit rationale and owner. +7. Check whether HEAD or the worktree changed since Stage 0. Rerun affected + stages or clearly pin the report to the original revision. +8. Validate the report and roadmap document links and run `git diff --check`. + If implementation changed during the audit, rerun the full Stage 0 validation + baseline against the final audited revision. + +### Final Deliverable + +`docs/roadmap/audit-findings.md` must contain: + +- an executive assessment without unsupported quality scores; +- the audited revision and validation baseline; +- coverage and scenario completion summaries; +- confirmed findings ordered by dependency and risk; +- rejected candidate themes where their recurrence would otherwise waste work; +- the test-suite sufficiency assessment; +- positive conclusions and accepted risks; and +- a recommended remediation order, without implementing the remediation. + +### Exit Gate + +- Every completion criterion in the audit plan is satisfied or explicitly + marked limited with rationale. +- Every finding is evidence-backed, deduplicated, actionable, and assigned a + stable ID. +- No production change is included in the audit output. +- The report is sufficient to prepare a separate remediation sequence without + repeating discovery.