Establish the codebase audit baseline and contract map

This commit is contained in:
2026-08-10 12:12:48 +00:00
parent 74e2d21de5
commit 9850767a8a
3 changed files with 1451 additions and 0 deletions

View File

@@ -0,0 +1,385 @@
# Codebase Audit Findings
Status: in progress
This document is the working ledger and final report for the audit defined by
the [audit plan](audit-plan.md) and [audit sequence](audit-sequence.md). The
audit is investigative: candidate findings below are not remediation changes.
## Audit Identity And Baseline
| Field | Value |
| --- | --- |
| Audited revision | `74e2d21de5fb2ada0be5ef3fe9333e0d48ac7fb3` (`Close the completed roadmap documents`) |
| Branch | `main`, attached worktree |
| Initial worktree state | Untracked `docs/roadmap/audit-plan.md` and `docs/roadmap/audit-sequence.md`; no production or test changes |
| Audit date | 2026-08-10 UTC |
| Toolchain | `go version go1.26.5 linux/amd64` |
| Platform | `GOOS=linux`, `GOARCH=amd64` |
| Repository root | `/home/eric/Workspace/narratio` |
The two initial untracked files are the audit specification supplied for this
run. Adding this ledger and tracking those documents changes documentation
only; all implementation and test evidence remains pinned to the revision
above. If implementation or tests change, affected audit stages must be rerun
and this section must record the new revision.
### Baseline Commands
| Command | Result | Wall time | Evidence or limitation |
| --- | --- | --- | --- |
| `go test -count=1 ./...` | pass | 3.34 s | All 23 packages passed; `cmd/narratio` has no test files. |
| `go test -race -count=1 ./...` | fail | 55.65 s | Race in `internal/adapters/whisperx.(*FakeClient).Transcribe` at `fake.go:45`, reached concurrently by `TestTranscribeStageTranscribesPreparedAudio`; candidate `TST-001`. All packages reported before `internal/stage` passed. |
| `go vet ./...` | pass | 0.47 s | No diagnostics. |
| `go build -o "$audit_build_dir/narratio" ./cmd/narratio` | pass | 1.03 s | Built outside the repository in `/tmp/tmp.x11pJL7014`. |
| `go test -coverprofile="$audit_build_dir/coverage.out" ./...` | pass | 11.05 s | Diagnostic coverage only; no percentage is treated as a gate. |
Coverage ranged from 69.8% (`internal/manifest`) to 100% (`internal/logging`)
among tested packages. `cmd/narratio` reported 0% because it has no tests. The
remaining package results ranged from 70.1% to 88.1%. Stage 12 owns the
risk-based interpretation; these numbers are inventory signals only.
### Code Graph Freshness And Structural Inventory
The `narratio` graph was rebuilt in `moderate` mode after the revision was
pinned. Its branch record reports the exact audited HEAD, `main`, and the
repository root above. The index contains 2,407 nodes and 13,181 edges across
224 modeled files: 1,494 functions, 136 methods, 226 structs, 12 interfaces,
and 20 modeled package nodes. The moderate filter excluded documentation,
examples, `.git`, `.codex`, and `cmd/narratio`; the executable entry point was
therefore verified through `go list` and direct inspection instead of graph
evidence. Internal production code is represented at the pinned revision.
Repository inventory at that revision:
- 23 Go packages, including `cmd/narratio`;
- 221 tracked Go files and 95 tracked `_test.go` files;
- 278 tracked files total;
- one process entry point, `cmd/narratio/main.go`, delegating to
`internal/app.Execute`;
- 11 canonical stages returned by `internal/stage.All`; and
- 12 modeled interfaces, of which 11 are Narratio boundaries and one is the
private AWS S3 client seam.
Graph call tracing from `internal/app.Execute` confirms command dispatch into
run, single-stage, clean, and session-helper paths, followed by configuration,
artifact/path, manifest, stage, storage, restore, and cleanup owners. The
production import inventory shows no lower-level package importing
`internal/app`; apparent graph rollups such as `stage -> app`, `adapters -> app`,
and `config -> app` came from test relationships or graph classification and
are rejected as production dependency reversals at this mapping stage.
### Metric Signals For Later Review
These are prioritization signals, not findings:
| Signal | Evidence | Assigned review |
| --- | --- | --- |
| High fan-in | `app.Error` (207), stage `Run` symbols (151), `app.Execute` (108), `stage.sessionPathsForEnv` (105), `manifest.New` (72), `manifest.MarkStageSucceeded` (56), `app.executeStages` (48), `artifacts.S3SessionPrefix` (41), and `artifacts.SessionWorkDirForCampaign` (37) | Owning behavior stages, then Stage 11 |
| High complexity | `app.executeStages` cyclomatic 54/cognitive 96; `previouscache.BuildPlan` 22/38; `analyzeStage.Run` 20/27; `audita.NewSubprocessRunner` 17/25; `app.SessionInit` 20/21 | Stages 2, 5, 7, 10, then 11 |
| Exact similarity | `app.Analyze`/`app.Publish`, `manifest.Load`/`LoadRun`, `manifest.Create`/`CreateRun`, adapter constructors, and Seriatim fake methods | Owning behavior stages, then Stage 11 |
| Test-heavy hotspot noise | Several test functions and fakes rank highly in transitive-depth and fan-in results | Stage 12; do not infer production risk from the metric |
### Automation And Fixture Inventory
- `.woodpecker/release.yml` is tag-only release automation. It cross-builds
Linux, macOS, and Windows binaries with Go 1.25, then publishes release
assets. It does not run tests, race tests, vet, or example validation.
- `examples/` contains 19 maintained files: pipeline, campaign, session,
template, stable-input, and placeholder-audio fixtures. Configuration tests
are documented as their validation owner.
- No fuzz tests, golden files, golden-update switches, opt-in/live test tags, or
`go:generate` test mechanisms were found.
- Platform build constraints exist for the native no-replace directory tests
and unsupported-platform fallback in `internal/fileops`.
## Execution Coverage Ledger
| Stage | Status | Evidence and result |
| --- | --- | --- |
| 0: baseline | complete | Revision/environment pinned; graph refreshed; inventories and every prescribed baseline command recorded. `TST-001` owns the non-blocking race limitation. |
| 1: contract and boundary map | complete | Canonical contracts and focused internal docs read; ownership, stage-contract, lifecycle, scenario, area, and preliminary risk-to-test matrices recorded below. |
| 2: runner and manifest | not_started | Assigned lifecycle and dual-ledger questions below. |
| 3: paths and filesystem | not_started | Assigned path, lock, artifact, and mutation questions below. |
| 4: publish and cleanup | not_started | Assigned remote commit and cleanup scenarios below. |
| 5: restore and previous state | not_started | Assigned restore and previous-cache scenarios below. |
| 6: configuration and composition | not_started | Assigned configuration, CLI composition, and process-boundary areas below. |
| 7: adapters and shared support | not_started | Assigned external-boundary and cancellation areas below. |
| 8: ordinary stages | not_started | Assigned prepare/transcript behavior and disabled-outcome questions below. |
| 9: extraction | not_started | Assigned extraction promotion, provenance, and resume scenario below. |
| 10: analyze and dependencies | not_started | Assigned artifact dependency/source and selection scenario below. |
| 11: maintainability | not_started | Seeded by graph complexity, similarity, and fan-in signals only. |
| 12: test policy | not_started | Seeded by intended owners and baseline execution observations. |
| 13: synthesis | not_started | No final ranking or accepted-risk decisions yet. |
## Area Coverage And Ownership
Every area in the audit plan has a primary execution owner. `assigned` means it
has been mapped but not behaviorally audited.
| Inspection area | Canonical implementation owner | Primary audit stage | Status |
| --- | --- | --- | --- |
| Process and application boundary | `cmd/narratio`, `internal/app` | 6 (runner lifecycle portions in 2; publish/restore portions in 4-5) | assigned |
| Stage registry and runner | `internal/stage`, `internal/app` | 2 | assigned |
| Configuration | `internal/config` | 6 | assigned |
| Prepare and audio | `internal/stage`, `internal/audio`, `internal/previouscache` | 8 | assigned |
| Transcript stages | `internal/stage` plus tool adapters | 8 | assigned |
| Extraction | `internal/stage`, Notarius adapter, `internal/fileops` | 9 | assigned |
| Analyze and artifact dependencies | `internal/stage`, `internal/artifacts`, `internal/artifactpolicy` | 10 | assigned |
| Publish and cleanup | `internal/stage`, `internal/app` | 4 | assigned |
| Manifest state | `internal/manifest`, transition policy in `internal/app` | 2 | assigned |
| Artifacts, paths, and policy | `internal/artifacts`, `internal/artifactpolicy`, `internal/pathsafe` | 3 (resolution consumption revisited in 10) | assigned |
| Restore | `internal/app`, `internal/artifacts`, `internal/previouscache`, `internal/audio` | 5 | assigned |
| File operations | `internal/fileops`, `internal/pathsafe`, local artifact store | 3 (promotion vertical slice in 9) | assigned |
| External adapters and storage | `internal/adapters`, `internal/audio` | 7 | assigned |
| Shared models and diagnostics | `internal/artifactmodel`, `internal/contracts`, `internal/logging` | 7 (maintainability revisited in 11) | assigned |
| Tests, examples, and automation | package test owners, `examples/`, `.woodpecker/` | 12 | assigned |
## Package And Interface Ownership Map
| Package | Owned contract or policy | Important boundaries | Audit owner |
| --- | --- | --- | --- |
| `cmd/narratio` | Process entry and exit; CLI delegates behavior to app | `main -> app.Execute` | 6 |
| `internal/app` | Command dispatch, composition, locking, planning, lifecycle, restore, cleanup, reporting | `Execute`, `executeStages`; consumes stage/artifact/manifest/adapter contracts | 2, 4-6 |
| `internal/config` | Strict discovery, defaults, resolve, template, and validation rules | Config models and load/resolve/validate functions | 6 |
| `internal/stage` | Canonical order and stage behavior | `Stage`, `ResumeValidator`, `Env`; adapter interfaces are injected | 2, 4, 8-10 |
| `internal/manifest` | Session/run models, transitions, validation, atomic persistence | `Store`; transition methods record but do not choose policy | 2 |
| `internal/artifacts` | Artifact identity/resolution, paths/keys, local store, remote current-state mechanics | `Store`; consumes explicit storage keys | 3, 5, 10 |
| `internal/artifactpolicy` | Configured source/destination identity and safety policy | Narrow validators used by config, artifacts, app, and stages | 3, 10 |
| `internal/artifactmodel` | Shared serialized artifact, contract, and provenance models | Data contract only | 3, 7 |
| `internal/pathsafe` | Confined relative path and destination mechanics | Narrow validation helpers; no stage policy | 3 |
| `internal/fileops` | Atomic files, copies, hashing, no-replace directory promotion | Filesystem mechanics receive explicit paths | 3, 9 |
| `internal/previouscache` | Deterministic previous-session requirement planning/materialization | Uses explicit object-store and artifact contracts | 5, 8, 10 |
| `internal/audio` | S3 audio spool/cache materialization | Uses `storage.ObjectStore`; no stage ordering | 5, 8 |
| `internal/contracts` | Bounds and shared JSON validation models | Data contract only | 7, 8 |
| `internal/logging` | Shared logger construction | `slog` composition | 7, 11 |
| `internal/adapters/whisperx` | WhisperX HTTP protocol | `Client` | 7 |
| `internal/adapters/seriatim` | Merge/normalize/trim/render subprocess protocol | `Runner` | 7 |
| `internal/adapters/audita` | Audita subprocess protocol | `Runner` | 7 |
| `internal/adapters/scriptorium` | Scriptorium run/render subprocess protocol | `Runner` | 7 |
| `internal/adapters/notarius` | Notarius invocation and receipt boundary | `Runner` | 7 (vertical behavior in 9) |
| `internal/adapters/notify` | Notification transport | `Sender` | 7 |
| `internal/adapters/storage` | Explicit bucket-relative object-store operations and S3 mechanics | `ObjectStore`; private `s3API` test seam | 7 |
| `internal/adapters/subprocess` | Shared bounded subprocess/config/log mechanics | Concrete helper package, not stage policy | 7 |
The graph reported no inbound production callers of `Stage.Declares`; text
search found definitions and test stubs but no production invocation. This
reduces the current impact of `ARC-001` but makes the interface's intended owner
and future use an explicit question rather than resolving the mismatch.
## Stage Contract Matrix
The table separates declared/static contracts from dynamic behavior. All
executed stages use the runner's session/run transitions. Unless noted, a
successful result records returned outputs, diagnostics, generated
configuration, and metadata; a different effective executed outcome can stale
succeeded downstream work, while force pre-stales succeeded downstream work.
| Order and stage | Inputs and outputs | Configuration and adapters | Skip/resume behavior | Materialization and manifest effects |
| --- | --- | --- | --- | --- |
| 1 `prepare` | Config, stable inputs, one audio mode, optional previous requirements -> canonical `inputs/**`, `audio/**`, optional `previous/**`, `manifest.inputs` | All resolved config; storage for S3/current previous state; audio/artifact/previous-cache services | No stage-specific resume validator or explicit self-skip | Writes canonical session inputs and deterministic input records; unlike processing stages, `Declares` labels produced canonical files as inputs. |
| 2 `transcribe` | Prepared FLAC files -> raw per-speaker JSON | WhisperX URL/language/retry/timeout/concurrency; `whisperx.Client` | Ordinary succeeded-record skip; no validator/self-skip | Bounded concurrent run-local writes, validation, then canonical transcript materialization. |
| 3 `merge` | Raw transcripts, speakers, autocorrect -> base transcript, optional report | Seriatim merge fields; `seriatim.Runner` | Ordinary succeeded-record skip | Normalized scratch inputs and run-local results validate before canonical transcript/report materialization. |
| 4 `polish` | Base transcript, glossary -> polished transcript, optional report | Audita fields/credential reference; `audita.Runner` | Ordinary succeeded-record skip | Run-local output, report, logs, and generated config; validates before canonical materialization. |
| 5 `normalize` | Polished transcript -> final transcript, optional report | Normalize plus Seriatim fields; `seriatim.Runner` | Ordinary succeeded-record skip | Manifest-first input; run-local validation then configurable canonical output/report. |
| 6 `trim` | Final transcript -> final-trimmed transcript and, when enabled, bounds | Trim, bounds, Scriptorium, and Seriatim fields; both runners when enabled | Disabled trim copies input and still succeeds; no explicit self-skip or resume validator | Run-local bounds/trim result validates then materializes; debug render is diagnostic, not output. |
| 7 `extract` | Final-trimmed source -> immutable index and configured lane outputs | Notarius executable/config/pipeline/timeout/output contracts; `notarius.Runner` | Disabled is explicit `notarius_disabled` self-skip; only current `ResumeValidator`; obsolete reruns, unsafe validation errors | Validates run-local receipt/bundle completely, promotes to unique immutable bundle, records checksums/contracts/provenance; identical repeated self-skip is stable. |
| 8 `render` | Final and final-trimmed JSON -> two Markdown transcripts | Render and Seriatim fields; `seriatim.Runner` | Disabled returns a zero-disposition result with skip metadata, therefore runner-level success rather than explicit self-skip; no validator | Run-local text validates non-empty before canonical materialization when enabled. |
| 9 `analyze` | Dynamic built-in, prepared, extraction, configured, and previous sources -> selected configured artifact outputs | Scriptorium artifact graph/selection; `scriptorium.Runner` | Missing config or no executable artifacts returns success with skip metadata; no validator | Topological run-local generation/reuse, validation, canonical outputs, deterministic metadata; static `Declares` omits dynamic outputs and several input families. |
| 10 `publish` | Session/run state, selected output rules, locks, previous cache -> remote run/output/current objects | Publish/storage/selection fields; `storage.ObjectStore` | Disabled publish or run upload returns success with skip metadata; force cannot bypass locks; no validator | Deterministic uploads; `current/manifest.json` before `current/run_id.txt`; commit metadata gates cleanup. Static prerequisites omit extract because disabled extraction is valid and lane resolution enforces required extraction state when selected. |
| 11 `notify` | No implemented persisted pipeline input/output | Optional `notify.Sender`; default no-op | Ordinary succeeded-record skip; no explicit self-skip or validator | Placeholder metadata and optional notification call; no returned output. `Declares` nevertheless advertises placeholder input/output paths. |
Configuration, adapters, skip policy, and dynamic outputs are not represented
by `IODecl`; their current canonical owners are the focused stage,
configuration, and integration contracts. Whether `IODecl` should remain a
partial display type or become an enforceable declaration is deferred as
`ARC-001`.
## Lifecycle Matrix
This is the intended contract map to be walked through both durable ledgers in
Stage 2. Cells marked `unknown` are not treated as implementation conclusions.
| Outcome | Session manifest intent | Invocation manifest intent | Downstream and next-invocation intent |
| --- | --- | --- | --- |
| First run | Pending/non-succeeded stage becomes running, then succeeded/failed/skipped; executing clears older result payload first | New run record; action `run`; terminal status records this invocation | Success enables later stages; failure stops current execution and an effective outcome change may stale succeeded downstream records. |
| Already-succeeded skip | Existing succeeded session record and payload remain unchanged, subject to resume validation | Action/status record a skip and stable reason for this invocation | Reusable result remains authoritative; pipeline continues. |
| Explicit self-skip | Session stage becomes skipped, clears older result payload, and may record bounded current skip details | Action was `run`, outcome is skipped with reason | Reconsidered later; a changed effective upstream outcome stales succeeded downstream work; identical extraction disabled skip is stable. |
| Failure | Current stage becomes failed with error; current output/log/config/metadata payload is cleared | Action `run`, failed outcome and overall failed run | Current execution stops; affected succeeded downstream work is intended to stale; later invocation reruns non-succeeded stages. |
| Interruption | Model admits `interrupted`, but no production transition reference has yet been found; process death can leave persisted `running` state | A process death can leave the run non-terminal; exact recovery semantics are unknown | CLI promises continuation of interrupted/partial sessions because non-succeeded stages run; explicit status normalization is `RSK-001` for Stage 2. |
| Forced replacement | Target execution starts fresh; succeeded downstream records are pre-marked stale; current target payload clears on running | Force flag and `run` action recorded | Replacement result determines later execution; locks and safety policy remain authoritative. |
| Non-resumable success | Prior success becomes stale while retaining details long enough for diagnosis/validation, then running clears them | Current invocation records execution after validation rejects skip | Obsolete result reruns; unsafe inability to decide stops without silently replacing current success. |
| Successful rerun | Target becomes succeeded with only new outputs/diagnostics/config/metadata | Current invocation records its own new success; earlier run manifests remain immutable | Changed effective outcome stales succeeded downstream work; identical effective outcome should avoid unnecessary invalidation. |
Stage 2 must separately verify save failures before and after each session/run
transition, first- and last-stage behavior, and which fields are retained in
historical invocation records.
## Cross-Boundary Scenario Assignments
| Scenario | Primary audit stage | Supporting packages and focused tests |
| --- | --- | --- |
| 1. Success becomes non-resumable, rerun fails, later reuse decision | 2 | `internal/app`, `internal/manifest`, `stage.ResumeValidator`; `runner_test.go`, `extract_lifecycle_test.go`, manifest transition tests |
| 2. Forced/changed upstream outcome with succeeded, self-skipped, disabled downstream | 2 | `internal/app`, `internal/stage`, `internal/manifest`; run-control, runner, extraction-lifecycle tests; disabled-stage detail revisited in 8 |
| 3. Extraction bundle followed by configuration/transitive-input change | 9 | Extract/resume, Notarius adapter, artifacts/fileops tests; downstream resolution revisited in 10 |
| 4. Published/restored/prepared previous state consumed locally by analyze | 5 | `internal/app`, `internal/previouscache`, `internal/audio`, `internal/artifacts`; restore/prepare tests; analyze consumption revisited in 10 |
| 5. Publish failure at every upload boundary, then status/restore/retry | 4 | Publish stage, storage fake/adapter, app status/restore; publish and operator-helper tests; restore interpretation revisited in 5 |
| 6. Restore identical/conflict/unsafe/cache/pre-manifest-install cases | 5 | Restore discovery/plan/execute/report, artifacts, previouscache, audio; restore test suite |
| 7. Cleanup after skipped/failed/locked/partial/committed publish | 4 | Publish metadata, post-publish cleanup, cleanup targets, pathsafe; publish/cleanup tests |
| 8. Cancellation through workers, HTTP, subprocess, storage, manifests | 7 | Adapter and subprocess tests; transcribe/stage tests in 8; runner reporting in 2 |
| 9. Disabled/unselected/reused/generated/extraction/previous source then publish filtering | 10 | Analyze, artifact catalog/resolver/policy, publish tests; config ownership in 6 and publish result in 4 |
| 10. Concurrent same-session invocation and lock cleanup failures | 3 | Runner lock lifetime in 2; local artifact store, path/file cleanup and lock tests in 3 |
## Preliminary Risk-To-Test Matrix
This matrix identifies intended owners only. It makes no sufficiency judgment.
| Architectural invariant or risk | Implementation owner | Intended test owner |
| --- | --- | --- |
| One deterministic canonical stage order | `internal/stage`, planner in `internal/app` | `internal/app/planner_test.go`, narrow registry tests |
| Session manifest is cross-invocation authority; run manifest is immutable invocation audit | `internal/app`, `internal/manifest` | Manifest transition tests plus assembled runner/run-stage tests |
| First run, skip, self-skip, failure, force, invalidation, and rerun transitions | `internal/app`, `internal/manifest` | App lifecycle tests as primary; manifest helpers own field mutation |
| Obsolete versus unsafe resume validation | Stage-specific `ResumeValidator`, runner | Extract resume tests plus runner integration tests |
| Run-local validation before canonical materialization | Individual stages and `run_local.go` | Focused stage package behavioral tests; fileops owns atomic mechanism |
| Strict config, defaults, identity, and cross-field validation | `internal/config` | Config package tests; example load/validation test samples assembly |
| Canonical path/key ownership and traversal confinement | Artifacts, artifactpolicy, pathsafe | Owning package tests; app/stage tests only for assembled policy |
| Immutable extraction promotion and provenance/checksum validation | Extract, fileops, artifacts, Notarius adapter | Fileops mechanism, extract behavior, artifact hydration, adapter contract tests |
| Deterministic artifact dependency and source resolution | Artifacts, artifactpolicy, analyze | Artifact/package tests and analyze package behavior tests |
| Previous-session consumption remains local in analyze | Previouscache/prepare/artifacts/analyze | Previouscache and prepare tests; one analyze boundary test for no remote call |
| Remote current pointer is publish's final commit point | Publish stage | Publish tests with stateful object-store fake; storage tests own transport only |
| Restore is confined, deterministic, conflict-safe, and installs manifest last | Restore app modules, artifacts/previouscache/audio | Restore plan/execution/workflow tests plus low-level path/file tests |
| Cleanup requires explicit scope and committed publish metadata | App cleanup modules, pathsafe | Cleanup-target and post-publish integration tests |
| Session single-writer lock and safe release | Local artifact store, app lifetime | Artifact local-store tests plus assembled concurrent runner tests |
| Adapter cancellation, error adaptation, and resource closure | Each adapter and shared subprocess package | Focused adapter boundary tests; stage tests sample propagation |
| Bounded deterministic transcription concurrency | Transcribe stage and WhisperX client | Stage concurrency/result-order tests; HTTP adapter retry/cancel tests |
| Secrets never persist or appear in diagnostics | Config/app composition and each adapter/logging boundary | Owning config/adapter tests plus selected assembled redaction checks |
| Default suite remains deterministic, offline, and credential-free | Every package; automation | Stage 12 repository-wide execution and test-policy audit |
## Candidate Register
No candidate is confirmed by Stage 1. Later owning stages must inspect the
implementation, focused tests, canonical contract, realistic scenario, and
callers before promoting or rejecting it.
### `ARC-001`: `IODecl` is not a complete or consistently classified stage contract
- Category: architectural boundary/ownership candidate.
- Evidence: `prepare.Declares` lists files it produces under `Inputs`;
`analyze.Declares` omits dynamic input families and has no outputs;
`publish.Declares` exposes only the manifest; and `notify.Declares` advertises
placeholder paths although its result has no persisted output. No production
caller of `Declares` was found.
- Contract tension: architecture says every stage declares required inputs,
produced output state, configuration, adapters, lifecycle, and failure
behavior; the Go interface declares only partial static artifacts.
- Realistic risk: a future planner, validator, or operator view could treat the
interface as authoritative and make incorrect dependency or readiness
decisions. Current likelihood appears low because the method has no
production caller.
- Confirmation owners: Stages 8-10 for dynamic contracts, then Stage 11 for
interface purpose/simplification. Smallest plausible outcome may be clearer
naming/documentation, a complete contract, or removal; do not choose yet.
### `ARC-002`: disabled-stage “skip” terminology spans two different durable outcomes
- Category: architectural/lifecycle ownership candidate.
- Evidence: production use of `StageDispositionSkipped` was found only in
extraction. Disabled render, absent/no-op analyze, and disabled publish return
zero-disposition results with skip metadata, which the runner treats as
success. Focused and operator docs use “skip” for several of these cases,
while manifest docs reserve self-skip for a durable skipped state.
- Realistic risk: maintainers or operator features may assume all disabled
outcomes clear state, are reconsidered, and invalidate downstream work in the
same way. Conversely, changing them to explicit self-skip could break valid
pipeline continuation or cleanup semantics.
- Confirmation owners: Stage 2 for runner truth table, Stage 4 for publish and
cleanup, Stage 8 for ordinary disabled stages. Treat wording and behavior as
unresolved until those flows are traced.
### `RSK-001`: interruption has a documented state but no mapped production transition
- Category: correctness/operational risk candidate.
- Evidence: `manifest.StatusInterrupted` is admitted and external docs promise
continuation of interrupted sessions, but graph-augmented code search found
the constant only in its declaration and an artifact rejection test. No
production transition to it was found during mapping.
- Realistic risk: a killed process may leave session/run records as `running`,
producing confusing status or dual-ledger interpretation even though the
planner reruns all non-succeeded stages.
- Confirmation owner: Stage 2 must trace load normalization, process failure
boundaries, status reporting, and next-invocation behavior before deciding
whether this is a defect, compatibility state, or unused model value.
### `TST-001`: full race baseline fails in the concurrent transcribe test
- Category: test-suite execution candidate.
- Evidence: the race detector reported concurrent slice access in
`internal/adapters/whisperx/fake.go:45` from transcribe workers in
`TestTranscribeStageTranscribesPreparedAudio`.
- Observed impact: the canonical full race command exits nonzero, weakening its
signal for other packages. The report currently points to a test fake, not a
production data race.
- Confirmation owners: Stage 8 should inspect the worker/fake contract; Stage
12 should classify suite impact and the smallest durable fix. Do not change
the fake during this investigative stage.
## Candidate Classification Log
| Candidate signal | Classification | Reason |
| --- | --- | --- |
| Graph rollups `stage -> app`, `adapters -> app`, `config -> app` | rejected as a production reversal at Stage 1 | `go list` production imports contain no lower-level import of `internal/app`; graph connections include tests and ambiguous package grouping. Reopen only with a concrete production edge. |
| Similar wrapper/manifest/adapter functions | deferred metric signals, not findings | Similarity alone does not establish duplicated policy; owning behavior stages must first establish contracts. |
| Coverage percentages | deferred diagnostic signals, not findings | Stage 12 must reason from risk and test ownership, not a numeric target. |
## Unresolved Questions And Follow-Up
- Does loading a persisted `running` stage or run normalize it to
`interrupted`, or is `interrupted` only a compatibility value?
- What exact session/run disagreement states are possible when either save
fails at each transition boundary?
- Are disabled render/analyze/publish outcomes intentionally successful so
pipeline continuation and optional outputs work, and do all operator views
describe that distinction accurately?
- Is `IODecl` intended only for display/tests, or should it own enforceable
dependency declarations?
- Does publish's omission of extract from its static prerequisite list combine
safely with every configured extraction output rule and disabled extraction?
- Which native CI runner limitations explain the absence of validation jobs in
the tag-only release workflow? Stage 12 owns the automation conclusion.
No accepted risks or final audit conclusions are recorded yet.
## Completed-Stage Evidence
### Stage 0
- Contracts and records: development guide, audit plan and sequence, all policy
documents, repository/branch/toolchain state.
- Graph evidence: refreshed moderate index at exact HEAD; architecture,
interface, complexity, similarity, fan-in, and `Execute` call trace queries.
- Commands: every baseline command listed above; Go/package/file/test and
automation inventories.
- Candidates: `TST-001`; metric signals assigned to later owners.
- Explicit no-finding conclusion: no production dependency reversal into
`internal/app` was found in the package import inventory.
- Limitation disposition: the graph excludes the executable entry point, which
was verified directly; the race failure is owned by Stages 8 and 12 and does
not prevent read-only audit work.
### Stage 1
- Contracts reviewed: architecture, testing and documentation policy; internal
overview and every focused internal document; CLI, configuration,
operations, and every integration contract.
- Code/evidence reviewed: canonical registry and stage declarations; all
modeled interfaces; production import graph; application dispatch trace;
explicit self-skip usages; interrupted-state usages; focused test ownership
references.
- Outputs: package/interface ownership, area coverage, stage contract,
lifecycle, cross-boundary scenario, and preliminary risk-to-test matrices.
- Candidates: `ARC-001`, `ARC-002`, `RSK-001`; no candidate was confirmed from
mapping evidence alone.
- Explicit no-finding conclusion: the canonical stage order agrees across the
registry, internal overview, CLI, and operations contract.
- Follow-up: all unresolved behavior has a named owner in Stages 2-12; every
area and invariant has an implementation owner and intended test owner.

332
docs/roadmap/audit-plan.md Normal file
View File

@@ -0,0 +1,332 @@
# Codebase Audit Plan
Status: proposed
## Purpose
This audit will evaluate Narratio for correctness, efficiency, maintainability,
and test-suite value. It will identify defects and credible risks, duplicated or
near-duplicated behavior, code that can be made smaller or more idiomatic, and
complex code whose remaining invariants need focused explanation.
The audit is investigative. It should produce evidence-backed findings and a
prioritized remediation backlog, not make opportunistic production changes as
it proceeds. The [Audit Sequence](audit-sequence.md) assigns this scope to
concrete execution stages.
## Authoritative Baseline
Review implemented behavior against its canonical owner rather than treating
the current implementation or tests as the specification:
- [Architecture](../policy/architecture.md) for system boundaries, dependency
direction, state and path ownership, safety properties, and pipeline
invariants;
- [Internal Overview](../internal/overview.md) and its focused internal
documents for implemented ownership and mechanics;
- [Testing Policy](../policy/testing.md) for risk-based sufficiency, durable
boundaries, test-double guidance, and test lifecycle decisions;
- the [CLI](../cli.md), [Configuration](../config.md),
[Operations](../operations.md), and [integration contracts](../integrations/)
for externally observable behavior; and
- the [Documentation Policy](../policy/documentation.md) for canonical ownership
and the distinction between current and proposed behavior.
Where code, tests, and documentation disagree, record the disagreement. Do not
assume which one is wrong until the canonical contract and caller expectations
have been traced.
## Audit Principles
1. Review correctness before cleanup. A shorter implementation is not an
improvement if it weakens a state transition, safety check, or external
contract.
2. Trace behavior across boundaries. Narratio's most important properties often
emerge from the interaction of application orchestration, stages, manifests,
artifact resolution, filesystem operations, and adapters.
3. Distinguish repeated syntax from repeated policy. Extract a helper only when
the behavior has one stable owner and the shared abstraction makes that
ownership clearer. Similar stage code may be intentionally explicit.
4. Prefer narrow, idiomatic Go over generic frameworks. In particular, proposed
refactors must preserve the explicit canonical stage sequence and must not
turn Narratio into a workflow engine or a second configuration system for
downstream tools.
5. Optimize credible work. Flag repeated I/O, hashing, serialization, remote
calls, subprocess work, allocation, or poor asymptotic behavior when the
relevant path can matter. Require a benchmark or workload argument for
performance changes whose benefit is not evident.
6. Treat comments as explanations of intent. Recommend comments for invariants,
ordering constraints, non-obvious failure policy, or security reasoning—not
as narration of ordinary Go or a substitute for simplifying code.
7. Judge tests as a suite. A test can be locally reasonable and still add no
marginal protection, while a compact test can be inadequate for a
consequential cross-component failure.
## Evidence And Finding Standard
Begin from a cleanly identified revision and record toolchain and platform
assumptions. Use the code knowledge graph to find ownership, callers, callees,
similarity candidates, high-complexity functions, and weakly protected
boundaries. Confirm every candidate by reading the implementation, its focused
tests, and the applicable contract. Text search and static analysis supplement
the graph for literals, configuration, generated files, and patterns that are
not modeled reliably.
Each finding should record:
- category: correctness defect, correctness risk, duplication, simplification,
efficiency, architectural boundary, comment/clarity, or test-suite issue;
- source locations and the affected contract or invariant;
- concrete evidence and a realistic failure or maintenance scenario;
- impact, likelihood, confidence, and estimated remediation scope separately;
- the smallest plausible improvement and its intended owner;
- tests that already protect the behavior, tests that should change or be
added, and tests that may become redundant; and
- dependencies on, or conflicts with, other findings.
Do not report a metric alone as a finding. Complexity, similarity, coverage,
fan-in, file size, and test count are prioritization signals that require manual
confirmation. Consolidate findings that share one root cause.
## Cross-Cutting Review Lenses
### Correctness And Pipeline Semantics
Construct an explicit lifecycle matrix for every stage outcome: first run,
already-succeeded skip, self-skip, failure, interruption, forced replacement,
non-resumable result, and successful rerun. Trace how each outcome changes the
session manifest, invocation manifest, downstream stage state, artifacts,
diagnostics, and cleanup eligibility.
Across the pipeline, verify:
- the registry exposes one deterministic canonical order;
- each stage's declared inputs, outputs, configuration, adapters, and manifest
effects agree with its implementation and focused documentation;
- inputs are resolved through manifest and artifact contracts rather than
incidental directory contents;
- run-local outputs are fully validated before canonical materialization;
- failure, cancellation, or process interruption cannot advertise partial work
as successful;
- force and changed outcomes invalidate exactly the intended succeeded
downstream work;
- repeated execution is idempotent where promised, and ordering is stable
wherever maps, directory reads, remote listings, or dependency graphs are
involved;
- session, campaign, run, source, checksum, contract, and external provenance
identities cannot be confused across runs; and
- errors preserve useful causes and do not expose secrets or private content.
Use fault-oriented reasoning at durability boundaries: fail immediately before
and after manifest saves, canonical renames, external process completion,
uploads, current-manifest publication, the current-run commit marker, restore
manifest installation, and cleanup. Determine which state is authoritative and
whether the next invocation recovers safely.
### Duplication And Helper Ownership
Search for exact and semantic duplication in production and tests, including:
- repeated stage setup, input resolution, output validation, run-local
materialization, metadata construction, and error adaptation;
- repeated manifest create/load/save and session/run transition handling;
- repeated adapter construction, timeout parsing, command execution, generated
configuration, log handling, and output checks;
- repeated source-ID, destination, remote-key, and path validation policy;
- repeated sorting, deduplication, checksum, copy, and atomic-write mechanics;
and
- repeated test fixtures and assertions that encode the same policy at several
layers.
For each candidate, decide whether it is coincidental similarity, a repeated
mechanism, or duplicated policy. Recommend extraction only when the helper can
have a clear package owner, a narrow contract, and callers that become easier
to understand. Prefer an unexported local helper when sharing is package-local.
Do not create a broad utility package, force unlike stage results into one data
model, or move policy into storage/file-operation helpers.
Initial similarity and complexity signals should seed, but not predetermine,
inspection of the single-stage command wrappers, session/run manifest
persistence pairs, adapter constructors, Scriptorium operations, stage fakes,
and common stage materialization paths.
### Simplification, Go Idioms, And Efficiency
Review long or branch-heavy functions for separable decisions, state
transitions, or data transformations. Pay particular attention to orchestration,
configuration validation, artifact dependency resolution, resume verification,
restore/previous-cache planning, and analyze/publish selection logic. A useful
refactor should reduce cognitive load while leaving the important ordering
visible.
Check for:
- unnecessary nesting, defensive branches made unreachable by earlier
validation, repeated normalization, and overly wide parameter lists;
- interfaces defined for hypothetical extensibility rather than a demonstrated
consumer boundary;
- manual slice, map, string, error, and filesystem logic with a clearer standard
library form;
- incorrect or inconsistent `errors.Is`/`errors.As`, wrapping, context
propagation, deferred cleanup, response-body closure, process waiting, and
goroutine/channel ownership;
- redundant filesystem scans, `stat`/checksum passes, whole-file buffering,
copying, YAML/JSON round trips, sorting, remote listings, downloads, uploads,
or adapter initialization;
- linear searches nested in loops and repeated dependency or artifact lookup
that should use an indexed map or a single planning pass;
- unbounded concurrency, leaked work after cancellation, serialized independent
work, and nondeterministic result collection; and
- obsolete dependencies, portability assumptions, and platform-sensitive path
or atomic-rename behavior.
Keep correctness and diagnosability ahead of micro-optimization. When a simpler
algorithm changes performance characteristics, specify the representative
input size and validation method.
### Comments And Local Explanation
Review high fan-in, high-complexity, security-sensitive, and commit-boundary
code after likely simplifications have been identified. Add a comment
recommendation when a maintainer needs to know why:
- state transitions or persistence operations occur in a specific order;
- a stale record intentionally retains data while another transition clears it;
- a path is checked more than once to resist traversal, symlink replacement, or
time-of-check/time-of-use hazards;
- an artifact is accepted only with particular manifest, checksum, contract, or
provenance evidence;
- a partial operation is intentionally not rolled back;
- a remote pointer or local manifest must be installed last; or
- concurrency, cancellation, compatibility, or downstream-tool behavior makes
an apparently simpler approach unsafe.
Prefer a named helper, typed state, or smaller control flow when that removes the
need for explanation. Check existing comments for stale claims as well as
missing rationale.
### Test Suite Against The Canonical Policy
Build a risk-to-test matrix rather than auditing tests file by file in
isolation. For each important behavior, identify its proper owner—parser,
validator, domain package, adapter, orchestrator, CLI, integration, or end to
end—and identify all tests that claim to protect it.
Evaluate:
- protection of data integrity, destructive operations, compatibility,
security, concurrency, idempotency, recovery, and partial failure;
- manifest transitions, force/invalidation, resume validation, atomic
materialization, publish commit order, restore install order, and cleanup
gates as assembled behaviors;
- realistic HTTP, subprocess, filesystem, and object-store boundary behavior,
including cancellation and malformed responses;
- whether higher-level tests intentionally sample lower-level behavior or
redundantly reproduce its full policy;
- whether tests assert durable outcomes or private constants, exact error text,
incidental paths, call choreography, or oversized snapshots;
- whether real fast collaborators could replace elaborate doubles, and whether
stateful fakes are realistic enough for the risk they protect;
- fixture/helper duplication, oversized test cases, and setup that obscures the
behavior under test without introducing a heavyweight test framework;
- deterministic, offline, credential-free, order-independent execution and
safe handling of environment and process-global state;
- focused fuzz candidates in parsing, normalization, source IDs, remote/local
path mapping, manifest decoding, and configuration boundaries; and
- the presence and value of a small number of representative assembled
workflows.
Use coverage only to locate unexpectedly weak consequential branches. Also
inspect packages with extensive coverage for redundant tests and refactoring
friction. For every proposed addition, deletion, or consolidation, state the
realistic defect and marginal confidence involved.
The audit baseline should include the repository's canonical commands plus
targeted diagnostic runs where supported:
```sh
go test ./...
go test -race ./...
go vet ./...
go build ./cmd/narratio
```
Use focused repeated or shuffled runs to investigate state leakage and
flakiness, and collect package/branch coverage for diagnosis. Review continuous
integration to determine whether the appropriate offline validation is enforced;
do not turn coverage percentage into a gate merely for this audit.
## Area-By-Area Inspection Map
| Area | Primary locations | What to inspect |
| --- | --- | --- |
| Process and application boundary | `cmd/narratio`, `internal/app` | Command dispatch, configuration selection, production composition, secret loading, lock lifetime, object-store initialization, context/error propagation, and separation of CLI reporting from orchestration policy. Review operator commands for consistent current-state authority and shared read-only mechanics. |
| Stage registry and runner | `internal/stage/placeholders.go`, `internal/stage/stage.go`, `internal/app/planner.go`, `internal/app/runner.go`, `internal/app/run_stage.go` | Canonical order, action decisions, resume/force/self-skip/failure transitions, downstream invalidation, session/run manifest consistency, resource lifecycle, cleanup triggering, and opportunities to decompose the runner without hiding its state machine. |
| Configuration | `internal/config` | Strict decoding, discovery and precedence, centralized defaults, normalization, templating, validation order, unknown fields, empty-value behavior, secret references, cross-field constraints, path confinement, deterministic errors, duplicated validator policy, and compatibility with maintained examples. |
| Prepare and audio | `internal/stage/prepare.go`, `internal/audio`, `internal/previouscache` | Local/S3 exclusivity, cache and spool identity, partial downloads, checksum/reuse policy, previous-session required/optional planning, deterministic input records, clearing semantics, traversal safety, and avoiding repeated remote or filesystem work. |
| Transcript stages | `internal/stage/transcribe.go`, `merge.go`, `polish.go`, `normalize.go`, `trim.go`, `render.go` | Contract parity across similar stages, bounded concurrency and cancellation, deterministic speaker/input ordering, run-local validation and canonical promotion, report/diagnostic classification, disabled behavior, and narrow opportunities for shared mechanics. |
| Extraction | `internal/stage/extract.go`, `extract_resume.go`, `internal/adapters/notarius`, `internal/fileops/directory.go` | External receipt and lane validation, configuration fingerprint limits, immutable promotion, symlink/root replacement defenses, provenance and checksum checks, immediate and cross-invocation reuse, obsolete versus unsafe outcomes, failure residue, and whether dense verification logic can be clarified without weakening it. |
| Analyze and artifact dependencies | `internal/stage/analyze.go`, `internal/artifacts`, `internal/artifactpolicy` | Source-family validation, runtime catalog state, enabled/selected/reused distinctions, topological ordering and cycle handling, required/optional inputs, local-only previous sources, deterministic metadata, repeated lookup/scanning, and ownership shared with config and publish. |
| Publish and cleanup | `internal/stage/publish.go`, `internal/app/post_publish_cleanup.go`, `internal/app/cleanup_targets.go` | Prerequisite success, output selection, locks, required/optional behavior, exclusion rules, deterministic upload set, retry/idempotency implications, current-manifest then commit-marker ordering, metadata gates, and destructive path confinement. |
| Manifest state | `internal/manifest` | Validation and backward compatibility, atomic persistence, timestamps, session/run identity, transition truth table, clearing versus retaining payload, create/load/save duplication, failure during dual-manifest updates, and whether state mutation has a single owner. |
| Artifacts, paths, and policy | `internal/artifacts`, `internal/artifactpolicy`, `internal/pathsafe` | Canonical helper coverage, ad hoc reconstruction by callers, source-ID ownership, manifest-first resolution, extraction/current-state identity, destination normalization, stable ordering, typed missing-state errors, symlink/traversal defenses, and duplicate policy across config/stages/app. |
| Restore | `internal/app/restore*.go`, `internal/previouscache`, `internal/audio` | Remote authority, confined mapping, deterministic plan actions, local conflict and force behavior, dry-run purity, temp-file installation, manifest-last ordering, partial failure/retry behavior, report accuracy, cache reuse, and shared current-state mechanics. |
| File operations | `internal/fileops`, `internal/pathsafe`, local-store code in `internal/artifacts` | Atomic-write and promotion guarantees, permissions, close/sync/rename error handling, temp cleanup, same-filesystem assumptions, replacement policy, regular-file-only traversal, symlink and root-swap resistance, lock cleanup, and portability. |
| External adapters and storage | `internal/adapters`, `internal/audio` | Transport isolation, shared subprocess mechanics versus adapter-specific policy, command/config duplication, quoting and working directories, timeouts/cancellation, stdout/stderr separation, HTTP body and retry behavior, S3 pagination/streaming/not-found mapping, credential independence, and external error adaptation. |
| Shared models and diagnostics | `internal/artifactmodel`, `internal/contracts`, `internal/logging` | Serialization and validation invariants, unnecessary conversions, ownership of shared types, stable diagnostic structure, redaction, and whether small shared packages remain cohesive. |
| Tests, examples, and automation | all `*_test.go`, `examples/`, `.woodpecker/` | Risk ownership, semantic duplication, fixture cost, policy-coupled assertions, realistic boundary tests, end-to-end sufficiency, default-suite isolation, example validation, diagnostic coverage, flakiness, runtime cost, and enforcement of canonical validation. |
## Narratio-Specific Cross-Boundary Scenarios
In addition to package-local review, trace these complete scenarios because a
modular pipeline can look correct within every package while violating an
end-to-end invariant:
1. A stage succeeds, its result becomes non-resumable, the rerun fails, and a
later invocation decides what remains usable.
2. An upstream forced or changed outcome interacts with already-succeeded,
self-skipped, and disabled downstream stages.
3. Extraction produces a valid immutable bundle, then configuration or
transitive Notarius inputs change before analyze or publish.
4. Previous-session state is published, restored or prepared into the local
cache, and consumed by analyze without an unintended remote read.
5. Publish fails at each upload boundary, especially between current manifest
and current-run pointer, followed by status, restore, and retry.
6. Restore encounters identical files, conflicting files, unsafe remote keys,
cache hits, and a failure immediately before manifest installation.
7. Automatic or manual cleanup is requested after skipped, failed, locked,
partially uploaded, and fully committed publish outcomes.
8. Cancellation reaches bounded transcription work, HTTP requests,
subprocesses, object storage, and manifest reporting without leaks or false
success.
9. A configured artifact is disabled, unselected, reused, generated from
another artifact, sourced from extraction, or sourced from a previous
session, then filtered for publish.
10. The same session is invoked concurrently, including lock contention and
cleanup/release failures.
## Completion Criteria
The audit is complete when:
- every area in the inspection map has been reviewed against its canonical
contracts and focused tests;
- the stage lifecycle matrix and cross-boundary scenarios have explicit
conclusions;
- duplication candidates have been classified rather than merely counted;
- simplification and performance recommendations explain their correctness
constraints and expected benefit;
- comment recommendations identify the non-obvious rationale to preserve;
- the test suite has a risk-based sufficiency assessment, including gaps,
redundancy, durability, execution properties, and automation;
- findings are deduplicated, evidence-backed, and ranked by risk and dependency;
and
- unresolved questions and intentionally accepted risks are recorded rather
than silently omitted.
## Execution
The [Audit Sequence](audit-sequence.md) is the canonical owner of execution
order, stage boundaries, checkpoints, validation, and audit deliverables. This
document remains the canonical owner of audit scope, review criteria, and the
finding standard.

View File

@@ -0,0 +1,734 @@
# Codebase Audit Sequence
Status: proposed
## Purpose And Relationship To The Audit Plan
This document turns the [Codebase Audit Plan](audit-plan.md) into a bounded,
execution-ready sequence. The plan owns scope, review criteria, and the finding
standard. This document owns ordering, dependencies, working records,
validation, and exit gates.
The sequence is for investigation only. Do not mix production refactors or bug
fixes into the audit. A confirmed urgent defect may justify stopping to request
a separate remediation change, but its fix is not part of this sequence.
## Audit Run Records
Create `docs/roadmap/audit-findings.md` when the audit begins. It is the single
working ledger and final audit report. Initialize it with:
- the audited revision, branch/worktree state, Go version, platform, and audit
date;
- baseline command results and timings;
- an area coverage ledger;
- the stage lifecycle matrix;
- the cross-boundary scenario matrix from the audit plan;
- a risk-to-test matrix;
- candidate and confirmed finding registers; and
- unresolved questions, accepted risks, and final conclusions.
Track each execution stage in the coverage ledger with one of `not_started`,
`in_progress`, `complete`, or `blocked`. For a completed stage, record:
- contracts, packages, files, and important symbols reviewed;
- graph traces, commands, tests, or other evidence used;
- finding and candidate IDs produced;
- explicit no-finding conclusions for reviewed high-risk behavior; and
- follow-up questions assigned to later stages.
Use stable finding IDs with these prefixes:
| Prefix | Category |
| --- | --- |
| `COR` | Confirmed correctness defect |
| `RSK` | Correctness or operational risk |
| `ARC` | Ownership or architectural-boundary issue |
| `DUP` | Duplicated mechanism or policy |
| `SIM` | Simplification or idiomatic-Go opportunity |
| `EFF` | Efficiency or resource-use issue |
| `COM` | Missing, misleading, or stale explanatory comment |
| `TST` | Test-suite gap, redundancy, brittleness, or execution issue |
Candidate IDs remain candidates until manual inspection confirms the behavior,
contract, realistic scenario, and affected callers. Rejected candidates remain
in a short classification log so later stages do not reopen them without new
evidence.
## Execution Rules
1. Pin the audit to the revision recorded in Stage 0. If the worktree or HEAD
changes, record the change and rerun every affected stage; do not silently
combine evidence from different implementations.
2. Use codebase graph search and call/data-flow traces before broad source
search. Read the exact implementation, focused tests, and canonical contract
before confirming a finding.
3. Record test-policy observations during every behavior pass. Stage 12 owns the
suite-wide conclusion but must not rediscover the suite from scratch.
4. Record cross-area observations as candidates for the stage that owns the
conclusion. Avoid producing duplicate findings from several review passes.
5. Treat baseline failures as evidence, not automatic blockers. Continue when
read-only inspection remains sound, and state the limitation. Stop only when
the repository cannot be identified, required sources are unavailable, or a
failure makes later evidence unreliable.
6. Do not exercise a suspected destructive, credentialed, paid, or live-service
path merely to prove a defect. Use source reasoning, existing safe fakes, or
a narrowly controlled offline reproduction.
7. Escalate a credible active data-loss, secret-exposure, or unsafe-cleanup
defect immediately. Preserve the evidence and do not wait for final
synthesis before reporting it.
8. A stage is complete only when its exit gate is met. A package test passing is
evidence, not proof that the review is complete.
## Sequence Overview
| Stage | Focus | Depends on | Primary result |
| --- | --- | --- | --- |
| 0 | Pin revision and establish baseline | None | Reproducible audit record |
| 1 | Contract, boundary, and lifecycle map | 0 | Review matrices and ownership map |
| 2 | Runner and manifest state machine | 1 | Lifecycle and dual-ledger conclusions |
| 3 | Paths, artifacts, and filesystem safety | 1-2 | State/path authority and mutation conclusions |
| 4 | Publish, remote commit, and cleanup | 2-3 | Commit-boundary and destructive-operation conclusions |
| 5 | Restore and remote/previous state | 2-4 | Restore authority and recovery conclusions |
| 6 | Configuration and application composition | 1-5 | Validation and wiring conclusions |
| 7 | External adapters and shared support | 3, 6 | Boundary, cancellation, and resource conclusions |
| 8 | Prepare and transcript-processing stages | 2-3, 6-7 | Ordinary stage-contract conclusions |
| 9 | Extraction vertical slice | 2-3, 6-7 | Promotion, provenance, and resume conclusions |
| 10 | Analyze and artifact dependency slice | 3, 6, 8-9 | Dependency and source-resolution conclusions |
| 11 | Cross-codebase duplication, simplicity, efficiency, and comments | 2-10 | Classified maintainability candidates |
| 12 | Test-suite policy audit | 2-11 | Risk-based suite sufficiency assessment |
| 13 | Synthesis and audit closeout | 0-12 | Final deduplicated audit report |
Stages are intentionally ordered. Later stages may resolve candidates raised by
earlier ones, but they must not invalidate an earlier stage silently. Return to
the owning stage, update its coverage record, and note the new evidence.
## Stage 0: Pin Revision And Establish Baseline
### Entry
- Repository root and `docs/development.md` are available.
- The audit plan and canonical policy documents can be read.
### Execute
1. Record `git rev-parse HEAD`, branch/detached state, `git status --short`,
`go version`, `go env GOOS GOARCH`, and the current date.
2. Confirm that the code knowledge graph represents the recorded repository and
revision; refresh the index if it is missing or stale.
3. Capture the package/file/test inventory, entry points, architecture
boundaries, high fan-in symbols, complexity signals, and similarity signals.
4. Run the default offline baseline and record wall time and failures:
```sh
go test -count=1 ./...
go test -race -count=1 ./...
go vet ./...
```
5. Build into an external temporary directory so validation does not add a
workspace binary:
```sh
audit_build_dir="$(mktemp -d)"
go build -o "$audit_build_dir/narratio" ./cmd/narratio
go test -coverprofile="$audit_build_dir/coverage.out" ./...
```
6. Inventory the repository's CI/release validation, maintained examples, fuzz
tests, golden data, opt-in tests, and generated-test update mechanisms.
### Output
- Baseline and inventory sections in `audit-findings.md`.
- Initial coverage ledger containing Stages 0-13.
- Unconfirmed metric-driven candidates, clearly labeled as such.
### Exit Gate
- Revision and environment are reproducible.
- Every baseline command has a recorded result.
- Graph freshness is known.
- Any limitation that affects later stages has an owner and disposition.
## Stage 1: Build The Contract, Boundary, And Lifecycle Map
### Entry
- Stage 0 is complete.
### Execute
1. Read the architecture, internal overview, testing policy, focused internal
documents, and the relevant CLI/configuration/operations/integration
contracts using the development guide's routing rules.
2. Map each package and important interface to its owned policy. Mark every
cross-package dependency that appears to reverse or blur the intended
direction for later confirmation.
3. Build a stage-contract matrix with canonical order, declared inputs,
outputs, configuration, adapters, skip behavior, resume validation,
materialization boundary, manifest effects, and downstream invalidation.
4. Build the lifecycle matrix required by the audit plan: first run,
already-succeeded skip, self-skip, failure, interruption, forced replacement,
non-resumable result, and successful rerun.
5. Assign each of the ten cross-boundary scenarios in the audit plan to its
primary execution stage and list supporting packages/tests.
6. Seed the risk-to-test matrix with the intended test owner for each
architectural invariant. Do not judge sufficiency yet.
### Output
- Package ownership, stage-contract, lifecycle, scenario, and preliminary
risk-to-test matrices.
- `ARC` and `RSK` candidates for apparent disagreements, without deciding from
documentation alone which artifact is wrong.
### Exit Gate
- Every area in the audit plan's inspection map has an assigned stage.
- Every architectural invariant has an implementation owner and intended test
owner.
- Unknown or contradictory contracts are explicitly recorded.
## Stage 2: Audit The Runner And Manifest State Machine
### Entry
- Stage 1 matrices are complete.
### Execute
1. Trace the entry paths into full-run and single-stage execution through
`internal/app/planner.go`, `runner.go`, `run_stage.go`, and related helpers.
2. Inspect `internal/manifest` models, validation, session/run creation,
loading, normalization, atomic saves, and all transition methods.
3. Walk every lifecycle-matrix cell through both manifests. Verify clearing
versus retention of outputs, diagnostics, generated configuration, metadata,
errors, actions, timestamps, and downstream state.
4. Reason about failures before and after each session-manifest and run-manifest
save. Determine which disagreement states are possible and how a later
invocation interprets them.
5. Review force, changed-result, self-skip, failed-result, and non-resumable
invalidation separately. Confirm behavior at the first and last canonical
stage.
6. Review session lock acquisition/release and concurrent invocation behavior,
while leaving path implementation details to Stage 3.
7. Classify the runner's complexity and repeated session/run persistence paths:
state-machine clarity, justified explicitness, candidate local helpers, and
comments that preserve ordering rationale.
8. Review focused app/manifest tests against the matrix and add observations to
the risk-to-test ledger.
### Validation
```sh
go test -count=1 ./internal/app ./internal/manifest
go test -race -count=1 ./internal/app ./internal/manifest
```
### Exit Gate
- Every lifecycle cell has a source-backed conclusion for both manifests.
- Cross-boundary scenarios 1, 2, and the lock portion of 10 are resolved or
carry explicit questions.
- All runner/manifest candidates are confirmed, rejected, or assigned to a
named later stage.
## Stage 3: Audit Paths, Artifacts, And Filesystem Safety
### Entry
- Stages 1-2 are complete.
### Execute
1. Review `internal/artifacts`, `internal/artifactpolicy`, `internal/pathsafe`,
`internal/fileops`, and local-store filesystem code.
2. Inventory canonical path and key helpers, then search callers for ad hoc
reconstruction, double normalization, mixed slash/filesystem semantics, or
policy implemented outside its owner.
3. Trace built-in, configured, extraction, previous-session, and current-state
artifact resolution. Verify identity, checksum, contract, provenance,
deterministic ordering, and typed missing-state behavior.
4. Review atomic file writes, copies, directory promotion, temp cleanup,
permission preservation, close/sync/rename errors, existing-destination
behavior, same-filesystem assumptions, and platform sensitivity.
5. Walk traversal, absolute path, broad root, symlink component, inspected-root
replacement, non-regular file, and time-of-check/time-of-use scenarios.
6. Confirm that low-level file/storage helpers receive explicit destinations
and do not infer stage, campaign, session, run, or publish policy.
7. Inspect lock-file implementation and cleanup errors to finish scenario 10.
8. Record focused test ownership and gaps without duplicating Stage 2's state
conclusions.
### Validation
```sh
go test -count=1 ./internal/artifacts ./internal/artifactpolicy ./internal/pathsafe ./internal/fileops
go test -race -count=1 ./internal/artifacts ./internal/fileops
```
### Exit Gate
- Every canonical path/key family has one identified owner.
- Every material filesystem mutation has documented confinement and atomicity
conclusions.
- Scenario 10 is resolved.
- Safety checks that appear repetitive are classified before any simplification
recommendation is made.
## Stage 4: Audit Publish, Remote Commit, And Cleanup
### Entry
- Stages 2-3 are complete.
### Execute
1. Trace publish from stage selection through object-store calls, manifest
metadata, commit-marker publication, run completion, and post-publish
cleanup.
2. Verify prerequisite stage-state checks, selected/configured/extraction
output resolution, required versus optional outputs, static and remote
locks, run-file exclusions, previous-cache inclusion, and deterministic
upload order.
3. Enumerate failures before and after every upload. Prove that
`current/run_id.txt` is written last and is the only remote-current commit
point.
4. Review retry/idempotency behavior, existing remote objects, partial uploads,
pointer/manifest disagreement, and status/restore interpretation after each
partial outcome.
5. Trace automatic and manual cleanup gates. Confirm publish execution,
`uploaded`, `current_pointer_written`, explicit policy, and confined targets
are all required at the correct boundary.
6. Confirm that `--force` cannot override publish locks or cleanup safety.
7. Review duplication between publish planning, artifact destination policy,
operator views, and cleanup metadata only after ownership is established.
### Validation
```sh
go test -count=1 ./internal/stage ./internal/app ./internal/artifacts ./internal/adapters/storage
```
### Exit Gate
- Cross-boundary scenarios 5 and 7 are resolved for every relevant failure
boundary.
- Remote-current authority and local-cleanup eligibility have explicit truth
tables.
- Publish findings distinguish stage policy from storage mechanics.
## Stage 5: Audit Restore And Remote/Previous State
### Entry
- Stages 2-4 are complete.
### Execute
1. Trace restore discovery, planning, execution, reporting, audio
materialization, and previous-cache planning through `internal/app`,
`internal/artifacts`, `internal/previouscache`, `internal/audio`, and storage.
2. Confirm remote pointer/manifest identity and campaign/session/run authority,
including missing and inconsistent current state.
3. Verify remote-to-local confinement, deterministic action ordering,
`download`/`skip_same`/`conflict` decisions, force semantics, and dry-run
purity.
4. Walk failures during download, checksum or manifest validation, atomic
install, report persistence, and the manifest-last boundary. Record the
intentional lack of rollback and retry consequences.
5. Review audio spool/cache identity, cache-hit verification, partial download
behavior, and duplicate remote/filesystem work.
6. Review previous-session requirement planning, required/optional behavior,
identity checks, published-path fallback, and deterministic local mapping.
7. Confirm which mechanics are shared with status/validate/operator commands
and which caller-specific missing-state policies must remain separate.
### Validation
```sh
go test -count=1 ./internal/app ./internal/previouscache ./internal/audio ./internal/artifacts ./internal/adapters/storage
```
### Exit Gate
- Cross-boundary scenarios 4 and 6 are resolved through retry/recovery.
- Restore authority, manifest-last installation, and partial-write behavior are
explicit.
- Previous-cache conclusions are ready for the prepare and analyze passes.
## Stage 6: Audit Configuration And Application Composition
### Entry
- Stage 1 is complete and Stages 2-5 have identified the policies that
configuration and composition must supply.
### Execute
1. Review `internal/config`, `cmd/narratio`, application command dispatch,
configuration selection, secret-file environment loading, and production
collaborator construction.
2. Trace discovery, precedence, strict YAML decoding, defaults, empty values,
normalization, session templating, and validation order across pipeline,
campaign, and session configuration.
3. Verify cross-field constraints for stage enablement, paths, timeouts,
concurrency, artifacts, Notarius, Scriptorium, publish, storage, cleanup,
audio, and previous-session behavior.
4. Compare validation logic with maintained examples and the public
configuration contract. Record contract drift rather than silently choosing
code or docs.
5. Check that filesystem secrets are loaded before the boundary that consumes
them and are excluded from logs, manifests, reports, generated files, and
errors.
6. Review conditional construction of expensive/external collaborators and
cleanup of anything with a lifecycle. Confirm test injection cannot create a
behavior different from production composition.
7. Classify repeated validators, path checks, timeout parsing, constructor
wrappers, and single-stage command wrappers by policy owner.
### Validation
```sh
go test -count=1 ./internal/config ./internal/app ./cmd/narratio
go vet ./...
```
### Exit Gate
- Every operator-visible field used by audited behavior has a traced default,
normalization, validation, and consumer.
- Composition conclusions cover enabled and disabled stages without requiring
live services or credentials.
- Maintained examples have an explicit validity conclusion.
## Stage 7: Audit External Adapters And Shared Support
### Entry
- Stages 3 and 6 are complete.
### Execute
1. Review `internal/adapters`, `internal/audio`, `internal/logging`,
`internal/contracts`, and `internal/artifactmodel` at their public package
boundaries.
2. For each HTTP, subprocess, notification, and object-storage adapter, compare
implementation with its integration contract and trace all production
callers.
3. Verify context cancellation, timeout ownership, process termination and
waiting, goroutine/channel closure, HTTP response-body closure, retries,
malformed responses, streaming, pagination, not-found mapping, and local
file cleanup.
4. Confirm command argument construction, working directory, environment,
generated configuration, stdout/stderr separation, output validation, and
external error adaptation stay inside the owning adapter.
5. Compare subprocess implementations to the shared subprocess package. Classify
repeated constructor/config/log/output mechanics separately from
adapter-specific protocol policy.
6. Review fakes for realistic state and concurrency behavior, but defer their
suite-wide value judgment to Stage 12.
7. Check shared models for avoidable conversions, stable serialization,
validation ownership, and redaction-sensitive diagnostic fields.
### Validation
```sh
go test -count=1 ./internal/adapters/... ./internal/audio ./internal/logging ./internal/contracts ./internal/artifactmodel
go test -race -count=1 ./internal/adapters/... ./internal/audio
```
### Exit Gate
- Every external resource has an explicit acquisition, cancellation, and
release conclusion.
- Transport types and protocol policy have not leaked into stages.
- Adapter duplication candidates identify the correct shared or specific
owner.
## Stage 8: Audit Prepare And Transcript-Processing Stages
### Entry
- Stages 2-3 and 6-7 are complete.
### Execute
1. Review `prepare`, `transcribe`, `merge`, `polish`, `normalize`, `trim`, and
`render` as vertical slices from resolved configuration and manifest input
through adapter call, run-local output, validation, canonical
materialization, and recorded result.
2. Verify each implementation against the Stage 1 contract matrix and focused
internal document. Record any undeclared input, output, diagnostic, config,
adapter, or skip/failure behavior.
3. For prepare, confirm local/S3 exclusivity, stable input copying,
previous-cache clearing/hydration, and deterministic manifest input records.
4. For transcribe, confirm unique speaker identities, bounded runtime
concurrency, cancellation, deterministic result ordering, adapter-returned
path identity, and partial failure behavior.
5. For transformation/render stages, confirm manifest-first resolution,
run-local paths, schema/report validation, disabled/default behavior,
canonical promotion, and diagnostic-versus-artifact classification.
6. Compare similar stage implementations for shared mechanisms only after
listing meaningful differences. Avoid a generic stage framework.
7. Add stage-focused test ownership, gaps, and redundancy candidates to the
risk-to-test matrix.
### Validation
```sh
go test -count=1 ./internal/stage ./internal/audio ./internal/previouscache ./internal/adapters/whisperx ./internal/adapters/seriatim ./internal/adapters/audita ./internal/adapters/scriptorium
go test -race -count=1 ./internal/stage ./internal/audio
```
### Exit Gate
- Every reviewed stage has a completed contract-matrix row.
- Cross-boundary scenario 8 is resolved for transcription and subprocess-backed
transformation stages.
- Similarity candidates are classified as intentional explicitness, local
helper candidates, or shared-owner findings.
## Stage 9: Audit The Extraction Vertical Slice
### Entry
- Stages 2-3 and 6-7 are complete.
### Execute
1. Trace extraction from configuration validation and composition through
transcript resolution, invocation fingerprint, Notarius execution, receipt
and lane validation, directory promotion, manifest recording, catalog
hydration, resume validation, analyze, and publish consumers.
2. Verify run-local isolation, regular-file and confined-index requirements,
required-lane policy, contract/provenance construction, checksum timing,
immutable destination identity, and no-replacement promotion.
3. Enumerate failures before and after subprocess completion, receipt parsing,
payload inspection, promotion, and manifest persistence. Determine what
remains diagnostic, durable, advertised, and reusable.
4. Walk every resume validation branch. Distinguish obsolete/missing outcomes
that trigger rerun from unsafe conditions that must stop execution.
5. Evaluate the fingerprint's intentionally observable and unobservable inputs
against documentation and force guidance.
6. Review the dense validation code for named sub-decisions and comments while
preserving the visible security proof and check ordering.
7. Confirm focused tests cover immediate reuse, cross-invocation reuse,
configuration change, payload tampering, provenance mismatch, symlinks/root
replacement, failure residue, and downstream invalidation at the correct
layers.
### Validation
```sh
go test -count=1 ./internal/stage ./internal/artifacts ./internal/fileops ./internal/adapters/notarius ./internal/app
```
### Exit Gate
- Cross-boundary scenario 3 is resolved, including transitive-input limits.
- Promotion, advertisement, and resume each have a distinct authority and
failure conclusion.
- Every proposed simplification states which security or compatibility checks
it preserves.
## Stage 10: Audit Analyze And Artifact Dependencies
### Entry
- Stages 3, 6, 8, and 9 are complete.
### Execute
1. Trace all analyze source families from configuration validation through
runtime catalog registration, availability, resolution, Scriptorium
execution/reuse, materialization, metadata, and publish selection.
2. Verify enabled, selected, executable, reused, generated, and unavailable
states are distinct and deterministic.
3. Review configured-artifact dependency validation and runtime topological
ordering for cycles, missing dependencies, stable ordering, and consistency
between configuration and execution.
4. Confirm required/optional behavior and guidance for built-in transcripts,
prepared stable inputs, configured artifacts, extraction sources, and
previous-session sources.
5. Prove previous-session resolution is local-only during analyze and that
disabled artifacts are reused only under the documented conditions.
6. Inspect repeated resolution branches, parameter width, nested lookup, and
ordering work for a smaller representation or indexed plan without merging
distinct source policies.
7. Review tests for each state transition and source family at the narrowest
stable owner, noting semantic duplication across config, artifacts, stage,
publish, and assembled runner tests.
### Validation
```sh
go test -count=1 ./internal/stage ./internal/artifacts ./internal/artifactpolicy ./internal/config ./internal/adapters/scriptorium ./internal/app
```
### Exit Gate
- Cross-boundary scenario 9 is resolved for every source family and selection
state.
- Dependency ordering and source availability have explicit determinism and
complexity conclusions.
- Config, artifact-policy, catalog, stage, and publish ownership is unambiguous
or represented by an `ARC` finding.
## Stage 11: Audit Duplication, Simplicity, Efficiency, And Comments
### Entry
- Behavior stages 2-10 are complete, so structural candidates can be judged
against known contracts.
### Execute
1. Rerun graph similarity, complexity, fan-in/fan-out, call-path, loop-depth,
scan-in-loop, and change-coupling analyses on production code. Add targeted
text/static searches for patterns the graph cannot represent.
2. Revisit all `DUP`, `SIM`, `EFF`, and `COM` candidates collected earlier.
Search for additional occurrences and trace all callers before assigning an
owner.
3. For duplication, classify coincidental syntax, shared mechanism, duplicated
policy, or deliberately explicit security/state logic. Propose only the
narrowest helper that improves ownership and comprehension.
4. For complexity, sketch the smaller control flow or data model and verify it
leaves state transitions, validation order, and commit boundaries visible.
5. For efficiency, state the input scale or call frequency, current and proposed
complexity/I/O behavior, expected benefit, and benchmark or measurement
needed. Reject micro-optimizations without a credible workload.
6. Review standard-library usage, errors, slices/maps, allocations, copying,
sorting, serialization, filesystem passes, adapter initialization, remote
calls, goroutines/channels, and interface breadth across the complete codebase.
7. Review comments only after simplification decisions. Recommend why-comments
for remaining invariants, compatibility limits, safety checks, partial
failure, and ordering; flag comments that restate code or no longer match it.
8. Check dependencies and platform assumptions for clear correctness,
portability, complexity, or maintenance consequences.
### Validation
- Run focused package tests for any behavior used to disprove or confirm a
candidate.
- Run existing benchmarks where relevant. Propose a benchmark rather than
inventing performance claims when representative measurement is absent.
### Exit Gate
- Every structural candidate is confirmed, rejected with a reason, or merged
into a stronger root-cause finding.
- No helper recommendation creates a generic workflow abstraction or moves
policy into a low-level utility.
- Every efficiency finding has a credible workload and validation method.
- Every comment finding states the non-obvious rationale that should be
preserved.
## Stage 12: Audit The Test Suite Against Policy
### Entry
- Stages 2-11 have populated the risk-to-test matrix and test observations.
### Execute
1. Complete the risk-to-test matrix. For every consequential invariant, list
the current tests, proper owner, protected defect, missing failure modes, and
overlap with other layers.
2. Review tests by behavior cluster rather than filename: parsing/validation,
domain/state, filesystem, adapters, orchestration, CLI, integration, and
representative assembled workflows.
3. Classify gaps for data integrity, destructive operations, compatibility,
security, concurrency, idempotency, recovery, cancellation, and partial
success. Confirm the gap is not credibly protected elsewhere.
4. Classify redundancy and brittleness: private constants/defaults, exact error
wording, incidental formatting/paths, mock choreography, oversized
snapshots, helper-level duplication, and the same policy repeated across
layers.
5. Review doubles using the policy order: real deterministic collaborator,
stateful fake, stub, then mock when interaction is contractual. Check that
fakes model the failure and state semantics used by the tests.
6. Inspect test helpers and large test functions for simplification and
meaningful table-driven boundaries without creating a fixture framework
whose maintenance cost exceeds its value.
7. Review determinism and isolation: credentials, network access, paid APIs,
environment, working directory, clocks, randomness, ports, temp paths,
process-global state, ordering, cleanup, and parallel execution.
8. Use coverage to investigate consequential weak branches, not as a score.
Review heavily covered behavior for marginal-value duplication as well.
9. Identify focused fuzz opportunities for parsers, YAML/JSON normalization,
source IDs, confined paths, remote/local mapping, and manifest decoding.
10. Compare local requirements with `.woodpecker/` and other automation. Record
missing enforcement as a risk/cost decision, not an assumption that every
diagnostic command belongs in CI.
11. Investigate order dependence and flakiness with bounded runs, recording
runtime and any reproducible seed:
```sh
go test -shuffle=on -count=3 ./...
go test -race -shuffle=on -count=1 ./...
```
### Exit Gate
- Every important risk has a sufficiency conclusion and one intended test
owner.
- Every proposed test addition names the realistic defect and marginal value.
- Every deletion/consolidation names the stronger remaining protection.
- Default-suite determinism, offline behavior, runtime, flakiness, and CI
enforcement have explicit conclusions.
## Stage 13: Synthesize And Close The Audit
### Entry
- Stages 0-12 meet their exit gates or have explicitly accepted limitations.
### Execute
1. Reconcile candidates and findings across stages. Merge shared root causes and
remove repeated symptoms while retaining all affected locations and
contracts.
2. Recheck every confirmed finding against current source, callers, tests, and
canonical documentation. Downgrade or reject anything supported only by a
metric or hypothetical preference.
3. Rank impact, likelihood, confidence, and remediation scope separately. Order
the recommended backlog by dependency: correctness/data safety first,
architectural ownership next, then simplification/duplication, tests,
efficiency, and comments where they remain necessary.
4. Record positive conclusions for high-risk areas where the current design and
tests are sufficient. The report should not imply that only defective areas
were reviewed.
5. Reconcile the area coverage ledger, lifecycle matrix, cross-boundary scenario
matrix, and risk-to-test matrix with the audit plan's completion criteria.
6. Record any accepted risks, ambiguous contracts, environmental limitations,
and deferred investigations with an explicit rationale and owner.
7. Check whether HEAD or the worktree changed since Stage 0. Rerun affected
stages or clearly pin the report to the original revision.
8. Validate the report and roadmap document links and run `git diff --check`.
If implementation changed during the audit, rerun the full Stage 0 validation
baseline against the final audited revision.
### Final Deliverable
`docs/roadmap/audit-findings.md` must contain:
- an executive assessment without unsupported quality scores;
- the audited revision and validation baseline;
- coverage and scenario completion summaries;
- confirmed findings ordered by dependency and risk;
- rejected candidate themes where their recurrence would otherwise waste work;
- the test-suite sufficiency assessment;
- positive conclusions and accepted risks; and
- a recommended remediation order, without implementing the remediation.
### Exit Gate
- Every completion criterion in the audit plan is satisfied or explicitly
marked limited with rationale.
- Every finding is evidence-backed, deduplicated, actionable, and assigned a
stable ID.
- No production change is included in the audit output.
- The report is sufficient to prepare a separate remediation sequence without
repeating discovery.