Document test suite policy audit conclusions
This commit is contained in:
@@ -36,8 +36,8 @@ and this section must record the new revision.
|
|||||||
|
|
||||||
Coverage ranged from 69.8% (`internal/manifest`) to 100% (`internal/logging`)
|
Coverage ranged from 69.8% (`internal/manifest`) to 100% (`internal/logging`)
|
||||||
among tested packages. `cmd/narratio` reported 0% because it has no tests. The
|
among tested packages. `cmd/narratio` reported 0% because it has no tests. The
|
||||||
remaining package results ranged from 70.1% to 88.1%. Stage 12 owns the
|
remaining package results ranged from 70.1% to 88.1%. The Stage 12 risk-based
|
||||||
risk-based interpretation; these numbers are inventory signals only.
|
interpretation is recorded below; these numbers are diagnostic signals only.
|
||||||
|
|
||||||
### Code Graph Freshness And Structural Inventory
|
### Code Graph Freshness And Structural Inventory
|
||||||
|
|
||||||
@@ -109,7 +109,7 @@ These are prioritization signals, not findings:
|
|||||||
| 9: extraction | complete | Configuration, transcript resolution, fingerprinting, Notarius execution, receipt/index/lane validation, immutable promotion, manifest advertisement, catalog hydration, resume, and explicit analyze/publish consumption were traced as one slice. The exact focused command passed. Confirmed `COR-020`; added `DUP-007`, `SIM-002`, and `TST-009`, and refined `ARC-001`, `RSK-013`, and scenario 3. |
|
| 9: extraction | complete | Configuration, transcript resolution, fingerprinting, Notarius execution, receipt/index/lane validation, immutable promotion, manifest advertisement, catalog hydration, resume, and explicit analyze/publish consumption were traced as one slice. The exact focused command passed. Confirmed `COR-020`; added `DUP-007`, `SIM-002`, and `TST-009`, and refined `ARC-001`, `RSK-013`, and scenario 3. |
|
||||||
| 10: analyze and dependencies | complete | All five source-policy families, seven built-in catalog entries, configured-artifact execution/reuse, dependency validation/order, previous-cache locality, lifecycle, and publish selection were traced as one slice. The exact focused command passed. Confirmed `COR-021` through `COR-024` and `RSK-015`; added `ARC-007`, `DUP-008`, `SIM-003`, `COM-005`, and `TST-010`; and refined `ARC-001`, `ARC-002`, `RSK-013`, and scenario 9. |
|
| 10: analyze and dependencies | complete | All five source-policy families, seven built-in catalog entries, configured-artifact execution/reuse, dependency validation/order, previous-cache locality, lifecycle, and publish selection were traced as one slice. The exact focused command passed. Confirmed `COR-021` through `COR-024` and `RSK-015`; added `ARC-007`, `DUP-008`, `SIM-003`, `COM-005`, and `TST-010`; and refined `ARC-001`, `ARC-002`, `RSK-013`, and scenario 9. |
|
||||||
| 11: maintainability | complete | Production graph metrics, change coupling, dead-code/static patterns, all structural candidates, efficiency workloads, comments, dependencies, and platform assumptions reviewed. Confirmed the narrow `ARC`, `DUP`, and `SIM` corrections recorded below, merged lifecycle wording and extraction proof candidates, rejected generic abstractions and micro-optimizations, added `SIM-004` and `COM-006`, and passed the full normal test suite and vet. |
|
| 11: maintainability | complete | Production graph metrics, change coupling, dead-code/static patterns, all structural candidates, efficiency workloads, comments, dependencies, and platform assumptions reviewed. Confirmed the narrow `ARC`, `DUP`, and `SIM` corrections recorded below, merged lifecycle wording and extraction proof candidates, rejected generic abstractions and micro-optimizations, added `SIM-004` and `COM-006`, and passed the full normal test suite and vet. |
|
||||||
| 12: test policy | not_started | Seeded by intended owners and baseline execution observations. |
|
| 12: test policy | complete | All 749 tests were inventoried by behavior owner and consequential risk; every prior `TST` candidate was classified, five suite-wide candidates were added, and coverage, doubles, helpers, determinism, offline behavior, runtime, fuzzing, and automation were assessed. The prescribed shuffled run exposed `TST-011`; the prescribed race run reproduced only `TST-001`. |
|
||||||
| 13: synthesis | not_started | No final ranking or accepted-risk decisions yet. |
|
| 13: synthesis | not_started | No final ranking or accepted-risk decisions yet. |
|
||||||
|
|
||||||
## Area Coverage And Ownership
|
## Area Coverage And Ownership
|
||||||
@@ -133,7 +133,7 @@ has been mapped but not behaviorally audited.
|
|||||||
| File operations | `internal/fileops`, `internal/pathsafe`, local artifact store | 3 (promotion vertical slice in 9) | reviewed |
|
| File operations | `internal/fileops`, `internal/pathsafe`, local artifact store | 3 (promotion vertical slice in 9) | reviewed |
|
||||||
| External adapters and storage | `internal/adapters`, `internal/audio` | 7 | reviewed |
|
| External adapters and storage | `internal/adapters`, `internal/audio` | 7 | reviewed |
|
||||||
| Shared models and diagnostics | `internal/artifactmodel`, `internal/contracts`, `internal/logging` | 7 (maintainability revisited in 11) | reviewed |
|
| Shared models and diagnostics | `internal/artifactmodel`, `internal/contracts`, `internal/logging` | 7 (maintainability revisited in 11) | reviewed |
|
||||||
| Tests, examples, and automation | package test owners, `examples/`, `.woodpecker/` | 12 | assigned |
|
| Tests, examples, and automation | package test owners, `examples/`, `.woodpecker/` | 12 | reviewed |
|
||||||
|
|
||||||
## Package And Interface Ownership Map
|
## Package And Interface Ownership Map
|
||||||
|
|
||||||
@@ -397,9 +397,10 @@ no-output metadata so the pipeline can continue to publish; it is durably
|
|||||||
reused until forced, closing analyze's behavior portion of `ARC-002` and leaving
|
reused until forced, closing analyze's behavior portion of `ARC-002` and leaving
|
||||||
the wording gap in `COM-005`.
|
the wording gap in `COM-005`.
|
||||||
|
|
||||||
## Preliminary Risk-To-Test Matrix
|
## Intended Risk-To-Test Ownership Matrix
|
||||||
|
|
||||||
This matrix identifies intended owners only. It makes no sufficiency judgment.
|
This Stage 1 matrix identifies intended owners. The final Stage 12 sufficiency
|
||||||
|
assessment follows the accumulated behavior-pass observations below.
|
||||||
|
|
||||||
| Architectural invariant or risk | Implementation owner | Intended test owner |
|
| Architectural invariant or risk | Implementation owner | Intended test owner |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
@@ -469,6 +470,119 @@ Stage 10 test observations for this matrix:
|
|||||||
| Previous-session locality and guidance | Requirement collection covers enabled/disabled, deduplication, required-wins, and stable ordering. Prepare/restore/app tests cover planned requirements; analyze proves manifest/fallback local resolution and explicitly asserts no object-store call. | Selection is absent from the collector API (`COR-022`), while the guidance test asserts only a fragment of the malformed command and therefore preserves `COR-023`. |
|
| Previous-session locality and guidance | Requirement collection covers enabled/disabled, deduplication, required-wins, and stable ordering. Prepare/restore/app tests cover planned requirements; analyze proves manifest/fallback local resolution and explicitly asserts no object-store call. | Selection is absent from the collector API (`COR-022`), while the guidance test asserts only a fragment of the malformed command and therefore preserves `COR-023`. |
|
||||||
| Scriptorium request and output materialization | Stage tests inspect exact named input paths, vars, generated/reused metadata, render-debug, logs/configuration, and canonical materialization. Adapter tests assert deterministic flags and generated invocation configuration. | Input `artifact`/`path` fields have no consumer or request representation (`COR-024`); no stage-side link/non-regular/oversized result case covers analyze, extending `RSK-013`; and the broad execution/resolution shape remains `SIM-003`. |
|
| Scriptorium request and output materialization | Stage tests inspect exact named input paths, vars, generated/reused metadata, render-debug, logs/configuration, and canonical materialization. Adapter tests assert deterministic flags and generated invocation configuration. | Input `artifact`/`path` fields have no consumer or request representation (`COR-024`); no stage-side link/non-regular/oversized result case covers analyze, extending `RSK-013`; and the broad execution/resolution shape remains `SIM-003`. |
|
||||||
|
|
||||||
|
## Test-Suite Policy Conclusions
|
||||||
|
|
||||||
|
The graph inventory contains 904 functions in 95 test files, including 749
|
||||||
|
`Test` functions and no fuzz tests or benchmarks. Review grouped those tests by
|
||||||
|
the policy owner they protect rather than by filename. Parsing and validation
|
||||||
|
are concentrated in `config`, `artifactpolicy`, manifest decoding, and adapter
|
||||||
|
protocol tests; domain and durable state in `manifest`, `artifacts`, and app
|
||||||
|
lifecycle tests; filesystem safety in `pathsafe`, `fileops`, and the local
|
||||||
|
store; adapter contracts in their owning packages; orchestration and CLI in
|
||||||
|
`app`; and representative assembled behavior in app/stage workflow tests.
|
||||||
|
|
||||||
|
### Final risk-to-test sufficiency matrix
|
||||||
|
|
||||||
|
| Consequential invariant or risk | Current protection | Proper owner and realistic protected defect | Missing modes or cross-layer overlap | Sufficiency conclusion |
|
||||||
|
| --- | --- | --- | --- | --- |
|
||||||
|
| Canonical order, planning, and lifecycle transitions | Registry/planner tables, manifest transition tests, and app first-run/skip/force/failure/rerun cases | App lifecycle tests protect durable state-machine outcomes; manifest tests own field mutation | Run-store failures and resume-validation terminalization cannot be injected; several assembled tests repeat per-stage metadata | Strong ordinary protection; add the narrow persistence seam in `TST-002` and consolidate overlap under `TST-015` |
|
||||||
|
| Session/run identity and immutable invocation audit | Per-invocation runner tests and typed manifest load/save tests | App owns identity synchronization; manifest owns serialization | Derived identity conflicts and disagreement states remain visible only through `COR-001`/`TST-002` | Insufficient at the cross-manifest boundary; `TST-002` is the intended addition |
|
||||||
|
| Strict configuration, defaults, identity, and examples | Extensive strict-load/default/validation tables plus maintained-example loading | Config tests protect rejected language and normalized values; one example test protects assembly | A 949-line loader/validator table repeats complete YAML and broad default assertions already owned elsewhere | Behavior breadth is strong; restructure for clearer ownership under `TST-014` without reducing contract cases |
|
||||||
|
| Confined identities, paths, locks, and atomic persistence | `pathsafe`, `fileops`, artifact, cleanup, and local-store tests | Low-level owners protect no-escape/no-follow/atomicity; one app or stage case protects composition | Unsafe identity, destination-ancestor links, directory sync, stale/release lock faults, and concurrent runner cases are absent; duplicating each at every caller would be wasteful | Insufficient for confirmed safety roots; add one owner-level case per root plus representative composition under `TST-003` |
|
||||||
|
| Remote publish commit and cleanup recovery | Publish source/order/failure tests and cleanup path/effect tests | Publish owns pointer-last snapshot visibility; cleanup owns durable retry evidence | Current doubles record calls but do not model prior readable versions, accepted-with-error, barriers, or cleanup reload/retry | Insufficient for recovery/idempotency/concurrency; stateful fake and focused cases in `TST-004` |
|
||||||
|
| Restore committed-state authority and partial replacement | Restore plan/execution/workflow tables, conflict/force, malformed manifest, cache tests | Restore package owns snapshot selection, conflict policy, manifest-last commit, and resumable failure | No generation change, committed-vs-stray scope, partial forced overwrite, plan/lock race, same-size replacement, or foreign-path case | Insufficient at durable transition boundaries; package-level stateful cases in `TST-005` |
|
||||||
|
| Composition has one resolved configuration authority | Production assembly is exercised; injected environments generally leave `Env.Config` nil | App composition tests protect production-shaped dependency injection | No test intentionally diverges the two configs, so the impossible split fixture remains allowed | Insufficient seam fidelity; one authority test in `TST-006` |
|
||||||
|
| Adapter protocol, cancellation, resource, output, and secret boundaries | Strong argument/schema, retry/status, direct-child timeout, ordinary output, S3 one-page, and override-secret tests | Each adapter owns protocol behavior; shared subprocess/file acquisition owns mechanical limits and redaction | Descendant kill, inherited/raw secret leakage, scheme restriction, streaming cancellation, close observation, pagination progress, and non-regular/oversized outputs are absent | Insufficient at adversarial boundaries; targeted additions in `TST-007`, not duplicated protocol matrices |
|
||||||
|
| Bounded transcription concurrency and cancellation | Concurrency bound/order and ordinary adapter-error tests | Transcribe owns dispatch/cancel completion; WhisperX fake must honor concurrent interface use | The fake races and no barrier-controlled mid-dispatch cancellation exists | Insufficient; `TST-001` restores race signal and `TST-008` protects cancellation |
|
||||||
|
| Ordinary-stage input transitions and run-local output authority | Broad prepare and transcript-stage behavior suites | Prepare owns source replacement/deduplication; stage/shared acquisition owns requested path and multi-output commit behavior | Stale previous state, duplicate audio, alternate adapter path, and partial multi-output materialization are absent; schema tables already overlap heavily | Add only the transition/authority cases in `TST-008`; current schema and protocol coverage is otherwise sufficient |
|
||||||
|
| Extraction identity, immutable promotion, and consumer evidence | Strong adapter validation, lifecycle/reuse, catalog hydration, explicit-consumer, and low-level promotion tests | Extract owns direct-input fingerprint and lifecycle; artifact/file owners protect evidence and promotion | Direct input mutation and one assembled unsafe-root case are absent. Post-install orphan residue is not advertised authority and would require private choreography | Add the two marginal cases in `TST-009`; reject a dedicated orphan-residue test and reuse `TST-002`/`TST-003` owners |
|
||||||
|
| Analyze effective selection, optional inputs, dependencies, and passthrough contract | Strong source-family, catalog, order/cycle, reuse, local-previous, and publish-filter tests | Analyze/app composition owns effective selection; artifact owners retain source parsing | Optional built-ins, selected-disabled prerequisites, simultaneous errors, actionable guidance, and accepted passthrough fields are not protected | Insufficient at cross-owner seams; focused additions in `TST-010`, without repeating all source spellings |
|
||||||
|
| Artifact/source identifiers and remote/local mapping reject hostile structured input | Deterministic tables cover representative traversal, source families, and mappings | `pathsafe` and `artifactpolicy` are pure security-sensitive owners; restore mapping is a secondary seed source | No fuzz target probes arbitrary separators, normalization idempotence, round-trip mapping, or no-escape properties | Add focused seeded property fuzzing under `TST-013`; generic YAML/JSON parser fuzzing has lower marginal value |
|
||||||
|
| CLI parsing and operator rendering remain compatible and actionable | Command tables cover valid/invalid flags, dispatch, status/validate rendering, and semantic error fragments | App command tests own supported invocation language and actionable output | A few broad workflow cases repeat downstream policy, but no consequential CLI gap was established beyond `COR-023` guidance | Sufficient once `TST-010` asserts repaired guidance; retain semantic fragments rather than exact full prose |
|
||||||
|
| Representative assembled workflows prove boundary composition | App/stage tests run real filesystem/config/manifest collaborators with adapter fakes | App owns a small number of end-to-end success/failure cases; focused packages own details | Stage metadata checklists and a six-adapter generic failure matrix repeat focused suites and generic runner behavior | Overprotected and costly in places; consolidate under `TST-015`, retaining one assembled success and one generic failure |
|
||||||
|
| Default suite is deterministic, isolated, offline, and credential-free | Normal full suite passes in about 3.5 seconds using temp dirs, loopback HTTP, and test-binary subprocesses | Each test owns cleanup of process-global state; repository execution owns the aggregate signal | A config-secret test leaves environment values behind and fails under repeated execution; no tests use `t.Parallel` | Offline/credential isolation is sufficient; determinism is not until `TST-011` is fixed |
|
||||||
|
| Race diagnostics remain trustworthy | Baseline and shuffled race runs exercise all packages | Concurrent consumers and their doubles jointly own race-safe fixtures | WhisperX fake request capture races, causing the only observed race failure | Insufficient until `TST-001`; no second production race was observed |
|
||||||
|
| Repository automation enforces supported validation | Tag release automation cross-builds Linux, macOS, and Windows binaries | Normal change automation should protect test/vet/build compatibility before release | No PR/push test, vet, or build job exists; tag publishing is not tied to a validated revision | Insufficient; record the proportional enforcement decision in `TST-012` |
|
||||||
|
|
||||||
|
### Doubles, helpers, redundancy, and brittleness
|
||||||
|
|
||||||
|
Most tests follow the preferred collaborator order. Pure validation uses real
|
||||||
|
values; filesystem and workflow tests use real temporary directories and local
|
||||||
|
manifests; HTTP tests use loopback servers; and protocol subprocess tests run
|
||||||
|
the current test binary. Those subprocess argument assertions are interaction
|
||||||
|
contracts, not mock choreography, and should remain. Storage and publish are
|
||||||
|
the important exception: call-recording stubs cannot express version
|
||||||
|
visibility, accepted-with-error outcomes, pagination progress, or concurrent
|
||||||
|
commit barriers. `TST-004` and `TST-005` therefore call for one shared stateful
|
||||||
|
object-store fake. The WhisperX request-capture fake is the only double proven
|
||||||
|
unsafe for its actual concurrent consumer (`TST-001`); sequential fakes do not
|
||||||
|
need blanket synchronization.
|
||||||
|
|
||||||
|
The largest test, `config.TestLoadAndValidate`, spans roughly 949 lines and
|
||||||
|
combines strict YAML loading, normalization/default checks, and validation in
|
||||||
|
one repeated full-document table. It matches semantic error fragments rather
|
||||||
|
than exact full messages, but its fixture duplication obscures which layer
|
||||||
|
failed and raises change cost; `TST-014` owns a split into strict-load cases,
|
||||||
|
constructed validator tables, and a small assembly sample. By contrast, the
|
||||||
|
large artifact requirement table and protocol helper-process tests express
|
||||||
|
meaningful contract cases and should remain table-driven. Existing helpers are
|
||||||
|
mostly domain-specific setup; no generic fixture framework is justified.
|
||||||
|
|
||||||
|
The broad stage metadata checklist and six-adapter assembled failure matrix
|
||||||
|
repeat focused per-stage protocol/behavior suites and generic runner result
|
||||||
|
mapping. `TST-015` names the stronger protection that must remain before those
|
||||||
|
rows are deleted: focused owners plus one representative assembled success and
|
||||||
|
one representative terminal failure. No oversized snapshots or golden files
|
||||||
|
were found. Exact full `err.Error()` equality is rare; semantic fragments,
|
||||||
|
`errors.Is`, typed errors, and externally visible protocol arguments dominate,
|
||||||
|
so no repository-wide error-assertion rewrite is warranted.
|
||||||
|
|
||||||
|
### Determinism, coverage, fuzzing, and automation
|
||||||
|
|
||||||
|
`go test -count=1 -cover ./...` passed all packages in 3.49 seconds. Package
|
||||||
|
coverage ranged from 69.8% in `internal/manifest` to 100% in
|
||||||
|
`internal/logging`, with `cmd/narratio` at 0% because it has no test files.
|
||||||
|
These percentages are not a quality score: manifest's lowest result aligns
|
||||||
|
with the uninjectable save-failure branch in `TST-002`, while logging's 100%
|
||||||
|
does not justify more tests and the placeholder notification path remains an
|
||||||
|
architectural decision in `ARC-004`, not a coverage target.
|
||||||
|
|
||||||
|
The default suite uses no live services, paid APIs, ambient credentials, or
|
||||||
|
fixed external ports. It uses temporary paths, loopback servers, and helper
|
||||||
|
subprocesses, and no test calls `t.Parallel`; normal runtime does not justify
|
||||||
|
parallelizing process-global fixtures. A few bounded sleeps exercise actual
|
||||||
|
time/process behavior, but the observed order defect is instead deterministic:
|
||||||
|
`go test -shuffle=on -count=3 ./...` failed in 4.71 seconds with seed
|
||||||
|
`1786373771816345415` because
|
||||||
|
`TestLoadSecretsFromConfigLoadsValidFiles` leaves two secret environment
|
||||||
|
variables set. The isolated command
|
||||||
|
`go test -shuffle=1786373771816345415 -count=3 -run '^TestLoadSecretsFromConfigLoadsValidFiles$' ./internal/app`
|
||||||
|
reproduced failures on repetitions two and three. `TST-011` owns restoration of
|
||||||
|
the prior environment state.
|
||||||
|
|
||||||
|
`go test -race -shuffle=on -count=1 ./...` failed in 54.21 seconds with seed
|
||||||
|
`1786373816980315094`; the only reported race was the known WhisperX fake in
|
||||||
|
`TST-001`, now reached by `TestTranscribeStageConcurrencyBounded`. Every other
|
||||||
|
package passed. Barrier-controlled cancellation/concurrency tests should
|
||||||
|
replace additional timing dependence where `TST-008` touches that path, but no
|
||||||
|
general flakiness conclusion follows from the bounded evidence.
|
||||||
|
|
||||||
|
There are no fuzz tests. The highest-value additions are seeded property tests
|
||||||
|
for confined path normalization/join and artifact source identifiers, with
|
||||||
|
properties such as no panic, no root escape, stable normalization, and valid
|
||||||
|
remote/local round trips (`TST-013`). Existing YAML/JSON tables already protect
|
||||||
|
Narratio's own strict schemas; indiscriminate fuzzing of standard-library
|
||||||
|
decoders or every manifest/config wrapper would add less marginal value.
|
||||||
|
|
||||||
|
`.woodpecker/release.yml` runs only for tags and cross-builds release binaries
|
||||||
|
with Go 1.25 before publishing them. No automation runs the repository's local
|
||||||
|
test, vet, or build requirements on ordinary changes. `TST-012` recommends a
|
||||||
|
normal validation workflow and requires a release to consume or repeat the
|
||||||
|
validated revision. The 54-second race suite belongs in automation after
|
||||||
|
`TST-001` is fixed, at a frequency chosen against its cost; repeated shuffled
|
||||||
|
runs are suitable for scheduled/audit diagnostics rather than necessarily
|
||||||
|
every change.
|
||||||
|
|
||||||
## Path, Artifact, Filesystem, And Lock Conclusions
|
## Path, Artifact, Filesystem, And Lock Conclusions
|
||||||
|
|
||||||
### Canonical ownership and normalization
|
### Canonical ownership and normalization
|
||||||
@@ -2890,11 +3004,12 @@ credible workload large enough to justify a performance finding.
|
|||||||
|
|
||||||
## Candidate Register
|
## Candidate Register
|
||||||
|
|
||||||
The `TST` entries remain candidates for Stage 12. Structural entries retain
|
The `TST` entries now carry their Stage 12 classifications. Structural entries
|
||||||
their original locations for traceability, but their Stage 11 classifications
|
retain their original locations for traceability, but their Stage 11
|
||||||
are final for this investigative pass: confirmed recommendations remain for
|
classifications are also final for this investigative pass: confirmed
|
||||||
Stage 13 prioritization, merged entries defer to the named stronger root, and
|
recommendations remain for Stage 13 prioritization, merged entries defer to the
|
||||||
rejected entries require new evidence before reopening.
|
named stronger root, and rejected entries require new evidence before
|
||||||
|
reopening.
|
||||||
|
|
||||||
### `ARC-001`: `IODecl` is not a complete or consistently classified stage contract
|
### `ARC-001`: `IODecl` is not a complete or consistently classified stage contract
|
||||||
|
|
||||||
@@ -3082,6 +3197,9 @@ rejected entries require new evidence before reopening.
|
|||||||
valid.
|
valid.
|
||||||
- Confirmation owner: Stage 12 should classify suite impact and the smallest
|
- Confirmation owner: Stage 12 should classify suite impact and the smallest
|
||||||
durable fake fix. Do not change the fake during this investigative stage.
|
durable fake fix. Do not change the fake during this investigative stage.
|
||||||
|
- Stage 12 classification: confirmed. Protect request capture with the fake's
|
||||||
|
own synchronization and expose a safe snapshot accessor. This restores the
|
||||||
|
full race suite without changing legitimate production concurrency.
|
||||||
|
|
||||||
### `TST-002`: runner tests cannot exercise invocation-manifest save failures
|
### `TST-002`: runner tests cannot exercise invocation-manifest save failures
|
||||||
|
|
||||||
@@ -3097,6 +3215,10 @@ rejected entries require new evidence before reopening.
|
|||||||
- Confirmation owner: Stage 12 should decide the smallest persistence seam and
|
- Confirmation owner: Stage 12 should decide the smallest persistence seam and
|
||||||
representative boundary cases; avoid exhaustive choreography tests for every
|
representative boundary cases; avoid exhaustive choreography tests for every
|
||||||
mechanically identical save call.
|
mechanically identical save call.
|
||||||
|
- Stage 12 classification: confirmed. Inject the run-store/terminal-operation
|
||||||
|
boundary and retain one session-first terminal save failure plus one
|
||||||
|
resume-validation terminalization case. Repeating every `SaveRun` call order
|
||||||
|
would test private choreography rather than a distinct defect.
|
||||||
|
|
||||||
### `DUP-001`: session save duplicates the shared atomic JSON writer
|
### `DUP-001`: session save duplicates the shared atomic JSON writer
|
||||||
|
|
||||||
@@ -3302,6 +3424,12 @@ rejected entries require new evidence before reopening.
|
|||||||
- Confirmation owner: Stage 12 should select one narrow behavior-level case per
|
- Confirmation owner: Stage 12 should select one narrow behavior-level case per
|
||||||
distinct invariant and reuse shared low-level tests across callers. Avoid
|
distinct invariant and reuse shared low-level tests across callers. Avoid
|
||||||
duplicating every path spelling or persistence call sequence.
|
duplicating every path spelling or persistence call sequence.
|
||||||
|
- Stage 12 classification: confirmed and narrowed. Put unsafe identity and
|
||||||
|
symlink-ancestor cases at their low-level policy owners, retain one
|
||||||
|
destructive and one writer composition case, add a durability interaction
|
||||||
|
case at the injectable filesystem capability, and add stale/release failure
|
||||||
|
plus one concurrent assembled runner case. Do not repeat the same ancestry
|
||||||
|
table at every caller.
|
||||||
|
|
||||||
### `TST-004`: publish protocol tests do not preserve prior current state or exercise recovery
|
### `TST-004`: publish protocol tests do not preserve prior current state or exercise recovery
|
||||||
|
|
||||||
@@ -3322,6 +3450,11 @@ rejected entries require new evidence before reopening.
|
|||||||
with barriers/version semantics and one behavior-level case per distinct
|
with barriers/version semantics and one behavior-level case per distinct
|
||||||
invariant. Prefer extending current order/cleanup fixtures over duplicating
|
invariant. Prefer extending current order/cleanup fixtures over duplicating
|
||||||
all source-family tables, which already have good focused coverage.
|
all source-family tables, which already have good focused coverage.
|
||||||
|
- Stage 12 classification: confirmed. One versioned object-store fake should
|
||||||
|
prove prior-pair readability across the manifest-before-pointer window and
|
||||||
|
failure, accepted-with-error ambiguity, retry, cleanup evidence durability,
|
||||||
|
and concurrent mutation. Existing source-family tables remain the stronger
|
||||||
|
protection for selection and payload composition.
|
||||||
|
|
||||||
### `TST-005`: restore tests omit committed-snapshot and partial-transition invariants
|
### `TST-005`: restore tests omit committed-snapshot and partial-transition invariants
|
||||||
|
|
||||||
@@ -3345,6 +3478,10 @@ rejected entries require new evidence before reopening.
|
|||||||
consolidate caller rendering matrices. Avoid a mock for every mechanically
|
consolidate caller rendering matrices. Avoid a mock for every mechanically
|
||||||
identical download boundary; retain the existing manifest-invalid case as the
|
identical download boundary; retain the existing manifest-invalid case as the
|
||||||
representative pre-rename validation failure.
|
representative pre-rename validation failure.
|
||||||
|
- Stage 12 classification: confirmed. Reuse the stateful store and barriers
|
||||||
|
from `TST-004` for one package-level case per committed-scope, generation,
|
||||||
|
conflict/partial-transition, identity, and recovery root. Keep the existing
|
||||||
|
manifest-invalid case; do not mock every download.
|
||||||
|
|
||||||
### `TST-006`: injected runner environments can split configuration authority
|
### `TST-006`: injected runner environments can split configuration authority
|
||||||
|
|
||||||
@@ -3362,6 +3499,10 @@ rejected entries require new evidence before reopening.
|
|||||||
- Confirmation owner: Stage 12 should inventory intentional divergent fixtures
|
- Confirmation owner: Stage 12 should inventory intentional divergent fixtures
|
||||||
and either reject differing configs, always rebind to the resolved config, or
|
and either reject differing configs, always rebind to the resolved config, or
|
||||||
replace the dual input with a constructor that makes ownership singular.
|
replace the dual input with a constructor that makes ownership singular.
|
||||||
|
- Stage 12 classification: confirmed. No intentional divergent fixture was
|
||||||
|
found. One app composition test should prove that injected collaborators
|
||||||
|
cannot introduce a second configuration authority; production should always
|
||||||
|
bind the environment to the resolved configuration.
|
||||||
|
|
||||||
### `TST-007`: adapter tests omit adversarial liveness and output-trust boundaries
|
### `TST-007`: adapter tests omit adversarial liveness and output-trust boundaries
|
||||||
|
|
||||||
@@ -3385,6 +3526,11 @@ rejected entries require new evidence before reopening.
|
|||||||
retain protocol-specific argument/schema assertions. It should classify
|
retain protocol-specific argument/schema assertions. It should classify
|
||||||
automatic fake materialization by the stage risks it protects rather than
|
automatic fake materialization by the stage risks it protects rather than
|
||||||
mechanically rewriting every fake.
|
mechanically rewriting every fake.
|
||||||
|
- Stage 12 classification: confirmed. Add one focused case per process-tree,
|
||||||
|
inherited/raw secret, URL scheme, streaming cancellation, bounded regular
|
||||||
|
output, response closure, and pagination-progress root. Retain exact protocol
|
||||||
|
argument/schema assertions; they are observable contracts, not brittle mock
|
||||||
|
choreography.
|
||||||
|
|
||||||
### `TST-008`: ordinary-stage tests omit transition, cancellation, and output-authority boundaries
|
### `TST-008`: ordinary-stage tests omit transition, cancellation, and output-authority boundaries
|
||||||
|
|
||||||
@@ -3406,6 +3552,10 @@ rejected entries require new evidence before reopening.
|
|||||||
transcribe cancellation, and one cross-adapter output-authority contract.
|
transcribe cancellation, and one cross-adapter output-authority contract.
|
||||||
Reuse shared safe-output acquisition tests from `TST-007`; do not duplicate
|
Reuse shared safe-output acquisition tests from `TST-007`; do not duplicate
|
||||||
every schema error or stage fixture already covered.
|
every schema error or stage fixture already covered.
|
||||||
|
- Stage 12 classification: confirmed. Add exactly those transition,
|
||||||
|
duplicate-input, barrier-cancellation, requested-path, and representative
|
||||||
|
multi-output partial-materialization cases. Safe output shape/size stays at
|
||||||
|
the shared owner from `TST-007`; existing schema tables remain sufficient.
|
||||||
|
|
||||||
### `TST-009`: extraction tests omit direct-input identity and assembled late-failure boundaries
|
### `TST-009`: extraction tests omit direct-input identity and assembled late-failure boundaries
|
||||||
|
|
||||||
@@ -3431,6 +3581,12 @@ rejected entries require new evidence before reopening.
|
|||||||
downstream invalidation and catalog fixtures rather than creating a broad new
|
downstream invalidation and catalog fixtures rather than creating a broad new
|
||||||
end-to-end framework; coordinate manifest-persistence faults with `TST-002`
|
end-to-end framework; coordinate manifest-persistence faults with `TST-002`
|
||||||
and destination ancestry with `TST-003`/`COR-003`.
|
and destination ancestry with `TST-003`/`COR-003`.
|
||||||
|
- Stage 12 classification: confirmed with one rejection. Add the direct-input
|
||||||
|
lifecycle regression and one assembled unsafe-root case; keep lane size at
|
||||||
|
the shared bounded reader. Reject a dedicated post-install orphan-residue
|
||||||
|
assertion because an unadvertised orphan is not authority and exposing the
|
||||||
|
point would test private choreography. Persistence and ancestry remain owned
|
||||||
|
by `TST-002` and `TST-003`.
|
||||||
|
|
||||||
### `TST-010`: analyze tests omit effective-selection and optional built-in boundaries
|
### `TST-010`: analyze tests omit effective-selection and optional built-in boundaries
|
||||||
|
|
||||||
@@ -3456,6 +3612,92 @@ rejected entries require new evidence before reopening.
|
|||||||
rejection for the passthrough fields after their contract is chosen. Reuse
|
rejection for the passthrough fields after their contract is chosen. Reuse
|
||||||
existing fixtures for publish/source success and do not retest every
|
existing fixtures for publish/source success and do not retest every
|
||||||
artifact-policy spelling.
|
artifact-policy spelling.
|
||||||
|
- Stage 12 classification: confirmed. Add the optional-built-in table, one
|
||||||
|
selected-disabled assembled workflow, deterministic simultaneous-error
|
||||||
|
ordering, semantic actionable guidance, and the chosen passthrough behavior.
|
||||||
|
Existing source-family and publish tables remain the stronger protection and
|
||||||
|
should not be repeated end to end.
|
||||||
|
|
||||||
|
### `TST-011`: filesystem-secret test leaks process environment across repetitions
|
||||||
|
|
||||||
|
- Category: test-suite determinism defect.
|
||||||
|
- Stage 12 classification: confirmed by isolated reproduction.
|
||||||
|
- Evidence: `TestLoadSecretsFromConfigLoadsValidFiles` invokes the production
|
||||||
|
loader, which sets `NARRATIO_TEST_SECRET_A` and
|
||||||
|
`NARRATIO_TEST_SECRET_B`, but the test does not restore their prior state.
|
||||||
|
`go test -shuffle=on -count=3 ./...` failed with seed
|
||||||
|
`1786373771816345415`; an isolated same-seed, three-count invocation failed
|
||||||
|
on its second and third repetitions because the variables were already
|
||||||
|
present and correctly reported as preserved rather than loaded.
|
||||||
|
- Realistic defect and marginal value: leaked secrets make the package depend
|
||||||
|
on repetition/order and can conceal whether later cases loaded or preserved
|
||||||
|
values. Restore the exact prior environment state in cleanup, including
|
||||||
|
unset-versus-empty semantics. One cleanup repair protects the behavior; no
|
||||||
|
production change or broad environment fixture is warranted.
|
||||||
|
|
||||||
|
### `TST-012`: automation does not enforce repository validation before release
|
||||||
|
|
||||||
|
- Category: test-suite execution/automation risk.
|
||||||
|
- Stage 12 classification: confirmed.
|
||||||
|
- Evidence: `.woodpecker/release.yml` is tag-only and cross-builds/publishes
|
||||||
|
binaries. No repository automation runs `go test ./...`, `go vet ./...`, or a
|
||||||
|
normal build for pushes or proposed changes, and release publication is not
|
||||||
|
tied to a revision that passed those checks.
|
||||||
|
- Realistic defect and marginal value: a change or tagged revision can ship
|
||||||
|
despite breaking tests, static analysis, or a supported build. Add ordinary
|
||||||
|
test/vet/build validation and make release consume or repeat the validated
|
||||||
|
revision. Add race execution after `TST-001` is repaired at a frequency
|
||||||
|
justified by its roughly 54-second cost; reserve repeated shuffle runs for a
|
||||||
|
scheduled/audit diagnostic unless CI capacity supports them.
|
||||||
|
|
||||||
|
### `TST-013`: security-sensitive path and source parsers have no property fuzzing
|
||||||
|
|
||||||
|
- Category: test-suite sufficiency opportunity.
|
||||||
|
- Stage 12 classification: confirmed narrow addition.
|
||||||
|
- Evidence: no fuzz target exists. Deterministic tables cover known separators,
|
||||||
|
traversal spellings, source families, and remote/local mappings, but the
|
||||||
|
`pathsafe` and `artifactpolicy` boundaries accept attacker- or
|
||||||
|
configuration-controlled strings whose combinations are larger than those
|
||||||
|
tables.
|
||||||
|
- Realistic defect and marginal value: an unanticipated separator,
|
||||||
|
normalization sequence, or source token can escape a root, classify
|
||||||
|
inconsistently, or break a round trip. Seed the existing tables and assert no
|
||||||
|
panic, no root escape, normalization stability, and valid mapping round
|
||||||
|
trips. Do not blanket-fuzz standard YAML/JSON decoders or every typed manifest
|
||||||
|
wrapper; their Narratio-specific schema cases already have stronger tables.
|
||||||
|
|
||||||
|
### `TST-014`: configuration tests conflate loader, defaults, and validator ownership
|
||||||
|
|
||||||
|
- Category: test-suite redundancy and maintainability issue.
|
||||||
|
- Stage 12 classification: confirmed consolidation.
|
||||||
|
- Evidence: `config.TestLoadAndValidate` is roughly 949 lines with a large table
|
||||||
|
of repeated complete YAML documents followed by broad normalized/default
|
||||||
|
assertions and validator checks. It already uses semantic error substrings,
|
||||||
|
so exact prose is not the defect; fixture breadth and mixed ownership are.
|
||||||
|
- Realistic defect and marginal value: changing an unrelated required/default
|
||||||
|
field forces many validation fixtures to change and a failure does not
|
||||||
|
identify whether strict decoding, defaulting, or domain validation owns the
|
||||||
|
contract. Retain small strict-load cases, construct typed configurations for
|
||||||
|
validator tables, and keep one representative load/default/validate assembly
|
||||||
|
case. Those focused owners are the stronger remaining protection; do not
|
||||||
|
introduce a generic fixture framework.
|
||||||
|
|
||||||
|
### `TST-015`: assembled stage matrices duplicate focused behavior owners
|
||||||
|
|
||||||
|
- Category: test-suite redundancy and brittleness issue.
|
||||||
|
- Stage 12 classification: confirmed consolidation.
|
||||||
|
- Evidence: `TestStagesReturnExpectedMetadata` is a broad per-stage metadata
|
||||||
|
checklist that overlaps focused stage suites and incidental metadata, while
|
||||||
|
`TestAdapterBackedStageFailureMarksManifestFailed` repeats six adapters to
|
||||||
|
prove the runner's generic terminal-failure mapping. Similar broad success
|
||||||
|
assertions exist in `TestExecuteStagesPlaceholderSuccessUpdatesManifest`.
|
||||||
|
- Realistic defect and marginal value: a harmless metadata/default change
|
||||||
|
causes several layers to fail while real stage defects remain better
|
||||||
|
diagnosed by focused owners. Retain focused stage/adapter protocol and
|
||||||
|
behavior tests, one representative assembled success proving composition and
|
||||||
|
durable manifests, and one representative adapter failure proving generic
|
||||||
|
runner terminalization. Delete the per-stage checklist and repeated generic
|
||||||
|
failure rows only after those stronger protections are explicit.
|
||||||
|
|
||||||
### `SIM-001`: runner terminalization and persistence ordering lack a narrow owner
|
### `SIM-001`: runner terminalization and persistence ordering lack a narrow owner
|
||||||
|
|
||||||
@@ -3800,8 +4042,9 @@ rejected entries require new evidence before reopening.
|
|||||||
digest that covers profiles, prompts, modules, references, executable build,
|
digest that covers profiles, prompts, modules, references, executable build,
|
||||||
and other transitive inputs, or is documented operator-forced invalidation
|
and other transitive inputs, or is documented operator-forced invalidation
|
||||||
the intended long-term contract?
|
the intended long-term contract?
|
||||||
- Which native CI runner limitations explain the absence of validation jobs in
|
- Which automation service and cadence should carry the validation required by
|
||||||
the tag-only release workflow? Stage 12 owns the automation conclusion.
|
`TST-012`, particularly the roughly 54-second race run, and should tagged
|
||||||
|
release builds consume an already validated revision or repeat validation?
|
||||||
|
|
||||||
No accepted risks or final audit conclusions are recorded yet.
|
No accepted risks or final audit conclusions are recorded yet.
|
||||||
|
|
||||||
@@ -4275,3 +4518,62 @@ No accepted risks or final audit conclusions are recorded yet.
|
|||||||
test-policy audit. Stage 13 owns prioritization, dependency ordering, and
|
test-policy audit. Stage 13 owns prioritization, dependency ordering, and
|
||||||
accepted-risk decisions; it should treat the merged/rejected structural
|
accepted-risk decisions; it should treat the merged/rejected structural
|
||||||
dispositions above as resolved rather than reopening them from metric shape.
|
dispositions above as resolved rather than reopening them from metric shape.
|
||||||
|
|
||||||
|
### Stage 12
|
||||||
|
|
||||||
|
- Entry revision: `f387222` (`Document maintainability audit conclusions`).
|
||||||
|
Commits since the pinned audit revision modify audit documentation only, so
|
||||||
|
implementation/test evidence remains pinned to the identity recorded above.
|
||||||
|
- Policy and scope reviewed: development guide, exact Stage 12 sequence, every
|
||||||
|
repository policy document, the complete risk-to-test ledger and prior
|
||||||
|
behavior-pass observations, all ten accumulated `TST` candidates, and the
|
||||||
|
repository's test/example/automation inventory.
|
||||||
|
- Graph/source evidence: 904 test-file functions, 749 `Test` functions, 95 test
|
||||||
|
files, no fuzz tests, no benchmarks, no `t.Parallel` use, test-size and
|
||||||
|
complexity hotspots, process-global environment/directory mutation,
|
||||||
|
time-dependent cases, test doubles, helper subprocesses, loopback HTTP, and
|
||||||
|
assembled workflow overlap. Exact source inspection followed the graph for
|
||||||
|
large/mixed-owner tests, secret loading, fakes, and release automation.
|
||||||
|
- Coverage diagnostic: `go test -count=1 -cover ./...` passed all packages in
|
||||||
|
3.49 s. Tested-package coverage ranged from 69.8% to 100%; `cmd/narratio`
|
||||||
|
reported 0% because it has no tests. Percentages were used only to revisit
|
||||||
|
weak consequential branches and high-coverage duplication. Manifest's low
|
||||||
|
end aligns with `TST-002`; logging's 100% and the notification placeholder do
|
||||||
|
not justify percentage-driven additions.
|
||||||
|
- Determinism validation: `go test -shuffle=on -count=3 ./...` failed in 4.71 s
|
||||||
|
with seed `1786373771816345415` because
|
||||||
|
`TestLoadSecretsFromConfigLoadsValidFiles` leaked two environment values. The
|
||||||
|
isolated same-seed, same-test, three-count command reproduced failure on
|
||||||
|
repetitions two and three, confirming `TST-011` rather than random flakiness.
|
||||||
|
- Race validation: `go test -race -shuffle=on -count=1 ./...` failed in 54.21 s
|
||||||
|
with seed `1786373816980315094`. The only race was the already recorded
|
||||||
|
WhisperX fake request-slice mutation in `TST-001`, reached by the bounded
|
||||||
|
concurrent transcribe test; every other package passed and no second
|
||||||
|
production race was observed.
|
||||||
|
- Final documentation-only verification: `go test -count=1 ./...` passed all
|
||||||
|
packages in 3.50 s and `go vet ./...` passed in 0.78 s. `git diff --check`
|
||||||
|
reported no whitespace errors.
|
||||||
|
- Suite conclusions: normal execution is fast, offline, credential-free, and
|
||||||
|
based on temp files, loopback services, or the current test binary. Current
|
||||||
|
focused tests strongly protect ordinary parsing, protocol, lifecycle, and
|
||||||
|
source-family behavior. Confirmed additions target durable disagreement,
|
||||||
|
destructive recovery, adversarial output/resource boundaries, cancellation,
|
||||||
|
effective selection, and direct input identity. Stateful object-store
|
||||||
|
behavior is preferable to additional call-recording mocks; protocol argument
|
||||||
|
assertions remain contractual.
|
||||||
|
- Candidate disposition: confirmed `TST-001` through `TST-008` and `TST-010`;
|
||||||
|
confirmed `TST-009` after rejecting its low-value orphan-residue subcase; and
|
||||||
|
added confirmed `TST-011` through `TST-015` for environment isolation,
|
||||||
|
automation, focused fuzzing, configuration-test ownership, and assembled-test
|
||||||
|
consolidation. Every proposed addition names its realistic defect and owner;
|
||||||
|
every consolidation names the focused/representative protection that remains.
|
||||||
|
- Explicit no-finding conclusions: no live network, paid-service, ambient
|
||||||
|
credential, fixed-port, oversized snapshot, golden-file, pervasive exact-
|
||||||
|
error, generic fixture-framework, blanket fake-synchronization, or general
|
||||||
|
parallel-test opportunity was established. Existing YAML/JSON tables do not
|
||||||
|
justify indiscriminate parser fuzzing, and bounded evidence did not establish
|
||||||
|
flakiness beyond the reproducible environment leak and known fake race.
|
||||||
|
- Later-stage assignment: Stage 13 owns only deduplication, ranking,
|
||||||
|
remediation ordering, accepted-risk decisions, and final audit closeout. It
|
||||||
|
should not reopen the risk-based test ownership and marginal-value decisions
|
||||||
|
recorded here without new evidence.
|
||||||
|
|||||||
Reference in New Issue
Block a user