123 lines
4.5 KiB
Markdown
123 lines
4.5 KiB
Markdown
# Audita Release Checklist
|
|
|
|
Use this checklist before cutting a pre-1.0 or 1.0 release candidate.
|
|
|
|
## Core test pass
|
|
|
|
- Run:
|
|
- `go test ./...`
|
|
- Confirm tests pass without live LLM credentials and without Python dependencies.
|
|
|
|
## Config validation and precedence
|
|
|
|
- Validate a representative config:
|
|
- `audita config validate --config <path>`
|
|
- Inspect redacted effective config:
|
|
- `audita config print-effective --config <path>`
|
|
- Confirm precedence behavior:
|
|
- defaults -> file config -> environment -> CLI.
|
|
- Confirm missing `/etc/audita/config.yml` is non-fatal when `--config`/`AUDITA_CONFIG` are unset.
|
|
|
|
## Output schema checks
|
|
|
|
- Verify default output schema remains `bare-segments`.
|
|
- Verify `--output-schema audita-v1` emits object payload with `schema` and `version`.
|
|
- Verify unknown schema (for example `seriatim-intermediate`) fails clearly.
|
|
|
|
## Subprocess contract checks
|
|
|
|
- With `--output`, verify stdout is empty on success.
|
|
- Without `--output`, verify stdout contains transcript JSON only.
|
|
- Verify `--report-json` writes file output and does not write report JSON to stdout.
|
|
- Verify failure stderr remains human-readable and includes diagnostics path when available.
|
|
- Verify nonzero exit on failures.
|
|
|
|
## Structured LLM checks
|
|
|
|
- Verify runtime uses the Audita-owned OpenAI-compatible adapter.
|
|
- Verify structured response schemas are attached via `response_format.type=json_schema`.
|
|
- Verify diagnostics metadata includes structured schema `id/version/name/sha256`.
|
|
- Verify provider output is still locally decoded/validated before use.
|
|
|
|
## Report and diagnostics schema checks
|
|
|
|
- Verify report metadata fields:
|
|
- `report_schema_name`
|
|
- `report_schema_version`
|
|
- `output_schema`
|
|
- `config_version` when file config is used.
|
|
- Verify diagnostics artifact references exist in reports:
|
|
- transcript/normalization/chunking/invocation/effective-config artifacts
|
|
- utilization diagnostics artifact
|
|
- correction ledger artifact
|
|
- error log on failures.
|
|
|
|
## Redaction checks
|
|
|
|
- Verify secrets are redacted from:
|
|
- `effective-config.json`
|
|
- run-dir and `--report-json` reports
|
|
- LLM request/response/error diagnostics payloads.
|
|
- Verify no API keys/bearer tokens leak into fixtures or outputs.
|
|
|
|
## Prompt and validator metadata checks
|
|
|
|
- Verify prompt metadata appears in LLM request metadata diagnostics:
|
|
- `prompt_id`, `prompt_version`, `prompt_source`, `embedded_path`, `sha256`.
|
|
- Verify stable validator keys appear in report decisions/rejections.
|
|
- Verify built-in validator chains resolve and execute for default and explicit module runs.
|
|
|
|
## Utilization diagnostics checks
|
|
|
|
- Verify `utilization-diagnostics.json` exists on successful runs.
|
|
- Verify partial utilization artifact behavior on controlled failure paths.
|
|
- Verify utilization fields are structurally present and nonnegative:
|
|
- effective concurrency
|
|
- run timing
|
|
- module timing summaries
|
|
- per-validator timing summaries.
|
|
|
|
## Correction ledger checks
|
|
|
|
- Verify `correction-ledger.json` exists on successful runs.
|
|
- Verify report references ledger artifact path.
|
|
- Verify ledger dispositions include applied/rejected and skipped/failed where exercised.
|
|
- Verify validator rejection and proposal-application skip remain distinct.
|
|
|
|
## Pipeline behavior checks
|
|
|
|
- Verify default full pipeline run remains:
|
|
- `glossary`, `homophones`, `glossary`, `spoken_word`, `grammar`
|
|
- with deterministic repeated instance naming (`glossary_1`, `glossary_2`).
|
|
- Verify explicit module runs (`--modules`) still work.
|
|
|
|
## Failure and cancellation checks
|
|
|
|
- Verify controlled failure paths retain diagnostics and produce best-effort failure reports.
|
|
- Verify timeout/cancellation paths exit nonzero, do not hang, and retain failure diagnostics when initialized.
|
|
|
|
## Release fixture/idempotence checks
|
|
|
|
- Run release fixtures (`internal/cli/testdata/release`) through `go test ./...`.
|
|
- Confirm fixture checks cover:
|
|
- must-apply and must-not-apply expectations
|
|
- protected-term survival
|
|
- report and diagnostics contracts
|
|
- output-schema checks
|
|
- prompt/schema metadata diagnostics
|
|
- utilization/ledger artifacts
|
|
- idempotence-oriented second pass no-op behavior with deterministic fake responses.
|
|
|
|
## Deferred-feature guardrail
|
|
|
|
- Confirm release docs do not claim support for deferred items:
|
|
- filesystem prompt overrides
|
|
- user-configurable validator chains
|
|
- arbitrary user-supplied output schemas
|
|
- resume/start-at/stop-after execution
|
|
- diff/check/propose-only modes
|
|
- generated transcript descriptions enabled by default
|
|
- interactive review UI
|
|
- UI/server wrapper
|
|
- provider benchmarking harness.
|