# Audita Operations ## Scope This document covers operational behavior for `audita process` as currently implemented: - run lifecycle; - output and report files; - diagnostics artifacts; - run-directory retention behavior; - failure inspection and recovery. For command syntax, see [`docs/cli.md`](cli.md). ## Process Run Lifecycle A `process` run performs these high-level steps: 1. load effective config (defaults + optional file + env + CLI); 2. create a per-run diagnostics directory; 3. load transcript JSON and glossary YAML; 4. parse/validate input schemas; 5. normalize transcript and compute chunking; 6. run configured modules/validators; 7. serialize output schema and write transcript output; 8. build and write process report; 9. apply run-directory retention. If a failure happens after diagnostics initialization, the run writes failure details and returns nonzero. ## Output Files Transcript output: - when `--output ` is set, corrected transcript JSON is written to that file; - when `--output` is omitted, corrected transcript JSON is written to stdout. Report output: - when `--report-json ` is set, Audita writes a process report JSON file; - the run directory also writes its own `report.json` artifact. On success with `--output`, stdout is expected to be empty. ## Diagnostics Directory By default, runs use `work_dir` from effective config (default `/tmp/audita`). Each run directory is created under the work dir using a generated ID like `run-`. Top-level diagnostics artifacts: - `source-transcript.json` - `source-transcript-parsed.json` - `normalized-transcript.json` - `normalization-summary.json` - `chunking-summary.json` - `utilization-diagnostics.json` - `correction-ledger.json` - `invocation.json` - `effective-config.json` (redacted) - `report.json` - `error.log` (failure runs) Report diagnostics metadata includes resolved paths to these artifacts. ## Correction Ledger and Utilization Diagnostics `correction-ledger.json` records correction dispositions: - `applied` - `skipped` - `rejected` - `failed` `utilization-diagnostics.json` records effective concurrency and execution timing summaries for run/module/validator activity. ## Retention Behavior Retention is controlled by `work_dir_retention` (`auto|always|never`). Current behavior: - failed runs are always retained; - `always`: successful runs are retained; - `auto`: successful runs are retained only when skipped/rejected corrections occurred; clean successful runs are removed; - `never`: successful runs are currently retained (same net retention outcome as `always` in current implementation). Even when a successful run directory is removed under `auto`, an explicit `--report-json` file is still preserved at its target path. ## Failure Inspection For failed runs: 1. read stderr for the top-level failure and diagnostics path; 2. open `error.log` in the reported run directory; 3. inspect run `report.json` (`status`, `error_phase`, `error_message`); 4. inspect related artifacts referenced by report diagnostics metadata. Typical `error_phase` values include: - `transcript_read` - `glossary_read` - `transcript_schema` - `glossary_schema` - `chunking` - `runner_setup` - `runner_execution` - `output_schema` - `serialization` - `output_write` - `stdout_write` ## Recovery Guidance Safe recovery pattern: 1. correct the immediate input/config/output-path problem; 2. rerun with `--work-dir-retention always` during debugging; 3. once stable, restore your normal retention mode. Not implemented: - resume/checkpoint APIs - remote diagnostics/report storage