diff --git a/docs/architecture.md b/docs/architecture.md new file mode 100644 index 0000000..56d3466 --- /dev/null +++ b/docs/architecture.md @@ -0,0 +1,14 @@ +# Audita Architecture Index + +This file is the entrypoint for architecture documentation. + +Core architecture overview: +- [Architecture Overview](./architecture/architecture.md) + +Focused architecture contracts: +- [Public Contract](./architecture/public-contract.md) +- [Diagnostics](./architecture/diagnostics.md) +- [Structured LLM](./architecture/structured-llm.md) +- [Validators](./architecture/validators.md) +- [Prompts](./architecture/prompts.md) +- [Output Schemas](./architecture/output-schemas.md) diff --git a/docs/architecture/architecture.md b/docs/architecture/architecture.md index 96b76e4..70193af 100644 --- a/docs/architecture/architecture.md +++ b/docs/architecture/architecture.md @@ -1,861 +1,153 @@ # Audita Architecture -## Scope and intent -This document describes: -- the architecture used in production today. +## Scope +This document describes the production architecture implemented in this repository today. -Historical rewrite details live in `docs/rewrite-notes.md`. +Audita is a single-process Go CLI that: +- loads effective runtime configuration; +- reads transcript and glossary inputs; +- normalizes and sections transcripts; +- runs a built-in module pipeline with validator chains; +- writes transcript output and run diagnostics. -## Current implementation status -Implemented today: -- Go CLI entrypoint and `audita process` wiring. -- Config defaults, env loading, CLI override precedence, and validation. -- Transcript and glossary parsing/validation. -- Deterministic transcript normalization. -- Deterministic token estimation and transcript chunking. -- Per-run diagnostics directory creation plus process-level artifacts. -- Process report JSON output with diagnostics artifact references. -- Framework foundation packages for contracts and proposal application. -- Production runner orchestration package with deterministic sequential module execution. -- Module-level report structures with applied/skipped change records. -- Runtime validator models and deterministic validators. -- Deterministic validator-chain execution in the runner with cardinality enforcement. -- Module-level validator decision/rejection reporting. -- Internal structured LLM client contract plus an Audita-owned OpenAI-compatible structured LLM adapter package. -- Bounded FIFO LLM scheduler infrastructure with context-aware permit handling. -- Runtime primary/validation LLM effective-config resolution helpers with validation inheritance. -- Generic JSON prompt/response diagnostics writer primitives with secret redaction. -- LLM-backed validator models, prompt builders, batching, and runtime execution. -- Runner wiring for LLM validators via the internal structured LLM abstraction and scheduler hooks. -- LLM validator diagnostics artifacts and report-level decision metadata paths. -- Shared LLM proposal-generation helper with structured correction-set parsing. -- Deterministic proposal-index assignment and enriched proposal mapping for shared generation. -- Proposal-generation diagnostics artifacts with secret redaction. -- Production module registry with known-key recognition and explicit unsupported-module errors. -- Production `grammar` module implementation in `internal/modules/grammar`. -- Production `glossary` module implementation in `internal/modules/glossary`. -- Production `homophones` module implementation in `internal/modules/homophones`. -- Production `spoken_word` module implementation in `internal/modules/spoken_word`. -- Explicit runtime support for `--modules grammar` through the production runner path. -- Explicit runtime support for `--modules glossary`, including repeated stages such as `--modules glossary,glossary`. -- Explicit runtime support for `--modules homophones` through the production runner path. -- Explicit runtime support for `--modules spoken_word` through the production runner path. +## Runtime entrypoints +Primary CLI commands: +- `audita process --glossary [flags]` +- `audita config validate --config ` +- `audita config print-effective [--config ]` -Current reality: -- all production modules exist and are wired into the default runtime path. -- a normal `audita process` run without `--modules` now executes the full sequence: - - `glossary` - - `homophones` - - `glossary` - - `spoken_word` - - `grammar` -- repeated glossary stages are deterministic and reported distinctly as `glossary_1` and `glossary_2`. +Command ownership lives in `internal/cli/run.go`. -## Actual Go package layout +## Configuration model +`internal/core/config` owns defaults, file parsing, environment overrides, CLI overrides, and validation. -```text -cmd/audita/ - main.go +Effective-config loading for `process` and `config print-effective` is centralized in: +- `ResolveConfigPath` +- `LoadEffectiveConfig` -internal/cli/ - run.go - -internal/core/config/ - config.go - env.go - flags.go - redaction.go - validation.go - -internal/core/schema/ - transcript.go - glossary.go - errors.go - -internal/core/io/ - files.go - -internal/core/normalization/ - normalize.go - tokens.go - -internal/core/chunking/ - sections.go - summary.go - tokens.go - -internal/core/diagnostics/ - run_dir.go - -internal/core/reporting/ - report.go - -internal/framework/contracts/ - contracts.go - -internal/framework/proposals/ - proposal.go - policy.go - preview.go - apply.go - -internal/framework/runner/ - observability.go - runner.go - -internal/framework/proposal_generation/ - generate.go - -internal/framework/modules/ - registry.go - -internal/modules/grammar/ - module.go - prompt.go - -internal/modules/glossary/ - module.go - prompt.go - -internal/modules/homophones/ - module.go - prompt.go - -internal/modules/spoken_word/ - module.go - prompt.go - -internal/framework/validators/ - models.go - deterministic.go - llm_models.go - llm_prompt_builders.go - llm_batching.go - llm_validators.go - -internal/framework/warnings/ - warnings.go - -internal/validators/ - metadata/ - metadata.go - registry.go - chains.go - proposal_shape/ - validator.go - confidence_threshold/ - validator.go - original_text_presence/ - validator.go - non_empty_corrected_text/ - validator.go - no_effect/ - validator.go - protected_terms/ - validator.go - spoken_form_plausibility/ - validator.go - meaning_reversal_review/ - validator.go - editorial_review/ - validator.go - grammar_review/ - validator.go - spoken_word_review/ - validator.go - -internal/prompts/ - registry.go - render.go - assets/ - shared/ - modules/ - validators/ - -internal/framework/llm/ - openai_compatible_client.go - scheduler.go - effective_config.go - diagnostics.go - -internal/framework/responseschema/ - registry.go - registry_test.go - -internal/cli/ - review_artifacts.go - parity_test.go - release_fixtures_test.go - testdata/ - parity/ - release/ -``` - -## Current CLI behavior -Primary commands: - -```sh -audita process --glossary [flags] -audita config validate --config -audita config print-effective [--config ] -``` - -Current runtime flow (`internal/cli/run.go`): -1. Build runtime config from: - - defaults; - - file config source (`--config`, `AUDITA_CONFIG`, or default search paths when present: `/usr/local/etc/audita/config.yml`, then `/etc/audita/config.yml`); - - environment overrides; - - CLI overrides. -2. Parse flags and apply CLI overrides. -3. Validate transcript positional argument and required `--glossary`. -4. Create per-run diagnostics directory. -5. Read transcript and glossary files. -6. Parse/validate transcript and glossary. -7. Write source transcript artifacts. -8. Normalize transcript. -9. Write normalized transcript and normalization summary artifacts. -10. Chunk normalized transcript and compute chunk summaries. -11. Write chunking summary artifact. -12. Execute runner modules sequentially: - - default run path uses configured default sequence (`glossary,homophones,glossary,spoken_word,grammar`); - - explicit `--modules` overrides the default sequence; - - test/injected module factory path remains available for deterministic runtime tests. - - each module recomputes chunks from the current working transcript, runs chunk proposal work concurrently, aggregates deterministically, validates, and applies approved proposals once. -13. Output working transcript to `--output` file or stdout. -14. Build process report metadata. -15. Optionally write `--report-json`; always write run-dir `report.json`. -16. Apply work-dir retention. - -Config command behavior (`internal/cli/run.go`): -- `audita config validate --config `: - - loads and validates a versioned YAML config file; - - does not require transcript or glossary inputs. -- `audita config print-effective [--config ]`: - - builds effective config from defaults + file config + env overrides; - - prints redacted JSON to stdout; - - does not require transcript or glossary inputs. - -Parity fixture status: -- representative Python-parity fixture coverage exists under `internal/cli/testdata/parity`; -- parity tests use fake structured LLM responses for deterministic behavior, including default full-pipeline shape assertions; -- parity comparisons intentionally ignore nondeterministic metadata (timestamps, run IDs, temp paths, token usage) and remain strict for deterministic contract fields (transcript content, module order/instance naming, applied/skipped/rejected counts, and status). -- intentional Python-vs-Go differences and open parity gaps are documented in `docs/python-parity.md`. - -Important behavior details: -- Glossary is validated and is used for explicit glossary/grammar/homophones/spoken_word module correction paths. -- Default production CLI behavior now executes the full production module sequence unless `--modules` override is supplied. -- Explicit `--modules grammar`, `--modules glossary`, `--modules homophones`, and `--modules spoken_word` continue to run production module paths with LLM-backed proposal generation and validator-chain execution. -- Default runs (without explicit module selection) perform LLM calls through production module and validator paths. -- Success path is generally quiet on stderr. -- Malformed module-stage LLM payloads degrade to validator rejections and module warnings instead of aborting the run. -- Source IDs are preserved into a canonical transcript before normalization; normalization then reassigns output IDs sequentially from `1`. - -## Implemented data contracts - -### Transcript input -Accepted top-level forms: -- bare JSON array of segments -- object with `segments` array - -Source segment contract: -- `id` optional integer -- `speaker` non-empty string -- `start` finite non-negative number -- `end` finite non-negative number with `end >= start` -- `text` non-empty string -- `categories` optional array of non-empty strings - -Additional checks: -- duplicate explicit source IDs are rejected. - -### Transcript output -Transcript output is selected through an output schema registry (`internal/core/outputschema`). - -Supported output schemas: -- `bare-segments` (default): - - top-level JSON array of normalized segments; - - each segment includes `id`, `speaker`, `start`, `end`, `text`, optional `categories`. -- `audita-v1`: - - top-level object with: - - `schema: "audita-v1"` - - `version: "v1"` - - `segments: [...]` (same normalized segment payload). - -Current status: -- `seriatim-intermediate` is not implemented yet; selecting it fails clearly as an unsupported output schema. - -Selection behavior: -- CLI: `--output-schema ` -- file config: `output.schema: ` -- precedence remains runtime-wide defaults -> file config -> env -> CLI. - -Both stdout transcript output and `--output` file output use the same selected output encoder. - -### Glossary input -YAML with `glossary` entries. Required fields per entry: -- `name`, `category`, `summary` - -Optional: -- `aliases`, `plural` - -## Implemented config/env/flag behavior -Precedence for `audita process`: -1. defaults (`config.Default()`) -2. config file (if resolved from `--config`, `AUDITA_CONFIG`, or default path) +Effective precedence for `audita process`: +1. defaults +2. config file 3. environment overrides -4. CLI flags (`ApplyCLIOverrides`) +4. CLI overrides -File-config source behavior: -- explicit `--config `: - - required to exist, otherwise process fails clearly. -- `AUDITA_CONFIG` (when `--config` is not provided): - - required to exist, otherwise process fails clearly. -- default paths `/usr/local/etc/audita/config.yml`, then `/etc/audita/config.yml` (when neither explicit source is provided): - - first existing path in that order is used; - - both missing is silently ignored. +`audita config validate` is intentionally file-only validation: +- load versioned file; +- apply onto defaults; +- validate; +- do not apply environment overrides. -Versioned file-config behavior (`internal/core/config/file_config.go`): -- supported version: `version: 1`; -- missing version fails; -- unsupported version fails; -- strict unknown-field rejection is enabled. +Supported module and output-schema keys are validated through shared catalogs: +- module keys: `internal/core/modulecatalog` +- output schemas: `internal/core/outputschema` -`api_key_env` behavior: -- file config can declare API key environment variable names for proposal/validation LLM settings; -- runtime resolves those names from the process environment during config application; -- no direct API-key value field is supported in file config. +## Pipeline and module orchestration +The built-in module sequence is configured in runtime config and executed by `internal/framework/runner` through resolved module specs. -Redaction behavior: -- effective config artifacts and `audita config print-effective` both use the same redaction path (`Config.Redacted()`), so API keys are not emitted in plaintext. - -Implemented config surfaces include: -- module list -- primary and validation LLM settings -- total/proposal/validation LLM concurrency controls -- transcript description context (`--transcript-description`) -- section token controls and target sections -- confidence thresholds -- normalization controls -- work-dir and retention mode - -Current caveat: -- LLM/module-related settings are active for default and explicit module-run paths. -- compatibility environment variables and lower-level CLI tuning flags remain available while the preferred config-driven surface is adopted. - -Transcript description behavior: -- `--transcript-description` is a process-flag input for optional user-supplied background context. -- runtime config stores this value in `Config.TranscriptDescription` after CLI trimming and length validation. -- default value is empty; empty values produce no prompt context section. -- this value is intentionally non-secret and appears in effective config and invocation metadata artifacts. - -## Implemented transcript description prompt context -Transcript description context is wired through production prompt paths: -- proposal prompts for `glossary`, `homophones`, `spoken_word`, and `grammar`; -- LLM-backed validator prompts for spoken-form plausibility, meaning reversal, editorial review, grammar review, and spoken-word review. - -Prompt guardrail semantics are consistent across modules and validators: -- transcript description is labeled as "background context only"; -- it may help interpret ambiguous terms; -- it must not override transcript content; -- the model must not invent corrections, facts, names, events, motivations, or speaker intent from this description. - -Generated transcript descriptions remain deferred and are not implemented in the current runtime. - -## Implemented embedded prompt assets -Prompt assets are now built-in embedded Markdown files under `internal/prompts/assets`: -- `assets/modules/*` for production module proposal prompts; -- `assets/validators/*` for LLM-backed validator prompts; -- `assets/shared/prompt_hardening.md` for shared prompt-injection hardening text. - -Prompt source behavior: -- built-in embedded prompts are the only supported source in current runtime; -- filesystem prompt overrides and prompt-source selection flags are not implemented. - -`internal/prompts` registry responsibilities: -- register stable prompt IDs; -- register prompt version and source metadata; -- load embedded assets; -- compute deterministic SHA-256 prompt source hashes; -- render system/user prompts with `text/template` using missing-key errors. - -Prompt metadata fields: -- `prompt_id` -- `prompt_version` -- `prompt_source` (`builtin`) -- `embedded_path` -- `sha256` - -Prompt rendering flow: -- module proposal builders construct typed template data (section JSON, glossary JSON, transcript-description block) and render via `internal/prompts`; -- validator prompt builders construct typed template data (validation payload JSON, transcript-description block) and render via `internal/prompts`. - -Shared prompt hardening: -- the same centralized hardening fragment is included in every proposal and LLM-validator prompt; -- hardening text enforces untrusted transcript handling, no instruction-following from transcript content, and no invented facts/corrections. - -Prompt metadata diagnostics flow: -- proposal-generation diagnostics request metadata includes prompt metadata; -- LLM-validator diagnostics request metadata includes prompt metadata; -- detailed prompt metadata is diagnostics-scoped today and is not yet expanded into broad report-level prompt registries. - -## Implemented structured LLM infrastructure -`internal/framework/contracts` now defines a typed structured-completion contract: -- `StructuredLLMClient.CompleteStructured(ctx, req, out)` -- caller-owned typed decode target via `out` pointer. -- caller-selected structured response schema metadata via `StructuredCompletionRequest.ResponseSchema`. - -`internal/framework/llm` provides `OpenAICompatibleClient`, a direct `net/http` adapter over OpenAI-compatible chat completions: -- configurable `base_url`, model, optional API key, retries, HTTP client, and request timeout; -- OpenAI-compatible endpoint behavior (for example OpenAI/OpenRouter/local-compatible base URLs); -- request message translation from `contracts.LLMMessage` to chat-completions messages; -- strict `response_format.type = json_schema` with registered structured response schemas (`strict: true`, schema name, and schema body); -- response metadata mapping (provider/model/token usage) into Audita-owned response types; -- API-key redaction in adapter-returned errors; -- context cancellation and timeout propagation through request contexts and HTTP client timeouts; -- bounded retry behavior for transient request failures and malformed retryable structured responses. - -Structured response schemas are owned by Audita in `internal/framework/responseschema` and currently include: -- key `correction_set`: - - id `audita.correction_set` - - version `v1` - - name `audita_correction_set_v1` - - sha256 `05f8ff3fa04f68115c0cb1859d2656f51aa5c0bae8ff2470b2d4f6f531953195` -- key `validator_decision_set`: - - id `audita.validator_decision_set` - - version `v1` - - name `audita_validator_decision_set_v1` - - sha256 `b73f4790b98fbb955f0aec5496dd8ce9a8fe14aa2f35c700b4b4e5634f106fd5` - -Provider-level structured output is treated as a guardrail, not a trust boundary: -- the adapter decodes assistant message content into caller-owned structs; -- proposal-generation and validator layers continue local validation (shape, cardinality, confidence bounds, and proposal-index semantics) before changes can be applied. - -Current runtime boundary: -- the default CLI runtime path (without explicit module selection) instantiates the full production module sequence. -- LLM calls are exercised in production in both default full-pipeline runs and explicit `--modules` runs, and in tests when fake/injected clients are used. -- normal `go test ./...` does not require real LLM credentials or Python dependencies. - -`internal/framework/llm` also provides: -- a bounded FIFO `Scheduler` for controlled concurrent LLM calls with reliable permit release on success, error, and cancellation; -- primary/validation effective-config resolution helpers, including validation inheritance fallback to total LLM concurrency settings; -- generic interaction diagnostics primitives that write machine-readable JSON artifacts for request metadata, request payload, response payload, and optional error payload with secret redaction. - -Structured LLM diagnostics behavior: -- proposal-generation and validator diagnostics include structured response schema metadata (`id`, `version`, `name`, `sha256`) when schema-driven calls are made; -- API keys and bearer tokens are redacted from request/response/error diagnostics artifacts and surfaced errors. - -Dependency posture: -- the runtime no longer depends on `instructor-go`; -- structured LLM behavior is implemented through Audita-owned code paths behind `StructuredLLMClient`. - -LLM concurrency runtime behavior: -- `total` concurrency bounds all proposal and validation LLM calls. -- `proposal` concurrency adds a proposal-only sub-cap, composed with total. -- `validation` concurrency adds a validation-only sub-cap, composed with total. -- legacy `llm-concurrency` inputs remain compatibility aliases for total concurrency. -- modules execute serially, chunk proposals run concurrently within each module, and approved proposals are applied once per module in deterministic order. - -## Implemented normalization behavior -Normalization (`internal/core/normalization`) currently: -- sorts by segment start time; -- merges adjacent same-speaker segments when constraints pass; -- uses gap-based joiners: - - gap `< ellipsis_gap` -> single space join - - gap `>= ellipsis_gap` -> `... ` join -- enforces merged duration and token-limit constraints; -- reassigns output IDs sequentially from `1`; -- returns `NormalizationSummary` with merge and skip counters. - -Note: merged categories are concatenated (not deduplicated). - -## Implemented chunking behavior -Chunking (`internal/core/chunking`) currently provides: -- deterministic heuristic token estimation; -- contiguous sectioning with section metadata; -- max/min section token validation; -- optional `target_sections` override for section-count planning; -- summary and detailed summary generation. - -Current behavior details: -- if a single segment exceeds max tokens, it is emitted as its own section (not hard-failed); -- default section count is planned from `ceil(total_tokens / max_section_tokens)`; -- section sizing targets `ceil(total_tokens / section_count)` with a deterministic forward pass; -- sections remain contiguous and ordered, and segments are never split. - -## Implemented proposal/replacement infrastructure -`internal/framework/proposals` provides deterministic proposal composition logic: -- `CorrectionProposal` and `EnrichedCorrectionProposal` models; -- replacement policies: `require_unique`, `replace_all`; -- safe preview (`PreviewProposalForSegment`) with stable skip reasons; -- deterministic apply (`ApplyProposals`) in ascending `proposal_index` order; -- applied/skipped change records suitable for reporting. - -`internal/framework/contracts` provides interfaces and run-spec metadata scaffolding, including deterministic repeated module instance naming (`ResolveModuleRunSpecs`). - -These primitives are wired into the production runner and report model. The grammar, glossary, homophones, and spoken_word modules are implemented. - -## Implemented validator runtime infrastructure -`internal/framework/validators` provides deterministic validator infrastructure: -- runtime validation request/result models; -- stable validator reason codes; -- cardinality enforcement for validator decisions: - - missing proposal indexes fail - - duplicate proposal indexes fail - - unknown proposal indexes fail -- deterministic validators: - - confidence threshold by module key/config threshold - - original-text presence against current working transcript - - non-empty corrected text - - identical/no-effect rejection - - conservative protected glossary-term guard for non-glossary modules - -`internal/framework/runner` executes module pipelines with deterministic boundaries: -- modules still execute serially over the working transcript; -- section proposal work is launched promptly and can run concurrently; -- section-level validator-chain work starts as section proposals become available (deterministic validators before LLM-backed validators); -- proposal-generation and LLM-validator calls can overlap under composed scheduler limits; -- approved proposals are still applied once per module after section work settles. - -Validator rejections are reported distinctly from proposal-application skips. - -Validator composition is now explicit and registry-backed through `internal/validators`: -- built-in validator registry with stable keys and lookup/build failure for unknown keys; -- built-in chain definitions per production module key; -- production modules resolve validator chains from those built-in definitions. - -Package ownership boundary: -- `internal/validators/` owns built-in validator construction and stable key identity. -- `internal/framework/validators` remains shared runtime machinery: - - request/result models; - - decision cardinality enforcement; - - protected-vocabulary helpers; - - generic LLM-backed validator runtime, batching, and diagnostics glue. - -Validator execution classification metadata: -- `internal/validators/metadata` defines execution class markers: - - `deterministic` - - `llm_backed` -- runner ordering uses this metadata interface rather than concrete framework validator type assertions. -- validators without classification metadata default to deterministic ordering. - -`protected_terms` construction ownership: -- `internal/validators/protected_terms.New()` builds the general (non-glossary-stage) variant. -- `internal/validators/protected_terms.NewGlossaryStage()` builds the glossary-stage variant used by glossary chains. -- both variants preserve the stable key `protected_terms`. - -Stable built-in validator keys: -- deterministic: - - `proposal_shape` - - `confidence_threshold` - - `original_text_presence` - - `non_empty_corrected_text` (historical key name; current semantics reject empty resulting segment text) - - `no_effect` - - `protected_terms` -- LLM-backed: - - `spoken_form_plausibility` - - `meaning_reversal_review` - - `editorial_review` - -Built-in module chains: -- `glossary`: - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `spoken_form_plausibility` - - `meaning_reversal_review` -- `homophones`: - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `spoken_form_plausibility` - - `meaning_reversal_review` -- `spoken_word`: - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `editorial_review` - - `meaning_reversal_review` -- `grammar`: - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `editorial_review` - - `meaning_reversal_review` - -1.0 boundary: -- validator chains are built-in and not user-configurable from config/CLI. -- existing threshold and batching knobs remain configurable. - -## Implemented LLM-backed validator infrastructure -`internal/framework/validators` now includes LLM-backed validator support: -- typed request/response models for structured LLM validation; -- prompt builders for: - - spoken-form plausibility - - meaning reversal detection -- editorial review -- grammar review -- spoken-word review -- deterministic batching by `validation_max_prompt_tokens`; -- strict cardinality validation of synthesized validator decision sets; -- malformed validator payloads reject only the affected batch with warnings; -- oversized single-proposal validator inputs reject only the affected proposal; -- transport/provider/runtime LLM call failures remain fatal. - -`internal/framework/runner` wires LLM validators into existing validator chains using: -- the internal structured LLM client abstraction (`contracts.StructuredLLMClient`); -- bounded scheduler hooks for validator call execution; -- diagnostics writer hooks for machine-readable prompt/response artifacts with secret redaction. - -## Implemented shared proposal-generation infrastructure -`internal/framework/proposal_generation` provides a reusable, prompt-agnostic helper for future real modules: -- structured request model including module key/instance, replacement policy, working transcript context, optional section metadata, glossary, config, and diagnostics context; -- structured correction-set response model (`corrections`) mapped into existing `proposals.CorrectionProposal` and `proposals.EnrichedCorrectionProposal` models; -- deterministic proposal-index assignment through a caller-provided `start_index`; -- structured LLM calls through `contracts.StructuredLLMClient` only (no direct provider calls); -- optional bounded execution through scheduler hooks (`contracts.LLMScheduler`); -- prompt/response diagnostics artifact writing via the generic `internal/framework/llm` diagnostics primitives with redaction of API keys/secrets. - -This helper only produces candidate proposals; validator-chain execution and proposal application remain runner responsibilities. - -## Implemented production module-registry scaffolding -`internal/framework/modules` now provides a production registry scaffold: -- recognizes intended module keys: - - `glossary` - - `homophones` - - `spoken_word` - - `grammar` -- supports explicit constructor registration with dependency injection for: - - run spec - - config - - glossary - - proposal/validation structured LLM clients - - proposal/validation schedulers - - diagnostics directory context -- returns explicit errors for unknown keys (`unsupported_module`). - -The `grammar`, `glossary`, `homophones`, and `spoken_word` module keys are now registered and constructible. - -## Implemented grammar production module -`internal/modules/grammar` now provides the first production module: -- prompt builder faithfully constrained to punctuation/capitalization/spacing/article cleanup; -- explicit guardrails against meaning-changing rewrites, style rewrites, summarization, and invention; -- proposal generation through `internal/framework/proposal_generation` and `contracts.StructuredLLMClient`; -- scheduler-aware proposal calls through existing `contracts.LLMScheduler` hooks; -- replacement policy `require_unique` (current runtime policy); -- validator chain integration using existing deterministic + LLM-backed validators; -- grammar confidence threshold enforcement through existing validator/config infrastructure; -- module-level reporting and diagnostics capture through existing runner/reporting paths. - -## Implemented glossary production module -`internal/modules/glossary` now provides the second production module: -- prompt builder aligned to Python glossary-module intent, constrained to glossary-backed domain/acoustic corrections; -- prompt context includes glossary names, aliases, categories, summaries, and plural forms where available; -- guardrails against broad style rewriting and against replacing unrelated terms simply because they appear in glossary entries; -- proposal generation through `internal/framework/proposal_generation` and `contracts.StructuredLLMClient`; -- scheduler-aware proposal calls through existing `contracts.LLMScheduler` hooks; -- replacement policy `replace_all` (matching Python glossary behavior); -- validator chain integration using existing deterministic + LLM-backed validators; -- glossary confidence threshold enforcement through existing validator/config infrastructure; -- module-level reporting and diagnostics capture through existing runner/reporting paths; -- explicit support for repeated glossary stages with deterministic instance names (`glossary_1`, `glossary_2`, ...), where later stages see prior-stage working transcript changes. - -## Implemented protected-term behavior -`internal/framework/validators/protected_terms.go` provides deterministic glossary-derived protected vocabulary: -- extracts protected terms from glossary names and aliases; -- includes explicit plural fields and synthetic plural forms where safe; -- deduplicates and returns stable ordering for repeatable behavior/tests. - -This vocabulary is used by deterministic validators for both glossary-stage and non-glossary-stage protection checks, keeping protected-term guardrails active across modules. - -## Implemented homophones production module -`internal/modules/homophones` now provides the third production module: -- prompt builder aligned to Python homophones-module intent, constrained to conservative homophone/near-homophone/mistranscription corrections; -- prompt context includes protected glossary names/aliases/plurals to avoid damaging known terms; -- explicit guardrails against punctuation cleanup, grammar cleanup, style rewriting, summarization, and content invention; -- proposal generation through `internal/framework/proposal_generation` and `contracts.StructuredLLMClient`; -- scheduler-aware proposal calls through existing `contracts.LLMScheduler` hooks; -- replacement policy `require_unique` (matching Python homophones behavior); -- validator chain integration using existing deterministic + LLM-backed validators; -- homophones confidence threshold enforcement through existing validator/config infrastructure; -- protected-term guardrails for non-glossary modules remain active and are exercised through the homophones path; -- module-level reporting and diagnostics capture through existing runner/reporting paths. - -## Implemented spoken_word production module -`internal/modules/spoken_word` now provides the fourth production module: -- prompt builder aligned to Python spoken_word-module intent, constrained to conservative dysfluency cleanup; -- strong prompt guardrails preserving meaning/intent/voice/named entities/domain terms and substantive content; -- explicit guardrails against summarization, style rewriting, grammar-only cleanup, punctuation-only cleanup, invention, and meaning-changing rewrites; -- proposal generation through `internal/framework/proposal_generation` and `contracts.StructuredLLMClient`; -- scheduler-aware proposal calls through existing `contracts.LLMScheduler` hooks; -- replacement policy `require_unique` (matching Python spoken_word behavior); -- validator chain integration using existing deterministic + LLM-backed validators, including strong semantic guardrails (`spoken_word_review`, `meaning_reversal_review`); -- spoken_word confidence threshold enforcement through existing validator/config infrastructure; -- protected-term guardrails for non-glossary modules remain active and are exercised through the spoken_word path; -- module-level reporting and diagnostics capture through existing runner/reporting paths. - -## Reports and diagnostics (implemented) -Current per-run artifacts include: -- `source-transcript.json` -- `source-transcript-parsed.json` -- `normalized-transcript.json` -- `normalization-summary.json` -- `chunking-summary.json` -- `utilization-diagnostics.json` -- `correction-ledger.json` -- `invocation.json` -- `effective-config.json` (redacted credentials) -- `report.json` -- `error.log` on failure - -`--report-json` writes a separate report file when requested. - -Current process reports include diagnostics metadata references for: -- diagnostics directory path; -- source transcript artifact path; -- parsed source transcript artifact path; -- normalized transcript artifact path; -- normalization summary artifact path; -- chunking summary artifact path; -- utilization diagnostics artifact path; -- correction ledger artifact path; -- invocation metadata artifact path; -- redacted effective-config artifact path; -- error-log artifact path on failure. - -Current process reports also include: -- module-level results (when runner modules execute), including applied/skipped proposal changes; -- run-level module summary totals and failed module instance metadata. -- module-level warning records for malformed proposal-generation payloads and malformed validator batches. -- module-level validator decisions and validator rejections. -- optional decision-level diagnostic artifact paths for validator LLM interactions when available. -- stable validator keys in `validator_name` fields for validator decisions/rejections. -- explicit report metadata: - - report schema name; - - report schema version; - - selected output schema; - - config file version when config file input is used. -- review/observability artifacts: - - run-level and module-level utilization/timing summaries; - - flattened correction ledger entries for applied/rejected/skipped/failed correction dispositions. - -Utilization diagnostics collection: -- collection is performed in the runner path via lightweight instrumentation around LLM scheduler and structured-client execution (`internal/framework/runner`); -- instrumentation is observational only and does not change scheduler acquisition/release semantics or module execution order; -- serialized artifact: `utilization-diagnostics.json`. - -Utilization diagnostics high-level shape: -- `effective_concurrency`: - - `total_llm`, `proposal_llm`, `validation_llm`; -- `run_timing`: - - `run_wall_time_ms`; - - `scheduler_queue_wait_ms`; - - `llm_execution_time_ms`; - - `deterministic_validation_time_ms`; - - `max_in_flight_llm_calls`; - - `average_in_flight_llm_calls`; -- `llm_calls`: - - `total_proposal_calls`; - - `total_validation_calls`; -- `modules`: - - per-module key/instance timing summaries including module wall time and per-module call counts; -- `validators`: - - per-validator summaries keyed by stable validator key with elapsed time and LLM-backed marker. - -Correction ledger construction: -- ledger entries are built from runner module results in the CLI report/diagnostics path (`internal/cli/review_artifacts.go`); -- serialized artifact: `correction-ledger.json`; -- one flattened record per applied/validator-rejected/application-skipped outcome where data is available, plus module-failed records for failed module instances. - -Correction ledger high-level shape: -- run/module/proposal identity: - - `run_id`, `module_key`, `module_instance`, `proposal_index`, `segment_id`; -- correction payload: - - `original_text`, `proposed_corrected_text`, `applied_corrected_text` (when applied), `replacement_policy`; -- disposition: - - `disposition` in `{applied,rejected,skipped,failed}`; - - `disposition_reason_code`, `disposition_message`; -- validator decision snapshots: - - `deterministic_validator_decisions[]`; - - `llm_validator_decisions[]`; - - each decision uses stable validator keys and reason codes. - -Identity and metadata boundaries: -- stable module keys/instance names and stable validator keys are included directly in ledger records; -- prompt metadata and structured response schema metadata remain in LLM interaction diagnostics payloads and are not duplicated into every ledger row; -- reports reference artifact paths for utilization and ledger files through diagnostics metadata. - -Redaction and retention: -- secret redaction guarantees continue to apply to diagnostics/report artifacts; -- utilization and ledger artifacts are emitted within the existing run-directory retention model (`auto|always|never`) and are retained/removed with the run directory. - -Current report schema metadata values: -- `report_metadata.report_schema_name = "audita-process-report"` -- `report_metadata.report_schema_version = "v1"` - -Retention modes implemented in `ApplyRetention`: -- `always`: keep all run directories. -- `never`: keep successful run directories. -- `auto`: keep failed runs and successful runs with skipped corrections. -- failed runs are always retained. - -Current runtime note: -- default non-explicit runs usually have no module-level skipped corrections, so `auto` commonly removes clean successful run directories. -- explicit grammar/glossary/homophones/spoken_word runs can produce validator rejections and application skips, which are reflected in reports and retention input. - -## Current tests and quality posture -Implemented tests currently cover: -- CLI argument handling and behavior (`internal/cli/run_test.go`) -- subprocess stdout/stderr and exit-code behavior (`cmd/audita/main_integration_test.go`) -- config/env/override validation (`internal/core/config/*_test.go`) -- transcript and glossary schema validation (`internal/core/schema/*_test.go`) -- deterministic normalization (`internal/core/normalization/*_test.go`) -- deterministic chunking and summaries (`internal/core/chunking/*_test.go`) -- proposal preview/apply semantics (`internal/framework/proposals/*_test.go`) -- contracts/foundation composition tests (`internal/framework/contracts/*_test.go`) -- runner sequencing and failure behavior with deterministic fake modules (`internal/framework/runner/*_test.go`) -- CLI runner integration through injected fake module factories (`internal/cli/run_test.go`) -- validator models, cardinality enforcement, and deterministic validators (`internal/framework/validators/*_test.go`) -- LLM-backed validator batching, prompt builders, structured-response safety, scheduler hooks, and diagnostics redaction (`internal/framework/validators/*_test.go`, `internal/framework/runner/*_test.go`) -- shared proposal-generation request/response parsing, deterministic indexing, scheduler hooks, and diagnostics redaction (`internal/framework/proposal_generation/*_test.go`, `internal/framework/runner/*_test.go`) -- production module-registry known-key recognition and unsupported/internal-registry error behavior (`internal/framework/modules/*_test.go`, `internal/cli/run_test.go`) -- production grammar module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, and explicit CLI/runtime integration (`internal/modules/grammar/*_test.go`, `internal/cli/run_test.go`, `internal/framework/runner/*_test.go`) -- production glossary module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, repeated-stage behavior, and explicit CLI/runtime integration (`internal/modules/glossary/*_test.go`, `internal/cli/run_test.go`, `internal/framework/runner/*_test.go`) -- production homophones module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, protected-term behavior, and explicit CLI/runtime integration (`internal/modules/homophones/*_test.go`, `internal/cli/run_test.go`, `internal/framework/runner/*_test.go`) -- production spoken_word module prompt constraints, proposal mapping, validator-chain behavior, semantic guardrail behavior, confidence-threshold enforcement, diagnostics redaction, protected-term behavior, and explicit CLI/runtime integration (`internal/modules/spoken_word/*_test.go`, `internal/cli/run_test.go`, `internal/framework/runner/*_test.go`) -- glossary-derived protected-term extraction and stable behavior (`internal/framework/validators/protected_terms_test.go`) -- default full-pipeline runtime shape and ordering (`internal/cli/run_test.go`, `cmd/audita/main_integration_test.go`, `internal/cli/parity_test.go`) -- subprocess operational hardening behavior including large-input, failure-mode, timeout/cancellation, backend-failure, and partial-progress paths (`cmd/audita/main_integration_test.go`) -- report/diagnostics redaction and artifact-shape behavior across success and failure paths (`internal/cli/run_test.go`, `cmd/audita/main_integration_test.go`) -- curated release-fixture and idempotence-oriented readiness checks using fake structured LLM responses (`internal/cli/release_fixtures_test.go`, `internal/cli/testdata/release`) - -## Operational hardening status -The runtime now includes hardened subprocess behavior for parent-process callers: -- deterministic success/failure exit codes; -- strict stdout/stderr separation suitable for machine orchestration; -- failure stderr summaries that include diagnostics location when available; -- retained failure diagnostics (`report.json`, `error.log`, and artifacts written before failure); -- deterministic timeout/cancellation behavior in tests; -- redaction coverage for API keys/secrets across reports, diagnostics artifacts, and surfaced errors. -- stable output routing behavior: - - with `--output`, stdout remains empty on success; - - without `--output`, stdout contains only transcript JSON in the selected output schema; - - `--report-json` writes report data to file only (never stdout). - -Operational caller guidance is documented in [`docs/subprocess-operations.md`](docs/subprocess-operations.md). - -## Final status -- Audita's default full module-sequence runtime is implemented and tested. -- Parity fixtures and operational hardening coverage are in place. -- Historical migration context is documented in [`docs/migration-from-python.md`](docs/migration-from-python.md). +Current default sequence: +- `glossary` +- `homophones` +- `glossary` +- `spoken_word` +- `grammar` + +Execution behavior: +- modules execute serially over the working transcript; +- section proposal work can run concurrently within a module; +- validator execution happens on generated proposals before application; +- approved proposals are applied once per module in deterministic proposal-index order. + +Production modules remain separate packages: +- `internal/modules/glossary` +- `internal/modules/homophones` +- `internal/modules/spoken_word` +- `internal/modules/grammar` + +## Proposal generation and prompt context +Shared proposal plumbing is centralized in `internal/framework/proposal_generation`. + +Module packages provide: +- module identity and replacement policy; +- module-specific prompt message building; +- built-in validator chain selection. + +Shared prompt payload helpers are in `internal/framework/promptcontext`. + +## Validator architecture +Built-in validator construction and chain composition live in `internal/validators`. + +Shared validator runtime mechanics live in `internal/framework/validators`. + +Execution class metadata (deterministic vs LLM-backed) is centralized in `internal/validators/metadata` and used for ordering and reporting classification. + +## Structured LLM boundary +All production LLM calls go through the internal contract: +- `contracts.StructuredLLMClient` +- `CompleteStructured(ctx, req, out)` + +The OpenAI-compatible HTTP adapter is implemented in `internal/framework/llm`. + +Structured response schemas are registered in `internal/framework/responseschema` and attached to requests via `response_format` metadata. + +Malformed structured-output detection is centralized in `internal/framework/structuredoutput` and reused by proposal generation and validator execution so downgrade behavior stays consistent. + +## Stage naming and diagnostics metadata +Diagnostics stage naming is centralized in `internal/framework/stagename`: +- module proposal stage names; +- proposal-generation stage names; +- validator batch stage names. + +Prompt metadata and response-schema metadata each expose canonical diagnostics maps via: +- `prompts.Metadata.DiagnosticsMap()` +- `responseschema.Schema.DiagnosticsMap()` + +## Diagnostics and reporting +Run-directory artifacts are owned by `internal/core/diagnostics`. + +Stable artifact names are centralized constants (for example transcript artifacts, `invocation.json`, `effective-config.json`, `utilization-diagnostics.json`, `correction-ledger.json`, `report.json`, `error.log`). + +Report diagnostics path metadata is constructed through `BuildDiagnosticsMetadata`, which keeps run-directory artifact references consistent between success and failure reports. + +## Secret redaction +Redaction responsibilities are split by concern: +- structural config redaction: `config.Config.Redacted()` +- byte/string payload redaction for diagnostics and surfaced errors: framework redaction utilities. + +Configured LLM secret extraction is centralized in `llm.ConfiguredSecrets(cfg)` and reused across proposal and validator diagnostics paths. + +## Output contracts +Transcript output schema selection is owned by `internal/core/outputschema`. + +Supported schemas: +- `bare-segments` +- `audita-v1` + +Unknown schema keys fail validation and runtime resolution. + +## Key package map +Core packages: +- `internal/core/config` +- `internal/core/schema` +- `internal/core/normalization` +- `internal/core/chunking` +- `internal/core/diagnostics` +- `internal/core/reporting` +- `internal/core/modulecatalog` +- `internal/core/outputschema` + +Framework packages: +- `internal/framework/contracts` +- `internal/framework/proposals` +- `internal/framework/proposal_generation` +- `internal/framework/promptcontext` +- `internal/framework/runner` +- `internal/framework/validators` +- `internal/framework/llm` +- `internal/framework/responseschema` +- `internal/framework/stagename` +- `internal/framework/structuredoutput` + +Domain packages: +- `internal/modules/*` +- `internal/validators/*` +- `internal/prompts` diff --git a/docs/architecture/public-contract.md b/docs/architecture/public-contract.md index ac253ad..3b8846f 100644 --- a/docs/architecture/public-contract.md +++ b/docs/architecture/public-contract.md @@ -1,31 +1,25 @@ # Audita Public Contract -This document defines stability expectations for Audita's external process and data interfaces. - ## Scope +This document defines stability expectations for Audita's external runtime interfaces. -This contract covers: -- CLI invocation and behavior -- versioned config file behavior -- transcript/glossary input forms -- transcript output schema selection -- process report schema metadata -- stable validator key identifiers in report/diagnostics records -- prompt metadata identifiers in diagnostics -- diagnostics directory behavior -- utilization diagnostics and correction-ledger artifact presence/pathing in diagnostics metadata -- stdout/stderr and exit-code behavior -- secret redaction guarantees -- compatibility and deprecation policy - -## CLI stability expectations +Covered interfaces: +- CLI commands and major flags; +- versioned config behavior and precedence; +- transcript/glossary input forms; +- output schema selection; +- report schema metadata; +- diagnostics artifact path metadata; +- stdout/stderr and exit-code behavior; +- redaction guarantees. +## CLI contract Stable commands: - `audita process` - `audita config validate` - `audita config print-effective` -For `audita process`, stable high-value flags include: +Stable high-value `process` flags: - `--config` - `--glossary` - `--output` @@ -33,136 +27,95 @@ For `audita process`, stable high-value flags include: - `--modules` - `--output-schema` -Compatibility flags and lower-level tuning flags remain available; they may be narrowed over time with explicit compatibility notes. +## Config contract +Supported config format: +- YAML; +- `version: 1`; +- strict unknown-field rejection. -## Config file stability expectations +Path resolution for `process` and `config print-effective`: +1. `--config` +2. `AUDITA_CONFIG` +3. `/usr/local/etc/audita/config.yml` +4. `/etc/audita/config.yml` -Supported file format: -- YAML -- strict unknown-field rejection -- explicit `version` +Missing explicit path is an error. Missing default paths is non-fatal. -Supported version: -- `version: 1` - -Precedence for `audita process`: -1. built-in defaults +Precedence for `process`: +1. defaults 2. file config 3. environment overrides 4. CLI overrides -Config source behavior: -- `--config `: missing path is a clear failure -- `AUDITA_CONFIG`: missing path is a clear failure -- defaults `/usr/local/etc/audita/config.yml`, then `/etc/audita/config.yml`: both missing is non-fatal +`config validate` remains file-only validation (defaults + file config; no env overrides). -## Supported transcript input forms +Module and output-schema keys are validated against built-in catalogs. Unknown keys fail validation. -Audita accepts transcript JSON as either: -- a top-level array of segments -- an object with a `segments` array +## Input contract +Supported transcript JSON top-level forms: +- array of segments +- object with `segments` array -Segments must satisfy the schema and validation rules enforced by `internal/core/schema`. +Supported glossary YAML form: +- top-level `glossary` list with required entry fields validated by schema parsing. -## Supported glossary input form - -Audita accepts glossary YAML with a top-level `glossary` entry list and validates required fields per entry. - -## Supported output schema names - -Built-in output schema registry supports: +## Output schema contract +Supported transcript output schemas: - `bare-segments` (default) - `audita-v1` -`seriatim-intermediate` is planned but not implemented. +Unknown schema keys fail before output write. -Unknown output schema names fail clearly. - -## Report schema/versioning expectations - -Process report payloads include `report_metadata` with: +## Report metadata contract +Process reports include stable report metadata fields: - `report_schema_name` - `report_schema_version` - `output_schema` -- `config_version` when file config is used +- `config_version` (when file config is loaded) Current values: -- `report_schema_name`: `audita-process-report` -- `report_schema_version`: `v1` +- `report_schema_name = audita-process-report` +- `report_schema_version = v1` -`--report-json` output and diagnostics run-dir `report.json` use the same report schema metadata. +`--report-json` output and run-directory `report.json` use the same report schema metadata. -Validator decision/rejection records in reports use stable validator keys in `validator_name`. -Module results may also include warning records for malformed module-stage LLM payloads. -Report diagnostics metadata includes artifact-path fields for utilization diagnostics and correction ledger when diagnostics initialization succeeds. +Validator decision/rejection records use stable validator keys via `validator_name`. -## Diagnostics directory behavior +## Diagnostics metadata contract +When run-directory initialization succeeds, diagnostics metadata paths reference stable artifacts, including: +- transcript and normalization artifacts; +- chunking summary; +- invocation metadata; +- redacted effective config; +- utilization diagnostics; +- correction ledger; +- `error.log` on failures. -When diagnostics directory creation succeeds, Audita writes run artifacts including: -- invocation metadata -- redacted effective config -- transcript/normalization/chunking artifacts -- utilization diagnostics (`utilization-diagnostics.json`) -- correction ledger (`correction-ledger.json`) -- report and failure error log (when applicable) -- module/LLM diagnostics artifacts as available +LLM interaction diagnostics include stable prompt and structured-schema identifiers where applicable. -Retention behavior is controlled by configured retention mode; failed runs are retained. +## Stdout/stderr and exit codes +Success: +- with `--output`, stdout is empty; +- without `--output`, stdout contains transcript JSON only; +- report JSON is not written to stdout. -Diagnostics metadata for LLM interactions may include semi-public prompt identifiers: -- `prompt_id` -- `prompt_version` -- `prompt_source` -- `embedded_path` -- `sha256` +Failures: +- nonzero exit; +- human-readable stderr summary; +- diagnostics directory path on stderr when available. -These are diagnostic identifiers, not user-facing prompt override controls. +Exit codes: +- `0` success +- nonzero failure -## Stdout/stderr behavior +## Redaction contract +Configured secrets are redacted from: +- effective config outputs; +- diagnostics artifacts; +- report artifacts; +- surfaced adapter/runtime errors. -Success behavior: -- with `--output`, stdout is empty -- without `--output`, stdout contains only transcript JSON in selected output schema -- report JSON is not written to stdout -- success stderr remains empty even when reports/diagnostics contain module warnings +## Compatibility policy +Stable command behavior, schema names, report metadata keys, diagnostics-path field semantics, and validator key identities are treated as public contract. -Failure behavior: -- stderr contains human-readable error summary -- nonzero exit -- diagnostics path is printed when available - -## Exit-code behavior - -- `0`: success -- nonzero: failure - -Treat any nonzero exit as a failed invocation. - -## Secret redaction guarantees - -Audita redacts API keys and authorization secrets from: -- effective config outputs (`audita config print-effective`, diagnostics effective-config artifact) -- report artifacts -- LLM diagnostics artifacts -- surfaced request/response error messages - -Config files should reference secrets via environment variable names (`api_key_env`) rather than embedding secret values. - -## Compatibility and deprecation policy - -- Existing stable schema names, report metadata keys, and top-level command behavior are treated as public contract. -- Existing stable validator keys remain public contract values even when validator semantics are refined. -- Compatibility inputs (legacy flags/env aliases) may remain during transition windows. -- Any planned removal or behavior change should include clear compatibility notes and migration guidance. - -## Breaking changes after 1.0 - -After 1.0, breaking changes include, for example: -- changing default success/failure exit-code semantics -- changing stdout/stderr routing semantics -- silently changing default output schema shape -- removing supported output schema names without compatibility strategy -- changing report schema fields or meanings incompatibly -- changing config version semantics incompatibly without version bump - -Additive fields, additive diagnostics, and new optional schema names are generally non-breaking when existing behavior remains intact. +Additive fields are acceptable when existing fields and behavior remain compatible. diff --git a/docs/architecture/structured-llm.md b/docs/architecture/structured-llm.md index 2d83d3d..3d31332 100644 --- a/docs/architecture/structured-llm.md +++ b/docs/architecture/structured-llm.md @@ -1,90 +1,72 @@ # Structured LLM Architecture -## Purpose - +## Scope This document describes Audita's structured LLM runtime boundary and adapter behavior. -## Why Audita owns the adapter - -Audita owns a small structured LLM adapter so that core runtime behavior is controlled inside the repository: -- request construction and schema handling are explicit and testable; -- retries, timeouts, cancellation, and error redaction are consistent across modules and validators; -- provider SDK types are not exposed outside the adapter boundary; -- dependency weight and transitive provider-specific behavior are reduced. - -At runtime, the rest of Audita depends only on the internal contract: -- `StructuredLLMClient` +## Runtime boundary +Production LLM integration depends on the internal contract only: +- `contracts.StructuredLLMClient` - `CompleteStructured(ctx, req, out)` -## OpenAI-compatible request shape +Provider SDK types do not leak past this boundary. -At a conceptual level, Audita sends chat completion requests with: -- `model` -- `messages` (role/content pairs) -- `response_format`: - - `type = "json_schema"` - - `json_schema.name` (stable schema name) - - `json_schema.strict = true` - - `json_schema.schema` (registered JSON Schema payload) +## Adapter ownership +`internal/framework/llm` owns the OpenAI-compatible HTTP adapter and shared LLM runtime utilities. -The adapter uses OpenAI-compatible `POST {base_url}/chat/completions` over `net/http`. +Key responsibilities: +- request assembly; +- timeout/cancellation propagation; +- bounded retry behavior; +- scheduler integration; +- provider response decoding; +- error redaction. -## Structured response schema registry +## Structured schema registry +Structured response schemas are registered in `internal/framework/responseschema` and include stable metadata: +- `id` +- `version` +- `name` +- `json_schema` +- `sha256` -Structured response schemas are registered in `internal/framework/responseschema` with stable metadata: -- schema key -- schema ID -- schema version -- schema name (OpenAI-compatible `response_format` name) -- raw JSON Schema payload -- SHA-256 hash +Current schema keys: +- `correction_set` +- `validator_decision_set` -Current schemas: -- `correction_set`: - - id `audita.correction_set` - - version `v1` - - name `audita_correction_set_v1` -- `validator_decision_set`: - - id `audita.validator_decision_set` - - version `v1` - - name `audita_validator_decision_set_v1` +Schema metadata is attached to diagnostics through `Schema.DiagnosticsMap()`. -## Provider compatibility assumptions +## Request shape assumptions +Audita targets OpenAI-compatible chat-completions endpoints and sends structured requests with: +- model; +- chat messages; +- `response_format.type = json_schema`; +- schema name and JSON schema payload. -Audita assumes an OpenAI-compatible chat-completions endpoint that: -- accepts message arrays with model selection; -- accepts `response_format.type = json_schema`; -- returns a completion with assistant message content and optional usage metadata. +## Local validation remains mandatory +Provider schema enforcement is treated as transport-level guardrails. -Provider-specific differences are expected in strictness and error payload shapes, so the adapter treats provider output as untrusted until locally decoded. +Audita still validates output locally before applying behavior changes: +- proposal decoding and proposal invariants; +- validator decision decoding and cardinality checks; +- deterministic validation and apply-time rules. -## Local decode and validation remain mandatory +## Shared malformed-output policy +Malformed structured-output classification is centralized in `internal/framework/structuredoutput`. -Provider-level structured output is a transport guardrail, not final validation. +Proposal generation and validator execution both use this shared classifier so downgrade behavior cannot drift between the two paths. -After receiving a response, Audita still: -- decodes assistant content into typed request-specific structs; -- validates proposal and validator payload invariants locally; -- enforces deterministic validator/cardinality rules before any transcript application. +## Secrets and redaction +Secret extraction for LLM redaction is centralized in `llm.ConfiguredSecrets(cfg)` and reused by proposal and validator diagnostics writers. -This protects runtime correctness even when provider responses are malformed, partial, or semantically inconsistent. +Secrets are redacted from: +- diagnostics artifacts; +- report artifacts; +- surfaced adapter/runtime errors. -## Diagnostics and redaction +## Concurrency and scheduling +LLM execution is constrained by composed scheduler limits: +- total LLM concurrency; +- proposal LLM concurrency; +- validation LLM concurrency. -When structured schemas are used, diagnostics metadata records: -- schema ID -- schema version -- schema name -- schema hash - -Diagnostics and surfaced errors preserve secret redaction: -- API keys and bearer tokens are redacted from request/response/error artifacts; -- redaction is applied before diagnostic files are written. - -## Runtime behavior guarantees - -The structured LLM path preserves existing runtime guarantees: -- bounded LLM call execution through schedulers; -- context-aware cancellation and timeout propagation; -- retry behavior for transient failures and retryable malformed structured responses; -- deterministic module/chunk/proposal/validator behavior outside provider nondeterminism. +The scheduler is FIFO and context-aware so permits are released on success, failure, and cancellation. diff --git a/docs/architecture/validators.md b/docs/architecture/validators.md index e801416..08c84b2 100644 --- a/docs/architecture/validators.md +++ b/docs/architecture/validators.md @@ -1,160 +1,96 @@ # Audita Validators -This document describes Audita's built-in validator registry and module validator chains. - -For LLM-backed validator prompt asset details, see [`docs/prompts.md`](prompts.md). - -## Package ownership - -Built-in validator construction is package-owned under `internal/validators/`: -- `internal/validators/confidence_threshold` -- `internal/validators/proposal_shape` -- `internal/validators/original_text_presence` -- `internal/validators/non_empty_corrected_text` -- `internal/validators/no_effect` -- `internal/validators/protected_terms` -- `internal/validators/spoken_form_plausibility` -- `internal/validators/meaning_reversal_review` -- `internal/validators/editorial_review` - -Registry and chain wiring stay in: -- `internal/validators/registry.go` -- `internal/validators/chains.go` - -Shared validator runtime mechanics stay in `internal/framework/validators`: -- request/result/decision models -- decision cardinality helpers -- protected vocabulary helpers -- shared LLM validator runtime, batching, and diagnostics helpers - -Execution classification metadata is defined in `internal/validators/metadata`: -- `deterministic` -- `llm_backed` - -Runner ordering uses this metadata so deterministic validators run before LLM-backed validators without concrete framework type assertions. - ## Scope +This document defines the built-in validator system used by production module runs. -Validator chains are built-in runtime behavior. +## Ownership boundaries +Built-in validator keys, constructors, and module chains are owned by `internal/validators`. -Current 1.0 boundary: -- built-in validator keys and built-in module chains are stable runtime identifiers; -- thresholds and batching knobs remain configurable where already supported; -- arbitrary user-defined validator chains are deferred. +Shared runtime execution mechanics are owned by `internal/framework/validators`, including: +- validator request/result models; +- deterministic proposal checks; +- LLM validator batching and execution; +- decision-cardinality enforcement; +- diagnostics integration. -## Built-in validator keys - -### Deterministic validators +Execution class metadata is owned by `internal/validators/metadata`. +## Stable validator keys +Deterministic: - `proposal_shape` - - rejects malformed proposal fields before other validators run. - `confidence_threshold` - - checks proposal confidence against module-specific configured threshold. - `original_text_presence` - - ensures target segment exists and `original_text` exists in current working segment text. - `non_empty_corrected_text` - - rejects proposals whose previewed resulting segment text would be empty or whitespace-only. - `no_effect` - - rejects proposals where `original_text == corrected_text`. - `protected_terms` - - protects glossary-derived terms from unsafe mutations in non-glossary modules. - - glossary stages use glossary-specific protection logic but still report this same stable key. - -### LLM-backed validators +LLM-backed: - `spoken_form_plausibility` - - checks whether proposed spoken-form change remains plausible in transcript context. - `meaning_reversal_review` - - checks for likely meaning reversal or semantic contradiction. - `editorial_review` - - performs conservative editorial safety review. ## Built-in module chains +`glossary`: +- `proposal_shape` +- `no_effect` +- `original_text_presence` +- `confidence_threshold` +- `protected_terms` +- `non_empty_corrected_text` +- `spoken_form_plausibility` +- `meaning_reversal_review` -Current built-in chains resolved from `internal/validators/chains.go`: +`homophones`: +- `proposal_shape` +- `no_effect` +- `original_text_presence` +- `confidence_threshold` +- `protected_terms` +- `non_empty_corrected_text` +- `spoken_form_plausibility` +- `meaning_reversal_review` -- `glossary` - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `spoken_form_plausibility` - - `meaning_reversal_review` +`spoken_word`: +- `proposal_shape` +- `no_effect` +- `original_text_presence` +- `confidence_threshold` +- `protected_terms` +- `non_empty_corrected_text` +- `editorial_review` +- `meaning_reversal_review` -- `homophones` - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `spoken_form_plausibility` - - `meaning_reversal_review` +`grammar`: +- `proposal_shape` +- `no_effect` +- `original_text_presence` +- `confidence_threshold` +- `protected_terms` +- `non_empty_corrected_text` +- `editorial_review` +- `meaning_reversal_review` -- `spoken_word` - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `editorial_review` - - `meaning_reversal_review` +## Ordering and execution semantics +Validator ordering is based on canonical metadata: +- deterministic validators run before LLM-backed validators. -- `grammar` - - `proposal_shape` - - `no_effect` - - `original_text_presence` - - `confidence_threshold` - - `protected_terms` - - `non_empty_corrected_text` - - `editorial_review` - - `meaning_reversal_review` +Within each module stage: +- proposals are generated per section; +- validator chains execute on those proposals; +- approved proposals are applied once after section work settles. -## Protected terms construction +## Malformed payload behavior +Malformed structured-output from proposal generation and LLM validator calls is downgraded, not treated as a process-fatal transport error. -`protected_terms` has explicit constructors: -- general constructor used by non-glossary modules through the built-in registry -- glossary-stage constructor used by glossary chain resolution +Current outcomes: +- malformed proposal-generation payloads produce section/module warnings and zero proposals for the affected section; +- malformed validator decision payloads reject the affected validator batch with warnings; +- deterministic validator behavior and runner order remain unchanged. -Both variants preserve existing behavior and report the stable key `protected_terms`. +## Reporting identity +Reports and diagnostics use stable validator keys as identifiers. -## Execution semantics +Correction-ledger deterministic-vs-LLM classification is derived from canonical validator metadata, not package-local hardcoded maps. -- modules execute serially; -- section proposal work can run concurrently within a module; -- deterministic validators run before LLM-backed validators; -- malformed module proposal payloads are downgraded to section-scoped module warnings with zero proposals for the affected section rather than module failure; -- malformed/missing/duplicate/unknown LLM validator decisions reject the affected validator batch with warnings instead of failing the module; -- oversized single-proposal validator inputs reject only the affected proposal under that validator; -- approved proposals are applied once per module after section work settles. - -## Validator rejections vs proposal-application skips - -- validator rejection: - - proposal is denied by validator-chain review and appears in validator rejection reporting with validator key and reason code. -- proposal-application skip: - - proposal passed validators but could not be applied under replacement-policy semantics (for example no matching span at apply time). -- module warning: - - malformed proposal-generation payloads and malformed validator batches are recorded in module warning records and diagnostics without writing success stderr. - -These are separate outcomes and are reported separately. - -## Reporting and diagnostics identity - -- report validator decision/rejection entries use stable validator keys in `validator_name`. -- report module results include warning records for malformed module-stage LLM payloads. -- validator LLM diagnostics include validator identity in interaction metadata and structured response schema metadata. -- correction ledger entries include deterministic and LLM validator decision snapshots keyed by the same stable validator keys, and keep validator rejection distinct from application-level skip. - -Prompt assets are unchanged by the validator package-ownership refactor and remain built-in under `internal/prompts`. - -## Configurable knobs that remain supported - -- per-module confidence thresholds (`thresholds.*` / equivalent env+CLI overrides) -- validation batching limits (`validation_max_prompt_tokens` / equivalent env+CLI overrides) -- validation LLM model/base URL/timeout/retries/concurrency settings - -These tune validator behavior without exposing arbitrary user-defined chains. +## Prompt assets +LLM validator prompt assets and prompt metadata are documented in [Prompts](./prompts.md). diff --git a/docs/configuration.md b/docs/configuration.md index 416910d..f70d20d 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -1,51 +1,48 @@ # Audita Configuration -This document describes Audita's versioned YAML config support and related commands. - -## Purpose - -Audita's config file provides a stable place for pipeline defaults and runtime tuning that would otherwise require many environment variables or CLI flags. - -Use config files for baseline settings, then use environment variables and CLI flags for deployment and per-run overrides. - -## Supported version - -Current supported config version: +## Scope +This document defines the supported versioned YAML configuration model and runtime precedence behavior. +## Supported file version +Current supported config file version: - `version: 1` -Rules: - -- missing `version` fails validation; -- unknown versions fail validation; -- unknown fields fail validation (strict decoding). +Validation rules: +- missing `version` fails; +- unsupported version fails; +- unknown YAML fields fail (strict decoding). ## Config path resolution +For `audita process` and `audita config print-effective`, path resolution order is: +1. `--config ` +2. `AUDITA_CONFIG` +3. `/usr/local/etc/audita/config.yml` (if present) +4. `/etc/audita/config.yml` (if present) -For `audita process`, config path resolution is: +Missing-path behavior: +- missing `--config` path is an error; +- missing `AUDITA_CONFIG` path is an error; +- missing both default paths is non-fatal. -1. `--config ` if provided -2. `AUDITA_CONFIG` if set and `--config` is not provided -3. default `/usr/local/etc/audita/config.yml` if present -4. fallback default `/etc/audita/config.yml` if present - -Missing-file behavior: - -- missing `--config` path: hard failure; -- missing `AUDITA_CONFIG` path: hard failure; -- missing both default-path files: non-fatal, run continues. - -## Precedence model - -Effective config precedence is: - -1. built-in defaults +## Effective precedence +`audita process` effective precedence: +1. defaults 2. file config 3. environment overrides 4. CLI overrides -## Supported YAML fields +`audita config print-effective` uses: +1. defaults +2. file config +3. environment overrides +`audita config validate` intentionally uses file-only validation: +1. defaults +2. file config + +Environment overrides are not applied in `config validate`. + +## Supported top-level YAML fields ```yaml version: 1 @@ -62,7 +59,6 @@ llm: api_key_env: AUDITA_LLM_API_KEY timeout: 120s max_retries: 3 - validation: base_url: https://openrouter.ai/api/v1 model: openrouter/google/gemma-4-31b-it @@ -100,91 +96,54 @@ diagnostics: retention: auto ``` -`context.description` provides background-only transcript context for prompts. -If both config and CLI provide a description, `--transcript-description` takes precedence. +## Module and output-schema validation +`pipeline.modules` keys are validated against the built-in supported module catalog. -`output.schema` supports the built-in output schema registry values: -- `bare-segments` (default) +Supported module keys: +- `glossary` +- `homophones` +- `spoken_word` +- `grammar` + +Repeated supported module keys are allowed. + +`output.schema` is validated against the built-in output schema catalog. + +Supported output schema keys: +- `bare-segments` - `audita-v1` -Unknown schema names fail clearly before transcript output is written. +Unknown module keys and unknown output schema keys fail validation. -Duration-like fields accept either: +## Duration field parsing +Duration-like fields support: +- numeric seconds (for example `120`, `3.5`) +- duration strings (for example `120s`, `2m`) -- numeric seconds (for example `120`, `3.5`), or -- duration strings (for example `120s`, `2m`). - -For LLM timeouts, duration strings must resolve to whole seconds. +LLM timeout duration strings must resolve to whole seconds. ## Secret handling - -Use `api_key_env` for secrets: - +Use `api_key_env` fields for secrets: - `llm.proposal.api_key_env` - `llm.validation.api_key_env` -These fields must contain environment variable names, not secret values. +These fields store environment variable names, not secret values. -At runtime, Audita resolves those names from the process environment. - -Redaction behavior: - -- run diagnostics `effective-config.json` is redacted; -- `audita config print-effective` output is redacted; -- API keys are never emitted in plaintext by those outputs. - -## Config commands - -Validate a config file: +Resolved secret values are redacted from: +- `audita config print-effective` output; +- diagnostics `effective-config.json`; +- report and diagnostics payloads. +## Commands +Validate a file config: ```sh audita config validate --config ./audita.yml ``` Print redacted effective config: - ```sh audita config print-effective --config ./audita.yml ``` -`print-effective` loads defaults, then file config, then environment overrides. - -## Example: local OpenAI-compatible endpoint - -```yaml -version: 1 - -llm: - proposal: - base_url: http://localhost:8000/v1 - model: local/proposal-model - api_key_env: AUDITA_LLM_API_KEY - timeout: 90s - max_retries: 2 - - validation: - base_url: http://localhost:8000/v1 - model: local/validation-model - api_key_env: AUDITA_VALIDATION_LLM_API_KEY - timeout: 90s - max_retries: 2 - -pipeline: - modules: [glossary, homophones, glossary, spoken_word, grammar] - -diagnostics: - work_dir: /tmp/audita - retention: auto -``` - ## Compatibility notes - -Existing environment variables and lower-level CLI flags remain available for compatibility. - -Current guidance: - -- prefer file config for baseline behavior; -- keep environment variables for secrets/deployment-specific overrides; -- use CLI flags for per-run overrides. -- validator chains are built-in and are not user-configurable in config. -- prompt source selection and filesystem prompt overrides are not config options. +Legacy compatibility flags and environment aliases remain available where implemented, but the stable configuration surface is the versioned YAML model described above. diff --git a/docs/development.md b/docs/development.md new file mode 100644 index 0000000..614b211 --- /dev/null +++ b/docs/development.md @@ -0,0 +1,33 @@ +# Audita Development Workflow + +## Scope +This document defines the canonical contributor workflow and engineering conventions for this repository. + +## Workflow +1. Start from a clean understanding of scope and constraints. +2. Make focused changes that preserve existing public behavior unless behavior change is explicitly intended. +3. Run targeted tests for touched packages. +4. Run `go test ./...` before finalizing substantial changes. +5. Update affected documentation so it describes current behavior only. + +## Engineering conventions +- Keep module packages separate: `glossary`, `homophones`, `spoken_word`, `grammar`. +- Prefer narrow shared helpers and catalogs over broad abstractions. +- Preserve diagnostics artifact naming and report field contracts unless intentionally changed. +- Preserve CLI/config precedence semantics unless intentionally changed. +- Treat stable validator keys, prompt identifiers, and output-schema keys as contract surfaces. + +## Configuration and runtime expectations +- `audita process` precedence is defaults -> file -> env -> CLI. +- `audita config validate` validates file config merged onto defaults only. +- `audita config print-effective` includes environment overrides and prints redacted JSON. + +## Testing expectations +- Add tests for new behavior and for bug fixes. +- Keep deterministic fixtures stable. +- Do not reduce existing parity, release-fixture, subprocess, or module-specific coverage without equivalent replacement. + +## Commit discipline +- Keep commits scoped and reviewable. +- Avoid mixing unrelated refactors with behavior changes. +- Use clear plain-English commit messages. diff --git a/docs/documentation/policy.md b/docs/documentation/policy.md new file mode 100644 index 0000000..bfefc16 --- /dev/null +++ b/docs/documentation/policy.md @@ -0,0 +1,27 @@ +# Documentation Policy + +## Scope +This policy defines how project documentation should be authored and maintained. + +## Core rules +- Document the current behavior of the codebase. +- Remove stale behavior descriptions promptly when code changes. +- Do not describe development history in architecture or behavior docs unless a document is explicitly historical. +- Do not use architecture or behavior docs as changelogs. +- Prefer rewriting stale sections from scratch when substantial behavior or ownership changes occur. + +## Consistency requirements +- Keep command examples aligned with current CLI surfaces. +- Keep configuration examples aligned with supported fields and precedence. +- Keep architecture package ownership descriptions aligned with current code layout. +- Keep stable contract identifiers accurate (module keys, validator keys, output-schema keys, report metadata fields). + +## Cross-document expectations +- `docs/architecture/*` documents runtime behavior and package ownership. +- `docs/configuration.md` documents config schema and precedence. +- `docs/development.md` documents contributor workflow and engineering conventions. + +## Review expectations for documentation changes +- Verify referenced files and links exist. +- Verify examples match current behavior. +- Prefer concise, direct language and avoid speculative future claims.