418 lines
19 KiB
Markdown
418 lines
19 KiB
Markdown
# Audita Go Architecture
|
|
|
|
## Scope and intent
|
|
This document describes:
|
|
- the current implemented Go architecture; and
|
|
- the intended final architecture for later rewrite phases.
|
|
|
|
Status labels are explicit so future engineers and LLM agents do not assume unimplemented behavior exists.
|
|
|
|
## Current implementation status
|
|
Implemented today:
|
|
- Go CLI entrypoint and `audita process` wiring.
|
|
- Config defaults, env loading, CLI override precedence, and validation.
|
|
- Transcript and glossary parsing/validation.
|
|
- Deterministic transcript normalization.
|
|
- Deterministic token estimation and transcript chunking.
|
|
- Per-run diagnostics directory creation plus Phase 6 process-level artifacts.
|
|
- Process report JSON output with diagnostics artifact references.
|
|
- Framework foundation packages for contracts and proposal application.
|
|
- Production runner orchestration package with deterministic sequential module execution.
|
|
- Module-level report structures with applied/skipped change records.
|
|
- Runtime validator models and deterministic validators.
|
|
- Deterministic validator-chain execution in the runner with cardinality enforcement.
|
|
- Module-level validator decision/rejection reporting.
|
|
- Internal structured LLM client contract plus an `instructor-go`-backed adapter package.
|
|
- Bounded LLM scheduler/semaphore infrastructure with context-aware permit handling.
|
|
- Runtime primary/validation LLM effective-config resolution helpers with validation inheritance.
|
|
- Generic JSON prompt/response diagnostics writer primitives with secret redaction.
|
|
- LLM-backed validator models, prompt builders, batching, and runtime execution.
|
|
- Runner wiring for LLM validators via the internal structured LLM abstraction and scheduler hooks.
|
|
- LLM validator diagnostics artifacts and report-level decision metadata paths.
|
|
- Shared LLM proposal-generation helper with structured correction-set parsing.
|
|
- Deterministic proposal-index assignment and enriched proposal mapping for shared generation.
|
|
- Proposal-generation diagnostics artifacts with secret redaction.
|
|
- Production module registry scaffolding with known-key recognition and explicit unsupported/unimplemented errors.
|
|
- Production `grammar` module implementation in `internal/modules/grammar`.
|
|
- Explicit runtime support for `--modules grammar` through the production runner path.
|
|
|
|
Not implemented in CLI runtime path today:
|
|
- Real module execution pipeline for `glossary`, `homophones`, and `spoken_word`.
|
|
- Real domain proposal prompts for production modules.
|
|
- End-to-end transcript polishing with real module behavior.
|
|
|
|
Phase sequencing note:
|
|
- Phase 9 LLM infrastructure is complete (structured client, scheduler, effective config resolution, diagnostics primitives);
|
|
- Phase 10 LLM-backed validator runtime integration is complete;
|
|
- Phase 11 shared proposal-generation framework and module-registry scaffolding are complete;
|
|
- Phase 12 grammar module implementation and explicit runtime wiring are complete;
|
|
- next recommended phase is Phase 13 (glossary module and protected-term behavior).
|
|
|
|
## Actual Go package layout
|
|
|
|
```text
|
|
cmd/audita/
|
|
main.go
|
|
|
|
internal/cli/
|
|
run.go
|
|
|
|
internal/core/config/
|
|
config.go
|
|
env.go
|
|
flags.go
|
|
redaction.go
|
|
validation.go
|
|
|
|
internal/core/schema/
|
|
transcript.go
|
|
glossary.go
|
|
errors.go
|
|
|
|
internal/core/io/
|
|
files.go
|
|
|
|
internal/core/normalization/
|
|
normalize.go
|
|
tokens.go
|
|
|
|
internal/core/chunking/
|
|
sections.go
|
|
summary.go
|
|
tokens.go
|
|
|
|
internal/core/diagnostics/
|
|
run_dir.go
|
|
|
|
internal/core/reporting/
|
|
report.go
|
|
|
|
internal/framework/contracts/
|
|
contracts.go
|
|
|
|
internal/framework/proposals/
|
|
proposal.go
|
|
policy.go
|
|
preview.go
|
|
apply.go
|
|
|
|
internal/framework/runner/
|
|
runner.go
|
|
|
|
internal/framework/proposal_generation/
|
|
generate.go
|
|
|
|
internal/framework/modules/
|
|
registry.go
|
|
|
|
internal/modules/grammar/
|
|
module.go
|
|
prompt.go
|
|
|
|
internal/framework/validators/
|
|
models.go
|
|
deterministic.go
|
|
llm_models.go
|
|
llm_prompt_builders.go
|
|
llm_batching.go
|
|
llm_validators.go
|
|
|
|
internal/framework/llm/
|
|
instructor_client.go
|
|
scheduler.go
|
|
effective_config.go
|
|
diagnostics.go
|
|
```
|
|
|
|
## Current CLI behavior
|
|
Primary command:
|
|
|
|
```sh
|
|
audita process <transcript.json> --glossary <glossary.yaml> [flags]
|
|
```
|
|
|
|
Current runtime flow (`internal/cli/run.go`):
|
|
1. Load config from env.
|
|
2. Parse flags and apply CLI overrides.
|
|
3. Validate transcript positional argument and required `--glossary`.
|
|
4. Create per-run diagnostics directory.
|
|
5. Read transcript and glossary files.
|
|
6. Parse/validate transcript and glossary.
|
|
7. Write source transcript artifacts.
|
|
8. Normalize transcript.
|
|
9. Write normalized transcript and normalization summary artifacts.
|
|
10. Chunk normalized transcript and compute chunk summaries.
|
|
11. Write chunking summary artifact.
|
|
12. Execute runner modules sequentially when:
|
|
- `--modules` is explicitly provided (production grammar path); or
|
|
- a test/injected module factory is provided.
|
|
13. Output working transcript to `--output` file or stdout.
|
|
14. Build process report (`phase` currently set to `phase12-grammar-module`).
|
|
15. Optionally write `--report-json`; always write run-dir `report.json`.
|
|
16. Apply work-dir retention.
|
|
|
|
Important behavior details:
|
|
- Glossary is validated but not yet used for real correction module logic.
|
|
- Default production CLI behavior remains deterministic normalization/chunking/reporting unless modules are explicitly selected with `--modules`.
|
|
- Explicit `--modules grammar` runs the production grammar module path with LLM-backed proposal generation and validator-chain execution.
|
|
- Default runs (without explicit module selection) do not perform LLM calls.
|
|
- Success path is generally quiet on stderr.
|
|
- Source IDs are preserved into a canonical transcript before normalization; normalization then reassigns output IDs sequentially from `1`.
|
|
|
|
## Implemented data contracts
|
|
|
|
### Transcript input
|
|
Accepted top-level forms:
|
|
- bare JSON array of segments
|
|
- object with `segments` array
|
|
|
|
Source segment contract:
|
|
- `id` optional integer
|
|
- `speaker` non-empty string
|
|
- `start` finite non-negative number
|
|
- `end` finite non-negative number with `end >= start`
|
|
- `text` non-empty string
|
|
- `categories` optional array of non-empty strings
|
|
|
|
Additional checks:
|
|
- duplicate explicit source IDs are rejected.
|
|
|
|
### Transcript output
|
|
Current output uses `schema.TranscriptToJSON` and is a bare JSON array of normalized segments:
|
|
- `id`, `speaker`, `start`, `end`, `text`, optional `categories`.
|
|
|
|
### Glossary input
|
|
YAML with `glossary` entries. Required fields per entry:
|
|
- `name`, `category`, `summary`
|
|
|
|
Optional:
|
|
- `aliases`, `plural`
|
|
|
|
## Implemented config/env/flag behavior
|
|
Precedence:
|
|
1. defaults (`config.Default()`)
|
|
2. environment (`config.LoadFromEnv()`)
|
|
3. CLI flags (`ApplyCLIOverrides`)
|
|
|
|
Implemented config surfaces include:
|
|
- module list
|
|
- primary and validation LLM settings
|
|
- section token controls and target sections
|
|
- confidence thresholds
|
|
- normalization controls
|
|
- work-dir and retention mode
|
|
|
|
Current caveat:
|
|
- LLM/module-related settings are active for explicit grammar runs; the default non-explicit path remains deterministic.
|
|
|
|
## Implemented structured LLM infrastructure
|
|
`internal/framework/contracts` now defines a typed structured-completion contract:
|
|
- `StructuredLLMClient.CompleteStructured(ctx, req, out)`
|
|
- caller-owned typed decode target via `out` pointer.
|
|
|
|
`internal/framework/llm` provides `InstructorClient`, an internal adapter over `github.com/jxnl/instructor-go`:
|
|
- configurable `base_url`, model, optional API key, retries, mode, HTTP client, and request timeout;
|
|
- OpenAI-compatible endpoint behavior (for example OpenAI/OpenRouter/local-compatible base URLs);
|
|
- default mode is JSON mode (`ModeJSON`), with optional tool-call mode (`ModeToolCall`);
|
|
- request message translation from `contracts.LLMMessage` to chat-completions messages;
|
|
- response metadata mapping (provider/model/token usage) into Audita-owned response types;
|
|
- API-key redaction in adapter-returned errors.
|
|
|
|
Current runtime boundary:
|
|
- the default CLI runtime path (without explicit module selection) still does not instantiate the full production module sequence.
|
|
- LLM calls are exercised in production when `--modules grammar` is explicitly requested and in tests when fake/injected clients are used.
|
|
|
|
`internal/framework/llm` also provides:
|
|
- a bounded `Scheduler` for controlled concurrent LLM calls with reliable permit release;
|
|
- primary/validation effective-config resolution helpers, including validation inheritance fallback to primary settings;
|
|
- generic interaction diagnostics primitives that write machine-readable JSON artifacts for request metadata, request payload, response payload, and optional error payload with secret redaction.
|
|
|
|
## Implemented normalization behavior
|
|
Normalization (`internal/core/normalization`) currently:
|
|
- sorts by segment start time;
|
|
- merges adjacent same-speaker segments when constraints pass;
|
|
- uses gap-based joiners:
|
|
- gap `< ellipsis_gap` -> single space join
|
|
- gap `>= ellipsis_gap` -> `... ` join
|
|
- enforces merged duration and token-limit constraints;
|
|
- reassigns output IDs sequentially from `1`;
|
|
- returns `NormalizationSummary` with merge and skip counters.
|
|
|
|
Note: merged categories are concatenated (not deduplicated).
|
|
|
|
## Implemented chunking behavior
|
|
Chunking (`internal/core/chunking`) currently provides:
|
|
- deterministic heuristic token estimation;
|
|
- contiguous sectioning with section metadata;
|
|
- max/min section token validation;
|
|
- optional `target_sections` handling with target-aware merge/split logic;
|
|
- summary and detailed summary generation.
|
|
|
|
Current behavior details:
|
|
- if a single segment exceeds max tokens, it is emitted as its own section (not hard-failed);
|
|
- section balancing is deterministic but heuristic.
|
|
|
|
## Implemented proposal/replacement infrastructure
|
|
`internal/framework/proposals` provides deterministic foundation logic:
|
|
- `CorrectionProposal` and `EnrichedCorrectionProposal` models;
|
|
- replacement policies: `require_unique`, `replace_all`;
|
|
- safe preview (`PreviewProposalForSegment`) with stable skip reasons;
|
|
- deterministic apply (`ApplyProposals`) in ascending `proposal_index` order;
|
|
- applied/skipped change records suitable for reporting.
|
|
|
|
`internal/framework/contracts` provides interfaces and run-spec metadata scaffolding, including deterministic repeated module instance naming (`ResolveModuleRunSpecs`).
|
|
|
|
These primitives are wired into the production runner and report model. The grammar module is implemented; other production modules remain pending.
|
|
|
|
## Implemented validator runtime infrastructure
|
|
`internal/framework/validators` provides deterministic validator infrastructure:
|
|
- runtime validation request/result models;
|
|
- stable validator reason codes;
|
|
- cardinality enforcement for validator decisions:
|
|
- missing proposal indexes fail
|
|
- duplicate proposal indexes fail
|
|
- unknown proposal indexes fail
|
|
- deterministic validators:
|
|
- confidence threshold by module key/config threshold
|
|
- original-text presence against current working transcript
|
|
- non-empty corrected text
|
|
- identical/no-effect rejection
|
|
- conservative protected glossary-term guard for non-glossary modules
|
|
|
|
`internal/framework/runner` executes validator chains in order for each module and applies only validator-approved proposals.
|
|
Validator rejections are reported distinctly from proposal-application skips.
|
|
|
|
## Implemented LLM-backed validator infrastructure
|
|
`internal/framework/validators` now includes LLM-backed validator support:
|
|
- typed request/response models for structured LLM validation;
|
|
- prompt builders for:
|
|
- spoken-form plausibility
|
|
- meaning reversal detection
|
|
- editorial review
|
|
- grammar review
|
|
- spoken-word review
|
|
- deterministic batching by `validation_max_prompt_tokens`;
|
|
- strict cardinality validation of structured LLM decisions (missing/duplicate/unknown indexes fail);
|
|
- safe failure behavior for malformed/invalid structured responses.
|
|
|
|
`internal/framework/runner` wires LLM validators into existing validator chains using:
|
|
- the internal structured LLM client abstraction (`contracts.StructuredLLMClient`);
|
|
- bounded scheduler hooks for validator call execution;
|
|
- diagnostics writer hooks for machine-readable prompt/response artifacts with secret redaction.
|
|
|
|
## Implemented shared proposal-generation infrastructure
|
|
`internal/framework/proposal_generation` provides a reusable, prompt-agnostic helper for future real modules:
|
|
- structured request model including module key/instance, replacement policy, working transcript context, optional section metadata, glossary, config, and diagnostics context;
|
|
- structured correction-set response model (`corrections`) mapped into existing `proposals.CorrectionProposal` and `proposals.EnrichedCorrectionProposal` models;
|
|
- deterministic proposal-index assignment through a caller-provided `start_index`;
|
|
- structured LLM calls through `contracts.StructuredLLMClient` only (no direct provider calls);
|
|
- optional bounded execution through scheduler hooks (`contracts.LLMScheduler`);
|
|
- prompt/response diagnostics artifact writing via the generic `internal/framework/llm` diagnostics primitives with redaction of API keys/secrets.
|
|
|
|
This helper only produces candidate proposals; validator-chain execution and proposal application remain runner responsibilities.
|
|
|
|
## Implemented production module-registry scaffolding
|
|
`internal/framework/modules` now provides a production registry scaffold:
|
|
- recognizes intended module keys:
|
|
- `glossary`
|
|
- `homophones`
|
|
- `spoken_word`
|
|
- `grammar`
|
|
- supports explicit constructor registration with dependency injection for:
|
|
- run spec
|
|
- config
|
|
- glossary
|
|
- proposal/validation structured LLM clients
|
|
- proposal/validation schedulers
|
|
- diagnostics directory context
|
|
- returns explicit errors for unknown keys (`unsupported_module`) and recognized-but-unimplemented keys (`unimplemented_module`).
|
|
|
|
The `grammar` module key is now registered and constructible. `glossary`, `homophones`, and `spoken_word` remain recognized-but-unimplemented.
|
|
|
|
## Implemented grammar production module
|
|
`internal/modules/grammar` now provides the first production module:
|
|
- prompt builder faithfully constrained to punctuation/capitalization/spacing/article cleanup;
|
|
- explicit guardrails against meaning-changing rewrites, style rewrites, summarization, and invention;
|
|
- proposal generation through `internal/framework/proposal_generation` and `contracts.StructuredLLMClient`;
|
|
- scheduler-aware proposal calls through existing `contracts.LLMScheduler` hooks;
|
|
- replacement policy `require_unique` (matching Python implementation);
|
|
- validator chain integration using existing deterministic + LLM-backed validators;
|
|
- grammar confidence threshold enforcement through existing validator/config infrastructure;
|
|
- module-level reporting and diagnostics capture through existing runner/reporting paths.
|
|
|
|
## Reports and diagnostics (implemented)
|
|
Current per-run artifacts include:
|
|
- `source-transcript.json`
|
|
- `source-transcript-parsed.json`
|
|
- `normalized-transcript.json`
|
|
- `normalization-summary.json`
|
|
- `chunking-summary.json`
|
|
- `invocation.json`
|
|
- `effective-config.json` (redacted credentials)
|
|
- `report.json`
|
|
- `error.log` on failure
|
|
|
|
`--report-json` writes a separate report file when requested.
|
|
|
|
Current process reports include diagnostics metadata references for:
|
|
- diagnostics directory path;
|
|
- source transcript artifact path;
|
|
- parsed source transcript artifact path;
|
|
- normalized transcript artifact path;
|
|
- normalization summary artifact path;
|
|
- chunking summary artifact path;
|
|
- invocation metadata artifact path;
|
|
- redacted effective-config artifact path;
|
|
- error-log artifact path on failure.
|
|
|
|
Current process reports also include:
|
|
- module-level results (when runner modules execute), including applied/skipped proposal changes;
|
|
- run-level module summary totals and failed module instance metadata.
|
|
- module-level validator decisions and validator rejections.
|
|
- optional decision-level diagnostic artifact paths for validator LLM interactions when available.
|
|
|
|
Retention modes implemented in `ApplyRetention`:
|
|
- `always`: keep all run directories.
|
|
- `never`: keep successful run directories.
|
|
- `auto`: keep failed runs and successful runs with skipped corrections.
|
|
- failed runs are always retained.
|
|
|
|
Current runtime note:
|
|
- default non-explicit runs usually have no module-level skipped corrections, so `auto` commonly removes clean successful run directories.
|
|
- explicit grammar runs can produce validator rejections and application skips, which are reflected in reports and retention input.
|
|
|
|
Intentionally deferred to module/LLM phases:
|
|
- real domain proposal prompts and production module implementations remain tied to later module phases.
|
|
|
|
## Current tests and quality posture
|
|
Implemented tests currently cover:
|
|
- CLI argument handling and behavior (`internal/cli/run_test.go`)
|
|
- subprocess stdout/stderr and exit-code behavior (`cmd/audita/main_integration_test.go`)
|
|
- config/env/override validation (`internal/core/config/*_test.go`)
|
|
- transcript and glossary schema validation (`internal/core/schema/*_test.go`)
|
|
- deterministic normalization (`internal/core/normalization/*_test.go`)
|
|
- deterministic chunking and summaries (`internal/core/chunking/*_test.go`)
|
|
- proposal preview/apply semantics (`internal/framework/proposals/*_test.go`)
|
|
- contracts/foundation composition tests (`internal/framework/contracts/*_test.go`)
|
|
- runner sequencing and failure behavior with deterministic fake modules (`internal/framework/runner/*_test.go`)
|
|
- CLI runner integration through injected fake module factories (`internal/cli/run_test.go`)
|
|
- validator models, cardinality enforcement, and deterministic validators (`internal/framework/validators/*_test.go`)
|
|
- LLM-backed validator batching, prompt builders, structured-response safety, scheduler hooks, and diagnostics redaction (`internal/framework/validators/*_test.go`, `internal/framework/runner/*_test.go`)
|
|
- shared proposal-generation request/response parsing, deterministic indexing, scheduler hooks, and diagnostics redaction (`internal/framework/proposal_generation/*_test.go`, `internal/framework/runner/*_test.go`)
|
|
- production module-registry known-key recognition and unsupported/unimplemented error behavior (`internal/framework/modules/*_test.go`, `internal/cli/run_test.go`)
|
|
- production grammar module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, and explicit CLI/runtime integration (`internal/modules/grammar/*_test.go`, `internal/cli/run_test.go`, `internal/framework/runner/*_test.go`)
|
|
|
|
Not covered yet (because not implemented): production `glossary`, `homophones`, and `spoken_word` modules plus full default-sequence transcript-polishing runtime behavior.
|
|
|
|
## Intended final architecture (not yet implemented)
|
|
The intended end-state still matches the rewrite plan:
|
|
- sequential module pipeline over a mutable working transcript
|
|
- real module implementations (`glossary`, `homophones`, `spoken_word`, `grammar`)
|
|
- structured LLM proposal generation
|
|
- deterministic and LLM validators
|
|
- validator cardinality enforcement in pipeline execution
|
|
- proposal application integrated per module stage
|
|
- prompt/response diagnostics for LLM/module stages
|
|
|
|
Until those phases are implemented, documentation and external descriptions should treat the current Go CLI as deterministic preprocessing/reporting infrastructure, not a full LLM transcript polisher.
|