Files
audita/docs/architecture.md

19 KiB

Audita Go Architecture

Scope and intent

This document describes:

  • the current implemented Go architecture; and
  • the intended final architecture for later rewrite phases.

Status labels are explicit so future engineers and LLM agents do not assume unimplemented behavior exists.

Current implementation status

Implemented today:

  • Go CLI entrypoint and audita process wiring.
  • Config defaults, env loading, CLI override precedence, and validation.
  • Transcript and glossary parsing/validation.
  • Deterministic transcript normalization.
  • Deterministic token estimation and transcript chunking.
  • Per-run diagnostics directory creation plus Phase 6 process-level artifacts.
  • Process report JSON output with diagnostics artifact references.
  • Framework foundation packages for contracts and proposal application.
  • Production runner orchestration package with deterministic sequential module execution.
  • Module-level report structures with applied/skipped change records.
  • Runtime validator models and deterministic validators.
  • Deterministic validator-chain execution in the runner with cardinality enforcement.
  • Module-level validator decision/rejection reporting.
  • Internal structured LLM client contract plus an instructor-go-backed adapter package.
  • Bounded LLM scheduler/semaphore infrastructure with context-aware permit handling.
  • Runtime primary/validation LLM effective-config resolution helpers with validation inheritance.
  • Generic JSON prompt/response diagnostics writer primitives with secret redaction.
  • LLM-backed validator models, prompt builders, batching, and runtime execution.
  • Runner wiring for LLM validators via the internal structured LLM abstraction and scheduler hooks.
  • LLM validator diagnostics artifacts and report-level decision metadata paths.
  • Shared LLM proposal-generation helper with structured correction-set parsing.
  • Deterministic proposal-index assignment and enriched proposal mapping for shared generation.
  • Proposal-generation diagnostics artifacts with secret redaction.
  • Production module registry scaffolding with known-key recognition and explicit unsupported/unimplemented errors.
  • Production grammar module implementation in internal/modules/grammar.
  • Explicit runtime support for --modules grammar through the production runner path.

Not implemented in CLI runtime path today:

  • Real module execution pipeline for glossary, homophones, and spoken_word.
  • Real domain proposal prompts for production modules.
  • End-to-end transcript polishing with real module behavior.

Phase sequencing note:

  • Phase 9 LLM infrastructure is complete (structured client, scheduler, effective config resolution, diagnostics primitives);
  • Phase 10 LLM-backed validator runtime integration is complete;
  • Phase 11 shared proposal-generation framework and module-registry scaffolding are complete;
  • Phase 12 grammar module implementation and explicit runtime wiring are complete;
  • next recommended phase is Phase 13 (glossary module and protected-term behavior).

Actual Go package layout

cmd/audita/
  main.go

internal/cli/
  run.go

internal/core/config/
  config.go
  env.go
  flags.go
  redaction.go
  validation.go

internal/core/schema/
  transcript.go
  glossary.go
  errors.go

internal/core/io/
  files.go

internal/core/normalization/
  normalize.go
  tokens.go

internal/core/chunking/
  sections.go
  summary.go
  tokens.go

internal/core/diagnostics/
  run_dir.go

internal/core/reporting/
  report.go

internal/framework/contracts/
  contracts.go

internal/framework/proposals/
  proposal.go
  policy.go
  preview.go
  apply.go

internal/framework/runner/
  runner.go

internal/framework/proposal_generation/
  generate.go

internal/framework/modules/
  registry.go

internal/modules/grammar/
  module.go
  prompt.go

internal/framework/validators/
  models.go
  deterministic.go
  llm_models.go
  llm_prompt_builders.go
  llm_batching.go
  llm_validators.go

internal/framework/llm/
  instructor_client.go
  scheduler.go
  effective_config.go
  diagnostics.go

Current CLI behavior

Primary command:

audita process <transcript.json> --glossary <glossary.yaml> [flags]

Current runtime flow (internal/cli/run.go):

  1. Load config from env.
  2. Parse flags and apply CLI overrides.
  3. Validate transcript positional argument and required --glossary.
  4. Create per-run diagnostics directory.
  5. Read transcript and glossary files.
  6. Parse/validate transcript and glossary.
  7. Write source transcript artifacts.
  8. Normalize transcript.
  9. Write normalized transcript and normalization summary artifacts.
  10. Chunk normalized transcript and compute chunk summaries.
  11. Write chunking summary artifact.
  12. Execute runner modules sequentially when:
    • --modules is explicitly provided (production grammar path); or
    • a test/injected module factory is provided.
  13. Output working transcript to --output file or stdout.
  14. Build process report (phase currently set to phase12-grammar-module).
  15. Optionally write --report-json; always write run-dir report.json.
  16. Apply work-dir retention.

Important behavior details:

  • Glossary is validated but not yet used for real correction module logic.
  • Default production CLI behavior remains deterministic normalization/chunking/reporting unless modules are explicitly selected with --modules.
  • Explicit --modules grammar runs the production grammar module path with LLM-backed proposal generation and validator-chain execution.
  • Default runs (without explicit module selection) do not perform LLM calls.
  • Success path is generally quiet on stderr.
  • Source IDs are preserved into a canonical transcript before normalization; normalization then reassigns output IDs sequentially from 1.

Implemented data contracts

Transcript input

Accepted top-level forms:

  • bare JSON array of segments
  • object with segments array

Source segment contract:

  • id optional integer
  • speaker non-empty string
  • start finite non-negative number
  • end finite non-negative number with end >= start
  • text non-empty string
  • categories optional array of non-empty strings

Additional checks:

  • duplicate explicit source IDs are rejected.

Transcript output

Current output uses schema.TranscriptToJSON and is a bare JSON array of normalized segments:

  • id, speaker, start, end, text, optional categories.

Glossary input

YAML with glossary entries. Required fields per entry:

  • name, category, summary

Optional:

  • aliases, plural

Implemented config/env/flag behavior

Precedence:

  1. defaults (config.Default())
  2. environment (config.LoadFromEnv())
  3. CLI flags (ApplyCLIOverrides)

Implemented config surfaces include:

  • module list
  • primary and validation LLM settings
  • section token controls and target sections
  • confidence thresholds
  • normalization controls
  • work-dir and retention mode

Current caveat:

  • LLM/module-related settings are active for explicit grammar runs; the default non-explicit path remains deterministic.

Implemented structured LLM infrastructure

internal/framework/contracts now defines a typed structured-completion contract:

  • StructuredLLMClient.CompleteStructured(ctx, req, out)
  • caller-owned typed decode target via out pointer.

internal/framework/llm provides InstructorClient, an internal adapter over github.com/jxnl/instructor-go:

  • configurable base_url, model, optional API key, retries, mode, HTTP client, and request timeout;
  • OpenAI-compatible endpoint behavior (for example OpenAI/OpenRouter/local-compatible base URLs);
  • default mode is JSON mode (ModeJSON), with optional tool-call mode (ModeToolCall);
  • request message translation from contracts.LLMMessage to chat-completions messages;
  • response metadata mapping (provider/model/token usage) into Audita-owned response types;
  • API-key redaction in adapter-returned errors.

Current runtime boundary:

  • the default CLI runtime path (without explicit module selection) still does not instantiate the full production module sequence.
  • LLM calls are exercised in production when --modules grammar is explicitly requested and in tests when fake/injected clients are used.

internal/framework/llm also provides:

  • a bounded Scheduler for controlled concurrent LLM calls with reliable permit release;
  • primary/validation effective-config resolution helpers, including validation inheritance fallback to primary settings;
  • generic interaction diagnostics primitives that write machine-readable JSON artifacts for request metadata, request payload, response payload, and optional error payload with secret redaction.

Implemented normalization behavior

Normalization (internal/core/normalization) currently:

  • sorts by segment start time;
  • merges adjacent same-speaker segments when constraints pass;
  • uses gap-based joiners:
    • gap < ellipsis_gap -> single space join
    • gap >= ellipsis_gap -> ... join
  • enforces merged duration and token-limit constraints;
  • reassigns output IDs sequentially from 1;
  • returns NormalizationSummary with merge and skip counters.

Note: merged categories are concatenated (not deduplicated).

Implemented chunking behavior

Chunking (internal/core/chunking) currently provides:

  • deterministic heuristic token estimation;
  • contiguous sectioning with section metadata;
  • max/min section token validation;
  • optional target_sections handling with target-aware merge/split logic;
  • summary and detailed summary generation.

Current behavior details:

  • if a single segment exceeds max tokens, it is emitted as its own section (not hard-failed);
  • section balancing is deterministic but heuristic.

Implemented proposal/replacement infrastructure

internal/framework/proposals provides deterministic foundation logic:

  • CorrectionProposal and EnrichedCorrectionProposal models;
  • replacement policies: require_unique, replace_all;
  • safe preview (PreviewProposalForSegment) with stable skip reasons;
  • deterministic apply (ApplyProposals) in ascending proposal_index order;
  • applied/skipped change records suitable for reporting.

internal/framework/contracts provides interfaces and run-spec metadata scaffolding, including deterministic repeated module instance naming (ResolveModuleRunSpecs).

These primitives are wired into the production runner and report model. The grammar module is implemented; other production modules remain pending.

Implemented validator runtime infrastructure

internal/framework/validators provides deterministic validator infrastructure:

  • runtime validation request/result models;
  • stable validator reason codes;
  • cardinality enforcement for validator decisions:
    • missing proposal indexes fail
    • duplicate proposal indexes fail
    • unknown proposal indexes fail
  • deterministic validators:
    • confidence threshold by module key/config threshold
    • original-text presence against current working transcript
    • non-empty corrected text
    • identical/no-effect rejection
    • conservative protected glossary-term guard for non-glossary modules

internal/framework/runner executes validator chains in order for each module and applies only validator-approved proposals. Validator rejections are reported distinctly from proposal-application skips.

Implemented LLM-backed validator infrastructure

internal/framework/validators now includes LLM-backed validator support:

  • typed request/response models for structured LLM validation;
  • prompt builders for:
    • spoken-form plausibility
    • meaning reversal detection
    • editorial review
    • grammar review
    • spoken-word review
  • deterministic batching by validation_max_prompt_tokens;
  • strict cardinality validation of structured LLM decisions (missing/duplicate/unknown indexes fail);
  • safe failure behavior for malformed/invalid structured responses.

internal/framework/runner wires LLM validators into existing validator chains using:

  • the internal structured LLM client abstraction (contracts.StructuredLLMClient);
  • bounded scheduler hooks for validator call execution;
  • diagnostics writer hooks for machine-readable prompt/response artifacts with secret redaction.

Implemented shared proposal-generation infrastructure

internal/framework/proposal_generation provides a reusable, prompt-agnostic helper for future real modules:

  • structured request model including module key/instance, replacement policy, working transcript context, optional section metadata, glossary, config, and diagnostics context;
  • structured correction-set response model (corrections) mapped into existing proposals.CorrectionProposal and proposals.EnrichedCorrectionProposal models;
  • deterministic proposal-index assignment through a caller-provided start_index;
  • structured LLM calls through contracts.StructuredLLMClient only (no direct provider calls);
  • optional bounded execution through scheduler hooks (contracts.LLMScheduler);
  • prompt/response diagnostics artifact writing via the generic internal/framework/llm diagnostics primitives with redaction of API keys/secrets.

This helper only produces candidate proposals; validator-chain execution and proposal application remain runner responsibilities.

Implemented production module-registry scaffolding

internal/framework/modules now provides a production registry scaffold:

  • recognizes intended module keys:
    • glossary
    • homophones
    • spoken_word
    • grammar
  • supports explicit constructor registration with dependency injection for:
    • run spec
    • config
    • glossary
    • proposal/validation structured LLM clients
    • proposal/validation schedulers
    • diagnostics directory context
  • returns explicit errors for unknown keys (unsupported_module) and recognized-but-unimplemented keys (unimplemented_module).

The grammar module key is now registered and constructible. glossary, homophones, and spoken_word remain recognized-but-unimplemented.

Implemented grammar production module

internal/modules/grammar now provides the first production module:

  • prompt builder faithfully constrained to punctuation/capitalization/spacing/article cleanup;
  • explicit guardrails against meaning-changing rewrites, style rewrites, summarization, and invention;
  • proposal generation through internal/framework/proposal_generation and contracts.StructuredLLMClient;
  • scheduler-aware proposal calls through existing contracts.LLMScheduler hooks;
  • replacement policy require_unique (matching Python implementation);
  • validator chain integration using existing deterministic + LLM-backed validators;
  • grammar confidence threshold enforcement through existing validator/config infrastructure;
  • module-level reporting and diagnostics capture through existing runner/reporting paths.

Reports and diagnostics (implemented)

Current per-run artifacts include:

  • source-transcript.json
  • source-transcript-parsed.json
  • normalized-transcript.json
  • normalization-summary.json
  • chunking-summary.json
  • invocation.json
  • effective-config.json (redacted credentials)
  • report.json
  • error.log on failure

--report-json writes a separate report file when requested.

Current process reports include diagnostics metadata references for:

  • diagnostics directory path;
  • source transcript artifact path;
  • parsed source transcript artifact path;
  • normalized transcript artifact path;
  • normalization summary artifact path;
  • chunking summary artifact path;
  • invocation metadata artifact path;
  • redacted effective-config artifact path;
  • error-log artifact path on failure.

Current process reports also include:

  • module-level results (when runner modules execute), including applied/skipped proposal changes;
  • run-level module summary totals and failed module instance metadata.
  • module-level validator decisions and validator rejections.
  • optional decision-level diagnostic artifact paths for validator LLM interactions when available.

Retention modes implemented in ApplyRetention:

  • always: keep all run directories.
  • never: keep successful run directories.
  • auto: keep failed runs and successful runs with skipped corrections.
  • failed runs are always retained.

Current runtime note:

  • default non-explicit runs usually have no module-level skipped corrections, so auto commonly removes clean successful run directories.
  • explicit grammar runs can produce validator rejections and application skips, which are reflected in reports and retention input.

Intentionally deferred to module/LLM phases:

  • real domain proposal prompts and production module implementations remain tied to later module phases.

Current tests and quality posture

Implemented tests currently cover:

  • CLI argument handling and behavior (internal/cli/run_test.go)
  • subprocess stdout/stderr and exit-code behavior (cmd/audita/main_integration_test.go)
  • config/env/override validation (internal/core/config/*_test.go)
  • transcript and glossary schema validation (internal/core/schema/*_test.go)
  • deterministic normalization (internal/core/normalization/*_test.go)
  • deterministic chunking and summaries (internal/core/chunking/*_test.go)
  • proposal preview/apply semantics (internal/framework/proposals/*_test.go)
  • contracts/foundation composition tests (internal/framework/contracts/*_test.go)
  • runner sequencing and failure behavior with deterministic fake modules (internal/framework/runner/*_test.go)
  • CLI runner integration through injected fake module factories (internal/cli/run_test.go)
  • validator models, cardinality enforcement, and deterministic validators (internal/framework/validators/*_test.go)
  • LLM-backed validator batching, prompt builders, structured-response safety, scheduler hooks, and diagnostics redaction (internal/framework/validators/*_test.go, internal/framework/runner/*_test.go)
  • shared proposal-generation request/response parsing, deterministic indexing, scheduler hooks, and diagnostics redaction (internal/framework/proposal_generation/*_test.go, internal/framework/runner/*_test.go)
  • production module-registry known-key recognition and unsupported/unimplemented error behavior (internal/framework/modules/*_test.go, internal/cli/run_test.go)
  • production grammar module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, and explicit CLI/runtime integration (internal/modules/grammar/*_test.go, internal/cli/run_test.go, internal/framework/runner/*_test.go)

Not covered yet (because not implemented): production glossary, homophones, and spoken_word modules plus full default-sequence transcript-polishing runtime behavior.

Intended final architecture (not yet implemented)

The intended end-state still matches the rewrite plan:

  • sequential module pipeline over a mutable working transcript
  • real module implementations (glossary, homophones, spoken_word, grammar)
  • structured LLM proposal generation
  • deterministic and LLM validators
  • validator cardinality enforcement in pipeline execution
  • proposal application integrated per module stage
  • prompt/response diagnostics for LLM/module stages

Until those phases are implemented, documentation and external descriptions should treat the current Go CLI as deterministic preprocessing/reporting infrastructure, not a full LLM transcript polisher.