Files
audita/docs/architecture.md

6.9 KiB

Audita Go Architecture

Scope and intent

This document describes:

  • the current implemented Go architecture; and
  • the intended final architecture for later rewrite phases.

Status labels are explicit so future engineers and LLM agents do not assume unimplemented behavior exists.

Current implementation status

Implemented today:

  • Go CLI entrypoint and audita process wiring.
  • Config defaults, env loading, CLI override precedence, and validation.
  • Transcript and glossary parsing/validation.
  • Deterministic transcript normalization.
  • Deterministic token estimation and transcript chunking.
  • Per-run diagnostics directory creation plus basic artifacts.
  • Minimal process report JSON output.
  • Framework foundation packages for contracts and proposal application.

Not implemented in CLI runtime path today:

  • Real module execution pipeline (glossary, homophones, spoken_word, grammar).
  • Structured LLM proposal generation.
  • Validator chain execution.
  • Pipeline runner orchestration.
  • Production LLM scheduler behavior.
  • End-to-end transcript polishing beyond normalization.

Actual Go package layout

cmd/audita/
  main.go

internal/cli/
  run.go

internal/core/config/
  config.go
  env.go
  flags.go
  redaction.go
  validation.go

internal/core/schema/
  transcript.go
  glossary.go
  errors.go

internal/core/io/
  files.go

internal/core/normalization/
  normalize.go
  tokens.go

internal/core/chunking/
  sections.go
  summary.go
  tokens.go

internal/core/diagnostics/
  run_dir.go

internal/core/reporting/
  report.go

internal/framework/contracts/
  contracts.go

internal/framework/proposals/
  proposal.go
  policy.go
  preview.go
  apply.go

Current CLI behavior

Primary command:

audita process <transcript.json> --glossary <glossary.yaml> [flags]

Current runtime flow (internal/cli/run.go):

  1. Load config from env.
  2. Parse flags and apply CLI overrides.
  3. Validate transcript positional argument and required --glossary.
  4. Create per-run diagnostics directory.
  5. Read transcript and glossary files.
  6. Parse/validate transcript and glossary.
  7. Write source transcript artifacts.
  8. Normalize transcript.
  9. Write normalized transcript and normalization summary artifacts.
  10. Chunk normalized transcript and compute chunk summaries.
  11. Write chunking summary artifact.
  12. Output normalized transcript to --output file or stdout.
  13. Build process report (phase currently set to phase3-chunking).
  14. Optionally write --report-json; always write run-dir report.json.
  15. Apply work-dir retention.

Important behavior details:

  • Glossary is validated but not yet used for correction logic.
  • No module execution occurs despite --modules config support.
  • No LLM calls occur.
  • Success path is generally quiet on stderr.

Implemented data contracts

Transcript input

Accepted top-level forms:

  • bare JSON array of segments
  • object with segments array

Source segment contract:

  • id optional integer
  • speaker non-empty string
  • start finite non-negative number
  • end finite non-negative number with end >= start
  • text non-empty string
  • categories optional array of non-empty strings

Additional checks:

  • duplicate explicit source IDs are rejected.

Transcript output

Current output uses schema.TranscriptToJSON and is a bare JSON array of normalized segments:

  • id, speaker, start, end, text, optional categories.

Glossary input

YAML with glossary entries. Required fields per entry:

  • name, category, summary

Optional:

  • aliases, plural

Implemented config/env/flag behavior

Precedence:

  1. defaults (config.Default())
  2. environment (config.LoadFromEnv())
  3. CLI flags (ApplyCLIOverrides)

Implemented config surfaces include:

  • module list
  • primary and validation LLM settings
  • section token controls and target sections
  • confidence thresholds
  • normalization controls
  • work-dir and retention mode

Current caveat:

  • LLM/module-related settings are mostly infrastructure-only today; runtime path does not execute LLM or modules.

Implemented normalization behavior

Normalization (internal/core/normalization) currently:

  • sorts by segment start time;
  • merges adjacent same-speaker segments when constraints pass;
  • uses gap-based joiners:
    • gap < ellipsis_gap -> single space join
    • gap >= ellipsis_gap -> ... join
  • enforces merged duration and token-limit constraints;
  • reassigns output IDs sequentially from 1;
  • returns NormalizationSummary with merge and skip counters.

Note: merged categories are concatenated (not deduplicated).

Implemented chunking behavior

Chunking (internal/core/chunking) currently provides:

  • deterministic heuristic token estimation;
  • contiguous sectioning with section metadata;
  • max/min section token validation;
  • optional target_sections handling with target-aware merge/split logic;
  • summary and detailed summary generation.

Current behavior details:

  • if a single segment exceeds max tokens, it is emitted as its own section (not hard-failed);
  • section balancing is deterministic but heuristic.

Implemented proposal/replacement infrastructure

internal/framework/proposals provides deterministic foundation logic:

  • CorrectionProposal and EnrichedCorrectionProposal models;
  • replacement policies: require_unique, replace_all;
  • safe preview (PreviewProposalForSegment) with stable skip reasons;
  • deterministic apply (ApplyProposals) in ascending proposal_index order;
  • applied/skipped change records suitable for reporting.

internal/framework/contracts provides interfaces and run-spec metadata scaffolding, including deterministic repeated module instance naming (ResolveModuleRunSpecs).

These are framework primitives only; they are not yet wired into CLI runtime module execution.

Reports and diagnostics (implemented)

Current per-run artifacts include:

  • source-transcript.json
  • source-transcript-parsed.json
  • normalized-transcript.json
  • normalization-summary.json
  • chunking-summary.json
  • report.json
  • error.log on failure

--report-json writes a separate report file when requested.

Retention modes implemented in ApplyRetention:

  • always: keep successful runs
  • never: remove successful runs
  • auto: currently same as never for successful runs
  • failed runs are always retained

Intended final architecture (not yet implemented)

The intended end-state still matches the rewrite plan:

  • sequential module pipeline over a mutable working transcript
  • real module implementations (glossary, homophones, spoken_word, grammar)
  • structured LLM proposal generation
  • deterministic and LLM validators
  • validator cardinality enforcement in pipeline execution
  • proposal application integrated per module stage
  • richer run reporting with module-level applied/skipped changes

Until those phases are implemented, documentation and external descriptions should treat the current Go CLI as deterministic preprocessing/reporting infrastructure, not a full LLM transcript polisher.