7.9 KiB
Audita Go Architecture
Scope and intent
This document describes:
- the current implemented Go architecture; and
- the intended final architecture for later rewrite phases.
Status labels are explicit so future engineers and LLM agents do not assume unimplemented behavior exists.
Current implementation status
Implemented today:
- Go CLI entrypoint and
audita processwiring. - Config defaults, env loading, CLI override precedence, and validation.
- Transcript and glossary parsing/validation.
- Deterministic transcript normalization.
- Deterministic token estimation and transcript chunking.
- Per-run diagnostics directory creation plus basic artifacts.
- Minimal process report JSON output.
- Framework foundation packages for contracts and proposal application.
Not implemented in CLI runtime path today:
- Real module execution pipeline (
glossary,homophones,spoken_word,grammar). - Structured LLM proposal generation.
- Validator chain execution.
- Pipeline runner orchestration.
- Production LLM scheduler behavior.
- End-to-end transcript polishing beyond normalization.
Actual Go package layout
cmd/audita/
main.go
internal/cli/
run.go
internal/core/config/
config.go
env.go
flags.go
redaction.go
validation.go
internal/core/schema/
transcript.go
glossary.go
errors.go
internal/core/io/
files.go
internal/core/normalization/
normalize.go
tokens.go
internal/core/chunking/
sections.go
summary.go
tokens.go
internal/core/diagnostics/
run_dir.go
internal/core/reporting/
report.go
internal/framework/contracts/
contracts.go
internal/framework/proposals/
proposal.go
policy.go
preview.go
apply.go
Current CLI behavior
Primary command:
audita process <transcript.json> --glossary <glossary.yaml> [flags]
Current runtime flow (internal/cli/run.go):
- Load config from env.
- Parse flags and apply CLI overrides.
- Validate transcript positional argument and required
--glossary. - Create per-run diagnostics directory.
- Read transcript and glossary files.
- Parse/validate transcript and glossary.
- Write source transcript artifacts.
- Normalize transcript.
- Write normalized transcript and normalization summary artifacts.
- Chunk normalized transcript and compute chunk summaries.
- Write chunking summary artifact.
- Output normalized transcript to
--outputfile or stdout. - Build process report (
phasecurrently set tophase3-chunking). - Optionally write
--report-json; always write run-dirreport.json. - Apply work-dir retention.
Important behavior details:
- Glossary is validated but not yet used for correction logic.
- No module execution occurs despite
--modulesconfig support. - No LLM calls occur.
- Success path is generally quiet on stderr.
- Source IDs are preserved into a canonical transcript before normalization; normalization then reassigns output IDs sequentially from
1.
Implemented data contracts
Transcript input
Accepted top-level forms:
- bare JSON array of segments
- object with
segmentsarray
Source segment contract:
idoptional integerspeakernon-empty stringstartfinite non-negative numberendfinite non-negative number withend >= starttextnon-empty stringcategoriesoptional array of non-empty strings
Additional checks:
- duplicate explicit source IDs are rejected.
Transcript output
Current output uses schema.TranscriptToJSON and is a bare JSON array of normalized segments:
id,speaker,start,end,text, optionalcategories.
Glossary input
YAML with glossary entries. Required fields per entry:
name,category,summary
Optional:
aliases,plural
Implemented config/env/flag behavior
Precedence:
- defaults (
config.Default()) - environment (
config.LoadFromEnv()) - CLI flags (
ApplyCLIOverrides)
Implemented config surfaces include:
- module list
- primary and validation LLM settings
- section token controls and target sections
- confidence thresholds
- normalization controls
- work-dir and retention mode
Current caveat:
- LLM/module-related settings are mostly infrastructure-only today; runtime path does not execute LLM or modules.
Implemented normalization behavior
Normalization (internal/core/normalization) currently:
- sorts by segment start time;
- merges adjacent same-speaker segments when constraints pass;
- uses gap-based joiners:
- gap
< ellipsis_gap-> single space join - gap
>= ellipsis_gap->...join
- gap
- enforces merged duration and token-limit constraints;
- reassigns output IDs sequentially from
1; - returns
NormalizationSummarywith merge and skip counters.
Note: merged categories are concatenated (not deduplicated).
Implemented chunking behavior
Chunking (internal/core/chunking) currently provides:
- deterministic heuristic token estimation;
- contiguous sectioning with section metadata;
- max/min section token validation;
- optional
target_sectionshandling with target-aware merge/split logic; - summary and detailed summary generation.
Current behavior details:
- if a single segment exceeds max tokens, it is emitted as its own section (not hard-failed);
- section balancing is deterministic but heuristic.
Implemented proposal/replacement infrastructure
internal/framework/proposals provides deterministic foundation logic:
CorrectionProposalandEnrichedCorrectionProposalmodels;- replacement policies:
require_unique,replace_all; - safe preview (
PreviewProposalForSegment) with stable skip reasons; - deterministic apply (
ApplyProposals) in ascendingproposal_indexorder; - applied/skipped change records suitable for reporting.
internal/framework/contracts provides interfaces and run-spec metadata scaffolding, including deterministic repeated module instance naming (ResolveModuleRunSpecs).
These are framework primitives only; they are not yet wired into CLI runtime module execution.
Reports and diagnostics (implemented)
Current per-run artifacts include:
source-transcript.jsonsource-transcript-parsed.jsonnormalized-transcript.jsonnormalization-summary.jsonchunking-summary.jsonreport.jsonerror.logon failure
--report-json writes a separate report file when requested.
Retention modes implemented in ApplyRetention:
always: keep successful runsnever: remove successful runsauto: currently same asneverfor successful runs- failed runs are always retained
Current tests and quality posture
Implemented tests currently cover:
- CLI argument handling and behavior (
internal/cli/run_test.go) - subprocess stdout/stderr and exit-code behavior (
cmd/audita/main_integration_test.go) - config/env/override validation (
internal/core/config/*_test.go) - transcript and glossary schema validation (
internal/core/schema/*_test.go) - deterministic normalization (
internal/core/normalization/*_test.go) - deterministic chunking and summaries (
internal/core/chunking/*_test.go) - proposal preview/apply semantics (
internal/framework/proposals/*_test.go) - contracts/foundation composition tests (
internal/framework/contracts/*_test.go)
Not covered yet (because not implemented): end-to-end module runner behavior, validator runtime flow, and real LLM integration.
Intended final architecture (not yet implemented)
The intended end-state still matches the rewrite plan:
- sequential module pipeline over a mutable working transcript
- real module implementations (
glossary,homophones,spoken_word,grammar) - structured LLM proposal generation
- deterministic and LLM validators
- validator cardinality enforcement in pipeline execution
- proposal application integrated per module stage
- richer run reporting with module-level applied/skipped changes
Until those phases are implemented, documentation and external descriptions should treat the current Go CLI as deterministic preprocessing/reporting infrastructure, not a full LLM transcript polisher.