220 lines
6.9 KiB
Markdown
220 lines
6.9 KiB
Markdown
# Audita Go Architecture
|
|
|
|
## Scope and intent
|
|
This document describes:
|
|
- the current implemented Go architecture; and
|
|
- the intended final architecture for later rewrite phases.
|
|
|
|
Status labels are explicit so future engineers and LLM agents do not assume unimplemented behavior exists.
|
|
|
|
## Current implementation status
|
|
Implemented today:
|
|
- Go CLI entrypoint and `audita process` wiring.
|
|
- Config defaults, env loading, CLI override precedence, and validation.
|
|
- Transcript and glossary parsing/validation.
|
|
- Deterministic transcript normalization.
|
|
- Deterministic token estimation and transcript chunking.
|
|
- Per-run diagnostics directory creation plus basic artifacts.
|
|
- Minimal process report JSON output.
|
|
- Framework foundation packages for contracts and proposal application.
|
|
|
|
Not implemented in CLI runtime path today:
|
|
- Real module execution pipeline (`glossary`, `homophones`, `spoken_word`, `grammar`).
|
|
- Structured LLM proposal generation.
|
|
- Validator chain execution.
|
|
- Pipeline runner orchestration.
|
|
- Production LLM scheduler behavior.
|
|
- End-to-end transcript polishing beyond normalization.
|
|
|
|
## Actual Go package layout
|
|
|
|
```text
|
|
cmd/audita/
|
|
main.go
|
|
|
|
internal/cli/
|
|
run.go
|
|
|
|
internal/core/config/
|
|
config.go
|
|
env.go
|
|
flags.go
|
|
redaction.go
|
|
validation.go
|
|
|
|
internal/core/schema/
|
|
transcript.go
|
|
glossary.go
|
|
errors.go
|
|
|
|
internal/core/io/
|
|
files.go
|
|
|
|
internal/core/normalization/
|
|
normalize.go
|
|
tokens.go
|
|
|
|
internal/core/chunking/
|
|
sections.go
|
|
summary.go
|
|
tokens.go
|
|
|
|
internal/core/diagnostics/
|
|
run_dir.go
|
|
|
|
internal/core/reporting/
|
|
report.go
|
|
|
|
internal/framework/contracts/
|
|
contracts.go
|
|
|
|
internal/framework/proposals/
|
|
proposal.go
|
|
policy.go
|
|
preview.go
|
|
apply.go
|
|
```
|
|
|
|
## Current CLI behavior
|
|
Primary command:
|
|
|
|
```sh
|
|
audita process <transcript.json> --glossary <glossary.yaml> [flags]
|
|
```
|
|
|
|
Current runtime flow (`internal/cli/run.go`):
|
|
1. Load config from env.
|
|
2. Parse flags and apply CLI overrides.
|
|
3. Validate transcript positional argument and required `--glossary`.
|
|
4. Create per-run diagnostics directory.
|
|
5. Read transcript and glossary files.
|
|
6. Parse/validate transcript and glossary.
|
|
7. Write source transcript artifacts.
|
|
8. Normalize transcript.
|
|
9. Write normalized transcript and normalization summary artifacts.
|
|
10. Chunk normalized transcript and compute chunk summaries.
|
|
11. Write chunking summary artifact.
|
|
12. Output normalized transcript to `--output` file or stdout.
|
|
13. Build process report (`phase` currently set to `phase3-chunking`).
|
|
14. Optionally write `--report-json`; always write run-dir `report.json`.
|
|
15. Apply work-dir retention.
|
|
|
|
Important behavior details:
|
|
- Glossary is validated but not yet used for correction logic.
|
|
- No module execution occurs despite `--modules` config support.
|
|
- No LLM calls occur.
|
|
- Success path is generally quiet on stderr.
|
|
|
|
## Implemented data contracts
|
|
|
|
### Transcript input
|
|
Accepted top-level forms:
|
|
- bare JSON array of segments
|
|
- object with `segments` array
|
|
|
|
Source segment contract:
|
|
- `id` optional integer
|
|
- `speaker` non-empty string
|
|
- `start` finite non-negative number
|
|
- `end` finite non-negative number with `end >= start`
|
|
- `text` non-empty string
|
|
- `categories` optional array of non-empty strings
|
|
|
|
Additional checks:
|
|
- duplicate explicit source IDs are rejected.
|
|
|
|
### Transcript output
|
|
Current output uses `schema.TranscriptToJSON` and is a bare JSON array of normalized segments:
|
|
- `id`, `speaker`, `start`, `end`, `text`, optional `categories`.
|
|
|
|
### Glossary input
|
|
YAML with `glossary` entries. Required fields per entry:
|
|
- `name`, `category`, `summary`
|
|
|
|
Optional:
|
|
- `aliases`, `plural`
|
|
|
|
## Implemented config/env/flag behavior
|
|
Precedence:
|
|
1. defaults (`config.Default()`)
|
|
2. environment (`config.LoadFromEnv()`)
|
|
3. CLI flags (`ApplyCLIOverrides`)
|
|
|
|
Implemented config surfaces include:
|
|
- module list
|
|
- primary and validation LLM settings
|
|
- section token controls and target sections
|
|
- confidence thresholds
|
|
- normalization controls
|
|
- work-dir and retention mode
|
|
|
|
Current caveat:
|
|
- LLM/module-related settings are mostly infrastructure-only today; runtime path does not execute LLM or modules.
|
|
|
|
## Implemented normalization behavior
|
|
Normalization (`internal/core/normalization`) currently:
|
|
- sorts by segment start time;
|
|
- merges adjacent same-speaker segments when constraints pass;
|
|
- uses gap-based joiners:
|
|
- gap `< ellipsis_gap` -> single space join
|
|
- gap `>= ellipsis_gap` -> `... ` join
|
|
- enforces merged duration and token-limit constraints;
|
|
- reassigns output IDs sequentially from `1`;
|
|
- returns `NormalizationSummary` with merge and skip counters.
|
|
|
|
Note: merged categories are concatenated (not deduplicated).
|
|
|
|
## Implemented chunking behavior
|
|
Chunking (`internal/core/chunking`) currently provides:
|
|
- deterministic heuristic token estimation;
|
|
- contiguous sectioning with section metadata;
|
|
- max/min section token validation;
|
|
- optional `target_sections` handling with target-aware merge/split logic;
|
|
- summary and detailed summary generation.
|
|
|
|
Current behavior details:
|
|
- if a single segment exceeds max tokens, it is emitted as its own section (not hard-failed);
|
|
- section balancing is deterministic but heuristic.
|
|
|
|
## Implemented proposal/replacement infrastructure
|
|
`internal/framework/proposals` provides deterministic foundation logic:
|
|
- `CorrectionProposal` and `EnrichedCorrectionProposal` models;
|
|
- replacement policies: `require_unique`, `replace_all`;
|
|
- safe preview (`PreviewProposalForSegment`) with stable skip reasons;
|
|
- deterministic apply (`ApplyProposals`) in ascending `proposal_index` order;
|
|
- applied/skipped change records suitable for reporting.
|
|
|
|
`internal/framework/contracts` provides interfaces and run-spec metadata scaffolding, including deterministic repeated module instance naming (`ResolveModuleRunSpecs`).
|
|
|
|
These are framework primitives only; they are not yet wired into CLI runtime module execution.
|
|
|
|
## Reports and diagnostics (implemented)
|
|
Current per-run artifacts include:
|
|
- `source-transcript.json`
|
|
- `source-transcript-parsed.json`
|
|
- `normalized-transcript.json`
|
|
- `normalization-summary.json`
|
|
- `chunking-summary.json`
|
|
- `report.json`
|
|
- `error.log` on failure
|
|
|
|
`--report-json` writes a separate report file when requested.
|
|
|
|
Retention modes implemented in `ApplyRetention`:
|
|
- `always`: keep successful runs
|
|
- `never`: remove successful runs
|
|
- `auto`: currently same as `never` for successful runs
|
|
- failed runs are always retained
|
|
|
|
## Intended final architecture (not yet implemented)
|
|
The intended end-state still matches the rewrite plan:
|
|
- sequential module pipeline over a mutable working transcript
|
|
- real module implementations (`glossary`, `homophones`, `spoken_word`, `grammar`)
|
|
- structured LLM proposal generation
|
|
- deterministic and LLM validators
|
|
- validator cardinality enforcement in pipeline execution
|
|
- proposal application integrated per module stage
|
|
- richer run reporting with module-level applied/skipped changes
|
|
|
|
Until those phases are implemented, documentation and external descriptions should treat the current Go CLI as deterministic preprocessing/reporting infrastructure, not a full LLM transcript polisher.
|