30 KiB
Audita Architecture
Scope and intent
This document describes:
- the architecture used in production today.
Historical rewrite details live in docs/rewrite-notes.md.
Current implementation status
Implemented today:
- Go CLI entrypoint and
audita processwiring. - Config defaults, env loading, CLI override precedence, and validation.
- Transcript and glossary parsing/validation.
- Deterministic transcript normalization.
- Deterministic token estimation and transcript chunking.
- Per-run diagnostics directory creation plus process-level artifacts.
- Process report JSON output with diagnostics artifact references.
- Framework foundation packages for contracts and proposal application.
- Production runner orchestration package with deterministic sequential module execution.
- Module-level report structures with applied/skipped change records.
- Runtime validator models and deterministic validators.
- Deterministic validator-chain execution in the runner with cardinality enforcement.
- Module-level validator decision/rejection reporting.
- Internal structured LLM client contract plus an Audita-owned OpenAI-compatible structured LLM adapter package.
- Bounded FIFO LLM scheduler infrastructure with context-aware permit handling.
- Runtime primary/validation LLM effective-config resolution helpers with validation inheritance.
- Generic JSON prompt/response diagnostics writer primitives with secret redaction.
- LLM-backed validator models, prompt builders, batching, and runtime execution.
- Runner wiring for LLM validators via the internal structured LLM abstraction and scheduler hooks.
- LLM validator diagnostics artifacts and report-level decision metadata paths.
- Shared LLM proposal-generation helper with structured correction-set parsing.
- Deterministic proposal-index assignment and enriched proposal mapping for shared generation.
- Proposal-generation diagnostics artifacts with secret redaction.
- Production module registry with known-key recognition and explicit unsupported-module errors.
- Production
grammarmodule implementation ininternal/modules/grammar. - Production
glossarymodule implementation ininternal/modules/glossary. - Production
homophonesmodule implementation ininternal/modules/homophones. - Production
spoken_wordmodule implementation ininternal/modules/spoken_word. - Explicit runtime support for
--modules grammarthrough the production runner path. - Explicit runtime support for
--modules glossary, including repeated stages such as--modules glossary,glossary. - Explicit runtime support for
--modules homophonesthrough the production runner path. - Explicit runtime support for
--modules spoken_wordthrough the production runner path.
Current reality:
- all production modules exist and are wired into the default runtime path.
- a normal
audita processrun without--modulesnow executes the full sequence:glossaryhomophonesglossaryspoken_wordgrammar
- repeated glossary stages are deterministic and reported distinctly as
glossary_1andglossary_2.
Actual Go package layout
cmd/audita/
main.go
internal/cli/
run.go
internal/core/config/
config.go
env.go
flags.go
redaction.go
validation.go
internal/core/schema/
transcript.go
glossary.go
errors.go
internal/core/io/
files.go
internal/core/normalization/
normalize.go
tokens.go
internal/core/chunking/
sections.go
summary.go
tokens.go
internal/core/diagnostics/
run_dir.go
internal/core/reporting/
report.go
internal/framework/contracts/
contracts.go
internal/framework/proposals/
proposal.go
policy.go
preview.go
apply.go
internal/framework/runner/
runner.go
internal/framework/proposal_generation/
generate.go
internal/framework/modules/
registry.go
internal/modules/grammar/
module.go
prompt.go
internal/modules/glossary/
module.go
prompt.go
internal/modules/homophones/
module.go
prompt.go
internal/modules/spoken_word/
module.go
prompt.go
internal/framework/validators/
models.go
deterministic.go
llm_models.go
llm_prompt_builders.go
llm_batching.go
llm_validators.go
internal/framework/llm/
openai_compatible_client.go
scheduler.go
effective_config.go
diagnostics.go
Current CLI behavior
Primary command:
audita process <transcript.json> --glossary <glossary.yaml> [flags]
Current runtime flow (internal/cli/run.go):
- Load config from env.
- Parse flags and apply CLI overrides.
- Validate transcript positional argument and required
--glossary. - Create per-run diagnostics directory.
- Read transcript and glossary files.
- Parse/validate transcript and glossary.
- Write source transcript artifacts.
- Normalize transcript.
- Write normalized transcript and normalization summary artifacts.
- Chunk normalized transcript and compute chunk summaries.
- Write chunking summary artifact.
- Execute runner modules sequentially:
- default run path uses configured default sequence (
glossary,homophones,glossary,spoken_word,grammar); - explicit
--modulesoverrides the default sequence; - test/injected module factory path remains available for deterministic runtime tests.
- each module recomputes chunks from the current working transcript, runs chunk proposal work concurrently, aggregates deterministically, validates, and applies approved proposals once.
- default run path uses configured default sequence (
- Output working transcript to
--outputfile or stdout. - Build process report (
phasecurrently set todefault_pipeline). - Optionally write
--report-json; always write run-dirreport.json. - Apply work-dir retention.
Parity fixture status:
- representative Python-parity fixture coverage exists under
internal/cli/testdata/parity; - parity tests use fake structured LLM responses for deterministic behavior, including default full-pipeline shape assertions;
- parity comparisons intentionally ignore nondeterministic metadata (timestamps, run IDs, temp paths, token usage) and remain strict for deterministic contract fields (transcript content, module order/instance naming, applied/skipped/rejected counts, and status).
- intentional Python-vs-Go differences and open parity gaps are documented in
docs/python-parity.md.
Important behavior details:
- Glossary is validated and is used for explicit glossary/grammar/homophones/spoken_word module correction paths.
- Default production CLI behavior now executes the full production module sequence unless
--modulesoverride is supplied. - Explicit
--modules grammar,--modules glossary,--modules homophones, and--modules spoken_wordcontinue to run production module paths with LLM-backed proposal generation and validator-chain execution. - Default runs (without explicit module selection) perform LLM calls through production module and validator paths.
- Success path is generally quiet on stderr.
- Source IDs are preserved into a canonical transcript before normalization; normalization then reassigns output IDs sequentially from
1.
Implemented data contracts
Transcript input
Accepted top-level forms:
- bare JSON array of segments
- object with
segmentsarray
Source segment contract:
idoptional integerspeakernon-empty stringstartfinite non-negative numberendfinite non-negative number withend >= starttextnon-empty stringcategoriesoptional array of non-empty strings
Additional checks:
- duplicate explicit source IDs are rejected.
Transcript output
Current output uses schema.TranscriptToJSON and is a bare JSON array of normalized segments:
id,speaker,start,end,text, optionalcategories.
Glossary input
YAML with glossary entries. Required fields per entry:
name,category,summary
Optional:
aliases,plural
Implemented config/env/flag behavior
Precedence:
- defaults (
config.Default()) - environment (
config.LoadFromEnv()) - CLI flags (
ApplyCLIOverrides)
Implemented config surfaces include:
- module list
- primary and validation LLM settings
- total/proposal/validation LLM concurrency controls
- transcript description context (
--transcript-description) - section token controls and target sections
- confidence thresholds
- normalization controls
- work-dir and retention mode
Current caveat:
- LLM/module-related settings are active for default and explicit module-run paths.
Transcript description behavior:
--transcript-descriptionis a process-flag input for optional user-supplied background context.- runtime config stores this value in
Config.TranscriptDescriptionafter CLI trimming and length validation. - default value is empty; empty values produce no prompt context section.
- this value is intentionally non-secret and appears in effective config and invocation metadata artifacts.
Implemented transcript description prompt context
Transcript description context is wired through production prompt paths:
- proposal prompts for
glossary,homophones,spoken_word, andgrammar; - LLM-backed validator prompts for spoken-form plausibility, meaning reversal, editorial review, grammar review, and spoken-word review.
Prompt guardrail semantics are consistent across modules and validators:
- transcript description is labeled as "background context only";
- it may help interpret ambiguous terms;
- it must not override transcript content;
- the model must not invent corrections, facts, names, events, motivations, or speaker intent from this description.
Generated transcript descriptions remain deferred and are not implemented in the current runtime.
Implemented structured LLM infrastructure
internal/framework/contracts now defines a typed structured-completion contract:
StructuredLLMClient.CompleteStructured(ctx, req, out)- caller-owned typed decode target via
outpointer. - caller-selected structured response schema metadata via
StructuredCompletionRequest.ResponseSchema.
internal/framework/llm provides OpenAICompatibleClient, a direct net/http adapter over OpenAI-compatible chat completions:
- configurable
base_url, model, optional API key, retries, HTTP client, and request timeout; - OpenAI-compatible endpoint behavior (for example OpenAI/OpenRouter/local-compatible base URLs);
- request message translation from
contracts.LLMMessageto chat-completions messages; - strict
response_format.type = json_schemawith registered structured response schemas (strict: true, schema name, and schema body); - response metadata mapping (provider/model/token usage) into Audita-owned response types;
- API-key redaction in adapter-returned errors;
- context cancellation and timeout propagation through request contexts and HTTP client timeouts;
- bounded retry behavior for transient request failures and malformed retryable structured responses.
Structured response schemas are owned by Audita in internal/framework/responseschema and currently include:
- key
correction_set:- id
audita.correction_set - version
v1 - name
audita_correction_set_v1 - sha256
05f8ff3fa04f68115c0cb1859d2656f51aa5c0bae8ff2470b2d4f6f531953195
- id
- key
validator_decision_set:- id
audita.validator_decision_set - version
v1 - name
audita_validator_decision_set_v1 - sha256
b73f4790b98fbb955f0aec5496dd8ce9a8fe14aa2f35c700b4b4e5634f106fd5
- id
Provider-level structured output is treated as a guardrail, not a trust boundary:
- the adapter decodes assistant message content into caller-owned structs;
- proposal-generation and validator layers continue local validation (shape, cardinality, confidence bounds, and proposal-index semantics) before changes can be applied.
Current runtime boundary:
- the default CLI runtime path (without explicit module selection) instantiates the full production module sequence.
- LLM calls are exercised in production in both default full-pipeline runs and explicit
--modulesruns, and in tests when fake/injected clients are used. - normal
go test ./...does not require real LLM credentials or Python dependencies.
internal/framework/llm also provides:
- a bounded FIFO
Schedulerfor controlled concurrent LLM calls with reliable permit release on success, error, and cancellation; - primary/validation effective-config resolution helpers, including validation inheritance fallback to total LLM concurrency settings;
- generic interaction diagnostics primitives that write machine-readable JSON artifacts for request metadata, request payload, response payload, and optional error payload with secret redaction.
Structured LLM diagnostics behavior:
- proposal-generation and validator diagnostics include structured response schema metadata (
id,version,name,sha256) when schema-driven calls are made; - API keys and bearer tokens are redacted from request/response/error diagnostics artifacts and surfaced errors.
Dependency posture:
- the runtime no longer depends on
instructor-go; - structured LLM behavior is implemented through Audita-owned code paths behind
StructuredLLMClient.
LLM concurrency runtime behavior:
totalconcurrency bounds all proposal and validation LLM calls.proposalconcurrency adds a proposal-only sub-cap, composed with total.validationconcurrency adds a validation-only sub-cap, composed with total.- legacy
llm-concurrencyinputs remain compatibility aliases for total concurrency. - modules execute serially, chunk proposals run concurrently within each module, and approved proposals are applied once per module in deterministic order.
Implemented normalization behavior
Normalization (internal/core/normalization) currently:
- sorts by segment start time;
- merges adjacent same-speaker segments when constraints pass;
- uses gap-based joiners:
- gap
< ellipsis_gap-> single space join - gap
>= ellipsis_gap->...join
- gap
- enforces merged duration and token-limit constraints;
- reassigns output IDs sequentially from
1; - returns
NormalizationSummarywith merge and skip counters.
Note: merged categories are concatenated (not deduplicated).
Implemented chunking behavior
Chunking (internal/core/chunking) currently provides:
- deterministic heuristic token estimation;
- contiguous sectioning with section metadata;
- max/min section token validation;
- optional
target_sectionsoverride for section-count planning; - summary and detailed summary generation.
Current behavior details:
- if a single segment exceeds max tokens, it is emitted as its own section (not hard-failed);
- default section count is planned from
ceil(total_tokens / max_section_tokens); - section sizing targets
ceil(total_tokens / section_count)with a deterministic forward pass; - sections remain contiguous and ordered, and segments are never split.
Implemented proposal/replacement infrastructure
internal/framework/proposals provides deterministic proposal composition logic:
CorrectionProposalandEnrichedCorrectionProposalmodels;- replacement policies:
require_unique,replace_all; - safe preview (
PreviewProposalForSegment) with stable skip reasons; - deterministic apply (
ApplyProposals) in ascendingproposal_indexorder; - applied/skipped change records suitable for reporting.
internal/framework/contracts provides interfaces and run-spec metadata scaffolding, including deterministic repeated module instance naming (ResolveModuleRunSpecs).
These primitives are wired into the production runner and report model. The grammar, glossary, homophones, and spoken_word modules are implemented.
Implemented validator runtime infrastructure
internal/framework/validators provides deterministic validator infrastructure:
- runtime validation request/result models;
- stable validator reason codes;
- cardinality enforcement for validator decisions:
- missing proposal indexes fail
- duplicate proposal indexes fail
- unknown proposal indexes fail
- deterministic validators:
- confidence threshold by module key/config threshold
- original-text presence against current working transcript
- non-empty corrected text
- identical/no-effect rejection
- conservative protected glossary-term guard for non-glossary modules
internal/framework/runner executes module pipelines with deterministic boundaries:
- modules still execute serially over the working transcript;
- section proposal work is launched promptly and can run concurrently;
- section-level validator-chain work starts as section proposals become available (deterministic validators before LLM-backed validators);
- proposal-generation and LLM-validator calls can overlap under composed scheduler limits;
- approved proposals are still applied once per module after section work settles.
Validator rejections are reported distinctly from proposal-application skips.
Implemented LLM-backed validator infrastructure
internal/framework/validators now includes LLM-backed validator support:
- typed request/response models for structured LLM validation;
- prompt builders for:
- spoken-form plausibility
- meaning reversal detection
- editorial review
- grammar review
- spoken-word review
- deterministic batching by
validation_max_prompt_tokens; - strict cardinality validation of structured LLM decisions (missing/duplicate/unknown indexes fail);
- safe failure behavior for malformed/invalid structured responses.
internal/framework/runner wires LLM validators into existing validator chains using:
- the internal structured LLM client abstraction (
contracts.StructuredLLMClient); - bounded scheduler hooks for validator call execution;
- diagnostics writer hooks for machine-readable prompt/response artifacts with secret redaction.
Implemented shared proposal-generation infrastructure
internal/framework/proposal_generation provides a reusable, prompt-agnostic helper for future real modules:
- structured request model including module key/instance, replacement policy, working transcript context, optional section metadata, glossary, config, and diagnostics context;
- structured correction-set response model (
corrections) mapped into existingproposals.CorrectionProposalandproposals.EnrichedCorrectionProposalmodels; - deterministic proposal-index assignment through a caller-provided
start_index; - structured LLM calls through
contracts.StructuredLLMClientonly (no direct provider calls); - optional bounded execution through scheduler hooks (
contracts.LLMScheduler); - prompt/response diagnostics artifact writing via the generic
internal/framework/llmdiagnostics primitives with redaction of API keys/secrets.
This helper only produces candidate proposals; validator-chain execution and proposal application remain runner responsibilities.
Implemented production module-registry scaffolding
internal/framework/modules now provides a production registry scaffold:
- recognizes intended module keys:
glossaryhomophonesspoken_wordgrammar
- supports explicit constructor registration with dependency injection for:
- run spec
- config
- glossary
- proposal/validation structured LLM clients
- proposal/validation schedulers
- diagnostics directory context
- returns explicit errors for unknown keys (
unsupported_module).
The grammar, glossary, homophones, and spoken_word module keys are now registered and constructible.
Implemented grammar production module
internal/modules/grammar now provides the first production module:
- prompt builder faithfully constrained to punctuation/capitalization/spacing/article cleanup;
- explicit guardrails against meaning-changing rewrites, style rewrites, summarization, and invention;
- proposal generation through
internal/framework/proposal_generationandcontracts.StructuredLLMClient; - scheduler-aware proposal calls through existing
contracts.LLMSchedulerhooks; - replacement policy
require_unique(current runtime policy); - validator chain integration using existing deterministic + LLM-backed validators;
- grammar confidence threshold enforcement through existing validator/config infrastructure;
- module-level reporting and diagnostics capture through existing runner/reporting paths.
Implemented glossary production module
internal/modules/glossary now provides the second production module:
- prompt builder aligned to Python glossary-module intent, constrained to glossary-backed domain/acoustic corrections;
- prompt context includes glossary names, aliases, categories, summaries, and plural forms where available;
- guardrails against broad style rewriting and against replacing unrelated terms simply because they appear in glossary entries;
- proposal generation through
internal/framework/proposal_generationandcontracts.StructuredLLMClient; - scheduler-aware proposal calls through existing
contracts.LLMSchedulerhooks; - replacement policy
replace_all(matching Python glossary behavior); - validator chain integration using existing deterministic + LLM-backed validators;
- glossary confidence threshold enforcement through existing validator/config infrastructure;
- module-level reporting and diagnostics capture through existing runner/reporting paths;
- explicit support for repeated glossary stages with deterministic instance names (
glossary_1,glossary_2, ...), where later stages see prior-stage working transcript changes.
Implemented protected-term behavior
internal/framework/validators/protected_terms.go provides deterministic glossary-derived protected vocabulary:
- extracts protected terms from glossary names and aliases;
- includes explicit plural fields and synthetic plural forms where safe;
- deduplicates and returns stable ordering for repeatable behavior/tests.
This vocabulary is used by deterministic validators for both glossary-stage and non-glossary-stage protection checks, keeping protected-term guardrails active across modules.
Implemented homophones production module
internal/modules/homophones now provides the third production module:
- prompt builder aligned to Python homophones-module intent, constrained to conservative homophone/near-homophone/mistranscription corrections;
- prompt context includes protected glossary names/aliases/plurals to avoid damaging known terms;
- explicit guardrails against punctuation cleanup, grammar cleanup, style rewriting, summarization, and content invention;
- proposal generation through
internal/framework/proposal_generationandcontracts.StructuredLLMClient; - scheduler-aware proposal calls through existing
contracts.LLMSchedulerhooks; - replacement policy
require_unique(matching Python homophones behavior); - validator chain integration using existing deterministic + LLM-backed validators;
- homophones confidence threshold enforcement through existing validator/config infrastructure;
- protected-term guardrails for non-glossary modules remain active and are exercised through the homophones path;
- module-level reporting and diagnostics capture through existing runner/reporting paths.
Implemented spoken_word production module
internal/modules/spoken_word now provides the fourth production module:
- prompt builder aligned to Python spoken_word-module intent, constrained to conservative dysfluency cleanup;
- strong prompt guardrails preserving meaning/intent/voice/named entities/domain terms and substantive content;
- explicit guardrails against summarization, style rewriting, grammar-only cleanup, punctuation-only cleanup, invention, and meaning-changing rewrites;
- proposal generation through
internal/framework/proposal_generationandcontracts.StructuredLLMClient; - scheduler-aware proposal calls through existing
contracts.LLMSchedulerhooks; - replacement policy
require_unique(matching Python spoken_word behavior); - validator chain integration using existing deterministic + LLM-backed validators, including strong semantic guardrails (
spoken_word_review,meaning_reversal_review); - spoken_word confidence threshold enforcement through existing validator/config infrastructure;
- protected-term guardrails for non-glossary modules remain active and are exercised through the spoken_word path;
- module-level reporting and diagnostics capture through existing runner/reporting paths.
Reports and diagnostics (implemented)
Current per-run artifacts include:
source-transcript.jsonsource-transcript-parsed.jsonnormalized-transcript.jsonnormalization-summary.jsonchunking-summary.jsoninvocation.jsoneffective-config.json(redacted credentials)report.jsonerror.logon failure
--report-json writes a separate report file when requested.
Current process reports include diagnostics metadata references for:
- diagnostics directory path;
- source transcript artifact path;
- parsed source transcript artifact path;
- normalized transcript artifact path;
- normalization summary artifact path;
- chunking summary artifact path;
- invocation metadata artifact path;
- redacted effective-config artifact path;
- error-log artifact path on failure.
Current process reports also include:
- module-level results (when runner modules execute), including applied/skipped proposal changes;
- run-level module summary totals and failed module instance metadata.
- module-level validator decisions and validator rejections.
- optional decision-level diagnostic artifact paths for validator LLM interactions when available.
Retention modes implemented in ApplyRetention:
always: keep all run directories.never: keep successful run directories.auto: keep failed runs and successful runs with skipped corrections.- failed runs are always retained.
Current runtime note:
- default non-explicit runs usually have no module-level skipped corrections, so
autocommonly removes clean successful run directories. - explicit grammar/glossary/homophones/spoken_word runs can produce validator rejections and application skips, which are reflected in reports and retention input.
Current tests and quality posture
Implemented tests currently cover:
- CLI argument handling and behavior (
internal/cli/run_test.go) - subprocess stdout/stderr and exit-code behavior (
cmd/audita/main_integration_test.go) - config/env/override validation (
internal/core/config/*_test.go) - transcript and glossary schema validation (
internal/core/schema/*_test.go) - deterministic normalization (
internal/core/normalization/*_test.go) - deterministic chunking and summaries (
internal/core/chunking/*_test.go) - proposal preview/apply semantics (
internal/framework/proposals/*_test.go) - contracts/foundation composition tests (
internal/framework/contracts/*_test.go) - runner sequencing and failure behavior with deterministic fake modules (
internal/framework/runner/*_test.go) - CLI runner integration through injected fake module factories (
internal/cli/run_test.go) - validator models, cardinality enforcement, and deterministic validators (
internal/framework/validators/*_test.go) - LLM-backed validator batching, prompt builders, structured-response safety, scheduler hooks, and diagnostics redaction (
internal/framework/validators/*_test.go,internal/framework/runner/*_test.go) - shared proposal-generation request/response parsing, deterministic indexing, scheduler hooks, and diagnostics redaction (
internal/framework/proposal_generation/*_test.go,internal/framework/runner/*_test.go) - production module-registry known-key recognition and unsupported/internal-registry error behavior (
internal/framework/modules/*_test.go,internal/cli/run_test.go) - production grammar module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, and explicit CLI/runtime integration (
internal/modules/grammar/*_test.go,internal/cli/run_test.go,internal/framework/runner/*_test.go) - production glossary module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, repeated-stage behavior, and explicit CLI/runtime integration (
internal/modules/glossary/*_test.go,internal/cli/run_test.go,internal/framework/runner/*_test.go) - production homophones module prompt constraints, proposal mapping, validator-chain behavior, confidence-threshold enforcement, diagnostics redaction, protected-term behavior, and explicit CLI/runtime integration (
internal/modules/homophones/*_test.go,internal/cli/run_test.go,internal/framework/runner/*_test.go) - production spoken_word module prompt constraints, proposal mapping, validator-chain behavior, semantic guardrail behavior, confidence-threshold enforcement, diagnostics redaction, protected-term behavior, and explicit CLI/runtime integration (
internal/modules/spoken_word/*_test.go,internal/cli/run_test.go,internal/framework/runner/*_test.go) - glossary-derived protected-term extraction and stable behavior (
internal/framework/validators/protected_terms_test.go) - default full-pipeline runtime shape and ordering (
internal/cli/run_test.go,cmd/audita/main_integration_test.go,internal/cli/parity_test.go) - subprocess operational hardening behavior including large-input, failure-mode, timeout/cancellation, backend-failure, and partial-progress paths (
cmd/audita/main_integration_test.go) - report/diagnostics redaction and artifact-shape behavior across success and failure paths (
internal/cli/run_test.go,cmd/audita/main_integration_test.go)
Operational hardening status
The runtime now includes hardened subprocess behavior for parent-process callers:
- deterministic success/failure exit codes;
- strict stdout/stderr separation suitable for machine orchestration;
- failure stderr summaries that include diagnostics location when available;
- retained failure diagnostics (
report.json,error.log, and artifacts written before failure); - deterministic timeout/cancellation behavior in tests;
- redaction coverage for API keys/secrets across reports, diagnostics artifacts, and surfaced errors.
Operational caller guidance is documented in docs/subprocess-operations.md.
Final status
- Audita's default full module-sequence runtime is implemented and tested.
- Parity fixtures and operational hardening coverage are in place.
- Historical migration context is documented in
docs/migration-from-python.md.