Audita

Audita takes raw audio transcripts, deterministically merges short same-speaker segments into speaking turns, and uses an LLM to identify and fix misheard words, jargon, domain-specific terms, and conservative readability issues.

Development

This project is set up for uv.

uv sync --extra dev
uv run pytest

Usage

Set an OpenRouter API key, then process a transcript with a glossary:

export OPENROUTER_API_KEY=...
uv run audita process transcript.json --glossary glossary.yaml --output corrected.json

From a checked-out repository, you can also use the root launcher:

./audita process transcript.json --glossary glossary.yaml --output corrected.json

For a system-wide command, install the source tree under /usr/local/src/audita, sync dependencies there, and symlink the root launcher into your PATH:

cd /usr/local/src/audita
uv sync --extra dev
ln -s /usr/local/src/audita/audita /usr/local/bin/audita
audita process transcript.json --glossary glossary.yaml --output corrected.json

Without --output, Audita writes the corrected transcript JSON to stdout and progress logs to stderr.

Useful configuration can be supplied by CLI flag or environment variable. CLI flags take precedence over environment variables.

Environment variable CLI flag Default Purpose
AUDITA_MODEL --model openrouter/google/gemma-4-31b-it OpenRouter model to use
AUDITA_BASE_URL --base-url https://openrouter.ai/api/v1 OpenAI-compatible API base URL
AUDITA_MAX_SECTION_TOKENS --max-section-tokens 6144 Maximum estimated tokens per LLM transcript section
AUDITA_GLOSSARY_CONFIDENCE_THRESHOLD --glossary-confidence-threshold 0.80 Minimum confidence required to apply a glossary correction
AUDITA_GRAMMAR_CONFIDENCE_THRESHOLD --grammar-confidence-threshold 0.80 Minimum confidence required to apply a grammar correction
AUDITA_GRAMMAR_VALIDATION_ENABLED --grammar-validation-enabled / --no-grammar-validation-enabled true Whether grammar corrections are checked by the semantic validator
AUDITA_GRAMMAR_VALIDATION_CONFIDENCE_THRESHOLD --grammar-validation-confidence-threshold 0.80 Minimum validator confidence required for validated grammar corrections
AUDITA_MAX_RETRIES --max-retries 3 Maximum Instructor retries for structured responses
AUDITA_GLOSSARY_MAX_LLM_PASSES --glossary-max-llm-passes 3 Total glossary correction passes
AUDITA_GRAMMAR_MAX_LLM_PASSES --grammar-max-llm-passes 3 Total grammar/readability correction passes
AUDITA_NORMALIZE_MAX_SEGMENT_GAP --normalize-max-segment-gap 5.0 Same-speaker gaps eligible for deterministic merging
AUDITA_NORMALIZE_ELLIPSIS_GAP --normalize-ellipsis-gap 3.5 Same-speaker gaps above this value are joined with ...
AUDITA_NORMALIZE_MAX_SEGMENT_DURATION --normalize-max-segment-duration 60.0 Maximum merged segment duration
AUDITA_NORMALIZE_MAX_SEGMENT_TOKENS --normalize-max-segment-tokens 2048 Maximum merged segment prompt payload size
AUDITA_WORK_DIR --work-dir /tmp/audita Per-run scratch diagnostics directory

OPENROUTER_API_KEY is required and is read from the environment.

AUDITA_WORK_DIR stores per-run diagnostics while processing. Successful runs clean up their run directory unless corrections are skipped; failed runs and skipped-correction runs preserve diagnostics for debugging.

Description
Audita takes raw audio transcripts and uses an LLM to identify and fix misheard words, jargon, and domain-specific terms.
Readme BSD-3-Clause 3.4 MiB
v1.0.0 Latest
2026-05-24 11:54:30 +00:00
Languages
Go 100%