Audita
Audita is now a framework-first transcript correction application. The public audita package provides:
- deterministic transcript normalization
- token-batched module orchestration
- reusable module / filter / review-stage contracts
- structured run reporting and work-dir diagnostics
The previous working implementation has been preserved as audita_prototype inside this repository. Its full regression suite lives under tests/audita_prototype.
Development
This project is set up for uv.
uv sync --extra dev
uv run pytest
Usage
Process a transcript with the new framework skeleton:
uv run audita process transcript.json --glossary glossary.yaml --output corrected.json
The framework currently runs this default module sequence:
glossary_primaryhomophonesglossary_secondaryspoken_wordgrammar
At this stage the module implementations are stubs. The framework is executable end to end, performs deterministic normalization, runs the full module lifecycle, and emits structured reports, but does not yet apply substantive LLM-driven corrections.
To also write a structured JSON report:
uv run audita process transcript.json --glossary glossary.yaml --output corrected.json --report-json report.json
From a checked-out repository, you can also use the root launcher:
./audita process transcript.json --glossary glossary.yaml --output corrected.json
For a system-wide command, install the source tree under /usr/local/src/audita, sync dependencies there, and symlink the root launcher into your PATH:
cd /usr/local/src/audita
uv sync --extra dev
ln -s /usr/local/src/audita/audita /usr/local/bin/audita
audita process transcript.json --glossary glossary.yaml --output corrected.json
Without --output, Audita writes the corrected transcript JSON to stdout and progress logs to stderr.
--report-json writes a separate machine-readable run report and never mixes report data into stdout.
Useful configuration can be supplied by CLI flag or environment variable. CLI flags take precedence over environment variables. The framework does not currently require OPENROUTER_API_KEY, but the config fields remain available for future module implementations.
| Environment variable | CLI flag | Default | Purpose |
|---|---|---|---|
AUDITA_MODEL |
--model |
openrouter/google/gemma-4-31b-it |
Reserved LLM model setting for future module implementations |
AUDITA_BASE_URL |
--base-url |
https://openrouter.ai/api/v1 |
Reserved OpenAI-compatible API base URL |
AUDITA_MAX_RETRIES |
--max-retries |
3 |
Maximum Instructor retries for structured responses |
AUDITA_MAX_SECTION_TOKENS |
--max-section-tokens |
6144 |
Maximum estimated tokens per transcript batch |
AUDITA_NORMALIZE_MAX_SEGMENT_GAP |
--normalize-max-segment-gap |
4.0 |
Same-speaker gaps eligible for deterministic merging |
AUDITA_NORMALIZE_ELLIPSIS_GAP |
--normalize-ellipsis-gap |
3.5 |
Same-speaker gaps above this value are joined with ... |
AUDITA_NORMALIZE_MAX_SEGMENT_DURATION |
--normalize-max-segment-duration |
60.0 |
Maximum merged segment duration |
AUDITA_NORMALIZE_MAX_SEGMENT_TOKENS |
--normalize-max-segment-tokens |
2048 |
Maximum merged segment prompt payload size |
AUDITA_WORK_DIR |
--work-dir |
/tmp/audita |
Per-run scratch diagnostics directory |
AUDITA_WORK_DIR_RETENTION |
--work-dir-retention |
auto |
Whether to retain the per-run work directory: auto, always, or never |
AUDITA_WORK_DIR stores per-run diagnostics while processing. Under the default AUDITA_WORK_DIR_RETENTION=auto, clean successful runs are removed, while failed runs and successful runs with final skipped corrections are preserved. Use always to keep every run directory and never to remove successful run directories even when skips remain.
Prototype Archive
The archived prototype remains importable as audita_prototype and is still covered by its original regression suite. This is intentional: the new audita package is a fresh framework skeleton, not a thin wrapper around the old code.