103 lines
4.9 KiB
Markdown
103 lines
4.9 KiB
Markdown
# narratio
|
||
|
||
`narratio` is a Go-based orchestration application for processing D&D session audio into transcripts and downstream artifacts.
|
||
|
||
This repository currently contains a **working scaffold** with strict config loading, local workdir/manifest handling, resumable stage control, real WhisperX and Seriatim adapters, and real `prepare` + `transcribe` + `merge` stages.
|
||
|
||
## Expected Config Files
|
||
|
||
`narratio` expects two YAML files:
|
||
|
||
- `pipeline.yml`: pipeline/workspace settings (`workspace`, `storage`, `whisperx`, `seriatim`, `audita`, `analyzer`, `notification`)
|
||
- `session.yml`: per-session settings (`session_id`, `inputs`, optional metadata)
|
||
|
||
WhisperX config contract in `pipeline.yml`:
|
||
|
||
- required: `whisperx.transcribe_url`
|
||
- defaulted when omitted: `whisperx.language` (`en`), `whisperx.timeout` (`30m`), `whisperx.retries` (`3`), `whisperx.retry_delay` (`2s`), `whisperx.concurrency` (`2`)
|
||
|
||
Seriatim config contract in `pipeline.yml`:
|
||
|
||
- required: `seriatim.binary` (name or path; existence is checked at execution time, not config validation time)
|
||
- defaulted when omitted: `seriatim.timeout` (`10m`), `seriatim.output_schema` (`seriatim-intermediate`), `seriatim.coalesce_gap` (`3.0`), `seriatim.report` (`true`)
|
||
- allowed `seriatim.output_schema` values: `seriatim-minimal`, `seriatim-intermediate`, `seriatim-full`
|
||
- optional tuning: `seriatim.env.*` (`overlap_word_run_gap`, `overlap_word_run_reorder_window`, `backchannel_max_duration`, `filler_max_duration`) must be `> 0` when provided
|
||
|
||
Audita config contract in `pipeline.yml`:
|
||
|
||
- required: `audita.binary` (name or path; existence is checked at execution time, not config validation time)
|
||
- required: `audita.llm_api_key_env` (environment variable name holding the API key secret)
|
||
- defaulted when omitted: `audita.timeout` (`3h`), `audita.llm_api_key_env` (`AUDITA_LLM_API_KEY`), `audita.modules` (`glossary,homophones,glossary,spoken_word,grammar,homophones,glossary`), `audita.base_url` (`https://openrouter.ai/api/v1`), `audita.model` (`openrouter/google/gemma-4-31b-it`), `audita.llm_concurrency` (`1`), `audita.validation_model` (`""`), `audita.validation_llm_concurrency` (`1`), `audita.report` (`true`)
|
||
- allowed `audita.modules` values: `glossary`, `homophones`, `spoken_word`, `grammar` (order and repeats are allowed)
|
||
- `audita.base_url` must be a valid URL when provided
|
||
- `audita.llm_concurrency` and `audita.validation_llm_concurrency` must be `> 0`
|
||
|
||
Audita credentials note:
|
||
|
||
- store only the environment variable **name** in config (`audita.llm_api_key_env`), never the API key value itself
|
||
- API key values must not be written to pipeline config, generated configs, logs, or manifest metadata
|
||
|
||
Audita CLI compatibility note:
|
||
|
||
- Narratio models `audita.llm_concurrency` in config now, but real adapter wiring should verify whether Audita expects a direct CLI flag or env-based configuration for primary LLM concurrency before implementation.
|
||
|
||
`speakers.yml` note:
|
||
|
||
- use Seriatim’s documented `match:` format (not the legacy direct mapping style used by older scripts/scaffolds)
|
||
- see [`examples/speakers.yml`](examples/speakers.yml) for a concrete `match:` example.
|
||
|
||
Prepared inputs note:
|
||
|
||
- `prepare` copies session inputs into the session workdir (`work/<session_id>/inputs`), and downstream stages consume those prepared copies, not the original source files.
|
||
- If you edit `speakers.yml`, `autocorrect.yml`, `glossary.yml`, or session input paths after `prepare` has run, rerun `prepare` with `--force` before rerunning downstream stages.
|
||
|
||
Decoding is strict (`KnownFields(true)`), so unknown YAML fields fail fast.
|
||
|
||
Example minimal files are available under `examples/`:
|
||
|
||
- `examples/pipeline.minimal.yml`
|
||
- `examples/session.minimal.yml`
|
||
|
||
## Current Scaffold Status
|
||
|
||
Implemented now:
|
||
|
||
- strict config load + validation
|
||
- local artifact/workdir creation and locking
|
||
- manifest create/load/save and stage status tracking
|
||
- stage framework with real `prepare`, `transcribe`, and `merge` stages; placeholder downstream stages
|
||
- resumable run control (`run`, `resume`, `run-stage`, `plan` with run/skip decisions)
|
||
- real WhisperX HTTP adapter plus real Seriatim subprocess adapter (both with fake/no-op adapters for test/scaffold usage)
|
||
|
||
Not implemented yet:
|
||
|
||
- real Audita execution
|
||
- real analyzer integration
|
||
- real remote archive/storage backend
|
||
- real notification backend
|
||
|
||
## Run Tests
|
||
|
||
```bash
|
||
go test ./...
|
||
```
|
||
|
||
## Run Plan
|
||
|
||
```bash
|
||
go run ./cmd/narratio plan --config examples/pipeline.minimal.yml --session examples/session.minimal.yml
|
||
```
|
||
|
||
## Run Pipeline (Current State)
|
||
|
||
The current `run` command executes `prepare` + real `transcribe` + real `merge` + placeholder downstream stages and records progress in `manifest.json`.
|
||
|
||
Default CLI wiring builds and uses:
|
||
|
||
- real WhisperX HTTP adapter from `pipeline.whisperx`
|
||
- real Seriatim subprocess adapter from `pipeline.seriatim`
|
||
|
||
```bash
|
||
go run ./cmd/narratio run --config examples/pipeline.minimal.yml --session examples/session.minimal.yml
|
||
```
|