3 Commits

Author SHA1 Message Date
7995c41675 Update the audita integration documentation reference
All checks were successful
ci/woodpecker/tag/release Pipeline was successful
2026-05-16 08:02:18 -05:00
62551d43a0 Added filesystem-based secrets loading configuration 2026-05-16 07:56:22 -05:00
8395c12dd3 Update documentation to include an implementation roadmap for the archive stage 2026-05-16 07:36:25 -05:00
14 changed files with 1580 additions and 115 deletions

View File

@@ -35,6 +35,16 @@ Pipeline config lookup for CLI commands:
- `/usr/local/etc/narratio/pipeline.yml`
- `/etc/narratio/pipeline.yml`
Optional secrets-from-files config:
- `pipeline.secrets.env_dir` may point to a directory of secret files
- each top-level file with an env-var-style name is loaded as an environment variable:
- file name = env var name
- file contents = env var value (trailing newline/CRLF trimmed)
- process environment wins: existing env vars are not overwritten
- if configured, Narratio fails fast when `env_dir` is missing/unreadable
- relative `env_dir` values resolve from Narratios current working directory
YAML decoding is strict (`KnownFields(true)`), so unknown fields fail fast.
## Canonical Stage Order
@@ -199,6 +209,8 @@ Prompt IDs and profile IDs are configuration values. They are not hardcoded in a
Do not put secrets in `pipeline.yml`. If API-key behavior is configured, use env var names only.
If `pipeline.secrets.env_dir` is configured, keep only references and secret files there; secret values are still not written to manifests, generated configs, or Narratio-managed logs.
## Scriptorium Runtime Behavior
Narratio integrates with Scriptorium through the public CLI subprocess contract:

View File

@@ -101,6 +101,15 @@ CLI pipeline config path resolution:
- `/usr/local/etc/narratio/pipeline.yml`
- `/etc/narratio/pipeline.yml`
Optional pipeline secrets directory:
- `pipeline.secrets.env_dir` enables loading environment variables from local files before command execution
- file name = env var name; file contents = env var value (trailing newline/CRLF trimmed)
- only env-var-style file names are considered; other entries are ignored
- existing process environment values are preserved (not overwritten)
- if configured, unreadable/missing `env_dir` fails command execution early
- relative `env_dir` values are resolved from current working directory
`pipeline.scriptorium` is optional. Existing pipelines without Scriptorium continue to work.
`pipeline.trim` is optional. Existing pipelines without trim config continue to work.
@@ -288,6 +297,7 @@ Current expected paths for `session_recap`:
- do not store secrets in pipeline YAML, generated invocation YAML, logs, or manifest metadata
- if API-key integration is configured, pass env var names only (never raw key values)
- with `pipeline.secrets.env_dir`, secret file values are loaded into process env only and are not persisted in manifest metadata or generated configs
- avoid logging transcript content or rendered prompt content by default
- treat generated artifacts and logs as potentially sensitive session material

View File

@@ -1,147 +1,96 @@
# Audita
# Audita Subprocess Operations
Audita is a framework-first transcript correction application. The public `audita` package provides:
This document describes how parent processes should invoke `audita process` safely in production orchestration.
- deterministic transcript normalization
- token-batched module orchestration
- concrete `glossary`, `homophones`, `spoken_word`, and `grammar` modules built on reusable proposal / validator contracts
- structured run reporting and work-dir diagnostics
## Recommended command form
The previous working implementation has been preserved as `audita_prototype` inside this repository. Its full regression suite lives under `tests/audita_prototype`.
## Development
This project is set up for `uv`.
Use explicit file outputs for orchestrated runs:
```sh
uv sync --extra dev
uv run pytest
audita process <transcript.json> \
--transcript-description "Brief context that may help resolve ambiguous terms." \
--glossary <glossary.yaml> \
--output <output-transcript.json> \
--report-json <report.json>
```
## Usage
Additional flags that may be situationally appropriate:
- `--config <path>` to select an explicit versioned config file.
- `--output-schema <bare-segments|audita-v1>` to select transcript output shape.
- `--work-dir <dir>` to control diagnostics location.
- `--work-dir-retention <always|auto|never>` to control retained run directories.
- `--total-llm-concurrency`, `--proposal-llm-concurrency`, and `--validation-llm-concurrency` when orchestration needs to set explicit LLM throughput controls.
- `--modules ...` only when intentionally overriding the default sequence.
Process a transcript with the current framework implementation:
For config-driven orchestration, validate config files in CI/preflight:
```sh
uv run audita process transcript.json --glossary glossary.yaml --output corrected.json
audita config validate --config <path>
```
The framework currently runs this default module sequence:
## Stdout behavior
1. `glossary`
2. `homophones`
3. `glossary`
4. `spoken_word`
5. `grammar`
- With `--output`: stdout is expected to be empty on success.
- Without `--output`: stdout contains transcript JSON only on success.
- Report JSON is never written to stdout.
Resolved run instance names are auto-numbered for repeats, so the default report pipeline is:
## Stderr behavior
1. `glossary_1`
2. `homophones`
3. `glossary_2`
4. `spoken_word`
5. `grammar`
- Success path should be quiet or minimal human-readable logs.
- Failure path writes concise human-readable errors.
- When a diagnostics run directory exists, failure stderr includes its path.
- Prompt/response diagnostic payloads are not streamed to stderr.
The default module sequence is fully implemented today:
## Output file behavior
- `glossary` proposes glossary-supported acoustic corrections
- `homophones` proposes conservative homophone and mistranscription corrections
- `spoken_word` proposes conservative dysfluency cleanup
- `grammar` proposes punctuation, capitalization, and spacing cleanup only
- `--output` writes transcript JSON in the selected output schema to the provided path.
- Output write failures return nonzero and surface actionable errors.
- The command does not silently ignore output write errors.
To run a custom module sequence, pass `--modules`:
## Report JSON behavior
```sh
uv run audita process transcript.json --glossary glossary.yaml --modules grammar --output corrected.json
```
- `--report-json` writes a machine-readable process report to the requested path.
- Run-directory `report.json` is written independently under diagnostics.
- Best-effort failure reports are emitted when possible without masking the primary failure.
- Report write failures return nonzero with clear stderr messaging.
- Report diagnostics metadata references run-directory artifacts including utilization diagnostics and correction ledger paths when available.
To also write a structured JSON report:
## Diagnostics directory behavior
```sh
uv run audita process transcript.json --glossary glossary.yaml --output corrected.json --report-json report.json
```
- Each run creates (when possible) a per-run diagnostics directory.
- Typical artifacts include transcript, normalization, chunking, invocation, effective config, LLM diagnostics, `utilization-diagnostics.json`, `correction-ledger.json`, `report.json`, and `error.log` on failure.
- Failed runs retain diagnostics.
- Under `auto` retention, successful runs with skipped/rejected corrections are retained; clean successful runs may be removed.
From a checked-out repository, you can also use the root launcher:
## Exit codes
```sh
./audita process transcript.json --glossary glossary.yaml --output corrected.json
```
- `0`: success.
- Nonzero: failure (input/schema/config/module/LLM/runtime/output/report/diagnostics errors).
For a system-wide command, install the source tree under `/usr/local/src/audita`, sync dependencies there, and symlink the root launcher into your `PATH`:
Treat any nonzero as a failed subprocess invocation.
```sh
cd /usr/local/src/audita
uv sync --extra dev
ln -s /usr/local/src/audita/audita /usr/local/bin/audita
audita process transcript.json --glossary glossary.yaml --output corrected.json
```
## Timeout and cancellation
Without `--output`, Audita writes the corrected transcript JSON to stdout and progress logs to stderr.
`--report-json` writes a separate machine-readable run report and never mixes report data into stdout.
- Runtime operations propagate context cancellation and request timeouts through LLM/scheduler paths.
- On cancellation or timeout, the process exits nonzero and should not hang.
- If diagnostics were initialized before failure, failure artifacts remain available for debugging.
Useful configuration can be supplied by CLI flag or environment variable. CLI flags take precedence over environment variables. Normal runs now require LLM API credentials, because the `glossary`, `homophones`, `spoken_word`, and `grammar` modules make real LLM calls. `AUDITA_LLM_API_KEY` and `--llm-api-key` are the preferred provider-neutral credential surfaces, while `OPENROUTER_API_KEY` remains supported as a backward-compatible fallback.
## Secret redaction expectations
| Environment variable | CLI flag | Default | Purpose |
| --- | --- | --- | --- |
| `AUDITA_MODULES` | `--modules` | `glossary,homophones,glossary,spoken_word,grammar` | Comma-separated logical module keys to run; CLI overrides the environment value |
| `AUDITA_LLM_API_KEY` | `--llm-api-key` | unset | Preferred provider-neutral LLM API credential; CLI overrides both environment-key variants |
| `AUDITA_VALIDATION_LLM_API_KEY` | `--validation-llm-api-key` | unset | Validation-phase LLM API credential; defaults to the primary LLM API key |
| `AUDITA_MODEL` | `--model` | `openrouter/google/gemma-4-31b-it` | LLM model name sent to the configured OpenAI-compatible endpoint |
| `AUDITA_VALIDATION_MODEL` | `--validation-model` | unset | Validation-phase LLM model; defaults to `AUDITA_MODEL` |
| `AUDITA_BASE_URL` | `--base-url` | `https://openrouter.ai/api/v1` | OpenAI-compatible API base URL |
| `AUDITA_VALIDATION_BASE_URL` | `--validation-base-url` | unset | Validation-phase OpenAI-compatible API base URL; defaults to `AUDITA_BASE_URL` |
| `AUDITA_LLM_TIMEOUT_SECONDS` | `--llm-timeout-seconds` | `600` | Per-request timeout in seconds for LLM calls to the configured OpenAI-compatible endpoint |
| `AUDITA_VALIDATION_LLM_TIMEOUT_SECONDS` | `--validation-llm-timeout-seconds` | unset | Validation-phase per-request timeout in seconds; defaults to `AUDITA_LLM_TIMEOUT_SECONDS` |
| `AUDITA_VALIDATION_MAX_PROMPT_TOKENS` | `--validation-max-prompt-tokens` | `2048` | Maximum estimated tokens per validation-phase LLM prompt batch |
| `AUDITA_TARGET_SECTIONS` | `--target-sections` | unset | Exact number of contiguous proposal-stage transcript sections; errors if min/max token bounds cannot be satisfied |
| `AUDITA_MAX_RETRIES` | `--max-retries` | `3` | Maximum Instructor retries for structured responses |
| `AUDITA_VALIDATION_MAX_RETRIES` | `--validation-max-retries` | unset | Validation-phase structured-output retries; defaults to `AUDITA_MAX_RETRIES` |
| `AUDITA_VALIDATION_LLM_CONCURRENCY` | `--validation-llm-concurrency` | unset | Validation-phase LLM concurrency; defaults to `AUDITA_LLM_CONCURRENCY` |
| `AUDITA_MAX_SECTION_TOKENS` | `--max-section-tokens` | `8192` | Maximum estimated tokens per proposal-stage transcript section |
| `AUDITA_MIN_SECTION_TOKENS` | `--min-section-tokens` | `2048` | Minimum estimated tokens per proposal-stage transcript section when balancing for concurrency |
| `AUDITA_GLOSSARY_CONFIDENCE_THRESHOLD` | `--glossary-confidence-threshold` | `0.8` | Minimum confidence required for glossary proposals to survive validation |
| `AUDITA_GRAMMAR_CONFIDENCE_THRESHOLD` | `--grammar-confidence-threshold` | `0.8` | Minimum confidence required for grammar proposals to survive validation |
| `AUDITA_HOMOPHONES_CONFIDENCE_THRESHOLD` | `--homophones-confidence-threshold` | `0.8` | Minimum confidence required for homophone proposals to survive validation |
| `AUDITA_SPOKEN_WORD_CONFIDENCE_THRESHOLD` | `--spoken-word-confidence-threshold` | `0.8` | Minimum confidence required for spoken-word proposals to survive validation |
| `AUDITA_NORMALIZE_MAX_SEGMENT_GAP` | `--normalize-max-segment-gap` | `4.0` | Same-speaker gaps eligible for deterministic merging |
| `AUDITA_NORMALIZE_ELLIPSIS_GAP` | `--normalize-ellipsis-gap` | `3.5` | Same-speaker gaps above this value are joined with ` ... ` |
| `AUDITA_NORMALIZE_MAX_SEGMENT_DURATION` | `--normalize-max-segment-duration` | `60.0` | Maximum merged segment duration |
| `AUDITA_NORMALIZE_MAX_SEGMENT_TOKENS` | `--normalize-max-segment-tokens` | `2048` | Maximum merged segment prompt payload size |
| `AUDITA_WORK_DIR` | `--work-dir` | `/tmp/audita` | Per-run scratch diagnostics directory |
| `AUDITA_WORK_DIR_RETENTION` | `--work-dir-retention` | `auto` | Whether to retain the per-run work directory: `auto`, `always`, or `never` |
API keys and configured secret values are redacted from:
- reports (`--report-json` and run-dir `report.json`);
- diagnostics artifacts (including effective config and LLM interaction artifacts);
- surfaced adapter/runtime errors;
- test fixtures and regression outputs.
Set `AUDITA_MODULES=grammar` to run only the grammar module by default, or override it per command with `--modules`.
Parent-process logs should still avoid printing raw environment variables.
Validation-phase LLM settings inherit from the primary `AUDITA_*` LLM settings by default. Set any of the `AUDITA_VALIDATION_*` values only when you want LLM-backed validators to use a different model, endpoint, credential, timeout, retry budget, or concurrency level.
## Parent-process pipe guidance
OpenRouter remains the default out of the box:
To avoid deadlocks in orchestrators:
- always read both stdout and stderr concurrently when invoking as a subprocess;
- prefer file outputs (`--output`, `--report-json`) for machine workflows;
- treat stderr as human-readable diagnostics, not structured data;
- parse structured results from output/report files.
```sh
export AUDITA_LLM_API_KEY=your-openrouter-key
audita process transcript.json --glossary glossary.yaml --output corrected.json
```
You can point Audita at any OpenAI-compatible endpoint by changing `AUDITA_BASE_URL` and, if needed, `AUDITA_MODEL`. For example, a local vLLM server:
```sh
export AUDITA_LLM_API_KEY=local-dev-key
export AUDITA_BASE_URL=http://localhost:8000/v1
export AUDITA_MODEL=meta-llama/Llama-3.1-8B-Instruct
audita process transcript.json --glossary glossary.yaml --output corrected.json
```
Or the actual OpenAI API:
```sh
export AUDITA_LLM_API_KEY=your-openai-key
export AUDITA_BASE_URL=https://api.openai.com/v1
export AUDITA_MODEL=gpt-4.1-mini
audita process transcript.json --glossary glossary.yaml --output corrected.json
```
`AUDITA_WORK_DIR` stores per-run diagnostics while processing. Under the default `AUDITA_WORK_DIR_RETENTION=auto`, clean successful runs are removed, while failed runs and successful runs with final skipped corrections are preserved. Use `always` to keep every run directory and `never` to remove successful run directories even when skips remain.
Failed runs always preserve the run directory and include an authoritative `report.json` alongside normalization and prompt/response diagnostics.
## Prototype Archive
The archived prototype remains importable as `audita_prototype` and is still covered by its original regression suite. This is intentional: the new `audita` package is a framework-oriented rewrite, not a thin wrapper around the old code.
For Go callers, prefer `exec.CommandContext` with explicit timeout/cancellation and buffered/streamed readers for both pipes.

File diff suppressed because it is too large Load Diff

View File

@@ -4,6 +4,11 @@ workspace:
storage:
backend: local
secrets:
# Optional: load environment variables from files in this directory.
# File name = env var name; file contents = env var value.
env_dir: /var/local/narratio/secrets
whisperx:
transcribe_url: "https://transcription.example.com/transcribe"
language: "en"

View File

@@ -170,6 +170,131 @@ func TestExecuteRunStageTranscribeUsesConfiguredWhisperXServer(t *testing.T) {
}
}
func TestExecuteRunStagePolishLoadsCredentialFromSecretsDir(t *testing.T) {
workspaceRoot := t.TempDir()
configDir := t.TempDir()
sessionID := "2026-05-03"
secretsDir := filepath.Join(configDir, "secrets")
if err := os.MkdirAll(secretsDir, 0o755); err != nil {
t.Fatalf("MkdirAll(%q): %v", secretsDir, err)
}
if err := os.WriteFile(filepath.Join(secretsDir, "OPENROUTER_API_KEY"), []byte("from-secret-file\n"), 0o600); err != nil {
t.Fatalf("write OPENROUTER_API_KEY secret file: %v", err)
}
seriatimBinary := writeSeriatimAppTestWrapper(t)
auditaBinary := writeAuditaAppTestWrapper(t)
t.Setenv("GO_WANT_APP_SERIATIM_HELPER", "1")
t.Setenv("GO_WANT_APP_AUDITA_HELPER", "1")
pipelinePath := filepath.Join(configDir, "pipeline.yml")
sessionPath := filepath.Join(configDir, "session.yml")
pipelineYAML := `workspace:
root: ` + workspaceRoot + `
storage:
backend: local
secrets:
env_dir: ./secrets
whisperx:
transcribe_url: https://example.com/transcribe
seriatim:
binary: ` + seriatimBinary + `
audita:
binary: ` + auditaBinary + `
llm_api_key_env: OPENROUTER_API_KEY
analyzer:
timeout: 20m
notification:
timeout: 10s
`
sessionYAML := `session_id: ` + sessionID + `
inputs:
audio_dir: ./audio
speakers_file: ./speakers.yml
autocorrect_file: ./autocorrect.yml
glossary_file: ./glossary.yml
`
if err := os.WriteFile(pipelinePath, []byte(pipelineYAML), 0o644); err != nil {
t.Fatalf("write pipeline.yml: %v", err)
}
if err := os.WriteFile(sessionPath, []byte(sessionYAML), 0o644); err != nil {
t.Fatalf("write session.yml: %v", err)
}
originalWD, err := os.Getwd()
if err != nil {
t.Fatalf("Getwd(): %v", err)
}
if err := os.Chdir(configDir); err != nil {
t.Fatalf("Chdir(%q): %v", configDir, err)
}
t.Cleanup(func() {
_ = os.Chdir(originalWD)
})
workRoot := filepath.Join(workspaceRoot, "work", sessionID)
mustWriteTestFile(t, filepath.Join(workRoot, "transcripts", "merged.json"), `{"schema":"seriatim-intermediate","segments":[]}`)
mustWriteTestFile(t, filepath.Join(workRoot, "inputs", "glossary.yml"), "[]\n")
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"run-stage", "--config", pipelinePath, "--session", sessionPath, "--force", "polish"}, &stdout, &stderr)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, stderr.String())
}
if !strings.Contains(stdout.String(), "stage=polish executed=1 skipped=0") {
t.Fatalf("stdout = %q, want polish execution", stdout.String())
}
}
func TestExecuteRunFailsWhenConfiguredSecretsDirMissing(t *testing.T) {
workspaceRoot := t.TempDir()
configDir := t.TempDir()
pipelinePath := filepath.Join(configDir, "pipeline.yml")
sessionPath := filepath.Join(configDir, "session.yml")
pipelineYAML := `workspace:
root: ` + workspaceRoot + `
storage:
backend: local
secrets:
env_dir: ./missing-secrets
whisperx:
transcribe_url: https://example.com/transcribe
seriatim:
binary: seriatim
audita:
binary: audita
analyzer:
timeout: 20m
notification:
timeout: 10s
`
sessionYAML := `session_id: 2026-05-03
inputs:
audio_dir: ./audio
speakers_file: ./speakers.yml
autocorrect_file: ./autocorrect.yml
glossary_file: ./glossary.yml
`
if err := os.WriteFile(pipelinePath, []byte(pipelineYAML), 0o644); err != nil {
t.Fatalf("write pipeline.yml: %v", err)
}
if err := os.WriteFile(sessionPath, []byte(sessionYAML), 0o644); err != nil {
t.Fatalf("write session.yml: %v", err)
}
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"run", "--config", pipelinePath, "--session", sessionPath}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if !strings.Contains(stderr.String(), "read secrets env_dir") {
t.Fatalf("stderr = %q, want secrets read-dir error context", stderr.String())
}
}
func TestExecuteUsesDefaultPipelineConfigPathWhenConfigFlagOmitted(t *testing.T) {
workspaceRoot := t.TempDir()
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {

View File

@@ -5,9 +5,12 @@ import (
"flag"
"fmt"
"io"
"log/slog"
"os"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/logging"
"gitea.maximumdirect.net/eric/narratio/internal/manifest"
)
@@ -45,6 +48,9 @@ func Plan(ctx context.Context, args []string, out io.Writer) error {
if err := config.Validate(cfg); err != nil {
return fmt.Errorf("plan: %w", err)
}
if _, err := loadSecretsFromConfig(cfg, logging.NewLogger(os.Stderr, slog.LevelInfo)); err != nil {
return fmt.Errorf("plan: %w", err)
}
store := artifacts.NewLocalStore(cfg.Pipeline.Workspace.Root)
paths, err := store.EnsureLayout(cfg.Session.SessionID)

View File

@@ -89,6 +89,53 @@ func TestPlanShowsRunAndSkipFromManifest(t *testing.T) {
}
}
func TestPlanFailsWhenConfiguredSecretsDirMissing(t *testing.T) {
workspaceRoot := t.TempDir()
configDir := t.TempDir()
pipelinePath := filepath.Join(configDir, "pipeline.yml")
sessionPath := filepath.Join(configDir, "session.yml")
pipelineYAML := `workspace:
root: ` + workspaceRoot + `
storage:
backend: local
secrets:
env_dir: ./missing-secrets
whisperx:
transcribe_url: https://example.com/transcribe
seriatim:
binary: seriatim
audita:
binary: audita
analyzer:
timeout: 20m
notification:
timeout: 10s
`
sessionYAML := `session_id: 2026-05-03
inputs:
audio_dir: ./audio
speakers_file: ./speakers.yml
autocorrect_file: ./autocorrect.yml
glossary_file: ./glossary.yml
`
if err := os.WriteFile(pipelinePath, []byte(pipelineYAML), 0o644); err != nil {
t.Fatalf("write pipeline.yml: %v", err)
}
if err := os.WriteFile(sessionPath, []byte(sessionYAML), 0o644); err != nil {
t.Fatalf("write session.yml: %v", err)
}
var out bytes.Buffer
err := Plan(context.Background(), []string{"--config", pipelinePath, "--session", sessionPath}, &out)
if err == nil {
t.Fatal("expected error, got nil")
}
if !strings.Contains(err.Error(), "read secrets env_dir") {
t.Fatalf("error = %q, want secrets read error context", err.Error())
}
}
func assertDir(t *testing.T, path string) {
t.Helper()
info, err := os.Stat(path)

View File

@@ -51,6 +51,9 @@ func executeStages(ctx context.Context, cfg *config.Config, stages []stage.Stage
if env.Logger == nil {
env.Logger = logging.NewLogger(os.Stderr, slog.LevelInfo)
}
if _, err := loadSecretsFromConfig(env.Config, env.Logger); err != nil {
return nil, fmt.Errorf("load secrets from files: %w", err)
}
if env.WhisperX == nil {
client, err := buildDefaultWhisperXClient(env.Config)
if err != nil {

View File

@@ -0,0 +1,88 @@
package app
import (
"fmt"
"log/slog"
"os"
"path/filepath"
"regexp"
"strings"
"gitea.maximumdirect.net/eric/narratio/internal/config"
)
var envVarNamePattern = regexp.MustCompile(`^[A-Za-z_][A-Za-z0-9_]*$`)
type secretsLoadStats struct {
Dir string
Loaded int
PreservedExisting int
Skipped int
}
func loadSecretsFromConfig(cfg *config.Config, logger *slog.Logger) (*secretsLoadStats, error) {
if cfg == nil || cfg.Pipeline == nil || cfg.Pipeline.Secrets == nil {
return nil, nil
}
rawDir := strings.TrimSpace(cfg.Pipeline.Secrets.EnvDir)
if rawDir == "" {
return nil, nil
}
resolvedDir := rawDir
if !filepath.IsAbs(resolvedDir) {
cwd, err := os.Getwd()
if err != nil {
return nil, fmt.Errorf("resolve secrets env_dir %q from current working directory: %w", rawDir, err)
}
resolvedDir = filepath.Join(cwd, resolvedDir)
}
resolvedDir = filepath.Clean(resolvedDir)
entries, err := os.ReadDir(resolvedDir)
if err != nil {
return nil, fmt.Errorf("read secrets env_dir %q: %w", resolvedDir, err)
}
stats := &secretsLoadStats{Dir: resolvedDir}
for _, entry := range entries {
name := entry.Name()
if !envVarNamePattern.MatchString(name) {
stats.Skipped++
continue
}
if entry.IsDir() {
stats.Skipped++
continue
}
secretPath := filepath.Join(resolvedDir, name)
bytes, err := os.ReadFile(secretPath)
if err != nil {
return nil, fmt.Errorf("read secret file %q: %w", secretPath, err)
}
value := strings.TrimRight(string(bytes), "\r\n")
if _, exists := os.LookupEnv(name); exists {
stats.PreservedExisting++
continue
}
if err := os.Setenv(name, value); err != nil {
return nil, fmt.Errorf("set environment variable %q from %q: %w", name, secretPath, err)
}
stats.Loaded++
}
if logger != nil {
logger.Info(
"loaded secret environment variables from filesystem",
"secrets_env_dir", stats.Dir,
"loaded", stats.Loaded,
"preserved_existing", stats.PreservedExisting,
"skipped", stats.Skipped,
)
}
return stats, nil
}

View File

@@ -0,0 +1,155 @@
package app
import (
"os"
"path/filepath"
"runtime"
"strings"
"testing"
"gitea.maximumdirect.net/eric/narratio/internal/config"
)
func TestLoadSecretsFromConfigLoadsValidFiles(t *testing.T) {
dir := t.TempDir()
mustWriteSecretFile(t, filepath.Join(dir, "NARRATIO_TEST_SECRET_A"), "value-1\n")
mustWriteSecretFile(t, filepath.Join(dir, "NARRATIO_TEST_SECRET_B"), "value-2\r\n")
mustWriteSecretFile(t, filepath.Join(dir, "not-valid-name.txt"), "ignored")
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Secrets: &config.SecretsConfig{EnvDir: dir},
},
}
stats, err := loadSecretsFromConfig(cfg, nil)
if err != nil {
t.Fatalf("loadSecretsFromConfig() error = %v", err)
}
if stats == nil {
t.Fatal("stats = nil, want non-nil")
}
if stats.Loaded != 2 {
t.Fatalf("Loaded = %d, want 2", stats.Loaded)
}
if stats.PreservedExisting != 0 {
t.Fatalf("PreservedExisting = %d, want 0", stats.PreservedExisting)
}
if stats.Skipped == 0 {
t.Fatalf("Skipped = %d, want > 0 for invalid filename", stats.Skipped)
}
if got := os.Getenv("NARRATIO_TEST_SECRET_A"); got != "value-1" {
t.Fatalf("NARRATIO_TEST_SECRET_A = %q, want value-1", got)
}
if got := os.Getenv("NARRATIO_TEST_SECRET_B"); got != "value-2" {
t.Fatalf("NARRATIO_TEST_SECRET_B = %q, want value-2", got)
}
}
func TestLoadSecretsFromConfigPreservesExistingEnv(t *testing.T) {
t.Setenv("OBJECT_STORAGE_KEY", "existing")
dir := t.TempDir()
mustWriteSecretFile(t, filepath.Join(dir, "OBJECT_STORAGE_KEY"), "from-file\n")
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Secrets: &config.SecretsConfig{EnvDir: dir},
},
}
stats, err := loadSecretsFromConfig(cfg, nil)
if err != nil {
t.Fatalf("loadSecretsFromConfig() error = %v", err)
}
if stats.PreservedExisting != 1 {
t.Fatalf("PreservedExisting = %d, want 1", stats.PreservedExisting)
}
if got := os.Getenv("OBJECT_STORAGE_KEY"); got != "existing" {
t.Fatalf("OBJECT_STORAGE_KEY = %q, want existing", got)
}
}
func TestLoadSecretsFromConfigRelativeDirUsesCWD(t *testing.T) {
cwd := t.TempDir()
secretsDir := filepath.Join(cwd, "secrets")
if err := os.MkdirAll(secretsDir, 0o755); err != nil {
t.Fatalf("MkdirAll(%q): %v", secretsDir, err)
}
mustWriteSecretFile(t, filepath.Join(secretsDir, "OBJECT_STORAGE_KEY_ID"), "id-123\n")
originalWD, err := os.Getwd()
if err != nil {
t.Fatalf("Getwd(): %v", err)
}
if err := os.Chdir(cwd); err != nil {
t.Fatalf("Chdir(%q): %v", cwd, err)
}
t.Cleanup(func() {
_ = os.Chdir(originalWD)
})
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Secrets: &config.SecretsConfig{EnvDir: "./secrets"},
},
}
if _, err := loadSecretsFromConfig(cfg, nil); err != nil {
t.Fatalf("loadSecretsFromConfig() error = %v", err)
}
if got := os.Getenv("OBJECT_STORAGE_KEY_ID"); got != "id-123" {
t.Fatalf("OBJECT_STORAGE_KEY_ID = %q, want id-123", got)
}
}
func TestLoadSecretsFromConfigMissingDirFails(t *testing.T) {
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Secrets: &config.SecretsConfig{EnvDir: filepath.Join(t.TempDir(), "missing")},
},
}
_, err := loadSecretsFromConfig(cfg, nil)
if err == nil {
t.Fatal("expected error, got nil")
}
if !strings.Contains(err.Error(), "read secrets env_dir") {
t.Fatalf("error = %q, want read-dir context", err.Error())
}
}
func TestLoadSecretsFromConfigUnreadableValidEntryFails(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("symlink behavior differs on windows")
}
dir := t.TempDir()
broken := filepath.Join(dir, "OPENROUTER_API_KEY")
if err := os.Symlink(filepath.Join(dir, "does-not-exist"), broken); err != nil {
t.Fatalf("Symlink(%q): %v", broken, err)
}
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Secrets: &config.SecretsConfig{EnvDir: dir},
},
}
_, err := loadSecretsFromConfig(cfg, nil)
if err == nil {
t.Fatal("expected error, got nil")
}
if !strings.Contains(err.Error(), "read secret file") {
t.Fatalf("error = %q, want read secret file context", err.Error())
}
}
func mustWriteSecretFile(t *testing.T, path, contents string) {
t.Helper()
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
t.Fatalf("MkdirAll(%q): %v", path, err)
}
if err := os.WriteFile(path, []byte(contents), 0o600); err != nil {
t.Fatalf("WriteFile(%q): %v", path, err)
}
}

View File

@@ -12,6 +12,7 @@ type Config struct {
type PipelineConfig struct {
Workspace WorkspaceConfig `yaml:"workspace"`
Storage StorageConfig `yaml:"storage"`
Secrets *SecretsConfig `yaml:"secrets"`
WhisperX WhisperXConfig `yaml:"whisperx"`
Seriatim SeriatimConfig `yaml:"seriatim"`
Audita AuditaConfig `yaml:"audita"`
@@ -36,6 +37,11 @@ type WorkspaceConfig struct {
Root string `yaml:"root"`
}
// SecretsConfig configures optional local filesystem secret loading.
type SecretsConfig struct {
EnvDir string `yaml:"env_dir"`
}
// StorageConfig configures storage backends and related parameters.
type StorageConfig struct {
Backend string `yaml:"backend"`

View File

@@ -72,6 +72,46 @@ inputs:
`,
wantLoadErr: "strict decode failed",
},
{
name: "unknown secrets field fails",
pipelineYAML: `workspace:
root: /tmp/narratio
whisperx:
transcribe_url: https://transcription.ai.rakestrawhome.com/transcribe
seriatim:
binary: seriatim
secrets:
bogus: true
`,
sessionYAML: `session_id: 2026-05-03
inputs:
audio_dir: ./audio
speakers_file: ./speakers.yml
autocorrect_file: ./autocorrect.yml
glossary_file: ./glossary.yml
`,
wantLoadErr: "strict decode failed",
},
{
name: "empty secrets env_dir fails",
pipelineYAML: `workspace:
root: /tmp/narratio
whisperx:
transcribe_url: https://transcription.ai.rakestrawhome.com/transcribe
seriatim:
binary: seriatim
secrets:
env_dir: " "
`,
sessionYAML: `session_id: 2026-05-03
inputs:
audio_dir: ./audio
speakers_file: ./speakers.yml
autocorrect_file: ./autocorrect.yml
glossary_file: ./glossary.yml
`,
wantValidate: "pipeline config \"pipeline.yml\" invalid: pipeline.secrets.env_dir must be non-empty when pipeline.secrets is configured",
},
{
name: "unknown session field fails",
pipelineYAML: `workspace:

View File

@@ -33,6 +33,9 @@ func validatePipeline(cfg *PipelineConfig) error {
if strings.TrimSpace(cfg.Workspace.Root) == "" {
return fmt.Errorf("pipeline.workspace.root is required")
}
if err := validateSecrets(cfg.Secrets); err != nil {
return err
}
if err := validateWhisperX(cfg.WhisperX); err != nil {
return err
}
@@ -61,6 +64,16 @@ func validatePipeline(cfg *PipelineConfig) error {
return nil
}
func validateSecrets(cfg *SecretsConfig) error {
if cfg == nil {
return nil
}
if strings.TrimSpace(cfg.EnvDir) == "" {
return fmt.Errorf("pipeline.secrets.env_dir must be non-empty when pipeline.secrets is configured")
}
return nil
}
func validateNormalize(cfg *NormalizeConfig) error {
if cfg == nil {
return nil