Files
narratio/docs/roadmap/narratio-s3-archive-implementation-plan.md

1036 lines
24 KiB
Markdown

# Narratio S3 Input and Archive Implementation Plan
## 1. Purpose
This document defines the implementation plan for adding S3-based audio input and S3 archival/promotion to `narratio`.
The feature has two related responsibilities:
1. **Input acquisition**: load source audio files from an S3 bucket into local working storage before transcription.
2. **Archival and promotion**: after a successful run, upload the complete run record to S3 and promote selected outputs to stable session-level paths.
The design preserves the existing stage-based architecture:
```text
prepare
transcribe
merge
polish
normalize
trim
analyze
archive
notify
```
The new S3 behavior should fit into the existing modular design:
- `prepare` acquires input audio.
- intermediate stages operate on the local workdir.
- `archive` uploads successful run outputs and promotes configured artifacts.
- external storage details remain behind a storage backend abstraction.
- failed runs remain local for diagnostics and are not uploaded to S3.
## 2. Finalized Design Decisions
The following design choices are settled:
```text
S3 session prefix:
{root_prefix}/campaigns/{campaign}/sessions/{session_id}/
Example:
dnd/campaigns/forsaken/sessions/2026-04-19/
run_id format:
20260515T031522Z-a1b2c3d4
local work path:
/var/lib/narratio/work/{campaign}/{session_id}/{run_id}/
local spool path:
/var/spool/narratio/{campaign}/{session_id}/{run_id}/audio/
S3 audio source:
audio already exists in S3 before narratio runs
failed runs:
retained locally only; not uploaded to S3
archive:
real final pipeline stage
promotion:
performed during archive stage after all prior required stages succeed
default promoted outputs:
transcripts/trimmed.json
artifacts/session_recap.md
current/manifest.json
current/run_id.txt
raw WhisperX transcripts:
uploaded under runs/{run_id}/transcripts/raw/
audio re-upload:
original source audio is not re-uploaded by archive by default
current/run_id.txt:
written last as the effective S3 commit pointer
```
### 2.1 Implementation Status (2026-05-16)
Implemented in repository:
- storage/archive configuration and validation foundations:
- `storage.s3`
- `spool`
- `archive`
- promotion-rule safety checks
- `inputs.audio_s3` modeling
- run and path-model foundations:
- run ID generation (`YYYYMMDDTHHMMSSZ-xxxxxxxx`)
- S3 key builders for session/run/current/promoted destinations
- campaign/session/run local work and spool path helpers
- manifest run/path identity fields
- examples and tests for the above foundations
- remote storage backend layer:
- object-store abstraction with `List`, `Download`, `Upload`, and `Exists`
- fake storage backend for deterministic, no-network testing
- S3-compatible backend built from `storage.s3` config
- backend construction helper from resolved config
Not implemented yet:
- prepare-stage S3 list/download behavior
- archive-stage S3 upload behavior
- promotion uploads
- writing `current/manifest.json` and `current/run_id.txt` to S3
## 3. S3 Layout
The canonical S3 layout should be:
```text
s3://{bucket}/{root_prefix}/campaigns/{campaign}/sessions/{session_id}/
audio/
speaker-1.flac
speaker-2.flac
transcripts/
trimmed.json
artifacts/
session_recap.md
current/
manifest.json
run_id.txt
runs/
{run_id}/
inputs/
session.yml
speakers.yml
glossary.yml
autocorrect.yml
pipeline.resolved.yml
transcripts/
raw/
speaker-1.json
speaker-2.json
merged.json
processed.json
normalized.json
trimmed.json
artifacts/
session_bounds.json
session_recap.md
reports/
seriatim.merge.report.json
audita.report.json
seriatim.normalize.report.json
seriatim.trim.report.json
config/
seriatim.generated.yml
seriatim.normalize.generated.yml
seriatim.trim.generated.yml
audita.generated.yml
scriptorium.bounds.generated.yml
scriptorium.session_recap.generated.yml
logs/
whisperx.*.log
seriatim.*.log
audita.*.log
scriptorium.*.log
manifest.json
```
### 3.1 Session Root
The session root is:
```text
{root_prefix}/campaigns/{campaign}/sessions/{session_id}/
```
For example:
```text
dnd/campaigns/forsaken/sessions/2026-04-19/
```
The `campaigns/` path segment is intentional. It leaves room for future campaign-level material:
```text
dnd/campaigns/{campaign}/campaign.yml
dnd/campaigns/{campaign}/glossary.yml
dnd/campaigns/{campaign}/characters/
dnd/campaigns/{campaign}/sessions/
```
### 3.2 Session-Level Paths
The session-level root contains the durable, promoted, current view of the session:
```text
audio/
transcripts/
artifacts/
current/
runs/
```
The top-level `audio/` directory is the source of truth for original session audio. Audio is assumed already present in S3 and should not be re-uploaded by `archive` by default.
The top-level `transcripts/` and `artifacts/` directories should contain only configured promoted outputs.
### 3.3 Run-Specific Paths
Each successful run is uploaded under:
```text
runs/{run_id}/
```
This contains the full run record:
- materialized inputs
- intermediate transcripts
- raw WhisperX transcripts
- tool reports
- generated configs
- logs
- generated artifacts
- manifest
Inputs belong under `runs/{run_id}/inputs/`, not at the session root, because inputs are run-specific. A rerun may use different `speakers.yml`, `glossary.yml`, `autocorrect.yml`, `pipeline.resolved.yml`, Scriptorium prompts, trim settings, models, or runtime configuration.
### 3.4 Current Pointer
The effective commit pointer is:
```text
current/run_id.txt
```
This file should be written last during archive.
`current/manifest.json` should also be written during promotion so consumers can inspect the current promoted run without first resolving the run directory.
The archive stage should upload in this order:
1. run record under `runs/{run_id}/`
2. configured promoted outputs under top-level `transcripts/` and `artifacts/`
3. `current/manifest.json`
4. `current/run_id.txt` last
This makes `current/run_id.txt` the closest practical S3 equivalent of an atomic session commit marker.
## 4. Local Filesystem Layout
The production local layout should be:
```text
/var/lib/narratio/
work/
{campaign}/
{session_id}/
{run_id}/
inputs/
audio/
transcripts/
artifacts/
reports/
config/
logs/
manifest.json
.lock
/var/spool/narratio/
{campaign}/
{session_id}/
{run_id}/
audio/
speaker-1.flac
speaker-2.flac
```
For local development, these roots should be configurable. For example:
```yaml
workspace:
root: ./workspace
spool:
root: ./spool
```
Resulting in:
```text
./workspace/work/{campaign}/{session_id}/{run_id}/
./spool/{campaign}/{session_id}/{run_id}/audio/
```
## 5. Local Audio Handling
The intended local flow is:
```text
S3 audio
spool/audio
work/audio
transcribe
```
The `prepare` stage should:
1. list `.flac` objects under the configured S3 audio prefix,
2. fail clearly if none are found,
3. download audio files into the spool directory,
4. copy or materialize them into the run workdir `audio/`,
5. record provenance in the manifest.
The rest of the pipeline should use `work/audio/`, not S3 paths directly and not spool paths.
Audio cleanup should be conservative in the first implementation. It is acceptable to add a config knob such as:
```yaml
spool:
delete_audio_after_archive: true
```
For v1, prefer retaining local workdir audio until the run has successfully archived. The source of truth remains S3, but local diagnostics are valuable during development.
## 6. Configuration Design
### 6.1 Storage Config
Add or refine a storage section:
```yaml
storage:
s3:
bucket: "my-dnd-archive"
root_prefix: "dnd"
region: "us-east-1"
endpoint: ""
force_path_style: false
```
Requirements:
- `bucket` is required when S3 input/archive is enabled.
- `root_prefix` defaults to `dnd`.
- `region` may be optional depending on SDK behavior.
- `endpoint` is optional for S3-compatible storage.
- `force_path_style` is useful for MinIO/Garage/S3-compatible backends.
- Credentials must not be stored in config. Use standard AWS environment/profile/instance-role mechanisms.
### 6.2 Workspace and Spool Config
Use:
```yaml
workspace:
root: "/var/lib/narratio"
spool:
root: "/var/spool/narratio"
delete_audio_after_archive: false
```
If existing `workspace.root` currently points directly to a work root, the implementation should either preserve the existing semantics or migrate carefully with documentation. The new layout should include campaign/session/run path segments.
### 6.3 Session Config
The session config should include campaign and session ID:
```yaml
session:
id: "2026-04-19"
campaign: "forsaken"
```
If the current config shape uses top-level `session_id`, either migrate to the nested shape with compatibility or maintain the current shape while ensuring both campaign and session ID are available to the path builder.
### 6.4 S3 Audio Input Config
Use relative audio prefix resolution:
```yaml
inputs:
audio_s3:
prefix: "audio/"
```
This resolves relative to:
```text
{root_prefix}/campaigns/{campaign}/sessions/{session_id}/
```
For example:
```text
dnd/campaigns/forsaken/sessions/2026-04-19/audio/
```
The prepare stage should fail if no `.flac` files are found under this prefix.
Local audio input should continue working for development unless intentionally deprecated later.
### 6.5 Archive Config
Add an archive config section:
```yaml
archive:
enabled: true
upload_run: true
promote_artifacts:
- from: "transcripts/trimmed.json"
to: "transcripts/trimmed.json"
required: true
- from: "artifacts/session_recap.md"
to: "artifacts/session_recap.md"
required: true
```
If `promote_artifacts` is omitted, use the built-in default list:
```yaml
archive:
promote_artifacts:
- from: "transcripts/trimmed.json"
to: "transcripts/trimmed.json"
required: true
- from: "artifacts/session_recap.md"
to: "artifacts/session_recap.md"
required: true
```
Additionally, archive should always write:
```text
current/manifest.json
current/run_id.txt
```
Those current-pointer artifacts are part of archive semantics and should not need to be listed in `promote_artifacts`.
## 7. Run Identity
`run_id` should be first-class.
Use the format:
```text
YYYYMMDDTHHMMSSZ-xxxxxxxx
```
Example:
```text
20260515T031522Z-a1b2c3d4
```
Properties:
- sortable by timestamp
- human-readable
- collision-resistant via short random suffix
- safe for file paths and S3 keys
The manifest should include:
```json
{
"campaign": "forsaken",
"session_id": "2026-04-19",
"run_id": "20260515T031522Z-a1b2c3d4",
"local_workdir": "/var/lib/narratio/work/forsaken/2026-04-19/20260515T031522Z-a1b2c3d4",
"s3_session_prefix": "dnd/campaigns/forsaken/sessions/2026-04-19/",
"s3_run_prefix": "dnd/campaigns/forsaken/sessions/2026-04-19/runs/20260515T031522Z-a1b2c3d4/"
}
```
### 7.1 CLI Behavior
Recommended behavior:
```text
run:
creates a new run_id unless --run-id is supplied
resume:
uses --run-id when supplied
otherwise may discover latest local run for the campaign/session
run-stage:
uses --run-id when supplied
otherwise may discover latest local run for the campaign/session
```
For safety, implementation may choose to require `--run-id` for `resume` and `run-stage` when multiple local runs exist.
The exact CLI behavior should be documented.
## 8. Manifest Changes
Extend manifest data to include:
- `campaign`
- `session_id`
- `run_id`
- `local_workdir`
- `local_spool_dir`
- `s3_bucket`
- `s3_session_prefix`
- `s3_run_prefix`
- archive status and promoted outputs
- S3 source provenance for audio inputs
- S3 destination records for archived outputs
Audio input records should include:
```text
source = s3
s3_bucket
s3_key
local_path
size
etag
sha256 if computed
```
Use ETag as S3 metadata only, not as a reliable checksum.
SHA-256 should be computed after download if practical.
## 9. Storage Backend Abstraction
S3 code should not leak into stages.
Add or extend a storage backend interface with operations like:
```text
List(ctx, prefix) ([]ObjectInfo, error)
Download(ctx, key, localPath) error
Upload(ctx, localPath, key, metadata) error
Exists(ctx, key) (bool, error)
```
Potential object metadata:
```text
key
size
etag
last_modified
```
A future copy method may be useful, but v1 can upload from local paths.
The `prepare` stage should use the storage backend to list/download audio.
The `archive` stage should use the storage backend to upload run records and promoted outputs.
Tests should use a fake storage backend, not real S3.
## 10. Prepare Stage Changes
The `prepare` stage should support both existing local audio workflows and the new S3 audio source.
### 10.1 S3 Audio Flow
When `inputs.audio_s3.prefix` is configured:
1. compute the session root:
```text
{root_prefix}/campaigns/{campaign}/sessions/{session_id}/
```
2. resolve the audio prefix:
```text
{session_root}/{inputs.audio_s3.prefix}
```
3. list objects under that prefix
4. filter to `.flac`
5. fail clearly if no `.flac` files are found
6. download each file to:
```text
{spool.root}/{campaign}/{session_id}/{run_id}/audio/
```
7. copy or materialize each file to:
```text
{workspace.root}/work/{campaign}/{session_id}/{run_id}/audio/
```
8. compute local checksums if practical
9. record audio input provenance in manifest
### 10.2 Local Audio Flow
Existing local audio flow should continue to work unless intentionally changed later.
If both local audio and S3 audio are configured, fail clearly unless a precedence rule is explicitly documented. Prefer requiring exactly one audio input source.
## 11. Archive Stage
The archive stage becomes a real stage.
It should run after `analyze` and before `notify`.
Archive should only upload successful runs.
Since the normal runner is sequential, if any prior stage fails, archive will not run. If the user invokes `run-stage archive` manually, the archive stage should validate prerequisite stages before uploading.
### 11.1 Archive Prerequisites
For v1, require these stages to have succeeded before archive:
```text
prepare
transcribe
merge
polish
normalize
trim
analyze
```
If a stage is optional in a future config, this prerequisite list may become configurable. For now, hardcoded prerequisites are acceptable.
### 11.2 Run Upload
Upload the local workdir record to:
```text
{session_root}/runs/{run_id}/
```
Suggested mapping:
```text
workdir/inputs/ → runs/{run_id}/inputs/
workdir/transcripts/ → runs/{run_id}/transcripts/
workdir/artifacts/ → runs/{run_id}/artifacts/
workdir/reports/ → runs/{run_id}/reports/
workdir/config/ → runs/{run_id}/config/
workdir/logs/ → runs/{run_id}/logs/
workdir/manifest.json → runs/{run_id}/manifest.json
```
If reports currently live under `artifacts/`, implementation may either:
1. keep that local layout and upload them under `runs/{run_id}/artifacts/`, or
2. add a logical archive mapping into `runs/{run_id}/reports/`.
Avoid disruptive local layout changes unless they are already easy and well-tested.
### 11.3 Promotion
For each configured promotion rule:
```yaml
- from: "transcripts/trimmed.json"
to: "transcripts/trimmed.json"
required: true
```
Upload:
```text
local workdir/transcripts/trimmed.json
→ s3://bucket/{session_root}/transcripts/trimmed.json
```
Rules:
- `from` is local workdir-relative.
- `to` is session-root-relative.
- if `required: true` and the source is missing, archive fails.
- if `required: false` and the source is missing, archive records a skipped promotion.
### 11.4 Commit Pointer
Write these last:
```text
current/manifest.json
current/run_id.txt
```
`current/run_id.txt` should contain exactly the run ID plus a trailing newline.
Writing `current/run_id.txt` last is the effective S3 commit marker.
### 11.5 Failed Runs
Failed runs should not be uploaded to S3.
Failed workdirs should remain local for diagnostics.
The archive stage should never upload a run that does not satisfy its prerequisite success checks.
## 12. Audio Upload Policy
Audio is assumed already present under:
```text
{session_root}/audio/
```
Archive should not re-upload source audio by default.
The archive stage may record audio input provenance in manifest, but should avoid duplicating large FLAC files under `runs/{run_id}/`.
A future option may support uploading local audio into S3, but that is not part of this implementation.
## 13. Promotion Defaults
Default promoted outputs:
```text
transcripts/trimmed.json
artifacts/session_recap.md
```
Always write current pointers:
```text
current/manifest.json
current/run_id.txt
```
Do not promote `session_bounds.json` by default.
Do not promote raw transcripts, logs, tool reports, generated configs, or full intermediate transcript tiers by default. They remain available under `runs/{run_id}/`.
## 14. Testing Strategy
Tests should not require real S3.
Use a fake storage backend for:
- listing audio objects
- downloading objects
- uploading objects
- recording upload order
- simulating missing objects
- simulating upload failures
### 14.1 Config Tests
Test:
- valid S3 storage config
- missing bucket when S3 mode enabled
- default root prefix
- invalid archive promotion rules
- default promotion list
- audio_s3 prefix validation
- local audio config still works
- conflict when both S3 audio and local audio are configured, if that rule is implemented
### 14.2 Path Builder Tests
Test S3 key construction:
```text
dnd/campaigns/forsaken/sessions/2026-04-19/audio/
dnd/campaigns/forsaken/sessions/2026-04-19/runs/{run_id}/...
dnd/campaigns/forsaken/sessions/2026-04-19/current/run_id.txt
```
Test local paths:
```text
/var/lib/narratio/work/forsaken/2026-04-19/{run_id}/
/var/spool/narratio/forsaken/2026-04-19/{run_id}/audio/
```
### 14.3 Prepare Tests
Test:
- S3 audio prefix with `.flac` objects downloads files
- no `.flac` objects fails clearly
- non-FLAC objects are ignored
- downloaded audio is materialized in workdir audio
- manifest records S3 provenance
- local audio mode still works
- fake storage errors fail the stage clearly
### 14.4 Archive Tests
Test:
- archive refuses to run if prerequisites are missing or failed
- successful archive uploads run record
- promotion rules upload configured outputs
- default promotion list applies
- optional missing promotion is skipped
- required missing promotion fails
- `current/manifest.json` is uploaded near the end
- `current/run_id.txt` is uploaded last
- failed runs are not uploaded
- audio files are not uploaded by default
- manifest records archive metadata
- fake storage upload failure fails the stage clearly
### 14.5 CLI/Run-ID Tests
Test:
- `run` creates a run ID
- supplied `--run-id` is honored
- `resume` can find or require a run ID according to final CLI policy
- `run-stage` can find or require a run ID according to final CLI policy
- multiple local runs are handled deterministically
## 15. Implementation Sequence
### Config and Path Model (Implemented)
Implement:
- storage.s3 config
- spool config
- archive config
- promotion rules
- run ID generator
- campaign-aware local work/spool path builder
- S3 session/run key builder
No real S3 calls yet.
Expected commit:
```text
Add archive storage path configuration
```
### Storage Backend Interface and S3 Backend (Implemented)
Implemented:
- storage backend interface
- object metadata type
- fake backend
- real S3 backend using AWS SDK or existing project dependency policy
- backend construction from config
No prepare/archive stage behavior yet.
Expected commit:
```text
Add S3 storage backend abstraction
```
### Prepare Stage S3 Audio Download
Implement:
- S3 audio source support
- list/download `.flac` files
- fail on empty audio prefix
- materialize audio into workdir
- manifest S3 provenance
- preserve local audio mode
Expected commit:
```text
Download S3 audio during prepare"
```
### Real Archive Stage Run Upload
Implement:
- archive prerequisites
- upload successful run workdir to `runs/{run_id}/`
- manifest archive metadata
- no promotion yet, or minimal internal scaffolding only
Expected commit:
```text
Upload successful run records to S3
```
### Promotion Rules and Current Pointer
Implement:
- default promotion rules
- configurable promotion rules
- required/optional behavior
- top-level promoted uploads
- `current/manifest.json`
- `current/run_id.txt` written last
Expected commit:
```text
Promote current session artifacts to S3
```
### Documentation and Examples
Update:
- architecture.md
- README.md
- examples
- runbook instructions
- config samples
- S3 layout documentation
Expected commit:
```text
Document S3 archive workflow
```
### Architectural Review
Review:
- storage code isolation
- prepare/archive stage boundaries
- no S3 leakage into unrelated stages
- no failed-run upload
- current pointer semantics
- promotion config
- tests
- docs
Expected commit:
```text
Review S3 archive architecture
```
## 16. Operational Workflow
Expected production flow:
1. Upload source audio to:
```text
s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/audio/
```
2. Run narratio:
```text
narratio run --config pipeline.yml --session sessions/{session_id}/session.yml
```
3. `prepare` downloads audio into spool/workdir.
4. Pipeline runs locally.
5. `archive` uploads the successful run record.
6. `archive` promotes configured current outputs.
7. `archive` writes `current/manifest.json`.
8. `archive` writes `current/run_id.txt` last.
Consumers can then read:
```text
s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/current/run_id.txt
s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/transcripts/trimmed.json
s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/artifacts/session_recap.md
```
## 17. Security and Privacy
Rules:
- Do not store AWS credentials in config.
- Use standard AWS credential mechanisms.
- Do not log full environment variables.
- Do not store secrets in manifests or generated configs.
- Treat transcripts and artifacts as potentially sensitive.
- Do not upload failed runs to S3.
- Preserve local failed workdirs for diagnostics.
- Do not re-upload source audio by default.
- Be careful not to log transcript contents during archive.
## 18. Non-Goals
Do not implement in this feature:
- uploading local audio to S3
- uploading failed runs to S3
- remote deletion or cleanup policies
- S3 object lifecycle configuration
- remote locking
- multi-user concurrency control
- a database-backed run registry
- a generic artifact publishing framework beyond the configured promotion list
- checksum-based stale detection, except where checksums are recorded as metadata
- archive-time redaction of manifests; secrets should not enter manifests in the first place
## 19. Open Follow-Up Ideas
Potential future improvements:
- local retention policy for successful workdirs
- optional cleanup of spool audio after successful archive
- optional upload of human-readable transcript exports
- optional promotion of `session_bounds.json`
- S3-side run index by date/model/prompt version
- S3 object metadata for checksums and content types
- remote run discovery for `resume`
- support for local-audio-to-S3 ingestion mode
- optional failed-run diagnostic upload behind an explicit flag
- checksum-based stale detection and stage invalidation
- archive verification pass after upload