# Narratio S3 Input and Archive Implementation Plan ## 1. Purpose This document defines the implementation plan for adding S3-based audio input and S3 archival/promotion to `narratio`. The feature has two related responsibilities: 1. **Input acquisition**: load source audio files from an S3 bucket into local working storage before transcription. 2. **Archival and promotion**: after a successful run, upload the complete run record to S3 and promote selected outputs to stable session-level paths. The design preserves the existing stage-based architecture: ```text prepare transcribe merge polish normalize trim analyze archive notify ``` The new S3 behavior should fit into the existing modular design: - `prepare` acquires input audio. - intermediate stages operate on the local workdir. - `archive` uploads successful run outputs and promotes configured artifacts. - external storage details remain behind a storage backend abstraction. - failed runs remain local for diagnostics and are not uploaded to S3. ## 2. Finalized Design Decisions The following design choices are settled: ```text S3 session prefix: {root_prefix}/campaigns/{campaign}/sessions/{session_id}/ Example: dnd/campaigns/forsaken/sessions/2026-04-19/ run_id format: 20260515T031522Z-a1b2c3d4 local work path: /var/lib/narratio/work/{campaign}/{session_id}/{run_id}/ local spool path: /var/spool/narratio/{campaign}/{session_id}/{run_id}/audio/ S3 audio source: audio already exists in S3 before narratio runs failed runs: retained locally only; not uploaded to S3 archive: real final pipeline stage promotion: performed during archive stage after all prior required stages succeed default promoted outputs: transcripts/trimmed.json artifacts/session_recap.md current/manifest.json current/run_id.txt raw WhisperX transcripts: uploaded under runs/{run_id}/transcripts/raw/ audio re-upload: original source audio is not re-uploaded by archive by default current/run_id.txt: written last as the effective S3 commit pointer ``` ### 2.1 Implementation Status (2026-05-16) Implemented in repository: - storage/archive configuration and validation foundations: - `storage.s3` - `spool` - `archive` - promotion-rule safety checks - `inputs.audio_s3` modeling - run and path-model foundations: - run ID generation (`YYYYMMDDTHHMMSSZ-xxxxxxxx`) - S3 key builders for session/run/current/promoted destinations - campaign/session/run local work and spool path helpers - manifest run/path identity fields - examples and tests for the above foundations - remote storage backend layer: - object-store abstraction with `List`, `Download`, `Upload`, and `Exists` - fake storage backend for deterministic, no-network testing - S3-compatible backend built from `storage.s3` config - backend construction helper from resolved config Not implemented yet: - prepare-stage S3 list/download behavior - archive-stage S3 upload behavior - promotion uploads - writing `current/manifest.json` and `current/run_id.txt` to S3 ## 3. S3 Layout The canonical S3 layout should be: ```text s3://{bucket}/{root_prefix}/campaigns/{campaign}/sessions/{session_id}/ audio/ speaker-1.flac speaker-2.flac transcripts/ trimmed.json artifacts/ session_recap.md current/ manifest.json run_id.txt runs/ {run_id}/ inputs/ session.yml speakers.yml glossary.yml autocorrect.yml pipeline.resolved.yml transcripts/ raw/ speaker-1.json speaker-2.json merged.json processed.json normalized.json trimmed.json artifacts/ session_bounds.json session_recap.md reports/ seriatim.merge.report.json audita.report.json seriatim.normalize.report.json seriatim.trim.report.json config/ seriatim.generated.yml seriatim.normalize.generated.yml seriatim.trim.generated.yml audita.generated.yml scriptorium.bounds.generated.yml scriptorium.session_recap.generated.yml logs/ whisperx.*.log seriatim.*.log audita.*.log scriptorium.*.log manifest.json ``` ### 3.1 Session Root The session root is: ```text {root_prefix}/campaigns/{campaign}/sessions/{session_id}/ ``` For example: ```text dnd/campaigns/forsaken/sessions/2026-04-19/ ``` The `campaigns/` path segment is intentional. It leaves room for future campaign-level material: ```text dnd/campaigns/{campaign}/campaign.yml dnd/campaigns/{campaign}/glossary.yml dnd/campaigns/{campaign}/characters/ dnd/campaigns/{campaign}/sessions/ ``` ### 3.2 Session-Level Paths The session-level root contains the durable, promoted, current view of the session: ```text audio/ transcripts/ artifacts/ current/ runs/ ``` The top-level `audio/` directory is the source of truth for original session audio. Audio is assumed already present in S3 and should not be re-uploaded by `archive` by default. The top-level `transcripts/` and `artifacts/` directories should contain only configured promoted outputs. ### 3.3 Run-Specific Paths Each successful run is uploaded under: ```text runs/{run_id}/ ``` This contains the full run record: - materialized inputs - intermediate transcripts - raw WhisperX transcripts - tool reports - generated configs - logs - generated artifacts - manifest Inputs belong under `runs/{run_id}/inputs/`, not at the session root, because inputs are run-specific. A rerun may use different `speakers.yml`, `glossary.yml`, `autocorrect.yml`, `pipeline.resolved.yml`, Scriptorium prompts, trim settings, models, or runtime configuration. ### 3.4 Current Pointer The effective commit pointer is: ```text current/run_id.txt ``` This file should be written last during archive. `current/manifest.json` should also be written during promotion so consumers can inspect the current promoted run without first resolving the run directory. The archive stage should upload in this order: 1. run record under `runs/{run_id}/` 2. configured promoted outputs under top-level `transcripts/` and `artifacts/` 3. `current/manifest.json` 4. `current/run_id.txt` last This makes `current/run_id.txt` the closest practical S3 equivalent of an atomic session commit marker. ## 4. Local Filesystem Layout The production local layout should be: ```text /var/lib/narratio/ work/ {campaign}/ {session_id}/ {run_id}/ inputs/ audio/ transcripts/ artifacts/ reports/ config/ logs/ manifest.json .lock /var/spool/narratio/ {campaign}/ {session_id}/ {run_id}/ audio/ speaker-1.flac speaker-2.flac ``` For local development, these roots should be configurable. For example: ```yaml workspace: root: ./workspace spool: root: ./spool ``` Resulting in: ```text ./workspace/work/{campaign}/{session_id}/{run_id}/ ./spool/{campaign}/{session_id}/{run_id}/audio/ ``` ## 5. Local Audio Handling The intended local flow is: ```text S3 audio ↓ spool/audio ↓ work/audio ↓ transcribe ``` The `prepare` stage should: 1. list `.flac` objects under the configured S3 audio prefix, 2. fail clearly if none are found, 3. download audio files into the spool directory, 4. copy or materialize them into the run workdir `audio/`, 5. record provenance in the manifest. The rest of the pipeline should use `work/audio/`, not S3 paths directly and not spool paths. Audio cleanup should be conservative in the first implementation. It is acceptable to add a config knob such as: ```yaml spool: delete_audio_after_archive: true ``` For v1, prefer retaining local workdir audio until the run has successfully archived. The source of truth remains S3, but local diagnostics are valuable during development. ## 6. Configuration Design ### 6.1 Storage Config Add or refine a storage section: ```yaml storage: s3: bucket: "my-dnd-archive" root_prefix: "dnd" region: "us-east-1" endpoint: "" force_path_style: false ``` Requirements: - `bucket` is required when S3 input/archive is enabled. - `root_prefix` defaults to `dnd`. - `region` may be optional depending on SDK behavior. - `endpoint` is optional for S3-compatible storage. - `force_path_style` is useful for MinIO/Garage/S3-compatible backends. - Credentials must not be stored in config. Use standard AWS environment/profile/instance-role mechanisms. ### 6.2 Workspace and Spool Config Use: ```yaml workspace: root: "/var/lib/narratio" spool: root: "/var/spool/narratio" delete_audio_after_archive: false ``` If existing `workspace.root` currently points directly to a work root, the implementation should either preserve the existing semantics or migrate carefully with documentation. The new layout should include campaign/session/run path segments. ### 6.3 Session Config The session config should include campaign and session ID: ```yaml session: id: "2026-04-19" campaign: "forsaken" ``` If the current config shape uses top-level `session_id`, either migrate to the nested shape with compatibility or maintain the current shape while ensuring both campaign and session ID are available to the path builder. ### 6.4 S3 Audio Input Config Use relative audio prefix resolution: ```yaml inputs: audio_s3: prefix: "audio/" ``` This resolves relative to: ```text {root_prefix}/campaigns/{campaign}/sessions/{session_id}/ ``` For example: ```text dnd/campaigns/forsaken/sessions/2026-04-19/audio/ ``` The prepare stage should fail if no `.flac` files are found under this prefix. Local audio input should continue working for development unless intentionally deprecated later. ### 6.5 Archive Config Add an archive config section: ```yaml archive: enabled: true upload_run: true promote_artifacts: - from: "transcripts/trimmed.json" to: "transcripts/trimmed.json" required: true - from: "artifacts/session_recap.md" to: "artifacts/session_recap.md" required: true ``` If `promote_artifacts` is omitted, use the built-in default list: ```yaml archive: promote_artifacts: - from: "transcripts/trimmed.json" to: "transcripts/trimmed.json" required: true - from: "artifacts/session_recap.md" to: "artifacts/session_recap.md" required: true ``` Additionally, archive should always write: ```text current/manifest.json current/run_id.txt ``` Those current-pointer artifacts are part of archive semantics and should not need to be listed in `promote_artifacts`. ## 7. Run Identity `run_id` should be first-class. Use the format: ```text YYYYMMDDTHHMMSSZ-xxxxxxxx ``` Example: ```text 20260515T031522Z-a1b2c3d4 ``` Properties: - sortable by timestamp - human-readable - collision-resistant via short random suffix - safe for file paths and S3 keys The manifest should include: ```json { "campaign": "forsaken", "session_id": "2026-04-19", "run_id": "20260515T031522Z-a1b2c3d4", "local_workdir": "/var/lib/narratio/work/forsaken/2026-04-19/20260515T031522Z-a1b2c3d4", "s3_session_prefix": "dnd/campaigns/forsaken/sessions/2026-04-19/", "s3_run_prefix": "dnd/campaigns/forsaken/sessions/2026-04-19/runs/20260515T031522Z-a1b2c3d4/" } ``` ### 7.1 CLI Behavior Recommended behavior: ```text run: creates a new run_id unless --run-id is supplied resume: uses --run-id when supplied otherwise may discover latest local run for the campaign/session run-stage: uses --run-id when supplied otherwise may discover latest local run for the campaign/session ``` For safety, implementation may choose to require `--run-id` for `resume` and `run-stage` when multiple local runs exist. The exact CLI behavior should be documented. ## 8. Manifest Changes Extend manifest data to include: - `campaign` - `session_id` - `run_id` - `local_workdir` - `local_spool_dir` - `s3_bucket` - `s3_session_prefix` - `s3_run_prefix` - archive status and promoted outputs - S3 source provenance for audio inputs - S3 destination records for archived outputs Audio input records should include: ```text source = s3 s3_bucket s3_key local_path size etag sha256 if computed ``` Use ETag as S3 metadata only, not as a reliable checksum. SHA-256 should be computed after download if practical. ## 9. Storage Backend Abstraction S3 code should not leak into stages. Add or extend a storage backend interface with operations like: ```text List(ctx, prefix) ([]ObjectInfo, error) Download(ctx, key, localPath) error Upload(ctx, localPath, key, metadata) error Exists(ctx, key) (bool, error) ``` Potential object metadata: ```text key size etag last_modified ``` A future copy method may be useful, but v1 can upload from local paths. The `prepare` stage should use the storage backend to list/download audio. The `archive` stage should use the storage backend to upload run records and promoted outputs. Tests should use a fake storage backend, not real S3. ## 10. Prepare Stage Changes The `prepare` stage should support both existing local audio workflows and the new S3 audio source. ### 10.1 S3 Audio Flow When `inputs.audio_s3.prefix` is configured: 1. compute the session root: ```text {root_prefix}/campaigns/{campaign}/sessions/{session_id}/ ``` 2. resolve the audio prefix: ```text {session_root}/{inputs.audio_s3.prefix} ``` 3. list objects under that prefix 4. filter to `.flac` 5. fail clearly if no `.flac` files are found 6. download each file to: ```text {spool.root}/{campaign}/{session_id}/{run_id}/audio/ ``` 7. copy or materialize each file to: ```text {workspace.root}/work/{campaign}/{session_id}/{run_id}/audio/ ``` 8. compute local checksums if practical 9. record audio input provenance in manifest ### 10.2 Local Audio Flow Existing local audio flow should continue to work unless intentionally changed later. If both local audio and S3 audio are configured, fail clearly unless a precedence rule is explicitly documented. Prefer requiring exactly one audio input source. ## 11. Archive Stage The archive stage becomes a real stage. It should run after `analyze` and before `notify`. Archive should only upload successful runs. Since the normal runner is sequential, if any prior stage fails, archive will not run. If the user invokes `run-stage archive` manually, the archive stage should validate prerequisite stages before uploading. ### 11.1 Archive Prerequisites For v1, require these stages to have succeeded before archive: ```text prepare transcribe merge polish normalize trim analyze ``` If a stage is optional in a future config, this prerequisite list may become configurable. For now, hardcoded prerequisites are acceptable. ### 11.2 Run Upload Upload the local workdir record to: ```text {session_root}/runs/{run_id}/ ``` Suggested mapping: ```text workdir/inputs/ → runs/{run_id}/inputs/ workdir/transcripts/ → runs/{run_id}/transcripts/ workdir/artifacts/ → runs/{run_id}/artifacts/ workdir/reports/ → runs/{run_id}/reports/ workdir/config/ → runs/{run_id}/config/ workdir/logs/ → runs/{run_id}/logs/ workdir/manifest.json → runs/{run_id}/manifest.json ``` If reports currently live under `artifacts/`, implementation may either: 1. keep that local layout and upload them under `runs/{run_id}/artifacts/`, or 2. add a logical archive mapping into `runs/{run_id}/reports/`. Avoid disruptive local layout changes unless they are already easy and well-tested. ### 11.3 Promotion For each configured promotion rule: ```yaml - from: "transcripts/trimmed.json" to: "transcripts/trimmed.json" required: true ``` Upload: ```text local workdir/transcripts/trimmed.json → s3://bucket/{session_root}/transcripts/trimmed.json ``` Rules: - `from` is local workdir-relative. - `to` is session-root-relative. - if `required: true` and the source is missing, archive fails. - if `required: false` and the source is missing, archive records a skipped promotion. ### 11.4 Commit Pointer Write these last: ```text current/manifest.json current/run_id.txt ``` `current/run_id.txt` should contain exactly the run ID plus a trailing newline. Writing `current/run_id.txt` last is the effective S3 commit marker. ### 11.5 Failed Runs Failed runs should not be uploaded to S3. Failed workdirs should remain local for diagnostics. The archive stage should never upload a run that does not satisfy its prerequisite success checks. ## 12. Audio Upload Policy Audio is assumed already present under: ```text {session_root}/audio/ ``` Archive should not re-upload source audio by default. The archive stage may record audio input provenance in manifest, but should avoid duplicating large FLAC files under `runs/{run_id}/`. A future option may support uploading local audio into S3, but that is not part of this implementation. ## 13. Promotion Defaults Default promoted outputs: ```text transcripts/trimmed.json artifacts/session_recap.md ``` Always write current pointers: ```text current/manifest.json current/run_id.txt ``` Do not promote `session_bounds.json` by default. Do not promote raw transcripts, logs, tool reports, generated configs, or full intermediate transcript tiers by default. They remain available under `runs/{run_id}/`. ## 14. Testing Strategy Tests should not require real S3. Use a fake storage backend for: - listing audio objects - downloading objects - uploading objects - recording upload order - simulating missing objects - simulating upload failures ### 14.1 Config Tests Test: - valid S3 storage config - missing bucket when S3 mode enabled - default root prefix - invalid archive promotion rules - default promotion list - audio_s3 prefix validation - local audio config still works - conflict when both S3 audio and local audio are configured, if that rule is implemented ### 14.2 Path Builder Tests Test S3 key construction: ```text dnd/campaigns/forsaken/sessions/2026-04-19/audio/ dnd/campaigns/forsaken/sessions/2026-04-19/runs/{run_id}/... dnd/campaigns/forsaken/sessions/2026-04-19/current/run_id.txt ``` Test local paths: ```text /var/lib/narratio/work/forsaken/2026-04-19/{run_id}/ /var/spool/narratio/forsaken/2026-04-19/{run_id}/audio/ ``` ### 14.3 Prepare Tests Test: - S3 audio prefix with `.flac` objects downloads files - no `.flac` objects fails clearly - non-FLAC objects are ignored - downloaded audio is materialized in workdir audio - manifest records S3 provenance - local audio mode still works - fake storage errors fail the stage clearly ### 14.4 Archive Tests Test: - archive refuses to run if prerequisites are missing or failed - successful archive uploads run record - promotion rules upload configured outputs - default promotion list applies - optional missing promotion is skipped - required missing promotion fails - `current/manifest.json` is uploaded near the end - `current/run_id.txt` is uploaded last - failed runs are not uploaded - audio files are not uploaded by default - manifest records archive metadata - fake storage upload failure fails the stage clearly ### 14.5 CLI/Run-ID Tests Test: - `run` creates a run ID - supplied `--run-id` is honored - `resume` can find or require a run ID according to final CLI policy - `run-stage` can find or require a run ID according to final CLI policy - multiple local runs are handled deterministically ## 15. Implementation Sequence ### Config and Path Model (Implemented) Implement: - storage.s3 config - spool config - archive config - promotion rules - run ID generator - campaign-aware local work/spool path builder - S3 session/run key builder No real S3 calls yet. Expected commit: ```text Add archive storage path configuration ``` ### Storage Backend Interface and S3 Backend (Implemented) Implemented: - storage backend interface - object metadata type - fake backend - real S3 backend using AWS SDK or existing project dependency policy - backend construction from config No prepare/archive stage behavior yet. Expected commit: ```text Add S3 storage backend abstraction ``` ### Prepare Stage S3 Audio Download Implement: - S3 audio source support - list/download `.flac` files - fail on empty audio prefix - materialize audio into workdir - manifest S3 provenance - preserve local audio mode Expected commit: ```text Download S3 audio during prepare" ``` ### Real Archive Stage Run Upload Implement: - archive prerequisites - upload successful run workdir to `runs/{run_id}/` - manifest archive metadata - no promotion yet, or minimal internal scaffolding only Expected commit: ```text Upload successful run records to S3 ``` ### Promotion Rules and Current Pointer Implement: - default promotion rules - configurable promotion rules - required/optional behavior - top-level promoted uploads - `current/manifest.json` - `current/run_id.txt` written last Expected commit: ```text Promote current session artifacts to S3 ``` ### Documentation and Examples Update: - architecture.md - README.md - examples - runbook instructions - config samples - S3 layout documentation Expected commit: ```text Document S3 archive workflow ``` ### Architectural Review Review: - storage code isolation - prepare/archive stage boundaries - no S3 leakage into unrelated stages - no failed-run upload - current pointer semantics - promotion config - tests - docs Expected commit: ```text Review S3 archive architecture ``` ## 16. Operational Workflow Expected production flow: 1. Upload source audio to: ```text s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/audio/ ``` 2. Run narratio: ```text narratio run --config pipeline.yml --session sessions/{session_id}/session.yml ``` 3. `prepare` downloads audio into spool/workdir. 4. Pipeline runs locally. 5. `archive` uploads the successful run record. 6. `archive` promotes configured current outputs. 7. `archive` writes `current/manifest.json`. 8. `archive` writes `current/run_id.txt` last. Consumers can then read: ```text s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/current/run_id.txt s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/transcripts/trimmed.json s3://{bucket}/dnd/campaigns/{campaign}/sessions/{session_id}/artifacts/session_recap.md ``` ## 17. Security and Privacy Rules: - Do not store AWS credentials in config. - Use standard AWS credential mechanisms. - Do not log full environment variables. - Do not store secrets in manifests or generated configs. - Treat transcripts and artifacts as potentially sensitive. - Do not upload failed runs to S3. - Preserve local failed workdirs for diagnostics. - Do not re-upload source audio by default. - Be careful not to log transcript contents during archive. ## 18. Non-Goals Do not implement in this feature: - uploading local audio to S3 - uploading failed runs to S3 - remote deletion or cleanup policies - S3 object lifecycle configuration - remote locking - multi-user concurrency control - a database-backed run registry - a generic artifact publishing framework beyond the configured promotion list - checksum-based stale detection, except where checksums are recorded as metadata - archive-time redaction of manifests; secrets should not enter manifests in the first place ## 19. Open Follow-Up Ideas Potential future improvements: - local retention policy for successful workdirs - optional cleanup of spool audio after successful archive - optional upload of human-readable transcript exports - optional promotion of `session_bounds.json` - S3-side run index by date/model/prompt version - S3 object metadata for checksums and content types - remote run discovery for `resume` - support for local-audio-to-S3 ingestion mode - optional failed-run diagnostic upload behind an explicit flag - checksum-based stale detection and stage invalidation - archive verification pass after upload