10 Commits

36 changed files with 4129 additions and 958 deletions

View File

@@ -6,7 +6,7 @@
narratio run --session-id 2026-04-04 narratio run --session-id 2026-04-04
``` ```
This command uses default config discovery for `pipeline.yml` and `session.yml`; both files must be discoverable unless you pass explicit `--config` and `--session` paths. This command uses default discovery for `pipeline.yml` and `session.yml`; both files must be discoverable unless you pass explicit `--config` and `--session` paths.
## Command Overview ## Command Overview
@@ -17,6 +17,7 @@ Implemented commands:
- `resume`: continue from first non-succeeded stage unless forced. - `resume`: continue from first non-succeeded stage unless forced.
- `status`: read and print stage statuses from an existing manifest. - `status`: read and print stage statuses from an existing manifest.
- `run-stage`: execute exactly one stage. - `run-stage`: execute exactly one stage.
- `restore`: restore durable local session state from the committed remote archive state.
Unknown commands print usage and exit non-zero. Unknown commands print usage and exit non-zero.
@@ -68,6 +69,15 @@ Valid stage names:
- `archive` - `archive`
- `notify` - `notify`
### `restore`
- `--config <path>`
- `--session <path>`
- `--session-id <value>`
- `--dry-run`: plan restore actions without writing local files.
- `--force`: overwrite local conflicting files with remote archive files.
- `--include-audio`: include durable archived `audio/**` files in restore scope.
### `status` ### `status`
- `--manifest <path>`: required manifest path. - `--manifest <path>`: required manifest path.
@@ -178,6 +188,39 @@ Common failure cases:
- unknown stage name. - unknown stage name.
- using `--artifacts` with any non-`analyze` stage. - using `--artifacts` with any non-`analyze` stage.
### `restore`
Purpose:
- Restore durable session state (`manifest.json`, `transcripts/**`, `artifacts/**`, and optional `audio/**`) from the committed remote archive current state.
Syntax:
```bash
narratio restore [--config <pipeline.yml>] [--session <session.yml>] [--session-id <id>] [--dry-run] [--force] [--include-audio]
```
Success output (dry-run):
- `Restore plan for <campaign>/<session_id>`
- `Remote run: <run_id>`
- `Would download: <n>`
- `Would skip unchanged: <n>`
- `Conflicts: <n>`
Success output (non-dry-run):
- `Restored session archive for <campaign>/<session_id>`
- `Remote run: <run_id>`
- `Downloaded: <n>`
- `Skipped unchanged: <n>`
- `Conflicts: <n>`
Common failure cases:
- storage backend is not configured.
- remote `current/run_id.txt` missing/empty.
- remote `current/manifest.json` missing or invalid.
- remote manifest session/campaign mismatch.
- local conflicts without `--force`.
- session lock conflict.
## Common Workflows ## Common Workflows
Default-discovery run: Default-discovery run:
@@ -204,6 +247,19 @@ Run only analyze stage with selected artifacts:
narratio run-stage --session-id 2026-04-04 --artifacts player_handout analyze narratio run-stage --session-id 2026-04-04 --artifacts player_handout analyze
``` ```
Preview restore actions without writes:
```bash
narratio restore --session-id 2026-04-04 --dry-run
```
Restore and then force analyze:
```bash
narratio restore --session-id 2026-04-04
narratio run-stage --session-id 2026-04-04 --force analyze
```
## Diagnostic / Recovery Commands ## Diagnostic / Recovery Commands
Inspect stage status: Inspect stage status:
@@ -219,4 +275,4 @@ Get manifest path from previous output:
- `--artifacts` filters which configured artifacts are executable when analyze runs. - `--artifacts` filters which configured artifacts are executable when analyze runs.
- `--artifacts` does not imply `--force`. - `--artifacts` does not imply `--force`.
- If analyze is already `succeeded` and `--force` is not set, runner-level skip still applies. - if analyze is already `succeeded` and `--force` is not set, runner-level skip still applies.

View File

@@ -13,6 +13,7 @@ These commands load and validate both files before running:
- `narratio plan` - `narratio plan`
- `narratio resume` - `narratio resume`
- `narratio run-stage` - `narratio run-stage`
- `narratio restore`
Behavior: Behavior:
@@ -23,28 +24,28 @@ Behavior:
## 2. Config file discovery ## 2. Config file discovery
Pipeline config lookup for `run`, `plan`, `resume`, and `run-stage`: Pipeline config lookup for `run`, `plan`, `resume`, `run-stage`, and `restore`:
- If `--config <path>` is provided, that path is used. - if `--config <path>` is provided, that path is used.
- If omitted, Narratio searches in order: - if omitted, Narratio searches in order:
1. `/usr/local/etc/narratio/pipeline.yml` 1. `/usr/local/etc/narratio/pipeline.yml`
2. `/etc/narratio/pipeline.yml` 2. `/etc/narratio/pipeline.yml`
- First existing file wins. - first existing file wins.
## 3. Session file discovery and templating ## 3. Session file discovery and templating
Session config lookup for `run`, `plan`, `resume`, and `run-stage`: Session config lookup for `run`, `plan`, `resume`, `run-stage`, and `restore`:
- If `--session <path>` is provided, that path is used. - if `--session <path>` is provided, that path is used.
- If omitted, Narratio searches in order: - if omitted, Narratio searches in order:
1. `./session.yml` 1. `./session.yml`
2. `/usr/local/etc/narratio/session.yml` 2. `/usr/local/etc/narratio/session.yml`
3. `/etc/narratio/session.yml` 3. `/etc/narratio/session.yml`
- First existing file wins. - first existing file wins.
Template behavior: Template behavior:
- Supported placeholders: - supported placeholders:
- `{{session_id}}` - `{{session_id}}`
- `{{ session_id }}` - `{{ session_id }}`
- `--session-id <value>` supplies the placeholder value. - `--session-id <value>` supplies the placeholder value.
@@ -106,11 +107,11 @@ archive:
enabled: true enabled: true
upload_run: true upload_run: true
promote_artifacts: promote_artifacts:
- from: transcripts/trimmed.json - source: narratio.transcript.trimmed
to: transcripts/trimmed.json dest: transcripts/trimmed.json
required: true required: true
- from: artifacts/session_recap.md - source: narratio.artifact.session_recap
to: artifacts/session_recap.md dest: artifacts/session_recap.md
required: true required: true
whisperx: whisperx:
@@ -130,8 +131,10 @@ scriptorium:
Operational notes: Operational notes:
- archive promotion is explicit and path-based via `archive.promote_artifacts`. - archive promotion is explicit and source-based via `archive.promote_artifacts`.
- `source` is required; `dest` is optional and derived when omitted.
- Narratio does not auto-promote all generated analyze artifacts. - Narratio does not auto-promote all generated analyze artifacts.
- `restore` reads the same config/session inputs and restore scope is bounded by committed archive current state.
## 7. Full pipeline reference ## 7. Full pipeline reference
@@ -154,9 +157,9 @@ Operational notes:
| `pipeline.spool.delete_audio_after_archive` | bool | No | `false` | | `pipeline.spool.delete_audio_after_archive` | bool | No | `false` |
| `pipeline.archive.enabled` | bool | No | `true` | | `pipeline.archive.enabled` | bool | No | `true` |
| `pipeline.archive.upload_run` | bool | No | `true` | | `pipeline.archive.upload_run` | bool | No | `true` |
| `pipeline.archive.promote_artifacts[]` | list | No | trimmed + session_recap rules | | `pipeline.archive.promote_artifacts[]` | list | No | trimmed transcript rule |
| `pipeline.archive.promote_artifacts[].from` | string | Yes (per rule) | none | | `pipeline.archive.promote_artifacts[].source` | string | Yes (per rule) | none |
| `pipeline.archive.promote_artifacts[].to` | string | Yes (per rule) | none | | `pipeline.archive.promote_artifacts[].dest` | string | No | derived from source |
| `pipeline.archive.promote_artifacts[].required` | bool | No | `true` | | `pipeline.archive.promote_artifacts[].required` | bool | No | `true` |
| `pipeline.whisperx.transcribe_url` | string | Yes | none | | `pipeline.whisperx.transcribe_url` | string | Yes | none |
| `pipeline.whisperx.language` | string | No | `en` | | `pipeline.whisperx.language` | string | No | `en` |
@@ -248,6 +251,29 @@ Allowed `pipeline.scriptorium.artifacts.<name>.inputs.<key>.source` values:
- `narratio.bounds.session` - `narratio.bounds.session`
- `narratio.artifact.<configured_artifact_key>` - `narratio.artifact.<configured_artifact_key>`
`pipeline.archive.promote_artifacts[].source` values:
- `narratio.transcript.merged`
- `narratio.transcript.polished`
- `narratio.transcript.full`
- `narratio.transcript.trimmed`
- `narratio.bounds.session`
- `narratio.artifact.<configured_artifact_key>`
Archive promotion destination rules:
- `dest` must be a clean relative path (not absolute, no traversal).
- duplicate `dest` values are rejected.
- if `dest` is omitted:
- built-in sources derive their canonical destination path;
- configured sources derive from `pipeline.scriptorium.artifacts.<name>.output_path`;
- derivation failure is a config validation error.
Restore-related implications:
- restore remote identity requires archive S3 identity to resolve (`pipeline.storage.s3.bucket` and session prefix derivation inputs).
- restore scope considers only committed current state and durable paths (`manifest.json`, `transcripts/**`, `artifacts/**`, optional `audio/**`).
## 8. Full session reference ## 8. Full session reference
| Path | Type | Required | Default | | Path | Type | Required | Default |

View File

@@ -4,7 +4,7 @@
Developers and LLM coding agents changing Narratio internals. Developers and LLM coding agents changing Narratio internals.
## Scope ## Scope
Implementation-accurate contracts for workspace/state, manifests, stages, artifact resolution, and adapter boundaries. Implementation-accurate contracts for workspace/state, manifests, stages, artifact resolution, adapter boundaries, and restore command behavior.
## Component Docs ## Component Docs
- `adapters.md`: external adapter map, runtime wiring, and boundary ownership. - `adapters.md`: external adapter map, runtime wiring, and boundary ownership.
@@ -12,6 +12,7 @@ Implementation-accurate contracts for workspace/state, manifests, stages, artifa
- `manifest.md`: session/run manifest schemas, lifecycle transitions, and persistence semantics. - `manifest.md`: session/run manifest schemas, lifecycle transitions, and persistence semantics.
- `artifacts.md`: built-in artifact registry, runtime artifact catalog, and source-resolution behavior. - `artifacts.md`: built-in artifact registry, runtime artifact catalog, and source-resolution behavior.
- `workspace.md`: local state model, manifests, run-local layout, promotion, and cleanup invariants. - `workspace.md`: local state model, manifests, run-local layout, promotion, and cleanup invariants.
- `command-restore.md`: restore command discovery/planning/execution/reporting contract.
- `stage-prepare.md`: input materialization and provenance capture. - `stage-prepare.md`: input materialization and provenance capture.
- `stage-transcribe.md`: WhisperX transcript generation. - `stage-transcribe.md`: WhisperX transcript generation.
- `stage-merge.md`: Seriatim normalization + merge. - `stage-merge.md`: Seriatim normalization + merge.

View File

@@ -0,0 +1,86 @@
# Internal: Command Restore
## Purpose
Define the implemented `narratio restore` command contract: committed remote-state discovery, deterministic planning, safe file installation, conflict policy, and restore reporting.
## Inputs and outputs
Inputs:
- CLI flags: `--config`, `--session`, `--session-id`, `--dry-run`, `--force`, `--include-audio`.
- Resolved/validated `pipeline.yml` and `session.yml`.
- Configured remote object store.
- Remote committed current-state markers (`current/run_id.txt`, `current/manifest.json`).
Outputs:
- Dry-run summary to stdout (plan + counts).
- Non-dry-run completion summary to stdout.
- Local durable session files restored under canonical session root.
- Non-dry-run restore report at `reports/restore-latest.json`.
## Boundaries
Owns:
- Restore command flag parsing and command wiring.
- Remote current-state discovery and identity validation.
- Restore plan construction and conflict classification.
- Restore execution for planned downloads.
- Restore report model and persistence.
Does not own:
- Stage execution orchestration (`run`, `resume`, `run-stage`).
- Archive publish behavior (owned by archive stage).
- Storage transport implementation details (owned by storage adapters).
## Config fields used
- Config/session discovery and templating fields consumed by all commands.
- `pipeline.workspace.root` (local restore target root).
- `pipeline.storage.*` (remote backend + archive identity derivation).
- `pipeline.storage.s3.*` identity components used by archive prefix helpers.
- `session.session_id`
- `session.campaign`
## External adapters used
- `storage.ObjectStore` for `Exists`, `List`, `Download`.
- `artifacts.Store` (`LocalStore`) for layout and session lock management.
- `manifest.LocalStore` for manifest decode/validation and identity checks.
## State and manifest behavior
- Restore is not a pipeline run and does not create a run manifest.
- Restore uses committed remote current state only:
- `current/run_id.txt` must exist and be non-empty.
- `current/manifest.json` must decode and match requested session/campaign.
- Non-dry-run writes restore files to canonical session paths.
- Manifest install behavior:
- validated before replacement.
- installed last among download actions.
- existing local manifest is preserved if restored manifest validation/install fails.
- Non-dry-run report persists summary/action status metadata in `reports/restore-latest.json`.
## Skip and resume behavior
- Restore does not participate in stage skip/resume decisions.
- Restore provides durable local state so subsequent stage commands can resume or rerun based on restored manifest state.
- Dry-run is read-only and returns plan output only.
## Failure behavior
- Fails when storage backend is unavailable or archive identity cannot be resolved.
- Fails when remote current pointer/manifest is missing or invalid.
- Fails when remote manifest identity mismatches requested campaign/session.
- Fails on local conflicts unless `--force` is set.
- Fails fast on session lock acquisition conflict for non-dry-run execution.
- On execution failure, previously installed files remain; no rollback is performed.
## Tests to inspect before changing
- `internal/app/restore_test.go`
- `internal/app/restore_discovery_test.go`
- `internal/app/restore_plan_test.go`
- `internal/app/restore_execution_test.go`
- `internal/app/restore_workflow_test.go`
- `internal/artifacts/archive_identity_test.go`
## Architectural invariants
- Restore relies on centralized archive identity/key helpers (`internal/artifacts`) rather than ad hoc key building.
- `current/run_id.txt` is the remote commit marker; restore must not infer committed state from incidental files.
- Local path mapping is traversal-safe and constrained to session root.
- Restore scope is deterministic and path-classified:
- include `manifest.json`, `transcripts/**`, `artifacts/**`
- include `audio/**` only with `--include-audio`
- exclude `runs/**`, `logs/**`, `reports/**`, `config/**`, `inputs/**`
- Command remains standalone; no implicit `run --restore` behavior.

View File

@@ -7,7 +7,7 @@ Publish run records and promoted session artifacts to object storage, then atomi
Inputs: Inputs:
- session manifest and prerequisite stage records - session manifest and prerequisite stage records
- run root contents under `runs/{run_id}/` - run root contents under `runs/{run_id}/`
- promotion sources from session root (`archive.promote_artifacts`) - promotion rules with artifact `source` IDs and archive `dest` paths (`archive.promote_artifacts`)
Outputs: Outputs:
- uploaded run files under `{session_prefix}/runs/{run_id}/...` - uploaded run files under `{session_prefix}/runs/{run_id}/...`

View File

@@ -6,7 +6,7 @@ For field-level configuration, see [docs/config.md](./config.md). For full comma
## Normal workflow (S3-first path) ## Normal workflow (S3-first path)
1. Upload session `.flac` files to object storage under the session audio prefix. 1. Upload session `.flac` files to object storage under the configured session audio prefix.
2. Run Narratio: 2. Run Narratio:
```bash ```bash
@@ -21,6 +21,37 @@ Notes:
- default config/session discovery applies unless `--config` and `--session` are passed. - default config/session discovery applies unless `--config` and `--session` are passed.
- S3 audio mode requires `session.inputs.audio_s3.prefix` and valid object-store access. - S3 audio mode requires `session.inputs.audio_s3.prefix` and valid object-store access.
## Restore workflow
Use restore when local durable session state is missing or stale and archive current state is authoritative.
Dry-run (no local writes):
```bash
narratio restore --session-id 2026-04-04 --dry-run
```
Execution:
```bash
narratio restore --session-id 2026-04-04
```
Post-restore analyze rerun pattern:
```bash
narratio run-stage --session-id 2026-04-04 --force analyze
```
Restore source-of-truth:
- remote commit marker: `current/run_id.txt`
- remote current manifest: `current/manifest.json`
Restore default scope:
- includes `manifest.json`, `transcripts/**`, `artifacts/**`
- includes `audio/**` only with `--include-audio`
- excludes `runs/**`, `logs/**`, `reports/**`, `config/**`, `inputs/**`, and `current/**` (except remote `current/manifest.json` as source)
## Local filesystem layout and state artifacts ## Local filesystem layout and state artifacts
Session root: Session root:
@@ -29,7 +60,7 @@ Session root:
Primary state: Primary state:
- `manifest.json`: session-level stage state. - `manifest.json`: session-level stage state.
- `runs/{run_id}/manifest.json`: invocation-level state. - `runs/{run_id}/manifest.json`: invocation-level state.
- `.lock`: session lock while a run is active. - `.lock`: session lock while a modifying command is active.
Canonical session directories: Canonical session directories:
- `inputs/` - `inputs/`
@@ -48,6 +79,7 @@ Run-local stage directories:
Behavior: Behavior:
- directory creation is idempotent. - directory creation is idempotent.
- stage outputs are generally generated run-local first, then promoted to canonical paths on success. - stage outputs are generally generated run-local first, then promoted to canonical paths on success.
- restore installs downloaded files to canonical session paths and does not recreate historical run sandboxes.
## Analyze artifact execution lifecycle ## Analyze artifact execution lifecycle
@@ -84,12 +116,14 @@ Publish order:
`current/run_id.txt` is the remote commit marker. `current/run_id.txt` is the remote commit marker.
Archive promotion is explicit and path-based: Archive promotion is explicit and source-based:
- Narratio does not auto-promote all generated analyze artifacts. - Narratio does not auto-promote all generated analyze artifacts.
- each rule resolves `source` through the artifact resolver/catalog model, then uploads to `dest`.
- missing required promotion sources fail archive stage. - missing required promotion sources fail archive stage.
- missing optional promotion sources are skipped. - missing optional promotion sources are skipped.
- invalid resolved artifacts fail archive stage.
## Resume, retry, and safe rerun behavior ## Resume, retry, restore, and safe rerun behavior
Default skip: Default skip:
- `run` and `run-stage` skip already-succeeded stages unless `--force` is set. - `run` and `run-stage` skip already-succeeded stages unless `--force` is set.
@@ -98,6 +132,11 @@ Resume:
- `resume` starts at first non-succeeded stage. - `resume` starts at first non-succeeded stage.
- `resume --force` runs full stage order. - `resume --force` runs full stage order.
Restore conflict policy:
- restore classifies local differences as conflicts.
- without `--force`, restore fails when conflicts exist.
- with `--force`, conflicting local files are overwritten by remote archive files.
Forced reruns: Forced reruns:
- force-rerunning an upstream succeeded stage marks downstream succeeded stages as `stale`. - force-rerunning an upstream succeeded stage marks downstream succeeded stages as `stale`.
@@ -123,13 +162,18 @@ No cleanup for failed/incomplete/unarchived/archive-skipped runs.
## Failure and recovery playbooks ## Failure and recovery playbooks
After failure, Narratio keeps: After run failure, Narratio keeps:
- session manifest - session manifest
- run manifest - run manifest
- run-local artifacts/logs/config/reports - run-local artifacts/logs/config/reports
Failed or incomplete runs remain local-only. Failed or incomplete runs remain local-only.
After restore failure:
- already-installed restore files remain in place.
- restore does not roll back prior successful installs.
- existing local manifest is preserved if restored manifest validation/install fails.
Recommended recovery: Recommended recovery:
1. inspect state: 1. inspect state:
@@ -138,8 +182,27 @@ Recommended recovery:
narratio status --manifest <manifest-path> narratio status --manifest <manifest-path>
``` ```
2. fix root cause (config/input/credentials/service availability). 2. for restore-specific checks, run:
3. continue with `resume`, or targeted `run-stage --force` followed by `resume`.
```bash
narratio restore --session-id 2026-04-04 --dry-run
```
3. fix root cause (config/input/credentials/storage/service availability).
4. continue with `resume`, or targeted `run-stage --force` followed by `resume`.
## Restore report
Non-dry-run restore writes a durable report at:
- `reports/restore-latest.json`
Report content includes:
- identity (`campaign`, `session_id`, `run_id`)
- mode flags (`dry_run`, `force`, `include_audio`)
- plan counts and execution counts
- per-action status
Dry-run does not write restore report files.
## Operational caveats ## Operational caveats
@@ -147,3 +210,4 @@ narratio status --manifest <manifest-path>
- local and S3 audio input modes are mutually exclusive. - local and S3 audio input modes are mutually exclusive.
- archive publish requires upstream stages through `analyze` to be `succeeded`. - archive publish requires upstream stages through `analyze` to be `succeeded`.
- required promotion rules can fail when selected analyze artifacts did not generate a required file path. - required promotion rules can fail when selected analyze artifacts did not generate a required file path.
- restore requires configured remote object storage and committed remote current state.

715
docs/roadmap/restore.md Normal file
View File

@@ -0,0 +1,715 @@
# Roadmap: `narratio restore` Subcommand
## Status
Implemented through Step 8. This document remains as roadmap and design history for the restore feature, and as the home for future restore-related ideas (for example `run --restore`).
## Summary
Add a new `narratio restore` subcommand that hydrates a local session workspace from the current committed remote archive state.
The primary operator workflow is:
```bash
narratio restore --session-id 2026-04-04
narratio run-stage --force analyze
```
This should allow a new machine with no local workspace state to restore the durable session manifest, transcripts, and generated artifacts from S3, then generate new Scriptorium artifacts without re-running transcription, merge, normalize, polish, or trim.
This is intentionally a separate command. Do not fold this behavior into the `prepare` stage. The existing `prepare` stage should remain focused on materializing configured local/S3 inputs for a pipeline run.
## Goals
- Add a first-class `narratio restore` command.
- Restore the current committed remote session state into the canonical local session workspace.
- Use the existing object storage adapter boundary.
- Preserve archive commit semantics: only restore from a remote state that has a valid current commit marker.
- Restore durable session-level outputs needed for downstream stages, especially `analyze`.
- Provide safe conflict behavior by default.
- Support `--dry-run`, `--force`, and `--include-audio`.
- Keep the implementation explicit, testable, and narrow.
## Non-goals
- Do not make `restore` a pipeline stage.
- Do not change the `prepare` stage behavior as part of this work.
- Do not add implicit restore behavior to `narratio run` in this implementation.
- Do not restore historical run-local sandboxes by default.
- Do not implement a generic remote synchronization engine.
- Do not implement bidirectional sync.
- Do not delete local files merely because they are absent remotely.
- Do not merge remote and local manifests in the first implementation.
- Do not require live S3 for the ordinary unit test suite.
## Future work explicitly out of scope
A future change may add:
```bash
narratio run --restore
```
That future flag should run `narratio restore` before starting the normal pipeline. Mention this as future work in roadmap/docs if useful, but do not implement it now.
## Existing architecture to preserve
### `prepare` remains input materialization
The `prepare` stage currently materializes required session inputs into canonical local workspace paths and records input provenance. It owns local copying/materialization of config and audio inputs, including S3 audio download when `session.inputs.audio_s3.prefix` is configured. It does not own transcript generation/processing or archive publish behavior.
`restore` should not be implemented by expanding `prepare`. It should be an app-level command that reuses shared helpers where appropriate.
### Workspace model
The local durable session workspace is campaign-aware:
```text
{workspace.root}/work/{campaign}/{session_id}/
```
It contains durable session paths such as:
```text
manifest.json
inputs/
audio/
transcripts/
artifacts/
reports/
logs/
config/
current/
runs/
```
Run-local sandboxes live below:
```text
runs/{run_id}/
```
Restore should target durable session-level paths, not old run-local stage sandboxes.
### Storage boundary
The storage adapter owns object-store primitives only: `List`, `Download`, `Upload`, and `Exists`.
The storage adapter must not infer root prefixes, campaign names, session IDs, run IDs, or archive layout. Restore code must construct full bucket-relative keys before calling storage.
### Archive commit boundary
A remote run is current only after the archive stage has uploaded the run record, promoted outputs, `current/manifest.json`, and finally `current/run_id.txt`.
`current/run_id.txt` is the final remote commit marker and must be written last.
Restore must not treat incomplete, skipped, failed, or uncommitted archive attempts as current remote state.
## User-facing command
Add:
```bash
narratio restore [flags]
```
The command should use the same configuration/session discovery conventions as `run`, `plan`, `resume`, and `run-stage` where practical:
```bash
narratio restore --config /path/to/pipeline.yml --session ./session.yml --session-id 2026-04-04
```
Required effective inputs:
- resolved pipeline config;
- resolved session config;
- `session.campaign`;
- `session.session_id`;
- configured remote storage backend.
Supported flags:
```text
--config <path> Existing pipeline config path behavior.
--session <path> Existing session config path behavior.
--session-id <value> Existing session template behavior.
--dry-run Plan restore actions without writing local files.
--force Overwrite conflicting local files with remote files.
--include-audio Include archived session-level audio files.
```
Do not add `--restore` to `run` in this implementation.
## Default restore scope
By default, restore:
1. Validates and reads the current remote commit marker.
2. Downloads the current remote manifest into the local session manifest path.
3. Downloads durable transcript files.
4. Downloads durable generated artifact files.
Default included remote/local durable paths:
```text
manifest.json from remote current manifest
transcripts/**
artifacts/**
```
Default excluded paths:
```text
audio/** unless --include-audio is passed
runs/** always excluded for this implementation
logs/** excluded for this implementation
reports/** excluded for this implementation unless needed for current manifest validation
config/** excluded for this implementation
inputs/** excluded for this implementation
current/** remote control metadata only; do not mirror blindly
```
If the existing archive implementation stores promoted files in a different remote layout, use the existing archive/path helpers and current archive semantics rather than inventing a parallel layout.
## Remote state discovery
Implement restore around the current committed archive state.
Expected algorithm:
1. Resolve pipeline/session config.
2. Ensure storage is configured.
3. Ensure local workspace layout exists.
4. Acquire the session lock.
5. Build the remote session archive prefix using the same helpers/policy used by archive code.
6. Check for the remote `current/run_id.txt` commit marker.
7. Read the committed run ID.
8. Download `current/manifest.json` to a temporary file.
9. Validate that the manifest is parseable and belongs to the requested campaign/session.
10. Build a restore plan from the committed remote state.
11. Execute the restore plan unless `--dry-run` is set.
12. Emit a concise summary.
Important: `current/run_id.txt` is the commit marker. Do not restore from a remote session prefix merely because files exist under `transcripts/` or `artifacts/`.
## Restore planning
Create a planning layer before writing files.
A restore plan entry should include at least:
```go
type RestoreAction struct {
Kind RestoreActionKind
RemoteKey string
LocalPath string
Size int64
ETag string
ExistsLocal bool
SameLocal bool
Conflict bool
Reason string
}
```
Suggested action kinds:
```text
download
skip_same
skip_missing_optional
conflict
```
The restore planner should be deterministic:
- sort remote objects by key;
- sort planned actions by local path or stable restore priority;
- write/report stable output for tests.
## Conflict and overwrite policy
Default behavior should be safe.
For each planned file:
```text
local absent:
download
local present and same as remote:
skip
local present and different:
conflict; fail restore unless --force is set
--force:
overwrite local conflicting files with remote versions
--dry-run:
do not write any files; report what would happen
```
The first implementation may use size and checksum/hash comparison where available. If remote ETag cannot be treated as a content hash, compare by downloading to a temporary file and hashing locally before deciding whether a local file is the same. Prefer correctness over assuming provider-specific ETag semantics.
Do not delete local files that are not present remotely.
## File writing and transactionality
Restore should avoid partial writes.
Implementation requirements:
- download each remote object to a temporary file under the session workspace or OS temp dir;
- validate downloaded content where possible before replacing local files;
- create parent directories as needed;
- atomically rename/copy into place only after successful download;
- do not overwrite local files unless `--force` is set;
- if a later file fails, preserve already-restored files but return a failure summary;
- never corrupt an existing local manifest on failed manifest download/parse.
Manifest restore is especially sensitive:
- download remote `current/manifest.json` to a temporary file;
- parse and validate it;
- if no local manifest exists, install it;
- if a local manifest exists and is equivalent, skip;
- if a local manifest exists and differs, fail unless `--force` is set;
- with `--force`, replace the local manifest with the remote manifest after validation;
- do not attempt a manifest merge in the initial implementation.
## Manifest semantics
`restore` is not a pipeline run and should not mark stages as running/succeeded/failed.
The restored remote manifest becomes the local session manifest. That is what allows a subsequent command such as:
```bash
narratio run-stage --force analyze
```
to see existing upstream stage state and canonical durable outputs.
Do not create a new run manifest for `restore`.
It is acceptable to write a restore diagnostic report outside the manifest, for example:
```text
reports/restore-latest.json
```
or a timestamped report, if that pattern fits the existing codebase. The report must not contain secrets.
## Local workspace locking
`restore` should acquire the same session lock used by ordinary pipeline operations before modifying session workspace state.
If the lock is held, fail fast with the same lock-conflict behavior used elsewhere.
`--dry-run` may still acquire the lock for consistency, but it is acceptable to avoid the lock if the codebase already has a clear read-only command pattern. Prefer safety and simplicity.
## Audio behavior
By default, do not restore audio.
If `--include-audio` is passed:
- restore archived durable session-level audio files only;
- do not use run-scoped spool paths;
- do not mutate or delete spool state;
- do not infer original `session.inputs.audio_s3.prefix` behavior;
- respect the same conflict/force/dry-run behavior used for transcripts/artifacts.
If the archive does not contain durable audio files, `--include-audio` should report that no archived audio was found rather than failing, unless the final implementation chooses to treat explicit audio restore as required. Prefer non-failure for absent archived audio unless tests or existing archive semantics suggest otherwise.
## Remote object selection
Prefer using manifest/artifact metadata when it reliably identifies durable outputs.
Also support listing committed durable archive prefixes so restore can retrieve all top-level session artifacts that may not yet be fully represented in manifest metadata.
The implementation should inspect existing archive code before choosing the final object-selection method. Do not duplicate archive path construction.
Recommended selection priority:
1. Remote current manifest path.
2. Durable promoted transcript/artifact outputs recorded in the manifest or archive metadata, if available.
3. Objects under committed durable `transcripts/` and `artifacts/` archive prefixes.
4. Objects under durable `audio/` only when `--include-audio` is passed.
Always exclude:
```text
runs/**
```
for the first implementation.
## Package and file organization
Expected areas to inspect and update:
```text
cmd/narratio/
internal/app/
internal/adapters/storage/
internal/artifacts/
internal/manifest/
docs/
examples/
```
Suggested implementation shape:
```text
internal/app/restore.go
internal/app/restore_test.go
internal/archive/restore/
planner.go
executor.go
report.go
keys.go
*_test.go
```
The exact package name may vary. Use whatever best fits the existing repository, but keep these boundaries clear:
- `internal/app` owns CLI command handling, config/session loading, lock acquisition, and wiring.
- Restore planning/execution owns remote key discovery, conflict detection, downloads, and reporting.
- `internal/adapters/storage` remains a transport boundary only.
- Workspace/path helpers remain centralized; do not scatter string concatenation.
If the repository already has an `internal/archive` or archive-stage helper package, prefer extending that rather than creating a conflicting package layout.
## CLI output
`narratio restore` should print a concise operator summary.
Example successful output:
```text
Restored session archive for sample-campaign/2026-04-04
Remote run: 20260504T031500Z-a1b2c3
Downloaded: 4
Skipped unchanged: 2
Conflicts: 0
```
Example dry run:
```text
Restore plan for sample-campaign/2026-04-04
Remote run: 20260504T031500Z-a1b2c3
Would download: transcripts/processed.json
Would download: transcripts/trimmed.json
Would skip unchanged: artifacts/session_recap.md
```
Example conflict:
```text
restore conflict: local artifacts/session_recap.md differs from remote archive; rerun with --force to overwrite
```
Do not print transcript or artifact content.
## Error behavior
Fail clearly when:
- storage backend is not configured;
- S3 bucket/config is missing or invalid;
- remote current commit marker is missing;
- remote current manifest is missing;
- remote manifest is invalid;
- remote manifest does not match requested campaign/session;
- local file differs from remote and `--force` is not set;
- a required remote object download fails;
- a local path would escape the session workspace;
- a remote key maps to an unsafe local path.
Skip or report non-fatal conditions when:
- optional audio restore finds no archived audio;
- an included prefix has no objects;
- a local file already matches the remote file.
## Path safety
Every restored file must map to a safe path under the session root.
Validation rules:
- local restore paths must be relative to the session root;
- reject absolute paths;
- reject `..` traversal;
- reject paths that escape through symlinks if the codebase has symlink-safe path checks;
- do not restore remote keys directly without mapping/classification;
- do not mirror arbitrary remote keys.
## Testing plan
Add focused unit tests. Do not require live S3.
### CLI tests
Add or update `internal/app` command tests for:
- `narratio restore --help`;
- restore accepts `--config`, `--session`, and `--session-id`;
- restore accepts `--dry-run`;
- restore accepts `--force`;
- restore accepts `--include-audio`;
- restore fails when storage is not configured;
- restore does not run pipeline stages.
### Restore planner tests
Test:
- missing `current/run_id.txt` fails;
- missing `current/manifest.json` fails;
- invalid manifest fails;
- wrong campaign/session manifest fails;
- default scope includes manifest/transcripts/artifacts;
- default scope excludes audio/logs/reports/config/runs;
- `--include-audio` includes durable audio;
- run-local keys are excluded;
- keys are sorted deterministically;
- unsafe remote-to-local paths are rejected.
### Conflict policy tests
Test:
- absent local file downloads;
- matching local file skips;
- differing local file conflicts by default;
- `--force` overwrites conflicts;
- `--dry-run` writes nothing;
- partial failure does not corrupt an existing local manifest.
### Storage/fake tests
Use fake storage to simulate:
- object listing;
- object download;
- missing objects;
- download failures;
- metadata/ETag behavior.
### Workspace/lock tests
Test:
- session layout is created before restore;
- session lock conflict fails;
- restored files land under the expected campaign/session workspace;
- no files are written outside the session root.
### Follow-up command workflow test
Add at least one test that simulates:
```bash
narratio restore --session-id 2026-04-04
narratio run-stage --force analyze
```
The test does not need to run real Scriptorium. Use existing fake/stub behavior to verify that restored transcripts and manifest state are sufficient for analyze-stage input resolution.
## Documentation updates when implemented
When the feature is implemented, update current-behavior docs:
```text
docs/cli.md
docs/operations.md
docs/internal/storage.md or docs/internal/archive/restore.md
```
If the documentation set does not yet have an internal restore document, add one consistent with the existing internal-doc style:
```text
docs/internal/command-restore.md
```
or:
```text
docs/internal/archive-restore.md
```
Do not document future `narratio run --restore` behavior outside `docs/roadmap/` until implemented.
## Implementation phases
### Phase 1: Audit existing archive and path helpers (completed)
Before coding behavior, inspect:
```text
internal/app/
internal/stage/archive*
internal/adapters/storage/
internal/artifacts/
internal/manifest/
docs/internal/stage-archive.md, if present
```
Determine:
- exact remote archive key layout;
- how root prefix/campaign/session are modeled;
- how current commit marker keys are built;
- how current manifest is uploaded;
- where promoted outputs are uploaded;
- whether helper functions already exist for remote archive keys;
- whether local workspace path helpers can safely map restore destinations.
Deliverable:
- small code comments or internal helper selection;
- no large behavior change yet unless required by tests.
### Phase 2: Add CLI surface and command wiring (completed)
Add `narratio restore` command parsing.
Wire flags:
```text
--config
--session
--session-id
--dry-run
--force
--include-audio
```
Use the existing config/session load path where practical.
Deliverable:
- command exists;
- help output is sensible;
- command validates basic inputs;
- command returns a clear “not yet implemented” or calls an empty planner if phased commits are desired;
- CLI tests pass.
### Phase 3: Implement remote current-state discovery (completed)
Add restore code that:
- creates an object store from resolved config;
- builds remote current marker key;
- reads `current/run_id.txt`;
- reads/downloads `current/manifest.json`;
- validates manifest identity;
- returns remote current-state metadata.
Deliverable:
- fake-storage tests for current-state discovery;
- no local file writes beyond temporary files.
### Phase 4: Implement restore planning (completed)
Build deterministic restore plans for default scope and `--include-audio`.
Deliverable:
- plan lists manifest, transcript, artifact files;
- plan excludes run-local data;
- plan detects local same/conflict/missing states;
- dry-run output works;
- no real file overwrite yet except temp comparisons as needed.
### Phase 5: Implement restore execution (completed)
Execute the plan safely:
- create directories;
- download to temporary files;
- validate content where practical;
- atomically install files;
- enforce default conflict failure;
- support `--force`;
- preserve existing manifest unless safe to replace.
Deliverable:
- restore works end-to-end against fake storage;
- failures are clear and do not corrupt existing local manifest.
### Phase 6: Add restore report and operator summary (completed)
Add concise stdout summary and optional JSON restore report if consistent with project diagnostics.
Deliverable:
- user-friendly output;
- durable diagnostic report if implemented;
- no content leakage.
### Phase 7: Workflow integration test (completed)
Add a test for restoring a previous session and then forcing `analyze`.
Deliverable:
- restored manifest/transcripts/artifacts are sufficient for analyze input resolution;
- no upstream stages rerun;
- no reliance on live subprocesses or S3.
### Phase 8: Documentation update (completed)
Once implemented, update current-behavior docs and internal command docs.
Also leave future `narratio run --restore` in roadmap only.
## Definition of done
The feature is complete when:
- `narratio restore` exists and is documented.
- It uses the same config/session discovery semantics as other commands where practical.
- It requires configured remote storage.
- It restores only from a committed current archive state.
- It restores the current manifest, transcripts, and artifacts by default.
- It restores audio only with `--include-audio`.
- It excludes run-local sandboxes.
- It fails on local/remote conflicts by default.
- `--force` overwrites conflicts.
- `--dry-run` writes nothing.
- It uses fake storage in tests.
- It does not change `prepare` behavior.
- It does not implement `narratio run --restore`.
- It avoids AWS SDK leakage outside the storage adapter.
- It uses centralized path/key helpers rather than scattered string concatenation.
- `go test ./...` passes.
## Suggested test commands
Run focused tests first:
```bash
go test ./internal/app -run TestExecute -v
go test ./internal/adapters/storage -v
go test ./internal/artifacts -v
go test ./internal/manifest -v
```
Then run the full suite:
```bash
go test ./...
```
## Suggested commit message
```text
Add restore subcommand roadmap
```

View File

@@ -1,762 +0,0 @@
# Roadmap: Runtime-Defined Scriptorium Artifacts
## Status
Implementation roadmap for a pre-release hard cutover.
## Purpose
Narratio currently treats artifact generation as a narrow `analyze` stage that supports a hard-coded `session_recap` artifact. This roadmap describes how to generalize artifact generation so operators can define Scriptorium-backed output artifacts at runtime through `pipeline.yml`.
The goal is to keep Narratio as a fixed pipeline orchestrator while making the artifact generation step configurable, composable, deterministic, and easy to regenerate selectively.
## Desired Outcome
Operators should be able to define artifacts such as session recaps, player handouts, NPC summaries, quest logs, entity maps, or other campaign-specific outputs without changing Narratio code.
A configured artifact is declared under:
```text
pipeline.scriptorium.artifacts.<name>
```
Each configured artifact becomes a canonical runtime artifact source ID:
```text
narratio.artifact.<name>
```
For example:
```yaml
scriptorium:
artifacts:
session_recap:
enabled: true
prompt_id: dnd_session.session_recap
output_path: artifacts/session_recap.md
inputs:
transcript:
source: narratio.transcript.trimmed
required: true
```
This artifact is addressable by later artifacts as:
```text
narratio.artifact.session_recap
```
A dependent artifact can then consume it explicitly:
```yaml
scriptorium:
artifacts:
player_handout:
enabled: true
depends_on:
- session_recap
prompt_id: dnd_session.player_handout
output_path: artifacts/player_handout.md
inputs:
recap:
source: narratio.artifact.session_recap
required: true
transcript:
source: narratio.transcript.trimmed
required: true
```
## Resolved Design Decisions
The following decisions are settled for the initial implementation:
1. Configured artifact outputs must live under Narratio's internal artifact output directory, initially `artifacts/`.
2. The artifact output directory should be defined as an internal default in `internal/config/defaults.go`, but no public configuration knob should be exposed yet.
3. Artifact `output_path` should remain explicit in the initial implementation to avoid guessing file extensions or output formats.
4. A disabled artifact may still be referenced as an input if its declared output already exists on disk and passes basic validation.
5. A disabled artifact is not executable during the current analyze run.
6. Artifact-to-artifact references require an explicit `depends_on` entry. Narratio should fail fast if the dependency declaration is missing.
7. The manifest remains stage-oriented: `analyze` succeeds or fails as a full stage.
8. Analyze-stage metadata may record per-artifact output details for provenance and later resolution, but not for intra-stage resume semantics.
9. `--artifacts` should be added as a CLI filter for selective artifact generation.
10. `--artifacts` does not imply `--force`; it only changes which configured artifacts are treated as executable when `analyze` actually runs.
11. Because Narratio is still pre-release, the hard-coded `session_recap` behavior should be removed immediately rather than deprecated gradually.
## Scope
This roadmap covers:
- introducing a runtime artifact catalog;
- generalizing configured Scriptorium artifact execution;
- supporting `narratio.artifact.<name>` source IDs;
- adding explicit artifact dependencies;
- supporting disabled-but-resolvable artifact inputs;
- adding selective artifact execution via `--artifacts`;
- recording generated artifacts in analyze-stage metadata and/or manifest outputs;
- removing hard-coded `session_recap` behavior;
- updating tests and documentation.
## Non-Goals
This feature should not turn Narratio into a general workflow engine.
The initial implementation should not add:
- arbitrary shell-command artifacts;
- arbitrary user-defined stages;
- loops or conditional branching;
- automatic archive promotion of generated artifacts;
- semantic knowledge of particular artifact types;
- per-artifact resume semantics within a successful or failed analyze stage;
- automatic dependency inference without `depends_on`.
Narratio should continue to orchestrate a fixed pipeline. The configurable part is the set of Scriptorium artifact invocations performed during the `analyze` stage.
## Current State
Narratio already has several relevant pieces in place:
- `pipeline.scriptorium.artifacts` is modeled as a map of artifact definitions.
- The Scriptorium adapter already accepts generic run/render requests.
- The artifact resolver already understands canonical artifact source IDs.
- The `analyze` stage already resolves inputs, optionally runs render-debug, invokes Scriptorium, verifies output, and records metadata.
The main limitation is that `analyze` currently treats `session_recap` as the only executable artifact and rejects other enabled artifact definitions.
## Target Architecture
### Runtime Artifact Catalog
Introduce a per-run artifact catalog that tracks built-in artifacts and configured artifacts.
Conceptually:
```text
ArtifactCatalog
├── built-in artifacts
│ ├── narratio.transcript.merged
│ ├── narratio.transcript.polished
│ ├── narratio.transcript.full
│ ├── narratio.transcript.trimmed
│ └── narratio.bounds.session
└── configured artifacts
├── narratio.artifact.session_recap
├── narratio.artifact.player_handout
└── narratio.artifact.npc_summary
```
The catalog should distinguish between three states:
```text
planned valid configured or built-in artifact known to Narratio
available artifact has been produced or otherwise resolved
executable configured artifact selected for execution in this analyze run
```
Configured artifacts can be planned without being executable. This distinction is important for disabled artifacts and for `--artifacts` filtering.
### Configured Artifact Source IDs
Configured artifact keys map directly to source IDs:
```text
pipeline.scriptorium.artifacts.<name>
→ narratio.artifact.<name>
```
`session_recap` should no longer be a special built-in analyze artifact. Instead, it is just a conventional configured artifact key:
```yaml
scriptorium:
artifacts:
session_recap:
enabled: true
prompt_id: dnd_session.session_recap
output_path: artifacts/session_recap.md
```
`narratio.artifact.session_recap` remains valid only because `session_recap` is configured.
### Artifact Output Directory
Add an internal default artifact output directory, initially:
```text
artifacts
```
This default should live in `internal/config/defaults.go` or the existing equivalent defaults location.
For the initial implementation:
- expose no public config knob for the artifact output directory;
- require each configured artifact to provide an explicit `output_path`;
- validate that each configured artifact `output_path` is run-relative;
- validate that each configured artifact `output_path` is under the internal artifact output directory;
- reject output paths that escape the run workspace or use path traversal.
This preserves future configurability without forcing Narratio to guess output extensions or formats now.
### Enabled, Disabled, and Selected Artifacts
Configured artifacts should have three distinct execution states:
```text
enabled by config artifact has enabled: true
selected for execution artifact remains executable after --artifacts filtering
disabled for execution artifact is not executable, but may be resolvable from disk
```
Without `--artifacts`, all configured artifacts with `enabled: true` are selected for execution.
With `--artifacts`, only the named artifacts are selected for execution. All other configured artifacts are treated as disabled for the current analyze invocation, regardless of their configured `enabled` value.
Disabled artifacts may still be resolved as inputs if their configured `output_path` exists on disk and passes validation.
### Disabled Artifact Resolution
If artifact `B` references artifact `A`, and `A` is disabled for execution, Narratio should attempt to resolve `A` from disk.
This should succeed only when:
1. `A` is defined in `pipeline.scriptorium.artifacts`;
2. `A` has a valid `output_path`;
3. the output path exists in the current run workspace;
4. the output is non-empty, or otherwise passes any available artifact-specific validation.
The resolved provenance should make the source clear, for example:
```text
filesystem.disabled_artifact_output
```
If the file does not exist or fails validation, the dependent artifact should fail before invoking Scriptorium.
Example error wording:
```text
artifact player_handout requires narratio.artifact.session_recap, but session_recap is disabled for execution and artifacts/session_recap.md does not exist
```
### Explicit Dependencies
Artifact-to-artifact references require explicit `depends_on` entries.
If artifact `B` has an input source of `narratio.artifact.A`, then `B.depends_on` must include `A`.
This should fail:
```yaml
scriptorium:
artifacts:
player_handout:
enabled: true
prompt_id: dnd_session.player_handout
output_path: artifacts/player_handout.md
inputs:
recap:
source: narratio.artifact.session_recap
required: true
```
This should pass:
```yaml
scriptorium:
artifacts:
player_handout:
enabled: true
depends_on:
- session_recap
prompt_id: dnd_session.player_handout
output_path: artifacts/player_handout.md
inputs:
recap:
source: narratio.artifact.session_recap
required: true
```
`depends_on` values refer to configured artifact keys, not full source IDs.
Dependency validation should fail on:
- references to unknown artifact keys;
- missing `depends_on` entries for artifact-to-artifact input references;
- self-dependencies;
- dependency cycles among executable artifacts.
Dependencies on disabled artifacts are permitted, but the disabled dependency must resolve from disk before the dependent artifact runs.
### Execution Order
The analyze stage should execute selected artifacts in dependency order.
Rules:
- selected artifacts are executable;
- disabled artifacts are never executed;
- selected artifacts may depend on other selected artifacts;
- selected artifacts may depend on disabled artifacts if those disabled artifacts resolve from disk;
- independent selected artifacts run in deterministic sorted-name order.
Use topological sorting over selected artifacts, while validating dependency references across the full configured artifact set.
### Input Resolution
Input resolution should use the artifact catalog and existing artifact resolver behavior.
For each configured artifact input:
- built-in sources resolve through existing resolver behavior;
- `previous_session_artifact` preserves existing behavior;
- `narratio.artifact.<name>` resolves through the runtime artifact catalog;
- selected dependencies resolve after being produced earlier in the same analyze execution;
- disabled dependencies resolve from their configured output path on disk;
- optional missing inputs are omitted;
- required missing inputs fail before Scriptorium is invoked.
### Analyze Stage Generalization
The `analyze` stage should become the generic Scriptorium artifact stage.
High-level flow:
1. Load configured Scriptorium artifacts.
2. Apply the `--artifacts` filter, if present.
3. If no artifacts are selected for execution, return success metadata with `skipped=true`.
4. Build the runtime artifact catalog.
5. Validate artifact names, output paths, source IDs, dependencies, selected artifacts, and required fields.
6. Resolve any disabled dependencies that are required by selected artifacts.
7. Sort selected artifacts by dependency order.
8. For each selected artifact:
- resolve configured inputs;
- build the Scriptorium run request;
- optionally run Scriptorium render-debug;
- run Scriptorium;
- fail on validation-failed result;
- verify the output exists and is non-empty;
- record artifact output metadata;
- register `narratio.artifact.<name>` as available in the catalog.
9. Return aggregate analyze-stage metadata containing all generated and reused artifacts relevant to the run.
The Scriptorium adapter should remain generic. It should not decide which artifacts run, how dependencies work, or how artifacts are registered.
### Manifest and Metadata
The manifest should remain stage-oriented.
This means:
- `analyze` succeeds or fails as a full stage;
- if `analyze` has already succeeded and the user does not force it, the runner skips it as a full stage;
- Narratio should not implement per-artifact resume in the first version.
However, analyze-stage metadata should still record artifact outputs for provenance and future resolution.
Recommended metadata shape:
```json
{
"skipped": false,
"artifacts": [
{
"name": "session_recap",
"source_id": "narratio.artifact.session_recap",
"output_kind": "scriptorium_artifact",
"path": "artifacts/session_recap.md",
"prompt_id": "dnd_session.session_recap",
"profile_id": "local-gemma-31b",
"provenance": "generated.current_analyze_run"
},
{
"name": "player_handout",
"source_id": "narratio.artifact.player_handout",
"output_kind": "scriptorium_artifact",
"path": "artifacts/player_handout.md",
"prompt_id": "dnd_session.player_handout",
"profile_id": "local-gemma-31b",
"provenance": "generated.current_analyze_run"
}
],
"reused_artifacts": [
{
"name": "session_recap",
"source_id": "narratio.artifact.session_recap",
"path": "artifacts/session_recap.md",
"provenance": "filesystem.disabled_artifact_output"
}
]
}
```
The exact struct can differ from this example, but it should preserve:
- artifact name;
- canonical source ID;
- output path;
- prompt/profile provenance for generated artifacts;
- reused-vs-generated provenance.
### Resume and Force Behavior
Keep resume behavior stage-level.
Recommended semantics:
```text
No --force, analyze already succeeded:
runner skips analyze, regardless of --artifacts.
--force, no --artifacts:
analyze regenerates all configured artifacts with enabled: true.
--force --artifacts player_handout:
analyze treats only player_handout as executable.
all other configured artifacts are disabled for execution.
disabled dependencies may be reused from disk.
--artifacts player_handout on a not-yet-completed analyze stage:
analyze runs only player_handout.
disabled dependencies may be reused from disk.
```
`--artifacts` should not imply `--force`. It is an execution filter, not a resume override.
### `--artifacts` CLI Flag
Add an `--artifacts` flag to commands that can execute or resume the analyze stage.
The flag should accept one or more configured artifact names. Internally, normalize values to a set of artifact keys.
Recommended behavior:
- validate all requested artifact names against `pipeline.scriptorium.artifacts`;
- reject unknown artifact names before running stages;
- treat requested artifacts as the only executable artifacts for the analyze stage;
- treat all other configured artifacts as disabled for execution;
- allow disabled artifacts to satisfy dependencies from disk as described above;
- if `--artifacts` is used while executing a stage other than `analyze`, either reject it or ignore it with a clear validation error. Prefer rejection.
The exact CLI parsing style can follow Narratio's existing conventions. Both comma-separated and repeatable values are acceptable if the CLI package supports them cleanly, but the internal representation should be a set of artifact keys.
### Archive Behavior
Do not automatically archive every generated artifact.
Artifact generation and archive promotion should remain separate concerns. Operators should continue to use `archive.promote_artifacts` to decide which generated files should be promoted or uploaded.
Example:
```yaml
archive:
promote_artifacts:
- from: artifacts/session_recap.md
to: artifacts/session_recap.md
required: true
- from: artifacts/player_handout.md
to: artifacts/player_handout.md
required: false
```
A later enhancement may add opt-in automatic promotion of configured artifacts, but explicit promotion should remain the default.
## Implementation Plan
### Phase 1: Config Model and Defaults
Add or update the configured artifact model to include:
- `enabled`;
- `depends_on`;
- `prompt_id`;
- `profile_id`;
- `output_path`;
- `timeout`;
- `render_debug`;
- `inputs`;
- `vars`.
Add an internal default artifact output directory in `internal/config/defaults.go`, initially set to `artifacts`.
Validation rules:
- artifact names must match a conservative identifier pattern such as `^[a-z][a-z0-9_]*$`;
- selected/executable artifacts require `prompt_id` and `output_path`;
- configured artifacts that may be referenced while disabled require `output_path`;
- configured artifact output paths must be run-relative;
- configured artifact output paths must live under the internal artifact output directory;
- configured artifact output paths must not escape the run workspace;
- `narratio.artifact.<name>` input sources must refer to configured artifact keys;
- any `narratio.artifact.<name>` input source must have a matching `depends_on` entry;
- `depends_on` entries must refer to configured artifact keys;
- dependencies must not contain self-references or executable cycles;
- input names and var names must remain compatible with the Scriptorium adapter's validation rules;
- unknown YAML fields must continue to fail strict decode.
Tests:
- valid single configured artifact;
- valid multiple independent artifacts;
- valid artifact-to-artifact dependency;
- valid dependency on disabled artifact with output path;
- invalid artifact name;
- missing required fields;
- output path outside `artifacts/`;
- dependency on missing artifact;
- missing `depends_on` for artifact input source;
- self-dependency;
- cycle detection;
- typo in `narratio.artifact.<name>` source;
- unknown YAML fields still fail strict decode.
### Phase 2: CLI Filtering
Add the `--artifacts` flag and carry the selected artifact set into the run execution options.
Implementation notes:
- parse values according to existing CLI conventions;
- normalize to artifact key strings;
- validate against configured artifact definitions after config load;
- make the selected set available to the analyze stage;
- reject use with commands or stages where analyze cannot run.
Tests:
- no `--artifacts` means all enabled artifacts are selected;
- one requested artifact is selected;
- multiple requested artifacts are selected;
- unknown requested artifact fails;
- `--artifacts` does not imply `--force`;
- `--artifacts` with already-succeeded analyze stage is skipped unless forced;
- `--artifacts` on unsupported stage command fails clearly.
### Phase 3: Runtime Artifact Catalog
Introduce an internal artifact catalog abstraction.
Responsibilities:
- register built-in artifact definitions;
- register configured artifact definitions;
- map configured artifact keys to `narratio.artifact.<name>` IDs;
- track planned, available, and executable artifact states;
- expose lookup by canonical source ID;
- record generated provenance;
- record disabled-from-disk provenance.
Keep the catalog narrow. It should not execute Scriptorium and should not understand prompt semantics.
Tests:
- built-in source lookup;
- configured source registration;
- duplicate/conflicting source handling;
- planned but unavailable artifact lookup;
- selected artifact state;
- disabled artifact state;
- registering an artifact as available after generation;
- registering a disabled artifact as available from disk;
- resolving a configured artifact from analyze metadata if that behavior is implemented.
### Phase 4: Resolver Integration
Update artifact resolution so configured artifact IDs are resolved through the runtime catalog.
Resolution behavior:
- built-in sources continue using existing resolver behavior;
- configured artifact sources resolve from catalog availability/provenance;
- selected configured artifacts become available after generation;
- disabled configured artifacts may become available from disk;
- missing optional configured artifact inputs are omitted;
- missing required configured artifact inputs fail clearly.
Tests:
- configured artifact consumes a built-in transcript source;
- configured artifact consumes another configured artifact produced earlier in the same analyze run;
- configured artifact consumes a disabled artifact resolved from disk;
- required disabled artifact missing on disk fails;
- required configured artifact missing fails;
- optional missing configured artifact is omitted;
- reused artifact provenance is recorded distinctly from generated artifact provenance.
### Phase 5: Analyze Stage Generalization
Refactor `analyze` to execute selected configured artifacts.
Implementation notes:
- remove the hard-coded `session_recap` selection path;
- remove the hard-coded rejection of non-`session_recap` artifacts;
- preserve skip behavior when Scriptorium config is absent or no artifacts are selected;
- build the runtime artifact catalog;
- apply `--artifacts` filtering;
- validate selected artifacts and their dependencies;
- pre-resolve disabled dependencies from disk where required;
- compute deterministic dependency order;
- execute selected artifacts one at a time in dependency order;
- keep render-debug behavior at global and artifact levels;
- keep Scriptorium adapter invocation generic;
- after each successful run, register the artifact as available in the catalog;
- aggregate generated and reused artifact metadata.
Tests:
- no Scriptorium config skips;
- empty artifact map skips;
- no selected artifacts skips;
- disabled artifacts do not run;
- one selected artifact runs;
- multiple independent artifacts run in deterministic order;
- dependent selected artifact receives prior selected artifact as input;
- dependent selected artifact receives disabled-from-disk artifact as input;
- render-debug works for configured artifacts;
- Scriptorium validation failure fails the stage;
- missing required input fails the stage;
- successful outputs are non-empty and recorded;
- artifact filter executes only requested artifacts.
### Phase 6: Manifest and Stage Metadata
Update analyze-stage metadata and manifest output recording to support dynamic configured artifacts.
Recommended behavior:
- every generated configured artifact gets `source_id: narratio.artifact.<name>`;
- every generated configured artifact gets a generic output kind such as `scriptorium_artifact`;
- reused disabled artifacts are recorded separately from generated artifacts;
- metadata is sufficient for debugging, provenance, and future resolver support;
- metadata does not create per-artifact resume semantics.
Because this is a pre-release hard cutover, do not preserve a special legacy `session_recap` output kind unless a current internal test or archive path still requires it temporarily. Prefer updating tests and examples to treat `session_recap` as an ordinary configured artifact.
Tests:
- metadata records one generated configured artifact;
- metadata records multiple generated configured artifacts;
- metadata records reused disabled artifact provenance;
- `session_recap` is recorded as a normal configured artifact;
- manifest still treats `analyze` as a single succeeded or failed stage;
- runner skip behavior remains stage-level.
### Phase 7: Archive and Promotion Review
Review archive behavior after dynamic artifacts are recorded.
Implementation notes:
- do not automatically promote every configured artifact;
- keep `archive.promote_artifacts` explicit;
- update default or example promotion rules to use configured `session_recap` output path;
- ensure required promotion rules fail clearly when selected artifact generation did not produce a required file.
Tests:
- generated artifact can be promoted by explicit archive rule;
- required archive promotion fails if selected artifact was not generated and no file exists;
- optional archive promotion skips cleanly if file is absent;
- hard cutover does not rely on hard-coded `session_recap` generation.
### Phase 8: Documentation and Examples
Status: complete.
Update documentation after the implementation is complete.
Recommended documentation changes:
- update `docs/config.md` with the generalized artifact configuration model;
- update `docs/internal/artifacts.md` to describe the runtime artifact catalog;
- update `docs/stages/analyze.md` to describe generic Scriptorium artifact generation;
- update Scriptorium integration docs only if the adapter contract changes;
- update full annotated pipeline examples;
- add at least one example with multiple artifacts and one dependency;
- document `--artifacts` behavior and its relationship to `--force`;
- remove documentation stating that only `session_recap` is supported.
Documentation should make clear that:
- configured artifact source IDs use `narratio.artifact.<name>`;
- `depends_on` uses artifact keys, not full source IDs;
- artifact-to-artifact source references require explicit `depends_on`;
- disabled artifacts can be reused from disk when required by selected artifacts;
- `--artifacts` filters execution but does not imply `--force`;
- archive promotion remains explicit;
- per-artifact resume is not part of the initial implementation.
## Migration Strategy
Because Narratio is pre-release, perform a hard cutover.
Required changes:
1. Remove the hard-coded `session_recap` analyze behavior.
2. Require `session_recap` to be declared under `pipeline.scriptorium.artifacts.session_recap` if the operator wants a session recap.
3. Treat `narratio.artifact.session_recap` as valid only when `session_recap` is a configured artifact key.
4. Update config examples to show `session_recap` as a normal configured artifact.
5. Update tests to stop assuming that `session_recap` is a built-in analyze artifact.
6. Keep archive promotion explicit and path-based.
Example replacement config:
```yaml
scriptorium:
binary: scriptorium
config_path: /etc/scriptorium/config.yml
timeout: 10m
render_debug: false
artifacts:
session_recap:
enabled: true
prompt_id: dnd_session.session_recap
profile_id: local-gemma-31b
output_path: artifacts/session_recap.md
timeout: 20m
inputs:
transcript:
source: narratio.transcript.trimmed
required: true
prior_recap:
source: previous_session_artifact
artifact: artifacts/session_recap.md
required: false
vars:
artifact_title: Session Recap
```
## Acceptance Criteria
The feature is complete when:
- operators can define more than one enabled Scriptorium artifact in `pipeline.yml`;
- Narratio runs selected artifacts in deterministic dependency order;
- configured artifacts are addressable as `narratio.artifact.<name>`;
- one configured artifact can consume another configured artifact as an input;
- artifact-to-artifact input references require explicit `depends_on`;
- disabled artifacts can satisfy dependencies from existing on-disk outputs;
- missing required disabled artifacts fail clearly;
- optional missing inputs are omitted;
- `--artifacts` can selectively execute valid configured artifact names;
- `--artifacts` does not imply `--force`;
- render-debug behavior works for all configured artifacts;
- generated and reused artifacts are recorded in analyze-stage metadata;
- `session_recap` is no longer hard-coded and works as a normal configured artifact;
- archive promotion remains explicit;
- tests cover config validation, dependency sorting, disabled artifact resolution, resolver behavior, CLI filtering, analyze execution, archive interactions, and metadata.
## Suggested Implementation Order
1. Config model, defaults, and validation.
2. CLI parsing and propagation of `--artifacts` selection.
3. Runtime artifact catalog.
4. Resolver integration for configured artifacts.
5. Analyze stage generalization.
6. Stage metadata and manifest output recording.
7. Archive behavior review.
8. Documentation and examples.
This order keeps the most static pieces first, then moves into execution behavior once the configuration contract is explicit and well tested.

View File

@@ -6,7 +6,7 @@ Canonical operator troubleshooting guide for recurring implemented Narratio fail
## Config file discovery failure ## Config file discovery failure
Symptom: Symptom:
- `run`, `plan`, `resume`, or `run-stage` fails with config/session not found. - `run`, `plan`, `resume`, `run-stage`, or `restore` fails with config/session not found.
Likely Cause: Likely Cause:
- `pipeline.yml` or `session.yml` is missing from discovery paths. - `pipeline.yml` or `session.yml` is missing from discovery paths.
@@ -192,7 +192,7 @@ Links:
## Session lock conflict (`.lock`) ## Session lock conflict (`.lock`)
Symptom: Symptom:
- run fails with lock conflict for session workdir. - `run`, `resume`, `run-stage`, or `restore` fails with lock conflict for session workdir.
Likely Cause: Likely Cause:
- another Narratio process is running same session. - another Narratio process is running same session.
@@ -207,13 +207,104 @@ ps aux | grep narratio
``` ```
Safe Fix: Safe Fix:
- wait for active run to finish. - wait for active process to finish.
- if no process is active, remove only stale session `.lock` file. - if no process is active, remove only stale session `.lock` file.
Links: Links:
- [docs/operations.md](./operations.md) - [docs/operations.md](./operations.md)
- [docs/internal/workspace.md](./internal/workspace.md) - [docs/internal/workspace.md](./internal/workspace.md)
## Restore remote current pointer or manifest missing
Symptom:
- `restore` fails with remote current pointer or current manifest errors.
Likely Cause:
- `current/run_id.txt` was never published.
- `current/manifest.json` is missing for the session prefix.
- archive commit did not complete.
Diagnostics:
```bash
narratio restore --config /path/to/pipeline.yml --session /path/to/session.yml --session-id 2026-04-04 --dry-run
```
Safe Fix:
- verify archive stage succeeded for the target session.
- rerun/archive from a healthy source workspace so current pointers are published.
Links:
- [docs/operations.md](./operations.md)
- [docs/internal/stage-archive.md](./internal/stage-archive.md)
## Restore manifest identity mismatch
Symptom:
- `restore` fails because remote manifest session or campaign does not match requested values.
Likely Cause:
- wrong `--session-id` or wrong session config selected.
- archive prefix points to a different campaign/session.
Diagnostics:
```bash
narratio restore --config /path/to/pipeline.yml --session /path/to/session.yml --session-id 2026-04-04 --dry-run
```
Safe Fix:
- use the correct session config and `--session-id`.
- verify campaign/session identity in local config before restore.
Links:
- [docs/config.md](./config.md)
- [docs/operations.md](./operations.md)
## Restore conflict without `--force`
Symptom:
- `restore` fails with `restore conflict` and conflict counts.
Likely Cause:
- local durable file differs from remote file for one or more planned restore paths.
Diagnostics:
```bash
narratio restore --config /path/to/pipeline.yml --session /path/to/session.yml --session-id 2026-04-04 --dry-run
```
Safe Fix:
- review planned conflicts.
- rerun with `--force` only when remote state should overwrite local state.
Links:
- [docs/cli.md](./cli.md)
- [docs/operations.md](./operations.md)
## Restore report expectations
Symptom:
- operator expects restore report file but does not find one.
Likely Cause:
- restore was executed in `--dry-run` mode.
- restore failed before report persistence path (for example lock acquisition failure).
Diagnostics:
```bash
ls -l {workspace.root}/work/{campaign}/{session_id}/reports/restore-latest.json
```
Safe Fix:
- run non-dry-run restore for durable report output.
- resolve lock or early preflight failures and retry.
Links:
- [docs/operations.md](./operations.md)
## Secrets env-dir or credential-env failure ## Secrets env-dir or credential-env failure
Symptom: Symptom:
@@ -280,7 +371,7 @@ narratio run-stage --config /path/to/pipeline.yml --session /path/to/session.yml
Safe Fix: Safe Fix:
- rerun or resume upstream stages to generate required files. - rerun or resume upstream stages to generate required files.
- adjust promotion rules to match files that must exist. - adjust promotion `source`/`dest` rules to match artifacts that must exist.
- retry after storage issue is resolved. - retry after storage issue is resolved.
Links: Links:

View File

@@ -40,16 +40,16 @@ archive:
# Optional booleans; defaults are true. # Optional booleans; defaults are true.
enabled: true enabled: true
upload_run: true upload_run: true
# Optional promotion rules; required files fail archive if missing. # Optional promotion rules; sources use Narratio artifact source IDs.
promote_artifacts: promote_artifacts:
- from: transcripts/trimmed.json - source: narratio.transcript.trimmed
to: transcripts/trimmed.json dest: transcripts/trimmed.json
required: true required: true
- from: artifacts/session_recap.md - source: narratio.artifact.session_recap
to: artifacts/session_recap.md dest: artifacts/session_recap.md
required: true required: true
- from: artifacts/player_handout.md - source: narratio.artifact.player_handout
to: artifacts/player_handout.md dest: artifacts/player_handout.md
required: false required: false
whisperx: whisperx:

View File

@@ -19,14 +19,14 @@ archive:
enabled: true enabled: true
upload_run: true upload_run: true
promote_artifacts: promote_artifacts:
- from: transcripts/trimmed.json - source: narratio.transcript.trimmed
to: transcripts/trimmed.json dest: transcripts/trimmed.json
required: true required: true
- from: artifacts/session_recap.md - source: narratio.artifact.session_recap
to: artifacts/session_recap.md dest: artifacts/session_recap.md
required: true required: true
- from: artifacts/player_handout.md - source: narratio.artifact.player_handout
to: artifacts/player_handout.md dest: artifacts/player_handout.md
required: false required: false
whisperx: whisperx:

View File

@@ -7,7 +7,7 @@ import (
"strings" "strings"
) )
var supportedCommands = []string{"run", "plan", "status", "resume", "run-stage"} var supportedCommands = []string{"run", "plan", "status", "resume", "run-stage", "restore"}
// Execute dispatches CLI commands and returns a process exit code. // Execute dispatches CLI commands and returns a process exit code.
func Execute(args []string, stdout, stderr io.Writer) int { func Execute(args []string, stdout, stderr io.Writer) int {
@@ -32,6 +32,8 @@ func Execute(args []string, stdout, stderr io.Writer) int {
err = Resume(ctx, cmdArgs, stdout) err = Resume(ctx, cmdArgs, stdout)
case "run-stage": case "run-stage":
err = RunStage(ctx, cmdArgs, stdout) err = RunStage(ctx, cmdArgs, stdout)
case "restore":
err = Restore(ctx, cmdArgs, stdout)
default: default:
fmt.Fprintf(stderr, "unknown command: %q\n\n", cmd) fmt.Fprintf(stderr, "unknown command: %q\n\n", cmd)
printUsage(stderr) printUsage(stderr)

View File

@@ -195,7 +195,7 @@ func TestPostArchiveCleanupNotRunWhenPromotionIsMissing(t *testing.T) {
cfg.Pipeline.Spool.DeleteAudioAfterArchive = true cfg.Pipeline.Spool.DeleteAudioAfterArchive = true
cfg.Pipeline.Workspace.CleanupAfterArchive = true cfg.Pipeline.Workspace.CleanupAfterArchive = true
cfg.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{ cfg.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{
{From: "artifacts/missing.md", To: "artifacts/missing.md", Required: boolPtr(true)}, {Source: "narratio.transcript.merged", Dest: "transcripts/merged.json", Required: boolPtr(true)},
} }
archiveStageImpl, err := stage.Select("archive") archiveStageImpl, err := stage.Select("archive")
@@ -203,7 +203,7 @@ func TestPostArchiveCleanupNotRunWhenPromotionIsMissing(t *testing.T) {
t.Fatalf("Select(archive) error = %v", err) t.Fatalf("Select(archive) error = %v", err)
} }
_, err = executeStages(context.Background(), cfg, []stage.Stage{archiveStageImpl}, RunOptions{Env: &Env{ObjectStore: &storage.FakeBackend{}}}) _, err = executeStages(context.Background(), cfg, []stage.Stage{archiveStageImpl}, RunOptions{Env: &Env{ObjectStore: &storage.FakeBackend{}}})
if err == nil || !strings.Contains(err.Error(), "required promotion source missing") { if err == nil || !strings.Contains(err.Error(), "required promotion source unavailable") {
t.Fatalf("executeStages() error = %v, want promotion-missing failure", err) t.Fatalf("executeStages() error = %v, want promotion-missing failure", err)
} }
@@ -322,8 +322,15 @@ func archiveStageCleanupFixture(t *testing.T) (*config.Config, cleanupSeed, stri
Enabled: boolPtr(true), Enabled: boolPtr(true),
UploadRun: boolPtr(true), UploadRun: boolPtr(true),
PromoteArtifacts: []config.ArchivePromotionRule{ PromoteArtifacts: []config.ArchivePromotionRule{
{From: "transcripts/trimmed.json", To: "transcripts/trimmed.json", Required: boolPtr(true)}, {Source: "narratio.transcript.trimmed", Dest: "transcripts/trimmed.json", Required: boolPtr(true)},
{From: "artifacts/session_recap.md", To: "artifacts/session_recap.md", Required: boolPtr(true)}, {Source: "narratio.artifact.session_recap", Dest: "artifacts/session_recap.md", Required: boolPtr(true)},
},
}
cfg.Pipeline.Scriptorium = &config.ScriptoriumConfig{
Artifacts: map[string]config.ScriptoriumArtifactConfig{
"session_recap": {
OutputPath: "artifacts/session_recap.md",
},
}, },
} }
writeArchiveFixtureRunFiles( writeArchiveFixtureRunFiles(
@@ -353,14 +360,14 @@ func writeArchiveFixtureRunFiles(t *testing.T, runWorkDir, sessionRoot string) {
t.Helper() t.Helper()
mustWriteFile(t, filepath.Join(runWorkDir, "prepare", "inputs", "session.yml"), "session_id: 2026-05-03\n") mustWriteFile(t, filepath.Join(runWorkDir, "prepare", "inputs", "session.yml"), "session_id: 2026-05-03\n")
mustWriteFile(t, filepath.Join(runWorkDir, "transcribe", "outputs", "transcripts", "raw", "speaker.json"), "{}\n") mustWriteFile(t, filepath.Join(runWorkDir, "transcribe", "outputs", "transcripts", "raw", "speaker.json"), "{}\n")
mustWriteFile(t, filepath.Join(runWorkDir, "trim", "outputs", "transcripts", "trimmed.json"), "{}\n") mustWriteFile(t, filepath.Join(runWorkDir, "trim", "outputs", "transcripts", "trimmed.json"), "{\"segments\":[]}\n")
mustWriteFile(t, filepath.Join(runWorkDir, "analyze", "outputs", "artifacts", "session_recap.md"), "# recap\n") mustWriteFile(t, filepath.Join(runWorkDir, "analyze", "outputs", "artifacts", "session_recap.md"), "# recap\n")
mustWriteFile(t, filepath.Join(runWorkDir, "polish", "reports", "audita.report.json"), "{}\n") mustWriteFile(t, filepath.Join(runWorkDir, "polish", "reports", "audita.report.json"), "{}\n")
mustWriteFile(t, filepath.Join(runWorkDir, "merge", "config", "seriatim.generated.yml"), "key: value\n") mustWriteFile(t, filepath.Join(runWorkDir, "merge", "config", "seriatim.generated.yml"), "key: value\n")
mustWriteFile(t, filepath.Join(runWorkDir, "logs", "audita.stderr.log"), "stderr\n") mustWriteFile(t, filepath.Join(runWorkDir, "logs", "audita.stderr.log"), "stderr\n")
mustWriteFile(t, filepath.Join(runWorkDir, "manifest.json"), "{}\n") mustWriteFile(t, filepath.Join(runWorkDir, "manifest.json"), "{}\n")
mustWriteFile(t, filepath.Join(sessionRoot, "transcripts", "trimmed.json"), "{}\n") mustWriteFile(t, filepath.Join(sessionRoot, "transcripts", "trimmed.json"), "{\"segments\":[]}\n")
mustWriteFile(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "# recap\n") mustWriteFile(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "# recap\n")
} }

152
internal/app/restore.go Normal file
View File

@@ -0,0 +1,152 @@
package app
import (
"context"
"errors"
"flag"
"fmt"
"io"
"log/slog"
"os"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/logging"
)
var newObjectStoreFromConfigFn = storage.NewObjectStoreFromConfig
var discoverRemoteCurrentStateFn = discoverRemoteCurrentState
var buildRestorePlanFn = buildRestorePlan
var executeRestorePlanFn = executeRestorePlan
// Restore validates restore CLI/config inputs and storage preflight for future restore phases.
func Restore(ctx context.Context, args []string, out io.Writer) error {
fs := flag.NewFlagSet("restore", flag.ContinueOnError)
fs.SetOutput(out)
var pipelinePath string
var sessionPath string
var sessionID string
var dryRun bool
var force bool
var includeAudio bool
fs.StringVar(&pipelinePath, "config", "", "path to pipeline.yml (optional; defaults searched)")
fs.StringVar(&sessionPath, "session", "", "path to session.yml")
fs.StringVar(&sessionID, "session-id", "", "session identifier for session.yml templates")
fs.BoolVar(&dryRun, "dry-run", false, "plan restore actions without writing local files")
fs.BoolVar(&force, "force", false, "overwrite local conflicts with remote state")
fs.BoolVar(&includeAudio, "include-audio", false, "include archived session-level audio objects")
fs.Usage = func() {
_, _ = fmt.Fprintln(out, "Usage: narratio restore [--config <path>] [--session <path>] [--session-id <value>] [--dry-run] [--force] [--include-audio]")
_, _ = fmt.Fprintln(out)
_, _ = fmt.Fprintln(out, "Flags:")
fs.PrintDefaults()
}
if err := fs.Parse(args); err != nil {
if errors.Is(err, flag.ErrHelp) {
return nil
}
return fmt.Errorf("restore: invalid flags: %w", err)
}
if fs.NArg() != 0 {
return fmt.Errorf("restore: unexpected positional arguments")
}
resolvedPipelinePath, err := resolvePipelineConfigPath(pipelinePath)
if err != nil {
return fmt.Errorf("restore: %w", err)
}
resolvedSessionPath, err := resolveSessionConfigPath(sessionPath)
if err != nil {
return fmt.Errorf("restore: %w", err)
}
cfg, err := config.LoadWithSessionOptions(resolvedPipelinePath, resolvedSessionPath, config.SessionLoadOptions{
SessionID: sessionID,
})
if err != nil {
return fmt.Errorf("restore: %w", err)
}
if err := config.Validate(cfg); err != nil {
return fmt.Errorf("restore: %w", err)
}
if _, err := loadSecretsFromConfig(cfg, logging.NewLogger(os.Stderr, slog.LevelInfo)); err != nil {
return fmt.Errorf("restore: %w", err)
}
objectStore, err := newObjectStoreFromConfigFn(ctx, cfg)
if err != nil {
return fmt.Errorf("restore: %w", err)
}
current, err := discoverRemoteCurrentStateFn(ctx, cfg, objectStore)
if err != nil {
return fmt.Errorf("restore: %w", err)
}
plan, err := buildRestorePlanFn(ctx, cfg, current, objectStore, RestorePlanOptions{
IncludeAudio: includeAudio,
Force: force,
DryRun: dryRun,
})
if err != nil {
return fmt.Errorf("restore: %w", err)
}
report, err := newRestoreReport(current, plan, RestorePlanOptions{
IncludeAudio: includeAudio,
Force: force,
DryRun: dryRun,
})
if err != nil {
return fmt.Errorf("restore: %w", err)
}
if dryRun {
if err := writeRestoreDryRunSummary(out, report); err != nil {
return fmt.Errorf("restore: write plan output: %w", err)
}
return nil
}
artifactStore := artifacts.NewLocalStore(cfg.Pipeline.Workspace.Root)
if _, err := artifactStore.EnsureLayoutFor(cfg.Session.Campaign, cfg.Session.SessionID); err != nil {
return fmt.Errorf("restore: prepare workdir: %w", err)
}
lock, err := artifactStore.AcquireSessionLockFor(cfg.Session.Campaign, cfg.Session.SessionID)
if err != nil {
return fmt.Errorf("restore: acquire session lock: %w", err)
}
defer func() {
_ = artifactStore.ReleaseSessionLock(lock)
}()
if plan.ConflictCount > 0 && !force {
report.setFailed(fmt.Errorf("conflict: %d conflicting path(s)", plan.ConflictCount))
if _, reportErr := persistRestoreReport(artifactStore, cfg, report); reportErr != nil {
return fmt.Errorf("restore: report failure: %w", reportErr)
}
return fmt.Errorf(
"restore conflict: %d conflicting path(s); rerun with --force to overwrite (download=%d skip_same=%d conflicts=%d)",
plan.ConflictCount,
plan.DownloadCount,
plan.SkipSameCount,
plan.ConflictCount,
)
}
result, err := executeRestorePlanFn(ctx, cfg, current, plan, report, objectStore)
if err != nil {
report.setFailed(err)
if _, reportErr := persistRestoreReport(artifactStore, cfg, report); reportErr != nil {
return fmt.Errorf("restore: execute plan failed (%v) and report write failed (%v)", err, reportErr)
}
return fmt.Errorf("restore: execute plan: %w", err)
}
report.Execution.Downloaded = result.DownloadedCount
report.setSucceeded()
if _, err := persistRestoreReport(artifactStore, cfg, report); err != nil {
return fmt.Errorf("restore: write report: %w", err)
}
if err := writeRestoreSuccessSummary(out, report); err != nil {
return fmt.Errorf("restore: write summary: %w", err)
}
return nil
}

View File

@@ -0,0 +1,139 @@
package app
import (
"context"
"fmt"
"os"
"strings"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/manifest"
)
// RemoteCurrentState captures discovered committed remote archive state for one session.
type RemoteCurrentState struct {
Bucket string
SessionPrefix string
CurrentRunIDKey string
CurrentManifestKey string
RunID string
SessionID string
Campaign string
Manifest *manifest.Manifest
}
func discoverRemoteCurrentState(ctx context.Context, cfg *config.Config, store storage.ObjectStore) (*RemoteCurrentState, error) {
if cfg == nil || cfg.Pipeline == nil || cfg.Session == nil {
return nil, fmt.Errorf("resolved config with pipeline/session is required")
}
if store == nil {
return nil, fmt.Errorf("remote object store is required")
}
bucket := artifacts.ResolveArchiveBucket(cfg, nil)
if strings.TrimSpace(bucket) == "" {
return nil, fmt.Errorf("archive bucket is required")
}
sessionPrefix, err := artifacts.ResolveArchiveSessionPrefix(cfg, nil)
if err != nil {
return nil, fmt.Errorf("resolve archive session prefix: %w", err)
}
currentManifestKey, currentRunIDKey := artifacts.ResolveArchiveCurrentStateKeys(sessionPrefix)
exists, err := store.Exists(ctx, currentRunIDKey)
if err != nil {
return nil, fmt.Errorf("check remote current run pointer %q: %w", currentRunIDKey, err)
}
if !exists {
return nil, fmt.Errorf("remote current run pointer missing: %q", currentRunIDKey)
}
runIDPath, err := downloadObjectToTemp(ctx, store, currentRunIDKey, "narratio-restore-current-run-id-*.txt")
if err != nil {
return nil, fmt.Errorf("download remote current run pointer %q: %w", currentRunIDKey, err)
}
defer func() { _ = os.Remove(runIDPath) }()
runIDData, err := os.ReadFile(runIDPath)
if err != nil {
return nil, fmt.Errorf("read downloaded run pointer %q: %w", currentRunIDKey, err)
}
runID := strings.TrimSpace(string(runIDData))
if runID == "" {
return nil, fmt.Errorf("remote current run pointer %q is empty", currentRunIDKey)
}
exists, err = store.Exists(ctx, currentManifestKey)
if err != nil {
return nil, fmt.Errorf("check remote current manifest %q: %w", currentManifestKey, err)
}
if !exists {
return nil, fmt.Errorf("remote current manifest missing: %q", currentManifestKey)
}
manifestPath, err := downloadObjectToTemp(ctx, store, currentManifestKey, "narratio-restore-current-manifest-*.json")
if err != nil {
return nil, fmt.Errorf("download remote current manifest %q: %w", currentManifestKey, err)
}
defer func() { _ = os.Remove(manifestPath) }()
manifestStore := &manifest.LocalStore{}
remoteManifest, err := manifestStore.Load(ctx, manifestPath)
if err != nil {
return nil, fmt.Errorf("remote current manifest decode failed: %w", err)
}
requestedSession := strings.TrimSpace(cfg.Session.SessionID)
requestedCampaign := strings.TrimSpace(cfg.Session.Campaign)
manifestSession := strings.TrimSpace(remoteManifest.SessionID)
manifestCampaign := strings.TrimSpace(remoteManifest.Campaign)
if manifestSession != requestedSession {
return nil, fmt.Errorf(
"remote current manifest session_id %q does not match requested session_id %q",
manifestSession,
requestedSession,
)
}
if manifestCampaign == "" {
return nil, fmt.Errorf("remote current manifest campaign is required")
}
if manifestCampaign != requestedCampaign {
return nil, fmt.Errorf(
"remote current manifest campaign %q does not match requested campaign %q",
manifestCampaign,
requestedCampaign,
)
}
return &RemoteCurrentState{
Bucket: bucket,
SessionPrefix: sessionPrefix,
CurrentRunIDKey: currentRunIDKey,
CurrentManifestKey: currentManifestKey,
RunID: runID,
SessionID: manifestSession,
Campaign: manifestCampaign,
Manifest: remoteManifest,
}, nil
}
func downloadObjectToTemp(ctx context.Context, store storage.ObjectStore, key, pattern string) (string, error) {
tmp, err := os.CreateTemp("", pattern)
if err != nil {
return "", fmt.Errorf("create temp file: %w", err)
}
path := tmp.Name()
if err := tmp.Close(); err != nil {
_ = os.Remove(path)
return "", fmt.Errorf("close temp file: %w", err)
}
if err := store.Download(ctx, key, path); err != nil {
_ = os.Remove(path)
return "", err
}
return path, nil
}

View File

@@ -0,0 +1,239 @@
package app
import (
"context"
"encoding/json"
"fmt"
"strings"
"testing"
"time"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
)
func TestDiscoverRemoteCurrentStateSuccess(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
sessionPrefix, manifestKey, runIDKey := restoreDiscoveryKeys(cfg)
store.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte("20260519T010203Z-a1b2c3d4\n")})
store.SeedObject(storage.FakeObject{Key: manifestKey, Data: restoreManifestJSON(t, cfg.Session.SessionID, cfg.Session.Campaign)})
state, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err != nil {
t.Fatalf("discoverRemoteCurrentState() error = %v", err)
}
if state.RunID != "20260519T010203Z-a1b2c3d4" {
t.Fatalf("run id = %q, want 20260519T010203Z-a1b2c3d4", state.RunID)
}
if state.SessionPrefix != sessionPrefix {
t.Fatalf("session prefix = %q, want %q", state.SessionPrefix, sessionPrefix)
}
if state.CurrentRunIDKey != runIDKey {
t.Fatalf("current run id key = %q, want %q", state.CurrentRunIDKey, runIDKey)
}
if state.CurrentManifestKey != manifestKey {
t.Fatalf("current manifest key = %q, want %q", state.CurrentManifestKey, manifestKey)
}
if state.Manifest == nil {
t.Fatal("manifest is nil")
}
}
func TestDiscoverRemoteCurrentStateMissingRunPointerFails(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err == nil || !strings.Contains(err.Error(), "remote current run pointer missing") {
t.Fatalf("error = %v, want missing run pointer failure", err)
}
}
func TestDiscoverRemoteCurrentStateEmptyRunPointerFails(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
_, manifestKey, runIDKey := restoreDiscoveryKeys(cfg)
store.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte(" \n\t")})
store.SeedObject(storage.FakeObject{Key: manifestKey, Data: restoreManifestJSON(t, cfg.Session.SessionID, cfg.Session.Campaign)})
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err == nil || !strings.Contains(err.Error(), "is empty") {
t.Fatalf("error = %v, want empty run pointer failure", err)
}
}
func TestDiscoverRemoteCurrentStateMissingManifestFails(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
_, _, runIDKey := restoreDiscoveryKeys(cfg)
store.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte("20260519T010203Z-a1b2c3d4\n")})
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err == nil || !strings.Contains(err.Error(), "remote current manifest missing") {
t.Fatalf("error = %v, want missing manifest failure", err)
}
}
func TestDiscoverRemoteCurrentStateInvalidManifestFails(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
_, manifestKey, runIDKey := restoreDiscoveryKeys(cfg)
store.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte("20260519T010203Z-a1b2c3d4\n")})
store.SeedObject(storage.FakeObject{Key: manifestKey, Data: []byte("{invalid json")})
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err == nil || !strings.Contains(err.Error(), "remote current manifest decode failed") {
t.Fatalf("error = %v, want manifest decode failure", err)
}
}
func TestDiscoverRemoteCurrentStateSessionMismatchFails(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
_, manifestKey, runIDKey := restoreDiscoveryKeys(cfg)
store.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte("20260519T010203Z-a1b2c3d4\n")})
store.SeedObject(storage.FakeObject{Key: manifestKey, Data: restoreManifestJSON(t, "wrong-session", cfg.Session.Campaign)})
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err == nil || !strings.Contains(err.Error(), "does not match requested session_id") {
t.Fatalf("error = %v, want session mismatch failure", err)
}
}
func TestDiscoverRemoteCurrentStateCampaignMismatchFails(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
_, manifestKey, runIDKey := restoreDiscoveryKeys(cfg)
store.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte("20260519T010203Z-a1b2c3d4\n")})
store.SeedObject(storage.FakeObject{Key: manifestKey, Data: restoreManifestJSON(t, cfg.Session.SessionID, "wrong-campaign")})
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err == nil || !strings.Contains(err.Error(), "does not match requested campaign") {
t.Fatalf("error = %v, want campaign mismatch failure", err)
}
}
func TestDiscoverRemoteCurrentStateEmptyCampaignFails(t *testing.T) {
cfg := restoreDiscoveryConfig()
store := &storage.FakeBackend{}
_, manifestKey, runIDKey := restoreDiscoveryKeys(cfg)
store.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte("20260519T010203Z-a1b2c3d4\n")})
store.SeedObject(storage.FakeObject{Key: manifestKey, Data: restoreManifestJSON(t, cfg.Session.SessionID, "")})
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err == nil || !strings.Contains(err.Error(), "campaign is required") {
t.Fatalf("error = %v, want empty campaign failure", err)
}
}
func TestDiscoverRemoteCurrentStateUsesCurrentKeysUnderSessionPrefix(t *testing.T) {
cfg := restoreDiscoveryConfig()
sessionPrefix, manifestKey, runIDKey := restoreDiscoveryKeys(cfg)
base := &storage.FakeBackend{}
store := &captureObjectStore{delegate: base}
base.SeedObject(storage.FakeObject{Key: runIDKey, Data: []byte("20260519T010203Z-a1b2c3d4\n")})
base.SeedObject(storage.FakeObject{Key: manifestKey, Data: restoreManifestJSON(t, cfg.Session.SessionID, cfg.Session.Campaign)})
_, err := discoverRemoteCurrentState(context.Background(), cfg, store)
if err != nil {
t.Fatalf("discoverRemoteCurrentState() error = %v", err)
}
expectedRunKey := fmt.Sprintf("%scurrent/run_id.txt", sessionPrefix)
expectedManifestKey := fmt.Sprintf("%scurrent/manifest.json", sessionPrefix)
if !containsString(store.existsKeys, expectedRunKey) {
t.Fatalf("exists keys = %#v, want run pointer key %q", store.existsKeys, expectedRunKey)
}
if !containsString(store.existsKeys, expectedManifestKey) {
t.Fatalf("exists keys = %#v, want manifest key %q", store.existsKeys, expectedManifestKey)
}
if !containsString(store.downloadKeys, expectedRunKey) {
t.Fatalf("download keys = %#v, want run pointer key %q", store.downloadKeys, expectedRunKey)
}
if !containsString(store.downloadKeys, expectedManifestKey) {
t.Fatalf("download keys = %#v, want manifest key %q", store.downloadKeys, expectedManifestKey)
}
}
type captureObjectStore struct {
delegate storage.ObjectStore
existsKeys []string
downloadKeys []string
}
func (s *captureObjectStore) List(ctx context.Context, prefix string) ([]storage.ObjectInfo, error) {
return s.delegate.List(ctx, prefix)
}
func (s *captureObjectStore) Download(ctx context.Context, key, localPath string) error {
s.downloadKeys = append(s.downloadKeys, key)
return s.delegate.Download(ctx, key, localPath)
}
func (s *captureObjectStore) Upload(ctx context.Context, localPath, key string, opts storage.UploadOptions) (storage.ObjectInfo, error) {
return s.delegate.Upload(ctx, localPath, key, opts)
}
func (s *captureObjectStore) Exists(ctx context.Context, key string) (bool, error) {
s.existsKeys = append(s.existsKeys, key)
return s.delegate.Exists(ctx, key)
}
func restoreDiscoveryConfig() *config.Config {
return &config.Config{
Pipeline: &config.PipelineConfig{
Storage: config.StorageConfig{
S3: &config.StorageS3Config{
Bucket: "my-dnd-archive",
RootPrefix: "dnd",
},
},
},
Session: &config.SessionConfig{
SessionID: "2026-05-03",
Campaign: "sample-campaign",
},
}
}
func restoreDiscoveryKeys(cfg *config.Config) (sessionPrefix, manifestKey, runIDKey string) {
sessionPrefix = artifacts.S3SessionPrefix(cfg.Pipeline.Storage.S3.RootPrefix, cfg.Session.Campaign, cfg.Session.SessionID)
manifestKey, runIDKey = artifacts.ResolveArchiveCurrentStateKeys(sessionPrefix)
return sessionPrefix, manifestKey, runIDKey
}
func restoreManifestJSON(t *testing.T, sessionID, campaign string) []byte {
t.Helper()
now := time.Date(2026, 5, 19, 23, 0, 0, 0, time.UTC).Format(time.RFC3339Nano)
payload := map[string]any{
"session_id": sessionID,
"campaign": campaign,
"created_at": now,
"updated_at": now,
"stages": map[string]any{},
}
data, err := json.Marshal(payload)
if err != nil {
t.Fatalf("marshal manifest payload: %v", err)
}
return append(data, '\n')
}
func containsString(values []string, target string) bool {
for _, value := range values {
if value == target {
return true
}
}
return false
}

View File

@@ -0,0 +1,180 @@
package app
import (
"context"
"fmt"
"os"
"path/filepath"
"strings"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/manifest"
)
// RestoreExecutionResult captures concrete file-install results for one restore execution.
type RestoreExecutionResult struct {
DownloadedCount int
}
func executeRestorePlan(
ctx context.Context,
cfg *config.Config,
current *RemoteCurrentState,
plan *RestorePlan,
report *RestoreReport,
store storage.ObjectStore,
) (*RestoreExecutionResult, error) {
if cfg == nil || cfg.Pipeline == nil || cfg.Session == nil {
return nil, fmt.Errorf("resolved config with pipeline/session is required")
}
if current == nil {
return nil, fmt.Errorf("remote current state is required")
}
if plan == nil {
return nil, fmt.Errorf("restore plan is required")
}
if store == nil {
return nil, fmt.Errorf("remote object store is required")
}
sessionRoot := artifacts.SessionWorkDirForCampaign(cfg.Pipeline.Workspace.Root, cfg.Session.Campaign, cfg.Session.SessionID)
manifestActions := make([]RestoreAction, 0, 1)
actions := make([]RestoreAction, 0, len(plan.Actions))
for _, action := range plan.Actions {
if action.Kind != RestoreActionDownload {
continue
}
if action.LocalRelativePath == config.PathManifestFile {
manifestActions = append(manifestActions, action)
continue
}
actions = append(actions, action)
}
if len(manifestActions) > 1 {
return nil, fmt.Errorf("restore plan includes multiple manifest download actions")
}
if len(manifestActions) == 1 {
actions = append(actions, manifestActions[0])
}
result := &RestoreExecutionResult{}
for _, action := range actions {
if err := executeRestoreDownloadAction(ctx, cfg, sessionRoot, current, action, store); err != nil {
if report != nil {
report.markFailed(action, err)
}
return nil, fmt.Errorf("install %q from %q: %w", action.LocalRelativePath, action.RemoteKey, err)
}
if report != nil {
report.markDownloaded(action)
}
result.DownloadedCount++
}
return result, nil
}
func executeRestoreDownloadAction(
ctx context.Context,
cfg *config.Config,
sessionRoot string,
current *RemoteCurrentState,
action RestoreAction,
store storage.ObjectStore,
) error {
safeLocalPath, err := joinWithinSessionRoot(sessionRoot, action.LocalRelativePath)
if err != nil {
return fmt.Errorf("resolve safe local path: %w", err)
}
if strings.TrimSpace(action.LocalPath) != "" && filepath.Clean(action.LocalPath) != safeLocalPath {
return fmt.Errorf("restore plan local path mismatch for %q", action.LocalRelativePath)
}
tmpPath, err := downloadObjectToSiblingTemp(ctx, store, action.RemoteKey, safeLocalPath)
if err != nil {
return fmt.Errorf("download to temp file: %w", err)
}
removeTmp := true
defer func() {
if removeTmp {
_ = os.Remove(tmpPath)
}
}()
if action.LocalRelativePath == config.PathManifestFile {
if err := validateRestoredManifest(ctx, cfg, current, tmpPath); err != nil {
return err
}
}
if err := os.Chmod(tmpPath, 0o644); err != nil {
return fmt.Errorf("set file permissions: %w", err)
}
if err := os.Rename(tmpPath, safeLocalPath); err != nil {
return fmt.Errorf("install file atomically: %w", err)
}
removeTmp = false
return nil
}
func downloadObjectToSiblingTemp(ctx context.Context, store storage.ObjectStore, remoteKey, destPath string) (string, error) {
if strings.TrimSpace(destPath) == "" {
return "", fmt.Errorf("destination path is required")
}
dir := filepath.Dir(destPath)
if err := os.MkdirAll(dir, 0o755); err != nil {
return "", fmt.Errorf("create destination directory: %w", err)
}
base := filepath.Base(destPath)
tmp, err := os.CreateTemp(dir, "."+base+".restore-*.tmp")
if err != nil {
return "", fmt.Errorf("create temp file: %w", err)
}
tmpPath := tmp.Name()
if err := tmp.Close(); err != nil {
_ = os.Remove(tmpPath)
return "", fmt.Errorf("close temp file: %w", err)
}
if err := store.Download(ctx, remoteKey, tmpPath); err != nil {
_ = os.Remove(tmpPath)
return "", err
}
return tmpPath, nil
}
func validateRestoredManifest(ctx context.Context, cfg *config.Config, current *RemoteCurrentState, path string) error {
manifestStore := &manifest.LocalStore{}
m, err := manifestStore.Load(ctx, path)
if err != nil {
return fmt.Errorf("validate manifest decode: %w", err)
}
requestedSession := strings.TrimSpace(cfg.Session.SessionID)
requestedCampaign := strings.TrimSpace(cfg.Session.Campaign)
manifestSession := strings.TrimSpace(m.SessionID)
manifestCampaign := strings.TrimSpace(m.Campaign)
if manifestSession != requestedSession {
return fmt.Errorf("manifest session_id %q does not match requested session_id %q", manifestSession, requestedSession)
}
if manifestCampaign == "" {
return fmt.Errorf("manifest campaign is required")
}
if manifestCampaign != requestedCampaign {
return fmt.Errorf("manifest campaign %q does not match requested campaign %q", manifestCampaign, requestedCampaign)
}
if current != nil {
if expected := strings.TrimSpace(current.SessionID); expected != "" && manifestSession != expected {
return fmt.Errorf("manifest session_id %q does not match discovered session_id %q", manifestSession, expected)
}
if expected := strings.TrimSpace(current.Campaign); expected != "" && manifestCampaign != expected {
return fmt.Errorf("manifest campaign %q does not match discovered campaign %q", manifestCampaign, expected)
}
}
return nil
}

View File

@@ -0,0 +1,364 @@
package app
import (
"bytes"
"context"
"encoding/json"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/manifest"
)
func TestExecuteRestoreNonDryRunRestoresDurableFiles(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
fake := &storage.FakeBackend{}
cfg, sessionPrefix, manifestKey, runIDKey := seedRestoreCommittedState(t, fake, pipelinePath, sessionPath)
seedRestoreObject(fake, sessionPrefix+"transcripts/full.json", []byte(`{"segments":[1,2,3]}`))
seedRestoreObject(fake, sessionPrefix+"artifacts/session_recap.md", []byte("# recap\n"))
seedRestoreObject(fake, sessionPrefix+"audio/alice.flac", []byte("remote-audio"))
seedRestoreObject(fake, runIDKey, []byte("20260519T010203Z-a1b2c3d4\n"))
seedRestoreObject(fake, manifestKey, restoreManifestJSON(t, cfg.Session.SessionID, cfg.Session.Campaign))
restoreWithStoreAndRealPhases(t, fake)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath}, &stdout, &stderr)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, stderr.String())
}
if stderr.Len() != 0 {
t.Fatalf("stderr = %q, want empty", stderr.String())
}
if !strings.Contains(stdout.String(), "Restored session archive for sample-campaign/2026-05-03") {
t.Fatalf("stdout = %q, want completion summary", stdout.String())
}
sessionRoot := artifacts.SessionWorkDirForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
mustReadEquals(t, filepath.Join(sessionRoot, "transcripts", "full.json"), `{"segments":[1,2,3]}`)
mustReadEquals(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "# recap\n")
reportPath := filepath.Join(sessionRoot, "reports", "restore-latest.json")
report := mustReadRestoreReport(t, reportPath)
if report.Status != "succeeded" {
t.Fatalf("report status = %q, want succeeded", report.Status)
}
if report.Execution.Downloaded != 3 {
t.Fatalf("report execution.downloaded = %d, want 3", report.Execution.Downloaded)
}
if len(report.Actions) == 0 {
t.Fatal("report actions is empty")
}
if _, err := os.Stat(filepath.Join(sessionRoot, "audio", "alice.flac")); !os.IsNotExist(err) {
t.Fatalf("audio should not be restored by default; stat err=%v", err)
}
}
func TestExecuteRestoreIncludeAudioRestoresAudio(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
fake := &storage.FakeBackend{}
cfg, sessionPrefix, _, _ := seedRestoreCommittedState(t, fake, pipelinePath, sessionPath)
seedRestoreObject(fake, sessionPrefix+"audio/alice.flac", []byte("remote-audio"))
restoreWithStoreAndRealPhases(t, fake)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath, "--include-audio"}, &stdout, &stderr)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, stderr.String())
}
sessionRoot := artifacts.SessionWorkDirForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
mustReadEquals(t, filepath.Join(sessionRoot, "audio", "alice.flac"), "remote-audio")
report := mustReadRestoreReport(t, filepath.Join(sessionRoot, "reports", "restore-latest.json"))
if !report.IncludeAudio {
t.Fatalf("report include_audio = %v, want true", report.IncludeAudio)
}
}
func TestExecuteRestoreConflictWithoutForceDoesNotOverwrite(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
fake := &storage.FakeBackend{}
cfg, sessionPrefix, _, _ := seedRestoreCommittedState(t, fake, pipelinePath, sessionPath)
seedRestoreObject(fake, sessionPrefix+"transcripts/full.json", []byte("remote-transcript"))
sessionRoot := artifacts.SessionWorkDirForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
mustWriteTestFile(t, filepath.Join(sessionRoot, "transcripts", "full.json"), "local-transcript")
restoreWithStoreAndRealPhases(t, fake)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if !strings.Contains(stderr.String(), "conflicting path") {
t.Fatalf("stderr = %q, want conflict failure", stderr.String())
}
mustReadEquals(t, filepath.Join(sessionRoot, "transcripts", "full.json"), "local-transcript")
report := mustReadRestoreReport(t, filepath.Join(sessionRoot, "reports", "restore-latest.json"))
if report.Status != "failed" {
t.Fatalf("report status = %q, want failed", report.Status)
}
if report.Plan.Conflicts != 1 {
t.Fatalf("report plan.conflicts = %d, want 1", report.Plan.Conflicts)
}
}
func TestExecuteRestoreForceOverwritesDifferingFile(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
fake := &storage.FakeBackend{}
cfg, sessionPrefix, _, _ := seedRestoreCommittedState(t, fake, pipelinePath, sessionPath)
seedRestoreObject(fake, sessionPrefix+"transcripts/full.json", []byte("remote-transcript"))
sessionRoot := artifacts.SessionWorkDirForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
mustWriteTestFile(t, filepath.Join(sessionRoot, "transcripts", "full.json"), "local-transcript")
restoreWithStoreAndRealPhases(t, fake)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath, "--force"}, &stdout, &stderr)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, stderr.String())
}
mustReadEquals(t, filepath.Join(sessionRoot, "transcripts", "full.json"), "remote-transcript")
report := mustReadRestoreReport(t, filepath.Join(sessionRoot, "reports", "restore-latest.json"))
if !report.Force {
t.Fatalf("report force = %v, want true", report.Force)
}
}
func TestExecuteRestoreLockConflictFailsAndWritesNothing(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
fake := &storage.FakeBackend{}
cfg, sessionPrefix, _, _ := seedRestoreCommittedState(t, fake, pipelinePath, sessionPath)
seedRestoreObject(fake, sessionPrefix+"transcripts/full.json", []byte("remote-transcript"))
store := artifacts.NewLocalStore(workspaceRoot)
lock, err := store.AcquireSessionLockFor(cfg.Session.Campaign, cfg.Session.SessionID)
if err != nil {
t.Fatalf("AcquireSessionLockFor() error = %v", err)
}
defer func() { _ = store.ReleaseSessionLock(lock) }()
restoreWithStoreAndRealPhases(t, fake)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if !strings.Contains(stderr.String(), "acquire session lock") {
t.Fatalf("stderr = %q, want lock failure", stderr.String())
}
sessionRoot := artifacts.SessionWorkDirForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
if _, err := os.Stat(filepath.Join(sessionRoot, "transcripts", "full.json")); !os.IsNotExist(err) {
t.Fatalf("transcript should not be restored when lock acquisition fails; stat err=%v", err)
}
}
func TestExecuteRestoreInvalidManifestDoesNotCorruptExistingManifest(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
base := &storage.FakeBackend{}
cfg, sessionPrefix, manifestKey, _ := seedRestoreCommittedState(t, base, pipelinePath, sessionPath)
seedRestoreObject(base, sessionPrefix+"transcripts/full.json", []byte("remote-transcript"))
toggled := &stagedManifestDownloadStore{
delegate: base,
manifestKey: manifestKey,
firstManifest: restoreManifestJSON(t, cfg.Session.SessionID, cfg.Session.Campaign),
secondManifest: []byte("{invalid json"),
manifestReads: 0,
}
sessionRoot := artifacts.SessionWorkDirForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
existing := manifest.New(cfg.Session.SessionID, nowUTC())
existing.Campaign = cfg.Session.Campaign
existingPath := filepath.Join(sessionRoot, "manifest.json")
manifestStore := &manifest.LocalStore{}
if err := manifestStore.Save(context.Background(), existingPath, existing); err != nil {
t.Fatalf("save existing local manifest: %v", err)
}
existingData, err := os.ReadFile(existingPath)
if err != nil {
t.Fatalf("read existing local manifest: %v", err)
}
restoreWithStoreAndRealPhases(t, toggled)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath, "--force"}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if !strings.Contains(stderr.String(), "validate manifest decode") {
t.Fatalf("stderr = %q, want manifest validation failure", stderr.String())
}
mustReadEquals(t, filepath.Join(sessionRoot, "transcripts", "full.json"), "remote-transcript")
report := mustReadRestoreReport(t, filepath.Join(sessionRoot, "reports", "restore-latest.json"))
if report.Status != "failed" {
t.Fatalf("report status = %q, want failed", report.Status)
}
if strings.TrimSpace(report.Error) == "" {
t.Fatal("report error is empty, want failure context")
}
afterData, err := os.ReadFile(existingPath)
if err != nil {
t.Fatalf("read local manifest after failure: %v", err)
}
if string(afterData) != string(existingData) {
t.Fatalf("local manifest changed after failed restore; before=%q after=%q", string(existingData), string(afterData))
}
}
func TestExecuteRestorePlanPathMismatchFails(t *testing.T) {
cfg := restorePlanConfig(t)
current := restorePlanCurrentState(t, cfg)
store := &storage.FakeBackend{}
seedRestoreObject(store, current.SessionPrefix+"transcripts/full.json", []byte("remote-transcript"))
plan := &RestorePlan{Actions: []RestoreAction{{
Kind: RestoreActionDownload,
RemoteKey: current.SessionPrefix + "transcripts/full.json",
LocalRelativePath: "transcripts/full.json",
LocalPath: "/tmp/escape.txt",
}}}
report, err := newRestoreReport(current, plan, RestorePlanOptions{})
if err != nil {
t.Fatalf("newRestoreReport() error = %v", err)
}
_, err = executeRestorePlan(context.Background(), cfg, current, plan, report, store)
if err == nil {
t.Fatal("expected error, got nil")
}
if !strings.Contains(err.Error(), "local path mismatch") {
t.Fatalf("error = %v, want local path mismatch", err)
}
}
func mustReadRestoreReport(t *testing.T, path string) *RestoreReport {
t.Helper()
data, err := os.ReadFile(path)
if err != nil {
t.Fatalf("ReadFile(%q): %v", path, err)
}
var report RestoreReport
if err := json.Unmarshal(data, &report); err != nil {
t.Fatalf("Unmarshal restore report %q: %v", path, err)
}
return &report
}
func restoreWithStoreAndRealPhases(t *testing.T, objectStore storage.ObjectStore) {
t.Helper()
origStoreFn := newObjectStoreFromConfigFn
origDiscoverFn := discoverRemoteCurrentStateFn
origPlanFn := buildRestorePlanFn
origExecuteFn := executeRestorePlanFn
t.Cleanup(func() {
newObjectStoreFromConfigFn = origStoreFn
discoverRemoteCurrentStateFn = origDiscoverFn
buildRestorePlanFn = origPlanFn
executeRestorePlanFn = origExecuteFn
})
newObjectStoreFromConfigFn = func(context.Context, *config.Config) (storage.ObjectStore, error) {
return objectStore, nil
}
discoverRemoteCurrentStateFn = discoverRemoteCurrentState
buildRestorePlanFn = buildRestorePlan
executeRestorePlanFn = executeRestorePlan
}
func seedRestoreCommittedState(t *testing.T, fake *storage.FakeBackend, pipelinePath, sessionPath string) (*config.Config, string, string, string) {
t.Helper()
cfg, err := config.LoadWithSessionOptions(pipelinePath, sessionPath, config.SessionLoadOptions{})
if err != nil {
t.Fatalf("LoadWithSessionOptions() error = %v", err)
}
if err := config.Validate(cfg); err != nil {
t.Fatalf("Validate() error = %v", err)
}
sessionPrefix := artifacts.S3SessionPrefix(cfg.Pipeline.Storage.S3.RootPrefix, cfg.Session.Campaign, cfg.Session.SessionID)
manifestKey, runIDKey := artifacts.ResolveArchiveCurrentStateKeys(sessionPrefix)
seedRestoreObject(fake, runIDKey, []byte("20260519T010203Z-a1b2c3d4\n"))
seedRestoreObject(fake, manifestKey, restoreManifestJSON(t, cfg.Session.SessionID, cfg.Session.Campaign))
return cfg, sessionPrefix, manifestKey, runIDKey
}
func mustReadEquals(t *testing.T, path, want string) {
t.Helper()
data, err := os.ReadFile(path)
if err != nil {
t.Fatalf("ReadFile(%q): %v", path, err)
}
if string(data) != want {
t.Fatalf("file %q = %q, want %q", path, string(data), want)
}
}
type stagedManifestDownloadStore struct {
delegate *storage.FakeBackend
manifestKey string
firstManifest []byte
secondManifest []byte
manifestReads int
}
func (s *stagedManifestDownloadStore) List(ctx context.Context, prefix string) ([]storage.ObjectInfo, error) {
return s.delegate.List(ctx, prefix)
}
func (s *stagedManifestDownloadStore) Download(ctx context.Context, key, localPath string) error {
if strings.TrimSpace(key) == strings.TrimSpace(s.manifestKey) {
s.manifestReads++
payload := s.secondManifest
if s.manifestReads <= 1 {
payload = s.firstManifest
}
if err := os.MkdirAll(filepath.Dir(localPath), 0o755); err != nil {
return fmt.Errorf("download staged manifest: create parent: %w", err)
}
if err := os.WriteFile(localPath, payload, 0o644); err != nil {
return fmt.Errorf("download staged manifest: write local file: %w", err)
}
return nil
}
return s.delegate.Download(ctx, key, localPath)
}
func (s *stagedManifestDownloadStore) Upload(ctx context.Context, localPath, key string, opts storage.UploadOptions) (storage.ObjectInfo, error) {
return s.delegate.Upload(ctx, localPath, key, opts)
}
func (s *stagedManifestDownloadStore) Exists(ctx context.Context, key string) (bool, error) {
return s.delegate.Exists(ctx, key)
}

View File

@@ -0,0 +1,344 @@
package app
import (
"context"
"fmt"
"io"
"os"
"path"
"path/filepath"
"sort"
"strings"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
)
// RestoreActionKind identifies one restore planner action.
type RestoreActionKind string
const (
RestoreActionDownload RestoreActionKind = "download"
RestoreActionSkipSame RestoreActionKind = "skip_same"
RestoreActionConflict RestoreActionKind = "conflict"
)
// RestoreAction is one deterministic planner action.
type RestoreAction struct {
Kind RestoreActionKind
RemoteKey string
LocalRelativePath string
LocalPath string
Size int64
ETag string
ExistsLocal bool
SameLocal bool
Conflict bool
Reason string
}
// RestorePlan is the deterministic output of restore planning.
type RestorePlan struct {
Actions []RestoreAction
DownloadCount int
SkipSameCount int
ConflictCount int
}
// RestorePlanOptions control restore planning scope and classification.
type RestorePlanOptions struct {
IncludeAudio bool
Force bool
DryRun bool
}
func buildRestorePlan(ctx context.Context, cfg *config.Config, current *RemoteCurrentState, store storage.ObjectStore, opts RestorePlanOptions) (*RestorePlan, error) {
if cfg == nil || cfg.Pipeline == nil || cfg.Session == nil {
return nil, fmt.Errorf("resolved config with pipeline/session is required")
}
if current == nil {
return nil, fmt.Errorf("remote current state is required")
}
if store == nil {
return nil, fmt.Errorf("remote object store is required")
}
prefix := normalizeRemoteKey(current.SessionPrefix)
if strings.TrimSpace(prefix) == "" {
return nil, fmt.Errorf("remote session prefix is required")
}
if !strings.HasSuffix(prefix, "/") {
prefix += "/"
}
sessionPaths := artifacts.NewLocalStore(cfg.Pipeline.Workspace.Root).SessionPathsFor(cfg.Session.Campaign, cfg.Session.SessionID)
objects, err := store.List(ctx, prefix)
if err != nil {
return nil, fmt.Errorf("list remote session objects under %q: %w", prefix, err)
}
candidates := make(map[string]storage.ObjectInfo, len(objects)+1)
for _, obj := range objects {
key := normalizeRemoteKey(obj.Key)
if key == "" {
continue
}
obj.Key = key
candidates[key] = obj
}
if strings.TrimSpace(current.CurrentManifestKey) != "" {
key := normalizeRemoteKey(current.CurrentManifestKey)
if _, ok := candidates[key]; !ok {
candidates[key] = storage.ObjectInfo{Key: key}
}
}
actions := make([]RestoreAction, 0, len(candidates))
for key, obj := range candidates {
rel, include, err := restoreLocalRelativePathForKey(prefix, normalizeRemoteKey(current.CurrentManifestKey), key, opts.IncludeAudio)
if err != nil {
return nil, fmt.Errorf("map remote key %q: %w", key, err)
}
if !include {
continue
}
localPath, err := joinWithinSessionRoot(sessionPaths.Root, rel)
if err != nil {
return nil, fmt.Errorf("map remote key %q: %w", key, err)
}
action, err := classifyRestoreAction(ctx, store, obj, rel, localPath, opts.Force)
if err != nil {
return nil, fmt.Errorf("classify remote key %q: %w", key, err)
}
actions = append(actions, action)
}
sort.Slice(actions, func(i, j int) bool {
if actions[i].LocalRelativePath == actions[j].LocalRelativePath {
return actions[i].RemoteKey < actions[j].RemoteKey
}
return actions[i].LocalRelativePath < actions[j].LocalRelativePath
})
plan := &RestorePlan{Actions: actions}
for _, action := range actions {
switch action.Kind {
case RestoreActionDownload:
plan.DownloadCount++
case RestoreActionSkipSame:
plan.SkipSameCount++
case RestoreActionConflict:
plan.ConflictCount++
}
}
_ = opts.DryRun
return plan, nil
}
func normalizeRemoteKey(v string) string {
return strings.Trim(strings.ReplaceAll(strings.TrimSpace(v), "\\", "/"), "/")
}
func restoreLocalRelativePathForKey(sessionPrefix, currentManifestKey, key string, includeAudio bool) (string, bool, error) {
if key == "" {
return "", false, nil
}
if key == currentManifestKey {
return config.PathManifestFile, true, nil
}
if !strings.HasPrefix(key, sessionPrefix) {
return "", false, fmt.Errorf("key is outside resolved session prefix %q", sessionPrefix)
}
rel := strings.TrimPrefix(key, sessionPrefix)
rel = strings.TrimSpace(rel)
if rel == "" {
return "", false, nil
}
cleanRel := path.Clean(rel)
if cleanRel == "." || cleanRel == "" {
return "", false, nil
}
if cleanRel == ".." || strings.HasPrefix(cleanRel, "../") || strings.HasPrefix(cleanRel, "/") {
return "", false, fmt.Errorf("key relative path %q escapes session scope", rel)
}
if cleanRel == config.PathManifestFile {
return config.PathManifestFile, true, nil
}
if strings.HasPrefix(cleanRel, config.S3CurrentSegment+"/") {
return "", false, nil
}
if strings.HasPrefix(cleanRel, config.S3RunsSegment+"/") {
return "", false, nil
}
excludedRoots := []string{
config.PathLogsDirSegment,
config.PathReportsDirSegment,
config.PathConfigDirSegment,
config.PathInputsDirSegment,
}
for _, root := range excludedRoots {
if cleanRel == root || strings.HasPrefix(cleanRel, root+"/") {
return "", false, nil
}
}
if cleanRel == config.PathTranscriptsSegment || strings.HasPrefix(cleanRel, config.PathTranscriptsSegment+"/") {
return cleanRel, true, nil
}
if cleanRel == config.PathArtifactsDirSegment || strings.HasPrefix(cleanRel, config.PathArtifactsDirSegment+"/") {
return cleanRel, true, nil
}
if includeAudio && (cleanRel == config.PathAudioDirSegment || strings.HasPrefix(cleanRel, config.PathAudioDirSegment+"/")) {
return cleanRel, true, nil
}
return "", false, nil
}
func joinWithinSessionRoot(sessionRoot, relative string) (string, error) {
if strings.TrimSpace(sessionRoot) == "" {
return "", fmt.Errorf("session root is required")
}
cleanRel := path.Clean(strings.TrimSpace(relative))
if cleanRel == "." || cleanRel == "" {
return "", fmt.Errorf("relative path is required")
}
if cleanRel == ".." || strings.HasPrefix(cleanRel, "../") || strings.HasPrefix(cleanRel, "/") {
return "", fmt.Errorf("relative path escapes session root")
}
abs := filepath.Clean(filepath.Join(sessionRoot, filepath.FromSlash(cleanRel)))
root := filepath.Clean(sessionRoot)
if abs != root && !strings.HasPrefix(abs, root+string(filepath.Separator)) {
return "", fmt.Errorf("resolved local path escapes session root")
}
return abs, nil
}
func classifyRestoreAction(
ctx context.Context,
store storage.ObjectStore,
object storage.ObjectInfo,
localRelPath string,
localPath string,
force bool,
) (RestoreAction, error) {
action := RestoreAction{
RemoteKey: normalizeRemoteKey(object.Key),
LocalRelativePath: localRelPath,
LocalPath: localPath,
Size: object.Size,
ETag: object.ETag,
}
info, err := os.Stat(localPath)
if err != nil {
if os.IsNotExist(err) {
action.Kind = RestoreActionDownload
action.Reason = "local file missing"
return action, nil
}
return RestoreAction{}, fmt.Errorf("stat local file: %w", err)
}
action.ExistsLocal = true
if info.IsDir() {
action.Kind = RestoreActionConflict
action.Conflict = true
action.Reason = "local path is a directory"
return action, nil
}
if object.Size > 0 && info.Size() != object.Size {
if force {
action.Kind = RestoreActionDownload
action.Reason = "local file differs (size mismatch); overwrite with --force"
return action, nil
}
action.Kind = RestoreActionConflict
action.Conflict = true
action.Reason = "local file differs (size mismatch)"
return action, nil
}
localDigest, err := artifacts.SHA256File(localPath)
if err != nil {
return RestoreAction{}, fmt.Errorf("checksum local file: %w", err)
}
remotePath, err := downloadObjectToTemp(ctx, store, action.RemoteKey, "narratio-restore-plan-remote-*.tmp")
if err != nil {
return RestoreAction{}, fmt.Errorf("download remote object: %w", err)
}
defer func() { _ = os.Remove(remotePath) }()
remoteDigest, err := artifacts.SHA256File(remotePath)
if err != nil {
return RestoreAction{}, fmt.Errorf("checksum remote object: %w", err)
}
if remoteDigest == localDigest {
action.Kind = RestoreActionSkipSame
action.SameLocal = true
action.Reason = "local file matches remote content"
return action, nil
}
if force {
action.Kind = RestoreActionDownload
action.Reason = "local file differs; overwrite with --force"
return action, nil
}
action.Kind = RestoreActionConflict
action.Conflict = true
action.Reason = "local file differs"
return action, nil
}
func writeRestorePlan(out io.Writer, current *RemoteCurrentState, plan *RestorePlan, opts RestorePlanOptions) error {
if out == nil {
return fmt.Errorf("output writer is required")
}
if current == nil {
return fmt.Errorf("remote current state is required")
}
if plan == nil {
return fmt.Errorf("restore plan is required")
}
if _, err := fmt.Fprintf(
out,
"restore plan: session %s/%s run=%s actions=%d download=%d skip_same=%d conflict=%d dry_run=%t force=%t include_audio=%t\n",
current.Campaign,
current.SessionID,
current.RunID,
len(plan.Actions),
plan.DownloadCount,
plan.SkipSameCount,
plan.ConflictCount,
opts.DryRun,
opts.Force,
opts.IncludeAudio,
); err != nil {
return err
}
for _, action := range plan.Actions {
if _, err := fmt.Fprintf(out, "%s %s <- %s", action.Kind, action.LocalRelativePath, action.RemoteKey); err != nil {
return err
}
if strings.TrimSpace(action.Reason) != "" {
if _, err := fmt.Fprintf(out, " (%s)", action.Reason); err != nil {
return err
}
}
if _, err := fmt.Fprintln(out); err != nil {
return err
}
}
return nil
}

View File

@@ -0,0 +1,199 @@
package app
import (
"context"
"path/filepath"
"reflect"
"strings"
"testing"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
)
func TestRestorePlanDefaultScope(t *testing.T) {
cfg := restorePlanConfig(t)
current := restorePlanCurrentState(t, cfg)
store := &storage.FakeBackend{}
seedRestoreObject(store, current.CurrentManifestKey, []byte(`{"session_id":"2026-05-03"}`))
seedRestoreObject(store, current.SessionPrefix+"transcripts/full.json", []byte(`{"segments":[1]}`))
seedRestoreObject(store, current.SessionPrefix+"artifacts/session_recap.md", []byte("# recap\n"))
seedRestoreObject(store, current.SessionPrefix+"audio/alice.flac", []byte("audio"))
seedRestoreObject(store, current.SessionPrefix+"runs/20260519T010203Z-a1b2/manifest.json", []byte("{}"))
seedRestoreObject(store, current.SessionPrefix+"logs/archive.log", []byte("log"))
plan, err := buildRestorePlan(context.Background(), cfg, current, store, RestorePlanOptions{})
if err != nil {
t.Fatalf("buildRestorePlan() error = %v", err)
}
got := actionRelPaths(plan.Actions)
want := []string{"artifacts/session_recap.md", "manifest.json", "transcripts/full.json"}
if !reflect.DeepEqual(got, want) {
t.Fatalf("action local paths = %#v, want %#v", got, want)
}
if plan.DownloadCount != 3 || plan.SkipSameCount != 0 || plan.ConflictCount != 0 {
t.Fatalf("counts = download=%d skip_same=%d conflict=%d, want 3/0/0", plan.DownloadCount, plan.SkipSameCount, plan.ConflictCount)
}
}
func TestRestorePlanIncludeAudio(t *testing.T) {
cfg := restorePlanConfig(t)
current := restorePlanCurrentState(t, cfg)
store := &storage.FakeBackend{}
seedRestoreObject(store, current.CurrentManifestKey, []byte(`{"session_id":"2026-05-03"}`))
seedRestoreObject(store, current.SessionPrefix+"audio/alice.flac", []byte("audio"))
plan, err := buildRestorePlan(context.Background(), cfg, current, store, RestorePlanOptions{IncludeAudio: true})
if err != nil {
t.Fatalf("buildRestorePlan() error = %v", err)
}
got := actionRelPaths(plan.Actions)
want := []string{"audio/alice.flac", "manifest.json"}
if !reflect.DeepEqual(got, want) {
t.Fatalf("action local paths = %#v, want %#v", got, want)
}
}
func TestRestorePlanClassifiesSameAndConflict(t *testing.T) {
cfg := restorePlanConfig(t)
current := restorePlanCurrentState(t, cfg)
store := &storage.FakeBackend{}
seedRestoreObject(store, current.CurrentManifestKey, []byte(`{"session_id":"2026-05-03"}`))
seedRestoreObject(store, current.SessionPrefix+"transcripts/full.json", []byte(`{"segments":[1]}`))
seedRestoreObject(store, current.SessionPrefix+"artifacts/session_recap.md", []byte("remote-content\n"))
sessionRoot := artifacts.SessionWorkDirForCampaign(cfg.Pipeline.Workspace.Root, cfg.Session.Campaign, cfg.Session.SessionID)
mustWriteTestFile(t, filepath.Join(sessionRoot, "transcripts", "full.json"), `{"segments":[1]}`)
mustWriteTestFile(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "different\n")
plan, err := buildRestorePlan(context.Background(), cfg, current, store, RestorePlanOptions{})
if err != nil {
t.Fatalf("buildRestorePlan() error = %v", err)
}
if plan.SkipSameCount != 1 {
t.Fatalf("SkipSameCount = %d, want 1", plan.SkipSameCount)
}
if plan.ConflictCount != 1 {
t.Fatalf("ConflictCount = %d, want 1", plan.ConflictCount)
}
actionByRel := map[string]RestoreAction{}
for _, action := range plan.Actions {
actionByRel[action.LocalRelativePath] = action
}
if actionByRel["transcripts/full.json"].Kind != RestoreActionSkipSame {
t.Fatalf("transcripts/full.json kind = %q, want %q", actionByRel["transcripts/full.json"].Kind, RestoreActionSkipSame)
}
if actionByRel["artifacts/session_recap.md"].Kind != RestoreActionConflict {
t.Fatalf("artifacts/session_recap.md kind = %q, want %q", actionByRel["artifacts/session_recap.md"].Kind, RestoreActionConflict)
}
}
func TestRestorePlanForceTurnsConflictsIntoDownloads(t *testing.T) {
cfg := restorePlanConfig(t)
current := restorePlanCurrentState(t, cfg)
store := &storage.FakeBackend{}
seedRestoreObject(store, current.CurrentManifestKey, []byte(`{"session_id":"2026-05-03"}`))
seedRestoreObject(store, current.SessionPrefix+"artifacts/session_recap.md", []byte("remote-content\n"))
sessionRoot := artifacts.SessionWorkDirForCampaign(cfg.Pipeline.Workspace.Root, cfg.Session.Campaign, cfg.Session.SessionID)
mustWriteTestFile(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "different\n")
plan, err := buildRestorePlan(context.Background(), cfg, current, store, RestorePlanOptions{Force: true})
if err != nil {
t.Fatalf("buildRestorePlan() error = %v", err)
}
actionByRel := map[string]RestoreAction{}
for _, action := range plan.Actions {
actionByRel[action.LocalRelativePath] = action
}
recap := actionByRel["artifacts/session_recap.md"]
if recap.Kind != RestoreActionDownload {
t.Fatalf("artifacts/session_recap.md kind = %q, want %q", recap.Kind, RestoreActionDownload)
}
if plan.ConflictCount != 0 {
t.Fatalf("ConflictCount = %d, want 0", plan.ConflictCount)
}
}
func TestRestorePlanTraversalUnsafeKeyFails(t *testing.T) {
cfg := restorePlanConfig(t)
current := restorePlanCurrentState(t, cfg)
store := &storage.FakeBackend{}
seedRestoreObject(store, current.CurrentManifestKey, []byte(`{"session_id":"2026-05-03"}`))
seedRestoreObject(store, current.SessionPrefix+"artifacts/../../escape.txt", []byte("bad"))
_, err := buildRestorePlan(context.Background(), cfg, current, store, RestorePlanOptions{})
if err == nil {
t.Fatal("expected error, got nil")
}
if !strings.Contains(err.Error(), "escapes session scope") {
t.Fatalf("error = %v, want traversal safety failure", err)
}
}
func seedRestoreObject(store *storage.FakeBackend, key string, data []byte) {
store.SeedObject(storage.FakeObject{Key: key, Data: data})
}
func actionRelPaths(actions []RestoreAction) []string {
out := make([]string, 0, len(actions))
for _, action := range actions {
out = append(out, action.LocalRelativePath)
}
return out
}
func restorePlanConfig(t *testing.T) *config.Config {
t.Helper()
workspaceRoot := t.TempDir()
return &config.Config{
Pipeline: &config.PipelineConfig{
Workspace: config.WorkspaceConfig{Root: workspaceRoot},
},
Session: &config.SessionConfig{
SessionID: "2026-05-03",
Campaign: "sample-campaign",
},
}
}
func restorePlanCurrentState(t *testing.T, cfg *config.Config) *RemoteCurrentState {
t.Helper()
sessionPrefix := artifacts.S3SessionPrefix("dnd", cfg.Session.Campaign, cfg.Session.SessionID)
manifestKey, runIDKey := artifacts.ResolveArchiveCurrentStateKeys(sessionPrefix)
return &RemoteCurrentState{
Bucket: "test-bucket",
SessionPrefix: sessionPrefix,
CurrentManifestKey: manifestKey,
CurrentRunIDKey: runIDKey,
RunID: "20260519T010203Z-a1b2c3d4",
SessionID: cfg.Session.SessionID,
Campaign: cfg.Session.Campaign,
}
}
func TestWriteRestorePlan(t *testing.T) {
current := &RemoteCurrentState{Campaign: "sample-campaign", SessionID: "2026-05-03", RunID: "r-1"}
plan := &RestorePlan{Actions: []RestoreAction{{Kind: RestoreActionDownload, LocalRelativePath: "manifest.json", RemoteKey: "k", Reason: "local file missing"}}, DownloadCount: 1}
var out strings.Builder
if err := writeRestorePlan(&out, current, plan, RestorePlanOptions{DryRun: true}); err != nil {
t.Fatalf("writeRestorePlan() error = %v", err)
}
text := out.String()
if !strings.Contains(text, "restore plan: session sample-campaign/2026-05-03 run=r-1") {
t.Fatalf("output = %q, want plan summary", text)
}
if !strings.Contains(text, "download manifest.json <- k") {
t.Fatalf("output = %q, want action line", text)
}
}

View File

@@ -0,0 +1,243 @@
package app
import (
"encoding/json"
"fmt"
"io"
"path/filepath"
"strings"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
)
// RestoreReport is the durable restore diagnostic model.
type RestoreReport struct {
GeneratedAt string `json:"generated_at"`
SessionID string `json:"session_id"`
Campaign string `json:"campaign"`
RunID string `json:"run_id"`
DryRun bool `json:"dry_run"`
Force bool `json:"force"`
IncludeAudio bool `json:"include_audio"`
Status string `json:"status"`
Error string `json:"error,omitempty"`
Plan RestorePlanSummary `json:"plan"`
Execution RestoreExecutionStats `json:"execution"`
Actions []RestoreReportAction `json:"actions"`
reportPathRel string
}
type RestorePlanSummary struct {
Actions int `json:"actions"`
Download int `json:"download"`
SkipSame int `json:"skip_same"`
Conflicts int `json:"conflicts"`
}
type RestoreExecutionStats struct {
Downloaded int `json:"downloaded"`
Failed int `json:"failed"`
}
type RestoreReportAction struct {
Kind string `json:"kind"`
LocalRelativePath string `json:"local_relative_path"`
RemoteKey string `json:"remote_key"`
Reason string `json:"reason,omitempty"`
Status string `json:"status"`
Error string `json:"error,omitempty"`
}
func newRestoreReport(current *RemoteCurrentState, plan *RestorePlan, opts RestorePlanOptions) (*RestoreReport, error) {
if current == nil {
return nil, fmt.Errorf("remote current state is required")
}
if plan == nil {
return nil, fmt.Errorf("restore plan is required")
}
r := &RestoreReport{
GeneratedAt: nowUTC().Format("2006-01-02T15:04:05.999999999Z07:00"),
SessionID: current.SessionID,
Campaign: current.Campaign,
RunID: current.RunID,
DryRun: opts.DryRun,
Force: opts.Force,
IncludeAudio: opts.IncludeAudio,
Status: "planned",
Plan: RestorePlanSummary{
Actions: len(plan.Actions),
Download: plan.DownloadCount,
SkipSame: plan.SkipSameCount,
Conflicts: plan.ConflictCount,
},
Actions: make([]RestoreReportAction, 0, len(plan.Actions)),
reportPathRel: filepath.ToSlash(filepath.Join(config.PathReportsDirSegment, "restore-latest.json")),
}
for _, action := range plan.Actions {
r.Actions = append(r.Actions, RestoreReportAction{
Kind: string(action.Kind),
LocalRelativePath: action.LocalRelativePath,
RemoteKey: action.RemoteKey,
Reason: action.Reason,
Status: initialRestoreActionStatus(action.Kind),
})
}
return r, nil
}
func initialRestoreActionStatus(kind RestoreActionKind) string {
switch kind {
case RestoreActionDownload:
return "planned_download"
case RestoreActionSkipSame:
return "skipped_same"
case RestoreActionConflict:
return "conflict"
default:
return "planned"
}
}
func (r *RestoreReport) markDownloaded(action RestoreAction) {
if r == nil {
return
}
if idx := r.findAction(action); idx >= 0 {
r.Actions[idx].Status = "downloaded"
r.Actions[idx].Error = ""
}
r.Execution.Downloaded++
}
func (r *RestoreReport) markFailed(action RestoreAction, err error) {
if r == nil {
return
}
if idx := r.findAction(action); idx >= 0 {
r.Actions[idx].Status = "failed"
if err != nil {
r.Actions[idx].Error = err.Error()
}
}
r.Execution.Failed++
}
func (r *RestoreReport) setFailed(err error) {
if r == nil {
return
}
r.Status = "failed"
if err != nil {
r.Error = err.Error()
}
}
func (r *RestoreReport) setSucceeded() {
if r == nil {
return
}
r.Status = "succeeded"
r.Error = ""
}
func (r *RestoreReport) findAction(action RestoreAction) int {
if r == nil {
return -1
}
for i := range r.Actions {
if r.Actions[i].LocalRelativePath == action.LocalRelativePath && r.Actions[i].RemoteKey == action.RemoteKey {
return i
}
}
return -1
}
func writeRestoreDryRunSummary(out io.Writer, report *RestoreReport) error {
if out == nil {
return fmt.Errorf("output writer is required")
}
if report == nil {
return fmt.Errorf("restore report is required")
}
if _, err := fmt.Fprintf(out, "Restore plan for %s/%s\n", report.Campaign, report.SessionID); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Remote run: %s\n", report.RunID); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Would download: %d\n", report.Plan.Download); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Would skip unchanged: %d\n", report.Plan.SkipSame); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Conflicts: %d\n", report.Plan.Conflicts); err != nil {
return err
}
for _, action := range report.Actions {
line := ""
switch action.Status {
case "planned_download":
line = "Would download: " + action.LocalRelativePath
case "skipped_same":
line = "Would skip unchanged: " + action.LocalRelativePath
case "conflict":
line = "Conflict: " + action.LocalRelativePath
default:
line = strings.TrimSpace(action.Kind) + ": " + action.LocalRelativePath
}
if _, err := fmt.Fprintln(out, line); err != nil {
return err
}
}
return nil
}
func writeRestoreSuccessSummary(out io.Writer, report *RestoreReport) error {
if out == nil {
return fmt.Errorf("output writer is required")
}
if report == nil {
return fmt.Errorf("restore report is required")
}
if _, err := fmt.Fprintf(out, "Restored session archive for %s/%s\n", report.Campaign, report.SessionID); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Remote run: %s\n", report.RunID); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Downloaded: %d\n", report.Execution.Downloaded); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Skipped unchanged: %d\n", report.Plan.SkipSame); err != nil {
return err
}
if _, err := fmt.Fprintf(out, "Conflicts: %d\n", report.Plan.Conflicts); err != nil {
return err
}
return nil
}
func persistRestoreReport(store artifacts.Store, cfg *config.Config, report *RestoreReport) (string, error) {
if store == nil {
return "", fmt.Errorf("artifact store is required")
}
if cfg == nil || cfg.Pipeline == nil || cfg.Session == nil {
return "", fmt.Errorf("resolved config with pipeline/session is required")
}
if report == nil {
return "", fmt.Errorf("restore report is required")
}
sessionRoot := artifacts.SessionWorkDirForCampaign(cfg.Pipeline.Workspace.Root, cfg.Session.Campaign, cfg.Session.SessionID)
reportPath := filepath.Join(sessionRoot, filepath.FromSlash(report.reportPathRel))
payload, err := json.MarshalIndent(report, "", " ")
if err != nil {
return "", fmt.Errorf("marshal restore report: %w", err)
}
payload = append(payload, '\n')
if err := store.WriteFileAtomic(reportPath, payload, 0o644); err != nil {
return "", fmt.Errorf("write restore report %q: %w", reportPath, err)
}
return reportPath, nil
}

View File

@@ -0,0 +1,409 @@
package app
import (
"bytes"
"context"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
)
func TestExecuteRestoreHelp(t *testing.T) {
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--help"}, &stdout, &stderr)
if code != 0 {
t.Fatalf("exit code = %d, want 0", code)
}
if stderr.Len() != 0 {
t.Fatalf("stderr = %q, want empty", stderr.String())
}
out := stdout.String()
if !strings.Contains(out, "Usage: narratio restore") {
t.Fatalf("stdout = %q, want restore usage", out)
}
if !strings.Contains(out, "--include-audio") {
t.Fatalf("stdout = %q, want --include-audio flag", out)
}
}
func TestExecuteRestoreRecognizedAndReturnsNYI(t *testing.T) {
origStoreFn := newObjectStoreFromConfigFn
origDiscoverFn := discoverRemoteCurrentStateFn
origPlanFn := buildRestorePlanFn
origExecuteFn := executeRestorePlanFn
t.Cleanup(func() {
newObjectStoreFromConfigFn = origStoreFn
discoverRemoteCurrentStateFn = origDiscoverFn
buildRestorePlanFn = origPlanFn
executeRestorePlanFn = origExecuteFn
})
newObjectStoreFromConfigFn = func(context.Context, *config.Config) (storage.ObjectStore, error) {
return &storage.FakeBackend{}, nil
}
discoverRemoteCurrentStateFn = func(context.Context, *config.Config, storage.ObjectStore) (*RemoteCurrentState, error) {
return &RemoteCurrentState{
SessionID: "2026-05-03",
Campaign: "sample-campaign",
RunID: "20260519T010203Z-a1b2c3d4",
}, nil
}
buildRestorePlanFn = func(context.Context, *config.Config, *RemoteCurrentState, storage.ObjectStore, RestorePlanOptions) (*RestorePlan, error) {
return &RestorePlan{
Actions: []RestoreAction{
{
Kind: RestoreActionDownload,
LocalRelativePath: "manifest.json",
RemoteKey: "dnd/campaigns/sample-campaign/sessions/2026-05-03/current/manifest.json",
Reason: "local file missing",
},
},
DownloadCount: 1,
}, nil
}
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute(
[]string{
"restore",
"--config", pipelinePath,
"--session", sessionPath,
"--session-id", "2026-05-03",
"--dry-run",
"--force",
"--include-audio",
},
&stdout,
&stderr,
)
if code != 0 {
t.Fatalf("exit code = %d, want 0 for --dry-run restore planning; stderr=%q", code, stderr.String())
}
if stderr.Len() != 0 {
t.Fatalf("stderr = %q, want empty", stderr.String())
}
outText := stdout.String()
if !strings.Contains(outText, "Restore plan for sample-campaign/2026-05-03") {
t.Fatalf("stdout = %q, want restore plan summary", outText)
}
if !strings.Contains(outText, "Would download: 1") {
t.Fatalf("stdout = %q, want plan count output", outText)
}
if !strings.Contains(outText, "Would download: manifest.json") {
t.Fatalf("stdout = %q, want action output", outText)
}
manifestPath := artifacts.SessionManifestPathForCampaign(workspaceRoot, "sample-campaign", "2026-05-03")
if _, err := os.Stat(manifestPath); !os.IsNotExist(err) {
t.Fatalf("manifest should not be created during phase-4 restore planning; stat err=%v", err)
}
reportPath := filepath.Join(workspaceRoot, "work", "sample-campaign", "2026-05-03", "reports", "restore-latest.json")
if _, err := os.Stat(reportPath); !os.IsNotExist(err) {
t.Fatalf("restore report should not be written during dry-run; stat err=%v", err)
}
}
func TestExecuteRestoreRejectsUnexpectedPositionalArguments(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath, "extra"}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if !strings.Contains(stderr.String(), "restore: unexpected positional arguments") {
t.Fatalf("stderr = %q, want positional-args failure", stderr.String())
}
}
func TestExecuteRestoreFailsWhenStorageBackendNotConfigured(t *testing.T) {
origStoreFn := newObjectStoreFromConfigFn
origDiscoverFn := discoverRemoteCurrentStateFn
t.Cleanup(func() {
newObjectStoreFromConfigFn = origStoreFn
discoverRemoteCurrentStateFn = origDiscoverFn
})
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeRestoreConfigWithoutStorage(t, workspaceRoot)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if !strings.Contains(stderr.String(), "no remote object store backend is configured") {
t.Fatalf("stderr = %q, want storage backend preflight failure", stderr.String())
}
}
func TestExecuteRestoreDiscoveryErrorSurfaced(t *testing.T) {
origStoreFn := newObjectStoreFromConfigFn
origDiscoverFn := discoverRemoteCurrentStateFn
t.Cleanup(func() {
newObjectStoreFromConfigFn = origStoreFn
discoverRemoteCurrentStateFn = origDiscoverFn
})
newObjectStoreFromConfigFn = func(context.Context, *config.Config) (storage.ObjectStore, error) {
return &storage.FakeBackend{}, nil
}
discoverRemoteCurrentStateFn = func(context.Context, *config.Config, storage.ObjectStore) (*RemoteCurrentState, error) {
return nil, fmt.Errorf("remote current run pointer missing: %q", "dnd/campaigns/sample-campaign/sessions/2026-05-03/current/run_id.txt")
}
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if !strings.Contains(stderr.String(), "remote current run pointer missing") {
t.Fatalf("stderr = %q, want discovery error context", stderr.String())
}
}
func TestExecuteRestoreLoadsSecretsBeforeObjectStoreInit(t *testing.T) {
origStoreFn := newObjectStoreFromConfigFn
origDiscoverFn := discoverRemoteCurrentStateFn
origPlanFn := buildRestorePlanFn
origExecuteFn := executeRestorePlanFn
t.Cleanup(func() {
newObjectStoreFromConfigFn = origStoreFn
discoverRemoteCurrentStateFn = origDiscoverFn
buildRestorePlanFn = origPlanFn
executeRestorePlanFn = origExecuteFn
})
const accessKeyEnv = "OBJECT_STORAGE_KEY_ID"
const secretKeyEnv = "OBJECT_STORAGE_KEY"
restoreEnv := func(name string) {
value, exists := os.LookupEnv(name)
_ = os.Unsetenv(name)
t.Cleanup(func() {
if exists {
_ = os.Setenv(name, value)
return
}
_ = os.Unsetenv(name)
})
}
restoreEnv(accessKeyEnv)
restoreEnv(secretKeyEnv)
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
secretsDir := filepath.Join(t.TempDir(), "secrets")
mustWriteTestFile(t, filepath.Join(secretsDir, accessKeyEnv), "test-access-key-id\n")
mustWriteTestFile(t, filepath.Join(secretsDir, secretKeyEnv), "test-secret-key\n")
f, err := os.OpenFile(pipelinePath, os.O_APPEND|os.O_WRONLY, 0)
if err != nil {
t.Fatalf("open pipeline config for append: %v", err)
}
defer f.Close()
if _, err := f.WriteString("\nsecrets:\n env_dir: " + secretsDir + "\n"); err != nil {
t.Fatalf("append secrets config: %v", err)
}
storeInitCalled := false
newObjectStoreFromConfigFn = func(context.Context, *config.Config) (storage.ObjectStore, error) {
storeInitCalled = true
gotID, okID := os.LookupEnv(accessKeyEnv)
if !okID || gotID != "test-access-key-id" {
return nil, fmt.Errorf("missing or unexpected %s: %q (set=%t)", accessKeyEnv, gotID, okID)
}
gotSecret, okSecret := os.LookupEnv(secretKeyEnv)
if !okSecret || gotSecret != "test-secret-key" {
return nil, fmt.Errorf("missing or unexpected %s: %q (set=%t)", secretKeyEnv, gotSecret, okSecret)
}
return &storage.FakeBackend{}, nil
}
discoverRemoteCurrentStateFn = func(context.Context, *config.Config, storage.ObjectStore) (*RemoteCurrentState, error) {
return &RemoteCurrentState{
SessionID: "2026-05-03",
Campaign: "sample-campaign",
RunID: "20260519T010203Z-a1b2c3d4",
}, nil
}
buildRestorePlanFn = func(context.Context, *config.Config, *RemoteCurrentState, storage.ObjectStore, RestorePlanOptions) (*RestorePlan, error) {
return &RestorePlan{
Actions: []RestoreAction{
{
Kind: RestoreActionDownload,
LocalRelativePath: "manifest.json",
RemoteKey: "dnd/campaigns/sample-campaign/sessions/2026-05-03/current/manifest.json",
Reason: "local file missing",
},
},
DownloadCount: 1,
}, nil
}
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute(
[]string{
"restore",
"--config", pipelinePath,
"--session", sessionPath,
"--session-id", "2026-05-03",
"--dry-run",
},
&stdout,
&stderr,
)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, stderr.String())
}
if !storeInitCalled {
t.Fatal("expected object store initialization to be called")
}
}
func TestExecuteRestoreNonDryRunConflictFailsBeforeNYI(t *testing.T) {
origStoreFn := newObjectStoreFromConfigFn
origDiscoverFn := discoverRemoteCurrentStateFn
origPlanFn := buildRestorePlanFn
origExecuteFn := executeRestorePlanFn
t.Cleanup(func() {
newObjectStoreFromConfigFn = origStoreFn
discoverRemoteCurrentStateFn = origDiscoverFn
buildRestorePlanFn = origPlanFn
executeRestorePlanFn = origExecuteFn
})
newObjectStoreFromConfigFn = func(context.Context, *config.Config) (storage.ObjectStore, error) {
return &storage.FakeBackend{}, nil
}
discoverRemoteCurrentStateFn = func(context.Context, *config.Config, storage.ObjectStore) (*RemoteCurrentState, error) {
return &RemoteCurrentState{
SessionID: "2026-05-03",
Campaign: "sample-campaign",
RunID: "20260519T010203Z-a1b2c3d4",
}, nil
}
buildRestorePlanFn = func(context.Context, *config.Config, *RemoteCurrentState, storage.ObjectStore, RestorePlanOptions) (*RestorePlan, error) {
return &RestorePlan{
Actions: []RestoreAction{
{Kind: RestoreActionConflict, LocalRelativePath: "transcripts/full.json", RemoteKey: "k", Reason: "local file differs"},
},
ConflictCount: 1,
}, nil
}
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath}, &stdout, &stderr)
if code == 0 {
t.Fatal("exit code = 0, want non-zero")
}
if stdout.Len() != 0 {
t.Fatalf("stdout = %q, want empty on conflict failure", stdout.String())
}
if !strings.Contains(stderr.String(), "restore conflict: 1 conflicting path(s); rerun with --force to overwrite") {
t.Fatalf("stderr = %q, want conflict failure", stderr.String())
}
if strings.Contains(stderr.String(), "phase 4: restore execution") {
t.Fatalf("stderr = %q, should fail before phase-4 NYI boundary", stderr.String())
}
}
func TestExecuteRestoreNonDryRunForceExecutesPlan(t *testing.T) {
origStoreFn := newObjectStoreFromConfigFn
origDiscoverFn := discoverRemoteCurrentStateFn
origPlanFn := buildRestorePlanFn
origExecuteFn := executeRestorePlanFn
t.Cleanup(func() {
newObjectStoreFromConfigFn = origStoreFn
discoverRemoteCurrentStateFn = origDiscoverFn
buildRestorePlanFn = origPlanFn
executeRestorePlanFn = origExecuteFn
})
newObjectStoreFromConfigFn = func(context.Context, *config.Config) (storage.ObjectStore, error) {
return &storage.FakeBackend{}, nil
}
discoverRemoteCurrentStateFn = func(context.Context, *config.Config, storage.ObjectStore) (*RemoteCurrentState, error) {
return &RemoteCurrentState{
SessionID: "2026-05-03",
Campaign: "sample-campaign",
RunID: "20260519T010203Z-a1b2c3d4",
}, nil
}
buildRestorePlanFn = func(context.Context, *config.Config, *RemoteCurrentState, storage.ObjectStore, RestorePlanOptions) (*RestorePlan, error) {
return &RestorePlan{
Actions: []RestoreAction{
{Kind: RestoreActionDownload, LocalRelativePath: "transcripts/full.json", RemoteKey: "k", Reason: "local file differs; overwrite with --force"},
},
DownloadCount: 1,
}, nil
}
executeRestorePlanFn = func(context.Context, *config.Config, *RemoteCurrentState, *RestorePlan, *RestoreReport, storage.ObjectStore) (*RestoreExecutionResult, error) {
return &RestoreExecutionResult{DownloadedCount: 1}, nil
}
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFiles(t, workspaceRoot)
var stdout bytes.Buffer
var stderr bytes.Buffer
code := Execute([]string{"restore", "--config", pipelinePath, "--session", sessionPath, "--force"}, &stdout, &stderr)
if code != 0 {
t.Fatalf("exit code = %d, want 0; stderr=%q", code, stderr.String())
}
if !strings.Contains(stdout.String(), "Restored session archive for sample-campaign/2026-05-03") {
t.Fatalf("stdout = %q, want completion summary", stdout.String())
}
if stderr.Len() != 0 {
t.Fatalf("stderr = %q, want empty", stderr.String())
}
}
func writeRestoreConfigWithoutStorage(t *testing.T, workspaceRoot string) (string, string) {
t.Helper()
dir := t.TempDir()
pipelinePath := filepath.Join(dir, "pipeline.yml")
sessionPath := filepath.Join(dir, "session.yml")
pipelineYAML := `workspace:
root: ` + workspaceRoot + `
whisperx:
transcribe_url: https://example.com/transcribe
`
sessionYAML := `session_id: 2026-05-03
campaign: sample-campaign
inputs:
audio_dir: ./audio
speakers_file: ./speakers.yml
autocorrect_file: ./autocorrect.yml
glossary_file: ./glossary.yml
`
if err := os.WriteFile(pipelinePath, []byte(pipelineYAML), 0o644); err != nil {
t.Fatalf("write pipeline config: %v", err)
}
if err := os.WriteFile(sessionPath, []byte(sessionYAML), 0o644); err != nil {
t.Fatalf("write session config: %v", err)
}
mustWriteTestFile(t, filepath.Join(dir, "speakers.yml"), "alice: alice.flac\n")
mustWriteTestFile(t, filepath.Join(dir, "autocorrect.yml"), "[]\n")
mustWriteTestFile(t, filepath.Join(dir, "glossary.yml"), "[]\n")
mustWriteTestFile(t, filepath.Join(dir, "audio", "alice.flac"), "audio-bytes")
return pipelinePath, sessionPath
}

View File

@@ -0,0 +1,178 @@
package app
import (
"bytes"
"context"
"os"
"path/filepath"
"strings"
"testing"
"time"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/scriptorium"
"gitea.maximumdirect.net/eric/narratio/internal/adapters/storage"
"gitea.maximumdirect.net/eric/narratio/internal/artifacts"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/manifest"
"gitea.maximumdirect.net/eric/narratio/internal/stage"
)
func TestRestoreThenRunStageForceAnalyzeUsesRestoredDurableState(t *testing.T) {
workspaceRoot := t.TempDir()
pipelinePath, sessionPath := writeValidConfigFilesWithScriptoriumArtifacts(t, workspaceRoot)
fake := &storage.FakeBackend{}
cfg, sessionPrefix, manifestKey, runIDKey := seedRestoreCommittedState(t, fake, pipelinePath, sessionPath)
seedRestoreObject(fake, runIDKey, []byte("20260519T010203Z-a1b2c3d4\n"))
seedRestoreObject(fake, manifestKey, restoreWorkflowManifestJSON(t, cfg.Session.SessionID, cfg.Session.Campaign))
seedRestoreObject(fake, sessionPrefix+"transcripts/full.json", []byte(`{"segments":[1,2,3]}`+"\n"))
seedRestoreObject(fake, sessionPrefix+"artifacts/session_recap.md", []byte("# restored recap\n"))
restoreWithStoreAndRealPhases(t, fake)
origExecuteStagesFn := executeStagesFn
t.Cleanup(func() {
executeStagesFn = origExecuteStagesFn
})
executeStagesFn = func(ctx context.Context, cfg *config.Config, stages []stage.Stage, opts RunOptions) (*RunSummary, error) {
if opts.Env == nil {
opts.Env = &Env{}
}
opts.Env.Scriptorium = &scriptorium.NoopRunner{}
return executeStages(ctx, cfg, stages, opts)
}
var stdout bytes.Buffer
var stderr bytes.Buffer
restoreCode := Execute(
[]string{
"restore",
"--config", pipelinePath,
"--session", sessionPath,
"--session-id", cfg.Session.SessionID,
},
&stdout,
&stderr,
)
if restoreCode != 0 {
t.Fatalf("restore exit code = %d, want 0; stderr=%q", restoreCode, stderr.String())
}
if stderr.Len() != 0 {
t.Fatalf("restore stderr = %q, want empty", stderr.String())
}
sessionRoot := artifacts.SessionWorkDirForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
mustReadEquals(t, filepath.Join(sessionRoot, "transcripts", "full.json"), `{"segments":[1,2,3]}`+"\n")
mustReadEquals(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "# restored recap\n")
manifestStore := &manifest.LocalStore{}
sessionManifestPath := artifacts.SessionManifestPathForCampaign(workspaceRoot, cfg.Session.Campaign, cfg.Session.SessionID)
beforeAnalyze, err := manifestStore.Load(context.Background(), sessionManifestPath)
if err != nil {
t.Fatalf("load restored session manifest: %v", err)
}
upstreamCompletedAt := map[string]time.Time{}
for _, stageName := range []string{"prepare", "transcribe", "merge", "polish", "normalize", "trim"} {
rec := beforeAnalyze.Stages[stageName]
if rec == nil || rec.Status != manifest.StatusSucceeded || rec.CompletedAt == nil {
t.Fatalf("restored manifest stage %q = %#v, want succeeded with completion timestamp", stageName, rec)
}
upstreamCompletedAt[stageName] = *rec.CompletedAt
}
stdout.Reset()
stderr.Reset()
runStageCode := Execute(
[]string{
"run-stage",
"--config", pipelinePath,
"--session", sessionPath,
"--session-id", cfg.Session.SessionID,
"--force",
"--artifacts", "player_handout",
"analyze",
},
&stdout,
&stderr,
)
if runStageCode != 0 {
t.Fatalf("run-stage exit code = %d, want 0; stderr=%q", runStageCode, stderr.String())
}
if stderr.Len() != 0 {
t.Fatalf("run-stage stderr = %q, want empty", stderr.String())
}
if !strings.Contains(stdout.String(), "stage=analyze executed=1 skipped=0 force=true") {
t.Fatalf("run-stage stdout = %q, want analyze execution summary", stdout.String())
}
playerHandoutPath := filepath.Join(sessionRoot, "artifacts", "player_handout.md")
if _, err := os.Stat(playerHandoutPath); err != nil {
t.Fatalf("restored analyze output %q missing: %v", playerHandoutPath, err)
}
afterAnalyze, err := manifestStore.Load(context.Background(), sessionManifestPath)
if err != nil {
t.Fatalf("load session manifest after run-stage analyze: %v", err)
}
for _, stageName := range []string{"prepare", "transcribe", "merge", "polish", "normalize", "trim"} {
rec := afterAnalyze.Stages[stageName]
if rec == nil || rec.Status != manifest.StatusSucceeded || rec.CompletedAt == nil {
t.Fatalf("post-analyze manifest stage %q = %#v, want succeeded with completion timestamp", stageName, rec)
}
if !rec.CompletedAt.Equal(upstreamCompletedAt[stageName]) {
t.Fatalf(
"stage %q completion changed: before=%s after=%s",
stageName,
upstreamCompletedAt[stageName].Format(time.RFC3339Nano),
rec.CompletedAt.Format(time.RFC3339Nano),
)
}
}
analyzeRec := afterAnalyze.Stages["analyze"]
if analyzeRec == nil || analyzeRec.Status != manifest.StatusSucceeded {
t.Fatalf("post-analyze stage record = %#v, want succeeded", analyzeRec)
}
runManifestPaths, err := filepath.Glob(filepath.Join(sessionRoot, "runs", "*", "manifest.json"))
if err != nil {
t.Fatalf("glob run manifests: %v", err)
}
if len(runManifestPaths) != 1 {
t.Fatalf("run manifest count = %d, want 1; paths=%v", len(runManifestPaths), runManifestPaths)
}
runManifest, err := manifestStore.LoadRun(context.Background(), runManifestPaths[0])
if err != nil {
t.Fatalf("load run manifest %q: %v", runManifestPaths[0], err)
}
if len(runManifest.RequestedStages) != 1 || runManifest.RequestedStages[0] != "analyze" {
t.Fatalf("run manifest requested_stages = %#v, want [analyze]", runManifest.RequestedStages)
}
if runManifest.Stages["analyze"] == nil || runManifest.Stages["analyze"].Status != manifest.StatusSucceeded {
t.Fatalf("run manifest analyze stage = %#v, want succeeded", runManifest.Stages["analyze"])
}
if runManifest.Stages["prepare"] != nil {
t.Fatalf("run manifest should not include upstream prepare stage, got %#v", runManifest.Stages["prepare"])
}
}
func restoreWorkflowManifestJSON(t *testing.T, sessionID, campaign string) []byte {
t.Helper()
store := &manifest.LocalStore{}
now := time.Date(2026, 5, 19, 23, 0, 0, 0, time.UTC)
m := manifest.New(sessionID, now)
m.Campaign = campaign
m.RunID = "20260519T010203Z-a1b2c3d4"
stages := []string{"prepare", "transcribe", "merge", "polish", "normalize", "trim"}
for i, stageName := range stages {
m.MarkStageSucceeded(stageName, now.Add(time.Duration(i+1)*time.Minute), nil)
}
path := filepath.Join(t.TempDir(), "manifest.json")
if err := store.Save(context.Background(), path, m); err != nil {
t.Fatalf("save workflow manifest fixture: %v", err)
}
data, err := os.ReadFile(path)
if err != nil {
t.Fatalf("read workflow manifest fixture: %v", err)
}
return data
}

View File

@@ -76,7 +76,7 @@ func Resume(ctx context.Context, args []string, out io.Writer) error {
} }
} }
summary, err := executeStages(ctx, cfg, selected, RunOptions{ summary, err := executeStagesFn(ctx, cfg, selected, RunOptions{
Force: force, Force: force,
SelectedArtifacts: normalizedArtifacts, SelectedArtifacts: normalizedArtifacts,
}) })

View File

@@ -58,7 +58,7 @@ func Run(ctx context.Context, args []string, out io.Writer) error {
} }
stages := BuildFullPlan() stages := BuildFullPlan()
summary, err := executeStages(ctx, cfg, stages, RunOptions{ summary, err := executeStagesFn(ctx, cfg, stages, RunOptions{
Force: force, Force: force,
SelectedArtifacts: normalizedArtifacts, SelectedArtifacts: normalizedArtifacts,
}) })

View File

@@ -66,7 +66,7 @@ func RunStage(ctx context.Context, args []string, out io.Writer) error {
return fmt.Errorf("run-stage: %w", err) return fmt.Errorf("run-stage: %w", err)
} }
summary, err := executeStages(ctx, cfg, stages, RunOptions{ summary, err := executeStagesFn(ctx, cfg, stages, RunOptions{
Force: force, Force: force,
SelectedArtifacts: normalizedArtifacts, SelectedArtifacts: normalizedArtifacts,
}) })

View File

@@ -36,6 +36,8 @@ type RunSummary struct {
Skipped []string Skipped []string
} }
var executeStagesFn = executeStages
func executeStages(ctx context.Context, cfg *config.Config, stages []stage.Stage, opts RunOptions) (*RunSummary, error) { func executeStages(ctx context.Context, cfg *config.Config, stages []stage.Stage, opts RunOptions) (*RunSummary, error) {
env := opts.Env env := opts.Env
if env == nil { if env == nil {

View File

@@ -209,7 +209,14 @@ func TestExecuteStagesArchiveFailsWhenRequiredRecapPromotionMissingForSelectedAr
Enabled: boolPtr(true), Enabled: boolPtr(true),
UploadRun: boolPtr(true), UploadRun: boolPtr(true),
PromoteArtifacts: []config.ArchivePromotionRule{ PromoteArtifacts: []config.ArchivePromotionRule{
{From: "artifacts/session_recap.md", To: "artifacts/session_recap.md", Required: boolPtr(true)}, {Source: "narratio.artifact.session_recap", Dest: "artifacts/session_recap.md", Required: boolPtr(true)},
},
}
cfg.Pipeline.Scriptorium = &config.ScriptoriumConfig{
Artifacts: map[string]config.ScriptoriumArtifactConfig{
"session_recap": {
OutputPath: "artifacts/session_recap.md",
},
}, },
} }
@@ -247,8 +254,8 @@ func TestExecuteStagesArchiveFailsWhenRequiredRecapPromotionMissingForSelectedAr
if err == nil { if err == nil {
t.Fatal("expected archive promotion failure, got nil") t.Fatal("expected archive promotion failure, got nil")
} }
if !strings.Contains(err.Error(), "required promotion source missing") { if !strings.Contains(err.Error(), "required promotion source unavailable") {
t.Fatalf("error = %q, want required promotion source missing", err.Error()) t.Fatalf("error = %q, want required promotion source unavailable", err.Error())
} }
} }

View File

@@ -0,0 +1,77 @@
package artifacts
import (
"fmt"
"strings"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/manifest"
)
// ResolveArchiveBucket resolves archive bucket identity with manifest-first precedence.
func ResolveArchiveBucket(cfg *config.Config, m *manifest.Manifest) string {
if m != nil && strings.TrimSpace(m.S3Bucket) != "" {
return strings.TrimSpace(m.S3Bucket)
}
if cfg == nil || cfg.Pipeline == nil || cfg.Pipeline.Storage.S3 == nil {
return ""
}
return strings.TrimSpace(cfg.Pipeline.Storage.S3.Bucket)
}
// ResolveArchiveSessionPrefix resolves archive session prefix with manifest-first precedence.
func ResolveArchiveSessionPrefix(cfg *config.Config, m *manifest.Manifest) (string, error) {
if m != nil && strings.TrimSpace(m.S3SessionPrefix) != "" {
return strings.TrimSpace(m.S3SessionPrefix), nil
}
if cfg == nil || cfg.Session == nil || cfg.Pipeline == nil {
return "", fmt.Errorf("resolved config is required")
}
sessionID := strings.TrimSpace(cfg.Session.SessionID)
if sessionID == "" && m != nil {
sessionID = strings.TrimSpace(m.SessionID)
}
campaign := strings.TrimSpace(cfg.Session.Campaign)
if campaign == "" && m != nil {
campaign = strings.TrimSpace(m.Campaign)
}
if cfg.Pipeline.Storage.S3 == nil {
return "", fmt.Errorf("pipeline.storage.s3 configuration is required")
}
sessionPrefix := S3SessionPrefix(cfg.Pipeline.Storage.S3.RootPrefix, campaign, sessionID)
if strings.TrimSpace(sessionPrefix) == "" {
return "", fmt.Errorf("session prefix is required")
}
return sessionPrefix, nil
}
// ResolveArchiveRunPrefix resolves archive run prefix with manifest-first precedence.
func ResolveArchiveRunPrefix(cfg *config.Config, m *manifest.Manifest) (string, error) {
if m != nil {
runPrefix := strings.TrimSpace(m.S3RunPrefix)
if runPrefix != "" {
return runPrefix, nil
}
}
sessionPrefix, err := ResolveArchiveSessionPrefix(cfg, m)
if err != nil {
return "", err
}
runID := ""
if m != nil {
runID = strings.TrimSpace(m.RunID)
}
if runID == "" {
return "", fmt.Errorf("run id is required")
}
return S3RunPrefix(sessionPrefix, runID), nil
}
// ResolveArchiveCurrentStateKeys returns current pointer keys for a session prefix.
func ResolveArchiveCurrentStateKeys(sessionPrefix string) (manifestKey, runIDKey string) {
return S3CurrentManifestKey(sessionPrefix), S3CurrentRunPointerKey(sessionPrefix)
}

View File

@@ -0,0 +1,130 @@
package artifacts
import (
"strings"
"testing"
"gitea.maximumdirect.net/eric/narratio/internal/config"
"gitea.maximumdirect.net/eric/narratio/internal/manifest"
)
func TestResolveArchiveBucketPrefersManifestThenConfig(t *testing.T) {
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Storage: config.StorageConfig{
S3: &config.StorageS3Config{Bucket: "cfg-bucket"},
},
},
}
if got := ResolveArchiveBucket(cfg, &manifest.Manifest{S3Bucket: "manifest-bucket"}); got != "manifest-bucket" {
t.Fatalf("bucket = %q, want manifest-bucket", got)
}
if got := ResolveArchiveBucket(cfg, &manifest.Manifest{}); got != "cfg-bucket" {
t.Fatalf("bucket = %q, want cfg-bucket", got)
}
}
func TestResolveArchiveSessionPrefixPrefersManifestThenConfig(t *testing.T) {
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Storage: config.StorageConfig{
S3: &config.StorageS3Config{RootPrefix: "dnd"},
},
},
Session: &config.SessionConfig{
SessionID: "2026-04-19",
Campaign: "forsaken",
},
}
m := &manifest.Manifest{S3SessionPrefix: "manifest/session/prefix/"}
got, err := ResolveArchiveSessionPrefix(cfg, m)
if err != nil {
t.Fatalf("ResolveArchiveSessionPrefix() error = %v", err)
}
if got != "manifest/session/prefix/" {
t.Fatalf("session prefix = %q, want manifest/session/prefix/", got)
}
got, err = ResolveArchiveSessionPrefix(cfg, &manifest.Manifest{})
if err != nil {
t.Fatalf("ResolveArchiveSessionPrefix() error = %v", err)
}
want := "dnd/campaigns/forsaken/sessions/2026-04-19/"
if got != want {
t.Fatalf("session prefix = %q, want %q", got, want)
}
}
func TestResolveArchiveRunPrefixPrefersManifestThenDerived(t *testing.T) {
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Storage: config.StorageConfig{
S3: &config.StorageS3Config{RootPrefix: "dnd"},
},
},
Session: &config.SessionConfig{
SessionID: "2026-04-19",
Campaign: "forsaken",
},
}
m := &manifest.Manifest{
RunID: "20260516T010203Z-1a2b3c4d",
S3RunPrefix: "manifest/run/prefix/",
}
got, err := ResolveArchiveRunPrefix(cfg, m)
if err != nil {
t.Fatalf("ResolveArchiveRunPrefix() error = %v", err)
}
if got != "manifest/run/prefix/" {
t.Fatalf("run prefix = %q, want manifest/run/prefix/", got)
}
m = &manifest.Manifest{
RunID: "20260516T010203Z-1a2b3c4d",
}
got, err = ResolveArchiveRunPrefix(cfg, m)
if err != nil {
t.Fatalf("ResolveArchiveRunPrefix() error = %v", err)
}
want := "dnd/campaigns/forsaken/sessions/2026-04-19/runs/20260516T010203Z-1a2b3c4d/"
if got != want {
t.Fatalf("run prefix = %q, want %q", got, want)
}
}
func TestResolveArchiveIdentityErrorsAreDeterministic(t *testing.T) {
cfgNoS3 := &config.Config{
Pipeline: &config.PipelineConfig{},
Session: &config.SessionConfig{SessionID: "2026-04-19", Campaign: "forsaken"},
}
_, err := ResolveArchiveSessionPrefix(cfgNoS3, &manifest.Manifest{})
if err == nil || !strings.Contains(err.Error(), "pipeline.storage.s3 configuration is required") {
t.Fatalf("error = %v, want missing storage.s3", err)
}
cfg := &config.Config{
Pipeline: &config.PipelineConfig{
Storage: config.StorageConfig{
S3: &config.StorageS3Config{RootPrefix: "dnd"},
},
},
Session: &config.SessionConfig{SessionID: "2026-04-19", Campaign: "forsaken"},
}
_, err = ResolveArchiveRunPrefix(cfg, &manifest.Manifest{})
if err == nil || !strings.Contains(err.Error(), "run id is required") {
t.Fatalf("error = %v, want missing run id", err)
}
}
func TestResolveArchiveCurrentStateKeys(t *testing.T) {
manifestKey, runIDKey := ResolveArchiveCurrentStateKeys("dnd/campaigns/forsaken/sessions/2026-04-19/")
if manifestKey != "dnd/campaigns/forsaken/sessions/2026-04-19/current/manifest.json" {
t.Fatalf("manifest key = %q", manifestKey)
}
if runIDKey != "dnd/campaigns/forsaken/sessions/2026-04-19/current/run_id.txt" {
t.Fatalf("run id key = %q", runIDKey)
}
}

View File

@@ -79,8 +79,8 @@ type ArchiveConfig struct {
// ArchivePromotionRule configures one artifact promotion mapping. // ArchivePromotionRule configures one artifact promotion mapping.
type ArchivePromotionRule struct { type ArchivePromotionRule struct {
From string `yaml:"from"` Source string `yaml:"source"`
To string `yaml:"to"` Dest string `yaml:"dest"`
Required *bool `yaml:"required"` Required *bool `yaml:"required"`
} }

View File

@@ -74,8 +74,7 @@ const (
// DefaultArchivePromoteArtifacts defines the default archive promotion rules. // DefaultArchivePromoteArtifacts defines the default archive promotion rules.
// Callers should copy this slice before mutating. // Callers should copy this slice before mutating.
var DefaultArchivePromoteArtifacts = []ArchivePromotionRule{ var DefaultArchivePromoteArtifacts = []ArchivePromotionRule{
{From: PathTranscriptTrimmed, To: PathTranscriptTrimmed}, {Source: "narratio.transcript.trimmed", Dest: PathTranscriptTrimmed},
{From: "artifacts/session_recap.md", To: "artifacts/session_recap.md"},
} }
// DefaultPipelineConfigSearchPaths defines the default search order for // DefaultPipelineConfigSearchPaths defines the default search order for

View File

@@ -136,40 +136,86 @@ func TestSpoolAndArchiveDefaults(t *testing.T) {
if cfg.Pipeline.Archive.UploadRun == nil || !*cfg.Pipeline.Archive.UploadRun { if cfg.Pipeline.Archive.UploadRun == nil || !*cfg.Pipeline.Archive.UploadRun {
t.Fatalf("archive.upload_run = %#v, want true", cfg.Pipeline.Archive.UploadRun) t.Fatalf("archive.upload_run = %#v, want true", cfg.Pipeline.Archive.UploadRun)
} }
if len(cfg.Pipeline.Archive.PromoteArtifacts) != 2 { if len(cfg.Pipeline.Archive.PromoteArtifacts) != 1 {
t.Fatalf("archive.promote_artifacts len = %d, want 2 defaults", len(cfg.Pipeline.Archive.PromoteArtifacts)) t.Fatalf("archive.promote_artifacts len = %d, want 1 default", len(cfg.Pipeline.Archive.PromoteArtifacts))
} }
for i, item := range cfg.Pipeline.Archive.PromoteArtifacts { item := cfg.Pipeline.Archive.PromoteArtifacts[0]
if item.Required == nil || !*item.Required { if item.Required == nil || !*item.Required {
t.Fatalf("archive.promote_artifacts[%d].required = %#v, want true", i, item.Required) t.Fatalf("archive.promote_artifacts[0].required = %#v, want true", item.Required)
} }
if item.Source != "narratio.transcript.trimmed" {
t.Fatalf("archive.promote_artifacts[0].source = %q, want narratio.transcript.trimmed", item.Source)
}
if item.Dest != "transcripts/trimmed.json" {
t.Fatalf("archive.promote_artifacts[0].dest = %q, want transcripts/trimmed.json", item.Dest)
} }
} }
func TestArchivePromotionPathValidation(t *testing.T) { func TestArchivePromotionValidation(t *testing.T) {
tests := []struct { tests := []struct {
name string name string
ruleYML string ruleYML string
wantErr string wantErr string
}{ }{
{ {
name: "absolute from path rejected", name: "absolute dest path rejected",
ruleYML: `archive: ruleYML: `archive:
promote_artifacts: promote_artifacts:
- from: "/transcripts/trimmed.json" - source: "narratio.transcript.trimmed"
to: "transcripts/trimmed.json" dest: "/transcripts/trimmed.json"
`, `,
wantErr: "must be a relative path", wantErr: "must be a relative path",
}, },
{ {
name: "traversal to path rejected", name: "traversal dest path rejected",
ruleYML: `archive: ruleYML: `archive:
promote_artifacts: promote_artifacts:
- from: "transcripts/trimmed.json" - source: "narratio.transcript.trimmed"
to: "../trimmed.json" dest: "../trimmed.json"
`, `,
wantErr: "must not contain path traversal", wantErr: "must not contain path traversal",
}, },
{
name: "invalid source rejected",
ruleYML: `archive:
promote_artifacts:
- source: "narratio.unknown"
dest: "transcripts/trimmed.json"
`,
wantErr: "source \"narratio.unknown\" is unsupported",
},
{
name: "duplicate destination rejected",
ruleYML: `archive:
promote_artifacts:
- source: "narratio.transcript.trimmed"
dest: "artifacts/shared.md"
- source: "narratio.transcript.full"
dest: "artifacts/shared.md"
`,
wantErr: "duplicates another archive promotion destination",
},
{
name: "configured source requires configured artifact key",
ruleYML: `archive:
promote_artifacts:
- source: "narratio.artifact.session_recap"
dest: "artifacts/session_recap.md"
`,
wantErr: "configured artifact \"session_recap\" is not defined in pipeline.scriptorium.artifacts",
},
{
name: "configured source without output path fails when dest omitted",
ruleYML: `scriptorium:
artifacts:
session_recap:
enabled: false
archive:
promote_artifacts:
- source: "narratio.artifact.session_recap"
`,
wantErr: "destination cannot be derived",
},
} }
for _, tt := range tests { for _, tt := range tests {
@@ -189,6 +235,72 @@ func TestArchivePromotionPathValidation(t *testing.T) {
} }
} }
func TestArchivePromotionDerivesDestinationWhenOmitted(t *testing.T) {
tests := []struct {
name string
pipelineYML string
wantDest string
}{
{
name: "built in source derives canonical destination",
pipelineYML: testPipelineBaseYAML + `
archive:
promote_artifacts:
- source: narratio.transcript.full
`,
wantDest: "transcripts/normalized.json",
},
{
name: "configured source derives configured output path",
pipelineYML: testPipelineBaseYAML + `
scriptorium:
artifacts:
session_recap:
enabled: true
prompt_id: dnd.session_recap
output_path: artifacts/session_recap.md
archive:
promote_artifacts:
- source: narratio.artifact.session_recap
`,
wantDest: "artifacts/session_recap.md",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
pipelinePath, sessionPath := writeConfigFiles(t, tt.pipelineYML, testSessionBaseYAML)
cfg, err := Load(pipelinePath, sessionPath)
if err != nil {
t.Fatalf("Load() error = %v", err)
}
if err := Validate(cfg); err != nil {
t.Fatalf("Validate() error = %v", err)
}
if len(cfg.Pipeline.Archive.PromoteArtifacts) != 1 {
t.Fatalf("archive.promote_artifacts len = %d, want 1", len(cfg.Pipeline.Archive.PromoteArtifacts))
}
if cfg.Pipeline.Archive.PromoteArtifacts[0].Dest != tt.wantDest {
t.Fatalf("archive.promote_artifacts[0].dest = %q, want %q", cfg.Pipeline.Archive.PromoteArtifacts[0].Dest, tt.wantDest)
}
})
}
}
func TestArchiveLegacyFromToFailsStrictDecode(t *testing.T) {
pipelineYAML := testPipelineBaseYAML + `
archive:
promote_artifacts:
- from: transcripts/trimmed.json
to: transcripts/trimmed.json
`
pipelinePath, sessionPath := writeConfigFiles(t, pipelineYAML, testSessionBaseYAML)
_, err := Load(pipelinePath, sessionPath)
if err == nil || !strings.Contains(err.Error(), "strict decode failed") {
t.Fatalf("Load() error = %v, want strict decode failed", err)
}
}
func TestSessionAudioS3Validation(t *testing.T) { func TestSessionAudioS3Validation(t *testing.T) {
tests := []struct { tests := []struct {
name string name string

View File

@@ -47,7 +47,7 @@ func validatePipeline(cfg *PipelineConfig) error {
if err := validateSpool(cfg.Spool); err != nil { if err := validateSpool(cfg.Spool); err != nil {
return err return err
} }
if err := validateArchive(cfg.Archive); err != nil { if err := validateArchive(cfg.Archive, cfg.Scriptorium); err != nil {
return err return err
} }
if err := validateWhisperX(cfg.WhisperX); err != nil { if err := validateWhisperX(cfg.WhisperX); err != nil {
@@ -104,28 +104,91 @@ func validateSpool(cfg SpoolConfig) error {
return nil return nil
} }
func validateArchive(cfg *ArchiveConfig) error { func validateArchive(cfg *ArchiveConfig, scriptorium *ScriptoriumConfig) error {
if cfg == nil { if cfg == nil {
return nil return nil
} }
seenDest := map[string]struct{}{}
for i, item := range cfg.PromoteArtifacts { for i, item := range cfg.PromoteArtifacts {
prefix := fmt.Sprintf("pipeline.archive.promote_artifacts[%d]", i) prefix := fmt.Sprintf("pipeline.archive.promote_artifacts[%d]", i)
if strings.TrimSpace(item.From) == "" { source := strings.TrimSpace(item.Source)
return fmt.Errorf("%s.from is required", prefix) if source == "" {
return fmt.Errorf("%s.source is required", prefix)
} }
if strings.TrimSpace(item.To) == "" { if _, err := archiveSourceKnown(source, scriptorium); err != nil {
return fmt.Errorf("%s.to is required", prefix) return fmt.Errorf("%s.source %q is unsupported: %w", prefix, item.Source, err)
} }
if err := validateRelativeSafePath(prefix+".from", item.From); err != nil { dest := strings.TrimSpace(item.Dest)
if dest == "" {
derivedDest, err := deriveArchivePromotionDest(source, scriptorium)
if err != nil {
return fmt.Errorf("%s.dest is required when destination cannot be derived from %q: %w", prefix, source, err)
}
dest = derivedDest
cfg.PromoteArtifacts[i].Dest = derivedDest
}
if err := validateRelativeSafePath(prefix+".dest", dest); err != nil {
return err return err
} }
if err := validateRelativeSafePath(prefix+".to", item.To); err != nil { normalizedDest := filepath.ToSlash(filepath.Clean(dest))
return err if _, ok := seenDest[normalizedDest]; ok {
return fmt.Errorf("%s.dest %q duplicates another archive promotion destination", prefix, dest)
} }
seenDest[normalizedDest] = struct{}{}
} }
return nil return nil
} }
func archiveSourceKnown(source string, scriptorium *ScriptoriumConfig) (string, error) {
trimmed := strings.TrimSpace(source)
switch trimmed {
case "narratio.transcript.merged",
"narratio.transcript.polished",
"narratio.transcript.full",
"narratio.transcript.trimmed",
"narratio.bounds.session":
return "", nil
}
matches := narratioArtifactSourceRE.FindStringSubmatch(trimmed)
if len(matches) != 2 {
return "", fmt.Errorf("must be a built-in source id or narratio.artifact.<name>")
}
artifactKey := matches[1]
if scriptorium == nil || len(scriptorium.Artifacts) == 0 {
return "", fmt.Errorf("configured artifact %q is not defined in pipeline.scriptorium.artifacts", artifactKey)
}
if _, ok := scriptorium.Artifacts[artifactKey]; !ok {
return "", fmt.Errorf("configured artifact %q is not defined in pipeline.scriptorium.artifacts", artifactKey)
}
return artifactKey, nil
}
func deriveArchivePromotionDest(source string, scriptorium *ScriptoriumConfig) (string, error) {
trimmed := strings.TrimSpace(source)
switch trimmed {
case "narratio.transcript.merged":
return PathTranscriptMerged, nil
case "narratio.transcript.polished":
return PathTranscriptProcessed, nil
case "narratio.transcript.full":
return PathTranscriptNormalized, nil
case "narratio.transcript.trimmed":
return PathTranscriptTrimmed, nil
case "narratio.bounds.session":
return filepath.ToSlash(filepath.Join(PathArtifactsDirSegment, "session_bounds.json")), nil
}
artifactKey, err := archiveSourceKnown(trimmed, scriptorium)
if err != nil {
return "", err
}
artifactCfg := scriptorium.Artifacts[artifactKey]
outputPath := strings.TrimSpace(artifactCfg.OutputPath)
if outputPath == "" {
return "", fmt.Errorf("pipeline.scriptorium.artifacts.%s.output_path is empty", artifactKey)
}
return outputPath, nil
}
func validateSecrets(cfg *SecretsConfig) error { func validateSecrets(cfg *SecretsConfig) error {
if cfg == nil { if cfg == nil {
return nil return nil

View File

@@ -3,6 +3,7 @@ package stage
import ( import (
"context" "context"
"encoding/json" "encoding/json"
"errors"
"fmt" "fmt"
"io/fs" "io/fs"
"os" "os"
@@ -90,15 +91,15 @@ func (archiveStage) Run(ctx context.Context, env *Env, m *manifest.Manifest) (*S
return nil, fmt.Errorf("archive: run root %q is not a directory", runRoot) return nil, fmt.Errorf("archive: run root %q is not a directory", runRoot)
} }
runPrefix, err := archiveRunPrefix(env, m) runPrefix, err := artifacts.ResolveArchiveRunPrefix(env.Config, m)
if err != nil { if err != nil {
return nil, fmt.Errorf("archive: resolve s3 run prefix: %w", err) return nil, fmt.Errorf("archive: resolve s3 run prefix: %w", err)
} }
sessionPrefix, err := archiveSessionPrefix(env, m) sessionPrefix, err := artifacts.ResolveArchiveSessionPrefix(env.Config, m)
if err != nil { if err != nil {
return nil, fmt.Errorf("archive: resolve s3 session prefix: %w", err) return nil, fmt.Errorf("archive: resolve s3 session prefix: %w", err)
} }
bucket := archiveBucket(env, m) bucket := artifacts.ResolveArchiveBucket(env.Config, m)
if bucket == "" { if bucket == "" {
return nil, fmt.Errorf("archive: resolve s3 bucket: bucket is required") return nil, fmt.Errorf("archive: resolve s3 bucket: bucket is required")
} }
@@ -116,11 +117,12 @@ func (archiveStage) Run(ctx context.Context, env *Env, m *manifest.Manifest) (*S
if err != nil { if err != nil {
return nil, fmt.Errorf("archive: collect run files: %w", err) return nil, fmt.Errorf("archive: collect run files: %w", err)
} }
sessionRoot, err := resolveArchiveSessionRoot(env, m) sessionPaths := archiveSessionPaths(env, m)
runtimeCatalog, err := buildArchiveRuntimeArtifactCatalog(sessionPaths, env.Config.Pipeline.Scriptorium)
if err != nil { if err != nil {
return nil, fmt.Errorf("archive: resolve session root for promotions: %w", err) return nil, fmt.Errorf("archive: build runtime artifact catalog: %w", err)
} }
promotions, err := resolveArchivePromotions(sessionRoot, env.Config.Pipeline.Archive.PromoteArtifacts) promotions, skippedOptional, err := resolveArchivePromotions(sessionPaths, m, runtimeCatalog, env.Config.Pipeline.Archive.PromoteArtifacts)
if err != nil { if err != nil {
return nil, fmt.Errorf("archive: resolve promotion rules: %w", err) return nil, fmt.Errorf("archive: resolve promotion rules: %w", err)
} }
@@ -134,23 +136,15 @@ func (archiveStage) Run(ctx context.Context, env *Env, m *manifest.Manifest) (*S
} }
promotedUploaded := make([]string, 0, len(promotions)) promotedUploaded := make([]string, 0, len(promotions))
skippedOptional := make([]string, 0)
for _, promotion := range promotions { for _, promotion := range promotions {
if !promotion.Exists { key := artifacts.S3PromotedArtifactKey(sessionPrefix, promotion.Dest)
if promotion.Required {
return nil, fmt.Errorf("archive: required promotion source missing: %q", promotion.From)
}
skippedOptional = append(skippedOptional, promotion.To)
continue
}
key := artifacts.S3PromotedArtifactKey(sessionPrefix, promotion.To)
if _, err := env.ObjectStore.Upload(ctx, promotion.LocalPath, key, storage.UploadOptions{}); err != nil { if _, err := env.ObjectStore.Upload(ctx, promotion.LocalPath, key, storage.UploadOptions{}); err != nil {
return nil, fmt.Errorf("archive: upload promoted output %q to %q: %w", promotion.From, key, err) return nil, fmt.Errorf("archive: upload promoted output source %q to %q: %w", promotion.Source, key, err)
} }
promotedUploaded = append(promotedUploaded, promotion.To) promotedUploaded = append(promotedUploaded, promotion.Dest)
} }
currentManifestKey := artifacts.S3CurrentManifestKey(sessionPrefix) currentManifestKey, currentRunPointerKey := artifacts.ResolveArchiveCurrentStateKeys(sessionPrefix)
manifestTempPath, err := writeCurrentManifestSnapshot(m, archiveMetadataPreview( manifestTempPath, err := writeCurrentManifestSnapshot(m, archiveMetadataPreview(
bucket, bucket,
runPrefix, runPrefix,
@@ -171,7 +165,6 @@ func (archiveStage) Run(ctx context.Context, env *Env, m *manifest.Manifest) (*S
return nil, fmt.Errorf("archive: upload current manifest to %q: %w", currentManifestKey, err) return nil, fmt.Errorf("archive: upload current manifest to %q: %w", currentManifestKey, err)
} }
currentRunPointerKey := artifacts.S3CurrentRunPointerKey(sessionPrefix)
runIDTempPath, err := writeCurrentRunIDPointer(runID) runIDTempPath, err := writeCurrentRunIDPointer(runID)
if err != nil { if err != nil {
return nil, fmt.Errorf("archive: build current run id pointer: %w", err) return nil, fmt.Errorf("archive: build current run id pointer: %w", err)
@@ -204,11 +197,11 @@ func (archiveStage) Run(ctx context.Context, env *Env, m *manifest.Manifest) (*S
} }
type archivePromotion struct { type archivePromotion struct {
From string Source string
To string Dest string
Required bool Required bool
LocalPath string LocalPath string
Exists bool Provenance string
} }
func archiveDisabled(env *Env) bool { func archiveDisabled(env *Env) bool {
@@ -292,105 +285,151 @@ func resolveArchiveSessionRoot(env *Env, m *manifest.Manifest) (string, error) {
return filepath.Clean(artifacts.SessionWorkDirForCampaign(env.Config.Pipeline.Workspace.Root, campaign, sessionID)), nil return filepath.Clean(artifacts.SessionWorkDirForCampaign(env.Config.Pipeline.Workspace.Root, campaign, sessionID)), nil
} }
func archiveRunPrefix(env *Env, m *manifest.Manifest) (string, error) { func archiveSessionPaths(env *Env, m *manifest.Manifest) artifacts.SessionPaths {
runPrefix := strings.TrimSpace(m.S3RunPrefix)
if runPrefix != "" {
return runPrefix, nil
}
sessionPrefix, err := archiveSessionPrefix(env, m)
if err != nil {
return "", err
}
runID := strings.TrimSpace(m.RunID)
if runID == "" {
return "", fmt.Errorf("run id is required")
}
return artifacts.S3RunPrefix(sessionPrefix, runID), nil
}
func archiveSessionPrefix(env *Env, m *manifest.Manifest) (string, error) {
if m != nil && strings.TrimSpace(m.S3SessionPrefix) != "" {
return strings.TrimSpace(m.S3SessionPrefix), nil
}
sessionID := strings.TrimSpace(env.Config.Session.SessionID) sessionID := strings.TrimSpace(env.Config.Session.SessionID)
if sessionID == "" { if sessionID == "" && m != nil {
sessionID = strings.TrimSpace(m.SessionID) sessionID = strings.TrimSpace(m.SessionID)
} }
campaign := strings.TrimSpace(env.Config.Session.Campaign) campaign := strings.TrimSpace(env.Config.Session.Campaign)
if campaign == "" { if campaign == "" && m != nil {
campaign = strings.TrimSpace(m.Campaign) campaign = strings.TrimSpace(m.Campaign)
} }
if env.Config.Pipeline.Storage.S3 == nil { store := artifacts.NewLocalStore(env.Config.Pipeline.Workspace.Root)
return "", fmt.Errorf("pipeline.storage.s3 configuration is required") return store.SessionPathsFor(campaign, sessionID)
}
sessionPrefix := artifacts.S3SessionPrefix(env.Config.Pipeline.Storage.S3.RootPrefix, campaign, sessionID)
if strings.TrimSpace(sessionPrefix) == "" {
return "", fmt.Errorf("session prefix is required")
}
return sessionPrefix, nil
} }
func archiveBucket(env *Env, m *manifest.Manifest) string { func resolveArchivePromotions(
if m != nil && strings.TrimSpace(m.S3Bucket) != "" { paths artifacts.SessionPaths,
return strings.TrimSpace(m.S3Bucket) m *manifest.Manifest,
} catalog *artifacts.ArtifactCatalog,
if env == nil || env.Config == nil || env.Config.Pipeline == nil || env.Config.Pipeline.Storage.S3 == nil { rules []config.ArchivePromotionRule,
return "" ) ([]archivePromotion, []string, error) {
}
return strings.TrimSpace(env.Config.Pipeline.Storage.S3.Bucket)
}
func resolveArchivePromotions(sessionRoot string, rules []config.ArchivePromotionRule) ([]archivePromotion, error) {
sessionRoot = filepath.Clean(strings.TrimSpace(sessionRoot))
if sessionRoot == "" {
return nil, fmt.Errorf("session root is required")
}
out := make([]archivePromotion, 0, len(rules)) out := make([]archivePromotion, 0, len(rules))
skippedOptional := make([]string, 0)
for _, rule := range rules { for _, rule := range rules {
from := strings.TrimSpace(rule.From) source := strings.TrimSpace(rule.Source)
to := strings.TrimSpace(rule.To)
required := rule.Required == nil || *rule.Required required := rule.Required == nil || *rule.Required
dest, err := resolveArchivePromotionDest(rule, catalog)
resolvedPath, err := resolveWorkDirRelativePath(sessionRoot, from)
if err != nil { if err != nil {
return nil, fmt.Errorf("promotion from %q: %w", from, err) return nil, nil, fmt.Errorf("source %q: %w", source, err)
} }
info, err := os.Stat(resolvedPath) resolved, err := artifacts.ResolveSessionArtifactWithCatalog(paths, m, source, catalog)
exists := err == nil && !info.IsDir() if err != nil {
if err != nil && !os.IsNotExist(err) { if errors.Is(err, artifacts.ErrSessionArtifactNotFound) && !required {
return nil, fmt.Errorf("promotion source %q: %w", from, err) skippedOptional = append(skippedOptional, dest)
continue
}
if errors.Is(err, artifacts.ErrSessionArtifactNotFound) {
return nil, nil, fmt.Errorf("required promotion source unavailable: %q", source)
}
return nil, nil, fmt.Errorf("resolve source %q: %w", source, err)
} }
localPath := resolvedPath
out = append(out, archivePromotion{ out = append(out, archivePromotion{
From: from, Source: source,
To: to, Dest: dest,
Required: required, Required: required,
LocalPath: localPath, LocalPath: resolved.Path,
Exists: exists, Provenance: resolved.Provenance,
}) })
} }
return out, nil return out, skippedOptional, nil
} }
func resolveWorkDirRelativePath(workDir, rel string) (string, error) { func resolveArchivePromotionDest(rule config.ArchivePromotionRule, catalog *artifacts.ArtifactCatalog) (string, error) {
rel = filepath.Clean(filepath.FromSlash(strings.TrimSpace(rel))) dest := strings.TrimSpace(rule.Dest)
if rel == "." || rel == "" { if dest == "" {
entry, ok := catalog.Lookup(strings.TrimSpace(rule.Source))
if !ok {
return "", fmt.Errorf("destination omitted and source is unknown")
}
dest = strings.TrimSpace(entry.CanonicalRelPath)
if dest == "" {
return "", fmt.Errorf("destination omitted and no canonical destination is available")
}
}
return normalizeArchiveRelativePath(dest)
}
func normalizeArchiveRelativePath(rel string) (string, error) {
trimmed := strings.TrimSpace(rel)
if trimmed == "" {
return "", fmt.Errorf("relative path is required") return "", fmt.Errorf("relative path is required")
} }
full := filepath.Join(workDir, rel) cleaned := filepath.ToSlash(filepath.Clean(filepath.FromSlash(trimmed)))
cleanedWork := filepath.Clean(workDir) if cleaned == "." || cleaned == "" {
cleanedFull := filepath.Clean(full) return "", fmt.Errorf("relative path is required")
relative, err := filepath.Rel(cleanedWork, cleanedFull) }
if filepath.IsAbs(trimmed) || strings.HasPrefix(cleaned, "/") || cleaned == ".." || strings.HasPrefix(cleaned, "../") {
return "", fmt.Errorf("path must be a clean relative path")
}
return cleaned, nil
}
func buildArchiveRuntimeArtifactCatalog(
paths artifacts.SessionPaths,
scriptoriumCfg *config.ScriptoriumConfig,
) (*artifacts.ArtifactCatalog, error) {
catalog := artifacts.NewArtifactCatalog()
if err := catalog.RegisterBuiltIns(); err != nil {
return nil, err
}
if scriptoriumCfg == nil {
return catalog, nil
}
configured := map[string]artifacts.ConfiguredArtifactDefinition{}
for key, artifactCfg := range scriptoriumCfg.Artifacts {
configured[key] = artifacts.ConfiguredArtifactDefinition{
Enabled: artifactCfg.Enabled,
OutputPath: artifactCfg.OutputPath,
}
}
if err := catalog.RegisterConfiguredArtifacts(configured, nil); err != nil {
return nil, err
}
for _, entry := range catalog.ListConfigured() {
if strings.TrimSpace(entry.CanonicalRelPath) == "" {
continue
}
localPath, err := resolveConfiguredArtifactLocalPath(paths, entry.CanonicalRelPath)
if err != nil { if err != nil {
return "", fmt.Errorf("compute relative path: %w", err) continue
} }
if relative == ".." || strings.HasPrefix(relative, ".."+string(filepath.Separator)) { info, statErr := os.Stat(localPath)
return "", fmt.Errorf("path escapes workdir") if statErr != nil {
if os.IsNotExist(statErr) {
continue
} }
return cleanedFull, nil return nil, fmt.Errorf("stat configured artifact %q: %w", entry.SourceID, statErr)
}
if info.IsDir() {
continue
}
if err := catalog.MarkAvailableFromDisk(entry.SourceID, localPath); err != nil {
return nil, err
}
}
return catalog, nil
}
func resolveConfiguredArtifactLocalPath(paths artifacts.SessionPaths, configured string) (string, error) {
outputPath := strings.TrimSpace(configured)
if outputPath == "" {
return "", fmt.Errorf("configured artifact output path is required")
}
if filepath.IsAbs(outputPath) {
return filepath.Clean(outputPath), nil
}
rel := filepath.Clean(outputPath)
if rel == "." || rel == "" {
return "", fmt.Errorf("relative output path is required")
}
if rel == ".." || strings.HasPrefix(rel, ".."+string(filepath.Separator)) {
return "", fmt.Errorf("relative output path escapes session root: %q", configured)
}
return filepath.Join(paths.Root, rel), nil
} }
func collectArchiveRunFiles(runRoot, manifestPath string) ([]archiveUploadFile, error) { func collectArchiveRunFiles(runRoot, manifestPath string) ([]archiveUploadFile, error) {

View File

@@ -135,8 +135,8 @@ func TestArchiveUploadsRunRecordPromotionsAndCurrentPointer(t *testing.T) {
func TestArchiveUsesCustomPromotionRules(t *testing.T) { func TestArchiveUsesCustomPromotionRules(t *testing.T) {
env, m, _ := archiveFixture(t) env, m, _ := archiveFixture(t)
env.Config.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{ env.Config.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{
{From: "transcripts/trimmed.json", To: "published/trimmed.json", Required: boolPtr(true)}, {Source: "narratio.transcript.trimmed", Dest: "published/trimmed.json", Required: boolPtr(true)},
{From: "artifacts/session_recap.md", To: "published/recap.md", Required: boolPtr(true)}, {Source: "narratio.artifact.session_recap", Dest: "published/recap.md", Required: boolPtr(true)},
} }
_, err := archiveStage{}.Run(context.Background(), env, m) _, err := archiveStage{}.Run(context.Background(), env, m)
@@ -156,8 +156,8 @@ func TestArchiveUsesCustomPromotionRules(t *testing.T) {
func TestArchiveSkipsOptionalMissingPromotion(t *testing.T) { func TestArchiveSkipsOptionalMissingPromotion(t *testing.T) {
env, m, _ := archiveFixture(t) env, m, _ := archiveFixture(t)
env.Config.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{ env.Config.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{
{From: "transcripts/trimmed.json", To: "transcripts/trimmed.json", Required: boolPtr(true)}, {Source: "narratio.transcript.trimmed", Dest: "transcripts/trimmed.json", Required: boolPtr(true)},
{From: "artifacts/optional.md", To: "artifacts/optional.md", Required: boolPtr(false)}, {Source: "narratio.transcript.merged", Dest: "transcripts/merged.json", Required: boolPtr(false)},
} }
result, err := archiveStage{}.Run(context.Background(), env, m) result, err := archiveStage{}.Run(context.Background(), env, m)
@@ -165,7 +165,7 @@ func TestArchiveSkipsOptionalMissingPromotion(t *testing.T) {
t.Fatalf("Run() error = %v", err) t.Fatalf("Run() error = %v", err)
} }
got, _ := result.Metadata["skipped_optional_promotions"].([]string) got, _ := result.Metadata["skipped_optional_promotions"].([]string)
want := []string{"artifacts/optional.md"} want := []string{"transcripts/merged.json"}
if !reflect.DeepEqual(got, want) { if !reflect.DeepEqual(got, want) {
t.Fatalf("skipped_optional_promotions = %#v, want %#v", got, want) t.Fatalf("skipped_optional_promotions = %#v, want %#v", got, want)
} }
@@ -174,12 +174,12 @@ func TestArchiveSkipsOptionalMissingPromotion(t *testing.T) {
func TestArchiveFailsWhenRequiredPromotionMissing(t *testing.T) { func TestArchiveFailsWhenRequiredPromotionMissing(t *testing.T) {
env, m, _ := archiveFixture(t) env, m, _ := archiveFixture(t)
env.Config.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{ env.Config.Pipeline.Archive.PromoteArtifacts = []config.ArchivePromotionRule{
{From: "artifacts/missing.md", To: "artifacts/missing.md", Required: boolPtr(true)}, {Source: "narratio.transcript.merged", Dest: "transcripts/merged.json", Required: boolPtr(true)},
} }
_, err := archiveStage{}.Run(context.Background(), env, m) _, err := archiveStage{}.Run(context.Background(), env, m)
if err == nil || !strings.Contains(err.Error(), "required promotion source missing") { if err == nil || !strings.Contains(err.Error(), "required promotion source unavailable") {
t.Fatalf("Run() error = %v, want required promotion missing failure", err) t.Fatalf("Run() error = %v, want required promotion source unavailable failure", err)
} }
} }
@@ -259,13 +259,13 @@ func archiveFixture(t *testing.T) (*Env, *manifest.Manifest, string) {
sessionRoot := artifacts.SessionWorkDirForCampaign(root, campaign, sessionID) sessionRoot := artifacts.SessionWorkDirForCampaign(root, campaign, sessionID)
runRoot := artifacts.SessionRunRootForCampaign(root, campaign, sessionID, runID) runRoot := artifacts.SessionRunRootForCampaign(root, campaign, sessionID, runID)
writeStageTestFile(t, filepath.Join(sessionRoot, "transcripts", "trimmed.json"), "{}\n") writeStageTestFile(t, filepath.Join(sessionRoot, "transcripts", "trimmed.json"), "{\"segments\":[]}\n")
writeStageTestFile(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "# recap\n") writeStageTestFile(t, filepath.Join(sessionRoot, "artifacts", "session_recap.md"), "# recap\n")
writeStageTestFile(t, filepath.Join(runRoot, "prepare", "inputs", "session.yml"), "session_id: 2026-04-19\n") writeStageTestFile(t, filepath.Join(runRoot, "prepare", "inputs", "session.yml"), "session_id: 2026-04-19\n")
writeStageTestFile(t, filepath.Join(runRoot, "prepare", "outputs", "audio", "speaker.flac"), "flac\n") writeStageTestFile(t, filepath.Join(runRoot, "prepare", "outputs", "audio", "speaker.flac"), "flac\n")
writeStageTestFile(t, filepath.Join(runRoot, "transcribe", "outputs", "transcripts", "raw", "speaker.json"), "{}\n") writeStageTestFile(t, filepath.Join(runRoot, "transcribe", "outputs", "transcripts", "raw", "speaker.json"), "{}\n")
writeStageTestFile(t, filepath.Join(runRoot, "trim", "outputs", "transcripts", "trimmed.json"), "{}\n") writeStageTestFile(t, filepath.Join(runRoot, "trim", "outputs", "transcripts", "trimmed.json"), "{\"segments\":[]}\n")
writeStageTestFile(t, filepath.Join(runRoot, "analyze", "outputs", "artifacts", "session_recap.md"), "# recap\n") writeStageTestFile(t, filepath.Join(runRoot, "analyze", "outputs", "artifacts", "session_recap.md"), "# recap\n")
writeStageTestFile(t, filepath.Join(runRoot, "polish", "reports", "audita.report.json"), "{}\n") writeStageTestFile(t, filepath.Join(runRoot, "polish", "reports", "audita.report.json"), "{}\n")
writeStageTestFile(t, filepath.Join(runRoot, "merge", "config", "seriatim.generated.yml"), "key: value\n") writeStageTestFile(t, filepath.Join(runRoot, "merge", "config", "seriatim.generated.yml"), "key: value\n")
@@ -298,8 +298,17 @@ func archiveFixture(t *testing.T) (*Env, *manifest.Manifest, string) {
Enabled: boolPtr(true), Enabled: boolPtr(true),
UploadRun: boolPtr(true), UploadRun: boolPtr(true),
PromoteArtifacts: []config.ArchivePromotionRule{ PromoteArtifacts: []config.ArchivePromotionRule{
{From: "transcripts/trimmed.json", To: "transcripts/trimmed.json", Required: boolPtr(true)}, {Source: "narratio.transcript.trimmed", Dest: "transcripts/trimmed.json", Required: boolPtr(true)},
{From: "artifacts/session_recap.md", To: "artifacts/session_recap.md", Required: boolPtr(true)}, {Source: "narratio.artifact.session_recap", Dest: "artifacts/session_recap.md", Required: boolPtr(true)},
},
},
Scriptorium: &config.ScriptoriumConfig{
Artifacts: map[string]config.ScriptoriumArtifactConfig{
"session_recap": {
Enabled: true,
PromptID: "dnd.session_recap",
OutputPath: "artifacts/session_recap.md",
},
}, },
}, },
}, },