Files
narratio/docs/operations.md

11 KiB

Operations

This guide describes the implemented operator lifecycle for Narratio.

For field-level configuration, see docs/config.md. For full command/flag reference, see docs/cli.md.

Normal workflow (S3-first path)

  1. Create or upload session.yml, or pass a local session.yml explicitly.
  2. Upload session .flac files to object storage under the configured session audio prefix.
  3. Run Narratio:
narratio run --session-id 2026-04-04
  1. Read success output:
  • narratio run: session <session_id>; executed=<n> skipped=<n>; manifest=<path>
  • use manifest=<path> with status for inspection.

Notes:

  • default config/campaign/session discovery checks system config locations unless --config, --campaign, and --session are passed.
  • when local session.yml discovery misses, --session-id loads remote session.yml from {root_prefix}/campaigns/{campaign}/sessions/{session_id}/session.yml.
  • S3 audio mode requires session.inputs.audio_s3.prefix and valid object-store access.

Initialize a remote session skeleton:

narratio session init --config /etc/narratio/pipeline.yml --campaign /etc/narratio/campaign.yml --session-id 2026-04-04 --remote

Remote init writes {root_prefix}/campaigns/{campaign}/sessions/{session_id}/session.yml. It fails if the object already exists unless --force is passed.

Validate before running:

narratio session validate --session-id 2026-04-04

Restore workflow

Use restore when local durable session state is missing or stale and archive current state is authoritative.

Dry-run (no local writes):

narratio restore --session-id 2026-04-04 --dry-run

Execution:

narratio restore --session-id 2026-04-04

Post-restore analyze rerun pattern:

narratio analyze --session-id 2026-04-04

Restore source-of-truth:

  • remote commit marker: current/run_id.txt
  • remote current manifest: current/manifest.json
  • configured previous-session requirements are reconstructed from the previous session's remote current/ state, not from archived previous/** objects in the current session.

Restore default scope:

  • includes manifest.json, transcripts/**, and artifacts/** from the current session archive.
  • includes previous/** only when configured previous-session artifact inputs require it; restore hydrates those files the same way prepare would.
  • includes audio/** only with --include-audio
  • excludes runs/**, logs/**, reports/**, config/**, inputs/**, and current/** (except remote current/manifest.json as source)

Reset local state before restore testing:

narratio clean --session-id 2026-04-04 --dry-run
narratio clean --session-id 2026-04-04
narratio restore --session-id 2026-04-04 --include-audio

clean removes the local session work directory and session spool directory. It preserves the durable S3 audio cache by default, so repeated restore or forced prepare tests do not re-download large audio files.

Local filesystem layout and state artifacts

Session root:

  • {workspace.root}/work/{campaign}/{session_id}/

Primary state:

  • manifest.json: session-level stage state.
  • runs/{run_id}/manifest.json: invocation-level state.
  • .lock: session lock while a modifying command is active.
  • inputs/campaign.yml, inputs/session.yml, and inputs/pipeline.resolved.yml: materialized config inputs for the run.

Canonical session directories:

  • inputs/
  • audio/
  • transcripts/
  • artifacts/
  • previous/
  • reports/
  • logs/
  • config/
  • current/
  • runs/

Run-local stage directories:

  • runs/{run_id}/{stage}/ with stage-local outputs/, logs/, reports/, config/, scratch/.

Behavior:

  • directory creation is idempotent.
  • stage outputs are generally generated run-local first, then promoted to canonical paths on success.
  • restore installs downloaded files to canonical session paths and does not recreate historical run sandboxes.

Analyze artifact execution lifecycle

Analyze executes configured artifacts from pipeline.scriptorium.artifacts.

Execution model:

  • executable set = enabled artifacts, filtered by --artifacts when provided.
  • artifact-to-artifact dependencies are declared via depends_on.
  • selected artifacts run in deterministic dependency order.
  • after each successful artifact run, output is promoted to configured canonical output_path.

Configured artifact source reuse:

  • a non-executable configured artifact can satisfy inputs if its configured output file already exists and is valid.
  • reused configured artifact provenance is filesystem.disabled_artifact_output.

--artifacts behavior:

  • accepted on run, resume, run-stage analyze, and analyze.
  • filters analyze execution only.
  • does not imply force on run, resume, or run-stage; narratio analyze is force-by-design.
  • publish does not accept --artifacts; it is a force-by-design archive rerun.

Canonical previous-session input behavior:

  • canonical sources use narratio.previous_session.artifact.<artifact_key>.
  • these inputs are hydrated by prepare and by restore; analyze expects the local previous cache to already exist.
  • if analyze fails due to missing canonical previous cache, rerun:
    • narratio run-stage --session-id <id> --force prepare
    • or narratio restore --session-id <id> when remote archive current state is authoritative.

Remote archive layout and publish contract

Preferred manual publish command:

narratio publish --session-id <id>

publish is equivalent to narratio run-stage --force archive; use run-stage when you need the general single-stage command form.

When archive is enabled and run upload is enabled, archive publishes under:

  • session prefix: {root_prefix}/campaigns/{campaign}/sessions/{session_id}/
  • run prefix: {session_prefix}/runs/{run_id}/

Archive uploads:

  • run record files from run root (excluding audio/).
  • promoted files from explicit archive.promote_artifacts rules.
  • mutable session locks from helper commands live at {session_prefix}/locks.yml.

Publish order:

  1. upload current/manifest.json
  2. upload current/run_id.txt last

current/run_id.txt is the remote commit marker.

Archive promotion is explicit and source-based:

  • Narratio does not auto-promote all generated analyze artifacts.
  • each rule resolves source through the artifact resolver/catalog model, then uploads to dest.
  • missing required promotion sources fail archive stage.
  • missing optional promotion sources are skipped.
  • invalid resolved artifacts fail archive stage.
  • archive.locks skips top-level promotion overwrites for locked sources while run-local uploads still publish.
  • remote locks from {session_prefix}/locks.yml are merged with static archive.locks; static locks win on duplicate sources.
  • locked required promotions are treated as intentional successful skips and are recorded in archive metadata.

Lock helper behavior:

  • narratio locks --session-id <id> lists effective static and remote locks.
  • narratio locks add --session-id <id> --reason <text> <source> writes a remote lock.
  • narratio locks add --session-id <id> --force --reason <text> <source> updates an existing remote lock reason.
  • narratio locks remove --session-id <id> <source> removes only a remote lock.
  • locks remove cannot remove static pipeline locks.
  • remote lock writes check whether the lock store exists, but are not compare-and-swap atomic.

Resume, retry, restore, and safe rerun behavior

Default skip:

  • run and run-stage skip already-succeeded stages unless --force is set.

Resume:

  • resume starts at first non-succeeded stage.
  • resume --force runs full stage order.

Restore conflict policy:

  • restore classifies local differences as conflicts.
  • without --force, restore fails when conflicts exist.
  • with --force, conflicting local files are overwritten by remote archive files.

Forced reruns:

  • force-rerunning an upstream succeeded stage marks downstream succeeded stages as stale.
  • ordinary --force does not override archive locks.

Safe rerun pattern:

  1. rerun the changed stage with --force.
  2. run resume to rebuild downstream stages.

Cleanup behavior

Automatic post-archive cleanup is considered only when archive stage executed and succeeded.

Automatic cleanup toggles:

  • pipeline.spool.delete_audio_after_archive=true deletes run-scoped spool audio.
  • pipeline.workspace.cleanup_after_archive=true deletes run-scoped local run directory.

Manual cleanup:

  • narratio clean --session-id <id> deletes {workspace.root}/work/{campaign}/{session_id} and {spool.root}/{campaign}/{session_id}.
  • narratio clean --all deletes all local session work under {workspace.root}/work and all spool children under {spool.root}.
  • --dry-run prints targets without deleting.
  • --clear-cache also removes matching S3 audio cache files. Without it, cache is preserved.

The S3 audio cache under pipeline.cache.root is durable input cache state, not workspace or spool state. Automatic cleanup and default manual cleanup do not delete it.

Cleanup eligibility gates:

  • archive enabled
  • archive run upload enabled
  • run record upload completed
  • current pointer write completed (current/run_id.txt written)

No cleanup for failed/incomplete/unarchived/archive-skipped runs.

Failure and recovery playbooks

After run failure, Narratio keeps:

  • session manifest
  • run manifest
  • run-local artifacts/logs/config/reports

Failed or incomplete runs remain local-only.

After restore failure:

  • already-installed restore files remain in place.
  • restore does not roll back prior successful installs.
  • existing local manifest is preserved if restored manifest validation/install fails.

Recommended recovery:

  1. inspect state:
narratio status --session-id 2026-04-04

This reports local manifest state, committed remote current state, expected remote transcript/artifact availability, and archive locks.

  1. for one manifest file, run:
narratio status --manifest <manifest-path>
  1. for restore-specific checks, run:
narratio restore --session-id 2026-04-04 --dry-run
  1. fix root cause (config/input/credentials/storage/service availability).
  2. continue with resume, or targeted run-stage --force followed by resume.

Restore report

Non-dry-run restore writes a durable report at:

  • reports/restore-latest.json

Report content includes:

  • identity (campaign, session_id, run_id)
  • mode flags (dry_run, force, include_audio)
  • plan counts and execution counts
  • per-action status

Dry-run does not write restore report files.

Operational caveats

  • status with no config/session flags still requires explicit --manifest.
  • status --session-id <id> uses normal config/session loading, including remote session fallback.
  • status --session-id <id> includes the same promoted remote output availability view as artifacts list --remote when storage is configured.
  • local and S3 audio input modes are mutually exclusive.
  • archive publish requires upstream stages through analyze to be succeeded.
  • required promotion rules can fail when selected analyze artifacts did not generate a required file path.
  • restore requires configured remote object storage and committed remote current state.