211 lines
7.7 KiB
Markdown
211 lines
7.7 KiB
Markdown
# Roadmap: Transcript Artifact Naming
|
|
|
|
Status: Implemented
|
|
|
|
## Problem
|
|
|
|
Narratio's built-in transcript artifact names and canonical paths currently mix
|
|
operator-facing artifact meaning with historical stage and tool terminology:
|
|
|
|
- `narratio.transcript.merged` maps to `transcripts/merged.json`.
|
|
- `narratio.transcript.polished` maps to `transcripts/processed.json`.
|
|
- `narratio.transcript.full` maps to `transcripts/normalized.json`.
|
|
- `narratio.transcript.trimmed` maps to `transcripts/trimmed.json`.
|
|
|
|
This makes the public artifact surface harder to reason about. Operators see
|
|
`full`, `normalized`, `processed`, `polished`, `merged`, and `trimmed` used in
|
|
different places for the same transcript lineage.
|
|
|
|
The transcript source IDs, canonical paths, and manifest output kinds should
|
|
use one vocabulary based on each transcript's role in the session artifact
|
|
model.
|
|
|
|
## Target Model
|
|
|
|
Built-in transcript artifacts should use these public source IDs, canonical
|
|
paths, and manifest output kinds:
|
|
|
|
| Source ID | Canonical path | Output kind | Meaning |
|
|
| --- | --- | --- | --- |
|
|
| `narratio.transcript.base` | `transcripts/base.json` | `transcript_base` | First unified transcript produced by merging per-speaker raw transcripts. |
|
|
| `narratio.transcript.polished` | `transcripts/polished.json` | `transcript_polished` | Audita-polished transcript. |
|
|
| `narratio.transcript.final` | `transcripts/final.json` | `transcript_final` | Full final transcript after normalization. |
|
|
| `narratio.transcript.final_trimmed` | `transcripts/final.trimmed.json` | `transcript_final_trimmed` | Trimmed version of the final transcript. |
|
|
|
|
Stage names remain process-oriented and unchanged:
|
|
|
|
- `merge`
|
|
- `polish`
|
|
- `normalize`
|
|
- `trim`
|
|
|
|
Downstream adapter contracts also remain process-oriented. The rename changes
|
|
Narratio's artifact model, canonical paths, config examples, archive promotion
|
|
sources, lock sources, status output, and documentation. It should not rename
|
|
the stages themselves or move external integration details into stage logic.
|
|
|
|
## Compatibility Policy
|
|
|
|
This is a hard cutover.
|
|
|
|
After implementation, these old source IDs should be rejected:
|
|
|
|
- `narratio.transcript.merged`
|
|
- `narratio.transcript.full`
|
|
- `narratio.transcript.trimmed`
|
|
|
|
These old canonical paths should not be compatibility fallbacks:
|
|
|
|
- `transcripts/merged.json`
|
|
- `transcripts/processed.json`
|
|
- `transcripts/normalized.json`
|
|
- `transcripts/trimmed.json`
|
|
|
|
Existing remote archives are not migrated automatically. Operators who want
|
|
new promoted keys for old sessions should republish those sessions after
|
|
updating configuration.
|
|
|
|
## Implementation Stages
|
|
|
|
### Stage 1: Centralize Transcript Artifact Naming
|
|
|
|
Status: Implemented
|
|
|
|
Consolidate transcript artifact source IDs, canonical paths, and output kinds
|
|
in the artifact/path layer before changing runtime behavior.
|
|
|
|
Implementation requirements:
|
|
|
|
- Add or consolidate constants/helpers for built-in transcript source IDs.
|
|
- Add or consolidate constants/helpers for canonical transcript paths.
|
|
- Add or consolidate constants/helpers for transcript manifest output kinds.
|
|
- Keep source ID, path, and output-kind mappings in one registry or one
|
|
obviously shared artifact model.
|
|
- Update artifact registry tests to prove the target mapping.
|
|
- Avoid changing stage output behavior in this stage unless the implementation
|
|
is simpler and still reviewable.
|
|
|
|
Acceptance criteria:
|
|
|
|
- There is one clear source of truth for built-in transcript artifact names,
|
|
paths, and output kinds.
|
|
- Tests prove the new target mapping in the artifact layer.
|
|
- No generic workflow abstraction is introduced.
|
|
|
|
### Stage 2: Rename Runtime Outputs and Defaults
|
|
|
|
Status: Implemented
|
|
|
|
Switch runtime behavior to the new transcript artifact model.
|
|
|
|
Implementation requirements:
|
|
|
|
- Update `merge` to write and record `transcripts/base.json` with
|
|
`transcript_base`.
|
|
- Update `polish` to write and record `transcripts/polished.json` with
|
|
`transcript_polished`.
|
|
- Update `normalize` to write and record `transcripts/final.json` with
|
|
`transcript_final`.
|
|
- Update `trim` to write and record `transcripts/final.trimmed.json` with
|
|
`transcript_final_trimmed`.
|
|
- Update normalize and trim defaults to:
|
|
- `pipeline.normalize.output_path: transcripts/final.json`
|
|
- `pipeline.trim.output_path: transcripts/final.trimmed.json`
|
|
- Update built-in artifact resolution, archive promotion destination
|
|
derivation, archive locks, status output, artifact catalog output,
|
|
previous-cache resolution, restore planning, and restore execution to use
|
|
the new registry values.
|
|
- Ensure old source IDs fail config validation.
|
|
|
|
Acceptance criteria:
|
|
|
|
- New runs produce the target canonical transcript files.
|
|
- Manifest outputs use the target output kinds.
|
|
- Archive promotion and lock validation accept new source IDs and reject old
|
|
source IDs.
|
|
- Status and artifact listing display new source IDs.
|
|
- Restore uses the new canonical paths and does not restore old transcript
|
|
paths as canonical outputs.
|
|
|
|
### Stage 3: Update Tests, Examples, and Current Documentation
|
|
|
|
Status: Implemented
|
|
|
|
Update all implemented-behavior references after the runtime cutover lands.
|
|
|
|
Implementation requirements:
|
|
|
|
- Update examples to use `narratio.transcript.final_trimmed` and
|
|
`transcripts/final.trimmed.json` where trimmed final transcript is intended.
|
|
- Update examples that refer to full final transcripts to use
|
|
`narratio.transcript.final` and `transcripts/final.json`.
|
|
- Update `docs/config.md`, `docs/internal/artifacts.md`, stage docs,
|
|
CLI examples, operations examples, archive examples, lock examples, and
|
|
status/artifact-list examples.
|
|
- Add strict validation tests proving old source IDs are rejected.
|
|
- Mark roadmap stages implemented only after code, tests, examples, and
|
|
current-behavior docs agree.
|
|
|
|
Acceptance criteria:
|
|
|
|
- Maintained examples load and validate.
|
|
- Current-behavior docs describe only implemented new names.
|
|
- Old names remain only in this roadmap as historical/planning context until
|
|
this roadmap is retired or archived.
|
|
|
|
## Test Guidance
|
|
|
|
Run focused tests while implementing:
|
|
|
|
- `go test ./internal/artifacts -v`
|
|
- `go test ./internal/config -v`
|
|
- `go test ./internal/stage -v`
|
|
- `go test ./internal/app -v`
|
|
|
|
Run full validation before finishing:
|
|
|
|
- `go test ./...`
|
|
|
|
Run final searches:
|
|
|
|
- Old source IDs:
|
|
- `narratio.transcript.merged`
|
|
- `narratio.transcript.full`
|
|
- `narratio.transcript.trimmed`
|
|
- Old paths:
|
|
- `transcripts/merged.json`
|
|
- `transcripts/processed.json`
|
|
- `transcripts/normalized.json`
|
|
- `transcripts/trimmed.json`
|
|
- Old output kinds:
|
|
- `transcript_merged`
|
|
- `transcript_processed`
|
|
- `transcript_normalized`
|
|
- `transcript_trimmed`
|
|
|
|
Expected remaining matches should be limited to this roadmap's
|
|
historical/planning references until the roadmap is fully completed.
|
|
|
|
## Architecture Guardrails
|
|
|
|
- Keep Narratio explicit and stage-driven; do not introduce a generic workflow
|
|
or DAG abstraction.
|
|
- Keep path and artifact naming in centralized helpers rather than scattered
|
|
string concatenation.
|
|
- Preserve manifest-driven resume behavior.
|
|
- Keep storage details behind storage adapters.
|
|
- Do not move Seriatim, Audita, or Scriptorium command details out of their
|
|
adapter boundaries.
|
|
- Keep current-behavior documentation in sync only after implementation lands;
|
|
planned behavior belongs in this roadmap until then.
|
|
|
|
## Assumptions
|
|
|
|
- The cutover is intentionally not backward-compatible.
|
|
- Existing remote archive objects are not renamed or migrated automatically.
|
|
- Stage names and downstream adapter request field names remain unchanged.
|
|
- The term `base` is preferred over `merged` for the first unified transcript.
|
|
- The term `final` is preferred over `full` or `normalized` for the full final
|
|
transcript.
|
|
- The trimmed final path is `transcripts/final.trimmed.json`.
|