Compare commits
9 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| b1eb37d80d | |||
| 0fc92f3643 | |||
| 0dfd06c349 | |||
| da3720693d | |||
| 6dfc1ea527 | |||
| 761d70bbc6 | |||
| 451cc19418 | |||
| c37ea70dcb | |||
| a90859114a |
@@ -2,7 +2,10 @@
|
|||||||
|
|
||||||
`seriatim` is a Go CLI for transcript artifact processing.
|
`seriatim` is a Go CLI for transcript artifact processing.
|
||||||
|
|
||||||
It merges per-speaker WhisperX-style JSON into one deterministic transcript, trims existing seriatim artifacts by segment ID, and normalizes transcript-like JSON into standard seriatim output schemas.
|
It merges per-speaker WhisperX-style JSON into deterministic seriatim JSON,
|
||||||
|
trims existing seriatim artifacts by segment ID, normalizes transcript-like JSON
|
||||||
|
into supported output schemas, and renders existing seriatim artifacts as
|
||||||
|
human-readable Markdown.
|
||||||
|
|
||||||
## Quickstart
|
## Quickstart
|
||||||
|
|
||||||
@@ -20,6 +23,7 @@ go run ./cmd/seriatim merge \
|
|||||||
- `merge`: merge one or more input transcript JSON files.
|
- `merge`: merge one or more input transcript JSON files.
|
||||||
- `trim`: keep/remove segment IDs from an existing seriatim artifact.
|
- `trim`: keep/remove segment IDs from an existing seriatim artifact.
|
||||||
- `normalize`: canonicalize transcript-like JSON into a seriatim artifact.
|
- `normalize`: canonicalize transcript-like JSON into a seriatim artifact.
|
||||||
|
- `render`: render an existing seriatim artifact as Markdown.
|
||||||
|
|
||||||
## Documentation
|
## Documentation
|
||||||
|
|
||||||
|
|||||||
39
docs/cli.md
39
docs/cli.md
@@ -16,6 +16,7 @@ go run ./cmd/seriatim merge \
|
|||||||
| `merge` | Merge one or more raw transcript JSON inputs into one seriatim artifact. |
|
| `merge` | Merge one or more raw transcript JSON inputs into one seriatim artifact. |
|
||||||
| `trim` | Keep or remove segment IDs from an existing seriatim artifact. |
|
| `trim` | Keep or remove segment IDs from an existing seriatim artifact. |
|
||||||
| `normalize` | Canonicalize transcript-like JSON into a seriatim artifact. |
|
| `normalize` | Canonicalize transcript-like JSON into a seriatim artifact. |
|
||||||
|
| `render` | Render an existing seriatim artifact as Markdown. |
|
||||||
|
|
||||||
Root usage:
|
Root usage:
|
||||||
|
|
||||||
@@ -130,6 +131,35 @@ Flags:
|
|||||||
- Does not run merge modules.
|
- Does not run merge modules.
|
||||||
- When `--output-schema` is omitted, schema resolution is: `SERIATIM_OUTPUT_SCHEMA` -> default `seriatim-intermediate`.
|
- When `--output-schema` is omitted, schema resolution is: `SERIATIM_OUTPUT_SCHEMA` -> default `seriatim-intermediate`.
|
||||||
|
|
||||||
|
## `render`
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
|
||||||
|
```text
|
||||||
|
seriatim render [flags]
|
||||||
|
```
|
||||||
|
|
||||||
|
Flags:
|
||||||
|
|
||||||
|
| Flag | Required | Default | Description |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `--input-file string` | Yes | none | Input seriatim artifact JSON file. |
|
||||||
|
| `--output-file string` | Yes | none | Rendered output file path. |
|
||||||
|
| `--format string` | Yes | none | Output format. Current supported value: `markdown`. |
|
||||||
|
| `--title string` | No | `Transcript` | Markdown document title. |
|
||||||
|
| `--include-timestamps` | No | `true` | Include `[HH:MM:SS–HH:MM:SS]` per segment. |
|
||||||
|
| `--include-segment-ids` | No | `false` | Include `[#id]` marker per segment. |
|
||||||
|
| `--include-metadata` | No | `false` | Include artifact metadata block near the top. |
|
||||||
|
|
||||||
|
`render` behavior:
|
||||||
|
|
||||||
|
- Input must be a valid existing seriatim output artifact (`seriatim-minimal`, `seriatim-intermediate`, or `seriatim-full`).
|
||||||
|
- Raw WhisperX-style JSON is rejected.
|
||||||
|
- `render` does not execute merge/trim/normalize transformations.
|
||||||
|
- `render` has no `--report-file` output in the current implementation.
|
||||||
|
- Markdown output is deterministic for the same input artifact and render flags.
|
||||||
|
- Category names are not printed directly; `background`, `backchannel`, and `filler` only influence italics.
|
||||||
|
|
||||||
## Common workflows
|
## Common workflows
|
||||||
|
|
||||||
Merge with a speaker map and report output:
|
Merge with a speaker map and report output:
|
||||||
@@ -160,6 +190,15 @@ go run ./cmd/seriatim normalize \
|
|||||||
--output-file /tmp/seriatim-example-normalize-object.json
|
--output-file /tmp/seriatim-example-normalize-object.json
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Render an existing artifact as Markdown:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
go run ./cmd/seriatim render \
|
||||||
|
--input-file examples/render/input-intermediate.json \
|
||||||
|
--output-file /tmp/seriatim-example-render.md \
|
||||||
|
--format markdown
|
||||||
|
```
|
||||||
|
|
||||||
## Exit and errors
|
## Exit and errors
|
||||||
|
|
||||||
- Commands return exit code `0` on success.
|
- Commands return exit code `0` on success.
|
||||||
|
|||||||
@@ -23,6 +23,18 @@ For `trim`:
|
|||||||
- If `--output-schema` is omitted, output preserves the input artifact schema.
|
- If `--output-schema` is omitted, output preserves the input artifact schema.
|
||||||
- If `--output-schema` is set, it must be one of `seriatim-minimal`, `seriatim-intermediate`, `seriatim-full`.
|
- If `--output-schema` is set, it must be one of `seriatim-minimal`, `seriatim-intermediate`, `seriatim-full`.
|
||||||
|
|
||||||
|
## Render format and defaults
|
||||||
|
|
||||||
|
`render` requires `--input-file`, `--output-file`, and `--format`.
|
||||||
|
Current supported format value is `markdown`.
|
||||||
|
|
||||||
|
Render defaults:
|
||||||
|
|
||||||
|
- `--title`: `Transcript`
|
||||||
|
- `--include-timestamps`: `true`
|
||||||
|
- `--include-segment-ids`: `false`
|
||||||
|
- `--include-metadata`: `false`
|
||||||
|
|
||||||
## Merge module defaults
|
## Merge module defaults
|
||||||
|
|
||||||
Default merge module selections:
|
Default merge module selections:
|
||||||
@@ -152,6 +164,11 @@ All commands:
|
|||||||
- Validates `--output-schema` through the same schema set as `merge`.
|
- Validates `--output-schema` through the same schema set as `merge`.
|
||||||
- Currently accepts only `json` in `--output-modules`.
|
- Currently accepts only `json` in `--output-modules`.
|
||||||
|
|
||||||
|
`render`:
|
||||||
|
|
||||||
|
- Requires `--input-file`, `--output-file`, and `--format`.
|
||||||
|
- Validates `--format` as `markdown`.
|
||||||
|
|
||||||
## Related docs
|
## Related docs
|
||||||
|
|
||||||
- CLI reference: [cli.md](cli.md)
|
- CLI reference: [cli.md](cli.md)
|
||||||
|
|||||||
@@ -8,7 +8,8 @@ seriatim emits one of three public JSON output contracts:
|
|||||||
- `seriatim-intermediate`
|
- `seriatim-intermediate`
|
||||||
- `seriatim-full`
|
- `seriatim-full`
|
||||||
|
|
||||||
These are used by `merge`, `trim`, and `normalize`.
|
These are used by `merge`, `trim`, and `normalize`, and are accepted as input
|
||||||
|
by `render`.
|
||||||
|
|
||||||
## Schema roles
|
## Schema roles
|
||||||
|
|
||||||
|
|||||||
@@ -2,8 +2,8 @@
|
|||||||
|
|
||||||
## Purpose
|
## Purpose
|
||||||
|
|
||||||
Describes public artifact conversion and validation internals for merge output,
|
Describes implemented artifact parsing, conversion, validation, and render-model
|
||||||
trim, and normalize.
|
normalization internals.
|
||||||
|
|
||||||
## Artifact contracts
|
## Artifact contracts
|
||||||
|
|
||||||
@@ -19,35 +19,40 @@ Machine-readable schemas:
|
|||||||
- `schema/intermediate-output.schema.json`
|
- `schema/intermediate-output.schema.json`
|
||||||
- `schema/minimal-output.schema.json`
|
- `schema/minimal-output.schema.json`
|
||||||
|
|
||||||
## Schema selection
|
## Shared output-artifact parser
|
||||||
|
|
||||||
Merge pipeline conversion uses `internal/artifact.SelectedFromMerged`:
|
`internal/artifact/output_artifact.go` provides schema-aware parsing for
|
||||||
|
existing seriatim output artifacts.
|
||||||
|
|
||||||
|
Behavior:
|
||||||
|
|
||||||
|
- accepts only valid full, intermediate, or minimal seriatim output artifacts
|
||||||
|
- validates through `schema` semantic + JSON schema checks
|
||||||
|
- rejects malformed JSON
|
||||||
|
- rejects raw WhisperX-style JSON and other non-seriatim shapes
|
||||||
|
|
||||||
|
Consumers:
|
||||||
|
|
||||||
|
- `internal/trim` artifact-level trim flow
|
||||||
|
- `internal/render` artifact-level render flow
|
||||||
|
|
||||||
|
## Merge conversion behavior
|
||||||
|
|
||||||
|
`internal/artifact/transcript.go` converts `model.MergedTranscript` to public
|
||||||
|
contracts:
|
||||||
|
|
||||||
|
- full schema preserves source/provenance, overlap groups, and metadata module
|
||||||
|
lists
|
||||||
|
- intermediate schema emits segment timing/text/speaker with optional
|
||||||
|
categories and compact metadata
|
||||||
|
- minimal schema emits compact segment timing/text/speaker and compact metadata
|
||||||
|
|
||||||
|
Schema selection uses `internal/artifact.SelectedFromMerged`:
|
||||||
|
|
||||||
- `seriatim-full` -> `artifact.FromMerged`
|
- `seriatim-full` -> `artifact.FromMerged`
|
||||||
- `seriatim-intermediate` -> `artifact.IntermediateFromMerged`
|
- `seriatim-intermediate` -> `artifact.IntermediateFromMerged`
|
||||||
- `seriatim-minimal` -> `artifact.MinimalFromMerged`
|
- `seriatim-minimal` -> `artifact.MinimalFromMerged`
|
||||||
|
- unknown/empty -> intermediate fallback
|
||||||
Unknown/empty selection falls back to intermediate conversion.
|
|
||||||
|
|
||||||
## Merge conversion behavior
|
|
||||||
|
|
||||||
`internal/artifact` converts `model.MergedTranscript` to public contracts:
|
|
||||||
|
|
||||||
- full schema preserves source/provenance, overlap groups, and metadata module
|
|
||||||
lists.
|
|
||||||
- intermediate schema emits segment timing/text/speaker with optional
|
|
||||||
categories and compact metadata.
|
|
||||||
- minimal schema emits compact segment timing/text/speaker and compact
|
|
||||||
metadata.
|
|
||||||
|
|
||||||
## Validation behavior
|
|
||||||
|
|
||||||
`schema/output.go` validates both structure and semantics:
|
|
||||||
|
|
||||||
- embedded JSON Schema validation via `jsonschema/v6`
|
|
||||||
- semantic checks for sequential segment IDs starting at `1`
|
|
||||||
- semantic checks for non-inverted segment timing (`end >= start`)
|
|
||||||
- full schema overlap-group timing checks (`group.end >= group.start`)
|
|
||||||
|
|
||||||
## Trim internals
|
## Trim internals
|
||||||
|
|
||||||
@@ -72,9 +77,8 @@ Apply layer (`apply.go`):
|
|||||||
- schema-specific segment reconstruction for full/intermediate/minimal outputs
|
- schema-specific segment reconstruction for full/intermediate/minimal outputs
|
||||||
- overlap-group recomputation only for full-schema outputs
|
- overlap-group recomputation only for full-schema outputs
|
||||||
|
|
||||||
Artifact layer (`artifact.go`):
|
Artifact conversion layer (`artifact.go`):
|
||||||
|
|
||||||
- schema detection for full/intermediate/minimal artifacts
|
|
||||||
- schema-preserving trim application
|
- schema-preserving trim application
|
||||||
- supported schema conversions:
|
- supported schema conversions:
|
||||||
- full -> intermediate/minimal
|
- full -> intermediate/minimal
|
||||||
@@ -85,10 +89,10 @@ Artifact layer (`artifact.go`):
|
|||||||
|
|
||||||
Trim invariants:
|
Trim invariants:
|
||||||
|
|
||||||
- selected IDs must exist in input.
|
- selected IDs must exist in input
|
||||||
- input IDs must be positive, unique, sequential.
|
- input IDs must be positive, unique, sequential
|
||||||
- retained segment order follows input transcript order.
|
- retained segment order follows input transcript order
|
||||||
- output IDs are reassigned to `1..N`.
|
- output IDs are reassigned to `1..N`
|
||||||
|
|
||||||
## Normalize internals
|
## Normalize internals
|
||||||
|
|
||||||
@@ -117,13 +121,64 @@ Run layer (`normalize.go`):
|
|||||||
|
|
||||||
Normalize invariant:
|
Normalize invariant:
|
||||||
|
|
||||||
- report events do not embed transcript text.
|
- report events do not embed transcript text
|
||||||
|
|
||||||
|
## Render internals
|
||||||
|
|
||||||
|
`internal/render` is an artifact-level, downstream-only renderer.
|
||||||
|
|
||||||
|
Model normalization (`normalize.go`):
|
||||||
|
|
||||||
|
- converts full/intermediate/minimal artifacts into a common render model
|
||||||
|
- preserves segment order and segment IDs
|
||||||
|
- normalizes per-segment fields to ID, start, end, speaker, text, categories
|
||||||
|
- emits empty categories slice when categories are absent in input
|
||||||
|
|
||||||
|
Renderer registry (`registry.go`):
|
||||||
|
|
||||||
|
- resolves renderers by public format name
|
||||||
|
- currently registers `markdown`
|
||||||
|
|
||||||
|
Markdown renderer (`markdown.go`):
|
||||||
|
|
||||||
|
- writes title header `# {title}`
|
||||||
|
- renders optional `[HH:MM:SS–HH:MM:SS]` timestamps
|
||||||
|
- renders optional `[#id]` segment references
|
||||||
|
- renders `**speaker:** text`
|
||||||
|
- italicizes text when categories include `background`, `backchannel`, or
|
||||||
|
`filler`
|
||||||
|
- ignores unknown categories
|
||||||
|
- optionally includes metadata summary block
|
||||||
|
|
||||||
|
Run layer (`run.go`):
|
||||||
|
|
||||||
|
1. Read input artifact JSON.
|
||||||
|
2. Parse via shared output-artifact parser.
|
||||||
|
3. Normalize to render model.
|
||||||
|
4. Resolve renderer by `--format`.
|
||||||
|
5. Render text output.
|
||||||
|
6. Write output file.
|
||||||
|
|
||||||
|
Render invariants:
|
||||||
|
|
||||||
|
- does not run merge/trim/normalize modules
|
||||||
|
- does not expose report output
|
||||||
|
- deterministic for identical input artifact and render flags
|
||||||
|
|
||||||
|
## Validation behavior
|
||||||
|
|
||||||
|
`schema/output.go` validates both structure and semantics:
|
||||||
|
|
||||||
|
- embedded JSON Schema validation via `jsonschema/v6`
|
||||||
|
- semantic checks for sequential segment IDs starting at `1`
|
||||||
|
- semantic checks for non-inverted segment timing (`end >= start`)
|
||||||
|
- full schema overlap-group timing checks (`group.end >= group.start`)
|
||||||
|
|
||||||
## Boundaries
|
## Boundaries
|
||||||
|
|
||||||
- CLI flag semantics belong to `docs/cli.md`.
|
- CLI flag semantics belong to `docs/cli.md`.
|
||||||
- Runtime config/env surfaces belong to `docs/config.md`.
|
- Runtime config/env surfaces belong to `docs/config.md`.
|
||||||
- This doc describes internal conversion/validation behavior only.
|
- This document describes internal conversion/validation behavior only.
|
||||||
|
|
||||||
## Failure behavior
|
## Failure behavior
|
||||||
|
|
||||||
@@ -133,22 +188,29 @@ Representative failure classes:
|
|||||||
- schema validation failure for parsed artifact or built output
|
- schema validation failure for parsed artifact or built output
|
||||||
- unsupported schema conversion path (trim)
|
- unsupported schema conversion path (trim)
|
||||||
- selector or input-ID consistency errors (trim)
|
- selector or input-ID consistency errors (trim)
|
||||||
|
- unsupported renderer format (render)
|
||||||
- output/report file write failures from command paths
|
- output/report file write failures from command paths
|
||||||
|
|
||||||
## Tests to inspect before changes
|
## Tests to inspect before changes
|
||||||
|
|
||||||
- `schema/output_test.go`
|
- `schema/output_test.go`
|
||||||
- `internal/artifact/transcript_test.go`
|
- `internal/artifact/transcript_test.go`
|
||||||
|
- `internal/artifact/output_artifact_test.go`
|
||||||
- `internal/trim/selector_test.go`
|
- `internal/trim/selector_test.go`
|
||||||
- `internal/trim/artifact_test.go`
|
- `internal/trim/artifact_test.go`
|
||||||
- `internal/trim/apply_test.go`
|
- `internal/trim/apply_test.go`
|
||||||
- `internal/normalize/parse_test.go`
|
- `internal/normalize/parse_test.go`
|
||||||
|
- `internal/render/normalize_test.go`
|
||||||
|
- `internal/render/markdown_test.go`
|
||||||
|
- `internal/render/registry_test.go`
|
||||||
- `internal/cli/trim_test.go`
|
- `internal/cli/trim_test.go`
|
||||||
- `internal/cli/normalize_test.go`
|
- `internal/cli/normalize_test.go`
|
||||||
|
- `internal/cli/render_test.go`
|
||||||
|
|
||||||
## Invariants
|
## Invariants
|
||||||
|
|
||||||
- Public artifacts are validated through `schema` before acceptance.
|
- Public artifacts are validated through `schema` before acceptance.
|
||||||
- Segment IDs in emitted artifacts are sequential and deterministic.
|
- Segment IDs in emitted artifacts are sequential and deterministic.
|
||||||
- Internal-only fields are not emitted in minimal/intermediate contracts.
|
- Internal-only fields are not emitted in minimal/intermediate contracts.
|
||||||
- Trim and normalize stay artifact-level and do not execute merge modules.
|
- Trim, normalize, and render stay artifact-level and do not execute merge
|
||||||
|
modules.
|
||||||
|
|||||||
@@ -73,7 +73,8 @@ coalesce gap and overlap thresholds).
|
|||||||
- Pipeline does not parse CLI flags.
|
- Pipeline does not parse CLI flags.
|
||||||
- Pipeline does not normalize raw CLI strings.
|
- Pipeline does not normalize raw CLI strings.
|
||||||
- Pipeline delegates conversion to public output contracts to `internal/artifact`.
|
- Pipeline delegates conversion to public output contracts to `internal/artifact`.
|
||||||
- Artifact-level commands `trim` and `normalize` are outside this pipeline.
|
- Artifact-level commands `trim`, `normalize`, and `render` are outside this
|
||||||
|
pipeline.
|
||||||
|
|
||||||
## Failure behavior
|
## Failure behavior
|
||||||
|
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ This document covers runtime operation of the implemented CLI commands:
|
|||||||
- `merge`
|
- `merge`
|
||||||
- `trim`
|
- `trim`
|
||||||
- `normalize`
|
- `normalize`
|
||||||
|
- `render`
|
||||||
|
|
||||||
## Runtime model
|
## Runtime model
|
||||||
|
|
||||||
@@ -29,6 +30,7 @@ Command-specific expectations:
|
|||||||
- `merge`: requires at least one `--input-file`; optional `--speakers` and `--autocorrect` paths must exist when provided.
|
- `merge`: requires at least one `--input-file`; optional `--speakers` and `--autocorrect` paths must exist when provided.
|
||||||
- `trim`: input must be an existing valid seriatim artifact JSON file.
|
- `trim`: input must be an existing valid seriatim artifact JSON file.
|
||||||
- `normalize`: input must be a JSON object with `segments` or a top-level segment array.
|
- `normalize`: input must be a JSON object with `segments` or a top-level segment array.
|
||||||
|
- `render`: input must be an existing valid seriatim artifact JSON file.
|
||||||
|
|
||||||
## Normal workflow
|
## Normal workflow
|
||||||
|
|
||||||
@@ -80,11 +82,28 @@ go run ./cmd/seriatim normalize \
|
|||||||
--report-file normalize-report.json
|
--report-file normalize-report.json
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Render
|
||||||
|
|
||||||
|
1. Provide existing seriatim artifact with `--input-file`.
|
||||||
|
2. Provide `--output-file`.
|
||||||
|
3. Provide `--format markdown`.
|
||||||
|
4. Optionally provide `--title`, `--include-timestamps`, `--include-segment-ids`, and `--include-metadata`.
|
||||||
|
|
||||||
|
Example:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
go run ./cmd/seriatim render \
|
||||||
|
--input-file examples/render/input-intermediate.json \
|
||||||
|
--output-file /tmp/seriatim-example-render.md \
|
||||||
|
--format markdown
|
||||||
|
```
|
||||||
|
|
||||||
## Output and report artifacts
|
## Output and report artifacts
|
||||||
|
|
||||||
Primary output:
|
Primary outputs:
|
||||||
|
|
||||||
- `--output-file` writes JSON transcript artifact in selected schema.
|
- `merge`, `trim`, `normalize`: `--output-file` writes JSON transcript artifact in the selected schema.
|
||||||
|
- `render`: `--output-file` writes presentation Markdown.
|
||||||
|
|
||||||
Optional report output:
|
Optional report output:
|
||||||
|
|
||||||
@@ -92,6 +111,7 @@ Optional report output:
|
|||||||
- `merge` report metadata records reader/modules and event sequence.
|
- `merge` report metadata records reader/modules and event sequence.
|
||||||
- `trim` report includes a `trim-audit` event with mode/selector/counts and old-to-new ID mapping.
|
- `trim` report includes a `trim-audit` event with mode/selector/counts and old-to-new ID mapping.
|
||||||
- `normalize` report includes a `normalize-audit` event with input shape, repair stats, and output selection details.
|
- `normalize` report includes a `normalize-audit` event with input shape, repair stats, and output selection details.
|
||||||
|
- `render` has no report output in the current implementation.
|
||||||
|
|
||||||
## Failure and retry behavior
|
## Failure and retry behavior
|
||||||
|
|
||||||
@@ -106,9 +126,10 @@ Retry guidance:
|
|||||||
2. Re-run the same command.
|
2. Re-run the same command.
|
||||||
3. If a prior run created a partial or unwanted output/report file, remove it and rerun.
|
3. If a prior run created a partial or unwanted output/report file, remove it and rerun.
|
||||||
|
|
||||||
Operational note:
|
Operational notes:
|
||||||
|
|
||||||
- With identical inputs/config/version, merge behavior is deterministic and input files are sorted before processing.
|
- With identical inputs/config/version, `merge` behavior is deterministic and input files are sorted before processing.
|
||||||
|
- With identical input artifact and render flags, `render` output is deterministic.
|
||||||
|
|
||||||
## Cleanup
|
## Cleanup
|
||||||
|
|
||||||
@@ -124,6 +145,7 @@ Transcript artifacts and reports are local files and may contain sensitive conve
|
|||||||
- Store outputs in controlled directories with appropriate OS permissions.
|
- Store outputs in controlled directories with appropriate OS permissions.
|
||||||
- Share report files carefully; they include file paths and processing diagnostics.
|
- Share report files carefully; they include file paths and processing diagnostics.
|
||||||
- Normalize report events intentionally avoid embedding transcript text, but output artifacts contain transcript content.
|
- Normalize report events intentionally avoid embedding transcript text, but output artifacts contain transcript content.
|
||||||
|
- Rendered Markdown is human-readable transcript content and should be handled as sensitive output when applicable.
|
||||||
|
|
||||||
## Related docs
|
## Related docs
|
||||||
|
|
||||||
|
|||||||
@@ -14,7 +14,7 @@ must describe current behavior only; planned or speculative work belongs under
|
|||||||
## Project Shape
|
## Project Shape
|
||||||
|
|
||||||
seriatim is a Go CLI for transcript artifact processing. The implemented
|
seriatim is a Go CLI for transcript artifact processing. The implemented
|
||||||
commands are `merge`, `trim`, and `normalize`.
|
commands are `merge`, `trim`, `normalize`, and `render`.
|
||||||
|
|
||||||
`merge` reads one or more JSON transcript files, optionally maps input files to
|
`merge` reads one or more JSON transcript files, optionally maps input files to
|
||||||
canonical speakers, runs a registry-selected preprocessing chain, merges
|
canonical speakers, runs a registry-selected preprocessing chain, merges
|
||||||
@@ -22,11 +22,12 @@ canonical segments into deterministic chronological order, runs a
|
|||||||
registry-selected postprocessing chain, validates the selected output schema,
|
registry-selected postprocessing chain, validates the selected output schema,
|
||||||
and writes JSON output plus an optional JSON report.
|
and writes JSON output plus an optional JSON report.
|
||||||
|
|
||||||
`trim` and `normalize` are artifact-level commands outside the merge pipeline.
|
`trim`, `normalize`, and `render` are artifact-level commands outside the merge
|
||||||
`trim` reads an existing seriatim output artifact and projects it by segment ID.
|
pipeline. `trim` reads an existing seriatim output artifact and projects it by
|
||||||
`normalize` reads transcript-like JSON and emits one of seriatim's supported
|
segment ID. `normalize` reads transcript-like JSON and emits one of seriatim's
|
||||||
output schemas. Neither command runs merge preprocessing or postprocessing
|
supported output schemas. `render` reads an existing seriatim output artifact
|
||||||
modules.
|
and emits human-readable Markdown. None of these commands runs merge
|
||||||
|
preprocessing or postprocessing modules.
|
||||||
|
|
||||||
The supported public output schemas are `seriatim-minimal`,
|
The supported public output schemas are `seriatim-minimal`,
|
||||||
`seriatim-intermediate`, and `seriatim-full`. For command and flag details, use
|
`seriatim-intermediate`, and `seriatim-full`. For command and flag details, use
|
||||||
@@ -66,9 +67,9 @@ collects report events, converts the final transcript, and writes optional
|
|||||||
reports. Built-in adapters and modules are registered from `internal/builtin`.
|
reports. Built-in adapters and modules are registered from `internal/builtin`.
|
||||||
|
|
||||||
CLI code in `internal/cli` should parse flags, build validated config values,
|
CLI code in `internal/cli` should parse flags, build validated config values,
|
||||||
and delegate. `merge` delegates to `pipeline.Run`; `trim` and `normalize`
|
and delegate. `merge` delegates to `pipeline.Run`; `trim`, `normalize`, and
|
||||||
perform artifact-level orchestration and delegate deterministic parsing,
|
`render` perform artifact-level orchestration and delegate deterministic
|
||||||
validation, and transformation work to their internal packages.
|
parsing, validation, and transformation work to their internal packages.
|
||||||
|
|
||||||
Config loading and validation belongs in `internal/config`. Filesystem reads and
|
Config loading and validation belongs in `internal/config`. Filesystem reads and
|
||||||
writes are adapter concerns and should not spread into pure transformation
|
writes are adapter concerns and should not spread into pure transformation
|
||||||
@@ -166,10 +167,10 @@ correction or annotation modules, inspect the package tests for overlap,
|
|||||||
coalesce, danglers, backchannel, filler, and autocorrect behavior.
|
coalesce, danglers, backchannel, filler, and autocorrect behavior.
|
||||||
|
|
||||||
When changing artifact-level commands, inspect `internal/trim`,
|
When changing artifact-level commands, inspect `internal/trim`,
|
||||||
`internal/normalize`, and their CLI tests. When changing public output shape or
|
`internal/normalize`, `internal/render`, and their CLI tests. When changing
|
||||||
schema validation, inspect `schema` and `internal/artifact` tests. Report and
|
public output shape or schema validation, inspect `schema` and
|
||||||
diagnostic changes should be covered through the command or package tests that
|
`internal/artifact` tests. Report and diagnostic changes should be covered
|
||||||
emit the affected events.
|
through the command or package tests that emit the affected events.
|
||||||
|
|
||||||
## Dependency Policy
|
## Dependency Policy
|
||||||
|
|
||||||
@@ -202,8 +203,8 @@ free of secrets or private transcript data.
|
|||||||
registry name.
|
registry name.
|
||||||
- Preserve deterministic ordering, final segment ID assignment, and schema
|
- Preserve deterministic ordering, final segment ID assignment, and schema
|
||||||
validation before output acceptance.
|
validation before output acceptance.
|
||||||
- Keep `trim` and `normalize` artifact-level; do not run merge modules from
|
- Keep `trim`, `normalize`, and `render` artifact-level; do not run merge
|
||||||
those commands.
|
modules from those commands.
|
||||||
- Keep public output schemas validated through `schema`.
|
- Keep public output schemas validated through `schema`.
|
||||||
- Keep optional reports ordered, concise, and diagnostic.
|
- Keep optional reports ordered, concise, and diagnostic.
|
||||||
- Avoid broad dependencies without a concrete maintainability benefit.
|
- Avoid broad dependencies without a concrete maintainability benefit.
|
||||||
|
|||||||
@@ -16,6 +16,7 @@ It complements [architecture policy](architecture.md) and
|
|||||||
- `internal/artifact/`: conversion from internal merged model to public shapes.
|
- `internal/artifact/`: conversion from internal merged model to public shapes.
|
||||||
- `internal/trim/`: artifact-level trim logic.
|
- `internal/trim/`: artifact-level trim logic.
|
||||||
- `internal/normalize/`: artifact-level normalize parsing/building.
|
- `internal/normalize/`: artifact-level normalize parsing/building.
|
||||||
|
- `internal/render/`: artifact-level rendering and renderer registry.
|
||||||
- `internal/*` domain packages: overlap, coalesce, danglers, filler,
|
- `internal/*` domain packages: overlap, coalesce, danglers, filler,
|
||||||
backchannel, speaker, autocorrect, report, model.
|
backchannel, speaker, autocorrect, report, model.
|
||||||
- `schema/`: public structs plus embedded JSON Schemas and validation.
|
- `schema/`: public structs plus embedded JSON Schemas and validation.
|
||||||
@@ -36,6 +37,7 @@ go run ./cmd/seriatim --help
|
|||||||
go run ./cmd/seriatim merge --help
|
go run ./cmd/seriatim merge --help
|
||||||
go run ./cmd/seriatim trim --help
|
go run ./cmd/seriatim trim --help
|
||||||
go run ./cmd/seriatim normalize --help
|
go run ./cmd/seriatim normalize --help
|
||||||
|
go run ./cmd/seriatim render --help
|
||||||
```
|
```
|
||||||
|
|
||||||
Current toolchain note:
|
Current toolchain note:
|
||||||
|
|||||||
@@ -4,30 +4,31 @@ Each entry includes symptom, likely cause, inspection step, and safe fix.
|
|||||||
|
|
||||||
## Missing required flags
|
## Missing required flags
|
||||||
|
|
||||||
- Symptom: command fails with messages like `--input-file is required`, `--output-file is required`, or `exactly one of --keep or --remove is required`.
|
- Symptom: command fails with messages like `--input-file is required`, `--output-file is required`, `--format is required`, or `exactly one of --keep or --remove is required`.
|
||||||
- Likely cause: required command flags were omitted.
|
- Likely cause: one or more required flags were omitted.
|
||||||
- Inspection: run command help for the failing command:
|
- Inspection: run help for the failing command:
|
||||||
- `go run ./cmd/seriatim merge --help`
|
- `go run ./cmd/seriatim merge --help`
|
||||||
- `go run ./cmd/seriatim trim --help`
|
- `go run ./cmd/seriatim trim --help`
|
||||||
- `go run ./cmd/seriatim normalize --help`
|
- `go run ./cmd/seriatim normalize --help`
|
||||||
|
- `go run ./cmd/seriatim render --help`
|
||||||
- Safe fix: provide all required flags; for `trim`, provide exactly one selector mode (`--keep` or `--remove`).
|
- Safe fix: provide all required flags; for `trim`, provide exactly one selector mode (`--keep` or `--remove`).
|
||||||
|
|
||||||
## Invalid output or report path
|
## Invalid output or report path
|
||||||
|
|
||||||
- Symptom: errors like `--output-file parent directory ...` or `--report-file parent directory ...`.
|
- Symptom: errors like `--output-file parent directory ...` or `--report-file parent directory ...`.
|
||||||
- Likely cause: parent directory does not exist, is not a directory, or path points to an unusable target.
|
- Likely cause: parent directory does not exist, is not a directory, or the target path is unusable.
|
||||||
- Inspection: verify paths:
|
- Inspection: verify parent path and permissions:
|
||||||
- `dirname <path>`
|
- `dirname <path>`
|
||||||
- `ls -ld <parent-dir>`
|
- `ls -ld <parent-dir>`
|
||||||
- Safe fix: create/fix the parent directory and rerun; avoid using directory paths directly as output/report file targets.
|
- Safe fix: create or fix the parent directory and rerun. Use a file path (not a directory path) for output/report targets.
|
||||||
|
|
||||||
## Invalid merge input JSON
|
## Invalid merge input JSON
|
||||||
|
|
||||||
- Symptom: merge fails with messages like `parse input file`, `must contain top-level segments array`, `segment 0 missing numeric start`, or `segment 0 words must be an array`.
|
- Symptom: merge fails with messages like `parse input file`, `must contain top-level segments array`, `segment 0 missing numeric start`, or `segment 0 words must be an array`.
|
||||||
- Likely cause: malformed JSON or unsupported/missing fields in a merge input file.
|
- Likely cause: malformed JSON or unsupported/missing fields in a merge input file.
|
||||||
- Inspection: validate input JSON and required fields (`start`, `end`, `text`):
|
- Inspection: validate JSON and required segment fields (`start`, `end`, `text`):
|
||||||
- `jq . <input-file>`
|
- `jq . <input-file>`
|
||||||
- Safe fix: correct the JSON structure and segment/word field types, then rerun `merge`.
|
- Safe fix: correct JSON structure and segment/word field types, then rerun `merge`.
|
||||||
|
|
||||||
## Invalid normalize input shape
|
## Invalid normalize input shape
|
||||||
|
|
||||||
@@ -38,13 +39,22 @@ Each entry includes symptom, likely cause, inspection step, and safe fix.
|
|||||||
- `jq 'keys' <input-file>` (for object input)
|
- `jq 'keys' <input-file>` (for object input)
|
||||||
- Safe fix: reshape input into one supported form and rerun `normalize`.
|
- Safe fix: reshape input into one supported form and rerun `normalize`.
|
||||||
|
|
||||||
|
## Invalid render input artifact
|
||||||
|
|
||||||
|
- Symptom: render fails with messages like `input JSON is malformed` or `input JSON is not a valid seriatim output artifact`.
|
||||||
|
- Likely cause: input is malformed JSON or not one of the supported seriatim output schemas.
|
||||||
|
- Inspection:
|
||||||
|
- `jq . <input-file>`
|
||||||
|
- compare input shape against `schema/minimal-output.schema.json`, `schema/intermediate-output.schema.json`, and `schema/full-output.schema.json`
|
||||||
|
- Safe fix: render only a valid existing seriatim artifact (`seriatim-minimal`, `seriatim-intermediate`, or `seriatim-full`).
|
||||||
|
|
||||||
## Invalid speaker map or autocorrect YAML
|
## Invalid speaker map or autocorrect YAML
|
||||||
|
|
||||||
- Symptom: merge fails with errors such as `must contain at least one match rule`, `must include speaker`, `must include target`, or duplicate match/speaker validation failures.
|
- Symptom: merge fails with errors such as `must contain at least one match rule`, `must include speaker`, `must include target`, or duplicate match/speaker validation failures.
|
||||||
- Likely cause: YAML rule file structure/content does not match expected schema.
|
- Likely cause: YAML rule file structure/content does not match expected contract.
|
||||||
- Inspection: check YAML validity and required top-level keys:
|
- Inspection: check YAML validity and required top-level keys:
|
||||||
- `speakers.yml` requires top-level `match` rules.
|
- `speakers.yml` requires top-level `match` rules
|
||||||
- `autocorrect.yml` requires top-level `autocorrect` rules.
|
- `autocorrect.yml` requires top-level `autocorrect` rules
|
||||||
- Safe fix: correct YAML structure and rule content, then rerun `merge`.
|
- Safe fix: correct YAML structure and rule content, then rerun `merge`.
|
||||||
|
|
||||||
## Unknown module names
|
## Unknown module names
|
||||||
@@ -54,14 +64,18 @@ Each entry includes symptom, likely cause, inspection step, and safe fix.
|
|||||||
- Inspection: compare provided module names against defaults in CLI help and config docs.
|
- Inspection: compare provided module names against defaults in CLI help and config docs.
|
||||||
- Safe fix: use implemented module names only or remove unsupported modules from comma-separated lists.
|
- Safe fix: use implemented module names only or remove unsupported modules from comma-separated lists.
|
||||||
|
|
||||||
## Invalid output schema value
|
## Invalid format or schema values
|
||||||
|
|
||||||
- Symptom: errors like `--output-schema must be one of ...`.
|
- Symptom:
|
||||||
- Likely cause: unsupported schema value from flag or `SERIATIM_OUTPUT_SCHEMA`.
|
- render: `--format must be "markdown"`
|
||||||
- Inspection: check effective value:
|
- merge/normalize/trim: `--output-schema must be one of ...`
|
||||||
|
- Likely cause: unsupported `--format` or `--output-schema` value.
|
||||||
|
- Inspection:
|
||||||
- command flags
|
- command flags
|
||||||
- `echo "$SERIATIM_OUTPUT_SCHEMA"`
|
- `echo "$SERIATIM_OUTPUT_SCHEMA"` (for merge/normalize defaults)
|
||||||
- Safe fix: use one of `seriatim-minimal`, `seriatim-intermediate`, or `seriatim-full`.
|
- Safe fix:
|
||||||
|
- render: use `--format markdown`
|
||||||
|
- output schema: use `seriatim-minimal`, `seriatim-intermediate`, or `seriatim-full`
|
||||||
|
|
||||||
## Invalid trim selector
|
## Invalid trim selector
|
||||||
|
|
||||||
@@ -73,23 +87,23 @@ Each entry includes symptom, likely cause, inspection step, and safe fix.
|
|||||||
- list: `1-10,15,20-25`
|
- list: `1-10,15,20-25`
|
||||||
- Safe fix: correct selector syntax and rerun `trim`.
|
- Safe fix: correct selector syntax and rerun `trim`.
|
||||||
|
|
||||||
## Schema validation failures
|
## Artifact or schema validation failures
|
||||||
|
|
||||||
- Symptom: errors such as `validate-output: ...` in merge or `input JSON is not a valid seriatim output artifact` in trim.
|
- Symptom: errors such as `validate-output: ...`, `input JSON is not a valid seriatim output artifact`, or related schema-validation errors.
|
||||||
- Likely cause:
|
- Likely cause:
|
||||||
- merge module order/config produced invalid final artifact (for example, validating before IDs are assigned), or
|
- merge module order/config produced an invalid output artifact, or
|
||||||
- trim input is not a valid seriatim artifact.
|
- trim/render input is not a valid seriatim output artifact.
|
||||||
- Inspection:
|
- Inspection:
|
||||||
- for merge: inspect customized module ordering flags.
|
- for merge: inspect customized module ordering flags
|
||||||
- for trim: verify input artifact against known seriatim schema files in `schema/`.
|
- for trim/render: validate input against schema files in `schema/`
|
||||||
- Safe fix:
|
- Safe fix:
|
||||||
- restore valid merge postprocessing order ending with assigned IDs before validation, or
|
- merge: restore a valid postprocessing order ending with assigned IDs before output validation
|
||||||
- provide a valid seriatim artifact as trim input.
|
- trim/render: provide a valid seriatim artifact as input
|
||||||
|
|
||||||
## Report write failure
|
## Report write failure
|
||||||
|
|
||||||
- Symptom: errors like `write --report-file ...` or file-create failures when report writing is requested.
|
- Symptom: errors like `write --report-file ...` or file-create failures when report writing is requested.
|
||||||
- Likely cause: report path is not writable or is an invalid target (for example a directory path).
|
- Likely cause: report path is not writable or points to an invalid target.
|
||||||
- Inspection:
|
- Inspection:
|
||||||
- `ls -ld <report-parent-dir>`
|
- `ls -ld <report-parent-dir>`
|
||||||
- verify `--report-file` is a file path, not a directory
|
- verify `--report-file` is a file path, not a directory
|
||||||
|
|||||||
@@ -1,8 +1,7 @@
|
|||||||
# Examples
|
# Examples
|
||||||
|
|
||||||
These are small synthetic, copyable example assets for the implemented CLI
|
These are small synthetic, copyable example assets for the implemented CLI
|
||||||
commands.
|
commands. This directory is the canonical examples home for documentation.
|
||||||
This directory is the canonical examples home for documentation.
|
|
||||||
|
|
||||||
## Merge example
|
## Merge example
|
||||||
|
|
||||||
@@ -55,6 +54,25 @@ go run ./cmd/seriatim trim \
|
|||||||
--keep "1-2"
|
--keep "1-2"
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Render example
|
||||||
|
|
||||||
|
Input artifact:
|
||||||
|
|
||||||
|
- `render/input-intermediate.json`
|
||||||
|
|
||||||
|
Expected Markdown output shape:
|
||||||
|
|
||||||
|
- `render/output-markdown.md`
|
||||||
|
|
||||||
|
Run:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
go run ./cmd/seriatim render \
|
||||||
|
--input-file examples/render/input-intermediate.json \
|
||||||
|
--output-file /tmp/seriatim-example-render.md \
|
||||||
|
--format markdown
|
||||||
|
```
|
||||||
|
|
||||||
## YAML rule examples
|
## YAML rule examples
|
||||||
|
|
||||||
- `speakers.yml`
|
- `speakers.yml`
|
||||||
|
|||||||
33
examples/render/input-intermediate.json
Normal file
33
examples/render/input-intermediate.json
Normal file
@@ -0,0 +1,33 @@
|
|||||||
|
{
|
||||||
|
"metadata": {
|
||||||
|
"application": "seriatim",
|
||||||
|
"version": "v-test",
|
||||||
|
"output_schema": "seriatim-intermediate"
|
||||||
|
},
|
||||||
|
"segments": [
|
||||||
|
{
|
||||||
|
"id": 1,
|
||||||
|
"start": 1,
|
||||||
|
"end": 4,
|
||||||
|
"speaker": "Eric Rakestraw",
|
||||||
|
"text": "Hello there."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 2,
|
||||||
|
"start": 5,
|
||||||
|
"end": 8,
|
||||||
|
"speaker": "Mike Brown",
|
||||||
|
"text": "Welcome back, everyone."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 3,
|
||||||
|
"start": 9,
|
||||||
|
"end": 10,
|
||||||
|
"speaker": "Eric Rakestraw",
|
||||||
|
"text": "Yeah.",
|
||||||
|
"categories": [
|
||||||
|
"backchannel"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
7
examples/render/output-markdown.md
Normal file
7
examples/render/output-markdown.md
Normal file
@@ -0,0 +1,7 @@
|
|||||||
|
# Transcript
|
||||||
|
|
||||||
|
[00:00:01–00:00:04] **Eric Rakestraw:** Hello there.
|
||||||
|
|
||||||
|
[00:00:05–00:00:08] **Mike Brown:** Welcome back, everyone.
|
||||||
|
|
||||||
|
[00:00:09–00:00:10] **Eric Rakestraw:** *Yeah.*
|
||||||
178
internal/artifact/output_artifact.go
Normal file
178
internal/artifact/output_artifact.go
Normal file
@@ -0,0 +1,178 @@
|
|||||||
|
package artifact
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/schema"
|
||||||
|
)
|
||||||
|
|
||||||
|
const (
|
||||||
|
OutputSchemaMinimal = schema.OutputSchemaMinimal
|
||||||
|
OutputSchemaIntermediate = schema.OutputSchemaIntermediate
|
||||||
|
OutputSchemaFull = schema.OutputSchemaFull
|
||||||
|
)
|
||||||
|
|
||||||
|
// OutputArtifact stores a parsed seriatim output artifact of one supported schema.
|
||||||
|
type OutputArtifact struct {
|
||||||
|
Schema string
|
||||||
|
Full *schema.Transcript
|
||||||
|
Intermediate *schema.IntermediateTranscript
|
||||||
|
Minimal *schema.MinimalTranscript
|
||||||
|
}
|
||||||
|
|
||||||
|
// ParseOutputArtifactJSON parses and validates serialized seriatim output JSON.
|
||||||
|
func ParseOutputArtifactJSON(data []byte) (OutputArtifact, error) {
|
||||||
|
var decoded any
|
||||||
|
if err := json.Unmarshal(data, &decoded); err != nil {
|
||||||
|
return OutputArtifact{}, fmt.Errorf("input JSON is malformed: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
var full schema.Transcript
|
||||||
|
if err := json.Unmarshal(data, &full); err == nil {
|
||||||
|
if err := schema.ValidateTranscript(full); err == nil {
|
||||||
|
return OutputArtifact{
|
||||||
|
Schema: OutputSchemaFull,
|
||||||
|
Full: &full,
|
||||||
|
}, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
var intermediate schema.IntermediateTranscript
|
||||||
|
if err := json.Unmarshal(data, &intermediate); err == nil {
|
||||||
|
if err := schema.ValidateIntermediateTranscript(intermediate); err == nil {
|
||||||
|
return OutputArtifact{
|
||||||
|
Schema: OutputSchemaIntermediate,
|
||||||
|
Intermediate: &intermediate,
|
||||||
|
}, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
var minimal schema.MinimalTranscript
|
||||||
|
if err := json.Unmarshal(data, &minimal); err == nil {
|
||||||
|
if err := schema.ValidateMinimalTranscript(minimal); err == nil {
|
||||||
|
return OutputArtifact{
|
||||||
|
Schema: OutputSchemaMinimal,
|
||||||
|
Minimal: &minimal,
|
||||||
|
}, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return OutputArtifact{}, fmt.Errorf("input JSON is not a valid seriatim output artifact")
|
||||||
|
}
|
||||||
|
|
||||||
|
// Value returns the output payload value for serialization.
|
||||||
|
func (artifact OutputArtifact) Value() any {
|
||||||
|
switch artifact.Schema {
|
||||||
|
case OutputSchemaFull:
|
||||||
|
if artifact.Full == nil {
|
||||||
|
return schema.Transcript{}
|
||||||
|
}
|
||||||
|
return *artifact.Full
|
||||||
|
case OutputSchemaIntermediate:
|
||||||
|
if artifact.Intermediate == nil {
|
||||||
|
return schema.IntermediateTranscript{}
|
||||||
|
}
|
||||||
|
return *artifact.Intermediate
|
||||||
|
case OutputSchemaMinimal:
|
||||||
|
if artifact.Minimal == nil {
|
||||||
|
return schema.MinimalTranscript{}
|
||||||
|
}
|
||||||
|
return *artifact.Minimal
|
||||||
|
default:
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// SegmentCount returns the number of segments in the output artifact.
|
||||||
|
func (artifact OutputArtifact) SegmentCount() int {
|
||||||
|
switch artifact.Schema {
|
||||||
|
case OutputSchemaFull:
|
||||||
|
if artifact.Full == nil {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return len(artifact.Full.Segments)
|
||||||
|
case OutputSchemaIntermediate:
|
||||||
|
if artifact.Intermediate == nil {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return len(artifact.Intermediate.Segments)
|
||||||
|
case OutputSchemaMinimal:
|
||||||
|
if artifact.Minimal == nil {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return len(artifact.Minimal.Segments)
|
||||||
|
default:
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Application returns output artifact metadata application name.
|
||||||
|
func (artifact OutputArtifact) Application() string {
|
||||||
|
switch artifact.Schema {
|
||||||
|
case OutputSchemaFull:
|
||||||
|
if artifact.Full == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return artifact.Full.Metadata.Application
|
||||||
|
case OutputSchemaIntermediate:
|
||||||
|
if artifact.Intermediate == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return artifact.Intermediate.Metadata.Application
|
||||||
|
case OutputSchemaMinimal:
|
||||||
|
if artifact.Minimal == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return artifact.Minimal.Metadata.Application
|
||||||
|
default:
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Version returns output artifact metadata version.
|
||||||
|
func (artifact OutputArtifact) Version() string {
|
||||||
|
switch artifact.Schema {
|
||||||
|
case OutputSchemaFull:
|
||||||
|
if artifact.Full == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return artifact.Full.Metadata.Version
|
||||||
|
case OutputSchemaIntermediate:
|
||||||
|
if artifact.Intermediate == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return artifact.Intermediate.Metadata.Version
|
||||||
|
case OutputSchemaMinimal:
|
||||||
|
if artifact.Minimal == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return artifact.Minimal.Metadata.Version
|
||||||
|
default:
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// FullPayload returns the full-schema payload when present.
|
||||||
|
func (artifact OutputArtifact) FullPayload() (*schema.Transcript, error) {
|
||||||
|
if artifact.Full == nil {
|
||||||
|
return nil, fmt.Errorf("full artifact payload is missing")
|
||||||
|
}
|
||||||
|
return artifact.Full, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// IntermediatePayload returns the intermediate-schema payload when present.
|
||||||
|
func (artifact OutputArtifact) IntermediatePayload() (*schema.IntermediateTranscript, error) {
|
||||||
|
if artifact.Intermediate == nil {
|
||||||
|
return nil, fmt.Errorf("intermediate artifact payload is missing")
|
||||||
|
}
|
||||||
|
return artifact.Intermediate, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// MinimalPayload returns the minimal-schema payload when present.
|
||||||
|
func (artifact OutputArtifact) MinimalPayload() (*schema.MinimalTranscript, error) {
|
||||||
|
if artifact.Minimal == nil {
|
||||||
|
return nil, fmt.Errorf("minimal artifact payload is missing")
|
||||||
|
}
|
||||||
|
return artifact.Minimal, nil
|
||||||
|
}
|
||||||
134
internal/artifact/output_artifact_test.go
Normal file
134
internal/artifact/output_artifact_test.go
Normal file
@@ -0,0 +1,134 @@
|
|||||||
|
package artifact
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/schema"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestParseOutputArtifactJSONParsesFullIntermediateAndMinimal(t *testing.T) {
|
||||||
|
t.Run("full", func(t *testing.T) {
|
||||||
|
first := 0
|
||||||
|
value := schema.Transcript{
|
||||||
|
Metadata: schema.Metadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
InputReader: "json-files",
|
||||||
|
InputFiles: []string{"input.json"},
|
||||||
|
PreprocessingModules: []string{"validate-raw"},
|
||||||
|
PostprocessingModules: []string{"assign-ids", "validate-output"},
|
||||||
|
OutputModules: []string{"json"},
|
||||||
|
},
|
||||||
|
Segments: []schema.Segment{
|
||||||
|
{
|
||||||
|
ID: 1,
|
||||||
|
Source: "input.json",
|
||||||
|
SourceSegmentIndex: &first,
|
||||||
|
Speaker: "Alice",
|
||||||
|
Start: 1,
|
||||||
|
End: 2,
|
||||||
|
Text: "hello",
|
||||||
|
Categories: []string{"backchannel"},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
OverlapGroups: []schema.OverlapGroup{},
|
||||||
|
}
|
||||||
|
|
||||||
|
parsed := mustParseOutputArtifact(t, value)
|
||||||
|
if parsed.Schema != OutputSchemaFull {
|
||||||
|
t.Fatalf("schema = %q, want %q", parsed.Schema, OutputSchemaFull)
|
||||||
|
}
|
||||||
|
if parsed.Full == nil {
|
||||||
|
t.Fatal("expected full payload")
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
t.Run("intermediate", func(t *testing.T) {
|
||||||
|
value := schema.IntermediateTranscript{
|
||||||
|
Metadata: schema.IntermediateMetadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
OutputSchema: OutputSchemaIntermediate,
|
||||||
|
},
|
||||||
|
Segments: []schema.IntermediateSegment{
|
||||||
|
{ID: 1, Start: 1, End: 2, Speaker: "Alice", Text: "hello", Categories: []string{"filler"}},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
parsed := mustParseOutputArtifact(t, value)
|
||||||
|
if parsed.Schema != OutputSchemaIntermediate {
|
||||||
|
t.Fatalf("schema = %q, want %q", parsed.Schema, OutputSchemaIntermediate)
|
||||||
|
}
|
||||||
|
if parsed.Intermediate == nil {
|
||||||
|
t.Fatal("expected intermediate payload")
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
t.Run("minimal", func(t *testing.T) {
|
||||||
|
value := schema.MinimalTranscript{
|
||||||
|
Metadata: schema.MinimalMetadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
OutputSchema: OutputSchemaMinimal,
|
||||||
|
},
|
||||||
|
Segments: []schema.MinimalSegment{
|
||||||
|
{ID: 1, Start: 1, End: 2, Speaker: "Alice", Text: "hello"},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
parsed := mustParseOutputArtifact(t, value)
|
||||||
|
if parsed.Schema != OutputSchemaMinimal {
|
||||||
|
t.Fatalf("schema = %q, want %q", parsed.Schema, OutputSchemaMinimal)
|
||||||
|
}
|
||||||
|
if parsed.Minimal == nil {
|
||||||
|
t.Fatal("expected minimal payload")
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseOutputArtifactJSONRejectsMalformedJSON(t *testing.T) {
|
||||||
|
_, err := ParseOutputArtifactJSON([]byte(`{"metadata":`))
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected malformed JSON error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "input JSON is malformed") {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseOutputArtifactJSONRejectsRawWhisperXLikeInput(t *testing.T) {
|
||||||
|
data := []byte(`{
|
||||||
|
"segments": [
|
||||||
|
{
|
||||||
|
"id": 0,
|
||||||
|
"start": 0.1,
|
||||||
|
"end": 1.2,
|
||||||
|
"text": "hello",
|
||||||
|
"words": [{"word":"hello","start":0.1,"end":0.8}]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}`)
|
||||||
|
|
||||||
|
_, err := ParseOutputArtifactJSON(data)
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected artifact validation error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "not a valid seriatim output artifact") {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func mustParseOutputArtifact(t *testing.T, value any) OutputArtifact {
|
||||||
|
t.Helper()
|
||||||
|
data, err := json.Marshal(value)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("marshal: %v", err)
|
||||||
|
}
|
||||||
|
parsed, err := ParseOutputArtifactJSON(data)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("parse: %v", err)
|
||||||
|
}
|
||||||
|
return parsed
|
||||||
|
}
|
||||||
39
internal/cli/render.go
Normal file
39
internal/cli/render.go
Normal file
@@ -0,0 +1,39 @@
|
|||||||
|
package cli
|
||||||
|
|
||||||
|
import (
|
||||||
|
"github.com/spf13/cobra"
|
||||||
|
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/internal/config"
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/internal/render"
|
||||||
|
)
|
||||||
|
|
||||||
|
func newRenderCommand() *cobra.Command {
|
||||||
|
opts := config.RenderOptions{
|
||||||
|
Title: config.DefaultRenderTitle,
|
||||||
|
IncludeTimestamps: true,
|
||||||
|
}
|
||||||
|
|
||||||
|
cmd := &cobra.Command{
|
||||||
|
Use: "render",
|
||||||
|
Short: "Render a seriatim transcript artifact into human-readable output",
|
||||||
|
RunE: func(cmd *cobra.Command, args []string) error {
|
||||||
|
cfg, err := config.NewRenderConfig(opts)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
return render.Run(cmd.Context(), cfg)
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
flags := cmd.Flags()
|
||||||
|
flags.StringVar(&opts.InputFile, "input-file", "", "input seriatim transcript artifact JSON file")
|
||||||
|
flags.StringVar(&opts.OutputFile, "output-file", "", "rendered output file path")
|
||||||
|
flags.StringVar(&opts.Format, "format", "", "output format (markdown)")
|
||||||
|
flags.StringVar(&opts.Title, "title", config.DefaultRenderTitle, "document title")
|
||||||
|
flags.BoolVar(&opts.IncludeTimestamps, "include-timestamps", true, "include segment timestamps")
|
||||||
|
flags.BoolVar(&opts.IncludeSegmentIDs, "include-segment-ids", false, "include segment IDs")
|
||||||
|
flags.BoolVar(&opts.IncludeMetadata, "include-metadata", false, "include artifact metadata")
|
||||||
|
|
||||||
|
return cmd
|
||||||
|
}
|
||||||
274
internal/cli/render_test.go
Normal file
274
internal/cli/render_test.go
Normal file
@@ -0,0 +1,274 @@
|
|||||||
|
package cli
|
||||||
|
|
||||||
|
import (
|
||||||
|
"bytes"
|
||||||
|
"os"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/internal/config"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestRenderCommandIsRecognized(t *testing.T) {
|
||||||
|
cmd := NewRootCommand()
|
||||||
|
cmd.SetArgs([]string{"render", "--help"})
|
||||||
|
if err := cmd.Execute(); err != nil {
|
||||||
|
t.Fatalf("render command should be recognized: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRootHelpIncludesRender(t *testing.T) {
|
||||||
|
cmd := NewRootCommand()
|
||||||
|
var out bytes.Buffer
|
||||||
|
cmd.SetOut(&out)
|
||||||
|
cmd.SetErr(&out)
|
||||||
|
cmd.SetArgs([]string{"--help"})
|
||||||
|
if err := cmd.Execute(); err != nil {
|
||||||
|
t.Fatalf("help failed: %v", err)
|
||||||
|
}
|
||||||
|
if !strings.Contains(out.String(), "render") {
|
||||||
|
t.Fatalf("root help missing render command:\n%s", out.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRenderEndToEndMarkdownOutput(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeJSONFile(t, dir, "input.json", `{
|
||||||
|
"metadata": {
|
||||||
|
"application": "seriatim",
|
||||||
|
"version": "v-test",
|
||||||
|
"output_schema": "seriatim-intermediate"
|
||||||
|
},
|
||||||
|
"segments": [
|
||||||
|
{"id": 1, "start": 1, "end": 4, "speaker": "Eric", "text": "Hello there."},
|
||||||
|
{"id": 2, "start": 5, "end": 8, "speaker": "Mike", "text": "Yeah.", "categories": ["backchannel"]}
|
||||||
|
]
|
||||||
|
}`)
|
||||||
|
output := writeJSONFile(t, dir, "output.md", "")
|
||||||
|
|
||||||
|
err := executeRender(
|
||||||
|
"--input-file", input,
|
||||||
|
"--output-file", output,
|
||||||
|
"--format", config.RenderFormatMarkdown,
|
||||||
|
"--title", "Transcript",
|
||||||
|
)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render failed: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
data := readFile(t, output)
|
||||||
|
if !strings.Contains(data, "# Transcript") {
|
||||||
|
t.Fatalf("missing title:\n%s", data)
|
||||||
|
}
|
||||||
|
if !strings.Contains(data, "[00:00:01–00:00:04] **Eric:** Hello there.") {
|
||||||
|
t.Fatalf("missing first segment:\n%s", data)
|
||||||
|
}
|
||||||
|
if !strings.Contains(data, "[00:00:05–00:00:08] **Mike:** *Yeah.*") {
|
||||||
|
t.Fatalf("missing italicized backchannel segment:\n%s", data)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRenderWorksWithRequiredFlagsOnly(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeJSONFile(t, dir, "input.json", `{
|
||||||
|
"metadata": {
|
||||||
|
"application": "seriatim",
|
||||||
|
"version": "v-test",
|
||||||
|
"output_schema": "seriatim-minimal"
|
||||||
|
},
|
||||||
|
"segments": [
|
||||||
|
{"id": 1, "start": 1, "end": 2, "speaker": "Eric", "text": "Hello there."}
|
||||||
|
]
|
||||||
|
}`)
|
||||||
|
output := writeJSONFile(t, dir, "output.md", "")
|
||||||
|
|
||||||
|
err := executeRender(
|
||||||
|
"--input-file", input,
|
||||||
|
"--output-file", output,
|
||||||
|
"--format", config.RenderFormatMarkdown,
|
||||||
|
)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render with required flags failed: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
data := readFile(t, output)
|
||||||
|
if !strings.Contains(data, "# Transcript") {
|
||||||
|
t.Fatalf("missing default title:\n%s", data)
|
||||||
|
}
|
||||||
|
if !strings.Contains(data, "[00:00:01–00:00:02] **Eric:** Hello there.") {
|
||||||
|
t.Fatalf("missing rendered segment:\n%s", data)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRenderRejectsUnsupportedFormat(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeJSONFile(t, dir, "input.json", `{"metadata":{"application":"seriatim","version":"v-test","output_schema":"seriatim-minimal"},"segments":[]}`)
|
||||||
|
output := writeJSONFile(t, dir, "output.md", "")
|
||||||
|
|
||||||
|
err := executeRender(
|
||||||
|
"--input-file", input,
|
||||||
|
"--output-file", output,
|
||||||
|
"--format", "txt",
|
||||||
|
)
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected format error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "--format must be") {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRenderRejectsMalformedAndRawInput(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
output := writeJSONFile(t, dir, "output.md", "")
|
||||||
|
|
||||||
|
malformed := writeJSONFile(t, dir, "malformed.json", `{"metadata":`)
|
||||||
|
err := executeRender(
|
||||||
|
"--input-file", malformed,
|
||||||
|
"--output-file", output,
|
||||||
|
"--format", config.RenderFormatMarkdown,
|
||||||
|
)
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected malformed input error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "input JSON is malformed") {
|
||||||
|
t.Fatalf("unexpected malformed input error: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
raw := writeJSONFile(t, dir, "raw.json", `{"segments":[{"id":0,"start":0.1,"end":1.1,"text":"hello","words":[{"word":"hello"}]}]}`)
|
||||||
|
err = executeRender(
|
||||||
|
"--input-file", raw,
|
||||||
|
"--output-file", output,
|
||||||
|
"--format", config.RenderFormatMarkdown,
|
||||||
|
)
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected artifact validation error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "not a valid seriatim output artifact") {
|
||||||
|
t.Fatalf("unexpected raw input error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRenderSupportsMinimalIntermediateAndFullInputs(t *testing.T) {
|
||||||
|
tests := []struct {
|
||||||
|
name string
|
||||||
|
content string
|
||||||
|
}{
|
||||||
|
{
|
||||||
|
name: "minimal",
|
||||||
|
content: `{
|
||||||
|
"metadata": {
|
||||||
|
"application": "seriatim",
|
||||||
|
"version": "v-test",
|
||||||
|
"output_schema": "seriatim-minimal"
|
||||||
|
},
|
||||||
|
"segments": [{"id":1,"start":1,"end":2,"speaker":"A","text":"one"}]
|
||||||
|
}`,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
name: "intermediate",
|
||||||
|
content: `{
|
||||||
|
"metadata": {
|
||||||
|
"application": "seriatim",
|
||||||
|
"version": "v-test",
|
||||||
|
"output_schema": "seriatim-intermediate"
|
||||||
|
},
|
||||||
|
"segments": [{"id":1,"start":1,"end":2,"speaker":"A","text":"one","categories":["filler"]}]
|
||||||
|
}`,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
name: "full",
|
||||||
|
content: `{
|
||||||
|
"metadata": {
|
||||||
|
"application": "seriatim",
|
||||||
|
"version": "v-test",
|
||||||
|
"input_reader": "json-files",
|
||||||
|
"input_files": ["input.json"],
|
||||||
|
"preprocessing_modules": [],
|
||||||
|
"postprocessing_modules": [],
|
||||||
|
"output_modules": ["json"]
|
||||||
|
},
|
||||||
|
"segments": [{
|
||||||
|
"id":1,
|
||||||
|
"source":"input.json",
|
||||||
|
"source_segment_index":0,
|
||||||
|
"speaker":"A",
|
||||||
|
"start":1,
|
||||||
|
"end":2,
|
||||||
|
"text":"one"
|
||||||
|
}],
|
||||||
|
"overlap_groups": []
|
||||||
|
}`,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
for _, test := range tests {
|
||||||
|
t.Run(test.name, func(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeJSONFile(t, dir, "input.json", test.content)
|
||||||
|
output := writeJSONFile(t, dir, "output.md", "")
|
||||||
|
|
||||||
|
err := executeRender(
|
||||||
|
"--input-file", input,
|
||||||
|
"--output-file", output,
|
||||||
|
"--format", config.RenderFormatMarkdown,
|
||||||
|
)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render failed: %v", err)
|
||||||
|
}
|
||||||
|
data := readFile(t, output)
|
||||||
|
if !strings.Contains(data, "**A:**") {
|
||||||
|
t.Fatalf("missing rendered segment for %s input:\n%s", test.name, data)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRenderEmptyTranscriptIsDeterministic(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeJSONFile(t, dir, "input.json", `{
|
||||||
|
"metadata": {
|
||||||
|
"application": "seriatim",
|
||||||
|
"version": "v-test",
|
||||||
|
"output_schema": "seriatim-minimal"
|
||||||
|
},
|
||||||
|
"segments": []
|
||||||
|
}`)
|
||||||
|
output := writeJSONFile(t, dir, "output.md", "")
|
||||||
|
|
||||||
|
run := func() string {
|
||||||
|
err := executeRender(
|
||||||
|
"--input-file", input,
|
||||||
|
"--output-file", output,
|
||||||
|
"--format", config.RenderFormatMarkdown,
|
||||||
|
)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render failed: %v", err)
|
||||||
|
}
|
||||||
|
return readFile(t, output)
|
||||||
|
}
|
||||||
|
|
||||||
|
first := run()
|
||||||
|
second := run()
|
||||||
|
if first != second {
|
||||||
|
t.Fatalf("empty transcript render is not deterministic:\nfirst:\n%s\nsecond:\n%s", first, second)
|
||||||
|
}
|
||||||
|
if first != "# Transcript\n" {
|
||||||
|
t.Fatalf("unexpected empty transcript output:\n%s", first)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func executeRender(args ...string) error {
|
||||||
|
cmd := NewRootCommand()
|
||||||
|
cmd.SetArgs(append([]string{"render"}, args...))
|
||||||
|
return cmd.Execute()
|
||||||
|
}
|
||||||
|
|
||||||
|
func readFile(t *testing.T, path string) string {
|
||||||
|
t.Helper()
|
||||||
|
data, err := os.ReadFile(path)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("read %s: %v", path, err)
|
||||||
|
}
|
||||||
|
return string(data)
|
||||||
|
}
|
||||||
@@ -10,7 +10,7 @@ import (
|
|||||||
func NewRootCommand() *cobra.Command {
|
func NewRootCommand() *cobra.Command {
|
||||||
cmd := &cobra.Command{
|
cmd := &cobra.Command{
|
||||||
Use: "seriatim",
|
Use: "seriatim",
|
||||||
Short: "Merge, trim, and normalize transcript artifacts",
|
Short: "Merge, trim, normalize, and render transcript artifacts",
|
||||||
Version: buildinfo.Version,
|
Version: buildinfo.Version,
|
||||||
SilenceErrors: true,
|
SilenceErrors: true,
|
||||||
SilenceUsage: true,
|
SilenceUsage: true,
|
||||||
@@ -18,6 +18,7 @@ func NewRootCommand() *cobra.Command {
|
|||||||
|
|
||||||
cmd.AddCommand(newMergeCommand())
|
cmd.AddCommand(newMergeCommand())
|
||||||
cmd.AddCommand(newNormalizeCommand())
|
cmd.AddCommand(newNormalizeCommand())
|
||||||
|
cmd.AddCommand(newRenderCommand())
|
||||||
cmd.AddCommand(newTrimCommand())
|
cmd.AddCommand(newTrimCommand())
|
||||||
return cmd
|
return cmd
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -16,6 +16,8 @@ const (
|
|||||||
DefaultInputReader = "json-files"
|
DefaultInputReader = "json-files"
|
||||||
DefaultOutputModules = "json"
|
DefaultOutputModules = "json"
|
||||||
DefaultOutputSchema = OutputSchemaIntermediate
|
DefaultOutputSchema = OutputSchemaIntermediate
|
||||||
|
DefaultRenderTitle = "Transcript"
|
||||||
|
RenderFormatMarkdown = "markdown"
|
||||||
DefaultPreprocessingModules = "validate-raw,normalize-speakers,trim-text"
|
DefaultPreprocessingModules = "validate-raw,normalize-speakers,trim-text"
|
||||||
DefaultPostprocessingModules = "detect-overlaps,resolve-overlaps,backchannel,filler,resolve-danglers,coalesce,detect-overlaps,autocorrect,assign-ids,validate-output"
|
DefaultPostprocessingModules = "detect-overlaps,resolve-overlaps,backchannel,filler,resolve-danglers,coalesce,detect-overlaps,autocorrect,assign-ids,validate-output"
|
||||||
DefaultOverlapWordRunGap = 1.0
|
DefaultOverlapWordRunGap = 1.0
|
||||||
@@ -69,6 +71,17 @@ type NormalizeOptions struct {
|
|||||||
OutputModules string
|
OutputModules string
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// RenderOptions captures raw CLI option values before validation.
|
||||||
|
type RenderOptions struct {
|
||||||
|
InputFile string
|
||||||
|
OutputFile string
|
||||||
|
Format string
|
||||||
|
Title string
|
||||||
|
IncludeTimestamps bool
|
||||||
|
IncludeSegmentIDs bool
|
||||||
|
IncludeMetadata bool
|
||||||
|
}
|
||||||
|
|
||||||
// Config is the validated runtime configuration for a merge invocation.
|
// Config is the validated runtime configuration for a merge invocation.
|
||||||
type Config struct {
|
type Config struct {
|
||||||
InputFiles []string
|
InputFiles []string
|
||||||
@@ -108,6 +121,17 @@ type NormalizeConfig struct {
|
|||||||
OutputModules []string
|
OutputModules []string
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// RenderConfig is the validated runtime configuration for a render invocation.
|
||||||
|
type RenderConfig struct {
|
||||||
|
InputFile string
|
||||||
|
OutputFile string
|
||||||
|
Format string
|
||||||
|
Title string
|
||||||
|
IncludeTimestamps bool
|
||||||
|
IncludeSegmentIDs bool
|
||||||
|
IncludeMetadata bool
|
||||||
|
}
|
||||||
|
|
||||||
// NewMergeConfig validates raw merge options and returns normalized config.
|
// NewMergeConfig validates raw merge options and returns normalized config.
|
||||||
func NewMergeConfig(opts MergeOptions) (Config, error) {
|
func NewMergeConfig(opts MergeOptions) (Config, error) {
|
||||||
cfg := Config{
|
cfg := Config{
|
||||||
@@ -303,6 +327,42 @@ func NewNormalizeConfig(opts NormalizeOptions) (NormalizeConfig, error) {
|
|||||||
}, nil
|
}, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// NewRenderConfig validates raw render options and returns normalized config.
|
||||||
|
func NewRenderConfig(opts RenderOptions) (RenderConfig, error) {
|
||||||
|
inputFile, err := normalizeSingleInputFile(opts.InputFile, "--input-file")
|
||||||
|
if err != nil {
|
||||||
|
return RenderConfig{}, err
|
||||||
|
}
|
||||||
|
|
||||||
|
outputFile, err := normalizeOutputPath(opts.OutputFile, "--output-file")
|
||||||
|
if err != nil {
|
||||||
|
return RenderConfig{}, err
|
||||||
|
}
|
||||||
|
|
||||||
|
format := strings.TrimSpace(opts.Format)
|
||||||
|
if format == "" {
|
||||||
|
return RenderConfig{}, errors.New("--format is required")
|
||||||
|
}
|
||||||
|
if err := validateRenderFormat(format); err != nil {
|
||||||
|
return RenderConfig{}, err
|
||||||
|
}
|
||||||
|
|
||||||
|
title := strings.TrimSpace(opts.Title)
|
||||||
|
if title == "" {
|
||||||
|
title = DefaultRenderTitle
|
||||||
|
}
|
||||||
|
|
||||||
|
return RenderConfig{
|
||||||
|
InputFile: inputFile,
|
||||||
|
OutputFile: outputFile,
|
||||||
|
Format: format,
|
||||||
|
Title: title,
|
||||||
|
IncludeTimestamps: opts.IncludeTimestamps,
|
||||||
|
IncludeSegmentIDs: opts.IncludeSegmentIDs,
|
||||||
|
IncludeMetadata: opts.IncludeMetadata,
|
||||||
|
}, nil
|
||||||
|
}
|
||||||
|
|
||||||
func parseModuleList(value string) ([]string, error) {
|
func parseModuleList(value string) ([]string, error) {
|
||||||
value = strings.TrimSpace(value)
|
value = strings.TrimSpace(value)
|
||||||
if value == "" {
|
if value == "" {
|
||||||
@@ -485,3 +545,12 @@ func validateNormalizeOutputModules(modules []string) error {
|
|||||||
}
|
}
|
||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func validateRenderFormat(format string) error {
|
||||||
|
switch format {
|
||||||
|
case RenderFormatMarkdown:
|
||||||
|
return nil
|
||||||
|
default:
|
||||||
|
return fmt.Errorf("--format must be %q", RenderFormatMarkdown)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@@ -804,6 +804,137 @@ func TestNewNormalizeConfigTreatsWhitespaceReportFileAsOmitted(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestNewRenderConfigRequiresInputOutputAndFormat(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeTempFile(t, dir, "input.json")
|
||||||
|
output := filepath.Join(dir, "rendered.md")
|
||||||
|
|
||||||
|
_, err := NewRenderConfig(RenderOptions{
|
||||||
|
OutputFile: output,
|
||||||
|
Format: RenderFormatMarkdown,
|
||||||
|
})
|
||||||
|
if err == nil || !strings.Contains(err.Error(), "--input-file is required") {
|
||||||
|
t.Fatalf("expected input-file required error, got %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
_, err = NewRenderConfig(RenderOptions{
|
||||||
|
InputFile: input,
|
||||||
|
Format: RenderFormatMarkdown,
|
||||||
|
})
|
||||||
|
if err == nil || !strings.Contains(err.Error(), "--output-file is required") {
|
||||||
|
t.Fatalf("expected output-file required error, got %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
_, err = NewRenderConfig(RenderOptions{
|
||||||
|
InputFile: input,
|
||||||
|
OutputFile: output,
|
||||||
|
})
|
||||||
|
if err == nil || !strings.Contains(err.Error(), "--format is required") {
|
||||||
|
t.Fatalf("expected format required error, got %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNewRenderConfigRejectsUnknownFormat(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeTempFile(t, dir, "input.json")
|
||||||
|
output := filepath.Join(dir, "rendered.md")
|
||||||
|
|
||||||
|
opts := validRenderOptions(input, output)
|
||||||
|
opts.Format = "txt"
|
||||||
|
_, err := NewRenderConfig(opts)
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected format validation error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "--format must be") {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNewRenderConfigAppliesDefaultsAndFlags(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeTempFile(t, dir, "input.json")
|
||||||
|
output := filepath.Join(dir, "rendered.md")
|
||||||
|
|
||||||
|
cfg, err := NewRenderConfig(validRenderOptions(input, output))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("config failed: %v", err)
|
||||||
|
}
|
||||||
|
if cfg.Title != DefaultRenderTitle {
|
||||||
|
t.Fatalf("title = %q, want %q", cfg.Title, DefaultRenderTitle)
|
||||||
|
}
|
||||||
|
if !cfg.IncludeTimestamps {
|
||||||
|
t.Fatal("include timestamps should default true")
|
||||||
|
}
|
||||||
|
if cfg.IncludeSegmentIDs {
|
||||||
|
t.Fatal("include segment IDs should default false")
|
||||||
|
}
|
||||||
|
if cfg.IncludeMetadata {
|
||||||
|
t.Fatal("include metadata should default false")
|
||||||
|
}
|
||||||
|
|
||||||
|
opts := validRenderOptions(input, output)
|
||||||
|
opts.Title = "Meeting Notes"
|
||||||
|
opts.IncludeTimestamps = false
|
||||||
|
opts.IncludeSegmentIDs = true
|
||||||
|
opts.IncludeMetadata = true
|
||||||
|
cfg, err = NewRenderConfig(opts)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("config failed: %v", err)
|
||||||
|
}
|
||||||
|
if cfg.Title != "Meeting Notes" {
|
||||||
|
t.Fatalf("title = %q, want Meeting Notes", cfg.Title)
|
||||||
|
}
|
||||||
|
if cfg.IncludeTimestamps {
|
||||||
|
t.Fatal("include timestamps should be false")
|
||||||
|
}
|
||||||
|
if !cfg.IncludeSegmentIDs {
|
||||||
|
t.Fatal("include segment IDs should be true")
|
||||||
|
}
|
||||||
|
if !cfg.IncludeMetadata {
|
||||||
|
t.Fatal("include metadata should be true")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNewRenderConfigRejectsMissingAndDirectoryInputFile(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
output := filepath.Join(dir, "rendered.md")
|
||||||
|
|
||||||
|
missingInput := filepath.Join(dir, "missing.json")
|
||||||
|
_, err := NewRenderConfig(validRenderOptions(missingInput, output))
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected missing input-file error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "--input-file") {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
inputDir := filepath.Join(dir, "input-dir")
|
||||||
|
if err := os.MkdirAll(inputDir, 0o700); err != nil {
|
||||||
|
t.Fatalf("mkdir input dir: %v", err)
|
||||||
|
}
|
||||||
|
_, err = NewRenderConfig(validRenderOptions(inputDir, output))
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected directory input-file error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "is a directory, not a file") {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNewRenderConfigRejectsMissingOutputParent(t *testing.T) {
|
||||||
|
dir := t.TempDir()
|
||||||
|
input := writeTempFile(t, dir, "input.json")
|
||||||
|
output := filepath.Join(dir, "missing-parent", "rendered.md")
|
||||||
|
|
||||||
|
_, err := NewRenderConfig(validRenderOptions(input, output))
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected output parent directory error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "--output-file parent directory") {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func assertPositiveFloatEnvValidation(t *testing.T, envName string) {
|
func assertPositiveFloatEnvValidation(t *testing.T, envName string) {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
|
|
||||||
@@ -862,6 +993,18 @@ func validNormalizeOptions(inputFile string, outputFile string) NormalizeOptions
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func validRenderOptions(inputFile string, outputFile string) RenderOptions {
|
||||||
|
return RenderOptions{
|
||||||
|
InputFile: inputFile,
|
||||||
|
OutputFile: outputFile,
|
||||||
|
Format: RenderFormatMarkdown,
|
||||||
|
Title: DefaultRenderTitle,
|
||||||
|
IncludeTimestamps: true,
|
||||||
|
IncludeSegmentIDs: false,
|
||||||
|
IncludeMetadata: false,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func writeTempFile(t *testing.T, dir string, name string) string {
|
func writeTempFile(t *testing.T, dir string, name string) string {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
|
|
||||||
|
|||||||
97
internal/render/markdown.go
Normal file
97
internal/render/markdown.go
Normal file
@@ -0,0 +1,97 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"math"
|
||||||
|
"strings"
|
||||||
|
)
|
||||||
|
|
||||||
|
// MarkdownRenderer renders transcript artifacts as Markdown.
|
||||||
|
type MarkdownRenderer struct{}
|
||||||
|
|
||||||
|
// Render renders the transcript into deterministic Markdown.
|
||||||
|
func (MarkdownRenderer) Render(transcript Transcript, opts Options) (string, error) {
|
||||||
|
var lines []string
|
||||||
|
|
||||||
|
title := strings.TrimSpace(opts.Title)
|
||||||
|
if title == "" {
|
||||||
|
title = "Transcript"
|
||||||
|
}
|
||||||
|
lines = append(lines, "# "+escapeMarkdownInline(title), "")
|
||||||
|
|
||||||
|
if opts.IncludeMetadata {
|
||||||
|
lines = append(lines,
|
||||||
|
fmt.Sprintf("- Application: %s", escapeMarkdownInline(transcript.Metadata.Application)),
|
||||||
|
fmt.Sprintf("- Version: %s", escapeMarkdownInline(transcript.Metadata.Version)),
|
||||||
|
fmt.Sprintf("- Output schema: %s", escapeMarkdownInline(transcript.Schema)),
|
||||||
|
"",
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
for _, segment := range transcript.Segments {
|
||||||
|
parts := make([]string, 0, 4)
|
||||||
|
if opts.IncludeTimestamps {
|
||||||
|
parts = append(parts, fmt.Sprintf("[%s–%s]", formatTimestamp(segment.Start), formatTimestamp(segment.End)))
|
||||||
|
}
|
||||||
|
if opts.IncludeSegmentIDs {
|
||||||
|
parts = append(parts, fmt.Sprintf("[#%d]", segment.ID))
|
||||||
|
}
|
||||||
|
|
||||||
|
text := escapeMarkdownInline(segment.Text)
|
||||||
|
if shouldItalicize(segment.Categories) {
|
||||||
|
text = "*" + text + "*"
|
||||||
|
}
|
||||||
|
parts = append(parts, fmt.Sprintf("**%s:** %s", escapeMarkdownInline(segment.Speaker), text))
|
||||||
|
lines = append(lines, strings.Join(parts, " "))
|
||||||
|
lines = append(lines, "")
|
||||||
|
}
|
||||||
|
|
||||||
|
output := strings.Join(lines, "\n")
|
||||||
|
if !strings.HasSuffix(output, "\n") {
|
||||||
|
output += "\n"
|
||||||
|
}
|
||||||
|
return output, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func escapeMarkdownInline(value string) string {
|
||||||
|
replacer := strings.NewReplacer(
|
||||||
|
`\`, `\\`,
|
||||||
|
"`", "\\`",
|
||||||
|
"*", "\\*",
|
||||||
|
"_", "\\_",
|
||||||
|
"{", "\\{",
|
||||||
|
"}", "\\}",
|
||||||
|
"[", "\\[",
|
||||||
|
"]", "\\]",
|
||||||
|
"(", "\\(",
|
||||||
|
")", "\\)",
|
||||||
|
"#", "\\#",
|
||||||
|
"+", "\\+",
|
||||||
|
"!", "\\!",
|
||||||
|
"|", "\\|",
|
||||||
|
"<", "\\<",
|
||||||
|
">", "\\>",
|
||||||
|
)
|
||||||
|
return replacer.Replace(value)
|
||||||
|
}
|
||||||
|
|
||||||
|
func shouldItalicize(categories []string) bool {
|
||||||
|
for _, category := range categories {
|
||||||
|
switch category {
|
||||||
|
case "background", "backchannel", "filler":
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
func formatTimestamp(seconds float64) string {
|
||||||
|
total := int(math.Round(seconds))
|
||||||
|
if total < 0 {
|
||||||
|
total = 0
|
||||||
|
}
|
||||||
|
hours := total / 3600
|
||||||
|
minutes := (total % 3600) / 60
|
||||||
|
remainder := total % 60
|
||||||
|
return fmt.Sprintf("%02d:%02d:%02d", hours, minutes, remainder)
|
||||||
|
}
|
||||||
227
internal/render/markdown_test.go
Normal file
227
internal/render/markdown_test.go
Normal file
@@ -0,0 +1,227 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
import (
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestMarkdownRendererDefaultTranscriptShape(t *testing.T) {
|
||||||
|
transcript := Transcript{
|
||||||
|
Schema: "seriatim-intermediate",
|
||||||
|
Metadata: Metadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
},
|
||||||
|
Segments: []Segment{
|
||||||
|
{ID: 1, Start: 1, End: 4, Speaker: "Eric", Text: "Hello there."},
|
||||||
|
{ID: 2, Start: 5, End: 8, Speaker: "Mike", Text: "Welcome back, everyone."},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
output, err := MarkdownRenderer{}.Render(transcript, Options{
|
||||||
|
Title: "Transcript",
|
||||||
|
IncludeTimestamps: true,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render markdown: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
if !strings.Contains(output, "# Transcript") {
|
||||||
|
t.Fatalf("expected title in output:\n%s", output)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "[00:00:01–00:00:04] **Eric:** Hello there.") {
|
||||||
|
t.Fatalf("expected first segment in output:\n%s", output)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "[00:00:05–00:00:08] **Mike:** Welcome back, everyone.") {
|
||||||
|
t.Fatalf("expected second segment in output:\n%s", output)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestMarkdownRendererWithoutTimestamps(t *testing.T) {
|
||||||
|
transcript := Transcript{
|
||||||
|
Segments: []Segment{
|
||||||
|
{ID: 1, Start: 1, End: 4, Speaker: "Eric", Text: "Hello."},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
output, err := MarkdownRenderer{}.Render(transcript, Options{
|
||||||
|
Title: "Transcript",
|
||||||
|
IncludeTimestamps: false,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render markdown: %v", err)
|
||||||
|
}
|
||||||
|
if strings.Contains(output, "[00:00:01") {
|
||||||
|
t.Fatalf("timestamps should be omitted:\n%s", output)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "**Eric:** Hello.") {
|
||||||
|
t.Fatalf("expected speaker/text line:\n%s", output)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestMarkdownRendererWithSegmentIDs(t *testing.T) {
|
||||||
|
transcript := Transcript{
|
||||||
|
Segments: []Segment{
|
||||||
|
{ID: 17, Start: 1, End: 4, Speaker: "Eric", Text: "Hello."},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
output, err := MarkdownRenderer{}.Render(transcript, Options{
|
||||||
|
Title: "Transcript",
|
||||||
|
IncludeTimestamps: true,
|
||||||
|
IncludeSegmentIDs: true,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render markdown: %v", err)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "[#17]") {
|
||||||
|
t.Fatalf("expected segment ID in output:\n%s", output)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestMarkdownRendererMetadataOnlyWhenRequested(t *testing.T) {
|
||||||
|
transcript := Transcript{
|
||||||
|
Schema: "seriatim-full",
|
||||||
|
Metadata: Metadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
withMetadata, err := MarkdownRenderer{}.Render(transcript, Options{
|
||||||
|
Title: "Transcript",
|
||||||
|
IncludeMetadata: true,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render with metadata: %v", err)
|
||||||
|
}
|
||||||
|
if !strings.Contains(withMetadata, "- Application: seriatim") {
|
||||||
|
t.Fatalf("expected metadata block:\n%s", withMetadata)
|
||||||
|
}
|
||||||
|
|
||||||
|
withoutMetadata, err := MarkdownRenderer{}.Render(transcript, Options{Title: "Transcript"})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render without metadata: %v", err)
|
||||||
|
}
|
||||||
|
if strings.Contains(withoutMetadata, "- Application: seriatim") {
|
||||||
|
t.Fatalf("metadata should be omitted:\n%s", withoutMetadata)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestMarkdownRendererEscapesUserProvidedMarkdown(t *testing.T) {
|
||||||
|
transcript := Transcript{
|
||||||
|
Schema: "seriatim-intermediate",
|
||||||
|
Metadata: Metadata{
|
||||||
|
Application: "seriatim *cli*",
|
||||||
|
Version: "v[test]",
|
||||||
|
},
|
||||||
|
Segments: []Segment{
|
||||||
|
{
|
||||||
|
ID: 1,
|
||||||
|
Start: 1,
|
||||||
|
End: 2,
|
||||||
|
Speaker: "Dr. *A_[1]",
|
||||||
|
Text: "Use *literal* [link](target) and `code` \\ slash!",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
ID: 2,
|
||||||
|
Start: 2,
|
||||||
|
End: 3,
|
||||||
|
Speaker: "Narrator",
|
||||||
|
Text: "_aside_ with | pipe",
|
||||||
|
Categories: []string{"background"},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
output, err := MarkdownRenderer{}.Render(transcript, Options{
|
||||||
|
Title: "# Planning [notes]",
|
||||||
|
IncludeTimestamps: false,
|
||||||
|
IncludeMetadata: true,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render markdown: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
assertContains(t, output, "# \\# Planning \\[notes\\]")
|
||||||
|
assertContains(t, output, "- Application: seriatim \\*cli\\*")
|
||||||
|
assertContains(t, output, "- Version: v\\[test\\]")
|
||||||
|
assertContains(t, output, "**Dr. \\*A\\_\\[1\\]:** Use \\*literal\\* \\[link\\]\\(target\\) and \\`code\\` \\\\ slash\\!")
|
||||||
|
assertContains(t, output, "**Narrator:** *\\_aside\\_ with \\| pipe*")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestMarkdownRendererCategoryHintItalicsAndUnknownCategories(t *testing.T) {
|
||||||
|
transcript := Transcript{
|
||||||
|
Segments: []Segment{
|
||||||
|
{ID: 1, Start: 1, End: 2, Speaker: "A", Text: "bg", Categories: []string{"background"}},
|
||||||
|
{ID: 2, Start: 2, End: 3, Speaker: "B", Text: "bc", Categories: []string{"backchannel"}},
|
||||||
|
{ID: 3, Start: 3, End: 4, Speaker: "C", Text: "fill", Categories: []string{"filler"}},
|
||||||
|
{ID: 4, Start: 4, End: 5, Speaker: "D", Text: "plain", Categories: []string{"unknown-tag"}},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
output, err := MarkdownRenderer{}.Render(transcript, Options{
|
||||||
|
Title: "Transcript",
|
||||||
|
IncludeTimestamps: false,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("render markdown: %v", err)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "**A:** *bg*") {
|
||||||
|
t.Fatalf("expected background italics:\n%s", output)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "**B:** *bc*") {
|
||||||
|
t.Fatalf("expected backchannel italics:\n%s", output)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "**C:** *fill*") {
|
||||||
|
t.Fatalf("expected filler italics:\n%s", output)
|
||||||
|
}
|
||||||
|
if !strings.Contains(output, "**D:** plain") {
|
||||||
|
t.Fatalf("expected unknown category to be ignored:\n%s", output)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestMarkdownRendererIsDeterministic(t *testing.T) {
|
||||||
|
transcript := Transcript{
|
||||||
|
Schema: "seriatim-intermediate",
|
||||||
|
Metadata: Metadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
},
|
||||||
|
Segments: []Segment{
|
||||||
|
{ID: 1, Start: 1.2, End: 4.4, Speaker: "Eric", Text: "Hello there.", Categories: []string{"unknown-tag"}},
|
||||||
|
{ID: 2, Start: 65.1, End: 68.8, Speaker: "Mike", Text: "Yeah.", Categories: []string{"backchannel"}},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
opts := Options{
|
||||||
|
Title: "Transcript",
|
||||||
|
IncludeTimestamps: true,
|
||||||
|
IncludeSegmentIDs: true,
|
||||||
|
IncludeMetadata: true,
|
||||||
|
}
|
||||||
|
|
||||||
|
first, err := MarkdownRenderer{}.Render(transcript, opts)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("first render failed: %v", err)
|
||||||
|
}
|
||||||
|
second, err := MarkdownRenderer{}.Render(transcript, opts)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("second render failed: %v", err)
|
||||||
|
}
|
||||||
|
if first != second {
|
||||||
|
t.Fatalf("render output is not deterministic:\nfirst:\n%s\nsecond:\n%s", first, second)
|
||||||
|
}
|
||||||
|
if !strings.Contains(first, "[00:00:01–00:00:04] [#1] **Eric:** Hello there.") {
|
||||||
|
t.Fatalf("expected HH:MM:SS timestamp formatting:\n%s", first)
|
||||||
|
}
|
||||||
|
if !strings.Contains(first, "[00:01:05–00:01:09] [#2] **Mike:** *Yeah.*") {
|
||||||
|
t.Fatalf("expected HH:MM:SS timestamp formatting:\n%s", first)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func assertContains(t *testing.T, value string, want string) {
|
||||||
|
t.Helper()
|
||||||
|
if !strings.Contains(value, want) {
|
||||||
|
t.Fatalf("expected output to contain %q:\n%s", want, value)
|
||||||
|
}
|
||||||
|
}
|
||||||
24
internal/render/model.go
Normal file
24
internal/render/model.go
Normal file
@@ -0,0 +1,24 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
// Transcript is the render-normalized transcript model used by renderers.
|
||||||
|
type Transcript struct {
|
||||||
|
Schema string
|
||||||
|
Metadata Metadata
|
||||||
|
Segments []Segment
|
||||||
|
}
|
||||||
|
|
||||||
|
// Metadata is the render-relevant artifact metadata.
|
||||||
|
type Metadata struct {
|
||||||
|
Application string
|
||||||
|
Version string
|
||||||
|
}
|
||||||
|
|
||||||
|
// Segment is a normalized render segment.
|
||||||
|
type Segment struct {
|
||||||
|
ID int
|
||||||
|
Start float64
|
||||||
|
End float64
|
||||||
|
Speaker string
|
||||||
|
Text string
|
||||||
|
Categories []string
|
||||||
|
}
|
||||||
96
internal/render/normalize.go
Normal file
96
internal/render/normalize.go
Normal file
@@ -0,0 +1,96 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/internal/artifact"
|
||||||
|
)
|
||||||
|
|
||||||
|
// FromOutputArtifact converts a parsed output artifact into the internal render model.
|
||||||
|
func FromOutputArtifact(input artifact.OutputArtifact) (Transcript, error) {
|
||||||
|
switch input.Schema {
|
||||||
|
case artifact.OutputSchemaFull:
|
||||||
|
payload, err := input.FullPayload()
|
||||||
|
if err != nil {
|
||||||
|
return Transcript{}, err
|
||||||
|
}
|
||||||
|
segments := make([]Segment, len(payload.Segments))
|
||||||
|
for index, segment := range payload.Segments {
|
||||||
|
segments[index] = Segment{
|
||||||
|
ID: segment.ID,
|
||||||
|
Start: segment.Start,
|
||||||
|
End: segment.End,
|
||||||
|
Speaker: segment.Speaker,
|
||||||
|
Text: segment.Text,
|
||||||
|
Categories: normalizeCategories(segment.Categories),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return Transcript{
|
||||||
|
Schema: input.Schema,
|
||||||
|
Metadata: Metadata{
|
||||||
|
Application: payload.Metadata.Application,
|
||||||
|
Version: payload.Metadata.Version,
|
||||||
|
},
|
||||||
|
Segments: segments,
|
||||||
|
}, nil
|
||||||
|
case artifact.OutputSchemaIntermediate:
|
||||||
|
payload, err := input.IntermediatePayload()
|
||||||
|
if err != nil {
|
||||||
|
return Transcript{}, err
|
||||||
|
}
|
||||||
|
segments := make([]Segment, len(payload.Segments))
|
||||||
|
for index, segment := range payload.Segments {
|
||||||
|
segments[index] = Segment{
|
||||||
|
ID: segment.ID,
|
||||||
|
Start: segment.Start,
|
||||||
|
End: segment.End,
|
||||||
|
Speaker: segment.Speaker,
|
||||||
|
Text: segment.Text,
|
||||||
|
Categories: normalizeCategories(segment.Categories),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return Transcript{
|
||||||
|
Schema: input.Schema,
|
||||||
|
Metadata: Metadata{
|
||||||
|
Application: payload.Metadata.Application,
|
||||||
|
Version: payload.Metadata.Version,
|
||||||
|
},
|
||||||
|
Segments: segments,
|
||||||
|
}, nil
|
||||||
|
case artifact.OutputSchemaMinimal:
|
||||||
|
payload, err := input.MinimalPayload()
|
||||||
|
if err != nil {
|
||||||
|
return Transcript{}, err
|
||||||
|
}
|
||||||
|
segments := make([]Segment, len(payload.Segments))
|
||||||
|
for index, segment := range payload.Segments {
|
||||||
|
segments[index] = Segment{
|
||||||
|
ID: segment.ID,
|
||||||
|
Start: segment.Start,
|
||||||
|
End: segment.End,
|
||||||
|
Speaker: segment.Speaker,
|
||||||
|
Text: segment.Text,
|
||||||
|
Categories: []string{},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return Transcript{
|
||||||
|
Schema: input.Schema,
|
||||||
|
Metadata: Metadata{
|
||||||
|
Application: payload.Metadata.Application,
|
||||||
|
Version: payload.Metadata.Version,
|
||||||
|
},
|
||||||
|
Segments: segments,
|
||||||
|
}, nil
|
||||||
|
default:
|
||||||
|
return Transcript{}, fmt.Errorf("unsupported artifact schema %q", input.Schema)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func normalizeCategories(categories []string) []string {
|
||||||
|
if categories == nil {
|
||||||
|
return []string{}
|
||||||
|
}
|
||||||
|
out := make([]string, len(categories))
|
||||||
|
copy(out, categories)
|
||||||
|
return out
|
||||||
|
}
|
||||||
139
internal/render/normalize_test.go
Normal file
139
internal/render/normalize_test.go
Normal file
@@ -0,0 +1,139 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/internal/artifact"
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/schema"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestFromOutputArtifactNormalizesSupportedSchemas(t *testing.T) {
|
||||||
|
t.Run("full", func(t *testing.T) {
|
||||||
|
sourceIndex := 0
|
||||||
|
input := schema.Transcript{
|
||||||
|
Metadata: schema.Metadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
InputReader: "json-files",
|
||||||
|
InputFiles: []string{"a.json"},
|
||||||
|
PreprocessingModules: []string{"validate-raw"},
|
||||||
|
PostprocessingModules: []string{"assign-ids", "validate-output"},
|
||||||
|
OutputModules: []string{"json"},
|
||||||
|
},
|
||||||
|
Segments: []schema.Segment{
|
||||||
|
{
|
||||||
|
ID: 1,
|
||||||
|
Source: "a.json",
|
||||||
|
SourceSegmentIndex: &sourceIndex,
|
||||||
|
Speaker: "Alice",
|
||||||
|
Start: 1,
|
||||||
|
End: 2,
|
||||||
|
Text: "hello",
|
||||||
|
Categories: []string{"background"},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
OverlapGroups: []schema.OverlapGroup{},
|
||||||
|
}
|
||||||
|
model := mustNormalizeOutputArtifact(t, input)
|
||||||
|
if model.Schema != artifact.OutputSchemaFull {
|
||||||
|
t.Fatalf("schema = %q, want %q", model.Schema, artifact.OutputSchemaFull)
|
||||||
|
}
|
||||||
|
if len(model.Segments) != 1 {
|
||||||
|
t.Fatalf("segment count = %d, want 1", len(model.Segments))
|
||||||
|
}
|
||||||
|
if model.Segments[0].ID != 1 || model.Segments[0].Speaker != "Alice" || model.Segments[0].Text != "hello" {
|
||||||
|
t.Fatalf("unexpected segment: %#v", model.Segments[0])
|
||||||
|
}
|
||||||
|
if len(model.Segments[0].Categories) != 1 || model.Segments[0].Categories[0] != "background" {
|
||||||
|
t.Fatalf("categories = %#v, want [background]", model.Segments[0].Categories)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
t.Run("intermediate", func(t *testing.T) {
|
||||||
|
input := schema.IntermediateTranscript{
|
||||||
|
Metadata: schema.IntermediateMetadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
OutputSchema: artifact.OutputSchemaIntermediate,
|
||||||
|
},
|
||||||
|
Segments: []schema.IntermediateSegment{
|
||||||
|
{ID: 1, Start: 1, End: 2, Speaker: "Alice", Text: "one", Categories: []string{}},
|
||||||
|
{ID: 2, Start: 2, End: 3, Speaker: "Bob", Text: "two"},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
model := mustNormalizeOutputArtifact(t, input)
|
||||||
|
if model.Schema != artifact.OutputSchemaIntermediate {
|
||||||
|
t.Fatalf("schema = %q, want %q", model.Schema, artifact.OutputSchemaIntermediate)
|
||||||
|
}
|
||||||
|
if len(model.Segments[0].Categories) != 0 {
|
||||||
|
t.Fatalf("segment[0] categories = %#v, want empty slice", model.Segments[0].Categories)
|
||||||
|
}
|
||||||
|
if len(model.Segments[1].Categories) != 0 {
|
||||||
|
t.Fatalf("segment[1] categories = %#v, want empty slice", model.Segments[1].Categories)
|
||||||
|
}
|
||||||
|
if model.Segments[0].Categories == nil || model.Segments[1].Categories == nil {
|
||||||
|
t.Fatal("expected non-nil empty categories slices")
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
t.Run("minimal", func(t *testing.T) {
|
||||||
|
input := schema.MinimalTranscript{
|
||||||
|
Metadata: schema.MinimalMetadata{
|
||||||
|
Application: "seriatim",
|
||||||
|
Version: "v-test",
|
||||||
|
OutputSchema: artifact.OutputSchemaMinimal,
|
||||||
|
},
|
||||||
|
Segments: []schema.MinimalSegment{
|
||||||
|
{ID: 1, Start: 1, End: 2, Speaker: "Alice", Text: "one"},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
model := mustNormalizeOutputArtifact(t, input)
|
||||||
|
if model.Schema != artifact.OutputSchemaMinimal {
|
||||||
|
t.Fatalf("schema = %q, want %q", model.Schema, artifact.OutputSchemaMinimal)
|
||||||
|
}
|
||||||
|
if len(model.Segments[0].Categories) != 0 {
|
||||||
|
t.Fatalf("categories = %#v, want empty slice", model.Segments[0].Categories)
|
||||||
|
}
|
||||||
|
if model.Segments[0].Categories == nil {
|
||||||
|
t.Fatal("expected non-nil empty categories slice")
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestFromOutputArtifactRejectsMalformedAndRawInput(t *testing.T) {
|
||||||
|
_, err := artifact.ParseOutputArtifactJSON([]byte(`{"metadata":`))
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected malformed JSON error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "input JSON is malformed") {
|
||||||
|
t.Fatalf("unexpected malformed error: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
rawWhisper := []byte(`{"segments":[{"id":0,"start":0.1,"end":1.2,"text":"hello","words":[{"word":"hello"}]}]}`)
|
||||||
|
_, err = artifact.ParseOutputArtifactJSON(rawWhisper)
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected raw input artifact error")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "not a valid seriatim output artifact") {
|
||||||
|
t.Fatalf("unexpected raw input error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func mustNormalizeOutputArtifact(t *testing.T, value any) Transcript {
|
||||||
|
t.Helper()
|
||||||
|
data, err := json.Marshal(value)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("marshal: %v", err)
|
||||||
|
}
|
||||||
|
parsed, err := artifact.ParseOutputArtifactJSON(data)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("parse: %v", err)
|
||||||
|
}
|
||||||
|
model, err := FromOutputArtifact(parsed)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("normalize: %v", err)
|
||||||
|
}
|
||||||
|
return model
|
||||||
|
}
|
||||||
41
internal/render/registry.go
Normal file
41
internal/render/registry.go
Normal file
@@ -0,0 +1,41 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
import "fmt"
|
||||||
|
|
||||||
|
const FormatMarkdown = "markdown"
|
||||||
|
|
||||||
|
// Options configures rendering behavior across formats.
|
||||||
|
type Options struct {
|
||||||
|
Title string
|
||||||
|
IncludeTimestamps bool
|
||||||
|
IncludeSegmentIDs bool
|
||||||
|
IncludeMetadata bool
|
||||||
|
}
|
||||||
|
|
||||||
|
// Renderer turns a normalized render model into text output.
|
||||||
|
type Renderer interface {
|
||||||
|
Render(transcript Transcript, opts Options) (string, error)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Registry resolves renderers by public format name.
|
||||||
|
type Registry struct {
|
||||||
|
renderers map[string]Renderer
|
||||||
|
}
|
||||||
|
|
||||||
|
// NewRegistry returns a renderer registry with built-in renderers.
|
||||||
|
func NewRegistry() Registry {
|
||||||
|
return Registry{
|
||||||
|
renderers: map[string]Renderer{
|
||||||
|
FormatMarkdown: MarkdownRenderer{},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Resolve resolves a renderer by format name.
|
||||||
|
func (registry Registry) Resolve(format string) (Renderer, error) {
|
||||||
|
renderer, ok := registry.renderers[format]
|
||||||
|
if !ok {
|
||||||
|
return nil, fmt.Errorf("unsupported --format %q", format)
|
||||||
|
}
|
||||||
|
return renderer, nil
|
||||||
|
}
|
||||||
22
internal/render/registry_test.go
Normal file
22
internal/render/registry_test.go
Normal file
@@ -0,0 +1,22 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
import "testing"
|
||||||
|
|
||||||
|
func TestRegistryResolvesMarkdownRenderer(t *testing.T) {
|
||||||
|
registry := NewRegistry()
|
||||||
|
renderer, err := registry.Resolve(FormatMarkdown)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("resolve markdown renderer: %v", err)
|
||||||
|
}
|
||||||
|
if renderer == nil {
|
||||||
|
t.Fatal("expected renderer")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRegistryRejectsUnknownRenderer(t *testing.T) {
|
||||||
|
registry := NewRegistry()
|
||||||
|
_, err := registry.Resolve("txt")
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected unsupported format error")
|
||||||
|
}
|
||||||
|
}
|
||||||
72
internal/render/run.go
Normal file
72
internal/render/run.go
Normal file
@@ -0,0 +1,72 @@
|
|||||||
|
package render
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/internal/artifact"
|
||||||
|
"gitea.maximumdirect.net/eric/seriatim/internal/config"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Run executes artifact-level render orchestration.
|
||||||
|
func Run(ctx context.Context, cfg config.RenderConfig) error {
|
||||||
|
if err := ctx.Err(); err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
data, err := os.ReadFile(cfg.InputFile)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("read --input-file %q: %w", cfg.InputFile, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
inputArtifact, err := artifact.ParseOutputArtifactJSON(data)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("--input-file %q: %w", cfg.InputFile, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
model, err := FromOutputArtifact(inputArtifact)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("normalize artifact for render: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
registry := NewRegistry()
|
||||||
|
renderer, err := registry.Resolve(cfg.Format)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
rendered, err := renderer.Render(model, Options{
|
||||||
|
Title: cfg.Title,
|
||||||
|
IncludeTimestamps: cfg.IncludeTimestamps,
|
||||||
|
IncludeSegmentIDs: cfg.IncludeSegmentIDs,
|
||||||
|
IncludeMetadata: cfg.IncludeMetadata,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("render %q output: %w", cfg.Format, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
if err := writeFile(cfg.OutputFile, rendered); err != nil {
|
||||||
|
return fmt.Errorf("write --output-file %q: %w", cfg.OutputFile, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func writeFile(path string, content string) (err error) {
|
||||||
|
file, err := os.Create(path)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("create %q: %w", path, err)
|
||||||
|
}
|
||||||
|
defer func() {
|
||||||
|
closeErr := file.Close()
|
||||||
|
if err == nil && closeErr != nil {
|
||||||
|
err = fmt.Errorf("close %q: %w", path, closeErr)
|
||||||
|
}
|
||||||
|
}()
|
||||||
|
|
||||||
|
if _, err := file.WriteString(content); err != nil {
|
||||||
|
return fmt.Errorf("write %q: %w", path, err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
@@ -1,16 +1,16 @@
|
|||||||
package trim
|
package trim
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"encoding/json"
|
|
||||||
"fmt"
|
"fmt"
|
||||||
|
|
||||||
|
artifactpkg "gitea.maximumdirect.net/eric/seriatim/internal/artifact"
|
||||||
"gitea.maximumdirect.net/eric/seriatim/schema"
|
"gitea.maximumdirect.net/eric/seriatim/schema"
|
||||||
)
|
)
|
||||||
|
|
||||||
const (
|
const (
|
||||||
SchemaMinimal = schema.OutputSchemaMinimal
|
SchemaMinimal = artifactpkg.OutputSchemaMinimal
|
||||||
SchemaIntermediate = schema.OutputSchemaIntermediate
|
SchemaIntermediate = artifactpkg.OutputSchemaIntermediate
|
||||||
SchemaFull = schema.OutputSchemaFull
|
SchemaFull = artifactpkg.OutputSchemaFull
|
||||||
)
|
)
|
||||||
|
|
||||||
// Artifact stores a parsed seriatim output artifact of one supported schema.
|
// Artifact stores a parsed seriatim output artifact of one supported schema.
|
||||||
@@ -31,42 +31,16 @@ type ApplyArtifactResult struct {
|
|||||||
|
|
||||||
// ParseArtifactJSON parses and validates a serialized seriatim output artifact.
|
// ParseArtifactJSON parses and validates a serialized seriatim output artifact.
|
||||||
func ParseArtifactJSON(data []byte) (Artifact, error) {
|
func ParseArtifactJSON(data []byte) (Artifact, error) {
|
||||||
var decoded any
|
parsed, err := artifactpkg.ParseOutputArtifactJSON(data)
|
||||||
if err := json.Unmarshal(data, &decoded); err != nil {
|
if err != nil {
|
||||||
return Artifact{}, fmt.Errorf("input JSON is malformed: %w", err)
|
return Artifact{}, err
|
||||||
}
|
}
|
||||||
|
return Artifact{
|
||||||
var full schema.Transcript
|
Schema: parsed.Schema,
|
||||||
if err := json.Unmarshal(data, &full); err == nil {
|
Full: parsed.Full,
|
||||||
if err := schema.ValidateTranscript(full); err == nil {
|
Intermediate: parsed.Intermediate,
|
||||||
return Artifact{
|
Minimal: parsed.Minimal,
|
||||||
Schema: SchemaFull,
|
}, nil
|
||||||
Full: &full,
|
|
||||||
}, nil
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
var intermediate schema.IntermediateTranscript
|
|
||||||
if err := json.Unmarshal(data, &intermediate); err == nil {
|
|
||||||
if err := schema.ValidateIntermediateTranscript(intermediate); err == nil {
|
|
||||||
return Artifact{
|
|
||||||
Schema: SchemaIntermediate,
|
|
||||||
Intermediate: &intermediate,
|
|
||||||
}, nil
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
var minimal schema.MinimalTranscript
|
|
||||||
if err := json.Unmarshal(data, &minimal); err == nil {
|
|
||||||
if err := schema.ValidateMinimalTranscript(minimal); err == nil {
|
|
||||||
return Artifact{
|
|
||||||
Schema: SchemaMinimal,
|
|
||||||
Minimal: &minimal,
|
|
||||||
}, nil
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
return Artifact{}, fmt.Errorf("input JSON is not a valid seriatim output artifact")
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// ValidateArtifact validates an artifact against its declared schema.
|
// ValidateArtifact validates an artifact against its declared schema.
|
||||||
|
|||||||
Reference in New Issue
Block a user