10 KiB
Architecture Policy
Purpose
This document defines seriatim's development architecture and invariants for maintainers and automated coding agents. It describes how the implemented system is intended to be built and changed. It is not a user manual, CLI reference, config reference, or roadmap.
Keep this document aligned with documentation policy. It
must describe current behavior only; planned or speculative work belongs under
docs/roadmap/.
Project Shape
seriatim is a Go CLI for transcript artifact processing. The implemented
commands are merge, trim, and normalize.
merge reads one or more JSON transcript files, optionally maps input files to
canonical speakers, runs a registry-selected preprocessing chain, merges
canonical segments into deterministic chronological order, runs a
registry-selected postprocessing chain, validates the selected output schema,
and writes JSON output plus an optional JSON report.
trim and normalize are artifact-level commands outside the merge pipeline.
trim reads an existing seriatim output artifact and projects it by segment ID.
normalize reads transcript-like JSON and emits one of seriatim's supported
output schemas. Neither command runs merge preprocessing or postprocessing
modules.
The supported public output schemas are seriatim-minimal,
seriatim-intermediate, and seriatim-full. For command and flag details, use
CLI reference and configuration reference.
Core Design Principles
- Keep a hexagonal architecture boundary. Domain models, stage contracts, and deterministic transformations must stay separate from CLI parsing, filesystem access, config loading, reporting, and other external adapters.
- Keep stages and modules composable. Built-in modules are selected by canonical registry names and implement explicit interfaces for their pipeline role.
- Preserve deterministic behavior. Given the same inputs, configuration, and version, output ordering, segment IDs, schema validation, and report event ordering should remain stable.
- Current command execution is sequential. There is no scheduler, worker pool, or concurrent module execution in the implemented pipeline. Any concurrency added later must be bounded, observable, and must not make output handling nondeterministic.
- Prefer the Go standard library. Third-party dependencies should remain narrow and justified, such as Cobra for CLI structure, YAML parsing, and JSON Schema validation.
- Document current behavior. Architecture, user, and internal docs must not describe planned features as implemented behavior.
Architectural Boundaries
Core transcript data belongs in internal/model and public artifact contracts
belong in schema. Conversion from internal merged data to public JSON shapes
belongs at the artifact boundary, not inside CLI code or transformation
packages.
Pipeline orchestration belongs in internal/pipeline. It resolves registered
modules, validates preprocessing state transitions, executes stages in order,
collects report events, converts the final transcript, and writes optional
reports. Built-in adapters and modules are registered from internal/builtin.
CLI code in internal/cli should parse flags, build validated config values,
and delegate. merge delegates to pipeline.Run; trim and normalize
perform artifact-level orchestration and delegate deterministic parsing,
validation, and transformation work to their internal packages.
Config loading and validation belongs in internal/config. Filesystem reads and
writes are adapter concerns and should not spread into pure transformation
helpers. Existing built-in modules that load configured YAML files must keep
that I/O narrow and explicit.
Reports belong in internal/report. Modules and commands should emit concise
events for validation findings, corrections, and transformations without
turning report messages into a duplicate output artifact.
Tests and samples are supporting evidence for behavior. Tests should verify stable contracts and edge cases; samples should remain valid examples, not hidden architecture dependencies.
Modules or Stages
The merge pipeline has these implemented stages:
InputReader: reads configured external input into raw transcripts.Preprocessor: transformsPreprocessStatefrom raw to canonical state.Merger: combines canonical transcripts into one merged transcript.Postprocessor: transforms or annotates the merged transcript.OutputWriter: writes the selected output artifact.
Modules must keep narrow responsibilities, declare their stage through the
interface they implement, and use explicit config values. Preprocessors must
declare Requires() and Produces() states; the runner rejects invalid
raw/canonical ordering before processing completes.
Modules run in the configured order. Order-affecting modules must run before
assign-ids, and validate-output must see final IDs that match the selected
schema. Accepted and rejected transformations should be deterministic and, when
observable, recorded through report events.
Transformation helpers should avoid hidden global state. Shared caches, such as compiled JSON schemas, must be protected and must not affect output ordering.
State, Inputs, and Outputs
seriatim is file-based. It reads JSON inputs and optional YAML rule files, then writes JSON transcript artifacts and optional JSON reports.
The implemented application has no durable database, daemon state, resume state, remote storage, or background job state. Runtime state is held in memory for the current command invocation and serialized only through requested output and report files.
Input file paths are normalized and validated during config construction.
merge sorts input file paths before processing, then uses stable segment sort
keys. trim preserves transcript order while renumbering retained IDs.
normalize sorts by implemented deterministic keys and assigns fresh IDs.
Configuration and CLI Boundaries
The CLI surface is an adapter over validated config structs. Cobra command code should stay thin: parse flags, account for flag/default precedence, call config constructors, and delegate.
Config constructors validate required paths, output parent directories, module lists, selected schemas, mutually exclusive trim selector options, and supported environment-derived settings. Module name validation is split between config where command-specific names are fixed and the pipeline registry where module composition is resolved.
Do not duplicate full CLI or config reference material here. Use CLI reference and configuration reference for canonical user-facing details.
Errors, Logging, and Diagnostics
Commands return errors instead of printing inside deep logic. The root command
silences Cobra usage/error output, and cmd/seriatim/main.go prints one error
to stderr and exits with status 1.
Validation failures should fail fast with contextual errors. Correctable conditions should be deterministic and, where reports are requested, reflected as report events. Optional reports contain metadata and ordered events; they are not required for command success unless the report file itself cannot be written.
The implemented code does not use a logging subsystem. Diagnostics are returned as errors or written to optional report JSON. Normalize report events avoid embedding transcript text; keep that privacy-oriented behavior when changing normalize diagnostics.
Testing Expectations
go test ./... is the repository-wide check. There is currently no Makefile,
taskfile, linter config, or dedicated documentation check.
When changing config or CLI behavior, inspect internal/config and
internal/cli tests. When changing pipeline composition or stage contracts,
inspect internal/pipeline and internal/builtin tests. When changing
correction or annotation modules, inspect the package tests for overlap,
coalesce, danglers, backchannel, filler, and autocorrect behavior.
When changing artifact-level commands, inspect internal/trim,
internal/normalize, and their CLI tests. When changing public output shape or
schema validation, inspect schema and internal/artifact tests. Report and
diagnostic changes should be covered through the command or package tests that
emit the affected events.
Dependency Policy
Prefer the Go standard library for parsing, data transformation, concurrency primitives, filesystem work, and testing wherever it is reasonable.
Third-party dependencies must be narrow, justified, and preferably de facto
standard for their purpose. Existing examples include Cobra for CLI structure,
gopkg.in/yaml.v3 for YAML files, and jsonschema/v6 for validating embedded
public JSON schemas. Avoid broad framework dependencies for behavior that is
already simple and local.
Documentation Expectations
Architecture docs must stay aligned with documentation policy.
Current-behavior docs must not become aspirational. If code and docs disagree,
fix the inaccurate current-behavior doc or put planned work under
docs/roadmap/.
Prefer links to canonical docs instead of repeating full CLI, config, schema, or operations reference material. Keep examples real, tested where practical, and free of secrets or private transcript data.
Architectural Invariants
- Keep core/domain logic separate from CLI, config, filesystem, reporting, and other adapter concerns.
- Centralize default configuration values as constants defined in internal/config/config.go.
- Keep modules narrowly scoped, explicitly configured, and composable by registry name.
- Preserve deterministic ordering, final segment ID assignment, and schema validation before output acceptance.
- Keep
trimandnormalizeartifact-level; do not run merge modules from those commands. - Keep public output schemas validated through
schema. - Keep optional reports ordered, concise, and diagnostic.
- Avoid broad dependencies without a concrete maintainability benefit.
- Do not document unimplemented behavior outside
docs/roadmap/.
Non-Goals
The implemented application does not perform transcription, audio diarization, speaker inference from audio or text, summarization, daemon operation, remote storage, dynamic external plugin loading, or concurrent pipeline execution.
The architecture policy is not a package-by-package reference, CLI manual, config reference, schema reference, or roadmap.