Files
distributor/docs/roadmap/implementation.md

346 lines
14 KiB
Markdown

# Catalog State And Workflow Implementation Roadmap
This is the active staged implementation plan for
`docs/roadmap/catalog.md`. The feature roadmap defines the target state model
and policy decisions; this document defines the implementation sequence for an
LLM coding agent to follow stage by stage.
Future behavior must remain under `docs/roadmap/` until implemented. Update
current-behavior docs only in the stage that implements the corresponding
behavior.
## Current Baseline
`distributor` currently writes destination `.distributor.json` using
single-owner schema version `2` or shared-root schema version `3`. Destination
behavior is selected through several low-level knobs: `state.mode`,
`reconciliation.mode`, `takeover.mode`, and `transfer`.
The target behavior is a clean alpha break:
- all newly written destination state uses schema version `4`;
- all destinations use `state.mode: catalog` internally;
- user-facing destination behavior is selected by `workflow: additive` or
`workflow: replacement`;
- `workflow` defaults to `additive`;
- legacy destination policy fields are rejected, not aliased;
- legacy state schema versions `1`, `2`, and `3` are not migrated.
## Implementation Principles
- Keep the source manifest contract unchanged.
- Keep backend adapters unaware of catalog and workflow policy.
- Treat `workflow` as runtime config, not persisted state.
- Keep unmanaged content protected after catalog state exists.
- Preserve output `created_at` when a path changes owner; update only
`updated_at`.
- Prefer removing legacy state/config paths over compatibility shims.
- Preserve current public APIs outside destination state/config unless the
catalog roadmap explicitly changes them.
## Active Implementation Stages
## Stage 1: Destination Workflow Config Clean Break
Goal: replace low-level destination policy config with the user-facing workflow
switch.
Implementation scope:
- Add destination `workflow` config with accepted values `additive` and
`replacement`.
- Default omitted `workflow` to `additive`.
- Remove or make unsupported the destination-level config fields `state`,
`reconciliation`, `takeover`, and `transfer`.
- Ensure configs containing those legacy fields fail clearly. Prefer strict YAML
unknown-field failure by removing struct fields; add explicit validation only
if clearer errors are needed without weakening strict decoding.
- Preserve backend, publish, transform, path mapping, links, retention, and
other non-policy destination fields.
- Update example configs only when the implementation stage also updates
current docs; otherwise keep this stage focused on config code and tests.
Tests:
- `go test ./internal/config`
- Omitted workflow defaults to `additive`.
- `workflow: additive` and `workflow: replacement` validate.
- Unknown workflow values fail.
- Legacy `state`, `reconciliation`, `takeover`, and `transfer` fields fail.
- Existing valid examples are updated or tests are adjusted in the same change if
examples currently use legacy fields.
Completion criteria:
- New configs express destination replacement/additive intent with one field.
- No normal config path accepts legacy destination policy knobs.
## Stage 2: Catalog State Schema Version 4
Goal: implement the schema version `4` catalog state model in `internal/state`.
Implementation scope:
- Add catalog state types for top-level schema version `4`, `state.mode:
catalog`, and `outputs`.
- Add compact per-output source identity with `id`, `digest`, and `created`.
- Add catalog output records with required `path`, `pipeline_id`,
`destination_id`, `source`, `kind`, `sha256`, `size`, `created_at`, and
`updated_at`.
- Add optional `source_path`, `transform`, and `url`, present only where allowed
by the catalog roadmap.
- Validate duplicate paths, invalid owner ids, invalid source identities,
invalid output paths, invalid digests, negative sizes, invalid timestamps,
invalid kind/transform combinations, and invalid URLs.
- Marshal catalog state deterministically with the existing JSON formatting
conventions.
- Update `ParseDocument` so schema version `4` returns catalog state.
- Treat schema versions lower than `4` as superseded legacy state for publish
planning, not as readable/migrated active state. Keep enough detection to
identify legacy state and avoid treating it as arbitrary invalid JSON.
- Reject unsupported future schema versions.
Tests:
- `go test ./internal/state`
- Parse/validate/marshal valid catalog state.
- Reject malformed catalog state and invalid output records.
- Detect schema versions `1`, `2`, and `3` as superseded legacy state.
- Reject future schema versions.
- Prove no top-level `owners`, `sources`, workflow, source manifest, pipeline id,
destination id, or published timestamp is accepted for catalog state.
Completion criteria:
- State package has one canonical schema version `4` catalog contract for new
writes.
- Legacy state is detected but not migrated.
## Stage 3: Catalog Publish Planning
Goal: plan publish actions against catalog state and destination workflow.
Implementation scope:
- Replace publish request inputs that consume state/reconciliation/takeover and
transfer policy with destination workflow.
- Map workflow to catalog actions:
- `workflow: additive`: upsert planned outputs and retain all other managed
catalog outputs.
- `workflow: replacement`: upsert planned outputs and delete catalog outputs
owned by the current pipeline/destination that are omitted from the plan.
- Planned paths that already exist as catalog-managed outputs may be overwritten
and become owned by the current pipeline/destination.
- Planned path collisions with storage content not recorded in catalog state
fail as unmanaged content.
- Matching planned outputs may skip writes when source identity and output digest
already match, if that optimization can be implemented without changing
externally visible results; otherwise writing idempotently is acceptable.
- Existing schema `< 4` state is superseded:
- replacement workflow may plan a bounded destination-root clear before
writing planned outputs;
- additive workflow may plan overwrites for planned paths only and leave
unplanned files unmanaged.
- Invalid JSON or future schema state remains a conflict, not a superseded
legacy state.
- Preserve path mapping, publish policy, transforms, links, fixed-path
selection, and output collision checks.
Tests:
- `go test ./internal/publish`
- Additive workflow publishes new catalog state.
- Additive workflow overwrites an existing managed output and retains unrelated
outputs.
- Replacement workflow removes omitted outputs owned by the current
pipeline/destination.
- Replacement workflow preserves unrelated catalog outputs owned by other
pipeline/destination pairs.
- Managed output ownership moves to the current pipeline/destination when a
planned path is overwritten.
- `created_at` is preserved when an existing path changes owner; `updated_at`
changes.
- Unmanaged path collisions fail.
- Legacy schema `< 4` state follows the superseded-state rules above.
- Invalid/future state fails.
Completion criteria:
- Publish planning no longer depends on single-owner/shared-root comparison
semantics.
- Workflow behavior is fully driven by `workflow`.
## Stage 4: Catalog Publish Execution
Goal: execute catalog publish plans and write schema version `4` state.
Implementation scope:
- Write only schema version `4` catalog `.distributor.json`.
- Implement additive execution as managed upsert of planned outputs plus catalog
state update.
- Implement replacement execution as managed upsert plus deletion of omitted
current-owner catalog outputs.
- For replacement over superseded legacy state, clear the bounded destination
root before writing planned outputs and catalog state.
- For additive over superseded legacy state, overwrite planned paths and write a
catalog containing only planned outputs; leave unplanned files unmanaged.
- Preserve cleanup behavior on failed writes:
- additive cleanup removes outputs newly created by the failed attempt where
practical;
- replacement cleanup follows existing managed-replacement safety where
practical;
- state is written only after selected outputs are written.
- Preserve link metadata, generated output digests, source output digests,
content type behavior, and backend-safe writes.
Tests:
- `go test ./internal/publish ./internal/app`
- Additive execution writes planned outputs, overwrites managed planned paths,
retains unrelated managed outputs, and writes catalog state.
- Replacement execution deletes omitted current-owner outputs and preserves
unrelated owner outputs.
- Superseded legacy replacement clears bounded destination root only.
- Superseded legacy additive leaves unplanned files on disk but out of catalog
state.
- Failed writes do not leave misleading catalog state.
- Local, fake-backed SSH, and fake-backed S3 app paths exercise the same publish
behavior.
Completion criteria:
- Successful `run` writes only v4 catalog state.
- Destination outputs match additive/replacement workflow semantics.
## Stage 5: Run Reporting, Notifications, And CLI Surface
Goal: update user-visible run behavior to describe workflow/catalog actions
instead of legacy replacement/takeover actions.
Implementation scope:
- Replace legacy action labels tied to state/reconciliation/takeover/transfer
with these stable workflow-oriented labels:
- `publish_new`
- `upsert_additive`
- `replace_catalog`
- `skip_same`
- `force_replace`
- `fail_unmanaged`
- `fail_conflict`
- Ensure text and JSON run summaries include workflow-relevant counters.
- Include `workflow` in run action records where useful.
- Update fixed-path dry-run warnings to describe additive upsert or replacement
clearly.
- Ensure notifications use the new action labels.
- Remove reporting assumptions that depend on `replace_older`,
`replace_newer`, `replace_conflict`, or `replace_takeover`.
Tests:
- `go test ./internal/app ./internal/cli`
- Text dry-run output distinguishes additive from replacement workflow.
- JSON output includes workflow and stable action labels.
- Summary counters are deterministic.
- Notifications fire for additive and replacement writes.
- Existing CLI commands still parse and execute with the new config shape.
Completion criteria:
- Operators can understand from dry-run output whether a destination will upsert
or replace managed catalog outputs.
## Stage 6: Catalog Prune And Reconcile-State
Goal: update maintenance commands to operate on catalog state with the agreed
selector model.
Implementation scope:
- Prune:
- selected `--pipeline` and `--destination` prune only outputs currently owned
by that pipeline/destination;
- no `--all-owners` prune mode in the initial catalog implementation;
- do not add `--source-id`, `--path-prefix`, or `--kind` selectors.
- Reconcile-state:
- selected `--pipeline` and `--destination` repair only outputs currently
owned by that pipeline/destination;
- `--all-owners` repairs all catalog outputs;
- unmanaged reporting compares storage entries to all catalog output paths,
not just selected owner paths.
- Rewrite repaired/pruned state as schema version `4`.
- Remove legacy single-owner/shared-root maintenance branches.
Tests:
- `go test ./internal/state ./internal/app ./internal/cli`
- Prune selects only current-owner catalog outputs.
- Prune deletes selected outputs and removes their catalog records.
- Reconcile-state selected owner removes only missing outputs for that owner.
- `reconcile-state --all-owners` removes missing outputs for all owners.
- Unmanaged reporting excludes all catalog-managed paths and reports unrecorded
storage entries.
- JSON/text output remains stable and clear.
Completion criteria:
- Maintenance commands operate only on catalog state and respect the locked
selector rules.
## Stage 7: Clean Break Removal And Documentation
Goal: remove legacy state/config behavior and document the implemented catalog
workflow model.
Implementation scope:
- Remove dead code for writing schema version `2` single-owner and schema
version `3` shared-root state.
- Remove legacy config structs/constants/defaults/validation for destination
`state`, `reconciliation`, `takeover`, and `transfer` where no longer used.
- Remove or rewrite tests that only assert legacy state/config behavior.
- Update current-behavior docs:
- `docs/config.md`
- `docs/cli.md`
- `docs/operations.md`
- `docs/troubleshooting.md`
- `docs/integrations/destination-state.md`
- relevant `docs/internal/` docs
- `docs/policy/architecture.md`
- `docs/policy/development.md`
- Update examples to use `workflow` and remove legacy fields.
- Keep roadmap-only material out of current docs.
- Once implemented and documented, remove or rewrite
`docs/roadmap/catalog.md` so completed behavior is not described only as
future work.
Tests and checks:
- `go test ./...`
- `rg -n "state:|reconciliation:|takeover:|transfer:" examples docs --glob '!docs/roadmap/**'`
- `rg -n 'single_owner|shared_root|schema version `2`|schema version `3`' docs internal`
- `rg -n "workflow: additive|workflow: replacement|schema_version.*4" docs examples`
Completion criteria:
- Current docs and examples describe the catalog workflow model.
- Legacy destination policy fields and legacy write paths are gone.
- Full test suite passes.
## Refactors To Avoid
- Do not add top-level `owners` or `sources` catalogs.
- Do not persist workflow in `.distributor.json`.
- Do not keep deprecated config aliases for legacy destination policy fields.
- Do not implement state migration from schema versions `1`, `2`, or `3`.
- Do not add catalog-specific prune/reconcile selectors beyond the agreed
initial ownership scope.
- Do not move catalog policy into storage adapters.
## Open Questions
No open questions are known. The catalog roadmap locks the state shape, workflow
values, default workflow, clean-break policy, legacy-state behavior, output
field selection, `created_at` preservation, and maintenance command scope.