Files
distributor/docs/internal/state.md

4.6 KiB

Destination State Internals

Audience: developers and LLM coding agents changing internal/state.

Purpose

internal/state parses, validates, serializes, and updates .distributor.json destination catalog state records.

Inputs And Outputs

Inputs are destination state JSON, constructed catalog values, owner scopes, managed output paths, timestamps, and prune policy inputs. Outputs are validated catalog values, JSON bytes, managed path lists, owner-filtered output lists, missing-output repair results, and prune candidate plans.

Boundaries

The package does not inspect storage backends, mutate files, choose workflow actions, build publish outputs, generate URLs, or parse config. Publish planning consumes parsed catalog state and catalog output helpers.

The external destination state contract is documented in docs/integrations/destination-state.md.

Config Fields Used

internal/state uses shared constants for catalog mode, output kinds, slug-like id validation, link validation, storage path validation, and source manifest validation. Destination ids, pipeline ids, and link URLs originate from config but are supplied as values by callers.

Adapters Used

None.

State And Manifest Behavior

Current .distributor.json publish output uses schema version 4 catalog state. Required top-level fields are schema_version, created_at, updated_at, state.mode, and outputs; distributor_version is optional.

Each catalog output record requires a clean path, pipeline id, destination id, source identity, source or generated kind, lowercase SHA-256 digest, non-negative size, and created/updated timestamps. Generated outputs require source_path and transform; copied source outputs must omit both. Stored URLs are optional and must pass internal/link validation.

Embedded source identity records contain source manifest id, digest, and creation timestamp. Full source manifests are not embedded in catalog state.

The package identifies schema versions older than the current catalog schema as superseded legacy state for publish planning. It rejects invalid JSON, malformed catalog state, and unsupported future schema versions.

The package provides helpers for finding catalog outputs by path, filtering outputs by owner, listing managed output paths, removing missing output records for one owner or every owner, and building owner-scoped prune candidates.

Publish execution owns catalog output projection and timestamp preservation for rewritten outputs. State helpers only parse, validate, filter, and remove catalog records supplied by callers.

Skip And Resume Behavior

Catalog parsing and helper transformations are pure. State code does not decide whether to skip, upsert, replace, force, or fail; publish planning maps parsed state and storage observations to actions.

Missing-output removal helpers remove matching output records only and leave storage inspection, timestamp updates, validation, and state rewrites to callers.

Prune planning helpers are pure. They select managed output candidates, sort deterministically by updated_at and path, preserve the newest keep_latest candidates before evaluating older_than, and return planned prune/preserve lists without mutating state. App-level prune execution uses missing-output removal helpers to remove only confirmed deleted records after storage deletion succeeds.

Failure Behavior

Parsing rejects invalid JSON, trailing data, missing required fields, invalid timestamps, invalid catalog mode, duplicate outputs, invalid output paths, unsupported output kinds, missing generated transform metadata, invalid URLs, invalid digests, and negative sizes.

Tests To Inspect

  • internal/state/catalog_test.go
  • internal/state/prune_test.go
  • internal/app/reconcile_state_test.go
  • internal/cli/reconcile_state_test.go
  • internal/publish/*_test.go

Architectural Invariants

  • .distributor.json is the destination sentinel and state record.
  • State helpers do not inspect or mutate storage.
  • Source identity uses the source bundle contract.
  • Newly written publish state uses schema version 4.
  • Superseded legacy schema handling is limited to identifying older state for publish planning.
  • Missing-output repair helpers preserve unrelated owner records and outputs.
  • Prune planning uses output updated_at and preserves unrelated owners.
  • Generated outputs always record a transform id and source path.
  • Output records always carry created and updated timestamps after parsing.
  • Stored URLs are optional and must be absolute HTTP or HTTPS URLs when present.
  • distributor_version is diagnostic metadata, not a comparison key.