Files
distributor/docs/policy/architecture.md
Eric Rakestraw 8366af6fb6
All checks were successful
ci/woodpecker/tag/release Pipeline was successful
Close completed catalog roadmap
2026-06-19 17:15:24 +00:00

343 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Architecture
This document defines the development principles for `distributor`. It is inward-facing: developers and LLM coding agents should use it to preserve the projects shape, boundaries, and invariants as the code evolves.
## Project Scope
`distributor` is a domain-agnostic report bundle distribution tool.
Producer applications generate manifested bundles. `distributor` discovers those bundles, validates them, optionally derives publication artifacts such as HTML, and publishes selected source and generated artifacts to one or more configured destinations.
`distributor` does not generate domain reports, interpret domain-specific report content, run producer pipelines, edit reports, or act as a CMS. Weather reports, D&D recaps, calendar summaries, email digests, and additional report types should all enter `distributor` through the same bundle contract.
## Project Shape
Default to a small, explicit, dependency-light Go application. Keep the design modular enough to test and change safely, but do not add abstraction unless it protects a real boundary or enables a real extension point.
Business logic should live outside CLI, transport, and external-adapter packages. The core application should reason in terms of pipelines, bundles, destination state, transforms, and publish plans—not S3 SDK calls, SFTP sessions, shell commands, or filesystem details.
The current core workflow is:
1. load configured pipelines;
2. open the source backend;
3. discover source bundles beneath the source root;
4. validate each source bundle and its `manifest.json`;
5. select the source bundle or bundles for each destination according to that destination's path mapping policy;
6. open each destination backend independently;
7. inspect destination state at the resolved destination bundle path;
8. compare source state to destination state;
9. build a publish plan that selects source files, generated files, destination paths, and optional public URLs;
10. optionally transform Markdown to HTML for that destination;
11. publish selected source and generated artifacts;
12. write `.distributor.json` as the destination sentinel/state file;
13. run the notification hook, whose default implementation is currently a no-op.
## Pipeline Model
A pipeline has exactly one source and one or more destinations.
The source is discovered and validated once. Each destination has independent backend configuration, path mapping, publication policy, transform policy, public link policy, replacement behavior, state, and notification behavior.
The pipeline model is fan-out by design:
```text
source bundle
-> destination A: source files only
-> destination B: HTML only
-> destination C: source files + HTML
```
Destination-specific behavior must not leak back into the source bundle contract. A producer should not need to know whether a bundle will be published to local storage, another storage backend, a static site, email, RSS, or another notification channel.
## Source Bundle Contract
A source bundle is a directory containing `manifest.json`.
`manifest.json` is the sole producer-to-`distributor` contract. `distributor` must not rely on producer-specific work directory layouts, filenames, metadata, or conventions outside the configured source root and the source manifest.
The source manifest schema is intentionally minimal:
```json
{
"schema_version": 1,
"id": "weather.daily.brentwood.2026-05-30",
"digest": "sha256:...",
"created": "2026-05-30T11:10:00Z",
"files": [
{
"path": "report.md",
"sha256": "sha256:...",
"size": 12345
}
]
}
```
Required fields:
- `schema_version`: source manifest schema version. Current value: `1`.
- `id`: stable bundle identifier.
- `digest`: SHA-256 digest for the listed files.
- `created`: RFC3339 timestamp. UTC is preferred; explicit offsets are allowed.
- `files`: ordered file list.
- `files[].path`: relative path beneath the bundle root.
- `files[].sha256`: SHA-256 digest of the file contents.
- `files[].size`: file size in bytes.
Bundle paths must be relative, clean, and confined to the bundle root. Paths must not be absolute, empty, contain `..` path traversal, or otherwise escape the bundle root. Symlink handling must be explicit; unless documented otherwise, source bundle symlinks should be rejected.
Before processing a bundle, `distributor` must validate file existence, file size, each file SHA-256, and the bundle digest. Digest mismatch must fail in the MVP before any destination writes occur.
The source manifest should remain minimal. Routing, destination selection, publication format, credentials, static-site layout, notification recipients, and destination-specific metadata belong in `distributor` configuration and destination state, not in producer manifests.
## Destination State Contract
`manifest.json` from the source bundle is not copied to destinations as destination state.
Each destination bundle path is managed by `.distributor.json`. This file is both the destination sentinel and the destination state record.
Publish execution writes catalog destination state. One `.distributor.json` records all managed outputs under the destination bundle path, and each output carries its owning pipeline id and destination id.
Catalog state records:
- `distributor` state schema version;
- state creation and update timestamps;
- catalog state mode;
- owner identity for each managed output;
- compact source identity for each managed output;
- metadata for copied source outputs;
- metadata for generated outputs, such as HTML files;
- optional URL metadata for published outputs;
- any additional metadata required by `distributor`.
A representative destination state file is:
```json
{
"schema_version": 4,
"distributor_version": "0.1.0",
"created_at": "2026-05-30T11:12:00Z",
"updated_at": "2026-05-30T11:12:00Z",
"state": {
"mode": "catalog"
},
"outputs": [
{
"path": "index.html",
"pipeline_id": "weather-daily",
"destination_id": "static-html",
"source": {
"id": "weather.daily.brentwood.2026-05-30",
"digest": "sha256:...",
"created": "2026-05-30T11:10:00Z"
},
"kind": "generated",
"source_path": "report.md",
"transform": "markdown_to_html",
"sha256": "sha256:...",
"size": 23456,
"url": "https://reports.example.com/weather-daily/",
"created_at": "2026-05-30T11:12:00Z",
"updated_at": "2026-05-30T11:12:00Z"
}
]
}
```
Destination comparison rules are based on `.distributor.json`:
- No `.distributor.json`: publish normally only if the destination bundle path is empty.
- Existing catalog state with additive workflow: write planned outputs and retain unrelated managed outputs.
- Existing catalog state with replacement workflow: write planned outputs and remove omitted outputs for the current pipeline and destination owner.
- Planned paths that collide with unmanaged storage content fail by default.
- Invalid or unsupported destination state fails by default.
- Explicit forced replacement may clear the bounded destination bundle path after dry-run review.
## Publication and Transform Policy
Source files are canonical bundle artifacts. Transform outputs are derived publication artifacts.
Transforms are configured per destination. A destination may receive source files, generated HTML files, or both.
The MVP supports only Markdown-to-HTML transformation. HTML generation must not mutate the source bundle. Generated outputs must be deterministic from the source bundle and destination transform configuration, and must be recorded in `.distributor.json`.
Destination path mapping and public link generation are destination behavior. Source manifests do not declare where a bundle is published or which public URLs are recorded.
The application should distinguish:
- transform policy: how derived files are generated;
- publish policy: which source and generated files a destination receives.
For example, one destination may publish source files only as a long-term archive, while another publishes HTML only as a static site.
## Backend Abstraction
Sources and destinations use the same storage abstraction. Current runtime execution uses the local filesystem, SSH/SFTP, and S3-compatible backends. Additional storage backends should be peer implementations behind the same interface, and any backend-specific execution limitation must be documented.
Application logic must interact with storage through internal backend interfaces. Backend-specific behavior belongs in adapter packages. Pipeline, bundle, state, publish, and transform packages must not import service-specific or filesystem adapter implementation details.
Adapters should be thin. Backend adapters should implement storage operations and translate backend-specific errors, but should not make bundle comparison, transform, routing, or replacement decisions.
Remote file copy support should prefer native protocol implementations over shelling out, unless a later design document records a reason to differ.
## Dependency Policy
Prefer the Go standard library where practical.
Use external dependencies only when justified by correctness, security, interoperability, or substantial complexity reduction. Good reasons include YAML parsing, S3-compatible storage integration, SSH/SFTP integration, and Markdown rendering.
Avoid dependencies for small conveniences. Do not let external dependency types leak across internal package boundaries unless the dependency is itself the explicit public contract of that package.
## Package Layout
Use this current layout unless the project has a documented reason to differ:
- `cmd/distributor`: application entrypoint only.
- `pkg/bundle`: public producer-facing source manifest model, digest logic, parsing, manifest building, complete local bundle writing, and local validation helpers.
- `pkg/upload`: public producer-facing HTTP upload client built on `pkg/bundle`.
- `internal/app`: application orchestration and top-level use cases.
- `internal/cli`: CLI command definitions, flags, argument parsing, and command wiring.
- `internal/config`: configuration structs, defaults, loading, precedence, and validation.
- `internal/bundle`: storage-backed source bundle discovery and validation over the public manifest contract.
- `internal/state`: `.distributor.json` catalog parsing, validation, and output metadata.
- `internal/link`: shared HTTP URL validation for configured and persisted link metadata.
- `internal/storage`: backend interfaces, shared path/resource types, backend registry, and storage errors.
- `internal/adapters/local`: local filesystem backend.
- `internal/adapters/ssh`: SSH/SFTP backend.
- `internal/adapters/s3`: S3-compatible object storage backend.
- `internal/transform`: transform interfaces, registry, planning, and shared transform models.
- `internal/transform/markdown`: Markdown-to-HTML implementation.
- `internal/publish`: destination planning, catalog workflow safety checks, and publish execution.
- `internal/notify`: notification interface and MVP no-op notifier.
- `internal/logging`: logging setup and shared logging helpers.
New storage adapters should live under `internal/adapters/<name>` and stay thin.
Package-private implementation constants may live near the package that owns them, preferably in `constants.go` when useful.
## Configuration
Centralize configuration loading, processing, precedence, defaults, and validation in `internal/config`.
The goal is to make configuration discoverable and avoid implicit or hidden operational values. User-visible defaults and cross-package operational defaults should be defined in `internal/config/defaults.go`.
Unless documented otherwise, precedence is:
1. CLI flags
2. environment variables
3. configuration file
4. built-in defaults
Prefer YAML configuration. Config files should be discovered at `/usr/local/etc/distributor/config.yml`, with a CLI override via `--config`.
Configuration files should not contain raw secrets unless the application is explicitly designed for that. Prefer environment variables, secret files, SSH agent usage, standard AWS credential mechanisms, or explicitly named environment variable references for secrets.
Pipeline configuration should express:
- pipeline id;
- one source backend;
- one or more destinations;
- per-destination path mapping;
- per-destination publish policy;
- per-destination transform policy;
- per-destination public link policy;
- validation behavior;
- per-destination workflow and retention behavior.
## Modules and Registries
Each major workflow step should have an explicit input/output contract:
- source discovery;
- source validation;
- destination bundle selection;
- destination state inspection;
- destination comparison;
- transform planning/execution;
- publish planning;
- publish execution;
- notification.
If users can select backends, transforms, notifiers, or renderers, selection should go through a registry or equivalent mechanism rather than scattered conditionals.
The orchestrator should be able to plan, dry-run, and execute configured pipelines. Dry-run behavior should be first-class because the application may delete, overwrite, or publish files.
## Embedded Assets
Store embedded JSON schemas, Markdown templates, HTML templates, CSS, and similar assets as separate files, not inline string literals, unless there is a strong reason otherwise.
## Errors and Logging
Errors should be actionable and preserve context. Wrap errors with operation, pipeline id, destination id, backend, path, bundle id, and resource context where useful. CLI code should convert internal errors into concise user-facing messages.
Errors and logs must not expose secrets.
Use structured logging where practical. Logs should describe discovery, validation, planned actions, skipped copies, conflicts, replacements, external calls, retries, and failure causes, but should not include large report contents by default.
Skip and no-op decisions should be logged at an appropriate level so operators can distinguish successful publication from intentional no-op behavior.
## Context, Timeouts, and Cancellation
Long-running operations should accept `context.Context`. Storage operations, service requests, transforms, and multi-step workflows should respect cancellation and timeouts.
## State, Files, and Safety
If the application writes durable state, writes should be atomic where practical. Multi-step workflows should preserve enough state to support inspection, retry, or resume after failure.
Code that deletes, moves, or overwrites files must use narrow, explicit paths. Avoid broad parent-directory operations. Cleanup that can cause data loss must be opt-in.
`distributor` must never perform broad deletion against a configured source root. Destination deletion must be bounded to the resolved destination bundle path for the current source bundle and backend root.
Normal destructive replacement may occur only when a valid `.distributor.json` confirms that the destination bundle path is distributor-managed. Explicit forced replacement is a per-run CLI workflow for supported conflict and unmanaged-content cases; it must be dry-runnable, clearly reported, and constrained to the destination bundle path.
Replacement must be narrow, reported, test-covered, and configurable. Prefer normal replacement that deletes files recorded in `.distributor.json` and known generated outputs. Forced replacement may delete a bounded destination bundle prefix only when the operator explicitly requests it. Backend implementations must guard against path traversal, prefix confusion, and accidental deletion above the configured backend root.
Where practical, publish operations should use staging paths or temporary objects and promote them into place only after validation and transform steps succeed.
## Testing
Core logic should be testable without real external services. Use fakes, fixtures, or local test doubles for adapters where practical.
Config examples should be load-tested. Important CLI workflows should have parser or command tests. Component contracts should have focused tests that do not require running the full application unless end-to-end coverage is intentional.
Important tests include:
- source manifest parsing and validation;
- file size and SHA-256 validation;
- bundle digest validation;
- source bundle discovery beneath a source root;
- relative path safety and path traversal rejection;
- destination `.distributor.json` parsing and comparison;
- same/older/newer/conflict publish decisions;
- destination bundle path mapping;
- destructive replacement safety checks;
- transform output planning and metadata recording;
- public URL planning and state metadata;
- dry-run output;
- local backend behavior with temporary directories;
- fake backend behavior for storage-facing core logic.
## Documentation
Documentation should follow the project documentation policy. Keep user docs focused on implemented behavior. Put future, planned, or aspirational work only under `docs/roadmap/`.
When changing architecture, config, CLI behavior, adapters, manifest/state contracts, transform behavior, publish behavior, public package/API behavior, or component contracts, update the relevant docs and examples in the same change.
The source manifest and destination `.distributor.json` schemas should have canonical documentation once implemented. Producer-facing package and API workflows belong under `docs/consumers/`. Example configs should be valid and load-tested where practical.
## Non-Goals
`distributor` is not:
- a report generator;
- a domain-specific weather, D&D, calendar, or email summarization app;
- a workflow DAG engine;
- a CMS;
- a web authoring interface;
- a full-text search service;
- a general-purpose file synchronization tool;
- a backup system;
- a notification platform.
Additional notification, feed, template, or transform behavior must preserve the core bundle-distribution boundary.