Files
scriptorium/architecture.md

8.0 KiB

Scriptorium Architecture

1. Purpose and Non-Goals

Scriptorium is a prompt-profile execution engine.

It takes named input artifacts, renders prompt templates, calls an LLM, validates output, optionally performs bounded structured-output repair, and returns an artifact with metadata.

Scriptorium is not an orchestrator. It should not own transcription, transcript merging, transcript polishing, notification, or cross-step workflow control.

For the motivating D&D workflow:

  • Narratio orchestrates.
  • WhisperX transcribes.
  • Seriatim merges transcripts.
  • Audita polishes transcripts.
  • Scriptorium generates output artifacts from prepared inputs.

Core Go code must remain domain-generic.

2. Current Architecture

Current high-level structure:

  • cmd/scriptorium: binary entrypoint.
  • internal/domain: core domain types.
  • internal/usecase: Runner use case and repair loop orchestration.
  • internal/profile: filesystem prompt profile repository and profile validation.
  • internal/artifact: artifact reference readers (inline, file) and routing.
  • internal/prompt: Go template-based prompt renderer.
  • internal/llm: LLM client interface + OpenAI-compatible HTTP adapter.
  • internal/validate: output validator implementation (none/basic/json/json_schema).
  • internal/adapter/cli: CLI adapter.
  • internal/adapter/http: HTTP adapter (POST /v1/runs).

This is a practical ports-and-adapters implementation.

3. Run Data Flow

Runner.Run(ctx, RunRequest) currently executes:

  1. Validate minimum request requirements (profile_id).
  2. Load prompt profile by ID/version.
  3. Merge effective model target (profile defaults + request override).
  4. Resolve effective output contract (profile + optional request override).
  5. Resolve named artifact refs to loaded artifacts.
  6. Render prompt messages from templates.
  7. Hash rendered prompt for auditability.
  8. Call LLM client with provider-neutral GenerateRequest.
  9. Build output artifact from model content.
  10. Validate output.
  11. If structured validation failed and repair is enabled/bounded, run repair attempts and re-validate.
  12. Return RunResult with artifact, validation, raw output, and metadata.

Validation content failure remains a successful run result with validation.status=failed.

4. Package Responsibilities

  • domain

    • Owns core nouns and contracts.
    • Must not import adapters/provider SDK types.
  • usecase

    • Owns execution sequence and cross-port orchestration for a single run.
    • May coordinate validation and bounded repair.
    • Must not contain HTTP/CLI/wire concerns.
  • profile

    • Owns prompt profile loading/parsing/validation.
    • Handles YAML strict decoding and profile-level constraints.
  • artifact

    • Owns artifact ref resolution and content loading.
    • Produces normalized Artifact values with size/hash/content type.
  • prompt

    • Owns template rendering and required-input enforcement.
  • llm

    • Owns generation port and provider adapters.
    • Current adapter: OpenAI-compatible chat completions over net/http.
  • validate

    • Owns output validation semantics and JSON Schema integration.
  • adapter/http, adapter/cli

    • Owns transport/wire/flag concerns only.
    • Should stay thin and delegate business flow to usecase.Runner.

5. Domain Model (Current)

Key types in internal/domain:

  • RunRequest: profile selector, named input refs, vars, optional model override, optional validation override.
  • RunResult: output artifact, validation result, raw output, profile/model metadata, hashes, usage, timestamps, duration.
  • ArtifactRef: {type, uri, body} reference contract.
  • Artifact: loaded payload (name, content_type, body, uri, size, hash).
  • PromptProfile: YAML-backed profile definition.
  • RenderedPrompt / RenderedMessage: provider-neutral prompt structure.
  • GenerateRequest / GenerateResponse: provider-neutral model I/O.
  • ValidationResult: passed/failed/skipped + mode/errors/schema/repair attempts.

6. Interfaces and Adapters

Primary ports:

  • profile.Repository
  • artifact.Reader
  • prompt.Renderer
  • llm.Client
  • validate.Validator
  • usecase.OutputRepairer (usecase-local abstraction)

Current adapters:

  • Profile repository: filesystem YAML loader.
  • Artifact reader: composite reader for inline and file.
  • Prompt renderer: Go templates with input helper + vars.
  • LLM adapter: OpenAI-compatible /chat/completions.
  • Validator: standard validator with none/basic/json/json_schema.
  • CLI/HTTP adapters: thin request mapping and response mapping.

7. Validation and Repair Model

Validation modes:

  • none
  • basic
  • json
  • json_schema

Repair behavior:

  • Only applies to structured modes (json, json_schema).
  • Attempted only when validation fails, repairer exists, and repair_attempts > 0.
  • Bounded strictly by repair_attempts.
  • Uses a narrow JSON-repair prompt and re-validates each attempt.
  • If still invalid, run succeeds with failed validation and preserved final raw output.
  • Validator runtime/config errors are run errors.

8. Public Contracts

CLI

Commands:

  • scriptorium run
  • scriptorium serve

run:

  • Required: --profile-dir, --profile-id, --input.
  • Optional: model/endpoint overrides (--model, --llm-base-url), vars, output path, schema dir, timeout.
  • Artifact bytes go to stdout (or --out file); summaries/errors go to stderr.

serve:

  • Required: --profile-dir, --llm-base-url.
  • Exposes HTTP run endpoint.

HTTP

  • Endpoint: POST /v1/runs.
  • Request maps to RunRequest (profile_id, inputs, vars, optional model override).
  • Response includes artifact, validation, metadata, raw_model_output.
  • Validation content failures are represented as 200 with validation.status=failed.
  • Error responses are {error:{code,message}} with stable code mapping.

Prompt Profile YAML

  • id, version, expected_inputs, templates, model_defaults, output_format, validation.
  • Strict YAML decoding (KnownFields) rejects unknown fields.
  • validation.schema_path required when validation_mode=json_schema.
  • validation.repair_attempts must be non-negative.

Metadata

Current run metadata includes:

  • run_id (UUID v4)
  • profile_id, profile_version, profile_hash
  • model_name, endpoint, effective model_params
  • input_hashes, prompt_hash
  • token usage
  • start/end timestamps
  • duration
  • validation mode/status
  • repair attempts used

9. Extension Points (Future Work)

Future features should plug into existing boundaries, not bypass them.

Candidate extensions:

  • S3 artifact refs via artifact.Reader extension.
  • Token budgeting in usecase/model-target policy layer.
  • Streaming LLM output via additional llm.Client methods/adapters.
  • Batch execution as a separate use case (not hidden in single-run path).
  • Additional LLM providers implementing llm.Client.
  • Additional validators/modes in validate.
  • Additional profile repositories (embedded, remote, object storage).

These are future work, not part of current default behavior.

10. Architectural Guardrails

Contributors should preserve these constraints:

  • No D&D-specific behavior in core Go packages.
  • No orchestration creep into Scriptorium.
  • No unbounded repair loops.
  • No silent truncation/omission of rendered inputs or outputs.
  • Do not log full artifacts/prompts by default.
  • Keep provider-specific wire/SDK details out of domain types.
  • Keep adapter boundaries explicit and thin.

11. Testing Strategy

Protect behavior at boundaries and in usecase flow:

  • Profile loading/parsing/validation errors.
  • Artifact reading for inline/file + hash/content type behavior.
  • Prompt rendering required inputs/template error behavior.
  • LLM adapter request/response/error/timeout behavior.
  • Runner success path and metadata population.
  • Validation failure raw-output preservation.
  • Successful/failed/bounded repair flows.
  • HTTP request mapping, response shape, and error mapping.
  • CLI parsing helpers, required flags, and output stream separation.

Prefer focused unit tests and small integration-style tests with fake LLMs.