Update README and architecture documentation to reflect prompt/profile separation

This commit is contained in:
2026-05-05 11:16:33 -05:00
parent ddb4254124
commit 2b9658fb01
2 changed files with 237 additions and 388 deletions

View File

@@ -9,7 +9,6 @@ It accepts named input artifacts, renders prompt templates, calls an LLM, valida
Scriptorium is not an orchestrator. It must not own transcription, transcript merge/polish steps, notifications, or cross-step workflow control.
For the motivating D&D workflow:
- Narratio orchestrates.
- WhisperX transcribes.
- Seriatim merges transcripts.
@@ -20,207 +19,95 @@ Core Go code remains generic.
## 2. Current Architecture
Current implementation structure:
Scriptorium uses a ports-and-adapters architecture to decouple the core execution logic from external dependencies.
- `cmd/scriptorium`: binary entrypoint.
- `internal/domain`: core domain contracts.
- `internal/usecase`: `Runner` run flow, validation integration, bounded repair coordination.
- `internal/profile`: transitional filesystem prompt-definition repository (package rename deferred).
- `internal/artifact`: input artifact resolution (`inline`, `file`).
- `internal/prompt`: template rendering.
- `internal/llm`: provider-neutral client interface + OpenAI-compatible HTTP adapter.
- `internal/validate`: validation implementation (`none/basic/json/json_schema`).
- `internal/adapter/cli`: CLI adapter.
- `internal/adapter/http`: HTTP adapter (`POST /v1/runs`).
### Package Responsibilities
- `cmd/scriptorium`: Binary entrypoint for CLI and HTTP server.
- `internal/domain`: Core domain contracts, including `PromptDefinition`, `ExecutionProfile`, and `RunResult`.
- `internal/usecase`: `Runner` orchestration, including the logic for profile selection, runtime override resolution, and bounded repair.
- `internal/promptdef`: Repository for loading and validating Prompt Definitions from the filesystem.
- `internal/profile`: Repository for loading Execution Profiles from the filesystem.
- `internal/artifact`: Input artifact resolution (`inline`, `file`).
- `internal/prompt`: Template rendering via Go templates.
- `internal/llm`: Provider-neutral client interface and OpenAI-compatible HTTP adapter.
- `internal/validate`: Output validation implementation (`none/basic/json/json_schema`).
- `internal/adapter/cli`: CLI flag parsing and output handling.
- `internal/adapter/http`: HTTP request/response mapping.
## 3. Run Data Flow
`Runner.Run(ctx, RunRequest)` currently executes:
The `Runner.Run` flow executes the following steps:
1. Validate request (`prompt_id` required).
2. Load `PromptDefinition` by ID/version.
3. Determine selected profile ID (`request.profile_id` or prompt `default_profile`).
4. Resolve effective execution target from request override (execution-profile loading is deferred in this phase).
5. Resolve named input artifact refs.
6. Render prompt messages.
7. Hash prompt definition and rendered prompt.
8. Call LLM client with `GenerateRequest`.
9. Build output artifact.
10. Validate output.
11. If structured validation failed and repair is enabled, run bounded repair attempts and re-validate.
12. Return `RunResult` with artifact, raw output, validation result, metadata.
1. **Load Prompt Definition**: Retrieve the `PromptDefinition` by ID from the prompt repository.
2. **Select Profile**: Determine the `profile_id` using the precedence:
- Explicit `profile_id` in `RunRequest`.
- `default_profile` specified in the `PromptDefinition`.
- Error if neither is available.
3. **Load Execution Profile**: Retrieve the `ExecutionProfile` from the profile repository.
4. **Resolve Runtime Overrides**: Merge settings based on precedence (Highest to Lowest):
- Runtime overrides (CLI flags or HTTP `model` object).
- Execution Profile settings.
- Built-in application defaults.
5. **Resolve Artifacts**: Load all named input artifacts defined in the request.
6. **Render Prompt**: Apply template variables and input artifacts to the prompt templates.
7. **Call LLM**: Execute the generation request using the resolved `ExecutionTarget`.
8. **Validate/Repair**:
- Validate the model output against the output contract.
- If structured validation fails and `repair_attempts > 0`, perform bounded repair and re-validate.
9. **Return Result**: Produce a `RunResult` containing the final artifact, metadata, and validation status.
Validation content failures are returned as successful runs with `validation.status=failed`.
## 4. Domain Model
## 4. Package Responsibilities
Key domain types:
- `PromptDefinition`: Defines the "what" (templates, inputs, validation contract, and an optional `default_profile`).
- `ExecutionProfile`: Defines the "how" (endpoint, model, generation parameters, and `api_key_env`).
- `RunRequest`: The intent to execute a prompt, including `prompt_id`, optional `profile_id`, inputs, variables, and optional runtime overrides.
- `RunResult`: The outcome of a run, including the generated `Artifact`, `ValidationResult`, and auditing `RunMetadata`.
- `RunMetadata`: Detailed tracing info: `prompt_id`, `selected_profile_id`, model params, usage tokens, and hashes.
- `domain`
- Owns core nouns/contracts.
- Must not depend on adapters/provider SDK types.
## 5. Interfaces and Adapters
- `usecase`
- Owns single-run orchestration across ports.
- Owns bounded repair control flow.
- Must not own transport/wire concerns.
### Primary Ports
- `promptdef.Repository`: Lookup for prompt definitions.
- `profile.Repository`: Lookup for execution profiles.
- `artifact.Reader`: Loading of artifact content.
- `prompt.Renderer`: Template rendering.
- `llm.Client`: Model generation.
- `validate.Validator`: Output validation.
- `profile` (transitional)
- Currently loads prompt definitions from YAML.
- Package naming split (`prompt definition repo` vs `execution profile repo`) is deferred follow-up.
### Current Adapters
- **Repositories**: Filesystem YAML loaders for both prompts and profiles.
- **Artifact Reader**: Composite reader supporting `file` and `inline`.
- **Prompt Renderer**: Go templates with a custom `input` helper.
- **LLM Client**: OpenAI-compatible `/chat/completions` over HTTP.
- **Validator**: Standard validator supporting `none`, `basic`, `json`, and `json_schema`.
- `artifact`
- Loads artifacts from refs and normalizes payload metadata.
- `prompt`
- Renders templates and enforces required inputs.
- `llm`
- Defines generation client contract and protocol adapters.
- `validate`
- Owns output validation semantics and schema validation.
- `adapter/http`, `adapter/cli`
- Own request/response/flag mapping only.
- Delegate business flow to `usecase.Runner`.
## 5. Domain Model (Current)
Key types:
- `PromptDefinition`
- `id`, `version`, `default_profile`, `inputs`, `templates`, `output_format`, `validation`.
- `ExecutionProfile`
- Execution/runtime settings shape (`endpoint`, `model`, timeouts, `api_key_env`, etc.).
- Loading/persistence is deferred in this pass.
- `ExecutionTarget`
- Effective execution settings for a run.
- `RunRequest`
- `prompt_id`, `prompt_version`, optional `profile_id`, `inputs`, `vars`, optional `execution` override, optional validation override.
- `RunResult`
- Output artifact, validation, raw output, prompt/profile/model metadata, hashes, timing, usage.
- `ArtifactRef` / `Artifact`
- Input reference and loaded content contracts.
- `RenderedPrompt` / `RenderedMessage`
- Provider-neutral rendered prompt.
- `GenerateRequest` / `GenerateResponse`
- Provider-neutral model I/O.
## 6. Interfaces and Adapters
Primary ports:
- `profile.Repository` (transitional prompt-definition lookup)
- `artifact.Reader`
- `prompt.Renderer`
- `llm.Client`
- `validate.Validator`
- `usecase.OutputRepairer` (usecase-local)
Current adapters:
- Prompt definition repository: filesystem YAML loader.
- Artifact readers: `file`, `inline` via composite reader.
- Prompt renderer: Go templates with `input` helper.
- LLM adapter: OpenAI-compatible `/chat/completions` over `net/http`.
- Validator: standard validator (`none/basic/json/json_schema`).
- CLI/HTTP adapters.
## 7. Validation and Repair Model
Validation modes:
- `none`
- `basic`
- `json`
- `json_schema`
Repair behavior:
- Applies only to structured modes (`json`, `json_schema`).
- Triggered only on failed validation and only when `repair_attempts > 0`.
- Strictly bounded by `repair_attempts`.
- Uses a narrow repair prompt asking for corrected JSON only.
- Runtime validator/repair errors are run errors.
## 8. Public Contracts
## 6. Public Contracts
### CLI
- `run`: Executes a prompt. Uses flags like `--prompt`, `--profile`, `--input`, and various runtime overrides (e.g., `--model`, `--temperature`).
- `serve`: Starts the HTTP API.
Commands:
### HTTP API
- `POST /v1/runs`: Accepts `RunRequest` JSON and returns `RunResponse` JSON. No built-in auth.
- `scriptorium run`
- `scriptorium serve`
### YAML Shapes
- **Prompt YAML**: Includes `id`, `version`, `default_profile`, `inputs`, `templates`, and `validation`.
- **Profile YAML**: Includes `id`, `endpoint`, `model`, generation params, and `api_key_env`.
`run` flags:
## 7. Guardrails
- Required: `--profile-dir`, `--prompt-id`, `--input`.
- Optional: `--profile-id`, `--var`, `--out`, `--llm-base-url`, `--model`, `--api-key-env`, `--temperature`, `--max-tokens`, `--schema-dir`, `--timeout`.
- **Separation of Concerns**: Prompt content must not belong in execution profiles; model/API settings must not belong in prompt definitions.
- **Security**: Raw API keys are unsupported in all configuration and transport layers. Only `api_key_env` is used.
- **Path Resolution**: `content_file` paths in prompt definitions resolve relative to the prompt YAML file.
- **Integrity**: No silent prompt truncation or omission of content.
- **Reliability**: Repair loops are strictly bounded by `repair_attempts`.
Current transitional runtime behavior:
## 8. Extension Points
- Prompt definitions may provide `default_profile` selection.
- Execution-profile loading is deferred; execution settings must currently be supplied via run-time overrides.
### HTTP
- Endpoint: `POST /v1/runs`.
- Request maps to `RunRequest` with `prompt_id` (required), `inputs`, optional `profile_id`, `vars`, optional execution override (`model` object).
- Response includes `artifact`, `validation`, `metadata`, `raw_model_output`.
- Validation content failures return `200` with failed validation status.
- Error response shape: `{ "error": { "code": "...", "message": "..." } }`.
### Prompt Definition YAML
Current prompt-definition fields:
- `id`, `version`, optional `default_profile`, optional `description`
- `inputs[]` with `name`, `required`, optional `content_type`, optional `description`
- `templates[]` with `role` and either `content` or `content_file`
- `output_format`
- `validation` (`format`, `validation_mode`, `schema_path`, `repair_attempts`)
Strict YAML decoding (`KnownFields`) is enabled.
### API Key Policy
- Raw API keys are not accepted in YAML, CLI flags, HTTP body, or domain metadata.
- Auth is configured only by env var reference (`api_key_env`), resolved at request time by the LLM adapter.
## 9. Extension Points (Future Work)
Planned next extensions should reuse current boundaries:
- Execution-profile repository/loader implementation.
- Split transitional `internal/profile` into clearer prompt-definition/profile repositories.
- S3 artifact refs.
- Token budgeting/policy layer.
- Streaming generation.
- Batch run use case.
- Additional provider adapters.
- Additional validation modes.
## 10. Architectural Guardrails
- No D&D-specific logic in core Go packages.
- No orchestration creep into Scriptorium.
- No unbounded repair loops.
- No silent content truncation/omission.
- Do not log full prompts/artifacts by default.
- Keep provider-specific wire/SDK details out of domain types.
- Keep adapters thin.
## 11. Testing Strategy
Protect these behaviors with focused tests:
- Prompt-definition loading/validation errors.
- Artifact loading/hash/content-type behavior.
- Prompt rendering required-input and template error paths.
- LLM adapter request/response/auth/error/timeout behavior.
- Runner success/failure/metadata behavior.
- Validation failure raw-output preservation.
- Bounded repair behavior.
- HTTP mapping and error mapping.
- CLI parsing and output stream separation.
Prefer small unit tests and minimal integration-style tests with fake LLMs.
Future work should remain grounded in the current architecture:
- **Artifacts**: Add S3 artifact references via a new `artifact.Reader`.
- **LLM**: Implement additional provider adapters (e.g., Anthropic, Google).
- **Execution**: Add token budgeting, streaming generation, and batch execution capabilities.
- **Repositories**: Implement database-backed repositories for prompts and profiles.
- **Profiles**: Support more granular profile versioning and environment-specific profiles.