Update README and architecture documentation to reflect prompt/profile separation

This commit is contained in:
2026-05-05 11:16:33 -05:00
parent ddb4254124
commit 2b9658fb01
2 changed files with 237 additions and 388 deletions

366
README.md
View File

@@ -1,134 +1,159 @@
# scriptorium
Scriptorium is a generic prompt-definition execution engine written in Go.
Scriptorium is a generic prompt execution engine.
Given named input artifacts and a prompt definition, Scriptorium:
It takes:
- a prompt definition
- a selected or default execution profile
- named input artifacts
- template variables
- optional runtime overrides
1. Loads the prompt definition.
2. Resolves input artifact references.
3. Renders prompt messages from templates.
4. Calls an OpenAI-compatible LLM endpoint.
5. Validates output if configured.
6. Optionally performs bounded structured-output repair.
7. Returns a generated artifact plus run metadata.
It returns:
- generated artifact
- validation result
- metadata
## Where Scriptorium Fits
## Prompt vs Profile
Scriptorium is not an orchestrator.
Scriptorium separates **what** to do (Prompt) from **how** to do it (Profile).
In the D&D workflow:
### Prompt Definition
Defines the task logic and output contract.
- Task description and version.
- Message templates (system, user, etc.).
- Required and optional input artifacts.
- Output format and validation rules.
- Repair settings for structured output.
- Optional `default_profile` for convenience.
- Narratio orchestrates the full pipeline.
- WhisperX transcribes audio.
- Seriatim merges transcripts.
- Audita polishes transcripts.
- Scriptorium generates final artifacts from prepared inputs.
### Execution Profile
Defines the runtime environment and model settings.
- LLM endpoint (URL).
- Model name.
- Generation parameters: `temperature`, `max_tokens`, `top_p`.
- Runtime settings: `timeout`, `reasoning_effort`.
- API key source via `api_key_env`.
D&D-specific behavior belongs in profiles, schemas, fixtures, and caller inputs, not in core Go logic.
Callers can explicitly provide a `profile_id` to override the prompt's `default_profile`.
## Core Concepts
## Precedence
- Prompt definition: YAML config for templates, inputs, output format, and validation behavior.
- Execution profile: conceptual runtime config (endpoint/model/timeouts/auth source). In this transition, execution settings are supplied as run-time overrides.
- Named inputs: logical names (for example `transcript`, `glossary`) mapped to artifact references.
- Artifact refs: currently `file` and `inline` are supported.
- Template variables: key/value vars passed at run time and referenced as `{{.var_name}}`.
- Execution target: endpoint/model plus generation/runtime parameters (`temperature`, `max_tokens`, `top_p`, `timeout_seconds`, `reasoning_effort`, `api_key_env`).
- Output format: `text`, `markdown`, or `json`.
- Validation mode: `none`, `basic`, `json`, `json_schema`.
- Repair attempts: bounded retries for structured modes (`json`, `json_schema`) when output validation fails.
When resolving runtime settings, Scriptorium follows this precedence model (highest to lowest):
## Build and Test
1. **Runtime Overrides**: Provided via CLI flags or HTTP request `model` object.
2. **Execution Profile**: Settings defined in the selected profile.
3. **Application Defaults**: Built-in fallback values.
```bash
go build -o scriptorium ./cmd/scriptorium
go test ./...
```
### Profile Selection Logic
The engine determines which profile to use in this order:
1. Explicit `profile_id` (via `--profile` or HTTP request).
2. The `default_profile` named in the Prompt Definition.
3. Error: If neither is provided and no default exists.
## API Key Policy
To ensure security, Scriptorium does not support raw API keys in configuration files, CLI arguments, or HTTP requests.
- **`api_key_env`**: Profiles and overrides specify the name of an environment variable (e.g., `SCRIPTORIUM_API_KEY`).
- **Runtime Resolution**: The value of the environment variable is read directly from the process environment at runtime.
- **Zero Leakage**: API key values are never included in metadata, logs, or response bodies.
## CLI Usage
### `scriptorium run`
Required flags:
Runs a single prompt execution.
- `--prompt-dir`
- `--profile-dir`
- `--prompt`
- `--input` (repeatable `name=path`)
**Required Flags:**
- `--prompt-dir`: Directory containing prompt YAML files.
- `--profile-dir`: Directory containing profile YAML files.
- `--prompt`: The prompt ID to execute.
- `--input`: Input mapping `name=path` (repeatable).
Common optional flags:
**Optional Flags:**
- `--profile`: Override the prompt's default profile.
- `--var`: Template variable `name=value` (repeatable).
- `--out`: Write output to a file instead of stdout.
- `--llm-base-url`: Override endpoint.
- `--model`: Override model name.
- `--api-key-env`: Override API key environment variable name.
- `--temperature`: Override temperature.
- `--max-tokens`: Override max tokens.
- `--top-p`: Override top_p.
- `--timeout`: Override request timeout (e.g., `30s`, `1m`).
- `--profile` (execution profile selector; falls back to prompt `default_profile`)
- `--var` (repeatable `name=value`)
- `--out`
- `--llm-base-url`
- `--model`
- `--api-key-env`
- `--temperature`
- `--max-tokens`
- `--schema-dir`
- `--timeout`
Current transitional behavior: execution-profile loading is not implemented yet, so run-time execution settings must be supplied via overrides. In practice, provide at least endpoint and model (`--llm-base-url` and `--model`).
Example:
**Examples:**
Using the prompt's `default_profile`:
```bash
export SCRIPTORIUM_API_KEY="your-key"
go run ./cmd/scriptorium run \
export SCRIPTORIUM_API_KEY="sk-..."
scriptorium run \
--prompt-dir ./prompts \
--profile-dir ./profiles \
--prompt generic.markdown_summary \
--profile local-fast \
--input transcript=./examples/fixtures/transcript.md \
--input glossary=./examples/fixtures/glossary.yml \
--llm-base-url http://localhost:8000/v1 \
--model gpt-4o-mini \
--api-key-env SCRIPTORIUM_API_KEY \
--out ./out.md
--input transcript=./examples/fixtures/transcript.md
```
Output behavior:
Overriding the profile:
```bash
scriptorium run \
--prompt-dir ./prompts \
--profile-dir ./profiles \
--prompt generic.markdown_summary \
--profile local-quality \
--input transcript=./examples/fixtures/transcript.md
```
- Artifact content goes to stdout unless `--out` is set.
- Summaries and errors are written to stderr.
- Exit code `2` means the run succeeded but validation status is `failed`.
Overriding model and runtime values:
```bash
scriptorium run \
--prompt-dir ./prompts \
--profile-dir ./profiles \
--prompt generic.markdown_summary \
--model gpt-4o \
--temperature 0.7 \
--input transcript=./examples/fixtures/transcript.md
```
Using a local OpenAI-compatible vLLM endpoint:
```bash
scriptorium run \
--prompt-dir ./prompts \
--profile-dir ./profiles \
--prompt generic.markdown_summary \
--llm-base-url http://localhost:8000/v1 \
--model meta-llama-3-8b \
--input transcript=./examples/fixtures/transcript.md
```
### `scriptorium serve`
Starts HTTP API.
Starts the HTTP API.
Required flags:
**Required Flags:**
- `--prompt-dir`: Directory containing prompt YAML files.
- `--profile-dir`: Directory containing profile YAML files.
- `--prompt-dir`
- `--profile-dir`
Common optional flags:
- `--addr` (default `:8080`)
- `--schema-dir` (default `.`)
- `--model`
- `--timeout` (default `10m`)
**Optional Flags:**
- `--addr`: Listen address (default `:8080`).
- `--schema-dir`: Base directory for validation schemas.
- `--model`: Default model override.
- `--timeout`: Default request timeout.
## HTTP API
Endpoint:
### `POST /v1/runs`
- `POST /v1/runs`
No built-in authentication is provided by the server itself. Deploy behind a trusted boundary or gateway.
Request example:
Executes a prompt. No built-in authentication is provided; deploy behind a trusted gateway.
**Request Body:**
```json
{
"prompt_id": "generic.structured_events",
"prompt_version": "1.0.0",
"profile_id": "local-default",
"profile_id": "local-quality",
"inputs": {
"transcript": {"type": "file", "uri": "./examples/fixtures/transcript.md"},
"glossary": {"type": "file", "uri": "./examples/fixtures/glossary.yml"}
"transcript": {"type": "file", "uri": "./examples/fixtures/transcript.md"}
},
"vars": {
"session_date": "2026-05-04"
@@ -136,149 +161,86 @@ Request example:
"model": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0.0,
"max_tokens": 600,
"top_p": 1.0,
"timeout_seconds": 120,
"api_key_env": "SCRIPTORIUM_API_KEY"
"temperature": 0.0
}
}
```
Response shape:
**Response:**
Returns a `200 OK` with the generated artifact, validation results, and metadata including the `prompt_id` and the `selected_profile_id`.
```json
{
"artifact": {
"name": "output",
"content_type": "application/json",
"body": "{...}",
"uri": "",
"size": 123,
"hash": "..."
},
"validation": {
"status": "passed",
"mode": "json_schema",
"errors": [],
"schema_path": "structured_events.schema.json",
"repair_attempts": 0,
"is_valid": true
},
"metadata": {
"run_id": "xxxxxxxx-xxxx-4xxx-8xxx-xxxxxxxxxxxx",
"prompt_id": "generic.structured_events",
"prompt_version": "1.0.0",
"prompt_hash": "...",
"rendered_prompt_hash": "...",
"selected_profile_id": "local-default",
"model_name": "gpt-4o-mini",
"endpoint": "http://localhost:8000/v1",
"model_params": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0,
"max_tokens": 600,
"top_p": 1,
"timeout_seconds": 120,
"api_key_env": "SCRIPTORIUM_API_KEY"
},
"input_hashes": {"transcript": "...", "glossary": "..."},
"usage": {"prompt_tokens": 10, "completion_tokens": 20, "total_tokens": 30},
"start_time": "...",
"end_time": "...",
"duration_ms": 1523,
"validation_mode": "json_schema",
"validation_status": "passed",
"repair_attempts_used": 0
},
"raw_model_output": "{...}"
}
```
Validation content failures return `200` with `validation.status = "failed"` and preserve `raw_model_output`.
Error response shape:
```json
{
"error": {
"code": "artifact_read_failed",
"message": "failed to read input artifact"
}
}
```
**Validation Failures:**
If the model output fails validation (e.g., invalid JSON), the API returns `200 OK` with `validation.status = "failed"`. The original `raw_model_output` is preserved in the response to allow debugging.
## Prompt Definition Authoring
### Minimal Markdown prompt definition
```yaml
id: generic.markdown_summary
version: "1.0.0"
default_profile: local-default
inputs:
- name: transcript
required: true
templates:
- role: system
content: "You are a concise assistant."
- role: user
content: |
Summarize:
{{input "transcript"}}
output_format: markdown
validation:
validation_mode: basic
```
### Structured JSON prompt definition with schema validation
Prompts are defined in YAML.
### Canonical Shape
```yaml
id: generic.structured_events
version: "1.0.0"
default_profile: local-default
description: "Extracts structured events from a transcript"
default_profile: local-quality
inputs:
- name: transcript
required: true
description: "The raw session transcript"
- name: glossary
required: false
templates:
- role: system
content: "Return only JSON."
content: "You are a helpful assistant."
- role: user
content: |
Extract events from:
{{input "transcript"}}
content_file: messages/extract_events.tmpl
output_format: json
validation:
format: json
validation_mode: json_schema
schema_path: structured_events.schema.json
repair_attempts: 1
repair_attempts: 2
```
`repair_attempts` is strictly bounded and only applies to structured validation modes.
**Key Features:**
- **Inline vs File**: Use `content` for short prompts or `content_file` for larger templates.
- **Inputs**: Mark inputs as `required` to ensure the runner fails early if they are missing.
- **Validation**: Support `none`, `basic`, `json`, and `json_schema`.
- **Repair**: `repair_attempts` enables bounded retries to fix structured output.
## Validation Modes
## Execution Profile Authoring
- `none`: skipped validation result.
- `basic`: fails for empty/whitespace output.
- `json`: output must parse as JSON.
- `json_schema`: output must parse as JSON and satisfy configured schema.
Profiles are defined in YAML.
Validation content failures are returned in the structured result; raw model output is preserved.
### Canonical Shape
```yaml
id: local-quality
endpoint: http://localhost:8000/v1
model: gpt-4o
temperature: 0.0
max_tokens: 4096
top_p: 1.0
timeout_seconds: 300
reasoning_effort: high
api_key_env: SCRIPTORIUM_API_KEY
```
**Constraints:**
- **No Raw Keys**: Do not include actual API keys. Only specify the environment variable name in `api_key_env`.
- **Local Profiles**: For local endpoints that don't require auth, `api_key_env` can be omitted.
## Examples
- Prompt definitions: `prompts/`
- Execution profiles: `profiles/`
- Schemas: `schemas/`
- Fixtures: `examples/fixtures/`
- Local experimentation: `local-test/`
- **Prompt Definitions**: `prompts/`
- **Execution Profiles**: `profiles/`
- **Schemas**: `schemas/`
- **Fixtures**: `examples/fixtures/`
- **Local Experimentation**: `local-test/`
## Development Notes
## Build and Test
- Core follows ports-and-adapters and remains domain-generic.
- Domain/usecase packages do not depend on HTTP/CLI wire DTOs.
- To add a new LLM adapter: implement `internal/llm.Client`.
- To add a new artifact reader: extend `internal/artifact.Reader` routing.
- To add a new validation mode: extend `internal/validate` and preserve run semantics.
```bash
go build -o scriptorium ./cmd/scriptorium
go test ./...
```