Compare commits

6 Commits

21 changed files with 1149 additions and 2208 deletions

View File

@@ -22,6 +22,7 @@ go run ./cmd/scriptorium render \
```
This command renders the prepared prompt and effective runtime settings without calling an LLM.
For complete invocation and output behavior, see the [CLI reference](docs/cli.md).
## Documentation
@@ -29,7 +30,6 @@ This command renders the prepared prompt and effective runtime settings without
- [Configuration reference](docs/config.md)
- [HTTP API reference](docs/api.md)
- [Operations guide](docs/operations.md)
- [Troubleshooting](docs/troubleshooting.md)
- [Consumer integration overview](docs/consumers/api.md)
- [Go library package](docs/consumers/pkg-scriptorium.md)
- [Subprocess integration](docs/integrations/subprocess.md)
@@ -38,8 +38,8 @@ This command renders the prepared prompt and effective runtime settings without
## Examples
- `examples/config.yml`
- `examples/config.full.yml`
- `examples/render-markdown-summary.sh`
- `examples/http-run.json`
- `examples/go-library/prepare`
- [Minimal configuration](examples/config.yml) and [complete configuration](examples/config.full.yml)
- [Prompt definitions](examples/prompts/), [execution profiles](examples/profiles/), [schemas](examples/schemas/), and [synthetic input fixtures](examples/fixtures/)
- [Render script](examples/render-markdown-summary.sh)
- [HTTP request](examples/http-run.json)
- [Go library example](examples/go-library/prepare/main.go)

View File

@@ -0,0 +1,49 @@
# ADR 0001: Adopt Canonical Documentation Ownership
## Status
Accepted
## Date
2026-07-26
## Context
Scriptorium's documentation grew alongside its CLI, HTTP, public Go, and
integration interfaces. As a result, several documents repeated mutable
contracts such as flags, configuration fields, and status behavior. Those
parallel definitions made it unclear which document to update when behavior
changed and increased the risk of documentation drift.
## Decision
Assign each documentation topic one canonical owner, as defined in
[`docs/policy/documentation.md`](../policy/documentation.md). Non-owning
documents may provide short orientation and links, but do not redefine volatile
contracts. Current behavior is documented outside `docs/roadmap/`; roadmaps own
future work, sequencing, and implementation status.
## Alternatives Considered
- Keep broad reference material in several audience-specific documents. This
would preserve local convenience but leave conflicting contract definitions
likely.
- Consolidate all documentation into one reference. This would reduce duplicate
text but would not serve the distinct needs of users, operators, consumers,
and contributors.
## Rationale
Canonical ownership retains audience-specific guidance while making the source
of truth for each contract discoverable. It also makes documentation changes
reviewable alongside the implementation change that requires them.
## Consequences
- Changes to behavior must update the canonical owner in the same change.
- Cross-cutting documentation links to the owner instead of copying its
details.
- Documentation restructuring follows the implementation sequence in
[`docs/roadmap/documentation.md`](../roadmap/documentation.md); the roadmap,
not this ADR, records completion status.

View File

@@ -2,296 +2,142 @@
This is the canonical public HTTP contract for Scriptorium.
Implemented route:
## Service And Route
- `POST /v1/runs`
`POST /v1/runs` runs one prompt request and returns generated output,
validation, and metadata. The service has no built-in authentication or
authorization; deploy it behind appropriate network and authentication controls.
For CLI behavior, see [CLI reference](cli.md). For config and prompt/profile
file formats, see [Configuration reference](config.md).
The service address and HTTP limits are configured as described in the
[configuration reference](config.md). `serve` invocation is defined in the
[CLI reference](cli.md).
The maintained request-shape example is `examples/http-run.json`. It requires a
running `serve` process with an artifact root that can read the referenced
files, plus a reachable model endpoint for full execution.
## Base URL And Deployment
`scriptorium serve` listens on `server.addr` or `serve --addr`. The default is
`:8080`.
The route path is always:
```text
/v1/runs
```
The HTTP adapter has no built-in authentication or authorization. Deploy it
behind trusted network and authentication controls.
## Media Types
- Request body: JSON object.
- Response body: JSON object.
- Response `Content-Type`: `application/json`.
Requests are decoded as JSON regardless of the request `Content-Type` header.
There are no shared query parameters.
Requests and responses are JSON objects. Requests are decoded as JSON regardless
of their `Content-Type`; successful JSON responses use
`Content-Type: application/json`. There are no query parameters.
## Request Limits
HTTP limits are configured through `server.*` config fields or `serve` flags:
The configured request-body limit includes inline artifact bodies. The artifact
limit applies to HTTP `file` inputs. The response limit applies to the encoded
response, including the artifact body and optional raw output. A limit of zero
disables that limit.
- `server.max_request_bytes`: encoded JSON request body limit, including inline input bodies.
- `server.max_artifact_bytes`: file artifact limit for HTTP `file` input references.
- `server.max_response_bytes`: encoded JSON response limit, including artifact body and optional raw output.
Each limit defaults to `16777216` bytes. `0` disables that limit.
A request body over its limit returns `413 request_too_large`; an oversized
file input returns `413 artifact_too_large`; an oversized encoded response
returns `413 response_too_large`.
## `POST /v1/runs`
Runs one prompt request and returns the generated artifact, validation result,
and metadata.
### Request Body
The maintained [request example](../examples/http-run.json) is a complete
copyable shape. The smallest valid shape is:
```json
{
"prompt_id": "generic.markdown_summary",
"profile_id": "local-fast",
"prompt_version": "1.0.0",
"inputs": {
"transcript": {
"type": "file",
"uri": "./examples/fixtures/transcript.md"
},
"glossary": {
"type": "inline",
"body": "party:\n - Rin"
"transcript": {"type": "inline", "body": "Source text"}
}
},
"vars": {
"session_date": "2026-05-04"
},
"model": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0,
"max_tokens": 800,
"top_p": 1,
"timeout_seconds": 120,
"service_tier": "priority",
"reasoning_effort": "medium",
"api_key_env": "SCRIPTORIUM_API_KEY",
"extra_params": {
"provider_option": "enabled"
}
},
"include_raw_output": false
}
```
Request fields:
| Field | Required | Description |
| Field | Required | Meaning |
| --- | --- | --- |
| `prompt_id` | yes | Prompt ID. Must not be blank. |
| `prompt_id` | yes | Non-blank prompt ID. |
| `prompt_version` | no | Prompt version filter. |
| `profile_id` | no | Execution profile ID. If omitted, the prompt must define `default_profile`. |
| `inputs` | yes | Object mapping prompt input names to input references. Must contain at least one entry. |
| `vars` | no | Object mapping template variable names to string values. |
| `model` | no | Runtime model override object. |
| `include_raw_output` | no | When `true`, include `raw_model_output` in the response. |
| `profile_id` | no | Execution-profile ID; otherwise the prompt must set `default_profile`. |
| `inputs` | yes | Non-empty object mapping input names to references. |
| `vars` | no | Object mapping template-variable names to strings. |
| `model` | no | Runtime model-override object. |
| `include_raw_output` | no | Include `raw_model_output` when true. |
Input reference fields:
An input reference has a required `type` of `file` or `inline`. A `file`
reference requires `uri`; an `inline` reference requires `body`.
| Field | Required | Description |
| --- | --- | --- |
| `type` | yes | `file` or `inline`. |
| `uri` | for `file` | File URI/path. |
| `body` | for `inline` | Inline artifact body. |
HTTP file references require a configured artifact root. Relative paths resolve
within that root. Absolute paths must be lexically within it; traversal outside
it is rejected with `400 artifact_not_allowed`. This lexical check does not
resolve symlinks: the operating system follows symlinks inside the root,
including ones that target outside it. Keep the root narrow and inaccessible to
untrusted writers.
HTTP `file` references require `server.artifact_root` or `serve
--artifact-root`. Relative file URIs resolve against that root. Absolute file
URIs are accepted only when lexically inside the root. Relative traversal and
absolute paths outside the root return `400 artifact_not_allowed`.
The optional `model` object accepts `endpoint`, `model`, `temperature`,
`max_tokens`, `top_p`, `timeout_seconds`, `service_tier`,
`reasoning_effort`, `api_key_env`, and `extra_params`. Numeric ranges and
credential supply are defined by the [configuration reference](config.md).
Explicit zero values for the numeric fields are overrides; zero
`timeout_seconds` disables the outbound client timeout.
The containment check is lexical and does not resolve symlinks. Symlinks inside
the artifact root are followed by the operating system, including symlinks that
point outside the root. Keep the artifact root narrow and not writable by
untrusted users.
Raw API-key values are not accepted. `api_key` and any other unknown model
field cause `400 invalid_json`.
Model override fields:
### Strict JSON
| Field | Description |
| --- | --- |
| `endpoint` | Runtime endpoint override. |
| `model` | Runtime model override. |
| `temperature` | Number in range `0..2`. Explicit `0` is an override. |
| `max_tokens` | Integer greater than or equal to `0`. Explicit `0` is an override. |
| `top_p` | Number in range `0..1`. Explicit `0` is an override. |
| `timeout_seconds` | Integer greater than or equal to `0`. Explicit `0` disables the outbound client timeout. |
| `service_tier` | Provider-specific request tier. |
| `reasoning_effort` | Provider-specific reasoning setting. |
| `api_key_env` | Name of an environment variable containing the API key. |
| `extra_params` | JSON-compatible provider-specific top-level request fields. |
Raw API-key values are not accepted in HTTP payloads. A field such as
`api_key` is rejected as unknown JSON.
`extra_params` keys must not be empty and must not collide with reserved
outbound fields: `model`, `session_id`, `messages`, `temperature`,
`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or
`response_format`.
### Strict JSON Rules
Request decoding is strict:
- malformed JSON returns `400 invalid_json`
- unknown request fields return `400 invalid_json`
- unknown `inputs` item fields return `400 invalid_json`
- unknown `model` fields return `400 invalid_json`
- trailing JSON tokens after the request object return `400 invalid_json`
- request bodies above the configured limit return `413 request_too_large`
Request decoding rejects malformed JSON, unknown fields at every request level,
and trailing JSON tokens with `400 invalid_json`. A blank `prompt_id` or
empty `inputs` object returns `400 invalid_request`.
### Success Response
Status: `200 OK`
A completed run returns `200 OK`, including when generated content fails its
validation contract. The response contains:
```json
{
"artifact": {
"name": "output",
"content_type": "text/markdown",
"body": "Generated content",
"size": 17,
"hash": "..."
},
"validation": {
"status": "passed",
"mode": "basic",
"repair_attempts": 0,
"is_valid": true
},
"metadata": {
"run_id": "...",
"prompt_id": "generic.markdown_summary",
"prompt_version": "1.0.0",
"prompt_hash": "...",
"rendered_prompt_hash": "...",
"selected_profile_id": "local-fast",
"model_name": "gpt-4o-mini",
"endpoint": "http://localhost:8000/v1",
"model_params": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0.2,
"max_tokens": 500,
"top_p": 1,
"timeout_seconds": 90
},
"input_hashes": {
"transcript": "..."
},
"usage": {
"prompt_tokens": 11,
"completion_tokens": 22,
"total_tokens": 33,
"cached_tokens": 0,
"cache_write_tokens": 0
},
"start_time": "2026-05-04T12:00:00Z",
"end_time": "2026-05-04T12:00:01Z",
"duration_ms": 1000,
"validation_mode": "basic",
"validation_status": "passed",
"repair_attempts_used": 0
}
}
```
- `artifact`: `name`, `content_type`, `body`, `size`, `hash`, and
optional `uri`;
- `validation`: `status`, `mode`, `repair_attempts`, `is_valid`, plus
optional `errors` and `schema_path`;
- `metadata`: run, prompt, rendered-prompt, profile, model, input-hash, usage,
timing, validation, and repair-attempt metadata; and
- optional `raw_model_output` when requested.
Response fields:
`metadata.model_params` has `endpoint`, `model`, `temperature`,
`max_tokens`, `top_p`, and `timeout_seconds`, plus optional
`service_tier`, `reasoning_effort`, `api_key_env`, and `extra_params`.
`metadata.usage` always includes `prompt_tokens`, `completion_tokens`,
`total_tokens`, `cached_tokens`, and `cache_write_tokens`; unavailable
cache usage is reported as zero.
- `artifact`: generated output artifact.
- `validation`: validation result for the generated artifact.
- `metadata`: run and effective runtime metadata.
- `raw_model_output`: omitted unless `include_raw_output` is `true`.
`artifact.uri` is omitted when empty. `validation.errors` and
`validation.schema_path` are omitted when empty. `model_params.service_tier`,
`model_params.reasoning_effort`, `model_params.api_key_env`, and
`model_params.extra_params` are omitted when empty.
`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are
always present as numbers. They are `0` when the provider omits compatible cache
usage fields or reports no cache activity.
### Validation Failure Response
Generated-content validation failures still return `200 OK`.
```json
{
"validation": {
"status": "failed",
"mode": "json",
"errors": ["invalid JSON: ..."],
"repair_attempts": 0,
"is_valid": false
}
}
```
The response still includes `artifact` and `metadata`.
A validation failure has `validation.status: "failed"`, `is_valid: false`,
and any available diagnostic errors, while still returning the artifact and
metadata.
## Error Responses
Error body shape:
Errors have this shape:
```json
{
"error": {
"code": "invalid_request",
"message": "prompt_id is required"
}
}
{"error":{"code":"invalid_request","message":"prompt_id is required"}}
```
Current status/code mapping:
Messages are concise and do not expose wrapped internal causes.
| Status | Code | Meaning |
| --- | --- | --- |
| `400` | `invalid_json` | Malformed JSON, unknown JSON field, or trailing JSON token. |
| `400` | `invalid_request` | Missing/invalid request fields or invalid runtime overrides. |
| `400` | `profile_required` | No `profile_id` and prompt has no `default_profile`. |
| `400` | `prompt_load_failed` | Prompt definition YAML/contract failed to load. |
| `400` | `profile_load_failed` | Profile YAML/contract failed to load, including raw `api_key`. |
| `400` | `artifact_not_allowed` | HTTP file refs are disabled or requested path is outside artifact root. |
| `400` | `artifact_read_failed` | Input artifact could not be read or input ref was unsupported/invalid. |
| `400` | `invalid_json` | Malformed JSON, unknown field, or trailing JSON. |
| `400` | `invalid_request` | Missing or invalid request data or runtime override. |
| `400` | `profile_required` | No profile ID and no prompt default profile. |
| `400` | `prompt_load_failed` | Prompt definition failed to load. |
| `400` | `profile_load_failed` | Profile failed to load. |
| `400` | `artifact_not_allowed` | HTTP file input is disabled or outside the artifact root. |
| `400` | `artifact_read_failed` | Input artifact is invalid or cannot be read. |
| `400` | `prompt_render_failed` | Prompt template rendering failed. |
| `400` | `api_key_env_missing` | Selected `api_key_env` variable is unset or empty. |
| `404` | `not_found` | Route path is unknown. |
| `404` | `prompt_not_found` | Prompt ID/version was not found. |
| `404` | `profile_not_found` | Profile ID was not found. |
| `405` | `method_not_allowed` | Method is not `POST` on `/v1/runs`. |
| `413` | `request_too_large` | Encoded JSON request body exceeds configured request limit. |
| `413` | `artifact_too_large` | HTTP file input artifact exceeds configured artifact limit. |
| `413` | `response_too_large` | Encoded JSON response exceeds configured response limit. |
| `500` | `validation_runtime_failed` | Validator runtime/schema loading failed. |
| `500` | `internal_error` | Unclassified server error. |
| `400` | `api_key_env_missing` | The selected credential environment variable is unset or empty. |
| `404` | `not_found` | Route does not exist. |
| `404` | `prompt_not_found` | Prompt ID or version does not exist. |
| `404` | `profile_not_found` | Profile ID does not exist. |
| `405` | `method_not_allowed` | The route does not accept the method. |
| `413` | `request_too_large` | Encoded request exceeds its limit. |
| `413` | `artifact_too_large` | File input exceeds its limit. |
| `413` | `response_too_large` | Encoded response exceeds its limit. |
| `500` | `validation_runtime_failed` | Schema or validator runtime failure. |
| `500` | `internal_error` | Unclassified server failure. |
| `502` | `llm_failed` | Outbound model request failed. |
HTTP error messages are intentionally concise and do not include sensitive
internal causes.
## Retry And Idempotency
Scriptorium does not provide idempotency keys, pagination, caching headers, or
rate limiting.
Clients may retry transport failures or `5xx` responses when their surrounding
workflow can tolerate another model call. A retry can generate different output
and incur another provider request.
## Example File
- `examples/http-run.json`
Scriptorium provides no idempotency keys, pagination, caching headers, or rate
limits. Clients may retry transport failures or `5xx` responses only when
their workflow tolerates another model call: a retry can produce different
output and incur another provider request.

View File

@@ -1,5 +1,10 @@
# CLI Reference
This is the canonical contract for invoking Scriptorium. Configuration discovery,
precedence, directories, profiles, and schemas are defined in the
[configuration reference](config.md). The [HTTP API reference](api.md) owns
service request and response behavior.
## Shortest Useful Command
```bash
@@ -10,224 +15,132 @@ go run ./cmd/scriptorium render \
--input glossary=./examples/fixtures/glossary.yml
```
`render` prepares the prompt, loads input artifacts, resolves the execution
profile, and prints the prepared request without calling an LLM.
`render` prepares a request without calling an LLM.
## Command Overview
## Commands
- `scriptorium run`: prepare a prompt, call the configured LLM, write generated output, and print a run summary.
- `scriptorium render`: prepare a prompt only; write prepared-run output as `text` or `json`.
- `scriptorium serve`: start the HTTP server for `POST /v1/runs`.
- `scriptorium run`: prepare a prompt, call the configured LLM, and write the
generated artifact.
- `scriptorium render`: prepare a prompt and write prepared-run output.
- `scriptorium serve`: start the HTTP server.
Canonical related references:
All commands accept `--config <path>` and reject positional arguments. An
effective `prompt_dir` is required for every command. Supply it through the
configuration contract or the command's `--prompt-dir` flag.
- [Configuration reference](config.md)
- [HTTP API reference](api.md)
- [Subprocess integration](integrations/subprocess.md)
## `scriptorium run`
## Common Rules
- `--config` is supported by `run`, `render`, and `serve`.
- Positional arguments are rejected.
- `run` and `render` require `--prompt`, at least one `--input`, and an effective `prompt_dir`.
- `serve` requires an effective `prompt_dir`.
- `profile_dir` is optional. Without it, only built-in profiles are available.
- If `profile_dir` is set, custom profiles override built-in profiles with the same ID.
- Prompt cache control, `session_id`, structured output, and provider-specific profile fields are configured in YAML, not with CLI flags.
Config precedence is:
1. built-in defaults
2. config file values
3. CLI flags
## Flag Reference
### `scriptorium run`
```bash
```text
scriptorium run [flags]
```
Required through flags or config:
Required flags:
- `--prompt-dir <dir>`: prompt definition directory.
Required as flags:
- `--prompt <id>`: prompt ID to execute.
- `--input name=path`: input file mapping. Repeat or use comma-separated mappings.
| Flag | Meaning |
| --- | --- |
| `--prompt <id>` | Prompt ID to execute. |
| `--input name=path` | Input file mapping; repeat or use comma-separated mappings. |
Optional flags:
- `--config <path>`: application config file.
- `--profile-dir <dir>`: custom profile definition directory.
- `--schema-dir <dir>`: schema base directory for `json_schema` validation.
- `--profile <id>`: execution profile override. If omitted, the prompt `default_profile` is used.
- `--var name=value`: template variable mapping. Repeat or use comma-separated mappings.
- `--out <path>`: write generated artifact body to a file instead of stdout.
- `--llm-base-url <url>`: runtime endpoint override.
- `--model <name>`: runtime model override.
- `--api-key-env <name>`: runtime API-key environment variable name override.
- `--temperature <float>`: runtime temperature override.
- `--max-tokens <int>`: runtime max tokens override.
- `--top-p <float>`: runtime top-p override.
- `--timeout <duration>`: runtime timeout override using Go duration syntax, such as `30s` or `2m`.
| Flag | Meaning |
| --- | --- |
| `--config <path>` | Application configuration file. |
| `--prompt-dir <dir>` | Prompt-definition directory override. |
| `--profile-dir <dir>` | Custom profile-directory override. |
| `--schema-dir <dir>` | Schema base-directory override. |
| `--profile <id>` | Execution-profile override. |
| `--var name=value` | Template-variable mapping; repeat or use comma-separated mappings. |
| `--out <path>` | Write generated content to this file instead of stdout. |
| `--llm-base-url <url>` | Runtime endpoint override. |
| `--model <name>` | Runtime model override. |
| `--api-key-env <name>` | Runtime API-key environment-variable name override. |
| `--temperature <float>` | Runtime temperature override. |
| `--max-tokens <int>` | Runtime maximum-token override. |
| `--top-p <float>` | Runtime top-p override. |
| `--timeout <duration>` | Runtime timeout override using Go duration syntax. |
Deprecated aliases:
Deprecated aliases: `--prompt-id` for `--prompt`, and `--profile-id` for
`--profile`.
- `--prompt-id <id>`: alias for `--prompt`.
- `--profile-id <id>`: alias for `--profile`.
Omitted numeric runtime flags preserve the selected effective value; explicit
zero values override it. `--timeout 0s` disables the outbound HTTP-client
timeout. CLI durations are converted to whole seconds by truncation toward
zero, so any duration whose absolute value is below one second becomes an
explicit zero-second override.
Runtime override notes:
There is no raw API-key flag. Use `--api-key-env`.
- Omitted numeric override flags preserve the selected profile/default value.
- Explicit zero values override the selected profile/default value.
- `--timeout 0s` disables the outbound HTTP client timeout for that request.
- There is no raw API-key flag; use `--api-key-env`.
## `scriptorium render`
### `scriptorium render`
```bash
```text
scriptorium render [flags]
```
Required through flags or config:
`--prompt <id>` and at least one `--input name=path` are required. The
following optional flags are supported: `--config`, `--prompt-dir`,
`--profile-dir`, `--profile`, `--var`, `--out`, `--llm-base-url`,
`--model`, `--api-key-env`, `--temperature`, `--max-tokens`, `--top-p`,
`--timeout`, and `--format text|json`. Their meanings match the corresponding
`run` flags; `--format` selects prepared-run output and otherwise uses
`defaults.render_format`.
- `--prompt-dir <dir>`: prompt definition directory.
The same deprecated aliases and numeric/timeout behavior as `run` apply.
`render` does not accept `--schema-dir`; configure `schema_dir` through the
configuration file. It resolves profiles and schemas as part of preparation but
does not call an LLM.
Required as flags:
## `scriptorium serve`
- `--prompt <id>`: prompt ID to render.
- `--input name=path`: input file mapping. Repeat or use comma-separated mappings.
Optional flags:
- `--config <path>`: application config file.
- `--prompt-dir <dir>`: prompt definition directory.
- `--profile-dir <dir>`: custom profile definition directory.
- `--profile <id>`: execution profile override.
- `--var name=value`: template variable mapping. Repeat or use comma-separated mappings.
- `--out <path>`: write prepared-run output to a file instead of stdout.
- `--llm-base-url <url>`: runtime endpoint override for the prepared request.
- `--model <name>`: runtime model override for the prepared request.
- `--api-key-env <name>`: runtime API-key environment variable name override.
- `--temperature <float>`: runtime temperature override.
- `--max-tokens <int>`: runtime max tokens override.
- `--top-p <float>`: runtime top-p override.
- `--timeout <duration>`: runtime timeout override using Go duration syntax.
- `--format text|json`: prepared-run output format. Defaults to config `defaults.render_format`, then `text`.
Deprecated aliases:
- `--prompt-id <id>`: alias for `--prompt`.
- `--profile-id <id>`: alias for `--profile`.
Notes:
- `render` resolves profiles, loads schemas for `json_schema` prompts, and validates `api_key_env`.
- `render` does not accept `--schema-dir`; use config `schema_dir` for render-time schema lookup.
- `render` does not call the LLM.
### `scriptorium serve`
```bash
```text
scriptorium serve [flags]
```
Required through flags or config:
- `--prompt-dir <dir>`: prompt definition directory.
Optional flags:
- `--config <path>`: application config file.
- `--addr <listen-address>`: HTTP listen address.
- `--prompt-dir <dir>`: prompt definition directory.
- `--profile-dir <dir>`: custom profile definition directory.
- `--schema-dir <dir>`: schema base directory for `json_schema` validation.
- `--artifact-root <dir>`: base directory for HTTP `file` input references.
- `--max-request-bytes <n>`: maximum HTTP request body bytes; `0` disables the limit.
- `--max-artifact-bytes <n>`: maximum HTTP file artifact bytes; `0` disables the limit.
- `--max-response-bytes <n>`: maximum encoded HTTP response body bytes; `0` disables the limit.
| Flag | Meaning |
| --- | --- |
| `--config <path>` | Application configuration file. |
| `--addr <listen-address>` | HTTP listen-address override. |
| `--prompt-dir <dir>` | Prompt-definition directory override. |
| `--profile-dir <dir>` | Custom profile-directory override. |
| `--schema-dir <dir>` | Schema base-directory override. |
| `--artifact-root <dir>` | Root for HTTP `file` input references. |
| `--max-request-bytes <n>` | Maximum encoded HTTP request-body bytes; `0` disables the limit. |
| `--max-artifact-bytes <n>` | Maximum HTTP file-input artifact bytes; `0` disables the limit. |
| `--max-response-bytes <n>` | Maximum encoded HTTP response bytes; `0` disables the limit. |
Notes:
- `serve` does not accept runtime model override flags such as `--model` or `--llm-base-url`.
- HTTP request fields and error codes are documented in the [HTTP API reference](api.md).
- HTTP `file` input references are rejected unless an artifact root is configured.
- HTTP size-limit flags affect only `serve`.
`serve` accepts no runtime model override flags. HTTP request fields, response
schemas, and error codes are defined in the [HTTP API reference](api.md).
## Input And Variable Syntax
- `--input name=path` maps prompt input names to local file paths.
- `--var name=value` maps prompt template variables to string values.
- Both flags can be repeated.
- Both flags also accept comma-separated mappings, such as `--input transcript=./t.md,glossary=./g.yml`.
- Values may contain `=` after the first separator, such as `--var note=a=b=c`.
- Empty names and empty values are rejected.
`--input name=path` maps an input name to a local file; `--var name=value`
maps a template variable to a string. Both flags can be repeated or contain
comma-separated mappings. Values may contain `=` after the first separator.
Empty names and values are rejected.
CLI `run` and `render` convert every `--input` mapping to a `file` artifact
reference. HTTP also supports `inline` input references; see [HTTP API
reference](api.md).
CLI inputs are file references. HTTP inline inputs are defined by the
[HTTP API reference](api.md).
## Output Behavior
## Output And Exit Behavior
`run`:
- `run` writes generated content to stdout, or to `--out` when supplied, and
writes a concise summary to stderr.
- `render` writes prepared-run output to stdout, or to `--out` when supplied,
without a success summary.
- `serve` writes startup and server errors to stderr.
- Writes generated artifact content to stdout by default.
- Writes generated artifact content to `--out` when provided.
- Prints a success summary to stderr.
- Prints errors to stderr on failure.
Exit statuses:
`render`:
| Status | Meaning |
| --- | --- |
| `0` | Success. |
| `1` | Parse, configuration, loading, rendering, generation, output-write, or other runtime error. |
| `2` | `run` generated and wrote output, but validation failed. |
- Writes prepared-run output to stdout by default.
- Writes prepared-run output to `--out` when provided.
- Does not print a success summary.
## Workflows And Examples
`serve`:
- Logs startup and server errors to stderr.
## Exit Codes
- `0`: success.
- `1`: parse, config, load, render, generation, output-write, or runtime error.
- `2`: `run` completed and wrote output, but validation status is `failed`.
## Common Workflows
Render prompt inputs and variables as JSON:
```bash
go run ./cmd/scriptorium render \
--config ./examples/config.yml \
--prompt generic.markdown_summary \
--input transcript=./examples/fixtures/transcript.md \
--input glossary=./examples/fixtures/glossary.yml \
--var session_date=2026-05-04 \
--format json
```
Run a prompt with an explicit profile and file output:
```bash
go run ./cmd/scriptorium run \
--config ./examples/config.yml \
--prompt generic.markdown_summary \
--profile local-fast \
--input transcript=./examples/fixtures/transcript.md \
--input glossary=./examples/fixtures/glossary.yml \
--out ./summary.md
```
Start the HTTP server with example config:
```bash
go run ./cmd/scriptorium serve --config ./examples/config.yml
```
Copyable maintained script:
- `examples/render-markdown-summary.sh`
The [maintained render script](../examples/render-markdown-summary.sh) is a
copyable render workflow. The [HTTP request example](../examples/http-run.json)
is for a running `serve` process.

View File

@@ -1,324 +1,169 @@
# Configuration Reference
## Config Discovery And Precedence
This is the canonical reference for Scriptorium application settings and the
prompt, profile, and schema files those settings select. For command syntax,
see the [CLI reference](cli.md); for HTTP request shapes, limits, and outcomes,
see the [HTTP API reference](api.md).
## Discovery And Precedence
Application settings are resolved in this order:
1. built-in defaults
2. `config.yml` values
3. CLI overrides
1. built-in defaults;
2. a configuration file; then
3. CLI overrides.
When `--config` is omitted, Scriptorium searches:
When `--config` is omitted, Scriptorium searches
`/usr/local/etc/scriptorium/config.yml` and then `/etc/scriptorium/config.yml`.
If neither exists, it uses built-in defaults. An explicit `--config` path must
exist and decode successfully.
1. `/usr/local/etc/scriptorium/config.yml`
2. `/etc/scriptorium/config.yml`
The maintained [minimal configuration](../examples/config.yml) and
[full configuration](../examples/config.full.yml) are copyable examples.
If neither file exists, Scriptorium uses built-in defaults. When
`--config <path>` is provided, that file must exist and decode successfully.
## Application Configuration File
## Minimal Working Config
Configuration is strict YAML: unknown fields are rejected. Empty string values
do not override a prior value. Raw API-key fields are not accepted.
```yaml
prompt_dir: ./examples/prompts
```
This is enough for `run` and `render` when selected prompts use built-in
profiles. Set `profile_dir` when prompts or requests use custom profiles.
The maintained repository example is `examples/config.yml`.
## Production-Oriented Config
```yaml
prompt_dir: /opt/scriptorium/prompts
profile_dir: /opt/scriptorium/profiles
schema_dir: /opt/scriptorium/schemas
server:
addr: 127.0.0.1:8080
artifact_root: /var/lib/scriptorium/artifacts
max_request_bytes: 16777216
max_artifact_bytes: 16777216
max_response_bytes: 16777216
defaults:
render_format: text
```
The maintained full example is `examples/config.full.yml`.
## App Config Reference
Top-level fields:
| Field | Default | Description |
| Field | Default | Meaning |
| --- | --- | --- |
| `prompt_dir` | unset | Directory containing prompt definition YAML files. Required effectively by `run`, `render`, and `serve`. |
| `profile_dir` | unset | Directory containing custom profile YAML files. Built-in profiles remain available when unset. |
| `prompt_dir` | unset | Directory containing prompt-definition YAML. `run`, `render`, and `serve` require an effective value. |
| `profile_dir` | unset | Directory containing custom profile YAML. Built-in profiles remain available. |
| `schema_dir` | `.` | Base directory for relative JSON Schema paths. |
| `server` | `{}` | HTTP service settings used by `serve`. |
| `defaults` | `{}` | Adapter defaults. |
`server` fields:
| Field | Default | Description |
| --- | --- | --- |
| `server.addr` | `:8080` | Listen address for `serve`. |
| `server.artifact_root` | unset | Base directory for HTTP `file` input references. Without it, HTTP file refs are rejected. |
| `server.max_request_bytes` | `16777216` | Maximum encoded HTTP request body bytes. `0` disables the limit. |
| `server.max_artifact_bytes` | `16777216` | Maximum HTTP file artifact bytes. `0` disables the limit. |
| `server.max_response_bytes` | `16777216` | Maximum encoded HTTP response bytes. `0` disables the limit. |
`defaults` fields:
| Field | Default | Description |
| --- | --- | --- |
| `server.addr` | `:8080` | Address used by `serve`. |
| `server.artifact_root` | unset | Root that enables HTTP `file` input references. |
| `server.max_request_bytes` | `16777216` | Maximum encoded HTTP request body bytes; `0` disables the limit. |
| `server.max_artifact_bytes` | `16777216` | Maximum HTTP file-input artifact bytes; `0` disables the limit. |
| `server.max_response_bytes` | `16777216` | Maximum encoded HTTP response bytes; `0` disables the limit. |
| `defaults.render_format` | `text` | Default `render` output format: `text` or `json`. |
Config rules:
- YAML decoding is strict; unknown fields are rejected.
- HTTP size limits must be greater than or equal to `0`.
- Empty string config values are ignored.
- Raw API key fields are not supported in app config.
The three size fields must be zero or greater. The HTTP contract defines how
each limit is enforced and reported. `server.artifact_root` configures the
deployment boundary; see the [HTTP API reference](api.md) for request-path and
containment behavior, and [operations](operations.md) for deployment handling.
## Prompt Definition Files
Prompt definitions are YAML files anywhere under `prompt_dir`. Nested
directories are organizational; callers select prompts by YAML `id`, not file
path.
Prompt definitions are strict YAML files anywhere below `prompt_dir`. A prompt
is selected by its YAML `id`, not by file path; nested directories are only for
organization. See [maintained prompt examples](../examples/prompts/).
Example:
```yaml
id: generic.structured_events
version: "1.0.0"
default_profile: local-quality
description: Produce structured event JSON from a transcript.
inputs:
- name: transcript
required: true
content_type: text/markdown
description: Source transcript content
- name: glossary
required: false
content_type: text/yaml
description: Optional glossary context
messages:
- role: system
content_file: ./generic.structured_events.system.md
- role: user
content_file: ./generic.structured_events.user.md
output:
format: json
validation_mode: json_schema
schema_path: structured_events.schema.json
repair_attempts: 0
```
Prompt fields:
| Field | Required | Description |
| Field | Required | Meaning |
| --- | --- | --- |
| `id` | yes | Prompt identifier used by `--prompt` and HTTP `prompt_id`. |
| `id` | yes | Prompt identifier. |
| `version` | yes | Prompt version. |
| `default_profile` | no | Profile ID used when a request does not provide a profile. |
| `default_profile` | no | Profile used when a request omits a profile ID. |
| `description` | no | Human-readable description. |
| `session_id` | no | Go-template string rendered from request vars and forwarded as provider `session_id` when non-empty. |
| `inputs` | no | Named input declarations. |
| `messages` | yes | Chat message templates. |
| `session_id` | no | Go-template string rendered from request variables and sent to a compatible provider when non-empty. |
| `inputs` | no | Declared input metadata. |
| `messages` | yes | Chat-message templates. |
| `output` | yes | Output format and validation contract. |
`inputs[]` fields:
### Inputs And Messages
- `name` (required)
- `required` (optional boolean)
- `content_type` (optional metadata)
- `description` (optional)
Each `inputs` item has a required `name` and optional `required`,
`content_type`, and `description` fields. Input names must be unique.
`messages[]` fields:
Each message has a required `role`, exactly one of `content` or `content_file`,
and optional `cache_control`. A `content_file` path is relative to the prompt
file. `cache_control.type` must be `ephemeral`; its optional `ttl` is `1h`.
- `role` (required)
- exactly one of `content` or `content_file`
- `cache_control` (optional)
`session_id` uses the same template variables as messages. Empty rendered
values are omitted. A rendered value may contain at most 256 Unicode code
points.
Message rules:
### Output Contract
- `content_file` resolves relative to the prompt YAML file location.
- Repeated roles are allowed.
- Prompt YAML decoding is strict.
- Duplicate input names are invalid.
- Duplicate prompt IDs are invalid for a requested ID/version.
`messages[].cache_control` fields:
| Field | Required | Supported values |
| Field | Required | Values or behavior |
| --- | --- | --- |
| `type` | yes | `ephemeral` |
| `ttl` | no | `1h` |
`session_id` behavior:
- Rendered with the same variable context as message templates.
- Trimmed and omitted when empty.
- Rejected when longer than 256 Unicode code points.
- CLI callers pass variables with `--var`; HTTP callers use `vars`.
`output` fields:
| Field | Required | Supported values |
| --- | --- | --- |
| `format` | yes | `text`, `markdown`, `json` |
| `validation_mode` | yes | `none`, `basic`, `json`, `json_schema` |
| `schema_path` | only for `json_schema` | Relative to `schema_dir` unless absolute. |
| `repair_attempts` | yes | Integer greater than or equal to `0`. |
Repair boundary:
- `repair_attempts` is part of the prompt contract.
- The current CLI and HTTP wiring constructs the runner without a repairer, so normal `run` and `serve` execution does not perform repair attempts.
| `format` | yes | `text`, `markdown`, or `json`. |
| `validation_mode` | yes | `none`, `basic`, `json`, or `json_schema`. |
| `schema_path` | for `json_schema` | Schema path, relative to `schema_dir` unless absolute. |
| `repair_attempts` | no | Integer greater than or equal to `0`; omitted means `0`. |
## Profile Definition Files
Execution profiles are YAML files anywhere under `profile_dir`. Nested
directories are organizational; callers select profiles by YAML `id`, not file
path.
Profiles are strict YAML files anywhere below `profile_dir`. A profile is
selected by YAML `id`; nested directories are organizational. See the
[maintained profile examples](../examples/profiles/).
Scriptorium also ships built-in profiles. Custom profiles override built-ins
with the same ID.
Example:
```yaml
id: local-fast
endpoint: http://localhost:8000/v1
model: gpt-4o-mini
temperature: 0.2
max_tokens: 500
top_p: 1.0
timeout_seconds: 90
api_key_env: SCRIPTORIUM_API_KEY
service_tier: priority
reasoning_effort: medium
extra_params:
provider_route: primary
```
Profile fields:
| Field | Required | Description |
| Field | Required | Meaning |
| --- | --- | --- |
| `id` | yes | Profile identifier. |
| `endpoint` | yes | OpenAI-compatible base URL including `/v1`. |
| `endpoint` | yes | OpenAI-compatible base URL, including its API version path when needed. |
| `model` | yes | Provider model name. |
| `temperature` | no | Range `0..2`. |
| `max_tokens` | no | Integer greater than or equal to `0`. |
| `top_p` | no | Range `0..1`. |
| `timeout_seconds` | no | Integer greater than or equal to `0`. |
| `service_tier` | no | Provider-specific request tier. |
| `reasoning_effort` | no | Provider-specific reasoning setting. |
| `api_key_env` | no | Environment variable name containing the API key. |
| `extra_params` | no | JSON-compatible provider-specific top-level request fields. |
| `temperature` | no | Number from `0` through `2`. |
| `max_tokens` | no | Integer zero or greater. |
| `top_p` | no | Number from `0` through `1`. |
| `timeout_seconds` | no | Integer zero or greater. |
| `service_tier` | no | Non-empty provider-specific request tier. |
| `reasoning_effort` | no | Non-empty provider-specific reasoning setting. |
| `api_key_env` | no | Environment-variable name containing the API key. |
| `extra_params` | no | JSON-compatible provider-specific outbound request fields. |
Execution defaults before profile/request overrides:
Execution defaults before profile and request overrides are `temperature: 0`,
`max_tokens: 0`, `top_p: 1`, and `timeout_seconds: 600`. Profile numeric values
merge by non-zero value. Request overrides preserve presence, so an explicit
zero can override a profile value.
| Field | Default |
| --- | --- |
| `temperature` | `0.0` |
| `max_tokens` | `0` |
| `top_p` | `1.0` |
| `timeout_seconds` | `600` |
Custom profiles take precedence over built-ins with the same ID. Invalid custom
profiles are errors; they do not fall back to a built-in profile. Raw `api_key`
is rejected. Use `api_key_env`, or the public Go package's request-scoped key
mechanism described in the [package contract](consumers/pkg-scriptorium.md).
Profile rules:
`extra_params` keys must be non-empty and cannot be `model`, `session_id`,
`messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`,
`reasoning_effort`, or `response_format`.
- Profile YAML decoding is strict.
- Duplicate custom profile IDs are invalid.
- Matching custom and built-in IDs are valid override behavior.
- Raw `api_key` is rejected; use `api_key_env`.
- If `api_key_env` is set, the named environment variable must be set before `run`, `render`, or HTTP execution can prepare the request.
- Profile numeric fields merge by non-zero value. Request overrides are presence-aware, so explicit zero values are supported through CLI flags or HTTP model overrides.
- `extra_params` keys must not be empty and must not collide with reserved outbound fields: `model`, `session_id`, `messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or `response_format`.
### Built-In Profile Catalog
Built-in profile catalog:
Each embedded profile uses `OPENROUTER_API_KEY`.
| Provider | ID | Model | API key env |
| --- | --- | --- | --- |
| aion-labs | `aion-2` | `aion-labs/aion-2.0` | `OPENROUTER_API_KEY` |
| anthropic | `claude-fable-latest` | `~anthropic/claude-fable-latest` | `OPENROUTER_API_KEY` |
| anthropic | `claude-haiku-latest` | `~anthropic/claude-haiku-latest` | `OPENROUTER_API_KEY` |
| anthropic | `claude-opus-latest` | `~anthropic/claude-opus-latest` | `OPENROUTER_API_KEY` |
| anthropic | `claude-sonnet-latest` | `~anthropic/claude-sonnet-latest` | `OPENROUTER_API_KEY` |
| deepseek | `deepseek-3-2` | `deepseek/deepseek-v3.2` | `OPENROUTER_API_KEY` |
| deepseek | `deepseek-4-pro` | `deepseek/deepseek-v4-pro` | `OPENROUTER_API_KEY` |
| google | `gemini-2-flash` | `google/gemini-2.5-flash` | `OPENROUTER_API_KEY` |
| google | `gemini-2-flash-lite` | `google/gemini-2.5-flash-lite` | `OPENROUTER_API_KEY` |
| google | `gemini-2-pro` | `google/gemini-2.5-pro` | `OPENROUTER_API_KEY` |
| google | `gemini-3-flash-lite` | `google/gemini-3.1-flash-lite` | `OPENROUTER_API_KEY` |
| google | `gemini-flash-latest` | `~google/gemini-flash-latest` | `OPENROUTER_API_KEY` |
| google | `gemini-pro-latest` | `~google/gemini-pro-latest` | `OPENROUTER_API_KEY` |
| google | `gemma-4-31b` | `google/gemma-4-31b-it:exacto` | `OPENROUTER_API_KEY` |
| minimax | `minimax-m2` | `minimax/minimax-m2.5` | `OPENROUTER_API_KEY` |
| minimax | `minimax-m3` | `minimax/minimax-m3` | `OPENROUTER_API_KEY` |
| mistral | `mistral-large-2512` | `mistralai/mistral-large-2512` | `OPENROUTER_API_KEY` |
| mistral | `mistral-medium-3-5` | `mistralai/mistral-medium-3-5` | `OPENROUTER_API_KEY` |
| mistral | `mistral-small-3` | `mistralai/mistral-small-3.2-24b-instruct` | `OPENROUTER_API_KEY` |
| mistral | `mistral-small-4` | `mistralai/mistral-small-2603` | `OPENROUTER_API_KEY` |
| nvidia | `nemotron-3-ultra` | `nvidia/nemotron-3-ultra-550b-a55b` | `OPENROUTER_API_KEY` |
| openai | `gpt-5-mini` | `openai/gpt-5.4-mini` | `OPENROUTER_API_KEY` |
| openai | `gpt-5-nano` | `openai/gpt-5.4-nano` | `OPENROUTER_API_KEY` |
| Provider | ID | Model |
| --- | --- | --- |
| aion-labs | `aion-2` | `aion-labs/aion-2.0` |
| anthropic | `claude-fable-latest` | `~anthropic/claude-fable-latest` |
| anthropic | `claude-haiku-latest` | `~anthropic/claude-haiku-latest` |
| anthropic | `claude-opus-latest` | `~anthropic/claude-opus-latest` |
| anthropic | `claude-sonnet-latest` | `~anthropic/claude-sonnet-latest` |
| deepseek | `deepseek-3-2` | `deepseek/deepseek-v3.2` |
| deepseek | `deepseek-4-flash` | `deepseek/deepseek-v4-flash` |
| deepseek | `deepseek-4-pro` | `deepseek/deepseek-v4-pro` |
| google | `gemini-2-flash` | `google/gemini-2.5-flash` |
| google | `gemini-2-flash-lite` | `google/gemini-2.5-flash-lite` |
| google | `gemini-2-pro` | `google/gemini-2.5-pro` |
| google | `gemini-3-flash-lite` | `google/gemini-3.1-flash-lite` |
| google | `gemini-flash-latest` | `~google/gemini-flash-latest` |
| google | `gemini-pro-latest` | `~google/gemini-pro-latest` |
| google | `gemma-4-31b` | `google/gemma-4-31b-it:exacto` |
| minimax | `minimax-m2` | `minimax/minimax-m2.5` |
| minimax | `minimax-m3` | `minimax/minimax-m3` |
| mistral | `mistral-large-2512` | `mistralai/mistral-large-2512` |
| mistral | `mistral-medium-3-5` | `mistralai/mistral-medium-3-5` |
| mistral | `mistral-small-3` | `mistralai/mistral-small-3.2-24b-instruct` |
| mistral | `mistral-small-4` | `mistralai/mistral-small-2603` |
| nvidia | `nemotron-3-ultra` | `nvidia/nemotron-3-ultra-550b-a55b` |
| openai | `gpt-5-mini` | `openai/gpt-5.4-mini` |
| openai | `gpt-5-nano` | `openai/gpt-5.4-nano` |
## Schema Behavior
## Schemas
Schemas are JSON files, typically under `schema_dir`.
Schemas are JSON files, normally below `schema_dir`. `json_schema` output
requires a `schema_path`. Relative paths resolve from `schema_dir`; absolute
paths are used directly. Referenced nested schemas use relative paths and are
not discovered by basename. An unreadable or invalid schema is a runtime
validation error; generated content that fails JSON or schema validation is a
validation result.
Rules:
## Credentials
- `output.validation_mode: json_schema` requires `output.schema_path`.
- Relative `schema_path` values resolve from `schema_dir`.
- Absolute `schema_path` values are used directly.
- Nested schemas must be referenced by relative path; schemas are not searched recursively by basename.
- Missing or invalid schema documents are runtime validation errors.
- Invalid generated JSON produces validation status `failed`, not a runtime error.
Keep secrets in environment variables. Store only an environment-variable name
in `api_key_env`; do not place raw keys in configuration, prompt or profile
files, CLI arguments, examples, or HTTP payloads.
## Artifact References
Supported request input artifact reference types are:
- `file`
- `inline`
CLI `run` and `render` create `file` references from `--input name=path`.
HTTP `file` references require `server.artifact_root` or `serve
--artifact-root`. Relative file URIs resolve under that root. Absolute paths
and relative traversal outside the root are rejected by lexical checks. Symlinks
inside the root are followed by the operating system, including symlinks that
point outside the root.
HTTP `inline` references do not require an artifact root.
## Secrets Handling
- Keep secret values in environment variables.
- Store only environment-variable names in `api_key_env`.
- Do not put raw API keys in config, prompts, profiles, CLI arguments, examples, or HTTP request bodies.
## Maintained Examples
- Minimal app config: `examples/config.yml`
- Full app config: `examples/config.full.yml`
- Prompt examples: `examples/prompts/`
- Custom profile examples: `examples/profiles/`
- Schema examples: `examples/schemas/`
- Input fixtures: `examples/fixtures/`
- Render script: `examples/render-markdown-summary.sh`
- HTTP request-shape example: `examples/http-run.json`
## Integration References
## Related References
- [CLI reference](cli.md)
- [HTTP API reference](api.md)
- [Outbound OpenAI-compatible contract](integrations/openai-compatible-chat.md)
- [OpenAI-compatible outbound contract](integrations/openai-compatible-chat.md)

View File

@@ -1,65 +1,26 @@
# Consumer Integration Overview
This guide is for applications that call Scriptorium from another codebase.
This guide helps applications choose a Scriptorium interface and understand
their responsibilities. The linked contracts own interface syntax and wire
semantics.
Scriptorium exposes three integration surfaces:
| Surface | Use when |
| Interface | Use when |
| --- | --- |
| Go package | The consumer is Go, needs typed requests/results, or wants injected LLM clients for tests. |
| CLI subprocess | The consumer wants process isolation or is not written in Go. |
| HTTP API | The consumer needs a service boundary or remote access to `POST /v1/runs`. |
| Go package | The consumer is Go and needs typed requests, results, or an injected LLM client. |
| CLI subprocess | The consumer needs process isolation or is not written in Go. |
| HTTP API | The consumer needs a service boundary or remote access. |
Canonical references:
- Go package: [package contract](pkg-scriptorium.md)
- CLI subprocess: [subprocess integration](../integrations/subprocess.md)
- HTTP service: [HTTP API reference](../api.md)
- Prompt, profile, schema, and credential configuration: [configuration reference](../config.md)
- Go package: [Package scriptorium](pkg-scriptorium.md)
- CLI subprocess: [Subprocess integration](../integrations/subprocess.md)
- HTTP: [HTTP API reference](../api.md)
- File formats: [Configuration reference](../config.md)
## Required Deployment Inputs
Every integration needs operators to provide:
- prompt definitions;
- profile definitions or built-in profile IDs;
- schema files when prompts use `json_schema`;
- input artifacts or inline input bodies;
- API-key environment variables or direct per-request keys where supported.
Raw API keys do not belong in config, prompt files, profile YAML, CLI
arguments, or HTTP request bodies.
## Recommended Workflow
Use the Go package when:
- the consumer is a Go application;
- the application needs `context.Context` cancellation;
- repeated calls should avoid subprocess startup;
- tests need a fake LLM client;
- direct per-request `RunRequest.APIKey` is required.
Use the CLI subprocess when:
- the consumer is not Go;
- process isolation is useful;
- stdout/stderr separation and exit codes are enough;
- the consumer already manages local files and environment variables.
Use HTTP when:
- Scriptorium should run as a service;
- multiple clients need a shared prompt/profile deployment;
- clients can reach a trusted, protected HTTP boundary.
## Minimal Go Example
## Minimal Go Use
```go
engine, err := scriptorium.NewEngine(scriptorium.Config{
PromptDir: "./examples/prompts",
ProfileDir: "./examples/profiles",
SchemaDir: "./examples/schemas",
})
if err != nil {
return err
@@ -69,54 +30,30 @@ prepared, err := engine.Prepare(ctx, scriptorium.RunRequest{
PromptID: "generic.markdown_summary",
Inputs: map[string]scriptorium.ArtifactRef{
"transcript": scriptorium.File("./examples/fixtures/transcript.md"),
"glossary": scriptorium.File("./examples/fixtures/glossary.yml"),
},
})
if err != nil {
return err
}
_ = prepared.Messages
_ = prepared
```
Run the maintained package example:
```bash
go run ./examples/go-library/prepare
```
## Subprocess Workflow
Invoke `scriptorium render` for preflight and `scriptorium run` for generation.
Capture stdout and stderr separately. Treat exit code `2` from `run` as a
completed generation with failed validation.
See [Subprocess integration](../integrations/subprocess.md) for the stable
invocation contract.
## HTTP Workflow
Run `scriptorium serve` behind trusted controls and send JSON requests to
`POST /v1/runs`.
Do not duplicate endpoint schemas in consumers. Use the [HTTP API
reference](../api.md) as the authoritative contract.
For a maintained program, see
[`examples/go-library/prepare`](../../examples/go-library/prepare).
## Consumer Responsibilities
Consumers are responsible for:
- selecting prompt/profile IDs as deployment configuration;
- supplying all required inputs and vars;
- protecting generated artifacts and rendered prompts as sensitive data;
- deciding whether to keep output when validation fails;
- implementing retries only when another model call is acceptable.
- selecting and deploying prompt, profile, and schema assets;
- supplying required inputs and template variables;
- supplying credentials through the applicable interface;
- protecting rendered prompts and generated artifacts as potentially sensitive;
- deciding whether validation-failed output is usable; and
- retrying only when another model call is acceptable.
Scriptorium does not persist run state. Retrying a failed or timed-out request
can produce different output and can incur another provider request.
## Status Behavior
- Go package methods return typed results or errors that support `errors.Is`.
- CLI `run` exits `2` when generation succeeds but validation fails.
- HTTP returns `200 OK` for generated-content validation failures and exposes the failed status in the response body.
- Runtime validation failures are errors.
Scriptorium does not persist run state. A retry can produce different output and
can incur another provider request. CLI exit behavior belongs to the
[CLI reference](../cli.md); HTTP status behavior belongs to the
[HTTP API reference](../api.md); package errors and results belong to the
[package contract](pkg-scriptorium.md).

View File

@@ -1,4 +1,4 @@
# Package scriptorium
# Package `scriptorium`
Import path:
@@ -6,250 +6,158 @@ Import path:
import "gitea.maximumdirect.net/eric/scriptorium"
```
The root package is the public Go facade for Scriptorium's prompt prepare/run
workflow. It exposes typed requests, results, source options, injected LLM
clients, and stable public errors while keeping `internal/*` packages private.
This is the canonical public Go contract for in-process prompt preparation and
execution. Prompt, profile, and schema file formats are defined in the
[configuration reference](../config.md).
## Intended Use Cases
## Engine Construction
Use the package when a Go application needs:
`NewEngine(Config, ...Option)` constructs an engine. `Config` has these
fields:
- in-process prompt preparation or execution;
- typed request/result structs;
- direct `context.Context` cancellation;
- injected/fake LLM clients for tests;
- direct per-request `RunRequest.APIKey`.
| Field | Meaning |
| --- | --- |
| `PromptDir` | Prompt-definition directory, required unless a prompt source option is supplied. |
| `ProfileDir` | Optional custom profile directory over built-ins. |
| `SchemaDir` | Schema directory; empty uses `.`. |
| `Timeout` | Default timeout for the built-in OpenAI-compatible client. |
| `HTTPClient` | Optional HTTP client for that built-in client. |
Use [Subprocess integration](../integrations/subprocess.md) or the [HTTP API](../api.md)
when a process or service boundary is preferred.
Nil options are ignored. Invalid construction, including
`WithLLMClient(nil)`, returns an error matching `ErrInvalidConfig`.
## Construct An Engine
Source options replace their matching directory source:
- prompts: `WithPromptFS(fsys, root)`, `WithPromptFile(path)`;
- profiles: `WithProfileFS(fsys, root)`, `WithProfileFile(path)`, and
`WithProfiles(profiles...)`;
- schemas: `WithSchemaFS(fsys, root)`, `WithSchemaFile(path)`; and
- LLM client: `WithLLMClient(client)`.
`fs.FS` prompt-content and schema paths stay inside their configured roots.
A single-file option exposes that file by its base name. In-memory profiles take
precedence over an explicit or directory-backed profile source, which in turn
takes precedence over built-ins. File and filesystem sources use the format and
credential rules in the [configuration reference](../config.md).
## Prepare And Run
`Prepare(ctx, request)` resolves the prompt, profile, input artifacts,
validation contract, and rendered messages without calling an LLM.
`Run(ctx, request)` performs that preparation, calls the configured client,
and validates generated content.
```go
engine, err := scriptorium.NewEngine(scriptorium.Config{
PromptDir: "./examples/prompts",
ProfileDir: "./examples/profiles",
SchemaDir: "./examples/schemas",
})
if err != nil {
return err
}
```
`Config` fields:
| Field | Description |
| --- | --- |
| `PromptDir` | Prompt definition directory. Required unless `WithPromptFS` or `WithPromptFile` is used. |
| `ProfileDir` | Optional custom profile directory overlaid above built-in profiles. |
| `SchemaDir` | Schema directory. Defaults to `.` when empty. |
| `Timeout` | Default timeout for the built-in OpenAI-compatible client. |
| `HTTPClient` | Optional HTTP client for the built-in OpenAI-compatible client. |
`NewEngine` accepts `nil` options and ignores them. Invalid construction wraps
`ErrInvalidConfig`.
## Source Options
Directory fields are the compatibility path. Explicit source options override
the matching directory field.
Prompt sources:
- `WithPromptFS(fsys, root)`
- `WithPromptFile(path)`
Profile sources:
- `WithProfileFS(fsys, root)`
- `WithProfileFile(path)`
- `WithProfiles(profiles...)`
Schema sources:
- `WithSchemaFS(fsys, root)`
- `WithSchemaFile(path)`
LLM source:
- `WithLLMClient(client)`
Source behavior:
- Prompt and profile YAML use the same strict rules as directory loading.
- Prompt `content_file` values resolve relative to the prompt file.
- `fs.FS` roots are containment boundaries for prompt content files and schema paths.
- File options expose the selected file by its base name.
- Profile source precedence is in-memory profiles, then explicit profile file/FS/directory source, then built-ins.
- `WithLLMClient(nil)` returns `ErrInvalidConfig`.
## In-Memory Profiles
Use `WithProfiles` when the application already has typed model settings:
```go
profile := scriptorium.OpenAICompatibleProfile(scriptorium.OpenAICompatibleProfileConfig{
ID: "app.default",
Endpoint: "https://openrouter.ai/api/v1",
Model: "mistralai/mistral-small-3.2-24b-instruct",
APIKeyRequired: true,
})
engine, err := scriptorium.NewEngine(cfg, scriptorium.WithProfiles(profile))
```
`Profile` and `OpenAICompatibleProfileConfig` include:
- `ID`
- `Endpoint`
- `Model`
- `Temperature`
- `MaxTokens`
- `TopP`
- `TimeoutSeconds`
- `ServiceTier`
- `ReasoningEffort`
- `APIKeyRequired`
- `ExtraParams`
`WithProfiles` rejects duplicate IDs in one call. In-memory profiles do not
store raw keys. When `APIKeyRequired` is true, pass the secret on each request
with `RunRequest.APIKey`.
`ExtraParams` must be JSON-compatible: strings, booleans, finite numbers,
objects with string keys, arrays/slices, and nil. Unsupported values, non-string
map keys, non-finite floats, and cycles return `ErrInvalidConfig` for profiles
or `ErrInvalidRequest` for request overrides.
## Prepare Workflow
`Prepare` resolves prompt/profile/input/schema state and renders messages
without calling an LLM.
```go
prepared, err := engine.Prepare(ctx, scriptorium.RunRequest{
PromptID: "generic.markdown_summary",
Inputs: map[string]scriptorium.ArtifactRef{
"transcript": scriptorium.File("./examples/fixtures/transcript.md"),
"glossary": scriptorium.File("./examples/fixtures/glossary.yml"),
},
})
if err != nil {
return err
}
_ = prepared.EffectiveModelParams
_ = prepared.Messages
```
`PreparedRun` includes prompt ID/version/hash, selected profile, effective
model params, output contract, structured-output metadata, input hashes,
rendered prompt hash, rendered messages, and timing fields. It does not include
raw API-key values, model output, validation results, or internal target
presence metadata.
The maintained package example is
[`examples/go-library/prepare`](../../examples/go-library/prepare).
## Run Workflow
`PreparedRun` exposes prompt, selected-profile, effective-model, output
contract, structured-output, input-hash, rendered-message, and timing
information. It does not include a resolved API key, model output, validation
result, or target-presence metadata.
`Run` calls `Prepare`, invokes the configured LLM client, builds the output
artifact, and validates the output.
`RunResult` adds run ID, artifact, raw output, validation, model metadata,
usage, and duration. Generated-content validation failures return a result with
`Validation.Status == ValidationFailed`; schema or validator runtime failures
return an error matching `ErrValidation`.
```go
result, err := engine.Run(ctx, scriptorium.RunRequest{
PromptID: "generic.markdown_summary",
APIKey: apiKey,
Inputs: map[string]scriptorium.ArtifactRef{
"transcript": scriptorium.File("./examples/fixtures/transcript.md"),
"glossary": scriptorium.File("./examples/fixtures/glossary.yml"),
},
})
if err != nil {
return err
}
_ = result.Artifact
```
## Public Values
`RunResult` includes run ID, output artifact, raw output, validation result,
prompt/profile/model metadata, effective model params, input hashes, usage, and
timing fields.
`ArtifactRef` has `Type`, `URI`, and `Body`; `Artifact` has `Name`,
`ContentType`, `Body`, `URI`, `Size`, and `Hash`. `ExecutionTarget` exposes the
effective endpoint, model, numeric settings, credential-environment name,
service tier, reasoning effort, and extra parameters. `ValidationResult`
contains status, mode, errors, schema path, repair attempts, and validity.
Generated-content validation failures return a successful `RunResult` with
`Validation.Status == ValidationFailed`. Runtime/schema validation errors
return an error that matches `ErrValidation`.
The exported constants define these serialized values:
## Inputs
- artifact types: `inline` and `file`;
- output formats: `text`, `markdown`, and `json`;
- validation modes: `none`, `basic`, `json`, and `json_schema`; and
- validation statuses: `passed`, `failed`, and `skipped`.
Input helpers:
`TokenUsage` reports prompt, completion, total, cached, and cache-write token
counts. `RenderedPrompt`, `RenderedMessage`, `CacheControl`, and
`StructuredOutputSpec` are the public shapes used by injected LLM clients.
- `File(path)`: file-backed artifact reference.
- `Inline(body)`: inline artifact body.
- `InlineWithURI(uri, body)`: inline artifact body with URI metadata.
## Requests, Inputs, And Overrides
Input map keys must match the prompt's expected input names.
`RunRequest` fields are `PromptID`, `PromptVersion`, `ProfileID`,
`APIKey`, `Inputs`, `Vars`, `Execution`, `Validation`, and
`Metadata`.
Input helpers are:
- `File(path)` for a file-backed artifact;
- `Inline(body)` for inline content; and
- `InlineWithURI(uri, body)` for inline content with URI metadata.
Required declared inputs must be supplied. Template rendering must also resolve
every input name the prompt actually references. Extra entries in `Inputs`
are not rejected solely because they are undeclared.
`ExecutionTargetOverride` supplies endpoint, model, credential-environment,
service-tier, reasoning-effort, and extra-parameter overrides. Its numeric
fields (`Temperature`, `MaxTokens`, `TopP`, and `TimeoutSeconds`) are
pointers so explicit zero values are preserved. `OutputContract` supplies
`Format`, `ValidationMode`, `SchemaPath`, and `RepairAttempts`.
`ExtraParams` accepts JSON-compatible values: strings, booleans, finite
numbers, objects with string keys, arrays or slices, and nil. Unsupported
values, non-string map keys, non-finite floats, and cycles return
`ErrInvalidConfig` for profiles or `ErrInvalidRequest` for request
overrides.
## Profiles And Credentials
`OpenAICompatibleProfile(OpenAICompatibleProfileConfig)` creates an
in-memory `Profile`. Its public fields are `ID`, `Endpoint`, `Model`,
`Temperature`, `MaxTokens`, `TopP`, `TimeoutSeconds`, `ServiceTier`,
`ReasoningEffort`, `APIKeyRequired`, and `ExtraParams`.
`WithProfiles` rejects duplicate IDs in one call.
A direct `RunRequest.APIKey` is request-scoped and takes precedence over
`api_key_env` for the built-in client. It is excluded from JSON output and
from `PreparedRun` and `RunResult`. The package's `String` and
`GoString` methods report only whether a direct key is set. Do not use
reflection-based dumps of request structs, which can bypass that redaction.
## Injected LLM Clients
Use `WithLLMClient` for tests or custom model integrations:
`LLMClient` implements:
```go
type fakeLLM struct{}
func (fakeLLM) Generate(ctx context.Context, req scriptorium.GenerateRequest) (*scriptorium.GenerateResponse, error) {
return &scriptorium.GenerateResponse{
Content: "generated text",
Usage: scriptorium.TokenUsage{TotalTokens: 12},
}, nil
}
engine, err := scriptorium.NewEngine(cfg, scriptorium.WithLLMClient(fakeLLM{}))
Generate(context.Context, GenerateRequest) (*GenerateResponse, error)
```
Injected clients receive:
- rendered prompt;
- effective execution target;
- numeric target presence metadata;
- structured-output spec when applicable;
- direct request API key when provided.
Custom clients should not log raw prompts or API keys by default.
## Overrides And API Keys
`RunRequest` fields:
| Field | Description |
| --- | --- |
| `PromptID` | Prompt ID. |
| `PromptVersion` | Optional prompt version filter. |
| `ProfileID` | Optional profile override. |
| `APIKey` | Direct per-request API key. |
| `Inputs` | Input artifact references. |
| `Vars` | Template variables. |
| `Execution` | Per-request model overrides. |
| `Validation` | Per-request output contract override. |
| `Metadata` | Request metadata reserved for callers. |
`RunRequest.Execution` uses pointer fields for numeric values so explicit zero
overrides are preserved:
```go
zero := 0
req.Execution = &scriptorium.ExecutionTargetOverride{
MaxTokens: &zero,
}
```
Direct `RunRequest.APIKey` takes precedence over profile `api_key_env` for the
default OpenAI-compatible client. It is request-scoped, uses `json:"-"`, and is
not included in `PreparedRun` or `RunResult` JSON. Normal Go string formatting
of `RunRequest` and `GenerateRequest` reports only whether a direct key is set.
Raw API keys do not belong in profile YAML, in-memory profiles, or app config.
Avoid reflection-based debug dumps of request structs because exported fields
remain visible to tools that bypass `String` and `GoString`.
Injected clients receive the rendered prompt, effective execution target, numeric
target-presence metadata, optional structured-output specification, and direct
request API key. `GenerateResponse` returns content and `TokenUsage`.
Custom clients should avoid logging raw prompts or credentials.
## Errors
Public methods wrap context while preserving stable sentinel checks with
`errors.Is`:
Public methods preserve these sentinel checks through `errors.Is`:
- `ErrInvalidConfig`
- `ErrInvalidRequest`
@@ -262,23 +170,5 @@ Public methods wrap context while preserving stable sentinel checks with
- `ErrLLMGenerate`
- `ErrValidation`
Example:
```go
if errors.Is(err, scriptorium.ErrPromptNotFound) {
return err
}
```
## Examples
Run the maintained prepare-only example from the repository root:
```bash
go run ./examples/go-library/prepare
```
See also:
- [Configuration reference](../config.md)
- [Consumer integration overview](api.md)
For interface selection and operational responsibilities, see the
[consumer integration overview](api.md).

View File

@@ -18,6 +18,8 @@ Start with:
- [Architecture policy](policy/architecture.md) for system boundaries,
invariants, and non-goals;
- [Internal component overview](internal/overview.md) for the current package
and component map;
- [Documentation policy](policy/documentation.md) before changing
documentation;
- [Testing policy](policy/testing.md) before adding, rewriting, or deleting
@@ -27,17 +29,18 @@ Start with:
| Task | Read before changing |
| --- | --- |
| Public Go package or engine behavior | [Go package consumer contract](consumers/pkg-scriptorium.md), [runner internals](internal/runner.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
| CLI commands, flags, output, or exit behavior | [CLI contract](cli.md) and [adapter internals](internal/adapters.md) |
| HTTP routes, DTOs, limits, or status mapping | [HTTP API contract](api.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
| Application configuration | [Configuration contract](config.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
| Prompt, profile, schema, or artifact loading | [Configuration contract](config.md) and [source internals](internal/sources.md) |
| Repository orientation or component responsibility | [Internal component overview](internal/overview.md) and [architecture policy](policy/architecture.md) |
| Public Go package or engine behavior | [Go package consumer contract](consumers/pkg-scriptorium.md), [internal component overview](internal/overview.md), [runner internals](internal/runner.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
| CLI commands, flags, output, or exit behavior | [CLI contract](cli.md), [internal component overview](internal/overview.md), and [adapter internals](internal/adapters.md) |
| HTTP routes, DTOs, limits, or status mapping | [HTTP API contract](api.md), [internal component overview](internal/overview.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
| Application configuration | [Configuration contract](config.md), [internal component overview](internal/overview.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
| Prompt, profile, schema, or artifact loading | [Configuration contract](config.md), [internal component overview](internal/overview.md), and [source internals](internal/sources.md) |
| Runner orchestration, rendering, validation, or repair | [Runner internals](internal/runner.md) and [source internals](internal/sources.md) |
| OpenAI-compatible request or response behavior | [OpenAI-compatible integration](integrations/openai-compatible-chat.md), [runner internals](internal/runner.md), and [adapter internals](internal/adapters.md) |
| OpenAI-compatible request or response behavior | [OpenAI-compatible integration](integrations/openai-compatible-chat.md), [LLM internals](internal/llm.md), [runner internals](internal/runner.md), and [adapter internals](internal/adapters.md) |
| Subprocess behavior | [Subprocess integration](integrations/subprocess.md) and [CLI contract](cli.md) |
| Runtime operation, recovery, or troubleshooting | [Operations](operations.md) and [troubleshooting](troubleshooting.md) |
| Runtime operation or recovery | [Operations](operations.md) |
| Examples or copyable assets | The owning contract for the demonstrated behavior and the related files under `examples/` |
| Architecture decisions or future work | The [documentation policy](policy/documentation.md), relevant accepted ADRs under `adr/`, and relevant roadmap documents under `roadmap/` |
| Architecture decisions or future work | The [documentation policy](policy/documentation.md), relevant accepted ADRs such as [ADR 0001](adr/0001-adopt-canonical-documentation-ownership.md), and relevant roadmap documents under `roadmap/` |
For cross-cutting changes, follow every applicable row. Internal component
documents own detailed subsystem change recipes.

View File

@@ -1,118 +1,67 @@
# OpenAI-Compatible Chat Integration
## Scope
This is the outbound wire contract for Scriptorium's OpenAI-compatible
chat-completions client.
This document defines the outbound LLM contract implemented by `internal/llm/openai_compatible_client.go`.
## Endpoint And Method
It documents only fields and behaviors currently serialized by code.
Scriptorium uses the request endpoint override when present; otherwise it uses
the configured client base URL. It removes a trailing slash and sends
`POST /chat/completions`.
## Endpoint Construction
For example, `http://localhost:8000/v1` becomes
`http://localhost:8000/v1/chat/completions`.
Request endpoint is built as:
## Request Payload
1. choose base URL:
- `GenerateRequest.Target.Endpoint` if set
- otherwise client config `BaseURL`
2. trim trailing slash
3. append `/chat/completions`
The payload always contains `model` and rendered `messages`. It additionally
contains these fields when applicable:
Example:
| Field | Inclusion |
| --- | --- |
| `session_id` | Non-empty rendered prompt session ID. |
| `temperature` | Non-zero effective value or an explicit zero override. |
| `max_tokens` | Non-zero effective value or an explicit zero override. |
| `top_p` | Non-zero effective value or an explicit zero override. |
| `service_tier` | Any non-empty configured value. |
| `reasoning_effort` | Any non-empty configured value. |
| `response_format` | Structured output is requested. |
| provider-specific fields | Flattened from `extra_params`. |
- base URL: `http://localhost:8000/v1`
- final URL: `http://localhost:8000/v1/chat/completions`
`service_tier` and `reasoning_effort` are forwarded without a provider value
catalog; the selected backend decides which values it supports.
## Request Fields Sent
`extra_params` are top-level JSON fields, not a nested object. Keys cannot be
empty or collide with `model`, `session_id`, `messages`, `temperature`,
`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or
`response_format`. Values must be JSON-serializable.
Serialized JSON fields:
A rendered `session_id` is sent as a top-level JSON field, not as a header.
Empty values are omitted. The maximum length is 256 Unicode code points.
- `model` (required after fallback resolution)
- `session_id` (only when the rendered prompt includes a non-empty session ID)
- `messages` (rendered prompt messages)
- `temperature` (when non-zero, or when explicitly overridden to zero)
- `max_tokens` (when non-zero, or when explicitly overridden to zero)
- `top_p` (when non-zero, or when explicitly overridden to zero)
- `service_tier` (only when non-empty)
- `reasoning_effort` (only when non-empty)
- `response_format` (only when structured output is provided)
- profile/request `extra_params` as additional provider-specific top-level fields
`service_tier` is provider-specific. OpenRouter currently documents request values such as `flex` and `priority`; Scriptorium forwards any non-empty configured value and lets the backend validate support.
`reasoning_effort` is provider-specific. Scriptorium forwards any non-empty configured value as top-level `reasoning_effort` and lets the backend validate support.
`extra_params` are flattened into the outbound JSON object. They are not wrapped in an `extra_params` object:
```json
{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "rendered text"
}
],
"provider_route": "primary",
"provider_options": {
"retry_budget": 2
}
}
```
`extra_params` values must be JSON-compatible. Supported value shapes include strings, numbers, booleans, objects, and arrays.
Reserved `extra_params` keys are rejected before the HTTP request is made:
- `model`
- `session_id`
- `messages`
- `temperature`
- `max_tokens`
- `top_p`
- `service_tier`
- `reasoning_effort`
- `response_format`
Empty `extra_params` keys and values that cannot be encoded as JSON are also rejected before the HTTP request is made.
`session_id` is rendered from prompt YAML using request variables and serialized as a top-level JSON request field. Scriptorium does not send an `x-session-id` header. Empty rendered session IDs are omitted, and values longer than 256 characters are rejected before the HTTP request.
Messages without prompt cache control serialize with string `content`:
Messages without cache control use string `content`. A message with cache
control uses one text block:
```json
{
"role": "system",
"content": "rendered text"
}
```
Messages with prompt cache control serialize as a single text content-block array:
```json
{
"role": "system",
"content": [
{
"content": [{
"type": "text",
"text": "rendered text",
"cache_control": {
"type": "ephemeral",
"ttl": "1h"
}
}
]
"cache_control": {"type": "ephemeral", "ttl": "1h"}
}]
}
```
When cache-control `ttl` is unset in the prompt definition, `ttl` is omitted from the outbound payload.
Structured output is currently `json_schema` only, serialized as:
When the prompt omits cache-control `ttl`, the payload omits `ttl`.
Structured JSON Schema output is sent as:
```json
{
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "...",
"name": "schema name",
"strict": true,
"schema": {"type": "object"}
}
@@ -120,74 +69,40 @@ Structured output is currently `json_schema` only, serialized as:
}
```
## Authentication Header
## Authentication And Timeout
If `Target.APIKey` is set:
When a direct request API key is present, Scriptorium sends
`Authorization: Bearer <key>` and does not read `api_key_env`. Otherwise, it
resolves the configured non-empty `api_key_env` at request time and sends the
same header. If neither mechanism supplies a key, it sends no
`Authorization` header.
- set `Authorization: Bearer <value>`
- do not read `Target.APIKeyEnv`
The configured client timeout applies by default. A positive effective
`timeout_seconds` replaces it. An explicit request override of zero disables
the HTTP-client timeout; negative values are rejected before a request is sent.
If `Target.APIKey` is empty and `Target.APIKeyEnv` is set:
## Response Subset And Failures
- resolve environment variable value at request time
- set `Authorization: Bearer <value>`
A successful provider response must supply non-empty
`choices[0].message.content`. Scriptorium reads these optional or required
usage fields when present:
If the environment variable is unset/empty:
- request fails before HTTP call (`ErrInvalidRequest`)
If both `Target.APIKey` and `Target.APIKeyEnv` are empty:
- no `Authorization` header is sent
## Timeout Behavior
Base timeout comes from client configuration.
Per-request override:
- if `Target.TimeoutSeconds > 0`, use that value for request timeout
- if `Target.TimeoutSeconds == 0` and the value came from an explicit request override, disable the HTTP client timeout
- if `Target.TimeoutSeconds < 0`, request is rejected (`ErrInvalidRequest`)
## Response Expectations
Expected successful response shape (subset used):
- `choices[0].message.content`
- `usage.prompt_tokens`
- `usage.completion_tokens`
- `usage.total_tokens`
- `usage.prompt_tokens_details.cached_tokens` (optional)
- `usage.cache_write_tokens` (optional)
- `usage.prompt_tokens_details.cached_tokens`
- `usage.cache_write_tokens`
Absent cache usage fields are treated as zero. Parsed cache usage is exposed through run results and adapter response surfaces as:
Missing cache usage is reported as zero. Invalid JSON, an empty choices array,
or empty first-choice content is a malformed provider response. Network and
request-construction failures, non-2xx responses, and malformed responses fail
the outbound call. Provider response bodies are not exposed by this client.
- `cached_tokens`
- `cache_write_tokens`
The client does not implement built-in retries, tool calls, top-level
`cache_control`, or multi-request payload modes.
Malformed response conditions include:
## Related References
- invalid JSON
- empty `choices`
- empty `choices[0].message.content`
Malformed responses return `ErrMalformedResponse`.
## Error Handling
- network/request-construction failures: `ErrRequestFailed`
- non-2xx HTTP status: `ErrUnexpectedStatus` (includes status code; provider response bodies are not included)
- malformed response shape/content: `ErrMalformedResponse`
## Unsupported Or Non-Serialized Fields
The client does not serialize top-level `cache_control`.
No built-in retries, tool-calls, or multi-request payload modes are implemented in this client.
## Relationship To Runner
When prompt validation mode is `json_schema`, runner prepares a structured-output schema spec and passes it to the client as `StructuredOutput`.
The client only serializes the provider request payload; it does not load schema files itself.
Prompt schema preparation and runner orchestration are described in
[runner internals](../internal/runner.md). Prompt and profile configuration is
defined by the [configuration reference](../config.md).

View File

@@ -1,130 +1,41 @@
# Subprocess Integration
This document defines the supported subprocess contract for downstream
applications invoking Scriptorium through the public CLI.
This document covers process-boundary behavior for callers that invoke
Scriptorium as a child process. Command syntax, flags, output, and exit codes
are defined by the [CLI reference](../cli.md). Interface selection belongs in
the [consumer integration overview](../consumers/api.md).
This is a CLI contract. Go callers that want an in-process typed API should use
the [package guide](../consumers/pkg-scriptorium.md).
## Process Contract
## Supported Commands
Use `scriptorium render` when the caller needs prepared output without a model
call, and `scriptorium run` for generation. Pass an explicit `--config` or
make the configuration search paths available to the child process; configuration
discovery, fields, profile selection, and credential mechanisms are defined in
the [configuration reference](../config.md).
Downstream applications should invoke:
Pass required API-key environment variables through the child environment. Do
not place raw API keys in arguments. Keep the environment limited to the values
needed for the selected profile.
- `scriptorium render` for preflight/debug output without LLM execution.
- `scriptorium run` for generation.
## Streams And Output Ownership
`scriptorium serve` is an HTTP service command, not the recommended subprocess
contract for per-request execution.
Capture stdout and stderr separately. Stdout contains the requested artifact or
prepared output unless the caller selects an output file; stderr contains
summaries, diagnostics, and server messages. The exact destinations and status
meanings are part of the [CLI reference](../cli.md), not a stable stderr data
protocol.
## Recommended Invocation Shapes
When using `--out`, the caller owns the output path, its permissions, and
cleanup. Treat rendered prompts, generated artifacts, stdout, and stderr as
potentially sensitive.
Render:
## Cancellation And Recovery
```bash
scriptorium render \
--config <config_path> \
--prompt <prompt_id> \
--input transcript=<path> \
--format json
```
A CLI invocation performs one synchronous request and creates no durable run
state. A supervising process that needs cancellation must terminate the child
process according to its own process-management policy. A later invocation is a
new request and can make another model call; there is no resume or checkpoint
protocol.
Run:
```bash
scriptorium run \
--config <config_path> \
--prompt <prompt_id> \
--input transcript=<path> \
--out <artifact_path>
```
Callers may add:
- `--profile <profile_id>`
- repeatable `--input name=path`
- repeatable `--var name=value`
- runtime overrides when explicitly needed, such as `--model`, `--llm-base-url`, `--api-key-env`, and `--timeout`
Do not pass raw API keys as command arguments.
## Config And Directory Behavior
Callers can rely on resolved app config or pass explicit paths.
Default config search order:
1. `/usr/local/etc/scriptorium/config.yml`
2. `/etc/scriptorium/config.yml`
Rules:
- Explicit `--config` requires file existence and valid syntax.
- CLI flags override config values.
- `run` and `render` require an effective `prompt_dir`.
- `profile_dir` is optional because built-in profiles are available.
## Profile Selection
Profile selection follows runner behavior:
1. explicit `--profile`
2. prompt `default_profile`
3. error if neither is available
Treat prompt and profile IDs as deployment configuration, not hardcoded business
logic.
## Input And Variable Contract
- Inputs use repeated `--input name=path`.
- Input names must match prompt definition input names.
- Variables use repeated `--var name=value`.
- Both flags also accept comma-separated mappings.
- Prefer file inputs for large content.
CLI inputs are file references. HTTP-only `inline` references are documented in
the [HTTP API reference](../api.md).
## Environment Contract
- Pass through required API-key environment variables referenced by `api_key_env`.
- Keep subprocess environments scoped to required variables.
- Use `--api-key-env` only to name an environment variable.
- Never pass raw API keys via argv.
## Stdout And Stderr
`run`:
- stdout: generated artifact body unless `--out` is used.
- stderr: success summary and errors.
`render`:
- stdout: prepared-run output unless `--out` is used.
- stderr: errors.
Capture stdout and stderr separately. Do not parse stderr as a stable data
format beyond exit status handling.
## Exit Status Contract
- `0`: success.
- `1`: parse, config, load, render, generation, IO, or runtime error.
- `2`: `run` completed and output was written, but validation failed.
A `run` exit code `2` can still produce output on stdout or at `--out`.
Consumers must decide whether to keep or discard that output.
## Security Notes
- Treat generated artifacts, rendered prompts, stdout, and stderr as potentially sensitive.
- Use controlled output paths and access controls for persisted artifacts.
- Avoid logging full rendered prompts or generated artifacts by default.
## Canonical References
- CLI behavior: [CLI reference](../cli.md)
- Config and file formats: [Configuration reference](../config.md)
- Operations: [Operations guide](../operations.md)
- Troubleshooting: [Troubleshooting](../troubleshooting.md)
For deployment, filesystem permissions, and sensitive-artifact handling, see
the [operations guide](../operations.md).

View File

@@ -2,154 +2,112 @@
## Purpose
Adapters translate external interfaces into domain requests and translate domain results back out. They wire dependencies, apply app config, and own IO concerns, but they do not make runner decisions.
Adapters translate external inputs into domain requests, compose dependencies,
and translate domain results or errors back to their interface. They own IO and
presentation mechanics; use-case decisions remain in `internal/usecase`.
Source-loading behavior belongs in `docs/internal/sources.md`. User-facing CLI, HTTP, and package contracts belong in `docs/cli.md`, `docs/api.md`, and `docs/consumers/pkg-scriptorium.md`.
External contracts are canonical in the [CLI reference](../cli.md), [HTTP API
reference](../api.md), and [Go package contract](../consumers/pkg-scriptorium.md).
## Adapter Map
## Components And Collaborators
- `cmd/scriptorium`: process entrypoint.
- `internal/adapter/cli`: command parsing, config handoff, runner construction, stdout/stderr, exit codes.
- `internal/adapter/http`: `POST /v1/runs` request/response mapping and HTTP error/status mapping.
- root package `scriptorium`: public Go facade over internal runner types and dependencies.
- `cmd/scriptorium` passes process arguments and streams to
`internal/adapter/cli`.
- `internal/adapter/cli` parses commands, resolves application settings through
`internal/config`, constructs a runner, and owns process output handling.
- `internal/adapter/http` decodes DTOs, maps them to `domain.RunRequest`, calls
a runner interface, and maps errors and results to HTTP DTOs.
- The root `scriptorium` package maps its public types and options to internal
collaborators and maps selected internal errors to public sentinels.
- `internal/format` formats prepared runs for the CLI; `internal/llm`,
`internal/prompt`, and source packages supply runner dependencies.
Supporting implementation packages used during adapter wiring:
## Wiring Flows
- `internal/config`
- `internal/defaults`
- `internal/format`
- `internal/llm`
- `internal/prompt`
### CLI
## Inputs And Outputs
The CLI resolves configuration before constructing dependencies. `run` builds a
runner with the ordinary composite artifact reader and invokes `Runner.Run`;
`render` uses the same wiring and invokes `Runner.Prepare`; `serve` replaces the
file reader with the restricted artifact reader, builds an HTTP handler, and
starts the server.
CLI adapter:
Parser state records whether numeric runtime values were explicitly supplied.
That presence is carried into `domain.ExecutionTargetOverride`, allowing the
runner to distinguish omitted values from explicit zero overrides.
- Input: process args, optional config file, filesystem sources, environment variables.
- Output: process exit code, stdout artifact/prepared output, stderr summaries and errors.
### HTTP
HTTP adapter:
The handler first enforces transport limits, strict JSON decoding, and the
minimal request shape. It maps DTO values to domain types without deciding
prompt selection, source behavior, or validation semantics. On success it maps
the domain result to the response DTO; on failure it uses `errors.Is` over
runner, source, artifact, and profile errors to choose the public error mapping.
- Input: HTTP request method/path/headers/body for `POST /v1/runs`.
- Output: JSON success or error body with mapped status code.
The [HTTP API reference](../api.md) owns the route, DTO schema, status codes,
and externally observable limit behavior.
Public Go facade:
### Public Go Facade
- Input: typed `scriptorium.Config`, `Option`, and `RunRequest` values.
- Output: typed `PreparedRun` and `RunResult` values plus public sentinel errors.
`NewEngine` applies public options, selects filesystem, `fs.FS`, single-file,
or in-memory dependencies, and constructs a runner. The conversion functions
copy maps and slices across the boundary so callers do not receive internal
domain values. The facade maps selected internal errors to the public sentinel
set and keeps direct request API keys out of public results.
## Boundaries
## Package-Local Guarantees
- Adapters convert external shapes to `domain.RunRequest` and back.
- Runner orchestration remains in `internal/usecase`.
- Prompt/profile/schema/artifact source rules remain in repository, validator, and artifact packages.
- LLM provider request serialization remains in `internal/llm`.
- Public package types are facade types; internal domain types do not leak across the package boundary.
- Adapters do not embed runner orchestration or source-loading decisions.
- Configuration is resolved before adapter dependency composition.
- CLI and HTTP create runners without a repairer; a repairer is available only
through explicit internal runner construction.
- DTO conversion preserves explicit numeric-override presence.
- Error mapping matches error identities, not error text.
- No adapter creates durable run state; caller-selected output files are not
application state.
## Config Fields Used
## Failure And Verification Boundaries
Adapter app settings:
Keep external error payloads concise, preserve strict external decoding, and do
not serialize resolved secret values. Validation content failures remain result
state; runtime failures remain errors for the relevant adapter to map.
- `prompt_dir`
- `profile_dir`
- `schema_dir`
- `server.addr`
- `server.artifact_root`
- `server.max_request_bytes`
- `server.max_artifact_bytes`
- `server.max_response_bytes`
- `defaults.render_format`
Execution request/profile settings passed through the runner:
- `endpoint`
- `model`
- `temperature`
- `max_tokens`
- `top_p`
- `timeout_seconds`
- `service_tier`
- `api_key_env`
- `reasoning_effort`
- `extra_params`
CLI and HTTP preserve numeric override presence so omitted values and explicit zero values remain distinct.
## CLI Adapter
Implemented commands:
- `run`
- `render`
- `serve`
Behavior:
- `run` constructs a runner with direct filesystem artifact reading and calls `Runner.Run`.
- `render` constructs a runner and calls `Runner.Prepare`; it does not call the LLM.
- `serve` constructs a restricted artifact reader and HTTP handler, then starts an unauthenticated HTTP server.
- `run` exits `2` when generation succeeds but validation fails.
- parse, runtime, and output-write errors exit `1`.
- deprecated `--prompt-id` and `--profile-id` aliases are accepted.
## HTTP Adapter
Behavior:
- Accepts only `POST /v1/runs`.
- Decodes JSON strictly and rejects unknown fields and trailing JSON tokens.
- Rejects empty `prompt_id` and empty `inputs` before calling the runner.
- Does not accept raw API key values in the request body.
- Returns validation failures as `200` responses with failed validation details.
- Maps request-body, artifact, and encoded-response size failures to `413`.
- Maps domain and repository errors to stable error codes without returning wrapped internal cause text.
The HTTP adapter has no built-in authentication or authorization. Deployment controls must be provided outside the process.
## Public Go Facade
Behavior:
- `NewEngine` wires the same default runner components as CLI/HTTP unless options override them.
- Prompt, profile, and schema sources may come from directories, single files, or `fs.FS` roots.
- `WithProfiles` adds in-memory profiles ahead of file-backed and built-in profiles.
- `WithLLMClient` injects custom model behavior.
- `RunRequest.APIKey` is request-scoped and direct; it is used only for generation and is stripped from public results.
- internal errors are mapped to public sentinels in `errors.go`.
## Failure Behavior
Adapters should:
- keep external error payloads concise and stable.
- avoid leaking raw secret values.
- use sentinels and typed errors for mapping.
- preserve strict external input decoding.
- keep validation content failures distinct from runtime errors.
CLI writes human-readable summaries to stderr. HTTP writes JSON error envelopes. The public Go facade returns typed errors.
## State And Manifests
Adapters do not add durable run state.
- No adapter writes run manifests.
- No adapter implements checkpoint, skip, or resume behavior.
- CLI output files are caller-selected artifacts, not internal state.
## Tests To Inspect
Inspect focused tests when changing this area:
- `internal/adapter/cli/run_test.go`
- `internal/adapter/http/handler_test.go`
- `engine_test.go`
- `internal/format/prepared_run_test.go`
- `internal/llm/openai_compatible_client_test.go`
## Architectural Invariants
Run the affected adapter package tests and recheck the relevant canonical
contract. The [testing policy](../policy/testing.md) owns global test
sufficiency guidance.
- Adapter packages stay thin and translation-focused.
- App config is resolved before dependency construction.
- External input strictness is part of contract stability.
- CLI and HTTP construct runners without a repairer.
- HTTP endpoint details remain canonical in `docs/api.md`.
- Public Go package details remain canonical in `docs/consumers/pkg-scriptorium.md`.
## Change Recipes
### Application Configuration Fields
1. Add the field to the relevant `internal/config` shape and default handling.
2. Parse and validate it, then preserve configuration and CLI-override
precedence while wiring it through its consuming adapter.
3. Add focused configuration and adapter tests for parsing, mapping, and
effective behavior.
4. Update the [configuration contract](../config.md) and any affected external
contract.
### CLI Flags
1. Add the flag to the relevant parser in `internal/adapter/cli/run.go`.
2. Keep command scope and application-configuration precedence intentional.
3. Add or update parser and command tests in
`internal/adapter/cli/run_test.go`.
4. Update the [CLI contract](../cli.md) and affected maintained examples.
### Adapter Capabilities
1. Define or reuse the appropriate domain or use-case interface boundary.
2. Implement translation and IO behavior without moving use-case decisions out
of `internal/usecase`.
3. Add focused mapping, parsing, and error-behavior tests.
4. Update this document and the affected public or integration contract. Update
[source internals](sources.md) when source-loading behavior changes.

85
docs/internal/llm.md Normal file
View File

@@ -0,0 +1,85 @@
# LLM Internals
## Purpose
`internal/llm` defines the provider-neutral `Client` interface and the
OpenAI-compatible client implementation. The [OpenAI-compatible integration
contract](../integrations/openai-compatible-chat.md) owns the outbound HTTP wire
format and protocol behavior.
## Construction
`NewOpenAICompatibleClient` validates a non-empty configured base URL, records
an optional default model, and establishes the default timeout. A non-positive
configured timeout uses the internal default.
When callers supply an `http.Client`, construction clones it rather than
mutating the caller's instance. A supplied client with no timeout receives the
resolved default in the clone; a supplied non-zero timeout is retained. The
client stores the trimmed base URL, default model, timeout, and cloned client.
## Generate Flow
`Generate` receives a `domain.GenerateRequest` from the runner:
1. validate the effective timeout and choose the request endpoint;
2. map the domain request to the internal wire-request representation;
3. validate and flatten extra parameters, encode JSON, and create the HTTP
request;
4. prefer a direct API key, otherwise resolve the configured key environment
variable;
5. derive a request HTTP client when an explicit timeout changes the configured
client;
6. execute the request, reject non-success status responses without returning
provider response bodies; and
7. decode the response subset into `domain.GenerateResponse`.
`openAIChatRequestFromGenerateRequest` is the conversion boundary for effective
model defaults, explicit numeric-presence state, rendered messages, structured
output, and session-ID validation. `openAIChatRequestPayload` protects reserved
fields and JSON encoding before an HTTP call. The external payload shape is
defined only in the [integration contract](../integrations/openai-compatible-chat.md).
## Error Categories
The package uses these internal sentinels:
- `ErrInvalidConfig` for invalid client construction;
- `ErrInvalidRequest` for invalid effective generation input;
- `ErrRequestFailed` for request construction or transport failures;
- `ErrUnexpectedStatus` for non-success HTTP responses; and
- `ErrMalformedResponse` for invalid or incomplete successful-response data.
The runner maps an invalid LLM request to its invalid-request category and
other LLM failures to its generation category. Adapters then apply their public
error contracts.
## Package-Local Guarantees
- The default-model fallback happens before wire encoding.
- Per-request timeout handling clones a configured HTTP client when needed; it
does not mutate shared client state.
- Direct API keys take precedence over environment lookup within this client.
- Provider response bodies are discarded for non-success status responses.
- The client does not implement retries, tool calls, or a stateful session
store.
## Verification And Change Recipe
Inspect:
- `internal/llm/openai_compatible_client_test.go`
- `internal/usecase/runner_test.go`
- `internal/adapter/http/handler_test.go`
When changing the client:
1. keep domain-to-wire mapping inside `internal/llm` and preserve the `Client`
interface;
2. test construction, timeout selection, mapping, and error categorization;
3. update the [OpenAI-compatible integration contract](../integrations/openai-compatible-chat.md)
for any observable wire or protocol change; and
4. update [runner internals](runner.md) if the client boundary or structured
output handoff changes.
The [testing policy](../policy/testing.md) owns global test sufficiency.

48
docs/internal/overview.md Normal file
View File

@@ -0,0 +1,48 @@
# Internal Component Overview
## Purpose
This is the inventory of Scriptorium's implemented components for contributors.
The [architecture policy](../policy/architecture.md) owns normative boundaries
and invariants; public behavior belongs in the linked contracts.
## Public And Command Entrypoints
| Component | Implemented responsibility | References |
| --- | --- | --- |
| Root package `scriptorium` | Public Go facade that constructs the engine, exposes request/result types and options, and maps internal errors. | [Go package contract](../consumers/pkg-scriptorium.md), [adapter internals](adapters.md) |
| `cmd/scriptorium` | Process entrypoint that delegates command execution to the CLI adapter. | [CLI contract](../cli.md), [adapter internals](adapters.md) |
## Adapters, Domain, And Use Case
| Component | Implemented responsibility | References |
| --- | --- | --- |
| `internal/adapter/cli` | Parses CLI commands, applies application wiring, and handles process input and output. | [CLI contract](../cli.md), [adapter internals](adapters.md) |
| `internal/adapter/http` | Maps HTTP requests and responses to domain operations and maps public errors. | [HTTP API contract](../api.md), [adapter internals](adapters.md) |
| `internal/domain` | Defines core request, result, output-contract, and LLM-boundary types. | [runner internals](runner.md) |
| `internal/usecase` | Implements `Runner` preparation, execution, validation coordination, and the repairer boundary. | [runner internals](runner.md) |
## Configuration And Sources
| Component | Implemented responsibility | References |
| --- | --- | --- |
| `internal/config` | Loads application settings, applies defaults, and applies CLI overrides. | [configuration contract](../config.md), [adapter internals](adapters.md) |
| `internal/defaults` | Holds compile-time default values used when application settings are resolved. | [configuration contract](../config.md) |
| `internal/promptdef` | Loads prompt definitions from filesystem and `fs.FS` sources. | [configuration contract](../config.md), [source internals](sources.md) |
| `internal/profile` | Loads filesystem and `fs.FS` execution profiles and combines profile repositories. | [configuration contract](../config.md), [source internals](sources.md) |
| `internal/profile/builtin` | Provides embedded built-in execution profiles as a repository. | [configuration contract](../config.md), [source internals](sources.md) |
| `internal/filecatalog` | Provides shared YAML discovery and source-root helpers. | [source internals](sources.md) |
| `internal/artifact` | Reads inline and file-backed input artifacts. | [configuration contract](../config.md), [HTTP API contract](../api.md), [source internals](sources.md) |
| `internal/prompt` | Renders prompt templates into messages. | [runner internals](runner.md) |
## Formatting, Validation, And Model Access
| Component | Implemented responsibility | References |
| --- | --- | --- |
| `internal/format` | Formats prepared-run information for CLI output. | [CLI contract](../cli.md), [adapter internals](adapters.md) |
| `internal/validate` | Defines validation interfaces and provides standard filesystem and `fs.FS` schema validation. | [configuration contract](../config.md), [source internals](sources.md), [runner internals](runner.md) |
| `internal/llm` | Defines the provider-neutral LLM client boundary and its OpenAI-compatible implementation. | [OpenAI-compatible integration](../integrations/openai-compatible-chat.md), [LLM internals](llm.md), [runner internals](runner.md) |
Focused internal documents describe the components that have detailed
orchestration, adapter, or source behavior. Package tests live alongside the
implementation and are identified in those focused documents where relevant.

View File

@@ -2,145 +2,118 @@
## Purpose
`internal/usecase.Runner` is the core prompt-execution orchestrator. It prepares prompt requests, calls the configured LLM client for `Run`, validates generated output, and returns domain results.
`internal/usecase.Runner` is the prompt-execution orchestrator. It prepares
domain requests, invokes an injected LLM client, validates output, and returns
domain results. Transport parsing, response mapping, and public type conversion
remain outside this package.
Transport parsing, DTOs, CLI output, HTTP status mapping, and public package type conversion belong outside the runner.
The [configuration reference](../config.md) owns prompt, profile, schema, and
runtime-setting definitions. Public error behavior is defined by the
[HTTP API](../api.md) and [Go package](../consumers/pkg-scriptorium.md)
contracts.
## Inputs And Outputs
## Dependencies And Construction
Primary inputs:
`Runner` receives these collaborators:
- `domain.RunRequest`
- repositories/readers/renderers/validators injected at construction
- `context.Context` for cancellation
- `promptdef.Repository`;
- `profile.Repository`;
- `artifact.Reader`;
- `prompt.Renderer`;
- `llm.Client`;
- `validate.Validator`; and
- an optional `OutputRepairer`.
Primary outputs:
- `domain.PreparedRun` from `Prepare`
- `domain.RunResult` from `Run`
- wrapped sentinel errors for adapter mapping
LLM boundary types:
- `domain.GenerateRequest`
- `domain.GenerateResponse`
## Dependencies
`Runner` depends on package interfaces instead of concrete adapter types:
- `promptdef.Repository`
- `profile.Repository`
- `artifact.Reader`
- `prompt.Renderer`
- `llm.Client`
- `validate.Validator`
- optional `usecase.OutputRepairer`
The CLI, HTTP adapter, and public Go package construct these dependencies and pass them in.
## Config Fields
`Runner` does not read app config files. Effective behavior is determined by injected dependencies and the `domain.RunRequest`.
Adapter wiring commonly reflects these app config fields:
- `prompt_dir`
- `profile_dir`
- `schema_dir`
- `server.artifact_root`
- HTTP request/artifact/response size limits
Runtime model settings are resolved from the selected profile plus request overrides.
`NewRunner` constructs a runner without a repairer. `NewRunnerWithRepairer`
accepts one explicitly. Adapters and the public engine choose concrete
repositories and readers; the runner does not load application configuration.
## Prepare Flow
`Prepare`:
`Prepare` performs one deterministic preparation pass for a request:
1. requires a non-empty prompt ID.
2. loads the prompt definition and computes its hash.
3. selects the profile from request `profile_id`, then prompt `default_profile`.
4. loads the selected execution profile.
5. merges built-in execution defaults, profile values, and request overrides.
6. applies request-scoped direct API key values for public Go callers.
7. validates endpoint, model, and credential requirements.
8. resolves the output contract and JSON Schema document when required.
9. reads input artifacts.
10. renders prompt messages and hashes the rendered prompt.
11. returns a prepared run without calling the LLM.
1. validate the prompt ID and load the prompt definition;
2. hash the definition and select the explicit or default profile;
3. load the profile and resolve effective execution settings;
4. validate endpoint, model, and credential availability;
5. resolve the output contract and, for JSON Schema output, load a structured
schema document before model execution;
6. read and hash input artifacts;
7. render messages and the session ID; and
8. return a `PreparedRun` containing the effective state and rendered-prompt
hash.
Numeric request overrides are presence-aware: omitted values preserve the current effective value, while explicit zero values are real overrides.
Execution settings merge defaults, profile values, and a request override.
Numeric override presence is retained so explicit zero values are not confused
with omissions.
## Run Flow
## Run And Validation Flow
`Run`:
`Run` creates a run ID and timestamps, then calls `Prepare` rather than
duplicating preparation. It sends the prepared prompt, effective target,
target-presence state, and optional structured-output specification to the LLM
client. It converts the returned content to an output artifact, validates it,
and returns the artifact, validation, hashes, usage, and timing metadata.
1. creates a run ID and start timestamp.
2. calls `Prepare`.
3. calls the injected LLM client with rendered messages, effective target, target presence, and structured-output settings.
4. builds the output artifact.
5. validates the output.
6. optionally attempts bounded repair when a repairer is injected and the contract permits repair.
7. returns the run result with artifact, raw output, validation, hashes, selected profile/model metadata, usage, and timing.
A validator can return a content result or an operational error. Content
failures stay in the result; schema loading, compilation, and validator
operational failures are returned as `ErrValidation`. The canonical distinction
for callers is documented by the public contracts.
`Run` must reuse `Prepare`; prepare logic should not be duplicated elsewhere.
## Repair Boundary
## Validation And Repair
Repair is an internal optional loop. It starts only when a repairer is present,
the output contract permits one or more attempts, validation failed, and the
validation mode is JSON or JSON Schema. Each repair receives the previous
output, validation errors, effective target, structured-output specification,
and attempt metadata; every repaired result is validated again.
Validation content failures are returned as successful run results with `Validation.Status == failed`. They are not runtime errors.
`NewDefaultOutputRepairer` delegates to the injected LLM client. CLI, HTTP, and
the public engine use `NewRunner` and therefore do not inject this repairer.
Validation runtime failures, such as schema load or compile errors, return `ErrValidation`.
## Error Translation
Repair attempts occur only when all conditions are true:
- a repairer is injected
- `repair_attempts` is greater than zero
- validation status is `failed`
- validation mode is `json` or `json_schema`
CLI and HTTP wiring call `usecase.NewRunner(...)`, which does not inject a repairer. Normal CLI and HTTP execution therefore does not repair invalid output.
## Failure Behavior
Stable runner sentinels include:
Runner sentinels identify failure categories for adapters:
- `ErrInvalidRequest`
- `ErrProfileRequired`
- `ErrAPIKeyEnvMissing`
- `ErrAPIKeyRequired`
- `ErrPromptLoad`
- `ErrProfileLoad`
- `ErrArtifactLoad`
- `ErrAPIKeyEnvMissing` and `ErrAPIKeyRequired`
- `ErrPromptLoad`, `ErrProfileLoad`, and `ErrArtifactLoad`
- `ErrPromptRender`
- `ErrLLMGenerate`
- `ErrValidation`
Adapters should use `errors.Is` against sentinels and lower-level repository errors instead of matching message text.
Wrap errors with those sentinels and preserve their identities through
`errors.Is`; adapters must not classify errors by message text. The runner
passes direct keys only to the LLM boundary and never includes resolved key
values in prepared or run results.
Secret values must not appear in prepared output, run results, logs, HTTP responses, or serialized public package results. The effective API-key environment-variable name may appear.
## Package-Local Guarantees
## State And Manifests
- `Run` always reuses `Prepare`.
- Schema documents are loaded before the initial LLM call when structured output
is required.
- Output validation records attempts used, including repair attempts.
- Runner state is per request; the package does not create a durable run store
or manifest.
- Source, renderer, validator, and LLM implementations remain injected
boundaries.
The runner is stateless across requests.
## Verification And Change Recipe
- No durable run store.
- No manifest files.
- No checkpoint, skip, or resume behavior.
- Recovery is a new request after correcting inputs, config, or environment.
## Tests To Inspect
Inspect:
- `internal/usecase/runner_test.go`
- `internal/usecase/integration_test.go`
- `engine_test.go`
- `internal/adapter/cli/run_test.go`
- `internal/adapter/http/handler_test.go`
## Architectural Invariants
When changing orchestration:
- Use-case decisions stay in `internal/usecase`.
- `Run` reuses `Prepare`.
- Prompt/profile/artifact/schema loading remains behind injected boundaries.
- Validation content failures are result state; validation runtime failures are errors.
- Repair loops are bounded by `repair_attempts` and repairer presence.
- Resolved secret values are never serialized or emitted.
1. identify the collaborator boundary and the affected `Prepare` or `Run` state;
2. preserve the `Run`-through-`Prepare` path and error identity;
3. add focused runner or integration tests for changed state transitions,
validation, or repair behavior; and
4. update the owning external contract and any affected source or LLM internal
document.
The [testing policy](../policy/testing.md) owns global test sufficiency.

View File

@@ -2,142 +2,79 @@
## Purpose
This document covers implemented prompt, profile, schema, artifact, and catalog source behavior. It is for developers changing loaders or source wiring.
This document describes how source packages load prompt definitions, profiles,
schemas, and artifacts. The [configuration reference](../config.md) owns their
user-facing formats and settings. The [HTTP API reference](../api.md) owns
HTTP-visible artifact outcomes; [operations](../operations.md) owns deployment
handling.
Full user-facing YAML and config reference material belongs in `docs/config.md`.
## Prompt Definitions
## Prompt Definition Sources
`internal/promptdef` provides filesystem and `fs.FS` repositories. Both use
`internal/filecatalog` for recursive YAML discovery, deterministic ordering,
display paths, and root cleaning.
`internal/promptdef` provides directory-backed and `fs.FS` repositories.
Repositories select a prompt by YAML ID and optional version rather than by
path. They decode through strict YAML handling, reject duplicate matching
definitions, and resolve `content_file` relative to the definition. The `fs.FS`
implementation resolves content paths inside its source root; absolute paths and
traversal outside that root are rejected before file access.
Behavior:
## Profiles And Built-Ins
- recursively scans `.yaml` and `.yml` files.
- decodes YAML with known-fields checking.
- looks up prompts by YAML `id`, not by path.
- optionally filters by prompt `version`.
- rejects duplicate matching prompt IDs.
- requires `id`, `version`, and at least one message.
- requires each message to set exactly one of `content` or `content_file`.
- resolves filesystem `content_file` values relative to the prompt YAML file.
- resolves `fs.FS` `content_file` values inside the configured source root.
- permits prompt subdirectories only as organization; they are not part of prompt identity.
`internal/profile` provides filesystem, `fs.FS`, and overlay repositories.
`internal/profile/builtin` exposes embedded assets through the same repository
interface.
For `fs.FS` roots, absolute paths and relative traversal outside the source root are rejected by catalog path helpers.
An overlay asks its primary source first. It falls back only when the primary
reports `ErrProfileNotFound`; invalid YAML, duplicate IDs, validation failures,
and raw-key failures are returned rather than hidden by fallback. This makes a
custom ID override a built-in ID while retaining errors in the custom source.
## Profile Sources
The public engine can overlay in-memory profiles ahead of both file-backed and
built-in repositories. Profile field definitions, validation ranges, and the
built-in catalog remain in the [configuration reference](../config.md).
`internal/profile` provides directory-backed, `fs.FS`, and overlay repositories. `internal/profile/builtin` embeds built-in profile YAML assets and exposes them through the same repository interface.
## Schemas
Behavior:
`internal/validate` supplies `StandardValidator` for filesystem sources and
`FSValidator` for `fs.FS` sources. Directory-backed validation loads the named
schema path; it does not search directories by basename. `fs.FS` schema paths
are cleaned and checked against their configured root, while a single-file
source matches its file base name.
- recursively scans `.yaml` and `.yml` files.
- decodes YAML with known-fields checking.
- looks up profiles by YAML `id`, not by path.
- rejects duplicate IDs inside the same source.
- rejects raw `api_key` fields in YAML; file-backed profiles must use `api_key_env`.
- validates required `endpoint` and `model` values.
- validates numeric profile ranges.
The runner requests a schema document before generation when it needs
structured output. JSON and schema mismatches in generated content are
validation results; source access, decoding, registration, and compilation
failures are operational errors.
Overlay behavior:
## Artifacts
- custom profiles are primary.
- built-in profiles are fallback.
- fallback occurs only after a primary `ErrProfileNotFound`.
- primary validation, YAML, duplicate, and raw-key errors are returned directly.
- duplicate IDs across custom and built-in sources are allowed because the custom profile overrides the built-in one.
`internal/artifact` composes inline and file readers. The ordinary composite
reader used by CLI and the public engine reads file references from the process
filesystem. The restricted composite reader used by the HTTP adapter combines
inline reading with a rooted file reader and optional byte limit.
The public Go facade can add in-memory profiles ahead of file-backed and built-in profiles.
The rooted reader cleans paths and applies lexical containment without resolving
symlinks. It checks relative references against the configured root and accepts
absolute references only when they remain inside that lexical root. The OS still
follows symlinks after that check. The public containment outcome is documented
by the [HTTP API reference](../api.md); deployment permissions belong in
[operations](../operations.md).
## Schema Sources
## Failure Boundaries
`internal/validate` provides:
Source packages report repository, decoding, duplicate, validation, and read
failures to their callers. They do not select public status codes or response
schemas. The runner wraps source failures with use-case categories; adapters map
them to their own external contract.
- `StandardValidator` for filesystem paths.
- `FSValidator` for `fs.FS` roots and single-file public schema sources.
Source reads use current filesystem or `fs.FS` content for each request. These
packages create no manifests, checkpoints, or durable run state.
Behavior:
## Verification And Change Recipe
- `json_schema` validation requires a non-empty `schema_path`.
- filesystem schema paths resolve relative to `schema_dir` unless absolute.
- directory-backed schema lookup uses the explicit `schema_path`; it does not search recursively by basename.
- `fs.FS` schema paths must remain inside the configured source root.
- single-file schema sources match by the configured file base name.
- schema documents are loaded before the LLM call for structured output.
- JSON parse failures are validation content failures.
- schema access, decode, registration, and compile failures are runtime validation errors.
## Artifact Sources
`internal/artifact` supports two input artifact reference types:
- `inline`
- `file`
Inline behavior:
- requires a non-empty body.
- produces text/plain artifacts.
- hashes the body bytes.
Direct file behavior:
- used by CLI `run`, CLI `render`, and the public Go facade.
- requires a non-empty URI.
- reads from the process filesystem without HTTP artifact-root restrictions.
- infers content type from file extension, defaulting to text/plain.
Restricted file behavior:
- used by HTTP `serve`.
- allows inline artifacts even when no artifact root is configured.
- denies file artifacts when no artifact root is configured.
- resolves relative file URIs against `server.artifact_root`.
- accepts absolute file URIs only when they pass containment checks.
- applies `server.max_artifact_bytes` when configured.
Restricted containment is lexical. It cleans paths and checks the relative path against the configured root; it does not resolve symlinks. Symlinks inside the root are followed by the operating system, including symlinks that target files outside the root.
## Catalog Helpers
`internal/filecatalog` centralizes shared source helpers:
- recursive YAML discovery for filesystem and `fs.FS` roots.
- deterministic sorting.
- `.yaml` and `.yml` filtering.
- display paths for diagnostics.
- YAML file stems.
- `fs.FS` root cleaning and containment checks.
Repository code should use these helpers instead of reimplementing path traversal and containment rules.
## Failure Behavior
Common source failures:
- missing prompt/profile/schema/artifact files.
- invalid YAML or JSON.
- unknown YAML fields.
- duplicate prompt or profile IDs.
- prompt/profile validation errors.
- raw API key fields in profile YAML.
- unsupported artifact reference type.
- missing inline body or file URI.
- artifact outside HTTP root.
- artifact exceeding HTTP size limit.
- schema load or compile failure.
Prompt/profile repository lookup errors are mapped by adapters separately from runtime runner errors. Validation content failures remain result state; source and schema runtime failures return errors.
## State And Manifests
Source packages do not persist run state.
- No manifests are read or written.
- No source package implements skip or resume behavior.
- Source reads reflect the current filesystem or `fs.FS` state for each request.
## Tests To Inspect
Inspect:
- `internal/promptdef/repository_test.go`
- `internal/profile/repository_test.go`
@@ -147,11 +84,14 @@ Source packages do not persist run state.
- `internal/usecase/integration_test.go`
- `engine_test.go`
## Architectural Invariants
When updating prompt, profile, schema, or built-in assets:
- Prompt/profile identity comes from YAML `id`.
- External YAML decoding remains strict.
- File-backed profile YAML never accepts raw API key values.
- Built-in profiles are fallback, not a replacement for custom source validation.
- HTTP file artifacts remain rooted by lexical containment.
- Schema runtime failures remain errors, while JSON/schema content mismatches remain validation results.
1. keep assets valid for the strict loader and the relevant source boundary;
2. update the [configuration reference](../config.md) when a file-format,
catalog, or default changes;
3. run focused source and integration tests, including the built-in repository
test when embedded assets change; and
4. update this document when discovery, precedence, containment, or failure
mechanics change.
The [testing policy](../policy/testing.md) owns global test sufficiency.

View File

@@ -1,164 +1,157 @@
# Operations Guide
## Scope
## Scope And References
This guide covers operating the implemented CLI commands and HTTP service. It
does not replace the [CLI reference](cli.md), [Configuration reference](config.md),
or [HTTP API reference](api.md).
This runbook covers deployment, normal operation, capacity planning, and safe
recovery for Scriptorium. It does not redefine invocation syntax, configuration
fields, or HTTP wire behavior.
## Operational Model
- [CLI reference](cli.md): commands, output destinations, and exit codes.
- [Configuration reference](config.md): configuration, prompt/profile/schema
formats, defaults, and credentials.
- [HTTP API reference](api.md): route, request/response schema, status codes,
limits, and HTTP artifact access.
- [Consumer integration overview](consumers/api.md): caller responsibilities.
Scriptorium executes one prompt request per CLI invocation or HTTP request.
## Operational Model And State
Important boundaries:
Scriptorium handles one prompt request for each CLI invocation or HTTP request.
It has no durable run store, archive, checkpoint, cache, or resume mechanism.
A failed or interrupted request is recovered by correcting its inputs,
configuration, or environment and submitting a new request.
- No durable run state is stored.
- No manifest, archive, checkpoint, or built-in backup workflow is written.
- No built-in resume behavior exists.
- Recovery is rerun-based: correct inputs, config, or environment, then run again.
Generated artifacts, rendered prompts, model output, and run metadata are
caller-owned data. Retention, encryption, backup, and deletion are deployment
responsibilities.
## Filesystem Layout
## Deploy The Filesystem And Process
Operational deployments usually provide:
Provide the process with readable prompt, profile, and schema sources. Keep
prompt templates adjacent to the prompt definitions that reference them. For an
HTTP deployment that accepts file artifacts, use a dedicated, narrow artifact
directory rather than a general-purpose or sensitive filesystem tree.
- `prompt_dir`: prompt definition YAML files and adjacent `content_file` templates.
- `profile_dir`: optional custom profile YAML files.
- `schema_dir`: optional JSON Schema files.
- `server.artifact_root`: optional HTTP file-input root for `serve`.
Run Scriptorium under an identity that can:
Keep these directories readable by the Scriptorium process. Keep
`server.artifact_root` narrow and not writable by untrusted users.
- read only the prompt, profile, schema, and allowed input-artifact paths it
needs;
- read the required credential environment variables without writing them to
files or logs; and
- write only caller-selected output locations when CLI output files are used.
## Normal CLI Workflow
Do not make the HTTP artifact directory writable by untrusted users. The HTTP
artifact containment behavior is lexical and the operating system follows
symlinks; account for that when choosing ownership and mount boundaries. See
the [HTTP API reference](api.md) for the externally observable behavior.
Use `render` before `run` when changing prompt/profile/input wiring:
## Supply Credentials And Protect Runtime Data
```bash
go run ./cmd/scriptorium render \
--config ./examples/config.yml \
--prompt generic.markdown_summary \
--input transcript=./examples/fixtures/transcript.md \
--input glossary=./examples/fixtures/glossary.yml \
--format json
```
Set secret values in the process environment and configure only their
environment-variable names. Do not put raw keys in configuration, prompt or
profile files, process arguments, HTTP payloads, captured command lines, or
debug dumps.
Use `run` for generation after preflight:
Treat stdout, stderr, prepared-run output, generated artifacts, and HTTP
responses as potentially sensitive. Send service logs to a controlled collector
and apply the same retention and access rules as for model input and output.
```bash
go run ./cmd/scriptorium run \
--config ./examples/config.yml \
--prompt generic.markdown_summary \
--input transcript=./examples/fixtures/transcript.md \
--input glossary=./examples/fixtures/glossary.yml \
--out ./summary.md
```
## Run A Normal Workflow
Before production runs, confirm:
Before changing production inputs, profiles, or schemas:
- the effective config path is the intended one;
- prompt/profile/schema directories are readable;
- input file paths exist and match prompt input names;
- required API-key environment variables are set;
- the selected model endpoint is reachable from the process environment.
1. confirm the deployed configuration selects the intended sources and model
credentials;
2. use [`render`](cli.md) with the same request inputs and variables to confirm
preparation without a model call;
3. use [`run`](cli.md) for generation; and
4. retain or discard validation-failed output according to the caller's
policy.
## HTTP Service Operation
The [maintained render script](../examples/render-markdown-summary.sh) is a
copyable preflight example. The CLI reference owns its complete invocation and
exit semantics.
Start the service with:
## Expose The HTTP Service
```bash
go run ./cmd/scriptorium serve --config ./examples/config.yml
```
The HTTP service has no built-in authentication or authorization. Place it on a
trusted network or behind an authenticated reverse proxy, API gateway, or
equivalent access control. Restrict who can reach it and who can read the
artifact root.
The implemented HTTP route is `POST /v1/runs`; request and response fields are
defined in the [HTTP API reference](api.md).
Use a service manager or supervisor appropriate to the deployment to manage
process lifetime, restart policy, log capture, and environment injection. The
[HTTP API reference](api.md) owns client request shapes, status behavior, and
artifact-access outcomes.
The maintained HTTP request-shape example is `examples/http-run.json`.
## Plan Capacity And Limits
HTTP service notes:
Capacity is primarily determined by concurrent model calls, input and output
sizes, schema complexity, provider latency, and network behavior. Size limits
protect request bodies, HTTP file artifacts, and encoded responses; configure
them through the [configuration reference](config.md) and rely on the
[HTTP API reference](api.md) for their response effects.
- Unknown JSON fields are rejected.
- `inline` input references work without an artifact root.
- `file` input references require `server.artifact_root` or `serve --artifact-root`.
- Request bodies, HTTP file input artifacts, and encoded JSON responses are size-limited.
- Validation content failures return `200 OK` with `validation.status: "failed"`.
Before increasing a limit:
Security boundary:
1. measure representative input, generated-output, and optional raw-output
sizes;
2. confirm memory, network, and upstream-provider capacity;
3. retain an upstream request-size and authentication boundary; and
4. test the intended workload in a non-production environment.
- `serve` has no built-in authentication or authorization.
- Put it behind trusted controls such as a private network, authenticated reverse proxy, or API gateway.
- Do not expose an artifact root containing unrelated sensitive files.
- Symlinks inside the artifact root are followed by the operating system.
For large local inputs, prefer a controlled file-artifact directory over
placing arbitrary paths on the service host. Avoid disabling a limit unless an
equivalent trusted control exists elsewhere.
## Secrets Handling
## Diagnose And Recover
Raw API keys are not accepted in app config, profiles, CLI flags, or HTTP
request bodies.
### Preparation Or Configuration Failure
Use this pattern:
Capture the CLI diagnostic or HTTP error response, then verify the selected
configuration, prompt ID, profile selection, source readability, and input
mapping. Use `render` with the same request when it is unclear whether failure
occurs before model execution. Consult the [CLI reference](cli.md), the
[configuration reference](config.md), and the [HTTP API reference](api.md) for
the exact interface contract.
1. Set an environment variable containing the secret value.
2. Store only the variable name in profile `api_key_env` or request override `api_key_env`.
3. Scope the process environment to the minimum required variables.
### Credential Or Provider Failure
## Output, Logs, And Exit Codes
Confirm that the process environment contains the configured credential name
without printing the secret. Check endpoint reachability and provider health
from the process network. If preparation succeeds but generation fails, inspect
the selected model settings in prepared output and the service's controlled
logs. Correct the deployment or provider issue, then submit a new request.
`run`:
### Artifact Or Permission Failure
- stdout: generated artifact body unless `--out` is used.
- stderr: summary on success, errors on failure.
- exit `2`: generation completed and output was written, but validation failed.
Verify that the process can read the intended local input. For HTTP file
artifacts, verify the deployment's artifact root, ownership, path layout, and
file size. Do not widen filesystem permissions or the allowed root merely to
make an arbitrary path work; move or copy the required artifact into the
controlled location instead.
`render`:
### Validation Failure
- stdout: prepared-run output unless `--out` is used.
- stderr: errors.
- exit `0` on success, `1` on failure.
A generated-content validation failure is distinct from a runtime failure.
CLI `run` reports the validation result and error count in its success summary;
it does not print the individual validation messages. For HTTP, inspect the
validation object in the response according to the [HTTP API reference](api.md).
`serve`:
Use rendered input and generated output to determine whether prompt instructions,
the selected model, or the schema needs correction. If schema loading or
compilation itself fails, correct the source deployment or schema document
before rerunning.
- stderr: startup and server errors.
- HTTP response body: JSON success or error envelope.
### HTTP Limit Or Request Failure
## Validation Behavior
Compare the request, artifact, or expected response size with the deployed
configuration, and validate the request against the [HTTP API reference](api.md).
Reduce the payload, use an appropriate controlled artifact source, omit
unneeded raw output, or adjust the deployment limit after capacity review.
Prompt `output.validation_mode` controls validation:
## Cleanup And Reruns
- `none`: skipped.
- `basic`: output body must not be empty.
- `json`: output body must parse as JSON.
- `json_schema`: output body must parse as JSON and satisfy the configured schema.
Runtime/schema failures are hard failures (`run` exit `1`, HTTP error).
Generated-content validation failures are soft failures (`run` exit `2`, HTTP
`200 OK` with failed validation status).
## Size Limits
Defaults are documented in [Configuration reference](config.md). Operationally:
- Keep default HTTP limits unless larger payloads are measured and expected.
- Prefer `inline` HTTP inputs for small payloads.
- Prefer `file` HTTP inputs for larger local artifacts under a controlled artifact root.
- Increase `server.max_response_bytes` when generated artifacts or requested raw output are expected to be large.
- Use `0` only when another trusted layer enforces size limits.
## Maintained Examples
- `examples/config.yml`
- `examples/config.full.yml`
- `examples/render-markdown-summary.sh`
- `examples/http-run.json`
## Safe Recovery
For failed CLI commands or HTTP requests:
1. Capture stderr or the HTTP error `code` and `message`.
2. Confirm config path and effective directory settings.
3. Verify prompt ID, profile ID, schema path, and input mappings.
4. Verify required API-key environment variables.
5. Reproduce with `render --format json` when pre-LLM resolution is uncertain.
6. Rerun after correction.
Because Scriptorium does not persist run state, rerun is the supported recovery
path.
Because no run state is retained, cleanup concerns caller-owned output files,
logs, and artifacts only. Remove or rotate them using the deployment's normal
retention policy. After a correction, rerun the request from the beginning;
there is no safe resume point.

View File

@@ -4,14 +4,12 @@ This document is the development architecture policy for Scriptorium.
It is for developers and LLM coding agents. User-facing behavior belongs in `README.md` and the docs under `docs/` that target operators/users.
## Project Shape
## System Shape
Scriptorium is a narrow prompt-execution application with three entry paths:
- CLI `run`
- CLI `render`
- HTTP `POST /v1/runs` through `serve`
- public Go package `gitea.maximumdirect.net/eric/scriptorium`
Scriptorium is a narrow prompt-execution application with three executable
entry paths: CLI `run`, CLI `render`, and the HTTP service started by `serve`.
It also provides a public Go package for in-process use. Its current component
inventory is maintained in the [internal overview](../internal/overview.md).
Domain behavior is centralized in `internal/usecase` and `internal/domain`.
@@ -23,45 +21,17 @@ Domain behavior is centralized in `internal/usecase` and `internal/domain`.
- Keep config strict: YAML/JSON decoding for external inputs should reject unknown fields.
- Keep secrets out of payloads: raw API key values must not be accepted or emitted.
## Package Boundaries
## Dependency Direction
Current package map:
- root package `scriptorium`: public Go facade over engine construction, source options, request/result types, and error mapping.
- `cmd/scriptorium`: process entrypoint.
- `internal/adapter/cli`: command parsing, app wiring for CLI commands, output behavior.
- `internal/adapter/http`: HTTP DTO mapping and error/status mapping.
- `internal/config`: application settings loading and CLI override precedence.
- `internal/defaults`: compile-time default constants.
- `internal/domain`: core request/result and contract types.
- `internal/usecase`: `Runner` prepare/run orchestration and repair-hook boundary.
- `internal/promptdef`: filesystem prompt-definition repository.
- `internal/profile`: filesystem, `fs.FS`, and overlay execution-profile repositories.
- `internal/profile/builtin`: embedded built-in execution profiles.
- `internal/filecatalog`: shared YAML discovery and `fs.FS` source helpers.
- `internal/artifact`: artifact reference readers.
- `internal/prompt`: template renderer.
- `internal/llm`: provider-neutral LLM client interface and OpenAI-compatible implementation.
- `internal/validate`: validator interfaces and standard implementation.
- `internal/format`: prepared-run output formatting.
Detailed component behavior is documented in:
- `docs/internal/runner.md`
- `docs/internal/adapters.md`
- `docs/internal/sources.md`
## Configuration And Precedence
Application settings are resolved as:
1. built-in defaults
2. config file values
3. CLI overrides
`config.yml` is for application wiring (directories, server address, render default format), not prompt/profile runtime execution settings.
Profile selection and runtime model resolution remain use-case concerns.
- Adapters translate external shapes and IO concerns; they do not make
use-case decisions.
- Use-case and domain code depend on explicit repository, renderer, validator,
and LLM interfaces rather than adapter implementations.
- Source, rendering, validation, and LLM implementations remain behind their
package boundaries.
- Dependency-specific types must not leak across unrelated package boundaries.
- Prefer the standard library; add an external dependency only when it
materially reduces risk or complexity.
## State And Persistence Policy
@@ -70,44 +40,32 @@ Scriptorium has no durable run-state store.
- No built-in resume/checkpoint/archive behavior.
- Recovery model is rerun after correcting inputs/config/environment.
## External Integration Policy
## Contract Ownership
Current external contracts:
- inbound HTTP contract: `POST /v1/runs`, documented canonically in `docs/api.md`
- outbound model contract: OpenAI-compatible chat completions subset
- subprocess contract for integrators: CLI `run`/`render`
- public Go package contract: `docs/consumers/pkg-scriptorium.md`
Integration docs belong under `docs/integrations/`.
The [CLI](../cli.md), [configuration](../config.md), [HTTP API](../api.md),
[public Go package](../consumers/pkg-scriptorium.md), and
[integration](../integrations/) documents own their respective external
contracts. This policy keeps only the architectural boundaries that govern
their implementation.
## Error Handling And Logging
- Wrap errors with domain/operation context.
- Map domain errors to adapter-appropriate statuses/codes without leaking sensitive internals.
- Keep stderr summaries concise for CLI success/error paths.
- Never emit raw secret values.
## Testing Expectations
## Testing And Documentation
- Core runner behavior should be covered with isolated unit tests and fixture-based integration tests.
- Adapter behavior should be tested for parse/mapping/error semantics.
- Config parsing, prompt/profile loading, validator behavior, and LLM client error handling should remain covered by package tests.
- Repository-level docs/examples that claim runnable behavior should be validated by tests or smoke commands.
## Documentation Expectations
- Document implemented behavior only outside `docs/roadmap/`.
- Keep canonical reference locations stable (`docs/cli.md`, `docs/config.md`, `docs/operations.md`, `docs/troubleshooting.md`, `docs/internal/`).
- Update docs in the same change when architecture-relevant behavior changes.
Testing philosophy and change-validation expectations are defined by the
[testing policy](testing.md). Documentation ownership and maintenance rules are
defined by the [documentation policy](documentation.md).
## Architectural Invariants
- `Runner.Run` reuses `Runner.Prepare` flow.
- CLI and HTTP currently instantiate `Runner` without a repairer.
- Artifact reading supports `inline` and `file` references.
- Unknown input fields in config/prompt/profile/http JSON should be rejected by strict decoding.
- Raw API key values must not be accepted through config/HTTP payloads.
- Raw API key values must not be accepted through external configuration or
request payloads, and resolved secret values must not be emitted.
## Non-Goals

View File

@@ -2,9 +2,9 @@
## Status
Proposed implementation plan. This document records the findings of the
documentation audit performed after adoption of the canonical-ownership policy.
The revisions described here are not yet implemented.
Completed on 2026-07-26. This document records the findings of the
documentation audit performed after adoption of the canonical-ownership policy
and the completed refresh that addressed them.
## Objective
@@ -486,7 +486,7 @@ owning references, and not duplicated as complete files in prose docs.
`docs/roadmap/migration.md`.
**Gate:** Code, tests, examples, contracts, internal documentation, operations,
policies, and roadmap status agree.
policies, and roadmap status agree. Completed on 2026-07-26.
## Completion Criteria

View File

@@ -93,6 +93,10 @@ At minimum:
refresh and policy updates are merged and the repository has an agreed,
accurate baseline.
**Gate status:** Complete as of 2026-07-26. The completed documentation
refresh and its verification record are in the
[documentation compliance roadmap](documentation.md).
### Step 2: Record The Architectural Decision And Detailed Boundary
Create an ADR, under the policy established in Step 1, that records:

View File

@@ -1,361 +0,0 @@
# Troubleshooting
This guide lists common implemented failure modes and safe fixes.
Canonical references:
- [CLI reference](cli.md)
- [Configuration reference](config.md)
- [HTTP API reference](api.md)
- [Operations guide](operations.md)
## Missing Or Invalid Config
Symptom:
- CLI error includes `application config error`, `config file not found`, `invalid config YAML`, or `invalid config`.
Likely cause:
- `--config` points to a missing file.
- YAML syntax is invalid.
- Config contains unknown fields or negative HTTP size limits.
Diagnostic step:
```bash
go run ./cmd/scriptorium render --config /path/to/config.yml --prompt generic.markdown_summary --input transcript=./examples/fixtures/transcript.md --input glossary=./examples/fixtures/glossary.yml
```
Safe fix:
- Correct the config path.
- Fix YAML syntax.
- Remove unknown fields.
- Keep raw secrets out of config.
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
## Missing Prompt Directory
Symptom:
- CLI parse error says the prompt directory is required.
Likely cause:
- Neither config nor CLI flags provide an effective `prompt_dir`.
Diagnostic step:
- Re-run once with explicit `--prompt-dir`.
Safe fix:
- Set `prompt_dir` in config or pass `--prompt-dir`.
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
## Unknown Flags
Symptom:
- CLI parse error for an unknown flag.
Likely cause:
- Typo.
- Flag is valid for another command.
- `serve` was given runtime model override flags.
Diagnostic step:
- Compare the command with the command-specific flag list.
Safe fix:
- Remove unsupported flags.
- Use `run` or `render` for runtime model overrides.
Relevant links: [CLI reference](cli.md)
## Prompt Load Failures
Symptom:
- CLI run/render fails during prompt loading.
- HTTP returns `404 prompt_not_found` or `400 prompt_load_failed`.
Likely cause:
- Prompt ID/version does not exist.
- Prompt YAML is invalid or has unknown fields.
- Prompt contract is invalid, such as missing messages, invalid output mode, bad `content_file`, or missing `schema_path` for `json_schema`.
Diagnostic step:
```bash
go run ./cmd/scriptorium render --config ./examples/config.yml --prompt <prompt-id> --input transcript=./examples/fixtures/transcript.md --input glossary=./examples/fixtures/glossary.yml --format json
```
Safe fix:
- Correct prompt ID/version.
- Fix prompt YAML and referenced `content_file` paths.
- Fix output contract fields.
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
## Profile Load Failures
Symptom:
- CLI run/render fails during profile loading.
- HTTP returns `404 profile_not_found`, `400 profile_load_failed`, or `400 profile_required`.
Likely cause:
- Profile ID does not exist.
- Request omitted profile and prompt has no `default_profile`.
- Profile YAML is invalid or has unknown fields.
- Profile contains raw `api_key`.
Diagnostic step:
```bash
go run ./cmd/scriptorium render --config ./examples/config.yml --prompt generic.markdown_summary --profile <profile-id> --input transcript=./examples/fixtures/transcript.md --input glossary=./examples/fixtures/glossary.yml
```
Safe fix:
- Correct profile ID or prompt `default_profile`.
- Fix profile YAML and value ranges.
- Replace raw `api_key` with `api_key_env`.
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
## Input Artifact Failures
Symptom:
- CLI run/render fails while reading inputs.
- HTTP returns `400 artifact_read_failed`, `400 artifact_not_allowed`, or `413 artifact_too_large`.
Likely cause:
- Input file path is missing or unreadable.
- HTTP input type is unsupported or missing required fields.
- HTTP file refs are disabled because no artifact root is configured.
- HTTP file path is lexically outside the artifact root.
- HTTP file input exceeds `server.max_artifact_bytes`.
Diagnostic step:
- Verify each input path exists and is readable by the process.
- For HTTP, verify input refs use `file` or `inline`.
- For HTTP file refs, verify the artifact root and compare file size to `server.max_artifact_bytes`.
Safe fix:
- Correct paths and permissions.
- Configure a narrow artifact root for HTTP file refs.
- Use relative paths under the artifact root or switch to `inline`.
- Increase `server.max_artifact_bytes` only for expected larger inputs.
Relevant links: [HTTP API reference](api.md), [Configuration reference](config.md)
## Missing API-Key Environment Variable
Symptom:
- CLI render/run fails with an API-key environment error.
- HTTP returns `400 api_key_env_missing`.
Likely cause:
- Selected profile or runtime override sets `api_key_env`, but the environment variable is unset or empty.
Diagnostic step:
```bash
printenv SCRIPTORIUM_API_KEY
```
Safe fix:
- Set the required environment variable before starting the CLI command or HTTP service.
- Or use a profile that does not require provider API-key auth.
Relevant links: [Configuration reference](config.md), [Operations guide](operations.md)
## Prompt Template Render Failures
Symptom:
- CLI render/run fails during prompt rendering.
- HTTP returns `400 prompt_render_failed`.
Likely cause:
- Template references an input that was not supplied.
- Template syntax or variable reference is invalid.
Diagnostic step:
- Run `render --format json` with the same prompt, inputs, vars, and profile.
Safe fix:
- Align `{{input "name"}}` references with request input names.
- Fix template syntax and variable names.
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
## LLM Request Failures
Symptom:
- CLI `run` fails during generation.
- HTTP returns `502 llm_failed`.
Likely cause:
- Endpoint is unreachable.
- Provider returns non-2xx.
- Request times out.
- Provider response is malformed.
Diagnostic step:
- Run `render` first to confirm pre-LLM preparation works.
- Check selected endpoint/model in prepared output.
- Check network/provider logs for timeout or non-2xx details.
Safe fix:
- Correct endpoint/model/profile settings.
- Adjust timeout when appropriate.
- Resolve provider or network issue.
Relevant links: [Operations guide](operations.md), [Configuration reference](config.md)
## Validation Failed
Symptom:
- CLI `run` exits `2`.
- HTTP returns `200 OK` with `validation.status` set to `failed`.
Likely cause:
- Generated output failed `basic`, `json`, or `json_schema` content validation.
Diagnostic step:
- Inspect validation errors in CLI stderr or the HTTP response.
Safe fix:
- Refine prompt instructions.
- Adjust schema or model/profile settings.
- Rerun after correction.
Relevant links: [Operations guide](operations.md), [HTTP API reference](api.md)
## Validation Runtime Failure
Symptom:
- CLI `run` fails with validation runtime error.
- HTTP returns `500 validation_runtime_failed`.
Likely cause:
- `json_schema` schema file is missing or unreadable.
- Schema JSON is invalid.
Diagnostic step:
- Verify `schema_dir` and prompt `output.schema_path`.
- Check schema file readability and JSON syntax.
Safe fix:
- Correct schema path or permissions.
- Fix schema JSON.
- Rerun.
Relevant links: [Configuration reference](config.md), [Operations guide](operations.md)
## HTTP JSON Or Request Contract Errors
Symptom:
- HTTP returns `400 invalid_json` or `400 invalid_request`.
Likely cause:
- JSON body is malformed.
- Request has unknown fields or trailing JSON tokens.
- Required `prompt_id` or `inputs` is missing.
- Runtime override values are out of range.
- `extra_params` collides with reserved outbound fields.
Diagnostic step:
- Revalidate request JSON and compare fields with the API reference.
Safe fix:
- Send one JSON object with only supported fields.
- Include `prompt_id` and at least one input.
- Use valid model override ranges.
- Remove reserved `extra_params` keys.
Relevant links: [HTTP API reference](api.md)
## HTTP Size Limit Errors
Symptom:
- HTTP returns `413 request_too_large`, `413 artifact_too_large`, or `413 response_too_large`.
Likely cause:
- JSON request body exceeds `server.max_request_bytes`.
- HTTP file input exceeds `server.max_artifact_bytes`.
- Encoded JSON response exceeds `server.max_response_bytes`.
Diagnostic step:
- Compare request, file input, and expected response sizes with configured limits.
Safe fix:
- Use smaller inline inputs or switch to file inputs under the artifact root.
- Reduce generated output size.
- Omit `include_raw_output`.
- Increase limits only when the deployment expects larger payloads.
Relevant links: [HTTP API reference](api.md), [Operations guide](operations.md)
## HTTP Route Or Method Errors
Symptom:
- HTTP returns `404 not_found` or `405 method_not_allowed`.
Likely cause:
- Path is not `/v1/runs`.
- Method on `/v1/runs` is not `POST`.
Diagnostic step:
- Check the request URL and method.
Safe fix:
- Send `POST /v1/runs`.
Relevant links: [HTTP API reference](api.md)

View File

@@ -36,6 +36,40 @@ func (f *fakeRunner) Run(ctx context.Context, req domain.RunRequest) (*domain.Ru
return f.result, nil
}
func TestMaintainedHTTPRunExampleMatchesRequestContract(t *testing.T) {
body, err := os.ReadFile(filepath.Join("..", "..", "..", "examples", "http-run.json"))
if err != nil {
t.Fatalf("read maintained HTTP request example: %v", err)
}
runner := &fakeRunner{result: &domain.RunResult{}}
h := NewHandler(runner)
req := httptest.NewRequest(http.MethodPost, "/v1/runs", bytes.NewReader(body))
w := httptest.NewRecorder()
h.ServeHTTP(w, req)
if w.Code != http.StatusOK {
t.Fatalf("expected maintained HTTP request example to be accepted, got %d: %s", w.Code, w.Body.String())
}
if runner.last.PromptID != "generic.markdown_summary" {
t.Fatalf("unexpected prompt ID: %q", runner.last.PromptID)
}
if runner.last.ProfileID != "local-fast" {
t.Fatalf("unexpected profile ID: %q", runner.last.ProfileID)
}
wantInputs := map[string]domain.ArtifactRef{
"transcript": {Type: domain.ArtifactRefFile, URI: "./examples/fixtures/transcript.md"},
"glossary": {Type: domain.ArtifactRefFile, URI: "./examples/fixtures/glossary.yml"},
}
if !reflect.DeepEqual(runner.last.Inputs, wantInputs) {
t.Fatalf("unexpected inputs: got %#v, want %#v", runner.last.Inputs, wantInputs)
}
if runner.last.Vars["session_date"] != "2026-05-04" {
t.Fatalf("unexpected session_date: %#v", runner.last.Vars)
}
}
type handlerPromptRepo struct {
def *domain.PromptDefinition
}