diff --git a/docs/api.md b/docs/api.md index 1cfb779..1ae6fc8 100644 --- a/docs/api.md +++ b/docs/api.md @@ -2,296 +2,142 @@ This is the canonical public HTTP contract for Scriptorium. -Implemented route: +## Service And Route -- `POST /v1/runs` +`POST /v1/runs` runs one prompt request and returns generated output, +validation, and metadata. The service has no built-in authentication or +authorization; deploy it behind appropriate network and authentication controls. -For CLI behavior, see [CLI reference](cli.md). For config and prompt/profile -file formats, see [Configuration reference](config.md). +The service address and HTTP limits are configured as described in the +[configuration reference](config.md). `serve` invocation is defined in the +[CLI reference](cli.md). -The maintained request-shape example is `examples/http-run.json`. It requires a -running `serve` process with an artifact root that can read the referenced -files, plus a reachable model endpoint for full execution. - -## Base URL And Deployment - -`scriptorium serve` listens on `server.addr` or `serve --addr`. The default is -`:8080`. - -The route path is always: - -```text -/v1/runs -``` - -The HTTP adapter has no built-in authentication or authorization. Deploy it -behind trusted network and authentication controls. - -## Media Types - -- Request body: JSON object. -- Response body: JSON object. -- Response `Content-Type`: `application/json`. - -Requests are decoded as JSON regardless of the request `Content-Type` header. -There are no shared query parameters. +Requests and responses are JSON objects. Requests are decoded as JSON regardless +of their `Content-Type`; successful JSON responses use +`Content-Type: application/json`. There are no query parameters. ## Request Limits -HTTP limits are configured through `server.*` config fields or `serve` flags: +The configured request-body limit includes inline artifact bodies. The artifact +limit applies to HTTP `file` inputs. The response limit applies to the encoded +response, including the artifact body and optional raw output. A limit of zero +disables that limit. -- `server.max_request_bytes`: encoded JSON request body limit, including inline input bodies. -- `server.max_artifact_bytes`: file artifact limit for HTTP `file` input references. -- `server.max_response_bytes`: encoded JSON response limit, including artifact body and optional raw output. - -Each limit defaults to `16777216` bytes. `0` disables that limit. +A request body over its limit returns `413 request_too_large`; an oversized +file input returns `413 artifact_too_large`; an oversized encoded response +returns `413 response_too_large`. ## `POST /v1/runs` -Runs one prompt request and returns the generated artifact, validation result, -and metadata. - ### Request Body +The maintained [request example](../examples/http-run.json) is a complete +copyable shape. The smallest valid shape is: + ```json { "prompt_id": "generic.markdown_summary", - "profile_id": "local-fast", - "prompt_version": "1.0.0", "inputs": { - "transcript": { - "type": "file", - "uri": "./examples/fixtures/transcript.md" - }, - "glossary": { - "type": "inline", - "body": "party:\n - Rin" - } - }, - "vars": { - "session_date": "2026-05-04" - }, - "model": { - "endpoint": "http://localhost:8000/v1", - "model": "gpt-4o-mini", - "temperature": 0, - "max_tokens": 800, - "top_p": 1, - "timeout_seconds": 120, - "service_tier": "priority", - "reasoning_effort": "medium", - "api_key_env": "SCRIPTORIUM_API_KEY", - "extra_params": { - "provider_option": "enabled" - } - }, - "include_raw_output": false + "transcript": {"type": "inline", "body": "Source text"} + } } ``` -Request fields: - -| Field | Required | Description | +| Field | Required | Meaning | | --- | --- | --- | -| `prompt_id` | yes | Prompt ID. Must not be blank. | +| `prompt_id` | yes | Non-blank prompt ID. | | `prompt_version` | no | Prompt version filter. | -| `profile_id` | no | Execution profile ID. If omitted, the prompt must define `default_profile`. | -| `inputs` | yes | Object mapping prompt input names to input references. Must contain at least one entry. | -| `vars` | no | Object mapping template variable names to string values. | -| `model` | no | Runtime model override object. | -| `include_raw_output` | no | When `true`, include `raw_model_output` in the response. | +| `profile_id` | no | Execution-profile ID; otherwise the prompt must set `default_profile`. | +| `inputs` | yes | Non-empty object mapping input names to references. | +| `vars` | no | Object mapping template-variable names to strings. | +| `model` | no | Runtime model-override object. | +| `include_raw_output` | no | Include `raw_model_output` when true. | -Input reference fields: +An input reference has a required `type` of `file` or `inline`. A `file` +reference requires `uri`; an `inline` reference requires `body`. -| Field | Required | Description | -| --- | --- | --- | -| `type` | yes | `file` or `inline`. | -| `uri` | for `file` | File URI/path. | -| `body` | for `inline` | Inline artifact body. | +HTTP file references require a configured artifact root. Relative paths resolve +within that root. Absolute paths must be lexically within it; traversal outside +it is rejected with `400 artifact_not_allowed`. This lexical check does not +resolve symlinks: the operating system follows symlinks inside the root, +including ones that target outside it. Keep the root narrow and inaccessible to +untrusted writers. -HTTP `file` references require `server.artifact_root` or `serve ---artifact-root`. Relative file URIs resolve against that root. Absolute file -URIs are accepted only when lexically inside the root. Relative traversal and -absolute paths outside the root return `400 artifact_not_allowed`. +The optional `model` object accepts `endpoint`, `model`, `temperature`, +`max_tokens`, `top_p`, `timeout_seconds`, `service_tier`, +`reasoning_effort`, `api_key_env`, and `extra_params`. Numeric ranges and +credential supply are defined by the [configuration reference](config.md). +Explicit zero values for the numeric fields are overrides; zero +`timeout_seconds` disables the outbound client timeout. -The containment check is lexical and does not resolve symlinks. Symlinks inside -the artifact root are followed by the operating system, including symlinks that -point outside the root. Keep the artifact root narrow and not writable by -untrusted users. +Raw API-key values are not accepted. `api_key` and any other unknown model +field cause `400 invalid_json`. -Model override fields: +### Strict JSON -| Field | Description | -| --- | --- | -| `endpoint` | Runtime endpoint override. | -| `model` | Runtime model override. | -| `temperature` | Number in range `0..2`. Explicit `0` is an override. | -| `max_tokens` | Integer greater than or equal to `0`. Explicit `0` is an override. | -| `top_p` | Number in range `0..1`. Explicit `0` is an override. | -| `timeout_seconds` | Integer greater than or equal to `0`. Explicit `0` disables the outbound client timeout. | -| `service_tier` | Provider-specific request tier. | -| `reasoning_effort` | Provider-specific reasoning setting. | -| `api_key_env` | Name of an environment variable containing the API key. | -| `extra_params` | JSON-compatible provider-specific top-level request fields. | - -Raw API-key values are not accepted in HTTP payloads. A field such as -`api_key` is rejected as unknown JSON. - -`extra_params` keys must not be empty and must not collide with reserved -outbound fields: `model`, `session_id`, `messages`, `temperature`, -`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or -`response_format`. - -### Strict JSON Rules - -Request decoding is strict: - -- malformed JSON returns `400 invalid_json` -- unknown request fields return `400 invalid_json` -- unknown `inputs` item fields return `400 invalid_json` -- unknown `model` fields return `400 invalid_json` -- trailing JSON tokens after the request object return `400 invalid_json` -- request bodies above the configured limit return `413 request_too_large` +Request decoding rejects malformed JSON, unknown fields at every request level, +and trailing JSON tokens with `400 invalid_json`. A blank `prompt_id` or +empty `inputs` object returns `400 invalid_request`. ### Success Response -Status: `200 OK` +A completed run returns `200 OK`, including when generated content fails its +validation contract. The response contains: -```json -{ - "artifact": { - "name": "output", - "content_type": "text/markdown", - "body": "Generated content", - "size": 17, - "hash": "..." - }, - "validation": { - "status": "passed", - "mode": "basic", - "repair_attempts": 0, - "is_valid": true - }, - "metadata": { - "run_id": "...", - "prompt_id": "generic.markdown_summary", - "prompt_version": "1.0.0", - "prompt_hash": "...", - "rendered_prompt_hash": "...", - "selected_profile_id": "local-fast", - "model_name": "gpt-4o-mini", - "endpoint": "http://localhost:8000/v1", - "model_params": { - "endpoint": "http://localhost:8000/v1", - "model": "gpt-4o-mini", - "temperature": 0.2, - "max_tokens": 500, - "top_p": 1, - "timeout_seconds": 90 - }, - "input_hashes": { - "transcript": "..." - }, - "usage": { - "prompt_tokens": 11, - "completion_tokens": 22, - "total_tokens": 33, - "cached_tokens": 0, - "cache_write_tokens": 0 - }, - "start_time": "2026-05-04T12:00:00Z", - "end_time": "2026-05-04T12:00:01Z", - "duration_ms": 1000, - "validation_mode": "basic", - "validation_status": "passed", - "repair_attempts_used": 0 - } -} -``` +- `artifact`: `name`, `content_type`, `body`, `size`, `hash`, and + optional `uri`; +- `validation`: `status`, `mode`, `repair_attempts`, `is_valid`, plus + optional `errors` and `schema_path`; +- `metadata`: run, prompt, rendered-prompt, profile, model, input-hash, usage, + timing, validation, and repair-attempt metadata; and +- optional `raw_model_output` when requested. -Response fields: +`metadata.model_params` has `endpoint`, `model`, `temperature`, +`max_tokens`, `top_p`, and `timeout_seconds`, plus optional +`service_tier`, `reasoning_effort`, `api_key_env`, and `extra_params`. +`metadata.usage` always includes `prompt_tokens`, `completion_tokens`, +`total_tokens`, `cached_tokens`, and `cache_write_tokens`; unavailable +cache usage is reported as zero. -- `artifact`: generated output artifact. -- `validation`: validation result for the generated artifact. -- `metadata`: run and effective runtime metadata. -- `raw_model_output`: omitted unless `include_raw_output` is `true`. - -`artifact.uri` is omitted when empty. `validation.errors` and -`validation.schema_path` are omitted when empty. `model_params.service_tier`, -`model_params.reasoning_effort`, `model_params.api_key_env`, and -`model_params.extra_params` are omitted when empty. - -`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are -always present as numbers. They are `0` when the provider omits compatible cache -usage fields or reports no cache activity. - -### Validation Failure Response - -Generated-content validation failures still return `200 OK`. - -```json -{ - "validation": { - "status": "failed", - "mode": "json", - "errors": ["invalid JSON: ..."], - "repair_attempts": 0, - "is_valid": false - } -} -``` - -The response still includes `artifact` and `metadata`. +A validation failure has `validation.status: "failed"`, `is_valid: false`, +and any available diagnostic errors, while still returning the artifact and +metadata. ## Error Responses -Error body shape: +Errors have this shape: ```json -{ - "error": { - "code": "invalid_request", - "message": "prompt_id is required" - } -} +{"error":{"code":"invalid_request","message":"prompt_id is required"}} ``` -Current status/code mapping: +Messages are concise and do not expose wrapped internal causes. | Status | Code | Meaning | | --- | --- | --- | -| `400` | `invalid_json` | Malformed JSON, unknown JSON field, or trailing JSON token. | -| `400` | `invalid_request` | Missing/invalid request fields or invalid runtime overrides. | -| `400` | `profile_required` | No `profile_id` and prompt has no `default_profile`. | -| `400` | `prompt_load_failed` | Prompt definition YAML/contract failed to load. | -| `400` | `profile_load_failed` | Profile YAML/contract failed to load, including raw `api_key`. | -| `400` | `artifact_not_allowed` | HTTP file refs are disabled or requested path is outside artifact root. | -| `400` | `artifact_read_failed` | Input artifact could not be read or input ref was unsupported/invalid. | +| `400` | `invalid_json` | Malformed JSON, unknown field, or trailing JSON. | +| `400` | `invalid_request` | Missing or invalid request data or runtime override. | +| `400` | `profile_required` | No profile ID and no prompt default profile. | +| `400` | `prompt_load_failed` | Prompt definition failed to load. | +| `400` | `profile_load_failed` | Profile failed to load. | +| `400` | `artifact_not_allowed` | HTTP file input is disabled or outside the artifact root. | +| `400` | `artifact_read_failed` | Input artifact is invalid or cannot be read. | | `400` | `prompt_render_failed` | Prompt template rendering failed. | -| `400` | `api_key_env_missing` | Selected `api_key_env` variable is unset or empty. | -| `404` | `not_found` | Route path is unknown. | -| `404` | `prompt_not_found` | Prompt ID/version was not found. | -| `404` | `profile_not_found` | Profile ID was not found. | -| `405` | `method_not_allowed` | Method is not `POST` on `/v1/runs`. | -| `413` | `request_too_large` | Encoded JSON request body exceeds configured request limit. | -| `413` | `artifact_too_large` | HTTP file input artifact exceeds configured artifact limit. | -| `413` | `response_too_large` | Encoded JSON response exceeds configured response limit. | -| `500` | `validation_runtime_failed` | Validator runtime/schema loading failed. | -| `500` | `internal_error` | Unclassified server error. | +| `400` | `api_key_env_missing` | The selected credential environment variable is unset or empty. | +| `404` | `not_found` | Route does not exist. | +| `404` | `prompt_not_found` | Prompt ID or version does not exist. | +| `404` | `profile_not_found` | Profile ID does not exist. | +| `405` | `method_not_allowed` | The route does not accept the method. | +| `413` | `request_too_large` | Encoded request exceeds its limit. | +| `413` | `artifact_too_large` | File input exceeds its limit. | +| `413` | `response_too_large` | Encoded response exceeds its limit. | +| `500` | `validation_runtime_failed` | Schema or validator runtime failure. | +| `500` | `internal_error` | Unclassified server failure. | | `502` | `llm_failed` | Outbound model request failed. | -HTTP error messages are intentionally concise and do not include sensitive -internal causes. - ## Retry And Idempotency -Scriptorium does not provide idempotency keys, pagination, caching headers, or -rate limiting. - -Clients may retry transport failures or `5xx` responses when their surrounding -workflow can tolerate another model call. A retry can generate different output -and incur another provider request. - -## Example File - -- `examples/http-run.json` +Scriptorium provides no idempotency keys, pagination, caching headers, or rate +limits. Clients may retry transport failures or `5xx` responses only when +their workflow tolerates another model call: a retry can produce different +output and incur another provider request. diff --git a/docs/cli.md b/docs/cli.md index 94a4d22..ee88f2f 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -1,5 +1,10 @@ # CLI Reference +This is the canonical contract for invoking Scriptorium. Configuration discovery, +precedence, directories, profiles, and schemas are defined in the +[configuration reference](config.md). The [HTTP API reference](api.md) owns +service request and response behavior. + ## Shortest Useful Command ```bash @@ -10,224 +15,132 @@ go run ./cmd/scriptorium render \ --input glossary=./examples/fixtures/glossary.yml ``` -`render` prepares the prompt, loads input artifacts, resolves the execution -profile, and prints the prepared request without calling an LLM. +`render` prepares a request without calling an LLM. -## Command Overview +## Commands -- `scriptorium run`: prepare a prompt, call the configured LLM, write generated output, and print a run summary. -- `scriptorium render`: prepare a prompt only; write prepared-run output as `text` or `json`. -- `scriptorium serve`: start the HTTP server for `POST /v1/runs`. +- `scriptorium run`: prepare a prompt, call the configured LLM, and write the + generated artifact. +- `scriptorium render`: prepare a prompt and write prepared-run output. +- `scriptorium serve`: start the HTTP server. -Canonical related references: +All commands accept `--config ` and reject positional arguments. An +effective `prompt_dir` is required for every command. Supply it through the +configuration contract or the command's `--prompt-dir` flag. -- [Configuration reference](config.md) -- [HTTP API reference](api.md) -- [Subprocess integration](integrations/subprocess.md) +## `scriptorium run` -## Common Rules - -- `--config` is supported by `run`, `render`, and `serve`. -- Positional arguments are rejected. -- `run` and `render` require `--prompt`, at least one `--input`, and an effective `prompt_dir`. -- `serve` requires an effective `prompt_dir`. -- `profile_dir` is optional. Without it, only built-in profiles are available. -- If `profile_dir` is set, custom profiles override built-in profiles with the same ID. -- Prompt cache control, `session_id`, structured output, and provider-specific profile fields are configured in YAML, not with CLI flags. - -Config precedence is: - -1. built-in defaults -2. config file values -3. CLI flags - -## Flag Reference - -### `scriptorium run` - -```bash +```text scriptorium run [flags] ``` -Required through flags or config: +Required flags: -- `--prompt-dir `: prompt definition directory. - -Required as flags: - -- `--prompt `: prompt ID to execute. -- `--input name=path`: input file mapping. Repeat or use comma-separated mappings. +| Flag | Meaning | +| --- | --- | +| `--prompt ` | Prompt ID to execute. | +| `--input name=path` | Input file mapping; repeat or use comma-separated mappings. | Optional flags: -- `--config `: application config file. -- `--profile-dir `: custom profile definition directory. -- `--schema-dir `: schema base directory for `json_schema` validation. -- `--profile `: execution profile override. If omitted, the prompt `default_profile` is used. -- `--var name=value`: template variable mapping. Repeat or use comma-separated mappings. -- `--out `: write generated artifact body to a file instead of stdout. -- `--llm-base-url `: runtime endpoint override. -- `--model `: runtime model override. -- `--api-key-env `: runtime API-key environment variable name override. -- `--temperature `: runtime temperature override. -- `--max-tokens `: runtime max tokens override. -- `--top-p `: runtime top-p override. -- `--timeout `: runtime timeout override using Go duration syntax, such as `30s` or `2m`. +| Flag | Meaning | +| --- | --- | +| `--config ` | Application configuration file. | +| `--prompt-dir ` | Prompt-definition directory override. | +| `--profile-dir ` | Custom profile-directory override. | +| `--schema-dir ` | Schema base-directory override. | +| `--profile ` | Execution-profile override. | +| `--var name=value` | Template-variable mapping; repeat or use comma-separated mappings. | +| `--out ` | Write generated content to this file instead of stdout. | +| `--llm-base-url ` | Runtime endpoint override. | +| `--model ` | Runtime model override. | +| `--api-key-env ` | Runtime API-key environment-variable name override. | +| `--temperature ` | Runtime temperature override. | +| `--max-tokens ` | Runtime maximum-token override. | +| `--top-p ` | Runtime top-p override. | +| `--timeout ` | Runtime timeout override using Go duration syntax. | -Deprecated aliases: +Deprecated aliases: `--prompt-id` for `--prompt`, and `--profile-id` for +`--profile`. -- `--prompt-id `: alias for `--prompt`. -- `--profile-id `: alias for `--profile`. +Omitted numeric runtime flags preserve the selected effective value; explicit +zero values override it. `--timeout 0s` disables the outbound HTTP-client +timeout. CLI durations are converted to whole seconds by truncation toward +zero, so any duration whose absolute value is below one second becomes an +explicit zero-second override. -Runtime override notes: +There is no raw API-key flag. Use `--api-key-env`. -- Omitted numeric override flags preserve the selected profile/default value. -- Explicit zero values override the selected profile/default value. -- `--timeout 0s` disables the outbound HTTP client timeout for that request. -- There is no raw API-key flag; use `--api-key-env`. +## `scriptorium render` -### `scriptorium render` - -```bash +```text scriptorium render [flags] ``` -Required through flags or config: +`--prompt ` and at least one `--input name=path` are required. The +following optional flags are supported: `--config`, `--prompt-dir`, +`--profile-dir`, `--profile`, `--var`, `--out`, `--llm-base-url`, +`--model`, `--api-key-env`, `--temperature`, `--max-tokens`, `--top-p`, +`--timeout`, and `--format text|json`. Their meanings match the corresponding +`run` flags; `--format` selects prepared-run output and otherwise uses +`defaults.render_format`. -- `--prompt-dir `: prompt definition directory. +The same deprecated aliases and numeric/timeout behavior as `run` apply. +`render` does not accept `--schema-dir`; configure `schema_dir` through the +configuration file. It resolves profiles and schemas as part of preparation but +does not call an LLM. -Required as flags: +## `scriptorium serve` -- `--prompt `: prompt ID to render. -- `--input name=path`: input file mapping. Repeat or use comma-separated mappings. - -Optional flags: - -- `--config `: application config file. -- `--prompt-dir `: prompt definition directory. -- `--profile-dir `: custom profile definition directory. -- `--profile `: execution profile override. -- `--var name=value`: template variable mapping. Repeat or use comma-separated mappings. -- `--out `: write prepared-run output to a file instead of stdout. -- `--llm-base-url `: runtime endpoint override for the prepared request. -- `--model `: runtime model override for the prepared request. -- `--api-key-env `: runtime API-key environment variable name override. -- `--temperature `: runtime temperature override. -- `--max-tokens `: runtime max tokens override. -- `--top-p `: runtime top-p override. -- `--timeout `: runtime timeout override using Go duration syntax. -- `--format text|json`: prepared-run output format. Defaults to config `defaults.render_format`, then `text`. - -Deprecated aliases: - -- `--prompt-id `: alias for `--prompt`. -- `--profile-id `: alias for `--profile`. - -Notes: - -- `render` resolves profiles, loads schemas for `json_schema` prompts, and validates `api_key_env`. -- `render` does not accept `--schema-dir`; use config `schema_dir` for render-time schema lookup. -- `render` does not call the LLM. - -### `scriptorium serve` - -```bash +```text scriptorium serve [flags] ``` -Required through flags or config: - -- `--prompt-dir `: prompt definition directory. - Optional flags: -- `--config `: application config file. -- `--addr `: HTTP listen address. -- `--prompt-dir `: prompt definition directory. -- `--profile-dir `: custom profile definition directory. -- `--schema-dir `: schema base directory for `json_schema` validation. -- `--artifact-root `: base directory for HTTP `file` input references. -- `--max-request-bytes `: maximum HTTP request body bytes; `0` disables the limit. -- `--max-artifact-bytes `: maximum HTTP file artifact bytes; `0` disables the limit. -- `--max-response-bytes `: maximum encoded HTTP response body bytes; `0` disables the limit. +| Flag | Meaning | +| --- | --- | +| `--config ` | Application configuration file. | +| `--addr ` | HTTP listen-address override. | +| `--prompt-dir ` | Prompt-definition directory override. | +| `--profile-dir ` | Custom profile-directory override. | +| `--schema-dir ` | Schema base-directory override. | +| `--artifact-root ` | Root for HTTP `file` input references. | +| `--max-request-bytes ` | Maximum encoded HTTP request-body bytes; `0` disables the limit. | +| `--max-artifact-bytes ` | Maximum HTTP file-input artifact bytes; `0` disables the limit. | +| `--max-response-bytes ` | Maximum encoded HTTP response bytes; `0` disables the limit. | -Notes: - -- `serve` does not accept runtime model override flags such as `--model` or `--llm-base-url`. -- HTTP request fields and error codes are documented in the [HTTP API reference](api.md). -- HTTP `file` input references are rejected unless an artifact root is configured. -- HTTP size-limit flags affect only `serve`. +`serve` accepts no runtime model override flags. HTTP request fields, response +schemas, and error codes are defined in the [HTTP API reference](api.md). ## Input And Variable Syntax -- `--input name=path` maps prompt input names to local file paths. -- `--var name=value` maps prompt template variables to string values. -- Both flags can be repeated. -- Both flags also accept comma-separated mappings, such as `--input transcript=./t.md,glossary=./g.yml`. -- Values may contain `=` after the first separator, such as `--var note=a=b=c`. -- Empty names and empty values are rejected. +`--input name=path` maps an input name to a local file; `--var name=value` +maps a template variable to a string. Both flags can be repeated or contain +comma-separated mappings. Values may contain `=` after the first separator. +Empty names and values are rejected. -CLI `run` and `render` convert every `--input` mapping to a `file` artifact -reference. HTTP also supports `inline` input references; see [HTTP API -reference](api.md). +CLI inputs are file references. HTTP inline inputs are defined by the +[HTTP API reference](api.md). -## Output Behavior +## Output And Exit Behavior -`run`: +- `run` writes generated content to stdout, or to `--out` when supplied, and + writes a concise summary to stderr. +- `render` writes prepared-run output to stdout, or to `--out` when supplied, + without a success summary. +- `serve` writes startup and server errors to stderr. -- Writes generated artifact content to stdout by default. -- Writes generated artifact content to `--out` when provided. -- Prints a success summary to stderr. -- Prints errors to stderr on failure. +Exit statuses: -`render`: +| Status | Meaning | +| --- | --- | +| `0` | Success. | +| `1` | Parse, configuration, loading, rendering, generation, output-write, or other runtime error. | +| `2` | `run` generated and wrote output, but validation failed. | -- Writes prepared-run output to stdout by default. -- Writes prepared-run output to `--out` when provided. -- Does not print a success summary. +## Workflows And Examples -`serve`: - -- Logs startup and server errors to stderr. - -## Exit Codes - -- `0`: success. -- `1`: parse, config, load, render, generation, output-write, or runtime error. -- `2`: `run` completed and wrote output, but validation status is `failed`. - -## Common Workflows - -Render prompt inputs and variables as JSON: - -```bash -go run ./cmd/scriptorium render \ - --config ./examples/config.yml \ - --prompt generic.markdown_summary \ - --input transcript=./examples/fixtures/transcript.md \ - --input glossary=./examples/fixtures/glossary.yml \ - --var session_date=2026-05-04 \ - --format json -``` - -Run a prompt with an explicit profile and file output: - -```bash -go run ./cmd/scriptorium run \ - --config ./examples/config.yml \ - --prompt generic.markdown_summary \ - --profile local-fast \ - --input transcript=./examples/fixtures/transcript.md \ - --input glossary=./examples/fixtures/glossary.yml \ - --out ./summary.md -``` - -Start the HTTP server with example config: - -```bash -go run ./cmd/scriptorium serve --config ./examples/config.yml -``` - -Copyable maintained script: - -- `examples/render-markdown-summary.sh` +The [maintained render script](../examples/render-markdown-summary.sh) is a +copyable render workflow. The [HTTP request example](../examples/http-run.json) +is for a running `serve` process. diff --git a/docs/config.md b/docs/config.md index a6104f4..0e0e63c 100644 --- a/docs/config.md +++ b/docs/config.md @@ -1,324 +1,169 @@ # Configuration Reference -## Config Discovery And Precedence +This is the canonical reference for Scriptorium application settings and the +prompt, profile, and schema files those settings select. For command syntax, +see the [CLI reference](cli.md); for HTTP request shapes, limits, and outcomes, +see the [HTTP API reference](api.md). + +## Discovery And Precedence Application settings are resolved in this order: -1. built-in defaults -2. `config.yml` values -3. CLI overrides +1. built-in defaults; +2. a configuration file; then +3. CLI overrides. -When `--config` is omitted, Scriptorium searches: +When `--config` is omitted, Scriptorium searches +`/usr/local/etc/scriptorium/config.yml` and then `/etc/scriptorium/config.yml`. +If neither exists, it uses built-in defaults. An explicit `--config` path must +exist and decode successfully. -1. `/usr/local/etc/scriptorium/config.yml` -2. `/etc/scriptorium/config.yml` +The maintained [minimal configuration](../examples/config.yml) and +[full configuration](../examples/config.full.yml) are copyable examples. -If neither file exists, Scriptorium uses built-in defaults. When -`--config ` is provided, that file must exist and decode successfully. +## Application Configuration File -## Minimal Working Config +Configuration is strict YAML: unknown fields are rejected. Empty string values +do not override a prior value. Raw API-key fields are not accepted. -```yaml -prompt_dir: ./examples/prompts -``` - -This is enough for `run` and `render` when selected prompts use built-in -profiles. Set `profile_dir` when prompts or requests use custom profiles. - -The maintained repository example is `examples/config.yml`. - -## Production-Oriented Config - -```yaml -prompt_dir: /opt/scriptorium/prompts -profile_dir: /opt/scriptorium/profiles -schema_dir: /opt/scriptorium/schemas - -server: - addr: 127.0.0.1:8080 - artifact_root: /var/lib/scriptorium/artifacts - max_request_bytes: 16777216 - max_artifact_bytes: 16777216 - max_response_bytes: 16777216 - -defaults: - render_format: text -``` - -The maintained full example is `examples/config.full.yml`. - -## App Config Reference - -Top-level fields: - -| Field | Default | Description | +| Field | Default | Meaning | | --- | --- | --- | -| `prompt_dir` | unset | Directory containing prompt definition YAML files. Required effectively by `run`, `render`, and `serve`. | -| `profile_dir` | unset | Directory containing custom profile YAML files. Built-in profiles remain available when unset. | +| `prompt_dir` | unset | Directory containing prompt-definition YAML. `run`, `render`, and `serve` require an effective value. | +| `profile_dir` | unset | Directory containing custom profile YAML. Built-in profiles remain available. | | `schema_dir` | `.` | Base directory for relative JSON Schema paths. | -| `server` | `{}` | HTTP service settings used by `serve`. | -| `defaults` | `{}` | Adapter defaults. | - -`server` fields: - -| Field | Default | Description | -| --- | --- | --- | -| `server.addr` | `:8080` | Listen address for `serve`. | -| `server.artifact_root` | unset | Base directory for HTTP `file` input references. Without it, HTTP file refs are rejected. | -| `server.max_request_bytes` | `16777216` | Maximum encoded HTTP request body bytes. `0` disables the limit. | -| `server.max_artifact_bytes` | `16777216` | Maximum HTTP file artifact bytes. `0` disables the limit. | -| `server.max_response_bytes` | `16777216` | Maximum encoded HTTP response bytes. `0` disables the limit. | - -`defaults` fields: - -| Field | Default | Description | -| --- | --- | --- | +| `server.addr` | `:8080` | Address used by `serve`. | +| `server.artifact_root` | unset | Root that enables HTTP `file` input references. | +| `server.max_request_bytes` | `16777216` | Maximum encoded HTTP request body bytes; `0` disables the limit. | +| `server.max_artifact_bytes` | `16777216` | Maximum HTTP file-input artifact bytes; `0` disables the limit. | +| `server.max_response_bytes` | `16777216` | Maximum encoded HTTP response bytes; `0` disables the limit. | | `defaults.render_format` | `text` | Default `render` output format: `text` or `json`. | -Config rules: - -- YAML decoding is strict; unknown fields are rejected. -- HTTP size limits must be greater than or equal to `0`. -- Empty string config values are ignored. -- Raw API key fields are not supported in app config. +The three size fields must be zero or greater. The HTTP contract defines how +each limit is enforced and reported. `server.artifact_root` configures the +deployment boundary; see the [HTTP API reference](api.md) for request-path and +containment behavior, and [operations](operations.md) for deployment handling. ## Prompt Definition Files -Prompt definitions are YAML files anywhere under `prompt_dir`. Nested -directories are organizational; callers select prompts by YAML `id`, not file -path. +Prompt definitions are strict YAML files anywhere below `prompt_dir`. A prompt +is selected by its YAML `id`, not by file path; nested directories are only for +organization. See [maintained prompt examples](../examples/prompts/). -Example: - -```yaml -id: generic.structured_events -version: "1.0.0" -default_profile: local-quality -description: Produce structured event JSON from a transcript. - -inputs: - - name: transcript - required: true - content_type: text/markdown - description: Source transcript content - - name: glossary - required: false - content_type: text/yaml - description: Optional glossary context - -messages: - - role: system - content_file: ./generic.structured_events.system.md - - role: user - content_file: ./generic.structured_events.user.md - -output: - format: json - validation_mode: json_schema - schema_path: structured_events.schema.json - repair_attempts: 0 -``` - -Prompt fields: - -| Field | Required | Description | +| Field | Required | Meaning | | --- | --- | --- | -| `id` | yes | Prompt identifier used by `--prompt` and HTTP `prompt_id`. | +| `id` | yes | Prompt identifier. | | `version` | yes | Prompt version. | -| `default_profile` | no | Profile ID used when a request does not provide a profile. | +| `default_profile` | no | Profile used when a request omits a profile ID. | | `description` | no | Human-readable description. | -| `session_id` | no | Go-template string rendered from request vars and forwarded as provider `session_id` when non-empty. | -| `inputs` | no | Named input declarations. | -| `messages` | yes | Chat message templates. | +| `session_id` | no | Go-template string rendered from request variables and sent to a compatible provider when non-empty. | +| `inputs` | no | Declared input metadata. | +| `messages` | yes | Chat-message templates. | | `output` | yes | Output format and validation contract. | -`inputs[]` fields: +### Inputs And Messages -- `name` (required) -- `required` (optional boolean) -- `content_type` (optional metadata) -- `description` (optional) +Each `inputs` item has a required `name` and optional `required`, +`content_type`, and `description` fields. Input names must be unique. -`messages[]` fields: +Each message has a required `role`, exactly one of `content` or `content_file`, +and optional `cache_control`. A `content_file` path is relative to the prompt +file. `cache_control.type` must be `ephemeral`; its optional `ttl` is `1h`. -- `role` (required) -- exactly one of `content` or `content_file` -- `cache_control` (optional) +`session_id` uses the same template variables as messages. Empty rendered +values are omitted. A rendered value may contain at most 256 Unicode code +points. -Message rules: +### Output Contract -- `content_file` resolves relative to the prompt YAML file location. -- Repeated roles are allowed. -- Prompt YAML decoding is strict. -- Duplicate input names are invalid. -- Duplicate prompt IDs are invalid for a requested ID/version. - -`messages[].cache_control` fields: - -| Field | Required | Supported values | +| Field | Required | Values or behavior | | --- | --- | --- | -| `type` | yes | `ephemeral` | -| `ttl` | no | `1h` | - -`session_id` behavior: - -- Rendered with the same variable context as message templates. -- Trimmed and omitted when empty. -- Rejected when longer than 256 Unicode code points. -- CLI callers pass variables with `--var`; HTTP callers use `vars`. - -`output` fields: - -| Field | Required | Supported values | -| --- | --- | --- | -| `format` | yes | `text`, `markdown`, `json` | -| `validation_mode` | yes | `none`, `basic`, `json`, `json_schema` | -| `schema_path` | only for `json_schema` | Relative to `schema_dir` unless absolute. | -| `repair_attempts` | yes | Integer greater than or equal to `0`. | - -Repair boundary: - -- `repair_attempts` is part of the prompt contract. -- The current CLI and HTTP wiring constructs the runner without a repairer, so normal `run` and `serve` execution does not perform repair attempts. +| `format` | yes | `text`, `markdown`, or `json`. | +| `validation_mode` | yes | `none`, `basic`, `json`, or `json_schema`. | +| `schema_path` | for `json_schema` | Schema path, relative to `schema_dir` unless absolute. | +| `repair_attempts` | no | Integer greater than or equal to `0`; omitted means `0`. | ## Profile Definition Files -Execution profiles are YAML files anywhere under `profile_dir`. Nested -directories are organizational; callers select profiles by YAML `id`, not file -path. +Profiles are strict YAML files anywhere below `profile_dir`. A profile is +selected by YAML `id`; nested directories are organizational. See the +[maintained profile examples](../examples/profiles/). -Scriptorium also ships built-in profiles. Custom profiles override built-ins -with the same ID. - -Example: - -```yaml -id: local-fast -endpoint: http://localhost:8000/v1 -model: gpt-4o-mini -temperature: 0.2 -max_tokens: 500 -top_p: 1.0 -timeout_seconds: 90 -api_key_env: SCRIPTORIUM_API_KEY -service_tier: priority -reasoning_effort: medium -extra_params: - provider_route: primary -``` - -Profile fields: - -| Field | Required | Description | +| Field | Required | Meaning | | --- | --- | --- | | `id` | yes | Profile identifier. | -| `endpoint` | yes | OpenAI-compatible base URL including `/v1`. | +| `endpoint` | yes | OpenAI-compatible base URL, including its API version path when needed. | | `model` | yes | Provider model name. | -| `temperature` | no | Range `0..2`. | -| `max_tokens` | no | Integer greater than or equal to `0`. | -| `top_p` | no | Range `0..1`. | -| `timeout_seconds` | no | Integer greater than or equal to `0`. | -| `service_tier` | no | Provider-specific request tier. | -| `reasoning_effort` | no | Provider-specific reasoning setting. | -| `api_key_env` | no | Environment variable name containing the API key. | -| `extra_params` | no | JSON-compatible provider-specific top-level request fields. | +| `temperature` | no | Number from `0` through `2`. | +| `max_tokens` | no | Integer zero or greater. | +| `top_p` | no | Number from `0` through `1`. | +| `timeout_seconds` | no | Integer zero or greater. | +| `service_tier` | no | Non-empty provider-specific request tier. | +| `reasoning_effort` | no | Non-empty provider-specific reasoning setting. | +| `api_key_env` | no | Environment-variable name containing the API key. | +| `extra_params` | no | JSON-compatible provider-specific outbound request fields. | -Execution defaults before profile/request overrides: +Execution defaults before profile and request overrides are `temperature: 0`, +`max_tokens: 0`, `top_p: 1`, and `timeout_seconds: 600`. Profile numeric values +merge by non-zero value. Request overrides preserve presence, so an explicit +zero can override a profile value. -| Field | Default | -| --- | --- | -| `temperature` | `0.0` | -| `max_tokens` | `0` | -| `top_p` | `1.0` | -| `timeout_seconds` | `600` | +Custom profiles take precedence over built-ins with the same ID. Invalid custom +profiles are errors; they do not fall back to a built-in profile. Raw `api_key` +is rejected. Use `api_key_env`, or the public Go package's request-scoped key +mechanism described in the [package contract](consumers/pkg-scriptorium.md). -Profile rules: +`extra_params` keys must be non-empty and cannot be `model`, `session_id`, +`messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`, +`reasoning_effort`, or `response_format`. -- Profile YAML decoding is strict. -- Duplicate custom profile IDs are invalid. -- Matching custom and built-in IDs are valid override behavior. -- Raw `api_key` is rejected; use `api_key_env`. -- If `api_key_env` is set, the named environment variable must be set before `run`, `render`, or HTTP execution can prepare the request. -- Profile numeric fields merge by non-zero value. Request overrides are presence-aware, so explicit zero values are supported through CLI flags or HTTP model overrides. -- `extra_params` keys must not be empty and must not collide with reserved outbound fields: `model`, `session_id`, `messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or `response_format`. +### Built-In Profile Catalog -Built-in profile catalog: +Each embedded profile uses `OPENROUTER_API_KEY`. -| Provider | ID | Model | API key env | -| --- | --- | --- | --- | -| aion-labs | `aion-2` | `aion-labs/aion-2.0` | `OPENROUTER_API_KEY` | -| anthropic | `claude-fable-latest` | `~anthropic/claude-fable-latest` | `OPENROUTER_API_KEY` | -| anthropic | `claude-haiku-latest` | `~anthropic/claude-haiku-latest` | `OPENROUTER_API_KEY` | -| anthropic | `claude-opus-latest` | `~anthropic/claude-opus-latest` | `OPENROUTER_API_KEY` | -| anthropic | `claude-sonnet-latest` | `~anthropic/claude-sonnet-latest` | `OPENROUTER_API_KEY` | -| deepseek | `deepseek-3-2` | `deepseek/deepseek-v3.2` | `OPENROUTER_API_KEY` | -| deepseek | `deepseek-4-pro` | `deepseek/deepseek-v4-pro` | `OPENROUTER_API_KEY` | -| google | `gemini-2-flash` | `google/gemini-2.5-flash` | `OPENROUTER_API_KEY` | -| google | `gemini-2-flash-lite` | `google/gemini-2.5-flash-lite` | `OPENROUTER_API_KEY` | -| google | `gemini-2-pro` | `google/gemini-2.5-pro` | `OPENROUTER_API_KEY` | -| google | `gemini-3-flash-lite` | `google/gemini-3.1-flash-lite` | `OPENROUTER_API_KEY` | -| google | `gemini-flash-latest` | `~google/gemini-flash-latest` | `OPENROUTER_API_KEY` | -| google | `gemini-pro-latest` | `~google/gemini-pro-latest` | `OPENROUTER_API_KEY` | -| google | `gemma-4-31b` | `google/gemma-4-31b-it:exacto` | `OPENROUTER_API_KEY` | -| minimax | `minimax-m2` | `minimax/minimax-m2.5` | `OPENROUTER_API_KEY` | -| minimax | `minimax-m3` | `minimax/minimax-m3` | `OPENROUTER_API_KEY` | -| mistral | `mistral-large-2512` | `mistralai/mistral-large-2512` | `OPENROUTER_API_KEY` | -| mistral | `mistral-medium-3-5` | `mistralai/mistral-medium-3-5` | `OPENROUTER_API_KEY` | -| mistral | `mistral-small-3` | `mistralai/mistral-small-3.2-24b-instruct` | `OPENROUTER_API_KEY` | -| mistral | `mistral-small-4` | `mistralai/mistral-small-2603` | `OPENROUTER_API_KEY` | -| nvidia | `nemotron-3-ultra` | `nvidia/nemotron-3-ultra-550b-a55b` | `OPENROUTER_API_KEY` | -| openai | `gpt-5-mini` | `openai/gpt-5.4-mini` | `OPENROUTER_API_KEY` | -| openai | `gpt-5-nano` | `openai/gpt-5.4-nano` | `OPENROUTER_API_KEY` | +| Provider | ID | Model | +| --- | --- | --- | +| aion-labs | `aion-2` | `aion-labs/aion-2.0` | +| anthropic | `claude-fable-latest` | `~anthropic/claude-fable-latest` | +| anthropic | `claude-haiku-latest` | `~anthropic/claude-haiku-latest` | +| anthropic | `claude-opus-latest` | `~anthropic/claude-opus-latest` | +| anthropic | `claude-sonnet-latest` | `~anthropic/claude-sonnet-latest` | +| deepseek | `deepseek-3-2` | `deepseek/deepseek-v3.2` | +| deepseek | `deepseek-4-flash` | `deepseek/deepseek-v4-flash` | +| deepseek | `deepseek-4-pro` | `deepseek/deepseek-v4-pro` | +| google | `gemini-2-flash` | `google/gemini-2.5-flash` | +| google | `gemini-2-flash-lite` | `google/gemini-2.5-flash-lite` | +| google | `gemini-2-pro` | `google/gemini-2.5-pro` | +| google | `gemini-3-flash-lite` | `google/gemini-3.1-flash-lite` | +| google | `gemini-flash-latest` | `~google/gemini-flash-latest` | +| google | `gemini-pro-latest` | `~google/gemini-pro-latest` | +| google | `gemma-4-31b` | `google/gemma-4-31b-it:exacto` | +| minimax | `minimax-m2` | `minimax/minimax-m2.5` | +| minimax | `minimax-m3` | `minimax/minimax-m3` | +| mistral | `mistral-large-2512` | `mistralai/mistral-large-2512` | +| mistral | `mistral-medium-3-5` | `mistralai/mistral-medium-3-5` | +| mistral | `mistral-small-3` | `mistralai/mistral-small-3.2-24b-instruct` | +| mistral | `mistral-small-4` | `mistralai/mistral-small-2603` | +| nvidia | `nemotron-3-ultra` | `nvidia/nemotron-3-ultra-550b-a55b` | +| openai | `gpt-5-mini` | `openai/gpt-5.4-mini` | +| openai | `gpt-5-nano` | `openai/gpt-5.4-nano` | -## Schema Behavior +## Schemas -Schemas are JSON files, typically under `schema_dir`. +Schemas are JSON files, normally below `schema_dir`. `json_schema` output +requires a `schema_path`. Relative paths resolve from `schema_dir`; absolute +paths are used directly. Referenced nested schemas use relative paths and are +not discovered by basename. An unreadable or invalid schema is a runtime +validation error; generated content that fails JSON or schema validation is a +validation result. -Rules: +## Credentials -- `output.validation_mode: json_schema` requires `output.schema_path`. -- Relative `schema_path` values resolve from `schema_dir`. -- Absolute `schema_path` values are used directly. -- Nested schemas must be referenced by relative path; schemas are not searched recursively by basename. -- Missing or invalid schema documents are runtime validation errors. -- Invalid generated JSON produces validation status `failed`, not a runtime error. +Keep secrets in environment variables. Store only an environment-variable name +in `api_key_env`; do not place raw keys in configuration, prompt or profile +files, CLI arguments, examples, or HTTP payloads. -## Artifact References - -Supported request input artifact reference types are: - -- `file` -- `inline` - -CLI `run` and `render` create `file` references from `--input name=path`. - -HTTP `file` references require `server.artifact_root` or `serve ---artifact-root`. Relative file URIs resolve under that root. Absolute paths -and relative traversal outside the root are rejected by lexical checks. Symlinks -inside the root are followed by the operating system, including symlinks that -point outside the root. - -HTTP `inline` references do not require an artifact root. - -## Secrets Handling - -- Keep secret values in environment variables. -- Store only environment-variable names in `api_key_env`. -- Do not put raw API keys in config, prompts, profiles, CLI arguments, examples, or HTTP request bodies. - -## Maintained Examples - -- Minimal app config: `examples/config.yml` -- Full app config: `examples/config.full.yml` -- Prompt examples: `examples/prompts/` -- Custom profile examples: `examples/profiles/` -- Schema examples: `examples/schemas/` -- Input fixtures: `examples/fixtures/` -- Render script: `examples/render-markdown-summary.sh` -- HTTP request-shape example: `examples/http-run.json` - -## Integration References +## Related References - [CLI reference](cli.md) - [HTTP API reference](api.md) -- [Outbound OpenAI-compatible contract](integrations/openai-compatible-chat.md) +- [OpenAI-compatible outbound contract](integrations/openai-compatible-chat.md) diff --git a/docs/consumers/api.md b/docs/consumers/api.md index e7e21fd..c58d7d9 100644 --- a/docs/consumers/api.md +++ b/docs/consumers/api.md @@ -1,65 +1,26 @@ # Consumer Integration Overview -This guide is for applications that call Scriptorium from another codebase. +This guide helps applications choose a Scriptorium interface and understand +their responsibilities. The linked contracts own interface syntax and wire +semantics. -Scriptorium exposes three integration surfaces: - -| Surface | Use when | +| Interface | Use when | | --- | --- | -| Go package | The consumer is Go, needs typed requests/results, or wants injected LLM clients for tests. | -| CLI subprocess | The consumer wants process isolation or is not written in Go. | -| HTTP API | The consumer needs a service boundary or remote access to `POST /v1/runs`. | +| Go package | The consumer is Go and needs typed requests, results, or an injected LLM client. | +| CLI subprocess | The consumer needs process isolation or is not written in Go. | +| HTTP API | The consumer needs a service boundary or remote access. | -Canonical references: +- Go package: [package contract](pkg-scriptorium.md) +- CLI subprocess: [subprocess integration](../integrations/subprocess.md) +- HTTP service: [HTTP API reference](../api.md) +- Prompt, profile, schema, and credential configuration: [configuration reference](../config.md) -- Go package: [Package scriptorium](pkg-scriptorium.md) -- CLI subprocess: [Subprocess integration](../integrations/subprocess.md) -- HTTP: [HTTP API reference](../api.md) -- File formats: [Configuration reference](../config.md) - -## Required Deployment Inputs - -Every integration needs operators to provide: - -- prompt definitions; -- profile definitions or built-in profile IDs; -- schema files when prompts use `json_schema`; -- input artifacts or inline input bodies; -- API-key environment variables or direct per-request keys where supported. - -Raw API keys do not belong in config, prompt files, profile YAML, CLI -arguments, or HTTP request bodies. - -## Recommended Workflow - -Use the Go package when: - -- the consumer is a Go application; -- the application needs `context.Context` cancellation; -- repeated calls should avoid subprocess startup; -- tests need a fake LLM client; -- direct per-request `RunRequest.APIKey` is required. - -Use the CLI subprocess when: - -- the consumer is not Go; -- process isolation is useful; -- stdout/stderr separation and exit codes are enough; -- the consumer already manages local files and environment variables. - -Use HTTP when: - -- Scriptorium should run as a service; -- multiple clients need a shared prompt/profile deployment; -- clients can reach a trusted, protected HTTP boundary. - -## Minimal Go Example +## Minimal Go Use ```go engine, err := scriptorium.NewEngine(scriptorium.Config{ PromptDir: "./examples/prompts", ProfileDir: "./examples/profiles", - SchemaDir: "./examples/schemas", }) if err != nil { return err @@ -69,54 +30,30 @@ prepared, err := engine.Prepare(ctx, scriptorium.RunRequest{ PromptID: "generic.markdown_summary", Inputs: map[string]scriptorium.ArtifactRef{ "transcript": scriptorium.File("./examples/fixtures/transcript.md"), - "glossary": scriptorium.File("./examples/fixtures/glossary.yml"), }, }) if err != nil { return err } -_ = prepared.Messages +_ = prepared ``` -Run the maintained package example: - -```bash -go run ./examples/go-library/prepare -``` - -## Subprocess Workflow - -Invoke `scriptorium render` for preflight and `scriptorium run` for generation. -Capture stdout and stderr separately. Treat exit code `2` from `run` as a -completed generation with failed validation. - -See [Subprocess integration](../integrations/subprocess.md) for the stable -invocation contract. - -## HTTP Workflow - -Run `scriptorium serve` behind trusted controls and send JSON requests to -`POST /v1/runs`. - -Do not duplicate endpoint schemas in consumers. Use the [HTTP API -reference](../api.md) as the authoritative contract. +For a maintained program, see +[`examples/go-library/prepare`](../../examples/go-library/prepare). ## Consumer Responsibilities Consumers are responsible for: -- selecting prompt/profile IDs as deployment configuration; -- supplying all required inputs and vars; -- protecting generated artifacts and rendered prompts as sensitive data; -- deciding whether to keep output when validation fails; -- implementing retries only when another model call is acceptable. +- selecting and deploying prompt, profile, and schema assets; +- supplying required inputs and template variables; +- supplying credentials through the applicable interface; +- protecting rendered prompts and generated artifacts as potentially sensitive; +- deciding whether validation-failed output is usable; and +- retrying only when another model call is acceptable. -Scriptorium does not persist run state. Retrying a failed or timed-out request -can produce different output and can incur another provider request. - -## Status Behavior - -- Go package methods return typed results or errors that support `errors.Is`. -- CLI `run` exits `2` when generation succeeds but validation fails. -- HTTP returns `200 OK` for generated-content validation failures and exposes the failed status in the response body. -- Runtime validation failures are errors. +Scriptorium does not persist run state. A retry can produce different output and +can incur another provider request. CLI exit behavior belongs to the +[CLI reference](../cli.md); HTTP status behavior belongs to the +[HTTP API reference](../api.md); package errors and results belong to the +[package contract](pkg-scriptorium.md). diff --git a/docs/consumers/pkg-scriptorium.md b/docs/consumers/pkg-scriptorium.md index 1b43cc7..b90b2f1 100644 --- a/docs/consumers/pkg-scriptorium.md +++ b/docs/consumers/pkg-scriptorium.md @@ -1,4 +1,4 @@ -# Package scriptorium +# Package `scriptorium` Import path: @@ -6,250 +6,158 @@ Import path: import "gitea.maximumdirect.net/eric/scriptorium" ``` -The root package is the public Go facade for Scriptorium's prompt prepare/run -workflow. It exposes typed requests, results, source options, injected LLM -clients, and stable public errors while keeping `internal/*` packages private. +This is the canonical public Go contract for in-process prompt preparation and +execution. Prompt, profile, and schema file formats are defined in the +[configuration reference](../config.md). -## Intended Use Cases +## Engine Construction -Use the package when a Go application needs: +`NewEngine(Config, ...Option)` constructs an engine. `Config` has these +fields: -- in-process prompt preparation or execution; -- typed request/result structs; -- direct `context.Context` cancellation; -- injected/fake LLM clients for tests; -- direct per-request `RunRequest.APIKey`. +| Field | Meaning | +| --- | --- | +| `PromptDir` | Prompt-definition directory, required unless a prompt source option is supplied. | +| `ProfileDir` | Optional custom profile directory over built-ins. | +| `SchemaDir` | Schema directory; empty uses `.`. | +| `Timeout` | Default timeout for the built-in OpenAI-compatible client. | +| `HTTPClient` | Optional HTTP client for that built-in client. | -Use [Subprocess integration](../integrations/subprocess.md) or the [HTTP API](../api.md) -when a process or service boundary is preferred. +Nil options are ignored. Invalid construction, including +`WithLLMClient(nil)`, returns an error matching `ErrInvalidConfig`. -## Construct An Engine +Source options replace their matching directory source: + +- prompts: `WithPromptFS(fsys, root)`, `WithPromptFile(path)`; +- profiles: `WithProfileFS(fsys, root)`, `WithProfileFile(path)`, and + `WithProfiles(profiles...)`; +- schemas: `WithSchemaFS(fsys, root)`, `WithSchemaFile(path)`; and +- LLM client: `WithLLMClient(client)`. + +`fs.FS` prompt-content and schema paths stay inside their configured roots. +A single-file option exposes that file by its base name. In-memory profiles take +precedence over an explicit or directory-backed profile source, which in turn +takes precedence over built-ins. File and filesystem sources use the format and +credential rules in the [configuration reference](../config.md). + +## Prepare And Run + +`Prepare(ctx, request)` resolves the prompt, profile, input artifacts, +validation contract, and rendered messages without calling an LLM. +`Run(ctx, request)` performs that preparation, calls the configured client, +and validates generated content. ```go engine, err := scriptorium.NewEngine(scriptorium.Config{ PromptDir: "./examples/prompts", ProfileDir: "./examples/profiles", - SchemaDir: "./examples/schemas", }) if err != nil { return err } -``` -`Config` fields: - -| Field | Description | -| --- | --- | -| `PromptDir` | Prompt definition directory. Required unless `WithPromptFS` or `WithPromptFile` is used. | -| `ProfileDir` | Optional custom profile directory overlaid above built-in profiles. | -| `SchemaDir` | Schema directory. Defaults to `.` when empty. | -| `Timeout` | Default timeout for the built-in OpenAI-compatible client. | -| `HTTPClient` | Optional HTTP client for the built-in OpenAI-compatible client. | - -`NewEngine` accepts `nil` options and ignores them. Invalid construction wraps -`ErrInvalidConfig`. - -## Source Options - -Directory fields are the compatibility path. Explicit source options override -the matching directory field. - -Prompt sources: - -- `WithPromptFS(fsys, root)` -- `WithPromptFile(path)` - -Profile sources: - -- `WithProfileFS(fsys, root)` -- `WithProfileFile(path)` -- `WithProfiles(profiles...)` - -Schema sources: - -- `WithSchemaFS(fsys, root)` -- `WithSchemaFile(path)` - -LLM source: - -- `WithLLMClient(client)` - -Source behavior: - -- Prompt and profile YAML use the same strict rules as directory loading. -- Prompt `content_file` values resolve relative to the prompt file. -- `fs.FS` roots are containment boundaries for prompt content files and schema paths. -- File options expose the selected file by its base name. -- Profile source precedence is in-memory profiles, then explicit profile file/FS/directory source, then built-ins. -- `WithLLMClient(nil)` returns `ErrInvalidConfig`. - -## In-Memory Profiles - -Use `WithProfiles` when the application already has typed model settings: - -```go -profile := scriptorium.OpenAICompatibleProfile(scriptorium.OpenAICompatibleProfileConfig{ - ID: "app.default", - Endpoint: "https://openrouter.ai/api/v1", - Model: "mistralai/mistral-small-3.2-24b-instruct", - APIKeyRequired: true, -}) - -engine, err := scriptorium.NewEngine(cfg, scriptorium.WithProfiles(profile)) -``` - -`Profile` and `OpenAICompatibleProfileConfig` include: - -- `ID` -- `Endpoint` -- `Model` -- `Temperature` -- `MaxTokens` -- `TopP` -- `TimeoutSeconds` -- `ServiceTier` -- `ReasoningEffort` -- `APIKeyRequired` -- `ExtraParams` - -`WithProfiles` rejects duplicate IDs in one call. In-memory profiles do not -store raw keys. When `APIKeyRequired` is true, pass the secret on each request -with `RunRequest.APIKey`. - -`ExtraParams` must be JSON-compatible: strings, booleans, finite numbers, -objects with string keys, arrays/slices, and nil. Unsupported values, non-string -map keys, non-finite floats, and cycles return `ErrInvalidConfig` for profiles -or `ErrInvalidRequest` for request overrides. - -## Prepare Workflow - -`Prepare` resolves prompt/profile/input/schema state and renders messages -without calling an LLM. - -```go prepared, err := engine.Prepare(ctx, scriptorium.RunRequest{ PromptID: "generic.markdown_summary", Inputs: map[string]scriptorium.ArtifactRef{ "transcript": scriptorium.File("./examples/fixtures/transcript.md"), - "glossary": scriptorium.File("./examples/fixtures/glossary.yml"), }, }) if err != nil { return err } -_ = prepared.EffectiveModelParams +_ = prepared.Messages ``` -`PreparedRun` includes prompt ID/version/hash, selected profile, effective -model params, output contract, structured-output metadata, input hashes, -rendered prompt hash, rendered messages, and timing fields. It does not include -raw API-key values, model output, validation results, or internal target -presence metadata. +The maintained package example is +[`examples/go-library/prepare`](../../examples/go-library/prepare). -## Run Workflow +`PreparedRun` exposes prompt, selected-profile, effective-model, output +contract, structured-output, input-hash, rendered-message, and timing +information. It does not include a resolved API key, model output, validation +result, or target-presence metadata. -`Run` calls `Prepare`, invokes the configured LLM client, builds the output -artifact, and validates the output. +`RunResult` adds run ID, artifact, raw output, validation, model metadata, +usage, and duration. Generated-content validation failures return a result with +`Validation.Status == ValidationFailed`; schema or validator runtime failures +return an error matching `ErrValidation`. -```go -result, err := engine.Run(ctx, scriptorium.RunRequest{ - PromptID: "generic.markdown_summary", - APIKey: apiKey, - Inputs: map[string]scriptorium.ArtifactRef{ - "transcript": scriptorium.File("./examples/fixtures/transcript.md"), - "glossary": scriptorium.File("./examples/fixtures/glossary.yml"), - }, -}) -if err != nil { - return err -} -_ = result.Artifact -``` +## Public Values -`RunResult` includes run ID, output artifact, raw output, validation result, -prompt/profile/model metadata, effective model params, input hashes, usage, and -timing fields. +`ArtifactRef` has `Type`, `URI`, and `Body`; `Artifact` has `Name`, +`ContentType`, `Body`, `URI`, `Size`, and `Hash`. `ExecutionTarget` exposes the +effective endpoint, model, numeric settings, credential-environment name, +service tier, reasoning effort, and extra parameters. `ValidationResult` +contains status, mode, errors, schema path, repair attempts, and validity. -Generated-content validation failures return a successful `RunResult` with -`Validation.Status == ValidationFailed`. Runtime/schema validation errors -return an error that matches `ErrValidation`. +The exported constants define these serialized values: -## Inputs +- artifact types: `inline` and `file`; +- output formats: `text`, `markdown`, and `json`; +- validation modes: `none`, `basic`, `json`, and `json_schema`; and +- validation statuses: `passed`, `failed`, and `skipped`. -Input helpers: +`TokenUsage` reports prompt, completion, total, cached, and cache-write token +counts. `RenderedPrompt`, `RenderedMessage`, `CacheControl`, and +`StructuredOutputSpec` are the public shapes used by injected LLM clients. -- `File(path)`: file-backed artifact reference. -- `Inline(body)`: inline artifact body. -- `InlineWithURI(uri, body)`: inline artifact body with URI metadata. +## Requests, Inputs, And Overrides -Input map keys must match the prompt's expected input names. +`RunRequest` fields are `PromptID`, `PromptVersion`, `ProfileID`, +`APIKey`, `Inputs`, `Vars`, `Execution`, `Validation`, and +`Metadata`. + +Input helpers are: + +- `File(path)` for a file-backed artifact; +- `Inline(body)` for inline content; and +- `InlineWithURI(uri, body)` for inline content with URI metadata. + +Required declared inputs must be supplied. Template rendering must also resolve +every input name the prompt actually references. Extra entries in `Inputs` +are not rejected solely because they are undeclared. + +`ExecutionTargetOverride` supplies endpoint, model, credential-environment, +service-tier, reasoning-effort, and extra-parameter overrides. Its numeric +fields (`Temperature`, `MaxTokens`, `TopP`, and `TimeoutSeconds`) are +pointers so explicit zero values are preserved. `OutputContract` supplies +`Format`, `ValidationMode`, `SchemaPath`, and `RepairAttempts`. + +`ExtraParams` accepts JSON-compatible values: strings, booleans, finite +numbers, objects with string keys, arrays or slices, and nil. Unsupported +values, non-string map keys, non-finite floats, and cycles return +`ErrInvalidConfig` for profiles or `ErrInvalidRequest` for request +overrides. + +## Profiles And Credentials + +`OpenAICompatibleProfile(OpenAICompatibleProfileConfig)` creates an +in-memory `Profile`. Its public fields are `ID`, `Endpoint`, `Model`, +`Temperature`, `MaxTokens`, `TopP`, `TimeoutSeconds`, `ServiceTier`, +`ReasoningEffort`, `APIKeyRequired`, and `ExtraParams`. +`WithProfiles` rejects duplicate IDs in one call. + +A direct `RunRequest.APIKey` is request-scoped and takes precedence over +`api_key_env` for the built-in client. It is excluded from JSON output and +from `PreparedRun` and `RunResult`. The package's `String` and +`GoString` methods report only whether a direct key is set. Do not use +reflection-based dumps of request structs, which can bypass that redaction. ## Injected LLM Clients -Use `WithLLMClient` for tests or custom model integrations: +`LLMClient` implements: ```go -type fakeLLM struct{} - -func (fakeLLM) Generate(ctx context.Context, req scriptorium.GenerateRequest) (*scriptorium.GenerateResponse, error) { - return &scriptorium.GenerateResponse{ - Content: "generated text", - Usage: scriptorium.TokenUsage{TotalTokens: 12}, - }, nil -} - -engine, err := scriptorium.NewEngine(cfg, scriptorium.WithLLMClient(fakeLLM{})) +Generate(context.Context, GenerateRequest) (*GenerateResponse, error) ``` -Injected clients receive: - -- rendered prompt; -- effective execution target; -- numeric target presence metadata; -- structured-output spec when applicable; -- direct request API key when provided. - -Custom clients should not log raw prompts or API keys by default. - -## Overrides And API Keys - -`RunRequest` fields: - -| Field | Description | -| --- | --- | -| `PromptID` | Prompt ID. | -| `PromptVersion` | Optional prompt version filter. | -| `ProfileID` | Optional profile override. | -| `APIKey` | Direct per-request API key. | -| `Inputs` | Input artifact references. | -| `Vars` | Template variables. | -| `Execution` | Per-request model overrides. | -| `Validation` | Per-request output contract override. | -| `Metadata` | Request metadata reserved for callers. | - -`RunRequest.Execution` uses pointer fields for numeric values so explicit zero -overrides are preserved: - -```go -zero := 0 -req.Execution = &scriptorium.ExecutionTargetOverride{ - MaxTokens: &zero, -} -``` - -Direct `RunRequest.APIKey` takes precedence over profile `api_key_env` for the -default OpenAI-compatible client. It is request-scoped, uses `json:"-"`, and is -not included in `PreparedRun` or `RunResult` JSON. Normal Go string formatting -of `RunRequest` and `GenerateRequest` reports only whether a direct key is set. - -Raw API keys do not belong in profile YAML, in-memory profiles, or app config. -Avoid reflection-based debug dumps of request structs because exported fields -remain visible to tools that bypass `String` and `GoString`. +Injected clients receive the rendered prompt, effective execution target, numeric +target-presence metadata, optional structured-output specification, and direct +request API key. `GenerateResponse` returns content and `TokenUsage`. +Custom clients should avoid logging raw prompts or credentials. ## Errors -Public methods wrap context while preserving stable sentinel checks with -`errors.Is`: +Public methods preserve these sentinel checks through `errors.Is`: - `ErrInvalidConfig` - `ErrInvalidRequest` @@ -262,23 +170,5 @@ Public methods wrap context while preserving stable sentinel checks with - `ErrLLMGenerate` - `ErrValidation` -Example: - -```go -if errors.Is(err, scriptorium.ErrPromptNotFound) { - return err -} -``` - -## Examples - -Run the maintained prepare-only example from the repository root: - -```bash -go run ./examples/go-library/prepare -``` - -See also: - -- [Configuration reference](../config.md) -- [Consumer integration overview](api.md) +For interface selection and operational responsibilities, see the +[consumer integration overview](api.md). diff --git a/docs/integrations/openai-compatible-chat.md b/docs/integrations/openai-compatible-chat.md index 8e05f85..18f81bc 100644 --- a/docs/integrations/openai-compatible-chat.md +++ b/docs/integrations/openai-compatible-chat.md @@ -1,118 +1,67 @@ # OpenAI-Compatible Chat Integration -## Scope +This is the outbound wire contract for Scriptorium's OpenAI-compatible +chat-completions client. -This document defines the outbound LLM contract implemented by `internal/llm/openai_compatible_client.go`. +## Endpoint And Method -It documents only fields and behaviors currently serialized by code. +Scriptorium uses the request endpoint override when present; otherwise it uses +the configured client base URL. It removes a trailing slash and sends +`POST /chat/completions`. -## Endpoint Construction +For example, `http://localhost:8000/v1` becomes +`http://localhost:8000/v1/chat/completions`. -Request endpoint is built as: +## Request Payload -1. choose base URL: - - `GenerateRequest.Target.Endpoint` if set - - otherwise client config `BaseURL` -2. trim trailing slash -3. append `/chat/completions` +The payload always contains `model` and rendered `messages`. It additionally +contains these fields when applicable: -Example: +| Field | Inclusion | +| --- | --- | +| `session_id` | Non-empty rendered prompt session ID. | +| `temperature` | Non-zero effective value or an explicit zero override. | +| `max_tokens` | Non-zero effective value or an explicit zero override. | +| `top_p` | Non-zero effective value or an explicit zero override. | +| `service_tier` | Any non-empty configured value. | +| `reasoning_effort` | Any non-empty configured value. | +| `response_format` | Structured output is requested. | +| provider-specific fields | Flattened from `extra_params`. | -- base URL: `http://localhost:8000/v1` -- final URL: `http://localhost:8000/v1/chat/completions` +`service_tier` and `reasoning_effort` are forwarded without a provider value +catalog; the selected backend decides which values it supports. -## Request Fields Sent +`extra_params` are top-level JSON fields, not a nested object. Keys cannot be +empty or collide with `model`, `session_id`, `messages`, `temperature`, +`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or +`response_format`. Values must be JSON-serializable. -Serialized JSON fields: +A rendered `session_id` is sent as a top-level JSON field, not as a header. +Empty values are omitted. The maximum length is 256 Unicode code points. -- `model` (required after fallback resolution) -- `session_id` (only when the rendered prompt includes a non-empty session ID) -- `messages` (rendered prompt messages) -- `temperature` (when non-zero, or when explicitly overridden to zero) -- `max_tokens` (when non-zero, or when explicitly overridden to zero) -- `top_p` (when non-zero, or when explicitly overridden to zero) -- `service_tier` (only when non-empty) -- `reasoning_effort` (only when non-empty) -- `response_format` (only when structured output is provided) -- profile/request `extra_params` as additional provider-specific top-level fields - -`service_tier` is provider-specific. OpenRouter currently documents request values such as `flex` and `priority`; Scriptorium forwards any non-empty configured value and lets the backend validate support. - -`reasoning_effort` is provider-specific. Scriptorium forwards any non-empty configured value as top-level `reasoning_effort` and lets the backend validate support. - -`extra_params` are flattened into the outbound JSON object. They are not wrapped in an `extra_params` object: - -```json -{ - "model": "gpt-4o-mini", - "messages": [ - { - "role": "user", - "content": "rendered text" - } - ], - "provider_route": "primary", - "provider_options": { - "retry_budget": 2 - } -} -``` - -`extra_params` values must be JSON-compatible. Supported value shapes include strings, numbers, booleans, objects, and arrays. - -Reserved `extra_params` keys are rejected before the HTTP request is made: - -- `model` -- `session_id` -- `messages` -- `temperature` -- `max_tokens` -- `top_p` -- `service_tier` -- `reasoning_effort` -- `response_format` - -Empty `extra_params` keys and values that cannot be encoded as JSON are also rejected before the HTTP request is made. - -`session_id` is rendered from prompt YAML using request variables and serialized as a top-level JSON request field. Scriptorium does not send an `x-session-id` header. Empty rendered session IDs are omitted, and values longer than 256 characters are rejected before the HTTP request. - -Messages without prompt cache control serialize with string `content`: +Messages without cache control use string `content`. A message with cache +control uses one text block: ```json { "role": "system", - "content": "rendered text" + "content": [{ + "type": "text", + "text": "rendered text", + "cache_control": {"type": "ephemeral", "ttl": "1h"} + }] } ``` -Messages with prompt cache control serialize as a single text content-block array: - -```json -{ - "role": "system", - "content": [ - { - "type": "text", - "text": "rendered text", - "cache_control": { - "type": "ephemeral", - "ttl": "1h" - } - } - ] -} -``` - -When cache-control `ttl` is unset in the prompt definition, `ttl` is omitted from the outbound payload. - -Structured output is currently `json_schema` only, serialized as: +When the prompt omits cache-control `ttl`, the payload omits `ttl`. +Structured JSON Schema output is sent as: ```json { "response_format": { "type": "json_schema", "json_schema": { - "name": "...", + "name": "schema name", "strict": true, "schema": {"type": "object"} } @@ -120,74 +69,40 @@ Structured output is currently `json_schema` only, serialized as: } ``` -## Authentication Header +## Authentication And Timeout -If `Target.APIKey` is set: +When a direct request API key is present, Scriptorium sends +`Authorization: Bearer ` and does not read `api_key_env`. Otherwise, it +resolves the configured non-empty `api_key_env` at request time and sends the +same header. If neither mechanism supplies a key, it sends no +`Authorization` header. -- set `Authorization: Bearer ` -- do not read `Target.APIKeyEnv` +The configured client timeout applies by default. A positive effective +`timeout_seconds` replaces it. An explicit request override of zero disables +the HTTP-client timeout; negative values are rejected before a request is sent. -If `Target.APIKey` is empty and `Target.APIKeyEnv` is set: +## Response Subset And Failures -- resolve environment variable value at request time -- set `Authorization: Bearer ` +A successful provider response must supply non-empty +`choices[0].message.content`. Scriptorium reads these optional or required +usage fields when present: -If the environment variable is unset/empty: - -- request fails before HTTP call (`ErrInvalidRequest`) - -If both `Target.APIKey` and `Target.APIKeyEnv` are empty: - -- no `Authorization` header is sent - -## Timeout Behavior - -Base timeout comes from client configuration. - -Per-request override: - -- if `Target.TimeoutSeconds > 0`, use that value for request timeout -- if `Target.TimeoutSeconds == 0` and the value came from an explicit request override, disable the HTTP client timeout -- if `Target.TimeoutSeconds < 0`, request is rejected (`ErrInvalidRequest`) - -## Response Expectations - -Expected successful response shape (subset used): - -- `choices[0].message.content` - `usage.prompt_tokens` - `usage.completion_tokens` - `usage.total_tokens` -- `usage.prompt_tokens_details.cached_tokens` (optional) -- `usage.cache_write_tokens` (optional) +- `usage.prompt_tokens_details.cached_tokens` +- `usage.cache_write_tokens` -Absent cache usage fields are treated as zero. Parsed cache usage is exposed through run results and adapter response surfaces as: +Missing cache usage is reported as zero. Invalid JSON, an empty choices array, +or empty first-choice content is a malformed provider response. Network and +request-construction failures, non-2xx responses, and malformed responses fail +the outbound call. Provider response bodies are not exposed by this client. -- `cached_tokens` -- `cache_write_tokens` +The client does not implement built-in retries, tool calls, top-level +`cache_control`, or multi-request payload modes. -Malformed response conditions include: +## Related References -- invalid JSON -- empty `choices` -- empty `choices[0].message.content` - -Malformed responses return `ErrMalformedResponse`. - -## Error Handling - -- network/request-construction failures: `ErrRequestFailed` -- non-2xx HTTP status: `ErrUnexpectedStatus` (includes status code; provider response bodies are not included) -- malformed response shape/content: `ErrMalformedResponse` - -## Unsupported Or Non-Serialized Fields - -The client does not serialize top-level `cache_control`. - -No built-in retries, tool-calls, or multi-request payload modes are implemented in this client. - -## Relationship To Runner - -When prompt validation mode is `json_schema`, runner prepares a structured-output schema spec and passes it to the client as `StructuredOutput`. - -The client only serializes the provider request payload; it does not load schema files itself. +Prompt schema preparation and runner orchestration are described in +[runner internals](../internal/runner.md). Prompt and profile configuration is +defined by the [configuration reference](../config.md). diff --git a/docs/integrations/subprocess.md b/docs/integrations/subprocess.md index 8b6eeed..3426caf 100644 --- a/docs/integrations/subprocess.md +++ b/docs/integrations/subprocess.md @@ -1,130 +1,41 @@ # Subprocess Integration -This document defines the supported subprocess contract for downstream -applications invoking Scriptorium through the public CLI. +This document covers process-boundary behavior for callers that invoke +Scriptorium as a child process. Command syntax, flags, output, and exit codes +are defined by the [CLI reference](../cli.md). Interface selection belongs in +the [consumer integration overview](../consumers/api.md). -This is a CLI contract. Go callers that want an in-process typed API should use -the [package guide](../consumers/pkg-scriptorium.md). +## Process Contract -## Supported Commands +Use `scriptorium render` when the caller needs prepared output without a model +call, and `scriptorium run` for generation. Pass an explicit `--config` or +make the configuration search paths available to the child process; configuration +discovery, fields, profile selection, and credential mechanisms are defined in +the [configuration reference](../config.md). -Downstream applications should invoke: +Pass required API-key environment variables through the child environment. Do +not place raw API keys in arguments. Keep the environment limited to the values +needed for the selected profile. -- `scriptorium render` for preflight/debug output without LLM execution. -- `scriptorium run` for generation. +## Streams And Output Ownership -`scriptorium serve` is an HTTP service command, not the recommended subprocess -contract for per-request execution. +Capture stdout and stderr separately. Stdout contains the requested artifact or +prepared output unless the caller selects an output file; stderr contains +summaries, diagnostics, and server messages. The exact destinations and status +meanings are part of the [CLI reference](../cli.md), not a stable stderr data +protocol. -## Recommended Invocation Shapes +When using `--out`, the caller owns the output path, its permissions, and +cleanup. Treat rendered prompts, generated artifacts, stdout, and stderr as +potentially sensitive. -Render: +## Cancellation And Recovery -```bash -scriptorium render \ - --config \ - --prompt \ - --input transcript= \ - --format json -``` +A CLI invocation performs one synchronous request and creates no durable run +state. A supervising process that needs cancellation must terminate the child +process according to its own process-management policy. A later invocation is a +new request and can make another model call; there is no resume or checkpoint +protocol. -Run: - -```bash -scriptorium run \ - --config \ - --prompt \ - --input transcript= \ - --out -``` - -Callers may add: - -- `--profile ` -- repeatable `--input name=path` -- repeatable `--var name=value` -- runtime overrides when explicitly needed, such as `--model`, `--llm-base-url`, `--api-key-env`, and `--timeout` - -Do not pass raw API keys as command arguments. - -## Config And Directory Behavior - -Callers can rely on resolved app config or pass explicit paths. - -Default config search order: - -1. `/usr/local/etc/scriptorium/config.yml` -2. `/etc/scriptorium/config.yml` - -Rules: - -- Explicit `--config` requires file existence and valid syntax. -- CLI flags override config values. -- `run` and `render` require an effective `prompt_dir`. -- `profile_dir` is optional because built-in profiles are available. - -## Profile Selection - -Profile selection follows runner behavior: - -1. explicit `--profile` -2. prompt `default_profile` -3. error if neither is available - -Treat prompt and profile IDs as deployment configuration, not hardcoded business -logic. - -## Input And Variable Contract - -- Inputs use repeated `--input name=path`. -- Input names must match prompt definition input names. -- Variables use repeated `--var name=value`. -- Both flags also accept comma-separated mappings. -- Prefer file inputs for large content. - -CLI inputs are file references. HTTP-only `inline` references are documented in -the [HTTP API reference](../api.md). - -## Environment Contract - -- Pass through required API-key environment variables referenced by `api_key_env`. -- Keep subprocess environments scoped to required variables. -- Use `--api-key-env` only to name an environment variable. -- Never pass raw API keys via argv. - -## Stdout And Stderr - -`run`: - -- stdout: generated artifact body unless `--out` is used. -- stderr: success summary and errors. - -`render`: - -- stdout: prepared-run output unless `--out` is used. -- stderr: errors. - -Capture stdout and stderr separately. Do not parse stderr as a stable data -format beyond exit status handling. - -## Exit Status Contract - -- `0`: success. -- `1`: parse, config, load, render, generation, IO, or runtime error. -- `2`: `run` completed and output was written, but validation failed. - -A `run` exit code `2` can still produce output on stdout or at `--out`. -Consumers must decide whether to keep or discard that output. - -## Security Notes - -- Treat generated artifacts, rendered prompts, stdout, and stderr as potentially sensitive. -- Use controlled output paths and access controls for persisted artifacts. -- Avoid logging full rendered prompts or generated artifacts by default. - -## Canonical References - -- CLI behavior: [CLI reference](../cli.md) -- Config and file formats: [Configuration reference](../config.md) -- Operations: [Operations guide](../operations.md) -- Troubleshooting: [Troubleshooting](../troubleshooting.md) +For deployment, filesystem permissions, and sensitive-artifact handling, see +the [operations guide](../operations.md).