Compare commits
6 Commits
e0b1d6a0dc
...
31faaf4259
| Author | SHA1 | Date | |
|---|---|---|---|
| 31faaf4259 | |||
| 9932153b97 | |||
| f0ca233c25 | |||
| ff31f8daf8 | |||
| c927b7819d | |||
| 6d1fb66dd7 |
12
README.md
12
README.md
@@ -22,6 +22,7 @@ go run ./cmd/scriptorium render \
|
||||
```
|
||||
|
||||
This command renders the prepared prompt and effective runtime settings without calling an LLM.
|
||||
For complete invocation and output behavior, see the [CLI reference](docs/cli.md).
|
||||
|
||||
## Documentation
|
||||
|
||||
@@ -29,7 +30,6 @@ This command renders the prepared prompt and effective runtime settings without
|
||||
- [Configuration reference](docs/config.md)
|
||||
- [HTTP API reference](docs/api.md)
|
||||
- [Operations guide](docs/operations.md)
|
||||
- [Troubleshooting](docs/troubleshooting.md)
|
||||
- [Consumer integration overview](docs/consumers/api.md)
|
||||
- [Go library package](docs/consumers/pkg-scriptorium.md)
|
||||
- [Subprocess integration](docs/integrations/subprocess.md)
|
||||
@@ -38,8 +38,8 @@ This command renders the prepared prompt and effective runtime settings without
|
||||
|
||||
## Examples
|
||||
|
||||
- `examples/config.yml`
|
||||
- `examples/config.full.yml`
|
||||
- `examples/render-markdown-summary.sh`
|
||||
- `examples/http-run.json`
|
||||
- `examples/go-library/prepare`
|
||||
- [Minimal configuration](examples/config.yml) and [complete configuration](examples/config.full.yml)
|
||||
- [Prompt definitions](examples/prompts/), [execution profiles](examples/profiles/), [schemas](examples/schemas/), and [synthetic input fixtures](examples/fixtures/)
|
||||
- [Render script](examples/render-markdown-summary.sh)
|
||||
- [HTTP request](examples/http-run.json)
|
||||
- [Go library example](examples/go-library/prepare/main.go)
|
||||
|
||||
49
docs/adr/0001-adopt-canonical-documentation-ownership.md
Normal file
49
docs/adr/0001-adopt-canonical-documentation-ownership.md
Normal file
@@ -0,0 +1,49 @@
|
||||
# ADR 0001: Adopt Canonical Documentation Ownership
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Date
|
||||
|
||||
2026-07-26
|
||||
|
||||
## Context
|
||||
|
||||
Scriptorium's documentation grew alongside its CLI, HTTP, public Go, and
|
||||
integration interfaces. As a result, several documents repeated mutable
|
||||
contracts such as flags, configuration fields, and status behavior. Those
|
||||
parallel definitions made it unclear which document to update when behavior
|
||||
changed and increased the risk of documentation drift.
|
||||
|
||||
## Decision
|
||||
|
||||
Assign each documentation topic one canonical owner, as defined in
|
||||
[`docs/policy/documentation.md`](../policy/documentation.md). Non-owning
|
||||
documents may provide short orientation and links, but do not redefine volatile
|
||||
contracts. Current behavior is documented outside `docs/roadmap/`; roadmaps own
|
||||
future work, sequencing, and implementation status.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
- Keep broad reference material in several audience-specific documents. This
|
||||
would preserve local convenience but leave conflicting contract definitions
|
||||
likely.
|
||||
- Consolidate all documentation into one reference. This would reduce duplicate
|
||||
text but would not serve the distinct needs of users, operators, consumers,
|
||||
and contributors.
|
||||
|
||||
## Rationale
|
||||
|
||||
Canonical ownership retains audience-specific guidance while making the source
|
||||
of truth for each contract discoverable. It also makes documentation changes
|
||||
reviewable alongside the implementation change that requires them.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Changes to behavior must update the canonical owner in the same change.
|
||||
- Cross-cutting documentation links to the owner instead of copying its
|
||||
details.
|
||||
- Documentation restructuring follows the implementation sequence in
|
||||
[`docs/roadmap/documentation.md`](../roadmap/documentation.md); the roadmap,
|
||||
not this ADR, records completion status.
|
||||
334
docs/api.md
334
docs/api.md
@@ -2,296 +2,142 @@
|
||||
|
||||
This is the canonical public HTTP contract for Scriptorium.
|
||||
|
||||
Implemented route:
|
||||
## Service And Route
|
||||
|
||||
- `POST /v1/runs`
|
||||
`POST /v1/runs` runs one prompt request and returns generated output,
|
||||
validation, and metadata. The service has no built-in authentication or
|
||||
authorization; deploy it behind appropriate network and authentication controls.
|
||||
|
||||
For CLI behavior, see [CLI reference](cli.md). For config and prompt/profile
|
||||
file formats, see [Configuration reference](config.md).
|
||||
The service address and HTTP limits are configured as described in the
|
||||
[configuration reference](config.md). `serve` invocation is defined in the
|
||||
[CLI reference](cli.md).
|
||||
|
||||
The maintained request-shape example is `examples/http-run.json`. It requires a
|
||||
running `serve` process with an artifact root that can read the referenced
|
||||
files, plus a reachable model endpoint for full execution.
|
||||
|
||||
## Base URL And Deployment
|
||||
|
||||
`scriptorium serve` listens on `server.addr` or `serve --addr`. The default is
|
||||
`:8080`.
|
||||
|
||||
The route path is always:
|
||||
|
||||
```text
|
||||
/v1/runs
|
||||
```
|
||||
|
||||
The HTTP adapter has no built-in authentication or authorization. Deploy it
|
||||
behind trusted network and authentication controls.
|
||||
|
||||
## Media Types
|
||||
|
||||
- Request body: JSON object.
|
||||
- Response body: JSON object.
|
||||
- Response `Content-Type`: `application/json`.
|
||||
|
||||
Requests are decoded as JSON regardless of the request `Content-Type` header.
|
||||
There are no shared query parameters.
|
||||
Requests and responses are JSON objects. Requests are decoded as JSON regardless
|
||||
of their `Content-Type`; successful JSON responses use
|
||||
`Content-Type: application/json`. There are no query parameters.
|
||||
|
||||
## Request Limits
|
||||
|
||||
HTTP limits are configured through `server.*` config fields or `serve` flags:
|
||||
The configured request-body limit includes inline artifact bodies. The artifact
|
||||
limit applies to HTTP `file` inputs. The response limit applies to the encoded
|
||||
response, including the artifact body and optional raw output. A limit of zero
|
||||
disables that limit.
|
||||
|
||||
- `server.max_request_bytes`: encoded JSON request body limit, including inline input bodies.
|
||||
- `server.max_artifact_bytes`: file artifact limit for HTTP `file` input references.
|
||||
- `server.max_response_bytes`: encoded JSON response limit, including artifact body and optional raw output.
|
||||
|
||||
Each limit defaults to `16777216` bytes. `0` disables that limit.
|
||||
A request body over its limit returns `413 request_too_large`; an oversized
|
||||
file input returns `413 artifact_too_large`; an oversized encoded response
|
||||
returns `413 response_too_large`.
|
||||
|
||||
## `POST /v1/runs`
|
||||
|
||||
Runs one prompt request and returns the generated artifact, validation result,
|
||||
and metadata.
|
||||
|
||||
### Request Body
|
||||
|
||||
The maintained [request example](../examples/http-run.json) is a complete
|
||||
copyable shape. The smallest valid shape is:
|
||||
|
||||
```json
|
||||
{
|
||||
"prompt_id": "generic.markdown_summary",
|
||||
"profile_id": "local-fast",
|
||||
"prompt_version": "1.0.0",
|
||||
"inputs": {
|
||||
"transcript": {
|
||||
"type": "file",
|
||||
"uri": "./examples/fixtures/transcript.md"
|
||||
},
|
||||
"glossary": {
|
||||
"type": "inline",
|
||||
"body": "party:\n - Rin"
|
||||
"transcript": {"type": "inline", "body": "Source text"}
|
||||
}
|
||||
},
|
||||
"vars": {
|
||||
"session_date": "2026-05-04"
|
||||
},
|
||||
"model": {
|
||||
"endpoint": "http://localhost:8000/v1",
|
||||
"model": "gpt-4o-mini",
|
||||
"temperature": 0,
|
||||
"max_tokens": 800,
|
||||
"top_p": 1,
|
||||
"timeout_seconds": 120,
|
||||
"service_tier": "priority",
|
||||
"reasoning_effort": "medium",
|
||||
"api_key_env": "SCRIPTORIUM_API_KEY",
|
||||
"extra_params": {
|
||||
"provider_option": "enabled"
|
||||
}
|
||||
},
|
||||
"include_raw_output": false
|
||||
}
|
||||
```
|
||||
|
||||
Request fields:
|
||||
|
||||
| Field | Required | Description |
|
||||
| Field | Required | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `prompt_id` | yes | Prompt ID. Must not be blank. |
|
||||
| `prompt_id` | yes | Non-blank prompt ID. |
|
||||
| `prompt_version` | no | Prompt version filter. |
|
||||
| `profile_id` | no | Execution profile ID. If omitted, the prompt must define `default_profile`. |
|
||||
| `inputs` | yes | Object mapping prompt input names to input references. Must contain at least one entry. |
|
||||
| `vars` | no | Object mapping template variable names to string values. |
|
||||
| `model` | no | Runtime model override object. |
|
||||
| `include_raw_output` | no | When `true`, include `raw_model_output` in the response. |
|
||||
| `profile_id` | no | Execution-profile ID; otherwise the prompt must set `default_profile`. |
|
||||
| `inputs` | yes | Non-empty object mapping input names to references. |
|
||||
| `vars` | no | Object mapping template-variable names to strings. |
|
||||
| `model` | no | Runtime model-override object. |
|
||||
| `include_raw_output` | no | Include `raw_model_output` when true. |
|
||||
|
||||
Input reference fields:
|
||||
An input reference has a required `type` of `file` or `inline`. A `file`
|
||||
reference requires `uri`; an `inline` reference requires `body`.
|
||||
|
||||
| Field | Required | Description |
|
||||
| --- | --- | --- |
|
||||
| `type` | yes | `file` or `inline`. |
|
||||
| `uri` | for `file` | File URI/path. |
|
||||
| `body` | for `inline` | Inline artifact body. |
|
||||
HTTP file references require a configured artifact root. Relative paths resolve
|
||||
within that root. Absolute paths must be lexically within it; traversal outside
|
||||
it is rejected with `400 artifact_not_allowed`. This lexical check does not
|
||||
resolve symlinks: the operating system follows symlinks inside the root,
|
||||
including ones that target outside it. Keep the root narrow and inaccessible to
|
||||
untrusted writers.
|
||||
|
||||
HTTP `file` references require `server.artifact_root` or `serve
|
||||
--artifact-root`. Relative file URIs resolve against that root. Absolute file
|
||||
URIs are accepted only when lexically inside the root. Relative traversal and
|
||||
absolute paths outside the root return `400 artifact_not_allowed`.
|
||||
The optional `model` object accepts `endpoint`, `model`, `temperature`,
|
||||
`max_tokens`, `top_p`, `timeout_seconds`, `service_tier`,
|
||||
`reasoning_effort`, `api_key_env`, and `extra_params`. Numeric ranges and
|
||||
credential supply are defined by the [configuration reference](config.md).
|
||||
Explicit zero values for the numeric fields are overrides; zero
|
||||
`timeout_seconds` disables the outbound client timeout.
|
||||
|
||||
The containment check is lexical and does not resolve symlinks. Symlinks inside
|
||||
the artifact root are followed by the operating system, including symlinks that
|
||||
point outside the root. Keep the artifact root narrow and not writable by
|
||||
untrusted users.
|
||||
Raw API-key values are not accepted. `api_key` and any other unknown model
|
||||
field cause `400 invalid_json`.
|
||||
|
||||
Model override fields:
|
||||
### Strict JSON
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `endpoint` | Runtime endpoint override. |
|
||||
| `model` | Runtime model override. |
|
||||
| `temperature` | Number in range `0..2`. Explicit `0` is an override. |
|
||||
| `max_tokens` | Integer greater than or equal to `0`. Explicit `0` is an override. |
|
||||
| `top_p` | Number in range `0..1`. Explicit `0` is an override. |
|
||||
| `timeout_seconds` | Integer greater than or equal to `0`. Explicit `0` disables the outbound client timeout. |
|
||||
| `service_tier` | Provider-specific request tier. |
|
||||
| `reasoning_effort` | Provider-specific reasoning setting. |
|
||||
| `api_key_env` | Name of an environment variable containing the API key. |
|
||||
| `extra_params` | JSON-compatible provider-specific top-level request fields. |
|
||||
|
||||
Raw API-key values are not accepted in HTTP payloads. A field such as
|
||||
`api_key` is rejected as unknown JSON.
|
||||
|
||||
`extra_params` keys must not be empty and must not collide with reserved
|
||||
outbound fields: `model`, `session_id`, `messages`, `temperature`,
|
||||
`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or
|
||||
`response_format`.
|
||||
|
||||
### Strict JSON Rules
|
||||
|
||||
Request decoding is strict:
|
||||
|
||||
- malformed JSON returns `400 invalid_json`
|
||||
- unknown request fields return `400 invalid_json`
|
||||
- unknown `inputs` item fields return `400 invalid_json`
|
||||
- unknown `model` fields return `400 invalid_json`
|
||||
- trailing JSON tokens after the request object return `400 invalid_json`
|
||||
- request bodies above the configured limit return `413 request_too_large`
|
||||
Request decoding rejects malformed JSON, unknown fields at every request level,
|
||||
and trailing JSON tokens with `400 invalid_json`. A blank `prompt_id` or
|
||||
empty `inputs` object returns `400 invalid_request`.
|
||||
|
||||
### Success Response
|
||||
|
||||
Status: `200 OK`
|
||||
A completed run returns `200 OK`, including when generated content fails its
|
||||
validation contract. The response contains:
|
||||
|
||||
```json
|
||||
{
|
||||
"artifact": {
|
||||
"name": "output",
|
||||
"content_type": "text/markdown",
|
||||
"body": "Generated content",
|
||||
"size": 17,
|
||||
"hash": "..."
|
||||
},
|
||||
"validation": {
|
||||
"status": "passed",
|
||||
"mode": "basic",
|
||||
"repair_attempts": 0,
|
||||
"is_valid": true
|
||||
},
|
||||
"metadata": {
|
||||
"run_id": "...",
|
||||
"prompt_id": "generic.markdown_summary",
|
||||
"prompt_version": "1.0.0",
|
||||
"prompt_hash": "...",
|
||||
"rendered_prompt_hash": "...",
|
||||
"selected_profile_id": "local-fast",
|
||||
"model_name": "gpt-4o-mini",
|
||||
"endpoint": "http://localhost:8000/v1",
|
||||
"model_params": {
|
||||
"endpoint": "http://localhost:8000/v1",
|
||||
"model": "gpt-4o-mini",
|
||||
"temperature": 0.2,
|
||||
"max_tokens": 500,
|
||||
"top_p": 1,
|
||||
"timeout_seconds": 90
|
||||
},
|
||||
"input_hashes": {
|
||||
"transcript": "..."
|
||||
},
|
||||
"usage": {
|
||||
"prompt_tokens": 11,
|
||||
"completion_tokens": 22,
|
||||
"total_tokens": 33,
|
||||
"cached_tokens": 0,
|
||||
"cache_write_tokens": 0
|
||||
},
|
||||
"start_time": "2026-05-04T12:00:00Z",
|
||||
"end_time": "2026-05-04T12:00:01Z",
|
||||
"duration_ms": 1000,
|
||||
"validation_mode": "basic",
|
||||
"validation_status": "passed",
|
||||
"repair_attempts_used": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
- `artifact`: `name`, `content_type`, `body`, `size`, `hash`, and
|
||||
optional `uri`;
|
||||
- `validation`: `status`, `mode`, `repair_attempts`, `is_valid`, plus
|
||||
optional `errors` and `schema_path`;
|
||||
- `metadata`: run, prompt, rendered-prompt, profile, model, input-hash, usage,
|
||||
timing, validation, and repair-attempt metadata; and
|
||||
- optional `raw_model_output` when requested.
|
||||
|
||||
Response fields:
|
||||
`metadata.model_params` has `endpoint`, `model`, `temperature`,
|
||||
`max_tokens`, `top_p`, and `timeout_seconds`, plus optional
|
||||
`service_tier`, `reasoning_effort`, `api_key_env`, and `extra_params`.
|
||||
`metadata.usage` always includes `prompt_tokens`, `completion_tokens`,
|
||||
`total_tokens`, `cached_tokens`, and `cache_write_tokens`; unavailable
|
||||
cache usage is reported as zero.
|
||||
|
||||
- `artifact`: generated output artifact.
|
||||
- `validation`: validation result for the generated artifact.
|
||||
- `metadata`: run and effective runtime metadata.
|
||||
- `raw_model_output`: omitted unless `include_raw_output` is `true`.
|
||||
|
||||
`artifact.uri` is omitted when empty. `validation.errors` and
|
||||
`validation.schema_path` are omitted when empty. `model_params.service_tier`,
|
||||
`model_params.reasoning_effort`, `model_params.api_key_env`, and
|
||||
`model_params.extra_params` are omitted when empty.
|
||||
|
||||
`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are
|
||||
always present as numbers. They are `0` when the provider omits compatible cache
|
||||
usage fields or reports no cache activity.
|
||||
|
||||
### Validation Failure Response
|
||||
|
||||
Generated-content validation failures still return `200 OK`.
|
||||
|
||||
```json
|
||||
{
|
||||
"validation": {
|
||||
"status": "failed",
|
||||
"mode": "json",
|
||||
"errors": ["invalid JSON: ..."],
|
||||
"repair_attempts": 0,
|
||||
"is_valid": false
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The response still includes `artifact` and `metadata`.
|
||||
A validation failure has `validation.status: "failed"`, `is_valid: false`,
|
||||
and any available diagnostic errors, while still returning the artifact and
|
||||
metadata.
|
||||
|
||||
## Error Responses
|
||||
|
||||
Error body shape:
|
||||
Errors have this shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"error": {
|
||||
"code": "invalid_request",
|
||||
"message": "prompt_id is required"
|
||||
}
|
||||
}
|
||||
{"error":{"code":"invalid_request","message":"prompt_id is required"}}
|
||||
```
|
||||
|
||||
Current status/code mapping:
|
||||
Messages are concise and do not expose wrapped internal causes.
|
||||
|
||||
| Status | Code | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `400` | `invalid_json` | Malformed JSON, unknown JSON field, or trailing JSON token. |
|
||||
| `400` | `invalid_request` | Missing/invalid request fields or invalid runtime overrides. |
|
||||
| `400` | `profile_required` | No `profile_id` and prompt has no `default_profile`. |
|
||||
| `400` | `prompt_load_failed` | Prompt definition YAML/contract failed to load. |
|
||||
| `400` | `profile_load_failed` | Profile YAML/contract failed to load, including raw `api_key`. |
|
||||
| `400` | `artifact_not_allowed` | HTTP file refs are disabled or requested path is outside artifact root. |
|
||||
| `400` | `artifact_read_failed` | Input artifact could not be read or input ref was unsupported/invalid. |
|
||||
| `400` | `invalid_json` | Malformed JSON, unknown field, or trailing JSON. |
|
||||
| `400` | `invalid_request` | Missing or invalid request data or runtime override. |
|
||||
| `400` | `profile_required` | No profile ID and no prompt default profile. |
|
||||
| `400` | `prompt_load_failed` | Prompt definition failed to load. |
|
||||
| `400` | `profile_load_failed` | Profile failed to load. |
|
||||
| `400` | `artifact_not_allowed` | HTTP file input is disabled or outside the artifact root. |
|
||||
| `400` | `artifact_read_failed` | Input artifact is invalid or cannot be read. |
|
||||
| `400` | `prompt_render_failed` | Prompt template rendering failed. |
|
||||
| `400` | `api_key_env_missing` | Selected `api_key_env` variable is unset or empty. |
|
||||
| `404` | `not_found` | Route path is unknown. |
|
||||
| `404` | `prompt_not_found` | Prompt ID/version was not found. |
|
||||
| `404` | `profile_not_found` | Profile ID was not found. |
|
||||
| `405` | `method_not_allowed` | Method is not `POST` on `/v1/runs`. |
|
||||
| `413` | `request_too_large` | Encoded JSON request body exceeds configured request limit. |
|
||||
| `413` | `artifact_too_large` | HTTP file input artifact exceeds configured artifact limit. |
|
||||
| `413` | `response_too_large` | Encoded JSON response exceeds configured response limit. |
|
||||
| `500` | `validation_runtime_failed` | Validator runtime/schema loading failed. |
|
||||
| `500` | `internal_error` | Unclassified server error. |
|
||||
| `400` | `api_key_env_missing` | The selected credential environment variable is unset or empty. |
|
||||
| `404` | `not_found` | Route does not exist. |
|
||||
| `404` | `prompt_not_found` | Prompt ID or version does not exist. |
|
||||
| `404` | `profile_not_found` | Profile ID does not exist. |
|
||||
| `405` | `method_not_allowed` | The route does not accept the method. |
|
||||
| `413` | `request_too_large` | Encoded request exceeds its limit. |
|
||||
| `413` | `artifact_too_large` | File input exceeds its limit. |
|
||||
| `413` | `response_too_large` | Encoded response exceeds its limit. |
|
||||
| `500` | `validation_runtime_failed` | Schema or validator runtime failure. |
|
||||
| `500` | `internal_error` | Unclassified server failure. |
|
||||
| `502` | `llm_failed` | Outbound model request failed. |
|
||||
|
||||
HTTP error messages are intentionally concise and do not include sensitive
|
||||
internal causes.
|
||||
|
||||
## Retry And Idempotency
|
||||
|
||||
Scriptorium does not provide idempotency keys, pagination, caching headers, or
|
||||
rate limiting.
|
||||
|
||||
Clients may retry transport failures or `5xx` responses when their surrounding
|
||||
workflow can tolerate another model call. A retry can generate different output
|
||||
and incur another provider request.
|
||||
|
||||
## Example File
|
||||
|
||||
- `examples/http-run.json`
|
||||
Scriptorium provides no idempotency keys, pagination, caching headers, or rate
|
||||
limits. Clients may retry transport failures or `5xx` responses only when
|
||||
their workflow tolerates another model call: a retry can produce different
|
||||
output and incur another provider request.
|
||||
|
||||
277
docs/cli.md
277
docs/cli.md
@@ -1,5 +1,10 @@
|
||||
# CLI Reference
|
||||
|
||||
This is the canonical contract for invoking Scriptorium. Configuration discovery,
|
||||
precedence, directories, profiles, and schemas are defined in the
|
||||
[configuration reference](config.md). The [HTTP API reference](api.md) owns
|
||||
service request and response behavior.
|
||||
|
||||
## Shortest Useful Command
|
||||
|
||||
```bash
|
||||
@@ -10,224 +15,132 @@ go run ./cmd/scriptorium render \
|
||||
--input glossary=./examples/fixtures/glossary.yml
|
||||
```
|
||||
|
||||
`render` prepares the prompt, loads input artifacts, resolves the execution
|
||||
profile, and prints the prepared request without calling an LLM.
|
||||
`render` prepares a request without calling an LLM.
|
||||
|
||||
## Command Overview
|
||||
## Commands
|
||||
|
||||
- `scriptorium run`: prepare a prompt, call the configured LLM, write generated output, and print a run summary.
|
||||
- `scriptorium render`: prepare a prompt only; write prepared-run output as `text` or `json`.
|
||||
- `scriptorium serve`: start the HTTP server for `POST /v1/runs`.
|
||||
- `scriptorium run`: prepare a prompt, call the configured LLM, and write the
|
||||
generated artifact.
|
||||
- `scriptorium render`: prepare a prompt and write prepared-run output.
|
||||
- `scriptorium serve`: start the HTTP server.
|
||||
|
||||
Canonical related references:
|
||||
All commands accept `--config <path>` and reject positional arguments. An
|
||||
effective `prompt_dir` is required for every command. Supply it through the
|
||||
configuration contract or the command's `--prompt-dir` flag.
|
||||
|
||||
- [Configuration reference](config.md)
|
||||
- [HTTP API reference](api.md)
|
||||
- [Subprocess integration](integrations/subprocess.md)
|
||||
## `scriptorium run`
|
||||
|
||||
## Common Rules
|
||||
|
||||
- `--config` is supported by `run`, `render`, and `serve`.
|
||||
- Positional arguments are rejected.
|
||||
- `run` and `render` require `--prompt`, at least one `--input`, and an effective `prompt_dir`.
|
||||
- `serve` requires an effective `prompt_dir`.
|
||||
- `profile_dir` is optional. Without it, only built-in profiles are available.
|
||||
- If `profile_dir` is set, custom profiles override built-in profiles with the same ID.
|
||||
- Prompt cache control, `session_id`, structured output, and provider-specific profile fields are configured in YAML, not with CLI flags.
|
||||
|
||||
Config precedence is:
|
||||
|
||||
1. built-in defaults
|
||||
2. config file values
|
||||
3. CLI flags
|
||||
|
||||
## Flag Reference
|
||||
|
||||
### `scriptorium run`
|
||||
|
||||
```bash
|
||||
```text
|
||||
scriptorium run [flags]
|
||||
```
|
||||
|
||||
Required through flags or config:
|
||||
Required flags:
|
||||
|
||||
- `--prompt-dir <dir>`: prompt definition directory.
|
||||
|
||||
Required as flags:
|
||||
|
||||
- `--prompt <id>`: prompt ID to execute.
|
||||
- `--input name=path`: input file mapping. Repeat or use comma-separated mappings.
|
||||
| Flag | Meaning |
|
||||
| --- | --- |
|
||||
| `--prompt <id>` | Prompt ID to execute. |
|
||||
| `--input name=path` | Input file mapping; repeat or use comma-separated mappings. |
|
||||
|
||||
Optional flags:
|
||||
|
||||
- `--config <path>`: application config file.
|
||||
- `--profile-dir <dir>`: custom profile definition directory.
|
||||
- `--schema-dir <dir>`: schema base directory for `json_schema` validation.
|
||||
- `--profile <id>`: execution profile override. If omitted, the prompt `default_profile` is used.
|
||||
- `--var name=value`: template variable mapping. Repeat or use comma-separated mappings.
|
||||
- `--out <path>`: write generated artifact body to a file instead of stdout.
|
||||
- `--llm-base-url <url>`: runtime endpoint override.
|
||||
- `--model <name>`: runtime model override.
|
||||
- `--api-key-env <name>`: runtime API-key environment variable name override.
|
||||
- `--temperature <float>`: runtime temperature override.
|
||||
- `--max-tokens <int>`: runtime max tokens override.
|
||||
- `--top-p <float>`: runtime top-p override.
|
||||
- `--timeout <duration>`: runtime timeout override using Go duration syntax, such as `30s` or `2m`.
|
||||
| Flag | Meaning |
|
||||
| --- | --- |
|
||||
| `--config <path>` | Application configuration file. |
|
||||
| `--prompt-dir <dir>` | Prompt-definition directory override. |
|
||||
| `--profile-dir <dir>` | Custom profile-directory override. |
|
||||
| `--schema-dir <dir>` | Schema base-directory override. |
|
||||
| `--profile <id>` | Execution-profile override. |
|
||||
| `--var name=value` | Template-variable mapping; repeat or use comma-separated mappings. |
|
||||
| `--out <path>` | Write generated content to this file instead of stdout. |
|
||||
| `--llm-base-url <url>` | Runtime endpoint override. |
|
||||
| `--model <name>` | Runtime model override. |
|
||||
| `--api-key-env <name>` | Runtime API-key environment-variable name override. |
|
||||
| `--temperature <float>` | Runtime temperature override. |
|
||||
| `--max-tokens <int>` | Runtime maximum-token override. |
|
||||
| `--top-p <float>` | Runtime top-p override. |
|
||||
| `--timeout <duration>` | Runtime timeout override using Go duration syntax. |
|
||||
|
||||
Deprecated aliases:
|
||||
Deprecated aliases: `--prompt-id` for `--prompt`, and `--profile-id` for
|
||||
`--profile`.
|
||||
|
||||
- `--prompt-id <id>`: alias for `--prompt`.
|
||||
- `--profile-id <id>`: alias for `--profile`.
|
||||
Omitted numeric runtime flags preserve the selected effective value; explicit
|
||||
zero values override it. `--timeout 0s` disables the outbound HTTP-client
|
||||
timeout. CLI durations are converted to whole seconds by truncation toward
|
||||
zero, so any duration whose absolute value is below one second becomes an
|
||||
explicit zero-second override.
|
||||
|
||||
Runtime override notes:
|
||||
There is no raw API-key flag. Use `--api-key-env`.
|
||||
|
||||
- Omitted numeric override flags preserve the selected profile/default value.
|
||||
- Explicit zero values override the selected profile/default value.
|
||||
- `--timeout 0s` disables the outbound HTTP client timeout for that request.
|
||||
- There is no raw API-key flag; use `--api-key-env`.
|
||||
## `scriptorium render`
|
||||
|
||||
### `scriptorium render`
|
||||
|
||||
```bash
|
||||
```text
|
||||
scriptorium render [flags]
|
||||
```
|
||||
|
||||
Required through flags or config:
|
||||
`--prompt <id>` and at least one `--input name=path` are required. The
|
||||
following optional flags are supported: `--config`, `--prompt-dir`,
|
||||
`--profile-dir`, `--profile`, `--var`, `--out`, `--llm-base-url`,
|
||||
`--model`, `--api-key-env`, `--temperature`, `--max-tokens`, `--top-p`,
|
||||
`--timeout`, and `--format text|json`. Their meanings match the corresponding
|
||||
`run` flags; `--format` selects prepared-run output and otherwise uses
|
||||
`defaults.render_format`.
|
||||
|
||||
- `--prompt-dir <dir>`: prompt definition directory.
|
||||
The same deprecated aliases and numeric/timeout behavior as `run` apply.
|
||||
`render` does not accept `--schema-dir`; configure `schema_dir` through the
|
||||
configuration file. It resolves profiles and schemas as part of preparation but
|
||||
does not call an LLM.
|
||||
|
||||
Required as flags:
|
||||
## `scriptorium serve`
|
||||
|
||||
- `--prompt <id>`: prompt ID to render.
|
||||
- `--input name=path`: input file mapping. Repeat or use comma-separated mappings.
|
||||
|
||||
Optional flags:
|
||||
|
||||
- `--config <path>`: application config file.
|
||||
- `--prompt-dir <dir>`: prompt definition directory.
|
||||
- `--profile-dir <dir>`: custom profile definition directory.
|
||||
- `--profile <id>`: execution profile override.
|
||||
- `--var name=value`: template variable mapping. Repeat or use comma-separated mappings.
|
||||
- `--out <path>`: write prepared-run output to a file instead of stdout.
|
||||
- `--llm-base-url <url>`: runtime endpoint override for the prepared request.
|
||||
- `--model <name>`: runtime model override for the prepared request.
|
||||
- `--api-key-env <name>`: runtime API-key environment variable name override.
|
||||
- `--temperature <float>`: runtime temperature override.
|
||||
- `--max-tokens <int>`: runtime max tokens override.
|
||||
- `--top-p <float>`: runtime top-p override.
|
||||
- `--timeout <duration>`: runtime timeout override using Go duration syntax.
|
||||
- `--format text|json`: prepared-run output format. Defaults to config `defaults.render_format`, then `text`.
|
||||
|
||||
Deprecated aliases:
|
||||
|
||||
- `--prompt-id <id>`: alias for `--prompt`.
|
||||
- `--profile-id <id>`: alias for `--profile`.
|
||||
|
||||
Notes:
|
||||
|
||||
- `render` resolves profiles, loads schemas for `json_schema` prompts, and validates `api_key_env`.
|
||||
- `render` does not accept `--schema-dir`; use config `schema_dir` for render-time schema lookup.
|
||||
- `render` does not call the LLM.
|
||||
|
||||
### `scriptorium serve`
|
||||
|
||||
```bash
|
||||
```text
|
||||
scriptorium serve [flags]
|
||||
```
|
||||
|
||||
Required through flags or config:
|
||||
|
||||
- `--prompt-dir <dir>`: prompt definition directory.
|
||||
|
||||
Optional flags:
|
||||
|
||||
- `--config <path>`: application config file.
|
||||
- `--addr <listen-address>`: HTTP listen address.
|
||||
- `--prompt-dir <dir>`: prompt definition directory.
|
||||
- `--profile-dir <dir>`: custom profile definition directory.
|
||||
- `--schema-dir <dir>`: schema base directory for `json_schema` validation.
|
||||
- `--artifact-root <dir>`: base directory for HTTP `file` input references.
|
||||
- `--max-request-bytes <n>`: maximum HTTP request body bytes; `0` disables the limit.
|
||||
- `--max-artifact-bytes <n>`: maximum HTTP file artifact bytes; `0` disables the limit.
|
||||
- `--max-response-bytes <n>`: maximum encoded HTTP response body bytes; `0` disables the limit.
|
||||
| Flag | Meaning |
|
||||
| --- | --- |
|
||||
| `--config <path>` | Application configuration file. |
|
||||
| `--addr <listen-address>` | HTTP listen-address override. |
|
||||
| `--prompt-dir <dir>` | Prompt-definition directory override. |
|
||||
| `--profile-dir <dir>` | Custom profile-directory override. |
|
||||
| `--schema-dir <dir>` | Schema base-directory override. |
|
||||
| `--artifact-root <dir>` | Root for HTTP `file` input references. |
|
||||
| `--max-request-bytes <n>` | Maximum encoded HTTP request-body bytes; `0` disables the limit. |
|
||||
| `--max-artifact-bytes <n>` | Maximum HTTP file-input artifact bytes; `0` disables the limit. |
|
||||
| `--max-response-bytes <n>` | Maximum encoded HTTP response bytes; `0` disables the limit. |
|
||||
|
||||
Notes:
|
||||
|
||||
- `serve` does not accept runtime model override flags such as `--model` or `--llm-base-url`.
|
||||
- HTTP request fields and error codes are documented in the [HTTP API reference](api.md).
|
||||
- HTTP `file` input references are rejected unless an artifact root is configured.
|
||||
- HTTP size-limit flags affect only `serve`.
|
||||
`serve` accepts no runtime model override flags. HTTP request fields, response
|
||||
schemas, and error codes are defined in the [HTTP API reference](api.md).
|
||||
|
||||
## Input And Variable Syntax
|
||||
|
||||
- `--input name=path` maps prompt input names to local file paths.
|
||||
- `--var name=value` maps prompt template variables to string values.
|
||||
- Both flags can be repeated.
|
||||
- Both flags also accept comma-separated mappings, such as `--input transcript=./t.md,glossary=./g.yml`.
|
||||
- Values may contain `=` after the first separator, such as `--var note=a=b=c`.
|
||||
- Empty names and empty values are rejected.
|
||||
`--input name=path` maps an input name to a local file; `--var name=value`
|
||||
maps a template variable to a string. Both flags can be repeated or contain
|
||||
comma-separated mappings. Values may contain `=` after the first separator.
|
||||
Empty names and values are rejected.
|
||||
|
||||
CLI `run` and `render` convert every `--input` mapping to a `file` artifact
|
||||
reference. HTTP also supports `inline` input references; see [HTTP API
|
||||
reference](api.md).
|
||||
CLI inputs are file references. HTTP inline inputs are defined by the
|
||||
[HTTP API reference](api.md).
|
||||
|
||||
## Output Behavior
|
||||
## Output And Exit Behavior
|
||||
|
||||
`run`:
|
||||
- `run` writes generated content to stdout, or to `--out` when supplied, and
|
||||
writes a concise summary to stderr.
|
||||
- `render` writes prepared-run output to stdout, or to `--out` when supplied,
|
||||
without a success summary.
|
||||
- `serve` writes startup and server errors to stderr.
|
||||
|
||||
- Writes generated artifact content to stdout by default.
|
||||
- Writes generated artifact content to `--out` when provided.
|
||||
- Prints a success summary to stderr.
|
||||
- Prints errors to stderr on failure.
|
||||
Exit statuses:
|
||||
|
||||
`render`:
|
||||
| Status | Meaning |
|
||||
| --- | --- |
|
||||
| `0` | Success. |
|
||||
| `1` | Parse, configuration, loading, rendering, generation, output-write, or other runtime error. |
|
||||
| `2` | `run` generated and wrote output, but validation failed. |
|
||||
|
||||
- Writes prepared-run output to stdout by default.
|
||||
- Writes prepared-run output to `--out` when provided.
|
||||
- Does not print a success summary.
|
||||
## Workflows And Examples
|
||||
|
||||
`serve`:
|
||||
|
||||
- Logs startup and server errors to stderr.
|
||||
|
||||
## Exit Codes
|
||||
|
||||
- `0`: success.
|
||||
- `1`: parse, config, load, render, generation, output-write, or runtime error.
|
||||
- `2`: `run` completed and wrote output, but validation status is `failed`.
|
||||
|
||||
## Common Workflows
|
||||
|
||||
Render prompt inputs and variables as JSON:
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium render \
|
||||
--config ./examples/config.yml \
|
||||
--prompt generic.markdown_summary \
|
||||
--input transcript=./examples/fixtures/transcript.md \
|
||||
--input glossary=./examples/fixtures/glossary.yml \
|
||||
--var session_date=2026-05-04 \
|
||||
--format json
|
||||
```
|
||||
|
||||
Run a prompt with an explicit profile and file output:
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium run \
|
||||
--config ./examples/config.yml \
|
||||
--prompt generic.markdown_summary \
|
||||
--profile local-fast \
|
||||
--input transcript=./examples/fixtures/transcript.md \
|
||||
--input glossary=./examples/fixtures/glossary.yml \
|
||||
--out ./summary.md
|
||||
```
|
||||
|
||||
Start the HTTP server with example config:
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium serve --config ./examples/config.yml
|
||||
```
|
||||
|
||||
Copyable maintained script:
|
||||
|
||||
- `examples/render-markdown-summary.sh`
|
||||
The [maintained render script](../examples/render-markdown-summary.sh) is a
|
||||
copyable render workflow. The [HTTP request example](../examples/http-run.json)
|
||||
is for a running `serve` process.
|
||||
|
||||
393
docs/config.md
393
docs/config.md
@@ -1,324 +1,169 @@
|
||||
# Configuration Reference
|
||||
|
||||
## Config Discovery And Precedence
|
||||
This is the canonical reference for Scriptorium application settings and the
|
||||
prompt, profile, and schema files those settings select. For command syntax,
|
||||
see the [CLI reference](cli.md); for HTTP request shapes, limits, and outcomes,
|
||||
see the [HTTP API reference](api.md).
|
||||
|
||||
## Discovery And Precedence
|
||||
|
||||
Application settings are resolved in this order:
|
||||
|
||||
1. built-in defaults
|
||||
2. `config.yml` values
|
||||
3. CLI overrides
|
||||
1. built-in defaults;
|
||||
2. a configuration file; then
|
||||
3. CLI overrides.
|
||||
|
||||
When `--config` is omitted, Scriptorium searches:
|
||||
When `--config` is omitted, Scriptorium searches
|
||||
`/usr/local/etc/scriptorium/config.yml` and then `/etc/scriptorium/config.yml`.
|
||||
If neither exists, it uses built-in defaults. An explicit `--config` path must
|
||||
exist and decode successfully.
|
||||
|
||||
1. `/usr/local/etc/scriptorium/config.yml`
|
||||
2. `/etc/scriptorium/config.yml`
|
||||
The maintained [minimal configuration](../examples/config.yml) and
|
||||
[full configuration](../examples/config.full.yml) are copyable examples.
|
||||
|
||||
If neither file exists, Scriptorium uses built-in defaults. When
|
||||
`--config <path>` is provided, that file must exist and decode successfully.
|
||||
## Application Configuration File
|
||||
|
||||
## Minimal Working Config
|
||||
Configuration is strict YAML: unknown fields are rejected. Empty string values
|
||||
do not override a prior value. Raw API-key fields are not accepted.
|
||||
|
||||
```yaml
|
||||
prompt_dir: ./examples/prompts
|
||||
```
|
||||
|
||||
This is enough for `run` and `render` when selected prompts use built-in
|
||||
profiles. Set `profile_dir` when prompts or requests use custom profiles.
|
||||
|
||||
The maintained repository example is `examples/config.yml`.
|
||||
|
||||
## Production-Oriented Config
|
||||
|
||||
```yaml
|
||||
prompt_dir: /opt/scriptorium/prompts
|
||||
profile_dir: /opt/scriptorium/profiles
|
||||
schema_dir: /opt/scriptorium/schemas
|
||||
|
||||
server:
|
||||
addr: 127.0.0.1:8080
|
||||
artifact_root: /var/lib/scriptorium/artifacts
|
||||
max_request_bytes: 16777216
|
||||
max_artifact_bytes: 16777216
|
||||
max_response_bytes: 16777216
|
||||
|
||||
defaults:
|
||||
render_format: text
|
||||
```
|
||||
|
||||
The maintained full example is `examples/config.full.yml`.
|
||||
|
||||
## App Config Reference
|
||||
|
||||
Top-level fields:
|
||||
|
||||
| Field | Default | Description |
|
||||
| Field | Default | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `prompt_dir` | unset | Directory containing prompt definition YAML files. Required effectively by `run`, `render`, and `serve`. |
|
||||
| `profile_dir` | unset | Directory containing custom profile YAML files. Built-in profiles remain available when unset. |
|
||||
| `prompt_dir` | unset | Directory containing prompt-definition YAML. `run`, `render`, and `serve` require an effective value. |
|
||||
| `profile_dir` | unset | Directory containing custom profile YAML. Built-in profiles remain available. |
|
||||
| `schema_dir` | `.` | Base directory for relative JSON Schema paths. |
|
||||
| `server` | `{}` | HTTP service settings used by `serve`. |
|
||||
| `defaults` | `{}` | Adapter defaults. |
|
||||
|
||||
`server` fields:
|
||||
|
||||
| Field | Default | Description |
|
||||
| --- | --- | --- |
|
||||
| `server.addr` | `:8080` | Listen address for `serve`. |
|
||||
| `server.artifact_root` | unset | Base directory for HTTP `file` input references. Without it, HTTP file refs are rejected. |
|
||||
| `server.max_request_bytes` | `16777216` | Maximum encoded HTTP request body bytes. `0` disables the limit. |
|
||||
| `server.max_artifact_bytes` | `16777216` | Maximum HTTP file artifact bytes. `0` disables the limit. |
|
||||
| `server.max_response_bytes` | `16777216` | Maximum encoded HTTP response bytes. `0` disables the limit. |
|
||||
|
||||
`defaults` fields:
|
||||
|
||||
| Field | Default | Description |
|
||||
| --- | --- | --- |
|
||||
| `server.addr` | `:8080` | Address used by `serve`. |
|
||||
| `server.artifact_root` | unset | Root that enables HTTP `file` input references. |
|
||||
| `server.max_request_bytes` | `16777216` | Maximum encoded HTTP request body bytes; `0` disables the limit. |
|
||||
| `server.max_artifact_bytes` | `16777216` | Maximum HTTP file-input artifact bytes; `0` disables the limit. |
|
||||
| `server.max_response_bytes` | `16777216` | Maximum encoded HTTP response bytes; `0` disables the limit. |
|
||||
| `defaults.render_format` | `text` | Default `render` output format: `text` or `json`. |
|
||||
|
||||
Config rules:
|
||||
|
||||
- YAML decoding is strict; unknown fields are rejected.
|
||||
- HTTP size limits must be greater than or equal to `0`.
|
||||
- Empty string config values are ignored.
|
||||
- Raw API key fields are not supported in app config.
|
||||
The three size fields must be zero or greater. The HTTP contract defines how
|
||||
each limit is enforced and reported. `server.artifact_root` configures the
|
||||
deployment boundary; see the [HTTP API reference](api.md) for request-path and
|
||||
containment behavior, and [operations](operations.md) for deployment handling.
|
||||
|
||||
## Prompt Definition Files
|
||||
|
||||
Prompt definitions are YAML files anywhere under `prompt_dir`. Nested
|
||||
directories are organizational; callers select prompts by YAML `id`, not file
|
||||
path.
|
||||
Prompt definitions are strict YAML files anywhere below `prompt_dir`. A prompt
|
||||
is selected by its YAML `id`, not by file path; nested directories are only for
|
||||
organization. See [maintained prompt examples](../examples/prompts/).
|
||||
|
||||
Example:
|
||||
|
||||
```yaml
|
||||
id: generic.structured_events
|
||||
version: "1.0.0"
|
||||
default_profile: local-quality
|
||||
description: Produce structured event JSON from a transcript.
|
||||
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: text/markdown
|
||||
description: Source transcript content
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/yaml
|
||||
description: Optional glossary context
|
||||
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./generic.structured_events.system.md
|
||||
- role: user
|
||||
content_file: ./generic.structured_events.user.md
|
||||
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: structured_events.schema.json
|
||||
repair_attempts: 0
|
||||
```
|
||||
|
||||
Prompt fields:
|
||||
|
||||
| Field | Required | Description |
|
||||
| Field | Required | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `id` | yes | Prompt identifier used by `--prompt` and HTTP `prompt_id`. |
|
||||
| `id` | yes | Prompt identifier. |
|
||||
| `version` | yes | Prompt version. |
|
||||
| `default_profile` | no | Profile ID used when a request does not provide a profile. |
|
||||
| `default_profile` | no | Profile used when a request omits a profile ID. |
|
||||
| `description` | no | Human-readable description. |
|
||||
| `session_id` | no | Go-template string rendered from request vars and forwarded as provider `session_id` when non-empty. |
|
||||
| `inputs` | no | Named input declarations. |
|
||||
| `messages` | yes | Chat message templates. |
|
||||
| `session_id` | no | Go-template string rendered from request variables and sent to a compatible provider when non-empty. |
|
||||
| `inputs` | no | Declared input metadata. |
|
||||
| `messages` | yes | Chat-message templates. |
|
||||
| `output` | yes | Output format and validation contract. |
|
||||
|
||||
`inputs[]` fields:
|
||||
### Inputs And Messages
|
||||
|
||||
- `name` (required)
|
||||
- `required` (optional boolean)
|
||||
- `content_type` (optional metadata)
|
||||
- `description` (optional)
|
||||
Each `inputs` item has a required `name` and optional `required`,
|
||||
`content_type`, and `description` fields. Input names must be unique.
|
||||
|
||||
`messages[]` fields:
|
||||
Each message has a required `role`, exactly one of `content` or `content_file`,
|
||||
and optional `cache_control`. A `content_file` path is relative to the prompt
|
||||
file. `cache_control.type` must be `ephemeral`; its optional `ttl` is `1h`.
|
||||
|
||||
- `role` (required)
|
||||
- exactly one of `content` or `content_file`
|
||||
- `cache_control` (optional)
|
||||
`session_id` uses the same template variables as messages. Empty rendered
|
||||
values are omitted. A rendered value may contain at most 256 Unicode code
|
||||
points.
|
||||
|
||||
Message rules:
|
||||
### Output Contract
|
||||
|
||||
- `content_file` resolves relative to the prompt YAML file location.
|
||||
- Repeated roles are allowed.
|
||||
- Prompt YAML decoding is strict.
|
||||
- Duplicate input names are invalid.
|
||||
- Duplicate prompt IDs are invalid for a requested ID/version.
|
||||
|
||||
`messages[].cache_control` fields:
|
||||
|
||||
| Field | Required | Supported values |
|
||||
| Field | Required | Values or behavior |
|
||||
| --- | --- | --- |
|
||||
| `type` | yes | `ephemeral` |
|
||||
| `ttl` | no | `1h` |
|
||||
|
||||
`session_id` behavior:
|
||||
|
||||
- Rendered with the same variable context as message templates.
|
||||
- Trimmed and omitted when empty.
|
||||
- Rejected when longer than 256 Unicode code points.
|
||||
- CLI callers pass variables with `--var`; HTTP callers use `vars`.
|
||||
|
||||
`output` fields:
|
||||
|
||||
| Field | Required | Supported values |
|
||||
| --- | --- | --- |
|
||||
| `format` | yes | `text`, `markdown`, `json` |
|
||||
| `validation_mode` | yes | `none`, `basic`, `json`, `json_schema` |
|
||||
| `schema_path` | only for `json_schema` | Relative to `schema_dir` unless absolute. |
|
||||
| `repair_attempts` | yes | Integer greater than or equal to `0`. |
|
||||
|
||||
Repair boundary:
|
||||
|
||||
- `repair_attempts` is part of the prompt contract.
|
||||
- The current CLI and HTTP wiring constructs the runner without a repairer, so normal `run` and `serve` execution does not perform repair attempts.
|
||||
| `format` | yes | `text`, `markdown`, or `json`. |
|
||||
| `validation_mode` | yes | `none`, `basic`, `json`, or `json_schema`. |
|
||||
| `schema_path` | for `json_schema` | Schema path, relative to `schema_dir` unless absolute. |
|
||||
| `repair_attempts` | no | Integer greater than or equal to `0`; omitted means `0`. |
|
||||
|
||||
## Profile Definition Files
|
||||
|
||||
Execution profiles are YAML files anywhere under `profile_dir`. Nested
|
||||
directories are organizational; callers select profiles by YAML `id`, not file
|
||||
path.
|
||||
Profiles are strict YAML files anywhere below `profile_dir`. A profile is
|
||||
selected by YAML `id`; nested directories are organizational. See the
|
||||
[maintained profile examples](../examples/profiles/).
|
||||
|
||||
Scriptorium also ships built-in profiles. Custom profiles override built-ins
|
||||
with the same ID.
|
||||
|
||||
Example:
|
||||
|
||||
```yaml
|
||||
id: local-fast
|
||||
endpoint: http://localhost:8000/v1
|
||||
model: gpt-4o-mini
|
||||
temperature: 0.2
|
||||
max_tokens: 500
|
||||
top_p: 1.0
|
||||
timeout_seconds: 90
|
||||
api_key_env: SCRIPTORIUM_API_KEY
|
||||
service_tier: priority
|
||||
reasoning_effort: medium
|
||||
extra_params:
|
||||
provider_route: primary
|
||||
```
|
||||
|
||||
Profile fields:
|
||||
|
||||
| Field | Required | Description |
|
||||
| Field | Required | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `id` | yes | Profile identifier. |
|
||||
| `endpoint` | yes | OpenAI-compatible base URL including `/v1`. |
|
||||
| `endpoint` | yes | OpenAI-compatible base URL, including its API version path when needed. |
|
||||
| `model` | yes | Provider model name. |
|
||||
| `temperature` | no | Range `0..2`. |
|
||||
| `max_tokens` | no | Integer greater than or equal to `0`. |
|
||||
| `top_p` | no | Range `0..1`. |
|
||||
| `timeout_seconds` | no | Integer greater than or equal to `0`. |
|
||||
| `service_tier` | no | Provider-specific request tier. |
|
||||
| `reasoning_effort` | no | Provider-specific reasoning setting. |
|
||||
| `api_key_env` | no | Environment variable name containing the API key. |
|
||||
| `extra_params` | no | JSON-compatible provider-specific top-level request fields. |
|
||||
| `temperature` | no | Number from `0` through `2`. |
|
||||
| `max_tokens` | no | Integer zero or greater. |
|
||||
| `top_p` | no | Number from `0` through `1`. |
|
||||
| `timeout_seconds` | no | Integer zero or greater. |
|
||||
| `service_tier` | no | Non-empty provider-specific request tier. |
|
||||
| `reasoning_effort` | no | Non-empty provider-specific reasoning setting. |
|
||||
| `api_key_env` | no | Environment-variable name containing the API key. |
|
||||
| `extra_params` | no | JSON-compatible provider-specific outbound request fields. |
|
||||
|
||||
Execution defaults before profile/request overrides:
|
||||
Execution defaults before profile and request overrides are `temperature: 0`,
|
||||
`max_tokens: 0`, `top_p: 1`, and `timeout_seconds: 600`. Profile numeric values
|
||||
merge by non-zero value. Request overrides preserve presence, so an explicit
|
||||
zero can override a profile value.
|
||||
|
||||
| Field | Default |
|
||||
| --- | --- |
|
||||
| `temperature` | `0.0` |
|
||||
| `max_tokens` | `0` |
|
||||
| `top_p` | `1.0` |
|
||||
| `timeout_seconds` | `600` |
|
||||
Custom profiles take precedence over built-ins with the same ID. Invalid custom
|
||||
profiles are errors; they do not fall back to a built-in profile. Raw `api_key`
|
||||
is rejected. Use `api_key_env`, or the public Go package's request-scoped key
|
||||
mechanism described in the [package contract](consumers/pkg-scriptorium.md).
|
||||
|
||||
Profile rules:
|
||||
`extra_params` keys must be non-empty and cannot be `model`, `session_id`,
|
||||
`messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`,
|
||||
`reasoning_effort`, or `response_format`.
|
||||
|
||||
- Profile YAML decoding is strict.
|
||||
- Duplicate custom profile IDs are invalid.
|
||||
- Matching custom and built-in IDs are valid override behavior.
|
||||
- Raw `api_key` is rejected; use `api_key_env`.
|
||||
- If `api_key_env` is set, the named environment variable must be set before `run`, `render`, or HTTP execution can prepare the request.
|
||||
- Profile numeric fields merge by non-zero value. Request overrides are presence-aware, so explicit zero values are supported through CLI flags or HTTP model overrides.
|
||||
- `extra_params` keys must not be empty and must not collide with reserved outbound fields: `model`, `session_id`, `messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or `response_format`.
|
||||
### Built-In Profile Catalog
|
||||
|
||||
Built-in profile catalog:
|
||||
Each embedded profile uses `OPENROUTER_API_KEY`.
|
||||
|
||||
| Provider | ID | Model | API key env |
|
||||
| --- | --- | --- | --- |
|
||||
| aion-labs | `aion-2` | `aion-labs/aion-2.0` | `OPENROUTER_API_KEY` |
|
||||
| anthropic | `claude-fable-latest` | `~anthropic/claude-fable-latest` | `OPENROUTER_API_KEY` |
|
||||
| anthropic | `claude-haiku-latest` | `~anthropic/claude-haiku-latest` | `OPENROUTER_API_KEY` |
|
||||
| anthropic | `claude-opus-latest` | `~anthropic/claude-opus-latest` | `OPENROUTER_API_KEY` |
|
||||
| anthropic | `claude-sonnet-latest` | `~anthropic/claude-sonnet-latest` | `OPENROUTER_API_KEY` |
|
||||
| deepseek | `deepseek-3-2` | `deepseek/deepseek-v3.2` | `OPENROUTER_API_KEY` |
|
||||
| deepseek | `deepseek-4-pro` | `deepseek/deepseek-v4-pro` | `OPENROUTER_API_KEY` |
|
||||
| google | `gemini-2-flash` | `google/gemini-2.5-flash` | `OPENROUTER_API_KEY` |
|
||||
| google | `gemini-2-flash-lite` | `google/gemini-2.5-flash-lite` | `OPENROUTER_API_KEY` |
|
||||
| google | `gemini-2-pro` | `google/gemini-2.5-pro` | `OPENROUTER_API_KEY` |
|
||||
| google | `gemini-3-flash-lite` | `google/gemini-3.1-flash-lite` | `OPENROUTER_API_KEY` |
|
||||
| google | `gemini-flash-latest` | `~google/gemini-flash-latest` | `OPENROUTER_API_KEY` |
|
||||
| google | `gemini-pro-latest` | `~google/gemini-pro-latest` | `OPENROUTER_API_KEY` |
|
||||
| google | `gemma-4-31b` | `google/gemma-4-31b-it:exacto` | `OPENROUTER_API_KEY` |
|
||||
| minimax | `minimax-m2` | `minimax/minimax-m2.5` | `OPENROUTER_API_KEY` |
|
||||
| minimax | `minimax-m3` | `minimax/minimax-m3` | `OPENROUTER_API_KEY` |
|
||||
| mistral | `mistral-large-2512` | `mistralai/mistral-large-2512` | `OPENROUTER_API_KEY` |
|
||||
| mistral | `mistral-medium-3-5` | `mistralai/mistral-medium-3-5` | `OPENROUTER_API_KEY` |
|
||||
| mistral | `mistral-small-3` | `mistralai/mistral-small-3.2-24b-instruct` | `OPENROUTER_API_KEY` |
|
||||
| mistral | `mistral-small-4` | `mistralai/mistral-small-2603` | `OPENROUTER_API_KEY` |
|
||||
| nvidia | `nemotron-3-ultra` | `nvidia/nemotron-3-ultra-550b-a55b` | `OPENROUTER_API_KEY` |
|
||||
| openai | `gpt-5-mini` | `openai/gpt-5.4-mini` | `OPENROUTER_API_KEY` |
|
||||
| openai | `gpt-5-nano` | `openai/gpt-5.4-nano` | `OPENROUTER_API_KEY` |
|
||||
| Provider | ID | Model |
|
||||
| --- | --- | --- |
|
||||
| aion-labs | `aion-2` | `aion-labs/aion-2.0` |
|
||||
| anthropic | `claude-fable-latest` | `~anthropic/claude-fable-latest` |
|
||||
| anthropic | `claude-haiku-latest` | `~anthropic/claude-haiku-latest` |
|
||||
| anthropic | `claude-opus-latest` | `~anthropic/claude-opus-latest` |
|
||||
| anthropic | `claude-sonnet-latest` | `~anthropic/claude-sonnet-latest` |
|
||||
| deepseek | `deepseek-3-2` | `deepseek/deepseek-v3.2` |
|
||||
| deepseek | `deepseek-4-flash` | `deepseek/deepseek-v4-flash` |
|
||||
| deepseek | `deepseek-4-pro` | `deepseek/deepseek-v4-pro` |
|
||||
| google | `gemini-2-flash` | `google/gemini-2.5-flash` |
|
||||
| google | `gemini-2-flash-lite` | `google/gemini-2.5-flash-lite` |
|
||||
| google | `gemini-2-pro` | `google/gemini-2.5-pro` |
|
||||
| google | `gemini-3-flash-lite` | `google/gemini-3.1-flash-lite` |
|
||||
| google | `gemini-flash-latest` | `~google/gemini-flash-latest` |
|
||||
| google | `gemini-pro-latest` | `~google/gemini-pro-latest` |
|
||||
| google | `gemma-4-31b` | `google/gemma-4-31b-it:exacto` |
|
||||
| minimax | `minimax-m2` | `minimax/minimax-m2.5` |
|
||||
| minimax | `minimax-m3` | `minimax/minimax-m3` |
|
||||
| mistral | `mistral-large-2512` | `mistralai/mistral-large-2512` |
|
||||
| mistral | `mistral-medium-3-5` | `mistralai/mistral-medium-3-5` |
|
||||
| mistral | `mistral-small-3` | `mistralai/mistral-small-3.2-24b-instruct` |
|
||||
| mistral | `mistral-small-4` | `mistralai/mistral-small-2603` |
|
||||
| nvidia | `nemotron-3-ultra` | `nvidia/nemotron-3-ultra-550b-a55b` |
|
||||
| openai | `gpt-5-mini` | `openai/gpt-5.4-mini` |
|
||||
| openai | `gpt-5-nano` | `openai/gpt-5.4-nano` |
|
||||
|
||||
## Schema Behavior
|
||||
## Schemas
|
||||
|
||||
Schemas are JSON files, typically under `schema_dir`.
|
||||
Schemas are JSON files, normally below `schema_dir`. `json_schema` output
|
||||
requires a `schema_path`. Relative paths resolve from `schema_dir`; absolute
|
||||
paths are used directly. Referenced nested schemas use relative paths and are
|
||||
not discovered by basename. An unreadable or invalid schema is a runtime
|
||||
validation error; generated content that fails JSON or schema validation is a
|
||||
validation result.
|
||||
|
||||
Rules:
|
||||
## Credentials
|
||||
|
||||
- `output.validation_mode: json_schema` requires `output.schema_path`.
|
||||
- Relative `schema_path` values resolve from `schema_dir`.
|
||||
- Absolute `schema_path` values are used directly.
|
||||
- Nested schemas must be referenced by relative path; schemas are not searched recursively by basename.
|
||||
- Missing or invalid schema documents are runtime validation errors.
|
||||
- Invalid generated JSON produces validation status `failed`, not a runtime error.
|
||||
Keep secrets in environment variables. Store only an environment-variable name
|
||||
in `api_key_env`; do not place raw keys in configuration, prompt or profile
|
||||
files, CLI arguments, examples, or HTTP payloads.
|
||||
|
||||
## Artifact References
|
||||
|
||||
Supported request input artifact reference types are:
|
||||
|
||||
- `file`
|
||||
- `inline`
|
||||
|
||||
CLI `run` and `render` create `file` references from `--input name=path`.
|
||||
|
||||
HTTP `file` references require `server.artifact_root` or `serve
|
||||
--artifact-root`. Relative file URIs resolve under that root. Absolute paths
|
||||
and relative traversal outside the root are rejected by lexical checks. Symlinks
|
||||
inside the root are followed by the operating system, including symlinks that
|
||||
point outside the root.
|
||||
|
||||
HTTP `inline` references do not require an artifact root.
|
||||
|
||||
## Secrets Handling
|
||||
|
||||
- Keep secret values in environment variables.
|
||||
- Store only environment-variable names in `api_key_env`.
|
||||
- Do not put raw API keys in config, prompts, profiles, CLI arguments, examples, or HTTP request bodies.
|
||||
|
||||
## Maintained Examples
|
||||
|
||||
- Minimal app config: `examples/config.yml`
|
||||
- Full app config: `examples/config.full.yml`
|
||||
- Prompt examples: `examples/prompts/`
|
||||
- Custom profile examples: `examples/profiles/`
|
||||
- Schema examples: `examples/schemas/`
|
||||
- Input fixtures: `examples/fixtures/`
|
||||
- Render script: `examples/render-markdown-summary.sh`
|
||||
- HTTP request-shape example: `examples/http-run.json`
|
||||
|
||||
## Integration References
|
||||
## Related References
|
||||
|
||||
- [CLI reference](cli.md)
|
||||
- [HTTP API reference](api.md)
|
||||
- [Outbound OpenAI-compatible contract](integrations/openai-compatible-chat.md)
|
||||
- [OpenAI-compatible outbound contract](integrations/openai-compatible-chat.md)
|
||||
|
||||
@@ -1,65 +1,26 @@
|
||||
# Consumer Integration Overview
|
||||
|
||||
This guide is for applications that call Scriptorium from another codebase.
|
||||
This guide helps applications choose a Scriptorium interface and understand
|
||||
their responsibilities. The linked contracts own interface syntax and wire
|
||||
semantics.
|
||||
|
||||
Scriptorium exposes three integration surfaces:
|
||||
|
||||
| Surface | Use when |
|
||||
| Interface | Use when |
|
||||
| --- | --- |
|
||||
| Go package | The consumer is Go, needs typed requests/results, or wants injected LLM clients for tests. |
|
||||
| CLI subprocess | The consumer wants process isolation or is not written in Go. |
|
||||
| HTTP API | The consumer needs a service boundary or remote access to `POST /v1/runs`. |
|
||||
| Go package | The consumer is Go and needs typed requests, results, or an injected LLM client. |
|
||||
| CLI subprocess | The consumer needs process isolation or is not written in Go. |
|
||||
| HTTP API | The consumer needs a service boundary or remote access. |
|
||||
|
||||
Canonical references:
|
||||
- Go package: [package contract](pkg-scriptorium.md)
|
||||
- CLI subprocess: [subprocess integration](../integrations/subprocess.md)
|
||||
- HTTP service: [HTTP API reference](../api.md)
|
||||
- Prompt, profile, schema, and credential configuration: [configuration reference](../config.md)
|
||||
|
||||
- Go package: [Package scriptorium](pkg-scriptorium.md)
|
||||
- CLI subprocess: [Subprocess integration](../integrations/subprocess.md)
|
||||
- HTTP: [HTTP API reference](../api.md)
|
||||
- File formats: [Configuration reference](../config.md)
|
||||
|
||||
## Required Deployment Inputs
|
||||
|
||||
Every integration needs operators to provide:
|
||||
|
||||
- prompt definitions;
|
||||
- profile definitions or built-in profile IDs;
|
||||
- schema files when prompts use `json_schema`;
|
||||
- input artifacts or inline input bodies;
|
||||
- API-key environment variables or direct per-request keys where supported.
|
||||
|
||||
Raw API keys do not belong in config, prompt files, profile YAML, CLI
|
||||
arguments, or HTTP request bodies.
|
||||
|
||||
## Recommended Workflow
|
||||
|
||||
Use the Go package when:
|
||||
|
||||
- the consumer is a Go application;
|
||||
- the application needs `context.Context` cancellation;
|
||||
- repeated calls should avoid subprocess startup;
|
||||
- tests need a fake LLM client;
|
||||
- direct per-request `RunRequest.APIKey` is required.
|
||||
|
||||
Use the CLI subprocess when:
|
||||
|
||||
- the consumer is not Go;
|
||||
- process isolation is useful;
|
||||
- stdout/stderr separation and exit codes are enough;
|
||||
- the consumer already manages local files and environment variables.
|
||||
|
||||
Use HTTP when:
|
||||
|
||||
- Scriptorium should run as a service;
|
||||
- multiple clients need a shared prompt/profile deployment;
|
||||
- clients can reach a trusted, protected HTTP boundary.
|
||||
|
||||
## Minimal Go Example
|
||||
## Minimal Go Use
|
||||
|
||||
```go
|
||||
engine, err := scriptorium.NewEngine(scriptorium.Config{
|
||||
PromptDir: "./examples/prompts",
|
||||
ProfileDir: "./examples/profiles",
|
||||
SchemaDir: "./examples/schemas",
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -69,54 +30,30 @@ prepared, err := engine.Prepare(ctx, scriptorium.RunRequest{
|
||||
PromptID: "generic.markdown_summary",
|
||||
Inputs: map[string]scriptorium.ArtifactRef{
|
||||
"transcript": scriptorium.File("./examples/fixtures/transcript.md"),
|
||||
"glossary": scriptorium.File("./examples/fixtures/glossary.yml"),
|
||||
},
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
_ = prepared.Messages
|
||||
_ = prepared
|
||||
```
|
||||
|
||||
Run the maintained package example:
|
||||
|
||||
```bash
|
||||
go run ./examples/go-library/prepare
|
||||
```
|
||||
|
||||
## Subprocess Workflow
|
||||
|
||||
Invoke `scriptorium render` for preflight and `scriptorium run` for generation.
|
||||
Capture stdout and stderr separately. Treat exit code `2` from `run` as a
|
||||
completed generation with failed validation.
|
||||
|
||||
See [Subprocess integration](../integrations/subprocess.md) for the stable
|
||||
invocation contract.
|
||||
|
||||
## HTTP Workflow
|
||||
|
||||
Run `scriptorium serve` behind trusted controls and send JSON requests to
|
||||
`POST /v1/runs`.
|
||||
|
||||
Do not duplicate endpoint schemas in consumers. Use the [HTTP API
|
||||
reference](../api.md) as the authoritative contract.
|
||||
For a maintained program, see
|
||||
[`examples/go-library/prepare`](../../examples/go-library/prepare).
|
||||
|
||||
## Consumer Responsibilities
|
||||
|
||||
Consumers are responsible for:
|
||||
|
||||
- selecting prompt/profile IDs as deployment configuration;
|
||||
- supplying all required inputs and vars;
|
||||
- protecting generated artifacts and rendered prompts as sensitive data;
|
||||
- deciding whether to keep output when validation fails;
|
||||
- implementing retries only when another model call is acceptable.
|
||||
- selecting and deploying prompt, profile, and schema assets;
|
||||
- supplying required inputs and template variables;
|
||||
- supplying credentials through the applicable interface;
|
||||
- protecting rendered prompts and generated artifacts as potentially sensitive;
|
||||
- deciding whether validation-failed output is usable; and
|
||||
- retrying only when another model call is acceptable.
|
||||
|
||||
Scriptorium does not persist run state. Retrying a failed or timed-out request
|
||||
can produce different output and can incur another provider request.
|
||||
|
||||
## Status Behavior
|
||||
|
||||
- Go package methods return typed results or errors that support `errors.Is`.
|
||||
- CLI `run` exits `2` when generation succeeds but validation fails.
|
||||
- HTTP returns `200 OK` for generated-content validation failures and exposes the failed status in the response body.
|
||||
- Runtime validation failures are errors.
|
||||
Scriptorium does not persist run state. A retry can produce different output and
|
||||
can incur another provider request. CLI exit behavior belongs to the
|
||||
[CLI reference](../cli.md); HTTP status behavior belongs to the
|
||||
[HTTP API reference](../api.md); package errors and results belong to the
|
||||
[package contract](pkg-scriptorium.md).
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Package scriptorium
|
||||
# Package `scriptorium`
|
||||
|
||||
Import path:
|
||||
|
||||
@@ -6,250 +6,158 @@ Import path:
|
||||
import "gitea.maximumdirect.net/eric/scriptorium"
|
||||
```
|
||||
|
||||
The root package is the public Go facade for Scriptorium's prompt prepare/run
|
||||
workflow. It exposes typed requests, results, source options, injected LLM
|
||||
clients, and stable public errors while keeping `internal/*` packages private.
|
||||
This is the canonical public Go contract for in-process prompt preparation and
|
||||
execution. Prompt, profile, and schema file formats are defined in the
|
||||
[configuration reference](../config.md).
|
||||
|
||||
## Intended Use Cases
|
||||
## Engine Construction
|
||||
|
||||
Use the package when a Go application needs:
|
||||
`NewEngine(Config, ...Option)` constructs an engine. `Config` has these
|
||||
fields:
|
||||
|
||||
- in-process prompt preparation or execution;
|
||||
- typed request/result structs;
|
||||
- direct `context.Context` cancellation;
|
||||
- injected/fake LLM clients for tests;
|
||||
- direct per-request `RunRequest.APIKey`.
|
||||
| Field | Meaning |
|
||||
| --- | --- |
|
||||
| `PromptDir` | Prompt-definition directory, required unless a prompt source option is supplied. |
|
||||
| `ProfileDir` | Optional custom profile directory over built-ins. |
|
||||
| `SchemaDir` | Schema directory; empty uses `.`. |
|
||||
| `Timeout` | Default timeout for the built-in OpenAI-compatible client. |
|
||||
| `HTTPClient` | Optional HTTP client for that built-in client. |
|
||||
|
||||
Use [Subprocess integration](../integrations/subprocess.md) or the [HTTP API](../api.md)
|
||||
when a process or service boundary is preferred.
|
||||
Nil options are ignored. Invalid construction, including
|
||||
`WithLLMClient(nil)`, returns an error matching `ErrInvalidConfig`.
|
||||
|
||||
## Construct An Engine
|
||||
Source options replace their matching directory source:
|
||||
|
||||
- prompts: `WithPromptFS(fsys, root)`, `WithPromptFile(path)`;
|
||||
- profiles: `WithProfileFS(fsys, root)`, `WithProfileFile(path)`, and
|
||||
`WithProfiles(profiles...)`;
|
||||
- schemas: `WithSchemaFS(fsys, root)`, `WithSchemaFile(path)`; and
|
||||
- LLM client: `WithLLMClient(client)`.
|
||||
|
||||
`fs.FS` prompt-content and schema paths stay inside their configured roots.
|
||||
A single-file option exposes that file by its base name. In-memory profiles take
|
||||
precedence over an explicit or directory-backed profile source, which in turn
|
||||
takes precedence over built-ins. File and filesystem sources use the format and
|
||||
credential rules in the [configuration reference](../config.md).
|
||||
|
||||
## Prepare And Run
|
||||
|
||||
`Prepare(ctx, request)` resolves the prompt, profile, input artifacts,
|
||||
validation contract, and rendered messages without calling an LLM.
|
||||
`Run(ctx, request)` performs that preparation, calls the configured client,
|
||||
and validates generated content.
|
||||
|
||||
```go
|
||||
engine, err := scriptorium.NewEngine(scriptorium.Config{
|
||||
PromptDir: "./examples/prompts",
|
||||
ProfileDir: "./examples/profiles",
|
||||
SchemaDir: "./examples/schemas",
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
```
|
||||
|
||||
`Config` fields:
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `PromptDir` | Prompt definition directory. Required unless `WithPromptFS` or `WithPromptFile` is used. |
|
||||
| `ProfileDir` | Optional custom profile directory overlaid above built-in profiles. |
|
||||
| `SchemaDir` | Schema directory. Defaults to `.` when empty. |
|
||||
| `Timeout` | Default timeout for the built-in OpenAI-compatible client. |
|
||||
| `HTTPClient` | Optional HTTP client for the built-in OpenAI-compatible client. |
|
||||
|
||||
`NewEngine` accepts `nil` options and ignores them. Invalid construction wraps
|
||||
`ErrInvalidConfig`.
|
||||
|
||||
## Source Options
|
||||
|
||||
Directory fields are the compatibility path. Explicit source options override
|
||||
the matching directory field.
|
||||
|
||||
Prompt sources:
|
||||
|
||||
- `WithPromptFS(fsys, root)`
|
||||
- `WithPromptFile(path)`
|
||||
|
||||
Profile sources:
|
||||
|
||||
- `WithProfileFS(fsys, root)`
|
||||
- `WithProfileFile(path)`
|
||||
- `WithProfiles(profiles...)`
|
||||
|
||||
Schema sources:
|
||||
|
||||
- `WithSchemaFS(fsys, root)`
|
||||
- `WithSchemaFile(path)`
|
||||
|
||||
LLM source:
|
||||
|
||||
- `WithLLMClient(client)`
|
||||
|
||||
Source behavior:
|
||||
|
||||
- Prompt and profile YAML use the same strict rules as directory loading.
|
||||
- Prompt `content_file` values resolve relative to the prompt file.
|
||||
- `fs.FS` roots are containment boundaries for prompt content files and schema paths.
|
||||
- File options expose the selected file by its base name.
|
||||
- Profile source precedence is in-memory profiles, then explicit profile file/FS/directory source, then built-ins.
|
||||
- `WithLLMClient(nil)` returns `ErrInvalidConfig`.
|
||||
|
||||
## In-Memory Profiles
|
||||
|
||||
Use `WithProfiles` when the application already has typed model settings:
|
||||
|
||||
```go
|
||||
profile := scriptorium.OpenAICompatibleProfile(scriptorium.OpenAICompatibleProfileConfig{
|
||||
ID: "app.default",
|
||||
Endpoint: "https://openrouter.ai/api/v1",
|
||||
Model: "mistralai/mistral-small-3.2-24b-instruct",
|
||||
APIKeyRequired: true,
|
||||
})
|
||||
|
||||
engine, err := scriptorium.NewEngine(cfg, scriptorium.WithProfiles(profile))
|
||||
```
|
||||
|
||||
`Profile` and `OpenAICompatibleProfileConfig` include:
|
||||
|
||||
- `ID`
|
||||
- `Endpoint`
|
||||
- `Model`
|
||||
- `Temperature`
|
||||
- `MaxTokens`
|
||||
- `TopP`
|
||||
- `TimeoutSeconds`
|
||||
- `ServiceTier`
|
||||
- `ReasoningEffort`
|
||||
- `APIKeyRequired`
|
||||
- `ExtraParams`
|
||||
|
||||
`WithProfiles` rejects duplicate IDs in one call. In-memory profiles do not
|
||||
store raw keys. When `APIKeyRequired` is true, pass the secret on each request
|
||||
with `RunRequest.APIKey`.
|
||||
|
||||
`ExtraParams` must be JSON-compatible: strings, booleans, finite numbers,
|
||||
objects with string keys, arrays/slices, and nil. Unsupported values, non-string
|
||||
map keys, non-finite floats, and cycles return `ErrInvalidConfig` for profiles
|
||||
or `ErrInvalidRequest` for request overrides.
|
||||
|
||||
## Prepare Workflow
|
||||
|
||||
`Prepare` resolves prompt/profile/input/schema state and renders messages
|
||||
without calling an LLM.
|
||||
|
||||
```go
|
||||
prepared, err := engine.Prepare(ctx, scriptorium.RunRequest{
|
||||
PromptID: "generic.markdown_summary",
|
||||
Inputs: map[string]scriptorium.ArtifactRef{
|
||||
"transcript": scriptorium.File("./examples/fixtures/transcript.md"),
|
||||
"glossary": scriptorium.File("./examples/fixtures/glossary.yml"),
|
||||
},
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
_ = prepared.EffectiveModelParams
|
||||
_ = prepared.Messages
|
||||
```
|
||||
|
||||
`PreparedRun` includes prompt ID/version/hash, selected profile, effective
|
||||
model params, output contract, structured-output metadata, input hashes,
|
||||
rendered prompt hash, rendered messages, and timing fields. It does not include
|
||||
raw API-key values, model output, validation results, or internal target
|
||||
presence metadata.
|
||||
The maintained package example is
|
||||
[`examples/go-library/prepare`](../../examples/go-library/prepare).
|
||||
|
||||
## Run Workflow
|
||||
`PreparedRun` exposes prompt, selected-profile, effective-model, output
|
||||
contract, structured-output, input-hash, rendered-message, and timing
|
||||
information. It does not include a resolved API key, model output, validation
|
||||
result, or target-presence metadata.
|
||||
|
||||
`Run` calls `Prepare`, invokes the configured LLM client, builds the output
|
||||
artifact, and validates the output.
|
||||
`RunResult` adds run ID, artifact, raw output, validation, model metadata,
|
||||
usage, and duration. Generated-content validation failures return a result with
|
||||
`Validation.Status == ValidationFailed`; schema or validator runtime failures
|
||||
return an error matching `ErrValidation`.
|
||||
|
||||
```go
|
||||
result, err := engine.Run(ctx, scriptorium.RunRequest{
|
||||
PromptID: "generic.markdown_summary",
|
||||
APIKey: apiKey,
|
||||
Inputs: map[string]scriptorium.ArtifactRef{
|
||||
"transcript": scriptorium.File("./examples/fixtures/transcript.md"),
|
||||
"glossary": scriptorium.File("./examples/fixtures/glossary.yml"),
|
||||
},
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
_ = result.Artifact
|
||||
```
|
||||
## Public Values
|
||||
|
||||
`RunResult` includes run ID, output artifact, raw output, validation result,
|
||||
prompt/profile/model metadata, effective model params, input hashes, usage, and
|
||||
timing fields.
|
||||
`ArtifactRef` has `Type`, `URI`, and `Body`; `Artifact` has `Name`,
|
||||
`ContentType`, `Body`, `URI`, `Size`, and `Hash`. `ExecutionTarget` exposes the
|
||||
effective endpoint, model, numeric settings, credential-environment name,
|
||||
service tier, reasoning effort, and extra parameters. `ValidationResult`
|
||||
contains status, mode, errors, schema path, repair attempts, and validity.
|
||||
|
||||
Generated-content validation failures return a successful `RunResult` with
|
||||
`Validation.Status == ValidationFailed`. Runtime/schema validation errors
|
||||
return an error that matches `ErrValidation`.
|
||||
The exported constants define these serialized values:
|
||||
|
||||
## Inputs
|
||||
- artifact types: `inline` and `file`;
|
||||
- output formats: `text`, `markdown`, and `json`;
|
||||
- validation modes: `none`, `basic`, `json`, and `json_schema`; and
|
||||
- validation statuses: `passed`, `failed`, and `skipped`.
|
||||
|
||||
Input helpers:
|
||||
`TokenUsage` reports prompt, completion, total, cached, and cache-write token
|
||||
counts. `RenderedPrompt`, `RenderedMessage`, `CacheControl`, and
|
||||
`StructuredOutputSpec` are the public shapes used by injected LLM clients.
|
||||
|
||||
- `File(path)`: file-backed artifact reference.
|
||||
- `Inline(body)`: inline artifact body.
|
||||
- `InlineWithURI(uri, body)`: inline artifact body with URI metadata.
|
||||
## Requests, Inputs, And Overrides
|
||||
|
||||
Input map keys must match the prompt's expected input names.
|
||||
`RunRequest` fields are `PromptID`, `PromptVersion`, `ProfileID`,
|
||||
`APIKey`, `Inputs`, `Vars`, `Execution`, `Validation`, and
|
||||
`Metadata`.
|
||||
|
||||
Input helpers are:
|
||||
|
||||
- `File(path)` for a file-backed artifact;
|
||||
- `Inline(body)` for inline content; and
|
||||
- `InlineWithURI(uri, body)` for inline content with URI metadata.
|
||||
|
||||
Required declared inputs must be supplied. Template rendering must also resolve
|
||||
every input name the prompt actually references. Extra entries in `Inputs`
|
||||
are not rejected solely because they are undeclared.
|
||||
|
||||
`ExecutionTargetOverride` supplies endpoint, model, credential-environment,
|
||||
service-tier, reasoning-effort, and extra-parameter overrides. Its numeric
|
||||
fields (`Temperature`, `MaxTokens`, `TopP`, and `TimeoutSeconds`) are
|
||||
pointers so explicit zero values are preserved. `OutputContract` supplies
|
||||
`Format`, `ValidationMode`, `SchemaPath`, and `RepairAttempts`.
|
||||
|
||||
`ExtraParams` accepts JSON-compatible values: strings, booleans, finite
|
||||
numbers, objects with string keys, arrays or slices, and nil. Unsupported
|
||||
values, non-string map keys, non-finite floats, and cycles return
|
||||
`ErrInvalidConfig` for profiles or `ErrInvalidRequest` for request
|
||||
overrides.
|
||||
|
||||
## Profiles And Credentials
|
||||
|
||||
`OpenAICompatibleProfile(OpenAICompatibleProfileConfig)` creates an
|
||||
in-memory `Profile`. Its public fields are `ID`, `Endpoint`, `Model`,
|
||||
`Temperature`, `MaxTokens`, `TopP`, `TimeoutSeconds`, `ServiceTier`,
|
||||
`ReasoningEffort`, `APIKeyRequired`, and `ExtraParams`.
|
||||
`WithProfiles` rejects duplicate IDs in one call.
|
||||
|
||||
A direct `RunRequest.APIKey` is request-scoped and takes precedence over
|
||||
`api_key_env` for the built-in client. It is excluded from JSON output and
|
||||
from `PreparedRun` and `RunResult`. The package's `String` and
|
||||
`GoString` methods report only whether a direct key is set. Do not use
|
||||
reflection-based dumps of request structs, which can bypass that redaction.
|
||||
|
||||
## Injected LLM Clients
|
||||
|
||||
Use `WithLLMClient` for tests or custom model integrations:
|
||||
`LLMClient` implements:
|
||||
|
||||
```go
|
||||
type fakeLLM struct{}
|
||||
|
||||
func (fakeLLM) Generate(ctx context.Context, req scriptorium.GenerateRequest) (*scriptorium.GenerateResponse, error) {
|
||||
return &scriptorium.GenerateResponse{
|
||||
Content: "generated text",
|
||||
Usage: scriptorium.TokenUsage{TotalTokens: 12},
|
||||
}, nil
|
||||
}
|
||||
|
||||
engine, err := scriptorium.NewEngine(cfg, scriptorium.WithLLMClient(fakeLLM{}))
|
||||
Generate(context.Context, GenerateRequest) (*GenerateResponse, error)
|
||||
```
|
||||
|
||||
Injected clients receive:
|
||||
|
||||
- rendered prompt;
|
||||
- effective execution target;
|
||||
- numeric target presence metadata;
|
||||
- structured-output spec when applicable;
|
||||
- direct request API key when provided.
|
||||
|
||||
Custom clients should not log raw prompts or API keys by default.
|
||||
|
||||
## Overrides And API Keys
|
||||
|
||||
`RunRequest` fields:
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `PromptID` | Prompt ID. |
|
||||
| `PromptVersion` | Optional prompt version filter. |
|
||||
| `ProfileID` | Optional profile override. |
|
||||
| `APIKey` | Direct per-request API key. |
|
||||
| `Inputs` | Input artifact references. |
|
||||
| `Vars` | Template variables. |
|
||||
| `Execution` | Per-request model overrides. |
|
||||
| `Validation` | Per-request output contract override. |
|
||||
| `Metadata` | Request metadata reserved for callers. |
|
||||
|
||||
`RunRequest.Execution` uses pointer fields for numeric values so explicit zero
|
||||
overrides are preserved:
|
||||
|
||||
```go
|
||||
zero := 0
|
||||
req.Execution = &scriptorium.ExecutionTargetOverride{
|
||||
MaxTokens: &zero,
|
||||
}
|
||||
```
|
||||
|
||||
Direct `RunRequest.APIKey` takes precedence over profile `api_key_env` for the
|
||||
default OpenAI-compatible client. It is request-scoped, uses `json:"-"`, and is
|
||||
not included in `PreparedRun` or `RunResult` JSON. Normal Go string formatting
|
||||
of `RunRequest` and `GenerateRequest` reports only whether a direct key is set.
|
||||
|
||||
Raw API keys do not belong in profile YAML, in-memory profiles, or app config.
|
||||
Avoid reflection-based debug dumps of request structs because exported fields
|
||||
remain visible to tools that bypass `String` and `GoString`.
|
||||
Injected clients receive the rendered prompt, effective execution target, numeric
|
||||
target-presence metadata, optional structured-output specification, and direct
|
||||
request API key. `GenerateResponse` returns content and `TokenUsage`.
|
||||
Custom clients should avoid logging raw prompts or credentials.
|
||||
|
||||
## Errors
|
||||
|
||||
Public methods wrap context while preserving stable sentinel checks with
|
||||
`errors.Is`:
|
||||
Public methods preserve these sentinel checks through `errors.Is`:
|
||||
|
||||
- `ErrInvalidConfig`
|
||||
- `ErrInvalidRequest`
|
||||
@@ -262,23 +170,5 @@ Public methods wrap context while preserving stable sentinel checks with
|
||||
- `ErrLLMGenerate`
|
||||
- `ErrValidation`
|
||||
|
||||
Example:
|
||||
|
||||
```go
|
||||
if errors.Is(err, scriptorium.ErrPromptNotFound) {
|
||||
return err
|
||||
}
|
||||
```
|
||||
|
||||
## Examples
|
||||
|
||||
Run the maintained prepare-only example from the repository root:
|
||||
|
||||
```bash
|
||||
go run ./examples/go-library/prepare
|
||||
```
|
||||
|
||||
See also:
|
||||
|
||||
- [Configuration reference](../config.md)
|
||||
- [Consumer integration overview](api.md)
|
||||
For interface selection and operational responsibilities, see the
|
||||
[consumer integration overview](api.md).
|
||||
|
||||
@@ -18,6 +18,8 @@ Start with:
|
||||
|
||||
- [Architecture policy](policy/architecture.md) for system boundaries,
|
||||
invariants, and non-goals;
|
||||
- [Internal component overview](internal/overview.md) for the current package
|
||||
and component map;
|
||||
- [Documentation policy](policy/documentation.md) before changing
|
||||
documentation;
|
||||
- [Testing policy](policy/testing.md) before adding, rewriting, or deleting
|
||||
@@ -27,17 +29,18 @@ Start with:
|
||||
|
||||
| Task | Read before changing |
|
||||
| --- | --- |
|
||||
| Public Go package or engine behavior | [Go package consumer contract](consumers/pkg-scriptorium.md), [runner internals](internal/runner.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
|
||||
| CLI commands, flags, output, or exit behavior | [CLI contract](cli.md) and [adapter internals](internal/adapters.md) |
|
||||
| HTTP routes, DTOs, limits, or status mapping | [HTTP API contract](api.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
|
||||
| Application configuration | [Configuration contract](config.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
|
||||
| Prompt, profile, schema, or artifact loading | [Configuration contract](config.md) and [source internals](internal/sources.md) |
|
||||
| Repository orientation or component responsibility | [Internal component overview](internal/overview.md) and [architecture policy](policy/architecture.md) |
|
||||
| Public Go package or engine behavior | [Go package consumer contract](consumers/pkg-scriptorium.md), [internal component overview](internal/overview.md), [runner internals](internal/runner.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
|
||||
| CLI commands, flags, output, or exit behavior | [CLI contract](cli.md), [internal component overview](internal/overview.md), and [adapter internals](internal/adapters.md) |
|
||||
| HTTP routes, DTOs, limits, or status mapping | [HTTP API contract](api.md), [internal component overview](internal/overview.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
|
||||
| Application configuration | [Configuration contract](config.md), [internal component overview](internal/overview.md), [adapter internals](internal/adapters.md), and [source internals](internal/sources.md) |
|
||||
| Prompt, profile, schema, or artifact loading | [Configuration contract](config.md), [internal component overview](internal/overview.md), and [source internals](internal/sources.md) |
|
||||
| Runner orchestration, rendering, validation, or repair | [Runner internals](internal/runner.md) and [source internals](internal/sources.md) |
|
||||
| OpenAI-compatible request or response behavior | [OpenAI-compatible integration](integrations/openai-compatible-chat.md), [runner internals](internal/runner.md), and [adapter internals](internal/adapters.md) |
|
||||
| OpenAI-compatible request or response behavior | [OpenAI-compatible integration](integrations/openai-compatible-chat.md), [LLM internals](internal/llm.md), [runner internals](internal/runner.md), and [adapter internals](internal/adapters.md) |
|
||||
| Subprocess behavior | [Subprocess integration](integrations/subprocess.md) and [CLI contract](cli.md) |
|
||||
| Runtime operation, recovery, or troubleshooting | [Operations](operations.md) and [troubleshooting](troubleshooting.md) |
|
||||
| Runtime operation or recovery | [Operations](operations.md) |
|
||||
| Examples or copyable assets | The owning contract for the demonstrated behavior and the related files under `examples/` |
|
||||
| Architecture decisions or future work | The [documentation policy](policy/documentation.md), relevant accepted ADRs under `adr/`, and relevant roadmap documents under `roadmap/` |
|
||||
| Architecture decisions or future work | The [documentation policy](policy/documentation.md), relevant accepted ADRs such as [ADR 0001](adr/0001-adopt-canonical-documentation-ownership.md), and relevant roadmap documents under `roadmap/` |
|
||||
|
||||
For cross-cutting changes, follow every applicable row. Internal component
|
||||
documents own detailed subsystem change recipes.
|
||||
|
||||
@@ -1,118 +1,67 @@
|
||||
# OpenAI-Compatible Chat Integration
|
||||
|
||||
## Scope
|
||||
This is the outbound wire contract for Scriptorium's OpenAI-compatible
|
||||
chat-completions client.
|
||||
|
||||
This document defines the outbound LLM contract implemented by `internal/llm/openai_compatible_client.go`.
|
||||
## Endpoint And Method
|
||||
|
||||
It documents only fields and behaviors currently serialized by code.
|
||||
Scriptorium uses the request endpoint override when present; otherwise it uses
|
||||
the configured client base URL. It removes a trailing slash and sends
|
||||
`POST /chat/completions`.
|
||||
|
||||
## Endpoint Construction
|
||||
For example, `http://localhost:8000/v1` becomes
|
||||
`http://localhost:8000/v1/chat/completions`.
|
||||
|
||||
Request endpoint is built as:
|
||||
## Request Payload
|
||||
|
||||
1. choose base URL:
|
||||
- `GenerateRequest.Target.Endpoint` if set
|
||||
- otherwise client config `BaseURL`
|
||||
2. trim trailing slash
|
||||
3. append `/chat/completions`
|
||||
The payload always contains `model` and rendered `messages`. It additionally
|
||||
contains these fields when applicable:
|
||||
|
||||
Example:
|
||||
| Field | Inclusion |
|
||||
| --- | --- |
|
||||
| `session_id` | Non-empty rendered prompt session ID. |
|
||||
| `temperature` | Non-zero effective value or an explicit zero override. |
|
||||
| `max_tokens` | Non-zero effective value or an explicit zero override. |
|
||||
| `top_p` | Non-zero effective value or an explicit zero override. |
|
||||
| `service_tier` | Any non-empty configured value. |
|
||||
| `reasoning_effort` | Any non-empty configured value. |
|
||||
| `response_format` | Structured output is requested. |
|
||||
| provider-specific fields | Flattened from `extra_params`. |
|
||||
|
||||
- base URL: `http://localhost:8000/v1`
|
||||
- final URL: `http://localhost:8000/v1/chat/completions`
|
||||
`service_tier` and `reasoning_effort` are forwarded without a provider value
|
||||
catalog; the selected backend decides which values it supports.
|
||||
|
||||
## Request Fields Sent
|
||||
`extra_params` are top-level JSON fields, not a nested object. Keys cannot be
|
||||
empty or collide with `model`, `session_id`, `messages`, `temperature`,
|
||||
`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or
|
||||
`response_format`. Values must be JSON-serializable.
|
||||
|
||||
Serialized JSON fields:
|
||||
A rendered `session_id` is sent as a top-level JSON field, not as a header.
|
||||
Empty values are omitted. The maximum length is 256 Unicode code points.
|
||||
|
||||
- `model` (required after fallback resolution)
|
||||
- `session_id` (only when the rendered prompt includes a non-empty session ID)
|
||||
- `messages` (rendered prompt messages)
|
||||
- `temperature` (when non-zero, or when explicitly overridden to zero)
|
||||
- `max_tokens` (when non-zero, or when explicitly overridden to zero)
|
||||
- `top_p` (when non-zero, or when explicitly overridden to zero)
|
||||
- `service_tier` (only when non-empty)
|
||||
- `reasoning_effort` (only when non-empty)
|
||||
- `response_format` (only when structured output is provided)
|
||||
- profile/request `extra_params` as additional provider-specific top-level fields
|
||||
|
||||
`service_tier` is provider-specific. OpenRouter currently documents request values such as `flex` and `priority`; Scriptorium forwards any non-empty configured value and lets the backend validate support.
|
||||
|
||||
`reasoning_effort` is provider-specific. Scriptorium forwards any non-empty configured value as top-level `reasoning_effort` and lets the backend validate support.
|
||||
|
||||
`extra_params` are flattened into the outbound JSON object. They are not wrapped in an `extra_params` object:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "gpt-4o-mini",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "rendered text"
|
||||
}
|
||||
],
|
||||
"provider_route": "primary",
|
||||
"provider_options": {
|
||||
"retry_budget": 2
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`extra_params` values must be JSON-compatible. Supported value shapes include strings, numbers, booleans, objects, and arrays.
|
||||
|
||||
Reserved `extra_params` keys are rejected before the HTTP request is made:
|
||||
|
||||
- `model`
|
||||
- `session_id`
|
||||
- `messages`
|
||||
- `temperature`
|
||||
- `max_tokens`
|
||||
- `top_p`
|
||||
- `service_tier`
|
||||
- `reasoning_effort`
|
||||
- `response_format`
|
||||
|
||||
Empty `extra_params` keys and values that cannot be encoded as JSON are also rejected before the HTTP request is made.
|
||||
|
||||
`session_id` is rendered from prompt YAML using request variables and serialized as a top-level JSON request field. Scriptorium does not send an `x-session-id` header. Empty rendered session IDs are omitted, and values longer than 256 characters are rejected before the HTTP request.
|
||||
|
||||
Messages without prompt cache control serialize with string `content`:
|
||||
Messages without cache control use string `content`. A message with cache
|
||||
control uses one text block:
|
||||
|
||||
```json
|
||||
{
|
||||
"role": "system",
|
||||
"content": "rendered text"
|
||||
}
|
||||
```
|
||||
|
||||
Messages with prompt cache control serialize as a single text content-block array:
|
||||
|
||||
```json
|
||||
{
|
||||
"role": "system",
|
||||
"content": [
|
||||
{
|
||||
"content": [{
|
||||
"type": "text",
|
||||
"text": "rendered text",
|
||||
"cache_control": {
|
||||
"type": "ephemeral",
|
||||
"ttl": "1h"
|
||||
}
|
||||
}
|
||||
]
|
||||
"cache_control": {"type": "ephemeral", "ttl": "1h"}
|
||||
}]
|
||||
}
|
||||
```
|
||||
|
||||
When cache-control `ttl` is unset in the prompt definition, `ttl` is omitted from the outbound payload.
|
||||
|
||||
Structured output is currently `json_schema` only, serialized as:
|
||||
When the prompt omits cache-control `ttl`, the payload omits `ttl`.
|
||||
Structured JSON Schema output is sent as:
|
||||
|
||||
```json
|
||||
{
|
||||
"response_format": {
|
||||
"type": "json_schema",
|
||||
"json_schema": {
|
||||
"name": "...",
|
||||
"name": "schema name",
|
||||
"strict": true,
|
||||
"schema": {"type": "object"}
|
||||
}
|
||||
@@ -120,74 +69,40 @@ Structured output is currently `json_schema` only, serialized as:
|
||||
}
|
||||
```
|
||||
|
||||
## Authentication Header
|
||||
## Authentication And Timeout
|
||||
|
||||
If `Target.APIKey` is set:
|
||||
When a direct request API key is present, Scriptorium sends
|
||||
`Authorization: Bearer <key>` and does not read `api_key_env`. Otherwise, it
|
||||
resolves the configured non-empty `api_key_env` at request time and sends the
|
||||
same header. If neither mechanism supplies a key, it sends no
|
||||
`Authorization` header.
|
||||
|
||||
- set `Authorization: Bearer <value>`
|
||||
- do not read `Target.APIKeyEnv`
|
||||
The configured client timeout applies by default. A positive effective
|
||||
`timeout_seconds` replaces it. An explicit request override of zero disables
|
||||
the HTTP-client timeout; negative values are rejected before a request is sent.
|
||||
|
||||
If `Target.APIKey` is empty and `Target.APIKeyEnv` is set:
|
||||
## Response Subset And Failures
|
||||
|
||||
- resolve environment variable value at request time
|
||||
- set `Authorization: Bearer <value>`
|
||||
A successful provider response must supply non-empty
|
||||
`choices[0].message.content`. Scriptorium reads these optional or required
|
||||
usage fields when present:
|
||||
|
||||
If the environment variable is unset/empty:
|
||||
|
||||
- request fails before HTTP call (`ErrInvalidRequest`)
|
||||
|
||||
If both `Target.APIKey` and `Target.APIKeyEnv` are empty:
|
||||
|
||||
- no `Authorization` header is sent
|
||||
|
||||
## Timeout Behavior
|
||||
|
||||
Base timeout comes from client configuration.
|
||||
|
||||
Per-request override:
|
||||
|
||||
- if `Target.TimeoutSeconds > 0`, use that value for request timeout
|
||||
- if `Target.TimeoutSeconds == 0` and the value came from an explicit request override, disable the HTTP client timeout
|
||||
- if `Target.TimeoutSeconds < 0`, request is rejected (`ErrInvalidRequest`)
|
||||
|
||||
## Response Expectations
|
||||
|
||||
Expected successful response shape (subset used):
|
||||
|
||||
- `choices[0].message.content`
|
||||
- `usage.prompt_tokens`
|
||||
- `usage.completion_tokens`
|
||||
- `usage.total_tokens`
|
||||
- `usage.prompt_tokens_details.cached_tokens` (optional)
|
||||
- `usage.cache_write_tokens` (optional)
|
||||
- `usage.prompt_tokens_details.cached_tokens`
|
||||
- `usage.cache_write_tokens`
|
||||
|
||||
Absent cache usage fields are treated as zero. Parsed cache usage is exposed through run results and adapter response surfaces as:
|
||||
Missing cache usage is reported as zero. Invalid JSON, an empty choices array,
|
||||
or empty first-choice content is a malformed provider response. Network and
|
||||
request-construction failures, non-2xx responses, and malformed responses fail
|
||||
the outbound call. Provider response bodies are not exposed by this client.
|
||||
|
||||
- `cached_tokens`
|
||||
- `cache_write_tokens`
|
||||
The client does not implement built-in retries, tool calls, top-level
|
||||
`cache_control`, or multi-request payload modes.
|
||||
|
||||
Malformed response conditions include:
|
||||
## Related References
|
||||
|
||||
- invalid JSON
|
||||
- empty `choices`
|
||||
- empty `choices[0].message.content`
|
||||
|
||||
Malformed responses return `ErrMalformedResponse`.
|
||||
|
||||
## Error Handling
|
||||
|
||||
- network/request-construction failures: `ErrRequestFailed`
|
||||
- non-2xx HTTP status: `ErrUnexpectedStatus` (includes status code; provider response bodies are not included)
|
||||
- malformed response shape/content: `ErrMalformedResponse`
|
||||
|
||||
## Unsupported Or Non-Serialized Fields
|
||||
|
||||
The client does not serialize top-level `cache_control`.
|
||||
|
||||
No built-in retries, tool-calls, or multi-request payload modes are implemented in this client.
|
||||
|
||||
## Relationship To Runner
|
||||
|
||||
When prompt validation mode is `json_schema`, runner prepares a structured-output schema spec and passes it to the client as `StructuredOutput`.
|
||||
|
||||
The client only serializes the provider request payload; it does not load schema files itself.
|
||||
Prompt schema preparation and runner orchestration are described in
|
||||
[runner internals](../internal/runner.md). Prompt and profile configuration is
|
||||
defined by the [configuration reference](../config.md).
|
||||
|
||||
@@ -1,130 +1,41 @@
|
||||
# Subprocess Integration
|
||||
|
||||
This document defines the supported subprocess contract for downstream
|
||||
applications invoking Scriptorium through the public CLI.
|
||||
This document covers process-boundary behavior for callers that invoke
|
||||
Scriptorium as a child process. Command syntax, flags, output, and exit codes
|
||||
are defined by the [CLI reference](../cli.md). Interface selection belongs in
|
||||
the [consumer integration overview](../consumers/api.md).
|
||||
|
||||
This is a CLI contract. Go callers that want an in-process typed API should use
|
||||
the [package guide](../consumers/pkg-scriptorium.md).
|
||||
## Process Contract
|
||||
|
||||
## Supported Commands
|
||||
Use `scriptorium render` when the caller needs prepared output without a model
|
||||
call, and `scriptorium run` for generation. Pass an explicit `--config` or
|
||||
make the configuration search paths available to the child process; configuration
|
||||
discovery, fields, profile selection, and credential mechanisms are defined in
|
||||
the [configuration reference](../config.md).
|
||||
|
||||
Downstream applications should invoke:
|
||||
Pass required API-key environment variables through the child environment. Do
|
||||
not place raw API keys in arguments. Keep the environment limited to the values
|
||||
needed for the selected profile.
|
||||
|
||||
- `scriptorium render` for preflight/debug output without LLM execution.
|
||||
- `scriptorium run` for generation.
|
||||
## Streams And Output Ownership
|
||||
|
||||
`scriptorium serve` is an HTTP service command, not the recommended subprocess
|
||||
contract for per-request execution.
|
||||
Capture stdout and stderr separately. Stdout contains the requested artifact or
|
||||
prepared output unless the caller selects an output file; stderr contains
|
||||
summaries, diagnostics, and server messages. The exact destinations and status
|
||||
meanings are part of the [CLI reference](../cli.md), not a stable stderr data
|
||||
protocol.
|
||||
|
||||
## Recommended Invocation Shapes
|
||||
When using `--out`, the caller owns the output path, its permissions, and
|
||||
cleanup. Treat rendered prompts, generated artifacts, stdout, and stderr as
|
||||
potentially sensitive.
|
||||
|
||||
Render:
|
||||
## Cancellation And Recovery
|
||||
|
||||
```bash
|
||||
scriptorium render \
|
||||
--config <config_path> \
|
||||
--prompt <prompt_id> \
|
||||
--input transcript=<path> \
|
||||
--format json
|
||||
```
|
||||
A CLI invocation performs one synchronous request and creates no durable run
|
||||
state. A supervising process that needs cancellation must terminate the child
|
||||
process according to its own process-management policy. A later invocation is a
|
||||
new request and can make another model call; there is no resume or checkpoint
|
||||
protocol.
|
||||
|
||||
Run:
|
||||
|
||||
```bash
|
||||
scriptorium run \
|
||||
--config <config_path> \
|
||||
--prompt <prompt_id> \
|
||||
--input transcript=<path> \
|
||||
--out <artifact_path>
|
||||
```
|
||||
|
||||
Callers may add:
|
||||
|
||||
- `--profile <profile_id>`
|
||||
- repeatable `--input name=path`
|
||||
- repeatable `--var name=value`
|
||||
- runtime overrides when explicitly needed, such as `--model`, `--llm-base-url`, `--api-key-env`, and `--timeout`
|
||||
|
||||
Do not pass raw API keys as command arguments.
|
||||
|
||||
## Config And Directory Behavior
|
||||
|
||||
Callers can rely on resolved app config or pass explicit paths.
|
||||
|
||||
Default config search order:
|
||||
|
||||
1. `/usr/local/etc/scriptorium/config.yml`
|
||||
2. `/etc/scriptorium/config.yml`
|
||||
|
||||
Rules:
|
||||
|
||||
- Explicit `--config` requires file existence and valid syntax.
|
||||
- CLI flags override config values.
|
||||
- `run` and `render` require an effective `prompt_dir`.
|
||||
- `profile_dir` is optional because built-in profiles are available.
|
||||
|
||||
## Profile Selection
|
||||
|
||||
Profile selection follows runner behavior:
|
||||
|
||||
1. explicit `--profile`
|
||||
2. prompt `default_profile`
|
||||
3. error if neither is available
|
||||
|
||||
Treat prompt and profile IDs as deployment configuration, not hardcoded business
|
||||
logic.
|
||||
|
||||
## Input And Variable Contract
|
||||
|
||||
- Inputs use repeated `--input name=path`.
|
||||
- Input names must match prompt definition input names.
|
||||
- Variables use repeated `--var name=value`.
|
||||
- Both flags also accept comma-separated mappings.
|
||||
- Prefer file inputs for large content.
|
||||
|
||||
CLI inputs are file references. HTTP-only `inline` references are documented in
|
||||
the [HTTP API reference](../api.md).
|
||||
|
||||
## Environment Contract
|
||||
|
||||
- Pass through required API-key environment variables referenced by `api_key_env`.
|
||||
- Keep subprocess environments scoped to required variables.
|
||||
- Use `--api-key-env` only to name an environment variable.
|
||||
- Never pass raw API keys via argv.
|
||||
|
||||
## Stdout And Stderr
|
||||
|
||||
`run`:
|
||||
|
||||
- stdout: generated artifact body unless `--out` is used.
|
||||
- stderr: success summary and errors.
|
||||
|
||||
`render`:
|
||||
|
||||
- stdout: prepared-run output unless `--out` is used.
|
||||
- stderr: errors.
|
||||
|
||||
Capture stdout and stderr separately. Do not parse stderr as a stable data
|
||||
format beyond exit status handling.
|
||||
|
||||
## Exit Status Contract
|
||||
|
||||
- `0`: success.
|
||||
- `1`: parse, config, load, render, generation, IO, or runtime error.
|
||||
- `2`: `run` completed and output was written, but validation failed.
|
||||
|
||||
A `run` exit code `2` can still produce output on stdout or at `--out`.
|
||||
Consumers must decide whether to keep or discard that output.
|
||||
|
||||
## Security Notes
|
||||
|
||||
- Treat generated artifacts, rendered prompts, stdout, and stderr as potentially sensitive.
|
||||
- Use controlled output paths and access controls for persisted artifacts.
|
||||
- Avoid logging full rendered prompts or generated artifacts by default.
|
||||
|
||||
## Canonical References
|
||||
|
||||
- CLI behavior: [CLI reference](../cli.md)
|
||||
- Config and file formats: [Configuration reference](../config.md)
|
||||
- Operations: [Operations guide](../operations.md)
|
||||
- Troubleshooting: [Troubleshooting](../troubleshooting.md)
|
||||
For deployment, filesystem permissions, and sensitive-artifact handling, see
|
||||
the [operations guide](../operations.md).
|
||||
|
||||
@@ -2,154 +2,112 @@
|
||||
|
||||
## Purpose
|
||||
|
||||
Adapters translate external interfaces into domain requests and translate domain results back out. They wire dependencies, apply app config, and own IO concerns, but they do not make runner decisions.
|
||||
Adapters translate external inputs into domain requests, compose dependencies,
|
||||
and translate domain results or errors back to their interface. They own IO and
|
||||
presentation mechanics; use-case decisions remain in `internal/usecase`.
|
||||
|
||||
Source-loading behavior belongs in `docs/internal/sources.md`. User-facing CLI, HTTP, and package contracts belong in `docs/cli.md`, `docs/api.md`, and `docs/consumers/pkg-scriptorium.md`.
|
||||
External contracts are canonical in the [CLI reference](../cli.md), [HTTP API
|
||||
reference](../api.md), and [Go package contract](../consumers/pkg-scriptorium.md).
|
||||
|
||||
## Adapter Map
|
||||
## Components And Collaborators
|
||||
|
||||
- `cmd/scriptorium`: process entrypoint.
|
||||
- `internal/adapter/cli`: command parsing, config handoff, runner construction, stdout/stderr, exit codes.
|
||||
- `internal/adapter/http`: `POST /v1/runs` request/response mapping and HTTP error/status mapping.
|
||||
- root package `scriptorium`: public Go facade over internal runner types and dependencies.
|
||||
- `cmd/scriptorium` passes process arguments and streams to
|
||||
`internal/adapter/cli`.
|
||||
- `internal/adapter/cli` parses commands, resolves application settings through
|
||||
`internal/config`, constructs a runner, and owns process output handling.
|
||||
- `internal/adapter/http` decodes DTOs, maps them to `domain.RunRequest`, calls
|
||||
a runner interface, and maps errors and results to HTTP DTOs.
|
||||
- The root `scriptorium` package maps its public types and options to internal
|
||||
collaborators and maps selected internal errors to public sentinels.
|
||||
- `internal/format` formats prepared runs for the CLI; `internal/llm`,
|
||||
`internal/prompt`, and source packages supply runner dependencies.
|
||||
|
||||
Supporting implementation packages used during adapter wiring:
|
||||
## Wiring Flows
|
||||
|
||||
- `internal/config`
|
||||
- `internal/defaults`
|
||||
- `internal/format`
|
||||
- `internal/llm`
|
||||
- `internal/prompt`
|
||||
### CLI
|
||||
|
||||
## Inputs And Outputs
|
||||
The CLI resolves configuration before constructing dependencies. `run` builds a
|
||||
runner with the ordinary composite artifact reader and invokes `Runner.Run`;
|
||||
`render` uses the same wiring and invokes `Runner.Prepare`; `serve` replaces the
|
||||
file reader with the restricted artifact reader, builds an HTTP handler, and
|
||||
starts the server.
|
||||
|
||||
CLI adapter:
|
||||
Parser state records whether numeric runtime values were explicitly supplied.
|
||||
That presence is carried into `domain.ExecutionTargetOverride`, allowing the
|
||||
runner to distinguish omitted values from explicit zero overrides.
|
||||
|
||||
- Input: process args, optional config file, filesystem sources, environment variables.
|
||||
- Output: process exit code, stdout artifact/prepared output, stderr summaries and errors.
|
||||
### HTTP
|
||||
|
||||
HTTP adapter:
|
||||
The handler first enforces transport limits, strict JSON decoding, and the
|
||||
minimal request shape. It maps DTO values to domain types without deciding
|
||||
prompt selection, source behavior, or validation semantics. On success it maps
|
||||
the domain result to the response DTO; on failure it uses `errors.Is` over
|
||||
runner, source, artifact, and profile errors to choose the public error mapping.
|
||||
|
||||
- Input: HTTP request method/path/headers/body for `POST /v1/runs`.
|
||||
- Output: JSON success or error body with mapped status code.
|
||||
The [HTTP API reference](../api.md) owns the route, DTO schema, status codes,
|
||||
and externally observable limit behavior.
|
||||
|
||||
Public Go facade:
|
||||
### Public Go Facade
|
||||
|
||||
- Input: typed `scriptorium.Config`, `Option`, and `RunRequest` values.
|
||||
- Output: typed `PreparedRun` and `RunResult` values plus public sentinel errors.
|
||||
`NewEngine` applies public options, selects filesystem, `fs.FS`, single-file,
|
||||
or in-memory dependencies, and constructs a runner. The conversion functions
|
||||
copy maps and slices across the boundary so callers do not receive internal
|
||||
domain values. The facade maps selected internal errors to the public sentinel
|
||||
set and keeps direct request API keys out of public results.
|
||||
|
||||
## Boundaries
|
||||
## Package-Local Guarantees
|
||||
|
||||
- Adapters convert external shapes to `domain.RunRequest` and back.
|
||||
- Runner orchestration remains in `internal/usecase`.
|
||||
- Prompt/profile/schema/artifact source rules remain in repository, validator, and artifact packages.
|
||||
- LLM provider request serialization remains in `internal/llm`.
|
||||
- Public package types are facade types; internal domain types do not leak across the package boundary.
|
||||
- Adapters do not embed runner orchestration or source-loading decisions.
|
||||
- Configuration is resolved before adapter dependency composition.
|
||||
- CLI and HTTP create runners without a repairer; a repairer is available only
|
||||
through explicit internal runner construction.
|
||||
- DTO conversion preserves explicit numeric-override presence.
|
||||
- Error mapping matches error identities, not error text.
|
||||
- No adapter creates durable run state; caller-selected output files are not
|
||||
application state.
|
||||
|
||||
## Config Fields Used
|
||||
## Failure And Verification Boundaries
|
||||
|
||||
Adapter app settings:
|
||||
Keep external error payloads concise, preserve strict external decoding, and do
|
||||
not serialize resolved secret values. Validation content failures remain result
|
||||
state; runtime failures remain errors for the relevant adapter to map.
|
||||
|
||||
- `prompt_dir`
|
||||
- `profile_dir`
|
||||
- `schema_dir`
|
||||
- `server.addr`
|
||||
- `server.artifact_root`
|
||||
- `server.max_request_bytes`
|
||||
- `server.max_artifact_bytes`
|
||||
- `server.max_response_bytes`
|
||||
- `defaults.render_format`
|
||||
|
||||
Execution request/profile settings passed through the runner:
|
||||
|
||||
- `endpoint`
|
||||
- `model`
|
||||
- `temperature`
|
||||
- `max_tokens`
|
||||
- `top_p`
|
||||
- `timeout_seconds`
|
||||
- `service_tier`
|
||||
- `api_key_env`
|
||||
- `reasoning_effort`
|
||||
- `extra_params`
|
||||
|
||||
CLI and HTTP preserve numeric override presence so omitted values and explicit zero values remain distinct.
|
||||
|
||||
## CLI Adapter
|
||||
|
||||
Implemented commands:
|
||||
|
||||
- `run`
|
||||
- `render`
|
||||
- `serve`
|
||||
|
||||
Behavior:
|
||||
|
||||
- `run` constructs a runner with direct filesystem artifact reading and calls `Runner.Run`.
|
||||
- `render` constructs a runner and calls `Runner.Prepare`; it does not call the LLM.
|
||||
- `serve` constructs a restricted artifact reader and HTTP handler, then starts an unauthenticated HTTP server.
|
||||
- `run` exits `2` when generation succeeds but validation fails.
|
||||
- parse, runtime, and output-write errors exit `1`.
|
||||
- deprecated `--prompt-id` and `--profile-id` aliases are accepted.
|
||||
|
||||
## HTTP Adapter
|
||||
|
||||
Behavior:
|
||||
|
||||
- Accepts only `POST /v1/runs`.
|
||||
- Decodes JSON strictly and rejects unknown fields and trailing JSON tokens.
|
||||
- Rejects empty `prompt_id` and empty `inputs` before calling the runner.
|
||||
- Does not accept raw API key values in the request body.
|
||||
- Returns validation failures as `200` responses with failed validation details.
|
||||
- Maps request-body, artifact, and encoded-response size failures to `413`.
|
||||
- Maps domain and repository errors to stable error codes without returning wrapped internal cause text.
|
||||
|
||||
The HTTP adapter has no built-in authentication or authorization. Deployment controls must be provided outside the process.
|
||||
|
||||
## Public Go Facade
|
||||
|
||||
Behavior:
|
||||
|
||||
- `NewEngine` wires the same default runner components as CLI/HTTP unless options override them.
|
||||
- Prompt, profile, and schema sources may come from directories, single files, or `fs.FS` roots.
|
||||
- `WithProfiles` adds in-memory profiles ahead of file-backed and built-in profiles.
|
||||
- `WithLLMClient` injects custom model behavior.
|
||||
- `RunRequest.APIKey` is request-scoped and direct; it is used only for generation and is stripped from public results.
|
||||
- internal errors are mapped to public sentinels in `errors.go`.
|
||||
|
||||
## Failure Behavior
|
||||
|
||||
Adapters should:
|
||||
|
||||
- keep external error payloads concise and stable.
|
||||
- avoid leaking raw secret values.
|
||||
- use sentinels and typed errors for mapping.
|
||||
- preserve strict external input decoding.
|
||||
- keep validation content failures distinct from runtime errors.
|
||||
|
||||
CLI writes human-readable summaries to stderr. HTTP writes JSON error envelopes. The public Go facade returns typed errors.
|
||||
|
||||
## State And Manifests
|
||||
|
||||
Adapters do not add durable run state.
|
||||
|
||||
- No adapter writes run manifests.
|
||||
- No adapter implements checkpoint, skip, or resume behavior.
|
||||
- CLI output files are caller-selected artifacts, not internal state.
|
||||
|
||||
## Tests To Inspect
|
||||
Inspect focused tests when changing this area:
|
||||
|
||||
- `internal/adapter/cli/run_test.go`
|
||||
- `internal/adapter/http/handler_test.go`
|
||||
- `engine_test.go`
|
||||
- `internal/format/prepared_run_test.go`
|
||||
- `internal/llm/openai_compatible_client_test.go`
|
||||
|
||||
## Architectural Invariants
|
||||
Run the affected adapter package tests and recheck the relevant canonical
|
||||
contract. The [testing policy](../policy/testing.md) owns global test
|
||||
sufficiency guidance.
|
||||
|
||||
- Adapter packages stay thin and translation-focused.
|
||||
- App config is resolved before dependency construction.
|
||||
- External input strictness is part of contract stability.
|
||||
- CLI and HTTP construct runners without a repairer.
|
||||
- HTTP endpoint details remain canonical in `docs/api.md`.
|
||||
- Public Go package details remain canonical in `docs/consumers/pkg-scriptorium.md`.
|
||||
## Change Recipes
|
||||
|
||||
### Application Configuration Fields
|
||||
|
||||
1. Add the field to the relevant `internal/config` shape and default handling.
|
||||
2. Parse and validate it, then preserve configuration and CLI-override
|
||||
precedence while wiring it through its consuming adapter.
|
||||
3. Add focused configuration and adapter tests for parsing, mapping, and
|
||||
effective behavior.
|
||||
4. Update the [configuration contract](../config.md) and any affected external
|
||||
contract.
|
||||
|
||||
### CLI Flags
|
||||
|
||||
1. Add the flag to the relevant parser in `internal/adapter/cli/run.go`.
|
||||
2. Keep command scope and application-configuration precedence intentional.
|
||||
3. Add or update parser and command tests in
|
||||
`internal/adapter/cli/run_test.go`.
|
||||
4. Update the [CLI contract](../cli.md) and affected maintained examples.
|
||||
|
||||
### Adapter Capabilities
|
||||
|
||||
1. Define or reuse the appropriate domain or use-case interface boundary.
|
||||
2. Implement translation and IO behavior without moving use-case decisions out
|
||||
of `internal/usecase`.
|
||||
3. Add focused mapping, parsing, and error-behavior tests.
|
||||
4. Update this document and the affected public or integration contract. Update
|
||||
[source internals](sources.md) when source-loading behavior changes.
|
||||
|
||||
85
docs/internal/llm.md
Normal file
85
docs/internal/llm.md
Normal file
@@ -0,0 +1,85 @@
|
||||
# LLM Internals
|
||||
|
||||
## Purpose
|
||||
|
||||
`internal/llm` defines the provider-neutral `Client` interface and the
|
||||
OpenAI-compatible client implementation. The [OpenAI-compatible integration
|
||||
contract](../integrations/openai-compatible-chat.md) owns the outbound HTTP wire
|
||||
format and protocol behavior.
|
||||
|
||||
## Construction
|
||||
|
||||
`NewOpenAICompatibleClient` validates a non-empty configured base URL, records
|
||||
an optional default model, and establishes the default timeout. A non-positive
|
||||
configured timeout uses the internal default.
|
||||
|
||||
When callers supply an `http.Client`, construction clones it rather than
|
||||
mutating the caller's instance. A supplied client with no timeout receives the
|
||||
resolved default in the clone; a supplied non-zero timeout is retained. The
|
||||
client stores the trimmed base URL, default model, timeout, and cloned client.
|
||||
|
||||
## Generate Flow
|
||||
|
||||
`Generate` receives a `domain.GenerateRequest` from the runner:
|
||||
|
||||
1. validate the effective timeout and choose the request endpoint;
|
||||
2. map the domain request to the internal wire-request representation;
|
||||
3. validate and flatten extra parameters, encode JSON, and create the HTTP
|
||||
request;
|
||||
4. prefer a direct API key, otherwise resolve the configured key environment
|
||||
variable;
|
||||
5. derive a request HTTP client when an explicit timeout changes the configured
|
||||
client;
|
||||
6. execute the request, reject non-success status responses without returning
|
||||
provider response bodies; and
|
||||
7. decode the response subset into `domain.GenerateResponse`.
|
||||
|
||||
`openAIChatRequestFromGenerateRequest` is the conversion boundary for effective
|
||||
model defaults, explicit numeric-presence state, rendered messages, structured
|
||||
output, and session-ID validation. `openAIChatRequestPayload` protects reserved
|
||||
fields and JSON encoding before an HTTP call. The external payload shape is
|
||||
defined only in the [integration contract](../integrations/openai-compatible-chat.md).
|
||||
|
||||
## Error Categories
|
||||
|
||||
The package uses these internal sentinels:
|
||||
|
||||
- `ErrInvalidConfig` for invalid client construction;
|
||||
- `ErrInvalidRequest` for invalid effective generation input;
|
||||
- `ErrRequestFailed` for request construction or transport failures;
|
||||
- `ErrUnexpectedStatus` for non-success HTTP responses; and
|
||||
- `ErrMalformedResponse` for invalid or incomplete successful-response data.
|
||||
|
||||
The runner maps an invalid LLM request to its invalid-request category and
|
||||
other LLM failures to its generation category. Adapters then apply their public
|
||||
error contracts.
|
||||
|
||||
## Package-Local Guarantees
|
||||
|
||||
- The default-model fallback happens before wire encoding.
|
||||
- Per-request timeout handling clones a configured HTTP client when needed; it
|
||||
does not mutate shared client state.
|
||||
- Direct API keys take precedence over environment lookup within this client.
|
||||
- Provider response bodies are discarded for non-success status responses.
|
||||
- The client does not implement retries, tool calls, or a stateful session
|
||||
store.
|
||||
|
||||
## Verification And Change Recipe
|
||||
|
||||
Inspect:
|
||||
|
||||
- `internal/llm/openai_compatible_client_test.go`
|
||||
- `internal/usecase/runner_test.go`
|
||||
- `internal/adapter/http/handler_test.go`
|
||||
|
||||
When changing the client:
|
||||
|
||||
1. keep domain-to-wire mapping inside `internal/llm` and preserve the `Client`
|
||||
interface;
|
||||
2. test construction, timeout selection, mapping, and error categorization;
|
||||
3. update the [OpenAI-compatible integration contract](../integrations/openai-compatible-chat.md)
|
||||
for any observable wire or protocol change; and
|
||||
4. update [runner internals](runner.md) if the client boundary or structured
|
||||
output handoff changes.
|
||||
|
||||
The [testing policy](../policy/testing.md) owns global test sufficiency.
|
||||
48
docs/internal/overview.md
Normal file
48
docs/internal/overview.md
Normal file
@@ -0,0 +1,48 @@
|
||||
# Internal Component Overview
|
||||
|
||||
## Purpose
|
||||
|
||||
This is the inventory of Scriptorium's implemented components for contributors.
|
||||
The [architecture policy](../policy/architecture.md) owns normative boundaries
|
||||
and invariants; public behavior belongs in the linked contracts.
|
||||
|
||||
## Public And Command Entrypoints
|
||||
|
||||
| Component | Implemented responsibility | References |
|
||||
| --- | --- | --- |
|
||||
| Root package `scriptorium` | Public Go facade that constructs the engine, exposes request/result types and options, and maps internal errors. | [Go package contract](../consumers/pkg-scriptorium.md), [adapter internals](adapters.md) |
|
||||
| `cmd/scriptorium` | Process entrypoint that delegates command execution to the CLI adapter. | [CLI contract](../cli.md), [adapter internals](adapters.md) |
|
||||
|
||||
## Adapters, Domain, And Use Case
|
||||
|
||||
| Component | Implemented responsibility | References |
|
||||
| --- | --- | --- |
|
||||
| `internal/adapter/cli` | Parses CLI commands, applies application wiring, and handles process input and output. | [CLI contract](../cli.md), [adapter internals](adapters.md) |
|
||||
| `internal/adapter/http` | Maps HTTP requests and responses to domain operations and maps public errors. | [HTTP API contract](../api.md), [adapter internals](adapters.md) |
|
||||
| `internal/domain` | Defines core request, result, output-contract, and LLM-boundary types. | [runner internals](runner.md) |
|
||||
| `internal/usecase` | Implements `Runner` preparation, execution, validation coordination, and the repairer boundary. | [runner internals](runner.md) |
|
||||
|
||||
## Configuration And Sources
|
||||
|
||||
| Component | Implemented responsibility | References |
|
||||
| --- | --- | --- |
|
||||
| `internal/config` | Loads application settings, applies defaults, and applies CLI overrides. | [configuration contract](../config.md), [adapter internals](adapters.md) |
|
||||
| `internal/defaults` | Holds compile-time default values used when application settings are resolved. | [configuration contract](../config.md) |
|
||||
| `internal/promptdef` | Loads prompt definitions from filesystem and `fs.FS` sources. | [configuration contract](../config.md), [source internals](sources.md) |
|
||||
| `internal/profile` | Loads filesystem and `fs.FS` execution profiles and combines profile repositories. | [configuration contract](../config.md), [source internals](sources.md) |
|
||||
| `internal/profile/builtin` | Provides embedded built-in execution profiles as a repository. | [configuration contract](../config.md), [source internals](sources.md) |
|
||||
| `internal/filecatalog` | Provides shared YAML discovery and source-root helpers. | [source internals](sources.md) |
|
||||
| `internal/artifact` | Reads inline and file-backed input artifacts. | [configuration contract](../config.md), [HTTP API contract](../api.md), [source internals](sources.md) |
|
||||
| `internal/prompt` | Renders prompt templates into messages. | [runner internals](runner.md) |
|
||||
|
||||
## Formatting, Validation, And Model Access
|
||||
|
||||
| Component | Implemented responsibility | References |
|
||||
| --- | --- | --- |
|
||||
| `internal/format` | Formats prepared-run information for CLI output. | [CLI contract](../cli.md), [adapter internals](adapters.md) |
|
||||
| `internal/validate` | Defines validation interfaces and provides standard filesystem and `fs.FS` schema validation. | [configuration contract](../config.md), [source internals](sources.md), [runner internals](runner.md) |
|
||||
| `internal/llm` | Defines the provider-neutral LLM client boundary and its OpenAI-compatible implementation. | [OpenAI-compatible integration](../integrations/openai-compatible-chat.md), [LLM internals](llm.md), [runner internals](runner.md) |
|
||||
|
||||
Focused internal documents describe the components that have detailed
|
||||
orchestration, adapter, or source behavior. Package tests live alongside the
|
||||
implementation and are identified in those focused documents where relevant.
|
||||
@@ -2,145 +2,118 @@
|
||||
|
||||
## Purpose
|
||||
|
||||
`internal/usecase.Runner` is the core prompt-execution orchestrator. It prepares prompt requests, calls the configured LLM client for `Run`, validates generated output, and returns domain results.
|
||||
`internal/usecase.Runner` is the prompt-execution orchestrator. It prepares
|
||||
domain requests, invokes an injected LLM client, validates output, and returns
|
||||
domain results. Transport parsing, response mapping, and public type conversion
|
||||
remain outside this package.
|
||||
|
||||
Transport parsing, DTOs, CLI output, HTTP status mapping, and public package type conversion belong outside the runner.
|
||||
The [configuration reference](../config.md) owns prompt, profile, schema, and
|
||||
runtime-setting definitions. Public error behavior is defined by the
|
||||
[HTTP API](../api.md) and [Go package](../consumers/pkg-scriptorium.md)
|
||||
contracts.
|
||||
|
||||
## Inputs And Outputs
|
||||
## Dependencies And Construction
|
||||
|
||||
Primary inputs:
|
||||
`Runner` receives these collaborators:
|
||||
|
||||
- `domain.RunRequest`
|
||||
- repositories/readers/renderers/validators injected at construction
|
||||
- `context.Context` for cancellation
|
||||
- `promptdef.Repository`;
|
||||
- `profile.Repository`;
|
||||
- `artifact.Reader`;
|
||||
- `prompt.Renderer`;
|
||||
- `llm.Client`;
|
||||
- `validate.Validator`; and
|
||||
- an optional `OutputRepairer`.
|
||||
|
||||
Primary outputs:
|
||||
|
||||
- `domain.PreparedRun` from `Prepare`
|
||||
- `domain.RunResult` from `Run`
|
||||
- wrapped sentinel errors for adapter mapping
|
||||
|
||||
LLM boundary types:
|
||||
|
||||
- `domain.GenerateRequest`
|
||||
- `domain.GenerateResponse`
|
||||
|
||||
## Dependencies
|
||||
|
||||
`Runner` depends on package interfaces instead of concrete adapter types:
|
||||
|
||||
- `promptdef.Repository`
|
||||
- `profile.Repository`
|
||||
- `artifact.Reader`
|
||||
- `prompt.Renderer`
|
||||
- `llm.Client`
|
||||
- `validate.Validator`
|
||||
- optional `usecase.OutputRepairer`
|
||||
|
||||
The CLI, HTTP adapter, and public Go package construct these dependencies and pass them in.
|
||||
|
||||
## Config Fields
|
||||
|
||||
`Runner` does not read app config files. Effective behavior is determined by injected dependencies and the `domain.RunRequest`.
|
||||
|
||||
Adapter wiring commonly reflects these app config fields:
|
||||
|
||||
- `prompt_dir`
|
||||
- `profile_dir`
|
||||
- `schema_dir`
|
||||
- `server.artifact_root`
|
||||
- HTTP request/artifact/response size limits
|
||||
|
||||
Runtime model settings are resolved from the selected profile plus request overrides.
|
||||
`NewRunner` constructs a runner without a repairer. `NewRunnerWithRepairer`
|
||||
accepts one explicitly. Adapters and the public engine choose concrete
|
||||
repositories and readers; the runner does not load application configuration.
|
||||
|
||||
## Prepare Flow
|
||||
|
||||
`Prepare`:
|
||||
`Prepare` performs one deterministic preparation pass for a request:
|
||||
|
||||
1. requires a non-empty prompt ID.
|
||||
2. loads the prompt definition and computes its hash.
|
||||
3. selects the profile from request `profile_id`, then prompt `default_profile`.
|
||||
4. loads the selected execution profile.
|
||||
5. merges built-in execution defaults, profile values, and request overrides.
|
||||
6. applies request-scoped direct API key values for public Go callers.
|
||||
7. validates endpoint, model, and credential requirements.
|
||||
8. resolves the output contract and JSON Schema document when required.
|
||||
9. reads input artifacts.
|
||||
10. renders prompt messages and hashes the rendered prompt.
|
||||
11. returns a prepared run without calling the LLM.
|
||||
1. validate the prompt ID and load the prompt definition;
|
||||
2. hash the definition and select the explicit or default profile;
|
||||
3. load the profile and resolve effective execution settings;
|
||||
4. validate endpoint, model, and credential availability;
|
||||
5. resolve the output contract and, for JSON Schema output, load a structured
|
||||
schema document before model execution;
|
||||
6. read and hash input artifacts;
|
||||
7. render messages and the session ID; and
|
||||
8. return a `PreparedRun` containing the effective state and rendered-prompt
|
||||
hash.
|
||||
|
||||
Numeric request overrides are presence-aware: omitted values preserve the current effective value, while explicit zero values are real overrides.
|
||||
Execution settings merge defaults, profile values, and a request override.
|
||||
Numeric override presence is retained so explicit zero values are not confused
|
||||
with omissions.
|
||||
|
||||
## Run Flow
|
||||
## Run And Validation Flow
|
||||
|
||||
`Run`:
|
||||
`Run` creates a run ID and timestamps, then calls `Prepare` rather than
|
||||
duplicating preparation. It sends the prepared prompt, effective target,
|
||||
target-presence state, and optional structured-output specification to the LLM
|
||||
client. It converts the returned content to an output artifact, validates it,
|
||||
and returns the artifact, validation, hashes, usage, and timing metadata.
|
||||
|
||||
1. creates a run ID and start timestamp.
|
||||
2. calls `Prepare`.
|
||||
3. calls the injected LLM client with rendered messages, effective target, target presence, and structured-output settings.
|
||||
4. builds the output artifact.
|
||||
5. validates the output.
|
||||
6. optionally attempts bounded repair when a repairer is injected and the contract permits repair.
|
||||
7. returns the run result with artifact, raw output, validation, hashes, selected profile/model metadata, usage, and timing.
|
||||
A validator can return a content result or an operational error. Content
|
||||
failures stay in the result; schema loading, compilation, and validator
|
||||
operational failures are returned as `ErrValidation`. The canonical distinction
|
||||
for callers is documented by the public contracts.
|
||||
|
||||
`Run` must reuse `Prepare`; prepare logic should not be duplicated elsewhere.
|
||||
## Repair Boundary
|
||||
|
||||
## Validation And Repair
|
||||
Repair is an internal optional loop. It starts only when a repairer is present,
|
||||
the output contract permits one or more attempts, validation failed, and the
|
||||
validation mode is JSON or JSON Schema. Each repair receives the previous
|
||||
output, validation errors, effective target, structured-output specification,
|
||||
and attempt metadata; every repaired result is validated again.
|
||||
|
||||
Validation content failures are returned as successful run results with `Validation.Status == failed`. They are not runtime errors.
|
||||
`NewDefaultOutputRepairer` delegates to the injected LLM client. CLI, HTTP, and
|
||||
the public engine use `NewRunner` and therefore do not inject this repairer.
|
||||
|
||||
Validation runtime failures, such as schema load or compile errors, return `ErrValidation`.
|
||||
## Error Translation
|
||||
|
||||
Repair attempts occur only when all conditions are true:
|
||||
|
||||
- a repairer is injected
|
||||
- `repair_attempts` is greater than zero
|
||||
- validation status is `failed`
|
||||
- validation mode is `json` or `json_schema`
|
||||
|
||||
CLI and HTTP wiring call `usecase.NewRunner(...)`, which does not inject a repairer. Normal CLI and HTTP execution therefore does not repair invalid output.
|
||||
|
||||
## Failure Behavior
|
||||
|
||||
Stable runner sentinels include:
|
||||
Runner sentinels identify failure categories for adapters:
|
||||
|
||||
- `ErrInvalidRequest`
|
||||
- `ErrProfileRequired`
|
||||
- `ErrAPIKeyEnvMissing`
|
||||
- `ErrAPIKeyRequired`
|
||||
- `ErrPromptLoad`
|
||||
- `ErrProfileLoad`
|
||||
- `ErrArtifactLoad`
|
||||
- `ErrAPIKeyEnvMissing` and `ErrAPIKeyRequired`
|
||||
- `ErrPromptLoad`, `ErrProfileLoad`, and `ErrArtifactLoad`
|
||||
- `ErrPromptRender`
|
||||
- `ErrLLMGenerate`
|
||||
- `ErrValidation`
|
||||
|
||||
Adapters should use `errors.Is` against sentinels and lower-level repository errors instead of matching message text.
|
||||
Wrap errors with those sentinels and preserve their identities through
|
||||
`errors.Is`; adapters must not classify errors by message text. The runner
|
||||
passes direct keys only to the LLM boundary and never includes resolved key
|
||||
values in prepared or run results.
|
||||
|
||||
Secret values must not appear in prepared output, run results, logs, HTTP responses, or serialized public package results. The effective API-key environment-variable name may appear.
|
||||
## Package-Local Guarantees
|
||||
|
||||
## State And Manifests
|
||||
- `Run` always reuses `Prepare`.
|
||||
- Schema documents are loaded before the initial LLM call when structured output
|
||||
is required.
|
||||
- Output validation records attempts used, including repair attempts.
|
||||
- Runner state is per request; the package does not create a durable run store
|
||||
or manifest.
|
||||
- Source, renderer, validator, and LLM implementations remain injected
|
||||
boundaries.
|
||||
|
||||
The runner is stateless across requests.
|
||||
## Verification And Change Recipe
|
||||
|
||||
- No durable run store.
|
||||
- No manifest files.
|
||||
- No checkpoint, skip, or resume behavior.
|
||||
- Recovery is a new request after correcting inputs, config, or environment.
|
||||
|
||||
## Tests To Inspect
|
||||
Inspect:
|
||||
|
||||
- `internal/usecase/runner_test.go`
|
||||
- `internal/usecase/integration_test.go`
|
||||
- `engine_test.go`
|
||||
- `internal/adapter/cli/run_test.go`
|
||||
- `internal/adapter/http/handler_test.go`
|
||||
|
||||
## Architectural Invariants
|
||||
When changing orchestration:
|
||||
|
||||
- Use-case decisions stay in `internal/usecase`.
|
||||
- `Run` reuses `Prepare`.
|
||||
- Prompt/profile/artifact/schema loading remains behind injected boundaries.
|
||||
- Validation content failures are result state; validation runtime failures are errors.
|
||||
- Repair loops are bounded by `repair_attempts` and repairer presence.
|
||||
- Resolved secret values are never serialized or emitted.
|
||||
1. identify the collaborator boundary and the affected `Prepare` or `Run` state;
|
||||
2. preserve the `Run`-through-`Prepare` path and error identity;
|
||||
3. add focused runner or integration tests for changed state transitions,
|
||||
validation, or repair behavior; and
|
||||
4. update the owning external contract and any affected source or LLM internal
|
||||
document.
|
||||
|
||||
The [testing policy](../policy/testing.md) owns global test sufficiency.
|
||||
|
||||
@@ -2,142 +2,79 @@
|
||||
|
||||
## Purpose
|
||||
|
||||
This document covers implemented prompt, profile, schema, artifact, and catalog source behavior. It is for developers changing loaders or source wiring.
|
||||
This document describes how source packages load prompt definitions, profiles,
|
||||
schemas, and artifacts. The [configuration reference](../config.md) owns their
|
||||
user-facing formats and settings. The [HTTP API reference](../api.md) owns
|
||||
HTTP-visible artifact outcomes; [operations](../operations.md) owns deployment
|
||||
handling.
|
||||
|
||||
Full user-facing YAML and config reference material belongs in `docs/config.md`.
|
||||
## Prompt Definitions
|
||||
|
||||
## Prompt Definition Sources
|
||||
`internal/promptdef` provides filesystem and `fs.FS` repositories. Both use
|
||||
`internal/filecatalog` for recursive YAML discovery, deterministic ordering,
|
||||
display paths, and root cleaning.
|
||||
|
||||
`internal/promptdef` provides directory-backed and `fs.FS` repositories.
|
||||
Repositories select a prompt by YAML ID and optional version rather than by
|
||||
path. They decode through strict YAML handling, reject duplicate matching
|
||||
definitions, and resolve `content_file` relative to the definition. The `fs.FS`
|
||||
implementation resolves content paths inside its source root; absolute paths and
|
||||
traversal outside that root are rejected before file access.
|
||||
|
||||
Behavior:
|
||||
## Profiles And Built-Ins
|
||||
|
||||
- recursively scans `.yaml` and `.yml` files.
|
||||
- decodes YAML with known-fields checking.
|
||||
- looks up prompts by YAML `id`, not by path.
|
||||
- optionally filters by prompt `version`.
|
||||
- rejects duplicate matching prompt IDs.
|
||||
- requires `id`, `version`, and at least one message.
|
||||
- requires each message to set exactly one of `content` or `content_file`.
|
||||
- resolves filesystem `content_file` values relative to the prompt YAML file.
|
||||
- resolves `fs.FS` `content_file` values inside the configured source root.
|
||||
- permits prompt subdirectories only as organization; they are not part of prompt identity.
|
||||
`internal/profile` provides filesystem, `fs.FS`, and overlay repositories.
|
||||
`internal/profile/builtin` exposes embedded assets through the same repository
|
||||
interface.
|
||||
|
||||
For `fs.FS` roots, absolute paths and relative traversal outside the source root are rejected by catalog path helpers.
|
||||
An overlay asks its primary source first. It falls back only when the primary
|
||||
reports `ErrProfileNotFound`; invalid YAML, duplicate IDs, validation failures,
|
||||
and raw-key failures are returned rather than hidden by fallback. This makes a
|
||||
custom ID override a built-in ID while retaining errors in the custom source.
|
||||
|
||||
## Profile Sources
|
||||
The public engine can overlay in-memory profiles ahead of both file-backed and
|
||||
built-in repositories. Profile field definitions, validation ranges, and the
|
||||
built-in catalog remain in the [configuration reference](../config.md).
|
||||
|
||||
`internal/profile` provides directory-backed, `fs.FS`, and overlay repositories. `internal/profile/builtin` embeds built-in profile YAML assets and exposes them through the same repository interface.
|
||||
## Schemas
|
||||
|
||||
Behavior:
|
||||
`internal/validate` supplies `StandardValidator` for filesystem sources and
|
||||
`FSValidator` for `fs.FS` sources. Directory-backed validation loads the named
|
||||
schema path; it does not search directories by basename. `fs.FS` schema paths
|
||||
are cleaned and checked against their configured root, while a single-file
|
||||
source matches its file base name.
|
||||
|
||||
- recursively scans `.yaml` and `.yml` files.
|
||||
- decodes YAML with known-fields checking.
|
||||
- looks up profiles by YAML `id`, not by path.
|
||||
- rejects duplicate IDs inside the same source.
|
||||
- rejects raw `api_key` fields in YAML; file-backed profiles must use `api_key_env`.
|
||||
- validates required `endpoint` and `model` values.
|
||||
- validates numeric profile ranges.
|
||||
The runner requests a schema document before generation when it needs
|
||||
structured output. JSON and schema mismatches in generated content are
|
||||
validation results; source access, decoding, registration, and compilation
|
||||
failures are operational errors.
|
||||
|
||||
Overlay behavior:
|
||||
## Artifacts
|
||||
|
||||
- custom profiles are primary.
|
||||
- built-in profiles are fallback.
|
||||
- fallback occurs only after a primary `ErrProfileNotFound`.
|
||||
- primary validation, YAML, duplicate, and raw-key errors are returned directly.
|
||||
- duplicate IDs across custom and built-in sources are allowed because the custom profile overrides the built-in one.
|
||||
`internal/artifact` composes inline and file readers. The ordinary composite
|
||||
reader used by CLI and the public engine reads file references from the process
|
||||
filesystem. The restricted composite reader used by the HTTP adapter combines
|
||||
inline reading with a rooted file reader and optional byte limit.
|
||||
|
||||
The public Go facade can add in-memory profiles ahead of file-backed and built-in profiles.
|
||||
The rooted reader cleans paths and applies lexical containment without resolving
|
||||
symlinks. It checks relative references against the configured root and accepts
|
||||
absolute references only when they remain inside that lexical root. The OS still
|
||||
follows symlinks after that check. The public containment outcome is documented
|
||||
by the [HTTP API reference](../api.md); deployment permissions belong in
|
||||
[operations](../operations.md).
|
||||
|
||||
## Schema Sources
|
||||
## Failure Boundaries
|
||||
|
||||
`internal/validate` provides:
|
||||
Source packages report repository, decoding, duplicate, validation, and read
|
||||
failures to their callers. They do not select public status codes or response
|
||||
schemas. The runner wraps source failures with use-case categories; adapters map
|
||||
them to their own external contract.
|
||||
|
||||
- `StandardValidator` for filesystem paths.
|
||||
- `FSValidator` for `fs.FS` roots and single-file public schema sources.
|
||||
Source reads use current filesystem or `fs.FS` content for each request. These
|
||||
packages create no manifests, checkpoints, or durable run state.
|
||||
|
||||
Behavior:
|
||||
## Verification And Change Recipe
|
||||
|
||||
- `json_schema` validation requires a non-empty `schema_path`.
|
||||
- filesystem schema paths resolve relative to `schema_dir` unless absolute.
|
||||
- directory-backed schema lookup uses the explicit `schema_path`; it does not search recursively by basename.
|
||||
- `fs.FS` schema paths must remain inside the configured source root.
|
||||
- single-file schema sources match by the configured file base name.
|
||||
- schema documents are loaded before the LLM call for structured output.
|
||||
- JSON parse failures are validation content failures.
|
||||
- schema access, decode, registration, and compile failures are runtime validation errors.
|
||||
|
||||
## Artifact Sources
|
||||
|
||||
`internal/artifact` supports two input artifact reference types:
|
||||
|
||||
- `inline`
|
||||
- `file`
|
||||
|
||||
Inline behavior:
|
||||
|
||||
- requires a non-empty body.
|
||||
- produces text/plain artifacts.
|
||||
- hashes the body bytes.
|
||||
|
||||
Direct file behavior:
|
||||
|
||||
- used by CLI `run`, CLI `render`, and the public Go facade.
|
||||
- requires a non-empty URI.
|
||||
- reads from the process filesystem without HTTP artifact-root restrictions.
|
||||
- infers content type from file extension, defaulting to text/plain.
|
||||
|
||||
Restricted file behavior:
|
||||
|
||||
- used by HTTP `serve`.
|
||||
- allows inline artifacts even when no artifact root is configured.
|
||||
- denies file artifacts when no artifact root is configured.
|
||||
- resolves relative file URIs against `server.artifact_root`.
|
||||
- accepts absolute file URIs only when they pass containment checks.
|
||||
- applies `server.max_artifact_bytes` when configured.
|
||||
|
||||
Restricted containment is lexical. It cleans paths and checks the relative path against the configured root; it does not resolve symlinks. Symlinks inside the root are followed by the operating system, including symlinks that target files outside the root.
|
||||
|
||||
## Catalog Helpers
|
||||
|
||||
`internal/filecatalog` centralizes shared source helpers:
|
||||
|
||||
- recursive YAML discovery for filesystem and `fs.FS` roots.
|
||||
- deterministic sorting.
|
||||
- `.yaml` and `.yml` filtering.
|
||||
- display paths for diagnostics.
|
||||
- YAML file stems.
|
||||
- `fs.FS` root cleaning and containment checks.
|
||||
|
||||
Repository code should use these helpers instead of reimplementing path traversal and containment rules.
|
||||
|
||||
## Failure Behavior
|
||||
|
||||
Common source failures:
|
||||
|
||||
- missing prompt/profile/schema/artifact files.
|
||||
- invalid YAML or JSON.
|
||||
- unknown YAML fields.
|
||||
- duplicate prompt or profile IDs.
|
||||
- prompt/profile validation errors.
|
||||
- raw API key fields in profile YAML.
|
||||
- unsupported artifact reference type.
|
||||
- missing inline body or file URI.
|
||||
- artifact outside HTTP root.
|
||||
- artifact exceeding HTTP size limit.
|
||||
- schema load or compile failure.
|
||||
|
||||
Prompt/profile repository lookup errors are mapped by adapters separately from runtime runner errors. Validation content failures remain result state; source and schema runtime failures return errors.
|
||||
|
||||
## State And Manifests
|
||||
|
||||
Source packages do not persist run state.
|
||||
|
||||
- No manifests are read or written.
|
||||
- No source package implements skip or resume behavior.
|
||||
- Source reads reflect the current filesystem or `fs.FS` state for each request.
|
||||
|
||||
## Tests To Inspect
|
||||
Inspect:
|
||||
|
||||
- `internal/promptdef/repository_test.go`
|
||||
- `internal/profile/repository_test.go`
|
||||
@@ -147,11 +84,14 @@ Source packages do not persist run state.
|
||||
- `internal/usecase/integration_test.go`
|
||||
- `engine_test.go`
|
||||
|
||||
## Architectural Invariants
|
||||
When updating prompt, profile, schema, or built-in assets:
|
||||
|
||||
- Prompt/profile identity comes from YAML `id`.
|
||||
- External YAML decoding remains strict.
|
||||
- File-backed profile YAML never accepts raw API key values.
|
||||
- Built-in profiles are fallback, not a replacement for custom source validation.
|
||||
- HTTP file artifacts remain rooted by lexical containment.
|
||||
- Schema runtime failures remain errors, while JSON/schema content mismatches remain validation results.
|
||||
1. keep assets valid for the strict loader and the relevant source boundary;
|
||||
2. update the [configuration reference](../config.md) when a file-format,
|
||||
catalog, or default changes;
|
||||
3. run focused source and integration tests, including the built-in repository
|
||||
test when embedded assets change; and
|
||||
4. update this document when discovery, precedence, containment, or failure
|
||||
mechanics change.
|
||||
|
||||
The [testing policy](../policy/testing.md) owns global test sufficiency.
|
||||
|
||||
@@ -1,164 +1,157 @@
|
||||
# Operations Guide
|
||||
|
||||
## Scope
|
||||
## Scope And References
|
||||
|
||||
This guide covers operating the implemented CLI commands and HTTP service. It
|
||||
does not replace the [CLI reference](cli.md), [Configuration reference](config.md),
|
||||
or [HTTP API reference](api.md).
|
||||
This runbook covers deployment, normal operation, capacity planning, and safe
|
||||
recovery for Scriptorium. It does not redefine invocation syntax, configuration
|
||||
fields, or HTTP wire behavior.
|
||||
|
||||
## Operational Model
|
||||
- [CLI reference](cli.md): commands, output destinations, and exit codes.
|
||||
- [Configuration reference](config.md): configuration, prompt/profile/schema
|
||||
formats, defaults, and credentials.
|
||||
- [HTTP API reference](api.md): route, request/response schema, status codes,
|
||||
limits, and HTTP artifact access.
|
||||
- [Consumer integration overview](consumers/api.md): caller responsibilities.
|
||||
|
||||
Scriptorium executes one prompt request per CLI invocation or HTTP request.
|
||||
## Operational Model And State
|
||||
|
||||
Important boundaries:
|
||||
Scriptorium handles one prompt request for each CLI invocation or HTTP request.
|
||||
It has no durable run store, archive, checkpoint, cache, or resume mechanism.
|
||||
A failed or interrupted request is recovered by correcting its inputs,
|
||||
configuration, or environment and submitting a new request.
|
||||
|
||||
- No durable run state is stored.
|
||||
- No manifest, archive, checkpoint, or built-in backup workflow is written.
|
||||
- No built-in resume behavior exists.
|
||||
- Recovery is rerun-based: correct inputs, config, or environment, then run again.
|
||||
Generated artifacts, rendered prompts, model output, and run metadata are
|
||||
caller-owned data. Retention, encryption, backup, and deletion are deployment
|
||||
responsibilities.
|
||||
|
||||
## Filesystem Layout
|
||||
## Deploy The Filesystem And Process
|
||||
|
||||
Operational deployments usually provide:
|
||||
Provide the process with readable prompt, profile, and schema sources. Keep
|
||||
prompt templates adjacent to the prompt definitions that reference them. For an
|
||||
HTTP deployment that accepts file artifacts, use a dedicated, narrow artifact
|
||||
directory rather than a general-purpose or sensitive filesystem tree.
|
||||
|
||||
- `prompt_dir`: prompt definition YAML files and adjacent `content_file` templates.
|
||||
- `profile_dir`: optional custom profile YAML files.
|
||||
- `schema_dir`: optional JSON Schema files.
|
||||
- `server.artifact_root`: optional HTTP file-input root for `serve`.
|
||||
Run Scriptorium under an identity that can:
|
||||
|
||||
Keep these directories readable by the Scriptorium process. Keep
|
||||
`server.artifact_root` narrow and not writable by untrusted users.
|
||||
- read only the prompt, profile, schema, and allowed input-artifact paths it
|
||||
needs;
|
||||
- read the required credential environment variables without writing them to
|
||||
files or logs; and
|
||||
- write only caller-selected output locations when CLI output files are used.
|
||||
|
||||
## Normal CLI Workflow
|
||||
Do not make the HTTP artifact directory writable by untrusted users. The HTTP
|
||||
artifact containment behavior is lexical and the operating system follows
|
||||
symlinks; account for that when choosing ownership and mount boundaries. See
|
||||
the [HTTP API reference](api.md) for the externally observable behavior.
|
||||
|
||||
Use `render` before `run` when changing prompt/profile/input wiring:
|
||||
## Supply Credentials And Protect Runtime Data
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium render \
|
||||
--config ./examples/config.yml \
|
||||
--prompt generic.markdown_summary \
|
||||
--input transcript=./examples/fixtures/transcript.md \
|
||||
--input glossary=./examples/fixtures/glossary.yml \
|
||||
--format json
|
||||
```
|
||||
Set secret values in the process environment and configure only their
|
||||
environment-variable names. Do not put raw keys in configuration, prompt or
|
||||
profile files, process arguments, HTTP payloads, captured command lines, or
|
||||
debug dumps.
|
||||
|
||||
Use `run` for generation after preflight:
|
||||
Treat stdout, stderr, prepared-run output, generated artifacts, and HTTP
|
||||
responses as potentially sensitive. Send service logs to a controlled collector
|
||||
and apply the same retention and access rules as for model input and output.
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium run \
|
||||
--config ./examples/config.yml \
|
||||
--prompt generic.markdown_summary \
|
||||
--input transcript=./examples/fixtures/transcript.md \
|
||||
--input glossary=./examples/fixtures/glossary.yml \
|
||||
--out ./summary.md
|
||||
```
|
||||
## Run A Normal Workflow
|
||||
|
||||
Before production runs, confirm:
|
||||
Before changing production inputs, profiles, or schemas:
|
||||
|
||||
- the effective config path is the intended one;
|
||||
- prompt/profile/schema directories are readable;
|
||||
- input file paths exist and match prompt input names;
|
||||
- required API-key environment variables are set;
|
||||
- the selected model endpoint is reachable from the process environment.
|
||||
1. confirm the deployed configuration selects the intended sources and model
|
||||
credentials;
|
||||
2. use [`render`](cli.md) with the same request inputs and variables to confirm
|
||||
preparation without a model call;
|
||||
3. use [`run`](cli.md) for generation; and
|
||||
4. retain or discard validation-failed output according to the caller's
|
||||
policy.
|
||||
|
||||
## HTTP Service Operation
|
||||
The [maintained render script](../examples/render-markdown-summary.sh) is a
|
||||
copyable preflight example. The CLI reference owns its complete invocation and
|
||||
exit semantics.
|
||||
|
||||
Start the service with:
|
||||
## Expose The HTTP Service
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium serve --config ./examples/config.yml
|
||||
```
|
||||
The HTTP service has no built-in authentication or authorization. Place it on a
|
||||
trusted network or behind an authenticated reverse proxy, API gateway, or
|
||||
equivalent access control. Restrict who can reach it and who can read the
|
||||
artifact root.
|
||||
|
||||
The implemented HTTP route is `POST /v1/runs`; request and response fields are
|
||||
defined in the [HTTP API reference](api.md).
|
||||
Use a service manager or supervisor appropriate to the deployment to manage
|
||||
process lifetime, restart policy, log capture, and environment injection. The
|
||||
[HTTP API reference](api.md) owns client request shapes, status behavior, and
|
||||
artifact-access outcomes.
|
||||
|
||||
The maintained HTTP request-shape example is `examples/http-run.json`.
|
||||
## Plan Capacity And Limits
|
||||
|
||||
HTTP service notes:
|
||||
Capacity is primarily determined by concurrent model calls, input and output
|
||||
sizes, schema complexity, provider latency, and network behavior. Size limits
|
||||
protect request bodies, HTTP file artifacts, and encoded responses; configure
|
||||
them through the [configuration reference](config.md) and rely on the
|
||||
[HTTP API reference](api.md) for their response effects.
|
||||
|
||||
- Unknown JSON fields are rejected.
|
||||
- `inline` input references work without an artifact root.
|
||||
- `file` input references require `server.artifact_root` or `serve --artifact-root`.
|
||||
- Request bodies, HTTP file input artifacts, and encoded JSON responses are size-limited.
|
||||
- Validation content failures return `200 OK` with `validation.status: "failed"`.
|
||||
Before increasing a limit:
|
||||
|
||||
Security boundary:
|
||||
1. measure representative input, generated-output, and optional raw-output
|
||||
sizes;
|
||||
2. confirm memory, network, and upstream-provider capacity;
|
||||
3. retain an upstream request-size and authentication boundary; and
|
||||
4. test the intended workload in a non-production environment.
|
||||
|
||||
- `serve` has no built-in authentication or authorization.
|
||||
- Put it behind trusted controls such as a private network, authenticated reverse proxy, or API gateway.
|
||||
- Do not expose an artifact root containing unrelated sensitive files.
|
||||
- Symlinks inside the artifact root are followed by the operating system.
|
||||
For large local inputs, prefer a controlled file-artifact directory over
|
||||
placing arbitrary paths on the service host. Avoid disabling a limit unless an
|
||||
equivalent trusted control exists elsewhere.
|
||||
|
||||
## Secrets Handling
|
||||
## Diagnose And Recover
|
||||
|
||||
Raw API keys are not accepted in app config, profiles, CLI flags, or HTTP
|
||||
request bodies.
|
||||
### Preparation Or Configuration Failure
|
||||
|
||||
Use this pattern:
|
||||
Capture the CLI diagnostic or HTTP error response, then verify the selected
|
||||
configuration, prompt ID, profile selection, source readability, and input
|
||||
mapping. Use `render` with the same request when it is unclear whether failure
|
||||
occurs before model execution. Consult the [CLI reference](cli.md), the
|
||||
[configuration reference](config.md), and the [HTTP API reference](api.md) for
|
||||
the exact interface contract.
|
||||
|
||||
1. Set an environment variable containing the secret value.
|
||||
2. Store only the variable name in profile `api_key_env` or request override `api_key_env`.
|
||||
3. Scope the process environment to the minimum required variables.
|
||||
### Credential Or Provider Failure
|
||||
|
||||
## Output, Logs, And Exit Codes
|
||||
Confirm that the process environment contains the configured credential name
|
||||
without printing the secret. Check endpoint reachability and provider health
|
||||
from the process network. If preparation succeeds but generation fails, inspect
|
||||
the selected model settings in prepared output and the service's controlled
|
||||
logs. Correct the deployment or provider issue, then submit a new request.
|
||||
|
||||
`run`:
|
||||
### Artifact Or Permission Failure
|
||||
|
||||
- stdout: generated artifact body unless `--out` is used.
|
||||
- stderr: summary on success, errors on failure.
|
||||
- exit `2`: generation completed and output was written, but validation failed.
|
||||
Verify that the process can read the intended local input. For HTTP file
|
||||
artifacts, verify the deployment's artifact root, ownership, path layout, and
|
||||
file size. Do not widen filesystem permissions or the allowed root merely to
|
||||
make an arbitrary path work; move or copy the required artifact into the
|
||||
controlled location instead.
|
||||
|
||||
`render`:
|
||||
### Validation Failure
|
||||
|
||||
- stdout: prepared-run output unless `--out` is used.
|
||||
- stderr: errors.
|
||||
- exit `0` on success, `1` on failure.
|
||||
A generated-content validation failure is distinct from a runtime failure.
|
||||
CLI `run` reports the validation result and error count in its success summary;
|
||||
it does not print the individual validation messages. For HTTP, inspect the
|
||||
validation object in the response according to the [HTTP API reference](api.md).
|
||||
|
||||
`serve`:
|
||||
Use rendered input and generated output to determine whether prompt instructions,
|
||||
the selected model, or the schema needs correction. If schema loading or
|
||||
compilation itself fails, correct the source deployment or schema document
|
||||
before rerunning.
|
||||
|
||||
- stderr: startup and server errors.
|
||||
- HTTP response body: JSON success or error envelope.
|
||||
### HTTP Limit Or Request Failure
|
||||
|
||||
## Validation Behavior
|
||||
Compare the request, artifact, or expected response size with the deployed
|
||||
configuration, and validate the request against the [HTTP API reference](api.md).
|
||||
Reduce the payload, use an appropriate controlled artifact source, omit
|
||||
unneeded raw output, or adjust the deployment limit after capacity review.
|
||||
|
||||
Prompt `output.validation_mode` controls validation:
|
||||
## Cleanup And Reruns
|
||||
|
||||
- `none`: skipped.
|
||||
- `basic`: output body must not be empty.
|
||||
- `json`: output body must parse as JSON.
|
||||
- `json_schema`: output body must parse as JSON and satisfy the configured schema.
|
||||
|
||||
Runtime/schema failures are hard failures (`run` exit `1`, HTTP error).
|
||||
Generated-content validation failures are soft failures (`run` exit `2`, HTTP
|
||||
`200 OK` with failed validation status).
|
||||
|
||||
## Size Limits
|
||||
|
||||
Defaults are documented in [Configuration reference](config.md). Operationally:
|
||||
|
||||
- Keep default HTTP limits unless larger payloads are measured and expected.
|
||||
- Prefer `inline` HTTP inputs for small payloads.
|
||||
- Prefer `file` HTTP inputs for larger local artifacts under a controlled artifact root.
|
||||
- Increase `server.max_response_bytes` when generated artifacts or requested raw output are expected to be large.
|
||||
- Use `0` only when another trusted layer enforces size limits.
|
||||
|
||||
## Maintained Examples
|
||||
|
||||
- `examples/config.yml`
|
||||
- `examples/config.full.yml`
|
||||
- `examples/render-markdown-summary.sh`
|
||||
- `examples/http-run.json`
|
||||
|
||||
## Safe Recovery
|
||||
|
||||
For failed CLI commands or HTTP requests:
|
||||
|
||||
1. Capture stderr or the HTTP error `code` and `message`.
|
||||
2. Confirm config path and effective directory settings.
|
||||
3. Verify prompt ID, profile ID, schema path, and input mappings.
|
||||
4. Verify required API-key environment variables.
|
||||
5. Reproduce with `render --format json` when pre-LLM resolution is uncertain.
|
||||
6. Rerun after correction.
|
||||
|
||||
Because Scriptorium does not persist run state, rerun is the supported recovery
|
||||
path.
|
||||
Because no run state is retained, cleanup concerns caller-owned output files,
|
||||
logs, and artifacts only. Remove or rotate them using the deployment's normal
|
||||
retention policy. After a correction, rerun the request from the beginning;
|
||||
there is no safe resume point.
|
||||
|
||||
@@ -4,14 +4,12 @@ This document is the development architecture policy for Scriptorium.
|
||||
|
||||
It is for developers and LLM coding agents. User-facing behavior belongs in `README.md` and the docs under `docs/` that target operators/users.
|
||||
|
||||
## Project Shape
|
||||
## System Shape
|
||||
|
||||
Scriptorium is a narrow prompt-execution application with three entry paths:
|
||||
|
||||
- CLI `run`
|
||||
- CLI `render`
|
||||
- HTTP `POST /v1/runs` through `serve`
|
||||
- public Go package `gitea.maximumdirect.net/eric/scriptorium`
|
||||
Scriptorium is a narrow prompt-execution application with three executable
|
||||
entry paths: CLI `run`, CLI `render`, and the HTTP service started by `serve`.
|
||||
It also provides a public Go package for in-process use. Its current component
|
||||
inventory is maintained in the [internal overview](../internal/overview.md).
|
||||
|
||||
Domain behavior is centralized in `internal/usecase` and `internal/domain`.
|
||||
|
||||
@@ -23,45 +21,17 @@ Domain behavior is centralized in `internal/usecase` and `internal/domain`.
|
||||
- Keep config strict: YAML/JSON decoding for external inputs should reject unknown fields.
|
||||
- Keep secrets out of payloads: raw API key values must not be accepted or emitted.
|
||||
|
||||
## Package Boundaries
|
||||
## Dependency Direction
|
||||
|
||||
Current package map:
|
||||
|
||||
- root package `scriptorium`: public Go facade over engine construction, source options, request/result types, and error mapping.
|
||||
- `cmd/scriptorium`: process entrypoint.
|
||||
- `internal/adapter/cli`: command parsing, app wiring for CLI commands, output behavior.
|
||||
- `internal/adapter/http`: HTTP DTO mapping and error/status mapping.
|
||||
- `internal/config`: application settings loading and CLI override precedence.
|
||||
- `internal/defaults`: compile-time default constants.
|
||||
- `internal/domain`: core request/result and contract types.
|
||||
- `internal/usecase`: `Runner` prepare/run orchestration and repair-hook boundary.
|
||||
- `internal/promptdef`: filesystem prompt-definition repository.
|
||||
- `internal/profile`: filesystem, `fs.FS`, and overlay execution-profile repositories.
|
||||
- `internal/profile/builtin`: embedded built-in execution profiles.
|
||||
- `internal/filecatalog`: shared YAML discovery and `fs.FS` source helpers.
|
||||
- `internal/artifact`: artifact reference readers.
|
||||
- `internal/prompt`: template renderer.
|
||||
- `internal/llm`: provider-neutral LLM client interface and OpenAI-compatible implementation.
|
||||
- `internal/validate`: validator interfaces and standard implementation.
|
||||
- `internal/format`: prepared-run output formatting.
|
||||
|
||||
Detailed component behavior is documented in:
|
||||
|
||||
- `docs/internal/runner.md`
|
||||
- `docs/internal/adapters.md`
|
||||
- `docs/internal/sources.md`
|
||||
|
||||
## Configuration And Precedence
|
||||
|
||||
Application settings are resolved as:
|
||||
|
||||
1. built-in defaults
|
||||
2. config file values
|
||||
3. CLI overrides
|
||||
|
||||
`config.yml` is for application wiring (directories, server address, render default format), not prompt/profile runtime execution settings.
|
||||
|
||||
Profile selection and runtime model resolution remain use-case concerns.
|
||||
- Adapters translate external shapes and IO concerns; they do not make
|
||||
use-case decisions.
|
||||
- Use-case and domain code depend on explicit repository, renderer, validator,
|
||||
and LLM interfaces rather than adapter implementations.
|
||||
- Source, rendering, validation, and LLM implementations remain behind their
|
||||
package boundaries.
|
||||
- Dependency-specific types must not leak across unrelated package boundaries.
|
||||
- Prefer the standard library; add an external dependency only when it
|
||||
materially reduces risk or complexity.
|
||||
|
||||
## State And Persistence Policy
|
||||
|
||||
@@ -70,44 +40,32 @@ Scriptorium has no durable run-state store.
|
||||
- No built-in resume/checkpoint/archive behavior.
|
||||
- Recovery model is rerun after correcting inputs/config/environment.
|
||||
|
||||
## External Integration Policy
|
||||
## Contract Ownership
|
||||
|
||||
Current external contracts:
|
||||
|
||||
- inbound HTTP contract: `POST /v1/runs`, documented canonically in `docs/api.md`
|
||||
- outbound model contract: OpenAI-compatible chat completions subset
|
||||
- subprocess contract for integrators: CLI `run`/`render`
|
||||
- public Go package contract: `docs/consumers/pkg-scriptorium.md`
|
||||
|
||||
Integration docs belong under `docs/integrations/`.
|
||||
The [CLI](../cli.md), [configuration](../config.md), [HTTP API](../api.md),
|
||||
[public Go package](../consumers/pkg-scriptorium.md), and
|
||||
[integration](../integrations/) documents own their respective external
|
||||
contracts. This policy keeps only the architectural boundaries that govern
|
||||
their implementation.
|
||||
|
||||
## Error Handling And Logging
|
||||
|
||||
- Wrap errors with domain/operation context.
|
||||
- Map domain errors to adapter-appropriate statuses/codes without leaking sensitive internals.
|
||||
- Keep stderr summaries concise for CLI success/error paths.
|
||||
- Never emit raw secret values.
|
||||
|
||||
## Testing Expectations
|
||||
## Testing And Documentation
|
||||
|
||||
- Core runner behavior should be covered with isolated unit tests and fixture-based integration tests.
|
||||
- Adapter behavior should be tested for parse/mapping/error semantics.
|
||||
- Config parsing, prompt/profile loading, validator behavior, and LLM client error handling should remain covered by package tests.
|
||||
- Repository-level docs/examples that claim runnable behavior should be validated by tests or smoke commands.
|
||||
|
||||
## Documentation Expectations
|
||||
|
||||
- Document implemented behavior only outside `docs/roadmap/`.
|
||||
- Keep canonical reference locations stable (`docs/cli.md`, `docs/config.md`, `docs/operations.md`, `docs/troubleshooting.md`, `docs/internal/`).
|
||||
- Update docs in the same change when architecture-relevant behavior changes.
|
||||
Testing philosophy and change-validation expectations are defined by the
|
||||
[testing policy](testing.md). Documentation ownership and maintenance rules are
|
||||
defined by the [documentation policy](documentation.md).
|
||||
|
||||
## Architectural Invariants
|
||||
|
||||
- `Runner.Run` reuses `Runner.Prepare` flow.
|
||||
- CLI and HTTP currently instantiate `Runner` without a repairer.
|
||||
- Artifact reading supports `inline` and `file` references.
|
||||
- Unknown input fields in config/prompt/profile/http JSON should be rejected by strict decoding.
|
||||
- Raw API key values must not be accepted through config/HTTP payloads.
|
||||
- Raw API key values must not be accepted through external configuration or
|
||||
request payloads, and resolved secret values must not be emitted.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
|
||||
@@ -2,9 +2,9 @@
|
||||
|
||||
## Status
|
||||
|
||||
Proposed implementation plan. This document records the findings of the
|
||||
documentation audit performed after adoption of the canonical-ownership policy.
|
||||
The revisions described here are not yet implemented.
|
||||
Completed on 2026-07-26. This document records the findings of the
|
||||
documentation audit performed after adoption of the canonical-ownership policy
|
||||
and the completed refresh that addressed them.
|
||||
|
||||
## Objective
|
||||
|
||||
@@ -486,7 +486,7 @@ owning references, and not duplicated as complete files in prose docs.
|
||||
`docs/roadmap/migration.md`.
|
||||
|
||||
**Gate:** Code, tests, examples, contracts, internal documentation, operations,
|
||||
policies, and roadmap status agree.
|
||||
policies, and roadmap status agree. Completed on 2026-07-26.
|
||||
|
||||
## Completion Criteria
|
||||
|
||||
|
||||
@@ -93,6 +93,10 @@ At minimum:
|
||||
refresh and policy updates are merged and the repository has an agreed,
|
||||
accurate baseline.
|
||||
|
||||
**Gate status:** Complete as of 2026-07-26. The completed documentation
|
||||
refresh and its verification record are in the
|
||||
[documentation compliance roadmap](documentation.md).
|
||||
|
||||
### Step 2: Record The Architectural Decision And Detailed Boundary
|
||||
|
||||
Create an ADR, under the policy established in Step 1, that records:
|
||||
|
||||
@@ -1,361 +0,0 @@
|
||||
# Troubleshooting
|
||||
|
||||
This guide lists common implemented failure modes and safe fixes.
|
||||
|
||||
Canonical references:
|
||||
|
||||
- [CLI reference](cli.md)
|
||||
- [Configuration reference](config.md)
|
||||
- [HTTP API reference](api.md)
|
||||
- [Operations guide](operations.md)
|
||||
|
||||
## Missing Or Invalid Config
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI error includes `application config error`, `config file not found`, `invalid config YAML`, or `invalid config`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- `--config` points to a missing file.
|
||||
- YAML syntax is invalid.
|
||||
- Config contains unknown fields or negative HTTP size limits.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium render --config /path/to/config.yml --prompt generic.markdown_summary --input transcript=./examples/fixtures/transcript.md --input glossary=./examples/fixtures/glossary.yml
|
||||
```
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Correct the config path.
|
||||
- Fix YAML syntax.
|
||||
- Remove unknown fields.
|
||||
- Keep raw secrets out of config.
|
||||
|
||||
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
|
||||
|
||||
## Missing Prompt Directory
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI parse error says the prompt directory is required.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Neither config nor CLI flags provide an effective `prompt_dir`.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Re-run once with explicit `--prompt-dir`.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Set `prompt_dir` in config or pass `--prompt-dir`.
|
||||
|
||||
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
|
||||
|
||||
## Unknown Flags
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI parse error for an unknown flag.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Typo.
|
||||
- Flag is valid for another command.
|
||||
- `serve` was given runtime model override flags.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Compare the command with the command-specific flag list.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Remove unsupported flags.
|
||||
- Use `run` or `render` for runtime model overrides.
|
||||
|
||||
Relevant links: [CLI reference](cli.md)
|
||||
|
||||
## Prompt Load Failures
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI run/render fails during prompt loading.
|
||||
- HTTP returns `404 prompt_not_found` or `400 prompt_load_failed`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Prompt ID/version does not exist.
|
||||
- Prompt YAML is invalid or has unknown fields.
|
||||
- Prompt contract is invalid, such as missing messages, invalid output mode, bad `content_file`, or missing `schema_path` for `json_schema`.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium render --config ./examples/config.yml --prompt <prompt-id> --input transcript=./examples/fixtures/transcript.md --input glossary=./examples/fixtures/glossary.yml --format json
|
||||
```
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Correct prompt ID/version.
|
||||
- Fix prompt YAML and referenced `content_file` paths.
|
||||
- Fix output contract fields.
|
||||
|
||||
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
|
||||
|
||||
## Profile Load Failures
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI run/render fails during profile loading.
|
||||
- HTTP returns `404 profile_not_found`, `400 profile_load_failed`, or `400 profile_required`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Profile ID does not exist.
|
||||
- Request omitted profile and prompt has no `default_profile`.
|
||||
- Profile YAML is invalid or has unknown fields.
|
||||
- Profile contains raw `api_key`.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
```bash
|
||||
go run ./cmd/scriptorium render --config ./examples/config.yml --prompt generic.markdown_summary --profile <profile-id> --input transcript=./examples/fixtures/transcript.md --input glossary=./examples/fixtures/glossary.yml
|
||||
```
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Correct profile ID or prompt `default_profile`.
|
||||
- Fix profile YAML and value ranges.
|
||||
- Replace raw `api_key` with `api_key_env`.
|
||||
|
||||
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
|
||||
|
||||
## Input Artifact Failures
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI run/render fails while reading inputs.
|
||||
- HTTP returns `400 artifact_read_failed`, `400 artifact_not_allowed`, or `413 artifact_too_large`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Input file path is missing or unreadable.
|
||||
- HTTP input type is unsupported or missing required fields.
|
||||
- HTTP file refs are disabled because no artifact root is configured.
|
||||
- HTTP file path is lexically outside the artifact root.
|
||||
- HTTP file input exceeds `server.max_artifact_bytes`.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Verify each input path exists and is readable by the process.
|
||||
- For HTTP, verify input refs use `file` or `inline`.
|
||||
- For HTTP file refs, verify the artifact root and compare file size to `server.max_artifact_bytes`.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Correct paths and permissions.
|
||||
- Configure a narrow artifact root for HTTP file refs.
|
||||
- Use relative paths under the artifact root or switch to `inline`.
|
||||
- Increase `server.max_artifact_bytes` only for expected larger inputs.
|
||||
|
||||
Relevant links: [HTTP API reference](api.md), [Configuration reference](config.md)
|
||||
|
||||
## Missing API-Key Environment Variable
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI render/run fails with an API-key environment error.
|
||||
- HTTP returns `400 api_key_env_missing`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Selected profile or runtime override sets `api_key_env`, but the environment variable is unset or empty.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
```bash
|
||||
printenv SCRIPTORIUM_API_KEY
|
||||
```
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Set the required environment variable before starting the CLI command or HTTP service.
|
||||
- Or use a profile that does not require provider API-key auth.
|
||||
|
||||
Relevant links: [Configuration reference](config.md), [Operations guide](operations.md)
|
||||
|
||||
## Prompt Template Render Failures
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI render/run fails during prompt rendering.
|
||||
- HTTP returns `400 prompt_render_failed`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Template references an input that was not supplied.
|
||||
- Template syntax or variable reference is invalid.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Run `render --format json` with the same prompt, inputs, vars, and profile.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Align `{{input "name"}}` references with request input names.
|
||||
- Fix template syntax and variable names.
|
||||
|
||||
Relevant links: [Configuration reference](config.md), [CLI reference](cli.md)
|
||||
|
||||
## LLM Request Failures
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI `run` fails during generation.
|
||||
- HTTP returns `502 llm_failed`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Endpoint is unreachable.
|
||||
- Provider returns non-2xx.
|
||||
- Request times out.
|
||||
- Provider response is malformed.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Run `render` first to confirm pre-LLM preparation works.
|
||||
- Check selected endpoint/model in prepared output.
|
||||
- Check network/provider logs for timeout or non-2xx details.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Correct endpoint/model/profile settings.
|
||||
- Adjust timeout when appropriate.
|
||||
- Resolve provider or network issue.
|
||||
|
||||
Relevant links: [Operations guide](operations.md), [Configuration reference](config.md)
|
||||
|
||||
## Validation Failed
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI `run` exits `2`.
|
||||
- HTTP returns `200 OK` with `validation.status` set to `failed`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Generated output failed `basic`, `json`, or `json_schema` content validation.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Inspect validation errors in CLI stderr or the HTTP response.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Refine prompt instructions.
|
||||
- Adjust schema or model/profile settings.
|
||||
- Rerun after correction.
|
||||
|
||||
Relevant links: [Operations guide](operations.md), [HTTP API reference](api.md)
|
||||
|
||||
## Validation Runtime Failure
|
||||
|
||||
Symptom:
|
||||
|
||||
- CLI `run` fails with validation runtime error.
|
||||
- HTTP returns `500 validation_runtime_failed`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- `json_schema` schema file is missing or unreadable.
|
||||
- Schema JSON is invalid.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Verify `schema_dir` and prompt `output.schema_path`.
|
||||
- Check schema file readability and JSON syntax.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Correct schema path or permissions.
|
||||
- Fix schema JSON.
|
||||
- Rerun.
|
||||
|
||||
Relevant links: [Configuration reference](config.md), [Operations guide](operations.md)
|
||||
|
||||
## HTTP JSON Or Request Contract Errors
|
||||
|
||||
Symptom:
|
||||
|
||||
- HTTP returns `400 invalid_json` or `400 invalid_request`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- JSON body is malformed.
|
||||
- Request has unknown fields or trailing JSON tokens.
|
||||
- Required `prompt_id` or `inputs` is missing.
|
||||
- Runtime override values are out of range.
|
||||
- `extra_params` collides with reserved outbound fields.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Revalidate request JSON and compare fields with the API reference.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Send one JSON object with only supported fields.
|
||||
- Include `prompt_id` and at least one input.
|
||||
- Use valid model override ranges.
|
||||
- Remove reserved `extra_params` keys.
|
||||
|
||||
Relevant links: [HTTP API reference](api.md)
|
||||
|
||||
## HTTP Size Limit Errors
|
||||
|
||||
Symptom:
|
||||
|
||||
- HTTP returns `413 request_too_large`, `413 artifact_too_large`, or `413 response_too_large`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- JSON request body exceeds `server.max_request_bytes`.
|
||||
- HTTP file input exceeds `server.max_artifact_bytes`.
|
||||
- Encoded JSON response exceeds `server.max_response_bytes`.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Compare request, file input, and expected response sizes with configured limits.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Use smaller inline inputs or switch to file inputs under the artifact root.
|
||||
- Reduce generated output size.
|
||||
- Omit `include_raw_output`.
|
||||
- Increase limits only when the deployment expects larger payloads.
|
||||
|
||||
Relevant links: [HTTP API reference](api.md), [Operations guide](operations.md)
|
||||
|
||||
## HTTP Route Or Method Errors
|
||||
|
||||
Symptom:
|
||||
|
||||
- HTTP returns `404 not_found` or `405 method_not_allowed`.
|
||||
|
||||
Likely cause:
|
||||
|
||||
- Path is not `/v1/runs`.
|
||||
- Method on `/v1/runs` is not `POST`.
|
||||
|
||||
Diagnostic step:
|
||||
|
||||
- Check the request URL and method.
|
||||
|
||||
Safe fix:
|
||||
|
||||
- Send `POST /v1/runs`.
|
||||
|
||||
Relevant links: [HTTP API reference](api.md)
|
||||
@@ -36,6 +36,40 @@ func (f *fakeRunner) Run(ctx context.Context, req domain.RunRequest) (*domain.Ru
|
||||
return f.result, nil
|
||||
}
|
||||
|
||||
func TestMaintainedHTTPRunExampleMatchesRequestContract(t *testing.T) {
|
||||
body, err := os.ReadFile(filepath.Join("..", "..", "..", "examples", "http-run.json"))
|
||||
if err != nil {
|
||||
t.Fatalf("read maintained HTTP request example: %v", err)
|
||||
}
|
||||
|
||||
runner := &fakeRunner{result: &domain.RunResult{}}
|
||||
h := NewHandler(runner)
|
||||
req := httptest.NewRequest(http.MethodPost, "/v1/runs", bytes.NewReader(body))
|
||||
w := httptest.NewRecorder()
|
||||
|
||||
h.ServeHTTP(w, req)
|
||||
|
||||
if w.Code != http.StatusOK {
|
||||
t.Fatalf("expected maintained HTTP request example to be accepted, got %d: %s", w.Code, w.Body.String())
|
||||
}
|
||||
if runner.last.PromptID != "generic.markdown_summary" {
|
||||
t.Fatalf("unexpected prompt ID: %q", runner.last.PromptID)
|
||||
}
|
||||
if runner.last.ProfileID != "local-fast" {
|
||||
t.Fatalf("unexpected profile ID: %q", runner.last.ProfileID)
|
||||
}
|
||||
wantInputs := map[string]domain.ArtifactRef{
|
||||
"transcript": {Type: domain.ArtifactRefFile, URI: "./examples/fixtures/transcript.md"},
|
||||
"glossary": {Type: domain.ArtifactRefFile, URI: "./examples/fixtures/glossary.yml"},
|
||||
}
|
||||
if !reflect.DeepEqual(runner.last.Inputs, wantInputs) {
|
||||
t.Fatalf("unexpected inputs: got %#v, want %#v", runner.last.Inputs, wantInputs)
|
||||
}
|
||||
if runner.last.Vars["session_date"] != "2026-05-04" {
|
||||
t.Fatalf("unexpected session_date: %#v", runner.last.Vars)
|
||||
}
|
||||
}
|
||||
|
||||
type handlerPromptRepo struct {
|
||||
def *domain.PromptDefinition
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user