Clarify HTTP operations documentation

This commit is contained in:
2026-07-05 03:06:39 +00:00
parent 574f88bd6a
commit 879cb021b2
3 changed files with 448 additions and 395 deletions

View File

@@ -1,40 +1,68 @@
# HTTP API Reference
## Scope
This is the canonical public HTTP contract for Scriptorium.
This document is the canonical public HTTP contract for Scriptorium.
Current scope is only:
Implemented route:
- `POST /v1/runs`
For CLI behavior, see the [CLI reference](cli.md).
For CLI behavior, see [CLI reference](cli.md). For config and prompt/profile
file formats, see [Configuration reference](config.md).
## Endpoint
## Base URL And Deployment
- Method: `POST`
- Path: `/v1/runs`
- Content type: JSON request/response
`scriptorium serve` listens on `server.addr` or `serve --addr`. The default is
`:8080`.
Route behavior:
The route path is always:
- unknown path: `404 not_found`
- unsupported method on `/v1/runs`: `405 method_not_allowed`
```text
/v1/runs
```
Copyable request example file:
The HTTP adapter has no built-in authentication or authorization. Deploy it
behind trusted network and authentication controls.
- `examples/http-run.json`
## Media Types
## Request Body
- Request body: JSON object.
- Response body: JSON object.
- Response `Content-Type`: `application/json`.
Requests are decoded as JSON regardless of the request `Content-Type` header.
There are no shared query parameters.
## Request Limits
HTTP limits are configured through `server.*` config fields or `serve` flags:
- `server.max_request_bytes`: encoded JSON request body limit, including inline input bodies.
- `server.max_artifact_bytes`: file artifact limit for HTTP `file` input references.
- `server.max_response_bytes`: encoded JSON response limit, including artifact body and optional raw output.
Each limit defaults to `16777216` bytes. `0` disables that limit.
## `POST /v1/runs`
Runs one prompt request and returns the generated artifact, validation result,
and metadata.
### Request Body
```json
{
"prompt_id": "generic.structured_events",
"profile_id": "local-quality",
"prompt_id": "generic.markdown_summary",
"profile_id": "local-fast",
"prompt_version": "1.0.0",
"inputs": {
"transcript": {"type": "file", "uri": "./examples/fixtures/transcript.md"},
"glossary": {"type": "inline", "body": "party:\n - Rin"}
"transcript": {
"type": "file",
"uri": "./examples/fixtures/transcript.md"
},
"glossary": {
"type": "inline",
"body": "party:\n - Rin"
}
},
"vars": {
"session_date": "2026-05-04"
@@ -42,112 +70,120 @@ Copyable request example file:
"model": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0.0,
"temperature": 0,
"max_tokens": 800,
"top_p": 1.0,
"top_p": 1,
"timeout_seconds": 120,
"service_tier": "priority",
"reasoning_effort": "medium",
"api_key_env": "SCRIPTORIUM_API_KEY",
"extra_params": {
"route": "primary",
"provider_options": {
"retry_budget": 2
}
"provider_option": "enabled"
}
},
"include_raw_output": false
}
```
Required fields:
Request fields:
- `prompt_id`
- `inputs` (must contain at least one named input)
| Field | Required | Description |
| --- | --- | --- |
| `prompt_id` | yes | Prompt ID. Must not be blank. |
| `prompt_version` | no | Prompt version filter. |
| `profile_id` | no | Execution profile ID. If omitted, the prompt must define `default_profile`. |
| `inputs` | yes | Object mapping prompt input names to input references. Must contain at least one entry. |
| `vars` | no | Object mapping template variable names to string values. |
| `model` | no | Runtime model override object. |
| `include_raw_output` | no | When `true`, include `raw_model_output` in the response. |
Input reference types currently supported by runtime artifact loading:
Input reference fields:
- `file`
- `inline`
| Field | Required | Description |
| --- | --- | --- |
| `type` | yes | `file` or `inline`. |
| `uri` | for `file` | File URI/path. |
| `body` | for `inline` | Inline artifact body. |
HTTP `file` references require `server.artifact_root` or `serve --artifact-root`.
Relative file URIs resolve against that root. Absolute file URIs are accepted
only when they are lexically inside the root. Requests that escape the root by
lexical traversal, including `..` traversal and absolute paths outside the root,
return `400 artifact_not_allowed`. Symlinks inside the root are followed by the
operating system, including symlinks that point outside the root. The artifact
root must not be writable by untrusted users. `inline` references do not require
an artifact root.
HTTP file artifacts above the configured artifact limit return
`413 artifact_too_large`. Inline bodies are bounded by the request body limit.
HTTP `file` references require `server.artifact_root` or `serve
--artifact-root`. Relative file URIs resolve against that root. Absolute file
URIs are accepted only when lexically inside the root. Relative traversal and
absolute paths outside the root return `400 artifact_not_allowed`.
Model override notes:
The containment check is lexical and does not resolve symlinks. Symlinks inside
the artifact root are followed by the operating system, including symlinks that
point outside the root. Keep the artifact root narrow and not writable by
untrusted users.
- Numeric model override fields distinguish omitted values from explicit zero values. For example, omitting `temperature` preserves the selected profile/default value, while `"temperature": 0` explicitly sets the effective temperature to zero.
- `extra_params` accepts JSON-compatible values: strings, numbers, booleans, objects, and arrays.
- `extra_params` are passed through effective model metadata and flattened into top-level provider request fields by the OpenAI-compatible client.
- `extra_params` keys must not be empty and must not collide with reserved outbound fields: `model`, `session_id`, `messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or `response_format`.
- Raw API-key values are not accepted. Use `api_key_env` to name an environment variable.
Model override fields:
## Strict JSON Rules
| Field | Description |
| --- | --- |
| `endpoint` | Runtime endpoint override. |
| `model` | Runtime model override. |
| `temperature` | Number in range `0..2`. Explicit `0` is an override. |
| `max_tokens` | Integer greater than or equal to `0`. Explicit `0` is an override. |
| `top_p` | Number in range `0..1`. Explicit `0` is an override. |
| `timeout_seconds` | Integer greater than or equal to `0`. Explicit `0` disables the outbound client timeout. |
| `service_tier` | Provider-specific request tier. |
| `reasoning_effort` | Provider-specific reasoning setting. |
| `api_key_env` | Name of an environment variable containing the API key. |
| `extra_params` | JSON-compatible provider-specific top-level request fields. |
Request decoding uses strict JSON field checks:
Raw API-key values are not accepted in HTTP payloads. A field such as
`api_key` is rejected as unknown JSON.
- unknown request fields are rejected with `400 invalid_json`
- unknown `model` fields are rejected with `400 invalid_json`
- raw API-key payload fields such as `api_key` are rejected as unknown fields
- request bodies above the configured request limit are rejected with `413 request_too_large`
- trailing JSON tokens after the request object are rejected with `400 invalid_json`
`extra_params` keys must not be empty and must not collide with reserved
outbound fields: `model`, `session_id`, `messages`, `temperature`,
`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or
`response_format`.
## Success Response
### Strict JSON Rules
Request decoding is strict:
- malformed JSON returns `400 invalid_json`
- unknown request fields return `400 invalid_json`
- unknown `inputs` item fields return `400 invalid_json`
- unknown `model` fields return `400 invalid_json`
- trailing JSON tokens after the request object return `400 invalid_json`
- request bodies above the configured limit return `413 request_too_large`
### Success Response
Status: `200 OK`
Response shape:
```json
{
"artifact": {
"name": "output",
"content_type": "application/json",
"body": "{\"summary\":\"...\"}",
"uri": "",
"size": 123,
"content_type": "text/markdown",
"body": "Generated content",
"size": 17,
"hash": "..."
},
"validation": {
"status": "passed",
"mode": "json_schema",
"errors": [],
"schema_path": "structured_events.schema.json",
"mode": "basic",
"repair_attempts": 0,
"is_valid": true
},
"metadata": {
"run_id": "...",
"prompt_id": "generic.structured_events",
"prompt_id": "generic.markdown_summary",
"prompt_version": "1.0.0",
"prompt_hash": "...",
"rendered_prompt_hash": "...",
"selected_profile_id": "local-quality",
"selected_profile_id": "local-fast",
"model_name": "gpt-4o-mini",
"endpoint": "http://localhost:8000/v1",
"model_params": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0,
"max_tokens": 800,
"temperature": 0.2,
"max_tokens": 500,
"top_p": 1,
"timeout_seconds": 120,
"service_tier": "priority",
"reasoning_effort": "medium",
"api_key_env": "SCRIPTORIUM_API_KEY",
"extra_params": {
"route": "primary",
"provider_options": {
"retry_budget": 2
}
}
"timeout_seconds": 90
},
"input_hashes": {
"transcript": "..."
@@ -162,30 +198,46 @@ Response shape:
"start_time": "2026-05-04T12:00:00Z",
"end_time": "2026-05-04T12:00:01Z",
"duration_ms": 1000,
"validation_mode": "json_schema",
"validation_mode": "basic",
"validation_status": "passed",
"repair_attempts_used": 0
}
}
```
`raw_model_output` is omitted by default.
Response fields:
`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are always present as numbers. They are `0` when the provider omits compatible cache usage fields or reports no cache activity.
- `artifact`: generated output artifact.
- `validation`: validation result for the generated artifact.
- `metadata`: run and effective runtime metadata.
- `raw_model_output`: omitted unless `include_raw_output` is `true`.
To include it, send:
`artifact.uri` is omitted when empty. `validation.errors` and
`validation.schema_path` are omitted when empty. `model_params.service_tier`,
`model_params.reasoning_effort`, `model_params.api_key_env`, and
`model_params.extra_params` are omitted when empty.
- `"include_raw_output": true`
`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are
always present as numbers. They are `0` when the provider omits compatible cache
usage fields or reports no cache activity.
## Validation Failure Behavior
### Validation Failure Response
Validation content failures do not map to HTTP error status.
Generated-content validation failures still return `200 OK`.
Behavior:
```json
{
"validation": {
"status": "failed",
"mode": "json",
"errors": ["invalid JSON: ..."],
"repair_attempts": 0,
"is_valid": false
}
}
```
- status remains `200 OK`
- `validation.status` is `failed`
- validation errors are returned in `validation.errors`
The response still includes `artifact` and `metadata`.
## Error Responses
@@ -200,28 +252,38 @@ Error body shape:
}
```
Current error mapping (non-exhaustive):
Current status/code mapping:
- `400 invalid_json`: malformed JSON or unknown JSON fields
- `400 invalid_request`: missing/invalid request fields
- `400 profile_required`: no explicit `profile_id` and prompt has no `default_profile`
- `400 prompt_load_failed`: prompt definition invalid/unloadable
- `400 profile_load_failed`: profile invalid/unloadable
- `400 artifact_not_allowed`: file input artifact is outside the configured artifact root or file refs are not enabled
- `400 artifact_read_failed`: input artifact loading failed
- `400 prompt_render_failed`: template render failed
- `400 api_key_env_missing`: named API-key environment variable is missing
- `413 request_too_large`: request body exceeds the configured request limit
- `413 artifact_too_large`: HTTP file input artifact exceeds the configured artifact limit
- `413 response_too_large`: encoded JSON response exceeds the configured response limit
- `404 prompt_not_found`
- `404 profile_not_found`
- `502 llm_failed`: outbound model request failed
- `500 validation_runtime_failed`: validator runtime/schema-load failure
- `500 internal_error`
| Status | Code | Meaning |
| --- | --- | --- |
| `400` | `invalid_json` | Malformed JSON, unknown JSON field, or trailing JSON token. |
| `400` | `invalid_request` | Missing/invalid request fields or invalid runtime overrides. |
| `400` | `profile_required` | No `profile_id` and prompt has no `default_profile`. |
| `400` | `prompt_load_failed` | Prompt definition YAML/contract failed to load. |
| `400` | `profile_load_failed` | Profile YAML/contract failed to load, including raw `api_key`. |
| `400` | `artifact_not_allowed` | HTTP file refs are disabled or requested path is outside artifact root. |
| `400` | `artifact_read_failed` | Input artifact could not be read or input ref was unsupported/invalid. |
| `400` | `prompt_render_failed` | Prompt template rendering failed. |
| `400` | `api_key_env_missing` | Selected `api_key_env` variable is unset or empty. |
| `404` | `not_found` | Route path is unknown. |
| `404` | `prompt_not_found` | Prompt ID/version was not found. |
| `404` | `profile_not_found` | Profile ID was not found. |
| `405` | `method_not_allowed` | Method is not `POST` on `/v1/runs`. |
| `413` | `request_too_large` | Encoded JSON request body exceeds configured request limit. |
| `413` | `artifact_too_large` | HTTP file input artifact exceeds configured artifact limit. |
| `413` | `response_too_large` | Encoded JSON response exceeds configured response limit. |
| `500` | `validation_runtime_failed` | Validator runtime/schema loading failed. |
| `500` | `internal_error` | Unclassified server error. |
| `502` | `llm_failed` | Outbound model request failed. |
## Security And Deployment Note
HTTP error messages are intentionally concise and do not include sensitive
internal causes.
The HTTP adapter has no built-in authentication or authorization.
## Retry And Idempotency
Deploy behind trusted controls (for example authenticated gateway/reverse proxy and network boundaries).
Scriptorium does not provide idempotency keys, pagination, caching headers, or
rate limiting.
Clients may retry transport failures or `5xx` responses when their surrounding
workflow can tolerate another model call. A retry can generate different output
and incur another provider request.