298 lines
9.4 KiB
Markdown
298 lines
9.4 KiB
Markdown
# HTTP API Reference
|
|
|
|
This is the canonical public HTTP contract for Scriptorium.
|
|
|
|
Implemented route:
|
|
|
|
- `POST /v1/runs`
|
|
|
|
For CLI behavior, see [CLI reference](cli.md). For config and prompt/profile
|
|
file formats, see [Configuration reference](config.md).
|
|
|
|
The maintained request-shape example is `examples/http-run.json`. It requires a
|
|
running `serve` process with an artifact root that can read the referenced
|
|
files, plus a reachable model endpoint for full execution.
|
|
|
|
## Base URL And Deployment
|
|
|
|
`scriptorium serve` listens on `server.addr` or `serve --addr`. The default is
|
|
`:8080`.
|
|
|
|
The route path is always:
|
|
|
|
```text
|
|
/v1/runs
|
|
```
|
|
|
|
The HTTP adapter has no built-in authentication or authorization. Deploy it
|
|
behind trusted network and authentication controls.
|
|
|
|
## Media Types
|
|
|
|
- Request body: JSON object.
|
|
- Response body: JSON object.
|
|
- Response `Content-Type`: `application/json`.
|
|
|
|
Requests are decoded as JSON regardless of the request `Content-Type` header.
|
|
There are no shared query parameters.
|
|
|
|
## Request Limits
|
|
|
|
HTTP limits are configured through `server.*` config fields or `serve` flags:
|
|
|
|
- `server.max_request_bytes`: encoded JSON request body limit, including inline input bodies.
|
|
- `server.max_artifact_bytes`: file artifact limit for HTTP `file` input references.
|
|
- `server.max_response_bytes`: encoded JSON response limit, including artifact body and optional raw output.
|
|
|
|
Each limit defaults to `16777216` bytes. `0` disables that limit.
|
|
|
|
## `POST /v1/runs`
|
|
|
|
Runs one prompt request and returns the generated artifact, validation result,
|
|
and metadata.
|
|
|
|
### Request Body
|
|
|
|
```json
|
|
{
|
|
"prompt_id": "generic.markdown_summary",
|
|
"profile_id": "local-fast",
|
|
"prompt_version": "1.0.0",
|
|
"inputs": {
|
|
"transcript": {
|
|
"type": "file",
|
|
"uri": "./examples/fixtures/transcript.md"
|
|
},
|
|
"glossary": {
|
|
"type": "inline",
|
|
"body": "party:\n - Rin"
|
|
}
|
|
},
|
|
"vars": {
|
|
"session_date": "2026-05-04"
|
|
},
|
|
"model": {
|
|
"endpoint": "http://localhost:8000/v1",
|
|
"model": "gpt-4o-mini",
|
|
"temperature": 0,
|
|
"max_tokens": 800,
|
|
"top_p": 1,
|
|
"timeout_seconds": 120,
|
|
"service_tier": "priority",
|
|
"reasoning_effort": "medium",
|
|
"api_key_env": "SCRIPTORIUM_API_KEY",
|
|
"extra_params": {
|
|
"provider_option": "enabled"
|
|
}
|
|
},
|
|
"include_raw_output": false
|
|
}
|
|
```
|
|
|
|
Request fields:
|
|
|
|
| Field | Required | Description |
|
|
| --- | --- | --- |
|
|
| `prompt_id` | yes | Prompt ID. Must not be blank. |
|
|
| `prompt_version` | no | Prompt version filter. |
|
|
| `profile_id` | no | Execution profile ID. If omitted, the prompt must define `default_profile`. |
|
|
| `inputs` | yes | Object mapping prompt input names to input references. Must contain at least one entry. |
|
|
| `vars` | no | Object mapping template variable names to string values. |
|
|
| `model` | no | Runtime model override object. |
|
|
| `include_raw_output` | no | When `true`, include `raw_model_output` in the response. |
|
|
|
|
Input reference fields:
|
|
|
|
| Field | Required | Description |
|
|
| --- | --- | --- |
|
|
| `type` | yes | `file` or `inline`. |
|
|
| `uri` | for `file` | File URI/path. |
|
|
| `body` | for `inline` | Inline artifact body. |
|
|
|
|
HTTP `file` references require `server.artifact_root` or `serve
|
|
--artifact-root`. Relative file URIs resolve against that root. Absolute file
|
|
URIs are accepted only when lexically inside the root. Relative traversal and
|
|
absolute paths outside the root return `400 artifact_not_allowed`.
|
|
|
|
The containment check is lexical and does not resolve symlinks. Symlinks inside
|
|
the artifact root are followed by the operating system, including symlinks that
|
|
point outside the root. Keep the artifact root narrow and not writable by
|
|
untrusted users.
|
|
|
|
Model override fields:
|
|
|
|
| Field | Description |
|
|
| --- | --- |
|
|
| `endpoint` | Runtime endpoint override. |
|
|
| `model` | Runtime model override. |
|
|
| `temperature` | Number in range `0..2`. Explicit `0` is an override. |
|
|
| `max_tokens` | Integer greater than or equal to `0`. Explicit `0` is an override. |
|
|
| `top_p` | Number in range `0..1`. Explicit `0` is an override. |
|
|
| `timeout_seconds` | Integer greater than or equal to `0`. Explicit `0` disables the outbound client timeout. |
|
|
| `service_tier` | Provider-specific request tier. |
|
|
| `reasoning_effort` | Provider-specific reasoning setting. |
|
|
| `api_key_env` | Name of an environment variable containing the API key. |
|
|
| `extra_params` | JSON-compatible provider-specific top-level request fields. |
|
|
|
|
Raw API-key values are not accepted in HTTP payloads. A field such as
|
|
`api_key` is rejected as unknown JSON.
|
|
|
|
`extra_params` keys must not be empty and must not collide with reserved
|
|
outbound fields: `model`, `session_id`, `messages`, `temperature`,
|
|
`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or
|
|
`response_format`.
|
|
|
|
### Strict JSON Rules
|
|
|
|
Request decoding is strict:
|
|
|
|
- malformed JSON returns `400 invalid_json`
|
|
- unknown request fields return `400 invalid_json`
|
|
- unknown `inputs` item fields return `400 invalid_json`
|
|
- unknown `model` fields return `400 invalid_json`
|
|
- trailing JSON tokens after the request object return `400 invalid_json`
|
|
- request bodies above the configured limit return `413 request_too_large`
|
|
|
|
### Success Response
|
|
|
|
Status: `200 OK`
|
|
|
|
```json
|
|
{
|
|
"artifact": {
|
|
"name": "output",
|
|
"content_type": "text/markdown",
|
|
"body": "Generated content",
|
|
"size": 17,
|
|
"hash": "..."
|
|
},
|
|
"validation": {
|
|
"status": "passed",
|
|
"mode": "basic",
|
|
"repair_attempts": 0,
|
|
"is_valid": true
|
|
},
|
|
"metadata": {
|
|
"run_id": "...",
|
|
"prompt_id": "generic.markdown_summary",
|
|
"prompt_version": "1.0.0",
|
|
"prompt_hash": "...",
|
|
"rendered_prompt_hash": "...",
|
|
"selected_profile_id": "local-fast",
|
|
"model_name": "gpt-4o-mini",
|
|
"endpoint": "http://localhost:8000/v1",
|
|
"model_params": {
|
|
"endpoint": "http://localhost:8000/v1",
|
|
"model": "gpt-4o-mini",
|
|
"temperature": 0.2,
|
|
"max_tokens": 500,
|
|
"top_p": 1,
|
|
"timeout_seconds": 90
|
|
},
|
|
"input_hashes": {
|
|
"transcript": "..."
|
|
},
|
|
"usage": {
|
|
"prompt_tokens": 11,
|
|
"completion_tokens": 22,
|
|
"total_tokens": 33,
|
|
"cached_tokens": 0,
|
|
"cache_write_tokens": 0
|
|
},
|
|
"start_time": "2026-05-04T12:00:00Z",
|
|
"end_time": "2026-05-04T12:00:01Z",
|
|
"duration_ms": 1000,
|
|
"validation_mode": "basic",
|
|
"validation_status": "passed",
|
|
"repair_attempts_used": 0
|
|
}
|
|
}
|
|
```
|
|
|
|
Response fields:
|
|
|
|
- `artifact`: generated output artifact.
|
|
- `validation`: validation result for the generated artifact.
|
|
- `metadata`: run and effective runtime metadata.
|
|
- `raw_model_output`: omitted unless `include_raw_output` is `true`.
|
|
|
|
`artifact.uri` is omitted when empty. `validation.errors` and
|
|
`validation.schema_path` are omitted when empty. `model_params.service_tier`,
|
|
`model_params.reasoning_effort`, `model_params.api_key_env`, and
|
|
`model_params.extra_params` are omitted when empty.
|
|
|
|
`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are
|
|
always present as numbers. They are `0` when the provider omits compatible cache
|
|
usage fields or reports no cache activity.
|
|
|
|
### Validation Failure Response
|
|
|
|
Generated-content validation failures still return `200 OK`.
|
|
|
|
```json
|
|
{
|
|
"validation": {
|
|
"status": "failed",
|
|
"mode": "json",
|
|
"errors": ["invalid JSON: ..."],
|
|
"repair_attempts": 0,
|
|
"is_valid": false
|
|
}
|
|
}
|
|
```
|
|
|
|
The response still includes `artifact` and `metadata`.
|
|
|
|
## Error Responses
|
|
|
|
Error body shape:
|
|
|
|
```json
|
|
{
|
|
"error": {
|
|
"code": "invalid_request",
|
|
"message": "prompt_id is required"
|
|
}
|
|
}
|
|
```
|
|
|
|
Current status/code mapping:
|
|
|
|
| Status | Code | Meaning |
|
|
| --- | --- | --- |
|
|
| `400` | `invalid_json` | Malformed JSON, unknown JSON field, or trailing JSON token. |
|
|
| `400` | `invalid_request` | Missing/invalid request fields or invalid runtime overrides. |
|
|
| `400` | `profile_required` | No `profile_id` and prompt has no `default_profile`. |
|
|
| `400` | `prompt_load_failed` | Prompt definition YAML/contract failed to load. |
|
|
| `400` | `profile_load_failed` | Profile YAML/contract failed to load, including raw `api_key`. |
|
|
| `400` | `artifact_not_allowed` | HTTP file refs are disabled or requested path is outside artifact root. |
|
|
| `400` | `artifact_read_failed` | Input artifact could not be read or input ref was unsupported/invalid. |
|
|
| `400` | `prompt_render_failed` | Prompt template rendering failed. |
|
|
| `400` | `api_key_env_missing` | Selected `api_key_env` variable is unset or empty. |
|
|
| `404` | `not_found` | Route path is unknown. |
|
|
| `404` | `prompt_not_found` | Prompt ID/version was not found. |
|
|
| `404` | `profile_not_found` | Profile ID was not found. |
|
|
| `405` | `method_not_allowed` | Method is not `POST` on `/v1/runs`. |
|
|
| `413` | `request_too_large` | Encoded JSON request body exceeds configured request limit. |
|
|
| `413` | `artifact_too_large` | HTTP file input artifact exceeds configured artifact limit. |
|
|
| `413` | `response_too_large` | Encoded JSON response exceeds configured response limit. |
|
|
| `500` | `validation_runtime_failed` | Validator runtime/schema loading failed. |
|
|
| `500` | `internal_error` | Unclassified server error. |
|
|
| `502` | `llm_failed` | Outbound model request failed. |
|
|
|
|
HTTP error messages are intentionally concise and do not include sensitive
|
|
internal causes.
|
|
|
|
## Retry And Idempotency
|
|
|
|
Scriptorium does not provide idempotency keys, pagination, caching headers, or
|
|
rate limiting.
|
|
|
|
Clients may retry transport failures or `5xx` responses when their surrounding
|
|
workflow can tolerate another model call. A retry can generate different output
|
|
and incur another provider request.
|
|
|
|
## Example File
|
|
|
|
- `examples/http-run.json`
|