# HTTP API Reference This is the canonical public HTTP contract for Scriptorium. Implemented route: - `POST /v1/runs` For CLI behavior, see [CLI reference](cli.md). For config and prompt/profile file formats, see [Configuration reference](config.md). ## Base URL And Deployment `scriptorium serve` listens on `server.addr` or `serve --addr`. The default is `:8080`. The route path is always: ```text /v1/runs ``` The HTTP adapter has no built-in authentication or authorization. Deploy it behind trusted network and authentication controls. ## Media Types - Request body: JSON object. - Response body: JSON object. - Response `Content-Type`: `application/json`. Requests are decoded as JSON regardless of the request `Content-Type` header. There are no shared query parameters. ## Request Limits HTTP limits are configured through `server.*` config fields or `serve` flags: - `server.max_request_bytes`: encoded JSON request body limit, including inline input bodies. - `server.max_artifact_bytes`: file artifact limit for HTTP `file` input references. - `server.max_response_bytes`: encoded JSON response limit, including artifact body and optional raw output. Each limit defaults to `16777216` bytes. `0` disables that limit. ## `POST /v1/runs` Runs one prompt request and returns the generated artifact, validation result, and metadata. ### Request Body ```json { "prompt_id": "generic.markdown_summary", "profile_id": "local-fast", "prompt_version": "1.0.0", "inputs": { "transcript": { "type": "file", "uri": "./examples/fixtures/transcript.md" }, "glossary": { "type": "inline", "body": "party:\n - Rin" } }, "vars": { "session_date": "2026-05-04" }, "model": { "endpoint": "http://localhost:8000/v1", "model": "gpt-4o-mini", "temperature": 0, "max_tokens": 800, "top_p": 1, "timeout_seconds": 120, "service_tier": "priority", "reasoning_effort": "medium", "api_key_env": "SCRIPTORIUM_API_KEY", "extra_params": { "provider_option": "enabled" } }, "include_raw_output": false } ``` Request fields: | Field | Required | Description | | --- | --- | --- | | `prompt_id` | yes | Prompt ID. Must not be blank. | | `prompt_version` | no | Prompt version filter. | | `profile_id` | no | Execution profile ID. If omitted, the prompt must define `default_profile`. | | `inputs` | yes | Object mapping prompt input names to input references. Must contain at least one entry. | | `vars` | no | Object mapping template variable names to string values. | | `model` | no | Runtime model override object. | | `include_raw_output` | no | When `true`, include `raw_model_output` in the response. | Input reference fields: | Field | Required | Description | | --- | --- | --- | | `type` | yes | `file` or `inline`. | | `uri` | for `file` | File URI/path. | | `body` | for `inline` | Inline artifact body. | HTTP `file` references require `server.artifact_root` or `serve --artifact-root`. Relative file URIs resolve against that root. Absolute file URIs are accepted only when lexically inside the root. Relative traversal and absolute paths outside the root return `400 artifact_not_allowed`. The containment check is lexical and does not resolve symlinks. Symlinks inside the artifact root are followed by the operating system, including symlinks that point outside the root. Keep the artifact root narrow and not writable by untrusted users. Model override fields: | Field | Description | | --- | --- | | `endpoint` | Runtime endpoint override. | | `model` | Runtime model override. | | `temperature` | Number in range `0..2`. Explicit `0` is an override. | | `max_tokens` | Integer greater than or equal to `0`. Explicit `0` is an override. | | `top_p` | Number in range `0..1`. Explicit `0` is an override. | | `timeout_seconds` | Integer greater than or equal to `0`. Explicit `0` disables the outbound client timeout. | | `service_tier` | Provider-specific request tier. | | `reasoning_effort` | Provider-specific reasoning setting. | | `api_key_env` | Name of an environment variable containing the API key. | | `extra_params` | JSON-compatible provider-specific top-level request fields. | Raw API-key values are not accepted in HTTP payloads. A field such as `api_key` is rejected as unknown JSON. `extra_params` keys must not be empty and must not collide with reserved outbound fields: `model`, `session_id`, `messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or `response_format`. ### Strict JSON Rules Request decoding is strict: - malformed JSON returns `400 invalid_json` - unknown request fields return `400 invalid_json` - unknown `inputs` item fields return `400 invalid_json` - unknown `model` fields return `400 invalid_json` - trailing JSON tokens after the request object return `400 invalid_json` - request bodies above the configured limit return `413 request_too_large` ### Success Response Status: `200 OK` ```json { "artifact": { "name": "output", "content_type": "text/markdown", "body": "Generated content", "size": 17, "hash": "..." }, "validation": { "status": "passed", "mode": "basic", "repair_attempts": 0, "is_valid": true }, "metadata": { "run_id": "...", "prompt_id": "generic.markdown_summary", "prompt_version": "1.0.0", "prompt_hash": "...", "rendered_prompt_hash": "...", "selected_profile_id": "local-fast", "model_name": "gpt-4o-mini", "endpoint": "http://localhost:8000/v1", "model_params": { "endpoint": "http://localhost:8000/v1", "model": "gpt-4o-mini", "temperature": 0.2, "max_tokens": 500, "top_p": 1, "timeout_seconds": 90 }, "input_hashes": { "transcript": "..." }, "usage": { "prompt_tokens": 11, "completion_tokens": 22, "total_tokens": 33, "cached_tokens": 0, "cache_write_tokens": 0 }, "start_time": "2026-05-04T12:00:00Z", "end_time": "2026-05-04T12:00:01Z", "duration_ms": 1000, "validation_mode": "basic", "validation_status": "passed", "repair_attempts_used": 0 } } ``` Response fields: - `artifact`: generated output artifact. - `validation`: validation result for the generated artifact. - `metadata`: run and effective runtime metadata. - `raw_model_output`: omitted unless `include_raw_output` is `true`. `artifact.uri` is omitted when empty. `validation.errors` and `validation.schema_path` are omitted when empty. `model_params.service_tier`, `model_params.reasoning_effort`, `model_params.api_key_env`, and `model_params.extra_params` are omitted when empty. `metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are always present as numbers. They are `0` when the provider omits compatible cache usage fields or reports no cache activity. ### Validation Failure Response Generated-content validation failures still return `200 OK`. ```json { "validation": { "status": "failed", "mode": "json", "errors": ["invalid JSON: ..."], "repair_attempts": 0, "is_valid": false } } ``` The response still includes `artifact` and `metadata`. ## Error Responses Error body shape: ```json { "error": { "code": "invalid_request", "message": "prompt_id is required" } } ``` Current status/code mapping: | Status | Code | Meaning | | --- | --- | --- | | `400` | `invalid_json` | Malformed JSON, unknown JSON field, or trailing JSON token. | | `400` | `invalid_request` | Missing/invalid request fields or invalid runtime overrides. | | `400` | `profile_required` | No `profile_id` and prompt has no `default_profile`. | | `400` | `prompt_load_failed` | Prompt definition YAML/contract failed to load. | | `400` | `profile_load_failed` | Profile YAML/contract failed to load, including raw `api_key`. | | `400` | `artifact_not_allowed` | HTTP file refs are disabled or requested path is outside artifact root. | | `400` | `artifact_read_failed` | Input artifact could not be read or input ref was unsupported/invalid. | | `400` | `prompt_render_failed` | Prompt template rendering failed. | | `400` | `api_key_env_missing` | Selected `api_key_env` variable is unset or empty. | | `404` | `not_found` | Route path is unknown. | | `404` | `prompt_not_found` | Prompt ID/version was not found. | | `404` | `profile_not_found` | Profile ID was not found. | | `405` | `method_not_allowed` | Method is not `POST` on `/v1/runs`. | | `413` | `request_too_large` | Encoded JSON request body exceeds configured request limit. | | `413` | `artifact_too_large` | HTTP file input artifact exceeds configured artifact limit. | | `413` | `response_too_large` | Encoded JSON response exceeds configured response limit. | | `500` | `validation_runtime_failed` | Validator runtime/schema loading failed. | | `500` | `internal_error` | Unclassified server error. | | `502` | `llm_failed` | Outbound model request failed. | HTTP error messages are intentionally concise and do not include sensitive internal causes. ## Retry And Idempotency Scriptorium does not provide idempotency keys, pagination, caching headers, or rate limiting. Clients may retry transport failures or `5xx` responses when their surrounding workflow can tolerate another model call. A retry can generate different output and incur another provider request.