9.2 KiB
HTTP API Reference
This is the canonical public HTTP contract for Scriptorium.
Implemented route:
POST /v1/runs
For CLI behavior, see CLI reference. For config and prompt/profile file formats, see Configuration reference.
Base URL And Deployment
scriptorium serve listens on server.addr or serve --addr. The default is
:8080.
The route path is always:
/v1/runs
The HTTP adapter has no built-in authentication or authorization. Deploy it behind trusted network and authentication controls.
Media Types
- Request body: JSON object.
- Response body: JSON object.
- Response
Content-Type:application/json.
Requests are decoded as JSON regardless of the request Content-Type header.
There are no shared query parameters.
Request Limits
HTTP limits are configured through server.* config fields or serve flags:
server.max_request_bytes: encoded JSON request body limit, including inline input bodies.server.max_artifact_bytes: file artifact limit for HTTPfileinput references.server.max_response_bytes: encoded JSON response limit, including artifact body and optional raw output.
Each limit defaults to 16777216 bytes. 0 disables that limit.
POST /v1/runs
Runs one prompt request and returns the generated artifact, validation result, and metadata.
Request Body
{
"prompt_id": "generic.markdown_summary",
"profile_id": "local-fast",
"prompt_version": "1.0.0",
"inputs": {
"transcript": {
"type": "file",
"uri": "./examples/fixtures/transcript.md"
},
"glossary": {
"type": "inline",
"body": "party:\n - Rin"
}
},
"vars": {
"session_date": "2026-05-04"
},
"model": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0,
"max_tokens": 800,
"top_p": 1,
"timeout_seconds": 120,
"service_tier": "priority",
"reasoning_effort": "medium",
"api_key_env": "SCRIPTORIUM_API_KEY",
"extra_params": {
"provider_option": "enabled"
}
},
"include_raw_output": false
}
Request fields:
| Field | Required | Description |
|---|---|---|
prompt_id |
yes | Prompt ID. Must not be blank. |
prompt_version |
no | Prompt version filter. |
profile_id |
no | Execution profile ID. If omitted, the prompt must define default_profile. |
inputs |
yes | Object mapping prompt input names to input references. Must contain at least one entry. |
vars |
no | Object mapping template variable names to string values. |
model |
no | Runtime model override object. |
include_raw_output |
no | When true, include raw_model_output in the response. |
Input reference fields:
| Field | Required | Description |
|---|---|---|
type |
yes | file or inline. |
uri |
for file |
File URI/path. |
body |
for inline |
Inline artifact body. |
HTTP file references require server.artifact_root or serve --artifact-root. Relative file URIs resolve against that root. Absolute file
URIs are accepted only when lexically inside the root. Relative traversal and
absolute paths outside the root return 400 artifact_not_allowed.
The containment check is lexical and does not resolve symlinks. Symlinks inside the artifact root are followed by the operating system, including symlinks that point outside the root. Keep the artifact root narrow and not writable by untrusted users.
Model override fields:
| Field | Description |
|---|---|
endpoint |
Runtime endpoint override. |
model |
Runtime model override. |
temperature |
Number in range 0..2. Explicit 0 is an override. |
max_tokens |
Integer greater than or equal to 0. Explicit 0 is an override. |
top_p |
Number in range 0..1. Explicit 0 is an override. |
timeout_seconds |
Integer greater than or equal to 0. Explicit 0 disables the outbound client timeout. |
service_tier |
Provider-specific request tier. |
reasoning_effort |
Provider-specific reasoning setting. |
api_key_env |
Name of an environment variable containing the API key. |
extra_params |
JSON-compatible provider-specific top-level request fields. |
Raw API-key values are not accepted in HTTP payloads. A field such as
api_key is rejected as unknown JSON.
extra_params keys must not be empty and must not collide with reserved
outbound fields: model, session_id, messages, temperature,
max_tokens, top_p, service_tier, reasoning_effort, or
response_format.
Strict JSON Rules
Request decoding is strict:
- malformed JSON returns
400 invalid_json - unknown request fields return
400 invalid_json - unknown
inputsitem fields return400 invalid_json - unknown
modelfields return400 invalid_json - trailing JSON tokens after the request object return
400 invalid_json - request bodies above the configured limit return
413 request_too_large
Success Response
Status: 200 OK
{
"artifact": {
"name": "output",
"content_type": "text/markdown",
"body": "Generated content",
"size": 17,
"hash": "..."
},
"validation": {
"status": "passed",
"mode": "basic",
"repair_attempts": 0,
"is_valid": true
},
"metadata": {
"run_id": "...",
"prompt_id": "generic.markdown_summary",
"prompt_version": "1.0.0",
"prompt_hash": "...",
"rendered_prompt_hash": "...",
"selected_profile_id": "local-fast",
"model_name": "gpt-4o-mini",
"endpoint": "http://localhost:8000/v1",
"model_params": {
"endpoint": "http://localhost:8000/v1",
"model": "gpt-4o-mini",
"temperature": 0.2,
"max_tokens": 500,
"top_p": 1,
"timeout_seconds": 90
},
"input_hashes": {
"transcript": "..."
},
"usage": {
"prompt_tokens": 11,
"completion_tokens": 22,
"total_tokens": 33,
"cached_tokens": 0,
"cache_write_tokens": 0
},
"start_time": "2026-05-04T12:00:00Z",
"end_time": "2026-05-04T12:00:01Z",
"duration_ms": 1000,
"validation_mode": "basic",
"validation_status": "passed",
"repair_attempts_used": 0
}
}
Response fields:
artifact: generated output artifact.validation: validation result for the generated artifact.metadata: run and effective runtime metadata.raw_model_output: omitted unlessinclude_raw_outputistrue.
artifact.uri is omitted when empty. validation.errors and
validation.schema_path are omitted when empty. model_params.service_tier,
model_params.reasoning_effort, model_params.api_key_env, and
model_params.extra_params are omitted when empty.
metadata.usage.cached_tokens and metadata.usage.cache_write_tokens are
always present as numbers. They are 0 when the provider omits compatible cache
usage fields or reports no cache activity.
Validation Failure Response
Generated-content validation failures still return 200 OK.
{
"validation": {
"status": "failed",
"mode": "json",
"errors": ["invalid JSON: ..."],
"repair_attempts": 0,
"is_valid": false
}
}
The response still includes artifact and metadata.
Error Responses
Error body shape:
{
"error": {
"code": "invalid_request",
"message": "prompt_id is required"
}
}
Current status/code mapping:
| Status | Code | Meaning |
|---|---|---|
400 |
invalid_json |
Malformed JSON, unknown JSON field, or trailing JSON token. |
400 |
invalid_request |
Missing/invalid request fields or invalid runtime overrides. |
400 |
profile_required |
No profile_id and prompt has no default_profile. |
400 |
prompt_load_failed |
Prompt definition YAML/contract failed to load. |
400 |
profile_load_failed |
Profile YAML/contract failed to load, including raw api_key. |
400 |
artifact_not_allowed |
HTTP file refs are disabled or requested path is outside artifact root. |
400 |
artifact_read_failed |
Input artifact could not be read or input ref was unsupported/invalid. |
400 |
prompt_render_failed |
Prompt template rendering failed. |
400 |
api_key_env_missing |
Selected api_key_env variable is unset or empty. |
404 |
not_found |
Route path is unknown. |
404 |
prompt_not_found |
Prompt ID/version was not found. |
404 |
profile_not_found |
Profile ID was not found. |
405 |
method_not_allowed |
Method is not POST on /v1/runs. |
413 |
request_too_large |
Encoded JSON request body exceeds configured request limit. |
413 |
artifact_too_large |
HTTP file input artifact exceeds configured artifact limit. |
413 |
response_too_large |
Encoded JSON response exceeds configured response limit. |
500 |
validation_runtime_failed |
Validator runtime/schema loading failed. |
500 |
internal_error |
Unclassified server error. |
502 |
llm_failed |
Outbound model request failed. |
HTTP error messages are intentionally concise and do not include sensitive internal causes.
Retry And Idempotency
Scriptorium does not provide idempotency keys, pagination, caching headers, or rate limiting.
Clients may retry transport failures or 5xx responses when their surrounding
workflow can tolerate another model call. A retry can generate different output
and incur another provider request.