Files
scriptorium/docs/api.md

6.1 KiB

HTTP API Reference

This is the canonical public HTTP contract for Scriptorium.

Service And Route

POST /v1/runs runs one prompt request and returns generated output, validation, and metadata. The service has no built-in authentication or authorization; deploy it behind appropriate network and authentication controls.

The service address and HTTP limits are configured as described in the configuration reference. serve invocation is defined in the CLI reference.

Requests and responses are JSON objects. Requests are decoded as JSON regardless of their Content-Type; successful JSON responses use Content-Type: application/json. There are no query parameters.

Request Limits

The configured request-body limit includes inline artifact bodies. The artifact limit applies to HTTP file inputs. The response limit applies to the encoded response, including the artifact body and optional raw output. A limit of zero disables that limit.

A request body over its limit returns 413 request_too_large; an oversized file input returns 413 artifact_too_large; an oversized encoded response returns 413 response_too_large.

POST /v1/runs

Request Body

The maintained request example is a complete copyable shape. The smallest valid shape is:

{
  "prompt_id": "generic.markdown_summary",
  "inputs": {
    "transcript": {"type": "inline", "body": "Source text"}
  }
}
Field Required Meaning
prompt_id yes Non-blank prompt ID.
prompt_version no Prompt version filter.
profile_id no Execution-profile ID; otherwise the prompt must set default_profile.
inputs yes Non-empty object mapping input names to references.
vars no Object mapping template-variable names to strings.
model no Runtime model-override object.
include_raw_output no Include raw_model_output when true.

An input reference has a required type of file or inline. A file reference requires uri; an inline reference requires body.

HTTP file references require a configured artifact root. Relative paths resolve within that root. Absolute paths must be lexically within it; traversal outside it is rejected with 400 artifact_not_allowed. This lexical check does not resolve symlinks: the operating system follows symlinks inside the root, including ones that target outside it. Keep the root narrow and inaccessible to untrusted writers.

The optional model object accepts endpoint, model, temperature, max_tokens, top_p, timeout_seconds, service_tier, reasoning_effort, api_key_env, and extra_params. Numeric ranges and credential supply are defined by the configuration reference. Explicit zero values for the numeric fields are overrides; zero timeout_seconds disables the outbound client timeout.

Raw API-key values are not accepted. api_key and any other unknown model field cause 400 invalid_json.

Strict JSON

Request decoding rejects malformed JSON, unknown fields at every request level, and trailing JSON tokens with 400 invalid_json. A blank prompt_id or empty inputs object returns 400 invalid_request.

Success Response

A completed run returns 200 OK, including when generated content fails its validation contract. The response contains:

  • artifact: name, content_type, body, size, hash, and optional uri;
  • validation: status, mode, repair_attempts, is_valid, plus optional errors and schema_path;
  • metadata: run, prompt, rendered-prompt, profile, model, input-hash, usage, timing, validation, and repair-attempt metadata; and
  • optional raw_model_output when requested.

metadata.model_params has endpoint, model, temperature, max_tokens, top_p, and timeout_seconds, plus optional service_tier, reasoning_effort, api_key_env, and extra_params. metadata.usage always includes prompt_tokens, completion_tokens, total_tokens, cached_tokens, and cache_write_tokens; unavailable cache usage is reported as zero.

A validation failure has validation.status: "failed", is_valid: false, and any available diagnostic errors, while still returning the artifact and metadata.

Error Responses

Errors have this shape:

{"error":{"code":"invalid_request","message":"prompt_id is required"}}

Messages are concise and do not expose wrapped internal causes.

Status Code Meaning
400 invalid_json Malformed JSON, unknown field, or trailing JSON.
400 invalid_request Missing or invalid request data or runtime override.
400 profile_required No profile ID and no prompt default profile.
400 prompt_load_failed Prompt definition failed to load.
400 profile_load_failed Profile failed to load.
400 artifact_not_allowed HTTP file input is disabled or outside the artifact root.
400 artifact_read_failed Input artifact is invalid or cannot be read.
400 prompt_render_failed Prompt template rendering failed.
400 api_key_env_missing The selected credential environment variable is unset or empty.
404 not_found Route does not exist.
404 prompt_not_found Prompt ID or version does not exist.
404 profile_not_found Profile ID does not exist.
405 method_not_allowed The route does not accept the method.
413 request_too_large Encoded request exceeds its limit.
413 artifact_too_large File input exceeds its limit.
413 response_too_large Encoded response exceeds its limit.
500 validation_runtime_failed Schema or validator runtime failure.
500 internal_error Unclassified server failure.
502 llm_failed Outbound model request failed.

Retry And Idempotency

Scriptorium provides no idempotency keys, pagination, caching headers, or rate limits. Clients may retry transport failures or 5xx responses only when their workflow tolerates another model call: a retry can produce different output and incur another provider request.