Files
scriptorium/docs/api.md

9.4 KiB

HTTP API Reference

This is the canonical public HTTP contract for Scriptorium.

Implemented route:

  • POST /v1/runs

For CLI behavior, see CLI reference. For config and prompt/profile file formats, see Configuration reference.

The maintained request-shape example is examples/http-run.json. It requires a running serve process with an artifact root that can read the referenced files, plus a reachable model endpoint for full execution.

Base URL And Deployment

scriptorium serve listens on server.addr or serve --addr. The default is :8080.

The route path is always:

/v1/runs

The HTTP adapter has no built-in authentication or authorization. Deploy it behind trusted network and authentication controls.

Media Types

  • Request body: JSON object.
  • Response body: JSON object.
  • Response Content-Type: application/json.

Requests are decoded as JSON regardless of the request Content-Type header. There are no shared query parameters.

Request Limits

HTTP limits are configured through server.* config fields or serve flags:

  • server.max_request_bytes: encoded JSON request body limit, including inline input bodies.
  • server.max_artifact_bytes: file artifact limit for HTTP file input references.
  • server.max_response_bytes: encoded JSON response limit, including artifact body and optional raw output.

Each limit defaults to 16777216 bytes. 0 disables that limit.

POST /v1/runs

Runs one prompt request and returns the generated artifact, validation result, and metadata.

Request Body

{
  "prompt_id": "generic.markdown_summary",
  "profile_id": "local-fast",
  "prompt_version": "1.0.0",
  "inputs": {
    "transcript": {
      "type": "file",
      "uri": "./examples/fixtures/transcript.md"
    },
    "glossary": {
      "type": "inline",
      "body": "party:\n  - Rin"
    }
  },
  "vars": {
    "session_date": "2026-05-04"
  },
  "model": {
    "endpoint": "http://localhost:8000/v1",
    "model": "gpt-4o-mini",
    "temperature": 0,
    "max_tokens": 800,
    "top_p": 1,
    "timeout_seconds": 120,
    "service_tier": "priority",
    "reasoning_effort": "medium",
    "api_key_env": "SCRIPTORIUM_API_KEY",
    "extra_params": {
      "provider_option": "enabled"
    }
  },
  "include_raw_output": false
}

Request fields:

Field Required Description
prompt_id yes Prompt ID. Must not be blank.
prompt_version no Prompt version filter.
profile_id no Execution profile ID. If omitted, the prompt must define default_profile.
inputs yes Object mapping prompt input names to input references. Must contain at least one entry.
vars no Object mapping template variable names to string values.
model no Runtime model override object.
include_raw_output no When true, include raw_model_output in the response.

Input reference fields:

Field Required Description
type yes file or inline.
uri for file File URI/path.
body for inline Inline artifact body.

HTTP file references require server.artifact_root or serve --artifact-root. Relative file URIs resolve against that root. Absolute file URIs are accepted only when lexically inside the root. Relative traversal and absolute paths outside the root return 400 artifact_not_allowed.

The containment check is lexical and does not resolve symlinks. Symlinks inside the artifact root are followed by the operating system, including symlinks that point outside the root. Keep the artifact root narrow and not writable by untrusted users.

Model override fields:

Field Description
endpoint Runtime endpoint override.
model Runtime model override.
temperature Number in range 0..2. Explicit 0 is an override.
max_tokens Integer greater than or equal to 0. Explicit 0 is an override.
top_p Number in range 0..1. Explicit 0 is an override.
timeout_seconds Integer greater than or equal to 0. Explicit 0 disables the outbound client timeout.
service_tier Provider-specific request tier.
reasoning_effort Provider-specific reasoning setting.
api_key_env Name of an environment variable containing the API key.
extra_params JSON-compatible provider-specific top-level request fields.

Raw API-key values are not accepted in HTTP payloads. A field such as api_key is rejected as unknown JSON.

extra_params keys must not be empty and must not collide with reserved outbound fields: model, session_id, messages, temperature, max_tokens, top_p, service_tier, reasoning_effort, or response_format.

Strict JSON Rules

Request decoding is strict:

  • malformed JSON returns 400 invalid_json
  • unknown request fields return 400 invalid_json
  • unknown inputs item fields return 400 invalid_json
  • unknown model fields return 400 invalid_json
  • trailing JSON tokens after the request object return 400 invalid_json
  • request bodies above the configured limit return 413 request_too_large

Success Response

Status: 200 OK

{
  "artifact": {
    "name": "output",
    "content_type": "text/markdown",
    "body": "Generated content",
    "size": 17,
    "hash": "..."
  },
  "validation": {
    "status": "passed",
    "mode": "basic",
    "repair_attempts": 0,
    "is_valid": true
  },
  "metadata": {
    "run_id": "...",
    "prompt_id": "generic.markdown_summary",
    "prompt_version": "1.0.0",
    "prompt_hash": "...",
    "rendered_prompt_hash": "...",
    "selected_profile_id": "local-fast",
    "model_name": "gpt-4o-mini",
    "endpoint": "http://localhost:8000/v1",
    "model_params": {
      "endpoint": "http://localhost:8000/v1",
      "model": "gpt-4o-mini",
      "temperature": 0.2,
      "max_tokens": 500,
      "top_p": 1,
      "timeout_seconds": 90
    },
    "input_hashes": {
      "transcript": "..."
    },
    "usage": {
      "prompt_tokens": 11,
      "completion_tokens": 22,
      "total_tokens": 33,
      "cached_tokens": 0,
      "cache_write_tokens": 0
    },
    "start_time": "2026-05-04T12:00:00Z",
    "end_time": "2026-05-04T12:00:01Z",
    "duration_ms": 1000,
    "validation_mode": "basic",
    "validation_status": "passed",
    "repair_attempts_used": 0
  }
}

Response fields:

  • artifact: generated output artifact.
  • validation: validation result for the generated artifact.
  • metadata: run and effective runtime metadata.
  • raw_model_output: omitted unless include_raw_output is true.

artifact.uri is omitted when empty. validation.errors and validation.schema_path are omitted when empty. model_params.service_tier, model_params.reasoning_effort, model_params.api_key_env, and model_params.extra_params are omitted when empty.

metadata.usage.cached_tokens and metadata.usage.cache_write_tokens are always present as numbers. They are 0 when the provider omits compatible cache usage fields or reports no cache activity.

Validation Failure Response

Generated-content validation failures still return 200 OK.

{
  "validation": {
    "status": "failed",
    "mode": "json",
    "errors": ["invalid JSON: ..."],
    "repair_attempts": 0,
    "is_valid": false
  }
}

The response still includes artifact and metadata.

Error Responses

Error body shape:

{
  "error": {
    "code": "invalid_request",
    "message": "prompt_id is required"
  }
}

Current status/code mapping:

Status Code Meaning
400 invalid_json Malformed JSON, unknown JSON field, or trailing JSON token.
400 invalid_request Missing/invalid request fields or invalid runtime overrides.
400 profile_required No profile_id and prompt has no default_profile.
400 prompt_load_failed Prompt definition YAML/contract failed to load.
400 profile_load_failed Profile YAML/contract failed to load, including raw api_key.
400 artifact_not_allowed HTTP file refs are disabled or requested path is outside artifact root.
400 artifact_read_failed Input artifact could not be read or input ref was unsupported/invalid.
400 prompt_render_failed Prompt template rendering failed.
400 api_key_env_missing Selected api_key_env variable is unset or empty.
404 not_found Route path is unknown.
404 prompt_not_found Prompt ID/version was not found.
404 profile_not_found Profile ID was not found.
405 method_not_allowed Method is not POST on /v1/runs.
413 request_too_large Encoded JSON request body exceeds configured request limit.
413 artifact_too_large HTTP file input artifact exceeds configured artifact limit.
413 response_too_large Encoded JSON response exceeds configured response limit.
500 validation_runtime_failed Validator runtime/schema loading failed.
500 internal_error Unclassified server error.
502 llm_failed Outbound model request failed.

HTTP error messages are intentionally concise and do not include sensitive internal causes.

Retry And Idempotency

Scriptorium does not provide idempotency keys, pagination, caching headers, or rate limiting.

Clients may retry transport failures or 5xx responses when their surrounding workflow can tolerate another model call. A retry can generate different output and incur another provider request.

Example File

  • examples/http-run.json