Clarify HTTP operations documentation
This commit is contained in:
282
docs/api.md
282
docs/api.md
@@ -1,40 +1,68 @@
|
||||
# HTTP API Reference
|
||||
|
||||
## Scope
|
||||
This is the canonical public HTTP contract for Scriptorium.
|
||||
|
||||
This document is the canonical public HTTP contract for Scriptorium.
|
||||
|
||||
Current scope is only:
|
||||
Implemented route:
|
||||
|
||||
- `POST /v1/runs`
|
||||
|
||||
For CLI behavior, see the [CLI reference](cli.md).
|
||||
For CLI behavior, see [CLI reference](cli.md). For config and prompt/profile
|
||||
file formats, see [Configuration reference](config.md).
|
||||
|
||||
## Endpoint
|
||||
## Base URL And Deployment
|
||||
|
||||
- Method: `POST`
|
||||
- Path: `/v1/runs`
|
||||
- Content type: JSON request/response
|
||||
`scriptorium serve` listens on `server.addr` or `serve --addr`. The default is
|
||||
`:8080`.
|
||||
|
||||
Route behavior:
|
||||
The route path is always:
|
||||
|
||||
- unknown path: `404 not_found`
|
||||
- unsupported method on `/v1/runs`: `405 method_not_allowed`
|
||||
```text
|
||||
/v1/runs
|
||||
```
|
||||
|
||||
Copyable request example file:
|
||||
The HTTP adapter has no built-in authentication or authorization. Deploy it
|
||||
behind trusted network and authentication controls.
|
||||
|
||||
- `examples/http-run.json`
|
||||
## Media Types
|
||||
|
||||
## Request Body
|
||||
- Request body: JSON object.
|
||||
- Response body: JSON object.
|
||||
- Response `Content-Type`: `application/json`.
|
||||
|
||||
Requests are decoded as JSON regardless of the request `Content-Type` header.
|
||||
There are no shared query parameters.
|
||||
|
||||
## Request Limits
|
||||
|
||||
HTTP limits are configured through `server.*` config fields or `serve` flags:
|
||||
|
||||
- `server.max_request_bytes`: encoded JSON request body limit, including inline input bodies.
|
||||
- `server.max_artifact_bytes`: file artifact limit for HTTP `file` input references.
|
||||
- `server.max_response_bytes`: encoded JSON response limit, including artifact body and optional raw output.
|
||||
|
||||
Each limit defaults to `16777216` bytes. `0` disables that limit.
|
||||
|
||||
## `POST /v1/runs`
|
||||
|
||||
Runs one prompt request and returns the generated artifact, validation result,
|
||||
and metadata.
|
||||
|
||||
### Request Body
|
||||
|
||||
```json
|
||||
{
|
||||
"prompt_id": "generic.structured_events",
|
||||
"profile_id": "local-quality",
|
||||
"prompt_id": "generic.markdown_summary",
|
||||
"profile_id": "local-fast",
|
||||
"prompt_version": "1.0.0",
|
||||
"inputs": {
|
||||
"transcript": {"type": "file", "uri": "./examples/fixtures/transcript.md"},
|
||||
"glossary": {"type": "inline", "body": "party:\n - Rin"}
|
||||
"transcript": {
|
||||
"type": "file",
|
||||
"uri": "./examples/fixtures/transcript.md"
|
||||
},
|
||||
"glossary": {
|
||||
"type": "inline",
|
||||
"body": "party:\n - Rin"
|
||||
}
|
||||
},
|
||||
"vars": {
|
||||
"session_date": "2026-05-04"
|
||||
@@ -42,112 +70,120 @@ Copyable request example file:
|
||||
"model": {
|
||||
"endpoint": "http://localhost:8000/v1",
|
||||
"model": "gpt-4o-mini",
|
||||
"temperature": 0.0,
|
||||
"temperature": 0,
|
||||
"max_tokens": 800,
|
||||
"top_p": 1.0,
|
||||
"top_p": 1,
|
||||
"timeout_seconds": 120,
|
||||
"service_tier": "priority",
|
||||
"reasoning_effort": "medium",
|
||||
"api_key_env": "SCRIPTORIUM_API_KEY",
|
||||
"extra_params": {
|
||||
"route": "primary",
|
||||
"provider_options": {
|
||||
"retry_budget": 2
|
||||
}
|
||||
"provider_option": "enabled"
|
||||
}
|
||||
},
|
||||
"include_raw_output": false
|
||||
}
|
||||
```
|
||||
|
||||
Required fields:
|
||||
Request fields:
|
||||
|
||||
- `prompt_id`
|
||||
- `inputs` (must contain at least one named input)
|
||||
| Field | Required | Description |
|
||||
| --- | --- | --- |
|
||||
| `prompt_id` | yes | Prompt ID. Must not be blank. |
|
||||
| `prompt_version` | no | Prompt version filter. |
|
||||
| `profile_id` | no | Execution profile ID. If omitted, the prompt must define `default_profile`. |
|
||||
| `inputs` | yes | Object mapping prompt input names to input references. Must contain at least one entry. |
|
||||
| `vars` | no | Object mapping template variable names to string values. |
|
||||
| `model` | no | Runtime model override object. |
|
||||
| `include_raw_output` | no | When `true`, include `raw_model_output` in the response. |
|
||||
|
||||
Input reference types currently supported by runtime artifact loading:
|
||||
Input reference fields:
|
||||
|
||||
- `file`
|
||||
- `inline`
|
||||
| Field | Required | Description |
|
||||
| --- | --- | --- |
|
||||
| `type` | yes | `file` or `inline`. |
|
||||
| `uri` | for `file` | File URI/path. |
|
||||
| `body` | for `inline` | Inline artifact body. |
|
||||
|
||||
HTTP `file` references require `server.artifact_root` or `serve --artifact-root`.
|
||||
Relative file URIs resolve against that root. Absolute file URIs are accepted
|
||||
only when they are lexically inside the root. Requests that escape the root by
|
||||
lexical traversal, including `..` traversal and absolute paths outside the root,
|
||||
return `400 artifact_not_allowed`. Symlinks inside the root are followed by the
|
||||
operating system, including symlinks that point outside the root. The artifact
|
||||
root must not be writable by untrusted users. `inline` references do not require
|
||||
an artifact root.
|
||||
HTTP file artifacts above the configured artifact limit return
|
||||
`413 artifact_too_large`. Inline bodies are bounded by the request body limit.
|
||||
HTTP `file` references require `server.artifact_root` or `serve
|
||||
--artifact-root`. Relative file URIs resolve against that root. Absolute file
|
||||
URIs are accepted only when lexically inside the root. Relative traversal and
|
||||
absolute paths outside the root return `400 artifact_not_allowed`.
|
||||
|
||||
Model override notes:
|
||||
The containment check is lexical and does not resolve symlinks. Symlinks inside
|
||||
the artifact root are followed by the operating system, including symlinks that
|
||||
point outside the root. Keep the artifact root narrow and not writable by
|
||||
untrusted users.
|
||||
|
||||
- Numeric model override fields distinguish omitted values from explicit zero values. For example, omitting `temperature` preserves the selected profile/default value, while `"temperature": 0` explicitly sets the effective temperature to zero.
|
||||
- `extra_params` accepts JSON-compatible values: strings, numbers, booleans, objects, and arrays.
|
||||
- `extra_params` are passed through effective model metadata and flattened into top-level provider request fields by the OpenAI-compatible client.
|
||||
- `extra_params` keys must not be empty and must not collide with reserved outbound fields: `model`, `session_id`, `messages`, `temperature`, `max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or `response_format`.
|
||||
- Raw API-key values are not accepted. Use `api_key_env` to name an environment variable.
|
||||
Model override fields:
|
||||
|
||||
## Strict JSON Rules
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `endpoint` | Runtime endpoint override. |
|
||||
| `model` | Runtime model override. |
|
||||
| `temperature` | Number in range `0..2`. Explicit `0` is an override. |
|
||||
| `max_tokens` | Integer greater than or equal to `0`. Explicit `0` is an override. |
|
||||
| `top_p` | Number in range `0..1`. Explicit `0` is an override. |
|
||||
| `timeout_seconds` | Integer greater than or equal to `0`. Explicit `0` disables the outbound client timeout. |
|
||||
| `service_tier` | Provider-specific request tier. |
|
||||
| `reasoning_effort` | Provider-specific reasoning setting. |
|
||||
| `api_key_env` | Name of an environment variable containing the API key. |
|
||||
| `extra_params` | JSON-compatible provider-specific top-level request fields. |
|
||||
|
||||
Request decoding uses strict JSON field checks:
|
||||
Raw API-key values are not accepted in HTTP payloads. A field such as
|
||||
`api_key` is rejected as unknown JSON.
|
||||
|
||||
- unknown request fields are rejected with `400 invalid_json`
|
||||
- unknown `model` fields are rejected with `400 invalid_json`
|
||||
- raw API-key payload fields such as `api_key` are rejected as unknown fields
|
||||
- request bodies above the configured request limit are rejected with `413 request_too_large`
|
||||
- trailing JSON tokens after the request object are rejected with `400 invalid_json`
|
||||
`extra_params` keys must not be empty and must not collide with reserved
|
||||
outbound fields: `model`, `session_id`, `messages`, `temperature`,
|
||||
`max_tokens`, `top_p`, `service_tier`, `reasoning_effort`, or
|
||||
`response_format`.
|
||||
|
||||
## Success Response
|
||||
### Strict JSON Rules
|
||||
|
||||
Request decoding is strict:
|
||||
|
||||
- malformed JSON returns `400 invalid_json`
|
||||
- unknown request fields return `400 invalid_json`
|
||||
- unknown `inputs` item fields return `400 invalid_json`
|
||||
- unknown `model` fields return `400 invalid_json`
|
||||
- trailing JSON tokens after the request object return `400 invalid_json`
|
||||
- request bodies above the configured limit return `413 request_too_large`
|
||||
|
||||
### Success Response
|
||||
|
||||
Status: `200 OK`
|
||||
|
||||
Response shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"artifact": {
|
||||
"name": "output",
|
||||
"content_type": "application/json",
|
||||
"body": "{\"summary\":\"...\"}",
|
||||
"uri": "",
|
||||
"size": 123,
|
||||
"content_type": "text/markdown",
|
||||
"body": "Generated content",
|
||||
"size": 17,
|
||||
"hash": "..."
|
||||
},
|
||||
"validation": {
|
||||
"status": "passed",
|
||||
"mode": "json_schema",
|
||||
"errors": [],
|
||||
"schema_path": "structured_events.schema.json",
|
||||
"mode": "basic",
|
||||
"repair_attempts": 0,
|
||||
"is_valid": true
|
||||
},
|
||||
"metadata": {
|
||||
"run_id": "...",
|
||||
"prompt_id": "generic.structured_events",
|
||||
"prompt_id": "generic.markdown_summary",
|
||||
"prompt_version": "1.0.0",
|
||||
"prompt_hash": "...",
|
||||
"rendered_prompt_hash": "...",
|
||||
"selected_profile_id": "local-quality",
|
||||
"selected_profile_id": "local-fast",
|
||||
"model_name": "gpt-4o-mini",
|
||||
"endpoint": "http://localhost:8000/v1",
|
||||
"model_params": {
|
||||
"endpoint": "http://localhost:8000/v1",
|
||||
"model": "gpt-4o-mini",
|
||||
"temperature": 0,
|
||||
"max_tokens": 800,
|
||||
"temperature": 0.2,
|
||||
"max_tokens": 500,
|
||||
"top_p": 1,
|
||||
"timeout_seconds": 120,
|
||||
"service_tier": "priority",
|
||||
"reasoning_effort": "medium",
|
||||
"api_key_env": "SCRIPTORIUM_API_KEY",
|
||||
"extra_params": {
|
||||
"route": "primary",
|
||||
"provider_options": {
|
||||
"retry_budget": 2
|
||||
}
|
||||
}
|
||||
"timeout_seconds": 90
|
||||
},
|
||||
"input_hashes": {
|
||||
"transcript": "..."
|
||||
@@ -162,30 +198,46 @@ Response shape:
|
||||
"start_time": "2026-05-04T12:00:00Z",
|
||||
"end_time": "2026-05-04T12:00:01Z",
|
||||
"duration_ms": 1000,
|
||||
"validation_mode": "json_schema",
|
||||
"validation_mode": "basic",
|
||||
"validation_status": "passed",
|
||||
"repair_attempts_used": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`raw_model_output` is omitted by default.
|
||||
Response fields:
|
||||
|
||||
`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are always present as numbers. They are `0` when the provider omits compatible cache usage fields or reports no cache activity.
|
||||
- `artifact`: generated output artifact.
|
||||
- `validation`: validation result for the generated artifact.
|
||||
- `metadata`: run and effective runtime metadata.
|
||||
- `raw_model_output`: omitted unless `include_raw_output` is `true`.
|
||||
|
||||
To include it, send:
|
||||
`artifact.uri` is omitted when empty. `validation.errors` and
|
||||
`validation.schema_path` are omitted when empty. `model_params.service_tier`,
|
||||
`model_params.reasoning_effort`, `model_params.api_key_env`, and
|
||||
`model_params.extra_params` are omitted when empty.
|
||||
|
||||
- `"include_raw_output": true`
|
||||
`metadata.usage.cached_tokens` and `metadata.usage.cache_write_tokens` are
|
||||
always present as numbers. They are `0` when the provider omits compatible cache
|
||||
usage fields or reports no cache activity.
|
||||
|
||||
## Validation Failure Behavior
|
||||
### Validation Failure Response
|
||||
|
||||
Validation content failures do not map to HTTP error status.
|
||||
Generated-content validation failures still return `200 OK`.
|
||||
|
||||
Behavior:
|
||||
```json
|
||||
{
|
||||
"validation": {
|
||||
"status": "failed",
|
||||
"mode": "json",
|
||||
"errors": ["invalid JSON: ..."],
|
||||
"repair_attempts": 0,
|
||||
"is_valid": false
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- status remains `200 OK`
|
||||
- `validation.status` is `failed`
|
||||
- validation errors are returned in `validation.errors`
|
||||
The response still includes `artifact` and `metadata`.
|
||||
|
||||
## Error Responses
|
||||
|
||||
@@ -200,28 +252,38 @@ Error body shape:
|
||||
}
|
||||
```
|
||||
|
||||
Current error mapping (non-exhaustive):
|
||||
Current status/code mapping:
|
||||
|
||||
- `400 invalid_json`: malformed JSON or unknown JSON fields
|
||||
- `400 invalid_request`: missing/invalid request fields
|
||||
- `400 profile_required`: no explicit `profile_id` and prompt has no `default_profile`
|
||||
- `400 prompt_load_failed`: prompt definition invalid/unloadable
|
||||
- `400 profile_load_failed`: profile invalid/unloadable
|
||||
- `400 artifact_not_allowed`: file input artifact is outside the configured artifact root or file refs are not enabled
|
||||
- `400 artifact_read_failed`: input artifact loading failed
|
||||
- `400 prompt_render_failed`: template render failed
|
||||
- `400 api_key_env_missing`: named API-key environment variable is missing
|
||||
- `413 request_too_large`: request body exceeds the configured request limit
|
||||
- `413 artifact_too_large`: HTTP file input artifact exceeds the configured artifact limit
|
||||
- `413 response_too_large`: encoded JSON response exceeds the configured response limit
|
||||
- `404 prompt_not_found`
|
||||
- `404 profile_not_found`
|
||||
- `502 llm_failed`: outbound model request failed
|
||||
- `500 validation_runtime_failed`: validator runtime/schema-load failure
|
||||
- `500 internal_error`
|
||||
| Status | Code | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `400` | `invalid_json` | Malformed JSON, unknown JSON field, or trailing JSON token. |
|
||||
| `400` | `invalid_request` | Missing/invalid request fields or invalid runtime overrides. |
|
||||
| `400` | `profile_required` | No `profile_id` and prompt has no `default_profile`. |
|
||||
| `400` | `prompt_load_failed` | Prompt definition YAML/contract failed to load. |
|
||||
| `400` | `profile_load_failed` | Profile YAML/contract failed to load, including raw `api_key`. |
|
||||
| `400` | `artifact_not_allowed` | HTTP file refs are disabled or requested path is outside artifact root. |
|
||||
| `400` | `artifact_read_failed` | Input artifact could not be read or input ref was unsupported/invalid. |
|
||||
| `400` | `prompt_render_failed` | Prompt template rendering failed. |
|
||||
| `400` | `api_key_env_missing` | Selected `api_key_env` variable is unset or empty. |
|
||||
| `404` | `not_found` | Route path is unknown. |
|
||||
| `404` | `prompt_not_found` | Prompt ID/version was not found. |
|
||||
| `404` | `profile_not_found` | Profile ID was not found. |
|
||||
| `405` | `method_not_allowed` | Method is not `POST` on `/v1/runs`. |
|
||||
| `413` | `request_too_large` | Encoded JSON request body exceeds configured request limit. |
|
||||
| `413` | `artifact_too_large` | HTTP file input artifact exceeds configured artifact limit. |
|
||||
| `413` | `response_too_large` | Encoded JSON response exceeds configured response limit. |
|
||||
| `500` | `validation_runtime_failed` | Validator runtime/schema loading failed. |
|
||||
| `500` | `internal_error` | Unclassified server error. |
|
||||
| `502` | `llm_failed` | Outbound model request failed. |
|
||||
|
||||
## Security And Deployment Note
|
||||
HTTP error messages are intentionally concise and do not include sensitive
|
||||
internal causes.
|
||||
|
||||
The HTTP adapter has no built-in authentication or authorization.
|
||||
## Retry And Idempotency
|
||||
|
||||
Deploy behind trusted controls (for example authenticated gateway/reverse proxy and network boundaries).
|
||||
Scriptorium does not provide idempotency keys, pagination, caching headers, or
|
||||
rate limiting.
|
||||
|
||||
Clients may retry transport failures or `5xx` responses when their surrounding
|
||||
workflow can tolerate another model call. A retry can generate different output
|
||||
and incur another provider request.
|
||||
|
||||
Reference in New Issue
Block a user