Updated documentation to reflect current build state
This commit is contained in:
334
README.md
334
README.md
@@ -2,24 +2,87 @@
|
|||||||
|
|
||||||
Scriptorium is a generic prompt-profile execution engine written in Go.
|
Scriptorium is a generic prompt-profile execution engine written in Go.
|
||||||
|
|
||||||
It loads a prompt profile, resolves named input artifacts, renders a prompt, calls an OpenAI-compatible LLM endpoint, validates output, and returns an artifact plus metadata.
|
Given named input artifacts and a prompt profile, Scriptorium:
|
||||||
|
|
||||||
## Relationship to Narratio
|
1. Loads the profile.
|
||||||
|
2. Resolves input artifact references.
|
||||||
|
3. Renders prompt messages from templates.
|
||||||
|
4. Calls an OpenAI-compatible LLM endpoint.
|
||||||
|
5. Validates output if configured.
|
||||||
|
6. Optionally performs bounded structured-output repair.
|
||||||
|
7. Returns a generated artifact plus run metadata.
|
||||||
|
|
||||||
In the broader workflow, Narratio handles pipeline orchestration (transcription, cleanup, storage, notifications). Scriptorium handles only prompt-profile execution for a single run.
|
## Where Scriptorium Fits
|
||||||
|
|
||||||
## Repository Example Assets
|
Scriptorium is not an orchestrator.
|
||||||
|
|
||||||
- Profiles: `profiles/`
|
In the D&D workflow:
|
||||||
- Schemas: `schemas/`
|
|
||||||
- Tiny fixtures: `examples/fixtures/`
|
|
||||||
|
|
||||||
Included profiles:
|
- Narratio orchestrates the full pipeline.
|
||||||
- `generic.markdown_summary`
|
- WhisperX transcribes audio.
|
||||||
- `dnd.session_recap` (example content only; no D&D-specific Go logic)
|
- Seriatim merges transcripts.
|
||||||
- `generic.structured_events` (JSON + JSON Schema validation)
|
- Audita polishes transcripts.
|
||||||
|
- Scriptorium generates final artifacts from prepared inputs.
|
||||||
|
|
||||||
## Run a Local Profile (CLI)
|
D&D-specific behavior belongs in profiles, schemas, fixtures, and caller inputs, not in core Go logic.
|
||||||
|
|
||||||
|
## Core Concepts
|
||||||
|
|
||||||
|
- Prompt profile: YAML config that defines templates, model defaults, output format, and validation behavior.
|
||||||
|
- Named inputs: logical input names (for example `transcript`, `glossary`) mapped to artifact references.
|
||||||
|
- Artifact refs: currently `file` and `inline` are supported by readers used in v1 flows.
|
||||||
|
- Template variables: key/value vars provided at run time and accessed in templates as `{{.var_name}}`.
|
||||||
|
- Model target: endpoint/model and generation parameters (`temperature`, `max_tokens`, `top_p`, `timeout_seconds`).
|
||||||
|
- Output format: `text`, `markdown`, or `json`.
|
||||||
|
- Validation mode: `none`, `basic`, `json`, `json_schema`.
|
||||||
|
- Repair attempts: bounded retries for structured modes (`json`, `json_schema`) when output validation fails.
|
||||||
|
- Run metadata: IDs/hashes/model/timing/usage/validation details for auditability.
|
||||||
|
|
||||||
|
## Build and Test
|
||||||
|
|
||||||
|
Build:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
go build -o scriptorium ./cmd/scriptorium
|
||||||
|
```
|
||||||
|
|
||||||
|
Run tests:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
go test ./...
|
||||||
|
```
|
||||||
|
|
||||||
|
Run CLI locally:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
go run ./cmd/scriptorium run --help
|
||||||
|
```
|
||||||
|
|
||||||
|
## CLI Usage
|
||||||
|
|
||||||
|
### `scriptorium run`
|
||||||
|
|
||||||
|
Required flags:
|
||||||
|
|
||||||
|
- `--profile-dir`
|
||||||
|
- `--profile-id`
|
||||||
|
- `--input` (repeatable `name=path`)
|
||||||
|
|
||||||
|
Optional flags:
|
||||||
|
|
||||||
|
- `--var` (repeatable `name=value`)
|
||||||
|
- `--out`
|
||||||
|
- `--llm-base-url`
|
||||||
|
- `--llm-api-key`
|
||||||
|
- `--model`
|
||||||
|
- `--temperature`
|
||||||
|
- `--max-tokens`
|
||||||
|
- `--schema-dir`
|
||||||
|
- `--timeout`
|
||||||
|
|
||||||
|
If `--llm-base-url` and/or `--model` are omitted, profile `model_defaults` must provide them.
|
||||||
|
|
||||||
|
Markdown summary example:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
go run ./cmd/scriptorium run \
|
go run ./cmd/scriptorium run \
|
||||||
@@ -30,81 +93,222 @@ go run ./cmd/scriptorium run \
|
|||||||
--out ./out.md
|
--out ./out.md
|
||||||
```
|
```
|
||||||
|
|
||||||
This relies on `model_defaults.endpoint` and `model_defaults.model` in the selected profile.
|
Same run with explicit local OpenAI-compatible endpoint (for example vLLM):
|
||||||
You can override either at runtime with `--llm-base-url` and/or `--model`.
|
|
||||||
|
|
||||||
For schema-validated JSON output:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
go run ./cmd/scriptorium run \
|
go run ./cmd/scriptorium run \
|
||||||
--profile-dir ./profiles \
|
--profile-dir ./profiles \
|
||||||
--profile-id generic.structured_events \
|
--profile-id generic.markdown_summary \
|
||||||
--input transcript=./examples/fixtures/transcript.md \
|
--input transcript=./examples/fixtures/transcript.md \
|
||||||
--input glossary=./examples/fixtures/glossary.yml \
|
--input glossary=./examples/fixtures/glossary.yml \
|
||||||
--llm-base-url http://localhost:8000/v1 \
|
--llm-base-url http://localhost:8000/v1 \
|
||||||
--model gpt-4o-mini \
|
--model gpt-4o-mini \
|
||||||
--schema-dir ./schemas \
|
--out ./out.md
|
||||||
--out ./events.json
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Start Local HTTP API
|
Passing template variables:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
go run ./cmd/scriptorium serve \
|
go run ./cmd/scriptorium run \
|
||||||
--addr :8080 \
|
|
||||||
--profile-dir ./profiles \
|
--profile-dir ./profiles \
|
||||||
--schema-dir ./schemas \
|
--profile-id generic.markdown_summary \
|
||||||
--llm-base-url http://localhost:8000/v1 \
|
--input transcript=./examples/fixtures/transcript.md \
|
||||||
--model gpt-4o-mini
|
--input glossary=./examples/fixtures/glossary.yml \
|
||||||
|
--var session_date=2026-05-04 \
|
||||||
|
--var facilitator="Eris" \
|
||||||
|
--out ./out.md
|
||||||
```
|
```
|
||||||
|
|
||||||
## Call `POST /v1/runs`
|
Output behavior:
|
||||||
|
|
||||||
```bash
|
- Artifact content goes to stdout unless `--out` is set.
|
||||||
curl -sS http://localhost:8080/v1/runs \
|
- Summaries and errors are written to stderr.
|
||||||
-H 'Content-Type: application/json' \
|
- Exit code `2` indicates run succeeded but validation status is `failed`.
|
||||||
-d '{
|
|
||||||
"profile_id": "generic.structured_events",
|
### `scriptorium serve`
|
||||||
"inputs": {
|
|
||||||
"transcript": {"type": "file", "uri": "./examples/fixtures/transcript.md"},
|
Starts HTTP API.
|
||||||
"glossary": {"type": "file", "uri": "./examples/fixtures/glossary.yml"}
|
|
||||||
},
|
Required flags:
|
||||||
"model": {"model": "gpt-4o-mini"}
|
|
||||||
}'
|
- `--profile-dir`
|
||||||
|
- `--llm-base-url`
|
||||||
|
|
||||||
|
Common optional flags:
|
||||||
|
|
||||||
|
- `--addr` (default `:8080`)
|
||||||
|
- `--schema-dir` (default `.`)
|
||||||
|
- `--llm-api-key`
|
||||||
|
- `--model`
|
||||||
|
- `--timeout` (default `10m`)
|
||||||
|
|
||||||
|
## HTTP API
|
||||||
|
|
||||||
|
Run endpoint:
|
||||||
|
|
||||||
|
- `POST /v1/runs`
|
||||||
|
|
||||||
|
Request example:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"profile_id": "generic.structured_events",
|
||||||
|
"profile_version": "1.0.0",
|
||||||
|
"inputs": {
|
||||||
|
"transcript": {"type": "file", "uri": "./examples/fixtures/transcript.md"},
|
||||||
|
"glossary": {"type": "file", "uri": "./examples/fixtures/glossary.yml"}
|
||||||
|
},
|
||||||
|
"vars": {
|
||||||
|
"session_date": "2026-05-04"
|
||||||
|
},
|
||||||
|
"model": {
|
||||||
|
"endpoint": "http://localhost:8000/v1",
|
||||||
|
"model": "gpt-4o-mini",
|
||||||
|
"temperature": 0.0,
|
||||||
|
"max_tokens": 600,
|
||||||
|
"top_p": 1.0,
|
||||||
|
"timeout_seconds": 120
|
||||||
|
}
|
||||||
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Response shape:
|
Response shape:
|
||||||
- `artifact`
|
|
||||||
- `validation`
|
|
||||||
- `metadata`
|
|
||||||
- `raw_model_output`
|
|
||||||
|
|
||||||
`metadata` includes stable audit fields such as run/profile IDs, profile hash, effective model params, prompt/input hashes, timing, usage, and validation summary.
|
```json
|
||||||
|
{
|
||||||
|
"artifact": {
|
||||||
|
"name": "output",
|
||||||
|
"content_type": "application/json",
|
||||||
|
"body": "{...}",
|
||||||
|
"uri": "",
|
||||||
|
"size": 123,
|
||||||
|
"hash": "..."
|
||||||
|
},
|
||||||
|
"validation": {
|
||||||
|
"status": "passed",
|
||||||
|
"mode": "json_schema",
|
||||||
|
"errors": [],
|
||||||
|
"schema_path": "structured_events.schema.json",
|
||||||
|
"repair_attempts": 0,
|
||||||
|
"is_valid": true
|
||||||
|
},
|
||||||
|
"metadata": {
|
||||||
|
"run_id": "xxxxxxxx-xxxx-4xxx-8xxx-xxxxxxxxxxxx",
|
||||||
|
"profile_id": "generic.structured_events",
|
||||||
|
"profile_version": "1.0.0",
|
||||||
|
"profile_hash": "...",
|
||||||
|
"model_name": "gpt-4o-mini",
|
||||||
|
"endpoint": "http://localhost:8000/v1",
|
||||||
|
"model_params": {
|
||||||
|
"endpoint": "http://localhost:8000/v1",
|
||||||
|
"model": "gpt-4o-mini",
|
||||||
|
"temperature": 0,
|
||||||
|
"max_tokens": 600,
|
||||||
|
"top_p": 1,
|
||||||
|
"timeout_seconds": 120
|
||||||
|
},
|
||||||
|
"input_hashes": {"transcript": "...", "glossary": "..."},
|
||||||
|
"prompt_hash": "...",
|
||||||
|
"usage": {"prompt_tokens": 10, "completion_tokens": 20, "total_tokens": 30},
|
||||||
|
"start_time": "...",
|
||||||
|
"end_time": "...",
|
||||||
|
"duration_ms": 1523,
|
||||||
|
"validation_mode": "json_schema",
|
||||||
|
"validation_status": "passed",
|
||||||
|
"repair_attempts_used": 0
|
||||||
|
},
|
||||||
|
"raw_model_output": "{...}"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
## Add a New Prompt Profile
|
Validation content failures are returned as successful run responses (`200`) with `validation.status = "failed"`; raw model output is preserved in `raw_model_output`.
|
||||||
|
|
||||||
1. Add a YAML file under `profiles/` with:
|
Error response shape:
|
||||||
- `id`, `version`, `expected_inputs`, `templates`, `model_defaults`, `output_format`, `validation`
|
|
||||||
- optional model timeout via `model_defaults.timeout_seconds` (per-run LLM timeout override)
|
|
||||||
2. Ensure endpoint/model are available from either:
|
|
||||||
- profile defaults (`model_defaults.endpoint`, `model_defaults.model`), or
|
|
||||||
- request overrides (`--llm-base-url`, `--model`, or HTTP `model.endpoint`/`model.model`).
|
|
||||||
3. Use template helpers such as `{{input "transcript"}}` and template vars like `{{.session_date}}`.
|
|
||||||
4. For structured JSON output, set:
|
|
||||||
- `output_format: json`
|
|
||||||
- `validation.validation_mode: json_schema`
|
|
||||||
- `validation.schema_path: <schema file>`
|
|
||||||
5. Place schema files in `schemas/` and pass `--schema-dir ./schemas` for CLI/serve.
|
|
||||||
6. `validation.repair_attempts` is bounded and applies only to structured modes (`json`, `json_schema`).
|
|
||||||
|
|
||||||
## Validation Behavior
|
```json
|
||||||
|
{
|
||||||
|
"error": {
|
||||||
|
"code": "artifact_read_failed",
|
||||||
|
"message": "failed to read input artifact"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
Validation modes currently implemented:
|
## Prompt Profile Authoring
|
||||||
- `none`
|
|
||||||
- `basic` (non-empty output)
|
|
||||||
- `json` (must parse as JSON)
|
|
||||||
- `json_schema` (must parse JSON and satisfy schema)
|
|
||||||
|
|
||||||
Important behavior:
|
### Minimal Markdown profile
|
||||||
- Validation content failures are returned as structured run results (`validation.status = failed`) and preserve `raw_model_output`.
|
|
||||||
- Validation runtime/configuration failures are treated as run errors.
|
```yaml
|
||||||
|
id: generic.markdown_summary
|
||||||
|
version: "1.0.0"
|
||||||
|
expected_inputs:
|
||||||
|
- transcript
|
||||||
|
templates:
|
||||||
|
- role: system
|
||||||
|
content: "You are a concise assistant."
|
||||||
|
- role: user
|
||||||
|
content: |
|
||||||
|
Summarize:
|
||||||
|
{{input "transcript"}}
|
||||||
|
model_defaults:
|
||||||
|
endpoint: http://localhost:8000/v1
|
||||||
|
model: gpt-4o-mini
|
||||||
|
temperature: 0.2
|
||||||
|
max_tokens: 700
|
||||||
|
output_format: markdown
|
||||||
|
validation:
|
||||||
|
validation_mode: basic
|
||||||
|
```
|
||||||
|
|
||||||
|
### Structured JSON profile with schema validation
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
id: generic.structured_events
|
||||||
|
version: "1.0.0"
|
||||||
|
expected_inputs:
|
||||||
|
- transcript
|
||||||
|
templates:
|
||||||
|
- role: system
|
||||||
|
content: "Return only JSON."
|
||||||
|
- role: user
|
||||||
|
content: |
|
||||||
|
Extract events from:
|
||||||
|
{{input "transcript"}}
|
||||||
|
model_defaults:
|
||||||
|
endpoint: http://localhost:8000/v1
|
||||||
|
model: gpt-4o-mini
|
||||||
|
output_format: json
|
||||||
|
validation:
|
||||||
|
format: json
|
||||||
|
validation_mode: json_schema
|
||||||
|
schema_path: structured_events.schema.json
|
||||||
|
repair_attempts: 1
|
||||||
|
```
|
||||||
|
|
||||||
|
`repair_attempts` is bounded. Repair is attempted only for structured validation modes.
|
||||||
|
|
||||||
|
## Validation Modes
|
||||||
|
|
||||||
|
Supported modes:
|
||||||
|
|
||||||
|
- `none`: skipped validation result.
|
||||||
|
- `basic`: fails if output is empty/whitespace.
|
||||||
|
- `json`: output must parse as JSON.
|
||||||
|
- `json_schema`: output must parse as JSON and satisfy the configured schema.
|
||||||
|
|
||||||
|
Validation failures caused by output content are represented in `validation` and do not discard raw model output.
|
||||||
|
|
||||||
|
## Repository Examples
|
||||||
|
|
||||||
|
- Profiles: `profiles/`
|
||||||
|
- Schemas: `schemas/`
|
||||||
|
- Fixtures: `examples/fixtures/`
|
||||||
|
- Local experimentation: `local-test/`
|
||||||
|
|
||||||
|
## Development Notes
|
||||||
|
|
||||||
|
- Core is generic and follows a ports-and-adapters style.
|
||||||
|
- Domain/usecase packages do not depend on HTTP/CLI/wire types.
|
||||||
|
- To add a new LLM adapter: implement `internal/llm.Client`.
|
||||||
|
- To add a new artifact reader: implement/extend `internal/artifact.Reader` routing.
|
||||||
|
- To add a new validation mode: extend `internal/validate` and keep run semantics stable.
|
||||||
|
|||||||
852
architecture.md
852
architecture.md
@@ -1,752 +1,232 @@
|
|||||||
# Scriptorium Architecture
|
# Scriptorium Architecture
|
||||||
|
|
||||||
## Purpose
|
## 1. Purpose and Non-Goals
|
||||||
|
|
||||||
Scriptorium is a general-purpose prompt-profile execution service.
|
Scriptorium is a prompt-profile execution engine.
|
||||||
|
|
||||||
Its job is to take one or more named input artifacts, render a configured prompt profile, execute that prompt against an LLM endpoint, optionally validate the output, and return a generated artifact with useful metadata.
|
It takes named input artifacts, renders prompt templates, calls an LLM, validates output, optionally performs bounded structured-output repair, and returns an artifact with metadata.
|
||||||
|
|
||||||
The initial concrete use case is generating artifacts from cleaned Dungeons & Dragons session transcripts, such as session recaps, player analysis, structured event extraction, and glossary update suggestions.
|
Scriptorium is not an orchestrator. It should not own transcription, transcript merging, transcript polishing, notification, or cross-step workflow control.
|
||||||
|
|
||||||
However, Scriptorium must not be D&D-specific. D&D behavior belongs in prompt profiles, schemas, and caller-provided inputs. The Go application should remain a generic engine for prompt execution and output validation.
|
For the motivating D&D workflow:
|
||||||
|
|
||||||
## Intended Audience
|
- Narratio orchestrates.
|
||||||
|
- WhisperX transcribes.
|
||||||
|
- Seriatim merges transcripts.
|
||||||
|
- Audita polishes transcripts.
|
||||||
|
- Scriptorium generates output artifacts from prepared inputs.
|
||||||
|
|
||||||
This document is written for 5.3-Codex and future maintainers.
|
Core Go code must remain domain-generic.
|
||||||
|
|
||||||
When implementing this repository, prefer simple, idiomatic Go over elaborate framework code. The architecture should be modular, testable, and composable, but not over-engineered.
|
## 2. Current Architecture
|
||||||
|
|
||||||
The desired implementation style is:
|
Current high-level structure:
|
||||||
|
|
||||||
- Clear domain types.
|
- `cmd/scriptorium`: binary entrypoint.
|
||||||
- Small interfaces at architectural boundaries.
|
- `internal/domain`: core domain types.
|
||||||
- Explicit dependencies.
|
- `internal/usecase`: `Runner` use case and repair loop orchestration.
|
||||||
- No hidden global state.
|
- `internal/profile`: filesystem prompt profile repository and profile validation.
|
||||||
- No domain-specific D&D logic in core packages.
|
- `internal/artifact`: artifact reference readers (`inline`, `file`) and routing.
|
||||||
- Practical hexagonal / ports-and-adapters structure.
|
- `internal/prompt`: Go template-based prompt renderer.
|
||||||
- Boring, inspectable behavior.
|
- `internal/llm`: LLM client interface + OpenAI-compatible HTTP adapter.
|
||||||
|
- `internal/validate`: output validator implementation (`none/basic/json/json_schema`).
|
||||||
|
- `internal/adapter/cli`: CLI adapter.
|
||||||
|
- `internal/adapter/http`: HTTP adapter (`POST /v1/runs`).
|
||||||
|
|
||||||
## Core Concept
|
This is a practical ports-and-adapters implementation.
|
||||||
|
|
||||||
Scriptorium transforms:
|
## 3. Run Data Flow
|
||||||
|
|
||||||
- Prompt profile
|
`Runner.Run(ctx, RunRequest)` currently executes:
|
||||||
- Named input artifacts
|
|
||||||
- Template variables
|
|
||||||
- Model target
|
|
||||||
- Optional output contract
|
|
||||||
|
|
||||||
Into:
|
1. Validate minimum request requirements (`profile_id`).
|
||||||
|
2. Load prompt profile by ID/version.
|
||||||
|
3. Merge effective model target (profile defaults + request override).
|
||||||
|
4. Resolve effective output contract (profile + optional request override).
|
||||||
|
5. Resolve named artifact refs to loaded artifacts.
|
||||||
|
6. Render prompt messages from templates.
|
||||||
|
7. Hash rendered prompt for auditability.
|
||||||
|
8. Call LLM client with provider-neutral `GenerateRequest`.
|
||||||
|
9. Build output artifact from model content.
|
||||||
|
10. Validate output.
|
||||||
|
11. If structured validation failed and repair is enabled/bounded, run repair attempts and re-validate.
|
||||||
|
12. Return `RunResult` with artifact, validation, raw output, and metadata.
|
||||||
|
|
||||||
- Generated artifact
|
Validation content failure remains a successful run result with `validation.status=failed`.
|
||||||
- Validation result
|
|
||||||
- Prompt/model/input metadata
|
|
||||||
- Raw model output
|
|
||||||
- Structured error details, if applicable
|
|
||||||
|
|
||||||
Scriptorium should be thought of as a deterministic wrapper around a nondeterministic model call.
|
## 4. Package Responsibilities
|
||||||
|
|
||||||
The system should make the model call as auditable and reproducible as possible, even though LLM output itself may not be exactly reproducible.
|
- `domain`
|
||||||
|
- Owns core nouns and contracts.
|
||||||
|
- Must not import adapters/provider SDK types.
|
||||||
|
|
||||||
## Application Boundary
|
- `usecase`
|
||||||
|
- Owns execution sequence and cross-port orchestration for a single run.
|
||||||
|
- May coordinate validation and bounded repair.
|
||||||
|
- Must not contain HTTP/CLI/wire concerns.
|
||||||
|
|
||||||
Scriptorium is not an orchestrator.
|
- `profile`
|
||||||
|
- Owns prompt profile loading/parsing/validation.
|
||||||
|
- Handles YAML strict decoding and profile-level constraints.
|
||||||
|
|
||||||
The broader workflow may include audio transcription, transcript merging, transcript polishing, artifact persistence, and notifications. Those responsibilities belong to the external orchestrator, currently expected to be Narratio.
|
- `artifact`
|
||||||
|
- Owns artifact ref resolution and content loading.
|
||||||
|
- Produces normalized `Artifact` values with size/hash/content type.
|
||||||
|
|
||||||
Scriptorium should not know about WhisperX, Seriatim, Audita, or any other pipeline stage.
|
- `prompt`
|
||||||
|
- Owns template rendering and required-input enforcement.
|
||||||
|
|
||||||
Scriptorium only knows how to:
|
- `llm`
|
||||||
|
- Owns generation port and provider adapters.
|
||||||
|
- Current adapter: OpenAI-compatible chat completions over `net/http`.
|
||||||
|
|
||||||
1. Load a prompt profile.
|
- `validate`
|
||||||
2. Load or receive named input artifacts.
|
- Owns output validation semantics and JSON Schema integration.
|
||||||
3. Render a prompt.
|
|
||||||
4. Call an LLM.
|
|
||||||
5. Validate the output, if configured.
|
|
||||||
6. Return an output artifact and metadata.
|
|
||||||
|
|
||||||
## Initial Workflow Context
|
- `adapter/http`, `adapter/cli`
|
||||||
|
- Owns transport/wire/flag concerns only.
|
||||||
|
- Should stay thin and delegate business flow to `usecase.Runner`.
|
||||||
|
|
||||||
The initial D&D workflow is expected to look like this:
|
## 5. Domain Model (Current)
|
||||||
|
|
||||||
1. Narratio transcribes audio tracks using WhisperX.
|
Key types in `internal/domain`:
|
||||||
2. Narratio normalizes speaker names and saves per-speaker transcripts.
|
|
||||||
3. Narratio calls Seriatim to merge transcripts.
|
|
||||||
4. Narratio saves the merged transcript.
|
|
||||||
5. Narratio calls Audita to polish the transcript.
|
|
||||||
6. Narratio saves the processed transcript.
|
|
||||||
7. Narratio calls Scriptorium one or more times to generate output artifacts.
|
|
||||||
8. Narratio saves each generated artifact.
|
|
||||||
9. Narratio optionally sends a completion notification.
|
|
||||||
|
|
||||||
Scriptorium only owns step 7.
|
- `RunRequest`: profile selector, named input refs, vars, optional model override, optional validation override.
|
||||||
|
- `RunResult`: output artifact, validation result, raw output, profile/model metadata, hashes, usage, timestamps, duration.
|
||||||
|
- `ArtifactRef`: `{type, uri, body}` reference contract.
|
||||||
|
- `Artifact`: loaded payload (`name`, `content_type`, `body`, `uri`, `size`, `hash`).
|
||||||
|
- `PromptProfile`: YAML-backed profile definition.
|
||||||
|
- `RenderedPrompt` / `RenderedMessage`: provider-neutral prompt structure.
|
||||||
|
- `GenerateRequest` / `GenerateResponse`: provider-neutral model I/O.
|
||||||
|
- `ValidationResult`: passed/failed/skipped + mode/errors/schema/repair attempts.
|
||||||
|
|
||||||
Each Scriptorium request should initially produce one artifact. If multiple artifacts are needed, the orchestrator should call Scriptorium multiple times.
|
## 6. Interfaces and Adapters
|
||||||
|
|
||||||
Batch execution can be added later, but should not be part of the core v1 design unless there is an immediate need.
|
Primary ports:
|
||||||
|
|
||||||
## Primary Use Cases
|
- `profile.Repository`
|
||||||
|
- `artifact.Reader`
|
||||||
|
- `prompt.Renderer`
|
||||||
|
- `llm.Client`
|
||||||
|
- `validate.Validator`
|
||||||
|
- `usecase.OutputRepairer` (usecase-local abstraction)
|
||||||
|
|
||||||
Scriptorium should support the following v1 use cases:
|
Current adapters:
|
||||||
|
|
||||||
1. Generate a freeform Markdown artifact from a transcript and prompt profile.
|
- Profile repository: filesystem YAML loader.
|
||||||
2. Generate a structured JSON artifact from a transcript and prompt profile.
|
- Artifact reader: composite reader for `inline` and `file`.
|
||||||
3. Validate JSON output against a JSON Schema.
|
- Prompt renderer: Go templates with input helper + vars.
|
||||||
4. Return raw model output when validation fails.
|
- LLM adapter: OpenAI-compatible `/chat/completions`.
|
||||||
5. Optionally attempt one bounded repair pass for invalid structured output.
|
- Validator: standard validator with `none/basic/json/json_schema`.
|
||||||
6. Record metadata about the profile, model, inputs, prompt hash, and validation result.
|
- CLI/HTTP adapters: thin request mapping and response mapping.
|
||||||
7. Support local development through a CLI.
|
|
||||||
8. Support service usage through an HTTP API.
|
|
||||||
|
|
||||||
## Non-Goals for v1
|
## 7. Validation and Repair Model
|
||||||
|
|
||||||
Do not implement these in the initial version unless explicitly requested:
|
|
||||||
|
|
||||||
- Multi-agent workflows.
|
|
||||||
- Arbitrary DAG execution.
|
|
||||||
- Long-running job queues.
|
|
||||||
- Automatic RAG.
|
|
||||||
- Automatic prompt chaining.
|
|
||||||
- Automatic chunking and summarization.
|
|
||||||
- Model selection logic.
|
|
||||||
- Complex retry policies beyond basic HTTP/model retry and optional validation repair.
|
|
||||||
- D&D-specific Go packages.
|
|
||||||
- UI.
|
|
||||||
- Database persistence.
|
|
||||||
- Full artifact lifecycle management.
|
|
||||||
|
|
||||||
These may be valid future features, but v1 should remain a focused prompt-profile execution engine.
|
|
||||||
|
|
||||||
## Architectural Style
|
|
||||||
|
|
||||||
Use a practical hexagonal architecture.
|
|
||||||
|
|
||||||
The core domain and use case packages should not depend on infrastructure details such as HTTP, S3, local filesystems, or specific LLM providers.
|
|
||||||
|
|
||||||
External concerns should be implemented as adapters.
|
|
||||||
|
|
||||||
The central use case should be easy to test with fake prompt repositories, fake artifact readers, fake LLM clients, and fake validators.
|
|
||||||
|
|
||||||
Recommended high-level structure:
|
|
||||||
|
|
||||||
- cmd/scriptorium: application entrypoint
|
|
||||||
- internal/domain: core domain types
|
|
||||||
- internal/usecase: application use cases
|
|
||||||
- internal/profile: prompt profile loading and parsing
|
|
||||||
- internal/prompt: prompt rendering
|
|
||||||
- internal/llm: LLM client interfaces and adapters
|
|
||||||
- internal/validate: output validation implementations
|
|
||||||
- internal/artifact: artifact loading and storage adapters
|
|
||||||
- internal/adapter/http: HTTP API
|
|
||||||
- internal/adapter/cli: CLI interface
|
|
||||||
- internal/config: application configuration
|
|
||||||
- profiles: example prompt profiles
|
|
||||||
- schemas: example output schemas
|
|
||||||
- testdata: fixtures for tests
|
|
||||||
|
|
||||||
Exact package names may evolve, but the boundary principles should remain stable.
|
|
||||||
|
|
||||||
## Domain Model
|
|
||||||
|
|
||||||
The core domain should include these concepts.
|
|
||||||
|
|
||||||
### RunRequest
|
|
||||||
|
|
||||||
Represents one request to generate one artifact.
|
|
||||||
|
|
||||||
Fields should include:
|
|
||||||
|
|
||||||
- ProfileID
|
|
||||||
- Inputs
|
|
||||||
- Vars
|
|
||||||
- Optional model override
|
|
||||||
- Optional validation override, if needed
|
|
||||||
- Optional caller metadata
|
|
||||||
|
|
||||||
Inputs should be keyed by logical input name, not by filename.
|
|
||||||
|
|
||||||
Example logical input names:
|
|
||||||
|
|
||||||
- transcript
|
|
||||||
- glossary
|
|
||||||
- previous_recap
|
|
||||||
- campaign_notes
|
|
||||||
- source_document
|
|
||||||
|
|
||||||
### ArtifactRef
|
|
||||||
|
|
||||||
Represents a reference to an input artifact.
|
|
||||||
|
|
||||||
Artifact references should support at least inline content and local file paths in v1.
|
|
||||||
|
|
||||||
S3 references may be supported in v1 if needed, but should be implemented behind an interface.
|
|
||||||
|
|
||||||
Likely artifact reference types:
|
|
||||||
|
|
||||||
- inline
|
|
||||||
- file
|
|
||||||
- s3
|
|
||||||
|
|
||||||
The core use case should not care which reference type is used.
|
|
||||||
|
|
||||||
### Artifact
|
|
||||||
|
|
||||||
Represents loaded content.
|
|
||||||
|
|
||||||
Fields should include:
|
|
||||||
|
|
||||||
- Name
|
|
||||||
- ContentType
|
|
||||||
- Body
|
|
||||||
- Optional URI or source reference
|
|
||||||
- Optional size
|
|
||||||
- Optional SHA-256 hash
|
|
||||||
|
|
||||||
Artifacts are the actual input and output payloads after references have been resolved.
|
|
||||||
|
|
||||||
### PromptProfile
|
|
||||||
|
|
||||||
Represents a configured prompt execution profile.
|
|
||||||
|
|
||||||
A profile should include:
|
|
||||||
|
|
||||||
- ID
|
|
||||||
- Version
|
|
||||||
- Description
|
|
||||||
- Expected inputs
|
|
||||||
- Prompt templates
|
|
||||||
- Model defaults
|
|
||||||
- Output format
|
|
||||||
- Optional validation configuration
|
|
||||||
- Optional repair configuration
|
|
||||||
|
|
||||||
Prompt profiles should be serializable from YAML.
|
|
||||||
|
|
||||||
Prompt profiles are where domain-specific behavior belongs.
|
|
||||||
|
|
||||||
### RenderedPrompt
|
|
||||||
|
|
||||||
Represents the prompt after input artifacts and variables have been applied.
|
|
||||||
|
|
||||||
For OpenAI-compatible chat models, this should contain a list of chat messages.
|
|
||||||
|
|
||||||
At minimum, support system and user messages.
|
|
||||||
|
|
||||||
Future support for developer messages, assistant prefill, or multimodal parts can be added later.
|
|
||||||
|
|
||||||
### ModelTarget
|
|
||||||
|
|
||||||
Represents the LLM endpoint and model configuration.
|
|
||||||
|
|
||||||
Fields should include:
|
|
||||||
|
|
||||||
- Endpoint name or URL
|
|
||||||
- Model name
|
|
||||||
- Temperature
|
|
||||||
- Max tokens
|
|
||||||
- Top-p, if supported
|
|
||||||
- Additional provider-specific options, if needed
|
|
||||||
|
|
||||||
For v1, the main adapter should support OpenAI-compatible chat completion APIs.
|
|
||||||
|
|
||||||
### RunResult
|
|
||||||
|
|
||||||
Represents the complete result of a run.
|
|
||||||
|
|
||||||
Fields should include:
|
|
||||||
|
|
||||||
- Generated artifact
|
|
||||||
- Raw model output
|
|
||||||
- Validation result
|
|
||||||
- Profile ID and version
|
|
||||||
- Model name
|
|
||||||
- Endpoint name
|
|
||||||
- Input hashes
|
|
||||||
- Prompt hash
|
|
||||||
- Token usage, if available
|
|
||||||
- Start and end timestamps
|
|
||||||
- Error details, if applicable
|
|
||||||
|
|
||||||
### ValidationResult
|
|
||||||
|
|
||||||
Represents validation status.
|
|
||||||
|
|
||||||
Fields should include:
|
|
||||||
|
|
||||||
- Status: passed, failed, skipped
|
|
||||||
- Validation mode
|
|
||||||
- Error messages
|
|
||||||
- Schema path, if applicable
|
|
||||||
- Repair attempts used
|
|
||||||
- Final output validity
|
|
||||||
|
|
||||||
## Core Interfaces
|
|
||||||
|
|
||||||
Keep interfaces small and focused.
|
|
||||||
|
|
||||||
### PromptRepository
|
|
||||||
|
|
||||||
Responsible for loading prompt profiles.
|
|
||||||
|
|
||||||
Expected implementations:
|
|
||||||
|
|
||||||
- FilesystemPromptRepository
|
|
||||||
- EmbeddedPromptRepository, optional later
|
|
||||||
- S3PromptRepository, optional later
|
|
||||||
|
|
||||||
The use case should ask for a profile by ID or ID plus version.
|
|
||||||
|
|
||||||
### ArtifactReader
|
|
||||||
|
|
||||||
Responsible for resolving ArtifactRef values into Artifact values.
|
|
||||||
|
|
||||||
Expected implementations:
|
|
||||||
|
|
||||||
- InlineArtifactReader
|
|
||||||
- FileArtifactReader
|
|
||||||
- S3ArtifactReader
|
|
||||||
- CompositeArtifactReader
|
|
||||||
|
|
||||||
The CompositeArtifactReader can route by reference type.
|
|
||||||
|
|
||||||
### PromptRenderer
|
|
||||||
|
|
||||||
Responsible for rendering prompt templates using named artifacts and variables.
|
|
||||||
|
|
||||||
Use Go templates unless there is a strong reason to choose something else.
|
|
||||||
|
|
||||||
Renderer responsibilities:
|
|
||||||
|
|
||||||
- Verify required inputs exist.
|
|
||||||
- Expose safe template functions.
|
|
||||||
- Insert artifact content by logical name.
|
|
||||||
- Render system and user prompt sections.
|
|
||||||
- Return a RenderedPrompt.
|
|
||||||
|
|
||||||
Do not silently omit missing required inputs.
|
|
||||||
|
|
||||||
Do not silently truncate large inputs in v1.
|
|
||||||
|
|
||||||
### LLMClient
|
|
||||||
|
|
||||||
Responsible for executing a rendered prompt against a model endpoint.
|
|
||||||
|
|
||||||
The initial implementation should support OpenAI-compatible chat completions.
|
|
||||||
|
|
||||||
This should work with:
|
|
||||||
|
|
||||||
- vLLM
|
|
||||||
- LiteLLM
|
|
||||||
- OpenAI-compatible local endpoints
|
|
||||||
- OpenAI-compatible hosted endpoints, if configured
|
|
||||||
|
|
||||||
The domain should not depend on provider-specific SDK types.
|
|
||||||
|
|
||||||
### OutputValidator
|
|
||||||
|
|
||||||
Responsible for validating the generated artifact.
|
|
||||||
|
|
||||||
Expected validation modes:
|
|
||||||
|
|
||||||
- none
|
|
||||||
- basic
|
|
||||||
- json_schema
|
|
||||||
|
|
||||||
Basic validation may check things like non-empty output, required headings, or forbidden boilerplate.
|
|
||||||
|
|
||||||
JSON Schema validation should parse the output as JSON and validate it against the configured schema.
|
|
||||||
|
|
||||||
### OutputRepairer
|
|
||||||
|
|
||||||
Responsible for making a bounded attempt to repair invalid structured output.
|
|
||||||
|
|
||||||
This should be optional.
|
|
||||||
|
|
||||||
The repairer may use the same LLMClient with a repair prompt.
|
|
||||||
|
|
||||||
Repair attempts must be bounded by configuration. Default should be zero or one.
|
|
||||||
|
|
||||||
Do not implement unbounded repair loops.
|
|
||||||
|
|
||||||
## Prompt Profiles
|
|
||||||
|
|
||||||
Prompt profiles are the main extension mechanism.
|
|
||||||
|
|
||||||
The Go application should stay generic. Prompt profiles should define domain behavior.
|
|
||||||
|
|
||||||
A profile should be able to specify:
|
|
||||||
|
|
||||||
- ID
|
|
||||||
- Version
|
|
||||||
- Description
|
|
||||||
- Required and optional inputs
|
|
||||||
- System prompt template
|
|
||||||
- User prompt template
|
|
||||||
- Default model configuration
|
|
||||||
- Output format
|
|
||||||
- Validation mode
|
|
||||||
- Schema path, if applicable
|
|
||||||
- Repair attempts, if applicable
|
|
||||||
|
|
||||||
Profiles should live outside compiled Go code.
|
|
||||||
|
|
||||||
Example profile categories for the initial D&D use case:
|
|
||||||
|
|
||||||
- dnd.session_recap
|
|
||||||
- dnd.meta_analysis
|
|
||||||
- dnd.table_read
|
|
||||||
- dnd.structured_events
|
|
||||||
- dnd.glossary_update_suggestions
|
|
||||||
|
|
||||||
The code should not special-case these names.
|
|
||||||
|
|
||||||
## Template Rendering
|
|
||||||
|
|
||||||
Prompt rendering must be predictable and explicit.
|
|
||||||
|
|
||||||
Templates should be able to reference:
|
|
||||||
|
|
||||||
- Named input artifacts
|
|
||||||
- Template variables
|
|
||||||
- Profile metadata
|
|
||||||
|
|
||||||
The renderer should provide a helper equivalent to input(name), which inserts the content of a named artifact.
|
|
||||||
|
|
||||||
The renderer should fail when:
|
|
||||||
|
|
||||||
- A required input is missing.
|
|
||||||
- A template references an unknown input.
|
|
||||||
- A template references a missing required variable.
|
|
||||||
- The rendered prompt exceeds a configured token or size limit, if such a limit is configured.
|
|
||||||
|
|
||||||
In v1, do not silently truncate inputs.
|
|
||||||
|
|
||||||
If token counting is not implemented initially, use byte-size limits or leave token budgeting as a clearly marked future improvement.
|
|
||||||
|
|
||||||
## Output Formats
|
|
||||||
|
|
||||||
Scriptorium should support at least these output formats:
|
|
||||||
|
|
||||||
- markdown
|
|
||||||
- text
|
|
||||||
- json
|
|
||||||
|
|
||||||
For markdown and text, validation may be skipped or basic.
|
|
||||||
|
|
||||||
For JSON, validation should at minimum require valid JSON. If a schema is configured, validate against the schema.
|
|
||||||
|
|
||||||
The output artifact should preserve content type.
|
|
||||||
|
|
||||||
Suggested content types:
|
|
||||||
|
|
||||||
- text/markdown
|
|
||||||
- text/plain
|
|
||||||
- application/json
|
|
||||||
|
|
||||||
## Validation
|
|
||||||
|
|
||||||
Validation should be explicit and profile-driven.
|
|
||||||
|
|
||||||
Validation modes:
|
Validation modes:
|
||||||
|
|
||||||
- none: no validation beyond successful generation
|
- `none`
|
||||||
- basic: simple textual validation
|
- `basic`
|
||||||
- json: parse as JSON
|
- `json`
|
||||||
- json_schema: parse as JSON and validate against schema
|
- `json_schema`
|
||||||
|
|
||||||
For invalid structured output, Scriptorium should return:
|
Repair behavior:
|
||||||
|
|
||||||
- Validation status
|
- Only applies to structured modes (`json`, `json_schema`).
|
||||||
- Validation errors
|
- Attempted only when validation fails, repairer exists, and `repair_attempts > 0`.
|
||||||
- Raw model output
|
- Bounded strictly by `repair_attempts`.
|
||||||
- Repair attempts used
|
- Uses a narrow JSON-repair prompt and re-validates each attempt.
|
||||||
- Final output, if repair succeeded
|
- If still invalid, run succeeds with failed validation and preserved final raw output.
|
||||||
|
- Validator runtime/config errors are run errors.
|
||||||
|
|
||||||
Validation failure should not discard the raw output.
|
## 8. Public Contracts
|
||||||
|
|
||||||
## Repair
|
### CLI
|
||||||
|
|
||||||
Repair is only for structured output.
|
Commands:
|
||||||
|
|
||||||
The initial repair use case is invalid JSON or JSON that fails schema validation.
|
- `scriptorium run`
|
||||||
|
- `scriptorium serve`
|
||||||
|
|
||||||
The repair prompt should be deterministic and narrow:
|
`run`:
|
||||||
|
|
||||||
- Explain that the previous output failed validation.
|
- Required: `--profile-dir`, `--profile-id`, `--input`.
|
||||||
- Provide validation errors.
|
- Optional: model/endpoint overrides (`--model`, `--llm-base-url`), vars, output path, schema dir, timeout.
|
||||||
- Provide the previous output.
|
- Artifact bytes go to stdout (or `--out` file); summaries/errors go to stderr.
|
||||||
- Ask the model to return only corrected JSON.
|
|
||||||
- Do not ask the model to improve the answer substantively.
|
|
||||||
|
|
||||||
Repair must be bounded.
|
`serve`:
|
||||||
|
|
||||||
Recommended default:
|
- Required: `--profile-dir`, `--llm-base-url`.
|
||||||
|
- Exposes HTTP run endpoint.
|
||||||
|
|
||||||
- repair_attempts: 0 for freeform output
|
### HTTP
|
||||||
- repair_attempts: 1 for JSON schema output, if configured
|
|
||||||
|
|
||||||
## LLM Adapter
|
- Endpoint: `POST /v1/runs`.
|
||||||
|
- Request maps to `RunRequest` (`profile_id`, `inputs`, `vars`, optional `model` override).
|
||||||
|
- Response includes `artifact`, `validation`, `metadata`, `raw_model_output`.
|
||||||
|
- Validation content failures are represented as `200` with `validation.status=failed`.
|
||||||
|
- Error responses are `{error:{code,message}}` with stable code mapping.
|
||||||
|
|
||||||
The initial LLM adapter should target OpenAI-compatible chat completions.
|
### Prompt Profile YAML
|
||||||
|
|
||||||
The adapter should support:
|
- `id`, `version`, `expected_inputs`, `templates`, `model_defaults`, `output_format`, `validation`.
|
||||||
|
- Strict YAML decoding (`KnownFields`) rejects unknown fields.
|
||||||
|
- `validation.schema_path` required when `validation_mode=json_schema`.
|
||||||
|
- `validation.repair_attempts` must be non-negative.
|
||||||
|
|
||||||
- Base URL
|
### Metadata
|
||||||
- API key, optional for local endpoints
|
|
||||||
- Model name
|
|
||||||
- Temperature
|
|
||||||
- Max tokens
|
|
||||||
- Basic generation parameters
|
|
||||||
- Request timeout
|
|
||||||
- Token usage extraction, if returned by the endpoint
|
|
||||||
|
|
||||||
Do not couple the core domain to OpenAI SDK request or response structs.
|
Current run metadata includes:
|
||||||
|
|
||||||
The adapter should translate between internal GenerateRequest / GenerateResponse types and the provider wire format.
|
- `run_id` (UUID v4)
|
||||||
|
- `profile_id`, `profile_version`, `profile_hash`
|
||||||
|
- `model_name`, `endpoint`, effective `model_params`
|
||||||
|
- `input_hashes`, `prompt_hash`
|
||||||
|
- token usage
|
||||||
|
- start/end timestamps
|
||||||
|
- duration
|
||||||
|
- validation mode/status
|
||||||
|
- repair attempts used
|
||||||
|
|
||||||
## HTTP API
|
## 9. Extension Points (Future Work)
|
||||||
|
|
||||||
The HTTP API should be thin.
|
Future features should plug into existing boundaries, not bypass them.
|
||||||
|
|
||||||
It should translate HTTP requests into RunRequest values, call the use case, and translate RunResult values into HTTP responses.
|
Candidate extensions:
|
||||||
|
|
||||||
Suggested initial endpoint:
|
- S3 artifact refs via `artifact.Reader` extension.
|
||||||
|
- Token budgeting in usecase/model-target policy layer.
|
||||||
|
- Streaming LLM output via additional `llm.Client` methods/adapters.
|
||||||
|
- Batch execution as a separate use case (not hidden in single-run path).
|
||||||
|
- Additional LLM providers implementing `llm.Client`.
|
||||||
|
- Additional validators/modes in `validate`.
|
||||||
|
- Additional profile repositories (embedded, remote, object storage).
|
||||||
|
|
||||||
- POST /v1/runs
|
These are future work, not part of current default behavior.
|
||||||
|
|
||||||
The request should include:
|
## 10. Architectural Guardrails
|
||||||
|
|
||||||
- profile_id
|
Contributors should preserve these constraints:
|
||||||
- inputs
|
|
||||||
- vars
|
|
||||||
- optional model override
|
|
||||||
- optional caller metadata
|
|
||||||
|
|
||||||
The response should include:
|
- No D&D-specific behavior in core Go packages.
|
||||||
|
- No orchestration creep into Scriptorium.
|
||||||
|
- No unbounded repair loops.
|
||||||
|
- No silent truncation/omission of rendered inputs or outputs.
|
||||||
|
- Do not log full artifacts/prompts by default.
|
||||||
|
- Keep provider-specific wire/SDK details out of domain types.
|
||||||
|
- Keep adapter boundaries explicit and thin.
|
||||||
|
|
||||||
- artifact
|
## 11. Testing Strategy
|
||||||
- validation
|
|
||||||
- metadata
|
|
||||||
- raw_model_output
|
|
||||||
- error details, if applicable
|
|
||||||
|
|
||||||
The HTTP layer should not contain business logic.
|
Protect behavior at boundaries and in usecase flow:
|
||||||
|
|
||||||
## CLI
|
- Profile loading/parsing/validation errors.
|
||||||
|
- Artifact reading for inline/file + hash/content type behavior.
|
||||||
|
- Prompt rendering required inputs/template error behavior.
|
||||||
|
- LLM adapter request/response/error/timeout behavior.
|
||||||
|
- Runner success path and metadata population.
|
||||||
|
- Validation failure raw-output preservation.
|
||||||
|
- Successful/failed/bounded repair flows.
|
||||||
|
- HTTP request mapping, response shape, and error mapping.
|
||||||
|
- CLI parsing helpers, required flags, and output stream separation.
|
||||||
|
|
||||||
The CLI should be thin and call the same core use case as HTTP.
|
Prefer focused unit tests and small integration-style tests with fake LLMs.
|
||||||
|
|
||||||
Current command surface:
|
|
||||||
|
|
||||||
- scriptorium run
|
|
||||||
- scriptorium serve
|
|
||||||
|
|
||||||
The `run` command should accept:
|
|
||||||
|
|
||||||
- profile ID
|
|
||||||
- input mappings
|
|
||||||
- variable mappings
|
|
||||||
- output path, optional
|
|
||||||
- profile directory
|
|
||||||
- optional model/endpoint overrides
|
|
||||||
|
|
||||||
If model/endpoint overrides are omitted, profile model defaults should be used.
|
|
||||||
|
|
||||||
## Configuration
|
|
||||||
|
|
||||||
Application configuration should include:
|
|
||||||
|
|
||||||
- Prompt profile directory
|
|
||||||
- Schema directory
|
|
||||||
- LLM endpoints
|
|
||||||
- Default endpoint
|
|
||||||
- Timeout settings
|
|
||||||
- Optional artifact store settings
|
|
||||||
- Logging settings
|
|
||||||
|
|
||||||
Configuration should be file-based with environment variable overrides where appropriate.
|
|
||||||
|
|
||||||
Avoid hardcoding local paths.
|
|
||||||
|
|
||||||
Avoid hardcoding D&D-specific defaults.
|
|
||||||
|
|
||||||
## Artifact Storage
|
|
||||||
|
|
||||||
Scriptorium does not need to own artifact persistence in v1.
|
|
||||||
|
|
||||||
The default behavior should be:
|
|
||||||
|
|
||||||
- Read input artifacts.
|
|
||||||
- Return generated artifact to caller.
|
|
||||||
|
|
||||||
Narratio or another orchestrator can save the result to S3.
|
|
||||||
|
|
||||||
However, Scriptorium should be designed so that artifact readers and writers can be added later.
|
|
||||||
|
|
||||||
If an ArtifactWriter is added, it should be optional and should not change the core use case.
|
|
||||||
|
|
||||||
## Error Handling
|
|
||||||
|
|
||||||
Errors should be explicit and typed where useful.
|
|
||||||
|
|
||||||
Important error categories:
|
|
||||||
|
|
||||||
- Profile not found
|
|
||||||
- Invalid profile
|
|
||||||
- Required input missing
|
|
||||||
- Artifact read failure
|
|
||||||
- Template render failure
|
|
||||||
- LLM request failure
|
|
||||||
- LLM response parse failure
|
|
||||||
- Output validation failure
|
|
||||||
- Repair failure
|
|
||||||
|
|
||||||
Validation failure is not necessarily the same as application failure.
|
|
||||||
|
|
||||||
If the model returns output but the output fails validation, Scriptorium should return a structured RunResult with failed validation status when possible.
|
|
||||||
|
|
||||||
Transport-level errors, missing inputs, invalid profiles, and failed model calls should be returned as application errors.
|
|
||||||
|
|
||||||
## Observability
|
|
||||||
|
|
||||||
Use structured logging.
|
|
||||||
|
|
||||||
Log important lifecycle events:
|
|
||||||
|
|
||||||
- Run started
|
|
||||||
- Profile loaded
|
|
||||||
- Inputs loaded
|
|
||||||
- Prompt rendered
|
|
||||||
- LLM request started
|
|
||||||
- LLM response received
|
|
||||||
- Validation completed
|
|
||||||
- Repair attempted
|
|
||||||
- Run completed
|
|
||||||
|
|
||||||
Do not log full prompt content or full artifact content by default.
|
|
||||||
|
|
||||||
Do log hashes, sizes, profile IDs, model names, durations, and validation status.
|
|
||||||
|
|
||||||
## Metadata and Reproducibility
|
|
||||||
|
|
||||||
Every successful or partially successful run should include metadata.
|
|
||||||
|
|
||||||
Recommended metadata:
|
|
||||||
|
|
||||||
- Run ID
|
|
||||||
- Profile ID
|
|
||||||
- Profile version
|
|
||||||
- Profile hash
|
|
||||||
- Prompt hash
|
|
||||||
- Input artifact hashes
|
|
||||||
- Model endpoint
|
|
||||||
- Model name
|
|
||||||
- Generation parameters
|
|
||||||
- Created timestamp
|
|
||||||
- Duration
|
|
||||||
- Token usage, if available
|
|
||||||
- Validation mode
|
|
||||||
- Validation status
|
|
||||||
- Repair attempts used
|
|
||||||
|
|
||||||
This metadata is important for auditing and regeneration.
|
|
||||||
|
|
||||||
## Security and Safety Considerations
|
|
||||||
|
|
||||||
Scriptorium will often handle private transcripts or documents.
|
|
||||||
|
|
||||||
Default behavior should avoid accidental disclosure.
|
|
||||||
|
|
||||||
Guidelines:
|
|
||||||
|
|
||||||
- Do not log full inputs by default.
|
|
||||||
- Do not log full model outputs by default unless debug logging is explicitly enabled.
|
|
||||||
- Keep API keys in configuration or environment variables, not in prompt profiles.
|
|
||||||
- Avoid exposing local filesystem paths in public error messages when running as a service.
|
|
||||||
- Treat prompt profiles as trusted configuration.
|
|
||||||
- Treat input artifacts as untrusted content.
|
|
||||||
- Avoid shell execution entirely.
|
|
||||||
|
|
||||||
## Testing Strategy
|
|
||||||
|
|
||||||
Tests should focus on the use case and boundaries.
|
|
||||||
|
|
||||||
Recommended tests:
|
|
||||||
|
|
||||||
- Load valid prompt profile.
|
|
||||||
- Reject invalid prompt profile.
|
|
||||||
- Reject missing required input.
|
|
||||||
- Render prompt with named inputs.
|
|
||||||
- Render prompt with variables.
|
|
||||||
- Execute run with fake LLM client.
|
|
||||||
- Validate successful Markdown output.
|
|
||||||
- Validate successful JSON output.
|
|
||||||
- Detect invalid JSON output.
|
|
||||||
- Detect JSON Schema validation errors.
|
|
||||||
- Perform successful repair with fake LLM client.
|
|
||||||
- Preserve raw output on validation failure.
|
|
||||||
- Return useful metadata.
|
|
||||||
- HTTP handler maps request to use case correctly.
|
|
||||||
- CLI command maps flags to use case correctly.
|
|
||||||
|
|
||||||
Use fixtures in testdata.
|
|
||||||
|
|
||||||
The core use case should be testable without network access.
|
|
||||||
|
|
||||||
## Development Priorities
|
|
||||||
|
|
||||||
Implementation should proceed in this order:
|
|
||||||
|
|
||||||
1. Define domain types.
|
|
||||||
2. Define core interfaces.
|
|
||||||
3. Implement prompt profile loading from YAML.
|
|
||||||
4. Implement artifact loading for inline and local file inputs.
|
|
||||||
5. Implement prompt rendering.
|
|
||||||
6. Implement fake LLM client tests.
|
|
||||||
7. Implement OpenAI-compatible LLM client.
|
|
||||||
8. Implement basic validation.
|
|
||||||
9. Implement JSON validation.
|
|
||||||
10. Implement JSON Schema validation.
|
|
||||||
11. Implement optional repair.
|
|
||||||
12. Implement CLI.
|
|
||||||
13. Implement HTTP API.
|
|
||||||
14. Add example D&D profiles and schemas.
|
|
||||||
15. Add integration-style tests using fake adapters.
|
|
||||||
|
|
||||||
Do not start with HTTP or CLI. Start with the core use case.
|
|
||||||
|
|
||||||
## Design Principles
|
|
||||||
|
|
||||||
Prefer boring code.
|
|
||||||
|
|
||||||
Prefer explicit configuration.
|
|
||||||
|
|
||||||
Prefer small packages with clear responsibilities.
|
|
||||||
|
|
||||||
Prefer interfaces only at real boundaries.
|
|
||||||
|
|
||||||
Do not create abstractions before they are needed.
|
|
||||||
|
|
||||||
Do not let prompt profile complexity leak into Go code.
|
|
||||||
|
|
||||||
Do not let D&D assumptions leak into the core engine.
|
|
||||||
|
|
||||||
Do not silently truncate inputs.
|
|
||||||
|
|
||||||
Do not discard invalid model output.
|
|
||||||
|
|
||||||
Do not hide validation errors.
|
|
||||||
|
|
||||||
Do not implement an orchestrator inside Scriptorium.
|
|
||||||
|
|
||||||
## Summary
|
|
||||||
|
|
||||||
Scriptorium is a reusable prompt-profile execution engine.
|
|
||||||
|
|
||||||
It should provide this core transformation:
|
|
||||||
|
|
||||||
Named artifacts plus prompt profile plus model target produces generated artifact plus validation plus metadata.
|
|
||||||
|
|
||||||
The D&D transcript analysis workflow is the first use case, not the architecture itself.
|
|
||||||
|
|
||||||
The correct implementation is a small, modular Go service with a clean core use case and replaceable adapters for profiles, artifacts, prompt rendering, LLM calls, validation, CLI, and HTTP.
|
|
||||||
|
|||||||
Reference in New Issue
Block a user