Files
scriptorium/architecture.md

20 KiB

Scriptorium Architecture

Purpose

Scriptorium is a general-purpose prompt-profile execution service.

Its job is to take one or more named input artifacts, render a configured prompt profile, execute that prompt against an LLM endpoint, optionally validate the output, and return a generated artifact with useful metadata.

The initial concrete use case is generating artifacts from cleaned Dungeons & Dragons session transcripts, such as session recaps, player analysis, structured event extraction, and glossary update suggestions.

However, Scriptorium must not be D&D-specific. D&D behavior belongs in prompt profiles, schemas, and caller-provided inputs. The Go application should remain a generic engine for prompt execution and output validation.

Intended Audience

This document is written for 5.3-Codex and future maintainers.

When implementing this repository, prefer simple, idiomatic Go over elaborate framework code. The architecture should be modular, testable, and composable, but not over-engineered.

The desired implementation style is:

  • Clear domain types.
  • Small interfaces at architectural boundaries.
  • Explicit dependencies.
  • No hidden global state.
  • No domain-specific D&D logic in core packages.
  • Practical hexagonal / ports-and-adapters structure.
  • Boring, inspectable behavior.

Core Concept

Scriptorium transforms:

  • Prompt profile
  • Named input artifacts
  • Template variables
  • Model target
  • Optional output contract

Into:

  • Generated artifact
  • Validation result
  • Prompt/model/input metadata
  • Raw model output
  • Structured error details, if applicable

Scriptorium should be thought of as a deterministic wrapper around a nondeterministic model call.

The system should make the model call as auditable and reproducible as possible, even though LLM output itself may not be exactly reproducible.

Application Boundary

Scriptorium is not an orchestrator.

The broader workflow may include audio transcription, transcript merging, transcript polishing, artifact persistence, and notifications. Those responsibilities belong to the external orchestrator, currently expected to be Narratio.

Scriptorium should not know about WhisperX, Seriatim, Audita, or any other pipeline stage.

Scriptorium only knows how to:

  1. Load a prompt profile.
  2. Load or receive named input artifacts.
  3. Render a prompt.
  4. Call an LLM.
  5. Validate the output, if configured.
  6. Return an output artifact and metadata.

Initial Workflow Context

The initial D&D workflow is expected to look like this:

  1. Narratio transcribes audio tracks using WhisperX.
  2. Narratio normalizes speaker names and saves per-speaker transcripts.
  3. Narratio calls Seriatim to merge transcripts.
  4. Narratio saves the merged transcript.
  5. Narratio calls Audita to polish the transcript.
  6. Narratio saves the processed transcript.
  7. Narratio calls Scriptorium one or more times to generate output artifacts.
  8. Narratio saves each generated artifact.
  9. Narratio optionally sends a completion notification.

Scriptorium only owns step 7.

Each Scriptorium request should initially produce one artifact. If multiple artifacts are needed, the orchestrator should call Scriptorium multiple times.

Batch execution can be added later, but should not be part of the core v1 design unless there is an immediate need.

Primary Use Cases

Scriptorium should support the following v1 use cases:

  1. Generate a freeform Markdown artifact from a transcript and prompt profile.
  2. Generate a structured JSON artifact from a transcript and prompt profile.
  3. Validate JSON output against a JSON Schema.
  4. Return raw model output when validation fails.
  5. Optionally attempt one bounded repair pass for invalid structured output.
  6. Record metadata about the profile, model, inputs, prompt hash, and validation result.
  7. Support local development through a CLI.
  8. Support service usage through an HTTP API.

Non-Goals for v1

Do not implement these in the initial version unless explicitly requested:

  • Multi-agent workflows.
  • Arbitrary DAG execution.
  • Long-running job queues.
  • Automatic RAG.
  • Automatic prompt chaining.
  • Automatic chunking and summarization.
  • Model selection logic.
  • Complex retry policies beyond basic HTTP/model retry and optional validation repair.
  • D&D-specific Go packages.
  • UI.
  • Database persistence.
  • Full artifact lifecycle management.

These may be valid future features, but v1 should remain a focused prompt-profile execution engine.

Architectural Style

Use a practical hexagonal architecture.

The core domain and use case packages should not depend on infrastructure details such as HTTP, S3, local filesystems, or specific LLM providers.

External concerns should be implemented as adapters.

The central use case should be easy to test with fake prompt repositories, fake artifact readers, fake LLM clients, and fake validators.

Recommended high-level structure:

  • cmd/scriptorium: application entrypoint
  • internal/domain: core domain types
  • internal/usecase: application use cases
  • internal/profile: prompt profile loading and parsing
  • internal/prompt: prompt rendering
  • internal/llm: LLM client interfaces and adapters
  • internal/validate: output validation implementations
  • internal/artifact: artifact loading and storage adapters
  • internal/adapter/http: HTTP API
  • internal/adapter/cli: CLI interface
  • internal/config: application configuration
  • profiles: example prompt profiles
  • schemas: example output schemas
  • testdata: fixtures for tests

Exact package names may evolve, but the boundary principles should remain stable.

Domain Model

The core domain should include these concepts.

RunRequest

Represents one request to generate one artifact.

Fields should include:

  • ProfileID
  • Inputs
  • Vars
  • Optional model override
  • Optional validation override, if needed
  • Optional caller metadata

Inputs should be keyed by logical input name, not by filename.

Example logical input names:

  • transcript
  • glossary
  • previous_recap
  • campaign_notes
  • source_document

ArtifactRef

Represents a reference to an input artifact.

Artifact references should support at least inline content and local file paths in v1.

S3 references may be supported in v1 if needed, but should be implemented behind an interface.

Likely artifact reference types:

  • inline
  • file
  • s3

The core use case should not care which reference type is used.

Artifact

Represents loaded content.

Fields should include:

  • Name
  • ContentType
  • Body
  • Optional URI or source reference
  • Optional size
  • Optional SHA-256 hash

Artifacts are the actual input and output payloads after references have been resolved.

PromptProfile

Represents a configured prompt execution profile.

A profile should include:

  • ID
  • Version
  • Description
  • Expected inputs
  • Prompt templates
  • Model defaults
  • Output format
  • Optional validation configuration
  • Optional repair configuration

Prompt profiles should be serializable from YAML.

Prompt profiles are where domain-specific behavior belongs.

RenderedPrompt

Represents the prompt after input artifacts and variables have been applied.

For OpenAI-compatible chat models, this should contain a list of chat messages.

At minimum, support system and user messages.

Future support for developer messages, assistant prefill, or multimodal parts can be added later.

ModelTarget

Represents the LLM endpoint and model configuration.

Fields should include:

  • Endpoint name or URL
  • Model name
  • Temperature
  • Max tokens
  • Top-p, if supported
  • Additional provider-specific options, if needed

For v1, the main adapter should support OpenAI-compatible chat completion APIs.

RunResult

Represents the complete result of a run.

Fields should include:

  • Generated artifact
  • Raw model output
  • Validation result
  • Profile ID and version
  • Model name
  • Endpoint name
  • Input hashes
  • Prompt hash
  • Token usage, if available
  • Start and end timestamps
  • Error details, if applicable

ValidationResult

Represents validation status.

Fields should include:

  • Status: passed, failed, skipped
  • Validation mode
  • Error messages
  • Schema path, if applicable
  • Repair attempts used
  • Final output validity

Core Interfaces

Keep interfaces small and focused.

PromptRepository

Responsible for loading prompt profiles.

Expected implementations:

  • FilesystemPromptRepository
  • EmbeddedPromptRepository, optional later
  • S3PromptRepository, optional later

The use case should ask for a profile by ID or ID plus version.

ArtifactReader

Responsible for resolving ArtifactRef values into Artifact values.

Expected implementations:

  • InlineArtifactReader
  • FileArtifactReader
  • S3ArtifactReader
  • CompositeArtifactReader

The CompositeArtifactReader can route by reference type.

PromptRenderer

Responsible for rendering prompt templates using named artifacts and variables.

Use Go templates unless there is a strong reason to choose something else.

Renderer responsibilities:

  • Verify required inputs exist.
  • Expose safe template functions.
  • Insert artifact content by logical name.
  • Render system and user prompt sections.
  • Return a RenderedPrompt.

Do not silently omit missing required inputs.

Do not silently truncate large inputs in v1.

LLMClient

Responsible for executing a rendered prompt against a model endpoint.

The initial implementation should support OpenAI-compatible chat completions.

This should work with:

  • vLLM
  • LiteLLM
  • OpenAI-compatible local endpoints
  • OpenAI-compatible hosted endpoints, if configured

The domain should not depend on provider-specific SDK types.

OutputValidator

Responsible for validating the generated artifact.

Expected validation modes:

  • none
  • basic
  • json_schema

Basic validation may check things like non-empty output, required headings, or forbidden boilerplate.

JSON Schema validation should parse the output as JSON and validate it against the configured schema.

OutputRepairer

Responsible for making a bounded attempt to repair invalid structured output.

This should be optional.

The repairer may use the same LLMClient with a repair prompt.

Repair attempts must be bounded by configuration. Default should be zero or one.

Do not implement unbounded repair loops.

Prompt Profiles

Prompt profiles are the main extension mechanism.

The Go application should stay generic. Prompt profiles should define domain behavior.

A profile should be able to specify:

  • ID
  • Version
  • Description
  • Required and optional inputs
  • System prompt template
  • User prompt template
  • Default model configuration
  • Output format
  • Validation mode
  • Schema path, if applicable
  • Repair attempts, if applicable

Profiles should live outside compiled Go code.

Example profile categories for the initial D&D use case:

  • dnd.session_recap
  • dnd.meta_analysis
  • dnd.table_read
  • dnd.structured_events
  • dnd.glossary_update_suggestions

The code should not special-case these names.

Template Rendering

Prompt rendering must be predictable and explicit.

Templates should be able to reference:

  • Named input artifacts
  • Template variables
  • Profile metadata

The renderer should provide a helper equivalent to input(name), which inserts the content of a named artifact.

The renderer should fail when:

  • A required input is missing.
  • A template references an unknown input.
  • A template references a missing required variable.
  • The rendered prompt exceeds a configured token or size limit, if such a limit is configured.

In v1, do not silently truncate inputs.

If token counting is not implemented initially, use byte-size limits or leave token budgeting as a clearly marked future improvement.

Output Formats

Scriptorium should support at least these output formats:

  • markdown
  • text
  • json

For markdown and text, validation may be skipped or basic.

For JSON, validation should at minimum require valid JSON. If a schema is configured, validate against the schema.

The output artifact should preserve content type.

Suggested content types:

  • text/markdown
  • text/plain
  • application/json

Validation

Validation should be explicit and profile-driven.

Validation modes:

  • none: no validation beyond successful generation
  • basic: simple textual validation
  • json: parse as JSON
  • json_schema: parse as JSON and validate against schema

For invalid structured output, Scriptorium should return:

  • Validation status
  • Validation errors
  • Raw model output
  • Repair attempts used
  • Final output, if repair succeeded

Validation failure should not discard the raw output.

Repair

Repair is only for structured output.

The initial repair use case is invalid JSON or JSON that fails schema validation.

The repair prompt should be deterministic and narrow:

  • Explain that the previous output failed validation.
  • Provide validation errors.
  • Provide the previous output.
  • Ask the model to return only corrected JSON.
  • Do not ask the model to improve the answer substantively.

Repair must be bounded.

Recommended default:

  • repair_attempts: 0 for freeform output
  • repair_attempts: 1 for JSON schema output, if configured

LLM Adapter

The initial LLM adapter should target OpenAI-compatible chat completions.

The adapter should support:

  • Base URL
  • API key, optional for local endpoints
  • Model name
  • Temperature
  • Max tokens
  • Basic generation parameters
  • Request timeout
  • Token usage extraction, if returned by the endpoint

Do not couple the core domain to OpenAI SDK request or response structs.

The adapter should translate between internal GenerateRequest / GenerateResponse types and the provider wire format.

HTTP API

The HTTP API should be thin.

It should translate HTTP requests into RunRequest values, call the use case, and translate RunResult values into HTTP responses.

Suggested initial endpoint:

  • POST /v1/runs

The request should include:

  • profile_id
  • inputs
  • vars
  • optional model override
  • optional caller metadata

The response should include:

  • artifact
  • validation
  • metadata
  • raw_model_output, optionally controlled by request or config
  • error details, if applicable

The HTTP layer should not contain business logic.

CLI

The CLI should also be thin.

It should support local development and pipeline usage.

Suggested commands:

  • scriptorium run
  • scriptorium profiles list
  • scriptorium profiles inspect

The run command should accept:

  • profile ID
  • input mappings
  • variable mappings
  • output path, optional
  • profile directory
  • config path

The CLI should call the same use case used by the HTTP API.

Configuration

Application configuration should include:

  • Prompt profile directory
  • Schema directory
  • LLM endpoints
  • Default endpoint
  • Timeout settings
  • Optional artifact store settings
  • Logging settings

Configuration should be file-based with environment variable overrides where appropriate.

Avoid hardcoding local paths.

Avoid hardcoding D&D-specific defaults.

Artifact Storage

Scriptorium does not need to own artifact persistence in v1.

The default behavior should be:

  • Read input artifacts.
  • Return generated artifact to caller.

Narratio or another orchestrator can save the result to S3.

However, Scriptorium should be designed so that artifact readers and writers can be added later.

If an ArtifactWriter is added, it should be optional and should not change the core use case.

Error Handling

Errors should be explicit and typed where useful.

Important error categories:

  • Profile not found
  • Invalid profile
  • Required input missing
  • Artifact read failure
  • Template render failure
  • LLM request failure
  • LLM response parse failure
  • Output validation failure
  • Repair failure

Validation failure is not necessarily the same as application failure.

If the model returns output but the output fails validation, Scriptorium should return a structured RunResult with failed validation status when possible.

Transport-level errors, missing inputs, invalid profiles, and failed model calls should be returned as application errors.

Observability

Use structured logging.

Log important lifecycle events:

  • Run started
  • Profile loaded
  • Inputs loaded
  • Prompt rendered
  • LLM request started
  • LLM response received
  • Validation completed
  • Repair attempted
  • Run completed

Do not log full prompt content or full artifact content by default.

Do log hashes, sizes, profile IDs, model names, durations, and validation status.

Metadata and Reproducibility

Every successful or partially successful run should include metadata.

Recommended metadata:

  • Run ID
  • Profile ID
  • Profile version
  • Profile hash
  • Prompt hash
  • Input artifact hashes
  • Model endpoint
  • Model name
  • Generation parameters
  • Created timestamp
  • Duration
  • Token usage, if available
  • Validation mode
  • Validation status
  • Repair attempts used

This metadata is important for auditing and regeneration.

Security and Safety Considerations

Scriptorium will often handle private transcripts or documents.

Default behavior should avoid accidental disclosure.

Guidelines:

  • Do not log full inputs by default.
  • Do not log full model outputs by default unless debug logging is explicitly enabled.
  • Keep API keys in configuration or environment variables, not in prompt profiles.
  • Avoid exposing local filesystem paths in public error messages when running as a service.
  • Treat prompt profiles as trusted configuration.
  • Treat input artifacts as untrusted content.
  • Avoid shell execution entirely.

Testing Strategy

Tests should focus on the use case and boundaries.

Recommended tests:

  • Load valid prompt profile.
  • Reject invalid prompt profile.
  • Reject missing required input.
  • Render prompt with named inputs.
  • Render prompt with variables.
  • Execute run with fake LLM client.
  • Validate successful Markdown output.
  • Validate successful JSON output.
  • Detect invalid JSON output.
  • Detect JSON Schema validation errors.
  • Perform successful repair with fake LLM client.
  • Preserve raw output on validation failure.
  • Return useful metadata.
  • HTTP handler maps request to use case correctly.
  • CLI command maps flags to use case correctly.

Use fixtures in testdata.

The core use case should be testable without network access.

Development Priorities

Implementation should proceed in this order:

  1. Define domain types.
  2. Define core interfaces.
  3. Implement prompt profile loading from YAML.
  4. Implement artifact loading for inline and local file inputs.
  5. Implement prompt rendering.
  6. Implement fake LLM client tests.
  7. Implement OpenAI-compatible LLM client.
  8. Implement basic validation.
  9. Implement JSON validation.
  10. Implement JSON Schema validation.
  11. Implement optional repair.
  12. Implement CLI.
  13. Implement HTTP API.
  14. Add example D&D profiles and schemas.
  15. Add integration-style tests using fake adapters.

Do not start with HTTP or CLI. Start with the core use case.

Design Principles

Prefer boring code.

Prefer explicit configuration.

Prefer small packages with clear responsibilities.

Prefer interfaces only at real boundaries.

Do not create abstractions before they are needed.

Do not let prompt profile complexity leak into Go code.

Do not let D&D assumptions leak into the core engine.

Do not silently truncate inputs.

Do not discard invalid model output.

Do not hide validation errors.

Do not implement an orchestrator inside Scriptorium.

Summary

Scriptorium is a reusable prompt-profile execution engine.

It should provide this core transformation:

Named artifacts plus prompt profile plus model target produces generated artifact plus validation plus metadata.

The D&D transcript analysis workflow is the first use case, not the architecture itself.

The correct implementation is a small, modular Go service with a clean core use case and replaceable adapters for profiles, artifacts, prompt rendering, LLM calls, validation, CLI, and HTTP.