diff --git a/architecture.md b/architecture.md new file mode 100644 index 0000000..c896452 --- /dev/null +++ b/architecture.md @@ -0,0 +1,755 @@ +# Analyzer Architecture + +## Purpose + +Analyzer is a general-purpose prompt-profile execution service. + +Its job is to take one or more named input artifacts, render a configured prompt profile, execute that prompt against an LLM endpoint, optionally validate the output, and return a generated artifact with useful metadata. + +The initial concrete use case is generating artifacts from cleaned Dungeons & Dragons session transcripts, such as session recaps, player analysis, structured event extraction, and glossary update suggestions. + +However, Analyzer must not be D&D-specific. D&D behavior belongs in prompt profiles, schemas, and caller-provided inputs. The Go application should remain a generic engine for prompt execution and output validation. + +## Intended Audience + +This document is written for 5.3-Codex and future maintainers. + +When implementing this repository, prefer simple, idiomatic Go over elaborate framework code. The architecture should be modular, testable, and composable, but not over-engineered. + +The desired implementation style is: + +- Clear domain types. +- Small interfaces at architectural boundaries. +- Explicit dependencies. +- No hidden global state. +- No domain-specific D&D logic in core packages. +- Practical hexagonal / ports-and-adapters structure. +- Boring, inspectable behavior. + +## Core Concept + +Analyzer transforms: + +- Prompt profile +- Named input artifacts +- Template variables +- Model target +- Optional output contract + +Into: + +- Generated artifact +- Validation result +- Prompt/model/input metadata +- Raw model output +- Structured error details, if applicable + +Analyzer should be thought of as a deterministic wrapper around a nondeterministic model call. + +The system should make the model call as auditable and reproducible as possible, even though LLM output itself may not be exactly reproducible. + +## Application Boundary + +Analyzer is not an orchestrator. + +The broader workflow may include audio transcription, transcript merging, transcript polishing, artifact persistence, and notifications. Those responsibilities belong to the external orchestrator, currently expected to be Narratio. + +Analyzer should not know about WhisperX, Seriatim, Audita, or any other pipeline stage. + +Analyzer only knows how to: + +1. Load a prompt profile. +2. Load or receive named input artifacts. +3. Render a prompt. +4. Call an LLM. +5. Validate the output, if configured. +6. Return an output artifact and metadata. + +## Initial Workflow Context + +The initial D&D workflow is expected to look like this: + +1. Narratio transcribes audio tracks using WhisperX. +2. Narratio normalizes speaker names and saves per-speaker transcripts. +3. Narratio calls Seriatim to merge transcripts. +4. Narratio saves the merged transcript. +5. Narratio calls Audita to polish the transcript. +6. Narratio saves the processed transcript. +7. Narratio calls Analyzer one or more times to generate output artifacts. +8. Narratio saves each generated artifact. +9. Narratio optionally sends a completion notification. + +Analyzer only owns step 7. + +Each Analyzer request should initially produce one artifact. If multiple artifacts are needed, the orchestrator should call Analyzer multiple times. + +Batch execution can be added later, but should not be part of the core v1 design unless there is an immediate need. + +## Primary Use Cases + +Analyzer should support the following v1 use cases: + +1. Generate a freeform Markdown artifact from a transcript and prompt profile. +2. Generate a structured JSON artifact from a transcript and prompt profile. +3. Validate JSON output against a JSON Schema. +4. Return raw model output when validation fails. +5. Optionally attempt one bounded repair pass for invalid structured output. +6. Record metadata about the profile, model, inputs, prompt hash, and validation result. +7. Support local development through a CLI. +8. Support service usage through an HTTP API. + +## Non-Goals for v1 + +Do not implement these in the initial version unless explicitly requested: + +- Multi-agent workflows. +- Arbitrary DAG execution. +- Long-running job queues. +- Automatic RAG. +- Automatic prompt chaining. +- Automatic chunking and summarization. +- Model selection logic. +- Complex retry policies beyond basic HTTP/model retry and optional validation repair. +- D&D-specific Go packages. +- UI. +- Database persistence. +- Full artifact lifecycle management. + +These may be valid future features, but v1 should remain a focused prompt-profile execution engine. + +## Architectural Style + +Use a practical hexagonal architecture. + +The core domain and use case packages should not depend on infrastructure details such as HTTP, S3, local filesystems, or specific LLM providers. + +External concerns should be implemented as adapters. + +The central use case should be easy to test with fake prompt repositories, fake artifact readers, fake LLM clients, and fake validators. + +Recommended high-level structure: + +- cmd/analyzer: application entrypoint +- internal/domain: core domain types +- internal/usecase: application use cases +- internal/profile: prompt profile loading and parsing +- internal/prompt: prompt rendering +- internal/llm: LLM client interfaces and adapters +- internal/validate: output validation implementations +- internal/artifact: artifact loading and storage adapters +- internal/adapter/http: HTTP API +- internal/adapter/cli: CLI interface +- internal/config: application configuration +- profiles: example prompt profiles +- schemas: example output schemas +- testdata: fixtures for tests + +Exact package names may evolve, but the boundary principles should remain stable. + +## Domain Model + +The core domain should include these concepts. + +### RunRequest + +Represents one request to generate one artifact. + +Fields should include: + +- ProfileID +- Inputs +- Vars +- Optional model override +- Optional validation override, if needed +- Optional caller metadata + +Inputs should be keyed by logical input name, not by filename. + +Example logical input names: + +- transcript +- glossary +- previous_recap +- campaign_notes +- source_document + +### ArtifactRef + +Represents a reference to an input artifact. + +Artifact references should support at least inline content and local file paths in v1. + +S3 references may be supported in v1 if needed, but should be implemented behind an interface. + +Likely artifact reference types: + +- inline +- file +- s3 + +The core use case should not care which reference type is used. + +### Artifact + +Represents loaded content. + +Fields should include: + +- Name +- ContentType +- Body +- Optional URI or source reference +- Optional size +- Optional SHA-256 hash + +Artifacts are the actual input and output payloads after references have been resolved. + +### PromptProfile + +Represents a configured prompt execution profile. + +A profile should include: + +- ID +- Version +- Description +- Expected inputs +- Prompt templates +- Model defaults +- Output format +- Optional validation configuration +- Optional repair configuration + +Prompt profiles should be serializable from YAML. + +Prompt profiles are where domain-specific behavior belongs. + +### RenderedPrompt + +Represents the prompt after input artifacts and variables have been applied. + +For OpenAI-compatible chat models, this should contain a list of chat messages. + +At minimum, support system and user messages. + +Future support for developer messages, assistant prefill, or multimodal parts can be added later. + +### ModelTarget + +Represents the LLM endpoint and model configuration. + +Fields should include: + +- Endpoint name or URL +- Model name +- Temperature +- Max tokens +- Top-p, if supported +- Additional provider-specific options, if needed + +For v1, the main adapter should support OpenAI-compatible chat completion APIs. + +### RunResult + +Represents the complete result of a run. + +Fields should include: + +- Generated artifact +- Raw model output +- Validation result +- Profile ID and version +- Model name +- Endpoint name +- Input hashes +- Prompt hash +- Token usage, if available +- Start and end timestamps +- Error details, if applicable + +### ValidationResult + +Represents validation status. + +Fields should include: + +- Status: passed, failed, skipped +- Validation mode +- Error messages +- Schema path, if applicable +- Repair attempts used +- Final output validity + +## Core Interfaces + +Keep interfaces small and focused. + +### PromptRepository + +Responsible for loading prompt profiles. + +Expected implementations: + +- FilesystemPromptRepository +- EmbeddedPromptRepository, optional later +- S3PromptRepository, optional later + +The use case should ask for a profile by ID or ID plus version. + +### ArtifactReader + +Responsible for resolving ArtifactRef values into Artifact values. + +Expected implementations: + +- InlineArtifactReader +- FileArtifactReader +- S3ArtifactReader +- CompositeArtifactReader + +The CompositeArtifactReader can route by reference type. + +### PromptRenderer + +Responsible for rendering prompt templates using named artifacts and variables. + +Use Go templates unless there is a strong reason to choose something else. + +Renderer responsibilities: + +- Verify required inputs exist. +- Expose safe template functions. +- Insert artifact content by logical name. +- Render system and user prompt sections. +- Return a RenderedPrompt. + +Do not silently omit missing required inputs. + +Do not silently truncate large inputs in v1. + +### LLMClient + +Responsible for executing a rendered prompt against a model endpoint. + +The initial implementation should support OpenAI-compatible chat completions. + +This should work with: + +- vLLM +- LiteLLM +- OpenAI-compatible local endpoints +- OpenAI-compatible hosted endpoints, if configured + +The domain should not depend on provider-specific SDK types. + +### OutputValidator + +Responsible for validating the generated artifact. + +Expected validation modes: + +- none +- basic +- json_schema + +Basic validation may check things like non-empty output, required headings, or forbidden boilerplate. + +JSON Schema validation should parse the output as JSON and validate it against the configured schema. + +### OutputRepairer + +Responsible for making a bounded attempt to repair invalid structured output. + +This should be optional. + +The repairer may use the same LLMClient with a repair prompt. + +Repair attempts must be bounded by configuration. Default should be zero or one. + +Do not implement unbounded repair loops. + +## Prompt Profiles + +Prompt profiles are the main extension mechanism. + +The Go application should stay generic. Prompt profiles should define domain behavior. + +A profile should be able to specify: + +- ID +- Version +- Description +- Required and optional inputs +- System prompt template +- User prompt template +- Default model configuration +- Output format +- Validation mode +- Schema path, if applicable +- Repair attempts, if applicable + +Profiles should live outside compiled Go code. + +Example profile categories for the initial D&D use case: + +- dnd.session_recap +- dnd.meta_analysis +- dnd.table_read +- dnd.structured_events +- dnd.glossary_update_suggestions + +The code should not special-case these names. + +## Template Rendering + +Prompt rendering must be predictable and explicit. + +Templates should be able to reference: + +- Named input artifacts +- Template variables +- Profile metadata + +The renderer should provide a helper equivalent to input(name), which inserts the content of a named artifact. + +The renderer should fail when: + +- A required input is missing. +- A template references an unknown input. +- A template references a missing required variable. +- The rendered prompt exceeds a configured token or size limit, if such a limit is configured. + +In v1, do not silently truncate inputs. + +If token counting is not implemented initially, use byte-size limits or leave token budgeting as a clearly marked future improvement. + +## Output Formats + +Analyzer should support at least these output formats: + +- markdown +- text +- json + +For markdown and text, validation may be skipped or basic. + +For JSON, validation should at minimum require valid JSON. If a schema is configured, validate against the schema. + +The output artifact should preserve content type. + +Suggested content types: + +- text/markdown +- text/plain +- application/json + +## Validation + +Validation should be explicit and profile-driven. + +Validation modes: + +- none: no validation beyond successful generation +- basic: simple textual validation +- json: parse as JSON +- json_schema: parse as JSON and validate against schema + +For invalid structured output, Analyzer should return: + +- Validation status +- Validation errors +- Raw model output +- Repair attempts used +- Final output, if repair succeeded + +Validation failure should not discard the raw output. + +## Repair + +Repair is only for structured output. + +The initial repair use case is invalid JSON or JSON that fails schema validation. + +The repair prompt should be deterministic and narrow: + +- Explain that the previous output failed validation. +- Provide validation errors. +- Provide the previous output. +- Ask the model to return only corrected JSON. +- Do not ask the model to improve the answer substantively. + +Repair must be bounded. + +Recommended default: + +- repair_attempts: 0 for freeform output +- repair_attempts: 1 for JSON schema output, if configured + +## LLM Adapter + +The initial LLM adapter should target OpenAI-compatible chat completions. + +The adapter should support: + +- Base URL +- API key, optional for local endpoints +- Model name +- Temperature +- Max tokens +- Basic generation parameters +- Request timeout +- Token usage extraction, if returned by the endpoint + +Do not couple the core domain to OpenAI SDK request or response structs. + +The adapter should translate between internal GenerateRequest / GenerateResponse types and the provider wire format. + +## HTTP API + +The HTTP API should be thin. + +It should translate HTTP requests into RunRequest values, call the use case, and translate RunResult values into HTTP responses. + +Suggested initial endpoint: + +- POST /v1/runs + +The request should include: + +- profile_id +- inputs +- vars +- optional model override +- optional caller metadata + +The response should include: + +- artifact +- validation +- metadata +- raw_model_output, optionally controlled by request or config +- error details, if applicable + +The HTTP layer should not contain business logic. + +## CLI + +The CLI should also be thin. + +It should support local development and pipeline usage. + +Suggested commands: + +- analyzer run +- analyzer profiles list +- analyzer profiles inspect + +The run command should accept: + +- profile ID +- input mappings +- variable mappings +- output path, optional +- profile directory +- config path + +The CLI should call the same use case used by the HTTP API. + +## Configuration + +Application configuration should include: + +- Prompt profile directory +- Schema directory +- LLM endpoints +- Default endpoint +- Timeout settings +- Optional artifact store settings +- Logging settings + +Configuration should be file-based with environment variable overrides where appropriate. + +Avoid hardcoding local paths. + +Avoid hardcoding D&D-specific defaults. + +## Artifact Storage + +Analyzer does not need to own artifact persistence in v1. + +The default behavior should be: + +- Read input artifacts. +- Return generated artifact to caller. + +Narratio or another orchestrator can save the result to S3. + +However, Analyzer should be designed so that artifact readers and writers can be added later. + +If an ArtifactWriter is added, it should be optional and should not change the core use case. + +## Error Handling + +Errors should be explicit and typed where useful. + +Important error categories: + +- Profile not found +- Invalid profile +- Required input missing +- Artifact read failure +- Template render failure +- LLM request failure +- LLM response parse failure +- Output validation failure +- Repair failure + +Validation failure is not necessarily the same as application failure. + +If the model returns output but the output fails validation, Analyzer should return a structured RunResult with failed validation status when possible. + +Transport-level errors, missing inputs, invalid profiles, and failed model calls should be returned as application errors. + +## Observability + +Use structured logging. + +Log important lifecycle events: + +- Run started +- Profile loaded +- Inputs loaded +- Prompt rendered +- LLM request started +- LLM response received +- Validation completed +- Repair attempted +- Run completed + +Do not log full prompt content or full artifact content by default. + +Do log hashes, sizes, profile IDs, model names, durations, and validation status. + +## Metadata and Reproducibility + +Every successful or partially successful run should include metadata. + +Recommended metadata: + +- Run ID +- Profile ID +- Profile version +- Profile hash +- Prompt hash +- Input artifact hashes +- Model endpoint +- Model name +- Generation parameters +- Created timestamp +- Duration +- Token usage, if available +- Validation mode +- Validation status +- Repair attempts used + +This metadata is important for auditing and regeneration. + +## Security and Safety Considerations + +Analyzer will often handle private transcripts or documents. + +Default behavior should avoid accidental disclosure. + +Guidelines: + +- Do not log full inputs by default. +- Do not log full model outputs by default unless debug logging is explicitly enabled. +- Keep API keys in configuration or environment variables, not in prompt profiles. +- Avoid exposing local filesystem paths in public error messages when running as a service. +- Treat prompt profiles as trusted configuration. +- Treat input artifacts as untrusted content. +- Avoid shell execution entirely. + +## Testing Strategy + +Tests should focus on the use case and boundaries. + +Recommended tests: + +- Load valid prompt profile. +- Reject invalid prompt profile. +- Reject missing required input. +- Render prompt with named inputs. +- Render prompt with variables. +- Execute run with fake LLM client. +- Validate successful Markdown output. +- Validate successful JSON output. +- Detect invalid JSON output. +- Detect JSON Schema validation errors. +- Perform successful repair with fake LLM client. +- Preserve raw output on validation failure. +- Return useful metadata. +- HTTP handler maps request to use case correctly. +- CLI command maps flags to use case correctly. + +Use fixtures in testdata. + +The core use case should be testable without network access. + +## Development Priorities + +Implementation should proceed in this order: + +1. Define domain types. +2. Define core interfaces. +3. Implement prompt profile loading from YAML. +4. Implement artifact loading for inline and local file inputs. +5. Implement prompt rendering. +6. Implement fake LLM client tests. +7. Implement OpenAI-compatible LLM client. +8. Implement basic validation. +9. Implement JSON validation. +10. Implement JSON Schema validation. +11. Implement optional repair. +12. Implement CLI. +13. Implement HTTP API. +14. Add example D&D profiles and schemas. +15. Add integration-style tests using fake adapters. + +Do not start with HTTP or CLI. Start with the core use case. + +## Design Principles + +Prefer boring code. + +Prefer explicit configuration. + +Prefer small packages with clear responsibilities. + +Prefer interfaces only at real boundaries. + +Do not create abstractions before they are needed. + +Do not let prompt profile complexity leak into Go code. + +Do not let D&D assumptions leak into the core engine. + +Do not silently truncate inputs. + +Do not discard invalid model output. + +Do not hide validation errors. + +Do not implement an orchestrator inside Analyzer. + +## Summary + +Analyzer is a reusable prompt-profile execution engine. + +It should provide this core transformation: + +Named artifacts plus prompt profile plus model target produces generated artifact plus validation plus metadata. + +The D&D transcript analysis workflow is the first use case, not the architecture itself. + +The correct implementation is a small, modular Go service with a clean core use case and replaceable adapters for profiles, artifacts, prompt rendering, LLM calls, validation, CLI, and HTTP. diff --git a/go.mod b/go.mod new file mode 100644 index 0000000..aa9fdf7 --- /dev/null +++ b/go.mod @@ -0,0 +1,5 @@ +module gitea.maximumdirect.net/eric/scriptorium + +go 1.25.5 + +require gopkg.in/yaml.v3 v3.0.1 // indirect diff --git a/go.sum b/go.sum new file mode 100644 index 0000000..4bc0337 --- /dev/null +++ b/go.sum @@ -0,0 +1,3 @@ +gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= +gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA= +gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= diff --git a/internal/artifact/reader.go b/internal/artifact/reader.go new file mode 100644 index 0000000..584562f --- /dev/null +++ b/internal/artifact/reader.go @@ -0,0 +1,90 @@ +package artifact + +import ( + "context" + "crypto/sha256" + "errors" + "fmt" + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" + "mime" + "os" + "path/filepath" +) + +var ( + ErrUnsupportedRefType = errors.New("unsupported artifact reference type") + ErrMissingInlineBody = errors.New("missing body for inline artifact") + ErrMissingFilePath = errors.New("missing file path for file artifact") +) + +// Reader resolves artifact references into actual artifacts. +type Reader interface { + Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error) +} + +// CompositeReader routes artifact resolution based on the reference type. +type CompositeReader struct { + inlineReader *inlineReader + fileReader *fileReader +} + +func NewCompositeReader() Reader { + return &CompositeReader{ + inlineReader: &inlineReader{}, + fileReader: &fileReader{}, + } +} + +func (c *CompositeReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error) { + switch ref.Type { + case domain.ArtifactRefInline: + return c.inlineReader.Read(ctx, ref) + case domain.ArtifactRefFile: + return c.fileReader.Read(ctx, ref) + default: + return nil, fmt.Errorf("%w: %s", ErrUnsupportedRefType, ref.Type) + } +} + +type inlineReader struct{} + +func (r *inlineReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error) { + if ref.Body == "" { + return nil, ErrMissingInlineBody + } + + body := []byte(ref.Body) + return &domain.Artifact{ + Body: body, + Size: int64(len(body)), + Hash: fmt.Sprintf("%x", sha256.Sum256(body)), + URI: ref.URI, + }, nil +} + +type fileReader struct{} + +func (r *fileReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error) { + if ref.URI == "" { + return nil, ErrMissingFilePath + } + + data, err := os.ReadFile(ref.URI) + if err != nil { + return nil, fmt.Errorf("failed to read file %s: %w", ref.URI, err) + } + + contentType := mime.TypeByExtension(filepath.Ext(ref.URI)) + if contentType == "" { + contentType = "text/plain" // Default + } + + return &domain.Artifact{ + Name: filepath.Base(ref.URI), + ContentType: contentType, + Body: data, + URI: ref.URI, + Size: int64(len(data)), + Hash: fmt.Sprintf("%x", sha256.Sum256(data)), + }, nil +} diff --git a/internal/artifact/reader_test.go b/internal/artifact/reader_test.go new file mode 100644 index 0000000..943b896 --- /dev/null +++ b/internal/artifact/reader_test.go @@ -0,0 +1,101 @@ +package artifact + +import ( + "context" + "os" + "testing" + + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" +) + +func TestCompositeReader_Read(t *testing.T) { + reader := NewCompositeReader() + ctx := context.Background() + + t.Run("inline artifact", func(t *testing.T) { + ref := domain.ArtifactRef{ + Type: domain.ArtifactRefInline, + Body: "hello world", + } + art, err := reader.Read(ctx, ref) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if string(art.Body) != "hello world" { + t.Errorf("expected 'hello world', got %s", string(art.Body)) + } + if art.Hash == "" { + t.Error("expected hash to be computed") + } + }) + + t.Run("inline artifact missing body", func(t *testing.T) { + ref := domain.ArtifactRef{ + Type: domain.ArtifactRefInline, + Body: "", + } + _, err := reader.Read(ctx, ref) + if err == nil || err != ErrMissingInlineBody { + t.Errorf("expected ErrMissingInlineBody, got %v", err) + } + }) + + t.Run("unsupported ref type", func(t *testing.T) { + ref := domain.ArtifactRef{ + Type: domain.ArtifactRefS3, + URI: "s3://bucket/key", + } + _, err := reader.Read(ctx, ref) + if err == nil { + t.Error("expected error for unsupported type") + } + }) +} + +func TestFileReader_Read(t *testing.T) { + content := []byte("test file content") + tmpFile, err := os.CreateTemp("", "artifact_test_*.txt") + if err != nil { + t.Fatal(err) + } + defer os.Remove(tmpFile.Name()) + + if _, err := tmpFile.Write(content); err != nil { + t.Fatal(err) + } + tmpFile.Close() + + reader := NewCompositeReader() + ctx := context.Background() + + t.Run("file artifact loading", func(t *testing.T) { + ref := domain.ArtifactRef{ + Type: domain.ArtifactRefFile, + URI: tmpFile.Name(), + } + art, err := reader.Read(ctx, ref) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if string(art.Body) != string(content) { + t.Errorf("expected %s, got %s", string(content), string(art.Body)) + } + if art.Name == "" { + t.Error("expected name to be inferred from filename") + } + if art.Hash == "" { + t.Error("expected hash to be computed") + } + }) + + t.Run("missing file path", func(t *testing.T) { + ref := domain.ArtifactRef{ + Type: domain.ArtifactRefFile, + URI: "", + } + _, err := reader.Read(ctx, ref) + if err == nil || err != ErrMissingFilePath { + t.Errorf("expected ErrMissingFilePath, got %v", err) + } + }) +} diff --git a/internal/domain/domain.go b/internal/domain/domain.go new file mode 100644 index 0000000..9710e57 --- /dev/null +++ b/internal/domain/domain.go @@ -0,0 +1,179 @@ +package domain + +import ( + "time" +) + +// ArtifactRefType defines how an artifact is referenced. +type ArtifactRefType string + +const ( + ArtifactRefInline ArtifactRefType = "inline" + ArtifactRefFile ArtifactRefType = "file" + ArtifactRefS3 ArtifactRefType = "s3" +) + +// OutputFormat defines the desired format of the generated artifact. +type OutputFormat string + +const ( + FormatText OutputFormat = "text" + FormatMarkdown OutputFormat = "markdown" + FormatJSON OutputFormat = "json" +) + +// ValidationMode defines how the output should be validated. +type ValidationMode string + +const ( + ValidationNone ValidationMode = "none" + ValidationBasic ValidationMode = "basic" + ValidationJSON ValidationMode = "json" + ValidationJSONSchema ValidationMode = "json_schema" +) + +// ValidationStatus defines the result of a validation check. +type ValidationStatus string + +const ( + ValidationPassed ValidationStatus = "passed" + ValidationFailed ValidationStatus = "failed" + ValidationSkipped ValidationStatus = "skipped" +) + +// RunRequest represents a request to generate a single artifact. +type RunRequest struct { + ProfileID string + Inputs map[string]ArtifactRef + Vars map[string]string + Model *ModelTarget + Metadata map[string]string +} + +// RunResult represents the complete result of a prompt execution run. +type RunResult struct { + Artifact Artifact + RawOutput string + Validation ValidationResult + ProfileID string + ProfileVersion string + ModelName string + Endpoint string + InputHashes map[string]string + PromptHash string + Usage TokenUsage + StartTime time.Time + EndTime time.Time + Error error +} + +// ArtifactRef represents a reference to an input artifact. +type ArtifactRef struct { + Type ArtifactRefType + URI string + Body string // Used for inline +} + +// Artifact represents the actual loaded content of a reference. +type Artifact struct { + Name string + ContentType string + Body []byte + URI string + Size int64 + Hash string +} + +// PromptProfile represents a configured prompt execution profile. +type PromptProfile struct { + ID string `yaml:"id"` + Version string `yaml:"version"` + Description string `yaml:"description"` + ExpectedInputs []string `yaml:"expected_inputs"` + Templates []PromptMessageTemplate `yaml:"templates"` + ModelDefaults ModelTarget `yaml:"model_defaults"` + OutputFormat OutputFormat `yaml:"output_format"` + Validation OutputContract `yaml:"validation"` +} + +// PromptMessageTemplate defines a template for a chat message. +type PromptMessageTemplate struct { + Role string `yaml:"role"` + Content string `yaml:"content"` +} + +// ModelTarget represents the LLM endpoint and configuration. +type ModelTarget struct { + Endpoint string `yaml:"endpoint"` + Model string `yaml:"model"` + Temperature float64 `yaml:"temperature"` + MaxTokens int `yaml:"max_tokens"` + TopP float64 `yaml:"top_p"` +} + +// OutputContract defines the requirements for the output artifact. +type OutputContract struct { + Format OutputFormat `yaml:"format"` + ValidationMode ValidationMode `yaml:"validation_mode"` + SchemaPath string `yaml:"schema_path"` + RepairAttempts int `yaml:"repair_attempts"` +} + +// RenderedPrompt represents the prompt after template application. +type RenderedPrompt struct { + Messages []RenderedMessage +} + +// RenderedMessage is a single message in a rendered prompt. +type RenderedMessage struct { + Role string + Content string +} + +// GenerateRequest is the internal request passed to the LLM client. +type GenerateRequest struct { + Prompt RenderedPrompt + Target ModelTarget +} + +// GenerateResponse is the response received from the LLM client. +type GenerateResponse struct { + Content string + Usage TokenUsage +} + +// TokenUsage tracks token consumption. +type TokenUsage struct { + PromptTokens int + CompletionTokens int + TotalTokens int +} + +// ValidationResult represents the outcome of an output validation. +type ValidationResult struct { + Status ValidationStatus + Mode ValidationMode + Errors []string + SchemaPath string + RepairAttempts int + IsValid bool +} + +// RunMetadata contains auditing information for a run. +type RunMetadata struct { + RunID string + ProfileID string + ProfileVersion string + ProfileHash string + PromptHash string + InputHashes map[string]string + ModelEndpoint string + ModelName string + Params ModelTarget + Timestamp time.Time + Duration time.Duration + Usage TokenUsage + ValidationMode ValidationMode + ValidationStatus ValidationStatus + RepairAttempts int +} diff --git a/internal/llm/client.go b/internal/llm/client.go new file mode 100644 index 0000000..a1f799f --- /dev/null +++ b/internal/llm/client.go @@ -0,0 +1,11 @@ +package llm + +import ( + "context" + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" +) + +// Client executes a rendered prompt against an LLM endpoint. +type Client interface { + Generate(ctx context.Context, req domain.GenerateRequest) (*domain.GenerateResponse, error) +} diff --git a/internal/profile/filesystem_repository.go b/internal/profile/filesystem_repository.go new file mode 100644 index 0000000..9632ad5 --- /dev/null +++ b/internal/profile/filesystem_repository.go @@ -0,0 +1,113 @@ +package profile + +import ( + "context" + "errors" + "fmt" + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" + "gopkg.in/yaml.v3" + "os" + "path/filepath" + "strings" +) + +var ( + ErrProfileNotFound = errors.New("prompt profile not found") + ErrInvalidYAML = errors.New("invalid YAML format") + ErrInvalidProfile = errors.New("invalid profile configuration") +) + +type filesystemRepository struct { + dir string +} + +func NewFilesystemRepository(dir string) Repository { + return &filesystemRepository{dir: dir} +} + +func (r *filesystemRepository) GetProfile(ctx context.Context, id string, version string) (*domain.PromptProfile, error) { + files, err := os.ReadDir(r.dir) + if err != nil { + return nil, fmt.Errorf("failed to read profile directory: %w", err) + } + + for _, file := range files { + if file.IsDir() || (!strings.HasSuffix(file.Name(), ".yaml") && !strings.HasSuffix(file.Name(), ".yml")) { + continue + } + + fullPath := filepath.Join(r.dir, file.Name()) + data, err := os.ReadFile(fullPath) + if err != nil { + return nil, fmt.Errorf("failed to read profile file %s: %w", file.Name(), err) + } + + var prof domain.PromptProfile + if err := yaml.Unmarshal(data, &prof); err != nil { + if strings.Contains(file.Name(), id) { + return nil, fmt.Errorf("%w: %s: %v", ErrInvalidYAML, file.Name(), err) + } + continue + } + + if prof.ID == id { + if version == "" || prof.Version == version { + if err := validateProfile(&prof); err != nil { + return nil, fmt.Errorf("%w: %s: %v", ErrInvalidProfile, file.Name(), err) + } + return &prof, nil + } + } + + } + + return nil, ErrProfileNotFound +} + +func validateProfile(p *domain.PromptProfile) error { + if p.ID == "" { + return errors.New("profile id is required") + } + if p.Version == "" { + return errors.New("profile version is required") + } + if len(p.Templates) == 0 { + return errors.New("at least one prompt template message is required") + } + for i, t := range p.Templates { + if t.Role == "" { + return fmt.Errorf("template message %d is missing role", i) + } + if t.Content == "" { + return fmt.Errorf("template message %d is missing content", i) + } + } + if !isValidOutputFormat(p.OutputFormat) { + return fmt.Errorf("invalid output format: %s", p.OutputFormat) + } + if !isValidValidationMode(p.Validation.ValidationMode) { + return fmt.Errorf("invalid validation mode: %s", p.Validation.ValidationMode) + } + for i, input := range p.ExpectedInputs { + if strings.TrimSpace(input) == "" { + return fmt.Errorf("expected input %d has empty name", i) + } + } + return nil +} + +func isValidOutputFormat(f domain.OutputFormat) bool { + switch f { + case domain.FormatText, domain.FormatMarkdown, domain.FormatJSON: + return true + } + return false +} + +func isValidValidationMode(m domain.ValidationMode) bool { + switch m { + case domain.ValidationNone, domain.ValidationBasic, domain.ValidationJSON, domain.ValidationJSONSchema: + return true + } + return false +} diff --git a/internal/profile/repository.go b/internal/profile/repository.go new file mode 100644 index 0000000..2e13b5b --- /dev/null +++ b/internal/profile/repository.go @@ -0,0 +1,11 @@ +package profile + +import ( + "context" + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" +) + +// Repository handles loading and storing prompt profiles. +type Repository interface { + GetProfile(ctx context.Context, id string, version string) (*domain.PromptProfile, error) +} diff --git a/internal/profile/repository_test.go b/internal/profile/repository_test.go new file mode 100644 index 0000000..a99b48c --- /dev/null +++ b/internal/profile/repository_test.go @@ -0,0 +1,76 @@ +package profile + +import ( + "context" + "errors" + "os" + "path/filepath" + "testing" +) + +func TestFilesystemRepository_GetProfile(t *testing.T) { + tmpDir, err := os.MkdirTemp("", "profile_test") + if err != nil { + t.Fatal(err) + } + defer os.RemoveAll(tmpDir) + + testDataDir := "testdata" + files, err := os.ReadDir(testDataDir) + if err != nil { + t.Fatalf("failed to read testdata: %v", err) + } + + for _, f := range files { + src := filepath.Join(testDataDir, f.Name()) + dst := filepath.Join(tmpDir, f.Name()) + data, err := os.ReadFile(src) + if err != nil { + t.Fatal(err) + } + if err := os.WriteFile(dst, data, 0644); err != nil { + t.Fatal(err) + } + } + + repo := NewFilesystemRepository(tmpDir) + ctx := context.Background() + + t.Run("valid profile", func(t *testing.T) { + p, err := repo.GetProfile(ctx, "test-profile", "") + if err != nil { + t.Fatalf("expected no error, got %v", err) + } + if p == nil || p.ID != "test-profile" { + t.Errorf("expected profile test-profile, got %v", p) + } + }) + + t.Run("invalid YAML", func(t *testing.T) { + _, err := repo.GetProfile(ctx, "invalid_yaml", "") + if !errors.Is(err, ErrInvalidYAML) { + t.Errorf("expected ErrInvalidYAML, got %v", err) + } + }) + + t.Run("missing ID", func(t *testing.T) { + _, err := repo.GetProfile(ctx, "missing-id", "") + if !errors.Is(err, ErrProfileNotFound) { + t.Errorf("expected ErrProfileNotFound for profile with missing ID, got %v", err) + } + }) + + t.Run("no templates", func(t *testing.T) { + _, err := repo.GetProfile(ctx, "no-templates", "") + if !errors.Is(err, ErrInvalidProfile) { + t.Errorf("expected ErrInvalidProfile for profile with no templates, got %v", err) + } + }) + + t.Run("profile not found", func(t *testing.T) { + _, err := repo.GetProfile(ctx, "unknown", "") + if !errors.Is(err, ErrProfileNotFound) { + t.Errorf("expected ErrProfileNotFound, got %v", err) + } + }) +} diff --git a/internal/profile/testdata/invalid_yaml.yaml b/internal/profile/testdata/invalid_yaml.yaml new file mode 100644 index 0000000..801ce79 --- /dev/null +++ b/internal/profile/testdata/invalid_yaml.yaml @@ -0,0 +1,5 @@ +id: invalid-yaml +version: 1.0.0 +templates: + - role: system + content: [unclosed bracket diff --git a/internal/profile/testdata/missing_id.yaml b/internal/profile/testdata/missing_id.yaml new file mode 100644 index 0000000..e91fc3a --- /dev/null +++ b/internal/profile/testdata/missing_id.yaml @@ -0,0 +1,8 @@ +version: 1.0.0 +description: Missing ID +templates: + - role: system + content: Hello +output_format: text +validation: + validation_mode: none diff --git a/internal/profile/testdata/no_templates.yaml b/internal/profile/testdata/no_templates.yaml new file mode 100644 index 0000000..19fadd5 --- /dev/null +++ b/internal/profile/testdata/no_templates.yaml @@ -0,0 +1,6 @@ +id: no-templates +version: 1.0.0 +templates: [] +output_format: text +validation: + validation_mode: none diff --git a/internal/profile/testdata/valid.yaml b/internal/profile/testdata/valid.yaml new file mode 100644 index 0000000..9bbe30d --- /dev/null +++ b/internal/profile/testdata/valid.yaml @@ -0,0 +1,17 @@ +id: test-profile +version: 1.0.0 +description: A valid test profile +expected_inputs: + - transcript + - glossary +templates: + - role: system + content: You are a helpful assistant. + - role: user + content: Analyze this: {{.transcript}} +model_defaults: + model: gpt-4o + temperature: 0.7 +output_format: markdown +validation: + validation_mode: basic diff --git a/internal/prompt/go_renderer.go b/internal/prompt/go_renderer.go new file mode 100644 index 0000000..0047aae --- /dev/null +++ b/internal/prompt/go_renderer.go @@ -0,0 +1,71 @@ +package prompt + +import ( + "bytes" + "context" + "errors" + "fmt" + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" + "text/template" +) + +var ( + ErrMissingRequiredInput = errors.New("missing required input artifact") + ErrUnknownInput = errors.New("referenced unknown input artifact") + ErrInvalidTemplate = errors.New("invalid prompt template") + ErrInvalidMessageRole = errors.New("invalid or empty message role") +) + +type goRenderer struct{} + +func NewGoRenderer() Renderer { + return &goRenderer{} +} + +func (r *goRenderer) Render(ctx context.Context, profile *domain.PromptProfile, inputs map[string]*domain.Artifact, vars map[string]string) (*domain.RenderedPrompt, error) { + // 1. Verify required inputs + for _, req := range profile.ExpectedInputs { + if _, ok := inputs[req]; !ok { + return nil, fmt.Errorf("%w: %s", ErrMissingRequiredInput, req) + } + } + + // 2. Setup template functions + funcs := template.FuncMap{ + "input": func(name string) (string, error) { + art, ok := inputs[name] + if !ok { + return "", fmt.Errorf("%w: %s", ErrUnknownInput, name) + } + return string(art.Body), nil + }, + } + + var renderedMessages []domain.RenderedMessage + + for i, tmplMsg := range profile.Templates { + if tmplMsg.Role == "" { + return nil, fmt.Errorf("%w: message %d", ErrInvalidMessageRole, i) + } + + // Parse and execute template + tmpl, err := template.New(fmt.Sprintf("msg_%d", i)).Funcs(funcs).Parse(tmplMsg.Content) + if err != nil { + return nil, fmt.Errorf("%w: message %d: %v", ErrInvalidTemplate, i, err) + } + + var buf bytes.Buffer + if err := tmpl.Execute(&buf, vars); err != nil { + return nil, fmt.Errorf("execution failed for message %d: %v", i, err) + } + + renderedMessages = append(renderedMessages, domain.RenderedMessage{ + Role: tmplMsg.Role, + Content: buf.String(), + }) + } + + return &domain.RenderedPrompt{ + Messages: renderedMessages, + }, nil +} diff --git a/internal/prompt/renderer.go b/internal/prompt/renderer.go new file mode 100644 index 0000000..815ebeb --- /dev/null +++ b/internal/prompt/renderer.go @@ -0,0 +1,11 @@ +package prompt + +import ( + "context" + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" +) + +// Renderer renders prompt templates using named artifacts and variables. +type Renderer interface { + Render(ctx context.Context, profile *domain.PromptProfile, inputs map[string]*domain.Artifact, vars map[string]string) (*domain.RenderedPrompt, error) +} diff --git a/internal/prompt/renderer_test.go b/internal/prompt/renderer_test.go new file mode 100644 index 0000000..3712567 --- /dev/null +++ b/internal/prompt/renderer_test.go @@ -0,0 +1,90 @@ +package prompt + +import ( + "context" + "testing" + + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" +) + +func TestGoRenderer_Render(t *testing.T) { + renderer := NewGoRenderer() + ctx := context.Background() + + profile := &domain.PromptProfile{ + ID: "test-profile", + ExpectedInputs: []string{"transcript"}, + Templates: []domain.PromptMessageTemplate{ + {Role: "system", Content: "You are a {{.role}}."}, + {Role: "user", Content: "Analyze this: {{input \"transcript\"}}"}, + }, + } + + inputs := map[string]*domain.Artifact{ + "transcript": {Body: []byte("The quick brown fox.")}, + } + + vars := map[string]string{ + "role": "helpful assistant", + } + + t.Run("successful render", func(t *testing.T) { + res, err := renderer.Render(ctx, profile, inputs, vars) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if len(res.Messages) != 2 { + t.Errorf("expected 2 messages, got %d", len(res.Messages)) + } + if res.Messages[0].Content != "You are a helpful assistant." { + t.Errorf("unexpected system message: %s", res.Messages[0].Content) + } + if res.Messages[1].Content != "Analyze this: The quick brown fox." { + t.Errorf("unexpected user message: %s", res.Messages[1].Content) + } + }) + + t.Run("missing required input", func(t *testing.T) { + emptyInputs := map[string]*domain.Artifact{} + _, err := renderer.Render(ctx, profile, emptyInputs, vars) + if err == nil || (err != ErrMissingRequiredInput && err.Error() != "missing required input artifact: transcript") { + t.Errorf("expected ErrMissingRequiredInput, got %v", err) + } + }) + + t.Run("unknown input in template", func(t *testing.T) { + profileUnknown := &domain.PromptProfile{ + Templates: []domain.PromptMessageTemplate{ + {Role: "user", Content: "Hello {{input \"ghost\"}}"}, + }, + } + _, err := renderer.Render(ctx, profileUnknown, inputs, vars) + if err == nil { + t.Error("expected error for unknown input") + } + }) + + t.Run("invalid template syntax", func(t *testing.T) { + profileInvalid := &domain.PromptProfile{ + Templates: []domain.PromptMessageTemplate{ + {Role: "user", Content: "Hello {{.unclosed"}, + }, + } + _, err := renderer.Render(ctx, profileInvalid, inputs, vars) + if err == nil { + t.Error("expected error for invalid template syntax") + } + }) + + t.Run("empty message role", func(t *testing.T) { + profileNoRole := &domain.PromptProfile{ + Templates: []domain.PromptMessageTemplate{ + {Role: "", Content: "Hello"}, + }, + } + _, err := renderer.Render(ctx, profileNoRole, inputs, vars) + if err == nil { + t.Error("expected error for empty message role") + } + }) +} diff --git a/internal/validate/validator.go b/internal/validate/validator.go new file mode 100644 index 0000000..851af44 --- /dev/null +++ b/internal/validate/validator.go @@ -0,0 +1,11 @@ +package validate + +import ( + "context" + "gitea.maximumdirect.net/eric/scriptorium/internal/domain" +) + +// Validator validates the generated artifact based on the output contract. +type Validator interface { + Validate(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract) (domain.ValidationResult, error) +}