79 Commits

Author SHA1 Message Date
2e76003fd5 Prepare documentation for the v0.8.0 release 2026-08-25 12:19:17 +00:00
36ce5a5099 Fix output repair diagnostics and concurrency tests 2026-08-25 11:54:50 +00:00
465dc1389d Document bounded output repair 2026-08-25 09:58:10 +00:00
ae6f1a9865 Enable bounded output repair in the engine 2026-08-25 09:55:19 +00:00
ee99dc9478 Run bounded output repair attempts 2026-08-25 09:51:28 +00:00
00ee5893e9 Build bounded output repair requests 2026-08-25 09:46:56 +00:00
64d1cffd89 Preserve explicit empty model responses 2026-08-25 09:42:28 +00:00
f9e8afa2c3 Enforce bounded output repair contracts 2026-08-25 09:41:15 +00:00
e827631d8c Plan the public output repair implementation 2026-08-25 09:38:06 +00:00
115fe8ba58 Plan bounded output repair 2026-08-25 03:33:33 +00:00
c53250f023 Prepare documentation for the v0.7.0 release 2026-08-25 02:13:43 +00:00
3d99483219 Preserve cancellation after profile lookups 2026-08-25 02:06:47 +00:00
67f788b1e2 Document profile inheritance behavior 2026-08-25 01:46:16 +00:00
764103a2e2 Resolve inherited profiles in engine workflows 2026-08-25 01:41:21 +00:00
a08dd83d1f Add profile inheritance resolver 2026-08-25 01:38:15 +00:00
e8922d8ec5 Add profile inheritance definition support 2026-08-25 01:34:40 +00:00
2d44305a8a Document optional API key environment behavior 2026-08-25 00:44:26 +00:00
a11c80291e Allow unauthenticated optional API key requests 2026-08-25 00:40:49 +00:00
3239567297 Make optional API key environments nonblocking 2026-08-25 00:39:26 +00:00
c239304c2a Document structured generation errors 2026-08-23 19:04:00 +00:00
159b02116f Expose structured generation errors to consumers 2026-08-23 19:01:38 +00:00
e5b7adfb49 Return structured errors for provider status failures 2026-08-23 18:58:15 +00:00
fa2384e696 Bound provider error response body reads 2026-08-23 18:56:15 +00:00
af0bd3f31a Add internal structured provider error parsing 2026-08-23 18:54:42 +00:00
5fff8cd623 Retire the completed Rakestrawhome backend roadmaps 2026-08-23 17:34:14 +00:00
b2a6c47778 Document the Rakestrawhome built-in backend 2026-08-23 17:19:12 +00:00
a1805fe550 Add the Rakestrawhome Gemma profile 2026-08-23 17:12:17 +00:00
93af155254 Add the Rakestrawhome built-in backend 2026-08-23 17:10:34 +00:00
d783b687a5 Plan the built-in Rakestrawhome backend 2026-08-23 17:00:01 +00:00
4ca3be2c14 Finish audit remediation and prepare v0.6.0 2026-08-12 13:01:05 +00:00
227fb35f99 Centralize the maintainer validation workflow 2026-08-12 00:04:48 +00:00
e291b8bfe9 Consolidate provider transport test scaffolding 2026-08-11 23:58:31 +00:00
2b6a7f83c4 Bound and strictly decode provider responses 2026-08-11 23:46:16 +00:00
3a43550f70 Validate and compose provider endpoints 2026-08-11 23:38:45 +00:00
c281f721bc Preserve transport error identities 2026-08-11 23:28:00 +00:00
350b0e76d9 Preserve repair settings and cumulative usage 2026-08-11 23:21:34 +00:00
e43350fd0d Make validation cancellation authoritative 2026-08-11 23:14:01 +00:00
20d3e3b5ee Escape schema resources and reuse compiled plans 2026-08-11 23:05:11 +00:00
a93b799236 Preserve exact JSON validation semantics 2026-08-11 22:55:17 +00:00
e83a3ce179 Honor cancellation while rendering prompts 2026-08-11 22:48:20 +00:00
a04a3bbc5f Correct artifact file and empty input handling 2026-08-11 22:41:37 +00:00
731b66cff5 Avoid decoding unrelated profiles 2026-08-11 22:32:33 +00:00
70e0ea0cf0 Correct profile source validation and identity 2026-08-11 22:28:48 +00:00
d45c474c1e Unify prompt repository source handling 2026-08-11 22:18:50 +00:00
25f1ba0b30 Correct prompt definition selection and decoding 2026-08-11 22:11:39 +00:00
a718762da1 Contain prompt content paths within source roots 2026-08-11 22:03:32 +00:00
58ac3ce298 Correct engine construction edge cases 2026-08-11 21:54:47 +00:00
57f2ce1ce4 Harden public ownership and diagnostic contracts 2026-08-11 21:45:54 +00:00
c8b6d5c490 Consolidate public JSON serialization 2026-08-11 21:40:12 +00:00
abeb50b525 Bound JSON-compatible value copying 2026-08-11 21:32:24 +00:00
1cb07c7d91 Centralize output contract validation 2026-08-11 21:21:33 +00:00
8cfc71c351 Centralize execution setting and session validation 2026-08-11 21:12:56 +00:00
5ccfa4a345 Plan the codebase audit remediation 2026-08-11 20:10:14 +00:00
14e03f19d0 Consolidate and close the codebase audit 2026-08-11 17:21:29 +00:00
7c562a9374 Document cross-cutting architecture audit findings 2026-08-11 17:10:59 +00:00
ef97d85ac9 Document repository-wide test strategy audit 2026-08-11 17:00:59 +00:00
805e48c873 Document capacity scheduling audit results 2026-08-11 16:51:28 +00:00
e9e126dcba Document OpenAI transport audit findings 2026-08-11 16:42:55 +00:00
32e7a3557c Document prepared execution lifecycle audit 2026-08-11 16:31:05 +00:00
1d1b04e2e0 Record ordinary execution and repair audit findings 2026-08-11 16:22:09 +00:00
9748897751 Document inspection and target resolution audit findings 2026-08-11 16:09:50 +00:00
9d020039d5 Record output validation audit findings 2026-08-11 15:59:46 +00:00
5247ce0b73 Record artifact loading and prompt rendering audit findings 2026-08-11 15:41:57 +00:00
c434aa1dae Record profile source audit findings 2026-08-11 15:29:23 +00:00
ac9b3f3d80 Record prompt source audit findings 2026-08-11 15:18:17 +00:00
4f12a89a1b Record backend registry and defaults audit findings 2026-08-11 15:05:15 +00:00
df31e7f58e Record domain and JSON value audit findings 2026-08-11 14:54:56 +00:00
0678d242b9 Record engine operation audit findings 2026-08-11 14:43:20 +00:00
3b4ea21208 Record engine construction audit findings 2026-08-11 14:36:33 +00:00
1430e85147 Record configuration and adapter audit findings 2026-08-11 14:28:07 +00:00
34d7a19da5 Record public value and error audit findings 2026-08-11 14:17:55 +00:00
ebf1602635 Prepare the codebase audit plan 2026-08-11 14:05:30 +00:00
31f2ce3a09 Document Promptkit v0.5.0 2026-08-01 13:18:46 +00:00
fd06e4ca6b Clean up roadmap and fallback profile guidance 2026-08-01 13:16:26 +00:00
e63b8de1e9 Complete application fallback profile implementation 2026-08-01 12:39:01 +00:00
9354d2b373 Add application fallback profile sources 2026-08-01 12:37:01 +00:00
01ca5430bd Move profile composition to the engine facade 2026-08-01 12:31:50 +00:00
ae2179d103 Complete optional parameter omission 2026-08-01 02:43:02 +00:00
a248433d0f Omit unset optional request parameters 2026-08-01 02:41:44 +00:00
95 changed files with 12702 additions and 3748 deletions

View File

@@ -33,6 +33,18 @@ boundary and constraints that framework work must preserve.
## Release Guidance ## Release Guidance
Consumers upgrading from `v0.7.0` to `v0.8.0` should read the
[v0.8.0 changelog and migration guide](docs/releases/v0.8.0.md).
Earlier adopters can consult the
[v0.7.0 changelog and migration guide](docs/releases/v0.7.0.md).
Consumers upgrading from `v0.5.0` to `v0.6.0` can consult the
[v0.6.0 changelog and migration guide](docs/releases/v0.6.0.md).
Consumers upgrading from `v0.4.0` to `v0.5.0` can consult the
[v0.5.0 changelog and migration guide](docs/releases/v0.5.0.md).
Consumers upgrading from `v0.3.0` to `v0.4.0` should read the Consumers upgrading from `v0.3.0` to `v0.4.0` should read the
[v0.4.0 changelog and adoption guide](docs/releases/v0.4.0.md). [v0.4.0 changelog and adoption guide](docs/releases/v0.4.0.md).

View File

@@ -9,6 +9,10 @@ import (
// backend. // backend.
const BackendOpenRouter = backend.OpenRouterID const BackendOpenRouter = backend.OpenRouterID
// BackendRakestrawHome is the reserved ID of Promptkit's built-in
// Rakestrawhome backend.
const BackendRakestrawHome = backend.RakestrawHomeID
// BackendLocal is the case-sensitive conventional ID used by [LocalBackend]. // BackendLocal is the case-sensitive conventional ID used by [LocalBackend].
// It is not a built-in or reserved backend and must be registered with // It is not a built-in or reserved backend and must be registered with
// [WithBackend]. // [WithBackend].
@@ -20,21 +24,25 @@ const BackendLocal = "local"
// to this configuration value do not break source compatibility. // to this configuration value do not break source compatibility.
type Backend struct { type Backend struct {
// ID is the stable, case-sensitive registry key. NewEngine trims it and // ID is the stable, case-sensitive registry key. NewEngine trims it and
// requires a non-blank value. BackendOpenRouter is reserved. // requires a non-blank value. Built-in backend IDs are reserved.
ID string ID string
// Endpoint is the OpenAI-compatible base endpoint. NewEngine trims it and // Endpoint is the OpenAI-compatible base endpoint. NewEngine trims it and
// requires an absolute HTTP or HTTPS URL with a host and without user // requires an absolute HTTP or HTTPS URL with a host and without user
// information, a query string, or a fragment. Paths are allowed. // information, a query string, or a fragment. Paths are allowed.
Endpoint string Endpoint string
// APIKeyEnv optionally names the environment variable containing the API // APIKeyEnv optionally names an environment lookup source for an API key.
// key. NewEngine trims it and requires the portable form // NewEngine trims it and requires the portable form [A-Za-z_][A-Za-z0-9_]*.
// [A-Za-z_][A-Za-z0-9_]*. Store only the name, never a credential value. // A direct RunRequest.APIKey takes precedence. When no usable credential is
// available, the built-in client omits Authorization; injected clients own
// their own credential-resolution behavior. Store only the name, never a
// credential value.
APIKeyEnv string APIKeyEnv string
// ExtraParams contains backend-wide request defaults. Values must be // ExtraParams contains backend-wide request defaults. Values must be
// JSON-compatible, finite, acyclic, and keyed by non-empty strings. Keys // JSON-compatible, finite, acyclic, and keyed by non-empty strings. Keys
// must not be model, session_id, messages, temperature, max_tokens, top_p, // must not be model, session_id, messages, temperature, max_tokens, top_p,
// service_tier, reasoning_effort, or response_format. An empty map supplies // service_tier, reasoning_effort, or response_format. An empty map supplies
// no defaults. NewEngine deeply copies the map. // no defaults. NewEngine deeply copies the map and rejects excessively deep
// or large values for safety.
ExtraParams map[string]any ExtraParams map[string]any
// ConcurrencyLimit is the maximum number of simultaneous model-generation // ConcurrencyLimit is the maximum number of simultaneous model-generation
// calls allowed for this backend within one Engine. Zero leaves the backend // calls allowed for this backend within one Engine. Zero leaves the backend
@@ -73,11 +81,11 @@ func LocalBackend(endpoint string, concurrencyLimit int) Backend {
// //
// Registrations accumulate in option order. Every normalized ID must be unique // Registrations accumulate in option order. Every normalized ID must be unique
// across consumer registrations and built-ins; a duplicate or invalid // across consumer registrations and built-ins; a duplicate or invalid
// definition makes NewEngine fail with ErrInvalidConfig. In particular, // definition makes NewEngine fail with ErrInvalidConfig. Built-in IDs,
// BackendOpenRouter cannot be replaced. The immutable registration is scoped // including [BackendOpenRouter] and [BackendRakestrawHome], cannot be
// to the resulting Engine and cannot be enumerated, replaced, removed, or // replaced. The immutable registration is scoped to the resulting Engine and
// mutated after construction. WithBackend does not install package-global // cannot be enumerated, replaced, removed, or mutated after construction.
// state. // WithBackend does not install package-global state.
func WithBackend(backend Backend) Option { func WithBackend(backend Backend) Option {
queueCapacity := 0 queueCapacity := 0
queueCapacitySet := backend.QueueCapacity != nil queueCapacitySet := backend.QueueCapacity != nil

16
doc.go
View File

@@ -23,8 +23,9 @@
// InspectProfile return copied inspection values. Returned values and values // InspectProfile return copied inspection values. Returned values and values
// passed to extension interfaces are likewise isolated from engine state. // passed to extension interfaces are likewise isolated from engine state.
// Callers own those copies and may mutate them after the call that supplied or // Callers own those copies and may mutate them after the call that supplied or
// returned them. Returned structured errors are likewise caller-owned and may // returned them. [CapacityError] values are caller-owned and may be mutated
// be mutated without affecting engine state or another error. // without affecting engine state or another error. Immutable [GenerationError]
// values are also caller-owned and do not retain shared engine state.
// //
// # Security and sensitive data // # Security and sensitive data
// //
@@ -53,9 +54,14 @@
// Construction, inspection, handle, and error values, including [Config], // Construction, inspection, handle, and error values, including [Config],
// [Backend], [RunRequest], [ArtifactRef], [ExecutionTargetOverride], [Profile], // [Backend], [RunRequest], [ArtifactRef], [ExecutionTargetOverride], [Profile],
// [OpenAICompatibleProfileConfig], [ProfileInspection], // [OpenAICompatibleProfileConfig], [ProfileInspection],
// [PromptInputDefinition], [PromptInspection], [PreparedExecution], and // [PromptInputDefinition], [PromptInspection], [PreparedExecution],
// [CapacityError], do not have stable JSON representations. Direct API keys // [CapacityError], and [GenerationError], do not have stable JSON
// are nevertheless excluded from JSON for every public value. // representations. Direct API keys are nevertheless excluded from JSON for
// every public value.
// Provider-derived [GenerationError] accessor values are untrusted and can
// contain sensitive request or schema fragments. Applications must apply their
// own disclosure policy before logging, displaying, or returning them.
// //
// JSON timestamps use time.Time's RFC 3339 encoding and are omitted when zero. // JSON timestamps use time.Time's RFC 3339 encoding and are omitted when zero.
// PreparedRun and RunResult durations are encoded as integer milliseconds in // PreparedRun and RunResult durations are encoded as integer milliseconds in

View File

@@ -40,6 +40,36 @@ validation, and default transport behavior. Source discovery, format
validation, and profile precedence are defined by the validation, and profile precedence are defined by the
[framework format reference](../formats.md). [framework format reference](../formats.md).
## Supply Embedded Application Defaults
Use `WithFallbackProfileFS` when an application packages profile definitions
that should apply unless an operator provides an ordinary configured profile
with the same ID. For example, an application can embed its defaults while
continuing to use `ProfileDir` for operator overrides:
```go
//go:embed profiles/*.yaml
var applicationProfiles embed.FS
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: "prompts",
ProfileDir: operatorProfileDir,
},
promptkit.WithFallbackProfileFS(applicationProfiles, "profiles"),
)
```
Keep application-owned profile IDs and definitions in the embedded source.
Use the ordinary configured profile source for operator overrides. Leave
`operatorProfileDir` empty when the operator did not configure an override
directory; a non-empty path names an authoritative higher-precedence source,
so an unavailable or unreadable directory is an error rather than a reason to
fall back. The
[framework format reference](../formats.md#source-and-profile-precedence)
owns the exact profile format and lookup order; the
[`WithFallbackProfileFS` GoDoc](../../engine.go) owns its option contract and
validation rules.
## Inspect A Prompt Before Preparation ## Inspect A Prompt Before Preparation
Use [`Engine.InspectPrompt`](../../engine.go) to check one configured prompt's Use [`Engine.InspectPrompt`](../../engine.go) to check one configured prompt's
@@ -143,6 +173,29 @@ semantics. The
[OpenAI-compatible integration contract](../integrations/openai-compatible-chat.md) [OpenAI-compatible integration contract](../integrations/openai-compatible-chat.md)
owns the built-in client's outbound HTTP behavior. owns the built-in client's outbound HTTP behavior.
### Repair A Structured Result
Set a small additional-call budget when a structurally invalid result can be
corrected automatically:
```go
request.Validation = &promptkit.OutputContract{
Format: promptkit.FormatJSON,
ValidationMode: promptkit.ValidationJSONSchema,
SchemaPath: "events.schema.json",
RepairAttempts: 1,
}
```
Each repair attempt is another model call, so it can increase latency and
usage; `RunResult.Usage` is cumulative and `Validation.RepairAttempts` reports
calls actually started. Exhaustion still returns the final failed validation
result. `basic` validation can also repair an empty candidate, but structural
validity is not evidence of factual or domain correctness. See the
[output-contract format reference](../formats.md#output-contract) and
[`OutputContract` GoDoc](../../types.go) for the exact budget and eligibility
rules.
## Inputs, Profiles, And Overrides ## Inputs, Profiles, And Overrides
Use `File`, `Inline`, or `InlineWithURI` to construct request inputs. A request Use `File`, `Inline`, or `InlineWithURI` to construct request inputs. A request
@@ -152,13 +205,60 @@ replace execution settings or the complete output contract.
The [public value GoDoc](../../types.go) defines nil, empty, zero, replacement, The [public value GoDoc](../../types.go) defines nil, empty, zero, replacement,
copy, and credential behavior. The copy, and credential behavior. The
[framework format reference](../formats.md) defines how those request values [framework format reference](../formats.md) defines how those request values
interact with prompt definitions, file-backed profiles, built-ins, schemas, interact with prompt definitions, file-backed and application fallback
and framework defaults. profiles, built-ins, schemas, and framework defaults.
For programmatic profiles, For programmatic profiles,
[`OpenAICompatibleProfile`](../../profiles.go) converts ordinary [`OpenAICompatibleProfile`](../../profiles.go) converts ordinary
OpenAI-compatible settings into a value accepted by `WithProfiles`. OpenAI-compatible settings into a value accepted by `WithProfiles`.
### Alias A Built-In Profile
Give an application-owned profile ID a built-in base when prompts should select
the application ID while inheriting the built-in target. The child can override
only the setting it owns:
```go
promptkit.WithProfiles(promptkit.Profile{
ID: "weather-light",
BaseProfileID: "deepseek-4-flash",
ReasoningEffort: "high",
})
```
Select `weather-light` in a prompt or `RunRequest.ProfileID`; it remains the
reported selected profile. See the [profile inheritance format
reference](../formats.md#profile-inheritance) and the
[`Profile` GoDoc](../../types.go) for exact lookup, merging, and validation
behavior.
### Use The Rakestrawhome Built-In Profile
Set `RAKESTRAWHOME_INFERENCE_API_KEY` in the application environment, then
select `rakestrawhome-gemma-4-31b` as an ordinary profile ID. For example, a
prepared result identifies the selected built-in through
`BackendRakestrawHome`:
```go
prepared, err := engine.Prepare(ctx, promptkit.RunRequest{
PromptID: "meeting.summary",
ProfileID: "rakestrawhome-gemma-4-31b",
Inputs: inputs,
})
if err != nil {
return err
}
if prepared.SelectedBackendID != promptkit.BackendRakestrawHome {
return fmt.Errorf("unexpected backend %q", prepared.SelectedBackendID)
}
```
Do not register `rakestrawhome` manually. When adopting this built-in, remove
an existing `WithBackend` registration with that exact ID; retaining it causes
the intentional duplicate-ID configuration error. Direct request credentials
and runtime endpoint overrides remain supported under their ordinary GoDoc and
format contracts.
### Inspect A Profile Before Prompt Work ### Inspect A Profile Before Prompt Work
Use [`Engine.InspectProfile`](../../engine.go) to validate one configured Use [`Engine.InspectProfile`](../../engine.go) to validate one configured
@@ -174,7 +274,7 @@ if err != nil {
target := inspection.EffectiveModelParams target := inspection.EffectiveModelParams
if target.APIKeyEnv != "" { if target.APIKeyEnv != "" {
// Apply application policy for the named environment variable. // This is a configured optional environment lookup source.
} else if inspection.APIKeyRequired { } else if inspection.APIKeyRequired {
// Arrange a direct credential before later execution. // Arrange a direct credential before later execution.
} }
@@ -183,9 +283,11 @@ if target.APIKeyEnv != "" {
Use this configuration-time boundary when only the profile and its target need Use this configuration-time boundary when only the profile and its target need
checking. Use `Prepare` when the application also needs prompt, input, schema, checking. Use `Prepare` when the application also needs prompt, input, schema,
or rendering work; use prepared execution when that work must remain tied to a or rendering work; use prepared execution when that work must remain tied to a
later execution. Inspection reports credential requirements but leaves the later execution. A reported `APIKeyEnv` is a configured optional source, while
timing of credential enforcement to the application. The method's `APIKeyRequired` is the explicit local requirement. The
[GoDoc](../../engine.go) owns its exact result and error contract. [credential format reference](../formats.md#credentials) and the method's
[GoDoc](../../engine.go) own the exact precedence, timing, result, and error
contracts.
### Set A Per-Run Session And Reasoning ### Set A Per-Run Session And Reasoning
@@ -393,6 +495,24 @@ status; those choices remain with the consuming application. The
contract, while the [`Engine.Run` and error GoDoc](../../engine.go) owns broad contract, while the [`Engine.Run` and error GoDoc](../../engine.go) owns broad
error and cancellation identities. error and cancellation identities.
For a non-2xx response from the built-in OpenAI-compatible client, inspect the
status and deliberately selected provider diagnostic when useful:
```go
var generationErr *promptkit.GenerationError
if errors.As(err, &generationErr) {
status := generationErr.StatusCode()
message := generationErr.ProviderMessage()
_, _ = status, message // Apply application retry and presentation policy.
}
```
All provider fields are untrusted and can contain sensitive request or schema
fragments. Do not log, display, or return them without an application-specific
disclosure policy. Promptkit does not assign retry or presentation behavior.
The [`GenerationError` GoDoc](../../generation_error.go) owns the exact typed
error contract.
## Application Boundary ## Application Boundary
Promptkit is an importable library. It does not own a command, inbound HTTP Promptkit is an importable library. It does not own a command, inbound HTTP

View File

@@ -45,11 +45,17 @@ For cross-cutting changes, follow every applicable row. Do not create
placeholder documents for packages, APIs, or integrations that do not yet placeholder documents for packages, APIs, or integrations that do not yet
exist. exist.
## Maintainer-Run Validation ## Maintainer Validation
Promptkit does not currently use hosted CI. Maintainers are responsible for This section is the canonical local validation workflow for Promptkit. Run
running the documented checks before accepting changes. Run the default Go every command from the repository root before accepting a change. The test
validation from the Promptkit repository root: suite and maintained examples are deterministic, offline, and require no real
provider credentials.
### Tests, Analysis, Build, And Examples
Run the ordinary and race-enabled suites, static analysis, the build, and both
maintained consumer examples:
```sh ```sh
go test ./... go test ./...
@@ -57,85 +63,171 @@ go test -race ./...
go vet ./... go vet ./...
go build ./... go build ./...
go run ./examples/go-library/prepare go run ./examples/go-library/prepare
go run ./examples/go-library/run
``` ```
Check formatting across every tracked Go file: Both examples must exit successfully. Review their JSON output: preparation
must report the selected offline prompt, profile, model, and message count;
execution must report the deterministic generated output, successful
validation, selected offline model, and usage. Neither command may contact a
provider or require credentials.
### Go Formatting
Check every tracked Go file. The final command must succeed and the captured
list must be empty:
```sh ```sh
gofmt -l $(git ls-files '*.go') unformatted=$(
git ls-files '*.go' |
while IFS= read -r go_file
do
gofmt -l "$go_file"
done
)
test -z "$unformatted"
``` ```
The formatting command must produce no paths. Follow every added or changed ### Local Markdown Links
Markdown link and confirm its target exists. Finally, check whitespace:
Use the Python standard library to verify every repository-relative Markdown
target and local heading fragment. The check is offline and prints nothing on
success:
```sh
python3 - <<'PY'
from pathlib import Path
import re
import subprocess
import sys
from urllib.parse import unquote
root = Path.cwd().resolve()
markdown_files = [
root / name
for name in subprocess.check_output(
["git", "ls-files", "*.md"], text=True
).splitlines()
]
link_pattern = re.compile(r"!?\[[^]]*\]\(([^)]+)\)")
heading_pattern = re.compile(r"^#{1,6}\s+(.+?)\s*#*\s*$")
scheme_pattern = re.compile(r"^[a-z][a-z0-9+.-]*:", re.IGNORECASE)
def markdown_lines(path):
in_fence = False
fence = ""
for line in path.read_text(encoding="utf-8").splitlines():
stripped = line.lstrip()
marker = stripped[:3]
if marker in {"```", "~~~"}:
if not in_fence:
in_fence = True
fence = marker
elif marker == fence:
in_fence = False
fence = ""
continue
if not in_fence:
yield line
anchor_cache = {}
def anchors(path):
if path in anchor_cache:
return anchor_cache[path]
found = set()
counts = {}
for line in markdown_lines(path):
match = heading_pattern.match(line)
if not match:
continue
heading = re.sub(r"<[^>]+>", "", match.group(1)).replace("`", "")
base = re.sub(r"[^\w\- ]", "", heading.lower()).replace(" ", "-")
count = counts.get(base, 0)
counts[base] = count + 1
found.add(base if count == 0 else f"{base}-{count}")
anchor_cache[path] = found
return found
failures = []
for source in markdown_files:
text = "\n".join(markdown_lines(source))
for match in link_pattern.finditer(text):
target = match.group(1).strip()
if target.startswith("<") and target.endswith(">"):
target = target[1:-1]
if scheme_pattern.match(target) or target.startswith("//"):
continue
path_text, separator, fragment = target.partition("#")
destination = source if not path_text else source.parent / unquote(path_text)
try:
destination = destination.resolve()
destination.relative_to(root)
except ValueError:
failures.append(f"{source.relative_to(root)}: escapes repository: {target}")
continue
if not destination.exists():
failures.append(f"{source.relative_to(root)}: missing target: {target}")
continue
if separator and destination.suffix.lower() == ".md":
fragment = unquote(fragment).lower()
if fragment not in anchors(destination):
failures.append(f"{source.relative_to(root)}: missing anchor: {target}")
if failures:
print("\n".join(failures), file=sys.stderr)
raise SystemExit(1)
PY
```
### Repository Hygiene And Review
Reject an active Go workspace, tracked workspace files, a vendor tree, or a
module replacement:
```sh
case "$(go env GOWORK)" in
''|off) ;;
*) printf '%s\n' 'an active Go workspace is not allowed' >&2; exit 1 ;;
esac
test -z "$(git ls-files go.work go.work.sum)"
test ! -e vendor
if grep -Eq '^[[:space:]]*replace([[:space:]]|\()' go.mod
then
printf '%s\n' 'go.mod contains a replacement' >&2
exit 1
fi
```
Check whitespace in both unstaged and staged changes. List ignored files and
scan tracked content for common credential forms:
```sh ```sh
git diff --check git diff --check
git diff --cached --check
test -z "$(git ls-files --others --ignored --exclude-standard)"
credential_pattern='-----BEGIN ([A-Z0-9]+ )?PRIV''ATE KEY-----|AKI''A[0-9A-Z]{16}|gh[pousr]_[A-Za-z0-9]{36,}|sk-[A-Za-z0-9]{32,}'
if git grep -nEI -e "$credential_pattern" -- .
then
printf '%s\n' 'possible credential found' >&2
exit 1
fi
``` ```
Documentation-only work does not require unrelated new tests, but it still Inspect `git status --short --untracked-files=all` and the complete diff before
requires link validation and `git diff --check`. Run the Go validation whenever accepting a change. The status may contain only the intended source changes
documentation changes commands, examples, generated output, or another during development. Reject credentials, private keys, environment files,
behavior checked by the module. generated binaries, test or coverage output, downloaded assets, template
residue, and any other artifact that does not belong in source control. The
credential scan catches common forms but does not replace inspection of the
actual change.
## Focused Validation After committing the accepted change, require a clean candidate:
Use focused checks while iterating, then run the complete validation sequence
before accepting the change. The root package supports:
```sh ```sh
go test . test -z "$(git status --porcelain)"
go vet .
go build .
``` ```
Filter tests by name without assuming a fixed internal package layout:
```sh
go test ./... -run 'TestName'
```
Replace `TestName` with a useful regular expression. Target only paths that
exist, and consult the internal component overview for their owning
documentation. A filtered or package-specific run does not replace the
complete repository validation.
## Coordinated Work With Scriptorium
Promptkit and Scriptorium must remain independently valid. For temporary local
integration, use either a Go workspace outside both repositories or an
uncommitted replacement in the consuming module.
If the repositories are sibling directories, run the workspace commands from
their parent directory:
```sh
go work init ./promptkit ./scriptorium
go work sync
```
Use the workspace only for coordinated local checks. From the same parent
directory, remove it when finished:
```sh
rm -f go.work go.work.sum
```
Alternatively, from the Scriptorium repository root, temporarily point its
Promptkit dependency at the sibling checkout:
```sh
go mod edit -replace gitea.maximumdirect.net/eric/promptkit=../promptkit
```
After coordinated checks, remove the replacement and reconcile module
metadata:
```sh
go mod edit -dropreplace gitea.maximumdirect.net/eric/promptkit
go mod tidy
```
Never commit `go.work`, `go.work.sum`, or a local filesystem `replace`
directive. Before committing in either repository, inspect its module files and
working tree independently. Published consumer versions must depend on a tagged
Promptkit version, not a workspace, local replacement, or unpublished commit.

View File

@@ -9,9 +9,10 @@ explains how to select these sources and invoke the engine. The
owns the resulting outbound wire behavior. owns the resulting outbound wire behavior.
Prompt and profile sources recursively discover files ending in `.yaml` or Prompt and profile sources recursively discover files ending in `.yaml` or
`.yml`. YAML decoding is strict: unknown fields are errors for the selected `.yml`. Each prompt-definition and profile file contains exactly one YAML
definition. Definitions are selected by their YAML `id`, not their file name document; comments and trailing whitespace are allowed. YAML decoding is
or directory. strict: unknown fields are errors for the selected definition. Definitions are
selected by their YAML `id`, not their file name or directory.
## Prompt Definitions ## Prompt Definitions
@@ -86,9 +87,16 @@ Each message has a non-empty `role` and exactly one of:
- `content`, containing an inline Go template; or - `content`, containing an inline Go template; or
- `content_file`, naming a file whose contents are the Go template. - `content_file`, naming a file whose contents are the Go template.
For directory and `fs.FS` prompt sources, `content_file` resolves relative to `content_file` must be a relative path. It resolves from the directory that
the prompt file and remains within the source root. `WithPromptFile` also contains the prompt file and must remain within the configured prompt source
resolves it relative to that file. root; parent components are allowed only when the resolved target remains
inside that root. Absolute paths and paths that escape the root are rejected.
Operating-system directory and single-file sources also reject symlink targets
outside the root, while injected `fs.FS` sources apply containment in that
filesystem's relative path namespace. For `WithPromptFile`, the source root is
the directory containing the selected prompt file. Promptkit uses the parsed
path text exactly after checking separately that it is not blank, so leading
and trailing whitespace can name real filesystem entries.
Request variables are the template data, so a variable named `audience` is Request variables are the template data, so a variable named `audience` is
referenced as `{{.audience}}`. The `{{input "note"}}` helper renders the body referenced as `{{.audience}}`. The `{{input "note"}}` helper renders the body
@@ -118,7 +126,7 @@ outbound integration determines its wire representation.
| `format` | yes | `text`, `markdown`, or `json`. | | `format` | yes | `text`, `markdown`, or `json`. |
| `validation_mode` | yes | `none`, `basic`, `json`, or `json_schema`. | | `validation_mode` | yes | `none`, `basic`, `json`, or `json_schema`. |
| `schema_path` | for `json_schema` | Path to a schema in the configured schema source. | | `schema_path` | for `json_schema` | Path to a schema in the configured schema source. |
| `repair_attempts` | no | Integer zero or greater; omitted means zero. | | `repair_attempts` | no | Integer from zero through three; omitted means zero. A positive value requires `basic`, `json`, or `json_schema` validation. |
The validation modes behave as follows: The validation modes behave as follows:
@@ -128,14 +136,31 @@ The validation modes behave as follows:
- `json_schema` requires valid JSON that satisfies the selected schema. - `json_schema` requires valid JSON that satisfies the selected schema.
`format` controls output artifact metadata. JSON Schema mode also supplies the `format` controls output artifact metadata. JSON Schema mode also supplies the
schema to compatible model clients as structured-output metadata. The public schema to compatible model clients as structured-output metadata. Plain `json`
engine does not install an output repairer, so its validation is single-pass validation accepts every valid JSON value and does not request a provider-native
even when a positive `repair_attempts` value is present. JSON-object constraint.
`repair_attempts` counts additional generation calls after a failed validation.
Zero is single-pass. With a positive eligible budget, Promptkit stops at the
first valid candidate. If the budget is exhausted, it returns the final
candidate and its complete failed validation result; generation and operational
validation failures remain errors. `none` never permits repair.
A request-level `OutputContract` replaces the complete prompt output contract. A request-level `OutputContract` replaces the complete prompt output contract.
It does not merge individual fields. If its format is empty, Promptkit uses It does not merge individual fields. If its format is empty, Promptkit uses
`text`. `text`.
## Built-In Backends
Every engine provides these reserved OpenAI-compatible backend IDs. Consumers
must not register either ID with `WithBackend`; exact registration and
reservation behavior belongs to the [`Backend` GoDoc](../backends.go).
| ID | Base endpoint | API-key environment variable | Active generation limit | Default queue capacity |
| --- | --- | --- | ---: | ---: |
| `openrouter` | `https://openrouter.ai/api/v1` | `OPENROUTER_API_KEY` | 16 | 1024 |
| `rakestrawhome` | `https://inference.ai.rakestrawhome.com/v1` | `RAKESTRAWHOME_INFERENCE_API_KEY` | 4 | 1024 |
## Profile Definitions ## Profile Definitions
A profile supplies model execution settings: A profile supplies model execution settings:
@@ -154,11 +179,21 @@ extra_params:
provider_option: enabled provider_option: enabled
``` ```
A derived profile can use a named base and override only the settings it owns:
```yaml
id: local-summary-fast
base_profile: local-summary
timeout_seconds: 30
reasoning_effort: low
```
| Field | Required | Meaning | | Field | Required | Meaning |
| --- | --- | --- | | --- | --- | --- |
| `id` | yes | Non-empty profile identifier. IDs must be unique within one source. | | `id` | yes | Profile identifier, trimmed before selection and publication. It must be non-empty after trimming and unique within one source after normalization. |
| `base_profile` | no | One optional parent profile ID. A derived profile may inherit target fields from it. |
| `backend` | unless `endpoint` is present | Backend registry ID. It is trimmed and registry membership is checked when the profile is prepared or inspected. | | `backend` | unless `endpoint` is present | Backend registry ID. It is trimmed and registry membership is checked when the profile is prepared or inspected. |
| `endpoint` | unless `backend` is present | Non-empty OpenAI-compatible base URL, including an API version path when required. When both connection fields are present, this overrides the backend endpoint without changing backend identity. | | `endpoint` | unless `backend` is present | OpenAI-compatible base URL, including an API version path when required. A nonempty value is trimmed and must be absolute HTTP or HTTPS with a host and without user information, a query, or a fragment. When both connection fields are present, this overrides the backend endpoint without changing backend identity. |
| `model` | yes | Non-empty provider model name. | | `model` | yes | Non-empty provider model name. |
| `temperature` | no | Number from 0 through 2. | | `temperature` | no | Number from 0 through 2. |
| `max_tokens` | no | Integer zero or greater. | | `max_tokens` | no | Integer zero or greater. |
@@ -166,16 +201,21 @@ extra_params:
| `timeout_seconds` | no | Per-generation deadline in whole seconds; integer zero or greater. | | `timeout_seconds` | no | Per-generation deadline in whole seconds; integer zero or greater. |
| `service_tier` | no | Provider-specific request tier. | | `service_tier` | no | Provider-specific request tier. |
| `reasoning_effort` | no | Provider-specific reasoning setting. | | `reasoning_effort` | no | Provider-specific reasoning setting. |
| `api_key_env` | no | Name of an environment variable containing the API key. | | `api_key_env` | no | Optional environment-variable lookup source for an API key. |
| `extra_params` | no | JSON-compatible provider-specific outbound fields. | | `extra_params` | no | JSON-compatible provider-specific outbound fields. |
Raw `api_key` is prohibited in profile YAML. Store only an environment Raw `api_key` is prohibited in profile YAML. Store only an environment
variable name in `api_key_env`. variable name in `api_key_env`.
A standalone profile must provide a model and at least one of `backend` or
`endpoint`. A derived profile may omit those target fields because its selected
base chain can provide them. Local parsing still validates a derived profile's
own ID, supplied endpoint, execution-setting bounds, and `extra_params`.
Promptkit does not infer a backend from a model or endpoint. Endpoint-only Promptkit does not infer a backend from a model or endpoint. Endpoint-only
profiles remain supported and have no effective backend ID. profiles remain supported and have no effective backend ID.
The engine always provides the built-in `openrouter` ID. Consumers can add The engine always provides the built-in `openrouter` and `rakestrawhome` IDs.
engine-scoped IDs with Consumers can add engine-scoped IDs with
[`WithBackend`](../backends.go); exact registration validation belongs to its [`WithBackend`](../backends.go); exact registration validation belongs to its
GoDoc. GoDoc.
@@ -183,30 +223,33 @@ GoDoc.
objects with string keys. Keys must be non-empty. With the built-in client, objects with string keys. Keys must be non-empty. With the built-in client,
they also cannot collide with the standard fields listed in the they also cannot collide with the standard fields listed in the
[outbound request contract](integrations/openai-compatible-chat.md#request-body). [outbound request contract](integrations/openai-compatible-chat.md#request-body).
Excessively deep or large JSON-shaped values are rejected for safety.
### Defaults And Overrides ### Defaults And Overrides
Execution settings resolve in this order: Execution settings resolve in this order:
1. framework defaults; 1. the framework timeout baseline;
2. the selected backend, when the profile names one; 2. the selected backend, when the profile names one;
3. the selected profile; and 3. the selected profile; and
4. request `ExecutionTargetOverride` values. 4. request `ExecutionTargetOverride` values.
The framework defaults are: The framework baseline is:
| Setting | Default | | Setting | Default |
| --- | --- | | --- | --- |
| `temperature` | `0` | | `temperature` | Unspecified and omitted from compatible provider requests unless a profile or runtime override selects it. |
| `max_tokens` | `0` | | `max_tokens` | Unspecified and omitted from compatible provider requests unless a profile or runtime override selects it. |
| `top_p` | `1` | | `top_p` | Unspecified and omitted from compatible provider requests unless a profile or runtime override selects it. |
| `timeout_seconds` | `600` | | `timeout_seconds` | `600` |
Numeric zero in a file or in-memory profile means that the profile does not Numeric zero in a file or in-memory profile does not select a numeric value.
replace the framework default. Numeric request overrides use pointers, so an For `temperature`, `max_tokens`, and `top_p`, it leaves the provider control
explicit zero is preserved. In particular, an explicit request unspecified. For `timeout_seconds`, it retains the framework deadline. Numeric
`timeout_seconds` of zero disables the per-generation deadline while leaving request overrides use pointers, so an explicit zero is retained and sent to
the caller context and transport timeout intact. compatible providers. In particular, an explicit request `timeout_seconds` of
zero disables the per-generation deadline while leaving the caller context and
transport timeout intact.
Non-empty profile strings replace backend defaults, and non-empty request Non-empty profile strings replace backend defaults, and non-empty request
strings replace both. Request reasoning is the exception: a nil strings replace both. Request reasoning is the exception: a nil
@@ -231,50 +274,79 @@ default.
Profile sources resolve matching IDs in this order: Profile sources resolve matching IDs in this order:
1. in-memory profiles supplied with `WithProfiles`; 1. in-memory profiles supplied with `WithProfiles`;
2. a profile file, `fs.FS`, or configured profile directory; and 2. the ordinary configured source selected by a profile file, `fs.FS`, or
3. embedded built-in profiles. configured profile directory;
3. application fallback profiles supplied with `WithFallbackProfileFS`; and
4. embedded built-in profiles.
A higher-precedence source falls back only when the profile is absent. An A profile source supplies a complete definition; definitions and their fields
invalid matching profile is an error and does not fall back. In-memory are not merged across sources. A higher-precedence source falls back only when
`Profile` values follow the same ranges as YAML profiles. They use the requested profile ID is absent. An invalid matching profile is an error and
`APIKeyRequired` for request-scoped credentials instead of `api_key_env`. does not fall back. In-memory `Profile` values follow the same ranges as YAML
Preparation and exact profile inspection use this same source precedence. profiles. They use `APIKeyRequired` for request-scoped credentials instead of
`api_key_env`. Preparation and exact profile inspection use this same source
precedence.
When a selected definition names `base_profile`, every profile ID in that
chain is looked up through this same precedence order. A higher-precedence
definition therefore shadows a lower-precedence definition of the same base
ID, including a built-in. References are not source-qualified.
### Profile Inheritance
Promptkit resolves one linear base chain of at most 32 profiles, including the
selected profile. It merges settings from the root base to the selected leaf.
The leaf's `id` remains the selected profile identity. Nonblank string fields
(`backend`, `endpoint`, `model`, `service_tier`, `reasoning_effort`, and
`api_key_env`) and nonzero numeric fields replace inherited values. A nonempty
`extra_params` map replaces the complete inherited map rather than merging
keys, and `APIKeyRequired: true` remains true through the chain. Backend and
endpoint are independent: replacing one does not clear the other.
There is no profile-level clearing syntax. Blank strings, zero numbers, false,
and empty maps remain unspecified and inherit from a base. Use existing
presence-aware request overrides where an execution needs an explicit zero or
empty reasoning setting.
An absent directly selected profile reports the ordinary not-found error. Once
the selected profile exists, a missing base, cycle, overlong chain, or
incomplete resolved target is a profile-load failure. Ordinary operations
resolve chains afresh; prepared execution retains the fully resolved target.
## Built-In Profile Catalog ## Built-In Profile Catalog
Every built-in selects the `openrouter` backend. The engine's built-in backend Every built-in profile selects one maintained built-in backend and inherits
registry supplies `https://openrouter.ai/api/v1` and the environment-variable that backend's connection and credential metadata. Profile files do not repeat
name `OPENROUTER_API_KEY`, so individual profiles contain only model and those values. A configured, application fallback, or in-memory profile with
generation settings. Built-in profile files do not repeat those connection the same profile ID takes precedence.
values. A custom or in-memory profile with the same profile ID takes
precedence.
| Provider | ID | Model | | Provider | ID | Backend | Model |
| --- | --- | --- | | --- | --- | --- | --- |
| aion-labs | `aion-2` | `aion-labs/aion-2.0` | | aion-labs | `aion-2` | `openrouter` | `aion-labs/aion-2.0` |
| anthropic | `claude-fable-latest` | `~anthropic/claude-fable-latest` | | anthropic | `claude-fable-latest` | `openrouter` | `~anthropic/claude-fable-latest` |
| anthropic | `claude-haiku-latest` | `~anthropic/claude-haiku-latest` | | anthropic | `claude-haiku-latest` | `openrouter` | `~anthropic/claude-haiku-latest` |
| anthropic | `claude-opus-latest` | `~anthropic/claude-opus-latest` | | anthropic | `claude-opus-latest` | `openrouter` | `~anthropic/claude-opus-latest` |
| anthropic | `claude-sonnet-latest` | `~anthropic/claude-sonnet-latest` | | anthropic | `claude-sonnet-latest` | `openrouter` | `~anthropic/claude-sonnet-latest` |
| deepseek | `deepseek-3-2` | `deepseek/deepseek-v3.2` | | deepseek | `deepseek-3-2` | `openrouter` | `deepseek/deepseek-v3.2` |
| deepseek | `deepseek-4-flash` | `deepseek/deepseek-v4-flash` | | deepseek | `deepseek-4-flash` | `openrouter` | `deepseek/deepseek-v4-flash` |
| deepseek | `deepseek-4-pro` | `deepseek/deepseek-v4-pro` | | deepseek | `deepseek-4-pro` | `openrouter` | `deepseek/deepseek-v4-pro` |
| google | `gemini-2-flash` | `google/gemini-2.5-flash` | | google | `gemini-2-flash` | `openrouter` | `google/gemini-2.5-flash` |
| google | `gemini-2-flash-lite` | `google/gemini-2.5-flash-lite` | | google | `gemini-2-flash-lite` | `openrouter` | `google/gemini-2.5-flash-lite` |
| google | `gemini-2-pro` | `google/gemini-2.5-pro` | | google | `gemini-2-pro` | `openrouter` | `google/gemini-2.5-pro` |
| google | `gemini-3-flash-lite` | `google/gemini-3.1-flash-lite` | | google | `gemini-3-flash-lite` | `openrouter` | `google/gemini-3.1-flash-lite` |
| google | `gemini-flash-latest` | `~google/gemini-flash-latest` | | google | `gemini-flash-latest` | `openrouter` | `~google/gemini-flash-latest` |
| google | `gemini-pro-latest` | `~google/gemini-pro-latest` | | google | `gemini-pro-latest` | `openrouter` | `~google/gemini-pro-latest` |
| google | `gemma-4-31b` | `google/gemma-4-31b-it:exacto` | | google | `gemma-4-31b` | `openrouter` | `google/gemma-4-31b-it:exacto` |
| minimax | `minimax-m2` | `minimax/minimax-m2.5` | | google | `rakestrawhome-gemma-4-31b` | `rakestrawhome` | `google/gemma-4-31b-it` |
| minimax | `minimax-m3` | `minimax/minimax-m3` | | minimax | `minimax-m2` | `openrouter` | `minimax/minimax-m2.5` |
| mistral | `mistral-large-2512` | `mistralai/mistral-large-2512` | | minimax | `minimax-m3` | `openrouter` | `minimax/minimax-m3` |
| mistral | `mistral-medium-3-5` | `mistralai/mistral-medium-3-5` | | mistral | `mistral-large-2512` | `openrouter` | `mistralai/mistral-large-2512` |
| mistral | `mistral-small-3` | `mistralai/mistral-small-3.2-24b-instruct` | | mistral | `mistral-medium-3-5` | `openrouter` | `mistralai/mistral-medium-3-5` |
| mistral | `mistral-small-4` | `mistralai/mistral-small-2603` | | mistral | `mistral-small-3` | `openrouter` | `mistralai/mistral-small-3.2-24b-instruct` |
| nvidia | `nemotron-3-ultra` | `nvidia/nemotron-3-ultra-550b-a55b` | | mistral | `mistral-small-4` | `openrouter` | `mistralai/mistral-small-2603` |
| openai | `gpt-5-mini` | `openai/gpt-5.4-mini` | | nvidia | `nemotron-3-ultra` | `openrouter` | `nvidia/nemotron-3-ultra-550b-a55b` |
| openai | `gpt-5-nano` | `openai/gpt-5.4-nano` | | openai | `gpt-5-mini` | `openrouter` | `openai/gpt-5.4-mini` |
| openai | `gpt-5-nano` | `openrouter` | `openai/gpt-5.4-nano` |
## Schemas ## Schemas
@@ -293,16 +365,25 @@ schema produces a failed validation result.
Credential values belong at the request or environment boundary, never in Credential values belong at the request or environment boundary, never in
prompt, profile, schema, or example files: prompt, profile, schema, or example files:
- a file profile names an environment variable with `api_key_env`; - a backend or file profile can name an optional environment lookup source
- an in-memory profile may set `APIKeyRequired`; with `APIKeyEnv` or `api_key_env`;
- a request can provide a direct `APIKey` or override `APIKeyEnv`; and - an in-memory profile may set `APIKeyRequired` as an explicit local
requirement;
- a request can provide a direct `APIKey` or override the optional `APIKeyEnv`
source; and
- a direct request key takes precedence over environment lookup. - a direct request key takes precedence over environment lookup.
After a direct request key, the credential-source precedence is request After a direct request key, the credential-source precedence is request
`APIKeyEnv`, profile `api_key_env`, then the backend default. An in-memory `APIKeyEnv`, profile `api_key_env`, then the backend default. An in-memory
profile with `APIKeyRequired` clears an inherited backend environment name and profile with `APIKeyRequired` clears an inherited backend environment name and
requires a direct key unless the request explicitly supplies `APIKeyEnv`. requires a direct key unless the request explicitly supplies `APIKeyEnv`.
Promptkit validates required credential availability during preparation. Named environment sources are optional: when the selected source is absent,
empty, or whitespace-only, the built-in client omits the `Authorization`
header and handles the provider response normally. `APIKeyRequired` is the
only explicit local availability requirement. Promptkit validates required
credential availability during preparation and rechecks it when a prepared
execution runs. Injected clients receive resolved source metadata but define
their own credential-resolution behavior.
Direct keys are excluded from JSON results and redacted by public string Direct keys are excluded from JSON results and redacted by public string
formatters. Environment-variable names may appear in prepared metadata, but formatters. Environment-variable names may appear in prepared metadata, but
their values do not. their values do not.

View File

@@ -14,10 +14,16 @@ that produce these outbound settings.
Generation sends an HTTP `POST` with `Content-Type: application/json`. Generation sends an HTTP `POST` with `Content-Type: application/json`.
Before the client is called, the engine resolves framework, backend, profile, Before the client is called, the engine resolves framework, backend, profile,
and request values into one execution target. A non-empty endpoint from that and request values into one execution target. Endpoint configuration is trimmed
target overrides the client's configured base URL. After trailing slashes are and must be an absolute HTTP or HTTPS URL with a host and without user
removed, `/chat/completions` is appended. Generation fails before sending when information, a query, or a fragment. A non-empty endpoint from the target
neither source supplies an endpoint. overrides the client's configured base URL. The final selected endpoint is
validated again before transport.
The completion URL is composed through parsed URL path operations. Nested base
paths are retained, repeated trailing slashes are normalized, and the result
has exactly one appended `/chat/completions` suffix. Generation fails before
sending when neither source supplies a valid endpoint.
The target's backend ID is routing metadata for prepared values, results, and The target's backend ID is routing metadata for prepared values, results, and
injected clients. The built-in client does not derive the URL from that ID and injected clients. The built-in client does not derive the URL from that ID and
@@ -25,11 +31,13 @@ does not serialize it in the provider request.
## Authentication ## Authentication
A non-empty API key supplied directly on the execution target takes A usable API key supplied directly on the execution target takes precedence.
precedence. Otherwise, when an API-key environment-variable name is supplied, Otherwise, when an API-key environment-variable name is supplied, the client
the client reads that variable and requires a non-empty value. The selected reads and trims that variable. A bearer header is sent only when the resolved
key is sent as `Authorization: Bearer <key>`. No authorization header is sent direct or environment credential is non-empty. When neither source is usable,
when neither mechanism is configured. the client omits `Authorization` and handles the provider response normally.
An explicitly required target with no usable source is rejected before
transport.
The target contains the already resolved environment-variable name: an The target contains the already resolved environment-variable name: an
explicit request override takes precedence over profile metadata, which takes explicit request override takes precedence over profile metadata, which takes
@@ -53,12 +61,14 @@ never also sent as a session header.
The client conditionally includes: The client conditionally includes:
- `temperature`, `max_tokens`, and `top_p` when non-zero or explicitly - `temperature`, `max_tokens`, and `top_p` only when selected by a profile or
present; runtime override, including an explicit runtime zero; they are absent when
unspecified;
- non-empty `service_tier` and effective `reasoning_effort`; an explicitly - non-empty `service_tier` and effective `reasoning_effort`; an explicitly
disabled reasoning setting is empty and therefore omitted; and disabled reasoning setting is empty and therefore omitted; and
- `response_format` for JSON Schema structured output, including its name, - `response_format` for JSON Schema structured output, including its name,
strict flag, and schema document. strict flag, and schema document. Plain JSON validation does not add an
object-only response constraint.
The engine resolves backend, profile, and request extra-parameter maps by The engine resolves backend, profile, and request extra-parameter maps by
whole-map replacement rather than key merging. The resulting effective map is whole-map replacement rather than key merging. The resulting effective map is
@@ -81,13 +91,47 @@ request fields.
## Response Handling ## Response Handling
Any 2xx response is decoded as an OpenAI-compatible chat response. The client Any 2xx response body is limited to 16 MiB (16,777,216 bytes). A larger
returns the first choice's non-empty message content and maps prompt, declared `Content-Length` is rejected before the body is read, and streamed,
completion, total, cached, and cache-write token counts. chunked, or underreported bodies are read through the same bound with at most
one additional byte used to detect overflow. A body exactly at the limit is
allowed. The body is closed on every outcome and an oversized stream is not
drained.
Invalid JSON, absent choices, and empty first-choice content are malformed The bounded body must contain exactly one OpenAI-compatible JSON response
responses. For a non-2xx status, the error includes the status code but never object followed only by JSON whitespace and EOF. The client returns the first
the provider response body. choice's explicitly present string message content, including an empty or
whitespace-only string, and maps prompt, completion, total, cached, and
cache-write token counts. Invalid or truncated JSON, trailing non-whitespace
data, a second JSON value, absent choices, missing content, `null` content,
non-string content, and size overflow are malformed responses and return no
partial result.
For a non-2xx status, Promptkit recognizes one JSON document with a top-level
object-valued `error` member. Its optional `message` and `type` fields must be
strings, and `code` may be a string or JSON number. Valid supported fields are
handled independently, numeric codes retain their JSON number text, and
unknown fields are ignored. Missing, invalid, malformed, or multiply framed
envelopes contribute no provider detail.
Non-success bodies have a 65,536-byte limit. A larger declared
`Content-Length` is not read; otherwise the client reads at most one additional
byte to detect streamed or underreported overflow. Empty, unreadable,
oversized, malformed, and unrecognized bodies retain only the received status.
The body is always closed and no oversized stream is drained beyond that probe.
Extracted strings are made valid UTF-8, trimmed, and converted to one line by
collapsing Unicode whitespace, control, and format-character runs. Blank
values are omitted. Codes and types longer than 256 Unicode code points are
omitted; messages longer than 4,096 code points are truncated at a code-point
boundary with an ellipsis inside the limit. Promptkit never exposes raw bodies,
headers, endpoints, credentials, request data, schemas, generated content, or
unsupported provider metadata through this handling.
An outbound `http.Client.Do` failure retains both Promptkit's request-failure
identity and the exact transport error for `errors.Is` and `errors.As` checks.
The rendered error does not include the selected endpoint, request headers,
request content, credentials, or provider body.
## Timeout And Cancellation ## Timeout And Cancellation
@@ -102,5 +146,7 @@ Timeouts are layered:
timeout when the supplied value is not positive. timeout when the supplied value is not positive.
The earliest applicable caller, generation, or transport deadline controls the The earliest applicable caller, generation, or transport deadline controls the
request. Constructing the internal client does not mutate a supplied request. Caller cancellation retains `context.Canceled`; caller, generation,
`http.Client`. and whole-request timeout failures retain `context.DeadlineExceeded`, together
with the request-failure identity. Constructing the internal client does not
mutate a supplied `http.Client`.

View File

@@ -56,7 +56,10 @@ reserve another bounded slot.
`NewClient` wraps the engine's selected internal model client after public `NewClient` wraps the engine's selected internal model client after public
client adaptation or built-in client construction. Initial generation and the client adaptation or built-in client construction. Initial generation and the
default repairer receive the same wrapper. default repairer receive the same wrapper. Their requests retain the same
effective backend, credential, numeric-presence metadata, and structured-output
settings, so scheduling does not change provider omission semantics between
calls.
For each `Generate` call, the wrapper selects a pool from the request's For each `Generate` call, the wrapper selects a pool from the request's
effective backend ID. An unlimited call passes directly to the next client. A effective backend ID. An unlimited call passes directly to the next client. A
@@ -71,7 +74,9 @@ other backend IDs.
The wrapper passes generation requests, responses, and collaborator errors The wrapper passes generation requests, responses, and collaborator errors
through unchanged. It owns scheduling only; the concrete model client remains through unchanged. It owns scheduling only; the concrete model client remains
responsible for provider transport behavior. responsible for provider transport behavior. The runner, rather than the
capacity layer, sums all five usage fields from the initial response and every
completed repair response into the successful run result.
## Cancellation And Release ## Cancellation And Release

View File

@@ -25,16 +25,21 @@ and request precedence. The client uses its endpoint, credential metadata,
generation fields, and extra parameters. `BackendID` remains routing metadata generation fields, and extra parameters. `BackendID` remains routing metadata
for the generation boundary and is not mapped into the provider payload. for the generation boundary and is not mapped into the provider payload.
Construction validates the configured base URL and clones any supplied Construction trims and validates a nonempty configured base URL and clones any
`http.Client` so Promptkit can apply its timeout default without mutating the supplied `http.Client` so Promptkit can apply its timeout default without
caller's client. Generation then: mutating the caller's client. An empty configured base remains valid because a
resolved request target may supply the endpoint. Generation then:
1. validates request-level timeout and endpoint requirements; 1. validates shared execution-setting invariants and the final selected base
endpoint;
2. maps the internal request into the OpenAI-compatible chat payload; 2. maps the internal request into the OpenAI-compatible chat payload;
3. validates and merges extra parameters; 3. validates and merges extra parameters;
4. resolves authentication; 4. composes `/chat/completions` through parsed URL path operations;
5. performs the outbound request under the applicable deadlines; and 5. resolves authentication;
6. decodes the first response choice and token usage. 6. performs the outbound request under the applicable deadlines; and
7. decodes one strictly framed, size-bounded successful response object and
maps its first choice and token usage, or decodes bounded structured
non-success detail.
`internal/llm` owns the set of reserved OpenAI-compatible request fields used `internal/llm` owns the set of reserved OpenAI-compatible request fields used
when validating extra parameters. Backend registration consumes the same rule when validating extra parameters. Backend registration consumes the same rule
@@ -50,12 +55,14 @@ the target, rendered messages, and structured-output constraint retained by
executable preparation. Execution does not reopen or rerender consumer executable preparation. Execution does not reopen or rerender consumer
sources. sources.
Before backend admission, the runner rechecks that the frozen credential Before backend admission, the runner rechecks a frozen credential
environment-variable name is available. The handle does not retain the environment-variable name only when the target explicitly requires a
environment value; the model client resolves the value visible when generation credential. The handle does not retain the environment value; the model client
begins. A direct request key remains in private execution state only until the resolves the value visible when generation begins. For optional sources with no
claimed execution finishes or an unclaimed handle is discarded. Exact public usable value, the built-in client omits `Authorization` and continues to the
ownership and redaction semantics belong to the provider. A direct request key remains in private execution state only until
the claimed execution finishes or an unclaimed handle is discarded. Exact
public ownership and redaction semantics belong to the
[`PreparedExecution` GoDoc](../../prepared_execution.go). [`PreparedExecution` GoDoc](../../prepared_execution.go).
## Failure Categories ## Failure Categories
@@ -63,11 +70,47 @@ ownership and redaction semantics belong to the
The package preserves distinct error identities for invalid client The package preserves distinct error identities for invalid client
configuration, invalid generation requests, request execution failures, configuration, invalid generation requests, request execution failures,
non-success provider statuses, and malformed successful responses. Provider non-success provider statuses, and malformed successful responses. Provider
response bodies are not included in non-success errors. response bodies are never exposed in raw form through non-success errors.
Caller cancellation and deadline failures during the outbound request are Invalid nonempty configured endpoints are configuration failures. A missing or
reported as request execution failures. The runner classifies these identities invalid final selected endpoint is an invalid generation request and is
without depending on HTTP status mapping. rejected before transport.
Authentication resolves a trimmed direct key before a trimmed configured
environment value. Optional missing, empty, or whitespace-only sources do not
block transport and produce no `Authorization` header. An explicitly required
target with no usable source is rejected before transport with the existing
invalid-request diagnostics.
Successful response bodies have a fixed 16 MiB limit enforced by declared
length and by reading at most one byte beyond the boundary. The decoder accepts
exactly one JSON object plus trailing whitespace and EOF. Size overflow,
truncation, malformed JSON, trailing data, and a second value are malformed
responses with no partial result or provider content in the error. Every body
is closed, and an unbounded oversized stream is not drained.
After framing succeeds, the first choice must contain an explicitly present
string `message.content`. The string is returned exactly, including empty or
whitespace-only content. Missing choices, missing or `null` content, and
non-string content are malformed responses. Output validation and correction
eligibility remain outside this package.
For a non-success response, `ProviderHTTPError` retains the HTTP status and
only normalized detail from the bounded recognized envelope. It retains
`ErrUnexpectedStatus` through unwrapping. The client owns response closure;
its bounded reader and parser never close or drain a body themselves. The root
facade converts this concrete internal error into the public
[`GenerationError`](../../generation_error.go), while arbitrary injected-client
errors continue through the ordinary generation-error mapping unchanged.
An `http.Client.Do` failure is represented by a redacting multi-cause error:
the package request-failure sentinel and the exact returned transport error are
both available through `errors.Is` and `errors.As`, while the rendered text
does not expose the endpoint, headers, request content, credential, transport
detail, or provider body. Caller cancellation retains `context.Canceled`;
caller deadlines, generation deadlines, and whole-request client timeouts
retain `context.DeadlineExceeded`. The runner adds its generation category
without discarding those identities or depending on HTTP status mapping.
## Test Ownership ## Test Ownership
@@ -75,7 +118,15 @@ The
[OpenAI-compatible client tests](../../internal/llm/openai_compatible_client_test.go) [OpenAI-compatible client tests](../../internal/llm/openai_compatible_client_test.go)
own configuration, client cloning, deterministic deadline precedence, own configuration, client cloning, deterministic deadline precedence,
authentication, request and response mapping, malformed data, error identity, authentication, request and response mapping, malformed data, error identity,
cancellation, and response-body suppression. The root transport contract test cancellation, endpoint selection and composition, pre-transport rejection, and
also verifies that resolved backend settings reach this client without bounded single-document successful-response framing, closure, and
serializing backend identity. All use local test servers or test transports; response-body suppression. The focused
the default suite makes no live or paid provider requests. [provider HTTP error tests](../../internal/llm/provider_http_error_test.go)
own envelope parsing, normalization, and bounded-reader cases; their
[transport tests](../../internal/llm/provider_http_error_transport_test.go)
own non-success response closure and integration. Root transport contract tests
own public `GenerationError` conversion, while also verifying that resolved
backend settings reach this client without serializing backend identity and
that ordinary-run cancellation retains its public generation and context
identities. All use local test servers or controlled test transports; the
default suite makes no live or paid provider requests.

View File

@@ -11,23 +11,23 @@ contributor workflow and validation.
| Component | Implemented responsibility | References | | Component | Implemented responsibility | References |
| --- | --- | --- | | --- | --- | --- |
| Root `promptkit` package | Provides the supported engine facade, source, backend-registration, and injection options, public request, result, prompt-inspection, and profile-inspection values, opaque prepared-execution handles, profile construction, extension interfaces, value conversion, redacted formatting, typed capacity errors, public error mapping, and engine-local assembly. | [Package GoDoc](../../doc.go), [prepared execution](../../prepared_execution.go), [backend API](../../backends.go), [engine assembly](../../engine.go) | | Root `promptkit` package | Provides the supported engine facade, source, backend-registration, and injection options, public request, result, prompt-inspection, and profile-inspection values, opaque prepared-execution handles, profile construction, extension interfaces, value conversion, redacted formatting, typed capacity and generation error mapping, engine-local profile-source assembly including application fallbacks, and bounded output-repair assembly. | [Package GoDoc](../../doc.go), [prepared execution](../../prepared_execution.go), [backend API](../../backends.go), [engine assembly](../../engine.go) |
| `examples/go-library/prepare` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, and `Prepare`. It is not a public library package. | [Example program](../../examples/go-library/prepare/main.go) | | `examples/go-library/prepare` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, and `Prepare`. It is not a public library package. | [Example program](../../examples/go-library/prepare/main.go) |
| `examples/go-library/run` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, an injected deterministic model client, and `Run`. It is not a public library package. | [Example program](../../examples/go-library/run/main.go) | | `examples/go-library/run` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, an injected deterministic model client, and `Run`. It is not a public library package. | [Example program](../../examples/go-library/run/main.go) |
| `internal/backend` | Constructs each engine's immutable registry from the built-in OpenRouter definition and consumer additions, validates and defensively copies definitions through the shared JSON-value package, and consumes the LLM-owned OpenAI-compatible reserved request-field rule. | [Backend registry](../../internal/backend/registry.go) | | `internal/backend` | Constructs each engine's immutable registry from the maintained built-in definitions and consumer additions, validates and defensively copies definitions through the shared JSON-value package, and consumes the LLM-owned OpenAI-compatible reserved request-field rule. | [Backend registry](../../internal/backend/registry.go) |
| `internal/capacity` | Owns engine-local bounded execution admission and FIFO model-generation permits for limited backend IDs, including cancellation-safe waiter removal and client wrapping. | [Internal capacity management](capacity.md) | | `internal/capacity` | Owns engine-local bounded execution admission and FIFO model-generation permits for limited backend IDs, including cancellation-safe waiter removal and client wrapping. | [Internal capacity management](capacity.md) |
| `internal/domain` | Defines internal framework values for requests, artifacts, prompt definitions, profiles, execution targets, rendering, generation, and validation. | [Domain declarations](../../internal/domain/domain.go) | | `internal/domain` | Defines internal framework values for requests, artifacts, prompt definitions, profiles, execution targets, rendering, generation, and validation, and owns source-neutral invariants for shared execution settings, OpenAI-compatible base endpoints, session identifiers, and output contracts. Source parsing, required fields, other source-specific normalization, defaulting, and boundary-specific error classification remain with their callers. | [Domain declarations](../../internal/domain/domain.go), [endpoint invariant](../../internal/domain/endpoint.go) |
| `internal/defaults` | Defines application-neutral framework constants and constructs the default execution target. It contains no CLI, server, or inbound HTTP limits. | [Framework defaults](../../internal/defaults/defaults.go) | | `internal/defaults` | Defines application-neutral framework constants and constructs the default execution target. It contains no CLI, server, or inbound HTTP limits. | [Framework defaults](../../internal/defaults/defaults.go) |
| `internal/filecatalog` | Provides deterministic YAML discovery and path helpers for operating-system filesystems and `fs.FS` sources. | [File catalog](../../internal/filecatalog/catalog.go) | | `internal/filecatalog` | Provides deterministic YAML discovery and path helpers for operating-system filesystems and `fs.FS` sources. | [File catalog](../../internal/filecatalog/catalog.go) |
| `internal/jsonvalue` | Validates and deeply copies JSON-compatible extra-parameter and prepared-schema trees while preserving supported concrete value types. | [JSON values](../../internal/jsonvalue/jsonvalue.go) | | `internal/jsonvalue` | Validates and deeply copies bounded JSON-compatible extra-parameter and prepared-schema trees while preserving supported concrete value types and rejecting cycles or excessive depth and work. | [JSON values](../../internal/jsonvalue/jsonvalue.go) |
| `internal/promptdef` | Loads strictly decoded, validated prompt definitions from filesystem and `fs.FS` sources, including version selection and contained file-backed message content. | [Framework formats](../formats.md), [prompt-definition repository](../../internal/promptdef/filesystem_repository.go) | | `internal/promptdef` | Loads strictly decoded, validated prompt definitions from filesystem and `fs.FS` sources, including version selection and contained file-backed message content. | [Framework formats](../formats.md), [prompt-definition repository](../../internal/promptdef/filesystem_repository.go) |
| `internal/profile` | Loads strictly decoded, validated execution profiles, including backend selection, from filesystem and `fs.FS` sources and composes repositories with error-preserving fallback. | [Framework formats](../formats.md), [profile repositories](../../internal/profile/filesystem_repository.go) | | `internal/profile` | Loads strictly decoded, locally validated execution profiles from filesystem and `fs.FS` sources, overlays raw sources with error-preserving fallback, and resolves inherited profiles. | [Framework formats](../formats.md), [profile repositories](../../internal/profile/filesystem_repository.go), [internal sources](sources.md#profiles-and-built-ins) |
| `internal/profile/builtin` | Embeds the built-in profile catalog, whose entries select OpenRouter, and combines it with an optional primary repository. | [Built-in catalog](../formats.md#built-in-profile-catalog), [repository](../../internal/profile/builtin/repository.go) | | `internal/profile/builtin` | Embeds the built-in profile catalog, whose entries select maintained built-in backends. | [Built-in catalog](../formats.md#built-in-profile-catalog), [repository](../../internal/profile/builtin/repository.go) |
| `internal/prompt` | Renders prompt messages from Go templates with artifact, variable, session, and cache-control data. | [Go-template renderer](../../internal/prompt/go_renderer.go) | | `internal/prompt` | Renders prompt messages from Go templates with artifact, variable, session, and cache-control data. | [Go-template renderer](../../internal/prompt/go_renderer.go) |
| `internal/artifact` | Resolves ordinary inline and unrestricted caller-selected file references into copied artifacts with metadata and hashes. | [Internal sources and validation](sources.md) | | `internal/artifact` | Resolves ordinary inline and unrestricted caller-selected file references into copied artifacts with metadata and hashes. | [Internal sources and validation](sources.md) |
| `internal/validate` | Validates basic, JSON, and JSON Schema output using operating-system filesystem or `fs.FS` schema sources and creates frozen validation plans for prepared execution. | [Framework formats](../formats.md#schemas), [internal sources and validation](sources.md) | | `internal/validate` | Validates basic, JSON, and JSON Schema output using operating-system filesystem or `fs.FS` schema sources and creates operation-local validation plans with canonical contained schema resources. | [Framework formats](../formats.md#schemas), [internal sources and validation](sources.md) |
| `internal/llm` | Defines the internal generation boundary and implements outbound OpenAI-compatible chat requests from resolved execution targets, including response decoding, authentication, deadline handling, and ownership of the OpenAI-compatible reserved request-field policy. | [Internal model client](llm.md) | | `internal/llm` | Defines the internal generation boundary and implements outbound OpenAI-compatible chat requests from resolved execution targets, including bounded structured non-success response decoding, successful-response decoding, authentication, deadline handling, and ownership of the OpenAI-compatible reserved request-field policy. | [Internal model client](llm.md) |
| `internal/usecase` | Resolves prompt definitions and hashes, profiles, backends, and targets for exact inspection and request settings for preparation, and coordinates ordinary execution and one-attempt prepared execution across internal sources, rendering, artifact loading, generation, validation, capacity, and optional repair. | [Internal runner](runner.md), [prepared-execution implementation](../../internal/usecase/prepared_execution.go) | | `internal/usecase` | Resolves prompt definitions and hashes, profiles, backends, and targets for exact inspection and request settings for preparation, and coordinates ordinary execution and one-attempt prepared execution across internal sources, rendering, artifact loading, operation-local validation plans, generation, capacity, and bounded repair. | [Internal runner](runner.md), [prepared-execution implementation](../../internal/usecase/prepared_execution.go) |
The root package assembles these internal components without exposing their The root package assembles these internal components without exposing their
representations. Consumers depend only on the root facade. representations. Consumers depend only on the root facade.

View File

@@ -20,10 +20,13 @@ and override semantics consumed by the runner.
profiles, backend resolution, artifacts, rendering, model generation, and profiles, backend resolution, artifacts, rendering, model generation, and
validation. The root engine supplies one immutable registry containing the validation. The root engine supplies one immutable registry containing the
built-in backend and validated consumer additions, one engine-local run built-in backend and validated consumer additions, one engine-local run
admitter, and a model client wrapped by the same capacity manager. Schema admitter, and a model client wrapped by the same capacity manager. Validation
documents are loaded through the validator's optional schema-loader interface. plans and provider-facing schema metadata come from the validator's preparation
An output repairer can be injected internally, but the ordinary runner interface.
constructor does not enable one. The root engine supplies one default output repairer through the explicit
runner constructor, using the same capacity-wrapped client as initial
generation. The no-repair runner constructor remains available for focused
internal callers and tests.
Each invocation carries its state in request, prepared-run, and result values. Each invocation carries its state in request, prepared-run, and result values.
The runner has no durable run or session store. The runner has no durable run or session store.
@@ -69,14 +72,16 @@ performs only the work needed to validate routing and admission:
5. resolve application-neutral defaults, backend defaults, profile values, 5. resolve application-neutral defaults, backend defaults, profile values,
and explicit request overrides in that order; and explicit request overrides in that order;
6. validate endpoint, model, numeric overrides, and credential requirements; 6. validate endpoint, model, numeric overrides, and credential requirements;
7. resolve the effective output contract without loading its schema; and 7. resolve and validate the effective output contract without loading its
schema; and
8. retain the definition, source identities, effective settings, output 8. retain the definition, source identities, effective settings, output
contract, and preparation start time in invocation-local state. contract, and preparation start time in invocation-local state.
The completion phase consumes that state without reloading the prompt, The completion phase consumes that state without reloading the prompt,
profile, or backend: profile, or backend:
1. load structured-output schema metadata when required; 1. create one operation-local validation plan and derive structured-output
schema metadata from it when required;
2. load and hash input artifacts; 2. load and hash input artifacts;
3. render messages and the prompt-defined session; 3. render messages and the prompt-defined session;
4. apply any direct session ID; 4. apply any direct session ID;
@@ -87,7 +92,10 @@ profile, or backend:
`Run` performs backend admission between the phases. This structure preserves `Run` performs backend admission between the phases. This structure preserves
one execution-precedence and error-ordering implementation while allowing a one execution-precedence and error-ordering implementation while allowing a
full backend pool to reject work before expensive schema, artifact, and full backend pool to reject work before expensive schema, artifact, and
rendering operations. rendering operations. `Prepare` discards the plan after returning its public
metadata. `Run` retains the plan through initial and repaired-output validation
and discards it when the operation ends. Prepared execution stores the same
kind of plan only in its private payload.
Pointer-based numeric overrides preserve an explicit zero. Invalid negative or Pointer-based numeric overrides preserve an explicit zero. Invalid negative or
out-of-range values fail as invalid requests. Endpoint overrides do not change out-of-range values fail as invalid requests. Endpoint overrides do not change
@@ -118,8 +126,17 @@ its `RunAdmitter` to reserve capacity for the effective backend ID. A nil
admitter is an internal unlimited fallback. After successful admission, `Run` admitter is an internal unlimited fallback. After successful admission, `Run`
immediately defers the returned release function, performs the completion immediately defers the returned release function, performs the completion
phase, makes one initial generation call, builds the named output artifact, phase, makes one initial generation call, builds the named output artifact,
and validates that artifact. Invalid generated content remains a validation and validates that artifact with the plan compiled during completion. Invalid
result; an inability to perform validation is an operational error. generated content remains a validation result; an inability to generate or
perform validation is an operational error.
Validation preparation and execution honor cancellation at every
Promptkit-controlled boundary and do not publish a partial plan or result.
Schema reads are bounded and context-checked between chunks; JSON decoding,
schema compilation, and schema execution are checked immediately before and
after their synchronous calls. Promptkit does not move arbitrary filesystem or
JSON Schema work to background goroutines, so an already-blocked dependency
method must return before cancellation can take precedence over its outcome.
The admission lease covers completion-phase preparation, initial generation, The admission lease covers completion-phase preparation, initial generation,
validation, every repair, and every exit. It bounds accepted work without validation, every repair, and every exit. It bounds accepted work without
@@ -127,22 +144,37 @@ serializing preparation or validation behind the active-generation limit.
The wrapped model client separately acquires a FIFO active permit only around The wrapped model client separately acquires a FIFO active permit only around
each actual generation call. each actual generation call.
When an internal repairer is present, a JSON or JSON Schema content failure can After a failed `basic`, JSON, or JSON Schema validation with a positive frozen
trigger bounded repair attempts. Repair receives the effective execution budget, the installed repairer can make a bounded corrective call. Each request
target and session ID, validation errors, prior output, and structured-output starts with a fresh copy of the complete original rendered messages, includes
specification. The default repairer uses the same wrapped client as initial only the latest nonempty candidate as an assistant message, and appends one
corrective user message. Empty candidates omit that assistant message. The
correction carries validation diagnostics as JSON data bounded to 64 KiB; the
full diagnostics remain in the validation result.
Repair receives the effective execution target, explicit numeric-presence bits,
credential, backend identity, session ID, and structured-output specification.
The same request constructor supplies those common fields to initial and repair
generation. The default repairer uses the same wrapped client as initial
generation, so each repair reacquires the selected backend's active permit generation, so each repair reacquires the selected backend's active permit
while remaining inside its original admission lease. Repair never performs a while remaining inside its original admission lease. Repair never performs a
second bounded admission. This capability remains internal and is not a public second bounded admission, and repaired outputs use the operation's existing
option. validation plan. The runner stops at the first valid candidate, sums completed
generation usage, reports calls actually started, and returns the final failed
validation result on exhaustion. A repair generation failure follows the
ordinary generation-error category rather than becoming a validation error.
A successful result includes the output artifact and raw output, validation A successful result includes the output artifact and raw output, validation
state, effective session ID, prompt and rendered-prompt hashes, selected state, effective session ID, prompt and rendered-prompt hashes, selected
profile and backend, effective settings, input hashes, token usage, a generated profile and backend, effective settings, input hashes, token usage, a generated
run identifier, and UTC timing. The same effective session reaches initial run identifier, and UTC timing. The same effective session reaches initial
generation and any repair attempt through the rendered prompt. The same generation and any repair attempt through the rendered prompt. The same
effective target, including backend identity, reaches generation and any effective target and presence metadata, including backend identity and direct
repair attempt. credential during execution, reaches generation and every repair attempt.
Result usage is the field-wise sum of all five usage values from the initial
response and every completed repair response. Final raw output, artifact, and
validation state still come from the last candidate. A repair error returns no
partial run result or partial usage.
## Failure Categories ## Failure Categories
@@ -165,7 +197,8 @@ category. Deferred release restores the admission lease on preparation,
generation, validation, repair, and cancellation failures. generation, validation, repair, and cancellation failures.
Other context cancellation propagates through the invoked collaborator and is Other context cancellation propagates through the invoked collaborator and is
classified by the owning operation. classified by the owning operation. In particular, cancellation observed by
validation retains the context identity through the validation error category.
An overlong direct session is an invalid request before source loading, while An overlong direct session is an invalid request before source loading, while
an invalid or overlong prompt session template remains a prompt-render failure. an invalid or overlong prompt session template remains a prompt-render failure.
An unknown selected backend, or a selected backend with no configured resolver, An unknown selected backend, or a selected backend with no configured resolver,
@@ -177,8 +210,9 @@ The [runner tests](../../internal/usecase/runner_test.go) own preparation order,
selection and override precedence, the two-phase boundary, early admission, selection and override precedence, the two-phase boundary, early admission,
lease lifetime and release, direct-session resolution, schema-before-generation lease lifetime and release, direct-session resolution, schema-before-generation
behavior, hashing, generation and validation outcomes, backend propagation, behavior, hashing, generation and validation outcomes, backend propagation,
bounded repair, shared initial/repair capacity, credentials and redaction, bounded repair progression, initial/repair request parity, cumulative usage,
error categories, artifact metadata, usage, and timing. The shared initial/repair capacity, credentials and redaction, error categories,
artifact metadata, and timing. The
[capacity subsystem document](capacity.md) identifies the focused pool, [capacity subsystem document](capacity.md) identifies the focused pool,
waiter, and wrapped-client tests. waiter, and wrapped-client tests.

View File

@@ -12,9 +12,21 @@ validation modes, built-in catalog, and source precedence.
## Prompt Definitions ## Prompt Definitions
`internal/promptdef` discovers YAML deterministically, decodes and validates `internal/promptdef` uses one source-neutral flow for prompt selection and
definitions, selects an ID and optional version, and resolves file-backed normalization. That flow scans normalized YAML ID and version metadata,
message content within the selected operating-system or `fs.FS` source. requires one strictly decoded document per file, classifies errors for the
selected definition, detects duplicates, and normalizes the exact match.
Small operating-system and `fs.FS` adapters own discovery, byte reads, display
paths, content opening, and root containment. Each lookup remains a
point-in-time scan: definitions and catalogs are not cached, and file-backed
message content is opened only for the exact selected candidate.
Operating-system sources enforce containment against canonical roots and
targets so symlinks cannot escape. Injected `fs.FS` sources enforce containment
in their clean relative path namespace. A single-file source uses the selected
prompt file's containing directory as its root. Every content path must be
relative and is opened from its exact parsed text after a separate blank check;
contained parent components and whitespace-bearing names remain valid.
Exact prompt inspection performs one point-in-time lookup through that same Exact prompt inspection performs one point-in-time lookup through that same
repository and validates referenced message content before returning declared repository and validates referenced message content before returning declared
@@ -27,35 +39,69 @@ duplicate detection, and source containment:
## Profiles And Built-Ins ## Profiles And Built-Ins
`internal/profile` loads and validates execution profiles from an `internal/profile` loads, locally validates, overlays, and resolves execution
operating-system filesystem or an `fs.FS`. It supports a primary repository profiles from an operating-system filesystem or an `fs.FS`. A file contains
with fallback only when the primary reports that a profile is absent. Strict exactly one YAML document and its trimmed YAML `id` is its only selection
YAML decoding recognizes the optional `backend` field, trims its value, and identity; filenames do not confer authority. Each point lookup reads discovered
requires a model plus at least one non-blank backend or endpoint. Loading does files once for their metadata and reuses the selected file's bytes for strict
not check registry membership because the available registry belongs to the decoding; unrelated profiles are not fully decoded. Strict selected decoding
recognizes `base_profile` and the optional `backend` field, trims their values,
and permits inherited target fields only when a base is named. File-backed
`extra_params` values are validated and defensively copied through the shared
bounded JSON-value owner before a profile is published. OpenAI-compatible
reserved-field policy remains with the model-client and backend-registry owners.
The overlay repository consults the next repository only when the
higher-precedence repository reports that a profile is absent. A reliably
selected malformed profile stops fallback, while an unrelated malformed file
does not become authoritative through its filename. Loading does not check
backend registry membership because the available registry belongs to the
assembled engine; the runner checks membership during preparation and exact assembled engine; the runner checks membership during preparation and exact
profile inspection. profile inspection.
Exact profile inspection performs one point-in-time lookup through those The root engine assembles one raw composite catalog in precedence order:
profile sources and checks the resolved target without reading prompt, input, in-memory profiles, one ordinary configured source, an application fallback
or schema sources. It does not retain that lookup for a later execution. source, then the embedded built-in catalog. An explicit file or `fs.FS` profile
source replaces `Config.ProfileDir` within the ordinary configured-source
category. One outer resolving repository wraps that complete raw catalog, so
each base lookup observes the same precedence and shadowing rules.
`internal/profile/builtin` embeds the maintained built-in profile catalog and The resolving repository traverses every selected chain afresh, retains no
can place a caller-selected repository ahead of that catalog. Every embedded cache, detects cycles, limits a chain to 32 profiles, merges root-to-leaf into a
profile selects `openrouter` and inherits its endpoint and credential new caller-owned value, and validates the final target before publishing it. It
environment-variable name from the built-in backend registry rather than does not check backend registry membership. Exact `base_profile` syntax, merge
repeating those values. Profile behavior is owned by the rules, and consumer-visible failure behavior belong to the [framework format
[profile repository tests](../../internal/profile/repository_test.go), while reference](../formats.md#profile-inheritance).
catalog completeness, the backend-selection invariant, duplicate IDs, and
Exact profile inspection performs one point-in-time resolved lookup through
those profile sources and checks the final target without reading prompt, input,
or schema sources. It does not retain that lookup for a later execution.
Prepared execution instead freezes the fully resolved target; a later ordinary
operation performs a fresh traversal.
`internal/profile/builtin` embeds the maintained built-in profile catalog.
Every embedded profile selects a maintained built-in backend and inherits that
backend's endpoint and credential environment-variable name from the built-in
backend registry rather than repeating those values. Profile loading and
overlay behavior are owned by the overlay behavior are owned by the
[profile repository tests](../../internal/profile/repository_test.go), while
catalog completeness, the backend-selection invariant, and duplicate IDs are
owned by the
[built-in repository tests](../../internal/profile/builtin/repository_test.go). [built-in repository tests](../../internal/profile/builtin/repository_test.go).
## Ordinary Artifacts ## Ordinary Artifacts
`internal/artifact` resolves inline references and unrestricted, `internal/artifact` accepts explicitly typed inline references even when their
caller-selected file paths. It copies content into an artifact, records body is empty. It also resolves unrestricted, caller-selected paths only when
metadata and a content hash, applies a content-type fallback, and honors they identify regular operating-system files, checking that condition before
context cancellation. and after opening the file. It copies content into an artifact, records
metadata and an opaque content-equality value, and applies a content-type
fallback.
Regular files are read synchronously in bounded chunks. Cancellation is
checked before opening, before and after every read, and before publishing the
artifact, so a canceled read never publishes partial content. The ordinary
reader does not detach file reads into background goroutines.
This ordinary reader does not implement an inbound HTTP security boundary. In This ordinary reader does not implement an inbound HTTP security boundary. In
particular, it does not constrain files to an application root or impose an particular, it does not constrain files to an application root or impose an
@@ -67,8 +113,17 @@ implemented reader behavior and failures.
## Rendering ## Rendering
`internal/prompt` renders definition messages as Go templates using named `internal/prompt` renders definition messages as Go templates using named
artifacts and variables. It carries message roles, session IDs, and cache artifacts and variables. Within one render, each referenced artifact body is
control into the rendered prompt. The converted to text lazily and cached by input name for reuse across the session
and every message; the cache is not shared across renders. Conversion uses
bounded chunks and preserves the artifact bytes exactly.
Session and message parsing and execution remain synchronous. The renderer
checks cancellation before and after each parse and execution boundary,
between artifact conversion chunks, around each message, and before publishing
the complete prompt. It cannot interrupt template work already in progress and
never publishes a partial prompt after observing cancellation. It carries
message roles, session IDs, and cache control into the rendered prompt. The
[renderer tests](../../internal/prompt/renderer_test.go) own rendering behavior. [renderer tests](../../internal/prompt/renderer_test.go) own rendering behavior.
## Schemas And Output Validation ## Schemas And Output Validation
@@ -78,22 +133,35 @@ filesystem or an `fs.FS`. Invalid generated content is returned as a validation
result; inability to load, register, or compile a schema is an operational result; inability to load, register, or compile a schema is an operational
error. error.
For executable preparation, the built-in validators create a frozen validation Every preparation operation creates one operation-local validation plan. None,
plan. None, basic, and JSON modes retain the effective output contract without basic, and JSON modes retain the effective output contract without source
source access. JSON Schema mode loads the root document, resolves and compiles access. JSON Schema mode loads the root document once, resolves and compiles
every transitive reference during preparation, and retains the compiled each transitive reference, and retains the compiled validator. Schema compiler
validator. The provider-facing structured-output metadata uses that same resources use canonical escaped file or private-scheme URLs; loaders decode
captured root document. their paths once and enforce the configured source boundary. The
provider-facing structured-output metadata uses the root document captured by
the same plan.
`PrepareExecution` also completes prompt and profile selection, artifact Schema preparation and execution remain synchronous. Promptkit checks
loading and hashing, session and message rendering, and target resolution. cancellation before and after source resolution, JSON decoding, compilation,
`RunPrepared` uses the retained source-derived state and validation plan; it and validation, and between bounded schema-read chunks. Once cancellation is
does not reopen prompt, profile, input, or schema sources and does not rerender observed it returns the context error without publishing a partial plan or
the request. By contrast, ordinary `Prepare` produces a preparation value only: validation result, even when a compiler or validator has just returned a
a later `Run` performs its own source resolution and preparation. different error or a successful result. An `fs.FS` method or JSON Schema
dependency call already in progress cannot be preempted; Promptkit waits for
that call to return and then gives cancellation precedence. Validation does
not detach dependency work into background goroutines.
`Prepare` discards its validation plan after returning metadata. `Run` retains
its plan for initial and repaired-output validation, then discards it with the
operation. `PrepareExecution` retains the plan in its private frozen payload;
`RunPrepared` uses that plan without reopening prompt, profile, input, or
schema sources or rerendering the request. A later ordinary `Run` always
performs fresh source resolution and preparation.
The [validator tests](../../internal/validate/standard_validator_test.go) own The [validator tests](../../internal/validate/standard_validator_test.go) own
basic, JSON, JSON Schema, source resolution, schema loading, compilation, basic, JSON, JSON Schema, source resolution, schema loading, compilation,
frozen-reference behavior, and content-failure behavior. Prepared execution frozen-reference behavior, content-failure behavior, and the synchronous
cancellation boundary. Prepared execution
orchestration is owned by the orchestration is owned by the
[use-case tests](../../internal/usecase/prepared_execution_test.go). [use-case tests](../../internal/usecase/prepared_execution_test.go).

View File

@@ -19,10 +19,10 @@ results, public values, extension interfaces, profiles, and error sentinels.
The implemented internal components consist of: The implemented internal components consist of:
- `internal/domain`, which owns framework data values shared by later internal - `internal/domain`, which owns framework data values and source-neutral
components; invariants shared by later internal components;
- `internal/backend`, which owns validated immutable OpenAI-compatible backend - `internal/backend`, which owns validated immutable OpenAI-compatible backend
definitions and the built-in OpenRouter definition; definitions and the maintained built-in definitions;
- `internal/capacity`, which owns engine-local bounded run admission and - `internal/capacity`, which owns engine-local bounded run admission and
model-generation scheduling for limited backends; model-generation scheduling for limited backends;
- `internal/defaults`, which owns application-neutral framework defaults and - `internal/defaults`, which owns application-neutral framework defaults and
@@ -92,6 +92,13 @@ coordinates internal components and adapts the supported public extension
interfaces to narrow internal abstractions. Internal components must not depend interfaces to narrow internal abstractions. Internal components must not depend
on consumers or on Scriptorium. on consumers or on Scriptorium.
`internal/domain` owns source-neutral invariants for values shared across
multiple input and execution boundaries, including execution-setting bounds,
OpenAI-compatible base endpoints, session identifiers, and output-contract
legality. Callers retain source parsing, required-field rules, other
source-specific normalization, defaulting, error classification, and policy
specific to their own boundary.
## Repository And Consumer Boundary ## Repository And Consumer Boundary
Scriptorium is a downstream application that consumes Promptkit through Scriptorium is a downstream application that consumes Promptkit through

View File

@@ -50,19 +50,18 @@ Examples of appropriate seams include clocks, randomness, subprocesses, remote A
## Test execution requirements ## Test execution requirements
Promptkit currently uses maintainer-run validation rather than hosted CI. Promptkit currently uses maintainer-run validation rather than hosted CI.
Maintainers run the repository-documented test, vet, build, formatting, Maintainers run the complete local workflow in the
documentation-link, and repository-hygiene checks before accepting changes. [development guide](../development.md#maintainer-validation) before accepting
changes. That guide is the canonical owner of exact commands, formatting,
documentation-link validation, and repository-hygiene checks.
Introducing hosted CI later would supplement, not silently redefine, this Introducing hosted CI later would supplement, not silently redefine, this
documented validation model. documented validation model.
The complete test sequence includes ordinary and race-enabled package tests. Maintainer validation must include ordinary and race-enabled package tests,
The maintained offline consumer workflow is also run from the repository root: static analysis, a complete build, and execution of both maintained offline
consumer examples. The preparation example protects assembled preparation and
```sh inspection behavior. The execution example separately protects assembled
go test ./... `Run`, injected-client, validation, usage, and result behavior.
go test -race ./...
go run ./examples/go-library/prepare
```
Tests in the default suite must be deterministic, offline, and independent of Tests in the default suite must be deterministic, offline, and independent of
real credentials. They must not invoke paid APIs, use live network real credentials. They must not invoke paid APIs, use live network
@@ -88,8 +87,9 @@ Use each test type where it protects a distinct risk:
interaction, while replacing live or nondeterministic external boundaries. interaction, while replacing live or nondeterministic external boundaries.
- External-package root tests exercise the public facade as a Go consumer, - External-package root tests exercise the public facade as a Go consumer,
while internal package tests own focused implementation behavior. while internal package tests own focused implementation behavior.
- The maintained offline preparation example protects one representative - The maintained offline preparation and execution examples protect distinct
assembled consumer workflow without contacting a model provider. representative assembled consumer workflows without contacting a model
provider.
- Fixtures should be minimal, synthetic, versioned with the behavior they - Fixtures should be minimal, synthetic, versioned with the behavior they
exercise, and free of credentials or private data. exercise, and free of credentials or private data.
- Golden files are appropriate only when the complete output is intentionally - Golden files are appropriate only when the complete output is intentionally

View File

@@ -106,48 +106,12 @@ gitea.maximumdirect.net/eric/promptkit 1.25.5
promptkit gitea.maximumdirect.net/eric/promptkit promptkit gitea.maximumdirect.net/eric/promptkit
``` ```
Run the complete maintainer validation required by the As a release prerequisite, run the complete
[development guide](development.md): [maintainer validation workflow](development.md#maintainer-validation) against
the clean candidate. Do not substitute a partial command list: the development
```sh guide owns the tests, race checks, analysis, build, both offline examples,
go test ./... formatting, Markdown links, generated-output and credential review, and
go test -race ./... repository hygiene. Record the successful workflow result with the candidate.
go vet ./...
go build ./...
go run ./examples/go-library/prepare
```
Check every tracked Go file. This command must produce no output:
```sh
unformatted=$(
git ls-files '*.go' |
while IFS= read -r go_file
do
gofmt -l "$go_file"
done
)
test -z "$unformatted"
```
Follow every maintained Markdown link and confirm that its local or published
target exists. Review the repository for generated binaries, test or coverage
output, credentials, template residue, downloaded assets, and other files that
do not belong in source control.
Recheck module and repository hygiene, whitespace, and the clean checkout:
```sh
test -z "$(git ls-files go.work go.work.sum)"
test ! -e vendor
if grep -Eq '^[[:space:]]*replace([[:space:]]|\()' go.mod
then
printf '%s\n' 'go.mod contains a replacement' >&2
exit 1
fi
git diff --check
test -z "$(git status --porcelain)"
```
## Write The Release Note ## Write The Release Note

View File

@@ -126,9 +126,9 @@ later request must provide a direct credential. Environment-variable names may
be reported, but inspection does not read credential values or require the be reported, but inspection does not read credential values or require the
named variable to be populated. named variable to be populated.
Inspection applies the ordinary configured and built-in profile precedence and Inspection applies the engine's profile source precedence and resolves any
resolves any selected backend. It does not load a prompt, render content, selected backend. It does not load a prompt, render content, reserve capacity,
reserve capacity, or contact a model. or contact a model.
See the See the
[profile-inspection consumer guide](../consumers/pkg-promptkit.md#inspect-a-profile-before-prompt-work) [profile-inspection consumer guide](../consumers/pkg-promptkit.md#inspect-a-profile-before-prompt-work)

125
docs/releases/v0.5.0.md Normal file
View File

@@ -0,0 +1,125 @@
# Promptkit v0.5.0
This supplemental changelog and migration guide summarizes the consumer-facing
changes from `v0.4.0` to `v0.5.0`. The annotated `v0.5.0` tag is the
authoritative release record. Exact current contracts belong to the linked
GoDoc and durable documentation.
## Summary
`v0.5.0` makes provider requests less prescriptive and adds an application
fallback layer for profile definitions:
- unset optional provider controls are omitted from OpenAI-compatible request
bodies instead of being populated with framework values; and
- `WithFallbackProfileFS` lets an application package profile defaults that
operators can override through the existing ordinary profile sources.
These changes let compatible providers apply their own model defaults while
giving applications stable embedded profile IDs without weakening operator
configuration precedence.
## Compatibility
The release adds one public function and removes no public declaration.
Existing source code should continue to compile.
There is one intentional behavior change: when no profile or runtime override
selects `top_p`, Promptkit no longer sends the former framework value of `1`.
It omits `top_p` and lets the provider choose its behavior. Unset
`temperature` and `max_tokens` are likewise omitted. Explicit nonzero profile
values and runtime values—including explicit runtime zero values—retain their
precedence and wire effect.
Consumers that relied on Promptkit always sending `top_p: 1` should add that
value to the relevant profile or runtime override before upgrading. Consumers
that did not rely on the implicit sampling value require no migration.
Application fallback profiles are opt-in. Engines that do not call
`WithFallbackProfileFS` retain the previous profile-source behavior.
## Upgrade
Update the module dependency with:
```sh
go get gitea.maximumdirect.net/eric/promptkit@v0.5.0
go mod tidy
```
Run the consuming project's ordinary and race-enabled tests after upgrading.
If request payloads or model behavior are asserted in fixtures, review them for
the optional-parameter omission described below.
## Omitted Optional Provider Controls
The built-in OpenAI-compatible client now includes `temperature`,
`max_tokens`, and `top_p` only when a profile or runtime override selects the
value. An explicit runtime zero remains present because runtime override
pointers distinguish zero from an unspecified value.
Promptkit's positive generation deadline remains a framework concern and is
not a provider request-body default. Required request fields, session IDs,
structured output, reasoning selection, credentials, and explicit extra
parameters retain their existing behavior.
See the [framework default and precedence reference](../formats.md#defaults-and-overrides),
the [`ExecutionTargetOverride` GoDoc](../../types.go), and the
[OpenAI-compatible request-body contract](../integrations/openai-compatible-chat.md#request-body)
for current details.
## Embedded Application Fallback Profiles
Applications can package ordinary profile YAML in an `fs.FS` and register it
as a fallback source:
```go
//go:embed profiles/*.yaml
var applicationProfiles embed.FS
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: "prompts",
ProfileDir: operatorProfileDir,
},
promptkit.WithFallbackProfileFS(applicationProfiles, "profiles"),
)
```
Leave `operatorProfileDir` empty when no operator source is configured. A
configured ordinary source is authoritative: a matching definition overrides
the application fallback, while a read or validation failure remains an error
instead of silently reaching a lower layer.
Profile definitions resolve in this order:
1. in-memory profiles supplied with `WithProfiles`;
2. the ordinary configured source selected by `WithProfileFile`,
`WithProfileFS`, or `Config.ProfileDir`;
3. the application source supplied with `WithFallbackProfileFS`; and
4. Promptkit's embedded built-in profiles.
Only an absent profile ID falls through. Sources provide complete profiles and
do not merge fields. Loading remains lazy, and the new source uses the existing
strict profile YAML and credential rules.
See the
[embedded-default consumer guidance](../consumers/pkg-promptkit.md#supply-embedded-application-defaults),
the [`WithFallbackProfileFS` GoDoc](../../engine.go), and the
[profile source reference](../formats.md#source-and-profile-precedence) for
current details.
## Public API Changes
The release adds:
- `WithFallbackProfileFS`.
No public declaration was removed or changed.
## Consumer Action
- Review any workflow that depended on Promptkit's implicit `top_p: 1` and
configure the value explicitly when required.
- Optionally adopt `WithFallbackProfileFS` when an application should package
overridable profile defaults.
- Run consumer tests after updating the module dependency.

131
docs/releases/v0.6.0.md Normal file
View File

@@ -0,0 +1,131 @@
# Promptkit v0.6.0
This supplemental changelog and migration guide summarizes the consumer-facing
changes from `v0.5.0` to `v0.6.0`. The annotated `v0.6.0` tag is the
authoritative release record. Exact current contracts belong to the linked
GoDoc and durable documentation.
## Summary
`v0.6.0` is a broad correctness, safety, efficiency, and maintainability
release. It does not add or remove public declarations. The release:
- centralizes shared execution-setting, output-contract, endpoint, and
JSON-compatible-value rules;
- unifies prompt repository behavior and avoids unnecessary prompt and profile
decoding;
- bounds consumer-controlled JSON trees and successful provider responses;
- hardens prompt content paths, artifact files, provider URLs, JSON framing,
and error propagation;
- reuses compiled schema plans and rendered artifact text within an operation;
and
- improves cancellation behavior, prepared-value ownership, deterministic
transport testing, and maintainer validation.
## Compatibility
No public declaration was added, removed, or changed. Ordinary valid `v0.5.0`
configurations and requests should continue to compile and behave as before.
The release intentionally rejects or reports several inputs that were
previously accepted, altered, or misclassified:
- execution settings must be finite, within their documented ranges, and safe
to convert to Go durations;
- output formats, validation modes, repair counts, and JSON Schema dependencies
are validated consistently;
- file-backed prompt and profile identity comes from normalized YAML metadata,
not filenames;
- prompt `content_file` values must be exact relative paths contained by their
configured source root;
- built-in file artifacts must resolve to regular files;
- selected provider endpoints must be absolute HTTP or HTTPS URLs without user
information, query strings, or fragments;
- JSON documents and successful provider responses must contain exactly one
value, and successful provider bodies are limited to 16 MiB; and
- excessively deep or expansive JSON-compatible values fail with ordinary
validation errors.
These are compatibility corrections and safety boundaries rather than new
consumer configuration requirements. Consumers relying on an invalid or
ambiguous input should correct that input before upgrading.
## Upgrade
Update the module dependency with:
```sh
go get gitea.maximumdirect.net/eric/promptkit@v0.6.0
go mod tidy
```
Run the consuming project's ordinary and race-enabled tests after upgrading.
Applications with custom prompt/profile sources, local provider endpoints,
unusual artifact paths, or assertions over provider error identities should
pay particular attention to the compatibility notes below.
## Source Loading And Identity
Prompt definitions now share one source-neutral selection and normalization
flow across operating-system and `fs.FS` sources. YAML `id` and `version`
metadata are authoritative; filenames do not create a second identity system.
Only selected content bodies are loaded, malformed unrelated definitions do
not shadow valid exact matches, and per-file read failures are reported as
prompt-load failures rather than false absence.
File-backed profiles likewise use normalized YAML IDs, reuse their metadata
read for selected strict decoding, and avoid fully decoding unrelated files.
Selected malformed definitions remain authoritative and do not silently fall
through to a lower-precedence source.
Prompt `content_file` paths are opened exactly as declared after a separate
blank check. They must remain relative to and contained by the configured
prompt source root, including across operating-system symlinks.
See the [framework source and identity reference](../formats.md) and
[internal source overview](../internal/sources.md) for the current contracts.
## Validation, Cancellation, And Efficiency
JSON Schema documents preserve exact JSON-number representations. Schema
resource URLs safely escape legal filesystem names, and each operation loads
and compiles its schema graph once. `Run` and prepared execution reuse that
operation-local plan; Promptkit does not introduce a cross-operation cache.
Artifact reading, rendering, schema loading, compilation, and validation now
check cancellation at the synchronous boundaries Promptkit controls. Rendering
memoizes each artifact's text within one render operation, while plain JSON
validation avoids materializing an unnecessary generic tree.
The shared JSON-compatible-value owner now limits nesting and produced work so
unsafe consumer-controlled structures return errors instead of risking
unbounded recursion or allocation. See the
[architecture policy](../policy/architecture.md) for invariant ownership and
the [format reference](../formats.md) for validation behavior.
## Provider Transport Hardening
OpenAI-compatible endpoints are parsed and composed structurally, including
nested base paths. Underlying transport cancellation and deadline errors remain
discoverable with `errors.Is` through Promptkit's generation error category.
Successful provider bodies are read with a fixed 16 MiB bound and must contain
exactly one JSON response object followed only by whitespace. Oversized,
truncated, malformed, or multiply framed responses fail without returning a
partial result. See the
[OpenAI-compatible integration contract](../integrations/openai-compatible-chat.md)
for the canonical request, endpoint, error, and response behavior.
## Public API Changes
None.
## Consumer Action
- Correct any configuration or request that depends on the formerly permissive
cases described under Compatibility.
- Confirm custom local provider endpoints are absolute HTTP or HTTPS base URLs
without credentials, queries, or fragments.
- Confirm prompt content paths remain within their configured source root and
file artifacts resolve to regular files.
- Run ordinary and race-enabled consumer tests after updating the dependency.

155
docs/releases/v0.7.0.md Normal file
View File

@@ -0,0 +1,155 @@
# Promptkit v0.7.0
This supplemental changelog and migration guide summarizes the consumer-facing
changes from `v0.6.0` to `v0.7.0`. The annotated `v0.7.0` tag is the
authoritative release record. Exact current contracts belong to the linked
GoDoc and durable documentation.
## Summary
`v0.7.0` expands provider integration and profile composition while making
credential and generation-failure handling more flexible:
- Promptkit now includes the `rakestrawhome` backend and its Gemma profile;
- built-in generation failures expose bounded structured provider details;
- an unavailable optional API-key environment source no longer prevents a
request from reaching an upstream that permits unauthenticated access; and
- profiles can inherit from and selectively refine another profile.
## Compatibility
This release adds public declarations and fields but removes none. Existing
keyed configuration literals and ordinary `errors.Is` handling continue to
work.
Adding `BaseProfileID` to `Profile` and `OpenAICompatibleProfileConfig` changes
their struct shape. Consumers using positional composite literals for either
type must convert them to keyed literals. Existing keyed literals require no
change.
The `rakestrawhome` backend ID is now built in and reserved. A consumer that
previously registered that exact ID with `WithBackend` must remove its manual
registration before upgrading. Other custom backend registrations are
unchanged.
When an optional backend, profile, or request `APIKeyEnv` is unset, empty, or
whitespace-only, the built-in client now omits `Authorization` and sends the
request. Previously this condition could fail before transport. Set
`Profile.APIKeyRequired` when missing credentials must remain a local
preflight error.
Provider non-success responses continue to match `ErrLLMGenerate`. Their
rendered wording is not a compatibility contract; consumers can now use
`errors.As` with `*GenerationError` when structured status information is
needed.
## Upgrade
Update the module dependency with:
```sh
go get gitea.maximumdirect.net/eric/promptkit@v0.7.0
go mod tidy
```
Remove any manual `rakestrawhome` backend registration, convert positional
profile literals to keyed literals, and run the consuming project's ordinary
and race-enabled tests.
## Rakestrawhome Built-In Backend And Profile
Every engine now includes the reserved `rakestrawhome` backend, identified by
`BackendRakestrawHome`. The built-in `rakestrawhome-gemma-4-31b` profile
selects that backend. Consumers can use the maintained endpoint, credential,
capacity, and model defaults without registering either definition themselves.
See the [built-in backend and profile catalogs](../formats.md#built-in-backends)
and the [consumer adoption example](../consumers/pkg-promptkit.md#use-the-rakestrawhome-built-in-profile)
for the current contracts.
## Structured Generation Errors
Non-2xx responses from the built-in OpenAI-compatible client now return an
immutable `*GenerationError`. Consumers can inspect the HTTP status and any
safely extracted provider code, type, or message while retaining the ordinary
generation-error category:
```go
var generationErr *promptkit.GenerationError
if errors.As(err, &generationErr) {
status := generationErr.StatusCode()
_ = status
}
```
Provider fields are bounded and normalized but remain untrusted and may
contain sensitive request or schema details. Default and Go-syntax formatting
omit those fields. Applications must apply their own disclosure policy before
logging or presenting accessor values.
See the [`GenerationError` GoDoc](../../generation_error.go), the
[consumer error-handling guide](../consumers/pkg-promptkit.md#handle-errors),
and the [OpenAI-compatible response contract](../integrations/openai-compatible-chat.md#response-handling).
## Optional Credential Sources
`APIKeyEnv` names an optional environment lookup source unless the selected
profile explicitly sets `APIKeyRequired`. When neither a direct request key nor
a usable environment value exists, the built-in client omits the bearer header
and handles the upstream response normally. This supports local and other
OpenAI-compatible providers that permit unauthenticated requests without
hiding an authentication error returned by a provider that requires one.
The [credential format reference](../formats.md#credentials), the
[`Backend` GoDoc](../../backends.go), the
[`ExecutionTargetOverride` GoDoc](../../types.go), and the
[authentication integration contract](../integrations/openai-compatible-chat.md#authentication)
define the current precedence and availability rules.
## Profile Inheritance
YAML profiles can name one parent with `base_profile`; in-memory profiles use
`Profile.BaseProfileID`, and `OpenAICompatibleProfileConfig` forwards the same
field. A profile can act as an application-owned alias of a built-in or refine
selected inherited settings:
```go
promptkit.WithProfiles(promptkit.Profile{
ID: "weather-light",
BaseProfileID: "deepseek-4-flash",
ReasoningEffort: "high",
})
```
Base lookup observes the existing source precedence. Chains are linear,
cycle-safe, and resolved afresh for ordinary operations. Prepared execution
freezes the fully resolved target. The selected leaf ID remains public while
effective execution settings reflect the resolved chain.
See the [profile inheritance format reference](../formats.md#profile-inheritance),
the [consumer alias example](../consumers/pkg-promptkit.md#alias-a-built-in-profile),
and the [`Profile` GoDoc](../../types.go) for exact merge and validation
behavior.
## Public API Changes
The release adds:
- `BackendRakestrawHome`;
- `GenerationError`, including `StatusCode`, `ProviderCode`, `ProviderType`,
`ProviderMessage`, `Error`, `GoString`, and `Unwrap`;
- `Profile.BaseProfileID`; and
- `OpenAICompatibleProfileConfig.BaseProfileID`.
No public declaration was removed.
## Consumer Action
- Remove a manual backend registration whose ID is exactly `rakestrawhome`.
- Convert positional `Profile` or `OpenAICompatibleProfileConfig` literals to
keyed literals.
- Set `Profile.APIKeyRequired` where a missing credential must fail locally
instead of reaching the provider unauthenticated.
- Treat `GenerationError` provider fields as untrusted and potentially
sensitive when adopting the new accessors.
- Run consumer ordinary and race-enabled tests after updating the module.

109
docs/releases/v0.8.0.md Normal file
View File

@@ -0,0 +1,109 @@
# Promptkit v0.8.0
This supplemental changelog and migration guide summarizes the consumer-facing
changes from `v0.7.0` to `v0.8.0`. The annotated `v0.8.0` tag is the
authoritative release record. Exact current contracts belong to the linked
GoDoc and durable documentation.
## Summary
`v0.8.0` activates Promptkit's bounded output-repair workflow:
- failed nonempty-text, JSON, and JSON Schema validation can make a limited
number of corrective model calls;
- corrective calls preserve the original rendered conversation, effective
target, session, structured-output contract, and backend capacity policy;
- results report cumulative usage and the number of corrective calls actually
made; and
- explicitly empty OpenAI-compatible response content now reaches output
validation instead of being classified as a malformed provider envelope.
## Compatibility
This release adds no public declarations or fields and removes none. Existing
source code remains source-compatible.
The behavior of the existing `OutputContract.RepairAttempts` field and prompt
YAML `repair_attempts` field has changed. A positive value now authorizes real
additional model calls after eligible validation failures; earlier releases
accepted the field but the public engine remained single-pass. Consumers that
set a positive value should expect additional latency, token usage, and
provider cost when repair is needed.
Repair budgets must now be between zero and three. A positive budget requires
`basic`, `json`, or `json_schema` validation. Values above three and a positive
budget paired with `none` are invalid contracts rather than ignored settings.
An explicitly present empty or whitespace-only string returned by the built-in
OpenAI-compatible client is now a completed generation candidate. `none`
validation permits it, while `basic`, `json`, and `json_schema` classify it
under their ordinary validation rules and may repair it when configured.
Missing, `null`, or non-string content remains a malformed provider response.
## Upgrade
Update the module dependency with:
```sh
go get gitea.maximumdirect.net/eric/promptkit@v0.8.0
go mod tidy
```
Review every prompt definition and request override that sets a positive repair
budget. Use zero or omit the field to retain single-pass execution. Ensure each
positive budget is no greater than three and uses an eligible validation mode,
then run the consuming project's ordinary and race-enabled tests.
## Bounded Output Repair
`repair_attempts` counts corrective calls in addition to the initial model
call. Promptkit validates each completed candidate, stops at the first valid
one, and never exceeds the configured bound. If every candidate remains
invalid, the run completes successfully with the final candidate and its
failed validation result rather than returning an operational error.
Each correction starts from the original rendered messages and includes only
the latest invalid candidate and latest validation diagnostics. JSON Schema
mode retains the provider-native structured-output request as its first line of
defense. Promptkit performs only deterministic structural validation; a valid
response is not necessarily factual or correct for an application's domain.
Usage in the final result is cumulative across the initial response and every
completed corrective response. `ValidationResult.RepairAttempts` reports the
number of corrective calls actually made. Corrective generation failures use
the same public generation-error categories and structured provider details as
an initial generation failure.
See the [output-contract format reference](../formats.md#output-contract), the
[consumer repair example](../consumers/pkg-promptkit.md#repair-a-structured-result),
and the [`OutputContract` and `ValidationResult` GoDoc](../../types.go) for the
current contracts.
## Explicit Empty Content
The built-in OpenAI-compatible client now distinguishes an explicitly present
empty string from a missing or malformed `content` field. This aligns built-in
and injected clients by letting the selected output contract decide whether an
empty candidate is acceptable, invalid, or eligible for repair.
See the
[OpenAI-compatible response contract](../integrations/openai-compatible-chat.md#response-handling)
for the exact envelope behavior.
## Public API Changes
None. This release activates and tightens the documented behavior of existing
fields.
## Consumer Action
- Remove or set `repair_attempts` to zero where execution must remain
single-pass.
- Keep every positive repair budget at three or fewer and pair it with
`basic`, `json`, or `json_schema` validation.
- Account for additional latency, usage, and provider cost when enabling
repair.
- Continue checking the returned validation status because bounded repair can
exhaust without producing a valid candidate.
- Review workflows that previously treated explicit empty provider content as
a generation error.

61
docs/roadmap/deferred.md Normal file
View File

@@ -0,0 +1,61 @@
# Deferred Feature Ideas
## Purpose
This document catalogs feature ideas that remain potentially useful but have
been deliberately postponed. These ideas are not awaiting ordinary selection
from the [future feature catalog](future.md); each has a stated reason to wait
and should be reconsidered only when its trigger becomes relevant.
Deferred entries are not commitments, schedules, active implementation plans,
or descriptions of current behavior. When an entry is reactivated, move it to
`future.md` for evaluation or directly into a focused roadmap after its open
design dependencies have been resolved.
## Deferred Ideas
### Semantic Execution-Target Fingerprints
**Reason for deferral:** A stable digest requires a deliberate semantic-
equality and versioning design. Notarius can safely use conservative source
hashes and a Promptkit release marker today, while Weatherreporter does not
currently reuse LLM-dependent checkpoints.
Promptkit could expose an opaque equality value for a resolved profile and its
effective generation target. This would let checkpointing consumers detect
generation-affecting configuration changes without hashing YAML presentation
or depending on Promptkit's built-in catalog layout.
The digest should change with semantically relevant state such as the resolved
model, endpoint, backend routing identity, request defaults, extra parameters,
profile generation settings, and selected built-in profile semantics. It
should exclude credential values, concurrency and queue policy, source paths,
comments, formatting, and other representation-only changes. Whether a
credential environment-variable name affects equality must be decided
explicitly. The encoding should remain opaque and internally versioned so
Promptkit can deliberately invalidate earlier digests when its resolution
semantics change.
Reconsider this idea when a downstream consumer needs Promptkit-owned
checkpoint equality or when a broader semantic identity design is selected.
### Eager Source Validation
**Reason for deferral:** Exact prompt and profile inspection may already
provide a sufficiently small validation surface. Experience from downstream
adoption should establish whether an engine-wide operation would add enough
value to justify its broader contract.
Promptkit could provide an explicit offline operation that discovers and
structurally validates configured prompt, profile, and schema sources without
model generation. The normal `NewEngine` path would remain lazy.
An eager operation would need coherent handling for duplicate prompt IDs and
versions, strict YAML decoding, referenced content files, profile/backend
membership, schema syntax and transitive references, context cancellation,
and source-specific public errors. Credential declarations must remain
separate from credential values; checking current environment availability,
if supported at all, should be an explicit option and must not expose secrets.
Reconsider this idea after downstream use of `InspectPrompt`,
`InspectProfile`, and fixture-based preparation demonstrates a concrete gap.

View File

@@ -12,6 +12,9 @@ consumer value, and important scope boundaries. Defer API design,
implementation details, sequencing, and acceptance criteria until an idea is implementation details, sequencing, and acceptance criteria until an idea is
selected. selected.
Ideas that have been deliberately postponed rather than left available for
ordinary selection belong in the [deferred catalog](deferred.md).
## Using This Catalog ## Using This Catalog
- Add an idea when its purpose and likely value can be stated clearly. - Add an idea when its purpose and likely value can be stated clearly.
@@ -23,6 +26,8 @@ selected.
- When an idea is selected, move its active planning to a focused roadmap or, - When an idea is selected, move its active planning to a focused roadmap or,
when it requires a durable architectural decision, an ADR. Update when it requires a durable architectural decision, an ADR. Update
current-state documentation only when implementation lands. current-state documentation only when implementation lands.
- Move an idea to `deferred.md` when maintainers decide to retain it but wait
for a stated design dependency, demand signal, or reconsideration trigger.
- Remove ideas that are no longer relevant. Retain a rejected idea only when - Remove ideas that are no longer relevant. Retain a rejected idea only when
its rationale is likely to prevent repeated reconsideration. its rationale is likely to prevent repeated reconsideration.
@@ -33,7 +38,8 @@ consumers.
## Ideas ## Ideas
No ideas currently await selection. No ideas are currently awaiting selection. Active feature work belongs in its
focused roadmap rather than this catalog.
## Entry Format ## Entry Format

View File

@@ -1,245 +0,0 @@
# Notarius PromptKit Wishlist
## Purpose
This document records features and interface changes that would be useful
additions to PromptKit from the perspective of the maintainers of Notarius, a
downstream application that consumes PromptKit.
PromptKit now provides the capabilities Notarius currently needs. The
remaining deferred ideas are optional opportunities to improve checkpointing
and operational observability.
The examples are API sketches intended to communicate the desired capability,
not prescriptive names or finalized Go contracts.
## Priority 1: Atomic Execution With Prepared Details
**Disposition:** Implemented through [`Engine.PrepareExecution` and
`Engine.RunPrepared`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#prepare-now-and-execute-the-same-snapshot-later).
A separate `RunDetailed` method is not cataloged.
### Downstream need
Notarius needs both:
- the completed `RunResult`; and
- the rendered messages, effective output contract, hashes, and other
preparation details exposed by `PreparedRun`.
Notarius uses the prepared details to construct redaction-aware debug bundles
and retain enough information to diagnose model behavior.
### Implemented behavior
Notarius can prepare one frozen execution snapshot, retain a caller-owned and
credential-redacted `Details` value for its debug bundle, and execute the same
snapshot through `RunPrepared`. The opaque handle is engine-bound and
single-use; an unused handle can be released with `Discard`. The consumer
guide and exported GoDoc own the exact lifecycle and failure contracts.
### Value to Notarius
This removes duplicate work from PromptKit-backed calls and ensures that
retained debug material corresponds atomically to the actual execution.
## Priority 2: Prompt-Independent Profile Inspection
**Disposition:** Implemented as
[`Engine.InspectProfile`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#inspect-a-profile-before-prompt-work).
### Downstream need
Notarius validates configured pipeline profile IDs before beginning a run. It
needs to determine whether:
- a profile exists;
- its referenced backend is registered;
- its execution target can be resolved; and
- it declares a credential requirement that the application may need to
enforce.
This validation should not require model generation.
### Previous integration
Before profile inspection was available, Notarius constructed a synthetic
prompt using `testing/fstest.MapFS`, supplied a dummy transcript, and called
`Engine.Prepare` solely to exercise profile and backend resolution.
### Value to Notarius
The implemented interface eliminates a synthetic production-only prompt
fixture and establishes a direct, supported contract for configuration-time
profile and backend validation.
## Priority 3: Semantic Execution-Target Fingerprints
**Disposition:** Deferred pending a separate semantic-equality design for
resolved execution targets.
### Downstream need
Notarius checkpoints model-backed pipeline stages. A checkpoint must not be
reused when generation-affecting PromptKit configuration changes.
Notarius therefore needs a stable equality signal for the effective profile
and backend target used by a pipeline.
### Current integration
Notarius currently constructs this identity itself from:
- a manually maintained marker for the PromptKit release and built-in profile
catalog;
- raw hashes of configured profile files; and
- a separate hash of the configured conventional local-backend endpoint.
This is safe but conservative and coupled to PromptKit details. Raw file
hashing also invalidates checkpoints for semantically irrelevant YAML changes,
such as comments or formatting.
### Requested capability
Expose an opaque semantic digest for a resolved profile and its effective
generation target. It could be returned by the proposed profile-resolution
API:
```go
type ResolvedProfile struct {
ProfileID string
BackendID string
EffectiveTarget ExecutionTarget
ExecutionDigest string
}
```
Alternatively, PromptKit could expose a dedicated method such as
`ProfileExecutionDigest(profileID)`.
### Desired equality semantics
The digest should change when generation-affecting state changes, including:
- resolved model and endpoint;
- backend routing identity;
- backend request defaults and extra parameters;
- profile generation parameters; and
- the semantic identity of any selected built-in profile.
The digest should not incorporate:
- credential values;
- concurrency or queue capacity;
- filesystem source paths;
- YAML comments or formatting; or
- other settings that affect scheduling or source representation without
changing the generation target.
The credential environment-variable name may need to participate if changing
it can select a materially different provider account or target. PromptKit
should define this deliberately while continuing to exclude the resolved
secret value.
### Design considerations
- Treat the digest as an opaque equality value rather than a public encoding
of internal structures.
- Document which categories of change affect equality.
- Include a versioned semantic marker internally so PromptKit can deliberately
invalidate old digests when its resolution semantics change.
- Prefer a per-profile digest over a digest of every profile known to an
engine. Notarius generally knows which profiles a resolved pipeline uses.
- Do not require consumers to know PromptKit's built-in catalog version.
### Value to Notarius
This would let Notarius remove its PromptKit release marker and raw
profile-source fingerprinting, reduce unnecessary checkpoint invalidation, and
delegate execution-target equality to the component that owns target
resolution.
## Priority 4: Structured Capacity Errors
**Disposition:** Implemented behavior. See the consumer guide's
[Handle Errors](../consumers/pkg-promptkit.md#handle-errors) section.
### Downstream need
Notarius translates PromptKit backend-capacity rejection into a
provider-neutral application error. When multiple backends are active,
operators would benefit from knowing which backend rejected admission without
parsing an error string or exposing endpoint details.
### Implemented behavior
PromptKit now retains the broad capacity classification while allowing
Notarius to obtain the selected backend ID without parsing diagnostic text.
The consumer guide owns the application workflow, including retry and backoff
policy.
### Value to Notarius
This improves operational diagnostics and future metrics while preserving
the provider-neutral error boundary used by Notarius.
## Capabilities PromptKit Already Provides Well
The current PromptKit boundary is sufficient for Notarius's implemented
behavior. In particular, PromptKit already provides:
- filesystem, `fs.FS`, and programmatic prompt, profile, and schema sources;
- offline preparation without model execution;
- structured output and content validation;
- direct session propagation;
- tri-state per-run reasoning overrides;
- selected profile, backend, model, endpoint, effective parameters, hashes, and
token-usage provenance;
- endpoint-only profiles;
- the conventional `local` backend helper;
- arbitrary engine-scoped `Backend` registrations;
- backend authentication environment names, extra parameters, concurrency
limits, and queue-capacity policies;
- provider-client and artifact-reader extension interfaces;
- context cancellation; and
- useful public error sentinels, including profile absence and capacity
exhaustion.
The wishlist does not imply that Notarius needs PromptKit to broaden its core
responsibilities. It primarily asks for more direct access to information and
operations that PromptKit already computes internally.
## Responsibilities That Should Remain In Notarius
The following concerns belong to the downstream application and should not
move into PromptKit for the sake of Notarius:
- pipeline staging, dependencies, and generated references;
- application-wide scheduling across providers and backends;
- module and validation retry policy;
- checkpoints, resume, and recomputation;
- durable run artifacts and manifests;
- D&D prompts, schemas, extractors, validators, and normalizers;
- Notarius configuration-file parsing and precedence;
- domain-specific prompt-cache prefix policy; and
- application-specific redaction, retention, and debug-bundle policy.
PromptKit's complete `Backend` API already supports custom IDs, multiple local
endpoints, authentication, extra parameters, and explicit queue policies.
Whether Notarius exposes those capabilities in its own configuration is an
application-policy decision, not an upstream PromptKit gap.
## Suggested Upstream Sequence
For downstream adoption and any remaining upstream work, the useful order is:
1. Adopt prepared execution for atomic details and results.
2. Add a semantic execution-target digest, preferably alongside profile
inspection.
3. Use the implemented typed capacity error where backend admission diagnostics
are needed.
The first removes the concrete execution workaround. The second would improve
checkpoint correctness and reduce coupling. The third is operational polish.

View File

@@ -1,331 +0,0 @@
# Weatherreporter PromptKit Wishlist
## Purpose
This document records features and interface changes that would be useful
additions to PromptKit from the perspective of the maintainers of
Weatherreporter, a downstream application planning to replace its Scriptorium
CLI integration with PromptKit.
PromptKit now provides the capabilities Weatherreporter needs for the
migration. The remaining deferred ideas are optional opportunities to validate
configuration earlier and improve durable failure diagnostics.
The examples are API sketches intended to communicate the desired capability,
not prescriptive names or finalized Go contracts. The related
[Notarius PromptKit wishlist](notarius-promptkit-wishlist.md) proposes several
overlapping features from another downstream consumer's perspective.
## Priority 1: Executable Preparation Handles
**Disposition:** Implemented as [`Engine.PrepareExecution` and
`Engine.RunPrepared`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#prepare-now-and-execute-the-same-snapshot-later).
### Downstream need
Weatherreporter treats prompt preparation as a durable preflight boundary. It
needs to:
1. prepare the exact request that will be executed;
2. persist a safe preparation record before starting the provider call; and
3. execute without reloading or rerendering prompt, profile, schema, or input
sources.
Persisting preflight before generation leaves useful evidence when a provider
call fails or the process is interrupted during generation.
### Implemented behavior
Weatherreporter can prepare one frozen execution snapshot, persist a
caller-owned and credential-redacted `Details` value, and execute that same
snapshot through `RunPrepared`. The opaque handle is engine-bound and
single-use; an unused handle can be released with `Discard`. The consumer
guide and exported GoDoc own the exact lifecycle, credential, cancellation,
and capacity contracts.
### Value to Weatherreporter
This preserves Weatherreporter's durable preflight behavior, removes duplicate
work, eliminates the source-consistency window, and ensures that persisted
provenance describes the actual execution.
## Priority 2: Prompt-Definition Inspection
**Disposition:** Implemented as
[`Engine.InspectPrompt`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#inspect-a-prompt-before-preparation).
### Downstream need
Weatherreporter has a fixed registry of seven report definitions. Each report
selects a prompt ID and one of two output workflows:
- direct Markdown; or
- structured generated text followed by application-owned domain validation
and Markdown template rendering.
Weatherreporter will embed the PromptKit prompt definitions and private
response schemas that implement those reports. It needs to validate that the
report registry and embedded prompt corpus agree before weather collection or
provider execution.
### Previous integration option
Before prompt inspection was available, Weatherreporter could maintain
synthetic data-package fixtures and call `Engine.Prepare` for every report
prompt during tests. Runtime validation could also occur through the ordinary
per-report preparation stage.
This required complete placeholder inputs and profile resolution when the
application primarily wanted to inspect prompt identity and declared contracts.
### Value to Weatherreporter
The implemented interface lets Weatherreporter directly verify that every
report prompt exists, requires the curated `data_package` input, and declares
the expected Markdown or JSON Schema output contract. It reduces synthetic
test setup and moves failures ahead of weather collection.
## Priority 3: Prompt-Independent Profile Inspection
**Disposition:** Implemented as
[`Engine.InspectProfile`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#inspect-a-profile-before-prompt-work).
### Downstream need
Weatherreporter will allow operators to select an external PromptKit profile
source and may allow an explicit profile override. It needs to reject a missing
profile, unknown backend, or malformed execution target before collecting
weather data or writing report artifacts, and to apply its own policy to a
reported credential requirement.
### Value to Weatherreporter
The implemented interface improves fail-fast configuration validation and gives
operator-facing errors direct profile and backend context. It remains optional
for the initial migration.
## Priority 4: Eager Source Validation
**Disposition:** Deferred until prompt and profile inspection have been used
to determine whether a broader engine-wide validation operation is still
needed.
### Downstream need
PromptKit deliberately defers reading and validating filesystem and `fs.FS`
prompt, profile, and schema content until a request needs it. Weatherreporter
has a small fixed embedded prompt corpus and one optional external profile
source. It would benefit from an explicit offline validation operation for
tests, startup diagnostics, and configuration checks.
### Current integration option
Weatherreporter can prepare every report prompt with fixture inputs and inspect
any explicit profiles individually. That provides strong coverage but requires
consumer-maintained traversal and synthetic material.
### Requested capability
Consider an opt-in source-validation operation:
```go
type SourceValidationOptions struct {
RequireCredentials bool
}
func (e *Engine) ValidateSources(
ctx context.Context,
opts SourceValidationOptions,
) error
```
The operation should eagerly discover and structurally validate the configured
prompt, profile, and schema sources without model generation.
### Design considerations
- Keep deferred validation as the normal `NewEngine` behavior.
- Make eager validation an explicit consumer choice.
- Validate duplicate IDs and versions, strict YAML decoding, referenced content
files, profile/backend membership, schema syntax, and schema references.
- Distinguish structural credential declarations from current environment
availability.
- Do not read or expose credential values when credential availability is not
requested.
- Preserve source-specific public error identities and useful path context.
- Respect context cancellation during filesystem discovery and schema work.
- Consider whether exact prompt and profile inspection APIs already provide a
smaller sufficient surface before adding an engine-wide operation.
### Value to Weatherreporter
This would simplify offline corpus checks and catch malformed operator profile
sources before report work begins. It is helpful but lower priority than exact
prompt and profile inspection.
## Priority 5: Structured Generation Errors
**Disposition:** Deferred pending stronger downstream demand and a narrower
design that does not duplicate prepared provenance or impose HTTP-specific
fields on injected model clients.
### Downstream need
Weatherreporter preserves redacted, inspectable failure receipts for report
runs. When model generation fails operationally, it needs to classify the
failure and retain safe execution context without parsing error prose.
Prompt preparation already supplies selected profile, backend, and model
identity. Provider status classification would add useful operator context,
especially when the built-in OpenAI-compatible client receives a non-success
HTTP status.
### Current integration option
PromptKit exposes `ErrLLMGenerate` and preserves injected client errors through
`errors.Is`. Weatherreporter can reliably classify generation failure and use
its preparation record for profile, backend, and model provenance. Any further
diagnostic detail remains a redacted error string.
### Requested capability
Consider a typed generation error that continues to match `ErrLLMGenerate`:
```go
type GenerationError struct {
BackendID string
Model string
StatusCode int
}
```
The exact fields may differ. The useful contract is safe structured context
available through `errors.As`, while `errors.Is(err, ErrLLMGenerate)` remains
compatible.
### Design considerations
- Include only fields that PromptKit knows reliably and can expose safely.
- Treat an HTTP status as optional because injected model clients may not use
HTTP.
- Do not expose provider response bodies, endpoints, credential environment
names, credential values, request content, or generated content.
- Do not make a structured error a second source of prompt/profile provenance
already present in a prepared execution.
- Preserve injected client error identity.
- Keep retry and backoff policy with the consuming application.
### Value to Weatherreporter
This would improve durable failure receipts and troubleshooting, particularly
for built-in transport failures. It is not required if preparation details and
the existing sentinel remain available.
## Lower-Priority Shared Wishlist Items
### Structured Capacity Errors
**Disposition:** Implemented behavior. See the consumer guide's
[Handle Errors](../consumers/pkg-promptkit.md#handle-errors) section.
PromptKit now exposes the stable backend ID for capacity rejection without
requiring Weatherreporter to parse error text.
Weatherreporter currently generates batch reports sequentially and constructs
one engine per invocation, so engine-local capacity exhaustion is unlikely in
the initial design. The typed error becomes more valuable if report generation
later becomes concurrent or PromptKit engines become longer-lived.
It should not block adoption.
### Semantic Execution-Target Fingerprints
**Disposition:** Deferred pending a separate semantic-equality design for
resolved execution targets.
The semantic target digest proposed by the
[Notarius wishlist](notarius-promptkit-wishlist.md#priority-3-semantic-execution-target-fingerprints)
would provide a compact equality signal for audit metadata.
Weatherreporter does not currently reuse LLM-dependent checkpoints. Its Recent
Changes behavior compares deterministic module snapshots rather than generated
reports, so the digest has no immediate cache-correctness role. Existing
PromptKit result metadata is sufficient for the initial integration. A digest
would still be useful provenance and future-proofing, but it is not a
migration priority.
## Capabilities PromptKit Already Provides Well
PromptKit already provides the essential Weatherreporter integration surface:
- importable in-process engine construction;
- filesystem, `fs.FS`, and programmatic prompt, profile, and schema sources;
- offline preparation without model execution;
- prepared execution handles for a durable preflight boundary;
- exact prompt and profile inspection;
- versioned prompt selection;
- text, Markdown, JSON, and JSON Schema output contracts;
- single-pass output validation with raw output retained after completed
validation failure;
- inline input artifacts with provenance URIs and input hashes;
- selected profile, backend, model, effective target, prompt hashes, timing,
and token-usage provenance;
- endpoint-only profiles and engine-scoped backend registration;
- injected model-client and artifact-reader interfaces;
- caller cancellation, generation timeout, and transport timeout behavior; and
- public error sentinels for configuration, prompt, profile, artifact,
validation, capacity, and generation failures; and
- structured backend identity for capacity rejection.
These capabilities are sufficient for Weatherreporter to adopt PromptKit
without waiting for new upstream work.
## Responsibilities That Should Remain In Weatherreporter
The following concerns belong to Weatherreporter and should not move into
PromptKit:
- report definitions, valid periods, batches, and output naming;
- application prompt content and private report response schemas;
- deterministic weather facts, modules, and Recent Changes;
- curated `data_package` construction and persistence;
- generated-text domain validation and Markdown template rendering;
- managed artifact paths, atomic writes, metadata, and inspection commands;
- preparation, execution, raw-output, and failure-receipt schemas;
- CLI configuration loading and precedence;
- debug enablement, redaction, placement, sensitivity, and retention;
- distributor notification;
- batch continuation and any future retry policy; and
- application-level compatibility and migration policy.
## Suggested Upstream Sequence
For downstream adoption and any remaining upstream work, the useful order is:
1. Adopt the implemented executable preparation handles.
2. Consider eager source validation after evaluating whether the two exact
inspection APIs are sufficient.
3. Add structured generation errors.
4. Use structured capacity errors and consider semantic execution-target
fingerprints as lower-priority operational improvements.
The first item removes the material integration workaround. Prompt and profile
inspection improve fail-fast validation. The remaining items are optional
ergonomic and diagnostic improvements.
## Adoption Sequencing
Weatherreporter should not wait for the deferred wishlist items. The current
PromptKit interface is sufficient when Weatherreporter:
- embeds immutable prompt and schema assets;
- supplies immutable inline data-package bytes;
- constructs one engine per CLI invocation;
- prepares an execution, persists selected `Details`, and calls
`RunPrepared`; and
- keeps PromptKit behind a weatherreporter-owned adapter contract.
Prompt inspection, profile inspection, source validation, structured errors,
capacity details, and semantic fingerprints should not gate adoption.

191
engine.go
View File

@@ -53,9 +53,9 @@ var (
// an execution profile or resolve its backend, except for the profile // an execution profile or resolve its backend, except for the profile
// not-found case represented by ErrProfileNotFound. // not-found case represented by ErrProfileNotFound.
ErrProfileLoad = errors.New("failed to load execution profile") ErrProfileLoad = errors.New("failed to load execution profile")
// ErrAPIKeyEnvMissing identifies an APIKeyEnv whose environment variable is // ErrAPIKeyEnvMissing identifies an explicitly required APIKeyEnv whose
// unset or empty when no direct RunRequest.APIKey takes precedence. Such an // environment variable is unset or empty after direct RunRequest.APIKey
// error also matches ErrInvalidRequest. // precedence is applied. Such an error also matches ErrInvalidRequest.
ErrAPIKeyEnvMissing = errors.New("api_key_env points to an unset environment variable") ErrAPIKeyEnvMissing = errors.New("api_key_env points to an unset environment variable")
// ErrArtifactLoad identifies a failure to resolve an input artifact. Errors // ErrArtifactLoad identifies a failure to resolve an input artifact. Errors
// returned by an injected ArtifactReader remain available through errors.Is. // returned by an injected ArtifactReader remain available through errors.Is.
@@ -69,8 +69,9 @@ var (
// request, an LLM or provider rate-limit response, or ErrLLMGenerate. // request, an LLM or provider rate-limit response, or ErrLLMGenerate.
ErrCapacityExceeded = errors.New("backend capacity exceeded") ErrCapacityExceeded = errors.New("backend capacity exceeded")
// ErrLLMGenerate identifies a model-client failure or a nil successful // ErrLLMGenerate identifies a model-client failure or a nil successful
// response. Errors returned by an injected LLMClient remain available // response. A built-in OpenAI-compatible non-2xx response is available as a
// through errors.Is. // [GenerationError]. Errors returned by an injected LLMClient remain
// available through errors.Is.
ErrLLMGenerate = errors.New("failed to generate output") ErrLLMGenerate = errors.New("failed to generate output")
// ErrValidation identifies an operational failure to load or compile a // ErrValidation identifies an operational failure to load or compile a
// schema or validate output. A completed validation whose Status is // schema or validate output. A completed validation whose Status is
@@ -98,9 +99,10 @@ type Config struct {
// It is required unless a WithPromptFS or WithPromptFile option supplies the // It is required unless a WithPromptFS or WithPromptFile option supplies the
// prompt source. // prompt source.
PromptDir string PromptDir string
// ProfileDir is an optional directory whose profiles take precedence over // ProfileDir is an optional ordinary configured source whose profiles take
// embedded built-in profiles. An empty value selects only built-ins unless // precedence over application fallback and embedded built-in profiles. An
// profile options are also supplied. // empty value selects the lower-precedence sources unless a profile-source
// option supplies the ordinary source.
ProfileDir string ProfileDir string
// SchemaDir is the root for JSON Schema files. An empty value uses the // SchemaDir is the root for JSON Schema files. An empty value uses the
// current directory. WithSchemaFS or WithSchemaFile replaces this source. // current directory. WithSchemaFS or WithSchemaFile replaces this source.
@@ -119,12 +121,12 @@ type Config struct {
// Option customizes engine construction. // Option customizes engine construction.
// //
// NewEngine applies options in argument order and ignores nil options. Within // NewEngine applies options in argument order and ignores nil options. Within
// each prompt-source, profile-source, in-memory-profile, schema-source, // each prompt-source, ordinary-profile-source, fallback-profile-source,
// model-client, and artifact-reader category, the last non-nil valid option // in-memory-profile, schema-source, model-client, and artifact-reader
// replaces earlier options in that category. WithBackend is the additive // category, the last non-nil valid option replaces earlier options in that
// exception: unique registrations accumulate, and a repeated backend ID is an // category. WithBackend is the additive exception: unique registrations
// error rather than a replacement. An invalid option fails construction even // accumulate, and a repeated backend ID is an error rather than a replacement.
// if a later option would replace it. // An invalid option fails construction even if a later option would replace it.
type Option interface { type Option interface {
apply(*engineOptions) error apply(*engineOptions) error
} }
@@ -140,11 +142,13 @@ type engineOptions struct {
artifactReader artifactadapter.Reader artifactReader artifactadapter.Reader
promptDefs promptdef.Repository promptDefs promptdef.Repository
profiles profile.Repository profiles profile.Repository
fallbackProfiles profile.Repository
memoryProfiles profile.Repository memoryProfiles profile.Repository
backends []domain.Backend backends []domain.Backend
validator validate.Validator validator validate.Validator
promptSource bool promptSource bool
profileSource bool profileSource bool
fallbackProfileSource bool
memorySource bool memorySource bool
validatorSource bool validatorSource bool
artifactSource bool artifactSource bool
@@ -216,7 +220,7 @@ func WithPromptFile(path string) Option {
if err != nil { if err != nil {
return err return err
} }
options.promptDefs = promptdef.NewFSRepository(fsys, root) options.promptDefs = promptdef.NewFileRepository(fsys, root, filepath.Dir(path))
options.promptSource = true options.promptSource = true
return nil return nil
}) })
@@ -224,12 +228,12 @@ func WithPromptFile(path string) Option {
// WithProfileFS loads execution profiles from fsys under root. // WithProfileFS loads execution profiles from fsys under root.
// //
// Profiles from this source overlay built-in profiles. Profile YAML must use // Profiles from this ordinary configured source take precedence over
// api_key_env for environment-based credentials; raw API keys are rejected. // application fallback and built-in profiles. Profile YAML must use api_key_env
// fsys must be non-nil and root must be non-empty; otherwise NewEngine fails // for environment-based credentials; raw API keys are rejected. fsys must be
// with ErrInvalidConfig. This option replaces Config.ProfileDir and earlier // non-nil and root must be non-empty; otherwise NewEngine fails with
// file or FS profile-source options, but remains below WithProfiles in // ErrInvalidConfig. This option replaces Config.ProfileDir and earlier file or
// precedence. // FS profile-source options, but remains below WithProfiles in precedence.
func WithProfileFS(fsys fs.FS, root string) Option { func WithProfileFS(fsys fs.FS, root string) Option {
return optionFunc(func(options *engineOptions) error { return optionFunc(func(options *engineOptions) error {
if fsys == nil { if fsys == nil {
@@ -246,11 +250,12 @@ func WithProfileFS(fsys fs.FS, root string) Option {
// WithProfileFile loads execution profiles from the single profile file at path. // WithProfileFile loads execution profiles from the single profile file at path.
// //
// The profile overlays built-in profiles. Profile YAML must use api_key_env for // The profile takes precedence over application fallback and built-in profiles.
// environment-based credentials; raw API keys are rejected. path must name an // Profile YAML must use api_key_env for environment-based credentials; raw API
// existing non-directory file when NewEngine applies the option. This option // keys are rejected. path must name an existing non-directory file when
// replaces Config.ProfileDir and earlier file or FS profile-source options, // NewEngine applies the option. This option replaces Config.ProfileDir and
// but remains below WithProfiles in precedence. // earlier file or FS profile-source options, but remains below WithProfiles in
// precedence.
func WithProfileFile(path string) Option { func WithProfileFile(path string) Option {
return optionFunc(func(options *engineOptions) error { return optionFunc(func(options *engineOptions) error {
fsys, root, err := fileSource(path) fsys, root, err := fileSource(path)
@@ -263,13 +268,48 @@ func WithProfileFile(path string) Option {
}) })
} }
// WithProfiles configures in-memory profiles that take precedence over // WithFallbackProfileFS supplies application-owned fallback profile
// configured profile files and built-in profiles. // definitions from fsys under root.
// //
// NewEngine validates and copies every profile. IDs must be unique within one // Profile lookup checks, in order, profiles supplied by WithProfiles; the
// call. An invalid profile, duplicate ID, or unsupported ExtraParams value // ordinary configured source selected by WithProfileFile, WithProfileFS, or
// makes construction fail with ErrInvalidConfig. Repeating WithProfiles // Config.ProfileDir; this fallback source; and Promptkit's embedded built-in
// replaces the complete earlier in-memory set rather than merging it. // profiles. Each source supplies a complete profile definition; profile fields
// are not merged between sources. Only an absent profile ID proceeds to the
// next source. A matching read, parse, duplicate, validation, or credential
// format failure stops resolution.
//
// Files use the ordinary strict profile YAML and api_key_env credential rules.
// Loading and validation are lazy: NewEngine validates this option's arguments
// but does not read profile files. fsys must be non-nil and root must be
// nonblank; otherwise NewEngine returns an error matching ErrInvalidConfig.
// Repeating this option replaces the earlier valid fallback source.
//
// This option controls profile-definition lookup, not provider or generation
// failover.
func WithFallbackProfileFS(fsys fs.FS, root string) Option {
return optionFunc(func(options *engineOptions) error {
if fsys == nil {
return ErrInvalidConfig
}
if strings.TrimSpace(root) == "" {
return ErrInvalidConfig
}
options.fallbackProfiles = profile.NewFSRepository(fsys, root)
options.fallbackProfileSource = true
return nil
})
}
// WithProfiles configures in-memory profiles that take precedence over
// ordinary configured, application fallback, and built-in profiles.
//
// NewEngine locally validates and copies every profile. IDs must be unique
// within one call. An invalid local definition, duplicate ID, or unsupported
// ExtraParams value makes construction fail with ErrInvalidConfig. A derived
// profile's base reference and resolved target completeness are checked when it
// is selected or inspected. Repeating WithProfiles replaces the complete
// earlier in-memory set rather than merging it.
func WithProfiles(profiles ...Profile) Option { func WithProfiles(profiles ...Profile) Option {
return optionFunc(func(options *engineOptions) error { return optionFunc(func(options *engineOptions) error {
repo, err := newMemoryProfileRepository(profiles) repo, err := newMemoryProfileRepository(profiles)
@@ -350,13 +390,7 @@ func NewEngine(cfg Config, opts ...Option) (*Engine, error) {
promptDefs = promptdef.NewFilesystemRepository(cfg.PromptDir) promptDefs = promptdef.NewFilesystemRepository(cfg.PromptDir)
} }
profiles := builtin.NewRepositoryWithDirectory(cfg.ProfileDir) profiles := newProfileRepository(cfg.ProfileDir, options)
if options.profileSource {
profiles = builtin.NewRepositoryWithPrimary(options.profiles)
}
if options.memorySource {
profiles = profile.NewOverlayRepository(options.memoryProfiles, profiles)
}
backendRegistry, err := backend.NewRegistry(options.backends) backendRegistry, err := backend.NewRegistry(options.backends)
if err != nil { if err != nil {
@@ -396,7 +430,7 @@ func NewEngine(cfg Config, opts ...Option) (*Engine, error) {
} }
return &Engine{ return &Engine{
runner: usecase.NewRunner( runner: usecase.NewRunnerWithRepairer(
promptDefs, promptDefs,
profiles, profiles,
backendRegistry, backendRegistry,
@@ -404,27 +438,47 @@ func NewEngine(cfg Config, opts ...Option) (*Engine, error) {
prompt.NewGoRenderer(), prompt.NewGoRenderer(),
llmClient, llmClient,
validator, validator,
usecase.NewDefaultOutputRepairer(llmClient),
capacityManager, capacityManager,
), ),
}, nil }, nil
} }
func newProfileRepository(profileDir string, options engineOptions) profile.Repository {
repository := builtin.NewRepository()
if options.fallbackProfileSource {
repository = profile.NewOverlayRepository(options.fallbackProfiles, repository)
}
if options.profileSource {
repository = profile.NewOverlayRepository(options.profiles, repository)
} else if strings.TrimSpace(profileDir) != "" {
repository = profile.NewOverlayRepository(profile.NewFilesystemRepository(profileDir), repository)
}
if options.memorySource {
repository = profile.NewOverlayRepository(options.memoryProfiles, repository)
}
return profile.NewResolvingRepository(repository)
}
func fileSource(name string) (fs.FS, string, error) { func fileSource(name string) (fs.FS, string, error) {
cleanName := strings.TrimSpace(name) if strings.TrimSpace(name) == "" {
if cleanName == "" {
return nil, "", ErrInvalidConfig return nil, "", ErrInvalidConfig
} }
dir := filepath.Dir(cleanName) dir := filepath.Dir(name)
base := filepath.Base(cleanName) base := filepath.Base(name)
if base == "." || base == string(filepath.Separator) || strings.TrimSpace(base) == "" { if base == "." || base == string(filepath.Separator) {
return nil, "", ErrInvalidConfig return nil, "", ErrInvalidConfig
} }
info, err := os.Stat(cleanName) info, err := os.Stat(name)
if err != nil { if err != nil {
return nil, "", fmt.Errorf("%w: failed to access source file %q: %v", ErrInvalidConfig, cleanName, err) return nil, "", fmt.Errorf("%w: failed to access source file %q: %v", ErrInvalidConfig, name, err)
} }
if info.IsDir() { if info.IsDir() {
return nil, "", fmt.Errorf("%w: source path %q must be a file", ErrInvalidConfig, cleanName) return nil, "", fmt.Errorf("%w: source path %q must be a file", ErrInvalidConfig, name)
} }
return os.DirFS(dir), filepath.ToSlash(base), nil return os.DirFS(dir), filepath.ToSlash(base), nil
} }
@@ -481,10 +535,10 @@ func (e *Engine) InspectPrompt(
// //
// InspectProfile trims surrounding whitespace from profileID and looks up the // InspectProfile trims surrounding whitespace from profileID and looks up the
// resulting nonblank ID exactly and case-sensitively through the engine's // resulting nonblank ID exactly and case-sensitively through the engine's
// ordinary in-memory, configured-source, and built-in profile precedence. It // in-memory, ordinary configured-source, application fallback, and built-in
// applies framework defaults, the selected backend, and then the selected // profile precedence. It applies the framework timeout baseline, selected
// profile to EffectiveModelParams without a request override. BackendID is // backend, and then selected profile to EffectiveModelParams without a request
// empty for an endpoint-only profile. // override. BackendID is empty for an endpoint-only profile.
// //
// APIKeyEnv in the returned target is an environment-variable name, never its // APIKeyEnv in the returned target is an environment-variable name, never its
// value. APIKeyRequired instead reports a direct credential requirement and is // value. APIKeyRequired instead reports a direct credential requirement and is
@@ -586,21 +640,25 @@ func (e *Engine) PrepareExecution(ctx context.Context, req RunRequest) (*Prepare
// generated output. // generated output.
// //
// A content-validation failure is a successful run whose // A content-validation failure is a successful run whose
// RunResult.Validation has Status ValidationFailed. An inability to perform // RunResult.Validation has Status ValidationFailed. When its output contract
// validation returns an error matching ErrValidation and no partial result. // has a positive repair budget, a failed eligible validation can make bounded
// The public Engine does not perform output repair, so validation is // additional model calls and stops at the first valid candidate. Exhaustion
// single-pass even when OutputContract.RepairAttempts is positive. // returns the final failed validation result with cumulative usage and actual
// repair attempts. An inability to generate or validate returns an error and
// no partial result.
// //
// Run can return every error category documented by [Engine.Prepare], plus // Run can return every error category documented by [Engine.Prepare], plus
// ErrCapacityExceeded and ErrLLMGenerate. An engine admission rejection is // ErrCapacityExceeded and ErrLLMGenerate. An engine admission rejection is
// discoverable as [CapacityError] and still matches ErrCapacityExceeded. It // discoverable as [CapacityError] and still matches ErrCapacityExceeded. It
// occurs before artifacts, schemas, rendering, or model generation because the // occurs before artifacts, schemas, rendering, or model generation because the
// selected backend's admission capacity is full; it does not match // selected backend's admission capacity is full; it does not match
// ErrInvalidRequest or ErrLLMGenerate. Errors from injected clients remain // ErrInvalidRequest or ErrLLMGenerate. A built-in OpenAI-compatible non-2xx
// available through errors.Is. Cancellation while waiting for model-generation // response is discoverable as [GenerationError]. Errors from injected clients
// capacity matches both ErrLLMGenerate and the context error. Cancellation // remain available through errors.Is. Cancellation while waiting for
// otherwise follows the active collaborator's documented behavior. A nil // model-generation capacity matches both ErrLLMGenerate and the context error.
// Engine returns ErrInvalidConfig. Run returns no partial result on error. // Cancellation otherwise follows the active collaborator's documented
// behavior. A nil Engine returns ErrInvalidConfig. Run returns no partial
// result on error.
func (e *Engine) Run(ctx context.Context, req RunRequest) (*RunResult, error) { func (e *Engine) Run(ctx context.Context, req RunRequest) (*RunResult, error) {
if e == nil || e.runner == nil { if e == nil || e.runner == nil {
return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig) return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig)
@@ -631,16 +689,17 @@ func (e *Engine) Run(ctx context.Context, req RunRequest) (*RunResult, error) {
// //
// The supplied context governs this execution attempt independently of the // The supplied context governs this execution attempt independently of the
// preparation context. It covers credential revalidation, admission, // preparation context. It covers credential revalidation, admission,
// generation, validation, and any internal repair. Result timing begins after // generation, validation, and any bounded output repair. Result timing begins
// the claim and excludes preparation and consumer-held delay. // after the claim and excludes preparation and consumer-held delay.
// //
// RunPrepared can return ErrInvalidRequest, ErrAPIKeyEnvMissing, // RunPrepared can return ErrInvalidRequest, ErrAPIKeyEnvMissing,
// ErrCapacityExceeded, ErrLLMGenerate, or ErrValidation as applicable while // ErrCapacityExceeded, ErrLLMGenerate, or ErrValidation as applicable while
// preserving documented collaborator and context identities. An engine // preserving documented collaborator and context identities. An engine
// admission rejection is discoverable as [CapacityError] and still matches // admission rejection is discoverable as [CapacityError] and still matches
// ErrCapacityExceeded. A completed content-validation rejection is returned // ErrCapacityExceeded. A built-in OpenAI-compatible non-2xx response is
// in RunResult, not as an operational error. An operational error returns no // discoverable as [GenerationError]. A completed content-validation rejection,
// partial RunResult. // including repair exhaustion, is returned in RunResult, not as an operational
// error. An operational error returns no partial RunResult.
func (e *Engine) RunPrepared(ctx context.Context, prepared *PreparedExecution) (*RunResult, error) { func (e *Engine) RunPrepared(ctx context.Context, prepared *PreparedExecution) (*RunResult, error) {
if e == nil || e.runner == nil { if e == nil || e.runner == nil {
return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig) return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig)

File diff suppressed because it is too large Load Diff

View File

@@ -6,6 +6,7 @@ import (
"strings" "strings"
"gitea.maximumdirect.net/eric/promptkit/internal/capacity" "gitea.maximumdirect.net/eric/promptkit/internal/capacity"
"gitea.maximumdirect.net/eric/promptkit/internal/llm"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
"gitea.maximumdirect.net/eric/promptkit/internal/promptdef" "gitea.maximumdirect.net/eric/promptkit/internal/promptdef"
"gitea.maximumdirect.net/eric/promptkit/internal/usecase" "gitea.maximumdirect.net/eric/promptkit/internal/usecase"
@@ -21,6 +22,19 @@ func mapPublicError(err error) error {
return &CapacityError{BackendID: internalCapacityError.BackendID} return &CapacityError{BackendID: internalCapacityError.BackendID}
} }
publicErr := publicErrorFor(err) publicErr := publicErrorFor(err)
var providerHTTPError *llm.ProviderHTTPError
if errors.As(err, &providerHTTPError) && providerHTTPError != nil {
generationErr := newGenerationError(
providerHTTPError.StatusCode(),
providerHTTPError.ProviderCode(),
providerHTTPError.ProviderType(),
providerHTTPError.ProviderMessage(),
)
if publicErr != nil && !errors.Is(publicErr, ErrLLMGenerate) {
return fmt.Errorf("%w: %w", publicErr, generationErr)
}
return generationErr
}
if publicErr == nil { if publicErr == nil {
return err return err
} }

View File

@@ -6,6 +6,7 @@ import (
"fmt" "fmt"
"testing" "testing"
"gitea.maximumdirect.net/eric/promptkit/internal/llm"
"gitea.maximumdirect.net/eric/promptkit/internal/usecase" "gitea.maximumdirect.net/eric/promptkit/internal/usecase"
) )
@@ -48,3 +49,27 @@ func TestMapPublicErrorTranslatesCapacityError(t *testing.T) {
t.Fatalf("mapped backend ID changed with source error: %q", publicErr.BackendID) t.Fatalf("mapped backend ID changed with source error: %q", publicErr.BackendID)
} }
} }
func TestMapPublicErrorPreservesValidationAroundGenerationError(t *testing.T) {
internalErr := fmt.Errorf(
"%w: %w",
usecase.ErrValidation,
&llm.ProviderHTTPError{},
)
err := mapPublicError(internalErr)
if !errors.Is(err, ErrValidation) {
t.Fatalf("mapped error=%v, want ErrValidation", err)
}
if !errors.Is(err, ErrLLMGenerate) {
t.Fatalf("mapped error=%v, want ErrLLMGenerate", err)
}
var generationErr *GenerationError
if !errors.As(err, &generationErr) || generationErr == nil {
t.Fatalf("mapped error=%v, want GenerationError", err)
}
var leakedInternalErr *llm.ProviderHTTPError
if errors.As(err, &leakedInternalErr) {
t.Fatalf("mapped error exposes internal ProviderHTTPError: %v", err)
}
}

87
generation_error.go Normal file
View File

@@ -0,0 +1,87 @@
package promptkit
import "fmt"
// GenerationError reports a non-2xx response from Promptkit's built-in
// OpenAI-compatible client during [Engine.Run] or [Engine.RunPrepared].
//
// Engine-produced values are immutable, caller-owned values. Use errors.Is to
// match [ErrLLMGenerate] and errors.As with a *GenerationError target to obtain
// this type. The four provider accessors expose untrusted provider-controlled
// values that can contain sensitive request or schema fragments. Applications
// must apply their own disclosure policy before logging, displaying, or
// returning them to another caller.
//
// Accessors, Error, GoString, and Unwrap are safe on a nil receiver and a zero
// value. Default and Go-syntax formatting deliberately redact provider details.
// GenerationError has no stable JSON representation.
type GenerationError struct {
statusCode int
providerCode string
providerType string
providerMessage string
}
func newGenerationError(statusCode int, providerCode, providerType, providerMessage string) *GenerationError {
return &GenerationError{
statusCode: statusCode,
providerCode: providerCode,
providerType: providerType,
providerMessage: providerMessage,
}
}
// StatusCode returns the received provider HTTP status code, or zero for a nil
// receiver or zero value.
func (e *GenerationError) StatusCode() int {
if e == nil {
return 0
}
return e.statusCode
}
// ProviderCode returns the normalized provider error code, if present. Its
// value is untrusted and may contain sensitive data.
func (e *GenerationError) ProviderCode() string {
if e == nil {
return ""
}
return e.providerCode
}
// ProviderType returns the normalized provider error type, if present. Its
// value is untrusted and may contain sensitive data.
func (e *GenerationError) ProviderType() string {
if e == nil {
return ""
}
return e.providerType
}
// ProviderMessage returns the bounded normalized provider diagnostic, if
// present. Its value is untrusted and may contain sensitive data.
func (e *GenerationError) ProviderMessage() string {
if e == nil {
return ""
}
return e.providerMessage
}
// Error returns a redacted diagnostic that is not a parsing contract.
func (e *GenerationError) Error() string {
if e == nil || e.statusCode == 0 {
return ErrLLMGenerate.Error()
}
return fmt.Sprintf("%s: provider returned HTTP status %d", ErrLLMGenerate, e.statusCode)
}
// GoString returns the same redacted diagnostic as Error.
func (e *GenerationError) GoString() string {
return e.Error()
}
// Unwrap returns ErrLLMGenerate. It is safe to call on a nil receiver or zero
// value.
func (e *GenerationError) Unwrap() error {
return ErrLLMGenerate
}

View File

@@ -0,0 +1,133 @@
package promptkit_test
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"strings"
"testing"
"gitea.maximumdirect.net/eric/promptkit"
)
func TestBuiltInGenerationError(t *testing.T) {
const (
codeMarker = "provider-code-marker"
typeMarker = "provider-type-marker"
messageMarker = "provider-message-marker"
)
engine := newBuiltInGenerationErrorEngine(t, http.StatusUnprocessableEntity,
`{"error":{"code":"`+codeMarker+`","type":"`+typeMarker+`","message":"`+messageMarker+`"}}`)
result, err := engine.Run(context.Background(), generationErrorRunRequest())
if result != nil {
t.Fatalf("Run result = %#v, want nil", result)
}
assertGenerationError(t, err, http.StatusUnprocessableEntity, codeMarker, typeMarker, messageMarker)
preparedEngine := newBuiltInGenerationErrorEngine(t, http.StatusServiceUnavailable, `{"error":{}}`)
prepared, err := preparedEngine.PrepareExecution(context.Background(), generationErrorRunRequest())
if err != nil {
t.Fatalf("PrepareExecution: %v", err)
}
result, err = preparedEngine.RunPrepared(context.Background(), prepared)
if result != nil {
t.Fatalf("RunPrepared result = %#v, want nil", result)
}
assertGenerationError(t, err, http.StatusServiceUnavailable, "", "", "")
}
func TestBuiltInRepairGenerationError(t *testing.T) {
const (
codeMarker = "repair-code-marker"
typeMarker = "repair-type-marker"
messageMarker = "repair-message-marker"
)
calls := 0
config := contractConfig(frameworkSchemaDir)
config.HTTPClient = &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
calls++
if calls == 1 {
body := `{"choices":[{"message":{"content":"not-json"}}]}`
return &http.Response{StatusCode: http.StatusOK, ContentLength: int64(len(body)), Body: io.NopCloser(strings.NewReader(body))}, nil
}
body := `{"error":{"code":"` + codeMarker + `","type":"` + typeMarker + `","message":"` + messageMarker + `"}}`
return &http.Response{StatusCode: http.StatusUnprocessableEntity, ContentLength: int64(len(body)), Body: io.NopCloser(strings.NewReader(body))}, nil
})}
engine, err := promptkit.NewEngine(config)
if err != nil {
t.Fatalf("NewEngine: %v", err)
}
req := generationErrorRunRequest()
req.Validation = &promptkit.OutputContract{
Format: promptkit.FormatJSON,
ValidationMode: promptkit.ValidationJSON,
RepairAttempts: 1,
}
result, err := engine.Run(context.Background(), req)
if result != nil {
t.Fatalf("Run result = %#v, want nil", result)
}
if calls != 2 {
t.Fatalf("provider calls = %d, want 2", calls)
}
assertGenerationError(t, err, http.StatusUnprocessableEntity, codeMarker, typeMarker, messageMarker)
}
func assertGenerationError(t *testing.T, err error, statusCode int, code, providerType, message string) {
t.Helper()
if !errors.Is(err, promptkit.ErrLLMGenerate) {
t.Fatalf("errors.Is(%v, ErrLLMGenerate) = false", err)
}
var generationErr *promptkit.GenerationError
if !errors.As(err, &generationErr) || generationErr == nil {
t.Fatalf("error = %T, want *GenerationError", err)
}
if generationErr.StatusCode() != statusCode || generationErr.ProviderCode() != code || generationErr.ProviderType() != providerType || generationErr.ProviderMessage() != message {
t.Fatalf("GenerationError = %#v", generationErr)
}
wantFormatted := fmt.Sprintf("failed to generate output: provider returned HTTP status %d", statusCode)
for _, rendered := range []string{fmt.Sprintf("%v", generationErr), fmt.Sprintf("%+v", generationErr), fmt.Sprintf("%#v", generationErr)} {
if rendered != wantFormatted {
t.Fatalf("formatted error = %q, want %q", rendered, wantFormatted)
}
for _, marker := range []string{code, providerType, message} {
if marker != "" && strings.Contains(rendered, marker) {
t.Fatalf("formatted error exposed provider marker %q: %q", marker, rendered)
}
}
}
}
func newBuiltInGenerationErrorEngine(t *testing.T, statusCode int, body string) *promptkit.Engine {
t.Helper()
config := contractConfig(frameworkSchemaDir)
config.HTTPClient = &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
return &http.Response{
StatusCode: statusCode,
ContentLength: int64(len(body)),
Body: io.NopCloser(strings.NewReader(body)),
}, nil
})}
engine, err := promptkit.NewEngine(config)
if err != nil {
t.Fatalf("NewEngine: %v", err)
}
return engine
}
func generationErrorRunRequest() promptkit.RunRequest {
return promptkit.RunRequest{
PromptID: frameworkMarkdownSummaryPromptID,
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."),
},
}
}

View File

@@ -0,0 +1,32 @@
package promptkit
import (
"errors"
"fmt"
"testing"
)
func TestGenerationErrorNilAndZeroValue(t *testing.T) {
var nilError *GenerationError
zeroError := &GenerationError{}
for name, err := range map[string]*GenerationError{
"nil": nilError,
"zero": zeroError,
} {
t.Run(name, func(t *testing.T) {
if err.StatusCode() != 0 || err.ProviderCode() != "" || err.ProviderType() != "" || err.ProviderMessage() != "" {
t.Fatalf("accessors returned provider details: %#v", err)
}
if err.Error() != "failed to generate output" || err.GoString() != "failed to generate output" {
t.Fatalf("redacted formatting = (%q, %q)", err.Error(), err.GoString())
}
if fmt.Sprintf("%v", err) != "failed to generate output" || fmt.Sprintf("%#v", err) != "failed to generate output" {
t.Fatalf("formatted error = (%q, %q)", fmt.Sprintf("%v", err), fmt.Sprintf("%#v", err))
}
if !errors.Is(err, ErrLLMGenerate) {
t.Fatalf("errors.Is(%v, ErrLLMGenerate) = false", err)
}
})
}
}

View File

@@ -16,10 +16,12 @@ import (
var ( var (
ErrUnsupportedRefType = errors.New("unsupported artifact reference type") ErrUnsupportedRefType = errors.New("unsupported artifact reference type")
ErrMissingInlineBody = errors.New("missing body for inline artifact")
ErrMissingFilePath = errors.New("missing file path for file artifact") ErrMissingFilePath = errors.New("missing file path for file artifact")
ErrUnsupportedFile = errors.New("file artifact path is not a regular file")
) )
const fileReadChunkSize = 64 * 1024
// Reader resolves artifact references into actual artifacts. // Reader resolves artifact references into actual artifacts.
type Reader interface { type Reader interface {
Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error) Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error)
@@ -34,7 +36,7 @@ type CompositeReader struct {
func NewCompositeReader() Reader { func NewCompositeReader() Reader {
return &CompositeReader{ return &CompositeReader{
inlineReader: &inlineReader{}, inlineReader: &inlineReader{},
fileReader: &fileReader{}, fileReader: &fileReader{open: openArtifactFile},
} }
} }
@@ -64,10 +66,6 @@ func (r *inlineReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domai
default: default:
} }
if ref.Body == "" {
return nil, ErrMissingInlineBody
}
body := []byte(ref.Body) body := []byte(ref.Body)
return &domain.Artifact{ return &domain.Artifact{
ContentType: defaults.ContentTypeTextPlain, ContentType: defaults.ContentTypeTextPlain,
@@ -78,7 +76,15 @@ func (r *inlineReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domai
}, nil }, nil
} }
type fileReader struct{} type artifactFile interface {
Read([]byte) (int, error)
Stat() (os.FileInfo, error)
Close() error
}
type fileReader struct {
open func(string) (artifactFile, error)
}
func (r *fileReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error) { func (r *fileReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domain.Artifact, error) {
select { select {
@@ -91,25 +97,71 @@ func (r *fileReader) Read(ctx context.Context, ref domain.ArtifactRef) (*domain.
return nil, ErrMissingFilePath return nil, ErrMissingFilePath
} }
return readFileArtifact(ref.URI) return readFileArtifact(ctx, ref.URI, r.open)
} }
func readFileArtifact(path string) (*domain.Artifact, error) { func openArtifactFile(path string) (artifactFile, error) {
file, err := os.Open(path) return os.Open(path)
}
func readFileArtifact(ctx context.Context, path string, open func(string) (artifactFile, error)) (*domain.Artifact, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
info, err := os.Stat(path)
if err != nil {
return nil, fmt.Errorf("failed to read file %s: %w", path, err)
}
if !info.Mode().IsRegular() {
return nil, fmt.Errorf("%w: %s", ErrUnsupportedFile, path)
}
if err := ctx.Err(); err != nil {
return nil, err
}
file, err := open(path)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to read file %s: %w", path, err) return nil, fmt.Errorf("failed to read file %s: %w", path, err)
} }
defer file.Close() defer file.Close()
data, err := io.ReadAll(file) openedInfo, err := file.Stat()
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to read file %s: %w", path, err) return nil, fmt.Errorf("failed to inspect opened file %s: %w", path, err)
}
if !openedInfo.Mode().IsRegular() {
return nil, fmt.Errorf("%w: %s", ErrUnsupportedFile, path)
}
data := make([]byte, 0)
chunk := make([]byte, fileReadChunkSize)
for {
if err := ctx.Err(); err != nil {
return nil, err
}
n, readErr := file.Read(chunk)
if n > 0 {
data = append(data, chunk[:n]...)
}
if err := ctx.Err(); err != nil {
return nil, err
}
if errors.Is(readErr, io.EOF) {
break
}
if readErr != nil {
return nil, fmt.Errorf("failed to read file %s: %w", path, readErr)
}
} }
contentType := mime.TypeByExtension(filepath.Ext(path)) contentType := mime.TypeByExtension(filepath.Ext(path))
if contentType == "" { if contentType == "" {
contentType = defaults.ContentTypeTextPlain contentType = defaults.ContentTypeTextPlain
} }
hash := fmt.Sprintf("%x", sha256.Sum256(data))
if err := ctx.Err(); err != nil {
return nil, err
}
return &domain.Artifact{ return &domain.Artifact{
Name: filepath.Base(path), Name: filepath.Base(path),
@@ -117,6 +169,6 @@ func readFileArtifact(path string) (*domain.Artifact, error) {
Body: data, Body: data,
URI: path, URI: path,
Size: int64(len(data)), Size: int64(len(data)),
Hash: fmt.Sprintf("%x", sha256.Sum256(data)), Hash: hash,
}, nil }, nil
} }

View File

@@ -0,0 +1,43 @@
//go:build linux
package artifact
import (
"context"
"errors"
"path/filepath"
"syscall"
"testing"
"time"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
func TestFileReaderRejectsFIFOBeforeOpen(t *testing.T) {
path := filepath.Join(t.TempDir(), "artifact.fifo")
if err := syscall.Mkfifo(path, 0o600); err != nil {
t.Fatalf("create fifo: %v", err)
}
type result struct {
artifact *domain.Artifact
err error
}
done := make(chan result, 1)
go func() {
artifact, err := NewCompositeReader().Read(context.Background(), domain.ArtifactRef{
Type: domain.ArtifactRefFile,
URI: path,
})
done <- result{artifact: artifact, err: err}
}()
select {
case got := <-done:
if got.artifact != nil || !errors.Is(got.err, ErrUnsupportedFile) {
t.Fatalf("artifact=%#v err=%v, want nil/ErrUnsupportedFile", got.artifact, got.err)
}
case <-time.After(time.Second):
t.Fatal("FIFO read blocked instead of rejecting the non-regular file")
}
}

View File

@@ -1,6 +1,7 @@
package artifact package artifact
import ( import (
"bytes"
"context" "context"
"errors" "errors"
"os" "os"
@@ -11,54 +12,97 @@ import (
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
) )
func TestCompositeReader_Read(t *testing.T) { func TestCompositeReaderRejectsUnsupportedReferences(t *testing.T) {
reader := NewCompositeReader() _, err := NewCompositeReader().Read(context.Background(), domain.ArtifactRef{
ctx := context.Background()
t.Run("inline artifact", func(t *testing.T) {
ref := domain.ArtifactRef{
Type: domain.ArtifactRefInline,
Body: "hello world",
}
art, err := reader.Read(ctx, ref)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if string(art.Body) != "hello world" {
t.Errorf("expected 'hello world', got %s", string(art.Body))
}
if art.ContentType != "text/plain" {
t.Errorf("expected text/plain content type, got %q", art.ContentType)
}
if art.Hash != "b94d27b9934d3e08a52e52d7da7dabfac484efe37a5380ee9088f7ace2efcde9" {
t.Errorf("unexpected hash: %s", art.Hash)
}
if art.Size != int64(len(ref.Body)) {
t.Errorf("expected size %d, got %d", len(ref.Body), art.Size)
}
})
t.Run("inline artifact missing body", func(t *testing.T) {
ref := domain.ArtifactRef{
Type: domain.ArtifactRefInline,
Body: "",
}
_, err := reader.Read(ctx, ref)
if !errors.Is(err, ErrMissingInlineBody) {
t.Errorf("expected ErrMissingInlineBody, got %v", err)
}
})
t.Run("unsupported ref type", func(t *testing.T) {
ref := domain.ArtifactRef{
Type: domain.ArtifactRefType("unsupported"), Type: domain.ArtifactRefType("unsupported"),
URI: "unsupported://bucket/key", URI: "unsupported://bucket/key",
} })
_, err := reader.Read(ctx, ref)
if !errors.Is(err, ErrUnsupportedRefType) { if !errors.Is(err, ErrUnsupportedRefType) {
t.Error("expected error for unsupported type") t.Fatalf("expected ErrUnsupportedRefType, got %v", err)
}
}
func TestCompositeReaderSourceParityAndOpaqueHashes(t *testing.T) {
reader := NewCompositeReader()
hashes := make(map[string]string)
tests := []struct {
name string
content string
}{
{name: "empty", content: ""},
{name: "ordinary", content: "same content"},
{name: "changed", content: "changed content"},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
filePath := filepath.Join(t.TempDir(), "artifact.txt")
if err := os.WriteFile(filePath, []byte(tc.content), 0o600); err != nil {
t.Fatal(err)
}
sources := []struct {
name string
ref domain.ArtifactRef
wantURI string
}{
{
name: "inline",
ref: domain.ArtifactRef{Type: domain.ArtifactRefInline, Body: tc.content},
},
{
name: "inline with uri",
ref: domain.ArtifactRef{Type: domain.ArtifactRefInline, URI: "memory://input", Body: tc.content},
wantURI: "memory://input",
},
{
name: "file",
ref: domain.ArtifactRef{Type: domain.ArtifactRefFile, URI: filePath},
wantURI: filePath,
},
}
var sourceHash string
for _, source := range sources {
t.Run(source.name, func(t *testing.T) {
first, err := reader.Read(context.Background(), source.ref)
if err != nil {
t.Fatalf("first read: %v", err)
}
second, err := reader.Read(context.Background(), source.ref)
if err != nil {
t.Fatalf("second read: %v", err)
}
if string(first.Body) != tc.content || first.Size != int64(len(tc.content)) {
t.Fatalf("body=%q size=%d, want %q/%d", first.Body, first.Size, tc.content, len(tc.content))
}
if first.URI != source.wantURI {
t.Fatalf("URI = %q, want %q", first.URI, source.wantURI)
}
if first.Hash == "" || first.Hash != second.Hash {
t.Fatalf("hashes are not non-empty and stable: %q/%q", first.Hash, second.Hash)
}
if sourceHash == "" {
sourceHash = first.Hash
} else if first.Hash != sourceHash {
t.Fatalf("equal content hashes differ: %q/%q", sourceHash, first.Hash)
}
if source.ref.Type == domain.ArtifactRefFile {
if first.Name != filepath.Base(filePath) || !strings.HasPrefix(first.ContentType, "text/plain") {
t.Fatalf("unexpected file metadata: %+v", first)
}
} else if first.ContentType != "text/plain" {
t.Fatalf("inline content type = %q", first.ContentType)
} }
}) })
}
hashes[tc.name] = sourceHash
})
}
if hashes["empty"] == hashes["ordinary"] || hashes["ordinary"] == hashes["changed"] {
t.Fatalf("changed content did not change opaque hash: %#v", hashes)
}
} }
func TestCompositeReaderCopiesInlineData(t *testing.T) { func TestCompositeReaderCopiesInlineData(t *testing.T) {
@@ -87,94 +131,154 @@ func TestCompositeReaderCopiesInlineData(t *testing.T) {
} }
} }
func TestCompositeReaderHonorsCancellation(t *testing.T) { func TestCompositeReaderHonorsPreCancellation(t *testing.T) {
filePath := filepath.Join(t.TempDir(), "artifact.txt")
if err := os.WriteFile(filePath, []byte("ignored"), 0o600); err != nil {
t.Fatal(err)
}
tests := []struct {
name string
ref domain.ArtifactRef
}{
{name: "inline", ref: domain.ArtifactRef{Type: domain.ArtifactRefInline, Body: "ignored"}},
{name: "file", ref: domain.ArtifactRef{Type: domain.ArtifactRefFile, URI: filePath}},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background()) ctx, cancel := context.WithCancel(context.Background())
cancel() cancel()
_, err := NewCompositeReader().Read(ctx, domain.ArtifactRef{ artifact, err := NewCompositeReader().Read(ctx, tc.ref)
Type: domain.ArtifactRefInline, if artifact != nil || !errors.Is(err, context.Canceled) {
Body: "ignored", t.Fatalf("artifact=%#v err=%v, want nil/context.Canceled", artifact, err)
}
}) })
if !errors.Is(err, context.Canceled) {
t.Fatalf("expected context cancellation, got %v", err)
} }
} }
func TestFileReader_Read(t *testing.T) { func TestFileReaderFailuresAndMetadata(t *testing.T) {
content := []byte("test file content")
filePath := filepath.Join(t.TempDir(), "artifact.txt")
if err := os.WriteFile(filePath, content, 0o600); err != nil {
t.Fatal(err)
}
reader := NewCompositeReader() reader := NewCompositeReader()
ctx := context.Background()
t.Run("file artifact loading", func(t *testing.T) {
ref := domain.ArtifactRef{
Type: domain.ArtifactRefFile,
URI: filePath,
}
art, err := reader.Read(ctx, ref)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if string(art.Body) != string(content) {
t.Errorf("expected %s, got %s", string(content), string(art.Body))
}
if art.Name != filepath.Base(filePath) {
t.Errorf("expected name %q, got %q", filepath.Base(filePath), art.Name)
}
if !strings.HasPrefix(art.ContentType, "text/plain") {
t.Errorf("expected text content type, got %q", art.ContentType)
}
if art.URI != filePath {
t.Errorf("expected URI %q, got %q", filePath, art.URI)
}
if art.Size != int64(len(content)) {
t.Errorf("expected size %d, got %d", len(content), art.Size)
}
if art.Hash != "60f5237ed4049f0382661ef009d2bc42e48c3ceb3edb6600f7024e7ab3b838f3" {
t.Errorf("unexpected hash: %s", art.Hash)
}
})
t.Run("missing file path", func(t *testing.T) { t.Run("missing file path", func(t *testing.T) {
ref := domain.ArtifactRef{ _, err := reader.Read(context.Background(), domain.ArtifactRef{Type: domain.ArtifactRefFile})
Type: domain.ArtifactRefFile,
URI: "",
}
_, err := reader.Read(ctx, ref)
if !errors.Is(err, ErrMissingFilePath) { if !errors.Is(err, ErrMissingFilePath) {
t.Errorf("expected ErrMissingFilePath, got %v", err) t.Fatalf("expected ErrMissingFilePath, got %v", err)
} }
}) })
t.Run("missing file", func(t *testing.T) { t.Run("missing file", func(t *testing.T) {
ref := domain.ArtifactRef{ _, err := reader.Read(context.Background(), domain.ArtifactRef{
Type: domain.ArtifactRefFile, Type: domain.ArtifactRefFile,
URI: filepath.Join(t.TempDir(), "missing.txt"), URI: filepath.Join(t.TempDir(), "missing.txt"),
} })
if _, err := reader.Read(ctx, ref); err == nil { if err == nil {
t.Fatal("expected missing file error") t.Fatal("expected missing file error")
} }
}) })
t.Run("directory rejected before open", func(t *testing.T) {
artifact, err := reader.Read(context.Background(), domain.ArtifactRef{
Type: domain.ArtifactRefFile,
URI: t.TempDir(),
})
if artifact != nil || !errors.Is(err, ErrUnsupportedFile) {
t.Fatalf("artifact=%#v err=%v, want nil/ErrUnsupportedFile", artifact, err)
}
})
t.Run("non-regular opened target rejected", func(t *testing.T) {
filePath := filepath.Join(t.TempDir(), "artifact.txt")
if err := os.WriteFile(filePath, []byte("content"), 0o600); err != nil {
t.Fatal(err)
}
directoryInfo, err := os.Stat(t.TempDir())
if err != nil {
t.Fatal(err)
}
fileReader := &fileReader{open: func(path string) (artifactFile, error) {
file, err := os.Open(path)
if err != nil {
return nil, err
}
return &reportedInfoFile{artifactFile: file, info: directoryInfo}, nil
}}
artifact, err := fileReader.Read(context.Background(), domain.ArtifactRef{
Type: domain.ArtifactRefFile,
URI: filePath,
})
if artifact != nil || !errors.Is(err, ErrUnsupportedFile) {
t.Fatalf("artifact=%#v err=%v, want nil/ErrUnsupportedFile", artifact, err)
}
})
t.Run("unknown extension uses text fallback", func(t *testing.T) { t.Run("unknown extension uses text fallback", func(t *testing.T) {
path := filepath.Join(t.TempDir(), "artifact.unknownextension") filePath := filepath.Join(t.TempDir(), "artifact.unknownextension")
if err := os.WriteFile(path, content, 0o600); err != nil { if err := os.WriteFile(filePath, []byte("content"), 0o600); err != nil {
t.Fatal(err) t.Fatal(err)
} }
art, err := reader.Read(ctx, domain.ArtifactRef{ artifact, err := reader.Read(context.Background(), domain.ArtifactRef{
Type: domain.ArtifactRefFile, Type: domain.ArtifactRefFile,
URI: path, URI: filePath,
}) })
if err != nil { if err != nil {
t.Fatalf("unexpected error: %v", err) t.Fatalf("read artifact: %v", err)
} }
if art.ContentType != "text/plain" { if artifact.ContentType != "text/plain" {
t.Errorf("expected text/plain fallback, got %q", art.ContentType) t.Fatalf("content type = %q", artifact.ContentType)
} }
}) })
} }
func TestFileReaderCancelsAfterReadProgress(t *testing.T) {
filePath := filepath.Join(t.TempDir(), "artifact.bin")
content := bytes.Repeat([]byte("x"), fileReadChunkSize*2)
if err := os.WriteFile(filePath, content, 0o600); err != nil {
t.Fatal(err)
}
ctx, cancel := context.WithCancel(context.Background())
var opened *cancelAfterProgressFile
reader := &fileReader{open: func(path string) (artifactFile, error) {
file, err := os.Open(path)
if err != nil {
return nil, err
}
opened = &cancelAfterProgressFile{artifactFile: file, cancel: cancel}
return opened, nil
}}
artifact, err := reader.Read(ctx, domain.ArtifactRef{Type: domain.ArtifactRefFile, URI: filePath})
if artifact != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("artifact=%#v err=%v, want nil/context.Canceled", artifact, err)
}
if opened == nil || opened.reads != 1 {
t.Fatalf("read count = %v, want one progressing read", opened)
}
}
type reportedInfoFile struct {
artifactFile
info os.FileInfo
}
func (f *reportedInfoFile) Stat() (os.FileInfo, error) {
return f.info, nil
}
type cancelAfterProgressFile struct {
artifactFile
cancel context.CancelFunc
reads int
}
func (f *cancelAfterProgressFile) Read(buffer []byte) (int, error) {
n, err := f.artifactFile.Read(buffer)
if n > 0 {
f.reads++
f.cancel()
}
return n, err
}

View File

@@ -5,7 +5,6 @@ package backend
import ( import (
"errors" "errors"
"fmt" "fmt"
"net/url"
"regexp" "regexp"
"sort" "sort"
"strings" "strings"
@@ -19,11 +18,19 @@ const (
// OpenRouterID is the reserved ID of Promptkit's built-in OpenRouter // OpenRouterID is the reserved ID of Promptkit's built-in OpenRouter
// backend. // backend.
OpenRouterID = "openrouter" OpenRouterID = "openrouter"
// RakestrawHomeID is the reserved ID of Promptkit's built-in Rakestrawhome
// backend.
RakestrawHomeID = "rakestrawhome"
openRouterEndpoint = "https://openrouter.ai/api/v1" openRouterEndpoint = "https://openrouter.ai/api/v1"
openRouterAPIKeyEnv = "OPENROUTER_API_KEY" openRouterAPIKeyEnv = "OPENROUTER_API_KEY"
openRouterConcurrencyLimit = 16 openRouterConcurrencyLimit = 16
rakestrawHomeEndpoint = "https://inference.ai.rakestrawhome.com/v1"
rakestrawHomeAPIKeyEnv = "RAKESTRAWHOME_INFERENCE_API_KEY"
rakestrawHomeConcurrencyLimit = 4
defaultQueueCapacity = 1024 defaultQueueCapacity = 1024
) )
@@ -37,20 +44,16 @@ type Registry struct {
backends map[string]domain.Backend backends map[string]domain.Backend
} }
// NewRegistry constructs a registry containing the built-in OpenRouter // NewRegistry constructs a registry containing the built-in definitions
// definition followed by the supplied additions. Every ID must be unique. // followed by the supplied additions. Every ID must be unique.
func NewRegistry(additions []domain.Backend) (*Registry, error) { func NewRegistry(additions []domain.Backend) (*Registry, error) {
builtIns := builtInBackends()
registry := &Registry{ registry := &Registry{
backends: make(map[string]domain.Backend, len(additions)+1), backends: make(map[string]domain.Backend, len(builtIns)+len(additions)),
} }
definitions := make([]domain.Backend, 0, len(additions)+1) definitions := make([]domain.Backend, 0, len(builtIns)+len(additions))
definitions = append(definitions, domain.Backend{ definitions = append(definitions, builtIns...)
ID: OpenRouterID,
Endpoint: openRouterEndpoint,
APIKeyEnv: openRouterAPIKeyEnv,
ConcurrencyLimit: openRouterConcurrencyLimit,
})
definitions = append(definitions, additions...) definitions = append(definitions, additions...)
for _, definition := range definitions { for _, definition := range definitions {
@@ -72,6 +75,23 @@ func NewRegistry(additions []domain.Backend) (*Registry, error) {
return registry, nil return registry, nil
} }
func builtInBackends() []domain.Backend {
return []domain.Backend{
{
ID: OpenRouterID,
Endpoint: openRouterEndpoint,
APIKeyEnv: openRouterAPIKeyEnv,
ConcurrencyLimit: openRouterConcurrencyLimit,
},
{
ID: RakestrawHomeID,
Endpoint: rakestrawHomeEndpoint,
APIKeyEnv: rakestrawHomeAPIKeyEnv,
ConcurrencyLimit: rakestrawHomeConcurrencyLimit,
},
}
}
// GetBackend returns a defensive copy of the backend registered with id. // GetBackend returns a defensive copy of the backend registered with id.
func (r *Registry) GetBackend(id string) (domain.Backend, error) { func (r *Registry) GetBackend(id string) (domain.Backend, error) {
if r == nil { if r == nil {
@@ -109,10 +129,11 @@ func (r *Registry) CapacityPolicies() map[string]domain.BackendCapacityPolicy {
} }
func normalizeBackend(definition domain.Backend) (domain.Backend, error) { func normalizeBackend(definition domain.Backend) (domain.Backend, error) {
definition.Endpoint = strings.TrimSpace(definition.Endpoint) endpoint, err := domain.NormalizeOpenAICompatibleBaseEndpoint(definition.Endpoint)
if err := validateEndpoint(definition.Endpoint); err != nil { if err != nil {
return domain.Backend{}, fmt.Errorf("backend %q endpoint: %w", definition.ID, err) return domain.Backend{}, fmt.Errorf("backend %q endpoint: %w", definition.ID, err)
} }
definition.Endpoint = endpoint
definition.APIKeyEnv = strings.TrimSpace(definition.APIKeyEnv) definition.APIKeyEnv = strings.TrimSpace(definition.APIKeyEnv)
if definition.APIKeyEnv != "" && !environmentVariableName.MatchString(definition.APIKeyEnv) { if definition.APIKeyEnv != "" && !environmentVariableName.MatchString(definition.APIKeyEnv) {
@@ -182,31 +203,3 @@ func normalizeBackend(definition domain.Backend) (domain.Backend, error) {
definition.ExtraParams = extraParams definition.ExtraParams = extraParams
return definition, nil return definition, nil
} }
func validateEndpoint(endpoint string) error {
if endpoint == "" {
return errors.New("must not be blank")
}
if strings.Contains(endpoint, "#") {
return errors.New("must not contain a fragment")
}
parsed, err := url.Parse(endpoint)
if err != nil {
return fmt.Errorf("must be a valid URL: %w", err)
}
scheme := strings.ToLower(parsed.Scheme)
if scheme != "http" && scheme != "https" {
return errors.New("must use http or https")
}
if !parsed.IsAbs() || parsed.Hostname() == "" {
return errors.New("must be absolute and include a host")
}
if parsed.User != nil {
return errors.New("must not contain user information")
}
if parsed.RawQuery != "" || parsed.ForceQuery {
return errors.New("must not contain a query string")
}
return nil
}

View File

@@ -11,32 +11,63 @@ import (
const validEndpoint = "https://backend.example/v1" const validEndpoint = "https://backend.example/v1"
func TestRegistryIncludesExactOpenRouterDefinition(t *testing.T) { func TestRegistryIncludesExactBuiltInDefinitions(t *testing.T) {
registry, err := backend.NewRegistry(nil) registry, err := backend.NewRegistry(nil)
if err != nil { if err != nil {
t.Fatalf("construct registry: %v", err) t.Fatalf("construct registry: %v", err)
} }
definition, err := registry.GetBackend(backend.OpenRouterID) tests := []struct {
if err != nil { name string
t.Fatalf("look up OpenRouter: %v", err) id string
endpoint string
apiKeyEnv string
concurrent int
}{
{
name: "OpenRouter",
id: backend.OpenRouterID,
endpoint: "https://openrouter.ai/api/v1",
apiKeyEnv: "OPENROUTER_API_KEY",
concurrent: 16,
},
{
name: "Rakestrawhome",
id: backend.RakestrawHomeID,
endpoint: "https://inference.ai.rakestrawhome.com/v1",
apiKeyEnv: "RAKESTRAWHOME_INFERENCE_API_KEY",
concurrent: 4,
},
} }
if definition.ID != "openrouter" || for _, tc := range tests {
definition.Endpoint != "https://openrouter.ai/api/v1" || t.Run(tc.name, func(t *testing.T) {
definition.APIKeyEnv != "OPENROUTER_API_KEY" || definition, err := registry.GetBackend(tc.id)
definition.ConcurrencyLimit != 16 || if err != nil {
t.Fatalf("look up built-in: %v", err)
}
if definition.ID != tc.id ||
definition.Endpoint != tc.endpoint ||
definition.APIKeyEnv != tc.apiKeyEnv ||
definition.ConcurrencyLimit != tc.concurrent ||
definition.QueueCapacity != 1024 || definition.QueueCapacity != 1024 ||
!definition.QueueCapacitySet || !definition.QueueCapacitySet ||
definition.ExtraParams != nil { definition.ExtraParams != nil {
t.Fatalf("unexpected OpenRouter definition: %#v", definition) t.Fatalf("unexpected built-in definition: %#v", definition)
} }
})
}
policies := registry.CapacityPolicies() policies := registry.CapacityPolicies()
if len(policies) != 1 || if len(policies) != 2 ||
policies["openrouter"] != (domain.BackendCapacityPolicy{ policies[backend.OpenRouterID] != (domain.BackendCapacityPolicy{
ConcurrencyLimit: 16, ConcurrencyLimit: 16,
QueueCapacity: 1024, QueueCapacity: 1024,
}) ||
policies[backend.RakestrawHomeID] != (domain.BackendCapacityPolicy{
ConcurrencyLimit: 4,
QueueCapacity: 1024,
}) { }) {
t.Fatalf("unexpected OpenRouter capacity policies: %#v", policies) t.Fatalf("unexpected built-in capacity policies: %#v", policies)
} }
} }
@@ -109,11 +140,12 @@ func TestRegistryNormalizesUniqueAdditionsAndIsolatesMutations(t *testing.T) {
} }
policies := registry.CapacityPolicies() policies := registry.CapacityPolicies()
if len(policies) != 2 { if len(policies) != 3 {
t.Fatalf("unexpected capacity policy count: %#v", policies) t.Fatalf("unexpected capacity policy count: %#v", policies)
} }
policies["custom"] = domain.BackendCapacityPolicy{} policies["custom"] = domain.BackendCapacityPolicy{}
delete(policies, backend.OpenRouterID) delete(policies, backend.OpenRouterID)
delete(policies, backend.RakestrawHomeID)
againPolicies := registry.CapacityPolicies() againPolicies := registry.CapacityPolicies()
if againPolicies["custom"] != (domain.BackendCapacityPolicy{ if againPolicies["custom"] != (domain.BackendCapacityPolicy{
ConcurrencyLimit: 3, ConcurrencyLimit: 3,
@@ -122,7 +154,10 @@ func TestRegistryNormalizesUniqueAdditionsAndIsolatesMutations(t *testing.T) {
t.Fatalf("capacity policy map mutated registry state: %#v", againPolicies) t.Fatalf("capacity policy map mutated registry state: %#v", againPolicies)
} }
if _, ok := againPolicies[backend.OpenRouterID]; !ok { if _, ok := againPolicies[backend.OpenRouterID]; !ok {
t.Fatalf("capacity policy deletion mutated registry state: %#v", againPolicies) t.Fatalf("OpenRouter capacity policy deletion mutated registry state: %#v", againPolicies)
}
if _, ok := againPolicies[backend.RakestrawHomeID]; !ok {
t.Fatalf("Rakestrawhome capacity policy deletion mutated registry state: %#v", againPolicies)
} }
} }
@@ -242,11 +277,14 @@ func TestNewRegistryRejectsDuplicateIDs(t *testing.T) {
wantID string wantID string
}{ }{
{ {
name: "built-in collision after normalization", name: "OpenRouter collision after normalization",
additions: []domain.Backend{{ additions: []domain.Backend{{ID: " openrouter "}},
ID: " openrouter ", wantID: backend.OpenRouterID,
}}, },
wantID: "openrouter", {
name: "Rakestrawhome collision after normalization",
additions: []domain.Backend{{ID: " rakestrawhome "}},
wantID: backend.RakestrawHomeID,
}, },
{ {
name: "consumer collision after normalization", name: "consumer collision after normalization",

View File

@@ -14,21 +14,12 @@ const (
ContentTypeApplicationJSON = "application/json" ContentTypeApplicationJSON = "application/json"
OpenAIChatCompletionsPath = "/chat/completions" OpenAIChatCompletionsPath = "/chat/completions"
ExecutionDefaultTemperature = 0.0
ExecutionDefaultMaxTokens = 0
ExecutionDefaultTopP = 1.0
ExecutionDefaultTimeoutSeconds = 600 ExecutionDefaultTimeoutSeconds = 600
)
var (
LLMRequestTimeoutDefault = 10 * time.Minute LLMRequestTimeoutDefault = 10 * time.Minute
) )
func ExecutionTargetDefault() domain.ExecutionTarget { func ExecutionTargetDefault() domain.ExecutionTarget {
return domain.ExecutionTarget{ return domain.ExecutionTarget{
Temperature: ExecutionDefaultTemperature,
MaxTokens: ExecutionDefaultMaxTokens,
TopP: ExecutionDefaultTopP,
TimeoutSeconds: ExecutionDefaultTimeoutSeconds, TimeoutSeconds: ExecutionDefaultTimeoutSeconds,
} }
} }

View File

@@ -97,22 +97,22 @@ type RunResult struct {
// PreparedRun contains pre-LLM execution state from the prepare/render phase. // PreparedRun contains pre-LLM execution state from the prepare/render phase.
// It must never include resolved API key values, model output, or validation data. // It must never include resolved API key values, model output, or validation data.
type PreparedRun struct { type PreparedRun struct {
PromptID string `json:"prompt_id"` PromptID string
PromptVersion string `json:"prompt_version,omitempty"` PromptVersion string
PromptHash string `json:"prompt_hash,omitempty"` PromptHash string
SelectedProfileID string `json:"selected_profile_id"` SelectedProfileID string
SelectedBackendID string `json:"selected_backend_id,omitempty"` SelectedBackendID string
EffectiveModelParams ExecutionTarget `json:"effective_model_params"` EffectiveModelParams ExecutionTarget
TargetPresence ExecutionTargetPresence `json:"-"` TargetPresence ExecutionTargetPresence
OutputContract OutputContract `json:"output_contract"` OutputContract OutputContract
StructuredOutput *StructuredOutputSpec `json:"structured_output,omitempty"` StructuredOutput *StructuredOutputSpec
InputHashes map[string]string `json:"input_hashes,omitempty"` InputHashes map[string]string
SessionID string `json:"session_id,omitempty"` SessionID string
RenderedPromptHash string `json:"rendered_prompt_hash"` RenderedPromptHash string
Messages []RenderedMessage `json:"messages"` Messages []RenderedMessage
StartTime time.Time `json:"start_time,omitempty"` StartTime time.Time
EndTime time.Time `json:"end_time,omitempty"` EndTime time.Time
DurationMS int64 `json:"duration_ms,omitempty"` DurationMS int64
} }
// ArtifactRef represents a reference to an input artifact. // ArtifactRef represents a reference to an input artifact.
@@ -192,6 +192,7 @@ type BackendCapacityPolicy struct {
// ExecutionProfile describes how and where to execute a model. // ExecutionProfile describes how and where to execute a model.
type ExecutionProfile struct { type ExecutionProfile struct {
ID string `yaml:"id"` ID string `yaml:"id"`
BaseProfileID string `yaml:"base_profile"`
BackendID string `yaml:"backend"` BackendID string `yaml:"backend"`
Endpoint string `yaml:"endpoint"` Endpoint string `yaml:"endpoint"`
Model string `yaml:"model"` Model string `yaml:"model"`

View File

@@ -0,0 +1,39 @@
package domain
import (
"errors"
"net/url"
"strings"
)
// NormalizeOpenAICompatibleBaseEndpoint trims and validates a source-neutral
// OpenAI-compatible provider base endpoint.
func NormalizeOpenAICompatibleBaseEndpoint(endpoint string) (string, error) {
endpoint = strings.TrimSpace(endpoint)
if endpoint == "" {
return "", errors.New("endpoint must not be blank")
}
if strings.Contains(endpoint, "#") {
return "", errors.New("endpoint must not contain a fragment")
}
parsed, err := url.Parse(endpoint)
if err != nil {
return "", errors.New("endpoint must be a valid URL")
}
parsed.Scheme = strings.ToLower(parsed.Scheme)
if parsed.Scheme != "http" && parsed.Scheme != "https" {
return "", errors.New("endpoint must use http or https")
}
if !parsed.IsAbs() || parsed.Hostname() == "" {
return "", errors.New("endpoint must be absolute and include a host")
}
if parsed.User != nil {
return "", errors.New("endpoint must not contain user information")
}
if parsed.RawQuery != "" || parsed.ForceQuery {
return "", errors.New("endpoint must not contain a query string")
}
return parsed.String(), nil
}

View File

@@ -0,0 +1,47 @@
package domain
import "testing"
func TestNormalizeOpenAICompatibleBaseEndpoint(t *testing.T) {
tests := []struct {
name string
endpoint string
want string
wantErr bool
}{
{name: "http host", endpoint: "http://provider.example", want: "http://provider.example"},
{name: "https nested path and whitespace", endpoint: " HTTPS://provider.example/api/openai/v1 ", want: "https://provider.example/api/openai/v1"},
{name: "IPv4 host and port", endpoint: "http://127.0.0.1:8080/v1", want: "http://127.0.0.1:8080/v1"},
{name: "IPv6 host and port", endpoint: "https://[::1]:8443/v1", want: "https://[::1]:8443/v1"},
{name: "repeated trailing slashes", endpoint: "https://provider.example/v1///", want: "https://provider.example/v1///"},
{name: "blank", endpoint: " \t\n ", wantErr: true},
{name: "relative path", endpoint: "/api/v1", wantErr: true},
{name: "scheme relative", endpoint: "//provider.example/v1", wantErr: true},
{name: "missing host", endpoint: "https:///v1", wantErr: true},
{name: "unsupported scheme", endpoint: "ftp://provider.example/v1", wantErr: true},
{name: "user information", endpoint: "https://user:secret@provider.example/v1", wantErr: true},
{name: "query", endpoint: "https://provider.example/v1?mode=chat", wantErr: true},
{name: "empty query", endpoint: "https://provider.example/v1?", wantErr: true},
{name: "fragment", endpoint: "https://provider.example/v1#chat", wantErr: true},
{name: "empty fragment", endpoint: "https://provider.example/v1#", wantErr: true},
{name: "malformed URL", endpoint: "https://provider.example/%zz", wantErr: true},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
got, err := NormalizeOpenAICompatibleBaseEndpoint(tc.endpoint)
if tc.wantErr {
if err == nil {
t.Fatalf("expected endpoint error, got %q", got)
}
return
}
if err != nil {
t.Fatalf("normalize endpoint: %v", err)
}
if got != tc.want {
t.Fatalf("normalized endpoint = %q, want %q", got, tc.want)
}
})
}
}

View File

@@ -0,0 +1,34 @@
package domain
import (
"errors"
"math"
"time"
)
const maxExecutionTimeoutSeconds int64 = math.MaxInt64 / int64(time.Second)
// ValidateExecutionTargetSettings validates source-neutral execution-setting
// invariants on a resolved target.
func ValidateExecutionTargetSettings(target ExecutionTarget) error {
if !isFinite(target.Temperature) || target.Temperature < 0 || target.Temperature > 2 {
return errors.New("temperature must be finite and between 0 and 2")
}
if target.MaxTokens < 0 {
return errors.New("max_tokens must be greater than or equal to 0")
}
if !isFinite(target.TopP) || target.TopP < 0 || target.TopP > 1 {
return errors.New("top_p must be finite and between 0 and 1")
}
if target.TimeoutSeconds < 0 {
return errors.New("timeout_seconds must be greater than or equal to 0")
}
if int64(target.TimeoutSeconds) > maxExecutionTimeoutSeconds {
return errors.New("timeout_seconds exceeds the maximum supported duration")
}
return nil
}
func isFinite(value float64) bool {
return !math.IsNaN(value) && !math.IsInf(value, 0)
}

View File

@@ -0,0 +1,73 @@
package domain
import (
"math"
"strconv"
"strings"
"testing"
)
func TestValidateExecutionTargetSettings(t *testing.T) {
valid := ExecutionTarget{
Temperature: 1,
MaxTokens: 1,
TopP: 0.5,
TimeoutSeconds: 1,
}
type testCase struct {
name string
change func(*ExecutionTarget)
wantErr string
}
tests := []testCase{
{name: "temperature lower boundary", change: func(v *ExecutionTarget) { v.Temperature = 0 }},
{name: "temperature finite lower neighbor", change: func(v *ExecutionTarget) { v.Temperature = math.Nextafter(0, 1) }},
{name: "temperature finite upper neighbor", change: func(v *ExecutionTarget) { v.Temperature = math.Nextafter(2, 0) }},
{name: "temperature upper boundary", change: func(v *ExecutionTarget) { v.Temperature = 2 }},
{name: "temperature below lower boundary", change: func(v *ExecutionTarget) { v.Temperature = math.Nextafter(0, math.Inf(-1)) }, wantErr: "temperature"},
{name: "temperature above upper boundary", change: func(v *ExecutionTarget) { v.Temperature = math.Nextafter(2, math.Inf(1)) }, wantErr: "temperature"},
{name: "temperature NaN", change: func(v *ExecutionTarget) { v.Temperature = math.NaN() }, wantErr: "temperature"},
{name: "temperature positive infinity", change: func(v *ExecutionTarget) { v.Temperature = math.Inf(1) }, wantErr: "temperature"},
{name: "temperature negative infinity", change: func(v *ExecutionTarget) { v.Temperature = math.Inf(-1) }, wantErr: "temperature"},
{name: "max tokens lower boundary", change: func(v *ExecutionTarget) { v.MaxTokens = 0 }},
{name: "max tokens finite neighbor", change: func(v *ExecutionTarget) { v.MaxTokens = 1 }},
{name: "max tokens below lower boundary", change: func(v *ExecutionTarget) { v.MaxTokens = -1 }, wantErr: "max_tokens"},
{name: "top p lower boundary", change: func(v *ExecutionTarget) { v.TopP = 0 }},
{name: "top p finite lower neighbor", change: func(v *ExecutionTarget) { v.TopP = math.Nextafter(0, 1) }},
{name: "top p finite upper neighbor", change: func(v *ExecutionTarget) { v.TopP = math.Nextafter(1, 0) }},
{name: "top p upper boundary", change: func(v *ExecutionTarget) { v.TopP = 1 }},
{name: "top p below lower boundary", change: func(v *ExecutionTarget) { v.TopP = math.Nextafter(0, math.Inf(-1)) }, wantErr: "top_p"},
{name: "top p above upper boundary", change: func(v *ExecutionTarget) { v.TopP = math.Nextafter(1, math.Inf(1)) }, wantErr: "top_p"},
{name: "top p NaN", change: func(v *ExecutionTarget) { v.TopP = math.NaN() }, wantErr: "top_p"},
{name: "top p positive infinity", change: func(v *ExecutionTarget) { v.TopP = math.Inf(1) }, wantErr: "top_p"},
{name: "top p negative infinity", change: func(v *ExecutionTarget) { v.TopP = math.Inf(-1) }, wantErr: "top_p"},
{name: "timeout lower boundary", change: func(v *ExecutionTarget) { v.TimeoutSeconds = 0 }},
{name: "timeout finite neighbor", change: func(v *ExecutionTarget) { v.TimeoutSeconds = 1 }},
{name: "timeout below lower boundary", change: func(v *ExecutionTarget) { v.TimeoutSeconds = -1 }, wantErr: "timeout_seconds"},
}
if strconv.IntSize == 64 {
durationLimit := maxExecutionTimeoutSeconds
tests = append(tests,
testCase{name: "timeout duration boundary", change: func(v *ExecutionTarget) { v.TimeoutSeconds = int(durationLimit) }},
testCase{name: "timeout above duration boundary", change: func(v *ExecutionTarget) { v.TimeoutSeconds = int(durationLimit) + 1 }, wantErr: "timeout_seconds"},
)
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
target := valid
tt.change(&target)
err := ValidateExecutionTargetSettings(target)
if tt.wantErr == "" {
if err != nil {
t.Fatalf("validate execution settings: %v", err)
}
return
}
if err == nil || !strings.Contains(err.Error(), tt.wantErr) {
t.Fatalf("error = %v, want diagnostic containing %q", err, tt.wantErr)
}
})
}
}

View File

@@ -0,0 +1,38 @@
package domain
import (
"errors"
"fmt"
"strings"
)
const maxOutputRepairAttempts = 3
// ValidateOutputContract validates source-neutral output-contract invariants.
func ValidateOutputContract(contract OutputContract) error {
switch contract.Format {
case FormatText, FormatMarkdown, FormatJSON:
default:
return fmt.Errorf("invalid output format: %q", contract.Format)
}
switch contract.ValidationMode {
case ValidationNone, ValidationBasic, ValidationJSON, ValidationJSONSchema:
default:
return fmt.Errorf("invalid validation mode: %q", contract.ValidationMode)
}
if contract.ValidationMode == ValidationJSONSchema && strings.TrimSpace(contract.SchemaPath) == "" {
return errors.New("schema_path is required when validation_mode is json_schema")
}
if contract.RepairAttempts < 0 {
return errors.New("repair_attempts must be greater than or equal to 0")
}
if contract.RepairAttempts > maxOutputRepairAttempts {
return fmt.Errorf("repair_attempts must be less than or equal to %d", maxOutputRepairAttempts)
}
if contract.ValidationMode == ValidationNone && contract.RepairAttempts > 0 {
return errors.New("repair_attempts requires basic, json, or json_schema validation")
}
return nil
}

View File

@@ -0,0 +1,87 @@
package domain
import (
"strings"
"testing"
)
func TestValidateOutputContract(t *testing.T) {
valid := OutputContract{
Format: FormatText,
ValidationMode: ValidationNone,
}
tests := []struct {
name string
change func(*OutputContract)
wantErr string
}{
{name: "text format", change: func(c *OutputContract) { c.Format = FormatText }},
{name: "markdown format", change: func(c *OutputContract) { c.Format = FormatMarkdown }},
{name: "json format", change: func(c *OutputContract) { c.Format = FormatJSON }},
{name: "empty format", change: func(c *OutputContract) { c.Format = "" }, wantErr: "format"},
{name: "unsupported format", change: func(c *OutputContract) { c.Format = OutputFormat("binary") }, wantErr: "format"},
{name: "none validation", change: func(c *OutputContract) { c.ValidationMode = ValidationNone }},
{name: "basic validation", change: func(c *OutputContract) { c.ValidationMode = ValidationBasic }},
{name: "json validation", change: func(c *OutputContract) { c.ValidationMode = ValidationJSON }},
{name: "json schema validation", change: func(c *OutputContract) {
c.ValidationMode = ValidationJSONSchema
c.SchemaPath = "schema.json"
}},
{name: "empty validation mode", change: func(c *OutputContract) { c.ValidationMode = "" }, wantErr: "validation mode"},
{name: "unsupported validation mode", change: func(c *OutputContract) { c.ValidationMode = ValidationMode("unknown") }, wantErr: "validation mode"},
{name: "negative repair attempts", change: func(c *OutputContract) {
c.ValidationMode = ValidationBasic
c.RepairAttempts = -1
}, wantErr: "repair_attempts"},
{name: "zero repair attempts", change: func(c *OutputContract) { c.RepairAttempts = 0 }},
{name: "one repair attempt", change: func(c *OutputContract) {
c.ValidationMode = ValidationBasic
c.RepairAttempts = 1
}},
{name: "maximum repair attempts", change: func(c *OutputContract) {
c.ValidationMode = ValidationJSON
c.RepairAttempts = 3
}},
{name: "too many repair attempts", change: func(c *OutputContract) {
c.ValidationMode = ValidationJSONSchema
c.SchemaPath = "schema.json"
c.RepairAttempts = 4
}, wantErr: "repair_attempts"},
{name: "none validation with repair attempts", change: func(c *OutputContract) {
c.ValidationMode = ValidationNone
c.RepairAttempts = 1
}, wantErr: "repair_attempts"},
{name: "json schema empty path", change: func(c *OutputContract) {
c.ValidationMode = ValidationJSONSchema
c.SchemaPath = ""
}, wantErr: "schema_path"},
{name: "json schema whitespace path", change: func(c *OutputContract) {
c.ValidationMode = ValidationJSONSchema
c.SchemaPath = " \t "
}, wantErr: "schema_path"},
{name: "json schema nonblank path", change: func(c *OutputContract) {
c.ValidationMode = ValidationJSONSchema
c.SchemaPath = " schema.json "
}},
{name: "non-schema empty path", change: func(c *OutputContract) { c.SchemaPath = "" }},
{name: "non-schema populated path", change: func(c *OutputContract) { c.SchemaPath = "ignored.json" }},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
contract := valid
tt.change(&contract)
err := ValidateOutputContract(contract)
if tt.wantErr == "" {
if err != nil {
t.Fatalf("validate output contract: %v", err)
}
return
}
if err == nil || !strings.Contains(err.Error(), tt.wantErr) {
t.Fatalf("error = %v, want diagnostic containing %q", err, tt.wantErr)
}
})
}
}

View File

@@ -1,141 +0,0 @@
package domain
import (
"encoding/json"
"strings"
"testing"
)
func TestPreparedRunJSONDoesNotIncludeSecretValues(t *testing.T) {
const envName = "PROMPTKIT_TEST_API_KEY"
const secret = "super-secret-value"
t.Setenv(envName, secret)
prepared := PreparedRun{
PromptID: "prompt.id",
PromptVersion: "v1",
PromptHash: "prompt-hash",
SelectedProfileID: "local-fast",
EffectiveModelParams: ExecutionTarget{
Endpoint: "http://llm/v1",
Model: "gpt-test",
APIKeyEnv: envName,
APIKey: secret,
},
InputHashes: map[string]string{"transcript": "hash-1"},
RenderedPromptHash: "rendered-hash",
Messages: []RenderedMessage{
{Role: "system", Content: "You are helpful."},
{Role: "user", Content: "Summarize this."},
},
}
b, err := json.Marshal(prepared)
if err != nil {
t.Fatalf("marshal failed: %v", err)
}
out := string(b)
if strings.Contains(out, secret) {
t.Fatalf("prepared run JSON unexpectedly contains secret value: %s", out)
}
if !strings.Contains(out, `"api_key_env":"`+envName+`"`) {
t.Fatalf("prepared run JSON should include api_key_env name: %s", out)
}
var top map[string]any
if err := json.Unmarshal(b, &top); err != nil {
t.Fatalf("unmarshal failed: %v", err)
}
for _, forbidden := range []string{"raw_output", "validation", "artifact"} {
if _, ok := top[forbidden]; ok {
t.Fatalf("prepared run JSON should not include %q", forbidden)
}
}
}
func TestPreparedRunJSONIncludesMessageCacheControlOnlyWhenPresent(t *testing.T) {
prepared := PreparedRun{
PromptID: "prompt.id",
SelectedProfileID: "local-fast",
EffectiveModelParams: ExecutionTarget{
Endpoint: "http://llm/v1",
Model: "gpt-test",
},
RenderedPromptHash: "rendered-hash",
Messages: []RenderedMessage{
{
Role: "system",
Content: "You are helpful.",
CacheControl: &CacheControl{
Type: CacheControlEphemeral,
TTL: "1h",
},
},
{Role: "user", Content: "Summarize this."},
},
}
b, err := json.Marshal(prepared)
if err != nil {
t.Fatalf("marshal failed: %v", err)
}
var decoded struct {
Messages []map[string]any `json:"messages"`
}
if err := json.Unmarshal(b, &decoded); err != nil {
t.Fatalf("unmarshal failed: %v", err)
}
if len(decoded.Messages) != 2 {
t.Fatalf("expected 2 messages, got %d", len(decoded.Messages))
}
cacheControl, ok := decoded.Messages[0]["cache_control"].(map[string]any)
if !ok {
t.Fatalf("expected cache_control on first message, got %#v", decoded.Messages[0])
}
if cacheControl["type"] != string(CacheControlEphemeral) || cacheControl["ttl"] != "1h" {
t.Fatalf("unexpected cache_control payload: %#v", cacheControl)
}
if _, ok := decoded.Messages[1]["cache_control"]; ok {
t.Fatalf("expected second message to omit cache_control, got %#v", decoded.Messages[1])
}
}
func TestPreparedRunJSONIncludesSessionIDOnlyWhenPresent(t *testing.T) {
prepared := PreparedRun{
PromptID: "prompt.id",
SelectedProfileID: "local-fast",
EffectiveModelParams: ExecutionTarget{
Endpoint: "http://llm/v1",
Model: "gpt-test",
},
SessionID: "session-123",
RenderedPromptHash: "rendered-hash",
Messages: []RenderedMessage{{Role: "user", Content: "Summarize this."}},
}
b, err := json.Marshal(prepared)
if err != nil {
t.Fatalf("marshal failed: %v", err)
}
var decoded map[string]any
if err := json.Unmarshal(b, &decoded); err != nil {
t.Fatalf("unmarshal failed: %v", err)
}
if decoded["session_id"] != "session-123" {
t.Fatalf("expected session_id in prepared run JSON, got %#v", decoded["session_id"])
}
prepared.SessionID = ""
b, err = json.Marshal(prepared)
if err != nil {
t.Fatalf("marshal failed: %v", err)
}
if strings.Contains(string(b), "session_id") {
t.Fatalf("expected empty session_id to be omitted, got %s", b)
}
}

View File

@@ -8,6 +8,9 @@ import (
// NormalizeSessionID applies the shared session identifier rule. // NormalizeSessionID applies the shared session identifier rule.
func NormalizeSessionID(raw string) (string, error) { func NormalizeSessionID(raw string) (string, error) {
if !utf8.ValidString(raw) {
return "", fmt.Errorf("session_id must contain valid UTF-8")
}
normalized := strings.TrimSpace(raw) normalized := strings.TrimSpace(raw)
if normalized == "" { if normalized == "" {
return "", nil return "", nil

View File

@@ -10,7 +10,7 @@ func TestNormalizeSessionID(t *testing.T) {
name string name string
raw string raw string
want string want string
wantErr bool wantErrContains string
}{ }{
{ {
name: "trims surrounding Unicode whitespace", name: "trims surrounding Unicode whitespace",
@@ -30,19 +30,22 @@ func TestNormalizeSessionID(t *testing.T) {
{ {
name: "one Unicode code point over maximum is rejected", name: "one Unicode code point over maximum is rejected",
raw: strings.Repeat("界", SessionIDMaxLength+1), raw: strings.Repeat("界", SessionIDMaxLength+1),
wantErr: true, wantErrContains: "exceeds maximum",
}, },
{name: "invalid UTF-8 before valid content", raw: string([]byte{0xff}) + "session", wantErrContains: "valid UTF-8"},
{name: "invalid UTF-8 within valid content", raw: "ses" + string([]byte{0xff}) + "sion", wantErrContains: "valid UTF-8"},
{name: "invalid UTF-8 after valid content", raw: "session" + string([]byte{0xff}), wantErrContains: "valid UTF-8"},
} }
for _, tt := range tests { for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) { t.Run(tt.name, func(t *testing.T) {
got, err := NormalizeSessionID(tt.raw) got, err := NormalizeSessionID(tt.raw)
if tt.wantErr { if tt.wantErrContains != "" {
if err == nil { if err == nil {
t.Fatal("expected normalization error") t.Fatal("expected normalization error")
} }
if !strings.Contains(err.Error(), "exceeds maximum") { if !strings.Contains(err.Error(), tt.wantErrContains) {
t.Fatalf("expected useful length diagnostic, got %v", err) t.Fatalf("expected diagnostic containing %q, got %v", tt.wantErrContains, err)
} }
return return
} }

View File

@@ -36,7 +36,8 @@ func FindYAMLFiles(ctx context.Context, root string) ([]string, error) {
return files, err return files, err
} }
// FindFSYAMLFiles returns sorted paths for .yaml and .yml files under root in fsys. // FindFSYAMLFiles returns root itself when it names a file. For a directory
// root, it returns sorted paths for .yaml and .yml files beneath that root.
func FindFSYAMLFiles(ctx context.Context, fsys fs.FS, root string) ([]string, error) { func FindFSYAMLFiles(ctx context.Context, fsys fs.FS, root string) ([]string, error) {
cleanRoot := CleanFSRoot(root) cleanRoot := CleanFSRoot(root)
var files []string var files []string
@@ -52,6 +53,10 @@ func FindFSYAMLFiles(ctx context.Context, fsys fs.FS, root string) ([]string, er
if d.IsDir() { if d.IsDir() {
return nil return nil
} }
if name == cleanRoot {
files = append(files, name)
return nil
}
if !IsYAMLFile(d.Name()) { if !IsYAMLFile(d.Name()) {
return nil return nil
} }
@@ -71,10 +76,10 @@ func RelativePath(root string, filePath string) string {
return filepath.Clean(rel) return filepath.Clean(rel)
} }
// CleanFSRoot normalizes a root path for use with fs.FS. // CleanFSRoot normalizes a root path for use with fs.FS while preserving
// nonblank leading and trailing whitespace.
func CleanFSRoot(root string) string { func CleanFSRoot(root string) string {
root = strings.TrimSpace(root) if strings.TrimSpace(root) == "" || root == "." {
if root == "" || root == "." {
return "." return "."
} }
return path.Clean(root) return path.Clean(root)
@@ -97,22 +102,21 @@ func DisplayPath(root string, name string) string {
// ResolveFSPath resolves userPath from baseDir and keeps it inside root. // ResolveFSPath resolves userPath from baseDir and keeps it inside root.
func ResolveFSPath(root string, baseDir string, userPath string) (string, string, error) { func ResolveFSPath(root string, baseDir string, userPath string) (string, string, error) {
cleanRoot := CleanFSRoot(root) cleanRoot := CleanFSRoot(root)
cleanBase := path.Clean(strings.TrimSpace(baseDir)) cleanBase := path.Clean(baseDir)
if cleanBase == "" { if strings.TrimSpace(baseDir) == "" {
cleanBase = cleanRoot cleanBase = cleanRoot
} }
if !containsFSPath(cleanRoot, cleanBase) { if !containsFSPath(cleanRoot, cleanBase) {
return "", "", fmt.Errorf("base path %q is outside source root %q", cleanBase, cleanRoot) return "", "", fmt.Errorf("base path %q is outside source root %q", cleanBase, cleanRoot)
} }
cleanUserPath := strings.TrimSpace(userPath) if strings.TrimSpace(userPath) == "" {
if cleanUserPath == "" {
return "", "", fmt.Errorf("path is required") return "", "", fmt.Errorf("path is required")
} }
cleanUserPath = path.Clean(cleanUserPath) if path.IsAbs(userPath) {
if path.IsAbs(cleanUserPath) {
return "", "", fmt.Errorf("path %q must be relative", userPath) return "", "", fmt.Errorf("path %q must be relative", userPath)
} }
cleanUserPath := path.Clean(userPath)
resolved := path.Clean(path.Join(cleanBase, cleanUserPath)) resolved := path.Clean(path.Join(cleanBase, cleanUserPath))
if !containsFSPath(cleanRoot, resolved) { if !containsFSPath(cleanRoot, resolved) {
@@ -130,13 +134,6 @@ func containsFSPath(root string, name string) bool {
return name == root || strings.HasPrefix(name, strings.TrimSuffix(root, "/")+"/") return name == root || strings.HasPrefix(name, strings.TrimSuffix(root, "/")+"/")
} }
// Stem strips .yaml or .yml from a file name.
func Stem(name string) string {
name = strings.TrimSuffix(name, ".yaml")
name = strings.TrimSuffix(name, ".yml")
return name
}
func IsYAMLFile(name string) bool { func IsYAMLFile(name string) bool {
return strings.HasSuffix(name, ".yaml") || strings.HasSuffix(name, ".yml") return strings.HasSuffix(name, ".yaml") || strings.HasSuffix(name, ".yml")
} }

View File

@@ -54,7 +54,7 @@ func TestFindFSYAMLFilesNestedSortedAndFiltered(t *testing.T) {
"other/ignored.yaml": &fstest.MapFile{Data: []byte("id: ignored")}, "other/ignored.yaml": &fstest.MapFile{Data: []byte("id: ignored")},
} }
got, err := FindFSYAMLFiles(context.Background(), fsys, " prompts ") got, err := FindFSYAMLFiles(context.Background(), fsys, "prompts")
if err != nil { if err != nil {
t.Fatalf("expected no error, got %v", err) t.Fatalf("expected no error, got %v", err)
} }
@@ -98,8 +98,11 @@ func TestCleanFSRoot(t *testing.T) {
want string want string
}{ }{
{name: "empty", root: "", want: "."}, {name: "empty", root: "", want: "."},
{name: "whitespace only", root: " \t ", want: "."},
{name: "dot", root: ".", want: "."}, {name: "dot", root: ".", want: "."},
{name: "trimmed", root: " prompts/../profiles ", want: "profiles"}, {name: "cleaned", root: "prompts/../profiles", want: "profiles"},
{name: "leading whitespace preserved", root: " profiles", want: " profiles"},
{name: "trailing whitespace preserved", root: "profiles ", want: "profiles "},
} }
for _, tc := range tests { for _, tc := range tests {
@@ -158,6 +161,22 @@ func TestResolveFSPath(t *testing.T) {
wantPath: "prompts/shared/user.tmpl", wantPath: "prompts/shared/user.tmpl",
wantDisplay: "shared/user.tmpl", wantDisplay: "shared/user.tmpl",
}, },
{
name: "leading whitespace preserved",
root: "prompts",
baseDir: "prompts/nested",
userPath: " user.tmpl",
wantPath: "prompts/nested/ user.tmpl",
wantDisplay: "nested/ user.tmpl",
},
{
name: "trailing whitespace preserved",
root: "prompts",
baseDir: "prompts/nested",
userPath: "user.tmpl ",
wantPath: "prompts/nested/user.tmpl ",
wantDisplay: "nested/user.tmpl ",
},
{ {
name: "escape rejected", name: "escape rejected",
root: "prompts", root: "prompts",
@@ -218,26 +237,6 @@ func TestResolveFSPath(t *testing.T) {
} }
} }
func TestStemStripsYAMLExtensions(t *testing.T) {
tests := []struct {
name string
in string
want string
}{
{name: "yaml", in: "prompt.yaml", want: "prompt"},
{name: "yml", in: "profile.yml", want: "profile"},
{name: "other", in: "file.txt", want: "file.txt"},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
if got := Stem(tc.in); got != tc.want {
t.Fatalf("expected %q, got %q", tc.want, got)
}
})
}
}
func TestIsYAMLFile(t *testing.T) { func TestIsYAMLFile(t *testing.T) {
tests := []struct { tests := []struct {
name string name string

View File

@@ -1,5 +1,5 @@
// Package jsonvalue validates and defensively copies JSON-compatible value // Package jsonvalue validates and defensively copies bounded JSON-compatible
// trees used by configuration, request, and prepared-state boundaries. // value trees used by configuration, request, and prepared-state boundaries.
package jsonvalue package jsonvalue
import ( import (
@@ -8,29 +8,39 @@ import (
"math" "math"
"reflect" "reflect"
"sort" "sort"
"strconv"
) )
const maxSafeJSONInteger = 1<<53 - 1 const (
maxContainerDepth = 100
maxProducedNodes = 100_000
)
type visit struct { type visit struct {
typ reflect.Type typ reflect.Type
ptr uintptr ptr uintptr
} }
type traversalState struct {
active map[visit]struct{}
producedNodes int
}
// Copy validates and deeply copies a JSON-compatible value while preserving // Copy validates and deeply copies a JSON-compatible value while preserving
// compatible concrete map, slice, array, scalar, and number types. // compatible concrete map, slice, array, scalar, and number types. It rejects
// cycles and values that exceed the package's traversal limits.
func Copy(src any) (any, error) { func Copy(src any) (any, error) {
return copyValue(reflect.ValueOf(src), "value", make(map[visit]struct{}), true) return copyValue(reflect.ValueOf(src), "value", newTraversalState(), true, 0)
} }
// CopyMap validates and deeply copies an extra-parameter map while preserving // CopyMap validates and deeply copies an extra-parameter map while preserving
// compatible concrete map, slice, array, scalar, and number types. // compatible concrete map, slice, array, scalar, and number types. It rejects
// empty object keys, cycles, and values that exceed the package's traversal
// limits.
func CopyMap(src map[string]any) (map[string]any, error) { func CopyMap(src map[string]any) (map[string]any, error) {
if src == nil { if src == nil {
return nil, nil return nil, nil
} }
copied, err := copyValue(reflect.ValueOf(src), "extra_params", make(map[visit]struct{}), false) copied, err := copyValue(reflect.ValueOf(src), "extra_params", newTraversalState(), false, 0)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -44,71 +54,90 @@ func CopyMap(src map[string]any) (map[string]any, error) {
func copyValue( func copyValue(
value reflect.Value, value reflect.Value,
path string, path string,
seen map[visit]struct{}, state *traversalState,
allowEmptyMapKeys bool, allowEmptyMapKeys bool,
containerDepth int,
) (any, error) { ) (any, error) {
if !value.IsValid() { resolved, cleanup, isNull, err := state.resolveIndirection(value, path)
if err != nil {
return nil, err
}
defer cleanup()
if isNull {
if err := state.produceNode(path); err != nil {
return nil, err
}
return nil, nil return nil, nil
} }
if value.Kind() == reflect.Interface { value = resolved
if value.IsNil() {
return nil, nil
}
return copyValue(value.Elem(), path, seen, allowEmptyMapKeys)
}
if !value.CanInterface() { if !value.CanInterface() {
return nil, fmt.Errorf("%s: value cannot be copied", path) return nil, fmt.Errorf("%s: value cannot be copied", path)
} }
if number, ok := value.Interface().(json.Number); ok { if number, ok := value.Interface().(json.Number); ok {
if _, err := json.Marshal(number); err != nil { if !validJSONNumber(number) {
return nil, fmt.Errorf("%s: invalid JSON number", path) return nil, fmt.Errorf("%s: invalid JSON number", path)
} }
f, err := strconv.ParseFloat(number.String(), 64) if err := state.produceNode(path); err != nil {
if err != nil || math.IsNaN(f) || math.IsInf(f, 0) { return nil, err
return nil, fmt.Errorf("%s: invalid JSON number", path)
} }
return number, nil return number, nil
} }
switch value.Kind() { switch value.Kind() {
case reflect.Bool, reflect.String: case reflect.Bool, reflect.String:
if err := state.produceNode(path); err != nil {
return nil, err
}
return value.Interface(), nil return value.Interface(), nil
case reflect.Int, reflect.Int8, reflect.Int16, reflect.Int32, reflect.Int64: case reflect.Int, reflect.Int8, reflect.Int16, reflect.Int32, reflect.Int64:
if value.Int() < -maxSafeJSONInteger || value.Int() > maxSafeJSONInteger { if err := state.produceNode(path); err != nil {
return nil, fmt.Errorf("%s: integer is outside the JSON-safe range", path) return nil, err
} }
return value.Interface(), nil return value.Interface(), nil
case reflect.Uint, reflect.Uint8, reflect.Uint16, reflect.Uint32, reflect.Uint64, reflect.Uintptr: case reflect.Uint, reflect.Uint8, reflect.Uint16, reflect.Uint32, reflect.Uint64, reflect.Uintptr:
if value.Uint() > maxSafeJSONInteger { if err := state.produceNode(path); err != nil {
return nil, fmt.Errorf("%s: integer is outside the JSON-safe range", path) return nil, err
} }
return value.Interface(), nil return value.Interface(), nil
case reflect.Float32, reflect.Float64: case reflect.Float32, reflect.Float64:
number := value.Convert(reflect.TypeOf(float64(0))).Float() number := value.Float()
if math.IsNaN(number) || math.IsInf(number, 0) { if math.IsNaN(number) || math.IsInf(number, 0) {
return nil, fmt.Errorf("%s: floating-point value must be finite", path) return nil, fmt.Errorf("%s: floating-point value must be finite", path)
} }
if err := state.produceNode(path); err != nil {
return nil, err
}
return value.Interface(), nil return value.Interface(), nil
case reflect.Pointer: case reflect.Map:
if value.IsNil() { if value.IsNil() {
if err := state.produceNode(path); err != nil {
return nil, err
}
return nil, nil return nil, nil
} }
current := visit{typ: value.Type(), ptr: value.Pointer()} nextDepth, err := state.enterContainer(path, containerDepth)
if _, ok := seen[current]; ok { if err != nil {
return nil, fmt.Errorf("%s: cyclic value is not supported", path) return nil, err
} }
seen[current] = struct{}{} return copyMapValue(value, path, state, allowEmptyMapKeys, nextDepth)
defer delete(seen, current)
return copyValue(value.Elem(), path, seen, allowEmptyMapKeys)
case reflect.Map:
return copyMapValue(value, path, seen, allowEmptyMapKeys)
case reflect.Slice: case reflect.Slice:
if value.IsNil() { if value.IsNil() {
if err := state.produceNode(path); err != nil {
return nil, err
}
return nil, nil return nil, nil
} }
return copySequenceValue(value, path, seen, allowEmptyMapKeys) nextDepth, err := state.enterContainer(path, containerDepth)
if err != nil {
return nil, err
}
return copySequenceValue(value, path, state, allowEmptyMapKeys, nextDepth)
case reflect.Array: case reflect.Array:
return copySequenceValue(value, path, seen, allowEmptyMapKeys) nextDepth, err := state.enterContainer(path, containerDepth)
if err != nil {
return nil, err
}
return copySequenceValue(value, path, state, allowEmptyMapKeys, nextDepth)
default: default:
return nil, fmt.Errorf("%s: unsupported JSON value type %s", path, value.Type()) return nil, fmt.Errorf("%s: unsupported JSON value type %s", path, value.Type())
} }
@@ -117,22 +146,26 @@ func copyValue(
func copyMapValue( func copyMapValue(
value reflect.Value, value reflect.Value,
path string, path string,
seen map[visit]struct{}, state *traversalState,
allowEmptyMapKeys bool, allowEmptyMapKeys bool,
containerDepth int,
) (any, error) { ) (any, error) {
if value.IsNil() {
return nil, nil
}
if value.Type().Key().Kind() != reflect.String { if value.Type().Key().Kind() != reflect.String {
return nil, fmt.Errorf("%s: map key type %s is not supported", path, value.Type().Key()) return nil, fmt.Errorf("%s: map key type %s is not supported", path, value.Type().Key())
} }
if err := state.produceNode(path); err != nil {
return nil, err
}
if err := state.ensureChildCapacity(path, value.Len()); err != nil {
return nil, err
}
current := visit{typ: value.Type(), ptr: value.Pointer()} current := visit{typ: value.Type(), ptr: value.Pointer()}
if _, ok := seen[current]; ok { if _, ok := state.active[current]; ok {
return nil, fmt.Errorf("%s: cyclic value is not supported", path) return nil, fmt.Errorf("%s: cyclic value is not supported", path)
} }
seen[current] = struct{}{} state.active[current] = struct{}{}
defer delete(seen, current) defer delete(state.active, current)
keys := value.MapKeys() keys := value.MapKeys()
sort.Slice(keys, func(i, j int) bool { sort.Slice(keys, func(i, j int) bool {
@@ -152,7 +185,13 @@ func copyMapValue(
if name == "" && !allowEmptyMapKeys { if name == "" && !allowEmptyMapKeys {
return nil, fmt.Errorf("%s: map key must not be empty", path) return nil, fmt.Errorf("%s: map key must not be empty", path)
} }
copied, err := copyValue(value.MapIndex(key), path+"."+name, seen, allowEmptyMapKeys) copied, err := copyValue(
value.MapIndex(key),
path+"."+name,
state,
allowEmptyMapKeys,
containerDepth,
)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -190,17 +229,25 @@ func copyMapValue(
func copySequenceValue( func copySequenceValue(
value reflect.Value, value reflect.Value,
path string, path string,
seen map[visit]struct{}, state *traversalState,
allowEmptyMapKeys bool, allowEmptyMapKeys bool,
containerDepth int,
) (any, error) { ) (any, error) {
if err := state.produceNode(path); err != nil {
return nil, err
}
if err := state.ensureChildCapacity(path, value.Len()); err != nil {
return nil, err
}
var current visit var current visit
if value.Kind() == reflect.Slice { if value.Kind() == reflect.Slice {
current = visit{typ: value.Type(), ptr: value.Pointer()} current = visit{typ: value.Type(), ptr: value.Pointer()}
if _, ok := seen[current]; ok { if _, ok := state.active[current]; ok {
return nil, fmt.Errorf("%s: cyclic value is not supported", path) return nil, fmt.Errorf("%s: cyclic value is not supported", path)
} }
seen[current] = struct{}{} state.active[current] = struct{}{}
defer delete(seen, current) defer delete(state.active, current)
} }
values := make([]any, value.Len()) values := make([]any, value.Len())
@@ -210,8 +257,9 @@ func copySequenceValue(
copied, err := copyValue( copied, err := copyValue(
value.Index(i), value.Index(i),
fmt.Sprintf("%s[%d]", path, i), fmt.Sprintf("%s[%d]", path, i),
seen, state,
allowEmptyMapKeys, allowEmptyMapKeys,
containerDepth,
) )
if err != nil { if err != nil {
return nil, err return nil, err
@@ -248,6 +296,73 @@ func copySequenceValue(
return out, nil return out, nil
} }
func validJSONNumber(number json.Number) bool {
var parsed json.Number
if err := json.Unmarshal([]byte(number.String()), &parsed); err != nil {
return false
}
return parsed.String() == number.String()
}
func newTraversalState() *traversalState {
return &traversalState{active: make(map[visit]struct{})}
}
func (state *traversalState) produceNode(path string) error {
if state.producedNodes >= maxProducedNodes {
return fmt.Errorf("%s: JSON value work limit exceeded", path)
}
state.producedNodes++
return nil
}
func (state *traversalState) enterContainer(path string, depth int) (int, error) {
depth++
if depth > maxContainerDepth {
return 0, fmt.Errorf("%s: JSON container depth limit exceeded", path)
}
return depth, nil
}
func (state *traversalState) ensureChildCapacity(path string, count int) error {
if count > maxProducedNodes-state.producedNodes {
return fmt.Errorf("%s: JSON value work limit exceeded", path)
}
return nil
}
func (state *traversalState) resolveIndirection(
value reflect.Value,
path string,
) (reflect.Value, func(), bool, error) {
var visits []visit
cleanup := func() {
for _, current := range visits {
delete(state.active, current)
}
}
for value.IsValid() && (value.Kind() == reflect.Interface || value.Kind() == reflect.Pointer) {
if value.IsNil() {
return reflect.Value{}, cleanup, true, nil
}
if value.Kind() == reflect.Pointer {
current := visit{typ: value.Type(), ptr: value.Pointer()}
if _, ok := state.active[current]; ok {
cleanup()
return reflect.Value{}, nil, false, fmt.Errorf("%s: cyclic value is not supported", path)
}
state.active[current] = struct{}{}
visits = append(visits, current)
}
value = value.Elem()
}
if !value.IsValid() {
return reflect.Value{}, cleanup, true, nil
}
return value, cleanup, false, nil
}
func canAssignNil(typ reflect.Type) bool { func canAssignNil(typ reflect.Type) bool {
switch typ.Kind() { switch typ.Kind() {
case reflect.Chan, reflect.Func, reflect.Interface, reflect.Map, reflect.Pointer, reflect.Slice: case reflect.Chan, reflect.Func, reflect.Interface, reflect.Map, reflect.Pointer, reflect.Slice:

View File

@@ -1,112 +1,332 @@
package jsonvalue_test package jsonvalue
import ( import (
"encoding/json" "encoding/json"
"math" "math"
"reflect" "reflect"
"strings"
"testing" "testing"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
) )
func TestCopyMapPreservesTypesAndIsolatesMutations(t *testing.T) { type (
nested := map[string]int{"limit": 2} namedBool bool
sequence := []string{"one", "two"} namedString string
input := map[string]any{ namedInt64 int64
"count": int64(7), namedUint64 uint64
"number": json.Number("-1.25e+2"), namedFloat32 float32
"nested": nested, namedFloat64 float64
"sequence": sequence, namedKey string
} namedMap map[namedKey]namedInt64
namedSlice []namedString
copied, err := jsonvalue.CopyMap(input) namedArray [1]map[string]int
if err != nil { )
t.Fatalf("copy map: %v", err)
}
nested["limit"] = 99
sequence[0] = "changed"
input["added"] = true
if got, ok := copied["count"].(int64); !ok || got != 7 {
t.Fatalf("integer type or value changed: %#v", copied["count"])
}
if got, ok := copied["number"].(json.Number); !ok || got != "-1.25e+2" {
t.Fatalf("JSON number type or value changed: %#v", copied["number"])
}
if got := copied["nested"].(map[string]int)["limit"]; got != 2 {
t.Fatalf("nested map was not isolated: %d", got)
}
if got := copied["sequence"].([]string)[0]; got != "one" {
t.Fatalf("sequence was not isolated: %q", got)
}
if _, ok := copied["added"]; ok {
t.Fatalf("top-level map was not isolated: %#v", copied)
}
}
func TestCopyAllowsEmptyObjectKeysAndIsolatesMutations(t *testing.T) {
nested := map[string]any{"": []any{"original"}}
copiedValue, err := jsonvalue.Copy(nested)
if err != nil {
t.Fatalf("copy value: %v", err)
}
nested[""].([]any)[0] = "changed"
copied := copiedValue.(map[string]any)
if got := copied[""].([]any)[0]; got != "original" {
t.Fatalf("copied value was not isolated: %v", got)
}
}
func TestCopyMapRejectsInvalidValues(t *testing.T) {
cyclicMap := map[string]any{}
cyclicMap["self"] = cyclicMap
cyclicSlice := []any{nil}
cyclicSlice[0] = cyclicSlice
func TestCopyPreservesSupportedScalarAndNumberTypes(t *testing.T) {
maxInt := int(^uint(0) >> 1)
minInt := -maxInt - 1
tests := []struct { tests := []struct {
name string name string
value any value any
}{ }{
{name: "empty nested key", value: map[string]int{"": 1}}, {name: "bool", value: true},
{name: "non-string map key", value: map[int]string{1: "one"}}, {name: "named bool", value: namedBool(true)},
{name: "unsupported value", value: make(chan int)}, {name: "string", value: "value"},
{name: "cyclic map", value: cyclicMap}, {name: "named string", value: namedString("value")},
{name: "cyclic slice", value: cyclicSlice}, {name: "int", value: minInt},
{name: "NaN", value: math.NaN()}, {name: "int8", value: int8(-1 << 7)},
{name: "positive infinity", value: math.Inf(1)}, {name: "int16", value: int16(-1 << 15)},
{name: "unsafe signed integer", value: int64(1 << 53)}, {name: "int32", value: int32(-1 << 31)},
{name: "unsafe unsigned integer", value: uint64(1 << 53)}, {name: "int64", value: int64(-1 << 63)},
{name: "named int64", value: namedInt64(1<<63 - 1)},
{name: "uint", value: ^uint(0)},
{name: "uint8", value: ^uint8(0)},
{name: "uint16", value: ^uint16(0)},
{name: "uint32", value: ^uint32(0)},
{name: "uint64", value: ^uint64(0)},
{name: "uintptr", value: ^uintptr(0)},
{name: "named uint64", value: namedUint64(^uint64(0))},
{name: "float32", value: float32(1.25)},
{name: "float64", value: float64(-2.5e100)},
{name: "named float32", value: namedFloat32(3.5)},
{name: "named float64", value: namedFloat64(-4.5e200)},
{name: "JSON number integer", value: json.Number("18446744073709551615")},
{name: "JSON number fraction", value: json.Number("-1.25e+2")},
{name: "JSON number beyond float64", value: json.Number("1e9999")},
} }
for _, tc := range tests { for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) { t.Run(tc.name, func(t *testing.T) {
if _, err := jsonvalue.CopyMap(map[string]any{"value": tc.value}); err == nil { got, err := Copy(tc.value)
if err != nil {
t.Fatalf("copy value: %v", err)
}
if !reflect.DeepEqual(got, tc.value) {
t.Fatalf("value or concrete type changed: got %#v (%T), want %#v (%T)", got, got, tc.value, tc.value)
}
})
}
}
func TestCopyRejectsInvalidNumbers(t *testing.T) {
tests := []struct {
name string
value any
}{
{name: "float32 NaN", value: float32(math.NaN())},
{name: "float64 NaN", value: math.NaN()},
{name: "named float NaN", value: namedFloat64(math.NaN())},
{name: "positive infinity", value: math.Inf(1)},
{name: "negative infinity", value: math.Inf(-1)},
{name: "empty JSON number", value: json.Number("")},
{name: "leading zero JSON number", value: json.Number("01")},
{name: "leading plus JSON number", value: json.Number("+1")},
{name: "trailing decimal JSON number", value: json.Number("1.")},
{name: "leading decimal JSON number", value: json.Number(".1")},
{name: "non-number JSON number", value: json.Number("NaN")},
{name: "spaced JSON number", value: json.Number(" 1")},
{name: "quoted JSON number", value: json.Number(`"1"`)},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
if _, err := Copy(tc.value); err == nil {
t.Fatal("expected validation error") t.Fatal("expected validation error")
} }
}) })
} }
} }
func TestCopyMapValidatesJSONNumberSyntaxAndRange(t *testing.T) { func TestCopyPreservesCompatibleCollectionsAndNilEmptyDistinctions(t *testing.T) {
for _, number := range []json.Number{"0", "-1", "1.25", "-1.25e+2"} { collections := []struct {
t.Run("valid "+number.String(), func(t *testing.T) { name string
got, err := jsonvalue.CopyMap(map[string]any{"value": number}) value any
if err != nil { }{
t.Fatalf("copy valid JSON number: %v", err) {name: "unnamed map", value: map[string]int{"limit": 2}},
{name: "named map", value: namedMap{"limit": 2}},
{name: "unnamed slice", value: []string{"one", "two"}},
{name: "named slice", value: namedSlice{"one", "two"}},
{name: "unnamed array", value: [2]int{1, 2}},
{name: "named array", value: namedArray{{"limit": 2}}},
{name: "empty map", value: map[string]int{}},
{name: "empty named map", value: namedMap{}},
{name: "empty slice", value: []string{}},
{name: "empty named slice", value: namedSlice{}},
{name: "empty array", value: [0]string{}},
} }
if !reflect.DeepEqual(got["value"], number) {
t.Fatalf("JSON number changed: got %#v want %#v", got["value"], number) for _, tc := range collections {
t.Run(tc.name, func(t *testing.T) {
got, err := Copy(tc.value)
if err != nil {
t.Fatalf("copy collection: %v", err)
}
if !reflect.DeepEqual(got, tc.value) || reflect.TypeOf(got) != reflect.TypeOf(tc.value) {
t.Fatalf("collection changed: got %#v (%T), want %#v (%T)", got, got, tc.value, tc.value)
}
kind := reflect.ValueOf(got).Kind()
if (kind == reflect.Map || kind == reflect.Slice) && reflect.ValueOf(got).IsNil() {
t.Fatal("non-nil collection became nil")
} }
}) })
} }
for _, number := range []json.Number{"", "01", "+1", "1.", ".1", "1e9999", "not-a-number"} { var nilMap map[string]int
t.Run("invalid "+number.String(), func(t *testing.T) { var nilSlice []string
if _, err := jsonvalue.CopyMap(map[string]any{"value": number}); err == nil { var nilPointer *namedInt64
t.Fatal("expected invalid JSON number error") for _, value := range []any{nil, nilMap, nilSlice, nilPointer} {
got, err := Copy(value)
if err != nil {
t.Fatalf("copy null value: %v", err)
}
if got != nil {
t.Fatalf("null value became %#v (%T)", got, got)
}
}
gotNil, err := CopyMap(nil)
if err != nil || gotNil != nil {
t.Fatalf("nil CopyMap result = %#v, %v", gotNil, err)
}
gotEmpty, err := CopyMap(map[string]any{})
if err != nil || gotEmpty == nil || len(gotEmpty) != 0 {
t.Fatalf("empty CopyMap result = %#v, %v", gotEmpty, err)
}
}
func TestCopyHandlesIndirectionAndIsolatesNestedMutations(t *testing.T) {
integer := namedInt64(7)
nestedMap := namedMap{"limit": 2}
nestedSlice := namedSlice{"original"}
nestedArray := namedArray{{"limit": 3}}
shared := []any{map[string]int{"value": 4}}
input := map[string]any{
"integer": &integer,
"map": nestedMap,
"slice": nestedSlice,
"array": nestedArray,
"first": shared,
"second": shared,
}
copiedValue, err := Copy(input)
if err != nil {
t.Fatalf("copy mixed tree: %v", err)
}
copied := copiedValue.(map[string]any)
nestedMap["limit"] = 20
nestedSlice[0] = "changed"
nestedArray[0]["limit"] = 30
shared[0].(map[string]int)["value"] = 40
if got, ok := copied["integer"].(namedInt64); !ok || got != 7 {
t.Fatalf("pointer target changed: %#v", copied["integer"])
}
if got := copied["map"].(namedMap)["limit"]; got != 2 {
t.Fatalf("nested map aliased input: %d", got)
}
if got := copied["slice"].(namedSlice)[0]; got != "original" {
t.Fatalf("nested slice aliased input: %q", got)
}
if got := copied["array"].(namedArray)[0]["limit"]; got != 3 {
t.Fatalf("nested array aliased input: %d", got)
}
first := copied["first"].([]any)
second := copied["second"].([]any)
if got := first[0].(map[string]int)["value"]; got != 4 {
t.Fatalf("shared child aliased input: %d", got)
}
first[0].(map[string]int)["value"] = 99
if got := second[0].(map[string]int)["value"]; got != 4 {
t.Fatalf("repeated acyclic value shared copied output: %d", got)
}
}
func TestCopyAndCopyMapApplyDistinctEmptyKeyRules(t *testing.T) {
nested := map[string]any{"": []any{"original"}}
copiedValue, err := Copy(nested)
if err != nil {
t.Fatalf("Copy rejected empty schema key: %v", err)
}
nested[""].([]any)[0] = "changed"
if got := copiedValue.(map[string]any)[""].([]any)[0]; got != "original" {
t.Fatalf("copied schema value was not isolated: %v", got)
}
_, err = CopyMap(map[string]any{"nested": map[string]any{"": true}})
if err == nil || !strings.Contains(err.Error(), "extra_params.nested") {
t.Fatalf("CopyMap empty-key error = %v", err)
}
}
func TestCopyRejectsUnsupportedValuesAndActiveCycles(t *testing.T) {
cyclicMap := map[string]any{}
cyclicMap["self"] = cyclicMap
cyclicSlice := []any{nil}
cyclicSlice[0] = cyclicSlice
var cyclicPointer any
cyclicPointer = &cyclicPointer
tests := []struct {
name string
value any
wantPath string
}{
{name: "non-string map key", value: map[int]string{1: "one"}, wantPath: "value"},
{name: "unsupported channel", value: make(chan int), wantPath: "value"},
{name: "deterministic map path", value: map[string]any{"z": make(chan int), "a": make(chan int)}, wantPath: "value.a"},
{name: "cyclic map", value: cyclicMap, wantPath: "value.self"},
{name: "cyclic slice", value: cyclicSlice, wantPath: "value[0]"},
{name: "cyclic pointer", value: cyclicPointer, wantPath: "value"},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
_, err := Copy(tc.value)
if err == nil || !strings.Contains(err.Error(), tc.wantPath) {
t.Fatalf("error = %v, want structural path %q", err, tc.wantPath)
} }
}) })
} }
} }
func TestCopyEnforcesContainerDepth(t *testing.T) {
tests := []struct {
name string
depth int
wantErr bool
}{
{name: "just below", depth: maxContainerDepth - 1},
{name: "at limit", depth: maxContainerDepth},
{name: "over limit", depth: maxContainerDepth + 1, wantErr: true},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
_, err := Copy(alternatingContainers(tc.depth))
if tc.wantErr {
if err == nil || !strings.HasPrefix(err.Error(), "value") || !strings.Contains(err.Error(), "container depth limit") {
t.Fatalf("depth error = %v", err)
}
return
}
if err != nil {
t.Fatalf("copy depth %d: %v", tc.depth, err)
}
})
}
}
func TestCopyEnforcesProducedNodeBudgetForRepeatedAcyclicValues(t *testing.T) {
shared := []any{true}
sharedOccurrences := (maxProducedNodes - 2) / 2
justBelow := repeatedValues(shared, sharedOccurrences, 0)
atLimit := repeatedValues(shared, sharedOccurrences, 1)
overLimit := repeatedValues(shared, sharedOccurrences, 2)
for name, value := range map[string]any{
"just below": justBelow,
"at limit": atLimit,
} {
t.Run(name, func(t *testing.T) {
if _, err := Copy(value); err != nil {
t.Fatalf("copy value within work budget: %v", err)
}
})
}
_, err := Copy(overLimit)
if err == nil || !strings.HasPrefix(err.Error(), "value[") || !strings.Contains(err.Error(), "value work limit") {
t.Fatalf("work-budget error = %v", err)
}
_, err = Copy(make([]any, maxProducedNodes))
if err == nil || !strings.HasPrefix(err.Error(), "value:") || !strings.Contains(err.Error(), "value work limit") {
t.Fatalf("flat work-budget error = %v", err)
}
}
func alternatingContainers(depth int) any {
var value any = true
for level := 0; level < depth; level++ {
switch level % 3 {
case 0:
value = map[string]any{"child": value}
case 1:
value = []any{value}
default:
value = [1]any{value}
}
}
return value
}
func repeatedValues(shared []any, occurrences, leadingScalars int) []any {
values := make([]any, 0, leadingScalars+occurrences)
for i := 0; i < leadingScalars; i++ {
values = append(values, false)
}
for i := 0; i < occurrences; i++ {
values = append(values, shared)
}
return values
}

View File

@@ -25,6 +25,20 @@ var (
ErrMalformedResponse = errors.New("malformed llm response") ErrMalformedResponse = errors.New("malformed llm response")
) )
const maxOpenAIChatResponseBytes int64 = 16 << 20
type requestFailedError struct {
cause error
}
func (e *requestFailedError) Error() string {
return ErrRequestFailed.Error()
}
func (e *requestFailedError) Unwrap() []error {
return []error{ErrRequestFailed, e.cause}
}
type OpenAICompatibleConfig struct { type OpenAICompatibleConfig struct {
BaseURL string BaseURL string
Model string Model string
@@ -39,9 +53,11 @@ type OpenAICompatibleClient struct {
} }
func NewOpenAICompatibleClient(cfg OpenAICompatibleConfig) (*OpenAICompatibleClient, error) { func NewOpenAICompatibleClient(cfg OpenAICompatibleConfig) (*OpenAICompatibleClient, error) {
baseURL := strings.TrimSpace(cfg.BaseURL) baseURL := ""
if baseURL != "" { if strings.TrimSpace(cfg.BaseURL) != "" {
if _, err := url.ParseRequestURI(baseURL); err != nil { var err error
baseURL, err = domain.NormalizeOpenAICompatibleBaseEndpoint(cfg.BaseURL)
if err != nil {
return nil, fmt.Errorf("%w: invalid base URL: %v", ErrInvalidConfig, err) return nil, fmt.Errorf("%w: invalid base URL: %v", ErrInvalidConfig, err)
} }
} }
@@ -63,25 +79,29 @@ func NewOpenAICompatibleClient(cfg OpenAICompatibleConfig) (*OpenAICompatibleCli
} }
return &OpenAICompatibleClient{ return &OpenAICompatibleClient{
baseURL: strings.TrimRight(baseURL, "/"), baseURL: baseURL,
defaultModel: cfg.Model, defaultModel: cfg.Model,
httpClient: client, httpClient: client,
}, nil }, nil
} }
func (c *OpenAICompatibleClient) Generate(ctx context.Context, req domain.GenerateRequest) (*domain.GenerateResponse, error) { func (c *OpenAICompatibleClient) Generate(ctx context.Context, req domain.GenerateRequest) (*domain.GenerateResponse, error) {
if req.Target.TimeoutSeconds < 0 { if err := domain.ValidateExecutionTargetSettings(req.Target); err != nil {
return nil, fmt.Errorf("%w: timeout_seconds must be greater than or equal to 0", ErrInvalidRequest) return nil, fmt.Errorf("%w: %v", ErrInvalidRequest, err)
} }
endpoint := strings.TrimSpace(req.Target.Endpoint) selectedEndpoint := req.Target.Endpoint
if endpoint == "" { if strings.TrimSpace(selectedEndpoint) == "" {
endpoint = c.baseURL selectedEndpoint = c.baseURL
} }
if endpoint == "" { endpoint, err := domain.NormalizeOpenAICompatibleBaseEndpoint(selectedEndpoint)
return nil, fmt.Errorf("%w: endpoint is required", ErrInvalidRequest) if err != nil {
return nil, fmt.Errorf("%w: invalid endpoint: %v", ErrInvalidRequest, err)
}
endpoint, err = url.JoinPath(endpoint, defaults.OpenAIChatCompletionsPath)
if err != nil {
return nil, fmt.Errorf("%w: invalid endpoint path: %v", ErrInvalidRequest, err)
} }
endpoint = strings.TrimRight(endpoint, "/") + defaults.OpenAIChatCompletionsPath
wireReq, err := openAIChatRequestFromGenerateRequest(req, c.defaultModel) wireReq, err := openAIChatRequestFromGenerateRequest(req, c.defaultModel)
if err != nil { if err != nil {
@@ -113,13 +133,18 @@ func (c *OpenAICompatibleClient) Generate(ctx context.Context, req domain.Genera
return nil, fmt.Errorf("%w: failed to create request: %v", ErrRequestFailed, err) return nil, fmt.Errorf("%w: failed to create request: %v", ErrRequestFailed, err)
} }
httpReq.Header.Set("Content-Type", "application/json") httpReq.Header.Set("Content-Type", "application/json")
if apiKey := strings.TrimSpace(req.Target.APIKey); apiKey != "" { apiKey := strings.TrimSpace(req.Target.APIKey)
httpReq.Header.Set("Authorization", "Bearer "+apiKey) envName := strings.TrimSpace(req.Target.APIKeyEnv)
} else if envName := strings.TrimSpace(req.Target.APIKeyEnv); envName != "" { if apiKey == "" && envName != "" {
apiKey := strings.TrimSpace(os.Getenv(envName)) apiKey = strings.TrimSpace(os.Getenv(envName))
if apiKey == "" { }
if apiKey == "" && req.Target.APIKeyRequired {
if envName != "" {
return nil, fmt.Errorf("%w: api key environment variable %q is not set", ErrInvalidRequest, envName) return nil, fmt.Errorf("%w: api key environment variable %q is not set", ErrInvalidRequest, envName)
} }
return nil, fmt.Errorf("%w: api key is required", ErrInvalidRequest)
}
if apiKey != "" {
httpReq.Header.Set("Authorization", "Bearer "+apiKey) httpReq.Header.Set("Authorization", "Bearer "+apiKey)
} }
@@ -130,30 +155,36 @@ func (c *OpenAICompatibleClient) Generate(ctx context.Context, req domain.Genera
httpResp, err := httpClient.Do(httpReq) httpResp, err := httpClient.Do(httpReq)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %v", ErrRequestFailed, err) return nil, &requestFailedError{cause: err}
} }
defer httpResp.Body.Close() defer httpResp.Body.Close()
if httpResp.StatusCode < 200 || httpResp.StatusCode >= 300 { if httpResp.StatusCode < 200 || httpResp.StatusCode >= 300 {
_, _ = io.Copy(io.Discard, io.LimitReader(httpResp.Body, 4096)) return nil, providerHTTPErrorFromBody(
return nil, fmt.Errorf("%w: status=%d", ErrUnexpectedStatus, httpResp.StatusCode) httpResp.StatusCode,
httpResp.ContentLength,
httpResp.Body,
)
}
if httpResp.ContentLength > maxOpenAIChatResponseBytes {
return nil, openAIChatResponseTooLargeError()
} }
var wireResp openAIChatResponse wireResp, err := decodeOpenAIChatResponse(httpResp.Body)
if err := json.NewDecoder(httpResp.Body).Decode(&wireResp); err != nil { if err != nil {
return nil, fmt.Errorf("%w: failed to decode response: %v", ErrMalformedResponse, err) return nil, err
} }
if len(wireResp.Choices) == 0 { if len(wireResp.Choices) == 0 {
return nil, fmt.Errorf("%w: no choices returned", ErrMalformedResponse) return nil, fmt.Errorf("%w: no choices returned", ErrMalformedResponse)
} }
content := wireResp.Choices[0].Message.Content content := wireResp.Choices[0].Message.Content
if content == "" { if content == nil {
return nil, fmt.Errorf("%w: first choice has empty message content", ErrMalformedResponse) return nil, fmt.Errorf("%w: first choice has missing message content", ErrMalformedResponse)
} }
return &domain.GenerateResponse{ return &domain.GenerateResponse{
Content: content, Content: *content,
Usage: domain.TokenUsage{ Usage: domain.TokenUsage{
PromptTokens: wireResp.Usage.PromptTokens, PromptTokens: wireResp.Usage.PromptTokens,
CompletionTokens: wireResp.Usage.CompletionTokens, CompletionTokens: wireResp.Usage.CompletionTokens,
@@ -164,6 +195,46 @@ func (c *OpenAICompatibleClient) Generate(ctx context.Context, req domain.Genera
}, nil }, nil
} }
func decodeOpenAIChatResponse(body io.Reader) (openAIChatResponse, error) {
limited := &io.LimitedReader{
R: body,
N: maxOpenAIChatResponseBytes + 1,
}
decoder := json.NewDecoder(limited)
var response openAIChatResponse
if err := decoder.Decode(&response); err != nil {
if limited.N == 0 {
return openAIChatResponse{}, openAIChatResponseTooLargeError()
}
return openAIChatResponse{}, fmt.Errorf("%w: failed to decode response", ErrMalformedResponse)
}
if limited.N == 0 {
return openAIChatResponse{}, openAIChatResponseTooLargeError()
}
var trailing any
if err := decoder.Decode(&trailing); !errors.Is(err, io.EOF) {
if limited.N == 0 {
return openAIChatResponse{}, openAIChatResponseTooLargeError()
}
return openAIChatResponse{}, fmt.Errorf("%w: response contains trailing data", ErrMalformedResponse)
}
if limited.N == 0 {
return openAIChatResponse{}, openAIChatResponseTooLargeError()
}
return response, nil
}
func openAIChatResponseTooLargeError() error {
return fmt.Errorf(
"%w: response exceeds %d-byte limit",
ErrMalformedResponse,
maxOpenAIChatResponseBytes,
)
}
func openAIChatRequestFromGenerateRequest(req domain.GenerateRequest, defaultModel string) (openAIChatRequest, error) { func openAIChatRequestFromGenerateRequest(req domain.GenerateRequest, defaultModel string) (openAIChatRequest, error) {
model := strings.TrimSpace(req.Target.Model) model := strings.TrimSpace(req.Target.Model)
if model == "" { if model == "" {
@@ -309,7 +380,7 @@ type openAICacheControl struct {
type openAIChatResponseMessage struct { type openAIChatResponseMessage struct {
Role string `json:"role"` Role string `json:"role"`
Content string `json:"content"` Content *string `json:"content"`
} }
type openAIChatResponse struct { type openAIChatResponse struct {

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,189 @@
package llm
import (
"encoding/json"
"errors"
"fmt"
"io"
"strings"
"unicode"
)
const (
maxProviderErrorResponseBytes int64 = 64 << 10
maxProviderErrorIdentifierRunes = 256
maxProviderErrorMessageRunes = 4096
)
// ProviderHTTPError describes a non-success response from an LLM provider.
type ProviderHTTPError struct {
statusCode int
providerCode string
providerType string
providerMessage string
}
func (e *ProviderHTTPError) StatusCode() int {
if e == nil {
return 0
}
return e.statusCode
}
func (e *ProviderHTTPError) ProviderCode() string {
if e == nil {
return ""
}
return e.providerCode
}
func (e *ProviderHTTPError) ProviderType() string {
if e == nil {
return ""
}
return e.providerType
}
func (e *ProviderHTTPError) ProviderMessage() string {
if e == nil {
return ""
}
return e.providerMessage
}
func (e *ProviderHTTPError) Error() string {
if e == nil || e.statusCode == 0 {
return ErrUnexpectedStatus.Error()
}
return fmt.Sprintf("%s: status=%d", ErrUnexpectedStatus, e.statusCode)
}
func (e *ProviderHTTPError) GoString() string {
return e.Error()
}
func (e *ProviderHTTPError) Unwrap() error {
return ErrUnexpectedStatus
}
type providerErrorDetails struct {
providerCode string
providerType string
providerMessage string
}
func newProviderHTTPError(statusCode int, details providerErrorDetails) *ProviderHTTPError {
return &ProviderHTTPError{
statusCode: statusCode,
providerCode: details.providerCode,
providerType: details.providerType,
providerMessage: details.providerMessage,
}
}
func providerHTTPErrorFromBody(statusCode int, contentLength int64, body io.Reader) *ProviderHTTPError {
if contentLength > maxProviderErrorResponseBytes {
return newProviderHTTPError(statusCode, providerErrorDetails{})
}
limited := &io.LimitedReader{
R: body,
N: maxProviderErrorResponseBytes + 1,
}
contents, err := io.ReadAll(limited)
if err != nil || limited.N == 0 {
return newProviderHTTPError(statusCode, providerErrorDetails{})
}
return newProviderHTTPError(statusCode, parseProviderErrorEnvelope(contents))
}
func parseProviderErrorEnvelope(body []byte) providerErrorDetails {
decoder := json.NewDecoder(strings.NewReader(string(body)))
decoder.UseNumber()
var envelope map[string]json.RawMessage
if err := decoder.Decode(&envelope); err != nil {
return providerErrorDetails{}
}
var trailing any
if err := decoder.Decode(&trailing); !errors.Is(err, io.EOF) {
return providerErrorDetails{}
}
rawError, ok := envelope["error"]
if !ok {
return providerErrorDetails{}
}
var providerError map[string]json.RawMessage
if err := json.Unmarshal(rawError, &providerError); err != nil || providerError == nil {
return providerErrorDetails{}
}
var details providerErrorDetails
if raw, ok := providerError["message"]; ok {
var value string
if json.Unmarshal(raw, &value) == nil {
details.providerMessage = normalizeProviderErrorMessage(value)
}
}
if raw, ok := providerError["type"]; ok {
var value string
if json.Unmarshal(raw, &value) == nil {
details.providerType = normalizeProviderErrorIdentifier(value)
}
}
if raw, ok := providerError["code"]; ok {
var value any
fieldDecoder := json.NewDecoder(strings.NewReader(string(raw)))
fieldDecoder.UseNumber()
if fieldDecoder.Decode(&value) == nil {
switch value := value.(type) {
case string:
details.providerCode = normalizeProviderErrorIdentifier(value)
case json.Number:
details.providerCode = normalizeProviderErrorIdentifier(value.String())
}
}
}
return details
}
func normalizeProviderErrorIdentifier(value string) string {
normalized := normalizeProviderErrorText(value)
if len([]rune(normalized)) > maxProviderErrorIdentifierRunes {
return ""
}
return normalized
}
func normalizeProviderErrorMessage(value string) string {
normalized := normalizeProviderErrorText(value)
runes := []rune(normalized)
if len(runes) <= maxProviderErrorMessageRunes {
return normalized
}
return string(runes[:maxProviderErrorMessageRunes-1]) + "…"
}
func normalizeProviderErrorText(value string) string {
value = strings.ToValidUTF8(value, "<22>")
var result strings.Builder
result.Grow(len(value))
separatorPending := false
for _, r := range value {
if unicode.IsSpace(r) || unicode.IsControl(r) || unicode.In(r, unicode.Cf) {
if result.Len() > 0 {
separatorPending = true
}
continue
}
if separatorPending {
result.WriteByte(' ')
separatorPending = false
}
result.WriteRune(r)
}
return result.String()
}

View File

@@ -0,0 +1,255 @@
package llm
import (
"errors"
"fmt"
"io"
"reflect"
"strings"
"testing"
"unicode/utf8"
)
type guardedReader struct {
reader io.Reader
remaining int64
bytes int64
violated bool
}
func (r *guardedReader) Read(buffer []byte) (int, error) {
if int64(len(buffer)) > r.remaining {
r.violated = true
return 0, errors.New("reader was read past its allowed boundary")
}
n, err := r.reader.Read(buffer)
r.bytes += int64(n)
r.remaining -= int64(n)
return n, err
}
type failingReader struct {
err error
}
func (r failingReader) Read([]byte) (int, error) {
return 0, r.err
}
func TestProviderHTTPErrorEnvelopeParsing(t *testing.T) {
tests := []struct {
name string
body string
want providerErrorDetails
}{
{
name: "all supported string fields",
body: `{"error":{"message":"diagnostic","type":"invalid_request_error","code":"unsupported_parameter"}}`,
want: providerErrorDetails{providerMessage: "diagnostic", providerType: "invalid_request_error", providerCode: "unsupported_parameter"},
},
{
name: "integer code",
body: `{"error":{"code":17}}`,
want: providerErrorDetails{providerCode: "17"},
},
{
name: "fractional code",
body: `{"error":{"code":1.25}}`,
want: providerErrorDetails{providerCode: "1.25"},
},
{
name: "exponent code",
body: `{"error":{"code":6.02e+23}}`,
want: providerErrorDetails{providerCode: "6.02e+23"},
},
{
name: "invalid fields do not discard valid fields",
body: `{"error":{"message":null,"type":"invalid_request_error","code":false}}`,
want: providerErrorDetails{providerType: "invalid_request_error"},
},
{
name: "unknown fields are ignored",
body: `{"trace":"do not retain","error":{"param":"temperature","metadata":{"secret":"x"}}}`,
want: providerErrorDetails{},
},
{name: "missing error", body: `{}`, want: providerErrorDetails{}},
{name: "null error", body: `{"error":null}`, want: providerErrorDetails{}},
{name: "scalar error", body: `{"error":"nope"}`, want: providerErrorDetails{}},
{name: "empty error", body: `{"error":{}}`, want: providerErrorDetails{}},
{name: "malformed", body: `{"error":`, want: providerErrorDetails{}},
{name: "truncated", body: `{"error":{"message":"x"`, want: providerErrorDetails{}},
{name: "trailing garbage", body: `{"error":{"message":"x"}} garbage`, want: providerErrorDetails{}},
{name: "second document", body: `{"error":{"message":"x"}} {}`, want: providerErrorDetails{}},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
if got := parseProviderErrorEnvelope([]byte(tc.body)); !reflect.DeepEqual(got, tc.want) {
t.Fatalf("parseProviderErrorEnvelope() = %#v, want %#v", got, tc.want)
}
})
}
}
func TestProviderErrorTextNormalizationAndLimits(t *testing.T) {
validIdentifier := strings.Repeat("界", maxProviderErrorIdentifierRunes)
validMessage := strings.Repeat("界", maxProviderErrorMessageRunes)
tests := []struct {
name string
got string
want string
}{
{name: "multibyte text", got: "Grüße 世界", want: "Grüße 世界"},
{name: "invalid UTF-8", got: string([]byte{'a', 0xff, 'b'}), want: "a<>b"},
{name: "whitespace control and format runs", got: " \n\talpha\x00\u200b\u200bbeta \r ", want: "alpha beta"},
{name: "blank normalization", got: "\t\u200b\n", want: ""},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
if got := normalizeProviderErrorText(tc.got); got != tc.want {
t.Fatalf("normalizeProviderErrorText() = %q, want %q", got, tc.want)
}
})
}
if got := normalizeProviderErrorIdentifier(validIdentifier); got != validIdentifier {
t.Fatalf("exact identifier boundary = %q, want retained value", got)
}
if got := normalizeProviderErrorIdentifier(validIdentifier + "界"); got != "" {
t.Fatalf("overlong identifier = %q, want empty", got)
}
if got := normalizeProviderErrorMessage(validMessage); got != validMessage {
t.Fatalf("exact message boundary = %q, want retained value", got)
}
wantTruncatedMessage := strings.Repeat("界", maxProviderErrorMessageRunes-1) + "…"
if got := normalizeProviderErrorMessage(validMessage + "界"); got != wantTruncatedMessage {
t.Fatalf("overlong message length = %d, want %d", utf8.RuneCountInString(got), maxProviderErrorMessageRunes)
}
}
func TestProviderHTTPErrorIdentityAndFormatting(t *testing.T) {
const marker = "provider-secret-marker"
err := newProviderHTTPError(429, providerErrorDetails{
providerCode: marker + "-code",
providerType: marker + "-type",
providerMessage: marker + "-message",
})
if err.StatusCode() != 429 || err.ProviderCode() != marker+"-code" || err.ProviderType() != marker+"-type" || err.ProviderMessage() != marker+"-message" {
t.Fatalf("accessors returned unexpected values: %#v", err)
}
if !errors.Is(err, ErrUnexpectedStatus) {
t.Fatalf("errors.Is(%v, ErrUnexpectedStatus) = false", err)
}
for _, rendered := range []string{fmt.Sprintf("%v", err), fmt.Sprintf("%+v", err), fmt.Sprintf("%#v", err)} {
if rendered != "llm returned non-success status: status=429" {
t.Fatalf("formatted error = %q", rendered)
}
if strings.Contains(rendered, marker) {
t.Fatalf("formatted error exposed provider marker: %q", rendered)
}
}
var nilError *ProviderHTTPError
if nilError.StatusCode() != 0 || nilError.ProviderCode() != "" || nilError.ProviderType() != "" || nilError.ProviderMessage() != "" {
t.Fatal("nil accessors returned provider values")
}
if nilError.Error() != "llm returned non-success status" || nilError.GoString() != "llm returned non-success status" || !errors.Is(nilError, ErrUnexpectedStatus) {
t.Fatalf("nil error behavior is not safe: %v", nilError)
}
zero := &ProviderHTTPError{}
if zero.Error() != "llm returned non-success status" || zero.GoString() != "llm returned non-success status" || !errors.Is(zero, ErrUnexpectedStatus) {
t.Fatalf("zero error behavior is not safe: %v", zero)
}
}
func TestProviderHTTPErrorBodyBounds(t *testing.T) {
const (
statusCode = 502
marker = "provider-body-marker"
)
ordinaryBody := `{"error":{"message":"` + marker + `"}}`
exactLimitBody := ordinaryBody + strings.Repeat(" ", int(maxProviderErrorResponseBytes)-len(ordinaryBody))
overLimitBody := ordinaryBody + strings.Repeat(" ", int(maxProviderErrorResponseBytes)+1-len(ordinaryBody))
tests := []struct {
name string
contentLength int64
reader io.Reader
wantRead int64
wantMessage string
}{
{
name: "recognized envelope",
contentLength: int64(len(ordinaryBody)),
reader: strings.NewReader(ordinaryBody),
wantRead: int64(len(ordinaryBody)),
wantMessage: marker,
},
{
name: "exact limit",
contentLength: maxProviderErrorResponseBytes,
reader: strings.NewReader(exactLimitBody),
wantRead: maxProviderErrorResponseBytes,
wantMessage: marker,
},
{
name: "declared oversize does not read",
contentLength: maxProviderErrorResponseBytes + 1,
reader: strings.NewReader(ordinaryBody),
wantRead: 0,
},
{
name: "unknown length oversize",
contentLength: -1,
reader: strings.NewReader(overLimitBody),
wantRead: maxProviderErrorResponseBytes + 1,
},
{
name: "underreported oversize",
contentLength: maxProviderErrorResponseBytes,
reader: strings.NewReader(overLimitBody),
wantRead: maxProviderErrorResponseBytes + 1,
},
{
name: "read failure",
contentLength: -1,
reader: failingReader{err: errors.New("read failure")},
wantRead: 0,
},
{
name: "empty body",
contentLength: 0,
reader: strings.NewReader(""),
wantRead: 0,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
reader := &guardedReader{
reader: tc.reader,
remaining: maxProviderErrorResponseBytes + 1,
}
err := providerHTTPErrorFromBody(statusCode, tc.contentLength, reader)
if err == nil || err.StatusCode() != statusCode {
t.Fatalf("error status = %v, want %d", err, statusCode)
}
if reader.bytes != tc.wantRead {
t.Fatalf("body bytes read = %d, want %d", reader.bytes, tc.wantRead)
}
if reader.violated {
t.Fatal("body reader was asked to read beyond the overflow probe")
}
if got := err.ProviderMessage(); got != tc.wantMessage {
t.Fatalf("provider message = %q, want %q", got, tc.wantMessage)
}
if tc.wantMessage == "" {
if err.ProviderCode() != "" || err.ProviderType() != "" || strings.Contains(err.Error(), marker) {
t.Fatalf("discarded details were retained: %#v", err)
}
}
})
}
}

View File

@@ -0,0 +1,145 @@
package llm
import (
"context"
"errors"
"io"
"net/http"
"strings"
"testing"
)
func TestOpenAICompatibleClientStructuredNonSuccessResponse(t *testing.T) {
body := `{"error":{"message":" provider\nmessage\u200b","type":"invalid\ttype","code":1.5e+4}}`
responseBody := &countingReadCloser{reader: strings.NewReader(body)}
client := newNonSuccessResponseClient(t, http.StatusBadRequest, int64(len(body)), responseBody)
response, err := client.Generate(context.Background(), ordinaryGenerateRequest())
if response != nil {
t.Fatalf("response = %#v, want nil", response)
}
if !errors.Is(err, ErrUnexpectedStatus) {
t.Fatalf("errors.Is(%v, ErrUnexpectedStatus) = false", err)
}
var providerHTTPError *ProviderHTTPError
if !errors.As(err, &providerHTTPError) {
t.Fatalf("error = %T, want *ProviderHTTPError", err)
}
if providerHTTPError.StatusCode() != http.StatusBadRequest || providerHTTPError.ProviderCode() != "1.5e+4" || providerHTTPError.ProviderType() != "invalid type" || providerHTTPError.ProviderMessage() != "provider message" {
t.Fatalf("provider error = %#v", providerHTTPError)
}
if !responseBody.closed {
t.Fatal("non-success response body was not closed")
}
}
func TestOpenAICompatibleClientNonSuccessBodyOwnership(t *testing.T) {
const marker = "provider-body-marker"
normalBody := `{"error":{"message":"` + marker + `"}}`
overLimitBody := normalBody + strings.Repeat(" ", int(maxProviderErrorResponseBytes)+1-len(normalBody))
tests := []struct {
name string
contentLength int64
reader io.Reader
wantRead int64
wantMessage string
}{
{
name: "normal",
contentLength: int64(len(normalBody)),
reader: strings.NewReader(normalBody),
wantRead: int64(len(normalBody)),
wantMessage: marker,
},
{
name: "declared oversize",
contentLength: maxProviderErrorResponseBytes + 1,
reader: strings.NewReader(normalBody),
wantRead: 0,
},
{
name: "streamed oversize",
contentLength: -1,
reader: &guardedReader{
reader: strings.NewReader(overLimitBody),
remaining: maxProviderErrorResponseBytes + 1,
},
wantRead: maxProviderErrorResponseBytes + 1,
},
{
name: "underreported oversize",
contentLength: maxProviderErrorResponseBytes,
reader: &guardedReader{
reader: strings.NewReader(overLimitBody),
remaining: maxProviderErrorResponseBytes + 1,
},
wantRead: maxProviderErrorResponseBytes + 1,
},
{
name: "malformed",
contentLength: 1,
reader: strings.NewReader("{"),
wantRead: 1,
},
{
name: "read failure",
contentLength: -1,
reader: failingReader{err: errors.New("response read failed")},
wantRead: 0,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
body := &countingReadCloser{reader: tc.reader}
client := newNonSuccessResponseClient(t, http.StatusBadGateway, tc.contentLength, body)
response, err := client.Generate(context.Background(), ordinaryGenerateRequest())
if response != nil {
t.Fatalf("response = %#v, want nil", response)
}
var providerHTTPError *ProviderHTTPError
if !errors.As(err, &providerHTTPError) {
t.Fatalf("error = %T, want *ProviderHTTPError", err)
}
if !body.closed {
t.Fatal("response body was not closed")
}
if body.bytesRead != tc.wantRead {
t.Fatalf("body bytes read = %d, want %d", body.bytesRead, tc.wantRead)
}
if body.bytesRead > maxProviderErrorResponseBytes+1 {
t.Fatalf("body bytes read = %d, exceeds overflow probe", body.bytesRead)
}
if guarded, ok := tc.reader.(*guardedReader); ok && guarded.violated {
t.Fatal("body reader was asked to read beyond the overflow probe")
}
if got := providerHTTPError.ProviderMessage(); got != tc.wantMessage {
t.Fatalf("provider message = %q, want %q", got, tc.wantMessage)
}
if tc.wantMessage == "" && (providerHTTPError.ProviderCode() != "" || providerHTTPError.ProviderType() != "" || strings.Contains(providerHTTPError.Error(), marker)) {
t.Fatalf("discarded details were retained: %#v", providerHTTPError)
}
})
}
}
func newNonSuccessResponseClient(t *testing.T, statusCode int, contentLength int64, body io.ReadCloser) *OpenAICompatibleClient {
t.Helper()
client, err := NewOpenAICompatibleClient(OpenAICompatibleConfig{
BaseURL: "https://provider.example/v1",
Model: "m",
HTTPClient: &http.Client{Transport: roundTripFunc(func(*http.Request) (*http.Response, error) {
return &http.Response{
StatusCode: statusCode,
ContentLength: contentLength,
Body: body,
}, nil
})},
})
if err != nil {
t.Fatalf("construct client: %v", err)
}
return client
}

View File

@@ -0,0 +1,3 @@
id: rakestrawhome-gemma-4-31b
backend: rakestrawhome
model: google/gemma-4-31b-it

View File

@@ -2,7 +2,6 @@ package builtin
import ( import (
"embed" "embed"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
) )
@@ -15,17 +14,3 @@ var assets embed.FS
func NewRepository() profile.Repository { func NewRepository() profile.Repository {
return profile.NewFSRepository(assets, assetRoot) return profile.NewFSRepository(assets, assetRoot)
} }
func NewRepositoryWithPrimary(primary profile.Repository) profile.Repository {
if primary == nil {
return NewRepository()
}
return profile.NewOverlayRepository(primary, NewRepository())
}
func NewRepositoryWithDirectory(dir string) profile.Repository {
if strings.TrimSpace(dir) == "" {
return NewRepository()
}
return NewRepositoryWithPrimary(profile.NewFilesystemRepository(dir))
}

View File

@@ -2,14 +2,11 @@ package builtin
import ( import (
"context" "context"
"errors"
"io/fs" "io/fs"
"strings" "strings"
"testing" "testing"
"gitea.maximumdirect.net/eric/promptkit/internal/backend" "gitea.maximumdirect.net/eric/promptkit/internal/backend"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/profile"
"gopkg.in/yaml.v3" "gopkg.in/yaml.v3"
) )
@@ -29,8 +26,8 @@ func TestBuiltInProfilesValidateThroughRepository(t *testing.T) {
if p.ID != id { if p.ID != id {
t.Fatalf("expected profile id %q, got %q", id, p.ID) t.Fatalf("expected profile id %q, got %q", id, p.ID)
} }
if p.BackendID != backend.OpenRouterID { if !builtInBackendIDs[p.BackendID] {
t.Fatalf("expected profile %q to select %q, got %q", id, backend.OpenRouterID, p.BackendID) t.Fatalf("expected profile %q to select a maintained built-in, got %q", id, p.BackendID)
} }
if p.Endpoint != "" || p.APIKeyEnv != "" { if p.Endpoint != "" || p.APIKeyEnv != "" {
t.Fatalf("expected profile %q to inherit backend connection settings, got endpoint=%q api_key_env=%q", id, p.Endpoint, p.APIKeyEnv) t.Fatalf("expected profile %q to inherit backend connection settings, got endpoint=%q api_key_env=%q", id, p.Endpoint, p.APIKeyEnv)
@@ -43,6 +40,33 @@ func TestBuiltInProfilesDoNotContainDuplicateIDsOrRawAPIKeys(t *testing.T) {
loadBuiltInProfileIDs(t) loadBuiltInProfileIDs(t)
} }
func TestRakestrawhomeGemmaProfileUsesNativeDefaults(t *testing.T) {
p, err := NewRepository().GetProfile(context.Background(), "rakestrawhome-gemma-4-31b")
if err != nil {
t.Fatalf("load Rakestrawhome Gemma profile: %v", err)
}
if p.ID != "rakestrawhome-gemma-4-31b" ||
p.BackendID != backend.RakestrawHomeID ||
p.Model != "google/gemma-4-31b-it" ||
p.Endpoint != "" ||
p.Temperature != 0 ||
p.MaxTokens != 0 ||
p.TopP != 0 ||
p.TimeoutSeconds != 0 ||
p.ServiceTier != "" ||
p.ReasoningEffort != "" ||
p.APIKeyEnv != "" ||
p.APIKeyRequired ||
p.ExtraParams != nil {
t.Fatalf("unexpected Rakestrawhome Gemma profile: %#v", p)
}
}
var builtInBackendIDs = map[string]bool{
backend.OpenRouterID: true,
backend.RakestrawHomeID: true,
}
func loadBuiltInProfileIDs(t *testing.T) map[string]string { func loadBuiltInProfileIDs(t *testing.T) map[string]string {
t.Helper() t.Helper()
@@ -67,8 +91,9 @@ func loadBuiltInProfileIDs(t *testing.T) map[string]string {
if _, ok := raw["api_key"]; ok { if _, ok := raw["api_key"]; ok {
t.Fatalf("built-in profile %s contains raw api_key", name) t.Fatalf("built-in profile %s contains raw api_key", name)
} }
if raw["backend"] != backend.OpenRouterID { backendID, ok := raw["backend"].(string)
t.Fatalf("built-in profile %s does not select %q", name, backend.OpenRouterID) if !ok || !builtInBackendIDs[backendID] {
t.Fatalf("built-in profile %s does not select a maintained built-in: %#v", name, raw["backend"])
} }
if _, ok := raw["endpoint"]; ok { if _, ok := raw["endpoint"]; ok {
t.Fatalf("built-in profile %s repeats endpoint", name) t.Fatalf("built-in profile %s repeats endpoint", name)
@@ -91,53 +116,3 @@ func loadBuiltInProfileIDs(t *testing.T) map[string]string {
} }
return ids return ids
} }
func TestRepositoryWithPrimaryUsesPrimaryBeforeBuiltIns(t *testing.T) {
repo := NewRepositoryWithPrimary(staticProfileRepo{
profiles: map[string]string{"mistral-small-3": "custom-model"},
})
p, err := repo.GetProfile(context.Background(), "mistral-small-3")
if err != nil {
t.Fatalf("expected profile to load, got %v", err)
}
if p.Model != "custom-model" {
t.Fatalf("expected primary profile to override built-in, got %+v", p)
}
}
func TestRepositoryWithPrimaryFallsBackToBuiltIns(t *testing.T) {
repo := NewRepositoryWithPrimary(staticProfileRepo{})
p, err := repo.GetProfile(context.Background(), "mistral-small-3")
if err != nil {
t.Fatalf("expected built-in profile to load, got %v", err)
}
if p.ID != "mistral-small-3" {
t.Fatalf("unexpected profile: %+v", p)
}
}
func TestRepositoryWithPrimaryDoesNotFallBackAfterPrimaryError(t *testing.T) {
repo := NewRepositoryWithPrimary(staticProfileRepo{err: profile.ErrInvalidProfile})
_, err := repo.GetProfile(context.Background(), "mistral-small-3")
if !errors.Is(err, profile.ErrInvalidProfile) {
t.Fatalf("expected primary error, got %v", err)
}
}
type staticProfileRepo struct {
profiles map[string]string
err error
}
func (r staticProfileRepo) GetProfile(_ context.Context, id string) (*domain.ExecutionProfile, error) {
if r.err != nil {
return nil, r.err
}
if model, ok := r.profiles[id]; ok {
return &domain.ExecutionProfile{ID: id, Endpoint: "http://primary/v1", Model: model}, nil
}
return nil, profile.ErrProfileNotFound
}

View File

@@ -0,0 +1,47 @@
package profile
import (
"errors"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
// NormalizeAndValidateDefinition normalizes and validates one source-local
// profile definition without resolving a base profile.
func NormalizeAndValidateDefinition(profile *domain.ExecutionProfile) error {
if profile == nil {
return errors.New("profile is required")
}
profile.ID = strings.TrimSpace(profile.ID)
profile.BaseProfileID = strings.TrimSpace(profile.BaseProfileID)
profile.BackendID = strings.TrimSpace(profile.BackendID)
profile.Endpoint = strings.TrimSpace(profile.Endpoint)
if profile.ID == "" {
return errors.New("id is required")
}
if profile.Endpoint != "" {
endpoint, err := domain.NormalizeOpenAICompatibleBaseEndpoint(profile.Endpoint)
if err != nil {
return err
}
profile.Endpoint = endpoint
}
if profile.BaseProfileID == "" {
if profile.BackendID == "" && profile.Endpoint == "" {
return errors.New("backend or endpoint is required")
}
if strings.TrimSpace(profile.Model) == "" {
return errors.New("model is required")
}
}
return domain.ValidateExecutionTargetSettings(domain.ExecutionTarget{
Temperature: profile.Temperature,
MaxTokens: profile.MaxTokens,
TopP: profile.TopP,
TimeoutSeconds: profile.TimeoutSeconds,
})
}

View File

@@ -5,13 +5,14 @@ import (
"context" "context"
"errors" "errors"
"fmt" "fmt"
"io"
"io/fs" "io/fs"
"os" "os"
"path"
"strings" "strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/filecatalog" "gitea.maximumdirect.net/eric/promptkit/internal/filecatalog"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
"gopkg.in/yaml.v3" "gopkg.in/yaml.v3"
) )
@@ -73,7 +74,8 @@ func (r *overlayRepository) GetProfile(ctx context.Context, id string) (*domain.
} }
func loadProfile(ctx context.Context, fsys fs.FS, root string, id string) (*domain.ExecutionProfile, error) { func loadProfile(ctx context.Context, fsys fs.FS, root string, id string) (*domain.ExecutionProfile, error) {
if strings.TrimSpace(id) == "" { id = strings.TrimSpace(id)
if id == "" {
return nil, fmt.Errorf("%w: profile id is required", ErrInvalidProfile) return nil, fmt.Errorf("%w: profile id is required", ErrInvalidProfile)
} }
if fsys == nil { if fsys == nil {
@@ -94,42 +96,49 @@ func loadProfile(ctx context.Context, fsys fs.FS, root string, id string) (*doma
} }
relPath := filecatalog.DisplayPath(root, fullPath) relPath := filecatalog.DisplayPath(root, fullPath)
fileMatch := filecatalog.Stem(path.Base(fullPath)) == id
data, err := fs.ReadFile(fsys, fullPath) data, err := fs.ReadFile(fsys, fullPath)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to read profile file %s: %w", relPath, err) return nil, fmt.Errorf("failed to read profile file %s: %w", relPath, err)
} }
metadata := readProfileFileMetadata(data) metadata, metadataErr := readProfileFileMetadata(data)
idMatch := fileMatch || metadata.id == id idMatch := metadata.matchesID(id)
if metadataErr != nil {
if idMatch {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidYAML, relPath, metadataErr)
}
continue
}
if metadata.hasRawAPIKey { if metadata.hasRawAPIKey {
if idMatch { if idMatch {
return nil, fmt.Errorf("%w: %s", ErrRawAPIKeyNotAllowed, relPath) return nil, fmt.Errorf("%w: %s", ErrRawAPIKeyNotAllowed, relPath)
} }
continue continue
} }
if !idMatch {
var prof domain.ExecutionProfile
decoder := yaml.NewDecoder(bytes.NewReader(data))
decoder.KnownFields(true)
if err := decoder.Decode(&prof); err != nil {
if idMatch {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidYAML, relPath, err)
}
continue continue
} }
prof, err := decodeProfile(data)
if err != nil {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidYAML, relPath, err)
}
prof.ID = strings.TrimSpace(prof.ID)
if prof.ID != id { if prof.ID != id {
continue continue
} }
prof.BackendID = strings.TrimSpace(prof.BackendID) prof.ExtraParams, err = jsonvalue.CopyMap(prof.ExtraParams)
if err := validateProfile(&prof); err != nil { if err != nil {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidProfile, relPath, err)
}
if err := NormalizeAndValidateDefinition(prof); err != nil {
if errors.Is(err, ErrRawAPIKeyNotAllowed) { if errors.Is(err, ErrRawAPIKeyNotAllowed) {
return nil, fmt.Errorf("%w: %s", err, relPath) return nil, fmt.Errorf("%w: %s", err, relPath)
} }
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidProfile, relPath, err) return nil, fmt.Errorf("%w: %s: %v", ErrInvalidProfile, relPath, err)
} }
matches = append(matches, profileMatch{ matches = append(matches, profileMatch{
profile: &prof, profile: prof,
path: relPath, path: relPath,
}) })
} }
@@ -155,15 +164,36 @@ type profileMatch struct {
} }
type profileFileMetadata struct { type profileFileMetadata struct {
id string ids []string
hasRawAPIKey bool hasRawAPIKey bool
} }
func readProfileFileMetadata(data []byte) profileFileMetadata { func readProfileFileMetadata(data []byte) (profileFileMetadata, error) {
decoder := yaml.NewDecoder(bytes.NewReader(data))
var node yaml.Node var node yaml.Node
if err := yaml.NewDecoder(bytes.NewReader(data)).Decode(&node); err != nil { if err := decoder.Decode(&node); err != nil {
return profileFileMetadata{} return profileFileMetadata{}, err
} }
metadata := profileMetadataFromNode(&node)
documentCount := 1
for {
var trailing yaml.Node
err := decoder.Decode(&trailing)
if errors.Is(err, io.EOF) {
if documentCount == 1 {
return metadata, nil
}
return metadata, errors.New("profile file must contain exactly one YAML document")
}
if err != nil {
return metadata, err
}
documentCount++
metadata.merge(profileMetadataFromNode(&trailing))
}
}
func profileMetadataFromNode(node *yaml.Node) profileFileMetadata {
if node.Kind != yaml.DocumentNode || len(node.Content) == 0 { if node.Kind != yaml.DocumentNode || len(node.Content) == 0 {
return profileFileMetadata{} return profileFileMetadata{}
} }
@@ -178,7 +208,7 @@ func readProfileFileMetadata(data []byte) profileFileMetadata {
value := mapping.Content[i+1] value := mapping.Content[i+1]
switch key.Value { switch key.Value {
case "id": case "id":
metadata.id = strings.TrimSpace(value.Value) metadata.ids = append(metadata.ids, strings.TrimSpace(value.Value))
case "api_key": case "api_key":
metadata.hasRawAPIKey = true metadata.hasRawAPIKey = true
} }
@@ -186,29 +216,41 @@ func readProfileFileMetadata(data []byte) profileFileMetadata {
return metadata return metadata
} }
func validateProfile(p *domain.ExecutionProfile) error { func (m profileFileMetadata) matchesID(id string) bool {
if strings.TrimSpace(p.ID) == "" { for _, candidate := range m.ids {
return errors.New("id is required") if candidate == id {
return true
} }
if strings.TrimSpace(p.BackendID) == "" && strings.TrimSpace(p.Endpoint) == "" {
return errors.New("backend or endpoint is required")
} }
if strings.TrimSpace(p.Model) == "" { return false
return errors.New("model is required") }
}
func (m *profileFileMetadata) merge(other profileFileMetadata) {
if p.Temperature < 0 || p.Temperature > 2 { m.ids = append(m.ids, other.ids...)
return errors.New("temperature must be between 0 and 2") m.hasRawAPIKey = m.hasRawAPIKey || other.hasRawAPIKey
} }
if p.MaxTokens < 0 {
return errors.New("max_tokens must be greater than or equal to 0") func decodeProfile(data []byte) (*domain.ExecutionProfile, error) {
} var prof domain.ExecutionProfile
if p.TopP < 0 || p.TopP > 1 { decoder := yaml.NewDecoder(bytes.NewReader(data))
return errors.New("top_p must be between 0 and 1") decoder.KnownFields(true)
} if err := decoder.Decode(&prof); err != nil {
if p.TimeoutSeconds < 0 { return nil, err
return errors.New("timeout_seconds must be greater than or equal to 0") }
} if err := requireYAMLStreamEnd(decoder); err != nil {
return nil, err
return nil }
return &prof, nil
}
func requireYAMLStreamEnd(decoder *yaml.Decoder) error {
var trailing yaml.Node
err := decoder.Decode(&trailing)
if errors.Is(err, io.EOF) {
return nil
}
if err != nil {
return err
}
return errors.New("profile file must contain exactly one YAML document")
} }

View File

@@ -0,0 +1,58 @@
package profile
import (
"context"
"fmt"
"testing"
"testing/fstest"
)
func BenchmarkProfileRepositoryLookup(b *testing.B) {
for _, size := range []int{10, 1000} {
b.Run(fmt.Sprintf("catalog-%d", size), func(b *testing.B) {
files := fstest.MapFS{
"target.yaml": profileMapFile(`
id: target
endpoint: http://localhost:8000/v1
model: target-model
extra_params:
selected: true
`),
}
metadataNames := []string{"target.yaml"}
for i := 1; i < size; i++ {
name := fmt.Sprintf("profile-%04d.yaml", i)
files[name] = profileMapFile(fmt.Sprintf(`
id: profile-%04d
endpoint: http://localhost:8000/v1
model: unrelated-model
temperature: 0.5
max_tokens: 500
extra_params:
provider:
order:
- first
- second
`, i))
metadataNames = append(metadataNames, name)
}
fsys := &recordingProfileFS{FS: files}
repo := NewFSRepository(fsys, ".")
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
if _, err := repo.GetProfile(context.Background(), "target"); err != nil {
b.Fatal(err)
}
}
b.StopTimer()
for _, name := range metadataNames {
if got := fsys.openCount(name); got != b.N {
b.Fatalf("metadata %q opens = %d, want %d", name, got, b.N)
}
}
})
}
}

View File

@@ -2,11 +2,13 @@ package profile
import ( import (
"context" "context"
"encoding/json"
"errors" "errors"
"fmt"
"io/fs"
"os" "os"
"path/filepath" "path/filepath"
"strings" "strings"
"sync"
"testing" "testing"
"testing/fstest" "testing/fstest"
@@ -61,7 +63,7 @@ func TestFilesystemRepository_GetProfile(t *testing.T) {
wantErr bool wantErr bool
}{ }{
{name: "backend only", connection: "backend: ' openrouter '", wantBackend: "openrouter"}, {name: "backend only", connection: "backend: ' openrouter '", wantBackend: "openrouter"},
{name: "endpoint only", connection: "endpoint: http://localhost:8000/v1", wantEndpoint: "http://localhost:8000/v1"}, {name: "endpoint only", connection: "endpoint: ' https://localhost:8000/nested/v1 '", wantEndpoint: "https://localhost:8000/nested/v1"},
{name: "both", connection: "backend: openrouter\nendpoint: http://localhost:8000/v1", wantBackend: "openrouter", wantEndpoint: "http://localhost:8000/v1"}, {name: "both", connection: "backend: openrouter\nendpoint: http://localhost:8000/v1", wantBackend: "openrouter", wantEndpoint: "http://localhost:8000/v1"},
{name: "neither", wantErr: true}, {name: "neither", wantErr: true},
{name: "blank backend", connection: "backend: ' '", wantErr: true}, {name: "blank backend", connection: "backend: ' '", wantErr: true},
@@ -126,63 +128,6 @@ temperature: 0.1
} }
}) })
t.Run("valid profile with JSON-compatible extra params", func(t *testing.T) {
writeProfileTestFile(t, filepath.Join(tmpDir, "json-extra-params.yaml"), `
id: json-extra-params
endpoint: http://localhost:8000/v1
model: nested-model
extra_params:
string_value: enabled
number_value: 42
boolean_value: true
object_value:
nested: value
count: 2
array_value:
- first
- 3
- false
`)
p, err := repo.GetProfile(ctx, "json-extra-params")
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
var got map[string]any
encoded, err := json.Marshal(p.ExtraParams)
if err != nil {
t.Fatalf("expected extra_params to marshal as JSON, got %v", err)
}
if err := json.Unmarshal(encoded, &got); err != nil {
t.Fatalf("expected extra_params JSON to decode, got %v", err)
}
if got["string_value"] != "enabled" {
t.Fatalf("unexpected string extra param: %#v", got["string_value"])
}
if got["number_value"] != float64(42) {
t.Fatalf("unexpected number extra param: %#v", got["number_value"])
}
if got["boolean_value"] != true {
t.Fatalf("unexpected boolean extra param: %#v", got["boolean_value"])
}
objectValue, ok := got["object_value"].(map[string]any)
if !ok {
t.Fatalf("expected object extra param, got %#v", got["object_value"])
}
if objectValue["nested"] != "value" || objectValue["count"] != float64(2) {
t.Fatalf("unexpected object extra param: %#v", objectValue)
}
arrayValue, ok := got["array_value"].([]any)
if !ok {
t.Fatalf("expected array extra param, got %#v", got["array_value"])
}
if len(arrayValue) != 3 || arrayValue[0] != "first" || arrayValue[1] != float64(3) || arrayValue[2] != false {
t.Fatalf("unexpected array extra param: %#v", arrayValue)
}
})
t.Run("duplicate profile IDs fail as ambiguous", func(t *testing.T) { t.Run("duplicate profile IDs fail as ambiguous", func(t *testing.T) {
writeProfileTestFile(t, filepath.Join(tmpDir, "duplicate-profile-a.yaml"), ` writeProfileTestFile(t, filepath.Join(tmpDir, "duplicate-profile-a.yaml"), `
id: duplicate-profile id: duplicate-profile
@@ -245,10 +190,10 @@ api_key: secret
} }
}) })
t.Run("invalid yaml", func(t *testing.T) { t.Run("unidentifiable invalid yaml is unrelated", func(t *testing.T) {
_, err := repo.GetProfile(ctx, "invalid_yaml") _, err := repo.GetProfile(ctx, "invalid_yaml")
if !errors.Is(err, ErrInvalidYAML) { if !errors.Is(err, ErrProfileNotFound) {
t.Fatalf("expected ErrInvalidYAML, got %v", err) t.Fatalf("expected ErrProfileNotFound, got %v", err)
} }
}) })
@@ -274,14 +219,14 @@ api_key: secret
}) })
t.Run("unknown field", func(t *testing.T) { t.Run("unknown field", func(t *testing.T) {
_, err := repo.GetProfile(ctx, "unknown_field") _, err := repo.GetProfile(ctx, "unknown-field")
if !errors.Is(err, ErrInvalidYAML) { if !errors.Is(err, ErrInvalidYAML) {
t.Fatalf("expected ErrInvalidYAML for strict decode unknown field, got %v", err) t.Fatalf("expected ErrInvalidYAML for strict decode unknown field, got %v", err)
} }
}) })
t.Run("raw api_key rejected", func(t *testing.T) { t.Run("raw api_key rejected", func(t *testing.T) {
_, err := repo.GetProfile(ctx, "raw_api_key") _, err := repo.GetProfile(ctx, "raw-api-key")
if !errors.Is(err, ErrRawAPIKeyNotAllowed) { if !errors.Is(err, ErrRawAPIKeyNotAllowed) {
t.Fatalf("expected ErrRawAPIKeyNotAllowed, got %v", err) t.Fatalf("expected ErrRawAPIKeyNotAllowed, got %v", err)
} }
@@ -406,6 +351,614 @@ model: second
}) })
} }
func TestProfileRepositoriesRejectInvalidEndpoints(t *testing.T) {
tests := []struct {
name string
endpoint string
withBackend bool
}{
{name: "relative", endpoint: "/v1"},
{name: "missing host", endpoint: "https:///v1"},
{name: "unsupported scheme", endpoint: "ftp://provider.example/v1"},
{name: "user information", endpoint: "https://user@provider.example/v1"},
{name: "query", endpoint: "https://provider.example/v1?mode=chat"},
{name: "fragment", endpoint: "https://provider.example/v1#chat"},
{name: "backend with invalid override", endpoint: "/v1", withBackend: true},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
backend := ""
if tc.withBackend {
backend = "backend: openrouter\n"
}
repo := NewFSRepository(fstest.MapFS{
"profiles/invalid.yaml": profileMapFile(fmt.Sprintf(
"id: invalid-endpoint\nmodel: model\n%sendpoint: %q\n",
backend,
tc.endpoint,
)),
}, "profiles")
_, err := repo.GetProfile(context.Background(), "invalid-endpoint")
if !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("expected ErrInvalidProfile, got %v", err)
}
})
}
}
func TestProfileRepositoriesValidateExtraParams(t *testing.T) {
const validProfile = `
id: selected-profile
endpoint: http://localhost:8000/v1
model: model
extra_params:
string_value: enabled
object_value:
nested: true
array_value:
- first
- 3
`
tests := []struct {
name string
definition string
wantErr bool
diagnostics []string
}{
{name: "valid nested values", definition: validProfile},
{
name: "empty key",
definition: `
id: selected-profile
endpoint: http://localhost:8000/v1
model: model
extra_params:
"": value
`,
wantErr: true,
diagnostics: []string{"extra_params", "key must not be empty"},
},
{
name: "non-finite value",
definition: `
id: selected-profile
endpoint: http://localhost:8000/v1
model: model
extra_params:
invalid: .nan
`,
wantErr: true,
diagnostics: []string{"extra_params.invalid", "must be finite"},
},
{
name: "nested non-finite value",
definition: `
id: selected-profile
endpoint: http://localhost:8000/v1
model: model
extra_params:
outer:
invalid: .inf
`,
wantErr: true,
diagnostics: []string{"extra_params.outer.invalid", "must be finite"},
},
{
name: "unsupported decoded value",
definition: `
id: selected-profile
endpoint: http://localhost:8000/v1
model: model
extra_params:
timestamp: 2026-08-11T12:34:56Z
`,
wantErr: true,
diagnostics: []string{"extra_params.timestamp", "unsupported JSON value type"},
},
{
name: "excessive nesting",
definition: deeplyNestedExtraParamsProfile(101),
wantErr: true,
diagnostics: []string{"extra_params", "JSON container depth limit exceeded"},
},
}
for _, source := range profileRepositorySources() {
for _, tc := range tests {
t.Run(source.name+"/"+tc.name, func(t *testing.T) {
repo := source.newRepository(t, map[string]string{"selected.yaml": tc.definition})
got, err := repo.GetProfile(context.Background(), "selected-profile")
if tc.wantErr {
if !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("expected ErrInvalidProfile, got %v", err)
}
if !strings.Contains(err.Error(), "selected.yaml") {
t.Fatalf("expected source path in error, got %v", err)
}
for _, diagnostic := range tc.diagnostics {
if !strings.Contains(err.Error(), diagnostic) {
t.Fatalf("expected error to contain %q, got %v", diagnostic, err)
}
}
return
}
if err != nil {
t.Fatalf("load valid profile: %v", err)
}
if got.ExtraParams["string_value"] != "enabled" {
t.Fatalf("unexpected copied extra params: %#v", got.ExtraParams)
}
objectValue, objectOK := got.ExtraParams["object_value"].(map[string]any)
arrayValue, arrayOK := got.ExtraParams["array_value"].([]any)
if !objectOK || objectValue["nested"] != true ||
!arrayOK || len(arrayValue) != 2 || arrayValue[0] != "first" || arrayValue[1] != 3 {
t.Fatalf("unexpected copied nested extra params: %#v", got.ExtraParams)
}
})
}
}
}
func TestProfileRepositoriesSelectCanonicalYAMLID(t *testing.T) {
const validProfile = `
id: selected-profile
endpoint: http://localhost:8000/v1
model: selected-model
`
tests := []struct {
name string
files map[string]string
wantErr error
diagnostics []string
}{
{
name: "same stem unknown field with different id is unrelated",
files: map[string]string{
"selected-profile.yaml": `
id: unrelated-profile
endpoint: http://localhost:8000/v1
model: unrelated
unknown: true
`,
"valid.yaml": validProfile,
},
},
{
name: "same stem unidentifiable yaml is unrelated",
files: map[string]string{
"selected-profile.yaml": "id: [",
"valid.yaml": validProfile,
},
},
{
name: "same stem raw key with different id is unrelated",
files: map[string]string{
"selected-profile.yaml": `
id: unrelated-profile
endpoint: http://localhost:8000/v1
model: unrelated
api_key: secret
`,
"valid.yaml": validProfile,
},
},
{
name: "leading and trailing whitespace is normalized",
files: map[string]string{
"padded.yaml": `
id: " selected-profile "
endpoint: http://localhost:8000/v1
model: selected-model
`,
},
},
{
name: "blank id is unrelated",
files: map[string]string{
"selected-profile.yaml": `
id: " "
endpoint: http://localhost:8000/v1
model: unrelated
`,
},
wantErr: ErrProfileNotFound,
},
{
name: "normalized duplicates are ambiguous",
files: map[string]string{
"first.yaml": validProfile,
"nested/second.yaml": `
id: " selected-profile "
endpoint: http://localhost:8000/v1
model: duplicate
`,
},
wantErr: ErrInvalidProfile,
diagnostics: []string{"duplicate execution profile id", "first.yaml", "nested/second.yaml"},
},
{
name: "selected unknown field is authoritative",
files: map[string]string{
"malformed.yaml": `
id: selected-profile
endpoint: http://localhost:8000/v1
model: selected-model
unknown: true
`,
},
wantErr: ErrInvalidYAML,
diagnostics: []string{"malformed.yaml"},
},
{
name: "selected raw key is authoritative",
files: map[string]string{
"insecure.yaml": `
id: selected-profile
endpoint: http://localhost:8000/v1
model: selected-model
api_key: secret
`,
},
wantErr: ErrRawAPIKeyNotAllowed,
diagnostics: []string{"insecure.yaml"},
},
{
name: "selected identity in an additional document is authoritative",
files: map[string]string{
"additional-document.yaml": `
---
---
id: selected-profile
endpoint: http://localhost:8000/v1
model: selected-model
`,
},
wantErr: ErrInvalidYAML,
diagnostics: []string{"additional-document.yaml", "exactly one YAML document"},
},
}
for _, source := range profileRepositorySources() {
for _, tc := range tests {
t.Run(source.name+"/"+tc.name, func(t *testing.T) {
repo := source.newRepository(t, tc.files)
got, err := repo.GetProfile(context.Background(), " selected-profile ")
if tc.wantErr != nil {
if !errors.Is(err, tc.wantErr) {
t.Fatalf("expected %v, got %v", tc.wantErr, err)
}
for _, diagnostic := range tc.diagnostics {
if !strings.Contains(err.Error(), diagnostic) {
t.Fatalf("expected error to contain %q, got %v", diagnostic, err)
}
}
return
}
if err != nil {
t.Fatalf("load selected profile: %v", err)
}
if got.ID != "selected-profile" || got.Model != "selected-model" {
t.Fatalf("unexpected selected profile: %+v", got)
}
})
}
}
}
func TestProfileRepositoriesRequireOneYAMLDocument(t *testing.T) {
const profile = `
id: selected-profile
endpoint: http://localhost:8000/v1
model: selected-model
`
tests := []struct {
name string
suffix string
wantErr bool
}{
{name: "comments and trailing whitespace", suffix: "\n# trailing comment\n\n"},
{name: "second populated document", suffix: "\n---\nid: another\n", wantErr: true},
{name: "second empty document", suffix: "\n---\n", wantErr: true},
{name: "malformed trailing yaml", suffix: "\n---\n[", wantErr: true},
{name: "raw key in trailing document", suffix: "\n---\napi_key: secret\n", wantErr: true},
}
for _, source := range profileRepositorySources() {
for _, tc := range tests {
t.Run(source.name+"/"+tc.name, func(t *testing.T) {
repo := source.newRepository(t, map[string]string{"definition.yaml": profile + tc.suffix})
got, err := repo.GetProfile(context.Background(), "selected-profile")
if tc.wantErr {
if !errors.Is(err, ErrInvalidYAML) {
t.Fatalf("expected ErrInvalidYAML, got %v", err)
}
if !strings.Contains(err.Error(), "definition.yaml") {
t.Fatalf("expected source path in error, got %v", err)
}
return
}
if err != nil {
t.Fatalf("load one-document profile: %v", err)
}
if got.ID != "selected-profile" {
t.Fatalf("unexpected profile: %+v", got)
}
})
}
}
}
func TestProfileRepositoriesPreserveOverlayFallbackRules(t *testing.T) {
fallback := staticProfileRepo{profiles: map[string]*domain.ExecutionProfile{
"selected-profile": {ID: "selected-profile", Endpoint: "http://fallback", Model: "fallback-model"},
}}
tests := []struct {
name string
files map[string]string
wantModel string
wantErr error
}{
{
name: "same stem malformed different id falls back",
files: map[string]string{
"selected-profile.yaml": `
id: unrelated-profile
endpoint: http://localhost:8000/v1
model: unrelated
unknown: true
`,
},
wantModel: "fallback-model",
},
{
name: "blank id falls back",
files: map[string]string{
"selected-profile.yaml": `
id: " "
endpoint: http://localhost:8000/v1
model: unrelated
`,
},
wantModel: "fallback-model",
},
{
name: "selected malformed profile stops fallback",
files: map[string]string{
"other-name.yaml": `
id: selected-profile
endpoint: http://localhost:8000/v1
model: selected
unknown: true
`,
},
wantErr: ErrInvalidYAML,
},
}
for _, source := range profileRepositorySources() {
for _, tc := range tests {
t.Run(source.name+"/"+tc.name, func(t *testing.T) {
primary := source.newRepository(t, tc.files)
got, err := NewOverlayRepository(primary, fallback).GetProfile(context.Background(), "selected-profile")
if tc.wantErr != nil {
if !errors.Is(err, tc.wantErr) {
t.Fatalf("expected %v, got %v", tc.wantErr, err)
}
return
}
if err != nil {
t.Fatalf("load fallback profile: %v", err)
}
if got.Model != tc.wantModel {
t.Fatalf("model = %q, want %q", got.Model, tc.wantModel)
}
})
}
}
}
func TestProfileRepositoryReadsSourcesFreshOnEveryLookup(t *testing.T) {
newSource := func() (*recordingProfileFS, Repository) {
fsys := &recordingProfileFS{FS: fstest.MapFS{
"target.yaml": profileMapFile(`
id: target
endpoint: http://localhost:8000/v1
model: target-model
`),
"unrelated.yaml": profileMapFile(`
id: unrelated
endpoint: http://localhost:8000/v1
model: unrelated-model
`),
}}
return fsys, NewFSRepository(fsys, ".")
}
t.Run("selected source", func(t *testing.T) {
fsys, repo := newSource()
for lookup := 1; lookup <= 2; lookup++ {
got, err := repo.GetProfile(context.Background(), "target")
if err != nil {
t.Fatalf("lookup %d: %v", lookup, err)
}
if got.Model != "target-model" {
t.Fatalf("lookup %d model = %q", lookup, got.Model)
}
for _, name := range []string{"target.yaml", "unrelated.yaml"} {
if count := fsys.openCount(name); count != lookup {
t.Fatalf("%s opens after lookup %d = %d, want %d", name, lookup, count, lookup)
}
}
}
})
t.Run("overlay fallthrough", func(t *testing.T) {
primaryFS := &recordingProfileFS{FS: fstest.MapFS{
"unrelated.yaml": profileMapFile(`
id: unrelated
endpoint: http://localhost:8000/v1
model: unrelated-model
`),
}}
fallbackFS, fallback := newSource()
repo := NewOverlayRepository(NewFSRepository(primaryFS, "."), fallback)
for lookup := 1; lookup <= 2; lookup++ {
got, err := repo.GetProfile(context.Background(), "target")
if err != nil {
t.Fatalf("lookup %d: %v", lookup, err)
}
if got.Model != "target-model" {
t.Fatalf("lookup %d model = %q", lookup, got.Model)
}
if count := primaryFS.openCount("unrelated.yaml"); count != lookup {
t.Fatalf("primary opens after lookup %d = %d, want %d", lookup, count, lookup)
}
if count := fallbackFS.openCount("target.yaml"); count != lookup {
t.Fatalf("fallback opens after lookup %d = %d, want %d", lookup, count, lookup)
}
}
})
}
func TestProfileRepositoriesRejectInvalidExecutionSettings(t *testing.T) {
ctx := context.Background()
t.Run("operating-system filesystem", func(t *testing.T) {
dir := t.TempDir()
writeProfileTestFile(t, filepath.Join(dir, "invalid.yaml"), `
id: invalid
endpoint: http://localhost:8000/v1
model: model
temperature: .nan
`)
_, err := NewFilesystemRepository(dir).GetProfile(ctx, "invalid")
if !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("expected ErrInvalidProfile, got %v", err)
}
})
t.Run("fs.FS", func(t *testing.T) {
repo := NewFSRepository(fstest.MapFS{
"profiles/invalid.yaml": profileMapFile(`
id: invalid
endpoint: http://localhost:8000/v1
model: model
top_p: .inf
`),
}, "profiles")
_, err := repo.GetProfile(ctx, "invalid")
if !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("expected ErrInvalidProfile, got %v", err)
}
})
}
func TestProfileRepositoriesValidateDerivedDefinitions(t *testing.T) {
tests := []struct {
name string
files map[string]string
wantError error
wantBaseID string
wantProfile bool
}{
{
name: "alias is locally valid and normalizes base id",
files: map[string]string{"alias.yaml": `
id: selected-profile
base_profile: " base-profile "
`},
wantBaseID: "base-profile",
wantProfile: true,
},
{
name: "derived endpoint remains valid",
files: map[string]string{"invalid.yaml": `
id: selected-profile
base_profile: base-profile
endpoint: /v1
`},
wantError: ErrInvalidProfile,
},
{
name: "derived settings remain valid",
files: map[string]string{"invalid.yaml": `
id: selected-profile
base_profile: base-profile
top_p: 1.1
`},
wantError: ErrInvalidProfile,
},
{
name: "derived extra params remain valid",
files: map[string]string{"invalid.yaml": `
id: selected-profile
base_profile: base-profile
extra_params:
timestamp: 2026-08-11T12:34:56Z
`},
wantError: ErrInvalidProfile,
},
{
name: "derived raw key remains prohibited",
files: map[string]string{"invalid.yaml": `
id: selected-profile
base_profile: base-profile
api_key: secret
`},
wantError: ErrRawAPIKeyNotAllowed,
},
{
name: "derived duplicate id remains invalid",
files: map[string]string{
"first.yaml": "id: selected-profile\nbase_profile: first-base\n",
"second.yaml": "id: selected-profile\nbase_profile: second-base\n",
},
wantError: ErrInvalidProfile,
},
{
name: "derived extra document remains invalid",
files: map[string]string{"invalid.yaml": `
id: selected-profile
base_profile: base-profile
---
id: other
`},
wantError: ErrInvalidYAML,
},
{
name: "standalone profile remains complete",
files: map[string]string{"invalid.yaml": "id: selected-profile\n"},
wantError: ErrInvalidProfile,
},
}
for _, source := range profileRepositorySources() {
for _, tc := range tests {
t.Run(source.name+"/"+tc.name, func(t *testing.T) {
repo := source.newRepository(t, tc.files)
got, err := repo.GetProfile(context.Background(), "selected-profile")
if tc.wantError != nil {
if !errors.Is(err, tc.wantError) {
t.Fatalf("error = %v, want %v", err, tc.wantError)
}
return
}
if err != nil || !tc.wantProfile {
t.Fatalf("profile = %+v, error = %v, want valid derived definition", got, err)
}
if got.BaseProfileID != tc.wantBaseID {
t.Fatalf("BaseProfileID = %q, want %q", got.BaseProfileID, tc.wantBaseID)
}
})
}
}
}
func TestOverlayRepository(t *testing.T) { func TestOverlayRepository(t *testing.T) {
ctx := context.Background() ctx := context.Background()
primaryProfile := &domain.ExecutionProfile{ID: "shared", Endpoint: "http://primary", Model: "primary"} primaryProfile := &domain.ExecutionProfile{ID: "shared", Endpoint: "http://primary", Model: "primary"}
@@ -495,6 +1048,77 @@ func TestOverlayRepository(t *testing.T) {
}) })
} }
type profileRepositorySource struct {
name string
newRepository func(t *testing.T, files map[string]string) Repository
}
type recordingProfileFS struct {
fs.FS
mu sync.Mutex
opened []string
}
func (f *recordingProfileFS) Open(name string) (fs.File, error) {
f.mu.Lock()
f.opened = append(f.opened, name)
f.mu.Unlock()
return f.FS.Open(name)
}
func (f *recordingProfileFS) openCount(name string) int {
f.mu.Lock()
defer f.mu.Unlock()
count := 0
for _, opened := range f.opened {
if opened == name {
count++
}
}
return count
}
func profileRepositorySources() []profileRepositorySource {
return []profileRepositorySource{
{
name: "operating system",
newRepository: func(t *testing.T, files map[string]string) Repository {
t.Helper()
root := t.TempDir()
for name, content := range files {
filePath := filepath.Join(root, filepath.FromSlash(name))
if err := os.MkdirAll(filepath.Dir(filePath), 0o755); err != nil {
t.Fatalf("create profile directory: %v", err)
}
writeProfileTestFile(t, filePath, content)
}
return NewFilesystemRepository(root)
},
},
{
name: "filesystem",
newRepository: func(t *testing.T, files map[string]string) Repository {
t.Helper()
fsys := make(fstest.MapFS, len(files))
for name, content := range files {
fsys[name] = profileMapFile(content)
}
return NewFSRepository(fsys, ".")
},
},
}
}
func deeplyNestedExtraParamsProfile(depth int) string {
var definition strings.Builder
definition.WriteString("id: selected-profile\nendpoint: http://localhost:8000/v1\nmodel: model\nextra_params:\n")
for level := 0; level < depth; level++ {
fmt.Fprintf(&definition, "%slevel_%d:\n", strings.Repeat(" ", level+1), level)
}
fmt.Fprintf(&definition, "%svalue: true\n", strings.Repeat(" ", depth+1))
return definition.String()
}
func profileMapFile(content string) *fstest.MapFile { func profileMapFile(content string) *fstest.MapFile {
return &fstest.MapFile{Data: []byte(strings.TrimLeft(content, "\n"))} return &fstest.MapFile{Data: []byte(strings.TrimLeft(content, "\n"))}
} }

View File

@@ -0,0 +1,160 @@
package profile
import (
"context"
"errors"
"fmt"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
)
const maximumProfileChainLength = 32
type resolvingRepository struct {
source Repository
}
// NewResolvingRepository resolves inherited profile definitions from source.
func NewResolvingRepository(source Repository) Repository {
return &resolvingRepository{source: source}
}
func (r *resolvingRepository) GetProfile(ctx context.Context, id string) (*domain.ExecutionProfile, error) {
if r == nil || r.source == nil {
return nil, fmt.Errorf("%w: profile repository is required", ErrInvalidProfile)
}
requestedID := strings.TrimSpace(id)
if requestedID == "" {
return nil, fmt.Errorf("%w: profile id is required", ErrInvalidProfile)
}
if err := ctx.Err(); err != nil {
return nil, err
}
profile, err := r.getRawProfile(ctx, requestedID)
if err != nil {
return nil, err
}
if profile == nil {
return nil, fmt.Errorf("%w: selected profile %q is nil", ErrInvalidProfile, requestedID)
}
chain := []*domain.ExecutionProfile{profile}
chainIDs := []string{requestedID}
visited := map[string]struct{}{requestedID: {}}
current := profile
for {
baseID := strings.TrimSpace(current.BaseProfileID)
if baseID == "" {
break
}
if err := ctx.Err(); err != nil {
return nil, err
}
if _, seen := visited[baseID]; seen {
return nil, fmt.Errorf("%w: profile inheritance cycle %s", ErrInvalidProfile, joinProfileChain(chainIDs, baseID))
}
if len(chain) >= maximumProfileChainLength {
return nil, fmt.Errorf("%w: profile inheritance chain exceeds %d profiles: %s", ErrInvalidProfile, maximumProfileChainLength, joinProfileChain(chainIDs, baseID))
}
base, err := r.getRawProfile(ctx, baseID)
if err != nil {
if errors.Is(err, ErrProfileNotFound) {
return nil, fmt.Errorf("%w: base profile %q is missing in chain %s", ErrInvalidProfile, baseID, joinProfileChain(chainIDs, baseID))
}
return nil, fmt.Errorf("%w: failed to load base profile %q in chain %s: %w", ErrInvalidProfile, baseID, joinProfileChain(chainIDs, baseID), err)
}
if base == nil {
return nil, fmt.Errorf("%w: base profile %q is nil in chain %s", ErrInvalidProfile, baseID, joinProfileChain(chainIDs, baseID))
}
chain = append(chain, base)
chainIDs = append(chainIDs, baseID)
visited[baseID] = struct{}{}
current = base
}
resolved, err := mergeProfileChain(chain)
if err != nil {
return nil, fmt.Errorf("%w: resolved profile chain %s: %w", ErrInvalidProfile, strings.Join(chainIDs, " -> "), err)
}
if err := validateResolvedProfile(resolved); err != nil {
return nil, fmt.Errorf("%w: resolved profile chain %s: %w", ErrInvalidProfile, strings.Join(chainIDs, " -> "), err)
}
return resolved, nil
}
func (r *resolvingRepository) getRawProfile(ctx context.Context, id string) (*domain.ExecutionProfile, error) {
profile, err := r.source.GetProfile(ctx, id)
if err != nil {
return nil, err
}
if err := ctx.Err(); err != nil {
return nil, err
}
return profile, nil
}
func joinProfileChain(chain []string, next string) string {
return strings.Join(append(append([]string(nil), chain...), next), " -> ")
}
func mergeProfileChain(chain []*domain.ExecutionProfile) (*domain.ExecutionProfile, error) {
resolved := &domain.ExecutionProfile{ID: chain[0].ID}
for index := len(chain) - 1; index >= 0; index-- {
definition := chain[index]
if strings.TrimSpace(definition.BackendID) != "" {
resolved.BackendID = definition.BackendID
}
if strings.TrimSpace(definition.Endpoint) != "" {
resolved.Endpoint = definition.Endpoint
}
if strings.TrimSpace(definition.Model) != "" {
resolved.Model = definition.Model
}
if definition.Temperature != 0 {
resolved.Temperature = definition.Temperature
}
if definition.MaxTokens != 0 {
resolved.MaxTokens = definition.MaxTokens
}
if definition.TopP != 0 {
resolved.TopP = definition.TopP
}
if definition.TimeoutSeconds != 0 {
resolved.TimeoutSeconds = definition.TimeoutSeconds
}
if strings.TrimSpace(definition.ServiceTier) != "" {
resolved.ServiceTier = definition.ServiceTier
}
if strings.TrimSpace(definition.ReasoningEffort) != "" {
resolved.ReasoningEffort = definition.ReasoningEffort
}
if strings.TrimSpace(definition.APIKeyEnv) != "" {
resolved.APIKeyEnv = definition.APIKeyEnv
}
resolved.APIKeyRequired = resolved.APIKeyRequired || definition.APIKeyRequired
if len(definition.ExtraParams) != 0 {
extraParams, err := jsonvalue.CopyMap(definition.ExtraParams)
if err != nil {
return nil, err
}
resolved.ExtraParams = extraParams
}
}
resolved.ID = chain[0].ID
resolved.BaseProfileID = ""
return resolved, nil
}
func validateResolvedProfile(profile *domain.ExecutionProfile) error {
if profile == nil {
return errors.New("resolved profile is required")
}
profile.BaseProfileID = ""
return NormalizeAndValidateDefinition(profile)
}

View File

@@ -0,0 +1,418 @@
package profile
import (
"context"
"errors"
"fmt"
"io/fs"
"reflect"
"strings"
"sync"
"testing"
"testing/fstest"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
func TestResolvingRepositoryMergesProfileChain(t *testing.T) {
repo := &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"leaf": {
ID: "leaf",
BaseProfileID: "middle",
BackendID: "leaf-backend",
TopP: 0.8,
TimeoutSeconds: 45,
ReasoningEffort: "high",
},
"middle": {
ID: "middle",
BaseProfileID: "root",
Endpoint: "https://middle.example/v1",
Model: "middle-model",
MaxTokens: 256,
APIKeyEnv: "MIDDLE_API_KEY",
APIKeyRequired: true,
ExtraParams: map[string]any{"middle": map[string]any{"value": "middle"}},
},
"root": {
ID: "root",
BackendID: "root-backend",
Endpoint: "https://root.example/v1",
Model: "root-model",
Temperature: 0.3,
ServiceTier: "priority",
ExtraParams: map[string]any{"root": "value"},
},
}}
got, err := NewResolvingRepository(repo).GetProfile(context.Background(), "leaf")
if err != nil {
t.Fatalf("resolve profile: %v", err)
}
want := &domain.ExecutionProfile{
ID: "leaf",
BackendID: "leaf-backend",
Endpoint: "https://middle.example/v1",
Model: "middle-model",
Temperature: 0.3,
MaxTokens: 256,
TopP: 0.8,
TimeoutSeconds: 45,
ServiceTier: "priority",
ReasoningEffort: "high",
APIKeyEnv: "MIDDLE_API_KEY",
APIKeyRequired: true,
ExtraParams: map[string]any{"middle": map[string]any{"value": "middle"}},
}
if !reflect.DeepEqual(got, want) {
t.Fatalf("resolved profile:\n got %#v\nwant %#v", got, want)
}
}
func TestResolvingRepositoryRejectsMissingSourceAndProfileID(t *testing.T) {
if _, err := NewResolvingRepository(nil).GetProfile(context.Background(), "profile"); !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("nil source error = %v, want ErrInvalidProfile", err)
}
repo := &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{}}
if _, err := NewResolvingRepository(repo).GetProfile(context.Background(), " \t "); !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("blank id error = %v, want ErrInvalidProfile", err)
}
if got := repo.callCount(" "); got != 0 {
t.Fatalf("blank id looked up source %d times", got)
}
}
func TestResolvingRepositoryCopiesExtraParams(t *testing.T) {
baseParams := map[string]any{"nested": map[string]any{"value": "base"}}
repo := &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"child": {ID: "child", BaseProfileID: "base"},
"base": {
ID: "base",
Endpoint: "https://base.example/v1",
Model: "model",
ExtraParams: baseParams,
},
}}
resolver := NewResolvingRepository(repo)
first, err := resolver.GetProfile(context.Background(), "child")
if err != nil {
t.Fatalf("resolve inherited map: %v", err)
}
first.ExtraParams["nested"].(map[string]any)["value"] = "mutated"
second, err := resolver.GetProfile(context.Background(), "child")
if err != nil {
t.Fatalf("resolve inherited map again: %v", err)
}
if got := second.ExtraParams["nested"].(map[string]any)["value"]; got != "base" {
t.Fatalf("later result retained mutation: %v", got)
}
if got := baseParams["nested"].(map[string]any)["value"]; got != "base" {
t.Fatalf("source map retained mutation: %v", got)
}
repo.set("child", &domain.ExecutionProfile{
ID: "child",
BaseProfileID: "base",
ExtraParams: map[string]any{"child": "replacement"},
})
replaced, err := resolver.GetProfile(context.Background(), "child")
if err != nil {
t.Fatalf("resolve replacement map: %v", err)
}
if !reflect.DeepEqual(replaced.ExtraParams, map[string]any{"child": "replacement"}) {
t.Fatalf("extra params = %#v, want complete child replacement", replaced.ExtraParams)
}
}
func TestResolvingRepositoryUsesRawOverlayForEachLookup(t *testing.T) {
leafSource := NewFSRepository(profileTestFS(map[string]string{
"leaf.yaml": "id: leaf\nbase_profile: base\n",
}), ".")
fallback := NewFSRepository(profileTestFS(map[string]string{
"base.yaml": "id: base\nendpoint: https://fallback.example/v1\nmodel: fallback-model\n",
}), ".")
overlay := NewOverlayRepository(leafSource, fallback)
resolver := NewResolvingRepository(overlay)
got, err := resolver.GetProfile(context.Background(), "leaf")
if err != nil {
t.Fatalf("resolve fallback base: %v", err)
}
if got.Model != "fallback-model" {
t.Fatalf("fallback base model = %q", got.Model)
}
shadowing := NewOverlayRepository(NewFSRepository(profileTestFS(map[string]string{
"leaf.yaml": "id: leaf\nbase_profile: base\n",
"base.yaml": "id: base\nendpoint: https://primary.example/v1\nmodel: primary-model\n",
}), "."), fallback)
got, err = NewResolvingRepository(shadowing).GetProfile(context.Background(), "leaf")
if err != nil {
t.Fatalf("resolve shadowed base: %v", err)
}
if got.Model != "primary-model" || got.Endpoint != "https://primary.example/v1" {
t.Fatalf("shadowed base = %+v", got)
}
}
func TestResolvingRepositoryReportsSafetyAndSourceErrors(t *testing.T) {
sourceErr := errors.New("source failure")
tests := []struct {
name string
repo *resolvingTestRepository
id string
want []error
wantNot error
contains []string
}{
{
name: "missing selected profile preserves not found",
repo: &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{}},
id: "missing",
want: []error{ErrProfileNotFound},
wantNot: ErrInvalidProfile,
},
{
name: "missing base is invalid but not not found",
repo: &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"leaf": {ID: "leaf", BaseProfileID: "missing"},
}},
id: "leaf",
want: []error{ErrInvalidProfile},
wantNot: ErrProfileNotFound,
contains: []string{"missing", "leaf -> missing"},
},
{
name: "direct cycle",
repo: &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"a": {ID: "a", BaseProfileID: "a"},
}},
id: "a",
want: []error{ErrInvalidProfile},
contains: []string{"a -> a"},
},
{
name: "indirect cycle",
repo: &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"a": {ID: "a", BaseProfileID: "b"},
"b": {ID: "b", BaseProfileID: "c"},
"c": {ID: "c", BaseProfileID: "a"},
}},
id: "a",
want: []error{ErrInvalidProfile},
contains: []string{"a -> b -> c -> a"},
},
{
name: "nil result",
repo: &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"leaf": nil,
}},
id: "leaf",
want: []error{ErrInvalidProfile},
},
{
name: "incomplete resolved profile",
repo: &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"leaf": {ID: "leaf", BaseProfileID: "base"},
"base": {ID: "base", Model: "model"},
}},
id: "leaf",
want: []error{ErrInvalidProfile},
},
{
name: "base source error is retained",
repo: &resolvingTestRepository{
profiles: map[string]*domain.ExecutionProfile{"leaf": {ID: "leaf", BaseProfileID: "base"}},
errors: map[string]error{"base": sourceErr},
},
id: "leaf",
want: []error{ErrInvalidProfile, sourceErr},
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
_, err := NewResolvingRepository(tc.repo).GetProfile(context.Background(), tc.id)
for _, want := range tc.want {
if !errors.Is(err, want) {
t.Fatalf("error = %v, want %v", err, want)
}
}
if tc.wantNot != nil && errors.Is(err, tc.wantNot) {
t.Fatalf("error = %v, must not match %v", err, tc.wantNot)
}
for _, fragment := range tc.contains {
if !strings.Contains(err.Error(), fragment) {
t.Fatalf("error = %v, want %q", err, fragment)
}
}
})
}
}
func TestResolvingRepositoryEnforcesChainLength(t *testing.T) {
for _, count := range []int{maximumProfileChainLength, maximumProfileChainLength + 1} {
t.Run(fmt.Sprintf("%d profiles", count), func(t *testing.T) {
profiles := make(map[string]*domain.ExecutionProfile, count)
for index := 1; index <= count; index++ {
id := fmt.Sprintf("profile-%d", index)
definition := &domain.ExecutionProfile{ID: id}
if index == count {
definition.Endpoint = "https://root.example/v1"
definition.Model = "model"
} else {
definition.BaseProfileID = fmt.Sprintf("profile-%d", index+1)
}
profiles[id] = definition
}
got, err := NewResolvingRepository(&resolvingTestRepository{profiles: profiles}).GetProfile(context.Background(), "profile-1")
if count == maximumProfileChainLength {
if err != nil || got == nil {
t.Fatalf("profile = %+v, error = %v, want accepted chain", got, err)
}
return
}
if !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("error = %v, want ErrInvalidProfile", err)
}
})
}
}
func TestResolvingRepositoryIsFreshAndCancellationAware(t *testing.T) {
repo := &resolvingTestRepository{profiles: map[string]*domain.ExecutionProfile{
"leaf": {ID: "leaf", BaseProfileID: "base"},
"base": {ID: "base", Endpoint: "https://base.example/v1", Model: "first", ExtraParams: map[string]any{"nested": map[string]any{"value": "first"}}},
}}
resolver := NewResolvingRepository(repo)
first, err := resolver.GetProfile(context.Background(), "leaf")
if err != nil || first.Model != "first" {
t.Fatalf("first result=(%+v, %v)", first, err)
}
repo.set("base", &domain.ExecutionProfile{ID: "base", Endpoint: "https://base.example/v1", Model: "second", ExtraParams: map[string]any{"nested": map[string]any{"value": "second"}}})
second, err := resolver.GetProfile(context.Background(), "leaf")
if err != nil || second.Model != "second" {
t.Fatalf("second result=(%+v, %v)", second, err)
}
canceled, cancel := context.WithCancel(context.Background())
cancel()
if _, err := resolver.GetProfile(canceled, "leaf"); !errors.Is(err, context.Canceled) {
t.Fatalf("canceled lookup error = %v", err)
}
if got := repo.callCount("leaf"); got != 2 {
t.Fatalf("calls after canceled lookup = %d, want 2", got)
}
duringTraversal, cancelDuringTraversal := context.WithCancel(context.Background())
repo.afterGet = func(id string) {
if id == "leaf" {
cancelDuringTraversal()
}
}
if _, err := resolver.GetProfile(duringTraversal, "leaf"); !errors.Is(err, context.Canceled) {
t.Fatalf("during traversal error = %v", err)
}
if got := repo.callCount("base"); got != 2 {
t.Fatalf("base calls after cancellation = %d, want 2", got)
}
terminalLookup, cancelTerminalLookup := context.WithCancel(context.Background())
repo.afterGet = func(id string) {
if id == "base" {
cancelTerminalLookup()
}
}
if _, err := resolver.GetProfile(terminalLookup, "leaf"); !errors.Is(err, context.Canceled) {
t.Fatalf("terminal lookup cancellation error = %v", err)
}
if got := repo.callCount("base"); got != 3 {
t.Fatalf("base calls after terminal cancellation = %d, want 3", got)
}
repo.afterGet = nil
var wg sync.WaitGroup
errors := make(chan error, 8)
for index := 0; index < cap(errors); index++ {
wg.Add(1)
go func() {
defer wg.Done()
resolved, err := resolver.GetProfile(context.Background(), "leaf")
if err != nil {
errors <- err
return
}
resolved.ExtraParams["nested"].(map[string]any)["value"] = "mutated"
}()
}
wg.Wait()
close(errors)
for err := range errors {
t.Errorf("concurrent resolution: %v", err)
}
latest, err := resolver.GetProfile(context.Background(), "leaf")
if err != nil || latest.ExtraParams["nested"].(map[string]any)["value"] != "second" {
t.Fatalf("latest result=(%+v, %v)", latest, err)
}
}
type resolvingTestRepository struct {
mu sync.Mutex
profiles map[string]*domain.ExecutionProfile
errors map[string]error
calls map[string]int
afterGet func(string)
}
func (r *resolvingTestRepository) GetProfile(ctx context.Context, id string) (*domain.ExecutionProfile, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
r.mu.Lock()
if r.calls == nil {
r.calls = make(map[string]int)
}
r.calls[id]++
err := r.errors[id]
profile := r.profiles[id]
afterGet := r.afterGet
r.mu.Unlock()
if afterGet != nil {
afterGet(id)
}
if err != nil {
return nil, err
}
if profile == nil {
if _, exists := r.profiles[id]; exists {
return nil, nil
}
return nil, ErrProfileNotFound
}
copy := *profile
return &copy, nil
}
func (r *resolvingTestRepository) set(id string, profile *domain.ExecutionProfile) {
r.mu.Lock()
defer r.mu.Unlock()
r.profiles[id] = profile
}
func (r *resolvingTestRepository) callCount(id string) int {
r.mu.Lock()
defer r.mu.Unlock()
return r.calls[id]
}
func profileTestFS(files map[string]string) fs.FS {
fsys := make(fstest.MapFS, len(files))
for name, content := range files {
fsys[name] = profileMapFile(content)
}
return fsys
}

View File

@@ -5,6 +5,7 @@ import (
"context" "context"
"errors" "errors"
"fmt" "fmt"
"strings"
"text/template" "text/template"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
@@ -18,6 +19,8 @@ var (
ErrInvalidMessageRole = errors.New("invalid or empty message role") ErrInvalidMessageRole = errors.New("invalid or empty message role")
) )
const artifactTextChunkSize = 64 * 1024
type goRenderer struct{} type goRenderer struct{}
func NewGoRenderer() Renderer { func NewGoRenderer() Renderer {
@@ -25,11 +28,13 @@ func NewGoRenderer() Renderer {
} }
func (r *goRenderer) Render(ctx context.Context, definition *domain.PromptDefinition, inputs map[string]*domain.Artifact, vars map[string]string) (*domain.RenderedPrompt, error) { func (r *goRenderer) Render(ctx context.Context, definition *domain.PromptDefinition, inputs map[string]*domain.Artifact, vars map[string]string) (*domain.RenderedPrompt, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
if definition == nil { if definition == nil {
return nil, fmt.Errorf("%w: nil prompt definition", ErrRenderFailure) return nil, fmt.Errorf("%w: nil prompt definition", ErrRenderFailure)
} }
// 1. Verify required inputs
for _, in := range definition.Inputs { for _, in := range definition.Inputs {
if !in.Required { if !in.Required {
continue continue
@@ -40,44 +45,54 @@ func (r *goRenderer) Render(ctx context.Context, definition *domain.PromptDefini
} }
} }
// 2. Setup template functions resolver := newArtifactTextResolver(ctx, inputs)
funcs := template.FuncMap{ funcs := template.FuncMap{
"input": func(name string) (string, error) { "input": resolver.resolve,
art, ok := inputs[name]
if !ok || art == nil {
return "", fmt.Errorf("%w: %s", ErrUnknownInput, name)
}
return string(art.Body), nil
},
} }
sessionID, err := renderSessionID(definition.SessionID, funcs, vars) if err := ctx.Err(); err != nil {
return nil, err
}
sessionID, err := renderSessionID(ctx, definition.SessionID, funcs, vars)
if err != nil { if err != nil {
return nil, err return nil, err
} }
if err := ctx.Err(); err != nil {
return nil, err
}
var renderedMessages []domain.RenderedMessage renderedMessages := make([]domain.RenderedMessage, 0, len(definition.Templates))
for i, tmplMsg := range definition.Templates { for i, tmplMsg := range definition.Templates {
select { if err := ctx.Err(); err != nil {
case <-ctx.Done(): return nil, err
return nil, ctx.Err()
default:
} }
if tmplMsg.Role == "" { if tmplMsg.Role == "" {
return nil, fmt.Errorf("%w: message %d", ErrInvalidMessageRole, i) return nil, fmt.Errorf("%w: message %d", ErrInvalidMessageRole, i)
} }
// Parse and execute template if err := ctx.Err(); err != nil {
tmpl, err := template.New(fmt.Sprintf("msg_%d", i)).Funcs(funcs).Option("missingkey=error").Parse(tmplMsg.Content) return nil, err
if err != nil { }
return nil, fmt.Errorf("%w: message %d: %v", ErrInvalidTemplate, i, err) tmpl, parseErr := template.New(fmt.Sprintf("msg_%d", i)).Funcs(funcs).Option("missingkey=error").Parse(tmplMsg.Content)
if err := ctx.Err(); err != nil {
return nil, err
}
if parseErr != nil {
return nil, fmt.Errorf("%w: message %d: %v", ErrInvalidTemplate, i, parseErr)
} }
if err := ctx.Err(); err != nil {
return nil, err
}
var buf bytes.Buffer var buf bytes.Buffer
if err := tmpl.Execute(&buf, vars); err != nil { executeErr := tmpl.Execute(&buf, vars)
return nil, fmt.Errorf("%w: message %d: %w", ErrRenderFailure, i, err) if err := ctx.Err(); err != nil {
return nil, err
}
if executeErr != nil {
return nil, fmt.Errorf("%w: message %d: %w", ErrRenderFailure, i, executeErr)
} }
renderedMessages = append(renderedMessages, domain.RenderedMessage{ renderedMessages = append(renderedMessages, domain.RenderedMessage{
@@ -85,6 +100,13 @@ func (r *goRenderer) Render(ctx context.Context, definition *domain.PromptDefini
Content: buf.String(), Content: buf.String(),
CacheControl: cloneCacheControl(tmplMsg.CacheControl), CacheControl: cloneCacheControl(tmplMsg.CacheControl),
}) })
if err := ctx.Err(); err != nil {
return nil, err
}
}
if err := ctx.Err(); err != nil {
return nil, err
} }
return &domain.RenderedPrompt{ return &domain.RenderedPrompt{
@@ -93,15 +115,75 @@ func (r *goRenderer) Render(ctx context.Context, definition *domain.PromptDefini
}, nil }, nil
} }
func renderSessionID(raw string, funcs template.FuncMap, vars map[string]string) (string, error) { type artifactTextResolver struct {
tmpl, err := template.New("session_id").Funcs(funcs).Option("missingkey=error").Parse(raw) ctx context.Context
if err != nil { inputs map[string]*domain.Artifact
return "", fmt.Errorf("%w: session_id: %v", ErrInvalidTemplate, err) textByName map[string]string
}
func newArtifactTextResolver(ctx context.Context, inputs map[string]*domain.Artifact) *artifactTextResolver {
return &artifactTextResolver{
ctx: ctx,
inputs: inputs,
textByName: make(map[string]string),
}
}
func (r *artifactTextResolver) resolve(name string) (string, error) {
if err := r.ctx.Err(); err != nil {
return "", err
}
artifact, ok := r.inputs[name]
if !ok || artifact == nil {
return "", fmt.Errorf("%w: %s", ErrUnknownInput, name)
}
if text, ok := r.textByName[name]; ok {
return text, nil
} }
var builder strings.Builder
builder.Grow(len(artifact.Body))
for start := 0; start < len(artifact.Body); start += artifactTextChunkSize {
if err := r.ctx.Err(); err != nil {
return "", err
}
end := min(start+artifactTextChunkSize, len(artifact.Body))
_, _ = builder.Write(artifact.Body[start:end])
if err := r.ctx.Err(); err != nil {
return "", err
}
}
if err := r.ctx.Err(); err != nil {
return "", err
}
text := builder.String()
r.textByName[name] = text
return text, nil
}
func renderSessionID(ctx context.Context, raw string, funcs template.FuncMap, vars map[string]string) (string, error) {
if err := ctx.Err(); err != nil {
return "", err
}
tmpl, parseErr := template.New("session_id").Funcs(funcs).Option("missingkey=error").Parse(raw)
if err := ctx.Err(); err != nil {
return "", err
}
if parseErr != nil {
return "", fmt.Errorf("%w: session_id: %v", ErrInvalidTemplate, parseErr)
}
if err := ctx.Err(); err != nil {
return "", err
}
var buf bytes.Buffer var buf bytes.Buffer
if err := tmpl.Execute(&buf, vars); err != nil { executeErr := tmpl.Execute(&buf, vars)
return "", fmt.Errorf("%w: session_id: %w", ErrRenderFailure, err) if err := ctx.Err(); err != nil {
return "", err
}
if executeErr != nil {
return "", fmt.Errorf("%w: session_id: %w", ErrRenderFailure, executeErr)
} }
sessionID, err := domain.NormalizeSessionID(buf.String()) sessionID, err := domain.NormalizeSessionID(buf.String())

View File

@@ -1,6 +1,7 @@
package prompt package prompt
import ( import (
"bytes"
"context" "context"
"errors" "errors"
"strings" "strings"
@@ -228,6 +229,24 @@ func TestGoRenderer_Render(t *testing.T) {
} }
}) })
t.Run("malformed rendered session id fails rendering", func(t *testing.T) {
def := &domain.PromptDefinition{
SessionID: "{{ .session_id }}",
Inputs: []domain.PromptInput{{Name: "transcript", Required: true}},
Templates: []domain.PromptMessageTemplate{
{Role: "system", Content: "Speak in a {{.tone}} tone."},
},
}
_, err := renderer.Render(ctx, def, inputs, map[string]string{
"tone": "concise",
"session_id": "session" + string([]byte{0xff}),
})
if !errors.Is(err, ErrRenderFailure) {
t.Fatalf("expected ErrRenderFailure, got %v", err)
}
})
t.Run("inserting required input artifact", func(t *testing.T) { t.Run("inserting required input artifact", func(t *testing.T) {
def := &domain.PromptDefinition{ def := &domain.PromptDefinition{
Inputs: []domain.PromptInput{{Name: "transcript", Required: true}}, Inputs: []domain.PromptInput{{Name: "transcript", Required: true}},
@@ -343,3 +362,186 @@ func TestGoRenderer_Render(t *testing.T) {
} }
}) })
} }
func TestGoRendererCancellation(t *testing.T) {
t.Run("before session parsing", func(t *testing.T) {
definition := &domain.PromptDefinition{
SessionID: "{{ malformed",
Templates: []domain.PromptMessageTemplate{
{Role: "user", Content: "not rendered"},
},
}
ctx, cancel := context.WithCancel(context.Background())
cancel()
result, err := NewGoRenderer().Render(ctx, definition, nil, nil)
if result != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("result=%#v err=%v, want nil/context.Canceled", result, err)
}
if errors.Is(err, ErrInvalidTemplate) {
t.Fatalf("pre-canceled render parsed the malformed session: %v", err)
}
result, err = NewGoRenderer().Render(context.Background(), definition, nil, nil)
if result != nil || !errors.Is(err, ErrInvalidTemplate) {
t.Fatalf("active render result=%#v err=%v, want nil/ErrInvalidTemplate", result, err)
}
})
t.Run("during artifact text conversion", func(t *testing.T) {
ctx := newCancelOnCheckContext(3)
body := bytes.Repeat([]byte("x"), artifactTextChunkSize*2)
original := append([]byte(nil), body...)
resolver := newArtifactTextResolver(ctx, map[string]*domain.Artifact{
"document": {Body: body},
})
text, err := resolver.resolve("document")
if text != "" || !errors.Is(err, context.Canceled) {
t.Fatalf("text length=%d err=%v, want empty/context.Canceled", len(text), err)
}
if _, published := resolver.textByName["document"]; published {
t.Fatal("canceled conversion published partial artifact text")
}
if !bytes.Equal(body, original) {
t.Fatal("resolver mutated the artifact body")
}
})
t.Run("after final message execution", func(t *testing.T) {
definition := &domain.PromptDefinition{
Templates: []domain.PromptMessageTemplate{
{Role: "user", Content: "fully rendered"},
},
}
counter := &checkCountingContext{Context: context.Background()}
if _, err := NewGoRenderer().Render(counter, definition, nil, nil); err != nil {
t.Fatalf("count render checkpoints: %v", err)
}
// The final three checks occur after template execution, after the
// message is assembled, and immediately before publication.
ctx := newCancelOnCheckContext(counter.checks - 2)
result, err := NewGoRenderer().Render(ctx, definition, nil, nil)
if result != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("result=%#v err=%v, want nil/context.Canceled", result, err)
}
})
}
func TestGoRendererArtifactTextLifecycle(t *testing.T) {
body := []byte{'a', 0xff, 'b', 0xfe}
original := append([]byte(nil), body...)
artifact := &domain.Artifact{Body: body}
inputs := map[string]*domain.Artifact{"document": artifact}
definition := &domain.PromptDefinition{
Templates: []domain.PromptMessageTemplate{
{Role: "user", Content: "{{input \"document\"}}|{{input \"document\"}}"},
},
}
first, err := NewGoRenderer().Render(context.Background(), definition, inputs, nil)
if err != nil {
t.Fatalf("first render: %v", err)
}
wantFirst := append(append(append([]byte(nil), body...), '|'), body...)
if !bytes.Equal([]byte(first.Messages[0].Content), wantFirst) {
t.Fatalf("rendered bytes=%v, want %v", []byte(first.Messages[0].Content), wantFirst)
}
if !bytes.Equal(body, original) {
t.Fatalf("renderer mutated artifact body: got %v want %v", body, original)
}
body[0] = 'z'
if bytes.Equal([]byte(first.Messages[0].Content), append(append(append([]byte(nil), body...), '|'), body...)) {
t.Fatal("completed render aliases the artifact body")
}
second, err := NewGoRenderer().Render(context.Background(), definition, inputs, nil)
if err != nil {
t.Fatalf("second render: %v", err)
}
wantSecond := append(append(append([]byte(nil), body...), '|'), body...)
if !bytes.Equal([]byte(second.Messages[0].Content), wantSecond) {
t.Fatalf("second render reused text from another call: got %v want %v", []byte(second.Messages[0].Content), wantSecond)
}
nilInputs := map[string]*domain.Artifact{"document": nil}
result, err := NewGoRenderer().Render(context.Background(), definition, nilInputs, nil)
if result != nil || !errors.Is(err, ErrUnknownInput) || !errors.Is(err, ErrRenderFailure) {
t.Fatalf("nil input result=%#v err=%v, want ErrUnknownInput and ErrRenderFailure", result, err)
}
}
func BenchmarkGoRendererArtifactReferences(b *testing.B) {
body := bytes.Repeat([]byte("document content "), (artifactTextChunkSize*4)/len("document content "))
inputs := map[string]*domain.Artifact{
"document": {Body: body},
}
tests := []struct {
name string
definition *domain.PromptDefinition
}{
{
name: "one reference",
definition: &domain.PromptDefinition{
Templates: []domain.PromptMessageTemplate{
{Role: "user", Content: "{{input \"document\"}}"},
},
},
},
{
name: "repeated across session and messages",
definition: &domain.PromptDefinition{
SessionID: "document-{{len (input \"document\")}}",
Templates: []domain.PromptMessageTemplate{
{Role: "system", Content: "{{input \"document\"}}"},
{Role: "user", Content: "{{input \"document\"}} {{input \"document\"}}"},
},
},
},
}
for _, tc := range tests {
b.Run(tc.name, func(b *testing.B) {
renderer := NewGoRenderer()
b.ReportAllocs()
b.SetBytes(int64(len(body)))
for range b.N {
if _, err := renderer.Render(context.Background(), tc.definition, inputs, nil); err != nil {
b.Fatal(err)
}
}
})
}
}
type checkCountingContext struct {
context.Context
checks int
}
func (c *checkCountingContext) Err() error {
c.checks++
return c.Context.Err()
}
type cancelOnCheckContext struct {
context.Context
cancel context.CancelFunc
remaining int
}
func newCancelOnCheckContext(checks int) *cancelOnCheckContext {
ctx, cancel := context.WithCancel(context.Background())
return &cancelOnCheckContext{Context: ctx, cancel: cancel, remaining: checks}
}
func (c *cancelOnCheckContext) Err() error {
if c.Context.Err() == nil {
c.remaining--
if c.remaining == 0 {
c.cancel()
}
}
return c.Context.Err()
}

View File

@@ -0,0 +1,98 @@
package promptdef
import (
"fmt"
"io/fs"
"os"
"path"
"path/filepath"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/filecatalog"
)
type contentSourceRoot interface {
readContentFile(sourcePath string, contentFile string) (string, string, error)
}
type osContentSourceRoot struct {
root string
sourcePathsRelative bool
}
func (r osContentSourceRoot) readContentFile(sourcePath string, contentFile string) (string, string, error) {
if strings.TrimSpace(contentFile) == "" {
return "", "", fmt.Errorf("path is required")
}
if filepath.IsAbs(contentFile) {
return "", "", fmt.Errorf("path %q must be relative", contentFile)
}
root, err := filepath.Abs(r.root)
if err != nil {
return "", "", fmt.Errorf("resolve source root %q: %w", r.root, err)
}
canonicalRoot, err := filepath.EvalSymlinks(root)
if err != nil {
return "", "", fmt.Errorf("resolve source root %q: %w", r.root, err)
}
promptPath := sourcePath
if r.sourcePathsRelative && !filepath.IsAbs(promptPath) {
promptPath = filepath.Join(root, filepath.FromSlash(promptPath))
} else {
promptPath, err = filepath.Abs(promptPath)
if err != nil {
return "", "", fmt.Errorf("resolve prompt source %q: %w", sourcePath, err)
}
}
resolvedPath := filepath.Clean(filepath.Join(filepath.Dir(promptPath), contentFile))
if !containsOSPath(root, resolvedPath) {
return "", "", fmt.Errorf("path %q escapes source root %q", contentFile, r.root)
}
canonicalPath, err := filepath.EvalSymlinks(resolvedPath)
if err != nil {
return "", "", err
}
if !containsOSPath(canonicalRoot, canonicalPath) {
return "", "", fmt.Errorf("path %q escapes source root %q", contentFile, r.root)
}
body, err := os.ReadFile(canonicalPath)
if err != nil {
return "", "", err
}
return string(body), resolvedPath, nil
}
type fsContentSourceRoot struct {
fsys fs.FS
root string
}
func (r fsContentSourceRoot) readContentFile(sourcePath string, contentFile string) (string, string, error) {
root := filecatalog.CleanFSRoot(r.root)
cleanSourcePath := path.Clean(sourcePath)
if cleanSourcePath == root {
root = path.Dir(root)
}
resolvedPath, _, err := filecatalog.ResolveFSPath(root, path.Dir(cleanSourcePath), contentFile)
if err != nil {
return "", "", err
}
body, err := fs.ReadFile(r.fsys, resolvedPath)
if err != nil {
return "", "", err
}
return string(body), resolvedPath, nil
}
func containsOSPath(root string, name string) bool {
relative, err := filepath.Rel(root, name)
if err != nil || filepath.IsAbs(relative) {
return false
}
return relative != ".." && !strings.HasPrefix(relative, ".."+string(filepath.Separator))
}

View File

@@ -5,14 +5,11 @@ import (
"context" "context"
"errors" "errors"
"fmt" "fmt"
"io"
"io/fs" "io/fs"
"os"
"path"
"path/filepath"
"strings" "strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/filecatalog"
"gopkg.in/yaml.v3" "gopkg.in/yaml.v3"
) )
@@ -22,13 +19,8 @@ var (
ErrInvalidPromptDefinition = errors.New("invalid prompt definition configuration") ErrInvalidPromptDefinition = errors.New("invalid prompt definition configuration")
) )
type filesystemRepository struct { type sourceRepository struct {
dir string source promptDefinitionSource
}
type fsRepository struct {
fsys fs.FS
root string
} }
type promptDefinitionFile struct { type promptDefinitionFile struct {
@@ -69,19 +61,47 @@ type promptOutputContractFile struct {
} }
func NewFilesystemRepository(dir string) Repository { func NewFilesystemRepository(dir string) Repository {
return &filesystemRepository{dir: dir} return &sourceRepository{
source: osPromptSource{
root: dir,
contentRoot: osContentSourceRoot{root: dir},
},
}
} }
func NewFSRepository(fsys fs.FS, root string) Repository { func NewFSRepository(fsys fs.FS, root string) Repository {
return &fsRepository{fsys: fsys, root: root} return &sourceRepository{
source: fsPromptSource{
fsys: fsys,
root: root,
contentRoot: fsContentSourceRoot{fsys: fsys, root: root},
},
}
} }
func (r *filesystemRepository) GetPromptDefinition(ctx context.Context, id string, version string) (*domain.PromptDefinition, error) { // NewFileRepository constructs a repository for one operating-system prompt file.
func NewFileRepository(fsys fs.FS, file string, sourceDir string) Repository {
return &sourceRepository{
source: fsPromptSource{
fsys: fsys,
root: file,
contentRoot: osContentSourceRoot{
root: sourceDir,
sourcePathsRelative: true,
},
},
}
}
func (r *sourceRepository) GetPromptDefinition(ctx context.Context, id string, version string) (*domain.PromptDefinition, error) {
if strings.TrimSpace(id) == "" { if strings.TrimSpace(id) == "" {
return nil, fmt.Errorf("%w: prompt id is required", ErrInvalidPromptDefinition) return nil, fmt.Errorf("%w: prompt id is required", ErrInvalidPromptDefinition)
} }
if r == nil || r.source == nil {
return nil, errors.New("failed to read prompt definition directory: source is nil")
}
files, err := filecatalog.FindYAMLFiles(ctx, r.dir) files, err := r.source.findYAMLFiles(ctx)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to read prompt definition directory: %w", err) return nil, fmt.Errorf("failed to read prompt definition directory: %w", err)
} }
@@ -94,33 +114,24 @@ func (r *filesystemRepository) GetPromptDefinition(ctx context.Context, id strin
default: default:
} }
relPath := filecatalog.RelativePath(r.dir, fullPath) relPath := r.source.displayPath(fullPath)
fileMatch := filecatalog.Stem(filepath.Base(fullPath)) == id data, err := r.source.readDefinition(fullPath)
raw, err := loadPromptDefinitionFile(fullPath)
if err != nil { if err != nil {
if fileMatch || promptDefinitionFileHasID(fullPath, id) { return nil, fmt.Errorf("failed to read prompt definition file %s: %w", relPath, err)
}
raw, err := decodePromptDefinition(data)
if err != nil {
if promptDefinitionDataMatches(data, id, version) {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidYAML, relPath, err) return nil, fmt.Errorf("%w: %s: %v", ErrInvalidYAML, relPath, err)
} }
continue continue
} }
if !promptDefinitionMatches(raw, id, version) {
def, err := normalizePromptDefinition(raw, fullPath)
if err != nil {
if fileMatch || strings.TrimSpace(raw.ID) == id {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidPromptDefinition, relPath, err)
}
continue
}
if def.ID != id {
continue
}
if version != "" && def.Version != version {
continue continue
} }
matches = append(matches, promptDefinitionMatch{ matches = append(matches, promptDefinitionMatch{
def: def, raw: raw,
sourcePath: fullPath,
path: relPath, path: relPath,
}) })
} }
@@ -137,132 +148,23 @@ func (r *filesystemRepository) GetPromptDefinition(ctx context.Context, id strin
} }
if len(matches) == 1 { if len(matches) == 1 {
return matches[0].def, nil match := matches[0]
def, err := normalizePromptDefinition(match.raw, r.source, match.sourcePath)
if err != nil {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidPromptDefinition, match.path, err)
}
return def, nil
} }
return nil, ErrPromptDefinitionNotFound return nil, ErrPromptDefinitionNotFound
} }
func (r *fsRepository) GetPromptDefinition(ctx context.Context, id string, version string) (*domain.PromptDefinition, error) {
return loadPromptDefinition(ctx, r.fsys, r.root, id, version)
}
type promptDefinitionMatch struct { type promptDefinitionMatch struct {
def *domain.PromptDefinition raw *promptDefinitionFile
sourcePath string
path string path string
} }
func loadPromptDefinitionFile(path string) (*promptDefinitionFile, error) {
data, err := os.ReadFile(path)
if err != nil {
return nil, fmt.Errorf("failed to read prompt definition file: %w", err)
}
var raw promptDefinitionFile
decoder := yaml.NewDecoder(bytes.NewReader(data))
decoder.KnownFields(true)
if err := decoder.Decode(&raw); err != nil {
return nil, err
}
return &raw, nil
}
func promptDefinitionFileHasID(path string, id string) bool {
data, err := os.ReadFile(path)
if err != nil {
return false
}
var raw struct {
ID string `yaml:"id"`
}
if err := yaml.NewDecoder(bytes.NewReader(data)).Decode(&raw); err != nil {
return false
}
return strings.TrimSpace(raw.ID) == id
}
func loadPromptDefinition(ctx context.Context, fsys fs.FS, root string, id string, version string) (*domain.PromptDefinition, error) {
if strings.TrimSpace(id) == "" {
return nil, fmt.Errorf("%w: prompt id is required", ErrInvalidPromptDefinition)
}
if fsys == nil {
return nil, fmt.Errorf("failed to read prompt definition directory: filesystem is nil")
}
files, err := filecatalog.FindFSYAMLFiles(ctx, fsys, root)
if err != nil {
return nil, fmt.Errorf("failed to read prompt definition directory: %w", err)
}
cleanRoot := filecatalog.CleanFSRoot(root)
rootInfo, err := fs.Stat(fsys, cleanRoot)
if err != nil {
return nil, fmt.Errorf("failed to read prompt definition directory: %w", err)
}
var matches []promptDefinitionMatch
for _, fullPath := range files {
select {
case <-ctx.Done():
return nil, ctx.Err()
default:
}
relPath := filecatalog.DisplayPath(root, fullPath)
fileMatch := filecatalog.Stem(path.Base(fullPath)) == id
data, err := fs.ReadFile(fsys, fullPath)
if err != nil {
if fileMatch {
return nil, fmt.Errorf("%w: %s: failed to read prompt definition file: %v", ErrInvalidYAML, relPath, err)
}
continue
}
raw, err := decodePromptDefinition(data)
if err != nil {
if fileMatch || promptDefinitionDataHasID(data, id) {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidYAML, relPath, err)
}
continue
}
def, err := normalizePromptDefinitionFromFS(raw, fsys, root, fullPath, rootInfo.IsDir())
if err != nil {
if fileMatch || strings.TrimSpace(raw.ID) == id {
return nil, fmt.Errorf("%w: %s: %v", ErrInvalidPromptDefinition, relPath, err)
}
continue
}
if def.ID != id {
continue
}
if version != "" && def.Version != version {
continue
}
matches = append(matches, promptDefinitionMatch{
def: def,
path: relPath,
})
}
if len(matches) > 1 {
paths := make([]string, 0, len(matches))
for _, match := range matches {
paths = append(paths, match.path)
}
if version != "" {
return nil, fmt.Errorf("%w: duplicate prompt definition id %q version %q found in: %s", ErrInvalidPromptDefinition, id, version, strings.Join(paths, ", "))
}
return nil, fmt.Errorf("%w: duplicate prompt definition id %q found in: %s", ErrInvalidPromptDefinition, id, strings.Join(paths, ", "))
}
if len(matches) == 1 {
return matches[0].def, nil
}
return nil, ErrPromptDefinitionNotFound
}
func decodePromptDefinition(data []byte) (*promptDefinitionFile, error) { func decodePromptDefinition(data []byte) (*promptDefinitionFile, error) {
var raw promptDefinitionFile var raw promptDefinitionFile
decoder := yaml.NewDecoder(bytes.NewReader(data)) decoder := yaml.NewDecoder(bytes.NewReader(data))
@@ -270,59 +172,44 @@ func decodePromptDefinition(data []byte) (*promptDefinitionFile, error) {
if err := decoder.Decode(&raw); err != nil { if err := decoder.Decode(&raw); err != nil {
return nil, err return nil, err
} }
var additional yaml.Node
if err := decoder.Decode(&additional); err != io.EOF {
if err != nil {
return nil, err
}
return nil, errors.New("prompt definition file must contain exactly one YAML document")
}
return &raw, nil return &raw, nil
} }
func promptDefinitionDataHasID(data []byte, id string) bool { func promptDefinitionDataMatches(data []byte, id string, version string) bool {
var raw struct { var raw struct {
ID string `yaml:"id"` ID string `yaml:"id"`
Version string `yaml:"version"`
} }
if err := yaml.NewDecoder(bytes.NewReader(data)).Decode(&raw); err != nil { if err := yaml.NewDecoder(bytes.NewReader(data)).Decode(&raw); err != nil {
return false return false
} }
return strings.TrimSpace(raw.ID) == id return promptSelectorMatches(raw.ID, raw.Version, id, version)
} }
func normalizePromptDefinition(raw *promptDefinitionFile, sourcePath string) (*domain.PromptDefinition, error) { func promptDefinitionMatches(raw *promptDefinitionFile, id string, version string) bool {
promptDir := filepath.Dir(sourcePath) if raw == nil {
return normalizePromptDefinitionWithContent(raw, func(contentFile string) (string, string, error) { return false
resolvedPath := strings.TrimSpace(contentFile)
if !filepath.IsAbs(resolvedPath) {
resolvedPath = filepath.Join(promptDir, resolvedPath)
} }
resolvedPath = filepath.Clean(resolvedPath) return promptSelectorMatches(raw.ID, raw.Version, id, version)
body, err := os.ReadFile(resolvedPath)
if err != nil {
return "", "", err
}
return string(body), resolvedPath, nil
})
} }
func normalizePromptDefinitionFromFS(raw *promptDefinitionFile, fsys fs.FS, root string, sourcePath string, rootIsDir bool) (*domain.PromptDefinition, error) { func promptSelectorMatches(rawID string, rawVersion string, id string, version string) bool {
promptDir := path.Dir(sourcePath) if strings.TrimSpace(rawID) != id {
return normalizePromptDefinitionWithContent(raw, func(contentFile string) (string, string, error) { return false
var resolvedPath string
if rootIsDir {
var err error
resolvedPath, _, err = filecatalog.ResolveFSPath(root, promptDir, contentFile)
if err != nil {
return "", "", err
}
} else {
resolvedPath = strings.TrimSpace(contentFile)
if !path.IsAbs(resolvedPath) {
resolvedPath = path.Join(promptDir, resolvedPath)
}
resolvedPath = strings.TrimPrefix(path.Clean(resolvedPath), "/")
} }
return version == "" || strings.TrimSpace(rawVersion) == version
}
body, err := fs.ReadFile(fsys, resolvedPath) func normalizePromptDefinition(raw *promptDefinitionFile, sourceRoot contentSourceRoot, sourcePath string) (*domain.PromptDefinition, error) {
if err != nil { return normalizePromptDefinitionWithContent(raw, func(contentFile string) (string, string, error) {
return "", "", err return sourceRoot.readContentFile(sourcePath, contentFile)
}
return string(body), resolvedPath, nil
}) })
} }
@@ -402,17 +289,14 @@ func normalizePromptDefinitionWithContent(raw *promptDefinitionFile, readContent
}) })
} }
if !isValidOutputFormat(raw.Output.Format) { outputContract := domain.OutputContract{
return nil, fmt.Errorf("invalid output format: %q", raw.Output.Format) Format: raw.Output.Format,
ValidationMode: raw.Output.ValidationMode,
SchemaPath: strings.TrimSpace(raw.Output.SchemaPath),
RepairAttempts: raw.Output.RepairAttempts,
} }
if !isValidValidationMode(raw.Output.ValidationMode) { if err := domain.ValidateOutputContract(outputContract); err != nil {
return nil, fmt.Errorf("invalid validation mode: %q", raw.Output.ValidationMode) return nil, fmt.Errorf("output: %w", err)
}
if raw.Output.ValidationMode == domain.ValidationJSONSchema && strings.TrimSpace(raw.Output.SchemaPath) == "" {
return nil, errors.New("output.schema_path is required when output.validation_mode is json_schema")
}
if raw.Output.RepairAttempts < 0 {
return nil, errors.New("output.repair_attempts must be greater than or equal to 0")
} }
defaultProfile := "" defaultProfile := ""
@@ -432,12 +316,7 @@ func normalizePromptDefinitionWithContent(raw *promptDefinitionFile, readContent
Inputs: inputs, Inputs: inputs,
Templates: templates, Templates: templates,
OutputFormat: raw.Output.Format, OutputFormat: raw.Output.Format,
Validation: domain.OutputContract{ Validation: outputContract,
Format: raw.Output.Format,
ValidationMode: raw.Output.ValidationMode,
SchemaPath: strings.TrimSpace(raw.Output.SchemaPath),
RepairAttempts: raw.Output.RepairAttempts,
},
}, nil }, nil
} }
@@ -464,21 +343,3 @@ func normalizeCacheControl(raw *cacheControlFile) (*domain.CacheControl, error)
TTL: ttl, TTL: ttl,
}, nil }, nil
} }
func isValidOutputFormat(f domain.OutputFormat) bool {
switch f {
case domain.FormatText, domain.FormatMarkdown, domain.FormatJSON:
return true
default:
return false
}
}
func isValidValidationMode(m domain.ValidationMode) bool {
switch m {
case domain.ValidationNone, domain.ValidationBasic, domain.ValidationJSON, domain.ValidationJSONSchema:
return true
default:
return false
}
}

View File

@@ -3,23 +3,51 @@ package promptdef
import ( import (
"context" "context"
"errors" "errors"
"fmt"
"io/fs" "io/fs"
"os" "os"
"path/filepath" "path/filepath"
"strings" "strings"
"sync"
"testing" "testing"
"testing/fstest" "testing/fstest"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
) )
func TestFilesystemRepository_GetPromptDefinition(t *testing.T) { func TestPromptRepositoryDefinitionFixtures(t *testing.T) {
sources := []struct {
name string
newRepository func(string) Repository
contentPathsAreFull bool
}{
{
name: "operating system",
newRepository: NewFilesystemRepository,
contentPathsAreFull: true,
},
{
name: "filesystem",
newRepository: func(root string) Repository {
return NewFSRepository(os.DirFS(root), ".")
},
},
}
for _, source := range sources {
t.Run(source.name, func(t *testing.T) {
testPromptRepositoryDefinitionFixtures(t, source.newRepository, source.contentPathsAreFull)
})
}
}
func testPromptRepositoryDefinitionFixtures(t *testing.T, newRepository func(string) Repository, contentPathsAreFull bool) {
tmpDir := t.TempDir() tmpDir := t.TempDir()
if err := copyTree("testdata", tmpDir); err != nil { if err := copyTree("testdata", tmpDir); err != nil {
t.Fatalf("failed to copy testdata: %v", err) t.Fatalf("failed to copy testdata: %v", err)
} }
repo := NewFilesystemRepository(tmpDir) repo := newRepository(tmpDir)
ctx := context.Background() ctx := context.Background()
t.Run("valid inline prompt", func(t *testing.T) { t.Run("valid inline prompt", func(t *testing.T) {
@@ -64,8 +92,8 @@ func TestFilesystemRepository_GetPromptDefinition(t *testing.T) {
if p.Templates[1].ContentFile == "" { if p.Templates[1].ContentFile == "" {
t.Fatal("expected ContentFile source metadata to be preserved") t.Fatal("expected ContentFile source metadata to be preserved")
} }
if !filepath.IsAbs(p.Templates[1].ContentFile) { if filepath.IsAbs(p.Templates[1].ContentFile) != contentPathsAreFull {
t.Fatalf("expected resolved content_file path to be absolute, got %q", p.Templates[1].ContentFile) t.Fatalf("unexpected content_file path representation: %q", p.Templates[1].ContentFile)
} }
}) })
@@ -287,20 +315,20 @@ output:
targetErr error targetErr error
errSubstrs []string errSubstrs []string
}{ }{
{name: "invalid YAML", id: "invalid_yaml", targetErr: ErrInvalidYAML}, {name: "unidentifiable invalid YAML is unrelated", id: "invalid-yaml", targetErr: ErrPromptDefinitionNotFound},
{name: "missing id", id: "missing_id", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"id is required"}}, {name: "missing id is not selected by filename", id: "missing_id", targetErr: ErrPromptDefinitionNotFound},
{name: "no messages", id: "no_messages", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"at least one message is required"}}, {name: "no messages", id: "no-messages", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"at least one message is required"}},
{name: "both content and content_file", id: "both_content_and_content_file", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"exactly one"}}, {name: "both content and content_file", id: "both-content-and-content-file", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"exactly one"}},
{name: "neither content nor content_file", id: "neither_content_nor_content_file", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"exactly one"}}, {name: "neither content nor content_file", id: "neither-content-nor-content-file", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"exactly one"}},
{name: "missing content_file", id: "missing_content_file", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"failed to read content_file"}}, {name: "missing content_file", id: "missing-content-file", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"failed to read content_file"}},
{name: "duplicate input names", id: "duplicate_input_names", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"duplicate input name"}}, {name: "duplicate input names", id: "duplicate-input-names", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"duplicate input name"}},
{name: "invalid validation mode", id: "invalid_validation_mode", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"invalid validation mode"}}, {name: "invalid validation mode", id: "invalid-validation-mode", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"invalid validation mode"}},
{name: "json_schema without schema_path", id: "json_schema_without_schema_path", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"schema_path"}}, {name: "json_schema without schema_path", id: "json-schema-without-schema-path", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"schema_path"}},
{name: "unknown input field", id: "unknown_input_field", targetErr: ErrInvalidYAML, errSubstrs: []string{"field unknown_input_setting not found"}}, {name: "unknown input field", id: "unknown-input-field", targetErr: ErrInvalidYAML, errSubstrs: []string{"field unknown_input_setting not found"}},
{name: "empty cache control type", id: "empty_cache_control_type", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"cache_control", "type is required"}}, {name: "empty cache control type", id: "empty-cache-control-type", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"cache_control", "type is required"}},
{name: "unsupported cache control type", id: "unsupported_cache_control_type", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"cache_control", "unsupported type"}}, {name: "unsupported cache control type", id: "unsupported-cache-control-type", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"cache_control", "unsupported type"}},
{name: "unsupported cache control ttl", id: "unsupported_cache_control_ttl", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"cache_control", "unsupported ttl"}}, {name: "unsupported cache control ttl", id: "unsupported-cache-control-ttl", targetErr: ErrInvalidPromptDefinition, errSubstrs: []string{"cache_control", "unsupported ttl"}},
{name: "unknown cache control field", id: "unknown_cache_control_field", targetErr: ErrInvalidYAML, errSubstrs: []string{"field unexpected not found"}}, {name: "unknown cache control field", id: "unknown-cache-control-field", targetErr: ErrInvalidYAML, errSubstrs: []string{"field unexpected not found"}},
} }
for _, tc := range cases { for _, tc := range cases {
@@ -396,7 +424,7 @@ output:
for _, tc := range tests { for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) { t.Run(tc.name, func(t *testing.T) {
repo := NewFSRepository(fstest.MapFS{ fsys := &recordingFS{FS: fstest.MapFS{
"prompts/prompt.yaml": &fstest.MapFile{Data: []byte(` "prompts/prompt.yaml": &fstest.MapFile{Data: []byte(`
id: fs-escaped-prompt id: fs-escaped-prompt
version: "1.0.0" version: "1.0.0"
@@ -409,7 +437,8 @@ output:
repair_attempts: 0 repair_attempts: 0
`)}, `)},
"outside.tmpl": &fstest.MapFile{Data: []byte(`Outside root.`)}, "outside.tmpl": &fstest.MapFile{Data: []byte(`Outside root.`)},
}, "prompts") }}
repo := NewFSRepository(fsys, "prompts")
_, err := repo.GetPromptDefinition(context.Background(), "fs-escaped-prompt", "") _, err := repo.GetPromptDefinition(context.Background(), "fs-escaped-prompt", "")
if !errors.Is(err, ErrInvalidPromptDefinition) { if !errors.Is(err, ErrInvalidPromptDefinition) {
@@ -418,65 +447,644 @@ output:
if !strings.Contains(err.Error(), tc.wantErr) { if !strings.Contains(err.Error(), tc.wantErr) {
t.Fatalf("expected error to contain %q, got %v", tc.wantErr, err) t.Fatalf("expected error to contain %q, got %v", tc.wantErr, err)
} }
if fsys.wasOpened("outside.tmpl") {
t.Fatal("rejected content path opened the outside file")
}
}) })
} }
} }
func TestFSRepositoryRejectsDuplicatePromptIDs(t *testing.T) { type recordingFS struct {
repo := NewFSRepository(fstest.MapFS{ fs.FS
"one.yaml": &fstest.MapFile{Data: []byte(` mu sync.Mutex
id: duplicate-fs-prompt opened []string
version: "1.0.0" }
messages:
- role: user
content: First.
output:
format: text
validation_mode: none
repair_attempts: 0
`)},
"nested/two.yaml": &fstest.MapFile{Data: []byte(`
id: duplicate-fs-prompt
version: "1.0.0"
messages:
- role: user
content: Second.
output:
format: text
validation_mode: none
repair_attempts: 0
`)},
}, ".")
_, err := repo.GetPromptDefinition(context.Background(), "duplicate-fs-prompt", "") func TestPromptRepositoryReturnsDefinitionReadFailures(t *testing.T) {
if !errors.Is(err, ErrInvalidPromptDefinition) { readErr := errors.New("definition read failed")
t.Fatalf("expected ErrInvalidPromptDefinition, got %v", err) fsys := &definitionReadFailureFS{
FS: fstest.MapFS{
"prompts/target.yaml": &fstest.MapFile{Data: []byte("unread")},
},
target: "prompts/target.yaml",
err: readErr,
} }
if !strings.Contains(err.Error(), "one.yaml") || !strings.Contains(err.Error(), "nested/two.yaml") { repo := NewFSRepository(fsys, "prompts")
t.Fatalf("expected duplicate paths in error, got %v", err)
definition, err := repo.GetPromptDefinition(context.Background(), "target", "1")
if definition != nil || !errors.Is(err, readErr) {
t.Fatalf("GetPromptDefinition() = (%#v, %v), want nil and definition read error", definition, err)
}
if errors.Is(err, ErrPromptDefinitionNotFound) {
t.Fatalf("definition read error was classified as absence: %v", err)
}
if !strings.Contains(err.Error(), "target.yaml") {
t.Fatalf("definition read error lacks source context: %v", err)
} }
} }
func TestFSRepositoryRejectsUnknownYAMLFields(t *testing.T) { type definitionReadFailureFS struct {
repo := NewFSRepository(fstest.MapFS{ fs.FS
"not_named_like_id.yaml": &fstest.MapFile{Data: []byte(` target string
id: strict-fs-prompt err error
version: "1.0.0" }
unknown: true
func (f *definitionReadFailureFS) Open(name string) (fs.File, error) {
if name == f.target {
return nil, f.err
}
return f.FS.Open(name)
}
func (f *recordingFS) Open(name string) (fs.File, error) {
f.mu.Lock()
f.opened = append(f.opened, name)
f.mu.Unlock()
return f.FS.Open(name)
}
func (f *recordingFS) wasOpened(name string) bool {
return f.openCount(name) > 0
}
func (f *recordingFS) openCount(name string) int {
f.mu.Lock()
defer f.mu.Unlock()
count := 0
for _, opened := range f.opened {
if opened == name {
count++
}
}
return count
}
func TestPromptRepositorySelectionUsesYAMLMetadata(t *testing.T) {
const validDefinition = `
id: selected-prompt
version: "1"
messages: messages:
- role: user - role: user
content: Invalid. content: selected
output: output:
format: text format: text
validation_mode: none validation_mode: none
repair_attempts: 0 `
`)}, tests := []struct {
}, ".") name string
files map[string]string
wantErr error
diagnostics []string
wantContent string
}{
{
name: "same-stem strict error with different YAML ID is unrelated",
files: map[string]string{
"selected-prompt.yaml": `
id: another-prompt
version: "1"
unknown: true
`,
"valid.yaml": validDefinition,
},
},
{
name: "unidentifiable same-stem YAML is unrelated",
files: map[string]string{
"selected-prompt.yaml": "id: [",
"valid.yaml": validDefinition,
},
},
{
name: "same ID invalid different version is unrelated",
files: map[string]string{
"invalid-version.yaml": `
id: selected-prompt
version: "2"
output:
format: text
validation_mode: none
`,
"valid.yaml": validDefinition,
},
},
{
name: "selected content is resolved relative to its definition",
files: map[string]string{
"nested/selected.yaml": `
id: selected-prompt
version: "1"
messages:
- role: user
content_file: ./content/selected.tmpl
output:
format: text
validation_mode: none
`,
"nested/content/selected.tmpl": "selected from file",
},
wantContent: "selected from file",
},
{
name: "duplicate selected definitions are ambiguous",
files: map[string]string{
"selected-a.yaml": validDefinition,
"nested/selected-b.yaml": `
id: selected-prompt
version: "1"
messages:
- role: user
content: duplicate
output:
format: text
validation_mode: none
`,
},
wantErr: ErrInvalidPromptDefinition,
diagnostics: []string{"duplicate prompt definition id", "selected-a.yaml", "nested/selected-b.yaml"},
},
{
name: "selected strict error is authoritative",
files: map[string]string{
"selected-strict.yaml": `
id: selected-prompt
version: "1"
unknown: true
`,
},
wantErr: ErrInvalidYAML,
diagnostics: []string{"selected-strict.yaml"},
},
{
name: "selected semantic error is authoritative",
files: map[string]string{
"selected-invalid.yaml": `
id: selected-prompt
version: "1"
output:
format: text
validation_mode: none
`,
},
wantErr: ErrInvalidPromptDefinition,
diagnostics: []string{"selected-invalid.yaml", "at least one message"},
},
{
name: "selected content error includes definition context",
files: map[string]string{
"selected-missing-content.yaml": `
id: selected-prompt
version: "1"
messages:
- role: user
content_file: missing.tmpl
output:
format: text
validation_mode: none
`,
},
wantErr: ErrInvalidPromptDefinition,
diagnostics: []string{"selected-missing-content.yaml", "failed to read content_file", "missing.tmpl"},
},
}
_, err := repo.GetPromptDefinition(context.Background(), "strict-fs-prompt", "") for _, source := range promptRepositorySources() {
for _, tc := range tests {
t.Run(source.name+"/"+tc.name, func(t *testing.T) {
repo := source.newRepository(t, tc.files)
got, err := repo.GetPromptDefinition(context.Background(), "selected-prompt", "1")
if tc.wantErr != nil {
if !errors.Is(err, tc.wantErr) {
t.Fatalf("expected %v, got %v", tc.wantErr, err)
}
for _, diagnostic := range tc.diagnostics {
if !strings.Contains(err.Error(), diagnostic) {
t.Fatalf("expected error to contain %q, got %v", diagnostic, err)
}
}
return
}
if err != nil {
t.Fatalf("load selected prompt: %v", err)
}
wantContent := tc.wantContent
if wantContent == "" {
wantContent = "selected"
}
if got.ID != "selected-prompt" || got.Version != "1" || got.Templates[0].Content != wantContent {
t.Fatalf("unexpected selected definition: %+v", got)
}
})
}
}
}
func TestPromptRepositoryHonorsCancellation(t *testing.T) {
for _, source := range promptRepositorySources() {
t.Run(source.name, func(t *testing.T) {
repo := source.newRepository(t, map[string]string{
"definition.yaml": `
id: cancelled-prompt
version: "1"
messages:
- role: user
content: selected
output:
format: text
validation_mode: none
`,
})
ctx, cancel := context.WithCancel(context.Background())
cancel()
_, err := repo.GetPromptDefinition(ctx, "cancelled-prompt", "1")
if !errors.Is(err, context.Canceled) {
t.Fatalf("expected context cancellation, got %v", err)
}
})
}
}
func TestPromptRepositoryRequiresOneYAMLDocument(t *testing.T) {
const definition = `
id: one-document
version: "1"
messages:
- role: user
content: selected
output:
format: text
validation_mode: none
`
tests := []struct {
name string
suffix string
wantErr bool
}{
{name: "comments and trailing whitespace", suffix: "\n# trailing comment\n\n"},
{name: "second populated document", suffix: "\n---\nid: another\n", wantErr: true},
{name: "second empty document", suffix: "\n---\n", wantErr: true},
{name: "malformed trailing YAML", suffix: "\n---\n[", wantErr: true},
}
for _, source := range promptRepositorySources() {
for _, tc := range tests {
t.Run(source.name+"/"+tc.name, func(t *testing.T) {
repo := source.newRepository(t, map[string]string{"definition.yaml": definition + tc.suffix})
_, err := repo.GetPromptDefinition(context.Background(), "one-document", "1")
if tc.wantErr {
if !errors.Is(err, ErrInvalidYAML) { if !errors.Is(err, ErrInvalidYAML) {
t.Fatalf("expected ErrInvalidYAML, got %v", err) t.Fatalf("expected ErrInvalidYAML, got %v", err)
} }
if !strings.Contains(err.Error(), "definition.yaml") {
t.Fatalf("expected source path in error, got %v", err)
}
return
}
if err != nil {
t.Fatalf("load one document: %v", err)
}
})
}
}
}
func TestPromptRepositoryReadsOnlySelectedContent(t *testing.T) {
fsys := &recordingFS{FS: fstest.MapFS{
"target.yaml": &fstest.MapFile{Data: []byte(`
id: target
version: "1"
messages:
- role: user
content_file: target.tmpl
output:
format: text
validation_mode: none
`)},
"target.tmpl": &fstest.MapFile{Data: []byte("selected")},
"unrelated.yaml": &fstest.MapFile{Data: []byte(`
id: unrelated
version: "1"
messages:
- role: user
content_file: unrelated.tmpl
output:
format: text
validation_mode: none
`)},
"unrelated.tmpl": &fstest.MapFile{Data: []byte("unrelated")},
"other-version.yaml": &fstest.MapFile{Data: []byte(`
id: target
version: "2"
messages:
- role: user
content_file: other-version.tmpl
output:
format: text
validation_mode: none
`)},
"other-version.tmpl": &fstest.MapFile{Data: []byte("other version")},
}}
repo := NewFSRepository(fsys, ".")
for lookup := 1; lookup <= 2; lookup++ {
got, err := repo.GetPromptDefinition(context.Background(), "target", "1")
if err != nil {
t.Fatalf("lookup %d: %v", lookup, err)
}
if got.Templates[0].Content != "selected" {
t.Fatalf("lookup %d content = %q", lookup, got.Templates[0].Content)
}
if count := fsys.openCount("target.tmpl"); count != lookup {
t.Fatalf("selected content opens after lookup %d = %d, want %d", lookup, count, lookup)
}
for _, name := range []string{"unrelated.tmpl", "other-version.tmpl"} {
if count := fsys.openCount(name); count != 0 {
t.Fatalf("unselected content %q opened %d times", name, count)
}
}
for _, name := range []string{"target.yaml", "unrelated.yaml", "other-version.yaml"} {
if count := fsys.openCount(name); count != lookup {
t.Fatalf("metadata %q opens after lookup %d = %d, want %d", name, lookup, count, lookup)
}
}
}
}
func TestPromptDefinitionNormalizationRules(t *testing.T) {
tests := []struct {
name string
definition string
wantErr bool
wantDiagnostic string
wantSchemaPath string
}{
{
name: "missing version",
definition: `
id: normalization-rule
messages:
- role: user
content: test
output:
format: text
validation_mode: none
`,
wantErr: true,
wantDiagnostic: "version",
},
{
name: "blank input name",
definition: `
id: normalization-rule
version: "1"
inputs:
- name: " "
messages:
- role: user
content: test
output:
format: text
validation_mode: none
`,
wantErr: true,
wantDiagnostic: "input 0",
},
{
name: "blank message role",
definition: `
id: normalization-rule
version: "1"
messages:
- role: " "
content: test
output:
format: text
validation_mode: none
`,
wantErr: true,
wantDiagnostic: "role",
},
{
name: "invalid output format",
definition: `
id: normalization-rule
version: "1"
messages:
- role: user
content: test
output:
format: binary
validation_mode: none
`,
wantErr: true,
wantDiagnostic: "format",
},
{
name: "negative repair attempts",
definition: `
id: normalization-rule
version: "1"
messages:
- role: user
content: test
output:
format: text
validation_mode: none
repair_attempts: -1
`,
wantErr: true,
wantDiagnostic: "repair_attempts",
},
{
name: "repair attempts above maximum",
definition: `
id: normalization-rule
version: "1"
messages:
- role: user
content: test
output:
format: text
validation_mode: basic
repair_attempts: 4
`,
wantErr: true,
wantDiagnostic: "repair_attempts",
},
{
name: "none validation with repair attempts",
definition: `
id: normalization-rule
version: "1"
messages:
- role: user
content: test
output:
format: text
validation_mode: none
repair_attempts: 1
`,
wantErr: true,
wantDiagnostic: "repair_attempts",
},
{
name: "explicit blank default profile",
definition: `
id: normalization-rule
version: "1"
default_profile: " "
messages:
- role: user
content: test
output:
format: text
validation_mode: none
`,
wantErr: true,
wantDiagnostic: "default_profile",
},
{
name: "schema path normalization",
definition: `
id: normalization-rule
version: "1"
messages:
- role: user
content: test
output:
format: json
validation_mode: json_schema
schema_path: ' schema.json '
`,
wantSchemaPath: "schema.json",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
repo := NewFSRepository(fstest.MapFS{
"definition.yaml": &fstest.MapFile{Data: []byte(tt.definition)},
}, ".")
got, err := repo.GetPromptDefinition(context.Background(), "normalization-rule", "")
if tt.wantErr {
if !errors.Is(err, ErrInvalidPromptDefinition) {
t.Fatalf("expected ErrInvalidPromptDefinition, got %v", err)
}
if !strings.Contains(err.Error(), tt.wantDiagnostic) {
t.Fatalf("expected error containing %q, got %v", tt.wantDiagnostic, err)
}
return
}
if err != nil {
t.Fatalf("load prompt definition: %v", err)
}
if got.Validation.SchemaPath != tt.wantSchemaPath {
t.Fatalf("schema path = %q, want %q", got.Validation.SchemaPath, tt.wantSchemaPath)
}
})
}
}
type promptRepositorySource struct {
name string
newRepository func(t *testing.T, files map[string]string) Repository
}
func promptRepositorySources() []promptRepositorySource {
return []promptRepositorySource{
{
name: "operating system",
newRepository: func(t *testing.T, files map[string]string) Repository {
t.Helper()
root := t.TempDir()
for name, content := range files {
filePath := filepath.Join(root, filepath.FromSlash(name))
if err := os.MkdirAll(filepath.Dir(filePath), 0o755); err != nil {
t.Fatalf("create prompt directory: %v", err)
}
writePromptTestFile(t, filePath, content)
}
return NewFilesystemRepository(root)
},
},
{
name: "filesystem",
newRepository: func(t *testing.T, files map[string]string) Repository {
t.Helper()
fsys := make(fstest.MapFS, len(files))
for name, content := range files {
fsys[name] = &fstest.MapFile{Data: []byte(strings.TrimLeft(content, "\n"))}
}
return NewFSRepository(fsys, ".")
},
},
}
}
func BenchmarkPromptRepositoryLookup(b *testing.B) {
for _, size := range []int{10, 1000} {
b.Run(fmt.Sprintf("catalog-%d", size), func(b *testing.B) {
files := fstest.MapFS{
"target.yaml": &fstest.MapFile{Data: []byte(`
id: target
version: "1"
messages:
- role: user
content_file: target.tmpl
output:
format: text
validation_mode: none
`)},
"target.tmpl": &fstest.MapFile{Data: []byte("selected")},
}
metadataNames := []string{"target.yaml"}
contentNames := make([]string, 0, size-1)
for i := 1; i < size; i++ {
definitionName := fmt.Sprintf("prompt-%04d.yaml", i)
contentName := fmt.Sprintf("prompt-%04d.tmpl", i)
files[definitionName] = &fstest.MapFile{Data: []byte(fmt.Sprintf(`
id: prompt-%04d
version: "1"
messages:
- role: user
content_file: %s
output:
format: text
validation_mode: none
`, i, contentName))}
files[contentName] = &fstest.MapFile{Data: []byte("unrelated")}
metadataNames = append(metadataNames, definitionName)
contentNames = append(contentNames, contentName)
}
fsys := &recordingFS{FS: files}
repo := NewFSRepository(fsys, ".")
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
if _, err := repo.GetPromptDefinition(context.Background(), "target", "1"); err != nil {
b.Fatal(err)
}
}
b.StopTimer()
if got := fsys.openCount("target.tmpl"); got != b.N {
b.Fatalf("selected content opens = %d, want %d", got, b.N)
}
for _, name := range contentNames {
if got := fsys.openCount(name); got != 0 {
b.Fatalf("unrelated content %q opened %d times", name, got)
}
}
for _, name := range metadataNames {
if got := fsys.openCount(name); got != b.N {
b.Fatalf("metadata %q opens = %d, want %d", name, got, b.N)
}
}
})
}
} }
func assertCacheControl(t *testing.T, got *domain.CacheControl, wantType domain.CacheControlType, wantTTL string) { func assertCacheControl(t *testing.T, got *domain.CacheControl, wantType domain.CacheControlType, wantTTL string) {

View File

@@ -0,0 +1,63 @@
package promptdef
import (
"context"
"errors"
"io/fs"
"os"
"gitea.maximumdirect.net/eric/promptkit/internal/filecatalog"
)
type promptDefinitionSource interface {
contentSourceRoot
findYAMLFiles(context.Context) ([]string, error)
readDefinition(string) ([]byte, error)
displayPath(string) string
}
type osPromptSource struct {
root string
contentRoot osContentSourceRoot
}
func (s osPromptSource) findYAMLFiles(ctx context.Context) ([]string, error) {
return filecatalog.FindYAMLFiles(ctx, s.root)
}
func (s osPromptSource) readDefinition(name string) ([]byte, error) {
return os.ReadFile(name)
}
func (s osPromptSource) displayPath(name string) string {
return filecatalog.RelativePath(s.root, name)
}
func (s osPromptSource) readContentFile(sourcePath string, contentFile string) (string, string, error) {
return s.contentRoot.readContentFile(sourcePath, contentFile)
}
type fsPromptSource struct {
fsys fs.FS
root string
contentRoot contentSourceRoot
}
func (s fsPromptSource) findYAMLFiles(ctx context.Context) ([]string, error) {
if s.fsys == nil {
return nil, errors.New("filesystem is nil")
}
return filecatalog.FindFSYAMLFiles(ctx, s.fsys, s.root)
}
func (s fsPromptSource) readDefinition(name string) ([]byte, error) {
return fs.ReadFile(s.fsys, name)
}
func (s fsPromptSource) displayPath(name string) string {
return filecatalog.DisplayPath(s.root, name)
}
func (s fsPromptSource) readContentFile(sourcePath string, contentFile string) (string, string, error) {
return s.contentRoot.readContentFile(sourcePath, contentFile)
}

View File

@@ -0,0 +1,83 @@
package usecase
import (
"context"
"errors"
"math"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
func TestRunnerPrepareExecutionRejectsInvalidExecutionSettings(t *testing.T) {
runner := NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
defaultArtifactReader(),
defaultRenderer(),
&fakeLLM{forbid: true},
nil,
nil,
)
_, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
Execution: &domain.ExecutionTargetOverride{TopP: float64Ptr(math.Inf(-1))},
})
if !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("expected ErrInvalidRequest, got %v", err)
}
}
func TestRunnerPrepareExecutionValidatesAndNormalizesRequestEndpoints(t *testing.T) {
newRunner := func() *Runner {
return NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
defaultArtifactReader(),
defaultRenderer(),
&fakeLLM{forbid: true},
nil,
nil,
)
}
invalidEndpoints := []string{
"/v1",
"https:///v1",
"ftp://provider.example/v1",
"https://user@provider.example/v1",
"https://provider.example/v1?mode=chat",
"https://provider.example/v1#chat",
}
for _, endpoint := range invalidEndpoints {
t.Run(endpoint, func(t *testing.T) {
_, err := newRunner().PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
Execution: &domain.ExecutionTargetOverride{Endpoint: endpoint},
})
if !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("expected ErrInvalidRequest, got %v", err)
}
})
}
prepared, err := newRunner().PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
Execution: &domain.ExecutionTargetOverride{Endpoint: " https://provider.example/nested/v1 "},
})
if err != nil {
t.Fatalf("prepare normalized endpoint: %v", err)
}
if got := prepared.Details().EffectiveModelParams.Endpoint; got != "https://provider.example/nested/v1" {
t.Fatalf("effective endpoint = %q", got)
}
}

View File

@@ -0,0 +1,19 @@
package usecase
import "gitea.maximumdirect.net/eric/promptkit/internal/domain"
func newGenerationRequest(
prompt domain.RenderedPrompt,
sessionID string,
target domain.ExecutionTarget,
targetPresence domain.ExecutionTargetPresence,
structuredOutput *domain.StructuredOutputSpec,
) domain.GenerateRequest {
prompt.SessionID = sessionID
return domain.GenerateRequest{
Prompt: prompt,
Target: target,
TargetPresence: targetPresence,
StructuredOutput: structuredOutput,
}
}

View File

@@ -0,0 +1,187 @@
package usecase
import (
"context"
"errors"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
type outputContractTestCollaborators struct {
artifacts *fakeArtifactReader
renderer *fakeRenderer
llm *fakeLLM
validator *recordingValidationPreparer
admitter *fakeRunAdmitter
}
func newOutputContractTestRunner() (*Runner, outputContractTestCollaborators) {
collaborators := outputContractTestCollaborators{
artifacts: defaultArtifactReader(),
renderer: defaultRenderer(),
llm: &fakeLLM{forbid: true},
validator: &recordingValidationPreparer{plan: &recordingPreparedValidation{}},
admitter: &fakeRunAdmitter{},
}
return NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
collaborators.artifacts,
collaborators.renderer,
collaborators.llm,
collaborators.validator,
collaborators.admitter,
), collaborators
}
func TestRunnerPreparationNormalizesOutputContractConsistently(t *testing.T) {
tests := []struct {
name string
override domain.OutputContract
want domain.OutputContract
}{
{
name: "empty replacement format defaults to text",
override: domain.OutputContract{ValidationMode: domain.ValidationNone},
want: domain.OutputContract{Format: domain.FormatText, ValidationMode: domain.ValidationNone},
},
{
name: "markdown basic replacement",
override: domain.OutputContract{Format: domain.FormatMarkdown, ValidationMode: domain.ValidationBasic},
want: domain.OutputContract{Format: domain.FormatMarkdown, ValidationMode: domain.ValidationBasic},
},
{
name: "json replacement preserves non-schema fields",
override: domain.OutputContract{
Format: domain.FormatJSON,
ValidationMode: domain.ValidationJSON,
SchemaPath: "ignored.json",
RepairAttempts: 2,
},
want: domain.OutputContract{
Format: domain.FormatJSON,
ValidationMode: domain.ValidationJSON,
SchemaPath: "ignored.json",
RepairAttempts: 2,
},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
runner, _ := newOutputContractTestRunner()
req := domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
Validation: &tt.override,
}
prepared, err := runner.Prepare(context.Background(), req)
if err != nil {
t.Fatalf("prepare: %v", err)
}
preparedExecution, err := runner.PrepareExecution(context.Background(), req)
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
details := preparedExecution.Details()
if details == nil {
t.Fatal("prepared execution returned nil details")
}
if prepared.OutputContract != tt.want || details.OutputContract != tt.want {
t.Fatalf("output contracts = (%+v, %+v), want %+v", prepared.OutputContract, details.OutputContract, tt.want)
}
})
}
}
func TestRunnerPreparationRejectsInvalidOutputContractsBeforeCompletion(t *testing.T) {
tests := []struct {
name string
override domain.OutputContract
}{
{name: "unsupported format", override: domain.OutputContract{Format: "binary", ValidationMode: domain.ValidationNone}},
{name: "empty validation mode", override: domain.OutputContract{Format: domain.FormatText}},
{name: "negative repair attempts", override: domain.OutputContract{Format: domain.FormatText, ValidationMode: domain.ValidationNone, RepairAttempts: -1}},
{name: "repair attempts above maximum", override: domain.OutputContract{Format: domain.FormatText, ValidationMode: domain.ValidationBasic, RepairAttempts: 4}},
{name: "none validation with repair attempts", override: domain.OutputContract{Format: domain.FormatText, ValidationMode: domain.ValidationNone, RepairAttempts: 1}},
{name: "json schema without path", override: domain.OutputContract{Format: domain.FormatJSON, ValidationMode: domain.ValidationJSONSchema}},
}
for _, tt := range tests {
for _, operation := range []string{"Prepare", "PrepareExecution"} {
t.Run(tt.name+"/"+operation, func(t *testing.T) {
runner, collaborators := newOutputContractTestRunner()
req := domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
Validation: &tt.override,
}
var err error
switch operation {
case "Prepare":
var prepared *domain.PreparedRun
prepared, err = runner.Prepare(context.Background(), req)
if prepared != nil {
t.Fatalf("expected no partial prepared run, got %+v", prepared)
}
case "PrepareExecution":
var prepared *PreparedExecution
prepared, err = runner.PrepareExecution(context.Background(), req)
if prepared != nil {
t.Fatalf("expected no partial prepared execution, got %+v", prepared)
}
default:
t.Fatalf("unknown operation %q", operation)
}
if !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("expected ErrInvalidRequest, got %v", err)
}
assertOutputContractCompletionSkipped(t, collaborators)
})
}
}
}
func TestRunnerRunRejectsInvalidOutputContractBeforeAdmission(t *testing.T) {
runner, collaborators := newOutputContractTestRunner()
result, err := runner.Run(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
Validation: &domain.OutputContract{
Format: domain.OutputFormat("binary"),
ValidationMode: domain.ValidationNone,
},
})
if result != nil {
t.Fatalf("expected no partial result, got %+v", result)
}
if !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("expected ErrInvalidRequest, got %v", err)
}
assertOutputContractCompletionSkipped(t, collaborators)
}
func assertOutputContractCompletionSkipped(t *testing.T, collaborators outputContractTestCollaborators) {
t.Helper()
if collaborators.artifacts.calls != 0 || collaborators.renderer.calls != 0 ||
collaborators.validator.prepareCalls != 0 || collaborators.validator.directValidateCalls != 0 ||
len(collaborators.admitter.backendIDs) != 0 || collaborators.llm.calls != 0 {
t.Fatalf(
"invalid output contract reached downstream work: artifacts=%d renderer=%d prepare_validation=%d validation=%d admissions=%d generation=%d",
collaborators.artifacts.calls,
collaborators.renderer.calls,
collaborators.validator.prepareCalls,
collaborators.validator.directValidateCalls,
len(collaborators.admitter.backendIDs),
collaborators.llm.calls,
)
}
}

View File

@@ -42,25 +42,12 @@ func (r *Runner) PrepareExecution(ctx context.Context, req domain.RunRequest) (*
return nil, err return nil, err
} }
validationPlan, err := r.prepareValidation(ctx, state.effectiveContract) operation, err := r.completePreparation(ctx, req, state)
if err != nil { if err != nil {
return nil, err return nil, err
} }
structuredOutput, err := r.structuredOutputFromValidationPlan( executionSnapshot, err := clonePreparedRun(operation.run)
state.definition,
state.effectiveContract,
validationPlan,
)
if err != nil {
return nil, err
}
prepared, err := r.completePreparationWithStructuredOutput(ctx, req, state, structuredOutput)
if err != nil {
return nil, err
}
executionSnapshot, err := clonePreparedRun(prepared)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: failed to copy prepared execution: %v", ErrInvalidRequest, err) return nil, fmt.Errorf("%w: failed to copy prepared execution: %v", ErrInvalidRequest, err)
} }
@@ -76,52 +63,12 @@ func (r *Runner) PrepareExecution(ctx context.Context, req domain.RunRequest) (*
details: details, details: details,
payload: &preparedExecutionPayload{ payload: &preparedExecutionPayload{
prepared: executionSnapshot, prepared: executionSnapshot,
validation: validationPlan, validation: operation.validation,
directKey: state.effectiveModel.APIKey, directKey: state.effectiveModel.APIKey,
}, },
}, nil }, nil
} }
func (r *Runner) prepareValidation(
ctx context.Context,
contract domain.OutputContract,
) (validate.PreparedValidation, error) {
if r.validator == nil {
return noOpPreparedValidation{contract: contract}, nil
}
preparer, ok := r.validator.(validate.ValidationPreparer)
if !ok {
return nil, fmt.Errorf("%w: validator does not support prepared validation", ErrValidation)
}
plan, err := preparer.PrepareValidation(ctx, contract)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrValidation, err)
}
if plan == nil {
return nil, fmt.Errorf("%w: validator returned nil prepared validation", ErrValidation)
}
return plan, nil
}
func (r *Runner) structuredOutputFromValidationPlan(
def *domain.PromptDefinition,
contract domain.OutputContract,
plan validate.PreparedValidation,
) (*domain.StructuredOutputSpec, error) {
if contract.ValidationMode != domain.ValidationJSONSchema {
return nil, nil
}
schemaDocument := plan.SchemaDocument()
if schemaDocument == nil {
if r.validator == nil {
return nil, nil
}
return nil, fmt.Errorf("%w: prepared json_schema validation has no schema document", ErrValidation)
}
return structuredOutputSpec(def, schemaDocument), nil
}
// Details returns a fresh credential-redacted copy of the prepared run. // Details returns a fresh credential-redacted copy of the prepared run.
func (p *PreparedExecution) Details() *domain.PreparedRun { func (p *PreparedExecution) Details() *domain.PreparedRun {
if p == nil { if p == nil {

View File

@@ -3,11 +3,12 @@ package usecase
import ( import (
"context" "context"
"errors" "errors"
"os" "fmt"
"reflect" "reflect"
"testing" "testing"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/llm"
"gitea.maximumdirect.net/eric/promptkit/internal/validate" "gitea.maximumdirect.net/eric/promptkit/internal/validate"
) )
@@ -52,6 +53,12 @@ type recordingValidationPreparer struct {
directValidateCalls int directValidateCalls int
} }
type validationOnly struct{}
func (validationOnly) Validate(context.Context, *domain.Artifact, domain.OutputContract) (domain.ValidationResult, error) {
return domain.ValidationResult{}, nil
}
func (v *recordingValidationPreparer) Validate( func (v *recordingValidationPreparer) Validate(
context.Context, context.Context,
*domain.Artifact, *domain.Artifact,
@@ -164,15 +171,107 @@ func TestRunnerPrepareExecutionCompletesWithoutAdmissionOrGeneration(t *testing.
} }
} }
func TestRunnerRunPreparedRechecksEnvironmentCredentialBeforeAdmission(t *testing.T) { func TestRunnerPreparationRejectsExcessivelyDeepPreparedSchema(t *testing.T) {
operations := []struct {
name string
run func(*Runner, domain.RunRequest) error
}{
{
name: "Prepare",
run: func(runner *Runner, request domain.RunRequest) error {
_, err := runner.Prepare(context.Background(), request)
return err
},
},
{
name: "Run",
run: func(runner *Runner, request domain.RunRequest) error {
_, err := runner.Run(context.Background(), request)
return err
},
},
{
name: "PrepareExecution",
run: func(runner *Runner, request domain.RunRequest) error {
_, err := runner.PrepareExecution(context.Background(), request)
return err
},
},
}
for _, operation := range operations {
t.Run(operation.name, func(t *testing.T) {
def := promptDef(domain.FormatJSON, domain.ValidationJSONSchema, 0)
def.Validation.SchemaPath = "schema.json"
llmClient := &fakeLLM{forbid: true}
validator := &recordingValidationPreparer{
plan: &recordingPreparedValidation{schemaDocument: excessivelyDeepPreparedJSONValue()},
}
runner := NewRunner(
&fakePromptRepo{def: def},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
defaultArtifactReader(),
defaultRenderer(),
llmClient,
validator,
nil,
)
err := operation.run(runner, domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if !errors.Is(err, ErrValidation) {
t.Fatalf("expected ErrValidation, got %v", err)
}
if llmClient.calls != 0 {
t.Fatalf("invalid prepared schema reached generation: %d calls", llmClient.calls)
}
})
}
}
func excessivelyDeepPreparedJSONValue() any {
const clearlyUnsafeContainerDepth = 1_000
var value any = true
for level := 0; level < clearlyUnsafeContainerDepth; level++ {
value = map[string]any{"child": value}
}
return value
}
func TestRunnerRunPreparedCredentialAvailabilityBeforeAdmission(t *testing.T) {
const environmentName = "PROMPTKIT_PREPARED_EXECUTION_TEST_KEY" const environmentName = "PROMPTKIT_PREPARED_EXECUTION_TEST_KEY"
tests := []struct {
name string
apiKeyRequired bool
profileEnv bool
overrideEnv bool
wantFailure bool
}{
{name: "optional environment becomes unavailable", profileEnv: true},
{
name: "required request environment becomes unavailable",
apiKeyRequired: true,
overrideEnv: true,
wantFailure: true,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
t.Setenv(environmentName, "available-during-preparation") t.Setenv(environmentName, "available-during-preparation")
profile := defaultExecutionProfile() profile := defaultExecutionProfile()
profile.APIKeyRequired = tc.apiKeyRequired
if tc.profileEnv {
profile.APIKeyEnv = environmentName profile.APIKeyEnv = environmentName
}
validator := &recordingValidationPreparer{plan: &recordingPreparedValidation{}} validator := &recordingValidationPreparer{plan: &recordingPreparedValidation{}}
admitter := &fakeRunAdmitter{} admitter := &fakeRunAdmitter{}
llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: "unexpected"}} llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: "ok"}}
runner := NewRunner( runner := NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)}, &fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": profile}}, &fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": profile}},
@@ -183,30 +282,36 @@ func TestRunnerRunPreparedRechecksEnvironmentCredentialBeforeAdmission(t *testin
validator, validator,
admitter, admitter,
) )
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{ request := domain.RunRequest{PromptID: "p", ProfileID: "exec", Inputs: singleInputRef()}
PromptID: "p", if tc.overrideEnv {
ProfileID: "exec", request.Execution = &domain.ExecutionTargetOverride{APIKeyEnv: environmentName}
Inputs: singleInputRef(), }
}) prepared, err := runner.PrepareExecution(context.Background(), request)
if err != nil { if err != nil {
t.Fatalf("prepare execution: %v", err) t.Fatalf("prepare execution: %v", err)
} }
if err := os.Unsetenv(environmentName); err != nil { t.Setenv(environmentName, "")
t.Fatalf("unset credential environment: %v", err)
}
result, err := runner.RunPrepared(context.Background(), prepared) result, err := runner.RunPrepared(context.Background(), prepared)
if result != nil { if tc.wantFailure {
t.Fatalf("credential failure returned partial result: %+v", result) if result != nil || !errors.Is(err, ErrInvalidRequest) || !errors.Is(err, ErrAPIKeyEnvMissing) {
} t.Fatalf("required credential result = (%+v, %v)", result, err)
if !errors.Is(err, ErrInvalidRequest) || !errors.Is(err, ErrAPIKeyEnvMissing) {
t.Fatalf("credential error identities are missing: %v", err)
} }
if len(admitter.backendIDs) != 0 || llmClient.calls != 0 { if len(admitter.backendIDs) != 0 || llmClient.calls != 0 {
t.Fatalf("credential failure reached admission or generation: admission=%v generation=%d", admitter.backendIDs, llmClient.calls) t.Fatalf("required credential reached admission or generation: admission=%v generation=%d", admitter.backendIDs, llmClient.calls)
}
} else {
if result == nil || err != nil {
t.Fatalf("optional credential result = (%+v, %v), want success", result, err)
}
if len(admitter.backendIDs) != 1 || llmClient.calls != 1 {
t.Fatalf("optional credential admission=%v generation=%d, want one each", admitter.backendIDs, llmClient.calls)
}
} }
if _, err := runner.RunPrepared(context.Background(), prepared); !errors.Is(err, ErrInvalidRequest) { if _, err := runner.RunPrepared(context.Background(), prepared); !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("credential failure did not consume execution: %v", err) t.Fatalf("execution outcome did not consume handle: %v", err)
}
})
} }
} }
@@ -271,18 +376,31 @@ func TestRunnerRunPreparedUsesFrozenValidationForInitialAndRepairOutputs(t *test
}, },
} }
repairer := &fakeRepairer{ repairer := &fakeRepairer{
responses: []*domain.GenerateResponse{{Content: `{"repaired":true}`}}, responses: []*domain.GenerateResponse{{
Content: `{"repaired":true}`,
Usage: domain.TokenUsage{
PromptTokens: 2, CompletionTokens: 3, TotalTokens: 5,
CachedTokens: 7, CacheWriteTokens: 11,
},
}},
} }
admitter := &fakeRunAdmitter{} admitter := &fakeRunAdmitter{}
reader := defaultArtifactReader() reader := defaultArtifactReader()
renderer := defaultRenderer() renderer := defaultRenderer()
client := &sequenceLLM{responses: []*domain.GenerateResponse{{
Content: `{"broken":true}`,
Usage: domain.TokenUsage{
PromptTokens: 13, CompletionTokens: 17, TotalTokens: 19,
CachedTokens: 23, CacheWriteTokens: 29,
},
}}}
runner := NewRunnerWithRepairer( runner := NewRunnerWithRepairer(
&fakePromptRepo{def: promptDef(domain.FormatJSON, domain.ValidationJSON, 1)}, &fakePromptRepo{def: promptDef(domain.FormatJSON, domain.ValidationJSON, 1)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}}, &fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil, nil,
reader, reader,
renderer, renderer,
&fakeLLM{resp: &domain.GenerateResponse{Content: `{"broken":true}`}}, client,
validator, validator,
repairer, repairer,
admitter, admitter,
@@ -306,9 +424,20 @@ func TestRunnerRunPreparedUsesFrozenValidationForInitialAndRepairOutputs(t *test
if validator.directValidateCalls != 0 || repairer.calls != 1 { if validator.directValidateCalls != 0 || repairer.calls != 1 {
t.Fatalf("validation/repair calls=(direct=%d repair=%d), want (0, 1)", validator.directValidateCalls, repairer.calls) t.Fatalf("validation/repair calls=(direct=%d repair=%d), want (0, 1)", validator.directValidateCalls, repairer.calls)
} }
if len(client.requests) != 1 || len(repairer.reqs) != 1 ||
!reflect.DeepEqual(repairer.reqs[0].OriginalMessages, client.requests[0].Prompt.Messages) {
t.Fatalf("initial and repair messages = (%#v, %#v)", client.requests, repairer.reqs)
}
if result.Validation.Status != domain.ValidationPassed || result.Validation.RepairAttempts != 1 { if result.Validation.Status != domain.ValidationPassed || result.Validation.RepairAttempts != 1 {
t.Fatalf("unexpected repaired validation result: %+v", result.Validation) t.Fatalf("unexpected repaired validation result: %+v", result.Validation)
} }
wantUsage := domain.TokenUsage{
PromptTokens: 15, CompletionTokens: 20, TotalTokens: 24,
CachedTokens: 30, CacheWriteTokens: 40,
}
if result.Usage != wantUsage {
t.Fatalf("prepared cumulative usage = %+v, want %+v", result.Usage, wantUsage)
}
if admitter.releaseCalls != 1 { if admitter.releaseCalls != 1 {
t.Fatalf("admission releases=%d, want 1", admitter.releaseCalls) t.Fatalf("admission releases=%d, want 1", admitter.releaseCalls)
} }
@@ -321,6 +450,7 @@ func TestRunnerRunPreparedReleasesAdmissionAcrossExecutionErrors(t *testing.T) {
generationFailure := errors.New("generation failed") generationFailure := errors.New("generation failed")
validationFailure := errors.New("validation failed") validationFailure := errors.New("validation failed")
repairFailure := errors.New("repair failed") repairFailure := errors.New("repair failed")
repairInvalidRequest := fmt.Errorf("repair request: %w", llm.ErrInvalidRequest)
tests := []struct { tests := []struct {
name string name string
@@ -328,6 +458,7 @@ func TestRunnerRunPreparedReleasesAdmissionAcrossExecutionErrors(t *testing.T) {
validation *recordingPreparedValidation validation *recordingPreparedValidation
repairer *fakeRepairer repairer *fakeRepairer
wantError error wantError error
wantSource error
}{ }{
{ {
name: "generation failure", name: "generation failure",
@@ -353,7 +484,50 @@ func TestRunnerRunPreparedReleasesAdmissionAcrossExecutionErrors(t *testing.T) {
}}, }},
}, },
repairer: &fakeRepairer{err: repairFailure}, repairer: &fakeRepairer{err: repairFailure},
wantError: ErrValidation, wantError: ErrLLMGenerate,
wantSource: repairFailure,
},
{
name: "repair invalid request",
validation: &recordingPreparedValidation{
results: []domain.ValidationResult{{
Status: domain.ValidationFailed,
Mode: domain.ValidationJSON,
Errors: []string{"invalid"},
IsValid: false,
}},
},
repairer: &fakeRepairer{err: repairInvalidRequest},
wantError: ErrInvalidRequest,
wantSource: repairInvalidRequest,
},
{
name: "repair cancellation",
validation: &recordingPreparedValidation{
results: []domain.ValidationResult{{
Status: domain.ValidationFailed,
Mode: domain.ValidationJSON,
Errors: []string{"invalid"},
IsValid: false,
}},
},
repairer: &fakeRepairer{err: context.Canceled},
wantError: ErrLLMGenerate,
wantSource: context.Canceled,
},
{
name: "repair deadline",
validation: &recordingPreparedValidation{
results: []domain.ValidationResult{{
Status: domain.ValidationFailed,
Mode: domain.ValidationJSON,
Errors: []string{"invalid"},
IsValid: false,
}},
},
repairer: &fakeRepairer{err: context.DeadlineExceeded},
wantError: ErrLLMGenerate,
wantSource: context.DeadlineExceeded,
}, },
} }
@@ -392,6 +566,9 @@ func TestRunnerRunPreparedReleasesAdmissionAcrossExecutionErrors(t *testing.T) {
if result != nil || !errors.Is(err, test.wantError) { if result != nil || !errors.Is(err, test.wantError) {
t.Fatalf("run prepared=(%+v, %v), want %v", result, err, test.wantError) t.Fatalf("run prepared=(%+v, %v), want %v", result, err, test.wantError)
} }
if test.wantSource != nil && !errors.Is(err, test.wantSource) {
t.Fatalf("run prepared error = %v, want source %v", err, test.wantSource)
}
if len(admitter.backendIDs) != 1 || admitter.releaseCalls != 1 { if len(admitter.backendIDs) != 1 || admitter.releaseCalls != 1 {
t.Fatalf( t.Fatalf(
"admission calls=%#v releases=%d, want one each", "admission calls=%#v releases=%d, want one each",
@@ -492,7 +669,7 @@ func TestRunnerPrepareExecutionRequiresValidationPreparer(t *testing.T) {
reader, reader,
defaultRenderer(), defaultRenderer(),
&fakeLLM{forbid: true}, &fakeLLM{forbid: true},
&fakeValidator{}, validationOnly{},
nil, nil,
) )

View File

@@ -57,14 +57,19 @@ func (r *Runner) resolveProfileSelection(
}, nil }, nil
} }
func validateResolvedExecutionTarget(target domain.ExecutionTarget) error { func normalizeResolvedExecutionTarget(target domain.ExecutionTarget) (domain.ExecutionTarget, error) {
if strings.TrimSpace(target.Endpoint) == "" { endpoint, err := domain.NormalizeOpenAICompatibleBaseEndpoint(target.Endpoint)
return errors.New("execution endpoint is required") if err != nil {
return domain.ExecutionTarget{}, fmt.Errorf("execution endpoint: %w", err)
} }
target.Endpoint = endpoint
if strings.TrimSpace(target.Model) == "" { if strings.TrimSpace(target.Model) == "" {
return errors.New("execution model is required") return domain.ExecutionTarget{}, errors.New("execution model is required")
} }
return nil if err := domain.ValidateExecutionTargetSettings(target); err != nil {
return domain.ExecutionTarget{}, err
}
return target, nil
} }
// InspectProfile resolves one explicit profile without prompt or execution work. // InspectProfile resolves one explicit profile without prompt or execution work.
@@ -86,13 +91,11 @@ func (r *Runner) InspectProfile(
if err != nil { if err != nil {
return nil, err return nil, err
} }
target, _, err := resolveExecutionTarget(selection.backend, selection.profile, nil) target, _ := resolveExecutionTarget(selection.backend, selection.profile, nil)
target, err = normalizeResolvedExecutionTarget(target)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %w", ErrProfileLoad, err) return nil, fmt.Errorf("%w: %w", ErrProfileLoad, err)
} }
if err := validateResolvedExecutionTarget(target); err != nil {
return nil, fmt.Errorf("%w: %w", ErrProfileLoad, err)
}
target.APIKey = "" target.APIKey = ""
return &domain.ProfileInspection{ return &domain.ProfileInspection{

View File

@@ -2,23 +2,34 @@ package usecase
import ( import (
"context" "context"
"encoding/json"
"errors" "errors"
"fmt" "fmt"
"strings" "strings"
"unicode/utf8"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/llm" "gitea.maximumdirect.net/eric/promptkit/internal/llm"
) )
const (
maxRepairDiagnosticBytes = 64 * 1024
omittedRepairDiagnostics = "additional validation diagnostics were omitted"
)
// OutputRepairer generates a corrected candidate after validation fails.
type OutputRepairer interface { type OutputRepairer interface {
Repair(ctx context.Context, req RepairRequest) (*domain.GenerateResponse, error) Repair(ctx context.Context, req RepairRequest) (*domain.GenerateResponse, error)
} }
// RepairRequest contains the immutable execution state needed for one correction.
type RepairRequest struct { type RepairRequest struct {
OriginalMessages []domain.RenderedMessage
PreviousOutput string PreviousOutput string
ValidationErrors []string ValidationErrors []string
SessionID string SessionID string
Target domain.ExecutionTarget Target domain.ExecutionTarget
TargetPresence domain.ExecutionTargetPresence
StructuredOutput *domain.StructuredOutputSpec StructuredOutput *domain.StructuredOutputSpec
Attempt int Attempt int
MaxAttempts int MaxAttempts int
@@ -29,6 +40,7 @@ type defaultOutputRepairer struct {
llm llm.Client llm llm.Client
} }
// NewDefaultOutputRepairer constructs the standard internal output repairer.
func NewDefaultOutputRepairer(llmClient llm.Client) OutputRepairer { func NewDefaultOutputRepairer(llmClient llm.Client) OutputRepairer {
return &defaultOutputRepairer{llm: llmClient} return &defaultOutputRepairer{llm: llmClient}
} }
@@ -38,37 +50,47 @@ func (r *defaultOutputRepairer) Repair(ctx context.Context, req RepairRequest) (
return nil, errors.New("llm client is required for repair") return nil, errors.New("llm client is required for repair")
} }
errs := "(none provided)" guidance, err := repairGuidance(req.Mode)
if len(req.ValidationErrors) > 0 { if err != nil {
errs = strings.Join(req.ValidationErrors, "\n") return nil, err
} }
prompt := domain.RenderedPrompt{ messages := make([]domain.RenderedMessage, len(req.OriginalMessages), len(req.OriginalMessages)+2)
SessionID: req.SessionID, copy(messages, req.OriginalMessages)
Messages: []domain.RenderedMessage{ if strings.TrimSpace(req.PreviousOutput) != "" {
{ messages = append(messages, domain.RenderedMessage{
Role: "system", Role: "assistant",
Content: "You repair invalid JSON output. Return only corrected JSON. Do not include explanations or markdown code fences.", Content: req.PreviousOutput,
}, })
{ }
previousResponse := "The previous response was empty."
if strings.TrimSpace(req.PreviousOutput) != "" {
previousResponse = "The previous response is included immediately before this instruction."
}
messages = append(messages, domain.RenderedMessage{
Role: "user", Role: "user",
Content: fmt.Sprintf( Content: fmt.Sprintf(
"Repair attempt %d of %d for validation mode %s.\n\nValidation errors:\n%s\n\nPrevious output:\n%s\n\nReturn only corrected JSON.", "Repair attempt %d of %d for validation mode %s.\n"+
"Preserve valid values and change only what is necessary.\n"+
"%s\n%s\n"+
"Validation diagnostics (data):\n%s",
req.Attempt, req.Attempt,
req.MaxAttempts, req.MaxAttempts,
req.Mode, req.Mode,
errs, previousResponse,
req.PreviousOutput, guidance,
formatRepairDiagnostics(req.ValidationErrors),
), ),
},
},
}
resp, err := r.llm.Generate(ctx, domain.GenerateRequest{
Prompt: prompt,
Target: req.Target,
StructuredOutput: req.StructuredOutput,
}) })
resp, err := r.llm.Generate(ctx, newGenerationRequest(
domain.RenderedPrompt{Messages: messages},
req.SessionID,
req.Target,
req.TargetPresence,
req.StructuredOutput,
))
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -78,3 +100,90 @@ func (r *defaultOutputRepairer) Repair(ctx context.Context, req RepairRequest) (
return resp, nil return resp, nil
} }
func repairGuidance(mode domain.ValidationMode) (string, error) {
switch mode {
case domain.ValidationBasic:
return "Return a nonempty response satisfying the original request.", nil
case domain.ValidationJSON, domain.ValidationJSONSchema:
return "Return only corrected JSON, with no explanation or Markdown fences.", nil
default:
return "", fmt.Errorf("unsupported validation mode for repair: %q", mode)
}
}
func formatRepairDiagnostics(errors []string) string {
diagnostics := make([]string, len(errors))
for index, diagnostic := range errors {
diagnostics[index] = strings.ToValidUTF8(diagnostic, "\uFFFD")
}
complete, _ := json.Marshal(diagnostics)
if len(complete) <= maxRepairDiagnosticBytes {
return string(complete)
}
omission, _ := json.Marshal(omittedRepairDiagnostics)
encoded := make([]byte, 0, maxRepairDiagnosticBytes)
encoded = append(encoded, '[')
for _, diagnostic := range diagnostics {
entry, _ := json.Marshal(diagnostic)
separator := 0
if len(encoded) > 1 {
separator = 1
}
available := maxRepairDiagnosticBytes - len(encoded) - separator - 1 - len(omission) - 1
if len(entry) <= available {
if separator != 0 {
encoded = append(encoded, ',')
}
encoded = append(encoded, entry...)
continue
}
if available < len(`""`) {
break
}
if separator != 0 {
encoded = append(encoded, ',')
}
encoded = append(encoded, truncateDiagnosticJSONValue(diagnostic, available)...)
break
}
if len(encoded) > 1 {
encoded = append(encoded, ',')
}
encoded = append(encoded, omission...)
encoded = append(encoded, ']')
return string(encoded)
}
func truncateDiagnosticJSONValue(value string, maxBytes int) []byte {
if maxBytes < len(`""`) {
return nil
}
boundaries := []int{0}
for end := 0; end < len(value); {
_, size := utf8.DecodeRuneInString(value[end:])
end += size
if end+len(`""`) > maxBytes {
break
}
boundaries = append(boundaries, end)
}
low, high := 0, len(boundaries)-1
best := []byte(`""`)
for low <= high {
mid := low + (high-low)/2
candidate, _ := json.Marshal(value[:boundaries[mid]])
if len(candidate) <= maxBytes {
best = candidate
low = mid + 1
continue
}
high = mid - 1
}
return best
}

View File

@@ -0,0 +1,349 @@
package usecase
import (
"context"
"encoding/json"
"errors"
"fmt"
"reflect"
"strings"
"testing"
"unicode/utf8"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
type recordingRepairClient struct {
requests []domain.GenerateRequest
response *domain.GenerateResponse
err error
}
func (c *recordingRepairClient) Generate(_ context.Context, req domain.GenerateRequest) (*domain.GenerateResponse, error) {
c.requests = append(c.requests, req)
return c.response, c.err
}
func TestDefaultOutputRepairerBuildsFullContextRequest(t *testing.T) {
client := &recordingRepairClient{response: &domain.GenerateResponse{Content: "corrected"}}
repairer := NewDefaultOutputRepairer(client)
original := []domain.RenderedMessage{
{Role: "system", Content: "Follow the task.", CacheControl: &domain.CacheControl{Type: domain.CacheControlEphemeral, TTL: "1h"}},
{Role: "user", Content: "Summarize the report."},
}
before := append([]domain.RenderedMessage(nil), original...)
previous := strings.Repeat("candidate ", 12_000)
target := domain.ExecutionTarget{BackendID: "backend", Endpoint: "https://provider.example/v1", Model: "model"}
presence := domain.ExecutionTargetPresence{Temperature: true, TopP: true}
structured := &domain.StructuredOutputSpec{}
response, err := repairer.Repair(context.Background(), RepairRequest{
OriginalMessages: original,
PreviousOutput: previous,
ValidationErrors: []string{"invalid JSON"},
SessionID: "session",
Target: target,
TargetPresence: presence,
StructuredOutput: structured,
Attempt: 1,
MaxAttempts: 3,
Mode: domain.ValidationJSONSchema,
})
if err != nil {
t.Fatalf("repair: %v", err)
}
if response == nil || response.Content != "corrected" {
t.Fatalf("response = %+v", response)
}
if !reflect.DeepEqual(original, before) {
t.Fatalf("original messages changed: got %#v, want %#v", original, before)
}
if len(client.requests) != 1 {
t.Fatalf("generation requests = %d, want 1", len(client.requests))
}
request := client.requests[0]
if request.Prompt.SessionID != "session" || !reflect.DeepEqual(request.Target, target) || request.TargetPresence != presence || request.StructuredOutput != structured {
t.Fatalf("generation request fields = %+v", request)
}
if len(request.Prompt.Messages) != len(original)+2 {
t.Fatalf("message count = %d, want %d", len(request.Prompt.Messages), len(original)+2)
}
if !reflect.DeepEqual(request.Prompt.Messages[:len(original)], original) {
t.Fatalf("original messages = %#v, want %#v", request.Prompt.Messages[:len(original)], original)
}
assistant := request.Prompt.Messages[len(original)]
if assistant.Role != "assistant" || assistant.Content != previous {
t.Fatalf("assistant candidate = %+v", assistant)
}
correction := request.Prompt.Messages[len(original)+1]
if correction.Role != "user" || !strings.Contains(correction.Content, "Repair attempt 1 of 3") ||
!strings.Contains(correction.Content, "Preserve valid values") ||
!strings.Contains(correction.Content, "Return only corrected JSON") {
t.Fatalf("correction message = %q", correction.Content)
}
if diagnostics := repairDiagnosticsFromMessage(t, correction.Content); !reflect.DeepEqual(diagnostics, []string{"invalid JSON"}) {
t.Fatalf("diagnostics = %#v", diagnostics)
}
}
func TestDefaultOutputRepairerDoesNotAccumulateCandidates(t *testing.T) {
client := &recordingRepairClient{response: &domain.GenerateResponse{Content: "corrected"}}
repairer := NewDefaultOutputRepairer(client)
backing := make([]domain.RenderedMessage, 1, 4)
backing[0] = domain.RenderedMessage{Role: "user", Content: "Original task"}
before := append([]domain.RenderedMessage(nil), backing...)
for _, candidate := range []string{"first invalid", "second invalid"} {
_, err := repairer.Repair(context.Background(), RepairRequest{
OriginalMessages: backing,
PreviousOutput: candidate,
Attempt: 1,
MaxAttempts: 3,
Mode: domain.ValidationJSON,
})
if err != nil {
t.Fatalf("repair %q: %v", candidate, err)
}
}
if !reflect.DeepEqual(backing, before) {
t.Fatalf("caller messages changed: got %#v, want %#v", backing, before)
}
if len(client.requests) != 2 {
t.Fatalf("generation requests = %d, want 2", len(client.requests))
}
for index, request := range client.requests {
messages := request.Prompt.Messages
if len(messages) != 3 || messages[0] != backing[0] || messages[1].Role != "assistant" || messages[1].Content != []string{"first invalid", "second invalid"}[index] {
t.Fatalf("request %d messages = %#v", index, messages)
}
}
}
func TestDefaultOutputRepairerOmitsEmptyCandidateMessage(t *testing.T) {
for _, tc := range []struct {
name string
candidate string
}{
{name: "empty", candidate: ""},
{name: "whitespace", candidate: " \n\t "},
} {
t.Run(tc.name, func(t *testing.T) {
client := &recordingRepairClient{response: &domain.GenerateResponse{Content: "corrected"}}
repairer := NewDefaultOutputRepairer(client)
_, err := repairer.Repair(context.Background(), RepairRequest{
OriginalMessages: []domain.RenderedMessage{{Role: "user", Content: "Original task"}},
PreviousOutput: tc.candidate,
Attempt: 1,
MaxAttempts: 1,
Mode: domain.ValidationBasic,
})
if err != nil {
t.Fatalf("repair: %v", err)
}
messages := client.requests[0].Prompt.Messages
if len(messages) != 2 || messages[1].Role != "user" || !strings.Contains(messages[1].Content, "previous response was empty") {
t.Fatalf("messages = %#v", messages)
}
if strings.Contains(messages[1].Content, tc.candidate) && tc.candidate != "" {
t.Fatalf("correction message repeated whitespace candidate: %q", messages[1].Content)
}
})
}
}
func TestDefaultOutputRepairerUsesModeSpecificGuidance(t *testing.T) {
tests := []struct {
mode domain.ValidationMode
want string
}{
{mode: domain.ValidationBasic, want: "Return a nonempty response"},
{mode: domain.ValidationJSON, want: "Return only corrected JSON"},
{mode: domain.ValidationJSONSchema, want: "Return only corrected JSON"},
}
for _, tc := range tests {
t.Run(string(tc.mode), func(t *testing.T) {
client := &recordingRepairClient{response: &domain.GenerateResponse{Content: "corrected"}}
_, err := NewDefaultOutputRepairer(client).Repair(context.Background(), RepairRequest{Mode: tc.mode, Attempt: 1, MaxAttempts: 1})
if err != nil {
t.Fatalf("repair: %v", err)
}
if !strings.Contains(client.requests[0].Prompt.Messages[0].Content, tc.want) {
t.Fatalf("correction message = %q", client.requests[0].Prompt.Messages[0].Content)
}
})
}
client := &recordingRepairClient{response: &domain.GenerateResponse{Content: "unexpected"}}
_, err := NewDefaultOutputRepairer(client).Repair(context.Background(), RepairRequest{Mode: domain.ValidationNone})
if err == nil || !strings.Contains(err.Error(), "unsupported validation mode") {
t.Fatalf("unsupported mode error = %v", err)
}
if len(client.requests) != 0 {
t.Fatalf("generation requests = %d, want 0", len(client.requests))
}
}
func TestFormatRepairDiagnosticsBoundsAndPreservesData(t *testing.T) {
tests := []struct {
name string
errors []string
wantOmission bool
check func(*testing.T, []string)
}{
{
name: "below limit",
errors: []string{"first", "second"},
check: func(t *testing.T, got []string) {
t.Helper()
if !reflect.DeepEqual(got, []string{"first", "second"}) {
t.Fatalf("diagnostics = %#v", got)
}
},
},
{
name: "at limit",
errors: []string{strings.Repeat("x", maxRepairDiagnosticBytes-4)},
check: func(t *testing.T, got []string) {
t.Helper()
if len(got) != 1 || len(got[0]) != maxRepairDiagnosticBytes-4 {
t.Fatalf("diagnostics lengths = %#v", got)
}
},
},
{
name: "multibyte truncation preserves prior entries",
errors: []string{"first", strings.Repeat("界", maxRepairDiagnosticBytes)},
wantOmission: true,
check: func(t *testing.T, got []string) {
t.Helper()
if len(got) != 3 || got[0] != "first" || !strings.HasPrefix(strings.Repeat("界", maxRepairDiagnosticBytes), got[1]) {
t.Fatalf("diagnostics = %#v", got)
}
},
},
{
name: "many diagnostics",
errors: manyRepairDiagnostics(),
wantOmission: true,
check: func(t *testing.T, got []string) {
t.Helper()
if len(got) < 2 || got[0] != manyRepairDiagnostics()[0] {
t.Fatalf("diagnostics = %#v", got)
}
},
},
{
name: "no room for partial diagnostic",
errors: func() []string {
omission, _ := json.Marshal(omittedRepairDiagnostics)
return []string{
strings.Repeat("x", maxRepairDiagnosticBytes-len(omission)-5),
strings.Repeat("y", 128),
}
}(),
wantOmission: true,
check: func(t *testing.T, got []string) {
t.Helper()
if len(got) != 2 || got[1] != omittedRepairDiagnostics {
t.Fatalf("diagnostics = %#v", got)
}
},
},
{
name: "invalid UTF-8",
errors: []string{"broken\xffinput"},
check: func(t *testing.T, got []string) {
t.Helper()
if !reflect.DeepEqual(got, []string{"broken\uFFFDinput"}) {
t.Fatalf("diagnostics = %#v", got)
}
},
},
{
name: "one huge diagnostic",
errors: []string{strings.Repeat("x", maxRepairDiagnosticBytes*2)},
wantOmission: true,
check: func(t *testing.T, got []string) {
t.Helper()
if len(got) != 2 || !strings.HasPrefix(strings.Repeat("x", maxRepairDiagnosticBytes*2), got[0]) {
t.Fatalf("diagnostics = %#v", got)
}
},
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
before := append([]string(nil), tc.errors...)
encoded := formatRepairDiagnostics(tc.errors)
if len(encoded) > maxRepairDiagnosticBytes || !utf8.ValidString(encoded) {
t.Fatalf("encoded diagnostic length/UTF-8 = (%d, %v)", len(encoded), utf8.ValidString(encoded))
}
var got []string
if err := json.Unmarshal([]byte(encoded), &got); err != nil {
t.Fatalf("decode diagnostics: %v; encoded=%q", err, encoded)
}
if !reflect.DeepEqual(tc.errors, before) {
t.Fatalf("input diagnostics changed: got %#v, want %#v", tc.errors, before)
}
if hasOmission := len(got) > 0 && got[len(got)-1] == omittedRepairDiagnostics; hasOmission != tc.wantOmission {
t.Fatalf("omission = %v, want %v; diagnostics=%#v", hasOmission, tc.wantOmission, got)
}
tc.check(t, got)
})
}
}
func manyRepairDiagnostics() []string {
diagnostics := make([]string, 1_000)
for index := range diagnostics {
diagnostics[index] = fmt.Sprintf("diagnostic %04d %s", index, strings.Repeat("x", 128))
}
return diagnostics
}
func TestDefaultOutputRepairerPropagatesGenerationFailures(t *testing.T) {
expected := errors.New("generation failed")
for _, tc := range []struct {
name string
client *recordingRepairClient
want error
}{
{name: "nil client", want: nil},
{name: "generation error", client: &recordingRepairClient{err: expected}, want: expected},
{name: "nil response", client: &recordingRepairClient{}, want: nil},
} {
t.Run(tc.name, func(t *testing.T) {
var repairer OutputRepairer
if tc.client == nil {
repairer = NewDefaultOutputRepairer(nil)
} else {
repairer = NewDefaultOutputRepairer(tc.client)
}
response, err := repairer.Repair(context.Background(), RepairRequest{Mode: domain.ValidationJSON})
if response != nil || err == nil {
t.Fatalf("response/error = (%+v, %v)", response, err)
}
if tc.want != nil && !errors.Is(err, tc.want) {
t.Fatalf("error = %v, want %v", err, tc.want)
}
})
}
}
func repairDiagnosticsFromMessage(t *testing.T, message string) []string {
t.Helper()
const marker = "Validation diagnostics (data):\n"
index := strings.Index(message, marker)
if index < 0 {
t.Fatalf("missing diagnostics marker in %q", message)
}
encoded := message[index+len(marker):]
var diagnostics []string
if err := json.Unmarshal([]byte(encoded), &diagnostics); err != nil {
t.Fatalf("decode diagnostics: %v", err)
}
return diagnostics
}

View File

@@ -17,6 +17,7 @@ import (
"gitea.maximumdirect.net/eric/promptkit/internal/capacity" "gitea.maximumdirect.net/eric/promptkit/internal/capacity"
"gitea.maximumdirect.net/eric/promptkit/internal/defaults" "gitea.maximumdirect.net/eric/promptkit/internal/defaults"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
"gitea.maximumdirect.net/eric/promptkit/internal/llm" "gitea.maximumdirect.net/eric/promptkit/internal/llm"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
"gitea.maximumdirect.net/eric/promptkit/internal/prompt" "gitea.maximumdirect.net/eric/promptkit/internal/prompt"
@@ -71,6 +72,11 @@ type preparationState struct {
start time.Time start time.Time
} }
type preparedOperation struct {
run *domain.PreparedRun
validation validate.PreparedValidation
}
func NewRunner( func NewRunner(
promptDefs promptdef.Repository, promptDefs promptdef.Repository,
profiles profile.Repository, profiles profile.Repository,
@@ -137,18 +143,20 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
} }
defer release() defer release()
prepared, err := r.completePreparation(ctx, req, state) operation, err := r.completePreparation(ctx, req, state)
if err != nil { if err != nil {
return nil, err return nil, err
} }
directAPIKey := state.effectiveModel.APIKey directAPIKey := state.effectiveModel.APIKey
return r.executePreparedRun(ctx, prepared, directAPIKey, runID, start, func( return r.executePreparedRun(ctx, operation.run, directAPIKey, runID, start, func(
ctx context.Context, ctx context.Context,
artifact *domain.Artifact, artifact *domain.Artifact,
attemptsUsed int, attemptsUsed int,
) (domain.ValidationResult, error) { ) (domain.ValidationResult, error) {
return r.validateOutput(ctx, artifact, prepared.OutputContract, attemptsUsed) result, err := operation.validation.Validate(ctx, artifact)
result.RepairAttempts = attemptsUsed
return result, err
}) })
} }
@@ -168,18 +176,20 @@ func (r *Runner) executePreparedRun(
) (*domain.RunResult, error) { ) (*domain.RunResult, error) {
executionTarget := prepared.EffectiveModelParams executionTarget := prepared.EffectiveModelParams
executionTarget.APIKey = directAPIKey executionTarget.APIKey = directAPIKey
genResp, err := r.llm.Generate(ctx, domain.GenerateRequest{ genResp, err := r.llm.Generate(ctx, newGenerationRequest(
Prompt: domain.RenderedPrompt{SessionID: prepared.SessionID, Messages: prepared.Messages}, domain.RenderedPrompt{Messages: prepared.Messages},
Target: executionTarget, prepared.SessionID,
TargetPresence: prepared.TargetPresence, executionTarget,
StructuredOutput: prepared.StructuredOutput, prepared.TargetPresence,
}) prepared.StructuredOutput,
))
if err != nil { if err != nil {
if errors.Is(err, llm.ErrInvalidRequest) { return nil, wrapGenerationError(err)
return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
} }
return nil, fmt.Errorf("%w: %w", ErrLLMGenerate, err) if genResp == nil {
return nil, fmt.Errorf("%w: model returned nil response", ErrLLMGenerate)
} }
usage := genResp.Usage
outputArtifact := buildOutputArtifact(genResp.Content, prepared.OutputContract.Format) outputArtifact := buildOutputArtifact(genResp.Content, prepared.OutputContract.Format)
validationResult, err := validateArtifact(ctx, &outputArtifact, 0) validationResult, err := validateArtifact(ctx, &outputArtifact, 0)
@@ -193,23 +203,26 @@ func (r *Runner) executePreparedRun(
attemptsUsed++ attemptsUsed++
repairResp, repairErr := r.repairer.Repair(ctx, RepairRequest{ repairResp, repairErr := r.repairer.Repair(ctx, RepairRequest{
OriginalMessages: prepared.Messages,
PreviousOutput: genResp.Content, PreviousOutput: genResp.Content,
ValidationErrors: validationResult.Errors, ValidationErrors: validationResult.Errors,
SessionID: prepared.SessionID, SessionID: prepared.SessionID,
Target: executionTarget, Target: executionTarget,
TargetPresence: prepared.TargetPresence,
StructuredOutput: prepared.StructuredOutput, StructuredOutput: prepared.StructuredOutput,
Attempt: attemptsUsed, Attempt: attemptsUsed,
MaxAttempts: prepared.OutputContract.RepairAttempts, MaxAttempts: prepared.OutputContract.RepairAttempts,
Mode: prepared.OutputContract.ValidationMode, Mode: prepared.OutputContract.ValidationMode,
}) })
if repairErr != nil { if repairErr != nil {
return nil, fmt.Errorf("%w: %w", ErrValidation, repairErr) return nil, wrapGenerationError(repairErr)
} }
if repairResp == nil { if repairResp == nil {
return nil, fmt.Errorf("%w: repairer returned nil response", ErrValidation) return nil, fmt.Errorf("%w: repairer returned nil response", ErrLLMGenerate)
} }
genResp = repairResp genResp = repairResp
usage = addTokenUsage(usage, repairResp.Usage)
outputArtifact = buildOutputArtifact(genResp.Content, prepared.OutputContract.Format) outputArtifact = buildOutputArtifact(genResp.Content, prepared.OutputContract.Format)
validationResult, err = validateArtifact(ctx, &outputArtifact, attemptsUsed) validationResult, err = validateArtifact(ctx, &outputArtifact, attemptsUsed)
@@ -238,19 +251,40 @@ func (r *Runner) executePreparedRun(
Endpoint: prepared.EffectiveModelParams.Endpoint, Endpoint: prepared.EffectiveModelParams.Endpoint,
EffectiveModelParams: executionTarget, EffectiveModelParams: executionTarget,
InputHashes: prepared.InputHashes, InputHashes: prepared.InputHashes,
Usage: genResp.Usage, Usage: usage,
StartTime: start, StartTime: start,
EndTime: end, EndTime: end,
Duration: end.Sub(start), Duration: end.Sub(start),
}, nil }, nil
} }
func wrapGenerationError(err error) error {
if errors.Is(err, llm.ErrInvalidRequest) {
return fmt.Errorf("%w: %w", ErrInvalidRequest, err)
}
return fmt.Errorf("%w: %w", ErrLLMGenerate, err)
}
func addTokenUsage(total, next domain.TokenUsage) domain.TokenUsage {
return domain.TokenUsage{
PromptTokens: total.PromptTokens + next.PromptTokens,
CompletionTokens: total.CompletionTokens + next.CompletionTokens,
TotalTokens: total.TotalTokens + next.TotalTokens,
CachedTokens: total.CachedTokens + next.CachedTokens,
CacheWriteTokens: total.CacheWriteTokens + next.CacheWriteTokens,
}
}
func (r *Runner) Prepare(ctx context.Context, req domain.RunRequest) (*domain.PreparedRun, error) { func (r *Runner) Prepare(ctx context.Context, req domain.RunRequest) (*domain.PreparedRun, error) {
state, err := r.resolvePreparation(ctx, req, time.Now().UTC()) state, err := r.resolvePreparation(ctx, req, time.Now().UTC())
if err != nil { if err != nil {
return nil, err return nil, err
} }
return r.completePreparation(ctx, req, state) operation, err := r.completePreparation(ctx, req, state)
if err != nil {
return nil, err
}
return operation.run, nil
} }
func (r *Runner) resolvePreparation( func (r *Runner) resolvePreparation(
@@ -272,6 +306,10 @@ func (r *Runner) resolvePreparation(
} }
def := promptSelection.definition def := promptSelection.definition
promptDefinitionHash := promptSelection.hash promptDefinitionHash := promptSelection.hash
effectiveContract, err := resolveOutputContract(def, req.Validation)
if err != nil {
return nil, fmt.Errorf("%w: output contract: %v", ErrInvalidRequest, err)
}
selectedProfileID := strings.TrimSpace(req.ProfileID) selectedProfileID := strings.TrimSpace(req.ProfileID)
if selectedProfileID == "" { if selectedProfileID == "" {
@@ -286,19 +324,16 @@ func (r *Runner) resolvePreparation(
return nil, err return nil, err
} }
effectiveModel, targetPresence, err := resolveExecutionTarget(selection.backend, selection.profile, req.Execution) effectiveModel, targetPresence := resolveExecutionTarget(selection.backend, selection.profile, req.Execution)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
}
effectiveModel.APIKey = req.APIKey effectiveModel.APIKey = req.APIKey
if err := validateResolvedExecutionTarget(effectiveModel); err != nil { effectiveModel, err = normalizeResolvedExecutionTarget(effectiveModel)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err) return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
} }
if err := validateAPIKey(effectiveModel.APIKeyEnv, effectiveModel.APIKey, effectiveModel.APIKeyRequired); err != nil { if err := validateAPIKey(effectiveModel.APIKeyEnv, effectiveModel.APIKey, effectiveModel.APIKeyRequired); err != nil {
return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err) return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
} }
effectiveContract := resolveOutputContract(def, req.Validation)
return &preparationState{ return &preparationState{
definition: def, definition: def,
directSessionID: directSessionID, directSessionID: directSessionID,
@@ -315,16 +350,68 @@ func (r *Runner) completePreparation(
ctx context.Context, ctx context.Context,
req domain.RunRequest, req domain.RunRequest,
state *preparationState, state *preparationState,
) (*domain.PreparedRun, error) { ) (*preparedOperation, error) {
structuredOutput, err := r.resolveStructuredOutput( validationPlan, err := r.prepareValidation(ctx, state.effectiveContract)
ctx, if err != nil {
return nil, err
}
structuredOutput, err := r.structuredOutputFromValidationPlan(
state.definition, state.definition,
state.effectiveContract, state.effectiveContract,
validationPlan,
) )
if err != nil { if err != nil {
return nil, err return nil, err
} }
return r.completePreparationWithStructuredOutput(ctx, req, state, structuredOutput) prepared, err := r.completePreparationWithStructuredOutput(ctx, req, state, structuredOutput)
if err != nil {
return nil, err
}
return &preparedOperation{run: prepared, validation: validationPlan}, nil
}
func (r *Runner) prepareValidation(
ctx context.Context,
contract domain.OutputContract,
) (validate.PreparedValidation, error) {
if r.validator == nil {
return noOpPreparedValidation{contract: contract}, nil
}
preparer, ok := r.validator.(validate.ValidationPreparer)
if !ok {
return nil, fmt.Errorf("%w: validator does not support prepared validation", ErrValidation)
}
plan, err := preparer.PrepareValidation(ctx, contract)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrValidation, err)
}
if plan == nil {
return nil, fmt.Errorf("%w: validator returned nil prepared validation", ErrValidation)
}
return plan, nil
}
func (r *Runner) structuredOutputFromValidationPlan(
def *domain.PromptDefinition,
contract domain.OutputContract,
plan validate.PreparedValidation,
) (*domain.StructuredOutputSpec, error) {
if contract.ValidationMode != domain.ValidationJSONSchema {
return nil, nil
}
schemaDocument := plan.SchemaDocument()
if schemaDocument == nil {
if r.validator == nil {
return nil, nil
}
return nil, fmt.Errorf("%w: prepared json_schema validation has no schema document", ErrValidation)
}
schemaDocument, err := jsonvalue.Copy(schemaDocument)
if err != nil {
return nil, fmt.Errorf("%w: invalid prepared json_schema schema document: %v", ErrValidation, err)
}
return structuredOutputSpec(def, schemaDocument), nil
} }
func (r *Runner) completePreparationWithStructuredOutput( func (r *Runner) completePreparationWithStructuredOutput(
@@ -398,24 +485,6 @@ func (r *Runner) admitRun(ctx context.Context, backendID string) (func(), error)
return release, nil return release, nil
} }
func (r *Runner) resolveStructuredOutput(ctx context.Context, def *domain.PromptDefinition, contract domain.OutputContract) (*domain.StructuredOutputSpec, error) {
if contract.ValidationMode != domain.ValidationJSONSchema {
return nil, nil
}
loader, ok := r.validator.(validate.SchemaDocumentLoader)
if !ok || loader == nil {
return nil, fmt.Errorf("%w: json_schema output requires schema document loader", ErrValidation)
}
schemaDoc, err := loader.LoadSchemaDocument(ctx, contract.SchemaPath)
if err != nil {
return nil, fmt.Errorf("%w: failed to load json schema for structured output: %v", ErrValidation, err)
}
return structuredOutputSpec(def, schemaDoc), nil
}
func structuredOutputSpec(def *domain.PromptDefinition, schemaDocument any) *domain.StructuredOutputSpec { func structuredOutputSpec(def *domain.PromptDefinition, schemaDocument any) *domain.StructuredOutputSpec {
return &domain.StructuredOutputSpec{ return &domain.StructuredOutputSpec{
Type: domain.StructuredOutputJSONSchema, Type: domain.StructuredOutputJSONSchema,
@@ -453,25 +522,6 @@ func deriveStructuredSchemaName(promptID string, promptVersion string) string {
return name return name
} }
func (r *Runner) validateOutput(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract, attemptsUsed int) (domain.ValidationResult, error) {
if r.validator == nil || contract.ValidationMode == domain.ValidationNone {
return domain.ValidationResult{
Status: domain.ValidationSkipped,
Mode: contract.ValidationMode,
SchemaPath: contract.SchemaPath,
RepairAttempts: attemptsUsed,
IsValid: true,
}, nil
}
res, err := r.validator.Validate(ctx, artifact, contract)
if err != nil {
return domain.ValidationResult{}, err
}
res.RepairAttempts = attemptsUsed
return res, nil
}
func (r *Runner) shouldAttemptRepair(contract domain.OutputContract, validationResult domain.ValidationResult) bool { func (r *Runner) shouldAttemptRepair(contract domain.OutputContract, validationResult domain.ValidationResult) bool {
if r.repairer == nil { if r.repairer == nil {
return false return false
@@ -482,7 +532,12 @@ func (r *Runner) shouldAttemptRepair(contract domain.OutputContract, validationR
if validationResult.Status != domain.ValidationFailed { if validationResult.Status != domain.ValidationFailed {
return false return false
} }
return contract.ValidationMode == domain.ValidationJSON || contract.ValidationMode == domain.ValidationJSONSchema switch contract.ValidationMode {
case domain.ValidationBasic, domain.ValidationJSON, domain.ValidationJSONSchema:
return true
default:
return false
}
} }
func mergeExecutionTarget(base domain.ExecutionTarget, override domain.ExecutionTarget) domain.ExecutionTarget { func mergeExecutionTarget(base domain.ExecutionTarget, override domain.ExecutionTarget) domain.ExecutionTarget {
@@ -527,7 +582,7 @@ func mergeExecutionTarget(base domain.ExecutionTarget, override domain.Execution
return out return out
} }
func mergeExecutionTargetOverride(base domain.ExecutionTarget, override domain.ExecutionTargetOverride) (domain.ExecutionTarget, domain.ExecutionTargetPresence, error) { func mergeExecutionTargetOverride(base domain.ExecutionTarget, override domain.ExecutionTargetOverride) (domain.ExecutionTarget, domain.ExecutionTargetPresence) {
out := base out := base
var presence domain.ExecutionTargetPresence var presence domain.ExecutionTargetPresence
if override.Endpoint != "" { if override.Endpoint != "" {
@@ -537,30 +592,18 @@ func mergeExecutionTargetOverride(base domain.ExecutionTarget, override domain.E
out.Model = override.Model out.Model = override.Model
} }
if override.Temperature != nil { if override.Temperature != nil {
if *override.Temperature < 0 || *override.Temperature > 2 {
return domain.ExecutionTarget{}, domain.ExecutionTargetPresence{}, errors.New("temperature must be between 0 and 2")
}
out.Temperature = *override.Temperature out.Temperature = *override.Temperature
presence.Temperature = true presence.Temperature = true
} }
if override.MaxTokens != nil { if override.MaxTokens != nil {
if *override.MaxTokens < 0 {
return domain.ExecutionTarget{}, domain.ExecutionTargetPresence{}, errors.New("max_tokens must be greater than or equal to 0")
}
out.MaxTokens = *override.MaxTokens out.MaxTokens = *override.MaxTokens
presence.MaxTokens = true presence.MaxTokens = true
} }
if override.TopP != nil { if override.TopP != nil {
if *override.TopP < 0 || *override.TopP > 1 {
return domain.ExecutionTarget{}, domain.ExecutionTargetPresence{}, errors.New("top_p must be between 0 and 1")
}
out.TopP = *override.TopP out.TopP = *override.TopP
presence.TopP = true presence.TopP = true
} }
if override.TimeoutSeconds != nil { if override.TimeoutSeconds != nil {
if *override.TimeoutSeconds < 0 {
return domain.ExecutionTarget{}, domain.ExecutionTargetPresence{}, errors.New("timeout_seconds must be greater than or equal to 0")
}
out.TimeoutSeconds = *override.TimeoutSeconds out.TimeoutSeconds = *override.TimeoutSeconds
presence.TimeoutSeconds = true presence.TimeoutSeconds = true
} }
@@ -576,35 +619,31 @@ func mergeExecutionTargetOverride(base domain.ExecutionTarget, override domain.E
if len(override.ExtraParams) > 0 { if len(override.ExtraParams) > 0 {
out.ExtraParams = copyExtraParams(override.ExtraParams) out.ExtraParams = copyExtraParams(override.ExtraParams)
} }
return out, presence, nil return out, presence
} }
func resolveExecutionTarget(backendValue *domain.Backend, profileValue *domain.ExecutionProfile, override *domain.ExecutionTargetOverride) (domain.ExecutionTarget, domain.ExecutionTargetPresence, error) { func resolveExecutionTarget(backendValue *domain.Backend, profileValue *domain.ExecutionProfile, override *domain.ExecutionTargetOverride) (domain.ExecutionTarget, domain.ExecutionTargetPresence) {
out := defaults.ExecutionTargetDefault() out := defaults.ExecutionTargetDefault()
out = mergeExecutionTarget(out, backendToTarget(backendValue)) out = mergeExecutionTarget(out, backendToTarget(backendValue))
out = mergeExecutionTarget(out, executionProfileToTarget(profileValue)) out = mergeExecutionTarget(out, executionProfileToTarget(profileValue))
var presence domain.ExecutionTargetPresence var presence domain.ExecutionTargetPresence
if override != nil { if override != nil {
var err error out, presence = mergeExecutionTargetOverride(out, *override)
out, presence, err = mergeExecutionTargetOverride(out, *override)
if err != nil {
return domain.ExecutionTarget{}, domain.ExecutionTargetPresence{}, err
} }
} return out, presence
return out, presence, nil
} }
func validateAPIKey(apiKeyEnv string, apiKey string, apiKeyRequired bool) error { func validateAPIKey(apiKeyEnv string, apiKey string, apiKeyRequired bool) error {
if strings.TrimSpace(apiKey) != "" { if strings.TrimSpace(apiKey) != "" {
return nil return nil
} }
if !apiKeyRequired {
return nil
}
envName := strings.TrimSpace(apiKeyEnv) envName := strings.TrimSpace(apiKeyEnv)
if envName == "" { if envName == "" {
if apiKeyRequired {
return ErrAPIKeyRequired return ErrAPIKeyRequired
} }
return nil
}
if strings.TrimSpace(os.Getenv(envName)) == "" { if strings.TrimSpace(os.Getenv(envName)) == "" {
return fmt.Errorf("%w: api key environment variable %q is not set", ErrAPIKeyEnvMissing, envName) return fmt.Errorf("%w: api key environment variable %q is not set", ErrAPIKeyEnvMissing, envName)
} }
@@ -658,18 +697,21 @@ func copyExtraParams(src map[string]any) map[string]any {
return cp return cp
} }
func resolveOutputContract(def *domain.PromptDefinition, override *domain.OutputContract) domain.OutputContract { func resolveOutputContract(def *domain.PromptDefinition, override *domain.OutputContract) (domain.OutputContract, error) {
contract := def.Validation contract := def.Validation
if contract.Format == "" { if contract.Format == "" {
contract.Format = def.OutputFormat contract.Format = def.OutputFormat
} }
if override != nil { if override != nil {
contract = *override contract = *override
}
if contract.Format == "" { if contract.Format == "" {
contract.Format = domain.FormatText contract.Format = domain.FormatText
} }
return contract }
if err := domain.ValidateOutputContract(contract); err != nil {
return domain.OutputContract{}, err
}
return contract, nil
} }
func hashRenderedPrompt(p domain.RenderedPrompt) string { func hashRenderedPrompt(p domain.RenderedPrompt) string {

View File

@@ -6,9 +6,11 @@ import (
"encoding/hex" "encoding/hex"
"errors" "errors"
"fmt" "fmt"
"math"
"path/filepath" "path/filepath"
"reflect" "reflect"
"regexp" "regexp"
"strconv"
"strings" "strings"
"sync" "sync"
"testing" "testing"
@@ -189,16 +191,34 @@ func (f *fakeValidator) Validate(ctx context.Context, artifact *domain.Artifact,
return f.result, nil return f.result, nil
} }
func (f *fakeValidator) LoadSchemaDocument(ctx context.Context, schemaPath string) (any, error) { func (f *fakeValidator) PrepareValidation(_ context.Context, contract domain.OutputContract) (validate.PreparedValidation, error) {
var schemaDocument any
if contract.ValidationMode == domain.ValidationJSONSchema {
f.schemaLoads++ f.schemaLoads++
f.schemaLoadPath = schemaPath f.schemaLoadPath = contract.SchemaPath
if f.schemaErr != nil { if f.schemaErr != nil {
return nil, f.schemaErr return nil, f.schemaErr
} }
if f.schemaDoc != nil { schemaDocument = f.schemaDoc
return f.schemaDoc, nil if schemaDocument == nil {
schemaDocument = map[string]any{"type": "object"}
} }
return map[string]any{"type": "object"}, nil }
return &fakePreparedValidator{validator: f, contract: contract, schemaDocument: schemaDocument}, nil
}
type fakePreparedValidator struct {
validator *fakeValidator
contract domain.OutputContract
schemaDocument any
}
func (p *fakePreparedValidator) Validate(ctx context.Context, artifact *domain.Artifact) (domain.ValidationResult, error) {
return p.validator.Validate(ctx, artifact, p.contract)
}
func (p *fakePreparedValidator) SchemaDocument() any {
return p.schemaDocument
} }
type fakeRepairer struct { type fakeRepairer struct {
@@ -208,6 +228,30 @@ type fakeRepairer struct {
reqs []RepairRequest reqs []RepairRequest
} }
type sequenceLLM struct {
responses []*domain.GenerateResponse
requests []domain.GenerateRequest
}
func (c *sequenceLLM) Generate(_ context.Context, req domain.GenerateRequest) (*domain.GenerateResponse, error) {
c.requests = append(c.requests, req)
index := len(c.requests) - 1
if index >= len(c.responses) {
return nil, errors.New("no generation response configured")
}
return c.responses[index], nil
}
type recordingRepairer struct {
next OutputRepairer
reqs []RepairRequest
}
func (r *recordingRepairer) Repair(ctx context.Context, req RepairRequest) (*domain.GenerateResponse, error) {
r.reqs = append(r.reqs, req)
return r.next.Repair(ctx, req)
}
type fakeRunAdmitter struct { type fakeRunAdmitter struct {
backendIDs []string backendIDs []string
err error err error
@@ -251,8 +295,10 @@ func (c *controlledRepairLLM) Generate(
ctx context.Context, ctx context.Context,
req domain.GenerateRequest, req domain.GenerateRequest,
) (*domain.GenerateResponse, error) { ) (*domain.GenerateResponse, error) {
isRepair := len(req.Prompt.Messages) > 0 && messages := req.Prompt.Messages
strings.HasPrefix(req.Prompt.Messages[0].Content, "You repair invalid JSON") isRepair := len(messages) >= 2 &&
messages[len(messages)-2].Role == "assistant" &&
messages[len(messages)-1].Role == "user"
c.mu.Lock() c.mu.Lock()
c.calls++ c.calls++
c.active++ c.active++
@@ -502,6 +548,35 @@ func TestRunnerDirectSessionResolution(t *testing.T) {
t.Fatalf("invalid direct session invoked generation %d times", llmClient.calls) t.Fatalf("invalid direct session invoked generation %d times", llmClient.calls)
} }
}) })
t.Run("malformed direct value fails before loading or generation", func(t *testing.T) {
promptRepo := &fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)}
llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: "unexpected"}}
runner := NewRunner(
promptRepo,
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
defaultArtifactReader(),
defaultRenderer(),
llmClient,
nil, nil)
_, err := runner.Run(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
SessionID: "session" + string([]byte{0xff}),
Inputs: singleInputRef(),
})
if !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("expected ErrInvalidRequest, got %v", err)
}
if promptRepo.lastID != "" {
t.Fatalf("invalid direct session loaded prompt %q", promptRepo.lastID)
}
if llmClient.calls != 0 {
t.Fatalf("invalid direct session invoked generation %d times", llmClient.calls)
}
})
} }
func TestRunnerPrepareUsesPromptDefaultProfileWhenNoExplicitProfileID(t *testing.T) { func TestRunnerPrepareUsesPromptDefaultProfileWhenNoExplicitProfileID(t *testing.T) {
@@ -704,16 +779,23 @@ func TestRunnerPrepareRequestNumericOverridePresence(t *testing.T) {
} }
func TestRunnerPrepareInvalidRequestNumericOverridesFail(t *testing.T) { func TestRunnerPrepareInvalidRequestNumericOverridesFail(t *testing.T) {
tests := []struct { type testCase struct {
name string name string
override *domain.ExecutionTargetOverride override *domain.ExecutionTargetOverride
}{ }
tests := []testCase{
{name: "temperature below range", override: &domain.ExecutionTargetOverride{Temperature: float64Ptr(-0.1)}}, {name: "temperature below range", override: &domain.ExecutionTargetOverride{Temperature: float64Ptr(-0.1)}},
{name: "temperature above range", override: &domain.ExecutionTargetOverride{Temperature: float64Ptr(2.1)}}, {name: "temperature above range", override: &domain.ExecutionTargetOverride{Temperature: float64Ptr(2.1)}},
{name: "max tokens below range", override: &domain.ExecutionTargetOverride{MaxTokens: intPtr(-1)}}, {name: "max tokens below range", override: &domain.ExecutionTargetOverride{MaxTokens: intPtr(-1)}},
{name: "top p below range", override: &domain.ExecutionTargetOverride{TopP: float64Ptr(-0.1)}}, {name: "top p below range", override: &domain.ExecutionTargetOverride{TopP: float64Ptr(-0.1)}},
{name: "top p above range", override: &domain.ExecutionTargetOverride{TopP: float64Ptr(1.1)}}, {name: "top p above range", override: &domain.ExecutionTargetOverride{TopP: float64Ptr(1.1)}},
{name: "timeout below range", override: &domain.ExecutionTargetOverride{TimeoutSeconds: intPtr(-1)}}, {name: "timeout below range", override: &domain.ExecutionTargetOverride{TimeoutSeconds: intPtr(-1)}},
{name: "temperature is not finite", override: &domain.ExecutionTargetOverride{Temperature: float64Ptr(math.NaN())}},
{name: "top p is not finite", override: &domain.ExecutionTargetOverride{TopP: float64Ptr(math.Inf(1))}},
}
if strconv.IntSize == 64 {
durationLimit := int64(math.MaxInt64 / int64(time.Second))
tests = append(tests, testCase{name: "timeout cannot be represented as a duration", override: &domain.ExecutionTargetOverride{TimeoutSeconds: intPtr(int(durationLimit) + 1)}})
} }
for _, tc := range tests { for _, tc := range tests {
@@ -1690,36 +1772,6 @@ func TestRunnerRunSelectedProfileBeatsBuiltInDefault(t *testing.T) {
} }
} }
func TestRunnerRunBuiltInDefaultsUsedWhenProfileOmitsOptionalFields(t *testing.T) {
promptRepo := &fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)}
execRepo := &fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{
"exec": {ID: "exec", Endpoint: "http://profile/v1", Model: "profile-model"},
}}
llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: "ok"}}
runner := NewRunner(promptRepo, execRepo, nil, defaultArtifactReader(), defaultRenderer(), llmClient, nil, nil)
res, err := runner.Run(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
if res.EffectiveModelParams.Temperature != defaults.ExecutionDefaultTemperature {
t.Fatalf("expected default temperature %v, got %v", defaults.ExecutionDefaultTemperature, res.EffectiveModelParams.Temperature)
}
if res.EffectiveModelParams.TopP != defaults.ExecutionDefaultTopP {
t.Fatalf("expected default top_p %v, got %v", defaults.ExecutionDefaultTopP, res.EffectiveModelParams.TopP)
}
if res.EffectiveModelParams.MaxTokens != defaults.ExecutionDefaultMaxTokens {
t.Fatalf("expected default max_tokens %d, got %d", defaults.ExecutionDefaultMaxTokens, res.EffectiveModelParams.MaxTokens)
}
if res.EffectiveModelParams.TimeoutSeconds != defaults.ExecutionDefaultTimeoutSeconds {
t.Fatalf("expected default timeout_seconds %d, got %d", defaults.ExecutionDefaultTimeoutSeconds, res.EffectiveModelParams.TimeoutSeconds)
}
}
func TestRunnerRunAPIKeyEnvResolvesFromEnvironment(t *testing.T) { func TestRunnerRunAPIKeyEnvResolvesFromEnvironment(t *testing.T) {
t.Setenv("PROMPTKIT_TEST_API_KEY", "secret") t.Setenv("PROMPTKIT_TEST_API_KEY", "secret")
promptRepo := &fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)} promptRepo := &fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)}
@@ -1737,22 +1789,26 @@ func TestRunnerRunAPIKeyEnvResolvesFromEnvironment(t *testing.T) {
} }
} }
func TestRunnerRunAPIKeyEnvMissingEnvironmentValueFailsClearly(t *testing.T) { func TestRunnerRunOptionalAPIKeyEnvMissingEnvironmentValueReachesLLM(t *testing.T) {
const environmentName = "PROMPTKIT_MISSING_KEY"
t.Setenv(environmentName, "")
promptRepo := &fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)} promptRepo := &fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)}
execRepo := &fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{ execRepo := &fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{
"exec": {ID: "exec", Endpoint: "http://profile/v1", Model: "profile-model", APIKeyEnv: "PROMPTKIT_MISSING_KEY"}, "exec": {ID: "exec", Endpoint: "http://profile/v1", Model: "profile-model", APIKeyEnv: environmentName},
}} }}
runner := NewRunner(promptRepo, execRepo, nil, defaultArtifactReader(), defaultRenderer(), &fakeLLM{resp: &domain.GenerateResponse{Content: "ok"}}, nil, nil) llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: "ok"}}
runner := NewRunner(promptRepo, execRepo, nil, defaultArtifactReader(), defaultRenderer(), llmClient, nil, nil)
_, err := runner.Run(context.Background(), domain.RunRequest{PromptID: "p", ProfileID: "exec", Inputs: singleInputRef()}) result, err := runner.Run(context.Background(), domain.RunRequest{PromptID: "p", ProfileID: "exec", Inputs: singleInputRef()})
if !errors.Is(err, ErrInvalidRequest) { if err != nil || result == nil {
t.Fatalf("expected ErrInvalidRequest, got %v", err) t.Fatalf("optional credential run = (%+v, %v), want success", result, err)
} }
if !errors.Is(err, ErrAPIKeyEnvMissing) { if llmClient.calls != 1 {
t.Fatalf("expected ErrAPIKeyEnvMissing, got %v", err) t.Fatalf("LLM calls = %d, want 1", llmClient.calls)
} }
if !strings.Contains(err.Error(), "PROMPTKIT_MISSING_KEY") { if llmClient.lastReq.Target.APIKeyEnv != environmentName {
t.Fatalf("expected missing env name in error, got %v", err) t.Fatalf("LLM api_key_env = %q, want %q", llmClient.lastReq.Target.APIKeyEnv, environmentName)
} }
} }
@@ -2044,50 +2100,272 @@ func TestRunnerRunValidationStillWorks(t *testing.T) {
} }
} }
func TestRunnerRunStructuredRepairRemainsBoundedAndUsesEffectiveModelSettings(t *testing.T) { func TestRunnerRepairStateMachine(t *testing.T) {
repairer := &fakeRepairer{responses: []*domain.GenerateResponse{{Content: `{"broken":`}, {Content: `{"still":`}}} failed := func(mode domain.ValidationMode, diagnostic string) domain.ValidationResult {
llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: `{"initial":`}} return domain.ValidationResult{
Status: domain.ValidationFailed,
Mode: mode,
Errors: []string{diagnostic},
IsValid: false,
}
}
passed := func(mode domain.ValidationMode) domain.ValidationResult {
return domain.ValidationResult{
Status: domain.ValidationPassed,
Mode: mode,
IsValid: true,
}
}
responses := func(count int) []*domain.GenerateResponse {
values := make([]*domain.GenerateResponse, count)
for index := range values {
unit := index + 1
values[index] = &domain.GenerateResponse{
Content: fmt.Sprintf(`{"candidate":%d}`, index),
Usage: domain.TokenUsage{
PromptTokens: unit,
CompletionTokens: unit * 10,
TotalTokens: unit * 100,
CachedTokens: unit * 1000,
CacheWriteTokens: unit * 10000,
},
}
}
return values
}
zeroOverrides := &domain.ExecutionTargetOverride{
Temperature: float64Ptr(0),
MaxTokens: intPtr(0),
TopP: float64Ptr(0),
TimeoutSeconds: intPtr(0),
}
allPresent := domain.ExecutionTargetPresence{
Temperature: true,
MaxTokens: true,
TopP: true,
TimeoutSeconds: true,
}
emptyThenValid := responses(2)
emptyThenValid[0].Content = " \t "
tests := []struct {
name string
mode domain.ValidationMode
budget int
validationResults []domain.ValidationResult
responses []*domain.GenerateResponse
execution *domain.ExecutionTargetOverride
wantPresence domain.ExecutionTargetPresence
wantRepairs int
wantStatus domain.ValidationStatus
structured bool
}{
{
name: "initial success does not repair",
mode: domain.ValidationJSON,
budget: 3,
validationResults: []domain.ValidationResult{passed(domain.ValidationJSON)},
responses: responses(1),
wantStatus: domain.ValidationPassed,
},
{
name: "empty basic output repairs successfully",
mode: domain.ValidationBasic,
budget: 3,
validationResults: []domain.ValidationResult{failed(domain.ValidationBasic, "empty output"), passed(domain.ValidationBasic)},
responses: emptyThenValid,
wantRepairs: 1,
wantStatus: domain.ValidationPassed,
},
{
name: "inherited numeric values remain absent",
mode: domain.ValidationJSON,
budget: 1,
validationResults: []domain.ValidationResult{
failed(domain.ValidationJSON, "initial syntax"),
passed(domain.ValidationJSON),
},
responses: responses(2),
wantRepairs: 1,
wantStatus: domain.ValidationPassed,
},
{
name: "explicit numeric zeros remain present",
mode: domain.ValidationJSONSchema,
budget: 1,
execution: zeroOverrides,
structured: true,
validationResults: []domain.ValidationResult{
failed(domain.ValidationJSONSchema, "initial schema mismatch"),
passed(domain.ValidationJSONSchema),
},
responses: responses(2),
wantPresence: allPresent,
wantRepairs: 1,
wantStatus: domain.ValidationPassed,
},
{
name: "successful repair stops below larger budget",
mode: domain.ValidationJSON,
budget: 3,
validationResults: []domain.ValidationResult{
failed(domain.ValidationJSON, "candidate zero"),
failed(domain.ValidationJSON, "candidate one"),
passed(domain.ValidationJSON),
},
responses: responses(3),
wantRepairs: 2,
wantStatus: domain.ValidationPassed,
},
{
name: "failed repairs exhaust exact larger budget",
mode: domain.ValidationJSON,
budget: 3,
validationResults: []domain.ValidationResult{
failed(domain.ValidationJSON, "candidate zero"),
failed(domain.ValidationJSON, "candidate one"),
failed(domain.ValidationJSON, "candidate two"),
failed(domain.ValidationJSON, "candidate three"),
},
responses: responses(4),
wantRepairs: 3,
wantStatus: domain.ValidationFailed,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
definition := promptDef(domain.FormatJSON, tc.mode, tc.budget)
plan := &recordingPreparedValidation{results: tc.validationResults}
if tc.structured {
definition.Validation.SchemaPath = "schema.json"
plan.schemaDocument = map[string]any{"type": "object"}
}
validator := &recordingValidationPreparer{plan: plan}
client := &sequenceLLM{responses: tc.responses}
repairer := &recordingRepairer{next: NewDefaultOutputRepairer(client)}
runner := NewRunnerWithRepairer( runner := NewRunnerWithRepairer(
&fakePromptRepo{def: promptDef(domain.FormatJSON, domain.ValidationJSON, 1)}, &fakePromptRepo{def: definition},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{ &fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{
"exec": {ID: "exec", BackendID: "custom", Model: "profile-model", TimeoutSeconds: 55}, "exec": {ID: "exec", BackendID: "custom", Model: "profile-model"},
}}, fakeBackendResolver{backends: map[string]domain.Backend{ }},
"custom": {ID: "custom", Endpoint: "http://backend/v1"}, fakeBackendResolver{backends: map[string]domain.Backend{
"custom": {ID: "custom", Endpoint: "http://backend.example/v1"},
}}, }},
defaultArtifactReader(), defaultArtifactReader(),
defaultRenderer(), defaultRenderer(),
llmClient, client,
validate.NewStandardValidator("."), validator,
repairer, nil) repairer,
nil,
)
res, err := runner.Run(context.Background(), domain.RunRequest{ result, err := runner.Run(context.Background(), domain.RunRequest{
PromptID: "p", PromptID: "p",
ProfileID: "exec", ProfileID: "exec",
SessionID: " repair-session ",
APIKey: "direct-secret",
Inputs: singleInputRef(), Inputs: singleInputRef(),
Execution: &domain.ExecutionTargetOverride{Endpoint: "http://override/v1", Model: "override-model", TimeoutSeconds: intPtr(22)}, Execution: tc.execution,
}) })
if err != nil { if err != nil {
t.Fatalf("expected no error, got %v", err) t.Fatalf("run: %v", err)
} }
if repairer.calls != 1 || res.Validation.RepairAttempts != 1 { if len(client.requests) != tc.wantRepairs+1 || len(repairer.reqs) != tc.wantRepairs {
t.Fatalf("expected one bounded repair, calls=%d attempts=%d", repairer.calls, res.Validation.RepairAttempts) t.Fatalf(
"generation/repair calls = (%d, %d), want (%d, %d)",
len(client.requests), len(repairer.reqs), tc.wantRepairs+1, tc.wantRepairs,
)
} }
if len(repairer.reqs) != 1 { if len(plan.artifacts) != tc.wantRepairs+1 {
t.Fatalf("expected one repair request, got %d", len(repairer.reqs)) t.Fatalf("validation calls = %d, want %d", len(plan.artifacts), tc.wantRepairs+1)
} }
if repairer.reqs[0].Target.Endpoint != "http://override/v1" || repairer.reqs[0].Target.Model != "override-model" {
t.Fatalf("expected repair to use effective target, got %+v", repairer.reqs[0].Target) initialRequest := client.requests[0]
if initialRequest.TargetPresence != tc.wantPresence {
t.Fatalf("initial target presence = %+v, want %+v", initialRequest.TargetPresence, tc.wantPresence)
} }
if repairer.reqs[0].Target.TimeoutSeconds != 22 { if initialRequest.Target.APIKey != "direct-secret" || initialRequest.Target.BackendID != "custom" ||
t.Fatalf("expected repair to use effective timeout, got %d", repairer.reqs[0].Target.TimeoutSeconds) initialRequest.Prompt.SessionID != "repair-session" {
t.Fatalf("initial common request fields = %+v", initialRequest)
} }
if llmClient.lastReq.Target.BackendID != "custom" || if initialRequest.Target.Endpoint != "http://backend.example/v1" ||
repairer.reqs[0].Target.BackendID != "custom" || initialRequest.Target.Model != "profile-model" ||
res.SelectedBackendID != "custom" { initialRequest.Target.Temperature != 0 || initialRequest.Target.MaxTokens != 0 || initialRequest.Target.TopP != 0 {
t.Fatalf("expected backend identity in generation, repair, and result: generate=%q repair=%q result=%q", t.Fatalf("initial effective target = %+v", initialRequest.Target)
llmClient.lastReq.Target.BackendID, repairer.reqs[0].Target.BackendID, res.SelectedBackendID) }
if tc.structured != (initialRequest.StructuredOutput != nil) {
t.Fatalf("initial structured output = %+v, want present %v", initialRequest.StructuredOutput, tc.structured)
}
if tc.structured && (initialRequest.StructuredOutput.JSONSchema == nil ||
initialRequest.StructuredOutput.JSONSchema.Name != "p_1") {
t.Fatalf("initial JSON Schema metadata = %+v, want derived schema name p_1", initialRequest.StructuredOutput)
}
for index, req := range repairer.reqs {
if req.Attempt != index+1 || req.MaxAttempts != tc.budget || req.Mode != tc.mode {
t.Fatalf("repair request %d progression = %+v", index, req)
}
if req.PreviousOutput != tc.responses[index].Content ||
!reflect.DeepEqual(req.ValidationErrors, tc.validationResults[index].Errors) {
t.Fatalf("repair request %d prior state = %+v", index, req)
}
if !reflect.DeepEqual(req.OriginalMessages, initialRequest.Prompt.Messages) {
t.Fatalf("repair request %d original messages drifted: %#v", index, req.OriginalMessages)
}
if req.TargetPresence != tc.wantPresence || !reflect.DeepEqual(req.Target, initialRequest.Target) ||
req.SessionID != initialRequest.Prompt.SessionID ||
!reflect.DeepEqual(req.StructuredOutput, initialRequest.StructuredOutput) {
t.Fatalf("repair request %d common fields drifted: %+v", index, req)
}
generated := client.requests[index+1]
if generated.TargetPresence != initialRequest.TargetPresence ||
!reflect.DeepEqual(generated.Target, initialRequest.Target) ||
generated.Prompt.SessionID != initialRequest.Prompt.SessionID ||
!reflect.DeepEqual(generated.StructuredOutput, initialRequest.StructuredOutput) {
t.Fatalf("repair generation request %d common fields drifted: %+v", index, generated)
}
expectedMessages := len(initialRequest.Prompt.Messages) + 1
if strings.TrimSpace(tc.responses[index].Content) != "" {
expectedMessages++
}
if len(generated.Prompt.Messages) != expectedMessages ||
generated.Prompt.Messages[len(generated.Prompt.Messages)-1].Role != "user" {
t.Fatalf("repair generation request %d messages = %#v", index, generated.Prompt.Messages)
}
if strings.TrimSpace(tc.responses[index].Content) != "" {
assistant := generated.Prompt.Messages[len(generated.Prompt.Messages)-2]
if assistant.Role != "assistant" || assistant.Content != tc.responses[index].Content {
t.Fatalf("repair generation request %d candidate = %+v", index, assistant)
}
}
}
lastResponse := tc.responses[tc.wantRepairs]
if result.RawOutput != lastResponse.Content || string(result.Artifact.Body) != lastResponse.Content {
t.Fatalf("final output = (%q, %q), want %q", result.RawOutput, result.Artifact.Body, lastResponse.Content)
}
if result.Validation.Status != tc.wantStatus || result.Validation.RepairAttempts != tc.wantRepairs {
t.Fatalf("final validation = %+v, want status %q and %d repairs", result.Validation, tc.wantStatus, tc.wantRepairs)
}
if result.SelectedBackendID != "custom" || result.SessionID != "repair-session" ||
result.EffectiveModelParams.APIKey != "" {
t.Fatalf("result execution metadata = %+v", result)
}
var wantUsage domain.TokenUsage
for _, response := range tc.responses[:tc.wantRepairs+1] {
wantUsage.PromptTokens += response.Usage.PromptTokens
wantUsage.CompletionTokens += response.Usage.CompletionTokens
wantUsage.TotalTokens += response.Usage.TotalTokens
wantUsage.CachedTokens += response.Usage.CachedTokens
wantUsage.CacheWriteTokens += response.Usage.CacheWriteTokens
}
if result.Usage != wantUsage {
t.Fatalf("cumulative usage = %+v, want %+v", result.Usage, wantUsage)
}
})
} }
} }
@@ -2182,96 +2460,6 @@ func TestRunnerSchedulesInitialAndRepairGenerationThroughOneBackendPool(t *testi
} }
} }
func TestRunnerRunRepairCarriesEffectiveSessionID(t *testing.T) {
llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: `{"broken":`}}
runner := NewRunnerWithRepairer(
&fakePromptRepo{def: promptDef(domain.FormatJSON, domain.ValidationJSON, 1)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{
"exec": {ID: "exec", Endpoint: "http://example.test/v1", Model: "model"},
}},
nil,
defaultArtifactReader(),
defaultRenderer(),
llmClient,
validate.NewStandardValidator("."),
NewDefaultOutputRepairer(llmClient), nil)
result, err := runner.Run(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
SessionID: " repair-session ",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
if llmClient.calls != 2 {
t.Fatalf("expected initial generation and one repair, got %d calls", llmClient.calls)
}
if llmClient.lastReq.Prompt.SessionID != "repair-session" {
t.Fatalf("expected repair generation to retain effective session, got %q", llmClient.lastReq.Prompt.SessionID)
}
if result.SessionID != "repair-session" {
t.Fatalf("expected result to retain effective session, got %q", result.SessionID)
}
}
func TestRunnerRunJSONSchemaRepairCarriesStructuredOutputSpec(t *testing.T) {
def := promptDef(domain.FormatJSON, domain.ValidationJSONSchema, 1)
def.Validation.SchemaPath = "events.schema.json"
validator := &fakeValidator{
result: domain.ValidationResult{
Status: domain.ValidationFailed,
Mode: domain.ValidationJSONSchema,
Errors: []string{"schema mismatch"},
IsValid: false,
},
schemaDoc: map[string]any{
"type": "object",
"properties": map[string]any{
"events": map[string]any{"type": "array"},
},
},
}
repairer := &fakeRepairer{
responses: []*domain.GenerateResponse{
{Content: `{"events":[]}`},
},
}
llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: `{"events":[1]}`}}
runner := NewRunnerWithRepairer(
&fakePromptRepo{def: def},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}}, nil,
defaultArtifactReader(),
defaultRenderer(),
llmClient,
validator,
repairer, nil)
_, err := runner.Run(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
if llmClient.lastReq.StructuredOutput == nil || llmClient.lastReq.StructuredOutput.JSONSchema == nil {
t.Fatalf("expected initial llm request to include structured output, got %+v", llmClient.lastReq.StructuredOutput)
}
if len(repairer.reqs) != 1 {
t.Fatalf("expected one repair request, got %d", len(repairer.reqs))
}
if repairer.reqs[0].StructuredOutput == nil || repairer.reqs[0].StructuredOutput.JSONSchema == nil {
t.Fatalf("expected repair request structured output, got %+v", repairer.reqs[0].StructuredOutput)
}
if repairer.reqs[0].StructuredOutput.JSONSchema.Name != "p_1" {
t.Fatalf("expected derived schema name p_1, got %q", repairer.reqs[0].StructuredOutput.JSONSchema.Name)
}
}
func TestExecutionProfileToTargetPopulatesAllFieldsAndCopiesExtraParams(t *testing.T) { func TestExecutionProfileToTargetPopulatesAllFieldsAndCopiesExtraParams(t *testing.T) {
src := &domain.ExecutionProfile{ src := &domain.ExecutionProfile{
ID: "exec", ID: "exec",
@@ -2330,10 +2518,7 @@ func TestResolveExecutionTargetProfileValuesPopulateAllSupportedFields(t *testin
}, },
} }
target, presence, err := resolveExecutionTarget(nil, profileValue, nil) target, presence := resolveExecutionTarget(nil, profileValue, nil)
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
if presence != (domain.ExecutionTargetPresence{}) { if presence != (domain.ExecutionTargetPresence{}) {
t.Fatalf("expected no request override presence, got %+v", presence) t.Fatalf("expected no request override presence, got %+v", presence)
} }
@@ -2385,10 +2570,7 @@ func TestResolveExecutionTargetRuntimeOverridesBeatProfileForAllOverrideableFiel
}, },
} }
target, presence, err := resolveExecutionTarget(nil, profileValue, override) target, presence := resolveExecutionTarget(nil, profileValue, override)
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
if presence != (domain.ExecutionTargetPresence{Temperature: true, MaxTokens: true, TopP: true, TimeoutSeconds: true}) { if presence != (domain.ExecutionTargetPresence{Temperature: true, MaxTokens: true, TopP: true, TimeoutSeconds: true}) {
t.Fatalf("unexpected override presence: %+v", presence) t.Fatalf("unexpected override presence: %+v", presence)
} }
@@ -2435,12 +2617,9 @@ func TestResolveExecutionTargetReasoningOverrideStates(t *testing.T) {
for _, tt := range tests { for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) { t.Run(tt.name, func(t *testing.T) {
target, _, err := resolveExecutionTarget(nil, profileValue, &domain.ExecutionTargetOverride{ target, _ := resolveExecutionTarget(nil, profileValue, &domain.ExecutionTargetOverride{
ReasoningEffort: tt.override, ReasoningEffort: tt.override,
}) })
if err != nil {
t.Fatalf("resolve execution target: %v", err)
}
if target.ReasoningEffort != tt.want { if target.ReasoningEffort != tt.want {
t.Fatalf("reasoning effort = %q, want %q", target.ReasoningEffort, tt.want) t.Fatalf("reasoning effort = %q, want %q", target.ReasoningEffort, tt.want)
} }
@@ -2569,10 +2748,7 @@ func TestResolveExecutionTargetUsesBackendProfileAndRequestPrecedence(t *testing
ExtraParams: map[string]any{"request": true}, ExtraParams: map[string]any{"request": true},
} }
target, _, err := resolveExecutionTarget(backendValue, profileValue, override) target, _ := resolveExecutionTarget(backendValue, profileValue, override)
if err != nil {
t.Fatalf("resolve target: %v", err)
}
if target.BackendID != "custom" { if target.BackendID != "custom" {
t.Fatalf("endpoint override changed backend identity: %+v", target) t.Fatalf("endpoint override changed backend identity: %+v", target)
} }
@@ -2583,12 +2759,9 @@ func TestResolveExecutionTargetUsesBackendProfileAndRequestPrecedence(t *testing
t.Fatalf("expected whole-map request replacement, got %#v", target.ExtraParams) t.Fatalf("expected whole-map request replacement, got %#v", target.ExtraParams)
} }
target, _, err = resolveExecutionTarget(backendValue, &domain.ExecutionProfile{ target, _ = resolveExecutionTarget(backendValue, &domain.ExecutionProfile{
ID: "exec", BackendID: "custom", Model: "profile-model", ID: "exec", BackendID: "custom", Model: "profile-model",
}, nil) }, nil)
if err != nil {
t.Fatalf("resolve backend defaults: %v", err)
}
if target.Endpoint != backendValue.Endpoint || if target.Endpoint != backendValue.Endpoint ||
target.APIKeyEnv != backendValue.APIKeyEnv || target.APIKeyEnv != backendValue.APIKeyEnv ||
!reflect.DeepEqual(target.ExtraParams, backendValue.ExtraParams) { !reflect.DeepEqual(target.ExtraParams, backendValue.ExtraParams) {

View File

@@ -0,0 +1,311 @@
package validate
import (
"context"
"errors"
"io/fs"
"strings"
"sync"
"sync/atomic"
"testing"
"testing/fstest"
"time"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"github.com/santhosh-tekuri/jsonschema/v6"
)
func TestValidationCancellationBeforeWorkDoesNotOpenSchemaSource(t *testing.T) {
source := &countingSchemaFS{FS: fstest.MapFS{
"schema.json": {Data: []byte(`{"type":"object"}`)},
}}
validator := NewFSValidator(source, ".").(ValidationPreparer)
ctx, cancel := context.WithCancel(context.Background())
cancel()
plan, err := validator.PrepareValidation(ctx, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
if plan != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("PrepareValidation() = (%v, %v), want nil plan and context cancellation", plan, err)
}
if opens := source.opens.Load(); opens != 0 {
t.Fatalf("schema source opens = %d, want 0", opens)
}
}
func TestValidationCancellationBetweenReferencedSchemaReadChunks(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
source := &controlledSchemaFS{
FS: fstest.MapFS{
"root.json": {Data: []byte(`{"$ref":"child.json"}`)},
"child.json": {Data: []byte(`{"type":"string"}` + strings.Repeat(" ", schemaReadChunkSize*2))},
},
target: "child.json",
cancel: cancel,
}
validator := NewFSValidator(source, ".").(ValidationPreparer)
plan, err := validator.PrepareValidation(ctx, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "root.json",
})
if plan != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("PrepareValidation() = (%v, %v), want nil plan and context cancellation", plan, err)
}
if reads := source.reads.Load(); reads != 1 {
t.Fatalf("controlled child reads = %d, want 1", reads)
}
if closes := source.closes.Load(); closes != 1 {
t.Fatalf("controlled child closes = %d, want 1", closes)
}
}
func TestDecodeCancellationAfterSynchronousCallWins(t *testing.T) {
ctx := newCheckpointContext(2)
value, err := decodeJSONValue(ctx, []byte(`{"value":1}`))
if value != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("decodeJSONValue() = (%v, %v), want nil value and context cancellation", value, err)
}
}
func TestCompileCancellationAfterSynchronousCallWins(t *testing.T) {
for _, dependencyErr := range []error{nil, errors.New("compile failed")} {
ctx, cancel := context.WithCancel(context.Background())
compiled := &jsonschema.Schema{}
schema, err := compileJSONSchema(ctx, "promptkit-schema:/root.json", func(string) (*jsonschema.Schema, error) {
cancel()
return compiled, dependencyErr
})
if schema != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("compileJSONSchema() = (%v, %v), want authoritative context cancellation", schema, err)
}
if dependencyErr != nil && errors.Is(err, dependencyErr) {
t.Fatalf("compileJSONSchema() error = %v, dependency error should not win", err)
}
}
}
func TestExecutionCancellationAfterSynchronousCallWins(t *testing.T) {
for _, dependencyErr := range []error{nil, errors.New("schema mismatch")} {
ctx, cancel := context.WithCancel(context.Background())
executor := &controlledSchemaExecutor{cancel: cancel, err: dependencyErr}
validationErrors, err := executeJSONSchema(ctx, executor, map[string]any{"ok": true})
if validationErrors != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("executeJSONSchema() = (%v, %v), want no result and context cancellation", validationErrors, err)
}
if calls := executor.calls.Load(); calls != 1 {
t.Fatalf("schema execution calls = %d, want 1", calls)
}
}
}
func TestCancellationDoesNotDetachBlockedSchemaRead(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
started := make(chan struct{})
release := make(chan struct{})
source := &controlledSchemaFS{
FS: fstest.MapFS{
"schema.json": {Data: []byte(`{"type":"object"}`)},
},
target: "schema.json",
started: started,
release: release,
}
validator := NewFSValidator(source, ".").(ValidationPreparer)
type outcome struct {
plan PreparedValidation
err error
}
result := make(chan outcome, 1)
go func() {
plan, err := validator.PrepareValidation(ctx, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
result <- outcome{plan: plan, err: err}
}()
waitForSignal(t, started, "schema read to start")
cancel()
select {
case got := <-result:
t.Fatalf("blocked dependency returned before release: (%v, %v)", got.plan, got.err)
default:
}
close(release)
got := waitForOutcome(t, result)
if got.plan != nil || !errors.Is(got.err, context.Canceled) {
t.Fatalf("PrepareValidation() after release = (%v, %v), want nil plan and context cancellation", got.plan, got.err)
}
if reads := source.reads.Load(); reads != 1 {
t.Fatalf("blocked reads = %d, want 1 completed read", reads)
}
}
func TestCancellationDoesNotDetachBlockedSchemaExecution(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
started := make(chan struct{})
release := make(chan struct{})
executor := &controlledSchemaExecutor{started: started, release: release}
result := make(chan error, 1)
go func() {
_, err := executeJSONSchema(ctx, executor, map[string]any{"ok": true})
result <- err
}()
waitForSignal(t, started, "schema execution to start")
cancel()
select {
case err := <-result:
t.Fatalf("blocked dependency returned before release: %v", err)
default:
}
close(release)
select {
case err := <-result:
if !errors.Is(err, context.Canceled) {
t.Fatalf("executeJSONSchema() after release = %v, want context cancellation", err)
}
case <-time.After(2 * time.Second):
t.Fatal("timed out waiting for released schema execution")
}
if calls := executor.calls.Load(); calls != 1 {
t.Fatalf("schema execution calls = %d, want 1 completed call", calls)
}
}
type countingSchemaFS struct {
fs.FS
opens atomic.Int32
}
func (f *countingSchemaFS) Open(name string) (fs.File, error) {
f.opens.Add(1)
return f.FS.Open(name)
}
type controlledSchemaFS struct {
fs.FS
target string
cancel context.CancelFunc
started chan struct{}
release chan struct{}
once sync.Once
reads atomic.Int32
closes atomic.Int32
}
func (f *controlledSchemaFS) Open(name string) (fs.File, error) {
file, err := f.FS.Open(name)
if err != nil || name != f.target {
return file, err
}
return &controlledSchemaFile{File: file, owner: f}, nil
}
type controlledSchemaFile struct {
fs.File
owner *controlledSchemaFS
}
func (f *controlledSchemaFile) Read(buffer []byte) (int, error) {
f.owner.once.Do(func() {
if f.owner.started != nil {
close(f.owner.started)
}
if f.owner.release != nil {
<-f.owner.release
}
})
n, err := f.File.Read(buffer)
f.owner.reads.Add(1)
if f.owner.cancel != nil {
f.owner.cancel()
}
return n, err
}
func (f *controlledSchemaFile) Close() error {
f.owner.closes.Add(1)
return f.File.Close()
}
type controlledSchemaExecutor struct {
cancel context.CancelFunc
err error
started chan struct{}
release chan struct{}
calls atomic.Int32
}
func (e *controlledSchemaExecutor) Validate(any) error {
e.calls.Add(1)
if e.started != nil {
close(e.started)
}
if e.release != nil {
<-e.release
}
if e.cancel != nil {
e.cancel()
}
return e.err
}
type checkpointContext struct {
context.Context
mu sync.Mutex
remaining int
canceled bool
done chan struct{}
}
func newCheckpointContext(checksUntilCancel int) *checkpointContext {
return &checkpointContext{
Context: context.Background(),
remaining: checksUntilCancel,
done: make(chan struct{}),
}
}
func (c *checkpointContext) Done() <-chan struct{} {
return c.done
}
func (c *checkpointContext) Err() error {
c.mu.Lock()
defer c.mu.Unlock()
if c.canceled {
return context.Canceled
}
c.remaining--
if c.remaining == 0 {
c.canceled = true
close(c.done)
return context.Canceled
}
return nil
}
func waitForSignal(t *testing.T, signal <-chan struct{}, description string) {
t.Helper()
select {
case <-signal:
case <-time.After(2 * time.Second):
t.Fatalf("timed out waiting for %s", description)
}
}
func waitForOutcome[T any](t *testing.T, result <-chan T) T {
t.Helper()
select {
case value := <-result:
return value
case <-time.After(2 * time.Second):
t.Fatal("timed out waiting for released dependency")
var zero T
return zero
}
}

View File

@@ -1,15 +1,18 @@
package validate package validate
import ( import (
"bytes"
"context" "context"
"encoding/json" "encoding/json"
"errors" "errors"
"fmt" "fmt"
"io"
"io/fs" "io/fs"
"net/url" "net/url"
"os" "os"
"path" "path"
"path/filepath" "path/filepath"
"runtime"
"strings" "strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
@@ -19,6 +22,8 @@ import (
const jsonSchemaDraft2020 = "https://json-schema.org/draft/2020-12/schema" const jsonSchemaDraft2020 = "https://json-schema.org/draft/2020-12/schema"
const schemaReadChunkSize = 64 * 1024
// StandardValidator provides basic, JSON, and JSON Schema output validation. // StandardValidator provides basic, JSON, and JSON Schema output validation.
type StandardValidator struct { type StandardValidator struct {
schemaBaseDir string schemaBaseDir string
@@ -48,7 +53,7 @@ func (v *FSValidator) Validate(ctx context.Context, artifact *domain.Artifact, c
type preparedValidation struct { type preparedValidation struct {
contract domain.OutputContract contract domain.OutputContract
schemaDocument any schemaDocument any
schema *jsonschema.Schema schema schemaExecutor
} }
func (p *preparedValidation) Validate(ctx context.Context, artifact *domain.Artifact) (domain.ValidationResult, error) { func (p *preparedValidation) Validate(ctx context.Context, artifact *domain.Artifact) (domain.ValidationResult, error) {
@@ -59,14 +64,11 @@ func (p *preparedValidation) SchemaDocument() any {
return p.schemaDocument return p.schemaDocument
} }
func (p *preparedValidation) validateJSONSchema(instance any, _ string) ([]string, error) { func (p *preparedValidation) validateJSONSchema(ctx context.Context, instance any, _ string) ([]string, error) {
if p.schema == nil { if p.schema == nil {
return nil, errors.New("prepared JSON schema is unavailable") return nil, errors.New("prepared JSON schema is unavailable")
} }
if err := p.schema.Validate(instance); err != nil { return executeJSONSchema(ctx, p.schema, instance)
return []string{fmt.Sprintf("json schema validation failed: %v", err)}, nil
}
return nil, nil
} }
func (v *StandardValidator) PrepareValidation(ctx context.Context, contract domain.OutputContract) (PreparedValidation, error) { func (v *StandardValidator) PrepareValidation(ctx context.Context, contract domain.OutputContract) (PreparedValidation, error) {
@@ -79,11 +81,11 @@ func (v *StandardValidator) PrepareValidation(ctx context.Context, contract doma
return prepared, nil return prepared, nil
} }
resolvedSchemaPath, err := v.resolveSchemaPath(contract.SchemaPath) resolvedSchemaPath, err := v.resolveSchemaPath(ctx, contract.SchemaPath)
if err != nil { if err != nil {
return nil, err return nil, err
} }
schemaDocument, err := loadJSONSchemaFile(resolvedSchemaPath) schemaDocument, err := loadJSONSchemaFile(ctx, resolvedSchemaPath)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err) return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err)
} }
@@ -91,11 +93,24 @@ func (v *StandardValidator) PrepareValidation(ctx context.Context, contract doma
if err != nil { if err != nil {
return nil, err return nil, err
} }
compiler := newSchemaCompiler(standardSchemaLoader{root: schemaRoot}) if err := ctx.Err(); err != nil {
if err := compiler.AddResource(resolvedSchemaPath, schemaDocument); err != nil { return nil, err
}
compiler := newSchemaCompiler(standardSchemaLoader{ctx: ctx, root: schemaRoot})
resourceURL := fileSchemaResourceURL(resolvedSchemaPath)
if err := ctx.Err(); err != nil {
return nil, err
}
if err := compiler.AddResource(resourceURL.String(), schemaDocument); err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return nil, contextErr
}
return nil, fmt.Errorf("failed to register JSON schema %q: %w", resolvedSchemaPath, err) return nil, fmt.Errorf("failed to register JSON schema %q: %w", resolvedSchemaPath, err)
} }
schema, err := compiler.Compile(resolvedSchemaPath) if err := ctx.Err(); err != nil {
return nil, err
}
schema, err := compileJSONSchema(ctx, resourceURL.String(), compiler.Compile)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err) return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err)
} }
@@ -118,16 +133,25 @@ func (v *FSValidator) PrepareValidation(ctx context.Context, contract domain.Out
return prepared, nil return prepared, nil
} }
schemaName, schemaDocument, err := v.loadSchemaDocument(contract.SchemaPath) schemaName, schemaDocument, err := v.loadSchemaDocument(ctx, contract.SchemaPath)
if err != nil { if err != nil {
return nil, err return nil, err
} }
resourceURL := fsSchemaResourceURL(schemaName) resourceURL := fsSchemaResourceURL(schemaName)
compiler := newSchemaCompiler(fsSchemaLoader{fsys: v.fsys, root: filecatalog.CleanFSRoot(v.root)}) compiler := newSchemaCompiler(fsSchemaLoader{ctx: ctx, fsys: v.fsys, root: filecatalog.CleanFSRoot(v.root)})
if err := compiler.AddResource(resourceURL, schemaDocument); err != nil { if err := ctx.Err(); err != nil {
return nil, err
}
if err := compiler.AddResource(resourceURL.String(), schemaDocument); err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return nil, contextErr
}
return nil, fmt.Errorf("failed to register JSON schema %q: %w", schemaName, err) return nil, fmt.Errorf("failed to register JSON schema %q: %w", schemaName, err)
} }
schema, err := compiler.Compile(resourceURL) if err := ctx.Err(); err != nil {
return nil, err
}
schema, err := compileJSONSchema(ctx, resourceURL.String(), compiler.Compile)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", schemaName, err) return nil, fmt.Errorf("failed to compile JSON schema %q: %w", schemaName, err)
} }
@@ -140,7 +164,7 @@ func (v *FSValidator) PrepareValidation(ctx context.Context, contract domain.Out
return prepared, nil return prepared, nil
} }
type schemaValidatorFunc func(instance any, schemaPath string) ([]string, error) type schemaValidatorFunc func(ctx context.Context, instance any, schemaPath string) ([]string, error)
func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract, validateSchema schemaValidatorFunc) (domain.ValidationResult, error) { func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract, validateSchema schemaValidatorFunc) (domain.ValidationResult, error) {
select { select {
@@ -152,7 +176,6 @@ func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract d
res := domain.ValidationResult{ res := domain.ValidationResult{
Mode: contract.ValidationMode, Mode: contract.ValidationMode,
SchemaPath: contract.SchemaPath, SchemaPath: contract.SchemaPath,
RepairAttempts: contract.RepairAttempts,
} }
if artifact == nil { if artifact == nil {
@@ -165,7 +188,11 @@ func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract d
res.IsValid = true res.IsValid = true
return res, nil return res, nil
case domain.ValidationBasic: case domain.ValidationBasic:
if strings.TrimSpace(string(artifact.Body)) == "" { empty := strings.TrimSpace(string(artifact.Body)) == ""
if err := ctx.Err(); err != nil {
return domain.ValidationResult{}, err
}
if empty {
res.Status = domain.ValidationFailed res.Status = domain.ValidationFailed
res.IsValid = false res.IsValid = false
res.Errors = []string{"output is empty"} res.Errors = []string{"output is empty"}
@@ -175,26 +202,32 @@ func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract d
res.IsValid = true res.IsValid = true
return res, nil return res, nil
case domain.ValidationJSON: case domain.ValidationJSON:
_, jsonErr := parseJSON(artifact.Body) valid := json.Valid(artifact.Body)
if jsonErr != nil { if err := ctx.Err(); err != nil {
return domain.ValidationResult{}, err
}
if !valid {
res.Status = domain.ValidationFailed res.Status = domain.ValidationFailed
res.IsValid = false res.IsValid = false
res.Errors = []string{fmt.Sprintf("invalid JSON: %v", jsonErr)} res.Errors = []string{"invalid JSON"}
return res, nil return res, nil
} }
res.Status = domain.ValidationPassed res.Status = domain.ValidationPassed
res.IsValid = true res.IsValid = true
return res, nil return res, nil
case domain.ValidationJSONSchema: case domain.ValidationJSONSchema:
instance, jsonErr := parseJSON(artifact.Body) instance, jsonErr := decodeJSONValue(ctx, artifact.Body)
if jsonErr != nil { if jsonErr != nil {
if contextErr := ctx.Err(); contextErr != nil {
return domain.ValidationResult{}, contextErr
}
res.Status = domain.ValidationFailed res.Status = domain.ValidationFailed
res.IsValid = false res.IsValid = false
res.Errors = []string{fmt.Sprintf("invalid JSON: %v", jsonErr)} res.Errors = []string{fmt.Sprintf("invalid JSON: %v", jsonErr)}
return res, nil return res, nil
} }
validationErrors, err := validateSchema(instance, contract.SchemaPath) validationErrors, err := validateSchema(ctx, instance, contract.SchemaPath)
if err != nil { if err != nil {
return domain.ValidationResult{}, err return domain.ValidationResult{}, err
} }
@@ -213,8 +246,8 @@ func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract d
} }
} }
func (v *StandardValidator) validateJSONSchema(instance any, schemaPath string) ([]string, error) { func (v *StandardValidator) validateJSONSchema(ctx context.Context, instance any, schemaPath string) ([]string, error) {
resolvedSchemaPath, err := v.resolveSchemaPath(schemaPath) resolvedSchemaPath, err := v.resolveSchemaPath(ctx, schemaPath)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -223,20 +256,21 @@ func (v *StandardValidator) validateJSONSchema(instance any, schemaPath string)
if err != nil { if err != nil {
return nil, err return nil, err
} }
compiler := newSchemaCompiler(standardSchemaLoader{root: schemaRoot}) if err := ctx.Err(); err != nil {
schema, err := compiler.Compile(resolvedSchemaPath) return nil, err
}
compiler := newSchemaCompiler(standardSchemaLoader{ctx: ctx, root: schemaRoot})
resourceURL := fileSchemaResourceURL(resolvedSchemaPath)
schema, err := compileJSONSchema(ctx, resourceURL.String(), compiler.Compile)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err) return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err)
} }
if err := schema.Validate(instance); err != nil { return executeJSONSchema(ctx, schema, instance)
return []string{fmt.Sprintf("json schema validation failed: %v", err)}, nil
}
return nil, nil
} }
func (v *FSValidator) validateJSONSchema(instance any, schemaPath string) ([]string, error) { func (v *FSValidator) validateJSONSchema(ctx context.Context, instance any, schemaPath string) ([]string, error) {
schemaName, schemaDoc, err := v.loadSchemaDocument(schemaPath) schemaName, schemaDoc, err := v.loadSchemaDocument(ctx, schemaPath)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -245,71 +279,60 @@ func (v *FSValidator) validateJSONSchema(instance any, schemaPath string) ([]str
if err := validateSchemaDialect(schemaDoc); err != nil { if err := validateSchemaDialect(schemaDoc); err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", schemaName, err) return nil, fmt.Errorf("failed to compile JSON schema %q: %w", schemaName, err)
} }
compiler := newSchemaCompiler(fsSchemaLoader{fsys: v.fsys, root: filecatalog.CleanFSRoot(v.root)}) compiler := newSchemaCompiler(fsSchemaLoader{ctx: ctx, fsys: v.fsys, root: filecatalog.CleanFSRoot(v.root)})
if err := compiler.AddResource(resourceURL, schemaDoc); err != nil { if err := ctx.Err(); err != nil {
return nil, err
}
if err := compiler.AddResource(resourceURL.String(), schemaDoc); err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return nil, contextErr
}
return nil, fmt.Errorf("failed to register JSON schema %q: %w", schemaName, err) return nil, fmt.Errorf("failed to register JSON schema %q: %w", schemaName, err)
} }
schema, err := compiler.Compile(resourceURL) if err := ctx.Err(); err != nil {
return nil, err
}
schema, err := compileJSONSchema(ctx, resourceURL.String(), compiler.Compile)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", schemaName, err) return nil, fmt.Errorf("failed to compile JSON schema %q: %w", schemaName, err)
} }
if err := schema.Validate(instance); err != nil { return executeJSONSchema(ctx, schema, instance)
return []string{fmt.Sprintf("json schema validation failed: %v", err)}, nil
}
return nil, nil
} }
func parseJSON(body []byte) (any, error) { func decodeJSONValue(ctx context.Context, body []byte) (any, error) {
var v any if err := ctx.Err(); err != nil {
if err := json.Unmarshal(body, &v); err != nil {
return nil, err return nil, err
} }
return v, nil decoder := json.NewDecoder(bytes.NewReader(body))
} decoder.UseNumber()
func (v *StandardValidator) LoadSchemaDocument(ctx context.Context, schemaPath string) (any, error) { var value any
select { decodeErr := decoder.Decode(&value)
case <-ctx.Done(): if err := ctx.Err(); err != nil {
return nil, ctx.Err()
default:
}
resolved, err := v.resolveSchemaPath(schemaPath)
if err != nil {
return nil, err return nil, err
} }
if decodeErr != nil {
raw, err := os.ReadFile(resolved) return nil, decodeErr
if err != nil {
return nil, fmt.Errorf("failed to read schema file %q: %w", resolved, err)
} }
var doc any var trailing any
if err := json.Unmarshal(raw, &doc); err != nil { trailingErr := decoder.Decode(&trailing)
return nil, fmt.Errorf("failed to decode JSON schema %q: %w", resolved, err) if err := ctx.Err(); err != nil {
}
if err := validateSchemaDialect(doc); err != nil {
return nil, fmt.Errorf("failed to decode JSON schema %q: %w", resolved, err)
}
return doc, nil
}
func (v *FSValidator) LoadSchemaDocument(ctx context.Context, schemaPath string) (any, error) {
select {
case <-ctx.Done():
return nil, ctx.Err()
default:
}
_, doc, err := v.loadSchemaDocument(schemaPath)
if err != nil {
return nil, err return nil, err
} }
return doc, nil if errors.Is(trailingErr, io.EOF) {
return value, nil
} else if trailingErr != nil {
return nil, trailingErr
}
return nil, errors.New("multiple JSON values")
} }
func (v *StandardValidator) resolveSchemaPath(schemaPath string) (string, error) { func (v *StandardValidator) resolveSchemaPath(ctx context.Context, schemaPath string) (string, error) {
if err := ctx.Err(); err != nil {
return "", err
}
if strings.TrimSpace(schemaPath) == "" { if strings.TrimSpace(schemaPath) == "" {
return "", errors.New("schema path is required for json_schema validation") return "", errors.New("schema path is required for json_schema validation")
} }
@@ -318,13 +341,25 @@ func (v *StandardValidator) resolveSchemaPath(schemaPath string) (string, error)
if err != nil { if err != nil {
return "", err return "", err
} }
if err := ctx.Err(); err != nil {
return "", err
}
resolved, err := containedFilesystemPath(root, schemaPath) resolved, err := containedFilesystemPath(root, schemaPath)
if err != nil { if err != nil {
return "", err return "", err
} }
if err := ctx.Err(); err != nil {
return "", err
}
if _, err := os.Stat(resolved); err != nil { if _, err := os.Stat(resolved); err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return "", contextErr
}
return "", fmt.Errorf("failed to access schema file %q: %w", resolved, err) return "", fmt.Errorf("failed to access schema file %q: %w", resolved, err)
} }
if err := ctx.Err(); err != nil {
return "", err
}
return resolved, nil return resolved, nil
} }
@@ -345,28 +380,39 @@ func (v *StandardValidator) schemaRoot() (string, error) {
return resolved, nil return resolved, nil
} }
func (v *FSValidator) loadSchemaDocument(schemaPath string) (string, any, error) { func (v *FSValidator) loadSchemaDocument(ctx context.Context, schemaPath string) (string, any, error) {
resolved, err := v.resolveSchemaPath(schemaPath) resolved, err := v.resolveSchemaPath(ctx, schemaPath)
if err != nil { if err != nil {
return "", nil, err return "", nil, err
} }
raw, err := fs.ReadFile(v.fsys, resolved) raw, err := readSchemaFile(ctx, func() (fs.File, error) {
return v.fsys.Open(resolved)
})
if err != nil { if err != nil {
return "", nil, fmt.Errorf("failed to read schema file %q: %w", resolved, err) return "", nil, fmt.Errorf("failed to read schema file %q: %w", resolved, err)
} }
var doc any doc, err := decodeJSONValue(ctx, raw)
if err := json.Unmarshal(raw, &doc); err != nil { if err != nil {
return "", nil, fmt.Errorf("failed to decode JSON schema %q: %w", resolved, err) return "", nil, fmt.Errorf("failed to decode JSON schema %q: %w", resolved, err)
} }
if err := validateSchemaDialect(doc); err != nil { if err := validateSchemaDialect(doc); err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return "", nil, contextErr
}
return "", nil, fmt.Errorf("failed to decode JSON schema %q: %w", resolved, err) return "", nil, fmt.Errorf("failed to decode JSON schema %q: %w", resolved, err)
} }
if err := ctx.Err(); err != nil {
return "", nil, err
}
return resolved, doc, nil return resolved, doc, nil
} }
func (v *FSValidator) resolveSchemaPath(schemaPath string) (string, error) { func (v *FSValidator) resolveSchemaPath(ctx context.Context, schemaPath string) (string, error) {
if err := ctx.Err(); err != nil {
return "", err
}
if strings.TrimSpace(schemaPath) == "" { if strings.TrimSpace(schemaPath) == "" {
return "", errors.New("schema path is required for json_schema validation") return "", errors.New("schema path is required for json_schema validation")
} }
@@ -377,8 +423,14 @@ func (v *FSValidator) resolveSchemaPath(schemaPath string) (string, error) {
cleanRoot := filecatalog.CleanFSRoot(v.root) cleanRoot := filecatalog.CleanFSRoot(v.root)
rootInfo, err := fs.Stat(v.fsys, cleanRoot) rootInfo, err := fs.Stat(v.fsys, cleanRoot)
if err != nil { if err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return "", contextErr
}
return "", fmt.Errorf("failed to access schema source %q: %w", cleanRoot, err) return "", fmt.Errorf("failed to access schema source %q: %w", cleanRoot, err)
} }
if err := ctx.Err(); err != nil {
return "", err
}
var resolved string var resolved string
if rootInfo.IsDir() { if rootInfo.IsDir() {
@@ -392,15 +444,21 @@ func (v *FSValidator) resolveSchemaPath(schemaPath string) (string, error) {
if err != nil { if err != nil {
return "", err return "", err
} }
if cleanSchemaPath != path.Base(cleanRoot) { if cleanSchemaPath != strings.TrimSpace(path.Base(cleanRoot)) {
return "", fmt.Errorf("schema path %q does not match schema file %q", cleanSchemaPath, path.Base(cleanRoot)) return "", fmt.Errorf("schema path %q does not match schema file %q", cleanSchemaPath, path.Base(cleanRoot))
} }
resolved = cleanRoot resolved = cleanRoot
} }
if _, err := fs.Stat(v.fsys, resolved); err != nil { if _, err := fs.Stat(v.fsys, resolved); err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return "", contextErr
}
return "", fmt.Errorf("failed to access schema file %q: %w", resolved, err) return "", fmt.Errorf("failed to access schema file %q: %w", resolved, err)
} }
if err := ctx.Err(); err != nil {
return "", err
}
return resolved, nil return resolved, nil
} }
@@ -416,8 +474,19 @@ func cleanSchemaFSPath(schemaPath string) (string, error) {
return cleaned, nil return cleaned, nil
} }
func fsSchemaResourceURL(schemaName string) string { func fileSchemaResourceURL(schemaName string) *url.URL {
return "promptkit-schema:///" + strings.TrimPrefix(path.Clean(schemaName), "/") filePath := filepath.ToSlash(schemaName)
if runtime.GOOS == "windows" && !strings.HasPrefix(filePath, "/") {
filePath = "/" + filePath
}
return &url.URL{Scheme: "file", Path: filePath}
}
func fsSchemaResourceURL(schemaName string) *url.URL {
return &url.URL{
Scheme: "promptkit-schema",
Path: "/" + strings.TrimPrefix(path.Clean(schemaName), "/"),
}
} }
func newSchemaCompiler(loader jsonschema.URLLoader) *jsonschema.Compiler { func newSchemaCompiler(loader jsonschema.URLLoader) *jsonschema.Compiler {
@@ -427,6 +496,35 @@ func newSchemaCompiler(loader jsonschema.URLLoader) *jsonschema.Compiler {
return compiler return compiler
} }
type schemaExecutor interface {
Validate(instance any) error
}
func compileJSONSchema(ctx context.Context, resourceURL string, compile func(string) (*jsonschema.Schema, error)) (*jsonschema.Schema, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
schema, compileErr := compile(resourceURL)
if err := ctx.Err(); err != nil {
return nil, err
}
return schema, compileErr
}
func executeJSONSchema(ctx context.Context, schema schemaExecutor, instance any) ([]string, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
validationErr := schema.Validate(instance)
if err := ctx.Err(); err != nil {
return nil, err
}
if validationErr != nil {
return []string{fmt.Sprintf("json schema validation failed: %v", validationErr)}, nil
}
return nil, nil
}
func validateSchemaDialect(doc any) error { func validateSchemaDialect(doc any) error {
object, ok := doc.(map[string]any) object, ok := doc.(map[string]any)
if !ok { if !ok {
@@ -447,19 +545,33 @@ func validateSchemaDialect(doc any) error {
} }
type standardSchemaLoader struct { type standardSchemaLoader struct {
ctx context.Context
root string root string
} }
func (l standardSchemaLoader) Load(resourceURL string) (any, error) { func (l standardSchemaLoader) Load(resourceURL string) (any, error) {
if err := l.ctx.Err(); err != nil {
return nil, err
}
parsed, err := url.Parse(resourceURL)
if err != nil || parsed.Scheme != "file" || parsed.Host != "" || parsed.RawQuery != "" || parsed.Opaque != "" {
return nil, fmt.Errorf("schema reference %q is not a contained file reference", resourceURL)
}
fileName, err := (jsonschema.FileLoader{}).ToFile(resourceURL) fileName, err := (jsonschema.FileLoader{}).ToFile(resourceURL)
if err != nil { if err != nil {
return nil, fmt.Errorf("schema reference %q is not a contained file reference: %w", resourceURL, err) return nil, fmt.Errorf("schema reference %q is not a contained file reference: %w", resourceURL, err)
} }
resolved, err := containedFilesystemPath(l.root, fileName) resolved, err := containedFilesystemPath(l.root, fileName)
if err != nil { if err != nil {
if contextErr := l.ctx.Err(); contextErr != nil {
return nil, contextErr
}
return nil, err return nil, err
} }
return loadJSONSchemaFile(resolved) if err := l.ctx.Err(); err != nil {
return nil, err
}
return loadJSONSchemaFile(l.ctx, resolved)
} }
func containedFilesystemPath(root, name string) (string, error) { func containedFilesystemPath(root, name string) (string, error) {
@@ -485,38 +597,94 @@ func containedFilesystemPath(root, name string) (string, error) {
return candidate, nil return candidate, nil
} }
func loadJSONSchemaFile(name string) (any, error) { func loadJSONSchemaFile(ctx context.Context, name string) (any, error) {
raw, err := os.ReadFile(name) raw, err := readSchemaFile(ctx, func() (fs.File, error) {
return os.Open(name)
})
if err != nil { if err != nil {
return nil, err return nil, err
} }
var doc any doc, err := decodeJSONValue(ctx, raw)
if err := json.Unmarshal(raw, &doc); err != nil { if err != nil {
return nil, err return nil, err
} }
if err := validateSchemaDialect(doc); err != nil { if err := validateSchemaDialect(doc); err != nil {
if contextErr := ctx.Err(); contextErr != nil {
return nil, contextErr
}
return nil, err
}
if err := ctx.Err(); err != nil {
return nil, err return nil, err
} }
return doc, nil return doc, nil
} }
func readSchemaFile(ctx context.Context, open func() (fs.File, error)) ([]byte, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
file, openErr := open()
if err := ctx.Err(); err != nil {
if file != nil {
_ = file.Close()
}
return nil, err
}
if openErr != nil {
if file != nil {
_ = file.Close()
}
return nil, openErr
}
if file == nil {
return nil, errors.New("schema source returned a nil file")
}
defer file.Close()
var contents []byte
chunk := make([]byte, schemaReadChunkSize)
for {
if err := ctx.Err(); err != nil {
return nil, err
}
n, readErr := file.Read(chunk)
if n > 0 {
contents = append(contents, chunk[:n]...)
}
if err := ctx.Err(); err != nil {
return nil, err
}
if errors.Is(readErr, io.EOF) {
return contents, nil
}
if readErr != nil {
return nil, readErr
}
if n == 0 {
return nil, io.ErrNoProgress
}
}
}
type fsSchemaLoader struct { type fsSchemaLoader struct {
ctx context.Context
fsys fs.FS fsys fs.FS
root string root string
} }
func (l fsSchemaLoader) Load(resourceURL string) (any, error) { func (l fsSchemaLoader) Load(resourceURL string) (any, error) {
if err := l.ctx.Err(); err != nil {
return nil, err
}
parsed, err := url.Parse(resourceURL) parsed, err := url.Parse(resourceURL)
if err != nil { if err != nil {
return nil, fmt.Errorf("invalid schema reference %q: %w", resourceURL, err) return nil, fmt.Errorf("invalid schema reference %q: %w", resourceURL, err)
} }
if parsed.Scheme != "promptkit-schema" || parsed.Host != "" { if parsed.Scheme != "promptkit-schema" || parsed.Host != "" || parsed.RawQuery != "" || parsed.Opaque != "" {
return nil, fmt.Errorf("schema reference %q is not allowed", resourceURL) return nil, fmt.Errorf("schema reference %q is not allowed", resourceURL)
} }
name, err := url.PathUnescape(strings.TrimPrefix(parsed.Path, "/")) name := strings.TrimPrefix(parsed.Path, "/")
if err != nil {
return nil, fmt.Errorf("invalid schema reference %q: %w", resourceURL, err)
}
name = path.Clean(name) name = path.Clean(name)
if l.root == "." { if l.root == "." {
if strings.HasPrefix(name, "../") || name == ".." { if strings.HasPrefix(name, "../") || name == ".." {
@@ -528,21 +696,35 @@ func (l fsSchemaLoader) Load(resourceURL string) (any, error) {
rootInfo, err := fs.Stat(l.fsys, l.root) rootInfo, err := fs.Stat(l.fsys, l.root)
if err != nil { if err != nil {
if contextErr := l.ctx.Err(); contextErr != nil {
return nil, contextErr
}
return nil, err
}
if err := l.ctx.Err(); err != nil {
return nil, err return nil, err
} }
if !rootInfo.IsDir() && name != l.root { if !rootInfo.IsDir() && name != l.root {
return nil, fmt.Errorf("schema reference %q is outside the configured schema file", resourceURL) return nil, fmt.Errorf("schema reference %q is outside the configured schema file", resourceURL)
} }
raw, err := fs.ReadFile(l.fsys, name) raw, err := readSchemaFile(l.ctx, func() (fs.File, error) {
return l.fsys.Open(name)
})
if err != nil { if err != nil {
return nil, err return nil, err
} }
var doc any doc, err := decodeJSONValue(l.ctx, raw)
if err := json.Unmarshal(raw, &doc); err != nil { if err != nil {
return nil, err return nil, err
} }
if err := validateSchemaDialect(doc); err != nil { if err := validateSchemaDialect(doc); err != nil {
if contextErr := l.ctx.Err(); contextErr != nil {
return nil, contextErr
}
return nil, err
}
if err := l.ctx.Err(); err != nil {
return nil, err return nil, err
} }
return doc, nil return doc, nil

View File

@@ -3,9 +3,11 @@ package validate
import ( import (
"context" "context"
"encoding/json" "encoding/json"
"net/url"
"os" "os"
"path/filepath" "path/filepath"
"reflect" "reflect"
"runtime"
"strconv" "strconv"
"strings" "strings"
"testing" "testing"
@@ -62,6 +64,43 @@ func TestStandardValidatorBasicFailureEmpty(t *testing.T) {
} }
} }
func TestStandardValidatorReportsZeroRepairAttempts(t *testing.T) {
v := NewStandardValidator("")
for _, tc := range []struct {
name string
body string
contract domain.OutputContract
}{
{
name: "passed basic validation",
body: "answer",
contract: domain.OutputContract{
ValidationMode: domain.ValidationBasic,
RepairAttempts: 3,
},
},
{
name: "failed JSON validation",
body: `{"answer":`,
contract: domain.OutputContract{
ValidationMode: domain.ValidationJSON,
RepairAttempts: 3,
},
},
} {
t.Run(tc.name, func(t *testing.T) {
res, err := v.Validate(context.Background(), &domain.Artifact{Body: []byte(tc.body)}, tc.contract)
if err != nil {
t.Fatalf("validate: %v", err)
}
if res.RepairAttempts != 0 {
t.Fatalf("repair attempts = %d, want 0", res.RepairAttempts)
}
})
}
}
func TestStandardValidatorJSONSuccess(t *testing.T) { func TestStandardValidatorJSONSuccess(t *testing.T) {
v := NewStandardValidator("") v := NewStandardValidator("")
@@ -93,6 +132,51 @@ func TestStandardValidatorJSONFailure(t *testing.T) {
} }
} }
func TestJSONValidationChecksCompleteSyntaxWithoutChangingArtifact(t *testing.T) {
tests := []struct {
name string
body string
wantValid bool
}{
{name: "ordinary object", body: `{"count":2,"ok":true}`, wantValid: true},
{name: "integer at exact float boundary", body: `9007199254740992`, wantValid: true},
{name: "integer beyond exact float boundary", body: `9007199254740993`, wantValid: true},
{name: "large exponent", body: `1e400`, wantValid: true},
{name: "precise decimal", body: `0.123456789012345678901234567890`, wantValid: true},
{name: "surrounding whitespace", body: " \n [1,2,3] \t", wantValid: true},
{name: "malformed document", body: `{"count":`, wantValid: false},
{name: "trailing value", body: `1 2`, wantValid: false},
}
validator := NewStandardValidator("")
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
body := []byte(tc.body)
before := append([]byte(nil), body...)
artifact := &domain.Artifact{Body: body}
result, err := validator.Validate(context.Background(), artifact, domain.OutputContract{
ValidationMode: domain.ValidationJSON,
})
if err != nil {
t.Fatalf("validate JSON: %v", err)
}
if result.IsValid != tc.wantValid {
t.Fatalf("valid = %v, want %v; result=%+v", result.IsValid, tc.wantValid, result)
}
if tc.wantValid && result.Status != domain.ValidationPassed {
t.Fatalf("status = %q, want %q", result.Status, domain.ValidationPassed)
}
if !tc.wantValid && result.Status != domain.ValidationFailed {
t.Fatalf("status = %q, want %q", result.Status, domain.ValidationFailed)
}
if !reflect.DeepEqual(artifact.Body, before) {
t.Fatalf("artifact body changed: got %q, want %q", artifact.Body, before)
}
})
}
}
func TestStandardValidatorJSONSchemaSuccess(t *testing.T) { func TestStandardValidatorJSONSchemaSuccess(t *testing.T) {
tmp := t.TempDir() tmp := t.TempDir()
schemaPath := filepath.Join(tmp, "schema.json") schemaPath := filepath.Join(tmp, "schema.json")
@@ -285,55 +369,6 @@ func TestStandardValidatorJSONSchemaCompilationError(t *testing.T) {
} }
} }
func TestStandardValidatorLoadSchemaDocumentSuccess(t *testing.T) {
tmp := t.TempDir()
if err := os.WriteFile(filepath.Join(tmp, "schema.json"), []byte(`{
"type": "object",
"properties": {
"name": {"type": "string"}
}
}`), 0644); err != nil {
t.Fatal(err)
}
v := NewStandardValidator(tmp)
loader, ok := v.(SchemaDocumentLoader)
if !ok {
t.Fatal("standard validator must implement SchemaDocumentLoader")
}
doc, err := loader.LoadSchemaDocument(context.Background(), "schema.json")
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
obj, ok := doc.(map[string]any)
if !ok {
t.Fatalf("expected object document, got %#v", doc)
}
if obj["type"] != "object" {
t.Fatalf("expected schema type=object, got %#v", obj["type"])
}
}
func TestStandardValidatorLoadSchemaDocumentInvalidJSON(t *testing.T) {
tmp := t.TempDir()
if err := os.WriteFile(filepath.Join(tmp, "schema.json"), []byte(`{`), 0644); err != nil {
t.Fatal(err)
}
v := NewStandardValidator(tmp)
loader, ok := v.(SchemaDocumentLoader)
if !ok {
t.Fatal("standard validator must implement SchemaDocumentLoader")
}
_, err := loader.LoadSchemaDocument(context.Background(), "schema.json")
if err == nil {
t.Fatal("expected decode error")
}
}
func TestFSValidatorJSONSchemaSuccess(t *testing.T) { func TestFSValidatorJSONSchemaSuccess(t *testing.T) {
v := NewFSValidator(fstest.MapFS{ v := NewFSValidator(fstest.MapFS{
"schemas/events.schema.json": &fstest.MapFile{Data: []byte(`{ "schemas/events.schema.json": &fstest.MapFile{Data: []byte(`{
@@ -416,17 +451,70 @@ func TestFSValidatorPreparedSchemaSurvivesSourceMutation(t *testing.T) {
} }
} }
func TestFSValidatorJSONSchemaRegistrationError(t *testing.T) { func TestFSValidatorEscapesSchemaResourcePath(t *testing.T) {
for _, schemaName := range []string{"%zz.json", "space name.json", "hash#.json", "query?.json", "rún.json"} {
t.Run(schemaName, func(t *testing.T) {
v := NewFSValidator(fstest.MapFS{ v := NewFSValidator(fstest.MapFS{
"schemas/%zz.json": &fstest.MapFile{Data: []byte(`{"type":"object"}`)}, "schemas/" + schemaName: &fstest.MapFile{Data: []byte(`{"type":"object"}`)},
}, "schemas") }, "schemas")
res, err := v.Validate(context.Background(), &domain.Artifact{Body: []byte(`{}`)}, domain.OutputContract{ res, err := v.Validate(context.Background(), &domain.Artifact{Body: []byte(`{}`)}, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema, ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "%zz.json", SchemaPath: schemaName,
}) })
if err == nil || !strings.Contains(err.Error(), "failed to register JSON schema") { if err != nil || !res.IsValid {
t.Fatalf("expected schema registration error, got result=%#v error=%v", res, err) t.Fatalf("validate schema %q: result=%#v error=%v", schemaName, res, err)
}
})
}
}
func TestSchemaReferencesPreserveEscapedFilenames(t *testing.T) {
for _, name := range []string{"%2F.json", "space name.json", "hash#.json", "query?.json", "rún.json"} {
t.Run(name, func(t *testing.T) {
rootName := "root-" + name
childName := "child-" + name
childPath := "nested/" + childName
reference := (&url.URL{Path: childPath}).EscapedPath()
rootSchema := []byte(`{"$ref":` + strconv.Quote(reference) + `}`)
childSchema := []byte(`{"type":"integer","minimum":2}`)
t.Run("fs.FS", func(t *testing.T) {
validator := NewFSValidator(fstest.MapFS{
"schemas/" + rootName: &fstest.MapFile{Data: rootSchema},
"schemas/" + childPath: &fstest.MapFile{Data: childSchema},
}, "schemas")
assertSchemaValidation(t, validator, rootName)
})
t.Run("operating system files", func(t *testing.T) {
if runtime.GOOS == "windows" && strings.ContainsAny(rootName+childName, `<>:"/\|?*`) {
t.Skip("filename is not legal on Windows")
}
root := t.TempDir()
if err := os.Mkdir(filepath.Join(root, "nested"), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(root, rootName), rootSchema, 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(root, filepath.FromSlash(childPath)), childSchema, 0o644); err != nil {
t.Fatal(err)
}
assertSchemaValidation(t, NewStandardValidator(root), rootName)
})
})
}
}
func assertSchemaValidation(t *testing.T, validator Validator, schemaPath string) {
t.Helper()
result, err := validator.Validate(context.Background(), &domain.Artifact{Body: []byte(`2`)}, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: schemaPath,
})
if err != nil || !result.IsValid {
t.Fatalf("validate schema %q: result=%#v error=%v", schemaPath, result, err)
} }
} }
@@ -516,25 +604,6 @@ func TestFSValidatorSingleSchemaFileUsesBaseName(t *testing.T) {
} }
} }
func TestFSValidatorLoadSchemaDocument(t *testing.T) {
v := NewFSValidator(fstest.MapFS{
"schemas/schema.json": &fstest.MapFile{Data: []byte(`{"type":"object"}`)},
}, "schemas")
loader, ok := v.(SchemaDocumentLoader)
if !ok {
t.Fatal("fs validator must implement SchemaDocumentLoader")
}
doc, err := loader.LoadSchemaDocument(context.Background(), "schema.json")
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
obj, ok := doc.(map[string]any)
if !ok || obj["type"] != "object" {
t.Fatalf("unexpected schema document: %#v", doc)
}
}
func TestStandardValidatorJSONSchemaReferenceBoundaries(t *testing.T) { func TestStandardValidatorJSONSchemaReferenceBoundaries(t *testing.T) {
root := t.TempDir() root := t.TempDir()
if err := os.WriteFile(filepath.Join(root, "child.json"), []byte(`{ if err := os.WriteFile(filepath.Join(root, "child.json"), []byte(`{
@@ -629,6 +698,225 @@ func TestFSValidatorJSONSchemaReferenceBoundaries(t *testing.T) {
} }
} }
func TestJSONSchemaNumericConstraintsRetainJSONPrecision(t *testing.T) {
tests := []struct {
name string
schema string
instance string
wantValid bool
}{
{
name: "const distinguishes adjacent large integers",
schema: `{"const":9007199254740993}`,
instance: `9007199254740993`,
wantValid: true,
},
{
name: "const rejects adjacent large integer",
schema: `{"const":9007199254740993}`,
instance: `9007199254740992`,
wantValid: false,
},
{
name: "const accepts exponent beyond float range",
schema: `{"const":1e400}`,
instance: `1e400`,
wantValid: true,
},
{
name: "minimum accepts precise decimal boundary",
schema: `{"type":"number","minimum":0.123456789012345678901234567890}`,
instance: `0.123456789012345678901234567890`,
wantValid: true,
},
{
name: "minimum rejects lower precise decimal",
schema: `{"type":"number","minimum":0.123456789012345678901234567890}`,
instance: `0.123456789012345678901234567889`,
wantValid: false,
},
{
name: "maximum distinguishes adjacent large integers",
schema: `{"type":"number","maximum":9007199254740992}`,
instance: `9007199254740993`,
wantValid: false,
},
{
name: "multiple of accepts exact decimal multiple",
schema: `{"type":"number","multipleOf":0.0000000000000000001}`,
instance: `0.0000000000000000003`,
wantValid: true,
},
{
name: "multiple of rejects inexact decimal multiple",
schema: `{"type":"number","multipleOf":0.0000000000000000001}`,
instance: `0.00000000000000000031`,
wantValid: false,
},
{
name: "ordinary number remains supported",
schema: `{"type":"number","minimum":1,"maximum":3}`,
instance: `2`,
wantValid: true,
},
}
for _, source := range jsonSchemaValidatorSources() {
t.Run(source.name, func(t *testing.T) {
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
validator := source.new(t, []byte(tc.schema))
result, err := validator.Validate(context.Background(), &domain.Artifact{Body: []byte(tc.instance)}, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
if err != nil {
t.Fatalf("validate JSON Schema instance: %v", err)
}
if result.IsValid != tc.wantValid {
t.Fatalf("valid = %v, want %v; result=%+v", result.IsValid, tc.wantValid, result)
}
})
}
})
}
}
func TestJSONSchemaDecodingRequiresOneCompleteDocument(t *testing.T) {
for _, source := range jsonSchemaValidatorSources() {
t.Run(source.name, func(t *testing.T) {
for _, schema := range []string{`{"type":`, `{} {}`} {
validator := source.new(t, []byte(schema))
_, err := validator.Validate(context.Background(), &domain.Artifact{Body: []byte(`1`)}, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
if err == nil {
t.Fatalf("schema %q: expected decoding error", schema)
}
}
validator := source.new(t, []byte(`{}`))
for _, instance := range []string{`{"value":`, `1 2`} {
result, err := validator.Validate(context.Background(), &domain.Artifact{Body: []byte(instance)}, domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
if err != nil {
t.Fatalf("instance %q: expected completed validation, got %v", instance, err)
}
if result.Status != domain.ValidationFailed || result.IsValid || len(result.Errors) == 0 {
t.Fatalf("instance %q: expected failed validation, got %+v", instance, result)
}
}
})
}
}
func TestPreparedSchemaDocumentsRetainExactNumbers(t *testing.T) {
const schema = `{
"const": 9007199254740993,
"minimum": 0.123456789012345678901234567890,
"maximum": 1e400,
"multipleOf": 0.0000000000000000001
}`
want := map[string]string{
"const": "9007199254740993",
"minimum": "0.123456789012345678901234567890",
"maximum": "1e400",
"multipleOf": "0.0000000000000000001",
}
for _, source := range jsonSchemaValidatorSources() {
t.Run(source.name, func(t *testing.T) {
validator := source.new(t, []byte(schema))
preparer, ok := validator.(ValidationPreparer)
if !ok {
t.Fatal("validator does not support validation preparation")
}
prepared, err := preparer.PrepareValidation(context.Background(), domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
if err != nil {
t.Fatalf("prepare validation: %v", err)
}
document, ok := prepared.SchemaDocument().(map[string]any)
if !ok {
t.Fatalf("schema document = %#v, want object", prepared.SchemaDocument())
}
for name, wantNumber := range want {
got, ok := document[name].(json.Number)
if !ok {
t.Fatalf("schema field %q = %#v, want json.Number", name, document[name])
}
if got.String() != wantNumber {
t.Fatalf("schema field %q = %q, want %q", name, got, wantNumber)
}
}
})
}
}
type jsonSchemaValidatorSource struct {
name string
new func(*testing.T, []byte) Validator
}
func jsonSchemaValidatorSources() []jsonSchemaValidatorSource {
return []jsonSchemaValidatorSource{
{
name: "operating system files",
new: func(t *testing.T, schema []byte) Validator {
t.Helper()
root := t.TempDir()
if err := os.WriteFile(filepath.Join(root, "schema.json"), schema, 0o644); err != nil {
t.Fatalf("write schema: %v", err)
}
return NewStandardValidator(root)
},
},
{
name: "fs.FS",
new: func(t *testing.T, schema []byte) Validator {
t.Helper()
return NewFSValidator(fstest.MapFS{
"schema.json": &fstest.MapFile{Data: schema},
}, ".")
},
},
}
}
func BenchmarkJSONValidation(b *testing.B) {
largeArray := []byte(`[` + strings.Repeat(`12345678901234567890,`, 32*1024) + `0]`)
benchmarks := []struct {
name string
body []byte
}{
{name: "scalar", body: []byte(`1e400`)},
{name: "object", body: []byte(`{"name":"eris","count":9007199254740993,"enabled":true}`)},
{name: "large array", body: largeArray},
}
validator := NewStandardValidator("")
contract := domain.OutputContract{ValidationMode: domain.ValidationJSON}
for _, benchmark := range benchmarks {
b.Run(benchmark.name, func(b *testing.B) {
artifact := &domain.Artifact{Body: benchmark.body}
b.ReportAllocs()
b.SetBytes(int64(len(benchmark.body)))
b.ResetTimer()
for range b.N {
result, err := validator.Validate(context.Background(), artifact, contract)
if err != nil || !result.IsValid {
b.Fatalf("validate JSON: result=%+v error=%v", result, err)
}
}
})
}
}
func TestJSONSchemaDialectIsDraft2020(t *testing.T) { func TestJSONSchemaDialectIsDraft2020(t *testing.T) {
tests := []struct { tests := []struct {
name string name string
@@ -673,8 +961,8 @@ func TestJSONSchemaDialectIsDraft2020(t *testing.T) {
func assertSchemaDocument(t *testing.T, got any, expectedJSON []byte) { func assertSchemaDocument(t *testing.T, got any, expectedJSON []byte) {
t.Helper() t.Helper()
var expected any expected, err := decodeJSONValue(context.Background(), expectedJSON)
if err := json.Unmarshal(expectedJSON, &expected); err != nil { if err != nil {
t.Fatalf("decode expected schema document: %v", err) t.Fatalf("decode expected schema document: %v", err)
} }
if !reflect.DeepEqual(got, expected) { if !reflect.DeepEqual(got, expected) {

View File

@@ -7,11 +7,16 @@ import (
) )
// Validator validates the generated artifact based on the output contract. // Validator validates the generated artifact based on the output contract.
// Validation checks ctx around Promptkit-controlled work and synchronous
// dependency calls. A dependency call already in progress cannot be preempted;
// after it returns, cancellation takes precedence over its result.
type Validator interface { type Validator interface {
Validate(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract) (domain.ValidationResult, error) Validate(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract) (domain.ValidationResult, error)
} }
// PreparedValidation validates artifacts against one frozen output contract. // PreparedValidation validates artifacts against one frozen output contract.
// Its cancellation boundary is synchronous: Validate does not detach schema
// execution, and an observed context error prevents publication of a result.
type PreparedValidation interface { type PreparedValidation interface {
Validate(ctx context.Context, artifact *domain.Artifact) (domain.ValidationResult, error) Validate(ctx context.Context, artifact *domain.Artifact) (domain.ValidationResult, error)
// SchemaDocument returns the root JSON Schema document used for provider // SchemaDocument returns the root JSON Schema document used for provider
@@ -21,11 +26,9 @@ type PreparedValidation interface {
} }
// ValidationPreparer freezes validation resources for one output contract. // ValidationPreparer freezes validation resources for one output contract.
// Preparation reads schemas in context-checked chunks and checks ctx around
// decoding and compilation. Filesystem and compiler calls remain synchronous,
// so cancellation becomes authoritative when an in-progress call returns.
type ValidationPreparer interface { type ValidationPreparer interface {
PrepareValidation(ctx context.Context, contract domain.OutputContract) (PreparedValidation, error) PrepareValidation(ctx context.Context, contract domain.OutputContract) (PreparedValidation, error)
} }
// SchemaDocumentLoader loads JSON schema documents using validator path semantics.
type SchemaDocumentLoader interface {
LoadSchemaDocument(ctx context.Context, schemaPath string) (any, error)
}

124
json.go
View File

@@ -2,9 +2,16 @@ package promptkit
import ( import (
"encoding/json" "encoding/json"
"fmt"
"math"
"time" "time"
) )
const (
minDurationMilliseconds = int64(time.Duration(math.MinInt64) / time.Millisecond)
maxDurationMilliseconds = int64(time.Duration(math.MaxInt64) / time.Millisecond)
)
// MarshalJSON implements json.Marshaler for PreparedRun. It uses RFC 3339 // MarshalJSON implements json.Marshaler for PreparedRun. It uses RFC 3339
// timestamps, integer duration_ms, and omits zero timing values. // timestamps, integer duration_ms, and omits zero timing values.
func (r PreparedRun) MarshalJSON() ([]byte, error) { func (r PreparedRun) MarshalJSON() ([]byte, error) {
@@ -21,35 +28,8 @@ func (r PreparedRun) MarshalJSON() ([]byte, error) {
durationMS = &r.DurationMS durationMS = &r.DurationMS
} }
return json.Marshal(struct { return json.Marshal(preparedRunJSON{
PromptID string `json:"prompt_id"` preparedRunJSONFields: preparedRunJSONFields(r),
PromptVersion string `json:"prompt_version,omitempty"`
PromptHash string `json:"prompt_hash,omitempty"`
SelectedProfileID string `json:"selected_profile_id"`
SelectedBackendID string `json:"selected_backend_id,omitempty"`
EffectiveModelParams ExecutionTarget `json:"effective_model_params"`
OutputContract OutputContract `json:"output_contract"`
StructuredOutput *StructuredOutputSpec `json:"structured_output,omitempty"`
InputHashes map[string]string `json:"input_hashes,omitempty"`
SessionID string `json:"session_id,omitempty"`
RenderedPromptHash string `json:"rendered_prompt_hash"`
Messages []RenderedMessage `json:"messages"`
StartTime *time.Time `json:"start_time,omitempty"`
EndTime *time.Time `json:"end_time,omitempty"`
DurationMS *int64 `json:"duration_ms,omitempty"`
}{
PromptID: r.PromptID,
PromptVersion: r.PromptVersion,
PromptHash: r.PromptHash,
SelectedProfileID: r.SelectedProfileID,
SelectedBackendID: r.SelectedBackendID,
EffectiveModelParams: r.EffectiveModelParams,
OutputContract: r.OutputContract,
StructuredOutput: r.StructuredOutput,
InputHashes: r.InputHashes,
SessionID: r.SessionID,
RenderedPromptHash: r.RenderedPromptHash,
Messages: r.Messages,
StartTime: startTime, StartTime: startTime,
EndTime: endTime, EndTime: endTime,
DurationMS: durationMS, DurationMS: durationMS,
@@ -74,22 +54,7 @@ func (r RunResult) MarshalJSON() ([]byte, error) {
} }
return json.Marshal(runResultJSON{ return json.Marshal(runResultJSON{
RunID: r.RunID, runResultJSONFields: runResultJSONFields(r),
Artifact: r.Artifact,
RawOutput: r.RawOutput,
Validation: r.Validation,
PromptID: r.PromptID,
PromptVersion: r.PromptVersion,
PromptHash: r.PromptHash,
SessionID: r.SessionID,
RenderedPromptHash: r.RenderedPromptHash,
SelectedProfileID: r.SelectedProfileID,
SelectedBackendID: r.SelectedBackendID,
ModelName: r.ModelName,
Endpoint: r.Endpoint,
EffectiveModelParams: r.EffectiveModelParams,
InputHashes: r.InputHashes,
Usage: r.Usage,
StartTime: startTime, StartTime: startTime,
EndTime: endTime, EndTime: endTime,
DurationMS: durationMS, DurationMS: durationMS,
@@ -97,66 +62,49 @@ func (r RunResult) MarshalJSON() ([]byte, error) {
} }
// UnmarshalJSON implements json.Unmarshaler for RunResult. It decodes // UnmarshalJSON implements json.Unmarshaler for RunResult. It decodes
// duration_ms into Duration with millisecond precision. // duration_ms into Duration with millisecond precision. A duration_ms outside
// the range representable by time.Duration returns an error without changing
// the receiver.
func (r *RunResult) UnmarshalJSON(data []byte) error { func (r *RunResult) UnmarshalJSON(data []byte) error {
var wire runResultJSON var wire runResultJSON
if err := json.Unmarshal(data, &wire); err != nil { if err := json.Unmarshal(data, &wire); err != nil {
return err return err
} }
*r = RunResult{ result := RunResult(wire.runResultJSONFields)
RunID: wire.RunID, if wire.DurationMS != nil {
Artifact: wire.Artifact, if *wire.DurationMS < minDurationMilliseconds || *wire.DurationMS > maxDurationMilliseconds {
RawOutput: wire.RawOutput, return fmt.Errorf(
Validation: wire.Validation, "decode RunResult duration_ms: %d cannot be represented as time.Duration",
PromptID: wire.PromptID, *wire.DurationMS,
PromptVersion: wire.PromptVersion, )
PromptHash: wire.PromptHash, }
SessionID: wire.SessionID, result.Duration = time.Duration(*wire.DurationMS) * time.Millisecond
RenderedPromptHash: wire.RenderedPromptHash,
SelectedProfileID: wire.SelectedProfileID,
SelectedBackendID: wire.SelectedBackendID,
ModelName: wire.ModelName,
Endpoint: wire.Endpoint,
EffectiveModelParams: wire.EffectiveModelParams,
InputHashes: wire.InputHashes,
Usage: wire.Usage,
Duration: time.Duration(valueOrZero(wire.DurationMS)) * time.Millisecond,
} }
if wire.StartTime != nil { if wire.StartTime != nil {
r.StartTime = *wire.StartTime result.StartTime = *wire.StartTime
} }
if wire.EndTime != nil { if wire.EndTime != nil {
r.EndTime = *wire.EndTime result.EndTime = *wire.EndTime
} }
*r = result
return nil return nil
} }
type runResultJSON struct { type preparedRunJSONFields PreparedRun
RunID string `json:"run_id"`
Artifact Artifact `json:"artifact"` type preparedRunJSON struct {
RawOutput string `json:"raw_output"` preparedRunJSONFields
Validation ValidationResult `json:"validation"`
PromptID string `json:"prompt_id"`
PromptVersion string `json:"prompt_version,omitempty"`
PromptHash string `json:"prompt_hash,omitempty"`
SessionID string `json:"session_id,omitempty"`
RenderedPromptHash string `json:"rendered_prompt_hash"`
SelectedProfileID string `json:"selected_profile_id"`
SelectedBackendID string `json:"selected_backend_id,omitempty"`
ModelName string `json:"model_name"`
Endpoint string `json:"endpoint"`
EffectiveModelParams ExecutionTarget `json:"effective_model_params"`
InputHashes map[string]string `json:"input_hashes,omitempty"`
Usage TokenUsage `json:"usage"`
StartTime *time.Time `json:"start_time,omitempty"` StartTime *time.Time `json:"start_time,omitempty"`
EndTime *time.Time `json:"end_time,omitempty"` EndTime *time.Time `json:"end_time,omitempty"`
DurationMS *int64 `json:"duration_ms,omitempty"` DurationMS *int64 `json:"duration_ms,omitempty"`
} }
func valueOrZero(value *int64) int64 { type runResultJSONFields RunResult
if value == nil {
return 0 type runResultJSON struct {
} runResultJSONFields
return *value StartTime *time.Time `json:"start_time,omitempty"`
EndTime *time.Time `json:"end_time,omitempty"`
DurationMS *int64 `json:"duration_ms,omitempty"`
} }

328
json_contract_test.go Normal file
View File

@@ -0,0 +1,328 @@
package promptkit_test
import (
"encoding/json"
"math"
"reflect"
"strconv"
"strings"
"testing"
"time"
"gitea.maximumdirect.net/eric/promptkit"
)
func TestPreparedRunJSONContractRoundTripsAllFields(t *testing.T) {
value := fullyPopulatedPreparedRun()
payload, err := json.Marshal(value)
if err != nil {
t.Fatalf("marshal PreparedRun: %v", err)
}
object := decodeJSONObject(t, payload)
assertJSONFields(t, object,
"prompt_id",
"prompt_version",
"prompt_hash",
"selected_profile_id",
"selected_backend_id",
"effective_model_params",
"output_contract",
"structured_output",
"input_hashes",
"session_id",
"rendered_prompt_hash",
"messages",
"start_time",
"end_time",
"duration_ms",
)
var messages []map[string]json.RawMessage
if err := json.Unmarshal(object["messages"], &messages); err != nil {
t.Fatalf("decode PreparedRun messages: %v", err)
}
if len(messages) != 2 {
t.Fatalf("message count = %d, want 2", len(messages))
}
if _, ok := messages[0]["cache_control"]; !ok {
t.Fatalf("first message omitted cache_control: %s", object["messages"])
}
if _, ok := messages[1]["cache_control"]; ok {
t.Fatalf("second message included empty cache_control: %s", object["messages"])
}
var decoded promptkit.PreparedRun
if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal PreparedRun: %v", err)
}
if !reflect.DeepEqual(decoded, value) {
t.Fatalf("PreparedRun did not round trip:\ngot %#v\nwant %#v", decoded, value)
}
zeroPayload, err := json.Marshal(promptkit.PreparedRun{})
if err != nil {
t.Fatalf("marshal zero PreparedRun: %v", err)
}
zeroObject := decodeJSONObject(t, zeroPayload)
for _, field := range []string{
"prompt_version",
"prompt_hash",
"selected_backend_id",
"structured_output",
"input_hashes",
"session_id",
"start_time",
"end_time",
"duration_ms",
} {
if _, ok := zeroObject[field]; ok {
t.Fatalf("zero PreparedRun included %q: %s", field, zeroPayload)
}
}
}
func TestRunResultJSONContractRoundTripsAllFields(t *testing.T) {
value := fullyPopulatedRunResult()
payload, err := json.Marshal(value)
if err != nil {
t.Fatalf("marshal RunResult: %v", err)
}
object := decodeJSONObject(t, payload)
assertJSONFields(t, object,
"run_id",
"artifact",
"raw_output",
"validation",
"prompt_id",
"prompt_version",
"prompt_hash",
"session_id",
"rendered_prompt_hash",
"selected_profile_id",
"selected_backend_id",
"model_name",
"endpoint",
"effective_model_params",
"input_hashes",
"usage",
"start_time",
"end_time",
"duration_ms",
)
if _, ok := object["duration"]; ok {
t.Fatalf("RunResult included nanosecond duration field: %s", payload)
}
var decoded promptkit.RunResult
if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal RunResult: %v", err)
}
if !reflect.DeepEqual(decoded, value) {
t.Fatalf("RunResult did not round trip:\ngot %#v\nwant %#v", decoded, value)
}
zeroPayload, err := json.Marshal(promptkit.RunResult{})
if err != nil {
t.Fatalf("marshal zero RunResult: %v", err)
}
zeroObject := decodeJSONObject(t, zeroPayload)
for _, field := range []string{
"prompt_version",
"prompt_hash",
"session_id",
"selected_backend_id",
"input_hashes",
"start_time",
"end_time",
"duration_ms",
"duration",
} {
if _, ok := zeroObject[field]; ok {
t.Fatalf("zero RunResult included %q: %s", field, zeroPayload)
}
}
}
func TestRunResultJSONDurationMillisecondBoundaries(t *testing.T) {
maxMilliseconds := int64(time.Duration(math.MaxInt64) / time.Millisecond)
minMilliseconds := int64(time.Duration(math.MinInt64) / time.Millisecond)
for _, milliseconds := range []int64{minMilliseconds, maxMilliseconds} {
t.Run(strconv.FormatInt(milliseconds, 10), func(t *testing.T) {
var decoded promptkit.RunResult
if err := json.Unmarshal(durationPayload(milliseconds), &decoded); err != nil {
t.Fatalf("decode representable duration_ms: %v", err)
}
want := time.Duration(milliseconds) * time.Millisecond
if decoded.Duration != want {
t.Fatalf("Duration = %v, want %v", decoded.Duration, want)
}
})
}
for _, milliseconds := range []int64{
minMilliseconds - 1,
maxMilliseconds + 1,
math.MinInt64,
math.MaxInt64,
} {
t.Run(strconv.FormatInt(milliseconds, 10), func(t *testing.T) {
original := fullyPopulatedRunResult()
decoded := original
err := json.Unmarshal(durationPayload(milliseconds), &decoded)
if err == nil || !strings.Contains(err.Error(), "duration_ms") || !strings.Contains(err.Error(), "time.Duration") {
t.Fatalf("overflow error = %v", err)
}
if !reflect.DeepEqual(decoded, original) {
t.Fatalf("failed decode partially updated receiver:\ngot %#v\nwant %#v", decoded, original)
}
})
}
}
func fullyPopulatedPreparedRun() promptkit.PreparedRun {
start := time.Date(2026, time.August, 11, 12, 13, 14, 150_000_000, time.UTC)
return promptkit.PreparedRun{
PromptID: "prompt.prepared",
PromptVersion: "2.1.0",
PromptHash: "prompt-hash",
SelectedProfileID: "profile-prepared",
SelectedBackendID: "backend-prepared",
EffectiveModelParams: jsonContractExecutionTarget(),
OutputContract: promptkit.OutputContract{
Format: promptkit.FormatJSON,
ValidationMode: promptkit.ValidationJSONSchema,
SchemaPath: "schemas/prepared.json",
RepairAttempts: 2,
},
StructuredOutput: &promptkit.StructuredOutputSpec{
Type: promptkit.StructuredOutputJSONSchema,
JSONSchema: &promptkit.StructuredOutputJSONSpec{
Name: "prepared_schema",
Strict: true,
Schema: map[string]any{
"type": "object",
"required": []any{"value"},
"properties": map[string]any{
"value": map[string]any{"minimum": float64(1)},
},
},
},
},
InputHashes: map[string]string{"first": "hash-1", "second": "hash-2"},
SessionID: "session-prepared",
RenderedPromptHash: "rendered-hash",
Messages: []promptkit.RenderedMessage{
{
Role: "system",
Content: "System message",
CacheControl: &promptkit.CacheControl{
Type: promptkit.CacheControlEphemeral,
TTL: "1h",
},
},
{Role: "user", Content: "User message"},
},
StartTime: start,
EndTime: start.Add(1501 * time.Millisecond),
DurationMS: 1501,
}
}
func fullyPopulatedRunResult() promptkit.RunResult {
start := time.Date(2026, time.August, 11, 15, 16, 17, 250_000_000, time.UTC)
return promptkit.RunResult{
RunID: "run-id",
Artifact: promptkit.Artifact{
Name: "result.json",
ContentType: "application/json",
Body: []byte(`{"value":2}`),
URI: "memory://result.json",
Size: 11,
Hash: "artifact-hash",
},
RawOutput: `{"value":2}`,
Validation: promptkit.ValidationResult{
Status: promptkit.ValidationFailed,
Mode: promptkit.ValidationJSONSchema,
Errors: []string{"first error", "second error"},
SchemaPath: "schemas/result.json",
RepairAttempts: 2,
IsValid: false,
},
PromptID: "prompt.result",
PromptVersion: "3.2.1",
PromptHash: "result-prompt-hash",
SessionID: "session-result",
RenderedPromptHash: "result-rendered-hash",
SelectedProfileID: "profile-result",
SelectedBackendID: "backend-result",
ModelName: "model-result",
Endpoint: "https://result.example/v1",
EffectiveModelParams: jsonContractExecutionTarget(),
InputHashes: map[string]string{"input": "input-hash"},
Usage: promptkit.TokenUsage{
PromptTokens: 101,
CompletionTokens: 202,
TotalTokens: 303,
CachedTokens: 44,
CacheWriteTokens: 55,
},
StartTime: start,
EndTime: start.Add(1750 * time.Millisecond),
Duration: 1750 * time.Millisecond,
}
}
func jsonContractExecutionTarget() promptkit.ExecutionTarget {
return promptkit.ExecutionTarget{
BackendID: "backend-target",
Endpoint: "https://target.example/v1",
Model: "model-target",
Temperature: 0.75,
MaxTokens: 321,
TopP: 0.875,
TimeoutSeconds: 43,
ServiceTier: "priority",
ReasoningEffort: "high",
APIKeyEnv: "PROMPTKIT_JSON_CONTRACT_KEY",
ExtraParams: map[string]any{
"enabled": true,
"weight": float64(1.25),
"nested": map[string]any{"name": "value"},
},
}
}
func durationPayload(milliseconds int64) []byte {
return []byte(`{"duration_ms":` + strconv.FormatInt(milliseconds, 10) + `}`)
}
func decodeJSONObject(t *testing.T, payload []byte) map[string]json.RawMessage {
t.Helper()
var object map[string]json.RawMessage
if err := json.Unmarshal(payload, &object); err != nil {
t.Fatalf("decode JSON object: %v", err)
}
return object
}
func assertJSONFields(t *testing.T, object map[string]json.RawMessage, fields ...string) {
t.Helper()
want := make(map[string]struct{}, len(fields))
for _, field := range fields {
want[field] = struct{}{}
}
for field := range object {
if _, ok := want[field]; !ok {
t.Errorf("unexpected JSON field %q", field)
}
}
for field := range want {
if _, ok := object[field]; !ok {
t.Errorf("missing JSON field %q", field)
}
}
}

View File

@@ -0,0 +1,96 @@
package promptkit
import (
"context"
"reflect"
"strconv"
"sync"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
func TestPublicLLMClientAdapterGivesClientOwnedNestedValues(t *testing.T) {
source := adapterOwnershipRequest()
want := adapterOwnershipRequest()
client := &retainingMutatingLLMClient{}
adapter := publicLLMClientAdapter{client: client}
if _, err := adapter.Generate(context.Background(), source); err != nil {
t.Fatalf("generate: %v", err)
}
if !reflect.DeepEqual(source, want) {
t.Fatalf("client mutation changed prepared source:\ngot %#v\nwant %#v", source, want)
}
var mutations sync.WaitGroup
mutations.Add(1)
go func() {
defer mutations.Done()
for i := 0; i < 10_000; i++ {
mutateGenerateRequest(&client.retained, strconv.Itoa(i))
}
}()
for i := 0; i < 10_000; i++ {
laterRequest := fromDomainGenerateRequest(source)
if laterRequest.Prompt.Messages[0].Content != "source-message" ||
laterRequest.Prompt.Messages[0].CacheControl.TTL != "source-ttl" ||
laterRequest.Target.ExtraParams["nested"].([]any)[0] != "source-extra" ||
laterRequest.StructuredOutput.JSONSchema.Schema.(map[string]any)["enum"].([]any)[0] != "source-schema" {
t.Fatal("retained client mutation reached a later execution request")
}
}
mutations.Wait()
if !reflect.DeepEqual(source, want) {
t.Fatalf("retained client mutation changed prepared source:\ngot %#v\nwant %#v", source, want)
}
}
type retainingMutatingLLMClient struct {
retained GenerateRequest
}
func (c *retainingMutatingLLMClient) Generate(_ context.Context, request GenerateRequest) (*GenerateResponse, error) {
c.retained = request
mutateGenerateRequest(&c.retained, "client-mutation")
return &GenerateResponse{Content: "generated"}, nil
}
func adapterOwnershipRequest() domain.GenerateRequest {
return domain.GenerateRequest{
Prompt: domain.RenderedPrompt{
SessionID: "source-session",
Messages: []domain.RenderedMessage{{
Role: "user",
Content: "source-message",
CacheControl: &domain.CacheControl{
Type: domain.CacheControlEphemeral,
TTL: "source-ttl",
},
}},
},
Target: domain.ExecutionTarget{
Model: "source-model",
APIKey: "source-api-key",
ExtraParams: map[string]any{
"nested": []any{"source-extra"},
},
},
StructuredOutput: &domain.StructuredOutputSpec{
Type: domain.StructuredOutputJSONSchema,
JSONSchema: &domain.StructuredOutputJSONSpec{
Name: "source-schema-name",
Strict: true,
Schema: map[string]any{"enum": []any{"source-schema"}},
},
},
}
}
func mutateGenerateRequest(request *GenerateRequest, value string) {
request.Prompt.Messages[0].Content = value
request.Prompt.Messages[0].CacheControl.TTL = value
request.Target.ExtraParams["nested"].([]any)[0] = value
request.StructuredOutput.JSONSchema.Schema.(map[string]any)["enum"].([]any)[0] = value
}

View File

@@ -0,0 +1,73 @@
package promptkit_test
import (
"context"
"errors"
"testing"
"gitea.maximumdirect.net/eric/promptkit"
)
func TestPreparationRejectsInvalidOutputContractWithPublicError(t *testing.T) {
engine, err := promptkit.NewEngine(
promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile",
Endpoint: "http://example.test/v1",
Model: "model",
}),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
for _, tc := range []struct {
name string
contract promptkit.OutputContract
}{
{
name: "unsupported format",
contract: promptkit.OutputContract{
Format: promptkit.OutputFormat("binary"),
ValidationMode: promptkit.ValidationNone,
},
},
{
name: "repair attempts above maximum",
contract: promptkit.OutputContract{
Format: promptkit.FormatText,
ValidationMode: promptkit.ValidationBasic,
RepairAttempts: 4,
},
},
{
name: "none validation with repair attempts",
contract: promptkit.OutputContract{
Format: promptkit.FormatText,
ValidationMode: promptkit.ValidationNone,
RepairAttempts: 1,
},
},
} {
t.Run(tc.name, func(t *testing.T) {
req := promptkit.RunRequest{PromptID: "prompt", Validation: &tc.contract}
prepared, err := engine.Prepare(context.Background(), req)
if prepared != nil {
t.Fatalf("expected no partial prepared run, got %+v", prepared)
}
if !errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("prepare error = %v, want ErrInvalidRequest", err)
}
preparedExecution, err := engine.PrepareExecution(context.Background(), req)
if preparedExecution != nil {
t.Fatalf("expected no partial prepared execution, got %+v", preparedExecution)
}
if !errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("prepare execution error = %v, want ErrInvalidRequest", err)
}
})
}
}

View File

@@ -44,12 +44,12 @@ func (p *PreparedExecution) Discard() {
// String returns a constant representation that exposes no retained request, // String returns a constant representation that exposes no retained request,
// rendered content, or credential data. // rendered content, or credential data.
func (p *PreparedExecution) String() string { func (p PreparedExecution) String() string {
return preparedExecutionString return preparedExecutionString
} }
// GoString returns a constant Go-syntax representation that exposes no // GoString returns a constant Go-syntax representation that exposes no
// retained request, rendered content, or credential data. // retained request, rendered content, or credential data.
func (p *PreparedExecution) GoString() string { func (p PreparedExecution) GoString() string {
return preparedExecutionString return preparedExecutionString
} }

View File

@@ -5,7 +5,6 @@ import (
"encoding/json" "encoding/json"
"errors" "errors"
"fmt" "fmt"
"os"
"reflect" "reflect"
"strings" "strings"
"sync" "sync"
@@ -257,6 +256,50 @@ func TestPreparedExecutionLifecycleAndEngineBinding(t *testing.T) {
} }
} }
func TestPreparedExecutionRepairsEmptyBasicOutput(t *testing.T) {
client := &preparedRecordingClient{responses: []*promptkit.GenerateResponse{
{
Content: "",
Usage: promptkit.TokenUsage{PromptTokens: 3, CompletionTokens: 5, TotalTokens: 8},
},
{
Content: "Corrected summary.",
Usage: promptkit.TokenUsage{PromptTokens: 7, CompletionTokens: 11, TotalTokens: 18},
},
}}
engine := newPreparedContractEngine(t, client, "Summarize the source.")
prepared, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{
PromptID: "prepared",
Validation: &promptkit.OutputContract{
Format: promptkit.FormatMarkdown,
ValidationMode: promptkit.ValidationBasic,
RepairAttempts: 1,
},
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
details := prepared.Details()
result, err := engine.RunPrepared(context.Background(), prepared)
if err != nil {
t.Fatalf("run prepared: %v", err)
}
if result.RawOutput != "Corrected summary." || result.Validation.Status != promptkit.ValidationPassed ||
result.Validation.RepairAttempts != 1 || result.Usage != (promptkit.TokenUsage{PromptTokens: 10, CompletionTokens: 16, TotalTokens: 26}) {
t.Fatalf("repaired result = %+v", result)
}
requests := client.snapshot()
if len(requests) != 2 || len(requests[1].Prompt.Messages) != len(details.Messages)+1 ||
!reflect.DeepEqual(requests[1].Prompt.Messages[:len(details.Messages)], details.Messages) ||
requests[1].Prompt.Messages[len(requests[1].Prompt.Messages)-1].Role != "user" {
t.Fatalf("prepared repair requests = %#v", requests)
}
if _, err := engine.RunPrepared(context.Background(), prepared); !errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("second RunPrepared error = %v, want ErrInvalidRequest", err)
}
}
func TestPreparedExecutionConcurrentClaimAllowsOneGeneration(t *testing.T) { func TestPreparedExecutionConcurrentClaimAllowsOneGeneration(t *testing.T) {
release := make(chan struct{}) release := make(chan struct{})
client := &preparedRecordingClient{ client := &preparedRecordingClient{
@@ -409,14 +452,35 @@ func TestPreparedExecutionDiscardAndFormattingDoNotExposePrivateState(t *testing
t.Fatalf("prepare execution: %v", err) t.Fatalf("prepare execution: %v", err)
} }
formattedValues := []string{ copied := *prepared
fmt.Sprint(prepared), zeroValue := promptkit.PreparedExecution{}
fmt.Sprintf("%+v", prepared), var nilHandle *promptkit.PreparedExecution
fmt.Sprintf("%#v", prepared), for name, value := range map[string]any{
} "original pointer": prepared,
for _, formatted := range formattedValues { "copied value": copied,
"zero value": zeroValue,
"zero pointer": &zeroValue,
} {
for format, formatted := range map[string]string{
"String": fmt.Sprintf("%s", value),
"GoString": fmt.Sprintf("%#v", value),
"v": fmt.Sprintf("%v", value),
"+v": fmt.Sprintf("%+v", value),
} {
if formatted != "promptkit.PreparedExecution{opaque}" { if formatted != "promptkit.PreparedExecution{opaque}" {
t.Fatalf("unexpected opaque formatting: %q", formatted) t.Fatalf("%s %s formatting = %q, want opaque representation", name, format, formatted)
}
assertPreparedPrivateValuesAbsent(t, formatted, directCredential, renderedContent)
}
}
for format, formatted := range map[string]string{
"String": fmt.Sprintf("%s", nilHandle),
"GoString": fmt.Sprintf("%#v", nilHandle),
"v": fmt.Sprintf("%v", nilHandle),
"+v": fmt.Sprintf("%+v", nilHandle),
} {
if formatted != "<nil>" {
t.Fatalf("nil pointer %s formatting = %q, want <nil>", format, formatted)
} }
assertPreparedPrivateValuesAbsent(t, formatted, directCredential, renderedContent) assertPreparedPrivateValuesAbsent(t, formatted, directCredential, renderedContent)
} }
@@ -433,27 +497,9 @@ func TestPreparedExecutionDiscardAndFormattingDoNotExposePrivateState(t *testing
} }
assertPreparedPrivateValuesAbsent(t, string(detailsJSON), directCredential) assertPreparedPrivateValuesAbsent(t, string(detailsJSON), directCredential)
prepared.Discard() executionResult, err := engine.RunPrepared(context.Background(), &copied)
prepared.Discard()
result, lifecycleErr := engine.RunPrepared(context.Background(), prepared)
if result != nil || !errors.Is(lifecycleErr, promptkit.ErrInvalidRequest) {
t.Fatalf("discarded execution result=(%+v, %v), want ErrInvalidRequest", result, lifecycleErr)
}
assertPreparedPrivateValuesAbsent(t, lifecycleErr.Error(), directCredential, renderedContent)
if !reflect.DeepEqual(prepared.Details(), detailsBefore) {
t.Fatal("details changed after discard")
}
executed, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{
PromptID: "prepared",
APIKey: directCredential,
})
if err != nil { if err != nil {
t.Fatalf("prepare execution for request inspection: %v", err) t.Fatalf("run copied execution after formatting: %v", err)
}
executionResult, err := engine.RunPrepared(context.Background(), executed)
if err != nil {
t.Fatalf("run execution for request inspection: %v", err)
} }
requests := client.snapshot() requests := client.snapshot()
if len(requests) != 1 || requests[0].APIKey != directCredential { if len(requests) != 1 || requests[0].APIKey != directCredential {
@@ -478,16 +524,34 @@ func TestPreparedExecutionDiscardAndFormattingDoNotExposePrivateState(t *testing
} }
assertPreparedPrivateValuesAbsent(t, string(resultJSON), directCredential) assertPreparedPrivateValuesAbsent(t, string(resultJSON), directCredential)
var nilHandle *promptkit.PreparedExecution
nilHandle.Discard() nilHandle.Discard()
if !reflect.DeepEqual(nilHandle.Details(), promptkit.PreparedRun{}) { if !reflect.DeepEqual(nilHandle.Details(), promptkit.PreparedRun{}) {
t.Fatalf("nil handle details=%+v, want zero value", nilHandle.Details()) t.Fatalf("nil handle details=%+v, want zero value", nilHandle.Details())
} }
zeroHandle := &promptkit.PreparedExecution{} zeroHandle := &zeroValue
zeroHandle.Discard() zeroHandle.Discard()
if !reflect.DeepEqual(zeroHandle.Details(), promptkit.PreparedRun{}) { if !reflect.DeepEqual(zeroHandle.Details(), promptkit.PreparedRun{}) {
t.Fatalf("zero handle details=%+v, want zero value", zeroHandle.Details()) t.Fatalf("zero handle details=%+v, want zero value", zeroHandle.Details())
} }
discarded, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{
PromptID: "prepared",
APIKey: directCredential,
})
if err != nil {
t.Fatalf("prepare execution for discard: %v", err)
}
discardedDetails := discarded.Details()
discarded.Discard()
discarded.Discard()
result, lifecycleErr := engine.RunPrepared(context.Background(), discarded)
if result != nil || !errors.Is(lifecycleErr, promptkit.ErrInvalidRequest) {
t.Fatalf("discarded execution result=(%+v, %v), want ErrInvalidRequest", result, lifecycleErr)
}
assertPreparedPrivateValuesAbsent(t, lifecycleErr.Error(), directCredential, renderedContent)
if !reflect.DeepEqual(discarded.Details(), discardedDetails) {
t.Fatal("details changed after discard")
}
} }
func TestPreparedExecutionCredentialCapacityAndTimingBoundaries(t *testing.T) { func TestPreparedExecutionCredentialCapacityAndTimingBoundaries(t *testing.T) {
@@ -503,19 +567,25 @@ func TestPreparedExecutionCredentialCapacityAndTimingBoundaries(t *testing.T) {
engine, err := promptkit.NewEngine( engine, err := promptkit.NewEngine(
promptkit.Config{}, promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prepared", "profile", "content"), "."), promptkit.WithPromptFS(contractPromptFS("prepared", "profile", "content"), "."),
promptkit.WithProfileFS(preparedCredentialProfileSource(environmentName), "."), promptkit.WithProfiles(promptkit.Profile{
ID: "profile",
Endpoint: "http://example.test/v1",
Model: "model",
APIKeyRequired: true,
}),
promptkit.WithLLMClient(client), promptkit.WithLLMClient(client),
) )
if err != nil { if err != nil {
t.Fatalf("construct credential engine: %v", err) t.Fatalf("construct credential engine: %v", err)
} }
prepared, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{PromptID: "prepared"}) prepared, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{
PromptID: "prepared",
Execution: &promptkit.ExecutionTargetOverride{APIKeyEnv: environmentName},
})
if err != nil { if err != nil {
t.Fatalf("prepare credential execution: %v", err) t.Fatalf("prepare credential execution: %v", err)
} }
if err := os.Unsetenv(environmentName); err != nil { t.Setenv(environmentName, "")
t.Fatalf("unset credential environment: %v", err)
}
result, err := engine.RunPrepared(context.Background(), prepared) result, err := engine.RunPrepared(context.Background(), prepared)
if result != nil || if result != nil ||
@@ -647,6 +717,7 @@ func (r *mutablePreparedArtifactReader) callCount() int {
type preparedRecordingClient struct { type preparedRecordingClient struct {
mu sync.Mutex mu sync.Mutex
response *promptkit.GenerateResponse response *promptkit.GenerateResponse
responses []*promptkit.GenerateResponse
err error err error
requests []promptkit.GenerateRequest requests []promptkit.GenerateRequest
started chan struct{} started chan struct{}
@@ -674,6 +745,12 @@ func (c *preparedRecordingClient) Generate(
if c.err != nil { if c.err != nil {
return nil, c.err return nil, c.err
} }
if len(c.responses) > 0 {
if index := len(c.requests) - 1; index < len(c.responses) {
return c.responses[index], nil
}
return nil, fmt.Errorf("no response configured for generation %d", len(c.requests))
}
return c.response, nil return c.response, nil
} }
@@ -734,16 +811,6 @@ model: ` + model + `
} }
} }
func preparedCredentialProfileSource(environmentName string) fstest.MapFS {
return fstest.MapFS{
"profile.yaml": &fstest.MapFile{Data: []byte(`id: profile
endpoint: http://example.test/v1
model: model
api_key_env: ` + environmentName + `
`)},
}
}
func preparedSchemaSource() fstest.MapFS { func preparedSchemaSource() fstest.MapFS {
return fstest.MapFS{ return fstest.MapFS{
"schema.json": &fstest.MapFile{Data: []byte(`{ "schema.json": &fstest.MapFile{Data: []byte(`{

View File

@@ -0,0 +1,243 @@
package promptkit_test
import (
"context"
"errors"
"io/fs"
"strings"
"sync"
"testing"
"testing/fstest"
"gitea.maximumdirect.net/eric/promptkit"
)
func TestProfileInheritanceBuiltInAliasWorkflow(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: frameworkPromptDir,
SchemaDir: frameworkSchemaDir,
}, promptkit.WithProfiles(promptkit.Profile{
ID: "weather-light",
BaseProfileID: "deepseek-4-flash",
ReasoningEffort: "high",
TimeoutSeconds: 120,
}))
if err != nil {
t.Fatalf("construct alias engine: %v", err)
}
base, err := engine.InspectProfile(context.Background(), "deepseek-4-flash")
if err != nil {
t.Fatalf("inspect base: %v", err)
}
child, err := engine.InspectProfile(context.Background(), "weather-light")
if err != nil {
t.Fatalf("inspect alias: %v", err)
}
if child.ProfileID != "weather-light" ||
child.EffectiveModelParams.BackendID != base.EffectiveModelParams.BackendID ||
child.EffectiveModelParams.Model != base.EffectiveModelParams.Model ||
child.EffectiveModelParams.ReasoningEffort != "high" ||
child.EffectiveModelParams.TimeoutSeconds != 120 {
t.Fatalf("alias inspection = %+v, base = %+v", child, base)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: frameworkMarkdownSummaryPromptID,
ProfileID: "weather-light",
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."),
},
})
if err != nil {
t.Fatalf("prepare alias: %v", err)
}
if prepared.SelectedProfileID != "weather-light" {
t.Fatalf("SelectedProfileID = %q", prepared.SelectedProfileID)
}
timeout := 15
reasoning := "low"
overridden, err := engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: frameworkMarkdownSummaryPromptID,
ProfileID: "weather-light",
Execution: &promptkit.ExecutionTargetOverride{
TimeoutSeconds: &timeout,
ReasoningEffort: &reasoning,
},
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."),
},
})
if err != nil {
t.Fatalf("prepare override: %v", err)
}
if overridden.EffectiveModelParams.TimeoutSeconds != timeout ||
overridden.EffectiveModelParams.ReasoningEffort != reasoning {
t.Fatalf("runtime override target = %+v", overridden.EffectiveModelParams)
}
}
func TestProfileInheritanceYAMLAliasOfBuiltIn(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(fstest.MapFS{}, "."),
promptkit.WithProfileFS(fstest.MapFS{
"alias.yaml": &fstest.MapFile{Data: []byte("id: yaml-alias\nbase_profile: deepseek-4-flash\n")},
}, "."),
)
if err != nil {
t.Fatalf("construct YAML alias engine: %v", err)
}
inspection, err := engine.InspectProfile(context.Background(), "yaml-alias")
if err != nil {
t.Fatalf("inspect YAML alias: %v", err)
}
if inspection.ProfileID != "yaml-alias" || inspection.EffectiveModelParams.Model == "" {
t.Fatalf("YAML alias inspection = %+v", inspection)
}
}
func TestProfileInheritanceRetainsRequiredCredentialBehavior(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(fstest.MapFS{}, "."),
promptkit.WithBackend(promptkit.Backend{
ID: "credential-backend",
Endpoint: "https://credential.example/v1",
APIKeyEnv: "OPTIONAL_BACKEND_KEY",
}),
promptkit.WithProfiles(
promptkit.Profile{
ID: "credential-base",
BackendID: "credential-backend",
Model: "model",
APIKeyRequired: true,
},
promptkit.Profile{ID: "credential-child", BaseProfileID: "credential-base"},
),
)
if err != nil {
t.Fatalf("construct credential inheritance engine: %v", err)
}
inspection, err := engine.InspectProfile(context.Background(), "credential-child")
if err != nil {
t.Fatalf("inspect credential child: %v", err)
}
if !inspection.APIKeyRequired || inspection.EffectiveModelParams.APIKeyEnv != "" {
t.Fatalf("credential inspection = %+v", inspection)
}
}
func TestProfileInheritancePreservesPublicErrorIdentities(t *testing.T) {
newEngine := func(t *testing.T, profiles fs.FS) *promptkit.Engine {
t.Helper()
options := []promptkit.Option{promptkit.WithPromptFS(fstest.MapFS{}, ".")}
if profiles != nil {
options = append(options, promptkit.WithProfileFS(profiles, "."))
}
engine, err := promptkit.NewEngine(promptkit.Config{}, options...)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
return engine
}
tests := []struct {
name string
profiles fs.FS
profile string
contains []string
want error
wantNot error
}{
{
name: "missing selected profile",
profile: "missing",
want: promptkit.ErrProfileNotFound,
wantNot: promptkit.ErrProfileLoad,
},
{
name: "missing base",
profiles: fstest.MapFS{
"child.yaml": &fstest.MapFile{Data: []byte("id: child\nbase_profile: missing\n")},
},
profile: "child",
contains: []string{"child", "missing"},
want: promptkit.ErrProfileLoad,
wantNot: promptkit.ErrProfileNotFound,
},
{
name: "cycle",
profiles: fstest.MapFS{
"a.yaml": &fstest.MapFile{Data: []byte("id: a\nbase_profile: b\n")},
"b.yaml": &fstest.MapFile{Data: []byte("id: b\nbase_profile: a\n")},
},
profile: "a",
contains: []string{"a", "b"},
want: promptkit.ErrProfileLoad,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
result, err := newEngine(t, tc.profiles).InspectProfile(context.Background(), tc.profile)
if result != nil || !errors.Is(err, tc.want) || (tc.wantNot != nil && errors.Is(err, tc.wantNot)) {
t.Fatalf("inspection=(%+v, %v), want %v without %v", result, err, tc.want, tc.wantNot)
}
for _, fragment := range tc.contains {
if !strings.Contains(err.Error(), fragment) {
t.Fatalf("error = %v, want %q", err, fragment)
}
}
})
}
}
func TestProfileInheritanceFreezesPreparedExecution(t *testing.T) {
profiles := &mutableInheritanceProfileFS{files: fstest.MapFS{
"child.yaml": &fstest.MapFile{Data: []byte("id: child\nbase_profile: base\n")},
"base.yaml": &fstest.MapFile{Data: []byte("id: base\nendpoint: https://base.example/v1\nmodel: first-model\n")},
}}
client := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "ok"}}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prepared", "child", "content"), "."),
promptkit.WithProfileFS(profiles, "."),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
prepared, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{PromptID: "prepared"})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
profiles.set("base.yaml", "id: base\nendpoint: https://base.example/v1\nmodel: second-model\n")
result, err := engine.RunPrepared(context.Background(), prepared)
if err != nil || result == nil || len(client.requests) != 1 || client.requests[0].Target.Model != "first-model" {
t.Fatalf("prepared execution=(%+v, %v), requests=%+v", result, err, client.requests)
}
inspection, err := engine.InspectProfile(context.Background(), "child")
if err != nil || inspection.EffectiveModelParams.Model != "second-model" {
t.Fatalf("fresh inspection=(%+v, %v)", inspection, err)
}
}
type mutableInheritanceProfileFS struct {
mu sync.RWMutex
files fstest.MapFS
}
func (f *mutableInheritanceProfileFS) Open(name string) (fs.File, error) {
f.mu.RLock()
defer f.mu.RUnlock()
return f.files.Open(name)
}
func (f *mutableInheritanceProfileFS) set(name, content string) {
f.mu.Lock()
defer f.mu.Unlock()
f.files[name] = &fstest.MapFile{Data: []byte(content)}
}

View File

@@ -2,17 +2,16 @@ package promptkit
import ( import (
"context" "context"
"errors"
"fmt" "fmt"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue" "gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
) )
// OpenAICompatibleProfile returns an ordinary in-memory Profile for an // OpenAICompatibleProfile returns an in-memory Profile for an OpenAI-compatible
// OpenAI-compatible chat-completions endpoint. // chat-completions endpoint. A non-blank BaseProfileID permits its target
// fields to be inherited when the profile is selected or inspected.
// //
// It does not register global state, maintain a model catalog, or resolve // It does not register global state, maintain a model catalog, or resolve
// credentials. If APIKeyRequired is true, callers satisfy it with // credentials. If APIKeyRequired is true, callers satisfy it with
@@ -25,6 +24,7 @@ import (
func OpenAICompatibleProfile(cfg OpenAICompatibleProfileConfig) Profile { func OpenAICompatibleProfile(cfg OpenAICompatibleProfileConfig) Profile {
return Profile{ return Profile{
ID: cfg.ID, ID: cfg.ID,
BaseProfileID: cfg.BaseProfileID,
BackendID: cfg.BackendID, BackendID: cfg.BackendID,
Endpoint: cfg.Endpoint, Endpoint: cfg.Endpoint,
Model: cfg.Model, Model: cfg.Model,
@@ -87,8 +87,9 @@ func toDomainProfile(publicProfile Profile) (domain.ExecutionProfile, error) {
return domain.ExecutionProfile{}, err return domain.ExecutionProfile{}, err
} }
prof := domain.ExecutionProfile{ prof := domain.ExecutionProfile{
ID: strings.TrimSpace(publicProfile.ID), ID: publicProfile.ID,
BackendID: strings.TrimSpace(publicProfile.BackendID), BaseProfileID: publicProfile.BaseProfileID,
BackendID: publicProfile.BackendID,
Endpoint: publicProfile.Endpoint, Endpoint: publicProfile.Endpoint,
Model: publicProfile.Model, Model: publicProfile.Model,
Temperature: publicProfile.Temperature, Temperature: publicProfile.Temperature,
@@ -100,33 +101,8 @@ func toDomainProfile(publicProfile Profile) (domain.ExecutionProfile, error) {
APIKeyRequired: publicProfile.APIKeyRequired, APIKeyRequired: publicProfile.APIKeyRequired,
ExtraParams: extraParams, ExtraParams: extraParams,
} }
if err := validatePublicProfile(prof); err != nil { if err := profile.NormalizeAndValidateDefinition(&prof); err != nil {
return domain.ExecutionProfile{}, err return domain.ExecutionProfile{}, err
} }
return prof, nil return prof, nil
} }
func validatePublicProfile(prof domain.ExecutionProfile) error {
if strings.TrimSpace(prof.ID) == "" {
return errors.New("id is required")
}
if strings.TrimSpace(prof.BackendID) == "" && strings.TrimSpace(prof.Endpoint) == "" {
return errors.New("backend or endpoint is required")
}
if strings.TrimSpace(prof.Model) == "" {
return errors.New("model is required")
}
if prof.Temperature < 0 || prof.Temperature > 2 {
return errors.New("temperature must be between 0 and 2")
}
if prof.MaxTokens < 0 {
return errors.New("max_tokens must be greater than or equal to 0")
}
if prof.TopP < 0 || prof.TopP > 1 {
return errors.New("top_p must be between 0 and 1")
}
if prof.TimeoutSeconds < 0 {
return errors.New("timeout_seconds must be greater than or equal to 0")
}
return nil
}

View File

@@ -12,7 +12,6 @@ import (
"sync/atomic" "sync/atomic"
"testing" "testing"
"testing/fstest" "testing/fstest"
"time"
"gitea.maximumdirect.net/eric/promptkit" "gitea.maximumdirect.net/eric/promptkit"
) )
@@ -78,6 +77,41 @@ api_key_env: PROMPTKIT_INSPECTION_ABSENT_KEY
} }
} }
func TestFileProfileNormalizedIDMatchesInspectionAndPreparation(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "normalized-profile", "message"), "."),
promptkit.WithProfileFS(fstest.MapFS{
"profile.yaml": &fstest.MapFile{Data: []byte(`
id: " normalized-profile "
endpoint: http://profile.example/v1
model: normalized-model
`)},
}, "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
inspection, err := engine.InspectProfile(context.Background(), " normalized-profile ")
if err != nil {
t.Fatalf("inspect normalized profile: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: "prompt",
ProfileID: " normalized-profile ",
})
if err != nil {
t.Fatalf("prepare with normalized profile: %v", err)
}
if inspection.ProfileID != "normalized-profile" ||
prepared.SelectedProfileID != inspection.ProfileID ||
inspection.EffectiveModelParams.Model != "normalized-model" ||
prepared.EffectiveModelParams.Model != inspection.EffectiveModelParams.Model {
t.Fatalf("inspection=%#v prepared=%#v", inspection, prepared)
}
}
func TestInspectProfilePreservesPublicErrorIdentities(t *testing.T) { func TestInspectProfilePreservesPublicErrorIdentities(t *testing.T) {
newEngine := func(t *testing.T, options ...promptkit.Option) *promptkit.Engine { newEngine := func(t *testing.T, options ...promptkit.Option) *promptkit.Engine {
t.Helper() t.Helper()
@@ -110,7 +144,7 @@ func TestInspectProfilePreservesPublicErrorIdentities(t *testing.T) {
} }
malformed := newEngine(t, promptkit.WithProfileFS(fstest.MapFS{ malformed := newEngine(t, promptkit.WithProfileFS(fstest.MapFS{
"broken.yaml": &fstest.MapFile{Data: []byte("id: broken\nendpoint: http://broken.example/v1\nmodel: model\nunknown: value\n")}, "broken.yaml": &fstest.MapFile{Data: []byte("id: broken\nendpoint: http://broken.example/v1\nmodel: model\nextra_params:\n invalid: .nan\n")},
}, ".")) }, "."))
if result, err := malformed.InspectProfile(context.Background(), "broken"); result != nil || if result, err := malformed.InspectProfile(context.Background(), "broken"); result != nil ||
!errors.Is(err, promptkit.ErrProfileLoad) { !errors.Is(err, promptkit.ErrProfileLoad) {
@@ -391,18 +425,6 @@ output:
} }
} }
func TestPreparedRunJSONOmitsZeroTimingValues(t *testing.T) {
payload, err := json.Marshal(promptkit.PreparedRun{})
if err != nil {
t.Fatalf("marshal prepared run: %v", err)
}
for _, field := range []string{"start_time", "end_time", "duration_ms"} {
if strings.Contains(string(payload), `"`+field+`"`) {
t.Fatalf("expected zero %s to be omitted, got %s", field, payload)
}
}
}
func TestBackendIdentityJSONNamesAndOmission(t *testing.T) { func TestBackendIdentityJSONNamesAndOmission(t *testing.T) {
t.Run("execution target round trip", func(t *testing.T) { t.Run("execution target round trip", func(t *testing.T) {
value := promptkit.ExecutionTarget{BackendID: promptkit.BackendOpenRouter} value := promptkit.ExecutionTarget{BackendID: promptkit.BackendOpenRouter}
@@ -757,13 +779,17 @@ func TestBackendRegistrationRejectsInvalidAndDuplicateDefinitions(t *testing.T)
{name: "reserved extra parameter", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: map[string]any{"model": "override"}}}}, {name: "reserved extra parameter", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: map[string]any{"model": "override"}}}},
{name: "cyclic extra parameter", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: cycle}}}, {name: "cyclic extra parameter", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: cycle}}},
{name: "malformed JSON number", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: map[string]any{"value": json.Number("01")}}}}, {name: "malformed JSON number", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: map[string]any{"value": json.Number("01")}}}},
{name: "excessively deep extra parameter", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: map[string]any{"value": excessivelyDeepJSONValue()}}}},
{name: "duplicate consumer id", backends: []promptkit.Backend{ {name: "duplicate consumer id", backends: []promptkit.Backend{
{ID: " custom ", Endpoint: "http://one.example/v1"}, {ID: " custom ", Endpoint: "http://one.example/v1"},
{ID: "custom", Endpoint: "http://two.example/v1"}, {ID: "custom", Endpoint: "http://two.example/v1"},
}}, }},
{name: "reserved built-in id", backends: []promptkit.Backend{{ {name: "reserved OpenRouter ID", backends: []promptkit.Backend{{
ID: promptkit.BackendOpenRouter, Endpoint: "http://replacement.example/v1", ID: promptkit.BackendOpenRouter, Endpoint: "http://replacement.example/v1",
}}}, }}},
{name: "reserved Rakestrawhome ID", backends: []promptkit.Backend{{
ID: promptkit.BackendRakestrawHome, Endpoint: "http://replacement.example/v1",
}}},
} }
for _, tt := range tests { for _, tt := range tests {
@@ -818,93 +844,7 @@ func TestBackendExtraParamsAreDeeplyCopiedAtConstructionAndLookup(t *testing.T)
} }
} }
func TestPreparedRunJSONTimingRoundTrips(t *testing.T) { func TestEngineValidationWithZeroRepairBudgetIsSinglePass(t *testing.T) {
start := time.Date(2026, time.July, 29, 12, 0, 0, 0, time.UTC)
prepared := promptkit.PreparedRun{
PromptID: "prompt",
StartTime: start,
EndTime: start.Add(1250 * time.Millisecond),
DurationMS: 1250,
}
payload, err := json.Marshal(prepared)
if err != nil {
t.Fatalf("marshal prepared run: %v", err)
}
var decoded promptkit.PreparedRun
if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal prepared run: %v", err)
}
if decoded.DurationMS != prepared.DurationMS ||
!decoded.StartTime.Equal(prepared.StartTime) ||
!decoded.EndTime.Equal(prepared.EndTime) {
t.Fatalf("timing values did not round trip: got %#v, want %#v", decoded, prepared)
}
}
func TestRunResultJSONUsesMillisecondsAndRoundTrips(t *testing.T) {
start := time.Date(2026, time.July, 29, 12, 0, 0, 0, time.UTC)
result := promptkit.RunResult{
RunID: "opaque-run-id",
Artifact: promptkit.Artifact{Name: "output", ContentType: "text/plain", Body: []byte("ok")},
SessionID: "session-123",
StartTime: start,
EndTime: start.Add(1500 * time.Millisecond),
Duration: 1500 * time.Millisecond,
}
payload, err := json.Marshal(result)
if err != nil {
t.Fatalf("marshal run result: %v", err)
}
var object map[string]any
if err := json.Unmarshal(payload, &object); err != nil {
t.Fatalf("decode run result JSON: %v", err)
}
if got := object["duration_ms"]; got != float64(1500) {
t.Fatalf("expected duration_ms=1500, got %#v in %s", got, payload)
}
if _, exists := object["duration"]; exists {
t.Fatalf("unexpected nanosecond duration field in %s", payload)
}
if got := object["session_id"]; got != result.SessionID {
t.Fatalf("expected session_id=%q, got %#v in %s", result.SessionID, got, payload)
}
artifact, ok := object["artifact"].(map[string]any)
if !ok || artifact["content_type"] != "text/plain" {
t.Fatalf("expected stable artifact JSON fields, got %#v", object["artifact"])
}
var decoded promptkit.RunResult
if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal run result: %v", err)
}
if decoded.SessionID != result.SessionID ||
decoded.Duration != result.Duration ||
!decoded.StartTime.Equal(result.StartTime) ||
!decoded.EndTime.Equal(result.EndTime) {
t.Fatalf("timing values did not round trip: got %#v, want %#v", decoded, result)
}
payload, err = json.Marshal(promptkit.RunResult{})
if err != nil {
t.Fatalf("marshal zero run result: %v", err)
}
for _, field := range []string{"session_id", "start_time", "end_time", "duration_ms"} {
if strings.Contains(string(payload), `"`+field+`"`) {
t.Fatalf("expected zero %s to be omitted, got %s", field, payload)
}
}
var decodedEmpty promptkit.RunResult
if err := json.Unmarshal(payload, &decodedEmpty); err != nil {
t.Fatalf("unmarshal run result without session_id: %v", err)
}
if decodedEmpty.SessionID != "" {
t.Fatalf("expected absent session_id to decode empty, got %q", decodedEmpty.SessionID)
}
}
func TestEngineValidationIsSinglePass(t *testing.T) {
client := &fakeLLMClient{ client := &fakeLLMClient{
response: &promptkit.GenerateResponse{Content: "not-json"}, response: &promptkit.GenerateResponse{Content: "not-json"},
} }
@@ -919,7 +859,7 @@ func TestEngineValidationIsSinglePass(t *testing.T) {
Validation: &promptkit.OutputContract{ Validation: &promptkit.OutputContract{
Format: promptkit.FormatJSON, Format: promptkit.FormatJSON,
ValidationMode: promptkit.ValidationJSON, ValidationMode: promptkit.ValidationJSON,
RepairAttempts: 3, RepairAttempts: 0,
}, },
}) })
if err != nil { if err != nil {
@@ -934,6 +874,87 @@ func TestEngineValidationIsSinglePass(t *testing.T) {
} }
} }
func TestEngineRunRepairsJSONSchemaOutput(t *testing.T) {
client := &fakeLLMClient{responses: []*promptkit.GenerateResponse{
{
Content: "{}",
Usage: promptkit.TokenUsage{PromptTokens: 2, CompletionTokens: 3, TotalTokens: 5, CachedTokens: 7, CacheWriteTokens: 11},
},
{
Content: `{"events":[{"title":"Repaired event"}]}`,
Usage: promptkit.TokenUsage{PromptTokens: 13, CompletionTokens: 17, TotalTokens: 19, CachedTokens: 23, CacheWriteTokens: 29},
},
}}
engine := newContractEngineWithOptions(t, frameworkSchemaDir, promptkit.WithLLMClient(client))
result, err := engine.Run(context.Background(), promptkit.RunRequest{
PromptID: frameworkStructuredEventsPromptID,
SessionID: " repair-session ",
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."),
},
Validation: &promptkit.OutputContract{
Format: promptkit.FormatJSON,
ValidationMode: promptkit.ValidationJSONSchema,
SchemaPath: "structured_events.schema.json",
RepairAttempts: 1,
},
})
if err != nil {
t.Fatalf("run: %v", err)
}
if result.RawOutput != client.responses[1].Content || result.Validation.Status != promptkit.ValidationPassed ||
result.Validation.RepairAttempts != 1 {
t.Fatalf("repaired result = %+v", result)
}
wantUsage := promptkit.TokenUsage{PromptTokens: 15, CompletionTokens: 20, TotalTokens: 24, CachedTokens: 30, CacheWriteTokens: 40}
if result.Usage != wantUsage {
t.Fatalf("usage = %+v, want %+v", result.Usage, wantUsage)
}
if len(client.requests) != 2 {
t.Fatalf("generation calls = %d, want 2", len(client.requests))
}
initial, repaired := client.requests[0], client.requests[1]
if initial.Prompt.SessionID != "repair-session" || repaired.Prompt.SessionID != initial.Prompt.SessionID ||
!reflect.DeepEqual(repaired.Target, initial.Target) || repaired.TargetPresence != initial.TargetPresence ||
!reflect.DeepEqual(repaired.StructuredOutput, initial.StructuredOutput) {
t.Fatalf("generation request state drifted: initial=%+v repaired=%+v", initial, repaired)
}
if initial.StructuredOutput == nil || initial.StructuredOutput.JSONSchema == nil {
t.Fatalf("expected structured output on initial request: %+v", initial)
}
}
func TestEngineRunReturnsFinalResultAfterRepairExhaustion(t *testing.T) {
client := &fakeLLMClient{responses: []*promptkit.GenerateResponse{
{Content: "not-json", Usage: promptkit.TokenUsage{PromptTokens: 2, CompletionTokens: 3, TotalTokens: 5}},
{Content: "still-not-json", Usage: promptkit.TokenUsage{PromptTokens: 7, CompletionTokens: 11, TotalTokens: 18}},
}}
engine := newContractEngineWithOptions(t, frameworkSchemaDir, promptkit.WithLLMClient(client))
result, err := engine.Run(context.Background(), promptkit.RunRequest{
PromptID: frameworkMarkdownSummaryPromptID,
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."),
},
Validation: &promptkit.OutputContract{
Format: promptkit.FormatJSON,
ValidationMode: promptkit.ValidationJSON,
RepairAttempts: 1,
},
})
if err != nil || result == nil {
t.Fatalf("run = (%+v, %v), want exhausted result", result, err)
}
if result.RawOutput != "still-not-json" || result.Validation.Status != promptkit.ValidationFailed ||
result.Validation.RepairAttempts != 1 || len(result.Validation.Errors) == 0 ||
result.Usage != (promptkit.TokenUsage{PromptTokens: 9, CompletionTokens: 14, TotalTokens: 23}) {
t.Fatalf("exhausted result = %+v", result)
}
}
func TestRepeatedOptionsUseLastValueInEachCategory(t *testing.T) { func TestRepeatedOptionsUseLastValueInEachCategory(t *testing.T) {
profile := promptkit.Profile{ID: "profile", Endpoint: "http://example.test/v1", Model: "model"} profile := promptkit.Profile{ID: "profile", Endpoint: "http://example.test/v1", Model: "model"}
@@ -973,6 +994,24 @@ func TestRepeatedOptionsUseLastValueInEachCategory(t *testing.T) {
} }
}) })
t.Run("fallback profile source", func(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS("profile", "first-model"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS("profile", "second-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare from last fallback profile source: %v", err)
}
if prepared.EffectiveModelParams.Model != "second-model" {
t.Fatalf("expected last fallback profile source, got %q", prepared.EffectiveModelParams.Model)
}
})
t.Run("in-memory profiles", func(t *testing.T) { t.Run("in-memory profiles", func(t *testing.T) {
first := profile first := profile
first.Model = "first-model" first.Model = "first-model"
@@ -1057,6 +1096,227 @@ func TestRepeatedOptionsUseLastValueInEachCategory(t *testing.T) {
}) })
} }
func TestFallbackProfileSourcePrecedence(t *testing.T) {
const profileID = "application-profile"
prepareModel := func(t *testing.T, engine *promptkit.Engine, promptID string) string {
t.Helper()
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: promptID})
if err != nil {
t.Fatalf("prepare: %v", err)
}
return prepared.EffectiveModelParams.Model
}
t.Run("in-memory profiles override ordinary and fallback profiles", func(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", profileID, "message"), "."),
promptkit.WithProfiles(promptkit.Profile{ID: profileID, Endpoint: "http://example.test/v1", Model: "memory-model"}),
promptkit.WithProfileFS(contractProfileFS(profileID, "ordinary-model"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS(profileID, "fallback-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if model := prepareModel(t, engine, "prompt"); model != "memory-model" {
t.Fatalf("expected in-memory profile, got %q", model)
}
})
t.Run("ordinary filesystem source overrides fallback profile", func(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", profileID, "message"), "."),
promptkit.WithProfileFS(contractProfileFS(profileID, "ordinary-model"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS(profileID, "fallback-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if model := prepareModel(t, engine, "prompt"); model != "ordinary-model" {
t.Fatalf("expected ordinary profile, got %q", model)
}
})
t.Run("ordinary option replaces configured directory", func(t *testing.T) {
profileDir := t.TempDir()
writePublicProfileFile(t, profileDir, profileID, "http://example.test/v1", "directory-model")
engine, err := promptkit.NewEngine(promptkit.Config{ProfileDir: profileDir},
promptkit.WithPromptFS(contractPromptFS("prompt", profileID, "message"), "."),
promptkit.WithProfileFS(contractProfileFS(profileID, "option-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if model := prepareModel(t, engine, "prompt"); model != "option-model" {
t.Fatalf("expected ordinary option profile, got %q", model)
}
})
t.Run("configured directory overrides fallback profile", func(t *testing.T) {
profileDir := t.TempDir()
writePublicProfileFile(t, profileDir, profileID, "http://example.test/v1", "directory-model")
engine, err := promptkit.NewEngine(promptkit.Config{ProfileDir: profileDir},
promptkit.WithPromptFS(contractPromptFS("prompt", profileID, "message"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS(profileID, "fallback-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if model := prepareModel(t, engine, "prompt"); model != "directory-model" {
t.Fatalf("expected configured directory profile, got %q", model)
}
})
t.Run("fallback profile overrides built-in profile", func(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "mistral-small-3", "message"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS("mistral-small-3", "fallback-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if model := prepareModel(t, engine, "prompt"); model != "fallback-model" {
t.Fatalf("expected fallback profile, got %q", model)
}
})
t.Run("missing fallback profile uses built-in profile", func(t *testing.T) {
t.Setenv("OPENROUTER_API_KEY", "test-key")
baseline, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "mistral-small-3", "message"), "."),
)
if err != nil {
t.Fatalf("construct baseline engine: %v", err)
}
want := prepareModel(t, baseline, "prompt")
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "mistral-small-3", "message"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS(profileID, "fallback-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if model := prepareModel(t, engine, "prompt"); model != want {
t.Fatalf("expected built-in profile model %q, got %q", want, model)
}
})
}
func TestFallbackProfileSourcePreservesLazyLoadingAndErrors(t *testing.T) {
const profileID = "application-profile"
t.Run("construction defers malformed fallback profiles", func(t *testing.T) {
_, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(fstest.MapFS{}, "."),
promptkit.WithFallbackProfileFS(fstest.MapFS{
"broken.yaml": &fstest.MapFile{Data: []byte("id: broken\nunknown: value\n")},
}, "."),
)
if err != nil {
t.Fatalf("construct engine with malformed fallback profile: %v", err)
}
})
t.Run("unrelated malformed fallback profile does not block matching definition", func(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", profileID, "message"), "."),
promptkit.WithFallbackProfileFS(fstest.MapFS{
"broken.yaml": &fstest.MapFile{Data: []byte("id: unrelated\nunknown: value\n")},
"valid.yaml": &fstest.MapFile{Data: []byte("id: application-profile\nendpoint: http://example.test/v1\nmodel: fallback-model\n")},
}, "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare from valid fallback profile: %v", err)
}
if prepared.EffectiveModelParams.Model != "fallback-model" {
t.Fatalf("unexpected fallback profile model: %q", prepared.EffectiveModelParams.Model)
}
})
t.Run("matching malformed fallback profile does not reach built-in profile", func(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "mistral-small-3", "message"), "."),
promptkit.WithFallbackProfileFS(fstest.MapFS{
"mistral-small-3.yaml": &fstest.MapFile{Data: []byte("id: mistral-small-3\nendpoint: http://example.test/v1\nmodel: fallback-model\nunknown: value\n")},
}, "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if _, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"}); !errors.Is(err, promptkit.ErrProfileLoad) {
t.Fatalf("expected ErrProfileLoad, got %v", err)
}
})
t.Run("matching malformed ordinary profile does not reach fallback profile", func(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", profileID, "message"), "."),
promptkit.WithProfileFS(fstest.MapFS{
"application-profile.yaml": &fstest.MapFile{Data: []byte("id: application-profile\nendpoint: http://example.test/v1\nmodel: ordinary-model\nunknown: value\n")},
}, "."),
promptkit.WithFallbackProfileFS(contractProfileFS(profileID, "fallback-model"), "."),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
if _, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"}); !errors.Is(err, promptkit.ErrProfileLoad) {
t.Fatalf("expected ErrProfileLoad, got %v", err)
}
})
}
func TestFallbackProfileSourceWorksAcrossWorkflows(t *testing.T) {
const profileID = "application-profile"
client := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "ok"}}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", profileID, "message"), "."),
promptkit.WithFallbackProfileFS(contractProfileFS(profileID, "fallback-model"), "."),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
inspection, err := engine.InspectProfile(context.Background(), profileID)
if err != nil {
t.Fatalf("inspect fallback profile: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare fallback profile: %v", err)
}
preparedExecution, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare execution with fallback profile: %v", err)
}
preparedDetails := preparedExecution.Details()
preparedResult, err := engine.RunPrepared(context.Background(), preparedExecution)
if err != nil {
t.Fatalf("run prepared fallback profile: %v", err)
}
runResult, err := engine.Run(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("run fallback profile: %v", err)
}
for name, model := range map[string]string{
"inspection": inspection.EffectiveModelParams.Model,
"preparation": prepared.EffectiveModelParams.Model,
"prepared execution": preparedDetails.EffectiveModelParams.Model,
"prepared result": preparedResult.EffectiveModelParams.Model,
"run result": runResult.EffectiveModelParams.Model,
} {
if model != "fallback-model" {
t.Fatalf("%s model=%q, want fallback-model", name, model)
}
}
}
func TestEngineSupportsConcurrentPrepareAndRun(t *testing.T) { func TestEngineSupportsConcurrentPrepareAndRun(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{}, engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."), promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),

190
types.go
View File

@@ -82,7 +82,8 @@ const (
// Prepare, PrepareExecution, and Run copy the request's maps, pointers, and // Prepare, PrepareExecution, and Run copy the request's maps, pointers, and
// nested JSON-compatible values before using them. The caller may mutate the // nested JSON-compatible values before using them. The caller may mutate the
// request after any method returns. A successful PrepareExecution retains its // request after any method returns. A successful PrepareExecution retains its
// own private execution snapshot for RunPrepared. // own private execution snapshot for RunPrepared. Excessively deep or large
// JSON-shaped values are rejected for safety.
type RunRequest struct { type RunRequest struct {
// PromptID is the required non-empty prompt identifier. // PromptID is the required non-empty prompt identifier.
PromptID string PromptID string
@@ -116,7 +117,8 @@ type RunRequest struct {
// empty maps are equivalent. // empty maps are equivalent.
Vars map[string]string Vars map[string]string
// Execution optionally overrides individual execution settings. Nil uses // Execution optionally overrides individual execution settings. Nil uses
// the selected profile over its backend, when any, and framework defaults. // the selected profile over its backend, when any, and the framework
// baseline.
Execution *ExecutionTargetOverride Execution *ExecutionTargetOverride
// Validation optionally replaces the prompt's complete output contract. It // Validation optionally replaces the prompt's complete output contract. It
// does not merge individual fields. Nil uses the prompt contract. // does not merge individual fields. Nil uses the prompt contract.
@@ -144,9 +146,10 @@ type PreparedRun struct {
// SelectedBackendID equals EffectiveModelParams.BackendID. It is empty for // SelectedBackendID equals EffectiveModelParams.BackendID. It is empty for
// an endpoint-only profile. // an endpoint-only profile.
SelectedBackendID string `json:"selected_backend_id,omitempty"` SelectedBackendID string `json:"selected_backend_id,omitempty"`
// EffectiveModelParams contains framework defaults overlaid by the selected // EffectiveModelParams contains settings resolved from the framework timeout
// backend, profile, and then request overrides. It excludes resolved API-key // baseline, selected backend, profile, and then request overrides. Unset
// values. // optional provider controls remain zero rather than reporting a provider
// default. It excludes resolved API-key values.
EffectiveModelParams ExecutionTarget `json:"effective_model_params"` EffectiveModelParams ExecutionTarget `json:"effective_model_params"`
// OutputContract is the complete effective output contract. // OutputContract is the complete effective output contract.
OutputContract OutputContract `json:"output_contract"` OutputContract OutputContract `json:"output_contract"`
@@ -240,8 +243,8 @@ type ArtifactRef struct {
// URI is the file path for ArtifactRefFile and optional provenance metadata // URI is the file path for ArtifactRefFile and optional provenance metadata
// for ArtifactRefInline. // for ArtifactRefInline.
URI string URI string
// Body is the content for ArtifactRefInline and is ignored for // Body is the content for ArtifactRefInline, where an empty value is valid,
// ArtifactRefFile. // and is ignored for ArtifactRefFile.
Body string Body string
} }
@@ -294,17 +297,22 @@ type ExecutionTarget struct {
// empty for endpoint-only profiles. It is supplied to injected LLMClient // empty for endpoint-only profiles. It is supplied to injected LLMClient
// implementations as part of the effective target. // implementations as part of the effective target.
BackendID string `json:"backend_id,omitempty"` BackendID string `json:"backend_id,omitempty"`
// Endpoint is the model-provider base URL. // Endpoint is the normalized absolute HTTP or HTTPS model-provider base URL.
// It has a host and no user information, query, or fragment.
Endpoint string `json:"endpoint"` Endpoint string `json:"endpoint"`
// Model is the provider model identifier. // Model is the provider model identifier.
Model string `json:"model"` Model string `json:"model"`
// Temperature is the effective sampling temperature from 0 through 2. // Temperature is the resolved sampling temperature from 0 through 2. Zero
// leaves the field unspecified to compatible providers unless the
// corresponding ExecutionTargetPresence bit is true.
Temperature float64 `json:"temperature"` Temperature float64 `json:"temperature"`
// MaxTokens is the non-negative effective output-token limit. Zero leaves // MaxTokens is the non-negative resolved output-token limit. Zero leaves
// the limit unspecified to compatible providers unless it was an explicit // the limit unspecified to compatible providers unless the corresponding
// request override. // ExecutionTargetPresence bit is true.
MaxTokens int `json:"max_tokens"` MaxTokens int `json:"max_tokens"`
// TopP is the effective nucleus-sampling value from 0 through 1. // TopP is the resolved nucleus-sampling value from 0 through 1. Zero leaves
// the field unspecified to compatible providers unless the corresponding
// ExecutionTargetPresence bit is true.
TopP float64 `json:"top_p"` TopP float64 `json:"top_p"`
// TimeoutSeconds is the non-negative per-generation deadline. Zero disables // TimeoutSeconds is the non-negative per-generation deadline. Zero disables
// this deadline without disabling caller cancellation or the transport cap. // this deadline without disabling caller cancellation or the transport cap.
@@ -314,7 +322,10 @@ type ExecutionTarget struct {
// ReasoningEffort is the effective opaque provider-specific reasoning // ReasoningEffort is the effective opaque provider-specific reasoning
// setting. An empty value instructs model clients to omit reasoning. // setting. An empty value instructs model clients to omit reasoning.
ReasoningEffort string `json:"reasoning_effort"` ReasoningEffort string `json:"reasoning_effort"`
// APIKeyEnv is an environment-variable name, not its credential value. // APIKeyEnv is the resolved name of an optional environment lookup source,
// not its credential value. The built-in client omits Authorization when no
// usable direct or environment credential is available; injected clients may
// resolve this metadata differently.
APIKeyEnv string `json:"api_key_env"` APIKeyEnv string `json:"api_key_env"`
// ExtraParams contains copied JSON-compatible provider parameters. // ExtraParams contains copied JSON-compatible provider parameters.
ExtraParams map[string]any `json:"extra_params"` ExtraParams map[string]any `json:"extra_params"`
@@ -329,9 +340,11 @@ type ExecutionTarget struct {
type ProfileInspection struct { type ProfileInspection struct {
// ProfileID is the trimmed, exact profile ID inspected by the engine. // ProfileID is the trimmed, exact profile ID inspected by the engine.
ProfileID string ProfileID string
// EffectiveModelParams contains framework defaults overlaid by the selected // EffectiveModelParams contains settings resolved from the framework timeout
// backend and then the profile, without a request override. APIKeyEnv is an // baseline, selected backend, and then profile, without a request override.
// environment-variable name, never its credential value. // Unset optional provider controls remain zero rather than reporting a
// provider default. APIKeyEnv is an environment-variable name, never its
// credential value.
EffectiveModelParams ExecutionTarget EffectiveModelParams ExecutionTarget
// APIKeyRequired reports that a later execution must supply a direct API // APIKeyRequired reports that a later execution must supply a direct API
// key or an explicit request environment override. It is mutually exclusive // key or an explicit request environment override. It is mutually exclusive
@@ -382,19 +395,28 @@ type PromptInspection struct {
// fields replace profile values and preserve explicit zero or empty values. A // fields replace profile values and preserve explicit zero or empty values. A
// non-empty ExtraParams map replaces the complete profile or backend map // non-empty ExtraParams map replaces the complete profile or backend map
// rather than merging keys. Empty string fields, nil pointers, and a nil or // rather than merging keys. Empty string fields, nil pointers, and a nil or
// empty ExtraParams map inherit the selected profile over its backend, when // empty ExtraParams map inherit lower-precedence values. An optional provider
// any, and framework defaults. // control that remains zero is unspecified; TimeoutSeconds retains its
// framework deadline when no higher-precedence value is present.
type ExecutionTargetOverride struct { type ExecutionTargetOverride struct {
// Endpoint replaces the profile or backend endpoint when non-empty without // Endpoint replaces the profile or backend endpoint when non-empty without
// changing the effective BackendID. // changing the effective BackendID. Preparation trims it and requires an
// absolute HTTP or HTTPS URL with a host and no user information, query, or
// fragment.
Endpoint string Endpoint string
// Model replaces the profile model when non-empty. // Model replaces the profile model when non-empty.
Model string Model string
// Temperature, when non-nil, must point to a value from 0 through 2. // Temperature, when non-nil, must point to a value from 0 through 2. A
// pointed-to zero is explicitly present; nil inherits a lower-precedence
// value and otherwise leaves the provider control unspecified.
Temperature *float64 Temperature *float64
// MaxTokens, when non-nil, must point to a non-negative value. // MaxTokens, when non-nil, must point to a non-negative value. A pointed-to
// zero is explicitly present; nil inherits a lower-precedence value and
// otherwise leaves the provider control unspecified.
MaxTokens *int MaxTokens *int
// TopP, when non-nil, must point to a value from 0 through 1. // TopP, when non-nil, must point to a value from 0 through 1. A pointed-to
// zero is explicitly present; nil inherits a lower-precedence value and
// otherwise leaves the provider control unspecified.
TopP *float64 TopP *float64
// TimeoutSeconds, when non-nil, must point to a non-negative value. A // TimeoutSeconds, when non-nil, must point to a non-negative value. A
// pointed-to zero disables the per-generation deadline. // pointed-to zero disables the per-generation deadline.
@@ -407,61 +429,81 @@ type ExecutionTargetOverride struct {
// inherited value and disables reasoning for this run. Non-blank values // inherited value and disables reasoning for this run. Non-blank values
// are opaque and are not validated against a fixed vocabulary. // are opaque and are not validated against a fixed vocabulary.
ReasoningEffort *string ReasoningEffort *string
// APIKeyEnv replaces the profile or backend environment-variable name when // APIKeyEnv replaces the profile or backend optional environment lookup
// non-blank. A direct RunRequest.APIKey still takes precedence over // source when non-blank. A direct RunRequest.APIKey still takes precedence.
// environment lookup. // The built-in client omits Authorization when neither source has a usable
// value; injected clients may resolve this metadata differently.
APIKeyEnv string APIKeyEnv string
// ExtraParams, when non-empty, replaces the complete profile or backend map. // ExtraParams, when non-empty, replaces the complete profile or backend map.
// Values must be JSON-compatible: nil, booleans, finite numbers, strings, // Values must be JSON-compatible: nil, booleans, finite numbers, strings,
// arrays or slices, and maps with non-empty string keys. Cycles are invalid. // arrays or slices, and maps with non-empty string keys. Cycles and
// excessively deep or large values are invalid.
ExtraParams map[string]any ExtraParams map[string]any
} }
// Profile is an in-memory execution profile for library consumers. // Profile is an in-memory execution profile for library consumers.
// //
// It is equivalent to a loaded profile file after validation. Raw API keys do // A standalone Profile is equivalent to a loaded profile file after local
// not belong in profiles; use APIKeyRequired to require callers to provide a // validation. A derived profile names BaseProfileID and can inherit target
// RunRequest.APIKey or explicit request ExecutionTargetOverride.APIKeyEnv, or // fields when selected or inspected. Raw API keys do not belong in profiles;
// use profile YAML api_key_env with file and FS profile sources. Profile has no // use APIKeyRequired to require callers to provide a RunRequest.APIKey or
// stable JSON representation. // explicit request ExecutionTargetOverride.APIKeyEnv, or use profile YAML
// api_key_env with file and FS profile sources. Profile has no stable JSON
// representation.
// //
// WithProfiles validates and copies Profile values during NewEngine. Numeric // WithProfiles locally validates and copies Profile values during NewEngine.
// zero, blank strings, and an empty ExtraParams map inherit framework defaults; // It checks base-reference existence and resolved target completeness when a
// use ExecutionTargetOverride pointer fields to request explicit numeric zero. // derived profile is selected or inspected. Zero Temperature, MaxTokens, and
// TopP values and blank ServiceTier and ReasoningEffort values leave those
// provider controls unspecified. A zero TimeoutSeconds retains the framework
// deadline, while an empty ExtraParams map inherits backend request defaults.
// Use ExecutionTargetOverride pointer fields to request an explicit numeric
// zero.
type Profile struct { type Profile struct {
// ID is the required non-blank profile identifier. WithProfiles trims it. // ID is the required non-blank profile identifier. WithProfiles trims it.
ID string ID string
// BaseProfileID optionally names one base profile. WithProfiles trims it. A
// non-blank value permits required target fields to be inherited when the
// profile is selected or inspected, which is also when reference existence
// and resolved completeness are checked. A blank value leaves this as a
// standalone profile.
BaseProfileID string
// BackendID optionally selects an engine backend. WithProfiles trims it. // BackendID optionally selects an engine backend. WithProfiles trims it.
// Backend membership is checked when a request selects the profile; an // Backend membership is checked when a request selects the profile; an
// unknown ID makes preparation fail with ErrProfileLoad. // unknown ID makes preparation fail with ErrProfileLoad.
BackendID string BackendID string
// Endpoint is the model-provider base URL. It is required only when // Endpoint is the model-provider base URL. A standalone Profile requires an
// BackendID is blank and otherwise overrides the backend endpoint when // endpoint when BackendID is blank; a derived Profile may inherit either
// non-blank. // field. A non-blank endpoint overrides the backend endpoint when
// non-blank. WithProfiles trims it and requires an absolute HTTP or HTTPS URL
// with a host and no user information, query, or fragment.
Endpoint string Endpoint string
// Model is the required non-blank provider model identifier. // Model is the provider model identifier. It is required for a standalone
// Profile and may be inherited by a derived Profile.
Model string Model string
// Temperature is from 0 through 2. Zero inherits the framework default. // Temperature is from 0 through 2. Zero leaves the provider control
// unspecified.
Temperature float64 Temperature float64
// MaxTokens is non-negative. Zero inherits the framework default. // MaxTokens is non-negative. Zero leaves the provider control unspecified.
MaxTokens int MaxTokens int
// TopP is from 0 through 1. Zero inherits the framework default rather than // TopP is from 0 through 1. Zero leaves the provider control unspecified
// selecting an explicit zero. // rather than selecting an explicit zero.
TopP float64 TopP float64
// TimeoutSeconds is non-negative. Zero inherits the framework default. // TimeoutSeconds is non-negative. Zero retains the framework deadline.
TimeoutSeconds int TimeoutSeconds int
// ServiceTier is optional; a blank value inherits the framework default. // ServiceTier is optional; a blank value leaves it unspecified.
ServiceTier string ServiceTier string
// ReasoningEffort is optional; a blank value inherits the framework // ReasoningEffort is optional; a blank value leaves it unspecified.
// default.
ReasoningEffort string ReasoningEffort string
// APIKeyRequired clears a backend's inherited API-key environment name and // APIKeyRequired clears a backend's inherited API-key environment name and
// requires a non-blank RunRequest.APIKey unless the request explicitly // requires a non-blank RunRequest.APIKey unless the request explicitly
// supplies ExecutionTargetOverride.APIKeyEnv. It does not store a credential. // supplies ExecutionTargetOverride.APIKeyEnv. When false, a named
// environment source remains optional. It does not store a credential.
APIKeyRequired bool APIKeyRequired bool
// ExtraParams contains provider-specific JSON-compatible values. An empty // ExtraParams contains provider-specific JSON-compatible values. An empty
// map inherits backend request defaults, when any. WithProfiles validates // map inherits backend request defaults, when any. WithProfiles validates
// and deeply copies it during NewEngine. // and deeply copies it during NewEngine. Excessively deep or large values
// are rejected for safety.
ExtraParams map[string]any ExtraParams map[string]any
} }
@@ -469,13 +511,16 @@ type Profile struct {
// profile. // profile.
// //
// It contains ordinary profile fields for OpenAI-compatible chat-completions // It contains ordinary profile fields for OpenAI-compatible chat-completions
// endpoints. APIKeyRequired follows Profile.APIKeyRequired. Raw API keys do not // endpoints. BaseProfileID and APIKeyRequired follow Profile. Raw API keys do
// belong in this config. OpenAICompatibleProfileConfig has no stable JSON // not belong in this config. OpenAICompatibleProfileConfig has no stable JSON
// representation and is not validated until its resulting Profile is supplied // representation and is not validated until its resulting Profile is supplied
// through WithProfiles to NewEngine. // through WithProfiles to NewEngine.
type OpenAICompatibleProfileConfig struct { type OpenAICompatibleProfileConfig struct {
// ID becomes Profile.ID. // ID becomes Profile.ID.
ID string ID string
// BaseProfileID becomes Profile.BaseProfileID. A non-blank value permits the
// resulting Profile to inherit target fields when it is selected or inspected.
BaseProfileID string
// BackendID becomes Profile.BackendID. // BackendID becomes Profile.BackendID.
BackendID string BackendID string
// Endpoint becomes Profile.Endpoint. // Endpoint becomes Profile.Endpoint.
@@ -520,11 +565,11 @@ type ExecutionTargetPresence struct {
// JSON representation. // JSON representation.
// //
// A non-nil RunRequest.Validation replaces the complete prompt contract. It // A non-nil RunRequest.Validation replaces the complete prompt contract. It
// does not merge fields. The public Engine validates generated output once and // does not merge fields. The public Engine performs bounded correction after a
// does not install an output repairer. // failed eligible validation when RepairAttempts is positive.
type OutputContract struct { type OutputContract struct {
// Format selects generated artifact metadata. An empty effective value // Format selects generated artifact metadata. An empty value in a non-nil
// defaults to FormatText. // request replacement defaults to FormatText.
Format OutputFormat `json:"format"` Format OutputFormat `json:"format"`
// ValidationMode selects the content check. Use one of the declared // ValidationMode selects the content check. Use one of the declared
// ValidationMode constants. // ValidationMode constants.
@@ -532,9 +577,9 @@ type OutputContract struct {
// SchemaPath is required when ValidationMode is ValidationJSONSchema and is // SchemaPath is required when ValidationMode is ValidationJSONSchema and is
// ignored by other modes. // ignored by other modes.
SchemaPath string `json:"schema_path"` SchemaPath string `json:"schema_path"`
// RepairAttempts is a requested repair limit. A non-positive value requests // RepairAttempts is an additional generation-call budget from zero through
// no repairs. The public Engine performs no repairs even when this value is // three. Zero is single-pass. A positive value is valid only with basic,
// positive, so its runs report zero attempts used. // json, or json_schema validation.
RepairAttempts int `json:"repair_attempts"` RepairAttempts int `json:"repair_attempts"`
} }
@@ -550,8 +595,8 @@ type ValidationResult struct {
Errors []string `json:"errors,omitempty"` Errors []string `json:"errors,omitempty"`
// SchemaPath is the effective schema path for JSON Schema validation. // SchemaPath is the effective schema path for JSON Schema validation.
SchemaPath string `json:"schema_path,omitempty"` SchemaPath string `json:"schema_path,omitempty"`
// RepairAttempts is the number of repairs actually attempted. It is always // RepairAttempts is the number of corrective generation calls actually
// zero for the public Engine. // started for this result.
RepairAttempts int `json:"repair_attempts"` RepairAttempts int `json:"repair_attempts"`
// IsValid is true for ValidationPassed and ValidationSkipped and false for // IsValid is true for ValidationPassed and ValidationSkipped and false for
// ValidationFailed. // ValidationFailed.
@@ -640,10 +685,10 @@ type StructuredOutputJSONSpec struct {
// and retained copies. It is responsible for the cancellation behavior of any // and retained copies. It is responsible for the cancellation behavior of any
// work it starts and for synchronizing access to retained or shared data. // work it starts and for synchronizing access to retained or shared data.
// //
// A returned error makes Run or RunPrepared return ErrLLMGenerate while // An arbitrary returned error makes Run or RunPrepared return ErrLLMGenerate
// preserving the client error through errors.Is. A nil response with a nil // while preserving the client error through errors.Is rather than translating
// error also produces ErrLLMGenerate. Promptkit copies the non-nil response // it. A nil response with a nil error also produces ErrLLMGenerate. Promptkit
// before returning from either method. // copies the non-nil response before returning from either method.
type LLMClient interface { type LLMClient interface {
Generate(context.Context, GenerateRequest) (*GenerateResponse, error) Generate(context.Context, GenerateRequest) (*GenerateResponse, error)
} }
@@ -669,9 +714,8 @@ type GenerateRequest struct {
// GenerateResponse is returned by an injected LLM client and has a stable JSON // GenerateResponse is returned by an injected LLM client and has a stable JSON
// representation. // representation.
type GenerateResponse struct { type GenerateResponse struct {
// Content is the generated output. It must be non-empty when using the // Content is the generated output. It may be explicitly empty; Promptkit
// built-in client; injected clients may return empty content for Promptkit // applies the effective output contract to classify it.
// validation to classify.
Content string `json:"content"` Content string `json:"content"`
// Usage is the client's token accounting. // Usage is the client's token accounting.
Usage TokenUsage `json:"usage"` Usage TokenUsage `json:"usage"`
@@ -679,8 +723,12 @@ type GenerateResponse struct {
// File returns a file-backed artifact reference whose URI is path. // File returns a file-backed artifact reference whose URI is path.
// //
// The default artifact reader opens path as a caller-selected operating-system // The default artifact reader accepts path only when it resolves to a regular
// path without restricting it to an application root or imposing a size limit. // operating-system file, checking that condition before and after opening it.
// It reads synchronously in bounded chunks and checks context cancellation
// before opening, before and after each read, and before returning the
// artifact; it cannot interrupt a filesystem operation already in progress.
// It does not restrict path to an application root or impose a size limit.
// Applications accepting untrusted paths must validate them before calling // Applications accepting untrusted paths must validate them before calling
// Promptkit or use [WithArtifactReader] to enforce application policy. // Promptkit or use [WithArtifactReader] to enforce application policy.
func File(path string) ArtifactRef { func File(path string) ArtifactRef {
@@ -688,13 +736,13 @@ func File(path string) ArtifactRef {
} }
// Inline returns an inline artifact reference whose Body is body and whose URI // Inline returns an inline artifact reference whose Body is body and whose URI
// is empty. // is empty. An empty body is a valid, explicitly supplied input.
func Inline(body string) ArtifactRef { func Inline(body string) ArtifactRef {
return ArtifactRef{Type: ArtifactRefInline, Body: body} return ArtifactRef{Type: ArtifactRefInline, Body: body}
} }
// InlineWithURI returns an inline artifact reference with body content and uri // InlineWithURI returns an inline artifact reference with body content and uri
// provenance metadata. // provenance metadata. An empty body is a valid, explicitly supplied input.
func InlineWithURI(uri string, body string) ArtifactRef { func InlineWithURI(uri string, body string) ArtifactRef {
return ArtifactRef{Type: ArtifactRefInline, URI: uri, Body: body} return ArtifactRef{Type: ArtifactRefInline, URI: uri, Body: body}
} }