41 Commits

Author SHA1 Message Date
bd6cffc9d0 Prepare documentation for Promptkit v0.4.0 2026-07-30 23:48:25 +00:00
e40c4f182b Document structured capacity errors 2026-07-30 23:26:44 +00:00
7428e50c2c Expose structured capacity errors 2026-07-30 23:24:04 +00:00
63c67a4520 Add internal capacity error identity 2026-07-30 23:21:27 +00:00
25a7052a3d Clarify prompt input requirement documentation 2026-07-30 22:10:33 +00:00
fc3255967e Document prompt inspection API 2026-07-30 21:06:33 +00:00
e920168b30 Expose prompt inspection through the engine 2026-07-30 21:03:47 +00:00
272b6a4bc1 Add internal prompt inspection 2026-07-30 21:00:05 +00:00
dde48a31fc Document profile inspection API 2026-07-30 19:57:07 +00:00
242eace4a7 Expose profile inspection through the engine 2026-07-30 19:52:00 +00:00
0bf5f88136 Add internal profile inspection resolution 2026-07-30 19:48:09 +00:00
369ab5392d Add feature roadmap and implementation plan for profile inspection API 2026-07-30 19:42:51 +00:00
2ba0146e5d Tighten prepared execution credential handling 2026-07-30 19:06:17 +00:00
6112c2af0c Document prepared execution workflow 2026-07-30 18:27:03 +00:00
f5e12c00f5 Expose prepared execution handles 2026-07-30 18:20:35 +00:00
49fe402dd2 Add prepared execution lifecycle to the runner 2026-07-30 18:10:28 +00:00
c301eb8d55 Add frozen validation preparation plans 2026-07-30 17:59:45 +00:00
c13e9710d9 Add feature roadmap and implementation plan for downstream consumer wishlist items 2026-07-30 17:54:44 +00:00
87b5ec3d75 Organize downstream feature requests in the future roadmap 2026-07-30 17:13:26 +00:00
cb4028a637 Add feature roadmaps with wishlists from downstream consumers 2026-07-30 16:59:37 +00:00
5a1bff4529 Prepare the local backend convenience release 2026-07-30 04:01:19 +00:00
805a7f965d Document local backend configuration paths 2026-07-30 03:40:52 +00:00
147f5e5ff5 Add local backend convenience constructor 2026-07-30 03:38:23 +00:00
e361c97bb5 Document the v0.2.0 upgrade path 2026-07-29 23:22:38 +00:00
be67707582 Keep capacity cancellation tests independent of implementation details 2026-07-29 22:41:12 +00:00
e61ab700c7 Document backend capacity management 2026-07-29 21:16:49 +00:00
d2c4051dd0 Wire backend capacity into engine runtime 2026-07-29 21:11:08 +00:00
861da355d8 Add early backend run admission 2026-07-29 21:01:37 +00:00
a752f88166 Add backend capacity manager 2026-07-29 20:55:46 +00:00
238fa90bfa Add backend capacity policy configuration 2026-07-29 20:48:10 +00:00
ffe6d261a9 Add roadmap and implementation plan for backend concurrency scaling 2026-07-29 20:42:45 +00:00
dc39562ff7 Keep session IDs on repair requests and retire completed plans 2026-07-29 20:25:00 +00:00
0a839aa16d Document per-run session and reasoning overrides 2026-07-29 19:47:39 +00:00
f6ee18f6b3 Add direct per-run session overrides 2026-07-29 19:43:26 +00:00
eb8ab215e8 Add tri-state reasoning overrides 2026-07-29 19:36:48 +00:00
f89cb94ed2 Address backend implementation review findings 2026-07-29 19:08:05 +00:00
359b7313f4 Harden backend contracts and documentation 2026-07-29 17:22:55 +00:00
ae210b3c26 Add engine-scoped backend registration 2026-07-29 17:15:51 +00:00
810f80e7c9 Resolve profiles through built-in backends 2026-07-29 17:10:34 +00:00
8d00354c59 Add immutable backend registry foundation 2026-07-29 16:59:43 +00:00
d0010689f3 Prepare the backend registry implementation roadmap 2026-07-29 16:50:36 +00:00
86 changed files with 9368 additions and 780 deletions

View File

@@ -31,6 +31,17 @@ Contributors should start with the [development guide](docs/development.md).
The [architecture policy](docs/policy/architecture.md) defines the library The [architecture policy](docs/policy/architecture.md) defines the library
boundary and constraints that framework work must preserve. boundary and constraints that framework work must preserve.
## Release Guidance
Consumers upgrading from `v0.3.0` to `v0.4.0` should read the
[v0.4.0 changelog and adoption guide](docs/releases/v0.4.0.md).
Consumers upgrading from `v0.2.0` to `v0.3.0` should read the
[v0.3.0 changelog](docs/releases/v0.3.0.md).
Consumers moving from `v0.1.0` to `v0.2.0` should read the
[v0.2.0 changelog and migration guide](docs/releases/v0.2.0.md).
## Related Project ## Related Project
[Scriptorium](https://gitea.maximumdirect.net/eric/scriptorium) is the CLI and [Scriptorium](https://gitea.maximumdirect.net/eric/scriptorium) is the CLI and

99
backends.go Normal file
View File

@@ -0,0 +1,99 @@
package promptkit
import (
"gitea.maximumdirect.net/eric/promptkit/internal/backend"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
// BackendOpenRouter is the reserved ID of Promptkit's built-in OpenRouter
// backend.
const BackendOpenRouter = backend.OpenRouterID
// BackendLocal is the case-sensitive conventional ID used by [LocalBackend].
// It is not a built-in or reserved backend and must be registered with
// [WithBackend].
const BackendLocal = "local"
// Backend configures one engine-scoped OpenAI-compatible backend.
//
// Backend has no stable JSON representation. Use keyed literals so additions
// to this configuration value do not break source compatibility.
type Backend struct {
// ID is the stable, case-sensitive registry key. NewEngine trims it and
// requires a non-blank value. BackendOpenRouter is reserved.
ID string
// Endpoint is the OpenAI-compatible base endpoint. NewEngine trims it and
// requires an absolute HTTP or HTTPS URL with a host and without user
// information, a query string, or a fragment. Paths are allowed.
Endpoint string
// APIKeyEnv optionally names the environment variable containing the API
// key. NewEngine trims it and requires the portable form
// [A-Za-z_][A-Za-z0-9_]*. Store only the name, never a credential value.
APIKeyEnv string
// ExtraParams contains backend-wide request defaults. Values must be
// JSON-compatible, finite, acyclic, and keyed by non-empty strings. Keys
// must not be model, session_id, messages, temperature, max_tokens, top_p,
// service_tier, reasoning_effort, or response_format. An empty map supplies
// no defaults. NewEngine deeply copies the map.
ExtraParams map[string]any
// ConcurrencyLimit is the maximum number of simultaneous model-generation
// calls allowed for this backend within one Engine. Zero leaves the backend
// unlimited. A negative value makes NewEngine fail with ErrInvalidConfig.
ConcurrencyLimit int
// QueueCapacity controls how many additional Run or RunPrepared calls may
// be admitted beyond ConcurrencyLimit. Nil uses 1024 when ConcurrencyLimit
// is positive; a pointer uses its exact value, including zero. The pointed-to
// value must be non-negative, and QueueCapacity must be nil when
// ConcurrencyLimit is zero. Their sum must fit in an int. WithBackend copies
// the value and does not retain the pointer.
QueueCapacity *int
}
// LocalBackend returns a caller-owned Backend for a conventional local
// OpenAI-compatible endpoint. It sets ID to BackendLocal and copies endpoint
// and concurrencyLimit into Endpoint and ConcurrencyLimit without
// normalization or validation. APIKeyEnv, ExtraParams, and QueueCapacity keep
// their zero values.
//
// LocalBackend does not read environment variables, register the value, or
// mutate engine or package state. Supply the returned value through
// [WithBackend]; [NewEngine] then applies the ordinary backend validation and
// concurrency semantics, including default queue capacity for a positive
// limit, unlimited behavior for zero, and ErrInvalidConfig for a negative
// limit.
func LocalBackend(endpoint string, concurrencyLimit int) Backend {
return Backend{
ID: BackendLocal,
Endpoint: endpoint,
ConcurrencyLimit: concurrencyLimit,
}
}
// WithBackend adds one Backend registration to the constructed Engine.
//
// Registrations accumulate in option order. Every normalized ID must be unique
// across consumer registrations and built-ins; a duplicate or invalid
// definition makes NewEngine fail with ErrInvalidConfig. In particular,
// BackendOpenRouter cannot be replaced. The immutable registration is scoped
// to the resulting Engine and cannot be enumerated, replaced, removed, or
// mutated after construction. WithBackend does not install package-global
// state.
func WithBackend(backend Backend) Option {
queueCapacity := 0
queueCapacitySet := backend.QueueCapacity != nil
if queueCapacitySet {
queueCapacity = *backend.QueueCapacity
}
return optionFunc(func(options *engineOptions) error {
options.backends = append(options.backends, domain.Backend{
ID: backend.ID,
Endpoint: backend.Endpoint,
APIKeyEnv: backend.APIKeyEnv,
ExtraParams: backend.ExtraParams,
ConcurrencyLimit: backend.ConcurrencyLimit,
QueueCapacity: queueCapacity,
QueueCapacitySet: queueCapacitySet,
})
return nil
})
}

409
capacity_contract_test.go Normal file
View File

@@ -0,0 +1,409 @@
package promptkit_test
import (
"context"
"errors"
"sync"
"testing"
"time"
"gitea.maximumdirect.net/eric/promptkit"
)
func TestEngineLimitsInjectedClientConcurrency(t *testing.T) {
release := make(chan struct{})
client := newCapacityGateClient(release, 8)
engine := newBackendCapacityEngine(t, client, 2, capacityInt(4), nil)
results := make(chan capacityRunResult, 6)
for i := 0; i < 6; i++ {
go runCapacityRequest(engine, context.Background(), promptkit.RunRequest{
PromptID: "prompt",
Execution: &promptkit.ExecutionTargetOverride{
Endpoint: "http://request.example/v1",
},
}, results)
}
first := awaitCapacityRequest(t, client.started)
second := awaitCapacityRequest(t, client.started)
if first.Target.BackendID != "limited" || second.Target.BackendID != "limited" {
t.Fatalf("endpoint override changed backend pool: first=%q second=%q",
first.Target.BackendID, second.Target.BackendID)
}
if active, peak, _ := client.snapshot(); active != 2 || peak != 2 {
t.Fatalf("client concurrency before release=(active=%d peak=%d), want 2", active, peak)
}
close(release)
for i := 0; i < 6; i++ {
outcome := awaitCapacityRun(t, results)
if outcome.err != nil || outcome.result == nil {
t.Fatalf("run outcome=(%+v, %v), want success", outcome.result, outcome.err)
}
}
if _, peak, calls := client.snapshot(); peak > 2 || calls != 6 {
t.Fatalf("client observations=(peak=%d calls=%d), want peak <= 2 and 6 calls", peak, calls)
}
}
func TestEngineRejectsRunBeforeCompletionWhenAdmissionIsFull(t *testing.T) {
artifactRelease := make(chan struct{})
reader := &capacityArtifactReader{
entered: make(chan struct{}, 2),
release: artifactRelease,
}
client := newCapacityGateClient(closedCapacityChannel(), 2)
engine := newBackendCapacityEngine(t, client, 1, capacityInt(0), reader)
firstResult := make(chan capacityRunResult, 1)
go runCapacityRequest(engine, context.Background(), capacityInputRequest("http://first.example/v1"), firstResult)
awaitCapacitySignal(t, reader.entered, "first artifact read")
canceledContext, cancel := context.WithCancel(context.Background())
cancel()
result, err := engine.Run(canceledContext, capacityInputRequest("http://canceled.example/v1"))
if result != nil || !errors.Is(err, context.Canceled) {
t.Fatalf("canceled capacity admission=(%+v, %v), want context cancellation", result, err)
}
var canceledCapacityErr *promptkit.CapacityError
if errors.Is(err, promptkit.ErrCapacityExceeded) || errors.As(err, &canceledCapacityErr) {
t.Fatalf("canceled admission exposed capacity rejection: %v", err)
}
result, err = engine.Run(context.Background(), capacityInputRequest("http://second.example/v1"))
if result != nil {
t.Fatalf("capacity rejection returned partial result: %+v", result)
}
if !errors.Is(err, promptkit.ErrCapacityExceeded) {
t.Fatalf("capacity rejection=%v, want ErrCapacityExceeded", err)
}
if errors.Is(err, promptkit.ErrInvalidRequest) || errors.Is(err, promptkit.ErrLLMGenerate) {
t.Fatalf("capacity rejection had an unrelated category: %v", err)
}
var capacityErr *promptkit.CapacityError
if !errors.As(err, &capacityErr) || capacityErr == nil {
t.Fatalf("capacity rejection=%v, want CapacityError", err)
}
if capacityErr.BackendID != "limited" {
t.Fatalf("capacity backend ID=%q, want limited", capacityErr.BackendID)
}
capacityErr.BackendID = "changed"
result, err = engine.Run(context.Background(), capacityInputRequest("http://third.example/v1"))
var subsequentCapacityErr *promptkit.CapacityError
if result != nil || !errors.As(err, &subsequentCapacityErr) ||
subsequentCapacityErr == nil || subsequentCapacityErr.BackendID != "limited" {
t.Fatalf("subsequent capacity rejection=(%+v, %v), want independent limited CapacityError", result, err)
}
if calls := reader.callCount(); calls != 1 {
t.Fatalf("artifact calls=%d, want only the admitted run", calls)
}
if _, _, calls := client.snapshot(); calls != 0 {
t.Fatalf("client calls=%d before admitted run was released, want 0", calls)
}
close(artifactRelease)
outcome := awaitCapacityRun(t, firstResult)
if outcome.err != nil || outcome.result == nil {
t.Fatalf("first run outcome=(%+v, %v), want success", outcome.result, outcome.err)
}
}
func TestBackendCapacityIsIndependentBetweenEngines(t *testing.T) {
firstRelease := make(chan struct{})
firstClient := newCapacityGateClient(firstRelease, 1)
firstEngine := newBackendCapacityEngine(t, firstClient, 1, capacityInt(0), nil)
secondClient := newCapacityGateClient(closedCapacityChannel(), 1)
secondEngine := newBackendCapacityEngine(t, secondClient, 1, capacityInt(0), nil)
firstResult := make(chan capacityRunResult, 1)
go runCapacityRequest(firstEngine, context.Background(), promptkit.RunRequest{PromptID: "prompt"}, firstResult)
awaitCapacityRequest(t, firstClient.started)
result, err := secondEngine.Run(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil || result == nil {
t.Fatalf("second engine run=(%+v, %v), want independent success", result, err)
}
if _, _, calls := secondClient.snapshot(); calls != 1 {
t.Fatalf("second engine client calls=%d, want 1", calls)
}
close(firstRelease)
outcome := awaitCapacityRun(t, firstResult)
if outcome.err != nil || outcome.result == nil {
t.Fatalf("first engine run=(%+v, %v), want success", outcome.result, outcome.err)
}
}
func TestUnlimitedBackendsRetainInjectedClientConcurrency(t *testing.T) {
tests := []struct {
name string
configure func(*testing.T, promptkit.LLMClient) *promptkit.Engine
}{
{
name: "custom backend",
configure: func(t *testing.T, client promptkit.LLMClient) *promptkit.Engine {
return newBackendCapacityEngine(t, client, 0, nil, nil)
},
},
{
name: "endpoint-only profile",
configure: func(t *testing.T, client promptkit.LLMClient) *promptkit.Engine {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile", Endpoint: "http://endpoint.example/v1", Model: "model",
}),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct endpoint-only engine: %v", err)
}
return engine
},
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
release := make(chan struct{})
client := newCapacityGateClient(release, 2)
engine := tc.configure(t, client)
results := make(chan capacityRunResult, 2)
for i := 0; i < 2; i++ {
go runCapacityRequest(
engine,
context.Background(),
promptkit.RunRequest{PromptID: "prompt"},
results,
)
}
awaitCapacityRequest(t, client.started)
awaitCapacityRequest(t, client.started)
if active, peak, _ := client.snapshot(); active != 2 || peak != 2 {
t.Fatalf("unlimited concurrency=(active=%d peak=%d), want 2", active, peak)
}
close(release)
for i := 0; i < 2; i++ {
outcome := awaitCapacityRun(t, results)
if outcome.err != nil || outcome.result == nil {
t.Fatalf("run outcome=(%+v, %v), want success", outcome.result, outcome.err)
}
}
})
}
}
func TestCapacityExceededSentinelContract(t *testing.T) {
if promptkit.ErrCapacityExceeded == nil {
t.Fatal("ErrCapacityExceeded is nil")
}
var nilCapacityErr *promptkit.CapacityError
zeroCapacityErr := &promptkit.CapacityError{}
populatedCapacityErr := &promptkit.CapacityError{BackendID: "limited"}
for _, capacityErr := range []error{nilCapacityErr, zeroCapacityErr} {
if !errors.Is(capacityErr, promptkit.ErrCapacityExceeded) {
t.Fatalf("capacity error=%v, want ErrCapacityExceeded", capacityErr)
}
}
var discoveredCapacityErr *promptkit.CapacityError
if !errors.As(populatedCapacityErr, &discoveredCapacityErr) || discoveredCapacityErr != populatedCapacityErr {
t.Fatalf("populated capacity error is not discoverable: %v", populatedCapacityErr)
}
for _, unrelated := range []error{
promptkit.ErrInvalidConfig,
promptkit.ErrInvalidRequest,
promptkit.ErrLLMGenerate,
promptkit.ErrValidation,
} {
if errors.Is(promptkit.ErrCapacityExceeded, unrelated) ||
errors.Is(unrelated, promptkit.ErrCapacityExceeded) ||
errors.Is(populatedCapacityErr, unrelated) {
t.Fatalf("ErrCapacityExceeded aliases unrelated sentinel %v", unrelated)
}
}
}
type capacityRunResult struct {
result *promptkit.RunResult
err error
}
func runCapacityRequest(
engine *promptkit.Engine,
ctx context.Context,
request promptkit.RunRequest,
results chan<- capacityRunResult,
) {
result, err := engine.Run(ctx, request)
results <- capacityRunResult{result: result, err: err}
}
func newBackendCapacityEngine(
t *testing.T,
client promptkit.LLMClient,
limit int,
queueCapacity *int,
reader promptkit.ArtifactReader,
) *promptkit.Engine {
t.Helper()
promptFS := contractPromptFS("prompt", "profile", "message")
if reader != nil {
promptFS = contractInputPromptFS()
}
options := []promptkit.Option{
promptkit.WithPromptFS(promptFS, "."),
promptkit.WithBackend(promptkit.Backend{
ID: "limited",
Endpoint: "http://backend.example/v1",
ConcurrencyLimit: limit,
QueueCapacity: queueCapacity,
}),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile", BackendID: "limited", Model: "model",
}),
promptkit.WithLLMClient(client),
}
if reader != nil {
options = append(options, promptkit.WithArtifactReader(reader))
}
engine, err := promptkit.NewEngine(promptkit.Config{}, options...)
if err != nil {
t.Fatalf("construct capacity engine: %v", err)
}
return engine
}
func capacityInputRequest(endpoint string) promptkit.RunRequest {
return promptkit.RunRequest{
PromptID: "input-prompt",
Inputs: map[string]promptkit.ArtifactRef{
"input": promptkit.Inline("input"),
},
Execution: &promptkit.ExecutionTargetOverride{Endpoint: endpoint},
}
}
type capacityGateClient struct {
mu sync.Mutex
active int
peak int
calls int
started chan promptkit.GenerateRequest
release <-chan struct{}
}
func newCapacityGateClient(release <-chan struct{}, buffer int) *capacityGateClient {
return &capacityGateClient{
started: make(chan promptkit.GenerateRequest, buffer),
release: release,
}
}
func (c *capacityGateClient) Generate(
ctx context.Context,
request promptkit.GenerateRequest,
) (*promptkit.GenerateResponse, error) {
c.mu.Lock()
c.calls++
c.active++
if c.active > c.peak {
c.peak = c.active
}
c.mu.Unlock()
defer func() {
c.mu.Lock()
c.active--
c.mu.Unlock()
}()
c.started <- request
select {
case <-c.release:
return &promptkit.GenerateResponse{Content: "ok"}, nil
case <-ctx.Done():
return nil, ctx.Err()
}
}
func (c *capacityGateClient) snapshot() (active, peak, calls int) {
c.mu.Lock()
defer c.mu.Unlock()
return c.active, c.peak, c.calls
}
type capacityArtifactReader struct {
mu sync.Mutex
calls int
entered chan struct{}
release <-chan struct{}
}
func (r *capacityArtifactReader) Read(
ctx context.Context,
_ promptkit.ArtifactRef,
) (*promptkit.Artifact, error) {
r.mu.Lock()
r.calls++
r.mu.Unlock()
r.entered <- struct{}{}
select {
case <-r.release:
return &promptkit.Artifact{Body: []byte("input")}, nil
case <-ctx.Done():
return nil, ctx.Err()
}
}
func (r *capacityArtifactReader) callCount() int {
r.mu.Lock()
defer r.mu.Unlock()
return r.calls
}
func awaitCapacityRequest(
t *testing.T,
requests <-chan promptkit.GenerateRequest,
) promptkit.GenerateRequest {
t.Helper()
select {
case request := <-requests:
return request
case <-time.After(5 * time.Second):
t.Fatal("timed out waiting for client invocation")
return promptkit.GenerateRequest{}
}
}
func awaitCapacityRun(t *testing.T, results <-chan capacityRunResult) capacityRunResult {
t.Helper()
select {
case result := <-results:
return result
case <-time.After(5 * time.Second):
t.Fatal("timed out waiting for Run")
return capacityRunResult{}
}
}
func awaitCapacitySignal(t *testing.T, signal <-chan struct{}, name string) {
t.Helper()
select {
case <-signal:
case <-time.After(5 * time.Second):
t.Fatalf("timed out waiting for %s", name)
}
}
func capacityInt(value int) *int {
return &value
}
func closedCapacityChannel() <-chan struct{} {
channel := make(chan struct{})
close(channel)
return channel
}

40
capacity_error.go Normal file
View File

@@ -0,0 +1,40 @@
package promptkit
import (
"fmt"
"strings"
)
// CapacityError reports bounded admission rejected for a selected backend.
//
// Engine-produced values identify only rejection at Promptkit's bounded
// [Engine.Run] or [Engine.RunPrepared] admission boundary. BackendID is the
// normalized registered backend ID used for routing and capacity; endpoint
// overrides do not change it. Every engine-produced value is nonnil and has a
// nonblank BackendID. Provider errors, active-generation waiting, and caller
// cancellation are not represented by this type.
//
// Callers own returned values and may mutate BackendID without affecting engine
// state or another error. CapacityError and its default Go encoding have no
// stable JSON contract. Consumer-constructed values do not establish that an
// engine rejected work.
type CapacityError struct {
// BackendID is the normalized registered backend ID whose admission was
// rejected.
BackendID string
}
// Error returns diagnostic wording that is not a parsing contract. It is safe
// to call on a nil receiver or a value with a blank BackendID.
func (e *CapacityError) Error() string {
if e == nil || strings.TrimSpace(e.BackendID) == "" {
return ErrCapacityExceeded.Error()
}
return fmt.Sprintf("backend %q admission: %v", e.BackendID, ErrCapacityExceeded)
}
// Unwrap returns ErrCapacityExceeded so errors.Is and errors.As can be used
// together. It is safe to call on a nil receiver or a zero value.
func (e *CapacityError) Unwrap() error {
return ErrCapacityExceeded
}

View File

@@ -4,6 +4,7 @@ import (
"reflect" "reflect"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
) )
func toDomainRunRequest(req RunRequest) (domain.RunRequest, error) { func toDomainRunRequest(req RunRequest) (domain.RunRequest, error) {
@@ -15,6 +16,7 @@ func toDomainRunRequest(req RunRequest) (domain.RunRequest, error) {
PromptID: req.PromptID, PromptID: req.PromptID,
PromptVersion: req.PromptVersion, PromptVersion: req.PromptVersion,
ProfileID: req.ProfileID, ProfileID: req.ProfileID,
SessionID: req.SessionID,
APIKey: req.APIKey, APIKey: req.APIKey,
Inputs: toDomainArtifactRefMap(req.Inputs), Inputs: toDomainArtifactRefMap(req.Inputs),
Vars: copyStringMap(req.Vars), Vars: copyStringMap(req.Vars),
@@ -32,6 +34,7 @@ func fromDomainPreparedRun(prepared *domain.PreparedRun) *PreparedRun {
PromptVersion: prepared.PromptVersion, PromptVersion: prepared.PromptVersion,
PromptHash: prepared.PromptHash, PromptHash: prepared.PromptHash,
SelectedProfileID: prepared.SelectedProfileID, SelectedProfileID: prepared.SelectedProfileID,
SelectedBackendID: prepared.SelectedBackendID,
EffectiveModelParams: fromDomainExecutionTarget(prepared.EffectiveModelParams), EffectiveModelParams: fromDomainExecutionTarget(prepared.EffectiveModelParams),
OutputContract: fromDomainOutputContract(prepared.OutputContract), OutputContract: fromDomainOutputContract(prepared.OutputContract),
StructuredOutput: fromDomainStructuredOutputSpec(prepared.StructuredOutput), StructuredOutput: fromDomainStructuredOutputSpec(prepared.StructuredOutput),
@@ -57,8 +60,10 @@ func fromDomainRunResult(result *domain.RunResult) *RunResult {
PromptID: result.PromptID, PromptID: result.PromptID,
PromptVersion: result.PromptVersion, PromptVersion: result.PromptVersion,
PromptHash: result.PromptHash, PromptHash: result.PromptHash,
SessionID: result.SessionID,
RenderedPromptHash: result.RenderedPromptHash, RenderedPromptHash: result.RenderedPromptHash,
SelectedProfileID: result.SelectedProfileID, SelectedProfileID: result.SelectedProfileID,
SelectedBackendID: result.SelectedBackendID,
ModelName: result.ModelName, ModelName: result.ModelName,
Endpoint: result.Endpoint, Endpoint: result.Endpoint,
EffectiveModelParams: fromDomainExecutionTarget(result.EffectiveModelParams), EffectiveModelParams: fromDomainExecutionTarget(result.EffectiveModelParams),
@@ -131,7 +136,7 @@ func toDomainExecutionTargetOverride(override *ExecutionTargetOverride) (*domain
if override == nil { if override == nil {
return nil, nil return nil, nil
} }
extraParams, err := copyPublicJSONMap(override.ExtraParams) extraParams, err := jsonvalue.CopyMap(override.ExtraParams)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -143,7 +148,7 @@ func toDomainExecutionTargetOverride(override *ExecutionTargetOverride) (*domain
TopP: copyFloat64Ptr(override.TopP), TopP: copyFloat64Ptr(override.TopP),
TimeoutSeconds: copyIntPtr(override.TimeoutSeconds), TimeoutSeconds: copyIntPtr(override.TimeoutSeconds),
ServiceTier: override.ServiceTier, ServiceTier: override.ServiceTier,
ReasoningEffort: override.ReasoningEffort, ReasoningEffort: copyStringPtr(override.ReasoningEffort),
APIKeyEnv: override.APIKeyEnv, APIKeyEnv: override.APIKeyEnv,
ExtraParams: extraParams, ExtraParams: extraParams,
}, nil }, nil
@@ -151,6 +156,7 @@ func toDomainExecutionTargetOverride(override *ExecutionTargetOverride) (*domain
func fromDomainExecutionTarget(target domain.ExecutionTarget) ExecutionTarget { func fromDomainExecutionTarget(target domain.ExecutionTarget) ExecutionTarget {
return ExecutionTarget{ return ExecutionTarget{
BackendID: target.BackendID,
Endpoint: target.Endpoint, Endpoint: target.Endpoint,
Model: target.Model, Model: target.Model,
Temperature: target.Temperature, Temperature: target.Temperature,
@@ -164,6 +170,40 @@ func fromDomainExecutionTarget(target domain.ExecutionTarget) ExecutionTarget {
} }
} }
func fromDomainProfileInspection(inspection *domain.ProfileInspection) *ProfileInspection {
if inspection == nil {
return nil
}
return &ProfileInspection{
ProfileID: inspection.ProfileID,
EffectiveModelParams: fromDomainExecutionTarget(inspection.EffectiveModelParams),
APIKeyRequired: inspection.APIKeyRequired,
}
}
func fromDomainPromptInspection(inspection *domain.PromptInspection) *PromptInspection {
if inspection == nil {
return nil
}
inputs := make([]PromptInputDefinition, len(inspection.Inputs))
for i, input := range inspection.Inputs {
inputs[i] = PromptInputDefinition{
Name: input.Name,
Required: input.Required,
ContentType: input.ContentType,
Description: input.Description,
}
}
return &PromptInspection{
PromptID: inspection.PromptID,
PromptVersion: inspection.PromptVersion,
PromptHash: inspection.PromptHash,
DefaultProfileID: inspection.DefaultProfileID,
Inputs: inputs,
OutputContract: fromDomainOutputContract(inspection.OutputContract),
}
}
func fromDomainExecutionTargetPresence(presence domain.ExecutionTargetPresence) ExecutionTargetPresence { func fromDomainExecutionTargetPresence(presence domain.ExecutionTargetPresence) ExecutionTargetPresence {
return ExecutionTargetPresence{ return ExecutionTargetPresence{
Temperature: presence.Temperature, Temperature: presence.Temperature,
@@ -396,6 +436,14 @@ func copyFloat64Ptr(src *float64) *float64 {
return &v return &v
} }
func copyStringPtr(src *string) *string {
if src == nil {
return nil
}
v := *src
return &v
}
func copyIntPtr(src *int) *int { func copyIntPtr(src *int) *int {
if src == nil { if src == nil {
return nil return nil

40
doc.go
View File

@@ -2,21 +2,29 @@
// prompt-defined LLM workflows. // prompt-defined LLM workflows.
// //
// Applications construct an [Engine] with [NewEngine], select filesystem or // Applications construct an [Engine] with [NewEngine], select filesystem or
// in-memory sources with options, and call [Engine.Prepare] or [Engine.Run]. // in-memory sources and optional engine-scoped [Backend] registrations, and
// Concrete repositories, validators, and the built-in OpenAI-compatible client // call [Engine.InspectPrompt], [Engine.InspectProfile], [Engine.Prepare],
// remain internal implementation details. // [Engine.PrepareExecution], [Engine.Run], or [Engine.RunPrepared]. Concrete
// registries, repositories, validators, and the built-in OpenAI-compatible
// client remain internal implementation details.
// //
// # Concurrency and ownership // # Concurrency and ownership
// //
// An Engine supports concurrent Prepare and Run calls. An injected [LLMClient] // An Engine supports concurrent InspectPrompt, InspectProfile, Prepare,
// or [ArtifactReader] can therefore receive concurrent calls and must be safe // PrepareExecution, Run, and RunPrepared calls. Engine-local backend policies
// for that use. // bound admitted Run and RunPrepared calls and model generations where
// configured, while different backend pools and unlimited backends continue
// independently. An injected [LLMClient] or [ArtifactReader] can therefore
// still receive concurrent calls and must be safe for that use.
// //
// NewEngine copies in-memory profiles. Prepare and Run copy request maps, // NewEngine copies in-memory profiles and backend definitions. Prepare,
// slices, pointer values, and JSON-compatible extra parameters before using // PrepareExecution, and Run copy request maps, slices, pointer values, and
// them. Returned values and values passed to extension interfaces are likewise // JSON-compatible extra parameters before using them. InspectPrompt and
// isolated from engine state. Callers own those copies and may mutate them // InspectProfile return copied inspection values. Returned values and values
// after the call that supplied or returned them. // passed to extension interfaces are likewise isolated from engine state.
// Callers own those copies and may mutate them after the call that supplied or
// returned them. Returned structured errors are likewise caller-owned and may
// be mutated without affecting engine state or another error.
// //
// # Security and sensitive data // # Security and sensitive data
// //
@@ -42,10 +50,12 @@
// [GenerateResponse], [ExecutionTargetPresence], and the string value types // [GenerateResponse], [ExecutionTargetPresence], and the string value types
// used by those values. // used by those values.
// //
// Construction values, including [Config], [RunRequest], [ArtifactRef], // Construction, inspection, handle, and error values, including [Config],
// [ExecutionTargetOverride], [Profile], and // [Backend], [RunRequest], [ArtifactRef], [ExecutionTargetOverride], [Profile],
// [OpenAICompatibleProfileConfig], do not have stable JSON representations. // [OpenAICompatibleProfileConfig], [ProfileInspection],
// Direct API keys are nevertheless excluded from JSON for every public value. // [PromptInputDefinition], [PromptInspection], [PreparedExecution], and
// [CapacityError], do not have stable JSON representations. Direct API keys
// are nevertheless excluded from JSON for every public value.
// //
// JSON timestamps use time.Time's RFC 3339 encoding and are omitted when zero. // JSON timestamps use time.Time's RFC 3339 encoding and are omitted when zero.
// PreparedRun and RunResult durations are encoded as integer milliseconds in // PreparedRun and RunResult durations are encoded as integer milliseconds in

View File

@@ -33,18 +33,44 @@ engine, err := promptkit.NewEngine(promptkit.Config{
}) })
``` ```
Options support single-file or `fs.FS` sources, in-memory profiles, and Options support single-file or `fs.FS` sources, in-memory profiles,
injected artifact or model clients. Consult the engine-scoped backends, and injected artifact or model clients. Consult the
[constructor and option GoDoc](../../engine.go) for composition, precedence, [constructor and option GoDoc](../../engine.go) for composition, precedence,
validation, and default transport behavior. Source discovery, format validation, and default transport behavior. Source discovery, format
validation, and profile precedence are defined by the validation, and profile precedence are defined by the
[framework format reference](../formats.md). [framework format reference](../formats.md).
## Inspect A Prompt Before Preparation
Use [`Engine.InspectPrompt`](../../engine.go) to check one configured prompt's
declared inputs and output workflow without creating placeholder inputs or
resolving a profile:
```go
inspection, err := engine.InspectPrompt(ctx, "meeting.summary", "")
if err != nil {
return err
}
for _, input := range inspection.Inputs {
// Compare the declared input with application configuration.
}
```
Use this configuration-time boundary when the application needs only the
declared prompt interface. Use `InspectProfile` separately when it must also
check a configured profile. Use `Prepare` when it needs inputs, schemas, or
rendered messages, and use prepared execution when that work must remain tied
to later execution. The method's [GoDoc](../../engine.go) owns exact fields,
hash, ownership, and error semantics.
## Prepare Without Model Execution ## Prepare Without Model Execution
[`Engine.Prepare`](../../engine.go) resolves the selected prompt and profile, [`Engine.Prepare`](../../engine.go) resolves the selected prompt and profile,
loads inputs and any structured-output schema, and renders messages without loads inputs and any structured-output schema, and renders messages without
calling a model client: calling a model client. Choose it when the prepared value is the final
inspection or persistence result and no later execution must be tied to that
exact snapshot:
```go ```go
prepared, err := engine.Prepare(ctx, promptkit.RunRequest{ prepared, err := engine.Prepare(ctx, promptkit.RunRequest{
@@ -61,12 +87,48 @@ shows a complete runnable setup with a prompt file, in-memory profile, and
inline input. Exact request requirements and prepared-result fields belong to inline input. Exact request requirements and prepared-result fields belong to
the [`RunRequest` and `PreparedRun` GoDoc](../../types.go). the [`RunRequest` and `PreparedRun` GoDoc](../../types.go).
## Prepare Now And Execute The Same Snapshot Later
Use [`Engine.PrepareExecution`](../../engine.go) when an application must
inspect or persist preflight details before deciding whether to start model
work, while ensuring that later execution uses those exact rendered messages,
inputs, target settings, and validation resources:
```go
preparedExecution, err := engine.PrepareExecution(ctx, promptkit.RunRequest{
PromptID: "meeting.summary",
Inputs: map[string]promptkit.ArtifactRef{
"note": promptkit.Inline("Synthetic meeting notes"),
},
})
if err != nil {
return err
}
defer preparedExecution.Discard()
details := preparedExecution.Details()
// Inspect or persist an application-selected safe subset of details.
result, err := engine.RunPrepared(ctx, preparedExecution)
```
Preparation does not call the model or reserve backend capacity.
`RunPrepared` executes from the retained snapshot rather than reloading
consumer sources. The handle is opaque in-process state, while `Details`
contains rendered content and remains subject to the application's data
handling policy. The
[`PreparedExecution` and method GoDoc](../../prepared_execution.go) and
[engine operation GoDoc](../../engine.go) own exact lifecycle, engine-binding,
credential, cancellation, timing, and error semantics.
## Execute And Validate ## Execute And Validate
[`Engine.Run`](../../engine.go) performs the same preparation, invokes the [`Engine.Run`](../../engine.go) performs the same preparation, invokes the
configured model client, classifies the generated artifact, and validates the configured model client, classifies the generated artifact, and validates the
content. A completed content check may return `ValidationFailed` in the result; content in one call. Choose it when the application does not need a preflight
an operational inability to validate returns an error. boundary tied to the eventual execution. A completed content check may return
`ValidationFailed` in the result; an operational inability to validate returns
an error.
The maintained The maintained
[offline execution example](../../examples/go-library/run/main.go) injects a [offline execution example](../../examples/go-library/run/main.go) injects a
@@ -97,6 +159,175 @@ For programmatic profiles,
[`OpenAICompatibleProfile`](../../profiles.go) converts ordinary [`OpenAICompatibleProfile`](../../profiles.go) converts ordinary
OpenAI-compatible settings into a value accepted by `WithProfiles`. OpenAI-compatible settings into a value accepted by `WithProfiles`.
### Inspect A Profile Before Prompt Work
Use [`Engine.InspectProfile`](../../engine.go) to validate one configured
profile without constructing a synthetic prompt or placeholder inputs. It
resolves the profile's effective target but does not prepare or execute a
prompt:
```go
inspection, err := engine.InspectProfile(ctx, profileID)
if err != nil {
return err
}
target := inspection.EffectiveModelParams
if target.APIKeyEnv != "" {
// Apply application policy for the named environment variable.
} else if inspection.APIKeyRequired {
// Arrange a direct credential before later execution.
}
```
Use this configuration-time boundary when only the profile and its target need
checking. Use `Prepare` when the application also needs prompt, input, schema,
or rendering work; use prepared execution when that work must remain tied to a
later execution. Inspection reports credential requirements but leaves the
timing of credential enforcement to the application. The method's
[GoDoc](../../engine.go) owns its exact result and error contract.
### Set A Per-Run Session And Reasoning
Supply a direct session ID when one prompt should be correlated with a
consumer-managed conversation or workflow without changing prompt variables:
```go
reasoning := "high"
result, err := engine.Run(ctx, promptkit.RunRequest{
PromptID: "meeting.summary",
SessionID: "conversation-42",
Inputs: map[string]promptkit.ArtifactRef{
"note": promptkit.Inline("Synthetic meeting notes"),
},
Execution: &promptkit.ExecutionTargetOverride{
ReasoningEffort: &reasoning,
},
})
```
A nil reasoning pointer inherits the selected profile, a pointer to a
nonblank string replaces it, and a pointer to a blank string disables
reasoning for that run. Session IDs are correlation metadata, not credentials;
use stable, non-secret values that are safe to expose to collaborators and
providers. The
[`RunRequest` and `ExecutionTargetOverride` GoDoc](../../types.go) owns the
exact normalization, precedence, error, copying, and exposure contract.
### Configure A Local OpenAI-Compatible Endpoint
Choose the smallest configuration that fits how the endpoint will be reused.
#### Use An Endpoint-Only Profile
Put the endpoint directly on an in-memory profile when only that profile needs
it and shared backend identity or capacity policy is unnecessary:
```go
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: "prompts",
},
promptkit.WithProfiles(promptkit.Profile{
ID: "local-summary",
Endpoint: "http://localhost:8000/v1",
Model: "example-model",
}),
)
```
Endpoint-only profiles have an empty backend ID and remain unrestricted by
backend capacity policy.
#### Use The Conventional Local Backend
Use `LocalBackend` when profiles should share the conventional `local`
identity, endpoint, and concurrency limit:
```go
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: "prompts",
},
promptkit.WithBackend(
promptkit.LocalBackend("http://localhost:8000/v1", 2),
),
promptkit.WithProfiles(promptkit.Profile{
ID: "local-summary",
BackendID: promptkit.BackendLocal,
Model: "example-model",
}),
)
```
The helper is explicit: it does not pre-register a backend or read environment
variables. Supplying a positive limit leaves queue capacity omitted, so normal
backend registration selects the existing default waiting capacity of 1024.
The returned value still enters the engine through `WithBackend`.
#### Configure A Complete Backend
Use a keyed `Backend` value for authentication, extra request parameters, an
explicit queue capacity, a custom ID, or multiple local endpoints:
```go
noWaiting := 0
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: "prompts",
},
promptkit.WithBackend(promptkit.Backend{
ID: "local-gpu",
Endpoint: "http://gpu-host:8000/v1",
APIKeyEnv: "LOCAL_GPU_API_KEY",
ExtraParams: map[string]any{"provider_option": "enabled"},
ConcurrencyLimit: 2,
QueueCapacity: &noWaiting,
}),
promptkit.WithProfiles(promptkit.Profile{
ID: "gpu-summary",
BackendID: "local-gpu",
Model: "example-model",
}),
)
```
Use distinct custom IDs when registering multiple local endpoints.
Registrations belong to one engine and custom IDs cannot replace built-ins.
The [`Backend`, `LocalBackend`, and `WithBackend` GoDoc](../../backends.go)
defines exact construction, validation, copying, uniqueness, concurrency, and
request-default behavior.
Both file-backed and in-memory profiles select a registration through
`backend` or `Profile.BackendID`. Profile and request endpoint overrides retain
that routing and capacity identity. `PreparedRun.SelectedBackendID`,
`RunResult.SelectedBackendID`, and the effective `ExecutionTarget.BackendID`
expose it to consumers and injected model clients. Endpoint-only profiles
remain supported and expose an empty backend ID.
### Limit Backend Concurrency
Set `Backend.ConcurrencyLimit` when a backend needs protection from too many
simultaneous model calls. Leaving `QueueCapacity` nil, as in the local-backend
example above, selects the default waiting capacity of 1024.
To accept no waiting backlog beyond the active calls, provide an explicit
zero:
```go
noWaiting := 0
backend := promptkit.Backend{
ID: "local-gpu",
Endpoint: "http://gpu-host:8000/v1",
ConcurrencyLimit: 2,
QueueCapacity: &noWaiting,
}
```
The pointer distinguishes an explicit zero from omission. Keep using keyed
`Backend` literals so additive configuration fields remain source-compatible.
Capacity belongs to one engine and the selected backend ID; endpoint-only
profiles and custom backends without a configured limit remain unrestricted.
Exact validation, defaulting, ownership, and concurrency semantics belong to
the [`Backend` GoDoc](../../backends.go).
## Credentials ## Credentials
File-backed profiles name an environment variable; in-memory profiles can File-backed profiles name an environment variable; in-memory profiles can
@@ -137,7 +368,30 @@ distinguish invalid construction, invalid requests, absent sources,
source-loading failures, collaborator failures, and operational validation source-loading failures, collaborator failures, and operational validation
failures. Specific request conditions may also match the broader failures. Specific request conditions may also match the broader
`ErrInvalidRequest`, and injected collaborator identities are preserved where `ErrInvalidRequest`, and injected collaborator identities are preserved where
documented. documented. Invalid or duplicate backend registrations match
`ErrInvalidConfig`; selecting an unknown backend matches `ErrProfileLoad`.
When a limited backend has admitted all active and waiting calls, handle
`ErrCapacityExceeded` separately from request errors and provider failures:
```go
result, err := engine.Run(ctx, request)
if errors.Is(err, promptkit.ErrCapacityExceeded) {
var capacityErr *promptkit.CapacityError
if errors.As(err, &capacityErr) {
// Record capacityErr.BackendID using application-owned diagnostics.
}
// Apply application policy: shed work, report overload, or retry later.
}
```
A rejected call returns no partial result and does not invoke the model
client. Promptkit does not prescribe retries or map this error to an HTTP
status; those choices remain with the consuming application. The
[`CapacityError` GoDoc](../../capacity_error.go) owns the exact typed-error
contract, while the [`Engine.Run` and error GoDoc](../../engine.go) owns broad
error and cancellation identities.
## Application Boundary ## Application Boundary

View File

@@ -58,6 +58,11 @@ When a request omits a version, the selected prompt ID must identify exactly
one definition. When it supplies a version, the ID and version pair must be one definition. When it supplies a version, the ID and version pair must be
unique. unique.
Exact prompt inspection uses this same configured source, strict decoding,
referenced content-file resolution, and ID/version selection. It reports the
selected definition's declared metadata without changing the prompt format or
executing the definition.
### Inputs ### Inputs
Each `inputs` item has these fields: Each `inputs` item has these fields:
@@ -91,7 +96,8 @@ of a named input. Missing variables and input references are errors.
The optional `session_id` uses the same template data and input helper. Its The optional `session_id` uses the same template data and input helper. Its
rendered value is trimmed, omitted when empty, and limited to 256 Unicode code rendered value is trimmed, omitted when empty, and limited to 256 Unicode code
points. points. A nonblank direct request session ID bypasses this template completely;
a blank direct value leaves the template behavior unchanged.
### Cache Control ### Cache Control
@@ -136,7 +142,7 @@ A profile supplies model execution settings:
```yaml ```yaml
id: local-summary id: local-summary
endpoint: http://localhost:8000/v1 backend: openrouter
model: example-model model: example-model
temperature: 0.2 temperature: 0.2
max_tokens: 500 max_tokens: 500
@@ -144,7 +150,6 @@ top_p: 0.95
timeout_seconds: 90 timeout_seconds: 90
service_tier: flex service_tier: flex
reasoning_effort: medium reasoning_effort: medium
api_key_env: EXAMPLE_API_KEY
extra_params: extra_params:
provider_option: enabled provider_option: enabled
``` ```
@@ -152,7 +157,8 @@ extra_params:
| Field | Required | Meaning | | Field | Required | Meaning |
| --- | --- | --- | | --- | --- | --- |
| `id` | yes | Non-empty profile identifier. IDs must be unique within one source. | | `id` | yes | Non-empty profile identifier. IDs must be unique within one source. |
| `endpoint` | yes | Non-empty OpenAI-compatible base URL, including an API version path when required. | | `backend` | unless `endpoint` is present | Backend registry ID. It is trimmed and registry membership is checked when the profile is prepared or inspected. |
| `endpoint` | unless `backend` is present | Non-empty OpenAI-compatible base URL, including an API version path when required. When both connection fields are present, this overrides the backend endpoint without changing backend identity. |
| `model` | yes | Non-empty provider model name. | | `model` | yes | Non-empty provider model name. |
| `temperature` | no | Number from 0 through 2. | | `temperature` | no | Number from 0 through 2. |
| `max_tokens` | no | Integer zero or greater. | | `max_tokens` | no | Integer zero or greater. |
@@ -166,6 +172,13 @@ extra_params:
Raw `api_key` is prohibited in profile YAML. Store only an environment Raw `api_key` is prohibited in profile YAML. Store only an environment
variable name in `api_key_env`. variable name in `api_key_env`.
Promptkit does not infer a backend from a model or endpoint. Endpoint-only
profiles remain supported and have no effective backend ID.
The engine always provides the built-in `openrouter` ID. Consumers can add
engine-scoped IDs with
[`WithBackend`](../backends.go); exact registration validation belongs to its
GoDoc.
`extra_params` accepts null, booleans, finite numbers, strings, arrays, and `extra_params` accepts null, booleans, finite numbers, strings, arrays, and
objects with string keys. Keys must be non-empty. With the built-in client, objects with string keys. Keys must be non-empty. With the built-in client,
they also cannot collide with the standard fields listed in the they also cannot collide with the standard fields listed in the
@@ -176,8 +189,9 @@ they also cannot collide with the standard fields listed in the
Execution settings resolve in this order: Execution settings resolve in this order:
1. framework defaults; 1. framework defaults;
2. the selected profile; and 2. the selected backend, when the profile names one;
3. request `ExecutionTargetOverride` values. 3. the selected profile; and
4. request `ExecutionTargetOverride` values.
The framework defaults are: The framework defaults are:
@@ -194,15 +208,25 @@ explicit zero is preserved. In particular, an explicit request
`timeout_seconds` of zero disables the per-generation deadline while leaving `timeout_seconds` of zero disables the per-generation deadline while leaving
the caller context and transport timeout intact. the caller context and transport timeout intact.
Non-empty request strings replace profile strings. A non-empty request Non-empty profile strings replace backend defaults, and non-empty request
`ExtraParams` map replaces the profile map rather than merging keys. strings replace both. Request reasoning is the exception: a nil
`ReasoningEffort` pointer inherits the profile, a pointer to a nonblank string
trims and replaces it, and a pointer to a blank string clears it. Backend
identity is retained when either layer overrides the endpoint, so the override
also retains any engine-local capacity policy configured for that backend.
Capacity configuration belongs to the Go
[`Backend` API](../backends.go), not prompt or profile YAML. A non-empty
`extra_params` map at each layer replaces the entire lower-precedence map
rather than merging keys.
The [outbound integration contract](integrations/openai-compatible-chat.md) The [outbound integration contract](integrations/openai-compatible-chat.md)
defines how the effective settings are serialized. defines how the effective settings are serialized.
### Source And Profile Precedence ### Source And Profile Precedence
An explicit request profile ID takes precedence over the prompt's An explicit request profile ID takes precedence over the prompt's
`default_profile`. If neither is present, preparation fails. `default_profile`. If neither is present, preparation fails. Exact profile
inspection instead takes one explicit profile ID and does not use a prompt
default.
Profile sources resolve matching IDs in this order: Profile sources resolve matching IDs in this order:
@@ -214,11 +238,15 @@ A higher-precedence source falls back only when the profile is absent. An
invalid matching profile is an error and does not fall back. In-memory invalid matching profile is an error and does not fall back. In-memory
`Profile` values follow the same ranges as YAML profiles. They use `Profile` values follow the same ranges as YAML profiles. They use
`APIKeyRequired` for request-scoped credentials instead of `api_key_env`. `APIKeyRequired` for request-scoped credentials instead of `api_key_env`.
Preparation and exact profile inspection use this same source precedence.
## Built-In Profile Catalog ## Built-In Profile Catalog
Built-ins use the OpenRouter-compatible endpoint and Every built-in selects the `openrouter` backend. The engine's built-in backend
`OPENROUTER_API_KEY`. A custom or in-memory profile with the same ID takes registry supplies `https://openrouter.ai/api/v1` and the environment-variable
name `OPENROUTER_API_KEY`, so individual profiles contain only model and
generation settings. Built-in profile files do not repeat those connection
values. A custom or in-memory profile with the same profile ID takes
precedence. precedence.
| Provider | ID | Model | | Provider | ID | Model |
@@ -270,6 +298,10 @@ prompt, profile, schema, or example files:
- a request can provide a direct `APIKey` or override `APIKeyEnv`; and - a request can provide a direct `APIKey` or override `APIKeyEnv`; and
- a direct request key takes precedence over environment lookup. - a direct request key takes precedence over environment lookup.
After a direct request key, the credential-source precedence is request
`APIKeyEnv`, profile `api_key_env`, then the backend default. An in-memory
profile with `APIKeyRequired` clears an inherited backend environment name and
requires a direct key unless the request explicitly supplies `APIKeyEnv`.
Promptkit validates required credential availability during preparation. Promptkit validates required credential availability during preparation.
Direct keys are excluded from JSON results and redacted by public string Direct keys are excluded from JSON results and redacted by public string
formatters. Environment-variable names may appear in prepared metadata, but formatters. Environment-variable names may appear in prepared metadata, but

View File

@@ -13,10 +13,15 @@ that produce these outbound settings.
## Endpoint And Method ## Endpoint And Method
Generation sends an HTTP `POST` with `Content-Type: application/json`. Generation sends an HTTP `POST` with `Content-Type: application/json`.
A non-empty endpoint from the execution target overrides the client's Before the client is called, the engine resolves framework, backend, profile,
configured base URL. After trailing slashes are removed, and request values into one execution target. A non-empty endpoint from that
`/chat/completions` is appended. Generation fails before sending when neither target overrides the client's configured base URL. After trailing slashes are
source supplies an endpoint. removed, `/chat/completions` is appended. Generation fails before sending when
neither source supplies an endpoint.
The target's backend ID is routing metadata for prepared values, results, and
injected clients. The built-in client does not derive the URL from that ID and
does not serialize it in the provider request.
## Authentication ## Authentication
@@ -26,6 +31,12 @@ the client reads that variable and requires a non-empty value. The selected
key is sent as `Authorization: Bearer <key>`. No authorization header is sent key is sent as `Authorization: Bearer <key>`. No authorization header is sent
when neither mechanism is configured. when neither mechanism is configured.
The target contains the already resolved environment-variable name: an
explicit request override takes precedence over profile metadata, which takes
precedence over the backend default. Only the name reaches prepared metadata;
the environment value is read just before the provider call and is never added
to the JSON body.
## Request Body ## Request Body
The request body always contains `model` and `messages`. The execution The request body always contains `model` and `messages`. The execution
@@ -36,20 +47,24 @@ Each ordinary message contains its `role` and string `content`. A
cache-controlled message instead uses a text content block containing `type`, cache-controlled message instead uses a text content block containing `type`,
`text`, and `cache_control`; an empty cache-control TTL is omitted. `text`, and `cache_control`; an empty cache-control TTL is omitted.
A non-empty session ID is trimmed, checked against the internal domain limit, The effective direct or prompt-rendered session ID is trimmed, limited to 256
and sent as top-level `session_id`. It is not sent as a session header. Unicode code points, and sent when nonempty as top-level `session_id`. It is
never also sent as a session header.
The client conditionally includes: The client conditionally includes:
- `temperature`, `max_tokens`, and `top_p` when non-zero or explicitly - `temperature`, `max_tokens`, and `top_p` when non-zero or explicitly
present; present;
- non-empty `service_tier` and `reasoning_effort`; and - non-empty `service_tier` and effective `reasoning_effort`; an explicitly
disabled reasoning setting is empty and therefore omitted; and
- `response_format` for JSON Schema structured output, including its name, - `response_format` for JSON Schema structured output, including its name,
strict flag, and schema document. strict flag, and schema document.
Extra parameters are merged directly into the top-level body after JSON The engine resolves backend, profile, and request extra-parameter maps by
serialization is verified. Empty keys and collisions with these reserved whole-map replacement rather than key merging. The resulting effective map is
fields are rejected before any provider call: then merged directly into the top-level body after JSON serialization is
verified. Empty keys and collisions with these reserved fields are rejected
before any provider call:
- `model` - `model`
- `session_id` - `session_id`
@@ -61,6 +76,9 @@ fields are rejected before any provider call:
- `reasoning_effort` - `reasoning_effort`
- `response_format` - `response_format`
`backend_id`, `api_key_env`, and resolved credential values are not provider
request fields.
## Response Handling ## Response Handling
Any 2xx response is decoded as an OpenAI-compatible chat response. The client Any 2xx response is decoded as an OpenAI-compatible chat response. The client

113
docs/internal/capacity.md Normal file
View File

@@ -0,0 +1,113 @@
# Internal Capacity Management
## Purpose
This document describes the implemented engine-local capacity coordination in
`internal/capacity`. The [architecture policy](../policy/architecture.md) owns
component boundaries, the [backend GoDoc](../../backends.go) owns exact public
configuration semantics, and the
[internal runner document](runner.md) owns orchestration around admission.
Capacity scheduling is outside the provider wire contract. It does not add
fields to execution targets, generated requests, prompt or profile YAML, or
stable JSON values.
## Construction And Pool Lifecycle
Each root `NewEngine` call obtains a normalized capacity-policy snapshot from
its immutable backend registry and constructs a new `Manager`. The manager
creates one pool for each limited backend ID. It has no package-global mutable
state, background workers, shutdown protocol, or persistence, so engines with
the same registrations still have independent capacity.
Unlimited registered backends and endpoint-only profiles have no pool. Their
admission and generation calls take the unrestricted fast path. An endpoint
override does not change the selected backend ID and therefore does not change
the pool.
One pool owns immutable active and total limits plus mutex-protected admission
count, active count, and ordered waiter list. Pool state exists only for the
lifetime of its engine.
## Bounded Execution Admission
For ordinary `Run`, the runner asks the manager to admit after resolving the
prompt, profile, selected backend, effective execution target, credentials, and
output contract, but before schema loading, artifact loading, or rendering.
`PrepareExecution` performs no admission. `RunPrepared` claims its handle,
rechecks credential availability, and then asks the manager to admit the
frozen backend before generation.
Admission is immediate: a limited pool either reserves a slot or returns only
the internal `ErrCapacityExceeded` identity. The runner attaches the selected
backend identity at its use-case boundary, and the root facade translates that
typed value without treating it as an invalid request or generation failure.
The total admitted bound is the active-generation limit plus its configured
waiting capacity. The returned release function is idempotent. The runner
defers it as soon as admission succeeds. An ordinary run holds the lease across
remaining preparation, initial generation, validation, every repair attempt,
and all failure or cancellation exits. Prepared execution holds the normal
lease across generation, validation, every internal repair attempt, and all
execution exits. A repair is part of its original admission and does not
reserve another bounded slot.
## FIFO Generation Permits
`NewClient` wraps the engine's selected internal model client after public
client adaptation or built-in client construction. Initial generation and the
default repairer receive the same wrapper.
For each `Generate` call, the wrapper selects a pool from the request's
effective backend ID. An unlimited call passes directly to the next client. A
limited call acquires an active permit, invokes the next client, and defers
permit release so ordinary returns and panic unwinding both restore capacity.
Preparation and validation never hold an active permit.
When all active permits are occupied, calls join a mutex-protected FIFO waiter
list. Releasing a permit transfers it directly to the oldest remaining waiter
before making it generally available. Pools do not order work relative to
other backend IDs.
The wrapper passes generation requests, responses, and collaborator errors
through unchanged. It owns scheduling only; the concrete model client remains
responsible for provider transport behavior.
## Cancellation And Release
Admission checks the caller context before reserving a slot. A call canceled
while waiting for an active permit removes its waiter under the same pool lock
used to grant permits. If cancellation removes the waiter first, the wrapped
client is not invoked. If a concurrent grant wins first, the call owns the
permit and invokes the client with the original context, allowing the client
to observe cancellation normally.
This grant-or-cancel decision prevents lost and double-released permits.
Admission leases and active permits are released after success, collaborator
errors, validation failures, cancellation, and panic unwinding. Canceled
waiters are unlinked so their contexts and requests are not retained by the
pool.
## Test Ownership
The [manager tests](../../internal/capacity/manager_test.go) own policy
validation, bounded admission, idempotent release, context handling, and
unlimited admission. The
[client tests](../../internal/capacity/client_test.go) own peak enforcement,
FIFO transfer, canceled-waiter removal, grant/cancel races, independent pools,
unlimited calls, passthrough behavior, and panic release.
The [runner tests](../../internal/usecase/runner_test.go) own ordinary early
admission, lease lifetime, failure release, and shared initial/repair
scheduling. The
[prepared-execution use-case tests](../../internal/usecase/prepared_execution_test.go)
own deferred admission, credential ordering, and prepared-execution lease
release. The
[external package capacity tests](../../capacity_contract_test.go) own the
assembled public-engine behavior for configured limits, capacity errors,
endpoint identity, engine independence, and injected clients. The
[prepared-execution contract tests](../../prepared_execution_contract_test.go)
own the public prepared-capacity boundary. The
[root error-boundary tests](../../errors_internal_test.go) own preservation of
the public generation category and context identity when generation is
canceled.

View File

@@ -20,6 +20,11 @@ orchestration. `OpenAICompatibleClient` is the built-in implementation. It
uses internal domain values for rendered prompts, execution targets, uses internal domain values for rendered prompts, execution targets,
structured output, responses, and token usage. structured output, responses, and token usage.
The runner supplies a fully resolved target after applying backend, profile,
and request precedence. The client uses its endpoint, credential metadata,
generation fields, and extra parameters. `BackendID` remains routing metadata
for the generation boundary and is not mapped into the provider payload.
Construction validates the configured base URL and clones any supplied Construction validates the configured base URL and clones any supplied
`http.Client` so Promptkit can apply its timeout default without mutating the `http.Client` so Promptkit can apply its timeout default without mutating the
caller's client. Generation then: caller's client. Generation then:
@@ -31,9 +36,28 @@ caller's client. Generation then:
5. performs the outbound request under the applicable deadlines; and 5. performs the outbound request under the applicable deadlines; and
6. decodes the first response choice and token usage. 6. decodes the first response choice and token usage.
`internal/llm` owns the set of reserved OpenAI-compatible request fields used
when validating extra parameters. Backend registration consumes the same rule
without making the model client depend on registry configuration.
The implementation has no retry loop, tool-call support, provider catalog, The implementation has no retry loop, tool-call support, provider catalog,
inbound HTTP behavior, or durable session store. inbound HTTP behavior, or durable session store.
## Prepared Generation
For [`RunPrepared`](../../engine.go), the runner supplies the model client with
the target, rendered messages, and structured-output constraint retained by
executable preparation. Execution does not reopen or rerender consumer
sources.
Before backend admission, the runner rechecks that the frozen credential
environment-variable name is available. The handle does not retain the
environment value; the model client resolves the value visible when generation
begins. A direct request key remains in private execution state only until the
claimed execution finishes or an unclaimed handle is discarded. Exact public
ownership and redaction semantics belong to the
[`PreparedExecution` GoDoc](../../prepared_execution.go).
## Failure Categories ## Failure Categories
The package preserves distinct error identities for invalid client The package preserves distinct error identities for invalid client
@@ -51,5 +75,7 @@ The
[OpenAI-compatible client tests](../../internal/llm/openai_compatible_client_test.go) [OpenAI-compatible client tests](../../internal/llm/openai_compatible_client_test.go)
own configuration, client cloning, deterministic deadline precedence, own configuration, client cloning, deterministic deadline precedence,
authentication, request and response mapping, malformed data, error identity, authentication, request and response mapping, malformed data, error identity,
cancellation, and response-body suppression. They use local test servers and cancellation, and response-body suppression. The root transport contract test
test transports; the default suite makes no live or paid provider requests. also verifies that resolved backend settings reach this client without
serializing backend identity. All use local test servers or test transports;
the default suite makes no live or paid provider requests.

View File

@@ -11,20 +11,23 @@ contributor workflow and validation.
| Component | Implemented responsibility | References | | Component | Implemented responsibility | References |
| --- | --- | --- | | --- | --- | --- |
| Root `promptkit` package | Provides the supported engine facade, source and injection options, public request and result values, built-in profile construction, extension interfaces, value conversion, redacted formatting, and public error mapping. | [Package GoDoc](../../doc.go), [engine assembly](../../engine.go) | | Root `promptkit` package | Provides the supported engine facade, source, backend-registration, and injection options, public request, result, prompt-inspection, and profile-inspection values, opaque prepared-execution handles, profile construction, extension interfaces, value conversion, redacted formatting, typed capacity errors, public error mapping, and engine-local assembly. | [Package GoDoc](../../doc.go), [prepared execution](../../prepared_execution.go), [backend API](../../backends.go), [engine assembly](../../engine.go) |
| `examples/go-library/prepare` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, and `Prepare`. It is not a public library package. | [Example program](../../examples/go-library/prepare/main.go) | | `examples/go-library/prepare` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, and `Prepare`. It is not a public library package. | [Example program](../../examples/go-library/prepare/main.go) |
| `examples/go-library/run` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, an injected deterministic model client, and `Run`. It is not a public library package. | [Example program](../../examples/go-library/run/main.go) | | `examples/go-library/run` | Demonstrates an offline downstream consumer using a prompt file, in-memory profile, inline input, an injected deterministic model client, and `Run`. It is not a public library package. | [Example program](../../examples/go-library/run/main.go) |
| `internal/backend` | Constructs each engine's immutable registry from the built-in OpenRouter definition and consumer additions, validates and defensively copies definitions through the shared JSON-value package, and consumes the LLM-owned OpenAI-compatible reserved request-field rule. | [Backend registry](../../internal/backend/registry.go) |
| `internal/capacity` | Owns engine-local bounded execution admission and FIFO model-generation permits for limited backend IDs, including cancellation-safe waiter removal and client wrapping. | [Internal capacity management](capacity.md) |
| `internal/domain` | Defines internal framework values for requests, artifacts, prompt definitions, profiles, execution targets, rendering, generation, and validation. | [Domain declarations](../../internal/domain/domain.go) | | `internal/domain` | Defines internal framework values for requests, artifacts, prompt definitions, profiles, execution targets, rendering, generation, and validation. | [Domain declarations](../../internal/domain/domain.go) |
| `internal/defaults` | Defines application-neutral framework constants and constructs the default execution target. It contains no CLI, server, or inbound HTTP limits. | [Framework defaults](../../internal/defaults/defaults.go) | | `internal/defaults` | Defines application-neutral framework constants and constructs the default execution target. It contains no CLI, server, or inbound HTTP limits. | [Framework defaults](../../internal/defaults/defaults.go) |
| `internal/filecatalog` | Provides deterministic YAML discovery and path helpers for operating-system filesystems and `fs.FS` sources. | [File catalog](../../internal/filecatalog/catalog.go) | | `internal/filecatalog` | Provides deterministic YAML discovery and path helpers for operating-system filesystems and `fs.FS` sources. | [File catalog](../../internal/filecatalog/catalog.go) |
| `internal/jsonvalue` | Validates and deeply copies JSON-compatible extra-parameter and prepared-schema trees while preserving supported concrete value types. | [JSON values](../../internal/jsonvalue/jsonvalue.go) |
| `internal/promptdef` | Loads strictly decoded, validated prompt definitions from filesystem and `fs.FS` sources, including version selection and contained file-backed message content. | [Framework formats](../formats.md), [prompt-definition repository](../../internal/promptdef/filesystem_repository.go) | | `internal/promptdef` | Loads strictly decoded, validated prompt definitions from filesystem and `fs.FS` sources, including version selection and contained file-backed message content. | [Framework formats](../formats.md), [prompt-definition repository](../../internal/promptdef/filesystem_repository.go) |
| `internal/profile` | Loads strictly decoded, validated execution profiles from filesystem and `fs.FS` sources and composes repositories with error-preserving fallback. | [Framework formats](../formats.md), [profile repositories](../../internal/profile/filesystem_repository.go) | | `internal/profile` | Loads strictly decoded, validated execution profiles, including backend selection, from filesystem and `fs.FS` sources and composes repositories with error-preserving fallback. | [Framework formats](../formats.md), [profile repositories](../../internal/profile/filesystem_repository.go) |
| `internal/profile/builtin` | Embeds the built-in execution profile catalog and combines it with an optional primary repository. | [Built-in catalog](../formats.md#built-in-profile-catalog), [repository](../../internal/profile/builtin/repository.go) | | `internal/profile/builtin` | Embeds the built-in profile catalog, whose entries select OpenRouter, and combines it with an optional primary repository. | [Built-in catalog](../formats.md#built-in-profile-catalog), [repository](../../internal/profile/builtin/repository.go) |
| `internal/prompt` | Renders prompt messages from Go templates with artifact, variable, session, and cache-control data. | [Go-template renderer](../../internal/prompt/go_renderer.go) | | `internal/prompt` | Renders prompt messages from Go templates with artifact, variable, session, and cache-control data. | [Go-template renderer](../../internal/prompt/go_renderer.go) |
| `internal/artifact` | Resolves ordinary inline and unrestricted caller-selected file references into copied artifacts with metadata and hashes. | [Internal sources and validation](sources.md) | | `internal/artifact` | Resolves ordinary inline and unrestricted caller-selected file references into copied artifacts with metadata and hashes. | [Internal sources and validation](sources.md) |
| `internal/validate` | Validates basic, JSON, and JSON Schema output using operating-system filesystem or `fs.FS` schema sources. | [Framework formats](../formats.md#schemas), [internal sources and validation](sources.md) | | `internal/validate` | Validates basic, JSON, and JSON Schema output using operating-system filesystem or `fs.FS` schema sources and creates frozen validation plans for prepared execution. | [Framework formats](../formats.md#schemas), [internal sources and validation](sources.md) |
| `internal/llm` | Defines the internal generation boundary and implements outbound OpenAI-compatible chat requests, response decoding, authentication, and deadline handling. | [Internal model client](llm.md) | | `internal/llm` | Defines the internal generation boundary and implements outbound OpenAI-compatible chat requests from resolved execution targets, including response decoding, authentication, deadline handling, and ownership of the OpenAI-compatible reserved request-field policy. | [Internal model client](llm.md) |
| `internal/usecase` | Coordinates preparation and execution across internal sources, rendering, artifact loading, generation, validation, and optional repair. | [Internal runner](runner.md) | | `internal/usecase` | Resolves prompt definitions and hashes, profiles, backends, and targets for exact inspection and request settings for preparation, and coordinates ordinary execution and one-attempt prepared execution across internal sources, rendering, artifact loading, generation, validation, capacity, and optional repair. | [Internal runner](runner.md), [prepared-execution implementation](../../internal/usecase/prepared_execution.go) |
The root package assembles these internal components without exposing their The root package assembles these internal components without exposing their
representations. Consumers depend only on the root facade. representations. Consumers depend only on the root facade.

View File

@@ -17,7 +17,10 @@ and override semantics consumed by the runner.
## Collaborators ## Collaborators
`Runner` coordinates narrow internal interfaces for prompt definitions, `Runner` coordinates narrow internal interfaces for prompt definitions,
profiles, artifacts, rendering, model generation, and validation. Schema profiles, backend resolution, artifacts, rendering, model generation, and
validation. The root engine supplies one immutable registry containing the
built-in backend and validated consumer additions, one engine-local run
admitter, and a model client wrapped by the same capacity manager. Schema
documents are loaded through the validator's optional schema-loader interface. documents are loaded through the validator's optional schema-loader interface.
An output repairer can be injected internally, but the ordinary runner An output repairer can be injected internally, but the ordinary runner
constructor does not enable one. constructor does not enable one.
@@ -25,43 +28,121 @@ constructor does not enable one.
Each invocation carries its state in request, prepared-run, and result values. Each invocation carries its state in request, prepared-run, and result values.
The runner has no durable run or session store. The runner has no durable run or session store.
## Preparation Flow ## Shared Prompt Selection
`Prepare` performs the reusable pre-generation workflow: The runner uses one prompt-selection and hashing boundary for ordinary
preparation and exact prompt inspection. Preparation retains its early
request-ID check before direct-session normalization; both operations then use
the configured prompt repository to select one definition, load referenced
message content, and calculate the same prompt hash.
1. validate the prompt selection and load the prompt definition; Inspection stops after that structural lookup. It does not parse templates or
2. hash the loaded definition; touch profile, artifact, schema, renderer, validator, admission, or model
collaborators. The root [`Engine.InspectPrompt`](../../engine.go) GoDoc owns
the public operation's exact contract.
## Shared Profile Selection
The runner uses one profile-selection and target-resolution boundary for
ordinary preparation and exact profile inspection. Preparation first selects a
request profile or a prompt default; inspection begins with its required
explicit profile ID. Both then apply the ordinary source precedence, resolve a
named backend, and construct the effective target from framework, backend, and
profile values.
Inspection stops after the resulting endpoint and model are structurally
validated. It does not check credential availability or perform prompt,
artifact, schema, rendering, admission, or model-client work. The root
[`Engine.InspectProfile`](../../engine.go) GoDoc owns the public operation's
exact contract.
## Shared Preparation Pipeline
`Prepare` and `Run` share one private preparation pipeline split at the point
where a run can be assigned to its selected backend pool. The resolution phase
performs only the work needed to validate routing and admission:
1. validate the required prompt selection and normalize any direct session ID;
2. load the prompt definition and hash the original definition;
3. select the request profile or the prompt's default profile; 3. select the request profile or the prompt's default profile;
4. resolve application-neutral defaults, profile values, and explicit request 4. resolve the profile's backend ID, when present;
overrides in that order; 5. resolve application-neutral defaults, backend defaults, profile values,
5. validate endpoint, model, numeric overrides, and credential requirements; and explicit request overrides in that order;
6. resolve the output contract and load a structured-output schema when 6. validate endpoint, model, numeric overrides, and credential requirements;
required; 7. resolve the effective output contract without loading its schema; and
7. load and hash input artifacts; 8. retain the definition, source identities, effective settings, output
8. render and hash the prompt; and contract, and preparation start time in invocation-local state.
9. return the effective settings, source identities, messages, hashes, and
preparation timing. The completion phase consumes that state without reloading the prompt,
profile, or backend:
1. load structured-output schema metadata when required;
2. load and hash input artifacts;
3. render messages and the prompt-defined session;
4. apply any direct session ID;
5. hash the effective rendered prompt; and
6. construct the prepared value and preparation timing.
`Prepare` runs both phases consecutively and never performs capacity admission.
`Run` performs backend admission between the phases. This structure preserves
one execution-precedence and error-ordering implementation while allowing a
full backend pool to reject work before expensive schema, artifact, and
rendering operations.
Pointer-based numeric overrides preserve an explicit zero. Invalid negative or Pointer-based numeric overrides preserve an explicit zero. Invalid negative or
out-of-range values fail as invalid requests. A direct API key takes out-of-range values fail as invalid requests. Endpoint overrides do not change
precedence over environment lookup for execution; secret values remain the selected backend identity. Non-empty extra-parameter maps replace whole
excluded from serialized metadata. lower-precedence maps. A direct API key takes precedence over environment
lookup; otherwise request, profile, and backend environment-variable names
apply in that order. A profile requiring a direct key clears an inherited
backend environment name unless the request supplies its own name. Secret
values remain excluded from serialized metadata.
Reasoning overrides are tri-state: nil inherits the profile, a pointer to a
nonblank string trims and replaces it, and a pointer to a blank string clears
it. A nonblank direct session is normalized before source loading, bypasses
the prompt session template, and is applied after ordinary message rendering.
A blank direct value retains prompt-template behavior. The runner clears the
template only on a value copy of the definition, so the definition hash always
describes the original source while the rendered-prompt hash includes the
effective direct or rendered session.
The registry is read-only after engine construction. Concurrent `Prepare` and
`Run` calls resolve independent defensive backend values and keep all
invocation state local.
## Run Flow ## Run Flow
`Run` calls `Prepare` rather than maintaining a second preparation path. It `Run` records its start time, performs the shared resolution phase, and asks
performs one initial generation call, builds the named output artifact, and its `RunAdmitter` to reserve capacity for the effective backend ID. A nil
validates that artifact. Invalid generated content remains a validation result; admitter is an internal unlimited fallback. After successful admission, `Run`
an inability to perform validation is an operational error. immediately defers the returned release function, performs the completion
phase, makes one initial generation call, builds the named output artifact,
and validates that artifact. Invalid generated content remains a validation
result; an inability to perform validation is an operational error.
The admission lease covers completion-phase preparation, initial generation,
validation, every repair, and every exit. It bounds accepted work without
serializing preparation or validation behind the active-generation limit.
The wrapped model client separately acquires a FIFO active permit only around
each actual generation call.
When an internal repairer is present, a JSON or JSON Schema content failure can When an internal repairer is present, a JSON or JSON Schema content failure can
trigger bounded repair attempts. Repair receives the effective execution trigger bounded repair attempts. Repair receives the effective execution
target, validation errors, prior output, and structured-output specification. target and session ID, validation errors, prior output, and structured-output
This capability remains internal and is not a public option. specification. The default repairer uses the same wrapped client as initial
generation, so each repair reacquires the selected backend's active permit
while remaining inside its original admission lease. Repair never performs a
second bounded admission. This capability remains internal and is not a public
option.
A successful result includes the output artifact and raw output, validation A successful result includes the output artifact and raw output, validation
state, prompt and rendered-prompt hashes, selected profile, effective settings, state, effective session ID, prompt and rendered-prompt hashes, selected
input hashes, token usage, a generated run identifier, and UTC timing. profile and backend, effective settings, input hashes, token usage, a generated
run identifier, and UTC timing. The same effective session reaches initial
generation and any repair attempt through the rendered prompt. The same
effective target, including backend identity, reaches generation and any
repair attempt.
## Failure Categories ## Failure Categories
@@ -69,17 +150,40 @@ Package errors distinguish invalid requests, required profile selection,
credential failures, and prompt, profile, artifact, rendering, generation, and credential failures, and prompt, profile, artifact, rendering, generation, and
validation failures. Wrapping preserves the package identities mapped by the validation failures. Wrapping preserves the package identities mapped by the
public facade and retains collaborator identities where they are part of the public facade and retains collaborator identities where they are part of the
internal contract. Context cancellation propagates through the invoked internal contract.
collaborator and is classified by the owning operation.
Admission capacity exhaustion retains the internal capacity identity. At the
use-case boundary, the runner attaches the selected backend ID in an internal
typed error, and the root facade copies that value into the public
[`CapacityError`](../../capacity_error.go) without parsing diagnostic text. It
is not recategorized as an invalid request or generation failure, and no
partial result is returned. A context already done at admission retains its
context identity directly. Cancellation while waiting for an active generation
permit prevents client invocation when it wins the grant race; the model-client
boundary then preserves the context error through the generation-failure
category. Deferred release restores the admission lease on preparation,
generation, validation, repair, and cancellation failures.
Other context cancellation propagates through the invoked collaborator and is
classified by the owning operation.
An overlong direct session is an invalid request before source loading, while
an invalid or overlong prompt session template remains a prompt-render failure.
An unknown selected backend, or a selected backend with no configured resolver,
is classified as a profile-load failure.
## Test Ownership And Changes ## Test Ownership And Changes
The [runner tests](../../internal/usecase/runner_test.go) own preparation order, The [runner tests](../../internal/usecase/runner_test.go) own preparation order,
selection and override precedence, schema-before-generation behavior, hashing, selection and override precedence, the two-phase boundary, early admission,
generation and validation outcomes, bounded repair, credentials and redaction, lease lifetime and release, direct-session resolution, schema-before-generation
error categories, artifact metadata, usage, and timing. behavior, hashing, generation and validation outcomes, backend propagation,
bounded repair, shared initial/repair capacity, credentials and redaction,
error categories, artifact metadata, usage, and timing. The
[capacity subsystem document](capacity.md) identifies the focused pool,
waiter, and wrapped-client tests.
Changes to orchestration should continue to use the existing package Changes to orchestration should continue to use the existing package
interfaces, keep request state local to an invocation, and preserve `Run`'s use interfaces, keep request state local to an invocation, and preserve the shared
of `Prepare`. Source, renderer, validator, or model-client contract changes resolution and completion pipeline. Source, renderer, validator, or
belong first in their owning package and document. model-client contract changes belong first in their owning package and
document.

View File

@@ -16,6 +16,11 @@ validation modes, built-in catalog, and source precedence.
definitions, selects an ID and optional version, and resolves file-backed definitions, selects an ID and optional version, and resolves file-backed
message content within the selected operating-system or `fs.FS` source. message content within the selected operating-system or `fs.FS` source.
Exact prompt inspection performs one point-in-time lookup through that same
repository and validates referenced message content before returning declared
metadata. It does not parse templates or read profile, input, or schema
sources, and it does not retain the definition for a later execution.
Its package tests own prompt selection, strict decoding, definition validation, Its package tests own prompt selection, strict decoding, definition validation,
duplicate detection, and source containment: duplicate detection, and source containment:
[prompt-definition repository tests](../../internal/promptdef/repository_test.go). [prompt-definition repository tests](../../internal/promptdef/repository_test.go).
@@ -24,13 +29,25 @@ duplicate detection, and source containment:
`internal/profile` loads and validates execution profiles from an `internal/profile` loads and validates execution profiles from an
operating-system filesystem or an `fs.FS`. It supports a primary repository operating-system filesystem or an `fs.FS`. It supports a primary repository
with fallback only when the primary reports that a profile is absent. with fallback only when the primary reports that a profile is absent. Strict
YAML decoding recognizes the optional `backend` field, trims its value, and
requires a model plus at least one non-blank backend or endpoint. Loading does
not check registry membership because the available registry belongs to the
assembled engine; the runner checks membership during preparation and exact
profile inspection.
Exact profile inspection performs one point-in-time lookup through those
profile sources and checks the resolved target without reading prompt, input,
or schema sources. It does not retain that lookup for a later execution.
`internal/profile/builtin` embeds the maintained built-in profile catalog and `internal/profile/builtin` embeds the maintained built-in profile catalog and
can place a caller-selected repository ahead of that catalog. Profile behavior can place a caller-selected repository ahead of that catalog. Every embedded
is owned by the profile selects `openrouter` and inherits its endpoint and credential
environment-variable name from the built-in backend registry rather than
repeating those values. Profile behavior is owned by the
[profile repository tests](../../internal/profile/repository_test.go), while [profile repository tests](../../internal/profile/repository_test.go), while
catalog completeness, duplicate IDs, and overlay behavior are owned by the catalog completeness, the backend-selection invariant, duplicate IDs, and
overlay behavior are owned by the
[built-in repository tests](../../internal/profile/builtin/repository_test.go). [built-in repository tests](../../internal/profile/builtin/repository_test.go).
## Ordinary Artifacts ## Ordinary Artifacts
@@ -61,6 +78,22 @@ filesystem or an `fs.FS`. Invalid generated content is returned as a validation
result; inability to load, register, or compile a schema is an operational result; inability to load, register, or compile a schema is an operational
error. error.
For executable preparation, the built-in validators create a frozen validation
plan. None, basic, and JSON modes retain the effective output contract without
source access. JSON Schema mode loads the root document, resolves and compiles
every transitive reference during preparation, and retains the compiled
validator. The provider-facing structured-output metadata uses that same
captured root document.
`PrepareExecution` also completes prompt and profile selection, artifact
loading and hashing, session and message rendering, and target resolution.
`RunPrepared` uses the retained source-derived state and validation plan; it
does not reopen prompt, profile, input, or schema sources and does not rerender
the request. By contrast, ordinary `Prepare` produces a preparation value only:
a later `Run` performs its own source resolution and preparation.
The [validator tests](../../internal/validate/standard_validator_test.go) own The [validator tests](../../internal/validate/standard_validator_test.go) own
basic, JSON, JSON Schema, source resolution, schema loading, compilation, and basic, JSON, JSON Schema, source resolution, schema loading, compilation,
content-failure behavior. frozen-reference behavior, and content-failure behavior. Prepared execution
orchestration is owned by the
[use-case tests](../../internal/usecase/prepared_execution_test.go).

View File

@@ -21,10 +21,16 @@ The implemented internal components consist of:
- `internal/domain`, which owns framework data values shared by later internal - `internal/domain`, which owns framework data values shared by later internal
components; components;
- `internal/backend`, which owns validated immutable OpenAI-compatible backend
definitions and the built-in OpenRouter definition;
- `internal/capacity`, which owns engine-local bounded run admission and
model-generation scheduling for limited backends;
- `internal/defaults`, which owns application-neutral framework defaults and - `internal/defaults`, which owns application-neutral framework defaults and
constructs the default execution target; constructs the default execution target;
- `internal/filecatalog`, which discovers YAML files and provides source-path - `internal/filecatalog`, which discovers YAML files and provides source-path
helpers for filesystem and `fs.FS` consumers; helpers for filesystem and `fs.FS` consumers;
- `internal/jsonvalue`, which validates and defensively copies JSON-compatible
extra-parameter trees;
- `internal/promptdef`, which loads and validates prompt definitions from - `internal/promptdef`, which loads and validates prompt definitions from
filesystem and `fs.FS` sources; filesystem and `fs.FS` sources;
- `internal/profile`, which loads, validates, and overlays execution profiles - `internal/profile`, which loads, validates, and overlays execution profiles
@@ -45,17 +51,21 @@ The `examples/go-library/prepare` and `examples/go-library/run` packages are
maintained downstream consumers of the root facade. They do not expose library maintained downstream consumers of the root facade. They do not expose library
packages or participate in internal assembly. packages or participate in internal assembly.
The root facade assembles the internal repositories, renderer, validator, The root facade assembles one immutable backend registry, one capacity manager,
outbound client, and use-case runner while translating public values and the internal repositories, renderer, validator, outbound client, and use-case
errors at the library boundary. The defaults and renderer depend on the domain runner while translating public values and errors at the library boundary. The
model. Prompt-definition and profile repositories use the domain model, file registry contains built-ins plus validated engine-scoped consumer additions.
catalog, and YAML decoder. The built-in profile repository supplies an The facade constructs the capacity manager from the registry's immutable
embedded `fs.FS` to the profile package. Artifact reading uses the domain model policy snapshot, wraps the selected built-in or injected model client, and
and application-neutral defaults. Validation uses the domain model, file supplies bounded admission to the runner. The defaults and renderer depend on
the domain model. Prompt-definition and profile repositories use the domain
model, file catalog, and YAML decoder. The built-in profile repository supplies
an embedded `fs.FS` to the profile package. Artifact reading uses the domain
model and application-neutral defaults. Validation uses the domain model, file
catalog, and JSON Schema implementation. The model client uses the domain catalog, and JSON Schema implementation. The model client uses the domain
model, application-neutral defaults, and an injected or standard-library HTTP model, application-neutral defaults, and an injected or standard-library HTTP
client. The use-case runner depends on the narrow interfaces owned by each client. The use-case runner depends on the narrow interfaces owned by each
internal component. internal component, including backend lookup and run admission.
The current implementation follows this dependency direction: The current implementation follows this dependency direction:
@@ -72,9 +82,15 @@ downstream consumers, including Scriptorium
narrow injected abstractions narrow injected abstractions
``` ```
The facade coordinates internal components and adapts the supported public The backend registry depends on the domain model and shared JSON-value
extension interfaces to narrow internal abstractions. Internal components must validation, has no mutation API after construction, and consumes the
not depend on consumers or on Scriptorium. OpenAI-compatible reserved request-field rule owned by the model client. The
capacity component depends on the domain model and the narrow internal
model-client boundary, not on provider transport implementation. The model
client does not depend on registry or capacity configuration. The facade
coordinates internal components and adapts the supported public extension
interfaces to narrow internal abstractions. Internal components must not depend
on consumers or on Scriptorium.
## Repository And Consumer Boundary ## Repository And Consumer Boundary

View File

@@ -78,6 +78,7 @@ mechanisms, not secret values.
| Framework file formats | `docs/formats.md` | Prompt-definition and profile YAML fields, schema references, defaults, validation modes, built-in profiles, credentials, and file-to-request precedence. | Exported Go declarations, outbound wire behavior, internal parsing mechanics, and application configuration. | | Framework file formats | `docs/formats.md` | Prompt-definition and profile YAML fields, schema references, defaults, validation modes, built-in profiles, credentials, and file-to-request precedence. | Exported Go declarations, outbound wire behavior, internal parsing mechanics, and application configuration. |
| Consumer guidance | `docs/consumers/`, when consumer workflows require dedicated guidance | Task-oriented use of implemented public APIs, minimal examples, and consumer responsibilities. | Exact exported declarations and internal mechanics. | | Consumer guidance | `docs/consumers/`, when consumer workflows require dedicated guidance | Task-oriented use of implemented public APIs, minimal examples, and consumer responsibilities. | Exact exported declarations and internal mechanics. |
| Durable integration contracts | `docs/integrations/`, when integrations exist | External formats and protocols, compatibility behavior, and upstream or downstream responsibilities. | Internal transformations and public Go declarations. | | Durable integration contracts | `docs/integrations/`, when integrations exist | External formats and protocols, compatibility behavior, and upstream or downstream responsibilities. | Internal transformations and public Go declarations. |
| Supplemental release guidance | None. `docs/releases/` may be used when a release benefits from a changelog or migration guide. | No canonical content. These files may briefly summarize release-specific changes, compatibility, and consumer migration paths, and may be corrected, consolidated, archived, or removed when no longer useful. | Public API and behavior contracts, formats, integrations, architecture, release procedure, and the authoritative annotated-tag release record. |
| Implemented component inventory | `docs/internal/overview.md` | Current packages and components, their implemented responsibilities, and links to focused internal documents. | Normative architecture, contributor workflow, external contracts, and proposed components. | | Implemented component inventory | `docs/internal/overview.md` | Current packages and components, their implemented responsibilities, and links to focused internal documents. | Normative architecture, contributor workflow, external contracts, and proposed components. |
| Internal subsystem behavior | Other files under `docs/internal/`, when a subsystem needs durable detail | Implementation flow, internal collaborators and state transitions, package-local guarantees and failures, and relevant tests. | Global architecture invariants, public API definitions, and future package plans. | | Internal subsystem behavior | Other files under `docs/internal/`, when a subsystem needs durable detail | Implementation flow, internal collaborators and state transitions, package-local guarantees and failures, and relevant tests. | Global architecture invariants, public API definitions, and future package plans. |
| Architectural decision history | `docs/adr/`, when repository-local decisions require records | Significant decisions, context, alternatives, rationale, consequences, and supersession history. | Current behavior reference, implementation status, and task sequencing. | | Architectural decision history | `docs/adr/`, when repository-local decisions require records | Significant decisions, context, alternatives, rationale, consequences, and supersession history. | Current behavior reference, implementation status, and task sequencing. |
@@ -85,9 +86,9 @@ mechanisms, not secret values.
| Complete copyable artifacts | `examples/` | Valid inputs, Go programs, and other files intended to be copied or run. | Field-by-field reference, exact API declarations, and prose explanation. | | Complete copyable artifacts | `examples/` | Valid inputs, Go programs, and other files intended to be copied or run. | Field-by-field reference, exact API declarations, and prose explanation. |
Conditional owners do not require placeholder files or directories. Create a Conditional owners do not require placeholder files or directories. Create a
consumer, integration, subsystem, ADR, roadmap, or example document only when consumer, integration, release, subsystem, ADR, roadmap, or example document
the corresponding implemented interface, decision, planned effort, or only when the corresponding implemented interface, release, decision, planned
maintained artifact exists. effort, or maintained artifact exists.
## Boundary Rules ## Boundary Rules
@@ -109,6 +110,23 @@ but must link to its canonical definition rather than restate it.
The [framework format reference](../formats.md) owns exact prompt, profile, and The [framework format reference](../formats.md) owns exact prompt, profile, and
schema-file contracts. Integration documents own external wire formats. schema-file contracts. Integration documents own external wire formats.
### Supplemental Release Guidance
Files under `docs/releases/` may provide changelog-style summaries and
migration guidance for a particular release. They are navigation and
orientation aids, not canonical owners of public APIs, behavior, formats,
integrations, architecture, release procedure, or other durable facts. When a
reader needs detail beyond a short release-specific note, the release document
must link to the applicable canonical documentation rather than reproduce its
contract.
The annotated tag message required by the
[release procedure](../release.md#write-the-release-note) remains the
authoritative release record. Supplemental release documents may be corrected,
consolidated, archived, or removed at any time when they are no longer useful,
provided maintained documentation does not depend on them and the annotated
tag record remains intact.
### Security Topics ### Security Topics
This policy owns what documentation and examples may contain. Architecture owns This policy owns what documentation and examples may contain. Architecture owns
@@ -160,6 +178,10 @@ durable owners, update incoming links, and archive or remove the roadmap
according to repository practice. Do not preserve completed roadmaps as a according to repository practice. Do not preserve completed roadmaps as a
second current-state reference. second current-state reference.
Supplemental release documents may likewise be removed without preserving a
replacement. Before removal, update maintained incoming links so current
documentation does not depend on an optional historical guide.
Before completing documentation work: Before completing documentation work:
- verify affected behavior and examples; - verify affected behavior and examples;

243
docs/releases/v0.2.0.md Normal file
View File

@@ -0,0 +1,243 @@
# Promptkit v0.2.0
This supplemental changelog and migration guide summarizes the consumer-facing
changes from `v0.1.0` to `v0.2.0`. The annotated `v0.2.0` tag is the
authoritative release record. Exact current contracts belong to the linked
GoDoc and durable documentation.
## Summary
`v0.2.0` adds three major capabilities:
- an engine-scoped registry for reusable OpenAI-compatible backend
definitions;
- bounded, backend-specific run admission and model-generation concurrency;
and
- direct per-run session IDs and tri-state reasoning-effort overrides.
Existing endpoint-only profiles remain supported. Consumers can adopt backend
registration and runtime overrides incrementally rather than rewriting all
profiles during the upgrade.
## Compatibility At A Glance
Promptkit remains pre-`v1`, and this minor release includes source-level and
behavioral changes that deserve review.
| Area | `v0.1.0` consumer impact |
| --- | --- |
| Endpoint-only profiles | Continue to work without migration. |
| Built-in profiles | Continue to use OpenRouter and `OPENROUTER_API_KEY`; they now select the built-in `openrouter` backend. |
| Custom backends | Registration is optional. Existing profiles may keep their endpoint and credential configuration. |
| Reasoning overrides | String assignments must migrate to the new pointer field. |
| `RunRequest.Metadata` | Removed; delete assignments to this field. |
| OpenRouter concurrency | Now limited to 16 active generations with waiting capacity of 1024 per engine. |
| Public JSON | `v0.2.0` formalizes supported JSON representations; consumers relying on `v0.1.0` encodings should review the notes below. |
| Unkeyed public struct literals | May require updates because fields were added. Keyed literals are recommended. |
## Upgrade
After the `v0.2.0` tag is published, update the module dependency with:
```sh
go get gitea.maximumdirect.net/eric/promptkit@v0.2.0
go mod tidy
```
Run the consuming project's ordinary tests and race-enabled tests after the
upgrade, especially if it calls one engine concurrently or persists Promptkit
JSON values.
## Backend Registry
Consumers may now register reusable OpenAI-compatible backend definitions with
`WithBackend`, then select them by ID from file-backed or in-memory profiles.
A backend can supply its endpoint, API-key environment-variable name,
request-wide extra parameters, and optional capacity policy.
Registrations are immutable and belong to one engine. Consumer registrations
can add new IDs but cannot replace Promptkit's reserved `openrouter` backend.
Profiles that select a backend may still override its endpoint without losing
the backend's routing or capacity identity.
An existing endpoint-only in-memory profile remains valid:
```go
promptkit.Profile{
ID: "local",
Endpoint: "http://localhost:8000/v1",
Model: "example-model",
}
```
Adopting the registry is optional and can be done when several profiles should
share connection or capacity settings:
```go
engine, err := promptkit.NewEngine(
promptkit.Config{PromptDir: "prompts"},
promptkit.WithBackend(promptkit.Backend{
ID: "local",
Endpoint: "http://localhost:8000/v1",
APIKeyEnv: "LOCAL_LLM_API_KEY",
}),
promptkit.WithProfiles(promptkit.Profile{
ID: "local-summary",
BackendID: "local",
Model: "example-model",
}),
)
```
See the
[local-endpoint consumer guide](../consumers/pkg-promptkit.md#configure-a-local-openai-compatible-endpoint)
for task-oriented usage. The
[`Backend` and `WithBackend` GoDoc](../../backends.go) owns exact registration,
validation, copying, defaulting, and uniqueness semantics. The
[framework format reference](../formats.md) owns the profile `backend` field
and execution precedence.
## Backend-Specific Concurrency
Each registered backend may now define:
- an active model-generation limit; and
- a bounded number of additional admitted `Run` calls.
Promptkit owns scheduling for both its built-in model client and an injected
`LLMClient`. `Run` remains synchronous: an admitted caller waits for its
ordinary result, while a call beyond the bounded admission capacity returns
`ErrCapacityExceeded`. Capacity is engine-local and keyed by backend ID.
Endpoint-only profiles and custom backends without a configured limit remain
unlimited.
The built-in OpenRouter backend now permits 16 active generations and 1024
additional admitted calls per engine. Applications that can exceed this bound
should handle capacity exhaustion separately from provider and request
failures:
```go
result, err := engine.Run(ctx, request)
if errors.Is(err, promptkit.ErrCapacityExceeded) {
// Apply application-specific overload or retry policy.
}
```
Promptkit does not prescribe retries or map this error to an HTTP status. See
the
[concurrency consumer guidance](../consumers/pkg-promptkit.md#limit-backend-concurrency)
and the [`Backend` GoDoc](../../backends.go) for the canonical configuration
contract. Runtime behavior and public error identities belong to the
[`Engine.Run` GoDoc](../../engine.go).
## Per-Run Session IDs
`RunRequest.SessionID` can now supply a consumer-managed correlation ID for one
`Prepare` or `Run` invocation. A nonblank direct value overrides the prompt's
session template and is exposed in prepared values, results, injected-client
requests, and provider observability. Session IDs should therefore be stable,
non-secret values.
```go
result, err := engine.Run(ctx, promptkit.RunRequest{
PromptID: "meeting.summary",
SessionID: "conversation-42",
})
```
The built-in OpenAI-compatible client sends a nonempty effective session as the
top-level `session_id` request-body field, not as an `x-session-id` header. See
the
[session and reasoning consumer guide](../consumers/pkg-promptkit.md#set-a-per-run-session-and-reasoning),
the [`RunRequest` GoDoc](../../types.go), and the
[OpenAI-compatible request contract](../integrations/openai-compatible-chat.md#request-body)
for exact normalization, length, exposure, and wire behavior.
## Per-Run Reasoning Effort
`ExecutionTargetOverride.ReasoningEffort` changed from `string` to `*string` so
one request can distinguish inheritance, replacement, and explicit clearing.
Update a `v0.1.0` override like this:
```go
// v0.1.0
Execution: &promptkit.ExecutionTargetOverride{
ReasoningEffort: "high",
}
```
to:
```go
// v0.2.0
reasoning := "high"
Execution: &promptkit.ExecutionTargetOverride{
ReasoningEffort: &reasoning,
}
```
The three states are:
- `nil` inherits the selected profile's value;
- a pointer to a nonblank string replaces it for that invocation; and
- a pointer to an empty or whitespace-only string clears it for that
invocation.
This allows consumers to consolidate profiles that differed only by reasoning
effort. The [`ExecutionTargetOverride` GoDoc](../../types.go) owns the exact
override contract.
## Other Migration Notes
### Remove `RunRequest.Metadata`
`RunRequest.Metadata` is no longer part of the public request. Remove any
assignment to that field. Use application-owned state keyed by `RunResult.RunID`
or a direct `SessionID` when correlation is needed; these identifiers have
different purposes, so choose according to the application's lifecycle.
### Review Persisted JSON
`v0.2.0` defines stable JSON representations for the public result, artifact,
execution, validation, and model-client values listed in the
[package documentation](../../doc.go). Consumers that treated `v0.1.0`
reflection-derived encodings as stable should update fixtures and stored-data
adapters.
In particular:
- `RunResult` encodes elapsed time as integer milliseconds in `duration_ms`
instead of encoding `time.Duration` under `duration`;
- result JSON can include the new `session_id` and `selected_backend_id`
fields;
- execution-target JSON can include `backend_id`; and
- artifact and target-presence fields now use their documented lower-case
names.
The `v0.2.0` `RunResult` decoder reads `duration_ms`; it does not translate a
persisted `v0.1.0` `duration` field. Transform old payloads before decoding
when preserving their elapsed duration matters.
### Prefer Keyed Struct Literals
New fields were added to several public structs. Replace positional composite
literals with keyed literals so future additive fields do not cause another
source migration.
## Migration Checklist
- Update the module dependency and run the consumer's tests.
- Change reasoning overrides from strings to pointers.
- Remove uses of `RunRequest.Metadata`.
- Review unkeyed Promptkit struct literals.
- Decide whether shared endpoints should move into registered backends.
- If using built-in OpenRouter profiles at high concurrency, handle
`ErrCapacityExceeded` and review the new engine-local bound.
- Review stored JSON, fixtures, and downstream decoders.
- Optionally replace profile-specific session or reasoning variants with
per-run overrides.
For complete consumer workflows, use the
[package consumer guide](../consumers/pkg-promptkit.md) and maintained
[offline execution example](../../examples/go-library/run/main.go).

74
docs/releases/v0.3.0.md Normal file
View File

@@ -0,0 +1,74 @@
# Promptkit v0.3.0
This supplemental changelog summarizes the consumer-facing changes from
`v0.2.0` to `v0.3.0`. The annotated `v0.3.0` tag is the authoritative release
record. Exact current contracts belong to the linked GoDoc and durable
documentation.
## Summary
`v0.3.0` adds a concise way to register the common local OpenAI-compatible
backend configuration:
- `BackendLocal` provides the conventional, non-reserved backend ID `"local"`;
and
- `LocalBackend` constructs an ordinary `Backend` from an endpoint and
concurrency limit.
The helper is explicit and additive. It does not pre-register a backend, read
environment variables, select a model, or replace the complete `Backend`
configuration interface.
## Compatibility
Existing `v0.2.0` consumers require no migration. Endpoint-only profiles,
complete custom `Backend` values, the built-in OpenRouter backend, and existing
registrations using the literal ID `"local"` continue to work unchanged.
## Upgrade
Update the module dependency with:
```sh
go get gitea.maximumdirect.net/eric/promptkit@v0.3.0
go mod tidy
```
Run the consuming project's ordinary tests and race-enabled tests after the
upgrade.
## Configure A Local Backend
Register the convenience value through the existing `WithBackend` option and
select it from one or more profiles:
```go
engine, err := promptkit.NewEngine(
promptkit.Config{PromptDir: "prompts"},
promptkit.WithBackend(
promptkit.LocalBackend("http://localhost:8000/v1", 2),
),
promptkit.WithProfiles(promptkit.Profile{
ID: "local-summary",
BackendID: promptkit.BackendLocal,
Model: "example-model",
}),
)
```
Use an endpoint-only profile when shared backend identity and capacity policy
are unnecessary. Continue to use a complete keyed `Backend` value for custom
IDs, authentication, extra request parameters, explicit queue capacity, or
multiple local endpoints.
See the
[local-endpoint consumer guide](../consumers/pkg-promptkit.md#configure-a-local-openai-compatible-endpoint)
for task-oriented configuration choices. The
[`BackendLocal`, `LocalBackend`, and `WithBackend` GoDoc](../../backends.go)
owns their exact construction, registration, validation, and concurrency
semantics.
## Consumer Action
None. Adopt the convenience constructor when it simplifies local endpoint
configuration.

189
docs/releases/v0.4.0.md Normal file
View File

@@ -0,0 +1,189 @@
# Promptkit v0.4.0
This supplemental changelog and adoption guide summarizes the consumer-facing
changes from `v0.3.0` to `v0.4.0`. The annotated `v0.4.0` tag is the
authoritative release record. Exact current contracts belong to the linked
GoDoc and durable documentation.
## Summary
`v0.4.0` adds four complementary capabilities:
- opaque prepared-execution handles for preparing once, inspecting safe
details, and executing the same frozen snapshot;
- exact prompt-definition inspection without profile resolution or execution;
- exact profile inspection without selecting a prompt or checking credential
availability; and
- structured backend identity on engine admission-capacity rejection.
These APIs let consumers perform more precise preflight work and retain useful
operational context without reproducing Promptkit's internal resolution logic.
## Compatibility
The release is additive for `v0.3.0` consumers. Existing uses of `Prepare`,
`Run`, backend registration, endpoint-only profiles, local-backend helpers,
runtime overrides, public JSON values, and error sentinels continue to work
without migration.
Capacity rejection now returns a structured error while continuing to match
`ErrCapacityExceeded` through `errors.Is`. Error-string wording and direct
sentinel equality were not public contracts.
The new inspection values, capacity error, and prepared-execution handle do not
have stable JSON representations. `PreparedExecution.Details` returns the
existing stable `PreparedRun` value.
## Upgrade
Update the module dependency with:
```sh
go get gitea.maximumdirect.net/eric/promptkit@v0.4.0
go mod tidy
```
Run the consuming project's ordinary and race-enabled tests after upgrading.
No source migration is required.
## Prepare Once And Execute The Same Snapshot
Consumers that need to persist preparation details before generation can now
prepare an opaque, engine-bound execution:
```go
prepared, err := engine.PrepareExecution(ctx, request)
if err != nil {
// Handle preparation failure.
}
defer prepared.Discard()
details := prepared.Details()
// Persist a consumer-selected, appropriately protected preparation record.
result, err := engine.RunPrepared(ctx, prepared)
```
Preparation freezes the selected sources, rendered messages, effective
settings, input content, structured-output metadata, and validation resources
needed by execution. `Details` returns a fresh, caller-owned,
credential-redacted `PreparedRun`.
A handle belongs to its creating engine and permits one execution attempt.
`RunPrepared` consumes that attempt on success and on operational failure.
`Discard` is idempotent and releases an unclaimed handle's execution-only
state. Consumers should discard handles they will not execute, particularly
when a direct request API key may be retained privately until claim or
discard.
Prepared execution does not reserve backend admission during preparation.
Credential availability and backend admission are checked when execution
begins. The execution context is independent of the preparation context.
See the
[prepared-execution consumer guide](../consumers/pkg-promptkit.md#prepare-now-and-execute-the-same-snapshot-later),
the [`PreparedExecution` GoDoc](../../prepared_execution.go), and the
[`Engine.PrepareExecution` and `Engine.RunPrepared` GoDoc](../../engine.go)
for the exact lifecycle, ownership, cancellation, capacity, timing, and
failure contracts.
## Inspect A Prompt
`Engine.InspectPrompt` resolves one prompt ID and optional version through the
engine's configured prompt source:
```go
inspection, err := engine.InspectPrompt(ctx, "report.summary", "")
```
The result includes prompt identity, the opaque prompt hash, declared default
profile ID, declared input metadata, and normalized output contract. It
structurally loads the selected definition and referenced message content but
does not resolve a profile, load schemas or artifacts, render templates,
reserve capacity, or contact a model.
Use inspection for exact configuration checks and metadata discovery. Use
`PrepareExecution` rather than relying on a prior inspection when later
execution must freeze one exact source state, because filesystem-backed
inspection is only a point-in-time lookup.
See the
[prompt-inspection consumer guide](../consumers/pkg-promptkit.md#inspect-a-prompt-before-preparation)
and [`Engine.InspectPrompt` GoDoc](../../engine.go) for exact selection,
ownership, and error behavior.
## Inspect A Profile
`Engine.InspectProfile` resolves one explicit profile independently of a
prompt:
```go
inspection, err := engine.InspectProfile(ctx, "report-production")
```
The result includes the resolved effective execution target and whether a
later request must provide a direct credential. Environment-variable names may
be reported, but inspection does not read credential values or require the
named variable to be populated.
Inspection applies the ordinary configured and built-in profile precedence and
resolves any selected backend. It does not load a prompt, render content,
reserve capacity, or contact a model.
See the
[profile-inspection consumer guide](../consumers/pkg-promptkit.md#inspect-a-profile-before-prompt-work)
and [`Engine.InspectProfile` GoDoc](../../engine.go) for the exact resolution,
credential, ownership, and error contracts.
## Identify Capacity-Rejected Backends
Calls rejected at Promptkit's bounded engine admission boundary continue to
match `ErrCapacityExceeded`. Consumers can additionally obtain the selected
registered backend ID without parsing diagnostic text:
```go
result, err := engine.Run(ctx, request)
if errors.Is(err, promptkit.ErrCapacityExceeded) {
var capacityErr *promptkit.CapacityError
if errors.As(err, &capacityErr) {
// Record capacityErr.BackendID using application-owned diagnostics.
}
// Apply application-owned overload or retry policy.
}
```
The structured error applies to `Run` and `RunPrepared` admission rejection.
It does not represent provider throttling, quota exhaustion, cancellation
while waiting for generation capacity, or another model-client failure.
Promptkit does not prescribe retry timing or transport status mapping.
See the
[error-handling consumer guide](../consumers/pkg-promptkit.md#handle-errors),
the [`CapacityError` GoDoc](../../capacity_error.go), and the
[`ErrCapacityExceeded` GoDoc](../../engine.go) for the canonical contracts.
## Public API Additions
The release adds:
- `Engine.PrepareExecution`;
- `Engine.RunPrepared`;
- `PreparedExecution`, including `Details`, `Discard`, `String`, and
`GoString`;
- `Engine.InspectPrompt`;
- `PromptInspection`;
- `PromptInputDefinition`;
- `Engine.InspectProfile`;
- `ProfileInspection`; and
- `CapacityError`.
No public API was removed.
## Consumer Action
None. Existing `v0.3.0` workflows may upgrade without adopting the new APIs.
Consumers that adopt prepared execution should discard unused handles.
Consumers that need backend-specific capacity diagnostics may add an
`errors.As` check while retaining their existing `errors.Is` classification.

View File

@@ -1,222 +0,0 @@
# Documentation Hardening Roadmap
## Purpose
This roadmap coordinates a focused pass over Promptkit's documentation and
public contract. The work should close the gaps identified by the documentation
audit, strengthen canonical ownership, and leave consumers and contributors
with guidance that is accurate, navigable, and proportionate to their needs.
This is planning material, not a description of implemented behavior. Follow
the [documentation policy](../policy/documentation.md) throughout the work and
update current-state documents only when their claims are supported by the
implementation and tests.
## Scope And Principles
The work covers public GoDoc, consumer guidance, maintained examples, format
and integration references, architecture ownership, roadmap lifecycle, and
documentation validation.
- Resolve ambiguous behavior before documenting it as a contract.
- Keep exact exported API semantics in Go declarations and GoDoc.
- Keep task-oriented guidance in consumer documents and complete runnable
artifacts in `examples/`.
- Put security-relevant consumer responsibilities at the public boundary, not
only in internal contributor documents.
- Remove duplicated ownership instead of synchronizing parallel references.
- Make documentation checks reproducible where practical.
- Preserve existing behavior unless a stage explicitly selects and tests an
API change.
The roadmap does not add new Promptkit features, redesign framework formats,
or implement ideas from [the future feature catalog](future.md). If resolving
an ambiguity requires a behavioral change, treat that change as a separately
reviewable implementation unit and update its canonical documentation in the
same unit.
## Stage 1: Resolve Public Contract Questions
Before expanding prose, decide the intended contract for exported behavior
that is currently ambiguous.
- [x] Decide whether `RunRequest.Metadata` has a supported observable purpose.
Define its propagation and ownership, or remove or deprecate it through an
intentional public API change.
- [x] Decide and test whether one `Engine` supports concurrent `Prepare` and
`Run` calls.
- [x] Decide how repeated options of the same category behave, including
prompt, profile, schema, model-client, and artifact-reader options.
- [x] Define which public values have supported JSON representations.
- [x] Define JSON time units and omission behavior, including the relationship
between prepared-run and run-result durations.
- [x] Decide whether run IDs and exposed hashes have stable formats or must be
treated as opaque values.
- [x] Confirm the intended transport-timeout default and its zero or negative
configuration semantics.
- [x] Confirm the supported JSON Schema dialect and reference boundaries,
including whether remote references are allowed.
The selected exported API contracts are implemented and tested. Their durable
definitions now belong to the root package declarations and GoDoc.
One format-level decision remains here until Stage 5 moves it to the framework
format reference: JSON Schema uses Draft 2020-12, with that dialect selected
when `$schema` is omitted. Same-document fragments and relative references
contained by a directory or `fs.FS` schema root are supported. A single-file
source supports only references contained in that document. Absolute,
escaping, and remote references are rejected.
**Gate:** Each question has an explicit answer backed by existing behavior or
by an accepted implementation change and proportionate tests. No later stage
should invent a contract merely to fill a documentation gap.
## Stage 2: Make GoDoc The Canonical Public Contract
Strengthen the root package declarations so `go doc` is sufficient to
understand exact public behavior without relying on internal documents.
- [x] Add useful field-level GoDoc to configuration, request, profile,
execution-target, result, artifact, validation, structured-output, and model
client values.
- [x] Document required fields and nil, empty, and zero-value semantics.
- [x] Document override, replacement, profile-precedence, and copy-ownership
behavior where it belongs to the exported API.
- [x] Document credential inputs, redaction, and the values intentionally
excluded from serialization.
- [x] Give each public error sentinel an accurate comment and document the
supported `errors.Is` relationships.
- [x] Document engine concurrency and option-composition behavior selected in
Stage 1.
- [x] Document serialization, time, run-ID, and hash semantics selected in
Stage 1.
- [x] Review constructor, option, extension-interface, `Prepare`, and `Run`
GoDoc for complete failure and cancellation expectations.
Update the [consumer guide](../consumers/pkg-promptkit.md) to summarize and link
to these contracts instead of maintaining exhaustive copies of exported names
or exact semantics.
**Gate:** `go doc -all .` presents a coherent public contract, exported
declarations have accurate comments, and contract tests protect every newly
documented behavior whose compatibility risk warrants durable coverage.
## Stage 3: Improve Consumer Safety And Executable Guidance
Move consumer-relevant security boundaries to the places where consumers will
encounter them and add one representative execution workflow.
- [x] Explain in public GoDoc and the consumer guide that the default file
artifact reader accepts unrestricted caller-selected paths.
- [x] Make clear that Promptkit does not impose an application root, inbound
request-size policy, or untrusted-input security boundary.
- [x] Explain that rendered messages, artifact bodies, raw model output, and
validation details may be sensitive even when credentials are redacted.
- [x] Clarify the responsibilities of injected artifact readers and model
clients for cancellation, copying, logging, and secret handling.
- [x] Add a maintained offline `Run` example using an injected deterministic
model client, without credentials, live network access, or paid calls.
- [x] Link the consumer guide to the execution example and keep embedded
snippets smaller than the maintained artifact.
- [x] Decide whether the existing preparation example should remain separate
or share reusable fixtures without obscuring either workflow.
The preparation and execution examples remain separate, self-contained
workflows. Each keeps its own small prompt fixture so consumers can copy or run
one example without depending on the other.
**Gate:** Both preparation and execution have complete, secret-free,
deterministic consumer examples, and the consumer guide exposes the important
filesystem and data-sensitivity boundaries without leaking internal mechanics.
## Stage 4: Restore Canonical Ownership
Remove parallel definitions and make navigation follow the ownership model in
the documentation policy.
- [ ] Reduce the [architecture policy](../policy/architecture.md) to durable
boundaries, layers, dependency direction, invariants, and non-goals.
- [ ] Keep the exact implemented package and component inventory solely in the
[internal component overview](../internal/overview.md).
- [ ] Review the consumer guide's public error and option lists so they remain
task-oriented summaries rather than duplicate API references.
- [ ] Review internal documents for public-contract statements that should be
links to GoDoc or the format and integration owners.
- [ ] Reconcile the documentation policy's temporary-roadmap lifecycle with
the continuing idea-catalog role of `docs/roadmap/future.md`.
- [ ] Rephrase or link roadmap statements that depend on exact current API
behavior, particularly the runtime reasoning entry.
- [ ] Confirm that every document states its audience or purpose and links to
the canonical owner of adjacent topics.
**Gate:** Every authoritative fact has one clear owner, package inventory
changes no longer require edits to the architecture policy, and roadmaps
cannot be mistaken for current-state references.
## Stage 5: Refine Format And Integration References
Close compatibility gaps in the documents that own file formats and outbound
wire behavior.
- [ ] State or canonically link the exact session-ID limit enforced by the
OpenAI-compatible client.
- [ ] State the configured and default transport-timeout behavior without
referring to an unnamed internal default.
- [ ] Document the JSON Schema dialect and local, contained, and remote
reference behavior selected in Stage 1.
- [ ] Clarify structured-output naming and strictness when those values are
part of the public or integration contract.
- [ ] Add a caveat that built-in profiles are maintained configurations, not a
guarantee of continuing third-party model availability.
- [ ] Recheck every prompt, profile, schema, credential, request-body, response,
timeout, and precedence statement against its owning implementation and
tests.
**Gate:** A consumer can determine the supported file and wire compatibility
boundaries without consulting internal source code or relying on unspecified
defaults.
## Stage 6: Make Documentation Validation Reproducible
Align contributor and release procedures around a small, consistent set of
checks.
- [ ] Use one robust command for checking every tracked Go file with `gofmt`.
- [ ] Provide a repository-local or clearly documented command that validates
local Markdown targets and heading fragments.
- [ ] Decide how published external links are checked without making ordinary
validation depend on mutable network services.
- [ ] Reconcile the validation descriptions in the
[development guide](../development.md),
[testing policy](../policy/testing.md), and
[release procedure](../release.md) so one document owns each requirement.
- [ ] Ensure example validation covers every maintained example added by this
roadmap.
- [ ] Keep documentation-only validation proportionate while requiring full Go
validation when commands, examples, generated output, or checked behavior
changes.
**Gate:** A maintainer can run the documented formatting, link, example, Go,
and repository-hygiene checks exactly as written, with no hidden manual
procedure for local documentation.
## Completion Criteria
The roadmap is complete when:
- all Stage 1 contract questions are resolved;
- GoDoc is the authoritative and sufficient exported API reference;
- consumer guidance covers unrestricted file access and sensitive generated
data;
- maintained offline examples cover both `Prepare` and `Run`;
- architecture, inventory, consumer, internal, format, integration, and
roadmap documents follow their assigned ownership boundaries;
- schema, session, timeout, structured-output, and built-in-profile
compatibility statements are explicit;
- documentation validation is reproducible and consistent across contributor
and release workflows; and
- the complete maintainer validation passes.
After completion, move any durable decisions to GoDoc, policy, format,
integration, or ADR owners as appropriate. Remove this roadmap after incoming
links are updated; do not retain it as a second current-state reference.

View File

@@ -33,78 +33,7 @@ consumers.
## Ideas ## Ideas
### Extensible LLM backend registry No ideas currently await selection.
Introduce a registry that separates backend-specific connection,
authentication, and limited request defaults from model execution profiles.
Initial support would cover OpenAI-compatible backends and include a small
built-in catalog, potentially starting with OpenRouter. A profile could select
a backend while optionally overriding its default endpoint, and each backend
could name an optional environment variable for its API key without storing
the credential itself. Downstream consumers could register additional,
uniquely named backends, such as OpenAI or unauthenticated local-network
services, but could not replace built-in IDs. Model selection and generation
settings would remain profile concerns, and custom model clients would remain
available for behavior outside the registry's supported protocol.
### Backend-specific concurrency management
Extend the proposed LLM backend registry with optional per-backend concurrency
limits and bounded, buffered admission queues. Promptkit could then route
simultaneous generation requests according to backend capacity while
containing accidental runaway submission. Downstream consumers would continue
invoking synchronous `Run` calls, including concurrently from multiple
goroutines, and each admitted call would wait for and return its ordinary
result.
- Scope limits to an engine instance rather than hidden process-global state.
- Give different backend IDs independent capacity pools. A profile endpoint
override would remain part of its selected backend's pool.
- Configure active concurrency and waiting capacity separately. Concurrency
protects the backend, while queue capacity protects the process from
admitting an unbounded backlog.
- Give queue capacity a generous, configurable bounded default intended as a
safety ceiling for bugs or unintended loops rather than a routine
application constraint. Select an exact default during implementation
planning and measurement.
- Reject a call with a recognizable capacity error when its backend queue is
full rather than allowing it to wait outside the bounded queue.
- Admit requests before expensive preparation and artifact copying where
practical so queued work remains lightweight.
- Apply a limit to each actual generation request, including repair attempts,
without unnecessarily serializing prompt preparation.
- Make queued and active waits respect caller cancellation and deadlines.
- Treat concurrency as backend policy rather than a profile-level model
setting.
- Keep the queue ephemeral and in-process, with no survival guarantee across
engine or process shutdown.
- Preserve the existing execution model as far as practical. Durable jobs,
polling, priorities, application worker lifecycle, retries, and
cross-process coordination would be separate future capabilities.
### Explicit per-run session and reasoning controls
Allow consumers to associate a session ID with each run and to inherit,
replace, or explicitly disable the reasoning effort configured by its selected
profile. Prompt definitions can currently derive a session ID from a template,
and a non-empty runtime `ReasoningEffort` can replace the profile value, but
there is no direct request-level session ID and an empty reasoning value means
that no override was supplied. These controls would let consumers reuse one
prompt and model profile across sessions and reasoning levels without
maintaining duplicate definitions.
- Preserve a prompt's session ID template and a profile's reasoning effort as
reusable defaults.
- Let a directly supplied per-run session ID take precedence over a rendered
prompt default while retaining the existing validation limit and outbound
representation.
- Distinguish an omitted runtime reasoning choice from an explicit request to
disable reasoning.
- Ensure disabling reasoning omits the corresponding provider request setting
rather than relying on a provider-specific magic value.
- Keep the effective session ID and reasoning choice visible in prepared and
run metadata and available to injected model clients without introducing
additional prompt or profile selection mechanisms.
## Entry Format ## Entry Format

View File

@@ -0,0 +1,245 @@
# Notarius PromptKit Wishlist
## Purpose
This document records features and interface changes that would be useful
additions to PromptKit from the perspective of the maintainers of Notarius, a
downstream application that consumes PromptKit.
PromptKit now provides the capabilities Notarius currently needs. The
remaining deferred ideas are optional opportunities to improve checkpointing
and operational observability.
The examples are API sketches intended to communicate the desired capability,
not prescriptive names or finalized Go contracts.
## Priority 1: Atomic Execution With Prepared Details
**Disposition:** Implemented through [`Engine.PrepareExecution` and
`Engine.RunPrepared`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#prepare-now-and-execute-the-same-snapshot-later).
A separate `RunDetailed` method is not cataloged.
### Downstream need
Notarius needs both:
- the completed `RunResult`; and
- the rendered messages, effective output contract, hashes, and other
preparation details exposed by `PreparedRun`.
Notarius uses the prepared details to construct redaction-aware debug bundles
and retain enough information to diagnose model behavior.
### Implemented behavior
Notarius can prepare one frozen execution snapshot, retain a caller-owned and
credential-redacted `Details` value for its debug bundle, and execute the same
snapshot through `RunPrepared`. The opaque handle is engine-bound and
single-use; an unused handle can be released with `Discard`. The consumer
guide and exported GoDoc own the exact lifecycle and failure contracts.
### Value to Notarius
This removes duplicate work from PromptKit-backed calls and ensures that
retained debug material corresponds atomically to the actual execution.
## Priority 2: Prompt-Independent Profile Inspection
**Disposition:** Implemented as
[`Engine.InspectProfile`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#inspect-a-profile-before-prompt-work).
### Downstream need
Notarius validates configured pipeline profile IDs before beginning a run. It
needs to determine whether:
- a profile exists;
- its referenced backend is registered;
- its execution target can be resolved; and
- it declares a credential requirement that the application may need to
enforce.
This validation should not require model generation.
### Previous integration
Before profile inspection was available, Notarius constructed a synthetic
prompt using `testing/fstest.MapFS`, supplied a dummy transcript, and called
`Engine.Prepare` solely to exercise profile and backend resolution.
### Value to Notarius
The implemented interface eliminates a synthetic production-only prompt
fixture and establishes a direct, supported contract for configuration-time
profile and backend validation.
## Priority 3: Semantic Execution-Target Fingerprints
**Disposition:** Deferred pending a separate semantic-equality design for
resolved execution targets.
### Downstream need
Notarius checkpoints model-backed pipeline stages. A checkpoint must not be
reused when generation-affecting PromptKit configuration changes.
Notarius therefore needs a stable equality signal for the effective profile
and backend target used by a pipeline.
### Current integration
Notarius currently constructs this identity itself from:
- a manually maintained marker for the PromptKit release and built-in profile
catalog;
- raw hashes of configured profile files; and
- a separate hash of the configured conventional local-backend endpoint.
This is safe but conservative and coupled to PromptKit details. Raw file
hashing also invalidates checkpoints for semantically irrelevant YAML changes,
such as comments or formatting.
### Requested capability
Expose an opaque semantic digest for a resolved profile and its effective
generation target. It could be returned by the proposed profile-resolution
API:
```go
type ResolvedProfile struct {
ProfileID string
BackendID string
EffectiveTarget ExecutionTarget
ExecutionDigest string
}
```
Alternatively, PromptKit could expose a dedicated method such as
`ProfileExecutionDigest(profileID)`.
### Desired equality semantics
The digest should change when generation-affecting state changes, including:
- resolved model and endpoint;
- backend routing identity;
- backend request defaults and extra parameters;
- profile generation parameters; and
- the semantic identity of any selected built-in profile.
The digest should not incorporate:
- credential values;
- concurrency or queue capacity;
- filesystem source paths;
- YAML comments or formatting; or
- other settings that affect scheduling or source representation without
changing the generation target.
The credential environment-variable name may need to participate if changing
it can select a materially different provider account or target. PromptKit
should define this deliberately while continuing to exclude the resolved
secret value.
### Design considerations
- Treat the digest as an opaque equality value rather than a public encoding
of internal structures.
- Document which categories of change affect equality.
- Include a versioned semantic marker internally so PromptKit can deliberately
invalidate old digests when its resolution semantics change.
- Prefer a per-profile digest over a digest of every profile known to an
engine. Notarius generally knows which profiles a resolved pipeline uses.
- Do not require consumers to know PromptKit's built-in catalog version.
### Value to Notarius
This would let Notarius remove its PromptKit release marker and raw
profile-source fingerprinting, reduce unnecessary checkpoint invalidation, and
delegate execution-target equality to the component that owns target
resolution.
## Priority 4: Structured Capacity Errors
**Disposition:** Implemented behavior. See the consumer guide's
[Handle Errors](../consumers/pkg-promptkit.md#handle-errors) section.
### Downstream need
Notarius translates PromptKit backend-capacity rejection into a
provider-neutral application error. When multiple backends are active,
operators would benefit from knowing which backend rejected admission without
parsing an error string or exposing endpoint details.
### Implemented behavior
PromptKit now retains the broad capacity classification while allowing
Notarius to obtain the selected backend ID without parsing diagnostic text.
The consumer guide owns the application workflow, including retry and backoff
policy.
### Value to Notarius
This improves operational diagnostics and future metrics while preserving
the provider-neutral error boundary used by Notarius.
## Capabilities PromptKit Already Provides Well
The current PromptKit boundary is sufficient for Notarius's implemented
behavior. In particular, PromptKit already provides:
- filesystem, `fs.FS`, and programmatic prompt, profile, and schema sources;
- offline preparation without model execution;
- structured output and content validation;
- direct session propagation;
- tri-state per-run reasoning overrides;
- selected profile, backend, model, endpoint, effective parameters, hashes, and
token-usage provenance;
- endpoint-only profiles;
- the conventional `local` backend helper;
- arbitrary engine-scoped `Backend` registrations;
- backend authentication environment names, extra parameters, concurrency
limits, and queue-capacity policies;
- provider-client and artifact-reader extension interfaces;
- context cancellation; and
- useful public error sentinels, including profile absence and capacity
exhaustion.
The wishlist does not imply that Notarius needs PromptKit to broaden its core
responsibilities. It primarily asks for more direct access to information and
operations that PromptKit already computes internally.
## Responsibilities That Should Remain In Notarius
The following concerns belong to the downstream application and should not
move into PromptKit for the sake of Notarius:
- pipeline staging, dependencies, and generated references;
- application-wide scheduling across providers and backends;
- module and validation retry policy;
- checkpoints, resume, and recomputation;
- durable run artifacts and manifests;
- D&D prompts, schemas, extractors, validators, and normalizers;
- Notarius configuration-file parsing and precedence;
- domain-specific prompt-cache prefix policy; and
- application-specific redaction, retention, and debug-bundle policy.
PromptKit's complete `Backend` API already supports custom IDs, multiple local
endpoints, authentication, extra parameters, and explicit queue policies.
Whether Notarius exposes those capabilities in its own configuration is an
application-policy decision, not an upstream PromptKit gap.
## Suggested Upstream Sequence
For downstream adoption and any remaining upstream work, the useful order is:
1. Adopt prepared execution for atomic details and results.
2. Add a semantic execution-target digest, preferably alongside profile
inspection.
3. Use the implemented typed capacity error where backend admission diagnostics
are needed.
The first removes the concrete execution workaround. The second would improve
checkpoint correctness and reduce coupling. The third is operational polish.

View File

@@ -0,0 +1,331 @@
# Weatherreporter PromptKit Wishlist
## Purpose
This document records features and interface changes that would be useful
additions to PromptKit from the perspective of the maintainers of
Weatherreporter, a downstream application planning to replace its Scriptorium
CLI integration with PromptKit.
PromptKit now provides the capabilities Weatherreporter needs for the
migration. The remaining deferred ideas are optional opportunities to validate
configuration earlier and improve durable failure diagnostics.
The examples are API sketches intended to communicate the desired capability,
not prescriptive names or finalized Go contracts. The related
[Notarius PromptKit wishlist](notarius-promptkit-wishlist.md) proposes several
overlapping features from another downstream consumer's perspective.
## Priority 1: Executable Preparation Handles
**Disposition:** Implemented as [`Engine.PrepareExecution` and
`Engine.RunPrepared`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#prepare-now-and-execute-the-same-snapshot-later).
### Downstream need
Weatherreporter treats prompt preparation as a durable preflight boundary. It
needs to:
1. prepare the exact request that will be executed;
2. persist a safe preparation record before starting the provider call; and
3. execute without reloading or rerendering prompt, profile, schema, or input
sources.
Persisting preflight before generation leaves useful evidence when a provider
call fails or the process is interrupted during generation.
### Implemented behavior
Weatherreporter can prepare one frozen execution snapshot, persist a
caller-owned and credential-redacted `Details` value, and execute that same
snapshot through `RunPrepared`. The opaque handle is engine-bound and
single-use; an unused handle can be released with `Discard`. The consumer
guide and exported GoDoc own the exact lifecycle, credential, cancellation,
and capacity contracts.
### Value to Weatherreporter
This preserves Weatherreporter's durable preflight behavior, removes duplicate
work, eliminates the source-consistency window, and ensures that persisted
provenance describes the actual execution.
## Priority 2: Prompt-Definition Inspection
**Disposition:** Implemented as
[`Engine.InspectPrompt`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#inspect-a-prompt-before-preparation).
### Downstream need
Weatherreporter has a fixed registry of seven report definitions. Each report
selects a prompt ID and one of two output workflows:
- direct Markdown; or
- structured generated text followed by application-owned domain validation
and Markdown template rendering.
Weatherreporter will embed the PromptKit prompt definitions and private
response schemas that implement those reports. It needs to validate that the
report registry and embedded prompt corpus agree before weather collection or
provider execution.
### Previous integration option
Before prompt inspection was available, Weatherreporter could maintain
synthetic data-package fixtures and call `Engine.Prepare` for every report
prompt during tests. Runtime validation could also occur through the ordinary
per-report preparation stage.
This required complete placeholder inputs and profile resolution when the
application primarily wanted to inspect prompt identity and declared contracts.
### Value to Weatherreporter
The implemented interface lets Weatherreporter directly verify that every
report prompt exists, requires the curated `data_package` input, and declares
the expected Markdown or JSON Schema output contract. It reduces synthetic
test setup and moves failures ahead of weather collection.
## Priority 3: Prompt-Independent Profile Inspection
**Disposition:** Implemented as
[`Engine.InspectProfile`](../../engine.go). See the
[consumer guidance](../consumers/pkg-promptkit.md#inspect-a-profile-before-prompt-work).
### Downstream need
Weatherreporter will allow operators to select an external PromptKit profile
source and may allow an explicit profile override. It needs to reject a missing
profile, unknown backend, or malformed execution target before collecting
weather data or writing report artifacts, and to apply its own policy to a
reported credential requirement.
### Value to Weatherreporter
The implemented interface improves fail-fast configuration validation and gives
operator-facing errors direct profile and backend context. It remains optional
for the initial migration.
## Priority 4: Eager Source Validation
**Disposition:** Deferred until prompt and profile inspection have been used
to determine whether a broader engine-wide validation operation is still
needed.
### Downstream need
PromptKit deliberately defers reading and validating filesystem and `fs.FS`
prompt, profile, and schema content until a request needs it. Weatherreporter
has a small fixed embedded prompt corpus and one optional external profile
source. It would benefit from an explicit offline validation operation for
tests, startup diagnostics, and configuration checks.
### Current integration option
Weatherreporter can prepare every report prompt with fixture inputs and inspect
any explicit profiles individually. That provides strong coverage but requires
consumer-maintained traversal and synthetic material.
### Requested capability
Consider an opt-in source-validation operation:
```go
type SourceValidationOptions struct {
RequireCredentials bool
}
func (e *Engine) ValidateSources(
ctx context.Context,
opts SourceValidationOptions,
) error
```
The operation should eagerly discover and structurally validate the configured
prompt, profile, and schema sources without model generation.
### Design considerations
- Keep deferred validation as the normal `NewEngine` behavior.
- Make eager validation an explicit consumer choice.
- Validate duplicate IDs and versions, strict YAML decoding, referenced content
files, profile/backend membership, schema syntax, and schema references.
- Distinguish structural credential declarations from current environment
availability.
- Do not read or expose credential values when credential availability is not
requested.
- Preserve source-specific public error identities and useful path context.
- Respect context cancellation during filesystem discovery and schema work.
- Consider whether exact prompt and profile inspection APIs already provide a
smaller sufficient surface before adding an engine-wide operation.
### Value to Weatherreporter
This would simplify offline corpus checks and catch malformed operator profile
sources before report work begins. It is helpful but lower priority than exact
prompt and profile inspection.
## Priority 5: Structured Generation Errors
**Disposition:** Deferred pending stronger downstream demand and a narrower
design that does not duplicate prepared provenance or impose HTTP-specific
fields on injected model clients.
### Downstream need
Weatherreporter preserves redacted, inspectable failure receipts for report
runs. When model generation fails operationally, it needs to classify the
failure and retain safe execution context without parsing error prose.
Prompt preparation already supplies selected profile, backend, and model
identity. Provider status classification would add useful operator context,
especially when the built-in OpenAI-compatible client receives a non-success
HTTP status.
### Current integration option
PromptKit exposes `ErrLLMGenerate` and preserves injected client errors through
`errors.Is`. Weatherreporter can reliably classify generation failure and use
its preparation record for profile, backend, and model provenance. Any further
diagnostic detail remains a redacted error string.
### Requested capability
Consider a typed generation error that continues to match `ErrLLMGenerate`:
```go
type GenerationError struct {
BackendID string
Model string
StatusCode int
}
```
The exact fields may differ. The useful contract is safe structured context
available through `errors.As`, while `errors.Is(err, ErrLLMGenerate)` remains
compatible.
### Design considerations
- Include only fields that PromptKit knows reliably and can expose safely.
- Treat an HTTP status as optional because injected model clients may not use
HTTP.
- Do not expose provider response bodies, endpoints, credential environment
names, credential values, request content, or generated content.
- Do not make a structured error a second source of prompt/profile provenance
already present in a prepared execution.
- Preserve injected client error identity.
- Keep retry and backoff policy with the consuming application.
### Value to Weatherreporter
This would improve durable failure receipts and troubleshooting, particularly
for built-in transport failures. It is not required if preparation details and
the existing sentinel remain available.
## Lower-Priority Shared Wishlist Items
### Structured Capacity Errors
**Disposition:** Implemented behavior. See the consumer guide's
[Handle Errors](../consumers/pkg-promptkit.md#handle-errors) section.
PromptKit now exposes the stable backend ID for capacity rejection without
requiring Weatherreporter to parse error text.
Weatherreporter currently generates batch reports sequentially and constructs
one engine per invocation, so engine-local capacity exhaustion is unlikely in
the initial design. The typed error becomes more valuable if report generation
later becomes concurrent or PromptKit engines become longer-lived.
It should not block adoption.
### Semantic Execution-Target Fingerprints
**Disposition:** Deferred pending a separate semantic-equality design for
resolved execution targets.
The semantic target digest proposed by the
[Notarius wishlist](notarius-promptkit-wishlist.md#priority-3-semantic-execution-target-fingerprints)
would provide a compact equality signal for audit metadata.
Weatherreporter does not currently reuse LLM-dependent checkpoints. Its Recent
Changes behavior compares deterministic module snapshots rather than generated
reports, so the digest has no immediate cache-correctness role. Existing
PromptKit result metadata is sufficient for the initial integration. A digest
would still be useful provenance and future-proofing, but it is not a
migration priority.
## Capabilities PromptKit Already Provides Well
PromptKit already provides the essential Weatherreporter integration surface:
- importable in-process engine construction;
- filesystem, `fs.FS`, and programmatic prompt, profile, and schema sources;
- offline preparation without model execution;
- prepared execution handles for a durable preflight boundary;
- exact prompt and profile inspection;
- versioned prompt selection;
- text, Markdown, JSON, and JSON Schema output contracts;
- single-pass output validation with raw output retained after completed
validation failure;
- inline input artifacts with provenance URIs and input hashes;
- selected profile, backend, model, effective target, prompt hashes, timing,
and token-usage provenance;
- endpoint-only profiles and engine-scoped backend registration;
- injected model-client and artifact-reader interfaces;
- caller cancellation, generation timeout, and transport timeout behavior; and
- public error sentinels for configuration, prompt, profile, artifact,
validation, capacity, and generation failures; and
- structured backend identity for capacity rejection.
These capabilities are sufficient for Weatherreporter to adopt PromptKit
without waiting for new upstream work.
## Responsibilities That Should Remain In Weatherreporter
The following concerns belong to Weatherreporter and should not move into
PromptKit:
- report definitions, valid periods, batches, and output naming;
- application prompt content and private report response schemas;
- deterministic weather facts, modules, and Recent Changes;
- curated `data_package` construction and persistence;
- generated-text domain validation and Markdown template rendering;
- managed artifact paths, atomic writes, metadata, and inspection commands;
- preparation, execution, raw-output, and failure-receipt schemas;
- CLI configuration loading and precedence;
- debug enablement, redaction, placement, sensitivity, and retention;
- distributor notification;
- batch continuation and any future retry policy; and
- application-level compatibility and migration policy.
## Suggested Upstream Sequence
For downstream adoption and any remaining upstream work, the useful order is:
1. Adopt the implemented executable preparation handles.
2. Consider eager source validation after evaluating whether the two exact
inspection APIs are sufficient.
3. Add structured generation errors.
4. Use structured capacity errors and consider semantic execution-target
fingerprints as lower-priority operational improvements.
The first item removes the material integration workaround. Prompt and profile
inspection improve fail-fast validation. The remaining items are optional
ergonomic and diagnostic improvements.
## Adoption Sequencing
Weatherreporter should not wait for the deferred wishlist items. The current
PromptKit interface is sufficient when Weatherreporter:
- embeds immutable prompt and schema assets;
- supplies immutable inline data-package bytes;
- constructs one engine per CLI invocation;
- prepares an execution, persists selected `Details`, and calls
`RunPrepared`; and
- keeps PromptKit behind a weatherreporter-owned adapter contract.
Prompt inspection, profile inspection, source validation, structured errors,
capacity details, and semantic fingerprints should not gate adoption.

249
engine.go
View File

@@ -12,7 +12,10 @@ import (
"time" "time"
artifactadapter "gitea.maximumdirect.net/eric/promptkit/internal/artifact" artifactadapter "gitea.maximumdirect.net/eric/promptkit/internal/artifact"
"gitea.maximumdirect.net/eric/promptkit/internal/backend"
"gitea.maximumdirect.net/eric/promptkit/internal/capacity"
"gitea.maximumdirect.net/eric/promptkit/internal/defaults" "gitea.maximumdirect.net/eric/promptkit/internal/defaults"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/llm" "gitea.maximumdirect.net/eric/promptkit/internal/llm"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
"gitea.maximumdirect.net/eric/promptkit/internal/profile/builtin" "gitea.maximumdirect.net/eric/promptkit/internal/profile/builtin"
@@ -23,7 +26,8 @@ import (
) )
// ErrInvalidConfig identifies invalid engine construction, including missing // ErrInvalidConfig identifies invalid engine construction, including missing
// required configuration, invalid options, and a nil Engine receiver. // required configuration, invalid options or backend registrations, and a nil
// Engine receiver.
var ErrInvalidConfig = errors.New("invalid engine configuration") var ErrInvalidConfig = errors.New("invalid engine configuration")
var ( var (
@@ -46,8 +50,8 @@ var (
// ErrPromptNotFound. // ErrPromptNotFound.
ErrPromptLoad = errors.New("failed to load prompt definition") ErrPromptLoad = errors.New("failed to load prompt definition")
// ErrProfileLoad identifies a failure to read, decode, validate, or select // ErrProfileLoad identifies a failure to read, decode, validate, or select
// an execution profile, except for the not-found case represented by // an execution profile or resolve its backend, except for the profile
// ErrProfileNotFound. // not-found case represented by ErrProfileNotFound.
ErrProfileLoad = errors.New("failed to load execution profile") ErrProfileLoad = errors.New("failed to load execution profile")
// ErrAPIKeyEnvMissing identifies an APIKeyEnv whose environment variable is // ErrAPIKeyEnvMissing identifies an APIKeyEnv whose environment variable is
// unset or empty when no direct RunRequest.APIKey takes precedence. Such an // unset or empty when no direct RunRequest.APIKey takes precedence. Such an
@@ -59,6 +63,11 @@ var (
// ErrPromptRender identifies a failure to render prompt messages or the // ErrPromptRender identifies a failure to render prompt messages or the
// session ID from the resolved inputs and variables. // session ID from the resolved inputs and variables.
ErrPromptRender = errors.New("failed to render prompt") ErrPromptRender = errors.New("failed to render prompt")
// ErrCapacityExceeded identifies a Run or RunPrepared rejected because the
// selected backend already admitted ConcurrencyLimit + QueueCapacity calls.
// A [CapacityError] reports the selected backend ID. It is not an invalid
// request, an LLM or provider rate-limit response, or ErrLLMGenerate.
ErrCapacityExceeded = errors.New("backend capacity exceeded")
// ErrLLMGenerate identifies a model-client failure or a nil successful // ErrLLMGenerate identifies a model-client failure or a nil successful
// response. Errors returned by an injected LLMClient remain available // response. Errors returned by an injected LLMClient remain available
// through errors.Is. // through errors.Is.
@@ -69,10 +78,15 @@ var (
ErrValidation = errors.New("failed to validate output") ErrValidation = errors.New("failed to validate output")
) )
// Engine prepares and runs Promptkit prompt requests. // Engine inspects prompts and profiles and prepares and runs Promptkit prompt
// requests.
// //
// An Engine is safe for concurrent calls to [Engine.Prepare] and [Engine.Run]. // An Engine is safe for concurrent calls to [Engine.InspectPrompt],
// Injected collaborators may consequently be invoked concurrently. // [Engine.InspectProfile], [Engine.Prepare], [Engine.PrepareExecution],
// [Engine.Run], and [Engine.RunPrepared]. Each Engine owns independent
// backend-capacity pools that coordinate Run and RunPrepared admission and
// model generation. Injected collaborators may still be invoked concurrently
// across different backend pools or for unlimited backends.
type Engine struct { type Engine struct {
runner *usecase.Runner runner *usecase.Runner
} }
@@ -107,8 +121,10 @@ type Config struct {
// NewEngine applies options in argument order and ignores nil options. Within // NewEngine applies options in argument order and ignores nil options. Within
// each prompt-source, profile-source, in-memory-profile, schema-source, // each prompt-source, profile-source, in-memory-profile, schema-source,
// model-client, and artifact-reader category, the last non-nil valid option // model-client, and artifact-reader category, the last non-nil valid option
// replaces earlier options in that category. An invalid option fails // replaces earlier options in that category. WithBackend is the additive
// construction even if a later option would replace it. // exception: unique registrations accumulate, and a repeated backend ID is an
// error rather than a replacement. An invalid option fails construction even
// if a later option would replace it.
type Option interface { type Option interface {
apply(*engineOptions) error apply(*engineOptions) error
} }
@@ -125,6 +141,7 @@ type engineOptions struct {
promptDefs promptdef.Repository promptDefs promptdef.Repository
profiles profile.Repository profiles profile.Repository
memoryProfiles profile.Repository memoryProfiles profile.Repository
backends []domain.Backend
validator validate.Validator validator validate.Validator
promptSource bool promptSource bool
profileSource bool profileSource bool
@@ -133,10 +150,14 @@ type engineOptions struct {
artifactSource bool artifactSource bool
} }
// WithLLMClient replaces the built-in model client used by [Engine.Run]. // WithLLMClient replaces the built-in model client used by [Engine.Run] and
// [Engine.RunPrepared].
// //
// A nil client makes NewEngine fail with ErrInvalidConfig. The client may be // A nil client makes NewEngine fail with ErrInvalidConfig. The Engine schedules
// called concurrently and is not used by [Engine.Prepare]. // Generate calls according to the selected backend's capacity policy, but the
// client may still be called concurrently across different backend pools or for
// unlimited backends. The client is not used by [Engine.Prepare] or
// [Engine.PrepareExecution].
func WithLLMClient(client LLMClient) Option { func WithLLMClient(client LLMClient) Option {
return optionFunc(func(options *engineOptions) error { return optionFunc(func(options *engineOptions) error {
if client == nil { if client == nil {
@@ -302,12 +323,14 @@ func WithSchemaFile(path string) Option {
// //
// Options are applied in order according to [Option]. PromptDir is required // Options are applied in order according to [Option]. PromptDir is required
// unless a prompt-source option is present. Construction validates option // unless a prompt-source option is present. Construction validates option
// arguments and in-memory profiles but defers reading and validating prompt, // arguments, in-memory profiles, and backend registrations but defers reading
// file-backed profile, and schema contents until Prepare or Run needs them. // and validating prompt, file-backed profile, and schema contents until Prepare
// or Run needs them.
// //
// NewEngine returns an error matching ErrInvalidConfig for invalid // NewEngine returns an error matching ErrInvalidConfig for invalid
// configuration or options. It does not perform model requests or require // configuration, options, or backend-capacity policies. Each constructed
// credentials. // Engine has independent backend-capacity pools. Construction does not perform
// model requests or require credentials.
func NewEngine(cfg Config, opts ...Option) (*Engine, error) { func NewEngine(cfg Config, opts ...Option) (*Engine, error) {
var options engineOptions var options engineOptions
for _, opt := range opts { for _, opt := range opts {
@@ -335,6 +358,16 @@ func NewEngine(cfg Config, opts ...Option) (*Engine, error) {
profiles = profile.NewOverlayRepository(options.memoryProfiles, profiles) profiles = profile.NewOverlayRepository(options.memoryProfiles, profiles)
} }
backendRegistry, err := backend.NewRegistry(options.backends)
if err != nil {
return nil, fmt.Errorf("%w: failed to construct backend registry: %v", ErrInvalidConfig, err)
}
capacityManager, err := capacity.NewManager(backendRegistry.CapacityPolicies())
if err != nil {
return nil, fmt.Errorf("%w: failed to construct backend capacity manager: %v", ErrInvalidConfig, err)
}
validator := options.validator validator := options.validator
if !options.validatorSource { if !options.validatorSource {
schemaDir := cfg.SchemaDir schemaDir := cfg.SchemaDir
@@ -355,6 +388,7 @@ func NewEngine(cfg Config, opts ...Option) (*Engine, error) {
return nil, fmt.Errorf("%w: %v", ErrInvalidConfig, err) return nil, fmt.Errorf("%w: %v", ErrInvalidConfig, err)
} }
} }
llmClient = capacity.NewClient(capacityManager, llmClient)
artifacts := options.artifactReader artifacts := options.artifactReader
if !options.artifactSource { if !options.artifactSource {
@@ -365,10 +399,12 @@ func NewEngine(cfg Config, opts ...Option) (*Engine, error) {
runner: usecase.NewRunner( runner: usecase.NewRunner(
promptDefs, promptDefs,
profiles, profiles,
backendRegistry,
artifacts, artifacts,
prompt.NewGoRenderer(), prompt.NewGoRenderer(),
llmClient, llmClient,
validator, validator,
capacityManager,
), ),
}, nil }, nil
} }
@@ -393,13 +429,101 @@ func fileSource(name string) (fs.FS, string, error) {
return os.DirFS(dir), filepath.ToSlash(base), nil return os.DirFS(dir), filepath.ToSlash(base), nil
} }
// InspectPrompt resolves one explicit prompt definition without selecting a
// profile or starting execution work.
//
// InspectPrompt requires a nonblank promptID. It passes nonblank promptID and
// promptVersion values unchanged to the engine's ordinary, case-sensitive
// prompt selection. An empty version succeeds only when that source has one
// selected ID; a nonempty version selects one exact ID/version pair. The
// configured prompt source is used without merging, fallback, or enumeration.
//
// A successful result proves that the selected definition and any referenced
// message content files were structurally loaded. Inputs are returned in
// definition order. DefaultProfileID is declared metadata only and is not
// resolved. OutputContract is the normalized declared contract, with a JSON
// Schema path when declared but without loading or compiling that schema.
// PromptHash is the same opaque equality value as PreparedRun.PromptHash for
// the selected definition and observed source state; its spelling, length,
// encoding, algorithm, and security properties are not contracts.
//
// This method does not return prompt bodies, templates, source paths, schemas,
// rendered messages, or execution settings. It does not resolve a profile or
// credential, read artifacts or schemas, render, validate, admit capacity,
// contact a provider, or generate model output. The returned PromptInspection
// and its input slice are caller-owned. Filesystem-backed inspection is a
// point-in-time lookup and does not freeze a definition for later execution.
//
// A nil Engine returns an error matching ErrInvalidConfig. A blank prompt ID
// matches ErrInvalidRequest. An absent exact ID or version matches
// ErrPromptNotFound and not ErrPromptLoad. Malformed, unreadable, duplicate,
// ambiguous, referenced-content, or hashing failures match ErrPromptLoad.
// Cancellation during lookup matches ErrPromptLoad while preserving the
// context error. InspectPrompt returns no partial result on error.
func (e *Engine) InspectPrompt(
ctx context.Context,
promptID string,
promptVersion string,
) (*PromptInspection, error) {
if e == nil || e.runner == nil {
return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig)
}
inspection, err := e.runner.InspectPrompt(ctx, promptID, promptVersion)
if err != nil {
return nil, mapPublicError(err)
}
return fromDomainPromptInspection(inspection), nil
}
// InspectProfile resolves one explicit profile without selecting a prompt or
// starting execution work.
//
// InspectProfile trims surrounding whitespace from profileID and looks up the
// resulting nonblank ID exactly and case-sensitively through the engine's
// ordinary in-memory, configured-source, and built-in profile precedence. It
// applies framework defaults, the selected backend, and then the selected
// profile to EffectiveModelParams without a request override. BackendID is
// empty for an endpoint-only profile.
//
// APIKeyEnv in the returned target is an environment-variable name, never its
// value. APIKeyRequired instead reports a direct credential requirement and is
// mutually exclusive with a nonblank APIKeyEnv. InspectProfile neither derives
// an ID from a prompt default_profile nor checks credential availability, so an
// absent or blank named environment variable is not an error.
//
// The returned ProfileInspection and all nested mutable values are
// caller-owned. Filesystem-backed inspection is a point-in-time lookup and
// does not freeze the profile for a later execution. This method does not load
// a prompt, render, read artifacts or schemas, admit backend capacity, contact
// a provider, or generate model output.
//
// A nil Engine returns an error matching ErrInvalidConfig. A blank profile ID
// matches ErrInvalidRequest. An absent exact ID matches ErrProfileNotFound and
// not ErrProfileLoad. Malformed or unreadable profile data, an unknown backend,
// or an invalid resolved target matches ErrProfileLoad. Cancellation during
// profile loading matches ErrProfileLoad while preserving the context error.
// InspectProfile returns no partial result on error.
func (e *Engine) InspectProfile(ctx context.Context, profileID string) (*ProfileInspection, error) {
if e == nil || e.runner == nil {
return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig)
}
inspection, err := e.runner.InspectProfile(ctx, profileID)
if err != nil {
return nil, mapPublicError(err)
}
return fromDomainProfileInspection(inspection), nil
}
// Prepare resolves and renders a prompt request without calling an LLM. // Prepare resolves and renders a prompt request without calling an LLM.
// //
// Prepare selects the prompt and profile, resolves effective execution // Prepare selects the prompt and profile, resolves any selected backend and
// settings and the output contract, loads and hashes inputs, loads structured // effective execution settings, resolves the output contract, loads and hashes
// output schema metadata when required, and renders the session ID and // inputs, loads structured-output schema metadata when required, and renders
// messages. The returned PreparedRun is owned by the caller and never contains // the session ID and messages. The returned PreparedRun is owned by the caller
// a resolved API-key value, model output, or validation result. // and never contains a resolved API-key value, model output, or validation
// result.
// //
// A nil Engine returns an error matching ErrInvalidConfig. Request and // A nil Engine returns an error matching ErrInvalidConfig. Request and
// preparation failures may match ErrInvalidRequest, ErrPromptNotFound, // preparation failures may match ErrInvalidRequest, ErrPromptNotFound,
@@ -426,6 +550,38 @@ func (e *Engine) Prepare(ctx context.Context, req RunRequest) (*PreparedRun, err
return fromDomainPreparedRun(prepared), nil return fromDomainPreparedRun(prepared), nil
} }
// PrepareExecution completely prepares a prompt request without calling the
// configured LLMClient or reserving backend admission capacity.
//
// The returned opaque handle is bound to this Engine and permits one
// [Engine.RunPrepared] invocation. Preparation freezes the selected sources,
// rendered messages, effective settings, inputs, provider structured-output
// metadata, and validation resources needed by that invocation. The handle
// retains a direct RunRequest.APIKey only in private execution state;
// [PreparedExecution.Details] is credential-redacted.
//
// The context governs preparation only. Cancellation after this method
// returns does not invalidate the handle or propagate to RunPrepared.
// PrepareExecution returns the same error categories as [Engine.Prepare] and
// returns no handle on error. A nil Engine returns an error matching
// ErrInvalidConfig.
func (e *Engine) PrepareExecution(ctx context.Context, req RunRequest) (*PreparedExecution, error) {
if e == nil || e.runner == nil {
return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig)
}
domainReq, err := toDomainRunRequest(req)
if err != nil {
return nil, fmt.Errorf("%w: %v", ErrInvalidRequest, err)
}
prepared, err := e.runner.PrepareExecution(ctx, domainReq)
if err != nil {
return nil, mapPublicError(err)
}
return &PreparedExecution{internal: prepared}, nil
}
// Run prepares a request, invokes the configured LLMClient, and validates the // Run prepares a request, invokes the configured LLMClient, and validates the
// generated output. // generated output.
// //
@@ -436,11 +592,15 @@ func (e *Engine) Prepare(ctx context.Context, req RunRequest) (*PreparedRun, err
// single-pass even when OutputContract.RepairAttempts is positive. // single-pass even when OutputContract.RepairAttempts is positive.
// //
// Run can return every error category documented by [Engine.Prepare], plus // Run can return every error category documented by [Engine.Prepare], plus
// ErrLLMGenerate. Errors from injected clients remain available through // ErrCapacityExceeded and ErrLLMGenerate. An engine admission rejection is
// errors.Is. Cancellation is passed through the active collaborator and is // discoverable as [CapacityError] and still matches ErrCapacityExceeded. It
// reported in the applicable operation category; no general errors.Is // occurs before artifacts, schemas, rendering, or model generation because the
// relationship to ctx.Err is promised. A nil Engine returns ErrInvalidConfig. // selected backend's admission capacity is full; it does not match
// Run returns no partial result on error. // ErrInvalidRequest or ErrLLMGenerate. Errors from injected clients remain
// available through errors.Is. Cancellation while waiting for model-generation
// capacity matches both ErrLLMGenerate and the context error. Cancellation
// otherwise follows the active collaborator's documented behavior. A nil
// Engine returns ErrInvalidConfig. Run returns no partial result on error.
func (e *Engine) Run(ctx context.Context, req RunRequest) (*RunResult, error) { func (e *Engine) Run(ctx context.Context, req RunRequest) (*RunResult, error) {
if e == nil || e.runner == nil { if e == nil || e.runner == nil {
return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig) return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig)
@@ -457,3 +617,42 @@ func (e *Engine) Run(ctx context.Context, req RunRequest) (*RunResult, error) {
} }
return fromDomainRunResult(result), nil return fromDomainRunResult(result), nil
} }
// RunPrepared atomically claims and executes a handle created by
// [Engine.PrepareExecution].
//
// A valid owning-Engine invocation consumes the handle's one attempt before
// credential revalidation, backend admission, generation, or validation.
// Cancellation, capacity rejection, generation failure, operational
// validation failure, and success all leave the handle unusable. A nil,
// zero-value, foreign-Engine, discarded, claimed, or used handle returns an
// error matching ErrInvalidRequest; a nil Engine returns ErrInvalidConfig and
// does not claim the handle.
//
// The supplied context governs this execution attempt independently of the
// preparation context. It covers credential revalidation, admission,
// generation, validation, and any internal repair. Result timing begins after
// the claim and excludes preparation and consumer-held delay.
//
// RunPrepared can return ErrInvalidRequest, ErrAPIKeyEnvMissing,
// ErrCapacityExceeded, ErrLLMGenerate, or ErrValidation as applicable while
// preserving documented collaborator and context identities. An engine
// admission rejection is discoverable as [CapacityError] and still matches
// ErrCapacityExceeded. A completed content-validation rejection is returned
// in RunResult, not as an operational error. An operational error returns no
// partial RunResult.
func (e *Engine) RunPrepared(ctx context.Context, prepared *PreparedExecution) (*RunResult, error) {
if e == nil || e.runner == nil {
return nil, fmt.Errorf("%w: engine is nil", ErrInvalidConfig)
}
var internal *usecase.PreparedExecution
if prepared != nil {
internal = prepared.internal
}
result, err := e.runner.RunPrepared(ctx, internal)
if err != nil {
return nil, mapPublicError(err)
}
return fromDomainRunResult(result), nil
}

View File

@@ -281,6 +281,9 @@ func TestEngineExecutionSettingPrecedence(t *testing.T) {
intPointer := func(value int) *int { intPointer := func(value int) *int {
return &value return &value
} }
stringPointer := func(value string) *string {
return &value
}
defaultsProfile := executionProfileFixture{ defaultsProfile := executionProfileFixture{
id: "settings-defaults", id: "settings-defaults",
@@ -388,7 +391,7 @@ func TestEngineExecutionSettingPrecedence(t *testing.T) {
TopP: floatPointer(requestTarget.TopP), TopP: floatPointer(requestTarget.TopP),
TimeoutSeconds: intPointer(requestTarget.TimeoutSeconds), TimeoutSeconds: intPointer(requestTarget.TimeoutSeconds),
ServiceTier: requestTarget.ServiceTier, ServiceTier: requestTarget.ServiceTier,
ReasoningEffort: requestTarget.ReasoningEffort, ReasoningEffort: stringPointer(requestTarget.ReasoningEffort),
APIKeyEnv: requestTarget.APIKeyEnv, APIKeyEnv: requestTarget.APIKeyEnv,
ExtraParams: requestTarget.ExtraParams, ExtraParams: requestTarget.ExtraParams,
}, },
@@ -452,6 +455,43 @@ func TestEngineExecutionSettingPrecedence(t *testing.T) {
} }
}) })
} }
t.Run("blank request reasoning clears profile setting", func(t *testing.T) {
profile := executionProfileFixture{
id: "settings-reasoning-clear",
endpoint: "http://profile-reasoning.test/v1",
model: "profile-reasoning-model",
reasoningEffort: "medium",
}
profileDir := t.TempDir()
writeExecutionProfileFixture(t, profileDir, profile)
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: frameworkPromptDir,
ProfileDir: profileDir,
SchemaDir: frameworkSchemaDir,
})
if err != nil {
t.Fatalf("construct engine: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: frameworkMarkdownSummaryPromptID,
ProfileID: profile.id,
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Nia labels the archive."),
"glossary": promptkit.Inline("archive: A catalogued collection."),
},
Execution: &promptkit.ExecutionTargetOverride{
ReasoningEffort: stringPointer(" \t "),
},
})
if err != nil {
t.Fatalf("prepare engine: %v", err)
}
if prepared.EffectiveModelParams.ReasoningEffort != "" {
t.Fatalf("expected blank request reasoning to clear profile value, got %q", prepared.EffectiveModelParams.ReasoningEffort)
}
})
} }
func TestRunSucceedsWithInjectedLLMClient(t *testing.T) { func TestRunSucceedsWithInjectedLLMClient(t *testing.T) {
@@ -571,19 +611,26 @@ func TestEngineRunWithDirectorySourcesAndFileInputs(t *testing.T) {
func TestRunPassesPreparedRequestToInjectedLLMClient(t *testing.T) { func TestRunPassesPreparedRequestToInjectedLLMClient(t *testing.T) {
const directKey = "direct-injected-key" const directKey = "direct-injected-key"
const directSession = "assembled-session"
fake := &fakeLLMClient{ fake := &fakeLLMClient{
response: &promptkit.GenerateResponse{Content: `{"events":[{"title":"Archive labelled"}]}`}, response: &promptkit.GenerateResponse{Content: `{"events":[{"title":"Archive labelled"}]}`},
} }
engine := newContractEngineWithOptions(t, frameworkSchemaDir, promptkit.WithLLMClient(fake)) engine := newContractEngineWithOptions(t, frameworkSchemaDir, promptkit.WithLLMClient(fake))
_, err := engine.Run(context.Background(), promptkit.RunRequest{ runRequest := promptkit.RunRequest{
PromptID: frameworkStructuredEventsPromptID, PromptID: frameworkStructuredEventsPromptID,
SessionID: " " + directSession + " ",
APIKey: directKey, APIKey: directKey,
Inputs: map[string]promptkit.ArtifactRef{ Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."), "transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."), "glossary": promptkit.Inline("gate: A guarded passage."),
}, },
}) }
prepared, err := engine.Prepare(context.Background(), runRequest)
if err != nil {
t.Fatalf("expected prepare to succeed, got %v", err)
}
result, err := engine.Run(context.Background(), runRequest)
if err != nil { if err != nil {
t.Fatalf("expected run to succeed, got %v", err) t.Fatalf("expected run to succeed, got %v", err)
} }
@@ -594,6 +641,16 @@ func TestRunPassesPreparedRequestToInjectedLLMClient(t *testing.T) {
if len(req.Prompt.Messages) != 2 || !strings.Contains(req.Prompt.Messages[1].Content, "Rin opens the gate.") { if len(req.Prompt.Messages) != 2 || !strings.Contains(req.Prompt.Messages[1].Content, "Rin opens the gate.") {
t.Fatalf("expected rendered prompt in generate request, got %+v", req.Prompt) t.Fatalf("expected rendered prompt in generate request, got %+v", req.Prompt)
} }
if prepared.SessionID != directSession ||
req.Prompt.SessionID != directSession ||
result.SessionID != directSession {
t.Fatalf(
"direct session did not propagate consistently: prepared=%q generated=%q result=%q",
prepared.SessionID,
req.Prompt.SessionID,
result.SessionID,
)
}
if req.StructuredOutput == nil || req.StructuredOutput.Type != promptkit.StructuredOutputJSONSchema || req.StructuredOutput.JSONSchema == nil { if req.StructuredOutput == nil || req.StructuredOutput.Type != promptkit.StructuredOutputJSONSchema || req.StructuredOutput.JSONSchema == nil {
t.Fatalf("expected structured output handoff, got %+v", req.StructuredOutput) t.Fatalf("expected structured output handoff, got %+v", req.StructuredOutput)
} }
@@ -755,6 +812,92 @@ func TestRunUsesDirectAPIKeyWithDefaultLLMClient(t *testing.T) {
} }
} }
func TestRunUsesResolvedBackendWithBuiltInLLMClient(t *testing.T) {
const (
backendID = "local-test"
envName = "PROMPTKIT_BACKEND_TRANSPORT_KEY"
apiKey = "synthetic-backend-key"
)
t.Setenv(envName, apiKey)
var (
gotAuth string
gotBody map[string]any
)
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotAuth = r.Header.Get("Authorization")
if r.URL.Path != "/v1/chat/completions" {
t.Errorf("unexpected path: %s", r.URL.Path)
}
if err := json.NewDecoder(r.Body).Decode(&gotBody); err != nil {
t.Errorf("decode request body: %v", err)
}
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{
"choices": [{"message": {"role": "assistant", "content": "# Summary\n\nDone."}}],
"usage": {"prompt_tokens": 3, "completion_tokens": 4, "total_tokens": 7}
}`))
}))
defer server.Close()
engine, err := promptkit.NewEngine(promptkit.Config{
PromptDir: frameworkPromptDir,
SchemaDir: frameworkSchemaDir,
},
promptkit.WithBackend(promptkit.Backend{
ID: backendID,
Endpoint: server.URL + "/v1",
APIKeyEnv: envName,
ExtraParams: map[string]any{
"provider": "synthetic",
},
}),
promptkit.WithProfiles(promptkit.Profile{
ID: "backend-transport",
BackendID: backendID,
Model: "test-model",
}),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
result, err := engine.Run(context.Background(), promptkit.RunRequest{
PromptID: frameworkMarkdownSummaryPromptID,
ProfileID: "backend-transport",
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."),
},
})
if err != nil {
t.Fatalf("run with resolved backend: %v", err)
}
if gotAuth != "Bearer "+apiKey {
t.Fatalf("unexpected Authorization header: %q", gotAuth)
}
if gotBody["model"] != "test-model" || gotBody["provider"] != "synthetic" {
t.Fatalf("backend defaults did not reach provider payload: %#v", gotBody)
}
for _, field := range []string{"backend_id", "api_key_env"} {
if _, ok := gotBody[field]; ok {
t.Fatalf("internal metadata field %q was serialized to provider payload: %#v", field, gotBody)
}
}
bodyJSON, err := json.Marshal(gotBody)
if err != nil {
t.Fatalf("marshal captured provider payload: %v", err)
}
if strings.Contains(string(bodyJSON), apiKey) {
t.Fatalf("credential value was serialized to provider payload: %s", bodyJSON)
}
if result.SelectedBackendID != backendID ||
result.EffectiveModelParams.Endpoint != server.URL+"/v1" ||
result.EffectiveModelParams.APIKeyEnv != envName {
t.Fatalf("unexpected resolved backend metadata: %+v", result)
}
}
func TestPrepareDirectAPIKeyBypassesMissingEnvWithoutLeakingOrHashing(t *testing.T) { func TestPrepareDirectAPIKeyBypassesMissingEnvWithoutLeakingOrHashing(t *testing.T) {
const missingEnv = "PROMPTKIT_PUBLIC_PREPARE_MISSING" const missingEnv = "PROMPTKIT_PUBLIC_PREPARE_MISSING"
const firstKey = "first-direct-key" const firstKey = "first-direct-key"
@@ -1271,6 +1414,18 @@ func TestPrepareUsesBuiltInProfileWithoutProfileDir(t *testing.T) {
if prepared.SelectedProfileID != "mistral-small-3" { if prepared.SelectedProfileID != "mistral-small-3" {
t.Fatalf("unexpected selected profile: %q", prepared.SelectedProfileID) t.Fatalf("unexpected selected profile: %q", prepared.SelectedProfileID)
} }
if prepared.SelectedBackendID != promptkit.BackendOpenRouter {
t.Fatalf("unexpected selected backend: %q", prepared.SelectedBackendID)
}
if prepared.EffectiveModelParams.BackendID != promptkit.BackendOpenRouter {
t.Fatalf("unexpected effective backend: %q", prepared.EffectiveModelParams.BackendID)
}
if prepared.EffectiveModelParams.Endpoint != "https://openrouter.ai/api/v1" {
t.Fatalf("unexpected built-in endpoint: %q", prepared.EffectiveModelParams.Endpoint)
}
if prepared.EffectiveModelParams.APIKeyEnv != "OPENROUTER_API_KEY" {
t.Fatalf("unexpected built-in api key environment name: %q", prepared.EffectiveModelParams.APIKeyEnv)
}
if prepared.EffectiveModelParams.Model != "mistralai/mistral-small-3.2-24b-instruct" { if prepared.EffectiveModelParams.Model != "mistralai/mistral-small-3.2-24b-instruct" {
t.Fatalf("unexpected built-in model: %q", prepared.EffectiveModelParams.Model) t.Fatalf("unexpected built-in model: %q", prepared.EffectiveModelParams.Model)
} }
@@ -1641,6 +1796,7 @@ func TestOpenAICompatibleProfileRunsThroughNormalProfilePath(t *testing.T) {
fake := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "ok"}} fake := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "ok"}}
prof := promptkit.OpenAICompatibleProfile(promptkit.OpenAICompatibleProfileConfig{ prof := promptkit.OpenAICompatibleProfile(promptkit.OpenAICompatibleProfileConfig{
ID: "template-profile", ID: "template-profile",
BackendID: " openrouter ",
Endpoint: "http://template/v1", Endpoint: "http://template/v1",
Model: "template-model", Model: "template-model",
APIKeyRequired: true, APIKeyRequired: true,
@@ -1672,7 +1828,9 @@ func TestOpenAICompatibleProfileRunsThroughNormalProfilePath(t *testing.T) {
if len(fake.requests) != 1 { if len(fake.requests) != 1 {
t.Fatalf("expected one request, got %d", len(fake.requests)) t.Fatalf("expected one request, got %d", len(fake.requests))
} }
if fake.requests[0].Target.Model != "template-model" || fake.requests[0].APIKey != "template-key" { if fake.requests[0].Target.BackendID != promptkit.BackendOpenRouter ||
fake.requests[0].Target.Model != "template-model" ||
fake.requests[0].APIKey != "template-key" {
t.Fatalf("unexpected generated request: %+v", fake.requests[0]) t.Fatalf("unexpected generated request: %+v", fake.requests[0])
} }
if !reflect.DeepEqual(fake.requests[0].Target.ExtraParams, map[string]any{"provider": "template"}) { if !reflect.DeepEqual(fake.requests[0].Target.ExtraParams, map[string]any{"provider": "template"}) {
@@ -2256,6 +2414,11 @@ func TestExtraParamsTypedNestedValuesAreCopiedAcrossPublicBoundary(t *testing.T)
} }
func TestRunRejectsInvalidExtraParams(t *testing.T) { func TestRunRejectsInvalidExtraParams(t *testing.T) {
cyclicMap := map[string]any{}
cyclicMap["self"] = cyclicMap
cyclicSlice := []any{nil}
cyclicSlice[0] = cyclicSlice
tests := []struct { tests := []struct {
name string name string
extraParams map[string]any extraParams map[string]any
@@ -2267,42 +2430,10 @@ func TestRunRejectsInvalidExtraParams(t *testing.T) {
{name: "nan", extraParams: map[string]any{"bad": math.NaN()}}, {name: "nan", extraParams: map[string]any{"bad": math.NaN()}},
{name: "positive infinity", extraParams: map[string]any{"bad": math.Inf(1)}}, {name: "positive infinity", extraParams: map[string]any{"bad": math.Inf(1)}},
{name: "negative infinity", extraParams: map[string]any{"bad": math.Inf(-1)}}, {name: "negative infinity", extraParams: map[string]any{"bad": math.Inf(-1)}},
} {name: "cyclic map", extraParams: cyclicMap},
for _, tc := range tests { {name: "cyclic slice", extraParams: map[string]any{"cycle": cyclicSlice}},
t.Run(tc.name, func(t *testing.T) { {name: "malformed JSON number", extraParams: map[string]any{"value": json.Number("+1")}},
fake := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "ok"}} {name: "empty nested key", extraParams: map[string]any{"nested": map[string]any{"": true}}},
engine := newContractEngineWithOptions(t, frameworkSchemaDir, promptkit.WithLLMClient(fake))
_, err := engine.Run(context.Background(), promptkit.RunRequest{
PromptID: frameworkMarkdownSummaryPromptID,
Inputs: map[string]promptkit.ArtifactRef{
"transcript": promptkit.Inline("Rin opens the gate."),
"glossary": promptkit.Inline("gate: A guarded passage."),
},
Execution: &promptkit.ExecutionTargetOverride{ExtraParams: tc.extraParams},
})
if !errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("expected ErrInvalidRequest, got %v", err)
}
if len(fake.requests) != 0 {
t.Fatalf("expected invalid request to fail before LLM call, got %d requests", len(fake.requests))
}
})
}
}
func TestRunRejectsCyclicExtraParams(t *testing.T) {
cyclicMap := map[string]any{}
cyclicMap["self"] = cyclicMap
cyclicSlice := []any{nil}
cyclicSlice[0] = cyclicSlice
tests := []struct {
name string
extraParams map[string]any
}{
{name: "map", extraParams: cyclicMap},
{name: "slice", extraParams: map[string]any{"cycle": cyclicSlice}},
} }
for _, tc := range tests { for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) { t.Run(tc.name, func(t *testing.T) {

View File

@@ -3,7 +3,9 @@ package promptkit
import ( import (
"errors" "errors"
"fmt" "fmt"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/capacity"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
"gitea.maximumdirect.net/eric/promptkit/internal/promptdef" "gitea.maximumdirect.net/eric/promptkit/internal/promptdef"
"gitea.maximumdirect.net/eric/promptkit/internal/usecase" "gitea.maximumdirect.net/eric/promptkit/internal/usecase"
@@ -13,6 +15,11 @@ func mapPublicError(err error) error {
if err == nil { if err == nil {
return nil return nil
} }
var internalCapacityError *usecase.CapacityError
if errors.As(err, &internalCapacityError) && internalCapacityError != nil &&
strings.TrimSpace(internalCapacityError.BackendID) != "" {
return &CapacityError{BackendID: internalCapacityError.BackendID}
}
publicErr := publicErrorFor(err) publicErr := publicErrorFor(err)
if publicErr == nil { if publicErr == nil {
return err return err
@@ -38,6 +45,8 @@ func publicErrorFor(err error) error {
return ErrProfileLoad return ErrProfileLoad
case errors.Is(err, usecase.ErrAPIKeyEnvMissing): case errors.Is(err, usecase.ErrAPIKeyEnvMissing):
return errors.Join(ErrInvalidRequest, ErrAPIKeyEnvMissing) return errors.Join(ErrInvalidRequest, ErrAPIKeyEnvMissing)
case errors.Is(err, capacity.ErrCapacityExceeded):
return ErrCapacityExceeded
case errors.Is(err, usecase.ErrArtifactLoad): case errors.Is(err, usecase.ErrArtifactLoad):
return ErrArtifactLoad return ErrArtifactLoad
case errors.Is(err, usecase.ErrPromptRender): case errors.Is(err, usecase.ErrPromptRender):

50
errors_internal_test.go Normal file
View File

@@ -0,0 +1,50 @@
package promptkit
import (
"context"
"errors"
"fmt"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/usecase"
)
func TestMapPublicErrorPreservesGenerationCancellation(t *testing.T) {
internalErr := fmt.Errorf("%w: %w", usecase.ErrLLMGenerate, context.Canceled)
err := mapPublicError(internalErr)
if !errors.Is(err, ErrLLMGenerate) {
t.Fatalf("mapped error=%v, want ErrLLMGenerate", err)
}
if !errors.Is(err, context.Canceled) {
t.Fatalf("mapped error=%v, want context.Canceled", err)
}
}
func TestMapPublicErrorTranslatesCapacityError(t *testing.T) {
internalErr := &usecase.CapacityError{BackendID: "limited"}
err := mapPublicError(internalErr)
var publicErr *CapacityError
if !errors.As(err, &publicErr) || publicErr == nil {
t.Fatalf("mapped error=%v, want public CapacityError", err)
}
if publicErr.BackendID != "limited" {
t.Fatalf("mapped backend ID=%q, want limited", publicErr.BackendID)
}
if !errors.Is(err, ErrCapacityExceeded) {
t.Fatalf("mapped error=%v, want ErrCapacityExceeded", err)
}
if errors.Is(err, ErrInvalidRequest) || errors.Is(err, ErrLLMGenerate) {
t.Fatalf("mapped capacity error has an unrelated category: %v", err)
}
var leakedInternalErr *usecase.CapacityError
if errors.As(err, &leakedInternalErr) {
t.Fatalf("mapped error exposes internal CapacityError: %v", err)
}
internalErr.BackendID = "changed"
if publicErr.BackendID != "limited" {
t.Fatalf("mapped backend ID changed with source error: %q", publicErr.BackendID)
}
}

View File

@@ -0,0 +1,212 @@
// Package backend owns validated, immutable OpenAI-compatible backend
// definitions.
package backend
import (
"errors"
"fmt"
"net/url"
"regexp"
"sort"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
"gitea.maximumdirect.net/eric/promptkit/internal/llm"
)
const (
// OpenRouterID is the reserved ID of Promptkit's built-in OpenRouter
// backend.
OpenRouterID = "openrouter"
openRouterEndpoint = "https://openrouter.ai/api/v1"
openRouterAPIKeyEnv = "OPENROUTER_API_KEY"
openRouterConcurrencyLimit = 16
defaultQueueCapacity = 1024
)
// ErrBackendNotFound identifies a registry lookup for an unknown backend ID.
var ErrBackendNotFound = errors.New("backend not found")
var environmentVariableName = regexp.MustCompile(`^[A-Za-z_][A-Za-z0-9_]*$`)
// Registry is an immutable collection of validated backend definitions.
type Registry struct {
backends map[string]domain.Backend
}
// NewRegistry constructs a registry containing the built-in OpenRouter
// definition followed by the supplied additions. Every ID must be unique.
func NewRegistry(additions []domain.Backend) (*Registry, error) {
registry := &Registry{
backends: make(map[string]domain.Backend, len(additions)+1),
}
definitions := make([]domain.Backend, 0, len(additions)+1)
definitions = append(definitions, domain.Backend{
ID: OpenRouterID,
Endpoint: openRouterEndpoint,
APIKeyEnv: openRouterAPIKeyEnv,
ConcurrencyLimit: openRouterConcurrencyLimit,
})
definitions = append(definitions, additions...)
for _, definition := range definitions {
definition.ID = strings.TrimSpace(definition.ID)
if definition.ID == "" {
return nil, errors.New("backend ID must not be blank")
}
if _, exists := registry.backends[definition.ID]; exists {
return nil, fmt.Errorf("backend ID %q is already registered", definition.ID)
}
normalized, err := normalizeBackend(definition)
if err != nil {
return nil, err
}
registry.backends[normalized.ID] = normalized
}
return registry, nil
}
// GetBackend returns a defensive copy of the backend registered with id.
func (r *Registry) GetBackend(id string) (domain.Backend, error) {
if r == nil {
return domain.Backend{}, fmt.Errorf("%w: %q", ErrBackendNotFound, id)
}
definition, ok := r.backends[id]
if !ok {
return domain.Backend{}, fmt.Errorf("%w: %q", ErrBackendNotFound, id)
}
extraParams, err := jsonvalue.CopyMap(definition.ExtraParams)
if err != nil {
return domain.Backend{}, fmt.Errorf("copy backend %q: %w", id, err)
}
definition.ExtraParams = extraParams
return definition, nil
}
// CapacityPolicies returns a copy of the normalized policies for limited
// backends.
func (r *Registry) CapacityPolicies() map[string]domain.BackendCapacityPolicy {
policies := make(map[string]domain.BackendCapacityPolicy)
if r == nil {
return policies
}
for id, definition := range r.backends {
if definition.ConcurrencyLimit == 0 {
continue
}
policies[id] = domain.BackendCapacityPolicy{
ConcurrencyLimit: definition.ConcurrencyLimit,
QueueCapacity: definition.QueueCapacity,
}
}
return policies
}
func normalizeBackend(definition domain.Backend) (domain.Backend, error) {
definition.Endpoint = strings.TrimSpace(definition.Endpoint)
if err := validateEndpoint(definition.Endpoint); err != nil {
return domain.Backend{}, fmt.Errorf("backend %q endpoint: %w", definition.ID, err)
}
definition.APIKeyEnv = strings.TrimSpace(definition.APIKeyEnv)
if definition.APIKeyEnv != "" && !environmentVariableName.MatchString(definition.APIKeyEnv) {
return domain.Backend{}, fmt.Errorf(
"backend %q api key environment variable %q is invalid",
definition.ID,
definition.APIKeyEnv,
)
}
if definition.ConcurrencyLimit < 0 {
return domain.Backend{}, fmt.Errorf(
"backend %q concurrency limit must not be negative",
definition.ID,
)
}
if definition.QueueCapacity < 0 {
return domain.Backend{}, fmt.Errorf(
"backend %q queue capacity must not be negative",
definition.ID,
)
}
if definition.ConcurrencyLimit == 0 {
if definition.QueueCapacitySet {
return domain.Backend{}, fmt.Errorf(
"backend %q queue capacity requires a positive concurrency limit",
definition.ID,
)
}
definition.QueueCapacity = 0
} else {
if !definition.QueueCapacitySet {
definition.QueueCapacity = defaultQueueCapacity
definition.QueueCapacitySet = true
}
maxInt := int(^uint(0) >> 1)
if definition.QueueCapacity > maxInt-definition.ConcurrencyLimit {
return domain.Backend{}, fmt.Errorf(
"backend %q total capacity overflows int",
definition.ID,
)
}
}
keys := make([]string, 0, len(definition.ExtraParams))
for key := range definition.ExtraParams {
keys = append(keys, key)
}
sort.Strings(keys)
for _, key := range keys {
if key == "" {
return domain.Backend{}, fmt.Errorf("backend %q extra parameter key must not be empty", definition.ID)
}
if llm.IsReservedOpenAIChatRequestField(key) {
return domain.Backend{}, fmt.Errorf(
"backend %q extra parameter %q collides with a reserved request field",
definition.ID,
key,
)
}
}
extraParams, err := jsonvalue.CopyMap(definition.ExtraParams)
if err != nil {
return domain.Backend{}, fmt.Errorf("backend %q extra parameters: %w", definition.ID, err)
}
definition.ExtraParams = extraParams
return definition, nil
}
func validateEndpoint(endpoint string) error {
if endpoint == "" {
return errors.New("must not be blank")
}
if strings.Contains(endpoint, "#") {
return errors.New("must not contain a fragment")
}
parsed, err := url.Parse(endpoint)
if err != nil {
return fmt.Errorf("must be a valid URL: %w", err)
}
scheme := strings.ToLower(parsed.Scheme)
if scheme != "http" && scheme != "https" {
return errors.New("must use http or https")
}
if !parsed.IsAbs() || parsed.Hostname() == "" {
return errors.New("must be absolute and include a host")
}
if parsed.User != nil {
return errors.New("must not contain user information")
}
if parsed.RawQuery != "" || parsed.ForceQuery {
return errors.New("must not contain a query string")
}
return nil
}

View File

@@ -0,0 +1,368 @@
package backend_test
import (
"errors"
"strings"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/backend"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
const validEndpoint = "https://backend.example/v1"
func TestRegistryIncludesExactOpenRouterDefinition(t *testing.T) {
registry, err := backend.NewRegistry(nil)
if err != nil {
t.Fatalf("construct registry: %v", err)
}
definition, err := registry.GetBackend(backend.OpenRouterID)
if err != nil {
t.Fatalf("look up OpenRouter: %v", err)
}
if definition.ID != "openrouter" ||
definition.Endpoint != "https://openrouter.ai/api/v1" ||
definition.APIKeyEnv != "OPENROUTER_API_KEY" ||
definition.ConcurrencyLimit != 16 ||
definition.QueueCapacity != 1024 ||
!definition.QueueCapacitySet ||
definition.ExtraParams != nil {
t.Fatalf("unexpected OpenRouter definition: %#v", definition)
}
policies := registry.CapacityPolicies()
if len(policies) != 1 ||
policies["openrouter"] != (domain.BackendCapacityPolicy{
ConcurrencyLimit: 16,
QueueCapacity: 1024,
}) {
t.Fatalf("unexpected OpenRouter capacity policies: %#v", policies)
}
}
func TestRegistryNormalizesUniqueAdditionsAndIsolatesMutations(t *testing.T) {
nested := map[string]int{"limit": 2}
extraParams := map[string]any{
"count": int64(7),
"nested": nested,
}
registry, err := backend.NewRegistry([]domain.Backend{
{
ID: " custom ",
Endpoint: " https://custom.example/openai/v1 ",
APIKeyEnv: " CUSTOM_API_KEY ",
ExtraParams: extraParams,
ConcurrencyLimit: 3,
QueueCapacity: 2,
QueueCapacitySet: true,
},
{
ID: "Custom",
Endpoint: validEndpoint,
},
})
if err != nil {
t.Fatalf("construct registry: %v", err)
}
nested["limit"] = 99
extraParams["added"] = true
got, err := registry.GetBackend("custom")
if err != nil {
t.Fatalf("look up custom backend: %v", err)
}
if got.ID != "custom" ||
got.Endpoint != "https://custom.example/openai/v1" ||
got.APIKeyEnv != "CUSTOM_API_KEY" ||
got.ConcurrencyLimit != 3 ||
got.QueueCapacity != 2 ||
!got.QueueCapacitySet {
t.Fatalf("unexpected normalized definition: %#v", got)
}
if count, ok := got.ExtraParams["count"].(int64); !ok || count != 7 {
t.Fatalf("integer type or value changed: %#v", got.ExtraParams["count"])
}
gotNested, ok := got.ExtraParams["nested"].(map[string]int)
if !ok || gotNested["limit"] != 2 {
t.Fatalf("container type or value changed: %#v", got.ExtraParams["nested"])
}
if _, exists := got.ExtraParams["added"]; exists {
t.Fatalf("registry retained caller map: %#v", got.ExtraParams)
}
gotNested["limit"] = 100
got.ExtraParams["added"] = true
again, err := registry.GetBackend("custom")
if err != nil {
t.Fatalf("look up custom backend again: %v", err)
}
if again.ExtraParams["nested"].(map[string]int)["limit"] != 2 {
t.Fatalf("lookup exposed registry nested map: %#v", again.ExtraParams)
}
if _, exists := again.ExtraParams["added"]; exists {
t.Fatalf("lookup exposed registry map: %#v", again.ExtraParams)
}
if _, err := registry.GetBackend("Custom"); err != nil {
t.Fatalf("backend IDs should be case-sensitive: %v", err)
}
policies := registry.CapacityPolicies()
if len(policies) != 2 {
t.Fatalf("unexpected capacity policy count: %#v", policies)
}
policies["custom"] = domain.BackendCapacityPolicy{}
delete(policies, backend.OpenRouterID)
againPolicies := registry.CapacityPolicies()
if againPolicies["custom"] != (domain.BackendCapacityPolicy{
ConcurrencyLimit: 3,
QueueCapacity: 2,
}) {
t.Fatalf("capacity policy map mutated registry state: %#v", againPolicies)
}
if _, ok := againPolicies[backend.OpenRouterID]; !ok {
t.Fatalf("capacity policy deletion mutated registry state: %#v", againPolicies)
}
}
func TestNewRegistryNormalizesCapacityPolicy(t *testing.T) {
maxInt := int(^uint(0) >> 1)
tests := []struct {
name string
definition domain.Backend
want domain.BackendCapacityPolicy
wantSet bool
wantError bool
}{
{
name: "unlimited when omitted",
definition: domain.Backend{},
},
{
name: "default queue",
definition: domain.Backend{
ConcurrencyLimit: 2,
},
want: domain.BackendCapacityPolicy{
ConcurrencyLimit: 2,
QueueCapacity: 1024,
},
wantSet: true,
},
{
name: "explicit zero queue",
definition: domain.Backend{
ConcurrencyLimit: 2,
QueueCapacitySet: true,
},
want: domain.BackendCapacityPolicy{
ConcurrencyLimit: 2,
},
wantSet: true,
},
{
name: "negative concurrency limit",
definition: domain.Backend{
ConcurrencyLimit: -1,
},
wantError: true,
},
{
name: "negative queue capacity",
definition: domain.Backend{
ConcurrencyLimit: 1,
QueueCapacity: -1,
QueueCapacitySet: true,
},
wantError: true,
},
{
name: "queue without limit",
definition: domain.Backend{
QueueCapacitySet: true,
},
wantError: true,
},
{
name: "total overflow",
definition: domain.Backend{
ConcurrencyLimit: maxInt,
QueueCapacity: 1,
QueueCapacitySet: true,
},
wantError: true,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
tc.definition.ID = "custom"
tc.definition.Endpoint = validEndpoint
registry, err := backend.NewRegistry([]domain.Backend{tc.definition})
if tc.wantError {
if err == nil {
t.Fatal("expected invalid capacity policy error")
}
return
}
if err != nil {
t.Fatalf("construct registry: %v", err)
}
definition, err := registry.GetBackend("custom")
if err != nil {
t.Fatalf("look up custom backend: %v", err)
}
if definition.ConcurrencyLimit != tc.want.ConcurrencyLimit ||
definition.QueueCapacity != tc.want.QueueCapacity ||
definition.QueueCapacitySet != tc.wantSet {
t.Fatalf("normalized capacity=(%d, %d, %t), want (%d, %d, %t)",
definition.ConcurrencyLimit,
definition.QueueCapacity,
definition.QueueCapacitySet,
tc.want.ConcurrencyLimit,
tc.want.QueueCapacity,
tc.wantSet,
)
}
policies := registry.CapacityPolicies()
got, ok := policies["custom"]
if ok != tc.wantSet || got != tc.want {
t.Fatalf("capacity policy=(%#v, %t), want (%#v, %t)", got, ok, tc.want, tc.wantSet)
}
})
}
}
func TestNewRegistryRejectsDuplicateIDs(t *testing.T) {
tests := []struct {
name string
additions []domain.Backend
wantID string
}{
{
name: "built-in collision after normalization",
additions: []domain.Backend{{
ID: " openrouter ",
}},
wantID: "openrouter",
},
{
name: "consumer collision after normalization",
additions: []domain.Backend{
{ID: "custom", Endpoint: validEndpoint},
{ID: " custom ", Endpoint: "https://other.example/v1"},
},
wantID: "custom",
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
_, err := backend.NewRegistry(tc.additions)
if err == nil {
t.Fatal("expected duplicate ID error")
}
if !strings.Contains(err.Error(), tc.wantID) {
t.Fatalf("expected error to identify %q, got %v", tc.wantID, err)
}
})
}
}
func TestNewRegistryValidatesIDs(t *testing.T) {
for _, id := range []string{"", " \t\n "} {
t.Run(id, func(t *testing.T) {
_, err := backend.NewRegistry([]domain.Backend{{
ID: id,
Endpoint: validEndpoint,
}})
if err == nil {
t.Fatal("expected blank ID error")
}
})
}
}
func TestNewRegistryValidatesEndpoints(t *testing.T) {
tests := []struct {
name string
endpoint string
}{
{name: "blank", endpoint: ""},
{name: "relative", endpoint: "/v1"},
{name: "missing host", endpoint: "https:///v1"},
{name: "unsupported scheme", endpoint: "ftp://backend.example/v1"},
{name: "user information", endpoint: "https://user@backend.example/v1"},
{name: "query", endpoint: "https://backend.example/v1?mode=chat"},
{name: "empty query", endpoint: "https://backend.example/v1?"},
{name: "fragment", endpoint: "https://backend.example/v1#chat"},
{name: "empty fragment", endpoint: "https://backend.example/v1#"},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
_, err := backend.NewRegistry([]domain.Backend{{
ID: "custom",
Endpoint: tc.endpoint,
}})
if err == nil {
t.Fatal("expected invalid endpoint error")
}
})
}
}
func TestNewRegistryValidatesEnvironmentVariableNames(t *testing.T) {
for _, name := range []string{"1API_KEY", "API-KEY", "API KEY", "ÅPI_KEY"} {
t.Run(name, func(t *testing.T) {
_, err := backend.NewRegistry([]domain.Backend{{
ID: "custom",
Endpoint: validEndpoint,
APIKeyEnv: name,
}})
if err == nil {
t.Fatal("expected invalid environment-variable name error")
}
})
}
}
func TestNewRegistryRejectsInvalidAndReservedExtraParameters(t *testing.T) {
tests := []struct {
name string
extraParams map[string]any
}{
{name: "unsupported value", extraParams: map[string]any{"value": make(chan int)}},
{name: "reserved key", extraParams: map[string]any{"model": "override"}},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
_, err := backend.NewRegistry([]domain.Backend{{
ID: "custom",
Endpoint: validEndpoint,
ExtraParams: tc.extraParams,
}})
if err == nil {
t.Fatal("expected invalid extra parameters error")
}
})
}
}
func TestRegistryLookupReportsNotFound(t *testing.T) {
registry, err := backend.NewRegistry(nil)
if err != nil {
t.Fatalf("construct registry: %v", err)
}
_, err = registry.GetBackend("missing")
if !errors.Is(err, backend.ErrBackendNotFound) {
t.Fatalf("expected ErrBackendNotFound, got %v", err)
}
if !strings.Contains(err.Error(), "missing") {
t.Fatalf("expected error to identify backend, got %v", err)
}
}

View File

@@ -0,0 +1,40 @@
package capacity
import (
"context"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/llm"
)
type client struct {
manager *Manager
next llm.Client
}
// NewClient wraps next with configured active-generation limits. A nil manager
// leaves next unchanged.
func NewClient(manager *Manager, next llm.Client) llm.Client {
if manager == nil {
return next
}
return &client{
manager: manager,
next: next,
}
}
func (c *client) Generate(
ctx context.Context,
req domain.GenerateRequest,
) (*domain.GenerateResponse, error) {
pool := c.manager.getPool(req.Target.BackendID)
if pool == nil {
return c.next.Generate(ctx, req)
}
if err := pool.acquire(ctx); err != nil {
return nil, err
}
defer pool.releaseActive()
return c.next.Generate(ctx, req)
}

View File

@@ -0,0 +1,517 @@
package capacity
import (
"context"
"errors"
"reflect"
"runtime"
"sync"
"sync/atomic"
"testing"
"time"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/llm"
)
type generateResult struct {
response *domain.GenerateResponse
err error
}
type clientFunc func(
context.Context,
domain.GenerateRequest,
) (*domain.GenerateResponse, error)
func (f clientFunc) Generate(
ctx context.Context,
req domain.GenerateRequest,
) (*domain.GenerateResponse, error) {
return f(ctx, req)
}
type blockingClient struct {
mu sync.Mutex
active int
peak int
calls map[string]int
started chan string
releases map[string]chan struct{}
}
func newBlockingClient(releases map[string]chan struct{}) *blockingClient {
return &blockingClient{
calls: make(map[string]int),
started: make(chan string, 64),
releases: releases,
}
}
func (c *blockingClient) Generate(
ctx context.Context,
req domain.GenerateRequest,
) (*domain.GenerateResponse, error) {
id := req.Prompt.SessionID
c.mu.Lock()
c.active++
if c.active > c.peak {
c.peak = c.active
}
c.calls[id]++
c.mu.Unlock()
defer func() {
c.mu.Lock()
c.active--
c.mu.Unlock()
}()
c.started <- id
if release := c.releases[id]; release != nil {
select {
case <-release:
case <-ctx.Done():
return nil, ctx.Err()
}
}
return &domain.GenerateResponse{Content: id}, nil
}
func (c *blockingClient) callCount(id string) int {
c.mu.Lock()
defer c.mu.Unlock()
return c.calls[id]
}
func (c *blockingClient) peakConcurrency() int {
c.mu.Lock()
defer c.mu.Unlock()
return c.peak
}
func generateAsync(
client llm.Client,
ctx context.Context,
backendID string,
id string,
) <-chan generateResult {
result := make(chan generateResult, 1)
go func() {
response, err := client.Generate(ctx, domain.GenerateRequest{
Prompt: domain.RenderedPrompt{SessionID: id},
Target: domain.ExecutionTarget{BackendID: backendID},
})
result <- generateResult{response: response, err: err}
}()
return result
}
func waitForWaiterCount(t *testing.T, manager *Manager, backendID string, want int) {
t.Helper()
pool := manager.pools[backendID]
deadline := time.Now().Add(2 * time.Second)
for {
pool.mu.Lock()
got := pool.waiters.Len()
pool.mu.Unlock()
if got == want {
return
}
if time.Now().After(deadline) {
t.Fatalf("waiter count=%d, want %d", got, want)
}
runtime.Gosched()
}
}
func receiveStarted(t *testing.T, started <-chan string) string {
t.Helper()
select {
case id := <-started:
return id
case <-time.After(2 * time.Second):
t.Fatal("timed out waiting for wrapped client invocation")
return ""
}
}
func receiveResult(t *testing.T, result <-chan generateResult) generateResult {
t.Helper()
select {
case got := <-result:
return got
case <-time.After(2 * time.Second):
t.Fatal("timed out waiting for generation result")
return generateResult{}
}
}
func newTestManager(t *testing.T, policies map[string]domain.BackendCapacityPolicy) *Manager {
t.Helper()
manager, err := NewManager(policies)
if err != nil {
t.Fatalf("construct manager: %v", err)
}
return manager
}
func TestClientLimitsPeakConcurrencyAndServesWaitersFIFO(t *testing.T) {
manager := newTestManager(t, map[string]domain.BackendCapacityPolicy{
"limited": {ConcurrencyLimit: 1},
})
firstRelease := make(chan struct{})
secondRelease := make(chan struct{})
thirdRelease := make(chan struct{})
next := newBlockingClient(map[string]chan struct{}{
"first": firstRelease,
"second": secondRelease,
"third": thirdRelease,
})
client := NewClient(manager, next)
first := generateAsync(client, context.Background(), "limited", "first")
if got := receiveStarted(t, next.started); got != "first" {
t.Fatalf("first invocation=%q, want first", got)
}
second := generateAsync(client, context.Background(), "limited", "second")
waitForWaiterCount(t, manager, "limited", 1)
third := generateAsync(client, context.Background(), "limited", "third")
waitForWaiterCount(t, manager, "limited", 2)
close(firstRelease)
if got := receiveResult(t, first); got.err != nil {
t.Fatalf("first generation: %v", got.err)
}
if got := receiveStarted(t, next.started); got != "second" {
t.Fatalf("second invocation=%q, want second", got)
}
close(secondRelease)
if got := receiveResult(t, second); got.err != nil {
t.Fatalf("second generation: %v", got.err)
}
if got := receiveStarted(t, next.started); got != "third" {
t.Fatalf("third invocation=%q, want third", got)
}
close(thirdRelease)
if got := receiveResult(t, third); got.err != nil {
t.Fatalf("third generation: %v", got.err)
}
if peak := next.peakConcurrency(); peak != 1 {
t.Fatalf("peak concurrency=%d, want 1", peak)
}
}
func TestClientPeakConcurrencyDoesNotExceedConfiguredLimit(t *testing.T) {
const limit = 2
manager := newTestManager(t, map[string]domain.BackendCapacityPolicy{
"limited": {ConcurrencyLimit: limit},
})
gate := make(chan struct{})
releases := make(map[string]chan struct{})
for i := range 5 {
releases[string(rune('a'+i))] = gate
}
next := newBlockingClient(releases)
client := NewClient(manager, next)
results := make([]<-chan generateResult, 0, len(releases))
for id := range releases {
results = append(results, generateAsync(client, context.Background(), "limited", id))
}
for range limit {
receiveStarted(t, next.started)
}
waitForWaiterCount(t, manager, "limited", len(releases)-limit)
close(gate)
for _, result := range results {
if got := receiveResult(t, result); got.err != nil {
t.Fatalf("generation: %v", got.err)
}
}
if peak := next.peakConcurrency(); peak != limit {
t.Fatalf("peak concurrency=%d, want %d", peak, limit)
}
}
func TestClientRemovesCanceledWaiters(t *testing.T) {
tests := []struct {
name string
cancelID string
wantOrder []string
}{
{name: "first waiter", cancelID: "one", wantOrder: []string{"two", "three"}},
{name: "middle waiter", cancelID: "two", wantOrder: []string{"one", "three"}},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
manager := newTestManager(t, map[string]domain.BackendCapacityPolicy{
"limited": {ConcurrencyLimit: 1},
})
holderRelease := make(chan struct{})
releases := map[string]chan struct{}{
"holder": holderRelease,
"one": make(chan struct{}),
"two": make(chan struct{}),
"three": make(chan struct{}),
}
next := newBlockingClient(releases)
client := NewClient(manager, next)
holder := generateAsync(client, context.Background(), "limited", "holder")
if got := receiveStarted(t, next.started); got != "holder" {
t.Fatalf("initial invocation=%q, want holder", got)
}
contexts := make(map[string]context.Context)
cancels := make(map[string]context.CancelFunc)
results := make(map[string]<-chan generateResult)
for _, id := range []string{"one", "two", "three"} {
contexts[id], cancels[id] = context.WithCancel(context.Background())
results[id] = generateAsync(client, contexts[id], "limited", id)
waitForWaiterCount(t, manager, "limited", len(results))
}
cancels[tc.cancelID]()
if got := receiveResult(t, results[tc.cancelID]); !errors.Is(got.err, context.Canceled) {
t.Fatalf("canceled waiter error=%v, want context.Canceled", got.err)
}
waitForWaiterCount(t, manager, "limited", 2)
close(holderRelease)
if got := receiveResult(t, holder); got.err != nil {
t.Fatalf("holder generation: %v", got.err)
}
for _, id := range tc.wantOrder {
if got := receiveStarted(t, next.started); got != id {
t.Fatalf("next invocation=%q, want %q", got, id)
}
close(releases[id])
if got := receiveResult(t, results[id]); got.err != nil {
t.Fatalf("%s generation: %v", id, got.err)
}
}
if calls := next.callCount(tc.cancelID); calls != 0 {
t.Fatalf("canceled waiter invoked wrapped client %d times", calls)
}
for _, cancel := range cancels {
cancel()
}
})
}
}
func TestClientGrantCancellationRaceDoesNotLeakPermit(t *testing.T) {
const iterations = 200
for i := range iterations {
manager := newTestManager(t, map[string]domain.BackendCapacityPolicy{
"limited": {ConcurrencyLimit: 1},
})
holderRelease := make(chan struct{})
var waiterCalls atomic.Int64
next := clientFunc(func(
_ context.Context,
req domain.GenerateRequest,
) (*domain.GenerateResponse, error) {
if req.Prompt.SessionID == "holder" {
<-holderRelease
} else if req.Prompt.SessionID == "waiter" {
waiterCalls.Add(1)
}
return &domain.GenerateResponse{Content: req.Prompt.SessionID}, nil
})
client := NewClient(manager, next)
holder := generateAsync(client, context.Background(), "limited", "holder")
waitForActiveCount(t, manager, "limited", 1)
ctx, cancel := context.WithCancel(context.Background())
waiterResult := generateAsync(client, ctx, "limited", "waiter")
waitForWaiterCount(t, manager, "limited", 1)
start := make(chan struct{})
var race sync.WaitGroup
race.Add(2)
go func() {
defer race.Done()
<-start
cancel()
}()
go func() {
defer race.Done()
<-start
close(holderRelease)
}()
close(start)
race.Wait()
if got := receiveResult(t, holder); got.err != nil {
t.Fatalf("iteration %d holder generation: %v", i, got.err)
}
got := receiveResult(t, waiterResult)
switch calls := waiterCalls.Load(); {
case calls == 0 && errors.Is(got.err, context.Canceled):
case calls == 1 && got.err == nil:
default:
t.Fatalf("iteration %d waiter calls=%d error=%v", i, calls, got.err)
}
probe := generateAsync(client, context.Background(), "limited", "probe")
if got := receiveResult(t, probe); got.err != nil {
t.Fatalf("iteration %d probe generation: %v", i, got.err)
}
waitForActiveCount(t, manager, "limited", 0)
waitForWaiterCount(t, manager, "limited", 0)
}
}
func waitForActiveCount(t *testing.T, manager *Manager, backendID string, want int) {
t.Helper()
pool := manager.pools[backendID]
deadline := time.Now().Add(2 * time.Second)
for {
pool.mu.Lock()
got := pool.active
pool.mu.Unlock()
if got == want {
return
}
if time.Now().After(deadline) {
t.Fatalf("active count=%d, want %d", got, want)
}
runtime.Gosched()
}
}
func TestClientUsesIndependentPoolsAndUnlimitedFastPaths(t *testing.T) {
manager := newTestManager(t, map[string]domain.BackendCapacityPolicy{
"alpha": {ConcurrencyLimit: 1},
"beta": {ConcurrencyLimit: 1},
})
alphaRelease := make(chan struct{})
betaRelease := make(chan struct{})
next := newBlockingClient(map[string]chan struct{}{
"alpha": alphaRelease,
"beta": betaRelease,
})
client := NewClient(manager, next)
alpha := generateAsync(client, context.Background(), "alpha", "alpha")
beta := generateAsync(client, context.Background(), "beta", "beta")
started := map[string]bool{
receiveStarted(t, next.started): true,
receiveStarted(t, next.started): true,
}
if !started["alpha"] || !started["beta"] {
t.Fatalf("independent pools did not both start: %#v", started)
}
close(alphaRelease)
close(betaRelease)
if got := receiveResult(t, alpha); got.err != nil {
t.Fatalf("alpha generation: %v", got.err)
}
if got := receiveResult(t, beta); got.err != nil {
t.Fatalf("beta generation: %v", got.err)
}
for _, backendID := range []string{"", "unknown"} {
response, err := client.Generate(context.Background(), domain.GenerateRequest{
Prompt: domain.RenderedPrompt{SessionID: backendID},
Target: domain.ExecutionTarget{BackendID: backendID},
})
if err != nil || response == nil {
t.Fatalf("unlimited backend %q response=(%#v, %v)", backendID, response, err)
}
}
if got := NewClient(nil, next); got != next {
t.Fatal("nil manager did not return the wrapped client unchanged")
}
}
func TestClientPreservesRequestsResponsesAndErrors(t *testing.T) {
manager := newTestManager(t, map[string]domain.BackendCapacityPolicy{
"limited": {ConcurrencyLimit: 1},
})
request := domain.GenerateRequest{
Prompt: domain.RenderedPrompt{
SessionID: "session",
Messages: []domain.RenderedMessage{
{Role: "user", Content: "content"},
},
},
Target: domain.ExecutionTarget{
BackendID: "limited",
Model: "model",
ExtraParams: map[string]any{"key": "value"},
},
}
response := &domain.GenerateResponse{
Content: "output",
Usage: domain.TokenUsage{TotalTokens: 7},
}
collaboratorErr := errors.New("collaborator failure")
tests := []struct {
name string
response *domain.GenerateResponse
err error
}{
{name: "successful response", response: response},
{name: "nil response"},
{name: "collaborator error", response: response, err: collaboratorErr},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
var captured domain.GenerateRequest
next := clientFunc(func(
_ context.Context,
req domain.GenerateRequest,
) (*domain.GenerateResponse, error) {
captured = req
return tc.response, tc.err
})
gotResponse, gotErr := NewClient(manager, next).Generate(context.Background(), request)
if !reflect.DeepEqual(captured, request) {
t.Fatalf("request changed: %#v", captured)
}
if gotResponse != tc.response || gotErr != tc.err {
t.Fatalf("response=(%p, %v), want (%p, %v)",
gotResponse, gotErr, tc.response, tc.err)
}
})
}
}
func TestClientReleasesPermitDuringPanicUnwinding(t *testing.T) {
manager := newTestManager(t, map[string]domain.BackendCapacityPolicy{
"limited": {ConcurrencyLimit: 1},
})
var calls atomic.Int64
next := clientFunc(func(
_ context.Context,
_ domain.GenerateRequest,
) (*domain.GenerateResponse, error) {
if calls.Add(1) == 1 {
panic("test panic")
}
return &domain.GenerateResponse{Content: "recovered"}, nil
})
client := NewClient(manager, next)
request := domain.GenerateRequest{
Target: domain.ExecutionTarget{BackendID: "limited"},
}
func() {
defer func() {
if recover() == nil {
t.Fatal("expected wrapped client panic")
}
}()
_, _ = client.Generate(context.Background(), request)
}()
response, err := client.Generate(context.Background(), request)
if err != nil || response == nil || response.Content != "recovered" {
t.Fatalf("generation after panic=(%#v, %v)", response, err)
}
}

View File

@@ -0,0 +1,160 @@
// Package capacity coordinates engine-local run admission and model-generation
// concurrency for configured backends.
package capacity
import (
"container/list"
"context"
"errors"
"fmt"
"strings"
"sync"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
// ErrCapacityExceeded identifies an admission rejected because a backend's
// configured run capacity is full.
var ErrCapacityExceeded = errors.New("backend capacity exceeded")
// Manager owns independent backend capacity pools with immutable limits.
type Manager struct {
pools map[string]*pool
}
type pool struct {
mu sync.Mutex
concurrencyLimit int
totalCapacity int
admitted int
active int
waiters list.List
}
type waiter struct {
ready chan struct{}
element *list.Element
granted bool
}
// NewManager constructs independent pools from normalized backend policies.
func NewManager(policies map[string]domain.BackendCapacityPolicy) (*Manager, error) {
manager := &Manager{
pools: make(map[string]*pool, len(policies)),
}
maxInt := int(^uint(0) >> 1)
for id, policy := range policies {
if strings.TrimSpace(id) == "" {
return nil, errors.New("backend capacity policy ID must not be blank")
}
if policy.ConcurrencyLimit <= 0 {
return nil, fmt.Errorf(
"backend %q concurrency limit must be positive",
id,
)
}
if policy.QueueCapacity < 0 {
return nil, fmt.Errorf(
"backend %q queue capacity must not be negative",
id,
)
}
if policy.QueueCapacity > maxInt-policy.ConcurrencyLimit {
return nil, fmt.Errorf("backend %q total capacity overflows int", id)
}
manager.pools[id] = &pool{
concurrencyLimit: policy.ConcurrencyLimit,
totalCapacity: policy.ConcurrencyLimit + policy.QueueCapacity,
}
}
return manager, nil
}
// Admit immediately reserves one configured backend run slot. Backends without
// a configured pool are unlimited.
func (m *Manager) Admit(ctx context.Context, backendID string) (func(), error) {
pool := m.getPool(backendID)
if pool == nil {
return releaseNothing, nil
}
pool.mu.Lock()
defer pool.mu.Unlock()
if err := ctx.Err(); err != nil {
return nil, err
}
if pool.admitted >= pool.totalCapacity {
return nil, ErrCapacityExceeded
}
pool.admitted++
var once sync.Once
return func() {
once.Do(func() {
pool.mu.Lock()
pool.admitted--
pool.mu.Unlock()
})
}, nil
}
func releaseNothing() {}
func (m *Manager) getPool(backendID string) *pool {
if m == nil || backendID == "" {
return nil
}
return m.pools[backendID]
}
func (p *pool) acquire(ctx context.Context) error {
p.mu.Lock()
if err := ctx.Err(); err != nil {
p.mu.Unlock()
return err
}
if p.active < p.concurrencyLimit && p.waiters.Len() == 0 {
p.active++
p.mu.Unlock()
return nil
}
waiter := &waiter{ready: make(chan struct{})}
waiter.element = p.waiters.PushBack(waiter)
p.mu.Unlock()
select {
case <-waiter.ready:
return nil
case <-ctx.Done():
p.mu.Lock()
if !waiter.granted {
p.waiters.Remove(waiter.element)
waiter.element = nil
p.mu.Unlock()
return ctx.Err()
}
p.mu.Unlock()
return nil
}
}
func (p *pool) releaseActive() {
var ready chan struct{}
p.mu.Lock()
if element := p.waiters.Front(); element != nil {
waiter := element.Value.(*waiter)
p.waiters.Remove(element)
waiter.element = nil
waiter.granted = true
ready = waiter.ready
} else {
p.active--
}
p.mu.Unlock()
if ready != nil {
close(ready)
}
}

View File

@@ -0,0 +1,163 @@
package capacity
import (
"context"
"errors"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
func TestNewManagerRejectsInvalidPolicies(t *testing.T) {
maxInt := int(^uint(0) >> 1)
tests := []struct {
name string
id string
policy domain.BackendCapacityPolicy
}{
{
name: "blank ID",
id: " \t ",
policy: domain.BackendCapacityPolicy{ConcurrencyLimit: 1},
},
{
name: "zero concurrency",
id: "backend",
policy: domain.BackendCapacityPolicy{},
},
{
name: "negative concurrency",
id: "backend",
policy: domain.BackendCapacityPolicy{ConcurrencyLimit: -1},
},
{
name: "negative queue",
id: "backend",
policy: domain.BackendCapacityPolicy{
ConcurrencyLimit: 1,
QueueCapacity: -1,
},
},
{
name: "total overflow",
id: "backend",
policy: domain.BackendCapacityPolicy{
ConcurrencyLimit: maxInt,
QueueCapacity: 1,
},
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
_, err := NewManager(map[string]domain.BackendCapacityPolicy{
tc.id: tc.policy,
})
if err == nil {
t.Fatal("expected invalid policy error")
}
})
}
}
func TestManagerAdmissionIsBoundedAndReleaseIsIdempotent(t *testing.T) {
policies := map[string]domain.BackendCapacityPolicy{
"limited": {
ConcurrencyLimit: 2,
QueueCapacity: 1,
},
"independent": {
ConcurrencyLimit: 1,
},
}
manager, err := NewManager(policies)
if err != nil {
t.Fatalf("construct manager: %v", err)
}
policies["limited"] = domain.BackendCapacityPolicy{
ConcurrencyLimit: 100,
QueueCapacity: 100,
}
releases := make([]func(), 0, 3)
for range 3 {
release, err := manager.Admit(context.Background(), "limited")
if err != nil {
t.Fatalf("admit within configured capacity: %v", err)
}
releases = append(releases, release)
}
if release, err := manager.Admit(context.Background(), "limited"); release != nil ||
!errors.Is(err, ErrCapacityExceeded) {
t.Fatalf("admission beyond capacity=(release=%t, err=%v), want ErrCapacityExceeded",
release != nil, err)
}
independentRelease, err := manager.Admit(context.Background(), "independent")
if err != nil {
t.Fatalf("admit independent backend while first is full: %v", err)
}
independentRelease()
releases[0]()
releases[0]()
replacement, err := manager.Admit(context.Background(), "limited")
if err != nil {
t.Fatalf("admit after release: %v", err)
}
replacement()
releases[1]()
releases[2]()
pool := manager.pools["limited"]
pool.mu.Lock()
admitted := pool.admitted
pool.mu.Unlock()
if admitted != 0 {
t.Fatalf("admitted runs after releases=%d, want 0", admitted)
}
}
func TestManagerAdmissionHonorsContextAndUnlimitedBackends(t *testing.T) {
manager, err := NewManager(map[string]domain.BackendCapacityPolicy{
"limited": {ConcurrencyLimit: 1},
})
if err != nil {
t.Fatalf("construct manager: %v", err)
}
release, err := manager.Admit(context.Background(), "limited")
if err != nil {
t.Fatalf("fill limited pool: %v", err)
}
defer release()
ctx, cancel := context.WithCancel(context.Background())
cancel()
if release, err := manager.Admit(ctx, "limited"); release != nil ||
!errors.Is(err, context.Canceled) {
t.Fatalf("canceled limited admission=(release=%t, err=%v), want context cancellation",
release != nil, err)
}
var nilManager *Manager
for _, tc := range []struct {
name string
manager *Manager
backendID string
}{
{name: "nil manager", manager: nilManager, backendID: "limited"},
{name: "blank ID", manager: manager},
{name: "unknown ID", manager: manager, backendID: "unknown"},
} {
t.Run(tc.name, func(t *testing.T) {
release, err := tc.manager.Admit(ctx, tc.backendID)
if err != nil {
t.Fatalf("unlimited admission: %v", err)
}
if release == nil {
t.Fatal("unlimited admission returned nil release")
}
release()
release()
})
}
}

View File

@@ -63,6 +63,7 @@ type RunRequest struct {
PromptID string PromptID string
PromptVersion string PromptVersion string
ProfileID string ProfileID string
SessionID string
APIKey string `json:"-" yaml:"-"` APIKey string `json:"-" yaml:"-"`
Inputs map[string]ArtifactRef Inputs map[string]ArtifactRef
Vars map[string]string Vars map[string]string
@@ -79,8 +80,10 @@ type RunResult struct {
PromptID string PromptID string
PromptVersion string PromptVersion string
PromptHash string PromptHash string
SessionID string
RenderedPromptHash string RenderedPromptHash string
SelectedProfileID string SelectedProfileID string
SelectedBackendID string
ModelName string ModelName string
Endpoint string Endpoint string
EffectiveModelParams ExecutionTarget EffectiveModelParams ExecutionTarget
@@ -98,6 +101,7 @@ type PreparedRun struct {
PromptVersion string `json:"prompt_version,omitempty"` PromptVersion string `json:"prompt_version,omitempty"`
PromptHash string `json:"prompt_hash,omitempty"` PromptHash string `json:"prompt_hash,omitempty"`
SelectedProfileID string `json:"selected_profile_id"` SelectedProfileID string `json:"selected_profile_id"`
SelectedBackendID string `json:"selected_backend_id,omitempty"`
EffectiveModelParams ExecutionTarget `json:"effective_model_params"` EffectiveModelParams ExecutionTarget `json:"effective_model_params"`
TargetPresence ExecutionTargetPresence `json:"-"` TargetPresence ExecutionTargetPresence `json:"-"`
OutputContract OutputContract `json:"output_contract"` OutputContract OutputContract `json:"output_contract"`
@@ -141,6 +145,16 @@ type PromptDefinition struct {
Validation OutputContract `yaml:"validation"` Validation OutputContract `yaml:"validation"`
} }
// PromptInspection is the resolved result of exact prompt inspection.
type PromptInspection struct {
PromptID string
PromptVersion string
PromptHash string
DefaultProfileID string
Inputs []PromptInput
OutputContract OutputContract
}
// PromptInput describes one named input expected by a prompt definition. // PromptInput describes one named input expected by a prompt definition.
type PromptInput struct { type PromptInput struct {
Name string `yaml:"name"` Name string `yaml:"name"`
@@ -157,9 +171,28 @@ type PromptMessageTemplate struct {
CacheControl *CacheControl `yaml:"cache_control,omitempty" json:"cache_control,omitempty"` CacheControl *CacheControl `yaml:"cache_control,omitempty" json:"cache_control,omitempty"`
} }
// Backend describes reusable OpenAI-compatible connection defaults.
type Backend struct {
ID string
Endpoint string
APIKeyEnv string
ExtraParams map[string]any
ConcurrencyLimit int
QueueCapacity int
QueueCapacitySet bool
}
// BackendCapacityPolicy describes normalized run and generation capacity for
// one limited backend.
type BackendCapacityPolicy struct {
ConcurrencyLimit int
QueueCapacity int
}
// ExecutionProfile describes how and where to execute a model. // ExecutionProfile describes how and where to execute a model.
type ExecutionProfile struct { type ExecutionProfile struct {
ID string `yaml:"id"` ID string `yaml:"id"`
BackendID string `yaml:"backend"`
Endpoint string `yaml:"endpoint"` Endpoint string `yaml:"endpoint"`
Model string `yaml:"model"` Model string `yaml:"model"`
Temperature float64 `yaml:"temperature"` Temperature float64 `yaml:"temperature"`
@@ -182,7 +215,7 @@ type ExecutionTargetOverride struct {
TopP *float64 `json:"top_p,omitempty"` TopP *float64 `json:"top_p,omitempty"`
TimeoutSeconds *int `json:"timeout_seconds,omitempty"` TimeoutSeconds *int `json:"timeout_seconds,omitempty"`
ServiceTier string `json:"service_tier,omitempty"` ServiceTier string `json:"service_tier,omitempty"`
ReasoningEffort string `json:"reasoning_effort,omitempty"` ReasoningEffort *string `json:"reasoning_effort,omitempty"`
APIKeyEnv string `json:"api_key_env,omitempty"` APIKeyEnv string `json:"api_key_env,omitempty"`
ExtraParams map[string]any `json:"extra_params,omitempty"` ExtraParams map[string]any `json:"extra_params,omitempty"`
} }
@@ -198,6 +231,7 @@ type ExecutionTargetPresence struct {
// ExecutionTarget represents effective model runtime settings for a run. // ExecutionTarget represents effective model runtime settings for a run.
type ExecutionTarget struct { type ExecutionTarget struct {
BackendID string `yaml:"backend" json:"backend_id,omitempty"`
Endpoint string `yaml:"endpoint" json:"endpoint"` Endpoint string `yaml:"endpoint" json:"endpoint"`
Model string `yaml:"model" json:"model"` Model string `yaml:"model" json:"model"`
Temperature float64 `yaml:"temperature" json:"temperature"` Temperature float64 `yaml:"temperature" json:"temperature"`
@@ -212,6 +246,13 @@ type ExecutionTarget struct {
ExtraParams map[string]any `yaml:"extra_params" json:"extra_params"` ExtraParams map[string]any `yaml:"extra_params" json:"extra_params"`
} }
// ProfileInspection is the resolved result of exact profile inspection.
type ProfileInspection struct {
ProfileID string
EffectiveModelParams ExecutionTarget
APIKeyRequired bool
}
// OutputContract defines the requirements for the output artifact. // OutputContract defines the requirements for the output artifact.
type OutputContract struct { type OutputContract struct {
Format OutputFormat `yaml:"format"` Format OutputFormat `yaml:"format"`

View File

@@ -0,0 +1,19 @@
package domain
import (
"fmt"
"strings"
"unicode/utf8"
)
// NormalizeSessionID applies the shared session identifier rule.
func NormalizeSessionID(raw string) (string, error) {
normalized := strings.TrimSpace(raw)
if normalized == "" {
return "", nil
}
if length := utf8.RuneCountInString(normalized); length > SessionIDMaxLength {
return "", fmt.Errorf("session_id length %d exceeds maximum %d", length, SessionIDMaxLength)
}
return normalized, nil
}

View File

@@ -0,0 +1,57 @@
package domain
import (
"strings"
"testing"
)
func TestNormalizeSessionID(t *testing.T) {
tests := []struct {
name string
raw string
want string
wantErr bool
}{
{
name: "trims surrounding Unicode whitespace",
raw: "\u2003 session-123 \u2003",
want: "session-123",
},
{
name: "blank input is omitted",
raw: " \t\u2003 ",
want: "",
},
{
name: "maximum Unicode length is accepted",
raw: strings.Repeat("界", SessionIDMaxLength),
want: strings.Repeat("界", SessionIDMaxLength),
},
{
name: "one Unicode code point over maximum is rejected",
raw: strings.Repeat("界", SessionIDMaxLength+1),
wantErr: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, err := NormalizeSessionID(tt.raw)
if tt.wantErr {
if err == nil {
t.Fatal("expected normalization error")
}
if !strings.Contains(err.Error(), "exceeds maximum") {
t.Fatalf("expected useful length diagnostic, got %v", err)
}
return
}
if err != nil {
t.Fatalf("normalize session id: %v", err)
}
if got != tt.want {
t.Fatalf("normalized session id = %q, want %q", got, tt.want)
}
})
}
}

View File

@@ -1,25 +1,36 @@
package promptkit // Package jsonvalue validates and defensively copies JSON-compatible value
// trees used by configuration, request, and prepared-state boundaries.
package jsonvalue
import ( import (
"encoding/json" "encoding/json"
"fmt" "fmt"
"math" "math"
"reflect" "reflect"
"sort"
"strconv" "strconv"
) )
const maxSafeJSONInteger = 1<<53 - 1 const maxSafeJSONInteger = 1<<53 - 1
type jsonVisit struct { type visit struct {
typ reflect.Type typ reflect.Type
ptr uintptr ptr uintptr
} }
func copyPublicJSONMap(src map[string]any) (map[string]any, error) { // Copy validates and deeply copies a JSON-compatible value while preserving
// compatible concrete map, slice, array, scalar, and number types.
func Copy(src any) (any, error) {
return copyValue(reflect.ValueOf(src), "value", make(map[visit]struct{}), true)
}
// CopyMap validates and deeply copies an extra-parameter map while preserving
// compatible concrete map, slice, array, scalar, and number types.
func CopyMap(src map[string]any) (map[string]any, error) {
if src == nil { if src == nil {
return nil, nil return nil, nil
} }
copied, err := copyPublicJSONValue(reflect.ValueOf(src), "extra_params", make(map[jsonVisit]struct{})) copied, err := copyValue(reflect.ValueOf(src), "extra_params", make(map[visit]struct{}), false)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -30,7 +41,12 @@ func copyPublicJSONMap(src map[string]any) (map[string]any, error) {
return out, nil return out, nil
} }
func copyPublicJSONValue(value reflect.Value, path string, seen map[jsonVisit]struct{}) (any, error) { func copyValue(
value reflect.Value,
path string,
seen map[visit]struct{},
allowEmptyMapKeys bool,
) (any, error) {
if !value.IsValid() { if !value.IsValid() {
return nil, nil return nil, nil
} }
@@ -38,12 +54,15 @@ func copyPublicJSONValue(value reflect.Value, path string, seen map[jsonVisit]st
if value.IsNil() { if value.IsNil() {
return nil, nil return nil, nil
} }
return copyPublicJSONValue(value.Elem(), path, seen) return copyValue(value.Elem(), path, seen, allowEmptyMapKeys)
} }
if !value.CanInterface() { if !value.CanInterface() {
return nil, fmt.Errorf("%s: value cannot be copied", path) return nil, fmt.Errorf("%s: value cannot be copied", path)
} }
if number, ok := value.Interface().(json.Number); ok { if number, ok := value.Interface().(json.Number); ok {
if _, err := json.Marshal(number); err != nil {
return nil, fmt.Errorf("%s: invalid JSON number", path)
}
f, err := strconv.ParseFloat(number.String(), 64) f, err := strconv.ParseFloat(number.String(), 64)
if err != nil || math.IsNaN(f) || math.IsInf(f, 0) { if err != nil || math.IsNaN(f) || math.IsInf(f, 0) {
return nil, fmt.Errorf("%s: invalid JSON number", path) return nil, fmt.Errorf("%s: invalid JSON number", path)
@@ -65,8 +84,8 @@ func copyPublicJSONValue(value reflect.Value, path string, seen map[jsonVisit]st
} }
return value.Interface(), nil return value.Interface(), nil
case reflect.Float32, reflect.Float64: case reflect.Float32, reflect.Float64:
f := value.Convert(reflect.TypeOf(float64(0))).Float() number := value.Convert(reflect.TypeOf(float64(0))).Float()
if math.IsNaN(f) || math.IsInf(f, 0) { if math.IsNaN(number) || math.IsInf(number, 0) {
return nil, fmt.Errorf("%s: floating-point value must be finite", path) return nil, fmt.Errorf("%s: floating-point value must be finite", path)
} }
return value.Interface(), nil return value.Interface(), nil
@@ -74,28 +93,33 @@ func copyPublicJSONValue(value reflect.Value, path string, seen map[jsonVisit]st
if value.IsNil() { if value.IsNil() {
return nil, nil return nil, nil
} }
visit := jsonVisit{typ: value.Type(), ptr: value.Pointer()} current := visit{typ: value.Type(), ptr: value.Pointer()}
if _, ok := seen[visit]; ok { if _, ok := seen[current]; ok {
return nil, fmt.Errorf("%s: cyclic value is not supported", path) return nil, fmt.Errorf("%s: cyclic value is not supported", path)
} }
seen[visit] = struct{}{} seen[current] = struct{}{}
defer delete(seen, visit) defer delete(seen, current)
return copyPublicJSONValue(value.Elem(), path, seen) return copyValue(value.Elem(), path, seen, allowEmptyMapKeys)
case reflect.Map: case reflect.Map:
return copyPublicJSONMapValue(value, path, seen) return copyMapValue(value, path, seen, allowEmptyMapKeys)
case reflect.Slice: case reflect.Slice:
if value.IsNil() { if value.IsNil() {
return nil, nil return nil, nil
} }
return copyPublicJSONSequenceValue(value, path, seen) return copySequenceValue(value, path, seen, allowEmptyMapKeys)
case reflect.Array: case reflect.Array:
return copyPublicJSONSequenceValue(value, path, seen) return copySequenceValue(value, path, seen, allowEmptyMapKeys)
default: default:
return nil, fmt.Errorf("%s: unsupported JSON value type %s", path, value.Type()) return nil, fmt.Errorf("%s: unsupported JSON value type %s", path, value.Type())
} }
} }
func copyPublicJSONMapValue(value reflect.Value, path string, seen map[jsonVisit]struct{}) (any, error) { func copyMapValue(
value reflect.Value,
path string,
seen map[visit]struct{},
allowEmptyMapKeys bool,
) (any, error) {
if value.IsNil() { if value.IsNil() {
return nil, nil return nil, nil
} }
@@ -103,37 +127,43 @@ func copyPublicJSONMapValue(value reflect.Value, path string, seen map[jsonVisit
return nil, fmt.Errorf("%s: map key type %s is not supported", path, value.Type().Key()) return nil, fmt.Errorf("%s: map key type %s is not supported", path, value.Type().Key())
} }
visit := jsonVisit{typ: value.Type(), ptr: value.Pointer()} current := visit{typ: value.Type(), ptr: value.Pointer()}
if _, ok := seen[visit]; ok { if _, ok := seen[current]; ok {
return nil, fmt.Errorf("%s: cyclic value is not supported", path) return nil, fmt.Errorf("%s: cyclic value is not supported", path)
} }
seen[visit] = struct{}{} seen[current] = struct{}{}
defer delete(seen, visit) defer delete(seen, current)
keys := value.MapKeys()
sort.Slice(keys, func(i, j int) bool {
return keys[i].String() < keys[j].String()
})
type entry struct { type entry struct {
key reflect.Value key reflect.Value
name string name string
value any value any
} }
entries := make([]entry, 0, value.Len()) entries := make([]entry, 0, len(keys))
preserveType := true preserveType := true
elemType := value.Type().Elem() elementType := value.Type().Elem()
iter := value.MapRange() for _, key := range keys {
for iter.Next() {
key := iter.Key()
name := key.String() name := key.String()
copied, err := copyPublicJSONValue(iter.Value(), path+"."+name, seen) if name == "" && !allowEmptyMapKeys {
return nil, fmt.Errorf("%s: map key must not be empty", path)
}
copied, err := copyValue(value.MapIndex(key), path+"."+name, seen, allowEmptyMapKeys)
if err != nil { if err != nil {
return nil, err return nil, err
} }
entries = append(entries, entry{key: key, name: name, value: copied}) entries = append(entries, entry{key: key, name: name, value: copied})
if copied == nil { if copied == nil {
if !canAssignNil(elemType) { if !canAssignNil(elementType) {
preserveType = false preserveType = false
} }
continue continue
} }
if !reflect.TypeOf(copied).AssignableTo(elemType) { if !reflect.TypeOf(copied).AssignableTo(elementType) {
preserveType = false preserveType = false
} }
} }
@@ -142,7 +172,7 @@ func copyPublicJSONMapValue(value reflect.Value, path string, seen map[jsonVisit
out := reflect.MakeMapWithSize(value.Type(), len(entries)) out := reflect.MakeMapWithSize(value.Type(), len(entries))
for _, entry := range entries { for _, entry := range entries {
if entry.value == nil { if entry.value == nil {
out.SetMapIndex(entry.key, reflect.Zero(elemType)) out.SetMapIndex(entry.key, reflect.Zero(elementType))
continue continue
} }
out.SetMapIndex(entry.key, reflect.ValueOf(entry.value)) out.SetMapIndex(entry.key, reflect.ValueOf(entry.value))
@@ -157,33 +187,43 @@ func copyPublicJSONMapValue(value reflect.Value, path string, seen map[jsonVisit
return out, nil return out, nil
} }
func copyPublicJSONSequenceValue(value reflect.Value, path string, seen map[jsonVisit]struct{}) (any, error) { func copySequenceValue(
var visit jsonVisit value reflect.Value,
path string,
seen map[visit]struct{},
allowEmptyMapKeys bool,
) (any, error) {
var current visit
if value.Kind() == reflect.Slice { if value.Kind() == reflect.Slice {
visit = jsonVisit{typ: value.Type(), ptr: value.Pointer()} current = visit{typ: value.Type(), ptr: value.Pointer()}
if _, ok := seen[visit]; ok { if _, ok := seen[current]; ok {
return nil, fmt.Errorf("%s: cyclic value is not supported", path) return nil, fmt.Errorf("%s: cyclic value is not supported", path)
} }
seen[visit] = struct{}{} seen[current] = struct{}{}
defer delete(seen, visit) defer delete(seen, current)
} }
values := make([]any, value.Len()) values := make([]any, value.Len())
preserveType := true preserveType := true
elemType := value.Type().Elem() elementType := value.Type().Elem()
for i := 0; i < value.Len(); i++ { for i := 0; i < value.Len(); i++ {
copied, err := copyPublicJSONValue(value.Index(i), fmt.Sprintf("%s[%d]", path, i), seen) copied, err := copyValue(
value.Index(i),
fmt.Sprintf("%s[%d]", path, i),
seen,
allowEmptyMapKeys,
)
if err != nil { if err != nil {
return nil, err return nil, err
} }
values[i] = copied values[i] = copied
if copied == nil { if copied == nil {
if !canAssignNil(elemType) { if !canAssignNil(elementType) {
preserveType = false preserveType = false
} }
continue continue
} }
if !reflect.TypeOf(copied).AssignableTo(elemType) { if !reflect.TypeOf(copied).AssignableTo(elementType) {
preserveType = false preserveType = false
} }
} }
@@ -195,7 +235,7 @@ func copyPublicJSONSequenceValue(value reflect.Value, path string, seen map[json
} }
for i, copied := range values { for i, copied := range values {
if copied == nil { if copied == nil {
out.Index(i).Set(reflect.Zero(elemType)) out.Index(i).Set(reflect.Zero(elementType))
continue continue
} }
out.Index(i).Set(reflect.ValueOf(copied)) out.Index(i).Set(reflect.ValueOf(copied))

View File

@@ -0,0 +1,112 @@
package jsonvalue_test
import (
"encoding/json"
"math"
"reflect"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
)
func TestCopyMapPreservesTypesAndIsolatesMutations(t *testing.T) {
nested := map[string]int{"limit": 2}
sequence := []string{"one", "two"}
input := map[string]any{
"count": int64(7),
"number": json.Number("-1.25e+2"),
"nested": nested,
"sequence": sequence,
}
copied, err := jsonvalue.CopyMap(input)
if err != nil {
t.Fatalf("copy map: %v", err)
}
nested["limit"] = 99
sequence[0] = "changed"
input["added"] = true
if got, ok := copied["count"].(int64); !ok || got != 7 {
t.Fatalf("integer type or value changed: %#v", copied["count"])
}
if got, ok := copied["number"].(json.Number); !ok || got != "-1.25e+2" {
t.Fatalf("JSON number type or value changed: %#v", copied["number"])
}
if got := copied["nested"].(map[string]int)["limit"]; got != 2 {
t.Fatalf("nested map was not isolated: %d", got)
}
if got := copied["sequence"].([]string)[0]; got != "one" {
t.Fatalf("sequence was not isolated: %q", got)
}
if _, ok := copied["added"]; ok {
t.Fatalf("top-level map was not isolated: %#v", copied)
}
}
func TestCopyAllowsEmptyObjectKeysAndIsolatesMutations(t *testing.T) {
nested := map[string]any{"": []any{"original"}}
copiedValue, err := jsonvalue.Copy(nested)
if err != nil {
t.Fatalf("copy value: %v", err)
}
nested[""].([]any)[0] = "changed"
copied := copiedValue.(map[string]any)
if got := copied[""].([]any)[0]; got != "original" {
t.Fatalf("copied value was not isolated: %v", got)
}
}
func TestCopyMapRejectsInvalidValues(t *testing.T) {
cyclicMap := map[string]any{}
cyclicMap["self"] = cyclicMap
cyclicSlice := []any{nil}
cyclicSlice[0] = cyclicSlice
tests := []struct {
name string
value any
}{
{name: "empty nested key", value: map[string]int{"": 1}},
{name: "non-string map key", value: map[int]string{1: "one"}},
{name: "unsupported value", value: make(chan int)},
{name: "cyclic map", value: cyclicMap},
{name: "cyclic slice", value: cyclicSlice},
{name: "NaN", value: math.NaN()},
{name: "positive infinity", value: math.Inf(1)},
{name: "unsafe signed integer", value: int64(1 << 53)},
{name: "unsafe unsigned integer", value: uint64(1 << 53)},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
if _, err := jsonvalue.CopyMap(map[string]any{"value": tc.value}); err == nil {
t.Fatal("expected validation error")
}
})
}
}
func TestCopyMapValidatesJSONNumberSyntaxAndRange(t *testing.T) {
for _, number := range []json.Number{"0", "-1", "1.25", "-1.25e+2"} {
t.Run("valid "+number.String(), func(t *testing.T) {
got, err := jsonvalue.CopyMap(map[string]any{"value": number})
if err != nil {
t.Fatalf("copy valid JSON number: %v", err)
}
if !reflect.DeepEqual(got["value"], number) {
t.Fatalf("JSON number changed: got %#v want %#v", got["value"], number)
}
})
}
for _, number := range []json.Number{"", "01", "+1", "1.", ".1", "1e9999", "not-a-number"} {
t.Run("invalid "+number.String(), func(t *testing.T) {
if _, err := jsonvalue.CopyMap(map[string]any{"value": number}); err == nil {
t.Fatal("expected invalid JSON number error")
}
})
}
}

View File

@@ -12,7 +12,6 @@ import (
"os" "os"
"strings" "strings"
"time" "time"
"unicode/utf8"
"gitea.maximumdirect.net/eric/promptkit/internal/defaults" "gitea.maximumdirect.net/eric/promptkit/internal/defaults"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
@@ -177,12 +176,11 @@ func openAIChatRequestFromGenerateRequest(req domain.GenerateRequest, defaultMod
wireReq := openAIChatRequest{ wireReq := openAIChatRequest{
Model: model, Model: model,
} }
if sessionID := strings.TrimSpace(req.Prompt.SessionID); sessionID != "" { sessionID, err := domain.NormalizeSessionID(req.Prompt.SessionID)
if n := utf8.RuneCountInString(sessionID); n > domain.SessionIDMaxLength { if err != nil {
return openAIChatRequest{}, fmt.Errorf("session_id length %d exceeds maximum %d", n, domain.SessionIDMaxLength) return openAIChatRequest{}, err
} }
wireReq.SessionID = sessionID wireReq.SessionID = sessionID
}
wireReq.Messages = make([]openAIChatRequestMessage, 0, len(req.Prompt.Messages)) wireReq.Messages = make([]openAIChatRequestMessage, 0, len(req.Prompt.Messages))
for _, msg := range req.Prompt.Messages { for _, msg := range req.Prompt.Messages {
@@ -262,7 +260,7 @@ func openAIChatRequestPayload(req openAIChatRequest) (map[string]any, error) {
if key == "" { if key == "" {
return nil, errors.New("extra_params key must not be empty") return nil, errors.New("extra_params key must not be empty")
} }
if _, reserved := reservedOpenAIChatRequestFields[key]; reserved { if IsReservedOpenAIChatRequestField(key) {
return nil, fmt.Errorf("extra_params key %q collides with reserved request field", key) return nil, fmt.Errorf("extra_params key %q collides with reserved request field", key)
} }
if _, err := json.Marshal(value); err != nil { if _, err := json.Marshal(value); err != nil {
@@ -274,16 +272,23 @@ func openAIChatRequestPayload(req openAIChatRequest) (map[string]any, error) {
return out, nil return out, nil
} }
var reservedOpenAIChatRequestFields = map[string]struct{}{ // IsReservedOpenAIChatRequestField reports whether name is owned by the
"model": {}, // standard OpenAI-compatible chat request rather than extra parameters.
"session_id": {}, func IsReservedOpenAIChatRequestField(name string) bool {
"messages": {}, switch name {
"temperature": {}, case "model",
"max_tokens": {}, "session_id",
"top_p": {}, "messages",
"service_tier": {}, "temperature",
"reasoning_effort": {}, "max_tokens",
"response_format": {}, "top_p",
"service_tier",
"reasoning_effort",
"response_format":
return true
default:
return false
}
} }
type openAIChatRequestMessage struct { type openAIChatRequestMessage struct {

View File

@@ -1,9 +1,8 @@
id: aion-2 id: aion-2
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: aion-labs/aion-2.0 model: aion-labs/aion-2.0
temperature: 0.72 temperature: 0.72
reasoning_effort: high reasoning_effort: high
top_p: 0.95 top_p: 0.95
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: claude-fable-latest id: claude-fable-latest
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "~anthropic/claude-fable-latest" model: "~anthropic/claude-fable-latest"
reasoning_effort: high reasoning_effort: high
timeout_seconds: 600 timeout_seconds: 600
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: claude-haiku-latest id: claude-haiku-latest
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "~anthropic/claude-haiku-latest" model: "~anthropic/claude-haiku-latest"
reasoning_effort: medium reasoning_effort: medium
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: claude-opus-latest id: claude-opus-latest
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "~anthropic/claude-opus-latest" model: "~anthropic/claude-opus-latest"
reasoning_effort: high reasoning_effort: high
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: claude-sonnet-latest id: claude-sonnet-latest
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "~anthropic/claude-sonnet-latest" model: "~anthropic/claude-sonnet-latest"
reasoning_effort: high reasoning_effort: high
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: deepseek-3-2 id: deepseek-3-2
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: deepseek/deepseek-v3.2 model: deepseek/deepseek-v3.2
reasoning_effort: high reasoning_effort: high
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: deepseek-4-flash id: deepseek-4-flash
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: deepseek/deepseek-v4-flash model: deepseek/deepseek-v4-flash
#reasoning_effort: medium #reasoning_effort: medium
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: deepseek-4-pro id: deepseek-4-pro
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: deepseek/deepseek-v4-pro model: deepseek/deepseek-v4-pro
reasoning_effort: high reasoning_effort: high
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: gemini-2-flash-lite id: gemini-2-flash-lite
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "google/gemini-2.5-flash-lite" model: "google/gemini-2.5-flash-lite"
#temperature: 0.15 #temperature: 0.15
reasoning_effort: high reasoning_effort: high
#top_p: 0.98 #top_p: 0.98
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: gemini-2-flash id: gemini-2-flash
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "google/gemini-2.5-flash" model: "google/gemini-2.5-flash"
#temperature: 0.15 #temperature: 0.15
reasoning_effort: high reasoning_effort: high
#top_p: 0.98 #top_p: 0.98
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: gemini-2-pro id: gemini-2-pro
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "google/gemini-2.5-pro" model: "google/gemini-2.5-pro"
#temperature: 0.15 #temperature: 0.15
reasoning_effort: high reasoning_effort: high
#top_p: 0.98 #top_p: 0.98
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: gemini-3-flash-lite id: gemini-3-flash-lite
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "google/gemini-3.1-flash-lite" model: "google/gemini-3.1-flash-lite"
#temperature: 0.15 #temperature: 0.15
reasoning_effort: high reasoning_effort: high
#top_p: 0.98 #top_p: 0.98
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: gemini-flash-latest id: gemini-flash-latest
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "~google/gemini-flash-latest" model: "~google/gemini-flash-latest"
#temperature: 0.15 #temperature: 0.15
reasoning_effort: high reasoning_effort: high
#top_p: 0.98 #top_p: 0.98
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: gemini-pro-latest id: gemini-pro-latest
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "~google/gemini-pro-latest" model: "~google/gemini-pro-latest"
#temperature: 0.15 #temperature: 0.15
reasoning_effort: high reasoning_effort: high
#top_p: 0.98 #top_p: 0.98
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: gemma-4-31b id: gemma-4-31b
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: google/gemma-4-31b-it:exacto model: google/gemma-4-31b-it:exacto
temperature: 0.15 temperature: 0.15
reasoning_effort: high reasoning_effort: high
top_p: 0.98 top_p: 0.98
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: minimax-m2 id: minimax-m2
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: minimax/minimax-m2.5 model: minimax/minimax-m2.5
temperature: 0.5 temperature: 0.5
reasoning_effort: high reasoning_effort: high
top_p: 0.95 top_p: 0.95
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,9 +1,8 @@
id: minimax-m3 id: minimax-m3
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: minimax/minimax-m3 model: minimax/minimax-m3
#temperature: 0.5 #temperature: 0.5
reasoning_effort: high reasoning_effort: high
#top_p: 0.95 #top_p: 0.95
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: mistral-large-2512 id: mistral-large-2512
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: mistralai/mistral-large-2512 model: mistralai/mistral-large-2512
temperature: 0.15 temperature: 0.15
top_p: 0.98 top_p: 0.98
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY

View File

@@ -1,8 +1,7 @@
id: mistral-medium-3-5 id: mistral-medium-3-5
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: mistralai/mistral-medium-3-5 model: mistralai/mistral-medium-3-5
temperature: 0.15 temperature: 0.15
reasoning_effort: high reasoning_effort: high
top_p: 0.98 top_p: 0.98
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY

View File

@@ -1,7 +1,6 @@
id: mistral-small-3 id: mistral-small-3
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: mistralai/mistral-small-3.2-24b-instruct model: mistralai/mistral-small-3.2-24b-instruct
temperature: 0.05 temperature: 0.05
top_p: 1.0 top_p: 1.0
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY

View File

@@ -1,8 +1,7 @@
id: mistral-small-4 id: mistral-small-4
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: mistralai/mistral-small-2603 model: mistralai/mistral-small-2603
temperature: 0.1 temperature: 0.1
reasoning_effort: high reasoning_effort: high
top_p: 0.98 top_p: 0.98
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY

View File

@@ -1,7 +1,6 @@
id: nemotron-3-ultra id: nemotron-3-ultra
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: nvidia/nemotron-3-ultra-550b-a55b model: nvidia/nemotron-3-ultra-550b-a55b
reasoning_effort: high reasoning_effort: high
timeout_seconds: 180 timeout_seconds: 180
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: gpt-5-mini id: gpt-5-mini
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "openai/gpt-5.4-mini" model: "openai/gpt-5.4-mini"
reasoning_effort: high reasoning_effort: high
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -1,7 +1,6 @@
id: gpt-5-nano id: gpt-5-nano
endpoint: https://openrouter.ai/api/v1 backend: openrouter
model: "openai/gpt-5.4-nano" model: "openai/gpt-5.4-nano"
reasoning_effort: high reasoning_effort: high
timeout_seconds: 240 timeout_seconds: 240
api_key_env: OPENROUTER_API_KEY
service_tier: flex service_tier: flex

View File

@@ -7,6 +7,7 @@ import (
"strings" "strings"
"testing" "testing"
"gitea.maximumdirect.net/eric/promptkit/internal/backend"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
"gopkg.in/yaml.v3" "gopkg.in/yaml.v3"
@@ -28,6 +29,12 @@ func TestBuiltInProfilesValidateThroughRepository(t *testing.T) {
if p.ID != id { if p.ID != id {
t.Fatalf("expected profile id %q, got %q", id, p.ID) t.Fatalf("expected profile id %q, got %q", id, p.ID)
} }
if p.BackendID != backend.OpenRouterID {
t.Fatalf("expected profile %q to select %q, got %q", id, backend.OpenRouterID, p.BackendID)
}
if p.Endpoint != "" || p.APIKeyEnv != "" {
t.Fatalf("expected profile %q to inherit backend connection settings, got endpoint=%q api_key_env=%q", id, p.Endpoint, p.APIKeyEnv)
}
}) })
} }
} }
@@ -60,6 +67,15 @@ func loadBuiltInProfileIDs(t *testing.T) map[string]string {
if _, ok := raw["api_key"]; ok { if _, ok := raw["api_key"]; ok {
t.Fatalf("built-in profile %s contains raw api_key", name) t.Fatalf("built-in profile %s contains raw api_key", name)
} }
if raw["backend"] != backend.OpenRouterID {
t.Fatalf("built-in profile %s does not select %q", name, backend.OpenRouterID)
}
if _, ok := raw["endpoint"]; ok {
t.Fatalf("built-in profile %s repeats endpoint", name)
}
if _, ok := raw["api_key_env"]; ok {
t.Fatalf("built-in profile %s repeats api_key_env", name)
}
id, ok := raw["id"].(string) id, ok := raw["id"].(string)
if !ok || strings.TrimSpace(id) == "" { if !ok || strings.TrimSpace(id) == "" {
t.Fatalf("built-in profile %s has missing id", name) t.Fatalf("built-in profile %s has missing id", name)

View File

@@ -121,6 +121,7 @@ func loadProfile(ctx context.Context, fsys fs.FS, root string, id string) (*doma
if prof.ID != id { if prof.ID != id {
continue continue
} }
prof.BackendID = strings.TrimSpace(prof.BackendID)
if err := validateProfile(&prof); err != nil { if err := validateProfile(&prof); err != nil {
if errors.Is(err, ErrRawAPIKeyNotAllowed) { if errors.Is(err, ErrRawAPIKeyNotAllowed) {
return nil, fmt.Errorf("%w: %s", err, relPath) return nil, fmt.Errorf("%w: %s", err, relPath)
@@ -189,8 +190,8 @@ func validateProfile(p *domain.ExecutionProfile) error {
if strings.TrimSpace(p.ID) == "" { if strings.TrimSpace(p.ID) == "" {
return errors.New("id is required") return errors.New("id is required")
} }
if strings.TrimSpace(p.Endpoint) == "" { if strings.TrimSpace(p.BackendID) == "" && strings.TrimSpace(p.Endpoint) == "" {
return errors.New("endpoint is required") return errors.New("backend or endpoint is required")
} }
if strings.TrimSpace(p.Model) == "" { if strings.TrimSpace(p.Model) == "" {
return errors.New("model is required") return errors.New("model is required")

View File

@@ -52,6 +52,43 @@ func TestFilesystemRepository_GetProfile(t *testing.T) {
} }
}) })
t.Run("backend and endpoint connection matrix", func(t *testing.T) {
tests := []struct {
name string
connection string
wantBackend string
wantEndpoint string
wantErr bool
}{
{name: "backend only", connection: "backend: ' openrouter '", wantBackend: "openrouter"},
{name: "endpoint only", connection: "endpoint: http://localhost:8000/v1", wantEndpoint: "http://localhost:8000/v1"},
{name: "both", connection: "backend: openrouter\nendpoint: http://localhost:8000/v1", wantBackend: "openrouter", wantEndpoint: "http://localhost:8000/v1"},
{name: "neither", wantErr: true},
{name: "blank backend", connection: "backend: ' '", wantErr: true},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
id := "connection-" + strings.ReplaceAll(tt.name, " ", "-")
writeProfileTestFile(t, filepath.Join(tmpDir, id+".yaml"), "id: "+id+"\nmodel: model\n"+tt.connection+"\n")
p, err := repo.GetProfile(ctx, id)
if tt.wantErr {
if !errors.Is(err, ErrInvalidProfile) {
t.Fatalf("expected ErrInvalidProfile, got %v", err)
}
return
}
if err != nil {
t.Fatalf("expected profile to load, got %v", err)
}
if p.BackendID != tt.wantBackend || p.Endpoint != tt.wantEndpoint {
t.Fatalf("unexpected connection values: backend=%q endpoint=%q", p.BackendID, p.Endpoint)
}
})
}
})
t.Run("valid profile with api_key_env", func(t *testing.T) { t.Run("valid profile with api_key_env", func(t *testing.T) {
p, err := repo.GetProfile(ctx, "local-secure") p, err := repo.GetProfile(ctx, "local-secure")
if err != nil { if err != nil {

View File

@@ -5,10 +5,9 @@ import (
"context" "context"
"errors" "errors"
"fmt" "fmt"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"strings"
"text/template" "text/template"
"unicode/utf8"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
) )
var ( var (
@@ -95,10 +94,6 @@ func (r *goRenderer) Render(ctx context.Context, definition *domain.PromptDefini
} }
func renderSessionID(raw string, funcs template.FuncMap, vars map[string]string) (string, error) { func renderSessionID(raw string, funcs template.FuncMap, vars map[string]string) (string, error) {
if strings.TrimSpace(raw) == "" {
return "", nil
}
tmpl, err := template.New("session_id").Funcs(funcs).Option("missingkey=error").Parse(raw) tmpl, err := template.New("session_id").Funcs(funcs).Option("missingkey=error").Parse(raw)
if err != nil { if err != nil {
return "", fmt.Errorf("%w: session_id: %v", ErrInvalidTemplate, err) return "", fmt.Errorf("%w: session_id: %v", ErrInvalidTemplate, err)
@@ -109,9 +104,9 @@ func renderSessionID(raw string, funcs template.FuncMap, vars map[string]string)
return "", fmt.Errorf("%w: session_id: %w", ErrRenderFailure, err) return "", fmt.Errorf("%w: session_id: %w", ErrRenderFailure, err)
} }
sessionID := strings.TrimSpace(buf.String()) sessionID, err := domain.NormalizeSessionID(buf.String())
if n := utf8.RuneCountInString(sessionID); n > domain.SessionIDMaxLength { if err != nil {
return "", fmt.Errorf("%w: session_id length %d exceeds maximum %d", ErrRenderFailure, n, domain.SessionIDMaxLength) return "", fmt.Errorf("%w: session_id: %v", ErrRenderFailure, err)
} }
return sessionID, nil return sessionID, nil
} }

View File

@@ -0,0 +1,24 @@
package usecase
import (
"fmt"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/capacity"
)
// CapacityError identifies bounded admission rejected for one selected backend.
type CapacityError struct {
BackendID string
}
func (e *CapacityError) Error() string {
if e == nil || strings.TrimSpace(e.BackendID) == "" {
return capacity.ErrCapacityExceeded.Error()
}
return fmt.Sprintf("backend %q admission: %v", e.BackendID, capacity.ErrCapacityExceeded)
}
func (e *CapacityError) Unwrap() error {
return capacity.ErrCapacityExceeded
}

View File

@@ -0,0 +1,298 @@
package usecase
import (
"context"
"fmt"
"sync"
"time"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
"gitea.maximumdirect.net/eric/promptkit/internal/validate"
)
type preparedExecutionState uint8
const (
preparedExecutionReady preparedExecutionState = iota
preparedExecutionClaimed
preparedExecutionDiscarded
)
// PreparedExecution owns one frozen, single-use runner execution.
type PreparedExecution struct {
owner *Runner
mu sync.Mutex
state preparedExecutionState
details *domain.PreparedRun
payload *preparedExecutionPayload
}
type preparedExecutionPayload struct {
prepared *domain.PreparedRun
validation validate.PreparedValidation
directKey string
}
// PrepareExecution completes preparation without generation or admission and
// returns a runner-bound, single-use execution.
func (r *Runner) PrepareExecution(ctx context.Context, req domain.RunRequest) (*PreparedExecution, error) {
state, err := r.resolvePreparation(ctx, req, time.Now().UTC())
if err != nil {
return nil, err
}
validationPlan, err := r.prepareValidation(ctx, state.effectiveContract)
if err != nil {
return nil, err
}
structuredOutput, err := r.structuredOutputFromValidationPlan(
state.definition,
state.effectiveContract,
validationPlan,
)
if err != nil {
return nil, err
}
prepared, err := r.completePreparationWithStructuredOutput(ctx, req, state, structuredOutput)
if err != nil {
return nil, err
}
executionSnapshot, err := clonePreparedRun(prepared)
if err != nil {
return nil, fmt.Errorf("%w: failed to copy prepared execution: %v", ErrInvalidRequest, err)
}
details, err := clonePreparedRun(executionSnapshot)
if err != nil {
return nil, fmt.Errorf("%w: failed to copy prepared execution details: %v", ErrInvalidRequest, err)
}
return &PreparedExecution{
owner: r,
state: preparedExecutionReady,
details: details,
payload: &preparedExecutionPayload{
prepared: executionSnapshot,
validation: validationPlan,
directKey: state.effectiveModel.APIKey,
},
}, nil
}
func (r *Runner) prepareValidation(
ctx context.Context,
contract domain.OutputContract,
) (validate.PreparedValidation, error) {
if r.validator == nil {
return noOpPreparedValidation{contract: contract}, nil
}
preparer, ok := r.validator.(validate.ValidationPreparer)
if !ok {
return nil, fmt.Errorf("%w: validator does not support prepared validation", ErrValidation)
}
plan, err := preparer.PrepareValidation(ctx, contract)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrValidation, err)
}
if plan == nil {
return nil, fmt.Errorf("%w: validator returned nil prepared validation", ErrValidation)
}
return plan, nil
}
func (r *Runner) structuredOutputFromValidationPlan(
def *domain.PromptDefinition,
contract domain.OutputContract,
plan validate.PreparedValidation,
) (*domain.StructuredOutputSpec, error) {
if contract.ValidationMode != domain.ValidationJSONSchema {
return nil, nil
}
schemaDocument := plan.SchemaDocument()
if schemaDocument == nil {
if r.validator == nil {
return nil, nil
}
return nil, fmt.Errorf("%w: prepared json_schema validation has no schema document", ErrValidation)
}
return structuredOutputSpec(def, schemaDocument), nil
}
// Details returns a fresh credential-redacted copy of the prepared run.
func (p *PreparedExecution) Details() *domain.PreparedRun {
if p == nil {
return nil
}
p.mu.Lock()
detailsSnapshot := p.details
p.mu.Unlock()
details, err := clonePreparedRun(detailsSnapshot)
if err != nil {
return nil
}
return details
}
// Discard invalidates an unclaimed execution and drops its private payload.
func (p *PreparedExecution) Discard() {
if p == nil {
return
}
p.mu.Lock()
if p.state != preparedExecutionReady {
p.mu.Unlock()
return
}
p.state = preparedExecutionDiscarded
payload := p.payload
p.payload = nil
p.mu.Unlock()
payload.clear()
}
// RunPrepared claims and executes one prepared execution owned by this runner.
func (r *Runner) RunPrepared(ctx context.Context, prepared *PreparedExecution) (*domain.RunResult, error) {
payload, err := prepared.claim(r)
if err != nil {
return nil, err
}
defer payload.clear()
runID, err := newRunID()
if err != nil {
return nil, fmt.Errorf("failed to create run id: %w", err)
}
start := time.Now().UTC()
target := payload.prepared.EffectiveModelParams
if err := validateAPIKey(target.APIKeyEnv, payload.directKey, target.APIKeyRequired); err != nil {
return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
}
release, err := r.admitRun(ctx, target.BackendID)
if err != nil {
return nil, err
}
defer release()
return r.executePreparedRun(ctx, payload.prepared, payload.directKey, runID, start, func(
ctx context.Context,
artifact *domain.Artifact,
attemptsUsed int,
) (domain.ValidationResult, error) {
result, validationErr := payload.validation.Validate(ctx, artifact)
if validationErr != nil {
return domain.ValidationResult{}, validationErr
}
result.RepairAttempts = attemptsUsed
return result, nil
})
}
func (p *PreparedExecution) claim(owner *Runner) (*preparedExecutionPayload, error) {
if p == nil || owner == nil || p.owner != owner {
return nil, fmt.Errorf("%w: prepared execution does not belong to this runner", ErrInvalidRequest)
}
p.mu.Lock()
defer p.mu.Unlock()
if p.state != preparedExecutionReady || p.payload == nil {
return nil, fmt.Errorf("%w: prepared execution is not ready", ErrInvalidRequest)
}
p.state = preparedExecutionClaimed
payload := p.payload
p.payload = nil
return payload, nil
}
func (p *preparedExecutionPayload) clear() {
if p == nil {
return
}
if p.prepared != nil {
p.prepared.EffectiveModelParams.APIKey = ""
}
p.prepared = nil
p.validation = nil
p.directKey = ""
}
type noOpPreparedValidation struct {
contract domain.OutputContract
}
func (p noOpPreparedValidation) Validate(
ctx context.Context,
_ *domain.Artifact,
) (domain.ValidationResult, error) {
if err := ctx.Err(); err != nil {
return domain.ValidationResult{}, err
}
return domain.ValidationResult{
Status: domain.ValidationSkipped,
Mode: p.contract.ValidationMode,
SchemaPath: p.contract.SchemaPath,
RepairAttempts: p.contract.RepairAttempts,
IsValid: true,
}, nil
}
func (noOpPreparedValidation) SchemaDocument() any {
return nil
}
func clonePreparedRun(source *domain.PreparedRun) (*domain.PreparedRun, error) {
if source == nil {
return nil, nil
}
copied := *source
extraParams, err := jsonvalue.CopyMap(source.EffectiveModelParams.ExtraParams)
if err != nil {
return nil, err
}
copied.EffectiveModelParams.ExtraParams = extraParams
if source.InputHashes != nil {
copied.InputHashes = make(map[string]string, len(source.InputHashes))
for name, hash := range source.InputHashes {
copied.InputHashes[name] = hash
}
}
if source.Messages != nil {
copied.Messages = make([]domain.RenderedMessage, len(source.Messages))
for i, message := range source.Messages {
copied.Messages[i] = message
if message.CacheControl != nil {
cacheControl := *message.CacheControl
copied.Messages[i].CacheControl = &cacheControl
}
}
}
if source.StructuredOutput != nil {
structuredOutput := *source.StructuredOutput
copied.StructuredOutput = &structuredOutput
if source.StructuredOutput.JSONSchema != nil {
jsonSchema := *source.StructuredOutput.JSONSchema
copied.StructuredOutput.JSONSchema = &jsonSchema
schema, copyErr := cloneJSONValue(source.StructuredOutput.JSONSchema.Schema)
if copyErr != nil {
return nil, copyErr
}
copied.StructuredOutput.JSONSchema.Schema = schema
}
}
return &copied, nil
}
func cloneJSONValue(source any) (any, error) {
return jsonvalue.Copy(source)
}

View File

@@ -0,0 +1,510 @@
package usecase
import (
"context"
"errors"
"os"
"reflect"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/validate"
)
type recordingPreparedValidation struct {
contract domain.OutputContract
schemaDocument any
results []domain.ValidationResult
errs []error
artifacts []string
}
func (p *recordingPreparedValidation) Validate(
_ context.Context,
artifact *domain.Artifact,
) (domain.ValidationResult, error) {
p.artifacts = append(p.artifacts, string(artifact.Body))
index := len(p.artifacts) - 1
if index < len(p.errs) && p.errs[index] != nil {
return domain.ValidationResult{}, p.errs[index]
}
if len(p.results) == 0 {
return domain.ValidationResult{
Status: domain.ValidationPassed,
Mode: p.contract.ValidationMode,
IsValid: true,
}, nil
}
if index >= len(p.results) {
index = len(p.results) - 1
}
return p.results[index], nil
}
func (p *recordingPreparedValidation) SchemaDocument() any {
return p.schemaDocument
}
type recordingValidationPreparer struct {
plan *recordingPreparedValidation
prepareErr error
prepareCalls int
directValidateCalls int
}
func (v *recordingValidationPreparer) Validate(
context.Context,
*domain.Artifact,
domain.OutputContract,
) (domain.ValidationResult, error) {
v.directValidateCalls++
return domain.ValidationResult{}, errors.New("live validation must not be used")
}
func (v *recordingValidationPreparer) PrepareValidation(
_ context.Context,
contract domain.OutputContract,
) (validate.PreparedValidation, error) {
v.prepareCalls++
if v.prepareErr != nil {
return nil, v.prepareErr
}
v.plan.contract = contract
return v.plan, nil
}
func TestRunnerPrepareExecutionCompletesWithoutAdmissionOrGeneration(t *testing.T) {
schemaDocument := map[string]any{
"type": "object",
"properties": map[string]any{
"value": map[string]any{"type": "string"},
"": map[string]any{"type": "boolean"},
},
}
def := promptDef(domain.FormatJSON, domain.ValidationJSONSchema, 0)
def.Validation.SchemaPath = "schema.json"
reader := defaultArtifactReader()
renderer := &fakeRenderer{rendered: &domain.RenderedPrompt{
SessionID: "prepared-session",
Messages: []domain.RenderedMessage{{
Role: "user",
Content: "original message",
}},
}}
llmClient := &fakeLLM{forbid: true}
validator := &recordingValidationPreparer{
plan: &recordingPreparedValidation{schemaDocument: schemaDocument},
}
admitter := &fakeRunAdmitter{}
profile := defaultExecutionProfile()
profile.ExtraParams = map[string]any{
"metadata": map[string]any{"source": "original"},
}
runner := NewRunner(
&fakePromptRepo{def: def},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": profile}},
nil,
reader,
renderer,
llmClient,
validator,
admitter,
)
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
APIKey: "direct-test-key",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
defer prepared.Discard()
if validator.prepareCalls != 1 || validator.directValidateCalls != 0 {
t.Fatalf(
"validation calls=(prepare=%d direct=%d), want (1, 0)",
validator.prepareCalls,
validator.directValidateCalls,
)
}
if reader.calls != 1 || renderer.calls != 1 {
t.Fatalf("completion calls=(artifact=%d render=%d), want (1, 1)", reader.calls, renderer.calls)
}
if len(admitter.backendIDs) != 0 || llmClient.calls != 0 {
t.Fatalf("prepare invoked execution collaborators: admission=%v generation=%d", admitter.backendIDs, llmClient.calls)
}
first := prepared.Details()
if first == nil {
t.Fatal("prepared details are nil")
}
if first.EffectiveModelParams.APIKey != "" {
t.Fatal("prepared details retained the direct API key")
}
if first.StructuredOutput == nil ||
first.StructuredOutput.JSONSchema == nil ||
!reflect.DeepEqual(first.StructuredOutput.JSONSchema.Schema, schemaDocument) {
t.Fatalf("prepared details have unexpected structured output: %#v", first.StructuredOutput)
}
first.Messages[0].Content = "caller mutation"
first.InputHashes["input"] = "caller mutation"
first.EffectiveModelParams.ExtraParams["metadata"].(map[string]any)["source"] = "caller mutation"
first.StructuredOutput.JSONSchema.Schema.(map[string]any)["type"] = "string"
renderer.rendered.Messages[0].Content = "source mutation"
second := prepared.Details()
if second.Messages[0].Content != "original message" ||
second.InputHashes["input"] == "caller mutation" ||
second.EffectiveModelParams.ExtraParams["metadata"].(map[string]any)["source"] != "original" ||
second.StructuredOutput.JSONSchema.Schema.(map[string]any)["type"] != "object" {
t.Fatalf("details did not preserve an independent snapshot: %#v", second)
}
}
func TestRunnerRunPreparedRechecksEnvironmentCredentialBeforeAdmission(t *testing.T) {
const environmentName = "PROMPTKIT_PREPARED_EXECUTION_TEST_KEY"
t.Setenv(environmentName, "available-during-preparation")
profile := defaultExecutionProfile()
profile.APIKeyEnv = environmentName
validator := &recordingValidationPreparer{plan: &recordingPreparedValidation{}}
admitter := &fakeRunAdmitter{}
llmClient := &fakeLLM{resp: &domain.GenerateResponse{Content: "unexpected"}}
runner := NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": profile}},
nil,
defaultArtifactReader(),
defaultRenderer(),
llmClient,
validator,
admitter,
)
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
if err := os.Unsetenv(environmentName); err != nil {
t.Fatalf("unset credential environment: %v", err)
}
result, err := runner.RunPrepared(context.Background(), prepared)
if result != nil {
t.Fatalf("credential failure returned partial result: %+v", result)
}
if !errors.Is(err, ErrInvalidRequest) || !errors.Is(err, ErrAPIKeyEnvMissing) {
t.Fatalf("credential error identities are missing: %v", err)
}
if len(admitter.backendIDs) != 0 || llmClient.calls != 0 {
t.Fatalf("credential failure reached admission or generation: admission=%v generation=%d", admitter.backendIDs, llmClient.calls)
}
if _, err := runner.RunPrepared(context.Background(), prepared); !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("credential failure did not consume execution: %v", err)
}
}
func TestRunnerRunPreparedKeepsDirectCredentialOutOfMetadata(t *testing.T) {
const directKey = "direct-prepared-test-key"
profile := defaultExecutionProfile()
profile.APIKeyRequired = true
client := &fakeLLM{resp: &domain.GenerateResponse{Content: "ok"}}
runner := NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": profile}},
nil,
defaultArtifactReader(),
defaultRenderer(),
client,
&recordingValidationPreparer{plan: &recordingPreparedValidation{}},
nil,
)
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
APIKey: directKey,
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
if details := prepared.Details(); details.EffectiveModelParams.APIKey != "" {
t.Fatal("prepared details retained direct credential")
}
result, err := runner.RunPrepared(context.Background(), prepared)
if err != nil {
t.Fatalf("run prepared: %v", err)
}
if client.lastReq.Target.APIKey != directKey {
t.Fatal("generation did not receive direct credential")
}
if result.EffectiveModelParams.APIKey != "" {
t.Fatal("run result retained direct credential")
}
}
func TestRunnerRunPreparedUsesFrozenValidationForInitialAndRepairOutputs(t *testing.T) {
validator := &recordingValidationPreparer{
plan: &recordingPreparedValidation{
results: []domain.ValidationResult{
{
Status: domain.ValidationFailed,
Mode: domain.ValidationJSON,
Errors: []string{"invalid"},
IsValid: false,
},
{
Status: domain.ValidationPassed,
Mode: domain.ValidationJSON,
IsValid: true,
},
},
},
}
repairer := &fakeRepairer{
responses: []*domain.GenerateResponse{{Content: `{"repaired":true}`}},
}
admitter := &fakeRunAdmitter{}
reader := defaultArtifactReader()
renderer := defaultRenderer()
runner := NewRunnerWithRepairer(
&fakePromptRepo{def: promptDef(domain.FormatJSON, domain.ValidationJSON, 1)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
reader,
renderer,
&fakeLLM{resp: &domain.GenerateResponse{Content: `{"broken":true}`}},
validator,
repairer,
admitter,
)
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
result, err := runner.RunPrepared(context.Background(), prepared)
if err != nil {
t.Fatalf("run prepared: %v", err)
}
if !reflect.DeepEqual(validator.plan.artifacts, []string{`{"broken":true}`, `{"repaired":true}`}) {
t.Fatalf("prepared validation artifacts=%#v", validator.plan.artifacts)
}
if validator.directValidateCalls != 0 || repairer.calls != 1 {
t.Fatalf("validation/repair calls=(direct=%d repair=%d), want (0, 1)", validator.directValidateCalls, repairer.calls)
}
if result.Validation.Status != domain.ValidationPassed || result.Validation.RepairAttempts != 1 {
t.Fatalf("unexpected repaired validation result: %+v", result.Validation)
}
if admitter.releaseCalls != 1 {
t.Fatalf("admission releases=%d, want 1", admitter.releaseCalls)
}
if reader.calls != 1 || renderer.calls != 1 {
t.Fatalf("execution reopened preparation sources: artifact=%d render=%d", reader.calls, renderer.calls)
}
}
func TestRunnerRunPreparedReleasesAdmissionAcrossExecutionErrors(t *testing.T) {
generationFailure := errors.New("generation failed")
validationFailure := errors.New("validation failed")
repairFailure := errors.New("repair failed")
tests := []struct {
name string
generationErr error
validation *recordingPreparedValidation
repairer *fakeRepairer
wantError error
}{
{
name: "generation failure",
generationErr: generationFailure,
validation: &recordingPreparedValidation{},
wantError: ErrLLMGenerate,
},
{
name: "validation failure",
validation: &recordingPreparedValidation{
errs: []error{validationFailure},
},
wantError: ErrValidation,
},
{
name: "repair failure",
validation: &recordingPreparedValidation{
results: []domain.ValidationResult{{
Status: domain.ValidationFailed,
Mode: domain.ValidationJSON,
Errors: []string{"invalid"},
IsValid: false,
}},
},
repairer: &fakeRepairer{err: repairFailure},
wantError: ErrValidation,
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
def := promptDef(domain.FormatJSON, domain.ValidationJSON, 1)
if test.repairer == nil {
def.Validation.RepairAttempts = 0
}
validator := &recordingValidationPreparer{plan: test.validation}
admitter := &fakeRunAdmitter{}
runner := NewRunnerWithRepairer(
&fakePromptRepo{def: def},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
defaultArtifactReader(),
defaultRenderer(),
&fakeLLM{
resp: &domain.GenerateResponse{Content: `{"value":true}`},
err: test.generationErr,
},
validator,
test.repairer,
admitter,
)
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
result, err := runner.RunPrepared(context.Background(), prepared)
if result != nil || !errors.Is(err, test.wantError) {
t.Fatalf("run prepared=(%+v, %v), want %v", result, err, test.wantError)
}
if len(admitter.backendIDs) != 1 || admitter.releaseCalls != 1 {
t.Fatalf(
"admission calls=%#v releases=%d, want one each",
admitter.backendIDs,
admitter.releaseCalls,
)
}
})
}
}
func TestRunnerPreparedExecutionOwnershipUseAndDiscard(t *testing.T) {
newRunner := func() *Runner {
return NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationNone, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
defaultArtifactReader(),
defaultRenderer(),
&fakeLLM{resp: &domain.GenerateResponse{Content: "ok"}},
&recordingValidationPreparer{plan: &recordingPreparedValidation{}},
nil,
)
}
request := domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
}
owner := newRunner()
prepared, err := owner.PrepareExecution(context.Background(), request)
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
if _, err := newRunner().RunPrepared(context.Background(), prepared); !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("foreign runner error=%v, want ErrInvalidRequest", err)
}
if _, err := owner.RunPrepared(context.Background(), prepared); err != nil {
t.Fatalf("owner run prepared: %v", err)
}
if _, err := owner.RunPrepared(context.Background(), prepared); !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("second owner run error=%v, want ErrInvalidRequest", err)
}
if prepared.Details() == nil {
t.Fatal("details unavailable after execution")
}
discarded, err := owner.PrepareExecution(context.Background(), request)
if err != nil {
t.Fatalf("prepare discarded execution: %v", err)
}
discarded.Discard()
discarded.Discard()
if _, err := owner.RunPrepared(context.Background(), discarded); !errors.Is(err, ErrInvalidRequest) {
t.Fatalf("discarded execution error=%v, want ErrInvalidRequest", err)
}
if discarded.Details() == nil {
t.Fatal("details unavailable after discard")
}
}
func TestRunnerPreparedExecutionWithoutValidatorSkipsValidation(t *testing.T) {
runner := NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatJSON, domain.ValidationJSON, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
defaultArtifactReader(),
defaultRenderer(),
&fakeLLM{resp: &domain.GenerateResponse{Content: `{}`}},
nil,
nil,
)
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
result, err := runner.RunPrepared(context.Background(), prepared)
if err != nil {
t.Fatalf("run prepared: %v", err)
}
if result.Validation.Status != domain.ValidationSkipped || !result.Validation.IsValid {
t.Fatalf("unexpected no-validator result: %+v", result.Validation)
}
}
func TestRunnerPrepareExecutionRequiresValidationPreparer(t *testing.T) {
reader := defaultArtifactReader()
runner := NewRunner(
&fakePromptRepo{def: promptDef(domain.FormatText, domain.ValidationBasic, 0)},
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
reader,
defaultRenderer(),
&fakeLLM{forbid: true},
&fakeValidator{},
nil,
)
prepared, err := runner.PrepareExecution(context.Background(), domain.RunRequest{
PromptID: "p",
ProfileID: "exec",
Inputs: singleInputRef(),
})
if prepared != nil || !errors.Is(err, ErrValidation) {
t.Fatalf("prepare execution=(%+v, %v), want ErrValidation", prepared, err)
}
if reader.calls != 0 {
t.Fatalf("unsupported validator allowed completion, artifact calls=%d", reader.calls)
}
}

View File

@@ -0,0 +1,103 @@
package usecase
import (
"context"
"errors"
"fmt"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
type resolvedProfileSelection struct {
id string
profile *domain.ExecutionProfile
backend *domain.Backend
}
func (r *Runner) resolveProfileSelection(
ctx context.Context,
profileID string,
) (*resolvedProfileSelection, error) {
normalizedID := strings.TrimSpace(profileID)
if normalizedID == "" {
return nil, fmt.Errorf("%w: profile id is required", ErrInvalidRequest)
}
if r == nil || r.profiles == nil {
return nil, fmt.Errorf("%w: profile repository is not configured", ErrProfileLoad)
}
selectedProfile, err := r.profiles.GetProfile(ctx, normalizedID)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrProfileLoad, err)
}
if selectedProfile == nil {
return nil, fmt.Errorf("%w: profile repository returned nil profile", ErrProfileLoad)
}
profileValue := *selectedProfile
profileValue.BackendID = strings.TrimSpace(profileValue.BackendID)
var selectedBackend *domain.Backend
if profileValue.BackendID != "" {
if r.backends == nil {
return nil, fmt.Errorf("%w: backend %q cannot be resolved", ErrProfileLoad, profileValue.BackendID)
}
backendValue, err := r.backends.GetBackend(profileValue.BackendID)
if err != nil {
return nil, fmt.Errorf("%w: backend %q: %w", ErrProfileLoad, profileValue.BackendID, err)
}
selectedBackend = &backendValue
}
return &resolvedProfileSelection{
id: normalizedID,
profile: &profileValue,
backend: selectedBackend,
}, nil
}
func validateResolvedExecutionTarget(target domain.ExecutionTarget) error {
if strings.TrimSpace(target.Endpoint) == "" {
return errors.New("execution endpoint is required")
}
if strings.TrimSpace(target.Model) == "" {
return errors.New("execution model is required")
}
return nil
}
// InspectProfile resolves one explicit profile without prompt or execution work.
func (r *Runner) InspectProfile(
ctx context.Context,
profileID string,
) (*domain.ProfileInspection, error) {
normalizedID := strings.TrimSpace(profileID)
if normalizedID == "" {
return nil, fmt.Errorf("%w: profile id is required", ErrInvalidRequest)
}
select {
case <-ctx.Done():
return nil, fmt.Errorf("%w: %w", ErrProfileLoad, ctx.Err())
default:
}
selection, err := r.resolveProfileSelection(ctx, normalizedID)
if err != nil {
return nil, err
}
target, _, err := resolveExecutionTarget(selection.backend, selection.profile, nil)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrProfileLoad, err)
}
if err := validateResolvedExecutionTarget(target); err != nil {
return nil, fmt.Errorf("%w: %w", ErrProfileLoad, err)
}
target.APIKey = ""
return &domain.ProfileInspection{
ProfileID: selection.id,
EffectiveModelParams: target,
APIKeyRequired: target.APIKeyRequired,
}, nil
}

View File

@@ -0,0 +1,213 @@
package usecase
import (
"context"
"errors"
"reflect"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/defaults"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/profile"
)
type inspectionProfileRepository struct {
profile *domain.ExecutionProfile
err error
calls int
id string
}
func (r *inspectionProfileRepository) GetProfile(
_ context.Context,
id string,
) (*domain.ExecutionProfile, error) {
r.calls++
r.id = id
if r.err != nil {
return nil, r.err
}
return r.profile, nil
}
type inspectionBackendResolver struct {
backend domain.Backend
err error
calls int
id string
}
func (r *inspectionBackendResolver) GetBackend(id string) (domain.Backend, error) {
r.calls++
r.id = id
if r.err != nil {
return domain.Backend{}, r.err
}
return r.backend, nil
}
func TestRunnerInspectProfileResolvesProfileAndBackendOnce(t *testing.T) {
profiles := &inspectionProfileRepository{profile: &domain.ExecutionProfile{
ID: "profile",
BackendID: " backend ",
Model: "profile-model",
Temperature: 0.4,
MaxTokens: 32,
TimeoutSeconds: 45,
ServiceTier: "priority",
ReasoningEffort: "high",
ExtraParams: map[string]any{
"profile": "value",
},
}}
backends := &inspectionBackendResolver{backend: domain.Backend{
ID: "backend",
Endpoint: "https://backend.example/v1",
APIKeyEnv: "BACKEND_KEY",
ExtraParams: map[string]any{
"backend": "value",
},
}}
runner := &Runner{profiles: profiles, backends: backends}
inspection, err := runner.InspectProfile(context.Background(), " profile ")
if err != nil {
t.Fatalf("inspect profile: %v", err)
}
if profiles.calls != 1 || profiles.id != "profile" {
t.Fatalf("profile lookup=(calls=%d id=%q), want one exact lookup", profiles.calls, profiles.id)
}
if backends.calls != 1 || backends.id != "backend" {
t.Fatalf("backend lookup=(calls=%d id=%q), want one exact lookup", backends.calls, backends.id)
}
if profiles.profile.BackendID != " backend " {
t.Fatalf("inspection mutated repository profile backend: %q", profiles.profile.BackendID)
}
wantTarget := domain.ExecutionTarget{
BackendID: "backend",
Endpoint: "https://backend.example/v1",
Model: "profile-model",
Temperature: 0.4,
MaxTokens: 32,
TopP: defaults.ExecutionTargetDefault().TopP,
TimeoutSeconds: 45,
ServiceTier: "priority",
ReasoningEffort: "high",
APIKeyEnv: "BACKEND_KEY",
ExtraParams: map[string]any{
"profile": "value",
},
}
if inspection.ProfileID != "profile" || inspection.APIKeyRequired ||
!reflect.DeepEqual(inspection.EffectiveModelParams, wantTarget) {
t.Fatalf("inspection=%#v, want profile=%q target=%#v", inspection, "profile", wantTarget)
}
}
func TestRunnerInspectProfileDoesNotNeedExecutionCollaboratorsOrCredentials(t *testing.T) {
t.Setenv("PROMPTKIT_INSPECTION_TEST_KEY", "")
profiles := &inspectionProfileRepository{profile: &domain.ExecutionProfile{
ID: "endpoint-only",
Endpoint: "https://profile.example/v1",
Model: "profile-model",
APIKeyEnv: "PROMPTKIT_INSPECTION_TEST_KEY",
}}
runner := &Runner{profiles: profiles}
inspection, err := runner.InspectProfile(context.Background(), "endpoint-only")
if err != nil {
t.Fatalf("inspect endpoint-only profile: %v", err)
}
if inspection.EffectiveModelParams.BackendID != "" ||
inspection.EffectiveModelParams.APIKeyEnv != "PROMPTKIT_INSPECTION_TEST_KEY" ||
inspection.APIKeyRequired {
t.Fatalf("unexpected endpoint-only inspection: %#v", inspection)
}
}
func TestRunnerInspectProfileDirectCredentialRequirementClearsBackendEnvironment(t *testing.T) {
profiles := &inspectionProfileRepository{profile: &domain.ExecutionProfile{
ID: "direct-key",
BackendID: "backend",
Model: "profile-model",
APIKeyRequired: true,
}}
backends := &inspectionBackendResolver{backend: domain.Backend{
ID: "backend",
Endpoint: "https://backend.example/v1",
APIKeyEnv: "BACKEND_KEY",
}}
inspection, err := (&Runner{profiles: profiles, backends: backends}).InspectProfile(
context.Background(),
"direct-key",
)
if err != nil {
t.Fatalf("inspect direct-key profile: %v", err)
}
if !inspection.APIKeyRequired || inspection.EffectiveModelParams.APIKeyEnv != "" {
t.Fatalf("credential requirement was not resolved exclusively: %#v", inspection)
}
}
func TestRunnerInspectProfileClassifiesFailuresWithoutRepositoryWorkAfterCancellation(t *testing.T) {
t.Run("blank ID", func(t *testing.T) {
profiles := &inspectionProfileRepository{}
_, err := (&Runner{profiles: profiles}).InspectProfile(context.Background(), " \t ")
if !errors.Is(err, ErrInvalidRequest) || profiles.calls != 0 {
t.Fatalf("blank inspection=(%v, calls=%d), want invalid request without lookup", err, profiles.calls)
}
})
t.Run("canceled context", func(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
cancel()
profiles := &inspectionProfileRepository{}
_, err := (&Runner{profiles: profiles}).InspectProfile(ctx, "profile")
if !errors.Is(err, ErrProfileLoad) || !errors.Is(err, context.Canceled) || profiles.calls != 0 {
t.Fatalf("canceled inspection=(%v, calls=%d), want profile load and context identities without lookup", err, profiles.calls)
}
})
t.Run("missing profile", func(t *testing.T) {
profiles := &inspectionProfileRepository{err: profile.ErrProfileNotFound}
_, err := (&Runner{profiles: profiles}).InspectProfile(context.Background(), "missing")
if !errors.Is(err, ErrProfileLoad) || !errors.Is(err, profile.ErrProfileNotFound) {
t.Fatalf("missing profile error=%v, want profile load and not-found identities", err)
}
})
t.Run("unknown backend", func(t *testing.T) {
backendErr := errors.New("unknown backend")
profiles := &inspectionProfileRepository{profile: &domain.ExecutionProfile{
ID: "profile", BackendID: "backend", Model: "profile-model",
}}
backends := &inspectionBackendResolver{err: backendErr}
_, err := (&Runner{profiles: profiles, backends: backends}).InspectProfile(context.Background(), "profile")
if !errors.Is(err, ErrProfileLoad) || !errors.Is(err, backendErr) {
t.Fatalf("unknown backend error=%v, want profile load and backend identities", err)
}
})
t.Run("defensive invalid dependencies", func(t *testing.T) {
cases := []struct {
name string
runner *Runner
}{
{name: "nil repository", runner: &Runner{}},
{name: "nil profile", runner: &Runner{profiles: &inspectionProfileRepository{}}},
{name: "invalid target", runner: &Runner{profiles: &inspectionProfileRepository{
profile: &domain.ExecutionProfile{ID: "profile", Endpoint: "https://profile.example/v1"},
}}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
_, err := tc.runner.InspectProfile(context.Background(), "profile")
if !errors.Is(err, ErrProfileLoad) {
t.Fatalf("inspection error=%v, want profile load", err)
}
})
}
})
}

View File

@@ -0,0 +1,73 @@
package usecase
import (
"context"
"fmt"
"strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
)
type resolvedPromptDefinition struct {
definition *domain.PromptDefinition
hash string
}
func (r *Runner) resolvePromptDefinition(
ctx context.Context,
promptID string,
promptVersion string,
) (*resolvedPromptDefinition, error) {
if strings.TrimSpace(promptID) == "" {
return nil, fmt.Errorf("%w: prompt id is required", ErrInvalidRequest)
}
if r == nil || r.promptDefs == nil {
return nil, fmt.Errorf("%w: prompt repository is not configured", ErrPromptLoad)
}
definition, err := r.promptDefs.GetPromptDefinition(ctx, promptID, promptVersion)
if err != nil {
return nil, fmt.Errorf("%w: %w", ErrPromptLoad, err)
}
if definition == nil {
return nil, fmt.Errorf("%w: prompt repository returned nil definition", ErrPromptLoad)
}
hash, err := hashPromptDefinition(definition)
if err != nil {
return nil, fmt.Errorf("%w: failed to hash prompt definition: %v", ErrPromptLoad, err)
}
return &resolvedPromptDefinition{definition: definition, hash: hash}, nil
}
// InspectPrompt resolves one explicit prompt without execution work.
func (r *Runner) InspectPrompt(
ctx context.Context,
promptID string,
promptVersion string,
) (*domain.PromptInspection, error) {
if strings.TrimSpace(promptID) == "" {
return nil, fmt.Errorf("%w: prompt id is required", ErrInvalidRequest)
}
select {
case <-ctx.Done():
return nil, fmt.Errorf("%w: %w", ErrPromptLoad, ctx.Err())
default:
}
selection, err := r.resolvePromptDefinition(ctx, promptID, promptVersion)
if err != nil {
return nil, err
}
inputs := make([]domain.PromptInput, len(selection.definition.Inputs))
copy(inputs, selection.definition.Inputs)
return &domain.PromptInspection{
PromptID: selection.definition.ID,
PromptVersion: selection.definition.Version,
PromptHash: selection.hash,
DefaultProfileID: selection.definition.DefaultProfile,
Inputs: inputs,
OutputContract: selection.definition.Validation,
}, nil
}

View File

@@ -0,0 +1,157 @@
package usecase
import (
"context"
"errors"
"reflect"
"testing"
"gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/promptdef"
)
type inspectionPromptRepository struct {
definition *domain.PromptDefinition
err error
calls int
id string
version string
}
func (r *inspectionPromptRepository) GetPromptDefinition(
_ context.Context,
id string,
version string,
) (*domain.PromptDefinition, error) {
r.calls++
r.id = id
r.version = version
if r.err != nil {
return nil, r.err
}
return r.definition, nil
}
func TestRunnerInspectPromptResolvesOneDefinitionWithoutExecutionCollaborators(t *testing.T) {
definition := &domain.PromptDefinition{
ID: "normalized.prompt",
Version: "1.2.3",
DefaultProfile: "not-resolved",
Inputs: []domain.PromptInput{
{Name: "document", Required: true, ContentType: "text/plain", Description: "Source document."},
{Name: "audience", ContentType: "text/plain", Description: "Intended reader."},
},
Validation: domain.OutputContract{
Format: domain.FormatJSON,
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schemas/result.json",
RepairAttempts: 2,
},
}
repository := &inspectionPromptRepository{definition: definition}
runner := &Runner{promptDefs: repository}
inspection, err := runner.InspectPrompt(context.Background(), " prompt-id ", " version ")
if err != nil {
t.Fatalf("inspect prompt: %v", err)
}
wantHash, err := hashPromptDefinition(definition)
if err != nil {
t.Fatalf("hash prompt definition: %v", err)
}
if repository.calls != 1 || repository.id != " prompt-id " || repository.version != " version " {
t.Fatalf("prompt lookup=(calls=%d id=%q version=%q), want one unchanged lookup", repository.calls, repository.id, repository.version)
}
if inspection.PromptID != definition.ID ||
inspection.PromptVersion != definition.Version ||
inspection.PromptHash != wantHash ||
inspection.DefaultProfileID != definition.DefaultProfile ||
!reflect.DeepEqual(inspection.Inputs, definition.Inputs) ||
inspection.OutputContract != definition.Validation {
t.Fatalf("inspection=%#v, want definition metadata", inspection)
}
inspection.Inputs[0].Name = "changed"
second, err := runner.InspectPrompt(context.Background(), " prompt-id ", " version ")
if err != nil {
t.Fatalf("inspect prompt again: %v", err)
}
if definition.Inputs[0].Name != "document" || second.Inputs[0].Name != "document" {
t.Fatalf("inspection input mutation escaped caller result: definition=%#v next=%#v", definition.Inputs, second.Inputs)
}
}
func TestRunnerInspectPromptClassifiesFailuresWithoutRepositoryWorkAfterCancellation(t *testing.T) {
t.Run("blank ID", func(t *testing.T) {
repository := &inspectionPromptRepository{}
_, err := (&Runner{promptDefs: repository}).InspectPrompt(context.Background(), " \t ", "version")
if !errors.Is(err, ErrInvalidRequest) || repository.calls != 0 {
t.Fatalf("blank inspection=(%v, calls=%d), want invalid request without lookup", err, repository.calls)
}
})
t.Run("canceled context", func(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
cancel()
repository := &inspectionPromptRepository{}
_, err := (&Runner{promptDefs: repository}).InspectPrompt(ctx, "prompt", "version")
if !errors.Is(err, ErrPromptLoad) || !errors.Is(err, context.Canceled) || repository.calls != 0 {
t.Fatalf("canceled inspection=(%v, calls=%d), want prompt load and context identities without lookup", err, repository.calls)
}
})
t.Run("missing prompt", func(t *testing.T) {
repository := &inspectionPromptRepository{err: promptdef.ErrPromptDefinitionNotFound}
_, err := (&Runner{promptDefs: repository}).InspectPrompt(context.Background(), "missing", "version")
if !errors.Is(err, ErrPromptLoad) || !errors.Is(err, promptdef.ErrPromptDefinitionNotFound) {
t.Fatalf("missing prompt error=%v, want prompt load and not-found identities", err)
}
})
t.Run("defensive prompt dependencies", func(t *testing.T) {
var nilRunner *Runner
cases := []struct {
name string
runner *Runner
}{
{name: "nil runner", runner: nilRunner},
{name: "nil repository", runner: &Runner{}},
{name: "nil definition", runner: &Runner{promptDefs: &inspectionPromptRepository{}}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
_, err := tc.runner.InspectPrompt(context.Background(), "prompt", "version")
if !errors.Is(err, ErrPromptLoad) {
t.Fatalf("inspection error=%v, want prompt load", err)
}
})
}
})
}
func TestRunnerPrepareUsesThePromptInspectionSelectionAndHash(t *testing.T) {
definition := promptDef(domain.FormatMarkdown, domain.ValidationBasic, 0)
repository := &fakePromptRepo{def: definition}
runner := NewRunner(
repository,
&fakeExecutionProfileRepo{profiles: map[string]*domain.ExecutionProfile{"exec": defaultExecutionProfile()}},
nil,
&fakeArtifactReader{},
&fakeRenderer{rendered: &domain.RenderedPrompt{Messages: []domain.RenderedMessage{{Role: "user", Content: "hello"}}}},
nil,
nil,
nil,
)
inspection, err := runner.InspectPrompt(context.Background(), definition.ID, definition.Version)
if err != nil {
t.Fatalf("inspect prompt: %v", err)
}
prepared, err := runner.Prepare(context.Background(), domain.RunRequest{PromptID: definition.ID, PromptVersion: definition.Version, ProfileID: "exec"})
if err != nil {
t.Fatalf("prepare prompt: %v", err)
}
if inspection.PromptHash != prepared.PromptHash {
t.Fatalf("inspection hash=%q, preparation hash=%q", inspection.PromptHash, prepared.PromptHash)
}
}

View File

@@ -17,6 +17,7 @@ type OutputRepairer interface {
type RepairRequest struct { type RepairRequest struct {
PreviousOutput string PreviousOutput string
ValidationErrors []string ValidationErrors []string
SessionID string
Target domain.ExecutionTarget Target domain.ExecutionTarget
StructuredOutput *domain.StructuredOutputSpec StructuredOutput *domain.StructuredOutputSpec
Attempt int Attempt int
@@ -42,7 +43,9 @@ func (r *defaultOutputRepairer) Repair(ctx context.Context, req RepairRequest) (
errs = strings.Join(req.ValidationErrors, "\n") errs = strings.Join(req.ValidationErrors, "\n")
} }
prompt := domain.RenderedPrompt{Messages: []domain.RenderedMessage{ prompt := domain.RenderedPrompt{
SessionID: req.SessionID,
Messages: []domain.RenderedMessage{
{ {
Role: "system", Role: "system",
Content: "You repair invalid JSON output. Return only corrected JSON. Do not include explanations or markdown code fences.", Content: "You repair invalid JSON output. Return only corrected JSON. Do not include explanations or markdown code fences.",
@@ -58,7 +61,8 @@ func (r *defaultOutputRepairer) Repair(ctx context.Context, req RepairRequest) (
req.PreviousOutput, req.PreviousOutput,
), ),
}, },
}} },
}
resp, err := r.llm.Generate(ctx, domain.GenerateRequest{ resp, err := r.llm.Generate(ctx, domain.GenerateRequest{
Prompt: prompt, Prompt: prompt,

View File

@@ -14,6 +14,7 @@ import (
"unicode" "unicode"
"gitea.maximumdirect.net/eric/promptkit/internal/artifact" "gitea.maximumdirect.net/eric/promptkit/internal/artifact"
"gitea.maximumdirect.net/eric/promptkit/internal/capacity"
"gitea.maximumdirect.net/eric/promptkit/internal/defaults" "gitea.maximumdirect.net/eric/promptkit/internal/defaults"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/llm" "gitea.maximumdirect.net/eric/promptkit/internal/llm"
@@ -40,41 +41,80 @@ var (
type Runner struct { type Runner struct {
promptDefs promptdef.Repository promptDefs promptdef.Repository
profiles profile.Repository profiles profile.Repository
backends BackendResolver
artifacts artifact.Reader artifacts artifact.Reader
renderer prompt.Renderer renderer prompt.Renderer
llm llm.Client llm llm.Client
validator validate.Validator validator validate.Validator
repairer OutputRepairer repairer OutputRepairer
admitter RunAdmitter
}
// BackendResolver resolves one normalized backend ID.
type BackendResolver interface {
GetBackend(string) (domain.Backend, error)
}
// RunAdmitter reserves capacity for one resolved backend run.
type RunAdmitter interface {
Admit(context.Context, string) (func(), error)
}
type preparationState struct {
definition *domain.PromptDefinition
directSessionID string
promptDefinitionHash string
selectedProfileID string
effectiveModel domain.ExecutionTarget
targetPresence domain.ExecutionTargetPresence
effectiveContract domain.OutputContract
start time.Time
} }
func NewRunner( func NewRunner(
promptDefs promptdef.Repository, promptDefs promptdef.Repository,
profiles profile.Repository, profiles profile.Repository,
backends BackendResolver,
artifacts artifact.Reader, artifacts artifact.Reader,
renderer prompt.Renderer, renderer prompt.Renderer,
llmClient llm.Client, llmClient llm.Client,
validator validate.Validator, validator validate.Validator,
admitter RunAdmitter,
) *Runner { ) *Runner {
return NewRunnerWithRepairer(promptDefs, profiles, artifacts, renderer, llmClient, validator, nil) return NewRunnerWithRepairer(
promptDefs,
profiles,
backends,
artifacts,
renderer,
llmClient,
validator,
nil,
admitter,
)
} }
func NewRunnerWithRepairer( func NewRunnerWithRepairer(
promptDefs promptdef.Repository, promptDefs promptdef.Repository,
profiles profile.Repository, profiles profile.Repository,
backends BackendResolver,
artifacts artifact.Reader, artifacts artifact.Reader,
renderer prompt.Renderer, renderer prompt.Renderer,
llmClient llm.Client, llmClient llm.Client,
validator validate.Validator, validator validate.Validator,
repairer OutputRepairer, repairer OutputRepairer,
admitter RunAdmitter,
) *Runner { ) *Runner {
return &Runner{ return &Runner{
promptDefs: promptDefs, promptDefs: promptDefs,
profiles: profiles, profiles: profiles,
backends: backends,
artifacts: artifacts, artifacts: artifacts,
renderer: renderer, renderer: renderer,
llm: llmClient, llm: llmClient,
validator: validator, validator: validator,
repairer: repairer, repairer: repairer,
admitter: admitter,
} }
} }
@@ -86,14 +126,51 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
start := time.Now().UTC() start := time.Now().UTC()
prepared, err := r.Prepare(ctx, req) state, err := r.resolvePreparation(ctx, req, time.Now().UTC())
if err != nil { if err != nil {
return nil, err return nil, err
} }
release, err := r.admitRun(ctx, state.effectiveModel.BackendID)
if err != nil {
return nil, err
}
defer release()
prepared, err := r.completePreparation(ctx, req, state)
if err != nil {
return nil, err
}
directAPIKey := state.effectiveModel.APIKey
return r.executePreparedRun(ctx, prepared, directAPIKey, runID, start, func(
ctx context.Context,
artifact *domain.Artifact,
attemptsUsed int,
) (domain.ValidationResult, error) {
return r.validateOutput(ctx, artifact, prepared.OutputContract, attemptsUsed)
})
}
type preparedValidationFunc func(
context.Context,
*domain.Artifact,
int,
) (domain.ValidationResult, error)
func (r *Runner) executePreparedRun(
ctx context.Context,
prepared *domain.PreparedRun,
directAPIKey string,
runID string,
start time.Time,
validateArtifact preparedValidationFunc,
) (*domain.RunResult, error) {
executionTarget := prepared.EffectiveModelParams
executionTarget.APIKey = directAPIKey
genResp, err := r.llm.Generate(ctx, domain.GenerateRequest{ genResp, err := r.llm.Generate(ctx, domain.GenerateRequest{
Prompt: domain.RenderedPrompt{SessionID: prepared.SessionID, Messages: prepared.Messages}, Prompt: domain.RenderedPrompt{SessionID: prepared.SessionID, Messages: prepared.Messages},
Target: prepared.EffectiveModelParams, Target: executionTarget,
TargetPresence: prepared.TargetPresence, TargetPresence: prepared.TargetPresence,
StructuredOutput: prepared.StructuredOutput, StructuredOutput: prepared.StructuredOutput,
}) })
@@ -105,7 +182,7 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
} }
outputArtifact := buildOutputArtifact(genResp.Content, prepared.OutputContract.Format) outputArtifact := buildOutputArtifact(genResp.Content, prepared.OutputContract.Format)
validationResult, err := r.validateOutput(ctx, &outputArtifact, prepared.OutputContract, 0) validationResult, err := validateArtifact(ctx, &outputArtifact, 0)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %w", ErrValidation, err) return nil, fmt.Errorf("%w: %w", ErrValidation, err)
} }
@@ -118,7 +195,8 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
repairResp, repairErr := r.repairer.Repair(ctx, RepairRequest{ repairResp, repairErr := r.repairer.Repair(ctx, RepairRequest{
PreviousOutput: genResp.Content, PreviousOutput: genResp.Content,
ValidationErrors: validationResult.Errors, ValidationErrors: validationResult.Errors,
Target: prepared.EffectiveModelParams, SessionID: prepared.SessionID,
Target: executionTarget,
StructuredOutput: prepared.StructuredOutput, StructuredOutput: prepared.StructuredOutput,
Attempt: attemptsUsed, Attempt: attemptsUsed,
MaxAttempts: prepared.OutputContract.RepairAttempts, MaxAttempts: prepared.OutputContract.RepairAttempts,
@@ -134,7 +212,7 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
genResp = repairResp genResp = repairResp
outputArtifact = buildOutputArtifact(genResp.Content, prepared.OutputContract.Format) outputArtifact = buildOutputArtifact(genResp.Content, prepared.OutputContract.Format)
validationResult, err = r.validateOutput(ctx, &outputArtifact, prepared.OutputContract, attemptsUsed) validationResult, err = validateArtifact(ctx, &outputArtifact, attemptsUsed)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %w", ErrValidation, err) return nil, fmt.Errorf("%w: %w", ErrValidation, err)
} }
@@ -142,6 +220,7 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
} }
end := time.Now().UTC() end := time.Now().UTC()
executionTarget.APIKey = ""
return &domain.RunResult{ return &domain.RunResult{
RunID: runID, RunID: runID,
@@ -151,11 +230,13 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
PromptID: prepared.PromptID, PromptID: prepared.PromptID,
PromptVersion: prepared.PromptVersion, PromptVersion: prepared.PromptVersion,
PromptHash: prepared.PromptHash, PromptHash: prepared.PromptHash,
SessionID: prepared.SessionID,
RenderedPromptHash: prepared.RenderedPromptHash, RenderedPromptHash: prepared.RenderedPromptHash,
SelectedProfileID: prepared.SelectedProfileID, SelectedProfileID: prepared.SelectedProfileID,
SelectedBackendID: prepared.SelectedBackendID,
ModelName: prepared.EffectiveModelParams.Model, ModelName: prepared.EffectiveModelParams.Model,
Endpoint: prepared.EffectiveModelParams.Endpoint, Endpoint: prepared.EffectiveModelParams.Endpoint,
EffectiveModelParams: prepared.EffectiveModelParams, EffectiveModelParams: executionTarget,
InputHashes: prepared.InputHashes, InputHashes: prepared.InputHashes,
Usage: genResp.Usage, Usage: genResp.Usage,
StartTime: start, StartTime: start,
@@ -165,20 +246,32 @@ func (r *Runner) Run(ctx context.Context, req domain.RunRequest) (*domain.RunRes
} }
func (r *Runner) Prepare(ctx context.Context, req domain.RunRequest) (*domain.PreparedRun, error) { func (r *Runner) Prepare(ctx context.Context, req domain.RunRequest) (*domain.PreparedRun, error) {
state, err := r.resolvePreparation(ctx, req, time.Now().UTC())
if err != nil {
return nil, err
}
return r.completePreparation(ctx, req, state)
}
func (r *Runner) resolvePreparation(
ctx context.Context,
req domain.RunRequest,
start time.Time,
) (*preparationState, error) {
if strings.TrimSpace(req.PromptID) == "" { if strings.TrimSpace(req.PromptID) == "" {
return nil, fmt.Errorf("%w: prompt id is required", ErrInvalidRequest) return nil, fmt.Errorf("%w: prompt id is required", ErrInvalidRequest)
} }
directSessionID, err := domain.NormalizeSessionID(req.SessionID)
start := time.Now().UTC()
def, err := r.promptDefs.GetPromptDefinition(ctx, req.PromptID, req.PromptVersion)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %w", ErrPromptLoad, err) return nil, fmt.Errorf("%w: session_id: %v", ErrInvalidRequest, err)
} }
promptDefinitionHash, err := hashPromptDefinition(def)
promptSelection, err := r.resolvePromptDefinition(ctx, req.PromptID, req.PromptVersion)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: failed to hash prompt definition: %v", ErrPromptLoad, err) return nil, err
} }
def := promptSelection.definition
promptDefinitionHash := promptSelection.hash
selectedProfileID := strings.TrimSpace(req.ProfileID) selectedProfileID := strings.TrimSpace(req.ProfileID)
if selectedProfileID == "" { if selectedProfileID == "" {
@@ -188,32 +281,58 @@ func (r *Runner) Prepare(ctx context.Context, req domain.RunRequest) (*domain.Pr
return nil, fmt.Errorf("%w: %w: profile id is required either in request or prompt default_profile", ErrInvalidRequest, ErrProfileRequired) return nil, fmt.Errorf("%w: %w: profile id is required either in request or prompt default_profile", ErrInvalidRequest, ErrProfileRequired)
} }
execProfile, err := r.profiles.GetProfile(ctx, selectedProfileID) selection, err := r.resolveProfileSelection(ctx, selectedProfileID)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %w", ErrProfileLoad, err) return nil, err
} }
effectiveModel, targetPresence, err := resolveExecutionTarget(execProfile, req.Execution) effectiveModel, targetPresence, err := resolveExecutionTarget(selection.backend, selection.profile, req.Execution)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err) return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
} }
effectiveModel.APIKey = req.APIKey effectiveModel.APIKey = req.APIKey
if strings.TrimSpace(effectiveModel.Endpoint) == "" { if err := validateResolvedExecutionTarget(effectiveModel); err != nil {
return nil, fmt.Errorf("%w: execution endpoint is required", ErrInvalidRequest) return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
}
if strings.TrimSpace(effectiveModel.Model) == "" {
return nil, fmt.Errorf("%w: execution model is required", ErrInvalidRequest)
} }
if err := validateAPIKey(effectiveModel.APIKeyEnv, effectiveModel.APIKey, effectiveModel.APIKeyRequired); err != nil { if err := validateAPIKey(effectiveModel.APIKeyEnv, effectiveModel.APIKey, effectiveModel.APIKeyRequired); err != nil {
return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err) return nil, fmt.Errorf("%w: %w", ErrInvalidRequest, err)
} }
effectiveContract := resolveOutputContract(def, req.Validation) effectiveContract := resolveOutputContract(def, req.Validation)
structuredOutput, err := r.resolveStructuredOutput(ctx, def, effectiveContract) return &preparationState{
definition: def,
directSessionID: directSessionID,
promptDefinitionHash: promptDefinitionHash,
selectedProfileID: selection.id,
effectiveModel: effectiveModel,
targetPresence: targetPresence,
effectiveContract: effectiveContract,
start: start,
}, nil
}
func (r *Runner) completePreparation(
ctx context.Context,
req domain.RunRequest,
state *preparationState,
) (*domain.PreparedRun, error) {
structuredOutput, err := r.resolveStructuredOutput(
ctx,
state.definition,
state.effectiveContract,
)
if err != nil { if err != nil {
return nil, err return nil, err
} }
return r.completePreparationWithStructuredOutput(ctx, req, state, structuredOutput)
}
func (r *Runner) completePreparationWithStructuredOutput(
ctx context.Context,
req domain.RunRequest,
state *preparationState,
structuredOutput *domain.StructuredOutputSpec,
) (*domain.PreparedRun, error) {
resolvedInputs := make(map[string]*domain.Artifact, len(req.Inputs)) resolvedInputs := make(map[string]*domain.Artifact, len(req.Inputs))
inputHashes := make(map[string]string, len(req.Inputs)) inputHashes := make(map[string]string, len(req.Inputs))
for name, ref := range req.Inputs { for name, ref := range req.Inputs {
@@ -228,31 +347,57 @@ func (r *Runner) Prepare(ctx context.Context, req domain.RunRequest) (*domain.Pr
inputHashes[name] = art.Hash inputHashes[name] = art.Hash
} }
renderedPrompt, err := r.renderer.Render(ctx, def, resolvedInputs, req.Vars) definitionToRender := state.definition
if state.directSessionID != "" {
definitionCopy := *state.definition
definitionCopy.SessionID = ""
definitionToRender = &definitionCopy
}
renderedPrompt, err := r.renderer.Render(ctx, definitionToRender, resolvedInputs, req.Vars)
if err != nil { if err != nil {
return nil, fmt.Errorf("%w: %w", ErrPromptRender, err) return nil, fmt.Errorf("%w: %w", ErrPromptRender, err)
} }
if state.directSessionID != "" {
renderedPrompt.SessionID = state.directSessionID
}
end := time.Now().UTC() end := time.Now().UTC()
effectiveModel := state.effectiveModel
effectiveModel.APIKey = ""
return &domain.PreparedRun{ return &domain.PreparedRun{
PromptID: def.ID, PromptID: state.definition.ID,
PromptVersion: def.Version, PromptVersion: state.definition.Version,
PromptHash: promptDefinitionHash, PromptHash: state.promptDefinitionHash,
SelectedProfileID: selectedProfileID, SelectedProfileID: state.selectedProfileID,
SelectedBackendID: state.effectiveModel.BackendID,
EffectiveModelParams: effectiveModel, EffectiveModelParams: effectiveModel,
TargetPresence: targetPresence, TargetPresence: state.targetPresence,
OutputContract: effectiveContract, OutputContract: state.effectiveContract,
StructuredOutput: structuredOutput, StructuredOutput: structuredOutput,
InputHashes: inputHashes, InputHashes: inputHashes,
SessionID: renderedPrompt.SessionID, SessionID: renderedPrompt.SessionID,
RenderedPromptHash: hashRenderedPrompt(*renderedPrompt), RenderedPromptHash: hashRenderedPrompt(*renderedPrompt),
Messages: renderedPrompt.Messages, Messages: renderedPrompt.Messages,
StartTime: start, StartTime: state.start,
EndTime: end, EndTime: end,
DurationMS: end.Sub(start).Milliseconds(), DurationMS: end.Sub(state.start).Milliseconds(),
}, nil }, nil
} }
func (r *Runner) admitRun(ctx context.Context, backendID string) (func(), error) {
if r.admitter == nil {
return func() {}, nil
}
release, err := r.admitter.Admit(ctx, backendID)
if err != nil {
if errors.Is(err, capacity.ErrCapacityExceeded) && strings.TrimSpace(backendID) != "" {
return nil, &CapacityError{BackendID: backendID}
}
return nil, err
}
return release, nil
}
func (r *Runner) resolveStructuredOutput(ctx context.Context, def *domain.PromptDefinition, contract domain.OutputContract) (*domain.StructuredOutputSpec, error) { func (r *Runner) resolveStructuredOutput(ctx context.Context, def *domain.PromptDefinition, contract domain.OutputContract) (*domain.StructuredOutputSpec, error) {
if contract.ValidationMode != domain.ValidationJSONSchema { if contract.ValidationMode != domain.ValidationJSONSchema {
return nil, nil return nil, nil
@@ -268,14 +413,18 @@ func (r *Runner) resolveStructuredOutput(ctx context.Context, def *domain.Prompt
return nil, fmt.Errorf("%w: failed to load json schema for structured output: %v", ErrValidation, err) return nil, fmt.Errorf("%w: failed to load json schema for structured output: %v", ErrValidation, err)
} }
return structuredOutputSpec(def, schemaDoc), nil
}
func structuredOutputSpec(def *domain.PromptDefinition, schemaDocument any) *domain.StructuredOutputSpec {
return &domain.StructuredOutputSpec{ return &domain.StructuredOutputSpec{
Type: domain.StructuredOutputJSONSchema, Type: domain.StructuredOutputJSONSchema,
JSONSchema: &domain.StructuredOutputJSONSpec{ JSONSchema: &domain.StructuredOutputJSONSpec{
Name: deriveStructuredSchemaName(def.ID, def.Version), Name: deriveStructuredSchemaName(def.ID, def.Version),
Strict: true, Strict: true,
Schema: schemaDoc, Schema: schemaDocument,
}, },
}, nil }
} }
func deriveStructuredSchemaName(promptID string, promptVersion string) string { func deriveStructuredSchemaName(promptID string, promptVersion string) string {
@@ -338,7 +487,10 @@ func (r *Runner) shouldAttemptRepair(contract domain.OutputContract, validationR
func mergeExecutionTarget(base domain.ExecutionTarget, override domain.ExecutionTarget) domain.ExecutionTarget { func mergeExecutionTarget(base domain.ExecutionTarget, override domain.ExecutionTarget) domain.ExecutionTarget {
out := base out := base
if override.Endpoint != "" { if strings.TrimSpace(override.BackendID) != "" {
out.BackendID = override.BackendID
}
if strings.TrimSpace(override.Endpoint) != "" {
out.Endpoint = override.Endpoint out.Endpoint = override.Endpoint
} }
if override.Model != "" { if override.Model != "" {
@@ -367,6 +519,7 @@ func mergeExecutionTarget(base domain.ExecutionTarget, override domain.Execution
} }
if override.APIKeyRequired { if override.APIKeyRequired {
out.APIKeyRequired = true out.APIKeyRequired = true
out.APIKeyEnv = ""
} }
if len(override.ExtraParams) > 0 { if len(override.ExtraParams) > 0 {
out.ExtraParams = copyExtraParams(override.ExtraParams) out.ExtraParams = copyExtraParams(override.ExtraParams)
@@ -414,8 +567,8 @@ func mergeExecutionTargetOverride(base domain.ExecutionTarget, override domain.E
if strings.TrimSpace(override.ServiceTier) != "" { if strings.TrimSpace(override.ServiceTier) != "" {
out.ServiceTier = override.ServiceTier out.ServiceTier = override.ServiceTier
} }
if strings.TrimSpace(override.ReasoningEffort) != "" { if override.ReasoningEffort != nil {
out.ReasoningEffort = override.ReasoningEffort out.ReasoningEffort = strings.TrimSpace(*override.ReasoningEffort)
} }
if strings.TrimSpace(override.APIKeyEnv) != "" { if strings.TrimSpace(override.APIKeyEnv) != "" {
out.APIKeyEnv = override.APIKeyEnv out.APIKeyEnv = override.APIKeyEnv
@@ -426,8 +579,9 @@ func mergeExecutionTargetOverride(base domain.ExecutionTarget, override domain.E
return out, presence, nil return out, presence, nil
} }
func resolveExecutionTarget(profileValue *domain.ExecutionProfile, override *domain.ExecutionTargetOverride) (domain.ExecutionTarget, domain.ExecutionTargetPresence, error) { func resolveExecutionTarget(backendValue *domain.Backend, profileValue *domain.ExecutionProfile, override *domain.ExecutionTargetOverride) (domain.ExecutionTarget, domain.ExecutionTargetPresence, error) {
out := defaults.ExecutionTargetDefault() out := defaults.ExecutionTargetDefault()
out = mergeExecutionTarget(out, backendToTarget(backendValue))
out = mergeExecutionTarget(out, executionProfileToTarget(profileValue)) out = mergeExecutionTarget(out, executionProfileToTarget(profileValue))
var presence domain.ExecutionTargetPresence var presence domain.ExecutionTargetPresence
if override != nil { if override != nil {
@@ -461,8 +615,13 @@ func executionProfileToTarget(p *domain.ExecutionProfile) domain.ExecutionTarget
if p == nil { if p == nil {
return domain.ExecutionTarget{} return domain.ExecutionTarget{}
} }
endpoint := p.Endpoint
if strings.TrimSpace(endpoint) == "" {
endpoint = ""
}
return domain.ExecutionTarget{ return domain.ExecutionTarget{
Endpoint: p.Endpoint, BackendID: p.BackendID,
Endpoint: endpoint,
Model: p.Model, Model: p.Model,
Temperature: p.Temperature, Temperature: p.Temperature,
MaxTokens: p.MaxTokens, MaxTokens: p.MaxTokens,
@@ -476,6 +635,18 @@ func executionProfileToTarget(p *domain.ExecutionProfile) domain.ExecutionTarget
} }
} }
func backendToTarget(value *domain.Backend) domain.ExecutionTarget {
if value == nil {
return domain.ExecutionTarget{}
}
return domain.ExecutionTarget{
BackendID: value.ID,
Endpoint: value.Endpoint,
APIKeyEnv: value.APIKeyEnv,
ExtraParams: copyExtraParams(value.ExtraParams),
}
}
func copyExtraParams(src map[string]any) map[string]any { func copyExtraParams(src map[string]any) map[string]any {
if len(src) == 0 { if len(src) == 0 {
return nil return nil

File diff suppressed because it is too large Load Diff

View File

@@ -45,6 +45,101 @@ func (v *FSValidator) Validate(ctx context.Context, artifact *domain.Artifact, c
return validateArtifact(ctx, artifact, contract, v.validateJSONSchema) return validateArtifact(ctx, artifact, contract, v.validateJSONSchema)
} }
type preparedValidation struct {
contract domain.OutputContract
schemaDocument any
schema *jsonschema.Schema
}
func (p *preparedValidation) Validate(ctx context.Context, artifact *domain.Artifact) (domain.ValidationResult, error) {
return validateArtifact(ctx, artifact, p.contract, p.validateJSONSchema)
}
func (p *preparedValidation) SchemaDocument() any {
return p.schemaDocument
}
func (p *preparedValidation) validateJSONSchema(instance any, _ string) ([]string, error) {
if p.schema == nil {
return nil, errors.New("prepared JSON schema is unavailable")
}
if err := p.schema.Validate(instance); err != nil {
return []string{fmt.Sprintf("json schema validation failed: %v", err)}, nil
}
return nil, nil
}
func (v *StandardValidator) PrepareValidation(ctx context.Context, contract domain.OutputContract) (PreparedValidation, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
prepared := &preparedValidation{contract: contract}
if contract.ValidationMode != domain.ValidationJSONSchema {
return prepared, nil
}
resolvedSchemaPath, err := v.resolveSchemaPath(contract.SchemaPath)
if err != nil {
return nil, err
}
schemaDocument, err := loadJSONSchemaFile(resolvedSchemaPath)
if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err)
}
schemaRoot, err := v.schemaRoot()
if err != nil {
return nil, err
}
compiler := newSchemaCompiler(standardSchemaLoader{root: schemaRoot})
if err := compiler.AddResource(resolvedSchemaPath, schemaDocument); err != nil {
return nil, fmt.Errorf("failed to register JSON schema %q: %w", resolvedSchemaPath, err)
}
schema, err := compiler.Compile(resolvedSchemaPath)
if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", resolvedSchemaPath, err)
}
if err := ctx.Err(); err != nil {
return nil, err
}
prepared.schemaDocument = schemaDocument
prepared.schema = schema
return prepared, nil
}
func (v *FSValidator) PrepareValidation(ctx context.Context, contract domain.OutputContract) (PreparedValidation, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
prepared := &preparedValidation{contract: contract}
if contract.ValidationMode != domain.ValidationJSONSchema {
return prepared, nil
}
schemaName, schemaDocument, err := v.loadSchemaDocument(contract.SchemaPath)
if err != nil {
return nil, err
}
resourceURL := fsSchemaResourceURL(schemaName)
compiler := newSchemaCompiler(fsSchemaLoader{fsys: v.fsys, root: filecatalog.CleanFSRoot(v.root)})
if err := compiler.AddResource(resourceURL, schemaDocument); err != nil {
return nil, fmt.Errorf("failed to register JSON schema %q: %w", schemaName, err)
}
schema, err := compiler.Compile(resourceURL)
if err != nil {
return nil, fmt.Errorf("failed to compile JSON schema %q: %w", schemaName, err)
}
if err := ctx.Err(); err != nil {
return nil, err
}
prepared.schemaDocument = schemaDocument
prepared.schema = schema
return prepared, nil
}
type schemaValidatorFunc func(instance any, schemaPath string) ([]string, error) type schemaValidatorFunc func(instance any, schemaPath string) ([]string, error)
func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract, validateSchema schemaValidatorFunc) (domain.ValidationResult, error) { func validateArtifact(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract, validateSchema schemaValidatorFunc) (domain.ValidationResult, error) {

View File

@@ -5,6 +5,7 @@ import (
"encoding/json" "encoding/json"
"os" "os"
"path/filepath" "path/filepath"
"reflect"
"strconv" "strconv"
"strings" "strings"
"testing" "testing"
@@ -120,6 +121,68 @@ func TestStandardValidatorJSONSchemaSuccess(t *testing.T) {
} }
} }
func TestStandardValidatorPreparedSchemaSurvivesSourceRemoval(t *testing.T) {
tmp := t.TempDir()
rootSchema := []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "original root",
"type": "object",
"required": ["value"],
"properties": {
"value": {"$ref": "value.json"}
}
}`)
rootPath := filepath.Join(tmp, "schema.json")
referencePath := filepath.Join(tmp, "value.json")
if err := os.WriteFile(rootPath, rootSchema, 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(referencePath, []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "integer",
"minimum": 2
}`), 0o644); err != nil {
t.Fatal(err)
}
validator := NewStandardValidator(tmp)
preparer, ok := validator.(ValidationPreparer)
if !ok {
t.Fatal("standard validator does not support validation preparation")
}
prepared, err := preparer.PrepareValidation(context.Background(), domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
if err != nil {
t.Fatalf("prepare validation: %v", err)
}
assertSchemaDocument(t, prepared.SchemaDocument(), rootSchema)
if err := os.Remove(rootPath); err != nil {
t.Fatal(err)
}
if err := os.Remove(referencePath); err != nil {
t.Fatal(err)
}
valid, err := prepared.Validate(context.Background(), &domain.Artifact{Body: []byte(`{"value":3}`)})
if err != nil {
t.Fatalf("validate prepared artifact: %v", err)
}
if valid.Status != domain.ValidationPassed || !valid.IsValid {
t.Fatalf("expected passed/valid, got status=%q valid=%v errors=%v", valid.Status, valid.IsValid, valid.Errors)
}
invalid, err := prepared.Validate(context.Background(), &domain.Artifact{Body: []byte(`{"value":"changed"}`)})
if err != nil {
t.Fatalf("validate prepared artifact: %v", err)
}
if invalid.Status != domain.ValidationFailed || invalid.IsValid {
t.Fatalf("expected failed/invalid, got status=%q valid=%v", invalid.Status, invalid.IsValid)
}
}
func TestStandardValidatorJSONSchemaNestedSchemaPathSuccess(t *testing.T) { func TestStandardValidatorJSONSchemaNestedSchemaPathSuccess(t *testing.T) {
tmp := t.TempDir() tmp := t.TempDir()
nestedDir := filepath.Join(tmp, "dnd") nestedDir := filepath.Join(tmp, "dnd")
@@ -295,6 +358,64 @@ func TestFSValidatorJSONSchemaSuccess(t *testing.T) {
} }
} }
func TestFSValidatorPreparedSchemaSurvivesSourceMutation(t *testing.T) {
rootSchema := []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "original root",
"type": "object",
"required": ["value"],
"properties": {
"value": {"$ref": "value.json"}
}
}`)
fsys := fstest.MapFS{
"schema.json": &fstest.MapFile{Data: rootSchema},
"value.json": &fstest.MapFile{Data: []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "integer",
"minimum": 2
}`)},
}
validator := NewFSValidator(fsys, ".")
preparer, ok := validator.(ValidationPreparer)
if !ok {
t.Fatal("filesystem validator does not support validation preparation")
}
prepared, err := preparer.PrepareValidation(context.Background(), domain.OutputContract{
ValidationMode: domain.ValidationJSONSchema,
SchemaPath: "schema.json",
})
if err != nil {
t.Fatalf("prepare validation: %v", err)
}
assertSchemaDocument(t, prepared.SchemaDocument(), rootSchema)
fsys["schema.json"] = &fstest.MapFile{Data: []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "string"
}`)}
fsys["value.json"] = &fstest.MapFile{Data: []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "string"
}`)}
valid, err := prepared.Validate(context.Background(), &domain.Artifact{Body: []byte(`{"value":3}`)})
if err != nil {
t.Fatalf("validate prepared artifact: %v", err)
}
if valid.Status != domain.ValidationPassed || !valid.IsValid {
t.Fatalf("expected passed/valid, got status=%q valid=%v errors=%v", valid.Status, valid.IsValid, valid.Errors)
}
invalid, err := prepared.Validate(context.Background(), &domain.Artifact{Body: []byte(`{"value":"changed"}`)})
if err != nil {
t.Fatalf("validate prepared artifact: %v", err)
}
if invalid.Status != domain.ValidationFailed || invalid.IsValid {
t.Fatalf("expected failed/invalid, got status=%q valid=%v", invalid.Status, invalid.IsValid)
}
}
func TestFSValidatorJSONSchemaRegistrationError(t *testing.T) { func TestFSValidatorJSONSchemaRegistrationError(t *testing.T) {
v := NewFSValidator(fstest.MapFS{ v := NewFSValidator(fstest.MapFS{
"schemas/%zz.json": &fstest.MapFile{Data: []byte(`{"type":"object"}`)}, "schemas/%zz.json": &fstest.MapFile{Data: []byte(`{"type":"object"}`)},
@@ -548,3 +669,15 @@ func TestJSONSchemaDialectIsDraft2020(t *testing.T) {
}) })
} }
} }
func assertSchemaDocument(t *testing.T, got any, expectedJSON []byte) {
t.Helper()
var expected any
if err := json.Unmarshal(expectedJSON, &expected); err != nil {
t.Fatalf("decode expected schema document: %v", err)
}
if !reflect.DeepEqual(got, expected) {
t.Fatalf("schema document mismatch:\n got: %#v\nwant: %#v", got, expected)
}
}

View File

@@ -2,6 +2,7 @@ package validate
import ( import (
"context" "context"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
) )
@@ -10,6 +11,20 @@ type Validator interface {
Validate(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract) (domain.ValidationResult, error) Validate(ctx context.Context, artifact *domain.Artifact, contract domain.OutputContract) (domain.ValidationResult, error)
} }
// PreparedValidation validates artifacts against one frozen output contract.
type PreparedValidation interface {
Validate(ctx context.Context, artifact *domain.Artifact) (domain.ValidationResult, error)
// SchemaDocument returns the root JSON Schema document used for provider
// structured output, or nil for non-schema modes. Returned internal
// immutable state must not be mutated.
SchemaDocument() any
}
// ValidationPreparer freezes validation resources for one output contract.
type ValidationPreparer interface {
PrepareValidation(ctx context.Context, contract domain.OutputContract) (PreparedValidation, error)
}
// SchemaDocumentLoader loads JSON schema documents using validator path semantics. // SchemaDocumentLoader loads JSON schema documents using validator path semantics.
type SchemaDocumentLoader interface { type SchemaDocumentLoader interface {
LoadSchemaDocument(ctx context.Context, schemaPath string) (any, error) LoadSchemaDocument(ctx context.Context, schemaPath string) (any, error)

View File

@@ -26,6 +26,7 @@ func (r PreparedRun) MarshalJSON() ([]byte, error) {
PromptVersion string `json:"prompt_version,omitempty"` PromptVersion string `json:"prompt_version,omitempty"`
PromptHash string `json:"prompt_hash,omitempty"` PromptHash string `json:"prompt_hash,omitempty"`
SelectedProfileID string `json:"selected_profile_id"` SelectedProfileID string `json:"selected_profile_id"`
SelectedBackendID string `json:"selected_backend_id,omitempty"`
EffectiveModelParams ExecutionTarget `json:"effective_model_params"` EffectiveModelParams ExecutionTarget `json:"effective_model_params"`
OutputContract OutputContract `json:"output_contract"` OutputContract OutputContract `json:"output_contract"`
StructuredOutput *StructuredOutputSpec `json:"structured_output,omitempty"` StructuredOutput *StructuredOutputSpec `json:"structured_output,omitempty"`
@@ -41,6 +42,7 @@ func (r PreparedRun) MarshalJSON() ([]byte, error) {
PromptVersion: r.PromptVersion, PromptVersion: r.PromptVersion,
PromptHash: r.PromptHash, PromptHash: r.PromptHash,
SelectedProfileID: r.SelectedProfileID, SelectedProfileID: r.SelectedProfileID,
SelectedBackendID: r.SelectedBackendID,
EffectiveModelParams: r.EffectiveModelParams, EffectiveModelParams: r.EffectiveModelParams,
OutputContract: r.OutputContract, OutputContract: r.OutputContract,
StructuredOutput: r.StructuredOutput, StructuredOutput: r.StructuredOutput,
@@ -79,8 +81,10 @@ func (r RunResult) MarshalJSON() ([]byte, error) {
PromptID: r.PromptID, PromptID: r.PromptID,
PromptVersion: r.PromptVersion, PromptVersion: r.PromptVersion,
PromptHash: r.PromptHash, PromptHash: r.PromptHash,
SessionID: r.SessionID,
RenderedPromptHash: r.RenderedPromptHash, RenderedPromptHash: r.RenderedPromptHash,
SelectedProfileID: r.SelectedProfileID, SelectedProfileID: r.SelectedProfileID,
SelectedBackendID: r.SelectedBackendID,
ModelName: r.ModelName, ModelName: r.ModelName,
Endpoint: r.Endpoint, Endpoint: r.Endpoint,
EffectiveModelParams: r.EffectiveModelParams, EffectiveModelParams: r.EffectiveModelParams,
@@ -108,8 +112,10 @@ func (r *RunResult) UnmarshalJSON(data []byte) error {
PromptID: wire.PromptID, PromptID: wire.PromptID,
PromptVersion: wire.PromptVersion, PromptVersion: wire.PromptVersion,
PromptHash: wire.PromptHash, PromptHash: wire.PromptHash,
SessionID: wire.SessionID,
RenderedPromptHash: wire.RenderedPromptHash, RenderedPromptHash: wire.RenderedPromptHash,
SelectedProfileID: wire.SelectedProfileID, SelectedProfileID: wire.SelectedProfileID,
SelectedBackendID: wire.SelectedBackendID,
ModelName: wire.ModelName, ModelName: wire.ModelName,
Endpoint: wire.Endpoint, Endpoint: wire.Endpoint,
EffectiveModelParams: wire.EffectiveModelParams, EffectiveModelParams: wire.EffectiveModelParams,
@@ -134,8 +140,10 @@ type runResultJSON struct {
PromptID string `json:"prompt_id"` PromptID string `json:"prompt_id"`
PromptVersion string `json:"prompt_version,omitempty"` PromptVersion string `json:"prompt_version,omitempty"`
PromptHash string `json:"prompt_hash,omitempty"` PromptHash string `json:"prompt_hash,omitempty"`
SessionID string `json:"session_id,omitempty"`
RenderedPromptHash string `json:"rendered_prompt_hash"` RenderedPromptHash string `json:"rendered_prompt_hash"`
SelectedProfileID string `json:"selected_profile_id"` SelectedProfileID string `json:"selected_profile_id"`
SelectedBackendID string `json:"selected_backend_id,omitempty"`
ModelName string `json:"model_name"` ModelName string `json:"model_name"`
Endpoint string `json:"endpoint"` Endpoint string `json:"endpoint"`
EffectiveModelParams ExecutionTarget `json:"effective_model_params"` EffectiveModelParams ExecutionTarget `json:"effective_model_params"`

55
prepared_execution.go Normal file
View File

@@ -0,0 +1,55 @@
package promptkit
import "gitea.maximumdirect.net/eric/promptkit/internal/usecase"
const preparedExecutionString = "promptkit.PreparedExecution{opaque}"
// PreparedExecution is an opaque, in-process handle for one completely
// prepared execution. A handle is bound to the [Engine] that created it and
// permits one [Engine.RunPrepared] invocation.
//
// PreparedExecution contains no supported serializable state and cannot be
// used as a restartable job. Copying the value preserves the same shared
// lifecycle; it does not create another execution attempt.
type PreparedExecution struct {
internal *usecase.PreparedExecution
}
// Details returns a fresh caller-owned, credential-redacted copy of the
// prepared request details. Mutating the result cannot affect execution or a
// later Details call. Details remains available after execution or discard.
//
// A nil receiver or zero-value PreparedExecution returns a zero [PreparedRun].
func (p *PreparedExecution) Details() PreparedRun {
if p == nil || p.internal == nil {
return PreparedRun{}
}
details := fromDomainPreparedRun(p.internal.Details())
if details == nil {
return PreparedRun{}
}
return *details
}
// Discard invalidates an unclaimed handle and drops Promptkit's references to
// its execution-only state. Discard is nil-safe and idempotent. It does not
// cancel an execution that has already claimed the handle; use the
// [Engine.RunPrepared] context for cancellation.
func (p *PreparedExecution) Discard() {
if p == nil || p.internal == nil {
return
}
p.internal.Discard()
}
// String returns a constant representation that exposes no retained request,
// rendered content, or credential data.
func (p *PreparedExecution) String() string {
return preparedExecutionString
}
// GoString returns a constant Go-syntax representation that exposes no
// retained request, rendered content, or credential data.
func (p *PreparedExecution) GoString() string {
return preparedExecutionString
}

View File

@@ -0,0 +1,773 @@
package promptkit_test
import (
"context"
"encoding/json"
"errors"
"fmt"
"os"
"reflect"
"strings"
"sync"
"testing"
"testing/fstest"
"time"
"gitea.maximumdirect.net/eric/promptkit"
)
func TestPreparedExecutionFreezesSourcesAndReturnsIndependentDetails(t *testing.T) {
promptSource := preparedPromptSource("original")
profileSource := preparedProfileSource("original-model")
schemaSource := preparedSchemaSource()
reader := &mutablePreparedArtifactReader{
body: "original artifact",
hash: "original-input-hash",
}
client := &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: `{"value":3}`},
}
engine, err := promptkit.NewEngine(
promptkit.Config{},
promptkit.WithPromptFS(promptSource, "."),
promptkit.WithProfileFS(profileSource, "."),
promptkit.WithSchemaFS(schemaSource, "."),
promptkit.WithArtifactReader(reader),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
temperature := 0.25
extraParams := map[string]any{
"nested": map[string]any{"source": "original"},
}
request := promptkit.RunRequest{
PromptID: "prepared",
Inputs: map[string]promptkit.ArtifactRef{
"input": promptkit.Inline("original request input"),
},
Vars: map[string]string{"label": "original variable"},
Execution: &promptkit.ExecutionTargetOverride{
Temperature: &temperature,
ExtraParams: extraParams,
},
}
preparationContext, cancelPreparation := context.WithCancel(context.Background())
prepared, err := engine.PrepareExecution(preparationContext, request)
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
cancelPreparation()
request.PromptID = "changed"
request.Inputs["input"] = promptkit.Inline("changed request input")
request.Vars["label"] = "changed variable"
temperature = 1.5
extraParams["nested"].(map[string]any)["source"] = "changed"
promptSource["prompt.yaml"] = &fstest.MapFile{Data: []byte(`id: changed`)}
profileSource["profile.yaml"] = &fstest.MapFile{Data: []byte(`id: changed`)}
schemaSource["schema.json"] = &fstest.MapFile{Data: []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "changed root",
"type": "string"
}`)}
schemaSource["value.json"] = &fstest.MapFile{Data: []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "string"
}`)}
reader.set("changed artifact", "changed-input-hash")
first := prepared.Details()
first.Messages[0].Content = "changed details"
first.InputHashes["input"] = "changed-details-hash"
first.EffectiveModelParams.ExtraParams["nested"].(map[string]any)["source"] = "changed details"
first.StructuredOutput.JSONSchema.Schema.(map[string]any)["title"] = "changed details"
second := prepared.Details()
if second.Messages[0].Content != "Input=original artifact Label=original variable" {
t.Fatalf("details message changed: %q", second.Messages[0].Content)
}
if second.InputHashes["input"] != "original-input-hash" {
t.Fatalf("details input hash changed: %q", second.InputHashes["input"])
}
if second.EffectiveModelParams.Model != "original-model" ||
second.EffectiveModelParams.Temperature != 0.25 ||
second.EffectiveModelParams.ExtraParams["nested"].(map[string]any)["source"] != "original" {
t.Fatalf("details target changed: %+v", second.EffectiveModelParams)
}
schema := second.StructuredOutput.JSONSchema.Schema.(map[string]any)
if schema["title"] != "original root" {
t.Fatalf("details schema changed: %#v", schema)
}
result, err := engine.RunPrepared(context.Background(), prepared)
if err != nil {
t.Fatalf("run prepared after preparation-context cancellation: %v", err)
}
if result.Validation.Status != promptkit.ValidationPassed || !result.Validation.IsValid {
t.Fatalf("frozen schema did not validate original output: %+v", result.Validation)
}
if reader.callCount() != 1 {
t.Fatalf("execution reopened artifact source: calls=%d", reader.callCount())
}
requests := client.snapshot()
if len(requests) != 1 {
t.Fatalf("generation calls=%d, want 1", len(requests))
}
generated := requests[0]
if generated.Prompt.Messages[0].Content != second.Messages[0].Content ||
generated.Target.Model != second.EffectiveModelParams.Model ||
!reflect.DeepEqual(generated.Target.ExtraParams, second.EffectiveModelParams.ExtraParams) ||
!reflect.DeepEqual(generated.StructuredOutput, second.StructuredOutput) {
t.Fatalf("generation did not use frozen details:\nrequest=%+v\ndetails=%+v", generated, second)
}
if result.PromptID != second.PromptID ||
result.PromptVersion != second.PromptVersion ||
result.PromptHash != second.PromptHash ||
result.SessionID != second.SessionID ||
result.RenderedPromptHash != second.RenderedPromptHash ||
result.SelectedProfileID != second.SelectedProfileID ||
result.SelectedBackendID != second.SelectedBackendID ||
!reflect.DeepEqual(result.EffectiveModelParams, second.EffectiveModelParams) ||
!reflect.DeepEqual(result.InputHashes, second.InputHashes) {
t.Fatalf("result provenance does not match details:\nresult=%+v\ndetails=%+v", result, second)
}
}
func TestPreparedExecutionLifecycleAndEngineBinding(t *testing.T) {
ownerClient := &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: "ok"},
}
owner := newPreparedContractEngine(t, ownerClient, "owner content")
foreign := newPreparedContractEngine(t, &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: "unexpected"},
}, "foreign content")
prepared, err := owner.PrepareExecution(context.Background(), promptkit.RunRequest{PromptID: "prepared"})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
copied := *prepared
var nilEngine *promptkit.Engine
if result, err := nilEngine.RunPrepared(context.Background(), prepared); result != nil ||
!errors.Is(err, promptkit.ErrInvalidConfig) {
t.Fatalf("nil engine result=(%+v, %v), want ErrInvalidConfig", result, err)
}
if result, err := foreign.RunPrepared(context.Background(), prepared); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("foreign engine result=(%+v, %v), want ErrInvalidRequest", result, err)
}
if result, err := owner.RunPrepared(context.Background(), nil); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("nil handle result=(%+v, %v), want ErrInvalidRequest", result, err)
}
if result, err := owner.RunPrepared(context.Background(), &promptkit.PreparedExecution{}); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("zero handle result=(%+v, %v), want ErrInvalidRequest", result, err)
}
result, err := owner.RunPrepared(context.Background(), &copied)
if err != nil || result == nil {
t.Fatalf("owner run prepared=(%+v, %v), want success", result, err)
}
for name, handle := range map[string]*promptkit.PreparedExecution{
"original": prepared,
"copy": &copied,
} {
if result, err := owner.RunPrepared(context.Background(), handle); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("%s reused handle result=(%+v, %v), want ErrInvalidRequest", name, result, err)
}
if handle.Details().PromptID != "prepared" {
t.Fatalf("%s details unavailable after execution", name)
}
}
if len(ownerClient.snapshot()) != 1 {
t.Fatalf("owner generation calls=%d, want 1", len(ownerClient.snapshot()))
}
collaboratorFailure := errors.New("prepared collaborator failure")
failingClient := &preparedRecordingClient{err: collaboratorFailure}
failingEngine := newPreparedContractEngine(t, failingClient, "failure content")
failing, err := failingEngine.PrepareExecution(context.Background(), promptkit.RunRequest{PromptID: "prepared"})
if err != nil {
t.Fatalf("prepare failing execution: %v", err)
}
if result, err := failingEngine.RunPrepared(context.Background(), failing); result != nil ||
!errors.Is(err, promptkit.ErrLLMGenerate) ||
!errors.Is(err, collaboratorFailure) {
t.Fatalf("generation failure result=(%+v, %v), want public and collaborator identities", result, err)
}
if result, err := failingEngine.RunPrepared(context.Background(), failing); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("failed execution was reusable: result=(%+v, %v)", result, err)
}
cancellationRelease := make(chan struct{})
cancellationStarted := make(chan struct{}, 1)
cancelingEngine := newPreparedContractEngine(t, &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: "unexpected"},
started: cancellationStarted,
release: cancellationRelease,
}, "cancellation content")
canceling, err := cancelingEngine.PrepareExecution(
context.Background(),
promptkit.RunRequest{PromptID: "prepared"},
)
if err != nil {
t.Fatalf("prepare canceled execution: %v", err)
}
executionContext, cancelExecution := context.WithCancel(context.Background())
type canceledOutcome struct {
result *promptkit.RunResult
err error
}
canceledResult := make(chan canceledOutcome, 1)
go func() {
result, runErr := cancelingEngine.RunPrepared(executionContext, canceling)
canceledResult <- canceledOutcome{result: result, err: runErr}
}()
select {
case <-cancellationStarted:
case <-time.After(5 * time.Second):
t.Fatal("timed out waiting for cancelable generation")
}
cancelExecution()
select {
case outcome := <-canceledResult:
if outcome.result != nil ||
!errors.Is(outcome.err, promptkit.ErrLLMGenerate) ||
!errors.Is(outcome.err, context.Canceled) {
t.Fatalf(
"canceled execution=(%+v, %v), want generation and context identities",
outcome.result,
outcome.err,
)
}
case <-time.After(5 * time.Second):
t.Fatal("timed out waiting for canceled execution")
}
if result, err := cancelingEngine.RunPrepared(context.Background(), canceling); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("canceled execution was reusable: result=(%+v, %v)", result, err)
}
}
func TestPreparedExecutionConcurrentClaimAllowsOneGeneration(t *testing.T) {
release := make(chan struct{})
client := &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: "ok"},
started: make(chan struct{}, 1),
release: release,
}
engine := newPreparedContractEngine(t, client, "concurrent content")
prepared, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{PromptID: "prepared"})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
type outcome struct {
result *promptkit.RunResult
err error
}
outcomes := make(chan outcome, 2)
for i := 0; i < 2; i++ {
go func() {
result, runErr := engine.RunPrepared(context.Background(), prepared)
outcomes <- outcome{result: result, err: runErr}
}()
}
select {
case <-client.started:
case <-time.After(5 * time.Second):
t.Fatal("timed out waiting for generation")
}
select {
case loser := <-outcomes:
if loser.result != nil || !errors.Is(loser.err, promptkit.ErrInvalidRequest) {
t.Fatalf("concurrent loser=(%+v, %v), want ErrInvalidRequest", loser.result, loser.err)
}
case <-time.After(5 * time.Second):
t.Fatal("timed out waiting for rejected concurrent claim")
}
close(release)
select {
case winner := <-outcomes:
if winner.err != nil || winner.result == nil {
t.Fatalf("concurrent winner=(%+v, %v), want success", winner.result, winner.err)
}
case <-time.After(5 * time.Second):
t.Fatal("timed out waiting for successful concurrent claim")
}
if len(client.snapshot()) != 1 {
t.Fatalf("generation calls=%d, want 1", len(client.snapshot()))
}
}
func TestPreparedExecutionRunAndDiscardRaceHasOneWinner(t *testing.T) {
const attempts = 32
for i := 0; i < attempts; i++ {
client := &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: "ok"},
}
engine := newPreparedContractEngine(t, client, "race content")
prepared, err := engine.PrepareExecution(
context.Background(),
promptkit.RunRequest{PromptID: "prepared"},
)
if err != nil {
t.Fatalf("attempt %d prepare execution: %v", i, err)
}
start := make(chan struct{})
type outcome struct {
result *promptkit.RunResult
err error
}
runOutcome := make(chan outcome, 1)
discardDone := make(chan struct{})
go func() {
<-start
result, runErr := engine.RunPrepared(context.Background(), prepared)
runOutcome <- outcome{result: result, err: runErr}
}()
go func() {
<-start
prepared.Discard()
close(discardDone)
}()
close(start)
runResult := <-runOutcome
<-discardDone
calls := len(client.snapshot())
switch {
case runResult.err == nil:
if runResult.result == nil || calls != 1 {
t.Fatalf(
"attempt %d run won with outcome=(%+v, %v), generation calls=%d",
i,
runResult.result,
runResult.err,
calls,
)
}
case errors.Is(runResult.err, promptkit.ErrInvalidRequest):
if runResult.result != nil || calls != 0 {
t.Fatalf(
"attempt %d discard won with outcome=(%+v, %v), generation calls=%d",
i,
runResult.result,
runResult.err,
calls,
)
}
default:
t.Fatalf("attempt %d unexpected run outcome=(%+v, %v)", i, runResult.result, runResult.err)
}
if prepared.Details().PromptID != "prepared" {
t.Fatalf("attempt %d details unavailable after race", i)
}
}
}
func TestPreparedExecutionDiscardAndFormattingDoNotExposePrivateState(t *testing.T) {
const (
directCredential = "pk-test-direct-credential-41f7"
renderedContent = "rendered-content-sentinel-98d2"
)
client := &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: "generated output"},
}
engine, err := promptkit.NewEngine(
promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prepared", "profile", renderedContent), "."),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile",
Endpoint: "http://example.test/v1",
Model: "model",
APIKeyRequired: true,
}),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
prepared, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{
PromptID: "prepared",
APIKey: directCredential,
})
if err != nil {
t.Fatalf("prepare execution: %v", err)
}
formattedValues := []string{
fmt.Sprint(prepared),
fmt.Sprintf("%+v", prepared),
fmt.Sprintf("%#v", prepared),
}
for _, formatted := range formattedValues {
if formatted != "promptkit.PreparedExecution{opaque}" {
t.Fatalf("unexpected opaque formatting: %q", formatted)
}
assertPreparedPrivateValuesAbsent(t, formatted, directCredential, renderedContent)
}
payload, err := json.Marshal(prepared)
if err != nil {
t.Fatalf("marshal opaque handle: %v", err)
}
assertPreparedPrivateValuesAbsent(t, string(payload), directCredential, renderedContent)
detailsBefore := prepared.Details()
detailsJSON, err := json.Marshal(detailsBefore)
if err != nil {
t.Fatalf("marshal prepared details: %v", err)
}
assertPreparedPrivateValuesAbsent(t, string(detailsJSON), directCredential)
prepared.Discard()
prepared.Discard()
result, lifecycleErr := engine.RunPrepared(context.Background(), prepared)
if result != nil || !errors.Is(lifecycleErr, promptkit.ErrInvalidRequest) {
t.Fatalf("discarded execution result=(%+v, %v), want ErrInvalidRequest", result, lifecycleErr)
}
assertPreparedPrivateValuesAbsent(t, lifecycleErr.Error(), directCredential, renderedContent)
if !reflect.DeepEqual(prepared.Details(), detailsBefore) {
t.Fatal("details changed after discard")
}
executed, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{
PromptID: "prepared",
APIKey: directCredential,
})
if err != nil {
t.Fatalf("prepare execution for request inspection: %v", err)
}
executionResult, err := engine.RunPrepared(context.Background(), executed)
if err != nil {
t.Fatalf("run execution for request inspection: %v", err)
}
requests := client.snapshot()
if len(requests) != 1 || requests[0].APIKey != directCredential {
t.Fatalf("direct credential did not reach only the client credential field: %#v", requests)
}
requestJSON, err := json.Marshal(requests[0])
if err != nil {
t.Fatalf("marshal captured generate request: %v", err)
}
for _, value := range []string{
fmt.Sprint(requests[0]),
fmt.Sprintf("%+v", requests[0]),
fmt.Sprintf("%#v", requests[0]),
string(requestJSON),
fmt.Sprint(executionResult),
} {
assertPreparedPrivateValuesAbsent(t, value, directCredential)
}
resultJSON, err := json.Marshal(executionResult)
if err != nil {
t.Fatalf("marshal execution result: %v", err)
}
assertPreparedPrivateValuesAbsent(t, string(resultJSON), directCredential)
var nilHandle *promptkit.PreparedExecution
nilHandle.Discard()
if !reflect.DeepEqual(nilHandle.Details(), promptkit.PreparedRun{}) {
t.Fatalf("nil handle details=%+v, want zero value", nilHandle.Details())
}
zeroHandle := &promptkit.PreparedExecution{}
zeroHandle.Discard()
if !reflect.DeepEqual(zeroHandle.Details(), promptkit.PreparedRun{}) {
t.Fatalf("zero handle details=%+v, want zero value", zeroHandle.Details())
}
}
func TestPreparedExecutionCredentialCapacityAndTimingBoundaries(t *testing.T) {
t.Run("credential is rechecked before generation", func(t *testing.T) {
const (
environmentName = "PROMPTKIT_PREPARED_CONTRACT_KEY"
environmentKey = "environment-credential-sentinel"
)
t.Setenv(environmentName, environmentKey)
client := &preparedRecordingClient{
response: &promptkit.GenerateResponse{Content: "unexpected"},
}
engine, err := promptkit.NewEngine(
promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prepared", "profile", "content"), "."),
promptkit.WithProfileFS(preparedCredentialProfileSource(environmentName), "."),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct credential engine: %v", err)
}
prepared, err := engine.PrepareExecution(context.Background(), promptkit.RunRequest{PromptID: "prepared"})
if err != nil {
t.Fatalf("prepare credential execution: %v", err)
}
if err := os.Unsetenv(environmentName); err != nil {
t.Fatalf("unset credential environment: %v", err)
}
result, err := engine.RunPrepared(context.Background(), prepared)
if result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) ||
!errors.Is(err, promptkit.ErrAPIKeyEnvMissing) {
t.Fatalf("credential execution=(%+v, %v), want credential identities", result, err)
}
if len(client.snapshot()) != 0 {
t.Fatalf("credential failure reached generation: %d calls", len(client.snapshot()))
}
assertPreparedPrivateValuesAbsent(t, err.Error(), environmentKey)
if result, err := engine.RunPrepared(context.Background(), prepared); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("credential failure did not consume handle: result=(%+v, %v)", result, err)
}
})
t.Run("preparation does not admit and execution timing starts after retention", func(t *testing.T) {
release := make(chan struct{})
client := newCapacityGateClient(release, 4)
engine := newBackendCapacityEngine(t, client, 1, capacityInt(0), nil)
activeRun := make(chan capacityRunResult, 1)
go runCapacityRequest(
engine,
context.Background(),
promptkit.RunRequest{PromptID: "prompt"},
activeRun,
)
awaitCapacityRequest(t, client.started)
prepared, err := engine.PrepareExecution(
context.Background(),
promptkit.RunRequest{PromptID: "prompt"},
)
if err != nil {
t.Fatalf("prepare while capacity is full: %v", err)
}
if _, _, calls := client.snapshot(); calls != 1 {
t.Fatalf("preparation invoked generation: calls=%d", calls)
}
if result, err := engine.RunPrepared(context.Background(), prepared); result != nil ||
!errors.Is(err, promptkit.ErrCapacityExceeded) {
t.Fatalf("capacity execution=(%+v, %v), want ErrCapacityExceeded", result, err)
} else {
var capacityErr *promptkit.CapacityError
if !errors.As(err, &capacityErr) || capacityErr == nil || capacityErr.BackendID != "limited" {
t.Fatalf("capacity execution=%v, want limited CapacityError", err)
}
}
if result, err := engine.RunPrepared(context.Background(), prepared); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("capacity rejection did not consume handle: result=(%+v, %v)", result, err)
}
if prepared.Details().PromptID != "prompt" {
t.Fatal("details unavailable after capacity rejection")
}
close(release)
activeOutcome := awaitCapacityRun(t, activeRun)
if activeOutcome.err != nil || activeOutcome.result == nil {
t.Fatalf("active run outcome=(%+v, %v), want success", activeOutcome.result, activeOutcome.err)
}
timed, err := engine.PrepareExecution(
context.Background(),
promptkit.RunRequest{PromptID: "prompt"},
)
if err != nil {
t.Fatalf("prepare timed execution: %v", err)
}
details := timed.Details()
time.Sleep(25 * time.Millisecond)
executionFloor := time.Now().UTC()
result, err := engine.RunPrepared(context.Background(), timed)
if err != nil {
t.Fatalf("run timed execution: %v", err)
}
if result.StartTime.Before(executionFloor) ||
!result.StartTime.After(details.EndTime) ||
result.EndTime.Before(result.StartTime) ||
result.Duration != result.EndTime.Sub(result.StartTime) {
t.Fatalf(
"execution timing includes preparation or retention: details_end=%s floor=%s result=%+v",
details.EndTime,
executionFloor,
result,
)
}
if _, _, calls := client.snapshot(); calls != 2 {
t.Fatalf("generation calls=%d, want active and timed executions only", calls)
}
})
}
type mutablePreparedArtifactReader struct {
mu sync.Mutex
body string
hash string
calls int
}
func (r *mutablePreparedArtifactReader) Read(
_ context.Context,
_ promptkit.ArtifactRef,
) (*promptkit.Artifact, error) {
r.mu.Lock()
defer r.mu.Unlock()
r.calls++
return &promptkit.Artifact{
Body: []byte(r.body),
Hash: r.hash,
}, nil
}
func (r *mutablePreparedArtifactReader) set(body, hash string) {
r.mu.Lock()
defer r.mu.Unlock()
r.body = body
r.hash = hash
}
func (r *mutablePreparedArtifactReader) callCount() int {
r.mu.Lock()
defer r.mu.Unlock()
return r.calls
}
type preparedRecordingClient struct {
mu sync.Mutex
response *promptkit.GenerateResponse
err error
requests []promptkit.GenerateRequest
started chan struct{}
release <-chan struct{}
}
func (c *preparedRecordingClient) Generate(
ctx context.Context,
request promptkit.GenerateRequest,
) (*promptkit.GenerateResponse, error) {
c.mu.Lock()
c.requests = append(c.requests, request)
c.mu.Unlock()
if c.started != nil {
c.started <- struct{}{}
}
if c.release != nil {
select {
case <-c.release:
case <-ctx.Done():
return nil, ctx.Err()
}
}
if c.err != nil {
return nil, c.err
}
return c.response, nil
}
func (c *preparedRecordingClient) snapshot() []promptkit.GenerateRequest {
c.mu.Lock()
defer c.mu.Unlock()
return append([]promptkit.GenerateRequest(nil), c.requests...)
}
func newPreparedContractEngine(
t *testing.T,
client promptkit.LLMClient,
message string,
) *promptkit.Engine {
t.Helper()
engine, err := promptkit.NewEngine(
promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prepared", "profile", message), "."),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile",
Endpoint: "http://example.test/v1",
Model: "model",
}),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct prepared execution engine: %v", err)
}
return engine
}
func preparedPromptSource(label string) fstest.MapFS {
return fstest.MapFS{
"prompt.yaml": &fstest.MapFile{Data: []byte(`id: prepared
version: "1"
default_profile: profile
inputs:
- name: input
required: true
messages:
- role: user
content: 'Input={{input "input"}} Label={{.label}}'
description: ` + label + `
output:
format: json
validation_mode: json_schema
schema_path: schema.json
`)},
}
}
func preparedProfileSource(model string) fstest.MapFS {
return fstest.MapFS{
"profile.yaml": &fstest.MapFile{Data: []byte(`id: profile
endpoint: http://example.test/v1
model: ` + model + `
`)},
}
}
func preparedCredentialProfileSource(environmentName string) fstest.MapFS {
return fstest.MapFS{
"profile.yaml": &fstest.MapFile{Data: []byte(`id: profile
endpoint: http://example.test/v1
model: model
api_key_env: ` + environmentName + `
`)},
}
}
func preparedSchemaSource() fstest.MapFS {
return fstest.MapFS{
"schema.json": &fstest.MapFile{Data: []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "original root",
"type": "object",
"required": ["value"],
"properties": {
"value": {"$ref": "value.json"}
}
}`)},
"value.json": &fstest.MapFile{Data: []byte(`{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "integer",
"minimum": 2
}`)},
}
}
func assertPreparedPrivateValuesAbsent(t *testing.T, value string, privateValues ...string) {
t.Helper()
for _, privateValue := range privateValues {
if strings.Contains(value, privateValue) {
t.Fatalf("value exposed private data %q: %s", privateValue, value)
}
}
}

View File

@@ -7,6 +7,7 @@ import (
"strings" "strings"
"gitea.maximumdirect.net/eric/promptkit/internal/domain" "gitea.maximumdirect.net/eric/promptkit/internal/domain"
"gitea.maximumdirect.net/eric/promptkit/internal/jsonvalue"
"gitea.maximumdirect.net/eric/promptkit/internal/profile" "gitea.maximumdirect.net/eric/promptkit/internal/profile"
) )
@@ -15,7 +16,8 @@ import (
// //
// It does not register global state, maintain a model catalog, or resolve // It does not register global state, maintain a model catalog, or resolve
// credentials. If APIKeyRequired is true, callers satisfy it with // credentials. If APIKeyRequired is true, callers satisfy it with
// RunRequest.APIKey. Raw API keys do not belong in profiles. // RunRequest.APIKey or an explicit request ExecutionTargetOverride.APIKeyEnv.
// Raw API keys do not belong in profiles.
// //
// The function copies the ExtraParams map itself but does not recursively copy // The function copies the ExtraParams map itself but does not recursively copy
// nested values. Validation and a deep copy occur when NewEngine applies a // nested values. Validation and a deep copy occur when NewEngine applies a
@@ -23,6 +25,7 @@ import (
func OpenAICompatibleProfile(cfg OpenAICompatibleProfileConfig) Profile { func OpenAICompatibleProfile(cfg OpenAICompatibleProfileConfig) Profile {
return Profile{ return Profile{
ID: cfg.ID, ID: cfg.ID,
BackendID: cfg.BackendID,
Endpoint: cfg.Endpoint, Endpoint: cfg.Endpoint,
Model: cfg.Model, Model: cfg.Model,
Temperature: cfg.Temperature, Temperature: cfg.Temperature,
@@ -79,12 +82,13 @@ func (r *memoryProfileRepository) GetProfile(_ context.Context, id string) (*dom
} }
func toDomainProfile(publicProfile Profile) (domain.ExecutionProfile, error) { func toDomainProfile(publicProfile Profile) (domain.ExecutionProfile, error) {
extraParams, err := copyPublicJSONMap(publicProfile.ExtraParams) extraParams, err := jsonvalue.CopyMap(publicProfile.ExtraParams)
if err != nil { if err != nil {
return domain.ExecutionProfile{}, err return domain.ExecutionProfile{}, err
} }
prof := domain.ExecutionProfile{ prof := domain.ExecutionProfile{
ID: strings.TrimSpace(publicProfile.ID), ID: strings.TrimSpace(publicProfile.ID),
BackendID: strings.TrimSpace(publicProfile.BackendID),
Endpoint: publicProfile.Endpoint, Endpoint: publicProfile.Endpoint,
Model: publicProfile.Model, Model: publicProfile.Model,
Temperature: publicProfile.Temperature, Temperature: publicProfile.Temperature,
@@ -106,8 +110,8 @@ func validatePublicProfile(prof domain.ExecutionProfile) error {
if strings.TrimSpace(prof.ID) == "" { if strings.TrimSpace(prof.ID) == "" {
return errors.New("id is required") return errors.New("id is required")
} }
if strings.TrimSpace(prof.Endpoint) == "" { if strings.TrimSpace(prof.BackendID) == "" && strings.TrimSpace(prof.Endpoint) == "" {
return errors.New("endpoint is required") return errors.New("backend or endpoint is required")
} }
if strings.TrimSpace(prof.Model) == "" { if strings.TrimSpace(prof.Model) == "" {
return errors.New("model is required") return errors.New("model is required")

View File

@@ -3,7 +3,10 @@ package promptkit_test
import ( import (
"context" "context"
"encoding/json" "encoding/json"
"errors"
"fmt" "fmt"
"io/fs"
"reflect"
"strings" "strings"
"sync" "sync"
"sync/atomic" "sync/atomic"
@@ -14,6 +17,380 @@ import (
"gitea.maximumdirect.net/eric/promptkit" "gitea.maximumdirect.net/eric/promptkit"
) )
type inspectionCountingFS struct {
opens atomic.Int64
}
func (f *inspectionCountingFS) Open(string) (fs.File, error) {
f.opens.Add(1)
return nil, fs.ErrNotExist
}
func TestInspectProfileResolvesCredentialStatesWithoutPromptOrGeneration(t *testing.T) {
const environmentName = "PROMPTKIT_INSPECTION_ABSENT_KEY"
t.Setenv(environmentName, "")
client := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "unexpected"}}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(fstest.MapFS{}, "."),
promptkit.WithProfileFS(fstest.MapFS{
"environment.yaml": &fstest.MapFile{Data: []byte(`id: environment
endpoint: http://environment.example/v1
model: environment-model
api_key_env: PROMPTKIT_INSPECTION_ABSENT_KEY
`)},
}, "."),
promptkit.WithProfiles(
promptkit.Profile{ID: "direct", Endpoint: "http://direct.example/v1", Model: "direct-model", APIKeyRequired: true},
promptkit.Profile{ID: "none", Endpoint: "http://none.example/v1", Model: "none-model"},
),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct inspection engine: %v", err)
}
for _, tc := range []struct {
profileID string
wantEnv string
wantDirectKey bool
wantEndpoint string
}{
{profileID: " environment ", wantEnv: environmentName, wantEndpoint: "http://environment.example/v1"},
{profileID: "direct", wantDirectKey: true, wantEndpoint: "http://direct.example/v1"},
{profileID: "none", wantEndpoint: "http://none.example/v1"},
} {
t.Run(tc.profileID, func(t *testing.T) {
inspection, err := engine.InspectProfile(context.Background(), tc.profileID)
if err != nil {
t.Fatalf("inspect profile: %v", err)
}
if inspection.ProfileID != strings.TrimSpace(tc.profileID) ||
inspection.EffectiveModelParams.Endpoint != tc.wantEndpoint ||
inspection.EffectiveModelParams.BackendID != "" ||
inspection.EffectiveModelParams.APIKeyEnv != tc.wantEnv ||
inspection.APIKeyRequired != tc.wantDirectKey {
t.Fatalf("unexpected inspection: %#v", inspection)
}
})
}
if len(client.requests) != 0 {
t.Fatalf("inspection invoked the model client %d times", len(client.requests))
}
}
func TestInspectProfilePreservesPublicErrorIdentities(t *testing.T) {
newEngine := func(t *testing.T, options ...promptkit.Option) *promptkit.Engine {
t.Helper()
engine, err := promptkit.NewEngine(
promptkit.Config{},
append([]promptkit.Option{promptkit.WithPromptFS(fstest.MapFS{}, ".")}, options...)...,
)
if err != nil {
t.Fatalf("construct inspection engine: %v", err)
}
return engine
}
var nilEngine *promptkit.Engine
if result, err := nilEngine.InspectProfile(context.Background(), "profile"); result != nil ||
!errors.Is(err, promptkit.ErrInvalidConfig) {
t.Fatalf("nil engine result=(%#v, %v), want ErrInvalidConfig", result, err)
}
valid := newEngine(t, promptkit.WithProfiles(promptkit.Profile{
ID: "profile", Endpoint: "http://profile.example/v1", Model: "model",
}))
if result, err := valid.InspectProfile(context.Background(), " \t "); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("blank profile result=(%#v, %v), want ErrInvalidRequest", result, err)
}
if result, err := valid.InspectProfile(context.Background(), "missing"); result != nil ||
!errors.Is(err, promptkit.ErrProfileNotFound) || errors.Is(err, promptkit.ErrProfileLoad) {
t.Fatalf("missing profile result=(%#v, %v), want only ErrProfileNotFound", result, err)
}
malformed := newEngine(t, promptkit.WithProfileFS(fstest.MapFS{
"broken.yaml": &fstest.MapFile{Data: []byte("id: broken\nendpoint: http://broken.example/v1\nmodel: model\nunknown: value\n")},
}, "."))
if result, err := malformed.InspectProfile(context.Background(), "broken"); result != nil ||
!errors.Is(err, promptkit.ErrProfileLoad) {
t.Fatalf("malformed profile result=(%#v, %v), want ErrProfileLoad", result, err)
}
unknownBackend := newEngine(t, promptkit.WithProfiles(promptkit.Profile{
ID: "unknown-backend", BackendID: "unknown", Model: "model",
}))
if result, err := unknownBackend.InspectProfile(context.Background(), "unknown-backend"); result != nil ||
!errors.Is(err, promptkit.ErrProfileLoad) {
t.Fatalf("unknown backend result=(%#v, %v), want ErrProfileLoad", result, err)
}
countingFS := &inspectionCountingFS{}
canceled := newEngine(t, promptkit.WithProfileFS(countingFS, "."))
ctx, cancel := context.WithCancel(context.Background())
cancel()
if result, err := canceled.InspectProfile(ctx, "profile"); result != nil ||
!errors.Is(err, promptkit.ErrProfileLoad) || !errors.Is(err, context.Canceled) || countingFS.opens.Load() != 0 {
t.Fatalf("canceled inspection result=(%#v, %v), opens=%d", result, err, countingFS.opens.Load())
}
}
func TestInspectProfileReturnsIndependentTargetMatchingPreparation(t *testing.T) {
extraParams := map[string]any{
"nested": map[string]any{"value": "original"},
}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile", Endpoint: "http://profile.example/v1", Model: "model", ExtraParams: extraParams,
}),
)
if err != nil {
t.Fatalf("construct inspection engine: %v", err)
}
first, err := engine.InspectProfile(context.Background(), "profile")
if err != nil {
t.Fatalf("first inspection: %v", err)
}
first.EffectiveModelParams.ExtraParams["nested"].(map[string]any)["value"] = "changed"
first.EffectiveModelParams.ExtraParams["later"] = true
second, err := engine.InspectProfile(context.Background(), "profile")
if err != nil {
t.Fatalf("second inspection: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare after inspection mutation: %v", err)
}
for _, target := range []promptkit.ExecutionTarget{second.EffectiveModelParams, prepared.EffectiveModelParams} {
if target.ExtraParams["nested"].(map[string]any)["value"] != "original" || target.ExtraParams["later"] != nil {
t.Fatalf("inspection mutation reached engine-owned target: %#v", target.ExtraParams)
}
}
if !reflect.DeepEqual(second.EffectiveModelParams, prepared.EffectiveModelParams) {
t.Fatalf("inspection target=%#v, preparation target=%#v", second.EffectiveModelParams, prepared.EffectiveModelParams)
}
}
func TestInspectPromptReturnsDeclaredMetadataWithoutExecutionWork(t *testing.T) {
client := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "unexpected"}}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(fstest.MapFS{
"report-v1.yaml": &fstest.MapFile{Data: []byte(`id: report
version: "1.0.0"
messages:
- role: user
content: old report
output:
format: text
validation_mode: none
`)},
"report-v2.yaml": &fstest.MapFile{Data: []byte(`id: report
version: "2.0.0"
default_profile: missing-profile
inputs:
- name: location
required: true
content_type: text/plain
description: Forecast location.
- name: units
content_type: text/plain
description: Unit preference.
messages:
- role: user
content_file: messages/report.md
output:
format: json
validation_mode: json_schema
schema_path: schemas/report.json
`)},
"messages/report.md": &fstest.MapFile{Data: []byte("rendered report body is not returned")},
}, "."),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct prompt-inspection engine: %v", err)
}
inspection, err := engine.InspectPrompt(context.Background(), "report", "2.0.0")
if err != nil {
t.Fatalf("inspect prompt: %v", err)
}
if inspection.PromptID != "report" ||
inspection.PromptVersion != "2.0.0" ||
inspection.PromptHash == "" ||
inspection.DefaultProfileID != "missing-profile" ||
inspection.OutputContract != (promptkit.OutputContract{
Format: promptkit.FormatJSON,
ValidationMode: promptkit.ValidationJSONSchema,
SchemaPath: "schemas/report.json",
}) {
t.Fatalf("unexpected inspection metadata: %#v", inspection)
}
wantInputs := []promptkit.PromptInputDefinition{
{Name: "location", Required: true, ContentType: "text/plain", Description: "Forecast location."},
{Name: "units", ContentType: "text/plain", Description: "Unit preference."},
}
if !reflect.DeepEqual(inspection.Inputs, wantInputs) {
t.Fatalf("inspection inputs=%#v, want %#v", inspection.Inputs, wantInputs)
}
if len(client.requests) != 0 {
t.Fatalf("inspection invoked the model client %d times", len(client.requests))
}
}
func TestInspectPromptPreservesPublicErrorIdentities(t *testing.T) {
newEngine := func(t *testing.T, source fstest.MapFS) *promptkit.Engine {
t.Helper()
engine, err := promptkit.NewEngine(promptkit.Config{}, promptkit.WithPromptFS(source, "."))
if err != nil {
t.Fatalf("construct prompt-inspection engine: %v", err)
}
return engine
}
validSource := fstest.MapFS{
"prompt.yaml": &fstest.MapFile{Data: []byte(`id: prompt
version: "1"
messages:
- role: user
content: body
output:
format: text
validation_mode: none
`)},
}
var nilEngine *promptkit.Engine
if result, err := nilEngine.InspectPrompt(context.Background(), "prompt", "1"); result != nil ||
!errors.Is(err, promptkit.ErrInvalidConfig) {
t.Fatalf("nil engine result=(%#v, %v), want ErrInvalidConfig", result, err)
}
valid := newEngine(t, validSource)
if result, err := valid.InspectPrompt(context.Background(), " \t ", "1"); result != nil ||
!errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("blank prompt result=(%#v, %v), want ErrInvalidRequest", result, err)
}
if result, err := valid.InspectPrompt(context.Background(), "missing", "1"); result != nil ||
!errors.Is(err, promptkit.ErrPromptNotFound) || errors.Is(err, promptkit.ErrPromptLoad) {
t.Fatalf("missing prompt result=(%#v, %v), want only ErrPromptNotFound", result, err)
}
if result, err := valid.InspectPrompt(context.Background(), "prompt", "missing"); result != nil ||
!errors.Is(err, promptkit.ErrPromptNotFound) || errors.Is(err, promptkit.ErrPromptLoad) {
t.Fatalf("missing version result=(%#v, %v), want only ErrPromptNotFound", result, err)
}
ambiguous := newEngine(t, fstest.MapFS{
"one.yaml": &fstest.MapFile{Data: []byte(`id: prompt
version: "1"
messages:
- role: user
content: first
output:
format: text
validation_mode: none
`)},
"two.yaml": &fstest.MapFile{Data: []byte(`id: prompt
version: "2"
messages:
- role: user
content: second
output:
format: text
validation_mode: none
`)},
})
if result, err := ambiguous.InspectPrompt(context.Background(), "prompt", ""); result != nil ||
!errors.Is(err, promptkit.ErrPromptLoad) {
t.Fatalf("ambiguous prompt result=(%#v, %v), want ErrPromptLoad", result, err)
}
for name, source := range map[string]fstest.MapFS{
"malformed definition": {
"broken.yaml": &fstest.MapFile{Data: []byte("id: broken\nversion: \"1\"\nunknown: value\n")},
},
"missing content file": {
"broken.yaml": &fstest.MapFile{Data: []byte(`id: broken
version: "1"
messages:
- role: user
content_file: missing.md
output:
format: text
validation_mode: none
`)},
},
} {
t.Run(name, func(t *testing.T) {
if result, err := newEngine(t, source).InspectPrompt(context.Background(), "broken", "1"); result != nil ||
!errors.Is(err, promptkit.ErrPromptLoad) {
t.Fatalf("broken prompt result=(%#v, %v), want ErrPromptLoad", result, err)
}
})
}
countingFS := &inspectionCountingFS{}
canceled, err := promptkit.NewEngine(promptkit.Config{}, promptkit.WithPromptFS(countingFS, "."))
if err != nil {
t.Fatalf("construct canceled prompt-inspection engine: %v", err)
}
ctx, cancel := context.WithCancel(context.Background())
cancel()
if result, err := canceled.InspectPrompt(ctx, "prompt", "1"); result != nil ||
!errors.Is(err, promptkit.ErrPromptLoad) || !errors.Is(err, context.Canceled) || countingFS.opens.Load() != 0 {
t.Fatalf("canceled inspection result=(%#v, %v), opens=%d", result, err, countingFS.opens.Load())
}
}
func TestInspectPromptReturnsIndependentMetadataMatchingPreparation(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(fstest.MapFS{
"prompt.yaml": &fstest.MapFile{Data: []byte(`id: prompt
version: "1"
default_profile: profile
inputs:
- name: subject
content_type: text/plain
description: Summary subject.
messages:
- role: user
content: summarize
output:
format: markdown
validation_mode: basic
`)},
}, "."),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile", Endpoint: "http://profile.example/v1", Model: "model",
}),
)
if err != nil {
t.Fatalf("construct prompt-inspection engine: %v", err)
}
first, err := engine.InspectPrompt(context.Background(), "prompt", "1")
if err != nil {
t.Fatalf("first inspection: %v", err)
}
first.Inputs[0].Name = "changed"
first.OutputContract.SchemaPath = "changed.json"
second, err := engine.InspectPrompt(context.Background(), "prompt", "1")
if err != nil {
t.Fatalf("second inspection: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt", PromptVersion: "1"})
if err != nil {
t.Fatalf("prepare after inspection mutation: %v", err)
}
if second.Inputs[0].Name != "subject" || second.OutputContract.SchemaPath != "" ||
prepared.OutputContract.SchemaPath != "" || second.PromptHash != prepared.PromptHash {
t.Fatalf("inspection mutation reached engine-owned prompt metadata: inspection=%#v prepared=%#v", second, prepared)
}
}
func TestPreparedRunJSONOmitsZeroTimingValues(t *testing.T) { func TestPreparedRunJSONOmitsZeroTimingValues(t *testing.T) {
payload, err := json.Marshal(promptkit.PreparedRun{}) payload, err := json.Marshal(promptkit.PreparedRun{})
if err != nil { if err != nil {
@@ -26,6 +403,421 @@ func TestPreparedRunJSONOmitsZeroTimingValues(t *testing.T) {
} }
} }
func TestBackendIdentityJSONNamesAndOmission(t *testing.T) {
t.Run("execution target round trip", func(t *testing.T) {
value := promptkit.ExecutionTarget{BackendID: promptkit.BackendOpenRouter}
payload, err := json.Marshal(value)
if err != nil {
t.Fatalf("marshal execution target: %v", err)
}
var decoded promptkit.ExecutionTarget
if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal execution target: %v", err)
}
if decoded.BackendID != value.BackendID {
t.Fatalf("backend identity did not round trip: got %q want %q", decoded.BackendID, value.BackendID)
}
})
t.Run("prepared run round trip", func(t *testing.T) {
value := promptkit.PreparedRun{SelectedBackendID: promptkit.BackendOpenRouter}
payload, err := json.Marshal(value)
if err != nil {
t.Fatalf("marshal prepared run: %v", err)
}
var decoded promptkit.PreparedRun
if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal prepared run: %v", err)
}
if decoded.SelectedBackendID != value.SelectedBackendID {
t.Fatalf("backend identity did not round trip: got %q want %q", decoded.SelectedBackendID, value.SelectedBackendID)
}
})
t.Run("run result round trip", func(t *testing.T) {
value := promptkit.RunResult{SelectedBackendID: promptkit.BackendOpenRouter}
payload, err := json.Marshal(value)
if err != nil {
t.Fatalf("marshal run result: %v", err)
}
var decoded promptkit.RunResult
if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal run result: %v", err)
}
if decoded.SelectedBackendID != value.SelectedBackendID {
t.Fatalf("backend identity did not round trip: got %q want %q", decoded.SelectedBackendID, value.SelectedBackendID)
}
})
payload, err := json.Marshal(promptkit.ExecutionTarget{})
if err != nil {
t.Fatalf("marshal empty execution target: %v", err)
}
if strings.Contains(string(payload), `"backend_id"`) {
t.Fatalf("empty backend identity was not omitted: %s", payload)
}
}
func TestEndpointOnlyProfileOmitsBackendIdentityFromStableJSON(t *testing.T) {
client := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "ok"}}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithProfiles(promptkit.Profile{
ID: "profile", Endpoint: "http://example.test/v1", Model: "model",
}),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare endpoint-only profile: %v", err)
}
result, err := engine.Run(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("run endpoint-only profile: %v", err)
}
if prepared.SelectedBackendID != "" ||
prepared.EffectiveModelParams.BackendID != "" ||
result.SelectedBackendID != "" ||
result.EffectiveModelParams.BackendID != "" {
t.Fatalf("endpoint-only profile acquired backend identity: prepared=%+v result=%+v", prepared, result)
}
for _, value := range []any{prepared, result} {
payload, err := json.Marshal(value)
if err != nil {
t.Fatalf("marshal endpoint-only value: %v", err)
}
if strings.Contains(string(payload), `"backend_id"`) || strings.Contains(string(payload), `"selected_backend_id"`) {
t.Fatalf("endpoint-only backend identity was not omitted: %s", payload)
}
}
}
func TestUnknownProfileBackendHasProfileLoadIdentity(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithProfiles(promptkit.Profile{ID: "profile", BackendID: "unknown", Model: "model"}),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
_, err = engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if !errors.Is(err, promptkit.ErrProfileLoad) {
t.Fatalf("expected ErrProfileLoad, got %v", err)
}
if errors.Is(err, promptkit.ErrInvalidRequest) {
t.Fatalf("unknown backend should not have invalid-request identity: %v", err)
}
}
func TestLocalBackendConstructsAndRegistersConventionalBackend(t *testing.T) {
const (
localBackendID = "local"
limit = 2
)
if promptkit.BackendLocal != localBackendID {
t.Fatalf("BackendLocal=%q, want %q", promptkit.BackendLocal, localBackendID)
}
endpoint := "http://local.example/v1"
backend := promptkit.LocalBackend(endpoint, limit)
want := promptkit.Backend{
ID: localBackendID,
Endpoint: endpoint,
ConcurrencyLimit: limit,
}
if !reflect.DeepEqual(backend, want) {
t.Fatalf("LocalBackend()=%+v, want %+v", backend, want)
}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "local-profile", "message"), "."),
promptkit.WithBackend(backend),
promptkit.WithProfiles(promptkit.Profile{
ID: "local-profile",
BackendID: localBackendID,
Model: "model",
}),
)
if err != nil {
t.Fatalf("construct engine with local backend: %v", err)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare with local backend: %v", err)
}
if prepared.SelectedBackendID != localBackendID ||
prepared.EffectiveModelParams.Endpoint != endpoint {
t.Fatalf("unexpected local backend preparation: %+v", prepared)
}
}
func TestCustomBackendFlowsThroughProfilesOverridesAndInjectedClient(t *testing.T) {
t.Setenv("CUSTOM_LLM_KEY", "test-key")
client := &fakeLLMClient{response: &promptkit.GenerateResponse{Content: "ok"}}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "backend-profile", "message"), "."),
promptkit.WithBackend(promptkit.Backend{
ID: " custom ",
Endpoint: " http://backend.example/v1 ",
APIKeyEnv: " CUSTOM_LLM_KEY ",
ExtraParams: map[string]any{
"provider": "custom",
},
}),
promptkit.WithProfiles(
promptkit.Profile{ID: "backend-profile", BackendID: "custom", Model: "backend-model"},
promptkit.Profile{ID: "profile-endpoint", BackendID: "custom", Endpoint: "http://profile.example/v1", Model: "profile-model"},
promptkit.Profile{ID: "blank-profile-endpoint", BackendID: "custom", Endpoint: " \t ", Model: "profile-model"},
),
promptkit.WithLLMClient(client),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
result, err := engine.Run(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("run with custom backend: %v", err)
}
if len(client.requests) != 1 {
t.Fatalf("expected one injected-client request, got %d", len(client.requests))
}
target := client.requests[0].Target
if target.BackendID != "custom" ||
target.Endpoint != "http://backend.example/v1" ||
target.APIKeyEnv != "CUSTOM_LLM_KEY" ||
target.Model != "backend-model" ||
target.ExtraParams["provider"] != "custom" ||
result.SelectedBackendID != "custom" {
t.Fatalf("unexpected custom backend settings: target=%+v result_backend=%q", target, result.SelectedBackendID)
}
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: "prompt", ProfileID: "profile-endpoint",
})
if err != nil {
t.Fatalf("prepare profile endpoint override: %v", err)
}
if prepared.SelectedBackendID != "custom" || prepared.EffectiveModelParams.Endpoint != "http://profile.example/v1" {
t.Fatalf("profile endpoint override changed backend identity: %+v", prepared)
}
prepared, err = engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: "prompt", ProfileID: "blank-profile-endpoint",
})
if err != nil {
t.Fatalf("prepare blank profile endpoint: %v", err)
}
if prepared.SelectedBackendID != "custom" || prepared.EffectiveModelParams.Endpoint != "http://backend.example/v1" {
t.Fatalf("blank profile endpoint did not inherit backend endpoint: %+v", prepared)
}
prepared, err = engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: "prompt",
Execution: &promptkit.ExecutionTargetOverride{
Endpoint: "http://request.example/v1",
},
})
if err != nil {
t.Fatalf("prepare request endpoint override: %v", err)
}
if prepared.SelectedBackendID != "custom" || prepared.EffectiveModelParams.Endpoint != "http://request.example/v1" {
t.Fatalf("request endpoint override changed backend identity: %+v", prepared)
}
}
func TestCustomBackendSupportsFileProfileAndBothSelectionPaths(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "file-profile", "message"), "."),
promptkit.WithProfileFS(fstest.MapFS{
"profile.yaml": &fstest.MapFile{Data: []byte(`id: file-profile
backend: file-backend
endpoint: " "
model: file-model
`)},
}, "."),
promptkit.WithBackend(promptkit.Backend{
ID: "file-backend",
Endpoint: "http://file-backend.example/v1",
}),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
for _, request := range []promptkit.RunRequest{
{PromptID: "prompt"},
{PromptID: "prompt", ProfileID: "file-profile"},
} {
prepared, err := engine.Prepare(context.Background(), request)
if err != nil {
t.Fatalf("prepare file profile: %v", err)
}
if prepared.SelectedBackendID != "file-backend" ||
prepared.EffectiveModelParams.Endpoint != "http://file-backend.example/v1" {
t.Fatalf("unexpected file-profile backend resolution: %+v", prepared)
}
}
}
func TestBackendOptionsAccumulateAndRegistrationsAreEngineLocal(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "first-profile", "message"), "."),
promptkit.WithBackend(promptkit.Backend{ID: "first", Endpoint: "http://first.example/v1"}),
promptkit.WithBackend(promptkit.Backend{ID: "second", Endpoint: "http://second.example/v1"}),
promptkit.WithProfiles(
promptkit.Profile{ID: "first-profile", BackendID: "first", Model: "model"},
promptkit.Profile{ID: "second-profile", BackendID: "second", Model: "model"},
),
)
if err != nil {
t.Fatalf("construct engine with accumulated registrations: %v", err)
}
for profileID, wantEndpoint := range map[string]string{
"first-profile": "http://first.example/v1",
"second-profile": "http://second.example/v1",
} {
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{
PromptID: "prompt", ProfileID: profileID,
})
if err != nil {
t.Fatalf("prepare %s: %v", profileID, err)
}
if prepared.EffectiveModelParams.Endpoint != wantEndpoint {
t.Fatalf("profile %s endpoint=%q, want %q", profileID, prepared.EffectiveModelParams.Endpoint, wantEndpoint)
}
}
newEngine := func(endpoint string) *promptkit.Engine {
t.Helper()
value, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithBackend(promptkit.Backend{ID: "same-id", Endpoint: endpoint}),
promptkit.WithProfiles(promptkit.Profile{ID: "profile", BackendID: "same-id", Model: "model"}),
)
if err != nil {
t.Fatalf("construct isolated engine: %v", err)
}
return value
}
firstEngine := newEngine("http://one.example/v1")
secondEngine := newEngine("http://two.example/v1")
for engine, wantEndpoint := range map[*promptkit.Engine]string{
firstEngine: "http://one.example/v1",
secondEngine: "http://two.example/v1",
} {
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("prepare isolated engine: %v", err)
}
if prepared.EffectiveModelParams.Endpoint != wantEndpoint {
t.Fatalf("isolated engine endpoint=%q, want %q", prepared.EffectiveModelParams.Endpoint, wantEndpoint)
}
}
}
func TestWithBackendCopiesQueueCapacity(t *testing.T) {
queueCapacity := 4
option := promptkit.WithBackend(promptkit.Backend{
ID: "custom",
Endpoint: "http://custom.example/v1",
ConcurrencyLimit: 1,
QueueCapacity: &queueCapacity,
})
queueCapacity = -1
_, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
option,
promptkit.WithProfiles(promptkit.Profile{
ID: "profile", BackendID: "custom", Model: "model",
}),
)
if err != nil {
t.Fatalf("construct engine after mutating queue pointer: %v", err)
}
}
func TestBackendRegistrationRejectsInvalidAndDuplicateDefinitions(t *testing.T) {
cycle := map[string]any{}
cycle["self"] = cycle
tests := []struct {
name string
backends []promptkit.Backend
}{
{name: "blank id", backends: []promptkit.Backend{{Endpoint: "http://example.test/v1"}}},
{name: "invalid endpoint", backends: []promptkit.Backend{{ID: "custom", Endpoint: "ftp://example.test/v1"}}},
{name: "invalid environment", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", APIKeyEnv: "BAD-NAME"}}},
{name: "reserved extra parameter", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: map[string]any{"model": "override"}}}},
{name: "cyclic extra parameter", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: cycle}}},
{name: "malformed JSON number", backends: []promptkit.Backend{{ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: map[string]any{"value": json.Number("01")}}}},
{name: "duplicate consumer id", backends: []promptkit.Backend{
{ID: " custom ", Endpoint: "http://one.example/v1"},
{ID: "custom", Endpoint: "http://two.example/v1"},
}},
{name: "reserved built-in id", backends: []promptkit.Backend{{
ID: promptkit.BackendOpenRouter, Endpoint: "http://replacement.example/v1",
}}},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
options := []promptkit.Option{
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
}
for _, backend := range tt.backends {
options = append(options, promptkit.WithBackend(backend))
}
_, err := promptkit.NewEngine(promptkit.Config{}, options...)
if !errors.Is(err, promptkit.ErrInvalidConfig) {
t.Fatalf("expected ErrInvalidConfig, got %v", err)
}
})
}
}
func TestBackendExtraParamsAreDeeplyCopiedAtConstructionAndLookup(t *testing.T) {
nested := map[string]any{"value": "original"}
extraParams := map[string]any{"nested": nested}
engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithBackend(promptkit.Backend{
ID: "custom", Endpoint: "http://example.test/v1", ExtraParams: extraParams,
}),
promptkit.WithProfiles(promptkit.Profile{ID: "profile", BackendID: "custom", Model: "model"}),
)
if err != nil {
t.Fatalf("construct engine: %v", err)
}
nested["value"] = "mutated input"
extraParams["later"] = true
prepared, err := engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("first prepare: %v", err)
}
gotNested := prepared.EffectiveModelParams.ExtraParams["nested"].(map[string]any)
if gotNested["value"] != "original" || prepared.EffectiveModelParams.ExtraParams["later"] != nil {
t.Fatalf("backend retained caller mutations: %#v", prepared.EffectiveModelParams.ExtraParams)
}
gotNested["value"] = "mutated lookup"
prepared, err = engine.Prepare(context.Background(), promptkit.RunRequest{PromptID: "prompt"})
if err != nil {
t.Fatalf("second prepare: %v", err)
}
gotNested = prepared.EffectiveModelParams.ExtraParams["nested"].(map[string]any)
if gotNested["value"] != "original" {
t.Fatalf("backend retained lookup mutation: %#v", prepared.EffectiveModelParams.ExtraParams)
}
}
func TestPreparedRunJSONTimingRoundTrips(t *testing.T) { func TestPreparedRunJSONTimingRoundTrips(t *testing.T) {
start := time.Date(2026, time.July, 29, 12, 0, 0, 0, time.UTC) start := time.Date(2026, time.July, 29, 12, 0, 0, 0, time.UTC)
prepared := promptkit.PreparedRun{ prepared := promptkit.PreparedRun{
@@ -55,6 +847,7 @@ func TestRunResultJSONUsesMillisecondsAndRoundTrips(t *testing.T) {
result := promptkit.RunResult{ result := promptkit.RunResult{
RunID: "opaque-run-id", RunID: "opaque-run-id",
Artifact: promptkit.Artifact{Name: "output", ContentType: "text/plain", Body: []byte("ok")}, Artifact: promptkit.Artifact{Name: "output", ContentType: "text/plain", Body: []byte("ok")},
SessionID: "session-123",
StartTime: start, StartTime: start,
EndTime: start.Add(1500 * time.Millisecond), EndTime: start.Add(1500 * time.Millisecond),
Duration: 1500 * time.Millisecond, Duration: 1500 * time.Millisecond,
@@ -74,6 +867,9 @@ func TestRunResultJSONUsesMillisecondsAndRoundTrips(t *testing.T) {
if _, exists := object["duration"]; exists { if _, exists := object["duration"]; exists {
t.Fatalf("unexpected nanosecond duration field in %s", payload) t.Fatalf("unexpected nanosecond duration field in %s", payload)
} }
if got := object["session_id"]; got != result.SessionID {
t.Fatalf("expected session_id=%q, got %#v in %s", result.SessionID, got, payload)
}
artifact, ok := object["artifact"].(map[string]any) artifact, ok := object["artifact"].(map[string]any)
if !ok || artifact["content_type"] != "text/plain" { if !ok || artifact["content_type"] != "text/plain" {
t.Fatalf("expected stable artifact JSON fields, got %#v", object["artifact"]) t.Fatalf("expected stable artifact JSON fields, got %#v", object["artifact"])
@@ -83,7 +879,10 @@ func TestRunResultJSONUsesMillisecondsAndRoundTrips(t *testing.T) {
if err := json.Unmarshal(payload, &decoded); err != nil { if err := json.Unmarshal(payload, &decoded); err != nil {
t.Fatalf("unmarshal run result: %v", err) t.Fatalf("unmarshal run result: %v", err)
} }
if decoded.Duration != result.Duration || !decoded.StartTime.Equal(result.StartTime) || !decoded.EndTime.Equal(result.EndTime) { if decoded.SessionID != result.SessionID ||
decoded.Duration != result.Duration ||
!decoded.StartTime.Equal(result.StartTime) ||
!decoded.EndTime.Equal(result.EndTime) {
t.Fatalf("timing values did not round trip: got %#v, want %#v", decoded, result) t.Fatalf("timing values did not round trip: got %#v, want %#v", decoded, result)
} }
@@ -91,11 +890,18 @@ func TestRunResultJSONUsesMillisecondsAndRoundTrips(t *testing.T) {
if err != nil { if err != nil {
t.Fatalf("marshal zero run result: %v", err) t.Fatalf("marshal zero run result: %v", err)
} }
for _, field := range []string{"start_time", "end_time", "duration_ms"} { for _, field := range []string{"session_id", "start_time", "end_time", "duration_ms"} {
if strings.Contains(string(payload), `"`+field+`"`) { if strings.Contains(string(payload), `"`+field+`"`) {
t.Fatalf("expected zero %s to be omitted, got %s", field, payload) t.Fatalf("expected zero %s to be omitted, got %s", field, payload)
} }
} }
var decodedEmpty promptkit.RunResult
if err := json.Unmarshal(payload, &decodedEmpty); err != nil {
t.Fatalf("unmarshal run result without session_id: %v", err)
}
if decodedEmpty.SessionID != "" {
t.Fatalf("expected absent session_id to decode empty, got %q", decodedEmpty.SessionID)
}
} }
func TestEngineValidationIsSinglePass(t *testing.T) { func TestEngineValidationIsSinglePass(t *testing.T) {
@@ -254,9 +1060,13 @@ func TestRepeatedOptionsUseLastValueInEachCategory(t *testing.T) {
func TestEngineSupportsConcurrentPrepareAndRun(t *testing.T) { func TestEngineSupportsConcurrentPrepareAndRun(t *testing.T) {
engine, err := promptkit.NewEngine(promptkit.Config{}, engine, err := promptkit.NewEngine(promptkit.Config{},
promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."), promptkit.WithPromptFS(contractPromptFS("prompt", "profile", "message"), "."),
promptkit.WithBackend(promptkit.Backend{
ID: "concurrent",
Endpoint: "http://example.test/v1",
}),
promptkit.WithProfiles(promptkit.Profile{ promptkit.WithProfiles(promptkit.Profile{
ID: "profile", ID: "profile",
Endpoint: "http://example.test/v1", BackendID: "concurrent",
Model: "model", Model: "model",
}), }),
promptkit.WithLLMClient(countingLLMClient{}), promptkit.WithLLMClient(countingLLMClient{}),

219
types.go
View File

@@ -51,7 +51,8 @@ const (
// ValidationPassed means the generated output satisfied its contract. // ValidationPassed means the generated output satisfied its contract.
ValidationPassed ValidationStatus = "passed" ValidationPassed ValidationStatus = "passed"
// ValidationFailed means validation completed and rejected the generated // ValidationFailed means validation completed and rejected the generated
// output. Engine.Run returns this status in a result, not as an error. // output. Engine.Run and Engine.RunPrepared return this status in a result,
// not as an error.
ValidationFailed ValidationStatus = "failed" ValidationFailed ValidationStatus = "failed"
// ValidationSkipped means ValidationNone selected no content check. // ValidationSkipped means ValidationNone selected no content check.
ValidationSkipped ValidationStatus = "skipped" ValidationSkipped ValidationStatus = "skipped"
@@ -78,9 +79,10 @@ const (
// RunRequest selects one prompt execution. It has no stable JSON // RunRequest selects one prompt execution. It has no stable JSON
// representation. // representation.
// //
// Prepare and Run copy the request's maps, pointers, and nested // Prepare, PrepareExecution, and Run copy the request's maps, pointers, and
// JSON-compatible values before using them. The caller may mutate the request // nested JSON-compatible values before using them. The caller may mutate the
// after either method returns. // request after any method returns. A successful PrepareExecution retains its
// own private execution snapshot for RunPrepared.
type RunRequest struct { type RunRequest struct {
// PromptID is the required non-empty prompt identifier. // PromptID is the required non-empty prompt identifier.
PromptID string PromptID string
@@ -91,9 +93,21 @@ type RunRequest struct {
// profile is used; if both are empty, the error matches ErrProfileRequired // profile is used; if both are empty, the error matches ErrProfileRequired
// and ErrInvalidRequest. // and ErrInvalidRequest.
ProfileID string ProfileID string
// SessionID optionally supplies a direct per-run session identifier. A
// nonblank value is trimmed and overrides the prompt definition's
// session_id template. A blank value supplies no direct override. The
// maximum is 256 Unicode code points after trimming. A direct value is
// opaque consumer metadata, not a credential, and may be exposed in
// prepared values, results, collaborator requests, provider requests, and
// provider observability. Callers should use stable, non-sensitive
// identifiers. An overlong direct value makes Prepare, PrepareExecution, or
// Run return an error matching ErrInvalidRequest.
SessionID string
// APIKey is a request-scoped direct credential. It takes precedence over // APIKey is a request-scoped direct credential. It takes precedence over
// APIKeyEnv, is passed to the selected LLMClient, and is never included in // APIKeyEnv, is passed to the selected LLMClient, and is never included in
// prepared values, results, hashes, JSON, String, or GoString output. // prepared values, results, hashes, JSON, String, or GoString output. A
// successful PrepareExecution retains it only in the opaque handle until
// RunPrepared claims the handle or Discard invalidates it.
APIKey string `json:"-"` APIKey string `json:"-"`
// Inputs maps prompt input names to references. A nil or empty map is valid // Inputs maps prompt input names to references. A nil or empty map is valid
// only when the selected prompt and its templates require no inputs. // only when the selected prompt and its templates require no inputs.
@@ -101,17 +115,18 @@ type RunRequest struct {
// Vars supplies Go-template data for messages and the session ID. Nil and // Vars supplies Go-template data for messages and the session ID. Nil and
// empty maps are equivalent. // empty maps are equivalent.
Vars map[string]string Vars map[string]string
// Execution optionally overrides individual profile execution settings. // Execution optionally overrides individual execution settings. Nil uses
// Nil uses the selected profile over framework defaults. // the selected profile over its backend, when any, and framework defaults.
Execution *ExecutionTargetOverride Execution *ExecutionTargetOverride
// Validation optionally replaces the prompt's complete output contract. It // Validation optionally replaces the prompt's complete output contract. It
// does not merge individual fields. Nil uses the prompt contract. // does not merge individual fields. Nil uses the prompt contract.
Validation *OutputContract Validation *OutputContract
} }
// PreparedRun contains prepared prompt execution state. It does not include // PreparedRun contains prepared prompt execution state returned by
// resolved API key values, model output, validation results, or internal target // [Engine.Prepare] or [PreparedExecution.Details]. It does not include resolved
// presence metadata. PreparedRun has a stable JSON representation. // API key values, model output, validation results, or internal target presence
// metadata. PreparedRun has a stable JSON representation.
// //
// All maps, slices, pointers, and schema values are caller-owned copies. JSON // All maps, slices, pointers, and schema values are caller-owned copies. JSON
// timestamps use RFC 3339 and zero timing values are omitted. Hash formats are // timestamps use RFC 3339 and zero timing values are omitted. Hash formats are
@@ -126,8 +141,12 @@ type PreparedRun struct {
// SelectedProfileID is the explicit request profile or prompt default that // SelectedProfileID is the explicit request profile or prompt default that
// supplied execution settings. // supplied execution settings.
SelectedProfileID string `json:"selected_profile_id"` SelectedProfileID string `json:"selected_profile_id"`
// SelectedBackendID equals EffectiveModelParams.BackendID. It is empty for
// an endpoint-only profile.
SelectedBackendID string `json:"selected_backend_id,omitempty"`
// EffectiveModelParams contains framework defaults overlaid by the selected // EffectiveModelParams contains framework defaults overlaid by the selected
// profile and then request overrides. It excludes resolved API-key values. // backend, profile, and then request overrides. It excludes resolved API-key
// values.
EffectiveModelParams ExecutionTarget `json:"effective_model_params"` EffectiveModelParams ExecutionTarget `json:"effective_model_params"`
// OutputContract is the complete effective output contract. // OutputContract is the complete effective output contract.
OutputContract OutputContract `json:"output_contract"` OutputContract OutputContract `json:"output_contract"`
@@ -136,11 +155,12 @@ type PreparedRun struct {
StructuredOutput *StructuredOutputSpec `json:"structured_output,omitempty"` StructuredOutput *StructuredOutputSpec `json:"structured_output,omitempty"`
// InputHashes maps every supplied input name to its opaque artifact hash. // InputHashes maps every supplied input name to its opaque artifact hash.
InputHashes map[string]string `json:"input_hashes,omitempty"` InputHashes map[string]string `json:"input_hashes,omitempty"`
// SessionID is the trimmed rendered session identifier, if any. // SessionID is the effective direct or rendered session identifier, if any.
SessionID string `json:"session_id,omitempty"` SessionID string `json:"session_id,omitempty"`
// RenderedPromptHash is an opaque equality value for SessionID and Messages. // RenderedPromptHash is an opaque equality value for SessionID and Messages.
RenderedPromptHash string `json:"rendered_prompt_hash"` RenderedPromptHash string `json:"rendered_prompt_hash"`
// Messages are the rendered messages that Run passes to the LLM client. // Messages are the rendered messages that Run or RunPrepared passes to the
// LLM client.
Messages []RenderedMessage `json:"messages"` Messages []RenderedMessage `json:"messages"`
// StartTime is the UTC time at which preparation began. // StartTime is the UTC time at which preparation began.
StartTime time.Time `json:"start_time,omitempty"` StartTime time.Time `json:"start_time,omitempty"`
@@ -175,11 +195,17 @@ type RunResult struct {
// PromptHash is the same opaque definition equality value exposed by // PromptHash is the same opaque definition equality value exposed by
// PreparedRun. // PreparedRun.
PromptHash string `json:"prompt_hash,omitempty"` PromptHash string `json:"prompt_hash,omitempty"`
// SessionID is the effective direct or rendered session identifier, if any.
// JSON omits an empty value.
SessionID string `json:"session_id,omitempty"`
// RenderedPromptHash is the same opaque rendered-prompt equality value // RenderedPromptHash is the same opaque rendered-prompt equality value
// computed during preparation. // computed during preparation.
RenderedPromptHash string `json:"rendered_prompt_hash"` RenderedPromptHash string `json:"rendered_prompt_hash"`
// SelectedProfileID identifies the profile used for execution. // SelectedProfileID identifies the profile used for execution.
SelectedProfileID string `json:"selected_profile_id"` SelectedProfileID string `json:"selected_profile_id"`
// SelectedBackendID equals EffectiveModelParams.BackendID. It is empty for
// an endpoint-only profile.
SelectedBackendID string `json:"selected_backend_id,omitempty"`
// ModelName is the effective model name and equals // ModelName is the effective model name and equals
// EffectiveModelParams.Model. // EffectiveModelParams.Model.
ModelName string `json:"model_name"` ModelName string `json:"model_name"`
@@ -194,12 +220,15 @@ type RunResult struct {
InputHashes map[string]string `json:"input_hashes,omitempty"` InputHashes map[string]string `json:"input_hashes,omitempty"`
// Usage is the token accounting reported by the LLM client. // Usage is the token accounting reported by the LLM client.
Usage TokenUsage `json:"usage"` Usage TokenUsage `json:"usage"`
// StartTime is the UTC time immediately before preparation begins. // StartTime is the UTC time immediately before ordinary Run preparation or
// after RunPrepared claims its handle.
StartTime time.Time `json:"start_time,omitempty"` StartTime time.Time `json:"start_time,omitempty"`
// EndTime is the UTC time after generation and validation complete. // EndTime is the UTC time after generation and validation complete.
EndTime time.Time `json:"end_time,omitempty"` EndTime time.Time `json:"end_time,omitempty"`
// Duration covers preparation, generation, and validation. JSON represents // Duration covers preparation, generation, and validation for Run. For
// it as integer milliseconds in duration_ms and omits a zero value. // RunPrepared it covers only the execution attempt after claim and excludes
// preparation and consumer-held delay. JSON represents it as integer
// milliseconds in duration_ms and omits a zero value.
Duration time.Duration `json:"-"` Duration time.Duration `json:"-"`
} }
@@ -239,10 +268,11 @@ type Artifact struct {
// ArtifactReader resolves a prompt input reference into its content. // ArtifactReader resolves a prompt input reference into its content.
// //
// Read may be called concurrently. It must honor ctx cancellation to make // Read may be called concurrently. It must honor ctx cancellation to make
// Prepare and Run responsive to cancellation. The engine passes a copied ref // Prepare, PrepareExecution, and Run responsive to cancellation. The engine
// and immediately copies the returned Artifact.Body; it does not retain either // passes a copied ref and immediately copies the returned Artifact.Body; it
// value. Readers supply artifact metadata, and the engine assigns an input-map // does not retain either value. Readers supply artifact metadata, and the
// name only when the returned artifact name is empty. // engine assigns an input-map name only when the returned artifact name is
// empty.
// //
// An injected reader owns any application-specific path containment, // An injected reader owns any application-specific path containment,
// authorization, content-size, and content-type policy. It must protect // authorization, content-size, and content-type policy. It must protect
@@ -259,6 +289,11 @@ type ArtifactReader interface {
// ExecutionTarget represents effective model runtime settings and has a stable // ExecutionTarget represents effective model runtime settings and has a stable
// JSON representation. It never exposes a resolved API-key value. // JSON representation. It never exposes a resolved API-key value.
type ExecutionTarget struct { type ExecutionTarget struct {
// BackendID is the effective routing identity selected by the profile. It
// remains unchanged when a profile or request overrides Endpoint and is
// empty for endpoint-only profiles. It is supplied to injected LLMClient
// implementations as part of the effective target.
BackendID string `json:"backend_id,omitempty"`
// Endpoint is the model-provider base URL. // Endpoint is the model-provider base URL.
Endpoint string `json:"endpoint"` Endpoint string `json:"endpoint"`
// Model is the provider model identifier. // Model is the provider model identifier.
@@ -276,7 +311,8 @@ type ExecutionTarget struct {
TimeoutSeconds int `json:"timeout_seconds"` TimeoutSeconds int `json:"timeout_seconds"`
// ServiceTier is an optional provider-specific request tier. // ServiceTier is an optional provider-specific request tier.
ServiceTier string `json:"service_tier"` ServiceTier string `json:"service_tier"`
// ReasoningEffort is an optional provider-specific reasoning setting. // ReasoningEffort is the effective opaque provider-specific reasoning
// setting. An empty value instructs model clients to omit reasoning.
ReasoningEffort string `json:"reasoning_effort"` ReasoningEffort string `json:"reasoning_effort"`
// APIKeyEnv is an environment-variable name, not its credential value. // APIKeyEnv is an environment-variable name, not its credential value.
APIKeyEnv string `json:"api_key_env"` APIKeyEnv string `json:"api_key_env"`
@@ -284,16 +320,73 @@ type ExecutionTarget struct {
ExtraParams map[string]any `json:"extra_params"` ExtraParams map[string]any `json:"extra_params"`
} }
// ProfileInspection is the caller-owned result of [Engine.InspectProfile].
// It has no stable JSON representation.
//
// EffectiveModelParams contains a copied effective target. APIKeyRequired is
// separate from that target to preserve ExecutionTarget's general execution
// and stable JSON contracts.
type ProfileInspection struct {
// ProfileID is the trimmed, exact profile ID inspected by the engine.
ProfileID string
// EffectiveModelParams contains framework defaults overlaid by the selected
// backend and then the profile, without a request override. APIKeyEnv is an
// environment-variable name, never its credential value.
EffectiveModelParams ExecutionTarget
// APIKeyRequired reports that a later execution must supply a direct API
// key or an explicit request environment override. It is mutually exclusive
// with a nonblank EffectiveModelParams.APIKeyEnv.
APIKeyRequired bool
}
// PromptInputDefinition describes one declared prompt input.
// It has no stable JSON representation.
type PromptInputDefinition struct {
// Name is the normalized prompt input name.
Name string
// Required is the prompt definition's declared required flag. When true,
// preparation fails if the input is omitted. A false value does not account
// for input references in message or session-ID templates.
Required bool
// ContentType is the declared input media-type metadata.
ContentType string
// Description is the declared human-readable input description.
Description string
}
// PromptInspection is the caller-owned result of [Engine.InspectPrompt].
// It has no stable JSON representation.
//
// Inputs contains copied declared input metadata in definition order.
// OutputContract is the normalized contract declared by the prompt definition,
// rather than a request-level effective override. PromptHash is opaque.
type PromptInspection struct {
// PromptID is the normalized ID of the selected prompt definition.
PromptID string
// PromptVersion is the normalized version of the selected prompt definition.
PromptVersion string
// PromptHash is the opaque equality value for the selected definition.
PromptHash string
// DefaultProfileID is declared metadata and is not resolved by inspection.
DefaultProfileID string
// Inputs contains caller-owned declared input metadata in definition order.
Inputs []PromptInputDefinition
// OutputContract is the normalized contract declared by the definition.
OutputContract OutputContract
}
// ExecutionTargetOverride represents per-request runtime setting overrides and // ExecutionTargetOverride represents per-request runtime setting overrides and
// has no stable JSON representation. // has no stable JSON representation.
// //
// Non-empty string fields replace profile values. Non-nil numeric pointers // Non-empty string fields replace profile and backend values. Non-nil pointer
// replace profile values and preserve explicit zero. A non-empty ExtraParams // fields replace profile values and preserve explicit zero or empty values. A
// map replaces the complete profile map rather than merging keys. Empty string // non-empty ExtraParams map replaces the complete profile or backend map
// fields, nil pointers, and a nil or empty ExtraParams map inherit the selected // rather than merging keys. Empty string fields, nil pointers, and a nil or
// profile over framework defaults. // empty ExtraParams map inherit the selected profile over its backend, when
// any, and framework defaults.
type ExecutionTargetOverride struct { type ExecutionTargetOverride struct {
// Endpoint replaces the profile endpoint when non-empty. // Endpoint replaces the profile or backend endpoint when non-empty without
// changing the effective BackendID.
Endpoint string Endpoint string
// Model replaces the profile model when non-empty. // Model replaces the profile model when non-empty.
Model string Model string
@@ -308,24 +401,29 @@ type ExecutionTargetOverride struct {
TimeoutSeconds *int TimeoutSeconds *int
// ServiceTier replaces the profile value when non-blank. // ServiceTier replaces the profile value when non-blank.
ServiceTier string ServiceTier string
// ReasoningEffort replaces the profile value when non-blank. An empty value // ReasoningEffort controls the per-run reasoning setting. Nil inherits the
// cannot clear a profile setting. // profile value. A pointer to a non-blank string trims and replaces the
ReasoningEffort string // profile value. A pointer to an empty or whitespace-only string clears the
// APIKeyEnv replaces the profile environment-variable name when non-blank. // inherited value and disables reasoning for this run. Non-blank values
// A direct RunRequest.APIKey still takes precedence over environment lookup. // are opaque and are not validated against a fixed vocabulary.
ReasoningEffort *string
// APIKeyEnv replaces the profile or backend environment-variable name when
// non-blank. A direct RunRequest.APIKey still takes precedence over
// environment lookup.
APIKeyEnv string APIKeyEnv string
// ExtraParams, when non-empty, replaces the profile map. Values must be // ExtraParams, when non-empty, replaces the complete profile or backend map.
// JSON-compatible: nil, booleans, finite numbers, strings, arrays or slices, // Values must be JSON-compatible: nil, booleans, finite numbers, strings,
// and maps with non-empty string keys. Cycles are invalid. // arrays or slices, and maps with non-empty string keys. Cycles are invalid.
ExtraParams map[string]any ExtraParams map[string]any
} }
// Profile is an in-memory execution profile for library consumers. // Profile is an in-memory execution profile for library consumers.
// //
// It is equivalent to a loaded profile file after validation. Raw API keys do // It is equivalent to a loaded profile file after validation. Raw API keys do
// not belong in profiles; use APIKeyRequired to require callers to provide // not belong in profiles; use APIKeyRequired to require callers to provide a
// RunRequest.APIKey for each request, or use profile YAML api_key_env with file // RunRequest.APIKey or explicit request ExecutionTargetOverride.APIKeyEnv, or
// and FS profile sources. Profile has no stable JSON representation. // use profile YAML api_key_env with file and FS profile sources. Profile has no
// stable JSON representation.
// //
// WithProfiles validates and copies Profile values during NewEngine. Numeric // WithProfiles validates and copies Profile values during NewEngine. Numeric
// zero, blank strings, and an empty ExtraParams map inherit framework defaults; // zero, blank strings, and an empty ExtraParams map inherit framework defaults;
@@ -333,7 +431,13 @@ type ExecutionTargetOverride struct {
type Profile struct { type Profile struct {
// ID is the required non-blank profile identifier. WithProfiles trims it. // ID is the required non-blank profile identifier. WithProfiles trims it.
ID string ID string
// Endpoint is the required non-blank model-provider base URL. // BackendID optionally selects an engine backend. WithProfiles trims it.
// Backend membership is checked when a request selects the profile; an
// unknown ID makes preparation fail with ErrProfileLoad.
BackendID string
// Endpoint is the model-provider base URL. It is required only when
// BackendID is blank and otherwise overrides the backend endpoint when
// non-blank.
Endpoint string Endpoint string
// Model is the required non-blank provider model identifier. // Model is the required non-blank provider model identifier.
Model string Model string
@@ -351,12 +455,13 @@ type Profile struct {
// ReasoningEffort is optional; a blank value inherits the framework // ReasoningEffort is optional; a blank value inherits the framework
// default. // default.
ReasoningEffort string ReasoningEffort string
// APIKeyRequired requires a non-blank RunRequest.APIKey. It does not store a // APIKeyRequired clears a backend's inherited API-key environment name and
// credential or enable environment lookup. // requires a non-blank RunRequest.APIKey unless the request explicitly
// supplies ExecutionTargetOverride.APIKeyEnv. It does not store a credential.
APIKeyRequired bool APIKeyRequired bool
// ExtraParams contains provider-specific JSON-compatible values. An empty // ExtraParams contains provider-specific JSON-compatible values. An empty
// map inherits framework defaults. WithProfiles validates and deeply copies // map inherits backend request defaults, when any. WithProfiles validates
// it during NewEngine. // and deeply copies it during NewEngine.
ExtraParams map[string]any ExtraParams map[string]any
} }
@@ -364,13 +469,15 @@ type Profile struct {
// profile. // profile.
// //
// It contains ordinary profile fields for OpenAI-compatible chat-completions // It contains ordinary profile fields for OpenAI-compatible chat-completions
// endpoints. APIKeyRequired is satisfied by RunRequest.APIKey. Raw API keys do // endpoints. APIKeyRequired follows Profile.APIKeyRequired. Raw API keys do not
// not belong in this config. OpenAICompatibleProfileConfig has no stable JSON // belong in this config. OpenAICompatibleProfileConfig has no stable JSON
// representation and is not validated until its resulting Profile is supplied // representation and is not validated until its resulting Profile is supplied
// through WithProfiles to NewEngine. // through WithProfiles to NewEngine.
type OpenAICompatibleProfileConfig struct { type OpenAICompatibleProfileConfig struct {
// ID becomes Profile.ID. // ID becomes Profile.ID.
ID string ID string
// BackendID becomes Profile.BackendID.
BackendID string
// Endpoint becomes Profile.Endpoint. // Endpoint becomes Profile.Endpoint.
Endpoint string Endpoint string
// Model becomes Profile.Model. // Model becomes Profile.Model.
@@ -470,8 +577,8 @@ type TokenUsage struct {
// RenderedPrompt is the fully rendered prompt passed to an LLM client and has // RenderedPrompt is the fully rendered prompt passed to an LLM client and has
// a stable JSON representation. // a stable JSON representation.
type RenderedPrompt struct { type RenderedPrompt struct {
// SessionID is the optional trimmed session identifier rendered from the // SessionID is the optional effective direct or rendered session
// prompt definition. // identifier supplied to the model client.
SessionID string `json:"session_id,omitempty"` SessionID string `json:"session_id,omitempty"`
// Messages contains rendered messages in definition order. // Messages contains rendered messages in definition order.
Messages []RenderedMessage `json:"messages"` Messages []RenderedMessage `json:"messages"`
@@ -517,10 +624,14 @@ type StructuredOutputJSONSpec struct {
Schema any `json:"schema"` Schema any `json:"schema"`
} }
// LLMClient executes rendered prompts for [Engine.Run]. // LLMClient executes rendered prompts for [Engine.Run] and
// [Engine.RunPrepared].
// //
// Generate may be called concurrently. It must honor context cancellation to // Generate is scheduled according to the resolved backend's capacity policy.
// make Run responsive to cancellation. The request and all nested maps, // It may still be called concurrently for different backend pools or unlimited
// backends. Cancellation while waiting for capacity can prevent Generate from
// being called. Once invoked, it must honor context cancellation to make Run
// and RunPrepared responsive to cancellation. The request and all nested maps,
// slices, and pointers are client-owned copies and may be mutated or retained // slices, and pointers are client-owned copies and may be mutated or retained
// without affecting engine state. // without affecting engine state.
// //
@@ -529,10 +640,10 @@ type StructuredOutputJSONSpec struct {
// and retained copies. It is responsible for the cancellation behavior of any // and retained copies. It is responsible for the cancellation behavior of any
// work it starts and for synchronizing access to retained or shared data. // work it starts and for synchronizing access to retained or shared data.
// //
// A returned error makes Run return ErrLLMGenerate while preserving the client // A returned error makes Run or RunPrepared return ErrLLMGenerate while
// error through errors.Is. A nil response with a nil error also produces // preserving the client error through errors.Is. A nil response with a nil
// ErrLLMGenerate. Promptkit copies the non-nil response before returning from // error also produces ErrLLMGenerate. Promptkit copies the non-nil response
// Run. // before returning from either method.
type LLMClient interface { type LLMClient interface {
Generate(context.Context, GenerateRequest) (*GenerateResponse, error) Generate(context.Context, GenerateRequest) (*GenerateResponse, error)
} }