Files
promptkit/docs/consumers/pkg-promptkit.md

14 KiB

Package promptkit

Purpose

This guide helps Go consumers assemble Promptkit and choose the main preparation or execution workflow. The declarations and GoDoc in the root package own exact field, option, serialization, concurrency, ownership, failure, and cancellation semantics. The framework format reference owns prompt, profile, and schema file contracts.

Import the package as:

import "gitea.maximumdirect.net/eric/promptkit"

The following Go fragments are illustrative and omit surrounding package, import, and error-handling code. Use the maintained examples for complete programs.

Construct An Engine

Create an engine with NewEngine. A directory-backed setup supplies a prompt directory and may supply profile and schema directories:

engine, err := promptkit.NewEngine(promptkit.Config{
	PromptDir:  "prompts",
	ProfileDir: "profiles",
	SchemaDir:  "schemas",
})

Options support single-file or fs.FS sources, in-memory profiles, engine-scoped backends, and injected artifact or model clients. Consult the constructor and option GoDoc for composition, precedence, validation, and default transport behavior. Source discovery, format validation, and profile precedence are defined by the framework format reference.

Prepare Without Model Execution

Engine.Prepare resolves the selected prompt and profile, loads inputs and any structured-output schema, and renders messages without calling a model client. Choose it when the prepared value is the final inspection or persistence result and no later execution must be tied to that exact snapshot:

prepared, err := engine.Prepare(ctx, promptkit.RunRequest{
	PromptID: "meeting.summary",
	Inputs: map[string]promptkit.ArtifactRef{
		"note": promptkit.Inline("Synthetic meeting notes"),
	},
})

The maintained offline preparation example shows a complete runnable setup with a prompt file, in-memory profile, and inline input. Exact request requirements and prepared-result fields belong to the RunRequest and PreparedRun GoDoc.

Prepare Now And Execute The Same Snapshot Later

Use Engine.PrepareExecution when an application must inspect or persist preflight details before deciding whether to start model work, while ensuring that later execution uses those exact rendered messages, inputs, target settings, and validation resources:

preparedExecution, err := engine.PrepareExecution(ctx, promptkit.RunRequest{
	PromptID: "meeting.summary",
	Inputs: map[string]promptkit.ArtifactRef{
		"note": promptkit.Inline("Synthetic meeting notes"),
	},
})
if err != nil {
	return err
}
defer preparedExecution.Discard()

details := preparedExecution.Details()
// Inspect or persist an application-selected safe subset of details.

result, err := engine.RunPrepared(ctx, preparedExecution)

Preparation does not call the model or reserve backend capacity. RunPrepared executes from the retained snapshot rather than reloading consumer sources. The handle is opaque in-process state, while Details contains rendered content and remains subject to the application's data handling policy. The PreparedExecution and method GoDoc and engine operation GoDoc own exact lifecycle, engine-binding, credential, cancellation, timing, and error semantics.

Execute And Validate

Engine.Run performs the same preparation, invokes the configured model client, classifies the generated artifact, and validates the content in one call. Choose it when the application does not need a preflight boundary tied to the eventual execution. A completed content check may return ValidationFailed in the result; an operational inability to validate returns an error.

The maintained offline execution example injects a deterministic model client and exercises Run without credentials, network access, or paid calls. It is intentionally separate from the preparation example so each workflow and its small prompt fixture can be copied and run on its own.

Use the RunResult and ValidationResult GoDoc for the returned data and the Engine.Run GoDoc for failure and cancellation semantics. The OpenAI-compatible integration contract owns the built-in client's outbound HTTP behavior.

Inputs, Profiles, And Overrides

Use File, Inline, or InlineWithURI to construct request inputs. A request can select a profile explicitly or use the prompt's default profile, and can replace execution settings or the complete output contract.

The public value GoDoc defines nil, empty, zero, replacement, copy, and credential behavior. The framework format reference defines how those request values interact with prompt definitions, file-backed profiles, built-ins, schemas, and framework defaults.

For programmatic profiles, OpenAICompatibleProfile converts ordinary OpenAI-compatible settings into a value accepted by WithProfiles.

Inspect A Profile Before Prompt Work

Use Engine.InspectProfile to validate one configured profile without constructing a synthetic prompt or placeholder inputs. It resolves the profile's effective target but does not prepare or execute a prompt:

inspection, err := engine.InspectProfile(ctx, profileID)
if err != nil {
	return err
}

target := inspection.EffectiveModelParams
if target.APIKeyEnv != "" {
	// Apply application policy for the named environment variable.
} else if inspection.APIKeyRequired {
	// Arrange a direct credential before later execution.
}

Use this configuration-time boundary when only the profile and its target need checking. Use Prepare when the application also needs prompt, input, schema, or rendering work; use prepared execution when that work must remain tied to a later execution. Inspection reports credential requirements but leaves the timing of credential enforcement to the application. The method's GoDoc owns its exact result and error contract.

Set A Per-Run Session And Reasoning

Supply a direct session ID when one prompt should be correlated with a consumer-managed conversation or workflow without changing prompt variables:

reasoning := "high"
result, err := engine.Run(ctx, promptkit.RunRequest{
	PromptID: "meeting.summary",
	SessionID: "conversation-42",
	Inputs: map[string]promptkit.ArtifactRef{
		"note": promptkit.Inline("Synthetic meeting notes"),
	},
	Execution: &promptkit.ExecutionTargetOverride{
		ReasoningEffort: &reasoning,
	},
})

A nil reasoning pointer inherits the selected profile, a pointer to a nonblank string replaces it, and a pointer to a blank string disables reasoning for that run. Session IDs are correlation metadata, not credentials; use stable, non-secret values that are safe to expose to collaborators and providers. The RunRequest and ExecutionTargetOverride GoDoc owns the exact normalization, precedence, error, copying, and exposure contract.

Configure A Local OpenAI-Compatible Endpoint

Choose the smallest configuration that fits how the endpoint will be reused.

Use An Endpoint-Only Profile

Put the endpoint directly on an in-memory profile when only that profile needs it and shared backend identity or capacity policy is unnecessary:

engine, err := promptkit.NewEngine(promptkit.Config{
	PromptDir: "prompts",
},
	promptkit.WithProfiles(promptkit.Profile{
		ID:       "local-summary",
		Endpoint: "http://localhost:8000/v1",
		Model:    "example-model",
	}),
)

Endpoint-only profiles have an empty backend ID and remain unrestricted by backend capacity policy.

Use The Conventional Local Backend

Use LocalBackend when profiles should share the conventional local identity, endpoint, and concurrency limit:

engine, err := promptkit.NewEngine(promptkit.Config{
	PromptDir: "prompts",
},
	promptkit.WithBackend(
		promptkit.LocalBackend("http://localhost:8000/v1", 2),
	),
	promptkit.WithProfiles(promptkit.Profile{
		ID:        "local-summary",
		BackendID: promptkit.BackendLocal,
		Model:     "example-model",
	}),
)

The helper is explicit: it does not pre-register a backend or read environment variables. Supplying a positive limit leaves queue capacity omitted, so normal backend registration selects the existing default waiting capacity of 1024. The returned value still enters the engine through WithBackend.

Configure A Complete Backend

Use a keyed Backend value for authentication, extra request parameters, an explicit queue capacity, a custom ID, or multiple local endpoints:

noWaiting := 0
engine, err := promptkit.NewEngine(promptkit.Config{
	PromptDir: "prompts",
},
	promptkit.WithBackend(promptkit.Backend{
		ID:               "local-gpu",
		Endpoint:         "http://gpu-host:8000/v1",
		APIKeyEnv:        "LOCAL_GPU_API_KEY",
		ExtraParams:      map[string]any{"provider_option": "enabled"},
		ConcurrencyLimit: 2,
		QueueCapacity:    &noWaiting,
	}),
	promptkit.WithProfiles(promptkit.Profile{
		ID:        "gpu-summary",
		BackendID: "local-gpu",
		Model:     "example-model",
	}),
)

Use distinct custom IDs when registering multiple local endpoints. Registrations belong to one engine and custom IDs cannot replace built-ins. The Backend, LocalBackend, and WithBackend GoDoc defines exact construction, validation, copying, uniqueness, concurrency, and request-default behavior.

Both file-backed and in-memory profiles select a registration through backend or Profile.BackendID. Profile and request endpoint overrides retain that routing and capacity identity. PreparedRun.SelectedBackendID, RunResult.SelectedBackendID, and the effective ExecutionTarget.BackendID expose it to consumers and injected model clients. Endpoint-only profiles remain supported and expose an empty backend ID.

Limit Backend Concurrency

Set Backend.ConcurrencyLimit when a backend needs protection from too many simultaneous model calls. Leaving QueueCapacity nil, as in the local-backend example above, selects the default waiting capacity of 1024.

To accept no waiting backlog beyond the active calls, provide an explicit zero:

noWaiting := 0
backend := promptkit.Backend{
	ID:               "local-gpu",
	Endpoint:         "http://gpu-host:8000/v1",
	ConcurrencyLimit: 2,
	QueueCapacity:    &noWaiting,
}

The pointer distinguishes an explicit zero from omission. Keep using keyed Backend literals so additive configuration fields remain source-compatible. Capacity belongs to one engine and the selected backend ID; endpoint-only profiles and custom backends without a configured limit remain unrestricted. Exact validation, defaulting, ownership, and concurrency semantics belong to the Backend GoDoc.

Credentials

File-backed profiles name an environment variable; in-memory profiles can require a direct request key. Direct keys are request-scoped and are excluded from supported JSON values and the package's String and GoString summaries. The exact precedence and redaction guarantees belong to RunRequest, GenerateRequest, and the profile GoDoc.

Protect Files And Generated Data

The default artifact reader opens a File reference as a caller-selected operating-system path. It does not constrain paths to an application root, impose an inbound request-size policy, or establish an untrusted-input security boundary. Applications must validate and restrict untrusted paths and payloads before constructing a request, or inject an artifact reader that enforces their filesystem, authorization, and size policies.

Rendered messages, input and output artifact bodies, raw model output, and validation diagnostics can contain sensitive data. API-key redaction does not sanitize those values. Treat prepared values, results, collaborator requests, errors, and logs according to the application's data-access, retention, and secret-handling policies.

Extension Interfaces

Inject an LLMClient or ArtifactReader when the built-in behavior does not fit the application. Their GoDoc defines concurrent use, context handling, ownership of copied values, nil responses, and preservation of collaborator errors. Implementations must honor cancellation, safely manage copies they retain, avoid unsafe logging of content or credentials, and enforce the application policy that motivated the injection.

Handle Errors

Use errors.Is with the public error sentinels and operation GoDoc. The declarations distinguish invalid construction, invalid requests, absent sources, source-loading failures, collaborator failures, and operational validation failures. Specific request conditions may also match the broader ErrInvalidRequest, and injected collaborator identities are preserved where documented. Invalid or duplicate backend registrations match ErrInvalidConfig; selecting an unknown backend matches ErrProfileLoad.

When a limited backend has admitted all active and waiting calls, handle ErrCapacityExceeded separately from request errors and provider failures:

result, err := engine.Run(ctx, request)
if errors.Is(err, promptkit.ErrCapacityExceeded) {
	// Apply application policy: shed work, report overload, or retry later.
}

A rejected call returns no partial result and does not invoke the model client. Promptkit does not prescribe retries or map this error to an HTTP status; those choices remain with the consuming application. The Engine.Run and error GoDoc owns exact error and cancellation identities.

Application Boundary

Promptkit is an importable library. It does not own a command, inbound HTTP API, process configuration, or deployment policy. Applications map the root package's results and errors into those concerns, including inbound size and trust policy.