Files
promptkit/docs/internal/llm.md

3.7 KiB

Internal Model Client

Purpose

This document describes Promptkit's internal model-client implementation. The architecture policy owns the library boundary, and the OpenAI-compatible chat integration owns the observable outbound HTTP contract. The framework format reference owns the profile and prompt settings consumed by the client.

The concrete client remains under internal/llm. The root engine assembles it as the default implementation behind Promptkit's public client boundary.

Components And Flow

Client is the provider-neutral generation boundary consumed by later orchestration. OpenAICompatibleClient is the built-in implementation. It uses internal domain values for rendered prompts, execution targets, structured output, responses, and token usage.

The runner supplies a fully resolved target after applying backend, profile, and request precedence. The client uses its endpoint, credential metadata, generation fields, and extra parameters. BackendID remains routing metadata for the generation boundary and is not mapped into the provider payload.

Construction validates the configured base URL and clones any supplied http.Client so Promptkit can apply its timeout default without mutating the caller's client. Generation then:

  1. validates request-level timeout and endpoint requirements;
  2. maps the internal request into the OpenAI-compatible chat payload;
  3. validates and merges extra parameters;
  4. resolves authentication;
  5. performs the outbound request under the applicable deadlines; and
  6. decodes the first response choice and token usage.

internal/llm owns the set of reserved OpenAI-compatible request fields used when validating extra parameters. Backend registration consumes the same rule without making the model client depend on registry configuration.

The implementation has no retry loop, tool-call support, provider catalog, inbound HTTP behavior, or durable session store.

Prepared Generation

For RunPrepared, the runner supplies the model client with the target, rendered messages, and structured-output constraint retained by executable preparation. Execution does not reopen or rerender consumer sources.

Before backend admission, the runner rechecks that the frozen credential environment-variable name is available. The handle does not retain the environment value; the model client resolves the value visible when generation begins. A direct request key remains in private execution state only until the claimed execution finishes or an unclaimed handle is discarded. Exact public ownership and redaction semantics belong to the PreparedExecution GoDoc.

Failure Categories

The package preserves distinct error identities for invalid client configuration, invalid generation requests, request execution failures, non-success provider statuses, and malformed successful responses. Provider response bodies are not included in non-success errors.

Caller cancellation and deadline failures during the outbound request are reported as request execution failures. The runner classifies these identities without depending on HTTP status mapping.

Test Ownership

The OpenAI-compatible client tests own configuration, client cloning, deterministic deadline precedence, authentication, request and response mapping, malformed data, error identity, cancellation, and response-body suppression. The root transport contract test also verifies that resolved backend settings reach this client without serializing backend identity. All use local test servers or test transports; the default suite makes no live or paid provider requests.