97 lines
4.7 KiB
Markdown
97 lines
4.7 KiB
Markdown
# Internal Model Client
|
|
|
|
## Purpose
|
|
|
|
This document describes Promptkit's internal model-client implementation. The
|
|
[architecture policy](../policy/architecture.md) owns the library boundary,
|
|
and the
|
|
[OpenAI-compatible chat integration](../integrations/openai-compatible-chat.md)
|
|
owns the observable outbound HTTP contract. The
|
|
[framework format reference](../formats.md) owns the profile and prompt
|
|
settings consumed by the client.
|
|
|
|
The concrete client remains under `internal/llm`. The root engine assembles it
|
|
as the default implementation behind Promptkit's public client boundary.
|
|
|
|
## Components And Flow
|
|
|
|
`Client` is the provider-neutral generation boundary consumed by later
|
|
orchestration. `OpenAICompatibleClient` is the built-in implementation. It
|
|
uses internal domain values for rendered prompts, execution targets,
|
|
structured output, responses, and token usage.
|
|
|
|
The runner supplies a fully resolved target after applying backend, profile,
|
|
and request precedence. The client uses its endpoint, credential metadata,
|
|
generation fields, and extra parameters. `BackendID` remains routing metadata
|
|
for the generation boundary and is not mapped into the provider payload.
|
|
|
|
Construction trims and validates a nonempty configured base URL and clones any
|
|
supplied `http.Client` so Promptkit can apply its timeout default without
|
|
mutating the caller's client. An empty configured base remains valid because a
|
|
resolved request target may supply the endpoint. Generation then:
|
|
|
|
1. validates shared execution-setting invariants and the final selected base
|
|
endpoint;
|
|
2. maps the internal request into the OpenAI-compatible chat payload;
|
|
3. validates and merges extra parameters;
|
|
4. composes `/chat/completions` through parsed URL path operations;
|
|
5. resolves authentication;
|
|
6. performs the outbound request under the applicable deadlines; and
|
|
7. decodes the first response choice and token usage.
|
|
|
|
`internal/llm` owns the set of reserved OpenAI-compatible request fields used
|
|
when validating extra parameters. Backend registration consumes the same rule
|
|
without making the model client depend on registry configuration.
|
|
|
|
The implementation has no retry loop, tool-call support, provider catalog,
|
|
inbound HTTP behavior, or durable session store.
|
|
|
|
## Prepared Generation
|
|
|
|
For [`RunPrepared`](../../engine.go), the runner supplies the model client with
|
|
the target, rendered messages, and structured-output constraint retained by
|
|
executable preparation. Execution does not reopen or rerender consumer
|
|
sources.
|
|
|
|
Before backend admission, the runner rechecks that the frozen credential
|
|
environment-variable name is available. The handle does not retain the
|
|
environment value; the model client resolves the value visible when generation
|
|
begins. A direct request key remains in private execution state only until the
|
|
claimed execution finishes or an unclaimed handle is discarded. Exact public
|
|
ownership and redaction semantics belong to the
|
|
[`PreparedExecution` GoDoc](../../prepared_execution.go).
|
|
|
|
## Failure Categories
|
|
|
|
The package preserves distinct error identities for invalid client
|
|
configuration, invalid generation requests, request execution failures,
|
|
non-success provider statuses, and malformed successful responses. Provider
|
|
response bodies are not included in non-success errors.
|
|
|
|
Invalid nonempty configured endpoints are configuration failures. A missing or
|
|
invalid final selected endpoint is an invalid generation request and is
|
|
rejected before transport.
|
|
|
|
An `http.Client.Do` failure is represented by a redacting multi-cause error:
|
|
the package request-failure sentinel and the exact returned transport error are
|
|
both available through `errors.Is` and `errors.As`, while the rendered text
|
|
does not expose the endpoint, headers, request content, credential, transport
|
|
detail, or provider body. Caller cancellation retains `context.Canceled`;
|
|
caller deadlines, generation deadlines, and whole-request client timeouts
|
|
retain `context.DeadlineExceeded`. The runner adds its generation category
|
|
without discarding those identities or depending on HTTP status mapping.
|
|
|
|
## Test Ownership
|
|
|
|
The
|
|
[OpenAI-compatible client tests](../../internal/llm/openai_compatible_client_test.go)
|
|
own configuration, client cloning, deterministic deadline precedence,
|
|
authentication, request and response mapping, malformed data, error identity,
|
|
cancellation, endpoint selection and composition, pre-transport rejection, and
|
|
response-body suppression. The root
|
|
transport contract tests also verify that resolved backend settings reach this
|
|
client without serializing backend identity and that ordinary-run cancellation
|
|
retains its public generation and context identities. All use local test
|
|
servers or controlled test transports; the default suite makes no live or paid
|
|
provider requests.
|