Add OpenAI-compatible model client
This commit is contained in:
86
docs/integrations/openai-compatible-chat.md
Normal file
86
docs/integrations/openai-compatible-chat.md
Normal file
@@ -0,0 +1,86 @@
|
||||
# OpenAI-Compatible Chat Integration
|
||||
|
||||
## Purpose
|
||||
|
||||
This document defines the outbound HTTP behavior implemented by Promptkit's
|
||||
internal OpenAI-compatible model client. The
|
||||
[internal model-client document](../internal/llm.md) owns implementation flow,
|
||||
errors, and test ownership. The client is not yet available through a usable
|
||||
public Promptkit engine.
|
||||
|
||||
## Endpoint And Method
|
||||
|
||||
Generation sends an HTTP `POST` with `Content-Type: application/json`.
|
||||
A non-empty endpoint from the execution target overrides the client's
|
||||
configured base URL. After trailing slashes are removed,
|
||||
`/chat/completions` is appended. Generation fails before sending when neither
|
||||
source supplies an endpoint.
|
||||
|
||||
## Authentication
|
||||
|
||||
A non-empty API key supplied directly on the execution target takes
|
||||
precedence. Otherwise, when an API-key environment-variable name is supplied,
|
||||
the client reads that variable and requires a non-empty value. The selected
|
||||
key is sent as `Authorization: Bearer <key>`. No authorization header is sent
|
||||
when neither mechanism is configured.
|
||||
|
||||
## Request Body
|
||||
|
||||
The request body always contains `model` and `messages`. The execution
|
||||
target's model takes precedence over the client's configured model, and one
|
||||
must be available.
|
||||
|
||||
Each ordinary message contains its `role` and string `content`. A
|
||||
cache-controlled message instead uses a text content block containing `type`,
|
||||
`text`, and `cache_control`; an empty cache-control TTL is omitted.
|
||||
|
||||
A non-empty session ID is trimmed, checked against the internal domain limit,
|
||||
and sent as top-level `session_id`. It is not sent as a session header.
|
||||
|
||||
The client conditionally includes:
|
||||
|
||||
- `temperature`, `max_tokens`, and `top_p` when non-zero or explicitly
|
||||
present;
|
||||
- non-empty `service_tier` and `reasoning_effort`; and
|
||||
- `response_format` for JSON Schema structured output, including its name,
|
||||
strict flag, and schema document.
|
||||
|
||||
Extra parameters are merged directly into the top-level body after JSON
|
||||
serialization is verified. Empty keys and collisions with these reserved
|
||||
fields are rejected before any provider call:
|
||||
|
||||
- `model`
|
||||
- `session_id`
|
||||
- `messages`
|
||||
- `temperature`
|
||||
- `max_tokens`
|
||||
- `top_p`
|
||||
- `service_tier`
|
||||
- `reasoning_effort`
|
||||
- `response_format`
|
||||
|
||||
## Response Handling
|
||||
|
||||
Any 2xx response is decoded as an OpenAI-compatible chat response. The client
|
||||
returns the first choice's non-empty message content and maps prompt,
|
||||
completion, total, cached, and cache-write token counts.
|
||||
|
||||
Invalid JSON, absent choices, and empty first-choice content are malformed
|
||||
responses. For a non-2xx status, the error includes the status code but never
|
||||
the provider response body.
|
||||
|
||||
## Timeout And Cancellation
|
||||
|
||||
Timeouts are layered:
|
||||
|
||||
- the caller context remains the outer cancellation boundary;
|
||||
- a positive generation timeout adds a request context deadline;
|
||||
- zero adds no generation-specific deadline;
|
||||
- a negative generation timeout is invalid; and
|
||||
- the cloned `http.Client` supplies the whole-request transport cap, retaining
|
||||
a positive supplied-client timeout or applying the configured/default
|
||||
timeout when the supplied value is not positive.
|
||||
|
||||
The earliest applicable caller, generation, or transport deadline controls the
|
||||
request. Constructing the internal client does not mutate a supplied
|
||||
`http.Client`.
|
||||
53
docs/internal/llm.md
Normal file
53
docs/internal/llm.md
Normal file
@@ -0,0 +1,53 @@
|
||||
# Internal Model Client
|
||||
|
||||
## Purpose
|
||||
|
||||
This document describes Promptkit's internal model-client implementation. The
|
||||
[architecture policy](../policy/architecture.md) owns the library boundary,
|
||||
and the
|
||||
[OpenAI-compatible chat integration](../integrations/openai-compatible-chat.md)
|
||||
owns the observable outbound HTTP contract.
|
||||
|
||||
The client is implemented only under `internal/llm`. The root package does not
|
||||
yet assemble it into a usable public engine.
|
||||
|
||||
## Components And Flow
|
||||
|
||||
`Client` is the provider-neutral generation boundary consumed by later
|
||||
orchestration. `OpenAICompatibleClient` is the built-in implementation. It
|
||||
uses internal domain values for rendered prompts, execution targets,
|
||||
structured output, responses, and token usage.
|
||||
|
||||
Construction validates the configured base URL and clones any supplied
|
||||
`http.Client` so Promptkit can apply its timeout default without mutating the
|
||||
caller's client. Generation then:
|
||||
|
||||
1. validates request-level timeout and endpoint requirements;
|
||||
2. maps the internal request into the OpenAI-compatible chat payload;
|
||||
3. validates and merges extra parameters;
|
||||
4. resolves authentication;
|
||||
5. performs the outbound request under the applicable deadlines; and
|
||||
6. decodes the first response choice and token usage.
|
||||
|
||||
The implementation has no retry loop, tool-call support, provider catalog,
|
||||
inbound HTTP behavior, or durable session store.
|
||||
|
||||
## Failure Categories
|
||||
|
||||
The package preserves distinct error identities for invalid client
|
||||
configuration, invalid generation requests, request execution failures,
|
||||
non-success provider statuses, and malformed successful responses. Provider
|
||||
response bodies are not included in non-success errors.
|
||||
|
||||
Caller cancellation and deadline failures during the outbound request are
|
||||
reported as request execution failures. The future runner can classify these
|
||||
identities without depending on HTTP status mapping.
|
||||
|
||||
## Test Ownership
|
||||
|
||||
The
|
||||
[OpenAI-compatible client tests](../../internal/llm/openai_compatible_client_test.go)
|
||||
own configuration, client cloning, deterministic deadline precedence,
|
||||
authentication, request and response mapping, malformed data, error identity,
|
||||
cancellation, and response-body suppression. They use local test servers and
|
||||
test transports; the default suite makes no live or paid provider requests.
|
||||
@@ -21,10 +21,11 @@ contributor workflow and validation.
|
||||
| `internal/prompt` | Renders prompt messages from Go templates with artifact, variable, session, and cache-control data. | [Go-template renderer](../../internal/prompt/go_renderer.go) |
|
||||
| `internal/artifact` | Resolves ordinary inline and unrestricted caller-selected file references into copied artifacts with metadata and hashes. | [Internal sources and validation](sources.md) |
|
||||
| `internal/validate` | Validates basic, JSON, and JSON Schema output using operating-system filesystem or `fs.FS` schema sources. | [Internal sources and validation](sources.md) |
|
||||
| `internal/llm` | Defines the internal generation boundary and implements outbound OpenAI-compatible chat requests, response decoding, authentication, and deadline handling. | [Internal model client](llm.md) |
|
||||
|
||||
These packages provide the internal model, source, and rendering foundation.
|
||||
Model clients, orchestration, and a usable public engine are not implemented in
|
||||
Promptkit yet.
|
||||
These packages provide the internal model, source, rendering, validation, and
|
||||
model-client foundation. Orchestration and a usable public engine are not
|
||||
implemented in Promptkit yet.
|
||||
|
||||
## Maintenance
|
||||
|
||||
|
||||
@@ -21,7 +21,7 @@ The implemented internal components consist of:
|
||||
- `internal/domain`, which owns framework data values shared by later internal
|
||||
components;
|
||||
- `internal/defaults`, which owns application-neutral framework defaults and
|
||||
constructs the default execution target; and
|
||||
constructs the default execution target;
|
||||
- `internal/filecatalog`, which discovers YAML files and provides source-path
|
||||
helpers for filesystem and `fs.FS` consumers;
|
||||
- `internal/promptdef`, which loads and validates prompt definitions from
|
||||
@@ -32,17 +32,20 @@ The implemented internal components consist of:
|
||||
catalog;
|
||||
- `internal/prompt`, which renders prompt messages from Go templates;
|
||||
- `internal/artifact`, which resolves ordinary inline and unrestricted
|
||||
caller-selected file references; and
|
||||
caller-selected file references;
|
||||
- `internal/validate`, which validates basic, JSON, and JSON Schema output
|
||||
using filesystem and `fs.FS` schema sources.
|
||||
using filesystem and `fs.FS` schema sources; and
|
||||
- `internal/llm`, which defines the provider-neutral generation boundary and
|
||||
implements outbound OpenAI-compatible chat requests.
|
||||
|
||||
The defaults and renderer depend on the domain model. Prompt-definition and
|
||||
profile repositories use the domain model, file catalog, and YAML decoder. The
|
||||
built-in profile repository supplies an embedded `fs.FS` to the profile
|
||||
package. Artifact reading uses the domain model and application-neutral
|
||||
defaults. Validation uses the domain model, file catalog, and JSON Schema
|
||||
implementation. Model clients, orchestration, and the public engine have not
|
||||
yet been extracted.
|
||||
implementation. The model client uses the domain model, application-neutral
|
||||
defaults, and an injected or standard-library HTTP client. Orchestration and
|
||||
the public engine have not yet been extracted.
|
||||
|
||||
Future framework extraction must follow this dependency direction:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user