2.3 KiB
Structured LLM Architecture
Scope
This document describes Audita's structured LLM runtime boundary and adapter behavior.
Runtime boundary
Production LLM integration depends on the internal contract only:
contracts.StructuredLLMClientCompleteStructured(ctx, req, out)
Provider SDK types do not leak past this boundary.
Adapter ownership
internal/framework/llm owns the OpenAI-compatible HTTP adapter and shared LLM runtime utilities.
Key responsibilities:
- request assembly;
- timeout/cancellation propagation;
- bounded retry behavior;
- scheduler integration;
- provider response decoding;
- error redaction.
Structured schema registry
Structured response schemas are registered in internal/framework/responseschema and include stable metadata:
idversionnamejson_schemasha256
Current schema keys:
correction_setvalidator_decision_set
Schema metadata is attached to diagnostics through Schema.DiagnosticsMap().
Request shape assumptions
Audita targets OpenAI-compatible chat-completions endpoints and sends structured requests with:
- model;
- chat messages;
response_format.type = json_schema;- schema name and JSON schema payload.
Local validation remains mandatory
Provider schema enforcement is treated as transport-level guardrails.
Audita still validates output locally before applying behavior changes:
- proposal decoding and proposal invariants;
- validator decision decoding and cardinality checks;
- deterministic validation and apply-time rules.
Shared malformed-output policy
Malformed structured-output classification is centralized in internal/framework/structuredoutput.
Proposal generation and validator execution both use this shared classifier so downgrade behavior cannot drift between the two paths.
Secrets and redaction
Secret extraction for LLM redaction is centralized in llm.ConfiguredSecrets(cfg) and reused by proposal and validator diagnostics writers.
Secrets are redacted from:
- diagnostics artifacts;
- report artifacts;
- surfaced adapter/runtime errors.
Concurrency and scheduling
LLM execution is constrained by composed scheduler limits:
- total LLM concurrency;
- proposal LLM concurrency;
- validation LLM concurrency.
The scheduler is FIFO and context-aware so permits are released on success, failure, and cancellation.