47 KiB
Codebase Audit
Sequence Prerequisite
Stage 0 had not been executed when this artifact was created: the repository
contained no audit.md at the start of the Stage 1 review. Consequently, the
reproducible repository baseline, package inventory, validation matrix, graph
refresh, and initial coverage ledger required by Stage 0 remain outstanding.
This review did not backfill that out-of-scope work.
The Stage 1 review began from commit
ebf1602635e108e2a7ac1abd3a3ca24a620104ce on branch main, with a clean
working tree and Go go1.26.5 linux/amd64. These details identify this review
only; they are not a substitute for the Stage 0 baseline.
Stage 1: Public Values, Conversion, Errors, And Formatting
Scope Reviewed
The review covered doc.go, types.go, convert.go, errors.go,
capacity_error.go, formatting.go, and prepared_execution.go, plus the
directly relevant root tests. Internal domain declarations and callers were
consulted only to confirm field-complete conversion, ownership, and public
error mapping. Engine assembly, runtime orchestration, adapter implementation,
and JSON codec mechanics were not audited.
Accepted Findings
S01-F01: Run-request formatting tests do not protect input and variable redaction
- Category: testing
- Severity: medium
- Confidence: high
- Status: accepted
- Affected code:
formatting.go(RunRequest.String,RunRequest.GoString, andRunRequest.redactedString) andengine_test.go(TestRunRequestFormattingRedactsDirectAPIKey) - Contract at issue: The formatter GoDoc promises that
StringandGoStringomit direct credentials and input and variable contents. The documentation and testing policies treat prompt inputs and other private content as sensitive and give consequential disclosure behavior a strong presumption of durable test protection. - Evidence: The implementation currently satisfies the contract by
formatting only request identifiers, collection lengths, presence flags,
and whether an API key is set. The focused test supplies an inline input but
asserts only that the API-key sentinel is absent and
APIKeySet:trueis present; it supplies no variable values and never checks whether the input URI, input body, or variable values appear. The test would remain green if a later formatter change appended input or variable contents while continuing to omit the API key. - Failure mode: A logging or diagnostic-formatting refactor could disclose
prompt input or template-variable content through ordinary
%v,%+v, or%#vformatting without a contract-test failure. - Recommended direction: Extend the existing formatting test, rather than adding a parallel test, with distinct input URI, input body, and variable sentinels and assert that each is absent from all three supported formatting forms. Retain the positive structural assertions so the test continues to distinguish a useful summary from an empty formatter.
- Required verification: Run the focused request-formatting test and the root package tests. Confirm that a deliberate formatter mutation exposing any sentinel makes the focused test fail.
Unresolved Observations
None. Medium- or low-confidence concerns discovered during this review were not promoted to findings.
Coverage Ledger
- Public value declarations and zero values: Reviewed request, prepared, result, artifact, inspection, target, output, validation, rendered-prompt, structured-output, generation, and token-usage values. Nil maps, slices, pointers, optional values, and zero-value public enum strings cross the facade without panics or invented values.
- Request conversion:
toDomainRunRequestand its helpers preserve every public field, copy maps and pointer values, validate and deeply copy nested JSON-compatible overrides, and retain direct credentials only in the internal request field intended for execution. - Prepared and result conversion:
fromDomainPreparedRun,fromDomainRunResult, and their helpers preserve all public fields while copying artifact bytes, validation diagnostics, hashes, rendered messages, cache-control pointers, effective extra parameters, and structured-output schemas. Internal credential and target-presence fields do not escape. - Inspection and extension conversion: Profile and prompt inspection outputs are independent copies. Generation requests receive copied prompt, target, presence, and structured-output values; generation responses contain no mutable fields requiring additional copying.
- Copy-rule ownership: Caller-supplied JSON-compatible values enter through
internal/jsonvaluevalidation and copying. The outward conversion helpers copy already-validated domain snapshots without introducing a second acceptance policy. No consolidation finding was warranted in this scope. - Public errors: Not-found identities remain distinct from load failures;
profile-required and missing-credential errors retain their more specific
identity together with
ErrInvalidRequest; collaborator and cancellation identities remain discoverable; and typed capacity errors expose only a copied backend ID plusErrCapacityExceededrather than the internal error type. Nil and zeroCapacityErrorvalues are safe. - Diagnostic formatting:
RunRequestandGenerateRequestcurrently omit direct credentials and content fromString,GoString,%+v, and%#v.PreparedExecutionalways formats as an opaque constant, including through a copied handle. Prepared and result values intentionally expose rendered or generated content as documented in package GoDoc; applications retain responsibility for logging those content-bearing values. - Prepared handle values: Nil and zero handles return zero details and may be discarded safely. Details are fresh deep copies and remain stable after execution or discard. Copying a handle shares its single-use lifecycle without exposing the internal representation.
- Test ownership: Root external-package tests appropriately own public snapshots, structured error identity, and opaque-handle behavior. Focused internal error-mapping tests cover internal-type containment. The one material redaction gap is recorded as S01-F01.
Verification Performed
The code knowledge graph was used to discover the scoped symbols, trace their callers and callees through the facade and internal domain boundary, and locate the focused tests. Important conclusions were confirmed against source.
The following focused commands passed:
go test . -run 'Test(PreparedRunJSONDoesNotExposeSecretOrTargetPresence|RunRequestFormattingRedactsDirectAPIKey|GenerateRequestFormattingRedactsDirectAPIKey|MapPublicErrorPreservesGenerationCancellation|MapPublicErrorTranslatesCapacityError|CapacityExceededSentinelContract|InspectProfileReturnsIndependentTargetMatchingPreparation|InspectPromptReturnsIndependentMetadataMatchingPreparation|PreparedExecutionFreezesSourcesAndReturnsIndependentDetails|PreparedExecutionLifecycleAndEngineBinding|PreparedExecutionDiscardAndFormattingDoNotExposePrivateState|InMemoryProfileExtraParamsAreCopiedAcrossPublicBoundary|ExtraParamsTypedNestedValuesAreCopiedAcrossPublicBoundary|ArtifactReaderReceivesPublicReferenceAndPreparesArtifact|RunPassesPreparedRequestToInjectedLLMClient|PublicErrorsSupportErrorsIs)$'
go test -race . -run 'Test(PreparedExecutionFreezesSourcesAndReturnsIndependentDetails|PreparedExecutionConcurrentClaimAllowsOneGeneration|PreparedExecutionRunAndDiscardRaceHasOneWinner|PreparedExecutionDiscardAndFormattingDoNotExposePrivateState|MapPublicErrorTranslatesCapacityError|RunRequestFormattingRedactsDirectAPIKey|GenerateRequestFormattingRedactsDirectAPIKey)$' -count=3
Handoff
- Execute the missing Stage 0 baseline before relying on this file as a complete audit ledger or beginning the next component review.
- Stage 2 owns public configuration helpers,
json.go, extension-adapter implementation, and adapter-specific mutation and cancellation behavior. - Stages 3 and 4 own engine construction and runtime operations respectively; this review did not evaluate those paths beyond tracing their use of the scoped conversion and error boundary.
Stage 2: Public Configuration And Extension Adapters
Scope Reviewed
The review covered backends.go, profiles.go, artifact_reader.go,
json.go, and llm_adapter.go, together with directly relevant root tests,
external-package contract tests, and internal/profile validation tests.
Internal backend, profile, artifact, and LLM declarations were consulted only
to compare boundary contracts and policy ownership. NewEngine option
assembly, source composition, runtime orchestration, internal profile-source
behavior, and transport mechanics were not audited.
Accepted Findings
S02-F01: Out-of-range JSON durations silently overflow during decoding
- Category: correctness
- Severity: medium
- Confidence: confirmed
- Status: accepted
- Affected code:
json.go(RunResult.UnmarshalJSONandrunResultJSON.DurationMS) andpublic_contract_test.go(TestRunResultJSONUsesMillisecondsAndRoundTrips) - Contract at issue:
RunResulthas a stable JSON representation in whichduration_msis an integer millisecond count. Decoding must not silently turn an accepted wire value into unrelated duration metadata. - Evidence: The wire field accepts the full
int64range, then decoding multiplies that value bytime.Millisecondwithout checking whether the nanosecond-valuedtime.Durationcan represent the result. A focused probe decoded{"duration_ms":9223372036854775807}with a nil error and producedDuration == -1ms. The existing round-trip test exercises only1500msand does not cover either representable boundaries or overflow. - Failure mode: Malformed or untrusted persisted JSON can be accepted while corrupting a very large positive duration into a negative or otherwise wrapped value. Downstream timing displays, comparisons, or metrics then consume false data without a decode error.
- Recommended direction: Validate the millisecond value against the range
that can be safely converted to
time.Durationbefore multiplication and return a contextual JSON decoding error for values outside that range. - Required verification: Add boundary cases for the largest safely
representable positive and negative millisecond values and their first
out-of-range neighbors, plus the reproduced maximum-
int64input. Retain ordinary and zero-value round-trip coverage.
S02-F02: In-memory and filesystem profiles duplicate semantic validation
- Category: duplication
- Severity: medium
- Confidence: high
- Status: accepted
- Affected code:
profiles.go(validatePublicProfile,toDomainProfile, and the memory profile repository) andinternal/profile/filesystem_repository.go(validateProfileandloadProfile), plus their focused tests - Contract at issue: Profiles supplied in memory and profiles loaded from a filesystem are two sources for the same execution-profile domain value. Required identity, backend-or-endpoint selection, model presence, and numeric bounds are one semantic acceptance policy and need one owner.
- Evidence:
validatePublicProfileandvalidateProfileindependently implement the same seven conditions with the same error text: required ID, backend or endpoint, required model, temperature in[0,2], non-negative maximum tokens, top-p in[0,1], and non-negative timeout. The root tests do not exercise the required-field or scalar-bound cases for in-memory profiles, and the filesystem tests do not protect all scalar bounds. This is policy duplication rather than mere translation or error wrapping. - Failure mode: A future constraint or correction can be applied to one profile source but not the other, making an otherwise identical profile valid or invalid according to where it was stored. Sparse boundary tests would not reliably expose the divergence.
- Recommended direction: Give the domain-level profile acceptance rule one internal owner that both in-memory and filesystem repositories invoke, while leaving source-specific normalization and public error translation at their existing boundaries.
- Required verification: Protect the shared validator with a table covering
every required field and both sides of every numeric bound, then retain a
small integration check for each source and for the public
ErrInvalidConfigtranslation.
S02-F03: The injected LLM client's mutation-ownership contract is untested
- Category: testing
- Severity: medium
- Confidence: high
- Status: accepted
- Affected code:
llm_adapter.go(publicLLMClientAdapter.Generate),convert.go(fromDomainGenerateRequestand its nested conversions),types.go(LLMClient), and injected-client tests inengine_test.go - Contract at issue: The public
LLMClientcontract explicitly states that maps, slices, and pointers inGenerateRequestare client-owned copies that may be mutated or retained. The adapter is the boundary responsible for satisfying that ownership promise. - Evidence: The adapter currently constructs independent messages, cache-control pointers, target parameters, and structured-output schema values before invoking the client. Existing fakes retain requests and tests inspect field propagation, errors, and cancellation, but no test mutates the nested request received by the client and proves that the source domain request remains unchanged. A shallow-copy regression would therefore preserve all current field-equality assertions.
- Failure mode: A conforming injected client could mutate or asynchronously retain nested request data and thereby alter prepared engine state, affect a later operation, or introduce a race despite following the documented interface contract.
- Recommended direction: Add a focused adapter-boundary ownership test that has a client mutate and retain each mutable nested shape, then verifies that the domain request and its nested values remain unchanged. Keep engine-level tests focused on observable request propagation and error identity.
- Required verification: Exercise prompt messages and cache control, target extra parameters, and structured-output schema under the focused test; run it with the race detector as well as normally.
S02-F04: The profile convenience constructor's full mapping is unprotected
- Category: testing
- Severity: medium
- Confidence: high
- Status: accepted
- Affected code:
profiles.go(OpenAICompatibleProfileandOpenAICompatibleProfileConfig) and the threeTestOpenAICompatibleProfile...tests inengine_test.go - Contract at issue: The exported convenience constructor promises a
Profilesuitable for the generalWithProfilespath. Its consumer-visible behavior is the complete, field-for-field mapping of configuration values, followed by the documented deferred validation and copying rules. - Evidence: The implementation currently maps every configuration field. The principal integration test asserts backend ID, model, direct API-key behavior, and extra parameters, while the other tests cover deferred nested parameter validation and ownership. No test protects endpoint, temperature, maximum tokens, top-p, timeout, service tier, or reasoning effort as constructor output. Dropping any of those assignments would leave the current constructor-specific tests green.
- Failure mode: A maintenance edit can silently discard a supported model
setting from the convenience path while the equivalent general
Profileconfiguration continues to work, creating source-dependent behavior for consumers. - Recommended direction: Add one direct, table-like all-field mapping test for the constructor and keep only the integration assertions that establish its passage through ordinary profile validation and ownership boundaries.
- Required verification: Populate every scalar and string setting with a
distinct non-zero value, compare the complete returned
Profile, and retain the existing nested-extra-parameter and invalid-parameter integration cases.
S02-F05: Stable JSON field mappings have multiple manual owners
- Category: duplication
- Severity: medium
- Confidence: high
- Status: accepted
- Affected code:
json.go(PreparedRun.MarshalJSON,runResultJSON,RunResult.MarshalJSON, andRunResult.UnmarshalJSON) and JSON tests inpublic_contract_test.go - Contract at issue: The stable public JSON shape should preserve every public field except for intentional timing representation and omission rules. The list of ordinary fields is one serialization policy, not a separate rule for each encoding direction.
- Evidence:
PreparedRun.MarshalJSONredeclares and assigns every field in an anonymous wire struct.RunResultrepeats its field list in the public type,runResultJSON, the marshal literal, and the unmarshal literal. The custom handling is needed only for timing fields, but ordinary fields are manually synchronized around it. Existing JSON tests cover timing, artifact content type, session ID, and backend omission but do not round-trip fully populated values. Adding a public field to either value can therefore omit it from stable JSON without a compile failure or focused test failure. - Failure mode: Public Go values and their documented stable JSON form can drift, or marshal and unmarshal can become asymmetric, as fields evolve. Consumer data may be silently absent after persistence or interchange.
- Recommended direction: Structure the wire representation so ordinary fields derive from a single alias or embedded representation and only the timing exceptions require explicit mapping. Avoid changing existing JSON names or omission behavior while consolidating ownership.
- Required verification: Add fully populated
PreparedRunandRunResultJSON contract cases that check required names and omissions and compare all fields after round trip, alongside the timing-boundary regression from S02-F01.
Unresolved Observations
None. Questions belonging to engine assembly or internal component behavior were handed to their owning stages rather than promoted from partial traces.
Coverage Ledger
- Backend helpers:
LocalBackendis a side-effect-free conventional-value constructor.WithBackendcopies the queue-capacity pointer when the option is applied, and existing tests protect normalization, invalid and duplicate definitions, nested extra-parameter freezing at construction, lookup copy behavior, and engine isolation. Actual registry composition remains Stage 3 scope and internal registry policy remains Stage 6 scope. - Profile helpers:
OpenAICompatibleProfilecorrectly performs a shallow top-level extra-parameter copy and defers deep validation and freezing to the general profile path as documented. The full-mapping test gap is S02-F04; the duplicated acceptance policy is S02-F02. - Memory profile repository: Repository construction rejects duplicate IDs and invalid nested JSON-compatible values, stores domain copies, and returns independent profile copies. Its source-neutral validation rule lacks a single owner as recorded in S02-F02.
- Artifact reader adapter: The public and internal reader interfaces each
contain only
Read. The adapter passes the caller context and error identity through, rejects a nil successful artifact, translates references without policy duplication, and copies returned body bytes. Focused and integrated tests protect mutation isolation, nil handling, reference translation, cancellation identity, and collaborator error identity. - LLM client adapter: The public and internal client interfaces each
contain only
Generate. The adapter forwards the exact context, preserves client error identity for the use-case boundary, rejects a nil successful response, and translates the scalar response without extra policy. The request conversion currently deep-copies mutable data; its missing mutation regression protection is S02-F03. - Adapter cancellation and errors: Existing public tests establish caller cancellation identity for generation and artifact loading and preserve injected sentinel errors through public wrapping. No adapter adds an independent deadline or cancellation mechanism.
- Stable JSON: Intentional timestamp, millisecond-duration, zero-value, session, artifact, and backend-identity behavior is partly protected. The confirmed overflow is S02-F01 and manual mapping drift is S02-F05.
- Filesystem, reader, and client ownership: Reader and client values are
stored as narrow injected interfaces and mutable values crossing their
adapter calls are copied as described above. Filesystem option validation,
lifetime, and composition reside in
engine.goand are intentionally handed to Stage 3 rather than inferred from this stage's helper review.
Verification Performed
The code knowledge graph was used to locate each scoped helper and adapter, trace its callers and callees, compare public and internal interface widths, and confirm the duplicated profile rule. Source and focused tests were then read to verify the graph conclusions.
The following focused commands passed:
go test . -run 'Test(PublicArtifactReaderAdapterCopiesBody|RunSucceedsWithInjectedLLMClient|RunPassesPreparedRequestToInjectedLLMClient|EngineRunPropagatesCallerCancellation|WithArtifactReaderRejectsNilReader|ArtifactReaderReceivesPublicReferenceAndPreparesArtifact|ArtifactReaderFailuresPreserveArtifactLoadErrors|RunAddsLLMGenerateToCollaboratorPublicError|PublicErrorsSupportErrorsIs|OpenAICompatibleProfileRunsThroughNormalProfilePath|OpenAICompatibleProfileDefersExtraParamsValidation|OpenAICompatibleProfileNestedExtraParamsRunThroughWithProfiles|WithProfilesRejectsDuplicateIDs|WithProfilesRejectsInvalidExtraParams|WithProfilesRejectsCyclicExtraParams|LocalBackendConstructsAndRegistersConventionalBackend|WithBackendCopiesQueueCapacity|BackendRegistrationRejectsInvalidAndDuplicateDefinitions|BackendExtraParamsAreDeeplyCopiedAtConstructionAndLookup|PreparedRunJSONOmitsZeroTimingValues|BackendIdentityJSONNamesAndOmission|PreparedRunJSONTimingRoundTrips|RunResultJSONUsesMillisecondsAndRoundTrips)$'
go test ./internal/profile -run 'Test(FilesystemRepository_GetProfile|FSRepository)$'
go test -race . -run 'Test(EngineRunPropagatesCallerCancellation|ArtifactReaderFailuresPreserveArtifactLoadErrors|PublicArtifactReaderAdapterCopiesBody|OpenAICompatibleProfileNestedExtraParamsRunThroughWithProfiles)$' -count=3
A temporary program outside the repository decoded
{"duration_ms":9223372036854775807} into RunResult; go run reported
error=<nil> duration=-1ms nanoseconds=-1000000, confirming S02-F01. The
temporary source was removed and no probe output was added to the repository.
Handoff
- The Stage 0 baseline remains absent and was not backfilled during this component review.
- Stage 3 owns
NewEngineoption application and the actual composition, validation, and lifetime of configured filesystems, readers, clients, profiles, and backends. - Stage 6 owns internal backend registry and built-in profile policy. Stage 8 owns the broader filesystem profile repository review; it should use S02-F02 as established evidence rather than repeating the public-side audit.
- Stage 14 owns transport-specific request construction, deadlines, response decoding, and resource handling. This stage assessed only the public injection adapter.
Stage 3: Engine Construction, Options, And Source Assembly
Scope Reviewed
The review covered the construction and option portions of engine.go, the
construction effect of WithBackend in backends.go, and directly relevant
tests in engine_test.go, public_contract_test.go, and
capacity_contract_test.go. Narrow traces into backend registry snapshots,
capacity-manager construction, built-in client construction, repositories,
validators, and usecase.NewRunner were used only to confirm the values and
dependencies assembled by NewEngine. Runtime engine methods, repository
parsing mechanics, validation mechanics, transport behavior, and capacity
scheduling were not audited.
Accepted Findings
S03-F01: Single-file source options alter valid caller paths
- Category: correctness
- Severity: medium
- Confidence: confirmed
- Status: accepted
- Affected code:
engine.go(fileSource,WithPromptFile,WithProfileFile, andWithSchemaFile) andTestSourceOptionsRejectInvalidInputsplus the three single-file success tests inengine_test.go - Contract at issue: Each single-file option accepts a path naming an existing non-directory file. Filesystem paths are exact caller values; leading and trailing whitespace are legal filename characters and the option GoDoc does not define normalization.
- Evidence:
fileSourceassignsstrings.TrimSpace(name)tocleanNameand performs every path operation andos.Statagainst that altered value. A focused probe created an existing file namedprompt.yaml, confirmed thatos.Staton the supplied path succeeded, and passed the same value toWithPromptFile.NewEnginereturnedErrInvalidConfigbecause it instead attempted to statprompt.yamlwithout the trailing space. All three public file options share this helper. Existing tests cover ordinary paths, blank paths, one missing path, and one directory path, but no exact-path boundary. - Failure mode: A consumer cannot configure an otherwise valid prompt, profile, or schema file whose name begins or ends with whitespace. The error also reports the altered path, obscuring why the supplied existing file was rejected.
- Recommended direction: Use trimming only to enforce the chosen blank- input rule, then perform path decomposition, validation, error reporting, and filesystem access with the original caller-supplied path.
- Required verification: Add a compact shared regression that constructs engines through all three single-file options using existing paths with a leading or trailing whitespace character. Retain the ordinary missing-file and directory rejection cases.
S03-F02: Construction precedence tests do not isolate documented ordering rules
- Category: testing
- Severity: medium
- Confidence: high
- Status: accepted
- Affected code:
engine.go(Option,NewEngine, andnewProfileRepository) andTestSourceOptionsRejectInvalidInputs,TestRepeatedOptionsUseLastValueInEachCategory,TestInMemoryProfilesOverrideBuiltInsAndProfileSources, andTestFallbackProfileSourcePrecedence - Contract at issue: Option order selects the last valid value within a
category, but profile lookup has a fixed cross-category order independent of
argument order: in-memory, ordinary configured, application fallback, then
built-in. A file or FS ordinary-profile option replaces
Config.ProfileDir, and any invalid option must fail construction even if a later option would replace it. - Evidence: The implementation correctly stores categories separately and
assembles the fixed profile overlay after applying options. The main
precedence tests, however, pass fallback, ordinary, and in-memory options in
the same low-to-high order that a generic order-based overlay would use, so
they would remain green if argument order accidentally began controlling
cross-category precedence. No directly relevant test gives
Config.ProfileDirand an ordinary-profile option colliding IDs to protect the documented replacement, and invalid-option tests do not place a valid replacement after the invalid value. Same-category last-value behavior is well covered but does not protect these distinct rules. - Failure mode: An assembly refactor could make mixed profile-source order
depend on option order, allow
Config.ProfileDirto compete with its replacement option, or silently discard an earlier invalid option. The current tests could still pass while consumers observe different selected profiles or construction success. - Recommended direction: Extend the existing precedence coverage with a
small set of discriminating cases rather than a combinatorial matrix: reverse
the cross-category option order, collide
Config.ProfileDirwith its option replacement, and place a valid same-category option after an invalid one. - Required verification: Assert the selected model for the two profile-
source cases and
errors.Is(err, ErrInvalidConfig)for the invalid-then- valid case. Keep the test at the public construction boundary and avoid assertions about private repository nesting.
Unresolved Observations
None. Lower-confidence concerns about typed-nil interface values and unusual non-regular files were not promoted because the documented Go interface and file contracts do not establish stronger behavior.
Coverage Ledger
- Option application:
NewEngineapplies non-nil options once in argument order and stops on the first error. Nil options compose safely in an option slice. Same-category prompt, ordinary profile, fallback profile, in-memory profile, schema, client, and reader options use last-valid-value semantics; backend registrations alone accumulate. The unprotected ordering edges are recorded as S03-F02. - Required and default dependencies: A nonblank configured prompt directory or prompt-source option is required. Profiles always end with the embedded built-in repository; schema validation defaults to the documented directory; the artifact reader, renderer, and model client receive application-neutral defaults when not injected. Construction performs no provider request and requires no credential.
- Prompt and schema source selection: Prompt and schema FS or file options
replace their corresponding
Configdirectory, retain the injectedfs.FSfor lazy access, and validate nil filesystems and blank roots. Single-file exact-path handling is defective as recorded in S03-F01. Source contents remain lazy and their parsing and containment belong to Stages 7 and 10. - Profile composition:
newProfileRepositorybuilds one explicit overlay in the documented order: built-in, application fallback, one ordinary configured source, then in-memory profiles. Only one ordinary source is installed, and an ordinary option suppressesConfig.ProfileDir. Matching malformed higher-precedence definitions stop lookup rather than becoming failover. The implementation is clear; S03-F02 concerns discriminating test coverage, not current behavior. - Backend and capacity assembly: All consumer backend additions enter one
immutable registry with the built-in backend.
NewEnginetakes one capacity- policy snapshot, constructs a fresh manager, and wraps either the injected or built-in client with that same manager before passing both to the runner. Invalid definitions and capacity policies fail asErrInvalidConfig, and tests protect additive registrations, deep-copy isolation, limited and unlimited behavior, and independence between engines. Registry rules and scheduler mechanics remain Stages 6 and 15 scope. - Caller-owned values and collaborators: Queue-capacity pointers are copied
when
WithBackendis created; backend maps and in-memory profile values are deeply frozen during construction. Injected filesystems, readers, clients, and HTTP transports remain explicit collaborator references. The built-in LLM constructor clones the suppliedhttp.Clientand focused internal tests protect non-mutation for positive, zero, and negative timeouts. - Client and validator selection:
WithLLMClientprevents construction of the built-in client while retaining engine-local capacity wrapping. OtherwiseConfig.Timeoutand a clonedConfig.HTTPClientconfigure the built-in client. Schema options construct the matching validator, while an emptySchemaDiruses the application-neutral default. Transport and validation semantics remain Stages 14 and 10 scope. - Failure atomicity and global state: Every error path returns before an
Engineis published. Construction state is local, registry and capacity values are rebuilt for each engine, and there is no process-global mutable configuration. File-backed prompt, profile, and schema contents are read lazily; malformed or missing source content is classified only when an operation selects it. - Assembly clarity and cost: Construction is a single option pass followed by one repository, registry, manager, client, validator, reader, renderer, and runner assembly. No relevant repeated I/O, parsing, or copying cost was found, and the category flags make replacement and default selection explicit without duplicating internal component policy.
- Test ownership: Root external-package tests appropriately protect public option validity, source selection, profile precedence, copy isolation, default-client configuration, and engine-local backend and capacity behavior. Focused internal tests own registry normalization, manager policy, and HTTP-client cloning. S03-F02 identifies the material missing distinctions rather than recommending duplicate internal choreography tests.
Verification Performed
The code knowledge graph was used to find NewEngine, every option category,
source-assembly helpers, and directly relevant tests; trace construction into
the registry, capacity manager, repositories, validators, model client, and
runner; and confirm that later runtime mechanics were outside the reviewed
path. Important ownership and error conclusions were confirmed against source.
The following focused commands passed:
go test . -run 'Test(NewEngineRejectsMissingPromptDir|NewEngineAcceptsMissingProfileDir|SourceOptionsRejectInvalidInputs|PackageOptionsComposeFromSlice|RepeatedOptionsUseLastValueInEachCategory|FallbackProfileSourcePrecedence|FallbackProfileSourcePreservesLazyLoadingAndErrors|BackendOptionsAccumulateAndRegistrationsAreEngineLocal|BackendRegistrationRejectsInvalidAndDuplicateDefinitions|BackendExtraParamsAreDeeplyCopiedAtConstructionAndLookup|BackendCapacityIsIndependentBetweenEngines|WithLLMClientRejectsNilClient|WithArtifactReaderRejectsNilReader|PromptRepositoryReadFailureMapsToPromptLoad|SelectedProfileRepositoryReadFailureMapsToProfileLoad|PrepareWorksWithPromptFSAndRelativeContentFile|PrepareWorksWithPromptFile|PrepareWorksWithProfileFSOverBuiltIns|PrepareWorksWithProfileFileOverBuiltIns|RunStructuredOutputWorksWithSchemaFS|RunStructuredOutputWorksWithSchemaFile|EngineRunLayersTransportAndGenerationTimeouts)$'
go test ./internal/backend -run 'Test(RegistryIncludesExactOpenRouterDefinition|RegistryNormalizesUniqueAdditionsAndIsolatesMutations|NewRegistryNormalizesCapacityPolicy)$'
go test ./internal/capacity -run 'Test(NewManagerRejectsInvalidPolicies|ManagerAdmissionIsBoundedAndReleaseIsIdempotent|ManagerAdmissionHonorsContextAndUnlimitedBackends)$'
go test ./internal/llm -run 'TestNewOpenAICompatibleClientDoesNotMutateSupplied(Nonzero|Zero)TimeoutClient|TestNewOpenAICompatibleClientTreatsSuppliedNegativeTimeoutAsUnset'
go test -race . -run 'Test(BackendCapacityIsIndependentBetweenEngines|BackendOptionsAccumulateAndRegistrationsAreEngineLocal|EngineSupportsConcurrentPrepareAndRun|RepeatedOptionsUseLastValueInEachCategory)$' -count=3
A temporary program outside the repository created an existing
prompt.yaml file and called NewEngine with WithPromptFile using that
exact path. Direct os.Stat returned nil, while construction returned
ErrInvalidConfig after reporting the trimmed prompt.yaml path, confirming
S03-F01. The temporary source and generated file were removed.
Handoff
- The Stage 0 baseline remains absent and was not backfilled during this component review.
- Stage 4 owns operation entry points and public runtime error/result behavior; construction tests were read only through the behavior needed to observe assembled dependencies.
- Stages 6, 7, 8, 10, 14, and 15 own backend policy, prompt sources, profile
repositories, validators, provider transport, and capacity scheduling
respectively. This stage established only that
NewEngineselects and wires their boundaries consistently.
Stage 4: Engine Operations And Root Contract Coverage
Scope Reviewed
The review covered the public operation portion of engine.go
(InspectPrompt, InspectProfile, Prepare, PrepareExecution, Run, and
RunPrepared) and the directly relevant root tests in engine_test.go,
public_contract_test.go, prepared_execution_contract_test.go, and
errors_internal_test.go. prepared_execution.go and conversion helpers were
consulted only to confirm the operation boundary established in Stage 1.
Internal runner tests were consulted only to identify test ownership; runner,
transport, validation, capacity, and repository mechanics were treated as
black boxes.
Accepted Findings
S04-F01: Ordinary-run cancellation identity is protected only below the public boundary
- Category: testing
- Severity: medium
- Confidence: high
- Status: accepted
- Affected code:
engine.go(Engine.Run),errors.go(mapPublicError),engine_test.go(TestEngineRunPropagatesCallerCancellation),errors_internal_test.go(TestMapPublicErrorPreservesGenerationCancellation), andinternal/usecase/runner_test.go(TestRunnerRunCancellationPreservesGenerationCategory) - Contract at issue:
Engine.Runpasses the caller's context through the execution boundary and preserves the active collaborator's cancellation identity while adding the public operation category. Cancellation and injected dependency failures are consequential public behaviors that should be asserted through the consumer-visible boundary. - Evidence:
Engine.Runcurrently passesctxunchanged to the runner and maps its returned error without dropping wrapped identities. The external- package cancellation test drives a real request context through the built-in HTTP transport but asserts onlyerrors.Is(err, ErrLLMGenerate)after cancellation. Thecontext.Canceledidentity is asserted separately only against the unexportedmapPublicErrorhelper and internal runner. A search of the root operation tests found no other ordinary-run assertion for that identity. Prepared execution does assert both identities at the public boundary, but it exercises a different entry point. - Failure mode: A facade or ordinary-run composition change could replace
the caller context, stop wrapping the collaborator cancellation, or discard
it during public error mapping. The internal tests and existing public test
could all remain green while consumers lose the ability to distinguish
caller cancellation with
errors.Is(err, context.Canceled). - Recommended direction: Extend the existing external-package
TestEngineRunPropagatesCallerCancellationassertion to require bothErrLLMGenerateandcontext.Canceled. Retain the focused internal tests only for the distinct internal translation and runner responsibilities they protect; do not add a parallel end-to-end cancellation test. - Required verification: Run the focused public cancellation test normally and with the race detector. Confirm that deliberately removing caller-context propagation or cancellation wrapping at the facade boundary makes that test fail.
Unresolved Observations
None. The absence of package-level operation convenience functions was
confirmed and is not a consistency defect; the reviewed public API exposes
these operations only as Engine methods.
Coverage Ledger
- Facade shape and request translation: All six methods reject a nil or
uninitialized engine before delegation.
Prepare,PrepareExecution, andRunuse the same field-complete, defensive request conversion and classify conversion failures asErrInvalidRequest; inspections pass their scalar selectors with the documented prompt/profile normalization behavior.RunPreparedunwraps only the opaque handle reference. Each successful facade method delegates once and converts the returned domain snapshot. - Context propagation: Every method passes the supplied context directly to its matching runner operation. Public tests protect cancellation before inspection source work, active ordinary generation, and prepared generation, and prove that a completed preparation is independent of later cancellation of its preparation context. The missing consumer-boundary assertion for the ordinary-run cancellation identity is recorded as S04-F01.
- Inspection operations: Prompt inspection loads declared metadata and referenced content without profile, artifact, schema, capacity, or provider work. Profile inspection resolves the effective target and credential state without prompt or generation work. External-package tests protect nil and blank inputs, not-found versus load identities, cancellation, point-in-time behavior, agreement with preparation, and deep ownership of returned nested values.
- Preparation and ordinary execution:
Preparereturns a caller-owned, credential-redacted prepared snapshot and performs no model generation.Runreturns a caller-owned result after one execution path; content validation failure remains a successful result, while operational failures return no partial result. Root tests protect representative translation, prepared/generated metadata agreement, injected artifact and LLM behavior, validation-result semantics, and the documented public error categories. - Prepared execution boundary:
PrepareExecutionpublishes one opaque, engine-bound handle whose details are independent copies of frozen preparation state.RunPreparedpreserves owner binding, atomic single-use claim behavior, independent execution context, no-result-on-error semantics, collaborator identities, credential revalidation, capacity rejection, and execution-only timing. External-package tests also protect concurrent claim and run/discard behavior. The internal claim, admission, validation, and release mechanisms remain assigned to later component stages. - Public error mapping: Every runner error is routed through one facade
mapping point. Ordinary public identities and injected collaborator errors
remain discoverable with
errors.Is; capacity failures become publicCapacityErrorvalues without leaking the internal type; successful paths do not invent errors. Internal mapping tests appropriately own internal-type containment, while public operation tests own consumer-visible categories and collaborator identities except for S04-F01. - Ownership: Request conversion freezes caller maps, slices, pointers, and JSON-compatible values before internal use. Inspection, prepared, details, and result conversions return fresh mutable values. The prepared-execution contract tests demonstrate that later caller and source mutations do not change frozen execution and that mutating one returned snapshot does not change engine-owned state.
- Convenience-function consistency: The graph and source search found no
package-level
Prepare,PrepareExecution,Run,RunPrepared,InspectPrompt, orInspectProfilefunctions. There is therefore no second operation surface whose translation, errors, or ownership can drift from the methods. - Test ownership and duplication: Root external-package tests protect the exported facade and representative assembled workflows. Focused internal tests own runner coordination and public-error translation mechanics. Some lifecycle and error categories necessarily appear at both levels, but the assertions address different stable boundaries; no removable semantic duplication was found. S04-F01 is the one important identity currently asserted only below the applicable public operation boundary.
- Clarity and cost: The operation facade is a uniform sequence of guard, translation where needed, one delegation, public error mapping, and outward conversion. No duplicated orchestration policy, repeated I/O, avoidable copying, or operation-layer complexity was found. Costs inside preparation, generation, validation, transport, and capacity remain assigned to their owning later stages.
Verification Performed
The code knowledge graph was used to find every public engine operation, confirm their runner and conversion edges, inventory directly relevant root tests, locate the internal cancellation assertions, and verify that no package-level operation convenience functions exist. Important context, error, ownership, and no-partial-result conclusions were confirmed against source.
The following focused commands passed:
go test . -run 'Test(PrepareWorksWithFrameworkContractCorpus|RunSucceedsWithInjectedLLMClient|RunPassesPreparedRequestToInjectedLLMClient|EngineRunPropagatesCallerCancellation|ArtifactReaderFailuresPreserveArtifactLoadErrors|RunAddsLLMGenerateToCollaboratorPublicError|PrepareWithoutProfileMatchesSpecificPublicError|RunValidationFailureReturnsResult|PublicErrorsSupportErrorsIs|InspectProfileResolvesCredentialStatesWithoutPromptOrGeneration|InspectProfilePreservesPublicErrorIdentities|InspectProfileReturnsIndependentTargetMatchingPreparation|InspectPromptReturnsDeclaredMetadataWithoutExecutionWork|InspectPromptPreservesPublicErrorIdentities|InspectPromptReturnsIndependentMetadataMatchingPreparation|PreparedExecutionFreezesSourcesAndReturnsIndependentDetails|PreparedExecutionLifecycleAndEngineBinding|PreparedExecutionConcurrentClaimAllowsOneGeneration|PreparedExecutionRunAndDiscardRaceHasOneWinner|PreparedExecutionDiscardAndFormattingDoNotExposePrivateState|PreparedExecutionCredentialCapacityAndTimingBoundaries)$'
go test -race . -run 'Test(EngineRunPropagatesCallerCancellation|InspectProfilePreservesPublicErrorIdentities|InspectPromptPreservesPublicErrorIdentities|PreparedExecutionFreezesSourcesAndReturnsIndependentDetails|PreparedExecutionLifecycleAndEngineBinding|PreparedExecutionConcurrentClaimAllowsOneGeneration|PreparedExecutionRunAndDiscardRaceHasOneWinner)$' -count=3
go test ./internal/usecase -run 'TestRunnerRunCancellationPreservesGenerationCategory'
go test . -run 'Test(MapPublicErrorPreservesGenerationCancellation|MapPublicErrorTranslatesCapacityError)$'
Handoff
- The Stage 0 baseline remains absent and was not backfilled during this operation-boundary review.
- Stages 7, 8, 10, 11, 13, 14, and 15 own prompt loading, profile loading, validation, runner preparation/execution, prepared-handle lifecycle, provider transport, and capacity mechanics respectively. Those stages should treat the facade behavior recorded here as their outward contract rather than repeat this public-boundary audit.