diff --git a/docs/operations.md b/docs/operations.md index 967be96..6223d7f 100644 --- a/docs/operations.md +++ b/docs/operations.md @@ -19,6 +19,8 @@ persists the module snapshot and prompt data package, records Promptkit preparation provenance before provider execution, then persists raw output and execution provenance, validates the structured generated text, and renders the managed Markdown report from the validated text and deterministic values. +The current receipts are transitional workspace state, not a cross-version +profile-provenance contract. The managed report and its final metadata are saved before single-report Distributor notification is attempted. `--out` writes an extra operator copy; @@ -111,8 +113,11 @@ The generated-text and render-context artifacts are written for every completed single-report generation. A report's metadata links the module snapshot, data package, preparation and execution receipts, managed report, generated-text artifacts, and any available single-report -notification artifact. Batch notification artifacts are separate batch-level -records under `notifications/batches`. +notification artifact. These current-version receipts remain transitional; use +the active command's classified error and explicit secure debug capture for +prompt diagnosis rather than relying on them as a durable interface. Batch +notification artifacts are separate batch-level records under +`notifications/batches`. RunIDs begin with the UTC generation timestamp and report ID. A Daily RunID also contains its local valid date so multiple Daily reports in one batch have @@ -186,10 +191,12 @@ remain available where they can be safely persisted. - A batch notification failure preserves each report's artifacts and adds the top-level batch notification artifact. -Use the RunID from the action summary with the inspection commands above. For -a batch failure, inspect the summary first, then inspect the affected report -RunIDs or the batch notification path. Do not remove the whole workspace as a -first response; retain it until the failure is understood. +Use the action summary and its classified error first. For prompt or provider +diagnosis, prefer an explicitly enabled secure debug capture; current-version +receipt paths may provide supplemental context when available. For a batch +failure, inspect the summary first, then inspect the affected report RunIDs or +the batch notification path. Do not remove the whole workspace as a first +response; retain it until the failure is understood. ## Operational Caveats diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md index 2bd37c3..271b2c8 100644 --- a/docs/roadmap/implementation.md +++ b/docs/roadmap/implementation.md @@ -1,14 +1,15 @@ # Domain-Specific Prompt Profiles Implementation Plan -Status: Completed. +Status: Stages 1–7 completed; remediation Stage 8 ready. ## Purpose And Authority -This document records the completed implementation of the +This document records the implementation and post-implementation remediation +of the [domain-specific prompt profiles roadmap](domain-profiles.md). The roadmap is -authoritative for scope, user intent, policy choices, and the implemented end -state. This plan records the implementation sequence, verification, and exit -gates used to reach it. +authoritative for scope, user intent, policy choices, and the intended end +state. This plan records implementation sequence, verification, audit findings, +and exit gates. This plan follows the repository's [architecture](../policy/architecture.md), @@ -361,6 +362,236 @@ the working tree contains only intentional changes, and the canonical documentation describes the implemented state. The feature is ready for code review and release preparation. +## Post-Implementation Review + +Stages 1–6 implemented the intended production behavior and passed their +offline verification gates. A subsequent review found no high-severity runtime +defect, but identified three test-quality issues and one remaining validation +obligation: + +- one app test asserted durable preparation and execution artifact provenance, + contrary to the active-execution boundary and accepted ephemeral-state + direction; +- an adapter-package test depended upward on app orchestration and duplicated + test ownership; +- embedded fallback profiles were inspected but not exercised through one + prepared execution with a provider fake; and +- the roadmap's representative model-evaluation policy had no recorded + evidence. + +Stages 7 and 8 address those findings without changing the profile catalog, +selection precedence, report assignments, prompt content, generated-text +schemas, or default offline test contract. + +## Stage 7: Correct Test Ownership And Fallback Execution Coverage + +### Goal + +Remove accidental durable-state and cross-layer test commitments while adding +one focused offline execution test for the embedded fallback path. + +### Work + +1. Rewrite `TestGenerateDetailedPreservesSelectedProfileThroughExecution` so + it protects active workflow behavior only: + + - retain the Hourly default, day-scale default, and global-override cases; + - assert the profile ID sent in `promptexec.ExecuteRequest`; + - have the executor fake record the preparation and execution values it + emits, then assert their logical profile ID and effective backend/model; + - do not load preparation, execution, or metadata files to establish a + durable profile-provenance contract; and + - remove artifact-content scans whose fake inputs cannot contain an endpoint + or credential. + +2. Preserve meaningful safety coverage at the boundary that can expose the + sensitive value: + + - retain adapter mapping coverage proving an endpoint from a real Promptkit + profile does not enter `promptexec.ProfileInspection`; + - retain app error coverage proving dependency errors containing an endpoint + or credential are replaced by a bounded classified error; and + - do not add profile endpoints or credentials to project-owned execution + types merely to make a leakage test possible. + +3. Remove `internal/app`, app configuration, and report-registry dependencies + from `internal/adapters/promptkit/adapter_test.go`. Move the assembled + application-preflight test to a new app-owned external integration test, + such as `internal/app/prompt_profile_integration_test.go` with package + `app_test`: + + - construct the real Promptkit adapter through its public `New` function; + - call the public app prompt-inspection operation; + - supply a deterministic credential lookup rather than reading the process + environment; and + - cover Hourly, one representative day-scale default, the explicit + `weather-deep` global override, and a same-ID endpoint-only + `weather-light` override. The asset contract tests already own the exact + mapping for all four prompts, so the integration test need not repeat all + four. + +4. Add one adapter-owned, offline fake-client execution test using the real + embedded Hourly prompt at `1.1.0` and selected profile `weather-light`. + Execute through the normal prepared adapter path and assert: + + - the preparation callback runs before the fake provider; + - preparation and execution report logical profile `weather-light`, backend + `openrouter`, and model `deepseek/deepseek-v4-flash`; + - the fake provider request targets `deepseek/deepseek-v4-flash`; and + - schema validation completes without contacting a live service. + + One execution case is sufficient because Promptkit owns uniform source + precedence and the adapter's inspection tests already cover fallback, + operator file, operator directory, built-in, and explicit in-memory layers. + +5. Reconcile the profile-related current-state documentation: + + - it may accurately describe fields present in current preparation and + execution receipts; + - it must not promise cross-version readability or characterize those + receipts as the profile feature's durable target architecture; and + - troubleshooting should prefer active command errors and explicit secure + debug capture, mentioning current-version receipts only as transitional + state if they remain useful before the ephemeral-state refactor. + +6. Do not change production profile resolution, prompt definitions, state + schemas, artifact validators, or the ephemeral-state roadmap in this stage. + +### Tests + +Run: + +```sh +go test -count=1 ./internal/promptassets ./internal/adapters/promptkit ./internal/app ./internal/cli +go test -count=1 -race ./internal/promptassets ./internal/adapters/promptkit ./internal/app ./internal/cli +go test -count=1 ./... +go vet ./... +go run ./cmd/weatherreporter --help +git diff --check +``` + +Review the changed tests against the testing policy and confirm that adapter +tests own adapter behavior, app tests own orchestration, and state tests remain +the sole owner of durable artifact format and validation details. + +### Exit Gate + +Active profile selection and effective-model propagation remain protected +without adding a durable-provenance commitment; the adapter test package no +longer imports the app layer; one embedded fallback profile completes prepared +execution through a provider fake; and every required check passes offline. + +## Stage 8: Evaluate The Initial Model Ladder + +### Goal + +Produce explicit release-candidate evidence that the selected models are +acceptable for their intended report tiers and that a representative local +override provides the promised operator experience. + +This is an opt-in evaluation stage, not an ordinary automated-test stage. It +requires operator-approved provider credentials, network access, and a local +OpenAI-compatible endpoint. Do not mark it complete when those prerequisites +are unavailable; report the missing prerequisite instead. + +### Corpus + +Use four representative, secret-free YAML data packages: one each for Daily, +Today, Tomorrow, and Hourly. The set must include at least one package with +precipitation windows and at least one with none. Remove precise private +location identifiers or other operationally sensitive values without changing +the meteorological relationships being evaluated. + +Record a SHA-256 hash and a short, non-sensitive description for each package. +Do not commit full packages or generated prose unless the user separately +approves them as repository fixtures. + +### Execution Matrix + +Run these six evaluations from the exact package bytes: + +| Case | Package | Profile | +| --- | --- | --- | +| Hourly default | Hourly | `weather-light` | +| Daily default | Daily | `weather-balanced` | +| Today default | Today | `weather-balanced` | +| Tomorrow default | Tomorrow | `weather-balanced` | +| Deep comparison | The same Daily package used above | `weather-deep` | +| Local override | The same Hourly package used above | operator-defined `weather-light` endpoint profile | + +After the successful local-override case, stop or deliberately address an +unavailable test endpoint and repeat it as a negative control. Confirm that the +request fails visibly and does not call or select an embedded remote profile. +This negative control is not an additional quality-evaluation case. + +Use a temporary, untracked evaluation harness beneath the module when exact +package replay is needed. It should call the existing Promptkit adapter and +project-owned execution contract rather than duplicate prompt loading, +rendering, or schema validation. Remove the harness and all unapproved raw +outputs before completing the stage. Never print or record credentials. + +### Evaluation Record + +Add a concise `## Evaluation Record` section to +`docs/roadmap/domain-profiles.md`. For every case, record: + +- evaluation date, logical profile, effective backend, and exact model + reported by execution; +- corpus hash, validation outcome, latency, prompt/completion/total token use, + and provider-reported or contemporaneously calculated cost; +- whether every generated claim is supported by the deterministic package; +- whether hazards, periods, uncertainty, and precipitation timing are used + correctly; +- whether `precipitation_timing` is exactly an empty string for the no-window + case; +- a short usefulness assessment for summary and forecast discussion; and +- any provider, alias, or local-endpoint caveat observed. + +Do not include credentials, endpoints, complete effective parameter maps, +full data packages, rendered prompts, or full generated responses in the +record. The secure debug directory may be used temporarily for operator review +and remains operator-managed. + +### Acceptance Rules + +- Every case must complete strict JSON Schema validation without repair. +- Generated prose must contain no material unsupported weather claim or + contradiction of deterministic hazards, periods, or uncertainty. +- Precipitation timing must agree with the deterministic windows and use the + required empty-string representation when no window exists. +- The local override must select the operator model without modifying a prompt + or application code and must not fall back to a remote profile when the local + endpoint is unavailable. +- Latency, tokens, and cost must be recorded, but this initial evaluation does + not impose an invented numeric threshold. The operator decides whether the + observed tradeoff remains acceptable for the named tier. +- If a default case fails schema or factual acceptance, do not weaken the + schema or prompt to accommodate the model. Reopen the concrete model or + profile-setting decision in the feature roadmap and leave this stage + incomplete. + +### Verification + +After removing temporary evaluation material, run: + +```sh +go test -count=1 ./... +git diff --check +git status --short +``` + +Confirm that the only intended repository change from this stage is the +concise evaluation record and any roadmap status correction required by its +result. Do not add live credentials, provider-dependent tests, a permanent +benchmark framework, or release notes before a release version is selected. + +### Exit Gate + +All six cases satisfy the acceptance rules, the roadmap contains concise and +safe evaluation evidence, no temporary corpus or response material remains in +the repository, and the default suite remains offline. Set this plan back to +`Status: Completed` only after both Stages 7 and 8 have passed. + ## Open Questions None. The model identifiers, profile settings, report assignments, version diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 163c8b4..9b39aae 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -1,7 +1,10 @@ # Troubleshooting -Keep failed workspace artifacts in place. When a RunID is available, start -with `weatherreporter inspect metadata RUN_ID` and use the paths in its result. +Start with the command's classified error. When content-rich prompt diagnostics +are needed, enable a new run with `--llm-debug-dir` and handle the resulting +secure capture as sensitive. Current-version workspace receipts can provide +additional context when present, but are transitional state rather than a +long-term troubleshooting interface. ## Prompt inspection or credentials fail before collection @@ -20,9 +23,10 @@ the selected profile's YAML, ID, backend or endpoint, and model. If the model is unexpected, first check the global `promptkit.profile` selection and then look for a same-ID definition in the configured file or directory. -The preparation and execution receipts named by run metadata retain the -selected profile ID and effective backend/model for diagnosis, but not an -endpoint or credential. See the maintained +Current-version preparation and execution receipts may retain the selected +profile ID and effective backend/model, but not an endpoint or credential. +Use them only as supplemental context after the active command error or an +explicit secure debug capture. See the maintained [local `weather-light` profile example](../examples/weather-light-local-profile.yml). ## Local model endpoint is unavailable @@ -36,10 +40,11 @@ does not probe endpoints or automatically use a remote profile instead. ## Preparation, capacity, or execution fails A preparation failure occurs before provider work; an execution failure occurs -after preparation. Both leave safe provenance and metadata when reached. A -capacity error for one batch report does not retry that report or prevent later -independent reports. Inspect the preparation or execution path, correct the -profile/backend condition, and create a new run. See [operations](operations.md). +after preparation. A capacity error for one batch report does not retry that +report or prevent later independent reports. Correct the profile or backend +condition identified by the bounded command error, then create a new run. +Use explicit secure debug capture only when additional content-rich diagnostics +are necessary. See [operations](operations.md). ## Generated text fails validation diff --git a/internal/adapters/promptkit/adapter_test.go b/internal/adapters/promptkit/adapter_test.go index bd9535c..e502068 100644 --- a/internal/adapters/promptkit/adapter_test.go +++ b/internal/adapters/promptkit/adapter_test.go @@ -12,10 +12,7 @@ import ( "time" promptkit "gitea.maximumdirect.net/eric/promptkit" - "gitea.maximumdirect.net/eric/weatherreporter/internal/app" - appconfig "gitea.maximumdirect.net/eric/weatherreporter/internal/config" "gitea.maximumdirect.net/eric/weatherreporter/internal/promptexec" - "gitea.maximumdirect.net/eric/weatherreporter/internal/report" ) type fakeClient struct { @@ -161,44 +158,6 @@ func TestMaintainedWeatherLightLocalProfileExampleInspectsOffline(t *testing.T) assertProfile(t, adapter, "weather-light", "", "weather-local") } -func TestApplicationPreflightResolvesEmbeddedAndOverriddenProfilesOffline(t *testing.T) { - lookupEnv := func(string) (string, bool) { return "test-key", true } - inspect := func(t *testing.T, adapter *Adapter, id report.ID, profile string, wantID string, wantBackend string, wantModel string) { - t.Helper() - result, err := app.InspectPromptExecution(context.Background(), app.PromptInspectionRequest{ - Resolved: resolvedPromptProfile(t, id), - Executor: adapter, - Promptkit: appconfig.PromptkitConfig{Profile: profile}, - LookupEnv: lookupEnv, - }) - if err != nil { - t.Fatalf("InspectPromptExecution() error = %v", err) - } - if result.ProfileID != wantID || result.BackendID != wantBackend || result.ModelName != wantModel { - t.Fatalf("inspection = %#v, want profile/backend/model %q/%q/%q", result, wantID, wantBackend, wantModel) - } - } - - embedded, err := newAdapterForTest(Config{}, &fakeClient{}) - if err != nil { - t.Fatalf("newAdapterForTest(embedded) error = %v", err) - } - inspect(t, embedded, report.Hourly, "", "weather-light", "openrouter", "deepseek/deepseek-v4-flash") - inspect(t, embedded, report.Daily, "", "weather-balanced", "openrouter", "~google/gemini-flash-latest") - inspect(t, embedded, report.Today, "", "weather-balanced", "openrouter", "~google/gemini-flash-latest") - inspect(t, embedded, report.Tomorrow, "", "weather-balanced", "openrouter", "~google/gemini-flash-latest") - inspect(t, embedded, report.Daily, "weather-deep", "weather-deep", "openrouter", "~anthropic/claude-sonnet-latest") - - override, err := newAdapterForTest(Config{ProfileFile: writeProfileFile(t, `id: weather-light -endpoint: https://local.example/v1 -model: local-weather -`)}, &fakeClient{}) - if err != nil { - t.Fatalf("newAdapterForTest(override) error = %v", err) - } - inspect(t, override, report.Hourly, "", "weather-light", "", "local-weather") -} - func TestProfileResolutionFallsThroughOnlyWhenTheConfiguredIDIsAbsent(t *testing.T) { absentAdapter, err := New(Config{ProfileDirectory: testProfileDirectory(t, `id: other-profile backend: openrouter @@ -283,6 +242,44 @@ func TestExecuteUsesPreparedInlineDataPackage(t *testing.T) { } } +func TestExecuteEmbeddedHourlyProfileThroughPreparedPath(t *testing.T) { + t.Setenv("OPENROUTER_API_KEY", "test-openrouter-key") + client := &fakeClient{response: hourlyValidResponse()} + adapter, err := newAdapter(Config{}, promptkit.WithLLMClient(client)) + if err != nil { + t.Fatalf("newAdapter() error = %v", err) + } + request := promptexec.ExecuteRequest{ + PromptID: "weather.hourly_generated_text", + PromptVersion: "1.1.0", + ProfileID: "weather-light", + DataPackage: []byte("report:\n id: hourly\nbriefing: {}\n"), + DataPackagePath: "data-packages/hourly/data_package.yaml", + } + var preparation promptexec.Preparation + prepared := false + result, err := adapter.Execute(context.Background(), request, func(value promptexec.Preparation, _ *promptexec.PreparationDebug) error { + if client.callCount() != 0 { + t.Fatal("provider was called before preparation completed") + } + preparation = value + prepared = true + return nil + }) + if err != nil { + t.Fatalf("Execute() error = %v", err) + } + if !prepared || preparation.ProfileID != "weather-light" || preparation.BackendID != "openrouter" || preparation.ModelName != "deepseek/deepseek-v4-flash" { + t.Fatalf("preparation = %#v", preparation) + } + if result == nil || result.ProfileID != "weather-light" || result.BackendID != "openrouter" || result.ModelName != "deepseek/deepseek-v4-flash" || result.Validation.Status != promptexec.ValidationPassed { + t.Fatalf("execution = %#v", result) + } + if client.callCount() != 1 || client.request().Target.Model != "deepseek/deepseek-v4-flash" { + t.Fatalf("provider calls/request = %d/%#v", client.callCount(), client.request()) + } +} + func TestExecuteUsesExactInlineDataPackageProvenance(t *testing.T) { client := &fakeClient{response: validResponse()} reader := &recordingReader{} @@ -528,20 +525,6 @@ func writeProfileFile(t *testing.T, profile string) string { return path } -func resolvedPromptProfile(t *testing.T, id report.ID) report.Resolved { - t.Helper() - now := time.Date(2026, 5, 29, 12, 0, 0, 0, time.UTC) - request := report.ResolveRequest{Now: now, Location: time.UTC} - if id == report.Daily { - request.Date = now - } - resolved, err := report.DefaultRegistry().Resolve(id, request) - if err != nil { - t.Fatalf("Resolve(%q) error = %v", id, err) - } - return resolved -} - func testExecuteRequest() promptexec.ExecuteRequest { return promptexec.ExecuteRequest{ PromptID: "weather.daily_generated_text", @@ -558,3 +541,10 @@ func validResponse() *promptkit.GenerateResponse { Usage: promptkit.TokenUsage{PromptTokens: 12, CompletionTokens: 8, TotalTokens: 20}, } } + +func hourlyValidResponse() *promptkit.GenerateResponse { + return &promptkit.GenerateResponse{ + Content: `{"summary":"A quiet hour is expected.","forecast_discussion":"Conditions remain settled.","precipitation_timing":""}`, + Usage: promptkit.TokenUsage{PromptTokens: 12, CompletionTokens: 8, TotalTokens: 20}, + } +} diff --git a/internal/app/prompt_profile_integration_test.go b/internal/app/prompt_profile_integration_test.go new file mode 100644 index 0000000..6ed6b54 --- /dev/null +++ b/internal/app/prompt_profile_integration_test.go @@ -0,0 +1,73 @@ +package app_test + +import ( + "context" + "os" + "path/filepath" + "testing" + "time" + + promptkitadapter "gitea.maximumdirect.net/eric/weatherreporter/internal/adapters/promptkit" + "gitea.maximumdirect.net/eric/weatherreporter/internal/app" + "gitea.maximumdirect.net/eric/weatherreporter/internal/config" + "gitea.maximumdirect.net/eric/weatherreporter/internal/report" +) + +func TestPromptInspectionResolvesEmbeddedAndOverriddenProfilesOffline(t *testing.T) { + lookupEnv := func(string) (string, bool) { return "test-key", true } + inspect := func(t *testing.T, adapter *promptkitadapter.Adapter, id report.ID, profile string, wantID string, wantBackend string, wantModel string) { + t.Helper() + result, err := app.InspectPromptExecution(context.Background(), app.PromptInspectionRequest{ + Resolved: resolvedPromptProfile(t, id), + Executor: adapter, + Promptkit: config.PromptkitConfig{Profile: profile}, + LookupEnv: lookupEnv, + }) + if err != nil { + t.Fatalf("InspectPromptExecution() error = %v", err) + } + if result.ProfileID != wantID || result.BackendID != wantBackend || result.ModelName != wantModel { + t.Fatalf("inspection = %#v, want profile/backend/model %q/%q/%q", result, wantID, wantBackend, wantModel) + } + } + + embedded, err := promptkitadapter.New(promptkitadapter.Config{}) + if err != nil { + t.Fatalf("New(embedded) error = %v", err) + } + inspect(t, embedded, report.Hourly, "", "weather-light", "openrouter", "deepseek/deepseek-v4-flash") + inspect(t, embedded, report.Daily, "", "weather-balanced", "openrouter", "~google/gemini-flash-latest") + inspect(t, embedded, report.Daily, "weather-deep", "weather-deep", "openrouter", "~anthropic/claude-sonnet-latest") + + override, err := promptkitadapter.New(promptkitadapter.Config{ProfileFile: writeProfileFile(t, `id: weather-light +endpoint: https://local.example/v1 +model: local-weather +`)}) + if err != nil { + t.Fatalf("New(override) error = %v", err) + } + inspect(t, override, report.Hourly, "", "weather-light", "", "local-weather") +} + +func resolvedPromptProfile(t *testing.T, id report.ID) report.Resolved { + t.Helper() + now := time.Date(2026, 5, 29, 12, 0, 0, 0, time.UTC) + request := report.ResolveRequest{Now: now, Location: time.UTC} + if id == report.Daily { + request.Date = now + } + resolved, err := report.DefaultRegistry().Resolve(id, request) + if err != nil { + t.Fatalf("Resolve(%q) error = %v", id, err) + } + return resolved +} + +func writeProfileFile(t *testing.T, profile string) string { + t.Helper() + path := filepath.Join(t.TempDir(), "profile.yml") + if err := os.WriteFile(path, []byte(profile), 0o600); err != nil { + t.Fatalf("write profile: %v", err) + } + return path +} diff --git a/internal/app/single_report_workflow_test.go b/internal/app/single_report_workflow_test.go index de2789c..0540df6 100644 --- a/internal/app/single_report_workflow_test.go +++ b/internal/app/single_report_workflow_test.go @@ -51,6 +51,8 @@ type workflowExecutor struct { beforeProvider func() preparationDebug *promptexec.PreparationDebug executionDebug *promptexec.ExecutionDebug + preparation *promptexec.Preparation + execution *promptexec.Execution } func (e *workflowExecutor) InspectPrompt(_ context.Context, id, version string) (promptexec.PromptInspection, error) { @@ -93,6 +95,7 @@ func (e *workflowExecutor) Execute(_ context.Context, req promptexec.ExecuteRequ RenderedPromptHash: "rendered-hash", ProfileID: req.ProfileID, BackendID: profile.BackendID, ModelName: profile.ModelName, DataPackagePath: req.DataPackagePath, StartedAt: stamp, EndedAt: stamp, } + e.preparation = &preparation if err := callback(preparation, e.preparationDebug); err != nil { return nil, err } @@ -110,14 +113,16 @@ func (e *workflowExecutor) Execute(_ context.Context, req promptexec.ExecuteRequ if validation == "" { validation = promptexec.ValidationPassed } - return &promptexec.Execution{ + execution := &promptexec.Execution{ RunID: "provider-run", PromptID: req.PromptID, PromptVersion: req.PromptVersion, PromptHash: "prompt-hash", RenderedPromptHash: "rendered-hash", ProfileID: req.ProfileID, BackendID: profile.BackendID, ModelName: profile.ModelName, GeneratedHash: "generated-hash", StartedAt: stamp, EndedAt: stamp, DataPackagePath: req.DataPackagePath, RawOutput: e.raw, Debug: e.executionDebug, Validation: promptexec.NewValidation(validation, "json_schema", e.definition.GeneratedTextSchemaID+".generated_text.schema.json", nil), - }, nil + } + e.execution = execution + return execution, nil } type workflowNotifier struct { @@ -251,7 +256,7 @@ func TestGenerateDetailedPreservesSelectedProfileThroughExecution(t *testing.T) definition: definition, prompt: logicalPromptInspection(definition), profile: test.profile, raw: []byte(test.raw), } bundle := workflowBundle(t) - result, err := GenerateDetailed(context.Background(), GenerateRequest{ + _, err := GenerateDetailed(context.Background(), GenerateRequest{ Config: cfg, Report: test.kind, Date: workflowTime("2026-05-29T12:00:00-05:00"), Now: workflowTime("2026-05-29T08:30:00-05:00"), Collector: &workflowCollector{result: &collect.Result{Bundle: &bundle}}, Executor: executor, Notifier: &workflowNotifier{}, }) @@ -261,23 +266,11 @@ func TestGenerateDetailedPreservesSelectedProfileThroughExecution(t *testing.T) if executor.request.ProfileID != test.profile.ProfileID { t.Fatalf("execution profile = %q, want %q", executor.request.ProfileID, test.profile.ProfileID) } - store, err := state.NewFilesystemStore(cfg.Workspace) - if err != nil { - t.Fatalf("NewFilesystemStore() error = %v", err) + if executor.preparation == nil || executor.preparation.ProfileID != test.profile.ProfileID || executor.preparation.BackendID != test.profile.BackendID || executor.preparation.ModelName != test.profile.ModelName { + t.Fatalf("prepared profile = %#v, want %q/%q/%q", executor.preparation, test.profile.ProfileID, test.profile.BackendID, test.profile.ModelName) } - preparation, err := store.LoadPromptPreparation(context.Background(), result.PreparationPath) - if err != nil || preparation.Preparation == nil || preparation.Preparation.ProfileID != test.profile.ProfileID || preparation.Preparation.BackendID != test.profile.BackendID || preparation.Preparation.ModelName != test.profile.ModelName { - t.Fatalf("preparation/error = %#v/%v", preparation, err) - } - execution, err := store.LoadPromptExecution(context.Background(), result.ExecutionPath) - if err != nil || execution.Provenance == nil || execution.Provenance.ProfileID != test.profile.ProfileID || execution.Provenance.BackendID != test.profile.BackendID || execution.Provenance.ModelName != test.profile.ModelName { - t.Fatalf("execution/error = %#v/%v", execution, err) - } - for _, path := range []string{result.MetadataPath, result.PreparationPath, result.ExecutionPath} { - data, err := os.ReadFile(path) - if err != nil || strings.Contains(string(data), "https://") || strings.Contains(string(data), "api_key") { - t.Fatalf("ordinary artifact %q leaks sensitive profile details or could not be read: %v", path, err) - } + if executor.execution == nil || executor.execution.ProfileID != test.profile.ProfileID || executor.execution.BackendID != test.profile.BackendID || executor.execution.ModelName != test.profile.ModelName { + t.Fatalf("executed profile = %#v, want %q/%q/%q", executor.execution, test.profile.ProfileID, test.profile.BackendID, test.profile.ModelName) } }) }