Files
weatherreporter/docs/roadmap/implementation.md

43 KiB

Promptkit Migration Implementation Plan

Status: Decision-complete implementation plan; unimplemented.

Purpose And Authority

This document defines the ordered implementation procedure for the Promptkit migration roadmap. The feature roadmap is authoritative for scope, user intent, policy choices, and the desired end state. This plan is authoritative for sequencing, concrete package ownership, compatibility work, tests, documentation updates, and completion gates.

Implement the stages in order. Do not reinterpret a roadmap decision merely because the current Scriptorium implementation makes another path shorter. When implementation reveals a conflict with the roadmap, stop and update the roadmap and this plan deliberately rather than introducing an implicit policy change.

This plan follows the repository's architecture, documentation, and testing policies. All stages are parts of one migration change. Intermediate stages may temporarily retain code needed by a later cutover, but no intermediate state should be released or documented as the final current behavior.

Cross-Stage Rules

  • Pin gitea.maximumdirect.net/eric/promptkit at exactly v0.4.0. Do not commit a go.work, local replace, pseudo-version, or unpublished commit.
  • Keep Promptkit types inside internal/adapters/promptkit, that package's tests, and the external contract test that verifies internal/promptassets with the real Promptkit inspector. App, CLI, report, state, and domain packages use project-owned types.
  • Keep Promptkit's opaque prepared-execution handle inside its adapter. Do not expose it through an app interface, serialize it, or make it restartable.
  • Construct one Promptkit engine per generate or run CLI invocation. Every report in a batch shares that engine. Inspection commands that only read persisted state do not construct an engine.
  • Keep batches sequential. Do not add retries, output repair, direct Markdown generation, arbitrary backend registration, or live-provider tests.
  • Preserve the exact persisted data_package bytes as the only model input. Promptkit receives those bytes through an inline artifact with the managed YAML path as provenance; it never receives an unrestricted file reference.
  • Persist successful preparation before RunPrepared. If preparation-record persistence or enabled debug persistence fails, discard the handle and do not call the provider.
  • Keep normal artifacts, errors, logs, and CLI summaries free of credentials, rendered messages, schema bodies, data-package bodies, provider response bodies, and full effective parameter maps.
  • Use atomic writes for durable state and debug files. Preserve every non-secret artifact reached before a later failure when practical.
  • Use deterministic Promptkit model-client fakes at the adapter boundary and project-owned executor fakes at the app boundary. The default test suite remains offline and credential-free.
  • At each stage, update or delete existing tests according to the testing policy. Protect contracts and failure behavior; do not preserve tests that exist only to assert the retired subprocess implementation.
  • Run the focused checks named by the stage while iterating. Before completing every stage, run git diff --check. Run the full repository gate in the final stage.

Fixed Package And Contract Decisions

Use these ownership boundaries so later stages do not need to choose an architecture while editing:

Area Owner Decision
Embedded prompt definitions, referenced prompt content, and provider-facing schemas internal/promptassets Embed one centralized asset tree and expose read-only prompt and schema fs.FS sources plus schema lookup by report schema ID. This package does not import Promptkit.
Project-owned prompt execution contract internal/promptexec Own prompt inspection, profile inspection, preparation, execution, validation, usage, debug, and neutral error values. This package does not import Promptkit, app, CLI, report, or state.
Promptkit integration internal/adapters/promptkit Construct and own the Promptkit engine, translate project requests and results, classify public Promptkit errors, and enforce prepared-handle lifecycle.
Workflow ordering internal/app Resolve reports, inspect prompt/profile selections, collect weather, build and save inputs, record preparation, invoke execution, persist results, perform final generated-text validation, render Markdown, and notify.
Durable artifacts internal/state Own v2 paths, preparation and execution schemas, metadata compatibility, atomic writes, and inspection reads.
Engine composition and flags internal/cli Construct one executor per action invocation, pass it into app requests, parse --llm-debug-dir, and expose project-owned summary paths.

internal/promptexec.Executor must provide three operations:

  1. inspect an exact prompt ID and version;
  2. inspect one explicit profile ID; and
  3. prepare and execute one request.

The execution operation accepts a preparation callback. The adapter calls that callback exactly once after successful PrepareExecution and before RunPrepared. The callback receives a safe project-owned preparation value and, only when requested, a project-owned sensitive debug value. A callback error aborts execution. A preparation failure returns a classified project-owned attempt error without calling the callback; app orchestration uses that error to persist a failure receipt.

The adapter execution request contains the exact prompt ID and version, the optional configured profile override, the exact YAML bytes, the managed data-package path used only as inline provenance, and whether sensitive debug capture is enabled. It does not contain a provider output path.

Project-owned errors use stable categories for invalid configuration, invalid request, prompt not found/load, profile not found/load, missing credential, artifact load, prompt render, capacity, generation, operational validation, completed validation rejection, cancellation, and deadline. A capacity error also carries the non-secret backend ID. Its Error text is Weatherreporter-owned and safe for CLI output; an unexported or non-serialized cause may remain available to errors.Is and errors.As. Never persist arbitrary dependency error prose.

Stage 1: Correct The Implemented Report Surface

Goal

Remove the unfinished three-day, weekend, and storm products before changing the LLM integration. The repository should expose only Daily, Today, Tomorrow, and Hourly while continuing to use Scriptorium temporarily for those four reports.

Work

  1. Remove the three report definitions and their files under internal/report. Remove their IDs, command names, aliases, period inputs, generation modes, batch flags, registry order, and direct-Markdown-only declarations.
  2. Add PromptVersion string to report.Definition and set it to 1.0.0 on all four retained definitions. Since every retained report uses generated text plus a repository template, remove the now-redundant GenerationMode and Generated fields and their branches.
  3. Remove three-day, weekend, and storm command parsing, help text, app request fields/constants, resolution branches, Distributor path variables, and configuration report aliases.
  4. Remove report-specific code whose only caller was one of the retired products. Use graph traces before deletion. This includes the dedicated three-day/weekend Recent Changes comparators, storm-window fact derivation, multi-day report derivation used only by these definitions, storm time parsing, and their focused tests when no retained behavior depends on them. Do not remove general meteorological behavior merely because a fixture or prose string contains words such as “storm” or “weekend.”
  5. Keep morning and evening planning behavior exactly as implemented: morning plans Today and Tomorrow, evening plans Tomorrow, and both may add eligible future Daily reports from complete hourly coverage.
  6. Update the current CLI, configuration, report-registry, facts, modules, changes, operations, troubleshooting, and other affected documentation to describe only the four implemented reports. Update docs/roadmap/future.md so it no longer claims that a manual Storm Report exists; future versions of these products remain explicitly unimplemented roadmap work.

Tests

  • Rewrite report registry and CLI parser tests around exactly four commands.

  • Remove fake successful-generation tests for nonexistent prompts.

  • Preserve Daily/Today/Tomorrow/Hourly period, module, batch-planning, Distributor path, and output-name coverage.

  • Add negative CLI/config tests showing the retired command names and aliases are rejected.

  • Run:

    go test ./internal/report ./internal/config ./internal/facts \
      ./internal/changes ./internal/app ./internal/cli
    go test ./...
    go run ./cmd/weatherreporter --help
    git diff --check
    

Exit Gate

No implemented registry, CLI, configuration, app, current-state document, or test claims that three-day, weekend, or storm generation exists. The four retained reports still pass through the existing generated-text/template workflow, and each carries exact prompt version 1.0.0.

Stage 2: Promote And Reconcile Runtime Prompt Assets

Goal

Create the single embedded prompt/schema corpus used by Promptkit and by Weatherreporter's final generated-text boundary.

Work

  1. Add Promptkit v0.4.0 to go.mod and go.sum.

  2. Create internal/promptassets with an embedded tree rooted at:

    assets/
      prompts/
        common/
        daily/
        today/
        tomorrow/
        hourly/
      schemas/
    

    Expose copied/read-only prompt and schema filesystem views and a Schema(id) lookup for the four report schema IDs. Keep path constants private except where a stable project contract is necessary.

  3. Promote only the four *_generated_text prompt definitions from docs/roadmap/scriptorium. Do not promote weather.daily_report.

  4. Use the Scriptorium corpus's common system, data-package guidance, and report-specific files as the behavioral base. Reconcile rather than editorially rewrite:

    • preserve its source weighting, hazard, precipitation, and style rules;
    • add the optional confidence field instruction to Daily, Today, and Hourly so all prompt text agrees with the canonical domain types and schemas;
    • retain Tomorrow's existing confidence instruction;
    • prefer the corpus's correct Tomorrow daily framing over the stale hourly-style file currently under internal/reporttemplate/prompts;
    • make no unrelated prompt-tone or product-policy changes.
  5. Every prompt definition must use:

    • its existing prompt ID;
    • version 1.0.0;
    • default_profile: gemini-flash-latest;
    • one required data_package input with content_type: application/yaml;
    • JSON output with validation_mode: json_schema;
    • a path inside the embedded schema root; and
    • zero or omitted repair_attempts.
  6. Move the four canonical schemas from internal/reporttemplate/schemas into internal/promptassets/assets/schemas. Preserve their current application contracts, including optional confidence, required summary/discussion, array discussion for day-style reports, string discussion for Hourly, additionalProperties: false, and the corrected Daily $id and title.

  7. Change internal/generatedtext to load these canonical schemas from promptassets. Remove schema and prompt-fragment ownership from internal/reporttemplate; it should embed and render Markdown templates and partials only.

  8. Retain docs/roadmap/scriptorium as migration source evidence until the final cleanup stage.

Tests

  • Add table-driven asset tests for exactly four prompt IDs and versions, referenced-file resolution, default profiles, YAML input declarations, output contracts, and schema lookup.

  • Construct a Promptkit engine over the embedded sources in an external test package and call InspectPrompt for every report. Do not use a provider or credentials.

  • Retain generated-text schema and typed-validation tests, now reading the relocated canonical schemas.

  • Assert that no active embedded prompt uses local-heavy, pipeline-weather/, JSON input metadata, positive repair attempts, or weather.daily_report.

  • Run:

    go test ./internal/promptassets ./internal/generatedtext \
      ./internal/reporttemplate
    git diff --check
    

Exit Gate

One embedded source contains exactly four valid Promptkit prompt definitions and four canonical schemas. Prompt inspection succeeds offline for every report, and no duplicate runtime provider-facing schema remains.

Stage 3: Define The Neutral Prompt Execution Contract

Goal

Define the complete project-owned boundary between Weatherreporter and an LLM execution adapter. This stage introduces no Promptkit imports and makes no production orchestration changes.

Work

  1. Create internal/promptexec with the neutral contract fixed above. Include:

    • exact prompt and profile inspection values;
    • safe preparation provenance;
    • execution result, validation, token usage, and timing;
    • optional sensitive preparation/execution debug values;
    • stable error categories and a typed capacity error; and
    • small helpers for safe diagnostic bounding and copying.
  2. The safe preparation value includes prompt ID/version/hash, rendered prompt hash, input hashes, selected profile/backend, effective model name, output format/validation/schema path, preparation start/end/duration, and the managed data-package path. It excludes endpoints, API-key environment names, full parameters, rendered messages, schema bodies, and input bodies.

  3. The safe execution value includes Promptkit run ID, repeated prompt and input provenance, selected profile/backend/model, generated-content hash, usage, execution start/end/duration, completed validation status, bounded validation diagnostics, and raw output bytes. It excludes endpoint and full effective parameters.

  4. Bound routine validation diagnostics to at most 10 entries and at most 1,024 bytes per entry, truncating on a valid UTF-8 boundary. Bound a persisted safe error message to 2,048 bytes. Prefer stable categories and fields over diagnostic prose.

  5. Specify the execution lifecycle in interface comments and tests:

    • inspection is side-effect-free;
    • preparation invokes the callback exactly once only after successful preparation;
    • callback failure prevents provider execution;
    • completed validation rejection is a result, not an operational error;
    • operational failure does not fabricate a completed result; and
    • sensitive debug values are populated only when explicitly requested.
  6. Keep copying and truncation helpers internal to promptexec. Callers must not be able to mutate byte slices, maps, or diagnostics retained inside contract values.

Tests

  • Add compile-time test fakes that implement the complete interface without importing Promptkit.

  • Add table-driven tests for every error category, capacity fields, diagnostic count and byte limits, UTF-8-safe truncation, error-message bounds, and defensive copying.

  • Assert through representative fixtures that safe contract values contain no endpoint, credential, rendered message, schema body, input body, response body, or full parameter map.

  • Run:

    go test ./internal/promptexec
    git diff --check
    

Exit Gate

internal/promptexec provides a stable, documented, offline-tested contract that can represent every inspection, preparation, execution, validation, usage, debug, and failure outcome required by the roadmap. It has no Promptkit, app, CLI, report, or state dependency.

Stage 4: Implement The Promptkit Adapter

Goal

Implement and thoroughly test Promptkit behind the Stage 3 contract without modifying application or CLI orchestration.

Work

  1. Create internal/adapters/promptkit. Its constructor:

    • uses promptassets through WithPromptFS and WithSchemaFS;
    • selects one external profile directory through Promptkit config or one profile file through WithProfileFile;
    • optionally registers promptkit.LocalBackend(endpoint, concurrencyLimit);
    • sets the transport timeout;
    • accepts an injected Promptkit LLMClient only through an adapter-local test constructor or option; and
    • returns project-owned configuration errors.
  2. Implement prompt and profile inspection with exact versions and explicit field-by-field translation into promptexec values. Do not return or embed Promptkit values.

  3. Implement execution:

    • call PrepareExecution with the exact requested version and optional profile override;
    • pass the exact YAML as InlineWithURI(dataPackagePath, string(bytes));
    • immediately defer handle.Discard() after a successful prepare;
    • map safe preparation details and invoke the preparation callback;
    • call RunPrepared only after the callback succeeds;
    • treat ValidationFailed as a completed result rather than an operational error; and
    • return no invented execution result for operational errors.
  4. Translate Promptkit's public error sentinels with errors.Is, and CapacityError with errors.As. Preserve caller cancellation and deadline identities. Do not parse error strings.

  5. Capture sensitive debug values only when requested. Use project-owned debug structs and explicit field mapping rather than serializing whole Promptkit values, so a future upstream field cannot silently enter an artifact.

  6. Keep the prepared handle entirely within the adapter call. Discard it after every success or failure path and never make it serializable or reusable.

Tests

  • Use an injected deterministic Promptkit model client. Never use a live endpoint or credential.

  • Cover prompt/profile inspection, profile override precedence, built-in and external profiles, endpoint-only profiles, the conventional local backend, local concurrency/capacity rejection, missing credentials, and constructor validation.

  • Cover preparation callback ordering, callback failure preventing generation, exact inline bytes and URI, exact prompt version, successful execution, completed schema rejection with raw output, operational generation and validation failures with no partial result, cancellation, timeout, and handle discard.

  • Assert that safe values and errors do not contain rendered messages, endpoints, schema bodies, input bodies, response bodies, credentials, or full parameter maps.

  • Run:

    go test ./internal/promptexec ./internal/adapters/promptkit
    go test -race ./internal/adapters/promptkit
    git diff --check
    

Exit Gate

The Promptkit adapter is a complete offline-tested implementation of the neutral contract. No Promptkit type is referenced outside the adapter, its tests, or the external asset contract test, and no production path uses the adapter yet.

Stage 5: Introduce Durable State V2

Goal

Add the final normal-artifact model and backward-compatible metadata reading before app orchestration starts writing the new records. Sensitive debug storage remains out of scope until Stage 6.

Work

  1. Add the new-run path members Preparation and Execution:

    • preparation: preflight/<group>/<date>/prompt_preparation.<runID>.json;
    • execution: snapshots/<group>/<date>/prompt_execution.<runID>.json.

    Keep workspace.preflight_dir and its default physical directory unchanged. Retain Preflight and GeneratedTextResult only as temporary legacy write-path members required by the uncut Scriptorium workflows; mark them for deletion in Stage 10. New Promptkit code must use only the new members.

  2. Add PromptPreparationArtifact with schema weatherreporter.prompt_preparation.v1. It represents either successful preparation or a failed preparation receipt and includes schema version, status, report/RunID, prompt identity, safe preparation provenance when available, timing, data-package path, and a bounded classified error when failed.

  3. Add PromptExecutionArtifact with schema weatherreporter.prompt_execution.v1. It represents success, completed validation rejection, or operational failure and includes safe execution provenance, validation, usage, timing, reached artifact paths, and a bounded classified error when failed. It never embeds generated content.

  4. Advance newly written metadata to weatherreporter.metadata.v2. Replace preflightPath and generatedTextResultPath with preparationPath and executionPath. Populate paths only after the corresponding artifact has actually been saved.

  5. Implement explicit v1/v2 metadata decoding:

    • accept only the known v1 and v2 schema versions;
    • normalize v1 preflightPath and generatedTextResultPath internally for inspection;
    • preserve v1 field names when a loaded v1 record is marshaled by inspect metadata;
    • write only v2 through the new Promptkit state APIs; the temporary Scriptorium path may continue writing v1 until its Stage 10 removal; and
    • never dual-write legacy aliases.

    Direct inspection of v1 metadata, modules, data packages, sources, and referenced artifacts remains available for any historical report ID. Prior-snapshot reconstruction is required only for the four retained report IDs; do not restore retired definitions solely for legacy comparison.

  6. Add typed SavePromptPreparation, SavePromptExecution, and corresponding typed load methods where inspection needs them. Temporarily retain the Scriptorium write methods so the current production paths compile through Stage 9; delete those methods during Stage 10. Keep long-term legacy support read-only.

  7. Update app/CLI result structs and JSON field names to preparationPath, executionPath, and optional llmDebugPath; temporarily adapt old orchestration so the tree compiles until the cutover stages. Do not write v2 records through the Scriptorium path.

Tests

  • Add exact path, schema-defaulting, round-trip, required-field, atomic-write, and unknown-version tests.

  • Add v1 fixtures covering both generated-text and legacy preflight references; verify list, metadata, modules, data-package, and source inspection.

  • Verify that re-marshaled v1 inspection uses v1 field names and that v2 output contains no deprecated aliases.

  • Run:

    go test ./internal/state ./internal/app ./internal/cli
    git diff --check
    

Exit Gate

State can read historical v1 runs and write the complete normal v2 artifact contract through the new APIs. The existing production generation path still compiles and remains unchanged until cutover; no new Promptkit artifact uses a Scriptorium-specific filename, schema, or JSON field.

Stage 6: Add Secure LLM Debug Persistence

Goal

Implement the explicitly enabled sensitive-debug store as a small, independently auditable boundary before any CLI command can request debug capture.

Work

  1. Add a focused debug writer, owned by state or a narrow state-adjacent package, for an explicit operator root outside normal artifact derivation. It must:

    • validate or create the root before collection or provider work;
    • use <root>/<report-id>/<valid-date>/<run-id>/;
    • write preparation.json and execution.json atomically;
    • create directories with 0700 and files with 0600;
    • reject symlinks, path escape, non-directory roots, and unsafe run/report/date segments; and
    • return the per-run debug directory as the project-owned summary path.
  2. Define explicit project-owned debug wire structs. Preparation debug may contain rendered messages, structured-output schema, and effective endpoint/parameters. Execution debug may additionally contain raw generated output and completed validation details.

  3. Exclude direct credentials, resolved environment credential values, and arbitrary dependency objects. Use field-by-field mapping so future promptexec or Promptkit fields are not serialized implicitly.

  4. Make disabled debug behavior a no-op that performs no filesystem access. A requested debug initialization or write error is returned to the caller; it is never silently downgraded.

  5. Keep this writer independent of normal state path derivation. Normal state inspection must not discover or serve sensitive debug artifacts.

Tests

  • Verify exact grouping, atomic replacement, disabled behavior, and returned per-run paths.

  • Verify directory 0700 and file 0600 permissions on supported platforms.

  • Cover traversal, absolute-segment, symlink-root, symlink-component, non-directory, and invalid report/date/run segment rejection.

  • Marshal representative debug fixtures and verify credentials and resolved secret values are absent while the explicitly allowed diagnostic fields are retained.

  • Run:

    go test ./internal/state
    git diff --check
    

Exit Gate

Sensitive debug persistence is secure, explicitly rooted, atomic, isolated from normal state, and fully tested. No production CLI path enables it yet.

Stage 7: Add Promptkit Configuration, Composition, And Inspection

Goal

Add the configuration and project-owned composition seams needed for cutover, and centralize pre-collection prompt/profile validation. Retain Scriptorium as the production generator until Stage 8 so this stage does not create a half-cut-over runtime.

Work

  1. Add config.PromptkitConfig and nested local config with exactly:

    • profile;
    • profile_file;
    • profile_dir;
    • timeout, default 2m;
    • local.endpoint; and
    • local.concurrency_limit, default 1.

    Scriptorium config remains temporarily because production generation has not yet cut over. Promptkit validation rejects simultaneous profile sources, non-positive transport timeout, an invalid nonblank local endpoint, and negative concurrency. A blank local endpoint leaves local unregistered; concurrency zero means unlimited.

  2. Add an executor factory seam to cli.Runner using only project-owned types. Production construction delegates to internal/adapters/promptkit; CLI tests inject a fake factory. The factory creates one executor for an action, not one per report.

  3. Add a project-owned app inspection helper that:

    • inspects the exact prompt ID and report.Definition.PromptVersion;
    • verifies exactly one required data_package input with application/yaml;
    • verifies the expected JSON Schema output contract and declared default profile;
    • selects promptkit.profile when nonblank, otherwise the prompt default;
    • inspects that explicit profile;
    • rejects a profile requiring a direct API key, because Weatherreporter has no direct-key configuration; and
    • requires a nonblank environment value for any reported APIKeyEnv.
  4. Return only safe, project-owned inspection values and classified errors. The helper performs no collection, provider call, or durable write.

  5. Unit-test production factory argument mapping, but do not yet wire it into generate or run. Stage 8 performs that atomic production cutover.

Tests

  • Cover exact defaults, mutual profile-source exclusion, local endpoint validation, zero/unlimited and negative concurrency, and timeout validation.

  • Cover prompt/profile default and override precedence, prompt contract mismatch, unsupported direct-key profiles, missing environment credentials, and safe errors.

  • Verify the CLI factory maps embedded assets, external profile source, local backend, concurrency, and timeout exactly once without exposing Promptkit types.

  • Run:

    go test ./internal/config ./internal/app ./internal/cli \
      ./internal/adapters/promptkit
    git diff --check
    

Exit Gate

Configuration, engine construction, and pre-collection inspection are decision-complete and offline-tested behind project-owned seams. Production generation still follows the old path, so there is no dual execution mode.

Stage 8: Cut Over Single-Report Execution And Failure Persistence

Goal

Move all four generate commands to Promptkit prepared execution while preserving report output behavior and inspectable partial failure. Debug CLI enablement and final summary presentation are deferred to Stage 9.

Work

  1. Construct exactly one executor after configuration is loaded for a generate action and pass it through app.GenerateRequest.

  2. Isolate the current shared generation function before changing it:

    • move the Scriptorium implementation behind a temporary, unexported legacy batch helper;
    • make GenerateDetailed call the new Promptkit implementation described below;
    • leave RunBatchDetailed calling only the legacy helper until Stage 10; and
    • do not add a runtime switch, fallback, or dual invocation for the same report.

    This temporary split is solely a staging seam. Stage 10 moves batches to the Promptkit implementation and deletes the helper.

  3. Reorder GenerateDetailed:

    1. resolve the report and RunID;
    2. require the injected executor;
    3. run the Stage 7 exact prompt/profile/credential inspection; and only then
    4. collect weather.

    Inspection failures occur before managed run artifacts exist.

  4. Implement the new project-owned report-generation core with the injected executor and the Stage 5 artifact contract:

    • build, serialize once, and save the data package;
    • use those exact serialized bytes for execution;
    • persist a failed preparation receipt and v2 metadata when preparation fails;
    • use the preparation callback to save successful preparation provenance and metadata before provider work;
    • after execution success or completed validation rejection, save exact raw output, then the execution artifact, then metadata;
    • after operational execution failure, save a failure execution receipt and metadata without inventing raw output;
    • on Promptkit validation rejection, return a classified report failure after preserving raw output and bounded diagnostics;
    • on Promptkit validation success, run the existing typed internal/generatedtext validation, normalize/save generated text, build and save render context, and render managed Markdown as before.
  5. Make partial failures return a non-nil ReportResult whenever a RunID and inspectable paths exist. Preserve preparation/execution/raw/normalized/ context/report paths reached before preparation, capacity, generation, Promptkit validation, domain validation, template, copy, or notification failure.

  6. Remove Scriptorium result values from the single-report app requests, results, and test fakes used by the generate path. App tests depend only on promptexec.Executor. Retain the old adapter and configuration only for the still-uncut batch path.

  7. Do not expose --llm-debug-dir yet. Pass CaptureDebug: false during this stage; Stage 9 adds debug orchestration without changing provider execution semantics.

Tests

  • Add representative offline app workflows for each of the four reports, using real state, generated-text validation, contexts, and templates with a fake executor.

  • Verify inspection and credential checks occur before collection.

  • Verify exact profile override/default precedence and exact prompt versions.

  • Verify persistence ordering by observing that execution is not called until the successful preparation artifact and metadata exist.

  • Cover preparation failure, callback/state failure, capacity rejection, credential disappearance at execution, cancellation, deadline, generation failure, operational validation failure, completed schema rejection, generated-text domain failure, template failure, output-copy failure, and notification failure.

  • Verify partial results retain every reached normal-artifact path without prompt or response content.

  • Verify single-report commands invoke only Promptkit and batch commands invoke only the temporary legacy helper during this intermediate stage.

  • Run:

    go test ./internal/app ./internal/cli
    go test -race ./internal/app ./internal/adapters/promptkit
    git diff --check
    

Exit Gate

Every single-report command uses one Promptkit engine, performs inspection before collection, persists preparation before provider work, and produces the same managed Markdown/template and notification outcomes through project-owned contracts. Sensitive debug capture remains disabled.

Stage 9: Add Single-Report Debug And CLI Summary Behavior

Goal

Expose the opt-in debug workflow and finalize single-report CLI success and partial-failure summaries without expanding the normal artifact boundary.

Work

  1. Parse --llm-debug-dir PATH for all four generate commands. Do not add a YAML debug switch.

  2. Validate or create the debug root before prompt inspection or weather collection. A requested initialization failure is terminal and occurs before managed run artifacts exist.

  3. Pass debug intent through project-owned CLI and app request types. When enabled, request sensitive debug values from the executor:

    • write preparation debug inside the preparation callback, after the normal preparation artifact is saved and before metadata is finalized and provider work begins;
    • if the debug write fails, return the callback error so RunPrepared is not called;
    • write execution debug immediately after receiving a completed execution result and before continuing with downstream validation/rendering; and
    • treat every requested debug write failure as terminal while preserving already reached normal artifacts.
  4. Set llmDebugPath only after the per-run debug directory contains at least one successfully written debug artifact. Never copy sensitive debug content into a normal artifact, log, error, or summary.

  5. Finalize human and JSON CLI summaries around project-owned preparationPath, executionPath, and optional llmDebugPath. Include reached paths on partial failure; omit absent paths instead of inventing them.

  6. Preserve existing output behavior: quiet mode suppresses successful human output, JSON output remains machine-readable, and failures return a non-zero status with safe classified text.

Tests

  • Add CLI parser and help tests for --llm-debug-dir on all four generate commands and its absence from YAML configuration.

  • Verify debug initialization precedes inspection and collection.

  • Verify preparation debug persistence precedes provider execution and that a write failure prevents the model-client call.

  • Cover execution debug success/failure, partial normal-artifact retention, llmDebugPath population rules, quiet mode, JSON summaries, and safe failure text.

  • Assert that summaries and routine logs contain no rendered prompt, schema body, input body, generated body, endpoint, full parameters, or credential.

  • Run:

    go test ./internal/app ./internal/cli ./internal/state
    go test -race ./internal/app ./internal/adapters/promptkit
    git diff --check
    

Exit Gate

All single-report commands provide the complete Promptkit, v2 persistence, opt-in debug, and summary behavior required by the roadmap. Debug failure ordering is enforced without exposing sensitive content.

Stage 10: Cut Over Batches And Remove Scriptorium

Goal

Complete the production cutover, share one engine through each sequential batch, and delete the retired dependency boundary.

Work

  1. Construct one executor in the CLI for each run morning or run evening invocation and pass it through BatchRequest to every report.

  2. Parse and initialize --llm-debug-dir for run commands using the same policy as generate commands.

  3. Before collection, inspect the complete candidate set:

    • morning: Today, Tomorrow, and Daily;
    • evening: Tomorrow and Daily.

    Inspect exact prompt versions, validate declared input/output contracts, resolve unique effective profiles, and enforce credential availability. Daily inspection occurs before its collection-dependent future dates are known because every eligible Daily run uses the same exact prompt contract.

  4. Collect once, plan the batch as before, and execute every planned report sequentially through the shared executor, Stage 8 generation core, and store. Delete the temporary legacy batch helper after this call site moves. Preserve continuation after independent report failures and existing batch-notification gating.

  5. Include preparation, execution, and optional debug paths in each batch item, including failed items when those paths were reached.

  6. Replace Scriptorium configuration entirely with Promptkit configuration. Update defaults, validation, examples, config tests, and all construction sites. Explicitly reject a top-level scriptorium: key with an actionable migration error even though the general YAML loader is currently permissive; silently ignoring a former execution configuration is unsafe. Do not translate it or add a dual-run mode. The maintained examples contain only promptkit:.

  7. Delete:

    • internal/adapters/scriptorium;
    • the Scriptorium renderer interface and all result/request types;
    • the legacy state path members and Scriptorium write methods retained in Stage 5;
    • subprocess fakes and subprocess-specific tests;
    • direct-Markdown branches and remaining Scriptorium defaults; and
    • dead helpers used only by the old adapter.
  8. Run go mod tidy and verify the module graph contains Promptkit v0.4.0 and no Scriptorium module or local replacement.

Tests

  • Add batch tests proving one factory call/engine per CLI invocation, one executor shared across all reports, one collection, sequential ordering, later-report continuation, per-report path retention, and notification only after complete success.

  • Cover a capacity failure for one report followed by a later independent report and confirm no retry occurs.

  • Cover debug grouping for multiple reports and multiple Daily dates.

  • Update config tests for exact defaults, mutual exclusivity, local endpoint, zero/unlimited and negative concurrency, and maintained examples.

  • Run:

    go test ./internal/config ./internal/app ./internal/cli
    go test -race ./internal/app ./internal/adapters/promptkit
    go test ./...
    go run ./cmd/weatherreporter --help
    git diff --check
    

Exit Gate

All production generation paths use Promptkit. A batch owns one engine and continues sequentially under existing failure policy. No Scriptorium code, configuration field, subprocess path, or dependency remains.

Stage 11: Update Canonical Documentation And Complete Verification

Goal

Make current-state documentation match the completed implementation, remove migration-only source material, and perform repository-wide verification.

Work

  1. Update canonical current-state owners in the same migration change:

    • docs/development.md for repository orientation, package map, task routing, and validation language;
    • docs/policy/architecture.md for the Promptkit adapter boundary, generated-text-only flow, prepared execution, and four-report product;
    • docs/policy/testing.md for Promptkit/provider fakes instead of subprocess fakes and the offline external-boundary rule;
    • docs/cli.md for four commands, --llm-debug-dir, and renamed summary fields;
    • docs/config.md and examples/ for the exact Promptkit contract;
    • docs/operations.md for v2 paths, sensitive debug retention and permissions, preparation/execution lifecycle, and v1 inspection;
    • docs/troubleshooting.md for inspection, profile/credential, preparation, capacity, execution, validation, and debug failures;
    • relevant docs/internal/ files for app, CLI, report, prompt input, generated text, templates, state, collection, briefing, facts, changes, and package boundaries; and
    • docs/templates.md for Promptkit-generated prose and the relocated schema owner.
  2. Replace docs/integrations/scriptorium.md and docs/internal/scriptorium-adapter.md with canonical Promptkit integration and adapter documents. The integration document owns the logical prompt/profile/schema and durable compatibility contract; the internal document owns construction, mapping, lifecycle, and tests. Link rather than duplicate the full CLI/config/operations references.

  3. Search all non-roadmap current-state documentation, examples, Go code, tests, help output, and module metadata for stale scriptorium, local-heavy, retired report commands, old artifact field names, old filenames, and old metadata versions. Retain old names only in explicit v1 compatibility code/tests and historical roadmap discussion.

  4. Remove docs/roadmap/scriptorium after confirming every promoted runtime asset is represented under internal/promptassets and covered by asset tests. The unused legacy Daily Markdown prompt is deleted with this source directory.

  5. Keep the feature roadmap and this plan as temporary migration documents while implementation is under review. Once the implementation is accepted, mark both complete and move every still-useful contract to its canonical current-state owner before later archival or removal.

  6. Review all added links, fenced examples, flags, field names, defaults, schema IDs, file paths, and version strings against executable sources. Confirm examples contain no credentials or private infrastructure values.

Final Verification

Run:

gofmt -w <all changed Go files>
go mod tidy
go test ./...
go test -race ./...
go run ./cmd/weatherreporter --help
git diff --check

Then verify explicitly:

  • go list -m gitea.maximumdirect.net/eric/promptkit reports v0.4.0;
  • no committed go.work, replace, secret fixture, or live-provider test exists;
  • the maintained examples load through config tests;
  • all four embedded prompts inspect at exact version 1.0.0;
  • no runtime prompt requests repair attempts;
  • no ordinary artifact or CLI summary includes rendered prompts, schema bodies, input bodies, generated bodies, provider endpoints, or credentials;
  • v1 metadata fixtures remain inspectable and new runs write only v2;
  • help exposes only Daily, Today, Tomorrow, Hourly, morning, and evening; and
  • managed Markdown remains the only Distributor upload source.

Exit Gate

Every completion criterion in the feature roadmap is demonstrably satisfied. Current-state documentation describes the implementation rather than the migration, and the repository passes all final verification commands.

Open Questions

None. The feature roadmap and this implementation plan contain all product, architecture, configuration, compatibility, security, sequencing, and verification decisions required to begin implementation.