945 lines
43 KiB
Markdown
945 lines
43 KiB
Markdown
# Promptkit Migration Implementation Plan
|
|
|
|
Status: Decision-complete implementation plan; unimplemented.
|
|
|
|
## Purpose And Authority
|
|
|
|
This document defines the ordered implementation procedure for the
|
|
[Promptkit migration roadmap](promptkit.md). The feature roadmap is
|
|
authoritative for scope, user intent, policy choices, and the desired end
|
|
state. This plan is authoritative for sequencing, concrete package ownership,
|
|
compatibility work, tests, documentation updates, and completion gates.
|
|
|
|
Implement the stages in order. Do not reinterpret a roadmap decision merely
|
|
because the current Scriptorium implementation makes another path shorter.
|
|
When implementation reveals a conflict with the roadmap, stop and update the
|
|
roadmap and this plan deliberately rather than introducing an implicit policy
|
|
change.
|
|
|
|
This plan follows the repository's
|
|
[architecture](../policy/architecture.md),
|
|
[documentation](../policy/documentation.md), and
|
|
[testing](../policy/testing.md) policies. All stages are parts of one
|
|
migration change. Intermediate stages may temporarily retain code needed by a
|
|
later cutover, but no intermediate state should be released or documented as
|
|
the final current behavior.
|
|
|
|
## Cross-Stage Rules
|
|
|
|
- Pin `gitea.maximumdirect.net/eric/promptkit` at exactly `v0.4.0`. Do not
|
|
commit a `go.work`, local `replace`, pseudo-version, or unpublished commit.
|
|
- Keep Promptkit types inside `internal/adapters/promptkit`, that package's
|
|
tests, and the external contract test that verifies `internal/promptassets`
|
|
with the real Promptkit inspector. App, CLI, report, state, and domain
|
|
packages use project-owned types.
|
|
- Keep Promptkit's opaque prepared-execution handle inside its adapter. Do not
|
|
expose it through an app interface, serialize it, or make it restartable.
|
|
- Construct one Promptkit engine per `generate` or `run` CLI invocation.
|
|
Every report in a batch shares that engine. Inspection commands that only
|
|
read persisted state do not construct an engine.
|
|
- Keep batches sequential. Do not add retries, output repair, direct Markdown
|
|
generation, arbitrary backend registration, or live-provider tests.
|
|
- Preserve the exact persisted `data_package` bytes as the only model input.
|
|
Promptkit receives those bytes through an inline artifact with the managed
|
|
YAML path as provenance; it never receives an unrestricted file reference.
|
|
- Persist successful preparation before `RunPrepared`. If preparation-record
|
|
persistence or enabled debug persistence fails, discard the handle and do
|
|
not call the provider.
|
|
- Keep normal artifacts, errors, logs, and CLI summaries free of credentials,
|
|
rendered messages, schema bodies, data-package bodies, provider response
|
|
bodies, and full effective parameter maps.
|
|
- Use atomic writes for durable state and debug files. Preserve every
|
|
non-secret artifact reached before a later failure when practical.
|
|
- Use deterministic Promptkit model-client fakes at the adapter boundary and
|
|
project-owned executor fakes at the app boundary. The default test suite
|
|
remains offline and credential-free.
|
|
- At each stage, update or delete existing tests according to the testing
|
|
policy. Protect contracts and failure behavior; do not preserve tests that
|
|
exist only to assert the retired subprocess implementation.
|
|
- Run the focused checks named by the stage while iterating. Before completing
|
|
every stage, run `git diff --check`. Run the full repository gate in the
|
|
final stage.
|
|
|
|
## Fixed Package And Contract Decisions
|
|
|
|
Use these ownership boundaries so later stages do not need to choose an
|
|
architecture while editing:
|
|
|
|
| Area | Owner | Decision |
|
|
| --- | --- | --- |
|
|
| Embedded prompt definitions, referenced prompt content, and provider-facing schemas | `internal/promptassets` | Embed one centralized asset tree and expose read-only prompt and schema `fs.FS` sources plus schema lookup by report schema ID. This package does not import Promptkit. |
|
|
| Project-owned prompt execution contract | `internal/promptexec` | Own prompt inspection, profile inspection, preparation, execution, validation, usage, debug, and neutral error values. This package does not import Promptkit, app, CLI, report, or state. |
|
|
| Promptkit integration | `internal/adapters/promptkit` | Construct and own the Promptkit engine, translate project requests and results, classify public Promptkit errors, and enforce prepared-handle lifecycle. |
|
|
| Workflow ordering | `internal/app` | Resolve reports, inspect prompt/profile selections, collect weather, build and save inputs, record preparation, invoke execution, persist results, perform final generated-text validation, render Markdown, and notify. |
|
|
| Durable artifacts | `internal/state` | Own v2 paths, preparation and execution schemas, metadata compatibility, atomic writes, and inspection reads. |
|
|
| Engine composition and flags | `internal/cli` | Construct one executor per action invocation, pass it into app requests, parse `--llm-debug-dir`, and expose project-owned summary paths. |
|
|
|
|
`internal/promptexec.Executor` must provide three operations:
|
|
|
|
1. inspect an exact prompt ID and version;
|
|
2. inspect one explicit profile ID; and
|
|
3. prepare and execute one request.
|
|
|
|
The execution operation accepts a preparation callback. The adapter calls that
|
|
callback exactly once after successful `PrepareExecution` and before
|
|
`RunPrepared`. The callback receives a safe project-owned preparation value
|
|
and, only when requested, a project-owned sensitive debug value. A callback
|
|
error aborts execution. A preparation failure returns a classified
|
|
project-owned attempt error without calling the callback; app orchestration
|
|
uses that error to persist a failure receipt.
|
|
|
|
The adapter execution request contains the exact prompt ID and version, the
|
|
optional configured profile override, the exact YAML bytes, the managed
|
|
data-package path used only as inline provenance, and whether sensitive debug
|
|
capture is enabled. It does not contain a provider output path.
|
|
|
|
Project-owned errors use stable categories for invalid configuration, invalid
|
|
request, prompt not found/load, profile not found/load, missing credential,
|
|
artifact load, prompt render, capacity, generation, operational validation,
|
|
completed validation rejection, cancellation, and deadline. A capacity error
|
|
also carries the non-secret backend ID. Its `Error` text is
|
|
Weatherreporter-owned and safe for CLI output; an unexported or non-serialized
|
|
cause may remain available to `errors.Is` and `errors.As`. Never persist
|
|
arbitrary dependency error prose.
|
|
|
|
## Stage 1: Correct The Implemented Report Surface
|
|
|
|
### Goal
|
|
|
|
Remove the unfinished three-day, weekend, and storm products before changing
|
|
the LLM integration. The repository should expose only Daily, Today, Tomorrow,
|
|
and Hourly while continuing to use Scriptorium temporarily for those four
|
|
reports.
|
|
|
|
### Work
|
|
|
|
1. Remove the three report definitions and their files under `internal/report`.
|
|
Remove their IDs, command names, aliases, period inputs, generation modes,
|
|
batch flags, registry order, and direct-Markdown-only declarations.
|
|
2. Add `PromptVersion string` to `report.Definition` and set it to `1.0.0` on
|
|
all four retained definitions. Since every retained report uses generated
|
|
text plus a repository template, remove the now-redundant
|
|
`GenerationMode` and `Generated` fields and their branches.
|
|
3. Remove three-day, weekend, and storm command parsing, help text, app request
|
|
fields/constants, resolution branches, Distributor path variables, and
|
|
configuration report aliases.
|
|
4. Remove report-specific code whose only caller was one of the retired
|
|
products. Use graph traces before deletion. This includes the dedicated
|
|
three-day/weekend Recent Changes comparators, storm-window fact derivation,
|
|
multi-day report derivation used only by these definitions, storm time
|
|
parsing, and their focused tests when no retained behavior depends on them.
|
|
Do not remove general meteorological behavior merely because a fixture or
|
|
prose string contains words such as “storm” or “weekend.”
|
|
5. Keep morning and evening planning behavior exactly as implemented:
|
|
morning plans Today and Tomorrow, evening plans Tomorrow, and both may add
|
|
eligible future Daily reports from complete hourly coverage.
|
|
6. Update the current CLI, configuration, report-registry, facts, modules,
|
|
changes, operations, troubleshooting, and other affected documentation to
|
|
describe only the four implemented reports. Update `docs/roadmap/future.md`
|
|
so it no longer claims that a manual Storm Report exists; future versions
|
|
of these products remain explicitly unimplemented roadmap work.
|
|
|
|
### Tests
|
|
|
|
- Rewrite report registry and CLI parser tests around exactly four commands.
|
|
- Remove fake successful-generation tests for nonexistent prompts.
|
|
- Preserve Daily/Today/Tomorrow/Hourly period, module, batch-planning,
|
|
Distributor path, and output-name coverage.
|
|
- Add negative CLI/config tests showing the retired command names and aliases
|
|
are rejected.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/report ./internal/config ./internal/facts \
|
|
./internal/changes ./internal/app ./internal/cli
|
|
go test ./...
|
|
go run ./cmd/weatherreporter --help
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
No implemented registry, CLI, configuration, app, current-state document, or
|
|
test claims that three-day, weekend, or storm generation exists. The four
|
|
retained reports still pass through the existing generated-text/template
|
|
workflow, and each carries exact prompt version `1.0.0`.
|
|
|
|
## Stage 2: Promote And Reconcile Runtime Prompt Assets
|
|
|
|
### Goal
|
|
|
|
Create the single embedded prompt/schema corpus used by Promptkit and by
|
|
Weatherreporter's final generated-text boundary.
|
|
|
|
### Work
|
|
|
|
1. Add Promptkit `v0.4.0` to `go.mod` and `go.sum`.
|
|
2. Create `internal/promptassets` with an embedded tree rooted at:
|
|
|
|
```text
|
|
assets/
|
|
prompts/
|
|
common/
|
|
daily/
|
|
today/
|
|
tomorrow/
|
|
hourly/
|
|
schemas/
|
|
```
|
|
|
|
Expose copied/read-only prompt and schema filesystem views and a
|
|
`Schema(id)` lookup for the four report schema IDs. Keep path constants
|
|
private except where a stable project contract is necessary.
|
|
3. Promote only the four `*_generated_text` prompt definitions from
|
|
`docs/roadmap/scriptorium`. Do not promote `weather.daily_report`.
|
|
4. Use the Scriptorium corpus's common system, data-package guidance, and
|
|
report-specific files as the behavioral base. Reconcile rather than
|
|
editorially rewrite:
|
|
|
|
- preserve its source weighting, hazard, precipitation, and style rules;
|
|
- add the optional `confidence` field instruction to Daily, Today, and
|
|
Hourly so all prompt text agrees with the canonical domain types and
|
|
schemas;
|
|
- retain Tomorrow's existing confidence instruction;
|
|
- prefer the corpus's correct Tomorrow daily framing over the stale
|
|
hourly-style file currently under `internal/reporttemplate/prompts`;
|
|
- make no unrelated prompt-tone or product-policy changes.
|
|
|
|
5. Every prompt definition must use:
|
|
|
|
- its existing prompt ID;
|
|
- version `1.0.0`;
|
|
- `default_profile: gemini-flash-latest`;
|
|
- one required `data_package` input with
|
|
`content_type: application/yaml`;
|
|
- JSON output with `validation_mode: json_schema`;
|
|
- a path inside the embedded schema root; and
|
|
- zero or omitted `repair_attempts`.
|
|
|
|
6. Move the four canonical schemas from `internal/reporttemplate/schemas` into
|
|
`internal/promptassets/assets/schemas`. Preserve their current application
|
|
contracts, including optional `confidence`, required summary/discussion,
|
|
array discussion for day-style reports, string discussion for Hourly,
|
|
`additionalProperties: false`, and the corrected Daily `$id` and title.
|
|
7. Change `internal/generatedtext` to load these canonical schemas from
|
|
`promptassets`. Remove schema and prompt-fragment ownership from
|
|
`internal/reporttemplate`; it should embed and render Markdown templates and
|
|
partials only.
|
|
8. Retain `docs/roadmap/scriptorium` as migration source evidence until the
|
|
final cleanup stage.
|
|
|
|
### Tests
|
|
|
|
- Add table-driven asset tests for exactly four prompt IDs and versions,
|
|
referenced-file resolution, default profiles, YAML input declarations,
|
|
output contracts, and schema lookup.
|
|
- Construct a Promptkit engine over the embedded sources in an external test
|
|
package and call `InspectPrompt` for every report. Do not use a provider or
|
|
credentials.
|
|
- Retain generated-text schema and typed-validation tests, now reading the
|
|
relocated canonical schemas.
|
|
- Assert that no active embedded prompt uses `local-heavy`,
|
|
`pipeline-weather/`, JSON input metadata, positive repair attempts, or
|
|
`weather.daily_report`.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/promptassets ./internal/generatedtext \
|
|
./internal/reporttemplate
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
One embedded source contains exactly four valid Promptkit prompt definitions
|
|
and four canonical schemas. Prompt inspection succeeds offline for every
|
|
report, and no duplicate runtime provider-facing schema remains.
|
|
|
|
## Stage 3: Define The Neutral Prompt Execution Contract
|
|
|
|
### Goal
|
|
|
|
Define the complete project-owned boundary between Weatherreporter and an LLM
|
|
execution adapter. This stage introduces no Promptkit imports and makes no
|
|
production orchestration changes.
|
|
|
|
### Work
|
|
|
|
1. Create `internal/promptexec` with the neutral contract fixed above. Include:
|
|
|
|
- exact prompt and profile inspection values;
|
|
- safe preparation provenance;
|
|
- execution result, validation, token usage, and timing;
|
|
- optional sensitive preparation/execution debug values;
|
|
- stable error categories and a typed capacity error; and
|
|
- small helpers for safe diagnostic bounding and copying.
|
|
|
|
2. The safe preparation value includes prompt ID/version/hash, rendered prompt
|
|
hash, input hashes, selected profile/backend, effective model name, output
|
|
format/validation/schema path, preparation start/end/duration, and the
|
|
managed data-package path. It excludes endpoints, API-key environment
|
|
names, full parameters, rendered messages, schema bodies, and input bodies.
|
|
3. The safe execution value includes Promptkit run ID, repeated prompt and
|
|
input provenance, selected profile/backend/model, generated-content hash,
|
|
usage, execution start/end/duration, completed validation status, bounded
|
|
validation diagnostics, and raw output bytes. It excludes endpoint and full
|
|
effective parameters.
|
|
4. Bound routine validation diagnostics to at most 10 entries and at most
|
|
1,024 bytes per entry, truncating on a valid UTF-8 boundary. Bound a
|
|
persisted safe error message to 2,048 bytes. Prefer stable categories and
|
|
fields over diagnostic prose.
|
|
5. Specify the execution lifecycle in interface comments and tests:
|
|
|
|
- inspection is side-effect-free;
|
|
- preparation invokes the callback exactly once only after successful
|
|
preparation;
|
|
- callback failure prevents provider execution;
|
|
- completed validation rejection is a result, not an operational error;
|
|
- operational failure does not fabricate a completed result; and
|
|
- sensitive debug values are populated only when explicitly requested.
|
|
6. Keep copying and truncation helpers internal to `promptexec`. Callers must
|
|
not be able to mutate byte slices, maps, or diagnostics retained inside
|
|
contract values.
|
|
|
|
### Tests
|
|
|
|
- Add compile-time test fakes that implement the complete interface without
|
|
importing Promptkit.
|
|
- Add table-driven tests for every error category, capacity fields, diagnostic
|
|
count and byte limits, UTF-8-safe truncation, error-message bounds, and
|
|
defensive copying.
|
|
- Assert through representative fixtures that safe contract values contain no
|
|
endpoint, credential, rendered message, schema body, input body, response
|
|
body, or full parameter map.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/promptexec
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
`internal/promptexec` provides a stable, documented, offline-tested contract
|
|
that can represent every inspection, preparation, execution, validation,
|
|
usage, debug, and failure outcome required by the roadmap. It has no Promptkit,
|
|
app, CLI, report, or state dependency.
|
|
|
|
## Stage 4: Implement The Promptkit Adapter
|
|
|
|
### Goal
|
|
|
|
Implement and thoroughly test Promptkit behind the Stage 3 contract without
|
|
modifying application or CLI orchestration.
|
|
|
|
### Work
|
|
|
|
1. Create `internal/adapters/promptkit`. Its constructor:
|
|
|
|
- uses `promptassets` through `WithPromptFS` and `WithSchemaFS`;
|
|
- selects one external profile directory through Promptkit config or one
|
|
profile file through `WithProfileFile`;
|
|
- optionally registers
|
|
`promptkit.LocalBackend(endpoint, concurrencyLimit)`;
|
|
- sets the transport timeout;
|
|
- accepts an injected Promptkit `LLMClient` only through an adapter-local
|
|
test constructor or option; and
|
|
- returns project-owned configuration errors.
|
|
|
|
2. Implement prompt and profile inspection with exact versions and explicit
|
|
field-by-field translation into `promptexec` values. Do not return or embed
|
|
Promptkit values.
|
|
3. Implement execution:
|
|
|
|
- call `PrepareExecution` with the exact requested version and optional
|
|
profile override;
|
|
- pass the exact YAML as
|
|
`InlineWithURI(dataPackagePath, string(bytes))`;
|
|
- immediately `defer handle.Discard()` after a successful prepare;
|
|
- map safe preparation details and invoke the preparation callback;
|
|
- call `RunPrepared` only after the callback succeeds;
|
|
- treat `ValidationFailed` as a completed result rather than an operational
|
|
error; and
|
|
- return no invented execution result for operational errors.
|
|
|
|
4. Translate Promptkit's public error sentinels with `errors.Is`, and
|
|
`CapacityError` with `errors.As`. Preserve caller cancellation and deadline
|
|
identities. Do not parse error strings.
|
|
5. Capture sensitive debug values only when requested. Use project-owned debug
|
|
structs and explicit field mapping rather than serializing whole Promptkit
|
|
values, so a future upstream field cannot silently enter an artifact.
|
|
6. Keep the prepared handle entirely within the adapter call. Discard it after
|
|
every success or failure path and never make it serializable or reusable.
|
|
|
|
### Tests
|
|
|
|
- Use an injected deterministic Promptkit model client. Never use a live
|
|
endpoint or credential.
|
|
- Cover prompt/profile inspection, profile override precedence, built-in and
|
|
external profiles, endpoint-only profiles, the conventional local backend,
|
|
local concurrency/capacity rejection, missing credentials, and constructor
|
|
validation.
|
|
- Cover preparation callback ordering, callback failure preventing generation,
|
|
exact inline bytes and URI, exact prompt version, successful execution,
|
|
completed schema rejection with raw output, operational generation and
|
|
validation failures with no partial result, cancellation, timeout, and
|
|
handle discard.
|
|
- Assert that safe values and errors do not contain rendered messages,
|
|
endpoints, schema bodies, input bodies, response bodies, credentials, or
|
|
full parameter maps.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/promptexec ./internal/adapters/promptkit
|
|
go test -race ./internal/adapters/promptkit
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
The Promptkit adapter is a complete offline-tested implementation of the
|
|
neutral contract. No Promptkit type is referenced outside the adapter, its
|
|
tests, or the external asset contract test, and no production path uses the
|
|
adapter yet.
|
|
|
|
## Stage 5: Introduce Durable State V2
|
|
|
|
### Goal
|
|
|
|
Add the final normal-artifact model and backward-compatible metadata reading
|
|
before app orchestration starts writing the new records. Sensitive debug
|
|
storage remains out of scope until Stage 6.
|
|
|
|
### Work
|
|
|
|
1. Add the new-run path members `Preparation` and `Execution`:
|
|
|
|
- preparation:
|
|
`preflight/<group>/<date>/prompt_preparation.<runID>.json`;
|
|
- execution:
|
|
`snapshots/<group>/<date>/prompt_execution.<runID>.json`.
|
|
|
|
Keep `workspace.preflight_dir` and its default physical directory unchanged.
|
|
Retain `Preflight` and `GeneratedTextResult` only as temporary legacy
|
|
write-path members required by the uncut Scriptorium workflows; mark them
|
|
for deletion in Stage 10. New Promptkit code must use only the new members.
|
|
2. Add `PromptPreparationArtifact` with schema
|
|
`weatherreporter.prompt_preparation.v1`. It represents either successful
|
|
preparation or a failed preparation receipt and includes schema version,
|
|
status, report/RunID, prompt identity, safe preparation provenance when
|
|
available, timing, data-package path, and a bounded classified error when
|
|
failed.
|
|
3. Add `PromptExecutionArtifact` with schema
|
|
`weatherreporter.prompt_execution.v1`. It represents success, completed
|
|
validation rejection, or operational failure and includes safe execution
|
|
provenance, validation, usage, timing, reached artifact paths, and a bounded
|
|
classified error when failed. It never embeds generated content.
|
|
4. Advance newly written metadata to `weatherreporter.metadata.v2`. Replace
|
|
`preflightPath` and `generatedTextResultPath` with `preparationPath` and
|
|
`executionPath`. Populate paths only after the corresponding artifact has
|
|
actually been saved.
|
|
5. Implement explicit v1/v2 metadata decoding:
|
|
|
|
- accept only the known v1 and v2 schema versions;
|
|
- normalize v1 `preflightPath` and `generatedTextResultPath` internally for
|
|
inspection;
|
|
- preserve v1 field names when a loaded v1 record is marshaled by
|
|
`inspect metadata`;
|
|
- write only v2 through the new Promptkit state APIs; the temporary
|
|
Scriptorium path may continue writing v1 until its Stage 10 removal; and
|
|
- never dual-write legacy aliases.
|
|
|
|
Direct inspection of v1 metadata, modules, data packages, sources, and
|
|
referenced artifacts remains available for any historical report ID.
|
|
Prior-snapshot reconstruction is required only for the four retained report
|
|
IDs; do not restore retired definitions solely for legacy comparison.
|
|
6. Add typed `SavePromptPreparation`, `SavePromptExecution`, and corresponding
|
|
typed load methods where inspection needs them. Temporarily retain the
|
|
Scriptorium write methods so the current production paths compile through
|
|
Stage 9; delete those methods during Stage 10. Keep long-term legacy support
|
|
read-only.
|
|
7. Update app/CLI result structs and JSON field names to
|
|
`preparationPath`, `executionPath`, and optional `llmDebugPath`; temporarily
|
|
adapt old orchestration so the tree compiles until the cutover stages.
|
|
Do not write v2 records through the Scriptorium path.
|
|
|
|
### Tests
|
|
|
|
- Add exact path, schema-defaulting, round-trip, required-field, atomic-write,
|
|
and unknown-version tests.
|
|
- Add v1 fixtures covering both generated-text and legacy preflight references;
|
|
verify list, metadata, modules, data-package, and source inspection.
|
|
- Verify that re-marshaled v1 inspection uses v1 field names and that v2 output
|
|
contains no deprecated aliases.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/state ./internal/app ./internal/cli
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
State can read historical v1 runs and write the complete normal v2 artifact
|
|
contract through the new APIs. The existing production generation path still
|
|
compiles and remains unchanged until cutover; no new Promptkit artifact uses a
|
|
Scriptorium-specific filename, schema, or JSON field.
|
|
|
|
## Stage 6: Add Secure LLM Debug Persistence
|
|
|
|
### Goal
|
|
|
|
Implement the explicitly enabled sensitive-debug store as a small, independently
|
|
auditable boundary before any CLI command can request debug capture.
|
|
|
|
### Work
|
|
|
|
1. Add a focused debug writer, owned by state or a narrow state-adjacent
|
|
package, for an explicit operator root outside normal artifact derivation.
|
|
It must:
|
|
|
|
- validate or create the root before collection or provider work;
|
|
- use `<root>/<report-id>/<valid-date>/<run-id>/`;
|
|
- write `preparation.json` and `execution.json` atomically;
|
|
- create directories with `0700` and files with `0600`;
|
|
- reject symlinks, path escape, non-directory roots, and unsafe
|
|
run/report/date segments; and
|
|
- return the per-run debug directory as the project-owned summary path.
|
|
|
|
2. Define explicit project-owned debug wire structs. Preparation debug may
|
|
contain rendered messages, structured-output schema, and effective
|
|
endpoint/parameters. Execution debug may additionally contain raw generated
|
|
output and completed validation details.
|
|
3. Exclude direct credentials, resolved environment credential values, and
|
|
arbitrary dependency objects. Use field-by-field mapping so future
|
|
`promptexec` or Promptkit fields are not serialized implicitly.
|
|
4. Make disabled debug behavior a no-op that performs no filesystem access.
|
|
A requested debug initialization or write error is returned to the caller;
|
|
it is never silently downgraded.
|
|
5. Keep this writer independent of normal state path derivation. Normal state
|
|
inspection must not discover or serve sensitive debug artifacts.
|
|
|
|
### Tests
|
|
|
|
- Verify exact grouping, atomic replacement, disabled behavior, and returned
|
|
per-run paths.
|
|
- Verify directory `0700` and file `0600` permissions on supported platforms.
|
|
- Cover traversal, absolute-segment, symlink-root, symlink-component,
|
|
non-directory, and invalid report/date/run segment rejection.
|
|
- Marshal representative debug fixtures and verify credentials and resolved
|
|
secret values are absent while the explicitly allowed diagnostic fields are
|
|
retained.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/state
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
Sensitive debug persistence is secure, explicitly rooted, atomic, isolated
|
|
from normal state, and fully tested. No production CLI path enables it yet.
|
|
|
|
## Stage 7: Add Promptkit Configuration, Composition, And Inspection
|
|
|
|
### Goal
|
|
|
|
Add the configuration and project-owned composition seams needed for cutover,
|
|
and centralize pre-collection prompt/profile validation. Retain Scriptorium as
|
|
the production generator until Stage 8 so this stage does not create a
|
|
half-cut-over runtime.
|
|
|
|
### Work
|
|
|
|
1. Add `config.PromptkitConfig` and nested local config with exactly:
|
|
|
|
- `profile`;
|
|
- `profile_file`;
|
|
- `profile_dir`;
|
|
- `timeout`, default `2m`;
|
|
- `local.endpoint`; and
|
|
- `local.concurrency_limit`, default `1`.
|
|
|
|
Scriptorium config remains temporarily because production generation has
|
|
not yet cut over. Promptkit validation rejects simultaneous profile
|
|
sources, non-positive transport timeout, an invalid nonblank local
|
|
endpoint, and negative concurrency. A blank local endpoint leaves `local`
|
|
unregistered; concurrency zero means unlimited.
|
|
2. Add an executor factory seam to `cli.Runner` using only project-owned types.
|
|
Production construction delegates to `internal/adapters/promptkit`; CLI
|
|
tests inject a fake factory. The factory creates one executor for an action,
|
|
not one per report.
|
|
3. Add a project-owned app inspection helper that:
|
|
|
|
- inspects the exact prompt ID and `report.Definition.PromptVersion`;
|
|
- verifies exactly one required `data_package` input with
|
|
`application/yaml`;
|
|
- verifies the expected JSON Schema output contract and declared default
|
|
profile;
|
|
- selects `promptkit.profile` when nonblank, otherwise the prompt default;
|
|
- inspects that explicit profile;
|
|
- rejects a profile requiring a direct API key, because Weatherreporter has
|
|
no direct-key configuration; and
|
|
- requires a nonblank environment value for any reported `APIKeyEnv`.
|
|
|
|
4. Return only safe, project-owned inspection values and classified errors.
|
|
The helper performs no collection, provider call, or durable write.
|
|
5. Unit-test production factory argument mapping, but do not yet wire it into
|
|
`generate` or `run`. Stage 8 performs that atomic production cutover.
|
|
|
|
### Tests
|
|
|
|
- Cover exact defaults, mutual profile-source exclusion, local endpoint
|
|
validation, zero/unlimited and negative concurrency, and timeout validation.
|
|
- Cover prompt/profile default and override precedence, prompt contract
|
|
mismatch, unsupported direct-key profiles, missing environment credentials,
|
|
and safe errors.
|
|
- Verify the CLI factory maps embedded assets, external profile source, local
|
|
backend, concurrency, and timeout exactly once without exposing Promptkit
|
|
types.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/config ./internal/app ./internal/cli \
|
|
./internal/adapters/promptkit
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
Configuration, engine construction, and pre-collection inspection are
|
|
decision-complete and offline-tested behind project-owned seams. Production
|
|
generation still follows the old path, so there is no dual execution mode.
|
|
|
|
## Stage 8: Cut Over Single-Report Execution And Failure Persistence
|
|
|
|
### Goal
|
|
|
|
Move all four `generate` commands to Promptkit prepared execution while
|
|
preserving report output behavior and inspectable partial failure. Debug CLI
|
|
enablement and final summary presentation are deferred to Stage 9.
|
|
|
|
### Work
|
|
|
|
1. Construct exactly one executor after configuration is loaded for a
|
|
`generate` action and pass it through `app.GenerateRequest`.
|
|
2. Isolate the current shared generation function before changing it:
|
|
|
|
- move the Scriptorium implementation behind a temporary, unexported legacy
|
|
batch helper;
|
|
- make `GenerateDetailed` call the new Promptkit implementation described
|
|
below;
|
|
- leave `RunBatchDetailed` calling only the legacy helper until Stage 10;
|
|
and
|
|
- do not add a runtime switch, fallback, or dual invocation for the same
|
|
report.
|
|
|
|
This temporary split is solely a staging seam. Stage 10 moves batches to
|
|
the Promptkit implementation and deletes the helper.
|
|
3. Reorder `GenerateDetailed`:
|
|
|
|
1. resolve the report and RunID;
|
|
2. require the injected executor;
|
|
3. run the Stage 7 exact prompt/profile/credential inspection; and only then
|
|
4. collect weather.
|
|
|
|
Inspection failures occur before managed run artifacts exist.
|
|
4. Implement the new project-owned report-generation core with the injected
|
|
executor and the Stage 5 artifact contract:
|
|
|
|
- build, serialize once, and save the data package;
|
|
- use those exact serialized bytes for execution;
|
|
- persist a failed preparation receipt and v2 metadata when preparation
|
|
fails;
|
|
- use the preparation callback to save successful preparation provenance
|
|
and metadata before provider work;
|
|
- after execution success or completed validation rejection, save exact raw
|
|
output, then the execution artifact, then metadata;
|
|
- after operational execution failure, save a failure execution receipt and
|
|
metadata without inventing raw output;
|
|
- on Promptkit validation rejection, return a classified report failure
|
|
after preserving raw output and bounded diagnostics;
|
|
- on Promptkit validation success, run the existing typed
|
|
`internal/generatedtext` validation, normalize/save generated text, build
|
|
and save render context, and render managed Markdown as before.
|
|
|
|
5. Make partial failures return a non-nil `ReportResult` whenever a RunID and
|
|
inspectable paths exist. Preserve preparation/execution/raw/normalized/
|
|
context/report paths reached before preparation, capacity, generation,
|
|
Promptkit validation, domain validation, template, copy, or notification
|
|
failure.
|
|
6. Remove Scriptorium result values from the single-report app requests,
|
|
results, and test fakes used by the generate path. App tests depend only on
|
|
`promptexec.Executor`.
|
|
Retain the old adapter and configuration only for the still-uncut batch
|
|
path.
|
|
7. Do not expose `--llm-debug-dir` yet. Pass `CaptureDebug: false` during this
|
|
stage; Stage 9 adds debug orchestration without changing provider execution
|
|
semantics.
|
|
|
|
### Tests
|
|
|
|
- Add representative offline app workflows for each of the four
|
|
reports, using real state, generated-text validation, contexts, and
|
|
templates with a fake executor.
|
|
- Verify inspection and credential checks occur before collection.
|
|
- Verify exact profile override/default precedence and exact prompt versions.
|
|
- Verify persistence ordering by observing that execution is not called until
|
|
the successful preparation artifact and metadata exist.
|
|
- Cover preparation failure, callback/state failure, capacity rejection,
|
|
credential disappearance at execution, cancellation, deadline, generation
|
|
failure, operational validation failure, completed schema rejection,
|
|
generated-text domain failure, template failure, output-copy failure, and
|
|
notification failure.
|
|
- Verify partial results retain every reached normal-artifact path without
|
|
prompt or response content.
|
|
- Verify single-report commands invoke only Promptkit and batch commands invoke
|
|
only the temporary legacy helper during this intermediate stage.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/app ./internal/cli
|
|
go test -race ./internal/app ./internal/adapters/promptkit
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
Every single-report command uses one Promptkit engine, performs inspection
|
|
before collection, persists preparation before provider work, and produces the
|
|
same managed Markdown/template and notification outcomes through project-owned
|
|
contracts. Sensitive debug capture remains disabled.
|
|
|
|
## Stage 9: Add Single-Report Debug And CLI Summary Behavior
|
|
|
|
### Goal
|
|
|
|
Expose the opt-in debug workflow and finalize single-report CLI success and
|
|
partial-failure summaries without expanding the normal artifact boundary.
|
|
|
|
### Work
|
|
|
|
1. Parse `--llm-debug-dir PATH` for all four generate commands. Do not add a
|
|
YAML debug switch.
|
|
2. Validate or create the debug root before prompt inspection or weather
|
|
collection. A requested initialization failure is terminal and occurs
|
|
before managed run artifacts exist.
|
|
3. Pass debug intent through project-owned CLI and app request types. When
|
|
enabled, request sensitive debug values from the executor:
|
|
|
|
- write preparation debug inside the preparation callback, after the normal
|
|
preparation artifact is saved and before metadata is finalized and
|
|
provider work begins;
|
|
- if the debug write fails, return the callback error so `RunPrepared` is
|
|
not called;
|
|
- write execution debug immediately after receiving a completed execution
|
|
result and before continuing with downstream validation/rendering; and
|
|
- treat every requested debug write failure as terminal while preserving
|
|
already reached normal artifacts.
|
|
|
|
4. Set `llmDebugPath` only after the per-run debug directory contains at least
|
|
one successfully written debug artifact. Never copy sensitive debug content
|
|
into a normal artifact, log, error, or summary.
|
|
5. Finalize human and JSON CLI summaries around project-owned
|
|
`preparationPath`, `executionPath`, and optional `llmDebugPath`. Include
|
|
reached paths on partial failure; omit absent paths instead of inventing
|
|
them.
|
|
6. Preserve existing output behavior: quiet mode suppresses successful human
|
|
output, JSON output remains machine-readable, and failures return a
|
|
non-zero status with safe classified text.
|
|
|
|
### Tests
|
|
|
|
- Add CLI parser and help tests for `--llm-debug-dir` on all four generate
|
|
commands and its absence from YAML configuration.
|
|
- Verify debug initialization precedes inspection and collection.
|
|
- Verify preparation debug persistence precedes provider execution and that a
|
|
write failure prevents the model-client call.
|
|
- Cover execution debug success/failure, partial normal-artifact retention,
|
|
`llmDebugPath` population rules, quiet mode, JSON summaries, and safe
|
|
failure text.
|
|
- Assert that summaries and routine logs contain no rendered prompt, schema
|
|
body, input body, generated body, endpoint, full parameters, or credential.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/app ./internal/cli ./internal/state
|
|
go test -race ./internal/app ./internal/adapters/promptkit
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
All single-report commands provide the complete Promptkit, v2 persistence,
|
|
opt-in debug, and summary behavior required by the roadmap. Debug failure
|
|
ordering is enforced without exposing sensitive content.
|
|
|
|
## Stage 10: Cut Over Batches And Remove Scriptorium
|
|
|
|
### Goal
|
|
|
|
Complete the production cutover, share one engine through each sequential
|
|
batch, and delete the retired dependency boundary.
|
|
|
|
### Work
|
|
|
|
1. Construct one executor in the CLI for each `run morning` or `run evening`
|
|
invocation and pass it through `BatchRequest` to every report.
|
|
2. Parse and initialize `--llm-debug-dir` for run commands using the same
|
|
policy as generate commands.
|
|
3. Before collection, inspect the complete candidate set:
|
|
|
|
- morning: Today, Tomorrow, and Daily;
|
|
- evening: Tomorrow and Daily.
|
|
|
|
Inspect exact prompt versions, validate declared input/output contracts,
|
|
resolve unique effective profiles, and enforce credential availability.
|
|
Daily inspection occurs before its collection-dependent future dates are
|
|
known because every eligible Daily run uses the same exact prompt contract.
|
|
4. Collect once, plan the batch as before, and execute every planned report
|
|
sequentially through the shared executor, Stage 8 generation core, and
|
|
store. Delete the temporary legacy batch helper after this call site moves.
|
|
Preserve continuation after independent report failures and existing
|
|
batch-notification gating.
|
|
5. Include preparation, execution, and optional debug paths in each batch
|
|
item, including failed items when those paths were reached.
|
|
6. Replace Scriptorium configuration entirely with Promptkit configuration.
|
|
Update defaults, validation, examples, config tests, and all construction
|
|
sites. Explicitly reject a top-level `scriptorium:` key with an actionable
|
|
migration error even though the general YAML loader is currently
|
|
permissive; silently ignoring a former execution configuration is unsafe.
|
|
Do not translate it or add a dual-run mode. The maintained examples contain
|
|
only `promptkit:`.
|
|
7. Delete:
|
|
|
|
- `internal/adapters/scriptorium`;
|
|
- the Scriptorium renderer interface and all result/request types;
|
|
- the legacy state path members and Scriptorium write methods retained in
|
|
Stage 5;
|
|
- subprocess fakes and subprocess-specific tests;
|
|
- direct-Markdown branches and remaining Scriptorium defaults; and
|
|
- dead helpers used only by the old adapter.
|
|
|
|
8. Run `go mod tidy` and verify the module graph contains Promptkit `v0.4.0`
|
|
and no Scriptorium module or local replacement.
|
|
|
|
### Tests
|
|
|
|
- Add batch tests proving one factory call/engine per CLI invocation, one
|
|
executor shared across all reports, one collection, sequential ordering,
|
|
later-report continuation, per-report path retention, and notification only
|
|
after complete success.
|
|
- Cover a capacity failure for one report followed by a later independent
|
|
report and confirm no retry occurs.
|
|
- Cover debug grouping for multiple reports and multiple Daily dates.
|
|
- Update config tests for exact defaults, mutual exclusivity, local endpoint,
|
|
zero/unlimited and negative concurrency, and maintained examples.
|
|
- Run:
|
|
|
|
```sh
|
|
go test ./internal/config ./internal/app ./internal/cli
|
|
go test -race ./internal/app ./internal/adapters/promptkit
|
|
go test ./...
|
|
go run ./cmd/weatherreporter --help
|
|
git diff --check
|
|
```
|
|
|
|
### Exit Gate
|
|
|
|
All production generation paths use Promptkit. A batch owns one engine and
|
|
continues sequentially under existing failure policy. No Scriptorium code,
|
|
configuration field, subprocess path, or dependency remains.
|
|
|
|
## Stage 11: Update Canonical Documentation And Complete Verification
|
|
|
|
### Goal
|
|
|
|
Make current-state documentation match the completed implementation, remove
|
|
migration-only source material, and perform repository-wide verification.
|
|
|
|
### Work
|
|
|
|
1. Update canonical current-state owners in the same migration change:
|
|
|
|
- `docs/development.md` for repository orientation, package map, task
|
|
routing, and validation language;
|
|
- `docs/policy/architecture.md` for the Promptkit adapter boundary,
|
|
generated-text-only flow, prepared execution, and four-report product;
|
|
- `docs/policy/testing.md` for Promptkit/provider fakes instead of
|
|
subprocess fakes and the offline external-boundary rule;
|
|
- `docs/cli.md` for four commands, `--llm-debug-dir`, and renamed summary
|
|
fields;
|
|
- `docs/config.md` and `examples/` for the exact Promptkit contract;
|
|
- `docs/operations.md` for v2 paths, sensitive debug retention and
|
|
permissions, preparation/execution lifecycle, and v1 inspection;
|
|
- `docs/troubleshooting.md` for inspection, profile/credential,
|
|
preparation, capacity, execution, validation, and debug failures;
|
|
- relevant `docs/internal/` files for app, CLI, report, prompt input,
|
|
generated text, templates, state, collection, briefing, facts, changes,
|
|
and package boundaries; and
|
|
- `docs/templates.md` for Promptkit-generated prose and the relocated schema
|
|
owner.
|
|
|
|
2. Replace `docs/integrations/scriptorium.md` and
|
|
`docs/internal/scriptorium-adapter.md` with canonical Promptkit integration
|
|
and adapter documents. The integration document owns the logical
|
|
prompt/profile/schema and durable compatibility contract; the internal
|
|
document owns construction, mapping, lifecycle, and tests. Link rather than
|
|
duplicate the full CLI/config/operations references.
|
|
3. Search all non-roadmap current-state documentation, examples, Go code,
|
|
tests, help output, and module metadata for stale `scriptorium`,
|
|
`local-heavy`, retired report commands, old artifact field names, old
|
|
filenames, and old metadata versions. Retain old names only in explicit v1
|
|
compatibility code/tests and historical roadmap discussion.
|
|
4. Remove `docs/roadmap/scriptorium` after confirming every promoted runtime
|
|
asset is represented under `internal/promptassets` and covered by asset
|
|
tests. The unused legacy Daily Markdown prompt is deleted with this source
|
|
directory.
|
|
5. Keep the feature roadmap and this plan as temporary migration documents
|
|
while implementation is under review. Once the implementation is accepted,
|
|
mark both complete and move every still-useful contract to its canonical
|
|
current-state owner before later archival or removal.
|
|
6. Review all added links, fenced examples, flags, field names, defaults,
|
|
schema IDs, file paths, and version strings against executable sources.
|
|
Confirm examples contain no credentials or private infrastructure values.
|
|
|
|
### Final Verification
|
|
|
|
Run:
|
|
|
|
```sh
|
|
gofmt -w <all changed Go files>
|
|
go mod tidy
|
|
go test ./...
|
|
go test -race ./...
|
|
go run ./cmd/weatherreporter --help
|
|
git diff --check
|
|
```
|
|
|
|
Then verify explicitly:
|
|
|
|
- `go list -m gitea.maximumdirect.net/eric/promptkit` reports `v0.4.0`;
|
|
- no committed `go.work`, `replace`, secret fixture, or live-provider test
|
|
exists;
|
|
- the maintained examples load through config tests;
|
|
- all four embedded prompts inspect at exact version `1.0.0`;
|
|
- no runtime prompt requests repair attempts;
|
|
- no ordinary artifact or CLI summary includes rendered prompts, schema
|
|
bodies, input bodies, generated bodies, provider endpoints, or credentials;
|
|
- v1 metadata fixtures remain inspectable and new runs write only v2;
|
|
- help exposes only Daily, Today, Tomorrow, Hourly, morning, and evening; and
|
|
- managed Markdown remains the only Distributor upload source.
|
|
|
|
### Exit Gate
|
|
|
|
Every completion criterion in the feature roadmap is demonstrably satisfied.
|
|
Current-state documentation describes the implementation rather than the
|
|
migration, and the repository passes all final verification commands.
|
|
|
|
## Open Questions
|
|
|
|
None. The feature roadmap and this implementation plan contain all product,
|
|
architecture, configuration, compatibility, security, sequencing, and
|
|
verification decisions required to begin implementation.
|