Files
weatherreporter/docs/roadmap/profile-comparison.md

399 lines
18 KiB
Markdown

# LLM Profile Comparison Roadmap
Status: Implemented; retained as the feature roadmap.
## Purpose
Prompt development currently requires separate Weatherreporter invocations to
compare several LLM profiles. Those invocations may collect different weather
snapshots or rebuild inputs at different times, making model output harder to
compare and slowing prompt iteration.
Weatherreporter should provide a first-class `compare` command that resolves
one report, prepares one exact data package, executes the same prompt and data
package concurrently through several explicitly selected Promptkit profiles,
and publishes a self-contained local comparison bundle.
An illustrative invocation is:
```sh
weatherreporter compare daily \
--date 2026-08-24 \
--profile weather-light \
--profile weather-balanced \
--profile weather-deep
```
This is a prompt-development workflow, not an automated model evaluator. Its
output gives a maintainer consistent evidence for human comparison without
assigning scores or selecting a winner.
## Prerequisite
Configurable output directories are implemented. Profile comparison must reuse
the current [configuration reference](../config.md) and [operations
guide](../operations.md) rather than introduce a second destination policy.
## User Intent
The command is intended for deliberate evaluation of multiple profiles,
including sets of eight to twelve candidate models. Concurrency is part of the
feature, not a future optimization. Promptkit should retain ownership of
backend-specific capacity, while Weatherreporter owns comparison-wide
coordination, cancellation, deterministic results, and artifact publication.
Every profile must receive byte-for-byte identical prompt input. Weather data,
derived facts, modules, prompt metadata, and serialized YAML must not be
recollected or rebuilt separately for individual profiles.
Comparison bundles are explicitly requested, operator-owned development
outputs. They are not Weatherreporter state, are never read implicitly by a
later run, and do not weaken the ordinary stateless execution model.
## Command Contract
The command form is:
```text
weatherreporter compare REPORT [options]
```
`REPORT` accepts the implemented generated-text reports: `daily`, `today`,
`tomorrow`, and `hourly`. Report-date behavior matches `generate`: `daily`
requires `--date`, `today` may accept an explicit date or use the current local
date, and the remaining report types retain their existing period policies.
The command accepts the applicable common generation options, including
`--config`, `--units`, `--tz`, `--date`, `--llm-debug-dir`, and `--quiet`, plus:
- repeatable `--profile PROFILE_ID` selections;
- `--out-dir PATH` for the exact comparison-bundle directory; and
- `--replace` to authorize guarded replacement of a recognized existing
comparison bundle.
At least two distinct, nonblank profile IDs are required. Their command-line
order is significant and is preserved in filenames, summaries, and
`comparison.json`. Duplicate profile IDs are rejected rather than silently
deduplicated or executed twice.
Profiles are always explicit for this command. `promptkit.profile` does not add
or replace a comparison selection, but all other effective Promptkit settings,
profile-source precedence, local backend configuration, credential lookup, and
profile overrides remain in force.
The initial feature has no Weatherreporter-specific concurrency flag or
artificial profile-count ceiling. The explicit profile list bounds the
comparison, and Promptkit owns capacity enforcement for each selected backend.
## Preparation And Execution Invariants
A comparison has this logical lifecycle:
1. Parse and validate the report, date, profile list, configuration, output
destination, and replacement authorization.
2. Resolve the report definition, valid period, prompt identity, and default
output name once.
3. Inspect the exact prompt once and preflight every selected profile,
including its effective backend, model, and required credential
availability, before weather collection.
4. Collect weather data exactly once.
5. Build collected and derived facts, the module snapshot, briefing metadata,
and the prompt data package exactly once.
6. Marshal the data package to one immutable YAML byte sequence exactly once.
7. Execute the exact prompt version concurrently for every selected profile,
passing the same immutable YAML bytes to every execution.
8. Validate and render each profile result independently from the shared
deterministic inputs.
9. Assemble results in requested-profile order and publish one coherent
comparison bundle.
This lifecycle describes the required end-state behavior rather than an
implementation-stage sequence.
No profile execution may cause recollection, report re-resolution, module
rebuilding, or data-package remarshalling. Prompt execution may perform
Promptkit-owned validation or repair behavior, but Weatherreporter does not
retry a failed comparison execution independently.
## Concurrency And Cancellation
Weatherreporter starts one execution for each preflighted profile and permits
them to run concurrently through one shared, concurrency-safe Promptkit
executor. Promptkit's engine-local backend pools remain authoritative for
backend concurrency and waiting capacity. Profiles routed to a limited local
backend therefore respect its configured limit, while profiles routed to
other backends may proceed independently.
Weatherreporter must not add a second semaphore that obscures or overrides
Promptkit's backend policy. It must safely coordinate goroutine lifecycles,
result collection, debug callbacks, and output assembly without data races.
One profile failure does not cancel its peers. Provider, capacity, validation,
and rendering failures are recorded for that profile while other executions
continue. Cancellation or deadline expiration of the comparison command is
propagated to every outstanding execution, prevents new publication, and is
joined without leaking goroutines.
Completion order must not affect filenames, manifest order, CLI summaries, or
error aggregation. Those outputs always follow the original `--profile`
order.
## Output Destination
Without `--out-dir`, Weatherreporter derives a comparison directory from the
resolved report's existing default Markdown filename by removing `.md` and
prefixing `comparison-`:
| Report output | Comparison directory |
| --- | --- |
| `today.md` | `comparison-today/` |
| `tomorrow.md` | `comparison-tomorrow/` |
| `hourly.md` | `comparison-hourly/` |
| `daily-2026-08-24.md` | `comparison-daily-2026-08-24/` |
The derived directory is created beneath `output.directory` when configured,
or beneath the present working directory otherwise. An explicit `--out-dir`
is the exact bundle directory, resolves relative to the present working
directory when necessary, and overrides `output.directory` completely.
All destination selection and validation completes before weather collection.
The resolved comparison directory is returned in the command's structured
result.
## Comparison Bundle
A successful three-profile comparison has a flat layout:
```text
comparison-daily-2026-08-24/
├── comparison.json
├── data-package.yml
├── 01-weather-light.md
├── 02-weather-balanced.md
└── 03-weather-deep.md
```
`data-package.yml` contains the exact YAML bytes passed to every Promptkit
execution. It is written once and its SHA-256 digest is recorded in the
manifest.
Each report filename begins with its one-based, zero-padded selection position
and a filesystem-safe representation of the requested logical profile ID. The
safe representation must not permit absolute paths, traversal, separators, or
control characters. The manifest retains the exact case-sensitive profile ID,
so filename normalization never becomes the authority for profile identity.
`comparison.json` is the authoritative index for the bundle. It uses an
explicit schema version and records safe comparison information including:
- comparison identity and start and finish timestamps;
- report ID, resolved valid period, and effective timezone;
- prompt ID, version, and inspected prompt hash;
- the relative data-package filename and SHA-256 digest;
- total, succeeded, and failed profile counts; and
- one ordered result per requested profile containing the exact profile ID,
resolved backend and model, relative report filename when present,
execution and validation status, and safe error information when failed.
The manifest and normal command summary must not contain credentials, provider
request bodies, raw model output, rendered prompts, schemas, provider
endpoints, or other content-rich diagnostics. The explicit data package and
generated reports contain the development material the user requested and
must be handled as operator-owned potentially sensitive output.
## Failure And Publication Policy
Failure before concurrent execution, including invalid profiles, missing
credentials, collection failure, preparation failure, or unsafe destination,
publishes no comparison bundle and performs no model calls where the failure
is discoverable during preflight.
After execution begins, Weatherreporter waits for every non-cancelled profile.
If one or more profiles fail, it still publishes a coherent partial bundle
containing `data-package.yml`, every successfully rendered report, and a
manifest describing all successes and failures. It then returns a non-zero
exit status. A failed profile has no report file unless a future contract
explicitly introduces a separately named diagnostic artifact.
Bundle contents are staged outside the destination and published only after
the manifest is complete. Ordinary publication accepts only an absent or empty
target directory. A nonempty existing directory fails without modification
unless `--replace` is present.
`--replace` may replace only the exact resolved target and must reject broad or
unsafe targets such as a filesystem root, the present working directory, a
symlink, or an unrecognized nonempty directory. A recognized prior bundle must
contain a valid Weatherreporter comparison manifest. Replacement publishes the
new complete or coherent partial bundle as a unit, prevents stale reports from
the prior comparison from surviving, and preserves or restores the prior
bundle if the final replacement operation fails.
An interrupted or cancelled comparison does not replace an existing bundle.
Temporary staging artifacts are cleaned up on ordinary failure and
cancellation without scanning or modifying unrelated directories.
## Prompt Debugging
The existing `--llm-debug-dir` mechanism remains available. Concurrent
comparison executions require distinct, deterministic debug identities that
include the comparison and exact profile selection so callbacks cannot collide
or overwrite another profile's artifacts.
Debug writing must be concurrency-safe and retain the existing permission,
redaction, explicit-opt-in, and path-containment guarantees. Debug artifacts
remain separate from the comparison bundle; the bundle does not implicitly
enable full Promptkit diagnostics.
## Notification Policy
Profile comparisons never invoke Distributor notification, even when
notification is enabled in the effective configuration. Comparison reports
are local development artifacts rather than ordinary report publications.
Adding comparison publication or upload behavior would require a separate
accepted feature scope and explicit operator authorization.
## Architectural End State
Application orchestration exposes a reusable prepared-report boundary that
contains the resolved report, shared collected and derived facts, module
snapshot, briefing metadata, generated-text handler, render inputs, and exact
serialized data package. That boundary is immutable during concurrent profile
execution.
Ordinary `generate` behavior continues to prepare once and execute once.
`compare` prepares once and executes many without duplicating the generation
workflow or calling `GenerateDetailed` in a loop. Shared preparation,
profile-specific Promptkit execution, structured-output validation, rendering,
and artifact publication remain distinct responsibilities.
The Promptkit adapter remains the only owner of dependency-specific types and
engine calls. The CLI owns parsing and user-facing summaries. The configuration
package owns configuration. Application orchestration owns comparison order,
concurrency lifecycle, failure aggregation, and bundle publication. Domain,
prompt-input, generated-text, and template packages retain their existing
deterministic contracts.
## Scope
The completed feature includes:
- the `compare` CLI command for every implemented generated-text report;
- repeatable explicit profile selection and validation;
- configured and CLI output-directory integration through the implemented
destination policy;
- one-time report resolution, collection, deterministic preparation, and YAML
serialization;
- concurrent execution through one Promptkit executor with backend capacity
respected;
- independent validation and rendering with deterministic ordered results;
- the flat, versioned comparison-bundle contract;
- safe filename derivation and data-package hashing;
- coherent partial-result publication and non-zero failure behavior;
- guarded whole-bundle replacement through `--replace`;
- comparison-aware, concurrency-safe optional prompt debugging;
- explicit suppression of Distributor notification;
- structured normal and quiet-mode CLI behavior consistent with existing
commands;
- focused race-safe tests across configuration, CLI, application,
Promptkit-adapter, rendering, and filesystem boundaries; and
- updates to every affected canonical user, operator, architecture,
integration, and internal document.
## Compatibility
The feature is additive. Existing `generate` and `run` commands, report
definitions, profile defaults, configuration, output filenames, notification
behavior, and exit contracts remain unchanged.
The comparison manifest and bundle layout begin as versioned contracts. They
do not become inputs accepted by Weatherreporter, and no backward-compatible
replay or long-term archive guarantee is implied beyond identifying the schema
used to interpret a produced bundle.
## Testing Expectations
Tests should provide durable coverage for:
- report and date parsing consistent with `generate`;
- rejection of fewer than two profiles, blanks, and duplicates;
- inspection of the exact prompt and every profile before collection;
- no collection or model execution after a preflight failure;
- exactly one weather collection and one preparation for several profiles;
- byte-for-byte identical data-package input in every execution;
- observable concurrent execution through a concurrency-safe fake executor;
- respect for Promptkit-owned backend capacity in an assembled adapter test
where that integration adds distinct confidence;
- deterministic filenames, manifest order, summaries, and errors under varied
completion order;
- continuation and coherent partial publication after one profile fails;
- cancellation propagation, goroutine completion, and preservation of an
existing destination;
- destination precedence and each derived default directory;
- safe filename handling for unusual valid profile IDs;
- absent, empty, occupied, symlinked, unsafe, recognized, and unrecognized
replacement targets;
- removal of stale prior report files during authorized whole-bundle
replacement;
- exact package digest and manifest/result consistency;
- concurrency-safe, non-colliding opt-in debug artifacts; and
- absence of Distributor calls for complete and partial comparisons.
Concurrency and replacement behavior require race-enabled and consequential
failure-path coverage. Tests must remain deterministic, offline, credential
free, and independent of real Promptkit providers or machine-specific paths.
## Documentation End State
Once implemented, the [CLI reference](../cli.md) owns command syntax, flags,
summary, and exit behavior. The [operations guide](../operations.md) owns the
bundle lifecycle, replacement procedure, sensitivity guidance, and practical
prompt-comparison workflow. The [architecture policy](../policy/architecture.md)
owns the statelessness, concurrency, notification, and publication invariants.
The [Promptkit integration guide](../integrations/promptkit.md) should describe
the consumer-visible multi-profile execution boundary without duplicating
Promptkit's backend-capacity reference. The versioned manifest and flat bundle
format belong in a focused contract under `docs/integrations/`. App
orchestration, prompt input, generated text, prompt debugging, and bundle
publication mechanics belong in focused documents under `docs/internal/`.
Current-state documentation must not describe profile comparison as available
until the implementation lands.
## Non-Goals
This roadmap does not introduce:
- automatic model scoring, ranking, recommendation, or winner selection;
- semantic or textual diff generation between reports;
- repeated sampling of one profile or statistical evaluation;
- prompt or profile editing through Weatherreporter;
- replaying a saved data package as command input;
- comparing several report types in one command;
- Weatherreporter-owned backend concurrency or queue configuration;
- automatic retries beyond Promptkit's existing execution contract;
- Distributor upload or other external publication;
- comparison history, indexing, retention, cleanup schedules, or implicit
discovery of prior bundles; or
- changes to ordinary report content or normal generation behavior.
Any later automated evaluation, replay, sampling, or publication feature
requires a separate accepted roadmap.
## Completion Criteria
The feature is complete when a maintainer can select several Promptkit
profiles, have them execute concurrently against one exact prepared report
package, and receive a safe, flat, deterministic comparison bundle whose
manifest accurately describes every success and failure. Configured and
explicit destinations must follow the accepted output policy, replacement must
never mix or silently destroy unrelated contents, cancellation and partial
failure must be race-safe, ordinary notification must remain disabled, and all
affected canonical documentation must describe the implemented behavior.
## Open Questions
None. The scope, prerequisites, user intent, and target behavior are defined
above.