18 KiB
LLM Profile Comparison Roadmap
Status: Implemented; retained temporarily for post-implementation review.
Purpose
Prompt development currently requires separate Weatherreporter invocations to compare several LLM profiles. Those invocations may collect different weather snapshots or rebuild inputs at different times, making model output harder to compare and slowing prompt iteration.
Weatherreporter should provide a first-class compare command that resolves
one report, prepares one exact data package, executes the same prompt and data
package concurrently through several explicitly selected Promptkit profiles,
and publishes a self-contained local comparison bundle.
An illustrative invocation is:
weatherreporter compare daily \
--date 2026-08-24 \
--profile weather-light \
--profile weather-balanced \
--profile weather-deep
This is a prompt-development workflow, not an automated model evaluator. Its output gives a maintainer consistent evidence for human comparison without assigning scores or selecting a winner.
Prerequisite
Configurable output directories are implemented. Profile comparison must reuse the current configuration reference and operations guide rather than introduce a second destination policy.
User Intent
The command is intended for deliberate evaluation of multiple profiles, including sets of eight to twelve candidate models. Concurrency is part of the feature, not a future optimization. Promptkit should retain ownership of backend-specific capacity, while Weatherreporter owns comparison-wide coordination, cancellation, deterministic results, and artifact publication.
Every profile must receive byte-for-byte identical prompt input. Weather data, derived facts, modules, prompt metadata, and serialized YAML must not be recollected or rebuilt separately for individual profiles.
Comparison bundles are explicitly requested, operator-owned development outputs. They are not Weatherreporter state, are never read implicitly by a later run, and do not weaken the ordinary stateless execution model.
Command Contract
The command form is:
weatherreporter compare REPORT [options]
REPORT accepts the implemented generated-text reports: daily, today,
tomorrow, and hourly. Report-date behavior matches generate: daily
requires --date, today may accept an explicit date or use the current local
date, and the remaining report types retain their existing period policies.
The command accepts the applicable common generation options, including
--config, --units, --tz, --date, --llm-debug-dir, and --quiet, plus:
- repeatable
--profile PROFILE_IDselections; --out-dir PATHfor the exact comparison-bundle directory; and--replaceto authorize guarded replacement of a recognized existing comparison bundle.
At least two distinct, nonblank profile IDs are required. Their command-line
order is significant and is preserved in filenames, summaries, and
comparison.json. Duplicate profile IDs are rejected rather than silently
deduplicated or executed twice.
Profiles are always explicit for this command. promptkit.profile does not add
or replace a comparison selection, but all other effective Promptkit settings,
profile-source precedence, local backend configuration, credential lookup, and
profile overrides remain in force.
The initial feature has no Weatherreporter-specific concurrency flag or artificial profile-count ceiling. The explicit profile list bounds the comparison, and Promptkit owns capacity enforcement for each selected backend.
Preparation And Execution Invariants
A comparison has this logical lifecycle:
- Parse and validate the report, date, profile list, configuration, output destination, and replacement authorization.
- Resolve the report definition, valid period, prompt identity, and default output name once.
- Inspect the exact prompt once and preflight every selected profile, including its effective backend, model, and required credential availability, before weather collection.
- Collect weather data exactly once.
- Build collected and derived facts, the module snapshot, briefing metadata, and the prompt data package exactly once.
- Marshal the data package to one immutable YAML byte sequence exactly once.
- Execute the exact prompt version concurrently for every selected profile, passing the same immutable YAML bytes to every execution.
- Validate and render each profile result independently from the shared deterministic inputs.
- Assemble results in requested-profile order and publish one coherent comparison bundle.
This lifecycle describes the required end-state behavior rather than an implementation-stage sequence.
No profile execution may cause recollection, report re-resolution, module rebuilding, or data-package remarshalling. Prompt execution may perform Promptkit-owned validation or repair behavior, but Weatherreporter does not retry a failed comparison execution independently.
Concurrency And Cancellation
Weatherreporter starts one execution for each preflighted profile and permits them to run concurrently through one shared, concurrency-safe Promptkit executor. Promptkit's engine-local backend pools remain authoritative for backend concurrency and waiting capacity. Profiles routed to a limited local backend therefore respect its configured limit, while profiles routed to other backends may proceed independently.
Weatherreporter must not add a second semaphore that obscures or overrides Promptkit's backend policy. It must safely coordinate goroutine lifecycles, result collection, debug callbacks, and output assembly without data races.
One profile failure does not cancel its peers. Provider, capacity, validation, and rendering failures are recorded for that profile while other executions continue. Cancellation or deadline expiration of the comparison command is propagated to every outstanding execution, prevents new publication, and is joined without leaking goroutines.
Completion order must not affect filenames, manifest order, CLI summaries, or
error aggregation. Those outputs always follow the original --profile
order.
Output Destination
Without --out-dir, Weatherreporter derives a comparison directory from the
resolved report's existing default Markdown filename by removing .md and
prefixing comparison-:
| Report output | Comparison directory |
|---|---|
today.md |
comparison-today/ |
tomorrow.md |
comparison-tomorrow/ |
hourly.md |
comparison-hourly/ |
daily-2026-08-24.md |
comparison-daily-2026-08-24/ |
The derived directory is created beneath output.directory when configured,
or beneath the present working directory otherwise. An explicit --out-dir
is the exact bundle directory, resolves relative to the present working
directory when necessary, and overrides output.directory completely.
All destination selection and validation completes before weather collection. The resolved comparison directory is returned in the command's structured result.
Comparison Bundle
A successful three-profile comparison has a flat layout:
comparison-daily-2026-08-24/
├── comparison.json
├── data-package.yml
├── 01-weather-light.md
├── 02-weather-balanced.md
└── 03-weather-deep.md
data-package.yml contains the exact YAML bytes passed to every Promptkit
execution. It is written once and its SHA-256 digest is recorded in the
manifest.
Each report filename begins with its one-based, zero-padded selection position and a filesystem-safe representation of the requested logical profile ID. The safe representation must not permit absolute paths, traversal, separators, or control characters. The manifest retains the exact case-sensitive profile ID, so filename normalization never becomes the authority for profile identity.
comparison.json is the authoritative index for the bundle. It uses an
explicit schema version and records safe comparison information including:
- comparison identity and start and finish timestamps;
- report ID, resolved valid period, and effective timezone;
- prompt ID, version, and inspected prompt hash;
- the relative data-package filename and SHA-256 digest;
- total, succeeded, and failed profile counts; and
- one ordered result per requested profile containing the exact profile ID, resolved backend and model, relative report filename when present, execution and validation status, and safe error information when failed.
The manifest and normal command summary must not contain credentials, provider request bodies, raw model output, rendered prompts, schemas, provider endpoints, or other content-rich diagnostics. The explicit data package and generated reports contain the development material the user requested and must be handled as operator-owned potentially sensitive output.
Failure And Publication Policy
Failure before concurrent execution, including invalid profiles, missing credentials, collection failure, preparation failure, or unsafe destination, publishes no comparison bundle and performs no model calls where the failure is discoverable during preflight.
After execution begins, Weatherreporter waits for every non-cancelled profile.
If one or more profiles fail, it still publishes a coherent partial bundle
containing data-package.yml, every successfully rendered report, and a
manifest describing all successes and failures. It then returns a non-zero
exit status. A failed profile has no report file unless a future contract
explicitly introduces a separately named diagnostic artifact.
Bundle contents are staged outside the destination and published only after
the manifest is complete. Ordinary publication accepts only an absent or empty
target directory. A nonempty existing directory fails without modification
unless --replace is present.
--replace may replace only the exact resolved target and must reject broad or
unsafe targets such as a filesystem root, the present working directory, a
symlink, or an unrecognized nonempty directory. A recognized prior bundle must
contain a valid Weatherreporter comparison manifest. Replacement publishes the
new complete or coherent partial bundle as a unit, prevents stale reports from
the prior comparison from surviving, and preserves or restores the prior
bundle if the final replacement operation fails.
An interrupted or cancelled comparison does not replace an existing bundle. Temporary staging artifacts are cleaned up on ordinary failure and cancellation without scanning or modifying unrelated directories.
Prompt Debugging
The existing --llm-debug-dir mechanism remains available. Concurrent
comparison executions require distinct, deterministic debug identities that
include the comparison and exact profile selection so callbacks cannot collide
or overwrite another profile's artifacts.
Debug writing must be concurrency-safe and retain the existing permission, redaction, explicit-opt-in, and path-containment guarantees. Debug artifacts remain separate from the comparison bundle; the bundle does not implicitly enable full Promptkit diagnostics.
Notification Policy
Profile comparisons never invoke Distributor notification, even when notification is enabled in the effective configuration. Comparison reports are local development artifacts rather than ordinary report publications.
Adding comparison publication or upload behavior would require a separate accepted feature scope and explicit operator authorization.
Architectural End State
Application orchestration exposes a reusable prepared-report boundary that contains the resolved report, shared collected and derived facts, module snapshot, briefing metadata, generated-text handler, render inputs, and exact serialized data package. That boundary is immutable during concurrent profile execution.
Ordinary generate behavior continues to prepare once and execute once.
compare prepares once and executes many without duplicating the generation
workflow or calling GenerateDetailed in a loop. Shared preparation,
profile-specific Promptkit execution, structured-output validation, rendering,
and artifact publication remain distinct responsibilities.
The Promptkit adapter remains the only owner of dependency-specific types and engine calls. The CLI owns parsing and user-facing summaries. The configuration package owns configuration. Application orchestration owns comparison order, concurrency lifecycle, failure aggregation, and bundle publication. Domain, prompt-input, generated-text, and template packages retain their existing deterministic contracts.
Scope
The completed feature includes:
- the
compareCLI command for every implemented generated-text report; - repeatable explicit profile selection and validation;
- configured and CLI output-directory integration through the implemented destination policy;
- one-time report resolution, collection, deterministic preparation, and YAML serialization;
- concurrent execution through one Promptkit executor with backend capacity respected;
- independent validation and rendering with deterministic ordered results;
- the flat, versioned comparison-bundle contract;
- safe filename derivation and data-package hashing;
- coherent partial-result publication and non-zero failure behavior;
- guarded whole-bundle replacement through
--replace; - comparison-aware, concurrency-safe optional prompt debugging;
- explicit suppression of Distributor notification;
- structured normal and quiet-mode CLI behavior consistent with existing commands;
- focused race-safe tests across configuration, CLI, application, Promptkit-adapter, rendering, and filesystem boundaries; and
- updates to every affected canonical user, operator, architecture, integration, and internal document.
Compatibility
The feature is additive. Existing generate and run commands, report
definitions, profile defaults, configuration, output filenames, notification
behavior, and exit contracts remain unchanged.
The comparison manifest and bundle layout begin as versioned contracts. They do not become inputs accepted by Weatherreporter, and no backward-compatible replay or long-term archive guarantee is implied beyond identifying the schema used to interpret a produced bundle.
Testing Expectations
Tests should provide durable coverage for:
- report and date parsing consistent with
generate; - rejection of fewer than two profiles, blanks, and duplicates;
- inspection of the exact prompt and every profile before collection;
- no collection or model execution after a preflight failure;
- exactly one weather collection and one preparation for several profiles;
- byte-for-byte identical data-package input in every execution;
- observable concurrent execution through a concurrency-safe fake executor;
- respect for Promptkit-owned backend capacity in an assembled adapter test where that integration adds distinct confidence;
- deterministic filenames, manifest order, summaries, and errors under varied completion order;
- continuation and coherent partial publication after one profile fails;
- cancellation propagation, goroutine completion, and preservation of an existing destination;
- destination precedence and each derived default directory;
- safe filename handling for unusual valid profile IDs;
- absent, empty, occupied, symlinked, unsafe, recognized, and unrecognized replacement targets;
- removal of stale prior report files during authorized whole-bundle replacement;
- exact package digest and manifest/result consistency;
- concurrency-safe, non-colliding opt-in debug artifacts; and
- absence of Distributor calls for complete and partial comparisons.
Concurrency and replacement behavior require race-enabled and consequential failure-path coverage. Tests must remain deterministic, offline, credential free, and independent of real Promptkit providers or machine-specific paths.
Documentation End State
Once implemented, the CLI reference owns command syntax, flags, summary, and exit behavior. The operations guide owns the bundle lifecycle, replacement procedure, sensitivity guidance, and practical prompt-comparison workflow. The architecture policy owns the statelessness, concurrency, notification, and publication invariants.
The Promptkit integration guide should describe
the consumer-visible multi-profile execution boundary without duplicating
Promptkit's backend-capacity reference. The versioned manifest and flat bundle
format belong in a focused contract under docs/integrations/. App
orchestration, prompt input, generated text, prompt debugging, and bundle
publication mechanics belong in focused documents under docs/internal/.
Current-state documentation must not describe profile comparison as available until the implementation lands.
Non-Goals
This roadmap does not introduce:
- automatic model scoring, ranking, recommendation, or winner selection;
- semantic or textual diff generation between reports;
- repeated sampling of one profile or statistical evaluation;
- prompt or profile editing through Weatherreporter;
- replaying a saved data package as command input;
- comparing several report types in one command;
- Weatherreporter-owned backend concurrency or queue configuration;
- automatic retries beyond Promptkit's existing execution contract;
- Distributor upload or other external publication;
- comparison history, indexing, retention, cleanup schedules, or implicit discovery of prior bundles; or
- changes to ordinary report content or normal generation behavior.
Any later automated evaluation, replay, sampling, or publication feature requires a separate accepted roadmap.
Completion Criteria
The feature is complete when a maintainer can select several Promptkit profiles, have them execute concurrently against one exact prepared report package, and receive a safe, flat, deterministic comparison bundle whose manifest accurately describes every success and failure. Configured and explicit destinations must follow the accepted output policy, replacement must never mix or silently destroy unrelated contents, cancellation and partial failure must be race-safe, ordinary notification must remain disabled, and all affected canonical documentation must describe the implemented behavior.
Open Questions
None. The scope, prerequisites, user intent, and target behavior are defined above.