26 KiB
PromptKit v0.5 Implementation Plan
Objective
Implement the target state in PromptKit v0.5 Integration And LLM Profile Policy. Each numbered stage is intended to be one implementation prompt for a GPT-5.6-Terra coding agent. Complete stages in order and leave the repository buildable, tested, and internally coherent after every stage.
Follow Architecture,
Testing Policy, and
Documentation Policy throughout. Preserve
unrelated user changes. Use apply_patch for source and documentation edits,
run gofmt on changed Go files, and add only tests that protect the behaviors
and risks assigned to that stage.
Do not implement the separate deterministic session-ID or default-concurrency roadmap items as part of this plan. Do not perform paid or credentialed LLM calls.
Background Summary
Notarius currently pins PromptKit v0.3.0, calls Prepare and then Run for one
completion, validates profiles through a synthetic prompt, has no application
fallback profile source, and accepts LLM profiles only at individual bindings
or through the run-wide CLI override. PromptKit v0.5.0 is source-compatible
with the current tree; a temporary v0.5.0 module override has already passed
go test ./....
The implementation must nevertheless treat the upstream optional-parameter
change as intentional: unset temperature, max_tokens, and top_p remain
unset and are omitted from compatible provider requests. Do not restore the old
implicit top_p: 1 default.
Stage 1: Upgrade The PromptKit Dependency
Goal
Establish a clean PromptKit v0.5.0 baseline before adopting its new APIs.
Work
- Update
go.modandgo.sumfrom PromptKit v0.3.0 to v0.5.0 and rungo mod tidy. - Change the PromptKit built-in profile-catalog marker in
internal/framework/llm/promptkit_profile_fingerprint.goto identify v0.5.0. This deliberately invalidates LLM checkpoints tied to the prior catalog identity. - Review PromptKit-facing compile errors or test failures against the v0.4.0 and v0.5.0 release guides. Do not adopt prepared execution, inspection, or fallback profiles in this stage.
- Replace the existing test assertion for one exact built-in fingerprint hash with durable assertions that the fingerprint is deterministic, non-empty, non-secret, and changes when a semantic profile source changes. Do not add a new version-constant or exact-hash change detector.
- Update
docs/integrations/pkg-promptkit.mdto pin and link v0.5.0 and state the implemented dependency-level behavior: unset optional sampling controls are provider defaults. Do not document later stages as implemented. - Update any other canonical text that explicitly claims the dependency is v0.3.0, but defer descriptions of unimplemented v0.5 APIs.
Tests And Validation
go test ./internal/framework/llm ./internal/cligo test ./...go vet ./...go build ./cmd/notariusrg -n 'promptkit v0\.3\.0|promptkit@v0\.3\.0|PromptKit v0\.3\.0' .git diff --check
Completion Criteria
- The repository directly pins v0.5.0 and all default offline checks pass.
- The profile-source fingerprint identifies the new upstream catalog without a brittle literal-hash test.
- Current documentation no longer identifies v0.3.0 as the supported version.
Stage 2: Execute One Frozen Prepared Snapshot
Goal
Make Notarius debug details and generation use one exact PromptKit preparation.
Work
- Refactor
PromptKitClient.CompleteStructuredto callPrepareExecution, immediately deferDiscard, obtain a caller-ownedDetailsvalue, and execute withRunPrepared. - Preserve the existing Notarius request mapping, cancellation precedence, validation classification, raw structured bytes, response decoding, profile recording, usage reporting, and credential redaction.
- Ensure every preparation, execution, validation, empty-result, and decode error retains useful Notarius prompt context without exposing prepared handle state or secrets.
- Use
errors.Asto obtain*promptkit.CapacityErroron admission rejection. Preservecontracts.ErrLLMCapacityExceededas the stable classification and add a nonblank backend ID only to safe application-owned diagnostic context. Do not exposepromptkit.CapacityErroroutside the LLM adapter. - Update
docs/internal/llm.mdand the implemented-mechanics portion ofdocs/integrations/pkg-promptkit.mdto describe the single frozen execution snapshot and structured capacity adaptation.
Tests And Validation
- Adapt existing PromptKit client tests to the prepared-execution path.
- Retain or add one behavioral test proving that the debug prompt details match the request actually passed to generation when a backing prompt source could otherwise change between independent preparations. Test the resulting snapshot consistency, not a private helper call count.
- Retain capacity tests proving
errors.Isreachescontracts.ErrLLMCapacityExceeded, the selected backend can appear in safe diagnostic context, and provider calls are not made after rejected admission. - Run
go test ./internal/framework/llmandgo test -race ./internal/framework/llm. - Run
go test ./...andgit diff --check.
Completion Criteria
CompleteStructuredno longer calls independentPrepareandRunoperations for one request.- Debug prompt material and generation result originate from the same frozen PromptKit snapshot.
- Capacity remains a provider-neutral Notarius error classification.
Stage 3: Replace Synthetic Profile Validation With Inspection
Goal
Validate profiles through PromptKit's exact profile-inspection boundary and centralize engine profile-source construction.
Work
- Introduce a small provider-adapter-owned profile inspection or validation
function in
internal/framework/llm. Its public internal signature must use Notarius-owned configuration and result/error types rather than returning PromptKit types to the CLI. - Share the code that applies
profile_dir,profile_file, and registered backend options between the production PromptKit engine and the inspection engine. Preserve the mutual-exclusion and local-backend rules. - Change CLI explicit-profile preflight to use
Engine.InspectProfilethrough that LLM boundary. - Remove
profileCheckPromptID,profileCheckPromptFS, thetesting/fstestproduction dependency, and the syntheticPreparerequest. - Preserve distinct, useful errors for an absent profile, invalid profile, unknown backend registration, cancellation, and invalid profile source.
- Do not require
api_key_envto be populated during configuration validation. Inspection may report credential requirements internally, but actual preparation remains responsible for credential availability before a model call. - Update current-behavior sections in
docs/internal/cli.mdanddocs/internal/llm.md. Keep field definitions indocs/config.md.
Tests And Validation
- Replace synthetic-prompt tests with profile inspection tests covering: configured local backend success; missing local backend failure; absent profile; malformed profile; and an otherwise valid profile whose credential environment variable is intentionally unset.
- Prove validation performs no provider HTTP call and remains offline.
- Run
go test ./internal/framework/llm ./internal/cliandgo test ./.... - Run
git diff --check.
Completion Criteria
- No production synthetic profile-check prompt remains.
- Profile validation uses the same ordinary profile source and backend registrations as execution.
- Configuration validation succeeds for structurally valid profiles without reading credential values.
Stage 4: Add Application Fallback Profile Asset Plumbing
Goal
Allow module families to register application-owned fallback profile YAML without placing domain policy in generic LLM code.
Work
- Extend
internal/framework/llm.AssetRegistrywith a separate fallback profile source collection, registration method, flattened filesystem, and safe content digest. - Reuse the existing asset-source path validation and flattening behavior where appropriate. Reject invalid roots, unreadable assets, and duplicate flattened paths. Do not parse PromptKit profile YAML in Notarius.
- Add
promptkit.WithFallbackProfileFSto production engine options only when at least one fallback profile source is registered. - Supply the identical assembled fallback source to the profile-inspection engine. Adjust CLI composition so pipeline-aware profile validation can use the production LLM asset registry without exposing PromptKit types.
- Extend profile-source checkpoint identity to include the exact fallback profile asset digest in addition to the PromptKit catalog marker and operator source. Keep the resulting fingerprint hash-only and path/content/credential free.
- Keep operator source precedence owned by PromptKit. Do not implement profile merging or duplicate PromptKit source resolution in Notarius.
- Update
docs/internal/llm.mdonly for the new implemented generic asset and fingerprint mechanics. No domain fallback exists until Stage 5.
Tests And Validation
- Add focused AssetRegistry tests for successful flattening, invalid roots, duplicate paths, and hash changes when fallback bytes change.
- Add adapter-level tests showing that the fallback filesystem reaches both execution construction and inspection construction.
- Extend checkpoint tests to prove fallback content changes profile-source identity without exposing raw YAML or paths. Use relational comparisons, not a fixed hash literal.
- Run
go test ./internal/framework/llm ./internal/cliandgo test ./.... - Run
git diff --check.
Completion Criteria
- Generic plumbing can carry application fallback profiles while remaining unaware of D&D IDs or model settings.
- Inspection, execution, and checkpoint identity use the same fallback asset source.
Stage 5: Adopt The D&D dnd-extraction Fallback
Goal
Give the D&D module family one stable embedded workload profile that operators can replace.
Work
- Add a D&D-owned embedded PromptKit profile asset with ID
dnd-extractionunderinternal/modules/dnd. Use the exact baseline defined inpromptkit.md: OpenRouter,openai/gpt-5.6-luna, no explicit reasoning effort, a 240-second timeout, flex service tier, and no selected temperature, token limit, ortop_p. The omitted reasoning value intentionally allows OpenAI's backend to apply itsmediumdefault. - Register the profile filesystem from the D&D registrar through the generic
fallback profile asset boundary. Keep D&D policy out of
internal/framework/llmand the CLI composition root. - Change every maintained D&D LLM prompt definition—including scene chunking,
all D&D extractors, and NPC normalization—from the model-named default to
default_profile: dnd-extraction. - Add an integration-level profile-resolution test proving that:
- the fallback resolves when no operator source defines the ID;
- a valid operator profile with the same ID wins completely; and
- an invalid matching operator profile fails rather than falling through.
- Test through Notarius's assembled production assets and PromptKit boundary; do not duplicate every upstream source-precedence case.
- Update the implemented profile ownership and prompt-default behavior in
docs/internal/dnd.md,docs/internal/llm.md, anddocs/integrations/pkg-promptkit.md. Defer the complete operator walkthrough and examples to Stage 10.
Tests And Validation
- Run focused D&D prompt preparation tests and the production composition tests.
- Run
go test ./internal/modules/dnd/... ./internal/framework/llm ./internal/cli. - Run
go test ./.... - Verify
rg -n 'default_profile: gemini-2-flash' internal/modules/dndreturns no matches. - Run
git diff --check.
Completion Criteria
- All maintained D&D prompts use the application-owned logical profile ID.
- The fallback works without an operator profile and remains authoritatively overridable by a matching valid operator definition.
Stage 6: Introduce Module Execution-Class Metadata
Goal
Make each production module's ability to use an LLM statically discoverable without yet changing profile inheritance.
Work
- Add
ExecutionClass contracts.ExecutionClasstopipeline.ModuleSpecand preserve it through normalization, cloning, catalogs, registries, JSON/debug views, and lookup helpers. - In this transitional stage only, allow an omitted execution class to normalize to deterministic so existing test-only fixtures can be migrated in Stage 7 without breaking the repository midway.
- Explicitly classify every production module:
- D&D scene chunking, every D&D extractor, and D&D NPC normalization as
llm_backed; - all other current production input, chunk, merge, normalize, and output
modules as
deterministic.
- D&D scene chunking, every D&D extractor, and D&D NPC normalization as
- Update production module specification tests and production catalog tests to assert the semantic class alongside stage, artifact kind, and capabilities.
- Add catalog lookup support needed by later resolution to retrieve a selected module's execution class by stage and key without constructing it.
- Do not implement pipeline-level profile inheritance or reject deterministic profiles yet.
- Update
docs/internal/modules.mdanddocs/internal/dnd.mdto identify execution class as registered module metadata, while noting only implemented uses.
Tests And Validation
- Run module registration/spec tests across generic, Seriatim, and D&D families.
- Run
go test ./internal/framework/pipeline ./internal/modules/.... - Run
go test ./...andgit diff --check.
Completion Criteria
- Every production module has an explicit correct execution class.
- Catalog consumers can retrieve that class without a concrete module instance.
- Test-only omitted classes remain the only temporary compatibility behavior.
Stage 7: Enforce Execution Metadata And Remove Runtime Probing
Goal
Finish the execution-class contract so missing metadata cannot cause future profile drift.
Work
- Update every framework, CLI, and integration test module specification to declare an explicit execution class appropriate to the fake behavior.
- Change module-spec validation so an empty or unsupported execution class is a registration error. Remove the transitional deterministic default from Stage 6.
- Replace the chunk runner's special
ChunkExecutionClassProviderprobe with specification-derived behavior. Remove the now-redundant provider interface, implementation methods, and tests when they have no remaining consumer. - Ensure chunk producer provenance remains unchanged: it records a non-empty effective binding profile for an LLM-backed chunker, while a deterministic chunker records no profile. A profile selected only through the prompt default remains represented by PromptKit's actual-profile manifest rather than being invented as an explicit chunk binding.
- Review helper constructors and fixtures for opportunities to set execution class once without obscuring the class under test. Do not introduce an elaborate test-spec framework.
- Update internal documentation if the removal changes any described runtime mechanics.
Tests And Validation
- Add or retain focused registration tests for missing and invalid execution classes.
- Retain chunk-plan provenance tests for LLM-backed and deterministic chunkers.
- Run
go test ./internal/framework/pipeline ./internal/modules/.... - Run
go test ./...,go vet ./..., andgit diff --check.
Completion Criteria
- No registered module specification relies on an implicit execution class.
- Pipeline metadata, not a concrete runtime type assertion, owns module execution classification.
Stage 8: Resolve Programmatic Pipeline Profile Defaults
Goal
Implement profile inheritance and precedence inside the pipeline resolver before exposing the field through YAML configuration.
Work
- Add an optional trimmed
LLMProfilefield topipeline.PipelineProfile. Add a non-empty runtime override field topipeline.ResolveOptionsso all precedence decisions occur in the resolver rather than through pre-resolution mutation. - After module selection,
--onlyfiltering, default validator-chain selection, and validator compatibility resolution, apply effective profiles to every selected input, chunk, extract, merge, normalize, output, and validator binding according to the precedence inpromptkit.md. - Apply profiles only when the selected module or validator execution class is
llm_backed. - Reject a binding-specific
llm_profileon any deterministic module or validator. Do not reject or inspect an unused pipeline default when no selected LLM-backed binding consumes it. - Leave an LLM-backed binding empty when no CLI, binding, or pipeline profile is
selected so PromptKit can use the prompt's
default_profile. - Store the effective values on resolved bindings before digest construction. Do not add a second inheritance decision to execution.
- Ensure semantically equivalent repeated binding profiles and one inherited default produce the same resolved pipeline digest. Ensure any changed effective profile changes the digest.
- Do not modify file configuration or CLI parsing in this stage.
Tests And Validation
- Add pipeline package tests for the complete precedence matrix: runtime override; binding-specific exception; pipeline default; prompt fallback; and deterministic bindings.
- Cover default and explicitly configured validator chains, all relevant stage
categories,
--onlylane selection, unused defaults, deterministic-profile rejection, and semantic digest equivalence. - Prefer table-driven package-level tests over assertions on private traversal helpers.
- Run
go test ./internal/framework/pipelineandgo test ./.... - Run
git diff --check.
Completion Criteria
- Programmatic pipelines resolve one canonical effective profile policy.
- Only LLM-backed resolved bindings can contain a profile.
- Runtime override, binding, pipeline, and prompt precedence is unambiguous and digest-stable.
Stage 9: Expose Pipeline Defaults Through Configuration And CLI
Goal
Make the profile-default workflow available to operators while preserving validation and override behavior.
Work
- Add optional
pipelines.<id>.llm_profilesupport to the version 4 file configuration model. Use presence-aware decoding so an explicitly set blank value is rejected, while omission remains valid. - Preserve the field through file application, configuration cloning, effective configuration, and programmatic profile copies without aliasing or trimming drift.
- Remove
applyLLMProfileOverride. Pass the CLI override through the resolver's runtime-override input so deterministic bindings are never populated. - Update effective profile-ID collection to cover every selected LLM-backed module stage and LLM-backed validator, including future LLM-backed input and output modules. Do not inspect deterministic or unselected profiles.
- Ensure
run,config validate --pipeline, resume/checkpoint identity, and relevant dry preflight paths all use the same resolved effective profiles. - Preserve
--llm-profileas the highest-precedence non-empty run-wide override and preserve binding-specific profiles as exceptions when no CLI override is present. - Do not increment the configuration version.
- Update current configuration and CLI contracts in
docs/config.mdanddocs/cli.mdin the same stage. Link to operations for the deployment workflow rather than duplicating it prematurely.
Tests And Validation
- Add file-config tests for omission, trimming, explicit blank rejection, unknown-key behavior, cloning, and round-trip application.
- Add effective-config and CLI contract tests for precedence, LLM-only
application, inherited-profile inspection failure before factory execution,
--only, and digest changes. - Retain offline operation and do not require credentials for
config validate --pipeline. - Run
go test ./internal/core/config ./internal/framework/pipeline ./internal/cli. - Run
go test ./...,go vet ./..., andgit diff --check.
Completion Criteria
- Operators can select
dnd-extractiononce per pipeline. - Configuration and CLI paths share the resolver's precedence policy.
- Unknown effective profiles fail preflight, while deterministic and unused profiles do not cause spurious inspection.
Stage 10: Complete Operator Documentation, Examples, And Decision Record
Goal
Make the implemented workflow understandable, copyable, and maintainable without duplicating canonical facts.
Work
- Create an ADR using the next sequential number for the durable decision to use workload-oriented pipeline defaults with operator-overridable application fallback profiles. Record context, decision, alternatives, and consequences; do not turn the ADR into a field reference or implementation log.
- Complete
docs/config.mdas the canonical owner of profile-source fields,pipelines.<id>.llm_profile, validation, and precedence. - Complete
docs/operations.mdwith an operator workflow that distinguishes Notarius embedded prompts, Notarius fallback profiles, PromptKit built-ins, and deployment filesystem profiles. Include production/development/local use of the samednd-extractionID, credential handling, absolute-path guidance, and the fact that current relative profile paths use the process working directory rather than the configuration file's directory. - Complete
docs/integrations/pkg-promptkit.mdwith the v0.5.0 boundary, prepared execution, inspection, fallback and ordinary source precedence, optional provider controls, capacity adaptation, and compatibility policy. - Update
docs/internal/configuration.md,docs/internal/pipeline.md,docs/internal/cli.md,docs/internal/llm.md,docs/internal/modules.md, anddocs/internal/dnd.mdonly for their owned implementation details. Link to canonical configuration, operations, and upstream format contracts rather than restating them. - Keep exactly the existing two D&D configuration examples. Add
llm_profile: dnd-extractionto the minimal and complete pipelines and remove the now-redundant model-named binding override from the complete example. - Add one secret-free maintained operator profile at
examples/profiles/dnd-extraction.yml. It should be a complete valid profile for the same logical ID and may mirror the embedded baseline; its purpose is to demonstrate file ownership and format, not claim automatic environment detection. Link it from the configuration and operations documentation. - If the complete example selects the external profile file, use a path that is valid for the documented repository-root invocation and explicitly note the working-directory rule. Keep the minimal example dependent only on the embedded fallback.
- Add or extend maintained-example validation so both configuration examples and the profile YAML are checked without generation or credentials.
- Remove the now-implemented
Pipeline-Level LLM Profile Defaultssection fromdocs/roadmap/future.md. Preserve the unrelated deterministic session and concurrency items. - Do not delete
promptkit.mdor this implementation plan during the feature implementation; retire them only after post-implementation review.
Tests And Validation
- Run maintained example/configuration tests and relevant CLI help/parser tests.
- Run
go test ./.... - Run
rg -n 'gemini-2-flash' examples docsand review every remaining match for intentional model-policy or historical context. - Run
rg -n 'v0\.3\.0|profileCheckPrompt|applyLLMProfileOverride' .and resolve stale production or current-documentation matches. - Verify all new links and
git diff --check.
Completion Criteria
- Every current fact has one canonical documentation owner.
- Operators can distinguish and deploy all profile layers without reading Go source.
- Both maintained configurations and the maintained external profile are valid, secret-free, and tested offline.
- Implemented profile work no longer remains in
future.md.
Stage 11: Final Verification And Quality Review
Goal
Verify the complete migration as one integrated change and correct only defects or omissions found during that review.
Work
- Review the final diff against every acceptance criterion in
promptkit.md. - Confirm provider-specific PromptKit types remain inside the LLM integration boundary and D&D policy remains inside the D&D module family.
- Confirm execution and inspection receive identical ordinary, fallback, and backend configuration.
- Confirm no paths, profile YAML, endpoints, credentials, or prepared handle state leak into fingerprints or ordinary diagnostics.
- Confirm all production module specs have explicit correct execution classes and every resolved deterministic binding is profile-free.
- Confirm prompt default, pipeline default, binding override, and CLI override behavior through representative assembled configurations.
- Review tests for redundancy and remove obsolete synthetic-prompt, runtime-probe, exact-hash, or duplicated upstream-behavior tests superseded by stronger contract tests.
- Perform an optional manual D&D quality comparison if credentials and an evaluation transcript are deliberately supplied. Record no private input or credential material, and do not make this comparison a completion gate.
Validation Commands
gofmt -w <changed-go-files>
go test ./...
go test -race ./internal/framework/llm ./internal/core/config ./internal/framework/pipeline ./internal/cli
go vet ./...
go build ./cmd/notarius
git diff --check
Also run focused stale-contract searches:
rg -n 'gitea.maximumdirect.net/eric/promptkit v0\.3\.0|PromptKit v0\.3\.0' .
rg -n 'default_profile: gemini-2-flash|profileCheckPrompt|applyLLMProfileOverride' internal docs examples
Review any matches rather than deleting intentional historical references blindly.
Completion Criteria
- All automated checks pass offline and without real credentials.
- The implemented behavior matches
promptkit.mdwith no known architecture, provenance, checkpoint, profile-precedence, or documentation gap. - Any optional live evaluation is clearly separate from correctness testing.
Open Questions
None. The roadmap decisions are sufficient to implement every stage without an additional product or architecture choice.