Update future roadmap document with ideas for new feature developments

This commit is contained in:
2026-07-30 16:27:36 +00:00
parent 9171b66a41
commit b5aaeb1c78

View File

@@ -53,6 +53,120 @@ not as committed release dates.
spell, combat, interaction, and scene-description lanes after real-world use.
Add more complex chunking only in response to demonstrated failures.
## Cross-Cutting LLM Runtime
### Deterministic Prompt Session Identity
- Replace the source-document-ID default for prompt sessions with one
predictable, procedurally generated session ID for the complete
source-processing workload.
- Preserve an explicit non-empty `--session-id` as the highest-precedence
override. Otherwise, derive the default only from the effective input module
identity and the exact raw input bytes.
- Use a versioned, bounded representation such as
`notarius:v1:<sha256(input-module + NUL + raw-input)>`. The exact encoding
must fit PromptKit's session length contract and must not embed source
content.
- Keep the derived session stable across runs, pipelines, selected lanes,
ordered steps, retries, resume, recomputation, LLM profiles, reasoning
overrides, and output, debug, or cache settings.
- Do not include file-backed references, generated references, reference
contents, or the composition of a reference bundle in session derivation.
References may change between prompt calls within one pipeline without
changing routing affinity.
- Resolve the authoritative session before checkpoint construction and use the
same value for checkpoint runtime identity, every prompt-facing module,
PromptKit's direct session field, the compatibility `session_id` prompt
variable, run-manifest metadata, and debug metadata.
- Keep routing identity separate from cache and checkpoint content identity.
Exact prompt prefixes, reference contents, model settings, and other
generation-affecting inputs must continue to participate in their existing
hashes and checkpoint fingerprints even though they do not change the
session.
- Treat the generated value as a provider-visible, stable pseudonymous
correlation identifier. Do not introduce an installation-specific HMAC or
secret unless a concrete multi-tenant or privacy requirement justifies
sacrificing deterministic identity across installations.
### Pipeline-Level LLM Profile Defaults
- Add an optional pipeline-level `llm_profile` default so an operator can
select one PromptKit execution policy for the pipeline without repeating the
same profile ID on every LLM-backed module binding.
- Apply the following precedence consistently: an explicit run-wide
`--llm-profile` override, then a binding-specific `llm_profile`, then the
pipeline-level default, then the prompt definition's embedded
`default_profile`.
- Apply inheritance only to bindings whose resolved modules are LLM-backed,
including applicable chunk, extraction, merge, normalization, and validation
bindings. Do not attach an inherited profile to deterministic modules or
weaken existing validation that rejects profiles where generation is not
supported.
- Resolve inherited profiles before effective-pipeline validation, digest and
checkpoint construction, execution, and provenance capture. Validate every
resulting explicit profile ID against the selected PromptKit profile source,
and ensure manifests and debug output report the profile actually selected
for each generation target.
- Preserve binding-specific profiles as intentional exceptions for modules
that require a different quality, latency, cost, provider, or reasoning
policy. Preserve `--llm-profile` as the convenient highest-precedence
experiment or incident-response override for an entire run.
- Treat PromptKit profiles as execution-policy configuration and prefer stable,
workload-oriented IDs such as `dnd-extraction` over model names. A pipeline
should express the kind of work it performs rather than encode a particular
provider, model, or deployment environment.
- Establish deployment-managed PromptKit profile files as the recommended
environment-specific configuration mechanism. Production, development, and
local deployments may each define the same logical `dnd-extraction` ID with
different model and generation settings, while retaining one unchanged
Notarius pipeline definition.
- Keep deployment profiles distinct from both Notarius prompt assets embedded
in the application binary and PromptKit's built-in profile catalog. Document
that `promptkit.profile_dir` or `promptkit.profile_file` selects an external
filesystem source whose definitions overlay PromptKit built-ins, and
recommend application-owned IDs rather than silently replacing built-in
profile IDs.
- Document an operator-friendly layout in which the Notarius configuration and
a profile subdirectory are deployed together. Until profile paths are
explicitly resolved relative to the configuration file, clearly state that
relative paths use the process working directory and recommend absolute
paths for services and containers.
- Update the configuration reference, operations guide, PromptKit integration
boundary, relevant internal configuration and pipeline documentation, and
maintained D&D examples together. Include one concrete external-profile
example and explain the profile precedence and environment-neutral pipeline
pattern without duplicating PromptKit's complete profile-format reference.
### Raise The Default Application-Wide LLM Limit
- Raise the default `concurrency.total_llm` value from 1 to 16 so ordinary
single-backend runs can use PromptKit's expected OpenRouter capacity and
lower-capacity local backends without an unnecessarily narrower Notarius
limit.
- Keep the Notarius application-wide scheduler mandatory and require
`total_llm` to remain a positive integer. Do not make the default unlimited:
endpoint-only profiles, an unrestricted local backend, injected clients, and
aggregate work across several backends may have no narrower PromptKit limit.
- Continue defaulting `concurrency.stage_workers.extract` to the effective
`total_llm`, making its default 16 as part of the same change. Preserve an
explicit lower extract-worker setting when an operator wants less queued or
concurrent extraction work.
- Define effective provider concurrency as the intersection of the Notarius
application-wide limit, the selected PromptKit backend limit when present,
and the work made available by stage execution. A Notarius limit of 16 does
not narrow a backend already limited to 16, while a local backend limited to
4 remains bounded at 4.
- Treat the default as an application-wide safety ceiling across profiles,
backends, modules, retries, and validators. A run that intentionally needs
the combined capacity of several backends may configure a higher
`total_llm` and an appropriate extract-worker count explicitly.
- Retain the existing configuration and environment override surfaces. Update
canonical configuration, operations, and internal documentation together
when the default changes.
- Reconsider decoupling the extract-worker default from `total_llm` only after
mixed-backend workloads demonstrate a need for a high global emergency
ceiling with a lower default work-production rate.
## Shared Normalization And Quality Work
### Generic LLM-Assisted Deduplication