12 KiB
Future Work
Current Notarius behavior is documented in the canonical README, CLI, configuration, operations, internal, and integration docs. This roadmap records future work only. Items are ordered roughly by current value and specificity, not as committed release dates.
Near-Term: Ordered Pipeline Steps
Allow one configured pipeline to contain multiple ordered execution steps so accepted artifacts from an earlier step can become generated references for later steps in the same run. This is a bounded extension of the fixed pipeline model, not an arbitrary DAG or general workflow language.
Input parsing and chunk planning remain pipeline-wide. Each step selects one or more artifact lanes; every selected lane completes extraction, validation, merge, normalization, and validation before dependent later steps begin. Lanes within the same step remain independent and may execute concurrently. The runner exposes only accepted normalized artifacts across a step boundary; raw extracts, rejected outputs, and intermediate merge results cannot become downstream references.
Generated-reference bindings must be explicit in resolved configuration. A binding identifies an earlier producing lane and one declared reference slot on a later consuming lane. Resolution must reject missing producers, references to the same or a later step, incompatible artifact kinds or media types, undeclared consumer slots, cycles, and ambiguous bindings. A configured external reference and a generated reference cannot bind the same effective target slot; reject that pipeline or runtime override instead of applying a precedence rule. Each effective target slot accepts at most one producer. One generated artifact may fan out to multiple compatible target slots in a later step; aggregation from multiple producers into one slot is deferred until a concrete use case defines deterministic semantics.
Step-scoped reference defaults mirror existing pipeline-level reference defaults. A step binds an earlier artifact once, and the binding automatically applies to every target in that step that declares the named slot. Target-local bindings remain available when only one module should consume the artifact; do not combine a target-local and step-scoped binding for the same effective slot. A generated source uses a structured, unambiguous form rather than encoding a producer into a path-like string. The target configuration shape is:
steps:
- id: identify-npcs
artifacts:
npcs:
extract: dnd/npcs
normalize: dnd/npcs
- id: grounded-events
references:
npcs:
artifact:
step: identify-npcs
lane: npcs
artifacts:
spells:
extract: dnd/spells
normalize: dnd/spells
combat:
extract: dnd/combat-turns
normalize: dnd/combat-turns
This snippet shows the step and reference portion of a pipeline; pipeline-wide
input, chunk, output, and other references are omitted. Existing scalar
reference values continue to mean external file paths; the structured
artifact form means an accepted normalized artifact from the named earlier
step and lane. Do not infer generated bindings from module keys, matching lane
names, or D&D-specific knowledge in the framework.
The handoff should use the producer artifact's canonical codec representation and retain its artifact kind, schema identity, media type, content digest, and producer provenance. Generated references provide context or disambiguation, not source evidence. They use the existing module-facing reference contract where possible; typed or domain-specific adapters may validate and prepare a reference without moving domain concepts into the pipeline framework.
Pipeline identity, manifests, checkpoint dependencies, debug records, and errors must include step identity and generated-reference provenance. A downstream checkpoint is reusable only when the upstream artifact identity and content digest match. Resume should reconstruct an accepted upstream artifact through its registered codec instead of requiring the producing lane to run again when its checkpoint is reusable.
Dependency-aware resume and recomputation
Treat generated-reference bindings as checkpoint dependencies. Reusing an earlier step is safe only when its existing checkpoint and codec identity are valid. Reusing a dependent step additionally requires an exact match for every upstream artifact kind, schema identity, media type, and canonical content digest it consumed. A changed, missing, rejected, corrupt, or incompatible producer artifact invalidates all transitive dependent checkpoints; the runner must never combine a newly produced upstream artifact with stale downstream output.
Support selectively recomputing one configured step and all of its transitive dependents while retaining reusable independent and predecessor work. The operator-facing selection mechanism should identify a stable configured step, not individual internal stage operations. Resolution must reject a selection that would omit a required predecessor without a reusable accepted artifact.
Manifests, checkpoint events, and diagnostics should distinguish work that was executed, reused, or invalidated and record a bounded non-secret reason for dependency-driven invalidation. Completion order must not affect invalidation, public artifact ordering, or the set of dependent steps selected for rerun.
If a required producer finishes without an accepted normalized artifact, fail the entire run with a deterministic dependency error. Do not start any dependent step. Preserve the upstream rejection or empty-result outcome and step provenance in the failed run manifest so the cause remains auditable.
The first production workflow is D&D NPC grounding:
- the first step runs the NPC lane through accepted normalized output; and
- the second step runs spell and combat-turn lanes, binding that NPC artifact
to the spell extractor and to the combat extractor and normalizer through
their existing
npcsreference slots.
Spell and combat-turn extraction may run concurrently after the NPC handoff is available. The NPC artifact may disambiguate participant identity, but it does not prove that a spell cast or combat turn occurred.
Near-Term D&D Pipeline
Evaluate Spell Extraction And Normalization
- Evaluate ordinary extraction retries and the completed normalization path against a human-reviewed transcript set before adding repair-aware retries or an LLM-backed semantic validator.
- Maintain a small set of human-reviewed transcripts and outputs for prompt, validator, and normalizer development. Treat model-quality review as an iterative human evaluation aid, not a deterministic correctness gate.
Expand Sequential D&D Artifacts
- Add narrative extraction for scene summaries, party actions, and NPCs encountered when that output proves useful beyond the dedicated NPC artifact.
- Use ordered pipeline steps when a later artifact needs an accepted earlier artifact as context. Keep independent lanes in the same step and do not introduce a general DAG or concurrent cross-lane reconciliation model.
Improve D&D Scene Classification
- Extend scene annotations with classifications that downstream extractors can use, including reliable combat and narrative indicators.
- Strengthen the scene prompt so every scene containing combat turns is marked as combat, and add validation capable of detecting missing or inconsistent combat classifications.
- Allow the combat extractor to no-op for chunks that are not classified as combat, avoiding unnecessary model calls where practical.
- Allow a narrative extractor to select the corresponding scene classification rather than processing every chunk indiscriminately.
- Reassess whether one shared scene plan provides enough context for NPC, spell, combat, and narrative pipelines after these extractors have real-world usage. Add more complex chunking only in response to demonstrated failures.
Shared Normalization And Quality Work
Generic LLM-Assisted Deduplication
- Add a reusable normalizer that asks an LLM to identify duplicate sets in a list and propose one replacement element for each set.
- Define the minimum domain-neutral input contract, initially an ordered list whose elements have stable unique IDs. Artifact-kind registrations or adapters may expose that structure without moving domain rules into the generic package.
- Keep mutation deterministic: parse and validate the model's duplicate groups, require every referenced ID to exist, reject overlapping or malformed groups, prevent unrelated insertion or deletion, and apply only approved replacement operations in code.
- Preserve provenance needed for audit and downstream validation, and emit warnings describing every collapsed group.
- Evaluate batching and context-window limits before applying the normalizer to large artifact collections.
The model may use its own domain knowledge to judge semantic duplication; the generic implementation is responsible only for the common proposal contract, safety checks, and deterministic application of accepted changes.
Validation And Review
- Add domain validators and production default chains alongside each new D&D artifact.
- Add production LLM-backed validators only when a concrete review policy benefits from model judgment and deterministic checks are insufficient.
- Add validator diagnostics and timing summaries if operators need more detail than the current durable output bundle provides.
- Add validator compatibility metadata if deployments need config-time proof that a validator is suitable for a particular stage, module, or artifact kind.
- Add media-type validators when non-JSON artifact representations are introduced.
Further Reference Evolution
- Make prior-run artifacts easier to bind as references without changing the existing module-facing reference-item contract.
- Add structured or parsed references, such as typed NPC registries, rosters, or spell catalogs, when opaque UTF-8 prompt material is no longer sufficient.
- Add per-slot or per-chunk inclusion policies so large references are not repeated in every prompt unnecessarily.
- Add token budgeting and model context-window management for reference content.
- Add reference caching, preprocessing, summarization, embedding, or retrieval only when reference size and observed model behavior justify them.
- Extend generated references to prior-run artifacts or derived summaries only after same-run ordered handoffs establish the required provenance and lifecycle semantics.
Design Considerations To Revisit
These concerns are relevant to ordered artifact dependencies but are not committed near-term features.
Cross-artifact identity links
Evaluate whether downstream D&D artifacts should retain canonical NPC IDs from the generated NPC reference in addition to normalized display names. Any such contract must define player-character, unknown-actor, missing-NPC, and superseded-identity behavior before implementation. Deterministic validation may confirm that a linked ID exists in the consumed NPC artifact, but the link must never substitute for transcript evidence that the downstream event occurred.
Artifact contract evolution
Define compatibility and migration policy before generated-reference chains must span multiple schema versions or long-lived historical artifacts. The policy should address stable identifier semantics, which schema changes permit checkpoint reuse, when an older artifact may be decoded or adapted, and when a producer or all dependents must be recomputed. Do not add a general migration framework until an actual contract change requires one.
Blue-Sky Platform And Operations
These ideas are intentionally less specified. Promote one into an earlier section only after a concrete workflow, contract, and priority emerge.
Platform Extensions
- Additional input adapters, such as Markdown or note-export formats.
- Additional output encoders.
- Concurrent cross-lane entity normalization or broader workflow composition.
- Batching or specialized context-window controls for LLM-backed validators.
Distribution And Operations
- Packaged release artifacts for alpha distribution.
- A documented versioning and release process.
- Optional generated example-output fixtures with a regeneration procedure.
- Additional diagnostics or reporting views.
Workspace And Storage
- Default-idempotent run behavior with an explicit force override.
- Remote workspace storage.
- Workspace garbage collection and archival policies.
- Cross-machine checkpoint reuse.