Files
notarius/docs/roadmap/future.md

10 KiB

Future Work

Current Notarius behavior is documented in the canonical README, CLI, configuration, operations, internal, and integration docs. This roadmap records future work only. Items are ordered roughly by current value and specificity, not as committed release dates.

Near-Term D&D Pipeline

Evaluate Spell Extraction And Normalization

  • Evaluate ordinary extraction retries and the completed normalization path against a human-reviewed transcript set before adding repair-aware retries or an LLM-backed semantic validator.
  • Maintain a small set of human-reviewed transcripts and outputs for prompt, validator, and normalizer development. Treat model-quality review as an iterative human evaluation aid, not a deterministic correctness gate.

Extract NPC Interactions

  • Add an ordered NPC-interaction artifact that records how an identifiable NPC participates in the session without adding occurrence-level state to the normalized NPC registry. Use npc-interactions as the working lane and product name; the exact artifact kind may be finalized with its contract.
  • Keep each record minimal: canonical NPC name, one bounded interaction kind, and transcript source_refs supporting both the identity and classification.
  • Start with the mutually exclusive vocabulary mentioned, noncombat_presence, dialogue, combat_ally, combat_opponent, and other. Define narrow inclusion rules and category precedence before implementation so overlapping activity does not produce arbitrary labels.
  • Model interactions as ordered occurrences rather than one scalar NPC category. The same NPC may therefore have separate records when the transcript establishes distinct interactions, such as dialogue followed by hostile combat.
  • Run NPC identity extraction first and provide its accepted names-only projection to the interaction extractor for grounding. Registry names may disambiguate identity but never establish that an interaction occurred, and registry source references must not be copied into interaction evidence.
  • Evaluate category agreement, evidence sufficiency, duplicate behavior, and smaller-model reliability on human-reviewed transcripts before expanding the enum or adding additional fields.

Export Accepted Chunk Maps

  • Add an option to emit the accepted materialized chunk map as a proper, framework-owned artifact with a documented schema identity, version, media type, canonical encoding, and source/chunker provenance.
  • Export the exact ordered chunks used for lane execution, including stable chunk IDs and current-source ranges. Do not expose a model's raw boundary proposal or require a second model call to reconstruct information already owned by the framework.
  • Treat chunk-map export as an output concern rather than an ordinary extraction lane. Chunking is pipeline-wide and precedes lane extraction; a pseudo-extractor would duplicate work and obscure that ownership boundary.
  • Keep the generic contract independent of D&D interpretation. Namespaced annotations may be preserved when they are part of the accepted chunk plan, but downstream applications should not need domain-specific annotations to understand chunk identity, order, or source coverage.

Extract D&D Scene Descriptions

  • Add a dnd/scene-descriptions extractor that runs once for each accepted scene chunk and explicitly owns the small amount of scene synthesis useful to downstream applications.
  • Keep the private model response to exactly kind, title, and summary. Use the enum combat, narrative, recap, and meta: narrative means current-session in-world gameplay that is not combat, recap, or sustained out-of-character discussion.
  • Treat brief table talk or rules clarification as incidental to the enclosing gameplay scene. A sustained transition between kinds should normally create a chunk boundary; define a primary-kind rule for residual mixed chunks before implementation.
  • Deterministically attach the accepted chunk ID and its exact source range while mapping the private response into the durable artifact. Do not ask the model to reproduce IDs or segment boundaries, and do not defer required identity or evidence until normalization.
  • Keep normalization limited to stable ordering, exact deduplication, and canonical invariant enforcement. Titles and summaries are explicit, source-bounded synthesis owned by this artifact rather than by the chunker.
  • Do not add participants. Derive participant-oriented views by joining scene ranges with NPC evidence or, preferably, NPC-interaction occurrences. An NPC registry reference proves identity, not exhaustive presence in every scene.

Minimize And Use D&D Scene Chunking

  • Reduce the D&D scene chunker toward its narrow responsibility: identifying coherent scene boundaries. Retain boundary confidence or caveats only when a demonstrated validator or operator workflow consumes them; move title, summary, scene kind, and participant duties to dedicated artifacts.
  • Allow the combat extractor to no-op for chunks classified as non-combat only after the scene-description artifact can be supplied through an explicit ordered dependency. Do not make generic chunk materialization depend on a D&D classification.
  • Use ordered pipeline steps whenever a later artifact needs an accepted earlier artifact as context. Keep independent lanes in the same step and do not introduce a general DAG or concurrent cross-lane reconciliation model.
  • Reassess whether one shared scene plan provides enough context for NPC, spell, combat, interaction, and scene-description lanes after real-world use. Add more complex chunking only in response to demonstrated failures.

Shared Normalization And Quality Work

Generic LLM-Assisted Deduplication

  • Add a reusable normalizer that asks an LLM to identify duplicate sets in a list and propose one replacement element for each set.
  • Define the minimum domain-neutral input contract, initially an ordered list whose elements have stable unique IDs. Artifact-kind registrations or adapters may expose that structure without moving domain rules into the generic package.
  • Keep mutation deterministic: parse and validate the model's duplicate groups, require every referenced ID to exist, reject overlapping or malformed groups, prevent unrelated insertion or deletion, and apply only approved replacement operations in code.
  • Preserve provenance needed for audit and downstream validation, and emit warnings describing every collapsed group.
  • Evaluate batching and context-window limits before applying the normalizer to large artifact collections.

The model may use its own domain knowledge to judge semantic duplication; the generic implementation is responsible only for the common proposal contract, safety checks, and deterministic application of accepted changes.

Validation And Review

  • Add domain validators and production default chains alongside each new D&D artifact.
  • Add production LLM-backed validators only when a concrete review policy benefits from model judgment and deterministic checks are insufficient.
  • Add validator diagnostics and timing summaries if operators need more detail than the current durable output bundle provides.
  • Add validator compatibility metadata if deployments need config-time proof that a validator is suitable for a particular stage, module, or artifact kind.
  • Add media-type validators when non-JSON artifact representations are introduced.

Further Reference Evolution

  • Make prior-run artifacts easier to bind as references without changing the existing module-facing reference-item contract.
  • Add structured or parsed references, such as typed NPC registries, rosters, or spell catalogs, when opaque UTF-8 prompt material is no longer sufficient.
  • Add per-slot or per-chunk inclusion policies so large references are not repeated in every prompt unnecessarily.
  • Add token budgeting and model context-window management for reference content.
  • Add reference caching, preprocessing, summarization, embedding, or retrieval only when reference size and observed model behavior justify them.
  • Extend generated references to prior-run artifacts or derived summaries only after same-run ordered handoffs establish the required provenance and lifecycle semantics.

Design Considerations To Revisit

These concerns are relevant to ordered artifact dependencies but are not committed near-term features.

Evaluate whether downstream D&D artifacts should retain canonical NPC IDs from the generated NPC reference in addition to normalized display names. Any such contract must define player-character, unknown-actor, missing-NPC, and superseded-identity behavior before implementation. Deterministic validation may confirm that a linked ID exists in the consumed NPC artifact, but the link must never substitute for transcript evidence that the downstream event occurred.

Artifact contract evolution

Define compatibility and migration policy before generated-reference chains must span multiple schema versions or long-lived historical artifacts. The policy should address stable identifier semantics, which schema changes permit checkpoint reuse, when an older artifact may be decoded or adapted, and when a producer or all dependents must be recomputed. Do not add a general migration framework until an actual contract change requires one.

Blue-Sky Platform And Operations

These ideas are intentionally less specified. Promote one into an earlier section only after a concrete workflow, contract, and priority emerge.

Platform Extensions

  • Additional input adapters, such as Markdown or note-export formats.
  • Additional output encoders.
  • Concurrent cross-lane entity normalization or broader workflow composition.
  • Batching or specialized context-window controls for LLM-backed validators.

Distribution And Operations

  • Packaged release artifacts for alpha distribution.
  • A documented versioning and release process.
  • Optional generated example-output fixtures with a regeneration procedure.
  • Additional diagnostics or reporting views.

Workspace And Storage

  • Default-idempotent run behavior with an explicit force override.
  • Remote workspace storage.
  • Workspace garbage collection and archival policies.
  • Cross-machine checkpoint reuse.