Files
notarius/docs/roadmap/future.md

5.9 KiB

Future Work

Current Notarius behavior is documented in the canonical README, CLI, configuration, operations, internal, and integration docs. This roadmap records future work only. Items are ordered roughly by current value and specificity, not as committed release dates.

Near-Term D&D Pipeline

Evaluate Spell Extraction And Normalization

  • Evaluate ordinary extraction retries and the completed normalization path against a human-reviewed transcript set before adding repair-aware retries or an LLM-backed semantic validator.
  • Maintain a small set of human-reviewed transcripts and outputs for prompt, validator, and normalizer development. Treat model-quality review as an iterative human evaluation aid, not a deterministic correctness gate.

Add Sequential D&D Artifacts

  • The next proposed increment is D&D combat-turn extraction, using earlier NPC output as an explicit identity reference while preserving independent runs.
  • Add narrative extraction for scene summaries, party actions, and NPCs encountered when that output proves useful beyond the dedicated NPC artifact.
  • Define the preferred operational sequence for independent pipelines on the same transcript. The initial direction is NPCs first, followed by spells and combat turns as appropriate, with earlier JSON artifacts supplied to later runs as references.
  • Keep this sequencing operator- or script-driven initially. Do not require a general DAG or concurrent cross-lane reconciliation model.

Improve D&D Scene Classification

  • Extend scene annotations with classifications that downstream extractors can use, including reliable combat and narrative indicators.
  • Strengthen the scene prompt so every scene containing combat turns is marked as combat, and add validation capable of detecting missing or inconsistent combat classifications.
  • Allow the combat extractor to no-op for chunks that are not classified as combat, avoiding unnecessary model calls where practical.
  • Allow a narrative extractor to select the corresponding scene classification rather than processing every chunk indiscriminately.
  • Reassess whether one shared scene plan provides enough context for NPC, spell, combat, and narrative pipelines after these extractors have real-world usage. Add more complex chunking only in response to demonstrated failures.

Shared Normalization And Quality Work

Generic LLM-Assisted Deduplication

  • Add a reusable normalizer that asks an LLM to identify duplicate sets in a list and propose one replacement element for each set.
  • Define the minimum domain-neutral input contract, initially an ordered list whose elements have stable unique IDs. Artifact-kind registrations or adapters may expose that structure without moving domain rules into the generic package.
  • Keep mutation deterministic: parse and validate the model's duplicate groups, require every referenced ID to exist, reject overlapping or malformed groups, prevent unrelated insertion or deletion, and apply only approved replacement operations in code.
  • Preserve provenance needed for audit and downstream validation, and emit warnings describing every collapsed group.
  • Evaluate batching and context-window limits before applying the normalizer to large artifact collections.

The model may use its own domain knowledge to judge semantic duplication; the generic implementation is responsible only for the common proposal contract, safety checks, and deterministic application of accepted changes.

Validation And Review

  • Add domain validators and production default chains alongside each new D&D artifact.
  • Add production LLM-backed validators only when a concrete review policy benefits from model judgment and deterministic checks are insufficient.
  • Add validator diagnostics and timing summaries if operators need more detail than the current durable output bundle provides.
  • Add validator compatibility metadata if deployments need config-time proof that a validator is suitable for a particular stage, module, or artifact kind.
  • Add media-type validators when non-JSON artifact representations are introduced.

Reference And Sequential-Pipeline Evolution

  • Make prior-run artifacts easier to bind as references without changing the existing module-facing reference-item contract.
  • Add structured or parsed references, such as typed NPC registries, rosters, or spell catalogs, when opaque UTF-8 prompt material is no longer sufficient.
  • Add per-slot or per-chunk inclusion policies so large references are not repeated in every prompt unnecessarily.
  • Add token budgeting and model context-window management for reference content.
  • Add reference caching, preprocessing, summarization, embedding, or retrieval only when reference size and observed model behavior justify them.
  • Consider non-file reference producers for prior-run artifacts, derived summaries, or entity registries after manual sequential composition becomes burdensome.

Blue-Sky Platform And Operations

These ideas are intentionally less specified. Promote one into an earlier section only after a concrete workflow, contract, and priority emerge.

Platform Extensions

  • Additional input adapters, such as Markdown or note-export formats.
  • Additional output encoders.
  • Concurrent cross-lane entity normalization or broader workflow composition.
  • Batching or specialized context-window controls for LLM-backed validators.

Distribution And Operations

  • Packaged release artifacts for alpha distribution.
  • A documented versioning and release process.
  • Optional generated example-output fixtures with a regeneration procedure.
  • Additional diagnostics or reporting views.

Workspace And Storage

  • Default-idempotent run behavior with an explicit force override.
  • Remote workspace storage.
  • Workspace garbage collection and archival policies.
  • Cross-machine checkpoint reuse.