Compare commits
521 Commits
v0.1.0
...
f120be1cb4
| Author | SHA1 | Date | |
|---|---|---|---|
| f120be1cb4 | |||
| 17673d74ea | |||
| ef19a03cbf | |||
| 0546f6eb4f | |||
| d28d1062e0 | |||
| b3ebfcef37 | |||
| b70d9f77e3 | |||
| 2a75f40871 | |||
| a705ba74a1 | |||
| 8d9c9e7c87 | |||
| 0b5cc4f251 | |||
| 82ffe85f2d | |||
| b3644abc0e | |||
| d653bf1b90 | |||
| 3e66127b94 | |||
| 5d086c13ca | |||
| d36d4e7689 | |||
| 1456aa51cc | |||
| ffc179c822 | |||
| 8e669a1f14 | |||
| 14bfae216d | |||
| 5a58d87995 | |||
| 557809f364 | |||
| ee600975f0 | |||
| 0d8017e23f | |||
| 37b18edf3d | |||
| cda7a61b47 | |||
| 2ad9283148 | |||
| 41a8a80dda | |||
| 90c7fa6381 | |||
| e3839f8620 | |||
| 0fc2f9ee01 | |||
| 5d6305f21a | |||
| 551e4daea2 | |||
| 3589d33468 | |||
| a22c1a7f59 | |||
| ad85d71b0f | |||
| f3506240c2 | |||
| 70c199aa31 | |||
| e2b82746ab | |||
| 4235507f7b | |||
| b346670cc7 | |||
| 7868c26be7 | |||
| 92e89076a2 | |||
| d9b87347b8 | |||
| 20397ef710 | |||
| fc449863f2 | |||
| 51d62de1f3 | |||
| fc76805075 | |||
| 8e680cf96e | |||
| ece1bca460 | |||
| 516af12916 | |||
| 1015d61b2d | |||
| 4192aa8584 | |||
| 8cf03a2a44 | |||
| 2ee6b495e1 | |||
| f1b120b590 | |||
| 2ec17f5b4f | |||
| 6de470d541 | |||
| a6e176e160 | |||
| 3dfefd0e14 | |||
| f91237e9c0 | |||
| ad9ee076b5 | |||
| ce04387dbc | |||
| e8965ebbbb | |||
| c006b163d5 | |||
| 2f61118e78 | |||
| 5e2ccffc0f | |||
| 3f4a1f2647 | |||
| 9653e06297 | |||
| 916a32195b | |||
| 299c110267 | |||
| 4093ff2e8b | |||
| 39563f3ea0 | |||
| 4b52cace76 | |||
| 16d1b29b14 | |||
| 856b26b718 | |||
| ab70347c5d | |||
| 8ed99ceaee | |||
| 257bca0dc2 | |||
| bad5db3ef3 | |||
| 406ad1d362 | |||
| 5f207b2ab1 | |||
| c613b306ae | |||
| 1c7291e17d | |||
| 2345da106a | |||
| 03761c97dd | |||
| 9cb9ee48a7 | |||
| 08954f17e2 | |||
| a57f83e30d | |||
| f4c05c34ef | |||
| 29fcad6e9b | |||
| f5fd115046 | |||
| 7f28899730 | |||
| 84c0758455 | |||
| 5002864e88 | |||
| 55b188fd84 | |||
| d1f43df88e | |||
| d52387c1f7 | |||
| 9c5e3cff14 | |||
| 811d5b8bd9 | |||
| a168c13b85 | |||
| dd61a4efda | |||
| 228cc6ee83 | |||
| 06170e1f65 | |||
| 7715baa1f6 | |||
| 98506db1a9 | |||
| bb4855f0c6 | |||
| c51934d5c6 | |||
| c3513da880 | |||
| c7d853ea52 | |||
| da7fcdaffd | |||
| b6aad4fa98 | |||
| 9c6af28d02 | |||
| 04eabdfcb9 | |||
| a8a99c1037 | |||
| e15007fffb | |||
| e6b7c61f45 | |||
| 42973215fa | |||
| ba1d112d1f | |||
| b722131d57 | |||
| a92d2c0885 | |||
| 02ec10d66b | |||
| 9dd57dbfa4 | |||
| c164a3fc69 | |||
| 8834df617f | |||
| db2adb52da | |||
| fc3c128171 | |||
| 8a15b083a0 | |||
| d9dae2b639 | |||
| 7a4fd7be7a | |||
| 39388e96d4 | |||
| 12ac25bd63 | |||
| 394278e1f2 | |||
| 5cd7f8e737 | |||
| bf3fadf9ae | |||
| 58815aaf33 | |||
| ce857966f1 | |||
| a3bd0c1867 | |||
| b05634ee86 | |||
| 4829f94157 | |||
| 67b315099d | |||
| b5c86de4d7 | |||
| 2eeca2ed5a | |||
| b5aaeb1c78 | |||
| 9171b66a41 | |||
| b4363b3b73 | |||
| 241e9d2a89 | |||
| 715fff7b72 | |||
| d627b91b4f | |||
| a67b3aa76d | |||
| a16dcdfa52 | |||
| 46e4466d28 | |||
| 71a004bfc8 | |||
| f8333f2c15 | |||
| f603f7ac64 | |||
| 7a00e7049c | |||
| 2a9db9a957 | |||
| de046a8f13 | |||
| f1a6574013 | |||
| 4bca6d3103 | |||
| 7c569a3d8c | |||
| 8e04ef9e2b | |||
| 53a330587b | |||
| 7cfab8ada0 | |||
| 5c82b62856 | |||
| de8ed41b34 | |||
| d1eaec4dad | |||
| c0ec068f53 | |||
| 5cbd9e56e4 | |||
| 53490cdb59 | |||
| 7c94b5eeed | |||
| 4f2864fc96 | |||
| 893b03fccf | |||
| 256cc98ddb | |||
| e61e522662 | |||
| a4c7eca87b | |||
| 224a8292c4 | |||
| 64d461fc18 | |||
| fb1134e591 | |||
| 1da29e6788 | |||
| 0d947549fb | |||
| 950fba17ce | |||
| 678d2c6099 | |||
| db8db5ffc5 | |||
| 94b3eafb1a | |||
| f6981e2264 | |||
| 2f506f4985 | |||
| fdf8c4afd4 | |||
| 74c793e6a1 | |||
| b5835fbc37 | |||
| fd3f7b85cc | |||
| d86b74f485 | |||
| 59cbf1eb27 | |||
| ee43add75c | |||
| 46761706a2 | |||
| 5a968b64eb | |||
| 7b077c269d | |||
| 80ec939383 | |||
| d3c4d6f133 | |||
| 5ad661f95f | |||
| d63e5c6852 | |||
| fbb8e0d241 | |||
| 8d9a496935 | |||
| d1c48db4bc | |||
| 6bd781d344 | |||
| 8a12c56971 | |||
| 26bd59a5a2 | |||
| 927a7beb88 | |||
| f7059607af | |||
| 63de44c347 | |||
| f0ede9dacc | |||
| 5711f8b9e3 | |||
| da83510234 | |||
| f51b22bea7 | |||
| f320c2fcee | |||
| 4ba1e50a89 | |||
| 2a7e025251 | |||
| 3da20e9d6a | |||
| b1c0faa748 | |||
| 989f2c220b | |||
| 7e35915b3e | |||
| 24238d249e | |||
| 9614469b45 | |||
| 29ee68824d | |||
| aeaaf44ae0 | |||
| d752c51aec | |||
| 8199d95dc1 | |||
| 97cdb01357 | |||
| 7a66095912 | |||
| 9d1356a20e | |||
| e4471fc300 | |||
| 84a2854b5e | |||
| a1b76093ce | |||
| 1aa30a73db | |||
| dc7c0e2f9e | |||
| e2cb0d901a | |||
| 8e0b029f5f | |||
| 9bbf2535dd | |||
| 83fde83a58 | |||
| bef3d1359d | |||
| 1ff449435f | |||
| 6e21c83fd8 | |||
| 6dc9d522b1 | |||
| 8adcf6840d | |||
| f5ed30e455 | |||
| cacf3f24e7 | |||
| f08ca4ddfa | |||
| e1c2f3c202 | |||
| 9614eb540d | |||
| 1b46596a39 | |||
| e043d61a99 | |||
| ad89782c9b | |||
| cd29265d5d | |||
| 2f36b7c3b6 | |||
| b8b3f3abfa | |||
| b490297cde | |||
| 06148074a2 | |||
| 16a998055c | |||
| 97c9a8e5ce | |||
| 66415fd1fa | |||
| bfe25609a7 | |||
| 36e0512454 | |||
| b02f667107 | |||
| ed2b6f4580 | |||
| cb7f145c76 | |||
| 2b9d2eaeaa | |||
| 61016671ab | |||
| b2c076946b | |||
| 250c5c22b8 | |||
| 4b0b166143 | |||
| 90481a0e4b | |||
| 2cbaf20e55 | |||
| a263a0840c | |||
| 14991cf58b | |||
| ab0b4e350c | |||
| 7b2fb0880d | |||
| 748e02db80 | |||
| 23c55f8925 | |||
| 906d97b391 | |||
| f15fd4f9c1 | |||
| 7aadb088a6 | |||
| 9de399432e | |||
| 5bdd56cfb1 | |||
| 64ea23c21f | |||
| 7071102ab7 | |||
| 9184072839 | |||
| c437682407 | |||
| 22d4f29670 | |||
| afb7ed3cf1 | |||
| f846f252c0 | |||
| f5618d1f0c | |||
| f94ab0a6bf | |||
| 41b52aae74 | |||
| ed36f7d7fd | |||
| 3ba2bfd7f6 | |||
| b344d16dc1 | |||
| 6e12c09952 | |||
| 07460341e3 | |||
| 732b13669f | |||
| 3a8a82ebc9 | |||
| 1c9819f08e | |||
| 447c4f73f9 | |||
| e01b8d1b6d | |||
| 3d70920f3d | |||
| 110593ece1 | |||
| a1f5dce405 | |||
| 50aa60e0b8 | |||
| 2fbb3813aa | |||
| 6dd695611c | |||
| 92acb45775 | |||
| d3e171aa82 | |||
| c6f330eb06 | |||
| 20cfbfd311 | |||
| fb043325e1 | |||
| 06c0259788 | |||
| 3d5fd9dc05 | |||
| 3ba2c62cc1 | |||
| e2ab01f9d2 | |||
| 5186e061a8 | |||
| c4c907d421 | |||
| fa5076f5f1 | |||
| 8b5a4e0efd | |||
| 3eb68baca6 | |||
| ae97adb8b0 | |||
| 79b9fffcaf | |||
| be22852daa | |||
| f5107045c3 | |||
| 2c98763b9b | |||
| d2eb763b9b | |||
| 0f25e7339f | |||
| 87c57681f6 | |||
| 3d0d79360e | |||
| f08b407b72 | |||
| 4ff2c7795f | |||
| 3bfe05ab56 | |||
| 7806dba509 | |||
| ac53f83ac8 | |||
| 385e4593f4 | |||
| f64bb7c883 | |||
| 9d3175d36a | |||
| 4f96abf42c | |||
| d88bcb6070 | |||
| 0cca3b1f5d | |||
| bbc83ab042 | |||
| 2cba6d4512 | |||
| e70450c401 | |||
| 4d3351c774 | |||
| a586257d5e | |||
| b8163091cc | |||
| c7b3af82b4 | |||
| a42b06ba20 | |||
| 8d62973627 | |||
| 8cdefc72a1 | |||
| d3a8dc7930 | |||
| 86ebb62f84 | |||
| 50191ee694 | |||
| 3c35124db4 | |||
| e4ec521bed | |||
| 2111e01142 | |||
| a39eea7ed6 | |||
| 7bcce9953e | |||
| 9746a42e04 | |||
| 47bacc7abb | |||
| 8824948910 | |||
| 8cb11e60e4 | |||
| 26142f0e05 | |||
| 1542a12497 | |||
| a9250206d5 | |||
| 5bd0ba7a72 | |||
| 286fb9dce7 | |||
| 9fa9154dda | |||
| 604c7a7945 | |||
| a0f5e6e2b9 | |||
| 561d65a505 | |||
| 205e2a9908 | |||
| 8c59b6af14 | |||
| b3328b93e5 | |||
| 96a49bb7cd | |||
| 6d0a19c94c | |||
| 6fc6ce0adb | |||
| 8ba5228c01 | |||
| 51a36efb6b | |||
| ebd449d847 | |||
| 7844c0a93f | |||
| 3bfac14397 | |||
| 1c13e1d64a | |||
| 60b86dc40c | |||
| 68481804a7 | |||
| 236ccc62ad | |||
| 35bffdf336 | |||
| 3772b308e9 | |||
| ef6926322d | |||
| 2df7084d5d | |||
| 3d3cc0c08e | |||
| fbc3d9add6 | |||
| 35e45f0914 | |||
| 3e4fa923eb | |||
| 3013ee044d | |||
| adfd3bd052 | |||
| 4023c66508 | |||
| adfe3825ee | |||
| 814fcdc6ba | |||
| 66de1a5520 | |||
| 52e6b31408 | |||
| 142ba36695 | |||
| b949e9bbc0 | |||
| ce3a07512f | |||
| 1c84d19e5f | |||
| fc1b57bde2 | |||
| 075888c97f | |||
| 40709e4ad8 | |||
| 15c369c509 | |||
| a81b9f1e1f | |||
| 0327659355 | |||
| c99bad19ae | |||
| 35f9446ed8 | |||
| 21888d625f | |||
| feb03c3f8e | |||
| 3b07b64a0f | |||
| 6e6375521d | |||
| b1fe9dc5a7 | |||
| 6db2dc8d2a | |||
| 98b03a4629 | |||
| 610bdb4fea | |||
| 68ec69f2e4 | |||
| 451f6c0bb9 | |||
| 3011dd91ca | |||
| ae65b95374 | |||
| a5bbfea9b9 | |||
| ae9c2e1d5e | |||
| 1d3a444df8 | |||
| f044c00a7c | |||
| 7d89c2702b | |||
| 93653cccb8 | |||
| a024492dbf | |||
| 304c68f9fc | |||
| c5f2b14ff4 | |||
| fc8e03f98c | |||
| 16de4b6437 | |||
| 0f30888b00 | |||
| 3e67be6ac3 | |||
| 5ef027b6f0 | |||
| 666b4bf801 | |||
| d593bfee0a | |||
| b7ad66f0e0 | |||
| 249e49c928 | |||
| e54e74ed88 | |||
| 582c5dceed | |||
| a9d8505cdb | |||
| 7c95791e94 | |||
| aa14faa3cb | |||
| cc6b050367 | |||
| bcedf19a08 | |||
| c05ecb58d8 | |||
| 9e3f8809b3 | |||
| 4f057b99ac | |||
| aec807fcb0 | |||
| 79a585d17e | |||
| 671ff6d132 | |||
| 9b2d0297b7 | |||
| f91e643932 | |||
| 35fe405448 | |||
| 524f2ffb8e | |||
| 68b426cdb0 | |||
| 223f3751e8 | |||
| 7861d040df | |||
| 47cf7e76ec | |||
| 8cafa64174 | |||
| ecba0ad725 | |||
| b3757dcf7b | |||
| 3e456ec4d4 | |||
| 5b1efc89f6 | |||
| aee48d011e | |||
| 3217bb3e12 | |||
| 3df686f474 | |||
| 7d4c027d09 | |||
| 31d70a2dd7 | |||
| c9fbb331e2 | |||
| f6224dcbee | |||
| de6689bc1d | |||
| 0fc740470f | |||
| 49d94cc2e9 | |||
| 291298cf7b | |||
| 9532ae8121 | |||
| 7601731a2c | |||
| 3aa88ab9d3 | |||
| c1ba94192d | |||
| 22032dfd6d | |||
| 4cafde2502 | |||
| 8c623b7ad8 | |||
| 43dc954440 | |||
| 51053d390d | |||
| 9278797aa9 | |||
| 39e49d7f77 | |||
| 84c4c06712 | |||
| eab640aa21 | |||
| a516944086 | |||
| be6803ffa1 | |||
| ef4bdd4f9f | |||
| 2f97895732 | |||
| 9e89b88efc | |||
| a57c6397e3 | |||
| 39e071f5ca | |||
| 70d733edaf | |||
| 1c31f56af1 | |||
| f9999a73df | |||
| 11d8187052 | |||
| 86bff552c1 | |||
| d3f790095e | |||
| 95218218e2 | |||
| e700df82d8 | |||
| e19cc02c4d | |||
| 8a5419448f | |||
| c8217549a8 | |||
| 2130414899 | |||
| 7f83a20fa6 | |||
| 317ab0472d | |||
| e5eb0ba5c8 | |||
| b95af4f87d | |||
| 11073b613c |
10
.gitignore
vendored
10
.gitignore
vendored
@@ -1,3 +1,9 @@
|
||||
# build and testing artifacts
|
||||
notarius
|
||||
notarius-output
|
||||
workspace/
|
||||
.codebase-memory/
|
||||
|
||||
# ---> Go
|
||||
# If you prefer the allow list template instead of the deny list, see community template:
|
||||
# https://github.com/github/gitignore/blob/main/community/Golang/Go.AllowList.gitignore
|
||||
@@ -47,7 +53,8 @@ go.work.sum
|
||||
.LSOverride
|
||||
|
||||
# Icon must end with two \r
|
||||
Icon
|
||||
Icon
|
||||
|
||||
|
||||
# Thumbnails
|
||||
._*
|
||||
@@ -67,4 +74,3 @@ Icon
|
||||
Network Trash Folder
|
||||
Temporary Items
|
||||
.apdisk
|
||||
.apdisk
|
||||
|
||||
@@ -1,3 +1,2 @@
|
||||
Please carefully review the documents in `docs/policy` before making any changes to this repository.
|
||||
- `architecture.md` provides the canonical high-level architecture policy for this repository.
|
||||
- `documentation.md` provides the canonical documentation policy for this repository.
|
||||
Please review `docs/development.md` for initial orientation in this repository
|
||||
and follow its task-specific reading guide.
|
||||
|
||||
68
README.md
68
README.md
@@ -1,36 +1,46 @@
|
||||
# Notarius
|
||||
|
||||
Notarius is a Go CLI for extracting structured artifacts from source material
|
||||
with explicit, configurable pipeline modules.
|
||||
Notarius is a Go CLI for turning source material into structured artifacts with
|
||||
configured extraction pipelines. The implemented D&D workflow reads Seriatim
|
||||
transcript JSON and can produce NPC, location, and item registries; their
|
||||
source-grounded occurrences; scene descriptions, combat turns, enemy events,
|
||||
and spell casts.
|
||||
|
||||
The current implementation reads Seriatim transcript JSON, chunks the source
|
||||
units, extracts D&D spell-cast artifacts with an OpenAI-compatible LLM, and
|
||||
writes JSON output plus diagnostics for each run.
|
||||
## Quickstart
|
||||
|
||||
```sh
|
||||
NOTARIUS_LLM_DEFAULT_BASE_URL=http://127.0.0.1:8080/v1 \
|
||||
NOTARIUS_LLM_DEFAULT_MODEL=your-model \
|
||||
go run ./cmd/notarius run dnd-session \
|
||||
--config examples/dnd-spells.config.yml \
|
||||
--input examples/seriatim-minimal-transcript.json
|
||||
```
|
||||
Provide an OpenRouter API key through the environment, then run the maintained
|
||||
minimal example:
|
||||
|
||||
If the provider requires authentication, set
|
||||
`NOTARIUS_LLM_DEFAULT_API_KEY` in the environment before running the command.
|
||||
Outputs are written under `./notarius-output/<run-id>/` unless `--output-dir`
|
||||
is provided.
|
||||
~~~
|
||||
OPENROUTER_API_KEY=your-api-key \
|
||||
go run ./cmd/notarius run dnd-session \
|
||||
--config examples/dnd-minimal.config.yml \
|
||||
--input examples/seriatim-minimal-transcript.json
|
||||
~~~
|
||||
|
||||
Useful references:
|
||||
The command publishes a JSON output bundle. Its command syntax and exit
|
||||
behavior are documented in the [CLI reference](docs/cli.md); configuration,
|
||||
credentials, and module selection are owned by the
|
||||
[configuration reference](docs/config.md).
|
||||
|
||||
- [CLI reference](docs/cli.md)
|
||||
- [Configuration reference](docs/config.md)
|
||||
- [Operations](docs/operations.md)
|
||||
- [Troubleshooting](docs/troubleshooting.md)
|
||||
- [Seriatim input contract](docs/integrations/seriatim.md)
|
||||
- [OpenAI-compatible provider contract](docs/integrations/openai-compatible.md)
|
||||
- [JSON output contract](docs/integrations/json-output.md)
|
||||
- [D&D spell artifact contract](docs/integrations/dnd-spell-artifacts.md)
|
||||
- [Developer workflow](docs/policy/development.md)
|
||||
- [Internal architecture docs](docs/internal/overview.md)
|
||||
- [Maintained example config](examples/dnd-spells.config.yml)
|
||||
- [Maintained example input](examples/seriatim-minimal-transcript.json)
|
||||
For the complete ordered D&D workflow, use
|
||||
[the complete configuration](examples/dnd-complete.config.yml) with
|
||||
[its synthetic transcript](examples/dnd-complete-transcript.json). It
|
||||
demonstrates all implemented D&D lanes and the supporting campaign references.
|
||||
|
||||
## Documentation
|
||||
|
||||
- [CLI reference](docs/cli.md) — commands, flags, output streams, and exits.
|
||||
- [Configuration reference](docs/config.md) — configuration files, profiles,
|
||||
validation, and module selection.
|
||||
- [Operations](docs/operations.md) — output, state, recovery, and debug
|
||||
handling.
|
||||
- [Integration contracts](docs/integrations/) — Seriatim input and published
|
||||
artifact formats.
|
||||
- [Subprocess consumer guide](docs/consumers/subprocess.md) — invoke Notarius
|
||||
from an orchestrator and consume a published result.
|
||||
- [Internal overview](docs/internal/overview.md) — implemented component map
|
||||
for maintainers.
|
||||
- [Developer guide](docs/development.md) — contributor orientation and
|
||||
validation guidance.
|
||||
- [Future work](docs/roadmap/future.md) — unimplemented ideas and priorities.
|
||||
|
||||
18
assets/dnd/combat-turns/prompts/instructions.md
Normal file
18
assets/dnd/combat-turns/prompts/instructions.md
Normal file
@@ -0,0 +1,18 @@
|
||||
Extract Dungeons & Dragons combat-turn artifacts from the supplied transcript.
|
||||
Include a record only when the transcript establishes that an in-world
|
||||
participant takes a combat turn or performs a discrete interrupting combat
|
||||
event. Keep events in transcript chronology; place an interrupting event where
|
||||
it occurs.
|
||||
|
||||
Exclude initiative setup without a turn or combat event, tactical planning,
|
||||
table talk, rules lookup, hypothetical events, abandoned intentions, recaps
|
||||
outside the current passage, and downstream consequences. Do not infer combat
|
||||
events from Dungeons & Dragons rules knowledge. Preserve the session as played
|
||||
and attribute relevant nonstandard rulings to the GM or table. Unmatched actors
|
||||
remain permitted.
|
||||
|
||||
Treat each record as one turn-level event and keep its supporting transcript
|
||||
evidence together. Use `turn` for a regular combat turn, `reaction` for an
|
||||
off-turn reaction, `legendary_action` for a legendary action,
|
||||
`lair_action` for a lair action, and `other` for another discrete combat
|
||||
event that does not fit those categories.
|
||||
45
assets/dnd/combat-turns/prompts/prompt.yaml
Normal file
45
assets/dnd/combat-turns/prompts/prompt.yaml
Normal file
@@ -0,0 +1,45 @@
|
||||
id: dnd.combat_turns
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: npc_registry
|
||||
required: false
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-npc-registry.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_combat_turns_llm.v1.json
|
||||
repair_attempts: 0
|
||||
41
assets/dnd/combat-turns/schemas/dnd_combat_turns_llm.v1.json
Normal file
41
assets/dnd/combat-turns/schemas/dnd_combat_turns_llm.v1.json
Normal file
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.combat_turns.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["combat_turns"],
|
||||
"properties": {
|
||||
"combat_turns": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["actor", "turn_kind", "source_refs"],
|
||||
"properties": {
|
||||
"actor": {
|
||||
"type": "string"
|
||||
},
|
||||
"turn_kind": {
|
||||
"type": "string"
|
||||
},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {
|
||||
"type": "integer"
|
||||
},
|
||||
"end_unit_id": {
|
||||
"type": "integer"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
12
assets/dnd/enemy-events/prompts/combat-grounding.md
Normal file
12
assets/dnd/enemy-events/prompts/combat-grounding.md
Normal file
@@ -0,0 +1,12 @@
|
||||
Compact combat grounding is supplied below. It can guide attention and
|
||||
disambiguation, but it is not evidence. Do not derive an event, subject,
|
||||
outcome, or source range from either list. The current transcript alone must
|
||||
directly establish every returned event.
|
||||
|
||||
Combat-turn grounding:
|
||||
|
||||
{{ input "combat_turns" }}
|
||||
|
||||
Named combat-opponent grounding:
|
||||
|
||||
{{ input "npc_occurrences" }}
|
||||
25
assets/dnd/enemy-events/prompts/instructions.md
Normal file
25
assets/dnd/enemy-events/prompts/instructions.md
Normal file
@@ -0,0 +1,25 @@
|
||||
Extract Dungeons & Dragons enemy events from the supplied combat transcript.
|
||||
An `engaged` event requires direct establishment that a subject is actively
|
||||
opposing the party in combat. A `killed`, `fled`, `captured`, or
|
||||
`incapacitated` event requires explicit establishment of that outcome. An
|
||||
outcome may share evidence with an engagement, and a later engagement or
|
||||
outcome for the same subject remains a separate observation. Emit at most one
|
||||
`engaged` observation for the same subject in this combat scene.
|
||||
|
||||
For `killed`, direct death or killing is required. For `fled`, the subject
|
||||
must explicitly escape, retreat, or leave combat to avoid continued engagement.
|
||||
For `captured`, the subject must be explicitly taken prisoner or secured
|
||||
under the party's control. For `incapacitated`, the subject must be explicitly
|
||||
unable to continue acting without being established as killed or captured.
|
||||
|
||||
When the transcript identifies a named NPC, use its normalized registry
|
||||
spelling. A hostile creature without a registry entry is allowed. For unnamed
|
||||
individuals or groups, use only the narrowest transcript-grounded label, such
|
||||
as `Orcs`, `One orc`, or `Remaining orcs`; never invent member names, IDs,
|
||||
or quantities.
|
||||
|
||||
Exclude party members, allies, neutral observers, mentioned-but-absent enemies,
|
||||
hazards, traps, environmental effects, uncertain allegiance, table talk,
|
||||
planning, hypotheses, recaps outside this passage, and downstream inference.
|
||||
Do not infer an engagement or outcome from initiative, turn absence, damage,
|
||||
defeat, movement, or a scene ending.
|
||||
53
assets/dnd/enemy-events/prompts/prompt.yaml
Normal file
53
assets/dnd/enemy-events/prompts/prompt.yaml
Normal file
@@ -0,0 +1,53 @@
|
||||
id: dnd.enemy_events
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: npc_registry
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: combat_turns
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: npc_occurrences
|
||||
required: true
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-npc-registry.md
|
||||
- role: user
|
||||
content_file: ./combat-grounding.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_enemy_events_llm.v1.json
|
||||
repair_attempts: 0
|
||||
33
assets/dnd/enemy-events/schemas/dnd_enemy_events_llm.v1.json
Normal file
33
assets/dnd/enemy-events/schemas/dnd_enemy_events_llm.v1.json
Normal file
@@ -0,0 +1,33 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.enemy_events.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["events"],
|
||||
"properties": {
|
||||
"events": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["name", "kind", "source_refs"],
|
||||
"properties": {
|
||||
"name": {"type": "string"},
|
||||
"kind": {"type": "string"},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {"type": "integer"},
|
||||
"end_unit_id": {"type": "integer"}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,47 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.entity_reconcile.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["duplicate_groups"],
|
||||
"properties": {
|
||||
"duplicate_groups": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["members", "canonical"],
|
||||
"properties": {
|
||||
"members": {
|
||||
"type": "array",
|
||||
"items": {"$ref": "#/$defs/selector"}
|
||||
},
|
||||
"canonical": {"$ref": "#/$defs/selector"}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"$defs": {
|
||||
"selector": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["name", "source_refs"],
|
||||
"properties": {
|
||||
"name": {"type": "string", "minLength": 1},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {"type": "integer", "minimum": 1},
|
||||
"end_unit_id": {"type": "integer", "minimum": 1}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
30
assets/dnd/item-occurrences/prompts/instructions.md
Normal file
30
assets/dnd/item-occurrences/prompts/instructions.md
Normal file
@@ -0,0 +1,30 @@
|
||||
Extract meaningful Dungeons & Dragons item and currency occurrences: discoveries and changes
|
||||
in party possession established by the transcript. This is an occurrence history,
|
||||
not an inventory or ledger: do not calculate balances, resolve item identity
|
||||
across records, or infer ownership that the transcript does not establish.
|
||||
|
||||
For every occurrence, use the supplied canonical item `name`. Record a stated
|
||||
quantity as an integer and leave it null when the transcript does not state
|
||||
one. Preserve the stated currency denomination through the selected canonical
|
||||
registry name.
|
||||
|
||||
Use `discovered` when the party learns of or encounters an item without
|
||||
establishing possession. Use `acquired` when the party or a party member gains
|
||||
possession. Use `lost` when party possession ends through a gift, sale, payment,
|
||||
theft, abandonment, or destruction not caused by intended use. Use `consumed`
|
||||
when intended use depletes an expendable item. Monetary spending, purchases, and
|
||||
payments are always `lost`, not `consumed`. Classify currency as `consumed` only
|
||||
when the transcript explicitly describes it being physically destroyed or
|
||||
expended as a non-payment component. Use `transferred` only when possession
|
||||
moves between two distinct named party members.
|
||||
|
||||
Return both `from` and `to` for every occurrence, using `null` when a holder does not
|
||||
apply. For `discovered`, set both holders to `null`. For `acquired`, set `from`
|
||||
to `null` and provide `to`; for `lost` and `consumed`, provide `from` and set
|
||||
`to` to `null`; and for `transferred`, provide both holders. Use `party` only
|
||||
for collective or unresolved party possession, never for either side of a
|
||||
transfer. Do not emit a transfer for a gift, sale, or payment outside the party.
|
||||
|
||||
Ordinary non-depleting use is not an occurrence. Do not infer acquisition from a
|
||||
discovery, or discovery from an acquisition: emit both only when each is
|
||||
independently established.
|
||||
6
assets/dnd/item-occurrences/prompts/item-registry.md
Normal file
6
assets/dnd/item-occurrences/prompts/item-registry.md
Normal file
@@ -0,0 +1,6 @@
|
||||
Use the supplied item registry only to ground each occurrence. Every record
|
||||
must use one registry item's canonical `name`; do not invent, rename, merge,
|
||||
or infer registry items. The registry is not transcript evidence: cite only the
|
||||
current transcript chunk in `source_refs`.
|
||||
|
||||
{{ input "item_registry" }}
|
||||
45
assets/dnd/item-occurrences/prompts/prompt.yaml
Normal file
45
assets/dnd/item-occurrences/prompts/prompt.yaml
Normal file
@@ -0,0 +1,45 @@
|
||||
id: dnd.item_occurrences
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: item_registry
|
||||
required: true
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./item-registry.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_item_occurrences_llm.v1.json
|
||||
repair_attempts: 0
|
||||
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.item_occurrences.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["occurrences"],
|
||||
"properties": {
|
||||
"occurrences": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["name", "kind", "quantity", "from", "to", "source_refs"],
|
||||
"properties": {
|
||||
"name": {"type": "string"},
|
||||
"kind": {"type": "string"},
|
||||
"quantity": {"type": ["integer", "null"]},
|
||||
"from": {"type": ["string", "null"]},
|
||||
"to": {"type": ["string", "null"]},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {"type": "integer"},
|
||||
"end_unit_id": {"type": "integer"}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
12
assets/dnd/item-registry/extract/prompts/instructions.md
Normal file
12
assets/dnd/item-registry/extract/prompts/instructions.md
Normal file
@@ -0,0 +1,12 @@
|
||||
Extract only items established by the provided Dungeons & Dragons transcript.
|
||||
|
||||
Include named unique items, concrete reusable item types, and stable unique
|
||||
designations. Record each currency denomination separately when it is
|
||||
established, such as copper pieces, silver pieces, gold pieces, or platinum
|
||||
pieces. Do not use capitalization as an eligibility test. Keep distinct names
|
||||
and designations as separate candidates; do not merge aliases or invent
|
||||
qualifiers.
|
||||
|
||||
Do not record vague categories such as "loot", "treasure", or "some gear";
|
||||
generic weapons; inferred properties; quantities; or inferred uniqueness. Omit
|
||||
uncertain or unsupported items.
|
||||
40
assets/dnd/item-registry/extract/prompts/prompt.yaml
Normal file
40
assets/dnd/item-registry/extract/prompts/prompt.yaml
Normal file
@@ -0,0 +1,40 @@
|
||||
id: dnd.item_registry
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_item_registry_llm.v1.json
|
||||
repair_attempts: 0
|
||||
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.item_registry.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["items"],
|
||||
"properties": {
|
||||
"items": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["name", "source_refs"],
|
||||
"properties": {
|
||||
"name": {"type": "string"},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {"type": "integer"},
|
||||
"end_unit_id": {"type": "integer"}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
2
assets/dnd/item-registry/normalize/prompts/candidates.md
Normal file
2
assets/dnd/item-registry/normalize/prompts/candidates.md
Normal file
@@ -0,0 +1,2 @@
|
||||
Item candidates:
|
||||
{{ input "candidates" }}
|
||||
@@ -0,0 +1,8 @@
|
||||
Use candidate names and cited transcript windows only to determine whether
|
||||
candidates identify the same item type or unique designation. Do not treat
|
||||
nearby evidence, similar objects, or a shared owner as sufficient. Keep
|
||||
currency denominations, materially different item types, and uncertain aliases
|
||||
separate. Do not infer an item property or uniqueness.
|
||||
|
||||
When selecting a canonical display name, choose one supplied candidate name
|
||||
that is the clearest established designation.
|
||||
30
assets/dnd/item-registry/normalize/prompts/prompt.yaml
Normal file
30
assets/dnd/item-registry/normalize/prompts/prompt.yaml
Normal file
@@ -0,0 +1,30 @@
|
||||
id: dnd.item_registry.normalize
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: candidates
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-entity-reconciliation.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./candidates.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-windows.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_entity_reconcile_llm.v1.json
|
||||
repair_attempts: 0
|
||||
34
assets/dnd/location-occurrences/prompts/instructions.md
Normal file
34
assets/dnd/location-occurrences/prompts/instructions.md
Normal file
@@ -0,0 +1,34 @@
|
||||
Extract Dungeons & Dragons location occurrences from the supplied transcript.
|
||||
Include an occurrence only when the transcript establishes one supplied
|
||||
location, one occurrence kind, and a coherent passage supporting both.
|
||||
|
||||
Use exactly one kind per occurrence:
|
||||
|
||||
- visited: party members are physically present, arrive, remain, or depart;
|
||||
- planned: the party explicitly proposes, intends, or agrees to future travel;
|
||||
- recalled: the transcript explicitly recounts an earlier party visit; or
|
||||
- mentioned: the location is explicitly referenced without stronger support,
|
||||
including non-actionable speculation or a mere hypothetical reference.
|
||||
|
||||
A mere hypothetical or speculative reference is not planned unless the
|
||||
transcript also establishes an actual proposal, intention, or agreement to
|
||||
travel. When the hypothetical explicitly names a supplied location, it may be
|
||||
mentioned.
|
||||
|
||||
A generic phrase in the current chunk may refer to a supplied named registry
|
||||
location only when the chunk's context supports that coreference. It must not
|
||||
create a registry location, and registry content or provenance must never
|
||||
replace current-chunk evidence.
|
||||
|
||||
For every occurrence, return the exact selector from the location registry:
|
||||
the canonical `name`, plus an empty `registry_refs` array for a unique name or
|
||||
the complete ordered `registry_refs` array for a repeated name. Registry ranges
|
||||
and context identify the location only; they are not occurrence evidence.
|
||||
|
||||
For overlapping support, visited outranks planned, recalled, and mentioned;
|
||||
planned outranks recalled and mentioned; recalled outranks mentioned. A passage
|
||||
may produce multiple records when it independently establishes separate facts,
|
||||
such as recalling an earlier visit while planning a return. Omit inferred,
|
||||
unstated, uncertain, or unsupported places and occurrences. Do not infer a
|
||||
location or occurrence from surrounding events when the transcript does not
|
||||
state it. Do not summarize location descriptions.
|
||||
11
assets/dnd/location-occurrences/prompts/location-registry.md
Normal file
11
assets/dnd/location-occurrences/prompts/location-registry.md
Normal file
@@ -0,0 +1,11 @@
|
||||
A contextual location registry is provided below for identity grounding. It may
|
||||
be empty. Every record supplies a canonical display name. A name that appears
|
||||
once is selected with that name and an empty `registry_refs` array. A repeated
|
||||
name is selected only by copying both its name and its complete, ordered
|
||||
`registry_refs` array exactly as supplied.
|
||||
|
||||
Registry content is context, not occurrence evidence. Do not derive an
|
||||
occurrence or `source_refs` range from the registry. Do not invent a location
|
||||
or selector that is absent from it.
|
||||
|
||||
{{ input "location_registry" }}
|
||||
45
assets/dnd/location-occurrences/prompts/prompt.yaml
Normal file
45
assets/dnd/location-occurrences/prompts/prompt.yaml
Normal file
@@ -0,0 +1,45 @@
|
||||
id: dnd.location_occurrences
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: location_registry
|
||||
required: true
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./location-registry.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_location_occurrences_llm.v1.json
|
||||
repair_attempts: 0
|
||||
@@ -0,0 +1,45 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.location_occurrences.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["occurrences"],
|
||||
"properties": {
|
||||
"occurrences": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["name", "registry_refs", "kind", "source_refs"],
|
||||
"properties": {
|
||||
"name": {"type": "string"},
|
||||
"registry_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {"type": "integer", "minimum": 1},
|
||||
"end_unit_id": {"type": "integer", "minimum": 1}
|
||||
}
|
||||
}
|
||||
},
|
||||
"kind": {"enum": ["visited", "planned", "recalled", "mentioned"]},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {"type": "integer"},
|
||||
"end_unit_id": {"type": "integer"}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
13
assets/dnd/location-registry/extract/prompts/instructions.md
Normal file
13
assets/dnd/location-registry/extract/prompts/instructions.md
Normal file
@@ -0,0 +1,13 @@
|
||||
Extract only physical places established by the provided Dungeons & Dragons
|
||||
transcript that have a stable proper name or unique in-world designation. This
|
||||
includes named planes, regions, settlements, districts, buildings, rooms,
|
||||
landmarks, routes, and geographic features.
|
||||
|
||||
Do not create a registry location for generic, temporary, relative, or merely
|
||||
descriptive phrases, including "the room", "the bar", "the hallway",
|
||||
"outside", and "upstairs". Do not use capitalization as an eligibility test.
|
||||
Keep aliases and nested places when the transcript identifies them; do not merge
|
||||
or invent qualifiers for similarly named places.
|
||||
|
||||
Exclude people, creatures, objects, organizations, abstract concepts, and
|
||||
places merely inferred from an event. Omit uncertain or unsupported places.
|
||||
40
assets/dnd/location-registry/extract/prompts/prompt.yaml
Normal file
40
assets/dnd/location-registry/extract/prompts/prompt.yaml
Normal file
@@ -0,0 +1,40 @@
|
||||
id: dnd.location_registry
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_location_registry_llm.v1.json
|
||||
repair_attempts: 0
|
||||
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.location_registry.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["locations"],
|
||||
"properties": {
|
||||
"locations": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["name", "source_refs"],
|
||||
"properties": {
|
||||
"name": {"type": "string"},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {"type": "integer"},
|
||||
"end_unit_id": {"type": "integer"}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,2 @@
|
||||
Location candidates:
|
||||
{{ input "candidates" }}
|
||||
@@ -0,0 +1,6 @@
|
||||
Use candidate names and their cited transcript windows to determine whether
|
||||
candidates identify the same physical place. Do not treat matching names,
|
||||
nearby evidence, nested places, or generic labels as sufficient. Keep parent
|
||||
and child places, similarly named places, and uncertain aliases separate.
|
||||
|
||||
When selecting a canonical display name, prefer the clearest established name.
|
||||
30
assets/dnd/location-registry/normalize/prompts/prompt.yaml
Normal file
30
assets/dnd/location-registry/normalize/prompts/prompt.yaml
Normal file
@@ -0,0 +1,30 @@
|
||||
id: dnd.location_registry.normalize
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: candidates
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-entity-reconciliation.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./candidates.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-windows.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_entity_reconcile_llm.v1.json
|
||||
repair_attempts: 0
|
||||
26
assets/dnd/npc-occurrences/prompts/instructions.md
Normal file
26
assets/dnd/npc-occurrences/prompts/instructions.md
Normal file
@@ -0,0 +1,26 @@
|
||||
Extract Dungeons & Dragons NPC occurrences from the supplied
|
||||
transcript. Include an occurrence only when the transcript establishes one
|
||||
supplied NPC, one occurrence kind, and a coherent passage supporting both.
|
||||
Use the supplied canonical NPC `name`; never invent or substitute a similar
|
||||
name. Cite current-transcript evidence for every occurrence.
|
||||
|
||||
Do not summarize, infer relationships, sentiment, factions, motives, aliases,
|
||||
or persistent state. Do not identify player characters, anonymous groups, or
|
||||
invented NPCs. Split records when an NPC's occurrence kind changes, when
|
||||
combat alignment changes, or when an NPC is first mentioned and later becomes
|
||||
present.
|
||||
|
||||
Use exactly one kind per occurrence:
|
||||
|
||||
- mentioned: the NPC is referred to but is not established as present or communicating;
|
||||
- noncombat_presence: the NPC is present and relevant but does not meaningfully participate in dialogue or combat;
|
||||
- dialogue: the NPC speaks, responds, or is directly engaged in a meaningful non-combat exchange;
|
||||
- combat_ally: the NPC actively participates in combat on the party's side;
|
||||
- combat_opponent: the NPC actively participates in combat against the party; or
|
||||
- other: the transcript clearly establishes a direct NPC occurrence that fits none of the preceding kinds.
|
||||
|
||||
When activities overlap, active combat participation outranks dialogue,
|
||||
presence, and mention; dialogue outranks noncombat presence and mention; and
|
||||
noncombat presence outranks mention. Other is only for directly evidenced
|
||||
activity outside those categories. Split an occurrence rather than assigning
|
||||
both combat alignments.
|
||||
45
assets/dnd/npc-occurrences/prompts/prompt.yaml
Normal file
45
assets/dnd/npc-occurrences/prompts/prompt.yaml
Normal file
@@ -0,0 +1,45 @@
|
||||
id: dnd.npc_occurrences
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: npc_registry
|
||||
required: true
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-npc-registry.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_npc_occurrences_llm.v1.json
|
||||
repair_attempts: 0
|
||||
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.npc_occurrences.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["occurrences"],
|
||||
"properties": {
|
||||
"occurrences": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["name", "kind", "source_refs"],
|
||||
"properties": {
|
||||
"name": {
|
||||
"type": "string"
|
||||
},
|
||||
"kind": {
|
||||
"type": "string"
|
||||
},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {
|
||||
"type": "integer"
|
||||
},
|
||||
"end_unit_id": {
|
||||
"type": "integer"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
19
assets/dnd/npc-registry/extract/prompts/instructions.md
Normal file
19
assets/dnd/npc-registry/extract/prompts/instructions.md
Normal file
@@ -0,0 +1,19 @@
|
||||
Extract the individually identifiable Dungeons & Dragons non-player characters
|
||||
established by the provided transcript.
|
||||
|
||||
Include an in-world non-PC only when the transcript factually establishes a
|
||||
proper name or a stable, individually distinguishing title or alias. A factual
|
||||
third-party mention establishes that identity even when the NPC is not
|
||||
physically present, does not speak, and takes no direct action in this chunk.
|
||||
Record only the NPC identity and the transcript evidence that establishes it;
|
||||
do not infer or classify a separate occurrence.
|
||||
|
||||
Exclude human players, transcript speakers, and the GM as out-of-world people;
|
||||
player characters identified by the player or party references; names used only
|
||||
in hypothetical, speculative, or imagined examples; corrected transcription
|
||||
mistakes; anonymous or generic roles; indistinguishable crowds or groups;
|
||||
invented descriptive labels; and temporary summoned creatures or spell effects
|
||||
without a persistent individual identity.
|
||||
|
||||
Preserve observed display spelling. Do not invent a label for an anonymous
|
||||
creature, crowd, or generic role.
|
||||
40
assets/dnd/npc-registry/extract/prompts/prompt.yaml
Normal file
40
assets/dnd/npc-registry/extract/prompts/prompt.yaml
Normal file
@@ -0,0 +1,40 @@
|
||||
id: dnd.npc_registry
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_npc_registry_llm.v1.json
|
||||
repair_attempts: 0
|
||||
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.npc_registry.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["npcs"],
|
||||
"properties": {
|
||||
"npcs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": [
|
||||
"name",
|
||||
"source_refs"
|
||||
],
|
||||
"properties": {
|
||||
"name": {
|
||||
"type": "string"
|
||||
},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {
|
||||
"type": "integer"
|
||||
},
|
||||
"end_unit_id": {
|
||||
"type": "integer"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
3
assets/dnd/npc-registry/normalize/prompts/candidates.md
Normal file
3
assets/dnd/npc-registry/normalize/prompts/candidates.md
Normal file
@@ -0,0 +1,3 @@
|
||||
NPC candidates for identity comparison:
|
||||
|
||||
{{ input "candidates" }}
|
||||
10
assets/dnd/npc-registry/normalize/prompts/instructions.md
Normal file
10
assets/dnd/npc-registry/normalize/prompts/instructions.md
Normal file
@@ -0,0 +1,10 @@
|
||||
Use candidate aliases and their cited transcript windows to determine whether
|
||||
candidates refer to the same individual. Preserve distinct individuals even
|
||||
when their names are similar.
|
||||
|
||||
When selecting a canonical display name, prefer a complete, stable proper name
|
||||
over an abbreviation. Prefer an unadorned proper name over that name plus a
|
||||
contextual class, role, title, or relationship descriptor unless the transcript
|
||||
establishes the descriptor as part of the person's name. A longer display name
|
||||
is not inherently more canonical; for example, do not prefer `Captain Aria`
|
||||
over `Aria` solely because it includes the contextual title `Captain`.
|
||||
30
assets/dnd/npc-registry/normalize/prompts/prompt.yaml
Normal file
30
assets/dnd/npc-registry/normalize/prompts/prompt.yaml
Normal file
@@ -0,0 +1,30 @@
|
||||
id: dnd.npc_registry.normalize
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: candidates
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-entity-reconciliation.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./candidates.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-windows.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_entity_reconcile_llm.v1.json
|
||||
repair_attempts: 0
|
||||
5
assets/dnd/profiles/dnd-extraction.yaml
Normal file
5
assets/dnd/profiles/dnd-extraction.yaml
Normal file
@@ -0,0 +1,5 @@
|
||||
id: dnd-extraction
|
||||
backend: openrouter
|
||||
model: openai/gpt-5.6-luna
|
||||
timeout_seconds: 240
|
||||
service_tier: flex
|
||||
45
assets/dnd/scene-descriptions/prompts/instructions.md
Normal file
45
assets/dnd/scene-descriptions/prompts/instructions.md
Normal file
@@ -0,0 +1,45 @@
|
||||
Describe exactly one accepted Dungeons & Dragons scene from the supplied
|
||||
transcript chunk. The complete chunk is the evidence boundary: do not split it
|
||||
into multiple scenes or use facts that are not supported by it.
|
||||
|
||||
Return one kind, one concise title, and one concise summary. Choose exactly one
|
||||
kind:
|
||||
|
||||
- combat: active combat materially organizes the scene, including
|
||||
initiative-like exchanges or sustained hostile action. Planning a fight or
|
||||
discussing a completed fight is not combat by itself.
|
||||
- narrative: current-session in-world play that is not principally active
|
||||
combat, a prior-session recap, or sustained out-of-character session
|
||||
discussion. This includes exploration, travel, dialogue, investigation,
|
||||
in-character planning, and aftermath.
|
||||
- recap: the scene's organizing purpose is to recount events from a previous
|
||||
session for the table. An in-world character recounting history during
|
||||
current play remains narrative.
|
||||
- meta: the scene's organizing purpose is sustained out-of-character
|
||||
discussion about the game or session rather than advancing current in-world
|
||||
play.
|
||||
|
||||
Narrative is the default for actual current-session gameplay that does not meet
|
||||
another definition. When the accepted chunk is mixed:
|
||||
|
||||
1. use combat when active combat is a substantive central activity, even with
|
||||
brief setup, rules clarification, or immediate aftermath;
|
||||
2. otherwise use recap when recounting a previous session is the chunk's
|
||||
primary table purpose;
|
||||
3. otherwise use meta when sustained out-of-character session discussion is
|
||||
primary and in-world progression is no more than incidental; and
|
||||
4. use narrative for all remaining current-session in-world play.
|
||||
|
||||
Brief table talk, dice resolution, rules clarification, jokes, or
|
||||
administrative comments do not make a gameplay scene meta. A short recollection
|
||||
used to orient current action does not make a scene recap.
|
||||
|
||||
The title must be a short, distinguishing phrase rather than a sentence,
|
||||
chapter number, or generic label such as "Scene." It may use names and places
|
||||
established by the transcript or disambiguated by campaign references, but it
|
||||
must not invent a proper noun.
|
||||
|
||||
The summary must briefly state the main activity and material transition or
|
||||
outcome established within the accepted chunk. Do not add analysis, inferred
|
||||
motives, hidden state, future consequences, relationship claims, or facts from
|
||||
outside the chunk.
|
||||
38
assets/dnd/scene-descriptions/prompts/prompt.yaml
Normal file
38
assets/dnd/scene-descriptions/prompts/prompt.yaml
Normal file
@@ -0,0 +1,38 @@
|
||||
id: dnd.scene_descriptions
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_scene_descriptions_llm.v1.json
|
||||
repair_attempts: 0
|
||||
@@ -0,0 +1,18 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.scene_descriptions.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["kind", "title", "summary"],
|
||||
"properties": {
|
||||
"kind": {
|
||||
"type": "string"
|
||||
},
|
||||
"title": {
|
||||
"type": "string"
|
||||
},
|
||||
"summary": {
|
||||
"type": "string"
|
||||
}
|
||||
}
|
||||
}
|
||||
19
assets/dnd/scenes/prompts/instructions.md
Normal file
19
assets/dnd/scenes/prompts/instructions.md
Normal file
@@ -0,0 +1,19 @@
|
||||
Divide the complete provided transcript into coherent Dungeons & Dragons scenes
|
||||
for the `dnd/scenes` chunk module.
|
||||
|
||||
A scene is a coherent unit of play. Start a new scene when the transcript
|
||||
establishes a meaningful change in location, objective, threat, activity,
|
||||
encounter, or mode of play. Good reasons include a material move, beginning or
|
||||
ending combat, a substantially different encounter phase, a shift between
|
||||
combat, exploration, social interaction, planning, travel, rest, or downtime,
|
||||
a change in the central NPC, faction, threat, or objective, or a sustained
|
||||
table-level interruption that materially changes the activity.
|
||||
|
||||
Do not split a scene merely because a speaker or combat round changes, a
|
||||
routine turn occurs, or the table briefly digresses. Prefer fewer coherent
|
||||
scenes over speculative or fine-grained boundaries.
|
||||
|
||||
Cover the complete transcript from its first source unit to its last. Return
|
||||
scenes in source-unit order with no gaps or overlaps. Use only positive integer
|
||||
source-unit IDs from the transcript, and give every scene one inclusive
|
||||
`start_unit_id` and one inclusive `end_unit_id`.
|
||||
34
assets/dnd/scenes/prompts/prompt.yaml
Normal file
34
assets/dnd/scenes/prompts/prompt.yaml
Normal file
@@ -0,0 +1,34 @@
|
||||
id: dnd.scenes
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-full.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_scenes_llm.v1.json
|
||||
repair_attempts: 0
|
||||
28
assets/dnd/scenes/schemas/dnd_scenes_llm.v1.json
Normal file
28
assets/dnd/scenes/schemas/dnd_scenes_llm.v1.json
Normal file
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.scenes.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["scenes"],
|
||||
"properties": {
|
||||
"scenes": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
},
|
||||
"end_unit_id": {
|
||||
"type": "integer",
|
||||
"minimum": 1
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
Identify only well-supported duplicate groups among the supplied candidates.
|
||||
|
||||
Return each selected candidate's supplied contextual descriptor exactly: its
|
||||
`name` and complete ordered `source_refs`. A group must contain at least two
|
||||
supplied descriptors, and its `canonical` descriptor must be one of its
|
||||
members. Do not invent names, ranges, records, evidence, or replacement values.
|
||||
Omit any uncertain or unsafe group.
|
||||
@@ -0,0 +1,6 @@
|
||||
Transcript units are the only evidence for extracted events and factual claims.
|
||||
Every reported factual claim must be supported by cited transcript units. Use
|
||||
integer `start_unit_id` and `end_unit_id` values from the transcript.
|
||||
|
||||
When supporting evidence is non-contiguous, use multiple narrow ranges rather
|
||||
than a broad range that bridges unrelated conversation.
|
||||
7
assets/dnd/shared/prompts/common-dnd-identity.md
Normal file
7
assets/dnd/shared/prompts/common-dnd-identity.md
Normal file
@@ -0,0 +1,7 @@
|
||||
Use the most specific supported in-world character or creature identity for
|
||||
each actor or participant. Do not identify a human player, transcript speaker,
|
||||
or the GM as an out-of-world person when an in-world identity is supported.
|
||||
|
||||
Player, party, glossary, campaign, and NPC registry references may disambiguate
|
||||
an identity, but a reference alone cannot establish that the identity
|
||||
participated in the transcript.
|
||||
10
assets/dnd/shared/prompts/common-dnd-npc-registry.md
Normal file
10
assets/dnd/shared/prompts/common-dnd-npc-registry.md
Normal file
@@ -0,0 +1,10 @@
|
||||
A normalized Dungeons & Dragons NPC registry is provided below as grounding
|
||||
material. It may be empty. Use it only to prefer exact canonical participant
|
||||
names when the transcript identifies a participant.
|
||||
|
||||
Registry content is context, not event evidence. Do not extract events,
|
||||
participants, effects, or source references from the registry. Registry source
|
||||
references describe registry provenance and may belong to another session; they
|
||||
are never evidence for the current transcript.
|
||||
|
||||
{{ input "npc_registry" }}
|
||||
12
assets/dnd/shared/prompts/common-dnd-references.md
Normal file
12
assets/dnd/shared/prompts/common-dnd-references.md
Normal file
@@ -0,0 +1,12 @@
|
||||
Optional reference material for this Dungeons & Dragons campaign is provided
|
||||
below. Use it only to disambiguate names, aliases, speakers, campaign terms, or
|
||||
spell names already present in the transcript.
|
||||
|
||||
Player list reference:
|
||||
{{ input "players" }}
|
||||
|
||||
Party roster reference:
|
||||
{{ input "party" }}
|
||||
|
||||
Glossary reference:
|
||||
{{ input "glossary" }}
|
||||
5
assets/dnd/shared/prompts/common-dnd-system.md
Normal file
5
assets/dnd/shared/prompts/common-dnd-system.md
Normal file
@@ -0,0 +1,5 @@
|
||||
You process Dungeons & Dragons gameplay transcripts.
|
||||
|
||||
As input, you will receive one or more portions of a transcript. The transcript may contain transcription errors, repeated lines, incomplete sentences, and misheard proper nouns.
|
||||
|
||||
Return exactly one JSON object that conforms to the configured response schema, with no explanatory prose.
|
||||
3
assets/dnd/shared/prompts/common-dnd-transcript-chunk.md
Normal file
3
assets/dnd/shared/prompts/common-dnd-transcript-chunk.md
Normal file
@@ -0,0 +1,3 @@
|
||||
One extraction chunk from a Dungeons & Dragons gameplay transcript is provided below. Report and infer only what is within this chunk. Its unit IDs retain their source-wide meaning.
|
||||
|
||||
{{ input "transcript" }}
|
||||
3
assets/dnd/shared/prompts/common-dnd-transcript-full.md
Normal file
3
assets/dnd/shared/prompts/common-dnd-transcript-full.md
Normal file
@@ -0,0 +1,3 @@
|
||||
The complete ordered transcript of this Dungeons & Dragons gameplay session is provided below.
|
||||
|
||||
{{ input "transcript" }}
|
||||
@@ -0,0 +1,6 @@
|
||||
Selected Dungeons & Dragons gameplay transcript evidence windows are provided
|
||||
below. They may be incomplete, non-contiguous, or overlapping. Use them to
|
||||
evaluate candidate identity, but do not treat absence outside these windows as
|
||||
evidence.
|
||||
|
||||
{{ input "transcript" }}
|
||||
14
assets/dnd/spells/prompts/instructions.md
Normal file
14
assets/dnd/spells/prompts/instructions.md
Normal file
@@ -0,0 +1,14 @@
|
||||
Extract Dungeons & Dragons spell-cast artifacts from the provided transcript.
|
||||
Include an actual casting event or an unambiguous declared casting attempt.
|
||||
Exclude spell mentions, hypothetical plans, rules discussion, and catalog
|
||||
matches that do not establish a casting event in the transcript.
|
||||
|
||||
For every extracted cast, the transcript evidence must collectively support the
|
||||
in-world caster, the spell, and the fact that the cast or declared attempt
|
||||
occurred.
|
||||
|
||||
Attribute every cast to its in-world caster. Map first-person player speech to
|
||||
the associated player character, and attribute a spell narrated by the GM to
|
||||
the in-world creature that casts it. If the caster cannot be resolved, use only
|
||||
the most specific in-world identity supported by the transcript; do not invent
|
||||
a name.
|
||||
50
assets/dnd/spells/prompts/prompt.yaml
Normal file
50
assets/dnd/spells/prompts/prompt.yaml
Normal file
@@ -0,0 +1,50 @@
|
||||
id: dnd.spells
|
||||
version: "v1"
|
||||
default_profile: dnd-extraction
|
||||
inputs:
|
||||
- name: transcript
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: spell_catalog
|
||||
required: true
|
||||
content_type: application/json
|
||||
- name: players
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: party
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: glossary
|
||||
required: false
|
||||
content_type: text/plain
|
||||
- name: npc_registry
|
||||
required: false
|
||||
content_type: application/json
|
||||
messages:
|
||||
- role: system
|
||||
content_file: ./sharedassets/common-dnd-system.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-identity.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-references.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-transcript-chunk.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-extraction-evidence.md
|
||||
- role: user
|
||||
content_file: ./sharedassets/common-dnd-npc-registry.md
|
||||
- role: user
|
||||
content_file: ./spell-catalog.md
|
||||
- role: user
|
||||
content_file: ./instructions.md
|
||||
cache_control:
|
||||
type: ephemeral
|
||||
output:
|
||||
format: json
|
||||
validation_mode: json_schema
|
||||
schema_path: dnd_spells_llm.v1.json
|
||||
repair_attempts: 0
|
||||
6
assets/dnd/spells/prompts/spell-catalog.md
Normal file
6
assets/dnd/spells/prompts/spell-catalog.md
Normal file
@@ -0,0 +1,6 @@
|
||||
The spell catalog for this extraction is provided below as JSON. Each entry
|
||||
lists a `canonical_name` and its recognized `aliases`. If the transcript uses
|
||||
an alias, select that entry's `canonical_name`. Return spell names using the
|
||||
canonical spelling exactly; never return an alias as a spell name.
|
||||
|
||||
{{ input "spell_catalog" }}
|
||||
48
assets/dnd/spells/schemas/dnd_spells_llm.v1.json
Normal file
48
assets/dnd/spells/schemas/dnd_spells_llm.v1.json
Normal file
@@ -0,0 +1,48 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "notarius.dnd.spells.llm",
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["spell_casts"],
|
||||
"properties": {
|
||||
"spell_casts": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": [
|
||||
"caster",
|
||||
"spell",
|
||||
"source_refs"
|
||||
],
|
||||
"properties": {
|
||||
"caster": {
|
||||
"type": "string",
|
||||
"description": "Canonical in-world character or creature that casts the spell, never the human player, transcript speaker, or GM when the in-world caster can be identified."
|
||||
},
|
||||
"spell": {
|
||||
"type": "string",
|
||||
"description": "Canonical spell name from the provided spell-name catalog."
|
||||
},
|
||||
"source_refs": {
|
||||
"type": "array",
|
||||
"description": "Transcript ranges offered as evidence for the caster, spell name, and casting event in this spell-cast object.",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["start_unit_id", "end_unit_id"],
|
||||
"properties": {
|
||||
"start_unit_id": {
|
||||
"type": "integer"
|
||||
},
|
||||
"end_unit_id": {
|
||||
"type": "integer"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
3
assets/generic/normalize/deduplication/candidates.md
Normal file
3
assets/generic/normalize/deduplication/candidates.md
Normal file
@@ -0,0 +1,3 @@
|
||||
Candidates for identity comparison:
|
||||
|
||||
{{ input "candidates" }}
|
||||
5
assets/generic/normalize/deduplication/instructions.md
Normal file
5
assets/generic/normalize/deduplication/instructions.md
Normal file
@@ -0,0 +1,5 @@
|
||||
Your task is to identify duplicate entities withi the provided list of candidates.
|
||||
|
||||
Please review the listed candidates and their underlying evidence, and determine whether any of the listed candidates refer to the same underlying entity. Preserve distinct candidates that appear to refer to different underlying entities, even when their names are similar.
|
||||
|
||||
When selecting a canonical display name, prefer a complete, stable proper name over an abbreviation. Prefer an unadorned proper name over that name plus additional descriptors, unless the underlying evidence stablishes the descriptors as part of the entity's name. A longer display name is not inherently more canonical.
|
||||
15
assets/package.go
Normal file
15
assets/package.go
Normal file
@@ -0,0 +1,15 @@
|
||||
// Package assets exposes embedded LLM-facing content.
|
||||
package assets
|
||||
|
||||
import (
|
||||
"embed"
|
||||
"io/fs"
|
||||
)
|
||||
|
||||
//go:embed dnd generic
|
||||
var embedded embed.FS
|
||||
|
||||
// FS returns the embedded read-only asset filesystem.
|
||||
func FS() fs.FS {
|
||||
return embedded
|
||||
}
|
||||
23
docs/adr/0001-record-architecture-decisions.md
Normal file
23
docs/adr/0001-record-architecture-decisions.md
Normal file
@@ -0,0 +1,23 @@
|
||||
# ADR-0001: Record architecture decisions as ADRs
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-07-13
|
||||
|
||||
## Context
|
||||
Architectural reasoning made during design (pattern choices, rejected
|
||||
alternatives, trigger conditions for revisiting) is lost if only the final
|
||||
state is documented.
|
||||
|
||||
## Decision
|
||||
We keep a living overview in docs/policy/architecture.md describing current
|
||||
intended state, and immutable, numbered ADRs (Nygard format) in docs/adr/
|
||||
recording each significant decision, its alternatives, and its consequences.
|
||||
Changed decisions get a new ADR that marks the old one Superseded.
|
||||
|
||||
## Alternatives considered
|
||||
- Overview doc only: loses the "why" and the rejected options.
|
||||
- arc42 / RFC-style design docs: heavier than warranted for a solo repo.
|
||||
|
||||
## Consequences
|
||||
Small ongoing writing cost; durable reasoning trail; cheap onboarding for
|
||||
future contributors (including future-us).
|
||||
51
docs/adr/0002-linear-pipes-and-filters-pipeline.md
Normal file
51
docs/adr/0002-linear-pipes-and-filters-pipeline.md
Normal file
@@ -0,0 +1,51 @@
|
||||
# ADR-0002: Linear pipes-and-filters pipeline, not a general DAG
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-07-13
|
||||
|
||||
## Context
|
||||
|
||||
Notarius processes source material through one known workflow:
|
||||
|
||||
```text
|
||||
input -> chunk -> extract -> merge -> normalize -> output
|
||||
```
|
||||
|
||||
Input and chunking apply to the source as a whole. Each selected artifact lane
|
||||
then performs extract, merge, and normalize, after which output aggregates the
|
||||
lane outcomes. Chunk extraction has a natural scatter-gather shape, but no
|
||||
current use case requires arbitrary branches, joins, or user-defined stage
|
||||
topology.
|
||||
|
||||
## Decision
|
||||
|
||||
Notarius implements a fixed six-stage pipes-and-filters pipeline. Configuration
|
||||
selects implementations for these stages but cannot add stages, reorder them,
|
||||
or define an arbitrary graph.
|
||||
|
||||
The framework owns stage sequencing and the scatter-gather boundary between
|
||||
chunk, extract, and merge. Extract results are handed to merge in deterministic
|
||||
source-chunk order regardless of execution strategy. Each artifact lane remains
|
||||
logically linear. Output runs after every selected lane has either produced an
|
||||
accepted normalized artifact or reached a recorded rejection. A framework
|
||||
execution failure aborts the pipeline.
|
||||
|
||||
The runner's concrete internal representation and stage-specific scheduling
|
||||
policies are implementation details. Concurrency must preserve the pipeline's
|
||||
deterministic handoffs, validation behavior, and provenance, and all execution
|
||||
strategies must continue to honor context cancellation.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Build a general DAG engine now. This would support hypothetical branching
|
||||
topologies, but would add scheduling, topology validation, configuration, and
|
||||
state-management complexity without a current consumer. Revisit this choice
|
||||
only when a concrete workflow requires a topology the fixed pipeline cannot
|
||||
express.
|
||||
|
||||
## Consequences
|
||||
|
||||
The runner, configuration model, and operator mental model remain small. Stage
|
||||
ownership stays visible, and general chunking, merging, or normalization cannot
|
||||
be hidden inside extractors. A future DAG requirement will require an explicit
|
||||
architectural change rather than incremental exceptions to the fixed pipeline.
|
||||
119
docs/adr/0003-typed-interfaces-with-two-zone-data-model.md
Normal file
119
docs/adr/0003-typed-interfaces-with-two-zone-data-model.md
Normal file
@@ -0,0 +1,119 @@
|
||||
# ADR-0003: Strongly typed stage interfaces with a two-zone data model
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-07-13
|
||||
|
||||
## Context
|
||||
|
||||
Pipeline stages must exchange source data and extracted artifacts. Universal
|
||||
source data has one engine-wide meaning, while extracted artifacts have
|
||||
domain-specific shapes. Passing opaque bytes or `any` between all stages would
|
||||
make invalid wiring and merge behavior runtime concerns. Requiring JSON at
|
||||
every handoff would preserve interoperability but discard useful Go type safety
|
||||
while all modules are in-process.
|
||||
|
||||
The framework must also support multiple configured artifact domains, durable
|
||||
checkpoints, diagnostics, and output encoders without making those consumers
|
||||
depend on every domain's Go types.
|
||||
|
||||
## Decision
|
||||
|
||||
Notarius uses two typed data zones followed by one serialized boundary.
|
||||
|
||||
### Source zone
|
||||
|
||||
Input and chunk stages use conservative, engine-owned document, segment, chunk,
|
||||
and source-reference types. Their exact Go names are implementation details.
|
||||
Every segment carries engine-owned source provenance identifying the source
|
||||
location from which it was produced. Chunks preserve the ordered provenance of
|
||||
their segments.
|
||||
|
||||
Source-format-specific fields remain in input modules or explicitly namespaced
|
||||
metadata; they do not become framework contracts.
|
||||
|
||||
### Domain artifact zone
|
||||
|
||||
Each artifact lane has one domain-owned Go artifact type `T`. Its extract,
|
||||
merge, normalize, and domain-aware validation implementations use generic,
|
||||
strongly typed contracts over the same `T`. Raw JSON, opaque bytes, and `any`
|
||||
are not stage-handoff contracts within a lane.
|
||||
|
||||
Each registered domain artifact type supplies a codec for `T`. The codec owns:
|
||||
|
||||
- stable schema identity and an explicit schema version;
|
||||
- JSON serialization and deserialization;
|
||||
- the media type and schema metadata required at serialized boundaries; and
|
||||
- rejection of data that cannot be represented by the declared artifact
|
||||
schema.
|
||||
|
||||
An artifact type's JSON representation is a maintained domain contract.
|
||||
Changing it incompatibly requires a new schema version.
|
||||
|
||||
Extract, merge, and normalize may change the contents of `T`, but they do not
|
||||
change the lane's canonical Go artifact type or artifact schema identity. An
|
||||
extractor maps any provider- or prompt-specific response type into `T` before
|
||||
returning. A future lane that requires different artifact types at different
|
||||
stages requires a new architectural decision.
|
||||
|
||||
### Serialized boundary
|
||||
|
||||
After normalization, each typed artifact is converted into an engine-owned
|
||||
serialized artifact containing bytes, media type, and schema metadata. Output
|
||||
aggregation and output encoders consume this type-erased form. Intermediate
|
||||
checkpoint and debug encodings do not become stage-handoff contracts.
|
||||
|
||||
LLM transport, checkpoints, and opt-in debug recording are also explicit
|
||||
serialization boundaries. They may encode or decode a typed artifact through
|
||||
its domain codec, but they do not change the in-memory type used between
|
||||
extract, merge, normalize, and typed validators. Checkpoint reuse requires a
|
||||
compatible schema identity and version.
|
||||
|
||||
An LLM structured-response schema is a module transport contract and may differ
|
||||
from the domain artifact schema. The calling module owns the response type and
|
||||
maps it into the canonical `T`; the artifact codec remains authoritative for
|
||||
artifact checkpoints and output serialization.
|
||||
|
||||
The framework may use private type-erased adapters to store heterogeneous lane
|
||||
registrations and execute configured domains. Such an adapter must assemble a
|
||||
type-consistent lane before execution and must not expose `any` or raw payloads
|
||||
as module-facing handoffs inside the domain artifact zone.
|
||||
|
||||
### Construction and dependencies
|
||||
|
||||
Every module operation accepts `context.Context`. Modules receive stable runtime
|
||||
collaborators through an injected dependency set at construction time. In
|
||||
particular, LLM-using modules receive the application-provided structured LLM
|
||||
client and do not construct provider clients or bypass shared scheduling.
|
||||
|
||||
The application boundary enforces one configurable global upper bound on
|
||||
in-flight LLM calls across all stages, lanes, retries, and validators.
|
||||
|
||||
Configuration options are parsed and validated while a module is constructed,
|
||||
before that module executes. Per-run data such as source material, references,
|
||||
session identity, and lane identity remains operation input rather than a
|
||||
construction dependency.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Pass raw bytes between stages. This maximizes decoupling but moves wiring,
|
||||
parsing, and merge errors to runtime and prevents domain types from being the
|
||||
canonical in-process contract.
|
||||
- Require JSON plus schemas at every stage boundary. This is appropriate for an
|
||||
out-of-process boundary, but adds serialization and parsing inside the current
|
||||
in-process pipeline. The stable codec contract preserves this upgrade path if
|
||||
remote plugins are introduced.
|
||||
- Use a uniform `Process(any) (any, error)` contract. This simplifies a fully
|
||||
dynamic engine but turns incompatible module composition into type assertions
|
||||
and runtime failures. The fixed topology does not require that tradeoff.
|
||||
|
||||
## Consequences
|
||||
|
||||
Domain pipelines gain compile-time handoff safety and explicit merge semantics.
|
||||
Serialization, schema compatibility, checkpoint decoding, and output erasure
|
||||
have named owners. Dynamic registration requires a small erased adapter around
|
||||
each typed lane, and generic stage implementations must be instantiated for a
|
||||
specific artifact type or behavior rather than manipulating arbitrary JSON.
|
||||
|
||||
The engine-owned source model becomes a long-lived contract and must evolve
|
||||
conservatively. Domain authors must maintain a codec and versioned schema in
|
||||
addition to their Go artifact type.
|
||||
81
docs/adr/0004-package-modules-by-domain.md
Normal file
81
docs/adr/0004-package-modules-by-domain.md
Normal file
@@ -0,0 +1,81 @@
|
||||
# ADR-0004: Package modules by domain, not by stage
|
||||
|
||||
**Status:** Accepted — its asset-co-location rule is superseded by [ADR-0011](0011-centralize-llm-assets.md); its domain-first module packaging decision remains accepted.
|
||||
**Date:** 2026-07-13
|
||||
|
||||
## Context
|
||||
|
||||
Module packages can be grouped first by pipeline stage, such as
|
||||
`modules/chunk/dnd/scenes`, or first by domain, such as
|
||||
`modules/dnd/chunk/scenes`. A domain's extract, merge, normalize, validation,
|
||||
schema, prompt, and artifact-codec implementations collaborate around the same
|
||||
artifact types and are likely to evolve together.
|
||||
|
||||
Go package dependencies also constrain registration. If shared types live in a
|
||||
domain root package, that package cannot import child implementation packages
|
||||
to register them because the children already import the root types.
|
||||
|
||||
## Decision
|
||||
|
||||
Production extensions are grouped by domain under:
|
||||
|
||||
```text
|
||||
internal/modules/<domain>/<stage>/<name>
|
||||
```
|
||||
|
||||
Shared artifact types live at the domain root, for example
|
||||
`internal/modules/dnd/types.go`. Domain-specific validators, prompt fragments,
|
||||
schemas, reference helpers, and codecs also live within that domain tree.
|
||||
|
||||
Each domain exposes one production registration entry point from a sibling
|
||||
registrar package, for example `internal/modules/dnd/register`. The registrar
|
||||
may import the domain root and its child implementations; the domain root does
|
||||
not import its registrar or child packages. This keeps shared types available
|
||||
as `dnd.SpellList` without creating a Go import cycle.
|
||||
|
||||
The `generic` tree is a peer extension family for reusable implementations that
|
||||
contain no concrete source-format or artifact-domain knowledge. Source-format
|
||||
and output-format families, such as Seriatim and JSON output, follow the same
|
||||
domain-first organization even when they do not define a type in the
|
||||
[domain artifact zone](0003-typed-interfaces-with-two-zone-data-model.md#domain-artifact-zone).
|
||||
|
||||
Concrete domain implementation packages do not import another concrete domain.
|
||||
Generic extension packages never import concrete domains. A domain registrar
|
||||
may import domain-neutral generic extension packages to instantiate a reusable
|
||||
strategy for that domain's artifact type; the generic implementation remains
|
||||
unaware of the concrete type's domain semantics. Reuse needed directly by a
|
||||
domain implementation lives in a domain-neutral framework or helper package,
|
||||
not in a peer extension package.
|
||||
|
||||
The application composition root may import multiple registrar packages, and
|
||||
black-box integration tests may compose multiple domains. Other cross-domain
|
||||
reuse occurs through engine contracts and composition-time registration rather
|
||||
than concrete peer-domain imports.
|
||||
|
||||
A domain registrar owns registration of that domain's modules, validators,
|
||||
default validator chains, artifact codecs, schemas, and prompt assets. It does
|
||||
not take ownership of application execution or process behavior.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Group modules by stage. This keeps interchangeable strategies side by side,
|
||||
but scatters a domain's shared artifact model and collaborating extensions
|
||||
across the repository. It is preferable when generic strategy libraries
|
||||
dominate or when the project is primarily a stage-extension framework rather
|
||||
than an application composed from domain suites.
|
||||
- Put both shared types and `Register` in the domain root. This gives the
|
||||
shortest import path but creates an import cycle once child implementations
|
||||
import the root artifact types.
|
||||
|
||||
## Consequences
|
||||
|
||||
The repository layout makes supported domains immediately visible, and adding
|
||||
or extracting a domain affects one cohesive subtree. The CLI composition root
|
||||
depends on a small set of domain registrars instead of every leaf package.
|
||||
|
||||
Package moves must preserve user-visible module and validator keys unless a
|
||||
separate compatibility decision changes them. Shared behavior that cannot be
|
||||
expressed through framework contracts may need to move into a domain-neutral
|
||||
framework package rather than creating a concrete peer-domain import. Registrar
|
||||
packages become explicit composition points for instantiating generic typed
|
||||
strategies, in addition to registering domain-owned implementations.
|
||||
141
docs/adr/0005-cache-canonical-chunk-plans-by-source.md
Normal file
141
docs/adr/0005-cache-canonical-chunk-plans-by-source.md
Normal file
@@ -0,0 +1,141 @@
|
||||
# ADR-0005: Cache one canonical chunk plan per source
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-07-17
|
||||
|
||||
## Context
|
||||
|
||||
Notarius may run several extraction passes over the same source. A D&D
|
||||
transcript, for example, may first produce NPC artifacts and later produce
|
||||
spell or combat artifacts, with output from an earlier pass supplied as a
|
||||
reference to a later pass.
|
||||
|
||||
An LLM-backed chunker may process an entire, potentially large source in one
|
||||
expensive request. Recomputing boundaries for every pipeline or pass repeats
|
||||
that cost and can make otherwise comparable extraction runs use different
|
||||
source partitions. Stable chunk material also gives later extraction requests
|
||||
a better opportunity to benefit from provider-side prompt caching.
|
||||
|
||||
Chunk boundaries can affect extraction quality. Evidence may span a boundary,
|
||||
overlap may produce duplicates, and different partitions may change the context
|
||||
available to a model. Merge and normalization should remove structural signs
|
||||
of chunking from durable output, but they cannot guarantee recovery of evidence
|
||||
that an extractor did not receive.
|
||||
|
||||
Notarius therefore needs an explicit policy for choosing between automatically
|
||||
applying the latest chunking configuration and preserving one stable partition
|
||||
for repeated work on the same source.
|
||||
|
||||
## Decision
|
||||
|
||||
Notarius assigns one active canonical chunk plan to a source and reuses that
|
||||
plan by default across pipelines and invocations.
|
||||
|
||||
The canonical source identity is derived from the validated generic source
|
||||
document and covers the source-unit identity, order, and content needed to
|
||||
interpret plan boundaries. Input-adapter and chunk-producer identities are
|
||||
recorded as provenance, but the active-plan lookup does not vary with:
|
||||
|
||||
- pipeline identity or selected artifact lanes;
|
||||
- the configured chunk module or its options;
|
||||
- references;
|
||||
- LLM provider, model, profile, prompt, or response schema; or
|
||||
- configuration for later pipeline stages.
|
||||
|
||||
When an active plan exists, Notarius uses it even if the current pipeline
|
||||
configures a different chunk module or different chunk-module settings. The
|
||||
configured chunk module generates a plan only when none exists or when the
|
||||
operator explicitly requests recomputation.
|
||||
|
||||
The framework-owned minimum plan contract is an ordered, non-empty set of
|
||||
source-unit ranges. Each range identifies the inclusive start and end unit for
|
||||
one chunk. A chunk module may also provide namespaced, domain-specific
|
||||
annotations at plan or range scope. Those annotations are stored with the plan
|
||||
and passed through the pipeline when present, but they remain optional.
|
||||
Downstream stages must not assume that annotations associated with the
|
||||
currently configured chunk module are present on a reused plan produced by a
|
||||
different module.
|
||||
|
||||
The cache stores the plan rather than fully materialized chunks. The framework
|
||||
validates a reused plan against the current source and deterministically
|
||||
materializes its ranges into chunks. The same source and plan must produce
|
||||
byte-stable chunk input for later stages.
|
||||
|
||||
Canonical plan storage is a distinct cache surface with an independently
|
||||
configurable location. It is not coupled to the roots or lifecycles of
|
||||
invocation checkpoints, diagnostics, debug artifacts, or durable output. This
|
||||
allows per-user and system-service deployments to apply cache-specific
|
||||
ownership, permissions, placement, and cleanup policy without relocating other
|
||||
Notarius state.
|
||||
|
||||
One mutable active plan is stored under the canonical source identity and
|
||||
retains provenance for the module and relevant runtime inputs that produced it.
|
||||
Refreshing the active plan atomically replaces that one mutable record; readers
|
||||
must observe either the previous complete plan or the replacement complete
|
||||
plan, never a partial update.
|
||||
The effective plan producer is reported separately from the chunk module
|
||||
requested by the current pipeline; reuse must not attribute cached boundaries
|
||||
or annotations to a module that did not produce them.
|
||||
|
||||
Reuse is enabled by default. Operators can explicitly:
|
||||
|
||||
- bypass cached plans for an invocation without changing the active plan; or
|
||||
- recompute a plan with the configured chunk module and make it active for
|
||||
later work.
|
||||
|
||||
Exact storage layout, configuration fields, CLI syntax, publication mechanics,
|
||||
recovery behavior, and diagnostics are implementation and operational
|
||||
contracts rather than part of this decision.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Recompute chunks on every invocation. This always applies the current
|
||||
chunking configuration, but repeats the most expensive stage and weakens
|
||||
provider-side caching and cross-pass comparability.
|
||||
- Cache every distinct chunking request by including module options,
|
||||
references, prompts, profiles, and other runtime inputs in its identity. This
|
||||
closely associates a cached result with its producing request, but reduces
|
||||
reuse and permits boundary drift across operationally different passes.
|
||||
- Key plans by source plus chunk module and options. This shares plans across
|
||||
pipelines using the same strategy, but changing the configured strategy
|
||||
silently selects a different partition rather than preserving one canonical
|
||||
partition for the source.
|
||||
- Require operators to name or supply a plan for every run. Explicit selection
|
||||
is reproducible and may be useful as an advanced operation, but adds friction
|
||||
to the default workflow and does not provide automatic reuse.
|
||||
- Store fully materialized chunks. This simplifies loading, but duplicates
|
||||
source content and couples durable state to the current chunk representation
|
||||
rather than the stable boundary decision.
|
||||
- Store canonical plans beneath the general workspace root. This would reuse an
|
||||
existing location setting, but it couples a reusable application cache to
|
||||
checkpoint, diagnostic, and debug state that have different ownership,
|
||||
sensitivity, retention, and deployment requirements.
|
||||
|
||||
## Consequences
|
||||
|
||||
Independent pipelines and passes over the same source use stable boundaries by
|
||||
default. This reduces repeated LLM work, improves cross-pass comparability, and
|
||||
increases the opportunity for cached provider reads.
|
||||
|
||||
The configured chunk module may not execute, and its settings may have no
|
||||
effect, when an active plan already exists. Domain-specific annotations reflect
|
||||
the plan's original producer and may be absent or differ from those the current
|
||||
module would produce. User-visible provenance must make the effective plan
|
||||
clear.
|
||||
|
||||
A poor or outdated partition remains active until an operator replaces it.
|
||||
This can preserve suboptimal context boundaries and affect extraction recall or
|
||||
duplication even when merge and normalization hide the partition structure in
|
||||
durable output. Stable reuse is an intentional priority over automatically
|
||||
incorporating later chunk-strategy changes.
|
||||
|
||||
The framework gains a durable minimal chunk-plan contract and deterministic
|
||||
materialization responsibility. Chunk modules must separate required boundary
|
||||
output from optional annotations, and downstream modules may rely only on the
|
||||
minimal boundary contract unless a future decision introduces an explicit plan
|
||||
compatibility mechanism.
|
||||
|
||||
Operators must configure and secure canonical plan storage independently from
|
||||
other workspace state when the per-user default is not appropriate. Removing
|
||||
that cache remains recoverable because Notarius can regenerate it from the
|
||||
source, but doing so may repeat an expensive LLM operation.
|
||||
114
docs/adr/0006-separate-output-cache-and-debug-state.md
Normal file
114
docs/adr/0006-separate-output-cache-and-debug-state.md
Normal file
@@ -0,0 +1,114 @@
|
||||
# ADR-0006: Separate output, cache, and debug state
|
||||
|
||||
**Status:** Superseded by [ADR-0007](0007-separate-checkpoint-recording-from-reuse.md)
|
||||
**Date:** 2026-07-17
|
||||
|
||||
## Context
|
||||
|
||||
Notarius currently exposes a workspace as a shared parent for checkpoints,
|
||||
debug artifacts, and preferred diagnostics settings. Diagnostics are a second
|
||||
inspection surface with their own enablement, directory, retention, and legacy
|
||||
configuration. Durable output uses a separate CLI-selected root, while the
|
||||
canonical chunk-plan cache introduced by ADR-0005 correctly uses an independent
|
||||
cache root.
|
||||
|
||||
These concepts reflect implementation history more than operator intent. A user
|
||||
must understand differences among workspace state, diagnostics, debug artifacts,
|
||||
checkpoints, and chunk plans before deciding where Notarius may write. Some of
|
||||
those distinctions are important internally: a redacted run summary has a
|
||||
different sensitivity from a trace containing source material, prompts, and
|
||||
model responses. They do not require separate public filesystem categories.
|
||||
|
||||
Notarius needs a smaller state model that communicates why data exists, how it
|
||||
may be treated, and whether it is reconstructible.
|
||||
|
||||
## Decision
|
||||
|
||||
Notarius exposes three filesystem surfaces: output, cache, and debug. The
|
||||
public workspace concept and diagnostics as a separate output surface are
|
||||
removed.
|
||||
|
||||
### Output
|
||||
|
||||
Output is the durable result of a run and the only surface intended for normal
|
||||
consumption. It contains the logical files produced by the output stage,
|
||||
including the maintained result, manifest, warning, and rejection contracts.
|
||||
Output is not cache or inspection state.
|
||||
|
||||
### Cache
|
||||
|
||||
Cache contains reconstructible state used to avoid repeated work or resume an
|
||||
interrupted workflow. Canonical chunk plans and invocation checkpoints are
|
||||
distinct cache families with independent identities, compatibility rules,
|
||||
enablement policies, locations, and cleanup lifecycles.
|
||||
|
||||
ADR-0005 continues to govern canonical chunk-plan selection and reuse. Grouping
|
||||
chunk plans and checkpoints under the public cache category does not permit a
|
||||
checkpoint to compete with canonical plan reuse or couple their storage roots.
|
||||
|
||||
Checkpointing is an invocation policy rather than a prerequisite hidden in
|
||||
persistent workspace configuration. An explicit resume invocation may read
|
||||
compatible checkpoints and record replacement checkpoint state for work it
|
||||
executes. Runs that do not request resume perform no checkpoint I/O.
|
||||
|
||||
### Debug
|
||||
|
||||
Debug is an explicitly requested per-run inspection bundle intended for
|
||||
developers and troubleshooting. It is off by default. When enabled, one bundle
|
||||
contains both redacted run summaries and detailed stage and LLM traces. The
|
||||
internal distinction between a safe summary and a sensitive trace remains, but
|
||||
there is one public enablement and location model.
|
||||
|
||||
Debug data is never a cache input and has no automatic retention policy.
|
||||
Notarius does not create a debug directory unless debug is requested, and it
|
||||
does not automatically delete a requested bundle. Credentials remain redacted
|
||||
at every level, while the bundle as a whole is treated as potentially sensitive
|
||||
because traces may contain source, reference, prompt, model-response, and
|
||||
intermediate artifact content.
|
||||
|
||||
Concise progress, warnings, and failures continue to use stdout or stderr. A
|
||||
run without debug may fail without producing a filesystem inspection record.
|
||||
|
||||
Exact configuration fields, CLI flags, default paths, layouts, compatibility
|
||||
handling, and migration mechanics are configuration and operational contracts
|
||||
rather than part of this decision.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Keep workspace, diagnostics, checkpoints, debug, and chunk-plan cache as
|
||||
separate public concepts. This preserves compatibility and the current safe
|
||||
default-on failure records, but retains overlapping configuration and asks
|
||||
operators to reason about implementation-specific categories.
|
||||
- Keep diagnostics as an always-available redacted operational surface and use
|
||||
debug only for sensitive traces. This distinction is useful for a daemon or
|
||||
managed service with an operational logging contract, but the current CLI can
|
||||
report concise failures on stderr and provide inspection data when explicitly
|
||||
requested.
|
||||
- Put all non-output state beneath one physical root. This minimizes path
|
||||
configuration, but couples reconstructible caches to per-run inspection data
|
||||
and couples cache families whose identity, sensitivity, and cleanup policies
|
||||
differ.
|
||||
- Treat checkpoints as durable run state rather than cache. This emphasizes
|
||||
resumability, but checkpoints are derived, compatibility-checked data that may
|
||||
be deleted and recomputed. Cache more accurately describes their lifecycle.
|
||||
|
||||
## Consequences
|
||||
|
||||
The operator model becomes smaller: normal runs produce output and may use
|
||||
cache; developers explicitly request debug. Public configuration no longer
|
||||
exposes a workspace or overlapping diagnostics and debug systems.
|
||||
|
||||
The implementation retains separate collaborators and serializers where their
|
||||
security or lifecycle boundaries differ. Redacted summaries remain useful as
|
||||
the index to a debug bundle, and chunk plans and checkpoints retain separate
|
||||
stores even though both are cache.
|
||||
|
||||
Existing configuration, environment variables, flags, examples, and
|
||||
documentation require a deliberate compatibility transition. Default-on
|
||||
diagnostic directories disappear, so failures without debug are inspectable
|
||||
only through stderr and any durable output completed before the failure.
|
||||
|
||||
Debug becomes easier to request and substantially more complete, but enabling
|
||||
it creates sensitive files that the operator must protect and remove. Cache
|
||||
cleanup is recoverable but may repeat expensive work, while deleting output is
|
||||
data loss from the user's perspective.
|
||||
50
docs/adr/0007-separate-checkpoint-recording-from-reuse.md
Normal file
50
docs/adr/0007-separate-checkpoint-recording-from-reuse.md
Normal file
@@ -0,0 +1,50 @@
|
||||
# ADR-0007: Separate checkpoint recording from reuse
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-07-19
|
||||
|
||||
## Context
|
||||
|
||||
ADR-0006 made checkpoint I/O conditional on an explicit `--resume` invocation.
|
||||
That policy requires an operator to anticipate the need for recovery before a
|
||||
run begins. A failed ordinary run cannot reuse completed work because it did not
|
||||
record checkpoints.
|
||||
|
||||
Recording reconstructible state and authorizing reuse are separate operational
|
||||
decisions. Recording consumes storage and retains sensitive derived application
|
||||
data, while reuse may change which module operations execute during a run.
|
||||
|
||||
## Decision
|
||||
|
||||
ADR-0006's separation of output, cache, and debug surfaces remains in effect;
|
||||
this decision supersedes only its checkpoint invocation policy.
|
||||
|
||||
Checkpoint recording is controlled by an explicit persistent Boolean
|
||||
configuration setting and remains disabled by default. When recording is
|
||||
enabled, every run records checkpoint transitions and reusable approved stage
|
||||
results.
|
||||
|
||||
Checkpoint loading remains an invocation policy. Only a run with `--resume`
|
||||
loads and reuses compatible completed work. A recording-enabled run without
|
||||
`--resume` executes every stage normally and never loads checkpoints. A resume
|
||||
request while recording is disabled is rejected.
|
||||
|
||||
The existing checkpoint identities, compatibility rules, payload format,
|
||||
filesystem root behavior, and pipeline collaborator contracts remain unchanged.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Continue coupling reads and writes to `--resume`. This is safe by default but
|
||||
prevents recovery unless resume was anticipated on the earlier run.
|
||||
- Always record checkpoints. This maximizes recovery but creates potentially
|
||||
sensitive state without explicit operator consent.
|
||||
- Add a multi-value recording policy. This preserves the old behavior as an
|
||||
option but adds configuration complexity without a current need.
|
||||
|
||||
## Consequences
|
||||
|
||||
Operators can opt into recovery-ready runs while keeping checkpoint reuse
|
||||
explicit. Enabled successful, rejected, and failed runs may all leave sensitive
|
||||
checkpoint state, so operators remain responsible for access and retention.
|
||||
Disabled configurations perform no checkpoint I/O, and `--resume` requires the
|
||||
operator to enable recording first.
|
||||
50
docs/adr/0008-ordered-pipeline-steps.md
Normal file
50
docs/adr/0008-ordered-pipeline-steps.md
Normal file
@@ -0,0 +1,50 @@
|
||||
# ADR-0008: Bounded ordered pipeline steps and explicit artifact references
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-07-21
|
||||
|
||||
## Context
|
||||
|
||||
Notarius currently models one pipeline-wide input, chunking plan, artifact
|
||||
lanes, and output boundary. Some workflows need a deterministic handoff from
|
||||
one set of normalized artifacts to a later set of artifacts, such as using
|
||||
extracted NPC records while grounding later combat events. The workflow needs
|
||||
an explicit topology without turning the pipeline into a general-purpose
|
||||
workflow engine.
|
||||
|
||||
## Decision
|
||||
|
||||
Add an ordered collection of pipeline steps. Each step owns one or more
|
||||
artifact lanes, and lanes within a step retain the existing independent
|
||||
execution model. The pipeline continues to have one input, chunk plan, output,
|
||||
and failure boundary. Steps are barriers: a later step may consume only
|
||||
normalized artifacts from an earlier step.
|
||||
|
||||
Generated references use an explicit step-and-lane selector. Reference slots
|
||||
declare the generated artifact kinds and media types they accept. The resolver
|
||||
validates the topology, ordering, lane identity, artifact kind, schema, and
|
||||
codec compatibility before execution. External references remain supported as
|
||||
path sources, and the legacy top-level artifact map is interpreted as an
|
||||
implicit `default` step.
|
||||
|
||||
Pipeline-level references may not select generated artifacts. General DAGs,
|
||||
branches, loops, conditional execution, joins, and inferred dependencies are
|
||||
not part of this model.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- A general DAG would provide more flexibility but would also require a new
|
||||
scheduler, lifecycle model, failure semantics, and provenance model.
|
||||
- Separate pipeline runs connected through filesystem paths would lose the
|
||||
static topology and typed compatibility checks.
|
||||
- Inferring dependencies from module or lane names would make ordering and
|
||||
configuration errors difficult to detect reliably.
|
||||
|
||||
## Consequences
|
||||
|
||||
The resolved pipeline has a deterministic, inspectable topology and can
|
||||
include it in its identity digest. Configuration validation can reject invalid
|
||||
generated bindings before any work begins. Existing single-step profiles keep
|
||||
their behavior through the implicit `default` step. Execution handoff and
|
||||
multi-step scheduling require follow-up work in the runner and checkpoint
|
||||
layers.
|
||||
@@ -0,0 +1,98 @@
|
||||
# ADR-0009: Prefer minimal evidence-grounded extraction artifacts
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-07-22
|
||||
|
||||
## Context
|
||||
|
||||
Notarius is intended to extract structured facts from source material. Several
|
||||
early D&D artifacts grew to include descriptive prose, inferred relationships,
|
||||
immediate outcomes, summaries, and other enrichment alongside the facts that
|
||||
identify an event or entity. Those fields make one model call responsible for
|
||||
both extraction and synthesis.
|
||||
|
||||
In practice, the richer contracts have produced overlapping or weakly grounded
|
||||
fields and have made structurally valid, semantically coherent output harder for
|
||||
cost-effective smaller models. They also increase prompt size, validation and
|
||||
normalization policy, durable schema surface, downstream coupling, and the
|
||||
number of claims whose provenance must be evaluated.
|
||||
|
||||
The application needs a consistent rule for deciding what belongs in an
|
||||
extractor before redesigning the current D&D spell, NPC, and combat-turn
|
||||
contracts or adding new artifact families.
|
||||
|
||||
## Decision
|
||||
|
||||
An extraction module answers one narrowly stated question and returns the
|
||||
smallest durable structured artifact that usefully answers it.
|
||||
|
||||
Every model-produced field in an extraction artifact must:
|
||||
|
||||
- be necessary to answer the extractor's stated question or serve a known
|
||||
downstream consumer;
|
||||
- represent a fact or bounded classification that can be supported directly by
|
||||
cited source ranges;
|
||||
- remain independently meaningful without model-generated explanatory prose;
|
||||
and
|
||||
- justify the additional prompt, schema, validation, normalization, and
|
||||
compatibility surface it creates.
|
||||
|
||||
Source references are required provenance for extracted records. Auxiliary
|
||||
references may disambiguate identities or canonical names, but they do not
|
||||
establish source facts and are not copied into evidence.
|
||||
|
||||
Extraction artifacts do not include narrative summaries, general analysis,
|
||||
speculative enrichment, inferred biography or relationships, or redundant
|
||||
free-text descriptions by default. When such output has a demonstrated use, it
|
||||
belongs in an explicitly named extraction, classification, enrichment, or
|
||||
analysis module with its own contract and evidence policy.
|
||||
|
||||
Occurrence-level facts are not forced into entity-level attributes. A fact
|
||||
that can change between encounters, such as an NPC's role in a scene, belongs
|
||||
on an occurrence artifact rather than as one scalar property of a normalized
|
||||
NPC registry entry.
|
||||
|
||||
Deterministic mapping and normalization may assign application-owned
|
||||
identifiers, canonicalize known catalog values, order and deduplicate evidence,
|
||||
and collapse records under an explicit identity rule. They must not manufacture
|
||||
removed descriptive fields or synthesize missing claims to satisfy an older
|
||||
contract.
|
||||
|
||||
This is a default design rule, not a prohibition on rich artifacts. A richer
|
||||
field is appropriate when its consumer, evidence semantics, and ownership are
|
||||
explicit.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Keep rich schemas and improve prompts or use larger models. This retains
|
||||
potentially convenient prose but does not resolve overlapping field
|
||||
responsibilities, weak provenance, higher cost, or unnecessary downstream
|
||||
coupling.
|
||||
- Make enrichment fields optional. This reduces rejection pressure but leaves
|
||||
ambiguous artifact semantics and inconsistent records, and many strict
|
||||
structured-output providers still require nullable placeholders.
|
||||
- Keep minimal private LLM schemas while preserving rich durable artifacts.
|
||||
Deterministic code would have to invent, default, or separately derive the
|
||||
missing fields, hiding synthesis behind the extraction boundary.
|
||||
- Use one broad session-analysis module. This reduces the number of lanes but
|
||||
couples unrelated facts, schemas, retries, evaluation, and downstream
|
||||
consumers into one model call.
|
||||
|
||||
## Consequences
|
||||
|
||||
Extraction prompts and response schemas become smaller, more focused, and more
|
||||
suitable for lower-cost models. Artifacts carry fewer unsupported claims, and
|
||||
their evidence and validation policies become easier to explain and evaluate.
|
||||
Independent extractors can evolve, retry, and be consumed without requiring
|
||||
unrelated enrichment.
|
||||
|
||||
Some descriptive convenience fields will disappear from primary artifacts.
|
||||
Consumers that genuinely need them may require a separate module and explicit
|
||||
pipeline step. Entity registries may no longer resolve aliases or relationships
|
||||
unless a dedicated, evidence-grounded capability supplies them.
|
||||
|
||||
Removing durable fields is a schema compatibility change. Each affected
|
||||
artifact requires an explicit version and reference policy; private prompt
|
||||
changes alone are insufficient. Current-behavior integration and internal
|
||||
documentation must change with implementation, while the roadmap owns the
|
||||
proposed contract until then.
|
||||
49
docs/adr/0010-workload-oriented-llm-profile-defaults.md
Normal file
49
docs/adr/0010-workload-oriented-llm-profile-defaults.md
Normal file
@@ -0,0 +1,49 @@
|
||||
# ADR-0010: Use workload-oriented LLM profile defaults
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-08-03
|
||||
|
||||
## Context
|
||||
|
||||
LLM-backed D&D operations share an execution-policy choice, but repeating a
|
||||
provider or model-named profile on every module binding ties pipeline structure
|
||||
to a deployment decision. Different environments may require different model,
|
||||
backend, timeout, or reasoning settings while retaining the same workload.
|
||||
|
||||
Notarius also needs a usable default for maintained D&D prompts without making
|
||||
an operator profile mandatory. That default must remain owned by the D&D
|
||||
family, while generic LLM infrastructure stays unaware of domain-specific
|
||||
policy.
|
||||
|
||||
## Decision
|
||||
|
||||
Pipelines may name one workload-oriented default profile, inherited only by
|
||||
selected LLM-backed bindings and validators. Binding-level profile IDs remain
|
||||
intentional exceptions, and the run-wide CLI profile override has highest
|
||||
precedence.
|
||||
|
||||
The D&D family owns an embedded fallback profile named `dnd-extraction`.
|
||||
Operators may provide a complete profile with the same ID through a PromptKit
|
||||
filesystem source. PromptKit selects the higher-precedence matching definition;
|
||||
Notarius does not merge profile documents. Production, development, and local
|
||||
deployments can therefore use different execution policy behind one unchanged
|
||||
pipeline ID.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Repeat a model-named profile on every binding. This makes routine deployment
|
||||
policy changes noisy and obscures the shared workload intent.
|
||||
- Require every deployment to install a profile file. This adds configuration
|
||||
friction and leaves maintained D&D prompts without an application-owned
|
||||
fallback.
|
||||
- Put D&D profile policy in generic LLM infrastructure. This breaks domain
|
||||
ownership and makes generic code depend on one workload.
|
||||
|
||||
## Consequences
|
||||
|
||||
Pipeline configuration expresses workload intent rather than a specific
|
||||
provider or model. Operators can replace the complete execution policy without
|
||||
editing bindings, while binding-level and run-wide exceptions remain available.
|
||||
Profile changes affect resolved pipeline and checkpoint identity, so they may
|
||||
intentionally cause work to be recomputed. The D&D fallback becomes a
|
||||
maintained application execution-policy asset.
|
||||
69
docs/adr/0011-centralize-llm-assets.md
Normal file
69
docs/adr/0011-centralize-llm-assets.md
Normal file
@@ -0,0 +1,69 @@
|
||||
# ADR-0011: Centralize LLM-facing assets in a content-only package
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-08-05
|
||||
|
||||
## Context
|
||||
|
||||
LLM prompts, private response schemas, generic schemas, and fallback profiles
|
||||
are authored and reviewed as content, but package-local embedding scattered that
|
||||
content across implementation trees. Finding all of the assets that contribute
|
||||
to a prompt family required navigating code ownership boundaries rather than a
|
||||
single discoverable content boundary.
|
||||
|
||||
The repository must retain module ownership of prompt semantics, schema
|
||||
identities, registration, and prompt-cache behavior. Durable artifact schemas
|
||||
and non-LLM domain data have different compatibility and ownership rules, so
|
||||
they must not move merely because they are embedded files.
|
||||
|
||||
## Decision
|
||||
|
||||
LLM-facing content is embedded by the root `assets` package. It is a data-only
|
||||
dependency leaf: its single `FS() fs.FS` API returns the read-only embedded
|
||||
filesystem, and the package contains no business logic or internal or PromptKit
|
||||
dependencies. The accepted import path is
|
||||
`gitea.maximumdirect.net/eric/notarius/assets`; it makes repository-owned
|
||||
content available to its consumers, not a public extension contract.
|
||||
|
||||
Consumers scope that filesystem to the subtree they own before reading or
|
||||
registering content. Modules continue to own their manifests, prompt ordering,
|
||||
private response-schema identity, and registration. Centralizing physical files
|
||||
does not centralize domain semantics or transfer those responsibilities to the
|
||||
root package.
|
||||
|
||||
The root package contains prompt content, private LLM response schemas, generic
|
||||
LLM schemas, shared fragments, and fallback profiles. Durable artifact schemas
|
||||
and non-LLM domain data remain with their current owners. A module fingerprint
|
||||
is derived from its manifest-selected module and shared files, rather than from
|
||||
an entire asset tree. The relocation is accepted to cause a one-time checkpoint
|
||||
invalidation.
|
||||
|
||||
This decision supersedes only the physical asset-co-location portion of
|
||||
ADR-0004's decision that places domain-specific prompt fragments and schemas
|
||||
within the domain tree. ADR-0004's domain-first packaging and registrar
|
||||
ownership decisions remain accepted.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
- Keep package-local assets. This preserves physical co-location with code but
|
||||
makes prompt-author discovery and cross-family review unnecessarily costly.
|
||||
- Use `internal/llmassets`. This would hide content from legitimate owners
|
||||
outside the `internal` subtree and would make the root asset boundary depend
|
||||
on implementation-layer placement.
|
||||
- Build a behavioral central registry. This would mix content discovery with
|
||||
prompt selection and registration behavior, moving module semantics into a
|
||||
shared registry.
|
||||
- Use runtime filesystem overlays. This would add runtime configuration and
|
||||
failure modes where compile-time embedded content is sufficient.
|
||||
|
||||
## Consequences
|
||||
|
||||
Prompt authors can find in-scope LLM content in one top-level tree while module
|
||||
packages continue to define its meaning and registration. Consumers have an
|
||||
explicit, narrow dependency on only the content they need. The root package is
|
||||
intentionally importable but must remain a stable, content-only leaf rather
|
||||
than becoming a general extension API.
|
||||
|
||||
The initial relocation invalidates existing checkpoints once. Later checkpoint
|
||||
identity changes remain limited to the manifest-selected prompt and shared
|
||||
content, so unrelated files do not trigger recomputation.
|
||||
@@ -0,0 +1,64 @@
|
||||
# ADR-0012: Resolve opaque entity identifiers deterministically
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-08-08
|
||||
|
||||
## Context
|
||||
|
||||
Entity IDs in durable Notarius artifacts are application-owned, deterministic
|
||||
identifiers. They are useful to artifact consumers, but their hash-based form
|
||||
does not help a model distinguish entities and would make the model reproduce
|
||||
an opaque implementation detail. A plain name is likewise insufficient where
|
||||
multiple supplied records share that name.
|
||||
|
||||
The LLM boundary must preserve the typed artifact and durable-schema ownership
|
||||
of [ADR-0003](0003-typed-interfaces-with-two-zone-data-model.md) and the distinction
|
||||
between disambiguating references and source evidence in
|
||||
[ADR-0009](0009-minimal-evidence-grounded-extraction-artifacts.md).
|
||||
|
||||
## Decision
|
||||
|
||||
Callers present a model with semantic selections: a canonical name when it is
|
||||
unique in the request, or a contextual descriptor containing the name and
|
||||
source coordinates when that context is needed to distinguish supplied
|
||||
records. The model returns only those supplied selections. The caller resolves
|
||||
each accepted selection against the request-local supplied records and attaches
|
||||
the opaque application ID deterministically.
|
||||
|
||||
Source coordinates are permitted in a selection solely as identity context.
|
||||
They neither establish an occurrence fact nor replace that occurrence's
|
||||
current-transcript evidence. A selector must resolve exactly; unknown,
|
||||
ambiguous, partial, reordered, or otherwise unsafe selections are not mapped.
|
||||
Where an operation requires a complete grounded artifact, that failure rejects
|
||||
the complete artifact rather than accepting a partially mapped result.
|
||||
|
||||
An explicitly scoped request-local short label is permitted only when a
|
||||
contextual descriptor would be impractical and the caller can deterministically
|
||||
map the label within that one request. Such a label is not a durable ID, must
|
||||
not escape the request boundary, and requires a concrete justification in its
|
||||
own module contract.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- Ask the model to return durable IDs. This exposes opaque implementation
|
||||
state, does not improve semantic disambiguation, and makes model output
|
||||
depend on hash formatting.
|
||||
- Select by name alone. This cannot safely distinguish same-name records.
|
||||
- Make request-local labels durable identifiers. This would turn prompt
|
||||
presentation into a public identity contract and create avoidable migration
|
||||
pressure.
|
||||
- Let the model invent identifiers or resolve ambiguity. This makes identity
|
||||
assignment non-deterministic and weakens validation.
|
||||
|
||||
## Consequences
|
||||
|
||||
Durable integration contracts retain their exact ID/name pairs while models
|
||||
operate on readable contextual selections. Calling modules must own selector
|
||||
construction, exact resolution, ambiguity handling, and conversion into their
|
||||
durable artifact type; PromptKit and its adapter remain transport-only.
|
||||
|
||||
Some ambiguous or invalid proposals are deliberately omitted, retried, or
|
||||
rejected according to the caller's existing failure policy. Internal candidate
|
||||
keys may support deterministic request-local mapping, but they are not
|
||||
model-visible selectors or durable data. This adds local validation work while
|
||||
keeping identity assignment auditable and stable.
|
||||
236
docs/cli.md
236
docs/cli.md
@@ -1,129 +1,175 @@
|
||||
# CLI Reference
|
||||
|
||||
This is the canonical reference for the implemented Notarius command-line
|
||||
interface.
|
||||
interface. For the shortest successful run, see the [README](../README.md).
|
||||
Configuration fields, discovery rules, and selectable module keys are defined
|
||||
in [Configuration](config.md); runtime state and recovery procedures are
|
||||
defined in [Operations](operations.md).
|
||||
|
||||
## Quick Run
|
||||
## Command Summary
|
||||
|
||||
```sh
|
||||
NOTARIUS_LLM_DEFAULT_BASE_URL=http://127.0.0.1:8080/v1 \
|
||||
NOTARIUS_LLM_DEFAULT_MODEL=your-model \
|
||||
go run ./cmd/notarius run dnd-session \
|
||||
--config examples/dnd-spells.config.yml \
|
||||
--input examples/seriatim-minimal-transcript.json
|
||||
```
|
||||
|
||||
Set `NOTARIUS_LLM_DEFAULT_API_KEY` if the OpenAI-compatible provider requires
|
||||
a bearer token.
|
||||
|
||||
## Commands
|
||||
|
||||
```text
|
||||
~~~
|
||||
notarius help
|
||||
notarius run <pipeline-id> --input path/to/source.json [--config path/to/config.yml] [--only lane-a,lane-b]
|
||||
notarius config validate --config path/to/config.yml [--pipeline pipeline-id] [--only lane-a,lane-b]
|
||||
notarius pipelines list --config path/to/config.yml [--json]
|
||||
```
|
||||
notarius run <pipeline-id> --input path/to/source.json [--json] [flags]
|
||||
notarius config validate [--config path/to/config.yml] [--pipeline pipeline-id] [--only lane-a,lane-b]
|
||||
notarius pipelines list [--config path/to/config.yml] [--json]
|
||||
~~~
|
||||
|
||||
Running `notarius` with no arguments, `notarius help`, `notarius --help`, or
|
||||
`notarius -h` prints usage and exits successfully.
|
||||
Running Notarius without arguments, or with **help**, **--help**, or **-h**,
|
||||
writes the command summary to standard output and exits with status 0.
|
||||
|
||||
## `run`
|
||||
## run
|
||||
|
||||
`notarius run <pipeline-id>` executes a configured pipeline against one input
|
||||
file.
|
||||
~~~
|
||||
notarius run <pipeline-id> --input path/to/source.json [--json] [flags]
|
||||
~~~
|
||||
|
||||
Flags:
|
||||
The **run** command executes the named pipeline for one input file. The
|
||||
pipeline ID and **--input** are required.
|
||||
|
||||
- `--input path`: required source input file.
|
||||
- `--config path`: config file path. If omitted, Notarius checks
|
||||
`NOTARIUS_CONFIG`, then `/usr/local/etc/notarius/config.yml`.
|
||||
- `--only lane-a,lane-b`: run only the named artifact lanes. Values are
|
||||
comma-separated and must be non-empty.
|
||||
- `--output-dir path`: output root. The run writes to `<path>/<run-id>/`.
|
||||
Defaults to `./notarius-output`.
|
||||
- `--diagnostics-dir path`: diagnostics work directory override for this
|
||||
invocation.
|
||||
- `--llm-profile id`: override every effective module binding to use one LLM
|
||||
profile.
|
||||
| Flag | Meaning |
|
||||
| --- | --- |
|
||||
| **--config path** | Use this configuration file. When omitted, configuration discovery applies; see [Configuration](config.md). |
|
||||
| **--input path** | Source input file to process. Required. |
|
||||
| **--output-dir path** | Override the configured output root for this run. |
|
||||
| **--json** | Write the successful run-result receipt as JSON to standard output. |
|
||||
| **--chunk_cache auto\|bypass\|refresh** | Override chunk-plan cache handling for this run. |
|
||||
| **--resume** | Reuse compatible recorded checkpoints when checkpoint recording is enabled. |
|
||||
| **--recompute-step step-id** | With **--resume**, recompute the selected ordered step and its dependent lanes. It cannot be combined with **--only**. |
|
||||
| **--debug** | Retain a debug bundle for this run. |
|
||||
| **--debug-dir path** | Override the debug-bundle root. Requires **--debug**. |
|
||||
| **--only lane-a,lane-b** | Run only the selected comma-separated artifact lanes when that selection is valid for the configured pipeline. |
|
||||
| **--llm-profile id** | Highest-precedence configured profile for selected LLM-backed bindings and validators; it replaces binding and [pipeline](config.md#pipelines) defaults. |
|
||||
| **--session-id id** | Override the generated prompt session identifier with a non-empty value for LLM-backed module calls. |
|
||||
| **--reasoning-effort value** | Replace the selected PromptKit profile's reasoning effort for every LLM-backed call in this run. The value must be non-empty and the flag may be specified only once. |
|
||||
| **--clear-reasoning-effort** | Clear reasoning effort inherited from the selected PromptKit profile for every LLM-backed call in this run. |
|
||||
| **--reference selector=path** | Add or replace a file reference binding. Repeatable. |
|
||||
| **--without-reference selector** | Remove a configured optional reference binding. Repeatable. |
|
||||
|
||||
On success, the command prints the completed pipeline ID, approved and rejected
|
||||
artifact counts, and the output directory. If the run completes with warnings,
|
||||
the warning count is printed to stderr.
|
||||
**--chunk_cache** accepts only **auto**, **bypass**, or **refresh**.
|
||||
**--debug-dir**, **--output-dir**, **--session-id**, and
|
||||
**--reasoning-effort**, and **--recompute-step** reject explicit empty values.
|
||||
**--reasoning-effort** and **--clear-reasoning-effort** are mutually exclusive.
|
||||
When neither is present, reasoning effort comes from the selected PromptKit
|
||||
profile. These controls apply to the shared run client, including retries and
|
||||
LLM-backed validators, and do not modify configuration or profile files.
|
||||
Persistent reasoning settings remain a PromptKit profile concern.
|
||||
**--recompute-step** requires **--resume**; checkpoint requirements and reuse
|
||||
behavior are documented in [Operations](operations.md).
|
||||
|
||||
For durable output, diagnostics, retention, and failure inspection, see
|
||||
[Operations](operations.md).
|
||||
Every run uses one effective prompt session. Without **--session-id**, Notarius
|
||||
generates a stable `notarius:v1:` identifier from the trimmed resolved input
|
||||
module key and the input file's exact raw bytes. The same module and bytes
|
||||
therefore produce the same identifier, regardless of pipeline, references,
|
||||
profile, retries, or run settings. An explicit non-empty value replaces that
|
||||
default. Session identifiers are visible to providers; they are non-secret
|
||||
correlation identifiers, not credential storage. See
|
||||
[Operations](operations.md#operational-limits) for privacy and workflow
|
||||
guidance.
|
||||
|
||||
The current `run` command requires the resolved pipeline to use exactly one
|
||||
distinct LLM profile after defaults and overrides are applied.
|
||||
### Reference selectors
|
||||
|
||||
## `config validate`
|
||||
Use **--reference** only for a reference slot declared by the selected
|
||||
configured target. The accepted selector forms are:
|
||||
|
||||
`notarius config validate` loads and validates configuration.
|
||||
| Form | Target |
|
||||
| --- | --- |
|
||||
| slot=path | The unique selected target that declares slot. |
|
||||
| chunk.slot=path | The chunker. |
|
||||
| merge.slot=path | The unique selected merger that declares slot. |
|
||||
| lane.slot=path | The unique extractor, merger, or normalizer in lane that declares slot. |
|
||||
| lane.extract.slot=path | The extractor in lane. |
|
||||
| lane.merge.slot=path | The merger in lane. |
|
||||
| lane.normalize.slot=path | The normalizer in lane. |
|
||||
|
||||
Flags:
|
||||
**--without-reference** uses the same selector forms without =path. Slot
|
||||
names, requiredness, and configured bindings are part of the
|
||||
[configuration contract](config.md).
|
||||
|
||||
- `--config path`: config file path. If omitted, discovery uses
|
||||
`NOTARIUS_CONFIG`, then `/usr/local/etc/notarius/config.yml`.
|
||||
- `--pipeline pipeline-id`: additionally resolve one configured pipeline against
|
||||
the production module catalog.
|
||||
- `--only lane-a,lane-b`: validate resolution for selected artifact lanes. This
|
||||
flag requires `--pipeline`.
|
||||
### Run output
|
||||
|
||||
Without **--json**, standard output contains the completed pipeline ID, counts
|
||||
of normalized and rejected outputs, and the output directory. A debug-enabled
|
||||
run also prints its debug-bundle path to standard output. A successful run with
|
||||
warnings reports the warning count to standard error. The published JSON bundle
|
||||
is defined by the [JSON output contract](integrations/json-output.md).
|
||||
|
||||
With **--json**, successful standard output is exactly one
|
||||
`notarius.run-result.v1` JSON document followed by a newline, with no
|
||||
human-oriented status or debug-path line. Its fields and compatibility policy
|
||||
are defined by the [run-result contract](integrations/run-result.md). A caller
|
||||
must check for exit status 0 before decoding this output; a failed write can
|
||||
leave incomplete standard-output bytes that are not a result document.
|
||||
|
||||
Example:
|
||||
|
||||
~~~
|
||||
OPENROUTER_API_KEY=your-api-key \
|
||||
go run ./cmd/notarius run dnd-session \
|
||||
--config examples/dnd-minimal.config.yml \
|
||||
--input examples/seriatim-minimal-transcript.json
|
||||
~~~
|
||||
|
||||
## config validate
|
||||
|
||||
~~~
|
||||
notarius config validate [--config path/to/config.yml] [--pipeline pipeline-id] [--only lane-a,lane-b]
|
||||
~~~
|
||||
|
||||
This command loads and validates a configuration. With **--pipeline**, it also
|
||||
resolves that pipeline against the production module catalog. **--only** selects
|
||||
lanes during that resolution and requires **--pipeline**.
|
||||
|
||||
Success is written to standard output as either config "<path>" is valid or
|
||||
config "<path>" is valid for pipeline "<pipeline-id>".
|
||||
|
||||
Examples:
|
||||
|
||||
```sh
|
||||
~~~
|
||||
go run ./cmd/notarius config validate \
|
||||
--config examples/dnd-spells.config.yml
|
||||
--config examples/dnd-minimal.config.yml \
|
||||
--pipeline dnd-session
|
||||
|
||||
go run ./cmd/notarius config validate \
|
||||
--config examples/dnd-spells.config.yml \
|
||||
--pipeline dnd-session \
|
||||
--only spells
|
||||
```
|
||||
OPENROUTER_API_KEY=validation-placeholder \
|
||||
go run ./cmd/notarius config validate \
|
||||
--config examples/dnd-complete.config.yml \
|
||||
--pipeline dnd-session
|
||||
~~~
|
||||
|
||||
## `pipelines list`
|
||||
The placeholder in the second command is sufficient only for offline
|
||||
validation; it cannot run a provider-backed pipeline.
|
||||
|
||||
`notarius pipelines list` prints configured pipeline IDs in sorted order.
|
||||
## pipelines list
|
||||
|
||||
Flags:
|
||||
~~~
|
||||
notarius pipelines list [--config path/to/config.yml] [--json]
|
||||
~~~
|
||||
|
||||
- `--config path`: config file path. If omitted, discovery uses
|
||||
`NOTARIUS_CONFIG`, then `/usr/local/etc/notarius/config.yml`.
|
||||
- `--json`: print `{"pipelines":[...]}` instead of one ID per line.
|
||||
This command lists configured pipeline IDs in sorted order. By default, it
|
||||
writes one ID per line to standard output. **--json** writes an object shaped as
|
||||
{"pipelines":[...]} instead.
|
||||
|
||||
Examples:
|
||||
|
||||
```sh
|
||||
~~~
|
||||
go run ./cmd/notarius pipelines list \
|
||||
--config examples/dnd-spells.config.yml
|
||||
--config examples/dnd-minimal.config.yml
|
||||
~~~
|
||||
|
||||
go run ./cmd/notarius pipelines list \
|
||||
--config examples/dnd-spells.config.yml \
|
||||
--json
|
||||
```
|
||||
## Output Streams And Exit Statuses
|
||||
|
||||
## Exit Codes
|
||||
Successful commands write their primary result to standard output. Warnings and
|
||||
errors are written to standard error.
|
||||
|
||||
- `0`: command succeeded.
|
||||
- `1`: command syntax was valid, but loading config, resolving modules, running
|
||||
the pipeline, calling the provider, writing output, or writing diagnostics
|
||||
failed.
|
||||
- `2`: command syntax was invalid, a command was unknown, a required argument
|
||||
was missing, or a flag value was malformed.
|
||||
For **run --json**, warnings remain on standard error and standard output is a
|
||||
machine-readable success result only. Syntax and runtime diagnostics remain on
|
||||
standard error. Parse the result only after the process exits with status 0.
|
||||
|
||||
## Implemented Production Pipeline Modules
|
||||
| Status | Meaning |
|
||||
| --- | --- |
|
||||
| 0 | The command completed successfully, including root help. |
|
||||
| 1 | Command syntax was valid but configuration loading or validation, pipeline resolution or execution, provider use, output, or requested debug handling failed. |
|
||||
| 2 | The command or flag syntax was invalid, including unknown commands, missing required arguments, invalid flag values, or invalid flag combinations. |
|
||||
|
||||
The production CLI currently registers these module keys:
|
||||
|
||||
- input: `seriatim`
|
||||
- chunk: `generic`
|
||||
- extract: `dnd/spells`
|
||||
- merge: `appendorder`
|
||||
- normalize: `noop`
|
||||
- output: `json`
|
||||
|
||||
The production CLI does not currently register validator modules.
|
||||
|
||||
For YAML structure, defaults, environment overrides, and module binding syntax,
|
||||
see [Configuration](config.md).
|
||||
The root help spellings are the supported help path. Invoking **--help** on
|
||||
**run**, **config validate**, or **pipelines list** is handled by the flag
|
||||
parser as a usage error: it writes an error to standard error and exits with
|
||||
status 2.
|
||||
|
||||
600
docs/config.md
600
docs/config.md
@@ -1,213 +1,489 @@
|
||||
# Configuration
|
||||
|
||||
This is the canonical reference for implemented Notarius configuration.
|
||||
This is the canonical reference for Notarius configuration. Configuration files
|
||||
are YAML and must declare version 4. They select pipelines and their modules;
|
||||
the [CLI reference](cli.md) owns invocation syntax, and
|
||||
[Operations](operations.md) owns run-state procedures.
|
||||
|
||||
Notarius reads YAML config files with `version: 1`. File config is applied over
|
||||
built-in defaults, then environment overrides are applied.
|
||||
## Configuration Discovery And Precedence
|
||||
|
||||
## Discovery
|
||||
Commands that load configuration choose a file in this order:
|
||||
|
||||
Commands that accept `--config` load configuration in this order:
|
||||
1. a non-empty **--config** CLI value;
|
||||
2. a non-empty **NOTARIUS_CONFIG** environment value;
|
||||
3. the installed default file at **/usr/local/etc/notarius/config.yml**, when
|
||||
it exists.
|
||||
|
||||
1. the `--config` path, when provided;
|
||||
2. `NOTARIUS_CONFIG`, when set to a non-empty path;
|
||||
3. `/usr/local/etc/notarius/config.yml`.
|
||||
The command fails if none of these paths provides a configuration file.
|
||||
|
||||
If none is available, the command fails with a config file not found error.
|
||||
For configuration values, precedence is:
|
||||
|
||||
## Minimal Example
|
||||
1. built-in defaults;
|
||||
2. the selected YAML file;
|
||||
3. supported operational environment variables; and
|
||||
4. the CLI run overrides that apply to a command.
|
||||
|
||||
```yaml
|
||||
version: 1
|
||||
llm_profiles:
|
||||
default:
|
||||
provider: openai-compatible
|
||||
base_url: http://127.0.0.1:8080/v1
|
||||
model: your-model
|
||||
pipelines:
|
||||
dnd-session:
|
||||
input: seriatim
|
||||
chunk:
|
||||
module: generic
|
||||
options:
|
||||
max_units: 50
|
||||
artifacts:
|
||||
spells:
|
||||
extract: dnd/spells
|
||||
```
|
||||
Environment variables do not provide a second configuration schema. They only
|
||||
override the fields listed below.
|
||||
|
||||
The maintained fixture is [examples/dnd-spells.config.yml](../examples/dnd-spells.config.yml).
|
||||
## Maintained Examples
|
||||
|
||||
## Top-Level Fields
|
||||
- [Minimal D&D configuration](../examples/dnd-minimal.config.yml) is a
|
||||
single-lane Seriatim-to-spell pipeline.
|
||||
- [Complete D&D configuration](../examples/dnd-complete.config.yml) uses
|
||||
ordered steps, all implemented D&D lanes, generated references, state
|
||||
settings, bounded LLM concurrency, and the maintained
|
||||
[operator profile](../examples/profiles/dnd-extraction.yml).
|
||||
|
||||
- `version`: required. The only supported value is `1`.
|
||||
- `llm_profiles`: optional map of LLM profile IDs to profile settings.
|
||||
- `pipelines`: optional map of pipeline IDs to pipeline definitions.
|
||||
- `concurrency`: optional global concurrency settings.
|
||||
- `diagnostics`: optional diagnostics settings.
|
||||
Use these complete files as starting points rather than combining the
|
||||
illustrative fragments in this reference.
|
||||
|
||||
Unknown YAML fields are rejected.
|
||||
## File Shape And Defaults
|
||||
|
||||
## Defaults
|
||||
Unknown fields, duplicate mapping keys, empty identifiers, and identifiers that
|
||||
become duplicates after trimming whitespace are rejected. Every top-level field
|
||||
other than **version** is optional.
|
||||
|
||||
Built-in defaults:
|
||||
| Field | Type | Default | Rules |
|
||||
| --- | --- | --- | --- |
|
||||
| **version** | integer | none | Required; must be 4. |
|
||||
| **promptkit** | object | none | Profile source and optional local-backend configuration. |
|
||||
| **pipelines** | map | empty | Maps pipeline IDs to pipeline definitions. |
|
||||
| **concurrency** | object | see below | Global LLM and extraction limits. |
|
||||
| **output** | object | see below | Published output settings. |
|
||||
| **cache** | object | see below | Chunk-plan and checkpoint settings. |
|
||||
| **debug** | object | see below | Debug-bundle root only; it does not enable capture. |
|
||||
|
||||
```yaml
|
||||
llm_profiles:
|
||||
default:
|
||||
provider: openai-compatible
|
||||
timeout: 600
|
||||
max_retries: 3
|
||||
max_concurrency: 1
|
||||
Built-in defaults are:
|
||||
|
||||
| Field | Default |
|
||||
| --- | --- |
|
||||
| **concurrency.total_llm** | 16 |
|
||||
| **concurrency.stage_workers.extract** | Effective **total_llm** |
|
||||
| **output.directory** | **./notarius-output** |
|
||||
| **cache.chunk_plans.mode** | **auto** |
|
||||
| **cache.chunk_plans.directory** | Empty, selecting the per-user chunk-plan root |
|
||||
| **cache.checkpoints.enabled** | false |
|
||||
| **cache.checkpoints.directory** | Empty, selecting the per-user checkpoint root |
|
||||
| **debug.directory** | **./notarius-debug** |
|
||||
|
||||
An empty cache directory in YAML deliberately selects the corresponding
|
||||
per-user root. An explicit empty output or debug directory is invalid.
|
||||
|
||||
## PromptKit Profiles
|
||||
|
||||
The optional **promptkit** object selects one source of profile definitions and
|
||||
may register one conventional local OpenAI-compatible backend:
|
||||
|
||||
~~~yaml
|
||||
version: 4
|
||||
|
||||
promptkit:
|
||||
profile_dir: ./profiles
|
||||
# profile_file: ./profiles.yml
|
||||
local_backend:
|
||||
endpoint: http://localhost:8000/v1
|
||||
concurrency_limit: 2
|
||||
~~~
|
||||
|
||||
| Field | Type | Rules |
|
||||
| --- | --- | --- |
|
||||
| **profile_dir** | string | Non-empty directory containing profile files. |
|
||||
| **profile_file** | string | Non-empty profile file. |
|
||||
| **local_backend** | object | Optional registration for the conventional PromptKit backend ID **local**. |
|
||||
| **local_backend.endpoint** | string | Required when **local_backend** is present; absolute HTTP or HTTPS URL with a host. |
|
||||
| **local_backend.concurrency_limit** | integer | Optional non-negative limit; defaults to 0. |
|
||||
|
||||
Set at most one of **profile_dir** and **profile_file**. Relative values use
|
||||
the process working directory, not the configuration file's directory. The
|
||||
complete example's `./examples/profiles/dnd-extraction.yml` value is therefore
|
||||
valid when Notarius is launched from the repository root; use an absolute path
|
||||
for services and containers.
|
||||
|
||||
An operator source is optional. For a requested ID, PromptKit checks the
|
||||
configured operator source first, then Notarius's embedded fallback profiles,
|
||||
then its own built-in catalog. A matching profile is complete: it replaces a
|
||||
lower-precedence definition rather than merging with it. The maintained
|
||||
[`dnd-extraction` operator profile](../examples/profiles/dnd-extraction.yml)
|
||||
is a secret-free deployment artifact; production, development, and local
|
||||
deployments can each provide a complete definition with that same workload ID.
|
||||
Use workload-oriented IDs for new profiles instead of model names.
|
||||
[Operations](operations.md#promptkit-profile-deployment) owns the deployment
|
||||
workflow and credential-handling guidance.
|
||||
|
||||
When **local_backend** is present, its endpoint is trimmed and must use HTTP or
|
||||
HTTPS case-insensitively, be absolute, and have a non-empty host. URL paths are
|
||||
allowed. User information, queries, and fragments are rejected. A zero
|
||||
**concurrency_limit** leaves the local backend unrestricted inside PromptKit;
|
||||
a positive value limits simultaneous local generations. The application-wide
|
||||
**concurrency.total_llm** limit still applies in both cases. Neither local
|
||||
backend field has an environment override. Omitting **local_backend** registers
|
||||
nothing and preserves existing built-in and endpoint-only profile behavior.
|
||||
|
||||
A file-backed PromptKit profile selects the registration by its case-sensitive
|
||||
backend ID:
|
||||
|
||||
~~~yaml
|
||||
id: local-summary
|
||||
backend: local
|
||||
model: example-model
|
||||
~~~
|
||||
|
||||
Keep credentials out of the local-backend object. A PromptKit profile may name
|
||||
its credential environment variable through `api_key_env`; set that variable
|
||||
only in the run environment. PromptKit owns the
|
||||
[pinned profile-file format](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.5.0/docs/formats.md).
|
||||
The [PromptKit upstream boundary](integrations/pkg-promptkit.md) identifies the
|
||||
supported package API, and [Operations](operations.md#operational-limits)
|
||||
describes the effective concurrency layers.
|
||||
|
||||
`notarius config validate --pipeline <id>` resolves the selected pipeline and
|
||||
inspects every explicit effective profile without contacting a provider or
|
||||
requiring credential values. It rejects absent, malformed, or incompatible
|
||||
profiles before a run prepares modules. Credential availability is checked only
|
||||
when a generation is prepared.
|
||||
|
||||
## Migrating Version 3 Configuration
|
||||
|
||||
Version 3 files are not decoded or rewritten. Change **version: 3** to
|
||||
**version: 4** and rename the top-level **scriptorium:** section to
|
||||
**promptkit:**. Version 4 decoding is strict, so a remaining **scriptorium**
|
||||
field is rejected as unknown.
|
||||
|
||||
## Operational Environment Variables
|
||||
|
||||
These variables are applied after YAML values:
|
||||
|
||||
| Variable | Overrides | Rules |
|
||||
| --- | --- | --- |
|
||||
| **NOTARIUS_TOTAL_LLM_CONCURRENCY** | **concurrency.total_llm** | Integer. |
|
||||
| **NOTARIUS_STAGE_WORKERS_EXTRACT** | **concurrency.stage_workers.extract** | Integer. |
|
||||
| **NOTARIUS_OUTPUT_DIR** | **output.directory** | Non-empty path. |
|
||||
| **NOTARIUS_CACHE_CHUNK_PLANS_MODE** | **cache.chunk_plans.mode** | **auto**, **bypass**, or **refresh**. |
|
||||
| **NOTARIUS_CACHE_CHUNK_PLANS_DIR** | **cache.chunk_plans.directory** | Non-empty path. |
|
||||
| **NOTARIUS_CACHE_CHECKPOINTS_DIR** | **cache.checkpoints.directory** | Non-empty path. |
|
||||
| **NOTARIUS_DEBUG_DIR** | **debug.directory** | Non-empty path. |
|
||||
|
||||
Integer values are trimmed then parsed as base-10 integers. Directory and
|
||||
output values reject NUL characters. **NOTARIUS_CONFIG** participates only in
|
||||
configuration discovery.
|
||||
|
||||
## Concurrency, Output, Cache, And Debug
|
||||
|
||||
~~~yaml
|
||||
concurrency:
|
||||
total_llm: 1
|
||||
diagnostics:
|
||||
work_dir: /tmp/notarius
|
||||
retention: auto
|
||||
```
|
||||
total_llm: 2
|
||||
stage_workers:
|
||||
extract: 2
|
||||
output:
|
||||
directory: ./notarius-output
|
||||
cache:
|
||||
chunk_plans:
|
||||
mode: auto
|
||||
directory: ./notarius-cache/chunk-plans
|
||||
checkpoints:
|
||||
enabled: true
|
||||
directory: ./notarius-cache/checkpoints
|
||||
debug:
|
||||
directory: ./notarius-debug
|
||||
~~~
|
||||
|
||||
No pipelines are built in. A run requires a configured pipeline.
|
||||
**concurrency.total_llm** must be greater than zero. The only supported
|
||||
**concurrency.stage_workers** key is **extract**; its value must be from 1
|
||||
through **total_llm**. When omitted, it is recalculated from the effective
|
||||
**total_llm** after YAML and environment precedence.
|
||||
|
||||
## LLM Profiles
|
||||
|
||||
Each `llm_profiles` entry may contain:
|
||||
|
||||
- `provider`: optional provider key. Empty means `openai-compatible`; any other
|
||||
non-empty value must be `openai-compatible`.
|
||||
- `base_url`: provider base URL. Required for actual LLM calls.
|
||||
- `model`: provider model name. Required for actual LLM calls.
|
||||
- `api_key_env`: environment variable name to read for the API key.
|
||||
- `timeout`: request timeout as whole seconds or a Go-style duration string such
|
||||
as `10m`.
|
||||
- `max_retries`: retry count for provider calls. Must be zero or greater.
|
||||
- `max_concurrency`: per-profile LLM concurrency. Must be zero or greater; when
|
||||
zero, Notarius uses `concurrency.total_llm`.
|
||||
|
||||
Raw API keys are not accepted as file config fields. Use `api_key_env` or an
|
||||
environment override.
|
||||
|
||||
## Environment Overrides
|
||||
|
||||
These environment variables are applied after the config file:
|
||||
|
||||
- `NOTARIUS_CONFIG`: config discovery path.
|
||||
- `NOTARIUS_LLM_DEFAULT_API_KEY`: API key for the `default` LLM profile.
|
||||
- `NOTARIUS_LLM_DEFAULT_BASE_URL`: base URL for the `default` LLM profile.
|
||||
- `NOTARIUS_LLM_DEFAULT_MODEL`: model for the `default` LLM profile.
|
||||
- `NOTARIUS_LLM_DEFAULT_TIMEOUT_SECONDS`: integer timeout seconds for the
|
||||
`default` LLM profile.
|
||||
- `NOTARIUS_LLM_DEFAULT_MAX_RETRIES`: integer retry count for the `default` LLM
|
||||
profile.
|
||||
- `NOTARIUS_LLM_DEFAULT_MAX_CONCURRENCY`: integer max concurrency for the
|
||||
`default` LLM profile.
|
||||
- `NOTARIUS_TOTAL_LLM_CONCURRENCY`: integer global LLM concurrency.
|
||||
- `NOTARIUS_WORK_DIR`: diagnostics work directory.
|
||||
- `NOTARIUS_DIAGNOSTICS_RETENTION`: diagnostics retention mode.
|
||||
|
||||
Integer environment values must parse as base-10 integers.
|
||||
**cache.chunk_plans.mode** accepts **auto**, **bypass**, or **refresh**.
|
||||
**cache.checkpoints.enabled** is a boolean. The CLI can override the output
|
||||
directory and chunk-plan mode for one run; see [CLI reference](cli.md#run).
|
||||
|
||||
## Pipelines
|
||||
|
||||
A pipeline defines the fixed Notarius workflow:
|
||||
Each **pipelines** entry has a unique, non-empty ID and the following shape:
|
||||
|
||||
```text
|
||||
input -> chunk -> extract -> merge -> normalize -> output
|
||||
```
|
||||
~~~yaml
|
||||
pipelines:
|
||||
dnd-session:
|
||||
llm_profile: dnd-extraction
|
||||
input: seriatim
|
||||
chunk: generic
|
||||
output: json
|
||||
artifacts:
|
||||
spells:
|
||||
extract: dnd/spells
|
||||
merge: appendorder
|
||||
normalize: dnd/spells
|
||||
~~~
|
||||
|
||||
Pipeline fields:
|
||||
| Field | Type | Default | Rules |
|
||||
| --- | --- | --- | --- |
|
||||
| **llm_profile** | string | none | Optional non-empty default PromptKit profile ID for selected LLM-backed bindings and validators. An explicitly present blank value is invalid. |
|
||||
| **input** | module binding | none | Required. |
|
||||
| **chunk** | module binding | **generic** | Optional. |
|
||||
| **output** | module binding | **json** | Optional. |
|
||||
| **artifacts** | map | none | Compact single-step lane map. |
|
||||
| **steps** | list | none | Ordered lane definitions. Mutually exclusive with **artifacts**. |
|
||||
| **references** | map | none | External reference defaults for eligible targets. |
|
||||
|
||||
- `input`: required module binding.
|
||||
- `chunk`: optional module binding. Default module is `generic`.
|
||||
- `artifacts`: required for pipeline resolution. It maps artifact lane IDs to
|
||||
lane definitions.
|
||||
- `output`: optional module binding. Default module is `json`.
|
||||
Use either **artifacts** or **steps**. The compact **artifacts** form is an
|
||||
implicit single step. An explicit **steps** list must be non-empty; every step
|
||||
needs a unique non-empty **id**, an **artifacts** map, and may have
|
||||
**references**. A lane ID must not appear more than once in a pipeline,
|
||||
including across explicit steps.
|
||||
|
||||
Artifact lane fields:
|
||||
For each selected LLM-backed binding or validator, profile selection occurs
|
||||
after module, validator, and `--only` lane selection. It uses the
|
||||
run-level **--llm-profile** value first, then the binding's **llm_profile**,
|
||||
then the pipeline's **llm_profile**, and finally the PromptKit default.
|
||||
Deterministic bindings do not receive these defaults or run overrides.
|
||||
|
||||
- `extract`: required module binding.
|
||||
- `merge`: optional module binding. Default module is `appendorder`.
|
||||
- `normalize`: optional module binding. Default module is `noop`.
|
||||
- `validators`: optional list of module bindings. The production CLI currently
|
||||
does not register validator modules.
|
||||
A lane has these fields:
|
||||
|
||||
`notarius run` and `notarius config validate --pipeline` resolve the pipeline
|
||||
against the production module catalog and fail fast for unknown or incompatible
|
||||
module keys.
|
||||
| Field | Type | Default | Rules |
|
||||
| --- | --- | --- | --- |
|
||||
| **extract** | module binding | none | Required. |
|
||||
| **merge** | module binding | **appendorder** | Optional. |
|
||||
| **normalize** | module binding | **noop** | Optional. |
|
||||
| **references** | map | none | Supported compatibility alias for **extract.references**. |
|
||||
| **validators** | list | none | Non-empty lane-level lists are rejected. Set validator overrides on a binding instead. |
|
||||
|
||||
## Module Bindings
|
||||
The lane-level **references** alias remains accepted. When the alias and
|
||||
**extract.references** bind the same slot, **extract.references** wins. Use
|
||||
the binding-local form in new configurations.
|
||||
|
||||
Every module binding may use shorthand:
|
||||
## Module Bindings And Validators
|
||||
|
||||
```yaml
|
||||
Use a module key directly when no other binding fields are needed:
|
||||
|
||||
~~~yaml
|
||||
input: seriatim
|
||||
```
|
||||
~~~
|
||||
|
||||
or object form:
|
||||
Use an object for fields:
|
||||
|
||||
```yaml
|
||||
chunk:
|
||||
module: generic
|
||||
llm_profile: default
|
||||
~~~yaml
|
||||
extract:
|
||||
module: dnd/spells
|
||||
llm_profile: dnd-extraction
|
||||
retries: 2
|
||||
references:
|
||||
spell_catalog: ./dnd-spell-catalog.json
|
||||
~~~
|
||||
|
||||
| Binding field | Type | Default | Rules |
|
||||
| --- | --- | --- | --- |
|
||||
| **module** | string | none | Required for an object binding. Must be a registered compatible key. |
|
||||
| **llm_profile** | string | none | Optional non-empty PromptKit profile ID for an LLM-backed binding. It overrides the pipeline default unless the run supplies **--llm-profile**. |
|
||||
| **retries** | integer | 0 | Non-negative additional attempts for chunk, extract, merge, and normalize bindings. |
|
||||
| **options** | object | none | Must satisfy the selected module. |
|
||||
| **references** | map | none | Valid only on chunk, extract, merge, and normalize bindings. |
|
||||
| **validators** | list | production chain | Valid only on chunk, extract, merge, and normalize bindings. |
|
||||
|
||||
Omitting **validators** uses the registered chain. **validators: []** selects
|
||||
an empty chain; a non-empty list replaces the chain in the listed order.
|
||||
Validator bindings accept only **module**, **llm_profile**, and **options**.
|
||||
They reject **references**, **retries**, and nested **validators**. Deterministic
|
||||
validators reject an explicit **llm_profile**. Deterministic module bindings
|
||||
also reject an explicit **llm_profile**.
|
||||
|
||||
The **json** output module accepts optional **include_chunk_map** and
|
||||
**evidence_context** settings:
|
||||
|
||||
~~~yaml
|
||||
output:
|
||||
module: json
|
||||
options:
|
||||
max_units: 50
|
||||
```
|
||||
include_chunk_map: true
|
||||
evidence_context:
|
||||
enabled: true
|
||||
window_units: 3
|
||||
lanes:
|
||||
- npc-registry
|
||||
- spells
|
||||
~~~
|
||||
|
||||
Binding fields:
|
||||
**include_chunk_map** is a boolean and defaults to false. It adds the accepted
|
||||
chunk map when one exists; its wire format is defined in the
|
||||
[chunk-map contract](integrations/chunk-map.md).
|
||||
|
||||
- `module`: module key.
|
||||
- `llm_profile`: optional LLM profile ID. Empty means `default`.
|
||||
- `options`: optional module-specific settings.
|
||||
Omitting **evidence_context** disables evidence publication. When present, it
|
||||
is an object with these strict fields:
|
||||
|
||||
The `--llm-profile` run flag overrides every effective module binding to use
|
||||
one configured profile.
|
||||
|
||||
## Implemented Production Modules
|
||||
|
||||
| Slot | Key | Notes |
|
||||
| Field | Type | Rules |
|
||||
| --- | --- | --- |
|
||||
| input | `seriatim` | Reads Seriatim transcript JSON. |
|
||||
| chunk | `generic` | Splits source units into ordered chunks. |
|
||||
| extract | `dnd/spells` | Extracts `dnd.spell_cast` artifacts. |
|
||||
| merge | `appendorder` | Keeps candidates in append order. |
|
||||
| normalize | `noop` | Passes merged artifacts through unchanged. |
|
||||
| output | `json` | Produces JSON output files. |
|
||||
| **enabled** | boolean | Required. `false` permits no other evidence fields. |
|
||||
| **lanes** | array of strings | Required and non-empty when enabled. Each value is trimmed and must be unique; every value must name a configured pipeline lane. |
|
||||
| **window_units** | non-negative integer | Optional when enabled; defaults to 3. Zero retains only directly cited units. |
|
||||
|
||||
The `generic` chunker accepts:
|
||||
Unknown outer or nested option fields are rejected, as are incompatible YAML
|
||||
types. The allowlist remains valid when a run uses lane filtering: a configured
|
||||
lane that is not active for that invocation simply contributes no evidence.
|
||||
Evidence publication is opt-in because it can persist source text and metadata.
|
||||
Its payload contract is [Published Evidence Context](integrations/evidence-context.md).
|
||||
|
||||
- `max_units`: positive integer, default `50`;
|
||||
- `overlap_units`: non-negative integer, default `0`, and must be less than
|
||||
`max_units`.
|
||||
## References And Ordered Handoffs
|
||||
|
||||
## Diagnostics
|
||||
Reference maps bind named slots that the selected target declares. A scalar is
|
||||
an external path. Pipeline-level maps accept only external paths; step-local
|
||||
and binding-local maps may also select a normalized artifact from an earlier
|
||||
step:
|
||||
|
||||
`diagnostics` fields:
|
||||
~~~yaml
|
||||
steps:
|
||||
- id: describe-session
|
||||
artifacts:
|
||||
npc-registry:
|
||||
extract: dnd/npc-registry
|
||||
normalize: dnd/npc-registry
|
||||
- id: extract-events
|
||||
references:
|
||||
npc_registry:
|
||||
artifact:
|
||||
step: describe-session
|
||||
lane: npc-registry
|
||||
artifacts:
|
||||
spells:
|
||||
extract: dnd/spells
|
||||
normalize: dnd/spells
|
||||
~~~
|
||||
|
||||
- `work_dir`: directory for per-run diagnostics. Default: `/tmp/notarius`.
|
||||
- `retention`: `auto`, `always`, or `never`. Empty uses `auto`.
|
||||
An artifact selector contains only **step** and **lane**. The producer must be
|
||||
an earlier step and the selected artifact must be compatible with the consumer
|
||||
slot. A generated binding supplies one accepted normalized artifact; it does
|
||||
not name a file. A configured generated dependency remains required even when
|
||||
that consumer slot is otherwise optional.
|
||||
|
||||
`auto` retains diagnostics for failed runs and successful runs with warnings.
|
||||
`always` retains diagnostics for every run. `never` removes diagnostics for
|
||||
successful runs without regard to warnings; failed runs are retained.
|
||||
Pipeline references are defaults. A matching step-local or binding-local
|
||||
external path overrides a pipeline default. Required slots must be bound after
|
||||
these configuration values and any CLI reference overrides are applied.
|
||||
Reference paths in YAML are resolved relative to the configuration file.
|
||||
|
||||
The `--diagnostics-dir` run flag overrides `diagnostics.work_dir` for that
|
||||
invocation.
|
||||
### D&D Reference Slots
|
||||
|
||||
The following slot names are accepted by the implemented D&D modules when the
|
||||
selected target declares them:
|
||||
|
||||
| Slot | Source and use |
|
||||
| --- | --- |
|
||||
| **party** | Optional text campaign context. This is the canonical party-roster spelling. |
|
||||
| **roster** | Accepted compatibility alias for **party**. |
|
||||
| **players** | Optional text player context. |
|
||||
| **glossary** | Optional text campaign glossary. |
|
||||
| **spell_catalog** | Optional JSON spell-catalog overlay for spell extraction and normalization. See [spell-catalog overlays](integrations/dnd-spell-catalog-overlays.md). |
|
||||
| **location_registry** | Required normalized location registry for location-occurrence extraction and normalization. |
|
||||
| **item_registry** | Required normalized item registry for item-occurrence extraction and normalization. |
|
||||
| **npc_registry** | Normalized NPC registry. Optional for spells and combat turns; required for NPC occurrences and enemy-event extraction and normalization. |
|
||||
| **scene_descriptions** | Required normalized scene-description artifact for combat-turn and enemy-event extraction. |
|
||||
| **combat_turns** | Required normalized combat-turn artifact for enemy-event extraction. |
|
||||
| **npc_occurrences** | Required normalized NPC-occurrence artifact for enemy-event extraction. |
|
||||
|
||||
Registry-backed occurrence and enemy-event artifact slots have the following
|
||||
exact binding contracts. Durable semantics and wire shapes remain in their
|
||||
[NPC occurrence](integrations/dnd-npc-occurrence-artifacts.md),
|
||||
[location occurrence](integrations/dnd-location-occurrence-artifacts.md),
|
||||
[item occurrence](integrations/dnd-item-occurrence-artifacts.md), and
|
||||
[enemy-event](integrations/dnd-enemy-event-artifacts.md) contracts.
|
||||
|
||||
| Slot | Accepted artifact kind | Media type | Maximum size | Required stage |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `npc_registry` | `dnd/npc-registry` | `application/json` | 1,048,576 bytes | extract and normalize |
|
||||
| `scene_descriptions` | `dnd/scene-description-list` | `application/json` | 1,048,576 bytes | extract only |
|
||||
| `combat_turns` | `dnd/combat-turn-list` | `application/json` | 1,048,576 bytes | extract only |
|
||||
| `npc_occurrences` | `dnd/npc-occurrence-list` | `application/json` | 1,048,576 bytes | extract only |
|
||||
| `location_registry` | `dnd/location-registry` | `application/json` | 1,048,576 bytes | location-occurrence extract and normalize |
|
||||
| `item_registry` | `dnd/item-registry` | `application/json` | 1,048,576 bytes | item-occurrence extract and normalize |
|
||||
|
||||
Scene descriptions accept **party**, **players**, and **glossary**, but not
|
||||
**roster**. NPC occurrences require **npc_registry** for both extraction and
|
||||
normalization. Combat turns require **scene_descriptions** for extraction; the
|
||||
normalized combat-turn module may use optional **npc_registry**. Location occurrences
|
||||
require **location_registry** for extraction and normalization. Item occurrences require
|
||||
**item_registry** for extraction and normalization. Enemy-event extraction requires all
|
||||
four of its JSON artifact slots; its normalizer requires **npc_registry**.
|
||||
The [complete example](../examples/dnd-complete.config.yml) shows the ordered
|
||||
generated bindings.
|
||||
|
||||
## Production Module Keys
|
||||
|
||||
| Kind | Keys |
|
||||
| --- | --- |
|
||||
| Input | **seriatim** |
|
||||
| Chunk | **generic**, **dnd/scenes** |
|
||||
| Extract | **dnd/spells**, **dnd/npc-registry**, **dnd/combat-turns**, **dnd/item-occurrences**, **dnd/item-registry**, **dnd/npc-occurrences**, **dnd/scene-descriptions**, **dnd/enemy-events**, **dnd/location-registry**, **dnd/location-occurrences** |
|
||||
| Merge | **appendorder** |
|
||||
| Normalize | **noop**, **dnd/spells**, **dnd/npc-registry**, **dnd/combat-turns**, **dnd/item-occurrences**, **dnd/item-registry**, **dnd/npc-occurrences**, **dnd/scene-descriptions**, **dnd/enemy-events**, **dnd/location-registry**, **dnd/location-occurrences** |
|
||||
| Output | **json** |
|
||||
|
||||
`dnd/scenes` and every D&D extractor are `llm_backed`. The
|
||||
`dnd/npc-registry`, `dnd/location-registry`, and `dnd/item-registry`
|
||||
normalizers are also `llm_backed` for bounded duplicate proposals; every other
|
||||
D&D normalizer is `deterministic`. LLM-backed bindings use the effective
|
||||
[PromptKit profile](#promptkit-profiles). The complete example binds each
|
||||
registry in an earlier step before its occurrence consumer.
|
||||
|
||||
The D&D artifact contracts define each emitted schema:
|
||||
[spells](integrations/dnd-spell-artifacts.md),
|
||||
[NPC registry](integrations/dnd-npc-registry-artifacts.md),
|
||||
[NPC occurrences](integrations/dnd-npc-occurrence-artifacts.md),
|
||||
[combat turns](integrations/dnd-combat-turn-artifacts.md),
|
||||
[item registry](integrations/dnd-item-registry-artifacts.md),
|
||||
[item occurrences](integrations/dnd-item-occurrence-artifacts.md),
|
||||
[scene descriptions](integrations/dnd-scene-description-artifacts.md),
|
||||
[enemy events](integrations/dnd-enemy-event-artifacts.md),
|
||||
[location registry](integrations/dnd-location-registry-artifacts.md), and
|
||||
[location occurrences](integrations/dnd-location-occurrence-artifacts.md).
|
||||
|
||||
## Production Validator Keys And Default Chains
|
||||
|
||||
Available validator keys are:
|
||||
|
||||
| Family | Keys |
|
||||
| --- | --- |
|
||||
| Generic | **generic/always_accept**, **generic/always_reject**, **generic/valid_json**, **generic/valid_json_schema** |
|
||||
| Spells | **extract/dnd/spells/shape**, **extract/dnd/spells/catalog**, **extract/dnd/spells/source_refs**, **extract/dnd/spells/source_relatedness** |
|
||||
| NPC registry | **extract/dnd/npc-registry/shape**, **extract/dnd/npc-registry/source_refs**, **extract/dnd/npc-registry/source_relatedness**, **normalize/dnd/npc-registry/identity** |
|
||||
| Combat turns | **extract/dnd/combat-turns/shape**, **extract/dnd/combat-turns/source_refs**, **extract/dnd/combat-turns/source_relatedness**, **normalize/dnd/combat-turns/invariants** |
|
||||
| Item occurrences | **extract/dnd/item-occurrences/shape**, **extract/dnd/item-occurrences/registry**, **extract/dnd/item-occurrences/source_refs**, **extract/dnd/item-occurrences/source_relatedness**, **normalize/dnd/item-occurrences/invariants** |
|
||||
| Item registry | **extract/dnd/item-registry/shape**, **extract/dnd/item-registry/source_refs**, **extract/dnd/item-registry/source_relatedness**, **normalize/dnd/item-registry/identity** |
|
||||
| NPC occurrences | **extract/dnd/npc-occurrences/shape**, **extract/dnd/npc-occurrences/registry**, **extract/dnd/npc-occurrences/source_refs**, **extract/dnd/npc-occurrences/source_relatedness**, **normalize/dnd/npc-occurrences/invariants** |
|
||||
| Scene descriptions | **extract/dnd/scene-descriptions/shape**, **extract/dnd/scene-descriptions/source_refs**, **extract/dnd/scene-descriptions/source_relatedness**, **normalize/dnd/scene-descriptions/invariants** |
|
||||
| Enemy events | **extract/dnd/enemy-events/shape**, **extract/dnd/enemy-events/engagements**, **extract/dnd/enemy-events/source_refs**, **extract/dnd/enemy-events/source_relatedness**, **normalize/dnd/enemy-events/invariants** |
|
||||
| Location registry | **extract/dnd/location-registry/shape**, **extract/dnd/location-registry/source_refs**, **extract/dnd/location-registry/source_relatedness**, **normalize/dnd/location-registry/identity** |
|
||||
| Location occurrences | **extract/dnd/location-occurrences/shape**, **extract/dnd/location-occurrences/registry**, **extract/dnd/location-occurrences/source_refs**, **extract/dnd/location-occurrences/source_relatedness**, **normalize/dnd/location-occurrences/invariants** |
|
||||
|
||||
When no override is configured, production D&D bindings use the following
|
||||
ordered chains. Each row lists extract then normalize; spell chains are the
|
||||
same at both stages.
|
||||
|
||||
| Lane | Extract | Normalize |
|
||||
| --- | --- | --- |
|
||||
| Spells | generic/valid_json, extract/dnd/spells/shape, extract/dnd/spells/catalog, extract/dnd/spells/source_refs, generic/valid_json_schema, extract/dnd/spells/source_relatedness | Same as extract |
|
||||
| NPC registry | generic/valid_json, extract/dnd/npc-registry/shape, extract/dnd/npc-registry/source_refs, generic/valid_json_schema, extract/dnd/npc-registry/source_relatedness | generic/valid_json, extract/dnd/npc-registry/shape, normalize/dnd/npc-registry/identity, extract/dnd/npc-registry/source_refs, generic/valid_json_schema, extract/dnd/npc-registry/source_relatedness |
|
||||
| Combat turns | generic/valid_json, extract/dnd/combat-turns/shape, extract/dnd/combat-turns/source_refs, generic/valid_json_schema, extract/dnd/combat-turns/source_relatedness | generic/valid_json, extract/dnd/combat-turns/shape, normalize/dnd/combat-turns/invariants, extract/dnd/combat-turns/source_refs, generic/valid_json_schema, extract/dnd/combat-turns/source_relatedness |
|
||||
| Item occurrences | generic/valid_json, extract/dnd/item-occurrences/shape, extract/dnd/item-occurrences/registry, extract/dnd/item-occurrences/source_refs, generic/valid_json_schema, extract/dnd/item-occurrences/source_relatedness | generic/valid_json, extract/dnd/item-occurrences/shape, extract/dnd/item-occurrences/registry, normalize/dnd/item-occurrences/invariants, extract/dnd/item-occurrences/source_refs, generic/valid_json_schema, extract/dnd/item-occurrences/source_relatedness |
|
||||
| Item registry | generic/valid_json, extract/dnd/item-registry/shape, extract/dnd/item-registry/source_refs, generic/valid_json_schema, extract/dnd/item-registry/source_relatedness | generic/valid_json, extract/dnd/item-registry/shape, normalize/dnd/item-registry/identity, extract/dnd/item-registry/source_refs, generic/valid_json_schema, extract/dnd/item-registry/source_relatedness |
|
||||
| NPC occurrences | generic/valid_json, extract/dnd/npc-occurrences/shape, extract/dnd/npc-occurrences/registry, extract/dnd/npc-occurrences/source_refs, generic/valid_json_schema, extract/dnd/npc-occurrences/source_relatedness | generic/valid_json, extract/dnd/npc-occurrences/shape, extract/dnd/npc-occurrences/registry, normalize/dnd/npc-occurrences/invariants, extract/dnd/npc-occurrences/source_refs, generic/valid_json_schema, extract/dnd/npc-occurrences/source_relatedness |
|
||||
| Scene descriptions | generic/valid_json, extract/dnd/scene-descriptions/shape, extract/dnd/scene-descriptions/source_refs, generic/valid_json_schema, extract/dnd/scene-descriptions/source_relatedness | generic/valid_json, extract/dnd/scene-descriptions/shape, normalize/dnd/scene-descriptions/invariants, extract/dnd/scene-descriptions/source_refs, generic/valid_json_schema, extract/dnd/scene-descriptions/source_relatedness |
|
||||
| Enemy events | generic/valid_json, extract/dnd/enemy-events/shape, extract/dnd/enemy-events/engagements, extract/dnd/enemy-events/source_refs, generic/valid_json_schema, extract/dnd/enemy-events/source_relatedness | generic/valid_json, extract/dnd/enemy-events/shape, normalize/dnd/enemy-events/invariants, extract/dnd/enemy-events/source_refs, generic/valid_json_schema, extract/dnd/enemy-events/source_relatedness |
|
||||
| Location registry | generic/valid_json, extract/dnd/location-registry/shape, extract/dnd/location-registry/source_refs, generic/valid_json_schema, extract/dnd/location-registry/source_relatedness | generic/valid_json, extract/dnd/location-registry/shape, normalize/dnd/location-registry/identity, extract/dnd/location-registry/source_refs, generic/valid_json_schema, extract/dnd/location-registry/source_relatedness |
|
||||
| Location occurrences | generic/valid_json, extract/dnd/location-occurrences/shape, extract/dnd/location-occurrences/registry, extract/dnd/location-occurrences/source_refs, generic/valid_json_schema, extract/dnd/location-occurrences/source_relatedness | generic/valid_json, extract/dnd/location-occurrences/shape, extract/dnd/location-occurrences/registry, normalize/dnd/location-occurrences/invariants, extract/dnd/location-occurrences/source_refs, generic/valid_json_schema, extract/dnd/location-occurrences/source_relatedness |
|
||||
|
||||
Chains are only registered for the D&D extract and normalize modules shown
|
||||
above; select an explicit override when a different compatible chain is
|
||||
required.
|
||||
|
||||
## Validation
|
||||
|
||||
Configuration validation checks:
|
||||
Validate a file and one pipeline before running it:
|
||||
|
||||
- supported config version and known YAML fields;
|
||||
- non-empty, non-duplicated IDs after trimming;
|
||||
- supported LLM provider and non-negative profile limits;
|
||||
- positive global LLM concurrency;
|
||||
- supported diagnostics retention and non-empty work directory;
|
||||
- module binding LLM profiles refer to configured profiles.
|
||||
~~~sh
|
||||
go run ./cmd/notarius config validate \
|
||||
--config examples/dnd-minimal.config.yml \
|
||||
--pipeline dnd-session
|
||||
~~~
|
||||
|
||||
Pipeline resolution additionally checks:
|
||||
|
||||
- the pipeline ID exists;
|
||||
- at least one artifact lane is declared and selected;
|
||||
- selected lanes exist when `--only` is used;
|
||||
- required module keys are present;
|
||||
- module keys are registered for the expected slot;
|
||||
- module capability requirements are satisfied.
|
||||
Configuration validation rejects invalid YAML, unsupported fields, invalid
|
||||
defaults or environment overrides, incompatible module keys, unknown options,
|
||||
invalid reference bindings, missing required reference slots, invalid validator
|
||||
overrides, and incompatible generated artifact handoffs. Use
|
||||
[pipelines list](cli.md#pipelines-list) to inspect configured IDs.
|
||||
|
||||
77
docs/consumers/subprocess.md
Normal file
77
docs/consumers/subprocess.md
Normal file
@@ -0,0 +1,77 @@
|
||||
# Using Notarius As A Subprocess
|
||||
|
||||
Use this workflow when an orchestrator runs Notarius and consumes its published
|
||||
artifacts. The [CLI reference](../cli.md) owns invocation syntax and exit
|
||||
statuses, while the [run-result receipt](../integrations/run-result.md) and
|
||||
[Published JSON Output contract](../integrations/json-output.md) own the
|
||||
durable result formats.
|
||||
|
||||
## Run And Check The Process
|
||||
|
||||
Optionally preflight a selected configuration and pipeline before work starts:
|
||||
|
||||
```sh
|
||||
notarius config validate --config /path/to/notarius.yml --pipeline pipeline-id
|
||||
```
|
||||
|
||||
Invoke the run with explicit paths and machine-readable output. Capture
|
||||
standard output and standard error separately; do not combine them before
|
||||
processing the result.
|
||||
|
||||
```sh
|
||||
notarius run pipeline-id \
|
||||
--config /path/to/notarius.yml \
|
||||
--input /path/to/source.json \
|
||||
--output-dir /path/to/output-root \
|
||||
--json
|
||||
```
|
||||
|
||||
Use absolute paths for supplied input, configuration, output-root, and
|
||||
reference files. Notarius generates a stable prompt session for the resolved
|
||||
input module and exact input bytes. Pass **--session-id** only when intentionally
|
||||
grouping different invocations under a different session. Supply credentials
|
||||
through Notarius's documented configuration and environment mechanisms, never
|
||||
as command-line arguments or generated secret-bearing configuration. In
|
||||
particular, a session identifier is provider-visible and is not a credential
|
||||
mechanism.
|
||||
|
||||
Wait for the process before interpreting standard output. Only an exit status
|
||||
of 0 permits decoding the receipt. On a nonzero exit, retain standard error for
|
||||
diagnosis and ignore all standard-output bytes: a failed receipt write may have
|
||||
left a partial document.
|
||||
|
||||
## Discover Required Artifacts
|
||||
|
||||
Decode the successful receipt and accept the schema versions supported by the
|
||||
caller. Use its `output_directory` as the bundle root. For the production JSON
|
||||
output, resolve `index_file` under that root with a confinement check and reject
|
||||
an absolute path or a result that escapes the root.
|
||||
|
||||
Read the resulting `index.json` and locate each artifact by `lane_id`, not by a
|
||||
guessed filename. Before decoding a selected payload, verify its descriptor's
|
||||
media type and schema identity against the relevant published artifact
|
||||
contract. The JSON bundle contract links to the available lane contracts.
|
||||
|
||||
If `index.json` has an `evidence_context` descriptor, treat it as a
|
||||
pipeline-wide artifact rather than a lane entry. Verify its six descriptor
|
||||
fields before decoding the linked file according to the [Published Evidence
|
||||
Context contract](../integrations/evidence-context.md). Use each
|
||||
`evidence_refs` entry as the citation to source material. Its surrounding
|
||||
context range and included units explain the citation, but do not widen or
|
||||
replace the cited source reference.
|
||||
|
||||
A zero exit status may still report rejected outputs, warnings, or absent
|
||||
lanes. The caller decides which lane IDs are required for its own work and
|
||||
which are optional; it should make that decision explicitly rather than infer
|
||||
failure from the receipt counts alone.
|
||||
|
||||
## Preserve Provenance And Handle Data Carefully
|
||||
|
||||
Keep the receipt with the published `manifest.json`, and retain
|
||||
`rejected.json` and `warnings.json` when review or later provenance requires
|
||||
them. Treat the input, output bundle, cache, debug bundle, and captured process
|
||||
logs as potentially sensitive data. Apply the caller's access controls and
|
||||
retention policy, and avoid copying secrets into arguments, logs, or
|
||||
provenance records. An evidence-context artifact contains source-unit text and
|
||||
metadata, and selected lanes can cover most of an input; preserve and share it
|
||||
only when that source content is authorized for the recipient.
|
||||
43
docs/development.md
Normal file
43
docs/development.md
Normal file
@@ -0,0 +1,43 @@
|
||||
# Development
|
||||
|
||||
This is the first-read landing page for people and LLM coding agents working on
|
||||
Notarius. It provides a concise repository orientation and routes each kind of
|
||||
change to its canonical documentation.
|
||||
|
||||
Notarius is a Go CLI for configured structured extraction workflows. Start with
|
||||
the [README](../README.md) for product context, [Architecture](policy/architecture.md)
|
||||
for system boundaries, and [Internal Overview](internal/overview.md) for the
|
||||
implemented component map.
|
||||
|
||||
## What To Read
|
||||
|
||||
| When working on | Read | Why |
|
||||
| --- | --- | --- |
|
||||
| Finding the package or component that owns current behavior | [Internal Overview](internal/overview.md) | It is the implemented component inventory and routes to focused internals. |
|
||||
| Application shape, package boundaries, contracts, dependency direction, runtime guarantees, or safety properties | [Architecture](policy/architecture.md) and relevant [ADRs](adr/) | Architecture defines the intended system and its invariants; ADRs preserve significant decision rationale. |
|
||||
| Any documentation addition or revision | [Documentation Policy](policy/documentation.md) | It defines canonical homes, audiences, current-behavior rules, and maintenance requirements. |
|
||||
| Adding, changing, reviewing, or deleting tests | [Testing Policy](policy/testing.md) | It defines risk-based sufficiency, durable test boundaries, test-double guidance, and criteria for retaining tests. |
|
||||
| CLI composition or command behavior | [CLI Internals](internal/cli.md) and [CLI Reference](cli.md) | The internal guide owns composition and command flow; the reference owns public syntax. |
|
||||
| Building a subprocess caller or changing its result protocol | [Subprocess Consumer Guide](consumers/subprocess.md), [Run Result Receipt](integrations/run-result.md), and [CLI Internals](internal/cli.md) | These separate caller workflow, durable receipt contract, and CLI implementation behavior. |
|
||||
| Configuration loading, resolution, or user-visible configuration behavior | [Configuration Internals](internal/configuration.md) and [Configuration](config.md) | The internal guide owns loading and resolution mechanics; the reference owns the configuration contract. |
|
||||
| Pipeline resolution or execution | [Pipeline Internals](internal/pipeline.md) | It documents profiles, references, validation, retries, checkpoints, and runner behavior. |
|
||||
| Production modules or validators | [Module Internals](internal/modules.md), [D&D Module Internals](internal/dnd.md), and [D&D integration contracts](integrations/) | The generic guide owns extension mechanics, the D&D guide owns shared family conventions, and the contracts own durable output shapes. |
|
||||
| LLM clients, prompts, schemas, profiles, or scheduling | [LLM Runtime](internal/llm.md) | It documents the transport boundary and PromptKit integration. |
|
||||
| Output, cache, resume, or debug artifacts | [Run State Internals](internal/state.md), [Operations](operations.md), and [Configuration](config.md) | These separate implementation details, operator behavior, and configuration contracts. |
|
||||
| External input formats, artifact schemas, or durable output files | [Integration Contracts](integrations/) | Integration documents define external and durable data contracts. |
|
||||
| Proposed or unimplemented behavior | [Roadmap](roadmap/) | Future work belongs only in roadmap documentation until implemented. |
|
||||
|
||||
For an existing subsystem, also inspect its focused tests and the package-local
|
||||
types and contracts before changing behavior.
|
||||
|
||||
## Validation
|
||||
|
||||
Use focused package tests while iterating. Run the repository-wide checks when
|
||||
a change affects shared contracts, application behavior, or maintained
|
||||
documentation examples:
|
||||
|
||||
```sh
|
||||
go test ./...
|
||||
go vet ./...
|
||||
go build ./cmd/notarius
|
||||
```
|
||||
81
docs/integrations/chunk-map.md
Normal file
81
docs/integrations/chunk-map.md
Normal file
@@ -0,0 +1,81 @@
|
||||
# Accepted Chunk Map
|
||||
|
||||
This document defines the optional durable `chunk-map.json` artifact in a
|
||||
[published JSON bundle](json-output.md). It describes the accepted,
|
||||
materialized chunk plan used by one run. It is not a lane payload and is never
|
||||
an input to a later pipeline step.
|
||||
|
||||
## Contract Identity
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `source/chunk-map` |
|
||||
| Logical file | `chunk-map.json` |
|
||||
| Media type | `application/json` |
|
||||
| Schema ID | `notarius.source.chunk_map` |
|
||||
| Schema name | `notarius_source_chunk_map_v1` |
|
||||
| Schema version | `v1` |
|
||||
|
||||
The optional `chunk_map` descriptor in `index.json` identifies this artifact.
|
||||
Export is controlled by the JSON output binding described in
|
||||
[Configuration](../config.md#module-bindings-and-validators).
|
||||
|
||||
## Wire Shape
|
||||
|
||||
Every payload has these required fields:
|
||||
|
||||
| Field | Meaning |
|
||||
| --- | --- |
|
||||
| `source_id` | Accepted source-document identity. |
|
||||
| `source_digest` | Lower-case `sha256:` digest of that source document. |
|
||||
| `plan_digest` | Lower-case `sha256:` digest of the logical chunk plan. |
|
||||
| `requested_chunker` | Chunk module selected by the resolved pipeline. |
|
||||
| `producer` | Original accepted-plan producer. `input_module` and `chunk_module` are required; `llm_profile` is optional. |
|
||||
| `plan_annotations` | Plan-level annotation namespace map; `{}` when none are present. |
|
||||
| `chunks` | Non-empty execution-order chunk collection. |
|
||||
|
||||
Each `chunks` entry contains non-empty `id`, zero-based `index`, `source_ref`,
|
||||
positive `unit_count`, and an explicit `annotations` map. `source_ref` contains
|
||||
the same `source_id` as the top-level value plus positive inclusive
|
||||
`start_unit_id` and `end_unit_id` values. Endpoints identify source units; their
|
||||
numeric values do not by themselves establish source-document order.
|
||||
|
||||
Annotation namespaces are non-empty trimmed strings. Their values are arbitrary
|
||||
valid JSON and are retained without interpreting a module-specific namespace.
|
||||
|
||||
## Ordering And Validation
|
||||
|
||||
`chunks` are in execution order. Their indexes are contiguous, start at zero,
|
||||
and equal their array positions; chunk IDs are unique. The emitted map is built
|
||||
only after the selected plan has been accepted and materialized against the
|
||||
source document, so its ranges, unit counts, annotations, and digests describe
|
||||
that exact plan.
|
||||
|
||||
The codec rejects malformed JSON, trailing content, unknown fixed-object
|
||||
fields, invalid identities or digests, invalid annotations, duplicate chunk
|
||||
IDs, non-contiguous indexes, and a `plan_digest` that does not match the
|
||||
reconstructed logical plan. The checked-in
|
||||
[schema](../../internal/framework/chunkmap/assets/schemas/source_chunk_map.v1.json)
|
||||
defines the strict JSON shape.
|
||||
|
||||
## Valid Example
|
||||
|
||||
The compact
|
||||
[source chunk-map fixture](../../internal/framework/chunkmap/testdata/source_chunk_map.v1.json)
|
||||
is decoded by the production codec and demonstrates an accepted map with
|
||||
annotations, producer identity, and ordered chunks.
|
||||
|
||||
## Publication And Compatibility
|
||||
|
||||
The map is present only when a chunk plan was accepted and its export is
|
||||
enabled. It remains publishable if a later lane is rejected, but is absent when
|
||||
chunk-plan validation rejects the plan. `requested_chunker` identifies the
|
||||
current pipeline selection, while `producer` identifies the component that
|
||||
originally produced the accepted plan; they may differ when an accepted plan is
|
||||
reused.
|
||||
|
||||
The map contains structure rather than source content: it excludes transcript
|
||||
bytes, source-unit metadata, chunk text, private model output, reference
|
||||
content, debug data, and filesystem paths. Treat the exported map with the
|
||||
same care as other published output. Publication location and retention are
|
||||
defined in [Operations](../operations.md#output-bundles).
|
||||
71
docs/integrations/dnd-combat-turn-artifacts.md
Normal file
71
docs/integrations/dnd-combat-turn-artifacts.md
Normal file
@@ -0,0 +1,71 @@
|
||||
# D&D Combat-Turn Artifact
|
||||
|
||||
This contract defines the durable combat-action occurrence list produced by
|
||||
`dnd/combat-turns`. It records source-grounded turns and actions; it is not a
|
||||
complete initiative tracker, combat summary, or state model.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/combat-turn-list` |
|
||||
| Schema ID | `notarius.dnd.combat_turns` |
|
||||
| Schema name | `notarius_dnd_combat_turns_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
|
||||
`v1` is a strict JSON object with required `combat_turns`; the array may be
|
||||
empty. Turn and source-reference objects reject unknown fields. An incompatible
|
||||
shape change requires a new schema version.
|
||||
|
||||
## Wire shape
|
||||
|
||||
Each combat turn has these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `actor` | Non-empty acting character or creature name. |
|
||||
| `turn_kind` | `turn`, `reaction`, `legendary_action`, `lair_action`, or `other`. |
|
||||
| `source_refs` | One or more transcript evidence ranges. |
|
||||
|
||||
Each source reference has exactly `source_id`, `start_unit_id`, and
|
||||
`end_unit_id`. It identifies an inclusive current-transcript range; unit IDs
|
||||
are positive and the start may not follow the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"combat_turns": [
|
||||
{
|
||||
"actor": "Mira Thorn",
|
||||
"turn_kind": "turn",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 31, "end_unit_id": 32}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Eligibility, evidence, and normalized form
|
||||
|
||||
The extractor requires an approved [scene-description artifact](dnd-scene-description-artifacts.md).
|
||||
It emits combat turns only for a chunk with an exact matching scene classified
|
||||
`combat`; an exact non-combat scene produces an accepted empty list. The scene
|
||||
record controls eligibility only: its title, summary, and reference do not
|
||||
become turn evidence. No exact matching scene also produces an empty list and
|
||||
the `scene_classification_unavailable` warning.
|
||||
|
||||
An optional normalized [NPC registry artifact](dnd-npc-registry-artifacts.md) can ground an
|
||||
actor name. Its registry references are provenance, never combat evidence.
|
||||
Normalization trims and, where possible, canonicalizes actor names; orders and
|
||||
deduplicates exact source references; orders valid-evidence turns by source
|
||||
chronology; and collapses only duplicates with the same actor identity, turn
|
||||
kind, and complete valid evidence. It does not infer turns, initiative, or
|
||||
actions from registry or scene data.
|
||||
|
||||
The [NPC-occurrence artifact](dnd-npc-occurrence-artifacts.md) records
|
||||
broader NPC occurrences. The [enemy-event artifact](dnd-enemy-event-artifacts.md)
|
||||
uses combat turns as grounding only; turns do not establish an enemy event or
|
||||
its outcome. The [JSON output contract](json-output.md) defines publication,
|
||||
and [D&D module internals](../internal/dnd.md) describes routing and validation
|
||||
mechanics.
|
||||
114
docs/integrations/dnd-enemy-event-artifacts.md
Normal file
114
docs/integrations/dnd-enemy-event-artifacts.md
Normal file
@@ -0,0 +1,114 @@
|
||||
# D&D Enemy-Event Artifact
|
||||
|
||||
This contract defines the durable, source-grounded enemy-event occurrence list.
|
||||
It records enemies directly established as opposing the party and explicitly
|
||||
observed combat outcomes. It is an ordered observation artifact from which a
|
||||
consumer may derive a ledger; it is not a ledger, encounter roster, or terminal
|
||||
state model.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/enemy-event-list` |
|
||||
| Schema ID | `notarius.dnd.enemy_events` |
|
||||
| Schema name | `notarius_dnd_enemy_events_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
|
||||
`v1` is a strict JSON object with required `events`; the array may be empty.
|
||||
Event and source-reference objects reject unknown fields. An incompatible shape
|
||||
change requires a new schema version.
|
||||
|
||||
## Wire shape
|
||||
|
||||
Every event has these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `name` | Non-empty display name or directly grounded collective subject label. |
|
||||
| `kind` | `engaged`, `killed`, `fled`, `captured`, or `incapacitated`. |
|
||||
| `source_refs` | One or more current-transcript evidence ranges. |
|
||||
|
||||
Each source reference has exactly `source_id`, `start_unit_id`, and
|
||||
`end_unit_id`. It identifies an inclusive current-transcript range; unit IDs
|
||||
are positive and the start may not follow the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"events": [
|
||||
{
|
||||
"name": "Ashfang",
|
||||
"kind": "engaged",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 41, "end_unit_id": 42}
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "Ashfang",
|
||||
"kind": "fled",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 57, "end_unit_id": 58}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Event semantics and evidence
|
||||
|
||||
| Kind | Required evidence |
|
||||
| --- | --- |
|
||||
| `engaged` | The subject is directly established as actively opposing the party in combat. At most one engagement is emitted for one subject in one combat scene. |
|
||||
| `killed` | The transcript explicitly establishes that the subject died or was killed. Damage, defeat, disappearance, or combat ending is insufficient. |
|
||||
| `fled` | The subject explicitly escapes, retreats, or otherwise leaves combat to avoid continued engagement. Movement or absence from later turns is insufficient. |
|
||||
| `captured` | The subject is explicitly taken prisoner or secured under the party's control. A grapple or temporary restraint alone is insufficient. |
|
||||
| `incapacitated` | The subject is explicitly rendered unable to continue acting without being established as killed or captured. A missed turn is insufficient. |
|
||||
|
||||
The current transcript is the only event evidence. Campaign context and
|
||||
normalized NPC, scene-description, combat-turn, and NPC-occurrence artifacts
|
||||
can ground names or control combat eligibility, but none may supply event
|
||||
evidence. An outcome may share evidence with an engagement, in which case both
|
||||
events are retained.
|
||||
|
||||
Extraction is limited to chunks with an exact combat-scene classification. An
|
||||
exact non-combat classification produces an accepted empty list. Missing or
|
||||
mismatched classification also produces an accepted empty list and a
|
||||
`scene_classification_unavailable` warning.
|
||||
|
||||
## Subjects, normalization, and order
|
||||
|
||||
A subject matching the normalized NPC registry uses that registry's canonical
|
||||
display name. Unmatched hostile creatures, summoned entities, and directly
|
||||
grounded groups remain valid subjects. An unnamed homogeneous group uses the
|
||||
narrowest transcript-grounded label, such as `Orcs`, `One orc`, or `Remaining
|
||||
orcs`; the artifact never invents synthetic member identities or quantities.
|
||||
Party members, allies, neutral observers, mentioned-but-absent enemies, hazards,
|
||||
traps, and environmental effects are excluded.
|
||||
|
||||
Normalization collapses surrounding and repeated internal whitespace in subject
|
||||
display values, canonicalizes recognized registry names, canonicalizes and
|
||||
deduplicates exact source ranges, then orders events by valid evidence
|
||||
chronology, normalized subject identity, display name, kind, and reference
|
||||
sequence. The deterministic kind tie order is `engaged`,
|
||||
`incapacitated`, `captured`, `fled`, then `killed`. Only entries with the same
|
||||
normalized name, kind, and complete canonical evidence sequence are collapsed.
|
||||
Different kinds, evidence, repeated engagement in separate scenes, and later
|
||||
outcomes remain separate. A later engagement for the same named subject is
|
||||
preserved after an earlier outcome because the artifact does not assert an
|
||||
irreversible state transition.
|
||||
|
||||
## Non-goals
|
||||
|
||||
The artifact has no NPC or scene ID, quantity, confidence, description,
|
||||
rationale, summary, current state, or inferred terminal outcome. It does not
|
||||
emit `active` or `unresolved`; consumers may derive an unresolved ledger view
|
||||
only when an engagement has no later explicit outcome. It never infers an
|
||||
outcome from turn absence, scene termination, initiative order, hit-point
|
||||
guesses, or other artifacts.
|
||||
|
||||
The [JSON output contract](json-output.md) defines publication. Configuration
|
||||
keys, required generated-reference slots, and validator-chain selection are
|
||||
defined in the [configuration reference](../config.md). Implementation and
|
||||
prompt-grounding mechanics are described in the
|
||||
[D&D module internals](../internal/dnd.md).
|
||||
72
docs/integrations/dnd-item-occurrence-artifacts.md
Normal file
72
docs/integrations/dnd-item-occurrence-artifacts.md
Normal file
@@ -0,0 +1,72 @@
|
||||
# D&D Item-Occurrence Artifact
|
||||
|
||||
`dnd/item-occurrences` currently produces this source-grounded item and currency
|
||||
occurrence list. It records discoveries and possession changes, not an
|
||||
inventory, balance, or ledger.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/item-occurrence-list` |
|
||||
| Schema ID | `notarius.dnd.item_occurrences` |
|
||||
| Schema name | `notarius_dnd_item_occurrences_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
|
||||
`v1` accepts one strict JSON object with required `occurrences`; the array may
|
||||
be empty. Each occurrence has required `item_id`, `name`, `kind`, and
|
||||
`source_refs`, and occurrence and source-reference objects reject unknown
|
||||
fields. `quantity`, `from`, and `to` appear only when their kind permits them.
|
||||
An incompatible shape change requires a new schema version.
|
||||
|
||||
## Registry grounding
|
||||
|
||||
Both extraction and normalization require an `item_registry` reference bound to
|
||||
an earlier normalized `dnd/item-registry` artifact. The registry is immutable
|
||||
for an operation and contributes names-only grounding after the shared evidence
|
||||
message. Notarius resolves the model's selected name into the unchanged exact
|
||||
durable ID/name pair. It is never occurrence evidence.
|
||||
|
||||
Each occurrence must use one exact registry ID/name pair. An extraction response
|
||||
with an unknown or ambiguous selected name is rejected as invalid model output;
|
||||
the configured pipeline may retry it and never accepts a partial artifact.
|
||||
Normalization and validation remain defense in depth for artifacts entering
|
||||
through other boundaries: normalization canonicalizes a recognized name by ID,
|
||||
preserves unknown values for the registry validator, and the registry validator
|
||||
rejects unknown or mismatched pairs.
|
||||
|
||||
## Wire shape
|
||||
|
||||
Each source reference has exactly `source_id`, `start_unit_id`, and
|
||||
`end_unit_id`. It identifies an inclusive range in the current transcript;
|
||||
unit IDs are positive and the start may not follow the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"occurrences": [
|
||||
{
|
||||
"item_id": "item:sha256:…",
|
||||
"name": "Silver Pieces",
|
||||
"kind": "acquired",
|
||||
"quantity": 20,
|
||||
"to": "party",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 2, "end_unit_id": 2}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The five kinds remain `discovered`, `acquired`, `lost`, `consumed`, and
|
||||
`transferred`. Holder, quantity, currency, ordering, and exact-duplicate rules
|
||||
are unchanged: discovered has no holder; acquired requires `to`; lost and
|
||||
consumed require `from`; transferred requires distinct non-`party` holders.
|
||||
The only current downstream compatibility requirement is its registry handoff;
|
||||
the normalized occurrence list is otherwise published for callers. See
|
||||
[Configuration](../config.md#d-d-reference-slots) for the binding and
|
||||
[JSON output](json-output.md) for publication.
|
||||
|
||||
See [item registry](dnd-item-registry-artifacts.md) for the grounding artifact
|
||||
and [D&D module internals](../internal/dnd.md) for implementation details.
|
||||
95
docs/integrations/dnd-item-registry-artifacts.md
Normal file
95
docs/integrations/dnd-item-registry-artifacts.md
Normal file
@@ -0,0 +1,95 @@
|
||||
# D&D Item Registry Artifact
|
||||
|
||||
This contract defines the durable, source-grounded item registry produced by
|
||||
`dnd/item-registry`. It records transcript-established item types and unique
|
||||
designations for one source document; it is not an inventory, holder record,
|
||||
quantity ledger, or item-occurrence artifact.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/item-registry` |
|
||||
| Schema ID | `notarius.dnd.item_registry` |
|
||||
| Schema name | `notarius_dnd_item_registry_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
| Identity policy | `dnd.item_registry.identity.v1` |
|
||||
|
||||
`v1` accepts one strict JSON object with required `items`; the array may be
|
||||
empty. Item and source-reference objects reject unknown fields. An incompatible
|
||||
artifact shape or identity-policy change uses a new version or policy.
|
||||
|
||||
## Wire shape and identity
|
||||
|
||||
Each item has these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `id` | `item:sha256:` followed by 64 lowercase hexadecimal characters. |
|
||||
| `name` | Non-empty transcript-established item type or unique designation. |
|
||||
| `source_refs` | One or more transcript evidence ranges that establish the item. |
|
||||
|
||||
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
|
||||
The source ID identifies the transcript, unit IDs are positive inclusive unit
|
||||
identifiers, and the start may not follow the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"items": [
|
||||
{
|
||||
"id": "item:sha256:31e73b6280ef98e4d8070e07fd4de9b2c3e842cc03af1a09ca631cb95b73e3b3",
|
||||
"name": "Star Compass",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 4, "end_unit_id": 5}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The ID is deterministic for an item name or type, rather than for one physical
|
||||
instance. Notarius normalizes the display name for comparison with Unicode
|
||||
NFKC, supported apostrophe normalization, collapsed whitespace, and case
|
||||
folding. It hashes compact JSON for this array:
|
||||
|
||||
```text
|
||||
["dnd.item_registry.identity.v1", comparison_name]
|
||||
```
|
||||
|
||||
The canonical ID is the lowercase SHA-256 digest of those bytes with the
|
||||
`item:sha256:` prefix. Equal comparison names represent one item identity;
|
||||
normalization unions their transcript evidence when it safely consolidates a
|
||||
candidate group.
|
||||
|
||||
## Scope, reconciliation, and evidence
|
||||
|
||||
The registry includes named unique items, concrete reusable item types, stable
|
||||
unique designations, and separately established currency denominations. It
|
||||
excludes vague loot or treasure, generic weapons, quantities, inferred
|
||||
properties, and inferred uniqueness. Capitalization alone does not establish
|
||||
eligibility.
|
||||
|
||||
Normalization first applies deterministic display, evidence, and ID rules. It
|
||||
then may use a bounded LLM-assisted proposal to reconcile semantically duplicate
|
||||
records. The proposal may choose only a supplied candidate display name;
|
||||
invalid, uncertain, overlapping, or unsafe proposals retain the deterministic
|
||||
result with retry or fallback diagnostics. A proposal that mixes a recognized
|
||||
currency denomination with a non-currency item, or combines recognized
|
||||
denominations, is unsafe and retains every deterministic record. Currency
|
||||
denominations, materially different item types, and merely nearby objects
|
||||
remain distinct. Source references establish registry provenance, not evidence
|
||||
for later artifacts.
|
||||
|
||||
## Consumers and publication
|
||||
|
||||
`dnd/item-occurrences` requires one approved item registry through its
|
||||
`item_registry` reference slot for both extraction and normalization. Its
|
||||
consumer receives names-only grounding; Notarius resolves the selected name
|
||||
into the unchanged exact durable ID/name pair. The registry’s source references
|
||||
are never occurrence evidence. Unknown or ambiguous selections are rejected by
|
||||
the occurrence contract. See the
|
||||
[item-occurrence artifact](dnd-item-occurrence-artifacts.md) for that strict
|
||||
wire contract, [Configuration](../config.md#d-d-reference-slots) for binding
|
||||
rules and validator selection, and the [JSON output contract](json-output.md)
|
||||
for publication.
|
||||
86
docs/integrations/dnd-location-occurrence-artifacts.md
Normal file
86
docs/integrations/dnd-location-occurrence-artifacts.md
Normal file
@@ -0,0 +1,86 @@
|
||||
# D&D Location-Occurrence Artifact
|
||||
|
||||
This contract defines the durable occurrence list produced by
|
||||
`dnd/location-occurrences`. It records source-grounded ways the party relates
|
||||
to locations in a required normalized location registry; it does not extend
|
||||
that registry or infer a place absent from it.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/location-occurrence-list` |
|
||||
| Schema ID | `notarius.dnd.location_occurrences` |
|
||||
| Schema name | `notarius_dnd_location_occurrences_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
|
||||
`v1` accepts one strict JSON object with required `occurrences`; the array may
|
||||
be empty. Occurrence and source-reference objects reject unknown fields. An
|
||||
incompatible shape change requires a new schema version.
|
||||
|
||||
## Wire shape
|
||||
|
||||
Each occurrence has these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `location_id` | Exact ID from the required normalized [location registry](dnd-location-registry-artifacts.md). |
|
||||
| `name` | Exact canonical display name for `location_id` in that registry. |
|
||||
| `kind` | One of `visited`, `planned`, `recalled`, or `mentioned`. |
|
||||
| `source_refs` | One or more current-transcript evidence ranges for this occurrence. |
|
||||
|
||||
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
|
||||
It identifies an inclusive range in the current transcript; unit IDs are
|
||||
positive and the start may not follow the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"occurrences": [
|
||||
{
|
||||
"location_id": "location:sha256:fb05475da0fc7debf994b517e1906ffe7209887a6a1ec306356d84de820b1a24",
|
||||
"name": "Moon Gate",
|
||||
"kind": "visited",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 12, "end_unit_id": 13}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Occurrence categories
|
||||
|
||||
| Kind | Meaning |
|
||||
| --- | --- |
|
||||
| `visited` | The transcript establishes physical party presence, including arrival, continuing presence, or departure. |
|
||||
| `planned` | The party explicitly proposes, intends, or agrees to future travel; speculation alone is not enough. |
|
||||
| `recalled` | The transcript explicitly recounts prior party presence before the current live events. |
|
||||
| `mentioned` | The location is explicit but no stronger category applies, including lore, directions, third-party activity, non-actionable speculation, a mere hypothetical reference, or out-of-character discussion. |
|
||||
|
||||
For overlapping evidence, precedence is `visited`, then `planned`, then
|
||||
`recalled`, then `mentioned`. For example, “What if we went to Moon Gate?” is
|
||||
eligible as `mentioned` when its narrow evidence explicitly references that
|
||||
registry location, but it is not `planned` without an actual proposal,
|
||||
intention, or agreement to travel. Inferred, unstated, uncertain, and
|
||||
unsupported places or occurrences are omitted. Normalization
|
||||
canonicalizes the registry name, orders and deduplicates source references, and
|
||||
orders occurrences by source chronology, location ID, name, kind, and reference
|
||||
sequence. It collapses only exact duplicates with the same ID, kind, and
|
||||
complete canonical evidence sequence.
|
||||
|
||||
## Required grounding and evidence
|
||||
|
||||
Both extraction and normalization require exactly one `location_registry` reference of
|
||||
kind `dnd/location-registry`, media type `application/json`, and at most 1 MiB. The
|
||||
registry provides identity grounding only. The model selects a supplied
|
||||
contextual name-and-registry-reference descriptor, and Notarius resolves it
|
||||
into the exact durable ID/name pair. Unknown, partial, or ambiguous selections
|
||||
are rejected rather than guessed or reassigned. The current transcript is the
|
||||
only evidence source for an occurrence; registry evidence and provenance never
|
||||
become occurrence evidence.
|
||||
|
||||
See [Configuration](../config.md#d-d-reference-slots) for the selectable slot
|
||||
and generated-handoff compatibility, [D&D module internals](../internal/dnd.md)
|
||||
for implementation behavior, and the [JSON output contract](json-output.md)
|
||||
for publication.
|
||||
93
docs/integrations/dnd-location-registry-artifacts.md
Normal file
93
docs/integrations/dnd-location-registry-artifacts.md
Normal file
@@ -0,0 +1,93 @@
|
||||
# D&D Location Registry Artifact
|
||||
|
||||
This contract defines the durable, source-grounded location registry produced
|
||||
by `dnd/location-registry`. It records transcript-established physical places for one
|
||||
source document; it is not a map, location hierarchy, campaign-wide world
|
||||
registry, or location description.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/location-registry` |
|
||||
| Schema ID | `notarius.dnd.location_registry` |
|
||||
| Schema name | `notarius_dnd_location_registry_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
| Identity policy | `dnd.location_registry.identity.v1` |
|
||||
|
||||
`v1` accepts one strict JSON object with required `locations`; the array may be
|
||||
empty. Location and source-reference objects reject unknown fields. An
|
||||
incompatible artifact shape or identity-policy change uses a new version or
|
||||
policy.
|
||||
|
||||
## Wire shape and identity
|
||||
|
||||
Each location has these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `id` | `location:sha256:` followed by 64 lowercase hexadecimal characters. |
|
||||
| `name` | Non-empty transcript-established display name. |
|
||||
| `source_refs` | One or more transcript evidence ranges that identify the place. |
|
||||
|
||||
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
|
||||
The source ID identifies the transcript, unit IDs are positive inclusive unit
|
||||
identifiers, and the start may not follow the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"locations": [
|
||||
{
|
||||
"id": "location:sha256:fb05475da0fc7debf994b517e1906ffe7209887a6a1ec306356d84de820b1a24",
|
||||
"name": "Moon Gate",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 4, "end_unit_id": 5}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The ID is deterministic and scoped to the source document. Notarius normalizes
|
||||
the display name for comparison with Unicode NFKC, supported apostrophe
|
||||
normalization, collapsed whitespace, and case folding. It hashes compact JSON
|
||||
for this array, using the earliest canonical source reference as the anchor:
|
||||
|
||||
```text
|
||||
["dnd.location_registry.identity.v1", comparison_name, source_id, start_unit_id, end_unit_id]
|
||||
```
|
||||
|
||||
The canonical ID is the lowercase SHA-256 digest of those bytes with the
|
||||
`location:sha256:` prefix. Equal display names are allowed when their evidence
|
||||
anchors differ, so a generic name does not force distinct places to collapse.
|
||||
|
||||
## Scope, reconciliation, and evidence
|
||||
|
||||
Locations are physical or spatial places established by the transcript with a
|
||||
stable proper name or unique in-world designation, such as named planes,
|
||||
regions, settlements, districts, buildings, rooms, landmarks, routes, and
|
||||
geographic features. Generic, temporary, relative, and descriptive phrases
|
||||
such as “the room,” “the bar,” “the hallway,” “outside,” and “upstairs” are not
|
||||
registry locations. Capitalization alone does not establish eligibility.
|
||||
Notarius does not infer an unstated place or add hierarchy, coordinates,
|
||||
descriptions, participants, or ownership.
|
||||
|
||||
Normalization first applies deterministic display, evidence, and ID rules. It
|
||||
then may use a bounded LLM-assisted proposal to reconcile semantically duplicate
|
||||
records. The proposal is validated and applied conservatively; invalid or
|
||||
unusable proposals retain the deterministic result with retry or fallback
|
||||
diagnostics. The registry's source references establish registry provenance,
|
||||
not evidence for later artifacts.
|
||||
|
||||
## Consumers and publication
|
||||
|
||||
`dnd/location-occurrences` requires one approved location registry through its
|
||||
`location_registry` reference slot. Its prompt receives contextual selectors
|
||||
containing a canonical name and registry references; Notarius resolves a
|
||||
selection into the unchanged exact durable ID/name pair. Registry references
|
||||
must not be treated as occurrence evidence. See the
|
||||
[location-occurrence artifact](dnd-location-occurrence-artifacts.md)
|
||||
for that contract, [Configuration](../config.md#references-and-ordered-handoffs)
|
||||
for binding rules, and the [JSON output contract](json-output.md) for
|
||||
publication.
|
||||
90
docs/integrations/dnd-npc-occurrence-artifacts.md
Normal file
90
docs/integrations/dnd-npc-occurrence-artifacts.md
Normal file
@@ -0,0 +1,90 @@
|
||||
# D&D NPC Occurrence Artifact
|
||||
|
||||
This contract defines the durable occurrence list produced by
|
||||
`dnd/npc-occurrences`. It records discrete, source-grounded occurrences with
|
||||
NPCs already present in a normalized registry; it does not extend that registry
|
||||
or summarize the session.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/npc-occurrence-list` |
|
||||
| Schema ID | `notarius.dnd.npc_occurrences` |
|
||||
| Schema name | `notarius_dnd_npc_occurrences_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
|
||||
`v1` is a strict JSON object with required `occurrences`; the array may be
|
||||
empty. Occurrence and source-reference objects reject unknown fields. An
|
||||
incompatible shape change requires a new schema version.
|
||||
|
||||
## Wire shape
|
||||
|
||||
Each occurrence has these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `npc_id` | Exact durable ID from the required NPC registry. |
|
||||
| `name` | Non-empty canonical display name from the required NPC registry. |
|
||||
| `kind` | One of the occurrence categories below. |
|
||||
| `source_refs` | One or more transcript evidence ranges. |
|
||||
|
||||
Each source reference has exactly `source_id`, `start_unit_id`, and
|
||||
`end_unit_id`. It identifies an inclusive range in the current transcript;
|
||||
unit IDs are positive and the start may not follow the end. Extraction evidence
|
||||
for an occurrence is confined to its accepted chunk.
|
||||
|
||||
```json
|
||||
{
|
||||
"occurrences": [
|
||||
{
|
||||
"npc_id": "npc:sha256:example",
|
||||
"name": "Mira Thorn",
|
||||
"kind": "dialogue",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 12, "end_unit_id": 13}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Occurrence categories
|
||||
|
||||
| Kind | Meaning |
|
||||
| --- | --- |
|
||||
| `mentioned` | The NPC is referred to but is not established as present or communicating. |
|
||||
| `noncombat_presence` | The NPC is present and relevant without meaningful dialogue or combat participation. |
|
||||
| `dialogue` | The NPC speaks, responds, or meaningfully participates in a non-combat exchange. |
|
||||
| `combat_ally` | The NPC actively participates in combat on the party's side. |
|
||||
| `combat_opponent` | The NPC actively participates in combat against the party. |
|
||||
| `other` | A clearly evidenced direct occurrence not covered by another category. |
|
||||
|
||||
The categories do not represent motives, relationships, state, or events that
|
||||
the cited transcript does not establish. An `other` entry is not a substitute
|
||||
for uncertain classification.
|
||||
|
||||
## Identity, evidence, and order
|
||||
|
||||
The required normalized [NPC registry artifact](dnd-npc-registry-artifacts.md)
|
||||
supplies names-only contextual grounding to the model. Notarius resolves the
|
||||
selected name and writes the exact `{npc_id, name}` pair. An unknown or
|
||||
ambiguous selection rejects the complete model result; normalization does not
|
||||
repair names by similarity. Registry references are provenance only and never
|
||||
replace an occurrence's own evidence.
|
||||
The registry may include an identity established by a factual third-party
|
||||
mention; that provenance alone does not create a `mentioned` occurrence. Each
|
||||
occurrence remains a separately cited fact in the current transcript.
|
||||
Normalization validates the exact pair, orders and
|
||||
deduplicates exact source references, then orders occurrences by valid source
|
||||
chronology, NPC comparison identity, display name, kind, and reference sequence.
|
||||
Only entries with the same NPC ID, canonical name, kind, and complete valid evidence
|
||||
sequence are collapsed; distinct categories or evidence remain separate.
|
||||
|
||||
See the [combat-turn artifact](dnd-combat-turn-artifacts.md) for combat-action
|
||||
occurrences. The [enemy-event artifact](dnd-enemy-event-artifacts.md) consumes
|
||||
only `combat_opponent` occurrences as grounding; they never establish an enemy
|
||||
event or outcome. The [JSON output contract](json-output.md) defines
|
||||
publication. Pipeline mechanics are described in
|
||||
[D&D module internals](../internal/dnd.md).
|
||||
92
docs/integrations/dnd-npc-registry-artifacts.md
Normal file
92
docs/integrations/dnd-npc-registry-artifacts.md
Normal file
@@ -0,0 +1,92 @@
|
||||
# D&D NPC Registry Artifact
|
||||
|
||||
This contract defines the durable NPC registry produced by `dnd/npc-registry`. It is a
|
||||
minimal, source-grounded identity registry for other D&D artifacts, not a
|
||||
character sheet or a relationship summary.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/npc-registry` |
|
||||
| Schema ID | `notarius.dnd.npc_registry` |
|
||||
| Schema name | `notarius_dnd_npc_registry_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
| Identity policy | `dnd.npc_registry.identity.v1` |
|
||||
|
||||
`v1` accepts one strict JSON object with required `npcs`; the array may be
|
||||
empty. NPC and source-reference objects reject unknown fields. An incompatible
|
||||
artifact shape or identity-policy change uses a new version or policy.
|
||||
|
||||
## Wire shape and identity
|
||||
|
||||
Each NPC has these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `id` | `npc:sha256:` followed by 64 lowercase hexadecimal characters. |
|
||||
| `name` | Non-empty canonical display name. |
|
||||
| `source_refs` | One or more transcript evidence ranges for the identity. |
|
||||
|
||||
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
|
||||
The source ID identifies the transcript, unit IDs are positive inclusive unit
|
||||
identifiers, and the start may not follow the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"npcs": [
|
||||
{
|
||||
"id": "npc:sha256:35ba5f679aee69e07ae3bd65c44278f29539d5dc9bb5225db1c0060555b23221",
|
||||
"name": "Mira Thorn",
|
||||
"source_refs": [
|
||||
{"source_id": "session-7", "start_unit_id": 4, "end_unit_id": 5}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The ID is deterministic: normalize the name to Unicode NFKC, normalize the
|
||||
supported apostrophe forms, collapse whitespace, case-fold it, then serialize
|
||||
`["dnd.npc_registry.identity.v1", comparison_name]` as compact JSON. SHA-256
|
||||
those UTF-8 bytes and prefix the lowercase hexadecimal digest with
|
||||
`npc:sha256:`. Each canonical identity and ID appears at most once.
|
||||
Normalization collapses records with the same canonical identity, retains their
|
||||
earliest position, and merges
|
||||
their canonicalized evidence; it does not add aliases, roles, descriptions, or
|
||||
relationship fields.
|
||||
|
||||
When evidence supports a semantically duplicate group, the canonical display
|
||||
name is one of that group's supplied candidates. A complete, stable proper name
|
||||
is preferred over an abbreviation. An unadorned proper name is preferred over
|
||||
the same name plus a contextual class, role, title, or relationship descriptor
|
||||
unless the transcript establishes that descriptor as part of the person's
|
||||
name. A longer candidate is not preferred solely because it includes such a
|
||||
descriptor.
|
||||
|
||||
## Scope and consumers
|
||||
|
||||
Only individually identifiable NPC names with transcript evidence belong in
|
||||
this artifact. A factual third-party mention can establish an identity even if
|
||||
the NPC is not present, speaking, or acting in the cited passage. Names used
|
||||
only in hypothetical, speculative, or imagined examples are excluded, as are
|
||||
groups, generic roles, invented labels, and descriptive enrichment. Its source
|
||||
references prove registry provenance; they do not become evidence for a spell,
|
||||
occurrence, combat, or enemy-event occurrence.
|
||||
|
||||
Registry evidence establishes an identity, not an [NPC occurrence](dnd-npc-occurrence-artifacts.md).
|
||||
That later artifact independently records any current-transcript occurrence
|
||||
with its own cited evidence and category.
|
||||
|
||||
This registry can ground actor or caster names in the [spell](dnd-spell-artifacts.md)
|
||||
and [combat-turn](dnd-combat-turn-artifacts.md) artifacts. It is required to
|
||||
resolve the canonical `name` in an [NPC occurrence](dnd-npc-occurrence-artifacts.md).
|
||||
Occurrence consumers receive names-only grounding; Notarius resolves the
|
||||
selected canonical name and writes the unchanged exact durable ID/name pair.
|
||||
Spells, combat turns, and the [enemy-event artifact](dnd-enemy-event-artifacts.md)
|
||||
also receive names-only grounding for actor or subject display. None of these
|
||||
projections supply later-artifact evidence. [Configuration](../config.md#d-d-reference-slots)
|
||||
owns the `npc_registry` binding rules.
|
||||
The [JSON output contract](json-output.md) defines publication, and
|
||||
[D&D module internals](../internal/dnd.md) owns pipeline mechanics.
|
||||
70
docs/integrations/dnd-scene-description-artifacts.md
Normal file
70
docs/integrations/dnd-scene-description-artifacts.md
Normal file
@@ -0,0 +1,70 @@
|
||||
# D&D Scene-Description Artifact
|
||||
|
||||
This contract defines the durable output of `dnd/scene-descriptions`. Each
|
||||
record classifies one accepted transcript chunk and gives it a minimal
|
||||
source-grounded title and summary.
|
||||
|
||||
## Identity and compatibility
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/scene-description-list` |
|
||||
| Schema ID | `notarius.dnd.scene_descriptions` |
|
||||
| Schema name | `notarius_dnd_scene_descriptions_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
|
||||
`v1` is a strict JSON object with required non-empty `scenes`. Scene and
|
||||
source-reference objects reject unknown fields. An incompatible shape change
|
||||
requires a new schema version.
|
||||
|
||||
## Wire shape
|
||||
|
||||
Each scene has exactly these required fields:
|
||||
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `id` | Non-empty accepted chunk ID, assigned by Notarius. |
|
||||
| `source_ref` | The assigned inclusive source range for that chunk. |
|
||||
| `kind` | `combat`, `narrative`, `recap`, or `meta`. |
|
||||
| `title` | Non-empty, trimmed, source-grounded title. |
|
||||
| `summary` | Non-empty, trimmed, source-grounded summary. |
|
||||
|
||||
`source_ref` has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
|
||||
Its source ID identifies the input transcript; its positive unit IDs identify
|
||||
the chunk's inclusive range, with the start no later than the end.
|
||||
|
||||
```json
|
||||
{
|
||||
"scenes": [
|
||||
{
|
||||
"id": "chunk-000001",
|
||||
"source_ref": {"source_id": "session-7", "start_unit_id": 1, "end_unit_id": 3},
|
||||
"kind": "narrative",
|
||||
"title": "Arrival at the watchtower",
|
||||
"summary": "The party reaches the ruined watchtower and begins to investigate it."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Meaning and normalized form
|
||||
|
||||
`combat` identifies a chunk where active combat is the central activity.
|
||||
`narrative` is current in-world play that is not principally combat, recap, or
|
||||
meta discussion. `recap` is primarily a recounting of an earlier session, and
|
||||
`meta` is primarily out-of-character discussion. The artifact does not add
|
||||
participants, confidence, events, or information absent from the chunk.
|
||||
|
||||
Normalization trims title and summary, orders scenes by source position and
|
||||
then ID, and removes exact duplicate records. A reused ID with different
|
||||
durable fields, or the same source range with different kind, title, or
|
||||
summary, is invalid. It does not merge adjacent ranges, alter prose, or infer
|
||||
missing scenes.
|
||||
|
||||
The [combat-turn artifact](dnd-combat-turn-artifacts.md) and
|
||||
[enemy-event artifact](dnd-enemy-event-artifacts.md) use an exact matching
|
||||
`combat` scene only as eligibility control; scene title, summary, and source
|
||||
reference never become their evidence. Publication is defined by the
|
||||
[JSON output contract](json-output.md); implementation details live in
|
||||
[D&D module internals](../internal/dnd.md).
|
||||
@@ -1,152 +1,73 @@
|
||||
# D&D Spell-Cast Artifacts
|
||||
# D&D Spell Artifact
|
||||
|
||||
This document is the durable artifact contract for approved
|
||||
`dnd.spell_cast` artifacts produced by the implemented `dnd/spells` extractor.
|
||||
This contract defines the durable output of the `dnd/spells` extractor and
|
||||
normalizer. It records source-grounded spell-casting occurrences; it is not a
|
||||
spellbook, a rules lookup result, or a record of hypothetical casts.
|
||||
|
||||
## Artifact Identity
|
||||
## Identity and compatibility
|
||||
|
||||
- Extractor key: `dnd/spells`
|
||||
- Artifact type: `dnd.spell_cast`
|
||||
- Schema version: `v1`
|
||||
- Prompt ID: `dnd.spells`
|
||||
- Response schema key: `dnd_spells`
|
||||
- Response schema ID: `notarius.dnd.spells`
|
||||
- Response schema name: `notarius_dnd_spells_v1`
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `dnd/spell-list` |
|
||||
| Schema ID | `notarius.dnd.spells` |
|
||||
| Schema name | `notarius_dnd_spells_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Media type | `application/json` |
|
||||
|
||||
The extractor requires source chunks and transcript source capability. It
|
||||
returns generic artifact candidates that are serialized by the JSON output
|
||||
module.
|
||||
`v1` is a single strict JSON object. It requires `spell_casts`; the array may
|
||||
be empty. Each spell-cast object and source-reference object rejects unknown
|
||||
fields. An incompatible shape change requires a new schema version.
|
||||
|
||||
## Artifact Envelope
|
||||
## Wire shape
|
||||
|
||||
Approved artifacts use the generic artifact envelope documented in
|
||||
[JSON Output](json-output.md#artifact-files):
|
||||
Each `spell_casts` entry has these required fields:
|
||||
|
||||
```json
|
||||
{
|
||||
"extractor_key": "dnd/spells",
|
||||
"artifact_type": "dnd.spell_cast",
|
||||
"schema_version": "v1",
|
||||
"payload": {
|
||||
"caster": "Aria",
|
||||
"spell": "Cure Wounds",
|
||||
"effect": "heals an injured ally",
|
||||
"narrative_description": "Aria raises her holy symbol and casts Cure Wounds."
|
||||
},
|
||||
"source_refs": [
|
||||
{
|
||||
"source_id": "session-alpha",
|
||||
"start_unit_id": "seg-001",
|
||||
"end_unit_id": "seg-001"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
| Field | Contract |
|
||||
| --- | --- |
|
||||
| `caster` | Non-empty in-world character or creature name. |
|
||||
| `spell` | Non-empty spell name. |
|
||||
| `source_refs` | One or more transcript evidence ranges. |
|
||||
|
||||
## Payload Fields
|
||||
|
||||
The `payload` object contains:
|
||||
|
||||
- `caster`: in-world character or creature casting the spell;
|
||||
- `spell`: spell name;
|
||||
- `effect`: concise spell effect in the scene;
|
||||
- `narrative_description`: short description of the spell cast in context.
|
||||
|
||||
All payload fields are strings and must be non-empty after trimming.
|
||||
|
||||
`caster` is the in-world caster, not the transcript speaker.
|
||||
|
||||
## Source References
|
||||
|
||||
Source references live on the artifact envelope as `source_refs`; they are not
|
||||
duplicated inside the `payload`.
|
||||
|
||||
Each source reference uses the generic source-reference shape:
|
||||
|
||||
- `source_id`
|
||||
- `start_unit_id`
|
||||
- `end_unit_id`
|
||||
|
||||
Validation requires:
|
||||
|
||||
- at least one source reference;
|
||||
- non-empty source ID and unit IDs;
|
||||
- source ID matching the source document ID;
|
||||
- start and end unit IDs existing in the source document;
|
||||
- start unit appearing before or at the same position as end unit.
|
||||
|
||||
## Structured LLM Response Shape
|
||||
|
||||
The extractor asks the LLM for this top-level response shape:
|
||||
Every source reference has exactly `source_id`, `start_unit_id`, and
|
||||
`end_unit_id`. The source ID identifies the input transcript; the unit IDs are
|
||||
positive inclusive unit identifiers, and the start may not follow the end in
|
||||
that source. References are evidence for the cast, not campaign-reference or
|
||||
NPC-registry provenance.
|
||||
|
||||
```json
|
||||
{
|
||||
"spell_casts": [
|
||||
{
|
||||
"caster": "Aria",
|
||||
"spell": "Cure Wounds",
|
||||
"effect": "heals an injured ally",
|
||||
"narrative_description": "Aria raises her holy symbol and casts Cure Wounds.",
|
||||
"caster": "Mira Thorn",
|
||||
"spell": "Fireball",
|
||||
"source_refs": [
|
||||
{
|
||||
"source_id": "session-alpha",
|
||||
"start_unit_id": "seg-001",
|
||||
"end_unit_id": "seg-001"
|
||||
}
|
||||
{"source_id": "session-7", "start_unit_id": 12, "end_unit_id": 13}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
`spell_casts` must be present. It may be empty when no spell casts are found.
|
||||
## Evidence and normalized form
|
||||
|
||||
The response schema asset is embedded at
|
||||
`internal/modules/extract/dnd/spells/assets/schemas/dnd_spells.v1.json`.
|
||||
An entry represents an actual cast or an unambiguous declared attempt. A spell
|
||||
mention, rules discussion, plan, or catalog match alone is not an occurrence.
|
||||
The configured catalog checks the name; it does not establish evidence.
|
||||
|
||||
## Validators
|
||||
When normalization is selected, recognized spell names use the effective
|
||||
catalog's canonical display name. Source references are put in canonical source
|
||||
order and exact duplicate references are removed. A later entry is collapsed
|
||||
only when it has the same canonical spell, the same case- and
|
||||
whitespace-insensitive caster identity, and the same complete valid reference
|
||||
sequence. Remaining entries retain their merged order.
|
||||
|
||||
The extractor supplies two deterministic validators by default:
|
||||
The optional normalized [NPC registry artifact](dnd-npc-registry-artifacts.md) can ground a
|
||||
caster name. Its own references remain registry provenance and are never copied
|
||||
into `source_refs`.
|
||||
|
||||
- `dnd/spells/shape`
|
||||
- `dnd/spells/source_refs`
|
||||
## Related contracts
|
||||
|
||||
Rejection reason codes:
|
||||
|
||||
- `invalid_payload`: payload JSON cannot be decoded as a spell-cast payload.
|
||||
- `missing_required_field`: `caster`, `spell`, `effect`, or
|
||||
`narrative_description` is blank.
|
||||
- `missing_source_ref`: candidate has no source references.
|
||||
- `invalid_source_ref`: at least one source reference fails generic source
|
||||
reference validation.
|
||||
|
||||
Rejected candidates are written to `rejected.json` by the JSON output module.
|
||||
|
||||
## Manifest Metadata
|
||||
|
||||
The extractor adds prompt and response-schema provenance under the artifact lane
|
||||
manifest metadata:
|
||||
|
||||
```json
|
||||
{
|
||||
"metadata": {
|
||||
"extractor": {
|
||||
"prompt_id": "dnd.spells",
|
||||
"prompt_version": "v1",
|
||||
"prompt_sha256": "sha256:...",
|
||||
"response_schema_key": "dnd_spells",
|
||||
"response_schema_id": "notarius.dnd.spells",
|
||||
"response_schema_name": "notarius_dnd_spells_v1",
|
||||
"response_schema_version": "v1",
|
||||
"response_schema_sha256": "sha256:..."
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Raw prompt and schema content are not included in manifest metadata.
|
||||
|
||||
## Compatibility Limit
|
||||
|
||||
This contract covers only `dnd.spell_cast` artifacts produced by the
|
||||
implemented spell-cast extractor.
|
||||
The [spell-catalog overlay contract](dnd-spell-catalog-overlays.md) defines
|
||||
the configured catalog additions. The [JSON output contract](json-output.md)
|
||||
defines where this logical artifact is published; [D&D module internals](../internal/dnd.md)
|
||||
describes extraction and validation mechanics.
|
||||
|
||||
78
docs/integrations/dnd-spell-catalog-overlays.md
Normal file
78
docs/integrations/dnd-spell-catalog-overlays.md
Normal file
@@ -0,0 +1,78 @@
|
||||
# D&D Spell-Catalog Overlays
|
||||
|
||||
This document defines the optional JSON overlay consumed by the D&D spell
|
||||
extractor. An overlay contributes campaign spell names and aliases for
|
||||
recognition. It does not define spell rules, effects, levels, classes, or
|
||||
transcript evidence. Bind the optional `spell_catalog` reference as described
|
||||
in [Configuration](../config.md#references-and-ordered-handoffs).
|
||||
|
||||
## Contract Identity
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Consumer | D&D spell extraction and normalization |
|
||||
| Reference slot | `spell_catalog` |
|
||||
| Media type | `application/json` |
|
||||
| Required schema version | `notarius.dnd.spell-catalog-overlay.v1` |
|
||||
| Base catalog | Embedded D&D 5e 2014 SRD catalog |
|
||||
|
||||
At most one overlay document may be bound. The maintained example is
|
||||
[dnd-spell-catalog.json](../../examples/dnd-spell-catalog.json).
|
||||
|
||||
## Wire Shape
|
||||
|
||||
This is a minimal valid overlay:
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "notarius.dnd.spell-catalog-overlay.v1",
|
||||
"catalogs": [
|
||||
{
|
||||
"id": "campaign.example",
|
||||
"ruleset": "dnd-5e-2014",
|
||||
"source": {"title": "Example campaign spells"},
|
||||
"spells": [{"name": "Aegis of Emberfall"}]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Required | Meaning and constraints |
|
||||
| --- | --- | --- |
|
||||
| `schema_version` | Yes | Exactly `notarius.dnd.spell-catalog-overlay.v1`. |
|
||||
| `catalogs` | Yes | Non-empty array of catalog objects with unique IDs. |
|
||||
| `catalogs[].id` | Yes | Non-empty trimmed string. |
|
||||
| `catalogs[].ruleset` | Yes | Exactly `dnd-5e-2014`. |
|
||||
| `catalogs[].source.title` | Yes | Non-empty trimmed string. |
|
||||
| `catalogs[].source.version` | No | String when present. |
|
||||
| `catalogs[].source.url` | No | String when present. |
|
||||
| `catalogs[].source.license` | No | String when present. |
|
||||
| `catalogs[].spells` | Yes | Non-empty array of spell objects. |
|
||||
| `catalogs[].spells[].name` | Yes | Non-empty trimmed string. |
|
||||
| `catalogs[].spells[].aliases` | No | Array of non-empty trimmed strings when present. |
|
||||
|
||||
Unknown fields are rejected at every object level. The document must contain
|
||||
one JSON value; `null` is not accepted for optional strings or aliases.
|
||||
|
||||
## Composition And Compatibility
|
||||
|
||||
Notarius starts with the embedded base catalog, then applies overlay catalogs
|
||||
in ascending catalog-ID order. A new canonical spell name adds a recognition
|
||||
entry. If an overlay names an existing canonical spell, it augments that spell
|
||||
with aliases while retaining the established display spelling.
|
||||
|
||||
Repeated aliases for the same spell are accepted. A canonical-name, canonical-
|
||||
to-alias, or alias-to-alias collision between different spells is rejected,
|
||||
including a collision with the embedded catalog. Matching uses the catalog’s
|
||||
case, whitespace, and apostrophe normalization, so authors should avoid names
|
||||
or aliases that normalize to another spell.
|
||||
|
||||
Spell extraction receives the effective catalog as deterministic canonical-name
|
||||
and alias pairs. An alias in the transcript selects its associated canonical
|
||||
name; the extractor is instructed to return that canonical spelling. The
|
||||
projection contains no catalog source metadata or provenance, and aliases
|
||||
remain recognition context rather than transcript evidence.
|
||||
|
||||
The overlay is a recognition aid only. The durable spell-artifact schema and
|
||||
source-evidence rules are defined by the
|
||||
[D&D spell artifact contract](dnd-spell-artifacts.md).
|
||||
116
docs/integrations/evidence-context.md
Normal file
116
docs/integrations/evidence-context.md
Normal file
@@ -0,0 +1,116 @@
|
||||
# Published Evidence Context
|
||||
|
||||
This contract defines the optional `source/evidence-context` artifact emitted
|
||||
by the production JSON output. Its configuration is owned by
|
||||
[Configuration](../config.md#module-bindings-and-validators); its logical-file
|
||||
discovery is owned by [Published JSON Output](json-output.md).
|
||||
|
||||
## Identity And Discovery
|
||||
|
||||
When enabled, the JSON bundle contains `evidence-context.json` and an
|
||||
`index.json` `evidence_context` descriptor with the same six fields as other
|
||||
pipeline-wide artifact descriptors.
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Artifact kind | `source/evidence-context` |
|
||||
| Media type | `application/json` |
|
||||
| Schema ID | `notarius.source.evidence_context` |
|
||||
| Schema name | `notarius_source_evidence_context_v1` |
|
||||
| Schema version | `v1` |
|
||||
| Logical file | `evidence-context.json` |
|
||||
|
||||
Consumers must discover the file from the descriptor, verify all six descriptor
|
||||
fields, and decode only a supported schema version. The descriptor is optional:
|
||||
its absence means evidence publication was not enabled for that bundle.
|
||||
|
||||
## Payload
|
||||
|
||||
The v1 payload is a JSON object with required `source_id`, `source_digest`,
|
||||
`window_units`, `selected_lanes`, and `contexts` fields. `selected_lanes` and
|
||||
`contexts` are always arrays; an enabled configuration with no accepted direct
|
||||
evidence publishes `contexts: []`.
|
||||
|
||||
```json
|
||||
{
|
||||
"source_id": "session-alpha",
|
||||
"source_digest": "sha256:0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef",
|
||||
"window_units": 1,
|
||||
"selected_lanes": ["npc_registry", "spells"],
|
||||
"contexts": [
|
||||
{
|
||||
"context_ref": {
|
||||
"source_id": "session-alpha",
|
||||
"start_unit_id": 10,
|
||||
"end_unit_id": 20
|
||||
},
|
||||
"evidence_refs": [
|
||||
{
|
||||
"lane_id": "spells",
|
||||
"source_ref": {
|
||||
"source_id": "session-alpha",
|
||||
"start_unit_id": 10,
|
||||
"end_unit_id": 10
|
||||
}
|
||||
}
|
||||
],
|
||||
"units": [
|
||||
{
|
||||
"id": 10,
|
||||
"kind": "transcript_segment",
|
||||
"text": "Aria casts Cure Wounds.",
|
||||
"ref": {
|
||||
"source_id": "session-alpha",
|
||||
"start_unit_id": 10,
|
||||
"end_unit_id": 10
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 20,
|
||||
"kind": "transcript_segment",
|
||||
"text": "The party regroups.",
|
||||
"ref": {
|
||||
"source_id": "session-alpha",
|
||||
"start_unit_id": 20,
|
||||
"end_unit_id": 20
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Each context requires a `context_ref` object and `evidence_refs` and `units`
|
||||
arrays. `context_ref` identifies the first and last included unit. Each
|
||||
evidence entry contains a selected `lane_id` and an original `source_ref`. A
|
||||
unit uses the existing source-unit shape: required `id`, `kind`, `text`, and
|
||||
self `ref`, plus optional JSON-object `metadata`. Fixed payload objects reject
|
||||
unknown fields; unit metadata may contain application-defined JSON values.
|
||||
|
||||
## Citations And Context
|
||||
|
||||
`evidence_refs` are the authoritative citations. They identify the direct
|
||||
references emitted by accepted normalized artifacts. `context_ref` and the
|
||||
units collection include those cited units plus nearby source units selected by
|
||||
the configured window. They are explanatory context, not widened citations.
|
||||
|
||||
Only accepted outputs from the configured lane allowlist contribute. Rejected,
|
||||
failed, absent, and lane-filtered outputs do not contribute. The artifact never
|
||||
contains raw input bytes, prompts, model responses, auxiliary reference
|
||||
content, credentials, or filesystem paths.
|
||||
|
||||
## Ordering And Compatibility
|
||||
|
||||
The selected lane allowlist is lexical. Contexts and units are in source
|
||||
document position order, not numeric unit-ID order. Direct evidence entries
|
||||
are deterministically ordered by lane and source reference. Overlapping or
|
||||
contiguous windows merge, and each source unit appears at most once in the
|
||||
resulting contexts.
|
||||
|
||||
The artifact is additive to the JSON bundle and is not a lane payload,
|
||||
normalized-output count, checkpoint, or generated reference. Consumers that
|
||||
do not need it must tolerate the absent optional descriptor. Consumers that do
|
||||
use it should preserve the artifact and its schema identity with the run
|
||||
provenance, and should treat its source text and metadata as sensitive durable
|
||||
content.
|
||||
@@ -1,192 +1,142 @@
|
||||
# JSON Output
|
||||
# Published JSON Output
|
||||
|
||||
This document is the durable JSON output file-format contract produced by the
|
||||
implemented `json` output module and written by the CLI.
|
||||
This document defines the logical JSON bundle emitted by the production JSON
|
||||
output encoder. The bundle’s physical destination, atomic publication, and
|
||||
retention are operational concerns; see [Operations](../operations.md#output-bundles).
|
||||
Output configuration, including chunk-map and evidence-context publication, belongs in
|
||||
[Configuration](../config.md#module-bindings-and-validators).
|
||||
|
||||
## Output Directory
|
||||
## Bundle Layout
|
||||
|
||||
The CLI writes logical output files under:
|
||||
All paths below are logical, relative, slash-separated bundle paths. The
|
||||
encoder always emits the first four JSON files below and adds lane or
|
||||
pipeline-wide artifact files when their corresponding artifacts are available:
|
||||
|
||||
```text
|
||||
<output-root>/<run-id>/
|
||||
```
|
||||
A subprocess caller first obtains the physical bundle root from the
|
||||
[run-result receipt](run-result.md), then resolves `index.json` beneath that
|
||||
root for the logical discovery described here.
|
||||
|
||||
The default output root is `./notarius-output`. Operational behavior is covered
|
||||
in [Operations](../operations.md).
|
||||
| Path | Purpose |
|
||||
| --- | --- |
|
||||
| `index.json` | Entry point that names the other published files and lane payloads. |
|
||||
| `manifest.json` | Run provenance and result summaries. |
|
||||
| `rejected.json` | Rejected pipeline outputs. |
|
||||
| `warnings.json` | Accepted-output and run warnings. |
|
||||
| `lanes/<safe-lane-id>.json` | One normalized artifact payload for each lane. |
|
||||
| `chunk-map.json` | Optional accepted chunk map, when its export is enabled and available. |
|
||||
| `evidence-context.json` | Optional source-context artifact, when evidence publication is enabled. |
|
||||
|
||||
## Files
|
||||
|
||||
The `json` output module writes:
|
||||
|
||||
- `index.json`
|
||||
- `manifest.json`
|
||||
- `artifacts/<artifact-type>.json`, one file per approved artifact type
|
||||
- `rejected.json`
|
||||
- `warnings.json`
|
||||
|
||||
Files are pretty-printed JSON with a trailing newline.
|
||||
JSON files are pretty-printed with a trailing newline. Lane payloads are
|
||||
accepted only when their media type is `application/json`.
|
||||
|
||||
## `index.json`
|
||||
|
||||
Shape:
|
||||
`index.json` is the bundle’s discovery document. An approved run with no
|
||||
normalized lanes has this valid minimal index:
|
||||
|
||||
```json
|
||||
{
|
||||
"manifest_file": "manifest.json",
|
||||
"artifact_files": [
|
||||
{
|
||||
"artifact_type": "dnd.spell_cast",
|
||||
"file": "artifacts/dnd.spell_cast.json"
|
||||
}
|
||||
],
|
||||
"output_files": [],
|
||||
"rejected_file": "rejected.json",
|
||||
"warnings_file": "warnings.json"
|
||||
}
|
||||
```
|
||||
|
||||
`artifact_files` is sorted by artifact type. It is empty when no artifacts are
|
||||
approved.
|
||||
| Field | Required | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `manifest_file` | Yes | Always `manifest.json`. |
|
||||
| `output_files` | Yes | Lane descriptors sorted by `lane_id`. |
|
||||
| `rejected_file` | Yes | Always `rejected.json`. |
|
||||
| `warnings_file` | Yes | Always `warnings.json`. |
|
||||
| `chunk_map` | No | Descriptor for the pipeline-wide `chunk-map.json`; never a lane descriptor. |
|
||||
| `evidence_context` | No | Descriptor for the pipeline-wide `evidence-context.json`; never a lane descriptor. |
|
||||
|
||||
Each lane descriptor has required `lane_id` and `file`. It may also include
|
||||
`media_type`, `module_key`, `schema_id`, `schema_name`, and `schema_version`
|
||||
when supplied by the normalized artifact. Each pipeline-wide artifact
|
||||
descriptor (`chunk_map` or `evidence_context`) contains `artifact_kind`,
|
||||
`file`, `media_type`, `schema_id`, `schema_name`, and `schema_version`. Their
|
||||
payloads are defined by the [Accepted Chunk Map contract](chunk-map.md) and
|
||||
[Published Evidence Context](evidence-context.md), respectively.
|
||||
|
||||
The lane path is derived from its lane ID. Characters outside letters, digits,
|
||||
periods, underscores, and hyphens become underscores; `..` sequences are
|
||||
neutralized; leading and trailing periods and underscores are removed. A lane
|
||||
that produces an empty name, or two lanes that produce the same path, makes
|
||||
output encoding fail.
|
||||
|
||||
## Lane Payloads
|
||||
|
||||
Each `lanes/<safe-lane-id>.json` file is the codec-owned normalized JSON for
|
||||
that lane. Consumers should use the index descriptor’s schema identity rather
|
||||
than infer a lane schema from its name. The current D&D payload contracts are
|
||||
[spells](dnd-spell-artifacts.md), [NPC registry](dnd-npc-registry-artifacts.md),
|
||||
[NPC occurrences](dnd-npc-occurrence-artifacts.md),
|
||||
[combat turns](dnd-combat-turn-artifacts.md),
|
||||
[item registry](dnd-item-registry-artifacts.md),
|
||||
[item occurrences](dnd-item-occurrence-artifacts.md),
|
||||
[scene descriptions](dnd-scene-description-artifacts.md),
|
||||
[enemy events](dnd-enemy-event-artifacts.md),
|
||||
[location registry](dnd-location-registry-artifacts.md), and
|
||||
[location occurrences](dnd-location-occurrence-artifacts.md).
|
||||
|
||||
## `manifest.json`
|
||||
|
||||
`manifest.json` contains a run manifest:
|
||||
`manifest.json` is published provenance, not a copy of lane payloads or a
|
||||
checkpoint store. Fields without a value may be omitted. Its top-level fields
|
||||
group into the following externally observable summaries:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "run-123",
|
||||
"pipeline_id": "dnd-session",
|
||||
"pipeline_digest": "sha256:...",
|
||||
"input_module": "seriatim",
|
||||
"chunker": "generic",
|
||||
"source_digests": ["sha256:..."],
|
||||
"extractors": ["dnd/spells"],
|
||||
"merger": "appendorder",
|
||||
"normalizer": "noop",
|
||||
"output_encoder": "json",
|
||||
"artifact_lanes": [
|
||||
{
|
||||
"id": "spells",
|
||||
"extractor": "dnd/spells",
|
||||
"merger": "appendorder",
|
||||
"normalizer": "noop"
|
||||
}
|
||||
],
|
||||
"llm_profiles": [
|
||||
{
|
||||
"id": "default",
|
||||
"provider": "openai-compatible",
|
||||
"model": "configured-model"
|
||||
}
|
||||
],
|
||||
"validation_status": "approved",
|
||||
"started_at": "2026-01-01T00:00:00Z",
|
||||
"completed_at": "2026-01-01T00:00:01Z"
|
||||
}
|
||||
```
|
||||
| Group | Fields |
|
||||
| --- | --- |
|
||||
| Run identity and result | `run_id`, `pipeline_id`, `pipeline_digest`, `schema_version`, `validation_status`, `started_at`, `completed_at` |
|
||||
| Resolved components | `input_module`, `chunker`, `extractors`, `merger`, `normalizer`, `output_encoder`, `artifact_lanes`, `validator_chains`, `module_metadata` |
|
||||
| Source and references | `source_digests`, `references` |
|
||||
| Published result summaries | `normalized_outputs`, `rejected_outputs` |
|
||||
| Execution summaries | `chunk_plan`, `checkpoint_decisions`, `llm_profiles`, `metadata` |
|
||||
|
||||
Fields with empty values may be omitted by JSON encoding.
|
||||
`references` records provenance such as the target, slot, origin, digest,
|
||||
media type, size, and generated-artifact identity. It does not contain
|
||||
reference content. `normalized_outputs` and `rejected_outputs` likewise
|
||||
summarize results without embedding lane payload bytes. A chunk-plan summary is
|
||||
provenance for the plan used by this run; cache records, debug artifacts, and
|
||||
other operational state are not published as bundle files.
|
||||
|
||||
`validation_status` is `approved` when no candidates were rejected and
|
||||
`rejected` when one or more candidates were rejected.
|
||||
When present, `metadata.session_id` is the effective non-secret routing
|
||||
correlation identifier used for the run. It can be visible to providers and is
|
||||
not a substitute for a cache or checkpoint identity. Its generation and
|
||||
override behavior are defined by the [CLI reference](../cli.md#run).
|
||||
|
||||
## Artifact Files
|
||||
Each `llm_profiles` entry identifies effective, non-secret LLM execution
|
||||
provenance:
|
||||
|
||||
Each artifact file has this shape:
|
||||
| Field | Required | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `id` | Yes | Selected PromptKit profile identifier. |
|
||||
| `provider` | No | Notarius adapter provider identifier. |
|
||||
| `model` | No | Effective provider model identifier. |
|
||||
| `backend_id` | No | Effective PromptKit backend registration identifier. Endpoint-only profiles omit it. |
|
||||
| `reasoning_effort` | No | Effective opaque provider reasoning setting. An empty or explicitly cleared setting is omitted. |
|
||||
|
||||
```json
|
||||
{
|
||||
"artifact_type": "dnd.spell_cast",
|
||||
"artifacts": [
|
||||
{
|
||||
"extractor_key": "dnd/spells",
|
||||
"artifact_type": "dnd.spell_cast",
|
||||
"schema_version": "v1",
|
||||
"payload": {},
|
||||
"source_refs": [
|
||||
{
|
||||
"source_id": "session-alpha",
|
||||
"start_unit_id": "seg-001",
|
||||
"end_unit_id": "seg-001"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
These values describe observed execution; they are not a backend-registration
|
||||
interface. Entries that differ by backend or effective reasoning remain
|
||||
distinct even when their profile, provider, and model are otherwise equal.
|
||||
|
||||
Artifact envelope fields:
|
||||
## Rejections And Warnings
|
||||
|
||||
- `extractor_key`: extractor module key.
|
||||
- `artifact_type`: artifact type.
|
||||
- `schema_version`: artifact schema version.
|
||||
- `payload`: artifact-type-specific JSON payload.
|
||||
- `source_refs`: optional generic source references.
|
||||
- `metadata`: optional artifact metadata.
|
||||
`rejected.json` is always an object with a `rejected` array. Each entry has
|
||||
required `stage` and `message`; `step_id`, `lane_id`, `module_key`, `chunk_id`,
|
||||
`chunk_index`, `validator_name`, `reason_code`, `attempt_count`, and
|
||||
`diagnostic_artifact_path` are present only when applicable.
|
||||
|
||||
Artifact file names are produced by sanitizing the artifact type:
|
||||
`warnings.json` is always an object with a `warnings` array. Each warning has
|
||||
`reason_code` and `message`; `scope` is optional. Both arrays are empty when
|
||||
there is nothing to report.
|
||||
|
||||
- characters outside `A-Z`, `a-z`, `0-9`, `.`, `_`, and `-` become `_`;
|
||||
- repeated `..` sequences are replaced;
|
||||
- leading and trailing `.`, `_`, and `-` are trimmed;
|
||||
- empty sanitized names are rejected.
|
||||
## Compatibility
|
||||
|
||||
For current D&D spell-cast artifacts, the file is
|
||||
`artifacts/dnd.spell_cast.json`.
|
||||
|
||||
## `rejected.json`
|
||||
|
||||
Shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"rejected": [
|
||||
{
|
||||
"candidate": {
|
||||
"index": 0,
|
||||
"extractor_key": "dnd/spells",
|
||||
"artifact_type": "dnd.spell_cast",
|
||||
"schema_version": "v1",
|
||||
"payload": {},
|
||||
"source_refs": []
|
||||
},
|
||||
"validator_name": "dnd/spells/source_refs",
|
||||
"reason_code": "missing_source_ref",
|
||||
"message": "spell cast candidate must include at least one source ref"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
`rejected` is an empty array when no candidates are rejected.
|
||||
|
||||
## `warnings.json`
|
||||
|
||||
Shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"warnings": [
|
||||
{
|
||||
"scope": "output",
|
||||
"reason_code": "example_warning",
|
||||
"message": "warning message"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
`warnings` is an empty array when no warnings are reported.
|
||||
|
||||
## Path Safety
|
||||
|
||||
The output module returns slash-separated logical paths. The CLI also validates
|
||||
logical output names before writing:
|
||||
|
||||
- names must be non-empty;
|
||||
- names must be relative;
|
||||
- names must be clean;
|
||||
- names must use `/`, not `\`;
|
||||
- names must not contain `..`;
|
||||
- resolved paths must stay under the run output directory.
|
||||
|
||||
Durable writes are atomic per file.
|
||||
The index is the authoritative map from a logical lane to its published
|
||||
payload. Consumers must tolerate omitted optional manifest and descriptor
|
||||
fields, and should rely on the linked artifact contract for each lane’s JSON
|
||||
shape. This contract describes the published logical bundle only; it does not
|
||||
promise a filesystem layout or expose internal state formats.
|
||||
|
||||
@@ -1,128 +0,0 @@
|
||||
# OpenAI-Compatible Structured Output
|
||||
|
||||
This document describes the external LLM provider contract implemented by the
|
||||
production Notarius LLM client.
|
||||
|
||||
## Provider
|
||||
|
||||
- Provider key: `openai-compatible`
|
||||
- HTTP method: `POST`
|
||||
- Endpoint: `<base_url>/chat/completions`
|
||||
- Request body: JSON
|
||||
- Response mode: chat completions with structured JSON schema output
|
||||
|
||||
`base_url` is trimmed of trailing slashes before `/chat/completions` is
|
||||
appended. Configure provider settings in [Configuration](../config.md).
|
||||
|
||||
## Request
|
||||
|
||||
The client sends a JSON object with:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "configured-model",
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
"content": "..."
|
||||
},
|
||||
{
|
||||
"role": "user",
|
||||
"content": "..."
|
||||
}
|
||||
],
|
||||
"response_format": {
|
||||
"type": "json_schema",
|
||||
"json_schema": {
|
||||
"name": "schema_name",
|
||||
"strict": true,
|
||||
"schema": {}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Implemented request behavior:
|
||||
|
||||
- `model` comes from the structured completion request when set, otherwise from
|
||||
the configured LLM profile.
|
||||
- `messages` must be non-empty; each role and content must be non-empty after
|
||||
trimming.
|
||||
- `response_format.type` is always `json_schema`.
|
||||
- `response_format.json_schema.strict` is always `true`.
|
||||
- `response_format.json_schema.name` and `schema` come from the extractor or
|
||||
validator making the call.
|
||||
|
||||
If an API key is configured, the client sends:
|
||||
|
||||
```text
|
||||
Authorization: Bearer <api-key>
|
||||
```
|
||||
|
||||
The client always sends `Content-Type: application/json`.
|
||||
|
||||
## Response
|
||||
|
||||
The client expects a JSON response with at least one choice:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "provider-model",
|
||||
"choices": [
|
||||
{
|
||||
"message": {
|
||||
"content": "{\"field\":\"value\"}"
|
||||
}
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 10,
|
||||
"completion_tokens": 5,
|
||||
"total_tokens": 15
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`choices[0].message.content` may be either:
|
||||
|
||||
- a JSON string whose contents are valid JSON; or
|
||||
- raw JSON.
|
||||
|
||||
The decoded content is unmarshaled into the caller-provided structured output
|
||||
target. If `usage` is present, prompt, completion, and total token counts are
|
||||
copied into the completion response.
|
||||
|
||||
## Errors And Retries
|
||||
|
||||
The client validates base URL, model, response schema name, response schema
|
||||
JSON, messages, and output target before or during the call.
|
||||
|
||||
Retryable failures:
|
||||
|
||||
- HTTP request failure;
|
||||
- response body read failure;
|
||||
- HTTP `429`;
|
||||
- HTTP `5xx`;
|
||||
- malformed provider response envelope;
|
||||
- missing choices;
|
||||
- missing, empty, or invalid assistant JSON content;
|
||||
- structured-output decode failure.
|
||||
|
||||
Non-retryable provider status codes include non-`429` `4xx` responses.
|
||||
|
||||
Provider error bodies are parsed for `error.message` or `message` when present.
|
||||
Configured API key values and bearer-token values are redacted from returned
|
||||
provider errors.
|
||||
|
||||
## Timeouts And Concurrency
|
||||
|
||||
The configured profile timeout is applied per provider request when greater
|
||||
than zero. Context cancellation is respected.
|
||||
|
||||
The production CLI wraps the provider client with the LLM scheduler. Effective
|
||||
concurrency is described in [LLM runtime internals](../internal/llm.md).
|
||||
|
||||
## Limits
|
||||
|
||||
This contract documents only the fields the implemented client sends and reads.
|
||||
Provider-specific extensions are ignored unless they affect those fields.
|
||||
108
docs/integrations/pkg-promptkit.md
Normal file
108
docs/integrations/pkg-promptkit.md
Normal file
@@ -0,0 +1,108 @@
|
||||
# PromptKit Integration
|
||||
|
||||
Notarius pins
|
||||
[`gitea.maximumdirect.net/eric/promptkit` v0.5.0](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.5.0)
|
||||
as its in-process prompt engine. The upstream
|
||||
[Go package consumer guide](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.5.0/docs/consumers/pkg-promptkit.md)
|
||||
owns the public engine API, and the upstream
|
||||
[format reference](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.5.0/docs/formats.md)
|
||||
owns prompt, profile, and schema file contracts.
|
||||
|
||||
## Supported Boundary
|
||||
|
||||
Notarius relies on the root `promptkit` package to:
|
||||
|
||||
- construct an `Engine` with filesystem-backed prompt, schema, and optional
|
||||
operator and application-fallback profile sources;
|
||||
- prepare one frozen execution from a `RunRequest` with named inline artifacts,
|
||||
variables, a direct session ID, prompt identity, and profile selection, then
|
||||
record credential-redacted details and run that exact execution;
|
||||
- return rendered debug material, validated structured output, selected
|
||||
profile, backend, effective model metadata, and token usage;
|
||||
- register the optional conventional `local` backend through `BackendLocal`,
|
||||
`LocalBackend`, and `WithBackend`;
|
||||
- distinguish structured-output validation failure from execution failure; and
|
||||
- identify a missing explicit profile through `ErrProfileNotFound` and backend
|
||||
admission exhaustion through `ErrCapacityExceeded`.
|
||||
|
||||
The pinned
|
||||
[`BackendLocal`, `LocalBackend`, and `WithBackend` API](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.5.0/backends.go)
|
||||
owns the registration and backend-capacity contract.
|
||||
|
||||
For one completion, the adapter calls `PrepareExecution`, takes a
|
||||
caller-owned `Details` snapshot, and calls `RunPrepared` for that same opaque
|
||||
prepared execution. It defers `Discard` for every unexecuted handle. Explicit
|
||||
profile preflight uses `Engine.InspectProfile`; it does not prepare a synthetic
|
||||
prompt. PromptKit's prepared handle, inspection result, and capacity-error
|
||||
types stay inside the Notarius LLM adapter.
|
||||
|
||||
When a PromptKit profile and runtime override leave `temperature`, `max_tokens`,
|
||||
or `top_p` unset, Notarius leaves that control unset as well. Compatible
|
||||
providers therefore apply their own defaults; an operator that requires a
|
||||
specific sampling value must select it explicitly in the profile or runtime
|
||||
override.
|
||||
|
||||
Notarius does not use PromptKit's optional `ArtifactReader`. It materializes
|
||||
source and reference content itself and supplies owned inline artifacts at the
|
||||
adapter boundary. It also retains responsibility for pipeline retries,
|
||||
scheduling, debug persistence, redaction, profile provenance, and conversion
|
||||
from private model responses into durable domain artifacts.
|
||||
|
||||
Notarius sends one stable effective session through PromptKit's direct session
|
||||
field, which is authoritative for provider session behavior. It also retains
|
||||
the same value as the `session_id` prompt variable for maintained prompt
|
||||
compatibility. The generated identifier is 76 ASCII characters, within
|
||||
PromptKit v0.5.0's 256-code-point session limit. Session IDs are non-secret
|
||||
correlation identifiers and may be exposed to providers and provider
|
||||
observability. The CLI contract owns generation and override behavior.
|
||||
|
||||
Notarius records PromptKit's selected backend ID and effective reasoning
|
||||
setting as optional run-manifest provenance. Endpoint-only profiles have no
|
||||
backend ID. Debug prompt material also retains the selected backend ID and
|
||||
PromptKit's stable lower-case `effective_model_params` JSON, which may include
|
||||
`backend_id`. Notarius production configuration exposes one optional
|
||||
conventional `local` registration. It does not expose a general user-defined
|
||||
PromptKit backend registry. Endpoint-only profiles remain supported unchanged.
|
||||
|
||||
Notarius retains its application-wide scheduled client around the PromptKit
|
||||
adapter. PromptKit may apply a narrower limit for the selected backend;
|
||||
endpoint-only profiles have no such backend limit. The adapter translates
|
||||
PromptKit capacity rejection into the provider-neutral Notarius
|
||||
`ErrLLMCapacityExceeded` contract. It may include the normalized selected
|
||||
backend ID in safe diagnostic context, without exposing PromptKit's capacity
|
||||
error type, and leaves retries to the calling pipeline stage.
|
||||
|
||||
## Profile Sources And Compatibility
|
||||
|
||||
Notarius gives PromptKit the configured operator profile source, registered
|
||||
application fallback profile assets, and optional backend registration through
|
||||
the same construction path for inspection and execution. PromptKit owns the
|
||||
resulting source precedence and strict profile parsing: a matching operator
|
||||
profile is a complete replacement for a fallback or built-in profile, while an
|
||||
invalid matching document fails instead of falling through. The operator
|
||||
configuration and deployment workflow are defined in
|
||||
[Configuration](../config.md#promptkit-profiles) and
|
||||
[Operations](../operations.md#promptkit-profile-deployment).
|
||||
|
||||
Notarius supports this boundary against PromptKit v0.5.0. Its fallback source,
|
||||
prepared-execution, inspection, and typed capacity APIs are used as public
|
||||
upstream contracts; other PromptKit APIs or file-format behavior are not
|
||||
implicitly supported. A dependency upgrade requires reviewing the adapter,
|
||||
profile-source construction, and this compatibility statement against the
|
||||
pinned upstream documentation.
|
||||
|
||||
## Notarius Ownership
|
||||
|
||||
[LLM Runtime Internals](../internal/llm.md) describes how Notarius mounts
|
||||
module assets, maps its transport-neutral completion contract, prepares and
|
||||
executes requests, validates output, records provenance, captures debug
|
||||
material, redacts errors, and preserves timeout ownership.
|
||||
[D&D Module Internals](../internal/dnd.md) owns the embedded
|
||||
`dnd-extraction` fallback profile and the maintained D&D prompt defaults.
|
||||
[Configuration](../config.md#promptkit-profiles) defines how a Notarius
|
||||
configuration selects one PromptKit profile source and optionally registers
|
||||
the conventional local backend.
|
||||
|
||||
PromptKit API or format changes outside this boundary are not implicitly
|
||||
supported. Updating the pinned version requires reviewing the adapter and
|
||||
profile/configuration contracts against the upstream documentation.
|
||||
68
docs/integrations/run-result.md
Normal file
68
docs/integrations/run-result.md
Normal file
@@ -0,0 +1,68 @@
|
||||
# Run Result Receipt
|
||||
|
||||
`notarius run --json` writes this receipt to standard output when a run
|
||||
completes successfully. It lets a subprocess caller discover the physical root
|
||||
of the published output bundle without parsing interactive command output.
|
||||
Command syntax, streams, and exit statuses are defined in the
|
||||
[CLI reference](../cli.md); logical files within the bundle are defined in the
|
||||
[Published JSON Output contract](json-output.md).
|
||||
|
||||
## Schema
|
||||
|
||||
The current schema version is `notarius.run-result.v1`.
|
||||
|
||||
| Field | Required | Meaning |
|
||||
| --- | --- | --- |
|
||||
| `schema_version` | Yes | Exactly `notarius.run-result.v1`. |
|
||||
| `run_id` | Yes | The finalized Notarius run identifier. |
|
||||
| `pipeline_id` | Yes | The effective pipeline identifier. |
|
||||
| `output_directory` | Yes | Absolute path to the published, run-specific output bundle. |
|
||||
| `index_file` | For the production JSON output | Logical path `index.json`; omitted for other output modules. |
|
||||
| `normalized_output_count` | Yes | Number of final normalized outputs. |
|
||||
| `rejected_output_count` | Yes | Number of recorded rejected outputs. |
|
||||
| `warning_count` | Yes | Number of final run warnings. |
|
||||
| `validation_status` | Yes | The final run manifest validation status. |
|
||||
| `debug_directory` | No | Absolute path to the run-specific debug bundle when requested debug capture completed. |
|
||||
|
||||
For the production `json` output module, `index_file` is present only when the
|
||||
completed run returned exactly one logical output file named `index.json`.
|
||||
For another output module, its absence does not indicate a failed run.
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "notarius.run-result.v1",
|
||||
"run_id": "run-1770000000000000000-0123456789abcdef0123456789abcdef",
|
||||
"pipeline_id": "dnd-session",
|
||||
"output_directory": "/work/results/run-1770000000000000000-0123456789abcdef0123456789abcdef",
|
||||
"index_file": "index.json",
|
||||
"normalized_output_count": 6,
|
||||
"rejected_output_count": 2,
|
||||
"warning_count": 1,
|
||||
"validation_status": "rejected"
|
||||
}
|
||||
```
|
||||
|
||||
## Paths And Bundle Discovery
|
||||
|
||||
`output_directory` and `debug_directory`, when present, are lexical absolute
|
||||
paths. They identify the paths used by Notarius and do not resolve symlinks.
|
||||
`output_directory` is the run-specific bundle, not the configured output root.
|
||||
|
||||
The receipt is a summary and discovery document. It does not contain lane
|
||||
descriptors, payloads, manifest data, rejections, warnings, or file contents.
|
||||
For the production JSON output, resolve `index_file` beneath
|
||||
`output_directory`, reject path escapes, and use the
|
||||
[Published JSON Output contract](json-output.md) to discover logical files and
|
||||
lane payloads.
|
||||
|
||||
## Delivery And Compatibility
|
||||
|
||||
Notarius writes the receipt only after the output bundle has been published and
|
||||
any requested debug terminal reporting has completed. Standard output is not
|
||||
transactional: a result-write failure returns a nonzero status and can leave
|
||||
partial bytes. Consumers must ignore standard output unless the process exits
|
||||
with status 0.
|
||||
|
||||
Future versions may add optional fields to this schema. Consumers must tolerate
|
||||
unknown fields. An incompatible field or semantic change requires a new
|
||||
`schema_version` value.
|
||||
@@ -1,120 +1,73 @@
|
||||
# Seriatim Transcript JSON
|
||||
# Seriatim Transcript Input
|
||||
|
||||
This document is the external input contract for the implemented `seriatim`
|
||||
input adapter.
|
||||
This document defines the JSON transcript accepted by the production Seriatim
|
||||
input adapter. It is a source input, not a durable lane artifact. Configure the
|
||||
input adapter through [Configuration](../config.md#production-module-keys).
|
||||
|
||||
## Adapter
|
||||
## Contract Identity
|
||||
|
||||
- Module key: `seriatim`
|
||||
- Document kind: `transcript`
|
||||
- Unit kind: `transcript_segment`
|
||||
- Source format: `application/vnd.seriatim+json`
|
||||
|
||||
The adapter parses raw Seriatim JSON into a generic source document. It owns
|
||||
transcript-specific JSON parsing and metadata mapping; core source and pipeline
|
||||
code stay source-format agnostic.
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Consumer | Seriatim input adapter |
|
||||
| Media type | `application/vnd.seriatim+json` |
|
||||
| Source document kind | `transcript` |
|
||||
| Source-unit kind | `transcript_segment` |
|
||||
|
||||
## Accepted Shape
|
||||
|
||||
The input must be one JSON object with top-level `metadata` and `segments`
|
||||
fields. This covers the maintained minimal fixture and Seriatim intermediate
|
||||
output that provides the same required segment fields.
|
||||
The input is one JSON object containing `metadata` and a non-empty `segments`
|
||||
array. This minimal document is valid:
|
||||
|
||||
```json
|
||||
{
|
||||
"metadata": {
|
||||
"id": "session-alpha",
|
||||
"title": "Synthetic D&D spell session"
|
||||
},
|
||||
"metadata": {"id": "session-alpha"},
|
||||
"segments": [
|
||||
{
|
||||
"id": "seg-001",
|
||||
"id": 1,
|
||||
"start": 0,
|
||||
"end": 4,
|
||||
"speaker": "Aria",
|
||||
"text": "Aria raises her holy symbol and casts Cure Wounds."
|
||||
"text": "Aria casts Cure Wounds."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The maintained example is
|
||||
[examples/seriatim-minimal-transcript.json](../../examples/seriatim-minimal-transcript.json).
|
||||
The maintained two-segment input is
|
||||
[seriatim-minimal-transcript.json](../../examples/seriatim-minimal-transcript.json).
|
||||
|
||||
Top-level metadata entries are preserved. Other segment fields, such as
|
||||
`categories`, are ignored.
|
||||
| Field | Required | Meaning and constraints |
|
||||
| --- | --- | --- |
|
||||
| `metadata` | Yes | JSON object. Its entries become source metadata; no particular metadata key is otherwise required. |
|
||||
| `segments` | Yes | Non-empty array of segment objects, kept in input order. |
|
||||
| `segments[].id` | Yes | Positive canonical decimal integer, supplied as a JSON number or string. IDs must be unique. |
|
||||
| `segments[].start` | Yes | Finite, non-negative numeric value, supplied as a JSON number or string. |
|
||||
| `segments[].end` | Yes | Finite, non-negative numeric value that is not earlier than `start`. |
|
||||
| `segments[].speaker` | Yes | String that is non-empty after trimming. |
|
||||
| `segments[].text` | Yes | String that is non-empty after trimming. Its original text is retained. |
|
||||
|
||||
Multiple top-level JSON values are rejected.
|
||||
Additional top-level and segment fields are ignored. A missing required field,
|
||||
`null` in place of an object or array, malformed JSON, or more than one
|
||||
top-level JSON value is rejected.
|
||||
|
||||
## Validation
|
||||
## Source Identity And References
|
||||
|
||||
The adapter rejects:
|
||||
The adapter chooses the source ID in this order:
|
||||
|
||||
- empty raw input;
|
||||
- malformed JSON;
|
||||
- top-level JSON that is not an object;
|
||||
- missing, null, or non-object `metadata`;
|
||||
- missing, null, non-array, or empty `segments`;
|
||||
- segment values that are not objects;
|
||||
- segment `id` values that are neither strings nor numbers;
|
||||
- non-string `speaker` or `text`;
|
||||
- empty segment IDs;
|
||||
- segment IDs with leading or trailing whitespace;
|
||||
- duplicate segment IDs;
|
||||
- missing or empty `speaker`;
|
||||
- missing, empty, invalid, non-finite, or negative `start`;
|
||||
- missing, empty, invalid, non-finite, or negative `end`;
|
||||
- `end` values before `start`;
|
||||
- missing or empty `text`.
|
||||
1. a non-empty source ID supplied by the calling request;
|
||||
2. non-empty string `metadata.id`;
|
||||
3. non-empty string `metadata.source_id`;
|
||||
4. `seriatim:` followed by the first 16 hexadecimal characters of the raw
|
||||
input’s SHA-256 digest.
|
||||
|
||||
Segment text is preserved as provided, but it must not be empty after trimming.
|
||||
Each accepted segment becomes one source unit whose unit ID is `segments[].id`.
|
||||
Its self-reference uses the derived source ID and the same segment ID for both
|
||||
range endpoints. Artifact contracts use those segment IDs when they cite
|
||||
transcript evidence.
|
||||
|
||||
## Source Mapping
|
||||
## Compatibility
|
||||
|
||||
The adapter maps input to `SourceDocument`:
|
||||
|
||||
- `metadata` becomes `SourceDocument.Metadata`;
|
||||
- `SourceDocument.Kind` is `transcript`;
|
||||
- `SourceDocument.Format` is `application/vnd.seriatim+json`;
|
||||
- `SourceDocument.Digest` is `sha256:<hex>` of the exact raw input bytes.
|
||||
|
||||
`SourceDocument.ID` is selected in this order:
|
||||
|
||||
1. the parse request source ID, after trimming;
|
||||
2. `metadata.id`, when it is a non-empty string after trimming;
|
||||
3. `metadata.source_id`, when it is a non-empty string after trimming;
|
||||
4. `seriatim:<first-16-hex-chars-of-raw-sha256>`.
|
||||
|
||||
Each segment becomes one `SourceUnit`:
|
||||
|
||||
- `segment.id` becomes `SourceUnit.ID`; numeric IDs are converted to their JSON
|
||||
number text, so `1` becomes `"1"`;
|
||||
- `segment.text` becomes `SourceUnit.Text`;
|
||||
- `SourceUnit.Kind` is `transcript_segment`;
|
||||
- `speaker`, `start`, and `end` are stored in source-unit metadata.
|
||||
|
||||
## Metadata Keys
|
||||
|
||||
Seriatim unit metadata uses these keys:
|
||||
|
||||
- `speaker`: string speaker label;
|
||||
- `start`: `json.Number` start value;
|
||||
- `end`: `json.Number` end value.
|
||||
|
||||
The `internal/modules/input/seriatim` package exposes typed accessors for these
|
||||
values.
|
||||
|
||||
## Capabilities
|
||||
|
||||
The module declares these provided capabilities:
|
||||
|
||||
- `source.transcript`
|
||||
- `transcript.speaker`
|
||||
- `transcript.timestamps`
|
||||
|
||||
## Compatibility Limit
|
||||
|
||||
This contract covers only Seriatim transcript JSON with the top-level
|
||||
`metadata` object and `segments` array described here. Broader Seriatim output
|
||||
schemas are compatible only when they provide these required fields with the
|
||||
accepted types.
|
||||
This adapter accepts only the shape described here. A broader Seriatim export
|
||||
is usable only when it supplies this object, metadata, and segment shape with
|
||||
the stated types and constraints. Unknown additional fields do not add
|
||||
Notarius behavior.
|
||||
|
||||
172
docs/internal/cli.md
Normal file
172
docs/internal/cli.md
Normal file
@@ -0,0 +1,172 @@
|
||||
# CLI Internals
|
||||
|
||||
This document describes **internal/cli**, Notarius's production composition
|
||||
root. The [CLI reference](../cli.md) owns command syntax and exit statuses;
|
||||
[Configuration](../config.md) owns configuration values; and
|
||||
[Operations](../operations.md) owns filesystem layout, recovery, and operator
|
||||
procedures.
|
||||
|
||||
## Inputs, Outputs, And Boundaries
|
||||
|
||||
The CLI accepts process arguments, standard streams, and injectable options
|
||||
used by tests and embedding code. It writes command results to the supplied
|
||||
streams and returns a process exit status. For a run, it also creates the
|
||||
production catalog and runtime collaborators, hands a prepared pipeline and
|
||||
source bytes to the framework, and places the logical files returned by the
|
||||
runner.
|
||||
|
||||
It is the only boundary allowed to compose concrete registries, LLM clients,
|
||||
cache/checkpoint collaborators, debug recorders, and physical output paths.
|
||||
Pipeline modules receive interfaces and request data rather than CLI streams or
|
||||
filesystem roots. The [Architecture](../policy/architecture.md) defines this
|
||||
composition-root boundary; [Pipeline Internals](pipeline.md) owns resolution,
|
||||
preparation, and runner mechanics after their inputs are supplied.
|
||||
|
||||
## Dispatch And Configuration Handoff
|
||||
|
||||
The root dispatcher handles help, configuration validation, pipeline listing,
|
||||
and a pipeline run. It normalizes injectable options before dispatch so that a
|
||||
missing production dependency fails as a command error rather than reaching
|
||||
execution.
|
||||
|
||||
Commands that need configuration use one shared loader. The CLI discovers the
|
||||
file, parses it through **internal/core/config**, starts from defaults, applies
|
||||
the file and supported environment overrides, and then validates it for the
|
||||
command. The configured discovery and precedence contract is in
|
||||
[Configuration](../config.md), while the loading and resolution mechanics are
|
||||
in [Configuration Internals](configuration.md).
|
||||
|
||||
Configuration validation without a selected pipeline checks structural
|
||||
configuration only. Validation with a selected pipeline also builds the
|
||||
effective catalog, resolves the pipeline, and verifies every explicit effective
|
||||
PromptKit profile. Selected LLM-backed input, chunk, lane, output, and validator
|
||||
profiles are inspected
|
||||
against the configured PromptKit source and backend registrations without
|
||||
loading a prompt or performing generation, so an unknown or invalid profile
|
||||
fails before pipeline preparation. Credential availability remains an
|
||||
execution-time concern. Pipeline listing validates configuration before
|
||||
returning normalized, sorted identifiers.
|
||||
|
||||
## Production Composition
|
||||
|
||||
The production composition helper allocates every framework registry and the
|
||||
prompt-asset registry, then registers the generic, Seriatim, and D&D module
|
||||
families in that order. The resulting registries provide both the module
|
||||
catalog used for resolution and the concrete constructors used for preparation.
|
||||
Tests may provide a catalog or registries instead; production code must not
|
||||
silently merge an injected partial catalog with production registrations.
|
||||
|
||||
The production LLM factory builds one PromptKit-backed client from the resolved
|
||||
**promptkit.profile_dir** or **promptkit.profile_file** source, attaches the
|
||||
profile-provenance recorder, creates one scheduler from the effective global
|
||||
LLM limit, and wraps the client before it reaches modules. Registration and LLM
|
||||
construction errors are returned before a pipeline is prepared. Configuration
|
||||
field definitions remain in [Configuration](../config.md#promptkit-profiles);
|
||||
the D&D registrar's fallback profile assets and the adapter mechanics remain in
|
||||
[LLM Runtime](llm.md).
|
||||
|
||||
The factory also accepts `LLMRuntimeOverrides`, whose reasoning pointer
|
||||
preserves inherit, replace, and clear states across the composition boundary.
|
||||
Run orchestration constructs this value from the mutually exclusive
|
||||
`--reasoning-effort` and `--clear-reasoning-effort` controls. Absence preserves
|
||||
a nil pointer, replacement is trimmed, and clear uses a non-nil empty string.
|
||||
The same override reaches the one shared production client, checkpoint
|
||||
identity, and debug invocation metadata. Persistent reasoning configuration
|
||||
remains owned by PromptKit profiles; Notarius configuration has no reasoning
|
||||
field.
|
||||
|
||||
## Run Orchestration
|
||||
|
||||
After parsing and validating a run invocation, the CLI performs this ordered
|
||||
handoff:
|
||||
|
||||
1. load and validate configuration, then apply command-level operational
|
||||
overrides;
|
||||
2. create and validate a safe run identity, then allocate a debug bundle only
|
||||
when requested;
|
||||
3. build the effective catalog, resolve requested reference changes, resolve
|
||||
the effective pipeline, and inspect its explicit effective PromptKit
|
||||
profiles;
|
||||
4. materialize external or generated references and record redacted invocation
|
||||
and resolution provenance when debug capture is enabled;
|
||||
5. construct registries, the scheduled LLM client, and prepared modules;
|
||||
6. read the source input once, resolve its effective session from the explicit
|
||||
override or resolved input module and raw bytes, then construct requested
|
||||
checkpoint collaborators and invoke the framework runner with that same
|
||||
value; and
|
||||
7. write the runner's logical output files only after a successful run, then
|
||||
complete the command report and user-facing result.
|
||||
|
||||
Preparation happens before source parsing, so module construction and
|
||||
dependency failures cannot begin stage execution. The CLI also preserves the
|
||||
framework's result and warning information when it writes summaries and the
|
||||
final command result. Detailed state lifecycle, resume handling, and physical
|
||||
path confinement are maintained in [Run State Internals](state.md) and
|
||||
[Operations](../operations.md).
|
||||
|
||||
The CLI owns the versioned generated-session policy and resolves the sole
|
||||
effective value before checkpoint construction. It records that value in the
|
||||
final debug invocation summary when capture is enabled and passes it unchanged
|
||||
to checkpoint identity and `pipeline.RunInput`. The public flag and stability
|
||||
contract are defined by the [CLI reference](../cli.md#run); framework and LLM
|
||||
packages only transport the supplied value.
|
||||
|
||||
For `run --json`, the CLI constructs and encodes its private run-result receipt
|
||||
after a successful runner result is available, before it publishes logical
|
||||
output files. It writes the prepared receipt to standard output only after
|
||||
output publication and requested debug terminalization succeed. A receipt-write
|
||||
failure exits with runtime status 1 and may leave partial standard-output bytes,
|
||||
but the already-published output bundle remains complete and requested debug
|
||||
reporting remains successfully terminalized. The CLI reports a bounded
|
||||
command-owned error and does not repeat terminal reporting. The receipt remains
|
||||
a CLI reporting concern rather than a framework or output-module responsibility;
|
||||
its public contract is the
|
||||
[run-result receipt](../integrations/run-result.md).
|
||||
|
||||
## Failure Mapping And Terminal Reporting
|
||||
|
||||
Argument, flag, and invocation-combination failures are reported to standard
|
||||
error before runtime composition and use the syntax error class. Once an
|
||||
invocation is syntactically valid, configuration loading and validation,
|
||||
resolution, registration, profile checks, reference materialization, module
|
||||
construction, input reads, runner failures, output publication, and requested
|
||||
debug handling use the runtime failure class. The public status numbers and
|
||||
stream contract are defined in the [CLI reference](../cli.md#output-streams-and-exit-statuses).
|
||||
|
||||
When debug capture has been allocated, one command-state value records the
|
||||
known run result. Guarded terminalization writes a success report once, or
|
||||
attempts a failure report and error record once. A persistence failure is
|
||||
reported in addition to the original failure and never replaces it. If a debug
|
||||
path exists, failure output includes that path so the retained diagnostic data
|
||||
is discoverable.
|
||||
|
||||
## Invariants To Preserve
|
||||
|
||||
- Only the CLI composes production implementations and physical runtime roots.
|
||||
- Configuration and resolved composition failures occur before module
|
||||
preparation or source parsing.
|
||||
- A runner's logical files are published only after a successful run.
|
||||
- Production registries and a caller-supplied catalog or registries are
|
||||
alternative composition sources, not an implicit mixture.
|
||||
- A requested debug bundle has one terminal report attempt; its persistence
|
||||
errors supplement rather than obscure the primary command error.
|
||||
- User-facing flags, paths, exit codes, and configuration fields are defined
|
||||
by their public documentation, not duplicated here.
|
||||
|
||||
## Focused Tests
|
||||
|
||||
- **internal/cli/command_contract_test.go** covers dispatch, help, syntax and
|
||||
runtime error classes, discovery, validation, and listing.
|
||||
- **internal/cli/run_contract_test.go** covers the run handoff, publication,
|
||||
debug reporting, and command-owned state collaborators.
|
||||
- **internal/cli/production_contract_test.go** covers registrar composition,
|
||||
production catalog contents, assets, and representative configuration
|
||||
validation.
|
||||
- **internal/cli/reference_contract_test.go** covers CLI reference overrides,
|
||||
origin separation, and materialization boundaries.
|
||||
- **internal/cli/state_hardening_test.go** covers safe run identity, state
|
||||
roots, and failure ordering.
|
||||
|
||||
Run **go test ./internal/cli** after changing command composition or command
|
||||
behavior. Pair it with **go test ./internal/core/config** when the configuration
|
||||
handoff changes.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user