5 Commits

690 changed files with 13287 additions and 51010 deletions

3
.codebase-memory/.gitattributes vendored Normal file
View File

@@ -0,0 +1,3 @@
# Auto-generated by codebase-memory-mcp
# Prevent merge conflicts on compressed artifact
graph.db.zst merge=ours binary

View File

@@ -0,0 +1,11 @@
{
"schema_version": 2,
"commit": "9e4b989e53d65efa614b5eedd8530caed12f60b5",
"indexed_at": "2026-07-27T18:27:55Z",
"project": "home-eric-Workspace-notarius",
"nodes": 6356,
"edges": 35132,
"original_size": 26017792,
"compressed_size": 4386397,
"compression_level": 3
}

Binary file not shown.

2
.gitignore vendored
View File

@@ -2,7 +2,6 @@
notarius notarius
notarius-output notarius-output
workspace/ workspace/
.codebase-memory/
# ---> Go # ---> Go
# If you prefer the allow list template instead of the deny list, see community template: # If you prefer the allow list template instead of the deny list, see community template:
@@ -74,3 +73,4 @@ Icon
Network Trash Folder Network Trash Folder
Temporary Items Temporary Items
.apdisk .apdisk

View File

@@ -1,33 +0,0 @@
when:
- event: tag
steps:
- name: validate-release
image: golang:1.25.5
commands:
- |
set -eu
version="$CI_COMMIT_TAG"
release_note="docs/releases/$version.md"
if ! printf '%s\n' "$version" | grep -E -x 'v(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)' >/dev/null; then
printf '%s\n' "invalid release tag: $version" >&2
exit 1
fi
if [ ! -s "$release_note" ]; then
printf '%s\n' "missing release note: $release_note" >&2
exit 1
fi
if ! grep -F -x "# Notarius $version" "$release_note" >/dev/null; then
printf '%s\n' "release note heading does not match $version" >&2
exit 1
fi
for heading in '## Summary' '## Compatibility' '## Upgrade' '## Changes'; do
if ! grep -F -x "$heading" "$release_note" >/dev/null; then
printf '%s\n' "release note is missing heading: $heading" >&2
exit 1
fi
done
./scripts/check-release-source.sh "$version"

View File

@@ -2,9 +2,8 @@
Notarius is a Go CLI for turning source material into structured artifacts with Notarius is a Go CLI for turning source material into structured artifacts with
configured extraction pipelines. The implemented D&D workflow reads Seriatim configured extraction pipelines. The implemented D&D workflow reads Seriatim
transcript JSON and can produce NPC, location, and item registries; their transcript JSON and can produce scene descriptions, item and currency events,
source-grounded occurrences; scene descriptions, combat turns, enemy events, NPC identities, combat turns, NPC interactions, and spell casts.
and spell casts.
## Quickstart ## Quickstart
@@ -28,20 +27,6 @@ For the complete ordered D&D workflow, use
[its synthetic transcript](examples/dnd-complete-transcript.json). It [its synthetic transcript](examples/dnd-complete-transcript.json). It
demonstrates all implemented D&D lanes and the supporting campaign references. demonstrates all implemented D&D lanes and the supporting campaign references.
## Install A Source Release
Install a pinned source release with Go:
~~~
GOWORK=off go install \
gitea.maximumdirect.net/eric/notarius/cmd/notarius@<tag>
~~~
Replace `<tag>` with a stable release tag such as `vMAJOR.MINOR.PATCH`. The
installed command's diagnostic version is described in the [CLI
reference](docs/cli.md); maintainers preparing a release should follow [Source
Releases](docs/release.md).
## Documentation ## Documentation
- [CLI reference](docs/cli.md) — commands, flags, output streams, and exits. - [CLI reference](docs/cli.md) — commands, flags, output streams, and exits.
@@ -53,8 +38,6 @@ Releases](docs/release.md).
artifact formats. artifact formats.
- [Subprocess consumer guide](docs/consumers/subprocess.md) — invoke Notarius - [Subprocess consumer guide](docs/consumers/subprocess.md) — invoke Notarius
from an orchestrator and consume a published result. from an orchestrator and consume a published result.
- [Complete D&D consumer guide](docs/consumers/dnd-pipeline.md) — run the full
D&D pipeline as a subprocess and discover its structured artifacts.
- [Internal overview](docs/internal/overview.md) — implemented component map - [Internal overview](docs/internal/overview.md) — implemented component map
for maintainers. for maintainers.
- [Developer guide](docs/development.md) — contributor orientation and - [Developer guide](docs/development.md) — contributor orientation and

View File

@@ -1,18 +0,0 @@
Extract Dungeons & Dragons combat-turn artifacts from the supplied transcript.
Include a record only when the transcript establishes that an in-world
participant takes a combat turn or performs a discrete interrupting combat
event. Keep events in transcript chronology; place an interrupting event where
it occurs.
Exclude initiative setup without a turn or combat event, tactical planning,
table talk, rules lookup, hypothetical events, abandoned intentions, recaps
outside the current passage, and downstream consequences. Do not infer combat
events from Dungeons & Dragons rules knowledge. Preserve the session as played
and attribute relevant nonstandard rulings to the GM or table. Unmatched actors
remain permitted.
Treat each record as one turn-level event and keep its supporting transcript
evidence together. Use `turn` for a regular combat turn, `reaction` for an
off-turn reaction, `legendary_action` for a legendary action,
`lair_action` for a lair action, and `other` for another discrete combat
event that does not fit those categories.

View File

@@ -1,12 +0,0 @@
Compact combat grounding is supplied below. It can guide attention and
disambiguation, but it is not evidence. Do not derive an event, subject,
outcome, or source range from either list. The current transcript alone must
directly establish every returned event.
Combat-turn grounding:
{{ input "combat_turns" }}
Named combat-opponent grounding:
{{ input "npc_occurrences" }}

View File

@@ -1,25 +0,0 @@
Extract Dungeons & Dragons enemy events from the supplied combat transcript.
An `engaged` event requires direct establishment that a subject is actively
opposing the party in combat. A `killed`, `fled`, `captured`, or
`incapacitated` event requires explicit establishment of that outcome. An
outcome may share evidence with an engagement, and a later engagement or
outcome for the same subject remains a separate observation. Emit at most one
`engaged` observation for the same subject in this combat scene.
For `killed`, direct death or killing is required. For `fled`, the subject
must explicitly escape, retreat, or leave combat to avoid continued engagement.
For `captured`, the subject must be explicitly taken prisoner or secured
under the party's control. For `incapacitated`, the subject must be explicitly
unable to continue acting without being established as killed or captured.
When the transcript identifies a named NPC, use its normalized registry
spelling. A hostile creature without a registry entry is allowed. For unnamed
individuals or groups, use only the narrowest transcript-grounded label, such
as `Orcs`, `One orc`, or `Remaining orcs`; never invent member names, IDs,
or quantities.
Exclude party members, allies, neutral observers, mentioned-but-absent enemies,
hazards, traps, environmental effects, uncertain allegiance, table talk,
planning, hypotheses, recaps outside this passage, and downstream inference.
Do not infer an engagement or outcome from initiative, turn absence, damage,
defeat, movement, or a scene ending.

View File

@@ -1,53 +0,0 @@
id: dnd.enemy_events
version: "v1"
default_profile: dnd-extraction
inputs:
- name: transcript
required: true
content_type: application/json
- name: players
required: false
content_type: text/plain
- name: party
required: false
content_type: text/plain
- name: glossary
required: false
content_type: text/plain
- name: npc_registry
required: true
content_type: application/json
- name: combat_turns
required: true
content_type: application/json
- name: npc_occurrences
required: true
content_type: application/json
messages:
- role: system
content_file: ./sharedassets/common-dnd-system.md
- role: user
content_file: ./sharedassets/common-dnd-identity.md
- role: user
content_file: ./sharedassets/common-dnd-references.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-transcript-chunk.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-extraction-evidence.md
- role: user
content_file: ./sharedassets/common-dnd-npc-registry.md
- role: user
content_file: ./combat-grounding.md
- role: user
content_file: ./instructions.md
cache_control:
type: ephemeral
output:
format: json
validation_mode: json_schema
schema_path: dnd_enemy_events_llm.v1.json
repair_attempts: 1

View File

@@ -1,30 +0,0 @@
Extract meaningful Dungeons & Dragons item and currency occurrences: discoveries and changes
in party possession established by the transcript. This is an occurrence history,
not an inventory or ledger: do not calculate balances, resolve item identity
across records, or infer ownership that the transcript does not establish.
For every occurrence, use the supplied canonical item `name`. Record a stated
quantity as an integer and leave it null when the transcript does not state
one. Preserve the stated currency denomination through the selected canonical
registry name.
Use `discovered` when the party learns of or encounters an item without
establishing possession. Use `acquired` when the party or a party member gains
possession. Use `lost` when party possession ends through a gift, sale, payment,
theft, abandonment, or destruction not caused by intended use. Use `consumed`
when intended use depletes an expendable item. Monetary spending, purchases, and
payments are always `lost`, not `consumed`. Classify currency as `consumed` only
when the transcript explicitly describes it being physically destroyed or
expended as a non-payment component. Use `transferred` only when possession
moves between two distinct named party members.
Return both `from` and `to` for every occurrence, using `null` when a holder does not
apply. For `discovered`, set both holders to `null`. For `acquired`, set `from`
to `null` and provide `to`; for `lost` and `consumed`, provide `from` and set
`to` to `null`; and for `transferred`, provide both holders. Use `party` only
for collective or unresolved party possession, never for either side of a
transfer. Do not emit a transfer for a gift, sale, or payment outside the party.
Ordinary non-depleting use is not an occurrence. Do not infer acquisition from a
discovery, or discovery from an acquisition: emit both only when each is
independently established.

View File

@@ -1,6 +0,0 @@
Use the supplied item registry only to ground each occurrence. Every record
must use one registry item's canonical `name`; do not invent, rename, merge,
or infer registry items. The registry is not transcript evidence: cite only the
current transcript chunk in `source_refs`.
{{ input "item_registry" }}

View File

@@ -1,45 +0,0 @@
id: dnd.item_occurrences
version: "v1"
default_profile: dnd-extraction
inputs:
- name: transcript
required: true
content_type: application/json
- name: players
required: false
content_type: text/plain
- name: party
required: false
content_type: text/plain
- name: glossary
required: false
content_type: text/plain
- name: item_registry
required: true
content_type: application/json
messages:
- role: system
content_file: ./sharedassets/common-dnd-system.md
- role: user
content_file: ./sharedassets/common-dnd-identity.md
- role: user
content_file: ./sharedassets/common-dnd-references.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-transcript-chunk.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-extraction-evidence.md
- role: user
content_file: ./item-registry.md
- role: user
content_file: ./instructions.md
cache_control:
type: ephemeral
output:
format: json
validation_mode: json_schema
schema_path: dnd_item_occurrences_llm.v1.json
repair_attempts: 1

View File

@@ -1,36 +0,0 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "notarius.dnd.item_occurrences.llm",
"type": "object",
"additionalProperties": false,
"required": ["occurrences"],
"properties": {
"occurrences": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["name", "kind", "quantity", "from", "to", "source_refs"],
"properties": {
"name": {"type": "string"},
"kind": {"type": "string"},
"quantity": {"type": ["integer", "null"]},
"from": {"type": ["string", "null"]},
"to": {"type": ["string", "null"]},
"source_refs": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["start_unit_id", "end_unit_id"],
"properties": {
"start_unit_id": {"type": "integer"},
"end_unit_id": {"type": "integer"}
}
}
}
}
}
}
}
}

View File

@@ -1,12 +0,0 @@
Extract only items established by the provided Dungeons & Dragons transcript.
Include named unique items, concrete reusable item types, and stable unique
designations. Record each currency denomination separately when it is
established, such as copper pieces, silver pieces, gold pieces, or platinum
pieces. Do not use capitalization as an eligibility test. Keep distinct names
and designations as separate candidates; do not merge aliases or invent
qualifiers.
Do not record vague categories such as "loot", "treasure", or "some gear";
generic weapons; inferred properties; quantities; or inferred uniqueness. Omit
uncertain or unsupported items.

View File

@@ -1,40 +0,0 @@
id: dnd.item_registry
version: "v1"
default_profile: dnd-extraction
inputs:
- name: transcript
required: true
content_type: application/json
- name: players
required: false
content_type: text/plain
- name: party
required: false
content_type: text/plain
- name: glossary
required: false
content_type: text/plain
messages:
- role: system
content_file: ./sharedassets/common-dnd-system.md
- role: user
content_file: ./sharedassets/common-dnd-identity.md
- role: user
content_file: ./sharedassets/common-dnd-references.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-transcript-chunk.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-extraction-evidence.md
- role: user
content_file: ./instructions.md
cache_control:
type: ephemeral
output:
format: json
validation_mode: json_schema
schema_path: dnd_item_registry_llm.v1.json
repair_attempts: 1

View File

@@ -1,32 +0,0 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "notarius.dnd.item_registry.llm",
"type": "object",
"additionalProperties": false,
"required": ["items"],
"properties": {
"items": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["name", "source_refs"],
"properties": {
"name": {"type": "string"},
"source_refs": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["start_unit_id", "end_unit_id"],
"properties": {
"start_unit_id": {"type": "integer"},
"end_unit_id": {"type": "integer"}
}
}
}
}
}
}
}
}

View File

@@ -1,9 +0,0 @@
Determine whether candidates identify the same item type or unique designation
using their contextual labels and cited transcript windows. Do not treat nearby
evidence, similar objects, or a shared owner as sufficient.
Keep currency denominations and materially different item types separate. Keep
uncertain aliases separate. Do not infer an item property or uniqueness.
When selecting a canonical display name, choose one supplied candidate name
that is the clearest established designation.

View File

@@ -1,30 +0,0 @@
id: dnd.item_registry.normalize
version: "v1"
default_profile: dnd-extraction
inputs:
- name: candidates
required: true
content_type: application/json
- name: transcript
required: true
content_type: application/json
messages:
- role: system
content_file: ./sharedassets/common-dnd-system.md
- role: user
content_file: ./sharedassets/protocol.md
- role: user
content_file: ./instructions.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/candidates.md
- role: user
content_file: ./sharedassets/transcript-windows.md
cache_control:
type: ephemeral
output:
format: json
validation_mode: json_schema
schema_path: semantic_reconciliation_llm.v1.json
repair_attempts: 1

View File

@@ -1,34 +0,0 @@
Extract Dungeons & Dragons location occurrences from the supplied transcript.
Include an occurrence only when the transcript establishes one supplied
location, one occurrence kind, and a coherent passage supporting both.
Use exactly one kind per occurrence:
- visited: party members are physically present, arrive, remain, or depart;
- planned: the party explicitly proposes, intends, or agrees to future travel;
- recalled: the transcript explicitly recounts an earlier party visit; or
- mentioned: the location is explicitly referenced without stronger support,
including non-actionable speculation or a mere hypothetical reference.
A mere hypothetical or speculative reference is not planned unless the
transcript also establishes an actual proposal, intention, or agreement to
travel. When the hypothetical explicitly names a supplied location, it may be
mentioned.
A generic phrase in the current chunk may refer to a supplied named registry
location only when the chunk's context supports that coreference. It must not
create a registry location, and registry content or provenance must never
replace current-chunk evidence.
For every occurrence, return the exact selector from the location registry:
the canonical `name`, plus an empty `registry_refs` array for a unique name or
the complete ordered `registry_refs` array for a repeated name. Registry ranges
and context identify the location only; they are not occurrence evidence.
For overlapping support, visited outranks planned, recalled, and mentioned;
planned outranks recalled and mentioned; recalled outranks mentioned. A passage
may produce multiple records when it independently establishes separate facts,
such as recalling an earlier visit while planning a return. Omit inferred,
unstated, uncertain, or unsupported places and occurrences. Do not infer a
location or occurrence from surrounding events when the transcript does not
state it. Do not summarize location descriptions.

View File

@@ -1,11 +0,0 @@
A contextual location registry is provided below for identity grounding. It may
be empty. Every record supplies a canonical display name. A name that appears
once is selected with that name and an empty `registry_refs` array. A repeated
name is selected only by copying both its name and its complete, ordered
`registry_refs` array exactly as supplied.
Registry content is context, not occurrence evidence. Do not derive an
occurrence or `source_refs` range from the registry. Do not invent a location
or selector that is absent from it.
{{ input "location_registry" }}

View File

@@ -1,45 +0,0 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "notarius.dnd.location_occurrences.llm",
"type": "object",
"additionalProperties": false,
"required": ["occurrences"],
"properties": {
"occurrences": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["name", "registry_refs", "kind", "source_refs"],
"properties": {
"name": {"type": "string"},
"registry_refs": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["start_unit_id", "end_unit_id"],
"properties": {
"start_unit_id": {"type": "integer", "minimum": 1},
"end_unit_id": {"type": "integer", "minimum": 1}
}
}
},
"kind": {"enum": ["visited", "planned", "recalled", "mentioned"]},
"source_refs": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["start_unit_id", "end_unit_id"],
"properties": {
"start_unit_id": {"type": "integer"},
"end_unit_id": {"type": "integer"}
}
}
}
}
}
}
}
}

View File

@@ -1,13 +0,0 @@
Extract only physical places established by the provided Dungeons & Dragons
transcript that have a stable proper name or unique in-world designation. This
includes named planes, regions, settlements, districts, buildings, rooms,
landmarks, routes, and geographic features.
Do not create a registry location for generic, temporary, relative, or merely
descriptive phrases, including "the room", "the bar", "the hallway",
"outside", and "upstairs". Do not use capitalization as an eligibility test.
Keep aliases and nested places when the transcript identifies them; do not merge
or invent qualifiers for similarly named places.
Exclude people, creatures, objects, organizations, abstract concepts, and
places merely inferred from an event. Omit uncertain or unsupported places.

View File

@@ -1,32 +0,0 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "notarius.dnd.location_registry.llm",
"type": "object",
"additionalProperties": false,
"required": ["locations"],
"properties": {
"locations": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["name", "source_refs"],
"properties": {
"name": {"type": "string"},
"source_refs": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["start_unit_id", "end_unit_id"],
"properties": {
"start_unit_id": {"type": "integer"},
"end_unit_id": {"type": "integer"}
}
}
}
}
}
}
}
}

View File

@@ -1,8 +0,0 @@
Determine whether candidates identify the same physical place using their
contextual labels and cited transcript windows. Do not treat matching names,
nearby evidence, nested places, or generic labels as sufficient.
Keep parent and child places separate, as well as similarly named places and
uncertain aliases.
When selecting a canonical display name, prefer the clearest established name.

View File

@@ -1,30 +0,0 @@
id: dnd.location_registry.normalize
version: "v1"
default_profile: dnd-extraction
inputs:
- name: candidates
required: true
content_type: application/json
- name: transcript
required: true
content_type: application/json
messages:
- role: system
content_file: ./sharedassets/common-dnd-system.md
- role: user
content_file: ./sharedassets/protocol.md
- role: user
content_file: ./instructions.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/candidates.md
- role: user
content_file: ./sharedassets/transcript-windows.md
cache_control:
type: ephemeral
output:
format: json
validation_mode: json_schema
schema_path: semantic_reconciliation_llm.v1.json
repair_attempts: 1

View File

@@ -1,45 +0,0 @@
id: dnd.npc_occurrences
version: "v1"
default_profile: dnd-extraction
inputs:
- name: transcript
required: true
content_type: application/json
- name: players
required: false
content_type: text/plain
- name: party
required: false
content_type: text/plain
- name: glossary
required: false
content_type: text/plain
- name: npc_registry
required: true
content_type: application/json
messages:
- role: system
content_file: ./sharedassets/common-dnd-system.md
- role: user
content_file: ./sharedassets/common-dnd-identity.md
- role: user
content_file: ./sharedassets/common-dnd-references.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-transcript-chunk.md
cache_control:
type: ephemeral
- role: user
content_file: ./sharedassets/common-dnd-extraction-evidence.md
- role: user
content_file: ./sharedassets/common-dnd-npc-registry.md
- role: user
content_file: ./instructions.md
cache_control:
type: ephemeral
output:
format: json
validation_mode: json_schema
schema_path: dnd_npc_occurrences_llm.v1.json
repair_attempts: 1

View File

@@ -1,19 +0,0 @@
Extract the individually identifiable Dungeons & Dragons non-player characters
established by the provided transcript.
Include an in-world non-PC only when the transcript factually establishes a
proper name or a stable, individually distinguishing title or alias. A factual
third-party mention establishes that identity even when the NPC is not
physically present, does not speak, and takes no direct action in this chunk.
Record only the NPC identity and the transcript evidence that establishes it;
do not infer or classify a separate occurrence.
Exclude human players, transcript speakers, and the GM as out-of-world people;
player characters identified by the player or party references; names used only
in hypothetical, speculative, or imagined examples; corrected transcription
mistakes; anonymous or generic roles; indistinguishable crowds or groups;
invented descriptive labels; and temporary summoned creatures or spell effects
without a persistent individual identity.
Preserve observed display spelling. Do not invent a label for an anonymous
creature, crowd, or generic role.

View File

@@ -1,11 +0,0 @@
Determine whether candidates refer to the same individual using their
contextual labels and cited transcript windows. Preserve distinct individuals
even when their names are similar or their contextual descriptions are
identical.
When selecting a canonical display name, prefer a complete, stable proper name
over an abbreviation. Prefer an unadorned proper name over that name plus a
contextual class, role, title, or relationship descriptor unless the transcript
establishes the descriptor as part of the person's name. A longer display name
is not inherently more canonical; for example, do not prefer `Captain Aria`
over `Aria` solely because it includes the contextual title `Captain`.

View File

@@ -1,5 +0,0 @@
id: dnd-extraction
backend: openrouter
model: openai/gpt-5.6-luna
timeout_seconds: 240
service_tier: flex

View File

@@ -1,6 +0,0 @@
Transcript units are the only evidence for extracted events and factual claims.
Every reported factual claim must be supported by cited transcript units. Use
integer `start_unit_id` and `end_unit_id` values from the transcript.
When supporting evidence is non-contiguous, use multiple narrow ranges rather
than a broad range that bridges unrelated conversation.

View File

@@ -1,5 +0,0 @@
You process Dungeons & Dragons gameplay transcripts.
As input, you will receive one or more portions of a transcript. The transcript may contain transcription errors, repeated lines, incomplete sentences, and misheard proper nouns.
Return exactly one JSON object that conforms to the configured response schema, with no explanatory prose.

View File

@@ -1,3 +0,0 @@
One extraction chunk from a Dungeons & Dragons gameplay transcript is provided below. Report and infer only what is within this chunk. Its unit IDs retain their source-wide meaning.
{{ input "transcript" }}

View File

@@ -1,3 +0,0 @@
The complete ordered transcript of this Dungeons & Dragons gameplay session is provided below.
{{ input "transcript" }}

View File

@@ -1,14 +0,0 @@
Extract Dungeons & Dragons spell-cast artifacts from the provided transcript.
Include an actual casting event or an unambiguous declared casting attempt.
Exclude spell mentions, hypothetical plans, rules discussion, and catalog
matches that do not establish a casting event in the transcript.
For every extracted cast, the transcript evidence must collectively support the
in-world caster, the spell, and the fact that the cast or declared attempt
occurred.
Attribute every cast to its in-world caster. Map first-person player speech to
the associated player character, and attribute a spell narrated by the GM to
the in-world creature that casts it. If the caster cannot be resolved, use only
the most specific in-world identity supported by the transcript; do not invent
a name.

View File

@@ -1,6 +0,0 @@
The spell catalog for this extraction is provided below as JSON. Each entry
lists a `canonical_name` and its recognized `aliases`. If the transcript uses
an alias, select that entry's `canonical_name`. Return spell names using the
canonical spelling exactly; never return an alias as a spell name.
{{ input "spell_catalog" }}

View File

@@ -1,3 +0,0 @@
Candidate material:
{{ input "candidates" }}

View File

@@ -1,5 +0,0 @@
Identify only high-confidence duplicate entities among the supplied candidates.
Preserve distinct entities even when their names are similar. Treat contextual descriptions and transcript evidence as supporting material, not as permission to merge ambiguous records.
When several records are duplicates, choose as canonical the candidate with the clearest stable identity. Prefer a complete proper name over an abbreviation, and prefer an unadorned proper name over one with incidental descriptors unless the evidence establishes those descriptors as part of the name. A longer name is not inherently more canonical.

View File

@@ -1,27 +0,0 @@
id: generic.semantic_reconciliation
version: "v1"
inputs:
- name: candidates
required: true
content_type: application/json
- name: transcript
required: true
content_type: application/json
messages:
- role: system
content_file: ./system.md
- role: user
content_file: ./protocol.md
- role: user
content_file: ./instructions.md
cache_control:
type: ephemeral
- role: user
content_file: ./candidates.md
- role: user
content_file: ./transcript-windows.md
output:
format: json
validation_mode: json_schema
schema_path: semantic_reconciliation_llm.v1.json
repair_attempts: 1

View File

@@ -1,7 +0,0 @@
Use only the positive integer `candidate_id` values supplied in the candidate material.
Return a duplicate group only when the evidence supports that every selected candidate describes the same underlying entity. Each group must contain at least two distinct candidate IDs, and its `canonical_candidate_id` must be one of those IDs. A candidate may appear in at most one group.
Omit uncertain matches and candidates that should remain distinct. Do not invent candidates or infer an ID from list position. An empty `duplicate_groups` array is valid.
The response must conform exactly to the selected JSON schema. Return IDs only: do not copy candidate names, evidence, transcript text, source identifiers, or source ranges into the response.

View File

@@ -1,2 +0,0 @@
You reconcile structured records that may describe the same underlying entity.
Follow the supplied protocol and return only the requested structured result.

View File

@@ -1,3 +0,0 @@
Transcript evidence windows:
{{ input "transcript" }}

View File

@@ -1,32 +0,0 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "notarius.generic.semantic_reconciliation.llm",
"title": "notarius_semantic_reconciliation_llm_v1",
"type": "object",
"additionalProperties": false,
"required": ["duplicate_groups"],
"properties": {
"duplicate_groups": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["candidate_ids", "canonical_candidate_id"],
"properties": {
"candidate_ids": {
"type": "array",
"minItems": 2,
"items": {
"type": "integer",
"minimum": 1
}
},
"canonical_candidate_id": {
"type": "integer",
"minimum": 1
}
}
}
}
}
}

View File

@@ -1,15 +0,0 @@
// Package assets exposes embedded LLM-facing content.
package assets
import (
"embed"
"io/fs"
)
//go:embed dnd generic
var embedded embed.FS
// FS returns the embedded read-only asset filesystem.
func FS() fs.FS {
return embedded
}

View File

@@ -1,6 +1,6 @@
# ADR-0004: Package modules by domain, not by stage # ADR-0004: Package modules by domain, not by stage
**Status:** Accepted — its asset-co-location rule is superseded by [ADR-0011](0011-centralize-llm-assets.md); its domain-first module packaging decision remains accepted. **Status:** Accepted
**Date:** 2026-07-13 **Date:** 2026-07-13
## Context ## Context

View File

@@ -1,49 +0,0 @@
# ADR-0010: Use workload-oriented LLM profile defaults
**Status:** Accepted
**Date:** 2026-08-03
## Context
LLM-backed D&D operations share an execution-policy choice, but repeating a
provider or model-named profile on every module binding ties pipeline structure
to a deployment decision. Different environments may require different model,
backend, timeout, or reasoning settings while retaining the same workload.
Notarius also needs a usable default for maintained D&D prompts without making
an operator profile mandatory. That default must remain owned by the D&D
family, while generic LLM infrastructure stays unaware of domain-specific
policy.
## Decision
Pipelines may name one workload-oriented default profile, inherited only by
selected LLM-backed bindings and validators. Binding-level profile IDs remain
intentional exceptions, and the run-wide CLI profile override has highest
precedence.
The D&D family owns an embedded fallback profile named `dnd-extraction`.
Operators may provide a complete profile with the same ID through a PromptKit
filesystem source. PromptKit selects the higher-precedence matching definition;
Notarius does not merge profile documents. Production, development, and local
deployments can therefore use different execution policy behind one unchanged
pipeline ID.
## Alternatives considered
- Repeat a model-named profile on every binding. This makes routine deployment
policy changes noisy and obscures the shared workload intent.
- Require every deployment to install a profile file. This adds configuration
friction and leaves maintained D&D prompts without an application-owned
fallback.
- Put D&D profile policy in generic LLM infrastructure. This breaks domain
ownership and makes generic code depend on one workload.
## Consequences
Pipeline configuration expresses workload intent rather than a specific
provider or model. Operators can replace the complete execution policy without
editing bindings, while binding-level and run-wide exceptions remain available.
Profile changes affect resolved pipeline and checkpoint identity, so they may
intentionally cause work to be recomputed. The D&D fallback becomes a
maintained application execution-policy asset.

View File

@@ -1,69 +0,0 @@
# ADR-0011: Centralize LLM-facing assets in a content-only package
**Status:** Accepted
**Date:** 2026-08-05
## Context
LLM prompts, private response schemas, generic schemas, and fallback profiles
are authored and reviewed as content, but package-local embedding scattered that
content across implementation trees. Finding all of the assets that contribute
to a prompt family required navigating code ownership boundaries rather than a
single discoverable content boundary.
The repository must retain module ownership of prompt semantics, schema
identities, registration, and prompt-cache behavior. Durable artifact schemas
and non-LLM domain data have different compatibility and ownership rules, so
they must not move merely because they are embedded files.
## Decision
LLM-facing content is embedded by the root `assets` package. It is a data-only
dependency leaf: its single `FS() fs.FS` API returns the read-only embedded
filesystem, and the package contains no business logic or internal or PromptKit
dependencies. The accepted import path is
`gitea.maximumdirect.net/eric/notarius/assets`; it makes repository-owned
content available to its consumers, not a public extension contract.
Consumers scope that filesystem to the subtree they own before reading or
registering content. Modules continue to own their manifests, prompt ordering,
private response-schema identity, and registration. Centralizing physical files
does not centralize domain semantics or transfer those responsibilities to the
root package.
The root package contains prompt content, private LLM response schemas, generic
LLM schemas, shared fragments, and fallback profiles. Durable artifact schemas
and non-LLM domain data remain with their current owners. A module fingerprint
is derived from its manifest-selected module and shared files, rather than from
an entire asset tree. The relocation is accepted to cause a one-time checkpoint
invalidation.
This decision supersedes only the physical asset-co-location portion of
ADR-0004's decision that places domain-specific prompt fragments and schemas
within the domain tree. ADR-0004's domain-first packaging and registrar
ownership decisions remain accepted.
## Alternatives Considered
- Keep package-local assets. This preserves physical co-location with code but
makes prompt-author discovery and cross-family review unnecessarily costly.
- Use `internal/llmassets`. This would hide content from legitimate owners
outside the `internal` subtree and would make the root asset boundary depend
on implementation-layer placement.
- Build a behavioral central registry. This would mix content discovery with
prompt selection and registration behavior, moving module semantics into a
shared registry.
- Use runtime filesystem overlays. This would add runtime configuration and
failure modes where compile-time embedded content is sufficient.
## Consequences
Prompt authors can find in-scope LLM content in one top-level tree while module
packages continue to define its meaning and registration. Consumers have an
explicit, narrow dependency on only the content they need. The root package is
intentionally importable but must remain a stable, content-only leaf rather
than becoming a general extension API.
The initial relocation invalidates existing checkpoints once. Later checkpoint
identity changes remain limited to the manifest-selected prompt and shared
content, so unrelated files do not trigger recomputation.

View File

@@ -1,64 +0,0 @@
# ADR-0012: Resolve opaque entity identifiers deterministically
**Status:** Accepted
**Date:** 2026-08-08
## Context
Entity IDs in durable Notarius artifacts are application-owned, deterministic
identifiers. They are useful to artifact consumers, but their hash-based form
does not help a model distinguish entities and would make the model reproduce
an opaque implementation detail. A plain name is likewise insufficient where
multiple supplied records share that name.
The LLM boundary must preserve the typed artifact and durable-schema ownership
of [ADR-0003](0003-typed-interfaces-with-two-zone-data-model.md) and the distinction
between disambiguating references and source evidence in
[ADR-0009](0009-minimal-evidence-grounded-extraction-artifacts.md).
## Decision
Callers present a model with semantic selections: a canonical name when it is
unique in the request, or a contextual descriptor containing the name and
source coordinates when that context is needed to distinguish supplied
records. The model returns only those supplied selections. The caller resolves
each accepted selection against the request-local supplied records and attaches
the opaque application ID deterministically.
Source coordinates are permitted in a selection solely as identity context.
They neither establish an occurrence fact nor replace that occurrence's
current-transcript evidence. A selector must resolve exactly; unknown,
ambiguous, partial, reordered, or otherwise unsafe selections are not mapped.
Where an operation requires a complete grounded artifact, that failure rejects
the complete artifact rather than accepting a partially mapped result.
An explicitly scoped request-local short label is permitted only when a
contextual descriptor would be impractical and the caller can deterministically
map the label within that one request. Such a label is not a durable ID, must
not escape the request boundary, and requires a concrete justification in its
own module contract.
## Alternatives considered
- Ask the model to return durable IDs. This exposes opaque implementation
state, does not improve semantic disambiguation, and makes model output
depend on hash formatting.
- Select by name alone. This cannot safely distinguish same-name records.
- Make request-local labels durable identifiers. This would turn prompt
presentation into a public identity contract and create avoidable migration
pressure.
- Let the model invent identifiers or resolve ambiguity. This makes identity
assignment non-deterministic and weakens validation.
## Consequences
Durable integration contracts retain their exact ID/name pairs while models
operate on readable contextual selections. Calling modules must own selector
construction, exact resolution, ambiguity handling, and conversion into their
durable artifact type; PromptKit and its adapter remain transport-only.
Some ambiguous or invalid proposals are deliberately omitted, retried, or
rejected according to the caller's existing failure policy. Internal candidate
keys may support deterministic request-local mapping, but they are not
model-visible selectors or durable data. This adds local validation work while
keeping identity assignment auditable and stable.

View File

@@ -1,89 +0,0 @@
# ADR-0013: Use request-local candidate handles for semantic reconciliation
**Status:** Accepted
**Date:** 2026-08-09
## Context
Several typed normalize stage modules need semantic reconciliation after
deterministic preprocessing: a model can judge whether source-backed candidates
refer to the same underlying entity, while application code remains responsible
for constructing the normalized artifact. Requiring the model to reproduce a
candidate's full contextual selector makes the response larger and introduces
avoidable formatting, ordering, and transcription failure modes.
Reconciliation must preserve the exact typed artifact boundary established by
[ADR-0003](0003-typed-interfaces-with-two-zone-data-model.md), the domain-neutral
framework and concrete-domain dependency direction established by
[ADR-0004](0004-package-modules-by-domain.md), and the distinction in
[ADR-0009](0009-minimal-evidence-grounded-extraction-artifacts.md) between source
evidence and auxiliary identity context. It also needs a concrete, narrowly
scoped application of the request-local-label exception allowed by
[ADR-0012](0012-resolve-opaque-entity-identifiers-deterministically.md).
## Decision
Semantic reconciliation will be a domain-neutral framework mechanism used by
typed normalize stage modules. A consuming artifact family will retain
ownership of its typed records, identity rules, consolidation policy, durable
IDs, and domain warnings; the framework mechanism will not infer those rules
from arbitrary data.
For each reconciliation request, deterministic code will assign every eligible
model-visible candidate a contiguous, one-based integer handle. The model may
receive the candidate's contextual label, source references, and bounded source
context needed to judge identity, but its structured response will identify
candidates only by those supplied handles. A handle is local to one request,
does not represent entity identity, and must never enter a durable artifact or
be used to derive a durable ID.
The model will propose duplicate groups and select one supplied member of each
group as canonical. Deterministic code will resolve the handles through the
retained request mapping, validate the complete proposal, discard unsafe
groups, and apply only validated groups through typed domain-owned policy. The
model will not synthesize replacement records or directly mutate an artifact.
Every reconciliation prompt will combine a mandatory framework-owned protocol
and safety policy with an explicitly selected semantic policy. The semantic
policy may be the conservative generic policy or a domain-owned policy, but it
cannot replace the shared response protocol or deterministic safety boundary.
## Alternatives considered
- Return durable application IDs. Opaque IDs do not help semantic judgment,
expose application identity mechanics, and make model output reproduce data
that deterministic code already owns.
- Return names alone or copied contextual selectors. Names can be ambiguous,
while reproducing labels and source ranges adds response complexity and
creates mismatches without adding semantic information. Request-local
handles preserve exact selection without either failure mode.
- Ask the model to return synthesized canonical replacement records. This
would transfer typed artifact construction, provenance consolidation, and
durable identity policy to a probabilistic boundary.
- Reconcile reflection-discovered fields or arbitrary JSON. This would weaken
the typed artifact contract and move domain semantics into generic code.
- Hide reconciliation inside extraction or another stage. This would obscure
stage ownership and create cross-stage behavior outside the fixed pipeline;
reconciliation remains explicit normalize-stage behavior.
- Let each domain replace the complete prompt protocol. This would duplicate
safety mechanics and allow domain policy to bypass the common response and
validation contract.
## Consequences
Model responses become smaller and easier to validate, while deterministic
application code retains authority over identity, provenance, ordering, and
typed artifact construction. The framework requires a request-local mapping,
bounded context preparation, a private integer response contract, proposal
assessment, and shared prompt assets. Each consuming artifact family still
requires a typed adapter for its irreducibly domain-specific rules.
Request-local handles are deliberately unsuitable for persistence, logging as
entity identity, checkpoint contracts, or cross-request correlation. Changes
to shared protocol and policy assets must participate in the normal prompt,
schema, and checkpoint fingerprint mechanisms.
The shared mechanism and its initial D&D registry consumers are now
implemented. Current behavior is documented in
[Module Internals](../internal/modules.md#semantic-reconciliation) and
[D&D Module Internals](../internal/dnd.md#semantic-registry-reconciliation).

View File

@@ -1,59 +0,0 @@
# ADR-0014: Use feedback-aware validation retries
**Status:** Accepted
**Date:** 2026-08-26
## Context
Validation can identify a candidate defect after a producer has returned an
otherwise well-formed result. Retrying without the validator's deterministic,
bounded feedback wastes the useful diagnosis, while treating validator
execution failures as defects would ask a producer to repair conditions it
cannot control. The mechanism must preserve typed producer ownership,
checkpoint safety, and the repository's sensitive-data boundaries.
## Decision
The implementation will keep three independent budgets: the producer binding's
outer `retries` budget, PromptKit's structured-output repair budget, and each
validator's execution-retry budget. Validators will run sequentially in their
configured order and aggregate both rejections and execution failures before a
candidate disposition is selected.
A correction-capable producer will provide the exact single LLM response that
controlled its candidate using the `single_response_v1` protocol. A correction
attempt will reconstruct the ordinary request and append exactly two fresh
messages: that latest response as `assistant`, followed by one deterministic
aggregate correction request as `user`. Earlier turns will not accumulate.
Validator failures will not recurse into correction. Pipeline policy owns
terminal disposition, with field-by-field producer overrides over pipeline
defaults: structural failure and semantic rejection default to `fail_run`, and
validator execution failure defaults to `warn_continue`. Validators can report
facts and bounded corrective guidance, but never decide disposition.
Rejected and structurally invalid candidates will not advance. A candidate
allowed through after a validator execution failure will retain explicit
incomplete-validation provenance and will not be checkpointed. Exact response
and correction text remain attempt-local: they are excluded from ordinary
errors, warnings, manifests, receipts, caches, checkpoints, and default debug
summaries.
## Alternatives considered
- Retry every producer after any validation outcome. This conflates producer
defects with validator operational failures and wastes retry budget.
- Let validators decide whether to continue. This would distribute pipeline
disposition policy across validators and undermine consistent defaults.
- Reuse the full prior conversation. Accumulated turns introduce unbounded
prompt growth and make correction behavior depend on incidental history.
- Persist raw responses to simplify diagnosis. Raw model output and correction
guidance may be sensitive and do not belong in durable pipeline records.
## Consequences
The framework gains transport-neutral correction and candidate contracts,
producer capability checks, policy resolution, aggregated validation outcomes,
and conservative checkpoint handling. Prompt construction remains inside the
LLM adapter, while modules remain responsible for accurately exposing the
single response that directly controlled a candidate.

View File

@@ -1,67 +0,0 @@
# ADR-0015: Separate process warnings from quality diagnostics
**Status:** Accepted
**Date:** 2026-08-27
## Context
Notarius currently represents process degradation, incomplete validation,
extraction-quality doubt, and routine normalization with one flat warning
record. That makes ordinary successful runs noisy, loses the framework context
needed to explain a finding, and gives `warning_count` no stable operational
meaning. It also permits output encoders to add a warning after the durable
warning file has already been written.
The application needs one bounded diagnostic model that preserves exact
occurrence counts while retaining only safe, representative samples. Fresh and
resumed logical runs must present the same groups. The model must not alter
validation decisions, retry budgets, rejected-output behavior, or process exit
policy.
## Decision
Warnings are reserved for a completed run that advanced under an allowed
process-level degradation or incomplete-work policy. Extraction-quality signals
are advisories, and routine accepted transformations are observations. A
non-degraded successful run therefore has zero actionable warnings.
Modules and validators own a diagnostic's disposition, category, reason code,
scope, and safe message. The framework adds pipeline origin, including stage,
step, lane, module, validator, and chunk context where applicable. It then
aggregates deterministically by disposition, category, reason code, and full
origin. Chunk context remains on representative samples so equivalent findings
across chunks aggregate together.
Diagnostics carry exact occurrence counts, at most three distinct samples, and
numeric omitted-sample metadata. Producers and validators are bounded to 64
local groups. Final actionable warning groups are bounded without truncation;
the non-warning collection may truncate represented groups while preserving an
exact total occurrence count and explicit truncation metadata.
The public contracts will be versioned: grouped actionable warnings use
`notarius.warnings.v2`, grouped advisories and observations use
`notarius.diagnostics.v1`, and the run receipt uses
`notarius.run-result.v2`. Successful output encoders return logical files or
an error; they do not add post-encoding warnings.
## Alternatives considered
- Keep one warning list and filter only CLI output. This would leave durable
consumers with the same semantically mixed, unbounded contract.
- Map reason codes to severity in a central framework registry. This would
split module-owned meaning between synchronized policy tables and make new
diagnostic meaning implicit.
- Preserve local omission warning records. They inflate visible group counts
and lose exact occurrence semantics.
- Keep output-encoder warnings. A one-pass encoder cannot include those
records consistently in files it has already serialized; a two-phase encoder
protocol is deferred until a demonstrated need exists.
## Consequences
The framework gains validated diagnostic primitives, local collection,
origin-aware aggregation, and versioned durable presentation. Existing warning
transport remains temporarily while producers migrate. Current architecture,
operator, integration, and internal documentation will describe the behavior
only as each implementation step lands; this accepted decision does not claim
that the migration is complete.

View File

@@ -10,7 +10,6 @@ defined in [Operations](operations.md).
~~~ ~~~
notarius help notarius help
notarius --version
notarius run <pipeline-id> --input path/to/source.json [--json] [flags] notarius run <pipeline-id> --input path/to/source.json [--json] [flags]
notarius config validate [--config path/to/config.yml] [--pipeline pipeline-id] [--only lane-a,lane-b] notarius config validate [--config path/to/config.yml] [--pipeline pipeline-id] [--only lane-a,lane-b]
notarius pipelines list [--config path/to/config.yml] [--json] notarius pipelines list [--config path/to/config.yml] [--json]
@@ -19,17 +18,6 @@ notarius pipelines list [--config path/to/config.yml] [--json]
Running Notarius without arguments, or with **help**, **--help**, or **-h**, Running Notarius without arguments, or with **help**, **--help**, or **-h**,
writes the command summary to standard output and exits with status 0. writes the command summary to standard output and exits with status 0.
`notarius --version` is valid only as the sole root argument. It writes exactly
`notarius <version>` followed by a newline to standard output and exits with
status 0. A tagged `go install` build can report its main-module stable tag,
and controlled builds can inject a stable tag at link time through
`gitea.maximumdirect.net/eric/notarius/internal/buildinfo.Override`; an ordinary
unversioned checkout reports `development`. Invalid injected version content is
a runtime error with exit status 1, while extra `--version` arguments are a
syntax error with exit status 2. This diagnostic does not replace the
[run-result](integrations/run-result.md) or artifact contracts for downstream
compatibility decisions.
## run ## run
~~~ ~~~
@@ -51,33 +39,16 @@ pipeline ID and **--input** are required.
| **--debug** | Retain a debug bundle for this run. | | **--debug** | Retain a debug bundle for this run. |
| **--debug-dir path** | Override the debug-bundle root. Requires **--debug**. | | **--debug-dir path** | Override the debug-bundle root. Requires **--debug**. |
| **--only lane-a,lane-b** | Run only the selected comma-separated artifact lanes when that selection is valid for the configured pipeline. | | **--only lane-a,lane-b** | Run only the selected comma-separated artifact lanes when that selection is valid for the configured pipeline. |
| **--llm-profile id** | Highest-precedence configured profile for selected LLM-backed bindings and validators; it replaces binding and [pipeline](config.md#pipelines) defaults. | | **--llm-profile id** | Override effective LLM-capable module bindings with one configured profile. |
| **--session-id id** | Override the generated prompt session identifier with a non-empty value for LLM-backed module calls. | | **--session-id id** | Supply a non-empty prompt session identifier to LLM-backed module calls. |
| **--reasoning-effort value** | Replace the selected PromptKit profile's reasoning effort for every LLM-backed call in this run. The value must be non-empty and the flag may be specified only once. |
| **--clear-reasoning-effort** | Clear reasoning effort inherited from the selected PromptKit profile for every LLM-backed call in this run. |
| **--reference selector=path** | Add or replace a file reference binding. Repeatable. | | **--reference selector=path** | Add or replace a file reference binding. Repeatable. |
| **--without-reference selector** | Remove a configured optional reference binding. Repeatable. | | **--without-reference selector** | Remove a configured optional reference binding. Repeatable. |
**--chunk_cache** accepts only **auto**, **bypass**, or **refresh**. **--chunk_cache** accepts only **auto**, **bypass**, or **refresh**.
**--debug-dir**, **--output-dir**, **--session-id**, and **--debug-dir**, **--output-dir**, **--session-id**, and
**--reasoning-effort**, and **--recompute-step** reject explicit empty values. **--recompute-step** reject explicit empty values. **--recompute-step**
**--reasoning-effort** and **--clear-reasoning-effort** are mutually exclusive. requires **--resume**; checkpoint requirements and reuse behavior are
When neither is present, reasoning effort comes from the selected PromptKit documented in [Operations](operations.md).
profile. These controls apply to the shared run client, including retries and
LLM-backed validators, and do not modify configuration or profile files.
Persistent reasoning settings remain a PromptKit profile concern.
**--recompute-step** requires **--resume**; checkpoint requirements and reuse
behavior are documented in [Operations](operations.md).
Every run uses one effective prompt session. Without **--session-id**, Notarius
generates a stable `notarius:v1:` identifier from the trimmed resolved input
module key and the input file's exact raw bytes. The same module and bytes
therefore produce the same identifier, regardless of pipeline, references,
profile, retries, or run settings. An explicit non-empty value replaces that
default. Session identifiers are visible to providers; they are non-secret
correlation identifiers, not credential storage. See
[Operations](operations.md#operational-limits) for privacy and workflow
guidance.
### Reference selectors ### Reference selectors
@@ -103,14 +74,11 @@ names, requiredness, and configured bindings are part of the
Without **--json**, standard output contains the completed pipeline ID, counts Without **--json**, standard output contains the completed pipeline ID, counts
of normalized and rejected outputs, and the output directory. A debug-enabled of normalized and rejected outputs, and the output directory. A debug-enabled
run also prints its debug-bundle path to standard output. A successful run with run also prints its debug-bundle path to standard output. A successful run with
actionable process warnings reports their group and occurrence counts to warnings reports the warning count to standard error. The published JSON bundle
standard error. When the selected output module publishes `warnings.json`, the
summary also reports that durable file's path. Advisory and observation findings
do not produce a warning line. The published JSON bundle
is defined by the [JSON output contract](integrations/json-output.md). is defined by the [JSON output contract](integrations/json-output.md).
With **--json**, successful standard output is exactly one With **--json**, successful standard output is exactly one
`notarius.run-result.v2` JSON document followed by a newline, with no `notarius.run-result.v1` JSON document followed by a newline, with no
human-oriented status or debug-path line. Its fields and compatibility policy human-oriented status or debug-path line. Its fields and compatibility policy
are defined by the [run-result contract](integrations/run-result.md). A caller are defined by the [run-result contract](integrations/run-result.md). A caller
must check for exit status 0 before decoding this output; a failed write can must check for exit status 0 before decoding this output; a failed write can
@@ -174,7 +142,7 @@ go run ./cmd/notarius pipelines list \
Successful commands write their primary result to standard output. Warnings and Successful commands write their primary result to standard output. Warnings and
errors are written to standard error. errors are written to standard error.
For **run --json**, actionable process warnings remain on standard error and standard output is a For **run --json**, warnings remain on standard error and standard output is a
machine-readable success result only. Syntax and runtime diagnostics remain on machine-readable success result only. Syntax and runtime diagnostics remain on
standard error. Parse the result only after the process exits with status 0. standard error. Parse the result only after the process exits with status 0.

View File

@@ -1,7 +1,7 @@
# Configuration # Configuration
This is the canonical reference for Notarius configuration. Configuration files This is the canonical reference for Notarius configuration. Configuration files
are YAML and must declare version 4. They select pipelines and their modules; are YAML and must declare version 3. They select pipelines and their modules;
the [CLI reference](cli.md) owns invocation syntax, and the [CLI reference](cli.md) owns invocation syntax, and
[Operations](operations.md) owns run-state procedures. [Operations](operations.md) owns run-state procedures.
@@ -32,8 +32,7 @@ override the fields listed below.
single-lane Seriatim-to-spell pipeline. single-lane Seriatim-to-spell pipeline.
- [Complete D&D configuration](../examples/dnd-complete.config.yml) uses - [Complete D&D configuration](../examples/dnd-complete.config.yml) uses
ordered steps, all implemented D&D lanes, generated references, state ordered steps, all implemented D&D lanes, generated references, state
settings, bounded LLM concurrency, and the maintained settings, and bounded LLM concurrency.
[operator profile](../examples/profiles/dnd-extraction.yml).
Use these complete files as starting points rather than combining the Use these complete files as starting points rather than combining the
illustrative fragments in this reference. illustrative fragments in this reference.
@@ -46,8 +45,8 @@ other than **version** is optional.
| Field | Type | Default | Rules | | Field | Type | Default | Rules |
| --- | --- | --- | --- | | --- | --- | --- | --- |
| **version** | integer | none | Required; must be 4. | | **version** | integer | none | Required; must be 3. |
| **promptkit** | object | none | Profile source and optional local-backend configuration. | | **scriptorium** | object | none | Profile source configuration. |
| **pipelines** | map | empty | Maps pipeline IDs to pipeline definitions. | | **pipelines** | map | empty | Maps pipeline IDs to pipeline definitions. |
| **concurrency** | object | see below | Global LLM and extraction limits. | | **concurrency** | object | see below | Global LLM and extraction limits. |
| **output** | object | see below | Published output settings. | | **output** | object | see below | Published output settings. |
@@ -58,7 +57,7 @@ Built-in defaults are:
| Field | Default | | Field | Default |
| --- | --- | | --- | --- |
| **concurrency.total_llm** | 16 | | **concurrency.total_llm** | 1 |
| **concurrency.stage_workers.extract** | Effective **total_llm** | | **concurrency.stage_workers.extract** | Effective **total_llm** |
| **output.directory** | **./notarius-output** | | **output.directory** | **./notarius-output** |
| **cache.chunk_plans.mode** | **auto** | | **cache.chunk_plans.mode** | **auto** |
@@ -70,89 +69,20 @@ Built-in defaults are:
An empty cache directory in YAML deliberately selects the corresponding An empty cache directory in YAML deliberately selects the corresponding
per-user root. An explicit empty output or debug directory is invalid. per-user root. An explicit empty output or debug directory is invalid.
## PromptKit Profiles ## Scriptorium Profiles
The optional **promptkit** object selects one source of profile definitions and The optional **scriptorium** object selects one source of profile definitions:
may register one conventional local OpenAI-compatible backend:
~~~yaml
version: 4
promptkit:
profile_dir: ./profiles
# profile_file: ./profiles.yml
local_backend:
endpoint: http://localhost:8000/v1
concurrency_limit: 2
~~~
| Field | Type | Rules | | Field | Type | Rules |
| --- | --- | --- | | --- | --- | --- |
| **profile_dir** | string | Non-empty directory containing profile files. | | **profile_dir** | string | Non-empty directory containing profile files. |
| **profile_file** | string | Non-empty profile file. | | **profile_file** | string | Non-empty profile file. |
| **local_backend** | object | Optional registration for the conventional PromptKit backend ID **local**. |
| **local_backend.endpoint** | string | Required when **local_backend** is present; absolute HTTP or HTTPS URL with a host. |
| **local_backend.concurrency_limit** | integer | Optional non-negative limit; defaults to 0. |
Set at most one of **profile_dir** and **profile_file**. Relative values use Set at most one of these fields. Profile IDs used by a binding must be available
the process working directory, not the configuration file's directory. The from the selected Scriptorium profile source when the pipeline is resolved.
complete example's `./examples/profiles/dnd-extraction.yml` value is therefore Keep credentials out of this file: configure a profile to read its credential
valid when Notarius is launched from the repository root; use an absolute path from an environment variable, then set that environment variable only in the
for services and containers. run environment.
An operator source is optional. For a requested ID, PromptKit checks the
configured operator source first, then Notarius's embedded fallback profiles,
then its own built-in catalog. A matching profile is complete: it replaces a
lower-precedence definition rather than merging with it. The maintained
[`dnd-extraction` operator profile](../examples/profiles/dnd-extraction.yml)
is a secret-free deployment artifact; production, development, and local
deployments can each provide a complete definition with that same workload ID.
Use workload-oriented IDs for new profiles instead of model names.
[Operations](operations.md#promptkit-profile-deployment) owns the deployment
workflow and credential-handling guidance.
When **local_backend** is present, its endpoint is trimmed and must use HTTP or
HTTPS case-insensitively, be absolute, and have a non-empty host. URL paths are
allowed. User information, queries, and fragments are rejected. A zero
**concurrency_limit** leaves the local backend unrestricted inside PromptKit;
a positive value limits simultaneous local generations. The application-wide
**concurrency.total_llm** limit still applies in both cases. Neither local
backend field has an environment override. Omitting **local_backend** registers
nothing and preserves existing built-in and endpoint-only profile behavior.
A file-backed PromptKit profile selects the registration by its case-sensitive
backend ID:
~~~yaml
id: local-summary
backend: local
model: example-model
~~~
Keep credentials out of the local-backend object. A PromptKit profile may name
its credential environment variable through `api_key_env`; set that variable
only in the run environment. PromptKit owns the
[pinned profile-file format](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.9.0/docs/formats.md),
including `base_profile` inheritance. Notarius passes profiles through without
merging them. Filesystem profiles cannot express PromptKit's in-memory
`APIKeyRequired` setting; an unset `api_key_env` is optional and may reach the
provider without authorization.
The [PromptKit upstream boundary](integrations/pkg-promptkit.md) identifies the
supported package API, and [Operations](operations.md#operational-limits)
describes the effective concurrency layers.
`notarius config validate --pipeline <id>` resolves the selected pipeline and
inspects every explicit effective profile without contacting a provider or
requiring credential values. It rejects absent, malformed, or incompatible
profiles before a run prepares modules. Credential availability is checked only
when a generation is prepared.
## Migrating Version 3 Configuration
Version 3 files are not decoded or rewritten. Change **version: 3** to
**version: 4** and rename the top-level **scriptorium:** section to
**promptkit:**. Version 4 decoding is strict, so a remaining **scriptorium**
field is rejected as unknown.
## Operational Environment Variables ## Operational Environment Variables
@@ -208,7 +138,6 @@ Each **pipelines** entry has a unique, non-empty ID and the following shape:
~~~yaml ~~~yaml
pipelines: pipelines:
dnd-session: dnd-session:
llm_profile: dnd-extraction
input: seriatim input: seriatim
chunk: generic chunk: generic
output: json output: json
@@ -221,9 +150,6 @@ pipelines:
| Field | Type | Default | Rules | | Field | Type | Default | Rules |
| --- | --- | --- | --- | | --- | --- | --- | --- |
| **llm_profile** | string | none | Optional non-empty default PromptKit profile ID for selected LLM-backed bindings and validators. An explicitly present blank value is invalid. |
| **structured_output_repair_attempts** | integer | prompt-owned (1 in maintained production prompts) | Optional structural-repair limit from 0 through 3 for selected LLM-backed bindings and validators. Omission leaves the prompt's declared policy in control; explicit 0 disables structural repair at that scope. |
| **validation_policy** | object | see below | Optional terminal policy defaults for producer validation. Its fields inherit independently into chunk, extract, merge, and normalize bindings. |
| **input** | module binding | none | Required. | | **input** | module binding | none | Required. |
| **chunk** | module binding | **generic** | Optional. | | **chunk** | module binding | **generic** | Optional. |
| **output** | module binding | **json** | Optional. | | **output** | module binding | **json** | Optional. |
@@ -237,48 +163,6 @@ needs a unique non-empty **id**, an **artifacts** map, and may have
**references**. A lane ID must not appear more than once in a pipeline, **references**. A lane ID must not appear more than once in a pipeline,
including across explicit steps. including across explicit steps.
For each selected LLM-backed binding or validator, profile selection occurs
after module, validator, and `--only` lane selection. It uses the
run-level **--llm-profile** value first, then the binding's **llm_profile**,
then the pipeline's **llm_profile**, and finally the PromptKit default.
Deterministic bindings do not receive these defaults or run overrides.
Structural output repair is resolved after module, validator, and `--only` lane
selection. An object's **structured_output_repair_attempts** value takes
precedence over the pipeline value; otherwise, an LLM-backed binding or
validator inherits the pipeline value. If both are omitted, PromptKit uses the
prompt's declared repair policy. The value must be an integer from 0 through 3;
explicit `null` and non-integer values are invalid. An explicit value on a
deterministic binding or validator is invalid, while a pipeline value simply
does not apply to deterministic selections.
`validation_policy` controls terminal disposition for one complete producer
attempt and validator chain. It may appear on a pipeline or a **chunk**,
**extract**, **merge**, or **normalize** module binding; input, output, and
validator bindings reject it. Every field is optional and resolves in binding,
pipeline, then application-default order:
| Field | Values | Default |
| --- | --- | --- |
| **producer_structural_failure** | **fail_run**, **reject_output** | **fail_run** |
| **semantic_rejection** | **fail_run**, **reject_output** | **fail_run** |
| **validator_failure** | **warn_continue**, **fail_run** | **warn_continue** |
The policy object and its fields must be non-null, and unknown fields are
rejected. A deterministic producer may not explicitly set
**producer_structural_failure** on its binding, although a pipeline-level
default remains valid for pipelines that include LLM-backed producers.
After the producer binding's retry budget is exhausted, an invalid structured
response uses **producer_structural_failure**. One or more semantic validator
rejections use **semantic_rejection**; rejection takes precedence over an
exhausted validator failure or skip. With no rejection, an exhausted validator
failure or skip uses **validator_failure**. `reject_output` records the
terminal rejection without advancing that candidate. `warn_continue` is valid
only for validator execution failure: it advances a structurally valid,
otherwise unrejected result with incomplete-validation provenance and without
making it reusable checkpoint state.
A lane has these fields: A lane has these fields:
| Field | Type | Default | Rules | | Field | Type | Default | Rules |
@@ -306,7 +190,7 @@ Use an object for fields:
~~~yaml ~~~yaml
extract: extract:
module: dnd/spells module: dnd/spells
llm_profile: dnd-extraction llm_profile: gemini-2-flash
retries: 2 retries: 2
references: references:
spell_catalog: ./dnd-spell-catalog.json spell_catalog: ./dnd-spell-catalog.json
@@ -315,33 +199,17 @@ extract:
| Binding field | Type | Default | Rules | | Binding field | Type | Default | Rules |
| --- | --- | --- | --- | | --- | --- | --- | --- |
| **module** | string | none | Required for an object binding. Must be a registered compatible key. | | **module** | string | none | Required for an object binding. Must be a registered compatible key. |
| **llm_profile** | string | none | Optional non-empty PromptKit profile ID for an LLM-backed binding. It overrides the pipeline default unless the run supplies **--llm-profile**. | | **llm_profile** | string | none | Optional non-empty Scriptorium profile ID. |
| **structured_output_repair_attempts** | integer | pipeline or prompt-owned (1 in maintained production prompts) | Optional structural-repair limit from 0 through 3 for an LLM-backed binding. It overrides the pipeline value; explicit 0 disables structural repair. | | **retries** | integer | 0 | Non-negative additional attempts for chunk, extract, merge, and normalize bindings. |
| **validation_policy** | object | pipeline or application defaults | Optional field-by-field terminal-policy override for a chunk, extract, merge, or normalize binding. |
| **retries** | integer | 0 | Non-negative additional complete producer attempts for chunk, extract, merge, and normalize bindings. This single budget covers operational errors, invalid structured output, module-requested normalization retry, and semantic correction. |
| **options** | object | none | Must satisfy the selected module. | | **options** | object | none | Must satisfy the selected module. |
| **references** | map | none | Valid only on chunk, extract, merge, and normalize bindings. | | **references** | map | none | Valid only on chunk, extract, merge, and normalize bindings. |
| **validators** | list | production chain | Valid only on chunk, extract, merge, and normalize bindings. | | **validators** | list | production chain | Valid only on chunk, extract, merge, and normalize bindings. |
Omitting **validators** uses the registered chain. **validators: []** selects Omitting **validators** uses the registered chain. **validators: []** selects
an empty chain; a non-empty list replaces the chain in the listed order. an empty chain; a non-empty list replaces the chain in the listed order.
Validator bindings accept only **module**, **llm_profile**, Validator bindings accept only **module**, **llm_profile**, and **options**.
**structured_output_repair_attempts**, **retries**, and **options**. Their They reject **references**, **retries**, and nested **validators**. Deterministic
**retries** value is a non-negative additional validator-execution budget and validators reject an explicit **llm_profile**.
is valid only when the selected validator is LLM-backed. A validator retry
rechecks the same immutable candidate; it never regenerates the producer.
They reject
**validation_policy**, **references**, and nested **validators**. Deterministic
validators reject explicit **llm_profile** and
**structured_output_repair_attempts**.
Deterministic module bindings also reject those explicit fields.
An LLM-backed chunk, extract, merge, or normalize producer with both a
non-empty validator chain and positive **retries** must declare the supported
single-response correction capability. Preparation rejects a configuration
that could require semantic correction from a producer that cannot provide an
exact prior response. A deterministic producer, or an LLM attempt that did
not make a model call, cannot consume a semantic retry after rejection.
The **json** output module accepts optional **include_chunk_map** and The **json** output module accepts optional **include_chunk_map** and
**evidence_context** settings: **evidence_context** settings:
@@ -355,7 +223,7 @@ output:
enabled: true enabled: true
window_units: 3 window_units: 3
lanes: lanes:
- npc-registry - npcs
- spells - spells
~~~ ~~~
@@ -376,8 +244,7 @@ Unknown outer or nested option fields are rejected, as are incompatible YAML
types. The allowlist remains valid when a run uses lane filtering: a configured types. The allowlist remains valid when a run uses lane filtering: a configured
lane that is not active for that invocation simply contributes no evidence. lane that is not active for that invocation simply contributes no evidence.
Evidence publication is opt-in because it can persist source text and metadata. Evidence publication is opt-in because it can persist source text and metadata.
When enabled, it publishes the selected source-unit excerpt defined by the Its payload contract is [Published Evidence Context](integrations/evidence-context.md).
[Published Evidence Context contract](integrations/evidence-context.md).
## References And Ordered Handoffs ## References And Ordered Handoffs
@@ -390,15 +257,15 @@ step:
steps: steps:
- id: describe-session - id: describe-session
artifacts: artifacts:
npc-registry: npcs:
extract: dnd/npc-registry extract: dnd/npcs
normalize: dnd/npc-registry normalize: dnd/npcs
- id: extract-events - id: extract-events
references: references:
npc_registry: npcs:
artifact: artifact:
step: describe-session step: describe-session
lane: npc-registry lane: npcs
artifacts: artifacts:
spells: spells:
extract: dnd/spells extract: dnd/spells
@@ -428,38 +295,14 @@ selected target declares them:
| **players** | Optional text player context. | | **players** | Optional text player context. |
| **glossary** | Optional text campaign glossary. | | **glossary** | Optional text campaign glossary. |
| **spell_catalog** | Optional JSON spell-catalog overlay for spell extraction and normalization. See [spell-catalog overlays](integrations/dnd-spell-catalog-overlays.md). | | **spell_catalog** | Optional JSON spell-catalog overlay for spell extraction and normalization. See [spell-catalog overlays](integrations/dnd-spell-catalog-overlays.md). |
| **location_registry** | Required normalized location registry for location-occurrence extraction and normalization. | | **npcs** | Normalized NPC registry. Optional for spells and combat turns; required for NPC interactions. |
| **item_registry** | Required normalized item registry for item-occurrence extraction and normalization. | | **scene_descriptions** | Required normalized scene-description artifact for combat-turn extraction. |
| **npc_registry** | Normalized NPC registry. Optional for spells and combat turns; required for NPC occurrences and enemy-event extraction and normalization. |
| **scene_descriptions** | Required normalized scene-description artifact for combat-turn and enemy-event extraction. |
| **combat_turns** | Required normalized combat-turn artifact for enemy-event extraction. |
| **npc_occurrences** | Required normalized NPC-occurrence artifact for enemy-event extraction. |
Registry-backed occurrence and enemy-event artifact slots have the following
exact binding contracts. Durable semantics and wire shapes remain in their
[NPC occurrence](integrations/dnd-npc-occurrence-artifacts.md),
[location occurrence](integrations/dnd-location-occurrence-artifacts.md),
[item occurrence](integrations/dnd-item-occurrence-artifacts.md), and
[enemy-event](integrations/dnd-enemy-event-artifacts.md) contracts.
| Slot | Accepted artifact kind | Media type | Maximum size | Required stage |
| --- | --- | --- | --- | --- |
| `npc_registry` | `dnd/npc-registry` | `application/json` | 1,048,576 bytes | extract and normalize |
| `scene_descriptions` | `dnd/scene-description-list` | `application/json` | 1,048,576 bytes | extract only |
| `combat_turns` | `dnd/combat-turn-list` | `application/json` | 1,048,576 bytes | extract only |
| `npc_occurrences` | `dnd/npc-occurrence-list` | `application/json` | 1,048,576 bytes | extract only |
| `location_registry` | `dnd/location-registry` | `application/json` | 1,048,576 bytes | location-occurrence extract and normalize |
| `item_registry` | `dnd/item-registry` | `application/json` | 1,048,576 bytes | item-occurrence extract and normalize |
Scene descriptions accept **party**, **players**, and **glossary**, but not Scene descriptions accept **party**, **players**, and **glossary**, but not
**roster**. NPC occurrences require **npc_registry** for both extraction and **roster**. NPC interactions require **npcs** for both extraction and
normalization. Combat turns require **scene_descriptions** for extraction; the normalization. Combat turns require **scene_descriptions** for extraction; the
normalized combat-turn module may use optional **npc_registry**. Location occurrences normalized combat-turn module may use optional **npcs**. The complete example
require **location_registry** for extraction and normalization. Item occurrences require shows generated **npcs** and **scene_descriptions** bindings.
**item_registry** for extraction and normalization. Enemy-event extraction requires all
four of its JSON artifact slots; its normalizer requires **npc_registry**.
The [complete example](../examples/dnd-complete.config.yml) shows the ordered
generated bindings.
## Production Module Keys ## Production Module Keys
@@ -467,29 +310,18 @@ generated bindings.
| --- | --- | | --- | --- |
| Input | **seriatim** | | Input | **seriatim** |
| Chunk | **generic**, **dnd/scenes** | | Chunk | **generic**, **dnd/scenes** |
| Extract | **dnd/spells**, **dnd/npc-registry**, **dnd/combat-turns**, **dnd/item-occurrences**, **dnd/item-registry**, **dnd/npc-occurrences**, **dnd/scene-descriptions**, **dnd/enemy-events**, **dnd/location-registry**, **dnd/location-occurrences** | | Extract | **dnd/spells**, **dnd/npcs**, **dnd/combat-turns**, **dnd/item-events**, **dnd/npc-interactions**, **dnd/scene-descriptions** |
| Merge | **appendorder** | | Merge | **appendorder** |
| Normalize | **noop**, **dnd/spells**, **dnd/npc-registry**, **dnd/combat-turns**, **dnd/item-occurrences**, **dnd/item-registry**, **dnd/npc-occurrences**, **dnd/scene-descriptions**, **dnd/enemy-events**, **dnd/location-registry**, **dnd/location-occurrences** | | Normalize | **noop**, **dnd/spells**, **dnd/npcs**, **dnd/combat-turns**, **dnd/item-events**, **dnd/npc-interactions**, **dnd/scene-descriptions** |
| Output | **json** | | Output | **json** |
`dnd/scenes` and every D&D extractor are `llm_backed`. The
`dnd/npc-registry`, `dnd/location-registry`, and `dnd/item-registry`
normalizers are also `llm_backed` for bounded duplicate proposals; every other
D&D normalizer is `deterministic`. LLM-backed bindings use the effective
[PromptKit profile](#promptkit-profiles). The complete example binds each
registry in an earlier step before its occurrence consumer.
The D&D artifact contracts define each emitted schema: The D&D artifact contracts define each emitted schema:
[spells](integrations/dnd-spell-artifacts.md), [spells](integrations/dnd-spell-artifacts.md),
[NPC registry](integrations/dnd-npc-registry-artifacts.md), [NPCs](integrations/dnd-npc-artifacts.md),
[NPC occurrences](integrations/dnd-npc-occurrence-artifacts.md), [NPC interactions](integrations/dnd-npc-interaction-artifacts.md),
[combat turns](integrations/dnd-combat-turn-artifacts.md), [combat turns](integrations/dnd-combat-turn-artifacts.md),
[item registry](integrations/dnd-item-registry-artifacts.md), [item events](integrations/dnd-item-event-artifacts.md), and
[item occurrences](integrations/dnd-item-occurrence-artifacts.md), [scene descriptions](integrations/dnd-scene-description-artifacts.md).
[scene descriptions](integrations/dnd-scene-description-artifacts.md),
[enemy events](integrations/dnd-enemy-event-artifacts.md),
[location registry](integrations/dnd-location-registry-artifacts.md), and
[location occurrences](integrations/dnd-location-occurrence-artifacts.md).
## Production Validator Keys And Default Chains ## Production Validator Keys And Default Chains
@@ -499,15 +331,11 @@ Available validator keys are:
| --- | --- | | --- | --- |
| Generic | **generic/always_accept**, **generic/always_reject**, **generic/valid_json**, **generic/valid_json_schema** | | Generic | **generic/always_accept**, **generic/always_reject**, **generic/valid_json**, **generic/valid_json_schema** |
| Spells | **extract/dnd/spells/shape**, **extract/dnd/spells/catalog**, **extract/dnd/spells/source_refs**, **extract/dnd/spells/source_relatedness** | | Spells | **extract/dnd/spells/shape**, **extract/dnd/spells/catalog**, **extract/dnd/spells/source_refs**, **extract/dnd/spells/source_relatedness** |
| NPC registry | **extract/dnd/npc-registry/shape**, **extract/dnd/npc-registry/source_refs**, **extract/dnd/npc-registry/source_relatedness**, **normalize/dnd/npc-registry/identity** | | NPCs | **extract/dnd/npcs/shape**, **extract/dnd/npcs/source_refs**, **extract/dnd/npcs/source_relatedness**, **normalize/dnd/npcs/identity** |
| Combat turns | **extract/dnd/combat-turns/shape**, **extract/dnd/combat-turns/source_refs**, **extract/dnd/combat-turns/source_relatedness**, **normalize/dnd/combat-turns/invariants** | | Combat turns | **extract/dnd/combat-turns/shape**, **extract/dnd/combat-turns/source_refs**, **extract/dnd/combat-turns/source_relatedness**, **normalize/dnd/combat-turns/invariants** |
| Item occurrences | **extract/dnd/item-occurrences/shape**, **extract/dnd/item-occurrences/registry**, **extract/dnd/item-occurrences/source_refs**, **extract/dnd/item-occurrences/source_relatedness**, **normalize/dnd/item-occurrences/invariants** | | Item events | **extract/dnd/item-events/shape**, **extract/dnd/item-events/source_refs**, **extract/dnd/item-events/source_relatedness**, **normalize/dnd/item-events/invariants** |
| Item registry | **extract/dnd/item-registry/shape**, **extract/dnd/item-registry/source_refs**, **extract/dnd/item-registry/source_relatedness**, **normalize/dnd/item-registry/identity** | | NPC interactions | **extract/dnd/npc-interactions/shape**, **extract/dnd/npc-interactions/registry**, **extract/dnd/npc-interactions/source_refs**, **extract/dnd/npc-interactions/source_relatedness**, **normalize/dnd/npc-interactions/invariants** |
| NPC occurrences | **extract/dnd/npc-occurrences/shape**, **extract/dnd/npc-occurrences/registry**, **extract/dnd/npc-occurrences/source_refs**, **extract/dnd/npc-occurrences/source_relatedness**, **normalize/dnd/npc-occurrences/invariants** |
| Scene descriptions | **extract/dnd/scene-descriptions/shape**, **extract/dnd/scene-descriptions/source_refs**, **extract/dnd/scene-descriptions/source_relatedness**, **normalize/dnd/scene-descriptions/invariants** | | Scene descriptions | **extract/dnd/scene-descriptions/shape**, **extract/dnd/scene-descriptions/source_refs**, **extract/dnd/scene-descriptions/source_relatedness**, **normalize/dnd/scene-descriptions/invariants** |
| Enemy events | **extract/dnd/enemy-events/shape**, **extract/dnd/enemy-events/engagements**, **extract/dnd/enemy-events/source_refs**, **extract/dnd/enemy-events/source_relatedness**, **normalize/dnd/enemy-events/invariants** |
| Location registry | **extract/dnd/location-registry/shape**, **extract/dnd/location-registry/source_refs**, **extract/dnd/location-registry/source_relatedness**, **normalize/dnd/location-registry/identity** |
| Location occurrences | **extract/dnd/location-occurrences/shape**, **extract/dnd/location-occurrences/registry**, **extract/dnd/location-occurrences/source_refs**, **extract/dnd/location-occurrences/source_relatedness**, **normalize/dnd/location-occurrences/invariants** |
When no override is configured, production D&D bindings use the following When no override is configured, production D&D bindings use the following
ordered chains. Each row lists extract then normalize; spell chains are the ordered chains. Each row lists extract then normalize; spell chains are the
@@ -516,15 +344,11 @@ same at both stages.
| Lane | Extract | Normalize | | Lane | Extract | Normalize |
| --- | --- | --- | | --- | --- | --- |
| Spells | generic/valid_json, extract/dnd/spells/shape, extract/dnd/spells/catalog, extract/dnd/spells/source_refs, generic/valid_json_schema, extract/dnd/spells/source_relatedness | Same as extract | | Spells | generic/valid_json, extract/dnd/spells/shape, extract/dnd/spells/catalog, extract/dnd/spells/source_refs, generic/valid_json_schema, extract/dnd/spells/source_relatedness | Same as extract |
| NPC registry | generic/valid_json, extract/dnd/npc-registry/shape, extract/dnd/npc-registry/source_refs, generic/valid_json_schema, extract/dnd/npc-registry/source_relatedness | generic/valid_json, extract/dnd/npc-registry/shape, normalize/dnd/npc-registry/identity, extract/dnd/npc-registry/source_refs, generic/valid_json_schema, extract/dnd/npc-registry/source_relatedness | | NPCs | generic/valid_json, extract/dnd/npcs/shape, extract/dnd/npcs/source_refs, generic/valid_json_schema, extract/dnd/npcs/source_relatedness | generic/valid_json, extract/dnd/npcs/shape, normalize/dnd/npcs/identity, extract/dnd/npcs/source_refs, generic/valid_json_schema, extract/dnd/npcs/source_relatedness |
| Combat turns | generic/valid_json, extract/dnd/combat-turns/shape, extract/dnd/combat-turns/source_refs, generic/valid_json_schema, extract/dnd/combat-turns/source_relatedness | generic/valid_json, extract/dnd/combat-turns/shape, normalize/dnd/combat-turns/invariants, extract/dnd/combat-turns/source_refs, generic/valid_json_schema, extract/dnd/combat-turns/source_relatedness | | Combat turns | generic/valid_json, extract/dnd/combat-turns/shape, extract/dnd/combat-turns/source_refs, generic/valid_json_schema, extract/dnd/combat-turns/source_relatedness | generic/valid_json, extract/dnd/combat-turns/shape, normalize/dnd/combat-turns/invariants, extract/dnd/combat-turns/source_refs, generic/valid_json_schema, extract/dnd/combat-turns/source_relatedness |
| Item occurrences | generic/valid_json, extract/dnd/item-occurrences/shape, extract/dnd/item-occurrences/registry, extract/dnd/item-occurrences/source_refs, generic/valid_json_schema, extract/dnd/item-occurrences/source_relatedness | generic/valid_json, extract/dnd/item-occurrences/shape, extract/dnd/item-occurrences/registry, normalize/dnd/item-occurrences/invariants, extract/dnd/item-occurrences/source_refs, generic/valid_json_schema, extract/dnd/item-occurrences/source_relatedness | | Item events | generic/valid_json, extract/dnd/item-events/shape, extract/dnd/item-events/source_refs, generic/valid_json_schema, extract/dnd/item-events/source_relatedness | generic/valid_json, extract/dnd/item-events/shape, normalize/dnd/item-events/invariants, extract/dnd/item-events/source_refs, generic/valid_json_schema, extract/dnd/item-events/source_relatedness |
| Item registry | generic/valid_json, extract/dnd/item-registry/shape, extract/dnd/item-registry/source_refs, generic/valid_json_schema, extract/dnd/item-registry/source_relatedness | generic/valid_json, extract/dnd/item-registry/shape, normalize/dnd/item-registry/identity, extract/dnd/item-registry/source_refs, generic/valid_json_schema, extract/dnd/item-registry/source_relatedness | | NPC interactions | generic/valid_json, extract/dnd/npc-interactions/shape, extract/dnd/npc-interactions/registry, extract/dnd/npc-interactions/source_refs, generic/valid_json_schema, extract/dnd/npc-interactions/source_relatedness | generic/valid_json, extract/dnd/npc-interactions/shape, extract/dnd/npc-interactions/registry, normalize/dnd/npc-interactions/invariants, extract/dnd/npc-interactions/source_refs, generic/valid_json_schema, extract/dnd/npc-interactions/source_relatedness |
| NPC occurrences | generic/valid_json, extract/dnd/npc-occurrences/shape, extract/dnd/npc-occurrences/registry, extract/dnd/npc-occurrences/source_refs, generic/valid_json_schema, extract/dnd/npc-occurrences/source_relatedness | generic/valid_json, extract/dnd/npc-occurrences/shape, extract/dnd/npc-occurrences/registry, normalize/dnd/npc-occurrences/invariants, extract/dnd/npc-occurrences/source_refs, generic/valid_json_schema, extract/dnd/npc-occurrences/source_relatedness |
| Scene descriptions | generic/valid_json, extract/dnd/scene-descriptions/shape, extract/dnd/scene-descriptions/source_refs, generic/valid_json_schema, extract/dnd/scene-descriptions/source_relatedness | generic/valid_json, extract/dnd/scene-descriptions/shape, normalize/dnd/scene-descriptions/invariants, extract/dnd/scene-descriptions/source_refs, generic/valid_json_schema, extract/dnd/scene-descriptions/source_relatedness | | Scene descriptions | generic/valid_json, extract/dnd/scene-descriptions/shape, extract/dnd/scene-descriptions/source_refs, generic/valid_json_schema, extract/dnd/scene-descriptions/source_relatedness | generic/valid_json, extract/dnd/scene-descriptions/shape, normalize/dnd/scene-descriptions/invariants, extract/dnd/scene-descriptions/source_refs, generic/valid_json_schema, extract/dnd/scene-descriptions/source_relatedness |
| Enemy events | generic/valid_json, extract/dnd/enemy-events/shape, extract/dnd/enemy-events/engagements, extract/dnd/enemy-events/source_refs, generic/valid_json_schema, extract/dnd/enemy-events/source_relatedness | generic/valid_json, extract/dnd/enemy-events/shape, normalize/dnd/enemy-events/invariants, extract/dnd/enemy-events/source_refs, generic/valid_json_schema, extract/dnd/enemy-events/source_relatedness |
| Location registry | generic/valid_json, extract/dnd/location-registry/shape, extract/dnd/location-registry/source_refs, generic/valid_json_schema, extract/dnd/location-registry/source_relatedness | generic/valid_json, extract/dnd/location-registry/shape, normalize/dnd/location-registry/identity, extract/dnd/location-registry/source_refs, generic/valid_json_schema, extract/dnd/location-registry/source_relatedness |
| Location occurrences | generic/valid_json, extract/dnd/location-occurrences/shape, extract/dnd/location-occurrences/registry, extract/dnd/location-occurrences/source_refs, generic/valid_json_schema, extract/dnd/location-occurrences/source_relatedness | generic/valid_json, extract/dnd/location-occurrences/shape, extract/dnd/location-occurrences/registry, normalize/dnd/location-occurrences/invariants, extract/dnd/location-occurrences/source_refs, generic/valid_json_schema, extract/dnd/location-occurrences/source_relatedness |
Chains are only registered for the D&D extract and normalize modules shown Chains are only registered for the D&D extract and normalize modules shown
above; select an explicit override when a different compatible chain is above; select an explicit override when a different compatible chain is

View File

@@ -1,209 +0,0 @@
# Consuming The Complete D&D Pipeline
Use this workflow when an orchestrator runs the maintained complete D&D
pipeline and consumes its structured JSON artifacts. The generic
[subprocess consumer guide](subprocess.md) owns process-level responsibilities;
this guide connects that workflow to the complete D&D configuration, its
Seriatim input, and its artifact inventory.
The [CLI reference](../cli.md), [configuration reference](../config.md),
[run-result receipt](../integrations/run-result.md), and
[published JSON output contract](../integrations/json-output.md) remain the
canonical definitions of those public interfaces.
## Prepare And Validate The Deployment
Start from the maintained
[complete D&D configuration](../../examples/dnd-complete.config.yml). It uses
the `dnd-session` pipeline and demonstrates every implemented D&D lane, ordered
artifact handoffs, campaign references, chunk-map publication, and evidence
context.
A deployment must provide its own PromptKit profile and campaign reference
files. Use absolute paths for service and subprocess deployments. In
particular, observe these different resolution rules:
- reference paths in YAML are resolved relative to the Notarius configuration
file; and
- `promptkit.profile_file` is resolved relative to the Notarius process working
directory.
Do not copy the repository example's relative profile path into a deployment
without also controlling that working directory. The complete path and profile
rules are defined in [Configuration](../config.md).
Preflight the deployed configuration before processing sessions and whenever
it changes:
```sh
notarius config validate \
--config /absolute/path/to/notarius.yml \
--pipeline dnd-session
```
Provide credentials through the environment or the documented configuration
mechanism. Do not put credentials in command arguments, generated
configuration, or logs.
## Supply The Transcript
The complete pipeline consumes a Seriatim JSON document. The
[Seriatim input contract](../integrations/seriatim.md) defines its required
metadata, segments, and validation rules. Preserve segment IDs: D&D artifact
citations use those segment IDs as source-unit ranges.
When the caller maintains several transcript tiers, use the final trimmed JSON
transcript so extraction operates on the same session content presented to
later consumers. For example, Narratio identifies this implemented artifact as
`narratio.transcript.final_trimmed` and normally stores it at
`transcripts/final.trimmed.json`.
Notarius generates a stable prompt session from the resolved input module and
the exact input bytes. An ordinary orchestrator should not pass `--session-id`.
Use that override only when intentionally changing the routing relationship
between invocations; it is not a credential or output identity.
## Run Notarius
Invoke the pipeline with explicit absolute paths and request its
machine-readable receipt:
```sh
notarius run dnd-session \
--config /absolute/path/to/notarius.yml \
--input /absolute/path/to/transcripts/final.trimmed.json \
--output-dir /absolute/path/to/notarius-output \
--json
```
The caller should:
- capture stdout and stderr separately;
- propagate cancellation and impose an operator-appropriate timeout;
- wait for process completion before interpreting stdout; and
- retain stderr for diagnosis without copying secrets or transcript content
into other logs.
Only exit status 0 permits decoding stdout as a receipt. Ignore stdout after a
nonzero exit because a failed receipt write can leave partial bytes. The
[CLI reference](../cli.md#output-streams-and-exit-statuses) defines the complete
stream and exit-status contract.
## Discover The Published Bundle
Decode the successful stdout document as a supported run-result schema. For
the current contract, `schema_version` is `notarius.run-result.v2`. Tolerate
unknown fields allowed by that version, but reject an unsupported schema
version.
Use the receipt's absolute `output_directory` as the exact run-specific bundle
root. Do not scan the output root for its newest directory, guess a run ID, or
construct a bundle path. Resolve `index_file` beneath `output_directory` and
reject an absolute logical path or any result that escapes the bundle root.
The complete configuration uses the application validation defaults. A caller
that requires fully validated D&D artifacts must also require receipt
`validation_status: approved`; a successful `incomplete` result reflects the
configured validator-failure continuation policy and carries its bounded
validator provenance in `validation_summaries`.
Read `index.json` and locate each requested lane in `output_files` by its exact
`lane_id`. Do not guess a lane filename. Before decoding a payload:
1. resolve its descriptor's relative `file` beneath the bundle root with the
same confinement check;
2. verify the descriptor's media type and schema identity against the linked
artifact contract; and
3. decode the payload according to that contract.
The [published JSON output contract](../integrations/json-output.md) defines
the index and bundle layout. Treat all paths obtained from a decoded external
document as untrusted until confined to their documented root.
## Complete Artifact Inventory
When every configured lane is accepted, the complete example publishes these
lane artifacts:
| Lane ID | Purpose | Canonical contract |
| --- | --- | --- |
| `item-registry` | Canonical registry of encountered items and currency. | [Item registry](../integrations/dnd-item-registry-artifacts.md) |
| `npc-registry` | Canonical registry of named NPCs. | [NPC registry](../integrations/dnd-npc-registry-artifacts.md) |
| `location-registry` | Canonical registry of named locations. | [Location registry](../integrations/dnd-location-registry-artifacts.md) |
| `scene-descriptions` | Classification, title, and summary for each scene. | [Scene descriptions](../integrations/dnd-scene-description-artifacts.md) |
| `item-occurrences` | Source-grounded item discovery, acquisition, use, transfer, and loss events. | [Item occurrences](../integrations/dnd-item-occurrence-artifacts.md) |
| `spells` | Source-grounded spell casts and casters. | [Spell casts](../integrations/dnd-spell-artifacts.md) |
| `combat-turns` | Source-grounded combat turn participation. | [Combat turns](../integrations/dnd-combat-turn-artifacts.md) |
| `npc-occurrences` | Source-grounded NPC interaction occurrences. | [NPC occurrences](../integrations/dnd-npc-occurrence-artifacts.md) |
| `location-occurrences` | Source-grounded location occurrences. | [Location occurrences](../integrations/dnd-location-occurrence-artifacts.md) |
| `enemy-events` | Source-grounded enemy combat events. | [Enemy events](../integrations/dnd-enemy-event-artifacts.md) |
The JSON encoder always publishes these bundle-management files:
| File | Purpose |
| --- | --- |
| `index.json` | Discovery document for lane and pipeline-wide artifacts. |
| `manifest.json` | Run provenance and result summaries. |
| `rejected.json` | Rejected pipeline outputs. |
| `warnings.json` | Actionable process-degradation warnings. |
| `diagnostics.json` | Advisory and observation findings for accepted artifacts. |
The complete configuration also requests two pipeline-wide artifacts:
- [`chunk-map.json`](../integrations/chunk-map.md), the accepted chunk plan and
chunk metadata; and
- [`evidence-context.json`](../integrations/evidence-context.md), a reading
excerpt containing the union of selected cited source units and the
configured surrounding window.
Discover both from their top-level `index.json` descriptors rather than
treating them as lanes. Evidence context is convenient reading material, not
authoritative provenance; citations in the normalized lane payloads remain the
evidence contract.
Every optional or lane file is published only when its corresponding artifact
is available. A successful process does not guarantee that all configured
lanes were accepted.
## Decide What Counts As Consumer Success
Exit status 0 means Notarius completed the pipeline and published its result
bundle. The receipt or bundle may still report warnings, rejected outputs, or
missing lane descriptors. A downstream consumer must define its own required
artifact set explicitly.
A caller that claims to consume the complete D&D workflow should normally
require all ten lane IDs in the table and verify each descriptor's expected
contract. If any required lane is missing, rejected, or incompatible, fail the
caller's extraction step while retaining the Notarius bundle for diagnosis. A
consumer that needs only a subset may define and document a narrower policy.
Keep the successful receipt with the complete published bundle. Retain
`manifest.json`, `rejected.json`, `warnings.json`, and captured process logs as
required by the caller's provenance, diagnosis, and retention policies. Avoid
selectively copying payload files without also preserving enough index and
manifest information to identify their originating run and contracts.
The transcript, lane artifacts, evidence context, manifest, debug data, and
logs can all contain private campaign information. Apply the same access,
publication, and retention controls used for the source transcript.
## Consumer Checklist
- Validate the deployed Notarius configuration and `dnd-session` pipeline.
- Pass the final trimmed Seriatim JSON transcript with stable segment IDs.
- Use absolute configuration, input, output-root, profile, and reference paths
in service deployments.
- Capture stdout and stderr separately and enforce cancellation and timeout.
- Parse stdout only after exit status 0.
- Accept only supported receipt, index, and artifact schema versions while
tolerating permitted unknown fields.
- Use the receipt's `output_directory`; never guess the run directory.
- Confine `index_file` and every descriptor path to the published bundle root.
- Discover lanes by `lane_id` and verify descriptor compatibility before
decoding payloads.
- Enforce an explicit required-lane policy and inspect rejections and warnings.
- Preserve the receipt and sufficient bundle provenance for every retained
artifact.
- Protect all transcript-derived files and diagnostic streams as sensitive
campaign data.

View File

@@ -6,10 +6,6 @@ statuses, while the [run-result receipt](../integrations/run-result.md) and
[Published JSON Output contract](../integrations/json-output.md) own the [Published JSON Output contract](../integrations/json-output.md) own the
durable result formats. durable result formats.
For the maintained complete D&D workflow, including its transcript input,
configured lane inventory, and downstream acceptance checklist, see
[Consuming The Complete D&D Pipeline](dnd-pipeline.md).
## Run And Check The Process ## Run And Check The Process
Optionally preflight a selected configuration and pipeline before work starts: Optionally preflight a selected configuration and pipeline before work starts:
@@ -31,13 +27,10 @@ notarius run pipeline-id \
``` ```
Use absolute paths for supplied input, configuration, output-root, and Use absolute paths for supplied input, configuration, output-root, and
reference files. Notarius generates a stable prompt session for the resolved reference files. When a stable prompt session identifier or references are
input module and exact input bytes. Pass **--session-id** only when intentionally needed, pass the supported CLI flags. Supply credentials through Notarius's
grouping different invocations under a different session. Supply credentials documented configuration and environment mechanisms, never as command-line
through Notarius's documented configuration and environment mechanisms, never arguments or generated secret-bearing configuration.
as command-line arguments or generated secret-bearing configuration. In
particular, a session identifier is provider-visible and is not a credential
mechanism.
Wait for the process before interpreting standard output. Only an exit status Wait for the process before interpreting standard output. Only an exit status
of 0 permits decoding the receipt. On a nonzero exit, retain standard error for of 0 permits decoding the receipt. On a nonzero exit, retain standard error for
@@ -59,30 +52,23 @@ contract. The JSON bundle contract links to the available lane contracts.
If `index.json` has an `evidence_context` descriptor, treat it as a If `index.json` has an `evidence_context` descriptor, treat it as a
pipeline-wide artifact rather than a lane entry. Verify its six descriptor pipeline-wide artifact rather than a lane entry. Verify its six descriptor
fields before decoding the linked file according to the [Published Evidence fields before decoding the linked file according to the [Published Evidence
Context contract](../integrations/evidence-context.md). Decode its top-level Context contract](../integrations/evidence-context.md). Use each
source-unit array as a reading excerpt. Obtain authoritative citations and lane `evidence_refs` entry as the citation to source material. Its surrounding
provenance from the normalized lane artifacts; the excerpt has neither and its context range and included units explain the citation, but do not widen or
nearby units do not widen a lane artifact's cited source reference. replace the cited source reference.
A zero exit status may still report rejected outputs, warnings, or absent A zero exit status may still report rejected outputs, warnings, or absent
lanes. The caller decides which lane IDs are required for its own work and lanes. The caller decides which lane IDs are required for its own work and
which are optional; it should make that decision explicitly rather than infer which are optional; it should make that decision explicitly rather than infer
failure from the receipt counts alone. failure from the receipt counts alone.
When complete validation is required, also require receipt
`validation_status: approved` and inspect `validation_summaries`. A successful
run with `validation_status: incomplete` contains a structurally valid result
that advanced after validator execution could not complete under the configured
`warn_continue` policy. It is not reusable checkpoint state and should not be
silently treated as fully reviewed by the caller.
## Preserve Provenance And Handle Data Carefully ## Preserve Provenance And Handle Data Carefully
Keep the receipt with the published `manifest.json`, and retain Keep the receipt with the published `manifest.json`, and retain
`rejected.json`, `warnings.json`, and `diagnostics.json` when review or later provenance requires `rejected.json` and `warnings.json` when review or later provenance requires
them. Treat the input, output bundle, cache, debug bundle, and captured process them. Treat the input, output bundle, cache, debug bundle, and captured process
logs as potentially sensitive data. Apply the caller's access controls and logs as potentially sensitive data. Apply the caller's access controls and
retention policy, and avoid copying secrets into arguments, logs, or retention policy, and avoid copying secrets into arguments, logs, or
provenance records. An evidence-context artifact contains source-unit text and provenance records. An evidence-context artifact contains source-unit text and
metadata and can cover most of an input; preserve and share it only when that metadata, and selected lanes can cover most of an input; preserve and share it
source content is authorized for the recipient. only when that source content is authorized for the recipient.

View File

@@ -18,14 +18,13 @@ implemented component map.
| Any documentation addition or revision | [Documentation Policy](policy/documentation.md) | It defines canonical homes, audiences, current-behavior rules, and maintenance requirements. | | Any documentation addition or revision | [Documentation Policy](policy/documentation.md) | It defines canonical homes, audiences, current-behavior rules, and maintenance requirements. |
| Adding, changing, reviewing, or deleting tests | [Testing Policy](policy/testing.md) | It defines risk-based sufficiency, durable test boundaries, test-double guidance, and criteria for retaining tests. | | Adding, changing, reviewing, or deleting tests | [Testing Policy](policy/testing.md) | It defines risk-based sufficiency, durable test boundaries, test-double guidance, and criteria for retaining tests. |
| CLI composition or command behavior | [CLI Internals](internal/cli.md) and [CLI Reference](cli.md) | The internal guide owns composition and command flow; the reference owns public syntax. | | CLI composition or command behavior | [CLI Internals](internal/cli.md) and [CLI Reference](cli.md) | The internal guide owns composition and command flow; the reference owns public syntax. |
| Building a subprocess caller or changing its result protocol | [Subprocess Consumer Guide](consumers/subprocess.md), [Complete D&D Consumer Guide](consumers/dnd-pipeline.md), [Run Result Receipt](integrations/run-result.md), and [CLI Internals](internal/cli.md) | These separate generic caller workflow, the complete D&D workflow, the durable receipt contract, and CLI implementation behavior. | | Building a subprocess caller or changing its result protocol | [Subprocess Consumer Guide](consumers/subprocess.md), [Run Result Receipt](integrations/run-result.md), and [CLI Internals](internal/cli.md) | These separate caller workflow, durable receipt contract, and CLI implementation behavior. |
| Configuration loading, resolution, or user-visible configuration behavior | [Configuration Internals](internal/configuration.md) and [Configuration](config.md) | The internal guide owns loading and resolution mechanics; the reference owns the configuration contract. | | Configuration loading, resolution, or user-visible configuration behavior | [Configuration Internals](internal/configuration.md) and [Configuration](config.md) | The internal guide owns loading and resolution mechanics; the reference owns the configuration contract. |
| Pipeline resolution or execution | [Pipeline Internals](internal/pipeline.md) | It documents profiles, references, validation, retries, checkpoints, and runner behavior. | | Pipeline resolution or execution | [Pipeline Internals](internal/pipeline.md) | It documents profiles, references, validation, retries, checkpoints, and runner behavior. |
| Production modules or validators | [Module Internals](internal/modules.md), [D&D Module Internals](internal/dnd.md), and [D&D integration contracts](integrations/) | The generic guide owns extension mechanics, the D&D guide owns shared family conventions, and the contracts own durable output shapes. | | Production modules or validators | [Module Internals](internal/modules.md), [D&D Module Internals](internal/dnd.md), and [D&D integration contracts](integrations/) | The generic guide owns extension mechanics, the D&D guide owns shared family conventions, and the contracts own durable output shapes. |
| LLM clients, prompts, schemas, profiles, or scheduling | [LLM Runtime](internal/llm.md) | It documents the transport boundary and PromptKit integration. | | LLM clients, prompts, schemas, profiles, or scheduling | [LLM Runtime](internal/llm.md) | It documents the transport boundary and Scriptorium integration. |
| Output, cache, resume, or debug artifacts | [Run State Internals](internal/state.md), [Operations](operations.md), and [Configuration](config.md) | These separate implementation details, operator behavior, and configuration contracts. | | Output, cache, resume, or debug artifacts | [Run State Internals](internal/state.md), [Operations](operations.md), and [Configuration](config.md) | These separate implementation details, operator behavior, and configuration contracts. |
| External input formats, artifact schemas, or durable output files | [Integration Contracts](integrations/) | Integration documents define external and durable data contracts. | | External input formats, artifact schemas, or durable output files | [Integration Contracts](integrations/) | Integration documents define external and durable data contracts. |
| Release preparation, tagging, publication, or verification | [Source Releases](release.md) and [Documentation Policy](policy/documentation.md) | The release procedure owns maintainer guards and immutable-tag recovery; the policy assigns release-note ownership. |
| Proposed or unimplemented behavior | [Roadmap](roadmap/) | Future work belongs only in roadmap documentation until implemented. | | Proposed or unimplemented behavior | [Roadmap](roadmap/) | Future work belongs only in roadmap documentation until implemented. |
For an existing subsystem, also inspect its focused tests and the package-local For an existing subsystem, also inspect its focused tests and the package-local

View File

@@ -55,7 +55,7 @@ record controls eligibility only: its title, summary, and reference do not
become turn evidence. No exact matching scene also produces an empty list and become turn evidence. No exact matching scene also produces an empty list and
the `scene_classification_unavailable` warning. the `scene_classification_unavailable` warning.
An optional normalized [NPC registry artifact](dnd-npc-registry-artifacts.md) can ground an An optional normalized [NPC artifact](dnd-npc-artifacts.md) can ground an
actor name. Its registry references are provenance, never combat evidence. actor name. Its registry references are provenance, never combat evidence.
Normalization trims and, where possible, canonicalizes actor names; orders and Normalization trims and, where possible, canonicalizes actor names; orders and
deduplicates exact source references; orders valid-evidence turns by source deduplicates exact source references; orders valid-evidence turns by source
@@ -63,9 +63,7 @@ chronology; and collapses only duplicates with the same actor identity, turn
kind, and complete valid evidence. It does not infer turns, initiative, or kind, and complete valid evidence. It does not infer turns, initiative, or
actions from registry or scene data. actions from registry or scene data.
The [NPC-occurrence artifact](dnd-npc-occurrence-artifacts.md) records The [NPC-interaction artifact](dnd-npc-interaction-artifacts.md) records
broader NPC occurrences. The [enemy-event artifact](dnd-enemy-event-artifacts.md) broader NPC occurrences. The [JSON output contract](json-output.md) defines
uses combat turns as grounding only; turns do not establish an enemy event or publication, and [D&D module internals](../internal/dnd.md) describes routing
its outcome. The [JSON output contract](json-output.md) defines publication, and validation mechanics.
and [D&D module internals](../internal/dnd.md) describes routing and validation
mechanics.

View File

@@ -1,114 +0,0 @@
# D&D Enemy-Event Artifact
This contract defines the durable, source-grounded enemy-event occurrence list.
It records enemies directly established as opposing the party and explicitly
observed combat outcomes. It is an ordered observation artifact from which a
consumer may derive a ledger; it is not a ledger, encounter roster, or terminal
state model.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/enemy-event-list` |
| Schema ID | `notarius.dnd.enemy_events` |
| Schema name | `notarius_dnd_enemy_events_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
`v1` is a strict JSON object with required `events`; the array may be empty.
Event and source-reference objects reject unknown fields. An incompatible shape
change requires a new schema version.
## Wire shape
Every event has these required fields:
| Field | Contract |
| --- | --- |
| `name` | Non-empty display name or directly grounded collective subject label. |
| `kind` | `engaged`, `killed`, `fled`, `captured`, or `incapacitated`. |
| `source_refs` | One or more current-transcript evidence ranges. |
Each source reference has exactly `source_id`, `start_unit_id`, and
`end_unit_id`. It identifies an inclusive current-transcript range; unit IDs
are positive and the start may not follow the end.
```json
{
"events": [
{
"name": "Ashfang",
"kind": "engaged",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 41, "end_unit_id": 42}
]
},
{
"name": "Ashfang",
"kind": "fled",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 57, "end_unit_id": 58}
]
}
]
}
```
## Event semantics and evidence
| Kind | Required evidence |
| --- | --- |
| `engaged` | The subject is directly established as actively opposing the party in combat. At most one engagement is emitted for one subject in one combat scene. |
| `killed` | The transcript explicitly establishes that the subject died or was killed. Damage, defeat, disappearance, or combat ending is insufficient. |
| `fled` | The subject explicitly escapes, retreats, or otherwise leaves combat to avoid continued engagement. Movement or absence from later turns is insufficient. |
| `captured` | The subject is explicitly taken prisoner or secured under the party's control. A grapple or temporary restraint alone is insufficient. |
| `incapacitated` | The subject is explicitly rendered unable to continue acting without being established as killed or captured. A missed turn is insufficient. |
The current transcript is the only event evidence. Campaign context and
normalized NPC, scene-description, combat-turn, and NPC-occurrence artifacts
can ground names or control combat eligibility, but none may supply event
evidence. An outcome may share evidence with an engagement, in which case both
events are retained.
Extraction is limited to chunks with an exact combat-scene classification. An
exact non-combat classification produces an accepted empty list. Missing or
mismatched classification also produces an accepted empty list and a
`scene_classification_unavailable` warning.
## Subjects, normalization, and order
A subject matching the normalized NPC registry uses that registry's canonical
display name. Unmatched hostile creatures, summoned entities, and directly
grounded groups remain valid subjects. An unnamed homogeneous group uses the
narrowest transcript-grounded label, such as `Orcs`, `One orc`, or `Remaining
orcs`; the artifact never invents synthetic member identities or quantities.
Party members, allies, neutral observers, mentioned-but-absent enemies, hazards,
traps, and environmental effects are excluded.
Normalization collapses surrounding and repeated internal whitespace in subject
display values, canonicalizes recognized registry names, canonicalizes and
deduplicates exact source ranges, then orders events by valid evidence
chronology, normalized subject identity, display name, kind, and reference
sequence. The deterministic kind tie order is `engaged`,
`incapacitated`, `captured`, `fled`, then `killed`. Only entries with the same
normalized name, kind, and complete canonical evidence sequence are collapsed.
Different kinds, evidence, repeated engagement in separate scenes, and later
outcomes remain separate. A later engagement for the same named subject is
preserved after an earlier outcome because the artifact does not assert an
irreversible state transition.
## Non-goals
The artifact has no NPC or scene ID, quantity, confidence, description,
rationale, summary, current state, or inferred terminal outcome. It does not
emit `active` or `unresolved`; consumers may derive an unresolved ledger view
only when an engagement has no later explicit outcome. It never infers an
outcome from turn absence, scene termination, initiative order, hit-point
guesses, or other artifacts.
The [JSON output contract](json-output.md) defines publication. Configuration
keys, required generated-reference slots, and validator-chain selection are
defined in the [configuration reference](../config.md). Implementation and
prompt-grounding mechanics are described in the
[D&D module internals](../internal/dnd.md).

View File

@@ -0,0 +1,78 @@
# D&D Item-Event Artifact
This contract defines the durable item and currency occurrence list produced by
`dnd/item-events`. It records source-grounded discoveries and possession
changes; it does not maintain an inventory, balance, or ledger.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/item-event-list` |
| Schema ID | `notarius.dnd.item_events` |
| Schema name | `notarius_dnd_item_events_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
`v1` is a strict JSON object with required `events`; the array may be empty.
Event and source-reference objects reject unknown fields. An incompatible
shape change requires a new schema version.
## Wire shape
Every event has required `name`, `kind`, and `source_refs`. `quantity`, `from`,
and `to` are optional where the event kind permits them.
| Field | Contract |
| --- | --- |
| `name` | Non-empty item or currency display name. |
| `kind` | `discovered`, `acquired`, `lost`, `consumed`, or `transferred`. |
| `quantity` | Optional positive integer; omit it when no count is established. |
| `from` | Optional non-empty losing holder, when allowed by `kind`. |
| `to` | Optional non-empty gaining holder, when allowed by `kind`. |
| `source_refs` | One or more transcript evidence ranges. |
Each source reference has exactly `source_id`, `start_unit_id`, and
`end_unit_id`. It identifies an inclusive current-transcript range; unit IDs
are positive and the start may not follow the end.
```json
{
"events": [
{
"name": "Silver Pieces",
"kind": "acquired",
"quantity": 20,
"to": "party",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 2, "end_unit_id": 2}
]
}
]
}
```
## Holder rules and minimal extraction
`discovered` has neither holder; `acquired` requires `to` and forbids `from`;
`lost` and `consumed` require `from` and forbid `to`; `transferred` requires
both holders. `party` denotes collective possession. A transfer cannot use
`party` for either holder and its two normalized holders must differ.
Only an evidenced discovery or possession change belongs in this artifact.
It does not infer quantities or holders, convert currency denominations,
calculate balances, or merge nearby events. Campaign references may
disambiguate names but are never event evidence. Currency uses the ordinary
`name` field and an explicit `quantity` only when the transcript establishes
one; each denomination remains a separate event.
Normalization trims display whitespace, orders and removes exact duplicate
source references, then orders events by valid source chronology, name identity
and display value, kind, holders, quantity, and reference sequence. It
collapses only entries with the same normalized durable fields and complete
valid evidence.
The [JSON output contract](json-output.md) defines publication. See
[D&D module internals](../internal/dnd.md) for implementation details and the
[NPC-interaction artifact](dnd-npc-interaction-artifacts.md) for a distinct
kind of occurrence.

View File

@@ -1,72 +0,0 @@
# D&D Item-Occurrence Artifact
`dnd/item-occurrences` currently produces this source-grounded item and currency
occurrence list. It records discoveries and possession changes, not an
inventory, balance, or ledger.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/item-occurrence-list` |
| Schema ID | `notarius.dnd.item_occurrences` |
| Schema name | `notarius_dnd_item_occurrences_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
`v1` accepts one strict JSON object with required `occurrences`; the array may
be empty. Each occurrence has required `item_id`, `name`, `kind`, and
`source_refs`, and occurrence and source-reference objects reject unknown
fields. `quantity`, `from`, and `to` appear only when their kind permits them.
An incompatible shape change requires a new schema version.
## Registry grounding
Both extraction and normalization require an `item_registry` reference bound to
an earlier normalized `dnd/item-registry` artifact. The registry is immutable
for an operation and contributes names-only grounding after the shared evidence
message. Notarius resolves the model's selected name into the unchanged exact
durable ID/name pair. It is never occurrence evidence.
Each occurrence must use one exact registry ID/name pair. An extraction response
with an unknown or ambiguous selected name is rejected as invalid model output;
the configured pipeline may retry it and never accepts a partial artifact.
Normalization and validation remain defense in depth for artifacts entering
through other boundaries: normalization canonicalizes a recognized name by ID,
preserves unknown values for the registry validator, and the registry validator
rejects unknown or mismatched pairs.
## Wire shape
Each source reference has exactly `source_id`, `start_unit_id`, and
`end_unit_id`. It identifies an inclusive range in the current transcript;
unit IDs are positive and the start may not follow the end.
```json
{
"occurrences": [
{
"item_id": "item:sha256:…",
"name": "Silver Pieces",
"kind": "acquired",
"quantity": 20,
"to": "party",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 2, "end_unit_id": 2}
]
}
]
}
```
The five kinds remain `discovered`, `acquired`, `lost`, `consumed`, and
`transferred`. Holder, quantity, currency, ordering, and exact-duplicate rules
are unchanged: discovered has no holder; acquired requires `to`; lost and
consumed require `from`; transferred requires distinct non-`party` holders.
The only current downstream compatibility requirement is its registry handoff;
the normalized occurrence list is otherwise published for callers. See
[Configuration](../config.md#d-d-reference-slots) for the binding and
[JSON output](json-output.md) for publication.
See [item registry](dnd-item-registry-artifacts.md) for the grounding artifact
and [D&D module internals](../internal/dnd.md) for implementation details.

View File

@@ -1,95 +0,0 @@
# D&D Item Registry Artifact
This contract defines the durable, source-grounded item registry produced by
`dnd/item-registry`. It records transcript-established item types and unique
designations for one source document; it is not an inventory, holder record,
quantity ledger, or item-occurrence artifact.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/item-registry` |
| Schema ID | `notarius.dnd.item_registry` |
| Schema name | `notarius_dnd_item_registry_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
| Identity policy | `dnd.item_registry.identity.v1` |
`v1` accepts one strict JSON object with required `items`; the array may be
empty. Item and source-reference objects reject unknown fields. An incompatible
artifact shape or identity-policy change uses a new version or policy.
## Wire shape and identity
Each item has these required fields:
| Field | Contract |
| --- | --- |
| `id` | `item:sha256:` followed by 64 lowercase hexadecimal characters. |
| `name` | Non-empty transcript-established item type or unique designation. |
| `source_refs` | One or more transcript evidence ranges that establish the item. |
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
The source ID identifies the transcript, unit IDs are positive inclusive unit
identifiers, and the start may not follow the end.
```json
{
"items": [
{
"id": "item:sha256:31e73b6280ef98e4d8070e07fd4de9b2c3e842cc03af1a09ca631cb95b73e3b3",
"name": "Star Compass",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 4, "end_unit_id": 5}
]
}
]
}
```
The ID is deterministic for an item name or type, rather than for one physical
instance. Notarius normalizes the display name for comparison with Unicode
NFKC, supported apostrophe normalization, collapsed whitespace, and case
folding. It hashes compact JSON for this array:
```text
["dnd.item_registry.identity.v1", comparison_name]
```
The canonical ID is the lowercase SHA-256 digest of those bytes with the
`item:sha256:` prefix. Equal comparison names represent one item identity;
normalization unions their transcript evidence when it safely consolidates a
candidate group.
## Scope, reconciliation, and evidence
The registry includes named unique items, concrete reusable item types, stable
unique designations, and separately established currency denominations. It
excludes vague loot or treasure, generic weapons, quantities, inferred
properties, and inferred uniqueness. Capitalization alone does not establish
eligibility.
Normalization first applies deterministic display, evidence, and ID rules. It
then may use a bounded LLM-assisted proposal to reconcile semantically duplicate
records. The proposal may choose only a supplied candidate display name;
invalid, uncertain, overlapping, or unsafe proposals retain the deterministic
result with retry or fallback diagnostics. A proposal that mixes a recognized
currency denomination with a non-currency item, or combines recognized
denominations, is unsafe and retains every deterministic record. Currency
denominations, materially different item types, and merely nearby objects
remain distinct. Source references establish registry provenance, not evidence
for later artifacts.
## Consumers and publication
`dnd/item-occurrences` requires one approved item registry through its
`item_registry` reference slot for both extraction and normalization. Its
consumer receives names-only grounding; Notarius resolves the selected name
into the unchanged exact durable ID/name pair. The registrys source references
are never occurrence evidence. Unknown or ambiguous selections are rejected by
the occurrence contract. See the
[item-occurrence artifact](dnd-item-occurrence-artifacts.md) for that strict
wire contract, [Configuration](../config.md#d-d-reference-slots) for binding
rules and validator selection, and the [JSON output contract](json-output.md)
for publication.

View File

@@ -1,86 +0,0 @@
# D&D Location-Occurrence Artifact
This contract defines the durable occurrence list produced by
`dnd/location-occurrences`. It records source-grounded ways the party relates
to locations in a required normalized location registry; it does not extend
that registry or infer a place absent from it.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/location-occurrence-list` |
| Schema ID | `notarius.dnd.location_occurrences` |
| Schema name | `notarius_dnd_location_occurrences_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
`v1` accepts one strict JSON object with required `occurrences`; the array may
be empty. Occurrence and source-reference objects reject unknown fields. An
incompatible shape change requires a new schema version.
## Wire shape
Each occurrence has these required fields:
| Field | Contract |
| --- | --- |
| `location_id` | Exact ID from the required normalized [location registry](dnd-location-registry-artifacts.md). |
| `name` | Exact canonical display name for `location_id` in that registry. |
| `kind` | One of `visited`, `planned`, `recalled`, or `mentioned`. |
| `source_refs` | One or more current-transcript evidence ranges for this occurrence. |
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
It identifies an inclusive range in the current transcript; unit IDs are
positive and the start may not follow the end.
```json
{
"occurrences": [
{
"location_id": "location:sha256:fb05475da0fc7debf994b517e1906ffe7209887a6a1ec306356d84de820b1a24",
"name": "Moon Gate",
"kind": "visited",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 12, "end_unit_id": 13}
]
}
]
}
```
## Occurrence categories
| Kind | Meaning |
| --- | --- |
| `visited` | The transcript establishes physical party presence, including arrival, continuing presence, or departure. |
| `planned` | The party explicitly proposes, intends, or agrees to future travel; speculation alone is not enough. |
| `recalled` | The transcript explicitly recounts prior party presence before the current live events. |
| `mentioned` | The location is explicit but no stronger category applies, including lore, directions, third-party activity, non-actionable speculation, a mere hypothetical reference, or out-of-character discussion. |
For overlapping evidence, precedence is `visited`, then `planned`, then
`recalled`, then `mentioned`. For example, “What if we went to Moon Gate?” is
eligible as `mentioned` when its narrow evidence explicitly references that
registry location, but it is not `planned` without an actual proposal,
intention, or agreement to travel. Inferred, unstated, uncertain, and
unsupported places or occurrences are omitted. Normalization
canonicalizes the registry name, orders and deduplicates source references, and
orders occurrences by source chronology, location ID, name, kind, and reference
sequence. It collapses only exact duplicates with the same ID, kind, and
complete canonical evidence sequence.
## Required grounding and evidence
Both extraction and normalization require exactly one `location_registry` reference of
kind `dnd/location-registry`, media type `application/json`, and at most 1 MiB. The
registry provides identity grounding only. The model selects a supplied
contextual name-and-registry-reference descriptor, and Notarius resolves it
into the exact durable ID/name pair. Unknown, partial, or ambiguous selections
are rejected rather than guessed or reassigned. The current transcript is the
only evidence source for an occurrence; registry evidence and provenance never
become occurrence evidence.
See [Configuration](../config.md#d-d-reference-slots) for the selectable slot
and generated-handoff compatibility, [D&D module internals](../internal/dnd.md)
for implementation behavior, and the [JSON output contract](json-output.md)
for publication.

View File

@@ -1,93 +0,0 @@
# D&D Location Registry Artifact
This contract defines the durable, source-grounded location registry produced
by `dnd/location-registry`. It records transcript-established physical places for one
source document; it is not a map, location hierarchy, campaign-wide world
registry, or location description.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/location-registry` |
| Schema ID | `notarius.dnd.location_registry` |
| Schema name | `notarius_dnd_location_registry_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
| Identity policy | `dnd.location_registry.identity.v1` |
`v1` accepts one strict JSON object with required `locations`; the array may be
empty. Location and source-reference objects reject unknown fields. An
incompatible artifact shape or identity-policy change uses a new version or
policy.
## Wire shape and identity
Each location has these required fields:
| Field | Contract |
| --- | --- |
| `id` | `location:sha256:` followed by 64 lowercase hexadecimal characters. |
| `name` | Non-empty transcript-established display name. |
| `source_refs` | One or more transcript evidence ranges that identify the place. |
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
The source ID identifies the transcript, unit IDs are positive inclusive unit
identifiers, and the start may not follow the end.
```json
{
"locations": [
{
"id": "location:sha256:fb05475da0fc7debf994b517e1906ffe7209887a6a1ec306356d84de820b1a24",
"name": "Moon Gate",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 4, "end_unit_id": 5}
]
}
]
}
```
The ID is deterministic and scoped to the source document. Notarius normalizes
the display name for comparison with Unicode NFKC, supported apostrophe
normalization, collapsed whitespace, and case folding. It hashes compact JSON
for this array, using the earliest canonical source reference as the anchor:
```text
["dnd.location_registry.identity.v1", comparison_name, source_id, start_unit_id, end_unit_id]
```
The canonical ID is the lowercase SHA-256 digest of those bytes with the
`location:sha256:` prefix. Equal display names are allowed when their evidence
anchors differ, so a generic name does not force distinct places to collapse.
## Scope, reconciliation, and evidence
Locations are physical or spatial places established by the transcript with a
stable proper name or unique in-world designation, such as named planes,
regions, settlements, districts, buildings, rooms, landmarks, routes, and
geographic features. Generic, temporary, relative, and descriptive phrases
such as “the room,” “the bar,” “the hallway,” “outside,” and “upstairs” are not
registry locations. Capitalization alone does not establish eligibility.
Notarius does not infer an unstated place or add hierarchy, coordinates,
descriptions, participants, or ownership.
Normalization first applies deterministic display, evidence, and ID rules. It
then may use a bounded LLM-assisted proposal to reconcile semantically duplicate
records. The proposal is validated and applied conservatively; invalid or
unusable proposals retain the deterministic result with retry or fallback
diagnostics. The registry's source references establish registry provenance,
not evidence for later artifacts.
## Consumers and publication
`dnd/location-occurrences` requires one approved location registry through its
`location_registry` reference slot. Its prompt receives contextual selectors
containing a canonical name and registry references; Notarius resolves a
selection into the unchanged exact durable ID/name pair. Registry references
must not be treated as occurrence evidence. See the
[location-occurrence artifact](dnd-location-occurrence-artifacts.md)
for that contract, [Configuration](../config.md#references-and-ordered-handoffs)
for binding rules, and the [JSON output contract](json-output.md) for
publication.

View File

@@ -0,0 +1,69 @@
# D&D NPC Artifact
This contract defines the durable NPC registry produced by `dnd/npcs`. It is a
minimal, source-grounded identity registry for other D&D artifacts, not a
character sheet or a relationship summary.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/npc-list` |
| Schema ID | `notarius.dnd.npcs` |
| Schema name | `notarius_dnd_npcs_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
| Identity policy | `dnd.npcs.identity.v1` |
`v1` accepts one strict JSON object with required `npcs`; the array may be
empty. NPC and source-reference objects reject unknown fields. An incompatible
artifact shape or identity-policy change uses a new version or policy.
## Wire shape and identity
Each NPC has these required fields:
| Field | Contract |
| --- | --- |
| `id` | `npc:sha256:` followed by 64 lowercase hexadecimal characters. |
| `name` | Non-empty canonical display name. |
| `source_refs` | One or more transcript evidence ranges for the identity. |
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
The source ID identifies the transcript, unit IDs are positive inclusive unit
identifiers, and the start may not follow the end.
```json
{
"npcs": [
{
"id": "npc:sha256:99a16589618a04f535a7d21fdcc71a0b1c05d22f752cd492065b1086d97bc3d7",
"name": "Mira Thorn",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 4, "end_unit_id": 5}
]
}
]
}
```
The ID is deterministic: normalize the name to Unicode NFKC, normalize the
supported apostrophe forms, collapse whitespace, case-fold it, SHA-256 the
result, then prefix the lowercase hexadecimal digest with `npc:sha256:`. Each
canonical identity and ID appears at most once. Normalization collapses records
with the same canonical identity, retains their earliest position, and merges
their canonicalized evidence; it does not add aliases, roles, descriptions, or
relationship fields.
## Scope and consumers
Only individually identifiable NPC names with transcript evidence belong in
this artifact. Groups, generic roles, invented labels, and descriptive
enrichment are excluded. Its source references prove registry provenance; they
do not become evidence for a spell, interaction, or combat occurrence.
This registry can ground actor or caster names in the [spell](dnd-spell-artifacts.md)
and [combat-turn](dnd-combat-turn-artifacts.md) artifacts. It is required to
resolve the canonical `name` in an [NPC interaction](dnd-npc-interaction-artifacts.md).
The [JSON output contract](json-output.md) defines publication, and
[D&D module internals](../internal/dnd.md) owns pipeline mechanics.

View File

@@ -1,7 +1,7 @@
# D&D NPC Occurrence Artifact # D&D NPC Interaction Artifact
This contract defines the durable occurrence list produced by This contract defines the durable occurrence list produced by
`dnd/npc-occurrences`. It records discrete, source-grounded occurrences with `dnd/npc-interactions`. It records discrete, source-grounded interactions with
NPCs already present in a normalized registry; it does not extend that registry NPCs already present in a normalized registry; it does not extend that registry
or summarize the session. or summarize the session.
@@ -9,37 +9,35 @@ or summarize the session.
| Property | Value | | Property | Value |
| --- | --- | | --- | --- |
| Artifact kind | `dnd/npc-occurrence-list` | | Artifact kind | `dnd/npc-interaction-list` |
| Schema ID | `notarius.dnd.npc_occurrences` | | Schema ID | `notarius.dnd.npc_interactions` |
| Schema name | `notarius_dnd_npc_occurrences_v1` | | Schema name | `notarius_dnd_npc_interactions_v1` |
| Schema version | `v1` | | Schema version | `v1` |
| Media type | `application/json` | | Media type | `application/json` |
`v1` is a strict JSON object with required `occurrences`; the array may be `v1` is a strict JSON object with required `interactions`; the array may be
empty. Occurrence and source-reference objects reject unknown fields. An empty. Interaction and source-reference objects reject unknown fields. An
incompatible shape change requires a new schema version. incompatible shape change requires a new schema version.
## Wire shape ## Wire shape
Each occurrence has these required fields: Each interaction has these required fields:
| Field | Contract | | Field | Contract |
| --- | --- | | --- | --- |
| `npc_id` | Exact durable ID from the required NPC registry. |
| `name` | Non-empty canonical display name from the required NPC registry. | | `name` | Non-empty canonical display name from the required NPC registry. |
| `kind` | One of the occurrence categories below. | | `kind` | One of the interaction categories below. |
| `source_refs` | One or more transcript evidence ranges. | | `source_refs` | One or more transcript evidence ranges. |
Each source reference has exactly `source_id`, `start_unit_id`, and Each source reference has exactly `source_id`, `start_unit_id`, and
`end_unit_id`. It identifies an inclusive range in the current transcript; `end_unit_id`. It identifies an inclusive range in the current transcript;
unit IDs are positive and the start may not follow the end. Extraction evidence unit IDs are positive and the start may not follow the end. Extraction evidence
for an occurrence is confined to its accepted chunk. for an interaction is confined to its accepted chunk.
```json ```json
{ {
"occurrences": [ "interactions": [
{ {
"npc_id": "npc:sha256:example",
"name": "Mira Thorn", "name": "Mira Thorn",
"kind": "dialogue", "kind": "dialogue",
"source_refs": [ "source_refs": [
@@ -50,7 +48,7 @@ for an occurrence is confined to its accepted chunk.
} }
``` ```
## Occurrence categories ## Interaction categories
| Kind | Meaning | | Kind | Meaning |
| --- | --- | | --- | --- |
@@ -67,24 +65,14 @@ for uncertain classification.
## Identity, evidence, and order ## Identity, evidence, and order
The required normalized [NPC registry artifact](dnd-npc-registry-artifacts.md) The required normalized [NPC artifact](dnd-npc-artifacts.md) resolves `name`.
supplies names-only contextual grounding to the model. Notarius resolves the Registry references are provenance only and never replace an interaction's own
selected name and writes the exact `{npc_id, name}` pair. An unknown or evidence. Normalization canonicalizes recognized registry names, orders and
ambiguous selection rejects the complete model result; normalization does not deduplicates exact source references, then orders interactions by valid source
repair names by similarity. Registry references are provenance only and never
replace an occurrence's own evidence.
The registry may include an identity established by a factual third-party
mention; that provenance alone does not create a `mentioned` occurrence. Each
occurrence remains a separately cited fact in the current transcript.
Normalization validates the exact pair, orders and
deduplicates exact source references, then orders occurrences by valid source
chronology, NPC comparison identity, display name, kind, and reference sequence. chronology, NPC comparison identity, display name, kind, and reference sequence.
Only entries with the same NPC ID, canonical name, kind, and complete valid evidence Only entries with the same canonical name, kind, and complete valid evidence
sequence are collapsed; distinct categories or evidence remain separate. sequence are collapsed; distinct categories or evidence remain separate.
See the [combat-turn artifact](dnd-combat-turn-artifacts.md) for combat-action See the [combat-turn artifact](dnd-combat-turn-artifacts.md) for combat-action
occurrences. The [enemy-event artifact](dnd-enemy-event-artifacts.md) consumes occurrences and the [JSON output contract](json-output.md) for publication.
only `combat_opponent` occurrences as grounding; they never establish an enemy Pipeline mechanics are described in [D&D module internals](../internal/dnd.md).
event or outcome. The [JSON output contract](json-output.md) defines
publication. Pipeline mechanics are described in
[D&D module internals](../internal/dnd.md).

View File

@@ -1,92 +0,0 @@
# D&D NPC Registry Artifact
This contract defines the durable NPC registry produced by `dnd/npc-registry`. It is a
minimal, source-grounded identity registry for other D&D artifacts, not a
character sheet or a relationship summary.
## Identity and compatibility
| Property | Value |
| --- | --- |
| Artifact kind | `dnd/npc-registry` |
| Schema ID | `notarius.dnd.npc_registry` |
| Schema name | `notarius_dnd_npc_registry_v1` |
| Schema version | `v1` |
| Media type | `application/json` |
| Identity policy | `dnd.npc_registry.identity.v1` |
`v1` accepts one strict JSON object with required `npcs`; the array may be
empty. NPC and source-reference objects reject unknown fields. An incompatible
artifact shape or identity-policy change uses a new version or policy.
## Wire shape and identity
Each NPC has these required fields:
| Field | Contract |
| --- | --- |
| `id` | `npc:sha256:` followed by 64 lowercase hexadecimal characters. |
| `name` | Non-empty canonical display name. |
| `source_refs` | One or more transcript evidence ranges for the identity. |
A source reference has exactly `source_id`, `start_unit_id`, and `end_unit_id`.
The source ID identifies the transcript, unit IDs are positive inclusive unit
identifiers, and the start may not follow the end.
```json
{
"npcs": [
{
"id": "npc:sha256:35ba5f679aee69e07ae3bd65c44278f29539d5dc9bb5225db1c0060555b23221",
"name": "Mira Thorn",
"source_refs": [
{"source_id": "session-7", "start_unit_id": 4, "end_unit_id": 5}
]
}
]
}
```
The ID is deterministic: normalize the name to Unicode NFKC, normalize the
supported apostrophe forms, collapse whitespace, case-fold it, then serialize
`["dnd.npc_registry.identity.v1", comparison_name]` as compact JSON. SHA-256
those UTF-8 bytes and prefix the lowercase hexadecimal digest with
`npc:sha256:`. Each canonical identity and ID appears at most once.
Normalization collapses records with the same canonical identity, retains their
earliest position, and merges
their canonicalized evidence; it does not add aliases, roles, descriptions, or
relationship fields.
When evidence supports a semantically duplicate group, the canonical display
name is one of that group's supplied candidates. A complete, stable proper name
is preferred over an abbreviation. An unadorned proper name is preferred over
the same name plus a contextual class, role, title, or relationship descriptor
unless the transcript establishes that descriptor as part of the person's
name. A longer candidate is not preferred solely because it includes such a
descriptor.
## Scope and consumers
Only individually identifiable NPC names with transcript evidence belong in
this artifact. A factual third-party mention can establish an identity even if
the NPC is not present, speaking, or acting in the cited passage. Names used
only in hypothetical, speculative, or imagined examples are excluded, as are
groups, generic roles, invented labels, and descriptive enrichment. Its source
references prove registry provenance; they do not become evidence for a spell,
occurrence, combat, or enemy-event occurrence.
Registry evidence establishes an identity, not an [NPC occurrence](dnd-npc-occurrence-artifacts.md).
That later artifact independently records any current-transcript occurrence
with its own cited evidence and category.
This registry can ground actor or caster names in the [spell](dnd-spell-artifacts.md)
and [combat-turn](dnd-combat-turn-artifacts.md) artifacts. It is required to
resolve the canonical `name` in an [NPC occurrence](dnd-npc-occurrence-artifacts.md).
Occurrence consumers receive names-only grounding; Notarius resolves the
selected canonical name and writes the unchanged exact durable ID/name pair.
Spells, combat turns, and the [enemy-event artifact](dnd-enemy-event-artifacts.md)
also receive names-only grounding for actor or subject display. None of these
projections supply later-artifact evidence. [Configuration](../config.md#d-d-reference-slots)
owns the `npc_registry` binding rules.
The [JSON output contract](json-output.md) defines publication, and
[D&D module internals](../internal/dnd.md) owns pipeline mechanics.

View File

@@ -62,9 +62,8 @@ durable fields, or the same source range with different kind, title, or
summary, is invalid. It does not merge adjacent ranges, alter prose, or infer summary, is invalid. It does not merge adjacent ranges, alter prose, or infer
missing scenes. missing scenes.
The [combat-turn artifact](dnd-combat-turn-artifacts.md) and The [combat-turn artifact](dnd-combat-turn-artifacts.md) uses an exact matching
[enemy-event artifact](dnd-enemy-event-artifacts.md) use an exact matching
`combat` scene only as eligibility control; scene title, summary, and source `combat` scene only as eligibility control; scene title, summary, and source
reference never become their evidence. Publication is defined by the reference never become combat evidence. Publication is defined by the
[JSON output contract](json-output.md); implementation details live in [JSON output contract](json-output.md); implementation details live in
[D&D module internals](../internal/dnd.md). [D&D module internals](../internal/dnd.md).

View File

@@ -61,7 +61,7 @@ only when it has the same canonical spell, the same case- and
whitespace-insensitive caster identity, and the same complete valid reference whitespace-insensitive caster identity, and the same complete valid reference
sequence. Remaining entries retain their merged order. sequence. Remaining entries retain their merged order.
The optional normalized [NPC registry artifact](dnd-npc-registry-artifacts.md) can ground a The optional normalized [NPC artifact](dnd-npc-artifacts.md) can ground a
caster name. Its own references remain registry provenance and are never copied caster name. Its own references remain registry provenance and are never copied
into `source_refs`. into `source_refs`.

View File

@@ -67,12 +67,6 @@ including a collision with the embedded catalog. Matching uses the catalogs
case, whitespace, and apostrophe normalization, so authors should avoid names case, whitespace, and apostrophe normalization, so authors should avoid names
or aliases that normalize to another spell. or aliases that normalize to another spell.
Spell extraction receives the effective catalog as deterministic canonical-name
and alias pairs. An alias in the transcript selects its associated canonical
name; the extractor is instructed to return that canonical spelling. The
projection contains no catalog source metadata or provenance, and aliases
remain recognition context rather than transcript evidence.
The overlay is a recognition aid only. The durable spell-artifact schema and The overlay is a recognition aid only. The durable spell-artifact schema and
source-evidence rules are defined by the source-evidence rules are defined by the
[D&D spell artifact contract](dnd-spell-artifacts.md). [D&D spell artifact contract](dnd-spell-artifacts.md).

View File

@@ -1,11 +1,9 @@
# Published Evidence Context # Published Evidence Context
This contract defines the optional `source/evidence-context` artifact emitted This contract defines the optional `source/evidence-context` artifact emitted
by the production JSON output. It is a selected source-unit excerpt for by the production JSON output. Its configuration is owned by
convenient reading alongside normalized lane artifacts; it is not a second [Configuration](../config.md#module-bindings-and-validators); its logical-file
citation or provenance model. Its configuration is owned by discovery is owned by [Published JSON Output](json-output.md).
[Configuration](../config.md#module-bindings-and-validators), and its
logical-file discovery is owned by [Published JSON Output](json-output.md).
## Identity And Discovery ## Identity And Discovery
@@ -28,80 +26,91 @@ its absence means evidence publication was not enabled for that bundle.
## Payload ## Payload
The v1 payload is a top-level JSON array of generic source units. There is no The v1 payload is a JSON object with required `source_id`, `source_digest`,
wrapper, source-level metadata, context grouping, lane identifier, or evidence `window_units`, `selected_lanes`, and `contexts` fields. `selected_lanes` and
reference in the payload. An enabled configuration with no contributing `contexts` are always arrays; an enabled configuration with no accepted direct
accepted evidence publishes `[]`. evidence publishes `contexts: []`.
```json ```json
[ {
{ "source_id": "session-alpha",
"id": 10, "source_digest": "sha256:0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef",
"kind": "transcript_segment", "window_units": 1,
"text": "Aria casts Cure Wounds.", "selected_lanes": ["npcs", "spells"],
"ref": { "contexts": [
"source_id": "session-alpha", {
"start_unit_id": 10, "context_ref": {
"end_unit_id": 10 "source_id": "session-alpha",
"start_unit_id": 10,
"end_unit_id": 20
},
"evidence_refs": [
{
"lane_id": "spells",
"source_ref": {
"source_id": "session-alpha",
"start_unit_id": 10,
"end_unit_id": 10
}
}
],
"units": [
{
"id": 10,
"kind": "transcript_segment",
"text": "Aria casts Cure Wounds.",
"ref": {
"source_id": "session-alpha",
"start_unit_id": 10,
"end_unit_id": 10
}
},
{
"id": 20,
"kind": "transcript_segment",
"text": "The party regroups.",
"ref": {
"source_id": "session-alpha",
"start_unit_id": 20,
"end_unit_id": 20
}
}
]
} }
}, ]
{ }
"id": 20,
"kind": "transcript_segment",
"text": "The party regroups.",
"ref": {
"source_id": "session-alpha",
"start_unit_id": 20,
"end_unit_id": 20
}
}
]
``` ```
Each source unit has required `id`, `kind`, `text`, and self `ref` fields. Each context requires `context_ref`, `evidence_refs`, and `units` arrays.
`ref` contains `source_id`, `start_unit_id`, and `end_unit_id`, and both unit `context_ref` identifies the first and last included unit. Each evidence entry
endpoints identify that unit's `id`. A unit may also contain source-owned contains a selected `lane_id` and an original `source_ref`. A unit uses the
`metadata`, an open-ended JSON object. Fixed unit and reference fields are existing source-unit shape: required `id`, `kind`, `text`, and self `ref`, plus
strict: consumers must reject unknown fixed fields, malformed units, invalid optional JSON-object `metadata`. Fixed payload objects reject unknown fields;
self-references, units whose `source_id` differs from other units in the same unit metadata may contain application-defined JSON values.
excerpt, and a payload that is not the array described here.
The excerpt preserves each selected unit exactly as represented by the ## Citations And Context
validated generic source document. It does not add evidence-context-specific
annotations or reshape source-owned metadata.
## Selection And Citations `evidence_refs` are the authoritative citations. They identify the direct
references emitted by accepted normalized artifacts. `context_ref` and the
units collection include those cited units plus nearby source units selected by
the configured window. They are explanatory context, not widened citations.
The framework obtains direct source references only through typed evidence Only accepted outputs from the configured lane allowlist contribute. Rejected,
projections of accepted normalized artifacts in the configured lane allowlist. failed, absent, and lane-filtered outputs do not contribute. The artifact never
It validates each reference against the current source document, expands its contains raw input bytes, prompts, model responses, auxiliary reference
range by `window_units` source-unit positions on each side, clamps at document content, credentials, or filesystem paths.
boundaries, and takes the union of all expanded ranges. The output contains
each selected source unit once in source-document position order, regardless
of numeric unit IDs. Repeated references, overlapping windows, and citations
from multiple lanes do not duplicate a unit. Rejected, failed, absent,
inactive, and unselected lanes contribute nothing.
Normalized lane artifacts remain authoritative for citations and for which lane ## Ordering And Compatibility
cited a range. The excerpt has no lane attribution and must not be used to
reconstruct it. Its included nearby units provide reading context only; they
do not widen any citation in a lane artifact.
The excerpt contains at most every generic source unit once. It can therefore The selected lane allowlist is lexical. Contexts and units are in source
equal the complete generic source document when coverage is broad or the document position order, not numeric unit-ID order. Direct evidence entries
window is large. No byte-, token-, or compression-size guarantee is made, and are deterministically ordered by lane and source reference. Overlapping or
the framework does not truncate the excerpt to meet an arbitrary size limit. contiguous windows merge, and each source unit appears at most once in the
resulting contexts.
## Consumer Responsibilities And Data Handling
The artifact is additive to the JSON bundle and is not a lane payload, The artifact is additive to the JSON bundle and is not a lane payload,
normalized-output count, checkpoint, or generated reference. Consumers that normalized-output count, checkpoint, or generated reference. Consumers that
do not need it must tolerate an absent descriptor. Consumers that do use it do not need it must tolerate the absent optional descriptor. Consumers that do
should validate the descriptor and payload before use, retain the artifact with use it should preserve the artifact and its schema identity with the run
its schema identity when needed for a run record, and read citations from the provenance, and should treat its source text and metadata as sensitive durable
corresponding normalized lane artifacts. content.
The excerpt contains source-unit text and source-owned metadata and is durable
output. Treat it as sensitive source content, apply appropriate access controls
and retention, and do not assume its selected form is materially smaller or
less sensitive than the original input.

View File

@@ -9,7 +9,7 @@ Output configuration, including chunk-map and evidence-context publication, belo
## Bundle Layout ## Bundle Layout
All paths below are logical, relative, slash-separated bundle paths. The All paths below are logical, relative, slash-separated bundle paths. The
encoder always emits the first five JSON files below and adds lane or encoder always emits the first four JSON files below and adds lane or
pipeline-wide artifact files when their corresponding artifacts are available: pipeline-wide artifact files when their corresponding artifacts are available:
A subprocess caller first obtains the physical bundle root from the A subprocess caller first obtains the physical bundle root from the
@@ -21,11 +21,10 @@ root for the logical discovery described here.
| `index.json` | Entry point that names the other published files and lane payloads. | | `index.json` | Entry point that names the other published files and lane payloads. |
| `manifest.json` | Run provenance and result summaries. | | `manifest.json` | Run provenance and result summaries. |
| `rejected.json` | Rejected pipeline outputs. | | `rejected.json` | Rejected pipeline outputs. |
| `warnings.json` | Actionable process-degradation warnings. | | `warnings.json` | Accepted-output and run warnings. |
| `diagnostics.json` | Accepted-artifact quality advisories and normalization observations. |
| `lanes/<safe-lane-id>.json` | One normalized artifact payload for each lane. | | `lanes/<safe-lane-id>.json` | One normalized artifact payload for each lane. |
| `chunk-map.json` | Optional accepted chunk map, when its export is enabled and available. | | `chunk-map.json` | Optional accepted chunk map, when its export is enabled and available. |
| `evidence-context.json` | Optional selected source-unit excerpt, when evidence publication is enabled. | | `evidence-context.json` | Optional source-context artifact, when evidence publication is enabled. |
JSON files are pretty-printed with a trailing newline. Lane payloads are JSON files are pretty-printed with a trailing newline. Lane payloads are
accepted only when their media type is `application/json`. accepted only when their media type is `application/json`.
@@ -40,8 +39,7 @@ normalized lanes has this valid minimal index:
"manifest_file": "manifest.json", "manifest_file": "manifest.json",
"output_files": [], "output_files": [],
"rejected_file": "rejected.json", "rejected_file": "rejected.json",
"warnings_file": "warnings.json", "warnings_file": "warnings.json"
"diagnostics_file": "diagnostics.json"
} }
``` ```
@@ -51,7 +49,6 @@ normalized lanes has this valid minimal index:
| `output_files` | Yes | Lane descriptors sorted by `lane_id`. | | `output_files` | Yes | Lane descriptors sorted by `lane_id`. |
| `rejected_file` | Yes | Always `rejected.json`. | | `rejected_file` | Yes | Always `rejected.json`. |
| `warnings_file` | Yes | Always `warnings.json`. | | `warnings_file` | Yes | Always `warnings.json`. |
| `diagnostics_file` | Yes | Always `diagnostics.json`. |
| `chunk_map` | No | Descriptor for the pipeline-wide `chunk-map.json`; never a lane descriptor. | | `chunk_map` | No | Descriptor for the pipeline-wide `chunk-map.json`; never a lane descriptor. |
| `evidence_context` | No | Descriptor for the pipeline-wide `evidence-context.json`; never a lane descriptor. | | `evidence_context` | No | Descriptor for the pipeline-wide `evidence-context.json`; never a lane descriptor. |
@@ -74,15 +71,11 @@ output encoding fail.
Each `lanes/<safe-lane-id>.json` file is the codec-owned normalized JSON for Each `lanes/<safe-lane-id>.json` file is the codec-owned normalized JSON for
that lane. Consumers should use the index descriptors schema identity rather that lane. Consumers should use the index descriptors schema identity rather
than infer a lane schema from its name. The current D&D payload contracts are than infer a lane schema from its name. The current D&D payload contracts are
[spells](dnd-spell-artifacts.md), [NPC registry](dnd-npc-registry-artifacts.md), [spells](dnd-spell-artifacts.md), [NPCs](dnd-npc-artifacts.md),
[NPC occurrences](dnd-npc-occurrence-artifacts.md), [NPC interactions](dnd-npc-interaction-artifacts.md),
[combat turns](dnd-combat-turn-artifacts.md), [combat turns](dnd-combat-turn-artifacts.md),
[item registry](dnd-item-registry-artifacts.md), [item events](dnd-item-event-artifacts.md), and
[item occurrences](dnd-item-occurrence-artifacts.md), [scene descriptions](dnd-scene-description-artifacts.md).
[scene descriptions](dnd-scene-description-artifacts.md),
[enemy events](dnd-enemy-event-artifacts.md),
[location registry](dnd-location-registry-artifacts.md), and
[location occurrences](dnd-location-occurrence-artifacts.md).
## `manifest.json` ## `manifest.json`
@@ -95,7 +88,7 @@ group into the following externally observable summaries:
| Run identity and result | `run_id`, `pipeline_id`, `pipeline_digest`, `schema_version`, `validation_status`, `started_at`, `completed_at` | | Run identity and result | `run_id`, `pipeline_id`, `pipeline_digest`, `schema_version`, `validation_status`, `started_at`, `completed_at` |
| Resolved components | `input_module`, `chunker`, `extractors`, `merger`, `normalizer`, `output_encoder`, `artifact_lanes`, `validator_chains`, `module_metadata` | | Resolved components | `input_module`, `chunker`, `extractors`, `merger`, `normalizer`, `output_encoder`, `artifact_lanes`, `validator_chains`, `module_metadata` |
| Source and references | `source_digests`, `references` | | Source and references | `source_digests`, `references` |
| Published result summaries | `normalized_outputs`, `rejected_outputs`, `validation_summaries` | | Published result summaries | `normalized_outputs`, `rejected_outputs` |
| Execution summaries | `chunk_plan`, `checkpoint_decisions`, `llm_profiles`, `metadata` | | Execution summaries | `chunk_plan`, `checkpoint_decisions`, `llm_profiles`, `metadata` |
`references` records provenance such as the target, slot, origin, digest, `references` records provenance such as the target, slot, origin, digest,
@@ -105,87 +98,16 @@ summarize results without embedding lane payload bytes. A chunk-plan summary is
provenance for the plan used by this run; cache records, debug artifacts, and provenance for the plan used by this run; cache records, debug artifacts, and
other operational state are not published as bundle files. other operational state are not published as bundle files.
Each `validation_summaries` entry is a bounded outcome for one producer result. ## Rejections And Warnings
It has required `status`, `producer_attempt_count`, and `terminal_action`;
the stage and affected step, lane, module, or chunk identity are present when
applicable. `status` is `complete`, `rejected`, or `incomplete`.
`rejecting_validators`, `reason_codes`, and `incomplete_validators` preserve
configured validator order and omit later duplicates. Entries contain no raw
candidate response, correction guidance, validator diagnostic message, or
artifact payload. The same shape may appear as `validation` on an affected
rejection entry.
When present, `metadata.session_id` is the effective non-secret routing
correlation identifier used for the run. It can be visible to providers and is
not a substitute for a cache or checkpoint identity. Its generation and
override behavior are defined by the [CLI reference](../cli.md#run).
Each `llm_profiles` entry identifies effective, non-secret LLM execution
provenance:
| Field | Required | Meaning |
| --- | --- | --- |
| `id` | Yes | Selected PromptKit profile identifier. |
| `provider` | No | Notarius adapter provider identifier. |
| `model` | No | Effective provider model identifier. |
| `backend_id` | No | Effective PromptKit backend registration identifier. Endpoint-only profiles omit it. |
| `reasoning_effort` | No | Effective opaque provider reasoning setting. An empty or explicitly cleared setting is omitted. |
These values describe observed execution; they are not a backend-registration
interface. Entries that differ by backend or effective reasoning remain
distinct even when their profile, provider, and model are otherwise equal.
## Rejections, Warnings, And Diagnostics
`rejected.json` is always an object with a `rejected` array. Each entry has `rejected.json` is always an object with a `rejected` array. Each entry has
required `stage` and `message`; `step_id`, `lane_id`, `module_key`, `chunk_id`, required `stage` and `message`; `step_id`, `lane_id`, `module_key`, `chunk_id`,
`chunk_index`, `validator_name`, `reason_code`, `attempt_count`, and `chunk_index`, `validator_name`, `reason_code`, `attempt_count`, and
`diagnostic_artifact_path` are present only when applicable. An entry may also `diagnostic_artifact_path` are present only when applicable.
contain the bounded `validation` summary described above; the existing singular
validator and reason fields remain the first configured rejection for
compatibility.
`warnings.json` is always the `notarius.warnings.v2` envelope: `warnings.json` is always an object with a `warnings` array. Each warning has
`reason_code` and `message`; `scope` is optional. Both arrays are empty when
```json there is nothing to report.
{
"schema_version": "notarius.warnings.v2",
"group_count": 0,
"occurrence_count": 0,
"groups": []
}
```
It contains only process warnings. `group_count` is exact, and
`occurrence_count` is the exact sum of its group occurrence counts.
`diagnostics.json` is always the `notarius.diagnostics.v1` envelope:
```json
{
"schema_version": "notarius.diagnostics.v1",
"group_count": 0,
"occurrence_count": 0,
"truncated": false,
"unrepresented_occurrence_count": 0,
"groups": []
}
```
It contains only advisory and observation groups. `group_count` counts groups
represented in `groups`; `occurrence_count` includes both represented and
unrepresented occurrences. When `truncated` is true,
`unrepresented_occurrence_count` is the exact number omitted from group
representation.
Each group has `disposition`, `category`, `reason_code`, framework-owned
`origin`, exact `occurrence_count`, bounded `samples`, and
`omitted_sample_count`. Samples carry safe `scope` and `message`, plus a chunk
ID and zero-based chunk index when applicable. A group retains at most three
distinct samples. The framework fails rather than truncating actionable
warnings beyond 128 groups; it represents at most 256 advisory/observation
groups and records further occurrences through the diagnostic truncation
fields above.
## Compatibility ## Compatibility

View File

@@ -1,148 +0,0 @@
# PromptKit Integration
Notarius pins
[`gitea.maximumdirect.net/eric/promptkit` v0.9.0](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.9.0)
as its in-process prompt engine. The upstream
[Go package consumer guide](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.9.0/docs/consumers/pkg-promptkit.md)
owns the public engine API, and the upstream
[format reference](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.9.0/docs/formats.md)
owns prompt, profile, and schema file contracts.
## Supported Boundary
Notarius relies on the root `promptkit` package to:
- construct an `Engine` with filesystem-backed prompt, schema, and optional
operator and application-fallback profile sources;
- prepare one frozen execution from a `RunRequest` with named inline artifacts,
variables, a direct session ID, prompt identity, profile selection, and
optional appended rendered messages, then
record credential-redacted details and run that exact execution;
- return rendered debug material, validated structured output, selected
profile, backend, effective model metadata, and token usage;
- register the optional conventional `local` backend through `BackendLocal`,
`LocalBackend`, and `WithBackend`;
- distinguish structured-output validation failure from execution failure; and
- identify a missing explicit profile through `ErrProfileNotFound` and backend
admission exhaustion through `ErrCapacityExceeded`.
The pinned
[`BackendLocal`, `LocalBackend`, and `WithBackend` API](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.9.0/backends.go)
owns the registration and backend-capacity contract.
For one completion, the adapter calls `PrepareExecution`, takes a
caller-owned `Details` snapshot, and calls `RunPrepared` for that same opaque
prepared execution. It defers `Discard` for every unexecuted handle. Explicit
profile preflight uses `Engine.InspectProfile`; it does not prepare a synthetic
prompt. PromptKit's prepared handle, inspection result, and capacity-error
types stay inside the Notarius LLM adapter.
When a PromptKit profile and runtime override leave `temperature`, `max_tokens`,
or `top_p` unset, Notarius leaves that control unset as well. Compatible
providers therefore apply their own defaults; an operator that requires a
specific sampling value must select it explicitly in the profile or runtime
override.
Notarius does not use PromptKit's optional `ArtifactReader`. It materializes
source and reference content itself and supplies owned inline artifacts at the
adapter boundary. It also retains responsibility for pipeline retries,
scheduling, debug persistence, redaction, profile provenance, and conversion
from private model responses into durable domain artifacts.
Notarius sends one stable effective session through PromptKit's direct session
field, which is authoritative for provider session behavior. It also retains
the same value as the `session_id` prompt variable for maintained prompt
compatibility. The generated identifier is 76 ASCII characters, within
PromptKit v0.9.0's 256-code-point session limit. Session IDs are non-secret
correlation identifiers and may be exposed to providers and provider
observability. The CLI contract owns generation and override behavior.
Notarius records PromptKit's selected backend ID and effective reasoning
setting as optional run-manifest provenance. Endpoint-only profiles have no
backend ID. Debug prompt material also retains the selected backend ID and
PromptKit's stable lower-case `effective_model_params` JSON, which may include
`backend_id`. Notarius production configuration exposes one optional
conventional `local` registration. It does not expose a general user-defined
PromptKit backend registry. Endpoint-only profiles remain supported unchanged.
Notarius retains its application-wide scheduled client around the PromptKit
adapter. PromptKit may apply a narrower limit for the selected backend;
endpoint-only profiles have no such backend limit. The adapter translates
PromptKit capacity rejection into the provider-neutral Notarius
`ErrLLMCapacityExceeded` contract. It may include the normalized selected
backend ID in safe diagnostic context, without exposing PromptKit's capacity
error type, and leaves retries to the calling pipeline stage.
## Profile Sources And Compatibility
Notarius gives PromptKit the configured operator profile source, registered
application fallback profile assets, and optional backend registration through
the same construction path for inspection and execution. PromptKit owns the
resulting source precedence and strict profile parsing: a matching operator
profile is a complete replacement for a fallback or built-in profile, while an
invalid matching document fails instead of falling through. The operator
configuration and deployment workflow are defined in
[Configuration](../config.md#promptkit-profiles) and
[Operations](../operations.md#promptkit-profile-deployment).
PromptKit owns `base_profile` resolution under its
[pinned format rules](https://gitea.maximumdirect.net/eric/promptkit/src/tag/v0.9.0/docs/formats.md).
Notarius records the selected leaf identity and resolved target without parsing
or merging inheritance. An unset filesystem `api_key_env` is optional and may
reach the provider without authorization, which can result in a 401 or 403.
PromptKit v0.9.0 accepts only the `developer`, `system`, `user`, and
`assistant` text-chat roles after normalizing case and surrounding whitespace.
Maintained Notarius prompt definitions use only `system` and `user`.
For application-owned semantic correction, Notarius uses PromptKit v0.9.0's
`RunRequest.AppendedMessages` after the ordinary rendered prompt. It supplies
exactly two messages in order: the latest validated producer response with
role `assistant`, then deterministic validation guidance with role `user`.
It never exposes a general caller-selected role API, accumulates earlier
correction turns, or changes the ordinary prompt prefix. Ordinary requests
leave appended messages unset.
PromptKit preserves supplied content but does not own Notarius's correction
bounds. Notarius rejects invalid UTF-8, blank, or oversized assistant material
(at most 1 MiB), guidance (at most 64 KiB), and combined content (at most
1,114,112 bytes) before preparing the request. The transport-neutral
application contract owns defensive copying and these limits. Default request
and terminal summaries retain only safe counts, digests, identities, and usage;
complete appended messages remain limited to the explicitly requested detailed
debug trace.
PromptKit now obtains its maintained OpenRouter and Rakestrawhome backend and
profile catalogs from independently versioned transitive modules. Notarius
does not import or register either catalog; PromptKit retains catalog source,
identity, precedence, credential, and capacity ownership.
Notarius supports this boundary against PromptKit v0.9.0. Its fallback source,
prepared-execution, inspection, and typed capacity APIs are used as public
upstream contracts; other PromptKit APIs or file-format behavior are not
implicitly supported. A dependency upgrade requires reviewing the adapter,
profile-source construction, and this compatibility statement against the
pinned upstream documentation.
## Notarius Ownership
[LLM Runtime Internals](../internal/llm.md) describes how Notarius mounts
module assets, maps its transport-neutral completion contract, prepares and
executes requests, validates output, records provenance, captures debug
material, redacts errors, and preserves timeout ownership.
PromptKit provider error details do not cross the ordinary completion boundary.
Notarius exposes a provider-neutral generation category and optional status;
redacted provider details are retained only in requested debug material.
[D&D Module Internals](../internal/dnd.md) owns the embedded
`dnd-extraction` fallback profile and the maintained D&D prompt defaults.
[Configuration](../config.md#promptkit-profiles) defines how a Notarius
configuration selects one PromptKit profile source and optionally registers
the conventional local backend.
PromptKit API or format changes outside this boundary are not implicitly
supported. Updating the pinned version requires reviewing the adapter and
profile/configuration contracts against the upstream documentation. Maintained
production prompts use PromptKit's bounded structural-repair contract; their
current declaration is one additional repair attempt. Notarius retains the
transport-neutral boundary and does not expose PromptKit types to modules.

View File

@@ -9,58 +9,36 @@ Command syntax, streams, and exit statuses are defined in the
## Schema ## Schema
The current schema version is `notarius.run-result.v2`. The current schema version is `notarius.run-result.v1`.
| Field | Required | Meaning | | Field | Required | Meaning |
| --- | --- | --- | | --- | --- | --- |
| `schema_version` | Yes | Exactly `notarius.run-result.v2`. | | `schema_version` | Yes | Exactly `notarius.run-result.v1`. |
| `run_id` | Yes | The finalized Notarius run identifier. | | `run_id` | Yes | The finalized Notarius run identifier. |
| `pipeline_id` | Yes | The effective pipeline identifier. | | `pipeline_id` | Yes | The effective pipeline identifier. |
| `output_directory` | Yes | Absolute path to the published, run-specific output bundle. | | `output_directory` | Yes | Absolute path to the published, run-specific output bundle. |
| `index_file` | For the production JSON output | Logical path `index.json`; omitted for other output modules. | | `index_file` | For the production JSON output | Logical path `index.json`; omitted for other output modules. |
| `normalized_output_count` | Yes | Number of final normalized outputs. | | `normalized_output_count` | Yes | Number of final normalized outputs. |
| `rejected_output_count` | Yes | Number of recorded rejected outputs. | | `rejected_output_count` | Yes | Number of recorded rejected outputs. |
| `warning_group_count` | Yes | Exact number of actionable warning groups. | | `warning_count` | Yes | Number of final run warnings. |
| `warning_occurrence_count` | Yes | Exact occurrences represented by actionable warning groups. |
| `diagnostic_group_count` | Yes | Number of represented advisory and observation groups. |
| `diagnostic_occurrence_count` | Yes | Advisory and observation occurrences, including unrepresented occurrences. |
| `diagnostics_truncated` | Yes | Whether advisory/observation group representation was truncated. |
| `validation_status` | Yes | The final run manifest validation status. | | `validation_status` | Yes | The final run manifest validation status. |
| `validation_summaries` | No | Bounded per-producer validation outcomes; present when producer work ran. |
| `debug_directory` | No | Absolute path to the run-specific debug bundle when requested debug capture completed. | | `debug_directory` | No | Absolute path to the run-specific debug bundle when requested debug capture completed. |
For the production `json` output module, `index_file` is present only when the For the production `json` output module, `index_file` is present only when the
completed run returned exactly one logical output file named `index.json`. completed run returned exactly one logical output file named `index.json`.
For another output module, its absence does not indicate a failed run. For another output module, its absence does not indicate a failed run.
`validation_status` is `approved`, `rejected`, or `incomplete`; `incomplete`
means one or more otherwise accepted results advanced under validator-failure
`warn_continue`.
```json ```json
{ {
"schema_version": "notarius.run-result.v2", "schema_version": "notarius.run-result.v1",
"run_id": "run-1770000000000000000-0123456789abcdef0123456789abcdef", "run_id": "run-1770000000000000000-0123456789abcdef0123456789abcdef",
"pipeline_id": "dnd-session", "pipeline_id": "dnd-session",
"output_directory": "/work/results/run-1770000000000000000-0123456789abcdef0123456789abcdef", "output_directory": "/work/results/run-1770000000000000000-0123456789abcdef0123456789abcdef",
"index_file": "index.json", "index_file": "index.json",
"normalized_output_count": 6, "normalized_output_count": 6,
"rejected_output_count": 2, "rejected_output_count": 2,
"warning_group_count": 1, "warning_count": 1,
"warning_occurrence_count": 2, "validation_status": "rejected"
"diagnostic_group_count": 3,
"diagnostic_occurrence_count": 5,
"diagnostics_truncated": false,
"validation_status": "incomplete",
"validation_summaries": [
{
"stage": "extract",
"lane_id": "spells",
"status": "incomplete",
"incomplete_validators": ["dnd/spells/source_refs"],
"producer_attempt_count": 1,
"terminal_action": "warn_continue"
}
]
} }
``` ```
@@ -70,11 +48,8 @@ means one or more otherwise accepted results advanced under validator-failure
paths. They identify the paths used by Notarius and do not resolve symlinks. paths. They identify the paths used by Notarius and do not resolve symlinks.
`output_directory` is the run-specific bundle, not the configured output root. `output_directory` is the run-specific bundle, not the configured output root.
The receipt is a summary and discovery document. Its optional validation The receipt is a summary and discovery document. It does not contain lane
summaries contain only stable status, identity, validator names, reason codes, descriptors, payloads, manifest data, rejections, warnings, or file contents.
attempt counts, and terminal actions. It does not contain lane descriptors,
payloads, manifest payloads, rejection messages, warnings, raw model responses,
correction guidance, or file contents.
For the production JSON output, resolve `index_file` beneath For the production JSON output, resolve `index_file` beneath
`output_directory`, reject path escapes, and use the `output_directory`, reject path escapes, and use the
[Published JSON Output contract](json-output.md) to discover logical files and [Published JSON Output contract](json-output.md) to discover logical files and

View File

@@ -24,14 +24,10 @@ preparation, and runner mechanics after their inputs are supplied.
## Dispatch And Configuration Handoff ## Dispatch And Configuration Handoff
The root dispatcher handles help, version reporting, configuration validation, The root dispatcher handles help, configuration validation, pipeline listing,
pipeline listing, and a pipeline run. Version reporting resolves build and a pipeline run. It normalizes injectable options before dispatch so that a
information through `internal/buildinfo` before production composition, so the missing production dependency fails as a command error rather than reaching
diagnostic remains available without configuration or runtime collaborators. execution.
The public syntax, streams, exit classes, and version semantics are defined by
the [CLI reference](../cli.md). Other root commands normalize injectable
options before dispatch so that a missing production dependency fails as a
command error rather than reaching execution.
Commands that need configuration use one shared loader. The CLI discovers the Commands that need configuration use one shared loader. The CLI discovers the
file, parses it through **internal/core/config**, starts from defaults, applies file, parses it through **internal/core/config**, starts from defaults, applies
@@ -42,14 +38,9 @@ in [Configuration Internals](configuration.md).
Configuration validation without a selected pipeline checks structural Configuration validation without a selected pipeline checks structural
configuration only. Validation with a selected pipeline also builds the configuration only. Validation with a selected pipeline also builds the
effective catalog, resolves the pipeline, and verifies every explicit effective effective catalog, resolves the pipeline, and verifies explicitly selected
PromptKit profile. Selected LLM-backed input, chunk, lane, output, and validator Scriptorium profiles. Pipeline listing validates configuration before returning
profiles are inspected normalized, sorted identifiers.
against the configured PromptKit source and backend registrations without
loading a prompt or performing generation, so an unknown or invalid profile
fails before pipeline preparation. Credential availability remains an
execution-time concern. Pipeline listing validates configuration before
returning normalized, sorted identifiers.
## Production Composition ## Production Composition
@@ -60,24 +51,12 @@ catalog used for resolution and the concrete constructors used for preparation.
Tests may provide a catalog or registries instead; production code must not Tests may provide a catalog or registries instead; production code must not
silently merge an injected partial catalog with production registrations. silently merge an injected partial catalog with production registrations.
The production LLM factory builds one PromptKit-backed client from the resolved The production LLM factory builds the Scriptorium-backed client from resolved
**promptkit.profile_dir** or **promptkit.profile_file** source, attaches the configuration, creates one scheduler from the effective global LLM limit, and
profile-provenance recorder, creates one scheduler from the effective global wraps the client before it reaches modules. Registration and LLM construction
LLM limit, and wraps the client before it reaches modules. Registration and LLM errors are returned before a pipeline is prepared. Concrete module keys and
construction errors are returned before a pipeline is prepared. Configuration validator chains are public configuration choices and remain documented in
field definitions remain in [Configuration](../config.md#promptkit-profiles); [Configuration](../config.md).
the D&D registrar's fallback profile assets and the adapter mechanics remain in
[LLM Runtime](llm.md).
The factory also accepts `LLMRuntimeOverrides`, whose reasoning pointer
preserves inherit, replace, and clear states across the composition boundary.
Run orchestration constructs this value from the mutually exclusive
`--reasoning-effort` and `--clear-reasoning-effort` controls. Absence preserves
a nil pointer, replacement is trimmed, and clear uses a non-nil empty string.
The same override reaches the one shared production client, checkpoint
identity, and debug invocation metadata. Persistent reasoning configuration
remains owned by PromptKit profiles; Notarius configuration has no reasoning
field.
## Run Orchestration ## Run Orchestration
@@ -89,15 +68,12 @@ handoff:
2. create and validate a safe run identity, then allocate a debug bundle only 2. create and validate a safe run identity, then allocate a debug bundle only
when requested; when requested;
3. build the effective catalog, resolve requested reference changes, resolve 3. build the effective catalog, resolve requested reference changes, resolve
the effective pipeline, and inspect its explicit effective PromptKit the effective pipeline, and verify explicit Scriptorium profiles;
profiles;
4. materialize external or generated references and record redacted invocation 4. materialize external or generated references and record redacted invocation
and resolution provenance when debug capture is enabled; and resolution provenance when debug capture is enabled;
5. construct registries, the scheduled LLM client, and prepared modules; 5. construct registries, the scheduled LLM client, prepared modules, and the
6. read the source input once, resolve its effective session from the explicit requested cache/checkpoint collaborators;
override or resolved input module and raw bytes, then construct requested 6. read the source input and invoke the framework runner; and
checkpoint collaborators and invoke the framework runner with that same
value; and
7. write the runner's logical output files only after a successful run, then 7. write the runner's logical output files only after a successful run, then
complete the command report and user-facing result. complete the command report and user-facing result.
@@ -108,13 +84,6 @@ final command result. Detailed state lifecycle, resume handling, and physical
path confinement are maintained in [Run State Internals](state.md) and path confinement are maintained in [Run State Internals](state.md) and
[Operations](../operations.md). [Operations](../operations.md).
The CLI owns the versioned generated-session policy and resolves the sole
effective value before checkpoint construction. It records that value in the
final debug invocation summary when capture is enabled and passes it unchanged
to checkpoint identity and `pipeline.RunInput`. The public flag and stability
contract are defined by the [CLI reference](../cli.md#run); framework and LLM
packages only transport the supplied value.
For `run --json`, the CLI constructs and encodes its private run-result receipt For `run --json`, the CLI constructs and encodes its private run-result receipt
after a successful runner result is available, before it publishes logical after a successful runner result is available, before it publishes logical
output files. It writes the prepared receipt to standard output only after output files. It writes the prepared receipt to standard output only after

View File

@@ -39,22 +39,14 @@ This establishes the public precedence order without giving environment input a
second file schema. Loading and application reject malformed YAML, unsupported second file schema. Loading and application reject malformed YAML, unsupported
file versions, unknown fields, invalid values, and identifiers that are empty file versions, unknown fields, invalid values, and identifiers that are empty
or collide after whitespace normalization. The file application also makes the or collide after whitespace normalization. The file application also makes the
effective extraction-worker default follow the effective LLM limit. A present effective extraction-worker default follow the effective LLM limit.
PromptKit local-backend object requires and trims its endpoint, defaults its
omitted concurrency limit to zero, and is copied so the parsed file model
cannot alias the populated **Config**. A pipeline `llm_profile` is
presence-aware: omission remains empty, while a present blank value is
rejected and a non-empty file value is trimmed before it reaches **Config**.
**Config.Validate** checks configuration-only invariants before resolution. It **Config.Validate** checks configuration-only invariants before resolution. It
rejects incompatible profile sources, invalid state-surface values, unsupported rejects incompatible profile sources, invalid state-surface values, unsupported
concurrency settings, malformed bindings and references, invalid retries, and concurrency settings, malformed bindings and references, invalid retries, and
invalid pipeline, step, or lane structure. PromptKit local-backend validation invalid pipeline, step, or lane structure. Its errors retain the closest known
accepts only an absolute HTTP or HTTPS endpoint with a host and no user pipeline, lane, and binding context. It deliberately does not require modules
information, query, or fragment, and rejects a negative local concurrency to be registered: that requires a catalog and belongs to resolution.
limit. Its errors retain the closest known pipeline, lane, and binding context.
It deliberately does not require modules to be registered: that requires a
catalog and belongs to resolution.
The exact user-selectable values and validation rules are defined in The exact user-selectable values and validation rules are defined in
[Configuration](../config.md). Keep additions to the file model, an [Configuration](../config.md). Keep additions to the file model, an
@@ -64,34 +56,22 @@ environment override, its validation, and that reference in the same change.
**Config.Resolve** first recomputes derived concurrency defaults and validates **Config.Resolve** first recomputes derived concurrency defaults and validates
the configuration. It normalizes the requested pipeline ID, copies the selected the configuration. It normalizes the requested pipeline ID, copies the selected
profile, and passes the non-empty command-level LLM profile override, requested profile, applies a non-empty command-level LLM profile override to the
lane selection, and reference changes to the framework resolver. LLM-capable stage bindings, and calls the framework resolver with the requested
lane selection and reference changes.
After module and validator selection, the resolver applies the effective The command-level override does not replace an explicitly selected validator
profile policy to LLM-backed bindings only: command override, binding profile, profile. Validator bindings remain part of the resolved validator chain and
pipeline profile, then the prompt default. Deterministic bindings remain are resolved under their own declared configuration.
profile-free, and no second inheritance decision occurs during execution. The
public field definitions and precedence are owned by
[Configuration](../config.md#pipelines).
The resolver retains configured `validation_policy` overrides and derives one
detached concrete terminal policy for the chunk producer and every lane's
extract, merge, and normalize producers. That field-by-field inheritance is
complete before preparation, and the effective values contribute to pipeline
and checkpoint identity; execution does not interpret configuration defaults.
The framework resolver supplies defaults, selects lanes, resolves validator The framework resolver supplies defaults, selects lanes, resolves validator
chains, checks registered module and artifact compatibility, validates module chains, checks registered module and artifact compatibility, validates module
options, and returns the fixed ordered pipeline shape. Positive validator retry options, and returns the fixed ordered pipeline shape. The resulting
budgets require an LLM-backed selected validator; deterministic validators are **EffectiveConfig** retains the selected ID, requested selection and reference
rejected during resolution. Eligible LLM-backed producer specifications also
contribute their declared correction protocol to the resolved metadata. The
resulting **EffectiveConfig** retains the selected ID, requested selection and reference
changes, a clone of the input configuration, and the resolved pipeline. changes, a clone of the input configuration, and the resolved pipeline.
Callers may therefore retain or modify their input slices and maps without Callers may therefore retain or modify their input slices and maps without
changing the resolved result, and later consumers cannot mutate the original changing the resolved result, and later consumers cannot mutate the original
configuration through the effective value. This ownership includes the nested configuration through the effective value.
PromptKit local-backend value.
Resolution failures stop before module construction and source parsing. They Resolution failures stop before module construction and source parsing. They
include an error path for an unconfigured pipeline, missing module, missing include an error path for an unconfigured pipeline, missing module, missing
@@ -103,8 +83,7 @@ runtime error class described in the [CLI reference](../cli.md#output-streams-an
The framework assigns the resolved pipeline a deterministic SHA-256 digest The framework assigns the resolved pipeline a deterministic SHA-256 digest
after defaults, lane selection, module bindings, reference bindings, validator after defaults, lane selection, module bindings, reference bindings, validator
chains, selected correction protocols, effective LLM profiles, and artifact chains, and artifact schema identity have been resolved. The digest excludes
schema identity have been resolved. The digest excludes
its own stored value. It identifies resolved composition rather than raw YAML its own stored value. It identifies resolved composition rather than raw YAML
bytes, a debug payload, or all runtime state. The CLI records it as invocation bytes, a debug payload, or all runtime state. The CLI records it as invocation
provenance before execution; cache and checkpoint identity have additional provenance before execution; cache and checkpoint identity have additional
@@ -115,12 +94,9 @@ Configuration summaries must use **Redacted**, **RedactedSummaryPayload**, or
Those methods copy every binding and nested option container, replace values Those methods copy every binding and nested option container, replace values
whose key is credential-shaped with **[REDACTED]**, and omit materialized whose key is credential-shaped with **[REDACTED]**, and omit materialized
reference content while retaining safe binding and reference provenance. The reference content while retaining safe binding and reference provenance. The
payload must not alias the source configuration or resolved pipeline. payload must not alias the source configuration or resolved pipeline. This
PromptKit's local endpoint and concurrency limit are preserved as non-secret redaction is deliberately narrow: it protects configuration summaries and does
configuration metadata in the independently owned summary; the object contains not authorize recording arbitrary environment values or provider requests.
no credential value. This redaction is deliberately narrow: it protects
configuration summaries and does not authorize recording arbitrary environment
values or provider requests.
## Invariants To Preserve ## Invariants To Preserve

View File

@@ -7,36 +7,25 @@ selectable keys, bindings, reference syntax, and default validator chains.
## Durable Artifact Contracts ## Durable Artifact Contracts
The ten lanes have separate durable wire contracts. This guide deliberately The six lanes have separate durable wire contracts. This guide deliberately
does not repeat their JSON shapes or schemas. does not repeat their JSON shapes or schemas.
| Lane | Durable contract | | Lane | Durable contract |
| --- | --- | | --- | --- |
| Spells | [spell artifacts](../integrations/dnd-spell-artifacts.md) | | Spells | [spell artifacts](../integrations/dnd-spell-artifacts.md) |
| NPC registry | [NPC registry artifacts](../integrations/dnd-npc-registry-artifacts.md) | | NPCs | [NPC artifacts](../integrations/dnd-npc-artifacts.md) |
| Combat turns | [combat-turn artifacts](../integrations/dnd-combat-turn-artifacts.md) | | Combat turns | [combat-turn artifacts](../integrations/dnd-combat-turn-artifacts.md) |
| Item occurrences | [item-occurrence artifacts](../integrations/dnd-item-occurrence-artifacts.md) | | Item events | [item-event artifacts](../integrations/dnd-item-event-artifacts.md) |
| Item registry | [item-registry artifacts](../integrations/dnd-item-registry-artifacts.md) | | NPC interactions | [NPC-interaction artifacts](../integrations/dnd-npc-interaction-artifacts.md) |
| NPC occurrences | [NPC-occurrence artifacts](../integrations/dnd-npc-occurrence-artifacts.md) |
| Scene descriptions | [scene-description artifacts](../integrations/dnd-scene-description-artifacts.md) | | Scene descriptions | [scene-description artifacts](../integrations/dnd-scene-description-artifacts.md) |
| Enemy events | [enemy-event artifacts](../integrations/dnd-enemy-event-artifacts.md) |
| Location registry | [location-registry artifacts](../integrations/dnd-location-registry-artifacts.md) |
| Location occurrences | [location-occurrence artifacts](../integrations/dnd-location-occurrence-artifacts.md) |
## Family Composition ## Family Composition
The D&D registrar registers the familys artifact codecs, extractors, typed The D&D registrar registers the familys artifact codecs, extractors, typed
append-order mergers, normalizers, validators, prompt assets, fallback LLM append-order mergers, normalizers, validators, prompt assets, and default
profile asset, and default validator chains. Each extractor and normalizer has validator chains. Each extractor and normalizer has a stable module spec,
a stable module spec, explicit execution class, strict option decoding, and a strict option decoding, and a typed builder. Configuration remains the
typed builder. Scene chunking, every extractor, and NPC, location, and item-registry canonical owner of the exact keys and validator order.
normalization are registered as `llm_backed`; the remaining current D&D mergers
and normalizers are `deterministic`. The metadata is available to catalog inspection and
resolved-pipeline debug data and determines which selected bindings inherit the
pipeline profile. The registry normalizers use `single_response_v1`, forwarding
corrections to their reconciliation completion and retaining the accepted raw
proposal only as an owned model candidate. Configuration remains the canonical owner of the exact keys,
profile precedence, and validator order.
Private structured-LLM response schemas are deliberately minimal. They reject Private structured-LLM response schemas are deliberately minimal. They reject
invalid JSON structure, missing required fields, incompatible types, and invalid JSON structure, missing required fields, incompatible types, and
@@ -44,78 +33,21 @@ unknown fields, while preserving semantic candidates for deterministic
validation. Do not promote a private response envelope into a durable schema; validation. Do not promote a private response envelope into a durable schema;
the contracts above define durable data. the contracts above define durable data.
A D&D producer that declares `single_response_v1` forwards any supplied
semantic correction to its structured completion and returns an owned copy of
that completion's exact validated raw response as its model candidate. It does
not serialize normalized artifacts to create that candidate, so deterministic
identity, evidence, warning, and durable-schema behavior remains separate from
the model transport material.
## Prompt Construction ## Prompt Construction
D&D LLM-facing content lives beneath `assets/dnd/`. Each module contributes a D&D extractors assemble prompts from an ordered manifest of shared and
local `prompt.yaml` declaration and `instructions.md`; input-specific files module-owned assets. Reuse the shared D&D system, evidence, identity,
such as a catalog, registry, grounding projection, or candidate collection are reference, and transcript assets instead of copying their text into individual
local only when that module needs them. New extractor content uses its feature modules. A manifests declared sequence, including cache-control placement, is
subtree, while families with both extraction and normalization content use their part of the prompt behavior, and the chunk transcript is the final message.
`extract` and `normalize` subtrees. Shared visual-provenance fragments use Preserve that order when changing an extractor or its assets so prompt-cache
the `common-dnd-` prefix. Production lane code belongs with its D&D codec, behavior remains stable.
extractor, normalizer, and validator packages; registry projections and
identity helpers remain in their owning entity packages rather than in a
consumer lane.
The owning modules manifest is the source of truth for which local and shared All extractors use the shared prompt-input preparation rules. The current chunk
assets are selected, their mount paths, their message order, cache controls, is copied into transcript material; player, party, glossary, and compatible
and the files included in its prompt fingerprint. Shared fragments belong to campaign references are context for disambiguation, not source evidence.
the D&D shared implementation and are selected by name rather than copied into Reference prompt material is canonically ordered before it is rendered, which
module directories. The root `assets` package is a content-only boundary; its keeps equivalent inputs stable across runs.
physical ownership and rationale are defined by
[ADR-0011](../adr/0011-centralize-llm-assets.md).
Put each rule at its narrowest owner:
- universal behavior belongs in the shared system asset;
- D&D-family behavior belongs in a selected `common-dnd-` asset;
- rules for an input projection belong with that input asset;
- lane-specific policy belongs in the modules `instructions.md`; and
- transport-envelope shape belongs in the private response schema.
A rule is eligible for the system prompt only when every D&D LLM prompt needs
it regardless of lane, inputs, or response shape. Module instructions must not
repeat rules selected from shared assets or schemas. Reintroduce such repetition
only after observational evaluation with representative transcripts shows that
it improves results at the intended target models and cost; structural prompt
tests alone are not that evidence.
Every maintained D&D LLM prompt selects `dnd-extraction` as its default
profile. The D&D registrar registers the fallback, while an operator can
replace it with a complete profile of the same ID from the configured PromptKit
source. Deployment profile selection is documented in
[Configuration](../config.md#promptkit-profiles).
The D&D transcript assets have distinct consumers. Scene chunking consumes the
complete-session `common-dnd-transcript-full.md`, while extraction prompts
consume the current-chunk `common-dnd-transcript-chunk.md`. NPC, location, and
item normalization instead mount the generic semantic-reconciliation
candidate and transcript-window presentation assets. Player, party, glossary,
and compatible campaign references provide disambiguating context only when
declared by the active prompt; they never establish evidence. Reference
material is canonically ordered before rendering so equivalent inputs remain
stable.
Extraction prompts render the common system and identity messages first, then
cached campaign references and the cached chunk transcript. Evidence policy and
any lane-specific registry, catalog, or grounding projection follow that
prefix. The final module instructions message is ephemeral. This keeps the
reusable extraction prefix identical while preserving the lane-specific suffix.
Scene chunking intentionally uses a different order: system, cached campaign
references, uncached module instructions, then the final ephemeral full
transcript. Entity normalization also has its own order: D&D system, mandatory
generic protocol, ephemeral domain semantic instructions, generic candidate
presentation, and final ephemeral generic transcript windows. These orders and
cache controls are prompt behavior; change them only through the owning
manifest and prompt declaration.
## Evidence, Candidates, And Normalization ## Evidence, Candidates, And Normalization
@@ -128,70 +60,19 @@ result.
Default chains keep responsibilities separate: structural validators assess the Default chains keep responsibilities separate: structural validators assess the
candidate, source-reference validators resolve cited ranges against the current candidate, source-reference validators resolve cited ranges against the current
source and require extraction evidence to stay within the current chunk, source, durable-schema validation checks an approved representation, and
durable-schema validation checks an approved representation, and
relatedness validators report advisory evidence concerns. The configured order relatedness validators report advisory evidence concerns. The configured order
is documented in is documented in
[Configuration](../config.md#production-validator-keys-and-default-chains). [Configuration](../config.md#production-validator-keys-and-default-chains).
Every D&D rejection describes the correction in transcript-grounded domain Normalizers are deterministic for spells, combat turns, item events, NPC
terms, using contextual names, artifact fields, and source segment ranges when interactions, and scene descriptions. They canonicalize display values and
useful. The guidance must not ask the model to reproduce durable entity IDs, evidence, use source-document order for stable output, and issue bounded
hashes, validator module keys, or reason codes. Those identifiers remain in warnings for changes or collapsed duplicates. The NPC normalizer is the
ordinary validation provenance; only the actionable semantic guidance is intentional exception: it first produces a deterministic candidate set, then
eligible for the correction prompt. uses a bounded structured-LLM proposal to reconcile identity groups. Invalid
or unusable proposals retain the deterministic result and surface retry or
Enemy-event extraction additionally rejects a second `engaged` observation for fallback diagnostics; the model does not directly replace durable records.
the same comparison identity within one scene-scoped result. Normalization may
combine results from distinct scenes, so it intentionally does not apply that
rule. Configuration owns the exact validator key and chain position.
Normalizers are deterministic for spells, combat turns, item occurrences, NPC
occurrences, scene descriptions, enemy events, and location occurrences. They
canonicalize display values and evidence, use source-document order for stable
output, and emit bounded normalization observations for changes or collapsed duplicates. NPC,
item, and location registry normalizers are intentional exceptions: each first
produces a deterministic candidate set, then may use a bounded structured-LLM
proposal to reconcile identity groups.
## Semantic Registry Reconciliation
The three registry normalizers instantiate the domain-neutral
`internal/framework/semanticreconcile` engine with default bounds. Each
eligible candidate receives a contiguous, one-based `candidate_id` for that
request. The model sees that handle, the candidate label and source-free
evidence ranges, plus bounded transcript windows; it returns only duplicate
groups of supplied handles and one supplied canonical handle per group. It
never returns names, evidence, durable IDs, or replacement records. Identical
labels and evidence remain independently selectable because their handles are
distinct.
The generic core owns the mandatory handle protocol, candidate and transcript
presentation, the private response schema, source-reference validation,
candidate and combined-material limits, structured completion, proposal
assessment, stable group ordering, and typed plan-application mechanics. The
D&D prompt contributes its system message and registry-specific semantic
instructions. The generic registrar registers the shared prompt and schema;
the D&D registrar registers each consuming prompt and the fallback profile.
Fewer than two eligible candidates skips the LLM without a semantic warning.
An exceeded bound also skips the call and preserves the deterministic
preprocessed registry, adding the registry's bounded fallback warning. Invalid
structured output or discarded proposal groups use the normalizer's existing
retry contract; retry exhaustion preserves the safe deterministic or
partially applied result and emits its bounded fallback warning. Provider,
transport, cancellation, and context-material failures remain execution
errors.
Application remains typed and registry-owned. All three policies select the
canonical member's normalized display name, union member evidence in source
order, preserve ungrouped records, and derive durable identity only after
consolidation. NPC IDs derive from the final name. Item IDs also derive from
the final name, and a typed guard prevents currency aliases from crossing
denominations or mixing currency with non-currency records. Location IDs
derive from the final name and final evidence, preserving same-name,
parent/child, and distinct physical-place identities. Registry warning scopes,
reason codes, and postconditions remain outside the generic core.
## Generated References And Grounding ## Generated References And Grounding
@@ -201,55 +82,30 @@ producer provenance; consumers resolve the handed-off artifact into an
immutable, validated projection for each operation. External files are checked immutable, validated projection for each operation. External files are checked
during preparation, while generated artifacts are resolved at the handoff. during preparation, while generated artifacts are resolved at the handoff.
NPC and item registry consumers receive names-only grounding. Location NPC registries are names-only grounding projections: they may canonicalize
consumers receive a contextual selector containing the canonical name and the actors for spells and combat turns and are required for NPC interactions, but
registry references needed to distinguish same-name places. The calling module they do not supply evidence. Scene-description registries are eligibility-only
resolves those supplied selections locally and maps them into the unchanged projections: they retain the current chunks classification data, not scene
durable ID/name pair; an unknown or ambiguous selection rejects the complete prose or evidence, and exist to route combat extraction.
occurrence result rather than accepting a partial mapping. The NPC registry
additionally supplies names-only actor grounding to spells, combat turns, and
enemy events.
Registry references establish a registry identity and may disambiguate a
selection, but never become occurrence evidence. Each occurrence keeps its own
current-transcript source references, even when it was grounded through the
same registry record.
Scene descriptions are eligibility-only projections: they retain current-chunk
classification data, not scene prose or evidence, and exist to route combat
extraction. Enemy-event extraction also projects combat turns to `actor` and
`turn_kind` and filters NPC occurrences to `combat_opponent` names and kinds.
These projections are guidance only and never event evidence.
## Lane-Specific Rules ## Lane-Specific Rules
The following differences are intentional and should remain explicit when a The following differences are intentional and should remain explicit when a
shared helper changes. shared helper changes.
Shared D&D text comparison is identified by `dnd.text_comparison.v1`. Any
semantic change requires an explicit policy-version review for every affected
identity, mapping, normalization, and validator policy; helper source is not a
checkpoint fingerprint.
| Lane | Intentional behavior | | Lane | Intentional behavior |
| --- | --- | | --- | --- |
| Spells | May use a spell-catalog overlay and optional NPC grounding; the catalog validator supplies domain-specific semantic checks. | | Spells | May use a spell-catalog overlay and optional NPC grounding; the catalog validator supplies domain-specific semantic checks. |
| NPC registry | Establishes transcript-grounded NPC identities, including factual third-party mentions, without assigning occurrence categories. It does not consume an NPC registry, and its normalizer is the LLM-assisted reconciliation exception described above. | | NPCs | Does not consume an NPC registry. Its normalizer is the LLM-assisted reconciliation exception described above. |
| Combat turns | Requires a scene-description artifact. It calls the LLM only for an exact `combat` classification; exact non-combat classifications return an accepted empty result, while missing or mismatched classifications return an empty result with a bounded warning. Optional NPC grounding never becomes evidence. | | Combat turns | Requires a scene-description artifact. It calls the LLM only for an exact `combat` classification; exact non-combat classifications return an accepted empty result, while missing or mismatched classifications return an empty result with a bounded warning. Optional NPC grounding never becomes evidence. |
| Item occurrences | Requires the normalized item registry for exact deterministic grounding at extraction and normalization. Campaign context may disambiguate, but the registry never becomes occurrence evidence. | | Item events | Uses campaign context for disambiguation but has no NPC-registry or scene-description dependency. |
| Item registry | Produces source-grounded item types and unique designations. Its LLM-assisted reconciliation is proposal-only, preserves distinct currency denominations and item types, and does not create per-instance identities. | | NPC interactions | Requires the normalized NPC registry at extraction and normalization, using it for canonical actor grounding only. |
| NPC occurrences | Requires the normalized NPC registry at extraction and normalization, using it for canonical actor grounding only. It separately emits cited current-transcript occurrence facts, including `mentioned`, rather than deriving them from registry provenance. |
| Scene descriptions | Produces the classifications consumed by combat routing; it does not consume an NPC registry or provide evidence for combat artifacts. | | Scene descriptions | Produces the classifications consumed by combat routing; it does not consume an NPC registry or provide evidence for combat artifacts. |
| Enemy events | Requires NPC, scene-description, combat-turn, and NPC-occurrence artifacts. It calls the LLM only for an exact `combat` classification, records ordered observations rather than terminal state, and normalizes recognized names through the NPC registry while preserving grounded collective labels. |
| Location registry | Produces a source-anchored, session-scoped registry from stable proper names or unique in-world designations. Its LLM-assisted reconciliation is proposal-only and never collapses same-name places without validated identity and evidence rules. |
| Location occurrences | Requires the normalized location registry for both extraction and normalization. Its [durable occurrence categories](../integrations/dnd-location-occurrence-artifacts.md#occurrence-categories) distinguish explicit speculation from unsupported inference; the deterministic normalizer enforces exact registry grounding and never turns registry provenance into occurrence evidence. |
The combat and scene-description contracts describe their exact handoff and The combat and scene-description contracts describe their exact handoff and
empty-result behavior in more detail: empty-result behavior in more detail:
[combat turns](../integrations/dnd-combat-turn-artifacts.md) and [combat turns](../integrations/dnd-combat-turn-artifacts.md) and
[scene descriptions](../integrations/dnd-scene-description-artifacts.md). [scene descriptions](../integrations/dnd-scene-description-artifacts.md).
The [enemy-event contract](../integrations/dnd-enemy-event-artifacts.md)
defines its durable semantics; [Configuration](../config.md) owns its
selectable bindings and validation chains.
## Focused Verification ## Focused Verification

View File

@@ -1,12 +1,12 @@
# LLM Runtime Internals # LLM Runtime Internals
`internal/framework/llm` is Notariuss provider-independent structured `internal/framework/llm` is Notariuss provider-independent structured
completion boundary. It adapts framework requests to PromptKit, bounds completion boundary. It adapts framework requests to Scriptorium, bounds
provider calls, assembles registered prompt and schema assets, records selected provider calls, assembles registered prompt and schema assets, records selected
profiles, and redacts provider errors. The architectural boundary is defined in profiles, and redacts provider errors. The architectural boundary is defined in
[Architecture](../policy/architecture.md#llm-boundary); profile sources, [Architecture](../policy/architecture.md#llm-boundary); profile sources,
credentials, and concurrency settings belong in credentials, and concurrency settings belong in
[Configuration](../config.md#promptkit-profiles) and [Configuration](../config.md#scriptorium-profiles) and
[Configuration](../config.md#concurrency-output-cache-and-debug). [Configuration](../config.md#concurrency-output-cache-and-debug).
## Structured Completion Boundary ## Structured Completion Boundary
@@ -24,102 +24,25 @@ adapter does not own source evidence, artifact conversion, normalization, or
durable schemas. Those responsibilities remain with the module and its durable schemas. Those responsibilities remain with the module and its
[integration contract](../integrations/). [integration contract](../integrations/).
The calling module also resolves contextual entity selections and attaches any `ScriptoriumClient` validates the request target and prompt identity, maps each
application identity; PromptKit and this adapter do not own entity identity. named material to a Scriptorium inline artifact while preserving its origin URI,
forwards session and profile selection, then prepares and runs the prompt. It
returns Scriptoriums validated raw bytes rather than re-encoding the decoded
target. An empty optional material is represented as one space so its named
input is retained by Scriptorium.
`PromptKitClient` validates the request target and prompt identity, maps each An empty request profile lets the prompt select its configured default. The CLI
named material to a PromptKit inline artifact while preserving its origin URI, prepares every explicitly selected binding profile before a run begins, so a
passes the supplied request session through to PromptKit's direct per-run missing explicit profile fails before stage execution. Calls record the profile
session field, retains the same value as the `session_id` prompt variable for actually selected by Scriptorium; the recorder deduplicates non-secret profile
maintained prompt compatibility, and forwards profile selection. It does not identity, provider, and model values for manifest use.
derive or replace session values; the CLI owns that policy. It then creates one
frozen prepared execution, captures its caller-owned credential-redacted
details for debug material, and executes that exact snapshot through
PromptKit's prepared-execution boundary. The direct field
is authoritative for provider session behavior. A session ID is a stable,
non-secret correlation identifier and may be exposed to providers and provider
observability. The adapter returns PromptKits validated raw bytes rather than
re-encoding the decoded target. An empty optional material is represented as
one space so its named input is retained by PromptKit.
When a request includes semantic correction material, the adapter validates and
defensively copies it before preparation, then appends exactly two messages
after the ordinarily rendered prompt: the prior response as an assistant
message and the correction guidance as a user message. Requests without a
correction do not add messages or introduce caller roles. Ordinary request
summaries record correction byte counts and digests only; complete messages are
available solely in an explicitly requested debug trace.
The adapter leaves the ordinary rendered message prefix, named inputs,
variables, session, profile, execution overrides, prepared-execution path, and
PromptKit repair policy unchanged for a corrected request. It never imports a
PromptKit message type into a module or pipeline contract. PromptKit reports
actual structural repair count and cumulative token usage per completion; the
pipeline's safe terminal debug record projects those values without copying
message content.
Client construction may also receive a run-wide reasoning-effort override from
the CLI factory boundary. The adapter copies the caller-owned pointer and
creates a fresh PromptKit execution override for each request: a nil pointer
inherits the selected profile, a non-empty value replaces it, and an empty
value clears inherited reasoning. The CLI's mutually exclusive
`--reasoning-effort` and `--clear-reasoning-effort` controls select those
states. With neither flag, profile behavior remains unchanged. Because
production constructs one shared client, the selected state applies uniformly
to module calls, retries, and LLM-backed validators for the whole run.
An empty request profile lets the prompt select its configured default. Before a
run begins, the CLI asks the adapter to inspect every explicit profile on the
resolved selected LLM-backed bindings and validators, including inherited
pipeline profiles. Inspection resolves the profile and its selected backend and
target without loading a prompt, reading credentials, admitting capacity, or
contacting a provider, so a missing or invalid explicit profile fails before
stage execution while a valid `api_key_env` may remain unset. Calls record the
profile actually selected by PromptKit. The recorder trims and deduplicates
non-secret profile identity, provider, model, selected backend ID, and
effective reasoning values for manifest use. Entries that differ in backend or
reasoning remain distinct and deterministically ordered. Endpoint-only profiles
retain an empty backend ID, which the published JSON omits. Successful
completion responses and recorded profile manifests identify the adapter
provider as `promptkit`.
The CLI's profile-inspection engine and the production adapter use the same
profile-source construction to apply the configured profile directory or file,
the optional registered fallback profile assets, and the optional conventional
`local` backend. Preflight therefore resolves the same profile sources and
backend membership as runtime without performing generation. Fallback assets
are mounted only when at least one source is registered. The production D&D
registrar contributes its `dnd-extraction` fallback, and the maintained D&D
prompts select that logical ID by default. PromptKit owns source precedence and
profile parsing and inheritance: an operator-provided matching profile takes precedence over a
fallback profile without Notarius merging either document.
When the registration is absent, a profile selecting `backend: local` fails
inspection instead of falling back to a built-in or endpoint-only target.
Before execution, the adapter also contributes a non-secret checkpoint
fingerprint for the effective PromptKit profile source. It combines the
identity of PromptKit's compiled-in profile catalog with a deterministic digest
of every YAML profile in the configured profile directory, or of the configured
profile file, and a deterministic digest of the flattened fallback profile
assets. The fingerprint contains neither profile content nor source paths. It
covers inherited pipeline profiles, explicit binding profiles, and
prompt-selected defaults, so changing a model or other profile setting cannot
reuse checkpoints created under the
prior profile source. This cache identity is independent of durable
profile provenance: run manifests continue to list only profiles actually
observed during LLM calls. When the local backend is registered, a second
fingerprint hashes its trimmed endpoint behind a stable marker. Changing that
semantic execution target invalidates checkpoint reuse. The raw endpoint is not
stored in checkpoint identity, and the local concurrency limit is excluded
because it changes scheduling rather than execution semantics.
## Shared Provider-Call Limit ## Shared Provider-Call Limit
Production construction creates one PromptKit client and wraps it in one Production construction creates one Scriptorium client and wraps it in one
scheduled client. The scheduler has a fixed, positive permit limit, serves scheduled client. The scheduler has a fixed, positive permit limit, serves
queued calls in FIFO order, and removes a queued call when its context is queued calls in FIFO order, and removes a queued call when its context is
cancelled. It rechecks the caller context after admission and before dispatch. cancelled. A granted permit is released exactly once on every completion path.
A granted permit is released exactly once on every completion path.
The scheduled wrapper surrounds every `CompleteStructured` call, so concurrent The scheduled wrapper surrounds every `CompleteStructured` call, so concurrent
lanes, pipeline retries, and LLM-backed validators share the same provider-call lanes, pipeline retries, and LLM-backed validators share the same provider-call
@@ -128,57 +51,22 @@ worker counts cannot exceed the configured LLM limit. The configuration field
and its effective default are owned by and its effective default are owned by
[Configuration](../config.md#concurrency-output-cache-and-debug). [Configuration](../config.md#concurrency-output-cache-and-debug).
PromptKit applies a second, independent admission limit when the selected
profile names a limited backend. It sits beneath the Notarius scheduled client,
so it may narrow but cannot expand the application-wide limit. Built-in
OpenRouter profiles select PromptKit's reserved backend and its upstream
capacity policy. A positive configured local-backend limit bounds active local
generations inside PromptKit; zero leaves that backend unlimited there.
Endpoint-only profiles do not select a PromptKit backend and remain limited
only by the Notarius scheduler.
## Prompt And Schema Assets ## Prompt And Schema Assets
An `AssetRegistry` collects prompt, schema, and optional fallback-profile An `AssetRegistry` collects prompt and schema filesystems from production module
filesystems from production module families. It flattens registered roots into families. It flattens registered roots into the Scriptorium filesystems and
the corresponding PromptKit filesystems and rejects invalid roots, unreadable rejects invalid roots, unreadable assets, duplicate paths, and missing prompt
assets, duplicate paths, and missing prompt or schema files during preparation. or schema files during preparation. The frameworks `promptfs` helper combines
Fallback assets receive a safe content digest for checkpoint identity; raw module-owned prompt files with reusable domain fragments without making the
paths and bytes are never included. The frameworks `promptfs` helper combines
module-selected prompt files with reusable domain fragments without making the
framework depend on D&D content. framework depend on D&D content.
LLM-facing content is embedded once by the root `assets` package. Each consumer Each LLM-backed module owns its prompt declaration, package-specific assets,
uses only its scoped subtree, while the module retains ownership of its prompt and private response schema. Shared D&D wording is owned by the D&D shared
declaration, ordered manifest, private response-schema identity, and asset package; the detailed D&D conventions are in
registration. Shared D&D fragments are selected by D&D's shared implementation; [D&D Module Internals](dnd.md). The mounted prompt assets used by a module also
the detailed convention is in [D&D Module Internals](dnd.md). This physical determine its prompt fingerprint. Schema loaders validate JSON, attach identity
arrangement and its data-only boundary are defined by and digest metadata, make defensive copies, and expose diagnostics without raw
[Architecture](../policy/architecture.md) and schema bytes.
[ADR-0011](../adr/0011-centralize-llm-assets.md), rather than by this runtime
guide.
The generic registrar is the sole production registration owner for the
semantic-reconciliation default prompt and private response schema. The
domain-neutral reconciliation package also exposes only its mandatory protocol
and candidate/transcript presentation files for domain prompt manifests. D&D
registry normalizers mount those files while retaining ownership and hashing
of their D&D system message, semantic instructions, and complete prompt
declaration. The response schema is therefore registered once even though
several typed normalizers select it.
Mounted prompt assets determine a module's fingerprint. The fingerprint hashes
only the module and shared files explicitly selected by its manifest, so an
unrelated asset does not invalidate a checkpoint. Schema loaders validate JSON,
attach identity and digest metadata, make defensive copies, and expose
diagnostics without raw schema bytes.
Semantic-reconciliation normalizers extend this identity with the shared
response-schema digest, framework policy version, and complete limit-policy
digest. Their manifest metadata records the same content-free prompt, schema,
policy, and limit identities together with domain identity and normalization
policies. Request-local handles, source material, proposal content, and raw
asset bytes are not checkpoint metadata.
Private response schemas validate a model transport envelope. They are not the Private response schemas validate a model transport envelope. They are not the
durable artifact schema and should not be documented as an external wire durable artifact schema and should not be documented as an external wire
@@ -188,115 +76,54 @@ contract. Durable formats and compatibility rules remain in the
## Prompt Maintenance And Backend Caching ## Prompt Maintenance And Backend Caching
Prompt message order and shared asset bytes are runtime behavior. Backend cache Prompt message order and shared asset bytes are runtime behavior. Backend cache
reuse depends on identical preceding roles, rendered bytes, and cache-control reuse depends on the same preceding messages and content, not merely equivalent
metadata—not merely equivalent meaning. Keep reusable shared assets meaning. Keep reusable shared assets byte-identical and keep stable material
byte-identical and preserve each prompts declared ordering and cache controls before the inputs that vary per request wherever a prompts declared sequence
when editing it. supports caching. Preserve the existing manifest order and cache-control hints
when editing a prompt.
For sibling prompts that can reuse the same source material, order universal D&D extraction manifests place the changing chunk transcript at the end of the
shared context first, request source material next, and module-specific prompt after their reusable context. Scene chunking and NPC normalization use
suffixes last. Put a cache boundary at a reusable prefix that is useful to the their own declared message sequences because their inputs and work differ. The
backend. Redundant intermediate cache boundaries do not extend that reusable family-specific asset and ordering rules belong in [D&D Module Internals](dnd.md).
prefix and add no value. Do not add tests that enforce a fixed message-prefix length; prompt-asset tests
should instead verify the meaningful asset sequence, inputs, and cache controls
Prompt-family owners may choose a different sequence when their inputs and of the prompt being changed.
reuse pattern differ. The D&D familys extraction, scene-chunking, and NPC
normalization policies are maintained in [D&D Module Internals](dnd.md#prompt-construction).
Do not add tests that enforce prompt prose; prompt tests should verify the
meaningful input placement and cache controls of the prompt being changed.
## Validation, Repair, And Retries ## Validation, Repair, And Retries
PromptKit performs prompt rendering, provider execution, and the prompts Scriptorium performs prompt rendering, provider execution, and the prompts
structured-output validation. The adapter reports an empty result, validation structured-output validation. The adapter reports an empty result, validation
failure, empty structured body, or decode failure as failure, empty structured body, or decode failure as
`ErrInvalidStructuredOutput`, while retaining the returned raw bytes and debug `ErrInvalidStructuredOutput`, while retaining the returned raw bytes and debug
material when they exist. Provider failures remain operational errors rather material when they exist. Provider failures remain operational errors rather
than output-validation failures. Apart from documented context, capacity, and than output-validation failures.
invalid-output categories, provider error values and types do not cross the
adapter error chain; callers receive only a credential-redacted diagnostic.
When PromptKit rejects backend admission before generation, the adapter maps Prompt-declared repair is executed within Scriptoriums structured-output flow.
`promptkit.ErrCapacityExceeded` to The current production D&D prompt manifests set repair attempts to zero. That
`contracts.ErrLLMCapacityExceeded`, retaining prompt context and a redacted setting does not replace pipeline retry behavior: a bindings configured retry
upstream diagnostic without exposing the PromptKit sentinel or capacity-error count reruns its stage attempt after an error or rejection, and an exhausted
type as a framework contract. When supplied, the normalized selected backend rejection is a recorded output rather than a provider error. The pipeline owns
ID appears only in that safe application-owned diagnostic context. A canceled attempt lifecycle, validation chains, and retry diagnostics; see
caller context takes precedence. The adapter does not retry capacity failures; [Pipeline Internals](pipeline.md#validation-retries-and-output) and the
the pipeline's existing binding attempt policy sees the operational error and
decides whether to rerun the complete operation.
PromptKit executes structural repair within its structured-output flow. The
maintained production prompt manifests declare one additional repair attempt.
When a resolved binding supplies a repair value, the adapter inspects the
prompt, copies its complete output contract, changes only the repair limit, and
passes that complete replacement contract to PromptKit. This preserves the
prompt's output format, validation mode, schema, and provider structured-output
settings.
A successful repair is an ordinary successful completion, not a warning. The
adapter reports PromptKit's actual repair count and its cumulative usage
directly, without adding the initial and corrective counts again. Debug prompt
material records the configured complete contract; debug response material
records the repaired response and actual validation result. If the repair
budget is exhausted, the adapter retains the final raw bytes and debug material
and reports `ErrInvalidStructuredOutput`. Generation failures during an initial
or corrective call remain provider-neutral operational errors with the same
redaction boundary.
Structural repair does not replace pipeline retry behavior: a binding's
configured retry count reruns its complete stage attempt after an operational
or structural error, module-requested retry, or actionable semantic rejection.
The pipeline owns attempt lifecycle, validation chains, and retry diagnostics;
see [Pipeline Internals](pipeline.md#validation-retries-and-output) and the
[binding reference](../config.md#module-bindings-and-validators). [binding reference](../config.md#module-bindings-and-validators).
## Timeout Ownership
The caller context remains the outer cancellation authority. PromptKit applies
a positive effective generation timeout as an inner request deadline; an
explicit zero disables only that generation deadline. The HTTP client timeout
is a separate transport-wide cap. Notarius forwards the caller context and
does not install another timeout wrapper around PromptKit.
The selected PromptKit profile owns generation settings. Notarius binding
retries remain outside the adapter and repeat the complete module operation
and validation chain. PromptKit does not add a provider retry loop.
Operator-facing behavior is summarized in
[Operations](../operations.md#operational-limits), and the pinned upstream
contract is identified in
[PromptKit Integration](../integrations/pkg-promptkit.md).
## Observability And Redaction ## Observability And Redaction
When debug recording is enabled, the pipeline decorates the shared client. The When debug recording is enabled, the pipeline decorates the shared client. The
wrapper records prepared prompt and response material, timing, selected profile wrapper records prepared prompt and response material, timing, selected profile
and backend, effective model parameters, and call identifiers in the runs and model, and call identifiers in the runs debug bundle, including material
debug bundle, including material available from a failed structured completion. available from a failed structured completion. For a successful completion, a
Effective parameters use PromptKit's stable lower-case JSON field names and may debug-write failure is surfaced; when the completion already failed, its call
include `backend_id`. For a successful completion, a debug-write failure is error remains the result. Debug-bundle location, retention, and handling are
surfaced; when the completion already failed, its call error remains the operational concerns documented in [Operations](../operations.md#debug-bundles).
result. Debug-bundle location, retention, and handling are operational concerns
documented in [Operations](../operations.md#debug-bundles).
The attempt-terminal summary is a separate safe trace record: it contains Run manifests receive selected profile summaries and component identities, not
attempt kinds, validator outcome counts and reason codes, effective policy, prompt, schema, source, reference, or response content. Provider error text is
terminal action, and repair/usage references. It excludes raw assistant wrapped with prompt context and bearer credentials are redacted before it
responses and correction text. Those values can appear only in the explicitly crosses the runtime boundary. Known-secret redaction is available to other
requested detailed prompt and response artifacts, which require sensitive-data runtime collaborators; it does not make prompt or response contents safe for
handling. general logging.
Run manifests receive selected profile summaries, including optional effective
backend and reasoning provenance, and component identities—not prompt, schema,
source, reference, or response content. The published field semantics belong
to the [JSON output contract](../integrations/json-output.md#manifestjson).
Provider error text is wrapped with prompt context and bearer credentials are
redacted before it crosses the runtime boundary. Known-secret redaction is
available to other runtime collaborators; it does not make prompt or response
contents safe for general logging.
Generation failures expose an application-owned category and optional HTTP
status. Provider code, type, and message remain debug-only, after redaction.
## Failure Boundaries ## Failure Boundaries
@@ -304,8 +131,6 @@ status. Provider code, type, and message remain debug-only, after redaction.
sources, invalid asset registration, or a non-positive scheduler limit. sources, invalid asset registration, or a non-positive scheduler limit.
- Preparation failures, unavailable explicit profiles, provider failures, and - Preparation failures, unavailable explicit profiles, provider failures, and
context cancellation propagate to the calling stage with context. context cancellation propagate to the calling stage with context.
- Backend admission exhaustion is a provider-neutral operational error and is
not classified as invalid structured output or validator rejection.
- Malformed or schema-invalid provider output is classified separately as - Malformed or schema-invalid provider output is classified separately as
invalid structured output so the module or pipeline can apply its own retry invalid structured output so the module or pipeline can apply its own retry
and rejection policy. and rejection policy.

View File

@@ -12,24 +12,9 @@ exceptions. See [D&D Module Internals](dnd.md) rather than adding them here.
A module is a typed implementation registered for one pipeline stage. Its A module is a typed implementation registered for one pipeline stage. Its
`ModuleSpec` is the public-to-the-framework declaration of its stable key, `ModuleSpec` is the public-to-the-framework declaration of its stable key,
stage, execution class, required and provided capabilities, artifact kind, and stage, required and provided capabilities, artifact kind, and accepted
accepted reference slots. The execution class states whether a module is reference slots. The framework uses that declaration to resolve a configured
`deterministic` or `llm_backed`; registries retain it for catalog inspection and binding before it builds the implementation.
resolved-pipeline debug data without constructing the module. The framework
uses the declaration to resolve a configured binding before it builds the
implementation. After selection, the resolver applies profile inheritance only
to bindings whose declared execution class is `llm_backed` and rejects a
binding-specific profile on a deterministic module. The user-facing precedence
contract belongs in [Configuration](../config.md#pipelines).
An eligible LLM-backed chunk, extract, merge, or normalize producer may also
declare correction protocol `single_response_v1`. That declaration is a
promise that the implementation accepts one attempt-local semantic correction
and returns an owned copy of the exact one model response that directly
controlled the candidate. It must forward correction only to its structured
completion request; it must not manufacture prior-response material by
serializing a normalized artifact or expose opaque application IDs. Input,
output, validator, and deterministic specs cannot declare the protocol.
Implementations that accept options must provide both an option validator and Implementations that accept options must provide both an option validator and
a builder. The validator is used while resolving configuration; the builder a builder. The validator is used while resolving configuration; the builder
@@ -45,103 +30,45 @@ they need, register each leaf implementation, and add any family-owned assets
or default validator chains. They return contextual errors so production or default validator chains. They return contextual errors so production
composition fails at startup rather than at the first run. composition fails at startup rather than at the first run.
A validator that returns a completed rejection must supply two separate
bounded values: a stable `ReasonCode` for provenance and actionable
`CorrectionGuidance` for the producer. Guidance identifies the semantic defect
and the constraints on one complete corrected replacement. It must not contain
validator keys, diagnostic paths, opaque application IDs, or other internal
identifiers. An operator-facing `Message` may explain the same event, but the
framework never copies it into a model request. Missing or invalid guidance is
a validator contract failure.
An artifact family can register an optional typed evidence projector alongside An artifact family can register an optional typed evidence projector alongside
its codec. The projector returns defensive copies of the artifact's direct its codec. The projector returns defensive copies of the artifact's direct
generic source references and must use the codec's exact Go type. It does not generic source references and must use the codec's exact Go type. It does not
interpret surrounding context or publish files; the pipeline validates the interpret surrounding context or publish files; the pipeline validates the
capability during preparation and the output boundary owns publication. See capability during preparation and the output boundary owns publication. See
the [Published Evidence Context contract](../integrations/evidence-context.md) the [Published Evidence Context contract](../integrations/evidence-context.md)
for the durable source-unit excerpt. Lane artifacts retain citation and lane for the durable result.
provenance; the framework does not add either to that published excerpt.
An artifact family is broader than a module: it owns the cohesive domain
feature across its artifact type, codec, stage modules, validators, prompt
policy, schemas, identity helpers, and reference projections. An extractor and
normalizer in one artifact family remain independently registered modules in
their respective pipeline stages. This ownership vocabulary does not create a
new registry or change the fixed pipeline.
## Production Composition ## Production Composition
Production composition is intentionally split by family: Production composition is intentionally split by family:
- The generic registrar provides the unit chunker, generic JSON validators, - The generic registrar provides the unit chunker, generic JSON validators,
JSON output encoder, and shared semantic-reconciliation prompt and response and JSON output encoder.
schema assets.
- The Seriatim registrar provides the transcript input adapter. Its external - The Seriatim registrar provides the transcript input adapter. Its external
input behavior is defined by the [Seriatim contract](../integrations/seriatim.md). input behavior is defined by the [Seriatim contract](../integrations/seriatim.md).
- The D&D registrar provides its codecs, extractors, mergers, normalizers, - The D&D registrar provides its codecs, extractors, mergers, normalizers,
validators, prompt assets, fallback profile asset, and default chains. Its behavioral conventions validators, prompt assets, and default chains. Its behavioral conventions
are documented in [D&D Module Internals](dnd.md). are documented in [D&D Module Internals](dnd.md).
The CLI owns the composition that invokes these registrars. A module package The CLI owns the composition that invokes these registrars. A module package
may register its own family but must not assemble the CLI or make framework may register its own family but must not assemble the CLI or make framework
packages depend on production extensions. packages depend on production extensions.
## Semantic Reconciliation
`internal/framework/semanticreconcile` is a domain-neutral strategy used by a
typed normalize module; it is not itself a selectable stage module. A
source-backed artifact-family normalizer projects its deterministic records
into contextual candidates and owned typed record envelopes, supplies its
chosen prompt identity and resolved LLM profile, and constructs an engine with
explicit limits. The core filters invalid evidence, assigns contiguous
request-local integer handles, renders bounded candidate and transcript
materials, invokes the structured-completion boundary, and assesses the
returned duplicate groups into a stable non-overlapping plan.
The normalizer then applies that plan through a typed `ApplicationPolicy`. The
core preserves ungrouped records, contribution order, and provenance while the
artifact family owns group guards, field and evidence consolidation, durable
ID derivation, retry and fallback presentation, classified diagnostics, and postconditions.
Request-local handles do not enter the typed value or durable artifact. Fewer
than two eligible candidates skips model invocation; exceeding a candidate or
combined-material bound preserves the deterministic result under the family's
fallback policy. Provider, transport, cancellation, and context-construction
failures remain execution errors.
When the engine actually makes a proposal call, its typed result carries the
owned exact proposal response under the same correction contract as other
eligible producers. Deterministic skip, limit, and fallback outcomes carry no
model candidate, so a later rejection applies terminal policy without spending
an ineffective semantic retry.
The core supplies a conservative generic prompt and the single private
response schema. A domain prompt may substitute its semantic instructions but
mounts the core-owned protocol and candidate/transcript presentation assets.
Prompt, schema, policy, and limit identities participate in manifest metadata
and checkpoint fingerprints. The generic registrar owns production
registration of those shared assets; a consuming domain registrar owns only
its domain prompt.
## Adding Or Changing A Module ## Adding Or Changing A Module
1. Choose the pipeline stage and the typed artifact boundary. Put external 1. Choose the pipeline stage and the typed artifact boundary. Put external
input or durable artifact formats in the relevant integration contract, input or durable artifact formats in the relevant integration contract,
not in this guide or in a private LLM response type. not in this guide or in a private LLM response type.
2. Define a stable `ModuleSpec` with an explicit execution class, the exact 2. Define a stable `ModuleSpec` with the exact capabilities and reference
capabilities, and reference slots needed for the operation. Model a slots needed for the operation. Model a producer/consumer handoff as an
producer/consumer handoff as an artifact-compatible slot; configuration artifact-compatible slot; configuration then chooses an external file or a
then chooses an external file or a generated binding. generated binding.
3. Implement strict option decoding, construction, and the typed stage 3. Implement strict option decoding, construction, and the typed stage
interface. Preserve caller ownership: do not retain mutable request data interface. Preserve caller ownership: do not retain mutable request data
and return defensive copies where an implementation exposes stored data. and return defensive copies where an implementation exposes stored data.
If declaring correction capability, forward the request correction and
retain only the exact validated response that controlled the result.
4. Register the module through its typed registry helper and add it to the 4. Register the module through its typed registry helper and add it to the
owning family registrar. Add a default validator chain only when that owning family registrar. Add a default validator chain only when that
family owns the behavior; otherwise require an explicit compatible chain. family owns the behavior; otherwise require an explicit compatible chain.
Every rejection path in a validator must provide actionable correction
guidance while retaining its stable internal reason code.
5. Update the selectable-key and chain reference in 5. Update the selectable-key and chain reference in
[Configuration](../config.md#production-module-keys), the applicable [Configuration](../config.md#production-module-keys), the applicable
integration contract, and focused tests. Keep the configuration document integration contract, and focused tests. Keep the configuration document

View File

@@ -25,13 +25,10 @@ physical state roots.
| Area | Implemented owners | Responsibility | | Area | Implemented owners | Responsibility |
| --- | --- | --- | | --- | --- | --- |
| Executable and command boundary | **cmd/notarius**, **internal/cli** | Process entry, command dispatch, configuration discovery, production composition, runtime collaborator setup, durable file placement, and user-facing reporting. | | Executable and command boundary | **cmd/notarius**, **internal/cli** | Process entry, command dispatch, configuration discovery, production composition, runtime collaborator setup, durable file placement, and user-facing reporting. |
| Build information | **internal/buildinfo** | Resolves a stable linked release tag or build metadata for the diagnostic root version command. |
| Configuration | **internal/core/config** | Defaults, strict YAML parsing, environment overrides, structural validation, effective resolution, redaction, and resolved-composition summaries. | | Configuration | **internal/core/config** | Defaults, strict YAML parsing, environment overrides, structural validation, effective resolution, redaction, and resolved-composition summaries. |
| Generic models | **internal/core/source**, **internal/core/artifacts**, **internal/framework/contracts** | Source documents and chunks, manifests and provenance, plus typed artifact, reference, validation, output, and structured-completion contracts. | | Generic models | **internal/core/source**, **internal/core/artifacts**, **internal/framework/contracts** | Source documents and chunks, manifests and provenance, plus typed artifact, reference, validation, output, and structured-completion contracts. |
| Pipeline framework | **internal/framework/pipeline** | Registries, profile and reference resolution, typed preparation, validation, retry coordination, ordered execution, handoff, and result assembly. | | Pipeline framework | **internal/framework/pipeline** | Registries, profile and reference resolution, typed preparation, validation, retry coordination, ordered execution, handoff, and result assembly. |
| LLM and prompt runtime | **internal/framework/llm**, **internal/framework/promptfs** | Provider-neutral structured completions, scheduling, profile recording, prompt assets, schema registration, and credential-shaped-value redaction. | | LLM and prompt runtime | **internal/framework/llm**, **internal/framework/promptfs** | Provider-neutral structured completions, scheduling, profile recording, prompt assets, schema registration, and credential-shaped-value redaction. |
| Semantic reconciliation | **internal/framework/semanticreconcile** | Bounded source-backed candidate preparation, request-local handle proposals, deterministic assessment, typed plan application, and reconciliation identity metadata; see [Module Internals](modules.md#semantic-reconciliation) and [D&D Module Internals](dnd.md#semantic-registry-reconciliation). |
| Embedded LLM content | **assets** | Read-only centralized LLM-facing content, scoped by its consuming package; see [LLM Runtime](llm.md#prompt-and-schema-assets) and [D&D Module Internals](dnd.md#prompt-construction). |
| Runtime state | **internal/core/fileio**, **internal/core/debugbundle**, **internal/framework/checkpoint**, **internal/framework/chunkplan**, **internal/framework/chunkmap**, **internal/framework/debug** | Confined atomic files, debug bundles, checkpoint and chunk-plan state, accepted chunk maps, and pipeline-facing debug recording. | | Runtime state | **internal/core/fileio**, **internal/core/debugbundle**, **internal/framework/checkpoint**, **internal/framework/chunkplan**, **internal/framework/chunkmap**, **internal/framework/debug** | Confined atomic files, debug bundles, checkpoint and chunk-plan state, accepted chunk maps, and pipeline-facing debug recording. |
| Production extensions | **internal/modules/generic**, **internal/modules/seriatim**, **internal/modules/dnd** | Domain-neutral extensions, Seriatim input support, and D&D extraction families registered into the production catalog. | | Production extensions | **internal/modules/generic**, **internal/modules/seriatim**, **internal/modules/dnd** | Domain-neutral extensions, Seriatim input support, and D&D extraction families registered into the production catalog. |
@@ -51,9 +48,8 @@ the CLI composition boundary.
composition, and path safety. composition, and path safety.
- [LLM Runtime](llm.md): structured completion, scheduling, prompt assets, - [LLM Runtime](llm.md): structured completion, scheduling, prompt assets,
profiles, and secret handling. profiles, and secret handling.
- [Module Internals](modules.md): generic extension registration, artifact - [Module Internals](modules.md): generic extension registration, module
families, module construction, semantic reconciliation, validation, and construction, validation, and reference mechanics.
reference mechanics.
- [D&D Module Internals](dnd.md): shared D&D extractor conventions, generated - [D&D Module Internals](dnd.md): shared D&D extractor conventions, generated
reference projections, and lane-specific exceptions. Durable D&D and reference projections, and lane-specific exceptions. Durable D&D and
Seriatim data shapes remain in the [integration contracts](../integrations/). Seriatim data shapes remain in the [integration contracts](../integrations/).

View File

@@ -10,11 +10,11 @@ own durable output shapes. Concrete production extensions are covered by
## Boundary ## Boundary
The pipeline framework accepts a resolved composition, registries, shared The pipeline framework accepts a resolved composition, registries, shared
dependencies, input bytes, a supplied prompt session, and state/debug dependencies, input bytes, and state/debug collaborators. It returns logical
collaborators. It returns logical output files, normalized artifacts, recorded output files, normalized artifacts, recorded rejections and warnings, manifest
rejections, grouped diagnostics, manifest provenance, and checkpoint decisions. The provenance, and checkpoint decisions. The CLI owns process arguments,
CLI owns process arguments, configuration discovery, session resolution, configuration discovery, physical roots, and placement of returned output
physical roots, and placement of returned output files. files.
The framework has one fixed shape: The framework has one fixed shape:
@@ -32,26 +32,8 @@ Resolution turns a configured pipeline profile into a **ResolvedPipeline**.
It normalizes the pipeline and lane identities, applies stage defaults, selects It normalizes the pipeline and lane identities, applies stage defaults, selects
requested lanes where that is supported, resolves validator chains, checks requested lanes where that is supported, resolves validator chains, checks
module capabilities and typed artifact compatibility, validates options, and module capabilities and typed artifact compatibility, validates options, and
assigns a deterministic resolved-composition digest. A correction protocol is assigns a deterministic resolved-composition digest. The resolved pipeline
selected from each eligible LLM-backed producer specification and becomes part contains bindings and declared reference targets, not external reference bytes.
of that resolved identity; only `single_response_v1` is currently supported.
Preparation rejects an LLM-backed producer that combines a non-empty validator
chain with positive producer retries unless it declares that protocol. Producers
without validators or without retries remain valid without correction support.
The resolved pipeline contains bindings and declared reference targets, not
external reference bytes.
After selection, the resolver applies command, binding, and pipeline profile
precedence to LLM-backed bindings and validators only; prompt defaults remain
an empty resolved binding profile. It resolves structural output repair
separately: a binding's `structured_output_repair_attempts` value wins, then a
pipeline value applies to LLM-backed bindings and validators, and omission
leaves the prompt-owned policy intact. An explicit repair value on a
deterministic binding is rejected. Resolved bindings own copied repair values,
and these effective values are part of the digest, so execution and checkpoint
consumers do not repeat profile inheritance or configuration resolution.
Each LLM request receives its own copy of that resolved value. PromptKit spends
it only for structural correction inside one completion; the runner's binding
retry policy remains the separate outer budget for complete stage attempts.
Configuration resolution supplies the selected profile and catalog; see Configuration resolution supplies the selected profile and catalog; see
[Configuration Internals](configuration.md). [Configuration Internals](configuration.md).
@@ -59,20 +41,16 @@ External reference materialization happens before preparation. The materializer
checks that each slot is declared by the selected module, resolves a file path checks that each slot is declared by the selected module, resolves a file path
relative to the correct configuration or working-directory origin, reads relative to the correct configuration or working-directory origin, reads
UTF-8 text, verifies media type and size limits, and retains bounded UTF-8 text, verifies media type and size limits, and retains bounded
provenance. For a positive slot limit, it reads at most the limit plus one byte provenance. A generated-artifact selector remains declared but has no bytes
and rejects overflow before retaining content. A generated-artifact selector until its producing step completes.
remains declared but has no bytes until its producing step completes.
Preparation is the construction boundary. It validates the resolved shape and Preparation is the construction boundary. It validates the resolved shape and
registry set, clones the resolved data, then constructs the input adapter, registry set, clones the resolved data, then constructs the input adapter,
chunker, stage-local validators, every typed lane, and output encoder. The chunker, stage-local validators, every typed lane, and output encoder with
prepared producer metadata preserves each selected correction protocol, and cloned options, references, and shared dependencies. It also collects stable
the resolved digest carrying that metadata participates in checkpoint identity. checkpoint fingerprints. Missing registrations, incompatible typed entries,
Each registered builder receives its own cloned build request immediately nil implementations, and constructor failures are reported before source
before its module-owned code runs. Preparation also collects stable checkpoint parsing or any stage operation begins.
fingerprints. Missing registrations, incompatible typed entries, nil
implementations, and constructor failures are reported before source parsing
or any stage operation begins.
An output encoder can opt into source-evidence publication through its output An output encoder can opt into source-evidence publication through its output
policy. Preparation keeps the configured lane allowlist and active lanes policy. Preparation keeps the configured lane allowlist and active lanes
@@ -101,10 +79,6 @@ incompatible producer prevents the consumer step from starting.
The runner validates its input, installs no-op state collaborators when none The runner validates its input, installs no-op state collaborators when none
were supplied, and serially performs source parsing and chunk-plan selection. were supplied, and serially performs source parsing and chunk-plan selection.
It transports the supplied session unchanged to prompt-facing operations and
run-manifest metadata; it neither derives a session nor substitutes a parsed
source document identifier. The public session contract is owned by the
[CLI reference](../cli.md#run).
An accepted plan is materialized into source-addressed chunks and passes the An accepted plan is materialized into source-addressed chunks and passes the
configured chunk validators before any lane runs. A chunk rejection is a configured chunk validators before any lane runs. A chunk rejection is a
recorded pipeline outcome: lanes do not start, but the output stage can encode recorded pipeline outcome: lanes do not start, but the output stage can encode
@@ -132,52 +106,18 @@ for started workers, and prevents output encoding.
Every chunk, extract, merge, and normalize candidate passes its resolved Every chunk, extract, merge, and normalize candidate passes its resolved
validator chain. Validators receive immutable canonical input appropriate to validator chain. Validators receive immutable canonical input appropriate to
their target: chunks, codec-decoded typed candidates, or serialized codec their target: chunks, typed values, or serialized codec bytes. They may
bytes. Each typed validator receives a newly decoded value from the one approve, approve with warnings, reject, or fail. A rejection is an ordinary
candidate serialization for that attempt, while serialized validators receive pipeline result; a validator error is a framework error.
separately owned representation bytes and schema metadata. They may approve,
approve with warnings, reject, fail, or be skipped when a runtime prerequisite
is unavailable. The shared executor settles every configured validator in
order. A failed LLM-backed validator retries only itself against the same
immutable candidate; it does not regenerate the producer or alter the
validator request. Rejections stop that validator, while other configured
validators still run. The executor retains ordered results, bounded
deduplicated correction guidance from rejections, and only the final exhausted
failure outcome for each validator. The correction builder keeps first
occurrence order, omits internal reason codes, validator names, and operator
messages, and requests one complete replacement. Missing guidance or an
oversized aggregate is a framework contract error; guidance is never inferred
or truncated.
The runner applies the binding's retry policy around a stage operation and its The runner applies the binding's retry policy around a stage operation and its
complete validation chain. It preserves terminal diagnostics only from the final accepted complete validation chain. It preserves warnings only from the final accepted
or rejected attempt, plus one fixed validation-incomplete warning per validator whose execution or rejected attempt. Cancellation stops retries. Normalizer-specific retry
budget was exhausted under `warn_continue`. Cancellation stops retries. directives consume this same budget and validate any final safe fallback through
Normalizer-specific retry directives consume this same budget and validate any the normalizer chain.
final safe fallback through the normalizer chain.
The artifact-neutral producer-attempt state machine owns that shared budget,
attempt provenance, semantic-correction material, and terminal-policy
selection. It accepts producer and complete-validation closures, so artifact
materialization, cache handling, checkpoints, and debug output stay at the
operation boundary. It distinguishes operational, structural, module-requested,
and semantic retries. A semantic retry is available only for a valid latest
`single_response_v1` candidate; a deterministic or no-model rejection instead
settles the semantic policy immediately. Structural-output errors alone use the
structural policy, and validation failure without rejection settles the
validator-failure policy without regenerating the producer.
Chunk planning uses this state machine for generated plans. A rejected or
validation-incomplete automatic cache hit is not model material and therefore
falls through to a fresh initial generation at producer attempt one; it neither
receives a correction, consumes retry budget, promotes cached-candidate
diagnostics, nor overwrites the stored record. An incomplete cache validation
under `fail_run` terminates instead. Only a newly generated, completely
validated plan is published to the chunk-plan store. Rejected plans never
advance, and validation-incomplete plans remain unpublishable.
After terminal lane work, the runner assembles manifest provenance, normalized After terminal lane work, the runner assembles manifest provenance, normalized
artifacts, rejections, final grouped diagnostics, and an optional accepted chunk map. When an artifacts, rejections, warnings, and an optional accepted chunk map. When an
output policy selected evidence lanes, it decodes accepted serialized normalize output policy selected evidence lanes, it decodes accepted serialized normalize
outputs through their registered codecs and invokes the prepared typed outputs through their registered codecs and invokes the prepared typed
projectors. Rejected or absent lanes contribute nothing. This reconstruction is projectors. Rejected or absent lanes contribute nothing. This reconstruction is
@@ -188,15 +128,6 @@ The CLI publishes those files only after the runner returns without a framework
error. Logical file names and schemas are defined by the [output integration error. Logical file names and schemas are defined by the [output integration
contracts](../integrations/). contracts](../integrations/).
For every completed producer disposition, the runner projects one bounded
validation summary to the manifest, the affected rejection when present, and
the CLI result receipt. The summary records status, configured-order rejecting
validators and reason codes, incomplete validators, producer-attempt count,
and terminal action. It contains no operator message, correction guidance, or
model response. `complete`, `rejected`, and `incomplete` describe the final
candidate disposition; a run-level `incomplete` status indicates at least one
current-run output advanced under `warn_continue`.
## Checkpoint And Debug Hooks ## Checkpoint And Debug Hooks
The runner receives checkpoint and debug interfaces rather than roots. It The runner receives checkpoint and debug interfaces rather than roots. It
@@ -206,20 +137,6 @@ handoff. Generated-reference dependencies participate in checkpoint decisions.
Selective recomputation can require a canonical accepted normalized predecessor Selective recomputation can require a canonical accepted normalized predecessor
before a dependent lane starts. before a dependent lane starts.
The runner writes successful checkpoint artifacts only after complete accepted
validation. Chunk plans follow the same rule for publication. A rejection,
invalid structured response, or incomplete validation is never reusable state;
the current run may still hand off an otherwise valid `warn_continue` result
according to its terminal policy. The runner carries private reuse eligibility
through extract, merge, normalize, and generated-reference handoff. Any stage
derived from incomplete validation skips both checkpoint lookup and all
checkpoint publication even when that stage's own validation completes.
External references and fully validated generated references remain eligible.
Attempt debug records retain safe kind,
validator, repair-usage, policy, and terminal-decision provenance. Full
assistant and correction content remains confined to the requested detailed
LLM trace.
Debug recording is attempt-scoped and application-owned. A failure to persist Debug recording is attempt-scoped and application-owned. A failure to persist
required debug data is a framework error. State roots, persistence, reason-code required debug data is a framework error. State roots, persistence, reason-code
meanings, resume, and cleanup are intentionally owned by meanings, resume, and cleanup are intentionally owned by

View File

@@ -39,17 +39,12 @@ codecs, loader, and recorder. The CLI constructs a recorder whenever checkpoint
recording is enabled and constructs a loader only for a `--resume` invocation. recording is enabled and constructs a loader only for a `--resume` invocation.
Identity incorporates explicit stable semantic fingerprints collected from Identity incorporates explicit stable semantic fingerprints collected from
prepared modules and validators in addition to configuration, input, prepared modules and validators in addition to configuration, input,
references, runtime overrides, observed LLM profiles, and the LLM runtime's references, runtime overrides, and LLM profiles.
non-secret effective profile-source identity. A profile source change therefore
causes a cold miss even when the configured profile ID remains unchanged.
The serialized The serialized
`workspace_schema_version` identifiers are frozen wire-compatibility fields; `workspace_schema_version` identifiers are frozen wire-compatibility fields;
they do not describe a current public state surface. they do not describe a current public state surface.
Ordered-step lane checkpoints include the step identity in their storage scope. Ordered-step lane checkpoints include the step identity in their storage scope.
Accepted step and lane identities are encoded injectively before becoming
filesystem path components, while ordinary safe identifiers retain their
readable paths.
When a later lane consumes a generated artifact, its dependency fingerprints When a later lane consumes a generated artifact, its dependency fingerprints
include the producer's artifact kind, complete schema identity, media type, include the producer's artifact kind, complete schema identity, media type,
canonical content digest, and size. Ordinary resume compares those fingerprints canonical content digest, and size. Ordinary resume compares those fingerprints
@@ -64,13 +59,12 @@ Ordinary resume loads extract, merge, and normalize checkpoints progressively
and may execute later lane stages after an earlier cache miss. Selective and may execute later lane stages after an earlier cache miss. Selective
recomputation instead asks the loader for the required producer's accepted recomputation instead asks the loader for the required producer's accepted
normalize artifact. That lookup reuses the existing normalize files, requires normalize artifact. That lookup reuses the existing normalize files, requires
workspace schema v4 plus an exact non-empty invocation identity, and deliberately workspace schema v3 plus an exact non-empty invocation identity, and deliberately
does not require extract or merge checkpoint files or dependency fingerprints. does not require extract or merge checkpoint files or dependency fingerprints.
The runner performs canonical codec and producer-provenance validation before The runner performs canonical codec and producer-provenance validation before
cloning the artifact into normal step output. Success restores only stored cloning the artifact into normal step output. Success restores only stored
normalize diagnostics and emits one normalize decision; failure retains the normalize warnings and emits one normalize decision; failure retains the files,
files, records the decision, and stops without executing the producer or records the decision, and stops without executing the producer or consumer.
consumer.
The loader assigns a typed category and reason code at each validation site; The loader assigns a typed category and reason code at each validation site;
diagnostic prose is not classified after the fact. The runner then applies diagnostic prose is not classified after the fact. The runner then applies
@@ -98,7 +92,7 @@ owns the operator workflow and stable reason-code meanings.
`internal/core/debugbundle` allocates an explicitly requested per-run bundle `internal/core/debugbundle` allocates an explicitly requested per-run bundle
with `summary/` and `trace/` roots. `SummaryWriter` persists redacted command, with `summary/` and `trace/` roots. `SummaryWriter` persists redacted command,
resolution, run, final grouped diagnostic, and failure artifacts. `internal/framework/debug` resolution, run, warning, and failure artifacts. `internal/framework/debug`
implements the pipeline-facing trace recorder under the trace root. implements the pipeline-facing trace recorder under the trace root.
The CLI allocates a bundle before pipeline resolution and treats requested The CLI allocates a bundle before pipeline resolution and treats requested

View File

@@ -5,23 +5,6 @@ This is the canonical guide for operating Notarius runtime state. The
[Configuration](config.md) owns fields, defaults, and precedence. Maintainers [Configuration](config.md) owns fields, defaults, and precedence. Maintainers
who need implementation mechanics should read [Run State Internals](internal/state.md). who need implementation mechanics should read [Run State Internals](internal/state.md).
## Source Deployment
Linux is the supported deployment platform. Install a pinned source release
with the Go version declared in `go.mod` (currently Go 1.25.5):
~~~sh
GOWORK=off go install \
gitea.maximumdirect.net/eric/notarius/cmd/notarius@vMAJOR.MINOR.PATCH
~~~
Pin the exact tag in deployment automation rather than following a branch.
Use [`notarius --version`](cli.md#command-summary) as a diagnostic after
installation; its syntax and semantics are owned by the [CLI reference](cli.md).
The maintainer publication process, including tag guards and verification,
belongs to [Source Releases](release.md). macOS builds are best-effort for
development, and Windows is unsupported.
## State Surfaces ## State Surfaces
Each run can use independent roots with different retention and access-control Each run can use independent roots with different retention and access-control
@@ -59,45 +42,6 @@ evidence publication. Apply an appropriate umask and output-root access policy
before enabling that option; the requested output modes alone may not be before enabling that option; the requested output modes alone may not be
suitable for transcript-bearing bundles. suitable for transcript-bearing bundles.
## PromptKit Profile Deployment
Profile deployment has four distinct layers:
| Layer | Owner | Operational role |
| --- | --- | --- |
| Prompts and schemas | Notarius module families | Embedded request and structured-output definitions. They are not deployment profile files. |
| Fallback profiles | Notarius module families | Embedded application defaults, including D&D's `dnd-extraction` profile. |
| Built-in profiles | PromptKit | Upstream catalog entries available when no higher-precedence source defines an ID. |
| Operator profiles | Deployment filesystem | Complete environment-specific definitions selected by `promptkit.profile_file` or `promptkit.profile_dir`. |
The maintained D&D pipeline uses the workload ID `dnd-extraction`. The
embedded fallback makes that ID usable without an operator file. Production,
development, and local deployments can each install a different complete
definition for the same ID, retaining the pipeline while choosing their own
model, backend, timeout, or reasoning policy. An operator definition wins over
the fallback; it is not merged with it. The configuration field and full
precedence rules are owned by [Configuration](config.md#promptkit-profiles).
Use a profile source owned by the service account, keep it readable only by
the intended operator, and supply provider credentials through the service
environment—not in the Notarius configuration or profile YAML. The maintained
[operator profile](../examples/profiles/dnd-extraction.yml) is secret-free and
can be copied as a format starting point. Validate a deployment without a
provider call or credentials:
~~~sh
notarius config validate --config /etc/notarius/config.yml --pipeline dnd-session
~~~
An unset optional `api_key_env` reaches the provider without authorization and
may receive a 401 or 403 response.
Profile paths are currently resolved from the process working directory, not
from the configuration file. The complete example's
`./examples/profiles/dnd-extraction.yml` path is valid for a repository-root
invocation only. Use absolute paths such as
`/etc/notarius/profiles/dnd-extraction.yml` for services and containers.
## Run Lifecycle ## Run Lifecycle
Use the [run command](cli.md#run) to start a pipeline. A valid invocation loads Use the [run command](cli.md#run) to start a pipeline. A valid invocation loads
@@ -105,40 +49,10 @@ and resolves configuration before module preparation and source parsing. It
then performs any permitted cache lookup, executes the pipeline, and publishes then performs any permitted cache lookup, executes the pipeline, and publishes
logical output files only after a successful runner result. logical output files only after a successful runner result.
On success, the command reports the output bundle path. A run with actionable On success, the command reports the output bundle path. A warning-bearing run
process warnings still succeeds and reports warning-group and occurrence counts still succeeds and reports its warning count on standard error. Errors and
on standard error; advisory and observation findings do not produce a warning
line. Errors and
their exit classes are defined in the [CLI reference](cli.md#output-streams-and-exit-statuses). their exit classes are defined in the [CLI reference](cli.md#output-streams-and-exit-statuses).
## Validation Retries And Terminal Outcomes
Each producer binding has one outer **retries** budget. It covers complete
producer attempts for operational failures, invalid structured output,
normalizer fallback retry, and semantic correction. It is independent from
PromptKit's structural-repair calls inside one completion and from an
LLM-backed validator's own retry budget. A semantic correction rebuilds the
ordinary producer request and supplies only the latest rejected model response
plus aggregated validator guidance; it is not a conversation replay.
After the applicable budgets are exhausted, the resolved
[`validation_policy`](config.md#pipelines) determines the result. Structural
failure and semantic rejection normally fail the run; an explicit
`reject_output` records a rejection and allows unrelated work to finish. A
validator execution failure normally uses `warn_continue`, which keeps an
otherwise accepted result in the current run with `incomplete` validation
provenance. It emits one bounded warning for every validator whose execution
budget was exhausted. A corrected result that later passes validation does not
retain diagnostics from abandoned attempts.
Treat a successful process exit as a completed run, not as proof that every
candidate was fully validated. Inspect the receipt's `validation_status`,
`validation_summaries`, rejection count, and warning group and occurrence
counts when an orchestrator requires complete validation. The durable fields
and their meanings are owned by the
[run-result receipt](integrations/run-result.md) and
[published JSON output contract](integrations/json-output.md).
## Output Bundles ## Output Bundles
Each successful run receives a generated safe run identifier and writes beneath: Each successful run receives a generated safe run identifier and writes beneath:
@@ -160,9 +74,9 @@ are defined in [Accepted Chunk Map](integrations/chunk-map.md). An optional
[evidence context](integrations/evidence-context.md) contains source-unit text [evidence context](integrations/evidence-context.md) contains source-unit text
and metadata. It is not a cache or debug artifact: retain it with the output and metadata. It is not a cache or debug artifact: retain it with the output
bundle only for as long as consumers need it, and apply source-content access bundle only for as long as consumers need it, and apply source-content access
controls to the entire bundle. Its selected source-unit excerpt may include controls to the entire bundle. Selected lanes may collectively cite most of a
every source unit once when coverage is broad or its configured window is transcript, so a broad allowlist can make the evidence artifact nearly as
large, so do not assume a byte or token reduction or reduced sensitivity. sensitive and large as the source itself.
## Chunk-Plan Cache ## Chunk-Plan Cache
@@ -188,9 +102,7 @@ The configured cache mode controls one invocation:
A reused plan is still materialized and validated against the current source. A reused plan is still materialized and validated against the current source.
If a prior plan no longer gives acceptable results, use a refresh run rather If a prior plan no longer gives acceptable results, use a refresh run rather
than editing cache files. Deleting a plan is recoverable but can repeat costly than editing cache files. Deleting a plan is recoverable but can repeat costly
chunking work. A plan accepted only under incomplete validation is not chunking work.
published, and a rejected cache hit falls through to ordinary generation rather
than becoming a correction candidate.
## Checkpoint Recording, Resume, And Recompute ## Checkpoint Recording, Resume, And Recompute
@@ -205,21 +117,9 @@ compatible recorded work. A resume request fails when checkpoint recording is
disabled. Without **--resume**, a recording-enabled run executes normally and disabled. Without **--resume**, a recording-enabled run executes normally and
does not load checkpoint state. Compatibility includes the resolved pipeline, does not load checkpoint state. Compatibility includes the resolved pipeline,
input, selected lanes, runtime overrides, reference provenance, LLM-profile input, selected lanes, runtime overrides, reference provenance, LLM-profile
provenance, the effective PromptKit profile-source fingerprint, and provenance, and prepared-component fingerprints. A changed identity produces a
prepared-component fingerprints. When a local PromptKit backend is configured, cold miss; Notarius does not migrate, rewrite, or delete older checkpoint
compatibility also includes a non-secret fingerprint of its endpoint. Changing directories.
profile content or the local endpoint causes a cold miss; changing only the
local concurrency limit does not. A changed identity produces a cold miss;
Notarius does not migrate, rewrite, or delete older checkpoint directories.
Reasoning-effort inheritance, replacement, and explicit clearing are distinct
runtime identities, so checkpoints created under one state are not reused by
either of the others.
Only accepted, completely validated chunk, extract, merge, and normalize
results are checkpointed for reuse. Rejected, structurally invalid, and
validation-incomplete producer results remain non-reusable, even when a
`warn_continue` result advanced during its original run. A resumed invocation
therefore reruns that producer rather than treating degraded state as accepted.
Checkpoint state is confined below an identity-specific path: Checkpoint state is confined below an identity-specific path:
@@ -275,16 +175,11 @@ Only a [debug-enabled run](cli.md#run) creates a bundle:
~~~ ~~~
The summary contains redacted invocation and resolution information plus run, The summary contains redacted invocation and resolution information plus run,
final grouped diagnostic, checkpoint, chunk-plan, and terminal reporting artifacts. Attempt warning, checkpoint, chunk-plan, and terminal reporting artifacts. The trace
terminal records contain bounded attempt kinds, validator outcomes, policy, contains allowlisted application diagnostic records and can include source or
decision, PromptKit repair count, and usage; they do not contain assistant derived application data. Neither surface is a cache input. Do not treat a
responses or complete correction messages. The trace contains allowlisted debug bundle as safe to share merely because its configuration summary is
application diagnostic records and can include source, model, and correction redacted.
content. Neither surface is a cache input. Do not treat a debug bundle as safe
to share merely because its configuration summary is redacted. Invocation
metadata omits reasoning effort when it is inherited, records the replacement
value when one is supplied, and records an empty value when inherited reasoning
was explicitly cleared.
Notarius never creates debug state without an explicit request and never Notarius never creates debug state without an explicit request and never
automatically deletes a requested bundle. If allocation succeeds, the command automatically deletes a requested bundle. If allocation succeeds, the command
@@ -312,71 +207,10 @@ or automatic cleanup command.
## Operational Limits ## Operational Limits
Provider execution settings and the generation timeout come from the selected Provider retries and timeouts are supplied by the selected Scriptorium profile.
PromptKit profile. The invocation-only **--reasoning-effort** and Module retry settings and concurrency limits are configuration contracts; see
**--clear-reasoning-effort** controls may replace or clear that profile setting [module bindings](config.md#module-bindings-and-validators) and
for all LLM-backed calls in one run without changing the profile. PromptKit
structural output repair happens within one structured-completion call. Its
effective `structured_output_repair_attempts` limit is resolved from the
selected binding, then the pipeline, then the prompt declaration; see
[module bindings](config.md#module-bindings-and-validators). This is distinct
from Notarius binding **retries**, which rerun the complete module operation
and validation chain and do not consume or replenish the structural-repair
limit. The maintained production prompts declare one repair attempt, paid only
after a structural failure. One structured completion with repair budget **R**
makes at most **R + 1** serial provider calls. If one stage attempt performs
**C** structured completions, a binding with **retries: N** has a maximum of
**(N + 1) * C * (R + 1)** provider calls; LLM-backed validators have their own
corresponding invocation counts and budgets. This is an upper bound, not a
promise that every call reaches a provider.
Timeouts are layered. Caller cancellation is the outer authority. A positive
effective generation timeout adds an inner request deadline, while zero
disables only that generation deadline. The HTTP client timeout remains a
transport-wide cap. Notarius does not add another timeout around PromptKit.
Repairs are serial within the same caller context, so their worst-case latency
and cost follow the provider-call bound above; provision run deadlines and
provider budgets accordingly. Credentials remain optional unless the selected
PromptKit profile requires one, in which case preparation fails before a
provider call when its configured credential is unavailable.
The pinned upstream boundary and profile-format links are in
[PromptKit Integration](integrations/pkg-promptkit.md).
Concurrency has two independent layers. Notarius **total_llm** defaults to 16
and is the application-wide provider-call limit shared by all backends,
modules, retries, and validators. PromptKit may impose a narrower admission
limit for the selected backend. The effective active-generation bound is the
intersection of the Notarius limit, any PromptKit backend limit, and work made
available by the pipeline. Built-in OpenRouter profiles use PromptKit's
upstream backend limit; endpoint-only profiles have no PromptKit backend limit
and remain bounded by Notarius. For the configured local backend, a zero
**concurrency_limit** leaves only the Notarius scheduler as a call limit. A
positive value makes the effective active local-generation bound the smaller
of **total_llm** and that local limit, so a local limit of four permits no more
than four active local generations.
The Notarius scheduler admits one logical structured completion and holds that
permit while PromptKit performs its serial corrective calls. PromptKit applies
its selected-backend admission to each provider call; Notarius does not
reacquire a permit or add another scheduler for a repair.
For a positive local limit, PromptKit owns its default waiting capacity and
admission behavior. When a PromptKit backend has admitted all active and queued
work, a new call fails as capacity exhaustion before generation. The adapter
maps that failure to Notarius's existing provider-neutral capacity error and
does not retry it. The calling stage's configured retry policy applies
normally, and the run fails if those attempts are exhausted. Caller
cancellation remains authoritative. Configuration contracts are documented
under [PromptKit profiles](config.md#promptkit-profiles) and
[concurrency](config.md#concurrency-output-cache-and-debug). Extract-worker [concurrency](config.md#concurrency-output-cache-and-debug). Extract-worker
limits and actual provider-call limits are independent. Notarius writes local limits and actual provider-call limits are independent. Notarius writes local
filesystem state only; remote storage, archival, and retention automation are filesystem state only; remote storage, archival, and retention automation are
outside the implemented CLI. outside the implemented CLI.
Every run has an effective prompt session used for provider routing and run
provenance. The generated default is stable for the same input module and raw
input bytes; use [**--session-id**](cli.md#run) only when intentionally grouping
different invocations. Both generated and explicit values can be visible to
providers, manifests, checkpoints, and requested debug bundles. Do not put
credentials or other secrets in an explicit session identifier; command-line
values are not a credential mechanism.

View File

@@ -24,12 +24,6 @@ DAGs or a general workflow language. Every stage remains explicit; general
chunking, merging, or normalization behavior must not be hidden inside an chunking, merging, or normalization behavior must not be hidden inside an
extractor. extractor.
A stage module is one configured implementation of one pipeline stage. An
artifact family is the cohesive domain feature that owns an artifact across
the explicit stages and supporting codecs, validators, prompts, identity
rules, and reference projections. Artifact-family ownership does not combine
stages or alter the fixed pipeline.
Input and chunking are pipeline-wide. Each selected artifact lane owns its Input and chunking are pipeline-wide. Each selected artifact lane owns its
extract, merge, and normalize stages, and the output stage aggregates the run's extract, merge, and normalize stages, and the output stage aggregates the run's
lane outcomes. lane outcomes.
@@ -45,24 +39,11 @@ implementations. Domain-neutral model and framework layers provide reusable
policy, contracts, and orchestration. Concrete input, pipeline, output, and policy, contracts, and orchestration. Concrete input, pipeline, output, and
validation extensions depend inward on those generic layers. validation extensions depend inward on those generic layers.
Semantic reconciliation is one such domain-neutral framework mechanism. It
prepares bounded source context, invokes a shared model-judgment protocol,
validates proposals, and applies safe plans through typed policies supplied by
the consuming artifact family. It does not own domain identity, durable IDs,
warning semantics, or artifact construction rules.
Generic layers must not depend on production extensions. Concrete extensions Generic layers must not depend on production extensions. Concrete extensions
must not compose the application or take ownership of process behavior. The must not compose the application or take ownership of process behavior. The
current packages implementing these layers are inventoried in current packages implementing these layers are inventoried in
[Internal Overview](../internal/overview.md). [Internal Overview](../internal/overview.md).
The root `assets` package is a content-only dependency leaf. It may expose a
read-only embedded filesystem, but it must contain no business logic and must
not depend on `internal` packages or PromptKit. Consumers scope that filesystem
to the content they own; the root package is not a behavioral registry or a
public extension contract. The rationale and compatibility consequence are
recorded in [ADR-0011](../adr/0011-centralize-llm-assets.md).
The following dependency boundaries are mandatory: The following dependency boundaries are mandatory:
- extractors and validators do not depend on concrete input adapters; - extractors and validators do not depend on concrete input adapters;
@@ -92,12 +73,6 @@ Extract modules own artifact semantics, prompt use, response schemas, and
domain interpretation. Domain-specific concepts remain in the relevant module, domain interpretation. Domain-specific concepts remain in the relevant module,
validator, shared domain helper, and artifact contract. validator, shared domain helper, and artifact contract.
Physical centralization of LLM-facing content does not transfer semantic
ownership from those modules. Modules retain their manifests, response-schema
identities, prompt ordering, and registration, while reading only their scoped
content subtree. Generic framework code remains domain-neutral when it reads
its own scoped generic assets from the shared content container.
Typed artifact registrations declare one stable artifact kind and exact Go Typed artifact registrations declare one stable artifact kind and exact Go
type from extraction through merge, normalization, and semantic validation. type from extraction through merge, normalization, and semantic validation.
Pipeline resolution requires a compatible codec and matching kind-specific Pipeline resolution requires a compatible codec and matching kind-specific
@@ -146,9 +121,8 @@ lanes, validators, and LLM profile: the canonical source digest selects the
plan, while the current run still applies its configured chunk validators to plan, while the current run still applies its configured chunk validators to
the materialized chunks. the materialized chunks.
The framework owns orchestration, origin enrichment, aggregation, and handoff The framework owns orchestration and handoff provenance. Modules return logical
provenance. Modules return logical results and classified diagnostics; they do results and warnings; they do not own CLI reporting, physical output, cache, or
not own CLI reporting, physical output, cache, or
debug roots, durable file placement, or checkpoint and debug lifecycle. debug roots, durable file placement, or checkpoint and debug lifecycle.
After pipeline-wide chunking, extraction uses bounded framework concurrency. After pipeline-wide chunking, extraction uses bounded framework concurrency.
@@ -160,46 +134,26 @@ may overlap. The framework must not create unbounded goroutines per lane or
chunk. chunk.
Completion timing does not choose public ordering or errors. The coordinator Completion timing does not choose public ordering or errors. The coordinator
orders accepted artifacts, grouped diagnostics, rejections, checkpoint events, and orders accepted artifacts, warnings, rejections, checkpoint events, and
framework errors by stable pipeline scope. Rejections do not cancel unrelated framework errors by stable pipeline scope. Rejections do not cancel unrelated
work. A framework error cancels derived work, prevents undispatched work from work. A framework error cancels derived work, prevents undispatched work from
starting, waits for started work, and prevents output encoding. starting, waits for started work, and prevents output encoding.
Warnings are process-only signals: configuration degradation, approved fallback,
or incomplete configured validation. Quality uncertainty and grounding findings
are advisories; successful canonicalization and cleanup are observations.
Modules choose that semantic classification, while the framework attaches
origin, aggregates groups, enforces bounds, and presents final collections.
An ordinary successful run therefore has zero warnings. See
[ADR-0015](../adr/0015-separate-process-warnings-from-quality-diagnostics.md)
for the decision rationale.
## Validation ## Validation
Validation is a framework-managed boundary around outputs from chunk, extract, Validation is a framework-managed boundary around outputs from chunk, extract,
merge, and normalize stages. Validators receive immutable stage output and merge, and normalize stages. Validators receive immutable stage output
make an explicit whole-output decision: approve, reject, fail, or skip when a and make an explicit whole-output decision: approve, approve with warnings, or
runtime prerequisite is unavailable. reject.
Typed artifact validators receive the domain value directly. Chunk validators Typed artifact validators receive the domain value directly. Chunk validators
receive source-zone chunks, while serialized validators receive immutable receive source-zone chunks, while serialized validators receive immutable
representation bytes and declared schema metadata. A validator registered for representation bytes and declared schema metadata. A validator registered for
one target or artifact kind cannot satisfy an incompatible selection. one target or artifact kind cannot satisfy an incompatible selection.
The framework runs every applicable validator sequentially in configured order. Rejection is a recorded pipeline outcome, not a framework execution error.
It aggregates rejections, exhausted validator failures, and skips before the Validator execution failures are framework errors. Rejected output does not
producer policy chooses a disposition. A completed rejection never advances. advance to the next stage.
With no rejection, an exhausted validator failure may fail the run or, under
the configured `warn_continue` policy, advance a structurally valid candidate
with explicit incomplete-validation provenance. Validators report findings;
they do not choose candidate disposition.
A completed rejection supplies a stable reason code for internal provenance
and bounded actionable correction guidance for the candidate producer. Reason
codes, validator keys, and operator-facing messages remain diagnostic data;
they are not model instructions. The framework constructs model-facing retry
text only from the semantic guidance and fails the contract rather than
inventing or truncating missing guidance.
Default validator chains are production composition policy and are registered Default validator chains are production composition policy and are registered
centrally by stage and module. Configuration may replace a stage-local default, centrally by stage and module. Configuration may replace a stage-local default,
@@ -216,30 +170,6 @@ The caller of the LLM owns prompt selection, prompt inputs, response schema,
and interpretation of structured output. Provider adapters do not own source- and interpretation of structured output. Provider adapters do not own source-
or domain-specific prompt logic. or domain-specific prompt logic.
PromptKit owns bounded structural correction within one structured completion.
Notarius owns outer stage attempts, semantic validation, and acceptance policy;
the two budgets must remain separate.
An LLM-backed producer can participate in semantic correction only when it
declares `single_response_v1` and returns the exact one response that directly
controlled its candidate. On an actionable rejection, the framework rebuilds
the ordinary request and appends only the latest defective response as an
`assistant` message plus one aggregated `user` correction message. This is a
fresh replacement request, not a growing conversation. The retry budgets,
terminal policy, and sensitive-data rationale are recorded in
[ADR-0014](../adr/0014-feedback-aware-validation-retries.md).
When a model selects an application entity, callers must supply a contextual
selection and deterministically attach the opaque application identity whenever
the selection resolves exactly. Models do not receive or reproduce opaque
application identifiers. Semantic reconciliation may instead expose
contiguous, one-based candidate handles that exist only for one request;
deterministic code resolves them before typed application, and they never
become durable identity. This is the approved request-local-label application
of [ADR-0012](../adr/0012-resolve-opaque-entity-identifiers-deterministically.md)
recorded by
[ADR-0013](../adr/0013-use-request-local-candidate-handles-for-semantic-reconciliation.md).
LLM calls and other external operations accept cancellation and respect LLM calls and other external operations accept cancellation and respect
timeouts. Concurrency control belongs in shared runtime plumbing rather than in timeouts. Concurrency control belongs in shared runtime plumbing rather than in
individual modules. individual modules.
@@ -247,8 +177,7 @@ individual modules.
The application-wide LLM scheduler bounds actual provider calls independently The application-wide LLM scheduler bounds actual provider calls independently
of framework worker limits. Every LLM-backed module, retry, and validator uses of framework worker limits. Every LLM-backed module, retry, and validator uses
the single injected scheduled client, including work performed by overlapping the single injected scheduled client, including work performed by overlapping
lanes. Provider runtime adapters may enforce a narrower backend-specific limit lanes.
beneath this mandatory application-wide scheduler.
## Configuration And Provenance ## Configuration And Provenance
@@ -263,8 +192,7 @@ invalid or incompatible.
Run manifests record enough resolved pipeline, module, source, reference, and Run manifests record enough resolved pipeline, module, source, reference, and
LLM provenance to make a run auditable after configuration changes. Manifests LLM provenance to make a run auditable after configuration changes. Manifests
record identities and bounded validation summaries rather than secret, raw record identities and summaries rather than secret or large payload content.
model, correction, or large payload content.
## State, Output, And Safety ## State, Output, And Safety
@@ -284,39 +212,18 @@ an invocation that explicitly requests resume. Debug is never a cache input and
is never created without an explicit request. Pipeline modules receive is never created without an explicit request. Pipeline modules receive
collaborator interfaces and never physical roots. collaborator interfaces and never physical roots.
Only accepted, completely validated producer output is reusable checkpoint or
chunk-plan state. Rejected, structurally invalid, and validation-incomplete
results cannot become cache or checkpoint inputs, even when a
`warn_continue` result is allowed to advance in the current run. This
ineligibility follows derived merge and normalize results and generated
references for the remainder of the run: current-run handoff remains allowed,
but no dependent cache or checkpoint may be loaded or published.
Writes are atomic where practical. Paths for writes, moves, overwrites, and Writes are atomic where practical. Paths for writes, moves, overwrites, and
deletion must be narrow and explicit. Notarius never automatically deletes deletion must be narrow and explicit. Notarius never automatically deletes
output or requested debug bundles; cache cleanup is explicit and recoverable. output or requested debug bundles; cache cleanup is explicit and recoverable.
Secrets must not appear in errors, logs, output, cache, debug summaries, Secrets must not appear in errors, logs, output, cache, debug summaries,
traces, manifests, documentation, examples, or redacted configuration. Raw traces, manifests, documentation, examples, or redacted configuration. Debug
assistant responses and complete correction messages are attempt-local and are
excluded from ordinary durable records and summaries; the requested detailed
debug trace is the sole diagnostic surface allowed to retain them. Debug
collection is allowlisted to application-owned payloads and must not capture collection is allowlisted to application-owned payloads and must not capture
unrelated process environment values or filesystem content. Trace data may unrelated process environment values or filesystem content. Trace data may
contain application data and therefore inherits its sensitivity; operators own contain application data and therefore inherits its sensitivity; operators own
access controls and retention. Physical layout and operation are defined in access controls and retention. Physical layout and operation are defined in
[Operations](../operations.md). [Operations](../operations.md).
## Platform And Distribution
Linux is the supported deployment platform. macOS is supported only as a
best-effort development and compilation environment, while Windows is
unsupported. Notarius distributes source releases only: an immutable source
tag and its checked-in release note identify a release. The project does not
publish executable binaries, archives, installers, container images,
checksums, signatures, or package-manager entries. Maintainer release commands
and tag guards belong to [Source Releases](../release.md).
## Architectural Non-Goals ## Architectural Non-Goals
Notarius does not aim to provide: Notarius does not aim to provide:

View File

@@ -65,8 +65,6 @@ secret values.
| CLI contract | `docs/cli.md` | Commands, arguments, flags, invocation semantics, and exit codes. | End-to-end operating procedures, configuration field definitions, runtime filesystem layout, module implementation details. | | CLI contract | `docs/cli.md` | Commands, arguments, flags, invocation semantics, and exit codes. | End-to-end operating procedures, configuration field definitions, runtime filesystem layout, module implementation details. |
| Configuration contract | `docs/config.md` | Discovery and precedence, file schema, fields, defaults, environment overrides, validation rules, and user-selectable module or validator keys. | Complete example files, CLI syntax, runtime state lifecycle, module implementation details. | | Configuration contract | `docs/config.md` | Discovery and precedence, file schema, fields, defaults, environment overrides, validation rules, and user-selectable module or validator keys. | Complete example files, CLI syntax, runtime state lifecycle, module implementation details. |
| Operations | `docs/operations.md` | Runtime workflows, physical filesystem and state layout, output, cache, and debug handling, resume, cleanup, permissions, recovery, and operational limits. | CLI flag syntax, configuration field definitions, logical output schemas, implementation mechanics. | | Operations | `docs/operations.md` | Runtime workflows, physical filesystem and state layout, output, cache, and debug handling, resume, cleanup, permissions, recovery, and operational limits. | CLI flag syntax, configuration field definitions, logical output schemas, implementation mechanics. |
| Source release procedure | `docs/release.md` | Maintainer release selection, candidate validation, tagging, publication guards, verification, and immutable-tag recovery. | Product installation summary, CLI version semantics, historical release summaries, CI implementation detail. |
| Release-note history | `docs/releases/` | One checked-in historical summary for each source release made under the procedure. The note at the immutable tag is that release's record. | Current commands, behavior, contracts, and compatibility definitions. |
| Public HTTP contract, if introduced | `docs/api.md` | Routes, authentication, media types, request and response schemas, status codes, pagination, caching, idempotency, rate limits, and HTTP retry semantics. | Client walkthroughs, upstream or downstream integration internals, implementation detail. | | Public HTTP contract, if introduced | `docs/api.md` | Routes, authentication, media types, request and response schemas, status codes, pagination, caching, idempotency, rate limits, and HTTP retry semantics. | Client walkthroughs, upstream or downstream integration internals, implementation detail. |
| Consumer guidance, if a public package or API is introduced | `docs/consumers/` | Task-oriented use of the public interface, minimal client examples, and consumer responsibilities. | HTTP wire semantics, external protocol contracts, internal implementation detail. | | Consumer guidance, if a public package or API is introduced | `docs/consumers/` | Task-oriented use of the public interface, minimal client examples, and consumer responsibilities. | HTTP wire semantics, external protocol contracts, internal implementation detail. |
| External and durable integration contracts | `docs/integrations/` | External file formats and protocols, upstream and downstream contracts, logical output bundle paths and schemas, media types, and compatibility behavior. | Physical runtime placement and lifecycle, internal transformations, CLI syntax, configuration defaults. | | External and durable integration contracts | `docs/integrations/` | External file formats and protocols, upstream and downstream contracts, logical output bundle paths and schemas, media types, and compatibility behavior. | Physical runtime placement and lifecycle, internal transformations, CLI syntax, configuration defaults. |
@@ -97,14 +95,6 @@ runtime state and how to operate or recover the application. When a workflow
crosses these topics, choose the document that owns the task and link to the crosses these topics, choose the document that owns the task and link to the
other contracts. other contracts.
### Releases
`docs/release.md` owns the source-release procedure. Release notes are
historical summaries, not current-state contract owners: the checked-in note at
an immutable tag records that release, while current canonical documentation
must change with the behavior it describes. Do not use a release note to defer
or replace current documentation updates.
### Contracts And Implementation ### Contracts And Implementation
Integration and API documents define externally observable shapes and Integration and API documents define externally observable shapes and

View File

@@ -1,164 +0,0 @@
# Source Releases
This procedure is for maintainers publishing Notarius source releases. A
release is an immutable lightweight `vMAJOR.MINOR.PATCH` tag on `main` together
with its checked-in `docs/releases/<tag>.md` note. Tag CI validates that source
candidate after publication; it does not publish or repair a release.
Notarius publishes no binaries, archives, checksums, signatures, containers,
package-manager entries, or Gitea release objects. Windows is not supported.
Do not create retrospective notes for the pre-procedure `v0.1.0`, `v0.2.0`, or
`v0.3.0` tags.
## Select And Describe The Release
Choose an unused stable semantic version in the form `vMAJOR.MINOR.PATCH`.
Prereleases are not supported. Before `v1.0.0`, a minor release may change a
documented CLI, configuration, durable artifact, integration, or operating
contract when its note explains the impact and required operator action. A
patch release must not intentionally break those documented contracts within
its minor line.
Create the version-matched note as part of the candidate. Every new note uses
this structure, with concise, truthful content in each section:
```markdown
# Notarius vMAJOR.MINOR.PATCH
This release ...
## Summary
## Compatibility
## Upgrade
## Changes
```
The note is a historical summary. Link to current canonical documentation for
exact behavior, and update that documentation in the candidate rather than
using the note as a substitute.
## Prepare The Candidate
Set the selected release version and disable Go workspace use for every
candidate command:
```sh
RELEASE_VERSION=vMAJOR.MINOR.PATCH
export RELEASE_VERSION GOWORK=off
```
Run the shared source-candidate checks from the repository. They cover module
hygiene, tests, race tests, vet, builds, formatting, whitespace, maintained
configuration validation, and the Linux and Darwin command-build matrix:
```sh
./scripts/check-release-source.sh "$RELEASE_VERSION"
```
Before committing, manually follow every changed local Markdown link and
review the candidate for unintended files, generated output, credentials, or
other unrelated changes. Commit the release note and all affected current
documentation, then run the shared checker against that exact candidate. Push
the candidate commit to `main` only after it succeeds. Record the exact commit
only after that push:
```sh
RELEASE_COMMIT=$(git rev-parse 'HEAD^{commit}')
export RELEASE_COMMIT
```
For private-module installation, configure standard `GOPRIVATE` matching this
module and ordinary Git authentication for the hosting service before running
the verification below. The exact authentication mechanism belongs to the
maintainer environment; never record credentials or environment dumps in a
release note, command history, or repository file.
## Guard And Publish The Tag
Fetch current remote references, then run this guard without editing the
candidate. It requires `main`, a clean worktree and index, disabled workspace
use, a stable release version, the recorded and pushed commit, a matching note,
and unused local and remote tags:
```sh
git fetch origin main --tags
if ! printf '%s\n' "$RELEASE_VERSION" |
grep -E -x 'v(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)' >/dev/null
then
printf '%s\n' "invalid release version: $RELEASE_VERSION" >&2
exit 1
fi
test "$GOWORK" = off
test "$(git branch --show-current)" = main
test -z "$(git status --porcelain)"
test "$RELEASE_COMMIT" = "$(git rev-parse 'HEAD^{commit}')"
test "$RELEASE_COMMIT" = "$(git rev-parse 'origin/main^{commit}')"
test -s "docs/releases/$RELEASE_VERSION.md"
grep -F -x "# Notarius $RELEASE_VERSION" "docs/releases/$RELEASE_VERSION.md"
for heading in '## Summary' '## Compatibility' '## Upgrade' '## Changes'; do
grep -F -x "$heading" "docs/releases/$RELEASE_VERSION.md"
done
if git rev-parse -q --verify "refs/tags/$RELEASE_VERSION" >/dev/null; then
printf '%s\n' "local tag already exists: $RELEASE_VERSION" >&2
exit 1
fi
if git ls-remote --exit-code --tags origin "refs/tags/$RELEASE_VERSION" >/dev/null 2>&1; then
printf '%s\n' "remote tag already exists: $RELEASE_VERSION" >&2
exit 1
fi
```
Create an explicitly lightweight tag against the guarded commit, verify its
target, and push only that tag ref:
```sh
git -c tag.gpgSign=false tag "$RELEASE_VERSION" "$RELEASE_COMMIT"
test "$(git cat-file -t "$RELEASE_VERSION")" = commit
test "$(git rev-parse "$RELEASE_VERSION^{commit}")" = "$RELEASE_COMMIT"
git push origin "refs/tags/$RELEASE_VERSION:refs/tags/$RELEASE_VERSION"
```
Never use `git push --tags`, move a published tag, or delete a published tag.
## Verify The Published Release
Confirm that the remote tag still points at the guarded commit and that the
note is available from the tagged tree:
```sh
REMOTE_TAG_COMMIT=$(git ls-remote origin "refs/tags/$RELEASE_VERSION" | awk '{print $1}')
test "$REMOTE_TAG_COMMIT" = "$RELEASE_COMMIT"
git show "$RELEASE_VERSION:docs/releases/$RELEASE_VERSION.md" >/dev/null
```
Verify a fresh source installation and its diagnostic version. The temporary
directory confines the installed command to this check:
```sh
release_verification_dir=$(mktemp -d)
trap 'rm -rf "$release_verification_dir"' 0 HUP INT TERM
mkdir -p "$release_verification_dir/bin"
GOWORK=off GOBIN="$release_verification_dir/bin" go install \
"gitea.maximumdirect.net/eric/notarius/cmd/notarius@$RELEASE_VERSION"
test "$("$release_verification_dir/bin/notarius" --version)" = "notarius $RELEASE_VERSION"
```
An exact fresh checkout and `GOWORK=off go build ./cmd/notarius` is an
equivalent source verification when local installation policy requires it.
`notarius --version` is diagnostic only; downstream compatibility remains
defined by the published receipt and artifact contracts.
## Failure And Correction Policy
If candidate validation fails before publication, fix the candidate on `main`,
rerun the shared checker, and repeat the guards. An unpublished local tag may
be deleted after inspection.
If the remote tag or tag CI reveals a defect, leave the published tag intact.
Fix the defect on `main`, choose a new patch version, write a new matching
note, and repeat this procedure. Do not weaken tag immutability or add release
assets as a workaround.

View File

@@ -1,83 +0,0 @@
# Notarius v0.4.0
This release strengthens LLM reliability and validation throughout the
configured pipeline, upgrades the PromptKit integration, and establishes the
source-release and downstream-consumer workflows needed for broader D&D
pipeline integration.
## Summary
Notarius now distinguishes PromptKit structural-output repair from
application-owned semantic validation retries. Producer candidates can run
through complete deterministic validator chains, receive bounded semantic
correction guidance, and retry under explicit stage policies. Final run
receipts and manifests preserve bounded validation provenance, while outputs
that advance with incomplete validation remain available to the current run
without entering reusable checkpoint state.
The release also adds a maintained complete D&D subprocess-consumer workflow,
diagnostic build versions, and the source-only release procedure used to
publish this version.
## Compatibility
- Configuration files must use schema version 4. Version 3 is not decoded or
rewritten; rename the top-level `scriptorium` section to `promptkit` when
migrating. See [Configuration](../config.md#migrating-version-3-configuration).
- PromptKit is pinned to v0.9.0. Operator profile files use PromptKit's v0.9.0
format and may use its profile-inheritance support. Notarius continues to
resolve operator profiles before embedded fallbacks.
- Structural-output repair and semantic stage retries are separate bounded
mechanisms. Maintained production prompts request one structural repair by
default; explicit configuration can override the supported repair count.
- Validation policy can now fail a run, reject an output, or permit an
otherwise valid candidate to advance with incomplete-validation provenance.
The application defaults are documented in
[Configuration](../config.md#pipelines).
- The `notarius.run-result.v1` receipt remains at schema version 1 and adds
optional validation summaries plus a required validation-status field.
Consumers of this pre-release contract should follow the current
[run-result receipt](../integrations/run-result.md).
- Existing D&D artifact schema identities remain unchanged. Validation and
producer-policy changes can nevertheless cause previously accepted weak
candidates to retry, reject, or fail instead.
## Upgrade
1. Migrate every Notarius configuration to version 4 and rename `scriptorium`
to `promptkit`.
2. Review deployed PromptKit profiles against the pinned v0.9.0 profile format
and ensure their credential environment variables are available at run
time.
3. Run `notarius config validate --config <path> --pipeline <id>` before the
first production invocation.
4. Review `structured_output_repair_attempts`, producer retry counts, and
`validation_policy` wherever the deployment needs behavior different from
the documented defaults.
5. Update subprocess consumers to inspect receipt `validation_status` and to
tolerate the optional bounded `validation_summaries` field. A consumer that
requires fully validated artifacts should require `approved`.
## Changes
- Upgraded PromptKit from v0.5.0 through v0.9.0 and adopted profile
inheritance, structured-output repair, typed error classification, and the
correction-aware completion protocol.
- Added pipeline and binding configuration for structural repair and terminal
validation policy, with strict startup validation and effective-setting
provenance.
- Added feedback-aware retries for chunking, extraction, merge, normalize, and
semantic reconciliation producers. Retry prompts contain the exact defective
response and actionable semantic correction guidance without exposing
internal reason codes or opaque entity identifiers.
- Added complete validator-chain execution, validator retry handling, bounded
warnings, terminal dispositions, and durable validation summaries.
- Prevented validation-incomplete artifacts and all derived lineage from
loading or publishing reusable checkpoints while preserving same-run
generated-reference handoff.
- Tightened cached chunk-plan validation so only completely validated plans are
reused or replace stored plans.
- Added a machine-readable subprocess receipt workflow and complete D&D
consumer documentation covering all maintained artifacts.
- Added source-release checks, immutable lightweight-tag guidance, Linux and
Darwin build verification, and diagnostic `notarius --version` output.

View File

@@ -1,85 +0,0 @@
# Notarius v0.5.0
This release separates actionable process warnings from extraction-quality
advisories and routine normalization observations, giving operators a quiet
warning channel without discarding durable diagnostic detail.
## Summary
Notarius now carries one validated, origin-aware diagnostic contract from
producers and validators through retries, reusable state, output publication,
debug summaries, run receipts, and CLI presentation. Warnings are reserved for
process degradation or incomplete configured work. Data-quality findings are
advisories, and successful deterministic cleanup is recorded as observations.
An ordinary successful run therefore reports zero warnings while retaining
bounded diagnostic provenance for later review.
The framework aggregates findings deterministically by their stable identity
and complete pipeline origin, preserves exact occurrence counts, and retains
bounded representative samples. Warning groups fail rather than truncate;
advisory and observation representation is bounded with explicit truncation
metadata and exact unrepresented-occurrence counts.
## Compatibility
- `warnings.json` now uses the incompatible grouped
`notarius.warnings.v2` envelope and contains process warnings only. Consumers
of the former flat warning payload must migrate to the current
[JSON output contract](../integrations/json-output.md).
- The new `diagnostics.json` file uses `notarius.diagnostics.v1` and contains
advisory and observation groups. Production `index.json` files always expose
both `warnings_file` and `diagnostics_file`.
- The machine-readable run receipt is now `notarius.run-result.v2`. It replaces
`warning_count` with exact warning group and occurrence counts and adds
advisory/observation group, occurrence, and truncation fields. See the
current [run-result receipt](../integrations/run-result.md).
- Custom output modules must return their complete logical file set or an
error. The former `OutputResult.Warnings` field has been removed; an output
module cannot report a warning after serializing its output.
- Reusable state now uses `notarius.workspace.v4` and chunk-plan records use
`notarius.chunk-plan.v3` so they can preserve structured diagnostics. Older
pre-release reusable state is not reused under these contracts; start with
clean state when deterministic continuity with an older workspace is not
required.
- Validation acceptance, semantic retry budgets, rejection policy, and D&D
artifact schema identities are unchanged by this release.
## Upgrade
1. Update subprocess consumers to require `notarius.run-result.v2` and read
`warning_group_count`, `warning_occurrence_count`,
`diagnostic_group_count`, `diagnostic_occurrence_count`, and
`diagnostics_truncated`.
2. Update output-bundle consumers to decode `notarius.warnings.v2`, discover
`diagnostics.json` through `index.json`, and treat diagnostics as review
information rather than process warnings.
3. Update any custom output module for the removal of
`OutputResult.Warnings`; return an error when encoding cannot complete.
4. Clear pre-release reusable state before the first upgraded production run
when deterministic continuity with an older workspace is not required.
5. Run `notarius config validate --config <path> --pipeline <id>` and perform
one representative run before promoting the release in an automated
pipeline.
## Changes
- Added validated diagnostic dispositions, categories, origins, stable reason
codes, exact occurrence counts, and bounded representative samples.
- Added deterministic run-level aggregation with separate limits for
actionable warning groups and advisory/observation groups.
- Reclassified D&D source-relatedness and unresolved-identity findings as
data-quality advisories and routine normalization changes as observations.
- Preserved structured diagnostics across producer retries, validation,
generated-reference handoff, checkpoints, chunk-plan reuse, and debug
summaries while discarding superseded-attempt findings.
- Added grouped `warnings.json`, a new grouped `diagnostics.json`, and the
corresponding production index entries.
- Upgraded the machine-readable run receipt and human CLI summary to report
exact warning and diagnostic counts without allowing advisory volume to
create warning output.
- Removed post-encoding output warnings and hardened diagnostic validation,
overflow handling, aggregate memory bounds, and warning-file path
presentation.
- Documented diagnostic ownership, classification, operator interpretation,
durable contracts, and architectural invariants in ADR-0015 and the
canonical CLI, operations, integration, and internal documentation.

File diff suppressed because it is too large Load Diff

View File

@@ -1,648 +0,0 @@
# Warning Signal And Presentation Audit
## Executive Assessment
Notarius warning execution is mechanically stronger than its operator-facing
presentation. Terminal-attempt promotion, stable ordering after concurrent
work, checkpoint replay, validation summaries, and debug retention are all
substantially correct. The audit found no general duplicate-append defect in
the extract, merge, or normalize handoffs and no leakage of abandoned-attempt
warnings into a successful result.
The warning channel itself is not coherent. One flat `contracts.Warning` type
currently represents at least four materially different concepts:
- actionable degradation or incomplete validation;
- heuristic data-quality doubt;
- successful but potentially reviewable fallback; and
- routine canonicalization, ordering, and deduplication observations.
That conflation is the primary reason successful runs produce a count that is
large but operationally weak. The maintained complete D&D example demonstrates
the problem without a live provider: an approved run with no rejected outputs
published 12 warning records, all from three advisory relatedness checks. An
operator separately reported a successful complete D&D run with 10 outputs,
one rejection, and 85 warnings. The production bundle for that run was not
available in this environment, so its reason-code distribution could not be
measured.
The current implementation also has four correctness or robustness gaps:
1. warning records lose stage, step, lane, module, validator, and chunk
provenance when promoted, which makes safe aggregation and diagnosis
impossible from `warnings.json` alone;
2. there is no framework-level validation or aggregate bound, and the NPC- and
spell-relatedness validators bypass the D&D warning limiter entirely;
3. warnings returned by an output encoder are added after `warnings.json` has
already been encoded, so the receipt, stderr, debug bundle, and published
warning file can disagree; and
4. a skipped validator contributes to `incomplete` validation but does not
receive the warning generated for an exhausted validator failure.
The recommended end state is a structured diagnostic contract with explicit
disposition, category, origin, occurrence count, and bounded samples. Warnings
are reserved for process-level degradation or incompleteness. LLM-judged or
deterministically inferred extraction-quality signals are advisories, never
warnings, and routine normalization observations remain inspectable without
being reported as top-level warnings. An ordinary successful run in which all
configured work completes normally should therefore report zero warnings. This
is an architectural and durable-contract change, not merely revised CLI prose.
## Evidence And Limits
The audit used:
- a complete static search of production `contracts.Warning` constructors,
reason-code constants, result fields, and promotion sites under `internal/`;
- call-path inspection through producer attempts, validators, lane
coordination, chunk-plan reuse, checkpoints, output encoding, debug output,
CLI presentation, and run-result construction;
- the maintained complete and minimal D&D examples with offline fake LLMs;
- focused deterministic tests for warning bounds, semantic-reconciliation
fallback, `warn_continue`, semantic retries, terminal rejection, concurrency
ordering, and checkpoint reuse; and
- the operator-provided observation of an 85-warning complete D&D run.
No provider-backed production run was attempted because this environment has
no API key. Consequently, the audit can establish warning mechanics, possible
multiplicity, synthetic volume, and obvious heuristic limitations, but cannot
estimate production frequency or the real false-positive rate of individual
D&D advisories. Those measurements are not required to choose the recommended
architecture; they are required before strengthening any heuristic advisory
into a rejection or setting a numerical production acceptance target.
## Complete Warning-Producer Inventory
### Framework And Generic Boundaries
| Producer | Reason code | Trigger and consequence | Multiplicity and bound | Current surfaces and coverage |
| --- | --- | --- | --- | --- |
| Reference materialization in `internal/framework/pipeline/references.go` | `empty_reference` | A bound external reference is a valid, accepted media type but contains zero bytes. The prompt may receive materially incomplete context. | One per empty bound file; finite by configuration but no shared run-level cap. | Enters `RunInput.Warnings`; reference tests protect contextual scope. |
| Producer-attempt policy in `internal/framework/pipeline/producer_attempts.go` | `validator_execution_incomplete` | An applicable validator exhausted its execution budget and `warn_continue` accepted the otherwise valid candidate. | One per failed validator on each terminal candidate. An extract chain can multiply this by chunks and lanes. There is no global cap. | Durable warning, receipt count, stderr, debug, and validation summary. `TestWarnContinueRecordsOneWarningForEachExhaustedValidator` covers failures. |
| Chunk, extract, merge, normalize, and output module result contracts | Module-defined | A module may return arbitrary warnings with its successful candidate. | No contract validation, message limit, per-result cap, or global cap. Current production modules are inventoried below. | Terminal-attempt filtering and concurrency ordering are well tested. |
| Production JSON output encoder | None | The encoder copies incoming warnings into `warnings.json`; it does not currently create warnings. | Same count as its input. | JSON encoder and assembled-pipeline tests compare the incoming run warnings with the published file. |
| Output encoder result contract | Module-defined | Any output encoder may return warnings discovered during encoding. | Unbounded by contract. No production encoder currently exercises this capability. | Appended to final `RunOutput` only after logical files were encoded; this is the cross-surface defect described in AUD-WARN-004. |
Input adapters and production mergers do not currently have independent
warning producers. Chunk-plan and checkpoint decisions are structured manifest
or debug provenance rather than warnings. Cancellation and hard persistence,
reference, parsing, serialization, and provider failures remain errors.
### D&D Extraction Gates
| Producer | Reason code | Trigger and consequence | Multiplicity and bound |
| --- | --- | --- | --- |
| `dnd/combat-turns` extractor | `scene_classification_unavailable` | The chunk has no exact matching scene-description classification. The extractor returns an empty accepted result and skips the LLM, so combat-turn output may be incomplete. | At most one per chunk for this lane. |
| `dnd/enemy-events` extractor | `scene_classification_unavailable` | The same missing or mismatched scene gate causes accepted empty enemy-event output. | At most one per chunk for this lane. |
An exact non-combat classification produces an intentional empty result without
a warning. An exact combat classification proceeds normally. The two producers
share a code and operator consequence but use different messages; their module
origins are not retained in the final warning record.
### D&D Source-Relatedness Validators
All ten relatedness validators are deterministic advisories: they approve the
candidate and warn when contextual prose or an entity name is not lexically
present in cited text. Shape and source-reference failures are deliberately
left to blocking validators earlier in the chain. The same relatedness
validator is registered in both the extract and normalize default chain for
each artifact family in `internal/modules/dnd/register/chains.go`.
| Artifact family | Warning reason | Per-record trigger | Local bound | Omission reason |
| --- | --- | --- | --- | --- |
| Combat turns | `combat_turn_not_near_source` | Actor token sequence absent | 20 per validator invocation | `combat_turn_relatedness_warnings_omitted` |
| Enemy events | `enemy_event_not_near_source` | Subject token sequence absent | 20 | `enemy_event_relatedness_warnings_omitted` |
| Item occurrences | `item_occurrence_source_unrelated` | Item name token sequence absent | 20 | `item_occurrence_relatedness_warnings_omitted` |
| Item registry | `item_not_near_source` | Item name token sequence absent | 20 | `item_relatedness_warnings_omitted` |
| Location occurrences | `location_occurrence_not_near_source` | Location name token sequence absent | 20 | `location_occurrence_relatedness_warnings_omitted` |
| Location registry | `location_not_near_source` | Location name token sequence absent | 20 | `location_relatedness_warnings_omitted` |
| NPC occurrences | `npc_occurrence_not_near_source` | NPC name token sequence absent | 20 | `npc_occurrence_relatedness_warnings_omitted` |
| NPC registry | `npc_not_near_source` | NPC name token sequence absent | **Unbounded** | None |
| Scene descriptions | `scene_description_not_near_source` | No significant title or summary token appears; up to two findings per scene | 20 | `scene_description_relatedness_warnings_omitted` |
| Spells | `spell_not_near_source` | Spell-name token sequence absent | **Unbounded** | None |
The eight limiter-generated omission records are presentation artifacts, not
new source-relatedness conditions. They occupy a warning slot and make list
length differ from the actual occurrence count.
### D&D Normalizers
Every production D&D normalizer bounds its returned warning slice to 20 through
`internal/modules/dnd/shared/diagnostics`, including a final omission record
when needed. Registry semantic retries reserve space for their fallback
warning. The following table is complete by semantically distinct condition;
codes listed together are parallel artifact-family variants.
| Condition | Reason codes | Result impact | Current classification assessment |
| --- | --- | --- | --- |
| Display or field whitespace/name canonicalization | `npc_fields_normalized`, `item_fields_normalized`, `location_fields_normalized`, `spell_name_canonicalized`, `combat_actor_canonicalized`, `enemy_event_name_canonicalized`, `item_occurrence_name_canonicalized`, `location_occurrence_name_canonicalized`, `scene_description_prose_normalized` | Deterministic successful mutation. The item-occurrence code can also describe `from`/`to` whitespace, not only the item name. | Routine observation. |
| Durable ID recomputation | `npc_id_recomputed`, `item_id_recomputed`, `location_id_recomputed` | Restores the deterministic name-derived ID. | Routine observation; invalid identity is separately rejected by default chains. |
| Source-reference sorting or deduplication | `source_references_normalized` | Sorts and removes exact duplicate references while deliberately preserving invalid references for their validators. | Routine observation. Shared code is useful but ambiguous without producer origin. |
| Canonical record ordering | `combat_turns_reordered`, `enemy_events_reordered`, `item_occurrences_reordered`, `location_occurrences_reordered`, `npc_occurrences_reordered`, `scene_description_order_normalized` | Deterministic order changes only. | Routine observation. |
| Exact or approved semantic duplicate consolidation | `duplicate_npc_collapsed`, `duplicate_item_collapsed`, `duplicate_location_collapsed`, `duplicate_spell_cast_collapsed`, `duplicate_combat_turn_collapsed`, `duplicate_enemy_event_collapsed`, `duplicate_item_occurrence_collapsed`, `duplicate_location_occurrence_collapsed`, `duplicate_npc_occurrence_collapsed`, `scene_description_duplicate_collapsed` | Removes duplicate records and preserves or combines canonical evidence according to the artifact policy. Registry codes cover both exact and accepted semantic consolidation. | Durable normalization observation; not normally operator-actionable. |
| Unresolved external membership | `spell_name_unresolved`, `item_occurrence_unknown_item_id`, `location_occurrence_unknown_location_id` | The value is preserved but is not grounded in the effective catalog or registry. Default chains normally reject the same condition before normalization; it remains reachable with validator overrides or defensive direct use. | Actionable data-quality warning. |
| Unsafe currency consolidation proposal | `item_semantic_proposal_invalid` | The proposed group is rejected and all records are preserved because denominations or currency/non-currency members are incompatible. The same code is also used internally as a retry reason. | Advisory about model proposal quality; no accepted-data loss. The control and diagnostic meanings should be separated. |
| Semantic reconciliation unavailable or exhausted | `npc_semantic_reconciliation_exhausted`, `item_semantic_reconciliation_exhausted`, `location_semantic_reconciliation_exhausted` | The safe deterministic result is accepted, but possible semantic duplicates remain. | Actionable fallback warning. |
| Local warning truncation | `npc_normalization_warnings_omitted`, `item_normalization_warnings_omitted`, `location_normalization_warnings_omitted`, `spell_normalization_warnings_omitted`, `combat_turn_normalization_warnings_omitted`, `enemy_event_normalization_warnings_omitted`, `item_occurrence_normalization_warnings_omitted`, `location_occurrence_normalization_warnings_omitted`, `npc_occurrence_normalization_warnings_omitted`, `scene_description_normalization_warnings_omitted` | Reports that individual records were omitted from presentation. | Group metadata, not an independent warning. |
No production D&D normalization warning exposes raw model responses,
correction guidance, or provider errors. Most dynamic names are quoted and
truncated by the shared helper. That local discipline is not enforced by the
generic warning contract, and the spell relatedness message does not use the
shared truncation helper.
## Warning Propagation And Surface Map
```text
external-reference warnings -----------------------------+
|
module candidate warnings -> validation chain warnings |
| | |
+---- producer-attempt terminal policy -------+
| |
accepted / terminal rejection only |
| |
chunk or lane result in canonical order |
| |
checkpoint record/replay and ordered step merge |
| |
RunOutput.Warnings <------------+
|
OutputRequest -> output encoder
| |
warnings.json OutputResult.Warnings
|
appended to final RunOutput only
|
receipt, stderr, final debug warning summary
```
### Attempts And Validation
- `runProducerAttempts` promotes only the terminal accepted or terminal
rejected candidate's module and completed-validator warnings. Operational,
structural, semantic, and module-directed attempts that are superseded are
retained in attempt debug artifacts but not in the final collection.
- A module-directed semantic-reconciliation retry adds its fallback warning
only when no retry remains. Earlier attempt warnings are discarded.
- On `warn_continue`, warnings from the otherwise accepted candidate and
completed approved or rejected validators are retained. One fixed,
non-sensitive `validator_execution_incomplete` warning is added for every
failed validator. Skipped validators affect the validation summary and final
`incomplete` status but do not receive such a warning.
- A semantic terminal rejection retains only warnings from that rejected
attempt. Structural rejection after producer failure cannot retain a
candidate warning because no valid candidate result exists.
These behaviors are protected by the producer-attempt, extract-handoff,
rejection-warning, normalize-retry, and attempt-debug tests. They are the right
foundation for the redesign and should not be replaced with early-exit or
all-attempt accumulation.
### Concurrency And Ordering
Extract jobs are dispatched chunk-first and lane-second. Results are stored by
chunk index, finalized in ascending chunk order, and lane continuations are
merged into a slice indexed by configured lane order. Pipeline steps run in
configured order. The resulting public order is therefore:
1. pre-run reference warnings;
2. chunk-stage warnings;
3. step order;
4. configured lane order within each step;
5. chunk order within extract;
6. merge warnings; then
7. normalize warnings; followed by any output-result warnings.
`TestRunnerBoundsExtractJobsAndStabilizesReverseCompletion` exercises warning
order under reversed completion. No completion-order leak was found.
### Chunk Plans And Checkpoints
- A reusable chunk plan stores producer warnings only. Current validators run
again, and their current warnings are appended. Warnings from a cached plan
candidate that fails current validation are discarded before regeneration.
- Accepted extract, merge, and normalize checkpoints store the terminal
warnings for their stage. Reuse loads and appends those warnings once at the
same logical handoff. Tests compare fresh and resumed warning collections and
preserve their order.
- Validation-incomplete accepted outputs are not reusable, preventing a later
run from silently treating incomplete validation as complete.
- Required accepted-normalize hydration replays that normalize checkpoint's
warnings; checkpoint decisions separately expose that reuse occurred.
The recommended redesign should keep fresh and resumed logical diagnostics
equivalent. Whether a result was reused belongs in checkpoint provenance, not
in the diagnostic grouping key; adding a `reused` distinction would fragment
groups and make equivalent runs present differently.
### Terminal Surfaces
| Surface | Current content | Audience | Audit result |
| --- | --- | --- | --- |
| `RunOutput.Warnings` | Flat final slice | Framework and CLI | Canonical in-memory list, but lacks origin and bounds. |
| Published `warnings.json` | Object containing the warnings passed into the output encoder | Durable consumers | Exact for the production JSON encoder unless the encoder itself returns warnings. |
| `index.json` | Path to `warnings.json` | Durable consumers | Stable discovery path; no separate diagnostic-detail path. |
| Run-result v1 | `warning_count = len(final RunOutput.Warnings)` | Subprocess callers | Count only; no group/occurrence distinction. |
| Human stderr | `run completed with N warning(s)` | Operators | Count only and no direct detail path. Successful exit remains zero. |
| Manifest | Validation and rejection summaries, no warning collection | Durable provenance | Correctly avoids duplicating the flat list. |
| Debug summary `warnings.json` | Raw final warning array | Operators/developers | Includes final output-result warnings and can therefore differ from published `warnings.json`. |
| Debug run report | Final warning count | Operators/developers | Same final slice length as receipt and stderr. |
| Attempt/stage debug | Candidate-local warning detail and origin in path/envelope | Forensics | Sufficient to diagnose provenance, but debug capture is optional and is not a durable consumer contract. |
## Empirical Measurements
### Offline And Synthetic Runs
| Scenario | Result | What it establishes |
| --- | --- | --- |
| Maintained complete D&D config and transcript with the repository's offline fake LLM | Approved, 10 normalized outputs, 0 rejected outputs, 12 warnings; receipt, stderr, and published file all reported 12 | An ordinary structurally successful workflow can be noisy without fallback or incomplete validation. |
| Same complete run, grouped after publication | Three reason codes, seven exact `(reason, scope, message)` tuples, maximum exact-tuple repetition of three | The flat count materially overstates distinct operator conditions. Scope resets within chunks and does not identify origin. |
| Maintained focused scene-description workflow | Approved, one normalized output, 0 warnings | The warning channel can be quiet when synthetic model text is lexically grounded. |
| Generic warning publication contract | One warning reaches successful stderr, durable output, and debug summary | The ordinary pre-output path is consistent. |
| NPC semantic-reconciliation candidate-limit fallback | No LLM call, all records preserved, one exhaustion warning, total warnings no greater than 20 | Fallback is bounded and materially different from routine normalization. |
| `warn_continue` with two failed validators and one skipped validator | Validation status contains all three incomplete validators; warning slice contains two execution-incomplete records | Current warning count does not describe all incomplete validation. |
| Retrying extract candidate | Two producer attempts; only the accepted attempt's one warning is final | Retry does not amplify abandoned warnings. |
| Terminal semantic rejection | Only the final rejected attempt's operation and validator warnings are final | Rejection diagnostics are retained without retaining superseded warnings. |
| Fresh versus reused extract checkpoint | Warning collections are deeply equal | Checkpoint replay does not itself amplify warnings. |
| Spell normalizer with 21 unresolved entries | 20 records: 19 samples plus one omission record saying two additional warnings were omitted | `warning_count` is neither exact occurrence count nor distinct-condition count. |
The focused audit tests passed in `internal/cli`,
`internal/framework/pipeline`, the NPC-registry and spell normalizers, and all
D&D packages.
### Bounded Sample Review
The complete offline D&D run produced:
| Reason | Count | Sample | Review |
| --- | ---: | --- | --- |
| `location_not_near_source` | 4 | `Moon Gate` was absent from cited text | Correctly identifies deliberately unsupported fake output. Three records shared the same exact tuple because chunk and stage origin were lost. |
| `location_occurrence_not_near_source` | 4 | A `Moon Gate` visit was absent from cited text | Correctly identifies the same unsupported registry-driven occurrence, but repeats the same operator concern across extraction and normalization. |
| `scene_description_not_near_source` | 4 | A title or summary had no significant exact token in cited text | Mixed value. Generic `session scene` prose is ungrounded, while `Arrival` versus transcript `arrive` illustrates an expected lexical false positive. |
This fake workflow is an integration fixture, not a model-quality benchmark.
It nonetheless proves that the checks carry useful evidence while being too
imprecise and repetitive to serve as one-warning-per-record operator alerts.
### Production Evidence Still Needed
The reported 85-warning run establishes that high volume occurs in practice,
but the following remain unknown:
- dominant production reason codes and stage/lane sources;
- unique group count versus repeated occurrence count;
- false-positive rate for each relatedness family;
- how much volume comes from normalization observations versus advisories;
- whether fresh and resumed production runs remain equivalent; and
- a defensible numerical acceptance target.
If further data is worthwhile, the operator can supply the v1 receipt,
`warnings.json`, and manifest validation summaries without supplying transcript
or lane artifacts. An initial privacy-preserving report should group by reason
code and normalized scope family, count exact repeated tuples, and omit message
text. Reviewing heuristic precision requires a separately approved bounded
sample with its cited source context.
## Classification Of Current Conditions
| Target disposition | Current families | Result impact | Operator action | Durable placement |
| --- | --- | --- | --- | --- |
| **Warning** | Empty reference; validator failure or skip accepted under `warn_continue`; unavailable required scene classification; exhausted semantic reconciliation | A configured process completed under an allowed degraded or incomplete policy rather than completing normally | Correct reference/configuration, inspect provider/validator, or rerun | Actionable `warnings.json`, receipt summary, stderr summary, debug |
| **Advisory** | Source-relatedness heuristics; unresolved spell or registry membership; guarded invalid semantic proposal; any future LLM-judged uncertainty or extraction-quality signal | Uncertain data quality or poor model proposal, but accepted data is structurally valid and deterministic guards prevented unsafe mutation | Optional model/source review; no routine action for every record | Durable diagnostic detail and debug; never a top-level warning |
| **Observation** | Whitespace/name/ID/source-reference canonicalization; canonical ordering; exact and approved semantic duplicate consolidation | Successful intended normalization | None under normal operation | Durable bounded normalization diagnostics or debug; no stderr warning |
| **Not a diagnostic** | Rejection, invalid structure, cancellation, persistence error, provider failure under fail-run policy | Candidate or run did not complete according to policy | Inspect rejection/error and retry or correct input/configuration | Existing rejection, validation summary, error, and debug contracts |
Exact and semantic duplicate consolidation should remain distinguishable in
category or reason metadata even though both are observations. Semantic
reconciliation exhaustion remains a warning because a capability was not
applied; successful approved consolidation is an observation because it is the
normalizer's intended work.
## Ranked Findings
### AUD-WARN-001 — The Flat Warning Type Destroys Signal Quality
- **Priority:** High operator impact; high implementation leverage.
- **Evidence:** `contracts.Warning` has only scope, reason, and message. Routine
normalizer changes, heuristic doubt, fallback, and incomplete validation all
enter the same slice and the same CLI count. The offline complete run's 12
records were all advisories; the operator observed 85 records in a successful
real run.
- **Impact:** Operators cannot tell whether a warning requires a rerun, a
configuration repair, optional review, or no action. Repeated routine output
trains them to ignore the channel.
- **Recommendation:** Replace the flat semantic contract with explicit
`warning`, `advisory`, and `observation` dispositions plus a small category
vocabulary. Do not infer disposition from message text or require every
downstream consumer to maintain a reason-code policy table.
### AUD-WARN-002 — Warning Records Lose The Origin Needed For Diagnosis And Aggregation
- **Priority:** High correctness and usability impact.
- **Evidence:** The runner knows stage, step, lane, module, validator, chunk ID,
and chunk index at promotion time, but `terminalWarnings` flattens module and
validator records into `[]contracts.Warning`. Per-chunk scopes such as
`locations[0]` and `occurrences[0]` then repeat without identifying their
chunk or producer. `source_references_normalized` is intentionally shared
across families and is therefore especially ambiguous.
- **Impact:** `warnings.json` cannot answer which stage or module produced a
record. Message- or scope-based deduplication would merge unrelated findings
or retain accidental duplicates.
- **Recommendation:** Keep module findings free of framework context, then have
the framework add a structured origin envelope before promotion. Validator
findings must retain validator identity instead of passing through
`validationReport.Warnings()` as a flat slice.
### AUD-WARN-003 — Warning Volume Is Not End-To-End Bounded Or Validated
- **Priority:** High robustness impact; medium immediate likelihood.
- **Evidence:** D&D's `LimitWarnings` caps most individual producers at 20, but
NPC- and spell-relatedness return one warning per record without the helper.
Every extract validator is invoked per chunk, all ten relatedness checks run
again after normalization, and there is no run-level collector. The generic
contract validates neither disposition nor reason/message size, UTF-8,
blankness, or total records.
- **Impact:** Warning memory and output grow with chunks, lanes, configured
validators, and record counts. Local omission records lose exact occurrence
semantics while still incrementing `warning_count`.
- **Recommendation:** Add a generic bounded diagnostic collector that preserves
exact occurrence counts and bounded samples. Validate all diagnostic fields
at the module/framework boundary. Immediately bring NPC and spell
relatedness under the existing cap if the full redesign is staged.
### AUD-WARN-004 — Output Encoder Warnings Make Durable Surfaces Disagree
- **Priority:** Medium current impact; high contract correctness risk.
- **Evidence:** `Runner.Run` passes existing warnings to `encoder.Encode`, then
the production encoder serializes `warnings.json`. Only after encoding does
the runner append `OutputResult.Warnings`. The receipt, stderr, debug summary,
and debug run report see the final slice; the already-created published file
cannot. No production encoder currently returns a warning, so ordinary JSON
runs do not trigger the defect.
- **Impact:** A valid output-module implementation can violate the documented
claim that `warning_count` describes the published warning collection.
- **Recommendation:** Remove successful output warnings from the output-module
contract unless a demonstrated use case requires them; encoding failures
should be errors and optional encoder observations should be debug data. A
two-phase finalize API is the viable but more complex alternative.
### AUD-WARN-005 — Validation Skips Are Incomplete But Not Warned
- **Priority:** Medium operator/correctness impact.
- **Evidence:** `firstIncompleteValidation` treats failed and skipped validators
alike, and validation summaries include both. `incompleteValidationWarnings`
emits records only for `validationFailed`. The focused test demonstrates
three incomplete validators but two warnings.
- **Impact:** A successful run can have `validation_status: incomplete` while
its warning count understates or even omits the affected validators. A caller
that checks only warnings receives a weaker signal than the manifest and
receipt status.
- **Recommendation:** Produce one aggregated incomplete-validation warning
group whose occurrences cover both failure and skip, while retaining typed
outcome and safe reason metadata in the validation summary. Do not expose
provider errors or arbitrary skip prose in model or operator messages.
### AUD-WARN-006 — Relatedness Checks Are Useful But Repetitive And Lexically Weak
- **Priority:** Medium operator impact; low acceptance-policy urgency.
- **Evidence:** Every family runs the advisory in both extract and normalize
chains. The complete fixture contains exact repeated tuples, and the checks
rely on exact normalized token sequences or a minimal significant-token
overlap. Reason naming drifts between `*_not_near_source` and
`*_source_unrelated`.
- **Impact:** The checks can catch unsupported entities, but aliases, pronouns,
inflection, and generic scene prose create predictable false positives. Flat
per-record presentation magnifies them.
- **Recommendation:** Retain the validators and their stage-local execution,
but classify and aggregate them as advisories. Normalize reason-code naming
when the diagnostic contract changes. Do not strengthen them into rejection
rules without a human-reviewed production evaluation.
### AUD-WARN-007 — `warning_count` Has No Stable Operational Meaning
- **Priority:** High downstream-contract impact.
- **Evidence:** The receipt and CLI use `len(output.Warnings)`. One list element
can be an omission summary representing several hidden occurrences; repeated
records can represent the same condition; and skipped validators can be
absent. A 21-occurrence spell test produces a list length of 20.
- **Impact:** The value is neither an exact occurrence count nor a distinct
warning-group count. Consumers cannot set policy or present a trustworthy
summary from it.
- **Recommendation:** Introduce explicit warning-group and warning-occurrence
counts in a versioned receipt. Do not silently redefine the v1 field.
## Recommended Target Contract And Presentation Model
### Diagnostic Model
Use one validated internal diagnostic model with these concepts:
- **disposition:** `warning`, `advisory`, or `observation`;
- **category:** a small enum such as `configuration`, `degradation`,
`validation_incomplete`, `data_quality`, `fallback`, or `normalization`;
- **reason code:** stable semantic identity owned by the producer;
- **origin:** framework-added phase/stage, step ID, lane ID, module key,
validator name, chunk ID, and chunk index when applicable;
- **occurrence count:** exact number of matching findings;
- **samples:** a small deterministic list of bounded scope/message pairs; and
- **omitted sample count:** `occurrence_count - len(samples)`, represented as
metadata rather than another diagnostic record.
Errors and rejected outputs must not become diagnostic dispositions. A warning
means that the run completed under policy despite a process-level degradation
or incomplete configured operation. Advisory and observation dispositions can
describe accepted artifact quality and transformation provenance, but no
LLM-judged extraction-quality signal may be promoted to a warning. Validation
status remains authoritative for approval, rejection, and incomplete
validation.
### Aggregation
The framework runner should own aggregation after it enriches findings with
origin and before public output construction. Modules and validators retain
semantic ownership of disposition, category, reason, scope, and message; they
must not own CLI or file presentation.
The default stable key should be:
```text
disposition + category + reason_code
+ phase/stage + step_id + lane_id + module_key + validator_name
```
Chunk, record scope, and message text belong in samples and must not be part of
the group key. This groups repeated per-chunk findings without merging the same
code across distinct producers or pipeline locations. Group order should be
the first occurrence in the runner's existing canonical order; sample order
should follow the same order. A final canonical sort by the complete origin key
is also viable, but completion timing must never choose either order.
Aggregation must be incremental and bounded. Producers should use a shared
collector that counts every occurrence while retaining only bounded samples;
the framework then merges producer groups without reconstructing counts from
omission prose. A global maximum group count is also required, with overflow
represented by structured aggregate metadata and with actionable groups given
priority over lower dispositions.
### Durable Files
Keep one canonical home for each class:
- `warnings.json` should contain versioned, grouped actionable warnings only;
- a new `diagnostics.json` should contain versioned advisory and observation
groups only, avoiding duplication of warning groups;
- `index.json` should link both files;
- `rejected.json` and manifest validation summaries should retain their current
separate responsibilities; and
- debug bundles should retain candidate-attempt detail plus the final grouped
projections.
This is preferable to keeping all detail in `warnings.json` and filtering only
the CLI: downstream consumers would otherwise continue to receive a semantically
mixed warning contract, and routine observations would still dominate the
durable file.
### CLI And Receipt
For a successful human run with actionable warnings, print a concise summary
such as:
```text
notarius: run completed with 2 warning groups (7 occurrences); details=/.../warnings.json
```
Advisories and observations should not produce the warning line. Their durable
path remains discoverable through `index.json`; a concise non-warning count can
be added to the ordinary success line only if operator testing shows value. An
ordinary successful run with no process degradation should write nothing to
the warning stream even when it publishes quality advisories or normalization
observations.
Create `notarius.run-result.v2` rather than redefining v1. It should expose at
least:
- `warning_group_count`;
- `warning_occurrence_count`; and
- `diagnostic_group_count` for non-warning durable groups.
The receipt should continue to expose validation status, validation summaries,
and rejected-output count independently. Process exit behavior should not
change as part of warning presentation reform.
### Checkpoint Semantics
Store the structured terminal diagnostic groups with accepted checkpoints and
replay them exactly once at their logical stage. Fresh and reused runs should
produce the same public groups and counts. Checkpoint events and debug records,
not diagnostic identity, should disclose whether computation was reused.
### Output Encoder Boundary
Prefer removing `OutputResult.Warnings`. A successful output encoder should
either return the complete logical files or fail. If future encoders genuinely
need to produce durable post-encoding warnings, introduce an explicit
two-phase prepare/finalize contract so those warnings can be included in the
same published collection. Do not retain the current self-inconsistent
one-pass capability.
## Resolution Of Required Design Questions
| Question | Recommendation | Viable alternative and tradeoff |
| --- | --- | --- |
| Explicit severity/disposition or external reason mapping? | Put validated disposition and category in the contract. | A central reason-code registry avoids payload fields but makes new modules depend on a second synchronized policy table and leaves downstream meaning implicit. |
| Keep all detail in `warnings.json` or separate it? | Separate grouped actionable warnings from grouped advisories/observations in `diagnostics.json`. | Keep the flat durable list and aggregate only CLI output; simpler migration, but it preserves the noisy downstream contract and ambiguous count. |
| Who owns aggregation? | Framework runner/coordinator after origin enrichment. | Output module aggregation keeps framework types smaller but duplicates policy across encoders and cannot repair missing validator origin. |
| Stable aggregation key? | Disposition, category, reason code, and full producer origin; exclude chunk/scope/message. | Explicit producer-supplied grouping keys offer flexibility but add another identity that can drift from reason codes. Message-template grouping is brittle and unsafe. |
| Samples and omissions? | Exact occurrence count plus deterministic bounded samples and numeric omitted-sample count. | Omission warning records preserve the current representation but inflate group counts and require prose parsing. |
| `warning_count` semantics? | Version receipt and replace ambiguity with group and occurrence counts. | Keep v1 count as published record length and add optional fields; compatible, but two competing warning counts remain easy to misuse. |
| Which normalization changes remain warnings? | Only exhausted process fallback. Unresolved membership is a data-quality advisory; successful canonicalization, reordering, ID repair, source-ref dedupe, and duplicate consolidation are observations. | Treat unresolved membership or semantic duplicate consolidation as warnings because they affect grounding or cardinality; more conservative, but it violates the process-only warning rule and reports accepted artifact quality as an operational failure. |
| Source-relatedness disposition? | Grouped advisory by default; preserve current approve behavior. | Retain warning disposition or make rejection configurable. Rejection requires production precision evidence; current lexical rules are not strong enough. |
| Checkpoint-loaded warnings? | Present the same logical groups as fresh execution and use checkpoint events for reuse provenance. | Mark groups as replayed; aids forensics but fragments aggregation and makes semantically equivalent runs differ. |
| ADR and schema versions? | Add an ADR and version the run receipt and diagnostic files. | Treat the work as CLI-only presentation and avoid an ADR; insufficient because module contracts, output files, checkpoint payloads, and downstream fields change. |
## Compatibility, Documentation, And ADR Implications
The target alters public and internal contracts enough to require a new ADR.
It should record:
- the distinction among warnings, advisories, observations, rejections, and
errors;
- the invariant that warnings are process-level signals, LLM-judged extraction
quality is never a warning, and ordinary non-degraded success has zero
warnings;
- module semantic ownership versus framework origin/aggregation ownership;
- bounded group and sample semantics;
- fresh/checkpoint equivalence; and
- the output-encoder decision.
Implementation should introduce `notarius.run-result.v2`. The grouped warning
and diagnostic envelopes should each carry their own schema version. Because
the content of `warnings.json` changes incompatibly from a flat array wrapper
to groups, release notes and the published JSON integration contract must call
out the migration. `index.json` gains the diagnostic file path.
Canonical documentation updates belong in:
- `docs/cli.md` for stderr presentation only;
- `docs/operations.md` for operator review and debug workflow;
- `docs/integrations/json-output.md` for warning and diagnostic file schemas;
- `docs/integrations/run-result.md` for v2 fields and compatibility;
- `docs/consumers/subprocess.md` and `docs/consumers/dnd-pipeline.md` for
downstream policy checks;
- `docs/internal/pipeline.md` for promotion, aggregation, retry, and checkpoint
mechanics;
- `docs/internal/modules.md` and `docs/internal/dnd.md` for producer rules and
the D&D classification matrix; and
- `docs/policy/architecture.md` for the durable ownership invariant after the
ADR is accepted and implemented.
No configuration knob is required for the first implementation. A fixed,
well-documented taxonomy is easier to reason about than per-reason display
overrides. Configurable escalation or suppression can be considered only after
production review demonstrates a concrete operator need.
## Test Coverage Assessment
Existing coverage worth preserving includes:
- accepted-attempt and terminal-rejection warning promotion;
- validator failure retry exhaustion and `warn_continue`;
- module semantic retry fallback;
- deterministic warning order under concurrent lane completion;
- chunk-plan invalidation and discarded-cache warning behavior;
- fresh/checkpoint warning equivalence;
- local D&D warning caps and safe dynamic-message quoting;
- JSON warning-file publication; and
- CLI stderr, debug, and receipt counts.
Material gaps are:
- no bound test for NPC- or spell-relatedness warnings;
- no generic warning-field or result-size validation;
- no test for output encoder warnings versus published `warnings.json`;
- no operator-level aggregation or bounded-sample tests;
- no fresh/resume test for grouped counts because groups do not yet exist; and
- no production evaluation of advisory precision.
Tests should protect the semantic relationships: exact occurrence counts,
bounded samples, deterministic group order, actionable-only warning
presentation, and cross-surface equality. They should not assert one exact
warning count for every complete D&D run or treat message wording as a public
API unless the wording itself enforces a security boundary.
## Audit Conclusion
The application is in a good position for warning reform. Its retry,
validation, checkpoint, and concurrency mechanics provide reliable points at
which to attach structured diagnostics. The most valuable change is not to
suppress individual reason codes; it is to replace the semantically flat,
origin-free collection with bounded typed groups and to reserve the word
“warning” for conditions that merit operator attention.
Provider-backed runs would improve prioritization and help tune the D&D
advisories, but they are not necessary to conclude that routine normalization
and heuristic doubt should not dominate stderr or the durable warning
contract. They should be gathered before changing heuristic acceptance policy
or adopting a numerical production warning-volume target.

226
docs/roadmap/evidence.md Normal file
View File

@@ -0,0 +1,226 @@
# Published Evidence Context
## Status
Implemented.
## Purpose
Let downstream consumers build narrative reports from normalized artifacts
without separately parsing the original transcript or resolving source-unit
references themselves.
The production JSON output optionally publishes one deterministic, deduplicated
evidence-context artifact containing the transcript units relevant to explicitly
selected normalized lanes. Existing lane payloads remain the canonical semantic
results and retain their precise source references.
## Desired End State
When evidence-context publication is enabled, a consumer can:
1. discover one versioned evidence-context document through `index.json`;
2. obtain the union of source units needed to understand evidence cited by the
selected normalized lanes;
3. distinguish each artifact's direct evidence references from surrounding
units included only for narrative context;
4. retain speaker, timestamp, and other accepted source-unit metadata needed to
interpret the transcript; and
5. produce a narrative report without receiving duplicated transcript text in
every lane payload.
This is deterministic output projection. It does not invoke an LLM, change
normalization, or make surrounding context part of an artifact's evidence.
## Configuration Policy
Evidence publication is configured on the production JSON output module. The
intended configuration shape is:
```yaml
output:
module: json
options:
evidence_context:
enabled: true
window_units: 3
lanes:
- combat-turns
- item-events
- npc-interactions
- npcs
- spells
```
- Omitting `evidence_context` disables publication. When the object is present,
`enabled` is required.
- `enabled: false` accepts no `lanes` or `window_units` fields, preventing
silently ignored configuration.
- `lanes` is a required, non-empty allowlist of configured final lane IDs when
evidence publication is enabled. Values are trimmed, unique, and normalized
to lexical order.
- `window_units` is a non-negative integer and defaults to `3`. Zero publishes
only directly referenced units.
- Unknown lanes, duplicate lane IDs, and selected lanes whose artifact kind
cannot expose source evidence fail configuration resolution or pipeline
preparation.
- A selected lane that completes without a normalized output contributes no
evidence and does not make an otherwise successful run fail.
- Invocation-level lane filtering does not invalidate the configured allowlist.
Allowlisted lanes excluded from the effective run contribute nothing, while
the evidence document still records the configured allowlist.
The allowlist is intentional safety and stability policy. Scene descriptions
and other broad-range lanes are excluded unless named expressly. Adding a new
pipeline lane never silently increases output size or publishes more transcript
content.
## Evidence Collection Boundary
Evidence collection applies to accepted final normalized artifacts from the
selected lanes. It must not inspect arbitrary serialized JSON for fields named
`source_ref` or `source_refs`, and the generic JSON output module must not
depend on D&D artifact types.
Artifact-kind registrations expose their source references through an explicit
typed projection contract. The framework uses that contract to assemble a
domain-neutral evidence request containing:
- the accepted generic source document;
- the selected lane and artifact identities; and
- defensive copies of their direct source references.
The output stage owns publication of the resulting logical artifact. Generic
framework code owns range validation, position-based expansion, and union
logic. Domain-specific adapters own only the extraction of evidence references
from their typed artifacts.
Both plural-reference artifacts and singular-reference artifacts, such as
scene descriptions, can participate through the same projection contract.
They do so only when their configured lane is allowlisted.
## Range Expansion And Deduplication
For every valid direct source reference:
1. resolve its endpoints through source-document positions, not numeric
unit-ID arithmetic;
2. expand the range by `window_units` positions on each side;
3. clip the expanded range at document boundaries; and
4. union overlapping or contiguous expanded ranges.
Published contexts and units remain in source-document order. Each source unit
appears at most once in a merged context. Original direct references remain
unchanged and are associated with their contributing lane IDs so consumers can
tell why a context was included.
The projector must not silently omit or repair an invalid reference that
reaches this boundary. Such a value violates the accepted normalized-artifact
contract and causes output projection to fail with a content-safe error.
No implicit coverage limit truncates selected evidence. If the allowlisted
lanes collectively cite most or all of a transcript, the evidence document may
contain most or all of it. The explicit lane allowlist is the control that
prevents a broad lane such as scene descriptions from doing so accidentally.
## Durable Evidence Artifact
The JSON bundle gains one optional, non-lane artifact with these durable
identities:
| Property | Value |
| --- | --- |
| Logical file | `evidence-context.json` |
| Index descriptor | `evidence_context` |
| Artifact kind | `source/evidence-context` |
| Media type | `application/json` |
| Schema ID | `notarius.source.evidence_context` |
| Schema name | `notarius_source_evidence_context_v1` |
| Schema version | `v1` |
The descriptor in `index.json` carries the artifact and schema identities,
analogous to the existing chunk-map descriptor. The artifact is present
whenever evidence publication is enabled, including when its context collection
is empty.
The document contains:
- the source document ID and semantic digest;
- the effective window size;
- the sorted configured lane allowlist;
- an ordered context collection;
- each context's expanded start and end unit IDs;
- the original direct references and contributing lane IDs covered by that
context; and
- the ordered accepted source units, including unit ID, kind, text,
self-reference, and metadata.
Expanded context bounds are navigation aids, not citations. The original
references embedded in each context remain the authoritative direct evidence.
The evidence artifact is discovered separately from lane payloads and does not
increase the normalized-lane count reported by the runner or subprocess
receipt.
## Failure And Publication Semantics
- Evidence projection occurs only after selected normalized outputs are known
and before the output encoder returns its logical files.
- Projection or encoding failure is an output-stage framework error; the CLI
does not publish a partially assembled output bundle.
- Rejected or absent lane outputs contribute nothing. Their attempted values
and source references must not be published through this artifact.
- Context generation is deterministic for the same source document, selected
normalized outputs, lane allowlist, and window size.
- Existing output, checkpoint, warning, rejection, debug, and subprocess
success semantics remain unchanged.
## Sensitivity And Size
Unlike the current chunk map, the evidence artifact contains transcript text
and source-unit metadata. Enabling it therefore creates additional durable
sensitive data and may materially increase bundle size.
The implemented configuration, operations, integration, and consumer documents
state that:
- evidence publication is opt-in;
- output permissions and retention must be appropriate for source content;
- selecting broad or numerous lanes can publish most of the transcript; and
- the artifact must not contain raw input bytes, LLM prompts or responses,
auxiliary reference content, credentials, debug-only data, or filesystem
paths.
## Acceptance Criteria
- Evidence publication is disabled by default and leaves existing bundles
unchanged.
- Enabling it requires an explicit non-empty lane allowlist.
- References from all selected successful lanes contribute to one deduplicated
document.
- Non-monotonic unit IDs are expanded and ordered correctly by document
position.
- Overlapping windows share one ordered copy of each included source unit.
- Direct references remain distinguishable from added context.
- Scene descriptions cannot contribute unless their lane is explicitly
allowlisted.
- Invalid selected lanes and unsupported artifact kinds fail before execution;
invalid accepted references fail output projection rather than being ignored.
- Empty selected-lane results produce a valid empty evidence artifact.
- Existing D&D lane schemas, normalized-output counts, and source-reference
semantics do not change.
- Generic framework and output packages do not depend on D&D types or parse
artifact JSON heuristically.
- The published contract and operational documentation clearly describe source
sensitivity, discovery, compatibility, and retention.
## Out Of Scope
- Embedding transcript units directly into each D&D record or lane payload.
- Replacing precise source references with expanded context ranges.
- Automatically including every configured lane.
- An explicit full-transcript publication mode.
- LLM summarization, retrieval, ranking, or narrative generation.
- Per-record window sizes or lane-specific window sizes.
- CLI overrides for evidence configuration.
- Reading rejected attempts, debug artifacts, auxiliary references, or prior
output bundles as evidence sources.

View File

@@ -5,64 +5,13 @@ configuration, operations, internal, and integration docs. This roadmap records
future work only. Items are ordered roughly by current value and specificity, future work only. Items are ordered roughly by current value and specificity,
not as committed release dates. not as committed release dates.
## Near-Term Validation And LLM Reliability
PromptKit now owns structural output repair within one completion. Notarius
owns stage candidates, validator chains, semantic rejection policy, bounded
feedback-aware stage retries, validation provenance, and reusable-state
eligibility, and the separation of actionable process warnings from quality
diagnostics. The remaining near-term work applies those completed foundations
to domain review and empirical evaluation.
### D&D Combat Scene Semantic Validation
- Add an optional production LLM-backed D&D validator that determines whether
proposed scene boundaries and classifications represent substantive active
combat correctly. Its central quality goal is that active combat is kept in
coherent scenes classified as `combat`, rather than split incorrectly or
hidden inside scenes classified as `narrative`, `recap`, or `meta`.
- Resolve the validator's exact target before implementation. The current
`dnd/scenes` chunker owns only complete, gap-free source ranges, while the
per-chunk `dnd/scene-descriptions` extractor owns the `combat`, `narrative`,
`recap`, and `meta` classification. The preferred initial placement is
therefore an extract-stage validator for `dnd/scene-descriptions`, where it
can compare one proposed kind with the corresponding transcript chunk.
- Consider a chunk-stage LLM validator only for a distinct boundary-coherence
question that can be answered from the complete transcript and proposed
range map, such as whether one continuous combat was fragmented across
inappropriate scene boundaries. Do not duplicate the same classification
judgment at both stages. Moving classification into chunk-plan annotations
would change the deliberately minimal, annotation-free chunk contract and
requires an explicit architecture review before it is selected.
- Validate both false negatives and false positives: a non-combat kind must not
omit substantive active combat, and a combat kind must be supported by such
combat. Keep the existing deterministic downstream rule that combat-turn
extraction runs only for an exact `combat` scene classification; semantic
review improves the upstream classification but does not replace that gate.
- Run the semantic validator through PromptKit, use a minimal required-field
structured response schema, and let PromptKit repair structural validator
output within its bounded budget. A contract-invalid final validator response
is a validator execution failure, not a semantic rejection and not a reason
to recursively validate the validator.
- Evaluate the prompt and decision policy against a small human-reviewed set
containing combat setup, active turns, interruptions, multi-phase encounters,
brief rules discussion, aftermath, recalled combat, and false-positive
hostile dialogue. Measure false acceptance, false rejection, retry success,
added calls, latency, and token cost before placing it in the production
default chain.
- An ADR is not required if classification remains owned by
`dnd/scene-descriptions` and the validator follows the generic validation ADR.
Create or supersede an ADR if the work transfers scene classification into
the chunker or otherwise changes stage ownership or the durable chunk-plan
contract.
## Near-Term D&D Pipeline ## Near-Term D&D Pipeline
### Evaluate Spell Extraction And Normalization ### Evaluate Spell Extraction And Normalization
- Evaluate ordinary extraction retries and the completed normalization path - Evaluate ordinary extraction retries and the completed normalization path
against a human-reviewed transcript set before and after adopting the shared against a human-reviewed transcript set before adding repair-aware retries or
PromptKit repair and Notarius validation-retry policies above. an LLM-backed semantic validator.
- Maintain a small set of human-reviewed transcripts and outputs for prompt, - Maintain a small set of human-reviewed transcripts and outputs for prompt,
validator, and normalizer development. Treat model-quality review as an validator, and normalizer development. Treat model-quality review as an
iterative human evaluation aid, not a deterministic correctness gate. iterative human evaluation aid, not a deterministic correctness gate.
@@ -75,51 +24,26 @@ to domain review and empirical evaluation.
## Shared Normalization And Quality Work ## Shared Normalization And Quality Work
The implemented source-backed core and initial D&D registry adoption are ### Generic LLM-Assisted Deduplication
described by [Module Internals](../internal/modules.md#semantic-reconciliation)
and
[D&D Module Internals](../internal/dnd.md#semantic-registry-reconciliation).
The sections below keep broader extensions deferred.
### Large-Collection Semantic Reconciliation - Add a reusable normalizer that asks an LLM to identify duplicate sets in a
list and propose one replacement element for each set.
- Define the minimum domain-neutral input contract, initially an ordered list
whose elements have stable unique IDs. Artifact-kind registrations or
adapters may expose that structure without moving domain rules into the
generic package.
- Keep mutation deterministic: parse and validate the model's duplicate groups,
require every referenced ID to exist, reject overlapping or malformed groups,
prevent unrelated insertion or deletion, and apply only approved replacement
operations in code.
- Preserve provenance needed for audit and downstream validation, and emit
warnings describing every collapsed group.
- Evaluate batching and context-window limits before applying the normalizer to
large artifact collections.
- Evaluate deterministic candidate blocking only after representative registry The model may use its own domain knowledge to judge semantic duplication; the
inputs exceed the active roadmap's bounded single-request limits. Blocking generic implementation is responsible only for the common proposal contract,
should use cheap, explainable signals to form plausible comparison sets while safety checks, and deterministic application of accepted changes.
preserving the possibility that a duplicate appears outside a lexical name
match.
- Define correctness for candidates that appear in more than one block,
conflicting canonical selections, transitive identity across blocks, retry
isolation, and deterministic final ordering before implementation.
- Prefer a reconciliation graph or union plan with explicit conflict checks
over arbitrary fixed-size slices. Never silently treat a batch boundary as
evidence that two candidates are distinct.
- Record per-request bounds, block provenance, model calls, discarded
proposals, and final group derivation well enough to audit a collapse.
### Operator-Selected Semantic Policies
- Consider allowing an operator to select an approved semantic-policy prompt
for a typed reconciliation module without replacing the shared protocol,
response schema, or deterministic safety rules.
- Define the trusted asset source, configuration syntax, compatibility checks,
startup validation, provenance, prompt fingerprinting, checkpoint effects,
and support boundary before exposing the option.
- Prefer selection among registered, typed-policy-compatible prompt assets over
arbitrary filesystem prompt paths. Do not add this flexibility until an
operator workflow requires it; artifact-family-owned policy remains simpler
and safer for the initial implementation.
### Broader Reconciliation Inputs And Module Selection
- Revisit alternate context providers when a concrete non-source-backed entity
collection needs semantic reconciliation. Any extension must preserve the
same request-local identity, deterministic proposal validation, provenance,
and typed application guarantees.
- Consider a selectable generic normalizer only if Notarius gains a real
domain-neutral typed artifact contract that can safely support it. Do not
weaken exact artifact registration or introduce reflection-based arbitrary
JSON mutation merely to expose a universal module key.
### Validation And Review ### Validation And Review
@@ -176,18 +100,6 @@ checkpoint reuse, when an older artifact may be decoded or adapted, and when a
producer or all dependents must be recomputed. Do not add a general migration producer or all dependents must be recomputed. Do not add a general migration
framework until an actual contract change requires one. framework until an actual contract change requires one.
### Artifact-family-oriented physical packaging
[ADR-0004](../adr/0004-package-modules-by-domain.md) currently groups production
extensions by domain and then by pipeline stage. After artifact-family
ownership terminology is established and more families span extraction,
normalization, validation, codecs, references, and assets, reassess whether a
feature-first physical layout would improve navigation and reduce scattered
changes enough to justify a repository-wide package migration. Any change must
address Go dependency cycles, registrar ownership, stable public module keys,
and supersession of the affected ADR-0004 decision. Conceptual artifact-family
ownership does not by itself require this move.
## Blue-Sky Platform And Operations ## Blue-Sky Platform And Operations
These ideas are intentionally less specified. Promote one into an earlier These ideas are intentionally less specified. Promote one into an earlier
@@ -203,6 +115,7 @@ section only after a concrete workflow, contract, and priority emerge.
### Distribution And Operations ### Distribution And Operations
- Packaged release artifacts for alpha distribution. - Packaged release artifacts for alpha distribution.
- A documented versioning and release process.
- Optional generated example-output fixtures with a regeneration procedure. - Optional generated example-output fixtures with a regeneration procedure.
- Additional diagnostics or reporting views. - Additional diagnostics or reporting views.

View File

@@ -0,0 +1,486 @@
# Published Evidence Context Implementation Plan
## Status
Completed.
## Objective
Implement the accepted [Published Evidence Context](evidence.md) roadmap as an
optional, deterministic extension of the production JSON output. The completed
work must publish one deduplicated source-context artifact for an explicit set
of successful normalized lanes while leaving existing lane payloads,
normalization, checkpointing, subprocess results, and disabled output bundles
unchanged.
Complete the stages below in order. Each stage must leave its affected packages
passing before the next begins. Do not implement per-record hydration, implicit
all-lane collection, a full-transcript mode, LLM processing, or any other
roadmap item marked out of scope.
## Decisions And Invariants
- Evidence publication is output policy. Normalizers continue to return
semantic artifacts with precise source references and do not receive
hydration responsibilities.
- The framework operates on the accepted generic `source.SourceDocument` and
typed artifact projections. It must not inspect serialized JSON for
`source_ref` or `source_refs`, and generic packages must not depend on D&D
types.
- The selected lane allowlist is explicit, non-empty, and globally addressed by
resolved lane ID. Pipeline resolution already guarantees lane IDs are unique
across ordered steps.
- The framework decodes accepted serialized normalize outputs through their
registered artifact codecs before invoking typed evidence projectors. This
supports both fresh and checkpoint-reused normalize outputs without retaining
a second typed result channel.
- Expansion uses source-document positions. Numeric unit IDs are identities,
not sequence numbers.
- Direct source references are never widened or rewritten. Expanded ranges are
context bounds only.
- Rejected, failed, and absent normalized lane outputs contribute no evidence.
- Evidence output is sensitive durable source content, not cache or debug
state. It contains accepted source units only and never raw input bytes,
prompts, model responses, auxiliary references, paths, or credentials.
- The optional `evidence_context` index field is an additive v1 JSON-bundle
change. Existing D&D artifact schemas and the subprocess receipt do not
change.
## Stage 1: Add Typed Evidence Capability And Resolve Output Policy
Add a dedicated artifact-evidence registry under the pipeline framework:
- `pipeline.ArtifactEvidenceRegistry` stores one typed projector per artifact
kind.
- `pipeline.ArtifactEvidenceProjector[T]` is
`func(T) []source.SourceRef`.
- `pipeline.RegisterArtifactEvidence[T](registry, kind, projector)` accepts a
non-empty kind and non-nil projector, records the exact Go type for `T`, and
rejects duplicate kinds.
- The erased projection boundary checks the exact registered type, invokes the
projector, and returns a defensive copy of its references.
- The registry exposes only the discovery and projection operations required by
resolution, preparation, and execution; do not expose its mutable entries.
Add the registry to `pipeline.Registries` and `pipeline.ModuleCatalog`, including
CLI catalog conversion, production construction, empty-set detection, and
test registry helpers. A nil evidence registry remains valid when evidence
publication is disabled. Production construction and the D&D registrar require
and populate it.
Define these framework-level output-policy contracts:
```go
type EvidenceContextPolicy struct {
Enabled bool
WindowUnits int
LaneIDs []string
}
type EvidenceContextPolicyProvider interface {
EvidenceContextPolicy() EvidenceContextPolicy
}
```
Provider and prepared-pipeline boundaries defensively copy `LaneIDs`.
Add an optional output-profile option-validation callback to
`OutputEncoderRegistry`:
```go
type OutputProfileOptionContext struct {
LaneIDs []string
}
type OutputProfileOptionValidator func(
OutputProfileOptionContext,
map[string]any,
) error
```
Add `RegisterBuilderWithProfileValidation(spec, validateOptions,
validateProfile, builder)` and make existing output registration methods
delegate to it with no profile callback. The registry passes defensive copies
to validation. The callback receives the complete configured lane-ID set before
invocation-level `--only` filtering. The production JSON output uses it only to
prove that every configured evidence lane exists. Keep the extension generic:
the pipeline supplies lane identities, while the output module interprets its
own options. Build that lane set from the normalized legacy-or-steps profile
before selection and reject duplicate configured lane IDs through the existing
pipeline identity rules.
Extend the production JSON output options with the nested
`evidence_context` object:
- omission disables the feature;
- `enabled` is required when the object is present;
- `enabled: false` permits no `lanes` or `window_units` fields;
- `enabled: true` requires a non-empty `lanes` array;
- lane values are strings, trimmed, non-empty, unique after trimming, and
normalized to lexical order;
- `window_units` is an optional non-negative integer with default `3`; and
- outer and nested unknown fields and incompatible YAML value types remain
strict configuration errors.
The JSON encoder implements the policy provider from its decoded immutable
options. Pipeline resolution invokes its profile validator against all
configured steps, so an unknown evidence lane fails even when another lane is
selected with `--only`.
During `pipeline.Prepare`, after constructing the output encoder:
1. obtain and defensively normalize an enabled policy;
2. intersect its configured IDs with the effective prepared lanes, treating
allowlisted lanes removed by invocation-level filtering as inactive;
3. require an artifact-evidence registration for each active lane kind;
4. prove that its projector Go type exactly matches the active lane's registered
artifact codec type; and
5. retain an immutable private evidence plan on `PreparedPipeline`.
Duplicate or empty provider values, a missing evidence registry for an active
lane, unsupported active artifact kinds, and type mismatches fail preparation
with pipeline/output/lane context. A disabled or non-participating output
encoder creates no evidence plan and preserves existing preparation behavior.
The private plan retains both the full configured allowlist for publication and
the active lane/projector intersection for execution.
Register D&D evidence projectors for all six current artifact kinds. Each
projector returns copies of the artifact's direct references in record order:
spells, NPCs, combat turns, item events, NPC interactions, and the singular
reference from each scene description. Scene descriptions gain capability but
remain excluded unless their configured lane ID is allowlisted.
Stage tests:
- Registry tests cover nil, blank, duplicate, exact-type, defensive-copy, and
deterministic discovery behavior.
- JSON option tests cover disabled, enabled/default-window, explicit zero
window, normalization, duplicates, unknown fields, and invalid types.
- Preparation tests cover selected lanes across steps, unknown lanes,
unsupported kinds, projector/codec type mismatch, disabled behavior, and
defensive policy ownership.
- Resolution/preparation tests prove a valid allowlist survives `--only`, an
excluded lane contributes no active projector, and a genuinely unknown
configured lane still fails profile resolution.
- D&D registration tests prove every production D&D artifact kind has the
expected evidence capability without testing individual field loops
redundantly.
- One table-driven D&D projector test supplies representative values for all
six artifact kinds and proves plural and singular references are copied
without aliasing or semantic rewriting.
Stage completion:
- `go test ./internal/framework/pipeline`
- `go test ./internal/modules/generic/output/json`
- `go test ./internal/modules/dnd/register`
- `go test ./internal/cli`
## Stage 2: Define And Build The Evidence-Context Artifact
Add a domain-neutral `internal/framework/evidencecontext` package that owns the
durable model, JSON Schema, strict codec, projection algorithm, and these exact
identities:
- artifact kind `source/evidence-context`;
- media type `application/json`;
- schema ID `notarius.source.evidence_context`;
- schema name `notarius_source_evidence_context_v1`; and
- schema version `v1`.
Use these package-level model and build contracts:
```go
type Document struct {
SourceID string
SourceDigest string
WindowUnits int
SelectedLanes []string
Contexts []Context
}
type Context struct {
ContextRef source.SourceRef
EvidenceRefs []EvidenceRef
Units []source.SourceUnit
}
type EvidenceRef struct {
LaneID string
SourceRef source.SourceRef
}
type LaneEvidence struct {
LaneID string
SourceRefs []source.SourceRef
}
type BuildRequest struct {
Source *source.SourceDocument
WindowUnits int
SelectedLanes []string
LaneEvidence []LaneEvidence
}
```
Apply the JSON field names shown below. Provide `Build(BuildRequest)`,
`Serialize(BuildRequest)`, and a `Codec` with the same identity/encode/decode
responsibilities as the chunk-map codec.
The v1 payload has this exact shape:
```json
{
"source_id": "session-alpha",
"source_digest": "sha256:0000000000000000000000000000000000000000000000000000000000000000",
"window_units": 0,
"selected_lanes": ["npcs"],
"contexts": [
{
"context_ref": {
"source_id": "session-alpha",
"start_unit_id": 13,
"end_unit_id": 13
},
"evidence_refs": [
{
"lane_id": "npcs",
"source_ref": {
"source_id": "session-alpha",
"start_unit_id": 13,
"end_unit_id": 13
}
}
],
"units": [
{
"id": 13,
"kind": "transcript_segment",
"text": "The party meets Rowan.",
"ref": {
"source_id": "session-alpha",
"start_unit_id": 13,
"end_unit_id": 13
}
}
]
}
]
}
```
All displayed fields are required. `selected_lanes`, `contexts`,
`evidence_refs`, and `units` encode as arrays rather than `null`; `contexts`
may be empty. Each unit uses the existing `source.SourceUnit` JSON shape with
required `id`, `kind`, `text`, and `ref`, plus optional JSON-shaped `metadata`.
Fixed objects reject unknown fields; metadata remains an arbitrary JSON object.
The builder accepts the validated source document, selected lane IDs, effective
window, and lane-attributed direct references, then:
1. requires a non-negative window and a non-empty, trimmed, unique selected
lane set, then stores that set in lexical order;
2. requires every `LaneEvidence.LaneID` to belong to the selected set;
3. validates the source document, recomputes its semantic digest, and requires
it to equal `SourceDocument.Digest`;
4. validates every direct reference against one `source.DocumentIndex`;
5. deduplicates exact `(lane_id, source_ref)` contributions;
6. resolves endpoints to document positions;
7. expands each side without integer overflow and clips at document bounds;
8. sorts by expanded document position with deterministic lane/reference
tie-breakers;
9. merges overlapping or position-contiguous expanded intervals;
10. unions and deterministically sorts each merged context's direct
contributions; and
11. deep-clones the corresponding source units and JSON-shaped metadata.
Contexts are disjoint and ordered by document position, so a source unit occurs
at most once in the document. `context_ref` identifies the first and last
included units; `evidence_refs` retains only original citations. An empty
reference collection produces the same source identity, window, sorted
allowlist, and an explicit empty contexts array.
Projection failures identify only structural scope such as lane and reference
position. They must not include source text, metadata values, raw serialized
artifacts, or unrelated paths.
The codec must validate its model before encoding, produce deterministic JSON,
strictly decode the checked-in schema, and return independently owned values.
The JSON output encoder remains responsible for pretty-printing the logical
file with its standard trailing newline. Follow the existing chunk-map
package's separation between model, builder, codec, schema asset, and contract
tests where useful, without coupling the two artifact formats.
Stage tests:
- A table-driven builder suite covers zero and nonzero windows, boundary
clipping, non-monotonic unit IDs, separate gaps, overlapping and contiguous
windows, duplicate contributions, multiple lanes, stable ordering, empty
contexts, invalid selected/contributing lanes, source-digest mismatch, and
invalid references.
- Ownership tests prove output mutation cannot affect the source document or
projector inputs, including nested metadata.
- Codec tests cover round trip, required arrays, schema identity, malformed and
trailing JSON, unknown fixed fields, invalid ordering/ranges, mismatched
source identities, and independently owned decoded metadata.
- Use structured assertions and a compact valid fixture; do not add a large
transcript golden file.
Stage completion:
- `go test ./internal/framework/evidencecontext`
## Stage 3: Integrate Projection With Runner Output
Extend `contracts.OutputRequest` with an optional
`EvidenceContext *SerializedArtifact` field and clone it at every ownership
handoff, following the existing chunk-map pointer pattern.
After lane execution and final manifest population, but before invoking the
output encoder, the runner must:
1. skip all work when the prepared evidence plan is absent;
2. index accepted `NormalizeOutputs` by their globally unique lane IDs and fail
on an internal duplicate rather than silently overwrite it;
3. for each selected lane with an output, verify its source and artifact kind,
decode it through the prepared artifact codec registry, and invoke the
prepared typed projector;
4. build and serialize the evidence document through the evidence-context
package; and
5. pass a defensive serialized-artifact copy to the output encoder.
Selected lanes without normalized output contribute nothing. Normalize
rejections remain successful pipeline outcomes; evidence projection does not
inspect rejected candidates. An invalid accepted reference, incompatible
serialized artifact, projection type failure, or evidence serialization failure
is an output-stage framework error before logical files are returned or
physically published.
At this external-content consumption boundary, do not propagate artifact-codec
or metadata-cloning errors with `%w` when their text could contain artifact
fields or source metadata. Return fixed, actionable categories scoped by lane
and operation; detailed codec errors remain available to direct trusted
callers and their focused tests.
Add an allowlisted debug summary containing only evidence artifact identity,
selected lanes, window, context count, unit count, and source digest. Do not
duplicate transcript text or source-unit metadata into a new evidence-specific
debug envelope. Existing normalized-output debug behavior remains unchanged.
The output artifact is not a normalized lane, generated reference, checkpoint,
or manifest normalized-output entry. It does not alter normalized-output,
rejection, or warning counts. Resume continues to reuse normalized checkpoints;
evidence is deterministically rebuilt during the always-executed output stage.
Stage tests:
- Runner tests use a real codec, evidence projector, and small capturing output
encoder to prove selected-lane union, absent/rejected lane omission, invalid
accepted-reference failure, output-request defensive ownership, and no work
when disabled.
- Include one checkpoint-reused normalized-output case to prove evidence is
reconstructed identically without retaining typed normalize values.
- Confirm projection failures prevent output encoding and return a failed
manifest without changing rejection semantics.
Stage completion:
- `go test ./internal/framework/pipeline`
## Stage 4: Publish Through The JSON Bundle
Teach the production JSON encoder to recognize the optional evidence-context
artifact, verify its exact kind, media type, schema identity, schema digest, and
payload validity through the evidence-context codec, and emit:
- logical file `evidence-context.json`; and
- optional `index.json` descriptor field `evidence_context`.
The descriptor uses the same six fields as `chunk_map`:
`artifact_kind`, `file`, `media_type`, `schema_id`, `schema_name`, and
`schema_version`. Refactor the encoder's private descriptor representation only
as needed to share that shape; do not change the existing `chunk_map` wire
contract. Evidence output is ordered with the encoder's other fixed logical
files, remains a non-lane artifact, and is present with an empty contexts array
when enabled but no selected lane produces references.
The JSON encoder's validation boundary returns a fixed content-safe evidence
artifact error rather than propagating decoder or schema diagnostics that could
echo transcript text or metadata. Direct evidence-context codec tests retain
detailed structural errors.
Update the maintained complete D&D configuration to enable evidence context
with window `3` for `item-events`, `npcs`, `spells`, `combat-turns`, and
`npc-interactions`. Deliberately omit `scene-descriptions`. Keep the minimal
configuration disabled by omission.
Stage tests:
- JSON encoder tests own descriptor shape, exact logical filename, identity
checking, empty evidence publication, and disabled bundle stability.
- One assembled production D&D test uses multiple selected lanes with
overlapping references and non-monotonic unit IDs, decodes the published
artifact through its production codec, and proves union/deduplication and
scene-description exclusion.
- A second narrow case explicitly allowlists a scene-description lane to prove
capability is opt-in rather than hard-coded exclusion.
- Existing index, chunk-map, lane, manifest, warning, and rejection tests remain
the owners of their current formats; do not repeat their full matrices.
Stage completion:
- `go test ./internal/modules/generic/output/json`
- `go test ./internal/modules/dnd/...`
- `go test ./internal/modules/integration`
- `go test ./internal/cli`
## Stage 5: Publish Current-Behavior Documentation
After implementation and behavioral tests pass, update canonical documentation:
- `docs/config.md` owns the nested JSON output options, defaults, strict
validation, required lane allowlist, and a small configuration snippet.
- A new `docs/integrations/evidence-context.md` owns the complete v1 payload,
identities, direct-evidence versus context semantics, ordering,
compatibility, and a compact valid example.
- `docs/integrations/json-output.md` owns the optional logical file and
`index.json` descriptor; link to the evidence contract rather than repeating
its payload.
- `docs/operations.md` owns durable source-content sensitivity, permissions,
retention, and the possibility that selected lanes cover most of a
transcript.
- `docs/consumers/subprocess.md` explains discovery through the optional index
descriptor and requires consumers to treat `evidence_refs`, not expanded
context bounds, as citations.
- `docs/policy/architecture.md` records the generic typed evidence-projection
boundary and output ownership without adding D&D or wire-format detail.
- `docs/internal/pipeline.md` and `docs/internal/modules.md` describe the typed
evidence registry, preparation checks, reconstruction from serialized
normalize outputs, and output-stage ownership without restating public wire
fields.
Update only the smallest orientation links needed for discoverability. Do not
add a CLI flag, configuration environment override, or duplicate the complete
configuration outside `examples/`.
After all current-behavior documentation is accurate:
- set [the feature roadmap](evidence.md) status to `Implemented`;
- set this plan's status to `Completed`; and
- leave the integration and configuration documents, not either roadmap, as
the canonical implemented contract.
Final verification:
- `git diff --check`
- `go test ./...`
- `go vet ./...`
- `go build ./cmd/notarius`
- `go test -race ./internal/framework/evidencecontext ./internal/framework/pipeline ./internal/modules/generic/output/json ./internal/modules/dnd/... ./internal/modules/integration ./internal/cli`
## Open Questions
None. The plan fixes the configuration shape and defaults, typed projection
boundary, preparation timing, durable schema and identities, range-union
algorithm, failure semantics, JSON discovery, D&D coverage, documentation
ownership, and test boundaries.

208
docs/roadmap/subprocess.md Normal file
View File

@@ -0,0 +1,208 @@
# Subprocess Integration Contract
## Status
Implemented.
## Purpose
Make Notarius straightforward to invoke as a subprocess from an orchestrator
such as Narratio. A caller should be able to run a configured pipeline, discover
the published output bundle without parsing human prose or scanning a
directory, and hand selected structured artifacts to a later stage.
This work strengthens the public CLI boundary. It does not turn Notarius into a
Go library, embed Narratio-specific behavior, or change pipeline execution and
artifact semantics.
## Desired End State
A subprocess caller can:
1. validate a Notarius configuration and selected pipeline before execution;
2. invoke `notarius run` with explicit input, output-root, session, and
reference arguments;
3. request one versioned, machine-readable success result on standard output;
4. use that result to locate the published output bundle;
5. discover normalized lane payloads through the bundle's authoritative
`index.json`;
6. distinguish process failure from successful partial pipeline outcomes; and
7. record Notarius run provenance in its own manifest without depending on
internal packages, cache formats, debug formats, or human-readable messages.
The existing human-oriented command output remains the default for interactive
use.
## Machine-Readable Run Result
`notarius run` supports `--json`. On success, the flag makes standard output
contain exactly one JSON object followed by a newline. No human-oriented status
line is mixed into that stream.
The result uses the schema identity `notarius.run-result.v1` and contains:
| Field | Presence | Meaning |
| --- | --- | --- |
| `schema_version` | Required | Exactly `notarius.run-result.v1`. |
| `run_id` | Required | The Notarius run identifier. |
| `pipeline_id` | Required | The effective pipeline identifier. |
| `output_directory` | Required | Absolute path to the successfully published output bundle. |
| `index_file` | Required for the production JSON output | Logical bundle path `index.json`. |
| `normalized_output_count` | Required | Number of final normalized lane outputs returned by the pipeline. |
| `rejected_output_count` | Required | Number of recorded rejected outputs. |
| `warning_count` | Required | Number of final run warnings returned by the pipeline. |
| `validation_status` | Required | The run manifest's final validation status without reinterpretation. |
| `debug_directory` | Optional | Absolute debug-bundle path when debug capture was requested and completed. |
An illustrative successful result is:
```json
{
"schema_version": "notarius.run-result.v1",
"run_id": "run-1770000000000000000-0123456789abcdef0123456789abcdef",
"pipeline_id": "dnd-session",
"output_directory": "/srv/narratio/runs/session-7/notarius/run-1770000000000000000-0123456789abcdef0123456789abcdef",
"index_file": "index.json",
"normalized_output_count": 6,
"rejected_output_count": 2,
"warning_count": 1,
"validation_status": "approved"
}
```
The receipt is a discovery and summary document, not a duplicate output
envelope. It does not embed lane payloads, rejection entries, warnings, the run
manifest, or output-file contents. Consumers use `index_file` and the existing
published JSON output contract for those records.
The result contract must tolerate future additive optional fields. Any
incompatible field or semantic change requires a new run-result schema version.
## Stream, Publication, And Failure Semantics
Machine-readable output is emitted only after:
- the pipeline has completed without a framework error;
- all logical output files have been successfully published;
- requested debug terminal reporting has completed; and
- all result fields are known.
Writing or encoding the machine-readable result is part of successful command
completion. Failure to write it produces the existing runtime-failure exit
class.
With `--json`:
- successful stdout is exclusively the run-result JSON document;
- successful warnings remain on stderr under the existing CLI contract;
- syntax and runtime errors retain their existing exit statuses and stderr
diagnostics;
- consumers treat stdout as a valid result only when the process exits with
status 0; failures before result writing emit no result, while a failure
during the stdout write may leave incomplete bytes that must be ignored; and
- human-readable diagnostic wording is not promoted into a machine contract.
Without `--json`, current interactive stdout and stderr behavior remains
unchanged.
Successful runs may contain rejected outputs or omit some normalized lanes.
That remains a valid pipeline outcome. The run result reports counts, while
`index.json`, `rejected.json`, and `warnings.json` remain authoritative for
details. Notarius will not add a generic `--fail-on-rejection` policy as part
of this work.
## Output Discovery And Consumer Responsibilities
The production JSON encoder's `index.json` remains the authoritative mapping
from lane IDs to published payloads. A subprocess consumer should:
- resolve `index_file` beneath `output_directory` and reject path escape;
- locate expected outputs by `lane_id`, not by guessing filenames;
- check each selected descriptor's media type and schema identity;
- decode payloads according to their published integration contracts;
- decide which lanes are required or optional for its own later stages; and
- retain rejection, warning, and manifest files when they are needed for
provenance or review.
For Narratio, required report inputs and partial-success policy remain Narratio
stage configuration and orchestration concerns. Notarius does not acquire
knowledge of Narratio stages, manifests, workspace layout, publication policy,
or report formats.
## Invocation Guidance
The consumer documentation recommends that subprocess callers:
- use `notarius config validate --pipeline` as an optional preflight;
- pass explicit absolute paths for the input, configuration, output root, and
CLI-supplied references;
- use a stable, non-secret prompt session identifier when useful for provider
routing or caching;
- capture stdout and stderr separately;
- supply credentials through the configured environment mechanism rather than
command arguments or generated configuration containing secret values;
- place output, cache, debug, and subprocess logs under intentional
sensitivity and retention policies; and
- treat the Notarius manifest and run-result receipt as provenance while
leaving the caller's own manifest authoritative for its stage lifecycle.
Notarius configuration remains owned by Notarius. An orchestrator may select a
configuration and pass supported operational overrides, but should not
duplicate the complete Notarius configuration schema.
## Documentation End State
- `docs/cli.md` owns `run --json`, stream behavior, and exit semantics;
- a new `docs/integrations/run-result.md` owns the versioned run-result wire
contract and compatibility policy;
- `docs/integrations/json-output.md` remains the sole owner of output-bundle
discovery and lane publication;
- a new `docs/consumers/subprocess.md` provides the task-oriented invocation and
consumption workflow; and
- `docs/internal/cli.md` describes how the CLI constructs and emits the result
only after successful publication.
Other documents should link to these owners instead of repeating volatile
fields or command details.
## Acceptance Criteria
- An ordinary successful `run` retains its existing human-readable output.
- A successful `run --json` emits one valid `notarius.run-result.v1` document
and no human prose on stdout.
- Relative configured or overridden output and debug roots are reported as
absolute bundle paths.
- The receipt identifies the production JSON bundle entry point without
copying its lane descriptors or payloads.
- Warning-bearing and rejection-bearing runs remain successful and report
accurate counts.
- Syntax, configuration, provider, pipeline, publication, debug, and result
writing failures retain the correct nonzero exit class. Consumers are
explicitly required to ignore stdout from a nonzero invocation.
- The implementation does not expose internal Go types or couple generic CLI
code to D&D or Narratio concepts.
- Public and internal documentation assigns each new contract to one canonical
owner.
- Offline behavioral tests protect the structured-output contract, default
human behavior, absolute path reporting, stream separation, and failure to
serialize or write the success result without duplicating lower-level output
encoder tests.
## Out Of Scope
The following may be useful later but are not prerequisites for the Narratio
integration:
- a result-file flag in addition to machine-readable stdout;
- a JSON failure envelope or stable machine-readable error taxonomy;
- caller-supplied Notarius run IDs or exact output-bundle paths;
- a generic `--fail-on-rejection` or required-lane CLI policy;
- a public Go client package or importable Narratio adapter;
- Narratio stage, configuration, manifest, or report-generation changes;
- `notarius version --json`;
- installable or queryable artifact JSON Schemas;
- signal-aware CLI contexts and graceful SIGINT or SIGTERM handling;
- packaged release artifacts and a broader application-versioning policy.
These items should be promoted only in response to a demonstrated integration
need rather than bundled into the initial subprocess contract.

View File

@@ -1,6 +1,4 @@
version: 4 version: 3
promptkit:
profile_file: ./examples/profiles/dnd-extraction.yml
concurrency: concurrency:
total_llm: 2 total_llm: 2
stage_workers: stage_workers:
@@ -18,7 +16,6 @@ debug:
directory: ./notarius-debug directory: ./notarius-debug
pipelines: pipelines:
dnd-session: dnd-session:
llm_profile: dnd-extraction
input: seriatim input: seriatim
# Stable campaign context is shared by every module that accepts these slots. # Stable campaign context is shared by every module that accepts these slots.
references: references:
@@ -35,40 +32,29 @@ pipelines:
enabled: true enabled: true
window_units: 3 window_units: 3
lanes: lanes:
- item-occurrences - item-events
- item-registry - npcs
- location-registry
- location-occurrences
- npc-registry
- spells - spells
- combat-turns - combat-turns
- npc-occurrences - npc-interactions
- enemy-events
steps: steps:
# Establish session-wide reference artifacts before their consumers. # Establish session-wide reference artifacts alongside independent item events.
- id: describe-session - id: describe-session
artifacts: artifacts:
item-registry: item-events:
extract: extract:
module: dnd/item-registry module: dnd/item-events
retries: 2 retries: 2
merge: appendorder merge: appendorder
normalize: dnd/item-registry normalize: dnd/item-events
npc-registry: npcs:
extract: extract:
module: dnd/npc-registry module: dnd/npcs
retries: 2 retries: 2
merge: appendorder merge: appendorder
normalize: normalize:
module: dnd/npc-registry module: dnd/npcs
retries: 2 llm_profile: gemini-2-flash
location-registry:
extract:
module: dnd/location-registry
retries: 2
merge: appendorder
normalize:
module: dnd/location-registry
retries: 2 retries: 2
scene-descriptions: scene-descriptions:
extract: extract:
@@ -77,32 +63,18 @@ pipelines:
merge: appendorder merge: appendorder
normalize: dnd/scene-descriptions normalize: dnd/scene-descriptions
- id: extract-events - id: extract-events
# Accepted registry artifacts and scene-description eligibility artifacts # Accepted NPC grounding and scene-description eligibility artifacts are
# are supplied in memory to their compatible consumers in this step. # supplied in memory to their compatible consumers in this step.
references: references:
location_registry: npcs:
artifact: artifact:
step: describe-session step: describe-session
lane: location-registry lane: npcs
npc_registry:
artifact:
step: describe-session
lane: npc-registry
scene_descriptions: scene_descriptions:
artifact: artifact:
step: describe-session step: describe-session
lane: scene-descriptions lane: scene-descriptions
item_registry:
artifact:
step: describe-session
lane: item-registry
artifacts: artifacts:
item-occurrences:
extract:
module: dnd/item-occurrences
retries: 2
merge: appendorder
normalize: dnd/item-occurrences
spells: spells:
extract: extract:
module: dnd/spells module: dnd/spells
@@ -121,40 +93,9 @@ pipelines:
retries: 2 retries: 2
merge: appendorder merge: appendorder
normalize: dnd/combat-turns normalize: dnd/combat-turns
npc-occurrences: npc-interactions:
extract: extract:
module: dnd/npc-occurrences module: dnd/npc-interactions
retries: 2 retries: 2
merge: appendorder merge: appendorder
normalize: dnd/npc-occurrences normalize: dnd/npc-interactions
location-occurrences:
extract:
module: dnd/location-occurrences
retries: 2
merge: appendorder
normalize: dnd/location-occurrences
- id: track-enemies
references:
npc_registry:
artifact:
step: describe-session
lane: npc-registry
scene_descriptions:
artifact:
step: describe-session
lane: scene-descriptions
combat_turns:
artifact:
step: extract-events
lane: combat-turns
npc_occurrences:
artifact:
step: extract-events
lane: npc-occurrences
artifacts:
enemy-events:
extract:
module: dnd/enemy-events
retries: 2
merge: appendorder
normalize: dnd/enemy-events

View File

@@ -1,7 +1,6 @@
version: 4 version: 3
pipelines: pipelines:
dnd-session: dnd-session:
llm_profile: dnd-extraction
input: seriatim input: seriatim
artifacts: artifacts:
spells: spells:

View File

@@ -1,5 +0,0 @@
id: dnd-extraction
backend: openrouter
model: openai/gpt-5.6-luna
timeout_seconds: 240
service_tier: flex

7
go.mod
View File

@@ -3,14 +3,9 @@ module gitea.maximumdirect.net/eric/notarius
go 1.25.5 go 1.25.5
require ( require (
gitea.maximumdirect.net/eric/promptkit v0.9.0 gitea.maximumdirect.net/eric/scriptorium v0.11.1
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 github.com/santhosh-tekuri/jsonschema/v6 v6.0.2
gopkg.in/yaml.v3 v3.0.1 gopkg.in/yaml.v3 v3.0.1
) )
require golang.org/x/text v0.40.0 require golang.org/x/text v0.40.0
require (
gitea.maximumdirect.net/eric/promptkit-backend-openrouter v1.0.0 // indirect
gitea.maximumdirect.net/eric/promptkit-backend-rakestrawhome v1.0.0 // indirect
)

12
go.sum
View File

@@ -1,13 +1,13 @@
gitea.maximumdirect.net/eric/promptkit v0.9.0 h1:IpvDRC8L6xRxQ9hpuyKOmMc5b6MeLTKYyx+h1YAjy08= gitea.maximumdirect.net/eric/scriptorium v0.11.0 h1:rjvbt9FTaWHxYlHq7QlUzmMVUt3QdbTmeCkmH81N//o=
gitea.maximumdirect.net/eric/promptkit v0.9.0/go.mod h1:oMJ/WUJImUtwJ5e+6MAGECPYAErAkOaKel0G+3T/b4E= gitea.maximumdirect.net/eric/scriptorium v0.11.0/go.mod h1:FQ5lEuNxmrQyNgIomkpZdxvfTC0jWjbXYuq3tbJWF64=
gitea.maximumdirect.net/eric/promptkit-backend-openrouter v1.0.0 h1:lc062euk2qseO//D762i3JaFyulDNML3eQQX7DkYTho= gitea.maximumdirect.net/eric/scriptorium v0.11.1 h1:zBKtB3+fP8FcHGI8DJD99CiTL6crAGitBhWtE+xYJHc=
gitea.maximumdirect.net/eric/promptkit-backend-openrouter v1.0.0/go.mod h1:AIa7kAu2mfrRQgcspe4L+DW51WqgnALQT60lqkEywJI= gitea.maximumdirect.net/eric/scriptorium v0.11.1/go.mod h1:FQ5lEuNxmrQyNgIomkpZdxvfTC0jWjbXYuq3tbJWF64=
gitea.maximumdirect.net/eric/promptkit-backend-rakestrawhome v1.0.0 h1:j9YY7wsTVjzke2kHH4YAzpU0oUpM+x+nXwl1IeS+2eg=
gitea.maximumdirect.net/eric/promptkit-backend-rakestrawhome v1.0.0/go.mod h1:4RNS+LILDg4JbS4Ts9Lwy1C92wauXJIbeQaalps4Koo=
github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI= github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI=
github.com/dlclark/regexp2 v1.11.0/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8= github.com/dlclark/regexp2 v1.11.0/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 h1:KRzFb2m7YtdldCEkzs6KqmJw4nqEVZGK7IN2kJkjTuQ= github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 h1:KRzFb2m7YtdldCEkzs6KqmJw4nqEVZGK7IN2kJkjTuQ=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2/go.mod h1:JXeL+ps8p7/KNMjDQk3TCwPpBy0wYklyWTfbkIzdIFU= github.com/santhosh-tekuri/jsonschema/v6 v6.0.2/go.mod h1:JXeL+ps8p7/KNMjDQk3TCwPpBy0wYklyWTfbkIzdIFU=
golang.org/x/text v0.14.0 h1:ScX5w1eTa3QqT8oi6+ziP7dTV1S2+ALU0bI+0zXKWiQ=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs= golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY= golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM=

View File

@@ -1,36 +0,0 @@
// Package buildinfo resolves the product version embedded in a Notarius build.
package buildinfo
import (
"fmt"
"regexp"
"runtime/debug"
)
var stableVersion = regexp.MustCompile(`^v(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)$`)
// Override is set at link time for controlled builds.
var Override string
// Version returns the release version embedded in the build, or development
// when the build does not carry a stable release tag.
func Version() (string, error) {
buildVersion := ""
if info, ok := debug.ReadBuildInfo(); ok {
buildVersion = info.Main.Version
}
return resolve(Override, buildVersion)
}
func resolve(override, buildVersion string) (string, error) {
if override != "" {
if !stableVersion.MatchString(override) {
return "", fmt.Errorf("build version override is not a stable release tag")
}
return override, nil
}
if stableVersion.MatchString(buildVersion) {
return buildVersion, nil
}
return "development", nil
}

View File

@@ -1,45 +0,0 @@
package buildinfo
import "testing"
func TestResolve(t *testing.T) {
tests := []struct {
name string
override string
buildVersion string
want string
wantErr bool
}{
{name: "stable main module version", buildVersion: "v1.2.3", want: "v1.2.3"},
{name: "zero version", buildVersion: "v0.0.0", want: "v0.0.0"},
{name: "override takes precedence", override: "v2.3.4", buildVersion: "v1.2.3", want: "v2.3.4"},
{name: "invalid override", override: "version", buildVersion: "v1.2.3", wantErr: true},
{name: "override with whitespace", override: " v1.2.3", wantErr: true},
{name: "leading zero major", buildVersion: "v01.2.3", want: "development"},
{name: "leading zero minor", buildVersion: "v1.02.3", want: "development"},
{name: "leading zero patch", buildVersion: "v1.2.03", want: "development"},
{name: "build version with whitespace", buildVersion: "v1.2.3 ", want: "development"},
{name: "prerelease", buildVersion: "v1.2.3-rc.1", want: "development"},
{name: "build suffix", buildVersion: "v1.2.3+build.1", want: "development"},
{name: "pseudo version", buildVersion: "v0.0.0-20260102030405-abcdef123456", want: "development"},
{name: "development build", buildVersion: "(devel)", want: "development"},
{name: "missing build information", want: "development"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, err := resolve(tt.override, tt.buildVersion)
if tt.wantErr {
if err == nil {
t.Fatal("resolve() error = nil, want error")
}
return
}
if err != nil {
t.Fatalf("resolve() error = %v", err)
}
if got != tt.want {
t.Fatalf("resolve() = %q, want %q", got, tt.want)
}
})
}
}

Some files were not shown because too many files have changed in this diff Show More