201 lines
7.1 KiB
Markdown
201 lines
7.1 KiB
Markdown
# JSON Output
|
|
|
|
This document is the durable JSON output file-format contract produced by the
|
|
production JSON encoder and written by the CLI. Selectable output-encoder keys
|
|
are cataloged in
|
|
[Configuration](../config.md#implemented-production-modules).
|
|
|
|
The output module produces the logical bundle described here. The CLI's
|
|
physical placement and lifecycle for that bundle are defined in
|
|
[Operations](../operations.md#output-directory).
|
|
|
|
## Files
|
|
|
|
The encoder writes:
|
|
|
|
- `index.json`
|
|
- `manifest.json`
|
|
- `lanes/<lane-id>.json`, one file per normalized serialized artifact
|
|
- `rejected.json`
|
|
- `warnings.json`
|
|
- `chunk-map.json`, only when the JSON output binding enables
|
|
`include_chunk_map` and the run has an accepted chunk map
|
|
|
|
Files are pretty-printed JSON with a trailing newline when the payload is JSON.
|
|
Logical file paths are relative, slash-separated, and may not contain `..`.
|
|
|
|
## `index.json`
|
|
|
|
Shape:
|
|
|
|
```json
|
|
{
|
|
"manifest_file": "manifest.json",
|
|
"output_files": [
|
|
{
|
|
"lane_id": "spells",
|
|
"media_type": "application/json",
|
|
"file": "lanes/spells.json",
|
|
"module_key": "noop",
|
|
"schema_id": "notarius.dnd.spells",
|
|
"schema_name": "notarius_dnd_spells_v1",
|
|
"schema_version": "v1"
|
|
}
|
|
],
|
|
"rejected_file": "rejected.json",
|
|
"warnings_file": "warnings.json"
|
|
}
|
|
```
|
|
|
|
`output_files` is sorted by lane ID. Output file names are produced by
|
|
sanitizing the lane ID:
|
|
|
|
- characters outside `A-Z`, `a-z`, `0-9`, `.`, `_`, and `-` become `_`;
|
|
- repeated `..` sequences are replaced;
|
|
- leading and trailing `.`, `_`, and `-` are trimmed;
|
|
- empty sanitized names are rejected;
|
|
- two lanes that sanitize to the same output file are rejected.
|
|
|
|
`manifest_file`, `rejected_file`, and `warnings_file` contain the fixed paths
|
|
shown above. Each `output_files` entry requires `lane_id` and `file`. It also
|
|
contains the normalized payload `media_type`, normalizer `module_key`, and
|
|
response `schema_id`, `schema_name`, and `schema_version` when those values are
|
|
available.
|
|
|
|
When present, the top-level optional `chunk_map` descriptor contains exactly
|
|
`artifact_kind`, `file`, `media_type`, `schema_id`, `schema_name`, and
|
|
`schema_version`. It identifies the pipeline-wide `chunk-map.json`; it is not
|
|
a lane output and never appears in `output_files`. The descriptor and file are
|
|
both absent when export is disabled or no chunk plan was accepted. Its payload
|
|
contract is defined by [Accepted Chunk Map](chunk-map.md).
|
|
|
|
## `manifest.json`
|
|
|
|
`manifest.json` contains a run manifest. This abridged example shows its core
|
|
structure:
|
|
|
|
```json
|
|
{
|
|
"run_id": "run-123",
|
|
"pipeline_id": "dnd-session",
|
|
"artifact_lanes": [
|
|
{
|
|
"id": "spells",
|
|
"extractor": "dnd/spells",
|
|
"merger": "appendorder",
|
|
"normalizer": "noop"
|
|
}
|
|
],
|
|
"validation_status": "approved",
|
|
"started_at": "2026-01-01T00:00:00Z",
|
|
"completed_at": "2026-01-01T00:00:01Z"
|
|
}
|
|
```
|
|
|
|
Fields with empty values may be omitted by JSON encoding.
|
|
|
|
The manifest fields are:
|
|
|
|
- `run_id`, `pipeline_id`, and `pipeline_digest`: run and resolved-pipeline
|
|
identity;
|
|
- `input_module`, `chunker`, `extractors`, `merger`, `normalizer`, and
|
|
`output_encoder`: resolved module keys;
|
|
- `chunk_plan`: payload-free provenance for the effective chunk plan. `mode`
|
|
is the effective cache mode; `action` is `reused`, `generated`,
|
|
`refreshed`, or `bypassed` when a plan was materialized. `requested_module`
|
|
is the current pipeline chunker, while `producer_input_module`,
|
|
`producer_module`, `producer_llm_profile`, `producer_references`,
|
|
`producer_metadata`, `source_digest`, `plan_digest`, `plan_schema_version`,
|
|
and `created_at` describe the stored or generated producer when available.
|
|
A cached plan can therefore identify a producer different from the requested
|
|
module. This object never embeds ranges, units, annotations, prompts,
|
|
responses, or reference content;
|
|
- `module_metadata` and `artifact_lanes`: module and per-lane provenance,
|
|
including prompt and response-schema provenance when provided;
|
|
- `validator_chains`: resolved validation points and validators;
|
|
- `source_digests` and `references`: source and reference provenance;
|
|
- `normalized_outputs` and `rejected_outputs`: payload-free result summaries;
|
|
- `llm_profiles`: selected profile IDs and provider or model names when
|
|
available;
|
|
- `metadata`: the effective prompt `session_id`;
|
|
- `validation_status`: `approved` or `rejected`;
|
|
- `started_at` and `completed_at`: UTC run timestamps.
|
|
|
|
`source_digests` contains source document digests only. Bound references are
|
|
recorded separately under `references`, which contains provenance only: target
|
|
stage, lane ID when present, slot name, origin type and URI, digest, media
|
|
type, byte size, and binding source. Reference content is not written to
|
|
durable output.
|
|
|
|
Reference `stage` is `chunk`, `extract`, `merge`, or `normalize`. `lane_id` is
|
|
omitted for chunk references and present for extract, merge, and normalize
|
|
references.
|
|
|
|
`validation_status` is `approved` when no outputs were rejected and `rejected`
|
|
when one or more outputs were rejected.
|
|
|
|
Producer warnings and the current run's chunk-validation warnings remain in
|
|
`warnings.json`. The manifest records only provenance and decision summaries;
|
|
empty producer-only values are omitted for compatibility with existing readers.
|
|
|
|
`validator_chains` records the resolved validator chain for each validation
|
|
point. Entries include stage, lane ID when applicable, module key, and validators
|
|
with key and execution class. Empty chains are recorded with an empty
|
|
`validators` array, including chains resolved from explicit empty config
|
|
overrides.
|
|
|
|
`normalized_outputs` summarizes each normalized lane output without embedding
|
|
payload bytes. Entries include lane ID, normalizer module key, source ID, media
|
|
type, and response schema provenance where available.
|
|
|
|
`rejected_outputs` summarizes rejected module outputs without embedding raw
|
|
payload bytes. Entries include stage, lane, module, chunk, validator or reason,
|
|
message, attempt count, and optional diagnostic artifact path.
|
|
|
|
## Output Payload Files
|
|
|
|
Each normalized serialized artifact is written to
|
|
`lanes/<sanitized-lane-id>.json`. The JSON output encoder is domain-neutral and
|
|
accepts only artifacts whose codec media type is `application/json`. The file
|
|
contains the codec-owned JSON bytes pretty-printed.
|
|
|
|
The schema of each lane payload is owned by that artifact contract. For the
|
|
current D&D lanes, see [D&D Spell Artifact](dnd-spell-artifacts.md),
|
|
[D&D NPC Artifact](dnd-npc-artifacts.md), and
|
|
[D&D Combat-Turn Artifact](dnd-combat-turn-artifacts.md).
|
|
|
|
## `rejected.json`
|
|
|
|
Shape:
|
|
|
|
```json
|
|
{
|
|
"rejected": []
|
|
}
|
|
```
|
|
|
|
When output validation rejects an output, each entry contains `stage` and
|
|
`message`. It includes `lane_id`, `module_key`, `chunk_id`, `chunk_index`,
|
|
`validator_name`, `reason_code`, `attempt_count`, and
|
|
`diagnostic_artifact_path` when applicable.
|
|
|
|
## `warnings.json`
|
|
|
|
Shape:
|
|
|
|
```json
|
|
{
|
|
"warnings": [
|
|
{
|
|
"scope": "extract",
|
|
"reason_code": "example",
|
|
"message": "human-readable warning"
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
`warnings` is an empty array when no warnings are reported.
|
|
Each warning requires `reason_code` and `message`; `scope` is omitted when it is
|
|
empty.
|