Files
notarius/docs/integrations/json-output.md

7.1 KiB

JSON Output

This document is the durable JSON output file-format contract produced by the production JSON encoder and written by the CLI. Selectable output-encoder keys are cataloged in Configuration.

The output module produces the logical bundle described here. The CLI's physical placement and lifecycle for that bundle are defined in Operations.

Files

The encoder writes:

  • index.json
  • manifest.json
  • lanes/<lane-id>.json, one file per normalized serialized artifact
  • rejected.json
  • warnings.json
  • chunk-map.json, only when the JSON output binding enables include_chunk_map and the run has an accepted chunk map

Files are pretty-printed JSON with a trailing newline when the payload is JSON. Logical file paths are relative, slash-separated, and may not contain ...

index.json

Shape:

{
  "manifest_file": "manifest.json",
  "output_files": [
    {
      "lane_id": "spells",
      "media_type": "application/json",
      "file": "lanes/spells.json",
      "module_key": "noop",
      "schema_id": "notarius.dnd.spells",
      "schema_name": "notarius_dnd_spells_v1",
      "schema_version": "v1"
    }
  ],
  "rejected_file": "rejected.json",
  "warnings_file": "warnings.json"
}

output_files is sorted by lane ID. Output file names are produced by sanitizing the lane ID:

  • characters outside A-Z, a-z, 0-9, ., _, and - become _;
  • repeated .. sequences are replaced;
  • leading and trailing ., _, and - are trimmed;
  • empty sanitized names are rejected;
  • two lanes that sanitize to the same output file are rejected.

manifest_file, rejected_file, and warnings_file contain the fixed paths shown above. Each output_files entry requires lane_id and file. It also contains the normalized payload media_type, normalizer module_key, and response schema_id, schema_name, and schema_version when those values are available.

When present, the top-level optional chunk_map descriptor contains exactly artifact_kind, file, media_type, schema_id, schema_name, and schema_version. It identifies the pipeline-wide chunk-map.json; it is not a lane output and never appears in output_files. The descriptor and file are both absent when export is disabled or no chunk plan was accepted. Its payload contract is defined by Accepted Chunk Map.

manifest.json

manifest.json contains a run manifest. This abridged example shows its core structure:

{
  "run_id": "run-123",
  "pipeline_id": "dnd-session",
  "artifact_lanes": [
    {
      "id": "spells",
      "extractor": "dnd/spells",
      "merger": "appendorder",
      "normalizer": "noop"
    }
  ],
  "validation_status": "approved",
  "started_at": "2026-01-01T00:00:00Z",
  "completed_at": "2026-01-01T00:00:01Z"
}

Fields with empty values may be omitted by JSON encoding.

The manifest fields are:

  • run_id, pipeline_id, and pipeline_digest: run and resolved-pipeline identity;
  • input_module, chunker, extractors, merger, normalizer, and output_encoder: resolved module keys;
  • chunk_plan: payload-free provenance for the effective chunk plan. mode is the effective cache mode; action is reused, generated, refreshed, or bypassed when a plan was materialized. requested_module is the current pipeline chunker, while producer_input_module, producer_module, producer_llm_profile, producer_references, producer_metadata, source_digest, plan_digest, plan_schema_version, and created_at describe the stored or generated producer when available. A cached plan can therefore identify a producer different from the requested module. This object never embeds ranges, units, annotations, prompts, responses, or reference content;
  • module_metadata and artifact_lanes: module and per-lane provenance, including prompt and response-schema provenance when provided;
  • validator_chains: resolved validation points and validators;
  • source_digests and references: source and reference provenance;
  • normalized_outputs and rejected_outputs: payload-free result summaries;
  • llm_profiles: selected profile IDs and provider or model names when available;
  • metadata: the effective prompt session_id;
  • validation_status: approved or rejected;
  • started_at and completed_at: UTC run timestamps.

source_digests contains source document digests only. Bound references are recorded separately under references, which contains provenance only: target stage, lane ID when present, slot name, origin type and URI, digest, media type, byte size, and binding source. Reference content is not written to durable output.

Reference stage is chunk, extract, merge, or normalize. lane_id is omitted for chunk references and present for extract, merge, and normalize references.

validation_status is approved when no outputs were rejected and rejected when one or more outputs were rejected.

Producer warnings and the current run's chunk-validation warnings remain in warnings.json. The manifest records only provenance and decision summaries; empty producer-only values are omitted for compatibility with existing readers.

validator_chains records the resolved validator chain for each validation point. Entries include stage, lane ID when applicable, module key, and validators with key and execution class. Empty chains are recorded with an empty validators array, including chains resolved from explicit empty config overrides.

normalized_outputs summarizes each normalized lane output without embedding payload bytes. Entries include lane ID, normalizer module key, source ID, media type, and response schema provenance where available.

rejected_outputs summarizes rejected module outputs without embedding raw payload bytes. Entries include stage, lane, module, chunk, validator or reason, message, attempt count, and optional diagnostic artifact path.

Output Payload Files

Each normalized serialized artifact is written to lanes/<sanitized-lane-id>.json. The JSON output encoder is domain-neutral and accepts only artifacts whose codec media type is application/json. The file contains the codec-owned JSON bytes pretty-printed.

The schema of each lane payload is owned by that artifact contract. For the current D&D lanes, see D&D Spell Artifact, D&D NPC Artifact, and D&D Combat-Turn Artifact, and D&D Scene Description Artifact.

rejected.json

Shape:

{
  "rejected": []
}

When output validation rejects an output, each entry contains stage and message. It includes lane_id, module_key, chunk_id, chunk_index, validator_name, reason_code, attempt_count, and diagnostic_artifact_path when applicable.

warnings.json

Shape:

{
  "warnings": [
    {
      "scope": "extract",
      "reason_code": "example",
      "message": "human-readable warning"
    }
  ]
}

warnings is an empty array when no warnings are reported. Each warning requires reason_code and message; scope is omitted when it is empty.