Files
notarius/docs/integrations/chunk-map.md

3.9 KiB

Accepted Chunk Map

This document defines the durable accepted chunk-map artifact that the JSON output encoder can write as chunk-map.json. It describes the exact accepted, materialized chunks used by a run; it is not a lane artifact and is never an input to later pipeline steps. Enable it with the JSON output option described in Configuration.

Identity

  • Artifact kind: source/chunk-map
  • Logical file: chunk-map.json
  • Schema ID: notarius.source.chunk_map
  • Schema name: notarius_source_chunk_map_v1
  • Schema version: v1
  • Media type: application/json

The checked-in JSON Schema defines the strict wire shape. chunk-map.json is listed by the optional chunk_map descriptor in the JSON output index, not by the lane-oriented output_files collection.

Payload

The payload has these required fields:

  • source_id: accepted source-document identity.
  • source_digest: canonical lower-case sha256: digest of that document.
  • plan_digest: canonical lower-case sha256: digest of the accepted logical plan.
  • requested_chunker: chunk module selected by the current resolved pipeline.
  • producer: the original accepted-plan producer, with required input_module and chunk_module; llm_profile is present only for an LLM-backed producer.
  • plan_annotations: accepted plan-level annotation namespace map. It is {} when no namespaces are present.
  • chunks: non-empty execution-order collection of accepted chunks.

Each chunk has id, zero-based index, source_ref, positive unit_count, and an explicit annotations namespace map. A source reference has the source ID and inclusive positive start_unit_id and end_unit_id endpoints. Annotation values are arbitrary valid JSON under non-empty namespaces. They are preserved as canonical JSON without interpreting any module-specific namespace.

Invariants

The framework constructs this artifact only after materializing the selected logical plan and accepting it through the configured chunk validator chain. Construction proves the source and plan digests, source-document range order, materialized chunk IDs and indexes, source references, unit membership and counts, and plan/range annotations agree exactly. Chunk IDs are unique and indexes are contiguous and agree with array order.

The codec rejects unknown fixed-object fields, malformed identities or digests, invalid annotation JSON, trailing JSON content, and any payload whose reconstructed logical plan does not reproduce plan_digest. It makes defensive copies at serialization and decoding boundaries.

Acceptance And Provenance

The artifact is available only when the chunk plan was accepted. It remains available when a later extraction, merge, or normalization result is rejected; it is absent when chunk validation rejects the candidate plan.

requested_chunker describes the current pipeline selection. producer describes who originally produced the accepted plan. On a cache hit these can differ: the accepted ranges, annotations, digests, and stable materialized IDs are reused, while the producer remains the stored producer. Cache paths, actions, references, metadata, warnings, timestamps, and detailed provenance remain in the run manifest rather than this payload.

Data Handling

The map contains structure, not source content. It excludes transcript bytes, materialized units, source-unit metadata, chunk content, private model responses, rejected proposals, debug data, external-reference content, and filesystem paths.

Annotations can nevertheless be source- or model-derived. Treat an enabled chunk-map.json with the same sensitivity and retention expectations as lane output. Physical placement, confined atomic writing, and permissions follow the ordinary output operation.