Plan the ordered pipeline steps implementation

This commit is contained in:
2026-07-21 20:28:04 +00:00
parent f94ab0a6bf
commit f5618d1f0c
4 changed files with 805 additions and 677 deletions

View File

@@ -5,121 +5,17 @@ configuration, operations, internal, and integration docs. This roadmap records
future work only. Items are ordered roughly by current value and specificity,
not as committed release dates.
## Near-Term: Ordered Pipeline Steps
## Proposed Next Scope: Ordered Pipeline Steps
Allow one configured pipeline to contain multiple ordered execution steps so
accepted artifacts from an earlier step can become generated references for
later steps in the same run. This is a bounded extension of the fixed pipeline
model, not an arbitrary DAG or general workflow language.
Add ordered groups of artifact lanes, canonical generated-artifact references,
and dependency-aware checkpoint reuse without introducing a general DAG. The
first production workflow runs normalized D&D NPC extraction before spell and
combat-turn extraction and supplies that NPC artifact to their declared
reference slots.
Input parsing and chunk planning remain pipeline-wide. Each step selects one or
more artifact lanes; every selected lane completes extraction, validation,
merge, normalization, and validation before dependent later steps begin. Lanes
within the same step remain independent and may execute concurrently. The
runner exposes only accepted normalized artifacts across a step boundary; raw
extracts, rejected outputs, and intermediate merge results cannot become
downstream references.
Generated-reference bindings must be explicit in resolved configuration. A
binding identifies an earlier producing lane and one declared reference slot on
a later consuming lane. Resolution must reject missing producers, references to
the same or a later step, incompatible artifact kinds or media types, undeclared
consumer slots, cycles, and ambiguous bindings. A configured external reference
and a generated reference cannot bind the same effective target slot; reject
that pipeline or runtime override instead of applying a precedence rule.
Each effective target slot accepts at most one producer. One generated artifact
may fan out to multiple compatible target slots in a later step; aggregation
from multiple producers into one slot is deferred until a concrete use case
defines deterministic semantics.
Step-scoped reference defaults mirror existing pipeline-level reference
defaults. A step binds an earlier artifact once, and the binding automatically
applies to every target in that step that declares the named slot. Target-local
bindings remain available when only one module should consume the artifact; do
not combine a target-local and step-scoped binding for the same effective slot.
A generated source uses a structured, unambiguous form rather than encoding a
producer into a path-like string. The target configuration shape is:
```yaml
steps:
- id: identify-npcs
artifacts:
npcs:
extract: dnd/npcs
normalize: dnd/npcs
- id: grounded-events
references:
npcs:
artifact:
step: identify-npcs
lane: npcs
artifacts:
spells:
extract: dnd/spells
normalize: dnd/spells
combat:
extract: dnd/combat-turns
normalize: dnd/combat-turns
```
This snippet shows the step and reference portion of a pipeline; pipeline-wide
input, chunk, output, and other references are omitted. Existing scalar
reference values continue to mean external file paths; the structured
`artifact` form means an accepted normalized artifact from the named earlier
step and lane. Do not infer generated bindings from module keys, matching lane
names, or D&D-specific knowledge in the framework.
The handoff should use the producer artifact's canonical codec representation
and retain its artifact kind, schema identity, media type, content digest, and
producer provenance. Generated references provide context or disambiguation,
not source evidence. They use the existing module-facing reference contract
where possible; typed or domain-specific adapters may validate and prepare a
reference without moving domain concepts into the pipeline framework.
Pipeline identity, manifests, checkpoint dependencies, debug records, and
errors must include step identity and generated-reference provenance. A
downstream checkpoint is reusable only when the upstream artifact identity and
content digest match. Resume should reconstruct an accepted upstream artifact
through its registered codec instead of requiring the producing lane to run
again when its checkpoint is reusable.
### Dependency-aware resume and recomputation
Treat generated-reference bindings as checkpoint dependencies. Reusing an
earlier step is safe only when its existing checkpoint and codec identity are
valid. Reusing a dependent step additionally requires an exact match for every
upstream artifact kind, schema identity, media type, and canonical content
digest it consumed. A changed, missing, rejected, corrupt, or incompatible
producer artifact invalidates all transitive dependent checkpoints; the runner
must never combine a newly produced upstream artifact with stale downstream
output.
Support selectively recomputing one configured step and all of its transitive
dependents while retaining reusable independent and predecessor work. The
operator-facing selection mechanism should identify a stable configured step,
not individual internal stage operations. Resolution must reject a selection
that would omit a required predecessor without a reusable accepted artifact.
Manifests, checkpoint events, and diagnostics should distinguish work that was
executed, reused, or invalidated and record a bounded non-secret reason for
dependency-driven invalidation. Completion order must not affect invalidation,
public artifact ordering, or the set of dependent steps selected for rerun.
If a required producer finishes without an accepted normalized artifact, fail
the entire run with a deterministic dependency error. Do not start any
dependent step. Preserve the upstream rejection or empty-result outcome and
step provenance in the failed run manifest so the cause remains auditable.
The first production workflow is D&D NPC grounding:
1. the first step runs the NPC lane through accepted normalized output; and
2. the second step runs spell and combat-turn lanes, binding that NPC artifact
to the spell extractor and to the combat extractor and normalizer through
their existing `npcs` reference slots.
Spell and combat-turn extraction may run concurrently after the NPC handoff is
available. The NPC artifact may disambiguate participant identity, but it does
not prove that a spell cast or combat turn occurred.
The bounded target, compatibility and architecture decisions, exclusions, and
acceptance criteria are defined in
[Proposed Scope: Ordered Pipeline Steps](ordered-pipeline-steps.md).
## Near-Term D&D Pipeline