Archive the completed codebase audit
This commit is contained in:
@@ -1,516 +0,0 @@
|
||||
# Codebase Audit Plan
|
||||
|
||||
## Purpose
|
||||
|
||||
This document defines a repository-wide audit of Notarius for correctness,
|
||||
efficiency, maintainability, and clarity. The audit should identify concrete
|
||||
improvements without treating abstraction, fewer lines, or higher test coverage
|
||||
as goals in themselves.
|
||||
|
||||
The audit is intentionally separate from implementation. Its findings should
|
||||
be evidence-backed and sufficiently specific to support a later remediation
|
||||
roadmap, but the audit should not modify production code, tests, assets, or
|
||||
current-behavior documentation.
|
||||
|
||||
## Governing Principles
|
||||
|
||||
The audit must preserve the architecture and testing policies in
|
||||
`docs/policy/architecture.md` and `docs/policy/testing.md`.
|
||||
|
||||
In particular:
|
||||
|
||||
- Notarius remains a fixed, staged pipeline rather than a general workflow
|
||||
engine.
|
||||
- Generic framework packages must remain domain-neutral, and production
|
||||
modules must not acquire CLI or physical-state responsibilities.
|
||||
- Typed artifact boundaries, exact codec compatibility, deterministic ordering,
|
||||
whole-output validation, and generated-reference provenance are correctness
|
||||
properties, not incidental complexity to be optimized away.
|
||||
- The root `assets` package remains a content-only dependency leaf.
|
||||
- Shared helpers should protect demonstrated common semantics. Similar-looking
|
||||
code with different ownership, error policy, identity rules, or type contracts
|
||||
should remain separate.
|
||||
- Tests should protect durable behavior and meaningful risks. The audit should
|
||||
not recommend tests merely to increase coverage or freeze implementation
|
||||
details.
|
||||
- Efficiency claims must distinguish measured or structurally credible costs
|
||||
from cosmetic line-count reductions. Optimizing local CPU work that is
|
||||
insignificant beside an LLM call is low priority unless it also simplifies
|
||||
correctness or applies to large inputs.
|
||||
|
||||
## Audit Questions
|
||||
|
||||
Every audited area should be examined through the following questions.
|
||||
|
||||
### Correctness
|
||||
|
||||
- Are documented architecture invariants enforced at the correct boundary?
|
||||
- Can invalid configuration, incompatible artifact types, malformed references,
|
||||
or unavailable dependencies reach execution when they could be rejected
|
||||
during resolution or preparation?
|
||||
- Are nil, empty, absent, rejected, failed, and canceled states distinguished
|
||||
consistently?
|
||||
- Are stored or returned slices, maps, byte slices, options, metadata, source
|
||||
documents, references, and artifacts defensively owned where required?
|
||||
- Are public ordering, selected errors, warnings, and checkpoint decisions
|
||||
deterministic regardless of map or goroutine completion order?
|
||||
- Do cancellation, retry, validation, and partial-work semantics match their
|
||||
documented ownership?
|
||||
- Do checkpoint and chunk-plan identities include every semantic dependency and
|
||||
exclude scheduling-only or diagnostic state?
|
||||
- Can auxiliary references accidentally become source evidence, or can
|
||||
generated references bypass codec, schema, provenance, or step-order checks?
|
||||
- Can provider-specific values, credentials, or source content escape through
|
||||
errors, manifests, debug summaries, cache state, or logs?
|
||||
- Do schemas, codecs, candidate decoders, normalizers, and validators agree on
|
||||
the exact durable contract without silently accepting incompatible shapes?
|
||||
|
||||
### Duplication And Shared Mechanics
|
||||
|
||||
- Which exact or near-duplicate implementations express the same invariant and
|
||||
failure policy?
|
||||
- Has duplicated code already drifted in naming, nil handling, canonicalization,
|
||||
metadata, fingerprints, validation, or diagnostics?
|
||||
- Would a helper have a natural owner and a smaller, clearer contract than the
|
||||
duplicated callers?
|
||||
- Can an extraction preserve static typing and package ownership, or would it
|
||||
require reflection, `any`, callbacks with many policy parameters, or a
|
||||
domain-neutral package importing domain concepts?
|
||||
- Is repeated code required by a small interface adapter or typed registration
|
||||
boundary and therefore clearer when left explicit?
|
||||
|
||||
As a default heuristic, prioritize a shared helper when identical semantics
|
||||
appear in three or more production sites, or in two sites where divergence
|
||||
would create a meaningful correctness risk. Do not use that heuristic as a
|
||||
quota: one substantial duplicate may warrant extraction, while widespread
|
||||
one-line interface methods may not.
|
||||
|
||||
### Simplicity And Idiomatic Go
|
||||
|
||||
- Does a function combine orchestration, policy, transformation, persistence,
|
||||
and reporting that could be separated along existing ownership boundaries?
|
||||
- Are repeated scans, sorts, conversions, clones, encodes, or decodes doing work
|
||||
that can safely occur once?
|
||||
- Are intermediate representations necessary, or can a value be validated,
|
||||
canonicalized, and mapped in one comprehensible pass?
|
||||
- Are maps, sets, stable sorts, generics, standard-library helpers, and error
|
||||
wrapping used idiomatically?
|
||||
- Are abstractions earning their complexity, or are interfaces, option layers,
|
||||
wrappers, aliases, compatibility paths, and private types left over after a
|
||||
completed migration?
|
||||
- Are there unreachable error branches, redundant fingerprints or digests,
|
||||
duplicated sources of truth, or accessors used only by tests?
|
||||
- Can a smaller implementation preserve exact observable behavior and safety
|
||||
properties?
|
||||
|
||||
### Explanatory Comments
|
||||
|
||||
Comments should be recommended where the code is necessarily complex because
|
||||
it preserves a non-obvious invariant. Good candidates include:
|
||||
|
||||
- concurrency coordination, cancellation, and stable error selection;
|
||||
- checkpoint identity, reuse, forced recomputation, and dependency invalidation;
|
||||
- typed erasure and restoration at framework boundaries;
|
||||
- generated-reference ordering and provenance;
|
||||
- canonicalization and identity resolution where registry evidence differs
|
||||
from occurrence evidence;
|
||||
- prompt ordering or input identity required for backend caching; and
|
||||
- path confinement, atomic publication, redaction, or terminal error precedence.
|
||||
|
||||
Recommend comments that explain *why* a step or ordering constraint exists and
|
||||
what would break if it changed. Do not recommend comments that narrate syntax,
|
||||
repeat a function name, duplicate current-behavior documentation, or preserve
|
||||
implementation history.
|
||||
|
||||
### Tests
|
||||
|
||||
- Is each consequential invariant protected at the narrowest stable boundary?
|
||||
- Are concurrency, cancellation, retries, recovery, compatibility, path safety,
|
||||
and data-integrity behavior credibly exercised?
|
||||
- Do higher-level contract tests duplicate lower-level cases without adding
|
||||
integration confidence?
|
||||
- Are tests coupled to private constants, helper shape, exact prose, full error
|
||||
strings, or collaborator choreography rather than behavior?
|
||||
- Can repetitive fixtures or fakes be simplified without creating a test
|
||||
framework more complex than the tests?
|
||||
- Would a focused fuzz test, race test, or package-level invariant test protect
|
||||
a realistic risk better than several example tests?
|
||||
|
||||
## Evidence And Finding Standards
|
||||
|
||||
Static metrics and textual similarity are discovery aids, not findings. A long
|
||||
function may be a clear linear coordinator; identical methods may be useful
|
||||
typed adapters. Every reported finding must include:
|
||||
|
||||
1. a concise title and severity;
|
||||
2. exact files and symbols;
|
||||
3. the observed behavior or structural evidence;
|
||||
4. the correctness, efficiency, maintenance, or comprehension impact;
|
||||
5. a concrete recommended direction;
|
||||
6. important invariants the remediation must preserve;
|
||||
7. focused validation that would demonstrate success; and
|
||||
8. whether the recommendation is independent or should be grouped with another
|
||||
finding.
|
||||
|
||||
Use these severities:
|
||||
|
||||
- **High:** a credible risk of corrupt output, unsafe state handling, secret
|
||||
exposure, stale reuse, deadlock, nondeterminism, or violated external
|
||||
contract.
|
||||
- **Medium:** a plausible behavioral defect, meaningful wasted work on common
|
||||
paths, or complexity/duplication likely to cause future correctness drift.
|
||||
- **Low:** a contained simplification, small efficiency improvement, dead code,
|
||||
naming issue, or missing explanation with no current behavioral failure.
|
||||
|
||||
The audit should explicitly record examined areas with no findings. This makes
|
||||
coverage visible and prevents later agents from repeatedly rediscovering the
|
||||
same safe design.
|
||||
|
||||
## Repository Areas
|
||||
|
||||
### 1. Architecture And Dependency Boundaries
|
||||
|
||||
Inspect `docs/policy/architecture.md`, `docs/adr/`, `docs/internal/overview.md`,
|
||||
package imports, module registrars, and the CLI composition root.
|
||||
|
||||
Look for:
|
||||
|
||||
- framework or core code depending on production modules;
|
||||
- modules depending on CLI, physical roots, or provider-specific types;
|
||||
- domain knowledge placed in generic helpers;
|
||||
- duplicated registries or composition policy outside the owning registrar;
|
||||
- abstractions that turn the fixed pipeline into an implicit general graph; and
|
||||
- current code that no longer matches an accepted ADR or documented invariant.
|
||||
|
||||
Graph-reported cross-layer calls must be traced before being classified because
|
||||
tests and interface implementations can resemble dependency inversions without
|
||||
creating a production import violation.
|
||||
|
||||
### 2. Configuration And CLI Composition
|
||||
|
||||
Inspect `internal/core/config`, `internal/cli`, configuration parsing and
|
||||
redaction tests, profile construction, session derivation, catalog assembly,
|
||||
reference overrides, run-result handling, terminal reporting, and maintained
|
||||
example contract tests.
|
||||
|
||||
Pay particular attention to the currently dense paths around
|
||||
`runPipelineCommand`, configuration profile validation, selected reference
|
||||
targets, recomputation policy, and option normalization. Determine whether
|
||||
their complexity reflects necessary composition or mixed responsibilities that
|
||||
can be separated without moving policy into the framework.
|
||||
|
||||
Verify:
|
||||
|
||||
- file, environment, CLI, pipeline, binding, and prompt-default precedence;
|
||||
- consistent strict option and unknown-field handling;
|
||||
- session identity independence from references and pipeline-local changes;
|
||||
- effective profile and runtime fingerprint consistency;
|
||||
- redaction before errors or debug/manifest boundaries;
|
||||
- output publication only after framework success; and
|
||||
- one guarded terminalization path that preserves the primary failure.
|
||||
|
||||
### 3. Pipeline Resolution, Preparation, And Typed Registries
|
||||
|
||||
Inspect `internal/framework/pipeline/profile.go`, `prepare.go`, registry files,
|
||||
`options.go`, `references.go`, `handoff.go`, `construction.go`, typed contracts,
|
||||
and their focused tests.
|
||||
|
||||
This area deserves a dedicated pass because the current graph identifies
|
||||
`ResolvePipeline`, generated-binding validation, reference-target resolution,
|
||||
and generated-reference construction as high-complexity or high-fan-in code.
|
||||
|
||||
Verify:
|
||||
|
||||
- static failures occur before source parsing;
|
||||
- selected and unselected lanes do not contaminate each other's requirements;
|
||||
- stage defaults and overrides have one canonical resolution path;
|
||||
- typed registration and private erasure cannot panic or accept near-matching
|
||||
artifact types;
|
||||
- generated bindings reject cycles, forward references, ambiguity, wrong kinds,
|
||||
and missing accepted normalized producers;
|
||||
- materialized reference bytes and options are cloned and bounded; and
|
||||
- resolved composition and prepared fingerprints include the complete semantic
|
||||
policy exactly once.
|
||||
|
||||
Compare input, chunker, extractor, merger, normalizer, output, validator, codec,
|
||||
evidence-projector, and validator-chain registries for shared mechanics and
|
||||
intentional differences. Repeated typed registration code is a candidate only
|
||||
if a helper can retain useful compile-time guarantees and stage-specific
|
||||
diagnostics.
|
||||
|
||||
### 4. Pipeline Execution, Validation, Retry, And Concurrency
|
||||
|
||||
Inspect `runner*.go`, `typed_execution.go`, `runner_typed.go`,
|
||||
`runner_concurrent.go`, validation-chain execution, normalize retry behavior,
|
||||
synchronized collaborators, and the concurrency, cancellation, retry, debug,
|
||||
and checkpoint tests.
|
||||
|
||||
Trace complete paths rather than reviewing helper files in isolation:
|
||||
|
||||
- source and chunk-plan selection through chunk validation;
|
||||
- deterministic chunk-first/lane-second dispatch;
|
||||
- lane extraction through merge and normalize continuations;
|
||||
- rejection versus framework-error propagation;
|
||||
- cancellation before dispatch, while queued, and while running;
|
||||
- retry attempts and warning retention;
|
||||
- stable error selection after concurrent completion;
|
||||
- checkpoint hydration back into typed execution; and
|
||||
- output suppression after a framework error.
|
||||
|
||||
Look for goroutine leaks, unbounded work, lock-order risks, double release or
|
||||
double recording, races on shared result state, unnecessary serialization,
|
||||
and repeated canonicalization. Comments are especially valuable here when they
|
||||
explain ordering or cancellation invariants that are not apparent from local
|
||||
control flow.
|
||||
|
||||
### 5. State, Checkpoints, Chunk Plans, Debugging, And File Safety
|
||||
|
||||
Inspect `internal/framework/checkpoint`, `chunkplan`, `chunkmap`, `debug`,
|
||||
`evidencecontext`, `internal/core/fileio`, `debugbundle`, and their CLI
|
||||
composition.
|
||||
|
||||
Verify:
|
||||
|
||||
- narrow path validation and symlink-resistant confinement;
|
||||
- atomic writes and recoverable explicit cleanup;
|
||||
- separation of checkpoint recording, resume loading, chunk-plan caching, and
|
||||
debug capture;
|
||||
- canonical encoding before content identity is trusted;
|
||||
- complete but non-secret checkpoint fingerprints;
|
||||
- correct ordinary-resume and selective-recompute behavior;
|
||||
- producer dependency invalidation across ordered steps;
|
||||
- immutable hydration and no aliasing with stored bytes;
|
||||
- debug data never influencing execution or reuse; and
|
||||
- terminal persistence failures never obscuring the primary error.
|
||||
|
||||
Review the repeated extract/merge/normalize recorder and loader methods, path
|
||||
component validators in multiple state packages, and clone/encode/decode paths.
|
||||
Determine which repetition is a clear stage adapter and which can share a
|
||||
private primitive without weakening reason-code ownership or diagnostics.
|
||||
|
||||
### 6. LLM Runtime, Prompt Filesystems, And Assets
|
||||
|
||||
Inspect `internal/framework/llm`, `promptfs`, the PromptKit integration,
|
||||
scheduler, profile-source construction, prompt/schema registries, root
|
||||
`assets`, module prompt manifests, and relevant D&D shared assets.
|
||||
|
||||
Verify:
|
||||
|
||||
- every provider call passes through the shared scheduler and cancellation
|
||||
removes queued calls safely;
|
||||
- PromptKit and Notarius concurrency limits compose as documented;
|
||||
- profile inspection and runtime use identical source precedence;
|
||||
- session IDs, profile-source fingerprints, prompt fingerprints, and schema
|
||||
fingerprints reflect the intended semantic inputs;
|
||||
- secrets and provider-specific error types do not cross the boundary;
|
||||
- prompt inputs and private outputs do not expose opaque entity IDs;
|
||||
- prompt ordering, stable prefixes, and cache controls remain intentional;
|
||||
- schema loaders and filesystem adapters validate once and return defensive
|
||||
data; and
|
||||
- the root assets package contains no business logic.
|
||||
|
||||
Compare the LLM asset registry and prompt-filesystem adapters for duplicated
|
||||
filesystem behavior. Review repeated prompt/schema loader and metadata code in
|
||||
module packages, but reject an extraction that would centralize domain prompt
|
||||
ownership or make unrelated assets share one invalidation boundary.
|
||||
|
||||
### 7. Generic And Seriatim Modules
|
||||
|
||||
Inspect `internal/modules/generic` and `internal/modules/seriatim`, including
|
||||
module specs, option decoding, chunk planning, input translation, validators,
|
||||
output encoding, evidence-context publication, registration, and tests.
|
||||
|
||||
Verify that:
|
||||
|
||||
- external Seriatim details end at the input boundary;
|
||||
- generic chunking and output remain domain-neutral;
|
||||
- chunk plans and source units preserve source-addressed invariants;
|
||||
- output logical names are safe and deterministic;
|
||||
- output options do not bypass preparation-time compatibility checks; and
|
||||
- option decoding is strict, small, and consistent with configuration
|
||||
validation.
|
||||
|
||||
The graph flags generic integer option parsing and JSON output policy decoding
|
||||
as relatively complex. Examine whether that is inherent strict decoding or an
|
||||
opportunity for a smaller typed parser with equally precise diagnostics.
|
||||
|
||||
### 8. D&D Domain Model, Codecs, And Shared Helpers
|
||||
|
||||
Inspect `internal/modules/dnd` domain types, codecs, candidate decoders,
|
||||
identity packages, registries, shared source-reference helpers, diagnostics,
|
||||
registry resolution, entity reconciliation, mergers, registrar, and assets.
|
||||
|
||||
Compare all ten current artifact families. Build a convention matrix covering:
|
||||
|
||||
- module specs and execution classes;
|
||||
- constructor and option behavior;
|
||||
- manifest metadata and checkpoint fingerprints;
|
||||
- response-schema loading and private-versus-durable types;
|
||||
- source-reference conversion, canonicalization, ordering, and deduplication;
|
||||
- nil versus present-empty output;
|
||||
- codecs and strict JSON behavior;
|
||||
- registry lookup, identity derivation, immutable projections, and resolution;
|
||||
- normalizer retry/fallback behavior;
|
||||
- validators and default chains; and
|
||||
- registration, prompt assets, and documentation ownership.
|
||||
|
||||
The graph reports many exact similarities among codec `Decode` methods,
|
||||
fingerprint/metadata methods, registry extractors, identity helpers, occurrence
|
||||
normalizers, and validators. Treat these as a prioritized review list, not an
|
||||
instruction to create one generic D&D engine. A worthwhile helper must preserve
|
||||
domain-specific identity, evidence, kind ordering, validation, diagnostics,
|
||||
and artifact typing.
|
||||
|
||||
### 9. D&D Extraction And Normalization Flows
|
||||
|
||||
Trace each lane end to end rather than auditing only similarly named files:
|
||||
|
||||
- spells;
|
||||
- NPC registry and NPC occurrences;
|
||||
- combat turns and enemy events;
|
||||
- item registry and item occurrences;
|
||||
- scene descriptions; and
|
||||
- location registry and location occurrences.
|
||||
|
||||
For registry/occurrence pairs, verify the complete semantic boundary: the model
|
||||
uses contextual evidence, deterministic code attaches opaque identity, registry
|
||||
evidence does not become occurrence evidence, and unresolved or ambiguous
|
||||
selections fail according to lane policy.
|
||||
|
||||
Review whether any lane resolves or canonicalizes the same entity, source
|
||||
reference, or response twice; constructs unnecessary intermediate response
|
||||
forms; performs repeated sorts or scans; or retains transitional paths. Compare
|
||||
registry normalizers and occurrence normalizers for genuinely identical
|
||||
mechanics, while keeping currency, same-name location, NPC ambiguity, spell
|
||||
catalog, scene eligibility, and combat-specific policy with their owners.
|
||||
|
||||
### 10. Test Suite And Comment Coverage
|
||||
|
||||
Review the tests associated with every preceding area after understanding the
|
||||
production contracts. This should be a cross-cutting pass, not a request to add
|
||||
tests for every flagged function.
|
||||
|
||||
Identify:
|
||||
|
||||
- consequential unprotected invariants;
|
||||
- duplicated policy assertions across layers;
|
||||
- brittle tests coupled to internal constants, prompt prose, or private helper
|
||||
shape;
|
||||
- oversized test harnesses and repeated fixtures that obscure intent;
|
||||
- race-sensitive code not exercised under `-race`;
|
||||
- parsers, canonicalizers, and path handlers where fuzzing would address a real
|
||||
input-space risk; and
|
||||
- complex production code whose tests reveal an unclear ownership boundary.
|
||||
|
||||
Also identify necessarily complex symbols that lack a concise invariant-level
|
||||
comment. Comment recommendations should name the exact symbol and the fact the
|
||||
comment should explain; “add more comments” is not an actionable finding.
|
||||
|
||||
## Efficiency Evaluation
|
||||
|
||||
The audit should consider both runtime and maintenance efficiency.
|
||||
|
||||
For runtime efficiency, examine algorithmic behavior relative to realistic
|
||||
input dimensions: source units, chunks, lanes, references, artifacts, registry
|
||||
records, checkpoint files, and prompt assets. Prioritize repeated full-input
|
||||
passes, nested linear lookup, unnecessary JSON round trips, repeated hashing,
|
||||
large defensive copies at adjacent ownership boundaries, and serialization on
|
||||
concurrent hot paths. Preserve a defensive copy when it establishes ownership;
|
||||
removing it solely to reduce allocation is not an improvement.
|
||||
|
||||
For maintenance efficiency, prioritize repeated policy, parallel type systems,
|
||||
duplicated error classification, scattered defaults, and migrations that left
|
||||
two ways to perform the same operation. Boilerplate is costly only when it can
|
||||
drift or obscures the semantic core. Small explicit typed adapters can be more
|
||||
maintainable than a generic abstraction.
|
||||
|
||||
Do not recommend caching, pooling, concurrency, or a benchmark without naming
|
||||
the workload and risk it addresses. Add a benchmark only when a proposed
|
||||
optimization concerns a repeatable local path and the result would influence
|
||||
the decision.
|
||||
|
||||
## Audit Method
|
||||
|
||||
Each area should use the same method:
|
||||
|
||||
1. Read its architecture/internal documentation and focused tests.
|
||||
2. Map public/package contracts and trace the main call paths.
|
||||
3. Inspect high-fan-in, high-cognitive-complexity, nested-loop, and repeated-
|
||||
conversion symbols.
|
||||
4. Review exact and near-duplicate code side by side, including callers and
|
||||
failure semantics.
|
||||
5. Check dependency direction, ownership, aliasing, deterministic order,
|
||||
cancellation, and error classification.
|
||||
6. Compare tests with the risks owned at that layer.
|
||||
7. Record findings and inspected-with-no-finding areas before moving on.
|
||||
8. Run focused read-only validation when it can confirm or refute a suspected
|
||||
problem.
|
||||
|
||||
Prefer the repository knowledge graph for symbol discovery, call tracing, and
|
||||
similarity candidates. Use textual search for literals, diagnostics, config
|
||||
keys, asset content, and stale names. Read complete implementations and tests
|
||||
before reporting a metric-derived candidate.
|
||||
|
||||
## Execution Sequence
|
||||
|
||||
This document owns audit scope, questions, evidence standards, and the quality
|
||||
bar. [Staged Codebase Audit Sequence](audit-sequence.md) is the sole canonical
|
||||
owner of prompt order, stage boundaries, per-stage reading, validation commands,
|
||||
and acceptance criteria. Do not derive or maintain a second sequence here.
|
||||
|
||||
The audit is executed as bounded prompts and writes its accumulated findings to
|
||||
`docs/roadmap/audit.md`. Later stages must build on and reconcile earlier
|
||||
evidence rather than concatenate independent reports. Implementation and
|
||||
roadmap retirement remain separate work after maintainers review the completed
|
||||
audit.
|
||||
|
||||
## Baseline And Validation
|
||||
|
||||
Before the first audit stage, record the commit under review and require a clean
|
||||
worktree. Refresh the code knowledge graph so renamed or deleted code does not
|
||||
produce false findings. Run the normal offline baseline:
|
||||
|
||||
```sh
|
||||
go test ./...
|
||||
go vet ./...
|
||||
go build ./cmd/notarius
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Run `go test -race` for packages with concurrency or mutable shared state,
|
||||
especially `internal/framework/pipeline`, `internal/framework/llm`, state
|
||||
packages, and D&D registries. A repository-wide race run is appropriate for
|
||||
final verification if its cost remains reasonable.
|
||||
|
||||
Optional diagnostic commands should be used only when relevant:
|
||||
|
||||
- `go test -count=1` to rule out cache-masked failures;
|
||||
- `go test -shuffle=on` to detect order coupling;
|
||||
- focused fuzzing for existing or newly justified fuzz targets; and
|
||||
- focused benchmarks or profiles for a specific efficiency finding.
|
||||
|
||||
The audit itself should not change tests to make the baseline pass. Record any
|
||||
pre-existing failure and distinguish it from an audit finding.
|
||||
|
||||
## Deliverable Quality Bar
|
||||
|
||||
The completed audit should:
|
||||
|
||||
- cover every repository area listed above;
|
||||
- distinguish defects from refactoring opportunities and comment requests;
|
||||
- distinguish credible performance costs from aesthetic simplification;
|
||||
- identify intentional duplication that should remain explicit;
|
||||
- avoid recommendations that violate dependency direction or weaken typing;
|
||||
- cite exact evidence and preserve named invariants for every finding;
|
||||
- consolidate root causes rather than report many symptoms;
|
||||
- rank independent work so a later implementation plan can stage it safely;
|
||||
- recommend no code change whose expected benefit is smaller than its added
|
||||
abstraction or test-maintenance cost; and
|
||||
- leave implementation and roadmap retirement to later work.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None are required to begin the audit. If a later stage cannot determine whether
|
||||
behavior is intentional from code, tests, policies, ADRs, or current
|
||||
documentation, it should record the uncertainty and a recommended resolution
|
||||
rather than silently treating preference as a defect.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,350 +0,0 @@
|
||||
# Contextual Entity Grounding
|
||||
|
||||
## Purpose
|
||||
|
||||
Notarius should use an LLM for semantic interpretation of source evidence, not
|
||||
for referential-integrity work that deterministic code can perform more
|
||||
reliably. D&D prompts must therefore stop requiring models to reproduce opaque
|
||||
machine identifiers such as hash-derived entity IDs. Models should identify
|
||||
entities through human-readable, evidence-grounded context, after which
|
||||
Notarius resolves the selection and attaches the canonical durable identity.
|
||||
|
||||
This roadmap defines the policy, affected D&D prompt families, and intended
|
||||
end state. The ordered work needed to reach that state is maintained in
|
||||
[Implementation Plan](implementation.md).
|
||||
|
||||
## User Intent
|
||||
|
||||
The change has two goals:
|
||||
|
||||
- prevent otherwise useful model responses from failing because a long,
|
||||
non-semantic string was copied incorrectly; and
|
||||
- avoid spending prompt space and model effort on exact-copy work that provides
|
||||
no semantic value.
|
||||
|
||||
The policy is not a ban on identifiers. Durable artifacts may continue to use
|
||||
application-owned IDs, and prompts may continue to request source-unit ranges
|
||||
that locate evidence. The policy governs which identity work is assigned to
|
||||
the model.
|
||||
|
||||
## Policy
|
||||
|
||||
An LLM-facing prompt input or private response schema must not require a model
|
||||
to reproduce an opaque machine identifier when Notarius can establish the same
|
||||
association deterministically.
|
||||
|
||||
Opaque machine identifiers include cryptographic hashes, UUIDs, digests,
|
||||
database keys, durable entity IDs, and other tokens whose characters do not
|
||||
carry source-grounded meaning for the model. These values may remain in
|
||||
application state, provenance, diagnostics, checkpoints, and durable artifact
|
||||
contracts, but should be omitted from model-visible material when they do not
|
||||
help the model make a semantic decision.
|
||||
|
||||
The intended responsibility boundary is:
|
||||
|
||||
- the model decides which contextual entity is supported by the supplied
|
||||
evidence and returns the bounded semantic facts requested by the module;
|
||||
- the calling module validates that the contextual selection resolves to
|
||||
exactly one supplied candidate;
|
||||
- deterministic code supplies the canonical display value and durable entity
|
||||
ID; and
|
||||
- existing validators continue to enforce referential integrity at later
|
||||
artifact boundaries.
|
||||
|
||||
Transcript `start_unit_id` and `end_unit_id` values are permitted. They are
|
||||
contextual source coordinates and form part of the evidence contract rather
|
||||
than arbitrary identity tokens. Prompt IDs, schema IDs, fingerprints, session
|
||||
IDs, and digests may also remain in runtime metadata that the model is not
|
||||
asked to reproduce.
|
||||
|
||||
Short request-local labels are a narrowly permitted fallback only when a
|
||||
contextual selector cannot uniquely represent the available choices without
|
||||
unreasonable prompt cost. Such a label must be compact, scoped to one request,
|
||||
validated against the supplied candidate set, and never reused as a durable
|
||||
identity. Current D&D occurrence and reconciliation prompts should be designed
|
||||
without this exception; adopting it later requires a concrete demonstrated
|
||||
need and documented rationale.
|
||||
|
||||
## Current State
|
||||
|
||||
The initial NPC, item, and location registry extractors already follow the
|
||||
desired pattern: the model returns contextual names and evidence, and Notarius
|
||||
derives durable IDs afterward. Spells, combat turns, and enemy events use
|
||||
contextual actor names rather than requiring hash-derived NPC IDs.
|
||||
|
||||
Two current prompt families diverge from that pattern:
|
||||
|
||||
1. `dnd/npc-occurrences`, `dnd/item-occurrences`, and
|
||||
`dnd/location-occurrences` place durable registry IDs in model-visible
|
||||
projections and require the private LLM response to repeat those IDs.
|
||||
2. NPC-, item-, and location-registry normalization use the shared entity
|
||||
reconciliation prompt, which labels candidates with opaque
|
||||
`candidate-000001`-style keys and requires the model to copy those keys into
|
||||
duplicate-group proposals.
|
||||
|
||||
The durable occurrence artifacts correctly retain canonical entity IDs. The
|
||||
problem is the private model transport contract, not the published artifact
|
||||
contract.
|
||||
|
||||
## Target Architecture
|
||||
|
||||
### Model proposals and durable artifacts
|
||||
|
||||
Private LLM response types must express contextual semantic proposals rather
|
||||
than reuse the durable artifact type when that type contains an opaque entity
|
||||
ID. The extractor maps a validated private response into the existing durable
|
||||
artifact only after identity resolution succeeds.
|
||||
|
||||
No affected durable artifact kind, media type, schema ID, schema version, or
|
||||
JSON field changes as part of this work. NPC, item, and location occurrence
|
||||
artifacts continue to publish their exact canonical ID/name pair. Registry
|
||||
artifacts likewise retain their IDs and evidence.
|
||||
|
||||
The private schemas and prompt declarations may remain at their current `v1`
|
||||
identities because Notarius is pre-release and these are not external
|
||||
contracts. Their content hashes, mapping-policy fingerprints, and affected
|
||||
prompt fingerprints must change so incompatible checkpoints are not reused.
|
||||
|
||||
### NPC occurrence grounding
|
||||
|
||||
The NPC occurrence prompt receives an ordered names-only projection of the
|
||||
normalized NPC registry. Its private response contains the canonical NPC name,
|
||||
occurrence kind, and current-transcript source ranges, but no `npc_id`.
|
||||
|
||||
The extractor resolves the returned name under the existing NPC comparison
|
||||
policy. Resolution must produce exactly one registry entry. It then writes that
|
||||
entry's canonical display name and durable ID into the `dnd.NPCOccurrence`.
|
||||
An unknown or ambiguous selection invalidates the extraction operation; the
|
||||
extractor must not guess, use fuzzy matching, silently omit the record, or
|
||||
accept a partial response.
|
||||
|
||||
### Item occurrence grounding
|
||||
|
||||
The item occurrence prompt receives an ordered names-only projection of the
|
||||
normalized item registry. Its private response contains the canonical item
|
||||
name, occurrence kind, kind-specific fields, and current-transcript source
|
||||
ranges, but no `item_id`.
|
||||
|
||||
The extractor resolves the returned name under the existing item comparison
|
||||
and identity policies. Resolution must produce exactly one registry entry,
|
||||
whose canonical name and durable ID are attached deterministically. Unknown or
|
||||
ambiguous selections invalidate the complete extraction operation rather than
|
||||
being guessed, repaired by similarity, or dropped.
|
||||
|
||||
### Location occurrence grounding
|
||||
|
||||
Location identity cannot always be resolved from a display name alone: the
|
||||
current registry intentionally permits same-name locations with distinct
|
||||
source anchors. The location occurrence prompt must therefore receive a
|
||||
contextual registry descriptor that contains the canonical display name plus
|
||||
the minimum source-grounded registry evidence needed to distinguish same-name
|
||||
records. It must not contain the durable `location:sha256:...` value.
|
||||
|
||||
The private response uses two required selector fields: `name` and
|
||||
`registry_refs`. For a comparison-unique canonical name, `registry_refs` is an
|
||||
empty array and Notarius resolves the name under the location comparison
|
||||
policy. For a name shared by multiple registry records, `registry_refs`
|
||||
contains that record's complete canonically ordered registry ranges as
|
||||
`start_unit_id` and `end_unit_id` pairs, without `source_id`.
|
||||
|
||||
Every model-facing registry entry uses one fixed shape with required `name`,
|
||||
`registry_refs`, and `context` fields. `context` is an array of strict objects
|
||||
containing only `unit_id` and `text`. Comparison-unique entries use empty
|
||||
`registry_refs` and `context` arrays. Same-name entries use the complete
|
||||
registry-range selector and the bounded context described below. The model
|
||||
returns only `name` and `registry_refs`; it does not reproduce `context`.
|
||||
|
||||
For same-name groups, the projection also supplies bounded transcript units
|
||||
covered by each record's registry ranges so the model receives meaningful
|
||||
identity context rather than coordinates alone. Those ranges must resolve
|
||||
against the current source document, and the resulting contextual selectors
|
||||
must be unique. An invalid range or selector collision prevents the LLM call
|
||||
and fails the operation. Unique-name entries do not repeat registry ranges or
|
||||
context in the selector, preserving compatibility with a valid registry from
|
||||
another source when the name alone is unambiguous.
|
||||
|
||||
The private response separately supplies current-transcript `source_refs` that
|
||||
prove the occurrence. Registry identity evidence and occurrence evidence must
|
||||
remain different fields and must never be merged. The model should omit an
|
||||
occurrence when the transcript does not support choosing among same-name
|
||||
locations. If a returned selector does not resolve to exactly one supplied
|
||||
registry record, the extractor invalidates the complete operation rather than
|
||||
guessing.
|
||||
|
||||
### Registry reconciliation
|
||||
|
||||
The shared entity-reconciliation input replaces opaque candidate keys with
|
||||
contextual candidate descriptors. At minimum, a descriptor contains the
|
||||
candidate's display name and its canonically ordered source-reference ranges;
|
||||
the existing transcript windows remain available for semantic judgment.
|
||||
|
||||
Duplicate-group members and the canonical member in the private response use
|
||||
the same contextual descriptor shape. The shared reconciliation helper maps
|
||||
each descriptor back to exactly one internal candidate before assessing the
|
||||
proposal. Exact deterministic duplicates should already be removed before the
|
||||
LLM call; any remaining descriptor collision makes the affected candidate
|
||||
ineligible for model-assisted reconciliation rather than authorizing an
|
||||
arbitrary choice.
|
||||
|
||||
Existing safety behavior remains in force: groups must contain at least two
|
||||
supplied candidates, the canonical candidate must be a member, groups must not
|
||||
overlap, and domain-specific eligibility rules remain authoritative. Invalid,
|
||||
ambiguous, or unsafe groups are discarded through the existing bounded
|
||||
fallback and diagnostic behavior. The model never directly mutates the
|
||||
durable registry.
|
||||
|
||||
The shared private reconciliation schema and helper must remain domain-neutral
|
||||
within the D&D family. NPC-, item-, and location-specific duplicate policy
|
||||
continues to live in the owning normalizer.
|
||||
|
||||
## Prompt And Asset Changes
|
||||
|
||||
The following LLM-facing assets are in scope:
|
||||
|
||||
- the prompt instructions, registry input fragments, and private response
|
||||
schemas for NPC, item, and location occurrences;
|
||||
- the prompt manifests where input shape or selected fragments change;
|
||||
- the shared D&D entity-reconciliation fragment and private response schema;
|
||||
and
|
||||
- the NPC-, item-, and location-registry normalization prompt inputs that use
|
||||
the shared reconciliation contract.
|
||||
|
||||
Affected projections must exclude durable entity IDs rather than merely stop
|
||||
mentioning them in prose. Prompt instructions should describe the contextual
|
||||
selection rule once at the narrowest owning asset and must preserve the current
|
||||
distinction between registry grounding and transcript evidence.
|
||||
|
||||
Prompt ordering and cache controls should remain unchanged unless the new
|
||||
contextual input requires an intentional manifest change. Unrelated shared
|
||||
prompt bytes should not be edited. Prompt and schema fingerprints should
|
||||
invalidate only the operations whose selected assets or mapping semantics
|
||||
changed.
|
||||
|
||||
## Code And Validation Changes
|
||||
|
||||
The occurrence extractors need private response types and deterministic
|
||||
registry-resolution paths appropriate to their domain. Shared code is
|
||||
appropriate only for demonstrated mechanics that have identical semantics;
|
||||
NPC, item, and location ambiguity policies must not be forced behind a generic
|
||||
resolver merely to reduce line count.
|
||||
|
||||
Registry projections should expose explicit model-facing methods whose names
|
||||
describe whether they are names-only or contextual identity projections. The
|
||||
existing ID/name projections may remain only for deterministic consumers that
|
||||
genuinely require them; they must no longer be wired to an LLM input.
|
||||
|
||||
Mapping-policy and normalization-policy identifiers must be reviewed and
|
||||
advanced wherever their semantics change. Checkpoint fingerprints must cover
|
||||
the new projection content, private schema, prompt assets, and mapping policy,
|
||||
while continuing to exclude irrelevant internal implementation details.
|
||||
|
||||
Durable occurrence normalizers and registry validators remain defense in
|
||||
depth. They continue to validate exact ID/name pairs on artifacts entering
|
||||
through checkpoints, codecs, or other boundaries even though the LLM no longer
|
||||
produces the ID directly.
|
||||
|
||||
## Testing And Evaluation
|
||||
|
||||
Tests should protect the behavioral boundary rather than prompt prose or
|
||||
private helper structure. The completed work should demonstrate that:
|
||||
|
||||
- affected model-facing registry projections do not contain durable entity
|
||||
IDs;
|
||||
- private occurrence schemas reject opaque ID fields and accept the intended
|
||||
contextual shape;
|
||||
- valid contextual selections map to the exact canonical durable ID/name pair;
|
||||
- unknown, mismatched, and ambiguous selections fail without fuzzy matching,
|
||||
partial acceptance, or arbitrary reassignment;
|
||||
- same-name locations remain distinguishable through contextual evidence;
|
||||
- reconciliation preserves equal-name candidates, resolves valid contextual
|
||||
groups, and discards ambiguous or unsafe proposals;
|
||||
- registry evidence never becomes occurrence evidence;
|
||||
- durable codec, normalization, and validator behavior remains compatible; and
|
||||
- representative assembled D&D pipelines still prepare and execute with fake
|
||||
structured-LLM responses.
|
||||
|
||||
Do not add repository-wide prompt-prose snapshots, exact-message-count tests,
|
||||
or a change-detector test that merely scans for today's field names. Focused
|
||||
projection, schema, mapping, fallback, and integration tests are the stable
|
||||
owners of these risks. Model-quality evaluation with representative
|
||||
transcripts remains a manual development aid rather than an offline test gate.
|
||||
|
||||
## Documentation And Architectural Record
|
||||
|
||||
This policy is durable and applies to future modules, so it warrants
|
||||
`docs/adr/0012-resolve-opaque-entity-identifiers-deterministically.md`, which
|
||||
records:
|
||||
|
||||
- the semantic-proposal versus referential-integrity boundary;
|
||||
- why durable opaque IDs are excluded from model response contracts;
|
||||
- why contextual evidence coordinates remain permitted;
|
||||
- the narrowly scoped request-local-label exception;
|
||||
- alternatives including durable IDs, names-only matching, and short opaque
|
||||
handles; and
|
||||
- the consequences for private schemas, deterministic resolution, debugging,
|
||||
and ambiguous identities.
|
||||
|
||||
`docs/policy/architecture.md` states the general LLM boundary invariant and
|
||||
links to the ADR. `docs/internal/dnd.md` describes the concrete occurrence
|
||||
projections, contextual reconciliation selectors, resolution and failure
|
||||
behavior, and the continued separation of registry grounding from occurrence
|
||||
evidence. `docs/internal/llm.md` contains only a short clarification that
|
||||
caller-owned modules, not PromptKit or the transport adapter, resolve
|
||||
contextual model selections into application identities.
|
||||
|
||||
The NPC, item, and location occurrence and registry integration documents must
|
||||
continue to own their durable wire contracts, while removing current claims
|
||||
that the model-facing consumer projection contains `{id,name}` or that the raw
|
||||
LLM response supplies the durable ID. They should instead explain that
|
||||
Notarius resolves contextual model output and publishes the same exact durable
|
||||
ID/name pair. No public schema examples need to remove those IDs.
|
||||
|
||||
The generic LLM-assisted deduplication entry in `docs/roadmap/future.md` must be
|
||||
reconciled with this policy: stable IDs may exist inside deterministic state,
|
||||
but a future model-facing proposal should use contextual selectors or a
|
||||
documented request-local-label exception rather than durable IDs.
|
||||
|
||||
## Compatibility And Operational Effects
|
||||
|
||||
This work intentionally changes private prompt inputs, private structured
|
||||
responses, and mapping semantics. It will invalidate affected checkpoints
|
||||
through existing prompt, schema, projection, and policy fingerprints. No
|
||||
manual checkpoint migration is required.
|
||||
|
||||
Durable D&D artifacts and generated-reference compatibility remain unchanged.
|
||||
Operators do not receive new configuration fields or CLI controls. The feature
|
||||
does not change PromptKit, provider routing, profile selection, retries,
|
||||
concurrency, or public output placement.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
This work does not:
|
||||
|
||||
- remove canonical IDs from durable registries or occurrence artifacts;
|
||||
- change occurrence categories, evidence rules, or registry identity policy;
|
||||
- add fuzzy, probabilistic, or embedding-based entity resolution;
|
||||
- allow registry provenance to substitute for occurrence evidence;
|
||||
- introduce a general entity graph or cross-artifact identity framework;
|
||||
- redesign unrelated D&D prompts or their schemas;
|
||||
- implement the future generic deduplication normalizer; or
|
||||
- add provider-specific prompt behavior.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
The target state is complete when:
|
||||
|
||||
- no maintained D&D prompt requires a model to reproduce a durable opaque
|
||||
entity ID;
|
||||
- current D&D reconciliation prompts no longer require opaque candidate keys;
|
||||
- NPC, item, and location occurrence LLM outputs are resolved
|
||||
deterministically into their unchanged durable artifacts;
|
||||
- same-name location and reconciliation cases remain safe and unambiguous;
|
||||
- invalid contextual selections preserve the existing extraction-failure or
|
||||
normalization-fallback semantics appropriate to their stage;
|
||||
- affected checkpoint identities change without altering public schema
|
||||
versions;
|
||||
- focused and repository-wide tests pass offline;
|
||||
- the ADR, architecture invariant, D&D internal guide, LLM internal guide,
|
||||
relevant integration contracts, and future roadmap accurately describe
|
||||
their canonical portions of the implemented policy; and
|
||||
- no unrelated code, prompt behavior, or public contract changes are included.
|
||||
@@ -1,612 +0,0 @@
|
||||
# Contextual Entity Grounding Implementation Plan
|
||||
|
||||
## Objective
|
||||
|
||||
Implement [Contextual Entity Grounding](contextual-entity-grounding.md) so D&D
|
||||
LLM prompts return evidence-grounded contextual selectors while Notarius owns
|
||||
canonical entity IDs and referential integrity. Preserve every durable D&D
|
||||
artifact contract and remove opaque IDs only from model-visible inputs and
|
||||
private model responses.
|
||||
|
||||
This plan is written for a gpt-5.6-terra coding agent. Implement the stages in
|
||||
numeric order. Each stage is intentionally scoped to one implementation prompt
|
||||
and must leave the repository buildable and its focused tests passing before
|
||||
the next stage begins.
|
||||
|
||||
## Plan-Wide Decisions
|
||||
|
||||
Apply these decisions throughout every stage:
|
||||
|
||||
- Read `docs/development.md`, all files under `docs/policy/`, the feature
|
||||
roadmap, and the focused implementation/tests named by the stage before
|
||||
editing.
|
||||
- Preserve the fixed pipeline, typed artifact boundaries, root `assets`
|
||||
content-only rule, module ownership, PromptKit boundary, and evidence rules.
|
||||
- Do not change the durable NPC-, item-, location-registry, or occurrence Go
|
||||
types, JSON schemas, schema IDs, schema versions, media types, reference-slot
|
||||
contracts, categories, or generated-reference compatibility.
|
||||
- Keep the affected prompt and private response-schema identities at `v1`.
|
||||
They are private pre-release transport contracts; their changed content
|
||||
hashes provide the required compatibility boundary.
|
||||
- Advance semantic policy identifiers exactly as directed in each stage. Do
|
||||
not bump unrelated policy identifiers.
|
||||
- A contextual name match uses the entity family's existing comparison policy,
|
||||
never fuzzy matching. A model selection must resolve to exactly one supplied
|
||||
record before a durable ID is attached.
|
||||
- An invalid NPC, item, or location selection invalidates the complete
|
||||
extraction operation. Do not silently drop one response record, accept a
|
||||
partial artifact, or defer a known mapping failure to a later validator.
|
||||
- Registry provenance remains grounding only. Only the current extraction
|
||||
chunk's `source_refs` become occurrence evidence.
|
||||
- Preserve prompt message order and cache controls unless a stage explicitly
|
||||
directs otherwise. Edit only the selected module or shared assets; do not
|
||||
rewrite unrelated shared prompt bytes.
|
||||
- Preserve internal opaque IDs where deterministic code needs them. The rule
|
||||
applies to material shown to the model or requested from it, not to maps,
|
||||
fingerprints, checkpoints, diagnostics, or durable artifacts.
|
||||
- Follow `docs/policy/testing.md`: test package-level behavior and meaningful
|
||||
failure modes, not prompt prose, exact message counts, private helper
|
||||
structure, or a repository-wide string-scanning change detector. All tests
|
||||
remain deterministic, offline, and credential-free.
|
||||
- Use `apply_patch` for edits, `gofmt` changed Go files, and preserve unrelated
|
||||
worktree changes.
|
||||
|
||||
## Final Private Selector Contracts
|
||||
|
||||
These shapes are implementation requirements, not public artifact schemas.
|
||||
|
||||
### NPC occurrence response
|
||||
|
||||
Each response record contains exactly the required fields `name`, `kind`, and
|
||||
`source_refs`. It does not contain `npc_id`. Notarius resolves `name` through
|
||||
the normalized NPC registry and writes the matched registry record's `ID` and
|
||||
canonical `Name` into the durable occurrence.
|
||||
|
||||
### Item occurrence response
|
||||
|
||||
Each response record contains the existing required `name`, `kind`,
|
||||
`quantity`, `from`, `to`, and `source_refs` fields. It does not contain
|
||||
`item_id`. Retain the current nullable representation and kind-specific
|
||||
semantics. Notarius resolves `name` through the normalized item registry and
|
||||
adds the matched `ID` and canonical `Name`.
|
||||
|
||||
### Location occurrence response
|
||||
|
||||
Each response record contains exactly the required fields `name`,
|
||||
`registry_refs`, `kind`, and `source_refs`. `registry_refs` is always an array
|
||||
of strict objects containing required integer `start_unit_id` and
|
||||
`end_unit_id`; it may be empty.
|
||||
|
||||
- When `name` has one comparison-identity match in the supplied registry,
|
||||
`registry_refs` must be empty and name resolution selects that record.
|
||||
- When multiple registry records share the comparison identity,
|
||||
`registry_refs` must equal one record's complete canonically ordered source
|
||||
ranges with `source_id` removed.
|
||||
- The model-facing registry projection uses the same `name` plus
|
||||
`registry_refs` selector and adds a required `context` array. Unique-name
|
||||
records project empty `registry_refs` and `context` arrays. The private
|
||||
response does not reproduce `context`.
|
||||
- Same-name records receive bounded identity context consisting of the ordered
|
||||
source units covered by their registry ranges. Each context element is a
|
||||
strict object with exactly the required fields `unit_id` (integer) and
|
||||
`text` (string); do not expose the durable location ID, source ID, digest,
|
||||
or a replacement token.
|
||||
- Building same-name grounding validates that every registry range belongs to
|
||||
and resolves against the current source. If two records still produce the
|
||||
same contextual selector, grounding construction fails before the LLM call.
|
||||
- `registry_refs` never flow into the durable occurrence's `source_refs`.
|
||||
|
||||
### Entity-reconciliation response
|
||||
|
||||
The shared response remains an object with required `duplicate_groups`.
|
||||
Every group has required `members` and `canonical`. A member and the canonical
|
||||
selection are strict contextual objects containing:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "Mira Thorn",
|
||||
"source_refs": [
|
||||
{"start_unit_id": 12, "end_unit_id": 12}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The candidate prompt input uses the same descriptor and contains no `key`.
|
||||
`source_refs` is required and non-empty for every eligible candidate. The
|
||||
shared helper may retain its existing `candidate-000001`-style keys strictly
|
||||
inside Go state to preserve input-position mapping; those keys must never be
|
||||
serialized into prompt input or accepted in the private response.
|
||||
|
||||
If two candidates produce an identical contextual descriptor, neither is
|
||||
eligible for model-assisted reconciliation because the model cannot identify
|
||||
them independently. Otherwise the helper converts returned descriptors to its
|
||||
internal candidate keys before applying all existing unknown-member,
|
||||
ineligible-member, duplicate-member, canonical-membership, overlap, retry, and
|
||||
fallback rules.
|
||||
|
||||
## Stage 1: Convert NPC Occurrences To Name-Based Resolution
|
||||
|
||||
### Goal
|
||||
|
||||
Remove durable NPC IDs from the NPC-occurrence prompt and private response,
|
||||
then resolve the model's contextual name deterministically without weakening
|
||||
checkpoint identity or downstream validation.
|
||||
|
||||
### Work
|
||||
|
||||
1. Inspect:
|
||||
- `assets/dnd/npc-occurrences/`;
|
||||
- `internal/modules/dnd/extract/npcoccurrences/`;
|
||||
- `internal/modules/dnd/npcs/registry/`;
|
||||
- NPC-occurrence normalizer and validator checkpoint fingerprints; and
|
||||
- their focused tests.
|
||||
2. Change `dnd_npc_occurrences_llm.v1.json` so every occurrence requires only
|
||||
`name`, `kind`, and `source_refs`, continues to reject unknown fields, and
|
||||
no longer declares `npc_id`.
|
||||
3. Revise the NPC-occurrence instructions to require a supplied canonical NPC
|
||||
name and current-chunk evidence, with no instruction to copy or invent an
|
||||
ID. Continue using the existing shared names-only NPC registry fragment and
|
||||
preserve manifest order/cache controls.
|
||||
4. Remove `NPCID` from the private `occurrenceResponse`. After canonicalizing
|
||||
response evidence, resolve every response name with the existing
|
||||
`npcregistry.Registry.Lookup` comparison-key lookup. On the first unknown
|
||||
or non-unique selection, return an extractor-scoped mapping error and no
|
||||
value. For a match, construct the durable occurrence with the registry
|
||||
record's exact `ID` and canonical `Name`.
|
||||
5. Change `mappingPolicy` to
|
||||
`dnd.npc_occurrences.extract_mapping.v3`.
|
||||
6. Stop passing `IdentityPromptInput()` to the LLM; use the existing
|
||||
names-only `PromptInput()`.
|
||||
7. Replace the misleading exported model-input API used only for identity
|
||||
fingerprints: retain the unexported ordered `{id,name}` projection and its
|
||||
digest, expose that value as `IdentityDigest() string`, remove
|
||||
`IdentityPromptInput()`, and update NPC-occurrence extractor, normalizer,
|
||||
invariant-validator, and registry-validator fingerprints to use
|
||||
`IdentityDigest()`. The digest must still distinguish ID/name identity from
|
||||
the names-only prompt projection.
|
||||
8. Rewrite existing focused tests around observable behavior: rendered NPC
|
||||
registry input is names-only; the private schema rejects `npc_id`; valid
|
||||
names acquire the registry ID; comparison-equivalent names canonicalize;
|
||||
unknown names fail the whole extraction; registry identity fingerprints
|
||||
remain distinct and defensive; empty registries accept only empty model
|
||||
results. Remove tests whose only purpose was requiring the model to return
|
||||
exact ID/name pairs.
|
||||
|
||||
### Acceptance Criteria
|
||||
|
||||
- No NPC-occurrence LLM input or private response contains a durable NPC ID.
|
||||
- Durable NPC occurrences still contain the exact registry ID/name pair.
|
||||
- Mapping failures remain extractor failures eligible for the configured
|
||||
pipeline retry behavior.
|
||||
- Deterministic consumers still fingerprint the ordered registry identity,
|
||||
while spells, combat turns, enemy events, and NPC occurrences share the
|
||||
names-only model projection.
|
||||
|
||||
### Validation
|
||||
|
||||
```sh
|
||||
go fmt ./internal/modules/dnd/npcs/registry ./internal/modules/dnd/extract/npcoccurrences ./internal/modules/dnd/normalize/npcoccurrences ./internal/modules/dnd/validate/npcoccurrences/...
|
||||
go test ./internal/modules/dnd/npcs/registry ./internal/modules/dnd/extract/npcoccurrences ./internal/modules/dnd/normalize/npcoccurrences ./internal/modules/dnd/validate/npcoccurrences/...
|
||||
```
|
||||
|
||||
This stage is suitable for one gpt-5.6-terra prompt.
|
||||
|
||||
## Stage 2: Convert Item Occurrences To Name-Based Resolution
|
||||
|
||||
### Goal
|
||||
|
||||
Give item occurrences the same contextual-name/deterministic-ID boundary while
|
||||
preserving item-specific nullable fields and kind rules.
|
||||
|
||||
### Work
|
||||
|
||||
1. Inspect `assets/dnd/item-occurrences/`, the item occurrence extractor, the
|
||||
item registry, the item occurrence normalizer and registry validator, and
|
||||
their focused tests.
|
||||
2. Change `dnd_item_occurrences_llm.v1.json` to remove `item_id` from required
|
||||
fields and properties. Preserve required `name`, `kind`, `quantity`, `from`,
|
||||
`to`, and `source_refs`, all current enums/nullability, and strict unknown
|
||||
field rejection.
|
||||
3. Rewrite the item registry fragment and module instructions to require the
|
||||
supplied canonical name and current-chunk evidence without mentioning an
|
||||
ID. Preserve prompt order and cache controls.
|
||||
4. Change the item registry's model projection from ordered `{id,name}` pairs
|
||||
to ordered names-only objects, add a comparison-key index, and expose a
|
||||
defensive `Lookup(name) (dnd.Item, bool)` analogous to the NPC registry.
|
||||
Retain exact `LookupID` for durable normalizers and validators. Because item
|
||||
IDs are derived from the item comparison identity, the names-only
|
||||
`ProjectionDigest` remains sufficient for model input and existing
|
||||
checkpoint consumers.
|
||||
5. Remove `ItemID` from the private response. During response canonicalization,
|
||||
resolve every contextual name, replace it with the registry record's
|
||||
canonical name, and attach its durable ID when constructing the final
|
||||
`dnd.ItemOccurrence`. Unknown selections fail the complete extraction; do
|
||||
not alter evidence or nullable-field validation ownership.
|
||||
6. Change `mappingPolicy` to
|
||||
`dnd.item_occurrences.extract_mapping.v2`.
|
||||
7. Update focused tests to cover names-only projection, defensive comparison
|
||||
lookup, schema rejection of `item_id`, deterministic durable mapping,
|
||||
unknown-name failure after an otherwise valid record, empty registry/result
|
||||
behavior, and preservation of nullable/kind-specific fields.
|
||||
|
||||
### Acceptance Criteria
|
||||
|
||||
- Model-visible item registry and response content contain no item hash.
|
||||
- Every accepted durable occurrence has the matched registry ID and canonical
|
||||
name.
|
||||
- Invalid selection remains all-or-nothing, and existing normalizer/validator
|
||||
defense in depth remains unchanged.
|
||||
|
||||
### Validation
|
||||
|
||||
```sh
|
||||
go fmt ./internal/modules/dnd/items/registry ./internal/modules/dnd/extract/itemoccurrences
|
||||
go test ./internal/modules/dnd/items/registry ./internal/modules/dnd/extract/itemoccurrences ./internal/modules/dnd/normalize/itemoccurrences ./internal/modules/dnd/validate/itemoccurrences/...
|
||||
```
|
||||
|
||||
This stage is suitable for one gpt-5.6-terra prompt.
|
||||
|
||||
## Stage 3: Add Contextual Location Grounding
|
||||
|
||||
### Goal
|
||||
|
||||
Replace the location registry's ID/name prompt projection with an immutable,
|
||||
source-aware grounding object that can represent same-name locations safely.
|
||||
Introduce the new path alongside the old occurrence input so this stage remains
|
||||
buildable; Stage 4 performs the atomic extractor cutover and removes the old
|
||||
path.
|
||||
|
||||
### Work
|
||||
|
||||
1. Inspect the location registry, location identity and source-reference
|
||||
helpers, the generic source document index, occurrence checkpoint consumers,
|
||||
and their focused tests.
|
||||
2. In `internal/modules/dnd/locations/registry`, define the private-model
|
||||
types needed by both grounding and the location occurrence extractor:
|
||||
- a returned selector with exactly `name` and `registry_refs`;
|
||||
- a registry projection entry with exactly `name`, `registry_refs`, and
|
||||
`context`;
|
||||
- a source-free range with exactly `start_unit_id` and `end_unit_id`; and
|
||||
- a context unit with exactly `unit_id` and `text`.
|
||||
All fields are required in their private JSON shapes, and constructors and
|
||||
accessors must make defensive copies.
|
||||
3. Add an operation-scoped immutable grounding type constructed from a resolved
|
||||
registry and the current `*source.SourceDocument`. Its API must provide:
|
||||
- a cloned `contracts.LLMInputMaterial` for the `location_registry` slot;
|
||||
- deterministic resolution of a returned selector to one cloned
|
||||
`dnd.Location`.
|
||||
The prompt material's existing `Digest` field owns the digest of the exact
|
||||
model projection; do not expose a second grounding-specific digest API.
|
||||
4. Construct the projection in registry order. Group entries by the existing
|
||||
location comparison key:
|
||||
- every projection entry has exactly the required fields `name`,
|
||||
`registry_refs`, and `context`;
|
||||
- comparison-unique entries use empty `registry_refs` and `context` arrays;
|
||||
- every same-name entry uses its complete canonical source ranges stripped
|
||||
of `source_id` and includes ordered context units covered by those ranges;
|
||||
- each context unit contains exactly required integer `unit_id` and string
|
||||
`text` fields, and units are deduplicated in source order; and
|
||||
- same-name ranges must have `SourceID == doc.ID` and pass
|
||||
`source.DocumentIndex.ValidateRef`.
|
||||
5. Fail grounding construction with a bounded, content-safe error if the
|
||||
source is nil, a same-name range is invalid or belongs to another source,
|
||||
a comparison key is empty, or two records produce the same selector. Do not
|
||||
expose transcript text in the error.
|
||||
6. Resolution uses the existing comparison key. A unique-name selector is
|
||||
accepted only with empty `registry_refs`; a same-name selector is accepted
|
||||
only on an exact canonical range match. Reject unknown names, a non-empty
|
||||
range list for a unique name, an empty/partial/reordered range list for an
|
||||
ambiguous name, or any selector not present in the grounding.
|
||||
7. Separate deterministic identity fingerprinting from LLM material. Add
|
||||
`IdentityDigest()` over the registry's ordered `{id,name}` identity
|
||||
projection, and update the location normalizer and registry-validator
|
||||
checkpoint consumers to use it. The new operation grounding carries the
|
||||
model projection digest in its `LLMInputMaterial`. Retain the old ID-bearing
|
||||
prompt accessor only as a documented transitional dependency of the
|
||||
still-unchanged location occurrence extractor; do not add new callers.
|
||||
8. Add focused tests for unique names, same-name context and selectors,
|
||||
canonical range order, deterministic projection/digest, defensive copies,
|
||||
exact selector resolution, nil/foreign/invalid references, selector
|
||||
collisions, empty registries, and identity fingerprint stability. Do not
|
||||
assert large rendered prompt strings; decode the JSON projection and assert
|
||||
its semantic shape.
|
||||
|
||||
### Acceptance Criteria
|
||||
|
||||
- The location package can build and resolve contextual selectors without
|
||||
exposing `location_id`, `source_id`, digests, or replacement labels.
|
||||
- Same-name locations remain distinct and receive meaningful bounded context.
|
||||
- Deterministic checkpoint consumers retain an ID-sensitive fingerprint.
|
||||
- Only the existing occurrence extractor remains wired to the legacy
|
||||
ID-bearing prompt path until Stage 4; the repository compiles and tests pass.
|
||||
|
||||
### Validation
|
||||
|
||||
```sh
|
||||
go fmt ./internal/modules/dnd/locations/registry ./internal/modules/dnd/normalize/locationoccurrences ./internal/modules/dnd/validate/locationoccurrences/...
|
||||
go test ./internal/modules/dnd/locations/registry ./internal/modules/dnd/normalize/locationoccurrences ./internal/modules/dnd/validate/locationoccurrences/...
|
||||
```
|
||||
|
||||
This stage is suitable for one gpt-5.6-terra prompt.
|
||||
|
||||
## Stage 4: Convert Location Occurrences To Contextual Resolution
|
||||
|
||||
### Goal
|
||||
|
||||
Wire the Stage 3 grounding object into location occurrence extraction and
|
||||
remove durable location IDs from the prompt and private response.
|
||||
|
||||
### Work
|
||||
|
||||
1. Inspect `assets/dnd/location-occurrences/`, the location occurrence model,
|
||||
schema loader, extractor, canonicalization, prompt tests, and Stage 3
|
||||
grounding tests.
|
||||
2. Change `dnd_location_occurrences_llm.v1.json` so each occurrence requires
|
||||
exactly `name`, `registry_refs`, `kind`, and `source_refs`; remove
|
||||
`location_id`. Keep all four occurrence kinds. Make `registry_refs` a
|
||||
required array, including an empty array, of strict required positive
|
||||
integer ranges. Keep occurrence `source_refs` separate and unchanged.
|
||||
3. Rewrite `location-registry.md` and module instructions to explain the two
|
||||
selector cases, require exact supplied contextual selectors, prohibit
|
||||
invented locations, and state that registry ranges/context are identity
|
||||
grounding rather than occurrence evidence. Preserve manifest order and
|
||||
cache controls.
|
||||
4. Change the private response type to `Name`, `RegistryRefs`, `Kind`, and
|
||||
`SourceRefs`. Do not reuse `source.SourceRef` for the source-free registry
|
||||
range type.
|
||||
5. In `Extract`, construct operation grounding from the resolved registry and
|
||||
`req.Source` before calling the LLM, put its projection in the
|
||||
`location_registry` input, and resolve every returned selector after
|
||||
completion. Attach the selected registry record's exact ID and canonical
|
||||
name to the durable occurrence while retaining only the response's
|
||||
current-source `source_refs` as evidence.
|
||||
6. Fail the whole extraction on grounding-construction failure or the first
|
||||
unknown, malformed, mismatched, or ambiguous selector. This replaces the
|
||||
current behavior that can preserve unknown ID/name pairs for later
|
||||
validators. Keep later normalizer and validator checks as defense in depth
|
||||
for artifacts entering other boundaries.
|
||||
7. Change `mappingPolicy` to
|
||||
`dnd.location_occurrences.extract_mapping.v2`.
|
||||
8. Remove the legacy ID-bearing registry `PromptInput` and its model-projection
|
||||
digest once the extractor uses operation grounding. Keep the occurrence's
|
||||
static module fingerprint based on `IdentityDigest()`, prompt/schema
|
||||
fingerprints, and mapping policy. The operation-scoped projection is already
|
||||
covered by source/chunk identity and configured or generated reference
|
||||
dependencies, while its `LLMInputMaterial.Digest` identifies the exact model
|
||||
input; do not add a second digest API or operation-aware static fingerprint.
|
||||
9. Update focused schema, prompt, extractor, canonicalization, checkpoint, and
|
||||
generated-reference tests. Cover unique-name empty selectors, successful
|
||||
same-name selection, failure for an unsupported ambiguous mention,
|
||||
partial/reordered ranges, no registry-to-occurrence evidence leakage,
|
||||
all-or-nothing failure, empty registry/result behavior, and unchanged
|
||||
durable ordering/deduplication.
|
||||
|
||||
### Acceptance Criteria
|
||||
|
||||
- The location prompt and private response contain no durable location ID.
|
||||
- Unique and same-name records resolve according to the final selector
|
||||
contract.
|
||||
- Accepted durable output is unchanged in shape and still contains an exact
|
||||
location ID/name pair.
|
||||
- Registry context cannot become durable occurrence evidence.
|
||||
|
||||
### Validation
|
||||
|
||||
```sh
|
||||
go fmt ./internal/modules/dnd/extract/locationoccurrences ./internal/modules/dnd/locations/registry
|
||||
go test ./internal/modules/dnd/locations/registry ./internal/modules/dnd/extract/locationoccurrences ./internal/modules/dnd/normalize/locationoccurrences ./internal/modules/dnd/validate/locationoccurrences/...
|
||||
```
|
||||
|
||||
This stage is suitable for one gpt-5.6-terra prompt.
|
||||
|
||||
## Stage 5: Replace Reconciliation Keys With Contextual Descriptors
|
||||
|
||||
### Goal
|
||||
|
||||
Change the shared NPC/item/location registry-normalization proposal contract so
|
||||
opaque candidate keys remain internal and the model sees and returns only
|
||||
names plus evidence coordinates.
|
||||
|
||||
### Work
|
||||
|
||||
1. Inspect:
|
||||
- `internal/modules/dnd/shared/entityreconcile/`;
|
||||
- `assets/dnd/shared/prompts/common-dnd-entity-reconciliation.md`;
|
||||
- `assets/dnd/entity-reconciliation/schemas/`;
|
||||
- all three registry normalization manifests and prompt tests; and
|
||||
- the NPC, item, and location registry normalizers and reconciliation tests.
|
||||
2. Introduce one exported, defensively copied contextual selector type in
|
||||
`entityreconcile` with JSON `name` and `source_refs`, plus a strict
|
||||
source-free range type. Use it for candidate input views and for
|
||||
`DuplicateGroup.Members` and `.Canonical`.
|
||||
3. Keep deterministic candidate keys only inside `Materials`. During
|
||||
`BuildContext`, validate and canonicalize candidate references as today,
|
||||
serialize candidate views without `key`, derive a stable internal lookup
|
||||
from canonical selector JSON to the corresponding internal candidate key,
|
||||
and detect descriptor collisions before eligibility is established.
|
||||
Colliding candidates must not appear in the prompt input or become
|
||||
eligible; their records remain in deterministic normalization output.
|
||||
4. Update `Materials.Assess` to resolve every returned selector through that
|
||||
internal lookup before running the existing group assessment. Preserve
|
||||
existing issue categories where their meaning still applies. Treat an
|
||||
unknown or collided descriptor as an unknown/ineligible selection, discard
|
||||
only the affected group, and retain existing overlap handling. `SafeGroup`
|
||||
may continue returning internal candidate keys so the three domain
|
||||
normalizers retain their position mapping; those keys are not model-facing.
|
||||
5. Rewrite `dnd_entity_reconcile_llm.v1.json` so members and canonical are
|
||||
strict selector objects. Require non-empty `name` structurally where the
|
||||
current schemas do so, require `source_refs`, and make each range strict
|
||||
with required positive integer endpoints. Preserve `duplicate_groups` and
|
||||
the existing semantic assessment of minimum group size, membership,
|
||||
duplicates, eligibility, and overlap rather than moving every semantic
|
||||
failure into JSON Schema.
|
||||
6. Rewrite the shared reconciliation fragment to tell the model to return
|
||||
supplied contextual descriptors and never invent names or ranges. Remove
|
||||
every instruction about opaque keys. Preserve all three manifests' message
|
||||
ordering and cache controls.
|
||||
7. Change registry normalization policy identifiers to:
|
||||
- `dnd.npc_registry.normalize.v4`;
|
||||
- `dnd.item_registry.normalize.v2`; and
|
||||
- `dnd.location_registry.normalize.v2`.
|
||||
8. Update shared and domain tests to cover candidate JSON without keys,
|
||||
contextual proposal decoding, valid selector-to-internal-key mapping,
|
||||
equal names with different evidence, descriptor collision exclusion,
|
||||
unknown/partial/reordered descriptors, overlapping groups, canonical
|
||||
membership, invalid structured-output fallback, currency safety, and
|
||||
preservation of every non-applied deterministic candidate. Update prompt
|
||||
asset fixtures to the new selector schema; do not snapshot prompt prose.
|
||||
|
||||
### Acceptance Criteria
|
||||
|
||||
- No registry normalization prompt input or private response contains a
|
||||
`candidate-*` key.
|
||||
- Internal keys remain inaccessible to the model but may still support safe
|
||||
deterministic position mapping.
|
||||
- All existing normalizer safety, retry, fallback, warning, currency, and
|
||||
same-name-location policies remain intact.
|
||||
- Identical contextual descriptors cannot be arbitrarily reconciled.
|
||||
|
||||
### Validation
|
||||
|
||||
```sh
|
||||
go fmt ./internal/modules/dnd/shared/entityreconcile ./internal/modules/dnd/normalize/npcregistry ./internal/modules/dnd/normalize/itemregistry ./internal/modules/dnd/normalize/locationregistry
|
||||
go test ./internal/modules/dnd/shared/entityreconcile ./internal/modules/dnd/normalize/npcregistry ./internal/modules/dnd/normalize/itemregistry ./internal/modules/dnd/normalize/locationregistry
|
||||
```
|
||||
|
||||
This is the largest stage, but it is one cohesive shared-contract migration
|
||||
and is suitable for one gpt-5.6-terra prompt when implemented exactly within
|
||||
the listed packages. Do not combine it with occurrence or documentation work.
|
||||
|
||||
## Stage 6: Record The Decision And Update Canonical Documentation
|
||||
|
||||
### Goal
|
||||
|
||||
Document the implemented policy in its durable architectural, internal, and
|
||||
integration homes without duplicating volatile details or presenting roadmap
|
||||
work as current behavior prematurely.
|
||||
|
||||
### Work
|
||||
|
||||
1. Re-read `docs/policy/documentation.md`, ADR-0003, ADR-0009, ADR-0011,
|
||||
`docs/internal/dnd.md`, `docs/internal/llm.md`, and the six affected registry
|
||||
and occurrence integration documents. Verify the code before describing it.
|
||||
2. Add
|
||||
`docs/adr/0012-resolve-opaque-entity-identifiers-deterministically.md` in
|
||||
the repository's Nygard ADR format with status `Accepted` and the actual
|
||||
implementation date. Record the model-semantic/deterministic-identity
|
||||
boundary, source-coordinate allowance, request-local-label exception,
|
||||
alternatives, ambiguity behavior, and consequences. Link ADR-0003 and
|
||||
ADR-0009 rather than repeating their complete decisions.
|
||||
3. Add a concise normative invariant under the LLM boundary in
|
||||
`docs/policy/architecture.md`: callers use contextual model selections and
|
||||
attach opaque application identities deterministically when possible. Link
|
||||
ADR-0012 for rationale.
|
||||
4. Update `docs/internal/dnd.md` to replace exact model-facing `{id,name}`
|
||||
claims with the implemented NPC/item names-only and location contextual
|
||||
selector behavior. Document reconciliation descriptors, internal-only keys,
|
||||
all-or-nothing occurrence mapping failures, normalization fallback, and the
|
||||
separation between registry and occurrence evidence. Do not duplicate the
|
||||
private JSON schemas.
|
||||
5. Add only a short ownership clarification to `docs/internal/llm.md`: the
|
||||
calling module resolves contextual selections; PromptKit and its adapter do
|
||||
not own entity identity.
|
||||
6. Update these durable integration contracts while preserving their public
|
||||
ID-bearing wire examples and schema statements:
|
||||
- `docs/integrations/dnd-npc-registry-artifacts.md`;
|
||||
- `docs/integrations/dnd-npc-occurrence-artifacts.md`;
|
||||
- `docs/integrations/dnd-item-registry-artifacts.md`;
|
||||
- `docs/integrations/dnd-item-occurrence-artifacts.md`;
|
||||
- `docs/integrations/dnd-location-registry-artifacts.md`; and
|
||||
- `docs/integrations/dnd-location-occurrence-artifacts.md`.
|
||||
Remove claims that LLM consumers receive `{id,name}` or that raw model
|
||||
output supplies an ID. State that Notarius maps contextual output into the
|
||||
unchanged exact durable pair.
|
||||
7. Revise the generic LLM-assisted deduplication entry in
|
||||
`docs/roadmap/future.md`: stable unique IDs remain internal deterministic
|
||||
state, while a future model proposal uses contextual descriptors or a
|
||||
specifically justified request-local short label.
|
||||
8. Do not change README, CLI, configuration, operations, examples, or public
|
||||
schema files; this feature has no user-selectable surface or public wire
|
||||
change.
|
||||
|
||||
### Acceptance Criteria
|
||||
|
||||
- ADR-0012 owns rationale; architecture owns the normative boundary; internal
|
||||
docs own mechanics; integration docs own unchanged durable contracts; and
|
||||
the future roadmap no longer proposes durable IDs as the default model
|
||||
selector.
|
||||
- No current-behavior document claims that a model copies hash-based entity
|
||||
IDs or opaque reconciliation keys.
|
||||
- Documentation does not duplicate private schemas or implementation history.
|
||||
|
||||
### Validation
|
||||
|
||||
```sh
|
||||
git diff --check
|
||||
rg -n '\{id,name\}|ID/name grounding|Candidate keys are opaque|candidate-[0-9]' docs assets/dnd
|
||||
```
|
||||
|
||||
Review every search result semantically; durable wire-contract ID/name
|
||||
requirements and internal test fixtures are not automatically errors.
|
||||
|
||||
This stage is suitable for one gpt-5.6-terra prompt.
|
||||
|
||||
## Stage 7: Integration Audit And Final Verification
|
||||
|
||||
### Goal
|
||||
|
||||
Verify the assembled D&D family, remove obsolete identity-copy paths, and
|
||||
finish with a clean, policy-compliant implementation.
|
||||
|
||||
### Work
|
||||
|
||||
1. Audit every maintained D&D prompt manifest, selected fragment, private
|
||||
schema, and constructed prompt projection. Confirm that no model is asked to
|
||||
reproduce `npc:sha256:...`, `item:sha256:...`,
|
||||
`location:sha256:...`, `candidate-*`, a UUID, a digest, or another opaque
|
||||
entity handle. Do not confuse runtime metadata or durable output contracts
|
||||
with model-visible material.
|
||||
2. Trace all former APIs and fields, including `IdentityPromptInput`,
|
||||
ID-bearing item/location prompt projections, private `NPCID`/`ItemID`/
|
||||
`LocationID` response fields, and model-visible candidate keys. Remove dead
|
||||
code, obsolete comments, stale test names, and unused assets. Retain
|
||||
identity-only digests and exact durable lookup APIs used by deterministic
|
||||
consumers.
|
||||
3. Review prompt fingerprint registration and checkpoint fingerprints. Confirm
|
||||
that each affected prompt/schema/policy/projection change invalidates the
|
||||
relevant operation and that unrelated D&D lanes retain their existing
|
||||
fingerprints.
|
||||
4. Run representative production registration and multi-step pipeline tests
|
||||
using existing fakes. Update only tests whose stable behavior changed.
|
||||
Confirm generated NPC/item/location registry handoffs still prepare and
|
||||
that final durable occurrences encode and validate under their existing
|
||||
`v1` codecs.
|
||||
5. Run formatting, focused suites, full tests, vet, build, and documentation
|
||||
whitespace checks. Fix only failures caused by this feature. Report any
|
||||
unrelated pre-existing failure without broadening scope.
|
||||
6. Review the feature roadmap acceptance criteria one by one. Do not delete
|
||||
`contextual-entity-grounding.md` or this implementation plan in this stage;
|
||||
roadmap retirement is a separate maintainer action after review.
|
||||
|
||||
### Acceptance Criteria
|
||||
|
||||
- All feature-roadmap acceptance criteria are met.
|
||||
- The repository contains no obsolete model-facing opaque-identity path.
|
||||
- Public artifacts and generated handoffs remain compatible.
|
||||
- Tests are focused on behavior rather than prose or implementation shape.
|
||||
- The worktree contains only intentional feature and documentation changes.
|
||||
|
||||
### Validation
|
||||
|
||||
```sh
|
||||
go fmt ./internal/modules/dnd/...
|
||||
go test ./internal/modules/dnd/...
|
||||
go test ./internal/modules/integration/...
|
||||
go test ./...
|
||||
go vet ./...
|
||||
go build ./cmd/notarius
|
||||
git diff --check
|
||||
git status --short
|
||||
```
|
||||
|
||||
This stage is suitable for one gpt-5.6-terra prompt.
|
||||
Reference in New Issue
Block a user