15 KiB
NWS AFD Section Parsing Resilience Implementation Plan
Purpose
Complete the remaining resilience work defined in
afd-section-heading-variants.md without
changing the canonical forecast-discussion contract. Stages 1-8 summarize work
already completed. Implement Stages 9-13 in order.
Cross-Stage Constraints
- Keep production parsing changes in
internal/providers/nws/forecast_discussion.go. - Use only the Go standard library and keep new helpers unexported.
- Do not change
model,standards, source configuration or polling, schema identifiers, event-envelope behavior, Postgres mapping, sinks, consumer docs, or integration contracts. - Continue exposing only key messages, short term, and long term.
- Map
KEY POINTSto the existing key-message role. KeepNEAR TERM,DISCUSSION, and every other unmapped identity boundary-only. - Preserve first-occurrence behavior by canonical role, including across
KEY MESSAGESandKEY POINTSaliases. - Preserve raw body lines during structural scanning. Apply provider-specific cleanup only while parsing a mapped block.
- Prefer small grammar, label, and list-marker helpers over a broad regular expression or general parsing framework.
- Keep parsing deterministic. Tests must use local strings and fixtures, never live NWS requests.
- Preserve all pre-existing user work and avoid unrelated refactors.
Stage 1: Centralize Known Heading Parsing — Completed
The provider parser gained a shared heading classifier and section-block model for the original key-message, short-term, long-term, and aviation identities. Discovery, boundary handling, and qualifier extraction stopped using separate per-section patterns.
Stage 2: Add Original Heading-Variant Regressions — Completed
Provider and normalizer tests covered ellipsis-first and slash-qualified headings, malformed forms, empty bodies, mixed heading families, and unchanged canonical wire shape.
Stage 3: Validate the Original Feature — Completed
Focused and repository-wide checks passed, production changes remained inside the NWS provider parser, and the original roadmap was reconciled with the implemented baseline.
Stage 4: Generalize Structural Heading Recognition — Completed
Heading syntax was decoupled from canonical roles. The parser now accepts a constrained generic uppercase identity grammar, normalizes identity whitespace, supports embedded slashes without confusing slash qualifiers, and recognizes unknown valid identities as structural headings.
Stage 5: Scan Ordered Section Blocks Once — Completed
Repeated extraction was replaced by a one-pass ordered block scanner. Generic
headings, &&, $$, and the watch/advisory safeguard terminate active blocks;
role lookup happens after scanning and retains first-occurrence behavior.
Stage 6: Normalize Multiline Preambles and Markers — Completed
Short- and long-term parsing gained next-line parenthesized qualifiers,
case-insensitive Issued at handling, and exact change-marker removal. Key
messages gained exact marker cleanup and removal of one leading issue/update
metadata line.
Stage 7: Add Initial Cross-Office Coverage — Completed
The existing LSX fixture was retained and a compact BOU-style fixture added. Provider and normalizer regressions cover multiline qualifiers, change markers, leading metadata, generic boundary isolation, envelope behavior, and unchanged JSON shape.
Stage 8: Validate the First Resilience Follow-up — Completed
Focused tests, the full repository suite, static analysis, and diff checks passed. The public model and schemas remained unchanged, and the roadmap was updated to describe the completed behavior.
Stage 9: Add Parenthesized-Terminal Heading Syntax
Recognize the observed .<IDENTITY> (<qualifier>)... family without weakening
generic identity validation.
-
In
internal/providers/nws/forecast_discussion.go, add a dedicatedparseForecastDiscussionParenthesizedTerminalHeadinghelper returning the existingforecastDiscussionSectionHeadingtype. -
Update
parseForecastDiscussionSectionHeadingto try forms in this order:- if the line ends
/..., try the slash-qualified helper and return its result; - otherwise, try the parenthesized-terminal helper for any line ending
...and return it when it succeeds; and - fall back to the existing ellipsis-first helper.
Do not merge the forms into a single regular expression.
- if the line ends
-
Implement the parenthesized-terminal helper with these exact rules:
- operate on the already outer-trimmed line and require leading
.plus a terminal ASCII...; - remove the leading dot and terminal ellipsis, then remove horizontal whitespace immediately before the ellipsis;
- require the remaining content to end with
); - locate the first
(that is preceded by horizontal whitespace; everything before that separator is the raw identity and everything from(through the final)is the qualifier; - normalize and validate the identity through the existing identity helper;
- require nonempty text after trimming inside the outer parentheses;
- return the qualifier with its outer parentheses and internal punctuation intact; and
- reject missing separator whitespace, empty qualifiers, missing or misplaced parentheses, trailing text after the ellipsis, unsupported identity punctuation, and lowercase or mixed-case identities.
Parentheses inside the qualifier are content; only the first separator
(and final)delimit the outer qualifier. - operate on the already outer-trimmed line and require leading
-
Extend the heading table in
internal/providers/nws/forecast_discussion_test.gowith:.DISCUSSION (Today through Thursday)...;- a mapped identity such as
.SHORT TERM (Tonight)...; - identity whitespace normalization;
- qualifier punctuation and nested parentheses;
- each rejected malformed form listed above; and
- compatibility assertions for every existing ellipsis-first and slash-qualified case.
-
Add a scanner regression in which a mapped section is followed immediately, without
&&, by.DISCUSSION (Today through Thursday).... Assert that the discussion heading starts a new boundary-only block and its prose cannot leak into the mapped section.
Run and pass:
gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/forecast_discussion_test.go
go test -count=1 ./internal/providers/nws
Do not proceed until all previous heading and scanner tests still pass.
Stage 10: Add Explicit Canonical Role Aliases
Make true wording synonyms cheap to support while preserving semantic distinctions.
-
Add
KEY POINTSto the existing identity-to-role registry with the key-message role. Keep the registry as the only production source of identity-to-role mappings; do not add a parallel alias collection or role-selection switch. -
Leave
KEY MESSAGES,SHORT TERM, andLONG TERMmappings unchanged. Explicitly do not add role mappings forNEAR TERM,DISCUSSION,AVIATION, or other generic identities. -
Update registry tests so they no longer assume one identity per role. Assert instead that:
- every registry identity parses as a heading;
KEY MESSAGESandKEY POINTSboth map to the key-message role;SHORT TERMandLONG TERMretain their roles; andNEAR TERM,DISCUSSION, andAVIATIONhave no role entry.
-
Add
ParseForecastDiscussionTexttests for role-level first occurrence:KEY POINTSalone populates key messages;KEY POINTSfollowed byKEY MESSAGESkeeps the first block; andKEY MESSAGESfollowed byKEY POINTSkeeps the first block.
Use hyphen messages in this stage so list-tokenization changes remain scoped to Stage 11.
-
Add a regression containing
NEAR TERM,SHORT TERM, andLONG TERMin one bulletin. Assert that near-term prose terminates adjacent blocks but does not populate or override either canonical text section.
Run and pass:
gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/forecast_discussion_test.go
go test -count=1 ./internal/providers/nws
Do not change the canonical model to expose near-term or discussion content.
Stage 11: Harden Key-Message Metadata and Item Parsing
Replace prefix-sensitive, hyphen-only parsing with conservative metadata and list classifiers.
-
Add a helper that recognizes a leading labeled metadata line only when:
- its trimmed text begins with exactly
Issued atorUpdated at, compared ASCII case-insensitively; - at least one horizontal-whitespace byte follows the label; and
- parsing the remaining text through the existing unlabeled NWS issue-time grammar succeeds.
Reuse
parseForecastDiscussionIssueTimefor the timestamp grammar after removing the label. Do not add a second timestamp parser. A label prefix with no boundary or with an invalid timestamp returns false and remains content. - its trimmed text begins with exactly
-
Update key-message cleanup to remove at most one such valid leading metadata line after presentation-marker removal and blank-line trimming. Later valid timestamp lines remain message content.
-
Replace the hyphen-only check with a helper that classifies and strips one of these markers from a trimmed line:
- one leading
-, preserving current compatibility whether or not whitespace follows it; - one leading
*, whether or not whitespace follows it; or - one or more ASCII digits forming an integer greater than zero, followed by
)or., and then either end-of-line or horizontal whitespace.
After recognizing and stripping a leading hyphen or asterisk plus horizontal whitespace, also strip one immediately following valid numeric marker. This supports composite forms such as
- 1.without retaining either decorator. Otherwise strip only the recognized marker and following horizontal whitespace. Reject zero, overflow, alphanumeric prefixes, and numeric punctuation without the required boundary. Do not interpret these markers outside key-message parsing. - one leading
-
Parse a cleaned key-message body deterministically:
- first determine whether any nonempty line has a recognized marker;
- when markers exist, each marker starts a new message; nonempty unmarked lines after a marker continue that message; blank lines after the first marker are ignored; and blank-line-separated prose before the first marker is flushed as preserved message content;
- when no marker exists, each nonempty paragraph separated by one or more blank lines becomes one message; and
- in both modes, join wrapped lines with one ASCII space, discard empty messages, and preserve source order.
-
Add table-driven marker tests for
-,*,1),2., composite- 1.and* 1), multi-digit values, wrapped continuations, marker-only lines, and rejected zero/malformed numeric prefixes. -
Add block-level tests covering:
- numbered and asterisk lists producing distinct messages;
- unmarked paragraphs producing distinct messages;
- mixed introductory prose and marked items without data loss;
- valid leading
Issued atandUpdated atmetadata removal; Updated atmospheric conditions...remaining content;Updated at not a timestampremaining content;- a later valid timestamp-like line remaining content; and
- all existing hyphen, change-marker, and continuation behavior.
Run and pass:
gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/forecast_discussion_test.go
go test -count=1 ./internal/providers/nws
Stage 12: Add Representative Numbered and Key-Points Fixtures
Prove the new extension points through provider and normalizer boundaries.
-
Retain the existing LSX and BOU fixtures and their regressions.
-
Add two compact HTML fixtures under
internal/providers/nws/testdata/:- a BGM/CTP-style fixture with a valid header and issue time, numbered
KEY MESSAGES, wrapped item lines, and a boundary-onlyDISCUSSIONsection; and - an MFR-style fixture with a valid header and issue time, asterisk
KEY POINTS, and an immediately following.DISCUSSION (Today through Thursday)...section without an intervening&&.
Use concise representative text rather than full web pages. Add a short HTML comment to each fixture naming the real office-format family it represents; do not claim that edited fixture prose is a verbatim archived product.
- a BGM/CTP-style fixture with a valid header and issue time, numbered
-
Add provider end-to-end tests that assert:
- exact ordered key messages and joined continuations;
KEY POINTSalias mapping;- correct office and top-level issue metadata;
- absence of numeric/asterisk markers from canonical values;
- absence of
DISCUSSIONheadings, qualifiers, and prose from key messages; and - nil short- and long-term fields when the fixture contains neither mapped role.
-
Add table-driven normalizer regressions using both fixtures. For each, verify input envelope preservation, event kind, canonical schema, effective time, exact key-message mapping, and successful JSON marshaling.
-
Assert the JSON payload still has no
nearTerm,discussion,aviation, or genericsectionsfield. Do not add those fields to provider or canonical structs. -
Keep syntax and tokenizer edge cases in focused unit tables; do not add more full fixtures for cases already proven locally.
Run and pass:
gofmt -w internal/providers/nws/forecast_discussion_test.go internal/normalizers/nws/forecast_discussion_test.go
go test -count=1 ./internal/providers/nws ./internal/normalizers/nws
Stage 13: Reconcile Documentation and Perform Final Validation
Close the second resilience follow-up only after all behavior is implemented and verified.
-
Review the final diff. Production changes must remain confined to
internal/providers/nws/forecast_discussion.go; other Go changes must be owning provider and normalizer tests. Preserve unrelated user work. -
Confirm the public contract is unchanged. If implementation appears to require a canonical near-term/discussion field, schema change, configuration, or downstream migration, stop and report the conflict rather than expanding scope.
-
Run:
gofmt -w \ internal/providers/nws/forecast_discussion.go \ internal/providers/nws/forecast_discussion_test.go \ internal/normalizers/nws/forecast_discussion_test.go go test -count=1 ./internal/providers/nws ./internal/normalizers/nws go test -count=1 ./... go vet ./... git diff --check -
Verify every acceptance criterion in the feature roadmap has a corresponding passing test, including compatibility with Stages 1-8.
-
Update the feature roadmap status to
Implemented.and rewrite its remaining proposed or future-tense language as implemented behavior only after all checks pass. -
Update this implementation plan so Stages 9-13 are marked
— Completedand the Purpose states that all stages are complete. Preserve the stage details as a historical implementation record.
Open Questions
None. The plan makes the required policy choices explicitly: KEY POINTS is a
true key-message alias; NEAR TERM remains semantically distinct and
boundary-only; parenthesized-terminal syntax is a third constrained grammar;
metadata is removed only after timestamp validation; and key-message variants
are handled by an isolated marker classifier plus paragraph fallback.