Support parenthesized NWS discussion headings

This commit is contained in:
2026-08-03 00:13:58 +00:00
parent a62cb87b78
commit b8c6708439
4 changed files with 526 additions and 296 deletions

View File

@@ -2,99 +2,154 @@
## Status ## Status
Implemented. The original heading-variant work and the first resilience follow-up are
implemented. A second resilience follow-up is proposed to cover additional
observed NWS heading and key-message formats without changing the canonical
forecast-discussion contract.
## Implemented Behavior ## Completed Baseline
The NWS Area Forecast Discussion parser recognizes these heading families for The NWS Area Forecast Discussion parser currently separates:
all structurally valid section identities:
```text
.<SECTION>...<optional qualifier>
.<SECTION> /<qualifier>/...
```
Discovery, qualifier extraction, and recognized-section boundary detection use
one provider-specific parser. Slash delimiters are removed from qualifiers,
legacy qualifier text is preserved, and aviation remains a boundary rather than
a canonical field.
The parser separates:
1. generic structural heading recognition; 1. generic structural heading recognition;
2. section boundary scanning; 2. one-pass section-boundary scanning;
3. canonical section-role selection; and 3. canonical section-role selection; and
4. section-preamble and presentation cleanup. 4. section-preamble and presentation cleanup.
Structurally valid topic names act as safe boundaries without becoming new It recognizes generic uppercase identities in ellipsis-first and
canonical fields. slash-qualified headings, treats structurally recognized but unmapped sections
as boundaries, supports same-line and next-line qualifiers, removes exact NWS
change markers, and keeps the first occurrence of each mapped role. Provider and
normalizer tests cover two office-format families while preserving the existing
wire contract.
## Implemented Details ## Remaining Problem
### Generic heading recognition Observed AFDs still expose common forms outside the completed baseline:
One NWS-specific heading parser recognizes structurally valid uppercase AFD - a qualifier may appear between the identity and terminal ellipsis, as in
topic lines independently of the canonical model. It supports: `.DISCUSSION (Today through Thursday)...`;
- the same key-message concept may be headed `KEY POINTS` rather than `KEY
MESSAGES`;
- key-message items may use hyphens, asterisks, numeric markers such as `1)` or
`1.`, or unmarked paragraphs;
- a bare prefix check for `Issued at` or `Updated at` can misclassify ordinary
prose such as “Updated atmospheric conditions...” as metadata; and
- offices may publish semantically distinct `NEAR TERM`, `SHORT TERM`, and
`LONG TERM` sections, so wording flexibility cannot safely rely on treating
every similar identity as an alias.
These gaps do not undermine the provider-local architecture, but they mean the
parser is still brittle at its syntax, semantic-alias, and list-tokenization
extension points. The current compact cross-office fixture does not exercise
these format families.
## Objective
Make common, minor NWS presentation changes local and inexpensive to support
while preserving strict structural recognition and the existing canonical
schema. A new observed heading form should require an isolated grammar helper;
a true synonym should require one role-registry entry; and a new list marker
should require one marker-classifier case. None should require changes to the
scanner, canonical model, normalizer architecture, or downstream contracts.
## Target End State
### Heading syntax
The provider-local heading parser recognizes these three explicit families:
```text ```text
.<IDENTITY>...<optional qualifier> .<IDENTITY>...<optional qualifier>
.<IDENTITY> /<qualifier>/... .<IDENTITY> (<nonempty qualifier>)...
.<IDENTITY> /<nonempty qualifier>/...
``` ```
The identity may contain uppercase ASCII letters, digits, horizontal whitespace, The existing identity grammar remains unchanged: uppercase ASCII letters,
`/`, `&`, apostrophes, and hyphens. Outer whitespace and whitespace immediately digits, horizontal whitespace, `/`, `&`, apostrophes, and hyphens, with at least
before the ellipsis are ignored, and internal identity whitespace is collapsed one letter or digit. Identity whitespace is normalized to one ASCII space.
to one space for role lookup. Slash-qualified headings require a nonempty
qualifier but may contain slash characters inside that qualifier. Lowercase
prose, ordinary dotted lines, and malformed delimiters remain non-headings.
Every structurally valid heading terminates the preceding section. This applies The parenthesized-terminal form:
to known boundary-only sections such as aviation and to future or office-specific
topic identities that weatherfeeder does not map.
### Canonical role selection - requires horizontal whitespace between the identity and opening `(`;
- requires a nonempty standalone parenthetical qualifier immediately before
the terminal ASCII `...`;
- retains the outer parentheses in the parsed qualifier;
- permits ordinary qualifier punctuation inside the parentheses; and
- rejects trailing text after the terminal ellipsis.
A single provider-local role registry maps only these identities into existing Each family is parsed by a small, ordered helper. Slash-qualified parsing remains
parsed fields: first because its terminal is otherwise ambiguous with the ellipsis-first form;
parenthesized-terminal parsing runs second, followed by the existing
ellipsis-first form. Malformed heading-like lines remain body content.
- `KEY MESSAGES` to key messages; Every successfully parsed heading is a section boundary regardless of whether
its identity has a canonical role. This includes `DISCUSSION`, `NEAR TERM`,
aviation, marine, hydrology, office-specific sections, and future structurally
compatible identities.
### Canonical role aliases
One provider-local identity-to-role registry remains the sole semantic mapping
source. It maps:
- `KEY MESSAGES` and `KEY POINTS` to key messages;
- `SHORT TERM` to the short-term section; and - `SHORT TERM` to the short-term section; and
- `LONG TERM` to the long-term section. - `LONG TERM` to the long-term section.
All other identities are boundary-only. The first occurrence of each mapped Aliases are explicit and evidence-based; identity similarity is never inferred.
identity wins, preserving current behavior if an unusual bulletin repeats a The first encountered identity for a role wins, so `KEY MESSAGES` and `KEY
section. Heading syntax and canonical roles do not duplicate identity lists. POINTS` participate in the same first-occurrence policy.
### Section scanning `NEAR TERM` remains boundary-only. It is a distinct section that may coexist
with `SHORT TERM`, so silently relabeling it would lose meaning. Exposing it
requires a separate canonical schema roadmap. `DISCUSSION` and other unmapped
sections likewise remain boundary-only.
The discussion text is scanned once into ordered section blocks. A new heading, ### Key-message normalization
`&&`, `$$`, or the existing watch/advisory terminator ends the active block.
Original body lines are preserved until provider presentation cleanup is
applied. Preamble text before the first heading and post-`$$` signatures are not
treated as section content.
### Preamble and presentation handling After exact presentation-marker removal and blank-line trimming, the parser
removes at most one leading metadata line only when all of these conditions are
true:
For short- and long-term sections: - the label is exactly `Issued at` or `Updated at`, compared ASCII
case-insensitively;
- the label is followed by horizontal whitespace; and
- the remainder is a valid timestamp under the existing NWS issue-time grammar.
- a qualifier on the heading line remains authoritative; An invalid or merely prefix-matching line remains content. Metadata recognition
- when the heading has no qualifier, a standalone parenthesized first content does not consume later message prose.
line becomes the qualifier and retains its parentheses;
- an optional case-insensitive `Issued at` line after the qualifier is parsed
into the existing section issue time; and
- exact NWS change-presentation marker lines are removed without removing
arbitrary dashed prose.
For key messages, exact change-presentation markers and one leading `Issued at` Key-message item recognition supports:
or `Updated at` metadata line are removed before bullet parsing. Those metadata
lines never become key messages. - hyphen markers;
- asterisk markers;
- positive ASCII numeric markers followed by `)` or `.`, then either end of line
or a horizontal-space boundary;
- composite hyphen-or-asterisk plus numeric markers such as `- 1.`; and
- unmarked, blank-line-separated paragraphs when the block contains no
recognized list marker.
Marker text is removed from canonical messages. Wrapped nonempty lines remain
continuations of the current item, and blank lines between marked items do not
create empty messages. If marked and unmarked content are mixed, nonempty prose
before the first marker is preserved as its own message and later unmarked lines
continue the active marked item. The parser never interprets numeric or
asterisk markers outside a mapped key-message block.
### Representative coverage ### Representative coverage
Tests include small structural tables and maintained local HTML fixtures for at Tests combine small table-driven grammar and tokenizer cases with maintained
least two real NWS formatting families: slash-qualified same-line headings and local HTML fixtures for distinct real NWS format families. In addition to the
ellipsis-first headings with multiline qualifiers or change-presentation existing LSX and BOU coverage, fixtures cover:
markers. Tests remain deterministic and never contact live services.
- numbered `KEY MESSAGES` followed by a boundary-only `DISCUSSION` section; and
- asterisk `KEY POINTS` plus a parenthesized-terminal `DISCUSSION` heading.
Fixtures remain compact, deterministic, attributable in test comments to the
format family they represent, and free of live network dependencies. Provider
and normalizer tests prove canonical values, boundary isolation, envelope
behavior, and unchanged JSON wire shape.
## Compatibility and Contracts ## Compatibility and Contracts
@@ -107,41 +162,46 @@ This remains a provider-parsing compatibility improvement. It does not change:
- Postgres tables or event-to-row mapping; or - Postgres tables or event-to-row mapping; or
- downstream sink and consumer responsibilities. - downstream sink and consumer responsibilities.
A single `DISCUSSION` section and other currently unmapped identities are The implementation remains confined to the NWS provider parser plus owning
recognized as boundaries but are not forced into short- or long-term fields. provider and normalizer tests. It uses the Go standard library and introduces no
Exposing such content would require a separate canonical schema decision. runtime configuration or general parser framework.
## Verified Coverage ## Acceptance Criteria
Automated tests demonstrate that: The follow-up is complete when automated tests demonstrate that:
- every previously accepted heading and canonical result remains compatible; - all previously supported heading, scanning, preamble, marker, and canonical
- generic structurally valid identities terminate preceding content without results remain compatible;
becoming canonical fields; - parenthesized-terminal headings are parsed with their qualifier and terminate
- malformed heading-like lines and lowercase prose remain body content; preceding content;
- identities containing `/` and slash qualifiers containing `/` are parsed - malformed variants of that heading family remain body content;
without ambiguity; - `KEY POINTS` populates the existing key-message field and shares
- next-line parenthesized qualifiers and following `Issued at` lines populate first-occurrence behavior with `KEY MESSAGES`;
the existing short- and long-term metadata correctly; - `NEAR TERM` and `DISCUSSION` remain boundary-only and never populate short- or
- exact NWS change markers and leading key-message timestamps do not leak into long-term fields;
canonical prose or messages; - hyphen, asterisk, `N)`, `N.`, and composite key-message lists produce distinct
- repeated mapped sections preserve first-occurrence behavior; ordered messages with wrapped continuations;
- genuine cross-office fixture styles propagate correctly through the NWS - unmarked key-message paragraphs produce distinct ordered messages;
normalizer with unchanged wire shape; and - valid leading issue/update metadata is removed, while prefix collisions,
- focused tests, the full repository suite, and static analysis pass. malformed timestamps, and later timestamp-like prose are retained;
- representative additional office-format fixtures parse end to end through the
provider and normalizer without adjacent-section leakage; and
- focused tests, the full repository suite, static analysis, and diff checks
pass with no public contract changes.
## Non-Goals ## Non-Goals
This follow-up does not: This follow-up does not:
- add canonical `discussion`, aviation, marine, hydrology, climate, fire - add canonical `nearTerm`, `discussion`, aviation, marine, hydrology, climate,
weather, update, or arbitrary-section fields; fire-weather, update, or arbitrary-section fields;
- infer short- or long-term semantics from an unknown heading; - map `NEAR TERM` to `SHORT TERM` or infer roles from similar words;
- accept free-form or lowercase prose as a heading; - parse arbitrary lowercase or free-form prose as headings;
- introduce heuristic summarization; - accept arbitrary punctuation as list markers;
- preserve complete raw AFD documents in canonical payloads; - remove malformed metadata-like prose;
- introduce heuristic summarization or preserve complete raw AFD documents;
- change schemas, persistence contracts, configuration, or downstream APIs; or - change schemas, persistence contracts, configuration, or downstream APIs; or
- fetch live NWS data during tests. - fetch live NWS data during tests.
Additional canonical support for AFD section identities requires a separate Additional canonical support for AFD section identities remains a separate
consumer requirement and schema roadmap. consumer and schema decision.

View File

@@ -2,108 +2,135 @@
## Purpose ## Purpose
Complete the resilience end state in Complete the remaining resilience work defined in
[`afd-section-heading-variants.md`](afd-section-heading-variants.md) while [`afd-section-heading-variants.md`](afd-section-heading-variants.md) without
preserving the current canonical forecast-discussion contract. Stages 1-3 below changing the canonical forecast-discussion contract. Stages 1-8 summarize work
summarize completed work. Implement Stages 4-8 in order. already completed. Implement Stages 9-13 in order.
## Cross-Stage Constraints ## Cross-Stage Constraints
- Keep all production parsing changes in - Keep production parsing changes in
`internal/providers/nws/forecast_discussion.go`. `internal/providers/nws/forecast_discussion.go`.
- Use only the Go standard library and keep helpers unexported. - Use only the Go standard library and keep new helpers unexported.
- Do not change `model`, `standards`, source configuration or polling, schemas, - Do not change `model`, `standards`, source configuration or polling, schema
event-envelope behavior, Postgres mapping, sinks, consumer docs, or current identifiers, event-envelope behavior, Postgres mapping, sinks, consumer docs,
integration contracts. or integration contracts.
- Continue exposing only key messages, short term, and long term. All other - Continue exposing only key messages, short term, and long term.
structurally valid identities are boundary-only. - Map `KEY POINTS` to the existing key-message role. Keep `NEAR TERM`,
- Preserve first-occurrence behavior for mapped sections. `DISCUSSION`, and every other unmapped identity boundary-only.
- Preserve raw body lines during structural scanning; apply provider-specific - Preserve first-occurrence behavior by canonical role, including across `KEY
cleanup only when parsing a block's content. MESSAGES` and `KEY POINTS` aliases.
- Preserve raw body lines during structural scanning. Apply provider-specific
cleanup only while parsing a mapped block.
- Prefer small grammar, label, and list-marker helpers over a broad regular
expression or general parsing framework.
- Keep parsing deterministic. Tests must use local strings and fixtures, never - Keep parsing deterministic. Tests must use local strings and fixtures, never
live NWS requests. live NWS requests.
- Preserve all pre-existing user work and avoid unrelated refactors. - Preserve all pre-existing user work and avoid unrelated refactors.
## Stage 1: Centralize Known Heading Parsing — Completed ## Stage 1: Centralize Known Heading Parsing — Completed
The provider parser gained one heading classifier, explicit section and block The provider parser gained a shared heading classifier and section-block model
types, and shared discovery, boundary, and qualifier handling for the original for the original key-message, short-term, long-term, and aviation identities.
four identities. Key-message and text-section consumers were moved to the block Discovery, boundary handling, and qualifier extraction stopped using separate
representation, and immediately adjacent recognized headings became valid per-section patterns.
boundaries.
## Stage 2: Add Heading-Variant Regressions — Completed ## Stage 2: Add Original Heading-Variant Regressions — Completed
Provider tests now cover ellipsis-first and slash-qualified headings, malformed Provider and normalizer tests covered ellipsis-first and slash-qualified
forms, empty bodies, mixed heading families, and slash-qualified aviation headings, malformed forms, empty bodies, mixed heading families, and unchanged
termination. Normalizer coverage verifies canonical propagation and unchanged canonical wire shape.
wire shape.
## Stage 3: Validate and Record the Baseline — Completed ## Stage 3: Validate the Original Feature — Completed
Focused and full tests passed, production changes remained inside the NWS Focused and repository-wide checks passed, production changes remained inside
provider parser, and the original feature roadmap was marked implemented. the NWS provider parser, and the original roadmap was reconciled with the
implemented baseline.
## Stage 4: Generalize Heading Syntax and Centralize Canonical Roles ## Stage 4: Generalize Structural Heading Recognition — Completed
Decouple structural heading recognition from canonical section selection. Heading syntax was decoupled from canonical roles. The parser now accepts a
constrained generic uppercase identity grammar, normalizes identity whitespace,
supports embedded slashes without confusing slash qualifiers, and recognizes
unknown valid identities as structural headings.
1. In `internal/providers/nws/forecast_discussion.go`, replace the four-name ## Stage 5: Scan Ordered Section Blocks Once — Completed
regular-expression alternation with a generic heading parser. Keep the
`forecastDiscussionSectionHeading` result, with normalized `section` and
`qualifier` fields.
2. Implement the generic parser with these decisions:
- trim outer whitespace and require a leading `.`; Repeated extraction was replaced by a one-pass ordered block scanner. Generic
- try the slash-qualified form before the legacy ellipsis-first form so the headings, `&&`, `$$`, and the watch/advisory safeguard terminate active blocks;
terminal ellipsis of an all-uppercase slash heading cannot be mistaken for role lookup happens after scanning and retains first-occurrence behavior.
a legacy identity;
- accept identities made only from uppercase ASCII letters, digits,
horizontal whitespace, `/`, `&`, apostrophes, and hyphens, with at least
one letter or digit;
- trim the identity and collapse every internal run of spaces or tabs to one
ASCII space before returning it or performing role lookup;
- accept optional horizontal whitespace between the identity and the legacy
`...` delimiter;
- for legacy headings, store trimmed text after the first `...` as the
qualifier, preserving parentheses and permitting an empty value;
- for slash-qualified headings, remove the terminal `/...`, then use the
first slash preceded by horizontal whitespace as the opening separator;
trim the identity before that separator and the qualifier after it, reject
an empty qualifier, and allow additional `/` characters inside the
qualifier; slashes inside an identity must therefore be adjacent to its
other identity characters rather than preceded by whitespace;
- reject trailing text after a slash-qualified terminal, lowercase or mixed
case identities, missing delimiters, and lines whose identity contains
other punctuation; and
- keep ASCII `...` as the only delimiter because that is the raw product
convention; do not interpret a Unicode ellipsis.
3. Replace the duplicated identity constants and identity-pattern string with:
- an unexported `forecastDiscussionSectionRole` enum for key messages, short ## Stage 6: Normalize Multiline Preambles and Markers — Completed
term, and long term; and
- one `map[string]forecastDiscussionSectionRole` containing exactly `KEY
MESSAGES`, `SHORT TERM`, and `LONG TERM`.
Unknown identities and `AVIATION` intentionally have no role entry; successful Short- and long-term parsing gained next-line parenthesized qualifiers,
structural parsing is sufficient for boundary behavior. case-insensitive `Issued at` handling, and exact change-marker removal. Key
4. Update the heading table test in messages gained exact marker cleanup and removal of one leading issue/update
`internal/providers/nws/forecast_discussion_test.go`. Preserve every legacy metadata line.
positive case and revise the old policy-specific negatives:
- `.SYNOPSIS...`, `.UPDATE...`, `.MARINE...`, `.HYDROLOGY...`, an office ## Stage 7: Add Initial Cross-Office Coverage — Completed
watch/advisory heading, and `.PRELIMINARY POINT TEMPS/POPS ...` are valid
generic headings; The existing LSX fixture was retained and a compact BOU-style fixture added.
- a slash-qualified heading whose qualifier contains `/` is valid; Provider and normalizer regressions cover multiline qualifiers, change markers,
- leading/trailing whitespace, multiple separator spaces, repeated internal leading metadata, generic boundary isolation, envelope behavior, and unchanged
identity whitespace, and whitespace before a legacy ellipsis are valid and JSON shape.
produce the normalized identity; and
- lowercase prose, an empty or punctuation-only identity, malformed slash ## Stage 8: Validate the First Resilience Follow-up — Completed
terminals, missing ellipses, and unsupported identity punctuation remain
invalid. Focused tests, the full repository suite, static analysis, and diff checks
5. Add assertions that every role-registry key parses successfully, and confirm passed. The public model and schemas remained unchanged, and the roadmap was
that no production switch or second collection repeats the mapped identity updated to describe the completed behavior.
list.
## Stage 9: Add Parenthesized-Terminal Heading Syntax
Recognize the observed `.<IDENTITY> (<qualifier>)...` family without weakening
generic identity validation.
1. In `internal/providers/nws/forecast_discussion.go`, add a dedicated
`parseForecastDiscussionParenthesizedTerminalHeading` helper returning the
existing `forecastDiscussionSectionHeading` type.
2. Update `parseForecastDiscussionSectionHeading` to try forms in this order:
1. if the line ends `/...`, try the slash-qualified helper and return its
result;
2. otherwise, try the parenthesized-terminal helper for any line ending
`...` and return it when it succeeds; and
3. fall back to the existing ellipsis-first helper.
Do not merge the forms into a single regular expression.
3. Implement the parenthesized-terminal helper with these exact rules:
- operate on the already outer-trimmed line and require leading `.` plus a
terminal ASCII `...`;
- remove the leading dot and terminal ellipsis, then remove horizontal
whitespace immediately before the ellipsis;
- require the remaining content to end with `)`;
- locate the first `(` that is preceded by horizontal whitespace; everything
before that separator is the raw identity and everything from `(` through
the final `)` is the qualifier;
- normalize and validate the identity through the existing identity helper;
- require nonempty text after trimming inside the outer parentheses;
- return the qualifier with its outer parentheses and internal punctuation
intact; and
- reject missing separator whitespace, empty qualifiers, missing or misplaced
parentheses, trailing text after the ellipsis, unsupported identity
punctuation, and lowercase or mixed-case identities.
Parentheses inside the qualifier are content; only the first separator `(`
and final `)` delimit the outer qualifier.
4. Extend the heading table in
`internal/providers/nws/forecast_discussion_test.go` with:
- `.DISCUSSION (Today through Thursday)...`;
- a mapped identity such as `.SHORT TERM (Tonight)...`;
- identity whitespace normalization;
- qualifier punctuation and nested parentheses;
- each rejected malformed form listed above; and
- compatibility assertions for every existing ellipsis-first and
slash-qualified case.
5. Add a scanner regression in which a mapped section is followed immediately,
without `&&`, by `.DISCUSSION (Today through Thursday)...`. Assert that the
discussion heading starts a new boundary-only block and its prose cannot leak
into the mapped section.
Run and pass: Run and pass:
@@ -112,46 +139,37 @@ gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/fo
go test -count=1 ./internal/providers/nws go test -count=1 ./internal/providers/nws
``` ```
Do not proceed until generic syntax tests pass without changing canonical Do not proceed until all previous heading and scanner tests still pass.
models or schemas.
## Stage 5: Scan Ordered Blocks Once and Use Generic Boundaries ## Stage 10: Add Explicit Canonical Role Aliases
Replace repeated per-identity extraction with one structural pass. Make true wording synonyms cheap to support while preserving semantic
distinctions.
1. Replace `extractForecastDiscussionSection` with 1. Add `KEY POINTS` to the existing identity-to-role registry with the key-message
`parseForecastDiscussionSectionBlocks(lines []string) []forecastDiscussionSectionBlock`. role. Keep the registry as the only production source of identity-to-role
The scanner must: mappings; do not add a parallel alias collection or role-selection switch.
2. Leave `KEY MESSAGES`, `SHORT TERM`, and `LONG TERM` mappings unchanged.
Explicitly do not add role mappings for `NEAR TERM`, `DISCUSSION`,
`AVIATION`, or other generic identities.
3. Update registry tests so they no longer assume one identity per role. Assert
instead that:
- traverse lines once in source order; - every registry identity parses as a heading;
- start a block on every structurally valid heading; - `KEY MESSAGES` and `KEY POINTS` both map to the key-message role;
- finish the active block before a new heading; - `SHORT TERM` and `LONG TERM` retain their roles; and
- finish and clear the active block on `&&`; - `NEAR TERM`, `DISCUSSION`, and `AVIATION` have no role entry.
- finish the active block and stop scanning on `$$`; 4. Add `ParseForecastDiscussionText` tests for role-level first occurrence:
- retain the existing watch/advisory termination safeguard even though a
normal dotted watch/advisory heading is structurally recognized;
- ignore preamble lines before the first heading and signature lines after
`$$`; and
- append original, untrimmed body lines to the active block.
2. Refactor `ParseForecastDiscussionText` to iterate over the ordered blocks
once and look up each heading identity in the role registry. Ignore blocks
with no role. Populate each mapped output only if it has not already been
populated, so the first occurrence wins. Track seen roles explicitly rather
than inferring them from output values, because an empty first key-message
block still counts as the first occurrence.
3. Preserve contextual errors when a mapped text block fails to parse; include
the parsed heading identity in the wrapped error.
4. Replace extraction tests with scanner tests covering:
- consecutive headings and empty bodies; - `KEY POINTS` alone populates key messages;
- `&&`, `$$`, and watch/advisory termination; - `KEY POINTS` followed by `KEY MESSAGES` keeps the first block; and
- unknown valid headings terminating short- or long-term content even when - `KEY MESSAGES` followed by `KEY POINTS` keeps the first block.
no `&&` is present;
- malformed heading-like lines remaining inside the active body; Use hyphen messages in this stage so list-tokenization changes remain scoped
- preamble and post-signature exclusion; to Stage 11.
- ordered block retention; and 5. Add a regression containing `NEAR TERM`, `SHORT TERM`, and `LONG TERM` in one
- duplicate mapped sections where the parser keeps the first canonical bulletin. Assert that near-term prose terminates adjacent blocks but does not
occurrence. populate or override either canonical text section.
Run and pass: Run and pass:
@@ -160,85 +178,110 @@ gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/fo
go test -count=1 ./internal/providers/nws go test -count=1 ./internal/providers/nws
``` ```
## Stage 6: Normalize Multiline Preambles and NWS Presentation Markers Do not change the canonical model to expose near-term or discussion content.
Handle real section-layout variation without weakening heading syntax. ## Stage 11: Harden Key-Message Metadata and Item Parsing
1. Add an exact provider-local marker predicate for lines whose trimmed value is Replace prefix-sensitive, hyphen-only parsing with conservative metadata and
`-- Changed Discussion --` or `-- End Changed Discussion --`, compared list classifiers.
case-insensitively. Add a helper that removes only those complete marker
lines from a block body. Do not remove arbitrary dashed lines or bullet text.
2. Apply marker removal before parsing both key-message and text-section bodies.
3. Refactor text-section preamble parsing in this exact order:
- trim blank lines after marker removal; 1. Add a helper that recognizes a leading labeled metadata line only when:
- start with the qualifier parsed from the heading;
- only when that qualifier is empty, consume the first content line as a
qualifier if its trimmed value is a nonempty standalone parenthetical
string beginning with `(` and ending with `)`; preserve the parentheses;
- after the optional qualifier, consume an optional `Issued at` line and
parse it into `ForecastDiscussionSection.IssuedAt`; and
- pass only the remaining lines to existing signature trimming and paragraph
joining.
4. Make recognition of the `Issued at` label ASCII case-insensitive while
retaining the existing timestamp grammar and error behavior. The top-level
header path, which passes an unlabeled timestamp, must remain compatible.
5. Before key-message bullet parsing, remove markers, trim blank lines, and
discard at most one leading metadata line beginning case-insensitively with
`Issued at` or `Updated at`. Do not discard timestamp-like lines after the
first message begins.
6. Add focused tests for:
- same-line legacy and slash qualifiers remaining unchanged; - its trimmed text begins with exactly `Issued at` or `Updated at`, compared
- next-line parenthesized qualifiers followed by `Issued at`; ASCII case-insensitively;
- uppercase `ISSUED AT`; - at least one horizontal-whitespace byte follows the label; and
- a heading qualifier taking precedence over a following parenthetical prose - parsing the remaining text through the existing unlabeled NWS issue-time
line; grammar succeeds.
- empty sections;
- invalid issue timestamps retaining contextual errors; Reuse `parseForecastDiscussionIssueTime` for the timestamp grammar after
- marker removal at the beginning and end of text sections; removing the label. Do not add a second timestamp parser. A label prefix with
- key-message markers and a leading `Updated at` line not becoming messages; no boundary or with an invalid timestamp returns false and remains content.
2. Update key-message cleanup to remove at most one such valid leading metadata
line after presentation-marker removal and blank-line trimming. Later valid
timestamp lines remain message content.
3. Replace the hyphen-only check with a helper that classifies and strips one of
these markers from a trimmed line:
- one leading `-`, preserving current compatibility whether or not whitespace
follows it;
- one leading `*`, whether or not whitespace follows it; or
- one or more ASCII digits forming an integer greater than zero, followed by
`)` or `.`, and then either end-of-line or horizontal whitespace.
After recognizing and stripping a leading hyphen or asterisk plus horizontal
whitespace, also strip one immediately following valid numeric marker. This
supports composite forms such as `- 1.` without retaining either decorator.
Otherwise strip only the recognized marker and following horizontal
whitespace. Reject zero, overflow, alphanumeric prefixes, and numeric
punctuation without the required boundary. Do not interpret these markers
outside key-message parsing.
4. Parse a cleaned key-message body deterministically:
- first determine whether any nonempty line has a recognized marker;
- when markers exist, each marker starts a new message; nonempty unmarked
lines after a marker continue that message; blank lines after the first
marker are ignored; and blank-line-separated prose before the first marker
is flushed as preserved message content;
- when no marker exists, each nonempty paragraph separated by one or more
blank lines becomes one message; and
- in both modes, join wrapped lines with one ASCII space, discard empty
messages, and preserve source order.
5. Add table-driven marker tests for `-`, `*`, `1)`, `2.`, composite `- 1.` and
`* 1)`, multi-digit values, wrapped continuations, marker-only lines, and
rejected zero/malformed numeric prefixes.
6. Add block-level tests covering:
- numbered and asterisk lists producing distinct messages;
- unmarked paragraphs producing distinct messages;
- mixed introductory prose and marked items without data loss;
- valid leading `Issued at` and `Updated at` metadata removal;
- `Updated atmospheric conditions...` remaining content;
- `Updated at not a timestamp` remaining content;
- a later valid timestamp-like line remaining content; and
- all existing hyphen, change-marker, and continuation behavior.
Run and pass:
```sh
gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/forecast_discussion_test.go
go test -count=1 ./internal/providers/nws
```
## Stage 12: Add Representative Numbered and Key-Points Fixtures
Prove the new extension points through provider and normalizer boundaries.
1. Retain the existing LSX and BOU fixtures and their regressions.
2. Add two compact HTML fixtures under `internal/providers/nws/testdata/`:
- a BGM/CTP-style fixture with a valid header and issue time, numbered `KEY
MESSAGES`, wrapped item lines, and a boundary-only `DISCUSSION` section; and
- an MFR-style fixture with a valid header and issue time, asterisk `KEY
POINTS`, and an immediately following
`.DISCUSSION (Today through Thursday)...` section without an intervening
`&&`.
Use concise representative text rather than full web pages. Add a short HTML
comment to each fixture naming the real office-format family it represents;
do not claim that edited fixture prose is a verbatim archived product.
3. Add provider end-to-end tests that assert:
- exact ordered key messages and joined continuations;
- `KEY POINTS` alias mapping;
- correct office and top-level issue metadata;
- absence of numeric/asterisk markers from canonical values;
- absence of `DISCUSSION` headings, qualifiers, and prose from key messages;
and and
- arbitrary dashed prose remaining content. - nil short- and long-term fields when the fixture contains neither mapped
role.
Run and pass: 4. Add table-driven normalizer regressions using both fixtures. For each, verify
input envelope preservation, event kind, canonical schema, effective time,
```sh exact key-message mapping, and successful JSON marshaling.
gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/forecast_discussion_test.go 5. Assert the JSON payload still has no `nearTerm`, `discussion`, `aviation`, or
go test -count=1 ./internal/providers/nws generic `sections` field. Do not add those fields to provider or canonical
``` structs.
6. Keep syntax and tokenizer edge cases in focused unit tables; do not add more
## Stage 7: Add Cross-Office Fixtures and Normalizer Regressions full fixtures for cases already proven locally.
Prove the resilient parser against representative source shapes rather than only
synthetic heading replacement.
1. Retain the existing LSX fixture and mixed-format test for regression
compatibility.
2. Add one compact, maintained HTML fixture under
`internal/providers/nws/testdata/` representing a second real NWS formatting
family. It must contain:
- a valid AFD header and issue time;
- key-message change markers plus a leading `Updated at` line;
- ellipsis-first short- and long-term headings with qualifiers on the next
line;
- `Issued at` lines following those qualifiers;
- at least one structurally valid boundary-only section; and
- representative prose sufficient to detect metadata or adjacent-section
leakage.
Keep the fixture concise; include no unrelated webpage content, secrets, or
private data.
3. Add a provider end-to-end test that parses the new fixture and asserts exact
key messages, qualifiers, section issue times, representative prose, and
absence of change markers, timestamp metadata, and boundary-only content.
4. Add a normalizer regression using the same fixture. Verify kind, canonical
schema, envelope/effective-time behavior, short- and long-term mapping, and
JSON wire shape without `aviation`, `discussion`, or generic `sections`
fields.
5. Keep small parser and scanner edge cases table-driven. Do not multiply full
fixtures for cases that a short line slice proves more clearly.
Run and pass: Run and pass:
@@ -247,17 +290,18 @@ gofmt -w internal/providers/nws/forecast_discussion_test.go internal/normalizers
go test -count=1 ./internal/providers/nws ./internal/normalizers/nws go test -count=1 ./internal/providers/nws ./internal/normalizers/nws
``` ```
## Stage 8: Reconcile Documentation and Perform Final Validation ## Stage 13: Reconcile Documentation and Perform Final Validation
Close the resilience follow-up only after all behavior is proven. Close the second resilience follow-up only after all behavior is implemented and
verified.
1. Review the final diff. Production changes must remain confined to the NWS 1. Review the final diff. Production changes must remain confined to
provider parser; other Go changes should be tests in the owning provider and `internal/providers/nws/forecast_discussion.go`; other Go changes must be
normalizer packages. Preserve all unrelated user work. owning provider and normalizer tests. Preserve unrelated user work.
2. Confirm the public contract is unchanged. Do not edit canonical model, 2. Confirm the public contract is unchanged. If implementation appears to
schema, Postgres, config, consumer, or integration documentation unless an require a canonical near-term/discussion field, schema change, configuration,
actual contract change is discovered. If one appears necessary, stop rather or downstream migration, stop and report the conflict rather than expanding
than expanding this plan. scope.
3. Run: 3. Run:
```sh ```sh
@@ -271,14 +315,18 @@ Close the resilience follow-up only after all behavior is proven.
git diff --check git diff --check
``` ```
4. Verify every acceptance criterion in the feature roadmap has a corresponding 4. Verify every acceptance criterion in the feature roadmap has a corresponding
passing automated test, including compatibility with all Stage 1-2 cases. passing test, including compatibility with Stages 1-8.
5. After all checks pass, change the feature roadmap status to `Implemented.` 5. Update the feature roadmap status to `Implemented.` and rewrite its remaining
and rewrite any remaining future-tense statements that would misdescribe the proposed or future-tense language as implemented behavior only after all
completed parser. Do not mark the follow-up implemented earlier. checks pass.
6. Update this implementation plan so Stages 9-13 are marked `— Completed` and
the Purpose states that all stages are complete. Preserve the stage details
as a historical implementation record.
## Open Questions ## Open Questions
None. This plan fixes the parser grammar, boundary policy, canonical role None. The plan makes the required policy choices explicitly: `KEY POINTS` is a
selection, multiline preamble rules, presentation cleanup, fixture strategy, true key-message alias; `NEAR TERM` remains semantically distinct and
and compatibility boundary. Generic or single-section discussion content remains boundary-only; parenthesized-terminal syntax is a third constrained grammar;
outside the canonical model by explicit policy. metadata is removed only after timestamp validation; and key-message variants
are handled by an isolated marker classifier plus paragraph fallback.

View File

@@ -426,6 +426,11 @@ func parseForecastDiscussionSectionHeading(line string) (forecastDiscussionSecti
if strings.HasSuffix(line, "/...") { if strings.HasSuffix(line, "/...") {
return parseForecastDiscussionSlashQualifiedHeading(line) return parseForecastDiscussionSlashQualifiedHeading(line)
} }
if strings.HasSuffix(line, "...") {
if heading, ok := parseForecastDiscussionParenthesizedTerminalHeading(line); ok {
return heading, true
}
}
return parseForecastDiscussionEllipsisHeading(line) return parseForecastDiscussionEllipsisHeading(line)
} }
@@ -453,6 +458,38 @@ func parseForecastDiscussionSlashQualifiedHeading(line string) (forecastDiscussi
return forecastDiscussionSectionHeading{section: section, qualifier: qualifier}, true return forecastDiscussionSectionHeading{section: section, qualifier: qualifier}, true
} }
func parseForecastDiscussionParenthesizedTerminalHeading(line string) (forecastDiscussionSectionHeading, bool) {
if len(line) < 4 || line[0] != '.' || !strings.HasSuffix(line, "...") {
return forecastDiscussionSectionHeading{}, false
}
content := strings.TrimRight(line[1:len(line)-3], " \t")
if !strings.HasSuffix(content, ")") {
return forecastDiscussionSectionHeading{}, false
}
separator := -1
for i := 1; i < len(content); i++ {
if content[i] == '(' && isForecastDiscussionHorizontalWhitespace(content[i-1]) {
separator = i
break
}
}
if separator < 0 {
return forecastDiscussionSectionHeading{}, false
}
section, ok := normalizeForecastDiscussionSectionIdentity(content[:separator])
if !ok {
return forecastDiscussionSectionHeading{}, false
}
qualifier := content[separator:]
if len(qualifier) <= 2 || strings.TrimSpace(qualifier[1:len(qualifier)-1]) == "" {
return forecastDiscussionSectionHeading{}, false
}
return forecastDiscussionSectionHeading{section: section, qualifier: qualifier}, true
}
func parseForecastDiscussionEllipsisHeading(line string) (forecastDiscussionSectionHeading, bool) { func parseForecastDiscussionEllipsisHeading(line string) (forecastDiscussionSectionHeading, bool) {
content := line[1:] content := line[1:]
delimiter := strings.Index(content, "...") delimiter := strings.Index(content, "...")

View File

@@ -129,6 +129,34 @@ func TestParseForecastDiscussionSectionHeading(t *testing.T) {
wantQualifier: "(Tonight)", wantQualifier: "(Tonight)",
wantOK: true, wantOK: true,
}, },
{
name: "parenthesized terminal discussion",
line: ".DISCUSSION (Today through Thursday)...",
wantSection: "DISCUSSION",
wantQualifier: "(Today through Thursday)",
wantOK: true,
},
{
name: "parenthesized terminal mapped section",
line: ".SHORT TERM (Tonight)...",
wantSection: "SHORT TERM",
wantQualifier: "(Tonight)",
wantOK: true,
},
{
name: "parenthesized terminal normalized identity whitespace",
line: " .SHORT\tTERM (Tonight)... ",
wantSection: "SHORT TERM",
wantQualifier: "(Tonight)",
wantOK: true,
},
{
name: "parenthesized terminal qualifier punctuation and nesting",
line: ".DISCUSSION (Today (and Thursday), 70% chance!)...",
wantSection: "DISCUSSION",
wantQualifier: "(Today (and Thursday), 70% chance!)",
wantOK: true,
},
{ {
name: "lowercase prose", name: "lowercase prose",
line: ".This is ordinary prose...", line: ".This is ordinary prose...",
@@ -194,6 +222,41 @@ func TestParseForecastDiscussionSectionHeading(t *testing.T) {
line: ".Short Term... (Tonight)", line: ".Short Term... (Tonight)",
wantOK: false, wantOK: false,
}, },
{
name: "parenthesized terminal missing separator whitespace",
line: ".DISCUSSION(Today)...",
wantOK: false,
},
{
name: "parenthesized terminal empty qualifier",
line: ".DISCUSSION ( )...",
wantOK: false,
},
{
name: "parenthesized terminal missing final parenthesis",
line: ".DISCUSSION (Today...",
wantOK: false,
},
{
name: "parenthesized terminal misplaced parentheses",
line: ".DISCUSSION Today)...",
wantOK: false,
},
{
name: "parenthesized terminal trailing text",
line: ".DISCUSSION (Today)... extra",
wantOK: false,
},
{
name: "parenthesized terminal unsupported identity punctuation",
line: ".DISCUSSION: (Today)...",
wantOK: false,
},
{
name: "parenthesized terminal lowercase identity",
line: ".discussion (Today)...",
wantOK: false,
},
} }
for _, tt := range tests { for _, tt := range tests {
@@ -310,6 +373,28 @@ func TestParseForecastDiscussionSectionBlocksUsesGenericHeadingsAsBoundaries(t *
} }
} }
func TestParseForecastDiscussionSectionBlocksUsesParenthesizedTerminalHeadingAsBoundary(t *testing.T) {
got := parseForecastDiscussionSectionBlocks([]string{
".SHORT TERM... (Tonight)",
"Short-term prose.",
".DISCUSSION (Today through Thursday)...",
"Discussion prose.",
})
want := []forecastDiscussionSectionBlock{
{
heading: forecastDiscussionSectionHeading{section: "SHORT TERM", qualifier: "(Tonight)"},
body: []string{"Short-term prose."},
},
{
heading: forecastDiscussionSectionHeading{section: "DISCUSSION", qualifier: "(Today through Thursday)"},
body: []string{"Discussion prose."},
},
}
if !reflect.DeepEqual(got, want) {
t.Fatalf("blocks = %#v, want %#v", got, want)
}
}
func TestParseForecastDiscussionSectionBlocksKeepsMalformedHeadingLikeLines(t *testing.T) { func TestParseForecastDiscussionSectionBlocksKeepsMalformedHeadingLikeLines(t *testing.T) {
got := parseForecastDiscussionSectionBlocks([]string{ got := parseForecastDiscussionSectionBlocks([]string{
".SHORT TERM... (Tonight)", ".SHORT TERM... (Tonight)",