From f720b6cdc076a06d4b5c026494db2b48e0a9be6a Mon Sep 17 00:00:00 2001 From: Eric Rakestraw Date: Sun, 2 Aug 2026 20:47:24 +0000 Subject: [PATCH] Plan support for NWS AFD section heading variants --- docs/roadmap/afd-section-heading-variants.md | 141 +++++++++++++ docs/roadmap/implementation.md | 207 +++++++++++++++++++ 2 files changed, 348 insertions(+) create mode 100644 docs/roadmap/afd-section-heading-variants.md create mode 100644 docs/roadmap/implementation.md diff --git a/docs/roadmap/afd-section-heading-variants.md b/docs/roadmap/afd-section-heading-variants.md new file mode 100644 index 0000000..716f0fd --- /dev/null +++ b/docs/roadmap/afd-section-heading-variants.md @@ -0,0 +1,141 @@ +# NWS AFD Section Heading Variants + +## Status + +Proposed and unimplemented. + +## Problem + +The NWS Area Forecast Discussion parser recognizes supported section headings +only when the section name is followed immediately by an ellipsis, for example: + +```text +.SHORT TERM... (Through Monday) +.LONG TERM... (Tuesday through Sunday) +.AVIATION... (For the 18z TAFs through 18z Monday) +``` + +NWS offices also publish slash-qualified headings: + +```text +.SHORT TERM /THROUGH MONDAY/... +.LONG TERM /MONDAY NIGHT THROUGH SUNDAY/... +.AVIATION /18Z TAFS THROUGH 18Z MONDAY/... +``` + +The current exact-prefix discovery logic does not find slash-qualified short- +or long-term sections. Its boundary detection also does not recognize a +slash-qualified aviation heading, so preceding section text can absorb aviation +content. The forecast discussion can still normalize successfully, leaving +downstream consumers with missing or incorrectly bounded structured prose. + +## Feature Objective + +Make supported AFD section recognition consistent across section discovery, +section-boundary detection, and qualifier extraction while preserving the +existing canonical forecast-discussion contract. + +## Targeted End State + +One NWS-specific heading parser classifies a trimmed AFD line as either a +supported heading or a non-heading. For a supported heading it provides: + +- the canonical section identity; +- the qualifier, if present; and +- enough information for discovery and boundary detection to use the same + recognition result. + +The supported section identities remain: + +- `KEY MESSAGES`; +- `SHORT TERM`; +- `LONG TERM`; and +- `AVIATION`. + +`KEY MESSAGES`, `SHORT TERM`, and `LONG TERM` participate in the existing +canonical extraction behavior. `AVIATION` remains a recognized boundary only; +it does not become a canonical payload field. + +For each supported identity, the parser accepts both heading families: + +```text +.
... +.
//... +``` + +Qualifier handling is deterministic: + +- leading and trailing whitespace is removed; +- the enclosing slash pair used by the slash-qualified form is removed; +- qualifier content otherwise retains its published text, including + parentheses in the existing ellipsis-first form; and +- a heading without a qualifier yields an empty qualifier. + +The same classification rules govern discovery and termination. A recognized +heading ends the preceding section without becoming part of its text. Unknown +section names, unrelated dotted lines, and malformed slash-qualified lines are +not treated as supported headings. + +## Scope of Work + +The feature includes: + +- consolidating supported-heading recognition and qualifier extraction in + `internal/providers/nws`; +- applying that recognition consistently to section lookup and section + boundaries; +- preserving key-message parsing and short- and long-term section mapping; +- covering legacy, slash-qualified, mixed-format, and malformed headings with + focused provider-parser tests; and +- verifying that parsed short- and long-term sections continue through the NWS + normalizer into the existing canonical payload. + +Existing ellipsis-first fixtures and behavior remain regression coverage. New +fixtures should be minimal and representative rather than copies of complete +upstream bulletins unless a full bulletin is needed to prove an interaction. + +## Compatibility and Contracts + +This is a provider-parsing compatibility improvement. It does not change: + +- event kinds or raw and canonical schema identifiers; +- canonical models or JSON field names; +- source configuration or polling behavior; +- event envelope or effective-time behavior; +- Postgres tables or event-to-row mapping; or +- downstream sink and consumer responsibilities. + +Current-behavior documentation should be updated only if implementation reveals +an externally observable contract change beyond the scope defined here. + +## Acceptance Criteria + +The feature is complete when automated tests demonstrate that: + +- all currently accepted ellipsis-first headings produce unchanged results; +- slash-qualified short- and long-term headings are discovered and their + qualifiers exclude slash delimiters; +- slash-qualified recognized headings correctly terminate a preceding section, + including an aviation heading following a long-term section; +- documents mixing the two heading families are parsed correctly; +- headings without qualifiers retain existing behavior; +- unknown or malformed heading-like lines do not create supported sections or + prematurely terminate one; +- the NWS forecast-discussion normalizer emits populated canonical short- and + long-term fields for representative slash-qualified input without adding an + aviation field; and +- the full repository test suite passes. + +## Non-Goals + +This feature does not: + +- add new canonical forecast-discussion sections; +- add aviation prose to the canonical model; +- recognize arbitrary or previously unsupported NWS section families; +- introduce heuristic section-name matching or AFD summarization; +- preserve complete raw AFD documents in canonical payloads; or +- change schemas, persistence contracts, configuration, or downstream APIs. + +Support for additional section identities or broader AFD structure should be +driven by a separate consumer requirement and roadmap. diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md new file mode 100644 index 0000000..ca7522d --- /dev/null +++ b/docs/roadmap/implementation.md @@ -0,0 +1,207 @@ +# NWS AFD Section Heading Variants Implementation Plan + +## Purpose + +Implement the end state defined in +[`afd-section-heading-variants.md`](afd-section-heading-variants.md): recognize +legacy ellipsis-first and slash-qualified NWS Area Forecast Discussion headings +consistently during section discovery, boundary detection, and qualifier +extraction without changing the canonical event contract. + +Complete the stages below in order. Keep each stage limited to the files and +behavior it names, and leave the repository passing its tests before proceeding. + +## Cross-Stage Constraints + +- Keep all provider-format parsing in `internal/providers/nws`. +- Use only the Go standard library; do not add a dependency. +- Do not change `model`, `standards`, source configuration, polling behavior, + schemas, event-envelope behavior, Postgres mapping, or sink behavior. +- Preserve the existing case-sensitive supported section identities: `KEY + MESSAGES`, `SHORT TERM`, `LONG TERM`, and `AVIATION`. +- Continue exposing only key messages, short term, and long term in the parsed + and canonical forecast-discussion payloads. `AVIATION` is a boundary marker, + not a new output field. +- Preserve source body lines for the existing prose parsers; heading + normalization must not alter section text, issue-time parsing, signature + trimming, or paragraph joining. +- Treat unknown section names and malformed slash-qualified lines as ordinary + body lines, not as recognized boundaries. +- Use deterministic unit tests and the checked-in fixture. Do not contact live + NWS services. + +## Stage 1: Centralize Heading Parsing and Section Extraction + +Implement the provider-level parsing primitive and make it the sole source of +heading identity and qualifier interpretation. + +1. In `internal/providers/nws/forecast_discussion.go`, define unexported string + constants for the four supported section identities. Use those constants in + `ParseForecastDiscussionText` instead of repeating string literals. +2. Add an unexported `forecastDiscussionSectionHeading` value with `section` + and `qualifier` fields, plus a + `parseForecastDiscussionSectionHeading(string) (forecastDiscussionSectionHeading, bool)` + helper. The helper must trim outer line whitespace and implement exactly + these two case-sensitive, whole-line grammars: + + ```text + ^\.(KEY MESSAGES|SHORT TERM|LONG TERM|AVIATION)\.\.\.(.*)$ + ^\.(KEY MESSAGES|SHORT TERM|LONG TERM|AVIATION)[ \t]+/([^/]+)/\.\.\.$ + ``` + + Keep the supported-identity alternation in one shared pattern constant used + to build both regular expressions, so the accepted identity list cannot + drift between the two forms. +3. For the ellipsis-first form, trim leading and trailing whitespace from the + text following the ellipsis and otherwise preserve it verbatim. This retains + current results such as `(Through Late Sunday Night)` and permits an empty + qualifier. +4. For the slash-qualified form, require at least one space or tab before the + opening slash, forbid embedded slash characters, require the closing slash + immediately before the final ellipsis, and reject a qualifier that becomes + empty after trimming. Return the trimmed inner text without either slash. + Do not case-fold identities or accept trailing text after the final + ellipsis. +5. Add an unexported `forecastDiscussionSectionBlock` containing the parsed + heading and its body lines. Change `extractForecastDiscussionSection` to + return this value. Find the requested section by calling the new heading + parser and comparing its `section` field to the requested identity; remove + the current constructed exact-prefix lookup. +6. While collecting a block body, retain the existing terminators `&&`, `$$`, + and lines containing `WATCHES/WARNINGS/ADVISORIES`. Also stop before every + subsequent line recognized by the new heading parser, including a heading + immediately following the current heading. Remove the existing + `j > i+1` exception so an empty section cannot absorb the next heading. + Preserve original, untrimmed body lines in the returned block. +7. Update the consumers of the extracted block: + + - pass only `block.body` to key-message parsing and change + `parseForecastDiscussionKeyMessages` so it no longer assumes the heading + occupies element zero; + - make `parseForecastDiscussionTextSection` consume the block, initialize + `Qualifier` directly from `block.heading.qualifier`, and process + `block.body` with the existing issue-time and prose logic; and + - remove `forecastDiscussionHeaderRE`, + `isForecastDiscussionSectionHeader`, and + `parseForecastDiscussionQualifier` after all callers use the centralized + helper. +8. In `internal/providers/nws/forecast_discussion_test.go`, add a table-driven + unit test for `parseForecastDiscussionSectionHeading` with, at minimum: + + - ellipsis-first headings for all four identities; + - an ellipsis-first heading with no qualifier; + - slash-qualified headings for all four identities; + - leading/trailing outer whitespace and multiple spaces before a slash; + - exact qualifier expectations showing that legacy parentheses remain and + slash delimiters are removed; and + - negative cases for an unknown identity, missing opening or closing slash, + an embedded slash, a whitespace-only slash qualifier, missing required + whitespace before the slash, missing final ellipsis, trailing text after a + slash-qualified heading, and a lowercase section identity. + +Run and pass: + +```sh +gofmt -w internal/providers/nws/forecast_discussion.go internal/providers/nws/forecast_discussion_test.go +go test ./internal/providers/nws +``` + +Do not proceed until legacy provider tests still pass and the centralized +heading test covers every accepted identity in both forms. + +## Stage 2: Add Provider and Normalizer Regression Coverage + +Prove section discovery, boundary behavior, and canonical propagation using the +centralized parser. No production normalizer changes should be necessary. + +1. In `internal/providers/nws/forecast_discussion_test.go`, add focused tests of + `extractForecastDiscussionSection` using small line slices without `&&` + between sections. Cover: + + - a slash-qualified `LONG TERM` block followed by a slash-qualified + `AVIATION` heading, asserting that the long-term qualifier is normalized, + only long-term prose is returned, and aviation heading/body text is + excluded; + - a recognized heading immediately following another recognized heading, + asserting that the first block has an empty body; and + - unknown and malformed heading-like lines inside a block, asserting that + they remain in the body and do not terminate it before a real terminator or + recognized heading. +2. Add a provider end-to-end regression based on + `testdata/forecast_discussion_sample.html`. Derive the input in the test by + replacing selected fixture headings rather than duplicating the full HTML + fixture. Guard each replacement with an assertion that its original heading + exists so fixture drift cannot make the test pass without exercising the new + syntax. +3. Make that derived bulletin intentionally mix heading families: retain the + legacy `KEY MESSAGES` heading, convert `SHORT TERM` and `LONG TERM` to + slash-qualified headings, and convert `AVIATION` to a slash-qualified + heading. Parse it through `ParseForecastDiscussionHTML` and assert: + + - key messages are unchanged; + - short- and long-term sections are non-nil; + - their qualifiers equal the inner slash text with no slash delimiters or + legacy parentheses; + - their existing issued times and representative prose remain intact; and + - long-term text contains no aviation heading or aviation prose. +4. In `internal/normalizers/nws/forecast_discussion_test.go`, add a regression + that derives the same mixed-format HTML from the shared fixture and passes it + through `ForecastDiscussionNormalizer.Normalize`. Assert the existing kind, + canonical schema, and effective time; exact normalized short- and long-term + qualifiers; representative section text; and absence of aviation content + from long-term text. +5. Marshal the slash-qualified normalizer result and assert that it has no + top-level `aviation` or generic `sections` key. Do not add or alter canonical + model fields to satisfy this test. + +Run and pass: + +```sh +gofmt -w internal/providers/nws/forecast_discussion_test.go internal/normalizers/nws/forecast_discussion_test.go +go test ./internal/providers/nws ./internal/normalizers/nws +``` + +Do not introduce a second full-bulletin fixture unless the checked-in fixture +cannot express a required interaction through guarded heading replacement. + +## Stage 3: Validate the Feature and Close the Roadmap + +Verify the complete change against repository policy and record completion only +after all behavior is proven. + +1. Review the final diff and confirm production changes are confined to the NWS + provider parser. Test changes should be confined to the owning provider and + NWS normalizer packages. Do not retain incidental model, schema, config, + source, sink, Postgres, or current-behavior documentation edits introduced + during implementation, and preserve all pre-existing user work. +2. Run formatting on every changed Go file, then run uncached focused tests and + the full repository suite: + + ```sh + gofmt -w \ + internal/providers/nws/forecast_discussion.go \ + internal/providers/nws/forecast_discussion_test.go \ + internal/normalizers/nws/forecast_discussion_test.go + go test -count=1 ./internal/providers/nws ./internal/normalizers/nws + go test -count=1 ./... + git diff --check + ``` +3. Confirm each acceptance criterion in + [`afd-section-heading-variants.md`](afd-section-heading-variants.md) is + represented by a passing automated test. In particular, verify legacy + compatibility, mixed-format input, slash-qualified aviation termination, + malformed-line non-recognition, normalized qualifiers, and unchanged + canonical wire shape. +4. Because this feature does not alter a public contract, do not change + `docs/integrations/events.md`, `docs/integrations/postgres.md`, consumer docs, + configuration docs, or examples. If implementation appears to require such + a change, stop and reassess the implementation against the feature roadmap + instead of expanding scope. +5. After all checks pass, change the feature roadmap status from `Proposed and + unimplemented.` to `Implemented.` Do not mark it implemented earlier. + +## Open Questions + +None. The feature roadmap and this plan fix the accepted grammar, normalization +rules, architectural boundary, test coverage, and canonical compatibility +requirements.