15 KiB
Testing Policy
Purpose
Our tests exist to make incorrect changes expensive and correct changes cheap.
We do not optimize for test count, line coverage, exhaustive isolation, or the fewest possible tests. We optimize for sufficient confidence in important behavior while imposing as little unnecessary friction as possible on future development.
Every Test Has A Cost
Every test has an immediate cost and a continuing lifetime cost. It must be written, reviewed, executed, understood, diagnosed when it fails, updated when legitimate behavior changes, and maintained as fixtures and dependencies evolve.
Tests also create cognitive and architectural friction. They can constrain refactoring, duplicate policy, slow feedback, add noise to failures, and cause harmless implementation changes to require unrelated suite edits.
A test is warranted when the confidence it provides justifies those costs. Apply that judgment at two levels:
- Per test: What realistic defect does this test detect, how consequential would it be, and is that protection worth the test's lifetime cost?
- Across the suite: Does this collection provide materially more confidence than a smaller, simpler suite would?
Prefer a lean suite that provides sufficient confidence in the risks that matter without redundant or low-value tests. Some friction is intentional: tests should make dangerous changes, such as corrupting state, breaking compatibility, violating security boundaries, or reintroducing subtle defects, require deliberate review. They should not make ordinary internal changes needlessly expensive.
Maintenance cost is not a reason to omit testing by default. When omitting a plausible test, be able to explain why the protected failure is low-risk, already covered, obvious, reversible, or cheaper to detect elsewhere. Favor testing when failure would be consequential, subtle, or difficult to observe.
Default Testing Style
Use a classical or Detroit-style approach:
- Test observable behavior, resulting state, contracts, and invariants.
- Use real internal collaborators when they are fast and deterministic.
- Use fakes, stubs, or mocks primarily at expensive, nondeterministic, destructive, or external boundaries.
- Prefer package-level behavioral tests over tests coupled to private helpers or internal call sequences.
- Test exact collaborator interactions only when the interaction itself is a requirement.
Weatherreporter's important seams include clocks, Promptkit executors, HTTP services, Distributor uploads, filesystem roots, environment-backed secrets, and any future source of randomness or nondeterminism.
Execution Requirements
The development guide owns baseline repository validation. The default test suite is:
go test ./...
Run race-enabled tests when a change affects concurrent execution, goroutine
lifecycle, shared mutable state, or cancellation coordination. Use a focused
package command while iterating and go test -race ./... when the risk crosses
package boundaries.
Tests in the default suite must be deterministic, offline, and independent of real credentials. They must not invoke live Weather API, Promptkit providers, or Distributor services or depend on other mutable external infrastructure. Tests that require live infrastructure must be explicitly opt-in and clearly separated from the default suite.
Control clocks, environment variables, filesystem roots, and machine-specific
state when they affect behavior. Tests must be safe to repeat and must not
depend on execution order or state left by an earlier test. Tests that modify
process-global state may remain serial; use t.Parallel() only when the test
and its collaborators are actually safe to run concurrently.
Test Types And Assets
Use each test type where it protects a distinct risk:
- Unit and package tests protect focused domain behavior and invariants through the narrowest stable boundary.
- Contract tests protect CLI behavior, configuration, durable artifacts, schemas, templates, integration formats, compatibility, and stable error identity.
- Integration tests use real deterministic collaborators when correctness depends on their interaction, while replacing live or nondeterministic external boundaries.
- App and CLI tests protect representative assembled generation, batch, inspection, persistence, and notification workflows.
- Fixtures must be minimal, synthetic, versioned with the behavior they exercise, and free of credentials or private data.
- Golden files are appropriate only when the complete output is intentionally stable and semantic review of updates is practical.
- Failure-path tests should cover consequential malformed input, dependency failure, cancellation, partial results, and recovery behavior.
What Deserves Tests
Prioritize tests for:
- CLI, configuration, artifact, template, integration, and package contracts.
- Meteorological domain rules and important invariants.
- Boundary conditions and malformed input.
- Failure handling, cancellation, retries, recovery, and partial success.
- Serialization, schemas, compatibility, and round trips.
- Previously observed or plausible regressions.
- Representative app and CLI workflows.
A package-level contract is behavior relied upon by another package or major collaborator, not every observable implementation detail.
For data integrity, destructive operations, compatibility, security, concurrency, idempotency, or recovery, presume that durable tests are required unless the behavior is already credibly protected at another layer.
Do not add tests merely because a function, branch, or line exists. Do not add a test when the same meaningful risk is already adequately protected elsewhere.
Choose The Right Boundary
Test through the narrowest stable boundary that expresses the behavior clearly. That may be:
- a small pure function when dense domain logic is clearest there;
- a package operation when several internal collaborators jointly produce the behavior; or
- a larger integration or app boundary when correctness emerges from interaction.
Do not force every behavior through oversized workflow tests. Do not test every private helper merely because it exists. Choose the boundary that provides durable confidence with the least incidental coupling.
Test Behavior, Not Implementation
A test should protect a decision, contract, or invariant, not memorialize the current implementation. Before adding or retaining a test, ask:
What realistic defect would this test catch?
A test is suspect when its main purpose is to detect that someone changed a private constant, renamed or split a helper, reordered equivalent operations, changed incidental formatting, replaced one correct algorithm with another, or refactored private structure without changing behavior.
Refactoring should normally require no test edits unless the changed structure is itself contractual. A test can be factually correct and still have negative value when the behavior it protects is too incidental to justify its future cost.
Use these expectations when evaluating failures:
| Change | Expected effect on tests |
|---|---|
| Internal refactor that preserves behavior | Existing tests should normally remain unchanged and pass. |
| Internal default change with no contractual significance | Tests should normally derive expectations from configuration or relationships rather than duplicate the old value. |
| Intentional change to user-visible behavior, policy, schema, or compatibility | Relevant tests should be reviewed and changed deliberately. |
| Accidental contract or invariant violation | Tests should fail; fix production code rather than rewriting tests to accept the defect. |
A failing test is not necessarily a test that should be edited. Many tests may correctly fail because of one production defect. The maintenance smell is a correct internal change that requires unrelated expectation changes throughout the suite.
Separate Mechanism From Policy
Do not duplicate configurable thresholds and defaults throughout the suite. Test mechanisms relationally: a configured valid value is accepted, a value outside the permitted relationship is rejected, and runtime behavior respects the configured value.
Test an exact default when its literal value is itself a documented user, operational, safety, protocol, or compatibility contract. The same distinction applies to timeouts, capacities, retry counts, ranges, thresholds, and output limits.
When concurrency limits are introduced, distinguish configuration enforcement from runtime enforcement. Validate accepted and rejected settings separately from measuring whether observed peak concurrency respects the configured limit.
Avoid Semantic Duplication
Each behavior should have a clear test owner:
- CLI parser tests own arguments, flags, and command construction.
- Config tests own loading, precedence, defaults, secrets, and validation.
- Domain tests own weather transformations and invariants.
- Adapter tests own HTTP, Promptkit/provider, and upload boundaries.
- Orchestrator tests own workflow ordering, persistence, partial success, and failure propagation.
- State tests own path derivation, atomic artifacts, lookup, and round trips.
- Template and generated-text tests own schemas, render contexts, and rendered output contracts.
Higher-level tests should not repeat every lower-level case. Tests that are individually reasonable may still be collectively redundant; assess the marginal protection of each additional test.
Use Test Doubles Deliberately
Choose the least elaborate double that provides the required control or observation:
- Prefer real collaborators when they are fast and deterministic.
- Use small in-memory fakes when realistic stateful behavior helps.
- Use stubs when a dependency only needs controlled responses.
- Use mocks when the interaction itself is contractual.
Mocks are appropriate for requirements such as uploading exactly once, saving metadata before notification, propagating cancellation to Promptkit, or avoiding an external call after an earlier workflow failure. Do not use mocks merely to isolate every object or reproduce the implementation's call graph.
Go-Specific Guidance
Use:
- table-driven tests for meaningful behavioral categories and boundaries;
t.TempDir()for real filesystem behavior;httptest.Serverfor realistic Weather API interactions;- test-controlled clocks for periods and RunIDs;
- fake Promptkit executors or provider clients for Promptkit behavior;
- fake upload clients for Distributor behavior;
- fuzz tests when parsers, normalization, or path handling have a broad and consequential input space;
- golden files only when complete output stability is intentional; and
- a small number of representative app and CLI workflow tests.
Avoid exact error-string assertions unless wording is contractual. Prefer
errors.Is, errors.As, typed errors, structured fields, or the smallest
stable semantic fragment that identifies the failure. At CLI boundaries,
prefer structured summaries, exit behavior, and stable classifications over
snapshots of complete diagnostic wording.
Golden-file updates must require an explicit local flag. Ordinary validation must never update golden files automatically, and maintainers must inspect the semantic diff before accepting an update.
Keep tests readable and direct. Helpers and fixture frameworks must earn their maintenance cost; do not build elaborate infrastructure for small or isolated needs.
Coverage
Coverage is a diagnostic, not a target. Use it to find untested critical branches and unexpectedly weak packages. Do not write low-value tests solely to increase a percentage or infer quality from coverage alone.
Pure domain logic will often warrant higher coverage than CLI wiring or thin external adapters. Uneven coverage is acceptable when it reflects risk.
Regression Tests
A bug fix should normally include a regression test that fails before the fix and passes afterward. Prefer the narrowest durable test of the violated contract or invariant.
Retain the test when the defect could realistically recur and its consequences justify the ongoing cost. Remove or consolidate it if the design makes recurrence implausible or a stronger invariant test subsumes it.
Deleting Or Rewriting Tests
Tests are maintained code, not permanent historical artifacts. Delete or rewrite a test when its maintenance cost exceeds the confidence it provides. Candidates include tests that:
- require edits after harmless internal changes;
- assert private constants without protecting a real contract;
- duplicate the same policy across several layers;
- verify mock choreography rather than outcomes;
- snapshot large amounts of incidental output;
- protect risks already covered more effectively elsewhere; or
- are flaky, misleading, obsolete, or no longer correspond to a plausible failure.
Test removal must be deliberate and within the scope of the change. Identify the behavior the test protected and show that the behavior is covered more effectively elsewhere or that the failure is no longer plausible enough to justify durable coverage. Replace several brittle tests with one stronger behavior or invariant test when appropriate.
Do not delete or weaken a test merely because it fails after a production change. First determine whether the failure exposes an accidental regression, an intentional contract change, or an implementation-coupled assertion.
Reviewing A Proposed Test
When a proposed test's value or durability is not self-evident, ask:
- What realistic defect would it catch, and how consequential is that defect?
- Is the behavior already protected elsewhere?
- Which layer should own the test?
- Does it assert a durable contract or incidental implementation detail?
- What should cause it to fail, and what legitimate changes should not?
- Could a smaller or more direct test protect the same risk?
- What ongoing maintenance, execution, and diagnostic cost will it impose?
Written answers are not required for every routine test. Do not add a test when its expected lifetime cost exceeds its expected protective value.
Definition Of Sufficient
A suite is sufficient when:
- important contracts and invariants are protected;
- meaningful boundaries and failure modes are exercised;
- consequential regressions are credibly protected against silent recurrence;
- data integrity, destructive operations, compatibility, security, concurrency, idempotency, and recovery receive risk-appropriate protection;
- external boundaries have realistic local integration coverage;
- representative complete workflows are tested;
- failures provide useful signal rather than redundant noise; and
- legitimate internal changes usually do not require test edits.
Sufficiency is a risk judgment, not a coverage percentage or test count. Reassess it as Weatherreporter, its users, and the consequences of failure evolve.
The governing rule is:
Test heavily where failure is consequential, subtle, or difficult to detect after the fact. Test lightly where failure is obvious, reversible, and inexpensive.