Skip to content

NSI (US flood exposure & loss) — API reference#

US object-level flood exposure and loss subpackage — earthlens.nsi. Background and usage are covered under the other pages in this section (Introduction, Usage); this page is the rendered API. All three sources are public, so there is no auth module.

earthlens.nsi #

US object-level flood exposure & loss backend (earthlens.nsi).

One backend over three keyless US-federal REST sources, selected by a source= discriminator:

  • structures (default) — USACE National Structure Inventory (NSI): building points with occupancy, replacement value, foundation type/height, and area; vector.
  • nfhl — FEMA National Flood Hazard Layer: regulatory flood zones (FLD_ZONE, SFHA_TF) from the ArcGIS S_Fld_Haz_Ar layer; vector.
  • nfip — FEMA NFIP redacted claims (v3): flood-insurance claim records with paid amounts, via the OpenFEMA OData endpoint; tabular.

Output is per instance: the resolved source's output_kind decides whether :meth:~earthlens.nsi.backend.NSI.download returns a pyramids :class:~pyramids.feature.collection.FeatureCollection (structures/nfhl) or a :class:pandas.DataFrame (nfip). All three are public-domain and keyless — no auth. A spatial/attribute bound is required (no unbounded national pull), and aggregate= is rejected (these are records, not gridded rasters).

The public surface is the :class:Catalog (source name -> endpoint + output kind + field map) and its :class:Source rows, the :class:NSI backend, and the pure query/geometry helpers.

Catalog #

Bases: AbstractCatalog[Source]

Source catalog for the NSI backend.

Reads the bundled nsi_data_catalog.yaml (shipped as package data) and exposes its sources: block as a map of :class:Source rows keyed by name. Instantiate with no arguments (Catalog()). Resolve one row with :meth:get and list the shipped keys with :meth:available.

Attributes:

Name Type Description
datasets dict[str, Source]

Map from source name to its :class:Source row.

Examples:

  • Resolve a source and read its output kind:
    >>> from earthlens.nsi import Catalog
    >>> cat = Catalog()
    >>> cat.get("structures").output_kind
    'vector'
    >>> cat.get("nfip").output_kind
    'tabular'
    >>> cat.available()
    ['nfhl', 'nfip', 'structures']
    
Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
class Catalog(AbstractCatalog[Source]):
    """Source catalog for the NSI backend.

    Reads the bundled `nsi_data_catalog.yaml` (shipped as package data) and
    exposes its `sources:` block as a map of :class:`Source` rows keyed by
    name. Instantiate with no arguments (`Catalog()`). Resolve one row with
    :meth:`get` and list the shipped keys with :meth:`available`.

    Attributes:
        datasets: Map from source name to its :class:`Source` row.

    Examples:
        - Resolve a source and read its output kind:
            ```python
            >>> from earthlens.nsi import Catalog
            >>> cat = Catalog()
            >>> cat.get("structures").output_kind
            'vector'
            >>> cat.get("nfip").output_kind
            'tabular'
            >>> cat.available()
            ['nfhl', 'nfip', 'structures']

            ```
    """

    _catalog_kind: str = "NSI catalog"
    _entry_noun: str = "sources"

    #: The source rows live in the base :attr:`datasets` field so the inherited
    #: dict surface and :meth:`get_dataset`'s did-you-mean hint work unchanged.
    datasets: dict[str, Source] = Field(default_factory=dict)

    @classmethod
    def _autoload(cls) -> dict[str, Any]:
        """Read the bundled catalog from disk.

        Returns:
            dict[str, Any]: The `datasets` map read from the bundled catalog.
        """
        return {"datasets": Catalog.load().datasets}

    @classmethod
    def load(cls, catalog_path: Path | None = None) -> Catalog:
        """Read the NSI catalog from disk.

        Args:
            catalog_path: Path to the catalog YAML. Defaults to the module-level
                :data:`CATALOG_PATH`.

        Returns:
            A fully-populated :class:`Catalog`.

        Raises:
            ValueError: If the file has no `sources:` block, or a row fails
                validation.
        """
        catalog_path = catalog_path if catalog_path is not None else CATALOG_PATH
        return cls(datasets=dict(_load_catalog_data(catalog_path)))

    def get(self, source_id: str) -> Source:
        """Resolve a source name to its :class:`Source` row.

        Thin wrapper over the inherited :meth:`get_dataset`, which raises a
        `ValueError` with a did-you-mean hint on an unknown key.

        Args:
            source_id: A shipped source name (`"structures"`, `"nfhl"`,
                `"nfip"`).

        Returns:
            Source: The matching catalog row.

        Raises:
            ValueError: If `source_id` is not a known source; the message names
                the catalog kind and, when a close match exists, adds a
                did-you-mean hint.
        """
        return cast("Source", self.get_dataset(source_id))

    def available(self) -> list[str]:
        """Return the sorted list of shipped source names.

        Returns:
            list[str]: Every catalog key, sorted.
        """
        return sorted(self.datasets)

available() #

Return the sorted list of shipped source names.

Returns:

Type Description
list[str]

list[str]: Every catalog key, sorted.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
def available(self) -> list[str]:
    """Return the sorted list of shipped source names.

    Returns:
        list[str]: Every catalog key, sorted.
    """
    return sorted(self.datasets)

get(source_id) #

Resolve a source name to its :class:Source row.

Thin wrapper over the inherited :meth:get_dataset, which raises a ValueError with a did-you-mean hint on an unknown key.

Parameters:

Name Type Description Default
source_id str

A shipped source name ("structures", "nfhl", "nfip").

required

Returns:

Name Type Description
Source Source

The matching catalog row.

Raises:

Type Description
ValueError

If source_id is not a known source; the message names the catalog kind and, when a close match exists, adds a did-you-mean hint.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
def get(self, source_id: str) -> Source:
    """Resolve a source name to its :class:`Source` row.

    Thin wrapper over the inherited :meth:`get_dataset`, which raises a
    `ValueError` with a did-you-mean hint on an unknown key.

    Args:
        source_id: A shipped source name (`"structures"`, `"nfhl"`,
            `"nfip"`).

    Returns:
        Source: The matching catalog row.

    Raises:
        ValueError: If `source_id` is not a known source; the message names
            the catalog kind and, when a close match exists, adds a
            did-you-mean hint.
    """
    return cast("Source", self.get_dataset(source_id))

load(catalog_path=None) classmethod #

Read the NSI catalog from disk.

Parameters:

Name Type Description Default
catalog_path Path | None

Path to the catalog YAML. Defaults to the module-level :data:CATALOG_PATH.

None

Returns:

Type Description
Catalog

A fully-populated :class:Catalog.

Raises:

Type Description
ValueError

If the file has no sources: block, or a row fails validation.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
@classmethod
def load(cls, catalog_path: Path | None = None) -> Catalog:
    """Read the NSI catalog from disk.

    Args:
        catalog_path: Path to the catalog YAML. Defaults to the module-level
            :data:`CATALOG_PATH`.

    Returns:
        A fully-populated :class:`Catalog`.

    Raises:
        ValueError: If the file has no `sources:` block, or a row fails
            validation.
    """
    catalog_path = catalog_path if catalog_path is not None else CATALOG_PATH
    return cls(datasets=dict(_load_catalog_data(catalog_path)))

NSI #

Bases: AbstractDataSource

US flood exposure / loss backend (per-instance output).

Resolves source= to its catalog row, issues the one bounded request that source needs, and returns a :class:~pyramids.feature.collection.FeatureCollection (structures/nfhl) or a :class:pandas.DataFrame (nfip). US-only: a non-US box returns an empty result, not an error (G4).

Attributes:

Name Type Description
OUTPUT_KIND OutputKind

Set per instance in :meth:__init__ from the resolved source's output_kind. The facade reads it to gate aggregate= and to know the return shape.

Source code in libs/providers/hazards/src/earthlens/nsi/backend.py
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
class NSI(AbstractDataSource):
    """US flood exposure / loss backend (per-instance output).

    Resolves `source=` to its catalog row, issues the one bounded request that
    source needs, and returns a
    :class:`~pyramids.feature.collection.FeatureCollection` (`structures`/`nfhl`)
    or a :class:`pandas.DataFrame` (`nfip`). US-only: a non-US box returns an
    empty result, not an error (`G4`).

    Attributes:
        OUTPUT_KIND: Set **per instance** in :meth:`__init__` from the resolved
            source's `output_kind`. The facade reads it to gate `aggregate=` and
            to know the return shape.
    """

    OUTPUT_KIND: OutputKind = "vector"

    AGGREGATE_REFUSAL_REASON = (
        "NSI/FEMA sources are object-level records (structures, flood zones, "
        "insurance claims), not gridded rasters, so there is no meaningful "
        "gridded reduction. Call download() without aggregate="
    )

    #: These sources are point-in-time inventories / claim records, not a time
    #: series, so a missing `start` / `end` is legal.
    REQUIRES_TIME_WINDOW = False

    def __init__(
        self,
        start: str | None = None,
        end: str | None = None,
        lat_lim: list[float] | None = None,
        lon_lim: list[float] | None = None,
        temporal_resolution: str = "snapshot",
        path: Path | str | None = None,
        fmt: str = "%Y-%m-%d",
        source: str = "structures",
        fips: str | None = None,
        filters: dict[str, str | int] | None = None,
        max_records: int | None = None,
        output_format: OutputFormat = "csv",
        session: requests.Session | None = None,
    ):
        """Initialise an NSI backend instance.

        Args:
            start: Inclusive start of an optional window, parsed with `fmt`;
                `None` allowed (these sources are snapshots, not a time series).
            end: Inclusive end of the optional window; `None` allowed.
            lat_lim: `[min_lat, max_lat]` box — required for `nfhl`, and one of
                the two ways to select `structures` (the other is `fips=`).
            lon_lim: `[min_lon, max_lon]` box; see `lat_lim`.
            temporal_resolution: Recorded as the resolution label only.
            path: Output directory for a written tabular (`nfip`) result.
            fmt: `strptime` format for `start` / `end`.
            source: Which source to query — `"structures"` (default), `"nfhl"`,
                or `"nfip"`.
            fips: A 2/5/11/15-digit FIPS code selecting `structures` by
                state/county/tract/block.
            filters: NFIP attribute filter mapping — recognised keys are
                `state` (two-letter), `county` (5-digit FIPS), `year` (loss
                year), and `flood_event` (named event). At least one is required
                for `source="nfip"`, e.g. `filters={"county": "22071", "year": 2005}`.
            max_records: Optional cap on the number of `nfip` records fetched.
            output_format: On-disk format for the `nfip` table — `"csv"`
                (default) or `"parquet"`.
            session: An existing `requests.Session` to reuse for every request.
                Defaults to a fresh session; injectable so the whole path is
                testable with a fake transport.

        Raises:
            ValueError: If `source` is unknown, `output_format` is
                unrecognised, or the source's required bound is missing (`G3`).
        """
        if output_format not in OUTPUT_FORMATS:
            raise ValueError(
                f"output_format must be one of {list(OUTPUT_FORMATS)}, "
                f"got {output_format!r}."
            )

        self._catalog = Catalog()
        self._source: Source = self._catalog.get(source)
        self._fips = fips
        self._filters = dict(filters) if filters else {}
        self._max_records = max_records
        self._output_format: OutputFormat = output_format
        # The structures endpoint is queried with a `POST` carrying the AOI
        # polygon and answers with GeoJSON — a read with no side effect, so
        # replaying it after a transient `5xx` or a dropped connection cannot
        # double-submit anything. Without this the client would decline to
        # retry it, since a `POST` is otherwise assumed unsafe to repeat.
        #
        # Set on the client rather than per call because there is no per-call
        # override, and it is equivalent here: every other request this client
        # makes is a `GET`, which the idempotency gate never blocks. A future
        # `POST` on this client would inherit the promise, so check it still
        # holds before adding one.
        self._http: HttpClient = HttpClient(session=session, retry_unsafe_methods=True)

        # G1 — the per-instance output shape comes from the resolved source.
        self.OUTPUT_KIND = self._source.output_kind

        self._has_box = lat_lim is not None and lon_lim is not None
        self._lat_lim = lat_lim if lat_lim is not None else _GLOBAL_LAT
        self._lon_lim = lon_lim if lon_lim is not None else _GLOBAL_LON
        self._validate_bound()

        super().__init__(
            start=cast("str", start),
            end=cast("str", end),
            variables=[self._source.id],
            temporal_resolution=temporal_resolution,
            lat_lim=self._lat_lim,
            lon_lim=self._lon_lim,
            fmt=fmt,
            path=path,
        )

    def _validate_bound(self) -> None:
        """Enforce the required spatial/attribute bound for the source (`G3`).

        Also validates the `fips` shape, validates only a box that will actually
        be used (fails fast on a malformed one), and warns when a supplied box is
        ignored (`nfip`, or `structures` alongside `fips`).

        Raises:
            ValueError: If `structures` has neither `fips=` nor a box, `fips` is
                not a 2/5/11/15-digit code, `nfhl` has no box, `nfip` has no
                `state`/`county`/`year`/`flood_event`, or a box that will be used
                is not a valid `[min, max]` pair.
        """
        self._validate_fips_shape()
        provider = self._source.provider
        if provider == "nsi":
            self._validate_structures_bound()
        elif provider == "fema-arcgis":
            self._validate_nfhl_bound()
        elif provider == "openfema":
            self._validate_nfip_bound()

    def _validate_fips_shape(self) -> None:
        """Reject a `fips` that is not a 2/5/11/15-digit code."""
        if self._fips is not None and not (
            self._fips.isdigit() and len(self._fips) in _FIPS_LENGTHS
        ):
            raise ValueError(
                "fips must be a 2/5/11/15-digit code (state/county/tract/block); "
                f"got {self._fips!r}."
            )

    def _validate_structures_bound(self) -> None:
        """`structures` needs a `fips` or a box; a box alongside `fips` is warned."""
        if not (self._fips or self._has_box):
            raise ValueError(
                "source='structures' needs a fips= (2/5/11/15-digit) or a "
                "[lat_lim, lon_lim] box; an unbounded national pull is refused."
            )
        if self._fips and self._has_box:
            logger.warning(
                "NSI structures: both fips= and a [lat_lim, lon_lim] box were "
                "given; using fips= and ignoring the box."
            )
        # Validate the box only when it is the selector (no fips).
        if self._has_box and not self._fips:
            bbox_from_limits(self._lat_lim, self._lon_lim)

    def _validate_nfhl_bound(self) -> None:
        """`nfhl` needs a valid box (the ArcGIS query envelope)."""
        if not self._has_box:
            raise ValueError(
                "source='nfhl' needs a [lat_lim, lon_lim] box (the ArcGIS "
                "query envelope); an unbounded national pull is refused."
            )
        bbox_from_limits(self._lat_lim, self._lon_lim)

    def _validate_nfip_bound(self) -> None:
        """`nfip` needs an attribute filter; a supplied box is warned + ignored."""
        # odata_filter validates the keys (raises on an unknown one) and returns
        # None for an empty mapping — so this both fail-fasts and refuses an
        # unbounded pull.
        if _helpers.odata_filter(self._filters) is None:
            raise ValueError(
                "source='nfip' needs a filters= mapping with at least one of "
                f"{sorted(_helpers._NFIP_FILTER_FIELDS)}; an unbounded national "
                "pull is refused."
            )
        if self._has_box:
            logger.warning(
                "NSI nfip: a [lat_lim, lon_lim] box is ignored; nfip filters "
                "by the filters= mapping (state/county/year/flood_event) only."
            )

    def _create_grid(self, lat_lim: list, lon_lim: list) -> SpatialExtent:
        """Return the request's :class:`SpatialExtent`.

        Args:
            lat_lim: `[min_lat, max_lat]` (global sentinel when unbounded).
            lon_lim: `[min_lon, max_lon]`.

        Returns:
            SpatialExtent: The extent for the resolved box (or the globe).
        """
        return SpatialExtent.from_pairs(lat_lim=lat_lim, lon_lim=lon_lim)

    def _check_input_dates(
        self,
        start: str | None,
        end: str | None,
        temporal_resolution: str,
        fmt: str,
    ) -> TemporalExtent:
        """Parse the optional `[start, end]` window into a :class:`TemporalExtent`.

        Dates are not part of the request (these sources are snapshots), so
        `None` bounds are allowed and yield a `None`-dated extent.

        Args:
            start: Inclusive start date string, or `None`.
            end: Inclusive end date string, or `None`.
            temporal_resolution: Recorded as the resolution label.
            fmt: `strptime` format tried first for a string `start` / `end`.

        Returns:
            TemporalExtent: Frozen model with the parsed (or `None`) endpoints.
        """
        start_dt = to_datetime(start, fmt) if start else None
        end_dt = to_datetime(end, fmt) if end else None
        dates = (
            pd.DatetimeIndex([start_dt, end_dt])
            if start_dt is not None and end_dt is not None
            else pd.DatetimeIndex([])
        )
        return TemporalExtent(
            start_date=start_dt,
            end_date=end_dt,
            resolution=temporal_resolution,
            dates=dates,
        )

    def _search(self) -> list[RemoteProduct]:
        """Pin the one product to fetch (the resolved source).

        Returns:
            list[RemoteProduct]: A single product identified by the source name.
        """
        return [RemoteProduct(id=self._source.id, metadata={})]

    def _fetch(self, products: list[RemoteProduct]) -> list[Any]:
        """Route the product to its source and parse the response.

        Args:
            products: The list from :meth:`_search` (one product).

        Returns:
            list: One element — a `FeatureCollection` (`structures`/`nfhl`) or a
                `DataFrame` (`nfip`).
        """
        return [self._fetch_source()]

    def _fetch_source(self) -> FeatureCollection | pd.DataFrame:
        """Fetch and parse the resolved source.

        Returns:
            A `FeatureCollection` (vector) or a `DataFrame` (tabular).

        Raises:
            requests.HTTPError: If the upstream source returns a non-2xx status.
        """
        provider = self._source.provider
        if provider == "nsi":
            return self._fetch_structures()
        if provider == "fema-arcgis":
            return self._fetch_nfhl()
        return self._fetch_nfip()

    def _fetch_structures(self) -> FeatureCollection:
        """Fetch NSI structures by `fips=` (GET) or a box (POST polygon)."""
        endpoint = self._source.endpoint
        if self._fips:
            if len(self._fips) == 2:
                logger.warning(
                    f"NSI structures fips={self._fips!r} is a whole-state pull "
                    "returned in a single response; this can be very large — "
                    "consider a county/tract/block FIPS or a box."
                )
            geojson = self._http.get_json(endpoint, params={"fips": self._fips})
        else:
            body = nsi_polygon_body(self._lat_lim, self._lon_lim)
            geojson = self._http.post(endpoint, json=body).json()
        return to_feature_collection(geojson)

    def _fetch_nfhl(self) -> FeatureCollection:
        """Fetch FEMA NFHL flood zones for the box via the paged ArcGIS query."""
        url = f"{self._source.endpoint}/{self._source.layer_id}/query"
        params = arcgis_envelope(self._lat_lim, self._lon_lim)
        geojson = _helpers.paginate_arcgis(self._http, url, params)
        return to_feature_collection(geojson)

    def _fetch_nfip(self) -> pd.DataFrame:
        """Fetch NFIP claims for the attribute filter, paged, as a `DataFrame`."""
        source = self._source
        filter_str = _helpers.odata_filter(self._filters)
        # Only the curated columns are kept, so project them server-side with
        # $select rather than pulling all ~84 provider fields per record.
        select = ",".join(source.fields.values()) or None
        records, total = _helpers.paginate_nfip(
            self._http,
            source.endpoint,
            cast("str", source.records_key),
            filter_str=filter_str,
            page_size=source.page_size or 1000,
            max_records=self._max_records,
            select=select,
        )
        logger.info(
            f"NSI nfip: {total} claim record(s) match {filter_str!r}; fetched "
            f"{len(records)}"
            + (f" (capped at {self._max_records})" if self._max_records else "")
        )
        if (
            self._max_records is None
            and total is not None
            and total > _LARGE_PULL_THRESHOLD
        ):
            logger.warning(
                f"NSI nfip: {total} records matched with no max_records= cap; "
                "this is a large pull — pass max_records= or a narrower selector "
                "(add year/county to filters=) to bound it."
            )
        return _helpers.records_to_frame(records, source.fields)

    def download(
        self,
        progress_bar: bool = True,
    ) -> pd.DataFrame | FeatureCollection:
        """Fetch the selected source and return its per-instance shape.

        Args:
            progress_bar: Accepted for signature parity; one logical request is
                issued (paged for `nfip`), so this is a no-op.

        Returns:
            A :class:`~pyramids.feature.collection.FeatureCollection` for
            `structures`/`nfhl`, or a :class:`pandas.DataFrame` (also written to
            `root_dir`) for `nfip`.

        Raises:
            requests.HTTPError: If the upstream source returns a non-2xx status.
        """
        results = self._api()
        result = results[0]
        self._log_citation()
        if self.OUTPUT_KIND == "vector":
            logger.info(
                f"NSI {self._source.id}: returned a FeatureCollection "
                f"({len(result)} feature(s))."
            )
            return result
        out_path = self._write_table(result)
        if len(result):
            logger.info(
                f"NSI {self._source.id}: {len(result)} row(s) written to {out_path}."
            )
        else:
            logger.warning(
                f"NSI {self._source.id}: no rows matched; wrote an empty "
                f"(schema-only) table to {out_path}."
            )
        return result

    def _write_table(self, df: pd.DataFrame) -> Path:
        """Write a tabular result to `root_dir` and return the path.

        Args:
            df: The result frame.

        Returns:
            Path: The written CSV / Parquet file path.

        Raises:
            ImportError: If `output_format="parquet"` but `pyarrow` is missing.
        """
        ext = "parquet" if self._output_format == "parquet" else "csv"
        out_path = self.root_dir / f"nsi_{self._source.id}.{ext}"
        if self._output_format == "parquet":
            try:
                df.to_parquet(out_path, index=False)
            except ImportError as exc:  # pragma: no cover - depends on env
                raise ImportError(
                    "Writing Parquet requires 'pyarrow'. Install it (pip "
                    "install pyarrow) or use output_format='csv'."
                ) from exc
        else:
            df.to_csv(out_path, index=False)
        return out_path

    def _log_citation(self) -> None:
        """Log the resolved source's citation once (info, not a warning)."""
        citation = self._source.citation
        if citation:
            logger.info(f"NSI source citation: {citation}")

__init__(start=None, end=None, lat_lim=None, lon_lim=None, temporal_resolution='snapshot', path=None, fmt='%Y-%m-%d', source='structures', fips=None, filters=None, max_records=None, output_format='csv', session=None) #

Initialise an NSI backend instance.

Parameters:

Name Type Description Default
start str | None

Inclusive start of an optional window, parsed with fmt; None allowed (these sources are snapshots, not a time series).

None
end str | None

Inclusive end of the optional window; None allowed.

None
lat_lim list[float] | None

[min_lat, max_lat] box — required for nfhl, and one of the two ways to select structures (the other is fips=).

None
lon_lim list[float] | None

[min_lon, max_lon] box; see lat_lim.

None
temporal_resolution str

Recorded as the resolution label only.

'snapshot'
path Path | str | None

Output directory for a written tabular (nfip) result.

None
fmt str

strptime format for start / end.

'%Y-%m-%d'
source str

Which source to query — "structures" (default), "nfhl", or "nfip".

'structures'
fips str | None

A 2/5/11/15-digit FIPS code selecting structures by state/county/tract/block.

None
filters dict[str, str | int] | None

NFIP attribute filter mapping — recognised keys are state (two-letter), county (5-digit FIPS), year (loss year), and flood_event (named event). At least one is required for source="nfip", e.g. filters={"county": "22071", "year": 2005}.

None
max_records int | None

Optional cap on the number of nfip records fetched.

None
output_format OutputFormat

On-disk format for the nfip table — "csv" (default) or "parquet".

'csv'
session Session | None

An existing requests.Session to reuse for every request. Defaults to a fresh session; injectable so the whole path is testable with a fake transport.

None

Raises:

Type Description
ValueError

If source is unknown, output_format is unrecognised, or the source's required bound is missing (G3).

Source code in libs/providers/hazards/src/earthlens/nsi/backend.py
def __init__(
    self,
    start: str | None = None,
    end: str | None = None,
    lat_lim: list[float] | None = None,
    lon_lim: list[float] | None = None,
    temporal_resolution: str = "snapshot",
    path: Path | str | None = None,
    fmt: str = "%Y-%m-%d",
    source: str = "structures",
    fips: str | None = None,
    filters: dict[str, str | int] | None = None,
    max_records: int | None = None,
    output_format: OutputFormat = "csv",
    session: requests.Session | None = None,
):
    """Initialise an NSI backend instance.

    Args:
        start: Inclusive start of an optional window, parsed with `fmt`;
            `None` allowed (these sources are snapshots, not a time series).
        end: Inclusive end of the optional window; `None` allowed.
        lat_lim: `[min_lat, max_lat]` box — required for `nfhl`, and one of
            the two ways to select `structures` (the other is `fips=`).
        lon_lim: `[min_lon, max_lon]` box; see `lat_lim`.
        temporal_resolution: Recorded as the resolution label only.
        path: Output directory for a written tabular (`nfip`) result.
        fmt: `strptime` format for `start` / `end`.
        source: Which source to query — `"structures"` (default), `"nfhl"`,
            or `"nfip"`.
        fips: A 2/5/11/15-digit FIPS code selecting `structures` by
            state/county/tract/block.
        filters: NFIP attribute filter mapping — recognised keys are
            `state` (two-letter), `county` (5-digit FIPS), `year` (loss
            year), and `flood_event` (named event). At least one is required
            for `source="nfip"`, e.g. `filters={"county": "22071", "year": 2005}`.
        max_records: Optional cap on the number of `nfip` records fetched.
        output_format: On-disk format for the `nfip` table — `"csv"`
            (default) or `"parquet"`.
        session: An existing `requests.Session` to reuse for every request.
            Defaults to a fresh session; injectable so the whole path is
            testable with a fake transport.

    Raises:
        ValueError: If `source` is unknown, `output_format` is
            unrecognised, or the source's required bound is missing (`G3`).
    """
    if output_format not in OUTPUT_FORMATS:
        raise ValueError(
            f"output_format must be one of {list(OUTPUT_FORMATS)}, "
            f"got {output_format!r}."
        )

    self._catalog = Catalog()
    self._source: Source = self._catalog.get(source)
    self._fips = fips
    self._filters = dict(filters) if filters else {}
    self._max_records = max_records
    self._output_format: OutputFormat = output_format
    # The structures endpoint is queried with a `POST` carrying the AOI
    # polygon and answers with GeoJSON — a read with no side effect, so
    # replaying it after a transient `5xx` or a dropped connection cannot
    # double-submit anything. Without this the client would decline to
    # retry it, since a `POST` is otherwise assumed unsafe to repeat.
    #
    # Set on the client rather than per call because there is no per-call
    # override, and it is equivalent here: every other request this client
    # makes is a `GET`, which the idempotency gate never blocks. A future
    # `POST` on this client would inherit the promise, so check it still
    # holds before adding one.
    self._http: HttpClient = HttpClient(session=session, retry_unsafe_methods=True)

    # G1 — the per-instance output shape comes from the resolved source.
    self.OUTPUT_KIND = self._source.output_kind

    self._has_box = lat_lim is not None and lon_lim is not None
    self._lat_lim = lat_lim if lat_lim is not None else _GLOBAL_LAT
    self._lon_lim = lon_lim if lon_lim is not None else _GLOBAL_LON
    self._validate_bound()

    super().__init__(
        start=cast("str", start),
        end=cast("str", end),
        variables=[self._source.id],
        temporal_resolution=temporal_resolution,
        lat_lim=self._lat_lim,
        lon_lim=self._lon_lim,
        fmt=fmt,
        path=path,
    )

download(progress_bar=True) #

Fetch the selected source and return its per-instance shape.

Parameters:

Name Type Description Default
progress_bar bool

Accepted for signature parity; one logical request is issued (paged for nfip), so this is a no-op.

True

Returns:

Name Type Description
A DataFrame | FeatureCollection

class:~pyramids.feature.collection.FeatureCollection for

DataFrame | FeatureCollection

structures/nfhl, or a :class:pandas.DataFrame (also written to

DataFrame | FeatureCollection

root_dir) for nfip.

Raises:

Type Description
HTTPError

If the upstream source returns a non-2xx status.

Source code in libs/providers/hazards/src/earthlens/nsi/backend.py
def download(
    self,
    progress_bar: bool = True,
) -> pd.DataFrame | FeatureCollection:
    """Fetch the selected source and return its per-instance shape.

    Args:
        progress_bar: Accepted for signature parity; one logical request is
            issued (paged for `nfip`), so this is a no-op.

    Returns:
        A :class:`~pyramids.feature.collection.FeatureCollection` for
        `structures`/`nfhl`, or a :class:`pandas.DataFrame` (also written to
        `root_dir`) for `nfip`.

    Raises:
        requests.HTTPError: If the upstream source returns a non-2xx status.
    """
    results = self._api()
    result = results[0]
    self._log_citation()
    if self.OUTPUT_KIND == "vector":
        logger.info(
            f"NSI {self._source.id}: returned a FeatureCollection "
            f"({len(result)} feature(s))."
        )
        return result
    out_path = self._write_table(result)
    if len(result):
        logger.info(
            f"NSI {self._source.id}: {len(result)} row(s) written to {out_path}."
        )
    else:
        logger.warning(
            f"NSI {self._source.id}: no rows matched; wrote an empty "
            f"(schema-only) table to {out_path}."
        )
    return result

Source #

Bases: SummarisedLeaf

One NSI catalog row (one of the three flood sources).

The source name is the parent key in :attr:Catalog.datasets and is stored on the row as :attr:id so a resolved :class:Source is self-describing. Which fields matter depends on :attr:provider; a cross-field validator enforces the source-specific ones (nfhl needs a layer, nfip a records key).

Attributes:

Name Type Description
id str

The source name ("structures", "nfhl", "nfip").

provider str

The upstream service — "nsi", "fema-arcgis", or "openfema".

endpoint str

The base REST URL.

output_kind OutputKind

"vector" (a FeatureCollection) or "tabular" (a DataFrame). Copied onto the backend's OUTPUT_KIND per instance.

long_name str

Human-readable label.

citation str

The source's citation string, logged once on use.

license str

The data licence (all three are US public domain).

fields dict[str, str]

Friendly -> provider field-name map. For the tabular nfip source this shapes the output — the DataFrame is subset to these columns, renamed friendly. For the vector structures / nfhl sources it is informational only: those return the full provider column set unchanged (renaming would drop the many other NSI/ArcGIS fields), and this map documents the notable ones.

layer_id int | None

ArcGIS MapServer layer id (nfhl only — S_Fld_Haz_Ar = 28).

layer_name str | None

ArcGIS layer name (nfhl only).

records_key str | None

JSON envelope key holding the record list (nfip only — NfipClaims).

page_size int | None

OData page size for paged fetches (nfip only).

Examples:

  • Build a source row directly:
    >>> from earthlens.nsi import Source
    >>> row = Source(
    ...     id="nfip",
    ...     provider="openfema",
    ...     endpoint="https://www.fema.gov/api/open/v3/NfipClaims",
    ...     output_kind="tabular",
    ...     records_key="NfipClaims",
    ... )
    >>> row.output_kind
    'tabular'
    
Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
class Source(SummarisedLeaf):
    """One NSI catalog row (one of the three flood sources).

    The source name is the parent key in :attr:`Catalog.datasets` and is stored
    on the row as :attr:`id` so a resolved :class:`Source` is self-describing.
    Which fields matter depends on :attr:`provider`; a cross-field validator
    enforces the source-specific ones (`nfhl` needs a layer, `nfip` a records
    key).

    Attributes:
        id: The source name (`"structures"`, `"nfhl"`, `"nfip"`).
        provider: The upstream service — `"nsi"`, `"fema-arcgis"`, or
            `"openfema"`.
        endpoint: The base REST URL.
        output_kind: `"vector"` (a `FeatureCollection`) or `"tabular"` (a
            `DataFrame`). Copied onto the backend's `OUTPUT_KIND` per instance.
        long_name: Human-readable label.
        citation: The source's citation string, logged once on use.
        license: The data licence (all three are US public domain).
        fields: Friendly -> provider field-name map. For the tabular `nfip`
            source this **shapes the output** — the `DataFrame` is subset to
            these columns, renamed friendly. For the vector `structures` / `nfhl`
            sources it is **informational only**: those return the full provider
            column set unchanged (renaming would drop the many other NSI/ArcGIS
            fields), and this map documents the notable ones.
        layer_id: ArcGIS MapServer layer id (`nfhl` only — `S_Fld_Haz_Ar` = 28).
        layer_name: ArcGIS layer name (`nfhl` only).
        records_key: JSON envelope key holding the record list (`nfip` only —
            `NfipClaims`).
        page_size: OData page size for paged fetches (`nfip` only).

    Examples:
        - Build a source row directly:
            ```python
            >>> from earthlens.nsi import Source
            >>> row = Source(
            ...     id="nfip",
            ...     provider="openfema",
            ...     endpoint="https://www.fema.gov/api/open/v3/NfipClaims",
            ...     output_kind="tabular",
            ...     records_key="NfipClaims",
            ... )
            >>> row.output_kind
            'tabular'

            ```
    """

    _summary_fields = (
        "id",
        "long_name",
        "output_kind",
    )

    model_config = ConfigDict(frozen=True, extra="forbid")

    id: str
    provider: str
    endpoint: str
    output_kind: OutputKind
    long_name: str = ""
    citation: str = ""
    license: str = ""
    fields: dict[str, str] = Field(default_factory=dict)

    # nfhl
    layer_id: int | None = None
    layer_name: str | None = None
    # nfip
    records_key: str | None = None
    page_size: int | None = None

    @model_validator(mode="after")
    def _check_provider_fields(self) -> Source:
        """Enforce the per-source required fields.

        Returns:
            The validated row.

        Raises:
            ValueError: If an `nfhl` row omits `layer_id`/`layer_name`, or an
                `nfip` row omits `records_key`, or `output_kind` disagrees with
                the source family.
        """
        if self.provider == "fema-arcgis":
            if self.layer_id is None or not self.layer_name:
                raise ValueError(
                    f"nfhl source {self.id!r} needs layer_id and layer_name"
                )
            if self.output_kind != "vector":
                raise ValueError(
                    f"nfhl source {self.id!r} must be output_kind 'vector'"
                )
        if self.provider == "openfema":
            if not self.records_key:
                raise ValueError(f"nfip source {self.id!r} needs records_key")
            if self.output_kind != "tabular":
                raise ValueError(
                    f"nfip source {self.id!r} must be output_kind 'tabular'"
                )
        if self.provider == "nsi" and self.output_kind != "vector":
            raise ValueError(
                f"structures source {self.id!r} must be output_kind 'vector'"
            )
        return self

arcgis_envelope(lat_lim, lon_lim, out_fields='*') #

Build the ArcGIS query parameters for the NFHL flood-zone layer.

Parameters:

Name Type Description Default
lat_lim list[float]

[min_lat, max_lat] in degrees.

required
lon_lim list[float]

[min_lon, max_lon] in degrees.

required
out_fields str

Comma-separated attribute list to return ("*" = all).

'*'

Returns:

Name Type Description
dict dict

The params for a GET against the layer's query endpoint — an esriGeometryEnvelope in WGS84, GeoJSON output, geometry on.

Examples:

  • Build the query params for a box:
    >>> from earthlens.nsi import arcgis_envelope
    >>> params = arcgis_envelope([29.95, 29.96], [-90.07, -90.06])
    >>> params["geometryType"]
    'esriGeometryEnvelope'
    >>> params["f"]
    'geojson'
    >>> params["outFields"]
    '*'
    
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def arcgis_envelope(
    lat_lim: list[float], lon_lim: list[float], out_fields: str = "*"
) -> dict:
    """Build the ArcGIS `query` parameters for the NFHL flood-zone layer.

    Args:
        lat_lim: `[min_lat, max_lat]` in degrees.
        lon_lim: `[min_lon, max_lon]` in degrees.
        out_fields: Comma-separated attribute list to return (`"*"` = all).

    Returns:
        dict: The `params` for a GET against the layer's `query` endpoint —
            an `esriGeometryEnvelope` in WGS84, GeoJSON output, geometry on.

    Examples:
        - Build the query params for a box:
            ```python
            >>> from earthlens.nsi import arcgis_envelope
            >>> params = arcgis_envelope([29.95, 29.96], [-90.07, -90.06])
            >>> params["geometryType"]
            'esriGeometryEnvelope'
            >>> params["f"]
            'geojson'
            >>> params["outFields"]
            '*'

            ```
    """
    xmin, ymin, xmax, ymax = bbox_from_limits(lat_lim, lon_lim)
    envelope = {"xmin": xmin, "ymin": ymin, "xmax": xmax, "ymax": ymax}
    return {
        "geometry": json.dumps(envelope),
        "geometryType": "esriGeometryEnvelope",
        "inSR": "4326",
        "outSR": "4326",
        "spatialRel": "esriSpatialRelIntersects",
        "outFields": out_fields,
        "returnGeometry": "true",
        "f": "geojson",
        "where": "1=1",
    }

bbox_from_limits(lat_lim, lon_lim) #

Return (xmin, ymin, xmax, ymax) from lat_lim / lon_lim.

Parameters:

Name Type Description Default
lat_lim list[float]

[min_lat, max_lat] in degrees.

required
lon_lim list[float]

[min_lon, max_lon] in degrees.

required

Returns:

Type Description
tuple[float, float, float, float]

tuple[float, float, float, float]: The envelope as (min_lon, min_lat, max_lon, max_lat).

Raises:

Type Description
ValueError

If either limit is not a two-element [min, max] with min < max.

Examples:

>>> from earthlens.nsi import bbox_from_limits
>>> bbox_from_limits([29.95, 29.96], [-90.07, -90.06])
(-90.07, 29.95, -90.06, 29.96)
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def bbox_from_limits(
    lat_lim: list[float], lon_lim: list[float]
) -> tuple[float, float, float, float]:
    """Return `(xmin, ymin, xmax, ymax)` from `lat_lim` / `lon_lim`.

    Args:
        lat_lim: `[min_lat, max_lat]` in degrees.
        lon_lim: `[min_lon, max_lon]` in degrees.

    Returns:
        tuple[float, float, float, float]: The envelope as
            `(min_lon, min_lat, max_lon, max_lat)`.

    Raises:
        ValueError: If either limit is not a two-element `[min, max]` with
            `min < max`.

    Examples:
        ```python
        >>> from earthlens.nsi import bbox_from_limits
        >>> bbox_from_limits([29.95, 29.96], [-90.07, -90.06])
        (-90.07, 29.95, -90.06, 29.96)

        ```
    """
    if not (lat_lim and lon_lim and len(lat_lim) == 2 and len(lon_lim) == 2):
        raise ValueError(
            "lat_lim and lon_lim must each be a two-element [min, max]; got "
            f"lat_lim={lat_lim!r}, lon_lim={lon_lim!r}."
        )
    ymin, ymax = float(lat_lim[0]), float(lat_lim[1])
    xmin, xmax = float(lon_lim[0]), float(lon_lim[1])
    # Allow a degenerate (min == max) axis — a point / zero-width AOI — to match
    # the base SpatialExtent; only an inverted bound (min > max) is an error.
    if ymin > ymax or xmin > xmax:
        raise ValueError(
            f"each bound needs min <= max; got lat_lim={lat_lim!r}, lon_lim={lon_lim!r}."
        )
    return xmin, ymin, xmax, ymax

clear_catalog_cache() #

Empty the module-level catalog parse cache.

Useful in tests that rewrite the catalog on disk and want to force a re-parse. Production callers do not need this — the cache key includes the file's st_mtime_ns, so any real edit invalidates the entry on its own.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
def clear_catalog_cache() -> None:
    """Empty the module-level catalog parse cache.

    Useful in tests that rewrite the catalog on disk and want to force a
    re-parse. Production callers do not need this — the cache key includes the
    file's `st_mtime_ns`, so any real edit invalidates the entry on its own.
    """
    _CATALOG_CACHE.clear()

nsi_polygon_body(lat_lim, lon_lim) #

Build the GeoJSON-polygon body for an NSI structures POST.

The NSI ?bbox= query returns empty; the working AOI selector is a POST of a GeoJSON FeatureCollection carrying one rectangle polygon.

Parameters:

Name Type Description Default
lat_lim list[float]

[min_lat, max_lat] in degrees.

required
lon_lim list[float]

[min_lon, max_lon] in degrees.

required

Returns:

Name Type Description
dict dict

A GeoJSON FeatureCollection mapping with a single closed rectangle Polygon, ready to pass as the JSON request body.

Examples:

  • Build a polygon body for a small box:
    >>> from earthlens.nsi import nsi_polygon_body
    >>> body = nsi_polygon_body([29.95, 29.96], [-90.07, -90.06])
    >>> body["type"]
    'FeatureCollection'
    >>> body["features"][0]["geometry"]["type"]
    'Polygon'
    >>> len(body["features"][0]["geometry"]["coordinates"][0])
    5
    
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def nsi_polygon_body(lat_lim: list[float], lon_lim: list[float]) -> dict:
    """Build the GeoJSON-polygon body for an NSI structures POST.

    The NSI `?bbox=` query returns empty; the working AOI selector is a POST of
    a GeoJSON `FeatureCollection` carrying one rectangle polygon.

    Args:
        lat_lim: `[min_lat, max_lat]` in degrees.
        lon_lim: `[min_lon, max_lon]` in degrees.

    Returns:
        dict: A GeoJSON `FeatureCollection` mapping with a single closed
            rectangle `Polygon`, ready to pass as the JSON request body.

    Examples:
        - Build a polygon body for a small box:
            ```python
            >>> from earthlens.nsi import nsi_polygon_body
            >>> body = nsi_polygon_body([29.95, 29.96], [-90.07, -90.06])
            >>> body["type"]
            'FeatureCollection'
            >>> body["features"][0]["geometry"]["type"]
            'Polygon'
            >>> len(body["features"][0]["geometry"]["coordinates"][0])
            5

            ```
    """
    xmin, ymin, xmax, ymax = bbox_from_limits(lat_lim, lon_lim)
    ring = [
        [xmin, ymin],
        [xmax, ymin],
        [xmax, ymax],
        [xmin, ymax],
        [xmin, ymin],
    ]
    return {
        "type": "FeatureCollection",
        "features": [
            {
                "type": "Feature",
                "properties": {},
                "geometry": {"type": "Polygon", "coordinates": [ring]},
            }
        ],
    }

to_feature_collection(geojson) #

Wrap a GeoJSON FeatureCollection mapping into a pyramids collection.

Parameters:

Name Type Description Default
geojson dict

A GeoJSON mapping carrying a features list (WGS84).

required

Returns:

Name Type Description
FeatureCollection FeatureCollection

The features tagged EPSG:4326. An empty/missing features list yields a schema-light empty collection.

Raises:

Type Description
ValueError

If geojson carries no features key at all.

Examples:

  • Wrap a one-feature GeoJSON and inspect the collection:
    >>> from earthlens.nsi import to_feature_collection
    >>> fc = to_feature_collection(
    ...     {
    ...         "type": "FeatureCollection",
    ...         "features": [
    ...             {
    ...                 "type": "Feature",
    ...                 "geometry": {"type": "Point", "coordinates": [-90.0, 29.9]},
    ...                 "properties": {"occtype": "RES1"},
    ...             }
    ...         ],
    ...     }
    ... )
    >>> len(fc)
    1
    
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def to_feature_collection(geojson: dict) -> FeatureCollection:
    """Wrap a GeoJSON `FeatureCollection` mapping into a pyramids collection.

    Args:
        geojson: A GeoJSON mapping carrying a `features` list (WGS84).

    Returns:
        FeatureCollection: The features tagged `EPSG:4326`. An empty/missing
            `features` list yields a schema-light empty collection.

    Raises:
        ValueError: If `geojson` carries no `features` key at all.

    Examples:
        - Wrap a one-feature GeoJSON and inspect the collection:
            ```python
            >>> from earthlens.nsi import to_feature_collection
            >>> fc = to_feature_collection(
            ...     {
            ...         "type": "FeatureCollection",
            ...         "features": [
            ...             {
            ...                 "type": "Feature",
            ...                 "geometry": {"type": "Point", "coordinates": [-90.0, 29.9]},
            ...                 "properties": {"occtype": "RES1"},
            ...             }
            ...         ],
            ...     }
            ... )
            >>> len(fc)
            1

            ```
    """
    if "features" not in geojson:
        raise ValueError(
            "to_feature_collection expects a GeoJSON mapping with a 'features' "
            f"key; got keys {sorted(geojson)}."
        )
    features = geojson["features"]
    if not features:
        # An empty result carries just an (empty) geometry column — no fabricated
        # attribute/id column that a populated result (built from the provider's
        # properties) would not also have.
        empty = gpd.GeoDataFrame(geometry=gpd.GeoSeries([], crs=_WGS84), crs=_WGS84)
        return FeatureCollection(empty)
    gdf = gpd.GeoDataFrame.from_features(features, crs=_WGS84)
    return FeatureCollection(gdf)

earthlens.nsi.backend #

Backend for US object-level flood exposure & loss (earthlens.nsi).

NSI(AbstractDataSource) serves three keyless US-federal REST sources chosen by a source= discriminator — structures (USACE National Structure Inventory), nfhl (FEMA National Flood Hazard Layer), and nfip (FEMA NFIP redacted claims v3) — and returns the shape that source declares.

Two design points carry this backend:

  • Per-instance OUTPUT_KIND (G1). The resolved source's output_kind is copied onto self.OUTPUT_KIND in __init__: structures/nfhl return a pyramids :class:~pyramids.feature.collection.FeatureCollection (vector), nfip returns a :class:pandas.DataFrame (tabular). The :class:earthlens.earthlens.EarthLens facade reads the instance attribute to gate aggregate= and to know the return shape.
  • Bounded requests, required (G3). No unbounded national pull: structures needs a fips= or a [lat_lim, lon_lim] box, nfhl needs the box, and nfip needs at least one of state / county / year / flood_event. NFIP paging logs the total record count so a large pull is visible.

All three are US public-domain, keyless — no auth. aggregate= is rejected (these are records, not gridded rasters), and the parse uses no gridded-array dependency.

NSI #

Bases: AbstractDataSource

US flood exposure / loss backend (per-instance output).

Resolves source= to its catalog row, issues the one bounded request that source needs, and returns a :class:~pyramids.feature.collection.FeatureCollection (structures/nfhl) or a :class:pandas.DataFrame (nfip). US-only: a non-US box returns an empty result, not an error (G4).

Attributes:

Name Type Description
OUTPUT_KIND OutputKind

Set per instance in :meth:__init__ from the resolved source's output_kind. The facade reads it to gate aggregate= and to know the return shape.

Source code in libs/providers/hazards/src/earthlens/nsi/backend.py
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
class NSI(AbstractDataSource):
    """US flood exposure / loss backend (per-instance output).

    Resolves `source=` to its catalog row, issues the one bounded request that
    source needs, and returns a
    :class:`~pyramids.feature.collection.FeatureCollection` (`structures`/`nfhl`)
    or a :class:`pandas.DataFrame` (`nfip`). US-only: a non-US box returns an
    empty result, not an error (`G4`).

    Attributes:
        OUTPUT_KIND: Set **per instance** in :meth:`__init__` from the resolved
            source's `output_kind`. The facade reads it to gate `aggregate=` and
            to know the return shape.
    """

    OUTPUT_KIND: OutputKind = "vector"

    AGGREGATE_REFUSAL_REASON = (
        "NSI/FEMA sources are object-level records (structures, flood zones, "
        "insurance claims), not gridded rasters, so there is no meaningful "
        "gridded reduction. Call download() without aggregate="
    )

    #: These sources are point-in-time inventories / claim records, not a time
    #: series, so a missing `start` / `end` is legal.
    REQUIRES_TIME_WINDOW = False

    def __init__(
        self,
        start: str | None = None,
        end: str | None = None,
        lat_lim: list[float] | None = None,
        lon_lim: list[float] | None = None,
        temporal_resolution: str = "snapshot",
        path: Path | str | None = None,
        fmt: str = "%Y-%m-%d",
        source: str = "structures",
        fips: str | None = None,
        filters: dict[str, str | int] | None = None,
        max_records: int | None = None,
        output_format: OutputFormat = "csv",
        session: requests.Session | None = None,
    ):
        """Initialise an NSI backend instance.

        Args:
            start: Inclusive start of an optional window, parsed with `fmt`;
                `None` allowed (these sources are snapshots, not a time series).
            end: Inclusive end of the optional window; `None` allowed.
            lat_lim: `[min_lat, max_lat]` box — required for `nfhl`, and one of
                the two ways to select `structures` (the other is `fips=`).
            lon_lim: `[min_lon, max_lon]` box; see `lat_lim`.
            temporal_resolution: Recorded as the resolution label only.
            path: Output directory for a written tabular (`nfip`) result.
            fmt: `strptime` format for `start` / `end`.
            source: Which source to query — `"structures"` (default), `"nfhl"`,
                or `"nfip"`.
            fips: A 2/5/11/15-digit FIPS code selecting `structures` by
                state/county/tract/block.
            filters: NFIP attribute filter mapping — recognised keys are
                `state` (two-letter), `county` (5-digit FIPS), `year` (loss
                year), and `flood_event` (named event). At least one is required
                for `source="nfip"`, e.g. `filters={"county": "22071", "year": 2005}`.
            max_records: Optional cap on the number of `nfip` records fetched.
            output_format: On-disk format for the `nfip` table — `"csv"`
                (default) or `"parquet"`.
            session: An existing `requests.Session` to reuse for every request.
                Defaults to a fresh session; injectable so the whole path is
                testable with a fake transport.

        Raises:
            ValueError: If `source` is unknown, `output_format` is
                unrecognised, or the source's required bound is missing (`G3`).
        """
        if output_format not in OUTPUT_FORMATS:
            raise ValueError(
                f"output_format must be one of {list(OUTPUT_FORMATS)}, "
                f"got {output_format!r}."
            )

        self._catalog = Catalog()
        self._source: Source = self._catalog.get(source)
        self._fips = fips
        self._filters = dict(filters) if filters else {}
        self._max_records = max_records
        self._output_format: OutputFormat = output_format
        # The structures endpoint is queried with a `POST` carrying the AOI
        # polygon and answers with GeoJSON — a read with no side effect, so
        # replaying it after a transient `5xx` or a dropped connection cannot
        # double-submit anything. Without this the client would decline to
        # retry it, since a `POST` is otherwise assumed unsafe to repeat.
        #
        # Set on the client rather than per call because there is no per-call
        # override, and it is equivalent here: every other request this client
        # makes is a `GET`, which the idempotency gate never blocks. A future
        # `POST` on this client would inherit the promise, so check it still
        # holds before adding one.
        self._http: HttpClient = HttpClient(session=session, retry_unsafe_methods=True)

        # G1 — the per-instance output shape comes from the resolved source.
        self.OUTPUT_KIND = self._source.output_kind

        self._has_box = lat_lim is not None and lon_lim is not None
        self._lat_lim = lat_lim if lat_lim is not None else _GLOBAL_LAT
        self._lon_lim = lon_lim if lon_lim is not None else _GLOBAL_LON
        self._validate_bound()

        super().__init__(
            start=cast("str", start),
            end=cast("str", end),
            variables=[self._source.id],
            temporal_resolution=temporal_resolution,
            lat_lim=self._lat_lim,
            lon_lim=self._lon_lim,
            fmt=fmt,
            path=path,
        )

    def _validate_bound(self) -> None:
        """Enforce the required spatial/attribute bound for the source (`G3`).

        Also validates the `fips` shape, validates only a box that will actually
        be used (fails fast on a malformed one), and warns when a supplied box is
        ignored (`nfip`, or `structures` alongside `fips`).

        Raises:
            ValueError: If `structures` has neither `fips=` nor a box, `fips` is
                not a 2/5/11/15-digit code, `nfhl` has no box, `nfip` has no
                `state`/`county`/`year`/`flood_event`, or a box that will be used
                is not a valid `[min, max]` pair.
        """
        self._validate_fips_shape()
        provider = self._source.provider
        if provider == "nsi":
            self._validate_structures_bound()
        elif provider == "fema-arcgis":
            self._validate_nfhl_bound()
        elif provider == "openfema":
            self._validate_nfip_bound()

    def _validate_fips_shape(self) -> None:
        """Reject a `fips` that is not a 2/5/11/15-digit code."""
        if self._fips is not None and not (
            self._fips.isdigit() and len(self._fips) in _FIPS_LENGTHS
        ):
            raise ValueError(
                "fips must be a 2/5/11/15-digit code (state/county/tract/block); "
                f"got {self._fips!r}."
            )

    def _validate_structures_bound(self) -> None:
        """`structures` needs a `fips` or a box; a box alongside `fips` is warned."""
        if not (self._fips or self._has_box):
            raise ValueError(
                "source='structures' needs a fips= (2/5/11/15-digit) or a "
                "[lat_lim, lon_lim] box; an unbounded national pull is refused."
            )
        if self._fips and self._has_box:
            logger.warning(
                "NSI structures: both fips= and a [lat_lim, lon_lim] box were "
                "given; using fips= and ignoring the box."
            )
        # Validate the box only when it is the selector (no fips).
        if self._has_box and not self._fips:
            bbox_from_limits(self._lat_lim, self._lon_lim)

    def _validate_nfhl_bound(self) -> None:
        """`nfhl` needs a valid box (the ArcGIS query envelope)."""
        if not self._has_box:
            raise ValueError(
                "source='nfhl' needs a [lat_lim, lon_lim] box (the ArcGIS "
                "query envelope); an unbounded national pull is refused."
            )
        bbox_from_limits(self._lat_lim, self._lon_lim)

    def _validate_nfip_bound(self) -> None:
        """`nfip` needs an attribute filter; a supplied box is warned + ignored."""
        # odata_filter validates the keys (raises on an unknown one) and returns
        # None for an empty mapping — so this both fail-fasts and refuses an
        # unbounded pull.
        if _helpers.odata_filter(self._filters) is None:
            raise ValueError(
                "source='nfip' needs a filters= mapping with at least one of "
                f"{sorted(_helpers._NFIP_FILTER_FIELDS)}; an unbounded national "
                "pull is refused."
            )
        if self._has_box:
            logger.warning(
                "NSI nfip: a [lat_lim, lon_lim] box is ignored; nfip filters "
                "by the filters= mapping (state/county/year/flood_event) only."
            )

    def _create_grid(self, lat_lim: list, lon_lim: list) -> SpatialExtent:
        """Return the request's :class:`SpatialExtent`.

        Args:
            lat_lim: `[min_lat, max_lat]` (global sentinel when unbounded).
            lon_lim: `[min_lon, max_lon]`.

        Returns:
            SpatialExtent: The extent for the resolved box (or the globe).
        """
        return SpatialExtent.from_pairs(lat_lim=lat_lim, lon_lim=lon_lim)

    def _check_input_dates(
        self,
        start: str | None,
        end: str | None,
        temporal_resolution: str,
        fmt: str,
    ) -> TemporalExtent:
        """Parse the optional `[start, end]` window into a :class:`TemporalExtent`.

        Dates are not part of the request (these sources are snapshots), so
        `None` bounds are allowed and yield a `None`-dated extent.

        Args:
            start: Inclusive start date string, or `None`.
            end: Inclusive end date string, or `None`.
            temporal_resolution: Recorded as the resolution label.
            fmt: `strptime` format tried first for a string `start` / `end`.

        Returns:
            TemporalExtent: Frozen model with the parsed (or `None`) endpoints.
        """
        start_dt = to_datetime(start, fmt) if start else None
        end_dt = to_datetime(end, fmt) if end else None
        dates = (
            pd.DatetimeIndex([start_dt, end_dt])
            if start_dt is not None and end_dt is not None
            else pd.DatetimeIndex([])
        )
        return TemporalExtent(
            start_date=start_dt,
            end_date=end_dt,
            resolution=temporal_resolution,
            dates=dates,
        )

    def _search(self) -> list[RemoteProduct]:
        """Pin the one product to fetch (the resolved source).

        Returns:
            list[RemoteProduct]: A single product identified by the source name.
        """
        return [RemoteProduct(id=self._source.id, metadata={})]

    def _fetch(self, products: list[RemoteProduct]) -> list[Any]:
        """Route the product to its source and parse the response.

        Args:
            products: The list from :meth:`_search` (one product).

        Returns:
            list: One element — a `FeatureCollection` (`structures`/`nfhl`) or a
                `DataFrame` (`nfip`).
        """
        return [self._fetch_source()]

    def _fetch_source(self) -> FeatureCollection | pd.DataFrame:
        """Fetch and parse the resolved source.

        Returns:
            A `FeatureCollection` (vector) or a `DataFrame` (tabular).

        Raises:
            requests.HTTPError: If the upstream source returns a non-2xx status.
        """
        provider = self._source.provider
        if provider == "nsi":
            return self._fetch_structures()
        if provider == "fema-arcgis":
            return self._fetch_nfhl()
        return self._fetch_nfip()

    def _fetch_structures(self) -> FeatureCollection:
        """Fetch NSI structures by `fips=` (GET) or a box (POST polygon)."""
        endpoint = self._source.endpoint
        if self._fips:
            if len(self._fips) == 2:
                logger.warning(
                    f"NSI structures fips={self._fips!r} is a whole-state pull "
                    "returned in a single response; this can be very large — "
                    "consider a county/tract/block FIPS or a box."
                )
            geojson = self._http.get_json(endpoint, params={"fips": self._fips})
        else:
            body = nsi_polygon_body(self._lat_lim, self._lon_lim)
            geojson = self._http.post(endpoint, json=body).json()
        return to_feature_collection(geojson)

    def _fetch_nfhl(self) -> FeatureCollection:
        """Fetch FEMA NFHL flood zones for the box via the paged ArcGIS query."""
        url = f"{self._source.endpoint}/{self._source.layer_id}/query"
        params = arcgis_envelope(self._lat_lim, self._lon_lim)
        geojson = _helpers.paginate_arcgis(self._http, url, params)
        return to_feature_collection(geojson)

    def _fetch_nfip(self) -> pd.DataFrame:
        """Fetch NFIP claims for the attribute filter, paged, as a `DataFrame`."""
        source = self._source
        filter_str = _helpers.odata_filter(self._filters)
        # Only the curated columns are kept, so project them server-side with
        # $select rather than pulling all ~84 provider fields per record.
        select = ",".join(source.fields.values()) or None
        records, total = _helpers.paginate_nfip(
            self._http,
            source.endpoint,
            cast("str", source.records_key),
            filter_str=filter_str,
            page_size=source.page_size or 1000,
            max_records=self._max_records,
            select=select,
        )
        logger.info(
            f"NSI nfip: {total} claim record(s) match {filter_str!r}; fetched "
            f"{len(records)}"
            + (f" (capped at {self._max_records})" if self._max_records else "")
        )
        if (
            self._max_records is None
            and total is not None
            and total > _LARGE_PULL_THRESHOLD
        ):
            logger.warning(
                f"NSI nfip: {total} records matched with no max_records= cap; "
                "this is a large pull — pass max_records= or a narrower selector "
                "(add year/county to filters=) to bound it."
            )
        return _helpers.records_to_frame(records, source.fields)

    def download(
        self,
        progress_bar: bool = True,
    ) -> pd.DataFrame | FeatureCollection:
        """Fetch the selected source and return its per-instance shape.

        Args:
            progress_bar: Accepted for signature parity; one logical request is
                issued (paged for `nfip`), so this is a no-op.

        Returns:
            A :class:`~pyramids.feature.collection.FeatureCollection` for
            `structures`/`nfhl`, or a :class:`pandas.DataFrame` (also written to
            `root_dir`) for `nfip`.

        Raises:
            requests.HTTPError: If the upstream source returns a non-2xx status.
        """
        results = self._api()
        result = results[0]
        self._log_citation()
        if self.OUTPUT_KIND == "vector":
            logger.info(
                f"NSI {self._source.id}: returned a FeatureCollection "
                f"({len(result)} feature(s))."
            )
            return result
        out_path = self._write_table(result)
        if len(result):
            logger.info(
                f"NSI {self._source.id}: {len(result)} row(s) written to {out_path}."
            )
        else:
            logger.warning(
                f"NSI {self._source.id}: no rows matched; wrote an empty "
                f"(schema-only) table to {out_path}."
            )
        return result

    def _write_table(self, df: pd.DataFrame) -> Path:
        """Write a tabular result to `root_dir` and return the path.

        Args:
            df: The result frame.

        Returns:
            Path: The written CSV / Parquet file path.

        Raises:
            ImportError: If `output_format="parquet"` but `pyarrow` is missing.
        """
        ext = "parquet" if self._output_format == "parquet" else "csv"
        out_path = self.root_dir / f"nsi_{self._source.id}.{ext}"
        if self._output_format == "parquet":
            try:
                df.to_parquet(out_path, index=False)
            except ImportError as exc:  # pragma: no cover - depends on env
                raise ImportError(
                    "Writing Parquet requires 'pyarrow'. Install it (pip "
                    "install pyarrow) or use output_format='csv'."
                ) from exc
        else:
            df.to_csv(out_path, index=False)
        return out_path

    def _log_citation(self) -> None:
        """Log the resolved source's citation once (info, not a warning)."""
        citation = self._source.citation
        if citation:
            logger.info(f"NSI source citation: {citation}")

__init__(start=None, end=None, lat_lim=None, lon_lim=None, temporal_resolution='snapshot', path=None, fmt='%Y-%m-%d', source='structures', fips=None, filters=None, max_records=None, output_format='csv', session=None) #

Initialise an NSI backend instance.

Parameters:

Name Type Description Default
start str | None

Inclusive start of an optional window, parsed with fmt; None allowed (these sources are snapshots, not a time series).

None
end str | None

Inclusive end of the optional window; None allowed.

None
lat_lim list[float] | None

[min_lat, max_lat] box — required for nfhl, and one of the two ways to select structures (the other is fips=).

None
lon_lim list[float] | None

[min_lon, max_lon] box; see lat_lim.

None
temporal_resolution str

Recorded as the resolution label only.

'snapshot'
path Path | str | None

Output directory for a written tabular (nfip) result.

None
fmt str

strptime format for start / end.

'%Y-%m-%d'
source str

Which source to query — "structures" (default), "nfhl", or "nfip".

'structures'
fips str | None

A 2/5/11/15-digit FIPS code selecting structures by state/county/tract/block.

None
filters dict[str, str | int] | None

NFIP attribute filter mapping — recognised keys are state (two-letter), county (5-digit FIPS), year (loss year), and flood_event (named event). At least one is required for source="nfip", e.g. filters={"county": "22071", "year": 2005}.

None
max_records int | None

Optional cap on the number of nfip records fetched.

None
output_format OutputFormat

On-disk format for the nfip table — "csv" (default) or "parquet".

'csv'
session Session | None

An existing requests.Session to reuse for every request. Defaults to a fresh session; injectable so the whole path is testable with a fake transport.

None

Raises:

Type Description
ValueError

If source is unknown, output_format is unrecognised, or the source's required bound is missing (G3).

Source code in libs/providers/hazards/src/earthlens/nsi/backend.py
def __init__(
    self,
    start: str | None = None,
    end: str | None = None,
    lat_lim: list[float] | None = None,
    lon_lim: list[float] | None = None,
    temporal_resolution: str = "snapshot",
    path: Path | str | None = None,
    fmt: str = "%Y-%m-%d",
    source: str = "structures",
    fips: str | None = None,
    filters: dict[str, str | int] | None = None,
    max_records: int | None = None,
    output_format: OutputFormat = "csv",
    session: requests.Session | None = None,
):
    """Initialise an NSI backend instance.

    Args:
        start: Inclusive start of an optional window, parsed with `fmt`;
            `None` allowed (these sources are snapshots, not a time series).
        end: Inclusive end of the optional window; `None` allowed.
        lat_lim: `[min_lat, max_lat]` box — required for `nfhl`, and one of
            the two ways to select `structures` (the other is `fips=`).
        lon_lim: `[min_lon, max_lon]` box; see `lat_lim`.
        temporal_resolution: Recorded as the resolution label only.
        path: Output directory for a written tabular (`nfip`) result.
        fmt: `strptime` format for `start` / `end`.
        source: Which source to query — `"structures"` (default), `"nfhl"`,
            or `"nfip"`.
        fips: A 2/5/11/15-digit FIPS code selecting `structures` by
            state/county/tract/block.
        filters: NFIP attribute filter mapping — recognised keys are
            `state` (two-letter), `county` (5-digit FIPS), `year` (loss
            year), and `flood_event` (named event). At least one is required
            for `source="nfip"`, e.g. `filters={"county": "22071", "year": 2005}`.
        max_records: Optional cap on the number of `nfip` records fetched.
        output_format: On-disk format for the `nfip` table — `"csv"`
            (default) or `"parquet"`.
        session: An existing `requests.Session` to reuse for every request.
            Defaults to a fresh session; injectable so the whole path is
            testable with a fake transport.

    Raises:
        ValueError: If `source` is unknown, `output_format` is
            unrecognised, or the source's required bound is missing (`G3`).
    """
    if output_format not in OUTPUT_FORMATS:
        raise ValueError(
            f"output_format must be one of {list(OUTPUT_FORMATS)}, "
            f"got {output_format!r}."
        )

    self._catalog = Catalog()
    self._source: Source = self._catalog.get(source)
    self._fips = fips
    self._filters = dict(filters) if filters else {}
    self._max_records = max_records
    self._output_format: OutputFormat = output_format
    # The structures endpoint is queried with a `POST` carrying the AOI
    # polygon and answers with GeoJSON — a read with no side effect, so
    # replaying it after a transient `5xx` or a dropped connection cannot
    # double-submit anything. Without this the client would decline to
    # retry it, since a `POST` is otherwise assumed unsafe to repeat.
    #
    # Set on the client rather than per call because there is no per-call
    # override, and it is equivalent here: every other request this client
    # makes is a `GET`, which the idempotency gate never blocks. A future
    # `POST` on this client would inherit the promise, so check it still
    # holds before adding one.
    self._http: HttpClient = HttpClient(session=session, retry_unsafe_methods=True)

    # G1 — the per-instance output shape comes from the resolved source.
    self.OUTPUT_KIND = self._source.output_kind

    self._has_box = lat_lim is not None and lon_lim is not None
    self._lat_lim = lat_lim if lat_lim is not None else _GLOBAL_LAT
    self._lon_lim = lon_lim if lon_lim is not None else _GLOBAL_LON
    self._validate_bound()

    super().__init__(
        start=cast("str", start),
        end=cast("str", end),
        variables=[self._source.id],
        temporal_resolution=temporal_resolution,
        lat_lim=self._lat_lim,
        lon_lim=self._lon_lim,
        fmt=fmt,
        path=path,
    )

download(progress_bar=True) #

Fetch the selected source and return its per-instance shape.

Parameters:

Name Type Description Default
progress_bar bool

Accepted for signature parity; one logical request is issued (paged for nfip), so this is a no-op.

True

Returns:

Name Type Description
A DataFrame | FeatureCollection

class:~pyramids.feature.collection.FeatureCollection for

DataFrame | FeatureCollection

structures/nfhl, or a :class:pandas.DataFrame (also written to

DataFrame | FeatureCollection

root_dir) for nfip.

Raises:

Type Description
HTTPError

If the upstream source returns a non-2xx status.

Source code in libs/providers/hazards/src/earthlens/nsi/backend.py
def download(
    self,
    progress_bar: bool = True,
) -> pd.DataFrame | FeatureCollection:
    """Fetch the selected source and return its per-instance shape.

    Args:
        progress_bar: Accepted for signature parity; one logical request is
            issued (paged for `nfip`), so this is a no-op.

    Returns:
        A :class:`~pyramids.feature.collection.FeatureCollection` for
        `structures`/`nfhl`, or a :class:`pandas.DataFrame` (also written to
        `root_dir`) for `nfip`.

    Raises:
        requests.HTTPError: If the upstream source returns a non-2xx status.
    """
    results = self._api()
    result = results[0]
    self._log_citation()
    if self.OUTPUT_KIND == "vector":
        logger.info(
            f"NSI {self._source.id}: returned a FeatureCollection "
            f"({len(result)} feature(s))."
        )
        return result
    out_path = self._write_table(result)
    if len(result):
        logger.info(
            f"NSI {self._source.id}: {len(result)} row(s) written to {out_path}."
        )
    else:
        logger.warning(
            f"NSI {self._source.id}: no rows matched; wrote an empty "
            f"(schema-only) table to {out_path}."
        )
    return result

earthlens.nsi.catalog #

Source catalog for the NSI flood-exposure backend.

The nsi backend serves three keyless US-federal REST sources selected by a source= discriminator: structures (USACE National Structure Inventory), nfhl (FEMA National Flood Hazard Layer), and nfip (FEMA NFIP redacted claims v3). This module is the bridge from a source key to its endpoint, the per-instance output kind, and the friendly -> provider field map used to shape the response.

:class:Catalog is a thin :class:earthlens.base.AbstractCatalog subclass that loads the bundled nsi_data_catalog.yaml and exposes each row as a :class:Source. Resolve one source with :meth:Catalog.get (a did-you-mean hint on an unknown key); list the shipped keys with :meth:Catalog.available.

:data:CATALOG_PATH is the path to the bundled YAML.

Catalog #

Bases: AbstractCatalog[Source]

Source catalog for the NSI backend.

Reads the bundled nsi_data_catalog.yaml (shipped as package data) and exposes its sources: block as a map of :class:Source rows keyed by name. Instantiate with no arguments (Catalog()). Resolve one row with :meth:get and list the shipped keys with :meth:available.

Attributes:

Name Type Description
datasets dict[str, Source]

Map from source name to its :class:Source row.

Examples:

  • Resolve a source and read its output kind:
    >>> from earthlens.nsi import Catalog
    >>> cat = Catalog()
    >>> cat.get("structures").output_kind
    'vector'
    >>> cat.get("nfip").output_kind
    'tabular'
    >>> cat.available()
    ['nfhl', 'nfip', 'structures']
    
Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
class Catalog(AbstractCatalog[Source]):
    """Source catalog for the NSI backend.

    Reads the bundled `nsi_data_catalog.yaml` (shipped as package data) and
    exposes its `sources:` block as a map of :class:`Source` rows keyed by
    name. Instantiate with no arguments (`Catalog()`). Resolve one row with
    :meth:`get` and list the shipped keys with :meth:`available`.

    Attributes:
        datasets: Map from source name to its :class:`Source` row.

    Examples:
        - Resolve a source and read its output kind:
            ```python
            >>> from earthlens.nsi import Catalog
            >>> cat = Catalog()
            >>> cat.get("structures").output_kind
            'vector'
            >>> cat.get("nfip").output_kind
            'tabular'
            >>> cat.available()
            ['nfhl', 'nfip', 'structures']

            ```
    """

    _catalog_kind: str = "NSI catalog"
    _entry_noun: str = "sources"

    #: The source rows live in the base :attr:`datasets` field so the inherited
    #: dict surface and :meth:`get_dataset`'s did-you-mean hint work unchanged.
    datasets: dict[str, Source] = Field(default_factory=dict)

    @classmethod
    def _autoload(cls) -> dict[str, Any]:
        """Read the bundled catalog from disk.

        Returns:
            dict[str, Any]: The `datasets` map read from the bundled catalog.
        """
        return {"datasets": Catalog.load().datasets}

    @classmethod
    def load(cls, catalog_path: Path | None = None) -> Catalog:
        """Read the NSI catalog from disk.

        Args:
            catalog_path: Path to the catalog YAML. Defaults to the module-level
                :data:`CATALOG_PATH`.

        Returns:
            A fully-populated :class:`Catalog`.

        Raises:
            ValueError: If the file has no `sources:` block, or a row fails
                validation.
        """
        catalog_path = catalog_path if catalog_path is not None else CATALOG_PATH
        return cls(datasets=dict(_load_catalog_data(catalog_path)))

    def get(self, source_id: str) -> Source:
        """Resolve a source name to its :class:`Source` row.

        Thin wrapper over the inherited :meth:`get_dataset`, which raises a
        `ValueError` with a did-you-mean hint on an unknown key.

        Args:
            source_id: A shipped source name (`"structures"`, `"nfhl"`,
                `"nfip"`).

        Returns:
            Source: The matching catalog row.

        Raises:
            ValueError: If `source_id` is not a known source; the message names
                the catalog kind and, when a close match exists, adds a
                did-you-mean hint.
        """
        return cast("Source", self.get_dataset(source_id))

    def available(self) -> list[str]:
        """Return the sorted list of shipped source names.

        Returns:
            list[str]: Every catalog key, sorted.
        """
        return sorted(self.datasets)

available() #

Return the sorted list of shipped source names.

Returns:

Type Description
list[str]

list[str]: Every catalog key, sorted.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
def available(self) -> list[str]:
    """Return the sorted list of shipped source names.

    Returns:
        list[str]: Every catalog key, sorted.
    """
    return sorted(self.datasets)

get(source_id) #

Resolve a source name to its :class:Source row.

Thin wrapper over the inherited :meth:get_dataset, which raises a ValueError with a did-you-mean hint on an unknown key.

Parameters:

Name Type Description Default
source_id str

A shipped source name ("structures", "nfhl", "nfip").

required

Returns:

Name Type Description
Source Source

The matching catalog row.

Raises:

Type Description
ValueError

If source_id is not a known source; the message names the catalog kind and, when a close match exists, adds a did-you-mean hint.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
def get(self, source_id: str) -> Source:
    """Resolve a source name to its :class:`Source` row.

    Thin wrapper over the inherited :meth:`get_dataset`, which raises a
    `ValueError` with a did-you-mean hint on an unknown key.

    Args:
        source_id: A shipped source name (`"structures"`, `"nfhl"`,
            `"nfip"`).

    Returns:
        Source: The matching catalog row.

    Raises:
        ValueError: If `source_id` is not a known source; the message names
            the catalog kind and, when a close match exists, adds a
            did-you-mean hint.
    """
    return cast("Source", self.get_dataset(source_id))

load(catalog_path=None) classmethod #

Read the NSI catalog from disk.

Parameters:

Name Type Description Default
catalog_path Path | None

Path to the catalog YAML. Defaults to the module-level :data:CATALOG_PATH.

None

Returns:

Type Description
Catalog

A fully-populated :class:Catalog.

Raises:

Type Description
ValueError

If the file has no sources: block, or a row fails validation.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
@classmethod
def load(cls, catalog_path: Path | None = None) -> Catalog:
    """Read the NSI catalog from disk.

    Args:
        catalog_path: Path to the catalog YAML. Defaults to the module-level
            :data:`CATALOG_PATH`.

    Returns:
        A fully-populated :class:`Catalog`.

    Raises:
        ValueError: If the file has no `sources:` block, or a row fails
            validation.
    """
    catalog_path = catalog_path if catalog_path is not None else CATALOG_PATH
    return cls(datasets=dict(_load_catalog_data(catalog_path)))

Source #

Bases: SummarisedLeaf

One NSI catalog row (one of the three flood sources).

The source name is the parent key in :attr:Catalog.datasets and is stored on the row as :attr:id so a resolved :class:Source is self-describing. Which fields matter depends on :attr:provider; a cross-field validator enforces the source-specific ones (nfhl needs a layer, nfip a records key).

Attributes:

Name Type Description
id str

The source name ("structures", "nfhl", "nfip").

provider str

The upstream service — "nsi", "fema-arcgis", or "openfema".

endpoint str

The base REST URL.

output_kind OutputKind

"vector" (a FeatureCollection) or "tabular" (a DataFrame). Copied onto the backend's OUTPUT_KIND per instance.

long_name str

Human-readable label.

citation str

The source's citation string, logged once on use.

license str

The data licence (all three are US public domain).

fields dict[str, str]

Friendly -> provider field-name map. For the tabular nfip source this shapes the output — the DataFrame is subset to these columns, renamed friendly. For the vector structures / nfhl sources it is informational only: those return the full provider column set unchanged (renaming would drop the many other NSI/ArcGIS fields), and this map documents the notable ones.

layer_id int | None

ArcGIS MapServer layer id (nfhl only — S_Fld_Haz_Ar = 28).

layer_name str | None

ArcGIS layer name (nfhl only).

records_key str | None

JSON envelope key holding the record list (nfip only — NfipClaims).

page_size int | None

OData page size for paged fetches (nfip only).

Examples:

  • Build a source row directly:
    >>> from earthlens.nsi import Source
    >>> row = Source(
    ...     id="nfip",
    ...     provider="openfema",
    ...     endpoint="https://www.fema.gov/api/open/v3/NfipClaims",
    ...     output_kind="tabular",
    ...     records_key="NfipClaims",
    ... )
    >>> row.output_kind
    'tabular'
    
Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
class Source(SummarisedLeaf):
    """One NSI catalog row (one of the three flood sources).

    The source name is the parent key in :attr:`Catalog.datasets` and is stored
    on the row as :attr:`id` so a resolved :class:`Source` is self-describing.
    Which fields matter depends on :attr:`provider`; a cross-field validator
    enforces the source-specific ones (`nfhl` needs a layer, `nfip` a records
    key).

    Attributes:
        id: The source name (`"structures"`, `"nfhl"`, `"nfip"`).
        provider: The upstream service — `"nsi"`, `"fema-arcgis"`, or
            `"openfema"`.
        endpoint: The base REST URL.
        output_kind: `"vector"` (a `FeatureCollection`) or `"tabular"` (a
            `DataFrame`). Copied onto the backend's `OUTPUT_KIND` per instance.
        long_name: Human-readable label.
        citation: The source's citation string, logged once on use.
        license: The data licence (all three are US public domain).
        fields: Friendly -> provider field-name map. For the tabular `nfip`
            source this **shapes the output** — the `DataFrame` is subset to
            these columns, renamed friendly. For the vector `structures` / `nfhl`
            sources it is **informational only**: those return the full provider
            column set unchanged (renaming would drop the many other NSI/ArcGIS
            fields), and this map documents the notable ones.
        layer_id: ArcGIS MapServer layer id (`nfhl` only — `S_Fld_Haz_Ar` = 28).
        layer_name: ArcGIS layer name (`nfhl` only).
        records_key: JSON envelope key holding the record list (`nfip` only —
            `NfipClaims`).
        page_size: OData page size for paged fetches (`nfip` only).

    Examples:
        - Build a source row directly:
            ```python
            >>> from earthlens.nsi import Source
            >>> row = Source(
            ...     id="nfip",
            ...     provider="openfema",
            ...     endpoint="https://www.fema.gov/api/open/v3/NfipClaims",
            ...     output_kind="tabular",
            ...     records_key="NfipClaims",
            ... )
            >>> row.output_kind
            'tabular'

            ```
    """

    _summary_fields = (
        "id",
        "long_name",
        "output_kind",
    )

    model_config = ConfigDict(frozen=True, extra="forbid")

    id: str
    provider: str
    endpoint: str
    output_kind: OutputKind
    long_name: str = ""
    citation: str = ""
    license: str = ""
    fields: dict[str, str] = Field(default_factory=dict)

    # nfhl
    layer_id: int | None = None
    layer_name: str | None = None
    # nfip
    records_key: str | None = None
    page_size: int | None = None

    @model_validator(mode="after")
    def _check_provider_fields(self) -> Source:
        """Enforce the per-source required fields.

        Returns:
            The validated row.

        Raises:
            ValueError: If an `nfhl` row omits `layer_id`/`layer_name`, or an
                `nfip` row omits `records_key`, or `output_kind` disagrees with
                the source family.
        """
        if self.provider == "fema-arcgis":
            if self.layer_id is None or not self.layer_name:
                raise ValueError(
                    f"nfhl source {self.id!r} needs layer_id and layer_name"
                )
            if self.output_kind != "vector":
                raise ValueError(
                    f"nfhl source {self.id!r} must be output_kind 'vector'"
                )
        if self.provider == "openfema":
            if not self.records_key:
                raise ValueError(f"nfip source {self.id!r} needs records_key")
            if self.output_kind != "tabular":
                raise ValueError(
                    f"nfip source {self.id!r} must be output_kind 'tabular'"
                )
        if self.provider == "nsi" and self.output_kind != "vector":
            raise ValueError(
                f"structures source {self.id!r} must be output_kind 'vector'"
            )
        return self

clear_catalog_cache() #

Empty the module-level catalog parse cache.

Useful in tests that rewrite the catalog on disk and want to force a re-parse. Production callers do not need this — the cache key includes the file's st_mtime_ns, so any real edit invalidates the entry on its own.

Source code in libs/providers/hazards/src/earthlens/nsi/catalog.py
def clear_catalog_cache() -> None:
    """Empty the module-level catalog parse cache.

    Useful in tests that rewrite the catalog on disk and want to force a
    re-parse. Production callers do not need this — the cache key includes the
    file's `st_mtime_ns`, so any real edit invalidates the entry on its own.
    """
    _CATALOG_CACHE.clear()

earthlens.nsi.geometry #

Geometry helpers for the NSI backend.

Pure functions that turn a [lat_lim, lon_lim] bounding box into the two spatial request shapes the vector sources need — a GeoJSON polygon body for the NSI structures POST, and an ArcGIS envelope query for the FEMA NFHL layer — and that wrap a returned GeoJSON FeatureCollection into a pyramids :class:~pyramids.feature.collection.FeatureCollection. No network, no state.

arcgis_envelope(lat_lim, lon_lim, out_fields='*') #

Build the ArcGIS query parameters for the NFHL flood-zone layer.

Parameters:

Name Type Description Default
lat_lim list[float]

[min_lat, max_lat] in degrees.

required
lon_lim list[float]

[min_lon, max_lon] in degrees.

required
out_fields str

Comma-separated attribute list to return ("*" = all).

'*'

Returns:

Name Type Description
dict dict

The params for a GET against the layer's query endpoint — an esriGeometryEnvelope in WGS84, GeoJSON output, geometry on.

Examples:

  • Build the query params for a box:
    >>> from earthlens.nsi import arcgis_envelope
    >>> params = arcgis_envelope([29.95, 29.96], [-90.07, -90.06])
    >>> params["geometryType"]
    'esriGeometryEnvelope'
    >>> params["f"]
    'geojson'
    >>> params["outFields"]
    '*'
    
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def arcgis_envelope(
    lat_lim: list[float], lon_lim: list[float], out_fields: str = "*"
) -> dict:
    """Build the ArcGIS `query` parameters for the NFHL flood-zone layer.

    Args:
        lat_lim: `[min_lat, max_lat]` in degrees.
        lon_lim: `[min_lon, max_lon]` in degrees.
        out_fields: Comma-separated attribute list to return (`"*"` = all).

    Returns:
        dict: The `params` for a GET against the layer's `query` endpoint —
            an `esriGeometryEnvelope` in WGS84, GeoJSON output, geometry on.

    Examples:
        - Build the query params for a box:
            ```python
            >>> from earthlens.nsi import arcgis_envelope
            >>> params = arcgis_envelope([29.95, 29.96], [-90.07, -90.06])
            >>> params["geometryType"]
            'esriGeometryEnvelope'
            >>> params["f"]
            'geojson'
            >>> params["outFields"]
            '*'

            ```
    """
    xmin, ymin, xmax, ymax = bbox_from_limits(lat_lim, lon_lim)
    envelope = {"xmin": xmin, "ymin": ymin, "xmax": xmax, "ymax": ymax}
    return {
        "geometry": json.dumps(envelope),
        "geometryType": "esriGeometryEnvelope",
        "inSR": "4326",
        "outSR": "4326",
        "spatialRel": "esriSpatialRelIntersects",
        "outFields": out_fields,
        "returnGeometry": "true",
        "f": "geojson",
        "where": "1=1",
    }

bbox_from_limits(lat_lim, lon_lim) #

Return (xmin, ymin, xmax, ymax) from lat_lim / lon_lim.

Parameters:

Name Type Description Default
lat_lim list[float]

[min_lat, max_lat] in degrees.

required
lon_lim list[float]

[min_lon, max_lon] in degrees.

required

Returns:

Type Description
tuple[float, float, float, float]

tuple[float, float, float, float]: The envelope as (min_lon, min_lat, max_lon, max_lat).

Raises:

Type Description
ValueError

If either limit is not a two-element [min, max] with min < max.

Examples:

>>> from earthlens.nsi import bbox_from_limits
>>> bbox_from_limits([29.95, 29.96], [-90.07, -90.06])
(-90.07, 29.95, -90.06, 29.96)
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def bbox_from_limits(
    lat_lim: list[float], lon_lim: list[float]
) -> tuple[float, float, float, float]:
    """Return `(xmin, ymin, xmax, ymax)` from `lat_lim` / `lon_lim`.

    Args:
        lat_lim: `[min_lat, max_lat]` in degrees.
        lon_lim: `[min_lon, max_lon]` in degrees.

    Returns:
        tuple[float, float, float, float]: The envelope as
            `(min_lon, min_lat, max_lon, max_lat)`.

    Raises:
        ValueError: If either limit is not a two-element `[min, max]` with
            `min < max`.

    Examples:
        ```python
        >>> from earthlens.nsi import bbox_from_limits
        >>> bbox_from_limits([29.95, 29.96], [-90.07, -90.06])
        (-90.07, 29.95, -90.06, 29.96)

        ```
    """
    if not (lat_lim and lon_lim and len(lat_lim) == 2 and len(lon_lim) == 2):
        raise ValueError(
            "lat_lim and lon_lim must each be a two-element [min, max]; got "
            f"lat_lim={lat_lim!r}, lon_lim={lon_lim!r}."
        )
    ymin, ymax = float(lat_lim[0]), float(lat_lim[1])
    xmin, xmax = float(lon_lim[0]), float(lon_lim[1])
    # Allow a degenerate (min == max) axis — a point / zero-width AOI — to match
    # the base SpatialExtent; only an inverted bound (min > max) is an error.
    if ymin > ymax or xmin > xmax:
        raise ValueError(
            f"each bound needs min <= max; got lat_lim={lat_lim!r}, lon_lim={lon_lim!r}."
        )
    return xmin, ymin, xmax, ymax

nsi_polygon_body(lat_lim, lon_lim) #

Build the GeoJSON-polygon body for an NSI structures POST.

The NSI ?bbox= query returns empty; the working AOI selector is a POST of a GeoJSON FeatureCollection carrying one rectangle polygon.

Parameters:

Name Type Description Default
lat_lim list[float]

[min_lat, max_lat] in degrees.

required
lon_lim list[float]

[min_lon, max_lon] in degrees.

required

Returns:

Name Type Description
dict dict

A GeoJSON FeatureCollection mapping with a single closed rectangle Polygon, ready to pass as the JSON request body.

Examples:

  • Build a polygon body for a small box:
    >>> from earthlens.nsi import nsi_polygon_body
    >>> body = nsi_polygon_body([29.95, 29.96], [-90.07, -90.06])
    >>> body["type"]
    'FeatureCollection'
    >>> body["features"][0]["geometry"]["type"]
    'Polygon'
    >>> len(body["features"][0]["geometry"]["coordinates"][0])
    5
    
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def nsi_polygon_body(lat_lim: list[float], lon_lim: list[float]) -> dict:
    """Build the GeoJSON-polygon body for an NSI structures POST.

    The NSI `?bbox=` query returns empty; the working AOI selector is a POST of
    a GeoJSON `FeatureCollection` carrying one rectangle polygon.

    Args:
        lat_lim: `[min_lat, max_lat]` in degrees.
        lon_lim: `[min_lon, max_lon]` in degrees.

    Returns:
        dict: A GeoJSON `FeatureCollection` mapping with a single closed
            rectangle `Polygon`, ready to pass as the JSON request body.

    Examples:
        - Build a polygon body for a small box:
            ```python
            >>> from earthlens.nsi import nsi_polygon_body
            >>> body = nsi_polygon_body([29.95, 29.96], [-90.07, -90.06])
            >>> body["type"]
            'FeatureCollection'
            >>> body["features"][0]["geometry"]["type"]
            'Polygon'
            >>> len(body["features"][0]["geometry"]["coordinates"][0])
            5

            ```
    """
    xmin, ymin, xmax, ymax = bbox_from_limits(lat_lim, lon_lim)
    ring = [
        [xmin, ymin],
        [xmax, ymin],
        [xmax, ymax],
        [xmin, ymax],
        [xmin, ymin],
    ]
    return {
        "type": "FeatureCollection",
        "features": [
            {
                "type": "Feature",
                "properties": {},
                "geometry": {"type": "Polygon", "coordinates": [ring]},
            }
        ],
    }

to_feature_collection(geojson) #

Wrap a GeoJSON FeatureCollection mapping into a pyramids collection.

Parameters:

Name Type Description Default
geojson dict

A GeoJSON mapping carrying a features list (WGS84).

required

Returns:

Name Type Description
FeatureCollection FeatureCollection

The features tagged EPSG:4326. An empty/missing features list yields a schema-light empty collection.

Raises:

Type Description
ValueError

If geojson carries no features key at all.

Examples:

  • Wrap a one-feature GeoJSON and inspect the collection:
    >>> from earthlens.nsi import to_feature_collection
    >>> fc = to_feature_collection(
    ...     {
    ...         "type": "FeatureCollection",
    ...         "features": [
    ...             {
    ...                 "type": "Feature",
    ...                 "geometry": {"type": "Point", "coordinates": [-90.0, 29.9]},
    ...                 "properties": {"occtype": "RES1"},
    ...             }
    ...         ],
    ...     }
    ... )
    >>> len(fc)
    1
    
Source code in libs/providers/hazards/src/earthlens/nsi/geometry.py
def to_feature_collection(geojson: dict) -> FeatureCollection:
    """Wrap a GeoJSON `FeatureCollection` mapping into a pyramids collection.

    Args:
        geojson: A GeoJSON mapping carrying a `features` list (WGS84).

    Returns:
        FeatureCollection: The features tagged `EPSG:4326`. An empty/missing
            `features` list yields a schema-light empty collection.

    Raises:
        ValueError: If `geojson` carries no `features` key at all.

    Examples:
        - Wrap a one-feature GeoJSON and inspect the collection:
            ```python
            >>> from earthlens.nsi import to_feature_collection
            >>> fc = to_feature_collection(
            ...     {
            ...         "type": "FeatureCollection",
            ...         "features": [
            ...             {
            ...                 "type": "Feature",
            ...                 "geometry": {"type": "Point", "coordinates": [-90.0, 29.9]},
            ...                 "properties": {"occtype": "RES1"},
            ...             }
            ...         ],
            ...     }
            ... )
            >>> len(fc)
            1

            ```
    """
    if "features" not in geojson:
        raise ValueError(
            "to_feature_collection expects a GeoJSON mapping with a 'features' "
            f"key; got keys {sorted(geojson)}."
        )
    features = geojson["features"]
    if not features:
        # An empty result carries just an (empty) geometry column — no fabricated
        # attribute/id column that a populated result (built from the provider's
        # properties) would not also have.
        empty = gpd.GeoDataFrame(geometry=gpd.GeoSeries([], crs=_WGS84), crs=_WGS84)
        return FeatureCollection(empty)
    gdf = gpd.GeoDataFrame.from_features(features, crs=_WGS84)
    return FeatureCollection(gdf)