Skip to content

Catalog & utility API#

The shared plumbing every provider backend builds on: catalog loading, the strict YAML parser, the provider registry, and the small filesystem helpers. See Base contracts for the rules these implement.

Catalog loading#

All 48 catalog loaders route through load_catalog, which owns the catalog glob, the (path, mtime_ns) cache key, and the cache registry.

earthlens.base.load_catalog(path, cache, parse, *, provider, shard_noun='') #

Return the parsed catalog at path, memoised on the files' mtimes.

The composition every provider loader repeats: resolve the contributing files, build the key, return a live cache hit, else call parse and store the result. parse receives the file list and owns everything provider-specific — the row models, the merge across shards, the duplicate-key checks.

Parameters:

Name Type Description Default
path Path

The catalog directory or single YAML file.

required
cache CatalogParseCache

The module's :class:CatalogParseCache.

required
parse Callable[[list[Path]], T]

Callable taking the contributing files and returning the parsed catalog. Called only on a cache miss.

required
provider str

Provider name for the not-found error.

required
shard_noun str

Optional sharding description for that error.

''

Returns:

Type Description
T

Whatever parse returned, from the cache when the mtimes are unchanged.

Raises:

Type Description
ValueError

If path does not exist (see :func:yaml_files_for).

Examples:

  • The parse runs once, then the cached value is reused:
    >>> import tempfile
    >>> from pathlib import Path
    >>> from earthlens.base.yaml_loader import CatalogParseCache
    >>> from earthlens.base.catalog_source import load_catalog
    >>> one = Path(tempfile.mkdtemp()) / "c.yaml"
    >>> _ = one.write_text("datasets: {}\n")
    >>> cache, calls = CatalogParseCache(), []
    >>> def parse(files):
    ...     calls.append(files)
    ...     return {"rows": len(files)}
    >>> load_catalog(one, cache, parse, provider="Demo")
    {'rows': 1}
    >>> load_catalog(one, cache, parse, provider="Demo")
    {'rows': 1}
    >>> len(calls)
    1
    
Source code in libs/core/src/earthlens/base/catalog_source.py
def load_catalog(
    path: Path,
    cache: CatalogParseCache,
    parse: Callable[[list[Path]], T],
    *,
    provider: str,
    shard_noun: str = "",
) -> T:
    """Return the parsed catalog at `path`, memoised on the files' mtimes.

    The composition every provider loader repeats: resolve the contributing
    files, build the key, return a live cache hit, else call `parse` and store
    the result. `parse` receives the file list and owns everything
    provider-specific — the row models, the merge across shards, the
    duplicate-key checks.

    Args:
        path: The catalog directory or single YAML file.
        cache: The module's :class:`CatalogParseCache`.
        parse: Callable taking the contributing files and returning the parsed
            catalog. Called only on a cache miss.
        provider: Provider name for the not-found error.
        shard_noun: Optional sharding description for that error.

    Returns:
        Whatever `parse` returned, from the cache when the mtimes are unchanged.

    Raises:
        ValueError: If `path` does not exist (see :func:`yaml_files_for`).

    Examples:
        - The parse runs once, then the cached value is reused:
            ```python
            >>> import tempfile
            >>> from pathlib import Path
            >>> from earthlens.base.yaml_loader import CatalogParseCache
            >>> from earthlens.base.catalog_source import load_catalog
            >>> one = Path(tempfile.mkdtemp()) / "c.yaml"
            >>> _ = one.write_text("datasets: {}\\n")
            >>> cache, calls = CatalogParseCache(), []
            >>> def parse(files):
            ...     calls.append(files)
            ...     return {"rows": len(files)}
            >>> load_catalog(one, cache, parse, provider="Demo")
            {'rows': 1}
            >>> load_catalog(one, cache, parse, provider="Demo")
            {'rows': 1}
            >>> len(calls)
            1

            ```
    """
    files = yaml_files_for(path, provider=provider, shard_noun=shard_noun)
    key = catalog_cache_key(path, files)
    cached = cache.get(key)
    if cached is not None:
        return cached  # type: ignore[no-any-return]
    parsed = parse(files)
    cache[key] = parsed
    return parsed

Strict YAML#

The duplicate-key-rejecting loader every catalog parses through — a mapping that declares the same key twice raises ValueError rather than silently keeping the last value.

earthlens.base.yaml_loader.load_yaml_strict(path) #

Parse a YAML file, rejecting duplicate mapping keys.

A thin wrapper over yaml.load(..., Loader=_StrictSafeLoader) so callers (the catalog loaders) never touch the loader class directly.

Parameters:

Name Type Description Default
path str | Path

Filesystem path to the YAML file.

required

Returns:

Type Description
Any

The parsed YAML (typically a dict), or None for an empty

Any

file.

Raises:

Type Description
ValueError

If any mapping in the file declares a key twice.

Examples:

  • Parse a small YAML file and read a value:
    >>> import os, tempfile, textwrap
    >>> p = os.path.join(tempfile.mkdtemp(), "ok.yaml")
    >>> _ = open(p, "w").write(textwrap.dedent('''
    ...     name: demo
    ...     items:
    ...       - a
    ...       - b
    ... '''))
    >>> data = load_yaml_strict(p)
    >>> data["name"]
    'demo'
    >>> data["items"]
    ['a', 'b']
    
  • A duplicate mapping key is rejected at parse time:
    >>> import os, tempfile, textwrap
    >>> p = os.path.join(tempfile.mkdtemp(), "dup.yaml")
    >>> _ = open(p, "w").write(textwrap.dedent('''
    ...     a: 1
    ...     a: 2
    ... '''))
    >>> load_yaml_strict(p)  # doctest: +ELLIPSIS
    Traceback (most recent call last):
        ...
    ValueError: duplicate YAML key 'a' at line 3, ...
    
See Also

earthlens.ecmwf.catalog.Catalog: Uses this to load the CDS catalog. earthlens.gee.catalog.Catalog: Uses this to load the GEE catalog.

Source code in libs/core/src/earthlens/base/yaml_loader.py
def load_yaml_strict(path: str | Path) -> Any:
    """Parse a YAML file, rejecting duplicate mapping keys.

    A thin wrapper over `yaml.load(..., Loader=_StrictSafeLoader)` so
    callers (the catalog loaders) never touch the loader class directly.

    Args:
        path: Filesystem path to the YAML file.

    Returns:
        The parsed YAML (typically a `dict`), or `None` for an empty
        file.

    Raises:
        ValueError: If any mapping in the file declares a key twice.

    Examples:
        - Parse a small YAML file and read a value:
            ```python
            >>> import os, tempfile, textwrap
            >>> p = os.path.join(tempfile.mkdtemp(), "ok.yaml")
            >>> _ = open(p, "w").write(textwrap.dedent('''
            ...     name: demo
            ...     items:
            ...       - a
            ...       - b
            ... '''))
            >>> data = load_yaml_strict(p)
            >>> data["name"]
            'demo'
            >>> data["items"]
            ['a', 'b']

            ```
        - A duplicate mapping key is rejected at parse time:
            ```python
            >>> import os, tempfile, textwrap
            >>> p = os.path.join(tempfile.mkdtemp(), "dup.yaml")
            >>> _ = open(p, "w").write(textwrap.dedent('''
            ...     a: 1
            ...     a: 2
            ... '''))
            >>> load_yaml_strict(p)  # doctest: +ELLIPSIS
            Traceback (most recent call last):
                ...
            ValueError: duplicate YAML key 'a' at line 3, ...

            ```

    See Also:
        earthlens.ecmwf.catalog.Catalog: Uses this to load the CDS catalog.
        earthlens.gee.catalog.Catalog: Uses this to load the GEE catalog.

    """
    with open(path, encoding="utf-8") as stream:
        # `_StrictSafeLoader` subclasses `yaml.SafeLoader` (no arbitrary
        # object instantiation); bandit's B506 flags any `yaml.load`.
        return yaml.load(stream, Loader=_StrictSafeLoader)  # nosec B506

Provider registry#

Backends that populate the base providers field load it from a per-backend providers.yaml.

earthlens.base.Provider #

Bases: BaseModel

One canonical data provider — a slug-id with a display name and parent.

Frozen value object loaded from a backend's providers.yaml. Datasets reference providers by slug via their provider: field; the catalog loader validates that every referenced slug is registered.

Attributes:

Name Type Description
slug str

Stable kebab-case identifier (e.g. "nasa-lp-daac", "copernicus-marine", "ucsb-chc"); injected from the YAML mapping key.

display_name str

Human-readable name to render in docs and UIs.

parent str | None

Slug of the parent provider, or None for top-level organisations. Used to group e.g. all NASA DAACs under the "nasa" umbrella.

Source code in libs/core/src/earthlens/base/providers.py
class Provider(BaseModel):
    """One canonical data provider — a slug-id with a display name and parent.

    Frozen value object loaded from a backend's `providers.yaml`.
    Datasets reference providers by slug via their `provider:` field;
    the catalog loader validates that every referenced slug is
    registered.

    Attributes:
        slug: Stable kebab-case identifier (e.g. `"nasa-lp-daac"`,
            `"copernicus-marine"`, `"ucsb-chc"`); injected from the
            YAML mapping key.
        display_name: Human-readable name to render in docs and UIs.
        parent: Slug of the parent provider, or `None` for top-level
            organisations. Used to group e.g. all NASA DAACs under
            the `"nasa"` umbrella.
    """

    model_config = ConfigDict(frozen=True, extra="forbid")

    slug: str
    display_name: str
    parent: str | None = None

earthlens.base.load_providers(path) #

Parse + cache providers.yaml at path, keyed on (path, mtime_ns).

Parameters:

Name Type Description Default
path Path

Filesystem path of a providers.yaml-shaped file (a top-level providers: map of slug -> {display_name, parent?}).

required

Returns:

Type Description
dict[str, Provider]

slug -> Provider mapping.

Raises:

Type Description
ValueError

If the file is missing, declares a slug whose parent is not itself a registered slug, or fails pydantic validation on any entry.

Source code in libs/core/src/earthlens/base/providers.py
def load_providers(path: Path) -> dict[str, Provider]:
    """Parse + cache `providers.yaml` at `path`, keyed on `(path, mtime_ns)`.

    Args:
        path: Filesystem path of a `providers.yaml`-shaped file (a
            top-level `providers:` map of `slug -> {display_name,
            parent?}`).

    Returns:
        `slug -> Provider` mapping.

    Raises:
        ValueError: If the file is missing, declares a slug whose
            `parent` is not itself a registered slug, or fails
            pydantic validation on any entry.
    """
    resolved = str(path.resolve())
    try:
        mtime_ns = path.stat().st_mtime_ns
    except FileNotFoundError as exc:
        raise ValueError(
            f"providers registry not found at {path}; L2 (provider "
            "normalisation) expects this file alongside the catalog."
        ) from exc
    key = (resolved, mtime_ns)
    cached = _PROVIDERS_CACHE.get(key)
    if cached is not None:
        return cached

    data = load_yaml_strict(path) or {}
    raw = data.get("providers") or {}
    out: dict[str, Provider] = {}
    for slug, body in raw.items():
        try:
            out[slug] = Provider(slug=slug, **dict(body or {}))
        except ValidationError as exc:
            raise ValueError(f"invalid provider {slug!r} in {path}: {exc}") from exc
    for slug, p in out.items():
        if p.parent is not None and p.parent not in out:
            raise ValueError(
                f"provider {slug!r} declares parent={p.parent!r}, "
                f"which is not a known provider slug in {path}"
            )
    _PROVIDERS_CACHE[key] = out
    return out

Filesystem helpers#

earthlens.base.safe_filename(value) #

Sanitise an id into a filesystem-safe file stem.

Replaces every maximal run of characters outside the whitelist (A-Z a-z 0-9 . _ -) with a single _, then strips any leading / trailing _. Dots are kept, so a dataset id like cmems_mod_glo_phy_my_0.083deg_P1D-m is returned unchanged while a path-bearing key like planetary-computer/sentinel-2-l2a flattens to planetary-computer_sentinel-2-l2a.

Parameters:

Name Type Description Default
value str

The raw provider id / key.

required

Returns:

Type Description
str

A filesystem-safe stem: only A-Z a-z 0-9 . _ -, no leading /

str

trailing _.

Examples:

  • Path separators and Windows-illegal characters collapse to _, while dots and hyphens survive:
    >>> from earthlens.base.naming import safe_filename
    >>> safe_filename("a/b\\c:d")
    'a_b_c_d'
    >>> safe_filename('a*b?c"d<e>f|g')
    'a_b_c_d_e_f_g'
    >>> safe_filename("cmems_mod_glo_phy_my_0.083deg_P1D-m")
    'cmems_mod_glo_phy_my_0.083deg_P1D-m'
    
Source code in libs/core/src/earthlens/base/naming.py
def safe_filename(value: str) -> str:
    r"""Sanitise an id into a filesystem-safe file stem.

    Replaces every maximal run of characters outside the whitelist
    (`A-Z a-z 0-9 . _ -`) with a single `_`, then strips any leading /
    trailing `_`. Dots are kept, so a dataset id like
    `cmems_mod_glo_phy_my_0.083deg_P1D-m` is returned unchanged while a
    path-bearing key like `planetary-computer/sentinel-2-l2a` flattens to
    `planetary-computer_sentinel-2-l2a`.

    Args:
        value: The raw provider id / key.

    Returns:
        A filesystem-safe stem: only `A-Z a-z 0-9 . _ -`, no leading /
        trailing `_`.

    Examples:
        - Path separators and Windows-illegal characters collapse to `_`,
          while dots and hyphens survive:
            ```python
            >>> from earthlens.base.naming import safe_filename
            >>> safe_filename("a/b\\c:d")
            'a_b_c_d'
            >>> safe_filename('a*b?c"d<e>f|g')
            'a_b_c_d_e_f_g'
            >>> safe_filename("cmems_mod_glo_phy_my_0.083deg_P1D-m")
            'cmems_mod_glo_phy_my_0.083deg_P1D-m'

            ```
    """
    return _UNSAFE.sub("_", value).strip("_")

Catalog row summaries#

Catalog rows print as one readable line — what the row is called, what units it is in, and how big or how recent it is — instead of pydantic's field-complete dump. A row opts in by inheriting SummarisedLeaf and declaring the fields worth showing:

from pydantic import Field

from earthlens.base import SummarisedLeaf


class Dataset(SummarisedLeaf):
    _summary_fields = ("id", "title", "bands")

    id: str
    title: str | None = None
    bands: dict[str, int] = Field(default_factory=dict)

print(row) then gives Dataset(A/B, A title, 3 bands). None, empty strings and empty collections are skipped so a sparse row stays short; a non-empty collection renders as a count, and a True boolean renders as its own field name rather than a bare True. Only __str__ is defined — __repr__ keeps pydantic's field-complete form, which is the debugging contract.

Over-long values are clipped to MAX_FRAGMENT characters, and the joined summary to MAX_SUMMARY. Where the cut falls depends on what the value reads as: an identifier (S2A_MSIL2A_20230101T…_T31UFT) is cut in the middle so the trailing segment that distinguishes it survives, while prose keeps three quarters of its head and a quarter of its tail, because a title's first words are what identify it. The summary itself is cut on a fragment boundary — it drops whole fragments and appends ... rather than truncating one mid-word.

Declare _summary_fields bare, or as an explicit ClassVar. Annotating it without ClassVar makes pydantic capture it as a private attribute, which is rejected at class creation rather than silently degrading the summary to ClassName().

A row needing a shape the declared fields cannot express — a composed a -> b identity, a value with a unit, or anything derived from a property — overrides summary_parts and calls super(). _summary_fields is read from the class it is declared on, so a subclass that declares its own replaces the parent's list rather than extending it; spell out the inherited names too when both are wanted.

earthlens.base.SummarisedLeaf #

Bases: BaseModel

Catalog row that prints as one readable line instead of a field dump.

A catalog row is normally read for three things — what it is called, what units it is in, and how big or how recent it is. Pydantic's default __str__ answers none of them quickly: it prints every field, including the Nones.

None, empty strings and empty collections are skipped, so a sparse row stays short; a non-empty collection renders as a count (12 bands). Fragments are capped at MAX_FRAGMENT characters and the joined summary at MAX_SUMMARY, both marked with an ellipsis when cut. A subclass needing a different shape overrides summary_parts.

Only __str__ is defined. __repr__ is deliberately left as pydantic's field-complete form: that is the debugging contract, and doctests and log lines can depend on its exact text.

Examples:

  • Declare the fields worth showing, in order, then print a row:
    >>> class Dataset(SummarisedLeaf):
    ...     _summary_fields = ("id", "title", "bands")
    ...     id: str
    ...     title: str | None = None
    ...     bands: dict[str, int] = {}
    >>> print(Dataset(id="A/B", title="A title", bands={"b1": 1}))
    Dataset(A/B, A title, 1 band)
    
  • A sparse row stays short instead of printing None placeholders:
    >>> class Dataset(SummarisedLeaf):
    ...     _summary_fields = ("id", "title")
    ...     id: str
    ...     title: str | None = None
    >>> print(Dataset(id="A/B"))
    Dataset(A/B)
    
  • __repr__ still carries every field, so debugging is unaffected:
    >>> class Dataset(SummarisedLeaf):
    ...     _summary_fields = ("id",)
    ...     id: str
    ...     title: str | None = None
    >>> row = Dataset(id="A/B")
    >>> str(row)
    'Dataset(A/B)'
    >>> repr(row)
    "Dataset(id='A/B', title=None)"
    
See Also

FluxableLeaf: Adds a flux / state marker for variable rows.

Source code in libs/core/src/earthlens/base/leaves.py
class SummarisedLeaf(BaseModel):
    """Catalog row that prints as one readable line instead of a field dump.

    A catalog row is normally read for three things — what it is called,
    what units it is in, and how big or how recent it is. Pydantic's default
    `__str__` answers none of them quickly: it prints every field, including
    the `None`s.

    `None`, empty strings and empty collections are skipped, so a sparse row
    stays short; a non-empty collection renders as a count (`12 bands`).
    Fragments are capped at `MAX_FRAGMENT` characters and the joined summary
    at `MAX_SUMMARY`, both marked with an ellipsis when cut. A subclass
    needing a different shape overrides `summary_parts`.

    Only `__str__` is defined. `__repr__` is deliberately left as pydantic's
    field-complete form: that is the debugging contract, and doctests and log
    lines can depend on its exact text.

    Examples:
        - Declare the fields worth showing, in order, then print a row:
            ```python
            >>> class Dataset(SummarisedLeaf):
            ...     _summary_fields = ("id", "title", "bands")
            ...     id: str
            ...     title: str | None = None
            ...     bands: dict[str, int] = {}
            >>> print(Dataset(id="A/B", title="A title", bands={"b1": 1}))
            Dataset(A/B, A title, 1 band)

            ```
        - A sparse row stays short instead of printing `None` placeholders:
            ```python
            >>> class Dataset(SummarisedLeaf):
            ...     _summary_fields = ("id", "title")
            ...     id: str
            ...     title: str | None = None
            >>> print(Dataset(id="A/B"))
            Dataset(A/B)

            ```
        - `__repr__` still carries every field, so debugging is unaffected:
            ```python
            >>> class Dataset(SummarisedLeaf):
            ...     _summary_fields = ("id",)
            ...     id: str
            ...     title: str | None = None
            >>> row = Dataset(id="A/B")
            >>> str(row)
            'Dataset(A/B)'
            >>> repr(row)
            "Dataset(id='A/B', title=None)"

            ```

    See Also:
        FluxableLeaf: Adds a `flux` / `state` marker for variable rows.
    """

    #: Field names to include in `__str__`, in order. Empty means no summary
    #: beyond the class name, which is the honest answer for a row that has
    #: not declared one. A subclass declaration **replaces** its parent's
    #: rather than extending it — to keep a parent's fragments, override
    #: :meth:`summary_parts` and call `super()`, as `FluxableLeaf` does.
    _summary_fields: ClassVar[tuple[str, ...]] = ()

    @classmethod
    def __pydantic_init_subclass__(cls, **kwargs: Any) -> None:
        """Reject a `_summary_fields` spelling pydantic would silently swallow.

        A leading underscore plus a bare annotation — `_summary_fields:
        tuple[str, ...] = (...)` — makes pydantic treat the declaration as a
        private attribute and drop it from the class namespace. Lookup then
        falls through to this base's empty default and the row prints as
        `ClassName()` with no error, no lint warning and no type error. That
        is the one failure this class cannot detect at render time, so it is
        caught at class-creation time instead.

        Args:
            **kwargs: Class-construction keywords, forwarded to the base.

        Raises:
            TypeError: If `_summary_fields` was captured as a private
                attribute, i.e. annotated without `ClassVar`.
        """
        super().__pydantic_init_subclass__(**kwargs)
        near_miss = [
            name
            for name in (*vars(cls), *cls.__private_attributes__)
            if name.startswith("_summary") and name != "_summary_fields"
        ]
        if near_miss:
            raise TypeError(
                f"{cls.__name__} declares {near_miss[0]!r}; the summary is "
                "read from _summary_fields, so this would be ignored and the "
                f"row would print as '{cls.__name__}()'."
            )
        if "_summary_fields" in cls.__private_attributes__:
            raise TypeError(
                f"{cls.__name__}._summary_fields is annotated without "
                "ClassVar, so pydantic captured it as a private attribute and "
                "the summary would silently degrade to "
                f"'{cls.__name__}()'. Declare it bare "
                '(`_summary_fields = ("id", ...)`) or as '
                "`ClassVar[tuple[str, ...]]`."
            )

    def summary_parts(self) -> list[str]:
        """Return the rendered fragments that make up the one-line summary.

        Override to prepend a composed fragment (an `a -> b` identity, say)
        or to append a derived one; call `super().summary_parts()` so a base
        class's contribution is kept.

        Returns:
            list[str]: One fragment per declared field that had a value.

        Examples:
            - Only the populated fields produce a fragment:
                ```python
                >>> class Dataset(SummarisedLeaf):
                ...     _summary_fields = ("id", "title", "provider")
                ...     id: str
                ...     title: str | None = None
                ...     provider: str | None = None
                >>> Dataset(id="A/B", provider="esa").summary_parts()
                ['A/B', 'esa']

                ```
            - Prepend a composed fragment by calling `super()`:
                ```python
                >>> class Variable(SummarisedLeaf):
                ...     _summary_fields = ("units",)
                ...     name: str
                ...     nc_name: str
                ...     units: str
                ...     def summary_parts(self) -> list[str]:
                ...         pair = f"{self.name} -> {self.nc_name}"
                ...         return [pair, *super().summary_parts()]
                >>> print(Variable(name="2m_temperature", nc_name="t2m", units="K"))
                Variable(2m_temperature -> t2m, K)

                ```
        """
        parts = []
        for field in self._summary_fields:
            rendered = render_fragment(getattr(self, field, None), field)
            if rendered:
                parts.append(rendered)
        return parts

    def __str__(self) -> str:
        """Return `ClassName(fragment, fragment, ...)` for `print(row)`.

        Returns:
            str: The one-line summary. The row's own text is reproduced as
            itself, so a title carrying `—` or `≥` keeps it — the same choice
            :meth:`earthlens.base.AbstractCatalog.__str__` makes, since
            escaping the text into ASCII sequences is what actually makes a
            summary unreadable. A console on a narrow codepage (`cp1252`) can
            therefore still raise `UnicodeEncodeError` on `print(row)`; that
            is a property of the console, not of this summary.

        Examples:
            - The class names itself, then lists its fragments:
                ```python
                >>> class Band(SummarisedLeaf):
                ...     _summary_fields = ("id", "units")
                ...     id: str
                ...     units: str | None = None
                >>> str(Band(id="B1", units="K"))
                'Band(B1, K)'

                ```
            - A row with nothing to show still names itself honestly:
                ```python
                >>> class Band(SummarisedLeaf):
                ...     id: str = "B1"
                >>> str(Band())
                'Band()'

                ```
        """
        parts = self.summary_parts()
        kept: list[str] = []
        used = 0
        for part in parts:
            extra = len(part) + (2 if kept else 0)
            if used + extra > MAX_SUMMARY:
                break
            kept.append(part)
            used += extra
        body = ", ".join(kept)
        if len(kept) < len(parts):
            # Cut on a fragment boundary rather than mid-word: splicing a
            # comma-separated list is what makes a clipped summary unreadable.
            body = f"{body}, ..." if body else "..."
        return f"{type(self).__name__}({body})"

__pydantic_init_subclass__(**kwargs) classmethod #

Reject a _summary_fields spelling pydantic would silently swallow.

A leading underscore plus a bare annotation — _summary_fields: tuple[str, ...] = (...) — makes pydantic treat the declaration as a private attribute and drop it from the class namespace. Lookup then falls through to this base's empty default and the row prints as ClassName() with no error, no lint warning and no type error. That is the one failure this class cannot detect at render time, so it is caught at class-creation time instead.

Parameters:

Name Type Description Default
**kwargs Any

Class-construction keywords, forwarded to the base.

{}

Raises:

Type Description
TypeError

If _summary_fields was captured as a private attribute, i.e. annotated without ClassVar.

Source code in libs/core/src/earthlens/base/leaves.py
@classmethod
def __pydantic_init_subclass__(cls, **kwargs: Any) -> None:
    """Reject a `_summary_fields` spelling pydantic would silently swallow.

    A leading underscore plus a bare annotation — `_summary_fields:
    tuple[str, ...] = (...)` — makes pydantic treat the declaration as a
    private attribute and drop it from the class namespace. Lookup then
    falls through to this base's empty default and the row prints as
    `ClassName()` with no error, no lint warning and no type error. That
    is the one failure this class cannot detect at render time, so it is
    caught at class-creation time instead.

    Args:
        **kwargs: Class-construction keywords, forwarded to the base.

    Raises:
        TypeError: If `_summary_fields` was captured as a private
            attribute, i.e. annotated without `ClassVar`.
    """
    super().__pydantic_init_subclass__(**kwargs)
    near_miss = [
        name
        for name in (*vars(cls), *cls.__private_attributes__)
        if name.startswith("_summary") and name != "_summary_fields"
    ]
    if near_miss:
        raise TypeError(
            f"{cls.__name__} declares {near_miss[0]!r}; the summary is "
            "read from _summary_fields, so this would be ignored and the "
            f"row would print as '{cls.__name__}()'."
        )
    if "_summary_fields" in cls.__private_attributes__:
        raise TypeError(
            f"{cls.__name__}._summary_fields is annotated without "
            "ClassVar, so pydantic captured it as a private attribute and "
            "the summary would silently degrade to "
            f"'{cls.__name__}()'. Declare it bare "
            '(`_summary_fields = ("id", ...)`) or as '
            "`ClassVar[tuple[str, ...]]`."
        )

__str__() #

Return ClassName(fragment, fragment, ...) for print(row).

Returns:

Name Type Description
str str

The one-line summary. The row's own text is reproduced as

str

itself, so a title carrying — or ≥ keeps it — the same choice

str

meth:earthlens.base.AbstractCatalog.__str__ makes, since

str

escaping the text into ASCII sequences is what actually makes a

str

summary unreadable. A console on a narrow codepage (cp1252) can

str

therefore still raise UnicodeEncodeError on print(row); that

str

is a property of the console, not of this summary.

Examples:

  • The class names itself, then lists its fragments:
    >>> class Band(SummarisedLeaf):
    ...     _summary_fields = ("id", "units")
    ...     id: str
    ...     units: str | None = None
    >>> str(Band(id="B1", units="K"))
    'Band(B1, K)'
    
  • A row with nothing to show still names itself honestly:
    >>> class Band(SummarisedLeaf):
    ...     id: str = "B1"
    >>> str(Band())
    'Band()'
    
Source code in libs/core/src/earthlens/base/leaves.py
def __str__(self) -> str:
    """Return `ClassName(fragment, fragment, ...)` for `print(row)`.

    Returns:
        str: The one-line summary. The row's own text is reproduced as
        itself, so a title carrying `—` or `≥` keeps it — the same choice
        :meth:`earthlens.base.AbstractCatalog.__str__` makes, since
        escaping the text into ASCII sequences is what actually makes a
        summary unreadable. A console on a narrow codepage (`cp1252`) can
        therefore still raise `UnicodeEncodeError` on `print(row)`; that
        is a property of the console, not of this summary.

    Examples:
        - The class names itself, then lists its fragments:
            ```python
            >>> class Band(SummarisedLeaf):
            ...     _summary_fields = ("id", "units")
            ...     id: str
            ...     units: str | None = None
            >>> str(Band(id="B1", units="K"))
            'Band(B1, K)'

            ```
        - A row with nothing to show still names itself honestly:
            ```python
            >>> class Band(SummarisedLeaf):
            ...     id: str = "B1"
            >>> str(Band())
            'Band()'

            ```
    """
    parts = self.summary_parts()
    kept: list[str] = []
    used = 0
    for part in parts:
        extra = len(part) + (2 if kept else 0)
        if used + extra > MAX_SUMMARY:
            break
        kept.append(part)
        used += extra
    body = ", ".join(kept)
    if len(kept) < len(parts):
        # Cut on a fragment boundary rather than mid-word: splicing a
        # comma-separated list is what makes a clipped summary unreadable.
        body = f"{body}, ..." if body else "..."
    return f"{type(self).__name__}({body})"

summary_parts() #

Return the rendered fragments that make up the one-line summary.

Override to prepend a composed fragment (an a -> b identity, say) or to append a derived one; call super().summary_parts() so a base class's contribution is kept.

Returns:

Type Description
list[str]

list[str]: One fragment per declared field that had a value.

Examples:

  • Only the populated fields produce a fragment:
    >>> class Dataset(SummarisedLeaf):
    ...     _summary_fields = ("id", "title", "provider")
    ...     id: str
    ...     title: str | None = None
    ...     provider: str | None = None
    >>> Dataset(id="A/B", provider="esa").summary_parts()
    ['A/B', 'esa']
    
  • Prepend a composed fragment by calling super():
    >>> class Variable(SummarisedLeaf):
    ...     _summary_fields = ("units",)
    ...     name: str
    ...     nc_name: str
    ...     units: str
    ...     def summary_parts(self) -> list[str]:
    ...         pair = f"{self.name} -> {self.nc_name}"
    ...         return [pair, *super().summary_parts()]
    >>> print(Variable(name="2m_temperature", nc_name="t2m", units="K"))
    Variable(2m_temperature -> t2m, K)
    
Source code in libs/core/src/earthlens/base/leaves.py
def summary_parts(self) -> list[str]:
    """Return the rendered fragments that make up the one-line summary.

    Override to prepend a composed fragment (an `a -> b` identity, say)
    or to append a derived one; call `super().summary_parts()` so a base
    class's contribution is kept.

    Returns:
        list[str]: One fragment per declared field that had a value.

    Examples:
        - Only the populated fields produce a fragment:
            ```python
            >>> class Dataset(SummarisedLeaf):
            ...     _summary_fields = ("id", "title", "provider")
            ...     id: str
            ...     title: str | None = None
            ...     provider: str | None = None
            >>> Dataset(id="A/B", provider="esa").summary_parts()
            ['A/B', 'esa']

            ```
        - Prepend a composed fragment by calling `super()`:
            ```python
            >>> class Variable(SummarisedLeaf):
            ...     _summary_fields = ("units",)
            ...     name: str
            ...     nc_name: str
            ...     units: str
            ...     def summary_parts(self) -> list[str]:
            ...         pair = f"{self.name} -> {self.nc_name}"
            ...         return [pair, *super().summary_parts()]
            >>> print(Variable(name="2m_temperature", nc_name="t2m", units="K"))
            Variable(2m_temperature -> t2m, K)

            ```
    """
    parts = []
    for field in self._summary_fields:
        rendered = render_fragment(getattr(self, field, None), field)
        if rendered:
            parts.append(rendered)
    return parts

earthlens.base.FluxableLeaf #

Bases: SummarisedLeaf

Catalog row that flags whether its quantity accumulates over time.

Both ECMWF :class:earthlens.ecmwf.Variable and CHIRPS :class:earthlens.chc.Variable carry an identical types field plus is_flux property — flux quantities (precipitation, evapotranspiration, radiation) are accumulated per timestep on the server side, so monthly aggregation has to multiply by the number of days in the month. State / instantaneous values (temperature, pressure) don't need that scaling.

GEE :class:earthlens.gee.Band does NOT inherit from this — its raster bands don't carry flux semantics (cloud-screened optical reflectance, NDVI, etc.).

Attributes:

Name Type Description
types str | None

"flux" for accumulated quantities, None (the default) for state / instantaneous values. Concrete subclasses may narrow to a Literal if they enumerate other values.

Examples:

  • A flux row is marked as such in its summary:
    >>> class Variable(FluxableLeaf):
    ...     _summary_fields = ("units",)
    ...     units: str
    >>> print(Variable(units="mm", types="flux"))
    Variable(mm, flux)
    
  • Anything that is not exactly "flux" reads as state:
    >>> class Variable(FluxableLeaf):
    ...     _summary_fields = ("units",)
    ...     units: str
    >>> row = Variable(units="K")
    >>> row.is_flux
    False
    >>> print(row)
    Variable(K, state)
    
See Also

SummarisedLeaf: The one-line __str__ this builds on.

Source code in libs/core/src/earthlens/base/leaves.py
class FluxableLeaf(SummarisedLeaf):
    """Catalog row that flags whether its quantity accumulates over time.

    Both ECMWF :class:`earthlens.ecmwf.Variable` and CHIRPS
    :class:`earthlens.chc.Variable` carry an identical `types` field
    plus `is_flux` property — flux quantities (precipitation,
    evapotranspiration, radiation) are accumulated per timestep on
    the server side, so monthly aggregation has to multiply by the
    number of days in the month. State / instantaneous values
    (temperature, pressure) don't need that scaling.

    GEE :class:`earthlens.gee.Band` does NOT inherit from this — its
    raster bands don't carry flux semantics (cloud-screened optical
    reflectance, NDVI, etc.).

    Attributes:
        types: `"flux"` for accumulated quantities, `None` (the
            default) for state / instantaneous values. Concrete
            subclasses may narrow to a `Literal` if they enumerate
            other values.

    Examples:
        - A flux row is marked as such in its summary:
            ```python
            >>> class Variable(FluxableLeaf):
            ...     _summary_fields = ("units",)
            ...     units: str
            >>> print(Variable(units="mm", types="flux"))
            Variable(mm, flux)

            ```
        - Anything that is not exactly `"flux"` reads as state:
            ```python
            >>> class Variable(FluxableLeaf):
            ...     _summary_fields = ("units",)
            ...     units: str
            >>> row = Variable(units="K")
            >>> row.is_flux
            False
            >>> print(row)
            Variable(K, state)

            ```

    See Also:
        SummarisedLeaf: The one-line `__str__` this builds on.
    """

    model_config = ConfigDict(frozen=True, extra="forbid")

    types: str | None = None

    @property
    def is_flux(self) -> bool:
        """`True` when `types == "flux"`; drives monthly accumulation scaling.

        Returns:
            bool: Whether the quantity accumulates over the timestep.

        Examples:
            - An accumulated quantity is a flux:
                ```python
                >>> FluxableLeaf(types="flux").is_flux
                True

                ```
            - An instantaneous quantity, and the default, are not:
                ```python
                >>> [FluxableLeaf(types=t).is_flux for t in ("state", None)]
                [False, False]

                ```
        """
        return self.types == "flux"

    def summary_parts(self) -> list[str]:
        """Append the `flux` / `state` marker to the declared fields.

        `is_flux` is a property rather than a model field, so nothing
        field-driven picks it up — it has to be added explicitly, and it is
        the single most load-bearing fact about a variable row.

        Returns:
            list[str]: The declared fragments, then `"flux"` or `"state"`.

        Examples:
            - The marker is appended, so declared fields keep their order:
                ```python
                >>> class Variable(FluxableLeaf):
                ...     _summary_fields = ("units",)
                ...     units: str
                >>> Variable(units="mm", types="flux").summary_parts()
                ['mm', 'flux']

                ```
            - A row with no declared fields still reports its kind:
                ```python
                >>> FluxableLeaf().summary_parts()
                ['state']

                ```
        """
        return [*super().summary_parts(), "flux" if self.is_flux else "state"]

is_flux property #

True when types == "flux"; drives monthly accumulation scaling.

Returns:

Name Type Description
bool bool

Whether the quantity accumulates over the timestep.

Examples:

  • An accumulated quantity is a flux:
    >>> FluxableLeaf(types="flux").is_flux
    True
    
  • An instantaneous quantity, and the default, are not:
    >>> [FluxableLeaf(types=t).is_flux for t in ("state", None)]
    [False, False]
    

summary_parts() #

Append the flux / state marker to the declared fields.

is_flux is a property rather than a model field, so nothing field-driven picks it up — it has to be added explicitly, and it is the single most load-bearing fact about a variable row.

Returns:

Type Description
list[str]

list[str]: The declared fragments, then "flux" or "state".

Examples:

  • The marker is appended, so declared fields keep their order:
    >>> class Variable(FluxableLeaf):
    ...     _summary_fields = ("units",)
    ...     units: str
    >>> Variable(units="mm", types="flux").summary_parts()
    ['mm', 'flux']
    
  • A row with no declared fields still reports its kind:
    >>> FluxableLeaf().summary_parts()
    ['state']
    
Source code in libs/core/src/earthlens/base/leaves.py
def summary_parts(self) -> list[str]:
    """Append the `flux` / `state` marker to the declared fields.

    `is_flux` is a property rather than a model field, so nothing
    field-driven picks it up — it has to be added explicitly, and it is
    the single most load-bearing fact about a variable row.

    Returns:
        list[str]: The declared fragments, then `"flux"` or `"state"`.

    Examples:
        - The marker is appended, so declared fields keep their order:
            ```python
            >>> class Variable(FluxableLeaf):
            ...     _summary_fields = ("units",)
            ...     units: str
            >>> Variable(units="mm", types="flux").summary_parts()
            ['mm', 'flux']

            ```
        - A row with no declared fields still reports its kind:
            ```python
            >>> FluxableLeaf().summary_parts()
            ['state']

            ```
    """
    return [*super().summary_parts(), "flux" if self.is_flux else "state"]

earthlens.base.render_fragment(value, field) #

Render one field value for a one-line summary, or "" to omit it.

Collapses runs of whitespace, clips to MAX_FRAGMENT characters (cutting the middle, so both ends survive), and renders a non-empty collection as a count labelled with the field name. A summary_parts override composing its own fragment should call this so it gets the same treatment as a declared field.

Parameters:

Name Type Description Default
value Any

The attribute's value. None, an empty string and an empty collection all render as ""; 0, 0.0 and False are real values and render as themselves.

required
field str

The attribute's name, used to label a collection's count.

required

Returns:

Name Type Description
str str

The rendered fragment, or "" when the value carries nothing

str

worth showing (None, an empty string, or an empty collection).

Examples:

  • A scalar renders as its stripped text:
    >>> render_fragment("  K  ", "units")
    'K'
    
  • A collection renders as a count labelled with the field:
    >>> render_fragment({"b1": 1, "b2": 2}, "bands")
    '2 bands'
    
  • Nothing worth showing renders as the empty string:
    >>> [render_fragment(v, "units") for v in (None, "", [])]
    ['', '', '']
    
  • Zero is a real value, so it survives:
    >>> render_fragment(0, "count")
    '0'
    
Source code in libs/core/src/earthlens/base/leaves.py
def render_fragment(value: Any, field: str) -> str:
    """Render one field value for a one-line summary, or `""` to omit it.

    Collapses runs of whitespace, clips to `MAX_FRAGMENT` characters (cutting
    the middle, so both ends survive), and renders a non-empty collection as a
    count labelled with the field name. A `summary_parts` override composing
    its own fragment should call this so it gets the same treatment as a
    declared field.

    Args:
        value: The attribute's value. `None`, an empty string and an empty
            collection all render as `""`; `0`, `0.0` and `False` are real
            values and render as themselves.
        field: The attribute's name, used to label a collection's count.

    Returns:
        str: The rendered fragment, or `""` when the value carries nothing
        worth showing (`None`, an empty string, or an empty collection).

    Examples:
        - A scalar renders as its stripped text:
            ```python
            >>> render_fragment("  K  ", "units")
            'K'

            ```
        - A collection renders as a count labelled with the field:
            ```python
            >>> render_fragment({"b1": 1, "b2": 2}, "bands")
            '2 bands'

            ```
        - Nothing worth showing renders as the empty string:
            ```python
            >>> [render_fragment(v, "units") for v in (None, "", [])]
            ['', '', '']

            ```
        - Zero is a real value, so it survives:
            ```python
            >>> render_fragment(0, "count")
            '0'

            ```
    """
    if value is None:
        return ""
    if isinstance(value, bool):
        # A flag reads as its own name when set; when clear it has nothing to
        # say, and `False` beside four other fragments only takes up room.
        return field if value else ""
    if isinstance(value, _SIZED):
        if not value:
            return ""
        count = len(value)
        return f"{count} {_singular(field, count)}"
    # Collapse runs of whitespace so an embedded newline cannot split the
    # summary across lines; no shipped row does this today, but the summary
    # promises to be one line and nothing else enforces it.
    return _clip(" ".join(str(value).split()), MAX_FRAGMENT)

earthlens.base.render_measure(value, unit='m') #

Render a bare numeric measure with its unit, or "" when unknown.

A nominal resolution stored as a plain number renders as a lone figure with nothing saying what it measured. Five row classes needed the same two lines to fix that, so it lives here once.

Parameters:

Name Type Description Default
value float | int | None

The measure, or None / 0 when the row does not carry one.

required
unit str

The unit to append. Defaults to metres.

'm'

Returns:

Name Type Description
str str

"<n> <unit>" with a trailing zero trimmed, or "".

Examples:

  • A whole number drops its trailing zero:
    >>> render_measure(30.0)
    '30 m'
    
  • A fractional value keeps its precision, and the unit is free:
    >>> render_measure(1113.2), render_measure(3.75, "arc-second")
    ('1113.2 m', '3.75 arc-second')
    
  • An absent measure contributes nothing:
    >>> render_measure(None)
    ''
    
Source code in libs/core/src/earthlens/base/leaves.py
def render_measure(value: float | int | None, unit: str = "m") -> str:
    """Render a bare numeric measure with its unit, or `""` when unknown.

    A nominal resolution stored as a plain number renders as a lone figure
    with nothing saying what it measured. Five row classes needed the same
    two lines to fix that, so it lives here once.

    Args:
        value: The measure, or `None` / `0` when the row does not carry one.
        unit: The unit to append. Defaults to metres.

    Returns:
        str: `"<n> <unit>"` with a trailing zero trimmed, or `""`.

    Examples:
        - A whole number drops its trailing zero:
            ```python
            >>> render_measure(30.0)
            '30 m'

            ```
        - A fractional value keeps its precision, and the unit is free:
            ```python
            >>> render_measure(1113.2), render_measure(3.75, "arc-second")
            ('1113.2 m', '3.75 arc-second')

            ```
        - An absent measure contributes nothing:
            ```python
            >>> render_measure(None)
            ''

            ```
    """
    if not value:
        return ""
    return f"{value:g} {unit}"

Catalog keys on rows#

Some rows are addressed only by the key they are filed under — a radar station by its ICAO id, an Argo family by its name — and carry no copy of it. The loader injects the key so the row is self-describing, through one shared helper rather than a per-loader merge. The key is authoritative: a body may repeat the field, but a body declaring a different value is rejected with a ValueError naming the catalog file, the row and both values. The injected field is declared Field(default="", exclude=True), so model_dump() does not repeat the row's own key.

earthlens.base.catalog_source.row_fields_with_key(body, key_field, key, *, noun='row', source=None) #

Return a row body with the catalog key set on it, rejecting a mismatch.

Some rows are addressed by the mapping key alone — a radar station by its ICAO id, an Argo family by its short name — and carry no field holding it, so a summary of the row cannot name it. The loader copies the key on, the way the GEE loader injects a band id.

The key is authoritative. A body may repeat it, but a body that declares a different value would leave the row misnaming how it is addressed, so that is refused rather than silently resolved either way.

Parameters:

Name Type Description Default
body Any

The raw row body from the catalog file, or None for an empty row.

required
key_field str

The field the key is written to.

required
key str

The mapping key the row is filed under.

required
noun str

What to call the row in the error message ("station", "family").

'row'
source Path | None

The catalog file the row came from, named in the error so the offending entry can be found without searching. Omitted when the caller has no path to hand.

None

Returns:

Type Description
dict[str, Any]

dict[str, Any]: The body's fields with key_field set to key.

Raises:

Type Description
ValueError

If body declares key_field with a different value.

Examples:

  • The key is copied onto a body that does not carry it:
    >>> row_fields_with_key({"name": "Norman"}, "code", "KTLX")
    {'name': 'Norman', 'code': 'KTLX'}
    
  • Repeating the key is redundant but accepted:
    >>> row_fields_with_key({"code": "KTLX"}, "code", "KTLX")
    {'code': 'KTLX'}
    
  • Contradicting it is refused:
    >>> row_fields_with_key({"code": "KOUN"}, "code", "KTLX", noun="station")
    Traceback (most recent call last):
        ...
    ValueError: station 'KTLX' declares code='KOUN', which does not match the key it is filed under. Remove the field or rename the entry.
    
  • Naming the source puts the offending file in front of the reader:
    >>> from pathlib import Path
    >>> row_fields_with_key(
    ...     {"name": "Coastal"},
    ...     "name",
    ...     "River",
    ...     noun="flood type",
    ...     source=Path("hanze_data_catalog.yaml"),
    ... )
    Traceback (most recent call last):
        ...
    ValueError: hanze_data_catalog.yaml flood type 'River' declares name='Coastal', which does not match the key it is filed under. Remove the field or rename the entry.
    
Source code in libs/core/src/earthlens/base/catalog_source.py
def row_fields_with_key(
    body: Any,
    key_field: str,
    key: str,
    *,
    noun: str = "row",
    source: Path | None = None,
) -> dict[str, Any]:
    """Return a row body with the catalog key set on it, rejecting a mismatch.

    Some rows are addressed by the mapping key alone — a radar station by its
    ICAO id, an Argo family by its short name — and carry no field holding it,
    so a summary of the row cannot name it. The loader copies the key on, the
    way the GEE loader injects a band id.

    The key is authoritative. A body may repeat it, but a body that declares a
    *different* value would leave the row misnaming how it is addressed, so
    that is refused rather than silently resolved either way.

    Args:
        body: The raw row body from the catalog file, or `None` for an empty
            row.
        key_field: The field the key is written to.
        key: The mapping key the row is filed under.
        noun: What to call the row in the error message (`"station"`,
            `"family"`).
        source: The catalog file the row came from, named in the error so the
            offending entry can be found without searching. Omitted when the
            caller has no path to hand.

    Returns:
        dict[str, Any]: The body's fields with `key_field` set to `key`.

    Raises:
        ValueError: If `body` declares `key_field` with a different value.

    Examples:
        - The key is copied onto a body that does not carry it:
            ```python
            >>> row_fields_with_key({"name": "Norman"}, "code", "KTLX")
            {'name': 'Norman', 'code': 'KTLX'}

            ```
        - Repeating the key is redundant but accepted:
            ```python
            >>> row_fields_with_key({"code": "KTLX"}, "code", "KTLX")
            {'code': 'KTLX'}

            ```
        - Contradicting it is refused:
            ```python
            >>> row_fields_with_key({"code": "KOUN"}, "code", "KTLX", noun="station")
            Traceback (most recent call last):
                ...
            ValueError: station 'KTLX' declares code='KOUN', which does not match the key it is filed under. Remove the field or rename the entry.

            ```
        - Naming the source puts the offending file in front of the reader:
            ```python
            >>> from pathlib import Path
            >>> row_fields_with_key(
            ...     {"name": "Coastal"},
            ...     "name",
            ...     "River",
            ...     noun="flood type",
            ...     source=Path("hanze_data_catalog.yaml"),
            ... )
            Traceback (most recent call last):
                ...
            ValueError: hanze_data_catalog.yaml flood type 'River' declares name='Coastal', which does not match the key it is filed under. Remove the field or rename the entry.

            ```
    """
    fields = dict(body or {})
    declared = fields.get(key_field, key)
    if declared != key:
        where = f"{source} " if source is not None else ""
        raise ValueError(
            f"{where}{noun} {key!r} declares {key_field}={declared!r}, which "
            "does not match the key it is filed under. Remove the field or "
            "rename the entry."
        )
    fields[key_field] = key
    return fields