Catalog & utility API#
The shared plumbing every provider backend builds on: catalog loading, the strict YAML parser, the provider registry, and the small filesystem helpers. See Base contracts for the rules these implement.
Catalog loading#
All 48 catalog loaders route through load_catalog, which owns the catalog glob, the (path, mtime_ns) cache
key, and the cache registry.
earthlens.base.load_catalog(path, cache, parse, *, provider, shard_noun='')
#
Return the parsed catalog at path, memoised on the files' mtimes.
The composition every provider loader repeats: resolve the contributing
files, build the key, return a live cache hit, else call parse and store
the result. parse receives the file list and owns everything
provider-specific — the row models, the merge across shards, the
duplicate-key checks.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The catalog directory or single YAML file. |
required |
cache
|
CatalogParseCache
|
The module's :class: |
required |
parse
|
Callable[[list[Path]], T]
|
Callable taking the contributing files and returning the parsed catalog. Called only on a cache miss. |
required |
provider
|
str
|
Provider name for the not-found error. |
required |
shard_noun
|
str
|
Optional sharding description for that error. |
''
|
Returns:
| Type | Description |
|---|---|
T
|
Whatever |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
- The parse runs once, then the cached value is reused:
>>> import tempfile >>> from pathlib import Path >>> from earthlens.base.yaml_loader import CatalogParseCache >>> from earthlens.base.catalog_source import load_catalog >>> one = Path(tempfile.mkdtemp()) / "c.yaml" >>> _ = one.write_text("datasets: {}\n") >>> cache, calls = CatalogParseCache(), [] >>> def parse(files): ... calls.append(files) ... return {"rows": len(files)} >>> load_catalog(one, cache, parse, provider="Demo") {'rows': 1} >>> load_catalog(one, cache, parse, provider="Demo") {'rows': 1} >>> len(calls) 1
Source code in libs/core/src/earthlens/base/catalog_source.py
Strict YAML#
The duplicate-key-rejecting loader every catalog parses through — a mapping that declares the same key twice
raises ValueError rather than silently keeping the last value.
earthlens.base.yaml_loader.load_yaml_strict(path)
#
Parse a YAML file, rejecting duplicate mapping keys.
A thin wrapper over yaml.load(..., Loader=_StrictSafeLoader) so
callers (the catalog loaders) never touch the loader class directly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | Path
|
Filesystem path to the YAML file. |
required |
Returns:
| Type | Description |
|---|---|
Any
|
The parsed YAML (typically a |
Any
|
file. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If any mapping in the file declares a key twice. |
Examples:
- Parse a small YAML file and read a value:
- A duplicate mapping key is rejected at parse time:
See Also
earthlens.ecmwf.catalog.Catalog: Uses this to load the CDS catalog. earthlens.gee.catalog.Catalog: Uses this to load the GEE catalog.
Source code in libs/core/src/earthlens/base/yaml_loader.py
Provider registry#
Backends that populate the base providers field load it from a per-backend providers.yaml.
earthlens.base.Provider
#
Bases: BaseModel
One canonical data provider — a slug-id with a display name and parent.
Frozen value object loaded from a backend's providers.yaml.
Datasets reference providers by slug via their provider: field;
the catalog loader validates that every referenced slug is
registered.
Attributes:
| Name | Type | Description |
|---|---|---|
slug |
str
|
Stable kebab-case identifier (e.g. |
display_name |
str
|
Human-readable name to render in docs and UIs. |
parent |
str | None
|
Slug of the parent provider, or |
Source code in libs/core/src/earthlens/base/providers.py
earthlens.base.load_providers(path)
#
Parse + cache providers.yaml at path, keyed on (path, mtime_ns).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
Filesystem path of a |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Provider]
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If the file is missing, declares a slug whose
|
Source code in libs/core/src/earthlens/base/providers.py
Filesystem helpers#
earthlens.base.safe_filename(value)
#
Sanitise an id into a filesystem-safe file stem.
Replaces every maximal run of characters outside the whitelist
(A-Z a-z 0-9 . _ -) with a single _, then strips any leading /
trailing _. Dots are kept, so a dataset id like
cmems_mod_glo_phy_my_0.083deg_P1D-m is returned unchanged while a
path-bearing key like planetary-computer/sentinel-2-l2a flattens to
planetary-computer_sentinel-2-l2a.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str
|
The raw provider id / key. |
required |
Returns:
| Type | Description |
|---|---|
str
|
A filesystem-safe stem: only |
str
|
trailing |
Examples:
- Path separators and Windows-illegal characters collapse to
_, while dots and hyphens survive:
Source code in libs/core/src/earthlens/base/naming.py
Catalog row summaries#
Catalog rows print as one readable line — what the row is called, what units it is in, and how big or how
recent it is — instead of pydantic's field-complete dump. A row opts in by inheriting SummarisedLeaf and
declaring the fields worth showing:
from pydantic import Field
from earthlens.base import SummarisedLeaf
class Dataset(SummarisedLeaf):
_summary_fields = ("id", "title", "bands")
id: str
title: str | None = None
bands: dict[str, int] = Field(default_factory=dict)
print(row) then gives Dataset(A/B, A title, 3 bands). None, empty strings and empty collections are
skipped so a sparse row stays short; a non-empty collection renders as a count, and a True boolean renders as
its own field name rather than a bare True. Only __str__ is defined — __repr__ keeps pydantic's
field-complete form, which is the debugging contract.
Over-long values are clipped to MAX_FRAGMENT characters, and the joined summary to MAX_SUMMARY. Where the cut
falls depends on what the value reads as: an identifier (S2A_MSIL2A_20230101T…_T31UFT) is cut in the middle so
the trailing segment that distinguishes it survives, while prose keeps three quarters of its head and a quarter of
its tail, because a title's first words are what identify it. The summary itself is cut on a fragment boundary —
it drops whole fragments and appends ... rather than truncating one mid-word.
Declare _summary_fields bare, or as an explicit ClassVar. Annotating it without ClassVar makes pydantic
capture it as a private attribute, which is rejected at class creation rather than silently degrading the
summary to ClassName().
A row needing a shape the declared fields cannot express — a composed a -> b identity, a value with a unit,
or anything derived from a property — overrides summary_parts and calls super(). _summary_fields is read
from the class it is declared on, so a subclass that declares its own replaces the parent's list rather than
extending it; spell out the inherited names too when both are wanted.
earthlens.base.SummarisedLeaf
#
Bases: BaseModel
Catalog row that prints as one readable line instead of a field dump.
A catalog row is normally read for three things — what it is called,
what units it is in, and how big or how recent it is. Pydantic's default
__str__ answers none of them quickly: it prints every field, including
the Nones.
None, empty strings and empty collections are skipped, so a sparse row
stays short; a non-empty collection renders as a count (12 bands).
Fragments are capped at MAX_FRAGMENT characters and the joined summary
at MAX_SUMMARY, both marked with an ellipsis when cut. A subclass
needing a different shape overrides summary_parts.
Only __str__ is defined. __repr__ is deliberately left as pydantic's
field-complete form: that is the debugging contract, and doctests and log
lines can depend on its exact text.
Examples:
- Declare the fields worth showing, in order, then print a row:
- A sparse row stays short instead of printing
Noneplaceholders: __repr__still carries every field, so debugging is unaffected:
See Also
FluxableLeaf: Adds a flux / state marker for variable rows.
Source code in libs/core/src/earthlens/base/leaves.py
221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 | |
__pydantic_init_subclass__(**kwargs)
classmethod
#
Reject a _summary_fields spelling pydantic would silently swallow.
A leading underscore plus a bare annotation — _summary_fields:
tuple[str, ...] = (...) — makes pydantic treat the declaration as a
private attribute and drop it from the class namespace. Lookup then
falls through to this base's empty default and the row prints as
ClassName() with no error, no lint warning and no type error. That
is the one failure this class cannot detect at render time, so it is
caught at class-creation time instead.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
**kwargs
|
Any
|
Class-construction keywords, forwarded to the base. |
{}
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
Source code in libs/core/src/earthlens/base/leaves.py
__str__()
#
Return ClassName(fragment, fragment, ...) for print(row).
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The one-line summary. The row's own text is reproduced as |
str
|
itself, so a title carrying |
|
str
|
meth: |
|
str
|
escaping the text into ASCII sequences is what actually makes a |
|
str
|
summary unreadable. A console on a narrow codepage ( |
|
str
|
therefore still raise |
|
str
|
is a property of the console, not of this summary. |
Examples:
- The class names itself, then lists its fragments:
- A row with nothing to show still names itself honestly:
Source code in libs/core/src/earthlens/base/leaves.py
summary_parts()
#
Return the rendered fragments that make up the one-line summary.
Override to prepend a composed fragment (an a -> b identity, say)
or to append a derived one; call super().summary_parts() so a base
class's contribution is kept.
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: One fragment per declared field that had a value. |
Examples:
- Only the populated fields produce a fragment:
- Prepend a composed fragment by calling
super():>>> class Variable(SummarisedLeaf): ... _summary_fields = ("units",) ... name: str ... nc_name: str ... units: str ... def summary_parts(self) -> list[str]: ... pair = f"{self.name} -> {self.nc_name}" ... return [pair, *super().summary_parts()] >>> print(Variable(name="2m_temperature", nc_name="t2m", units="K")) Variable(2m_temperature -> t2m, K)
Source code in libs/core/src/earthlens/base/leaves.py
earthlens.base.FluxableLeaf
#
Bases: SummarisedLeaf
Catalog row that flags whether its quantity accumulates over time.
Both ECMWF :class:earthlens.ecmwf.Variable and CHIRPS
:class:earthlens.chc.Variable carry an identical types field
plus is_flux property — flux quantities (precipitation,
evapotranspiration, radiation) are accumulated per timestep on
the server side, so monthly aggregation has to multiply by the
number of days in the month. State / instantaneous values
(temperature, pressure) don't need that scaling.
GEE :class:earthlens.gee.Band does NOT inherit from this — its
raster bands don't carry flux semantics (cloud-screened optical
reflectance, NDVI, etc.).
Attributes:
| Name | Type | Description |
|---|---|---|
types |
str | None
|
|
Examples:
- A flux row is marked as such in its summary:
- Anything that is not exactly
"flux"reads as state:
See Also
SummarisedLeaf: The one-line __str__ this builds on.
Source code in libs/core/src/earthlens/base/leaves.py
420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 | |
is_flux
property
#
True when types == "flux"; drives monthly accumulation scaling.
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Whether the quantity accumulates over the timestep. |
Examples:
- An accumulated quantity is a flux:
- An instantaneous quantity, and the default, are not:
summary_parts()
#
Append the flux / state marker to the declared fields.
is_flux is a property rather than a model field, so nothing
field-driven picks it up — it has to be added explicitly, and it is
the single most load-bearing fact about a variable row.
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: The declared fragments, then |
Examples:
- The marker is appended, so declared fields keep their order:
- A row with no declared fields still reports its kind:
Source code in libs/core/src/earthlens/base/leaves.py
earthlens.base.render_fragment(value, field)
#
Render one field value for a one-line summary, or "" to omit it.
Collapses runs of whitespace, clips to MAX_FRAGMENT characters (cutting
the middle, so both ends survive), and renders a non-empty collection as a
count labelled with the field name. A summary_parts override composing
its own fragment should call this so it gets the same treatment as a
declared field.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Any
|
The attribute's value. |
required |
field
|
str
|
The attribute's name, used to label a collection's count. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The rendered fragment, or |
str
|
worth showing ( |
Examples:
- A scalar renders as its stripped text:
- A collection renders as a count labelled with the field:
- Nothing worth showing renders as the empty string:
- Zero is a real value, so it survives:
Source code in libs/core/src/earthlens/base/leaves.py
earthlens.base.render_measure(value, unit='m')
#
Render a bare numeric measure with its unit, or "" when unknown.
A nominal resolution stored as a plain number renders as a lone figure with nothing saying what it measured. Five row classes needed the same two lines to fix that, so it lives here once.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
float | int | None
|
The measure, or |
required |
unit
|
str
|
The unit to append. Defaults to metres. |
'm'
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
|
Examples:
- A whole number drops its trailing zero:
- A fractional value keeps its precision, and the unit is free:
- An absent measure contributes nothing:
Source code in libs/core/src/earthlens/base/leaves.py
Catalog keys on rows#
Some rows are addressed only by the key they are filed under — a radar station by its ICAO id, an Argo family by
its name — and carry no copy of it. The loader injects the key so the row is self-describing, through one shared
helper rather than a per-loader merge. The key is authoritative: a body may repeat the field, but a body
declaring a different value is rejected with a ValueError naming the catalog file, the row and both values. The
injected field is declared Field(default="", exclude=True), so model_dump() does not repeat the row's own key.
earthlens.base.catalog_source.row_fields_with_key(body, key_field, key, *, noun='row', source=None)
#
Return a row body with the catalog key set on it, rejecting a mismatch.
Some rows are addressed by the mapping key alone — a radar station by its ICAO id, an Argo family by its short name — and carry no field holding it, so a summary of the row cannot name it. The loader copies the key on, the way the GEE loader injects a band id.
The key is authoritative. A body may repeat it, but a body that declares a different value would leave the row misnaming how it is addressed, so that is refused rather than silently resolved either way.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
body
|
Any
|
The raw row body from the catalog file, or |
required |
key_field
|
str
|
The field the key is written to. |
required |
key
|
str
|
The mapping key the row is filed under. |
required |
noun
|
str
|
What to call the row in the error message ( |
'row'
|
source
|
Path | None
|
The catalog file the row came from, named in the error so the offending entry can be found without searching. Omitted when the caller has no path to hand. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
dict[str, Any]: The body's fields with |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
- The key is copied onto a body that does not carry it:
- Repeating the key is redundant but accepted:
- Contradicting it is refused:
- Naming the source puts the offending file in front of the reader:
>>> from pathlib import Path >>> row_fields_with_key( ... {"name": "Coastal"}, ... "name", ... "River", ... noun="flood type", ... source=Path("hanze_data_catalog.yaml"), ... ) Traceback (most recent call last): ... ValueError: hanze_data_catalog.yaml flood type 'River' declares name='Coastal', which does not match the key it is filed under. Remove the field or rename the entry.
Source code in libs/core/src/earthlens/base/catalog_source.py
286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 | |