Skip to content

Using the Google Earth Engine backend#

This page is the hands-on guide to the earthlens GEE backend — picking a dataset from the catalog, building a download, and the trade-offs between the export modes. For background see the Introduction; for credentials see Registering a project and Service account setup; the rendered API is on the Reference page.

Install: the backend needs the Earth Engine SDK — pip install earthlens[gee] (which adds earthengine-api). The EarthLens facade imports without it; import earthlens.gee requires it.

1. Find a dataset and its bands#

The catalog (per-category src/earthlens/gee/catalog/*.yaml files, loaded and merged by earthlens.gee.Catalog) maps Earth Engine asset ids to their band and aggregation metadata — shaped by Earth Engine's own model (a dataset is an image, image_collection, or table; its addressable units are bands; each band may carry a scale/offset, units, wavelength, value range):

from earthlens.gee import Catalog

cat = Catalog()
"USGS/SRTMGL1_003" in cat.datasets          # True (a curated entry)
"COPERNICUS/S2_SR_HARMONIZED" in cat.available_datasets  # True (in the index)

ds = cat.get_dataset("UCSB-CHG/CHIRPS/DAILY")
ds.ee_type           # 'image_collection'
ds.cadence           # Cadence(interval=1, unit='day')
ds.default_reducer   # 'mean'  (how a temporal composite collapses)
list(ds.bands)       # ['precipitation']
cat.get_band("UCSB-CHG/CHIRPS/DAILY", "precipitation").units   # 'mm/d'

available_datasets is the full index of asset ids Earth Engine publishes (regenerated by earthlens datasets refresh gee --write); datasets is the curated subset the package models in detail. earthlens datasets audit gee --coverage reports which available_datasets entries are ready to be curated, and earthlens datasets curate gee <id> prints a ready-to-paste datasets: stanza for one (add --write to append it to the right per-family file automatically).

2. Download#

from earthlens.gee import GEE

gee = GEE(
    start="2020-06-01",
    end="2020-08-31",
    temporal_resolution="monthly",       # one composite image per month
    variables={"UCSB-CHG/CHIRPS/DAILY": ["precipitation"]},
    lat_lim=[28.0, 32.0],                # [lat_min, lat_max]
    lon_lim=[30.0, 34.0],                # [lon_min, lon_max]
    path="data/gee",
    scale=5566,                          # output pixel size in metres
).authenticate(
    service_account="my-sa@my-project.iam.gserviceaccount.com",
    service_key="/path/to/key.json",     # path, or the JSON content as a string
)
paths = gee.download()
# -> [PosixPath('data/gee/UCSB-CHG_CHIRPS_DAILY_precipitation_20200601.tif'),
#     PosixPath('data/gee/UCSB-CHG_CHIRPS_DAILY_precipitation_20200701.tif'),
#     PosixPath('data/gee/UCSB-CHG_CHIRPS_DAILY_precipitation_20200801.tif')]

The request is {asset_id: [band, ...]} — list every band you want from each dataset (one image carries many; ERA5-Land alone has ~150). download() returns one entry per (dataset, band-set, time-bucket): a Path for export_via="url" (below), or a destination string for the async exports.

Authentication#

service_account + service_key use a Google Cloud service-account key (the recommended, headless-friendly path — see Service account setup). The Cloud project is read from the key file's project_id, or pass project= explicitly. Without a key, pass project=<a registered project> and the backend runs the interactive ee.Authenticate() once. A project that isn't registered for Earth Engine, or that the service account lacks an IAM role on, raises AuthenticationError with a pointer at the fix.

temporal_resolution#

  • "raw" (default) — one image: the whole [start, end] window collapsed with the dataset's default_reducer.
  • "daily" / "monthly" / "yearly" — one image per day / month / year, each its sub-window collapsed with the reducer (mean for rates and continuous fields, median for cloud-screened optical scenes, mosaic for tiled or annual maps). Override per call with reducer="median" etc.

Static image datasets (e.g. USGS/SRTMGL1_003) ignore temporal_resolution — they always yield a single image.

Region#

By default the clip is the lat/lon bbox (ee.Geometry.Rectangle). Pass region=<GeoDataFrame> to clip to an exact polygon set (converted via earthlens.gee.create_feature); the bbox is then used only for the "url" size estimate.

3. Export modes (export_via)#

export_via How Limits Output
"url" (default) Synchronous ee.Image.getDownloadURL → streamed download ≤ 32768 px per axis (≈ (east−west)/(scale/111320)); roughly tens of MB a GeoTIFF in path/
"drive" Async ee.batch.Export.image.toDrive, polled to completion maxPixels (set to 1e13) — no 32768-px cap left in the Google Drive drive_folder (a "drive://…" string is returned)
"gcs" Async ee.batch.Export.image.toCloudStorage, polled to completion as "drive" left in the gcs_bucket (a "gs://…" string is returned); the service account needs roles/storage.objectAdmin on the bucket

If a "url" request would exceed the 32768-px limit, download() raises a ValueError telling you the estimated width×height and to use a coarser scale, a smaller bbox, or export_via="drive". For large AOIs use "drive" / "gcs":

gee = GEE(
    start="2023-01-01", end="2023-12-31", temporal_resolution="monthly",
    variables={"COPERNICUS/S2_SR_HARMONIZED": ["B4", "B8"]},
    lat_lim=[51.0, 53.0], lon_lim=[4.0, 7.0],
    scale=10, export_via="drive", drive_folder="ee_exports",
).authenticate(
    service_account="my-sa@my-project.iam.gserviceaccount.com",
    service_key="/path/to/key.json",
)
locations = gee.download()   # blocks while the batch tasks run; pull the files from Drive

"gcs" writes to Cloud Storage, which incurs normal GCP storage/egress charges; "drive" and "url" do not (see the cost notes in the Introduction).

4. Fetch engine (engine) and Cloud Optimized GeoTIFFs (cog)#

Raw reads of a single materialised asset can skip getDownloadURL entirely and pull pixels through GDAL's Earth Engine driver, via the optional pyramids-eo reader:

pip install "earthlens[eedai]"

earthlens[all] deliberately does not include it: installing the extra flips the default engine="auto" onto the reader, and the two engines sample differently (see below), so it is opt-in.

gee = GEE(
    start="2000-02-11", end="2000-02-12",
    variables={"USGS/SRTMGL1_003": ["elevation"]},
    lat_lim=[29.9, 30.0], lon_lim=[31.2, 31.3],
    path="data/gee", scale=90,
    engine="eedai",   # or "auto" (default) / "ee"
    cog=True,         # write a Cloud Optimized GeoTIFF
)
engine Behaviour
"auto" (default) Use the reader when the request qualifies and [eedai] is installed; otherwise Earth Engine.
"ee" Always getDownloadURL — the historical path.
"eedai" Force the reader; raises if the request does not qualify.

A request qualifies when nothing server-side has to shape the image: no cloud_mask, no filters, and export_via="url". That covers both a single ee_type="image" asset and an ee_type="image_collection", which the reader composites client-side per time bucket with the same reducer Earth Engine would have used. The output CRS may be "EPSG:4326" or any metre-based projected CRS. Everything else stays on Earth Engine, which is the only engine that can run a computation graph or export to Drive / GCS / an asset.

A collection is sized before it is served, because the reader downloads every scene in the bucket and holds them to reduce. It goes back to Earth Engine's server-side reduce when the bucket has too many scenes, when the scene stack would not fit one pass, or when the reducer is mosaic — Earth Engine's is last-wins, while the reader returns the first scene of the stack, so they are not the same composite.

With a collection you can also narrow which scenes are read:

EarthLens(
    data_source="gee",
    variables={"COPERNICUS/S2_SR_HARMONIZED": ["B4"]},
    reducer="median",
    property_filter="CLOUDY_PIXEL_PERCENTAGE < 20",
    ...
)

property_filter is an OGR attribute-filter string on the collection's own scene properties. It is separate from filters because those are Earth Engine closures with no string form. It applies only on this path — a single image, or a request Earth Engine ends up serving, ignores it and says so in a warning. Build it in code, never from untrusted input: it reaches the catalog query without escaping.

One caveat worth knowing: the reader takes each band's nodata from the scene's own dataset and the EEDAI driver declares none, so its statistical reducers fold a scene's fill pixels into the result where Earth Engine would mask them. The two agree wherever the scenes carry no fill over the AOI. There is no way to supply the fill from here — the composite read takes no nodata argument — so pass engine="ee" when you need Earth Engine's masking exactly.

What you gain: no 32768-px synchronous cap and no HTTP/zip round-trip, so auto_split is unnecessary. What to know before switching:

  • The extra does not replace [gee]. The request is still built through earthengine-api, so Earth Engine credentials are still required.
  • The reader fetches at the asset's native resolution (its overviews are unreliable) and downsamples locally, so a wide AOI over a fine asset is a large read. A window too big to hold in memory is streamed to disk one tile at a time and mosaicked, so auto_split is not needed here. Tiling is reserved for reads at or near the asset's own resolution — the case Earth Engine's 32768-px cap would otherwise refuse. A request falls back to Earth Engine when it is much coarser than the asset (Earth Engine aggregates server-side and returns a small raster instead of fetching many native pixels per output pixel), when the whole read would still fetch more than the total-work ceiling, when resample is not "nearest" (upstream forbids that with tiling), when a polygon cutline is set (likewise), or when the catalog does not record the asset's native resolution, which leaves the read unsizeable.
  • Pixels are not byte-identical to the Earth Engine path. Earth Engine reads scale in a geographic CRS as a uniform degree-equivalent, while the EEDAI grid is sized for square metres on the ground, so column counts differ away from the equator; and the reader resamples locally (nearest by default) where Earth Engine aggregates server-side. The AOI, CRS and values agree — the sampling does not.
  • cog=True applies to this path only. A request served by Earth Engine writes a plain GeoTIFF (and logs that it did).

5. Via the EarthLens facade#

Once the GEE backend is registered in the facade you'll also be able to do EarthLens(data_source="gee", variables={...}, ...).download(); until then use earthlens.gee.GEE directly as above. (Tracking: plan task H9.)