Using the Google Earth Engine backend#
This page is the hands-on guide to the earthlens GEE backend — picking
a dataset from the catalog, building a download, and the trade-offs
between the export modes. For background see the
Introduction; for credentials see
Registering a project and
Service account setup; the rendered API is
on the Reference page.
Install: the backend needs the Earth Engine SDK —
pip install earthlens[gee](which addsearthengine-api). TheEarthLensfacade imports without it;import earthlens.geerequires it.
1. Find a dataset and its bands#
The catalog (per-category src/earthlens/gee/catalog/*.yaml files,
loaded and merged by earthlens.gee.Catalog) maps Earth Engine asset
ids to their band and aggregation metadata —
shaped by Earth Engine's own model (a dataset is an image,
image_collection, or table; its addressable units are bands; each
band may carry a scale/offset, units, wavelength, value range):
from earthlens.gee import Catalog
cat = Catalog()
"USGS/SRTMGL1_003" in cat.datasets # True (a curated entry)
"COPERNICUS/S2_SR_HARMONIZED" in cat.available_datasets # True (in the index)
ds = cat.get_dataset("UCSB-CHG/CHIRPS/DAILY")
ds.ee_type # 'image_collection'
ds.cadence # Cadence(interval=1, unit='day')
ds.default_reducer # 'mean' (how a temporal composite collapses)
list(ds.bands) # ['precipitation']
cat.get_band("UCSB-CHG/CHIRPS/DAILY", "precipitation").units # 'mm/d'
available_datasets is the full index of asset ids Earth Engine
publishes (regenerated by earthlens datasets refresh gee --write); datasets
is the curated subset the package models in detail. earthlens datasets audit
gee --coverage reports which available_datasets entries are ready to be
curated, and earthlens datasets curate gee <id> prints a ready-to-paste
datasets: stanza for one (add --write to append it to the right per-family
file automatically).
2. Download#
from earthlens.gee import GEE
gee = GEE(
start="2020-06-01",
end="2020-08-31",
temporal_resolution="monthly", # one composite image per month
variables={"UCSB-CHG/CHIRPS/DAILY": ["precipitation"]},
lat_lim=[28.0, 32.0], # [lat_min, lat_max]
lon_lim=[30.0, 34.0], # [lon_min, lon_max]
path="data/gee",
scale=5566, # output pixel size in metres
).authenticate(
service_account="my-sa@my-project.iam.gserviceaccount.com",
service_key="/path/to/key.json", # path, or the JSON content as a string
)
paths = gee.download()
# -> [PosixPath('data/gee/UCSB-CHG_CHIRPS_DAILY_precipitation_20200601.tif'),
# PosixPath('data/gee/UCSB-CHG_CHIRPS_DAILY_precipitation_20200701.tif'),
# PosixPath('data/gee/UCSB-CHG_CHIRPS_DAILY_precipitation_20200801.tif')]
The request is {asset_id: [band, ...]} — list every band you want from
each dataset (one image carries many; ERA5-Land alone has ~150).
download() returns one entry per (dataset, band-set, time-bucket): a
Path for export_via="url" (below), or a destination string for the
async exports.
Authentication#
service_account + service_key use a Google Cloud service-account key
(the recommended, headless-friendly path — see
Service account setup). The Cloud project is
read from the key file's project_id, or pass project= explicitly.
Without a key, pass project=<a registered project> and the backend
runs the interactive ee.Authenticate() once. A project that isn't
registered for Earth Engine, or that the service account lacks an IAM
role on, raises AuthenticationError with a pointer at the fix.
temporal_resolution#
"raw"(default) — one image: the whole[start, end]window collapsed with the dataset'sdefault_reducer."daily"/"monthly"/"yearly"— one image per day / month / year, each its sub-window collapsed with the reducer (meanfor rates and continuous fields,medianfor cloud-screened optical scenes,mosaicfor tiled or annual maps). Override per call withreducer="median"etc.
Static image datasets (e.g. USGS/SRTMGL1_003) ignore
temporal_resolution — they always yield a single image.
Region#
By default the clip is the lat/lon bbox (ee.Geometry.Rectangle). Pass
region=<GeoDataFrame> to clip to an exact polygon set (converted via
earthlens.gee.create_feature); the bbox is then used only for the
"url" size estimate.
3. Export modes (export_via)#
export_via |
How | Limits | Output |
|---|---|---|---|
"url" (default) |
Synchronous ee.Image.getDownloadURL → streamed download |
≤ 32768 px per axis (≈ (east−west)/(scale/111320)); roughly tens of MB |
a GeoTIFF in path/ |
"drive" |
Async ee.batch.Export.image.toDrive, polled to completion |
maxPixels (set to 1e13) — no 32768-px cap |
left in the Google Drive drive_folder (a "drive://…" string is returned) |
"gcs" |
Async ee.batch.Export.image.toCloudStorage, polled to completion |
as "drive" |
left in the gcs_bucket (a "gs://…" string is returned); the service account needs roles/storage.objectAdmin on the bucket |
If a "url" request would exceed the 32768-px limit, download()
raises a ValueError telling you the estimated width×height and to
use a coarser scale, a smaller bbox, or export_via="drive". For
large AOIs use "drive" / "gcs":
gee = GEE(
start="2023-01-01", end="2023-12-31", temporal_resolution="monthly",
variables={"COPERNICUS/S2_SR_HARMONIZED": ["B4", "B8"]},
lat_lim=[51.0, 53.0], lon_lim=[4.0, 7.0],
scale=10, export_via="drive", drive_folder="ee_exports",
).authenticate(
service_account="my-sa@my-project.iam.gserviceaccount.com",
service_key="/path/to/key.json",
)
locations = gee.download() # blocks while the batch tasks run; pull the files from Drive
"gcs"writes to Cloud Storage, which incurs normal GCP storage/egress charges;"drive"and"url"do not (see the cost notes in the Introduction).
4. Fetch engine (engine) and Cloud Optimized GeoTIFFs (cog)#
Raw reads of a single materialised asset can skip getDownloadURL
entirely and pull pixels through GDAL's Earth Engine driver, via the
optional pyramids-eo reader:
earthlens[all] deliberately does not include it: installing the
extra flips the default engine="auto" onto the reader, and the two
engines sample differently (see below), so it is opt-in.
gee = GEE(
start="2000-02-11", end="2000-02-12",
variables={"USGS/SRTMGL1_003": ["elevation"]},
lat_lim=[29.9, 30.0], lon_lim=[31.2, 31.3],
path="data/gee", scale=90,
engine="eedai", # or "auto" (default) / "ee"
cog=True, # write a Cloud Optimized GeoTIFF
)
engine |
Behaviour |
|---|---|
"auto" (default) |
Use the reader when the request qualifies and [eedai] is installed; otherwise Earth Engine. |
"ee" |
Always getDownloadURL — the historical path. |
"eedai" |
Force the reader; raises if the request does not qualify. |
A request qualifies when nothing server-side has to shape the image: no
cloud_mask, no filters, and export_via="url". That covers both a
single ee_type="image" asset and an ee_type="image_collection", which
the reader composites client-side per time bucket with the same reducer
Earth Engine would have used. The output CRS may be "EPSG:4326" or any
metre-based projected CRS. Everything else stays on Earth Engine, which is
the only engine that can run a computation graph or export to Drive / GCS /
an asset.
A collection is sized before it is served, because the reader downloads
every scene in the bucket and holds them to reduce. It goes back to Earth
Engine's server-side reduce when the bucket has too many scenes, when the
scene stack would not fit one pass, or when the reducer is mosaic — Earth
Engine's is last-wins, while the reader returns the first scene of the
stack, so they are not the same composite.
With a collection you can also narrow which scenes are read:
EarthLens(
data_source="gee",
variables={"COPERNICUS/S2_SR_HARMONIZED": ["B4"]},
reducer="median",
property_filter="CLOUDY_PIXEL_PERCENTAGE < 20",
...
)
property_filter is an OGR attribute-filter string on the collection's own
scene properties. It is separate from filters because those are Earth
Engine closures with no string form. It applies only on this path — a single
image, or a request Earth Engine ends up serving, ignores it and says so in
a warning. Build it in code, never from untrusted input: it reaches the
catalog query without escaping.
One caveat worth knowing: the reader takes each band's nodata from the
scene's own dataset and the EEDAI driver declares none, so its statistical
reducers fold a scene's fill pixels into the result where Earth Engine would
mask them. The two agree wherever the scenes carry no fill over the AOI. There
is no way to supply the fill from here — the composite read takes no nodata
argument — so pass engine="ee" when you need Earth Engine's masking exactly.
What you gain: no 32768-px synchronous cap and no HTTP/zip round-trip,
so auto_split is unnecessary. What to know before switching:
- The extra does not replace
[gee]. The request is still built throughearthengine-api, so Earth Engine credentials are still required. - The reader fetches at the asset's native resolution (its overviews
are unreliable) and downsamples locally, so a wide AOI over a fine
asset is a large read. A window too big to hold in memory is streamed
to disk one tile at a time and mosaicked, so
auto_splitis not needed here. Tiling is reserved for reads at or near the asset's own resolution — the case Earth Engine's 32768-px cap would otherwise refuse. A request falls back to Earth Engine when it is much coarser than the asset (Earth Engine aggregates server-side and returns a small raster instead of fetching many native pixels per output pixel), when the whole read would still fetch more than the total-work ceiling, whenresampleis not"nearest"(upstream forbids that with tiling), when a polygon cutline is set (likewise), or when the catalog does not record the asset's native resolution, which leaves the read unsizeable. - Pixels are not byte-identical to the Earth Engine path. Earth
Engine reads
scalein a geographic CRS as a uniform degree-equivalent, while the EEDAI grid is sized for square metres on the ground, so column counts differ away from the equator; and the reader resamples locally (nearest by default) where Earth Engine aggregates server-side. The AOI, CRS and values agree — the sampling does not. cog=Trueapplies to this path only. A request served by Earth Engine writes a plain GeoTIFF (and logs that it did).
5. Via the EarthLens facade#
Once the GEE backend is registered in the facade you'll also be able to
do EarthLens(data_source="gee", variables={...}, ...).download(); until
then use earthlens.gee.GEE directly as above. (Tracking: plan task
H9.)