Other — CIESIN GPWv4 population density¶
Gridded Population of the World v4 (CIESIN, ~1 km) over the Nile delta, reduced to a single mean over the 2020 entry.
Setup¶
First the imports. pyramids provides Dataset (GeoTIFF/NetCDF reading + plotting); earthlens provides the unified EarthLens entry point and the GEE Catalog.
import os
from pathlib import Path
from pyramids.dataset import Dataset
from pyramids.plot import ColorBar, ColorScaling
from earthlens.core import EarthLens
from earthlens.gee import Catalog, cancel_task
Output directory¶
Downloads land under a per-notebook out/other/ directory (it is .gitignored, so re-running overwrites it).
OUT_DIR = Path('out') / 'other'
OUT_DIR.mkdir(parents=True, exist_ok=True)
print(f'output directory: {OUT_DIR.resolve()}')
Credentials¶
The notebook reads the GEE service-account credentials from the GEE_SERVICE_ACCOUNT / GEE_SERVICE_KEY environment variables. Both must be set before running this cell.
SERVICE_ACCOUNT = os.environ['GEE_SERVICE_ACCOUNT']
SERVICE_KEY = os.environ['GEE_SERVICE_KEY']
Inspect the catalog entry¶
Before downloading anything, look at what the bundled catalog knows about the asset — bands, cadence, license, provider.
cat = Catalog()
ds = cat.get_dataset('CIESIN/GPWv4/population-density')
print(ds)
# The summary clips long text and shows only a band count, so an explorer
# notebook still wants the untruncated title and the fields it omits:
print(f'title (full): {ds.title}')
print(f'ee_type: {ds.ee_type}')
print(f'default_reducer: {ds.default_reducer}')
print(f'license: {ds.license}')
print(f'band ids (first 5): {list(ds.bands)[:5]}')
Download¶
Tiny AOI ([29.0, 32.0] lat, [29.0, 33.0] lon) at 5000.0 m, raw cadence — keeps the synchronous download under EE's 32768-px per-axis cap.
Build the request, then authenticate it as a separate step (the construct -> authenticate split keeps each line easy to read and re-run). export_via defaults to "url", so this is a synchronous getDownloadURL round-trip.
gee = EarthLens(
data_source="gee",
start='2020-01-01',
end='2020-12-31',
dataset='CIESIN/GPWv4/population-density',
variables=['population-density'],
aoi=[29.0, 29.0, 33.0, 32.0],
cadence='raw',
path=OUT_DIR,
scale=5000.0,
reducer='mean',
)
gee.authenticate(service_account=SERVICE_ACCOUNT, service_key=SERVICE_KEY)
download() writes the composite to disk and returns the list of written paths.
paths = gee.download(progress_bar=False)
print(f'wrote {len(paths)} GeoTIFF(s):')
for p in paths:
print(f' {p} ({p.stat().st_size / 1024:.1f} KB)')
Quick preview¶
Load the first written GeoTIFF through pyramids. (pyramids.dataset.Dataset is the project's GeoTIFF/NetCDF
wrapper.)
Two things need saying before this renders. First, the band arrives declaring no nodata: the EEDAI reader
folds the source band's fill through the composite and never stamps it into the written raster, so the fill
would count as a real population density and drag the bottom of the scale to -4e14. Second, population
density is extremely skewed — Cairo is three orders of magnitude denser than the desert around it — so a
linear ramp paints everything except the city centre the same colour. Class breaks fix the second problem the
same way the DEM notebook handles elevation.
# Measured from the written GeoTIFF: the fill GPWv4 delivers through the
# compositing read, which the reader leaves undeclared. It is an empirical
# constant, so the assertion below turns "the sentinel changed" into a loud
# failure instead of a map that silently plots the fill as data.
GPW_FILL = -407649103380480.0
# Density spans four orders of magnitude, so the breaks crowd the low end.
DENSITY_BREAKS = [0, 10, 50, 100, 500, 1000, 5000, 10000, 25000, 75000]
preview = Dataset.read_file(paths[0], read_only=False)
print(f'declared : {preview.no_data_value}')
preview.no_data_value = [GPW_FILL]
print(f'after declare: {preview.no_data_value}')
glyph = preview.plot(
cmap='magma',
color=ColorScaling.boundary(bounds=DENSITY_BREAKS),
colorbar=ColorBar(label='persons / km²'),
title='Population density — the Nile valley and delta',
)
glyph.ax.title.set_fontsize(11)
stats = preview.stats(approx_ok=False)
low, high = stats['min'].iloc[0], stats['max'].iloc[0]
print(f'value range: [{low:.4g}, {high:.4g}] persons/km²')
populated = preview.count_domain_cells()
total_cells = preview.rows * preview.columns
print(f'valid cells: {populated:,} of {total_cells:,}')
# Population density cannot be negative, so a negative minimum here means the
# fill is no longer GPW_FILL and is being counted as real data.
assert low >= 0, f'GPW_FILL no longer matches the raster fill; min is {low}'
# The handle was opened writable to declare the nodata; release it so the
# GDAL lock does not outlive the cell.
preview.close()
Tracking submitted jobs (asynchronous export)¶
The download above uses export_via="url" — a synchronous getDownloadURL round-trip. Nothing was queued, so there's no Earth Engine job to track.
To track an export instead, switch to an asynchronous sink (drive / gcs / asset) and pass wait_for_export=False so .download() returns a TaskInfo at submission time rather than blocking until completion. The cells below submit the same (asset_id, band, AOI, scale) request as an export_via="asset" task into the service account's own asset folder, then walk the four jobs-API calls (list_recent_tasks → wait_for_task_id → ee.data.getAsset → ee.data.deleteAsset) to make the job finish and tidy up. See track-batch-exports.ipynb for a deeper worked example.
Imports and the demo folder¶
Bring in the jobs-API helpers and compute the asset folder we own. GEE._export_via_batch writes the image at <asset_id>/<prefix>, so asset_id is the parent folder, not the final image path.
import ee
from earthlens.gee import list_recent_tasks, wait_for_task_id
# The asset goes into a `Folder` asset that we own. `GEE._export_via_batch`
# writes the actual image at `<asset_id>/<prefix>`, so `asset_id` here is
# the parent FOLDER (not the final image path). Both must be cleaned up.
_proj = ee.data._get_projects_path().removeprefix('projects/')
PARENT = f'projects/{_proj}/assets'
DEMO_FOLDER = f'{PARENT}/earthlens-demo-other'
print(f'demo folder: {DEMO_FOLDER}')
Reset and create the folder¶
Listing the parent says whether a previous run left the folder behind, so the cleanup never has to swallow a "not found" from Earth Engine. Then create the folder EE requires before a child write.
# Listing the parent says whether a previous run left the folder behind, so the
# cleanup below never has to swallow a "not found" from Earth Engine.
listed = ee.data.listAssets({'parent': PARENT})
siblings = [asset['name'] for asset in listed.get('assets', [])]
if DEMO_FOLDER in siblings:
children = ee.data.listAssets({'parent': DEMO_FOLDER})
for child in children.get('assets', []):
ee.data.deleteAsset(child['name'])
print(f'cleared leftover child: {child["name"]}')
ee.data.deleteAsset(DEMO_FOLDER)
print(f'cleared leftover folder: {DEMO_FOLDER}')
# Create the parent folder — EE requires it to exist before a child write.
ee.data.createAsset({'type': 'Folder'}, DEMO_FOLDER)
print(f'created folder: {DEMO_FOLDER}')
Submit¶
Same (asset_id, band, AOI, scale) request as the sync download above, just routed through export_via="asset" + wait_for_export=False. download() returns a TaskInfo per submitted bucket at the moment the task is queued — no blocking.
Build the asynchronous request and authenticate it on its own line, mirroring the sync download above but routed through export_via="asset" + wait_for_export=False.
async_gee = EarthLens(
data_source="gee",
start='2020-01-01',
end='2020-12-31',
dataset='CIESIN/GPWv4/population-density',
variables=['population-density'],
aoi=[29.0, 29.0, 33.0, 32.0],
cadence='raw',
path=OUT_DIR,
scale=5000.0,
reducer='mean',
export_via='asset',
asset_id=DEMO_FOLDER,
wait_for_export=False,
)
async_gee.authenticate(service_account=SERVICE_ACCOUNT, service_key=SERVICE_KEY)
download() returns a TaskInfo per submitted bucket at the moment the task is queued — it does not block on completion.
submitted = async_gee.download(progress_bar=False)
task_info = submitted[0]
print(f'submitted: id={task_info.id} state={task_info.state}')
print(f' description={task_info.description}')
List + wait¶
list_recent_tasks(description_prefix=...) returns every matching task across the current project; wait_for_task_id blocks until the one we care about reaches a terminal state. A real workflow would just poll later from a separate process — the wait here exists so the notebook shows the full success path end-to-end.
recent = list_recent_tasks(
description_prefix=task_info.description,
max_age_min=10,
)
print(f'list_recent_tasks matched {len(recent)} task(s):')
for t in recent:
print(f' {t.id} {t.state:<12} {t.description}')
final = None
try:
final = wait_for_task_id(
task_info.id,
poll_seconds=10,
progress_bar=False,
)
print(f'\nfinal state: {final.state}')
finally:
if final is None:
# The wait raises on FAILED / CANCELLED *and on timeout* — and a
# timeout leaves the export still running. Cancel it so an aborted
# notebook does not leave a live task behind; cancel_task is a no-op
# on an already-terminal task, and the original error still
# propagates out of this finally.
cancel_task(task_info.id)
print(f'cancelled {task_info.id} after the wait failed')
Verify + clean up¶
Confirm the produced asset exists on Earth Engine, then delete it (and the surrounding demo folder) so we don't leak storage between notebook runs. The backend wrote the image at <DEMO_FOLDER>/<task description>.
produced = f'{DEMO_FOLDER}/{task_info.description}'
meta = ee.data.getAsset(produced)
print(f'asset exists: type={meta.get("type")} name={meta.get("name")}')
ee.data.deleteAsset(produced)
print('asset deleted')
# Tear down the parent folder.
ee.data.deleteAsset(DEMO_FOLDER)
print(f'folder deleted: {DEMO_FOLDER}')
What's on disk¶
The written GeoTIFF is left under the per-notebook out/ directory for you to inspect. That directory is
.gitignored — re-running the notebook overwrites it.
for p in sorted(OUT_DIR.iterdir()) if OUT_DIR.exists() else []:
print(f'{p} ({p.stat().st_size / 1024:.1f} KB)')