Cloud I/O#
Tile-window iteration and Cloud-Optimized GeoTIFF (COG) write helpers. Introduced as Phase-4 backfill P29 to give digital-rivers a Dask-style chunked-streaming story without requiring a Dask dependency outright.
Module-level functions#
Cloud-optimised raster I/O and chunked-tile streaming.
Two working halves ship today plus two umbrella stubs for the still-deferred features:
- :func:
tile_windows— chunked-iteration helper that yieldspyramids.dataset.Windowtiles for streaming a continental DEM through any per-tile algorithm without materialising the full raster in memory. - :func:
write_cog— Cloud-Optimised GeoTIFF writer; a thin convenience wrapper that delegates to pyramids'Dataset.to_cog.
Deferred (umbrella raises NotImplementedError with a deferral note):
- :func:
dask_backend— full Dask-graph integration on top oftile_windows. - :func:
cloud_storage— Zarr / S3 / GCS read & write factories.
tile_windows(dataset, tile_rows=1024, tile_cols=1024, overlap=0)
#
Iterate Window tiles over a Dataset.
Yields one pyramids.dataset.Window per tile so callers can stream a
continental DEM through any per-tile algorithm without ever
materialising the full raster in memory. A Window is what
Dataset.read_array(window=...) and Dataset.write_array(window=...)
both take, so a tile can be read and written back with the same object.
Tile size defaults match the COG / Cloud-Optimised GeoTIFF spec (512×512 or 1024×1024 internal tiles).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset
|
A pyramids |
required | |
tile_rows
|
int
|
Tile height in cells. Defaults to 1024. |
1024
|
tile_cols
|
int
|
Tile width in cells. Defaults to 1024. |
1024
|
overlap
|
int
|
Cells of overlap between adjacent tiles. Useful for algorithms that need neighbour context (slopes, flow direction, dilations). Default 0. |
0
|
Yields:
| Type | Description |
|---|---|
|
|
|
|
tiles are clipped to the dataset bounds. |
Note
The window is x-first — column before row — matching GDAL and
read_array. Earlier releases yielded a bare row-first
(row_off, col_off, n_rows, n_cols) tuple while documenting it as
a read_array window, so feeding it straight to read_array
silently read a transposed region. Unpack as
col_off, row_off, cols, rows if you still need the four numbers.
Examples:
-
Iterate a 5x5 dataset in 3x3 tiles with no overlap:
import numpy as np from pyramids.dataset import Dataset, GeoReference from digitalrivers.cloud_io import tile_windows ds = Dataset.from_array( ... np.zeros((5, 5), dtype=np.float32), ... geo_ref=GeoReference( ... top_left_corner=(0, 0), ... cell_size=1.0, ... epsg=4326, ... ), ... ) windows = list(tile_windows(ds, tile_rows=3, tile_cols=3)) [tuple(w) for w in windows][(0, 0, 3, 3), (3, 0, 2, 3), (0, 3, 3, 2), (3, 3, 2, 2)] windows[1] Window(col_off=3, row_off=0, cols=2, rows=3)
-
Each window reads the tile it names:
import numpy as np from pyramids.dataset import Dataset, GeoReference from digitalrivers.cloud_io import tile_windows arr = np.arange(25, dtype=np.float32).reshape(5, 5) ds = Dataset.from_array( ... arr, ... geo_ref=GeoReference( ... top_left_corner=(0, 0), ... cell_size=1.0, ... epsg=4326, ... ), ... ) win = list(tile_windows(ds, tile_rows=3, tile_cols=3))[1] tile = np.asarray(ds.read_array(window=win)) bool((tile == arr[0:3, 3:5]).all()) True
Source code in src/digitalrivers/cloud_io.py
51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | |
dask_backend(*args, **kwargs)
#
Dask / chunked-tile backend for continental DEMs — umbrella stub.
Full Dask-graph integration remains deferred. The chunked-iteration
half ships as :func:tile_windows — callers process continental DEMs
by looping for win in tile_windows(ds): chunk = ds.read_array(window=win)
without loading the full mosaic in memory.
References
Dask documentation: https://docs.dask.org/ rioxarray chunked I/O.
Source code in src/digitalrivers/cloud_io.py
write_cog(dataset, path, compress='deflate')
#
Cloud-Optimised GeoTIFF writer.
Thin convenience wrapper that delegates to pyramids' Dataset.to_cog,
the canonical COG writer. COG is the standard cloud-native format for
raster data: internally tiled, internally overviewed, and indexable by
HTTP range requests — the foundation of every modern STAC-based pipeline.
Reach for dataset.to_cog(...) directly when you need the full option
matrix (overviews, blocksize, tiling scheme, reprojection, etc.).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset
|
Any |
required | |
path
|
str
|
Output |
required |
compress
|
str
|
GDAL compression method — |
'deflate'
|
Returns:
| Type | Description |
|---|---|
str
|
The output path on success. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
DriverNotExistError
|
If the GDAL build lacks the COG driver. |
FileNotFoundError
|
If the parent directory does not exist. |
FailedToSaveError
|
If GDAL's COG |
Examples:
-
Write a 5x5 DEM as a COG:
import numpy as np from pyramids.dataset import Dataset, GeoReference from digitalrivers.cloud_io import write_cog import tempfile, os arr = np.arange(25, dtype=np.float32).reshape(5, 5) ds = Dataset.from_array( ... arr, ... geo_ref=GeoReference( ... top_left_corner=(0, 0), ... cell_size=1.0, ... epsg=4326, ... ), ... ) with tempfile.TemporaryDirectory() as tmpdir: ... out_path = os.path.join(tmpdir, "out.tif") ... result = write_cog(ds, out_path) ... os.path.exists(result) True
Source code in src/digitalrivers/cloud_io.py
cloud_storage(*args, **kwargs)
#
Zarr / S3 / GCS factories — umbrella stub.
The COG write half is shipped under :func:write_cog. Zarr writers
and S3 / GCS read factories remain deferred pending a follow-up PR.
Source code in src/digitalrivers/cloud_io.py
Surface map#
| Function | Purpose |
|---|---|
tile_windows(dataset, tile_rows, tile_cols, overlap=0) |
Generator yielding Window tiles for chunked I/O |
write_cog(dataset, path, compress="deflate") |
Write a pyramids Dataset as a Cloud-Optimized GeoTIFF (overviews + tile layout) |
dask_backend(*args, **kwargs) |
Umbrella stub — raises NotImplementedError with a pointer to tile_windows for current Dask interop |
cloud_storage(*args, **kwargs) |
Umbrella stub — raises NotImplementedError for cloud-storage adapters (s3://, gs://, …) |