OpenStreetMap features — usage#
This page walks through fetching OSM features with the osm backend. For
background and the named-query list see Introduction; the
rendered API is the Reference page.
Install#
The backend is installed with one extra (imported lazily — the base package imports without it):
pip install earthlens[osm] # all three download types
pip install pyrosm # opt-in: the richer in-memory read engine
[osm] is part of [all] and covers all three download types — the two
live types plus the default streaming read engine. The richer in-memory read
engine is opt-in: pip install pyrosm builds a source dependency (needs a C
compiler), so install it only when you want engine="pyrosm". There are no
credentials to configure — every download type is public and keyless.
Quickstart — current-state hospitals (live)#
from earthlens.core import EarthLens
hospitals = EarthLens(
data_source="osm",
variables=["live:hospitals"], # a named query — see below
lat_lim=[49.40, 49.42], # a small bbox (degrees)
lon_lim=[8.67, 8.71],
path="./out",
).download()
print(len(hospitals), "features")
print(hospitals[["osm_id", "osm_type", "geometry"]].head())
download() returns a FeatureCollection (a geopandas.GeoDataFrame subclass),
so every pandas / geopandas method works on it directly. It also writes
out/osm_live-hospitals.geojson. Keep the bbox small — the live types run
against shared public infrastructure.
Quickstart — building history at a snapshot (history)#
buildings = EarthLens(
data_source="osm",
variables=["history:buildings"],
lat_lim=[49.40, 49.42],
lon_lim=[8.67, 8.71],
start="2020-01-01", # history needs a time (see below)
path="./out",
).download()
print(buildings["@snapshotTimestamp"].iloc[0]) # the history timestamp
Quickstart — every building in a region (bulk)#
For a bulk ask — every building in a country — use the bulk type. It
downloads a regional extract for region= (cached on disk), reads the layer with
the default streaming engine, and clips to the request bbox:
buildings = EarthLens(
data_source="osm",
variables=["bulk:buildings"],
region="malta", # a region key (or a raw "europe/andorra" path)
lat_lim=[35.88, 35.94], # bbox clips the read; omit for the whole extract
lon_lim=[14.48, 14.54],
path="./out",
).download()
print(len(buildings), "building footprints")
The first call downloads the extract (Malta is ~8.8 MB) to a cross-run cache
(osm_pbf/ under the shared earthlens cache directory by default — see
Configuration — override with cache_dir=); repeat
calls reuse it. List the region keys with Catalog().region_ids(), or pass a
raw "continent/region" path (any string with a /). Omit lat_lim /
lon_lim to read the whole extract — the bbox-area cap does not apply to a
bulk read.
Engines — engine="pyosmium" (default) vs engine="pyrosm"
engine="pyosmium" (the default) streams the extract with bounded memory and
is installed with the backend, so it works out of the box and handles
continent- or planet-scale extracts. It is the coarser reader: it
returns a slimmer osm_id / osm_type / geometry schema and, per layer, a
single geometry kind under one representative tag (so it under-reports a row's
advertised geometry_types — e.g. bulk:pois yields only node points, and
bulk:roads approximates network_type="driving" rather than reproducing the
exact drivable filter).
engine="pyrosm" reads the whole extract in memory and gives the richest,
exact per-layer columns and mixed geometry; it refuses a file over 4 GB, so
never load a planet-wide extract with it. It is opt-in
(pip install pyrosm — see above); selecting it without the engine installed
raises a clear ImportError naming the command. The backend warns before
downloading a multi-GB extract.
Choosing the query — variables#
For this backend variables is the list of named-query ids, not
data-variable names. The <type>: prefix routes the request:
# one named query
EarthLens(data_source="osm", variables=["live:roads"], ...)
# several at once — combined into one FeatureCollection
EarthLens(data_source="osm", variables=["live:hospitals", "live:cafes"], ...)
The shipped named queries are live:hospitals, live:roads,
live:buildings, live:cafes, live:schools, history:buildings,
history:highways, history:amenities, and the bulk:* layers (bulk:buildings,
bulk:roads, bulk:pois, bulk:landuse, bulk:natural, bulk:boundaries). List
them with EarthLens.list_datasets("osm"). An unknown id raises with a
did-you-mean hint.
The facade keys "osm", "openstreetmap", "live", "history", "bulk", and
the back-compat aliases "overpass" / "ohsome" all resolve to the same
backend. The overpass: / ohsome: / pbf: prefixes are likewise accepted as
aliases of live: / history: / bulk: in variables (so overpass:hospitals
== live:hospitals). list_datasets lists the canonical names.
The bbox and the time window#
- bbox —
lat_lim/lon_lim(degrees). The backend hands the box to each type in the order it expects (OverpassS,W,N,E; ohsomeW,S,E,N); you always pass plainlat_lim/lon_lim. You can also use the ergonomicaoi=channel (a bbox, a point +buffer, or a geometry). - time — the
live:type returns current state and ignoresstart/end. Thehistory:type is history-aware and requires a time: passstart=for a single snapshot, orstart=+end=for a range (the backend builds the ohsometimeas"start/end"). Ahistory:query with nostartraises a helpfulValueError.
# history over a multi-year range
EarthLens(
data_source="osm",
variables=["history:highways"],
lat_lim=[49.40, 49.42], lon_lim=[8.67, 8.71],
start="2016-01-01", end="2022-01-01",
path="./out",
).download()
Raw query / filter overrides (power users)#
When the named queries aren't enough, pass your own:
# raw Overpass QL — {bbox} is filled with the request bbox (S,W,N,E)
EarthLens(
data_source="osm",
variables=["live:hospitals"], # still needed to route
lat_lim=[49.40, 49.42], lon_lim=[8.67, 8.71],
query='[out:json][timeout:180];(node["tourism"="museum"]({bbox}););out geom;',
path="./out",
).download()
# raw ohsome filter
EarthLens(
data_source="osm",
variables=["history:buildings"],
lat_lim=[49.40, 49.42], lon_lim=[8.67, 8.71],
start="2020-01-01",
filter="leisure=park and geometry:polygon",
path="./out",
).download()
A raw query= with no {bbox} placeholder is sent verbatim (you supply the
bbox in the QL yourself). A raw Overpass query= must request JSON output
([out:json]) — the response is parsed as JSON, so an [out:xml] / [out:csv]
override will not parse.
Other knobs#
| Keyword | Meaning | Default |
|---|---|---|
endpoint |
the live: request endpoint URL |
https://overpass-api.de/api/interpreter |
user_agent |
User-Agent sent on the live: request (a real one is required) |
earthlens (+…) |
timeout |
live: request timeout (s); also the QL [timeout:N] budget |
180.0 |
file_format |
"geojson" or "gpkg" |
"geojson" |
max_bbox_deg2 |
bbox-area cap (square degrees) — guards the planet-wide footgun (live types only) | 100.0 |
region |
region key or raw "continent/region" path — required for a bulk:* query |
None |
engine |
regional-extract read engine: "pyosmium" (streaming) or "pyrosm" (in-memory, opt-in) |
"pyosmium" |
cache_dir |
directory for cached .osm.pbf extracts |
<cache_dir()>/osm_pbf |
Keep the bbox small (live types)
The live: and history: types are for small/targeted queries. A box
larger than
max_bbox_deg2 (the default 100 square degrees comfortably covers a large
country) is rejected before any request — in particular the whole-Earth
default you get if you omit lat_lim / lon_lim through the facade, which
would hammer the shared public services. Raise max_bbox_deg2= for a
genuinely larger area. The cap does not apply to a bulk read (it hits a
local extract, not a shared service), so a bulk request with no bbox simply
reads the whole downloaded extract.
The returned FeatureCollection#
CRS EPSG:4326. live: features carry osm_id, osm_type, the element's OSM
tags as columns, and a Point / LineString / Polygon geometry (a node → a
point, an open way → a line, a closed way → a polygon; relations are skipped in
the MVP). history: features carry the geometry plus @osmId,
@snapshotTimestamp, and @other_tags. An empty result (a quiet box) comes
back as an empty FeatureCollection with the osm_id / osm_type schema, not an
error.
Writing to disk#
download() writes one file automatically (osm_<ids>.geojson). To write it
yourself:
hospitals.to_file("hospitals.geojson", driver="GeoJSON")
hospitals.to_file("hospitals.gpkg", driver="GPKG")
Licensing — you must attribute (ODbL)#
OSM is ODbL 1.0 (share-alike). Every download() emits a LicenseWarning:
OpenStreetMap data is licensed under the Open Database License (ODbL 1.0),
which carries attribution and share-alike obligations: credit
'(c) OpenStreetMap contributors' ...
Credit "© OpenStreetMap contributors" and license any derived database you redistribute under ODbL.
Aggregation is not supported#
OSM output is vector, so the aggregate= argument is rejected:
EarthLens(data_source="osm", variables=["live:roads"], ...).download(aggregate=cfg)
# NotImplementedError: OSM features are vector, not gridded ...
Post-process the returned FeatureCollection directly instead (it is a GeoDataFrame).
Out of scope#
The history: aggregation queries (counts / areas over time) are not part of
this backend. For bulk asks, reach for the bulk: type (above) rather than tiling
many live queries.