Skip to content

HTTP client (earthlens.base.http)#

HttpClient is the shared requests-based transport for the REST-style backends. It owns the chores every backend used to hand-roll — a pooled session, a sensible User-Agent, a per-request timeout, a Retry-After-aware 429/5xx back-off loop, JSON decoding, and a streamed download with a progress bar — so a new backend keeps only the API-shaped parts: endpoint paths, query params, pagination, auth-header values, and response parsing.

Don't hand-roll a session or a retry loop. If you find yourself writing requests.Session(), a while loop that inspects 429 / Retry-After, or an iter_content chunk loop, reach for HttpClient instead. Backends that pre-date it are being migrated onto it.

Consuming it#

Construct one client per backend with the default headers it needs, then call the verbs:

from earthlens.base.http import HttpClient

http = HttpClient(headers={"X-API-Key": api_key})

# JSON GET with automatic status + transport retry and Retry-After honouring:
payload = http.get_json("https://api.example.org/v1/things", params={"bbox": "1,2,3,4"})

# Streamed download to disk with a tqdm bar (sized from Content-Length):
http.download("https://example.org/big.tif", dest, progress=True)

Retry policy#

Two kinds of failure are retried, on separate budgets, because they warrant opposite policies.

Status retries cover 429 / 500 / 502 / 503 / 504 and honour Retry-After. A 413 / 429 / 503 is the server asking for a later attempt, so it is replayed for any method. A 500 / 502 / 504 means the server already had the request and may have acted on it, so those are replayed only for idempotent methods.

Transport retries cover the failures that never reach a status line — a refused or reset connection, a DNS blip, a read timeout, a body truncated mid-stream. These are on by default; they were opt-in before, which meant a TCP reset partway through a large granule threw away the whole transfer.

The two phases have their own budgets:

Failure Budget Replayed for a POST?
connect — refused, unresolvable, connect timeout connect_retries (1) yes — the request never reached the server
read — reset mid-response, read timeout, truncated body read_retries (= max_retries) no, unless retry_unsafe_methods=True
SSLError, ProxyError never retried under the default set —

A backend that already wraps its calls in its own retry or resilience loop now has two layers: the client retries the transport, and the backend retries the call. The budgets multiply rather than add, so an outer loop of 3 over a client of 5 is up to 18 attempts. If that is not what you want, pass retry_on_exceptions=() to opt the client's transport retry out and keep the outer loop as the single authority.

If your endpoint is a POST that is safe to replay — a search or query API, an idempotent RPC — pass retry_unsafe_methods=True. Without it neither a transport failure nor a 5xx is replayed for that verb, and the suppression is logged at debug level rather than being silent.

The connect budget is small on purpose: a host that refuses a connection rarely starts accepting one within a back-off window, so a generous budget only turns a clear failure into a slow one.

timeout is a (connect, read) pair by default — (10.0, 60.0) — so a dead host fails in ten seconds while a slow transfer keeps a full read budget. A bare float still works and applies to both phases.

The default User-Agent is earthlens/{version} — deliberately non-Mozilla, because the DIGITAL.CSIC Anubis anti-bot wall (SPEIbase) blocks browser-like agents. Pass user_agent= for a descriptive contact string (e.g. Overpass / ohsome etiquette).

Downloads read the whole object, and verify it#

By default download reads the object once, whole, on every attempt: it generates no Range header of its own, and a retry re-requests from byte 0 rather than appending to what is already on disk. Resuming is available but opt-in — see Resuming a single file below.

Restarting is the default because appending to a partial file is only safe if the new bytes provably belong to the same representation as the old ones, and most servers do not give you enough to prove it: Accept-Ranges: bytes is advertised by hosts that then ignore Range and send the whole body from zero; Last-Modified has one-second resolution, so a validator can match across a real change; and a server may answer 416 without the Content-Range that would say how much it actually has. Every one of those produces a file that is the right size and the wrong bytes — corruption that survives to the user rather than failing loudly. The opt-in path below refuses to resume unless the server clears every one of those hazards.

What replaces it is verification after the fact:

Response download does
a body matching its Content-Length publishes it
a body short of it raises IncompleteDownloadError, retrying from the start until the read budget or a repeated byte count stops it
a body longer than it, or the same short count twice raises IncompleteDownloadError without retrying — both repeat
no usable length (chunked, Content-Encoding, contradictory duplicates) publishes it unchecked; there is no claim to check
a 206 to a request that carried no Range raises UnsolicitedPartialContentError without retrying
a break after the last byte, when the size already matches keeps it; the equality is the whole proof

The table above is what verify_length=True (the default) buys. Pass verify_length=False for a server that misreports the size of a body it generates on the fly, where the check would fail a download that is actually fine. It distrusts the advertised length in both directions, so an over-long body is published too, and it disables the salvage, whose proof is the same size equality. It cannot be combined with resume=True, which is addressed by that same length — the pair raises ValueError.

Because the check compares against bytes as delivered, download sends Accept-Encoding: identity — but only when neither the call nor the constructor named that header in any casing, so a backend that needs gzip, or that sets identity itself to protect a magic check, keeps what it asked for.

A Range you pass yourself is honoured verbatim: the response is accepted at its own Content-Length, and download does not check that the server returned the range you asked for.

Interrupted multi-file jobs resume at file granularity — _is_complete() skips the granules already on disk (see contracts).

Resuming a single file (resume=True)#

Within one file, download(url, dest, resume=True) will continue a broken transfer instead of re-reading it. It is off by default: everything above is what the other callers get.

Resume is only attempted when the first response proved it is possible — a 200 advertising Accept-Ranges: bytes, a known Content-Length, no content coding, and a strong ETag. A weak tag (W/"...") never arms, and Last-Modified is never used: its one-second resolution matches across a change made inside the same second.

Arming is the cheap half. A strong ETag is not sufficient — measured, data.worldpop.org sends one and then answers a Range with 200 and the whole body. So the binding checks are on the reply, and all must hold before one byte is appended:

check why
status is 206 absorbs the 200-with-the-whole-object case, a 416, a redirect, an error status
Content-Range start, end and total all match what was asked the object has not been replaced or re-cut
ETag still names the anchored representation a 206 may not rename what If-Range selected
no content or transfer coding Range addresses the encoded octets; the staged count is of decoded ones
the first 64 KiB re-send a window already on disk and match it byte for byte see below

That last one is the only check that reads the body. Every other gate compares one server claim against another, so a server deriving a truthful Content-Range from the request while streaming a different slice satisfies all of them. Re-reading a window we already hold is what makes the append safe.

Anything else discards the staged bytes and re-reads the whole object, and resume then stays off for the rest of the call — so a badly-behaved server costs exactly one extra request, not a restart/resume cycle. Only bytes this call wrote are ever built on, so a .part left by a killed run is never mistaken for a prefix, and the assembled file passes the same length and expect_magic gates as a whole-object read.

Measured: geofabrik, DWD radklim (13.5 GB) and GHSL all emit strong ETags and honour Range exactly; worldpop and Zenodo do not arm or are refused.

Testability#

Both the transport and the wait are injectable, so the whole client is unit-testable with a fake session and no real delays:

client = HttpClient(session=fake_session, sleep=captured_waits.append)

Pagination and response-envelope parsing stay in each backend — the client owns only how bytes move, never the API shape.

API#

earthlens.base.http.HttpClient #

Reusable HTTP transport: session, headers, timeout, retry, download.

Wraps a :class:requests.Session, attaching default headers (a non-Mozilla User-Agent and Accept-Encoding: gzip, deflate) to every request and applying a shared timeout. The verbs (:meth:get, :meth:post, :meth:request) route through a Retry-After-aware back-off loop; :meth:get_json decodes the JSON body and :meth:download streams a response to disk. Both the session and the sleep function are injectable so the client is fully unit-testable without a live network.

Default headers set at construction are merged with (and overridden by) any per-request headers=, so a backend expresses its quirks — openaq's X-API-Key, a descriptive osm contact User-Agent — without subclassing.

Attributes:

Name Type Description
timeout

Per-request timeout in seconds — a single float, or a (connect, read) pair bounding the two phases separately.

max_retries

Maximum retries on a retryable status before raising.

backoff_factor

Base seconds for exponential back-off when no Retry-After header is present.

status_forcelist

HTTP statuses that trigger a retry.

max_backoff

Ceiling in seconds on any single retry wait.

retry_on_exceptions

Transport exception types that trigger a retry.

connect_retries

Retry budget for connect-phase failures.

read_retries

Retry budget for read-phase failures.

retry_unsafe_methods

Whether a read-phase failure may replay a non-idempotent verb such as POST.

raise_for_status

Whether the final response is raise_for_status-ed.

min_interval

Minimum seconds between consecutive requests.

Examples:

  • The default agent is non-Mozilla and version-stamped:
    >>> from earthlens.base.http import HttpClient
    >>> HttpClient().default_headers["User-Agent"].startswith("earthlens/")
    True
    
Source code in libs/core/src/earthlens/base/http.py
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
2341
2342
2343
2344
2345
2346
2347
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
2363
2364
2365
2366
2367
2368
2369
2370
2371
2372
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
2383
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
2400
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
2421
2422
2423
2424
2425
2426
2427
2428
2429
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
2446
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
2458
2459
2460
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
2471
2472
2473
2474
2475
2476
2477
2478
2479
2480
2481
2482
2483
2484
2485
2486
2487
2488
2489
2490
2491
2492
2493
2494
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
2507
2508
2509
2510
2511
2512
2513
2514
2515
2516
2517
2518
2519
2520
2521
2522
2523
2524
2525
2526
2527
2528
2529
2530
2531
2532
2533
2534
2535
2536
2537
2538
2539
2540
2541
2542
2543
2544
2545
2546
2547
2548
class HttpClient:
    """Reusable HTTP transport: session, headers, timeout, retry, download.

    Wraps a :class:`requests.Session`, attaching default headers (a
    non-Mozilla `User-Agent` and `Accept-Encoding: gzip, deflate`) to
    every request and applying a shared timeout. The verbs
    (:meth:`get`, :meth:`post`, :meth:`request`) route through a
    `Retry-After`-aware back-off loop; :meth:`get_json` decodes the JSON
    body and :meth:`download` streams a response to disk. Both the
    session and the sleep function are injectable so the client is fully
    unit-testable without a live network.

    Default headers set at construction are merged with (and overridden
    by) any per-request `headers=`, so a backend expresses its quirks —
    openaq's `X-API-Key`, a descriptive osm contact `User-Agent` — without
    subclassing.

    Attributes:
        timeout: Per-request timeout in seconds — a single float, or a
            `(connect, read)` pair bounding the two phases separately.
        max_retries: Maximum retries on a retryable status before raising.
        backoff_factor: Base seconds for exponential back-off when no
            `Retry-After` header is present.
        status_forcelist: HTTP statuses that trigger a retry.
        max_backoff: Ceiling in seconds on any single retry wait.
        retry_on_exceptions: Transport exception types that trigger a retry.
        connect_retries: Retry budget for connect-phase failures.
        read_retries: Retry budget for read-phase failures.
        retry_unsafe_methods: Whether a read-phase failure may replay a
            non-idempotent verb such as `POST`.
        raise_for_status: Whether the final response is `raise_for_status`-ed.
        min_interval: Minimum seconds between consecutive requests.

    Examples:
        - The default agent is non-Mozilla and version-stamped:
            ```python
            >>> from earthlens.base.http import HttpClient
            >>> HttpClient().default_headers["User-Agent"].startswith("earthlens/")
            True

            ```
    """

    def __init__(
        self,
        *,
        session: requests.Session | None = None,
        user_agent: str | None = None,
        headers: dict[str, str] | None = None,
        timeout: Timeout = DEFAULT_TIMEOUT,
        max_retries: int = DEFAULT_MAX_RETRIES,
        backoff_factor: float = DEFAULT_BACKOFF_FACTOR,
        status_forcelist: tuple[int, ...] = DEFAULT_STATUS_FORCELIST,
        max_backoff: float | None = DEFAULT_MAX_BACKOFF,
        retry_on_exceptions: tuple[type[BaseException], ...] | None = None,
        connect_retries: int | None = None,
        read_retries: int | None = DEFAULT_READ_RETRIES,
        retry_unsafe_methods: bool = False,
        retry_predicate: Callable[[requests.Response], bool] | None = None,
        raise_for_status: bool = True,
        min_interval: float = 0.0,
        clock: Callable[[], float] = time.monotonic,
        sleep: Callable[[float], None] = time.sleep,
    ) -> None:
        """Build a client with default headers, timeout, and retry policy.

        Args:
            session: An existing :class:`requests.Session` to reuse.
                Defaults to a fresh session. Injectable so tests can
                supply a fake transport.
            user_agent: The default `User-Agent` header value. Defaults
                to `earthlens/{version}` (non-Mozilla by design; see
                :func:`_default_user_agent`).
            headers: Extra default headers merged onto every request
                (e.g. `{"X-API-Key": ...}`). Override per call with a
                request-level `headers=`.
            timeout: Per-request timeout in seconds — a single float bounds
                both the connect and read phases, or pass a `(connect, read)`
                pair to bound them separately (a short connect budget fails a
                dead host fast without shortening the read budget).
            max_retries: Maximum retries on a retryable status before the
                last response's error is raised.
            backoff_factor: Base seconds for exponential back-off when a
                response carries no `Retry-After` header.
            status_forcelist: HTTP statuses that trigger a retry.
            max_backoff: Ceiling in seconds on any single retry wait, so a
                large `Retry-After` cannot pin the thread indefinitely.
                `None` disables the cap.
            retry_on_exceptions: Exception types that also trigger a retry
                when raised by the transport. `None` (the default) uses
                :data:`DEFAULT_RETRY_EXCEPTIONS` — a refused or reset
                connection, a timeout, a body truncated mid-stream — and is
                what selects the strict policy: the never-retry list applies
                and `connect_retries` takes its small default. Passing any
                tuple, including an equal-valued one, means the caller owns
                the policy: nothing is vetoed and `connect_retries` follows
                `max_retries`. Pass `()` to retry on status only.
            connect_retries: Retries allowed for a failure in the connect
                phase, counted separately from `read_retries`. `None` (the
                default) resolves to :data:`DEFAULT_CONNECT_RETRIES` while
                the client uses the default exception set — a host that will
                not accept a connection rarely starts doing so within a
                back-off window — and to `max_retries` when the caller
                supplied its own `retry_on_exceptions`, since that caller
                already chose how hard to try.
            read_retries: Retries allowed for a failure after the connection
                was established — a reset mid-response, a read timeout, a
                truncated body. `None` (the default) means `max_retries`.
                This is the transient case a retry actually helps.
            retry_unsafe_methods: Whether a *read*-phase failure may replay a
                non-idempotent verb. `False` (the default) replays only the
                methods in :data:`IDEMPOTENT_METHODS`, because a response
                that broke after the request was delivered may already have
                been acted on, so replaying a `POST` risks a double
                submission. Connect-phase failures are replayed for any
                method regardless, since the request was never delivered.
            retry_predicate: An optional callback `(response) -> bool`
                that, when it returns `True`, marks a response retryable
                even if its status is not in `status_forcelist` (e.g. a
                `200` whose body signals a rate-limit).
            raise_for_status: Whether to call `raise_for_status` on the
                final response. `False` returns the response unraised so
                the caller can inspect the status itself (e.g. to redact a
                secret-bearing URL from the error, or to branch on a
                `4xx`). Overridable per request.
            min_interval: Minimum seconds between consecutive requests
                (a proactive client-side rate limit). `0.0` (default)
                disables throttling.
            clock: Monotonic clock used for the `min_interval` throttle.
                Injectable so tests drive it deterministically.
            sleep: The sleep function used between retries and for the
                throttle. Defaults to :func:`time.sleep`; injectable so
                tests run without real delays.
        """
        self._session = session if session is not None else new_session()
        self._user_agent = user_agent or _default_user_agent()
        self.timeout = timeout
        self.max_retries = max_retries
        self.backoff_factor = backoff_factor
        self.status_forcelist = tuple(status_forcelist)
        self.max_backoff = max_backoff
        # `None` means "use the default set" — an explicit sentinel rather than
        # an identity test against the constant, so a caller who re-spells the
        # same tuple, or writes `DEFAULT_RETRY_EXCEPTIONS + (Extra,)`, gets the
        # behaviour the value implies instead of one that depends on which
        # object it is.
        self._default_retry_set = retry_on_exceptions is None
        self.retry_on_exceptions = tuple(
            DEFAULT_RETRY_EXCEPTIONS
            if retry_on_exceptions is None
            else retry_on_exceptions
        )
        # The small connect budget is a property of the *default* policy. A
        # caller that brought its own `retry_on_exceptions` already chose how
        # hard to try, and silently capping its `max_retries` for one phase
        # would change a number it set deliberately.
        if connect_retries is not None:
            self.connect_retries = connect_retries
        elif self._default_retry_set:
            self.connect_retries = DEFAULT_CONNECT_RETRIES
        else:
            self.connect_retries = max_retries
        self.read_retries = max_retries if read_retries is None else read_retries
        self.retry_unsafe_methods = retry_unsafe_methods
        self._retry_predicate = retry_predicate
        self.raise_for_status = raise_for_status
        self.min_interval = min_interval
        self._clock = clock
        self._sleep = sleep
        self._last_request: float | None = None
        self._throttle_lock = threading.Lock()
        self._default_headers: dict[str, str] = {
            "User-Agent": self._user_agent,
            "Accept-Encoding": "gzip, deflate",
        }
        if headers:
            self._default_headers.update(headers)
        # `_default_headers` stamps the client's own `gzip, deflate` above and
        # then lets the caller's headers overwrite it, so afterwards the two are
        # indistinguishable. `download()` places an `identity` default
        # *underneath* a caller's choice and has to know which it is looking at,
        # so the distinction is recorded here while it still exists.
        self._accept_encoding_is_explicit: bool = any(
            key.lower() == "accept-encoding" for key in (headers or {})
        )

    @property
    def default_headers(self) -> dict[str, str]:
        """Return a copy of the headers merged onto every request."""
        return dict(self._default_headers)

    @property
    def session(self) -> requests.Session:
        """Return the underlying :class:`requests.Session`."""
        return self._session

    def _merge_headers(self, headers: dict[str, str] | None) -> dict[str, str]:
        """Merge per-request `headers` over the client's defaults."""
        merged = dict(self._default_headers)
        if headers:
            merged.update(headers)
        return merged

    def request(
        self,
        method: str,
        url: str,
        *,
        headers: dict[str, str] | None = None,
        timeout: Timeout | None = None,
        raise_for_status: bool | None = None,
        **kwargs: Any,
    ) -> requests.Response:
        """Send one request with the default headers, timeout, and retry.

        Args:
            method: HTTP verb (`"GET"`, `"POST"`, ...).
            url: Absolute request URL.
            headers: Per-request headers merged over the client defaults, and
                passed through verbatim. `download` adds only one header of its
                own, `Accept-Encoding: identity`, and only when neither this
                argument nor the constructor's `headers=` named it in any
                casing — the length check compares against bytes as delivered,
                which a transparently-encoded body would not match. Pass
                `{"Accept-Encoding": "gzip"}` to override that.

                A `Range` here is honoured as written: the response is accepted
                at its own `Content-Length`, and `download` does **not** check
                that the server returned the range you asked for.
            timeout: Per-request timeout override (seconds), as a single
                float or a `(connect, read)` pair. Defaults to the client's
                `timeout`.
            raise_for_status: Per-request override of the client's
                `raise_for_status` policy. `None` (default) uses the
                client setting.
            **kwargs: Extra keyword arguments forwarded to `requests`
                (`params`, `data`, `json`, `stream`, ...).

        Returns:
            requests.Response: The response (after `raise_for_status`
                unless it is disabled).

        Raises:
            requests.HTTPError: On a non-retryable error status, or after
                the retryable status is exhausted (when `raise_for_status`
                is on).
        """
        merged = self._merge_headers(headers)
        effective_timeout = self.timeout if timeout is None else timeout
        return self._request_with_retry(
            method,
            url,
            headers=merged,
            timeout=effective_timeout,
            raise_for_status=raise_for_status,
            **kwargs,
        )

    def get(self, url: str, **kwargs: Any) -> requests.Response:
        """Send a `GET` request. See :meth:`request` for arguments."""
        return self.request("GET", url, **kwargs)

    def post(self, url: str, **kwargs: Any) -> requests.Response:
        """Send a `POST` request. See :meth:`request` for arguments."""
        return self.request("POST", url, **kwargs)

    def get_json(self, url: str, **kwargs: Any) -> Any:
        """Send a `GET` request and decode the JSON response body.

        Convenience over :meth:`get` for the REST endpoints that return
        JSON envelopes.

        Args:
            url: Absolute request URL.
            **kwargs: Keyword arguments forwarded to :meth:`get`
                (`params`, `headers`, `timeout`, ...).

        Returns:
            The parsed JSON body (typically a `dict` or `list`).

        Raises:
            requests.HTTPError: On a non-retryable error status, or after
                the retryable status is exhausted.
        """
        return self.get(url, **kwargs).json()

    def stream(self, url: str, **kwargs: Any) -> requests.Response:
        """Send a streaming `GET` (`stream=True`), retry-wrapped.

        Returns the open response without consuming its body, so the
        caller can iterate `iter_content`. Retries follow the same
        `Retry-After`/back-off policy as the other verbs; the retry
        decision reads only the status line, never the body.

        A body that breaks **after** this returns is therefore the
        caller's to handle: the iteration happens outside the retry loop,
        so a `ChunkedEncodingError` raised mid-stream escapes it. The
        default transport retry covers the non-streaming verbs, whose
        bodies `requests` materialises inside the loop, and
        :meth:`download`, which owns its own. A caller streaming a large
        body that needs the same protection should use :meth:`download`
        or re-request on failure itself.

        Args:
            url: Absolute request URL.
            **kwargs: Keyword arguments forwarded to :meth:`get`.

        Returns:
            requests.Response: The open streaming response.
        """
        return self.get(url, stream=True, **kwargs)

    def _stream_to_file(
        self,
        response: requests.Response,
        dest: Path,
        *,
        chunk: int,
        progress: bool,
        desc: str,
    ) -> None:
        """Write a streaming response's body to `dest` with a `tqdm` bar.

        Args:
            response: The open streaming response.
            dest: The file to write (typically a temp `.part` path).
            chunk: Streaming block size in bytes.
            progress: Whether to show the progress bar.
            desc: The bar label (the final file name).
        """
        total = _progress_total(response.headers)
        bar = tqdm(
            total=total,
            unit="B",
            unit_scale=True,
            unit_divisor=1024,
            disable=not progress,
            desc=desc,
        )
        try:
            with open(dest, "wb") as handle:
                for block in response.iter_content(chunk_size=chunk):
                    if block:
                        handle.write(block)
                        bar.update(len(block))
        finally:
            bar.close()

    def _append_resumed_body(
        self,
        response: requests.Response,
        dest: Path,
        *,
        overlap_at: int,
        overlap: int,
        total: int,
        chunk: int,
        progress: bool,
        desc: str,
    ) -> None:
        """Check the re-sent overlap against the staged file, then append.

        The overlap is read and compared **before the file is opened for
        writing**, so a server whose body does not match the bytes already on
        disk cannot modify the staging file at all. That ordering is the
        guarantee: every other resume gate compares one server claim against
        another, and a server that derives a truthful `Content-Range` from the
        request while streaming a different slice satisfies all of them.

        Args:
            response: The open `206` response, positioned at `overlap_at`.
            dest: The staging file holding the bytes downloaded so far.
            overlap_at: Offset the re-sent window starts at.
            overlap: Length of that window.
            total: The object's complete length, for the progress bar.
            chunk: Streaming block size in bytes.
            progress: Whether to show the progress bar.
            desc: The bar label.

        Raises:
            _ResumeRefused: When the body ends inside the overlap, when the
                overlap does not match the staged bytes, or when the staging
                file cannot be read back for a non-deterministic reason.
            OSError: For a deterministic filesystem refusal (a full disk, a
                read-only mount), which a restart would hit again.
        """
        blocks = response.iter_content(chunk_size=chunk)
        head = bytearray()
        remainder = b""
        for block in blocks:
            if not block:
                continue
            wanted = overlap - len(head)
            head.extend(block[:wanted])
            if len(head) >= overlap:
                remainder = block[wanted:]
                break
        if len(head) < overlap:
            raise _ResumeRefused(
                f"the resumed body ended after {len(head):,} bytes, inside the "
                f"{overlap:,}-byte overlap"
            )

        try:
            with open(dest, "rb") as handle:
                handle.seek(overlap_at)
                staged = handle.read(overlap)
        except OSError as exc:
            if exc.errno in _DETERMINISTIC_OS_ERRNOS:
                raise
            raise _ResumeRefused(
                f"the staged file could not be read back: {exc}"
            ) from exc
        if bytes(head) != staged:
            raise _ResumeRefused(
                f"the re-sent {overlap:,}-byte overlap at {overlap_at:,} does "
                f"not match the staged bytes; the server is not serving the "
                f"representation this transfer started on"
            )

        resume_at = overlap_at + overlap
        bar = tqdm(
            total=total,
            initial=resume_at,
            unit="B",
            unit_scale=True,
            unit_divisor=1024,
            disable=not progress,
            desc=desc,
        )
        try:
            with open(dest, "r+b") as handle:
                # `resume_at` is this call's own byte count for this file, so
                # the seek already lands at its end. The truncate states that
                # invariant rather than guarding a reachable state: were the
                # staged file ever longer, appending after it would leave a gap
                # no size check could see.
                handle.truncate(resume_at)
                handle.seek(resume_at)
                if remainder:
                    handle.write(remainder)
                    bar.update(len(remainder))
                for block in blocks:
                    if block:
                        handle.write(block)
                        bar.update(len(block))
        finally:
            bar.close()

    def download(
        self,
        url: str,
        dest: str | Path,
        *,
        chunk: int = DEFAULT_CHUNK_SIZE,
        progress: bool = True,
        atomic: bool = True,
        expect_magic: bytes | tuple[bytes, ...] | None = None,
        headers: dict[str, str] | None = None,
        timeout: Timeout | None = None,
        resume: bool = False,
        verify_length: bool = True,
        **kwargs: Any,
    ) -> Path:
        """Stream `url` to `dest` atomically, optionally showing a `tqdm` bar.

        Absorbs the chunk-loop the streamed-download backends each
        re-implement: streams with `stream=True`, sizes a progress bar
        from `Content-Length` when present, and writes 1 MiB blocks. When
        `atomic` (the default) it writes to a sibling `<dest>.part` and
        renames on success, and it removes the temp on any failure — so a
        crashed or interrupted download never leaves a truncated `dest`.
        The whole download is retry-wrapped: a status in `status_forcelist`
        or an exception in `retry_on_exceptions` retries the attempt (after
        cleaning the temp), honouring the `Retry-After`/back-off policy.

        By default every attempt reads the object **once, whole**, and the body
        is verified after the fact: a response that advertised a
        `Content-Length` must deliver exactly that many bytes, and a short read
        is retried from the start. Pass `resume=True` to let a retry continue a
        broken transfer instead of re-reading it; see that argument.

        Args:
            url: Absolute request URL.
            dest: Output file path. Parent directories are created.
            chunk: Streaming block size in bytes (default 1 MiB).
            progress: Show a `tqdm` progress bar. `False` (or a
                non-interactive / test context) suppresses it.
            atomic: Write to `<dest>.part` then rename on success, cleaning up
                the temp on failure — so a crashed download never leaves a
                truncated `dest` and never removes an existing one. Keep this on
                (the default) whenever an existing `dest` must survive a failed
                attempt. `False` streams straight into `dest`, which is opened
                `"wb"` and therefore **truncated up front**: a mid-stream failure
                leaves `dest` short, and any previous contents are gone. The
                failure path does not additionally delete it, but that is damage
                limitation, not a guarantee — `atomic=False` is only appropriate
                when `dest` is known to be disposable. The length check below
                runs in both modes; under `atomic=False` a failed check leaves
                the caller-owned `dest` short, because there is no staging file
                to discard.

                Two calls downloading the same `dest` concurrently share one
                `<dest>.part` and will interleave into it. Serialising that is
                the caller's job.
            expect_magic: One or more byte prefixes the body must start with
                (e.g. `b"CDF"` / `b"\\x89HDF"` for NetCDF). A body that starts
                with none of them raises `ValueError` and the partial write is
                discarded, so an HTML error page served with a 200 status never
                lands as a `.nc`. `None` (the default) skips the check.
            headers: Per-request headers merged over the client defaults.
            timeout: Per-request timeout override (seconds), as a single
                float or a `(connect, read)` pair — the latter fails a dead
                host on the short connect budget without shortening the long
                read budget a large download needs.
            resume: Continue a broken transfer with a ranged request instead
                of re-reading the object, when — and only when — the server has
                proved it can be done safely. Off by default; the whole-object
                path above is unchanged for every caller that does not opt in.
                It requires `verify_length=True`, the default — the pair
                `resume=True, verify_length=False` raises `ValueError`.

                Three things make it inert, silently, so they are listed
                together. It requires staging, so it does nothing under
                `atomic=False` unless `expect_magic` forces staging back on:
                resuming means keeping a partial file between attempts, and
                that is only safe in the private `<dest>.part` this method
                owns, never in a caller-supplied `dest` that a failed attempt
                must not damage. It stands down when a `Range` is already
                present in the merged headers — including one set on the
                client, which disables resume for every call on it — because a
                caller who owns the object boundaries should not have ours
                added on top. And a leg only arms once more than
                `_RESUME_OVERLAP` bytes have been banked since the previous leg
                started, because a leg that gains less than the overlap it
                re-fetches costs more than it saves.

                On a leg `download` also overrides two headers rather than
                merging under them: `Accept-Encoding: identity`, because a
                coded reply would make the byte offsets mean something else,
                and `If-Range`, which must name the representation the transfer
                anchored on. This is the one place caller headers are not
                honoured verbatim. A leg is likewise exempt from
                `status_forcelist` and from any `retry_predicate`: both are
                written against whole-object responses, and neither can judge —
                or even parse — a `206` fragment of one. The refusal gates
                below are the leg's own check in their place.

                A resumed request is only *attempted* when the first response
                was a `200` that advertised `Accept-Ranges: bytes`, a known
                `Content-Length`, no content coding, and a **strong** `ETag`
                (never `Last-Modified`, whose one-second resolution matches
                across a change made within the same second). It is only
                *appended* when the reply is a `206` whose `Content-Range`
                start, end and total all match what was asked for, whose `ETag`
                still names the anchored representation, and — the check none of
                the others can stand in for — whose first `_RESUME_OVERLAP`
                bytes re-send a window already on disk and match it byte for
                byte. Every other answer, a `200` with the whole body included,
                discards the staged bytes and re-reads the object; resume then
                stays off for the rest of the call.

                Only bytes *this call* wrote are ever built on, so a `.part`
                left by an earlier run is never treated as a prefix. The
                assembled file goes through the same length and `expect_magic`
                gates as a whole-object read.
            verify_length: Compare the bytes written against the
                `Content-Length` the response advertised, and raise
                `IncompleteDownloadError` on a mismatch. On by default.

                Pass `False` for a server that misreports the size of a body it
                generates on the fly, where the check would fail a download that
                is actually fine. This distrusts the advertised length in
                **both** directions: an over-long body — a mis-framed response,
                a spliced proxy body — is published too, because a length you
                do not believe cannot catch one. It also disables the salvage,
                whose whole proof is that same size equality. And it cannot be
                combined with `resume=True`, which is addressed by the length it
                distrusts: the pair raises `ValueError` before anything is sent.

                It is a separate switch from `Accept-Encoding` on purpose: a
                coded body already has no checkable length — the header counts
                encoded octets and the file counts decoded ones — so asking for
                `gzip` disables the check as a side effect, and that side effect
                should not be the only way to express the intent.
            **kwargs: Extra keyword arguments forwarded to `requests`.

        Returns:
            Path: The `dest` path the bytes were written to.

        Raises:
            ValueError: When `expect_magic` is given and the body does not
                start with any of the supplied prefixes, or when `resume=True`
                is combined with `verify_length=False` — the latter refused up
                front, before a parent directory is created or a request sent.
            UnsolicitedPartialContentError: When a `206` answers a request that
                carried no `Range` — the body is a fragment of the object, and
                repeating the request returns the same fragment, so this is
                raised rather than retried whatever the client's retry policy.
            IncompleteDownloadError: When the bytes written do not equal the
                `Content-Length` the response advertised. A short body is
                retried from the start; a long one, a repeat of the same byte
                count, a spent read or attempt budget, or a client that opted
                out of transport retry raises at once.
            requests.HTTPError: On an error status — `download` always
                calls `raise_for_status` (the client's `raise_for_status`
                flag governs the verb methods, not `download`; a file
                fetch never keeps an error body). Note the resulting
                `HTTPError` is itself subject to `retry_on_exceptions`:
                a client whose `retry_on_exceptions` includes a supertype
                of `requests.HTTPError` (e.g. `requests.RequestException`,
                as ghsl/glaciers pass) will **retry** an error status
                before raising it, mirroring their old download loops.
                This is also how a refused resume surfaces when the ranged
                reply carried an error status and the retry budget is spent:
                the status and `.response` reach the caller intact rather than
                being flattened into a transport error.
            requests.ConnectionError: When a resumed leg is refused, the retry
                budget is spent, and the ranged reply carried no error status
                to report in its place. A refusal with budget left never
                escapes at all: the staged bytes are discarded, the object is
                re-read whole, and resume stays off for the rest of the call.
            BaseException: Once the transport retry budget is spent, the last
                transport failure is re-raised as a fresh instance of its own
                type, with a message naming the attempt count and the largest
                byte count any attempt reached.
            OSError: From `mkdir`, the streaming write, `stat` or `replace`,
                unchanged. `RangeReadError` is never raised here.
        """
        if resume and not verify_length:
            # Contradictory, not merely risky. A resumed leg's `Range` end and
            # its `Content-Range` check are both derived from the advertised
            # total, and the length post-condition is the only thing that
            # proves the assembly reached it — `_resume_refusal` validates the
            # window the server claimed, never the bytes it delivered. Asking
            # to resume is asserting the total is trustworthy; asking to skip
            # the check is asserting it is not. Refuse rather than silently
            # publish a correct-prefix-but-short file.
            raise ValueError(
                "resume=True requires verify_length=True: a resumed transfer is "
                "addressed by the advertised Content-Length, so the length check "
                "is the only proof the assembled file is complete."
            )
        dest = Path(dest)
        dest.parent.mkdir(parents=True, exist_ok=True)
        # `expect_magic` promises a rejected body is discarded, which is only
        # possible if `dest` has not been written yet — so a magic check forces
        # staging even when the caller passed `atomic=False`. Otherwise the
        # validation would run on a file that had already replaced a good one.
        staged = atomic or expect_magic is not None
        tmp = dest.with_name(dest.name + ".part") if staged else dest
        # Bytes of `tmp` that *this call* wrote and that a resumed request
        # may therefore build on. It is never seeded from a file found on
        # disk: a `.part` left by a killed run belongs to an unknown object,
        # and resuming onto it would publish a file assembled from two.
        banked = 0
        # Captured from the first response; frozen thereafter, so every leg
        # is checked against the representation the transfer started on.
        anchor: _ResumeAnchor | None = None
        # One strike. A server that answers a ranged request badly costs a
        # single extra attempt, never a restart/resume cycle.
        resume_off = not resume

        def discard_partial() -> None:
            """Remove the partial write, but never a caller-owned `dest`.

            When the write was staged (`atomic`, or a magic check forcing it)
            the partial lives at a private `<dest>.part`, so removing it is
            always safe. Otherwise the stream wrote straight to `dest`, which
            `_stream_to_file` has already truncated by opening it `"wb"` —
            deleting it as well would only turn a truncated file into a missing
            one, and would destroy a file this call never owned.

            An unlink that fails is logged and swallowed, never raised: this
            runs from every failure path, so a second error here would replace
            the real one on its way out.
            """
            nonlocal banked
            # Reset together: the count means "bytes of `tmp` this call
            # wrote", so it cannot outlive the file it describes.
            banked = 0
            if staged:
                # This runs from every failure path, including
                # `except BaseException`, and on Windows the unlink can raise
                # `PermissionError` when the handle has only just been released.
                # Removing the temp is a promise, but it must not become a
                # second failure that replaces the real one — so it is caught
                # and logged rather than suppressed silently, since the file it
                # promised to remove is still there.
                try:
                    tmp.unlink(missing_ok=True)
                except OSError as exc:
                    logger.debug(f"could not remove {tmp.name}: {exc}")

        merged = self._merge_headers(headers)
        # Header names are case-insensitive, and `_merge_headers` returns a
        # plain dict, so every test on them is folded explicitly.
        per_call = {key.lower() for key in (headers or {})}
        caller_sent_range = "range" in {key.lower() for key in merged}
        if "accept-encoding" not in per_call and not self._accept_encoding_is_explicit:
            # The length check below compares against the bytes as delivered, so
            # a transparently-encoded body would fail it. This sits *underneath*
            # the caller: a backend that asked for `gzip`, or for `identity` to
            # protect its own magic check, keeps what it asked for.
            merged["Accept-Encoding"] = "identity"
        effective_timeout = self.timeout if timeout is None else timeout
        attempt = 0
        # Per-kind accounting, as in `_request_with_retry`: a dead host should
        # not cost a large download six full connect attempts.
        spent = {"connect": 0, "read": 0}
        # The byte count the previous attempt reached. A second attempt landing
        # on the identical number is deterministic, not a transport blip.
        last_written: int | None = None
        best_written = 0
        # What the previous leg had banked, so per-leg progress is checkable.
        last_banked = 0
        if staged and tmp.exists():
            # Suppressed for the same reason as the salvage stat below: this is
            # a log line, and it runs before the first request, so a filesystem
            # blip here must not cost the whole download.
            with contextlib.suppress(OSError):
                logger.debug(
                    f"{redact_url(url)}: discarding a pre-existing "
                    f"{tmp.stat().st_size:,}-byte {tmp.name}"
                )

        def salvaged(expected: int | None) -> bool:
            """Whether the bytes on disk already are the whole object.

            A transfer can break after the last byte but before the stream
            frames its end. The equality is the entire proof, so this asks no
            question of the server and consults no retry policy. It also never
            raises: it is the first thing a transport handler runs, and an
            exception from inside a handler escapes past its own `try`.

            Args:
                expected: The length the response advertised, if any.

            Returns:
                bool: True when the staged file is exactly that long. False
                when the check is switched off, when the response advertised no
                length, when nothing is staged, and when the size cannot be
                read at all — an unprovable salvage is not a salvage, so the
                ordinary retry path takes over.
            """
            if not verify_length or expected is None or not tmp.exists():
                return False
            try:
                return tmp.stat().st_size == expected
            except OSError:
                # This runs as the FIRST statement of both transport `except`
                # handlers, and an exception raised inside a handler is not
                # caught by its own `try` — so an unguarded failure here would
                # leave `download` with the filesystem's errno instead of the
                # transport error, and with none of the retry budget spent.
                # Unprovable is not salvageable: fall through to the ordinary
                # retry path.
                return False

        def salvage_publishable(expected: int | None) -> bool:
            """Whether the staged bytes are the whole object *and* pass every gate.

            The salvage runs `expect_magic` itself rather than falling through to
            the publish. Without it a body that breaks after its last byte
            reaches `dest` through a weaker door than one that does not, so a
            200-served HTML error page whose `Content-Length` happens to match
            would land under a `.nc` name with the magic check never run.

            Args:
                expected: The length the response advertised, if any.

            Returns:
                bool: True when the file may be published as it stands.

            Raises:
                ValueError: When `expect_magic` is set and the staged body does
                    not start with one of the prefixes. The partial is discarded
                    first, exactly as on the whole-object path.
            """
            if not salvaged(expected):
                return False
            if expect_magic is not None:
                try:
                    _check_magic(tmp, expect_magic, url)
                except BaseException:
                    discard_partial()
                    raise
            return True

        while True:
            self._throttle()
            response: requests.Response | None = None
            expected_total: int | None = None
            status_wait: float | None = None
            # A resumed leg re-fetches `_RESUME_OVERLAP` bytes it already
            # holds, so `banked` must exceed the overlap for the leg to make
            # forward progress at all.
            resume_at = 0
            leg: _ResumeAnchor | None = None
            if (
                not resume_off
                and staged
                and anchor is not None
                and not caller_sent_range
                and _RESUME_OVERLAP < banked < anchor.total
                # Each leg re-fetches the overlap, so a leg that advanced less
                # than that cost more than it saved. Without this a server
                # dribbling a byte per leg burns the whole read budget
                # re-sending 64 KiB apiece, where a plain restart is cheaper.
                and banked - last_banked > _RESUME_OVERLAP
            ):
                resume_at, leg = banked, anchor
                last_banked = banked
            attempt_headers = merged
            if leg is not None:
                attempt_headers = {
                    **merged,
                    "Range": f"bytes={resume_at - _RESUME_OVERLAP}-{leg.total - 1}",
                    "If-Range": leg.etag,
                    # `Range` addresses the *encoded* representation while
                    # the staged count is of decoded bytes, so a coded reply
                    # would make the offsets mean different things.
                    "Accept-Encoding": "identity",
                }
            try:
                # Inside the retrying `try`: a failure here is a connect-phase
                # failure and must be retried like any other, which is why the
                # send and the body processing share one handler.
                response = self._send(
                    "GET",
                    url,
                    stream=True,
                    headers=attempt_headers,
                    timeout=effective_timeout,
                    **kwargs,
                )
                try:
                    # A resumed leg is exempt: `status_forcelist` and a
                    # `retry_predicate` are written against whole-object
                    # responses (the existing ones call `r.json()`), and handing
                    # one a `206` byte fragment would have it judge — or fail to
                    # parse — a slice of the object. The leg has its own gate.
                    retryable = leg is None and (
                        response.status_code in self.status_forcelist
                        or (
                            self._retry_predicate is not None
                            and self._retry_predicate(response)
                        )
                    )
                    if retryable and attempt < self.max_retries:
                        # Recorded, not slept on: the sleep happens after the
                        # `finally` below has released the streamed connection,
                        # rather than holding a socket open across the back-off.
                        status_wait = self._backoff_wait(
                            _parse_retry_after(response.headers.get("Retry-After")),
                            attempt,
                        )
                        logger.warning(
                            f"HTTP {response.status_code} on {redact_url(url)}; "
                            f"retry {attempt + 1}/{self.max_retries} after "
                            f"{status_wait:.1f}s"
                        )
                    else:
                        if leg is not None:
                            refusal = _resume_refusal(response, leg, resume_at)
                            if refusal is not None:
                                raise _ResumeRefused(refusal)
                            expected_total = leg.total
                            self._append_resumed_body(
                                response,
                                tmp,
                                overlap_at=resume_at - _RESUME_OVERLAP,
                                overlap=_RESUME_OVERLAP,
                                total=leg.total,
                                chunk=chunk,
                                progress=progress,
                                desc=dest.name,
                            )
                        else:
                            if response.status_code == 206 and not caller_sent_range:
                                raise UnsolicitedPartialContentError(
                                    f"{redact_url(url)} answered 206 to a request "
                                    "carrying no Range; the body is a fragment",
                                    response=response,
                                )
                            response.raise_for_status()
                            expected_total = _progress_total(response.headers)
                            if not resume_off and anchor is None and staged:
                                # Recorded before the body is read, so a break
                                # part-way through still leaves an anchor to
                                # resume against.
                                anchor = _arm_resume_anchor(response, expected_total)
                            self._stream_to_file(
                                response,
                                tmp,
                                chunk=chunk,
                                progress=progress,
                                desc=dest.name,
                            )
                        # One verification for both paths: a resumed file is
                        # published through the same size and magic gates as a
                        # whole-object read, never a weaker door.
                        written = tmp.stat().st_size
                        best_written = max(best_written, written)
                        if (
                            verify_length
                            and expected_total is not None
                            and written != expected_total
                        ):
                            raise IncompleteDownloadError(
                                f"{redact_url(url)} delivered {written:,} of "
                                f"{expected_total:,} advertised bytes",
                                written=written,
                                expected=expected_total,
                            )
                        if expect_magic is not None:
                            _check_magic(tmp, expect_magic, url)
                finally:
                    if response is not None:
                        response.close()
            except _ResumeRefused as exc:
                # Never re-armed for the rest of the call, so a server that
                # answers a ranged request badly costs exactly one extra
                # attempt rather than a restart/resume cycle.
                resume_off = True
                logger.warning(
                    f"{redact_url(url)}: refusing to resume ({exc}); discarding "
                    f"{banked:,} staged bytes and re-reading the whole object"
                )
                discard_partial()
                if attempt >= self.max_retries:
                    # A backstop, not the primary bound: `resume_off` already
                    # caps refusals at one per call. Without it the loop's
                    # termination would rest on that single flag, and a refusal
                    # would otherwise re-read the object with no budget left.
                    if response is not None and response.status_code >= 400:
                        # Report what the server actually said. Backends branch
                        # on `exc.response.status_code` to tell a missing
                        # granule (404/410) from a real failure, and a
                        # synthesised `ConnectionError` would both drop the
                        # status and reclassify a permanent error as transport.
                        response.raise_for_status()
                    raise requests.ConnectionError(
                        f"{redact_url(url)}: resume refused ({exc}) with no "
                        f"retry budget left after {attempt + 1} attempts"
                    ) from exc
                attempt += 1
                continue
            except UnsolicitedPartialContentError:
                # Deterministic: a server that volunteers partial content to a
                # Range-less request returns the same fragment next time.
                discard_partial()
                raise
            except IncompleteDownloadError as exc:
                over_delivered = (
                    exc.written is not None
                    and exc.expected is not None
                    and exc.written > exc.expected
                )
                if (
                    over_delivered
                    or not self.retry_on_exceptions
                    or exc.written == last_written
                    or spent["read"] >= self.read_retries
                    or attempt >= self.max_retries
                ):
                    discard_partial()
                    raise
                spent["read"] += 1
                last_written = exc.written
                logger.warning(
                    f"{redact_url(url)} delivered {exc.written:,} of "
                    f"{exc.expected:,} bytes; discarding and restarting, read "
                    f"retry {spent['read']}/{self.read_retries} "
                    f"(attempt {attempt + 1}/{self.max_retries})"
                )
                discard_partial()
                self._sleep(self._backoff_wait(None, attempt))
                attempt += 1
                continue
            except self.retry_on_exceptions as exc:
                if salvage_publishable(expected_total):
                    logger.debug(
                        f"{redact_url(url)} broke after the last byte; the "
                        f"{expected_total:,}-byte body is already complete"
                    )
                else:
                    kind = classify_transport_error(exc, strict=self._default_retry_set)
                    if kind is None:
                        # Deterministic - the next attempt reproduces it.
                        discard_partial()
                        raise
                    # `response is None` is exactly "the failure happened before
                    # a response object existed", which is the connect phase.
                    if kind == "connect" or (kind == "unknown" and response is None):
                        key, budget = "connect", self.connect_retries
                    else:
                        key, budget = "read", self.read_retries
                    if spent[key] >= budget or attempt >= self.max_retries:
                        discard_partial()
                        raise type(exc)(
                            f"{redact_url(url)} failed after {attempt + 1} "
                            f"attempts; best read {best_written:,} bytes"
                        ) from exc
                    spent[key] += 1
                    try:
                        staged_now = tmp.stat().st_size if tmp.exists() else 0
                    except OSError:
                        staged_now = 0
                    best_written = max(best_written, staged_now)
                    kept = 0
                    if not resume_off and staged and anchor is not None:
                        # Bank only what THIS attempt wrote: `_stream_to_file`
                        # truncated `tmp` before writing, so whatever is there
                        # now came from this call and from this representation.
                        size = staged_now
                        if _RESUME_OVERLAP < size < anchor.total:
                            kept = size
                    banked = kept
                    logger.warning(
                        f"{type(exc).__name__} ({kind}) on {redact_url(url)} after "
                        # The size on disk right now, never `best_written`: that
                        # is a high-water mark across attempts and outlives the
                        # file that produced it, so banking it would point a
                        # resumed request past the end of the staged file.
                        f"{staged_now:,} bytes; "
                        + (
                            f"keeping {kept:,} staged bytes to resume from, "
                            if kept
                            else "discarding and restarting, "
                        )
                        + f"{key} retry {spent[key]}/{budget} "
                        f"(attempt {attempt + 1}/{self.max_retries})"
                    )
                    if not kept:
                        discard_partial()
                    self._sleep(self._backoff_wait(None, attempt))
                    attempt += 1
                    continue
            except requests.RequestException:
                # Wider than the client's retry set on purpose: the salvage
                # issues no request and reads no header, so the caller's retry
                # policy is not the right gate - the size equality is the proof.
                if not salvage_publishable(expected_total):
                    discard_partial()
                    raise
                logger.debug(
                    f"{redact_url(url)} broke after the last byte; the "
                    f"{expected_total:,}-byte body is already complete"
                )
            except BaseException:
                discard_partial()
                raise
            if status_wait is not None:
                # The response is closed by now, so nothing is held open.
                self._sleep(status_wait)
                attempt += 1
                continue
            if staged:
                # Guard the rename too, so the "removes the temp on any
                # failure" promise holds if the final replace fails.
                try:
                    tmp.replace(dest)
                except BaseException:
                    discard_partial()
                    raise
            return dest

    def _throttle(self) -> None:
        """Sleep so consecutive requests are >= `min_interval` apart.

        A no-op when `min_interval` is `0`. Uses the injected monotonic
        `clock`, records the send time, and sleeps via the injected
        `sleep` so tests drive the rate limit deterministically.

        Held under a lock for the whole read-sleep-write sequence, so
        `min_interval` bounds the *aggregate* request rate rather than the
        rate per thread. Read-then-write without it is a race: every thread
        sees the same `_last_request`, each concludes the interval has
        elapsed, and they all fire together — the burst the limit exists to
        prevent.

        No backend shares one client across threads *today*: the three that
        thread (worldpop, ghsl, hdx, all `prefer="threads"`) build a fresh
        client inside each worker, and `_run_items` is a sequential loop. The
        lock is cheap and makes a shared client safe whenever one appears,
        rather than leaving a latent race for that change to trip over.
        """
        if self.min_interval <= 0:
            return
        with self._throttle_lock:
            if self._last_request is not None:
                remaining = self.min_interval - (self._clock() - self._last_request)
                if remaining > 0:
                    self._sleep(remaining)
            self._last_request = self._clock()

    def _backoff_wait(self, retry_after: float | None, attempt: int) -> float:
        """Compute one retry wait: `Retry-After` else exponential back-off.

        Applies the `max_backoff` ceiling and a non-negative floor.

        Args:
            retry_after: Parsed `Retry-After` seconds, or `None`.
            attempt: The zero-based attempt index.

        Returns:
            The clamped wait in seconds.
        """
        wait = (
            retry_after
            if retry_after is not None
            else self.backoff_factor * (2**attempt)
        )
        if self.max_backoff is not None:
            wait = min(wait, self.max_backoff)
        return max(0.0, wait)

    def _request_with_retry(
        self,
        method: str,
        url: str,
        *,
        raise_for_status: bool | None = None,
        **kwargs: Any,
    ) -> requests.Response:
        """Send one request, retrying statuses/exceptions with back-off.

        Retries while the attempt budget holds and either the response
        status is in `status_forcelist`, the `retry_predicate` marks the
        response retryable, or the transport raised one of
        `retry_on_exceptions`. Waits `Retry-After` seconds when present
        and numeric, otherwise `backoff_factor * 2**attempt` (capped by
        `max_backoff`). Honours the `min_interval` throttle before every
        send. Once non-retryable or exhausted, the response is
        `raise_for_status`-ed unless that is disabled, then returned.

        Args:
            method: HTTP verb.
            url: Absolute request URL.
            raise_for_status: Per-request override; `None` uses the
                client policy.
            **kwargs: Keyword arguments forwarded to the session.

        Returns:
            requests.Response: The final response.

        Raises:
            requests.HTTPError: On a non-retryable error status when
                `raise_for_status` is on.
            BaseException: The last transport exception, re-raised after
                the `retry_on_exceptions` budget is exhausted.
        """
        effective_raise = (
            self.raise_for_status if raise_for_status is None else raise_for_status
        )
        attempt = 0
        # Keyed by *budget*, not by classification: `connect` and `unknown`
        # draw on the same allowance, so counting them apart would let a request
        # alternating between the two spend it twice.
        spent = {"connect": 0, "read": 0}
        while True:
            self._throttle()
            try:
                response = self._send(method, url, **kwargs)
            except self.retry_on_exceptions as exc:
                kind = classify_transport_error(exc, strict=self._default_retry_set)
                if kind is None:
                    # Deterministic — the next attempt reproduces it exactly.
                    raise
                # An unidentified failure spends the cheap budget but is
                # replayed as cautiously as a read one.
                cheap = kind in {"connect", "unknown"}
                key = "connect" if cheap else "read"
                budget = self.connect_retries if cheap else self.read_retries
                # Two guards, not one: the per-kind budget shapes *which* failure
                # is worth repeating, and `max_retries` still caps the total, so
                # a request alternating between the two phases cannot make
                # `connect_retries + read_retries` attempts.
                if spent[key] >= budget or attempt >= self.max_retries:
                    raise
                # A *read* failure means the request was delivered and the
                # server may already have acted, so a non-idempotent verb is
                # replayed only when the caller vouches for it. Only a *connect*
                # failure is exempt, because it provably never arrived.
                if kind != "connect" and not (
                    self.retry_unsafe_methods or method.upper() in IDEMPOTENT_METHODS
                ):
                    logger.debug(
                        f"{type(exc).__name__} ({kind}) on {redact_url(url)}; not "
                        f"retrying a {method.upper()}: the request may have been "
                        f"delivered (pass retry_unsafe_methods=True if it is "
                        f"replay-safe)"
                    )
                    raise
                wait = self._backoff_wait(None, attempt)
                logger.warning(
                    f"{type(exc).__name__} ({kind}) on {redact_url(url)}; retry "
                    f"{spent[key] + 1}/{budget} after {wait:.1f}s"
                )
                self._sleep(wait)
                spent[key] += 1
                attempt += 1
                continue
            by_status = response.status_code in self.status_forcelist
            by_predicate = self._retry_predicate is not None and self._retry_predicate(
                response
            )
            if (
                by_status
                and response.status_code not in RETRY_AFTER_STATUS_CODES
                and not (
                    self.retry_unsafe_methods or method.upper() in IDEMPOTENT_METHODS
                )
            ):
                # `500` / `502` / `504` report that the server already had the
                # request, so replaying a non-idempotent verb risks the same
                # double submission the read-phase gate prevents. A `429` /
                # `503` / `413` is the server asking for a later attempt, which
                # is safe for any method.
                logger.debug(
                    f"HTTP {response.status_code} on {redact_url(url)}; not "
                    f"retrying a {method.upper()}: the server already had the "
                    f"request (pass retry_unsafe_methods=True if it is "
                    f"replay-safe)"
                )
                by_status = False
            # A `retry_predicate` is the caller inspecting the response and
            # saying "this one is not really a success" — often on a `200` whose
            # body carries a rate-limit. That is their judgement about their own
            # endpoint, so the verb gate does not overrule it.
            retryable = by_status or by_predicate
            if retryable and attempt < self.max_retries:
                retry_after = _parse_retry_after(response.headers.get("Retry-After"))
                wait = self._backoff_wait(retry_after, attempt)
                logger.warning(
                    f"HTTP {response.status_code} on {redact_url(url)}; retry "
                    f"{attempt + 1}/{self.max_retries} after {wait:.1f}s"
                )
                # Release the (possibly streamed) connection before retrying;
                # a stream=True body is otherwise never consumed and its socket
                # leaks out of the pool.
                response.close()
                self._sleep(wait)
                attempt += 1
                continue
            if effective_raise:
                try:
                    response.raise_for_status()
                except requests.HTTPError:
                    # Close the final errored response too — for a streamed
                    # request the caller never receives it, so nothing else would.
                    response.close()
                    raise
            return response

    def _send(self, method: str, url: str, **kwargs: Any) -> requests.Response:
        """Dispatch to the session's verb method, falling back to `request`.

        Routing `GET` to `session.get` (rather than a generic
        `session.request`) keeps drop-in fake transports that implement
        only `get()` working — the shape the migrated backends' tests use.

        Args:
            method: HTTP verb.
            url: Absolute request URL.
            **kwargs: Keyword arguments forwarded to the session call.

        Returns:
            requests.Response: The raw response (no status check).
        """
        verb = getattr(self._session, method.lower(), None)
        if callable(verb):
            return cast("requests.Response", verb(url, **kwargs))
        return self._session.request(method, url, **kwargs)

default_headers property #

Return a copy of the headers merged onto every request.

session property #

Return the underlying :class:requests.Session.

__init__(*, session=None, user_agent=None, headers=None, timeout=DEFAULT_TIMEOUT, max_retries=DEFAULT_MAX_RETRIES, backoff_factor=DEFAULT_BACKOFF_FACTOR, status_forcelist=DEFAULT_STATUS_FORCELIST, max_backoff=DEFAULT_MAX_BACKOFF, retry_on_exceptions=None, connect_retries=None, read_retries=DEFAULT_READ_RETRIES, retry_unsafe_methods=False, retry_predicate=None, raise_for_status=True, min_interval=0.0, clock=time.monotonic, sleep=time.sleep) #

Build a client with default headers, timeout, and retry policy.

Parameters:

Name Type Description Default
session Session | None

An existing :class:requests.Session to reuse. Defaults to a fresh session. Injectable so tests can supply a fake transport.

None
user_agent str | None

The default User-Agent header value. Defaults to earthlens/{version} (non-Mozilla by design; see :func:_default_user_agent).

None
headers dict[str, str] | None

Extra default headers merged onto every request (e.g. {"X-API-Key": ...}). Override per call with a request-level headers=.

None
timeout Timeout

Per-request timeout in seconds — a single float bounds both the connect and read phases, or pass a (connect, read) pair to bound them separately (a short connect budget fails a dead host fast without shortening the read budget).

DEFAULT_TIMEOUT
max_retries int

Maximum retries on a retryable status before the last response's error is raised.

DEFAULT_MAX_RETRIES
backoff_factor float

Base seconds for exponential back-off when a response carries no Retry-After header.

DEFAULT_BACKOFF_FACTOR
status_forcelist tuple[int, ...]

HTTP statuses that trigger a retry.

DEFAULT_STATUS_FORCELIST
max_backoff float | None

Ceiling in seconds on any single retry wait, so a large Retry-After cannot pin the thread indefinitely. None disables the cap.

DEFAULT_MAX_BACKOFF
retry_on_exceptions tuple[type[BaseException], ...] | None

Exception types that also trigger a retry when raised by the transport. None (the default) uses :data:DEFAULT_RETRY_EXCEPTIONS — a refused or reset connection, a timeout, a body truncated mid-stream — and is what selects the strict policy: the never-retry list applies and connect_retries takes its small default. Passing any tuple, including an equal-valued one, means the caller owns the policy: nothing is vetoed and connect_retries follows max_retries. Pass () to retry on status only.

None
connect_retries int | None

Retries allowed for a failure in the connect phase, counted separately from read_retries. None (the default) resolves to :data:DEFAULT_CONNECT_RETRIES while the client uses the default exception set — a host that will not accept a connection rarely starts doing so within a back-off window — and to max_retries when the caller supplied its own retry_on_exceptions, since that caller already chose how hard to try.

None
read_retries int | None

Retries allowed for a failure after the connection was established — a reset mid-response, a read timeout, a truncated body. None (the default) means max_retries. This is the transient case a retry actually helps.

DEFAULT_READ_RETRIES
retry_unsafe_methods bool

Whether a read-phase failure may replay a non-idempotent verb. False (the default) replays only the methods in :data:IDEMPOTENT_METHODS, because a response that broke after the request was delivered may already have been acted on, so replaying a POST risks a double submission. Connect-phase failures are replayed for any method regardless, since the request was never delivered.

False
retry_predicate Callable[[Response], bool] | None

An optional callback (response) -> bool that, when it returns True, marks a response retryable even if its status is not in status_forcelist (e.g. a 200 whose body signals a rate-limit).

None
raise_for_status bool

Whether to call raise_for_status on the final response. False returns the response unraised so the caller can inspect the status itself (e.g. to redact a secret-bearing URL from the error, or to branch on a 4xx). Overridable per request.

True
min_interval float

Minimum seconds between consecutive requests (a proactive client-side rate limit). 0.0 (default) disables throttling.

0.0
clock Callable[[], float]

Monotonic clock used for the min_interval throttle. Injectable so tests drive it deterministically.

monotonic
sleep Callable[[float], None]

The sleep function used between retries and for the throttle. Defaults to :func:time.sleep; injectable so tests run without real delays.

sleep
Source code in libs/core/src/earthlens/base/http.py
def __init__(
    self,
    *,
    session: requests.Session | None = None,
    user_agent: str | None = None,
    headers: dict[str, str] | None = None,
    timeout: Timeout = DEFAULT_TIMEOUT,
    max_retries: int = DEFAULT_MAX_RETRIES,
    backoff_factor: float = DEFAULT_BACKOFF_FACTOR,
    status_forcelist: tuple[int, ...] = DEFAULT_STATUS_FORCELIST,
    max_backoff: float | None = DEFAULT_MAX_BACKOFF,
    retry_on_exceptions: tuple[type[BaseException], ...] | None = None,
    connect_retries: int | None = None,
    read_retries: int | None = DEFAULT_READ_RETRIES,
    retry_unsafe_methods: bool = False,
    retry_predicate: Callable[[requests.Response], bool] | None = None,
    raise_for_status: bool = True,
    min_interval: float = 0.0,
    clock: Callable[[], float] = time.monotonic,
    sleep: Callable[[float], None] = time.sleep,
) -> None:
    """Build a client with default headers, timeout, and retry policy.

    Args:
        session: An existing :class:`requests.Session` to reuse.
            Defaults to a fresh session. Injectable so tests can
            supply a fake transport.
        user_agent: The default `User-Agent` header value. Defaults
            to `earthlens/{version}` (non-Mozilla by design; see
            :func:`_default_user_agent`).
        headers: Extra default headers merged onto every request
            (e.g. `{"X-API-Key": ...}`). Override per call with a
            request-level `headers=`.
        timeout: Per-request timeout in seconds — a single float bounds
            both the connect and read phases, or pass a `(connect, read)`
            pair to bound them separately (a short connect budget fails a
            dead host fast without shortening the read budget).
        max_retries: Maximum retries on a retryable status before the
            last response's error is raised.
        backoff_factor: Base seconds for exponential back-off when a
            response carries no `Retry-After` header.
        status_forcelist: HTTP statuses that trigger a retry.
        max_backoff: Ceiling in seconds on any single retry wait, so a
            large `Retry-After` cannot pin the thread indefinitely.
            `None` disables the cap.
        retry_on_exceptions: Exception types that also trigger a retry
            when raised by the transport. `None` (the default) uses
            :data:`DEFAULT_RETRY_EXCEPTIONS` — a refused or reset
            connection, a timeout, a body truncated mid-stream — and is
            what selects the strict policy: the never-retry list applies
            and `connect_retries` takes its small default. Passing any
            tuple, including an equal-valued one, means the caller owns
            the policy: nothing is vetoed and `connect_retries` follows
            `max_retries`. Pass `()` to retry on status only.
        connect_retries: Retries allowed for a failure in the connect
            phase, counted separately from `read_retries`. `None` (the
            default) resolves to :data:`DEFAULT_CONNECT_RETRIES` while
            the client uses the default exception set — a host that will
            not accept a connection rarely starts doing so within a
            back-off window — and to `max_retries` when the caller
            supplied its own `retry_on_exceptions`, since that caller
            already chose how hard to try.
        read_retries: Retries allowed for a failure after the connection
            was established — a reset mid-response, a read timeout, a
            truncated body. `None` (the default) means `max_retries`.
            This is the transient case a retry actually helps.
        retry_unsafe_methods: Whether a *read*-phase failure may replay a
            non-idempotent verb. `False` (the default) replays only the
            methods in :data:`IDEMPOTENT_METHODS`, because a response
            that broke after the request was delivered may already have
            been acted on, so replaying a `POST` risks a double
            submission. Connect-phase failures are replayed for any
            method regardless, since the request was never delivered.
        retry_predicate: An optional callback `(response) -> bool`
            that, when it returns `True`, marks a response retryable
            even if its status is not in `status_forcelist` (e.g. a
            `200` whose body signals a rate-limit).
        raise_for_status: Whether to call `raise_for_status` on the
            final response. `False` returns the response unraised so
            the caller can inspect the status itself (e.g. to redact a
            secret-bearing URL from the error, or to branch on a
            `4xx`). Overridable per request.
        min_interval: Minimum seconds between consecutive requests
            (a proactive client-side rate limit). `0.0` (default)
            disables throttling.
        clock: Monotonic clock used for the `min_interval` throttle.
            Injectable so tests drive it deterministically.
        sleep: The sleep function used between retries and for the
            throttle. Defaults to :func:`time.sleep`; injectable so
            tests run without real delays.
    """
    self._session = session if session is not None else new_session()
    self._user_agent = user_agent or _default_user_agent()
    self.timeout = timeout
    self.max_retries = max_retries
    self.backoff_factor = backoff_factor
    self.status_forcelist = tuple(status_forcelist)
    self.max_backoff = max_backoff
    # `None` means "use the default set" — an explicit sentinel rather than
    # an identity test against the constant, so a caller who re-spells the
    # same tuple, or writes `DEFAULT_RETRY_EXCEPTIONS + (Extra,)`, gets the
    # behaviour the value implies instead of one that depends on which
    # object it is.
    self._default_retry_set = retry_on_exceptions is None
    self.retry_on_exceptions = tuple(
        DEFAULT_RETRY_EXCEPTIONS
        if retry_on_exceptions is None
        else retry_on_exceptions
    )
    # The small connect budget is a property of the *default* policy. A
    # caller that brought its own `retry_on_exceptions` already chose how
    # hard to try, and silently capping its `max_retries` for one phase
    # would change a number it set deliberately.
    if connect_retries is not None:
        self.connect_retries = connect_retries
    elif self._default_retry_set:
        self.connect_retries = DEFAULT_CONNECT_RETRIES
    else:
        self.connect_retries = max_retries
    self.read_retries = max_retries if read_retries is None else read_retries
    self.retry_unsafe_methods = retry_unsafe_methods
    self._retry_predicate = retry_predicate
    self.raise_for_status = raise_for_status
    self.min_interval = min_interval
    self._clock = clock
    self._sleep = sleep
    self._last_request: float | None = None
    self._throttle_lock = threading.Lock()
    self._default_headers: dict[str, str] = {
        "User-Agent": self._user_agent,
        "Accept-Encoding": "gzip, deflate",
    }
    if headers:
        self._default_headers.update(headers)
    # `_default_headers` stamps the client's own `gzip, deflate` above and
    # then lets the caller's headers overwrite it, so afterwards the two are
    # indistinguishable. `download()` places an `identity` default
    # *underneath* a caller's choice and has to know which it is looking at,
    # so the distinction is recorded here while it still exists.
    self._accept_encoding_is_explicit: bool = any(
        key.lower() == "accept-encoding" for key in (headers or {})
    )

download(url, dest, *, chunk=DEFAULT_CHUNK_SIZE, progress=True, atomic=True, expect_magic=None, headers=None, timeout=None, resume=False, verify_length=True, **kwargs) #

Stream url to dest atomically, optionally showing a tqdm bar.

Absorbs the chunk-loop the streamed-download backends each re-implement: streams with stream=True, sizes a progress bar from Content-Length when present, and writes 1 MiB blocks. When atomic (the default) it writes to a sibling <dest>.part and renames on success, and it removes the temp on any failure — so a crashed or interrupted download never leaves a truncated dest. The whole download is retry-wrapped: a status in status_forcelist or an exception in retry_on_exceptions retries the attempt (after cleaning the temp), honouring the Retry-After/back-off policy.

By default every attempt reads the object once, whole, and the body is verified after the fact: a response that advertised a Content-Length must deliver exactly that many bytes, and a short read is retried from the start. Pass resume=True to let a retry continue a broken transfer instead of re-reading it; see that argument.

Parameters:

Name Type Description Default
url str

Absolute request URL.

required
dest str | Path

Output file path. Parent directories are created.

required
chunk int

Streaming block size in bytes (default 1 MiB).

DEFAULT_CHUNK_SIZE
progress bool

Show a tqdm progress bar. False (or a non-interactive / test context) suppresses it.

True
atomic bool

Write to <dest>.part then rename on success, cleaning up the temp on failure — so a crashed download never leaves a truncated dest and never removes an existing one. Keep this on (the default) whenever an existing dest must survive a failed attempt. False streams straight into dest, which is opened "wb" and therefore truncated up front: a mid-stream failure leaves dest short, and any previous contents are gone. The failure path does not additionally delete it, but that is damage limitation, not a guarantee — atomic=False is only appropriate when dest is known to be disposable. The length check below runs in both modes; under atomic=False a failed check leaves the caller-owned dest short, because there is no staging file to discard.

Two calls downloading the same dest concurrently share one <dest>.part and will interleave into it. Serialising that is the caller's job.

True
expect_magic bytes | tuple[bytes, ...] | None

One or more byte prefixes the body must start with (e.g. b"CDF" / b"\x89HDF" for NetCDF). A body that starts with none of them raises ValueError and the partial write is discarded, so an HTML error page served with a 200 status never lands as a .nc. None (the default) skips the check.

None
headers dict[str, str] | None

Per-request headers merged over the client defaults.

None
timeout Timeout | None

Per-request timeout override (seconds), as a single float or a (connect, read) pair — the latter fails a dead host on the short connect budget without shortening the long read budget a large download needs.

None
resume bool

Continue a broken transfer with a ranged request instead of re-reading the object, when — and only when — the server has proved it can be done safely. Off by default; the whole-object path above is unchanged for every caller that does not opt in. It requires verify_length=True, the default — the pair resume=True, verify_length=False raises ValueError.

Three things make it inert, silently, so they are listed together. It requires staging, so it does nothing under atomic=False unless expect_magic forces staging back on: resuming means keeping a partial file between attempts, and that is only safe in the private <dest>.part this method owns, never in a caller-supplied dest that a failed attempt must not damage. It stands down when a Range is already present in the merged headers — including one set on the client, which disables resume for every call on it — because a caller who owns the object boundaries should not have ours added on top. And a leg only arms once more than _RESUME_OVERLAP bytes have been banked since the previous leg started, because a leg that gains less than the overlap it re-fetches costs more than it saves.

On a leg download also overrides two headers rather than merging under them: Accept-Encoding: identity, because a coded reply would make the byte offsets mean something else, and If-Range, which must name the representation the transfer anchored on. This is the one place caller headers are not honoured verbatim. A leg is likewise exempt from status_forcelist and from any retry_predicate: both are written against whole-object responses, and neither can judge — or even parse — a 206 fragment of one. The refusal gates below are the leg's own check in their place.

A resumed request is only attempted when the first response was a 200 that advertised Accept-Ranges: bytes, a known Content-Length, no content coding, and a strong ETag (never Last-Modified, whose one-second resolution matches across a change made within the same second). It is only appended when the reply is a 206 whose Content-Range start, end and total all match what was asked for, whose ETag still names the anchored representation, and — the check none of the others can stand in for — whose first _RESUME_OVERLAP bytes re-send a window already on disk and match it byte for byte. Every other answer, a 200 with the whole body included, discards the staged bytes and re-reads the object; resume then stays off for the rest of the call.

Only bytes this call wrote are ever built on, so a .part left by an earlier run is never treated as a prefix. The assembled file goes through the same length and expect_magic gates as a whole-object read.

False
verify_length bool

Compare the bytes written against the Content-Length the response advertised, and raise IncompleteDownloadError on a mismatch. On by default.

Pass False for a server that misreports the size of a body it generates on the fly, where the check would fail a download that is actually fine. This distrusts the advertised length in both directions: an over-long body — a mis-framed response, a spliced proxy body — is published too, because a length you do not believe cannot catch one. It also disables the salvage, whose whole proof is that same size equality. And it cannot be combined with resume=True, which is addressed by the length it distrusts: the pair raises ValueError before anything is sent.

It is a separate switch from Accept-Encoding on purpose: a coded body already has no checkable length — the header counts encoded octets and the file counts decoded ones — so asking for gzip disables the check as a side effect, and that side effect should not be the only way to express the intent.

True
**kwargs Any

Extra keyword arguments forwarded to requests.

{}

Returns:

Name Type Description
Path Path

The dest path the bytes were written to.

Raises:

Type Description
ValueError

When expect_magic is given and the body does not start with any of the supplied prefixes, or when resume=True is combined with verify_length=False — the latter refused up front, before a parent directory is created or a request sent.

UnsolicitedPartialContentError

When a 206 answers a request that carried no Range — the body is a fragment of the object, and repeating the request returns the same fragment, so this is raised rather than retried whatever the client's retry policy.

IncompleteDownloadError

When the bytes written do not equal the Content-Length the response advertised. A short body is retried from the start; a long one, a repeat of the same byte count, a spent read or attempt budget, or a client that opted out of transport retry raises at once.

HTTPError

On an error status — download always calls raise_for_status (the client's raise_for_status flag governs the verb methods, not download; a file fetch never keeps an error body). Note the resulting HTTPError is itself subject to retry_on_exceptions: a client whose retry_on_exceptions includes a supertype of requests.HTTPError (e.g. requests.RequestException, as ghsl/glaciers pass) will retry an error status before raising it, mirroring their old download loops. This is also how a refused resume surfaces when the ranged reply carried an error status and the retry budget is spent: the status and .response reach the caller intact rather than being flattened into a transport error.

ConnectionError

When a resumed leg is refused, the retry budget is spent, and the ranged reply carried no error status to report in its place. A refusal with budget left never escapes at all: the staged bytes are discarded, the object is re-read whole, and resume stays off for the rest of the call.

BaseException

Once the transport retry budget is spent, the last transport failure is re-raised as a fresh instance of its own type, with a message naming the attempt count and the largest byte count any attempt reached.

OSError

From mkdir, the streaming write, stat or replace, unchanged. RangeReadError is never raised here.

Source code in libs/core/src/earthlens/base/http.py
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
def download(
    self,
    url: str,
    dest: str | Path,
    *,
    chunk: int = DEFAULT_CHUNK_SIZE,
    progress: bool = True,
    atomic: bool = True,
    expect_magic: bytes | tuple[bytes, ...] | None = None,
    headers: dict[str, str] | None = None,
    timeout: Timeout | None = None,
    resume: bool = False,
    verify_length: bool = True,
    **kwargs: Any,
) -> Path:
    """Stream `url` to `dest` atomically, optionally showing a `tqdm` bar.

    Absorbs the chunk-loop the streamed-download backends each
    re-implement: streams with `stream=True`, sizes a progress bar
    from `Content-Length` when present, and writes 1 MiB blocks. When
    `atomic` (the default) it writes to a sibling `<dest>.part` and
    renames on success, and it removes the temp on any failure — so a
    crashed or interrupted download never leaves a truncated `dest`.
    The whole download is retry-wrapped: a status in `status_forcelist`
    or an exception in `retry_on_exceptions` retries the attempt (after
    cleaning the temp), honouring the `Retry-After`/back-off policy.

    By default every attempt reads the object **once, whole**, and the body
    is verified after the fact: a response that advertised a
    `Content-Length` must deliver exactly that many bytes, and a short read
    is retried from the start. Pass `resume=True` to let a retry continue a
    broken transfer instead of re-reading it; see that argument.

    Args:
        url: Absolute request URL.
        dest: Output file path. Parent directories are created.
        chunk: Streaming block size in bytes (default 1 MiB).
        progress: Show a `tqdm` progress bar. `False` (or a
            non-interactive / test context) suppresses it.
        atomic: Write to `<dest>.part` then rename on success, cleaning up
            the temp on failure — so a crashed download never leaves a
            truncated `dest` and never removes an existing one. Keep this on
            (the default) whenever an existing `dest` must survive a failed
            attempt. `False` streams straight into `dest`, which is opened
            `"wb"` and therefore **truncated up front**: a mid-stream failure
            leaves `dest` short, and any previous contents are gone. The
            failure path does not additionally delete it, but that is damage
            limitation, not a guarantee — `atomic=False` is only appropriate
            when `dest` is known to be disposable. The length check below
            runs in both modes; under `atomic=False` a failed check leaves
            the caller-owned `dest` short, because there is no staging file
            to discard.

            Two calls downloading the same `dest` concurrently share one
            `<dest>.part` and will interleave into it. Serialising that is
            the caller's job.
        expect_magic: One or more byte prefixes the body must start with
            (e.g. `b"CDF"` / `b"\\x89HDF"` for NetCDF). A body that starts
            with none of them raises `ValueError` and the partial write is
            discarded, so an HTML error page served with a 200 status never
            lands as a `.nc`. `None` (the default) skips the check.
        headers: Per-request headers merged over the client defaults.
        timeout: Per-request timeout override (seconds), as a single
            float or a `(connect, read)` pair — the latter fails a dead
            host on the short connect budget without shortening the long
            read budget a large download needs.
        resume: Continue a broken transfer with a ranged request instead
            of re-reading the object, when — and only when — the server has
            proved it can be done safely. Off by default; the whole-object
            path above is unchanged for every caller that does not opt in.
            It requires `verify_length=True`, the default — the pair
            `resume=True, verify_length=False` raises `ValueError`.

            Three things make it inert, silently, so they are listed
            together. It requires staging, so it does nothing under
            `atomic=False` unless `expect_magic` forces staging back on:
            resuming means keeping a partial file between attempts, and
            that is only safe in the private `<dest>.part` this method
            owns, never in a caller-supplied `dest` that a failed attempt
            must not damage. It stands down when a `Range` is already
            present in the merged headers — including one set on the
            client, which disables resume for every call on it — because a
            caller who owns the object boundaries should not have ours
            added on top. And a leg only arms once more than
            `_RESUME_OVERLAP` bytes have been banked since the previous leg
            started, because a leg that gains less than the overlap it
            re-fetches costs more than it saves.

            On a leg `download` also overrides two headers rather than
            merging under them: `Accept-Encoding: identity`, because a
            coded reply would make the byte offsets mean something else,
            and `If-Range`, which must name the representation the transfer
            anchored on. This is the one place caller headers are not
            honoured verbatim. A leg is likewise exempt from
            `status_forcelist` and from any `retry_predicate`: both are
            written against whole-object responses, and neither can judge —
            or even parse — a `206` fragment of one. The refusal gates
            below are the leg's own check in their place.

            A resumed request is only *attempted* when the first response
            was a `200` that advertised `Accept-Ranges: bytes`, a known
            `Content-Length`, no content coding, and a **strong** `ETag`
            (never `Last-Modified`, whose one-second resolution matches
            across a change made within the same second). It is only
            *appended* when the reply is a `206` whose `Content-Range`
            start, end and total all match what was asked for, whose `ETag`
            still names the anchored representation, and — the check none of
            the others can stand in for — whose first `_RESUME_OVERLAP`
            bytes re-send a window already on disk and match it byte for
            byte. Every other answer, a `200` with the whole body included,
            discards the staged bytes and re-reads the object; resume then
            stays off for the rest of the call.

            Only bytes *this call* wrote are ever built on, so a `.part`
            left by an earlier run is never treated as a prefix. The
            assembled file goes through the same length and `expect_magic`
            gates as a whole-object read.
        verify_length: Compare the bytes written against the
            `Content-Length` the response advertised, and raise
            `IncompleteDownloadError` on a mismatch. On by default.

            Pass `False` for a server that misreports the size of a body it
            generates on the fly, where the check would fail a download that
            is actually fine. This distrusts the advertised length in
            **both** directions: an over-long body — a mis-framed response,
            a spliced proxy body — is published too, because a length you
            do not believe cannot catch one. It also disables the salvage,
            whose whole proof is that same size equality. And it cannot be
            combined with `resume=True`, which is addressed by the length it
            distrusts: the pair raises `ValueError` before anything is sent.

            It is a separate switch from `Accept-Encoding` on purpose: a
            coded body already has no checkable length — the header counts
            encoded octets and the file counts decoded ones — so asking for
            `gzip` disables the check as a side effect, and that side effect
            should not be the only way to express the intent.
        **kwargs: Extra keyword arguments forwarded to `requests`.

    Returns:
        Path: The `dest` path the bytes were written to.

    Raises:
        ValueError: When `expect_magic` is given and the body does not
            start with any of the supplied prefixes, or when `resume=True`
            is combined with `verify_length=False` — the latter refused up
            front, before a parent directory is created or a request sent.
        UnsolicitedPartialContentError: When a `206` answers a request that
            carried no `Range` — the body is a fragment of the object, and
            repeating the request returns the same fragment, so this is
            raised rather than retried whatever the client's retry policy.
        IncompleteDownloadError: When the bytes written do not equal the
            `Content-Length` the response advertised. A short body is
            retried from the start; a long one, a repeat of the same byte
            count, a spent read or attempt budget, or a client that opted
            out of transport retry raises at once.
        requests.HTTPError: On an error status — `download` always
            calls `raise_for_status` (the client's `raise_for_status`
            flag governs the verb methods, not `download`; a file
            fetch never keeps an error body). Note the resulting
            `HTTPError` is itself subject to `retry_on_exceptions`:
            a client whose `retry_on_exceptions` includes a supertype
            of `requests.HTTPError` (e.g. `requests.RequestException`,
            as ghsl/glaciers pass) will **retry** an error status
            before raising it, mirroring their old download loops.
            This is also how a refused resume surfaces when the ranged
            reply carried an error status and the retry budget is spent:
            the status and `.response` reach the caller intact rather than
            being flattened into a transport error.
        requests.ConnectionError: When a resumed leg is refused, the retry
            budget is spent, and the ranged reply carried no error status
            to report in its place. A refusal with budget left never
            escapes at all: the staged bytes are discarded, the object is
            re-read whole, and resume stays off for the rest of the call.
        BaseException: Once the transport retry budget is spent, the last
            transport failure is re-raised as a fresh instance of its own
            type, with a message naming the attempt count and the largest
            byte count any attempt reached.
        OSError: From `mkdir`, the streaming write, `stat` or `replace`,
            unchanged. `RangeReadError` is never raised here.
    """
    if resume and not verify_length:
        # Contradictory, not merely risky. A resumed leg's `Range` end and
        # its `Content-Range` check are both derived from the advertised
        # total, and the length post-condition is the only thing that
        # proves the assembly reached it — `_resume_refusal` validates the
        # window the server claimed, never the bytes it delivered. Asking
        # to resume is asserting the total is trustworthy; asking to skip
        # the check is asserting it is not. Refuse rather than silently
        # publish a correct-prefix-but-short file.
        raise ValueError(
            "resume=True requires verify_length=True: a resumed transfer is "
            "addressed by the advertised Content-Length, so the length check "
            "is the only proof the assembled file is complete."
        )
    dest = Path(dest)
    dest.parent.mkdir(parents=True, exist_ok=True)
    # `expect_magic` promises a rejected body is discarded, which is only
    # possible if `dest` has not been written yet — so a magic check forces
    # staging even when the caller passed `atomic=False`. Otherwise the
    # validation would run on a file that had already replaced a good one.
    staged = atomic or expect_magic is not None
    tmp = dest.with_name(dest.name + ".part") if staged else dest
    # Bytes of `tmp` that *this call* wrote and that a resumed request
    # may therefore build on. It is never seeded from a file found on
    # disk: a `.part` left by a killed run belongs to an unknown object,
    # and resuming onto it would publish a file assembled from two.
    banked = 0
    # Captured from the first response; frozen thereafter, so every leg
    # is checked against the representation the transfer started on.
    anchor: _ResumeAnchor | None = None
    # One strike. A server that answers a ranged request badly costs a
    # single extra attempt, never a restart/resume cycle.
    resume_off = not resume

    def discard_partial() -> None:
        """Remove the partial write, but never a caller-owned `dest`.

        When the write was staged (`atomic`, or a magic check forcing it)
        the partial lives at a private `<dest>.part`, so removing it is
        always safe. Otherwise the stream wrote straight to `dest`, which
        `_stream_to_file` has already truncated by opening it `"wb"` —
        deleting it as well would only turn a truncated file into a missing
        one, and would destroy a file this call never owned.

        An unlink that fails is logged and swallowed, never raised: this
        runs from every failure path, so a second error here would replace
        the real one on its way out.
        """
        nonlocal banked
        # Reset together: the count means "bytes of `tmp` this call
        # wrote", so it cannot outlive the file it describes.
        banked = 0
        if staged:
            # This runs from every failure path, including
            # `except BaseException`, and on Windows the unlink can raise
            # `PermissionError` when the handle has only just been released.
            # Removing the temp is a promise, but it must not become a
            # second failure that replaces the real one — so it is caught
            # and logged rather than suppressed silently, since the file it
            # promised to remove is still there.
            try:
                tmp.unlink(missing_ok=True)
            except OSError as exc:
                logger.debug(f"could not remove {tmp.name}: {exc}")

    merged = self._merge_headers(headers)
    # Header names are case-insensitive, and `_merge_headers` returns a
    # plain dict, so every test on them is folded explicitly.
    per_call = {key.lower() for key in (headers or {})}
    caller_sent_range = "range" in {key.lower() for key in merged}
    if "accept-encoding" not in per_call and not self._accept_encoding_is_explicit:
        # The length check below compares against the bytes as delivered, so
        # a transparently-encoded body would fail it. This sits *underneath*
        # the caller: a backend that asked for `gzip`, or for `identity` to
        # protect its own magic check, keeps what it asked for.
        merged["Accept-Encoding"] = "identity"
    effective_timeout = self.timeout if timeout is None else timeout
    attempt = 0
    # Per-kind accounting, as in `_request_with_retry`: a dead host should
    # not cost a large download six full connect attempts.
    spent = {"connect": 0, "read": 0}
    # The byte count the previous attempt reached. A second attempt landing
    # on the identical number is deterministic, not a transport blip.
    last_written: int | None = None
    best_written = 0
    # What the previous leg had banked, so per-leg progress is checkable.
    last_banked = 0
    if staged and tmp.exists():
        # Suppressed for the same reason as the salvage stat below: this is
        # a log line, and it runs before the first request, so a filesystem
        # blip here must not cost the whole download.
        with contextlib.suppress(OSError):
            logger.debug(
                f"{redact_url(url)}: discarding a pre-existing "
                f"{tmp.stat().st_size:,}-byte {tmp.name}"
            )

    def salvaged(expected: int | None) -> bool:
        """Whether the bytes on disk already are the whole object.

        A transfer can break after the last byte but before the stream
        frames its end. The equality is the entire proof, so this asks no
        question of the server and consults no retry policy. It also never
        raises: it is the first thing a transport handler runs, and an
        exception from inside a handler escapes past its own `try`.

        Args:
            expected: The length the response advertised, if any.

        Returns:
            bool: True when the staged file is exactly that long. False
            when the check is switched off, when the response advertised no
            length, when nothing is staged, and when the size cannot be
            read at all — an unprovable salvage is not a salvage, so the
            ordinary retry path takes over.
        """
        if not verify_length or expected is None or not tmp.exists():
            return False
        try:
            return tmp.stat().st_size == expected
        except OSError:
            # This runs as the FIRST statement of both transport `except`
            # handlers, and an exception raised inside a handler is not
            # caught by its own `try` — so an unguarded failure here would
            # leave `download` with the filesystem's errno instead of the
            # transport error, and with none of the retry budget spent.
            # Unprovable is not salvageable: fall through to the ordinary
            # retry path.
            return False

    def salvage_publishable(expected: int | None) -> bool:
        """Whether the staged bytes are the whole object *and* pass every gate.

        The salvage runs `expect_magic` itself rather than falling through to
        the publish. Without it a body that breaks after its last byte
        reaches `dest` through a weaker door than one that does not, so a
        200-served HTML error page whose `Content-Length` happens to match
        would land under a `.nc` name with the magic check never run.

        Args:
            expected: The length the response advertised, if any.

        Returns:
            bool: True when the file may be published as it stands.

        Raises:
            ValueError: When `expect_magic` is set and the staged body does
                not start with one of the prefixes. The partial is discarded
                first, exactly as on the whole-object path.
        """
        if not salvaged(expected):
            return False
        if expect_magic is not None:
            try:
                _check_magic(tmp, expect_magic, url)
            except BaseException:
                discard_partial()
                raise
        return True

    while True:
        self._throttle()
        response: requests.Response | None = None
        expected_total: int | None = None
        status_wait: float | None = None
        # A resumed leg re-fetches `_RESUME_OVERLAP` bytes it already
        # holds, so `banked` must exceed the overlap for the leg to make
        # forward progress at all.
        resume_at = 0
        leg: _ResumeAnchor | None = None
        if (
            not resume_off
            and staged
            and anchor is not None
            and not caller_sent_range
            and _RESUME_OVERLAP < banked < anchor.total
            # Each leg re-fetches the overlap, so a leg that advanced less
            # than that cost more than it saved. Without this a server
            # dribbling a byte per leg burns the whole read budget
            # re-sending 64 KiB apiece, where a plain restart is cheaper.
            and banked - last_banked > _RESUME_OVERLAP
        ):
            resume_at, leg = banked, anchor
            last_banked = banked
        attempt_headers = merged
        if leg is not None:
            attempt_headers = {
                **merged,
                "Range": f"bytes={resume_at - _RESUME_OVERLAP}-{leg.total - 1}",
                "If-Range": leg.etag,
                # `Range` addresses the *encoded* representation while
                # the staged count is of decoded bytes, so a coded reply
                # would make the offsets mean different things.
                "Accept-Encoding": "identity",
            }
        try:
            # Inside the retrying `try`: a failure here is a connect-phase
            # failure and must be retried like any other, which is why the
            # send and the body processing share one handler.
            response = self._send(
                "GET",
                url,
                stream=True,
                headers=attempt_headers,
                timeout=effective_timeout,
                **kwargs,
            )
            try:
                # A resumed leg is exempt: `status_forcelist` and a
                # `retry_predicate` are written against whole-object
                # responses (the existing ones call `r.json()`), and handing
                # one a `206` byte fragment would have it judge — or fail to
                # parse — a slice of the object. The leg has its own gate.
                retryable = leg is None and (
                    response.status_code in self.status_forcelist
                    or (
                        self._retry_predicate is not None
                        and self._retry_predicate(response)
                    )
                )
                if retryable and attempt < self.max_retries:
                    # Recorded, not slept on: the sleep happens after the
                    # `finally` below has released the streamed connection,
                    # rather than holding a socket open across the back-off.
                    status_wait = self._backoff_wait(
                        _parse_retry_after(response.headers.get("Retry-After")),
                        attempt,
                    )
                    logger.warning(
                        f"HTTP {response.status_code} on {redact_url(url)}; "
                        f"retry {attempt + 1}/{self.max_retries} after "
                        f"{status_wait:.1f}s"
                    )
                else:
                    if leg is not None:
                        refusal = _resume_refusal(response, leg, resume_at)
                        if refusal is not None:
                            raise _ResumeRefused(refusal)
                        expected_total = leg.total
                        self._append_resumed_body(
                            response,
                            tmp,
                            overlap_at=resume_at - _RESUME_OVERLAP,
                            overlap=_RESUME_OVERLAP,
                            total=leg.total,
                            chunk=chunk,
                            progress=progress,
                            desc=dest.name,
                        )
                    else:
                        if response.status_code == 206 and not caller_sent_range:
                            raise UnsolicitedPartialContentError(
                                f"{redact_url(url)} answered 206 to a request "
                                "carrying no Range; the body is a fragment",
                                response=response,
                            )
                        response.raise_for_status()
                        expected_total = _progress_total(response.headers)
                        if not resume_off and anchor is None and staged:
                            # Recorded before the body is read, so a break
                            # part-way through still leaves an anchor to
                            # resume against.
                            anchor = _arm_resume_anchor(response, expected_total)
                        self._stream_to_file(
                            response,
                            tmp,
                            chunk=chunk,
                            progress=progress,
                            desc=dest.name,
                        )
                    # One verification for both paths: a resumed file is
                    # published through the same size and magic gates as a
                    # whole-object read, never a weaker door.
                    written = tmp.stat().st_size
                    best_written = max(best_written, written)
                    if (
                        verify_length
                        and expected_total is not None
                        and written != expected_total
                    ):
                        raise IncompleteDownloadError(
                            f"{redact_url(url)} delivered {written:,} of "
                            f"{expected_total:,} advertised bytes",
                            written=written,
                            expected=expected_total,
                        )
                    if expect_magic is not None:
                        _check_magic(tmp, expect_magic, url)
            finally:
                if response is not None:
                    response.close()
        except _ResumeRefused as exc:
            # Never re-armed for the rest of the call, so a server that
            # answers a ranged request badly costs exactly one extra
            # attempt rather than a restart/resume cycle.
            resume_off = True
            logger.warning(
                f"{redact_url(url)}: refusing to resume ({exc}); discarding "
                f"{banked:,} staged bytes and re-reading the whole object"
            )
            discard_partial()
            if attempt >= self.max_retries:
                # A backstop, not the primary bound: `resume_off` already
                # caps refusals at one per call. Without it the loop's
                # termination would rest on that single flag, and a refusal
                # would otherwise re-read the object with no budget left.
                if response is not None and response.status_code >= 400:
                    # Report what the server actually said. Backends branch
                    # on `exc.response.status_code` to tell a missing
                    # granule (404/410) from a real failure, and a
                    # synthesised `ConnectionError` would both drop the
                    # status and reclassify a permanent error as transport.
                    response.raise_for_status()
                raise requests.ConnectionError(
                    f"{redact_url(url)}: resume refused ({exc}) with no "
                    f"retry budget left after {attempt + 1} attempts"
                ) from exc
            attempt += 1
            continue
        except UnsolicitedPartialContentError:
            # Deterministic: a server that volunteers partial content to a
            # Range-less request returns the same fragment next time.
            discard_partial()
            raise
        except IncompleteDownloadError as exc:
            over_delivered = (
                exc.written is not None
                and exc.expected is not None
                and exc.written > exc.expected
            )
            if (
                over_delivered
                or not self.retry_on_exceptions
                or exc.written == last_written
                or spent["read"] >= self.read_retries
                or attempt >= self.max_retries
            ):
                discard_partial()
                raise
            spent["read"] += 1
            last_written = exc.written
            logger.warning(
                f"{redact_url(url)} delivered {exc.written:,} of "
                f"{exc.expected:,} bytes; discarding and restarting, read "
                f"retry {spent['read']}/{self.read_retries} "
                f"(attempt {attempt + 1}/{self.max_retries})"
            )
            discard_partial()
            self._sleep(self._backoff_wait(None, attempt))
            attempt += 1
            continue
        except self.retry_on_exceptions as exc:
            if salvage_publishable(expected_total):
                logger.debug(
                    f"{redact_url(url)} broke after the last byte; the "
                    f"{expected_total:,}-byte body is already complete"
                )
            else:
                kind = classify_transport_error(exc, strict=self._default_retry_set)
                if kind is None:
                    # Deterministic - the next attempt reproduces it.
                    discard_partial()
                    raise
                # `response is None` is exactly "the failure happened before
                # a response object existed", which is the connect phase.
                if kind == "connect" or (kind == "unknown" and response is None):
                    key, budget = "connect", self.connect_retries
                else:
                    key, budget = "read", self.read_retries
                if spent[key] >= budget or attempt >= self.max_retries:
                    discard_partial()
                    raise type(exc)(
                        f"{redact_url(url)} failed after {attempt + 1} "
                        f"attempts; best read {best_written:,} bytes"
                    ) from exc
                spent[key] += 1
                try:
                    staged_now = tmp.stat().st_size if tmp.exists() else 0
                except OSError:
                    staged_now = 0
                best_written = max(best_written, staged_now)
                kept = 0
                if not resume_off and staged and anchor is not None:
                    # Bank only what THIS attempt wrote: `_stream_to_file`
                    # truncated `tmp` before writing, so whatever is there
                    # now came from this call and from this representation.
                    size = staged_now
                    if _RESUME_OVERLAP < size < anchor.total:
                        kept = size
                banked = kept
                logger.warning(
                    f"{type(exc).__name__} ({kind}) on {redact_url(url)} after "
                    # The size on disk right now, never `best_written`: that
                    # is a high-water mark across attempts and outlives the
                    # file that produced it, so banking it would point a
                    # resumed request past the end of the staged file.
                    f"{staged_now:,} bytes; "
                    + (
                        f"keeping {kept:,} staged bytes to resume from, "
                        if kept
                        else "discarding and restarting, "
                    )
                    + f"{key} retry {spent[key]}/{budget} "
                    f"(attempt {attempt + 1}/{self.max_retries})"
                )
                if not kept:
                    discard_partial()
                self._sleep(self._backoff_wait(None, attempt))
                attempt += 1
                continue
        except requests.RequestException:
            # Wider than the client's retry set on purpose: the salvage
            # issues no request and reads no header, so the caller's retry
            # policy is not the right gate - the size equality is the proof.
            if not salvage_publishable(expected_total):
                discard_partial()
                raise
            logger.debug(
                f"{redact_url(url)} broke after the last byte; the "
                f"{expected_total:,}-byte body is already complete"
            )
        except BaseException:
            discard_partial()
            raise
        if status_wait is not None:
            # The response is closed by now, so nothing is held open.
            self._sleep(status_wait)
            attempt += 1
            continue
        if staged:
            # Guard the rename too, so the "removes the temp on any
            # failure" promise holds if the final replace fails.
            try:
                tmp.replace(dest)
            except BaseException:
                discard_partial()
                raise
        return dest

get(url, **kwargs) #

Send a GET request. See :meth:request for arguments.

Source code in libs/core/src/earthlens/base/http.py
def get(self, url: str, **kwargs: Any) -> requests.Response:
    """Send a `GET` request. See :meth:`request` for arguments."""
    return self.request("GET", url, **kwargs)

get_json(url, **kwargs) #

Send a GET request and decode the JSON response body.

Convenience over :meth:get for the REST endpoints that return JSON envelopes.

Parameters:

Name Type Description Default
url str

Absolute request URL.

required
**kwargs Any

Keyword arguments forwarded to :meth:get (params, headers, timeout, ...).

{}

Returns:

Type Description
Any

The parsed JSON body (typically a dict or list).

Raises:

Type Description
HTTPError

On a non-retryable error status, or after the retryable status is exhausted.

Source code in libs/core/src/earthlens/base/http.py
def get_json(self, url: str, **kwargs: Any) -> Any:
    """Send a `GET` request and decode the JSON response body.

    Convenience over :meth:`get` for the REST endpoints that return
    JSON envelopes.

    Args:
        url: Absolute request URL.
        **kwargs: Keyword arguments forwarded to :meth:`get`
            (`params`, `headers`, `timeout`, ...).

    Returns:
        The parsed JSON body (typically a `dict` or `list`).

    Raises:
        requests.HTTPError: On a non-retryable error status, or after
            the retryable status is exhausted.
    """
    return self.get(url, **kwargs).json()

post(url, **kwargs) #

Send a POST request. See :meth:request for arguments.

Source code in libs/core/src/earthlens/base/http.py
def post(self, url: str, **kwargs: Any) -> requests.Response:
    """Send a `POST` request. See :meth:`request` for arguments."""
    return self.request("POST", url, **kwargs)

request(method, url, *, headers=None, timeout=None, raise_for_status=None, **kwargs) #

Send one request with the default headers, timeout, and retry.

Parameters:

Name Type Description Default
method str

HTTP verb ("GET", "POST", ...).

required
url str

Absolute request URL.

required
headers dict[str, str] | None

Per-request headers merged over the client defaults, and passed through verbatim. download adds only one header of its own, Accept-Encoding: identity, and only when neither this argument nor the constructor's headers= named it in any casing — the length check compares against bytes as delivered, which a transparently-encoded body would not match. Pass {"Accept-Encoding": "gzip"} to override that.

A Range here is honoured as written: the response is accepted at its own Content-Length, and download does not check that the server returned the range you asked for.

None
timeout Timeout | None

Per-request timeout override (seconds), as a single float or a (connect, read) pair. Defaults to the client's timeout.

None
raise_for_status bool | None

Per-request override of the client's raise_for_status policy. None (default) uses the client setting.

None
**kwargs Any

Extra keyword arguments forwarded to requests (params, data, json, stream, ...).

{}

Returns:

Type Description
Response

requests.Response: The response (after raise_for_status unless it is disabled).

Raises:

Type Description
HTTPError

On a non-retryable error status, or after the retryable status is exhausted (when raise_for_status is on).

Source code in libs/core/src/earthlens/base/http.py
def request(
    self,
    method: str,
    url: str,
    *,
    headers: dict[str, str] | None = None,
    timeout: Timeout | None = None,
    raise_for_status: bool | None = None,
    **kwargs: Any,
) -> requests.Response:
    """Send one request with the default headers, timeout, and retry.

    Args:
        method: HTTP verb (`"GET"`, `"POST"`, ...).
        url: Absolute request URL.
        headers: Per-request headers merged over the client defaults, and
            passed through verbatim. `download` adds only one header of its
            own, `Accept-Encoding: identity`, and only when neither this
            argument nor the constructor's `headers=` named it in any
            casing — the length check compares against bytes as delivered,
            which a transparently-encoded body would not match. Pass
            `{"Accept-Encoding": "gzip"}` to override that.

            A `Range` here is honoured as written: the response is accepted
            at its own `Content-Length`, and `download` does **not** check
            that the server returned the range you asked for.
        timeout: Per-request timeout override (seconds), as a single
            float or a `(connect, read)` pair. Defaults to the client's
            `timeout`.
        raise_for_status: Per-request override of the client's
            `raise_for_status` policy. `None` (default) uses the
            client setting.
        **kwargs: Extra keyword arguments forwarded to `requests`
            (`params`, `data`, `json`, `stream`, ...).

    Returns:
        requests.Response: The response (after `raise_for_status`
            unless it is disabled).

    Raises:
        requests.HTTPError: On a non-retryable error status, or after
            the retryable status is exhausted (when `raise_for_status`
            is on).
    """
    merged = self._merge_headers(headers)
    effective_timeout = self.timeout if timeout is None else timeout
    return self._request_with_retry(
        method,
        url,
        headers=merged,
        timeout=effective_timeout,
        raise_for_status=raise_for_status,
        **kwargs,
    )

stream(url, **kwargs) #

Send a streaming GET (stream=True), retry-wrapped.

Returns the open response without consuming its body, so the caller can iterate iter_content. Retries follow the same Retry-After/back-off policy as the other verbs; the retry decision reads only the status line, never the body.

A body that breaks after this returns is therefore the caller's to handle: the iteration happens outside the retry loop, so a ChunkedEncodingError raised mid-stream escapes it. The default transport retry covers the non-streaming verbs, whose bodies requests materialises inside the loop, and :meth:download, which owns its own. A caller streaming a large body that needs the same protection should use :meth:download or re-request on failure itself.

Parameters:

Name Type Description Default
url str

Absolute request URL.

required
**kwargs Any

Keyword arguments forwarded to :meth:get.

{}

Returns:

Type Description
Response

requests.Response: The open streaming response.

Source code in libs/core/src/earthlens/base/http.py
def stream(self, url: str, **kwargs: Any) -> requests.Response:
    """Send a streaming `GET` (`stream=True`), retry-wrapped.

    Returns the open response without consuming its body, so the
    caller can iterate `iter_content`. Retries follow the same
    `Retry-After`/back-off policy as the other verbs; the retry
    decision reads only the status line, never the body.

    A body that breaks **after** this returns is therefore the
    caller's to handle: the iteration happens outside the retry loop,
    so a `ChunkedEncodingError` raised mid-stream escapes it. The
    default transport retry covers the non-streaming verbs, whose
    bodies `requests` materialises inside the loop, and
    :meth:`download`, which owns its own. A caller streaming a large
    body that needs the same protection should use :meth:`download`
    or re-request on failure itself.

    Args:
        url: Absolute request URL.
        **kwargs: Keyword arguments forwarded to :meth:`get`.

    Returns:
        requests.Response: The open streaming response.
    """
    return self.get(url, stream=True, **kwargs)