Summary
SparqlConstructReader retries a 502/503/504 three times with p-retry’s defaults: 1 s, 2 s, 4 s. An endpoint whose backend is restarting stays 502/503 for far longer than those 7 s, so all four attempts hit the same outage and the stage fails anyway.
Observed
Limburgs Museum on LDMax (QLever behind nginx), 2026‑09‑08. After the backend went down under the stage’s concurrent queries, nginx answered 502 and then 503 for between 30 s and several minutes; one recovery took over five minutes. The pipeline log for that run (limburg/lol#176) shows every stage after the first failing on 503 within seconds of each other – the retries were spent while the backend was still coming up.
Proposed
- Backoff that outlasts a restart: something like 5 s, 15 s, 45 s (a
maxTimeout in the minute range), or make retries/minTimeout configurable on the reader and expose them in @lde/search-indexer.
- A retry after a 502/503 should first wait for the endpoint to answer a trivial
ASK again, rather than resending the expensive CONSTRUCT into the outage.
Retrying on more status codes (#504) is a separate question; this one is about how long the retries wait. It matters less once stage concurrency is bounded per host (#838), since the backend then rarely goes down in the first place.
Summary
SparqlConstructReaderretries a 502/503/504 three times with p-retry’s defaults: 1 s, 2 s, 4 s. An endpoint whose backend is restarting stays 502/503 for far longer than those 7 s, so all four attempts hit the same outage and the stage fails anyway.Observed
Limburgs Museum on LDMax (QLever behind nginx), 2026‑09‑08. After the backend went down under the stage’s concurrent queries, nginx answered 502 and then 503 for between 30 s and several minutes; one recovery took over five minutes. The pipeline log for that run (limburg/lol#176) shows every stage after the first failing on 503 within seconds of each other – the retries were spent while the backend was still coming up.
Proposed
maxTimeoutin the minute range), or makeretries/minTimeoutconfigurable on the reader and expose them in@lde/search-indexer.ASKagain, rather than resending the expensive CONSTRUCT into the outage.Retrying on more status codes (#504) is a separate question; this one is about how long the retries wait. It matters less once stage concurrency is bounded per host (#838), since the backend then rarely goes down in the first place.