Summary
A Stage runs up to 10 extraction CONSTRUCTs against the dataset’s endpoint at once (maxConcurrency ?? 10, batchSize ?? 10), and nothing bounds that per host. @lde/host-limiter is wired into @lde/distribution-probe and @lde/resolver, but not into stage readers, and @lde/search-indexer exposes no setting for either value. A small public endpoint gets ten concurrent queries per stage, and the one behind sparql.ldmax.nl dies of it.
Observed
Linked Open Limburg, @lde/search-indexer 0.16.0, @lde/pipeline 0.36.4, Limburgs Museum (https://sparql.ldmax.nl/Q2798268, QLever, ~34.5K triples). Replaying the stage’s CreativeWork CONSTRUCT with the real VALUES batches of 10 roots on 2026‑09‑08:
| Concurrent requests |
Result |
| 1 |
all 56 batches HTTP 200, ~1.7 s each |
| 2 |
~4 s per batch, backend dead after 8 batches |
| 10 (the default) |
backend dead after 7.6 s, every request HTTP 502 |
Once the backend is down, nginx answers 502 and then 503 for 30 s to several minutes (five outages observed on 2026‑09‑08). The reader retries 502/503 three times with 1, 2 and 4 s backoff (#840), so every stage of the run fails on the corpse the first one left, and the reactive dump fallback fires (#804). limburg/lol#176 is the user-facing report: the public endpoint “seems just fine”, because the UI sends one query at a time.
A QLever with LDE’s own defaults (--memory-max-size 4G) survived the same query at 10 concurrent requests locally, on the same dump, at 4.4 s each, so the crash is also a memory setting on the host’s side. That does not change the requirement: a pipeline must not send a publisher’s endpoint more load than a person would.
Proposed
- Bound in-flight stage queries per host, reusing
mapHostLimited from @lde/host-limiter the way the probe and the resolver already do. The per-host limit should default low (1 or 2) for endpoints the pipeline does not own, and can stay at 10 for the QLever the pipeline spawns itself.
- Expose
batchSize and maxConcurrency (or the per-host limit) in @lde/search-indexer’s configuration, so a deployment can tune them without composing searchStages by hand.
Related
Summary
A
Stageruns up to 10 extraction CONSTRUCTs against the dataset’s endpoint at once (maxConcurrency ?? 10,batchSize ?? 10), and nothing bounds that per host.@lde/host-limiteris wired into@lde/distribution-probeand@lde/resolver, but not into stage readers, and@lde/search-indexerexposes no setting for either value. A small public endpoint gets ten concurrent queries per stage, and the one behindsparql.ldmax.nldies of it.Observed
Linked Open Limburg,
@lde/search-indexer0.16.0,@lde/pipeline0.36.4, Limburgs Museum (https://sparql.ldmax.nl/Q2798268, QLever, ~34.5K triples). Replaying the stage’sCreativeWorkCONSTRUCT with the real VALUES batches of 10 roots on 2026‑09‑08:Once the backend is down, nginx answers 502 and then 503 for 30 s to several minutes (five outages observed on 2026‑09‑08). The reader retries 502/503 three times with 1, 2 and 4 s backoff (#840), so every stage of the run fails on the corpse the first one left, and the reactive dump fallback fires (#804). limburg/lol#176 is the user-facing report: the public endpoint “seems just fine”, because the UI sends one query at a time.
A QLever with LDE’s own defaults (
--memory-max-size 4G) survived the same query at 10 concurrent requests locally, on the same dump, at 4.4 s each, so the crash is also a memory setting on the host’s side. That does not change the requirement: a pipeline must not send a publisher’s endpoint more load than a person would.Proposed
mapHostLimitedfrom@lde/host-limiterthe way the probe and the resolver already do. The per-host limit should default low (1 or 2) for endpoints the pipeline does not own, and can stay at 10 for the QLever the pipeline spawns itself.batchSizeandmaxConcurrency(or the per-host limit) in@lde/search-indexer’s configuration, so a deployment can tune them without composingsearchStagesby hand.Related