Skip to content

Reactive dump fallback imports the dump but never re-runs the stages, and the error is swallowed #804

Description

@ddeboer

Summary

When the reactive dump fallback (#445) kicks in, the dump is imported and reported as selected, but the stages are not re-run and nothing in the output says why. The dataset ends the run with no documents, is recorded failed under the dump’s fingerprint, and is then skipped as “Unchanged since last run” on every following run – a transient failure in the fallback path freezes the dataset out of the index until the dump changes or the pipeline version rotates.

The immediate defect is that the error is swallowed: ConsoleReporter.stageFailed only prints when a spinner is active, and the reactive-fallback catch block fires at a moment when there is none.

Observed

Linked Open Limburg consumer stack, codeberg.org/limburg/lol/search-indexer:latest on 2026‑09‑03 – @lde/search-indexer 0.16.0, @lde/pipeline 0.36.4, @lde/pipeline-console-reporter 0.27.4, @lde/search-typesense 0.29.2, IMPORT_STRATEGY=sparqlWithImportFallback, REBUILD_MODE=in-place, PROVENANCE_FILE set. Limburgs Museum’s LDMax endpoint passed probing but answered 502/503 to the stage queries.

Three consecutive runs produced this (abridged):

Dataset https://n2t.net/ark:/61567/dataset [1/3]
  ✔ SPARQL endpoint https://sparql.ldmax.nl/Q2798268 (HTTP 200)
  ✔ Data dump https://www.ldmax.nl/datasets/Q2798268/collectie/resources/collectie.nq.gz (HTTP 200)
  ✔ SPARQL endpoint https://sparql.ldmax.nl/Q2798268 (HTTP 200) (selected)
   - Stage CreativeWork
   ✖ Stage CreativeWork Invalid SPARQL endpoint response from https://sparql.ldmax.nl/Q2798268 (HTTP status 502):
   ✖ Stage Person Invalid SPARQL endpoint response from https://sparql.ldmax.nl/Q2798268 (HTTP status 503):
   … (Organization, Place, Term, Occupation likewise)
   ✔ Stage Dataset 1 items, 1 quads, took 163ms
   ✔ Stage Publisher 1 items, 1 quads, took 122ms
   - Importing…
   ✔ Imported https://www.ldmax.nl/datasets/Q2798268/collectie/resources/collectie.nq.gz (23.6K triples, to http://qlever-lol-search-indexer:7001/sparql) in 1.1s
  ✔ Completed in 32.5s (memory: 136 MB RSS, 54 MB heap)

Pipeline completed in 33.1s (memory: 140 MB RSS, 56 MB heap)

After “Imported” no stage starts, yet the dataset reports “Completed” and the process exits 0. The provenance record afterwards:

"https://n2t.net/ark:/61567/dataset": {
  "sourceFingerprint": "2026-07-31T23:48:03.348Z|386385",
  "pipelineVersion": "10",
  "status": "failed"
}

The fingerprint is the dump’s (the endpoint’s would be null), so the next run prints ✖ Skipped: Unchanged since last run and the dataset stays absent. Deleting the provenance record is the only way out.

A fourth run on the identical stack did take the intended path – “Imported” followed by every stage re-run against QLever, 388 works – so whatever fails is intermittent, and that is exactly why the error needs to be visible.

Where it goes wrong

pipeline.js, reactive-fallback block:

this.reportSelectedDistribution(dataset, fallback, fingerprint);
await runWriter.reset?.(dataset);
stageFailed = await this.runStages(dataset, fallback.distribution, timeout, runWriter, context);
…
} catch (error) {
  this.reporter?.stageFailed?.('reactive-dump-fallback', error …);
}

The “Imported” line is reportSelectedDistribution, and the first stage would print a spinner immediately, so the only thing that can throw between the two without a trace is runWriter.reset – for in-place that is documents.dropWrittenThisRun(dataset), a delete-by-filter (or a retract on a keyed collection) per collection, wrapped in settleAll.

Whatever is thrown then reaches ConsoleReporter.stageFailed:

stageFailed(_stage, error) {
    if (this.activeSpinner) {
        this.activeSpinner.suffixText = chalk.red(error.message);
        this.activeSpinner.fail();
        this.activeSpinner = undefined;
    }
}

distributionSelected has just called succeed() on the import spinner and set activeSpinner = undefined, so the message is dropped. The same happens for the 'write' failure the pipeline reports when runWriter.flush throws.

stageFailed stays true, so the outcome is recorded failed – but with fingerprint already reassigned to the dump’s, so shouldReprocess treats the next run as unchanged.

Proposed fix

  1. Reporter: stageFailed (and any other failure hook) must print when there is no active spinner, the way importFailed already does with printLine(logSymbols.error, …). A failure that leaves no trace in the log is worse than the failure itself.
  2. Fallback failure must not freeze the dataset out. When the fallback path throws, keep the endpoint’s fingerprint (null, always reprocess) rather than adopting the dump’s, or record the outcome in a way shouldReprocess retries. As it stands, a transient failure during reset or the re-run is retried only on a dump change or a PIPELINE_VERSION rotation.
  3. Investigate the intermittent reset failure in @lde/search-typesense’s in-place dropWrittenThisRun once (1) makes it visible. I can share the full logs of the three runs.

Related

Activity

  1. added theissue type on Sep 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions