Skip to content

Port GeoNames RDF to LDE #782

Description

@ddeboer

geonames-rdf converts the GeoNames data dumps to a ~13.5M-feature geonames.nt with a set of shell scripts (download.sh, map.sh, test.sh) driving the SPARQL Anything CLI. We want to port it to LDE. #511 is the first piece: @lde/sparql-anything with a SparqlAnythingConverter.

This issue tracks what has to change before the migration can start, split into musts (the port cannot run correctly or finish without them) and shoulds (worth fixing, but a migration could ship around them).

Status

Every must and should is done. @lde/sparql-anything is at 0.1.0.

Four of them landed differently from the description below, and two follow-ups came out of the work; see the status comment and the one after it.

Still open: publication, and keeping test.sh's golden-file diff – see Scope beyond the converter. The LDE side of the conversion itself is complete: all four of map.sh's invocation shapes are expressible against the current API, and @lde/sparql-anything is at 0.1.2.

What map.sh actually does

It makes four different SPARQL Anything invocations, not one:

# invocation query --load {SOURCE}
1 admin codes admin-codes.rq – – (paths inline in the query)
2 places chunks places.rq admin-codes.ttl yes
3 alternate-names chunks alternate-names.rq – yes
4 ontology ontology.rq ontology.rdf (as the input) –

SparqlAnythingConverter as first proposed expressed only #2.

Musts

  • 1. Parallelism. convert() is a sequential for loop with an await inside. map.sh runs an xargs -P worker pool sized from nproc, capped by the cgroup memory limit (or /proc/meminfo on a host), budgeting ~3 GB per worker. On allCountries that is ~14 places chunks plus ~4 alternate-names chunks; running them one at a time multiplies the conversion, which takes ~17 min of the ~40 min workflow (measured, see below).

    Needs a concurrency option; the memory-aware default is generic enough to belong in LDE rather than in the consumer. This is the one that decides whether the port is viable – everything else is correctness. Do it after 6+7, so the pool is not reworked when its unit of work changes shape.

  • 2. JVM heap control, and a CLI escape hatch. map.sh passes -Xmx2g per worker precisely because SPARQL Anything materialises a chunk’s whole result graph before writing it. The converter hardcodes java -jar … with no -Xmx, so each JVM takes a quarter of host RAM by default – which over-subscribes the moment concurrency lands. Also needs a passthrough for flags the converter does not model, such as -ad on SPARQL Anything 1.2-dev for #624/#625. In review: feat(sparql-anything): cap the JVM heap, and pass through CLI arguments #789 (javaOptions, cliArgs).

  • 3. A generic, optional load. adminCodesFile was required, single, and named after a GeoNames concept; it could not express invocation test: collect coverage #3 (no --load) or feat: add dataset-registry-client #4 (--load is the input). Now load?: string. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)

  • 4. Guard against an empty chunk output. SPARQL Anything exits 0 when it cannot read or parse an input: it logs, writes an empty --output and stops, which is how we once shipped a geonames.nt missing every ontology triple while the run stayed green. assertNonEmpty now fails the conversion, and an empty chunkPaths is rejected outright. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)

  • 5. convert() resolving before the output was flushed. pipeline(…, {end: false}) resolves when the source ends; output.end() ran un-awaited in finally, so a caller could read a truncated file. Now await finished(output), plus a newline between concatenated files so two triples can never share a line. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)

Shoulds

  • 6+7. One call over jobs, not one converter per query. Originally two separate points – a single worker pool, and returning the chunk outputs; 8’s workDir change merged them and made them more pressing rather than less.

    map.sh deliberately feeds both chunk sets through a single pool: an alternate-names chunk is roughly a quarter of the work of a places chunk (one triple per row, no admin-codes join), so the long jobs are listed first and the short ones fill the tail. One SparqlAnythingConverter per query means two sequential drains, which loses that packing.

    And the port has to get places + alternate names + ontology.nt into one file. convert() returns void, and now that the per-chunk outputs live in a run directory deleted in finally, the caller can no longer concatenate them itself – which it could when they sat next to the inputs.

    Both want the same change: convert() takes jobs of {chunk, query, load} in a single call, or gains an explicit append. Settle this before 1.

  • 8. Chunk outputs beside the inputs, never cleaned. Outputs and generated queries now go to a fresh per-run directory under a caller-supplied workDir, removed in finally, and are referred to by runner-relative paths so the same command works on the host and under DockerTaskRunner. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)

  • 9. Shell interpolation, and the wait() race. Interpolated values go through shellQuote (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511). wait() attached its close listener only after run() returned, so a process that failed fast could close first and wait() would never settle – invisible while the converter was sequential, fatal under a pool. Fixed in @lde/task-runner-native (fix(task-runner-native): settle wait() for a process that already closed #787).

Scope beyond the converter

Conversion is roughly a third of the migration. Still unhomed:

  • Chunking. download.sh’s chunk_with_header splits a TSV into CHUNK_SIZE-row chunks, each prefixed with the header row the queries expect. Tracked in Split a large delimited file into row chunks with a repeated header #793, shipped in feat(sparql-anything): split a file into the chunks a conversion can hold #799 as chunk().
  • The ontology step. Converting the RDF/XML to N-Triples works today: it is a job with a load and no chunks, which is map.sh's fourth invocation exactly, and the non-empty guard applies to it like any other job. Fetching the pinned ontology_v3.3.rdf stays in geonames-rdf – the pinned version is listed below as domain policy.
  • Publication. gzip -9 plus zip -9, uploaded to DigitalOcean Spaces. Probably stays in the consumer’s CI. Worth noting it is now the workflow's second-biggest cost, at ~15 min against the conversion's ~17.
  • The consumption bridge, on the geonames-rdf side: how a shell-and-Docker repo invokes the TypeScript converter.

What should not move into LDE, because it is GeoNames domain policy rather than plumbing: the awk synthesis of the adm1/adm2 foreign keys, the filtering of alternate names belonging to out-of-scope features, and the pinned ontology version.

Finally, test.sh is a golden-file test that runs the real map.sh over fixtures and diffs the result against test/expected/geonames.nt. Whatever replaces it has to keep that end-to-end diff: reviewing that file line by line is the actual test.

Activity

  1. added theissue type on Aug 31, 2026
  2. ddeboer commented on Aug 31, 2026

    @ddeboer
    MemberAuthor

    Musts 3, 4 and 5 are folded into #511 and are green there:

    • 3 – adminCodesFile is now an optional load. One correction to this issue: it stays a single string, not string | string[]. SPARQL Anything's -l takes one path, and a file (loaded into the default graph) is not interchangeable with a directory (one named graph per RDF file), so an array would have been both unsupported by the CLI and semantically different from what it suggests.
    • 4 – each chunk's --output is asserted non-empty before it is concatenated.
    • 5 – convert() awaits the concatenated output being flushed and closed.

    Must 3 moved forward deliberately: the package is not published yet, so removing the domain-specific name is free now and a breaking bump later. It also prompted a general rule in AGENTS.md (#783) – a package's public API must stay domain-agnostic.

    Remaining here: musts 1–2 (concurrency, JVM heap and a CLI passthrough) and shoulds 6–9.

  3. ddeboer commented on Aug 31, 2026

    @ddeboer
    MemberAuthor

    Status: every must and should is done or in review

    # Where
    Must 1 Parallelism #795 (open)
    Must 2 JVM heap, CLI escape hatch #789
    Must 3 Generic, optional load #511
    Must 4 Guard against an empty chunk output #511
    Must 5 convert() resolving before the flush #511
    Should 6 One worker pool, not two phases #791
    Should 7 convert() returning the chunk outputs #791, by removing the need
    Should 8 Chunk outputs beside the inputs, never cleaned #511
    Should 9 wait() race under concurrency #787

    Once #795 merges, what is left in this issue is the Scope beyond #511 section: chunking, the ontology step, and publication.

    Where the plan changed

    Four items did not land the way this issue described them, and the reasons are worth recording.

    Must 3 – load is a string, not string | string[]. SPARQL Anything's -l takes one path, and the CLI usage documents repetition only for -v and -c. A file is also not interchangeable with a directory: a file loads into the default graph, a directory loads each RDF file into its own named graph. An array would have been unsupported and semantically different from what it suggests.

    Must 1 – no memory-aware default. This issue asks for a default sized from nproc and the cgroup limit. The converter spawns through a TaskRunner, so with Docker or a remote runner those describe the wrong machine: concurrency × heap is spent where the runner puts the processes, which the converter cannot see. The default is therefore 1, and the arithmetic map.sh does belongs to the caller, who knows their deployment. Same reasoning made heap a plain option with a 2g default rather than something derived.

    Should 7 – settled by removing the need. Returning the chunk outputs was for a caller that had to concatenate places + alternate names + ontology itself. Since #511 those outputs live in a per-run directory the converter deletes, so returning them would hand back paths that no longer exist – and since #791 a single convert() call takes all the jobs and concatenates them in order, which is what the returning was for.

    Should 8 – closed by the workDir change in #511, not separately. Per-chunk outputs no longer land beside the inputs: they go into a fresh directory under workDir, removed when the run ends. That also fixed a hazard this issue did not name – a stale .nt from an earlier run could satisfy must 4's non-empty check with the previous run's triples.

    Things found on the way that were not in this issue

    Package

    @lde/sparql-anything is on npm, bootstrapped manually after #511 (0.0.1), with a Trusted Publisher attached; CI has published since.

  4. ddeboer commented on Aug 31, 2026

    @ddeboer
    MemberAuthor

    #795 is merged, so every must and should in this issue is done. @lde/sparql-anything is at 0.1.0, @lde/task-runner-native at 0.2.18.

    What changed in #795 after the status comment above, all from review:

    • Stopping the other chunks is best effort. A task runner cannot stop what has already exited – DockerTaskRunner.stop() reads logs and stops the container, both of which reject for one that is gone – which is the likely state when a sibling has just failed. Unguarded, that replaced the conversion failure with a "no such container" error and rejected the join, leaving the other workers unawaited while the run directory was deleted underneath them.
    • A chunk could be started after the run had failed. The abort guard sits before two awaits, so a worker could pass it and then start a process while another chunk was failing – a process not in the set that failure stopped, whose whole conversion the run then waited out. Each worker now checks once its process exists, and stops it.

    New follow-up: DockerTaskRunner cannot run tasks in parallel

    DockerTaskRunner.run() force-removes any container with the configured containerName before starting a task, so with concurrency above one each chunk destroys the container of the chunk before it. @lde/sparql-anything documents this as a warning, but the fix belongs in @lde/task-runner-docker: a per-task container name, with containerName remaining the stable name only where it is the addressable endpoint (as sparql-qlever needs for network).

    Worth doing before the port runs conversions in Docker, which is the deployment the concurrency docs point at.

    What is left in this issue

    Only the Scope beyond #511 section:

    • Chunking – download.sh's chunk_with_header, which splits a TSV into CHUNK_SIZE-row chunks each carrying the header row. Generic, and the converter deliberately takes chunks it does not create, so this is the natural next LDE piece.
    • The ontology step – fetch a pinned ontology_v3.3.rdf and convert RDF/XML to N-Triples. Expressible as a job today: a query with no chunks and load pointing at the ontology.
    • Publication – gzip -9, zip -9, upload to Spaces. Probably stays in the consumer's CI.

    And test.sh's golden-file test, which whatever replaces map.sh has to keep: it runs the real conversion over fixtures and diffs against test/expected/geonames.nt.

  5. ddeboer commented on Sep 1, 2026

    @ddeboer
    MemberAuthor

    Correction: the conversion takes ~17 minutes, not ~24 hours

    Must 1 above claimed the job "already takes ~24 h in CI". That is wrong by a factor of about 35, and I have corrected it in the description. The real numbers, from geonames-rdf's Transform GeoNames data to RDF workflow:

    The last seven successful runs took 36, 37, 38, 39, 39, 40 and 45 minutes. Broken down, for the run of 2026-08-30:

    step duration
    download.sh – fetch, unzip, chunk 2 min 09 s
    map.sh – the conversion itself, already an xargs -P pool 17 min 29 s
    upload artifact 2 min 25 s
    compression (gzip -9, zip -9) 15 min 24 s
    uploads to Spaces 2 min 40 s

    Two things follow that are worth carrying into the port.

    Parallelism is still worth having, for a different reason than stated. map.sh is already parallel, so the 17 minutes is the pooled figure; the risk was never a day-long job getting longer, it was that a sequential convert() would have made a 17-minute step into a much longer one. The measurement does not change must 1, only the size of what it protects.

    Chunking is a smaller share than any of us assumed, and compression is a bigger one. All of download.sh – downloading several hundred megabytes, unzipping, and chunking – is 2 minutes, so chunking itself is seconds. Meanwhile compression takes nearly as long as the conversion. If anything in this pipeline deserves attention on wall-clock grounds, it is the 15 minutes of gzip -9/zip -9 in the publication step, not the chunker.

    I raise it because I have been arguing about chunker implementations on a "it is nothing against a day-long job" basis. Against a 40-minute workflow the argument still holds – a chunker difference of ~50 s is ~2 % – but it holds by a much smaller margin than I claimed, and the honest reason to choose is correctness rather than speed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions