Repository navigation
Port GeoNames RDF to LDE #782
Description
Activity
Musts 3, 4 and 5 are folded into #511 and are green there:
- 3 –
adminCodesFileis now an optionalload. One correction to this issue: it stays a singlestring, notstring | string[]. SPARQL Anything's-ltakes one path, and a file (loaded into the default graph) is not interchangeable with a directory (one named graph per RDF file), so an array would have been both unsupported by the CLI and semantically different from what it suggests. - 4 – each chunk's
--outputis asserted non-empty before it is concatenated. - 5 –
convert()awaits the concatenated output being flushed and closed.
Must 3 moved forward deliberately: the package is not published yet, so removing the domain-specific name is free now and a breaking bump later. It also prompted a general rule in AGENTS.md (#783) – a package's public API must stay domain-agnostic.
Remaining here: musts 1–2 (concurrency, JVM heap and a CLI passthrough) and shoulds 6–9.
- 3 –
Status: every must and should is done or in review
# Where Must 1 Parallelism #795 (open) Must 2 JVM heap, CLI escape hatch #789 Must 3 Generic, optional load#511 Must 4 Guard against an empty chunk output #511 Must 5 convert()resolving before the flush#511 Should 6 One worker pool, not two phases #791 Should 7 convert()returning the chunk outputs#791, by removing the need Should 8 Chunk outputs beside the inputs, never cleaned #511 Should 9 wait()race under concurrency#787 Once #795 merges, what is left in this issue is the Scope beyond #511 section: chunking, the ontology step, and publication.
Where the plan changed
Four items did not land the way this issue described them, and the reasons are worth recording.
Must 3 –
loadis astring, notstring | string[]. SPARQL Anything's-ltakes one path, and the CLI usage documents repetition only for-vand-c. A file is also not interchangeable with a directory: a file loads into the default graph, a directory loads each RDF file into its own named graph. An array would have been unsupported and semantically different from what it suggests.Must 1 – no memory-aware default. This issue asks for a default sized from
nprocand the cgroup limit. The converter spawns through aTaskRunner, so with Docker or a remote runner those describe the wrong machine:concurrency × heapis spent where the runner puts the processes, which the converter cannot see. The default is therefore1, and the arithmeticmap.shdoes belongs to the caller, who knows their deployment. Same reasoning madeheapa plain option with a2gdefault rather than something derived.Should 7 – settled by removing the need. Returning the chunk outputs was for a caller that had to concatenate places + alternate names + ontology itself. Since #511 those outputs live in a per-run directory the converter deletes, so returning them would hand back paths that no longer exist – and since #791 a single
convert()call takes all the jobs and concatenates them in order, which is what the returning was for.Should 8 – closed by the
workDirchange in #511, not separately. Per-chunk outputs no longer land beside the inputs: they go into a fresh directory underworkDir, removed when the run ends. That also fixed a hazard this issue did not name – a stale.ntfrom an earlier run could satisfy must 4's non-empty check with the previous run's triples.Things found on the way that were not in this issue
wait()could hang in three ways, not one (fix(task-runner-native): settle wait() for a process that already closed #787). Beyond the race described here,wait()removed everycloselistener, so it stranded astop()in flight and any concurrentwait(); and waiting twice answered the second caller with an empty string. The class now keeps one record per task, created at spawn, which removed the pid-keyed output buffers as well – those were unbounded and unsound, since the OS recycles pids.- Shell interpolation is quoted (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511), using a
shellQuotepromoted fromsparql-qleverinto@lde/task-runnerso both consumers share one implementation. - Docker never actually worked (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511). The generated query file was written to the host temp directory, which a container cannot see, while the docs claimed otherwise. Query files and chunk outputs now live under
workDirand are named relative to it. - A job's chunks cannot be silently ignored (feat(sparql-anything)!: convert a list of jobs, not a list of chunks #791). A query naming
{SOURCE}without chunks, or chunks whose query never names it, is rejected: SPARQL Anything would report the first as a parse error and the second not at all. cliArgscannot repeat the converter's own flags (feat(sparql-anything)!: convert a list of jobs, not a list of chunks #791).--format=TTLslipping through would concatenate Turtle as if it were N-Triples, and the run would stay green.
Package
@lde/sparql-anythingis on npm, bootstrapped manually after #511 (0.0.1), with a Trusted Publisher attached; CI has published since.#795 is merged, so every must and should in this issue is done.
@lde/sparql-anythingis at0.1.0,@lde/task-runner-nativeat0.2.18.What changed in #795 after the status comment above, all from review:
- Stopping the other chunks is best effort. A task runner cannot stop what has already exited –
DockerTaskRunner.stop()reads logs and stops the container, both of which reject for one that is gone – which is the likely state when a sibling has just failed. Unguarded, that replaced the conversion failure with a "no such container" error and rejected the join, leaving the other workers unawaited while the run directory was deleted underneath them. - A chunk could be started after the run had failed. The abort guard sits before two awaits, so a worker could pass it and then start a process while another chunk was failing – a process not in the set that failure stopped, whose whole conversion the run then waited out. Each worker now checks once its process exists, and stops it.
New follow-up:
DockerTaskRunnercannot run tasks in parallelDockerTaskRunner.run()force-removes any container with the configuredcontainerNamebefore starting a task, so withconcurrencyabove one each chunk destroys the container of the chunk before it.@lde/sparql-anythingdocuments this as a warning, but the fix belongs in@lde/task-runner-docker: a per-task container name, withcontainerNameremaining the stable name only where it is the addressable endpoint (assparql-qleverneeds fornetwork).Worth doing before the port runs conversions in Docker, which is the deployment the concurrency docs point at.
What is left in this issue
Only the Scope beyond #511 section:
- Chunking –
download.sh'schunk_with_header, which splits a TSV intoCHUNK_SIZE-row chunks each carrying the header row. Generic, and the converter deliberately takes chunks it does not create, so this is the natural next LDE piece. - The ontology step – fetch a pinned
ontology_v3.3.rdfand convert RDF/XML to N-Triples. Expressible as a job today: a query with no chunks andloadpointing at the ontology. - Publication –
gzip -9,zip -9, upload to Spaces. Probably stays in the consumer's CI.
And
test.sh's golden-file test, which whatever replacesmap.shhas to keep: it runs the real conversion over fixtures and diffs againsttest/expected/geonames.nt.- Stopping the other chunks is best effort. A task runner cannot stop what has already exited –
Correction: the conversion takes ~17 minutes, not ~24 hours
Must 1 above claimed the job "already takes ~24 h in CI". That is wrong by a factor of about 35, and I have corrected it in the description. The real numbers, from geonames-rdf's Transform GeoNames data to RDF workflow:
The last seven successful runs took 36, 37, 38, 39, 39, 40 and 45 minutes. Broken down, for the run of 2026-08-30:
step duration download.sh– fetch, unzip, chunk2 min 09 s map.sh– the conversion itself, already anxargs -Ppool17 min 29 s upload artifact 2 min 25 s compression ( gzip -9,zip -9)15 min 24 s uploads to Spaces 2 min 40 s Two things follow that are worth carrying into the port.
Parallelism is still worth having, for a different reason than stated.
map.shis already parallel, so the 17 minutes is the pooled figure; the risk was never a day-long job getting longer, it was that a sequentialconvert()would have made a 17-minute step into a much longer one. The measurement does not change must 1, only the size of what it protects.Chunking is a smaller share than any of us assumed, and compression is a bigger one. All of
download.sh– downloading several hundred megabytes, unzipping, and chunking – is 2 minutes, so chunking itself is seconds. Meanwhile compression takes nearly as long as the conversion. If anything in this pipeline deserves attention on wall-clock grounds, it is the 15 minutes ofgzip -9/zip -9in the publication step, not the chunker.I raise it because I have been arguing about chunker implementations on a "it is nothing against a day-long job" basis. Against a 40-minute workflow the argument still holds – a chunker difference of ~50 s is ~2 % – but it holds by a much smaller margin than I claimed, and the honest reason to choose is correctness rather than speed.
geonames-rdf converts the GeoNames data dumps to a ~13.5M-feature
geonames.ntwith a set of shell scripts (download.sh,map.sh,test.sh) driving the SPARQL Anything CLI. We want to port it to LDE. #511 is the first piece:@lde/sparql-anythingwith aSparqlAnythingConverter.This issue tracks what has to change before the migration can start, split into musts (the port cannot run correctly or finish without them) and shoulds (worth fixing, but a migration could ship around them).
Status
Every must and should is done.
@lde/sparql-anythingis at0.1.0.adminCodesFile→ a generic, optionalload— feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511convert()can resolve before the output is flushed — feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511convert()should return the chunk outputs — feat(sparql-anything)!: convert a list of jobs, not a list of chunks #791, by removing the needwait()race — feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511 (quoting) and fix(task-runner-native): settle wait() for a process that already closed #787 (the race)chunk()in@lde/sparql-anythingFour of them landed differently from the description below, and two follow-ups came out of the work; see the status comment and the one after it.
Still open: publication, and keeping
test.sh's golden-file diff – see Scope beyond the converter. The LDE side of the conversion itself is complete: all four ofmap.sh's invocation shapes are expressible against the current API, and@lde/sparql-anythingis at0.1.2.What
map.shactually doesIt makes four different SPARQL Anything invocations, not one:
--load{SOURCE}admin-codes.rqplaces.rqadmin-codes.ttlalternate-names.rqontology.rqontology.rdf(as the input)SparqlAnythingConverteras first proposed expressed only #2.Musts
1. Parallelism.
convert()is a sequentialforloop with anawaitinside.map.shruns anxargs -Pworker pool sized fromnproc, capped by the cgroup memory limit (or/proc/meminfoon a host), budgeting ~3 GB per worker. OnallCountriesthat is ~14 places chunks plus ~4 alternate-names chunks; running them one at a time multiplies the conversion, which takes ~17 min of the ~40 min workflow (measured, see below).Needs a
concurrencyoption; the memory-aware default is generic enough to belong in LDE rather than in the consumer. This is the one that decides whether the port is viable – everything else is correctness. Do it after 6+7, so the pool is not reworked when its unit of work changes shape.2. JVM heap control, and a CLI escape hatch.
map.shpasses-Xmx2gper worker precisely because SPARQL Anything materialises a chunk’s whole result graph before writing it. The converter hardcodesjava -jar …with no-Xmx, so each JVM takes a quarter of host RAM by default – which over-subscribes the moment concurrency lands. Also needs a passthrough for flags the converter does not model, such as-adon SPARQL Anything 1.2-dev for #624/#625. In review: feat(sparql-anything): cap the JVM heap, and pass through CLI arguments #789 (javaOptions,cliArgs).3. A generic, optional
load.adminCodesFilewas required, single, and named after a GeoNames concept; it could not express invocation test: collect coverage #3 (no--load) or feat: add dataset-registry-client #4 (--loadis the input). Nowload?: string. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)4. Guard against an empty chunk output. SPARQL Anything exits 0 when it cannot read or parse an input: it logs, writes an empty
--outputand stops, which is how we once shipped ageonames.ntmissing every ontology triple while the run stayed green.assertNonEmptynow fails the conversion, and an emptychunkPathsis rejected outright. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)5.
convert()resolving before the output was flushed.pipeline(…, {end: false})resolves when the source ends;output.end()ran un-awaited infinally, so a caller could read a truncated file. Nowawait finished(output), plus a newline between concatenated files so two triples can never share a line. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)Shoulds
6+7. One call over jobs, not one converter per query. Originally two separate points – a single worker pool, and returning the chunk outputs; 8’s
workDirchange merged them and made them more pressing rather than less.map.shdeliberately feeds both chunk sets through a single pool: an alternate-names chunk is roughly a quarter of the work of a places chunk (one triple per row, no admin-codes join), so the long jobs are listed first and the short ones fill the tail. OneSparqlAnythingConverterper query means two sequential drains, which loses that packing.And the port has to get places + alternate names +
ontology.ntinto one file.convert()returnsvoid, and now that the per-chunk outputs live in a run directory deleted infinally, the caller can no longer concatenate them itself – which it could when they sat next to the inputs.Both want the same change:
convert()takes jobs of{chunk, query, load}in a single call, or gains an explicitappend. Settle this before 1.8. Chunk outputs beside the inputs, never cleaned. Outputs and generated queries now go to a fresh per-run directory under a caller-supplied
workDir, removed infinally, and are referred to by runner-relative paths so the same command works on the host and underDockerTaskRunner. (feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511)9. Shell interpolation, and the
wait()race. Interpolated values go throughshellQuote(feat(sparql-anything): add SparqlAnythingConverter for chunked non-RDF to RDF conversion #511).wait()attached itscloselistener only afterrun()returned, so a process that failed fast could close first andwait()would never settle – invisible while the converter was sequential, fatal under a pool. Fixed in@lde/task-runner-native(fix(task-runner-native): settle wait() for a process that already closed #787).Scope beyond the converter
Conversion is roughly a third of the migration. Still unhomed:
download.sh’schunk_with_headersplits a TSV intoCHUNK_SIZE-row chunks, each prefixed with the header row the queries expect. Tracked in Split a large delimited file into row chunks with a repeated header #793, shipped in feat(sparql-anything): split a file into the chunks a conversion can hold #799 aschunk().loadand no chunks, which ismap.sh's fourth invocation exactly, and the non-empty guard applies to it like any other job. Fetching the pinnedontology_v3.3.rdfstays ingeonames-rdf– the pinned version is listed below as domain policy.gzip -9pluszip -9, uploaded to DigitalOcean Spaces. Probably stays in the consumer’s CI. Worth noting it is now the workflow's second-biggest cost, at ~15 min against the conversion's ~17.What should not move into LDE, because it is GeoNames domain policy rather than plumbing: the
awksynthesis of theadm1/adm2foreign keys, the filtering of alternate names belonging to out-of-scope features, and the pinned ontology version.Finally,
test.shis a golden-file test that runs the realmap.shover fixtures and diffs the result againsttest/expected/geonames.nt. Whatever replaces it has to keep that end-to-end diff: reviewing that file line by line is the actual test.