Repository navigation
Known 1.0 limitation (library fixed; Go improved in 1.0.2; remaining gap open): Schema.Rows scaling across goroutines #456
Description
Activity
The consumer's full measurement setup (
measuredby them, not reproduced here):Software: go/v1.0.0 with the linux-arm64 26.8.15.10 artifact (build 20261003.231921), Go 1.27 (the
golang:1.27image), CGO on.Host:
- Docker Desktop on an Apple-silicon Mac, with 14 vCPU visible, uid 65532, and the artifact cache mounted read-only.
- The host was loaded (load average 6–12), so absolute rates are noisy.
- The robust comparison is the RATIO within the same container:
RowsagainstParseBlock.
Workload: 100-row bodies, 8 goroutines against 1, one batch per call.
call speed-up at 8 goroutines Rows, JSONEachRow in, JSONCompactEachRow export, 8 own handles1.44× the same, one shared handle 1.41× Rows, no export1.6–1.8× Rows, JSONCompactEachRow input1.8–2.0× ParseBlock(JSONCompactEachRow, one row), own handles7.1–7.5× ParseBlock, one shared handle3.2× Describe0.98× For
Rowsat 8 goroutines: CPU per wall second is 2.6, and the ceiling is about 1,300–2,000 batches/s per process.Same harness on darwin-arm64:
Rowswith own handles scales 4.1–4.9× (CPU per wall second 8.2); with a shared handle, 3.2×.Read with care (
inferredfrom the shape, not a finding):- On Linux, a CPU per wall second of 2.6 at 8 goroutines points to contention or serialization rather than to CPU-bound work.
Describenot scaling at all fits that too.- Whether the lock is in the library or in the binding is what the profiling has to answer.
⚠️ Corrected below (#456 (comment)): the reporter's numbers were already taken withdoc_flags0, so this cause and workaround do not explain them. A second cause, in the call shape, is open.The cause is found, and a fix is being built for a 1.0.x artifact. No SDK change is needed. (Cause
measuredby the artifact producer on Linux runners; not reproduced in this repository.)What is slow: a
rowscall that asks for per-row documents.- Every binding's default is all three document groups: Go
DocAll, PythonDocFlags.ALL, TypeScriptDocFlags.All, RustDocFlags::ALL. - The library builds those documents after ClickHouse's own work for the call has finished, outside the per-call thread context ClickHouse sets up.
- Memory allocated there is charged to ClickHouse's single process-wide memory tracker, so concurrent callers contend on it.
Speed-up at 8 threads, against one thread:
doc_flagsspeed-up all groups (7, the default) 1.46× none (0) 7.2× A second, smaller limit: short calls all take one shared lock over the library's table of live handles.
Workaround until the 1.0.x artifact: ask only for the document groups you read.
doc_flagsonly chooses which groups the per-row documents carry: values, transformations, and DEFAULT/MATERIALIZED results.- The measured contrast is all groups against none. Whether a single group scales is not measured yet (
unverified). If you read no document groups, pass none:
binding no groups values only Go s.Rows(f, body, chtypes.WithDocFlags(0))chtypes.WithDocFlags(chtypes.DocValues)Python schema.rows(f, body, doc_flags=DocFlags(0))doc_flags=DocFlags.VALUESTypeScript schema.rows(f, body, { docFlags: 0 }){ docFlags: DocFlags.Values }Rust RowsOptions { doc_flags: Some(DocFlags::from_bits(0)), ..Default::default() }Some(DocFlags::VALUES)Next: when the fixed artifact reaches the production registry, this issue gets its build and a real-host run against it. If the fix does not land in the next 1.0.x, this goes into the known gaps in
docs/limitations.md.🤖 Generated with Claude Code
- Every binding's default is all three document groups: Go
- changed the title
[-]Known 1.0 limitation (under investigation): chs_preview_batch (Schema.Rows) barely scales across goroutines on Linux; block_create does[/-][+]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows with full per-row documents barely scales across threads on Linux[/+]on Oct 4, 2026 Correction to the comment above: narrowing
doc_flagsdoes NOT explain the reported numbers. The reporter took every Linux measurement in this issue withdoc_flags0 already (measuredby them). They still saw 1.4–2.0×, against 7.1–7.5× forParseBlockin the same container.So there is a second cause, in the call shape, and it is not found yet. The reporter's shape:
- a table with a column whose DEFAULT every row in the body omits;
date_time_input_format=best_effortand three other per-call settings;- an attached row filter;
- a JSONCompactEachRow export.
The probe that measured 7.2× at
doc_flags0 had none of the DEFAULT, the filter or those settings. It ruled out only the export path and the zone path.What this changes:
- The artifact producer is re-running with the reporter's exact shape, both on the current artifact and on the candidate fix.
- This issue stays open until the fix is measured against that shape. A fix for the
doc_flagscause alone does not close it. - The workaround: narrower
doc_flagsstill helps a caller that asks for document groups it does not read. It does nothing for the shape above, and there is no measured workaround for that shape yet.
🤖 Generated with Claude Code
- changed the title
[-]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows with full per-row documents barely scales across threads on Linux[/-][+]Known 1.0 limitation (one cause found, one open): Schema.Rows barely scales across threads on Linux[/+]on Oct 4, 2026 Re-measured with the reporter's exact call shape: it has the SAME cause, and the candidate fix covers it.
measuredby the artifact producer on CI runners, 8 threads against 1, own handles:- the reporter's table;
- every row omitting the DEFAULT column;
- the four per-call settings;
- the row filter;
- the JSONCompactEachRow export;
doc_flags0.
build arm64 (64 vCPU) amd64 (64 vCPU) current production artifact (two rounds) 3.60× / 4.11× 5.27× / 4.69× candidate fix 7.77× 7.90× the 0.x batch call, same patch (control) 7.57–7.94× 7.57–7.94× Why
doc_flags0 did not avoid it here (inferredfrom the profile): even with no document groups, the library still rewrites the batch result outside ClickHouse's per-call thread context. That rewrite is what contends on the process-wide memory tracker.- The reporter's rows are cheap, about 2 ms per 100-row call against 23 ms for the earlier probe's rows, so the rewrite is a much larger share of each call.
- The profile on the current artifact puts 52% (arm64) and 36% (amd64) of CPU in the memory tracker, all under that rewrite.
None of the four suspected parts of the shape is the cause. Measured one at a time on the current artifact:
- the bare shape scales 3.55–5.00×;
- each setting alone gives 3.57–3.62×;
- supplying the DEFAULT column against omitting it gives 4.09× against 3.55×;
- the filter and the export make no difference.
With one SHARED handle, the fix gives 4.67× (arm64) and 6.37× (amd64), against 4.17× and 5.76× for the 0.x call. That limit is the same as 0.x's, so it is not a 1.0 regression; give each thread its own handle for the full speed-up.
Still open:
- The reporter's own 1.4× under Docker Desktop (CPU per wall second 2.6) was not reproduced on CI runners. They should re-measure in their container once the fixed artifact is published.
- There is no workaround for this shape on the current artifact. Narrower
doc_flagshelps only calls that ask for document groups they do not read.
Next: the fix ships as a 1.0.x artifact with no SDK change. When it reaches the production registry, this issue gets its build and a real-host run. It closes once the reporter's shape is measured on the published artifact.
🤖 Generated with Claude Code
- changed the title
[-]Known 1.0 limitation (one cause found, one open): Schema.Rows barely scales across threads on Linux[/-][+]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows barely scales across threads on Linux[/+]on Oct 4, 2026 - added a commit that references this issue
on Oct 4, 2026 The fix is published: artifact build
20261004.052404is on the production registry. No SDK change is needed: every binding at 1.0.x picks it up at its next fetch. The previous build stays fetchable by digest.Checked here (
measured):- The tags: all eight production tags (the exact and floating tag of 26.3, 26.7, 26.8 and 26.9) resolve to the new build's manifests. Each manifest carries exactly one signed goldens document, revision 3.
- Real-host runs against production, all four bindings × three platforms (12 of 12 legs green on each line), including the goldens:
line run goldens r3 26.9.8.3 37182619694 PASS ×4 26.3.38.2 37182788256 PASS ×4 - The only skips are the two float cases whose platforms exclude darwin/arm64. That is the same set as the previous build's runs.
Scaling on the reporter's shape is 7.78× (arm64) and 7.90× (amd64) at 8 threads on the artifact producer's CI, against 1.09–5.16× before. That was measured on the same source tree before publishing.
This issue closes on the reporter's own re-measurement in their container against this build. Their original numbers came from that container and were never reproduced on CI runners.
🤖 Generated with Claude Code
The reporter's re-measurement shows no visible improvement in their environment, so this issue stays open.
measuredby the reporter on linux-arm64 in a container: 14 CPUs, the same test binary for both builds, each fetched fresh and confirmed from its manifest. Runs alternated new, old, new, old. 8 goroutines against 1, on the reporter's exact shape:build 20261004.052404 (fixed) build 20261003.231921 Rows, own handles2.01×, 2.13× 1.53×, 2.36× Rows, one shared handle1.47×, 1.48× 1.50×, 1.77× ParseBlock, own handles5.05–6.85× 5.47–6.93× ParseBlock, one shared handle3.0–3.6× 3.4× CPU per wall second for
Rowswith own handles is about 2.0–2.1 on the fixed build, against 7.8–7.9 forParseBlockin the same container. So on that host,Rowsthreads are waiting rather than working, whileParseBlockthreads run.Caveat (from the reporter): the host was busy, with a 1-minute load average of 10.8–15.6 on 14 cores. Ratio differences under about 30% are not resolvable.
What this means:
- On dedicated Linux hosts, the fix holds: 7.78× and 7.90× at 8 threads on the artifact producer's CI, as above.
- In the reporter's containerized environment, something on the
Rowspath still serializes. That is not explained yet (inferredfrom the CPU-per-wall figures). - The artifact producer resumes profiling against the production build. The reporter keeps this as a known limitation.
- If no SDK-side cause turns up, this goes into the known gaps in
docs/limitations.md.
🤖 Generated with Claude Code
- changed the title
[-]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows barely scales across threads on Linux[/-][+]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; open in containers): Schema.Rows scaling across threads on Linux[/+]on Oct 4, 2026 - changed the title
[-]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; open in containers): Schema.Rows scaling across threads on Linux[/-][+]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; still open in the reporter's environment): Schema.Rows scaling across threads on Linux[/+]on Oct 4, 2026 Correction to the environment named above: the reporter's measurements ran in OrbStack's Linux VM (14 vCPU) on a loaded Apple-silicon Mac, not in Docker Desktop. The point stands: a VM on a busy laptop, not a dedicated host.
Next: the reporter will measure on a dedicated Linux CI runner. That decides whether the rest is a VM effect or something to profile further.
- From a source read of the Go binding (
inferred, not measured): theRowspath takes no Go-side lock thatParseBlockdoes not also take.
🤖 Generated with Claude Code
- From a source read of the Go binding (
The remaining limit is in the Go binding: garbage-collector churn from decoding each
Rowsresult. The library is not the cause.measuredby the reporter on build 20261004.052404: an OrbStack Linux VM, 14 vCPU, 100-row bodies, own handles, a 5 s window after a 2 s warm-up.Go Rows1 goroutine 8 goroutines speed-up CPU per wall second at 8 default GC (3 runs) 832–847 calls/s 1766–1820 calls/s 2.1× 2.8–2.9 GOGC=off GOMEMLIMIT=4GiB852 calls/s 6131 calls/s 7.2× 8.07 - The control:
ParseBlockgives 6.9× with the default GC and 7.4× with it off. - GC trace: about 157 GC cycles per second at 8 goroutines, on a heap of only a few MB. GC work itself is only about 10% of CPU.
- Why it loses scaling (
inferred): the loss is the scheduling around those frequent cycles, threads parked and woken, rather than the collection work. - Why
Rowsallocates so much: the Go binding decodes every batch result into a generic JSON tree and builds a string per row while doing it. - Why CI never saw it: the artifact producer's CI drove the library from C, never through the Go binding.
Workaround for Go until the patch: make GC cycles rarer.
- Measured:
GOGC=offtogether with aGOMEMLIMITsized for your process restores 7.2×. Never setGOGC=offwithout a memory limit. - Not measured (
inferred): a higherGOGC(for example 400–800), or a process whose live heap is already large, cycles less often and should be affected less.
The fix: the Go binding's batch decoding will allocate far less per call. It ships as a Go 1.0.x patch with no API change. The library fix in build 20261004.052404 stays in force: it removes the library-side contention.
- Only Go was measured. Whether the Python and TypeScript bindings' own runtimes show a similar effect is not known yet.
- This issue closes when the reporter's harness shows the patched Go binding scaling with the default GC.
🤖 Generated with Claude Code
- The control:
- changed the title
[-]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; still open in the reporter's environment): Schema.Rows scaling across threads on Linux[/-][+]Known 1.0 limitation (library fixed in build 20261004.052404; Go binding fix coming in 1.0.x): Schema.Rows scaling across goroutines[/+]on Oct 4, 2026 The Go decode fix halves the GC cost per call, which doubles
Rowsscaling. It does not reach the control yet.measuredby the reporter in the same OrbStack Linux VM, on the candidate fix commit a8ed759 against artifact build 20261004.052404, with the DEFAULT GC (noGOGC, noGOMEMLIMIT), 8 goroutines against 1, own handles, a 1-minute host load of 11–12 on 14 cores:1 goroutine 8 goroutines speed-up CPU per wall second at 8 GC cycles per second at 8 Rows, Go 1.0.1818 calls/s 1698 calls/s 2.08× 2.83 155 Rows, candidate fix (two runs)861–862 calls/s 3573–3638 calls/s 4.15–4.22× ~4.95 161–167 ParseBlock(control)1271 calls/s 8664 calls/s 6.82× 7.82 34 What the fix changed: the batch decoding itself now allocates far less (
measuredby the Go benchmark on real 100-row documents):doc_flagsallocations per call, before after 0 4,476 173 7 56,822 2,641 So throughput doubles at a similar GC rate (
inferred), and the GC cycles per call halve.Not done yet: about 0.045 GC cycles per call remain. On the small heaps of these processes, that is still enough to cost scaling (
inferred). The next step is a profile of the wholeRowscall to find the remaining per-call allocations outside the decoder. This issue stays open.🤖 Generated with Claude Code
A second Go fix, cutting the rest of the per-call allocations in
Rows, takes the reporter's shape to about 5.15× at 8 goroutines with the default GC.What it changes (
measuredhere on the reporter's exact shape, fullRowscall,-benchmem): 412 allocations / 75.2 KB per call drop to 76 / 45.4 KB.- The result document is copied out of the library into a reused buffer, and every returned value stays ordinary Go memory.
- Most per-row small allocations are gone.
The reporter's harness:
measuredin the same OrbStack Linux VM, artifact build 20261004.052404, DEFAULT GC, 8 goroutines against 1, own handles, host load about 5. Run back to back against the first fix:1 goroutine 8 goroutines speed-up CPU per wall second at 8 GC cycles per second at 8 Rows, first fix (inv1)879 calls/s 3643 calls/s 4.14× 4.88 165 Rows, second fix (two runs)865–876 calls/s 4466–4492 calls/s 5.13–5.16× 6.0 147–149 ParseBlock(control)1299 calls/s 9151 calls/s 7.04× 7.88 34 GC cycles per call at 8 goroutines fell from about 0.045 to about 0.033 (
inferredfrom the rates).What is left: mostly the returned rows themselves (about 23 KB for 100 rows), the stock JSON decoder's own buffer, and the returned export payload. Removing those needs an API or document-shape change, not a binding fix.
The second fix merges only after CI runs its new concurrency tests under the race detector. Both fixes then ship in Go 1.0.2. This issue stays open for the remaining gap and for whether the other bindings' runtimes show the same effect.
🤖 Generated with Claude Code
- added a commit that references this issue
on Oct 6, 2026 Go 1.0.2 is released (the Go proxy
v1.0.2, 2026-10-06) with both binding fixes. Upgrade to get them; no code change is needed.What's in it:
Rowsdecodes its result in one streaming pass;- most per-call allocations around the call are gone.
- On the reporter's 100-row shape, a full call went from 412 allocations / 75 KB to 76 / 45 KB.
The reporter's measurement (8 goroutines against 1, default GC): about 5.15×, up from about 2.1× on Go 1.0.1. The
ParseBlockcontrol reaches about 7×.This issue stays open for the rest of the gap and for the other bindings:
- What remains per call is mostly the returned rows themselves, the stock JSON decoder's buffer and the returned export payload.
- Python, TypeScript and Rust also decode each batch into a generic tree. Whether their runtimes show a similar effect is not measured yet.
Until then, for Go: a process with a larger live heap, or a higher
GOGC(orGOGC=offwith a sizedGOMEMLIMIT, as measured above), runs fewer GC cycles per call.🤖 Generated with Claude Code
- changed the title
[-]Known 1.0 limitation (library fixed in build 20261004.052404; Go binding fix coming in 1.0.x): Schema.Rows scaling across goroutines[/-][+]Known 1.0 limitation (library fixed; Go improved in 1.0.2; remaining gap open): Schema.Rows scaling across goroutines[/+]on Oct 6, 2026 Scaling measurement on GitHub-hosted Linux runners (standard 4-vCPU), same shape as the earlier runs: one table, 100-row body, the 4 settings, a tenant filter,
JSONCompactEachRowexport,WithDocFlags(0), default GC. Each point is a 2 s warm-up then a 5 s window. Run once.- chtypes go v1.0.2, artifact build 20261004.052404 (26.8.15.10; the manifest in the fetched cache contains that build string), Go 1.27.1,
nproc= 4,GOMAXPROCS= 4. ubuntu-24.04: linux-amd64, AMD EPYC 7763.ubuntu-24.04-arm: linux-arm64, Neoverse-N2.- Run: https://github.com/Wave-RF/WaveHouse/actions/runs/37459986581
Own handles unless noted. Ratio is calls/s relative to g=1 for the same work on the same runner. CPU/wall is process CPU seconds per wall second. GC/s is completed GC cycles per second (the count
gctrace=1prints, taken fromruntime.MemStats.NumGC).ubuntu-24.04 (x86_64, EPYC 7763)
work g calls/s ratio CPU/wall GC/s Rows 1 400 1.00 1.04 5.4 Rows 2 733 1.83 2.05 10.8 Rows 4 842 2.10 3.82 13.2 Rows 8 833 2.08 3.93 15.4 Rows, shared handle 8 822 n/a 3.94 15.0 ParseBlock 1 575 1.00 1.04 1.2 ParseBlock 2 1118 1.95 2.08 2.6 ParseBlock 4 1290 2.25 3.89 3.0 ParseBlock 8 1276 2.22 3.96 3.0 ubuntu-24.04-arm (aarch64, Neoverse-N2)
work g calls/s ratio CPU/wall GC/s Rows 1 384 1.00 1.01 5.6 Rows 2 741 1.93 2.01 11.0 Rows 4 1387 3.61 3.76 23.4 Rows 8 1379 3.59 3.84 27.0 Rows, shared handle 8 1157 n/a 3.92 22.0 ParseBlock 1 534 1.00 1.01 1.6 ParseBlock 2 1045 1.96 2.01 2.4 ParseBlock 4 1965 3.68 3.84 4.6 ParseBlock 8 1960 3.67 3.95 4.6 What it shows: on both runners
Rowsscales about as well asParseBlockon the same machine (arm64 3.61x vs 3.68x at g=4, close to the 4-way ceiling; amd64 2.10x vs 2.25x, with the process using ~3.8 CPUs in both, so the cores are busy rather than waiting), so on 4 vCPUs we see noRows-specific scaling gap, and GC rates here (5 to 27/s) are far below the ~80 to 160/s of the earlier 14-core runs. Whether the amd64 ceiling near 2x reflects SMT sharing of physical cores on that runner type is inferred, not measured, and this run does not test the 8-core-and-up regime where the gap showed.- chtypes go v1.0.2, artifact build 20261004.052404 (26.8.15.10; the manifest in the fetched cache contains that build string), Go 1.27.1,
Closing: at 8 and 16 goroutines, Go
Rowsnow scales nearly as well asParseBlock. This is the ≥8-core datapoint the 4-vCPU run above could not give.Setup (
measured). The reporter's exact work units:- one table, the 100-row body, the 4 settings, a tenant filter,
JSONCompactEachRowexport,WithDocFlags(0); - go v1.0.2, artifact build 20261004.052404 (26.8.15.10, read from the loaded library's
build_info), Go 1.27.1, default GC; - own handles unless marked; each point is a 2 s warm-up then a 5 s window; run once.
Two self-hosted Linux runners:
- arm64: Neoverse-V2, 192 CPUs;
- amd64: Xeon Platinum 8375C, 64 CPUs.
GOMAXPROCSwas 60 on both, so g=16 is not oversubscribed. The ratio is calls/s against g=1 for the same work.arm64 (Neoverse-V2)
work g calls/s ratio CPU/wall GC/s Rows 1 391 1.00 1.04 8.8 Rows 2 805 2.06 2.07 20.4 Rows 4 1571 4.01 4.14 56.2 Rows 8 2699 6.90 8.25 266.3 Rows 16 4446 11.36 15.33 439.5 Rows, shared handle 8 1103 2.82 8.13 83.2 ParseBlock 1 575 1.00 1.01 2.0 ParseBlock 2 1136 1.97 2.02 4.2 ParseBlock 4 2274 3.95 4.04 10.0 ParseBlock 8 4479 7.79 8.07 27.0 ParseBlock 16 7559 13.15 16.09 241.1 amd64 (Xeon 8375C)
work g calls/s ratio CPU/wall GC/s Rows 1 392 1.00 1.03 8.0 Rows 2 780 1.99 2.05 19.8 Rows 4 1524 3.88 4.13 60.2 Rows 8 2664 6.79 8.19 210.1 Rows 16 4285 10.92 14.48 280.2 Rows, shared handle 8 2385 6.08 8.25 183.7 ParseBlock 1 573 1.00 1.02 2.0 ParseBlock 2 1144 1.99 2.03 4.2 ParseBlock 4 2282 3.98 4.05 10.4 ParseBlock 8 4414 7.70 8.10 32.6 ParseBlock 16 6859 11.96 15.92 293.1 What it shows:
- g=8 (
measured):Rowsreaches 6.90× on arm64 and 6.79× on amd64, againstParseBlock's 7.79× and 7.70×. At g=16 it reaches 11.4× and 10.9×, against 13.2× and 12.0×. CPU/wall tracks g, so the workers stay busy. - The original report's ~2× ceiling is gone at this core count. The earlier gap came from the library's shared state (fixed in the 20261004 build) and the Go decode allocations (cut in 1.0.2, from 412 to 76 per call).
- What remains (
inferred, not isolated):Rowsruns at about 0.88× ofParseBlock's scaling at g=8, and that tracks its GC rate: 266 vs 27 GC/s on arm64.Rowsstill allocates per call whereParseBlockhands back a block. Cutting that further is a possible future Go optimisation, not part of this defect.
One more observation, outside this issue's shape: a single shared handle at g=8 reaches only 2.82× on arm64, but 6.08× on amd64. It was measured once; own handles are fine on both. The artifact producer already tracks a library-side result of the same arm-vs-x86 shape for one shared handle, and this Go datapoint has been passed on there. It is below the 4× to 6.5× range that
docs/reference/bindings-v1.md§3 quotes for a shared handle at 8 threads, so that range is due a recheck on arm64 once the library side is understood. The advice there stands, and is what this table measures: compile a schema per goroutine for the full speed-up.Closing as fixed by the 20261004 library build plus go v1.0.2.
🤖 Generated with Claude Code
- one table, the 100-row body, the 4 settings, a tenant filter,
- added a commit that references this issue
on Oct 6, 2026
Known 1.0 limitation, under investigation:
chs_preview_batch(GoSchema.Rows, and its equivalents) barely scales across goroutines or threads on Linux, whilechs_block_create(GoParseBlock) does.Reported numbers (
measuredby a downstream consumer on 1.0.0; this repository has not reproduced them yet). Each figure is throughput with N concurrent callers relative to one:chs_preview_batch(Schema.Rows)chs_block_create(ParseBlock)What it means for 1.0 users: calling
Rowsfrom many goroutines on Linux gives little more throughput than one caller. Results stay correct; only the parallel speed-up is missing. Where a workload allows it,ParseBlockscales.Status (updated 2026-10-06): partly fixed; the rest of the gap is open.
Rowsallocates far less per call (412 → 76 allocations, 75 → 45 KB on the reported shape). The reporter's 8-goroutine speed-up with the default GC rose from about 2.1× to about 5.15×, against about 7× for theParseBlockcontrol. Details: Known 1.0 limitation (library fixed; Go improved in 1.0.2; remaining gap open): Schema.Rows scaling across goroutines #456 (comment)Labelled a known 1.0 limitation. It belongs alongside
docs/limitations.md#known-gaps-in-10, and is added there if it is not fixed in the next 1.0.x.🤖 Generated with Claude Code