Skip to content

Known 1.0 limitation (library fixed; Go improved in 1.0.2; remaining gap open): Schema.Rows scaling across goroutines #456

Description

@EricAndrechek

Known 1.0 limitation, under investigation: chs_preview_batch (Go Schema.Rows, and its equivalents) barely scales across goroutines or threads on Linux, while chs_block_create (Go ParseBlock) does.

Reported numbers (measured by a downstream consumer on 1.0.0; this repository has not reproduced them yet). Each figure is throughput with N concurrent callers relative to one:

call linux-arm64 darwin-arm64
chs_preview_batch (Schema.Rows) ≈1.4× ≈4–5×
chs_block_create (ParseBlock) ≈7× —
the 0.x batch call, for comparison ≈4–6× —

What it means for 1.0 users: calling Rows from many goroutines on Linux gives little more throughput than one caller. Results stay correct; only the parallel speed-up is missing. Where a workload allows it, ParseBlock scales.

Status (updated 2026-10-06): partly fixed; the rest of the gap is open.

  1. Library (fixed): artifact build 20261004.052404 removes the library's contention on ClickHouse's process-wide memory tracker.
  2. Go binding (improved in Go 1.0.2, released 2026-10-06): Rows allocates far less per call (412 → 76 allocations, 75 → 45 KB on the reported shape). The reporter's 8-goroutine speed-up with the default GC rose from about 2.1× to about 5.15×, against about 7× for the ParseBlock control. Details: Known 1.0 limitation (library fixed; Go improved in 1.0.2; remaining gap open): Schema.Rows scaling across goroutines #456 (comment)
  3. Open: what remains is mostly the returned rows themselves. Whether Python, TypeScript and Rust show a similar runtime effect is not measured.

Labelled a known 1.0 limitation. It belongs alongside docs/limitations.md#known-gaps-in-10, and is added there if it is not fixed in the next 1.0.x.

🤖 Generated with Claude Code

Activity

  1. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    The consumer's full measurement setup (measured by them, not reproduced here):

    Software: go/v1.0.0 with the linux-arm64 26.8.15.10 artifact (build 20261003.231921), Go 1.27 (the golang:1.27 image), CGO on.

    Host:

    • Docker Desktop on an Apple-silicon Mac, with 14 vCPU visible, uid 65532, and the artifact cache mounted read-only.
    • The host was loaded (load average 6–12), so absolute rates are noisy.
    • The robust comparison is the RATIO within the same container: Rows against ParseBlock.

    Workload: 100-row bodies, 8 goroutines against 1, one batch per call.

    call speed-up at 8 goroutines
    Rows, JSONEachRow in, JSONCompactEachRow export, 8 own handles 1.44×
    the same, one shared handle 1.41×
    Rows, no export 1.6–1.8×
    Rows, JSONCompactEachRow input 1.8–2.0×
    ParseBlock (JSONCompactEachRow, one row), own handles 7.1–7.5×
    ParseBlock, one shared handle 3.2×
    Describe 0.98×

    For Rows at 8 goroutines: CPU per wall second is 2.6, and the ceiling is about 1,300–2,000 batches/s per process.

    Same harness on darwin-arm64: Rows with own handles scales 4.1–4.9× (CPU per wall second 8.2); with a shared handle, 3.2×.

    Read with care (inferred from the shape, not a finding):

    • On Linux, a CPU per wall second of 2.6 at 8 goroutines points to contention or serialization rather than to CPU-bound work.
    • Describe not scaling at all fits that too.
    • Whether the lock is in the library or in the binding is what the profiling has to answer.
  2. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    ⚠️ Corrected below (#456 (comment)): the reporter's numbers were already taken with doc_flags 0, so this cause and workaround do not explain them. A second cause, in the call shape, is open.

    The cause is found, and a fix is being built for a 1.0.x artifact. No SDK change is needed. (Cause measured by the artifact producer on Linux runners; not reproduced in this repository.)

    What is slow: a rows call that asks for per-row documents.

    • Every binding's default is all three document groups: Go DocAll, Python DocFlags.ALL, TypeScript DocFlags.All, Rust DocFlags::ALL.
    • The library builds those documents after ClickHouse's own work for the call has finished, outside the per-call thread context ClickHouse sets up.
    • Memory allocated there is charged to ClickHouse's single process-wide memory tracker, so concurrent callers contend on it.

    Speed-up at 8 threads, against one thread:

    doc_flags speed-up
    all groups (7, the default) 1.46×
    none (0) 7.2×

    A second, smaller limit: short calls all take one shared lock over the library's table of live handles.

    Workaround until the 1.0.x artifact: ask only for the document groups you read.

    • doc_flags only chooses which groups the per-row documents carry: values, transformations, and DEFAULT/MATERIALIZED results.
    • The measured contrast is all groups against none. Whether a single group scales is not measured yet (unverified). If you read no document groups, pass none:
    binding no groups values only
    Go s.Rows(f, body, chtypes.WithDocFlags(0)) chtypes.WithDocFlags(chtypes.DocValues)
    Python schema.rows(f, body, doc_flags=DocFlags(0)) doc_flags=DocFlags.VALUES
    TypeScript schema.rows(f, body, { docFlags: 0 }) { docFlags: DocFlags.Values }
    Rust RowsOptions { doc_flags: Some(DocFlags::from_bits(0)), ..Default::default() } Some(DocFlags::VALUES)

    Next: when the fixed artifact reaches the production registry, this issue gets its build and a real-host run against it. If the fix does not land in the next 1.0.x, this goes into the known gaps in docs/limitations.md.

    🤖 Generated with Claude Code

  3. changed the title [-]Known 1.0 limitation (under investigation): chs_preview_batch (Schema.Rows) barely scales across goroutines on Linux; block_create does[/-] [+]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows with full per-row documents barely scales across threads on Linux[/+] on Oct 4, 2026
  4. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    Correction to the comment above: narrowing doc_flags does NOT explain the reported numbers. The reporter took every Linux measurement in this issue with doc_flags 0 already (measured by them). They still saw 1.4–2.0×, against 7.1–7.5× for ParseBlock in the same container.

    So there is a second cause, in the call shape, and it is not found yet. The reporter's shape:

    • a table with a column whose DEFAULT every row in the body omits;
    • date_time_input_format=best_effort and three other per-call settings;
    • an attached row filter;
    • a JSONCompactEachRow export.

    The probe that measured 7.2× at doc_flags 0 had none of the DEFAULT, the filter or those settings. It ruled out only the export path and the zone path.

    What this changes:

    • The artifact producer is re-running with the reporter's exact shape, both on the current artifact and on the candidate fix.
    • This issue stays open until the fix is measured against that shape. A fix for the doc_flags cause alone does not close it.
    • The workaround: narrower doc_flags still helps a caller that asks for document groups it does not read. It does nothing for the shape above, and there is no measured workaround for that shape yet.

    🤖 Generated with Claude Code

  5. changed the title [-]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows with full per-row documents barely scales across threads on Linux[/-] [+]Known 1.0 limitation (one cause found, one open): Schema.Rows barely scales across threads on Linux[/+] on Oct 4, 2026
  6. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    Re-measured with the reporter's exact call shape: it has the SAME cause, and the candidate fix covers it. measured by the artifact producer on CI runners, 8 threads against 1, own handles:

    • the reporter's table;
    • every row omitting the DEFAULT column;
    • the four per-call settings;
    • the row filter;
    • the JSONCompactEachRow export;
    • doc_flags 0.
    build arm64 (64 vCPU) amd64 (64 vCPU)
    current production artifact (two rounds) 3.60× / 4.11× 5.27× / 4.69×
    candidate fix 7.77× 7.90×
    the 0.x batch call, same patch (control) 7.57–7.94× 7.57–7.94×

    Why doc_flags 0 did not avoid it here (inferred from the profile): even with no document groups, the library still rewrites the batch result outside ClickHouse's per-call thread context. That rewrite is what contends on the process-wide memory tracker.

    • The reporter's rows are cheap, about 2 ms per 100-row call against 23 ms for the earlier probe's rows, so the rewrite is a much larger share of each call.
    • The profile on the current artifact puts 52% (arm64) and 36% (amd64) of CPU in the memory tracker, all under that rewrite.

    None of the four suspected parts of the shape is the cause. Measured one at a time on the current artifact:

    • the bare shape scales 3.55–5.00×;
    • each setting alone gives 3.57–3.62×;
    • supplying the DEFAULT column against omitting it gives 4.09× against 3.55×;
    • the filter and the export make no difference.

    With one SHARED handle, the fix gives 4.67× (arm64) and 6.37× (amd64), against 4.17× and 5.76× for the 0.x call. That limit is the same as 0.x's, so it is not a 1.0 regression; give each thread its own handle for the full speed-up.

    Still open:

    • The reporter's own 1.4× under Docker Desktop (CPU per wall second 2.6) was not reproduced on CI runners. They should re-measure in their container once the fixed artifact is published.
    • There is no workaround for this shape on the current artifact. Narrower doc_flags helps only calls that ask for document groups they do not read.

    Next: the fix ships as a 1.0.x artifact with no SDK change. When it reaches the production registry, this issue gets its build and a real-host run. It closes once the reporter's shape is measured on the published artifact.

    🤖 Generated with Claude Code

  7. changed the title [-]Known 1.0 limitation (one cause found, one open): Schema.Rows barely scales across threads on Linux[/-] [+]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows barely scales across threads on Linux[/+] on Oct 4, 2026
  8. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    The fix is published: artifact build 20261004.052404 is on the production registry. No SDK change is needed: every binding at 1.0.x picks it up at its next fetch. The previous build stays fetchable by digest.

    Checked here (measured):

    • The tags: all eight production tags (the exact and floating tag of 26.3, 26.7, 26.8 and 26.9) resolve to the new build's manifests. Each manifest carries exactly one signed goldens document, revision 3.
    • Real-host runs against production, all four bindings × three platforms (12 of 12 legs green on each line), including the goldens:
    line run goldens r3
    26.9.8.3 37182619694 PASS ×4
    26.3.38.2 37182788256 PASS ×4
    • The only skips are the two float cases whose platforms exclude darwin/arm64. That is the same set as the previous build's runs.

    Scaling on the reporter's shape is 7.78× (arm64) and 7.90× (amd64) at 8 threads on the artifact producer's CI, against 1.09–5.16× before. That was measured on the same source tree before publishing.

    This issue closes on the reporter's own re-measurement in their container against this build. Their original numbers came from that container and were never reproduced on CI runners.

    🤖 Generated with Claude Code

  9. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    The reporter's re-measurement shows no visible improvement in their environment, so this issue stays open.

    measured by the reporter on linux-arm64 in a container: 14 CPUs, the same test binary for both builds, each fetched fresh and confirmed from its manifest. Runs alternated new, old, new, old. 8 goroutines against 1, on the reporter's exact shape:

    build 20261004.052404 (fixed) build 20261003.231921
    Rows, own handles 2.01×, 2.13× 1.53×, 2.36×
    Rows, one shared handle 1.47×, 1.48× 1.50×, 1.77×
    ParseBlock, own handles 5.05–6.85× 5.47–6.93×
    ParseBlock, one shared handle 3.0–3.6× 3.4×

    CPU per wall second for Rows with own handles is about 2.0–2.1 on the fixed build, against 7.8–7.9 for ParseBlock in the same container. So on that host, Rows threads are waiting rather than working, while ParseBlock threads run.

    Caveat (from the reporter): the host was busy, with a 1-minute load average of 10.8–15.6 on 14 cores. Ratio differences under about 30% are not resolvable.

    What this means:

    • On dedicated Linux hosts, the fix holds: 7.78× and 7.90× at 8 threads on the artifact producer's CI, as above.
    • In the reporter's containerized environment, something on the Rows path still serializes. That is not explained yet (inferred from the CPU-per-wall figures).
    • The artifact producer resumes profiling against the production build. The reporter keeps this as a known limitation.
    • If no SDK-side cause turns up, this goes into the known gaps in docs/limitations.md.

    🤖 Generated with Claude Code

  10. changed the title [-]Known 1.0 limitation (cause found, fix coming in 1.0.x): Schema.Rows barely scales across threads on Linux[/-] [+]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; open in containers): Schema.Rows scaling across threads on Linux[/+] on Oct 4, 2026
  11. changed the title [-]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; open in containers): Schema.Rows scaling across threads on Linux[/-] [+]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; still open in the reporter's environment): Schema.Rows scaling across threads on Linux[/+] on Oct 4, 2026
  12. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    Correction to the environment named above: the reporter's measurements ran in OrbStack's Linux VM (14 vCPU) on a loaded Apple-silicon Mac, not in Docker Desktop. The point stands: a VM on a busy laptop, not a dedicated host.

    Next: the reporter will measure on a dedicated Linux CI runner. That decides whether the rest is a VM effect or something to profile further.

    • From a source read of the Go binding (inferred, not measured): the Rows path takes no Go-side lock that ParseBlock does not also take.

    🤖 Generated with Claude Code

  13. EricAndrechek commented on Oct 4, 2026

    @EricAndrechek
    MemberAuthor

    The remaining limit is in the Go binding: garbage-collector churn from decoding each Rows result. The library is not the cause. measured by the reporter on build 20261004.052404: an OrbStack Linux VM, 14 vCPU, 100-row bodies, own handles, a 5 s window after a 2 s warm-up.

    Go Rows 1 goroutine 8 goroutines speed-up CPU per wall second at 8
    default GC (3 runs) 832–847 calls/s 1766–1820 calls/s 2.1× 2.8–2.9
    GOGC=off GOMEMLIMIT=4GiB 852 calls/s 6131 calls/s 7.2× 8.07
    • The control: ParseBlock gives 6.9× with the default GC and 7.4× with it off.
    • GC trace: about 157 GC cycles per second at 8 goroutines, on a heap of only a few MB. GC work itself is only about 10% of CPU.
    • Why it loses scaling (inferred): the loss is the scheduling around those frequent cycles, threads parked and woken, rather than the collection work.
    • Why Rows allocates so much: the Go binding decodes every batch result into a generic JSON tree and builds a string per row while doing it.
    • Why CI never saw it: the artifact producer's CI drove the library from C, never through the Go binding.

    Workaround for Go until the patch: make GC cycles rarer.

    • Measured: GOGC=off together with a GOMEMLIMIT sized for your process restores 7.2×. Never set GOGC=off without a memory limit.
    • Not measured (inferred): a higher GOGC (for example 400–800), or a process whose live heap is already large, cycles less often and should be affected less.

    The fix: the Go binding's batch decoding will allocate far less per call. It ships as a Go 1.0.x patch with no API change. The library fix in build 20261004.052404 stays in force: it removes the library-side contention.

    • Only Go was measured. Whether the Python and TypeScript bindings' own runtimes show a similar effect is not known yet.
    • This issue closes when the reporter's harness shows the patched Go binding scaling with the default GC.

    🤖 Generated with Claude Code

  14. changed the title [-]Known 1.0 limitation (fixed on dedicated hosts in build 20261004.052404; still open in the reporter's environment): Schema.Rows scaling across threads on Linux[/-] [+]Known 1.0 limitation (library fixed in build 20261004.052404; Go binding fix coming in 1.0.x): Schema.Rows scaling across goroutines[/+] on Oct 4, 2026
  15. EricAndrechek commented on Oct 6, 2026

    @EricAndrechek
    MemberAuthor

    The Go decode fix halves the GC cost per call, which doubles Rows scaling. It does not reach the control yet.

    measured by the reporter in the same OrbStack Linux VM, on the candidate fix commit a8ed759 against artifact build 20261004.052404, with the DEFAULT GC (no GOGC, no GOMEMLIMIT), 8 goroutines against 1, own handles, a 1-minute host load of 11–12 on 14 cores:

    1 goroutine 8 goroutines speed-up CPU per wall second at 8 GC cycles per second at 8
    Rows, Go 1.0.1 818 calls/s 1698 calls/s 2.08× 2.83 155
    Rows, candidate fix (two runs) 861–862 calls/s 3573–3638 calls/s 4.15–4.22× ~4.95 161–167
    ParseBlock (control) 1271 calls/s 8664 calls/s 6.82× 7.82 34

    What the fix changed: the batch decoding itself now allocates far less (measured by the Go benchmark on real 100-row documents):

    doc_flags allocations per call, before after
    0 4,476 173
    7 56,822 2,641

    So throughput doubles at a similar GC rate (inferred), and the GC cycles per call halve.

    Not done yet: about 0.045 GC cycles per call remain. On the small heaps of these processes, that is still enough to cost scaling (inferred). The next step is a profile of the whole Rows call to find the remaining per-call allocations outside the decoder. This issue stays open.

    🤖 Generated with Claude Code

  16. EricAndrechek commented on Oct 6, 2026

    @EricAndrechek
    MemberAuthor

    A second Go fix, cutting the rest of the per-call allocations in Rows, takes the reporter's shape to about 5.15× at 8 goroutines with the default GC.

    What it changes (measured here on the reporter's exact shape, full Rows call, -benchmem): 412 allocations / 75.2 KB per call drop to 76 / 45.4 KB.

    • The result document is copied out of the library into a reused buffer, and every returned value stays ordinary Go memory.
    • Most per-row small allocations are gone.

    The reporter's harness: measured in the same OrbStack Linux VM, artifact build 20261004.052404, DEFAULT GC, 8 goroutines against 1, own handles, host load about 5. Run back to back against the first fix:

    1 goroutine 8 goroutines speed-up CPU per wall second at 8 GC cycles per second at 8
    Rows, first fix (in v1) 879 calls/s 3643 calls/s 4.14× 4.88 165
    Rows, second fix (two runs) 865–876 calls/s 4466–4492 calls/s 5.13–5.16× 6.0 147–149
    ParseBlock (control) 1299 calls/s 9151 calls/s 7.04× 7.88 34

    GC cycles per call at 8 goroutines fell from about 0.045 to about 0.033 (inferred from the rates).

    What is left: mostly the returned rows themselves (about 23 KB for 100 rows), the stock JSON decoder's own buffer, and the returned export payload. Removing those needs an API or document-shape change, not a binding fix.

    The second fix merges only after CI runs its new concurrency tests under the race detector. Both fixes then ship in Go 1.0.2. This issue stays open for the remaining gap and for whether the other bindings' runtimes show the same effect.

    🤖 Generated with Claude Code

  17. added a commit that references this issue on Oct 6, 2026
  18. EricAndrechek commented on Oct 6, 2026

    @EricAndrechek
    MemberAuthor

    Go 1.0.2 is released (the Go proxy v1.0.2, 2026-10-06) with both binding fixes. Upgrade to get them; no code change is needed.

    What's in it:

    • Rows decodes its result in one streaming pass;
    • most per-call allocations around the call are gone.
    • On the reporter's 100-row shape, a full call went from 412 allocations / 75 KB to 76 / 45 KB.

    The reporter's measurement (8 goroutines against 1, default GC): about 5.15×, up from about 2.1× on Go 1.0.1. The ParseBlock control reaches about 7×.

    This issue stays open for the rest of the gap and for the other bindings:

    • What remains per call is mostly the returned rows themselves, the stock JSON decoder's buffer and the returned export payload.
    • Python, TypeScript and Rust also decode each batch into a generic tree. Whether their runtimes show a similar effect is not measured yet.

    Until then, for Go: a process with a larger live heap, or a higher GOGC (or GOGC=off with a sized GOMEMLIMIT, as measured above), runs fewer GC cycles per call.

    🤖 Generated with Claude Code

  19. changed the title [-]Known 1.0 limitation (library fixed in build 20261004.052404; Go binding fix coming in 1.0.x): Schema.Rows scaling across goroutines[/-] [+]Known 1.0 limitation (library fixed; Go improved in 1.0.2; remaining gap open): Schema.Rows scaling across goroutines[/+] on Oct 6, 2026
  20. EricAndrechek commented on Oct 6, 2026

    @EricAndrechek
    MemberAuthor

    Scaling measurement on GitHub-hosted Linux runners (standard 4-vCPU), same shape as the earlier runs: one table, 100-row body, the 4 settings, a tenant filter, JSONCompactEachRow export, WithDocFlags(0), default GC. Each point is a 2 s warm-up then a 5 s window. Run once.

    • chtypes go v1.0.2, artifact build 20261004.052404 (26.8.15.10; the manifest in the fetched cache contains that build string), Go 1.27.1, nproc = 4, GOMAXPROCS = 4.
    • ubuntu-24.04: linux-amd64, AMD EPYC 7763.
    • ubuntu-24.04-arm: linux-arm64, Neoverse-N2.
    • Run: https://github.com/Wave-RF/WaveHouse/actions/runs/37459986581

    Own handles unless noted. Ratio is calls/s relative to g=1 for the same work on the same runner. CPU/wall is process CPU seconds per wall second. GC/s is completed GC cycles per second (the count gctrace=1 prints, taken from runtime.MemStats.NumGC).

    ubuntu-24.04 (x86_64, EPYC 7763)

    work g calls/s ratio CPU/wall GC/s
    Rows 1 400 1.00 1.04 5.4
    Rows 2 733 1.83 2.05 10.8
    Rows 4 842 2.10 3.82 13.2
    Rows 8 833 2.08 3.93 15.4
    Rows, shared handle 8 822 n/a 3.94 15.0
    ParseBlock 1 575 1.00 1.04 1.2
    ParseBlock 2 1118 1.95 2.08 2.6
    ParseBlock 4 1290 2.25 3.89 3.0
    ParseBlock 8 1276 2.22 3.96 3.0

    ubuntu-24.04-arm (aarch64, Neoverse-N2)

    work g calls/s ratio CPU/wall GC/s
    Rows 1 384 1.00 1.01 5.6
    Rows 2 741 1.93 2.01 11.0
    Rows 4 1387 3.61 3.76 23.4
    Rows 8 1379 3.59 3.84 27.0
    Rows, shared handle 8 1157 n/a 3.92 22.0
    ParseBlock 1 534 1.00 1.01 1.6
    ParseBlock 2 1045 1.96 2.01 2.4
    ParseBlock 4 1965 3.68 3.84 4.6
    ParseBlock 8 1960 3.67 3.95 4.6

    What it shows: on both runners Rows scales about as well as ParseBlock on the same machine (arm64 3.61x vs 3.68x at g=4, close to the 4-way ceiling; amd64 2.10x vs 2.25x, with the process using ~3.8 CPUs in both, so the cores are busy rather than waiting), so on 4 vCPUs we see no Rows-specific scaling gap, and GC rates here (5 to 27/s) are far below the ~80 to 160/s of the earlier 14-core runs. Whether the amd64 ceiling near 2x reflects SMT sharing of physical cores on that runner type is inferred, not measured, and this run does not test the 8-core-and-up regime where the gap showed.

  21. EricAndrechek commented on Oct 6, 2026

    @EricAndrechek
    MemberAuthor

    Closing: at 8 and 16 goroutines, Go Rows now scales nearly as well as ParseBlock. This is the ≥8-core datapoint the 4-vCPU run above could not give.

    Setup (measured). The reporter's exact work units:

    • one table, the 100-row body, the 4 settings, a tenant filter, JSONCompactEachRow export, WithDocFlags(0);
    • go v1.0.2, artifact build 20261004.052404 (26.8.15.10, read from the loaded library's build_info), Go 1.27.1, default GC;
    • own handles unless marked; each point is a 2 s warm-up then a 5 s window; run once.

    Two self-hosted Linux runners:

    • arm64: Neoverse-V2, 192 CPUs;
    • amd64: Xeon Platinum 8375C, 64 CPUs.

    GOMAXPROCS was 60 on both, so g=16 is not oversubscribed. The ratio is calls/s against g=1 for the same work.

    arm64 (Neoverse-V2)

    work g calls/s ratio CPU/wall GC/s
    Rows 1 391 1.00 1.04 8.8
    Rows 2 805 2.06 2.07 20.4
    Rows 4 1571 4.01 4.14 56.2
    Rows 8 2699 6.90 8.25 266.3
    Rows 16 4446 11.36 15.33 439.5
    Rows, shared handle 8 1103 2.82 8.13 83.2
    ParseBlock 1 575 1.00 1.01 2.0
    ParseBlock 2 1136 1.97 2.02 4.2
    ParseBlock 4 2274 3.95 4.04 10.0
    ParseBlock 8 4479 7.79 8.07 27.0
    ParseBlock 16 7559 13.15 16.09 241.1

    amd64 (Xeon 8375C)

    work g calls/s ratio CPU/wall GC/s
    Rows 1 392 1.00 1.03 8.0
    Rows 2 780 1.99 2.05 19.8
    Rows 4 1524 3.88 4.13 60.2
    Rows 8 2664 6.79 8.19 210.1
    Rows 16 4285 10.92 14.48 280.2
    Rows, shared handle 8 2385 6.08 8.25 183.7
    ParseBlock 1 573 1.00 1.02 2.0
    ParseBlock 2 1144 1.99 2.03 4.2
    ParseBlock 4 2282 3.98 4.05 10.4
    ParseBlock 8 4414 7.70 8.10 32.6
    ParseBlock 16 6859 11.96 15.92 293.1

    What it shows:

    • g=8 (measured): Rows reaches 6.90× on arm64 and 6.79× on amd64, against ParseBlock's 7.79× and 7.70×. At g=16 it reaches 11.4× and 10.9×, against 13.2× and 12.0×. CPU/wall tracks g, so the workers stay busy.
    • The original report's ~2× ceiling is gone at this core count. The earlier gap came from the library's shared state (fixed in the 20261004 build) and the Go decode allocations (cut in 1.0.2, from 412 to 76 per call).
    • What remains (inferred, not isolated): Rows runs at about 0.88× of ParseBlock's scaling at g=8, and that tracks its GC rate: 266 vs 27 GC/s on arm64. Rows still allocates per call where ParseBlock hands back a block. Cutting that further is a possible future Go optimisation, not part of this defect.

    One more observation, outside this issue's shape: a single shared handle at g=8 reaches only 2.82× on arm64, but 6.08× on amd64. It was measured once; own handles are fine on both. The artifact producer already tracks a library-side result of the same arm-vs-x86 shape for one shared handle, and this Go datapoint has been passed on there. It is below the 4× to 6.5× range that docs/reference/bindings-v1.md §3 quotes for a shared handle at 8 threads, so that range is due a recheck on arm64 once the library side is understood. The advice there stands, and is what this table measures: compile a schema per goroutine for the full speed-up.

    Closing as fixed by the 20261004 library build plus go v1.0.2.

    🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions