Skip to content

One shared handle should scale across threads like N separate handles #480

Description

@EricAndrechek

One shared handle should scale with threads like N separate handles do.

Today (docs/reference/bindings-v1.md §3): calls on one compiled handle are safe to run concurrently, but they contend inside the library. A handle per thread is the documented way to get the full parallel speed-up.

Measured: go v1.0.2 Schema.Rows at 8 goroutines (full tables on #456).

arm64 (Neoverse-V2) amd64 (Xeon 8375C)
one shared handle 2.82× 6.08×
own handle per goroutine 6.90× 6.79×

The artifact producer measured the same arm-vs-x86 shape library-side. The cause is contention on the shared schema's column objects in the row path: atomic refcounts on shared columns.

The ask: one shared handle scales close to N handles, so a consumer need not keep a schema per thread. One direction: a per-thread working copy of the per-call block instead of shared column references.

Also: docs/reference/bindings-v1.md §3 quotes "about 4× to 6.5×" for a handle shared by 8 threads. The arm64 Go figure above is below that range, so §3 is to be corrected once this is resolved or re-measured.

Needs a library change from the artifact producer.

🤖 Generated with Claude Code

Activity

  1. EricAndrechek commented on Oct 6, 2026

    @EricAndrechek
    MemberAuthor

    Progress: the artifact producer has merged a library fix: an internal per-schema replica pool, so concurrent calls on one shared handle no longer contend on the same compiled objects. It reports one shared handle at about 7.6–7.9× on 8 threads vs 1 (one handle per thread: 7.7–7.9×), on arm64 and amd64 (measured there).

    It is not in a production build yet. When it ships, this repository will:

    • re-measure through the bindings;
    • update docs/reference/bindings-v1.md §3's scaling figures, plus an advanced knob to cap the pool;
    • close this issue with the table.
  2. EricAndrechek commented on Oct 6, 2026

    @EricAndrechek
    MemberAuthor

    Fixed in production library build 20261006.170903. Re-measured by this repository. We confirmed the build from the library's own build_info. Setup: Go binding 1.0.4, Rows on 100-row batches, default GC, on a 64-core arm64 runner and a 192-core amd64 runner. Throughput at eight goroutines against one:

    one shared schema one schema per goroutine shared, CHTYPES_V1_SCHEMA_REPLICAS=1 (pool off)
    arm64 7.55× 7.59× 4.75×
    amd64 6.75× 6.65× 3.78×

    A shared schema now scales like one schema per thread. The library's internal replica pool is what did it: the REPLICAS=1 control behaves the way earlier builds did. The docs are in #517, which has landed: docs/reference/bindings-v1.md §3, including the provisional CHTYPES_V1_SCHEMA_REPLICAS knob. Closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions