Skip to content

Partial-update merging compares Avro schemas field by field for every record #20268

Description

@yihua

Reading or compacting MOR file groups that carry partial-update log records (from a partial MERGE INTO) spends most of the per-record merge cost comparing Avro schemas: 16% of executor CPU and 47% of allocation in a 6M-row Spark SQL run. SparkRecordMergingUtils builds the merged schema of a full-plus-partial merge by converting the Spark struct type back to a HoodieSchema, which drops field defaults, so the result is unequal to the reader schema but hashes the same. Every schema-keyed lookup holding both then compares the two field by field, once per merged record.

Proposal: reuse the reader schema as the merged schema when the merge covers every reader field.

Activity

  1. yihua commented on Oct 9, 2026

    @yihua
    ContributorAuthor

    Fixed by #20269.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:performancePerformance optimizationsarea:readerReader core functionalitypriority:mediumModerate impact; usability gapstype:improvementImprovements to existing functionality

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions