Reading or compacting MOR file groups that carry partial-update log records (from a partial MERGE INTO) spends most of the per-record merge cost comparing Avro schemas: 16% of executor CPU and 47% of allocation in a 6M-row Spark SQL run. SparkRecordMergingUtils builds the merged schema of a full-plus-partial merge by converting the Spark struct type back to a HoodieSchema, which drops field defaults, so the result is unequal to the reader schema but hashes the same. Every schema-keyed lookup holding both then compares the two field by field, once per merged record.
Proposal: reuse the reader schema as the merged schema when the merge covers every reader field.
Reading or compacting MOR file groups that carry partial-update log records (from a partial
MERGE INTO) spends most of the per-record merge cost comparing Avro schemas: 16% of executor CPU and 47% of allocation in a 6M-row Spark SQL run.SparkRecordMergingUtilsbuilds the merged schema of a full-plus-partial merge by converting the Spark struct type back to aHoodieSchema, which drops field defaults, so the result is unequal to the reader schema but hashes the same. Every schema-keyed lookup holding both then compares the two field by field, once per merged record.Proposal: reuse the reader schema as the merged schema when the merge covers every reader field.