Skip to content

Data skipping returns too few rows on a METADATA_ONLY bootstrapped table once column stats are enabled #20266

Description

@yihua

Bug Description

What happened:

After a METADATA_ONLY bootstrap, a Spark snapshot query that filters on a data column silently returns too few rows (often zero) once the column stats index is enabled on the table. That happens on the first later write with Spark's defaults (column stats and partition stats on), and data skipping is on by default, so no special setting is needed.

The column stats of a bootstrapped file group come from its skeleton file, which only has the Hudi meta columns. Partition stats pruning keeps only partitions with a matching stats row, so a fully bootstrapped partition is pruned and a partition with bootstrapped and rewritten file groups is pruned by the rewritten files' value range alone. Column stats pruning keeps a file with no stats only if it has no record for any filtered column, so a filter that also references a meta column (_hoodie_record_key = 'trip_7' and rider = 'rider_7') prunes the skeleton file. Flink's FileStatsIndex and PartitionStatsIndex likely have the same gap (not verified).

What you expected:

The same rows with data skipping on and off.

Steps to reproduce:

  1. Write a partitioned parquet dataset (_row_key, rider, driver, timestamp, datestr, ...) to srcPath.
  2. Bootstrap it in METADATA_ONLY mode (the default selector) with hoodie.datasource.write.operation=bootstrap, hoodie.bootstrap.base.path=srcPath, record key _row_key, partition path datestr, and hoodie.metadata.index.column.stats.enable=false.
  3. Upsert a few rows with the Spark defaults, which builds column stats and partition stats over the skeleton files.
  4. Compare spark.read.format("hudi").load(basePath).filter("driver >= 'driver_5'").count() with the same query and hoodie.enable.data.skipping=false: 18 vs 56 rows.
  5. With partition stats disabled, single data-column filters are correct, but _hoodie_commit_time = '00000000000001' and driver >= 'driver_5' returns 1 row instead of 55. COW and MOR behave the same.

Environment

Hudi version: master (1.3.0-SNAPSHOT); reachable by default since 1.0.1
Query engine: Spark 3.5
Relevant configs: hoodie.metadata.index.column.stats.enable=true, hoodie.metadata.index.partition.stats.enable=true, hoodie.enable.data.skipping=true (all defaults), METADATA_ONLY bootstrap

Logs and Stack Trace

No response

Parent issue: #20257

Fixed by: #20267

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions