Skip to content

fix(spark): use an expression index only for filters on its own function - #20256

Draft
yihua wants to merge 1 commit into
apache:masterfrom
yihua:fix-expression-index-function-match
Draft

yihua wants to merge 1 commit into
apache:masterfrom
yihua:fix-expression-index-function-match

Conversation

@yihua

@yihua yihua commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Describe the issue this Pull Request addresses

closes #20255

Data skipping used an expression index for any unary function over the index's source column. With an index on lower(city), a filter upper(city) = 'X' was evaluated against the lower values and pruned every file holding matches (wrong result), and length(city) = 9 failed with a ClassCastException. The catch-all case expression: UnaryExpression => expression.child came in with #12290 while a separate function-name check still guarded it; #12455 replaced that check with per-function option checks and dropped the function-name comparison, leaving the catch-all unguarded.

Summary and Changelog

Each match arm in ExpressionIndexSupport.extractQueryAndLiterals now also requires the index definition's function (upper, lower, length, year, month, day, hour, plus the existing format-based functions; to_date / to_timestamp without a format, which the optimizer turns into a Cast, match only their own index). Filters on other functions fall back to column stats. TestExpressionIndex adds a COW and MOR (with log files) test that upper and length filters return the same rows as with data skipping off (wrong / failing on master), and a test that filters on the indexed function still prune for upper, length, year, hour and to_date.

Related: #17489 (an expression index matched by a direct column reference) touches the same method; this change does not cover that case.

Impact

Fixes wrong query results and query failures for tables with an expression index on a function of a column queried through a different function. Matches on the indexed function are unchanged.

Risk Level

low. With spark.sql.timestampType=TIMESTAMP_NTZ, a to_timestamp(col) filter (no format) becomes a cast to TimestampNTZType and no longer uses a to_timestamp expression index: results stay correct, only pruning is lost.

Documentation Update

none

Contributor's checklist

  • Read through contributor's guide
  • Enough context is provided in the sections above
  • Adequate tests were added if applicable

Data skipping unwrapped any unary expression over an expression index's
source column, so a filter such as upper(city) = 'X' was evaluated against
a lower(city) index (wrong result) and length(city) = 9 failed with a
ClassCastException. Each match in ExpressionIndexSupport now also requires
the index definition's function, and other filters fall back to the next
index.
@codecov-commenter

codecov-commenter commented Oct 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 12.90323% with 27 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.62%. Comparing base (308f78e) to head (c6091c3).

Files with missing lines Patch % Lines
...scala/org/apache/hudi/ExpressionIndexSupport.scala 12.90% 1 Missing and 26 partials ⚠️
Additional details and impacted files
@@            Coverage Diff            @@
##             master   #20256   +/-   ##
=========================================
  Coverage     80.62%   80.62%           
- Complexity    35055    35064    +9     
=========================================
  Files          2552     2552           
  Lines        143400   143411   +11     
  Branches      17454    17464   +10     
=========================================
+ Hits         115615   115624    +9     
+ Misses        19867    19859    -8     
- Partials       7918     7928   +10     
Components Coverage Δ
hudi-common 84.16% <ø> (ø)
hudi-client 83.82% <ø> (+<0.01%) ⬆️
hudi-flink 85.97% <ø> (+<0.01%) ⬆️
hudi-spark-datasource 74.03% <12.90%> (-0.03%) ⬇️
hudi-utilities 78.16% <ø> (-0.03%) ⬇️
hudi-cli 70.80% <ø> (ø)
hudi-hadoop 71.12% <ø> (+0.01%) ⬆️
hudi-sync 76.25% <ø> (+0.07%) ⬆️
hudi-io 81.52% <ø> (+0.09%) ⬆️
hudi-timeline-service 83.41% <ø> (ø)
hudi-cloud 81.00% <ø> (ø)
hudi-kafka-connect 53.20% <ø> (ø)
Flag Coverage Δ
common-and-other-modules 52.42% <0.00%> (+<0.01%) ⬆️
flink-integration-tests 49.54% <ø> (+<0.01%) ⬆️
hadoop-mr-java-client 43.98% <ø> (ø)
integration-tests 13.45% <0.00%> (-0.01%) ⬇️
spark-client-hadoop-common 38.69% <0.00%> (-0.01%) ⬇️
spark-java-tests 52.66% <0.00%> (-0.02%) ⬇️
spark-scala-tests 47.51% <12.90%> (-0.01%) ⬇️
utilities 36.88% <0.00%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...scala/org/apache/hudi/ExpressionIndexSupport.scala 66.31% <12.90%> (-1.94%) ⬇️

... and 22 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M PR with lines of changes in (100, 300]

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Expression index is used for filters on a different function of the same column, returning wrong results

2 participants