Repository navigation
Conversation
Data skipping unwrapped any unary expression over an expression index's source column, so a filter such as upper(city) = 'X' was evaluated against a lower(city) index (wrong result) and length(city) = 9 failed with a ClassCastException. Each match in ExpressionIndexSupport now also requires the index definition's function, and other filters fall back to the next index.
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #20256 +/- ##
=========================================
Coverage 80.62% 80.62%
- Complexity 35055 35064 +9
=========================================
Files 2552 2552
Lines 143400 143411 +11
Branches 17454 17464 +10
=========================================
+ Hits 115615 115624 +9
+ Misses 19867 19859 -8
- Partials 7918 7928 +10
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Describe the issue this Pull Request addresses
closes #20255
Data skipping used an expression index for any unary function over the index's source column. With an index on
lower(city), a filterupper(city) = 'X'was evaluated against thelowervalues and pruned every file holding matches (wrong result), andlength(city) = 9failed with aClassCastException. The catch-allcase expression: UnaryExpression => expression.childcame in with #12290 while a separate function-name check still guarded it; #12455 replaced that check with per-function option checks and dropped the function-name comparison, leaving the catch-all unguarded.Summary and Changelog
Each match arm in
ExpressionIndexSupport.extractQueryAndLiteralsnow also requires the index definition's function (upper,lower,length,year,month,day,hour, plus the existing format-based functions;to_date/to_timestampwithout a format, which the optimizer turns into aCast, match only their own index). Filters on other functions fall back to column stats.TestExpressionIndexadds a COW and MOR (with log files) test thatupperandlengthfilters return the same rows as with data skipping off (wrong / failing on master), and a test that filters on the indexed function still prune forupper,length,year,hourandto_date.Related: #17489 (an expression index matched by a direct column reference) touches the same method; this change does not cover that case.
Impact
Fixes wrong query results and query failures for tables with an expression index on a function of a column queried through a different function. Matches on the indexed function are unchanged.
Risk Level
low. With
spark.sql.timestampType=TIMESTAMP_NTZ, ato_timestamp(col)filter (no format) becomes a cast toTimestampNTZTypeand no longer uses ato_timestampexpression index: results stay correct, only pruning is lost.Documentation Update
none
Contributor's checklist