Skip to content

Use unbounded lead cost for prohibited clauses - #16715

Open
kkewwei wants to merge 7 commits into
apache:mainfrom
kkewwei:fix_mostnot_cost
Open

kkewwei wants to merge 7 commits into
apache:mainfrom
kkewwei:fix_mostnot_cost

Conversation

@kkewwei

@kkewwei kkewwei commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Description

BooleanScorerSupplier#bulkScorer used the positive scorer's cost as the lead cost for prohibited clauses. For an IndexOrDocValuesQuery , this could incorrectly select the doc-values implementation as the leader in ReqExclBulkScorer, With the doc-values implementation, which can require scanning all docs in the segment.

This change uses an unbounded lead cost for prohibited clauses, allowing IndexOrDocValuesQuery to select its index implementation for ReqExclBulkScorer.

Resolved #16702

@kkewwei

kkewwei commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

Since there was no existing task covering this case, I addedMostNotIODVRange.
https://github.com/mikemccand/luceneutil/compare/main...kkewwei:must_not_task?expand=1

I added a benchmark task set in luceneutil with wikimediumall, it shows an approximately 3x performance improvement for the query, with no performance impact on related tasks.

                            TaskQPS baseline      StdDevQPS my_modified_version      StdDev                Pct diff p-value
           BrowseMonthSSDVFacets        3.49     (24.6%)        3.40     (21.3%)   -2.7% ( -39% -   57%) 0.706
     BrowseRandomLabelTaxoFacets        2.26      (3.2%)        2.20      (3.5%)   -2.7% (  -9% -    4%) 0.012
       BrowseDayOfYearTaxoFacets        3.57     (10.2%)        3.49      (8.1%)   -2.1% ( -18% -   18%) 0.466
            BrowseDateTaxoFacets        3.48     (10.0%)        3.41      (8.2%)   -2.1% ( -18% -   17%) 0.473
     BrowseRandomLabelSSDVFacets        1.84      (9.8%)        1.81      (8.8%)   -1.9% ( -18% -   18%) 0.510
                          IntSet      255.29      (3.3%)      250.56      (3.3%)   -1.9% (  -8% -    4%) 0.075
            BrowseDateSSDVFacets        0.74      (9.1%)        0.73     (10.0%)   -1.2% ( -18% -   19%) 0.683
                     MedSpanNear       16.61      (1.7%)       16.46      (2.3%)   -0.9% (  -4% -    3%) 0.138
       BrowseDayOfYearSSDVFacets        3.20     (14.2%)        3.18     (20.2%)   -0.7% ( -30% -   39%) 0.902
        AndHighHighDayTaxoFacets        9.31      (5.7%)        9.25      (6.4%)   -0.6% ( -12% -   12%) 0.746
                    HighSpanNear        5.81      (2.5%)        5.77      (2.8%)   -0.6% (  -5% -    4%) 0.487
            HighTermTitleBDVSort        2.36      (9.5%)        2.35     (10.2%)   -0.4% ( -18% -   21%) 0.893
             LowIntervalsOrdered       32.26      (3.1%)       32.14      (6.5%)   -0.4% (  -9% -    9%) 0.810
            HighIntervalsOrdered        2.15      (5.9%)        2.15      (6.6%)   -0.2% ( -11% -   12%) 0.904
             MedIntervalsOrdered       16.45      (3.4%)       16.42      (4.1%)   -0.2% (  -7% -    7%) 0.860
                        BM25MSM2        3.16      (2.1%)        3.15      (3.1%)   -0.1% (  -5% -    5%) 0.873
                      AndHighMed      114.47      (2.4%)      114.36      (2.3%)   -0.1% (  -4% -    4%) 0.892
                       MedPhrase        6.33      (2.9%)        6.33      (4.4%)   -0.0% (  -7% -    7%) 0.999
                        PKLookup      166.12      (3.9%)      166.16      (4.2%)    0.0% (  -7% -    8%) 0.984
                     LowSpanNear        2.18      (1.8%)        2.18      (4.1%)    0.0% (  -5% -    6%) 0.965
               HighTermTitleSort       72.70      (3.3%)       72.77      (3.9%)    0.1% (  -6% -    7%) 0.933
                       OrHighMed       44.11      (1.9%)       44.16      (2.6%)    0.1% (  -4% -    4%) 0.877
                 LowSloppyPhrase        4.50      (3.5%)        4.51      (3.4%)    0.2% (  -6% -    7%) 0.878
                      HighPhrase       13.63      (2.4%)       13.66      (3.7%)    0.2% (  -5% -    6%) 0.812
                      AndHighLow      330.10      (5.8%)      330.89      (6.3%)    0.2% ( -11% -   13%) 0.901
         AndHighMedDayTaxoFacets       15.26      (5.1%)       15.30      (5.4%)    0.3% (  -9% -   11%) 0.858
          OrHighMedDayTaxoFacets        2.87      (3.5%)        2.88      (3.9%)    0.3% (  -6% -    7%) 0.778
                  MostNotLowHigh       43.05      (3.5%)       43.22      (4.4%)    0.4% (  -7% -    8%) 0.750
                         LowTerm      612.85      (2.5%)      615.33      (2.6%)    0.4% (  -4% -    5%) 0.621
                 MedSloppyPhrase       19.27      (4.0%)       19.36      (3.7%)    0.4% (  -7% -    8%) 0.724
                        HighTerm      393.62      (2.6%)      395.45      (3.5%)    0.5% (  -5% -    6%) 0.634
                        Wildcard        9.66      (2.8%)        9.71      (2.8%)    0.5% (  -5% -    6%) 0.594
                         MedTerm      438.19      (3.8%)      440.52      (4.3%)    0.5% (  -7% -    9%) 0.681
                       LowPhrase       46.58      (3.2%)       46.84      (3.5%)    0.6% (  -5% -    7%) 0.596
                  AndMissingHigh     2076.32      (4.0%)     2088.41      (4.0%)    0.6% (  -7% -    8%) 0.641
                HighSloppyPhrase        2.16      (4.9%)        2.18      (5.2%)    0.6% (  -9% -   11%) 0.705
                          Fuzzy1       37.40      (2.8%)       37.63      (3.0%)    0.6% (  -5% -    6%) 0.498
                    OrNotHighLow      324.53      (5.4%)      326.63      (5.7%)    0.6% (  -9% -   12%) 0.711
                          Fuzzy2       36.93      (3.2%)       37.18      (3.6%)    0.7% (  -5% -    7%) 0.527
                     AndHighHigh       24.50      (2.2%)       24.68      (2.6%)    0.8% (  -3% -    5%) 0.323
                      OrHighHigh       33.84      (1.9%)       34.12      (2.2%)    0.8% (  -3% -    5%) 0.211
                       OrHighLow      312.47      (3.2%)      315.16      (3.9%)    0.9% (  -6% -    8%) 0.444
                       ConstMSM2      966.22      (3.1%)      974.68      (3.7%)    0.9% (  -5% -    7%) 0.413
                         Prefix3      143.63      (1.9%)      145.01      (2.4%)    1.0% (  -3% -    5%) 0.166
                           range     4588.95      (6.1%)     4641.12      (5.9%)    1.1% ( -10% -   13%) 0.549
                         Respell       29.06      (2.9%)       29.39      (3.9%)    1.1% (  -5% -    8%) 0.294
               HighTermMonthSort     1582.51      (3.6%)     1607.89      (3.5%)    1.6% (  -5% -    9%) 0.154
                    OrNotHighMed      217.21      (4.2%)      220.74      (2.9%)    1.6% (  -5% -    9%) 0.156
            MedTermDayTaxoFacets        5.59      (4.5%)        5.69      (3.3%)    1.8% (  -5% -    9%) 0.151
                          IntNRQ      272.10      (3.7%)      279.02      (5.8%)    2.5% (  -6% -   12%) 0.097
           HighTermDayOfYearSort      159.11      (3.5%)      163.25      (3.1%)    2.6% (  -3% -    9%) 0.013
                      TermDTSort      120.11      (3.2%)      123.88      (3.6%)    3.1% (  -3% -   10%) 0.003
                   OrHighNotHigh      277.43      (6.1%)      288.03      (7.7%)    3.8% (  -9% -   18%) 0.082
                   OrNotHighHigh      165.88      (8.3%)      172.39      (7.4%)    3.9% ( -10% -   21%) 0.114
                    OrHighNotLow      226.74      (8.5%)      238.17      (7.9%)    5.0% ( -10% -   23%) 0.052
                    OrHighNotMed      184.67      (8.9%)      194.76      (9.9%)    5.5% ( -12% -   26%) 0.067
           BrowseMonthTaxoFacets        4.18     (43.8%)        4.79     (59.8%)   14.5% ( -61% -  210%) 0.382
                MostNotIODVRange        6.91     (39.0%)       27.56      (8.1%)  298.8% ( 181% -  566%) 0.000

@sgup432 sgup432 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see the core issue is in ReqExclBulkScorer, it always leads with the exclScorer, so when the positive side is very selective, picking the dvScorer for exclusion seems expensive.

I wonder in that selective case, can we fallback to ReqExclScorer? Seems like it leads with the required scorer and checks exclusion per candidate via exclScorer.matches(). When the lead is tiny, it is just O(leadCost) and likely cheaper than building a points scorer, which also costs extra memory for the bitset.

@kkewwei

kkewwei commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

@sgup432 That makes a lot of sense. When the required cost is much lower than the exclusion cost, usingReqExclScorer may be more appropriate.

I’ll add benchmarks with sparse required of 5%, 1%, 0.1%, 0.01%, 0.001% to find the threshold whereReqExclScorer becomes faster. I’ll test Term, BKD, andIndexOrDocValuesQuery required clauses separately, since the threshold may vary by operator.

If required is not very spare, ReqExclBulkScorer may be more appropriate, Since in ReqExclBulkScorer we just prebuild the acceptDocs for required BulkScorer.

@kkewwei

kkewwei commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

I add the benchmark(SparseRequiredMustNotBenchmark) for the spare required. The results are as follows:

Required term (ReqExclBulkScorer / ReqExclScorer) diff point (ReqExclBulkScorer / ReqExclScorer) diff IODV (ReqExclBulkScorer / ReqExclScorer) diff
10% 267 / 72 +272% 672 / 53 +1176% 649 / 51 +1163%
5% 356 / 126 +182% 394 / 104 +278% 388 / 102 +281%
1% 789 / 434 +82% 832 / 360 +131% 804 / 360 +123%
0.5% 855 / 632 +35% 883 / 533 +66% 892 / 517 +73%
0.1% 1074 / 1916 -44% 1079 / 2560 -58% 983 / 2069 -52%
0.01% 2031 / 19434 -90% 2250 / 17931 -87% 2332 / 18179 -87%
0.001% 16499 / 68120 -76% 16586 / 58921 -72% 16400 / 62531 -74%

There is a clear threshold between 0.1% and 0.5% from the benchmark:

  • When the required match rate is at least 0.5%,ReqExclBulkScorer is faster.

  • When the required match rate is at most 0.1%,ReqExclScorer is faster.

I will continue benchmarking 0.2%, 0.3%, and 0.4% to determine the best threshold.

@kkewwei

kkewwei commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

@sgup432 The results are as follows:

Required term (ReqExclBulkScorer / ReqExclScorer) diff point (ReqExclBulkScorer / ReqExclScorer) diff IODV (ReqExclBulkScorer / ReqExclScorer) diff
0.2% 908 / 1513 -39.9% 935 / 1363 -31.4% 1038 / 1382 -24.9%
0.3% 876 / 1244 -29.6% 923 / 1070 -13.7% 890 / 1098 -18.9%
0.4% 959 / 644 +48.9% 887 / 601 +47.5% 877 / 590 +48.8%

Under the current benchmark data distribution, there is a clear performance crossover between 0.3% and 0.4%:

  • When the required match rate is at most 0.3%,ReqExclScorer is faster.

  • When the required match rate reaches 0.4%,ReqExclBulkScorer is approximately 48% faster.

I will use 0.3% as the threshold for choosing betweenReqExclBulkScorer andReqExclScorer .

@sgup432

sgup432 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

@kkewwei Thanks for the numbers.
One clarification for the numbers in the table, is the required X% data represents the cost of the positive scorer in %?
Also I believe the number for respective scorer in represents the throughput? Higher, the better?

@sgup432 sgup432 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM. Some minor comments.

Comment thread lucene/core/src/java/org/apache/lucene/search/BooleanScorerSupplier.java Outdated
Comment thread lucene/CHANGES.txt Outdated
@kkewwei

kkewwei commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

One clarification for the numbers in the table, is the required X% data represents the cost of the positive scorer in %?
Also I believe the number for respective scorer in represents the throughput? Higher, the better?

@sgup432 Very thank you for the detail review.

Yes. The Required X% column(positiveScorer.cost() /maxDoc) represents the percentage of documents matched by the required/positive query.

The scorer numbers are throughput in operations per second (ops/s), higher values are better.

@sgup432

sgup432 commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

@romseygeek Do you think you can help with this one?

@sgup432

sgup432 commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

@kkewwei Also would be good to contribute your custom task MostNotIODVRange to luceneutil so that such queries get tracked. 🙂

@romseygeek romseygeek left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @kkewwei - I have one small suggestion before I merge.

Comment thread lucene/core/src/java/org/apache/lucene/search/BooleanScorerSupplier.java Outdated
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MUST_NOT + IndexOrDocValuesQuery: much slower on 10.x than 9.x/8.x

3 participants