Skip to content

perf: Adding support for LatestBaseFilesPathFilter to Spark File Index - #18136

Merged
nsivabalan merged 7 commits into
apache:masterfrom
suryaprasanna:path-filter-during-listing
Feb 26, 2026
Merged

nsivabalan merged 7 commits into
apache:masterfrom
suryaprasanna:path-filter-during-listing

Conversation

@suryaprasanna

@suryaprasanna suryaprasanna commented Feb 8, 2026 •

Copy link
Copy Markdown
Contributor

Describe the issue this Pull Request addresses

This PR adds an opt-in path filtering mechanism during file listing to prevent Spark driver OOM errors when querying large Hudi datasets with multiple file versions per partition.

Problem: When file listing is performed without filtering, all file versions (including older ones) are loaded into driver memory, causing OOM on large tables.

Solution: Added a new config hoodie.datasource.read.file.index.list.file.statuses.using.ro.path.filter (default: false) that enables HoodieROPathFilter during file listing to exclude older file versions.

Summary and Changelog

Users can now enable path filtering during file listing to avoid loading multiple file versions into memory on the driver. This is controlled by the new config hoodie.datasource.read.file.index.list.file.statuses.using.ro.path.filter.

Changes:

  • New Config: FILE_INDEX_LIST_FILE_STATUSES_USING_RO_PATH_FILTER (default: false)

    • Enables path filtering during file listing to reduce driver memory pressure
    • Filters out older file versions, keeping only the latest files needed for queries
  • API Extensions:

    • Extended HoodieTableMetadata, BaseTableMetadata, and FileSystemBackedTableMetadata to accept optional StoragePathFilter parameter
    • Added FSUtils.getAllDataFilesInPartition overload with path filter support
    • Created HoodieROTableStoragePathFilter wrapper to adapt Hadoop PathFilter to Hudi's StoragePathFilter interface
  • Spark Integration:

    • Updated BaseHoodieTableFileIndex to use path filter when enabled
    • Modified SparkHoodieTableFileIndex to apply HoodieROTablePathFilter during partition listing

Impact

Config Changes:

  • New config: hoodie.datasource.read.file.index.list.file.statuses.using.ro.path.filter (default: false)
    • When enabled, applies HoodieROTablePathFilter during file listing to filter out older file versions
    • Helps prevent OOM issues on driver for large tables with multiple file versions

Performance:

  • Reduces memory pressure on Spark driver for datasets with multiple file versions per partition
  • Enables successful queries on very large tables that previously failed with OOM errors

Risk Level

Low - The feature is behind a config flag (default: false) and does not change existing behavior unless explicitly enabled.

Verification:

  • Existing unit tests pass
  • Tested on large production datasets with OOM issues - queries now succeed with filter enabled

Documentation Update

Config documentation is included in the withDocumentation method of the new config property.

Contributor's checklist

  • Read through contributor's guide
  • Enough context is provided in the sections above
  • Adequate tests were added if applicable

@github-actions github-actions Bot added the size:M PR with lines of changes in (100, 300] label Feb 8, 2026
@suryaprasanna
suryaprasanna force-pushed the path-filter-during-listing branch 2 times, most recently from 95fda80 to 2dabf5d Compare February 9, 2026 16:16
@suryaprasanna

Copy link
Copy Markdown
Contributor Author

@nsivabalan seems like the checks are not running on the PR. Can you please chek?

@suryaprasanna
suryaprasanna force-pushed the path-filter-during-listing branch from 2dabf5d to 9b8395e Compare February 9, 2026 18:23
@github-actions github-actions Bot added size:L PR with lines of changes in (300, 1000] and removed size:M PR with lines of changes in (100, 300] labels Feb 9, 2026
@nsivabalan

Copy link
Copy Markdown
Contributor

can we consider both table types, and all query types and ensure we wire in the config only wherever applicable.

@apache apache deleted a comment from hudi-bot Feb 10, 2026
@@ -143,6 +146,7 @@ public abstract class BaseHoodieTableFileIndex implements AutoCloseable {
* @param configProperties unifying configuration (in the form of generic properties)
* @param queryType target query type
* @param queryPaths target DFS paths being queried

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add MOR condition.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added MOR table type condition.

.fromProperties(configProperties)
.enable(configProperties.getBoolean(ENABLE.key(), DEFAULT_METADATA_ENABLE_FOR_READERS)
&& HoodieTableMetadataUtil.isFilesPartitionAvailable(metaClient))
&& HoodieTableMetadataUtil.isFilesPartitionAvailable(metaClient)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hey @yihua : lets chat about this as well tomorrow.


if (useROPathFilterForListing && !shouldIncludePendingCommits) {
// Group files by partition path, then by file group ID
Map<String, PartitionPath> partitionsMap = new HashMap<>();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we move this to a private method.
generatePartitionFileSlicesPostROTablePathFilter

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

* By passing metaClient and completedTimeline, we can sync the view seen from this class against HoodieFileIndex class
*/
public HoodieROTablePathFilter(Configuration conf,
public HoodieROTablePathFilter(StorageConfiguration conf,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hey @yihua : can you review the changes in this patch

" them (if possible).")

val FILE_INDEX_LIST_FILE_STATUSES_USING_RO_PATH_FILTER: ConfigProperty[Boolean] =
ConfigProperty.key("hoodie.datasource.read.file.index.list.file.statuses.using.ro.path.filter")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hoodie.datasource.read.file.index.optimize.listing.using.path.filter

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Made the change.

properties.setProperty(DataSourceReadOptions.FILE_INDEX_LISTING_MODE_OVERRIDE.key, listingModeOverride)
}

var hoodieROTablePathFilterBasedFileListingEnabled = getConfigValue(options, sqlConf,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

once we fix the config key, lets fix these vars as well

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, updated.

val result = spark.sql(s"select id, name, price, ts from $tableName order by id").collect()
// Should have deleted records where id % 3 = 0 (3, 6, 9)
// Should have doubled price for even ids (2, 4, 8, 10)
assert(result.length == 7) // 10 - 3 deleted = 7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do you think below assertion would work.

we can rename one of the earlier versions of a file slice so that HoodieBaseFile parsing will fail.
so, if RO table path filter works as intended, listing files from a given partition should not fail, since we won't even try to parse the file.

but if RO table path filter did not work, it would fail.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@suryaprasanna : did you get a chance to address this?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

HoodieROTablePathFilter also creates FSV, so renaming the older file slice will also fail the latestBaseFiles API right?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha. sg.

@yihua yihua left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice work on adding an opt-in path filter to avoid driver OOM during file listing — the feature is well-motivated and the config is cleanly gated behind a default-off flag. The main concerns are correctness issues in the new loadFileSlicesForPartitions fast path: the partition path passed to FileSlice appears to be absolute rather than relative, and the partition map lookup can NPE if the key doesn't match. It's also worth clarifying MOR table compatibility and fixing the shared Hadoop config mutation in getPartitionPathFilter before merging.

Map<String, PartitionPath> partitionsMap = new HashMap<>();
partitions.forEach(p -> partitionsMap.put(p.path, p));
Map<PartitionPath, List<FileSlice>> partitionToFileSlices = new HashMap<>();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The partitionPathStr here is the absolute path (pathInfo.getPath().getParent().toString()), but FileSlice expects a relative partition path. The existing code path via HoodieTableFileSystemView always uses relative paths. This would cause mismatches downstream wherever FileSlice.getPartitionPath() is used. Should this be relPartitionPath instead?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@yihua PartitionPath object stores only relative partition path str, not absolute paths. Which code path are you referring it?

// Create FileSlice obj from StoragePathInfo.
String partitionPathStr = pathInfo.getPath().getParent().toString();
String relPartitionPath = FSUtils.getRelativePartitionPath(basePath, pathInfo.getPath().getParent());
HoodieBaseFile baseFile = new HoodieBaseFile(pathInfo);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If relPartitionPath doesn't exactly match a key in partitionsMap, partitionPathObj will be null and the computeIfAbsent call below will throw NPE. This could happen with path normalization differences (trailing slashes, scheme differences). Could you add a null check or use getRelativePartitionPath consistently with how PartitionPath.path was originally set?

List<StoragePathInfo> allFiles = listPartitionPathFiles(partitions, activeTimeline);
log.info("On {} with query instant as {}, it took {}ms to list all files {} Hudi partitions",
metaClient.getTableConfig().getTableName(), queryInstant.map(instant -> instant).orElse("N/A"),
timer.endTimer(), partitions.size());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Have you considered what happens with MOR tables here? The HoodieROTablePathFilter only returns base files (it calls fsView.getLatestBaseFiles()), so this path constructs FileSlices without log files. The !shouldIncludePendingCommits guard doesn't prevent MOR tables from reaching this code. It might be worth adding a table-type check (COW only) or documenting this limitation.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added table type equals COW condition.

@@ -146,6 +147,12 @@ public List<StoragePathInfo> getAllFilesInPartition(StoragePath partitionPath) t
}

@Override

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like the @Override annotation that belonged to getAllFilesInPartitions(Collection<String>) has been absorbed by the new method insertion. In the diff, the @Override on line 148 now applies to the new two-arg overload, while the original single-arg method (which is the actual interface abstract method) loses its @Override. Could you add @Override back to the original method?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes my bad, it is a refactoring mistake, fixed it now.

return getAllFilesInPartitions(partitions);
}

public Map<String, List<StoragePathInfo>> getAllFilesInPartitions(Collection<String> partitions)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Let this call getAllFilesInPartitions(partitions, Option.empty()) to be easier to read?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, makes sense. Refactored the code accordingly.

Comment on lines +157 to +164
Map<String, List<StoragePathInfo>> getAllFilesInPartitions(Collection<String> partitionPaths)
throws IOException;

default Map<String, List<StoragePathInfo>> getAllFilesInPartitions(Collection<String> partitionPaths,
Option<StoragePathFilter> pathFilterOption)
throws IOException {
return getAllFilesInPartitions(partitionPaths);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Make Map<String, List<StoragePathInfo>> getAllFilesInPartitions(Collection<String> partitionPaths) to have default implementation of getAllFilesInPartitions(partitionPaths, Option.empty()) so subclasses can avoid the repeating code? Then getAllFilesInPartitions(Collection<String> partitionPaths, Option<StoragePathFilter> pathFilterOption) becomes an abstract method.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good idea, made the suggested code changes.

HoodieTimer timer = HoodieTimer.start();
List<StoragePathInfo> allFiles = listPartitionPathFiles(partitions, activeTimeline);
log.info("On {} with query instant as {}, it took {}ms to list all files {} Hudi partitions",
metaClient.getTableConfig().getTableName(), queryInstant.map(instant -> instant).orElse("N/A"),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: queryInstant.map(instant -> instant).orElse("N/A") — the .map(instant -> instant) is a no-op. You can simplify to queryInstant.orElse("N/A").

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Made the suggested change.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wasn't the suggestion to go w/
queryInstant.orElse("N/A")

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bad, fixed it now, please check.


public HoodieROTablePathFilter() {
this(new Configuration());
this(HadoopFSUtils.getStorageConf());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

HoodieROTablePathFilter and BaseFileOnlyRelation should no longer be used based on the latest master; instead, HoodieCopyOnWriteSnapshotHadoopFsRelationFactory is used.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

lets TAL at all implementations extending from HoodieBaseHadoopFsRelationFactory and we write them in

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@yihua HoodieCopyOnWriteSnapshotHadoopFsRelationFactory uses HoodieFileIndex. To optimize the file system view calls in HoodieFileIndex we are using HoodieROPathFilter. I think originally HoodieROPathFilter is used at a Relation level, now we downgraded to PathFilter level. So, with the current setup it should be fine right?

@yihua yihua left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A better and general approach would be adding a file system view based on the latest snapshot only to limit the size of file slices in memory, which is used by the file index. That should solve the problem with better layering.

return getAllDataFilesInPartition(storage, partitionPath, Option.empty());
}

public static List<StoragePathInfo> getAllDataFilesInPartition(HoodieStorage storage,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we might need to change the naming, now that its not all files.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Renamed it to getAllDataFilesInPartitionByPathFilter

@nsivabalan nsivabalan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lets add tests for time travel query as well

// Add the FileSlice to partitionToFileSlices
PartitionPath partitionPathObj = partitionsMap.get(relPartitionPath);
List<FileSlice> fileSlices = partitionToFileSlices.computeIfAbsent(partitionPathObj, k -> new ArrayList<>());
fileSlices.add(fileSlice);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we avoid this special handling.
lets route all the files into FSV.
so that we maintain one flow for all cases.

just that the the input files could have already been filtered (if path filter is applied), or could be referring to all files(if no path filter).
much simpler from maintainability standpoint.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating a file system view is costly, we have actually created fsv via ropathfilter.
Adding fsv again again here means we are doing it twice so we should avoid it here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, I see that, but building FSV here is not invoking distributed spark context and there is no metadata table involved (NoOpTableMetadata is used)). its happening in driver and we have just 1 file slice per file group. So, constructing should not add much overhead.

On the plus side, we dont' need to maintain additional custom code for, when path filter is enabled.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is a tradeoff between simplicity and performance. Also, I feel creating FSV on a bunch of files itself is very error prone.
We should skip creating another FSV that way later if any more logic added to it, we dont need to go through all that logic.


public HoodieROTablePathFilter() {
this(new Configuration());
this(HadoopFSUtils.getStorageConf());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

lets TAL at all implementations extending from HoodieBaseHadoopFsRelationFactory and we write them in

@nsivabalan

Copy link
Copy Markdown
Contributor

A better and general approach would be adding a file system view based on the latest snapshot only to limit the size of file slices in memory, which is used by the file index. That should solve the problem with better layering.

Hey @yihua : based on latest state of the patch, I feel it nicely sits w/n HoodieTableMetadata and so, we can leverage this w/ any of FSV.
Just that I see some special handling of FileIndex after filtering which we can avoid (shared feedback above).
otherwise, current layering seems ok to me.

@suryaprasanna
suryaprasanna force-pushed the path-filter-during-listing branch from 9b8395e to de29445 Compare February 23, 2026 23:38
private final boolean shouldIncludePendingCommits;
protected final boolean shouldIncludePendingCommits;
private final boolean shouldValidateInstant;
protected final boolean useROPathFilterForListing;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

useLatestBasePathFilterForListing

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, renamed the variable.

HoodieTimer timer = HoodieTimer.start();
List<StoragePathInfo> allFiles = listPartitionPathFiles(partitions, activeTimeline);
log.info("On {} with query instant as {}, it took {}ms to list all files {} Hudi partitions",
metaClient.getTableConfig().getTableName(), queryInstant.map(instant -> instant).orElse("N/A"),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wasn't the suggestion to go w/
queryInstant.orElse("N/A")

// For MOR tables, we need log files which are not returned by HoodieROTablePathFilter
if (useROPathFilterForListing
&& !shouldIncludePendingCommits
&& metaClient.getTableConfig().getTableType() == HoodieTableType.COPY_ON_WRITE) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for MOR, if query type is RO, we could also afford to add the filtering right?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, also adding MOR with RO as well.

// Add the FileSlice to partitionToFileSlices
PartitionPath partitionPathObj = partitionsMap.get(relPartitionPath);
List<FileSlice> fileSlices = partitionToFileSlices.computeIfAbsent(partitionPathObj, k -> new ArrayList<>());
fileSlices.add(fileSlice);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, I see that, but building FSV here is not invoking distributed spark context and there is no metadata table involved (NoOpTableMetadata is used)). its happening in driver and we have just 1 file slice per file group. So, constructing should not add much overhead.

On the plus side, we dont' need to maintain additional custom code for, when path filter is enabled.

} else {
// Also allow passing in the path filter config via Spark session conf for convenience
pathFilterOptimizedListingEnabled = getConfigValue(options, sqlConf,
"spark." + DataSourceReadOptions.FILE_INDEX_LIST_FILE_STATUSES_USING_RO_PATH_FILTER.key, null)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't your other patch fix this already. we should fix all config key prefix stripping at some root level and not litter across the code base.
#18205

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we also need to clean other configs as well. Let me take that as a followup. I would like to get this landed soon so the other PRs can be unblocked.

val result = spark.sql(s"select id, name, price, ts from $tableName order by id").collect()
// Should have deleted records where id % 3 = 0 (3, 6, 9)
// Should have doubled price for even ids (2, 4, 8, 10)
assert(result.length == 7) // 10 - 3 deleted = 7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@suryaprasanna : did you get a chance to address this?

import org.apache.hudi.storage.StoragePath;
import org.apache.hudi.storage.StoragePathFilter;

public class HoodieROTableStoragePathFilter implements StoragePathFilter {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how about HoodieLatestBaseFilePathFilter

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, renaming the file.

…on the driver

Reviewers: O955 Project Hoodie Project Reviewer: Add blocking reviewers, pwason, jingli, meenalb, singh.sumit

Reviewed By: O955 Project Hoodie Project Reviewer: Add blocking reviewers, pwason

Tags: #has_java

JIRA Issues: HUDI-6646

Differential Revision: https://code.uberinternal.com/D17441111

Fix build failures

Fix checkstyle

Refactor code

Create unit tests
@suryaprasanna
suryaprasanna force-pushed the path-filter-during-listing branch from d34823b to 8d3be59 Compare February 25, 2026 06:56
@hudi-bot

Copy link
Copy Markdown
Collaborator

CI report:

Bot commands @hudi-bot supports the following commands:
  • @hudi-bot run azure re-run the last Azure build

val result = spark.sql(s"select id, name, price, ts from $tableName order by id").collect()
// Should have deleted records where id % 3 = 0 (3, 6, 9)
// Should have doubled price for even ids (2, 4, 8, 10)
assert(result.length == 7) // 10 - 3 deleted = 7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha. sg.

@nsivabalan

Copy link
Copy Markdown
Contributor

hey @suryaprasanna : can you fix the title and PR desc based on latest.
Essentially, Adding support for LatestBaseFilesPathFilter to Spark File Index

@suryaprasanna suryaprasanna changed the title perf: Add HoodieROPathFilter support during file listing to prevent driver OOM perf: Adding support for LatestBaseFilesPathFilter to Spark File Index Feb 26, 2026
@nsivabalan
nsivabalan merged commit 181af01 into apache:master Feb 26, 2026
87 of 88 checks passed
@nsivabalan

Copy link
Copy Markdown
Contributor

oh.
the PR desc is not fixed for the right config key
hoodie.datasource.read.file.index.optimize.listing.using.path.filter

dwshmilyss pushed a commit to dwshmilyss/hudi that referenced this pull request Mar 11, 2026
apache#18136)

This PR adds an opt-in path filtering mechanism during file listing to prevent Spark driver OOM errors when querying large Hudi datasets with multiple file versions per partition.

Problem: When file listing is performed without filtering, all file versions (including older ones) are loaded into driver memory, causing OOM on large tables.

Solution: Added a new config hoodie.datasource.read.file.index.list.file.statuses.using.ro.path.filter (default: false) that enables HoodieROPathFilter during file listing to exclude older file versions.

Summary and Changelog
Users can now enable path filtering during file listing to avoid loading multiple file versions into memory on the driver. This is controlled by the new config hoodie.datasource.read.file.index.list.file.statuses.using.ro.path.filter.

Changes:

New Config: FILE_INDEX_LIST_FILE_STATUSES_USING_RO_PATH_FILTER (default: false)

Enables path filtering during file listing to reduce driver memory pressure
Filters out older file versions, keeping only the latest files needed for queries
API Extensions:

Extended HoodieTableMetadata, BaseTableMetadata, and FileSystemBackedTableMetadata to accept optional StoragePathFilter parameter
Added FSUtils.getAllDataFilesInPartition overload with path filter support
Created HoodieROTableStoragePathFilter wrapper to adapt Hadoop PathFilter to Hudi's StoragePathFilter interface
Spark Integration:

Updated BaseHoodieTableFileIndex to use path filter when enabled
Modified SparkHoodieTableFileIndex to apply HoodieROTablePathFilter during partition listing
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 40.54054% with 44 lines in your changes missing coverage. Please review.
✅ Project coverage is 57.28%. Comparing base (7f4dfda) to head (8d3be59).
⚠️ Report is 77 commits behind head on master.

Files with missing lines Patch % Lines
...java/org/apache/hudi/BaseHoodieTableFileIndex.java 36.36% 19 Missing and 2 partials ⚠️
...e/hudi/hadoop/HoodieLatestBaseFilesPathFilter.java 0.00% 7 Missing ⚠️
...c/main/scala/org/apache/hudi/HoodieFileIndex.scala 40.00% 4 Missing and 2 partials ⚠️
...la/org/apache/hudi/SparkHoodieTableFileIndex.scala 33.33% 5 Missing and 1 partial ⚠️
...rg/apache/hudi/hadoop/HoodieROTablePathFilter.java 50.00% 2 Missing ⚠️
...c/main/java/org/apache/hudi/common/fs/FSUtils.java 66.66% 0 Missing and 1 partial ⚠️
...ache/hudi/common/table/view/NoOpTableMetadata.java 0.00% 1 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##             master   #18136      +/-   ##
============================================
- Coverage     57.29%   57.28%   -0.02%     
- Complexity    18555    18561       +6     
============================================
  Files          1945     1946       +1     
  Lines        106199   106262      +63     
  Branches      13128    13135       +7     
============================================
+ Hits          60852    60869      +17     
- Misses        39621    39662      +41     
- Partials       5726     5731       +5     
Flag Coverage Δ
hadoop-mr-java-client 45.41% <36.00%> (-0.02%) ⬇️
spark-java-tests 47.41% <37.83%> (-0.02%) ⬇️
spark-scala-tests 45.52% <36.48%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...va/org/apache/hudi/metadata/BaseTableMetadata.java 80.74% <ø> (ø)
...e/hudi/metadata/FileSystemBackedTableMetadata.java 81.88% <100.00%> (ø)
.../org/apache/hudi/metadata/HoodieTableMetadata.java 68.42% <100.00%> (+1.75%) ⬆️
...g/apache/hudi/hadoop/HiveHoodieTableFileIndex.java 100.00% <ø> (ø)
...main/scala/org/apache/hudi/DataSourceOptions.scala 95.29% <100.00%> (+0.04%) ⬆️
...c/main/java/org/apache/hudi/common/fs/FSUtils.java 65.29% <66.66%> (-0.11%) ⬇️
...ache/hudi/common/table/view/NoOpTableMetadata.java 12.50% <0.00%> (-0.55%) ⬇️
...rg/apache/hudi/hadoop/HoodieROTablePathFilter.java 81.25% <50.00%> (-0.86%) ⬇️
...c/main/scala/org/apache/hudi/HoodieFileIndex.scala 82.35% <40.00%> (-1.73%) ⬇️
...la/org/apache/hudi/SparkHoodieTableFileIndex.scala 72.22% <33.33%> (-1.21%) ⬇️
... and 2 more

... and 11 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@voonhous voonhous mentioned this pull request Sep 7, 2026
2 of 3 tasks
voonhous added a commit to ryux1/hudi that referenced this pull request Sep 7, 2026
The metadata config `BaseHoodieTableFileIndex` builds is gated on a
three-way conjunction, and nothing in the repo asserted how it
resolves. Both of the added conjuncts came from behavior fixes:
`isFilesPartitionAvailable` from HUDI-5403 (apache#7488, a Trino listing
regression) and `useLatestBaseFilesPathFilterForListing` from apache#18136.
The reflective test removed earlier in this PR planted the field and
asserted the getter returned it, so it discriminated nothing.

Parameterize the existing `TestLocalIndex` harness, which already
builds a real index over a real meta client, with the RO-path-filter
flag, and assert the full truth table through it.

Each row was checked against a mutant: dropping any one of the three
conjuncts from the production expression makes exactly one row fail.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L PR with lines of changes in (300, 1000]

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants