Repository navigation
Conversation
Add retry logic for queries that fail with "No files found" error on day-level paths (YYYY/MM/DD/*.parquet). When daily compaction hasn't run yet, these paths don't exist but the partition pruner includes them. Instead of checking upfront, retry the query without day-level paths on failure. - Add patternDayLevelPath regex to detect day-level path patterns - Add isDayLevelPathError() to check if error is about day-level paths - Add removeDayLevelPaths() to strip day-level paths from SQL - Retry query automatically when day-level path error detected
|
Thanks for investigating this! You've found a real edge case. After digging deeper, I believe the issue is in how the pruner validates day-level paths. Currently:
The fix should be in the pruner itself rather than retry logic. A couple of options: Option A: Change Option B: Only generate day-level paths if we know daily compaction has run (check for Would you be interested in implementing one of these fixes instead? The retry approach with SQL string manipulation is fragile and could break in edge cases (e.g., if only day-level paths exist, or if the regex doesn't match certain path formats). I can help review a pruner-level fix if you'd like to take this on! |
|
Thanks for the suggestions! I agree the fix should be in the pruner. Regarding the two options: Option 1 (check file presence) is thorough but could impact performance when users have many uncompacted files, Option 2 (only generate day-level paths after daily compaction) addresses this specific issue cleanly without the However, I'm thinking about other edge cases where paths might not have files—for example, if a user manually What do you think about combining both approaches?
This way we get the performance benefit of Option 2 for the common case, while Option 1 handles unexpected One more idea: when we detect missing files, we could create empty placeholder files to fill the gap. This would |
|
Thanks for the investigation, but after deeper analysis we're going to close this PR. Here's why: The Scenario Doesn't Happen in PracticeAfter tracing through the code:
In normal usage, queries include both hourly AND daily paths. The hourly paths have data, so DuckDB can establish the schema and the query succeeds. If Queries Return No Data, That's Correct BehaviorIf daily compacted files don't exist, the query returns 0 rows from that path, not a failure. The data is in the hourly paths. The Proposed Fixes Are Over-EngineeringThe suggestions (pruner-level validation, placeholder files, fallback layers) add complexity for an edge case that:
If you can provide a specific reproduction case with steps and logs showing the actual error, we'd reconsider. But based on code analysis, this appears to be a non-issue. Thanks for contributing - please don't let this discourage you from future PRs! |
|
This is actually a real issue I hit. I migrated about a year's worth of data (~700 million records), and after running compaction, all my queries started failing with "No files found" on day-level paths. I'll test more to provide reproduction steps for it. |
|
Ok, let's do this. Let's convert this to an issue, so, we can work and go deep on this. If you can provide how you ingest the data, how the dataset is stored, folders and sub folders and data structure, would be good to reproduce. |
|
After further debugging, I've identified that this issue is related to #131. |
|
Oh, OK, is the issue opened, that Im going to tackle today. I will keep you posted. |
Summary