Summary
Queries fail with No files found error when querying S3 storage and daily compaction hasn't created day-level files yet.
Reproduction
- Use S3 storage backend
- Ingest data (creates hourly files at
/year/month/day/hour/*.parquet)
- Query data before daily compaction runs:
SELECT COUNT(*) FROM measurement WHERE time >= '2026-01-20' AND time < '2026-01-21'
Error
IO Error: No files found that match the pattern "s3://bucket/db/measurement/2026/01/20/*.parquet"
Root Cause
-
GeneratePartitionPaths unconditionally adds day-level paths (/2026/01/20/*.parquet) for every day in the query range - no check for whether daily compaction has run or files exist.
-
filterExistingRemotePaths checks if the directory exists (it does - contains hour subdirs), not if files exist at that path (they don't).
-
Partition pruner has no access to daily_min_age_hours config to determine which days might have daily-compacted files.
File Structure
s3://bucket/db/measurement/2026/01/20/
├── 19/file_compacted.parquet ← Files here (hourly)
├── 20/file_compacted.parquet ← Files here (hourly)
└── (no files at day level) ← Query looks here, fails!
Generated Query
FROM read_parquet([
's3://.../2026/01/20/19/*.parquet', -- ✓ Has files
's3://.../2026/01/20/20/*.parquet', -- ✓ Has files
's3://.../2026/01/20/*.parquet', -- ✗ NO FILES - causes error
])
Affected Files
internal/pruning/partition_pruner.go:498-508 - Unconditional day path addition
internal/pruning/partition_pruner.go:778-791 - Directory check instead of file check
Workaround Limitation
Running daily compaction manually does NOT fully resolve this issue:
curl -X POST "http://server/api/v1/compaction/trigger?tier=daily"
Why: Daily compaction respects daily_min_age_hours (default: 24 hours). Data newer than this threshold will NOT be compacted, so day-level files won't exist for recent days.
Example with daily_min_age_hours = 25:
- Today's data: No day-level files (too recent) → Query fails
- Yesterday's data: No day-level files (only 24h old) → Query fails
- Data from 2+ days ago: Day-level files exist → Query works
Result: Queries on recent data (< daily_min_age_hours) will always fail on S3, regardless of manual compaction triggers.
Real-World Impact: Grafana Dashboards
Testing with Grafana dashboards querying last 6 hours of data:
SELECT COUNT(*) FROM downloads
WHERE $__timeFilter(time) AND response = 200
-- Expands to: WHERE time >= '2026-01-21T00:00:00Z' AND time < '2026-01-21T06:00:00Z'
This query will ALWAYS fail because:
- 6-hour-old data will never have daily compacted files
daily_min_age_hours (default 24h) prevents recent data from being daily-compacted
- Partition pruner still adds day-level paths for these recent hours
- Day-level paths have no files → query fails
Impact: Any Grafana dashboard querying recent data (< 24 hours) on S3 storage is broken.
Unexpected Behavior: Complex Queries Work
Simple query - FAILS:
SELECT COUNT(*) FROM downloads
WHERE time >= '2026-01-20' AND time < '2026-01-21' AND response = 200
Complex query - WORKS:
SELECT COUNT(*) FROM downloads
WHERE time >= '2026-01-20' AND time < '2026-01-21' AND response = 200
AND site NOT IN ('site1', 'site2')
Why: The WHERE clause extraction regex (whereClausePattern) can be affected by complex predicates, causing time range extraction to fail. When extraction fails, partition pruning falls back to **/*.parquet glob pattern which works correctly.
This means adding unrelated WHERE conditions can accidentally "fix" the query - indicating the partition pruning logic is fragile.
Related: #134
Summary
Queries fail with
No files founderror when querying S3 storage and daily compaction hasn't created day-level files yet.Reproduction
/year/month/day/hour/*.parquet)Error
Root Cause
GeneratePartitionPathsunconditionally adds day-level paths (/2026/01/20/*.parquet) for every day in the query range - no check for whether daily compaction has run or files exist.filterExistingRemotePathschecks if the directory exists (it does - contains hour subdirs), not if files exist at that path (they don't).Partition pruner has no access to
daily_min_age_hoursconfig to determine which days might have daily-compacted files.File Structure
Generated Query
Affected Files
internal/pruning/partition_pruner.go:498-508- Unconditional day path additioninternal/pruning/partition_pruner.go:778-791- Directory check instead of file checkWorkaround Limitation
Running daily compaction manually does NOT fully resolve this issue:
curl -X POST "http://server/api/v1/compaction/trigger?tier=daily"Why: Daily compaction respects
daily_min_age_hours(default: 24 hours). Data newer than this threshold will NOT be compacted, so day-level files won't exist for recent days.Example with
daily_min_age_hours = 25:Result: Queries on recent data (<
daily_min_age_hours) will always fail on S3, regardless of manual compaction triggers.Real-World Impact: Grafana Dashboards
Testing with Grafana dashboards querying last 6 hours of data:
This query will ALWAYS fail because:
daily_min_age_hours(default 24h) prevents recent data from being daily-compactedImpact: Any Grafana dashboard querying recent data (< 24 hours) on S3 storage is broken.
Unexpected Behavior: Complex Queries Work
Simple query - FAILS:
Complex query - WORKS:
Why: The WHERE clause extraction regex (
whereClausePattern) can be affected by complex predicates, causing time range extraction to fail. When extraction fails, partition pruning falls back to**/*.parquetglob pattern which works correctly.This means adding unrelated WHERE conditions can accidentally "fix" the query - indicating the partition pruning logic is fragile.
Related: #134