Skip to content

fix(pruning): verify files exist at day-level paths - #145

Merged
xe-nvdk merged 2 commits into
Basekick-Labs:mainfrom
khalid244:s3-partition-pruning-day-level
Jan 22, 2026
Merged

xe-nvdk merged 2 commits into
Basekick-Labs:mainfrom
khalid244:s3-partition-pruning-day-level

Conversation

@khalid244

@khalid244 khalid244 commented Jan 21, 2026 •

Copy link
Copy Markdown

Summary

  • Fixes S3 partition pruning bug where queries fail with "No files found" when daily compaction hasn't run
  • Day-level paths were included when directory existed but contained only hourly subdirectories
  • Adds file existence check for day-level paths (5 segments) while keeping fast directory check for hourly paths

Fixes #144

@xe-nvdk

xe-nvdk commented Jan 21, 2026 •

Copy link
Copy Markdown
Member

Hi @khalid244,

Thank you for this PR! I've tested your changes against the issue and found that your fix addresses part of the problem but misses a critical edge case.

What Your PR Fixes

Your PR correctly improves filterExistingRemotePaths() to verify that .parquet files exist at the day-level (5 segments: db/measurement/YYYY/MM/DD), not just that the directory exists. This is the right approach!

Before your fix:

  • Day-level path: s3://bucket/db/cpu/2026/01/21/*.parquet
  • Directory 2026/01/21/ exists (has hourly subdirs: 00/, 01/, etc.)
  • Filter passes (directory exists)
  • DuckDB tries to read 2026/01/21/*.parquet (no files) FAILS

After your fix:

  • Day-level path: s3://bucket/db/cpu/2026/01/21/*.parquet
  • Directory 2026/01/21/ exists (has hourly subdirs)
  • Your code checks for .parquet files directly at 2026/01/21/ level
  • No files found → filter excludes this path
  • Only hourly paths included

This is great progress!

What Your PR Doesn't Fix

However, there's an edge case your PR doesn't handle: queries that span time ranges with NO data at all (e.g., querying future dates or dates before any data was ingested).

The Problem

Testing with MinIO, I found this scenario:

  • Data exists: s3://arc-test/production/cpu/2026/01/21/17/*.parquet (today, hour 17)
  • Query: SELECT * FROM cpu WHERE time >= '2026-01-20' AND time < '2026-01-22' (yesterday to tomorrow)

What happens:

  1. Partition pruner generates paths for 2026-01-20, 2026-01-21, and 2026-01-22
  2. Your filterExistingRemotePaths() correctly filters out:
    • 2026/01/20/ (no data - before ingestion started)
    • 2026/01/21/*.parquet (day-level file doesn't exist)
    • 2026/01/22/ (future date, no data)
  3. Result: Empty array [] is returned
  4. In OptimizeTablePath() at line 578-582:
    if len(partitionPaths) == 0 {
        p.logger.Info().Msg("No data exists for time range, using fallback")
        return originalPath, false  // ← Returns s3://bucket/db/cpu/**/*.parquet
    }
  5. Query falls back to scanning all files: s3://bucket/db/cpu/**/*.parquet
  6. If the time range truly has no data, DuckDB still fails: "No files found that match the pattern"

Logs from Testing

INF Generated partition paths ... daily_paths=2 hourly_paths=48 total_paths=50
INF Filtered existing remote paths ... existing_count=0 original_count=50
INF No data exists for time range, using fallback
ERR Query failed error="No files found that match the pattern s3://arc-test/production/cpu/**/*.parquet"

Even though filtering works correctly, the fallback path still causes the query to fail.

Recommended Fix

You need to handle the case when len(existingPaths) == 0 before returning in filterExistingRemotePaths(). Here are two approaches:

Option 1: Return Sentinel Value (Recommended)

In filterExistingRemotePaths(), detect when all paths are filtered out and return a sentinel value to signal "no data":

func (p *PartitionPruner) filterExistingRemotePaths(paths []string) []string {
    // ... existing code ...

    existingPaths := []string{}
    for _, path := range paths {
        // ... your existing filtering logic ...
    }

    p.logger.Info().
        Int("existing_count", len(existingPaths)).
        Int("original_count", len(paths)).
        Msg("Filtered existing remote paths")

    // NEW: If all paths filtered out, return a special marker
    // This tells OptimizeTablePath to return empty result instead of fallback
    if len(existingPaths) == 0 {
        p.logger.Info().Msg("No partitions have data for the specified time range")
        return []string{"__EMPTY__"}  // Sentinel value
    }

    return existingPaths
}

Then in OptimizeTablePath(), check for the sentinel:

// Filter out non-existent paths
partitionPaths = p.filterExistingPaths(partitionPaths)

// NEW: Check if sentinel value returned (no data exists)
if len(partitionPaths) == 1 && partitionPaths[0] == "__EMPTY__" {
    p.logger.Info().Msg("No data exists for time range, will return empty result")
    p.partitionCache.set(cacheKey, []string{}, true)  // Cache empty result
    return []string{}, true  // Signal to query handler to return empty result
}

if len(partitionPaths) == 0 {
    p.logger.Info().Msg("No partitions found, using fallback")
    p.partitionCache.set(cacheKey, originalPath, false)
    return originalPath, false
}

Then in internal/api/query.go at the buildReadParquetExpr() functions (lines ~1477 and ~1511), add:

if pathList, ok := optimizedPath.([]string); ok {
    // NEW: Handle empty path list (no data in time range)
    if len(pathList) == 0 {
        h.logger.Info().Msg("No partitions for time range, returning empty result")
        // Return a query that produces an empty result with correct schema
        // This avoids DuckDB trying to read non-existent files
        return keyword + " (SELECT NULL AS time, NULL AS value LIMIT 0)"
    }

    // ... existing code for non-empty path lists ...
}

Option 2: Simpler Approach

Alternatively, just modify OptimizeTablePath() to not fall back when len(partitionPaths) == 0:

if len(partitionPaths) == 0 {
    p.logger.Info().Msg("No data exists for time range, will return empty result")
    p.partitionCache.set(cacheKey, []string{}, true)
    return []string{}, true  // Don't fall back - signal empty result
}

And handle the empty array in buildReadParquetExpr() as shown above.

Test Case to Add

Please add a test case in internal/pruning/partition_pruner_test.go:

func TestFilterExistingRemotePaths_NoDataInTimeRange(t *testing.T) {
    // Scenario: Query time range that has no data (future dates, or before ingestion)
    // Expected: Query should return empty result gracefully, not fail with "No files found"

    // Setup: Storage with only one hourly partition
    // Query: Time range that excludes that partition
    // Expected: filterExistingRemotePaths returns empty array
    // Expected: Query returns 0 rows instead of error
}

Summary

Your PR fixes the day-level file existence check, which is great! But it needs to also handle the empty result case where all paths are filtered out (no data in time range).

The fix requires changes in three places:

  1. filterExistingRemotePaths() - Your PR already improves this
  2. OptimizeTablePath() - Need to handle len(partitionPaths) == 0 better
  3. buildReadParquetExpr() in internal/api/query.go - Need to handle empty path list

Would you like to update your PR to include these changes? I'm happy to help if you have questions!

Testing

To reproduce the issue:

  1. Start MinIO: docker run -p 9000:9000 -e MINIO_ROOT_USER=minioadmin -e MINIO_ROOT_PASSWORD=minioadmin quay.io/minio/minio server /data
  2. Configure Arc with S3 storage (MinIO endpoint)
  3. Ingest data for today only
  4. Query: SELECT * FROM cpu WHERE time >= '2026-01-20' AND time < '2026-01-22' (includes tomorrow)
  5. Before full fix: Query fails with "No files found"
  6. After full fix: Query returns 0 rows (empty result)

Thank you again for your contribution! The file-level verification logic you added is exactly the right approach. We just need to extend it to handle the edge case where no files exist at all in the queried time range.

Khalid and others added 2 commits January 21, 2026 21:06
…in query

For S3/Azure storage, day-level partition paths (year/month/day/*.parquet)
were incorrectly included when the directory existed but contained only
hourly subdirectories (no actual .parquet files at day level).

This caused DuckDB queries to fail with "No files found that match the pattern"
when daily compaction hadn't run yet.

The fix adds a file existence check specifically for day-level paths (5 segments)
while keeping the fast directory check for hourly paths (6 segments).

Fixes #144
- Add error handling for storage.List() call
- Add debug logging when day-level path has no files
- Check for .parquet suffix to avoid matching non-parquet files
- Improve code comments explaining the fix
@xe-nvdk
xe-nvdk merged commit 4ed9a21 into Basekick-Labs:main Jan 22, 2026
xe-nvdk pushed a commit that referenced this pull request Jan 22, 2026
- Document the day-level file existence check fix
- Credit khalid244 for the contribution
xe-nvdk pushed a commit that referenced this pull request Feb 1, 2026
## New Features

### InfluxDB Client Compatibility
Arc's Line Protocol endpoints now use the same paths as InfluxDB, enabling drop-in compatibility with all official InfluxDB client libraries (Go, Python, JavaScript, Java, C#, PHP, Ruby, Telegraf, Node-RED).

- `/api/v1/write` → `/write` (InfluxDB 1.x clients, Telegraf)
- `/api/v1/write/influxdb` → `/api/v2/write` (InfluxDB 2.x clients)
- Arc-native endpoint `/api/v1/write/line-protocol` preserved
- Supports Bearer token, Token header, API key header, and query parameter authentication

### MQTT Ingestion Support
Native MQTT subscription for IoT and edge data ingestion. Connect directly to MQTT brokers without requiring additional infrastructure.

- Subscribe to multiple MQTT topics with wildcard support (`+`, `#`)
- Dynamic subscription management via REST API
- TLS/SSL connections with certificate validation
- Authentication via username/password or client certificates
- Connection auto-reconnect with exponential backoff
- Per-subscription statistics and monitoring
- Passwords encrypted at rest
- Auto-start subscriptions on server restart

### S3 File Caching via cache_httpfs Extension (PR #149)
Optional in-memory caching of S3 Parquet files via DuckDB's `cache_httpfs` community extension. 5-10x query performance improvement for workloads with repeated file access (CTEs, subqueries, Grafana dashboards).

- In-memory only — no disk caching, preserves stateless compute philosophy
- Opt-in, configurable cache size and TTL
- Graceful degradation if extension fails to load

*Contributed by @khalid244*

### Relative Time Expression Support in Partition Pruning
Queries using `NOW() - INTERVAL` now benefit from partition pruning. Supports seconds, minutes, hours, days, weeks, months.

## Bug Fixes

- **Control Characters in Measurement Names Break S3 (Issue #122)** — Added strict validation for measurement names across all ingestion endpoints (Line Protocol, MsgPack, Continuous Queries). Names must start with a letter, contain only alphanumeric/underscore/hyphen, max 128 chars.
- **Partition Pruner Fails on Non-Existent S3 Partitions (Issue #125, #144, PR #145)** — Fixed "No files found" errors when time range includes non-existent S3 partitions. Extended `filterExistingPaths()` for S3/Azure storage. Also fixed `filepath.Join()` mangling S3 URLs. *(Day-level fix by @khalid244)*
- **Server Timeout Config Values Ignored (Issue #126)** — `server.read_timeout` and `server.write_timeout` now correctly use configuration values instead of hardcoded 30s.
- **Large Payload Ingestion (413 Request Entity Too Large)** — `MaxPayloadSize` config now correctly passed to Fiber's `BodyLimit`.
- **Query Results Timestamp Timezone Inconsistency** — All timestamps in query results now normalized to UTC.
- **Azure Blob Storage Query SSL Certificate Errors on Linux (PR #92)** — Fixed by using system curl for SSL on Linux. *(Contributed by @schotime)*
- **UTC Consistency for Compaction Filenames (PR #132)** — Compacted filenames now use UTC timestamps consistently. *(Contributed by @schotime)*
- **S3 Subprocess Configuration Issues (Issue #131)** — Fixed missing credentials and SSL config forwarding to compaction subprocess.
- **Query Failures with Non-UTF8 Data (Issue #136)** — Added automatic UTF-8 sanitization during ingestion with optimized fast-path (~6-25ns overhead).
- **Nanosecond Timestamp Support for MessagePack Ingestion** — Extended timestamp detection to handle 19-digit nanosecond timestamps (important for InfluxDB migrations).
- **WHERE Clause Regex Fails to Match Multi-Line Queries (Issue #146, PR #148)** — Fixed partition pruner failing on multi-line SQL queries. *(Contributed by @khalid244)*
- **String Literals Containing SQL Keywords Break Partition Pruning** — Added string literal masking before regex matching.
- **Buffer Age-Based Flush Timing Under High Load (Issue #142)** — Ticker now fires at half the configured interval for more consistent flush timing (~25% improvement).
- **Arrow Writer Panic During High-Concurrency Writes (Issue #130)** — Fixed out-of-bounds panic during schema evolution with concurrent writes.
- **Empty Directories Not Cleaned Up After Daily Compaction** — Automatic cleanup of empty hour-level partition directories after compaction.
- **Compactor OOM and Segfaults with Large Datasets (Issue #102)** — Streaming I/O, memory limit passthrough, file batching (1000 max), and adaptive batch sizing on failure.
- **Orphaned Compaction Temp Directories (Issue #164, PR #165)** — Two-layer cleanup: startup cleanup + parent-side cleanup after subprocess completion.
- **Compaction Data Duplication on Crash (Issue #157, PR #163)** — Manifest-based tracking in S3 prevents re-compaction after crash. *(Contributed by @khalid244)*
- **WAL-Based S3 Recovery (Issue #159, PR #162)** — Fixed data loss during S3 outages with startup recovery, periodic recovery, and backpressure handling. *(Contributed by @khalid244)*
- **Tiered Storage Query Routing (Issue #166, PR #167)** — Fixed queries not routing to cold tier data and database listing not showing cold-only databases.
- **Retention Policies for S3/Azure Storage Backends (Issue #169, PR #170)** — Retention policies now work with all storage backends.
- **Retention Policy Empty Directory Cleanup (Issue #171)** — Empty directories cleaned up after retention policy file deletion.
- **Query Timeout for S3 Disconnection (Issue #151, PR #152)** — Configurable query timeout prevents indefinite hangs. Returns HTTP 504 when exceeded. *(Contributed by @khalid244)*

## Improvements

- **Configurable Server Idle and Shutdown Timeouts** — `server.idle_timeout` and `server.shutdown_timeout` now configurable.
- **Automatic Time Function Query Optimization** — `time_bucket()` and `date_trunc()` automatically rewritten to epoch arithmetic (2-2.5x faster GROUP BY queries).
- **Parallel Partition Scanning** — Queries spanning 3+ partitions now execute concurrently (2-4x speedup).
- **Two-Stage Distributed Aggregation (Enterprise)** — 5-20x speedup for cross-shard aggregations via scatter/gather.
- **DuckDB Query Engine Optimizations** — Parquet metadata caching, prefetching, pool-wide `SET GLOBAL` consistency (18-24% faster aggregations). *(SET GLOBAL fix by @khalid244)*
- **Automatic Regex-to-String Function Optimization** — URL domain extraction patterns rewritten to native string functions (2x+ faster).
- **Database Header for Query Optimization** — `x-arc-database` header skips regex parsing (5-17% faster queries).
- **MQTT Client Auto-Generated Client ID** — Prevents client ID collisions across multiple Arc instances.
- **MQTT Restart Endpoint** — `/api/v1/restart` to apply MQTT config changes without server restart.

## Security

- **Token Hashing Security Model** — New tokens hashed with bcrypt (cost 10). SHA256 prefixes for O(1) lookup. Legacy tokens supported.

## Breaking Changes

- `/api/v1/write` → `/write` and `/api/v1/write/influxdb` → `/api/v2/write`. Update client config if using old Arc-specific paths. InfluxDB client libraries need no changes.

## Contributors

Thanks to @schotime (Adam Schroder) and @khalid244 for their contributions to this release.
xe-nvdk pushed a commit that referenced this pull request Feb 1, 2026
## New Features

### InfluxDB Client Compatibility
Arc's Line Protocol endpoints now use the same paths as InfluxDB, enabling drop-in compatibility with all official InfluxDB client libraries (Go, Python, JavaScript, Java, C#, PHP, Ruby, Telegraf, Node-RED).

- `/api/v1/write` → `/write` (InfluxDB 1.x clients, Telegraf)
- `/api/v1/write/influxdb` → `/api/v2/write` (InfluxDB 2.x clients)
- Arc-native endpoint `/api/v1/write/line-protocol` preserved
- Supports Bearer token, Token header, API key header, and query parameter authentication

### MQTT Ingestion Support
Native MQTT subscription for IoT and edge data ingestion. Connect directly to MQTT brokers without requiring additional infrastructure.

- Subscribe to multiple MQTT topics with wildcard support (`+`, `#`)
- Dynamic subscription management via REST API
- TLS/SSL connections with certificate validation
- Authentication via username/password or client certificates
- Connection auto-reconnect with exponential backoff
- Per-subscription statistics and monitoring
- Passwords encrypted at rest
- Auto-start subscriptions on server restart

### S3 File Caching via cache_httpfs Extension (PR #149)
Optional in-memory caching of S3 Parquet files via DuckDB's `cache_httpfs` community extension. 5-10x query performance improvement for workloads with repeated file access (CTEs, subqueries, Grafana dashboards).

- In-memory only — no disk caching, preserves stateless compute philosophy
- Opt-in, configurable cache size and TTL
- Graceful degradation if extension fails to load

*Contributed by @khalid244*

### Relative Time Expression Support in Partition Pruning
Queries using `NOW() - INTERVAL` now benefit from partition pruning. Supports seconds, minutes, hours, days, weeks, months.

## Bug Fixes

- **Control Characters in Measurement Names Break S3 (Issue #122)** — Added strict validation for measurement names across all ingestion endpoints (Line Protocol, MsgPack, Continuous Queries). Names must start with a letter, contain only alphanumeric/underscore/hyphen, max 128 chars.
- **Partition Pruner Fails on Non-Existent S3 Partitions (Issue #125, #144, PR #145)** — Fixed "No files found" errors when time range includes non-existent S3 partitions. Extended `filterExistingPaths()` for S3/Azure storage. Also fixed `filepath.Join()` mangling S3 URLs. *(Day-level fix by @khalid244)*
- **Server Timeout Config Values Ignored (Issue #126)** — `server.read_timeout` and `server.write_timeout` now correctly use configuration values instead of hardcoded 30s.
- **Large Payload Ingestion (413 Request Entity Too Large)** — `MaxPayloadSize` config now correctly passed to Fiber's `BodyLimit`.
- **Query Results Timestamp Timezone Inconsistency** — All timestamps in query results now normalized to UTC.
- **Azure Blob Storage Query SSL Certificate Errors on Linux (PR #92)** — Fixed by using system curl for SSL on Linux. *(Contributed by @schotime)*
- **UTC Consistency for Compaction Filenames (PR #132)** — Compacted filenames now use UTC timestamps consistently. *(Contributed by @schotime)*
- **S3 Subprocess Configuration Issues (Issue #131)** — Fixed missing credentials and SSL config forwarding to compaction subprocess.
- **Query Failures with Non-UTF8 Data (Issue #136)** — Added automatic UTF-8 sanitization during ingestion with optimized fast-path (~6-25ns overhead).
- **Nanosecond Timestamp Support for MessagePack Ingestion** — Extended timestamp detection to handle 19-digit nanosecond timestamps (important for InfluxDB migrations).
- **WHERE Clause Regex Fails to Match Multi-Line Queries (Issue #146, PR #148)** — Fixed partition pruner failing on multi-line SQL queries. *(Contributed by @khalid244)*
- **String Literals Containing SQL Keywords Break Partition Pruning** — Added string literal masking before regex matching.
- **Buffer Age-Based Flush Timing Under High Load (Issue #142)** — Ticker now fires at half the configured interval for more consistent flush timing (~25% improvement).
- **Arrow Writer Panic During High-Concurrency Writes (Issue #130)** — Fixed out-of-bounds panic during schema evolution with concurrent writes.
- **Empty Directories Not Cleaned Up After Daily Compaction** — Automatic cleanup of empty hour-level partition directories after compaction.
- **Compactor OOM and Segfaults with Large Datasets (Issue #102)** — Streaming I/O, memory limit passthrough, file batching (1000 max), and adaptive batch sizing on failure.
- **Orphaned Compaction Temp Directories (Issue #164, PR #165)** — Two-layer cleanup: startup cleanup + parent-side cleanup after subprocess completion.
- **Compaction Data Duplication on Crash (Issue #157, PR #163)** — Manifest-based tracking in S3 prevents re-compaction after crash. *(Contributed by @khalid244)*
- **WAL-Based S3 Recovery (Issue #159, PR #162)** — Fixed data loss during S3 outages with startup recovery, periodic recovery, and backpressure handling. *(Contributed by @khalid244)*
- **Tiered Storage Query Routing (Issue #166, PR #167)** — Fixed queries not routing to cold tier data and database listing not showing cold-only databases.
- **Retention Policies for S3/Azure Storage Backends (Issue #169, PR #170)** — Retention policies now work with all storage backends.
- **Retention Policy Empty Directory Cleanup (Issue #171)** — Empty directories cleaned up after retention policy file deletion.
- **Query Timeout for S3 Disconnection (Issue #151, PR #152)** — Configurable query timeout prevents indefinite hangs. Returns HTTP 504 when exceeded. *(Contributed by @khalid244)*

## Improvements

- **Configurable Server Idle and Shutdown Timeouts** — `server.idle_timeout` and `server.shutdown_timeout` now configurable.
- **Automatic Time Function Query Optimization** — `time_bucket()` and `date_trunc()` automatically rewritten to epoch arithmetic (2-2.5x faster GROUP BY queries).
- **Parallel Partition Scanning** — Queries spanning 3+ partitions now execute concurrently (2-4x speedup).
- **Two-Stage Distributed Aggregation (Enterprise)** — 5-20x speedup for cross-shard aggregations via scatter/gather.
- **DuckDB Query Engine Optimizations** — Parquet metadata caching, prefetching, pool-wide `SET GLOBAL` consistency (18-24% faster aggregations). *(SET GLOBAL fix by @khalid244)*
- **Automatic Regex-to-String Function Optimization** — URL domain extraction patterns rewritten to native string functions (2x+ faster).
- **Database Header for Query Optimization** — `x-arc-database` header skips regex parsing (5-17% faster queries).
- **MQTT Client Auto-Generated Client ID** — Prevents client ID collisions across multiple Arc instances.
- **MQTT Restart Endpoint** — `/api/v1/restart` to apply MQTT config changes without server restart.

## Security

- **Token Hashing Security Model** — New tokens hashed with bcrypt (cost 10). SHA256 prefixes for O(1) lookup. Legacy tokens supported.

## Breaking Changes

- `/api/v1/write` → `/write` and `/api/v1/write/influxdb` → `/api/v2/write`. Update client config if using old Arc-specific paths. InfluxDB client libraries need no changes.

## Contributors

Thanks to @schotime (Adam Schroder) and @khalid244 for their contributions to this release.
@xe-nvdk xe-nvdk mentioned this pull request Feb 1, 2026
xe-nvdk added a commit that referenced this pull request Feb 1, 2026
## New Features

### InfluxDB Client Compatibility
Arc's Line Protocol endpoints now use the same paths as InfluxDB, enabling drop-in compatibility with all official InfluxDB client libraries (Go, Python, JavaScript, Java, C#, PHP, Ruby, Telegraf, Node-RED).

- `/api/v1/write` → `/write` (InfluxDB 1.x clients, Telegraf)
- `/api/v1/write/influxdb` → `/api/v2/write` (InfluxDB 2.x clients)
- Arc-native endpoint `/api/v1/write/line-protocol` preserved
- Supports Bearer token, Token header, API key header, and query parameter authentication

### MQTT Ingestion Support
Native MQTT subscription for IoT and edge data ingestion. Connect directly to MQTT brokers without requiring additional infrastructure.

- Subscribe to multiple MQTT topics with wildcard support (`+`, `#`)
- Dynamic subscription management via REST API
- TLS/SSL connections with certificate validation
- Authentication via username/password or client certificates
- Connection auto-reconnect with exponential backoff
- Per-subscription statistics and monitoring
- Passwords encrypted at rest
- Auto-start subscriptions on server restart

### S3 File Caching via cache_httpfs Extension (PR #149)
Optional in-memory caching of S3 Parquet files via DuckDB's `cache_httpfs` community extension. 5-10x query performance improvement for workloads with repeated file access (CTEs, subqueries, Grafana dashboards).

- In-memory only — no disk caching, preserves stateless compute philosophy
- Opt-in, configurable cache size and TTL
- Graceful degradation if extension fails to load

*Contributed by @khalid244*

### Relative Time Expression Support in Partition Pruning
Queries using `NOW() - INTERVAL` now benefit from partition pruning. Supports seconds, minutes, hours, days, weeks, months.

## Bug Fixes

- **Control Characters in Measurement Names Break S3 (Issue #122)** — Added strict validation for measurement names across all ingestion endpoints (Line Protocol, MsgPack, Continuous Queries). Names must start with a letter, contain only alphanumeric/underscore/hyphen, max 128 chars.
- **Partition Pruner Fails on Non-Existent S3 Partitions (Issue #125, #144, PR #145)** — Fixed "No files found" errors when time range includes non-existent S3 partitions. Extended `filterExistingPaths()` for S3/Azure storage. Also fixed `filepath.Join()` mangling S3 URLs. *(Day-level fix by @khalid244)*
- **Server Timeout Config Values Ignored (Issue #126)** — `server.read_timeout` and `server.write_timeout` now correctly use configuration values instead of hardcoded 30s.
- **Large Payload Ingestion (413 Request Entity Too Large)** — `MaxPayloadSize` config now correctly passed to Fiber's `BodyLimit`.
- **Query Results Timestamp Timezone Inconsistency** — All timestamps in query results now normalized to UTC.
- **Azure Blob Storage Query SSL Certificate Errors on Linux (PR #92)** — Fixed by using system curl for SSL on Linux. *(Contributed by @schotime)*
- **UTC Consistency for Compaction Filenames (PR #132)** — Compacted filenames now use UTC timestamps consistently. *(Contributed by @schotime)*
- **S3 Subprocess Configuration Issues (Issue #131)** — Fixed missing credentials and SSL config forwarding to compaction subprocess.
- **Query Failures with Non-UTF8 Data (Issue #136)** — Added automatic UTF-8 sanitization during ingestion with optimized fast-path (~6-25ns overhead).
- **Nanosecond Timestamp Support for MessagePack Ingestion** — Extended timestamp detection to handle 19-digit nanosecond timestamps (important for InfluxDB migrations).
- **WHERE Clause Regex Fails to Match Multi-Line Queries (Issue #146, PR #148)** — Fixed partition pruner failing on multi-line SQL queries. *(Contributed by @khalid244)*
- **String Literals Containing SQL Keywords Break Partition Pruning** — Added string literal masking before regex matching.
- **Buffer Age-Based Flush Timing Under High Load (Issue #142)** — Ticker now fires at half the configured interval for more consistent flush timing (~25% improvement).
- **Arrow Writer Panic During High-Concurrency Writes (Issue #130)** — Fixed out-of-bounds panic during schema evolution with concurrent writes.
- **Empty Directories Not Cleaned Up After Daily Compaction** — Automatic cleanup of empty hour-level partition directories after compaction.
- **Compactor OOM and Segfaults with Large Datasets (Issue #102)** — Streaming I/O, memory limit passthrough, file batching (1000 max), and adaptive batch sizing on failure.
- **Orphaned Compaction Temp Directories (Issue #164, PR #165)** — Two-layer cleanup: startup cleanup + parent-side cleanup after subprocess completion.
- **Compaction Data Duplication on Crash (Issue #157, PR #163)** — Manifest-based tracking in S3 prevents re-compaction after crash. *(Contributed by @khalid244)*
- **WAL-Based S3 Recovery (Issue #159, PR #162)** — Fixed data loss during S3 outages with startup recovery, periodic recovery, and backpressure handling. *(Contributed by @khalid244)*
- **Tiered Storage Query Routing (Issue #166, PR #167)** — Fixed queries not routing to cold tier data and database listing not showing cold-only databases.
- **Retention Policies for S3/Azure Storage Backends (Issue #169, PR #170)** — Retention policies now work with all storage backends.
- **Retention Policy Empty Directory Cleanup (Issue #171)** — Empty directories cleaned up after retention policy file deletion.
- **Query Timeout for S3 Disconnection (Issue #151, PR #152)** — Configurable query timeout prevents indefinite hangs. Returns HTTP 504 when exceeded. *(Contributed by @khalid244)*

## Improvements

- **Configurable Server Idle and Shutdown Timeouts** — `server.idle_timeout` and `server.shutdown_timeout` now configurable.
- **Automatic Time Function Query Optimization** — `time_bucket()` and `date_trunc()` automatically rewritten to epoch arithmetic (2-2.5x faster GROUP BY queries).
- **Parallel Partition Scanning** — Queries spanning 3+ partitions now execute concurrently (2-4x speedup).
- **Two-Stage Distributed Aggregation (Enterprise)** — 5-20x speedup for cross-shard aggregations via scatter/gather.
- **DuckDB Query Engine Optimizations** — Parquet metadata caching, prefetching, pool-wide `SET GLOBAL` consistency (18-24% faster aggregations). *(SET GLOBAL fix by @khalid244)*
- **Automatic Regex-to-String Function Optimization** — URL domain extraction patterns rewritten to native string functions (2x+ faster).
- **Database Header for Query Optimization** — `x-arc-database` header skips regex parsing (5-17% faster queries).
- **MQTT Client Auto-Generated Client ID** — Prevents client ID collisions across multiple Arc instances.
- **MQTT Restart Endpoint** — `/api/v1/restart` to apply MQTT config changes without server restart.

## Security

- **Token Hashing Security Model** — New tokens hashed with bcrypt (cost 10). SHA256 prefixes for O(1) lookup. Legacy tokens supported.

## Breaking Changes

- `/api/v1/write` → `/write` and `/api/v1/write/influxdb` → `/api/v2/write`. Update client config if using old Arc-specific paths. InfluxDB client libraries need no changes.

## Contributors

Thanks to @schotime (Adam Schroder) and @khalid244 for their contributions to this release.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

S3 partition pruning fails when daily compaction hasn't run

2 participants