Repository navigation
Questions regarding data durability and schema #1
Description
Activity
Hello! Thank you for the kind words and great questions. Let me address each one based on how Arc is currently architected:
Buffering and Data Loss Window
You're correct about the data flow:
- Data is buffered in memory (per-measurement buffers)
- Flushed to local Parquet files (temporary files)
- Uploaded to S3/MinIO in batch
Data loss window: Yes, there is a potential data loss window between receiving data and it being persisted to S3/MinIO. The window is controlled by two flush triggers:
- Size-based: Default 50,000 records per measurement (configurable via
WRITE_BUFFER_SIZE) - Time-based: Default 5 seconds (configurable via
WRITE_BUFFER_AGE)
Whichever limit is reached first triggers a flush. If the instance crashes before a flush completes, in-memory data is lost.
Mitigation strategies (not currently implemented, but on the roadmap):
- Write-Ahead Log (WAL) to local disk before buffering
- Configurable durability modes (performance vs. safety tradeoff)
- Clustering with replication
For the current alpha release, we're optimizing for throughput (1.89M records/sec in M3 Pro Max (14 cores, 36GB). If you need stronger durability guarantees, you could:
- Reduce
WRITE_BUFFER_AGEto 1-2 seconds - Reduce
WRITE_BUFFER_SIZEto force more frequent flushes - Use multiple Arc instances with client-side replication
Parquet File Management
File creation: You're right that shorter buffering intervals create more files. Each flush creates one Parquet file per measurement with naming:
{measurement}_{timestamp}_{record_count}.parquetPartitioning: Files are organized with hour-level time partitioning:
s3://bucket/cpu/2025/10/08/14/cpu_20251008_140530_50000.parquet ↑ └─── year/month/day/hour partitioning └─ measurement nameQuery performance: DuckDB handles many small files efficiently thanks to:
- Parquet statistics (min/max values for time-based pruning)
- Parallel file reads
- Columnar format with predicate pushdown
Compaction: Not currently implemented, but planned for a future release. For now:
- Smaller buffers = more files but lower latency
- Larger buffers = fewer files but higher data loss risk
- We recommend the defaults (50K records, 5 seconds) as a good balance
Schema and Collections Organization
Database model: Arc uses a flat namespace model:
- No explicit "database" concept (though the
db=parameter in Line Protocol is stored as metadata) - Each measurement is essentially a collection/table
- All measurements share the same S3 bucket but have separate directories
Example organization:
s3://arc/ ├── analytics_events/ │ └── 2025/10/08/14/.parquet ├── http_logs/ │ └── 2025/10/08/14/.parquet └── host_metrics/ └── 2025/10/08/14/*.parquetMultiple collections: No drawbacks to using multiple measurements on a single Arc instance:
- Each measurement has its own buffer (isolated flush behavior)
- Queries are per-measurement anyway (e.g.,
SELECT * FROM http_logs) - Storage is partitioned by measurement + time
- No cross-measurement overhead
Configuration Examples
For high-throughput, low-latency (accept 5s data loss risk):
WRITE_BUFFER_SIZE=50000 WRITE_BUFFER_AGE=5
# For high-durability (slower writes, <1s data loss risk): WRITE_BUFFER_SIZE=10000 WRITE_BUFFER_AGE=1I've added an ARCHITECTURE.md document to the repository that explains these internals in detail, including diagrams showing the data flow, buffering behavior, and storage layout. Let me know if you have more questions!
Awesome, thank you! I will for sure follow the evolution of this project.
I've opened #2 regarding data loss prevention, which seems today to be the biggest blocker for adoption.
I’ll keep it open until we land compaction in the upcoming builds.
Also, as of last night, we’ve started designing multi-database support, this will make Arc even more flexible for mixed workloads and multi-tenant setups.
Great news! Compaction is now fully implemented.
Compaction Implementation
Arc now includes automatic background compaction that merges small Parquet files into larger, optimized files. This significantly improves query performance and reduces storage costs.
How It Works
Architecture:
- Time-based partitioning: Files are organized by hour (
{measurement}/{year}/{month}/{day}/{hour}/) - Automatic merging: Compaction runs hourly, targeting partitions with 10+ small files
- DuckDB-powered: Uses DuckDB to read multiple Parquet files and write a single optimized file
- Concurrent processing: Supports up to 2 concurrent compaction jobs (configurable)
- Lock management: SQLite-based locking prevents concurrent compaction of the same partition
Configuration (
arc.conf):[compaction] enabled = true # Enable automatic compaction min_age_hours = 1 # Only compact partitions older than 1 hour min_files = 10 # Require at least 10 files to trigger compaction target_file_size_mb = 512 # Target size for compacted files max_concurrent_jobs = 2 # Max concurrent compaction jobs schedule = "5 * * * *" # Run every hour at :05 (cron syntax) compression = "zstd" # Better compression than snappy compression_level = 3 # Balance compression vs speed
Performance Impact
Compression:
- 80.4% compression ratio achieved (3.7 GB → 724 MB with ZSTD level 3)
- ZSTD compression provides better compression than Snappy
- Row groups optimized to ~120K rows for query performance
- Query speed improvements:
- Fewer files = less S3 API calls = faster query planning
- Better columnar compression = less data to scan
- Parquet statistics (min/max) enable partition pruning
Example from ClickBench:
Before compaction: 99.9M rows across 5,000+ small files (3.7 GB)
After compaction: Same 99.9M rows in 1 optimized file (724 MB)
Query speedup: Q0 (SELECT COUNT(*)) went from 2.3s → 0.18s (12.7x faster)API Endpoints
You can monitor and trigger compaction via REST API:
Get compaction status
GET /api/compaction/status
Trigger manual compaction
POST /api/compaction/trigger
View compaction candidates
GET /api/compaction/candidates
View compaction history
GET /api/compaction/history?limit=10
Multi-Database Support
Arc now supports multi-database architecture.
Each database is a separate namespace with:
- Isolated storage: {backend}/{database}/{measurement}/...
- Multi-tenant ready: Perfect for SaaS deployments or organizing data by environment
Storage structure:
/data/arc/ (or s3://bucket/) ├── production/ # Database: production │ ├── cpu/ │ │ └── 2025/10/14/15/*.parquet │ └── memory/ │ └── 2025/10/14/15/*.parquet ├── staging/ # Database: staging │ └── cpu/ │ └── 2025/10/14/15/*.parquet └── analytics/ # Database: analytics └── events/ └── 2025/10/14/15/*.parquetConfiguration (in arc.conf):
[storage.local] base_path = "./data/arc" database = "default" # Default database namespaceOr for S3:
[storage.s3] bucket = "arc-data" region = "us-east-1" database = "production" # Database namespace Usage: # Ingest to specific database (via header) POST /write X-Arc-Database: production Content-Type: application/x-msgpackQuery references database in table name
POST /query { "sql": "SELECT * FROM production.cpu WHERE time > NOW() - INTERVAL 1 HOUR" }This makes Arc much more flexible for mixed workloads and multi-tenant setups, while maintaining the same high-performance characteristics across all databases.
Let me know what do you think about this additions.
Thank you for your contributions, Open Source is feedback too.
- Time-based partitioning: Files are organized by hour (
Im closing this issue, features were implemented and are super stable at this point.
- added a commit that references this issue
on May 8, 2026 - added 4 commits that reference this issue
on May 15, 2026
Hello,
This is a very nice project!
Before using it, I would love to read more documentation about how the buffering and parquets files work. From my current understanding, data is buffered in memory, written to local parquet files and then uploaded to S3 in batch.
Is there some kind of data loss window if the instance crash between receiving the data, and having it sent to S3?
Also, how parquet files are managed? I suppose that using a shorted buffering interval will create more parquet files and thus increase query time? Is there some kind of compaction?
Regarding the schema, my current understanding is that there is currently a single "database" with multiple "collections" (measurement name in the README). Is it correct?
How these different collections are organized in the parquet files? Is there any drawbacks on using multiple "collections" with a single arc server? (e.g. "analytics_events", "http_logs", "host_metrics"...)