Skip to content

Questions regarding data durability and schema #1

Description

@notABot101010

Hello,

This is a very nice project!

Before using it, I would love to read more documentation about how the buffering and parquets files work. From my current understanding, data is buffered in memory, written to local parquet files and then uploaded to S3 in batch.

Is there some kind of data loss window if the instance crash between receiving the data, and having it sent to S3?

Also, how parquet files are managed? I suppose that using a shorted buffering interval will create more parquet files and thus increase query time? Is there some kind of compaction?

Regarding the schema, my current understanding is that there is currently a single "database" with multiple "collections" (measurement name in the README). Is it correct?

How these different collections are organized in the parquet files? Is there any drawbacks on using multiple "collections" with a single arc server? (e.g. "analytics_events", "http_logs", "host_metrics"...)

Activity

  1. xe-nvdk commented on Oct 8, 2025

    @xe-nvdk
    Member

    Hello! Thank you for the kind words and great questions. Let me address each one based on how Arc is currently architected:

    Buffering and Data Loss Window

    You're correct about the data flow:

    1. Data is buffered in memory (per-measurement buffers)
    2. Flushed to local Parquet files (temporary files)
    3. Uploaded to S3/MinIO in batch

    Data loss window: Yes, there is a potential data loss window between receiving data and it being persisted to S3/MinIO. The window is controlled by two flush triggers:

    • Size-based: Default 50,000 records per measurement (configurable via WRITE_BUFFER_SIZE)
    • Time-based: Default 5 seconds (configurable via WRITE_BUFFER_AGE)

    Whichever limit is reached first triggers a flush. If the instance crashes before a flush completes, in-memory data is lost.

    Mitigation strategies (not currently implemented, but on the roadmap):

    • Write-Ahead Log (WAL) to local disk before buffering
    • Configurable durability modes (performance vs. safety tradeoff)
    • Clustering with replication

    For the current alpha release, we're optimizing for throughput (1.89M records/sec in M3 Pro Max (14 cores, 36GB). If you need stronger durability guarantees, you could:

    • Reduce WRITE_BUFFER_AGE to 1-2 seconds
    • Reduce WRITE_BUFFER_SIZE to force more frequent flushes
    • Use multiple Arc instances with client-side replication

    Parquet File Management

    File creation: You're right that shorter buffering intervals create more files. Each flush creates one Parquet file per measurement with naming: {measurement}_{timestamp}_{record_count}.parquet

    Partitioning: Files are organized with hour-level time partitioning:

    s3://bucket/cpu/2025/10/08/14/cpu_20251008_140530_50000.parquet ↑ └─── year/month/day/hour partitioning └─ measurement name
    

    Query performance: DuckDB handles many small files efficiently thanks to:

    • Parquet statistics (min/max values for time-based pruning)
    • Parallel file reads
    • Columnar format with predicate pushdown

    Compaction: Not currently implemented, but planned for a future release. For now:

    • Smaller buffers = more files but lower latency
    • Larger buffers = fewer files but higher data loss risk
    • We recommend the defaults (50K records, 5 seconds) as a good balance

    Schema and Collections Organization

    Database model: Arc uses a flat namespace model:

    • No explicit "database" concept (though the db= parameter in Line Protocol is stored as metadata)
    • Each measurement is essentially a collection/table
    • All measurements share the same S3 bucket but have separate directories

    Example organization:

    s3://arc/ ├── analytics_events/ │ └── 2025/10/08/14/.parquet ├── http_logs/ │ └── 2025/10/08/14/.parquet └── host_metrics/ └── 2025/10/08/14/*.parquet
    

    Multiple collections: No drawbacks to using multiple measurements on a single Arc instance:

    • Each measurement has its own buffer (isolated flush behavior)
    • Queries are per-measurement anyway (e.g., SELECT * FROM http_logs)
    • Storage is partitioned by measurement + time
    • No cross-measurement overhead

    Configuration Examples

    For high-throughput, low-latency (accept 5s data loss risk):

    WRITE_BUFFER_SIZE=50000
    WRITE_BUFFER_AGE=5
    # For high-durability (slower writes, <1s data loss risk):
    WRITE_BUFFER_SIZE=10000
    WRITE_BUFFER_AGE=1
    

    I've added an ARCHITECTURE.md document to the repository that explains these internals in detail, including diagrams showing the data flow, buffering behavior, and storage layout. Let me know if you have more questions!

  2. notABot101010 commented on Oct 9, 2025

    @notABot101010
    Author

    Awesome, thank you! I will for sure follow the evolution of this project.

    I've opened #2 regarding data loss prevention, which seems today to be the biggest blocker for adoption.

  3. xe-nvdk commented on Oct 9, 2025

    @xe-nvdk
    Member

    I’ll keep it open until we land compaction in the upcoming builds.

    Also, as of last night, we’ve started designing multi-database support, this will make Arc even more flexible for mixed workloads and multi-tenant setups.

  4. reopened this on Oct 9, 2025
  5. xe-nvdk commented on Oct 14, 2025

    @xe-nvdk
    Member

    Great news! Compaction is now fully implemented.

    Compaction Implementation

    Arc now includes automatic background compaction that merges small Parquet files into larger, optimized files. This significantly improves query performance and reduces storage costs.

    How It Works

    Architecture:

    1. Time-based partitioning: Files are organized by hour ({measurement}/{year}/{month}/{day}/{hour}/)
    2. Automatic merging: Compaction runs hourly, targeting partitions with 10+ small files
    3. DuckDB-powered: Uses DuckDB to read multiple Parquet files and write a single optimized file
    4. Concurrent processing: Supports up to 2 concurrent compaction jobs (configurable)
    5. Lock management: SQLite-based locking prevents concurrent compaction of the same partition

    Configuration (arc.conf):

    [compaction]
    enabled = true              # Enable automatic compaction
    min_age_hours = 1          # Only compact partitions older than 1 hour
    min_files = 10             # Require at least 10 files to trigger compaction
    target_file_size_mb = 512  # Target size for compacted files
    max_concurrent_jobs = 2    # Max concurrent compaction jobs
    schedule = "5 * * * *"     # Run every hour at :05 (cron syntax)
    
    compression = "zstd"       # Better compression than snappy
    compression_level = 3      # Balance compression vs speed

    Performance Impact

    Compression:

    • 80.4% compression ratio achieved (3.7 GB → 724 MB with ZSTD level 3)
    • ZSTD compression provides better compression than Snappy
    • Row groups optimized to ~120K rows for query performance
    • Query speed improvements:
    • Fewer files = less S3 API calls = faster query planning
    • Better columnar compression = less data to scan
    • Parquet statistics (min/max) enable partition pruning

    Example from ClickBench:
    Before compaction: 99.9M rows across 5,000+ small files (3.7 GB)
    After compaction: Same 99.9M rows in 1 optimized file (724 MB)
    Query speedup: Q0 (SELECT COUNT(*)) went from 2.3s → 0.18s (12.7x faster)

    API Endpoints

    You can monitor and trigger compaction via REST API:

    Get compaction status

    GET /api/compaction/status

    Trigger manual compaction

    POST /api/compaction/trigger

    View compaction candidates

    GET /api/compaction/candidates

    View compaction history

    GET /api/compaction/history?limit=10

    Multi-Database Support

    Arc now supports multi-database architecture.

    Each database is a separate namespace with:

    • Isolated storage: {backend}/{database}/{measurement}/...
    • Multi-tenant ready: Perfect for SaaS deployments or organizing data by environment

    Storage structure:

    /data/arc/  (or s3://bucket/)
    ├── production/          # Database: production
    │   ├── cpu/
    │   │   └── 2025/10/14/15/*.parquet
    │   └── memory/
    │       └── 2025/10/14/15/*.parquet
    ├── staging/             # Database: staging
    │   └── cpu/
    │       └── 2025/10/14/15/*.parquet
    └── analytics/           # Database: analytics
        └── events/
            └── 2025/10/14/15/*.parquet
    

    Configuration (in arc.conf):

    [storage.local]
    base_path = "./data/arc"
    database = "default"  # Default database namespace
    

    Or for S3:

    [storage.s3]
    bucket = "arc-data"
    region = "us-east-1"
    database = "production"  # Database namespace
    Usage:
    # Ingest to specific database (via header)
    POST /write
    X-Arc-Database: production
    Content-Type: application/x-msgpack
    

    Query references database in table name

    POST /query
    {
      "sql": "SELECT * FROM production.cpu WHERE time > NOW() - INTERVAL 1 HOUR"
    }
    

    This makes Arc much more flexible for mixed workloads and multi-tenant setups, while maintaining the same high-performance characteristics across all databases.

    Let me know what do you think about this additions.

    Thank you for your contributions, Open Source is feedback too.

  6. xe-nvdk commented on Jan 14, 2026

    @xe-nvdk
    Member

    Im closing this issue, features were implemented and are super stable at this point.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions