Repository navigation
feat: Upgrade timestamp precision from milliseconds to microseconds - #22
Merged
Merged
Conversation
Improves timestamp storage from millisecond (ms) to microsecond (us) precision
across Arc for better observability and distributed tracing support.
- **Sub-millisecond precision**: Critical for distributed tracing spans
- **Industry standard**: Microseconds is standard for observability databases
(TimescaleDB, InfluxDB traces, OpenTelemetry)
- **No storage overhead**: Both ms and us use 64-bit timestamps
- **DuckDB native**: No conversion needed - reads timestamp[us] natively
- **Better precision**: 0.001ms resolution vs 1ms resolution
- Schema inference now uses `pa.timestamp('us')` instead of `pa.timestamp('ms')`
- Added intelligent timestamp unit auto-detection based on magnitude:
- < 1e10: seconds (e.g., 1730246400)
- < 1e13: milliseconds (e.g., 1730246400000) ← Telegraf, load tests
- ≥ 1e13: microseconds (e.g., 1730246400000000)
- Full backwards compatibility with existing millisecond timestamps
- Added auto-detection for columnar and single-record formats
- Maintains backwards compatibility with msgpack API (still accepts ms)
- Automatically converts to datetime for Arrow writer
- Updated to output datetime objects instead of converting to milliseconds
- Simplified timestamp handling (arrow_writer now accepts datetime directly)
- Supports seconds, milliseconds, and microseconds in query results
**NO BREAKING CHANGES** - All existing clients continue to work:
- ✅ **Telegraf Arc plugin**: Sends UnixMilli() → auto-detected as ms
- ✅ **Arc telemetry server**: Sends milliseconds → auto-detected as ms
- ✅ **Load test scripts**: Send ms timestamps → auto-detected as ms
- ✅ **Line protocol API**: Has precision parameter (ns/us/ms/s)
- ✅ **Existing data**: Old ms data still readable by DuckDB
- Verified new Parquet files store `timestamp[us]`
- Confirmed DuckDB reads `timestamp[us]` natively (no conversion)
- Tested JSON endpoint: Returns correct ISO timestamps
- Tested Arrow endpoint: Returns `timestamp[us]` with full precision
- Verified Telegraf compatibility with millisecond timestamps
- Confirmed mixed precision queries work (old ms + new us data)
No migration needed! New data uses microseconds, old data (ms) remains
readable. DuckDB automatically handles mixed precision when querying.
6 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Improves timestamp storage from millisecond (ms) to microsecond (us) precision across Arc for better observability and distributed tracing support.
Sub-millisecond precision: Critical for distributed tracing spans
Industry standard: Microseconds is standard for observability databases (TimescaleDB, InfluxDB traces, OpenTelemetry)
No storage overhead: Both ms and us use 64-bit timestamps
DuckDB native: No conversion needed - reads timestamp[us] natively
Better precision: 0.001ms resolution vs 1ms resolution
Schema inference now uses
pa.timestamp('us')instead ofpa.timestamp('ms')Added intelligent timestamp unit auto-detection based on magnitude:
Full backwards compatibility with existing millisecond timestamps
Added auto-detection for columnar and single-record formats
Maintains backwards compatibility with msgpack API (still accepts ms)
Automatically converts to datetime for Arrow writer
Updated to output datetime objects instead of converting to milliseconds
Simplified timestamp handling (arrow_writer now accepts datetime directly)
Supports seconds, milliseconds, and microseconds in query results
NO BREAKING CHANGES - All existing clients continue to work:
✅ Telegraf Arc plugin: Sends UnixMilli() → auto-detected as ms
✅ Arc telemetry server: Sends milliseconds → auto-detected as ms
✅ Load test scripts: Send ms timestamps → auto-detected as ms
✅ Line protocol API: Has precision parameter (ns/us/ms/s)
✅ Existing data: Old ms data still readable by DuckDB
Verified new Parquet files store
timestamp[us]Confirmed DuckDB reads
timestamp[us]natively (no conversion)Tested JSON endpoint: Returns correct ISO timestamps
Tested Arrow endpoint: Returns
timestamp[us]with full precisionVerified Telegraf compatibility with millisecond timestamps
Confirmed mixed precision queries work (old ms + new us data)
No migration needed! New data uses microseconds, old data (ms) remains readable. DuckDB automatically handles mixed precision when querying.