Skip to content

Commit b5f4c00

Browse files
Ignacio Van Droogenbroeckclaude
andcommitted
docs(compaction): document max_files_per_batch (v2026.09.1+)
Compaction splits large partitions into batches, each becoming an independent job with its own output file. That batch size was a hardcoded 30; it is configurable from v2026.09.1. Documents the valid range (2-500), why 1 is rejected (compaction's adaptive retry cannot process a single-file batch), and the edge/ constrained-link case that motivates lowering it — along with the trade-off of more compaction jobs and, in cluster mode, more Raft manifest entries. Gated behind a version admonition since the setting has no effect on earlier builds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent fc4cf60 commit b5f4c00

1 file changed

Lines changed: 21 additions & 0 deletions

File tree

‎docs/advanced/compaction.md‎

Lines changed: 21 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -184,6 +184,27 @@ max_concurrent_jobs = 2 # Run 2 compactions in parallel (default)
184184
# max_concurrent_jobs = 1 # Sequential (lower resource usage)
185185
```
186186

187+
#### Files Per Batch
188+
189+
:::info Available in v2026.09.1+
190+
`max_files_per_batch` is configurable starting in Arc **v2026.09.1**. On earlier versions the batch size is fixed at 30 files and this setting has no effect.
191+
:::
192+
193+
A partition with more files than this is split into several batches, each compacted as an independent job producing its own output file.
194+
195+
```toml
196+
[compaction]
197+
max_files_per_batch = 30 # Files per compaction job (default)
198+
# max_files_per_batch = 5 # Smaller outputs, more jobs per partition
199+
# max_files_per_batch = 60 # Fewer, larger outputs
200+
```
201+
202+
Valid range is **2–500**. Values outside it fall back to the default with a startup warning; `1` is rejected because compaction's adaptive retry cannot process a single-file batch.
203+
204+
This bounds the **file count** per job, not the output size in bytes — compacted file size tracks input file size, which follows your ingest buffer settings. The main reason to lower it is transferring compacted files over a constrained or intermittent link (edge deployments), where smaller, independently-transferable files resume better after an interruption. The trade-off is more compaction jobs per partition, and in cluster mode proportionally more Raft manifest entries.
205+
206+
The upper bound exists because DuckDB can abort when a single `read_parquet()` call spans too many files.
207+
187208
#### Compression
188209

189210
```toml

0 commit comments

Comments
 (0)