Summary
The compaction process creates temporary directories at ./data/compaction/{job_id}/ for each job. These are cleaned up via defer in job.go:202, but if a pod crashes mid-compaction, the defer never runs and orphaned directories accumulate.
Problem
./data/compaction/
├── metrics_cpu_2025_01_15_10_1737834567890/ <-- orphaned from crash
│ ├── file1.parquet
│ ├── file2.parquet
│ └── cpu_20250115_compacted.parquet
├── metrics_cpu_2025_01_15_11_1737838167890/ <-- another orphan
Over time, these accumulate and waste disk space on the compute node.
Proposed Solution
Add startup cleanup in Manager initialization:
// CleanupOrphanedTempDirs removes orphaned temp directories from previous runs.
// This handles cleanup after pod crashes where the defer cleanup didn't run.
func (m *Manager) CleanupOrphanedTempDirs() error {
entries, err := os.ReadDir(m.TempDirectory)
if err != nil {
if os.IsNotExist(err) {
return nil // Directory doesn't exist yet
}
return err
}
for _, entry := range entries {
if entry.IsDir() {
path := filepath.Join(m.TempDirectory, entry.Name())
if err := os.RemoveAll(path); err != nil {
m.logger.Warn().Err(err).Str("dir", entry.Name()).Msg("Failed to cleanup orphaned temp directory")
} else {
m.logger.Info().Str("dir", entry.Name()).Msg("Cleaned up orphaned temp directory")
}
}
}
return nil
}
Call this in NewManager() or at the start of each compaction cycle.
Context
Related to #157 and PR #163 (manifest-based recovery). The manifest PR handles S3 state recovery; this issue handles local temp state cleanup.
Labels
Summary
The compaction process creates temporary directories at
./data/compaction/{job_id}/for each job. These are cleaned up viadeferinjob.go:202, but if a pod crashes mid-compaction, the defer never runs and orphaned directories accumulate.Problem
Over time, these accumulate and waste disk space on the compute node.
Proposed Solution
Add startup cleanup in
Managerinitialization:Call this in
NewManager()or at the start of each compaction cycle.Context
Related to #157 and PR #163 (manifest-based recovery). The manifest PR handles S3 state recovery; this issue handles local temp state cleanup.
Labels
enhancementcompaction