Skip to content

Graceful Shutdown & Resource Cleanup #46

Description

@taitelee

Problem

Forcefully killing the application can leave NATS messages unacknowledged, ClickHouse batches half-written, or metrics unflushed, potentially leading to data inconsistencies.

Proposed Solution

Implement a signal listener (SIGTERM/SIGINT) that triggers a coordinated shutdown sequence. This will stop the API listeners first, then flush all pending data buffers to ClickHouse, and finally close database connections.

Alternatives Considered

None. Maybe some database crash recovery but we should not depend on that nor use it as a graceful shutdown method.

Additional Context

None

Activity

  1. self-assigned this
    on Apr 16, 2026
  2. added theissue type on Apr 16, 2026
  3. changed the title [-][feature] Graceful Shutdown & Resource Cleanup[/-] [+]Graceful Shutdown & Resource Cleanup[/+] on Apr 16, 2026
  4. moved this from Backlog to Ready in WaveHouse Task Boardon Apr 21, 2026
  5. moved this from Ready to In progress in WaveHouse Task Boardon Apr 21, 2026
  6. moved this from In progress to Backlog in WaveHouse Task Boardon Apr 27, 2026
  7. EricAndrechek commented on May 15, 2026

    @EricAndrechek
    Member

    Closing — implementation has shipped.

    The coordinated shutdown sequence requested in the issue body is in cmd/wavehouse/main.go:378-393:

    sigCh := make(chan os.Signal, 1)
    signal.Notify(sigCh, syscall.SIGINT, syscall.SIGTERM)
    <-sigCh
    
    shutCtx, shutCancel := context.WithTimeout(context.Background(), time.Duration(cfg.Server.ShutdownTimeout)*time.Second)
    
    _ = ingestStream.Stop(shutCtx)            // flush ingest buffers
    if err := srv.Shutdown(shutCtx); err != nil { ... }    // stop API listener
    if err := promSrv.Shutdown(shutCtx); err != nil { ... } // stop Prometheus exporter

    The earlier otelShutdown(ctx) (~line 123) flushes telemetry providers (traces/metrics/logs) on the same shutdown path.

    Matches the proposed-solution requirements:

    • ✅ SIGTERM/SIGINT signal listener
    • ✅ Bounded shutdown context (configurable timeout via cfg.Server.ShutdownTimeout)
    • ✅ Stops API listeners first (srv.Shutdown)
    • ✅ Flushes pending data buffers (ingestStream.Stop)
    • ✅ Flushes observability (otelShutdown)

    Reopen if a specific resource is leaking — e.g. if you find the ClickHouse driver isn't being closed via this path, that's a follow-up.

  8. moved this from In progress to Done in WaveHouse Task Boardon May 18, 2026
  9. added a commit that references this issue on May 20, 2026
    4bd92a8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions