Skip to content

Timer queue can't recover once the cleanup delete exceeds its 5s timeout (SQL persistence) #12341

Description

@tsurdilo

Each shard periodically deletes timer rows it has finished with, via single DELETE covering everything since the last successful cleanup. - https://github.com/temporalio/temporal/blob/main/service/history/queues/queue_base.go.
Under a hard-coded 5 second timeout (queueIOTimeout, queue_base.go:34).

If query exceeds 5 seconds it's cancelled and rolls back and nothing is deleted

log printed:

    {"msg":"Error range completing queue task","shard-id":3881,"component":"timer-queue-processor",
    "error":"RangeCompleteTimerTask operation failed. Error: pq: canceling statement due to user request"}

So once range becomes too wide to delete in 5s, it cannot shrink, but gets bigger actually.
This can prevent deletes and draining timer task backlog.

Ask:

  • allow deletion in chunks of X configurable via dynamic config
  • allow users to change the hard-coded 5s via dyn config
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions