Expected Behavior
After a task queue partition moves to another matching host, DescribeTaskQueue(ReportStats) keeps
reporting the tasks still in the backlog: ApproximateBacklogCount >= 1 while ApproximateBacklogAge > 0.
Actual Behavior
If the partition changes owner within ~1 minute of a task being written, the new owner reports
ApproximateBacklogCount = 0 while ApproximateBacklogAge keeps growing. It never corrects until the
task is dispatched. In production we saw a single activity task sit undispatched for 9+ hours
(approximateBacklogAge: 32855s, no approximateBacklogCount), because our autoscaler (KEDA's temporal
scaler) scales on the count and kept the only worker pool for that queue at 0 replicas.
From reading v1.30.6, what we think is happening (classic backlog, SQL persistence):
- taskQueueDB.CreateTasks increments ApproximateBacklogCount in memory. The SQL store returns
CreateTasksResponse{UpdatedMetadata: false} (common/persistence/sql/task_v1.go), so the new count
is only persisted by the periodic ack-level update (matching.updateAckInterval, default 1m).
- If ownership moves before that write, the new owner's takeOverTaskQueueLocked loads the
previously persisted count (without the new task), and bumps the range ID, so the old owner's
final write is rejected.
- The new owner's task reader still reads the task row, so the age is reported
(BacklogStatsByPriority uses the reader's in-memory oldest-task time), while the count comes from
the stale counter.
- Nothing recounts upward: the classic path resets the count only to 0 (when ackLevel ==
maxReadLevel), and updateBacklogStatsLocked clamps at 0 on dispatch ("ApproximateBacklogCount
could have under-counted.").
Note: the fair backlog manager has setKnownFairBacklogCount ("reset ApproximateBacklogCount when the
backlog count is known, e.g. when you're read to the end of the backlog"); the classic/pri readers
appear to have no equivalent upward correction. We have not tested whether matching.enableFairness
avoids this.
Related: during a version ramp, splitTaskQueueStatsByRampPercentage derives both shares from the
count and zeroes each share's age when its count is 0. So with a lost count the waiting task is
invisible in every version bucket (count 0, age 0), and a consumer can't even fall back on the age.
Steps to Reproduce the Problem
Deterministic on a cluster with SQL persistence and >= 2 matching pods:
- Start a workflow on a fresh task queue that has no pollers:
temporal workflow start --task-queue repro-hot --type NoWorker --workflow-id repro-hot-wf
(temporal task-queue describe --task-queue repro-hot --task-queue-type workflow -o json
shows approximateBacklogCount: 1)
- Within ~1s, force-delete all matching pods:
kubectl -n delete pod -l app.kubernetes.io/component=matching --grace-period=0 --force
- Once the new matching pods are up, describe the queue again:
approximateBacklogCount: 0, approximateBacklogAge grows (observed 36s -> 2m36s over 2+ minutes,
past the new owner's own 1m ack-level update).
Control: repeat step 1 on another queue, wait > 60s, then do step 2. That queue keeps
approximateBacklogCount: 1 after the kill, so the loss is specific to the unsaved window.
Starting a second workflow on the affected queue shows the count one short: count=1 while two
tasks are waiting (age = the older task's age).
Specifications
- Version: Temporal server v1.30.6 (classic backlog; matching.useNewMatcher and
matching.enableFairness at defaults (off); matching.numTaskqueueRead/WritePartitions = 5)
- Platform: Kubernetes (deployed via temporal-operator), PostgreSQL 12 persistence (postgres12
SQL plugin), linux/arm64
Expected Behavior
After a task queue partition moves to another matching host, DescribeTaskQueue(ReportStats) keeps
reporting the tasks still in the backlog: ApproximateBacklogCount >= 1 while ApproximateBacklogAge > 0.
Actual Behavior
If the partition changes owner within ~1 minute of a task being written, the new owner reports
ApproximateBacklogCount = 0 while ApproximateBacklogAge keeps growing. It never corrects until the
task is dispatched. In production we saw a single activity task sit undispatched for 9+ hours
(approximateBacklogAge: 32855s, no approximateBacklogCount), because our autoscaler (KEDA's temporal
scaler) scales on the count and kept the only worker pool for that queue at 0 replicas.
From reading v1.30.6, what we think is happening (classic backlog, SQL persistence):
CreateTasksResponse{UpdatedMetadata: false} (common/persistence/sql/task_v1.go), so the new count
is only persisted by the periodic ack-level update (matching.updateAckInterval, default 1m).
previously persisted count (without the new task), and bumps the range ID, so the old owner's
final write is rejected.
(BacklogStatsByPriority uses the reader's in-memory oldest-task time), while the count comes from
the stale counter.
maxReadLevel), and updateBacklogStatsLocked clamps at 0 on dispatch ("ApproximateBacklogCount
could have under-counted.").
Note: the fair backlog manager has setKnownFairBacklogCount ("reset ApproximateBacklogCount when the
backlog count is known, e.g. when you're read to the end of the backlog"); the classic/pri readers
appear to have no equivalent upward correction. We have not tested whether matching.enableFairness
avoids this.
Related: during a version ramp, splitTaskQueueStatsByRampPercentage derives both shares from the
count and zeroes each share's age when its count is 0. So with a lost count the waiting task is
invisible in every version bucket (count 0, age 0), and a consumer can't even fall back on the age.
Steps to Reproduce the Problem
Deterministic on a cluster with SQL persistence and >= 2 matching pods:
temporal workflow start --task-queue repro-hot --type NoWorker --workflow-id repro-hot-wf
(
temporal task-queue describe --task-queue repro-hot --task-queue-type workflow -o jsonshows approximateBacklogCount: 1)
kubectl -n delete pod -l app.kubernetes.io/component=matching --grace-period=0 --force
approximateBacklogCount: 0, approximateBacklogAge grows (observed 36s -> 2m36s over 2+ minutes,
past the new owner's own 1m ack-level update).
Control: repeat step 1 on another queue, wait > 60s, then do step 2. That queue keeps
approximateBacklogCount: 1 after the kill, so the loss is specific to the unsaved window.
Starting a second workflow on the affected queue shows the count one short: count=1 while two
tasks are waiting (age = the older task's age).
Specifications
matching.enableFairness at defaults (off); matching.numTaskqueueRead/WritePartitions = 5)
SQL plugin), linux/arm64