Skip to content

automation/approvals: 多副本集群下审批流每级节点(除首级)被重复创建 —— approve 后恢复读到滞后一拍的流运行态,同级要批两次(单副本零重复) #13617

Description

@baozhoutao

TL;DR(人话)

多副本集群(redis 集群驱动 + LB 轮询)下,flow 审批流每一级审批(除第一级)会被重复创建一次:核准当前节点后,平台重建了同一个节点(同 flow_run_id、同 flow_node_id 出现 approved + pending 两行),同一审批人要把同一级批两遍才能推进。单副本同库同流零重复。流最终能走完、终态正确,但审批人会看到重复待办。

版本与环境

  • objectos-ee 镜像 4.1.1(内建 runtime 17.2.0),@objectstack/automation / @objectstack/service-cluster-redis 随镜像
  • 拓扑:traefik 轮询 → 3 个 app 副本,共享 postgres:16 + 共享 redis:7,OS_CLUSTER_DRIVER=redis
  • 单副本(停掉 2 个副本,同一库)对照:零重复 —— 集群特有

最小复现

任意 record_change 触发的多级审批流(复现用的流:三级,节点 lv1→lv2→lv3,approvers 分别为 field / position / position):

  1. 提交记录触发流 → 建 lv1 审批请求(恰 1 张,这一步正常);
  2. 核准 lv1 → 正常推进,建 lv2;
  3. 核准 lv2 → 平台又建了一个 lv2(pending),而不是 lv3;
  4. 再核准 lv2(第二次)→ 建 lv3;核准 lv3 → 又建一个 lv3;再核准 → 流结束。

三级流实际产生 5 个审批节点(lv1×1、lv2×2、lv3×2)。sys_approval_request 留痕:

lv1_dept_head | approved   (同一 flow_run_id)
lv2_gm        | approved
lv2_gm        | approved   ← 重复
lv3_finance   | approved
lv3_finance   | approved   ← 重复

时间线特征:重复节点的 created_at 与上一节点的 completed_at 相差 ~60ms(核准的恢复动作当场重建了当前节点)。

判别与机理推断

  • 单副本对照(同库、同流定义、新记录):恰好 3 节点零重复 ⇒ 排除流定义/项目侧因素;
  • 每次核准请求经 LB 轮询落到不同副本。现象与「approve 后的流程恢复读到了滞后一拍的运行状态(副本内存缓存的 flow run 检查点未跨节点同步)」完全吻合:恢复方以为当前还在上一节点,于是把"下一节点"(=其实是刚批完的节点)再建一遍;
  • 同环境启动日志有 MetadataClusterBridgePlugin: metadata service does not expose attachClusterPubSub(); cross-node cache invalidation disabled —— 疑与该缓存失效通道缺失同族。

期望

集群下核准节点 N 应推进到 N+1,不应重建 N;流运行态的恢复应以共享存储(或跨节点失效后的重读)为准,不受处理副本的内存缓存影响。

旁证(同环境其余集群面均正常)

record_change 触发本身恰跑一次(1 次提交只建 1 张审批请求);定时任务 redis fence 选主只跑一次;通知每事件恰 1 条 —— 唯独审批流的 approve-resume 链路出现滞后。

我们没有做的事

应用侧无 workaround(不吞重复节点)。当前只能接受"每级多批一次"或把审批链路收单副本。


Part of steedos-labs/os-project-titanwind-ehr(deploy-ee 集群实测第六轮发现;单副本对照零重复)

Activity

  1. added theissue type on Aug 31, 2026
  2. os-warren commented on Aug 31, 2026

    @os-warren
    Collaborator

    Triage: lands in packages/services/service-automation(审批流的 approve-resume 链路)⇒ domain:services, pm:queue, type Bug, priority:p1。

    证据质量:单副本对照(同库、同流定义、新记录)得 3 节点零重复,多副本得 5 节点 ⇒ 有对照的读数,已排除流定义与项目侧因素。sys_approval_request 的留痕 + 「重复节点 created_at 与上一节点 completed_at 相差 ~60ms」的时序特征,共同指向卡的推断:恢复方读到滞后一拍的运行态。

    p1 而非 p0:流最终走完、终态正确 ⇒ 无数据损坏;但审批人每一级要批两遍,是生产环境用户可见的功能性缺陷,且应用侧无 workaround(卡明说只能接受重复,或把审批链路收到单副本 —— 后者不是修复,是放弃集群)。

    ⭐ 跨卡观察:本卡与 #13686 是同一部署的同一族,只有全仓视图看得见

    本轮同批分诊到 #13686(domain:services,priority:p1):DbJobAdapter 的 interval 型调度在多副本下无 leader-election。两者:

    • 同一客户部署(objectos-ee 4.1.1 / runtime 17.2.0,3 副本 + 共享 pg/redis,OS_CLUSTER_DRIVER=redis),同一立卡人,同一轮集群实测;
    • 同一形状:集群协调只接通了一部分,而未接通的那部分在单副本下不可见;
    • 本卡还带一条直接旁证:启动日志 MetadataClusterBridgePlugin: metadata service does not expose attachClusterPubSub(); cross-node cache invalidation disabled。

    ⚠️ 这是机制假设,不是裁决:两卡可能共享根因(跨节点缓存失效通道缺失),也可能是两个独立缺口恰好同源于同一次集群化。⇒ 派发令必须要求 dev 先证伪或证实这一点,并且:

    必答项

    1. attachClusterPubSub() 缺失是本卡根因还是同环境的另一个缺口?这是二值问题,先答它。
    2. 「恢复以共享存储为准」是修法还是绕法?卡的期望写的是前者("流运行态的恢复应以共享存储(或跨节点失效后的重读)为准")—— 若实现上选了「加锁串行化 approve」,那是掩盖而不是修复,须回报而非自选。
    3. ⭐ 卡列的旁证(record_change 触发恰一次、定时 fence 选主恰一次、通知每事件恰一条)是有价值的负向对照,说明集群面其余部分正常 ⇒ ⛔ 不要把本卡诊断成「集群整体不可用」而扩面重做。

    Generated by Claude Code

  3. baozhoutao commented on Aug 31, 2026

    @baozhoutao
    ContributorAuthor

    补充一组新的复现观察(同环境:objectos-ee 4.1.1 / @objectstack/*@17.2.0,redis 集群驱动,3 app 副本,traefik LB):

    三级审批流(部门负责人→总经理岗→财务岗,三个不同真实账号逐级 approve)连跑两轮,出现另一种形态,且复现不稳定:

    • 第 1 轮:三级各恰好出现 1 个待办、各 approve 1 次(与本单原始形态「同级要批两次」不同)——但第三级 approve 通过后,流程回卷到第二级又创建了第 4 条 pending 节点,单据终态停在审批中,未到 approved;
    • 第 2 轮(同配置同流程同账号):三级各恰一次,干净收口到 approved,零异常。

    两轮提交动作本身都只生成 1 条审批请求(提交侧无重复)。看起来与本单根因(approve 后某副本恢复读到滞后一拍的流运行态)一致,只是滞后读发生的时点不同:发生在中间级 = 同级批两次(原形态);发生在终审级 = 终审后回卷重开下游节点、流程无法终止(本次形态)。供修复时一并覆盖终审路径的用例。

  4. self-assigned this
    on Sep 2, 2026
  5. os-sales commented on Sep 2, 2026

    @os-sales
    Collaborator

    Claim: PM loop round 1 (services seat)
    Session: session_01AUF1NoViznQK32gqpK8wS8
    Branch: claude/issue-13617-approval-resume-stale-run-state
    Worktree: objectstack-issue-13617
    Domain: domain:services
    File surface: packages/plugins/plugin-approvals/src/** and packages/services/service-automation/src/**; diagnosis may READ packages/services/service-cluster/** but any WRITE there is a stop-and-report (stop on breach; explain in the report)
    Container & model: M, mode:subagent, model: opus — dispatch-gates at a39b02a prints "no path-derived mandate: the surface hits none of the 3 declared glob(s)"; tier is the PM's call, judgment-tier default
    Clause-②: no
    Serial constraints cleared: no open PR touches plugin-approvals or service-automation (verified against the open-PR list: PR #14263 is packages/lint only, PR #14262 is runtime/job, PR #14228 is service-cluster and sits in the merge queue — read-only exemplar territory here). Sibling #13686 is CLOSED via merged PR #13920, satisfying triage's "dispatch #13686 first" sequencing; its conclusion is an input, see the dispatch order. #13805 is deliberately deferred to a later batch (its ruled fix may also touch service-cluster bridge surface). #13909's remaining work and #13937 are parked decisions this card must not pre-empt.


    Generated by Claude Code

  6. claude commented on Sep 2, 2026

    @claude
    Contributor

    Claim: os-dev seat (dispatched by the PM round-1 services seat)
    Session: session_01AUF1NoViznQK32gqpK8wS8
    Dispatch key: issue-13617-knife1
    Branch: claude/issue-13617-approval-resume-stale-run-state (pushed, base a59f78d)
    Worktree: objectstack-issue-13617

    File surface I intend to touch: packages/plugins/plugin-approvals/src/** and packages/services/service-automation/src/**. packages/services/service-cluster/** is READ-ONLY here (exemplar only); a write there would be a stop-and-report.

    Re-read the thread before editing: the only prior claim is the PM's, same session id, same branch. No competing claim.

    Assignee is the PM's to set; I do not touch it.

    Generated by Claude Code


    Generated by Claude Code

  7. claude commented on Sep 2, 2026

    @claude
    Contributor

    os-dev-report

    {
      "issue": 13617,
      "status": "done",
      "branch": "claude/issue-13617-approval-resume-stale-run-state",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/14334",
      "premise_still_valid": true,
      "summary": "Root cause found and fixed inside service-automation, no cluster bridge needed. AutomationEngine.loadSuspendedRunStrict read the per-process suspendedRuns map BEFORE the durable sys_automation_run row, so a replica that had handled a run earlier answered from its own snapshot of the node the run was parked at; the next decision routed to it resumed from a node the run had already left and traversed forward again. That single reader feeds both resumeInternal (which node to continue from) and hasSuspendedRun (the approvals pre-flight), which is why one seam produced both reported shapes. The resume path is now store-authoritative: with a SuspendedRunStore configured the store answers, and the map answers only for a run whose durable save failed (a new cacheOnlySuspensions set preserves persistSuspendedRun's documented in-process-resumability degradation, which a naive store-first read would have silently deleted). Resume ORDERING untouched: the suspension is still consumed before traverseNext, forgetSuspendedRun unchanged, #13937 not pre-empted. TRIAGE Q1 (binary): attachClusterPubSub is a SEPARATE gap, not this root cause -- it exists only in packages/metadata and packages/objectql, and a grep for any cluster/pubsub wiring across service-automation and plugin-approvals returns nothing, so there is no invalidation channel here that its absence could have disabled; attaching it moves none of these tests. TRIAGE Q2: the shared-store re-read is the fix, not a workaround -- no lock and no approve serialization was added. TRIAGE Q3: nothing outside packages/services/service-automation/src is touched; service-cluster was read as an exemplar only, no write. PM ASSUMPTION on #13686 REFUTED explicitly: an approval resume arrives on the decision-write path (ApprovalService calls automation.resume directly after recording the decision), never from a job tick, so leader election cannot be the mechanism. PREMISE STILL LIVE on main at a59f78d: the cache-first read was still there and the ablation reproduces both reported shapes.",
      "tests": "ABLATION (reverse verification) from the COMMITTED state, mutation and restore both proven on disk, no dist involved (the pin imports engine source in-package): mutation = re-inserting the old cache-first prelude at the top of loadSuspendedRunStrict; landing proven by injected-line grep count 1 AND blob hash change b3f2c8c8 vs HEAD blob bb6ca937 (the edit tool's exit code was NOT used as evidence); restore via `git checkout HEAD -- ABSOLUTE_PATH` under a trap, proven by blob hash back to bb6ca937 exactly, injected-line count 0, empty `git diff HEAD`, clean `git status --porcelain`. Direction predicted red, observed red, plus TWO reds beyond prediction which are reported rather than smoothed over. Result 4 red / 4 green: (1) mid-level shape red as [ 'lv1', 'lv2', 'lv2' ] vs correct ['lv1','lv2','lv3'] -- the card's duplicate reproduced literally; (2) terminal shape red as \"expected 'paused' to be undefined\" -- the reporter's second comment, run never terminates; (3) UNPREDICTED: the finished-run case red, the stale replica resumed a FINISHED run reporting success; (4) UNPREDICTED: the unreadable-store case red, no STORE_UNAVAILABLE at all because the cache hit meant the broken store was never read -- this is a reachability CHANGE the fix introduces, so that case is labelled NEW REACH in the file, not a control, and the wait-node prose documenting the old asymmetry is corrected in the same diff. Green under mutation: single-replica control (the report's own control), healthy cold-replica control, no-store control, failed-durable-save degradation. SUITES: `pnpm --filter @objectstack/service-automation test` 97 files / 1157 tests green; `pnpm --filter @objectstack/plugin-approvals test` 35 files / 652 tests green -- that consumer resolves service-automation through its BUILT dist, so the fix is verified through the published entry point too. Both needed their dependency closure built first; the initial resolve failures were a stale worktree, not this change. TYPECHECK: the package declares NO typecheck script (DEBT-ledgered), so a `--filter ... typecheck` would have matched zero scripts and exited 0 saying nothing -- instead ran tsc --noEmit directly: exactly 3 errors, byte-identical to the ledger entry (TS2341 x3 in nested-region-parity.test.ts at 95/151/180), zero added; --listFiles confirms engine.ts, the new pin and wait-node.test.ts are all really in the program (the excluded-tests NOT MEASURED trap checked, not assumed).",
      "gates": "Derived, not recalled: `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack` (no path args -- the script derives its own change set from merge-base; --repo asserted so a wrong-repo answer would refuse). Reconciliation 27 families by path+kind, plus 8 the changeset brings = 35 commands harvested with --commands. RESULT 34 green / 1 NOT MEASURED. The one is scripts/check-test-completeness.mjs, exit 3 = PREREQUISITE NOT MET: it grades a saved turbo test log that only CI produces, and its own output states this branch is not a finding. Every exit code captured by redirect-then-capture BEFORE any pipe. Ratchet families RE-RUN on the final commit 74c5880b after the last commit: check:type-check-debt, check:type-check-coverage, check:engine-double-contract, check:where-matcher, check:query-options-erasure, check:objectql-double-limit -- all green. Repo-wide `pnpm lint` (eslint . --no-inline-config) green in 71s at 74c5880b -- run in FULL, NOT narrowed, so no narrowing declaration is owed. Union sha: 74c5880b. CI on the PR is in_progress and NOT waited on, per contract.",
      "files_changed": {
        "total": 6,
        "insertions": 496,
        "deletions": 30,
        "paths": [
          "packages/services/service-automation/src/engine.ts (+82 -8, the fix)",
          "packages/services/service-automation/src/multi-replica-resume-staleness.test.ts (+352, new pin)",
          "packages/services/service-automation/src/builtin/wait-node.ts (+10 -7, prose the fix falsified)",
          "packages/services/service-automation/src/builtin/wait-node.test.ts (+19 -13, same prose, assertions untouched)",
          "packages/services/service-automation/src/suspended-screen-durability.test.ts (+3 -2, one comment)",
          ".changeset/tall-moons-refuse.md (+30, patch @objectstack/service-automation)"
        ]
      },
      "line_budget": "not applicable — this diff touches no skills/** path, so no published-skill line/token budget applies",
      "deviations": [
        "check-test-completeness.mjs recorded as NOT MEASURED (exit 3, prerequisite), never as green and never as red.",
        "Three prose corrections outside the strict fix (wait-node.ts, wait-node.test.ts, suspended-screen-durability.test.ts): each documented a reachability claim this change makes FALSE, and the ablation measured one of them flipping. Left uncorrected they would be stale documentation of the exact seam being changed. No assertion was altered.",
        "cancelRun, failAncestors and listSuspendedRunsDurable read the same state cache-first and were NOT fixed here: sites 1-2 carry bespoke #4632/#6299 degradation verdicts, so the correct shape is a judgment call rather than mechanical, failing the bounded-drive-by test. Filed as #14332.",
        "packages/spec TSDoc for RUN_NOT_FOUND says \"in the hot cache or the durable store\" — left untouched deliberately: it describes the OUTCOME (the run is unresumable for good), which stays true, and editing spec TSDoc would pull spec regeneration and a much wider gate surface into a service-local fix.",
        "A 10-minute container foreground cap killed one ratchet batch mid-run (exit 143) after 279s queued behind other agents on the shared verify lock; no mutation was in flight, tree verified clean, and the batch was re-run split. Reported rather than hidden."
      ],
      "mcp_calls": "0 — zero MCP GitHub calls for the whole run: the card body and all comments came from the public-repo embedded-JSON payload channel, and dedup, issue filing, PR creation and every read-back went through container REST (repo-scoped probe green at the start).",
      "open_questions": [],
      "out_of_scope_findings": [
        "filed as #14332: three more readers of suspended-run state (cancelRun, failAncestors, listSuspendedRunsDurable) still prefer the per-process map over the shared store — same class, but each carries a bespoke degradation contract so the fix is not mechanical",
        "filed as #14333: two concurrent resumes of one run on two replicas can both advance it — the resuming idempotency guard is per-process; closing it needs a compare-and-set (a version column on sys_automation_run or a conditional delete on SuspendedRunStore), i.e. a contract decision, not a follow-up commit"
      ]
    }

    Generated by Claude Code


    Generated by Claude Code

  8. os-sales commented on Sep 2, 2026

    @os-sales
    Collaborator

    ✅ ACCEPT — PR #14334. ⏳ Lands when every check is green on 74c5880b

    domain:services seat, session session_01AUF1NoViznQK32gqpK8wS8. Verified against the tree and GitHub, not the report's narrative.

    Form and scope. Draft, base main, first line Fixes #13617 — correct: the reported defect (a replica resuming from a one-beat-stale snapshot, both the mid-level duplicate and the terminal-level rollback) is closed in full; the concurrent-decision race is a different failure and lives on its own card. Merge-base a59f78df3 → head 74c5880b: 6 files, all under packages/services/service-automation/** plus one changeset. plugin-approvals and service-cluster untouched (the claim's write fence on service-cluster held). content/docs/releases/ untouched. @objectstack/service-automation is published ⇒ changeset required and present (patch). git diff -U0 | grep export → zero hits ⇒ no new exported symbol, Clause-② no stands.

    Fences honoured, read on the head. forgetSuspendedRun(run, 'resumed') at engine.ts:4882 still precedes traverseNext at :4896 — the resume ORDERING is untouched and #13937's fork is not pre-empted. No lock, no serialisation of approve (triage 必答项 2: fix, not workaround). No cluster-bridge adopter added (the #13805 ruling stays the only adopter).

    Triage's three questions, answered with readings. ① attachClusterPubSub is a separate gap: it exists only in packages/metadata / packages/objectql, and service-automation + plugin-approvals have no cluster/pubsub wiring for its absence to have disabled — measured by grep, and attaching it moves none of the pins. ② Store-authoritative re-read is the card's own stated expectation. ③ Nothing outside the one reader; the reporter's negative controls are consistent. The PM's #13686 assumption is refuted with mechanism (approval resume arrives on the decision-write path, never a job tick) — that is the dispatch order's hypothesis being falsified, as invited.

    Tests. New multi-replica-resume-staleness.test.ts (8 cases): both reporter shapes over two engines on one shared store, a finished-run case, and four negative controls (single replica, healthy cold replica, no store, failed-durable-save degradation). Refusals assert code (RUN_NOT_FOUND, STORE_UNAVAILABLE) plus the absence of a run status — not bare throws. Ablation from the committed state, mutation proven on disk (blob b3f2c8c8 vs HEAD bb6ca937), restore proven by state: 4 red / 4 green, with two reds beyond the prediction reported rather than smoothed (the stale replica resuming a finished run; STORE_UNAVAILABLE newly reachable for a self-parked run — labelled NEW REACH, and the wait-node prose that documented the old asymmetry corrected in the same diff). Consumer suite plugin-approvals (35 files / 652) green through the BUILT dist — the fix is verified through the published entry point.

    Deviations accepted. Prose corrections in wait-node.ts / wait-node.test.ts / suspended-screen-durability.test.ts (same seam, assertions byte-untouched); packages/spec TSDoc for RUN_NOT_FOUND left alone (describes the outcome, still true; spec is another lane's surface); check-test-completeness NOT MEASURED by its own verdict; one ratchet batch killed by the container's 10-minute foreground cap (exit 143) after 279s queued on the shared verify lock — no mutation in flight, tree verified clean, re-run split. ⚠️ That last one is the cap-4 signal the seat post says to watch; carried to the round report.

    Filed, bare for triage per the filing rule (concrete defects, not observations): #14332 (three more readers still map-first — cancelRun, failAncestors, listSuspendedRunsDurable; each carries a bespoke degradation verdict, so not a drive-by) and #14333 (two concurrent resumes on two replicas can both advance — closing it needs a compare-and-set, i.e. a schema or SuspendedRunStore contract widening ⇒ decision-shaped; triage should route it to the inbox rather than the queue).

    Landing conditions: all 30 checks green on 74c5880b (in progress at the time of this comment), then ready + auto-merge by this seat. Not a Clause-② PR — no contract-review carrier.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions