Skip to content

time-relative/schedule cron flows permanently fail to re-bind after kernel rebuild — croner name leaked because DbJobAdapter.destroy() never destroys the cron adapter #8362

Description

@os-zhuang

现象

本地 magic-flow rig(cloud main 3f2884c4 × framework main cb43296ef,2026-08-13)实测:一个 timeRelative flow 首次绑定成功后,每一次 kernel 重建都无法重绑,且永不自愈(连续 4 次重建复现):

INFO [time-relative] bound flow 'xqao_contract_expiry_reminder_flow' → sweep 'xqao_contract.expiry_date' offsets [3]d on cron '0 8 * * *'
(kernel 被 freshness probe 驱逐、重建后)
WARN [time-relative] failed to schedule flow 'xqao_contract_expiry_reminder_flow': Cron: Tried to initialize new named job 'flow-time-relative:xqao_contract_expiry_reminder_flow', but name already taken.

之后 sweep 绑在空处,自动化静默死亡——verify_build、UI、DB 全部显示健康,唯一信号是这条 WARN。同一批重建里 record-change trigger 每次都正常重绑,对照明显。

两个对照实验钉住根因是残留的进程级 job 名:

  • 把 start 节点 schedule 换成 {type:'interval'} → 重绑成功(interval 走 setInterval,不经过 croner 命名注册表);换回 cron 型立刻 name already taken。
  • 在 sys_metadata 里改 flow 名(= 换 job 名)→ 下一次重建绑定成功。

根因链(四层)

  1. 绑定层 packages/triggers/trigger-schedule/src/time-relative-trigger.ts:214:job 名 flow-time-relative:<flowName>,不带 env/kernel 标识;绑定前的 this.stop(flowName) 读实例级 this.bound Map——新 kernel 的新实例 Map 为空,清理 no-op。schedule-trigger.ts(flow-schedule:<flowName>)结构完全相同,同样中招。
  2. 适配层 cron-job-adapter.ts:schedule() 先 await this.cancel(name),但也只查自己实例的 this.jobs Map。
  3. croner 层:new Cron(expression, { name }) 登记进 croner 进程级全局命名注册表——旧 kernel 的 Cron 对象还在,新建同名直接 throw。
  4. 销毁层(单点根因) db-job-adapter.ts:kernel 驱逐链路本身是完整的(KernelManager.evict() → kernel.shutdown() → plugin.destroy() → JobServicePlugin.destroy() → dbAdapter.destroy()),但 DbJobAdapter.destroy() 只有 await this.inner.destroy()——漏掉 this.cron(CronJobAdapter)。旧 kernel 的 croner 任务永不 stop,永久占名。

另外 trigger 的 schedule() 是 fire-and-forget(void Promise.resolve(...)),失败被降为一条 WARN。

为什么严重

  • 多租户云运行时里 kernel 驱逐是常规操作:freshness probe ~10s 一次,AI 每次 auto-publish 都会 bump freshness → 驱逐。所以「AI 建好定时自动化 → 用户改一次元数据 → 自动化死了」是常态路径。
  • job 名不带 env id:同一容器里两个环境同名 flow 不用等驱逐就冲突(AI 生成的 flow 名极易撞,如 contract_expiry_reminder_flow)。
  • 失败静默:唯一信号是没人看的 WARN,属于 valid-but-inert(ADR-0078)一类。
  • 旧 kernel 的 cron 任务不只占名字,它还活着——闭包握着已 shutdown 的旧 kernel 引擎(目前正因新绑定失败才没有双写;单修命名不修销毁会把静默死亡变成 zombie 双写)。

期望修复

  1. DbJobAdapter.destroy() 补上 this.cron 的销毁(CronJobAdapter.destroy() 已存在,只是没人调)。
  2. job 名加 environment/kernel 命名空间,杜绝跨 kernel/跨环境冲突。
  3. 重绑改为 replace 语义(bind 撞名时替换旧 job),而非 warn-and-give-up。
  4. 绑定失败升级为可观测状态(不只 WARN)。
  5. 回归测试:bind → 新插件实例模拟 kernel 重建 → 再 bind → job 恰好调度一次且能触发;覆盖 time-relative-trigger 与 schedule-trigger 两处。

注意:即使全部修完,「kernel 不驻留时定时任务不跑」仍是架构层问题(时钟应上移到常驻层),cloud 侧另行跟踪。

Activity

  1. os-zhuang commented on Aug 13, 2026

    @os-zhuang
    ContributorAuthor

    架构层跟踪(kernel 不驻留时定时任务不跑 + 付费分层):objectstack-ai/cloud#1288

  2. added theissue type on Aug 13, 2026
  3. hotlong commented on Aug 13, 2026

    @hotlong
    Contributor

    Triage: lands in packages/services/service-job (DbJobAdapter.destroy() missing this.cron destruction — the single-point root cause) + packages/triggers/trigger-schedule (job-name namespacing, replace-semantics rebind) — both families are domain:services per the lane table, single lane. → pm:queue, type Bug.

    target:v17 — release-board class ① (published-surface defect users hit today): kernel eviction is routine in the cloud runtime (freshness probe, every auto-publish), so "AI builds a scheduled automation → one metadata edit → the automation silently dies, permanently" is the normal path, not an edge case. The failure is silent (one WARN nobody watches) and un-namespaced job names additionally collide across environments in one container without any eviction involved.

    Dedup checked: #4174 and #4567 are closed, distinct defects (pre-kernel:ready registration; expression-envelope parsing). The architecture half (clock should live above the kernel) is already tracked at objectstack-ai/cloud#1288 per the filer's comment — out of this card's scope.

    Dispatch notes for the lane seat: the card carries a four-layer root-cause chain with two control experiments already run — premise quality is high, but re-verify line refs on origin/main at claim time (time-relative-trigger.ts:214 etc. will drift). Fix item 1 (destroy chain) is mechanical; items 2–3 (namespacing + replace semantics) restore the declared invariant; item 4 (observability escalation) should not silently expand into a new status surface — if it grows beyond a log/health signal, flag instead of inventing shape. Regression test per the card's item 5 must cover BOTH time-relative-trigger and schedule-trigger.

    Size/model suggestion: M, mode:subagent, model: opus.


    Generated by Claude Code

  4. self-assigned this
    on Aug 13, 2026
  5. os-zhuang commented on Aug 13, 2026

    @os-zhuang
    ContributorAuthor

    Claim: PM loop round 14
    Session: session_01ARidKDYSCD56LaygrvDPnk
    Branch: claude/issue-8362-cron-rebind-after-kernel-rebuild
    Worktree: objectstack-issue-8362
    Domain: domain:services
    File surface: packages/services/service-job/src/**, packages/triggers/trigger-schedule/src/** (stop on breach; explain in the report)
    Container & model: M, mode:subagent, model: opus
    Serial constraints cleared: in-flight 0 in this lane at claim time. No open claim touches service-job or trigger-schedule. The lane's other pm:queue cards (#8367 service-cluster seam, #8378 plugin read paths, #8231 examples i18n, #8069 service-messaging) are all file-disjoint from this surface.

    Selected first on the target: board rule — release-blocker class ① (published-surface defect users hit today) outranks the ordinary queue.

    Root cause re-verified on origin/main at claim time, ⛔ not taken from the card

    The card's single-point claim holds, exactly:

    db-job-adapter.ts:194   async destroy(): Promise<void> {
    db-job-adapter.ts:195     await this.inner.destroy();     ← only `inner`; `this.cron` never destroyed
    db-job-adapter.ts:196   }
    

    and the method that is never called already exists, with a doc comment naming the caller that does not call it:

    cron-job-adapter.ts:122   /** Stop all timers — call from plugin destroy. */
    cron-job-adapter.ts:123   async destroy(): Promise<void> {
    

    ⇒ Declared-but-not-wired, the same family as #8145 this lane closed today (a sink slot documented since Phase 3, passed undefined). Worth noting because it is now twice in one day that a documented wiring point was simply never connected, and both were invisible until someone measured runtime behaviour.

    ⚠️ Line refs for the trigger layer will drift — my grep for flow-time-relative: / flow-schedule: on origin/main returns only test files, so the production sites build those names by template literal and the card's time-relative-trigger.ts:214 is not directly greppable. ⛔ Re-derive every line ref at implementation time; ⛔ do not trust :214.

    Scope note carried from triage

    Both families (service-job, trigger-schedule) are domain:services per the lane table ⇒ single lane, no cross-domain split needed. The architecture half — "the clock should live above the kernel, so scheduled work does not depend on a resident kernel" — is out of scope and already tracked at objectstack-ai/cloud#1288. ⛔ Do not drift into it.


    Generated by Claude Code

  6. os-zhuang commented on Aug 13, 2026

    @os-zhuang
    ContributorAuthor
    {
      "issue": 8362,
      "status": "done",
      "branch": "claude/issue-8362-cron-rebind-after-kernel-rebuild",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/8462",
      "premise_still_valid": true,
      "summary": "All five asks implemented. (1) DbJobAdapter.destroy() now destroys the cron adapter — the single-point root cause, re-derived at db-job-adapter.ts:194-196 on my branch tip; JobServicePlugin also releases the cron adapter it owns on the `adapter:'cron'` path, where nothing else would (the same leak one branch over, found while fixing this and in scope). (2)+(3) CronJobAdapter scopes its entry in croner's process-global registry to the adapter INSTANCE and REPLACES a name a foreign job still holds by stopping that job. Namespacing is per-instance, NOT per-environment, deliberately: a rebuilt kernel reuses the environment id, so an env-scoped namespace would not fix the rebuild case at all — the per-instance key fixes both it and the cross-environment collision. Fix order handled as the card demands: the destroy chain and the reclaim path both STOP the old job (asserted via croner's isStopped()), so no configuration produces a zombie double-write. (4) Both triggers report a failed bind at `error` with consequence + remedy, grounded in AGENTS.md's degradation-log-level rule (persisted vs runtime state disagree ⇒ error), while the neighbouring 'no job service' composition branch stays at `warn` because that rule names it as correctly warn. I STOPPED there per triage's constraint — no new status surface. I considered and rejected writing bind failures into sys_job.last_status: that column is a run status with decisions recorded in its TSDoc, and overloading it is exactly the shape the constraint forbids. (5) Regression coverage for both triggers. TWO SCOPE NOTES for you: (a) one edit outside `src/**` — `croner` added to trigger-schedule's devDependencies (test-only), so the trigger-seam pins run against the real process-global registry instead of a hand-modelled Map; it mirrors the existing `service-automation` devDep backing that package's e2e test, and it avoids growing either shrink-only ledger in check:test-source-alias / check:type-source-resolution (croner is not a workspace package). (b) I did NOT wire the trigger tests to the real adapters cross-package: that would have required a vitest.config.ts AND a tsconfig paths block for trigger-schedule, and both gates carry set-equality shrink-only entries for it — churn in ledgers you did not name, on a release blocker. The full real-adapter chain is pinned in service-job instead, where the fix lives.",
      "tests": "Build closure first (fresh worktree): pnpm --filter '@objectstack/service-job^...' build, same for trigger-schedule (its pre-existing schedule-runas-e2e.test.ts fails on an unbuilt service-automation dist — the known stale-dist trap, not my change). RED, measured on pristine origin/main with the tests added and NO implementation: 'Test Files 2 failed | 6 passed (8) / Tests 4 failed | 67 passed (71)' — (i) destroy-chain case: 'expect(job.isStopped()).toBe(true)' -> 'AssertionError: expected false to be true'; (ii) rebuild-rebind case: \"Error: Cron: Tried to initialize new named job 'flow-time-relative:xqao_contract_expiry_reminder_flow', but name already taken.\" at CronJobAdapter.schedule src/cron-job-adapter.ts:75 — byte-identical to the WARN in the card; (iii) two-live-adapters case: same throw with NO eviction involved (the cross-environment defect); (iv) reclaim case: 'TypeError: adapterA.cronRegistryName is not a function' (new capability, no main-side equivalent — the main-side shape of that defect is (ii)). GREEN after the fix: service-job 'Test Files 8 passed (8) / Tests 71 passed (71)'; trigger-schedule 'Test Files 4 passed (4) / Tests 46 passed (46)'; both typecheck Done; eslint exit=0. SECOND REVERSE DIRECTION, predicted before running and reported as observed: reverting ONLY the two trigger sources to origin/main (service-job fix left in) turns the two observability cases RED and leaves the two trigger-seam rebind cases GREEN — 'Tests 2 failed | 44 passed (46)', both failures 'AssertionError: expected [] to have a length of 1' (main's error channel is empty; it logged warn). That is the honest direction: the rebind fix lives in the adapter, not in the triggers, so those two trigger-seam cases are coverage pins rather than red-on-main evidence — the red-on-main rebind evidence is (i)-(iii) against the real adapters and the real croner registry. Vacuity traps closed explicitly: every case uses the cron path (interval bypasses croner's named registry per the card's control experiment); the first bind is asserted registered before any second-bind assertion; 'exactly once' is asserted as a count AND the surviving croner job is then fired, asserting the NEW kernel's callback ran and the evicted kernel's did not; the zombie check holds the old Cron object across the rebuild and asserts oldJob.isStopped(), not merely that a new job exists. GATES re-derived for my actual diff with scripts/pm/dispatch-gates.mjs (it surfaced five beyond your list: check:changeset-gate-self-tests, check:objectui-changeset, check-changeset-no-major.mjs, check:query-options-erasure, check:type-check-coverage). Run locally, all PASS: check:nul-bytes, check:docs-audit-scope, check:test-source-alias, check:type-source-resolution, check:engine-double-contract, check:durability-log-level, check:startup-registry-verdict, check:query-options-erasure, check:changeset-gate-self-tests, check:objectui-changeset, check-changeset-fixed.mjs, check-changeset-no-major.mjs. check:engine-double-contract stayed green because I added cases to the existing db-job-adapter.test.ts rather than declaring a second engine double (a new (file,verb) pair would not be in its shrink-only baseline). CI status at report time: in_progress — reporting at draft-PR time per the 2026-08-10 ruling.",
      "open_questions": [],
      "out_of_scope_findings": []
    }

    Draft PR: #8462. Changeset included (.changeset/cron-rebind-after-kernel-rebuild.md, patch/patch), so no skip-changeset label applies.


    Generated by Claude Code

  7. os-zhuang commented on Aug 13, 2026

    @os-zhuang
    ContributorAuthor

    Review of PR #8462 — ACCEPT, pending CI convergence

    Diff read in full. Path face: 11 files, no docs/adr/** · .claude/skills/** · skills/** · content/docs/releases/**. Outside src/**: only trigger-schedule/package.json (+1, the croner devDep) and its pnpm-lock.yaml entry — declared, and the mechanical consequence of that one line.

    ⚠️ The card's fix item 2 was wrong, and the dev corrected it rather than implementing it

    The card asked for "job 名加 environment/kernel 命名空间". The implementation is scoped per adapter instance, not per environment, and the reason is decisive:

    a kernel rebuild produces a new adapter for the same environment id, which is exactly the collision an environment-scoped namespace would fail to prevent

    ⇒ An env-scoped namespace would not have fixed the reported bug at all — the four-consecutive-rebuild reproduction is one environment rebinding its own flow. Per-instance fixes both that and the cross-environment collision the card separately names. The namespace option survives as a cosmetic label for reading scheduledJobs in a multi-tenant container, explicitly documented as never load-bearing for uniqueness.

    That is the second prescription this lane has had corrected by measurement today, and both times the card was written by someone who had done real work — a good reminder that ⛔ a card's fix list is a hypothesis on the same footing as its root cause.

    The fix-order hazard I flagged is closed, and pinned as such

    I warned that fixing naming without fixing the destroy chain converts a silent death into zombie double-writes — the leaked croner job is not merely holding a string, its closure still references the shut-down kernel's engine. Both paths now stop the old job:

    • reclaimRegistryName() calls holder.stop() before taking the name, with the reasoning in the code: "Taking the name while leaving that timer running would turn a silent death into a zombie double-write … strictly worse than the bug being fixed."
    • The pins assert oldJob.isStopped() and fired === ['new-kernel'] — ⛔ not merely that a new job exists.

    A second leak, found and fixed in scope

    JobServicePlugin.destroy() now also releases the cron adapter it owns on the adapter: 'cron' path — where DbJobAdapter is not in the picture and nothing else would. The card did not name this branch. Same defect, one branch over.

    Fix 4 stopped exactly where triage told it to

    Bind failure now reports at error with consequence and remedy, and the level is argued from AGENTS.md's degradation rule rather than asserted: persisted and runtime state disagree while every outside surface keeps reporting healthy ⇒ error class. The neighbouring "no job service registered" branch deliberately stays warn, because the rule names that case as correctly warn — "the distinction is not the severity of the outcome, it is whether the outside can see it."

    And it stopped at the constraint: it considered writing bind failures into sys_job.last_status and rejected it — that column is a run status with its decisions recorded in its TSDoc, and overloading it "is exactly the shape the constraint forbids." ⛔ No new status surface invented. That is the STOP working as intended.

    Vacuity — all three traps I named are closed, explicitly

    • Cron path only. Every case goes through cron; the card's own control experiment showed interval bypasses croner's named registry entirely and would pass on a completely unfixed tree.
    • First bind asserted before second. expect(registeredFor(NAME)).toHaveLength(1) precedes every rebind assertion — "a rebind pin whose first bind registered nothing passes for the wrong reason."
    • "Scheduled" vs "fires". Count and trigger(), asserting the new kernel's callback ran and the evicted kernel's did not.

    Backed by the real process-global croner registry rather than a hand-modelled Map — which is what the croner devDep buys, and why that one line outside src/** is worth its scope note.

    Reverse verification, honestly apportioned

    RED on pristine origin/main with tests but no implementation: 4 failures, including the card's own string byte-identical — Cron: Tried to initialize new named job 'flow-time-relative:xqao_contract_expiry_reminder_flow', but name already taken.

    And the second direction is reported against the PR's own convenience: reverting only the two trigger sources leaves the trigger-seam rebind cases GREEN, because the rebind fix lives in the adapter. The dev says so plainly — "those two trigger-seam cases are coverage pins rather than red-on-main evidence" — and points at the service-job cases as the actual red-on-main proof. ⛔ It did not let a green-both-ways case sit in the table implying it proved something.

    Scope note (b), accepted

    It declined to wire the trigger tests to the real adapters cross-package, because that needs a vitest.config.ts and a tsconfig paths block, and both check:test-source-alias / check:type-source-resolution carry set-equality shrink-only entries for that package ⇒ ledger churn nobody asked for, on a release blocker. Right call: it pinned the real-adapter chain in service-job, where the fix lives, and left the ledgers alone.

    Next

    Flipping ready + enqueueing once ESLint, TypeScript Type Check, Test Core ×3, Dogfood ×3 and Temporal Conformance conclude success. ⛔ Not before. This is a target:v17 release-board card, so it also comes off the board on merge — I will confirm the board reading rather than assume it.


    Generated by Claude Code

  8. os-zhuang commented on Aug 13, 2026

    @os-zhuang
    ContributorAuthor

    Post-merge inventory — services seat #6021, session session_01ARidKDYSCD56LaygrvDPnk.

    PR #8462 MERGED 2026-08-13T16:02:07Z through the merge queue. All 26 checks concluded success before the flip; queue entry confirmed by the enqueued event, ⛔ not by the auto-merge receipt (that receipt has come back with empty method/enabled at fields every time this shift and carries no information in either direction).

    Auto-close verified, not assumed. The card closed completed with closed_by_pull_requests → #8462. That is correct here: the PR carried Fixes #8362 deliberately, because merging should close this card. ⚠️ I check this explicitly every time since #8131, where a Part of card auto-closed anyway and only a post-merge sweep caught it (mechanism now filed as #8293).

    Label state closed out: dropped pm:dispatched — the card is merged and closed, nothing is in flight, and a lingering dispatch label is exactly the stale state that reads as "someone is on it" to future sweeps. Kept target:v17: it records the release this landed in, and removing it would erase that.

    What landed

    1. DbJobAdapter.destroy() now destroys the cron adapter — the single-point root cause. CronJobAdapter.destroy() had existed all along with a doc comment naming a caller that never called it.
    2. JobServicePlugin releases the cron adapter it owns on the adapter: 'cron' path (same leak, one branch over).
    3. CronJobAdapter scopes its croner registry key per adapter instance — which independently fixes the second defect this card names (two environments in one container colliding on an AI-generated flow name, with no eviction involved).
    4. Claiming a name a foreign job still holds now stops the holder and retakes it. ⭐ The stopping half is load-bearing: taking the name while leaving the old timer running would have traded a silent death for a zombie double-write, which is worse than the bug.
    5. Both triggers report a failed bind at error, per the ADR-0078 valid-but-inert reading this card's body makes.

    ⭐ Worth recording that the card's own item 2 was corrected by measurement: it asked for an environment/kernel namespace, and the dev made it per-instance instead, on the ground that a rebuilt kernel reuses the environment id — which is the collision itself. Namespacing by environment would have looked like a fix and left the reported bug intact. This is the second time this shift a card's prescribed fix turned out to be a hypothesis on the same footing as its root cause.

    Still open, by design

    The architectural half — the clock living above the kernel so scheduled work does not depend on a resident kernel — is not addressed here and remains true after this merge: "scheduled jobs do not run while no kernel is resident." Tracked at objectstack-ai/cloud#1288, as this card's own closing note anticipated.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions