Repository navigation
time-relative/schedule cron flows permanently fail to re-bind after kernel rebuild — croner name leaked because DbJobAdapter.destroy() never destroys the cron adapter #8362
Description
Activity
架构层跟踪(kernel 不驻留时定时任务不跑 + 付费分层):objectstack-ai/cloud#1288
- addedbugSomething isn't workingSomething isn't working
on Aug 13, 2026 Triage: lands in
packages/services/service-job(DbJobAdapter.destroy()missingthis.crondestruction — the single-point root cause) +packages/triggers/trigger-schedule(job-name namespacing, replace-semantics rebind) — both families aredomain:servicesper the lane table, single lane. →pm:queue, type Bug.target:v17— release-board class ① (published-surface defect users hit today): kernel eviction is routine in the cloud runtime (freshness probe, every auto-publish), so "AI builds a scheduled automation → one metadata edit → the automation silently dies, permanently" is the normal path, not an edge case. The failure is silent (one WARN nobody watches) and un-namespaced job names additionally collide across environments in one container without any eviction involved.Dedup checked: #4174 and #4567 are closed, distinct defects (pre-
kernel:readyregistration; expression-envelope parsing). The architecture half (clock should live above the kernel) is already tracked at objectstack-ai/cloud#1288 per the filer's comment — out of this card's scope.Dispatch notes for the lane seat: the card carries a four-layer root-cause chain with two control experiments already run — premise quality is high, but re-verify line refs on
origin/mainat claim time (time-relative-trigger.ts:214etc. will drift). Fix item 1 (destroy chain) is mechanical; items 2–3 (namespacing + replace semantics) restore the declared invariant; item 4 (observability escalation) should not silently expand into a new status surface — if it grows beyond a log/health signal, flag instead of inventing shape. Regression test per the card's item 5 must cover BOTHtime-relative-triggerandschedule-trigger.Size/model suggestion: M,
mode:subagent,model: opus.
Generated by Claude Code
Claim: PM loop round 14
Session:session_01ARidKDYSCD56LaygrvDPnk
Branch:claude/issue-8362-cron-rebind-after-kernel-rebuild
Worktree:objectstack-issue-8362
Domain:domain:services
File surface:packages/services/service-job/src/**,packages/triggers/trigger-schedule/src/**(stop on breach; explain in the report)
Container & model:M,mode:subagent,model: opus
Serial constraints cleared: in-flight 0 in this lane at claim time. No open claim touchesservice-jobortrigger-schedule. The lane's otherpm:queuecards (#8367service-clusterseam, #8378 plugin read paths, #8231 examples i18n, #8069service-messaging) are all file-disjoint from this surface.Selected first on the
target:board rule — release-blocker class ① (published-surface defect users hit today) outranks the ordinary queue.Root cause re-verified on
origin/mainat claim time, ⛔ not taken from the cardThe card's single-point claim holds, exactly:
db-job-adapter.ts:194 async destroy(): Promise<void> { db-job-adapter.ts:195 await this.inner.destroy(); ← only `inner`; `this.cron` never destroyed db-job-adapter.ts:196 }and the method that is never called already exists, with a doc comment naming the caller that does not call it:
cron-job-adapter.ts:122 /** Stop all timers — call from plugin destroy. */ cron-job-adapter.ts:123 async destroy(): Promise<void> {⇒ Declared-but-not-wired, the same family as #8145 this lane closed today (a sink slot documented since Phase 3, passed
undefined). Worth noting because it is now twice in one day that a documented wiring point was simply never connected, and both were invisible until someone measured runtime behaviour.⚠️ Line refs for the trigger layer will drift — my grep forflow-time-relative:/flow-schedule:onorigin/mainreturns only test files, so the production sites build those names by template literal and the card'stime-relative-trigger.ts:214is not directly greppable. ⛔ Re-derive every line ref at implementation time; ⛔ do not trust:214.Scope note carried from triage
Both families (
service-job,trigger-schedule) aredomain:servicesper the lane table ⇒ single lane, no cross-domain split needed. The architecture half — "the clock should live above the kernel, so scheduled work does not depend on a resident kernel" — is out of scope and already tracked atobjectstack-ai/cloud#1288. ⛔ Do not drift into it.
Generated by Claude Code
{ "issue": 8362, "status": "done", "branch": "claude/issue-8362-cron-rebind-after-kernel-rebuild", "pr": "https://github.com/objectstack-ai/objectstack/pull/8462", "premise_still_valid": true, "summary": "All five asks implemented. (1) DbJobAdapter.destroy() now destroys the cron adapter — the single-point root cause, re-derived at db-job-adapter.ts:194-196 on my branch tip; JobServicePlugin also releases the cron adapter it owns on the `adapter:'cron'` path, where nothing else would (the same leak one branch over, found while fixing this and in scope). (2)+(3) CronJobAdapter scopes its entry in croner's process-global registry to the adapter INSTANCE and REPLACES a name a foreign job still holds by stopping that job. Namespacing is per-instance, NOT per-environment, deliberately: a rebuilt kernel reuses the environment id, so an env-scoped namespace would not fix the rebuild case at all — the per-instance key fixes both it and the cross-environment collision. Fix order handled as the card demands: the destroy chain and the reclaim path both STOP the old job (asserted via croner's isStopped()), so no configuration produces a zombie double-write. (4) Both triggers report a failed bind at `error` with consequence + remedy, grounded in AGENTS.md's degradation-log-level rule (persisted vs runtime state disagree ⇒ error), while the neighbouring 'no job service' composition branch stays at `warn` because that rule names it as correctly warn. I STOPPED there per triage's constraint — no new status surface. I considered and rejected writing bind failures into sys_job.last_status: that column is a run status with decisions recorded in its TSDoc, and overloading it is exactly the shape the constraint forbids. (5) Regression coverage for both triggers. TWO SCOPE NOTES for you: (a) one edit outside `src/**` — `croner` added to trigger-schedule's devDependencies (test-only), so the trigger-seam pins run against the real process-global registry instead of a hand-modelled Map; it mirrors the existing `service-automation` devDep backing that package's e2e test, and it avoids growing either shrink-only ledger in check:test-source-alias / check:type-source-resolution (croner is not a workspace package). (b) I did NOT wire the trigger tests to the real adapters cross-package: that would have required a vitest.config.ts AND a tsconfig paths block for trigger-schedule, and both gates carry set-equality shrink-only entries for it — churn in ledgers you did not name, on a release blocker. The full real-adapter chain is pinned in service-job instead, where the fix lives.", "tests": "Build closure first (fresh worktree): pnpm --filter '@objectstack/service-job^...' build, same for trigger-schedule (its pre-existing schedule-runas-e2e.test.ts fails on an unbuilt service-automation dist — the known stale-dist trap, not my change). RED, measured on pristine origin/main with the tests added and NO implementation: 'Test Files 2 failed | 6 passed (8) / Tests 4 failed | 67 passed (71)' — (i) destroy-chain case: 'expect(job.isStopped()).toBe(true)' -> 'AssertionError: expected false to be true'; (ii) rebuild-rebind case: \"Error: Cron: Tried to initialize new named job 'flow-time-relative:xqao_contract_expiry_reminder_flow', but name already taken.\" at CronJobAdapter.schedule src/cron-job-adapter.ts:75 — byte-identical to the WARN in the card; (iii) two-live-adapters case: same throw with NO eviction involved (the cross-environment defect); (iv) reclaim case: 'TypeError: adapterA.cronRegistryName is not a function' (new capability, no main-side equivalent — the main-side shape of that defect is (ii)). GREEN after the fix: service-job 'Test Files 8 passed (8) / Tests 71 passed (71)'; trigger-schedule 'Test Files 4 passed (4) / Tests 46 passed (46)'; both typecheck Done; eslint exit=0. SECOND REVERSE DIRECTION, predicted before running and reported as observed: reverting ONLY the two trigger sources to origin/main (service-job fix left in) turns the two observability cases RED and leaves the two trigger-seam rebind cases GREEN — 'Tests 2 failed | 44 passed (46)', both failures 'AssertionError: expected [] to have a length of 1' (main's error channel is empty; it logged warn). That is the honest direction: the rebind fix lives in the adapter, not in the triggers, so those two trigger-seam cases are coverage pins rather than red-on-main evidence — the red-on-main rebind evidence is (i)-(iii) against the real adapters and the real croner registry. Vacuity traps closed explicitly: every case uses the cron path (interval bypasses croner's named registry per the card's control experiment); the first bind is asserted registered before any second-bind assertion; 'exactly once' is asserted as a count AND the surviving croner job is then fired, asserting the NEW kernel's callback ran and the evicted kernel's did not; the zombie check holds the old Cron object across the rebuild and asserts oldJob.isStopped(), not merely that a new job exists. GATES re-derived for my actual diff with scripts/pm/dispatch-gates.mjs (it surfaced five beyond your list: check:changeset-gate-self-tests, check:objectui-changeset, check-changeset-no-major.mjs, check:query-options-erasure, check:type-check-coverage). Run locally, all PASS: check:nul-bytes, check:docs-audit-scope, check:test-source-alias, check:type-source-resolution, check:engine-double-contract, check:durability-log-level, check:startup-registry-verdict, check:query-options-erasure, check:changeset-gate-self-tests, check:objectui-changeset, check-changeset-fixed.mjs, check-changeset-no-major.mjs. check:engine-double-contract stayed green because I added cases to the existing db-job-adapter.test.ts rather than declaring a second engine double (a new (file,verb) pair would not be in its shrink-only baseline). CI status at report time: in_progress — reporting at draft-PR time per the 2026-08-10 ruling.", "open_questions": [], "out_of_scope_findings": [] }Draft PR: #8462. Changeset included (
.changeset/cron-rebind-after-kernel-rebuild.md, patch/patch), so noskip-changesetlabel applies.
Generated by Claude Code
Review of PR #8462 — ACCEPT, pending CI convergence
Diff read in full. Path face: 11 files, no
docs/adr/**·.claude/skills/**·skills/**·content/docs/releases/**. Outsidesrc/**: onlytrigger-schedule/package.json(+1, thecronerdevDep) and itspnpm-lock.yamlentry — declared, and the mechanical consequence of that one line.⚠️ The card's fix item 2 was wrong, and the dev corrected it rather than implementing itThe card asked for "job 名加 environment/kernel 命名空间". The implementation is scoped per adapter instance, not per environment, and the reason is decisive:
a kernel rebuild produces a new adapter for the same environment id, which is exactly the collision an environment-scoped namespace would fail to prevent
⇒ An env-scoped namespace would not have fixed the reported bug at all — the four-consecutive-rebuild reproduction is one environment rebinding its own flow. Per-instance fixes both that and the cross-environment collision the card separately names. The
namespaceoption survives as a cosmetic label for readingscheduledJobsin a multi-tenant container, explicitly documented as never load-bearing for uniqueness.That is the second prescription this lane has had corrected by measurement today, and both times the card was written by someone who had done real work — a good reminder that ⛔ a card's fix list is a hypothesis on the same footing as its root cause.
The fix-order hazard I flagged is closed, and pinned as such
I warned that fixing naming without fixing the destroy chain converts a silent death into zombie double-writes — the leaked croner job is not merely holding a string, its closure still references the shut-down kernel's engine. Both paths now stop the old job:
reclaimRegistryName()callsholder.stop()before taking the name, with the reasoning in the code: "Taking the name while leaving that timer running would turn a silent death into a zombie double-write … strictly worse than the bug being fixed."- The pins assert
oldJob.isStopped()andfired === ['new-kernel']— ⛔ not merely that a new job exists.
A second leak, found and fixed in scope
JobServicePlugin.destroy()now also releases the cron adapter it owns on theadapter: 'cron'path — whereDbJobAdapteris not in the picture and nothing else would. The card did not name this branch. Same defect, one branch over.Fix 4 stopped exactly where triage told it to
Bind failure now reports at
errorwith consequence and remedy, and the level is argued from AGENTS.md's degradation rule rather than asserted: persisted and runtime state disagree while every outside surface keeps reporting healthy ⇒errorclass. The neighbouring "no job service registered" branch deliberately stayswarn, because the rule names that case as correctly warn — "the distinction is not the severity of the outcome, it is whether the outside can see it."And it stopped at the constraint: it considered writing bind failures into
sys_job.last_statusand rejected it — that column is a run status with its decisions recorded in its TSDoc, and overloading it "is exactly the shape the constraint forbids." ⛔ No new status surface invented. That is the STOP working as intended.Vacuity — all three traps I named are closed, explicitly
- Cron path only. Every case goes through
cron; the card's own control experiment showedintervalbypasses croner's named registry entirely and would pass on a completely unfixed tree. - First bind asserted before second.
expect(registeredFor(NAME)).toHaveLength(1)precedes every rebind assertion — "a rebind pin whose first bind registered nothing passes for the wrong reason." - "Scheduled" vs "fires". Count and
trigger(), asserting the new kernel's callback ran and the evicted kernel's did not.
Backed by the real process-global croner registry rather than a hand-modelled Map — which is what the
cronerdevDep buys, and why that one line outsidesrc/**is worth its scope note.Reverse verification, honestly apportioned
RED on pristine
origin/mainwith tests but no implementation: 4 failures, including the card's own string byte-identical —Cron: Tried to initialize new named job 'flow-time-relative:xqao_contract_expiry_reminder_flow', but name already taken.And the second direction is reported against the PR's own convenience: reverting only the two trigger sources leaves the trigger-seam rebind cases GREEN, because the rebind fix lives in the adapter. The dev says so plainly — "those two trigger-seam cases are coverage pins rather than red-on-main evidence" — and points at the service-job cases as the actual red-on-main proof. ⛔ It did not let a green-both-ways case sit in the table implying it proved something.
Scope note (b), accepted
It declined to wire the trigger tests to the real adapters cross-package, because that needs a
vitest.config.tsand atsconfigpaths block, and bothcheck:test-source-alias/check:type-source-resolutioncarry set-equality shrink-only entries for that package ⇒ ledger churn nobody asked for, on a release blocker. Right call: it pinned the real-adapter chain inservice-job, where the fix lives, and left the ledgers alone.Next
Flipping ready + enqueueing once ESLint, TypeScript Type Check, Test Core ×3, Dogfood ×3 and Temporal Conformance conclude
success. ⛔ Not before. This is atarget:v17release-board card, so it also comes off the board on merge — I will confirm the board reading rather than assume it.
Generated by Claude Code
Post-merge inventory — services seat #6021, session
session_01ARidKDYSCD56LaygrvDPnk.PR #8462 MERGED 2026-08-13T16:02:07Z through the merge queue. All 26 checks concluded
successbefore the flip; queue entry confirmed by theenqueuedevent, ⛔ not by the auto-merge receipt (that receipt has come back with emptymethod/enabled atfields every time this shift and carries no information in either direction).Auto-close verified, not assumed. The card closed
completedwithclosed_by_pull_requests→ #8462. That is correct here: the PR carriedFixes #8362deliberately, because merging should close this card.⚠️ I check this explicitly every time since #8131, where aPart ofcard auto-closed anyway and only a post-merge sweep caught it (mechanism now filed as #8293).Label state closed out: dropped
pm:dispatched— the card is merged and closed, nothing is in flight, and a lingering dispatch label is exactly the stale state that reads as "someone is on it" to future sweeps. Kepttarget:v17: it records the release this landed in, and removing it would erase that.What landed
DbJobAdapter.destroy()now destroys the cron adapter — the single-point root cause.CronJobAdapter.destroy()had existed all along with a doc comment naming a caller that never called it.JobServicePluginreleases the cron adapter it owns on theadapter: 'cron'path (same leak, one branch over).CronJobAdapterscopes its croner registry key per adapter instance — which independently fixes the second defect this card names (two environments in one container colliding on an AI-generated flow name, with no eviction involved).- Claiming a name a foreign job still holds now stops the holder and retakes it. ⭐ The stopping half is load-bearing: taking the name while leaving the old timer running would have traded a silent death for a zombie double-write, which is worse than the bug.
- Both triggers report a failed bind at
error, per the ADR-0078 valid-but-inert reading this card's body makes.
⭐ Worth recording that the card's own item 2 was corrected by measurement: it asked for an environment/kernel namespace, and the dev made it per-instance instead, on the ground that a rebuilt kernel reuses the environment id — which is the collision itself. Namespacing by environment would have looked like a fix and left the reported bug intact. This is the second time this shift a card's prescribed fix turned out to be a hypothesis on the same footing as its root cause.
Still open, by design
The architectural half — the clock living above the kernel so scheduled work does not depend on a resident kernel — is not addressed here and remains true after this merge: "scheduled jobs do not run while no kernel is resident." Tracked at objectstack-ai/cloud#1288, as this card's own closing note anticipated.
Generated by Claude Code
- added a commit that references this issue
on Aug 17, 2026
现象
本地 magic-flow rig(cloud main
3f2884c4× framework maincb43296ef,2026-08-13)实测:一个timeRelativeflow 首次绑定成功后,每一次 kernel 重建都无法重绑,且永不自愈(连续 4 次重建复现):之后 sweep 绑在空处,自动化静默死亡——
verify_build、UI、DB 全部显示健康,唯一信号是这条 WARN。同一批重建里record-changetrigger 每次都正常重绑,对照明显。两个对照实验钉住根因是残留的进程级 job 名:
{type:'interval'}→ 重绑成功(interval 走setInterval,不经过 croner 命名注册表);换回 cron 型立刻name already taken。sys_metadata里改 flow 名(= 换 job 名)→ 下一次重建绑定成功。根因链(四层)
packages/triggers/trigger-schedule/src/time-relative-trigger.ts:214:job 名flow-time-relative:<flowName>,不带 env/kernel 标识;绑定前的this.stop(flowName)读实例级this.boundMap——新 kernel 的新实例 Map 为空,清理 no-op。schedule-trigger.ts(flow-schedule:<flowName>)结构完全相同,同样中招。cron-job-adapter.ts:schedule()先await this.cancel(name),但也只查自己实例的this.jobsMap。new Cron(expression, { name })登记进 croner 进程级全局命名注册表——旧 kernel 的 Cron 对象还在,新建同名直接 throw。db-job-adapter.ts:kernel 驱逐链路本身是完整的(KernelManager.evict()→kernel.shutdown()→plugin.destroy()→JobServicePlugin.destroy()→dbAdapter.destroy()),但DbJobAdapter.destroy()只有await this.inner.destroy()——漏掉this.cron(CronJobAdapter)。旧 kernel 的 croner 任务永不 stop,永久占名。另外 trigger 的
schedule()是 fire-and-forget(void Promise.resolve(...)),失败被降为一条 WARN。为什么严重
contract_expiry_reminder_flow)。期望修复
DbJobAdapter.destroy()补上this.cron的销毁(CronJobAdapter.destroy()已存在,只是没人调)。time-relative-trigger与schedule-trigger两处。注意:即使全部修完,「kernel 不驻留时定时任务不跑」仍是架构层问题(时钟应上移到常驻层),cloud 侧另行跟踪。