Repository navigation
time-relative sweep 非幂等:每次扫描对同一记录重复铸行(5s interval 下 70 秒 15 条重复提醒),需要落库幂等键 #10220
Description
Activity
Claim: PM loop round 1 (epic objectstack-ai/cloud#1472)
Session:session_3f79f27a-67a1-44c5-b9e3-f9e08b0aba46
Branch:claude/issue-10220-time-relative-idempotency
Worktree:objectstack-issue-10220
Domain: epic lane; lands inpackages/triggers/trigger-schedule
File surface:packages/triggers/trigger-schedule/**, plus the run-record store it reuses if one fits (packages/services/service-automation/**read-only unless the dedup record lands there — declare in report); tests + changeset (stop on breach; explain in the report)
Container & model:M/L, mode:subagent, model: opus
Clause-②: no (execution-dedup of an existing trigger; no contract accept/reject change intended — if the fix turns out to need a new spec key, stop and return needs_decision)
Serial constraints cleared: siblings #10219 (metadata-protocol/rest) and #10221 (settings read) — disjoint packages.Triage (labels only — card is already claimed and in flight):
domain:services, type Bug; rationale: the time-relative sweep lives inpackages/services/service-automation/src/engine.ts, and the fix (a durable idempotency key so a rescan does not re-mint rows for the same record/threshold) lands there.
本评论来自分诊座位 Routine
Generated by Claude Code
Generated by Claude Code
{
"issue": 10220,
"status": "needs_decision",
"branch": "claude/issue-10220-time-relative-idempotency",
"pr": null,
"premise_still_valid": true,
"summary": "Premise confirmed against origin/main (ce300c8): TimeRelativeTrigger.sweep() in packages/triggers/trigger-schedule/src/time-relative-trigger.ts holds ZERO cross-sweep memory — its only Set (seenIds, line 269) is declared inside sweep() and dedups across WINDOWS within one tick, so every tick re-dispatches the same (record, window) pair exactly as reported. PM mechanism assumption 1 is REFUTED by measurement, not by opinion:sys_automation_runcannot serve as the dedup store (evidence in open_questions Q1 option B). The clean fix therefore needs a persisted per-key ledger, which means a NEW platform object — and the repo mechanically forces that to be apackages/specedit, becausepackages/spec/src/system/constants/platform-object-names.test.tsscans every*.object.tsin the monorepo and fails unless PLATFORM_OBJECTS_BY_PACKAGE lists it. That is the contract-surface expansion PM assumption 2 says to stop on, and it is outside my declared file surface, so I stopped before the first edit: no code, no PR. A SECOND, independent contract question surfaced while deriving the key (open_questions Q2): the issue's prescribed key would silently change a semantic the spec prose explicitly promises forwithinDays.",
"tests": "No build/test was run and none is quoted: this run produced no code change, so any test output would be a measurement of origin/main, not of a fix. Reporting it as evidence would be false. All findings below are source-level reads of the fetched origin/main (ce300c8) with file:line citations, each independently checkable.",
"open_questions": [
{
"question": "Q1 (blocking, contract): where does the persisted idempotency record live? A persisted ledger is required by fix requirement (2), and in this repo any new*.object.tsmechanically forces apackages/specregistry edit (platform-object-names.test.ts: 'registers exactly the objects each package declares' — a package that declares an object with no registry group FAILS). That is the new-platform-object stop condition in PM assumption 2, and packages/spec is outside my declared file surface.",
"options": [
"A (RECOMMENDED) — New object, e.g.sys_flow_dispatch, owned by@objectstack/service-automation(which already has a spec registry group and already owns flow-run durability), exposed to triggers as a structurally-typedclaim(key): Promise<boolean>method on the automation service — the surface TimeRelativeTriggerPlugin ALREADY resolves (resolveService('automation'), time-relative-plugin.ts:51). trigger-schedule computes the key and calls claim(); it stays structurally decoupled (it has @objectstack/service-automation as a devDependency only — package.jsondependenciesis exactly {core, spec}) and never learns another service's table name. Delta: one *.object.ts + one name in PLATFORM_OBJECTS_BY_PACKAGE + one service method. Declare ADR-0057 telemetry retention on it (every key is producible on exactly ONE sweep day, so rows are dead weight after a few days) — but the window must be >= the catch-up horizon cloud#1288 will want, which is a decision input, not a default I should pick alone.",
"B (REFUTED — reusesys_automation_run, PM assumption 1) — Five independent refutations, all in packages/services/service-automation/src: (1) WRITE-TIME CAP: suspended-run-store.ts:42 DEFAULT_MAX_TERMINAL_RUNS_PER_FLOW = 100 and pruneFlowOverflow() (:317-337) deletes every terminal row of a flow beyond the newest 100 on each terminal write — a sweep over >100 records ERASES ITS OWN DEDUP EVIDENCE inside a single sweep. This alone is fatal. (2) NO WINDOW COLUMN: sys-automation-run.object.ts carries flow_name / trigger_object / trigger_record_id / started_at and nothing that identifies the matched window; started_at is the RUN day, not the window day, and reconstructing one from the other is exactly what a catch-up sweep (cloud#1288) breaks. (3) ROWS ARE OPTIONAL: plugin.ts:646-712 attaches ObjectStoreSuspendedRunStore only when an ObjectQL engine resolves AND the object registered; otherwise the run store is in-memory and persists nothing — i.e. dedup would silently not survive kernel rebuild, which is precisely fix requirement (2). (4) WRITE IS FIRE-AND-FORGET: engine.ts:4668void this.store.recordTerminal(record).catch(...)— not awaited, so a fast tick (the 5s repro) races the marker it would need to read. (5) LAYERING: reading sys_automation_run by table name from the trigger hard-couples trigger-schedule to another service's private storage schema, undoing the deliberate structural decoupling documented on TimeRelativeDataEngine (time-relative-trigger.ts:13-19).",
"C (viable, bigger) — Route dispatch through the queue service and letsys_job_queue.idempotency_keybe the ledger: db-queue-adapter.ts:249-265 already implements a persisted (queue, idempotency_key) claim, default window 24h (:161), with a constructor assert that the window never exceeds sys_job_queue's declared retention (:175-182) — so 'the reaper deleted a row the dedup needed' is unrepresentable. And the window suffices: for every key shape in Q2 option A, all duplicates of one key fall on ONE sweep day. Cost: dispatch moves OFF the sweep tick into a queue worker (the sweep's per-record error isolation, time-relative-trigger.ts:303-326, moves with it), trigger-schedule gains a queue dependency, and correctness becomes coupled to a tuning knob (idempotencyWindowMs) rather than to the key's own lifetime.",
"D (rejected — cheap but wrong) — Per-flow watermark ('this flow already swept window-day D'), storable without a new object. Rejected: it would skip records CREATED later the same day, and the card prescribes a per-record key."
],
"recommendation": "A, on all four axes. REAL BUSINESS NEED: measured, not speculative — 15 duplicate reminder rows in ~70s on a live rig, user-visible as spam (and as repeat notifications when the flow's action is a notify), plus a second named consumer already queued (cloud#1288's catch-up sweep, which the card says is blocked on this key). LONG-TERM SOUNDNESS: dedup belongs at the dispatch boundary as ONE ledger with an explicit contract; B makes correctness depend on a TELEMETRY table whose retention is designed to delete rows (the 100-row cap is a feature there and a silent correctness hole here), and C makes it depend on a queue tuning knob — both are the workaround shape, and B's failure mode is silent duplicate resurrection. AI-AUTHORED-METADATA SAFETY: A adds NO authorable key — the idempotency key is derived entirely from theconfig.timeRelativedescriptor the AI already writes, so declared = enforced and there is nothing new for a generator to mis-spell; B and C by contrast make correctness contingent on an OPERATOR having wired a durable run store / tuned a queue window, which is exactly the valid-but-inert class ADR-0078 exists to refuse. STARTUP SCOPE DISCIPLINE: A's delta is one object, one registry line, one service method — no new spec/authorable key, no new plugin, no config knob; C's 'route dispatch through the queue' is the strictly larger scope expansion. The single cost of A is the one PM assumption 2 reserves for the maintainer: a new platform object. I need the ruling on (i) go/no-go, (ii) object name, (iii) owning package — service-automation (smaller delta, better semantic home next to run history) vs trigger-schedule (needs a brand-new registry group), and (iv) the retention window, which must be >= cloud#1288's catch-up horizon."
},
{
"question": "Q2 (blocking, contract, NOT raised in the card): applying the card's key literally would silently change a documentedwithinDayssemantic. packages/spec/src/automation/time-relative-trigger.zod.ts:154-157 promises, in spec prose, that range mode 'Fires every day the record stays in range'. The card's key is (flowName, recordId, window-day of the dateField VALUE, offset); in range mode there is no offset and the dateField value's day is CONSTANT while the record sits in the window, so that key collapses 'fires every day in range' to 'fires once, ever'. In offset mode the same key is already correct and changes nothing (window day = today + offset, so the key implies the sweep day).",
"options": [
"A (RECOMMENDED) — Key on the MATCHED WINDOW's identity, derived from the DateWindow that computeDateWindows() already returns: offset mode ->(flowName, recordId, windowDay, offset)(identical to the card's key); range mode ->(flowName, recordId, sweepDay, rangeSpec). Kills intra-day duplicates in BOTH modes (the reported defect is intra-day in both), preserves both documented semantics, needs no spec prose change and no new spec key. Re-fire on edit still works in offset mode: changing due_date changes the window day, hence the key.",
"B — Take the card's key literally in both modes. Simpler string, but it is a behaviour change to a semantic the spec explicitly documents, so it requires apackages/specprose edit and a deliberate ruling — a range-mode flow that today reminds daily until the deadline would go quiet after one reminder."
],
"recommendation": "A. REAL BUSINESS NEED: 'remind me every day until it is due' is the documented purpose of range mode (the zod file's own example is 'expiring soon'), and no evidence was produced that anyone wants it silenced — the card's evidence is about DUPLICATES WITHIN ONE TICK, which A removes completely. LONG-TERM SOUNDNESS: A derives the key from the window the sweep already computes, so the key and the matching rule can never drift apart; B introduces a second, undocumented rule for what a window means. AI-SAFETY: under A a generator that writeswithinDaysgets exactly the behaviour the descriptor's own docs describe — under B the docs would be lying until someone edits them, and stale contract prose is the most reliable way to get AI-authored metadata wrong. SCOPE: A is contained in trigger-schedule; B pulls in a spec prose edit for a behaviour nobody asked to change. If the maintainer prefers B, the spec prose edit must ride the SAME change — declared and enforced must not diverge."
}
],
"out_of_scope_findings": []
}Maintainer ruling recorded (2026-08-20, via epic PM session
session_3f79f27a-67a1-44c5-b9e3-f9e08b0aba46, in-session decision UI):Q1 = Option A — new platform object
sys_flow_dispatchowned by@objectstack/service-automation, exposed to triggers as a structurally-typedclaim(key): Promise<boolean>on the automation service; one registry line inPLATFORM_OBJECTS_BY_PACKAGE; ADR-0057 retention declared on it (30-day window — comfortably >= any near-term catch-up horizon for objectstack-ai/cloud#1288; widen there if that work needs more).选项原文(维护者选择,未译):「A:新增 sys_flow_dispatch(推荐)— service-automation 名下新增专用对象 + claim(key) 服务方法;trigger-schedule 只算 key 不碰表名。增量最小(1 对象+1 注册行+1 方法),无新增可授权键(AI 无从写错),并为 cloud#1288 的 catch-up sweep 铺路。带 ADR-0057 保留策略(建议 30 天)。」
Q2 = Option A (PM-adjudicated within the epic lane: it is the option that does NOT change contract semantics — key on the matched window's identity; offset mode
(flowName, recordId, windowDay, offset), range mode(flowName, recordId, sweepDay, rangeSpec); spec prose forwithinDaysstays true). If the maintainer wants option B instead, that is a spec prose change and must be re-raised.Card returns to work state; re-dispatching with the ruling (model: fable per clause-② — contract surface expands by one platform object).
- added a commit that references this issue
on Aug 20, 2026 { "issue": 10220, "status": "done", "branch": "claude/issue-10220-flow-dispatch-ledger", "pr": "https://github.com/objectstack-ai/objectstack/pull/10316", "premise_still_valid": true, "summary": "Implemented per the 2026-08-20 ruling (Q1=A, Q2=A). New platform object sys_flow_dispatch owned by @objectstack/service-automation (ADR-0057 telemetry retention 30d), registered with the manifest alongside sys_automation_run; one authorized registry line in PLATFORM_OBJECTS_BY_PACKAGE. AutomationEngine.claim(key) is check-and-record against the persisted ledger (key = row id, so an insert race resolves to false via re-read); with no ledger it degrades to in-process dedup with a ONE-TIME warning, and a ledger ERROR falls back per-key so a store outage never blocks a dispatch (availability over strict-once — stated in the PR body, which also names cloud#1288 catch-up sweeps as the unblocked consumer). TimeRelativeTrigger claims a key derived from the matched window's identity via a single derivation (computeWindowClaimScopes) shared with the query windows: offset mode (flowName, recordId, windowDay, offset) — a dateField edit re-fires for the new window; range mode (flowName, recordId, sweepDay, rangeSpec) — withinDays prose stays true, never twice in one day. The trigger resolves the claim surface structurally from the automation service it already resolves (never learns the table name). Sibling #10219 is not addressed here; its seam (resyncFlowsFromProtocol / metadata:reloaded) is untouched.", "tests": "All at final head cc49bead0, clean tree, run directly per the dispatch prompt's macOS carve-out (os-verify-lock.sh error-loops on macOS bash 3.2 — noting the prompt-vs-role-file conflict as instructed; exit codes captured before any pipe throughout). trigger-schedule: 'Test Files 4 passed (4)' / 'Tests 57 passed (57)' incl. the four required regressions (same-window twice -> once; fresh trigger instance + surviving ledger -> still once; range next-day re-fire; offset dateField-move re-fire) plus claim-failure-still-dispatches and one-time degradation warn; typecheck exit 0. service-automation: 'Test Files 84 passed (84)' / 'Tests 998 passed (998)' incl. new flow-dispatch.test.ts (store check-and-record, insert-race -> false, store-error propagation, engine rebuild-over-shared-store, no-ledger warn-once, error-fallback). spec registry: platform-object-names 'Tests 7 passed (7)'. Full workspace build FULL_BUILD_EXIT=0 (turbo, 70 tasks, 0 cached). dispatch-gates.mjs battery at head all EXIT=0: cross-package-test-inputs, engine-double-contract, where-matcher, query-options-erasure, type-check-coverage, type-check-debt, nul-bytes, slot-lookup, test-source-alias, type-source-resolution, spec-parsed-alias, merge-driver, adr-0087, changeset-no-major, empty-changeset, affected-docs, dev-prereqs, spec liveness/empty-state/strictness-ledger/variant-docs, doc-formula-expressions. Two reds are PRE-EXISTING macOS-only self-test harness failures (check-adr-0087-registration --self-test I2, objectui-changeset-digest --self-test ENOENT) — reproduced identically on the pristine shared checkout at a different HEAD (68f65ff60), already filed as #10303 (#10086 names the symlink entry-guard class). Also pre-existing: service-automation has no typecheck script; direct tsc --noEmit shows 3 private-access errors confined to untouched nested-region-parity.test.ts.", "open_questions": [], "out_of_scope_findings": ["#10303 (already filed by a sibling, cited not duplicated): check-adr-0087-registration.mjs and objectui-changeset-digest.mjs --self-test fail on macOS unrelated to any diff"] }- added a commit that references this issue
on Aug 23, 2026 - added a commit that references this issue
on Oct 7, 2026
现象(P1:time-relative 扫描非幂等——每次扫描对同一记录重复铸行)
2026-08-20 rig 实测(objectstack
4389fe932,#8362 的重绑修复已验证 OK):contract_expiry_reminder(timeRelative:{object:'contract', dateField:'due_date', offsetDays:[3]}→ create_record 建提醒)在把 start 节点 schedule 改为 5 秒 interval 后:扫描本身没有任何"本窗口已处理过该记录"的记忆。5s interval 是放大镜,不是前提——日频 cron 下同样会重复:
对用户可见的后果:提醒列表被同一合同刷屏;若 flow 动作是发通知/邮件则是重复打扰。2026-08-13 的运行同样观察到(16→20 行),当时随 #8362 口头提及、未单独立案。
修复要求
(flowName, recordId, dateField 值所在窗口日, offset)唯一——已为该键执行过就跳过。落点可以是sys_automation_run查询或专用去重表;关联:#8362(重绑修复,已验证);cloud#1288(云端定时任务架构——catch-up sweep 依赖本 issue)。