Repository navigation
DbJobAdapter 把「handler 没抛错」记成 sys_job_run.status='success' —— 内部自行降级的 job(如 wait 唤醒打空)在作业审计面上仍显示成功 #5548
Description
Activity
发现分诊轮判级(services 车道 PM,2026-08-05):持有(留
finding)。为什么还留着:审计行失真真实,但 #5549 落地后卡住的 run 已有 error 日志 + 保留的active行两条可见通道,缺的只是sys_job_run那一格的准确性 —— 无用户今天撞到的损害面。wrap()判定口径(抛错 vs 结果对象)与 job 域的结果语义值得在下一个 job 域改动时顺带定形。下轮分诊轮复核。
Generated by Claude Code
发现分诊轮:维持持有**(
finding留)+ 把重启条件硬化(应 2026-08-05 那条判级里「下轮分诊轮复核」的明示要求)。**⚠️ 上一条判级(08-05 18:08Z,services 车道 PM)写着「下轮分诊轮复核」,此后 ~26 小时无复核。本轮补上。过时前提检查(
origin/main9e3709a)—— 前提完全成立,且 job 域自 08-05 起未动wrap()逐行实读:try { await handler(ctx); … finishRun(runId,'success') … } catch { … 'failed' … throw }一字未改 —— 仍然只按「是否抛错」二分;packages/spec/src/contracts/job-service.ts:50的JobHandler = (context) => Promise<void>仍无第三态 ⇒ 方向 B 的契约缺口原样;git log origin/main -- packages/services/service-job/src/db-job-adapter.ts与-- packages/spec/src/contracts/job-service.ts:自本单立单以来无任何 job 域改动触及这两处。
⇒ 上一条的重启条件(「下一个 job 域改动时顺带定形」)尚未触发,不是被忽略。
为什么仍是持有(维持上一条的判据,复核后同意):#5549 落地后,卡住的 run 已有 error 日志 + 刻意保留的
active行两条可见通道,缺的只是sys_job_run那一格的准确性 ⇒ 今天没有用户因此看不见故障,只是审计行不够精确。⛔ 重启条件硬化(上一条只有「下轮复核」这一软条件 —— 它不是可机械判定的退出口,而「A hold without a restart condition is a state nobody can ever legally exit」)。以下三条任一成立即改判:
- 任何 PR 改到
db-job-adapter.ts的wrap()/bumpJob()或job-service.ts的JobHandler⇒ 顺带按方向 C 收口(把「sys_job_run.status的语义 = handler 未抛错」写进契约注释),成本≈一段注释,可作该 PR 的顺带项;⛔ 不为方向 C 单独发起一次派发。 - 出现第二个「内部自行降级、不抛错」的 handler(今天只有 wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529 的 wait 唤醒这一个)⇒ 说明这是形状而非孤例,晋级
pm:queue,按方向 C 立即收口并重估 B。 - 有人依据
sys_job_run.status='success'做出过错误运维判断(Studio 作业界面误导的真实现场)⇒ 立即晋级,按方向 B 处置。
方向分流预判(供届时省一轮,不构成裁决):
- A(让这类 handler 抛错)—— wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529 刻意没这么做,理由是会改变 handler 与
IJobService的失败语义、第三方实现可能据此重试;⛔ 除非有新论据,不重开这条。 - B(给
JobHandler第三态回报)—— 动packages/spec/src/contracts/job-service.ts的公开契约面 ⇒ 届时按跨座位转移协议转domain:spec座位并升needs-user-decision,⛔ 不在 services 车道内自行拍板。 - C(认定语义即「未抛错」并写进契约注释)—— 最便宜,且它消掉的是「下一个读者重新推一遍」的成本(本单立单人自陈的价值)。
域:维持
domain:services(落点packages/services/service-job);⚠️ 若走方向 B,落点移到packages/spec⇒ 那一刻改域并转座位,不在今天预改。查重:三仓 open issue / PR 各搜一遍(
sys_job_run/DbJobAdapter/ job 审计 status)—— #5529 / PR #5549 是触发场景(已收官),无重复单,objectui / cloud 无影子。本评论来自分诊座位 Routine(#5474 试点),不构成认领。
Generated by Claude Code
Findings triage round (#4949 discipline): HOLD maintained (
findingkept), domain staysdomain:services. Stale-premise check @origin/main80f7dc6— the hardened restart conditions from 08-06 21:01Z are mechanically checkable, and all three read negative this round.Condition 1 — not fired (measured, not assumed)
REST commit list against
mainsince 2026-08-06T13:00Z:packages/services/service-job/src/db-job-adapter.ts→ 0 commitspackages/spec/src/contracts/job-service.ts→ 0 commits
And the code read straight off
origin/mainis verbatim what the body describes —wrap()still resolves the run by whether the handler threw:try { await handler(ctx); … finishRun(runId, 'success', …); await this.bumpJob(name, 'success'); } catch { … finishRun(runId, 'failed', msg, …); await this.bumpJob(name, 'failed', msg); throw err; }⇒ nobody has been "in there" to take direction C as a wrap-up item.
⚠️ Method note: the local checkout is shallow (50 commits), so its per-pathgit logattributes every file to the graft boundary — these are REST readings againstmain(Operational notes 6, fourth bullet). The file contents above are read fromorigin/maindirectly, which is reliable.Conditions 2 and 3 — no signal
- Still exactly one "handles its own failure, does not throw" handler on record (wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529's wait wake-up) ⇒ no evidence yet that this is a shape rather than a single instance.
- No report of anyone making a wrong operational call from
sys_job_run.status='success'(a real misleading-Studio-jobs-view incident).
Verdict: HOLD
Reason unchanged: after #5549, a stuck run is visible through two channels — the error log and the deliberately retained
activerow — so no user misses a failure today; what is missing is the accuracy of one cell in the audit row. Direction routing also unchanged: A stays closed absent new argument (#5529 deliberately declined it), B touchespackages/spec/src/contracts/job-service.ts⇒ at that moment it transfers to thedomain:specseat asneeds-user-decision, C is the cheap wrap-up whenever condition 1 fires.本评论来自分诊座位 Routine(#5474 试点),不构成认领。
Generated by Claude Code
os-project-manager commented
on Aug 8, 2026 CollaboratorMore actionsFindings sweep (maintainer-authorized one-off, 2026-08-07 — registered on #6015): promoted to the queue. The job audit surface records
successfor handlers that internally downgraded — an audit face that cannot distinguish "did the work" from "didn't throw" is lying to operators, and the wait-wake handler is a live producer of exactly that shape. Declared = enforced applies to audit truthfulness.finding→pm:queue.
Generated by Claude Code
os-project-manager commented
on Aug 8, 2026 CollaboratorMore actionsEscalation (services seat, session
session_01USNUyHEr7uaU6MoEWXitei): this card was promoted to the queue by the 2026-08-08 findings sweep, but the direction choice it needs is a public-contract-shape decision that the triage seat itself previously reserved for the maintainer ("B ⇒ transfer to thedomain:specseat asneeds-user-decision, do not decide inside the services lane" — 2026-08-06 triage comment above). Dispatching a dev before that ruling would only reproduce this analysis as aneeds_decisionreport. Labelingneeds-user-decisionand dropping it from the active queue.Premise refreshed before escalating (2026-08-08):
packages/services/service-job/src/db-job-adapter.ts— last commit 2026-07-27 (#3529).wrap()is verbatim as described: run status is resolved solely by whether the handler threw. Unchanged since filing and since every triage check.The concrete question
When a job handler completes without throwing but did not accomplish its work (today's live producer: the #5529 wait-wake handler, which reports
STORE_UNAVAILABLEviaAutomationResult.codeand deliberately does not throw), what should the job audit surface (sys_job_run.status/sys_job.last_status) record?Options
- B (minimal shape) — give handlers an explicit outcome-reporting channel without changing
JobHandler's signature: an optional method on the job context (e.g.reportOutcome('degraded', reason)) inpackages/spec/src/contracts/job-service.ts;DbJobAdaptermaps a reported degraded outcome to a distinctsys_job_run.statusvalue; the wait-wake handler adopts it. Backward compatible: handlers that never call it keep today's semantics exactly. Requires: spec half (contract + status vocabulary) at thedomain:specseat; services half (adapter mapping + wait-wake adoption) here, blocked-by the spec half. - C — declare the semantics to be "handler did not throw", write that into the contract's JSDoc, change no behavior. Cheapest; it contradicts the promotion's own rationale (an audit face that cannot distinguish "did the work" from "didn't throw"), and the text change lands in spec (surface lane), not here.
- A (make such handlers throw) — deliberately rejected in wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529 (changes failure semantics; third-party
IJobServiceimplementations may retry). Not reopened; no new argument since.
Three-axis analysis
- 实际业务需求: exactly one producer of the misleading shape exists today, and after fix(automation): keep a wait timer's one-shot armed when the shot never consumed the pause (#5529) #5549 the underlying failure is visible through two channels (error log + retained
activerow) — the harm is bounded to one inaccurate audit cell. But the read face is real: Studio's jobs view is the operator's first stop, and flow wait/resume is v17-adjacent infrastructure. Pull for B is emerging rather than zero; pull for C is "stop the next reader from re-deriving this". - 项目长远合理性: the sweep's promotion rationale — audit surfaces must not lie — is a North-Star statement. B is the contract-first shape for it. C permanently defines the audit column as "exception telemetry", which is honest but forecloses the column ever answering "did the work" without a second migration.
- 防 AI 写代码犯错: C is the tolerance-hides-errors shape: every future AI-written handler that swallows its own failure will show green in audit, and nothing structural pushes back. B makes the degraded path representable and loud, and the adapter maps it mechanically — declared = enforced extended to the audit face.
Recommendation
B (minimal shape), split contract-first: spec sub-issue at the
domain:specseat (context method + status vocabulary), services sub-issue here (Blocked-bythe spec half). If current appetite says not now, the honest alternative ispm:on-holdwith restart condition "second non-throwing degraded handler appears, or a real operator-misread incident" — not C: C spends the one-time chance to define the column correctly on documenting the defect. If the maintainer rules C regardless, the JSDoc change routes to the spec-surface lane per the existing seat boundaries.
Generated by Claude Code
- B (minimal shape) — give handlers an explicit outcome-reporting channel without changing
os-project-manager commented
on Aug 8, 2026 CollaboratorMore actionsMaintainer ruling — 2026-08-08. The maintainer reviewed the PM's three-axis analysis of the decision inbox and accepted the recommendations (「按照你的建议继续」). Recorded by the PM session;
needs-user-decisioncomes off with this comment.Decision: Option B — give
JobHandlera way to report "ran, but did not accomplish the work", and map it to asys_job_run.statusdistinct fromsuccess.Rationale across the three axes:
- A is correctly excluded (as wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529 already judged): making these handlers throw changes the failure semantics of
IJobServicefor third-party implementations that may retry on throw. That is a bigger contract change than B, in the riskier direction. - C is rejected: documenting
statusas "the handler did not throw" legitimises an audit surface that cannot answer the question it exists to answer. Studio's job view reads these rows; an audit face that reports success for a run that accomplished nothing is the declared ≠ enforced shape applied to observability. - B is additive:
JobHandlerkeeps returningPromise<void>as a legal outcome, so no existing handler or third-party implementation breaks; only handlers that want to report the third state do so.
Constraints:
- Vocabulary stays minimal — one additional outcome meaning "completed without accomplishing the work". ⛔ Do not open an enum family; a second key would need its own pull.
packages/spec/src/contracts/job-service.tsis the contract change (shared contract surface ⇒ spec seat owns that half);DbJobAdapter.wrap()is the consuming half. Contract-first: triage may split with aBlocked-by:line.- The motivating case is the acceptance test: wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529's wait-wake handler returning
STORE_UNAVAILABLEmust land asys_job_runrow that is visibly notsuccess, while the run staysactive(that part of wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529's fix is not reopened). - Refusal/observability assertions elsewhere stay verbatim; this widens the reportable outcomes, it does not relax any existing one.
Note on routing: this card was promoted to the queue by the findings sweep on 2026-08-07 and then correctly escalated back by the services lane because option B touches a public contract. That round trip is the protocol working — the sweep graded it dispatchable on the defect, the lane caught that the chosen fix crosses a contract line.
Generated by Claude Code
- A is correctly excluded (as wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529 already judged): making these handlers throw changes the failure semantics of
11 remaining items
os-project-manager commented
on Aug 9, 2026 CollaboratorMore actionsRound-18 dispatch outcome: a good BLOCKED, accepted as such (
domain:servicesseat, sessionsession_01USNUyHEr7uaU6MoEWXitei). No PR is the correct deliverable here — the dev hit the dispatch's own cross-lane stop condition, measured both failure directions instead of forcing either, filed #7072 with the full evidence, and tore down cleanly (zero files pushed).What the stop condition caught, in one line: #6617 shipped the producer side of the ruling; the consumer-side status vocabulary was never anyone's card. Three schema declarations pin
sys_job_run.statusto four values —JobExecutionStatusinpackages/spec(spec lane, hard single-owner) and the twoField.selects inpackages/platform-objects(metadata lane). The in-lane change both fails to compile (TS2345) and, if typed around, is refused at write time byrecord-validator(invalid_option) — with the refusal swallowed intologger.warnbyfinishRun's best-effort catch, leaving the run row onrunningforever. Forcing it would have shipped an audit face strictly worse than the one this card exists to fix.State moved accordingly: assignee cleared,
pm:dispatched→pm:blocked, body'sBlocked-by:updated#6617→#7072. When #7072 lands, this card returns to the queue as a small adapter wiring change plus the pinned matrix (throw ⇒failed+ retry;undefined/{outcome:'completed'}⇒success;{outcome:'degraded'}⇒ the new status, no retry;bumpJobin sync; #5529 wait-wake scenario end-to-end with the run stayingactive).Seat decision on the dev's open question 2 — where
reasonlands (veto window open, this is not an escalation): option A —failure_countstays flat on a degraded run, andJobRunOutcome.reasonis written to the existingsys_job_run.error/sys_job.last_errorcolumns. Grounds: the ruling's own "vocabulary stays minimal — do not open an enum family" argues identically against opening a column family, and option C (log-only) re-creates the exact silence the original ruling rejected. The honest cost travels with it: a column labelled "Error" will carry a non-error note wheneverstatus = 'degraded', so the label reads correctly only alongsidestatus— the #7072 implementer should record that in the TSDoc at site 1, and if the maintainer finds the overload unacceptable, a dedicatedreasoncolumn is the clean alternative decided on #7072, not discovered mid-implementation here.Routing of #7072 itself (one card vs split, and which seat) is with the triage seat — recommendation posted there.
Generated by Claude Code
UNLOCK — the blocking vocabulary card landed. Spec-lane PM seat (#6017), session
session_01PiRUoQkTSBBmpyXBY3cVn2.#7072 closed 2026-08-10T08:13Z via PR #7340 (merged, two-read confirmed). All three declaration sites now carry
degraded, moved together in one commit:JobExecutionStatus(packages/spec/src/system/job.zod.ts),sys_job_run.status, andsys_job.last_status— so this card'sDbJobAdapter/bumpJobwiring now compiles against the widened type AND its writes pass ObjectQL's enforced select validation (the silent-swallow trap #7072 measured is closed).Two decisions recorded on the spec side that this card's implementation should honor (both in the enum's TSDoc, per the 2026-08-08 ruling's minimal-vocabulary spirit): a degraded run's
reasonrides the existingerror/last_errorcolumns (no new column), andfailure_countstays flat —degradedis not a failure and never retries. Scheduling stays with thedomain:servicesseat; this is the unlock notice only. The objectui Studio jobs-view follow-up (fifth value needs label/colour/filter) is now fileable in that repo.
Generated by Claude Code
Unblocked —
Blocked-by: #7072is cleared, and the premise was re-verified at the merged ref rather than inferred from the close.#7072 landed as PR #7340. Checked against
origin/main, not a working tree:packages/spec/src/contracts/job-service.ts:69—JobRunOutcomeis{ outcome: 'completed' | 'degraded'; reason?: … }, and:85states plainly thatdegradedis not a failure and does not trigger a retry.packages/spec/src/system/job.test.ts:405—JobExecutionStatus.parse('degraded')no longer throws, so the consumer vocabulary that produced the TS2345 and theinvalid_optionwrite refusal is widened.packages/platform-objects/.../translations/*.objects.generated.ts— twodegradedoption labels, i.e. bothField.selectsites the second blockage named.packages/services/service-job/src/db-job-adapter.ts:128 / :167— stillfinishRun(runId, 'success')on the no-throw path. The work this card describes is genuinely still undone, so the card survives its own unblocking.
The contract file names the remaining step itself (
job-service.ts:108): mappingdegradedonto asys_job_run.statusdistinct fromsuccess. That is this card, and it is now a small adapter wiring change rather than the contract negotiation it was in round 18.
Claim: PM loop, services lane
Session:session_015fkdTyGmMD5s8ZtEifvuGy
Branch:claude/issue-5548-job-degraded-status
Worktree:objectstack-issue-5548
Domain:domain:services
File surface:packages/services/service-job/src/(db-job-adapter.ts+ its tests). ⛔ Notpackages/spec(the contract half landed as #6617/#7072 and is the spec seat's), ⛔ notpackages/platform-objects(the vocabulary half landed with #7340).
Container & model: S-grade wiring against a contract that already prescribes the shape,mode:subagent,model: opus— sized S because the vocabulary and the outcome type both already exist, but tiered up rather than down because the card carries two real judgement points (whetherbumpJobmirrors the run status ontosys_job.last_status, and whether adegradedrun counts toward the failure signal), and the tiering rule says borderline goes up.
Serial constraints cleared: no other in-flight claim or open PR touchespackages/services/service-job— the lane's previous card (#7385 / PR #7418,service-datasource) merged at 11:02Z and is a different package;#7072's own PR #7340 is merged, so there is no in-flight producer on the vocabulary either.
Generated by Claude Code
<!-- os-dev-report -->
{ "issue": 5548, "status": "done", "branch": "claude/issue-5548-job-degraded-status", "pr": "https://github.com/objectstack-ai/objectstack/pull/7446", "premise_still_valid": true, "summary": "Premise re-verified at origin/main 2ef18070e: DbJobAdapter.wrap() still resolved a run purely by whether the handler threw (finishRun(runId,'success') on the no-throw path), while JobRunOutcome (#6617) and the widened JobExecutionStatus / two Field.select sites (#7072/#7340) were already in place. Wired the services half: wrap() now reads the handler's resolved value and maps { outcome:'degraded', reason? } to sys_job_run.status='degraded' with the reason in `error`; bumpJob mirrors it onto sys_job.last_status/last_error while leaving failure_count flat and retry untouched; IntervalJobAdapter and CronJobAdapter map the same third state onto the in-memory JobExecution so getExecutions() cannot disagree with the persisted row; replay()'s synthetic row is settled from the execution just recorded instead of hard-coding 'success'. The ruling's named acceptance case adopts the channel: the wait node's timer wake-up resolves { outcome:'degraded', reason:'STORE_UNAVAILABLE' } on the store-unreachable path while still not throwing and still keeping its one-shot ARMED (#5529 untouched). Judgement calls, both explained in the PR body: `reason` rides the existing error/last_error columns (the #7072 seat decision recorded in the JobExecutionStatus TSDoc — deviating would contradict a shipped contract comment) and bumpJob DOES mirror degraded onto sys_job.last_status (leaving it 'success' would just move the defect one table over) while failure_count deliberately does not move. Scope note for the PM: packages/spec and packages/platform-objects were not touched, and nothing was found missing in either; the only edit outside service-job is the same-lane service-automation wake-up adoption the ruling itself prescribes ('the wait-wake handler adopts it'), plus a @objectstack/service-job devDependency so the end-to-end regression drives the real adapter rather than a spy.", "tests": "Build closure first (fresh worktree): pnpm --workspace-concurrency=2 --filter '@objectstack/service-job^...' --filter '@objectstack/service-automation^...' build — the ^... SUFFIX form, i.e. UPSTREAM dependencies, run before any typecheck. service-job: `pnpm --filter '@objectstack/service-job' test` => 'Test Files 7 passed (7) / Tests 56 passed (56)' (13 new in db-job-adapter.degraded-outcome.test.ts). service-automation: `pnpm --filter '@objectstack/service-automation' test` => 'Test Files 73 passed (73) / Tests 890 passed (890)' (3 new in wait-node-degraded-run.test.ts). Typecheck: `pnpm --filter '@objectstack/service-job' typecheck` => 'Done' (clean); service-automation has no typecheck script, and a manual `tsc --noEmit -p tsconfig.json` there reports only three PRE-EXISTING TS2341 errors in nested-region-parity.test.ts (file untouched by this PR, same on origin/main). REVERSE VERIFICATION, each half separately, taking the file out with `git checkout origin/main -- PATH` (never git stash) and restoring from the commit: (1) db-job-adapter.ts removed => 7 of the 13 new cases red, each naming the cell, e.g. \"AssertionError: expected 'success' to be 'degraded'\" and \"expected [ 'success', 'success' ] to deeply equal [ 'degraded', 'degraded' ]\" for the replay case; the 6 that stay green are exactly the additivity/legacy cases plus the two timer-adapter cases (different file, still fixed) — the predicted direction, recorded rather than assumed. (2) wait-node.ts removed => 2 of the 3 end-to-end cases red (\"expected 'success' to be 'degraded'\") with the success control still green, which is what proves the e2e pin exercises the producer adoption and not only the adapter. Additivity has its own reverse evidence: a handler resolving `undefined` is asserted to still land status 'success', error null, failure_count 0. Every status assertion reads the value written into the sys_job_run / sys_job cell — never 'the handler was called' or 'it did not throw', since that criterion IS the defect. Gates derived from the changed paths via `node scripts/pm/dispatch-gates.mjs`: check:nul-bytes ('OK, scanned 6830 text file(s)'), check:docs-audit-scope (self-tests 56 + 22 cases, scope in sync), check:changeset-gate-self-tests (118 + 142 + 117 assertions), check-changeset-fixed ('fixed group is in sync with 69 public workspace packages'), check-changeset-no-major ('introduces no major bump') — all green; ESLint on the six changed files: no output. Changeset added (.changeset/job-run-degraded-status.md, patch for both packages) so no skip-changeset label applies. CI at report time (reported honestly per the dispatch): ESLint, TypeScript Type Check, Check Changeset, Validate Package Dependencies, Test Core (1-3/3), Build Core, Build Docs, Dogfood gates all `in_progress`; completed and green so far: Auto Label, Check PR Size, Check Documentation Links, Console Pin Freshness, ADR maintainer approval, 'No other open PR may claim the same issue'. No red gate observed. PR body read back after creation: the sanitizer had eaten one `<`+letter token in the reverse-verification line; corrected and the session-URL footer survived the edit.", "open_questions": [], "out_of_scope_findings": [] }
Generated by Claude Code
Generated by Claude Code
ACCEPT — PR #7446, reviewed against GitHub rather than the report's own claims.
What shipped. The services half of the 2026-08-08 B-minimal ruling:
DbJobAdapter.wrap()now reads the handler's resolvedJobRunOutcomeand maps{ outcome: 'degraded', reason? }tosys_job_run.status: 'degraded'(reason inerror);bumpJobmirrors it ontosys_job.last_status/last_errorwhilefailure_countstays flat and retry stays keyed on a rejected promise only;IntervalJobAdapter/CronJobAdaptermap the same third state sogetExecutions()cannot disagree with the persisted row;replay()'s synthetic row is settled from the execution the inner adapter just recorded. The wait node's timer wake-up — the ruling's named acceptance case — resolves{ outcome: 'degraded', reason: 'STORE_UNAVAILABLE' }on the store-unreachable path while still not throwing and keeping the one-shot armed (#5529 untouched).What I verified myself (
get_files, 9 files):- Path surface: changeset + 6 source/test files +
package.json/lockfile. Nodocs/adr/**(no ACCEPT fork), nocontent/docs/releases/, and — the dispatch's hard boundary —packages/specandpackages/platform-objectsuntouched, confirmed in the diff, not from the report. - Every status assertion reads the written cell.
runRows()[0].status/jobRow().last_statusetc. — never "the handler was called" or "it did not throw", which is the criterion that was the defect. The dispatch made this the acceptance bar; the diff meets it literally. - The wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529 regression is end-to-end and honest. It parks a real run, cold-boots a second engine whose store
loadthrows, re-arms onto a realDbJobAdapter, fires the shot, and asserts: the run is still parked,sys_job_run.status === 'degraded',error === 'STORE_UNAVAILABLE',sys_job.active === true(the wait 定时唤醒 job 在 resume 没能消费掉暂停时也会自我取消 —— store 短暂不可达即丢掉这一次唤醒,run 挂到下次重启才被捞回 #5529 half),failure_count === 0. Plus a control where the store works and the row sayssuccesswith the one-shot disarmed. - Additivity has reverse evidence: a handler resolving
undefinedstill landssuccess/error: null/failure_count: 0;{ outcome: 'completed' }likewise; a throwing handler still landsfailedwithfailure_count+1. And adegradedrun under amaxRetries: 2policy runs exactly once. - Reverse verification ran in the right direction: adapter reverted → 7/13 new cases red naming the cell (
expected 'success' to be 'degraded'); wait-node reverted → 2/3 e2e cases red with the control green — which is what proves the e2e pin exercises the producer adoption, not just the adapter. - Both judgement calls the dispatch left open are decided with recorded grounds:
reasonrides the existingerror/last_errorcolumns (the Job status vocabulary has nodegradedvalue — #5548's services half cannot compile or persist without wideningJobExecutionStatus+ the twosys_job*selects #7072 seat decision recorded in theJobExecutionStatusTSDoc — deviating would contradict a shipped contract comment), andbumpJobmirrorsdegradedontosys_job.last_statuswhile the failure signal stays flat. Both match the contract's "degradedis NOT a failure" clause verbatim.
One out-of-surface extension, accepted under the three-criteria rule. The claim declared
packages/services/service-job/only; the diff also touchespackages/services/service-automation(wait-node adoption + a@objectstack/service-jobdevDependency + the lockfile row). (1) loudly declared in both the report and the PR body; (2) mandated by the dispatch itself — the required #5529 end-to-end regression cannot exist without the producer adopting the channel; (3) no atomicity-preserving split — landing the adapter without the named acceptance case would ship the mechanism with zero producers and the regression unprovable. Same lane, same package family, no cross-seat declaration owed.CI at review time: Check Changeset / Validate Package Dependencies / Check PR Size / Check Documentation Links / ADR maintainer approval / Auto Label all
success; ESLint and TypeScript Type Checkin_progress— per the L2 contract that is the expected state at report time, and the ready-flip waits for their individualconclusion: successreadings. Driving to landing next: gate conclusions → ready → auto-merge → merge queue, with the queue-branch read as the landing check.
Generated by Claude Code
- Path surface: changeset + 6 source/test files +
- added a commit that references this issue
on Aug 10, 2026 - added 3 commits that reference this issue
on Aug 17, 2026
发现于 #5529 的实测过程(PR 见该 issue),与该修复相邻但不同面 —— 这是 job 域的审计面问题,不是 automation 域的。
观察(实测,非推断)
DbJobAdapter.wrap()(packages/services/service-job/src/db-job-adapter.tsL161-176)只按 handler 是否抛错判定运行结果:对于「内部自行处理失败、不抛错」的 handler,这条路径把它记成成功。实测(临时 harness,已删除):一个 once job 的 handler 正常返回但什么也没完成,
sys_job行是sys_job_run只有一行status: 'success'。为什么可能值得看
sys_job/sys_job_run是机器可读面(Studio 的作业界面读它)。一种读法是sys_job_run.status本就只表示「handler 跑完没抛异常」,那它没说谎;另一种读法是它作为作业审计面在这里给不出「这一枪没干成事」的信号。两种读法都成立,所以这里只做记录、不预判严重度。需要说明的是 #5529 之后这个洞不是无声的:那条路径已经有一条 error 日志,并且 job 被刻意保留
active(#5529 就是这么修的),所以「run 卡住」本身是可见的 —— 缺的只是作业审计行里的那一格。因此按观察类归档(finding,不进pm:queue),严重度交分诊判定。可能的方向(未决,不在本 issue 范围内实施)
IJobService的失败语义(第三方实现可能据此重试),超出该 issue 派发范围。IJobService的 handler 一个「跑完了但没干成」的回报方式(而不是只有 throw / 不 throw 两态),由适配器映射到一个区别于success的sys_job_run.status。这会动IJobService契约面,属公开契约决定。sys_job_run.status的语义就是「handler 未抛错」,并把这一点写进契约注释,免得下一个读者重新推一遍。佐证
packages/services/service-job/src/db-job-adapter.tsL161-176(wrap),L277-299(bumpJob)。packages/spec/src/contracts/job-service.ts——JobHandler返回Promise,没有第三态。STORE_UNAVAILABLE时不取消 + 记 error,是本条观察的具体触发场景。2026-08-08 裁决落地(services 座位 PM,会话
session_01USNUyHEr7uaU6MoEWXitei):维护者批复三轴分析,采纳 B-minimal(可加性第三态;A 与 #5529 裁决冲突被否,C 把审计静默合法化被否)。按 shared-contract 规则 contract-first 拆分:spec 半边见 #6617(JobHandler 可选 degraded-outcome 回报通道,API 形状归 spec 座位);本单收窄为 services 半边(DbJobAdapter 把回报映射为区别于success的sys_job_run.status+bumpJob同步 + #5529 场景回归测试)。2026-08-09 二次阻塞(round 18 dev 实测,详见 #7072 与本线程 13:4xZ 评论):#6617 只交付了生产者侧(
JobRunOutcome+JobHandler返回值);sys_job_run.status的消费侧词表被三处 schema 钉死(spec 的JobExecutionStatusenum + platform-objects 的两个Field.select),全不在 services 车道 —— 适配器改动编译不过(TS2345),绕过类型则被record-validator以invalid_option拒写且被finishRun的 try/catch 吞掉。词表拓宽 = #7072;它落地后本单退化为一次小的适配器接线。Blocked-by: #7072