Repository navigation
A licensed max_nodes cap still cannot be enforced by any replica: the cluster has no membership view and no slot claim, so a count-carrying gate verdict stays advisory #8501
Description
Activity
Lane evidence for the grading round —
domain:servicesseat #6021, sessionsession_01ARidKDYSCD56LaygrvDPnk. ⛔ Not a claim, no label changed, not routed by me.Flagging this one for grading attention above the usual findings queue, for a reason that is not visible from the card alone:
⭐ This card is the only remaining work between a maintainer ruling and its enforcement. The maintainer ruled 2026-08-13 (on
objectstack-ai/cloud#1275) that a licensedmax_nodesoverflow refuses the excess replicas and runs up to the paid limit. #8367 — which I dispatched and whose PR #8503 is now in review — was believed by both its filer and triage to be the change that made that ruling enforceable.⚠️ That premise was falsified by measurement during implementation, and this card is where the falsification landed.So the state after PR #8503 merges is: the ruled verdict is expressible at the seam, documented as advisory, and not enforced. A cap of 3 with 5 replicas still starts 5. This card is the whole remainder.
Two notes for whoever grades it:
- ⛔ The decision half is already ruled — the refuse-excess-not-whole-cluster semantics come from the 2026-08-13 ruling and are not re-openable. But this card does carry genuinely new decisions of its own that the ruling does not answer: TTL semantics for a crashed replica's leaked slot, and what a replica does when the driver is reachable but its claim is refused (downgrade to in-memory vs refuse to serve). The existing seam's posture — "multi-node is an add-on, never brick" — points at downgrade, but that is an inference, not a ruling.
- The mechanism is not a signature change: it is a new runtime subsystem over the
ILock/ICounter/IKVprimitives this package already ships, plus an interaction with the split-brain guard. Size and model should be graded accordingly — this is not the S–M the parent card was.
Blocked-by: #8367is correct and will clear when PR #8503 lands.
Generated by Claude Code
Triage (first-touch grade): held as
finding, routeddomain:services(fix, if built, is the slot-claim mechanism overservice-cluster's ownILock/ICounter/IKVprimitives). Reasons for holding rather than promoting or escalating now: ① Blocked-by #8367 is still in flight — the seam this would consume isn't merged; ② the enforcement mechanism is a new runtime capability with its own design decisions (TTL semantics, refusal posture, split-brain interaction) — feature-class, human floor — but escalating it before #8367 lands would put a card in the maintainer's inbox that cannot be acted on yet. Re-grade when #8367 closes: expected path is aneeds-user-decisioncard asking appetite for the slot-claim mechanism, carrying the cloud#1275 ruling as its anchor (the ruled "refuse the excess replicas" outcome is unenforceable without it — that fact is this card's whole content, and it is measured, not speculative).本评论来自分诊座位 Routine。
Generated by Claude Code
Blocker cleared —
domain:servicesseat #6021. PR #8503 merged 2026-08-13T16:49:09Z; #8367 is closed. TheBlocked-by: #8367on this card no longer holds: the seam now carriesadmitted/refused/capped, so this card's mechanism has the vocabulary it needs to consume.⛔ Still unassigned, still
finding, still ungraded —⚠️ and now this card plus #8504 are the entire remaining distance between the 2026-08-13 ruling and its enforcement. #8367's closing comment says so in terms, so a reader arriving from the closed card lands here.
Generated by Claude Code
Escalated to
needs-user-decision(restart condition met: #8367 closed — PR #8503 merged 2026-08-13T16:49Z, andmulti-node-gate.tsonorigin/maincarries theadmitted/refused/cappedvocabulary; the body'sBlocked-by: #8367line is therefore exhausted — recorded here rather than edited out). This is the recorded expected path from the 2026-08-13 first-touch grade: the enforcement mechanism is feature-class, on the human floor.Already ruled (⛔ not re-openable): maintainer 2026-08-13 on objectstack-ai/cloud#1275 — a licensed
max_nodesoverflow refuses the excess replicas, runs up to the paid limit, and warns loudly; whole-cluster degrade is rejected. This card is (with #8504) the entire remaining distance between that ruling and reality: a cap of 3 with 5 replicas still starts 5, because every replica computes the same advisory verdict and none can tell whether it is the excess.The decision asked here: appetite for the atomic slot-claim mechanism (per-replica claim over the existing
ILock/ICounter/IKVprimitives, TTL expiry for a crashed replica's leaked seat, split-brain-guard interaction) — a new runtime subsystem, M/L-class.Four-prism block:
- Platform long-term coherence — closes a ruled-but-unenforced gap using primitives the package already ships; declared = enforced, no new dependency. The alternative (advisory forever) leaves a ruling permanently unenforced.
- Measured business pull — zero measured deployments overflowing a licensed cap today; the pull is the licensing programme itself (cloud#1275), not a bleeding user. Enforcement starts mattering with the first real licensed multi-node contract.
- AI-agent error-resistance — an enforced cap beats an advisory verdict an agent can misread as binding, and loud refusal beats silent overflow; but a half-built claim mechanism (wrong TTL, wrong refusal posture) can brick clusters, against the seam's own "multi-node is an add-on, never brick" posture.
- Startup scope discipline — a new distributed subsystem is a permanent maintenance obligation (TTL semantics, crash recovery, split-brain interplay). The cheap half — The 2026-08-13 max_nodes ruling's "loud warning" clause has no owner:
os servecalls the multi-node gate zero-arg, so a partial cap is silent even once the gate can express it #8504's loud-warning wiring atos serve— delivers visibility without the subsystem.
Recommendation: land #8504 (S–M) now; park this mechanism card with the named restart condition "first licensed multi-node deployment where the cap must bind, or the maintainer pulls it earlier". Prisms 2 and 4 dominate today, and prism 1's gap stays visible through the loud warning rather than silently. If you want full fidelity to the ruling now instead, this dispatches as M/L and needs the two sub-decisions below answered.
Sub-decisions the ruling does not answer (only if built now):
- Refused replica's posture: downgrade to single-node/in-memory (the "never brick" posture points here) vs refuse to serve.
- Slot TTL semantics for a crashed replica's leaked seat.
Premise re-check lines for a later reader:
git grep -n "admitted\|refused\|capped" origin/main -- packages/services/service-cluster/src/multi-node-gate.ts·git log --oneline -5 -- packages/services/service-cluster/src· state of #8504.本评论来自分诊座位 Routine。
Generated by Claude Code
Maintainer ruling (2026-08-14, verbatim: 「同意你的建议」, approving item 2 of the eight-item decision-box list: 「#8501:落 #8504、槽位机制挂起?」): the slot-claim mechanism is parked —
needs-user-decision→pm:on-hold. Restart condition (any one): the first real licensed multi-node deployment where the cap must bind, or the maintainer pulls it earlier. The deliverable-now half of the 2026-08-13 ruling lands via #8504 (loud advisory warning), which enters thetarget:v17workable set in the same stroke. Sub-decisions (refusal posture, slot TTL semantics) stay open until the restart fires — do not pre-decide them.
Generated by Claude Code
os-support-ai commented
on Aug 26, 2026 CollaboratorMore actions半状态巡查处置(H9)——hold 的 blocker 已经关了 13 天
domain:services座位,sessionsession_0157mMVAq9fjGe2kaSD2aJC8,轮次 1。巡查锚 #9857(02:12:33Z)对本卡报了 H9(
pm:on-hold无可点火的Restart-when:)。查下来比 H9 报的更糟一层:⚠️ 本卡正文的Blocked-by: #8367已经失效。 #8367 于 2026-08-13T16:49:09Z 关闭,由 PR #8503(feat(service-cluster): carry an admitted node count through the multi-node gate)合并关单。也就是说本卡等的那个 seam 加宽十三天前就落地了,而卡一直停在一个既没有Restart-when:也没有活 blocker 的 hold 里 —— 两个通道同时不可唤醒。本卡自己写了这一步的意义:「#8367 makes that verdict expressible … It does not make it enforceable」。所以 #8367 落地解除了阻塞但没有完成工作,剩下的就是本卡的正题:原子 slot 认领。
前提重验(带 re-check 命令)
git grep -n "allowMultiNode" origin/main -- packages/services/service-cluster/src git grep -n "generateNodeId" origin/main -- packages/services/service-cluster/src/cluster.ts git grep -n "OS_CLUSTER_REPLICAS" origin/main -- packages/services/service-cluster/src/split-brain-guard.ts git log --oneline -15 origin/main -- packages/services/service-cluster/src⚠️ 卡上的三条测量(单进程一次的 boot 期咨询、无成员视图、无序号)测于ff1e9b6a9,而 #8503 之后 seam 已经变了(现在带 admitted count)。⇒ 那三条必须重测,尤其第一条:如果 #8503 顺带改了调用点的形状,本卡的分析要跟着改。⛔ 不引用卡上写作时的描述。为什么转决策箱而不是回队
本卡 type 是 Feature,交付物是一套新的运行时机制(跨进程原子认领 + TTL 过期 + 优雅关闭释放 + 认领被拒时的姿态)。功能新增类是人工地板 —— 不代裁。派发它等于让 dev 顺手替你定"客户的第 4 个副本启动时到底发生什么",而那是产品姿态,不是实现细节。
一句话问题
客户买了"最多 3 个节点"的授权,起了 5 个。平台今天拦不住多出来的那 2 个 —— 每个副本自己算,算出来的答案都一样("允许 3 个"),但没有一个副本知道自己是不是那 3 个里的。所以要么 5 个全放进来(等于没有上限),要么 5 个全拒绝(整个集群停摆,而这是你 2026-08-13 明确否掉的做法)。
选项 × 真实代价
做什么 真实代价(客户可感知) A 建原子 slot 认领:每个副本启动时抢一个坑,抢不到的自己降级成单机并告警 一套新的集群机制(TTL、崩溃释放、关闭释放)。做错了客户会莫名其妙少几个副本 —— 比如一个副本崩了没释放坑,TTL 没到,新副本抢不到坑 B 维持现状:上限只是告警,不拦 超卖的客户照常跑满 5 个节点,只是日志里有一行警告。授权上限在技术上不存在 C 先只把告警做到运维看得见的地方(控制台/遥测),认领机制押后 同 B,但至少你能看见谁在超卖 —— 靠商务去谈,不靠代码拦 业务含义直译:A = 装闸机(有票才进站);B = 贴一张"限乘 3 人"的告示;C = 告示 + 摄像头,事后找超载的人算账。
四轴(从业务立场)
① 项目长远合理性 —— A 是唯一让"授权上限"这句话在技术上成立的选项;B/C 之下,任何按节点数计价的条款都只是君子协定。
⚠️ 但反面也真实:集群成员管理是一整块基础设施,建了就得永远维护(崩溃恢复、时钟、TTL 边界),而平台今天没有任何其它功能需要成员视图。为一个计费上限单独养一套集群机制,是本轴上最贵的那种"合理"。② 实际业务拉动 ——
⚠️ 我没有测到真实拉动:没有记录显示有客户超卖过节点数。本卡是 dev 在做别的卡(#8367)时发现的能力缺口,不是从客户事故来的。上游的裁决(cloud#1275)裁的是"如果发生,应该怎么办",不是"现在有人在这么干"。③ 防 AI 犯错 —— 中性偏 A:今天的形状会误导写部署文档的人(包括 AI)以为
max_nodes是被强制的;A 让声明与行为一致。但这一轴的分量比前两张卡轻 —— 错的后果是计费不准,不是数据泄露或静默越权。④ 创业阶段不扩散 —— 明确指向 B/C。这是典型的"能力扩张、无拉动":为一个还没发生的超卖场景,新增一套跨进程协调机制。按创业阶段聚焦原则,无拉动的能力面默认 defer,而且已经付出的 #8367 不构成豁免(沉没成本不买未来)。
推荐
C(先让超卖看得见,不建认领机制),回退 A —— 一旦出现第一个真实超卖客户,拉动就有了,那时再建,而且那时你会知道 TTL 和拒绝姿态该怎么定,因为有真实场景可依。
⚠️ 如果你的判断是"按节点计价马上要成为正式商务条款",那 ②④ 两轴的读数当场翻转,应该直接选 A —— 这正是我看不见的那部分,见置信缺口。① 项目长远合理性:A 是唯一让"授权上限"在技术上成立的选项,但平台今天没有第二个功能需要集群成员视图 —— 为一个计费上限单独养一套集群机制是本轴上最贵的"合理"。
② 实际业务拉动:未测到真实拉动 —— 无客户超卖记录;本卡来自 dev 做 #8367 时的发现,上游裁决裁的是"若发生怎么办"而非"现在有人这么干"。
③ 防 AI 犯错:中性偏 A(今天max_nodes看起来被强制其实不是,会误导写部署文档的人),但错的后果是计费不准,不是数据泄露 —— 本轴分量轻于前两张卡。
④ 创业阶段不扩散:明确指向 B/C —— 为尚未发生的超卖新增跨进程协调机制属无拉动的能力扩张,且已付出的 #8367 不构成沉没成本豁免。
推荐:C(先让超卖可见),回退 A;若按节点计价即将成为正式商务条款则②④翻转,直接 A。
置信缺口:本分析看不见按节点数计价在商务上有多硬 —— 有没有已签或在谈的合同以max_nodes计价,以及有没有客户已经超卖。那是商务事实,平台侧无遥测,而它一翻转就直接改结论。低摩擦裁决格式
回一个字母即可。「C」= 只做可见性;「A」= 建认领机制;「B」= 什么都不做,我会据此把卡关掉(not planned,理由与出处写进关单评论)。
裁后我会怎么执行(你不用管)
同笔四件:换标签、逐字抄裁决、把 #8367 已落地这件事写进卡面前提、派发令要求先对
origin/main重测那三条(seam 已经变过一次)。
Generated by Claude Code
Maintainer ruling recorded — Option C: visibility now, no slot-claim mechanism; card closed not planned with named reopen triggers
Provenance: maintainer, 2026-08-27, PM chat decision-inbox batch 4 (adjudication session
session_01DKWDdUJ2XNRESVVWUvcpnh), verbatim: 「其他接受」 — accepting this card's presented recommendation C (2026-08-26T05:29Z facet block), fallback A not taken.Ruled: oversell becomes VISIBLE (operator-facing console/telemetry surface over the
admitted/refused/cappedgate verdict) — filed as #12667, queued. The atomic slot-claim mechanism is deliberately NOT built: zero measured oversell, no second platform feature needs a cluster membership view, and a new distributed subsystem is a permanent maintenance obligation with no pull today. The 2026-08-13 overflow-semantics ruling (cloud#1275: refuse the excess, run to the paid limit, never whole-cluster degrade) is unchanged and ⛔ not re-opened — it stays the required shape IF the mechanism is ever built.Why closed rather than held: a hold is legal only with a machine-fireable restart condition, and this card's trigger is a commercial fact no scan can read (that is exactly the H9 finding that surfaced it). Closing not-planned keeps the record and makes reopening free.
Reopen triggers (any one — reopen this card and dispatch the mechanism as M/L, with the two open sub-decisions put to the maintainer: refused replica's posture — downgrade vs refuse-to-serve — and slot TTL semantics):
- Node-count pricing becomes a signed or actively-negotiated commercial term.
- The first real licensed multi-node deployment where the cap must bind.
- An operator reports harm traceable to oversell.
Standing constraints carried forward for that day: re-measure the three mechanism facts on the then-current tree (the seam already changed once at #8503); the "never brick" posture of the existing seam; interaction with the split-brain guard.
State:
needs-user-decisionremoved; closed as not planned. Visibility work: #12667.
Generated by Claude Code
Filed unassigned, observation-class from the
domain:servicesdev seat while implementing #8367 (sessionsession_01ARidKDYSCD56LaygrvDPnk). Not triaged or routed by me.Blocked-by: #8367 (the seam widening this would consume; it is necessary but not sufficient on its own)
Fact (measured on
origin/main@ff1e9b6a9)The maintainer ruled 2026-08-13 (recorded on
objectstack-ai/cloud#1275) that a licensedmax_nodesoverflow must refuse the excess replicas, run up to the paid limit, and warn loudly — explicitly not a whole-cluster degrade. #8367 makes that verdict expressible at theMultiNodeGateseam. It does not make it enforceable, and no consumer-side change can, because of three measured properties of the surrounding mechanism:packages/cli/src/commands/serve.ts:1234-1247, inside theOS_CLUSTER_DRIVERbranch ofos serve.nodeIdis generated randomly per process —generateNodeId()atpackages/services/service-cluster/src/cluster.ts:123, used viaparsed.nodeId ?? generateNodeId()atcluster.ts:60. Nothing registers a node, tracks liveness, or counts live members.OS_CLUSTER_REPLICAS(packages/services/service-cluster/src/split-brain-guard.ts:33-34,47-51) — an operator-declared desired count, identical in every replica, not a live membership count.Why that leaves the cap advisory
With a cap of 3 and 5 replicas booting, every replica calls the gate with the same input and computes the same verdict ("3 admitted"). None of them can determine whether it is one of the admitted 3 or one of the excess 2. The only outcomes reachable from that verdict locally are:
So "run N, refuse N+1" cannot be produced by any per-replica decision over a count alone.
What would make it binding
An atomic slot claim against the shared cluster primitives this package already ships (
ILock/ICounter/IKVon a remote driver): each booting replica claims a slot after connecting the driver; the claim that exceeds the cap fails, and that replica downgrades itself to single-node and warns. Needs, at minimum:Why this is filed rather than fixed
#8367's dispatch scoped the file surface to
packages/services/service-cluster/srcplus tests and a changeset, and the shape above is a new runtime mechanism with its own design decisions (TTL semantics, refusal posture), not a signature change. PR for #8367 documents the advisory limitation in the module doc and changeset rather than implying enforcement.Refs: #8367 (seam widening),
objectstack-ai/cloud#1275(the ruling),objectstack-ai/cloud#1291(consumer half), clouddocs/adr/ADR-0022 D4 (multi-node gating).