Skip to content

Commit ae226f6

Browse files
committed
chore: changeset for #13408
1 parent e8bb34a commit ae226f6

1 file changed

Lines changed: 65 additions & 0 deletions

File tree

Lines changed: 65 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,65 @@
1+
---
2+
"@objectstack/objectql": minor
3+
"@objectstack/runtime": minor
4+
---
5+
6+
fix(runtime,objectql): `/api/v1/ready` drains only on the PRIMARY datasource's failure; a secondary/tenant datasource is reported, not drained (#13408)
7+
8+
On a multi-datasource deployment, one datasource whose driver could not start
9+
pinned `/api/v1/ready` to 503 on **every replica** — so a readiness-checked load
10+
balancer drained every upstream and took the whole deployment offline, while
11+
Postgres and the app itself were healthy (`/api/v1/health` 200, direct reads
12+
working). One tenant's misconfiguration = total outage. Observed on a live
13+
3-replica EE deployment and recovered only by restarting every process.
14+
15+
Ruled 2026-08-31 (第 6 场总监席决裁批 #12, maintainer verbatim 「同意」), Option B:
16+
17+
> `/api/v1/ready` 只在**主/默认数据源**不健康时摘流量;次要/租户数据源的故障照常
18+
> **上报**(`/ready` 响应 body、日志、告警)但不 drain 节点。
19+
20+
**What changed.** When a driver reports itself unhealthy, `/ready` now asks
21+
which datasource is the deployment's primary one before choosing a status:
22+
23+
- primary unhealthy, or the criterion unresolvable ⇒ **503**, unchanged envelope;
24+
- primary healthy, only a secondary down ⇒ **200**, with the failed drivers named
25+
in a new `degraded` block: `{ status, state, degraded: { drivers, primaryDatasource } }`.
26+
27+
The failed driver is never hidden — the rejected alternative of filtering it out
28+
of the response stays rejected. `degraded` appears **only** on that branch; an
29+
all-healthy 200 body is byte-identical to before.
30+
31+
**The criterion is a readable fact, not a heuristic.** `ObjectQL.resolvePrimaryDatasource()`
32+
(new, exported with `PrimaryDatasourceVerdict` / `PrimaryDatasourceUnresolvedReason`)
33+
answers *where this deployment's platform system objects actually live*, resolved
34+
through the same five-step routing order every query uses — never registration
35+
order, never the driver flagged default at registration. The ADR-0057 §3.6 system
36+
ledgers (`audit` / `telemetry` / `event`) are excluded because they are
37+
deliberately routed off the primary. Disagreement, silence, or a name with no
38+
driver behind it return an **unresolved verdict**, never a guess.
39+
40+
**Fail toward draining.** Every way of not knowing — an engine that predates the
41+
probe, a probe that throws, a malformed verdict, a genuinely split deployment —
42+
falls back to the old whole-node 503. Staying in rotation requires a *positive*
43+
reading; the absence of a negative one is not permission. Pinned in both
44+
directions, including an inversion ablation.
45+
46+
**framework#3756 is not overturned.** Its quantified reason — "a replica that
47+
would fail 100% of its requests" — still holds in the single-datasource
48+
deployment it was measured on, where the primary *is* the only source; that
49+
deployment's answer is unchanged down to the response body. What #3756's
50+
reasoning never covered is the multi-datasource shape, and that is the only
51+
branch this carves out.
52+
53+
**Grade: `minor` for both, proposed rather than assumed** — this changes when a
54+
published operational probe drains a node, so an operator whose alerting keys on
55+
`/ready` returning 503 for *any* driver failure will see 200 + `degraded`
56+
instead, and `degraded` is a new response field. Nothing is removed or renamed,
57+
no declared contract key is added (the ruling explicitly declines that: 「⛔ 本裁不
58+
加契约键」), single-datasource behaviour is bit-identical, and no migration is
59+
required — so `major` overstates it and `patch` understates a deliberate change
60+
to an availability control surface.
61+
62+
⚠️ Out of scope, tracked separately as **#13578**: `DELETE` of a datasource still
63+
does not evict the stuck driver from the in-memory engine registry, so the
64+
datasource keeps appearing in `/ready`'s report until the process restarts. That
65+
half is a defect under either answer to this card and is queued independently.

0 commit comments

Comments
 (0)