|
| 1 | +--- |
| 2 | +"@objectstack/objectql": minor |
| 3 | +"@objectstack/runtime": minor |
| 4 | +--- |
| 5 | + |
| 6 | +fix(runtime,objectql): `/api/v1/ready` drains only on the PRIMARY datasource's failure; a secondary/tenant datasource is reported, not drained (#13408) |
| 7 | + |
| 8 | +On a multi-datasource deployment, one datasource whose driver could not start |
| 9 | +pinned `/api/v1/ready` to 503 on **every replica** — so a readiness-checked load |
| 10 | +balancer drained every upstream and took the whole deployment offline, while |
| 11 | +Postgres and the app itself were healthy (`/api/v1/health` 200, direct reads |
| 12 | +working). One tenant's misconfiguration = total outage. Observed on a live |
| 13 | +3-replica EE deployment and recovered only by restarting every process. |
| 14 | + |
| 15 | +Ruled 2026-08-31 (第 6 场总监席决裁批 #12, maintainer verbatim 「同意」), Option B: |
| 16 | + |
| 17 | +> `/api/v1/ready` 只在**主/默认数据源**不健康时摘流量;次要/租户数据源的故障照常 |
| 18 | +> **上报**(`/ready` 响应 body、日志、告警)但不 drain 节点。 |
| 19 | +
|
| 20 | +**What changed.** When a driver reports itself unhealthy, `/ready` now asks |
| 21 | +which datasource is the deployment's primary one before choosing a status: |
| 22 | + |
| 23 | +- primary unhealthy, or the criterion unresolvable ⇒ **503**, unchanged envelope; |
| 24 | +- primary healthy, only a secondary down ⇒ **200**, with the failed drivers named |
| 25 | + in a new `degraded` block: `{ status, state, degraded: { drivers, primaryDatasource } }`. |
| 26 | + |
| 27 | +The failed driver is never hidden — the rejected alternative of filtering it out |
| 28 | +of the response stays rejected. `degraded` appears **only** on that branch; an |
| 29 | +all-healthy 200 body is byte-identical to before. |
| 30 | + |
| 31 | +**The criterion is a readable fact, not a heuristic.** `ObjectQL.resolvePrimaryDatasource()` |
| 32 | +(new, exported with `PrimaryDatasourceVerdict` / `PrimaryDatasourceUnresolvedReason`) |
| 33 | +answers *where this deployment's platform system objects actually live*, resolved |
| 34 | +through the same five-step routing order every query uses — never registration |
| 35 | +order, never the driver flagged default at registration. The ADR-0057 §3.6 system |
| 36 | +ledgers (`audit` / `telemetry` / `event`) are excluded because they are |
| 37 | +deliberately routed off the primary. Disagreement, silence, or a name with no |
| 38 | +driver behind it return an **unresolved verdict**, never a guess. |
| 39 | + |
| 40 | +**Fail toward draining.** Every way of not knowing — an engine that predates the |
| 41 | +probe, a probe that throws, a malformed verdict, a genuinely split deployment — |
| 42 | +falls back to the old whole-node 503. Staying in rotation requires a *positive* |
| 43 | +reading; the absence of a negative one is not permission. Pinned in both |
| 44 | +directions, including an inversion ablation. |
| 45 | + |
| 46 | +**framework#3756 is not overturned.** Its quantified reason — "a replica that |
| 47 | +would fail 100% of its requests" — still holds in the single-datasource |
| 48 | +deployment it was measured on, where the primary *is* the only source; that |
| 49 | +deployment's answer is unchanged down to the response body. What #3756's |
| 50 | +reasoning never covered is the multi-datasource shape, and that is the only |
| 51 | +branch this carves out. |
| 52 | + |
| 53 | +**Grade: `minor` for both, proposed rather than assumed** — this changes when a |
| 54 | +published operational probe drains a node, so an operator whose alerting keys on |
| 55 | +`/ready` returning 503 for *any* driver failure will see 200 + `degraded` |
| 56 | +instead, and `degraded` is a new response field. Nothing is removed or renamed, |
| 57 | +no declared contract key is added (the ruling explicitly declines that: 「⛔ 本裁不 |
| 58 | +加契约键」), single-datasource behaviour is bit-identical, and no migration is |
| 59 | +required — so `major` overstates it and `patch` understates a deliberate change |
| 60 | +to an availability control surface. |
| 61 | + |
| 62 | +⚠️ Out of scope, tracked separately as **#13578**: `DELETE` of a datasource still |
| 63 | +does not evict the stuck driver from the in-memory engine registry, so the |
| 64 | +datasource keeps appearing in `/ready`'s report until the process restarts. That |
| 65 | +half is a defect under either answer to this card and is queued independently. |
0 commit comments