|
| 1 | +--- |
| 2 | +"@objectstack/spec": minor |
| 3 | +"@objectstack/platform-objects": minor |
| 4 | +"@objectstack/core": minor |
| 5 | +"@objectstack/metadata-protocol": patch |
| 6 | +--- |
| 7 | + |
| 8 | +feat(core,platform-objects,spec): the ADR-0119 D2 migration-journal runner — a migration killed mid-run is resumable to completion or compensable to clean, with journal rows proving which (#4617) |
| 9 | + |
| 10 | +**The gap D1 left open.** ADR-0119 D1 made `engine.transaction()` reachable |
| 11 | +through the contract, which is the right answer for multi-write atomicity that |
| 12 | +fits in one transaction. Migration-class work does not fit: a million-row |
| 13 | +backfill cannot hold one write-lock for its duration, `driver-memory`'s |
| 14 | +`beginTransaction` deep-clones the entire database (O(db) per begin), |
| 15 | +`ObjectQL.transaction()` binds the **default driver only** so a multi-datasource |
| 16 | +migration silently commits part of its work outside it, and a process **killed** |
| 17 | +— as distinct from a thrown error — defeats in-process rollback entirely. So the |
| 18 | +unit of atomicity is the *chunk*, and durability across chunks is a journal. |
| 19 | + |
| 20 | +Four consumers had each converged on the same four moves — dry-run preflight, |
| 21 | +undo journal, LIFO compensation, re-entrant forward recovery (ADR-0105 D13 |
| 22 | +promotion, ADR-0117 D8's ownership backfill, the org lifecycle transitions, and |
| 23 | +D10 master-data distribution #4585). One copy is engineering; four is platform |
| 24 | +debt, and the fourth author would have had to rediscover the invariant below |
| 25 | +from scratch. |
| 26 | + |
| 27 | +**New: `runMigrationJournal` (`@objectstack/core`).** Preflight runs every |
| 28 | +step's read-only validator before any step writes, so a plan that would fail at |
| 29 | +step 3 has not written step 1. Rows are chunked per the `bulk-write.ts` |
| 30 | +discipline; each chunk's writes run inside `engine.transaction()`. On failure, |
| 31 | +committed chunks are compensated newest-first, each in its own transaction. On |
| 32 | +restart, a rediscovered run resumes forward from the first chunk lacking |
| 33 | +`chunk_done`, or unwinds, per the plan's `onCrash` policy. Forward and |
| 34 | +compensate callbacks receive an `attempt` counter; `attempt > 1` means the prior |
| 35 | +outcome is UNKNOWN and the callback must recheck by natural key before |
| 36 | +re-writing — the same at-least-once contract `bulk-write.ts` already documents, |
| 37 | +reused rather than re-derived. |
| 38 | + |
| 39 | +**The invariant that carries the design:** `chunk_done(i)` is written **inside** |
| 40 | +the chunk's own transaction, so `done ⇔ committed` holds by construction; |
| 41 | +`chunk_started(i)` is written autonomously **before** it. That asymmetry is what |
| 42 | +gives `started ∧ ¬done` exactly one meaning — *the outcome is unknown* — which |
| 43 | +is the only state a crash can leave and the only state recovery reasons about. |
| 44 | +Making both writes symmetric would look tidier and would destroy recovery. |
| 45 | + |
| 46 | +**New: `sys_migration_journal` (`@objectstack/platform-objects`).** Rows keyed |
| 47 | +`(run_id, seq)` under a unique index, so a resumed run that miscomputes its next |
| 48 | +sequence fails loudly rather than double-recording an event. Registered |
| 49 | +unconditionally alongside `sys_migration` because recovery must be discoverable |
| 50 | +with **zero host wiring** — a journal some kernels compose and others do not is |
| 51 | +a journal a boot scanner cannot rely on (ADR-0078). Distinct in grain from |
| 52 | +`sys_migration`, which holds one durable verdict per named migration; this holds |
| 53 | +many rows per *run*. Read-only over the API; writes go through the runner in |
| 54 | +system context. |
| 55 | + |
| 56 | +**The runner refuses rather than degrades**, in four places: the runtime cannot |
| 57 | +roll back; any preflight fails; the plan declares `onCrash: 'compensate'` but a |
| 58 | +step cannot compensate; or a resume's plan hash disagrees with the journal |
| 59 | +(resuming a changed plan would apply chunk boundaries the journal never |
| 60 | +described). A compensation failure halts and is journalled — never swallowed — |
| 61 | +and the run ends `failed`, not `compensated`, because a database in a state no |
| 62 | +clean story covers must not be reported as a tidy rollback. |
| 63 | + |
| 64 | +**`engineCanRollBack` is now shared.** The two-level probe (engine method AND |
| 65 | +default-driver `beginTransaction`) was the same condition written twice — here |
| 66 | +and in `batchData`'s atomic gate. It now lives in `@objectstack/core` and |
| 67 | +`@objectstack/metadata-protocol` imports it, as a type predicate so callers do |
| 68 | +not each re-narrow the optional member by hand. Two copies of "can this runtime |
| 69 | +actually roll back?" drift by one clause and leave one caller believing it has |
| 70 | +atomicity it does not have. |
| 71 | + |
| 72 | +Boot reconciliation and `os migrate resume` land separately; `findInterruptedRuns` |
| 73 | +is the discovery primitive they will consume, and is exported here. |
| 74 | + |
| 75 | +**Docs:** ADR-0118 (plugin-reachable transactions) is renumbered **ADR-0119**. |
| 76 | +It merged one day after an unrelated ADR-0118 (非用户 actor 的平台契约) and the |
| 77 | +earlier merge holds the number; citations of "ADR-0118 D1/D2/D3/D4" written |
| 78 | +before 2026-08-03 mean the renumbered record. |
0 commit comments