Skip to content

fix(fork): keep the client usable while the attached server restarts - #1712

Open
donjor-agent[bot] wants to merge 2 commits into
hyprwsfrom
fix/attached-primary-keep-alive
Open

donjor-agent[bot] wants to merge 2 commits into
hyprwsfrom
fix/attached-primary-keep-alive

Conversation

@donjor-agent

@donjor-agent donjor-agent Bot commented Oct 11, 2026 •

Copy link
Copy Markdown

🤖 claude/anthropic claude-opus-5-5 (medium)

Problem

A server update restarts the attached backend, and the desktop client drops the environment.
Projects, threads, composer drafts, and the open route vanish until the server returns.

Change

Side Change
Desktop Reattach when a restarted server reports the same environment id
Renderer Keep a recently seen primary registered (offline) for a 5 min grace
Grace clock Monotonic (performance.now()), immune to NTP steps
Past the grace Removed as upstream does
Send Existing offline gate disables it

The renderer logic lives in primaryAbsenceGrace.fork.ts behind two marked hooks in platform.ts.

Why the monotonic clock

The first build still wiped in the VM.
CDP logpoints on the poll showed the cause:

  1. Server stops; the poll keeps the primary
  2. NTP steps the wall clock +54 min mid-restart
  3. Wall-clock grace reads as expired
  4. Poll removes the primary; drafts cleared

Scope and approval

Refs #1706, settled in the #1705 profile.
Its nightly proof on a large DB follows merge, so the Task closes after.

Verification

Claim

Type a draft, click Update, wait out the restart: the thread, sidebar, and draft stay.

Build

  • Before: official nightly 20261011.1011
  • After: this stack at 6150d67, stamped 1011
  • VM: isolated QA desktop, large DB copy
  • Server: in-app update, 1009 to 1011

Stress trials

Each trial reverts the VM snapshot; its stale guest clock forces an NTP step during the restart.

Build Trials Wiped
Official 1011 2 2
First PR build (wall clock) 6 5
This head (monotonic) 7 0

Cut table

Span Real Clip
Draft, click Update 10s 1x
Restart wait 36–60s 6–8x
Reconnect 11–14s 1x

Before

before-clip.mp4

After

after-clip.mp4

Checks

Check Result
Grace, platform, registry, desktop attach tests pass
tsc --noEmit on apps/web exit 0
vp run fork:ci exit 0

Claude Opus 5.5 in Claude Code, hosted by T3 Code.

Fork trailers

Fork-Domain: backend-attach
Fork-Tier: bugfix
Fork-Upstreamable: no

@rsi-pr-agent

rsi-pr-agent Bot commented Oct 11, 2026 •

Copy link
Copy Markdown

🩺 Aftercare 0b005de

follow-up 1 resolved 5

commit 25aab54

  • A1 Delayed restart refresh overwrites a newly minted bearer ✅ ebfebe3
  • A2 Reappearing primary at the same URL retains the wrong environment ✅ ebfebe3
  • A3 Mint completion restores the pre-restart server identity ✅ ebfebe3
  • 🔵 A4 Complete the large-DB nightly VM proof 🔁 ×3

commit ebfebe3

  • A5 Stale verification failure drops a newer verified attachment ✅ ca8272e
  • A6 Transient recovery discovery failure removes the primary before grace ✅ ca8272e

🕐 11 Oct 11:04 UTC ⏱️ 3 min

💬 Push to fix; reply A4: <reason> to dispute.

Details

A1 Delayed restart refresh overwrites a newly minted bearer apps/desktop/src/app/DesktopAttachedPrimary.ts:339
reverify publishes the attachment snapshot captured before asynchronous verification. A concurrent mint can publish token-2, then refresh restores token-1 so a late token-1 reject drops the successor. Fix: Update the current attachment with a stale-result guard instead of spreading captured held state. Add a gated refresh/mint test proving that a late old-token reject leaves the successor attached. Note: The original gated refresh/mint replay now preserves token-2 and ignores the late token-1 rejection; the head adds the matching regression test.

A2 Reappearing primary at the same URL retains the wrong environment apps/web/src/connection/platform.ts:623
Grace retention preserves A's endpoint signature. If environment B returns at the same URL, cache reuse skips descriptor discovery and keeps A registered indefinitely; the base correctly discovers B. Fix: Revalidate the descriptor when a grace-retained primary reappears. Retain identity only when its environment ID matches, and add an A-to-absent-to-B same-endpoint polling regression test. Note: Grace retention invalidates the endpoint signature. The actual polling replay rediscovers same-URL environment B instead of retaining A.

A3 Mint completion restores the pre-restart server identity apps/desktop/src/app/DesktopAttachedPrimary.ts:418
A mint spanning a restart writes its old live.value over the refreshed instance. An unauthorized reject then skips the old PID, allowing the same rejected restarted server to auto-attach immediately. Fix: Preserve the latest verified instance when publishing a mint result. Extend the gated restart test to reject the new token immediately and prove the restarted instance remains skipped until another restart. Note: Mint completion preserves the latest verified PID. Immediate unauthorized rejection skips PID 99, and only another restart to PID 100 reattaches.

A4 Complete the large-DB nightly VM proof
#1706 requires a large-DB clip covering retained UI state, cached navigation, Send recovery, and grace expiry. The PR explicitly defers this proof until a nightly contains the fix. Fix: After merge, run the large-DB VM scenario on a nightly containing the fix, attach the clip, and verify all task close conditions before closing the task. Note: Follow-up only: no new large-DB proof was found. The previously explained two-nightly requirement remains consistent with the deferred post-merge verification under #1705. V….

A5 Stale verification failure drops a newer verified attachment apps/desktop/src/app/DesktopAttachedPrimary.ts:336
The verification sequence guard runs after the failure branch calls dropNow. An older descriptor request failing after a newer refresh verifies the restarted server still clears that healthy attachment and its cache. Fix: Apply verification ordering before any failure-driven drop or successful state/cache update. Add a gated test where a newer restart verification succeeds before an older request fails. Note: The sequence/generation guard now runs before failure-driven removal. The gated replay confirms that an older verification failure preserves the newer verified attachment and cache; the head adds its….

A6 Transient recovery discovery failure removes the primary before grace apps/web/src/connection/platform.ts:628
After absence invalidates the signature, a returning target triggers descriptor discovery. If that request fails transiently, the poll emits no primary registration, clearing cached data and drafts before the grace expires. Fix: Retain the prior registration through transient recovery-validation failures until successful identity discovery or the original grace deadline. Keep rediscovery active and test present-to-absent-to-timeout-to-absent polling. Note: The hook now runs after discovery and retains the prior registration through failed recovery validation. Polling replays confirm continued rediscovery, removal at the original deadline, and grace ren….

Head Result Time Check
0b005de ✅ review passed: fixed: none; held: A4; new: none; withdrawn: none; no item resolved; files (3): apps/web/src/connection/platform.ts, …and 2 more 3 min run
ca8272e ✅ review passed: fixed: A5, A6; held: A4; new: none; withdrawn: none; files (5): apps/desktop/src/app/DesktopAttachedPrimary.fork.test.ts, …and 4 more 5 min run
ebfebe3 ⚠️ 2 open: fixed: A1, A2, A3; held: A4; new: A5, A6; withdrawn: none; files (5): .github/workflows/hyprws-ci.yml, …and 4 more 11 min run
25aab54 ⚠️ 3 open: fixed: none; held: none; new: A1, A2, A3, A4; withdrawn: none; no item resolved; files (6): apps/desktop/src/app/DesktopAttachedPrimary.fork.test.ts, …and 5 more 10 min run

@donjor-agent
donjor-agent Bot added this pull request to stack #1714 October 11, 2026 07:22
donjor added a commit that referenced this pull request Oct 11, 2026
## Problem

`hyprws CI` ran on pull requests only when the base was `hyprws` or `release/**`.
A native stack layer targets the layer below it, so it got no checks at all.
#1713 showed this: stacked on #1712, it had no CI run.

## Change

Drop the `branches` filter on the `pull_request` trigger.
The ledger already scopes a stacked head from its trunk merge base, so lower layers are judged too.

## Scope and approval

Very small, focused CI fix the maintainer chose (option B) while landing the #1705 stack.

## Verification

| Check | Result |
| --- | --- |
| `vp run fork:ci` | exit 0 |
| CI on this PR | runs, since base is `hyprws` |
| Stacked layer CI | proven once #1713 is pushed after merge |

Claude Opus 5.5 in Claude Code, hosted by T3 Code.

## Fork trailers

Fork-Domain: fork-meta
Fork-Tier: bugfix
Fork-Upstreamable: no
Co-authored-by: donjor <38745786+donjor@users.noreply.github.com>
@donjor
donjor force-pushed the fix/attached-primary-keep-alive branch from 25aab54 to ebfebe3 Compare October 11, 2026 07:51
@rsi-pr-agent rsi-pr-agent Bot added size:L and removed size:M labels Oct 11, 2026
@donjor-agent

donjor-agent Bot commented Oct 11, 2026

Copy link
Copy Markdown
Author

🤖 claude/anthropic claude-opus-5-5 (medium)

A4: preview builds carry no update feed, so the UI update flow needs two nightlies; the proof follows merge under #1705 (Refs, not Closes).

Desktop: a restarted server reporting the same environment id re-verifies
in place with its fresh pid/start time instead of dropping the attachment;
identity stays the env-id find plus descriptor check, and late verifies or
mints never drop or roll back a newer bearer or instance. Renderer: a
primary absent from the topology stays registered (offline) for a
five-minute grace before removal, so an update restart no longer wipes
caches, drafts and open routes; on return it re-runs descriptor discovery,
keeping the old registration through transient failures within the grace,
so a different environment at the same URL replaces it.
Covers #1706.

Fork-Domain: backend-attach
Fork-Tier: bugfix
Fork-Upstreamable: no
@donjor
donjor force-pushed the fix/attached-primary-keep-alive branch from ebfebe3 to ca8272e Compare October 11, 2026 08:10
An NTP step during the restart ended the wall-clock grace early and the
poll wiped the client anyway. Covers #1706.

Fork-Domain: backend-attach
Fork-Tier: bugfix
Fork-Upstreamable: no

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant