Repository navigation
Read users and roles through the record cache, ending user-cache refresh storms - #2776
Merged
Merged
Conversation
A replicated hdb_user/hdb_role write left a subscription-lifetime flag set, so every later system-database commit on that subscription signalled a user change, and the signal's .then had no rejection handler, so each failed commit also raised an unhandledRejection. Every signal rebuilds the user cache on every thread with full hdb_role + hdb_user scans, uncoalesced. - Table.ts: mark the commit's own context on an accepted user/role write and signal only when that commit succeeds; attach the continuation only for marked commits, including writes staged late into an open begin_txn. - user.ts: refresh through coalesceRefresh, one scan in flight per thread plus one trailing scan that starts after it, so a caller never resolves on a scan that read before its call. Dispatch-Task: harper-replicated-user-cache-refresh-storm Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M3df4Rq3g3EEaskdCiyiuh
A source-applied put with no record content is reported and skipped by _writeUpdate, so its commit stages nothing; marking it let a redelivered malformed hdb_user put still rebuild every thread's user cache. Pin exactly one signal per transaction however many user writes it carries, and drop the test comments that restated names. Dispatch-Task: harper-replicated-user-cache-refresh-storm Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M3df4Rq3g3EEaskdCiyiuh
A refresh that called the coalesced function synchronously saw `running` still unset and started a second, overlapping run (measured maxActive 2 by the v5.1 backport's review). Defer the refresh one microtask so `running` is set before it runs, and drop a comment that restated its condition. Dispatch-Task: harper-replicated-user-cache-refresh-storm Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M3df4Rq3g3EEaskdCiyiuh
… copy Authentication kept a per-thread, pre-joined copy of every hdb_user row with its hdb_role row and rebuilt it with a full scan whenever a USER ITC event arrived. Eight producers had to send that event (user and role operations, refresh-token updates, the replication apply loop), so a missed one left authorization stale and a spurious one cost a scan on every thread; the replication storm was the second kind. Lookups now point-read the user by name and its role by id through the primary store's record cache, which already stays coherent across threads. Nothing is rebuilt and nothing has to be announced: - readUserEntries re-checks the user's version after reading its role, so two unsnapshotted RocksDB reads cannot pair a user with a role from a different committed state. - The role's system-table permissions are memoized per role version, and every returned user gets its own role and permission objects. - auth.ts validates each authorizationCache hit against the user and role versions it was built from (isCurrentUser), so a password, role or activity change takes effect on the next request on every worker. - onUserChange, fed by per-thread hdb_user/hdb_role subscriptions, notifies the consumers that hold a user already: live-subscription revocation (which now runs a trailing sweep for a change that lands mid-sweep) and MCP list-changed. It never subscribes to an unaudited table, since subscribe() would enable and persist auditing. Deleted: usersWithRolesMap, setUsersWithRolesCache, getUsersWithRolesCache, signalUserChange, the ITC userHandler and USER event type, UserEventMsg, the apply loop's user/role mark, and the startup cache warm-ups. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A user record a resequenced write left on a reused version made the re-read loop in readUserEntries spin forever: its version can never be confirmed, so every re-read looked like a change. Such a record is now accepted as read, and is never current for isCurrentUser or the role memo. A metadata-less record keeps a stable version until it is next written. - A component's server.getUser principal in the authorization cache is checked against the user and role versions read for its name before it was resolved (trackUserRecords), instead of relying on the whole cache being cleared by a user-change notification. - A failed hdb_user/hdb_role subscription is retried with backoff. - The derived-role memo is bounded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A record a resequenced write left on a reused version was accepted without the torn-read re-check, and never counted as current, so its holders re-resolved on every request. Such a record is now stamped with its value instead of its version: the re-check, isCurrentUser and the role memo compare values for it. A listener that returns a non-promise value no longer logs a false failure. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kriszyp
force-pushed
the
fix/replicated-user-cache-refresh-storm
branch
from
September 24, 2026 17:41
c556abc to
d98465c
Compare
…roles readUserEntries let a Basic-auth username over LMDB's 1978-byte key limit reach store.getEntry, which throws instead of the intended clean 401; a long username is now treated as no such user. Its user/role recheck loop was also unbounded, so a sustained write storm on one user record could spin the event loop; it now falls back to no-user after 50 attempts. appendSystemTablesToRole assumed every hdb_role record had a permission object. A role with no permission — reachable only via a direct table write or a replicated write, since the operations API requires it — made listUsers() throw, taking down assertActiveSuperUserRemains with it (drop_user, alter_user, alter_role, drop_role) and getSuperUser's rescan. Also drops two comments that narrate the diff rather than the code. Found by round 7 of this branch's pre-push review (codex, gemini, cursor-kimi, domain adjudication); both are new minor/nit findings, no blockers survived adjudication. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016YrgfpWWVHjruNRZA7fyAg Dispatch-Task: pr-maint-f8a32eb70bff6d0d2b73f7c434344aa1
…guard The previous guard checked Buffer.byteLength against LMDB's 1978-byte limit, but primary-store keys are ordered-binary encoded and characters U+0000-U+0003 escape to two bytes each. A username like '\u0001'.repeat(1000) passed the byte-length check at 1000 bytes while its encoded key ran to about 2000, so it still reached store.getEntry and threw. Moved the guard into readEntry, mirroring resources/Table.ts's checkValidId two-tier check (a fast path under 659 characters, else measure with writeKey), so it covers every read through this module including isCurrentUser's re-check of a previously-resolved username. Also seed the malformed-role regression test through testUtils.seedUsers() instead of writing straight to hdb_role, so afterEach restores it instead of leaking it into later suites. Found by round 8 of this branch's pre-push review (codex, gemini, cursor-kimi, domain adjudication). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016YrgfpWWVHjruNRZA7fyAg Dispatch-Task: pr-maint-f8a32eb70bff6d0d2b73f7c434344aa1
…mments keyTooLargeForStore returned false for any non-string id, so a role reference set to an object or array by a direct or replicated write would still reach store.getEntry and throw on LMDB. It now treats anything but a string or number as unusable, same as the oversized-string case. Drops three comments that restated the following line or answered a reviewer rather than documenting a non-obvious invariant. Found by round 9 of this branch's pre-push review (codex, gemini, cursor-kimi, domain adjudication); the other majors it raised (retry exhaustion returning no-user, systemStore's error status) were dropped or downgraded by domain adjudication as intentional/factually wrong. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016YrgfpWWVHjruNRZA7fyAg Dispatch-Task: pr-maint-f8a32eb70bff6d0d2b73f7c434344aa1
Table.ts's checkValidId also validates object and bigint keys; keyTooLargeForStore rejects them outright instead. Restated the comment around what this guard actually covers. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016YrgfpWWVHjruNRZA7fyAg Dispatch-Task: pr-maint-f8a32eb70bff6d0d2b73f7c434344aa1
kriszyp
marked this pull request as ready for review
September 24, 2026 21:13
github-actions
Bot
requested review from
Ethan-Arrowood,
cb1kenobi and
heskew
September 24, 2026 21:13
kriszyp
added a commit
that referenced
this pull request
Sep 25, 2026
…v5.1) (#2777) * Stop replicated system writes from storming the user cache (v5.1) Backport of #2776 to v5.1. A replicated hdb_user/hdb_role write left a subscription-lifetime flag set, so every later system-database commit on that subscription signalled a user change, and the signal's .then had no rejection handler, so each failed commit also raised an unhandledRejection. Every signal rebuilds the user cache on every thread with full hdb_role + hdb_user scans, uncoalesced. - Table.ts: mark the commit's own context on a user/role write of a known source write type and signal only when that commit succeeds; attach the continuation only for marked commits, including writes staged late into an open begin_txn. v5.1 has no unknown-operation guard or valueless-put skip, so the mark checks the write type itself and a valueless put still counts. - user.ts: refresh through coalesceRefresh, one scan in flight per thread plus one trailing scan that starts after it. Dispatch-Task: harper-replicated-user-cache-refresh-storm Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M3df4Rq3g3EEaskdCiyiuh * Keep coalesceRefresh to one run under re-entry; check the database first A refresh that called the coalesced function synchronously saw `running` still unset and started a second, overlapping run. Defer the refresh one microtask so `running` is set before it runs. Test the database name before the write-type Set lookup so application-database writes skip it. Dispatch-Task: harper-replicated-user-cache-refresh-storm Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M3df4Rq3g3EEaskdCiyiuh --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Authentication now reads
hdb_userandhdb_rolethrough the primary store's record cache instead of keeping a per-thread copy of both tables, so a user or role change needs no broadcast and there is nothing left to rebuild. The replicated-write refresh storm this PR was opened for has no mechanism left: the replication apply loop no longer knows about users at all (theuserRoleUpdateflag and its signal are deleted).readUserEntriespoint-reads the user by name and its role by id, then re-checks the user after reading its role: RocksDB point reads share no snapshot, so one transaction that moves a user from role A to B and grants A super_user must not be read as "user on A, with A's grant". The re-check costs one nativeverifyVersionon RocksDB (isUnchanged).stampOf,holdsStamp) is compared by value, since its version no longer identifies one value.authorizationCachehit runsisCurrentUseragainst the versions the principal was built from (isCurrentUser), so a password, role or activity change takes effect on the next request on every worker. A component'sserver.getUserprincipal is checked against the versions read for its name before it was resolved (trackUserRecords).onUserChangefeeds live-subscription revocation and MCP list-changed from per-threadhdb_user/hdb_rolesubscriptions. It never subscribes to an unaudited table, becausesubscribe()would enable and persist auditing on it, and it retries a failed subscription with backoff.For the human reviewer
Framing (step-6 planning review):
Framing-Verdict: better-alternative-exists, resolved mostly by adopting it. The reviewer endorsed removing the copy, then asked for durable versions instead of object identity, a torn-read guard, no silent audit enablement, and a subscription failure contract. All four are in. Two parts were overruled, each on a fact:node_modules/lmdb/read.js:1062-1070), so each in-flight transaction would hold its own reader slot, and a reset costs ~1.5 µs plus a timer per lookup (measured below). Kept instead: LMDB lookups see another thread's commit from the next event-loop turn, like every other LMDB read. On a single busy turn, a request that arrives after an acknowledgement can still see the pre-change snapshot. RocksDB lookups have no such window.resources/transactionBroadcast.tsaddSubscription), so without auditing there is nothing to observe from.The requirement itself (your direction in this PR's thread): warranted. The storm was one instance of a design in which eight producers had to remember a broadcast, and a spurious broadcast cost an
O(users + roles)scan on every thread. Removing the copy removes both the missed-signal and the extra-signal failure classes. The earlier per-commit mark plus coalescer is gone rather than kept as a second layer;utility/coalesceRefresh.tssurvives only for the live-subscription sweep.Per-request cost, the price of reading records. Measured on this machine, unit environment, 50,000 iterations, RocksDB / LMDB:
authentication()Basic cache hitfindAndValidateUser(session, mTLS, token user, local bypass)The cache-hit check is two native
verifyVersioncalls on RocksDB (0.4 µs each) and two validated reads on LMDB. A change costs nothing extra on any thread. Tell me if the hit-path cost is too high: the alternative is to drop the per-hit check and rely ononUserChangeto flush the cache, which brings back the asynchronous window.Cross-worker completion semantics changed.
alter_user,alter_roleand the rest no longer await every worker. They do not need to for lookups: the first read after the commit on any worker sees it (subject to item 1 on LMDB). The push consumers (live-subscription revocation, MCP notifications) were never awaited onmaineither; they now fire from the table subscription.Audit-only push channel. On a node with
logging.auditLog: false, live-subscription revocation falls back to its 30-second backstop sweep, and MCP sessions get no user-changelist_changed. Each thread logs this once at info. Authentication itself is unaffected.mainenables auditing onhdb_certificateimplicitly (security/keys.tssubscribes unconditionally); I did not want to extend that to tables holding password hashes.Component principals are tracked by the Basic username. A component's
server.getUserprincipal is re-verified once thehdb_user/hdb_rolerecords named by the Basic username change. A component that maps that username to a different user, or takes the role from elsewhere, keeps its principal untilauthentication.cacheTTL(30 s) orserver.invalidateUser().mainflushed the whole cache on any user event; adding that flush back throughonUserChangeis a one-liner if you want both.Deleted outright, not deprecated: the
USERITC event type,signalUserChange,UserEventMsg,serverHandlers.userHandler,setUsersWithRolesCacheandgetUsersWithRolesCache.harper-prooutsidecorehas no callers (grep of itsmain). An external component calling them would fail at load.VERSION_REUSEDrecords. An out-of-order replicated write (for example,create_authentication_tokenson two nodes) leaves a user record on a reused version until its next in-order write. Such a record is compared by_.isEqualof its value, so cache hits and the role memo still work, at the cost of one decode plus a deep compare per hit. The first version of this redesign looped forever on such a record; the review caught it, and a test now pins it.Backports. Fix user-cache refresh storm from replicated system-database writes (v5.1) (#2777) stays the narrow per-commit fix, which suits a patch line. This redesign is
main(v5.3) only.v5.2still has the original defects; backport scope there is your call.Deliberately unchanged:
listUsers()still scans both tables through the legacy search, now only on cold paths (list_users, the last-super-user guard, and the firstgetSuperUsercall).getSuperUserremembers the super user it found and re-checks it per call, instead of iterating a map.Rebased onto
main's07b2c9015(the deferred-CF-reclamation merge). One textual conflict, inresources/Table.ts's import list: kept main's newdropColumnFamily/markDropInProgress/recordRetiredGeneration/sweepDroppedTableBlobs/storeNameFor/storeNamesForimports, droppedUserEventMsgper this PR's removal of the ITC user broadcast. No semantic overlap — that work is table-drop generation reclamation, disjoint from user/role lookups.Two correctness bugs the post-rebase review found, both fixed.
appendSystemTablesToRolenow tolerates a role record with nopermission(reachable only via a direct table write or a replicated write, since the operations API requires it) instead of throwing and taking downlistUsers()— and with itassertActiveSuperUserRemains(drop_user,alter_user,alter_role,drop_role).readEntrynow guards every user/role store read against an id that can't be a key: a username over LMDB's 1978-byte encoded-key limit (measuring the ordered-binary length, not UTF-8 bytes — escaped low control characters expand it) or a non-string/non-number id (a malformed stored role reference) previously reachedstore.getEntryand threw, turning an unauthenticatable credential into an internal fault instead of a clean 401.readUserEntries's consistency retry is now bounded at 50 attempts (MAX_USER_ENTRY_ATTEMPTS), falling back to "no such user" rather than looping the event loop forever under a sustained write storm on one user record. The alternative — throwing an internal fault instead of a credential rejection after exhaustion — is a one-line change if you'd rather distinguish the two; either fails closed.Three declined findings, left as-is:
hdb_user/hdb_roletable subscriptions (security/user.tsaround L649) make every commit tosystemtrigger an audit-log scan on a worker with no othersystemsubscription. In the default config, TLS listeners already subscribe tosystem.hdb_certificate, so the added cost there is per-record dispatch only; a node with no TLS listener pays the new scan.userViewallocates a provenance object and aWeakMapentry on every principal, including the session-cookie path, which never callsisCurrentUseron it.server.getUser(security/auth.tsaround L263); with the default resolver this is redundant, sincefindAndValidateUserre-reads the same records. Both are cheap per-request costs, not correctness issues.integrationTests/security/user-change-across-workers.test.ts'srejectedEverywheresamples 32 responses without identifying which worker answered (a 401 carries nothreadId), so the revocation half of the test is probabilistic rather than exhaustive (~1e-4 chance of missing a worker, not zero).Changes
security/user.ts: the map and its rebuild are gone (usersWithRolesMap,setUsersWithRolesCache). New:systemStore(throws a plain, untagged error when a system table is missing, so a storage outage is never read as a credential rejection);withSystemTablePermissions(now tolerates a role with nopermission), shared withlistUsers;userViewandgetUserWithRole;findAndValidateUseron the record;dropUser's existence check; the last-super-user guard on an explicit scan; the subscription runs outside any request context; notifications coalesced to one per event-loop turn.addUser/alterUser/dropUserno longer refresh or signal.readEntryrejects an id too large or the wrong shape to be a key before it reaches the store. Look hardest at the re-check inreadUserEntriesandholdsStamp: they are the whole consistency argument.security/auth.ts: the per-hitisCurrentUsercheck, versions read before a Basicserver.getUser, and the ITC listener that flushed the cache is removed.server/liveSubscriptionAuth.ts: the sweep is triggered byonUserChange(header), and a change that lands mid-sweep gets a trailing sweep instead of being dropped by thesweepingguard.components/mcp/listChanged.ts: user changes come fromonUserChange; schema and resource changes still arrive over ITC.security/impersonation.ts(scoped-token name check,lookupUser),security/authn/oidc/tokenExchange.ts(policy user) andsecurity/authn/oidc/trustPolicyOperations.ts(add, list):getUserWithRolereplaces map reads.security/role.ts(add, alter, drop),security/tokenAuthentication.ts(refresh-token update), andresources/Table.ts(the apply-loop mark, the post-commit signal).USERplumbing deleted:utility/signalling.ts(signalUserChange),server/itc/serverHandlers.js(userHandler, exports),server/threads/itc.js(UserEventMsg),server/threads/manageThreads.js(listener registration),utility/hdbTerms.ts(USERtype).server/operationsServer.ts(setUp),server/fastifyRoutes.ts(side-effect import kept forserver.getUser),server/jobs/jobProcess.ts(same).resources/Resource.ts: comment now says where live-subscription rechecks read from.utility/coalesceRefresh.ts: one run in flight plus one trailing run, carried over from the first version of this PR, now used by the sweep.security/DESIGN.mdreplaces the coalescing note with the lookup invariant; root index;server/DESIGN.mdwording (rejection provenance, system-table carrier).Verification
Route: new unit tests against real
hdb_user/hdb_rolerecords (no stubs of the lookup path), an in-process replication apply-loop test, and a new multi-worker integration test.unitTests/security/userRecordLookups.test.js(19 cases; 18 on LMDB, where the reused-version case does not apply):hdb_useris found with no refresh or signal; deactivation and deletion and a role permission change take effect on the next lookup; the derived role is shared until its record changes;hdb_userwrite returns (it hung the first revision) and is compared by value;getSuperUserre-checks the user it remembered;isCurrentUsertransitions;authentication()cache hits: password change, deactivation and role change, and a component-provided principal;onUserChangefires for both tables, contains a throwing or rejecting listener, and a change during a sweep gets its own sweep.unitTests/resources/replicatedUserWrites.test.jsdrivesTable.sourcedFromwith a fake source: a replicated user write is visible to the next lookup and notifies; later non-user commits do not notify; a replicated role change reaches the next lookup; a failed commit leaves no user, notification or unhandled rejection; a late write inbegin_txn.integrationTests/security/user-change-across-workers.test.ts, 4 workers, every worker's authorization cache primed first, with theWhoAmIfixture (config): a role change is seen on every worker, identified by thread id; after a password change and a deactivation, 32 fresh connections all get 401 (a 401 does not name its worker; the chance of missing one of 4 workers is ~1e-4). Passes on RocksDB and LMDB. Negative control: with both the per-hit check and the notification flush removed fromdist/, all 3 cases fail ("worker 3 served the role as it was before alter_role").dist/build with a base-compatible subset: a user written straight tohdb_useris not found ("Login failed"), and a cached Basic credential survives a password change (200, expected 401). Both pass here.testUtils.seedUsers, which writes in one transaction, restores what it replaced, and waits for the notification its writes cause:auth-fastify(role fields asserted instead of a deep-equal on the old cached shape),authn/oidc/tokenExchange,authn/oidc/trustPolicyOperations,impersonation,tokenAuthentication,tokenRejectionClassification(a missingsystem.hdb_userstill propagates untagged),tokenScopeRefresh;user.test.jsnow restores the module instances it re-requires, which otherwise left a seconduser.tsinstalled asserver.getUser. Removed with the APIs they tested:signalUserChange,UserEventMsg(and its ipc twin),userHandler;serverHandlers.test.jspreloadscomponents/status, whichserverHandlersno longer loads transitively.coalesceRefresh.test.jsis unchanged apart from a lint fix.c556abce9(pre-rebase):npm run test:unit:main: 5843 passing, 1 failing:gitCredentials.test.js:395, the dispatch environment'sGIT_CONFIG_GLOBAL/GIT_EDITOR(19/19 with them unset).npm run test:unit:resources: 3003 passing, 0 failing.npm run test:integration:all: 2138 pass, 0 fail, 6 cancelled (all inollama-backend.test.ts, which needs a local Ollama; the same on base).npm run typecheck,prettier --check,oxlinton changed files,npm run check:design-docs: clean. Repo-wideoxlint --deny-warningsreports 15 warnings, all in files this PR does not touch.9eff24291(post-rebase, current head):test:unit:main5873 passing, 1 failing (the same knowngitCredentialsartifact);test:unit:resources3034 passing, 0 failing; the security integration suite (integrationTests/security/**, 24 suites including the new one) 128 passing, 0 failing. CI green on every leg, including the two prior LMDB-alias-close segfaults (#2721, still open) that hit this PR's earlier heads — a clean re-run each time.listUsers, an unbounded consistency-retry loop, and (over two rounds, since the first attempt only checked UTF-8 byte length) the LMDB oversized/malformed-key guard inreadEntry; declined the four items in "For the human reviewer" Introduce dependency lockfile #14 as accepted trade-offs or out-of-scope test-coverage gaps. Converged with only previously-adjudicated and pre-existing findings remaining.— Claude Opus 5.5
Complexity: complicated
Origin — the dispatch brief this PR was written from
Fix user-cache refresh storm from replicated system-database writes
LIVE CONVERSATION about #2776.
You are answering a person, in a thread, one turn at a time. Every turn:
each of your previous turns is in it. Read the PR/issue and the code as needed.
they are talking to you, and a status template is not an answer.
Each turn arrives as ASK (answer it, change nothing) or PERFORM (do it, then say what you did) —
the person chose which when they sent it, and the run's own prompt tells you which one this is.
Never infer it from the wording: an unrequested commit in the middle of a discussion and a polite
description of work that was supposed to happen are the two failures this exists to prevent.
Never mark a PR ready and never merge from this conversation.
Dispatch: task
chat-pr-harper-2776-kriszyp· queued by unknown · ran by claude/opus/xhigh · worker kzyp-xps-1Review-Coverage: authored=claude; ran=cursor-kimi,codex,gemini; adjudicated=domain; declined=cursor-grok,cursor-composer,cursor-muse; rounds=11; full=3 @ 9eff242
Human-Review-Need: 4 (decisions: read-through-user-lookups, retry-exhaustion-as-unknown-user, unkeyable-id-as-absent, notification-requires-audit, remove-user-itc-surface, component-principal-provenance, impersonated-principal-shape) @ 9eff242