Summary
On RocksDB, redeclaring an existing table at runtime opens new __dbis__ and index handles and never closes the ones they replace. A common trigger is an insert that adds a new attribute. closeDatabase only walks the current graph, so those orphaned handles stay open for the life of the thread.
The visible symptom is that online restore_backup is refused for practically every RocksDB database on a running server. verifyDatabaseClosed still sees the orphans in rocksdb-js's process-global registry and fails with:
Cannot restore database '…' while Harper is running: it is held open by a loaded component (or is the system database). Restore it offline instead …
No component and no alias is involved, so this is a different mechanism from #2802.
Reproduction
Found while adding end-to-end coverage for #2965. Reproduces on main @ 16e5c663c.
- Start Harper;
threads.count 1 and 3 both reproduce.
- Run
create_table with database: restore_staging, table: Items, primary_key: id.
- Insert rows carrying an attribute the table didn't declare (
note).
- Run
create_backup, then restore_backup with that backup_id.
- The job ends in
ERROR with the message above.
The registry still shows refCount: 4 on <root>/database/restore_staging after the restore close broadcast. A probe that tracked handle opens and closes per thread showed this after every thread ran closeDatabase:
t3 job afterClose liveStoresThisThread=0 refCount=16
t1 http afterClose liveStoresThisThread=0 refCount=4
t0 main afterClose liveStoresThisThread=3 refCount=4
__dbis__ <- openRocksDatabase <- declareTable <- table <- ResourceBridge.createTable
Items/__createdtime__@… <- openIndex <- declareTable <- table <- ResourceBridge.createTable
Items/__updatedtime__@… <- openIndex <- declareTable <- table <- ResourceBridge.createTable
Their replacements were opened when the insert redeclared the table: Items.addAttributes → table() → declareTable. Force-closing just those three handles in the probe made the online restore complete.
Cause (resources/databases.ts, at 16e5c663c)
__dbis__: in declareTable, attributesDbi is only assigned on the new-table branch. On a redeclare, if (!attributesDbi) opens a new __dbis__ handle and overwrites rootStore.dbisDb without closing the previous one. initStores reuses rootStore.dbisDb; declareTable doesn't.
- Indexes: every indexed attribute is reopened through
openIndex, and indices[attribute.name] = dbi overwrites the previous handle without closing it.
- Nothing tracks them:
GLOBAL_TARGET.adopt is a no-op on the premise that closeDatabase "walks the graph". But closeDatabaseStores reaches only the current table.indices and rootStore.dbisDb, so replaced handles are unreachable. The branch target does record every opened store (branch.openedStores); the global one doesn't.
- LMDB: not affected, because
openDB returns a cached handle per name.
Each redeclare in a thread leaks 1 + (number of indexes) handles. The leak is silent and grows with schema churn.
Suggested direction
The invariant to restore: every handle opened for a root store is reachable from what closeDatabase closes.
- On the redeclare path, reuse
rootStore.dbisDb as initStores does, and keep an index handle whose store name and generation haven't changed.
- As a backstop for handles that are genuinely replaced, record them per root store (as
branch.openedStores does) and have closeDatabaseStores close them. Closing them immediately could break readers still holding the old Table.
- Regression test: create a table, insert a new attribute,
closeDatabase, then assert registryStatus() has no entry for the path. Also make the online-restore integration test run the success path online.
Related: #2802 (alias variant of the same 409), #2965 (found while fixing; its integration test covers the success path offline until this lands).
🤖 Generated by Anthropic Claude (Opus); posted via @cb1kenobi.
Summary
On RocksDB, redeclaring an existing table at runtime opens new
__dbis__and index handles and never closes the ones they replace. A common trigger is an insert that adds a new attribute.closeDatabaseonly walks the current graph, so those orphaned handles stay open for the life of the thread.The visible symptom is that online
restore_backupis refused for practically every RocksDB database on a running server.verifyDatabaseClosedstill sees the orphans in rocksdb-js's process-global registry and fails with:No component and no alias is involved, so this is a different mechanism from #2802.
Reproduction
Found while adding end-to-end coverage for #2965. Reproduces on
main@16e5c663c.threads.count1 and 3 both reproduce.create_tablewithdatabase: restore_staging,table: Items,primary_key: id.note).create_backup, thenrestore_backupwith thatbackup_id.ERRORwith the message above.The registry still shows
refCount: 4on<root>/database/restore_stagingafter the restore close broadcast. A probe that tracked handle opens and closes per thread showed this after every thread rancloseDatabase:Their replacements were opened when the insert redeclared the table:
Items.addAttributes→table()→declareTable. Force-closing just those three handles in the probe made the online restore complete.Cause (
resources/databases.ts, at16e5c663c)__dbis__: indeclareTable,attributesDbiis only assigned on the new-table branch. On a redeclare,if (!attributesDbi)opens a new__dbis__handle and overwritesrootStore.dbisDbwithout closing the previous one.initStoresreusesrootStore.dbisDb;declareTabledoesn't.openIndex, andindices[attribute.name] = dbioverwrites the previous handle without closing it.GLOBAL_TARGET.adoptis a no-op on the premise thatcloseDatabase"walks the graph". ButcloseDatabaseStoresreaches only the currenttable.indicesandrootStore.dbisDb, so replaced handles are unreachable. The branch target does record every opened store (branch.openedStores); the global one doesn't.openDBreturns a cached handle per name.Each redeclare in a thread leaks 1 + (number of indexes) handles. The leak is silent and grows with schema churn.
Suggested direction
The invariant to restore: every handle opened for a root store is reachable from what
closeDatabasecloses.rootStore.dbisDbasinitStoresdoes, and keep an index handle whose store name and generation haven't changed.branch.openedStoresdoes) and havecloseDatabaseStoresclose them. Closing them immediately could break readers still holding the oldTable.closeDatabase, then assertregistryStatus()has no entry for the path. Also make the online-restore integration test run the success path online.Related: #2802 (alias variant of the same 409), #2965 (found while fixing; its integration test covers the success path offline until this lands).
🤖 Generated by Anthropic Claude (Opus); posted via @cb1kenobi.