Problem
The model registry (resources/models/backendRegistry.ts) is per thread. A component using models.registerBackend on every worker builds a separate in-process model on each worker, repeating weight loading, native context creation and warmup. Backend objects contain closures and native handles and cannot be sent to another thread. Flair's in-process GGUF embedder motivates this API (flair#2052).
Proposed API
Add models.registerProcessBackend(kind, id, factory, options?) beside registerBackend, plus models.backendStatus(kind, id). The factory receives { kind, logicalName, signal }; it may return a ModelBackend or register one during construction. This factory shape avoids constructing a model in every worker before registration.
models.registerProcessBackend(
'embedding',
'local:nomic',
async ({ signal }) => {
const engine = await loadEngine(config, { signal });
await engine.warmup();
return {
...models.defineBackend({ name: 'local:nomic', embed: (input, opts) => engine.embed(input, opts) }),
dispose: () => engine.close(),
};
},
{ concurrency: 1, maxPending: 256, maxBatchInputs: 64, maxRestarts: 1 }
);
await models.embed('hello', { model: 'local:nomic' });
models.backendStatus('embedding', 'local:nomic');
Options: concurrency, maxPending, maxBatchInputs for embeddings, maxRestarts, ownerWaitMs and timeoutMs. The proxy API covers embed, generate, decide and scoreChoices; generateStream is deferred. The per-thread API and YAML bootstrap remain available.
For ownership and disposal, see resources/models/DESIGN.md § “A process-wide backend has one live instance per key at a time for what its runs are handed, and no per-thread fallback.”
For handover, election, readiness, status and failure, see resources/models/DESIGN.md § “Handover and election.”
For routing, resending, trust and load, see the Trust, Resend and Load paragraphs in that document's process-wide backend section.
Test plan
Use real worker threads through startWorker and the production port mesh. Cover shared and single-thread registration; concurrent calls; cancellation; queue bounds; embedding batches, split usage and accounting; owner exit; generation and handover; failed starts; duplicate and invalid registrations; ownership and disposal of returned and registered objects; late registration; trust and domain checks; refusal provenance; method forwarding; structured clone; factory signals; and the unchanged per-thread default. Cover constructBackend capture and guard behavior directly. Focused gaps to add: owner RELEASE with live: false, and DISPOSE_FAILED with thread exit before a newer-generation claim.
How tested in PR #3082
The prior PR text records that, at 48a703df, npx mocha unitTests/resources/models/processBackend.test.js passed 73 tests; the LMDB run of the same file passed 73. npx mocha "unitTests/resources/models/**/*.js" passed 868 with 0 failures. npm run test:unit:resources had 4109 passing, 54 pending and 3 failing; the merge base 4cceb720 had the same 3 failures with 4032 passing and 54 pending. At 4f4a1e3f, focused npx mocha over processBackend.test.js and backendRegistry.test.js passed 105. Prettier, design-doc, lint and type checks were clean at the revisions reported in the PR.
Problem
The model registry (
resources/models/backendRegistry.ts) is per thread. A component usingmodels.registerBackendon every worker builds a separate in-process model on each worker, repeating weight loading, native context creation and warmup. Backend objects contain closures and native handles and cannot be sent to another thread. Flair's in-process GGUF embedder motivates this API (flair#2052).Proposed API
Add
models.registerProcessBackend(kind, id, factory, options?)besideregisterBackend, plusmodels.backendStatus(kind, id). The factory receives{ kind, logicalName, signal }; it may return aModelBackendor register one during construction. This factory shape avoids constructing a model in every worker before registration.Options:
concurrency,maxPending,maxBatchInputsfor embeddings,maxRestarts,ownerWaitMsandtimeoutMs. The proxy API coversembed,generate,decideandscoreChoices;generateStreamis deferred. The per-thread API and YAML bootstrap remain available.For ownership and disposal, see
resources/models/DESIGN.md§ “A process-wide backend has one live instance per key at a time for what its runs are handed, and no per-thread fallback.”For handover, election, readiness, status and failure, see
resources/models/DESIGN.md§ “Handover and election.”For routing, resending, trust and load, see the Trust, Resend and Load paragraphs in that document's process-wide backend section.
Test plan
Use real worker threads through
startWorkerand the production port mesh. Cover shared and single-thread registration; concurrent calls; cancellation; queue bounds; embedding batches, split usage and accounting; owner exit; generation and handover; failed starts; duplicate and invalid registrations; ownership and disposal of returned and registered objects; late registration; trust and domain checks; refusal provenance; method forwarding; structured clone; factory signals; and the unchanged per-thread default. CoverconstructBackendcapture and guard behavior directly. Focused gaps to add: ownerRELEASEwithlive: false, andDISPOSE_FAILEDwith thread exit before a newer-generation claim.How tested in PR #3082
The prior PR text records that, at
48a703df,npx mocha unitTests/resources/models/processBackend.test.jspassed 73 tests; the LMDB run of the same file passed 73.npx mocha "unitTests/resources/models/**/*.js"passed 868 with 0 failures.npm run test:unit:resourceshad 4109 passing, 54 pending and 3 failing; the merge base4cceb720had the same 3 failures with 4032 passing and 54 pending. At4f4a1e3f, focusednpx mochaoverprocessBackend.test.jsandbackendRegistry.test.jspassed 105. Prettier, design-doc, lint and type checks were clean at the revisions reported in the PR.