[pull] master from mudler:master#1150
Merged
Merged
Conversation
…409a56b27ad92f` (#11009) ⬆️ Update PrismML-Eng/llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
…ac94e14fdbebb6bb138c` (#11006) ⬆️ Update ServeurpersoCom/qwentts.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
…23ed9d698a1ff` (#11007) ⬆️ Update CrispStrobe/CrispASR Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
…#11017) backendLoader logged "BackendLoader starting" at INFO as its very first statement, unconditionally. That reads as "a model is being loaded", but backendLoader is not only a load path: in distributed mode Load() deliberately bypasses the local cache and calls backendLoader on every inference request so SmartRouter can re-pick a replica per request. The model is already resident, no process is spawned, and nothing is loaded, yet the banner fires at request rate. On a live cluster this produced ~5 "BackendLoader starting" lines per second for a single embedding model, sustained, starting 22 seconds after the load had already completed. The model was state=loaded with in_flight=0 and exactly one backend process on the worker. It looked exactly like a retry storm and cost real debugging time during an unrelated production investigation. The adjacent "effective runtime tuning" banner, documented as "logged once per load", had the same problem for the same reason. Emit both banners at INFO only when the model is not already resident, and keep the per-call trace at DEBUG for anyone following the routing path. isResident is a plain store lookup with no health probe and no eviction, so it is safe on the per-request hot path (unlike checkIsLoaded, which probes and can evict). Same class of defect as #10985: a log line that sends the reader after the wrong thing. Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The Realtime guide incorrectly sent the Opus backend through the model gallery endpoint. Point users to the backend gallery API and document the UI and CLI alternatives. Assisted-by: Codex:gpt-5 Signed-off-by: Richard Palethorpe <io@richiejp.com>
…ll-clock (#11019) A 70 GB video checkpoint (longcat-video-avatar-1.5) could not be loaded on a distributed cluster. The request failed with HTTP 500 after 1499.98s - exactly the 25m00s cold-load ceiling - while staging was demonstrably healthy: 26 of 57 files and 39 GB transferred at a sustained ~26 MB/s, zero errors, no stalls. It was not wedged, it was killed by a timer. ModelLoadCeilingFor covers node selection, backend install, file staging and the remote LoadModel. Install and load carry their own budgets; staging was covered only by a FIXED 5-minute margin. But staging time is bytes over bandwidth, not a constant: 70 GB at 26 MB/s needs ~45m against a 25m ceiling, so the failure is deterministic for any sufficiently large model rather than a flake. Simply raising the constant moves the cliff to the next model size - the deployment target here is checkpoints of 600 GB and beyond. The ceiling's real purpose is that "a wedged worker can never pin the lock indefinitely". Progress, not elapsed time, is what distinguishes a wedged worker from a large one. The hold is now a deadline that extends whenever the transfer reports bytes and expires a 5-minute stall window after they stop: - A large model transferring fine continues, for hours if needed. - A worker that died mid-transfer still fails within the stall window. Progress is observed at byte level on the transfer itself, via the existing staging progress callback. Per-file completion would be too coarse - a single 600 GB shard would be indistinguishable from a stall for hours. The observation point is back-pressured by the socket, so it reflects the network rather than local disk reads. Observation is coarsened to one timer touch per stall/20 so the per-read callback stays cheap. The base budget (unchanged, and still derived from the install and load timeouts) continues to cover the steps that report no progress, so LOCALAI_NATS_MODEL_LOAD_TIMEOUT keeps working exactly as before. An absolute cap of 24h bounds the hold even while progress keeps arriving, so a peer trickling bytes forever cannot pin the advisory lock; 600 GB at the measured 26 MB/s is ~6.5h, so the cap sits far above any legitimate transfer. Also fixes the incoherent layering the same error exposed: the resumable upload carried a 1h retry budget nested inside the 25m ceiling, so the inner budget was unreachable and the message still blamed it ("failed after 1 attempts within 1h0m0s budget") while the 25m parent was the actual killer. The upload now adopts the caller's deadline when there is one, and applies its fixed budget only when nothing above bounded it - which also stops a fixed 1h from reintroducing the size cliff under the now-extendable parent. This is the successor to #10968, where a hardcoded 5-minute LoadModel gRPC timeout was replaced by this derived ceiling. Fixing the inner timeout exposed the outer ceiling as the new binding constraint. Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
…-ui in the npm_and_yarn group across 1 directory (#11016) chore(deps): bump body-parser Bumps the npm_and_yarn group with 1 update in the /core/http/react-ui directory: [body-parser](https://github.com/expressjs/body-parser). Updates `body-parser` from 2.2.2 to 2.3.0 - [Release notes](https://github.com/expressjs/body-parser/releases) - [Changelog](https://github.com/expressjs/body-parser/blob/master/HISTORY.md) - [Commits](expressjs/body-parser@v2.2.2...v2.3.0) --- updated-dependencies: - dependency-name: body-parser dependency-version: 2.3.0 dependency-type: indirect dependency-group: npm_and_yarn ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )