tag:github.com,2008:https://github.com/mudler/LocalAI/releasesRelease notes from LocalAI2026-08-07T15:51:31Ztag:github.com,2008:Repository/615869301/v4.8.22026-08-07T16:01:57Zv4.8.2
<h2>What's Changed</h2>
<h3>👒 Dependencies</h3>
<ul>
<li>chore(deps): bump actions/stale from 10.4.0 to 11.0.0 by <a class="user-mention notranslate" data-hovercard-type="organization" data-hovercard-url="/orgs/dependabot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dependabot">@dependabot</a>[bot] in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5084207787" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11395" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11395/hovercard" href="https://github.com/mudler/LocalAI/pull/11395">#11395</a></li>
<li>chore(deps): bump actions/checkout from 4 to 7 by <a class="user-mention notranslate" data-hovercard-type="organization" data-hovercard-url="/orgs/dependabot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dependabot">@dependabot</a>[bot] in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5084208384" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11396" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11396/hovercard" href="https://github.com/mudler/LocalAI/pull/11396">#11396</a></li>
</ul>
<h3>Other Changes</h3>
<ul>
<li>feat(gallery): fall back to mirrors and a cached index when the primary source fails by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5079739524" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11389" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11389/hovercard" href="https://github.com/mudler/LocalAI/pull/11389">#11389</a></li>
<li>chore(model-gallery): ⬆️ update checksum by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5085214191" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11405" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11405/hovercard" href="https://github.com/mudler/LocalAI/pull/11405">#11405</a></li>
<li>docs: ⬆️ update docs version mudler/LocalAI by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5085025747" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11397" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11397/hovercard" href="https://github.com/mudler/LocalAI/pull/11397">#11397</a></li>
<li>feat(swagger): update swagger by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5085100052" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11398" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11398/hovercard" href="https://github.com/mudler/LocalAI/pull/11398">#11398</a></li>
<li>feat(gallery): default to index.localai.io with GitHub as a mirror by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5087730220" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11409" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11409/hovercard" href="https://github.com/mudler/LocalAI/pull/11409">#11409</a></li>
<li>feat(nemo-speech-cpp): add the NVIDIA NeMo-Speech.cpp backend by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5085371145" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11406" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11406/hovercard" href="https://github.com/mudler/LocalAI/pull/11406">#11406</a></li>
<li>fix(ci): remove unsupported cosign bundle flag by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5091758656" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11413" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11413/hovercard" href="https://github.com/mudler/LocalAI/pull/11413">#11413</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.8.1...v4.8.2"><tt>v4.8.1...v4.8.2</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.8.12026-08-06T16:01:36Zv4.8.1
<h2>What's Changed</h2>
<h3>Other Changes</h3>
<ul>
<li>docs(blog): cover the terminal agent in the 4.8 post by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5072091582" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11372" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11372/hovercard" href="https://github.com/mudler/LocalAI/pull/11372">#11372</a></li>
<li>fix(vram): contain malformed GGUF metadata by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/richiejp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/richiejp">@richiejp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5073921563" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11374" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11374/hovercard" href="https://github.com/mudler/LocalAI/pull/11374">#11374</a></li>
<li>docs: ⬆️ update docs version mudler/LocalAI by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075079972" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11377" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11377/hovercard" href="https://github.com/mudler/LocalAI/pull/11377">#11377</a></li>
<li>chore: ⬆️ Update 0xShug0/audio.cpp to <code>7efbb58def443722ea540d931dd3debee3e4d5e8</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075114493" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11378" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11378/hovercard" href="https://github.com/mudler/LocalAI/pull/11378">#11378</a></li>
<li>chore: ⬆️ Update ikawrakow/ik_llama.cpp to <code>cf1aa57e1a0fabfd015831718fc99d1aec01ada5</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075278539" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11380" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11380/hovercard" href="https://github.com/mudler/LocalAI/pull/11380">#11380</a></li>
<li>chore: ⬆️ Update leejet/stable-diffusion.cpp to <code>c6beeef35526c6dc94b74a7fb69f9d2e6a2a7a12</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075600737" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11384" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11384/hovercard" href="https://github.com/mudler/LocalAI/pull/11384">#11384</a></li>
<li>chore: ⬆️ Update CrispStrobe/CrispASR to <code>21901d3f7c23554f072964828363e49ddbc2dc68</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075471621" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11383" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11383/hovercard" href="https://github.com/mudler/LocalAI/pull/11383">#11383</a></li>
<li>fix(vllm-cpp): mirror the engine's ABI v10 so the backend loads again by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075910394" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11386" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11386/hovercard" href="https://github.com/mudler/LocalAI/pull/11386">#11386</a></li>
<li>fix(react-ui): stop traces page crash when switching trace tabs by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nandanadileep/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nandanadileep">@nandanadileep</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5077535181" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11387" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11387/hovercard" href="https://github.com/mudler/LocalAI/pull/11387">#11387</a></li>
<li>chore: ⬆️ Update antirez/ds4 to <code>b0309611041655f4e45671cfd9c9886aff161406</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075325779" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11381" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11381/hovercard" href="https://github.com/mudler/LocalAI/pull/11381">#11381</a></li>
<li>chore(model-gallery): ⬆️ update checksum by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5075367622" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11382" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11382/hovercard" href="https://github.com/mudler/LocalAI/pull/11382">#11382</a></li>
<li>gallery: add Qwen3.5 9B HauhauCS variants by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-org-maint-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-org-maint-bot">@localai-org-maint-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5056582666" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11339" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11339/hovercard" href="https://github.com/mudler/LocalAI/pull/11339">#11339</a></li>
<li>gallery: add Qwen3.5 9B Defiant Fable variants by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-org-maint-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-org-maint-bot">@localai-org-maint-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5055281196" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11335" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11335/hovercard" href="https://github.com/mudler/LocalAI/pull/11335">#11335</a></li>
<li>feat(vllm-cpp): wire the full engine config surface through engine_args by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4996557962" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11159" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11159/hovercard" href="https://github.com/mudler/LocalAI/pull/11159">#11159</a></li>
<li>fix(react-ui): restore 3D Studio results and history by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/richiejp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/richiejp">@richiejp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5082737271" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11393" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11393/hovercard" href="https://github.com/mudler/LocalAI/pull/11393">#11393</a></li>
<li>fix(cli): ignore a half-populated socket activation environment by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5082797009" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11394" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11394/hovercard" href="https://github.com/mudler/LocalAI/pull/11394">#11394</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.8.0...v4.8.1"><tt>v4.8.0...v4.8.1</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.8.02026-08-05T10:32:06Zv4.8.0<h1>🎉 LocalAI 4.8.0 Release! 🚀</h1>
<h1 align="center">
<br>
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/refs/heads/master/core/http/static/logo.png"><img height="300" src="https://raw.githubusercontent.com/mudler/LocalAI/refs/heads/master/core/http/static/logo.png" style="max-width: 100%; height: auto; max-height: 300px;"></a>
<br>
<br>
</h1>
<p>LocalAI 4.8.0 is out!</p>
<p>Twenty-two days, 386 pull requests, and three new modalities. This release introduces <strong>vllm.cpp</strong>, a C++20 engine maintained by the LocalAI team, which began as a vLLM port and now carries its own featureset, shipping as the <code>vllm-cpp</code> backend in <strong>alpha development builds</strong>. Around it: <strong>3D generation</strong> as a new modality, a multi-family <strong>audio.cpp</strong> engine, gallery entries that install the build your hardware can actually run, and a deep reliability pass on distributed mode driven by production incidents.</p>
<p><strong>Highlights:</strong></p>
<ul>
<li>🚀 <strong>vllm.cpp (alpha)</strong> - a C++20 engine maintained by the LocalAI team, which began as a vLLM port and keeps vLLM as its reference implementation: V1 serving architecture (paged KV cache, continuous batching, prefix caching, scheduler, sampler) with no Python, PyTorch or ggml at inference. Measured at 1.045x vLLM on Qwen3.6-27B NVFP4 at concurrency 1, with token-for-token identical output. Loads safetensors and GGUF, enforces structured output in-engine, and runs on CPU, CUDA, Metal and Vulkan. The Apple Silicon build ships the MLX GEMM provider, measured at 1.5x to 2.2x on an M4. Shipping as alpha development builds: try it, do not depend on it.</li>
<li>🧊 <strong>3D generation</strong> - a new modality end to end: <code>Generate3D</code> RPC, <code>FLAG_3D</code> capability, <code>POST /v1/3d/generations</code>, the <code>trellis2cpp</code> image-to-3D backend, and a UI page with a native GLB viewer and print remeshing.</li>
<li>🔊 <strong>audio.cpp</strong> - one backend process serving six audio endpoints across many model families, picked from the GGUF's own metadata: speech, transcription, VAD, diarization, source separation and sound generation.</li>
<li>🎛️ <strong>One model, many builds</strong> - a gallery entry can declare <code>variants:</code>, and LocalAI installs the largest build that your host can actually run. No more hunting through the gallery for the right quantization.</li>
<li>⚡ <strong>A much lighter web UI</strong> - gzip on the wire, immutable caching for hashed assets, and paginated trace endpoints: the React bundle is 3.48x smaller and the trace poll dropped from 21 MB to 7 KB.</li>
<li>📊 <strong>An Activity page</strong> - the stacked operations bar collapses to one line, and a new admin Activity page keeps the record of what installed, failed or was cancelled, instead of dropping it the moment it finished.</li>
<li>📦 <strong>Hugging Face artifact materialization</strong> - immutable snapshot resolution, authenticated downloads with real progress, and staged artifacts that remote workers can bind to.</li>
<li>🎚️ <strong>VRAM budgets</strong> - cap how much of a card LocalAI may use, per node, as a percentage (<code>80%</code>) or an absolute amount (<code>12GB</code>).</li>
<li>🗣️ <strong>Two new TTS engines</strong> - <code>magpie-tts-cpp</code> (NVIDIA Magpie Multilingual, 5 voices, 9+ languages) and <code>moss-tts-cpp</code> (48 kHz stereo with reference-audio voice cloning).</li>
<li>🌳 <strong>Sub-2-bit models</strong> - a new <code>bonsai</code> backend serves the 1-bit and ternary Bonsai quantizations of Qwen3 and Qwen3.6-27B.</li>
<li>🖧 <strong>Distributed mode hardening</strong> - the reaper no longer deletes rows for backends that are alive and busy, phantom replicas are cleaned up, and <code>in_flight</code> counters stop leaking.</li>
</ul>
<p>Plus a Valkey vector store, systemd socket activation, persistent trace history, two security fixes, a documentation overhaul aimed squarely at onboarding, and a new localai.io.</p>
<p align="center">
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-ui-home.png"><img src="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-ui-home.png" alt="The LocalAI Home page" width="900" style="max-width: 100%;"></a>
</p>
<hr>
<h2>📊 This release in numbers</h2>
<table>
<thead>
<tr>
<th></th>
<th></th>
</tr>
</thead>
<tbody>
<tr>
<td>Pull requests merged</td>
<td><strong>386</strong></td>
</tr>
<tr>
<td>Commits</td>
<td>392</td>
</tr>
<tr>
<td>Files changed</td>
<td>1,204 (<strong>+152,918</strong> / -42,012)</td>
</tr>
<tr>
<td>Development window</td>
<td>22 days (2026-07-14 to 2026-08-05)</td>
</tr>
<tr>
<td>Contributors</td>
<td>25, of whom <strong>11 first-time</strong></td>
</tr>
<tr>
<td>New backends</td>
<td><strong>7</strong> (<code>vllm-cpp</code>, <code>audio-cpp</code>, <code>trellis2cpp</code>, <code>valkey-store</code>, <code>bonsai</code>, <code>magpie-tts-cpp</code>, <code>moss-tts-cpp</code>)</td>
</tr>
<tr>
<td>Gallery entries</td>
<td>1,221 to <strong>1,515</strong> (+294)</td>
</tr>
</tbody>
</table>
<p>Where the work landed:</p>
<table>
<thead>
<tr>
<th>Area</th>
<th>Change</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>core/</code></td>
<td>+54,974 / -8,291 across 546 files</td>
</tr>
<tr>
<td><code>gallery/</code></td>
<td>+39,193 / -29,119 (variant ladders rewrite most of the index)</td>
</tr>
<tr>
<td><code>backend/</code></td>
<td>+25,668 / -1,305 across 256 files</td>
</tr>
<tr>
<td><code>pkg/</code></td>
<td>+9,678 / -385</td>
</tr>
<tr>
<td><code>.github/</code></td>
<td>+7,526 / -61</td>
</tr>
<tr>
<td><code>docs/</code></td>
<td>+4,285 / -2,461 (near-flat by design: the dedup pass removed as much as it added)</td>
</tr>
<tr>
<td><code>website/</code></td>
<td>+4,479 (new project site)</td>
</tr>
<tr>
<td><code>scripts/</code></td>
<td>+3,373 / -175</td>
</tr>
</tbody>
</table>
<hr>
<h2>📌 TL;DR</h2>
<table>
<thead>
<tr>
<th>Area</th>
<th>Summary</th>
</tr>
</thead>
<tbody>
<tr>
<td>🚀 <strong>vllm.cpp</strong> (alpha)</td>
<td>An Apache-2.0 C++20 engine maintained by the LocalAI team and developed in its own repository, which began as a vLLM port, shipping as <strong>alpha development builds</strong> of the <code>vllm-cpp</code> backend over its stable C ABI v5. It uses vLLM as its reference implementation and benchmark, and implements vLLM's V1 architecture (paged KV cache, continuous batching, prefix caching, scheduler, sampler) with no Python, PyTorch or ggml at inference. Safetensors + GGUF, in-engine structured output (JSON schema / regex / choice / GBNF). Chat and tool calling ride the llama.cpp autoparser path: full minja templates, <code>tool_choice: auto</code> as a lazy structural-tag constraint, 30 tool dialects and 7 reasoning parsers, streamed <code>ChatDelta</code>/<code>ToolCallDelta</code>. CPU amd64/arm64, CUDA 12/13 (Blackwell), L4T, Vulkan and Darwin Metal, the last with the MLX GEMM provider vendored in (1.5x to 2.2x on an M4).</td>
</tr>
<tr>
<td>🧊 <strong>3D generation</strong></td>
<td>A new modality, wired end to end: <code>Generate3D</code> RPC, <code>FLAG_3D</code> capability, <code>POST /v1/3d/generations</code>, the <code>trellis2cpp</code> image-to-3D backend over TRELLIS.2, and a UI page with a native GLB viewer, IndexedDB history and previewable print remeshing.</td>
</tr>
<tr>
<td>🔊 <strong>audio.cpp</strong></td>
<td>New native C++ backend over <a href="https://github.com/0xShug0/audio.cpp">audio.cpp</a>, a multi-family ggml audio engine: one process serves <code>/v1/audio/speech</code> (supertonic, chatterbox, irodori-voicedesign), <code>/v1/audio/transcriptions</code> (citrinet, nemotron, forced-aligner), <code>/v1/audio/vad</code>, <code>/v1/audio/diarize</code> (sortformer), <code>/audio/transform</code> (htdemucs 4-stem separation, voice conversion, speech-to-speech) and <code>/v1/sound-generation</code>. Family comes from the GGUF's own <code>audiocpp.model_spec.family</code> key, so no per-model backend options. 13 gallery entries. CPU, CUDA 12/13, Vulkan, Metal.</td>
</tr>
<tr>
<td>📊 <strong>Activity page</strong></td>
<td>The stacked operations bar becomes a permanent one-line strip (<code>✕</code> now hides rather than cancels), with a new admin <code>/app/activity</code> page: in-progress detail with per-node breakdown, a "needs attention" lane with Cancel and Retry, and a bounded 50-entry record of what finished.</td>
</tr>
<tr>
<td>🗄️ <strong>Valkey vector store</strong></td>
<td>New <code>valkey-store</code> backend adding Valkey Search as a vector store option.</td>
</tr>
<tr>
<td>🌐 <strong>New localai.io</strong></td>
<td>The site splits into a project site at the root and docs under <code>/docs/</code>, with 214 generated redirect stubs so every published URL keeps working. Adds an engines page driven by YAML, a blog, an ecosystem band and <code>ADOPTERS.md</code>.</td>
</tr>
<tr>
<td>🎛️ <strong>Gallery variants</strong></td>
<td>An entry may declare <code>variants:</code> referencing other entries. Install-time selection drops builds the host cannot run (<code>IsBackendCompatible</code>) or cannot fit (VRAM, or cgroup-aware RAM on CPU hosts), then picks the largest that fits. Override with <code>variant</code> on <code>POST /models/apply</code>, <code>local-ai models install --variant</code>, the <code>install_model</code> MCP tool, or the UI split-button. <code>GET /api/models?has_variants=true</code> narrows the list. Older clients ignore the key and install the entry as before.</td>
</tr>
<tr>
<td>📦 <strong>HF artifacts</strong></td>
<td>Immutable snapshot resolution, authenticated downloads with progress, gallery install and preload materialization, runtime binding to staged artifacts, and UI progress reporting. Python backends reuse the Go download path.</td>
</tr>
<tr>
<td>⚡ <strong>HTTP performance</strong></td>
<td>gzip middleware (<code>--disable-http-compression</code>, <code>--http-compression-min-length</code>), with streaming paths explicitly skipped. <code>/assets/*</code> served <code>immutable</code>, <code>index.html</code> <code>no-cache</code>. <code>/api/traces</code> and <code>/api/backend-traces</code> accept <code>limit</code>/<code>offset</code>/<code>full</code> and summarize by default, with <code>GET /api/traces/{id}</code> for the full record. React bundle 2,815,513 B to 807,918 B; backend-trace poll 21,131,097 B to 7,201 B.</td>
</tr>
<tr>
<td>🎚️ <strong>VRAM budget</strong></td>
<td><code>LOCALAI_VRAM_BUDGET=80%</code> or <code>=12GB</code> (also <code>--vram-budget</code>), on <code>local-ai</code> and <code>local-ai worker</code>. Standalone it is a hard per-process cap inherited by context-fit, GGUF warnings and the watchdog; distributed it is a placement ceiling the scheduler respects. Admin override via <code>PUT</code>/<code>DELETE /api/nodes/:id/vram-budget</code> and the <code>set_node_vram_budget</code> MCP tool. Unset means all detected VRAM.</td>
</tr>
<tr>
<td>🗣️ <strong>magpie-tts-cpp</strong></td>
<td>New Go/purego backend over <a href="https://github.com/mudler/magpie-tts.cpp">magpie-tts.cpp</a>, a ggml port of NVIDIA Magpie TTS Multilingual 357M with NanoCodec embedded. 5 voices, 9+ languages, 22.05 kHz mono, one self-contained GGUF.</td>
</tr>
<tr>
<td>🗣️ <strong>moss-tts-cpp</strong></td>
<td>New Go/purego backend over <a href="https://github.com/mudler/moss-tts.cpp">moss-tts.cpp</a> for MOSS-TTS-Local v1.5. 48 kHz stereo, optional reference-audio voice cloning, no Python at inference.</td>
</tr>
<tr>
<td>🌳 <strong>bonsai</strong></td>
<td>New backend on the <a href="https://github.com/PrismML-Eng/llama.cpp">PrismML llama.cpp fork</a>, which is the only decoder for the Q1_0 and Q2_0 quant formats. Eight gallery entries across Bonsai 8B/27B and Ternary-Bonsai 8B/27B, from ~1.15 GB.</td>
</tr>
<tr>
<td>🖧 <strong>Distributed reliability</strong></td>
<td>A busy backend is no longer reaped: the worker is asked directly over a new <code>models.running</code> subject, and the port-probe fallback distinguishes <code>DeadlineExceeded</code> (busy) from <code>Unavailable</code> (gone), requiring three consecutive misses. Frontend model stubs are dropped when no healthy replica remains, <code>in_flight</code> leaks are closed, and model-load deadlines scale with checkpoint size.</td>
</tr>
<tr>
<td>🛡️ <strong>Security</strong></td>
<td>Inline GRPO reward code in <code>POST /api/fine-tuning/jobs</code> is refused unless the operator sets <code>LOCALAI_TRL_ALLOW_INLINE_REWARD=true</code>; the previous builtin allowlist was escapable to arbitrary code execution on an endpoint that is unauthenticated by default. Also picks up hono 4.12.25 for <a title="CVE-2026-54290" data-hovercard-type="advisory" data-hovercard-url="/advisories/GHSA-88fw-hqm2-52qc/hovercard" href="https://github.com/advisories/GHSA-88fw-hqm2-52qc">CVE-2026-54290</a>.</td>
</tr>
<tr>
<td>🧠 <strong>Models</strong></td>
<td>MiniMax-M3, Gemma 4 llama.cpp MTP variants, Qwen3.5-4B DFlash, MOSS-TTS-Local v1.5, the APEX families as variant ladders, and the Bonsai families. Duplicate entries removed and linted against recurring.</td>
</tr>
<tr>
<td>📖 <strong>Docs</strong></td>
<td>Onboarding overhaul: one model carried through install to first API call, a new "Build your first agent" walkthrough, a runtime-errors reference keyed on literal error strings, an agent actions catalog, and a new Operations section.</td>
</tr>
</tbody>
</table>
<hr>
<h2>🚀 New Features & Major Enhancements</h2>
<h3>🚀 Introducing vllm.cpp (alpha)</h3>
<p align="center">
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/vllm.cpp/master/benchmarks/media/concurrency_race.gif"><img src="https://raw.githubusercontent.com/mudler/vllm.cpp/master/benchmarks/media/concurrency_race.gif" alt="vllm.cpp and vLLM generating side by side on Qwen3.6-27B" width="900" style="max-width: 100%;"></a>
<br><em>vllm.cpp against vLLM on Qwen3.6-27B, identical output at every concurrency.</em>
</p>
<p><strong><a href="https://github.com/mudler/vllm.cpp">vllm.cpp</a> is Apache-2.0, maintained by the LocalAI team, and began as a C++20 port of vLLM.</strong> We want it community-first rather than a LocalAI-only engine, so it lives in its own repository with its own docs, benchmark record and issue tracker, and it is usable without LocalAI anywhere in the picture. It implements vLLM's V1 serving architecture (paged KV cache, continuous batching, prefix caching, scheduler, sampler) on a portable tensor runtime with <strong>no Python, no PyTorch and no ggml at inference time</strong>, and uses vLLM itself as its reference implementation: correctness is checked by comparing output against it, and the benchmark scoreboard is kept against it.</p>
<p>It has since grown a featureset vLLM does not have, which is what the port was for. It loads GGUF as well as Hugging Face safetensors, runs on CPU, Apple Metal and Vulkan alongside NVIDIA CUDA, ships speculative decoding and KV offload, and enforces structured output in-engine (JSON schema, regex, choice, GBNF). Its benchmark page now measures against llama.cpp, MLX-LM and DwarfStar as well as vLLM, because those are the engines it actually competes with on that hardware.</p>
<p><strong>The project is expected to be renamed</strong>, with the new name still to be decided. It is drifting far enough from vLLM that calling it a port undersells it and calling it vllm.cpp will eventually mislead.</p>
<h3>Numbers, from the project's own scoreboard</h3>
<p align="center">
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-vllm-cpp-scoreboard.png"><img src="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-vllm-cpp-scoreboard.png" alt="Throughput of vllm.cpp relative to each reference engine" width="900" style="max-width: 100%;"></a>
</p>
<p>These come from <a href="https://github.com/mudler/vllm.cpp/blob/master/docs/BENCHMARKS.md">vllm.cpp's BENCHMARKS.md</a>, which reports ties as ties and losses as losses. Throughput is vllm.cpp over the reference, so above 1.0 is ahead.</p>
<table>
<thead>
<tr>
<th>Reference</th>
<th>Workload</th>
<th>Result</th>
</tr>
</thead>
<tbody>
<tr>
<td>vLLM</td>
<td>Qwen3.6-27B NVFP4, GB10</td>
<td>1.045x at concurrency 1, 1.007x to 1.017x at c2 to c32, output token-for-token identical</td>
</tr>
<tr>
<td>vLLM</td>
<td>Qwen3.6-35B-A3B NVFP4, GB10</td>
<td>1.010x at c16 and 1.013x at c32; behind at c1 to c8 (0.817x at c1)</td>
</tr>
<tr>
<td>vLLM</td>
<td>DeepSeek-V2-Lite MLA, GB10</td>
<td>0.86x to 0.95x throughput, TTFT ahead at c4 and c8</td>
</tr>
<tr>
<td>llama.cpp</td>
<td>Qwen3.5-2B GGUF, CPU aarch64</td>
<td>prefill 1.18x, decode a tie, memory parity, byte-identical output</td>
</tr>
<tr>
<td>MLX-LM</td>
<td>Qwen3-0.6B, Apple M4</td>
<td>97.6% of warm total, prefill ahead</td>
</tr>
<tr>
<td>DwarfStar (ds4)</td>
<td>DeepSeek-V4-Flash IQ2_XXS, one DGX Spark</td>
<td>18.69 vs 16.33 tok/s decode, <strong>1.144x</strong>, same output</td>
</tr>
<tr>
<td>vLLM</td>
<td>Laguna-XS-2.1 NVFP4, GB10</td>
<td>44.46 vs 43.10 tok/s, <strong>1.03x</strong>, same output</td>
</tr>
</tbody>
</table>
<p>The upstream page is careful about its own noise band: on the 27B grid it calls c2 through c32 ties rather than wins, because the run-to-run spread is 0.5% and those margins land between 0.7% and 1.7%. The c1 result is the one it stands behind.</p>
<p>The DeepSeek-V4-Flash row is the one that shows how far the project has moved from being a vLLM port. It runs DeepSeek-V4-Flash at roughly 2-bit (IQ2_XXS mixed, about 80 GB) on a <strong>single DGX Spark</strong>, decoding at 18.69 tok/s against DwarfStar's 16.33. At 300B+ total parameters even a 4-bit checkpoint is 156 GB or more, so a 2-bit GGUF is what fits inside the Spark's 119 GiB unified pool, and reading GGUF is what makes that possible.</p>
<p>That figure moved twice in a week, and the second move came from one lever. The dense Q8_0 projection tower was being read from the GGUF mmap over unified memory, which the GB10 reads about 20% slower per-GEMV than device memory. Staging that ~6 GiB tower device-resident once at load, same bytes and same kernels, took decode from 16.23 to 18.69, generating the same tokens and using no more peak memory. The same change took Laguna-XS-2.1 from 87% of vLLM to 1.03x ahead of it.</p>
<p>Speculative decoding is in similar shape: MTP on Qwen3.6-27B NVFP4 generates the same tokens as vLLM's MTP and runs about 4% faster at concurrency 1.</p>
<p>It ships here as the <strong><code>vllm-cpp</code></strong> backend, which dlopens the engine's stable C ABI (v5) through purego. Concurrent requests batch continuously inside the engine's shared scheduler rather than serializing, so the backend runs on <code>base.Base</code> rather than <code>SingleThread</code>.</p>
<p><strong>Tool calling is at llama.cpp parity, by construction</strong>, because chat reuses the same autoparser path. With <code>use_tokenizer_template</code> the engine renders the model's own chat template (GGUF <code>tokenizer.chat_template</code> or <code>tokenizer_config.json</code>, full minja) and handles the rest itself:</p>
<ul>
<li><code>tool_choice: auto</code> lowers to a lazy structural-tag decode constraint; <code>required</code> and named choices force the family's native syntax where expressible.</li>
<li>Streaming per-dialect parsers cover <strong>30 tool dialects and 7 reasoning parsers</strong>, with <code><think></code> reasoning split before tool parsing.</li>
<li><code>ChatDelta</code>, <code>ToolCallDelta</code> and reasoning stream exactly as the llama-cpp backend does.</li>
</ul>
<p><code>tool_parser:</code> and <code>reasoning_parser:</code> are model options, auto-detected when unset.</p>
<p>Getting started is a normal backend install:</p>
<div class="highlight highlight-source-yaml notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="name: qwen3-vllm
backend: vllm-cpp
context_size: 8192
parameters:
model: Qwen3-4B # a safetensors directory or a .gguf file
options:
- max_num_seqs:16 # also: block_size:<n>, num_blocks:<n>"><pre><span class="pl-ent">name</span>: <span class="pl-s">qwen3-vllm</span>
<span class="pl-ent">backend</span>: <span class="pl-s">vllm-cpp</span>
<span class="pl-ent">context_size</span>: <span class="pl-c1">8192</span>
<span class="pl-ent">parameters</span>:
<span class="pl-ent">model</span>: <span class="pl-s">Qwen3-4B </span><span class="pl-c"><span class="pl-c">#</span> a safetensors directory or a .gguf file</span>
<span class="pl-ent">options</span>:
- <span class="pl-s">max_num_seqs:16 </span><span class="pl-c"><span class="pl-c">#</span> also: block_size:<n>, num_blocks:<n></span></pre></div>
<p>The build matrix covers CPU amd64/arm64, CUDA 12/13 (including Blackwell <code>120a;121a</code>), L4T arm64 for GB10, Vulkan and Darwin Metal, with a gallery meta plus 12 image entries. The llama-cpp GGUF and vllm safetensors importers gained preference swaps, so the backend can be chosen at import time.</p>
<p><strong>Apple Silicon gets the MLX GEMM provider.</strong> The darwin build vendors vllm.cpp's optional MLX backend, which upstream keeps off by default on the position that it has to earn its ~124 MB. Measured on an M4 (Qwen3-1.7B-bf16, p=512 g=128, arms toggled on one binary so there is no build-difference confound):</p>
<table>
<thead>
<tr>
<th align="right">Batch</th>
<th align="right">MLX tok/s</th>
<th align="right">native tok/s</th>
<th align="right">speedup</th>
<th align="right">MLX TTFT</th>
<th align="right">native TTFT</th>
</tr>
</thead>
<tbody>
<tr>
<td align="right">1</td>
<td align="right">5.79</td>
<td align="right">3.08</td>
<td align="right"><strong>1.88x</strong></td>
<td align="right">3.32 s</td>
<td align="right">7.68 s</td>
</tr>
<tr>
<td align="right">4</td>
<td align="right">15.75</td>
<td align="right">10.24</td>
<td align="right"><strong>1.54x</strong></td>
<td align="right">9.63 s</td>
<td align="right">18.77 s</td>
</tr>
<tr>
<td align="right">16</td>
<td align="right">38.65</td>
<td align="right">17.69</td>
<td align="right"><strong>2.19x</strong></td>
<td align="right">18.33 s</td>
<td align="right">54.48 s</td>
</tr>
</tbody>
</table>
<p>Read those as indicative rather than binding: two reps with a spread reaching 9.4%, so the multipliers carry about +/-10%. The gap is far larger than the noise, and time-to-first-token roughly halves across the range.</p>
<p><strong>These are alpha development builds, not a released backend.</strong> vllm.cpp is early. It ships in 4.8 so people who want to try it can, not because it is ready for anything you depend on, and <code>llama-cpp</code> stays the default for real use. Expect rough edges.</p>
<p>The CPU path is end-to-end verified against <code>Qwen3.5-2B-UD-Q8_K_XL.gguf</code> with the full Ginkgo suite: blocking and streaming byte-parity, greedy determinism, stop words, GBNF-constrained generation, concurrent streams, real template rendering, reasoning split, a <code>required</code> tool call returning schema-valid arguments, and an <code>auto</code> run where the engine engages the tool itself and streams parsed deltas. The GPU images build and ship, but their runtime behavior has not been through that gate. No throughput comparison against upstream vLLM is claimed. Please report what breaks.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4967928216" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11100" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11100/hovercard" href="https://github.com/mudler/LocalAI/pull/11100">#11100</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4982987263" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11137" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11137/hovercard" href="https://github.com/mudler/LocalAI/pull/11137">#11137</a></p>
</blockquote>
<h3>🧊 3D generation, end to end</h3>
<p align="center">
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/3d-generation.gif"><img src="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/3d-generation.gif" alt="A generated 3D llama turning in the GLB viewer" width="760" style="max-width: 100%;"></a>
<br><em>trellis2-4b, 2,502,928 vertices, turning in the browser.</em>
</p>
<p>LocalAI gains a new modality. Image-to-3D is wired through the whole stack rather than bolted onto an existing endpoint: a <code>Generate3D</code> RPC in <code>backend.proto</code>, a <code>FLAG_3D</code> capability so the loader knows which backends can serve it, and <code>POST /v1/3d/generations</code>.</p>
<p>The first engine behind it is <strong><code>trellis2cpp</code></strong>, a native image-to-3D backend over TRELLIS.2. The React UI gets a 3D generation page with a native GLB viewer, IndexedDB-backed history so your generations survive a reload, and previewable print remeshing for output you intend to actually print.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4929796935" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10979" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10979/hovercard" href="https://github.com/mudler/LocalAI/pull/10979">#10979</a></p>
</blockquote>
<h3>🔊 audio.cpp: one backend, six audio endpoints</h3>
<p><strong><code>audio-cpp</code></strong> wraps <a href="https://github.com/0xShug0/audio.cpp">audio.cpp</a>, a multi-family ggml audio engine. Rather than one backend per model family, a single backend process serves several unrelated families through one runtime vocabulary, and picks the family from the GGUF's own <code>audiocpp.model_spec.family</code> metadata key, so a model needs no backend-specific options to load.</p>
<table>
<thead>
<tr>
<th>Endpoint</th>
<th>Families</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>/v1/audio/speech</code> (batch + streaming)</td>
<td>supertonic, chatterbox (voice cloning), irodori-voicedesign (voice design via <code>instructions</code>)</td>
</tr>
<tr>
<td><code>/v1/audio/transcriptions</code> (batch, streaming, live)</td>
<td>citrinet, nemotron, forced-aligner</td>
</tr>
<tr>
<td><code>/v1/audio/vad</code></td>
<td>silero-vad, marblenet-vad</td>
</tr>
<tr>
<td><code>/v1/audio/diarize</code></td>
<td>sortformer</td>
</tr>
<tr>
<td><code>/audio/transform</code></td>
<td>htdemucs (4-stem separation), chatterbox (voice conversion), seedvc-singing, vevo2 (speech to speech)</td>
</tr>
<tr>
<td><code>/v1/sound-generation</code></td>
<td>stable-audio-sfx</td>
</tr>
</tbody>
</table>
<p>Thirteen gallery entries ship with it, one representative model per task kind the engine can actually serve. Where it cannot honestly back an RPC it returns <code>UNIMPLEMENTED</code> with a reason rather than an empty success, and a failed load is a gRPC error rather than <code>success: false</code>, so the loader's greedy backend probe never silently selects it for a model it cannot serve.</p>
<p>Two changes reach beyond the backend. <code>backend.proto</code> gains <code>AudioTransformStem</code> and <code>AudioTransformResult.stems</code>, so source separation can return the whole stem set instead of a single mixdown. And <code>/audio/transform</code> no longer hardcodes a 16 kHz mono fold: that fold made 4-stem separation unreachable by construction, so it became a per-backend capability, with existing backends keeping it explicitly and the default for an unregistered backend being to leave the upload alone.</p>
<p>Platforms: CPU (amd64 and arm64), CUDA 12, CUDA 13 and Vulkan on Linux, plus Metal on darwin-arm64. No ROCm, which upstream does not support.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4986104108" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11141" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11141/hovercard" href="https://github.com/mudler/LocalAI/pull/11141">#11141</a></p>
</blockquote>
<h3>📊 A one-line strip, and an Activity page</h3>
<p align="center">
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-ui-activity.png"><img src="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-ui-activity.png" alt="The Activity page with four installs running" width="900" style="max-width: 100%;"></a>
<br><em>Four backend installs in flight, with the record of what already finished.</em>
</p>
<p>The operations bar rendered one row per in-flight operation above every page. Queue four model installs and a backend and it took most of the viewport, on every route, until the last one finished. Two things were conflated: a global "something is happening" signal, which needs one line, and the detail of what is happening, which needs a page.</p>
<p>The strip now collapses to a single line permanently, showing one operation (a failure first, otherwise the least-advanced running one) with a <code>+N more</code> pill. Its <code>✕</code> <strong>hides the strip and never cancels</strong>, a deliberate change: the same glyph previously cancelled a 17 GB download in one row and dismissed a message in the next. Cancelling moved to the page, behind a labelled button.</p>
<p>The new admin-only Activity page at <code>/app/activity</code> carries the detail the strip has to drop (phase, bytes, derived time remaining, and a per-node breakdown for cluster installs), a "needs attention" lane for unacknowledged failures with Cancel and Retry, and a record of what finished. That record is a bounded 50-entry ring, which closes a real gap: <code>/api/operations</code> dropped an operation the moment it succeeded, so a user who stepped away had no way to learn whether an install finished, failed, or never started.</p>
<p>Several latent UI bugs were fixed along the way: retrying a failed removal re-downloaded the model, queued operations rendered as "Installing" with a spinner, a long error message pushed every page ~270px past the viewport, and the ETA blanked for every operation whenever one was verifying.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4999042321" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11163" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11163/hovercard" href="https://github.com/mudler/LocalAI/pull/11163">#11163</a></p>
</blockquote>
<h3>🎛️ One gallery entry, several builds</h3>
<p align="center">
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-ui-model-variants.png"><img src="https://raw.githubusercontent.com/mudler/LocalAI/master/website/static/media/v4-8-0-ui-model-variants.png" alt="The model detail pane listing every variant" width="900" style="max-width: 100%;"></a>
<br><em>One entry, four builds. LocalAI picks the largest that fits and marks it auto-selected.</em>
</p>
<p>A gallery entry can now declare <code>variants:</code>, a list of references to other gallery entries that are alternative builds of the same weights:</p>
<div class="highlight highlight-source-yaml notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="- name: nanbeige4.1-3b-q4 # still a normal, complete, installable entry
url: github:mudler/LocalAI/gallery/nanbeige4.1.yaml@master
overrides: {parameters: {model: nanbeige4.1-3b-q4_k_m.gguf}}
files: [...]
variants:
- model: nanbeige4.1-3b-q8"><pre>- <span class="pl-ent">name</span>: <span class="pl-s">nanbeige4.1-3b-q4 </span><span class="pl-c"><span class="pl-c">#</span> still a normal, complete, installable entry</span>
<span class="pl-ent">url</span>: <span class="pl-s">github:mudler/LocalAI/gallery/nanbeige4.1.yaml@master</span>
<span class="pl-ent">overrides</span>: <span class="pl-s">{parameters: {model: nanbeige4.1-3b-q4_k_m.gguf}}</span>
<span class="pl-ent">files</span>: <span class="pl-s">[...]</span>
<span class="pl-ent">variants</span>:
- <span class="pl-ent">model</span>: <span class="pl-s">nanbeige4.1-3b-q8</span></pre></div>
<p>Selection at install time, in order:</p>
<ol>
<li>Drop variants whose backend cannot run here. MLX disappears on Linux, CUDA on a Mac. Derived from the backend name, so authors never write hardware conditions.</li>
<li>Drop what does not fit: VRAM on GPU hosts, cgroup-aware system RAM on CPU hosts, so a container sees its own limit.</li>
<li>Take the largest that remains, on the basis that a bigger footprint is a better build of the same weights.</li>
</ol>
<p>The entry's own build competes in that ranking and is never filtered out, so selection always ends with something installable.</p>
<p>Sizes come from the existing <code>pkg/vram</code> estimator (remote GGUF header, HTTP <code>HEAD</code>, the declared <code>size:</code>, then the HF repo listing). Nothing is downloaded to decide, and a probe failure never fails an install.</p>
<p>Auto-selection is the default and every surface can override it: <code>variant</code> on <code>POST /models/apply</code> and <code>POST /api/models/install/:id</code>, <code>local-ai models install <name> --variant <variant></code>, the <code>variant</code> parameter on the <code>install_model</code> MCP tool, and a split-button menu in the models table. An explicit selection is honored even when it does not fit, with a warning, since that is a deliberate operator override.</p>
<p>Existing installations are unaffected: every released LocalAI reads <code>gallery/index.yaml</code> live and ignores keys it does not understand, so an older client drops <code>variants:</code> and installs the entry exactly as before. A spec re-parses the real index through a legacy-shaped struct to keep that true.</p>
<p>Known gaps worth stating: in distributed mode <code>InstallModel</code> resolves against the frontend rather than the worker that will serve the model, so a cluster with a small frontend and large workers selects conservatively. Probing within a single entry is still serial and uncapped.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4920773310" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10943" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10943/hovercard" href="https://github.com/mudler/LocalAI/pull/10943">#10943</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931263769" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10983" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10983/hovercard" href="https://github.com/mudler/LocalAI/pull/10983">#10983</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4932759425" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10992" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10992/hovercard" href="https://github.com/mudler/LocalAI/pull/10992">#10992</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4942170236" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11027" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11027/hovercard" href="https://github.com/mudler/LocalAI/pull/11027">#11027</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4984605739" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11139" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11139/hovercard" href="https://github.com/mudler/LocalAI/pull/11139">#11139</a></p>
</blockquote>
<h3>📦 Hugging Face model artifacts</h3>
<p>Model artifacts from Hugging Face are now materialized as a managed snapshot flow: immutable snapshot resolution, authenticated downloads with progress reporting, materialization on gallery install and preload, runtime binding to the staged artifacts, and progress surfaced in the UI. Python backends reuse the Go download path rather than fetching on their own.</p>
<p>A substantial run of follow-ups landed alongside it: per-file resume of interrupted materialization rather than starting over, each writer staging into its own partial tree, companion artifacts persisted so remote workers receive the <code>base_model</code> option, single-file HF snapshots loaded from the file rather than the directory, inferred materialization gated by backend, CIFS <code>EACCES</code> treated as lock contention rather than failure, and multi-file install progress kept proportional during verification.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4885461877" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10825" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10825/hovercard" href="https://github.com/mudler/LocalAI/pull/10825">#10825</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4915090324" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10908" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10908/hovercard" href="https://github.com/mudler/LocalAI/pull/10908">#10908</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4915202135" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10909" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10909/hovercard" href="https://github.com/mudler/LocalAI/pull/10909">#10909</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4915416184" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10910" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10910/hovercard" href="https://github.com/mudler/LocalAI/pull/10910">#10910</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921972992" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10949" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10949/hovercard" href="https://github.com/mudler/LocalAI/pull/10949">#10949</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931651715" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10986" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10986/hovercard" href="https://github.com/mudler/LocalAI/pull/10986">#10986</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4932858683" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10995" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10995/hovercard" href="https://github.com/mudler/LocalAI/pull/10995">#10995</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4957066833" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11071" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11071/hovercard" href="https://github.com/mudler/LocalAI/pull/11071">#11071</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4959754921" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11075" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11075/hovercard" href="https://github.com/mudler/LocalAI/pull/11075">#11075</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4972933657" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11117" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11117/hovercard" href="https://github.com/mudler/LocalAI/pull/11117">#11117</a></p>
</blockquote>
<h3>⚡ A much lighter web UI and trace API</h3>
<p>Three HTTP-layer problems, all measured on a live deployment, fixed together because they all shape what goes on the wire.</p>
<p>The server sent no <code>Content-Encoding</code> at all, regardless of <code>Accept-Encoding</code>. There is now gzip middleware, on by default and controllable with <code>--disable-http-compression</code> / <code>LOCALAI_DISABLE_HTTP_COMPRESSION</code> and <code>--http-compression-min-length</code> / <code>LOCALAI_HTTP_COMPRESSION_MIN_LENGTH</code> (default 1024). Streaming responses are skipped explicitly, since buffering them behind a gzip writer defeats incremental flushing and reads as a hung stream: SSE <code>Accept</code> headers, WebSocket upgrades, and the completion, realtime, speech, transcription, agent-job and log-tail path prefixes. Already-compressed formats are skipped too, because gzip made those marginally larger.</p>
<p>Vite content-hashes the bundle filenames, so an <code>/assets/</code> URL can never change content, yet they shipped with no <code>Cache-Control</code>, <code>ETag</code> or <code>Last-Modified</code>. They now carry <code>public, max-age=31536000, immutable</code>, <code>index.html</code> is explicitly <code>no-cache</code> so deploys are picked up, and unhashed locale JSONs get a 5 minute policy.</p>
<p><code>/api/traces</code> was returning a 21 MB unpaginated blob that the UI polled every 5 seconds. Both list endpoints now accept <code>limit</code> (default 50, max 1000, <code>0</code> for all), <code>offset</code> and <code>full</code>, and summarize by default: bodies and headers are dropped, the byte counters kept so the UI can still report what went missing. Every trace carries a process-lifetime <code>id</code>, and <code>GET /api/traces/{id}</code> serves the full record on expand or export. Paging metadata rides in <code>X-Total-Count</code>, <code>X-Trace-Offset</code> and <code>X-Trace-Limit</code>, so the list body stays a plain JSON array for existing consumers.</p>
<table>
<thead>
<tr>
<th></th>
<th>Before</th>
<th>After</th>
<th>Change</th>
</tr>
</thead>
<tbody>
<tr>
<td>React JS + CSS over the wire</td>
<td>2,815,513 B</td>
<td>807,918 B</td>
<td><strong>3.48x smaller</strong></td>
</tr>
<tr>
<td>All embedded assets (incl. fonts)</td>
<td>3,953,917 B</td>
<td>1,559,787 B</td>
<td>2.53x smaller</td>
</tr>
<tr>
<td>Repeat navigation asset transfer</td>
<td>full re-download</td>
<td>0 bytes</td>
<td>eliminated</td>
</tr>
<tr>
<td><code>/api/backend-traces</code> poll payload</td>
<td>21,131,097 B</td>
<td>7,201 B</td>
<td><strong>~2900x smaller</strong></td>
</tr>
</tbody>
</table>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4951173886" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11056" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11056/hovercard" href="https://github.com/mudler/LocalAI/pull/11056">#11056</a></p>
</blockquote>
<h3>🎚️ Per-node VRAM allocation budgets</h3>
<p>Operators can now cap how much VRAM LocalAI uses for model allocation on a node, as a percentage (<code>80%</code>) or an absolute amount (<code>12GB</code>). Everywhere LocalAI reads VRAM to make an allocation decision it now uses <code>min(detected, budget)</code>, a hard ceiling that never raises usable VRAM above physical. Percentages above 100% are rejected; absolute values above physical are clamped.</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="LOCALAI_VRAM_BUDGET=80%
LOCALAI_VRAM_BUDGET=12GB"><pre class="notranslate"><code>LOCALAI_VRAM_BUDGET=80%
LOCALAI_VRAM_BUDGET=12GB
</code></pre></div>
<p>It applies to both <code>local-ai</code> and <code>local-ai worker</code> (also <code>--vram-budget</code>), and is editable live from the standalone Settings page and per node in the distributed node UI.</p>
<p>The two paths are deliberately asymmetric:</p>
<ul>
<li><strong>Standalone</strong>, a hard per-process cap. <code>xsysinfo</code> holds it as a process-global default, so hardware defaults, context auto-fit, GGUF warnings and the watchdog all inherit it.</li>
<li><strong>Distributed</strong>, a placement ceiling. The worker reports raw VRAM plus its budget string; the registry resolves it and caps stored <code>available_vram</code> on registration and heartbeat, so the SQL scheduler needs no query change. The worker still sees its full card for its own context-fit.</li>
</ul>
<p>Admin overrides via <code>PUT</code>/<code>DELETE /api/nodes/:id/vram-budget</code> survive worker restarts, and are exposed as the <code>set_node_vram_budget</code> MCP tool.</p>
<p>Default unset means all detected VRAM, so existing deployments are unchanged.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4887571285" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10833" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10833/hovercard" href="https://github.com/mudler/LocalAI/pull/10833">#10833</a></p>
</blockquote>
<h3>🗣️ Two new text-to-speech engines</h3>
<p><strong>magpie-tts-cpp</strong> wraps <a href="https://github.com/mudler/magpie-tts.cpp">magpie-tts.cpp</a>, a C++17/ggml port of NVIDIA's Magpie TTS Multilingual 357M with its NanoCodec vocoder embedded: 5 voices (Aria, Jason, John, Leo, Sofia), 9+ languages, 22.05 kHz mono, from one self-contained GGUF with no Python or PyTorch at inference. GGUFs are published at <a href="https://huggingface.co/mudler/magpie-tts.cpp-gguf" rel="nofollow">mudler/magpie-tts.cpp-gguf</a>. A live gRPC check returns a valid non-silent WAV that round-trips exactly through ASR, and the upstream engine is parity-gated against NeMo per component (teacher-forced replay max abs diff 3.6e-5).</p>
<p><strong>moss-tts-cpp</strong> wraps <a href="https://github.com/mudler/moss-tts.cpp">moss-tts.cpp</a>, the ggml port of the OpenMOSS MOSS-TTS family, serving MOSS-TTS-Local v1.5 (a GPT-J local transformer decoded through MOSS-Audio-Tokenizer-v2). It produces 48 kHz stereo with optional reference-audio voice cloning. GGUFs are at <a href="https://huggingface.co/mudler/MOSS-TTS-Local-Transformer-v1.5-GGUF" rel="nofollow">mudler/MOSS-TTS-Local-Transformer-v1.5-GGUF</a>. Images cover CPU, CUDA 12/13, Intel SYCL f16/f32, Vulkan, ROCm, NVIDIA L4T and Darwin Metal.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4972768992" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11115" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11115/hovercard" href="https://github.com/mudler/LocalAI/pull/11115">#11115</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4900496924" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10860" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10860/hovercard" href="https://github.com/mudler/LocalAI/pull/10860">#10860</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4909181290" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10877" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10877/hovercard" href="https://github.com/mudler/LocalAI/pull/10877">#10877</a></p>
</blockquote>
<h3>🌳 Sub-2-bit models with the bonsai backend</h3>
<p>The <a href="https://huggingface.co/prism-ml" rel="nofollow">Bonsai models</a> are 1-bit (Q1_0) and ternary / 1.58-bit (Q2_0) quantizations of Qwen3 8B dense and Qwen3.6-27B hybrid attention. Their quant formats are only decodable by the <a href="https://github.com/PrismML-Eng/llama.cpp">PrismML fork of llama.cpp</a>, since stock llama.cpp has no Q1_0/Q2_0 kernels, so they need a dedicated fork backend in the same shape as <code>ik-llama-cpp</code> and <code>turboquant</code>.</p>
<p>The backend reuses <code>backend/cpp/llama-cpp/grpc-server.cpp</code> against the fork's <code>libllama</code> through a thin wrapper Makefile that only swaps <code>LLAMA_REPO</code> and <code>LLAMA_VERSION</code>, so these models are served over the same OpenAI-compatible API as stock <code>llama-cpp</code>. The reused server compiles against the fork with zero skew patches.</p>
<p>Eight gallery entries ship with it:</p>
<table>
<thead>
<tr>
<th>Family</th>
<th>Variants</th>
<th>Notes</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>bonsai-8b-1bit</code></td>
<td>Q1_0</td>
<td>Qwen3-8B, ~1.15 GB</td>
</tr>
<tr>
<td><code>ternary-bonsai-8b</code></td>
<td>Q2_0, <code>-q2-g64</code>, <code>-pq2</code></td>
<td>Qwen3-8B, ~2.18 GB</td>
</tr>
<tr>
<td><code>bonsai-27b-1bit</code></td>
<td>Q1_0</td>
<td>Qwen3.6-27B hybrid attention, vision, ~3.9 GB</td>
</tr>
<tr>
<td><code>ternary-bonsai-27b</code></td>
<td>Q2_0, <code>-pq2</code>, <code>-q2-g64</code></td>
<td>Qwen3.6-27B hybrid attention, vision, ~7.2 GB</td>
</tr>
</tbody>
</table>
<p>If Q1_0 and Q2_0 land in mainline llama.cpp, this backend can retire in favor of a routine <code>LLAMA_VERSION</code> bump on stock <code>llama-cpp</code>.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4887800908" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10834" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10834/hovercard" href="https://github.com/mudler/LocalAI/pull/10834">#10834</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4903975213" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10866" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10866/hovercard" href="https://github.com/mudler/LocalAI/pull/10866">#10866</a></p>
</blockquote>
<h3>🖧 Distributed mode reliability</h3>
<p>Investigating a model that showed as loaded on the home page but appeared on no node in the cluster turned up four separate bugs, all fixed here.</p>
<p>The reaper was deleting rows for backends that were alive and working. <code>probeLoadedModels</code> reaped a <code>node_models</code> row after one failed 1 second health check, and a busy backend cannot answer one: a single-threaded Python backend blocks for minutes inside a request.</p>
<ul>
<li>A new <code>models.running</code> subject asks the worker directly, since it holds the process handle and is not blocked by the backend. The reconciler diffs its process keys against the registry before any port probe.</li>
<li>A worker that does not answer is skipped, not assumed empty, so a NATS blip cannot delete a node's rows.</li>
<li>The port probe now separates <code>DeadlineExceeded</code> (busy) from <code>Unavailable</code> (gone), and only the latter counts, after three consecutive misses.</li>
</ul>
<p>Every routed model also left an in-process stub in the frontend's <code>ModelLoader</code>, and removal paths deleted only the database row, so the stub outlived the replica and the model was reported as loaded forever. The replica-removed hook became a list, and a new local-stub invalidator drops the stub once no healthy replica remains cluster-wide.</p>
<p>Alongside those: <code>in_flight</code> counters could leak high and pin a replica's VRAM against eviction; model-load deadlines now scale with checkpoint size and with progress rather than wall-clock; staging verification counts as progress rather than a stall; backend discovery no longer hides worker-installed or GPU-only backends behind the controller's filesystem and capability; the scheduler will not place a model on a node that cannot store it; and open responses are visible and cancellable across replicas.</p>
<p>Worker-side, a backend process whose directory a reinstall replaced is never reused, the gRPC port allocator is bounded and stops leaking dead backends' ports, deleted backends are reaped, and the worker has a real health endpoint with a mode-aware <code>HEALTHCHECK</code>.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4988542606" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11142" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11142/hovercard" href="https://github.com/mudler/LocalAI/pull/11142">#11142</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4976038430" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11121" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11121/hovercard" href="https://github.com/mudler/LocalAI/pull/11121">#11121</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4943004220" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11030" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11030/hovercard" href="https://github.com/mudler/LocalAI/pull/11030">#11030</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4942661675" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11029" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11029/hovercard" href="https://github.com/mudler/LocalAI/pull/11029">#11029</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4941769016" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11026" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11026/hovercard" href="https://github.com/mudler/LocalAI/pull/11026">#11026</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4938596558" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11019" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11019/hovercard" href="https://github.com/mudler/LocalAI/pull/11019">#11019</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4933022503" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11000" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11000/hovercard" href="https://github.com/mudler/LocalAI/pull/11000">#11000</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4933001118" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10999" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10999/hovercard" href="https://github.com/mudler/LocalAI/pull/10999">#10999</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931929199" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10990" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10990/hovercard" href="https://github.com/mudler/LocalAI/pull/10990">#10990</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924920294" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10970" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10970/hovercard" href="https://github.com/mudler/LocalAI/pull/10970">#10970</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924687474" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10968" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10968/hovercard" href="https://github.com/mudler/LocalAI/pull/10968">#10968</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924616250" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10967" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10967/hovercard" href="https://github.com/mudler/LocalAI/pull/10967">#10967</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924571166" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10966" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10966/hovercard" href="https://github.com/mudler/LocalAI/pull/10966">#10966</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4922508428" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10956" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10956/hovercard" href="https://github.com/mudler/LocalAI/pull/10956">#10956</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921690924" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10948" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10948/hovercard" href="https://github.com/mudler/LocalAI/pull/10948">#10948</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921687286" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10947" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10947/hovercard" href="https://github.com/mudler/LocalAI/pull/10947">#10947</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4890100883" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10838" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10838/hovercard" href="https://github.com/mudler/LocalAI/pull/10838">#10838</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4950809703" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11054" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11054/hovercard" href="https://github.com/mudler/LocalAI/pull/11054">#11054</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4757624550" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10551" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10551/hovercard" href="https://github.com/mudler/LocalAI/pull/10551">#10551</a></p>
</blockquote>
<h3>🧰 Smaller features worth knowing about</h3>
<ul>
<li><strong>Valkey Search vector store</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5010751998" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11196" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11196/hovercard" href="https://github.com/mudler/LocalAI/pull/11196">#11196</a>): a new <code>valkey-store</code> backend adds Valkey as a vector store option alongside the existing ones.</li>
<li><strong>systemd socket activation</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5001701347" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11169" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11169/hovercard" href="https://github.com/mudler/LocalAI/pull/11169">#11169</a>): <code>local-ai</code> consumes a TCP listener inherited through the systemd socket-activation protocol on Linux, so it can start on demand. Ordinary <code>--address</code> / <code>LOCALAI_ADDRESS</code> binding is unchanged when no activation listener is present, ambiguous multiple listeners are rejected, and the public-bind auth safety check runs against the actual inherited address. Documented alongside the Podman descriptor-passing requirement.</li>
<li><strong>Persistent trace history</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5015402002" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11203" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11203/hovercard" href="https://github.com/mudler/LocalAI/pull/11203">#11203</a>): API and backend trace histories now persist under the configured data path and survive a restart, as bounded per-record JSON files under <code>traces/api</code> and <code>traces/backend</code>. No database dependency, existing <code>tracing_max_items</code> bounds preserved, restored IDs advanced to avoid collisions, corrupt records skipped rather than blocking startup.</li>
<li><strong>Edit saved chat messages</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5007026265" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11189" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11189/hovercard" href="https://github.com/mudler/LocalAI/pull/11189">#11189</a>): inline Edit / Save / Cancel on saved user prompts and assistant responses, persisted through local chat history with no inference request, preserving structured content blocks and attachment metadata.</li>
<li><strong>Self-contained Intel SYCL backend</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4932492674" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10991" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10991/hovercard" href="https://github.com/mudler/LocalAI/pull/10991">#10991</a>): the Intel llama.cpp backend now runs on any host rather than requiring a matching oneAPI runtime.</li>
<li><strong>Configurable VAE tiling</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5017541960" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11216" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11216/hovercard" href="https://github.com/mudler/LocalAI/pull/11216">#11216</a>) for <code>stablediffusion-ggml</code>, and <strong>voice control on low-power devices</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4875032326" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10804" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10804/hovercard" href="https://github.com/mudler/LocalAI/pull/10804">#10804</a>) via the classifier/VAD path.</li>
<li><strong>Anthropic prompt-cache breakpoints</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4995621679" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11158" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11158/hovercard" href="https://github.com/mudler/LocalAI/pull/11158">#11158</a>): optional cache breakpoints in cloud-proxy translate mode.</li>
<li><strong><code>/v1/detokenize</code></strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4355857936" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/9620" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/9620/hovercard" href="https://github.com/mudler/LocalAI/pull/9620">#9620</a>), and <strong>deterministic, type-filtered backend auto-detection</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4652830357" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10286" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10286/hovercard" href="https://github.com/mudler/LocalAI/pull/10286">#10286</a>) so backend selection stops depending on probe order.</li>
<li><strong>MLX TTS routing</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5034271537" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11267" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11267/hovercard" href="https://github.com/mudler/LocalAI/pull/11267">#11267</a>): MLX TTS models now import to <code>mlx-audio</code> rather than a generic text backend.</li>
</ul>
<h3>🛡️ Inline fine-tuning reward code is now opt-in</h3>
<p><code>POST /api/fine-tuning/jobs</code> accepts <code>reward_functions[].code</code>, an inline Python body that was executed against a hand-rolled builtin allowlist. That allowlist was not a security boundary: standard CPython introspection reaches the real <code>os</code> module and yields arbitrary code execution on the host. Execution happened synchronously during a smoke test at job start, and the fine-tuning endpoint is unauthenticated by default.</p>
<p>Rather than trying to harden the allowlist, inline reward code is now refused unless the operator explicitly opts in with <code>LOCALAI_TRL_ALLOW_INLINE_REWARD=true</code> on the backend. Builtin reward functions are unaffected and keep working with no configuration. The documentation no longer describes the allowlist as a sandbox and states plainly that inline code is arbitrary execution.</p>
<p>Two further hardening fixes landed in the same cycle:</p>
<ul>
<li><strong>Tar hardlinks that escape the extraction root are rejected</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5033915721" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11266" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11266/hovercard" href="https://github.com/mudler/LocalAI/pull/11266">#11266</a>). <code>ExtractArchive</code> pre-scanned members and rejected symlinks, but tar hardlink entries carry a regular file mode and passed that check, and <code>Header.Linkname</code> was never validated, so an archive could create a link to a path outside the destination directory. Linkname now gets the same path check as member names; hardlinks resolving inside the root still extract, so ordinary archives are unaffected.</li>
<li><strong>Cyclic <code>$ref</code> in a JSON-schema grammar is rejected</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4944468246" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11041" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11041/hovercard" href="https://github.com/mudler/LocalAI/pull/11041">#11041</a>) rather than recursing into a stack-overflow crash.</li>
</ul>
<p>This release also picks up hono 4.12.25 for <a title="CVE-2026-54290" data-hovercard-type="advisory" data-hovercard-url="/advisories/GHSA-88fw-hqm2-52qc/hovercard" href="https://github.com/advisories/GHSA-88fw-hqm2-52qc">CVE-2026-54290</a>.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4955900053" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11068" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11068/hovercard" href="https://github.com/mudler/LocalAI/pull/11068">#11068</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5033915721" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11266" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11266/hovercard" href="https://github.com/mudler/LocalAI/pull/11266">#11266</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4944468246" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11041" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11041/hovercard" href="https://github.com/mudler/LocalAI/pull/11041">#11041</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4940334409" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11023" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11023/hovercard" href="https://github.com/mudler/LocalAI/pull/11023">#11023</a></p>
</blockquote>
<h3>🔍 Traces gain request identity</h3>
<p>The API Traces panel recorded who issued each request but never showed it, and never captured the caller's network identity. The table now has a sortable User column, and the expanded row carries User, Client IP and User Agent, from echo's <code>RealIP()</code> (honouring <code>X-Forwarded-For</code> / <code>X-Real-IP</code> behind a trusted proxy). Fields render only when present, so older buffered traces and unauthenticated local requests degrade cleanly. The time column now shows the date too.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4914779330" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10907" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10907/hovercard" href="https://github.com/mudler/LocalAI/pull/10907">#10907</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4914694638" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10905" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10905/hovercard" href="https://github.com/mudler/LocalAI/pull/10905">#10905</a></p>
</blockquote>
<hr>
<h2>🐛 Bug Fixes (recap)</h2>
<ul>
<li><code>fix(distributed)</code>: reaper reaps live backends, ghost model stubs, <code>in_flight</code> leak, sidecar staging runaway - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4988542606" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11142" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11142/hovercard" href="https://github.com/mudler/LocalAI/pull/11142">#11142</a></li>
<li><code>fix(distributed)</code>: scale the remote model-load deadline with checkpoint size - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4943004220" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11030" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11030/hovercard" href="https://github.com/mudler/LocalAI/pull/11030">#11030</a></li>
<li><code>fix(distributed)</code>: make the cold-load hold scale with progress, not wall-clock - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4938596558" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11019" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11019/hovercard" href="https://github.com/mudler/LocalAI/pull/11019">#11019</a></li>
<li><code>fix(distributed)</code>: count staging verification as progress, not as a stall - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4941769016" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11026" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11026/hovercard" href="https://github.com/mudler/LocalAI/pull/11026">#11026</a></li>
<li><code>fix(distributed)</code>: reject wrong-model requests at the backend and on the remaining modalities - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924920294" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10970" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10970/hovercard" href="https://github.com/mudler/LocalAI/pull/10970">#10970</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931929199" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10990" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10990/hovercard" href="https://github.com/mudler/LocalAI/pull/10990">#10990</a></li>
<li><code>fix(distributed)</code>: backend discovery hid worker-installed and GPU-only backends - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924616250" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10967" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10967/hovercard" href="https://github.com/mudler/LocalAI/pull/10967">#10967</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921687286" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10947" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10947/hovercard" href="https://github.com/mudler/LocalAI/pull/10947">#10947</a></li>
<li><code>fix(distributed)</code>: configurable remote model-load timeout, and reap the load when it times out - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921690924" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10948" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10948/hovercard" href="https://github.com/mudler/LocalAI/pull/10948">#10948</a></li>
<li><code>fix(distributed)</code>: make per-node backend upgrade actually upgrade - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4890100883" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10838" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10838/hovercard" href="https://github.com/mudler/LocalAI/pull/10838">#10838</a></li>
<li><code>fix(nodes)</code>: never schedule a model onto a node that cannot store it - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4950809703" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11054" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11054/hovercard" href="https://github.com/mudler/LocalAI/pull/11054">#11054</a></li>
<li><code>fix(openresponses)</code>: make responses visible and cancellable across replicas - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4933022503" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11000" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11000/hovercard" href="https://github.com/mudler/LocalAI/pull/11000">#11000</a></li>
<li><code>fix(worker)</code>: never reuse a backend process whose directory a reinstall replaced - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4942661675" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11029" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11029/hovercard" href="https://github.com/mudler/LocalAI/pull/11029">#11029</a></li>
<li><code>fix(worker)</code>: bound the gRPC port allocator and stop leaking dead backends' ports - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924687474" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10968" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10968/hovercard" href="https://github.com/mudler/LocalAI/pull/10968">#10968</a></li>
<li><code>fix(worker)</code>: reap deleted backends and stop models that live on a worker - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4922508428" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10956" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10956/hovercard" href="https://github.com/mudler/LocalAI/pull/10956">#10956</a></li>
<li><code>fix(worker)</code>: give the worker a real health endpoint and a mode-aware HEALTHCHECK - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4933001118" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10999" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10999/hovercard" href="https://github.com/mudler/LocalAI/pull/10999">#10999</a></li>
<li><code>fix(downloader)</code>: hash the partial file before issuing the resume request - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4967730602" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11099" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11099/hovercard" href="https://github.com/mudler/LocalAI/pull/11099">#11099</a></li>
<li><code>fix(downloader)</code>: bound the wait for response headers so a wedged origin cannot hang an install - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4950737277" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11053" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11053/hovercard" href="https://github.com/mudler/LocalAI/pull/11053">#11053</a></li>
<li><code>fix(downloader)</code>: distinguish read from write failures and retry transient ones - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931612325" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10985" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10985/hovercard" href="https://github.com/mudler/LocalAI/pull/10985">#10985</a></li>
<li><code>fix(modelartifacts)</code>: resume interrupted materialization per-file, not from scratch - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4957066833" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11071" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11071/hovercard" href="https://github.com/mudler/LocalAI/pull/11071">#11071</a></li>
<li><code>fix(modelartifacts)</code>: stage each writer's artifact in its own partial tree - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4932858683" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10995" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10995/hovercard" href="https://github.com/mudler/LocalAI/pull/10995">#10995</a></li>
<li><code>fix(modelartifacts)</code>: treat CIFS EACCES as lock contention, not failure - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931651715" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10986" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10986/hovercard" href="https://github.com/mudler/LocalAI/pull/10986">#10986</a></li>
<li><code>fix(model-artifacts)</code>: persist companion artifacts so remote workers get the <code>base_model</code> option - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4959754921" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11075" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11075/hovercard" href="https://github.com/mudler/LocalAI/pull/11075">#11075</a></li>
<li><code>fix(model-artifacts)</code>: load single-file HF snapshots from the file, not the directory - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4915202135" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10909" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10909/hovercard" href="https://github.com/mudler/LocalAI/pull/10909">#10909</a></li>
<li><code>fix(model-artifacts)</code>: gate inferred artifact materialization by backend - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4915416184" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10910" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10910/hovercard" href="https://github.com/mudler/LocalAI/pull/10910">#10910</a></li>
<li><code>fix(model-artifacts)</code>: materialize longcat-video on the controller, and support companion repos - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921972992" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10949" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10949/hovercard" href="https://github.com/mudler/LocalAI/pull/10949">#10949</a></li>
<li><code>fix(gallery)</code>: coalesce Hugging Face artifact progress - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4972933657" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11117" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11117/hovercard" href="https://github.com/mudler/LocalAI/pull/11117">#11117</a></li>
<li><code>fix(gallery)</code>: keep multi-file HF install progress proportional during verify - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4915090324" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10908" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10908/hovercard" href="https://github.com/mudler/LocalAI/pull/10908">#10908</a></li>
<li><code>fix(galleryop)</code>: make admitted operations queryable and survive a failed op - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4947914468" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11044" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11044/hovercard" href="https://github.com/mudler/LocalAI/pull/11044">#11044</a></li>
<li><code>fix(gpu-libs)</code>: bundle cuDNN only where it is used, and complete it when it is - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921685891" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10946" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10946/hovercard" href="https://github.com/mudler/LocalAI/pull/10946">#10946</a></li>
<li><code>fix(gpu)</code>: detect GPUs via sysfs when no pci.ids database is present - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4924571166" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10966" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10966/hovercard" href="https://github.com/mudler/LocalAI/pull/10966">#10966</a></li>
<li><code>fix(watchdog)</code>: force-kill stuck-busy backends instead of deadlocking the loader - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4762810651" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10578" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10578/hovercard" href="https://github.com/mudler/LocalAI/pull/10578">#10578</a></li>
<li><code>fix(watchdog)</code>: guard <code>StopWatchdog</code> with <code>watchdogMutex</code> to prevent double close - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4899656281" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10859" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10859/hovercard" href="https://github.com/mudler/LocalAI/pull/10859">#10859</a></li>
<li><code>fix(config)</code>: only inject llama.cpp serving options on the llama.cpp path - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4884685532" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10822" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10822/hovercard" href="https://github.com/mudler/LocalAI/pull/10822">#10822</a></li>
<li><code>fix(runtime-settings)</code>: apply persisted threads/context_size/f16 at startup - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4896186074" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10853" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10853/hovercard" href="https://github.com/mudler/LocalAI/pull/10853">#10853</a></li>
<li><code>fix(model)</code>: make backend shutdown model-scoped - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4902986100" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10865" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10865/hovercard" href="https://github.com/mudler/LocalAI/pull/10865">#10865</a></li>
<li><code>fix(model)</code>: only announce a load at INFO when a load actually happens - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4938332752" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11017" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11017/hovercard" href="https://github.com/mudler/LocalAI/pull/11017">#11017</a></li>
<li><code>fix(completions)</code>: reject empty <code>PromptStrings</code> in streaming to avoid an index-out-of-range panic - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4942281643" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11028" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11028/hovercard" href="https://github.com/mudler/LocalAI/pull/11028">#11028</a></li>
<li><code>fix(tts)</code>: forward the OpenAI <code>speed</code> field to the backend - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4975921380" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11120" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11120/hovercard" href="https://github.com/mudler/LocalAI/pull/11120">#11120</a></li>
<li><code>fix(realtime)</code>: accept the legacy <code>modalities</code> alias for <code>output_modalities</code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4970143323" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11104" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11104/hovercard" href="https://github.com/mudler/LocalAI/pull/11104">#11104</a></li>
<li><code>fix(vision)</code>: probe the media marker for pinned llama.cpp backend variants - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4922288264" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10955" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10955/hovercard" href="https://github.com/mudler/LocalAI/pull/10955">#10955</a></li>
<li><code>fix(audio-transform)</code>: serialize WebSocket writes to avoid a concurrent-write panic - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4897726483" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10857" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10857/hovercard" href="https://github.com/mudler/LocalAI/pull/10857">#10857</a></li>
<li><code>fix(qwen-asr)</code>: map ISO language codes to the names Qwen3-ASR expects - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4922819377" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10959" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10959/hovercard" href="https://github.com/mudler/LocalAI/pull/10959">#10959</a></li>
<li><code>fix(ollama)</code>: cap <code>num_ctx</code> so it cannot wrap negative when cast to int32 - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4943551948" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11032" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11032/hovercard" href="https://github.com/mudler/LocalAI/pull/11032">#11032</a></li>
<li><code>fix(ollama)</code>: set <code>ContextSize</code> via the embedded <code>LLMConfig</code> so the package builds - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4949924957" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11049" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11049/hovercard" href="https://github.com/mudler/LocalAI/pull/11049">#11049</a></li>
<li><code>fix(webui)</code>: use relative asset base so fonts and lazy chunks honor <code>X-Forwarded-Prefix</code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4914673922" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10904" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10904/hovercard" href="https://github.com/mudler/LocalAI/pull/10904">#10904</a></li>
<li><code>fix(agent-ui)</code>: reset streamed text at generation boundaries in agent chat - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4803925583" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10664" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10664/hovercard" href="https://github.com/mudler/LocalAI/pull/10664">#10664</a></li>
<li><code>fix(mcp)</code>: bound MCP session connect so an unreachable server cannot hang the widget - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4912779232" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10884" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10884/hovercard" href="https://github.com/mudler/LocalAI/pull/10884">#10884</a></li>
<li><code>fix(http)</code>: make <code>/readyz</code> reflect startup readiness - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931842920" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10989" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10989/hovercard" href="https://github.com/mudler/LocalAI/pull/10989">#10989</a></li>
<li><code>fix(upgrade-check)</code>: don't filter upgrade candidates by controller capability - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4940904340" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11024" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11024/hovercard" href="https://github.com/mudler/LocalAI/pull/11024">#11024</a></li>
<li><code>fix(cloud-proxy)</code>: publish backend gallery entries - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4899126142" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10858" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10858/hovercard" href="https://github.com/mudler/LocalAI/pull/10858">#10858</a></li>
<li><code>fix(backend)</code>: don't crash the whole process on an invalid <code>cutstrings</code>/<code>extract_regex</code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4896741157" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10855" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10855/hovercard" href="https://github.com/mudler/LocalAI/pull/10855">#10855</a></li>
<li><code>fix(backends)</code>: derive the protoc generator from the protobuf runtime - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4951328680" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11057" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11057/hovercard" href="https://github.com/mudler/LocalAI/pull/11057">#11057</a></li>
<li><code>fix(backend/python)</code>: don't await sync servicer behaviors in <code>AsyncModelIdentityInterceptor</code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4930915437" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10980" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10980/hovercard" href="https://github.com/mudler/LocalAI/pull/10980">#10980</a></li>
<li><code>fix(sglang)</code>: implement the Status RPC to unblock backend-monitor polling - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4905476276" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10867" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10867/hovercard" href="https://github.com/mudler/LocalAI/pull/10867">#10867</a></li>
<li><code>fix(vllm)</code>: generate protobuf 6 compatible stubs - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4921098476" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10944" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10944/hovercard" href="https://github.com/mudler/LocalAI/pull/10944">#10944</a></li>
<li><code>fix(vibevoice)</code>: install diffusers from PyPI instead of git main - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4925439664" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10972" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10972/hovercard" href="https://github.com/mudler/LocalAI/pull/10972">#10972</a></li>
<li><code>fix(kokoro)</code>: pin a compatible Intel XPU runtime - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4884958434" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10823" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10823/hovercard" href="https://github.com/mudler/LocalAI/pull/10823">#10823</a></li>
<li><code>fix(ace-step)</code>: drop nonexistent <code>Get*</code> proto accessors in <code>SoundGeneration</code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4957707339" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11072" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11072/hovercard" href="https://github.com/mudler/LocalAI/pull/11072">#11072</a></li>
<li><code>fix(trl)</code>: disable inline GRPO reward code by default (RCE) - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4955900053" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11068" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11068/hovercard" href="https://github.com/mudler/LocalAI/pull/11068">#11068</a></li>
<li><code>fix(turboquant)</code>: supersede stale dependency bump - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4954160283" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11064" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11064/hovercard" href="https://github.com/mudler/LocalAI/pull/11064">#11064</a></li>
<li><code>fix(turboquant,bonsai)</code>: do not apply vendored llama.cpp patches to fork trees - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4903975213" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10866" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10866/hovercard" href="https://github.com/mudler/LocalAI/pull/10866">#10866</a></li>
<li><code>fix(llama-cpp)</code>: retain CPU variants in GPU builds - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5030315652" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11255" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11255/hovercard" href="https://github.com/mudler/LocalAI/pull/11255">#11255</a>, and the same for turboquant - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5037346156" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11276" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11276/hovercard" href="https://github.com/mudler/LocalAI/pull/11276">#11276</a></li>
<li><code>fix(llama-cpp)</code>: preserve GPU layers during option passthrough - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5010206055" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11193" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11193/hovercard" href="https://github.com/mudler/LocalAI/pull/11193">#11193</a></li>
<li><code>fix(utils)</code>: reject tar hardlinks that escape the extraction root - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5033915721" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11266" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11266/hovercard" href="https://github.com/mudler/LocalAI/pull/11266">#11266</a></li>
<li><code>fix(grammars)</code>: reject cyclic <code>$ref</code> in JSON-schema grammar to prevent a stack-overflow crash - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4944468246" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11041" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11041/hovercard" href="https://github.com/mudler/LocalAI/pull/11041">#11041</a></li>
<li><code>fix(grammars)</code>: restore backslash escaping in the llama31 grammar fixture - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5024276464" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11242" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11242/hovercard" href="https://github.com/mudler/LocalAI/pull/11242">#11242</a></li>
<li><code>fix(model)</code>: deterministic, type-filtered backend auto-detection - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4652830357" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10286" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10286/hovercard" href="https://github.com/mudler/LocalAI/pull/10286">#10286</a></li>
<li><code>fix(oci)</code>: install backends on filesystems without symlinks - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5000662284" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11166" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11166/hovercard" href="https://github.com/mudler/LocalAI/pull/11166">#11166</a></li>
<li><code>fix(oci)</code>: identify signature verification requests - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5025690331" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11244" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11244/hovercard" href="https://github.com/mudler/LocalAI/pull/11244">#11244</a></li>
<li><code>fix(realtime)</code>: echo <code>response.metadata</code> on <code>response.created</code> and <code>response.done</code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5012360188" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11198" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11198/hovercard" href="https://github.com/mudler/LocalAI/pull/11198">#11198</a></li>
<li><code>fix(worker)</code>: report RAM alongside GPU memory - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5001169400" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11167" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11167/hovercard" href="https://github.com/mudler/LocalAI/pull/11167">#11167</a></li>
<li><code>fix(vllm)</code>: apply <code>Options[]</code> engine flags before engine init - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4990661613" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11147" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11147/hovercard" href="https://github.com/mudler/LocalAI/pull/11147">#11147</a></li>
<li><code>fix(mlx-vlm)</code>: install torch dependencies on Metal - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4999454095" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11164" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11164/hovercard" href="https://github.com/mudler/LocalAI/pull/11164">#11164</a></li>
<li><code>fix(kokoro)</code>: add a CPU backend fallback - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4997378672" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11161" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11161/hovercard" href="https://github.com/mudler/LocalAI/pull/11161">#11161</a></li>
<li><code>fix(chatterbox)</code>: pin cublas12 torch/transformers and setuptools so the backend loads - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4959081581" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11074" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11074/hovercard" href="https://github.com/mudler/LocalAI/pull/11074">#11074</a></li>
<li><code>fix(gallery)</code>: correct Nanbeige 4.2 artifacts - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5035185388" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11269" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11269/hovercard" href="https://github.com/mudler/LocalAI/pull/11269">#11269</a></li>
<li>Video: WAN 2.1 GGML entries never set <code>known_usecases</code>, so they resolved to image rather than video and <code>/video</code> rejected them - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5016899491" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11214" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11214/hovercard" href="https://github.com/mudler/LocalAI/pull/11214">#11214</a></li>
</ul>
<h3>🖧 P2P area</h3>
<ul>
<li><code>fix(p2p)</code>: serialize access to <code>p2pCtx</code>/<code>p2pCancel</code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4900955861" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10861" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10861/hovercard" href="https://github.com/mudler/LocalAI/pull/10861">#10861</a></li>
</ul>
<hr>
<h2>👒 Dependencies</h2>
<p>121 dependency updates landed this cycle, mostly nightly engine bumps:</p>
<table>
<thead>
<tr>
<th>Engine</th>
<th>Bumps</th>
</tr>
</thead>
<tbody>
<tr>
<td>CrispStrobe/CrispASR</td>
<td>17</td>
</tr>
<tr>
<td>ikawrakow/ik_llama.cpp</td>
<td>12</td>
</tr>
<tr>
<td>leejet/stable-diffusion.cpp</td>
<td>10</td>
</tr>
<tr>
<td>ServeurpersoCom/qwentts.cpp</td>
<td>8</td>
</tr>
<tr>
<td>ggml-org/llama.cpp</td>
<td>7</td>
</tr>
<tr>
<td>ServeurpersoCom/omnivoice.cpp</td>
<td>6</td>
</tr>
<tr>
<td>PrismML-Eng/llama.cpp</td>
<td>5</td>
</tr>
<tr>
<td>mudler/parakeet.cpp</td>
<td>4</td>
</tr>
<tr>
<td>ggml-org/whisper.cpp</td>
<td>4</td>
</tr>
<tr>
<td>antirez/ds4</td>
<td>3</td>
</tr>
<tr>
<td>0xShug0/audio.cpp</td>
<td>2</td>
</tr>
<tr>
<td>magpie-tts.cpp, locate-anything.cpp, depth-anything.cpp, trellis2cpp, rf-detr.cpp, ced.cpp, llama-cpp-turboquant</td>
<td>1 each</td>
</tr>
</tbody>
</table>
<p>Plus 20 dependabot updates across Python, JavaScript and GitHub Actions, and a <code>go-processmanager</code> bump for the concurrent-<code>Run</code> fix.</p>
<p>Note: one <code>ggml-org/llama.cpp</code> bump (<code>d2a8182</code>) was reverted within the cycle and is excluded from these notes.</p>
<hr>
<h2>📖 Documentation</h2>
<p>The documentation received an onboarding-focused overhaul (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4914393152" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10895" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10895/hovercard" href="https://github.com/mudler/LocalAI/pull/10895">#10895</a>), driven by a mapped per-page audit rather than page dates, with every factual claim verified against the code, backends, gallery and CLI:</p>
<ul>
<li><strong>Accuracy</strong>: the CPU image tag corrected to <code>localai/localai:latest</code> (there is no <code>latest-cpu</code>), the registry unified, the dead <code>llama-stable</code> backend removed from examples, the <code>mitm-proxy</code> flag documentation corrected, a non-existent <code>/sound</code> endpoint removed, the Voice Activity Detection example made runnable, and the CLI reference refreshed with <code>agent</code>, <code>mcp-server</code>, <code>agent-worker</code> and <code>p2p-worker</code>.</li>
<li><strong>Deduplication</strong>: duplicate and stale pages folded into canonical homes, with all inbound links repointed and old URLs preserved via aliases.</li>
<li><strong>Onboarding</strong>: one concrete model (<code>qwen3-4b</code>) now carries through install, Web UI chat and API curl, plus a new <strong>Build your first agent</strong> walkthrough that states plainly that LocalAGI is embedded.</li>
<li><strong>Errors</strong>: a new <strong>Runtime errors and troubleshooting</strong> reference keyed on the literal error strings users see, plus a new <strong>Agent actions catalog</strong> taken from the shipped action registry.</li>
<li><strong>Structure</strong>: installation merged under Getting started for one linear install-to-first-run spine, a new <strong>Operations</strong> section for operator-facing pages, and journey-ordered navigation.</li>
<li><strong>Process</strong>: a docs checkbox in the PR template and a docs-with-code rule in the agent instructions, so user-facing code changes update docs in the same change.</li>
</ul>
<p>Also: <code>grpc.attempts</code> timing and tuning guidance (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4905521646" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10868" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10868/hovercard" href="https://github.com/mudler/LocalAI/pull/10868">#10868</a>), a fix to the Opus backend installation instructions for realtime (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4938342157" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11018" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11018/hovercard" href="https://github.com/mudler/LocalAI/pull/11018">#11018</a>), reverse-proxy and long-inference timeout guidance (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5010444428" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11195" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11195/hovercard" href="https://github.com/mudler/LocalAI/pull/11195">#11195</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4954168785" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11065" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11065/hovercard" href="https://github.com/mudler/LocalAI/pull/11065">#11065</a>), persistent container storage clarified (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5009014702" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11190" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11190/hovercard" href="https://github.com/mudler/LocalAI/pull/11190">#11190</a>), and ROCm 7.x / RDNA 3.5 (Strix Halo, gfx1151) added to the GPU acceleration guide (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4205134226" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/9229" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/9229/hovercard" href="https://github.com/mudler/LocalAI/pull/9229">#9229</a>).</p>
<h3>🌐 A new localai.io</h3>
<p>The site splits in two: the project site at the root, and the documentation moved under <code>/docs/</code>. The Hugo docs site had always <em>been</em> localai.io, which left nowhere to explain what LocalAI is or to show what the team builds.</p>
<p>Every previously published URL keeps working. GitHub Pages has no server-side rewrites, so a generator walks the built docs output and leaves a meta refresh, a canonical link and a <code>noindex</code> at each old root path: <strong>214 redirect stubs</strong>, covering bare <code>.html</code> files as well as directory indexes, and never overwriting a path the root site owns.</p>
<p>The new site adds an <code>/engines/</code> page driven entirely by a YAML data file (so adding an engine is one edit, not hand-written HTML in two places), a <code>/blog/</code>, a real POSIX <code>install.sh</code> and a Kubernetes manifest wired to the actual <code>/readyz</code> and <code>/healthz</code> endpoints. An ecosystem band lists the companies whose engineers have contributed, the projects that integrate LocalAI, and where LocalAI has been written about, each backed by a different and stated standard of evidence, with <code>ADOPTERS.md</code> as the self-service mechanism for anyone who wants to be listed.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5025113405" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11243" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11243/hovercard" href="https://github.com/mudler/LocalAI/pull/11243">#11243</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5027372436" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11248" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11248/hovercard" href="https://github.com/mudler/LocalAI/pull/11248">#11248</a></p>
</blockquote>
<h3>🧹 CI cost and correctness</h3>
<p>A sustained pass on the build pipeline, most of it invisible to users but responsible for how quickly changes land: the full backend matrix now only rebuilds on breaking <code>backend.proto</code> edits (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5010048678" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11192" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11192/hovercard" href="https://github.com/mudler/LocalAI/pull/11192">#11192</a>), image and Go PR workflows skip content they cannot see (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5018208332" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11218" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11218/hovercard" href="https://github.com/mudler/LocalAI/pull/11218">#11218</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5019991410" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11223" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11223/hovercard" href="https://github.com/mudler/LocalAI/pull/11223">#11223</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5020081909" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11224" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11224/hovercard" href="https://github.com/mudler/LocalAI/pull/11224">#11224</a>), the native engine builds in a layer the registry cache can actually restore (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5019202978" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11221" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11221/hovercard" href="https://github.com/mudler/LocalAI/pull/11221">#11221</a>), and three workflows that stacked runs on every PR push were deduplicated (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4951471396" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11058" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11058/hovercard" href="https://github.com/mudler/LocalAI/pull/11058">#11058</a>). Go backends now rebuild on linked <code>pkg/</code> changes and matrix-entry edits (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4931811653" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10988" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10988/hovercard" href="https://github.com/mudler/LocalAI/pull/10988">#10988</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4928599113" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10975" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10975/hovercard" href="https://github.com/mudler/LocalAI/pull/10975">#10975</a>).</p>
<hr>
<h2>🙌 New Contributors</h2>
<p>Eleven people landed their first LocalAI contribution this cycle:</p>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ankit-aglawe/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ankit-aglawe">@ankit-aglawe</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4917513402" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10930" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10930/hovercard" href="https://github.com/mudler/LocalAI/pull/10930">#10930</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/anupamme/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/anupamme">@anupamme</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4940334409" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11023" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11023/hovercard" href="https://github.com/mudler/LocalAI/pull/11023">#11023</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/futurehua/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/futurehua">@futurehua</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4911214293" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10879" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10879/hovercard" href="https://github.com/mudler/LocalAI/pull/10879">#10879</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ghshhf/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ghshhf">@ghshhf</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4657056891" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10323" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10323/hovercard" href="https://github.com/mudler/LocalAI/pull/10323">#10323</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jimmykarily/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jimmykarily">@jimmykarily</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4932492674" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10991" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10991/hovercard" href="https://github.com/mudler/LocalAI/pull/10991">#10991</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/nandanadileep/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/nandanadileep">@nandanadileep</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4762810651" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10578" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10578/hovercard" href="https://github.com/mudler/LocalAI/pull/10578">#10578</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/owezzy/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/owezzy">@owezzy</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5010444428" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11195" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11195/hovercard" href="https://github.com/mudler/LocalAI/pull/11195">#11195</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ResearchForumOnline/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ResearchForumOnline">@ResearchForumOnline</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4982990980" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11138" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11138/hovercard" href="https://github.com/mudler/LocalAI/pull/11138">#11138</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/wuisabel-gif/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/wuisabel-gif">@wuisabel-gif</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4955900053" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11068" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11068/hovercard" href="https://github.com/mudler/LocalAI/pull/11068">#11068</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Zelys-DFKH/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Zelys-DFKH">@Zelys-DFKH</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5033915721" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/11266" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/11266/hovercard" href="https://github.com/mudler/LocalAI/pull/11266">#11266</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/zjuzhongwen/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/zjuzhongwen">@zjuzhongwen</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4922933002" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10960" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10960/hovercard" href="https://github.com/mudler/LocalAI/pull/10960">#10960</a></li>
</ul>
<p>Thank you all, and thanks to everyone who filed issues, tested builds and reported regressions this cycle.</p>
<hr>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.7.1...v4.8.0"><tt>v4.7.1...v4.8.0</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.7.12026-07-14T16:17:51Zv4.7.1
<h2>What's Changed</h2>
<h3>Other Changes</h3>
<ul>
<li>fix(config): only inject llama.cpp serving options on the llama.cpp path by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4884685532" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10822" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10822/hovercard" href="https://github.com/mudler/LocalAI/pull/10822">#10822</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.7.0...v4.7.1"><tt>v4.7.0...v4.7.1</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.7.02026-07-14T13:32:50Zv4.7.0<h1>🎉 LocalAI 4.7.0 Release! 🚀</h1>
<h1 align="center">
<br>
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/refs/heads/master/core/http/static/logo.png"><img height="300" src="https://raw.githubusercontent.com/mudler/LocalAI/refs/heads/master/core/http/static/logo.png" style="max-width: 100%; height: auto; max-height: 300px;"></a>
<br>
<br>
</h1>
<p>LocalAI 4.7.0 is out!</p>
<p>This release widens what LocalAI can generate and how you drive it: a UI-managed voice cloning library, local video and audio-driven avatar generation, and interleaved reasoning that travels with tool calls. It also lands new audio engines (one-pass diarized transcription, F5-TTS, true streaming TTS) and a batch of reliability fixes across auth, transcription, and the model gallery.</p>
<p><strong>Highlights:</strong></p>
<ul>
<li>🎙️ <strong>Managed voice cloning</strong> - record or upload a consented reference in the UI, save it as a reusable profile, and reference it from any cloning-capable TTS backend with a stable <code>localai://voice-profiles/<id></code> URI. No more hand-edited YAML or copying audio into model folders.</li>
<li>🎬 <strong>Local video & avatars</strong> - a new <code>longcat-video</code> backend brings text-to-video, image-to-video, and audio-driven talking-avatar generation, wired into the Studio UI with audio and reference-image controls.</li>
<li>🧠 <strong>Interleaved thinking with tool calls</strong> - an assistant turn can now carry <code>reasoning</code> and <code>tool_calls</code> together and keep the reasoning across the tool-result loop, with a <code>reasoning_content</code> inbound alias and Anthropic <code>thinking</code> block support.</li>
<li>🗣️ <strong>New audio engines</strong> - one-pass diarized transcription (<code>moss-transcribe-cpp</code>), F5-TTS voice cloning in CrispASR, and true streaming TTS in vibevoice-cpp (time-to-first-audio 2.38s vs 39.96s on CPU).</li>
<li>🎚️ <strong>Auto full context</strong> - <code>context_size: -1</code> runs a model at its full trained context window, read from GGUF metadata per-model, with a VRAM-fit warning.</li>
<li>🖥️ <strong>Sharper control</strong> - pick which GPUs llama.cpp offloads to (<code>devices:</code>), and a new model-load cooldown stops a deterministically-failing model from respawning its backend on every poll and leaking VRAM.</li>
</ul>
<p>Plus DFlash speculative-decoding gallery models, an OIDC fix for EC/PS/EdDSA-signed tokens, transcription language/translate settings that reach the backend, and the usual set of dependency updates.</p>
<p align="center">
<em><a target="_blank" rel="noopener noreferrer" href="https://private-user-images.githubusercontent.com/2420543/621494355-7a114654-5729-45b9-8437-22af7ab76e4f.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODYzMTk4ODgsIm5iZiI6MTc4NjMxOTU4OCwicGF0aCI6Ii8yNDIwNTQzLzYyMTQ5NDM1NS03YTExNDY1NC01NzI5LTQ1YjktODQzNy0yMmFmN2FiNzZlNGYucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MDgwOSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjA4MDlUMjM1MzA4WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NDJkMzdiNTI4YTQwOGJiM2I5OWQ2MWE4ZWE3MWUxN2Y1NmE3YzQxMjkwZWYzNmY3MGIzNzhkNTIxOTE0OGI1MiZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QmcmVzcG9uc2UtY29udGVudC10eXBlPWltYWdlJTJGcG5nIn0.noUxjT2tPHtgiqbCHXQ7nPVKz7vMM9DXy9vupk7BzYA"><img width="2000" height="1250" alt="ui-home" src="https://private-user-images.githubusercontent.com/2420543/621494355-7a114654-5729-45b9-8437-22af7ab76e4f.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODYzMTk4ODgsIm5iZiI6MTc4NjMxOTU4OCwicGF0aCI6Ii8yNDIwNTQzLzYyMTQ5NDM1NS03YTExNDY1NC01NzI5LTQ1YjktODQzNy0yMmFmN2FiNzZlNGYucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MDgwOSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjA4MDlUMjM1MzA4WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NDJkMzdiNTI4YTQwOGJiM2I5OWQ2MWE4ZWE3MWUxN2Y1NmE3YzQxMjkwZWYzNmY3MGIzNzhkNTIxOTE0OGI1MiZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QmcmVzcG9uc2UtY29udGVudC10eXBlPWltYWdlJTJGcG5nIn0.noUxjT2tPHtgiqbCHXQ7nPVKz7vMM9DXy9vupk7BzYA" content-type-secured-asset="image/png" style="max-width: 100%; height: auto; max-height: 1250px;"></a></em>
</p>
<hr>
<h2>📌 TL;DR</h2>
<table>
<thead>
<tr>
<th>Area</th>
<th>Summary</th>
</tr>
</thead>
<tbody>
<tr>
<td>🎙️ <strong>Voice Library</strong></td>
<td>Admin-managed voice cloning profiles: record/upload consented reference audio in the UI, preview, and reference via <code>localai://voice-profiles/<id></code>. New <code>GET/POST/DELETE /api/voice-profiles</code>, MCP tools (<code>list/create/delete_voice_profile</code>), and a typed <code>tts.voice_cloning</code> config. Cloning declared as a capability across 12+ TTS backends (self-discovered, no hardcoded backend names).</td>
</tr>
<tr>
<td>🎬 <strong>LongCat video & avatars</strong></td>
<td>New <code>longcat-video</code> Python backend: text-to-video, image-to-video, and LongCat-Video-Avatar 1.5 audio-driven avatars. Gallery entries <code>longcat-video</code> and <code>longcat-video-avatar-1.5</code>; new <code>known_input_modalities</code>/<code>known_output_modalities</code> config fields; <code>/video</code> endpoint extended with staged audio. CUDA 12/13 x86_64 + CUDA 13 ARM64 images.</td>
</tr>
<tr>
<td>🧠 <strong>Interleaved thinking</strong></td>
<td>Assistant turns carry <code>reasoning</code> + <code>tool_calls</code> together and preserve reasoning across the tool loop. <code>reasoning_content</code> accepted as an inbound alias; Anthropic Messages local path round-trips <code>thinking</code> blocks (gated on <code>thinking: {type: "enabled"}</code>).</td>
</tr>
<tr>
<td>🗣️ <strong>moss-transcribe-cpp</strong></td>
<td>New Go backend over the C++/ggml MOSS-Transcribe-Diarize port: joint multi-speaker transcription + diarization + timestamps in one pass. Gallery model <code>moss-transcribe-cpp-0.9b</code>. Offline; 1.6-2.2x faster than reference on CPU, bit-exact on CUDA.</td>
</tr>
<tr>
<td>⚡ <strong>vibevoice streaming TTS</strong></td>
<td>Real incremental streaming via <code>vv_capi_tts_stream</code>: TTFA 2.38s vs 39.96s batch (~17x) on CPU for <code>VibeVoice-Realtime-0.5B</code>.</td>
</tr>
<tr>
<td>🎨 <strong>F5-TTS</strong></td>
<td>F5-TTS linked into CrispASR + <code>f5-tts-crispasr</code> gallery model (24kHz, voice cloning via <code>voice:</code>/<code>voice_text:</code> options).</td>
</tr>
<tr>
<td>🎚️ <strong>context_size: -1</strong></td>
<td>Negative <code>context_size</code> (or <code>LOCALAI_CONTEXT_SIZE=-1</code>) resolves to the model's <code>n_ctx_train</code> from GGUF metadata, with a GPU-only VRAM-fit warning and a safe clamp so no backend ever sees a negative window.</td>
</tr>
<tr>
<td>🖥️ <strong>llama.cpp device selection</strong></td>
<td><code>options: [devices:CUDA1,CUDA2]</code> restricts offload to named GPUs (from <code>--list-devices</code>).</td>
</tr>
<tr>
<td>🛡️ <strong>Model-load cooldown</strong></td>
<td>A failed load enters a cooldown (default <code>10s</code>, geometric growth capped at 5m) so polling clients get <code>503 + Retry-After</code> instead of respawning a crashing backend and leaking GPU memory.</td>
</tr>
<tr>
<td>🧠 <strong>Speculative decoding models</strong></td>
<td>Four Qwen DFlash gallery entries (4B / 9B / 27B / 35B-A3B), each bundling target + drafter, <code>spec_type:draft-dflash</code>.</td>
</tr>
</tbody>
</table>
<hr>
<h2>🚀 New Features & Major Enhancements</h2>
<h3>🎙️ Managed voice cloning profiles (Voice Library)</h3>
<p align="center">
<em><a target="_blank" rel="noopener noreferrer" href="https://private-user-images.githubusercontent.com/2420543/621494591-ea99bb2d-42b5-42b8-8c33-761090290e3f.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODYzMTk4ODgsIm5iZiI6MTc4NjMxOTU4OCwicGF0aCI6Ii8yNDIwNTQzLzYyMTQ5NDU5MS1lYTk5YmIyZC00MmI1LTQyYjgtOGMzMy03NjEwOTAyOTBlM2YucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MDgwOSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjA4MDlUMjM1MzA4WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NjU3NmM0NGQzZGIzNDUxMTBiZDY3YzkwMGQ2ODc1MzcyNjk4MjY1ODM0MGU1YTgzNWExM2UwNjRiMzEyMmY0OCZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QmcmVzcG9uc2UtY29udGVudC10eXBlPWltYWdlJTJGcG5nIn0.y99qokhANeyChyKMVCuavtdA0e45C4eGQzXSp9mrvuc"><img width="2000" height="1250" alt="ui-voice-library" src="https://private-user-images.githubusercontent.com/2420543/621494591-ea99bb2d-42b5-42b8-8c33-761090290e3f.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODYzMTk4ODgsIm5iZiI6MTc4NjMxOTU4OCwicGF0aCI6Ii8yNDIwNTQzLzYyMTQ5NDU5MS1lYTk5YmIyZC00MmI1LTQyYjgtOGMzMy03NjEwOTAyOTBlM2YucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI2MDgwOSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNjA4MDlUMjM1MzA4WiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NjU3NmM0NGQzZGIzNDUxMTBiZDY3YzkwMGQ2ODc1MzcyNjk4MjY1ODM0MGU1YTgzNWExM2UwNjRiMzEyMmY0OCZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QmcmVzcG9uc2UtY29udGVudC10eXBlPWltYWdlJTJGcG5nIn0.y99qokhANeyChyKMVCuavtdA0e45C4eGQzXSp9mrvuc" content-type-secured-asset="image/png" style="max-width: 100%; height: auto; max-height: 1250px;"></a></em>
<br><em>Record or upload a consented reference, save it as a reusable profile, and reference it from any cloning-capable TTS backend.</em>
</p>
<p>Voice cloning becomes a first-class, UI-driven workflow. Record or upload consented reference audio in the React UI, normalize it to WAV, preview it, and save it as a named profile. Reference the profile from the TTS UI or <code>/v1/audio/speech</code> with a stable <code>localai://voice-profiles/<id></code> URI, no YAML editing and no copying audio into model directories.</p>
<ul>
<li>New REST surface: <code>GET /api/voice-profiles</code>, <code>GET /api/voice-profiles/:id/audio</code>, <code>POST /api/voice-profiles</code> (admin), <code>DELETE /api/voice-profiles/:id</code> (admin), surfaced in Swagger, <code>/api/instructions</code>, and the auth capability registry. MCP admin tools <code>list_voice_profiles</code>, <code>create_voice_profile</code>, <code>delete_voice_profile</code>.</li>
<li>New typed config <code>tts.voice_cloning</code> (bool) opts custom model names in or out of Voice Library compatibility; <code>false</code> also rejects saved profile references with HTTP 400. Request precedence is <code>voice</code> -> <code>tts.voice</code> -> <code>tts.audio_path</code>; existing options still work with no breaking change.</li>
<li>Compatibility is server-discovered from backend capabilities (the frontend hardcodes no backend names) and offers gallery models to install when none are present. Cloning is declared in the capability registry for <code>vllm-omni</code>, <code>vibevoice-cpp</code>, <code>coqui</code>, <code>pocket-tts</code>, <code>qwen-tts</code>, <code>qwen3-tts-cpp</code>, <code>faster-qwen3-tts</code>, <code>fish-speech</code>, <code>neutts</code>, <code>chatterbox</code>, <code>voxcpm</code>, <code>omnivoice-cpp</code>, and model-dependent <code>crispasr</code>.</li>
<li>Cross-backend contract: reference WAV + exact transcript via <code>params.ref_text</code>. E2E verified on qwen3-tts-cpp (4.47s reference produced a non-silent 3.44s 24kHz WAV).</li>
</ul>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4869267397" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10799" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10799/hovercard" href="https://github.com/mudler/LocalAI/pull/10799">#10799</a></p>
</blockquote>
<h3>🎬 LongCat video and avatar generation</h3>
<p>A new <code>longcat-video</code> Python backend brings local video generation to LocalAI: text-to-video, image-to-video, and LongCat-Video-Avatar 1.5 for audio-driven talking avatars. It is wired into the React Studio UI with audio and reference-image controls.</p>
<ul>
<li>Gallery entries <code>longcat-video</code> (text+image input, video output) and <code>longcat-video-avatar-1.5</code> (text+image+audio input, video output). Backend images: CUDA 12 and CUDA 13 x86_64, plus CUDA 13 ARM64 (DGX Spark / NVIDIA ARM64 guidance included).</li>
<li>Introduces declarative capability metadata: new <code>known_input_modalities</code> / <code>known_output_modalities</code> config fields so generic code discovers what a checkpoint accepts instead of branching on backend or checkpoint names. The HF importer emits matching self-describing recipe metadata.</li>
<li>The <code>/video</code> endpoint gains staged audio and video-generation parameters, with distributed audio staging (input bounded to 128 MiB). Avatar mode supports multi-segment generation, BF16 or optional INT8 quantization, and distillation settings, documented in a new <code>features/longcat-video.md</code> page.</li>
</ul>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4868172117" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10792" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10792/hovercard" href="https://github.com/mudler/LocalAI/pull/10792">#10792</a></p>
</blockquote>
<h3>🧠 Interleaved thinking with tool calls</h3>
<p>An assistant turn can now carry <code>reasoning</code> and <code>tool_calls</code> together, and the reasoning survives the tool-result loop across turns. This is uniform, tested, and documented behavior rather than a per-backend accident.</p>
<ul>
<li>OpenAI chat messages accept <code>reasoning_content</code> as an inbound alias for the canonical <code>reasoning</code> field (vLLM/DeepSeek/cogito-style clients emit it); emission is unchanged and the canonical field wins when both are present.</li>
<li>The Anthropic Messages local path round-trips <code>thinking</code> blocks: inbound <code>thinking</code> blocks parse into the message reasoning, and a <code>thinking</code> block is emitted before <code>tool_use</code> on both non-streaming and streaming responses, gated on the request param <code>thinking: {type: "enabled"}</code>. The cloud-proxy passthrough path is untouched.</li>
<li>Verified live: <code>lfm2.5-8b-a1b</code> and <code>gemma-4-e2b</code> return <code>reasoning</code> plus structured <code>tool_calls</code> in a single turn. A new <code>features/interleaved-thinking.md</code> doc is cross-linked from the model-configuration, text-generation, and functions guides.</li>
</ul>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4839297478" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10744" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10744/hovercard" href="https://github.com/mudler/LocalAI/pull/10744">#10744</a></p>
</blockquote>
<h3>🗣️ New audio engines: diarized transcription, F5-TTS, and streaming TTS</h3>
<p>Three additions widen the audio surface:</p>
<ul>
<li><strong>moss-transcribe-cpp</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4846804726" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10756" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10756/hovercard" href="https://github.com/mudler/LocalAI/pull/10756">#10756</a>): a new Go backend that dlopens the C++/ggml MOSS-Transcribe-Diarize port to do joint multi-speaker transcription, diarization, and timestamps in a single offline pass. Gallery model <code>moss-transcribe-cpp-0.9b</code> (default <code>q5_k</code> GGUF). The ggml port is byte-exact to the reference PyTorch and 1.6-2.2x faster on CPU, bit-exact on CUDA (verified on Blackwell / Jetson Thor). Builds across the full Linux matrix plus Darwin/Metal.</li>
<li><strong>F5-TTS in CrispASR</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4844315058" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10753" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10753/hovercard" href="https://github.com/mudler/LocalAI/pull/10753">#10753</a>): links the F5-TTS static runtime into the CrispASR build (SWivid, MIT; a 22-layer DiT flow-matching model with a built-in Vocos vocoder) and ships an <code>f5-tts-crispasr</code> gallery model. Produces 24kHz mono audio and auto-detects as <code>f5-tts</code> (no <code>backend:</code> selector). Voice cloning via <code>options: [voice:/path/ref.wav, voice_text:Transcript...]</code>; note F5-TTS runs a 32-step ODE solver, so CPU synthesis is compute-heavy.</li>
<li><strong>vibevoice-cpp true streaming</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4849938784" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10764" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10764/hovercard" href="https://github.com/mudler/LocalAI/pull/10764">#10764</a>): replaces whole-clip-then-chunk synthesis with real incremental streaming through the new <code>vv_capi_tts_stream</code> callback ABI. Time-to-first-audio drops from 39.96s (batch) to 2.38s (streaming), about 17x, on a CPU-only box for <code>VibeVoice-Realtime-0.5B</code>. Scope is the realtime-0.5B model; the non-streaming path is unchanged.</li>
</ul>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4846804726" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10756" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10756/hovercard" href="https://github.com/mudler/LocalAI/pull/10756">#10756</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4844315058" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10753" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10753/hovercard" href="https://github.com/mudler/LocalAI/pull/10753">#10753</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4849938784" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10764" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10764/hovercard" href="https://github.com/mudler/LocalAI/pull/10764">#10764</a></p>
</blockquote>
<h3>🎚️ Auto full context with <code>context_size: -1</code></h3>
<p>A <code>context_size: -1</code> sentinel (any negative value) now makes a model run at its full trained context, resolved per-model from the GGUF <code>n_ctx_train</code> metadata at load. This lets you opt a model into its true maximum window even when a gallery YAML already pins a value, without hardcoding a number. The global equivalents <code>LOCALAI_CONTEXT_SIZE=-1</code> / <code>--context-size -1</code> make every model resolve to its own trained max, while an explicit per-model value still wins.</p>
<p>Three defense layers keep it safe: GGUF resolution degrades to the default with a warning if metadata lacks a usable max; a GPU-only <code>warnIfContextExceedsVRAM</code> helper logs (never blocks) when the window likely will not fit; and the backend options layer clamps any residual negative to the default so no backend receives a negative <code>n_ctx</code>. The previously-only path (unset <code>context_size</code>) is unchanged.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4843960411" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10752" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10752/hovercard" href="https://github.com/mudler/LocalAI/pull/10752">#10752</a></p>
</blockquote>
<h3>🖥️ GPU device selection and model-load cooldown</h3>
<ul>
<li><strong>llama.cpp device selection</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4830504384" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10724" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10724/hovercard" href="https://github.com/mudler/LocalAI/pull/10724">#10724</a>): a new <code>device</code> / <code>devices</code> option in the llama.cpp <code>options:</code> array maps to upstream <code>--device</code>, so you can restrict offload to specific GPUs (for example excluding a display or debug GPU). Example: <code>options: [devices:CUDA1,CUDA2,CUDA3]</code>. Device names come from <code>llama-server --list-devices</code>.</li>
<li><strong>Model-load failure cooldown</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4831959076" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10728" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10728/hovercard" href="https://github.com/mudler/LocalAI/pull/10728">#10728</a>): after a load fails, new independent load triggers are refused during a cooldown window, returning a typed error mapped to <code>503 + Retry-After</code> instead of respawning the backend on every poll. This stops a deterministically-failing model from leaking GPU/CUDA state (a reporter saw leaked contexts climb to ~58 GB) and, under <code>LOCALAI_SINGLE_ACTIVE_BACKEND</code>, from stealing the active slot from healthy models. Configurable via <code>--model-load-failure-cooldown</code> / <code>LOCALAI_MODEL_LOAD_FAILURE_COOLDOWN</code> (default <code>10s</code>, <code>0</code> disables); the cooldown doubles per consecutive failure, caps at 5m, and resets on a successful load. Coalesced followers of a genuine concurrent burst still get their one retry.</li>
</ul>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4830504384" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10724" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10724/hovercard" href="https://github.com/mudler/LocalAI/pull/10724">#10724</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4831959076" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10728" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10728/hovercard" href="https://github.com/mudler/LocalAI/pull/10728">#10728</a></p>
</blockquote>
<hr>
<h2>🧠 Models</h2>
<ul>
<li><strong>Qwen DFlash speculative-decoding models</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4866104642" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10791" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10791/hovercard" href="https://github.com/mudler/LocalAI/pull/10791">#10791</a>): four gallery entries for the <code>llama-cpp</code> backend, each bundling a full target model plus its small block-diffusion drafter (drafters are not standalone chat models): <code>qwen3-4b-dflash</code>, <code>qwen3.5-9b-dflash</code>, <code>qwen3.6-27b-dflash</code>, <code>qwen3.6-35b-a3b-dflash</code>. Per-entry config is <code>flash_attention: on</code>, <code>draft_model:</code>, <code>use_jinja:true</code>, <code>spec_type:draft-dflash</code>, <code>spec_n_max:15</code>. Requires the pinned llama.cpp with upstream DFlash support; every bundled drafter is verified <code>general.architecture = dflash</code> (fork-only <code>dflash-draft</code> GGUFs are intentionally excluded because they fail to load). GPU recommended.</li>
<li><strong>Inference defaults</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4835040638" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10741" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10741/hovercard" href="https://github.com/mudler/LocalAI/pull/10741">#10741</a>): adds recommended sampling defaults for <code>deepseek-v4</code> (auto-generated from unsloth), so LocalAI applies correct generation parameters for the family automatically.</li>
<li>Plus new gallery models added via the gallery agent (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4837544982" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10743" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10743/hovercard" href="https://github.com/mudler/LocalAI/pull/10743">#10743</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4846535465" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10755" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10755/hovercard" href="https://github.com/mudler/LocalAI/pull/10755">#10755</a>).</li>
</ul>
<hr>
<h2>🐛 Bug Fixes (recap)</h2>
<ul>
<li><code>fix(auth)</code>: accept EC/PS/EdDSA-signed OIDC ID tokens, not just RS256 (OIDC login was 500ing at callback with EC-signed tokens, e.g. Authentik) - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4832382474" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10736" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10736/hovercard" href="https://github.com/mudler/LocalAI/pull/10736">#10736</a></li>
<li><code>fix(transcription)</code>: honor model-config <code>parameters.language</code>/<code>parameters.translate</code> and the OpenAI <code>language</code> form field, which were silently ignored for multipart uploads - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4832101655" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10731" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10731/hovercard" href="https://github.com/mudler/LocalAI/pull/10731">#10731</a></li>
<li><code>fix(backends)</code>: <code>opus</code> and <code>local-store</code> now refuse foreign model loads, so an LLM with no explicit backend cannot silently bind to the audio codec or vector store during backend probing - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4856015932" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10769" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10769/hovercard" href="https://github.com/mudler/LocalAI/pull/10769">#10769</a></li>
<li><code>fix(gallery)</code>: backend (re)install is now a clean atomic replace (stage, validate, swap, rollback) instead of an overlay, so stale files from a prior version no longer shadow the new one - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4831787233" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10726" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10726/hovercard" href="https://github.com/mudler/LocalAI/pull/10726">#10726</a></li>
<li><code>fix(vram)</code>: report the largest single GGUF quant instead of summing every quant in an HF repo, so a 9B model no longer shows as 71 GB / "May not fit" - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4822343105" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10707" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10707/hovercard" href="https://github.com/mudler/LocalAI/pull/10707">#10707</a></li>
<li><code>fix(logs)</code>: capture backend stdout/stderr into the log store by default in single mode, so the Backend Logs page is populated out of the box - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4837181801" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10742" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10742/hovercard" href="https://github.com/mudler/LocalAI/pull/10742">#10742</a></li>
<li><code>fix(vllm)</code>: pin the L4T arm64 backend to <code>vllm==0.24.0</code> for GB10 / DGX Spark stability (0.23 crashes deterministically on cold loads and pins GPU memory) - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4831785267" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10725" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10725/hovercard" href="https://github.com/mudler/LocalAI/pull/10725">#10725</a></li>
<li><code>fix(ds4)</code>: bundle the full transitive runtime dependency closure so the from-scratch DS4 image no longer exits 127 on a missing gRPC library - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4863448660" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10783" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10783/hovercard" href="https://github.com/mudler/LocalAI/pull/10783">#10783</a></li>
<li><code>fix(diffusers,vllm-omni,tinygrad)</code>: save generated images as PNG explicitly, fixing an <code>unknown file extension: .tmp</code> crash right after successful inference - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4832028621" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10729" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10729/hovercard" href="https://github.com/mudler/LocalAI/pull/10729">#10729</a></li>
<li><code>fix(react-ui)</code>: preserve uploaded file content when regenerating a non-last answer (attachments were silently dropped after forking and regenerating a file turn) - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4879846829" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10819" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10819/hovercard" href="https://github.com/mudler/LocalAI/pull/10819">#10819</a></li>
<li><code>fix(ui)</code>: prevent a large data table from breaking the flexbox layout and forcing a page-wide horizontal scrollbar - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4844834440" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10754" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10754/hovercard" href="https://github.com/mudler/LocalAI/pull/10754">#10754</a></li>
</ul>
<hr>
<h2>👒 Dependencies</h2>
<p>Submodule and backend bumps this cycle:</p>
<ul>
<li><code>ggml-org/llama.cpp</code> x7</li>
<li><code>CrispStrobe/CrispASR</code> x7</li>
<li><code>vllm-metal</code> (darwin) x6</li>
<li><code>ServeurpersoCom/qwentts.cpp</code> x4</li>
<li><code>ServeurpersoCom/omnivoice.cpp</code> x4</li>
<li><code>leejet/stable-diffusion.cpp</code> x3</li>
<li><code>ikawrakow/ik_llama.cpp</code> x3</li>
<li><code>ggml-org/whisper.cpp</code> x2</li>
<li><code>mudler/moss-transcribe.cpp</code> x1</li>
<li><code>mudler/locate-anything.cpp</code> x1</li>
<li><code>vllm-project/vllm</code> 0.24.0 -> 0.25.0 (Python backend) and cu130 wheel to <code>0.25.0</code></li>
</ul>
<p>Plus <code>grpcio</code> 1.81.1 -> 1.82.1, <code>charset-normalizer</code> >=3.4.9, and GitHub Actions bumps (<code>actions/cache</code> 4 -> 6, <code>actions/stale</code> 10.3.0 -> 10.4.0).</p>
<hr>
<h2>📖 Documentation</h2>
<ul>
<li>Refreshed the LocalAI homepage to frame the project as a modular multimodal AI runtime: breadth lanes (reason, listen and speak, create, see, act), the small-core plus on-demand-backends architecture, the native engines built by the LocalAI team, and the scale path from laptop to team server to cluster - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4859632833" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10780" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10780/hovercard" href="https://github.com/mudler/LocalAI/pull/10780">#10780</a></li>
<li>Docs-site version bump - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4822941114" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10709" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10709/hovercard" href="https://github.com/mudler/LocalAI/pull/10709">#10709</a></li>
</ul>
<hr>
<h2>🙌 New Contributors</h2>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/rvmz/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/rvmz">@rvmz</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4830504384" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10724" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10724/hovercard" href="https://github.com/mudler/LocalAI/pull/10724">#10724</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hogeheer499-commits/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hogeheer499-commits">@hogeheer499-commits</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4863448660" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10783" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10783/hovercard" href="https://github.com/mudler/LocalAI/pull/10783">#10783</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ajuijas/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ajuijas">@ajuijas</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4879846829" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10819" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10819/hovercard" href="https://github.com/mudler/LocalAI/pull/10819">#10819</a></li>
</ul>
<hr>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.6.2...v4.7.0"><tt>v4.6.2...v4.7.0</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.6.22026-07-06T19:52:52Zv4.6.2
<h2>What's Changed</h2>
<h3>👒 Dependencies</h3>
<ul>
<li>chore(deps): bump actions/cache from 4 to 6 by <a class="user-mention notranslate" data-hovercard-type="organization" data-hovercard-url="/orgs/dependabot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dependabot">@dependabot</a>[bot] in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4822038779" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10704" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10704/hovercard" href="https://github.com/mudler/LocalAI/pull/10704">#10704</a></li>
</ul>
<h3>Other Changes</h3>
<ul>
<li>refactor: use slices.Contains to simplify code by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/weifanglab/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/weifanglab">@weifanglab</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4821100201" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10702" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10702/hovercard" href="https://github.com/mudler/LocalAI/pull/10702">#10702</a></li>
<li>fix(ci): shard single-arch backend matrix under GitHub's 256-job limit by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4821701859" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10703" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10703/hovercard" href="https://github.com/mudler/LocalAI/pull/10703">#10703</a></li>
<li>chore(model gallery): add MiniCPM series models by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ZMXJJ/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ZMXJJ">@ZMXJJ</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4819773931" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10699" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10699/hovercard" href="https://github.com/mudler/LocalAI/pull/10699">#10699</a></li>
</ul>
<h2>New Contributors</h2>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/weifanglab/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/weifanglab">@weifanglab</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4821100201" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10702" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10702/hovercard" href="https://github.com/mudler/LocalAI/pull/10702">#10702</a></li>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ZMXJJ/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ZMXJJ">@ZMXJJ</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4819773931" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10699" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10699/hovercard" href="https://github.com/mudler/LocalAI/pull/10699">#10699</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.6.1...v4.6.2"><tt>v4.6.1...v4.6.2</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.6.12026-07-06T11:13:39Zv4.6.1
<h2>What's Changed</h2>
<h3>Other Changes</h3>
<ul>
<li>fix(auth): log the real cause of OIDC/OAuth user-info failures by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4809937540" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10679" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10679/hovercard" href="https://github.com/mudler/LocalAI/pull/10679">#10679</a></li>
<li>feat(api): add GET /v1/models/capabilities endpoint by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4810754624" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10687" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10687/hovercard" href="https://github.com/mudler/LocalAI/pull/10687">#10687</a></li>
<li>chore: ⬆️ Update CrispStrobe/CrispASR to <code>1109cb3fcae2e242c2b3d42ec0e3fd6e813f2ce7</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4810438436" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10685" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10685/hovercard" href="https://github.com/mudler/LocalAI/pull/10685">#10685</a></li>
<li>chore: ⬆️ Update ggml-org/llama.cpp to <code>665892536dfb1b7532161e3182304bd35c33e768</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4810438309" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10681" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10681/hovercard" href="https://github.com/mudler/LocalAI/pull/10681">#10681</a></li>
<li>chore(model-gallery): ⬆️ update checksum by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4810506143" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10686" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10686/hovercard" href="https://github.com/mudler/LocalAI/pull/10686">#10686</a></li>
<li>chore: ⬆️ Update vllm-metal (darwin) to <code>v0.3.0.dev20260704102955</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4806151058" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10668" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10668/hovercard" href="https://github.com/mudler/LocalAI/pull/10668">#10668</a></li>
<li>fix(ui): center the home empty-state wizard by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4812552375" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10691" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10691/hovercard" href="https://github.com/mudler/LocalAI/pull/10691">#10691</a></li>
<li>docs: ⬆️ update docs version mudler/LocalAI by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4810430436" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10680" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10680/hovercard" href="https://github.com/mudler/LocalAI/pull/10680">#10680</a></li>
<li>chore: ⬆️ Update CrispStrobe/CrispASR to <code>09df654e304947f7521e1f52992ceacccf03c300</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4814326888" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10693" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10693/hovercard" href="https://github.com/mudler/LocalAI/pull/10693">#10693</a></li>
<li>chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to <code>daedb763fd442e0916eb130a479fdd74947291c0</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4810438326" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10682" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10682/hovercard" href="https://github.com/mudler/LocalAI/pull/10682">#10682</a></li>
<li>chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to <code>73fe0c67bbf0898ba2999535e0680a02a7f8537d</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4810438334" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10683" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10683/hovercard" href="https://github.com/mudler/LocalAI/pull/10683">#10683</a></li>
<li>feat(agents): native Prometheus metrics for agent chat runs by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/walcz-de/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/walcz-de">@walcz-de</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4812089952" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10689" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10689/hovercard" href="https://github.com/mudler/LocalAI/pull/10689">#10689</a></li>
<li>chore: ⬆️ Update ggml-org/llama.cpp to <code>2da668617612d2df773f966e3b0ee22dc2beef7b</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4814326991" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10694" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10694/hovercard" href="https://github.com/mudler/LocalAI/pull/10694">#10694</a></li>
<li>fix(reasoning): don't persist request-scoped reasoning_effort as an operator disable (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4781088063" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10622" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10622/hovercard" href="https://github.com/mudler/LocalAI/issues/10622">#10622</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Anai-Guo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Anai-Guo">@Anai-Guo</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4782389591" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10623" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10623/hovercard" href="https://github.com/mudler/LocalAI/pull/10623">#10623</a></li>
<li>fix(startup): scope generated-content and upload dirs to the current user by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4818372591" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10698" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10698/hovercard" href="https://github.com/mudler/LocalAI/pull/10698">#10698</a></li>
<li>fix(config): cap auto-derived context to fit VRAM by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4814928998" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10696" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10696/hovercard" href="https://github.com/mudler/LocalAI/pull/10696">#10696</a></li>
<li>fix(llama-cpp): cap single-pass embedding batch to fit VRAM by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4814901970" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10695" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10695/hovercard" href="https://github.com/mudler/LocalAI/pull/10695">#10695</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.6.0...v4.6.1"><tt>v4.6.0...v4.6.1</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.6.02026-07-04T07:29:37Zv4.6.0<h1>🎉 LocalAI 4.6.0 Release! 🚀</h1>
<h1 align="center">
<br>
<a target="_blank" rel="noopener noreferrer nofollow" href="https://raw.githubusercontent.com/mudler/LocalAI/refs/heads/master/core/http/static/logo.png"><img height="300" src="https://raw.githubusercontent.com/mudler/LocalAI/refs/heads/master/core/http/static/logo.png" style="max-width: 100%; height: auto; max-height: 300px;"></a>
<br>
<br>
</h1>
<p>LocalAI 4.6.0 is out!</p>
<p>This is a reliability-focused release: AMD ROCm backends now run on-GPU at full speed, distributed model loads no longer wedge when a worker dies, and realtime sessions warm up predictably. It also brings conversation forking to the built-in chat UI, a Prometheus counter for PII/audit events, and an SSRF fix for the model gallery.</p>
<p><strong>Highlights:</strong></p>
<ul>
<li>🔴 <strong>AMD ROCm runs correctly</strong> - ggml audio backends offload to the GPU, hipBLASLt kernel-tuning data is bundled (no more slow generic kernels), <code>rocm-vllm</code> installs the right wheel, and the ASIC ID table is found.</li>
<li>🎙️ <strong>Predictable realtime</strong> - sessions eagerly warm the whole pipeline (VAD, ASR, LLM, TTS) up front, so the first turn no longer pays per-model cold-start stalls, plus a new <code>POST /backend/load</code> API and "Load into memory" UI button.</li>
<li>🌿 <strong>Forking chat</strong> - retry <em>any</em> assistant answer, branch a new chat from any point, duplicate, or copy the whole conversation, directly in the built-in UI.</li>
<li>🛡️ <strong>Distributed hardening</strong> - a dead worker can no longer pin the model-load advisory lock (the ~15-minute wedge is gone), and orphaned backend workers self-terminate instead of holding VRAM.</li>
<li>📊 <strong>PII/audit metrics</strong> - PII detections/masks/blocks are exported as a Prometheus counter, so you can alert when the filter stops firing.</li>
<li>🔒 <strong>Gallery SSRF fix</strong> - <code>POST /models/apply</code> config-URL fetches are validated against private/loopback/metadata addresses.</li>
</ul>
<p>Plus idempotent backend installs, tool-calling and reasoning fixes across the vLLM and Python/MLX backends, cloud-proxy compatibility with the newest reasoning models, and the usual set of dependency updates.</p>
<hr>
<h2>📌 TL;DR</h2>
<table>
<thead>
<tr>
<th>Area</th>
<th>Summary</th>
</tr>
</thead>
<tbody>
<tr>
<td>🔴 <strong>AMD ROCm reliability</strong></td>
<td>ggml audio backends now compile with <code>-DGGML_HIP=ON</code> and link HIP (real GPU offload); hipBLASLt <code>TensileLibrary</code> data bundled + <code>HIPBLASLT_TENSILE_LIBPATH</code> exported; <code>rocm-vllm</code> installs from the AMD wheel index on Python 3.12; <code>amdgpu.ids</code> symlinked so the ASIC table is found.</td>
</tr>
<tr>
<td>🎙️ <strong>Realtime warm-up + load API</strong></td>
<td>Sessions block-warm the full pipeline at start (errors surface up front); new <code>POST /backend/load</code> / <code>POST /v1/backend/load</code>, a "Load into memory" UI action, and a <code>load_model</code> MCP tool. Opt out per pipeline with <code>disable_warmup: true</code>.</td>
</tr>
<tr>
<td>🌿 <strong>Forking chat</strong></td>
<td>Regenerate any assistant answer (not just the last), branch a new chat from any turn, duplicate a chat, or copy it as Markdown - all client-side in the React UI.</td>
</tr>
<tr>
<td>🛡️ <strong>Process & distributed lifecycle</strong></td>
<td>A dead worker no longer pins the per-model PostgreSQL advisory lock (bounded load ceiling + context-scoped <code>lock_timeout</code>); backend workers self-terminate on parent death (<code>LOCALAI_BACKEND_PARENT_WATCH</code>); the watchdog stops logging optional <code>Free()</code> as an error.</td>
</tr>
<tr>
<td>⚙️ <strong>Idempotent backend installs</strong></td>
<td><code>POST /backends/apply</code> and the <code>LOCALAI_EXTERNAL_BACKENDS</code> boot loop no longer re-pull an already-installed backend unless <code>force: true</code>.</td>
</tr>
<tr>
<td>📊 <strong>PII/audit Prometheus counter</strong></td>
<td><code>localai_pii_events_total{kind,origin,action,direction}</code> on <code>/metrics</code>, complementing the <code>/api/pii/events</code> ring buffer.</td>
</tr>
<tr>
<td>🔒 <strong>Gallery SSRF hardening</strong></td>
<td>Gallery config URL fetches run through <code>ValidateExternalURL</code>, blocking private, loopback, link-local, and cloud-metadata addresses.</td>
</tr>
<tr>
<td>🧩 <strong>Tool-calling & reasoning fixes</strong></td>
<td>Non-streaming vLLM tool calls restored; MLX/Python backends decode tool-call arguments for chat templates and split closing-only <code></think></code> reasoning blocks.</td>
</tr>
</tbody>
</table>
<hr>
<h2>🚀 New Features & Major Enhancements</h2>
<h3>🔴 AMD ROCm backends run correctly on-GPU</h3>
<p>Four coupled fixes make ROCm/hipBLAS backends actually run on AMD hardware, and at full speed, instead of silently falling back to CPU or slow generic kernels:</p>
<ul>
<li><strong>GPU offload for ggml audio backends</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4806150311" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10667" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10667/hovercard" href="https://github.com/mudler/LocalAI/pull/10667">#10667</a>): <code>rocm-qwen3-tts-cpp</code>, <code>rocm-omnivoice-cpp</code>, <code>acestep-cpp</code>, and <code>vibevoice-cpp</code> were building CPU-only because their Makefiles passed the no-op <code>-DGGML_HIPBLAS=ON</code> (upstream ggml only understands <code>-DGGML_HIP=ON</code>) and the CMake link loop omitted <code>hip</code>. They now use the same hipblas recipe as llama-cpp and link the HIP backend.</li>
<li><strong>hipBLASLt kernel-tuning data</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4802636104" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10660" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10660/hovercard" href="https://github.com/mudler/LocalAI/issues/10660">#10660</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4806162734" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10672" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10672/hovercard" href="https://github.com/mudler/LocalAI/pull/10672">#10672</a>): the packager bundled rocBLAS data but not the parallel hipBLASLt <code>TensileLibrary_lazy_gfx*.dat</code> files, so every arch silently used slow kernels and logged <code>Cannot read "TensileLibrary_lazy_gfx*.dat"</code>. The data is now bundled and <code>HIPBLASLT_TENSILE_LIBPATH</code> is exported by the <code>llama-cpp</code> and <code>turboquant</code> <code>run.sh</code>.</li>
<li><strong><code>rocm-vllm</code> installs the right wheel</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4793183869" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10642" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10642/hovercard" href="https://github.com/mudler/LocalAI/issues/10642">#10642</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4798033771" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10651" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10651/hovercard" href="https://github.com/mudler/LocalAI/pull/10651">#10651</a>): the backend was pulling the CUDA-only PyPI <code>vllm</code> (fatal <code>ModuleNotFoundError: No module named 'vllm'</code> on AMD). It now pins CPython 3.12 and installs vLLM from the ROCm wheel index (<code>https://wheels.vllm.ai/rocm/</code>).</li>
<li><strong>ASIC ID table found</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4784825695" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10624" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10624/hovercard" href="https://github.com/mudler/LocalAI/issues/10624">#10624</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4788848735" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10627" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10627/hovercard" href="https://github.com/mudler/LocalAI/pull/10627">#10627</a>): the compute-only hipblas image lacks <code>/opt/amdgpu/share/libdrm/amdgpu.ids</code>, so every model load warned. Ubuntu's <code>libdrm-common</code> copy is now symlinked into place.</li>
</ul>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4806150311" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10667" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10667/hovercard" href="https://github.com/mudler/LocalAI/pull/10667">#10667</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4806162734" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10672" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10672/hovercard" href="https://github.com/mudler/LocalAI/pull/10672">#10672</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4798033771" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10651" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10651/hovercard" href="https://github.com/mudler/LocalAI/pull/10651">#10651</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4788848735" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10627" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10627/hovercard" href="https://github.com/mudler/LocalAI/pull/10627">#10627</a></p>
</blockquote>
<h3>🎙️ Realtime: eager pipeline warm-up + a load-into-memory API</h3>
<p>Realtime voice sessions now eagerly and blockingly warm the entire pipeline (VAD, transcription, LLM, TTS, sound detection, voice recognition) at session start instead of lazy-loading each sub-model on first use. The first turn no longer pays per-model cold-start stalls, and model-load errors surface up front at session start (as <code>model_load_error</code>) rather than mid-stream. Pipeline sub-models load concurrently, so a session warms in the time of its slowest stage, not the sum, and a failed stage names every broken model in a joined error.</p>
<p>This also adds a LocalAI-native <code>POST /backend/load</code> (and <code>/v1/backend/load</code>), the inverse of <code>/backend/shutdown</code>, exposed as a "Load into memory" UI action and a <code>load_model</code> MCP admin tool, so admins can pre-warm any model (including full pipelines) on demand. The <code>--load-to-memory</code> startup flag now routes through the same engine. Opt out per pipeline with <code>disable_warmup: true</code>.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4803567951" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10662" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10662/hovercard" href="https://github.com/mudler/LocalAI/pull/10662">#10662</a></p>
</blockquote>
<h3>🌿 Forking chat in the built-in UI</h3>
<p>The React chat UI gains conversation-management tools: regenerate <em>any</em> assistant answer (not just the last), branch a new chat from any answer, duplicate a chat into an independent copy, or copy the whole conversation to the clipboard as Markdown. Retrying a mid-conversation answer correctly truncates the conversation before re-asking, both in the DOM and in the request payload (this also fixes a latent stale-closure bug where a mid-conversation retry sent the downstream turns back to the model). All client-side, no backend changes.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4798333546" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10654" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10654/hovercard" href="https://github.com/mudler/LocalAI/pull/10654">#10654</a></p>
</blockquote>
<h3>🛡️ Sturdier process and distributed lifecycle</h3>
<ul>
<li><strong>Dead-worker advisory-lock wedge</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4772168973" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10600" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10600/hovercard" href="https://github.com/mudler/LocalAI/pull/10600">#10600</a>): a distributed worker going mid-load could pin a per-model PostgreSQL advisory lock and fail every subsequent request to that model with <code>55P03</code> for ~15 minutes. The detached load context is now bounded by a model-load ceiling, the install wait honors cancellation via <code>singleflight.DoChan</code>, and <code>lock_timeout</code> is scoped to the caller's context budget instead of a deployment-global GUC.</li>
<li><strong>Parent-death safety net</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4789980830" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10639" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10639/hovercard" href="https://github.com/mudler/LocalAI/pull/10639">#10639</a>): if LocalAI is <code>SIGKILL</code>ed before teardown, spawned backend workers used to get reparented to init and linger, holding VRAM and their port. Each backend now polls its parent PID and self-terminates on reparenting. Configurable via <code>LOCALAI_BACKEND_PARENT_WATCH</code> (default on, auto-off on Windows) and <code>LOCALAI_BACKEND_PARENT_WATCH_INTERVAL</code> (default <code>2s</code>). C++ coverage is llama-cpp for now; Python covers all backends.</li>
<li><strong>Quieter watchdog</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4772362980" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10602" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10602/hovercard" href="https://github.com/mudler/LocalAI/issues/10602">#10602</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4776160035" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10607" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10607/hovercard" href="https://github.com/mudler/LocalAI/pull/10607">#10607</a>): the optional <code>Free()</code> RPC returns gRPC <code>Unimplemented</code> for many backends and the federation proxy, so the watchdog no longer logs a misleading <code>Error freeing GPU resources</code> on eviction. A new <code>grpcerrors.IsUnimplemented</code> helper distinguishes it from genuine failures.</li>
<li><strong>Idempotent backend installs</strong> (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4793623386" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10643" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10643/hovercard" href="https://github.com/mudler/LocalAI/pull/10643">#10643</a>): <code>POST /backends/apply</code> and the <code>LOCALAI_EXTERNAL_BACKENDS</code> boot loop no longer re-download and re-extract an already-installed backend on every apply/boot. Pass <code>"force": true</code> (the UI's install button still does, doubling as "Reinstall").</li>
</ul>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4772168973" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10600" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10600/hovercard" href="https://github.com/mudler/LocalAI/pull/10600">#10600</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4789980830" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10639" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10639/hovercard" href="https://github.com/mudler/LocalAI/pull/10639">#10639</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4776160035" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10607" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10607/hovercard" href="https://github.com/mudler/LocalAI/pull/10607">#10607</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4793623386" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10643" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10643/hovercard" href="https://github.com/mudler/LocalAI/pull/10643">#10643</a></p>
</blockquote>
<h3>📊 PII/audit events as a Prometheus counter</h3>
<p>The PII middleware / MITM audit pipeline now emits a single monotonic counter, <code>localai_pii_events_total{kind, origin, action, direction}</code>, on <code>/metrics</code>, instrumented at the <code>EventStore.Record</code> choke point. Labels are cardinality-bounded (no pattern or user IDs). This complements the capacity-bound <code>/api/pii/events</code> ring buffer and, crucially, makes silent filter failure alertable: <code>rate()</code> on the counter detects that the PII filter stopped firing after a deploy.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4792378149" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10641" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10641/hovercard" href="https://github.com/mudler/LocalAI/pull/10641">#10641</a></p>
</blockquote>
<h3>🔒 Gallery SSRF hardening</h3>
<p><code>POST /models/apply</code> with an empty <code>id</code> fetches the supplied <code>url</code> directly; in a default Docker setup (no API key) any reachable client could probe internal services or cloud-metadata (<code>169.254.169.254</code>) and exfiltrate a slice via the job error. Gallery config fetches now run through the existing <code>ValidateExternalURL</code> guard (the same one protecting the CORS proxy and media downloads), blocking private, loopback, link-local, unspecified, and metadata addresses. Only plain <code>http(s)://</code> is validated; <code>huggingface://</code>, <code>github:</code>, <code>oci://</code>, <code>ollama://</code>, and <code>file://</code> are untouched.</p>
<blockquote>
<p>🔗 PRs: <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4806217291" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10673" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10673/hovercard" href="https://github.com/mudler/LocalAI/pull/10673">#10673</a></p>
</blockquote>
<hr>
<h2>🐛 Bug Fixes (recap)</h2>
<ul>
<li><code>fix(vllm)</code>: restore non-streaming tool-call extraction that regressed after <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4668958468" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10351" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10351/hovercard" href="https://github.com/mudler/LocalAI/pull/10351">#10351</a> (a capability flag was mistaken for run state) - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4789399895" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10638" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10638/hovercard" href="https://github.com/mudler/LocalAI/pull/10638">#10638</a></li>
<li><code>fix(python-backends)</code>: decode tool-call <code>arguments</code> for chat templates (unbreaks MLX/Qwen3.5 agent loops) and split reasoning when a model emits only a closing <code></think></code> - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4801637898" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10658" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10658/hovercard" href="https://github.com/mudler/LocalAI/pull/10658">#10658</a></li>
<li><code>fix(cloud-proxy)</code>: drop <code>temperature</code>/<code>top_p</code> and send <code>max_completion_tokens</code> so routing to the newest reasoning models (Claude Opus 4.x, GPT-5.x) stops 400ing - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4792353488" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10640" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10640/hovercard" href="https://github.com/mudler/LocalAI/pull/10640">#10640</a></li>
<li><code>fix(config)</code>: revert defaulting <code>swa_full:true</code> for sliding-window-attention models (restores the memory-light reduced KV cache; still available as an explicit per-model opt-in) - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4806219704" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10674" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10674/hovercard" href="https://github.com/mudler/LocalAI/pull/10674">#10674</a></li>
<li><code>fix(kokoros)</code>: implement the <code>AudioTranscriptionLive</code> trait stub so the backend compiles against the updated proto - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4779003416" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10612" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10612/hovercard" href="https://github.com/mudler/LocalAI/pull/10612">#10612</a></li>
<li><code>fix(launcher)</code>: keep the desktop launcher's data/config under <code>~/.localai</code> instead of the GUI's working directory - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4777514907" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10610" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10610/hovercard" href="https://github.com/mudler/LocalAI/issues/10610">#10610</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4779773804" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10613" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10613/hovercard" href="https://github.com/mudler/LocalAI/pull/10613">#10613</a></li>
</ul>
<hr>
<h2>👒 Dependencies</h2>
<p>Submodule and backend bumps this cycle:</p>
<ul>
<li><code>ggml-org/llama.cpp</code> x4</li>
<li><code>ikawrakow/ik_llama.cpp</code> x4</li>
<li><code>CrispStrobe/CrispASR</code> x4</li>
<li><code>leejet/stable-diffusion.cpp</code> x3</li>
<li><code>vllm-metal</code> (darwin) x3</li>
<li><code>ggml-org/whisper.cpp</code> x2</li>
<li><code>mudler/parakeet.cpp</code> x1</li>
<li><code>localai-org/privacy-filter.cpp</code> x1</li>
<li><code>vllm-project/vllm</code> cu130 wheel to <code>0.24.0</code></li>
</ul>
<p>Plus new gallery models added via the gallery agent (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4803581530" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10663" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10663/hovercard" href="https://github.com/mudler/LocalAI/pull/10663">#10663</a>, <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4794538863" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10644" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10644/hovercard" href="https://github.com/mudler/LocalAI/pull/10644">#10644</a>).</p>
<hr>
<h2>📖 Documentation</h2>
<ul>
<li>Docs version bump for the release - <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4780264826" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10614" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10614/hovercard" href="https://github.com/mudler/LocalAI/pull/10614">#10614</a></li>
</ul>
<hr>
<h2>🙌 New Contributors</h2>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/alaningtrump/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/alaningtrump">@alaningtrump</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4799825993" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10657" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10657/hovercard" href="https://github.com/mudler/LocalAI/pull/10657">#10657</a></li>
</ul>
<hr>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.5.6...v4.6.0"><tt>v4.5.6...v4.6.0</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.5.62026-06-30T16:07:11Zv4.5.6
<h2>What's Changed</h2>
<h3>👒 Dependencies</h3>
<ul>
<li>chore(deps): bump actions/cache from 4 to 6 by <a class="user-mention notranslate" data-hovercard-type="organization" data-hovercard-url="/orgs/dependabot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dependabot">@dependabot</a>[bot] in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4770657459" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10593" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10593/hovercard" href="https://github.com/mudler/LocalAI/pull/10593">#10593</a></li>
</ul>
<h3>Other Changes</h3>
<ul>
<li>docs: ⬆️ update docs version mudler/LocalAI by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4759823691" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10560" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10560/hovercard" href="https://github.com/mudler/LocalAI/pull/10560">#10560</a></li>
<li>fix(distributed): missing agent NATS permission by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ALameLlama/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ALameLlama">@ALameLlama</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4757470487" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10549" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10549/hovercard" href="https://github.com/mudler/LocalAI/pull/10549">#10549</a></li>
<li>feat(distributed): SyncedMap component + migrate finetune/quant/agent-tasks to cross-replica state by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4755936695" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10542" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10542/hovercard" href="https://github.com/mudler/LocalAI/pull/10542">#10542</a></li>
<li>chore(fish-speech): drop the darwin/metal build target by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4759903517" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10561" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10561/hovercard" href="https://github.com/mudler/LocalAI/pull/10561">#10561</a></li>
<li>fix(config): fall back to DefaultContextSize for unparseable GGUFs; pin NVFP4 gallery context_size by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4760016706" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10563" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10563/hovercard" href="https://github.com/mudler/LocalAI/pull/10563">#10563</a></li>
<li>ci(vibevoice): skip the ASR transcription e2e on release tag builds by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4760125316" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10567" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10567/hovercard" href="https://github.com/mudler/LocalAI/pull/10567">#10567</a></li>
<li>fix(gallery): match mmproj/model quant as a whole token so F16 no longer selects BF16 (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4759449084" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10559" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10559/hovercard" href="https://github.com/mudler/LocalAI/issues/10559">#10559</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4760090155" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10564" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10564/hovercard" href="https://github.com/mudler/LocalAI/pull/10564">#10564</a></li>
<li>fix(distributed): return empty backend list for agent nodes instead of failing backend.list (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4756856991" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10545" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10545/hovercard" href="https://github.com/mudler/LocalAI/issues/10545">#10545</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4760093987" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10565" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10565/hovercard" href="https://github.com/mudler/LocalAI/pull/10565">#10565</a></li>
<li>feat(distributed): add LOCALAI_DISTRIBUTED_SHARED_MODELS to skip staging on shared volumes (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4758502408" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10556" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10556/hovercard" href="https://github.com/mudler/LocalAI/issues/10556">#10556</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4760101515" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10566" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10566/hovercard" href="https://github.com/mudler/LocalAI/pull/10566">#10566</a></li>
<li>chore: ⬆️ Update leejet/stable-diffusion.cpp to <code>9956436c925a367daeab097598b1ea1f32d3503f</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4755165627" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10533" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10533/hovercard" href="https://github.com/mudler/LocalAI/pull/10533">#10533</a></li>
<li>fix(openresponses): bound resume-stream buffer and enforce response ownership by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4760286579" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10569" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10569/hovercard" href="https://github.com/mudler/LocalAI/pull/10569">#10569</a></li>
<li>chore: ⬆️ Update ggml-org/whisper.cpp to <code>0ae02cdb2c7317b50991367c165736ce42ed96ac</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4755144101" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10532" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10532/hovercard" href="https://github.com/mudler/LocalAI/pull/10532">#10532</a></li>
<li>chore: ⬆️ Update CrispStrobe/CrispASR to <code>6514c9da00b03a2f0f1b49a43fae4f3a01a41844</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4755235513" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10535" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10535/hovercard" href="https://github.com/mudler/LocalAI/pull/10535">#10535</a></li>
<li>chore: ⬆️ Update ggml-org/llama.cpp to <code>0ed235ea2c17a19fc8238668653946721ed136fd</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4755273251" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10536" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10536/hovercard" href="https://github.com/mudler/LocalAI/pull/10536">#10536</a></li>
<li>fix(ik-llama): port multimodal path to mtmd API and bump to f96eaddb (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4755203186" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10534" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10534/hovercard" href="https://github.com/mudler/LocalAI/pull/10534">#10534</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4760250713" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10568" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10568/hovercard" href="https://github.com/mudler/LocalAI/pull/10568">#10568</a></li>
<li>feat(backends): add voice-detect + face-detect ggml backends (replace Python insightface/speaker-recognition) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4714582456" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10441" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10441/hovercard" href="https://github.com/mudler/LocalAI/pull/10441">#10441</a></li>
<li>fix(kokoro): add explicit click dep so spacy CLI works on intel build by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4761745488" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10572" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10572/hovercard" href="https://github.com/mudler/LocalAI/pull/10572">#10572</a></li>
<li>fix(launcher): robust binary download/upgrade (resume, rate-limit, UX) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4761952381" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10575" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10575/hovercard" href="https://github.com/mudler/LocalAI/pull/10575">#10575</a></li>
<li>fix(distributed): missing agent NATS permissions by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ALameLlama/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ALameLlama">@ALameLlama</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4761197844" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10571" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10571/hovercard" href="https://github.com/mudler/LocalAI/pull/10571">#10571</a></li>
<li>fix(fish-speech): allow invalid_reference_casting so tokenizers builds on darwin by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4761746122" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10573" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10573/hovercard" href="https://github.com/mudler/LocalAI/pull/10573">#10573</a></li>
<li>fix(oci): retry layer downloads on transient network errors by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4763031495" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10579" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10579/hovercard" href="https://github.com/mudler/LocalAI/pull/10579">#10579</a></li>
<li>chore(model-gallery): ⬆️ update checksum by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4763594951" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10585" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10585/hovercard" href="https://github.com/mudler/LocalAI/pull/10585">#10585</a></li>
<li>chore: ⬆️ Update leejet/stable-diffusion.cpp to <code>c1790754d31bec0731ed5fddc9d5b9ff22ee19cd</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4763533039" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10584" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10584/hovercard" href="https://github.com/mudler/LocalAI/pull/10584">#10584</a></li>
<li>chore: ⬆️ Update CrispStrobe/CrispASR to <code>6b50f76e59700665358a1aabf5295597fa318e06</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4763532959" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10583" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10583/hovercard" href="https://github.com/mudler/LocalAI/pull/10583">#10583</a></li>
<li>chore: ⬆️ Update ggml-org/llama.cpp to <code>dbdaece23de9ac63f2e7ca9e6bfcdc4fc156a3fa</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4763532941" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10582" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10582/hovercard" href="https://github.com/mudler/LocalAI/pull/10582">#10582</a></li>
<li>chore: ⬆️ Update mudler/voice-detect.cpp to <code>3d510772357538c5182808ac7de2278b84824e24</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4763532862" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10581" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10581/hovercard" href="https://github.com/mudler/LocalAI/pull/10581">#10581</a></li>
<li>chore: ⬆️ Update mudler/face-detect.cpp to <code>06914b077d52f90d5421299138e7be6bdd06b5e8</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4763532855" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10580" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10580/hovercard" href="https://github.com/mudler/LocalAI/pull/10580">#10580</a></li>
<li>chore: ⬆️ Update vllm-metal (darwin) to <code>v0.3.0.dev20260628073537</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4759912405" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10562" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10562/hovercard" href="https://github.com/mudler/LocalAI/pull/10562">#10562</a></li>
<li>chore(recon): re-pin voice/face-detect to squashed release commits (+ graph-cache fix) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4769512343" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10591" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10591/hovercard" href="https://github.com/mudler/LocalAI/pull/10591">#10591</a></li>
<li>fix(sglang): parse tool_call function arguments before applying the chat template by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/pos-ei-don/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/pos-ei-don">@pos-ei-don</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4759285108" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10558" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10558/hovercard" href="https://github.com/mudler/LocalAI/pull/10558">#10558</a></li>
<li>feat(realtime): Semantic VAD EOU token by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/richiejp/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/richiejp">@richiejp</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4717218875" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10444" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10444/hovercard" href="https://github.com/mudler/LocalAI/pull/10444">#10444</a></li>
<li>fix(openai): stop max_tokens streaming retry loop on reasoning models (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4405371302" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/9716" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/9716/hovercard" href="https://github.com/mudler/LocalAI/issues/9716">#9716</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Dennisadira/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Dennisadira">@Dennisadira</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4718393994" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10448" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10448/hovercard" href="https://github.com/mudler/LocalAI/pull/10448">#10448</a></li>
<li>fix(import): derive model name from selected GGUF for repo-root URIs by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Dennisadira/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Dennisadira">@Dennisadira</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4768831290" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10589" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10589/hovercard" href="https://github.com/mudler/LocalAI/pull/10589">#10589</a></li>
<li>fix(functions): avoid quadratic-time debug logging in CleanupLLMResult / ParseFunctionCall by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/pos-ei-don/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/pos-ei-don">@pos-ei-don</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4770200812" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10592" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10592/hovercard" href="https://github.com/mudler/LocalAI/pull/10592">#10592</a></li>
<li>chore: ⬆️ Update leejet/stable-diffusion.cpp to <code>3b6c9ca97cfcda8e68e719e6670d06379fcbe943</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4771446376" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10594" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10594/hovercard" href="https://github.com/mudler/LocalAI/pull/10594">#10594</a></li>
<li>chore: ⬆️ Update ggml-org/llama.cpp to <code>6f4f53f2b7da54fcdbbecaaa734337c337ad6176</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4771446440" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10595" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10595/hovercard" href="https://github.com/mudler/LocalAI/pull/10595">#10595</a></li>
<li>chore: ⬆️ Update localai-org/privacy-filter.cpp to <code>595f59630c69d361b5196f2aba2c71c873d0c13c</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4771446489" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10596" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10596/hovercard" href="https://github.com/mudler/LocalAI/pull/10596">#10596</a></li>
<li>chore: ⬆️ Update CrispStrobe/CrispASR to <code>3b93758f9725d400eca82976f895e4cec3f31260</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4771446526" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10597" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10597/hovercard" href="https://github.com/mudler/LocalAI/pull/10597">#10597</a></li>
<li>chore: ⬆️ Update ikawrakow/ik_llama.cpp to <code>f74a6fb87b315b2c3154166e075360e15021a61d</code> by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4771446633" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10598" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10598/hovercard" href="https://github.com/mudler/LocalAI/pull/10598">#10598</a></li>
<li>fix(import): strip file:// scheme from model path for local imports by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/Dennisadira/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/Dennisadira">@Dennisadira</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4771688833" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10599" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10599/hovercard" href="https://github.com/mudler/LocalAI/pull/10599">#10599</a></li>
<li>fix(tests): align openresponses test model name with GGUF-derived naming (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4768831290" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10589" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10589/hovercard" href="https://github.com/mudler/LocalAI/pull/10589">#10589</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4777208332" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10609" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10609/hovercard" href="https://github.com/mudler/LocalAI/pull/10609">#10609</a></li>
<li>fix(macos): staple the notarization ticket to the .app, not just the dmg by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4775158325" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10606" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10606/hovercard" href="https://github.com/mudler/LocalAI/pull/10606">#10606</a></li>
<li>fix(watchdog): persist UI-saved Check Interval across restarts (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4772339164" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10601" data-hovercard-type="issue" data-hovercard-url="/mudler/LocalAI/issues/10601/hovercard" href="https://github.com/mudler/LocalAI/issues/10601">#10601</a>) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4774896904" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10605" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10605/hovercard" href="https://github.com/mudler/LocalAI/pull/10605">#10605</a></li>
<li>feat(config): default swa_full:true for sliding-window-attention models by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4778313664" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10611" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10611/hovercard" href="https://github.com/mudler/LocalAI/pull/10611">#10611</a></li>
</ul>
<h2>New Contributors</h2>
<ul>
<li><a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/ALameLlama/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/ALameLlama">@ALameLlama</a> made their first contribution in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4757470487" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10549" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10549/hovercard" href="https://github.com/mudler/LocalAI/pull/10549">#10549</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.5.5...v4.5.6"><tt>v4.5.5...v4.5.6</tt></a></p>mudlertag:github.com,2008:Repository/615869301/v4.5.52026-06-27T12:58:42Zv4.5.5
<h2>What's Changed</h2>
<h3>Other Changes</h3>
<ul>
<li>fix(backends): repair release CI build/test breaks (kokoros, fish-speech, llama-cpp-quantization, sglang) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4757348265" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10547" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10547/hovercard" href="https://github.com/mudler/LocalAI/pull/10547">#10547</a></li>
<li>chore(model gallery): 🤖 add 1 new models via gallery agent by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4756172283" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10544" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10544/hovercard" href="https://github.com/mudler/LocalAI/pull/10544">#10544</a></li>
<li>fix(backends): whisper darwin run.sh loads whichever fallback lib exists (.so/.dylib) by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/localai-bot/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/localai-bot">@localai-bot</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="4758018132" data-permission-text="Title is private" data-url="https://github.com/mudler/LocalAI/issues/10553" data-hovercard-type="pull_request" data-hovercard-url="/mudler/LocalAI/pull/10553/hovercard" href="https://github.com/mudler/LocalAI/pull/10553">#10553</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/mudler/LocalAI/compare/v4.5.4...v4.5.5"><tt>v4.5.4...v4.5.5</tt></a></p>mudler