tag:github.com,2008:https://github.com/ROCm/FastFlowLM/releases
Release notes from FastFlowLM
2026-08-19T19:26:15Z
tag:github.com,2008:Repository/1003248064/v1.0.2
2026-08-20T15:58:50Z
π FastFlowLM v1.0.2 β Faster Qwen3.5 & Qwen3.6-MoE
<p>This release brings a solid decoding and prefill speed boost across the entire Qwen3.5 family and Qwen3.6-MoE, plus a heads-up on a required weight update.</p>
<hr>
<h2>π₯ Weights Update Required</h2>
<p>Models in this release are quantized by FLM itself. If you're upgrading to v1.0.2, <strong>you'll need to re-download weights</strong> for the Qwen3.5 family and Qwen3.6-MoE β existing local copies from prior versions are not compatible.</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm pull qwen3.5:0.8b
flm pull qwen3.5:2b
flm pull qwen3.5:4b
flm pull qwen3.5:9b
flm pull qwen3.6-moe:35b-a3b"><pre>flm pull qwen3.5:0.8b
flm pull qwen3.5:2b
flm pull qwen3.5:4b
flm pull qwen3.5:9b
flm pull qwen3.6-moe:35b-a3b</pre></div>
<hr>
<h2>β‘ Performance Boost: Qwen3.5 Family & Qwen3.6-MoE</h2>
<p>Both prefill and decoding throughput have been improved across all context lengths (1kβ32k) for <strong>Qwen3.5</strong> (0.8B, 2B, 4B, 9B) and <strong>Qwen3.6-MoE</strong> (35B-A3B).</p>
<table>
<thead>
<tr>
<th>Model</th>
<th align="right">Decoding Gain (avg / peak)</th>
<th align="right">Prefill Gain (avg / peak)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Qwen3.5 0.8B</td>
<td align="right">+9.5% / +14.5%</td>
<td align="right">+34.8% / +44.3%</td>
</tr>
<tr>
<td>Qwen3.5 2B</td>
<td align="right">+10.2% / +11.8%</td>
<td align="right">+27.2% / +33.3%</td>
</tr>
<tr>
<td>Qwen3.5 4B</td>
<td align="right">+11.5% / +13.1%</td>
<td align="right">+31.5% / +37.4%</td>
</tr>
<tr>
<td>Qwen3.5 9B</td>
<td align="right">+12.0% / +13.0%</td>
<td align="right">+25.5% / +29.2%</td>
</tr>
<tr>
<td>Qwen3.6-MoE 35B-A3B</td>
<td align="right">+11.4% / +13.0%</td>
<td align="right">+15.9% / +19.7%</td>
</tr>
</tbody>
</table>
<p>Gains are largest on smaller models and at mid-to-long context lengths, with prefill benefiting more than decoding across the board.</p>
<hr>
<h2>π Summary</h2>
<table>
<thead>
<tr>
<th></th>
<th>Highlight</th>
</tr>
</thead>
<tbody>
<tr>
<td>π₯</td>
<td>Weights update required for Qwen3.5 family and Qwen3.6-MoE β re-pull models before running v1.0.2</td>
</tr>
<tr>
<td>β‘</td>
<td>Qwen3.5 family (0.8B/2B/4B/9B): up to <strong>+14.5%</strong> decoding and <strong>+44.3%</strong> prefill throughput</td>
</tr>
<tr>
<td>β‘</td>
<td>Qwen3.6-MoE 35B-A3B: up to <strong>+13.0%</strong> decoding and <strong>+19.7%</strong> prefill throughput</td>
</tr>
</tbody>
</table>
<p>Thanks for your support β see you in the next one! π</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v1.0.1
2026-08-11T22:30:45Z
π FastFlowLM v1.0.1 β Windows Installer Switch & SmolVLA Benchmark
<h2>π¦ Windows Installer: <code>flm-setup.exe</code> β <code>flm-setup.msi</code></h2>
<p>The Windows installer has been switched from <code>flm-setup.exe</code> to <code>flm-setup.msi</code>. This change is reflected across:</p>
<ul>
<li>Release artifacts</li>
<li>Download links</li>
<li>Documentation</li>
<li>Website</li>
</ul>
<hr>
<h2>π SmolVLA Benchmark</h2>
<p>We've provided the benchmark numbers for SmolVLA. Full results: <a href="https://fastflowlm.com/docs/benchmarks/smolvla_results/" rel="nofollow">https://fastflowlm.com/docs/benchmarks/smolvla_results/</a></p>
<p><strong>Test System:</strong> AMD Ryzenβ’ AI 9 370 (Strix Point) with 32 GB DRAM; performance is comparable to other Strix Point and Strix Halo Point systems.</p>
<p><strong>Inference Latency</strong> (ms per inference, with different camera input counts)</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>HW</th>
<th align="right">1 image</th>
<th align="right">2 images</th>
<th align="right">3 images</th>
</tr>
</thead>
<tbody>
<tr>
<td>SmolVLA</td>
<td>NPU (FLM)</td>
<td align="right">298</td>
<td align="right">363</td>
<td align="right">430</td>
</tr>
</tbody>
</table>
<hr>
<p>Thanks for your support β see you in the next one! π</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v1.0.0
2026-08-11T22:31:39Z
π FastFlowLM v1.0.0 β First Release Under ROCm
<p>This is a big one β v1.0.0 marks our first general release under the <strong>ROCm organization</strong>. Here's what's new π</p>
<hr>
<h2>π General Release v1.0.0 in the ROCm Org</h2>
<p>FastFlowLM is now officially maintained under <code>ROCm/FastFlowLM</code>. Everything from the old repo β issues, pull requests, and history β has been transferred over, so nothing is lost in the move. All future releases, issues, and contributions happen there β please make sure your bookmarks, forks, and remotes point to the new home.</p>
<hr>
<h2>π€ New Model: SmolVLA</h2>
<p>FLM now supports <strong><a href="https://fastflowlm.com/docs/models/smolvla" rel="nofollow">SmolVLA</a></strong>, a <strong>vision-language-action (VLA)</strong> model, adding robotics support to the lineup alongside our existing LLM and VLM models.</p>
<p>Model card and usage details:</p>
<ul>
<li><strong>Hugging Face:</strong> <a href="https://huggingface.co/FastFlowLM/smolvla" rel="nofollow">https://huggingface.co/FastFlowLM/smolvla</a></li>
<li><strong>ModelScope:</strong> <a href="https://modelscope.cn/models/amd/smolvla" rel="nofollow">https://modelscope.cn/models/amd/smolvla</a></li>
</ul>
<hr>
<h2>βοΈ Fine-Grained Control for <code>flm bench</code></h2>
<p>To keep benchmarking fast by default, <code>flm bench</code> now runs <strong>2 iterations</strong> at each context length from 1k to 32k.</p>
<p>Need more data points? Override the iteration count with a flag:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm bench --bench-iterations 4"><pre>flm bench --bench-iterations 4</pre></div>
<hr>
<h2>π Bug Fix: Qwen3-VL Two-Image Handling</h2>
<p>Fixed an issue in <code>qwen3vl-it</code> where passing two images in a single request could cause the model to fail to recognize either image. Multi-image prompts now resolve correctly.</p>
<hr>
<h2>π Summary</h2>
<table>
<thead>
<tr>
<th></th>
<th>Highlight</th>
</tr>
</thead>
<tbody>
<tr>
<td>π </td>
<td>General release v1.0.0 β first release under the <code>ROCm/FastFlowLM</code> org</td>
</tr>
<tr>
<td>π€</td>
<td>New model: <strong>SmolVLA</strong> support</td>
</tr>
<tr>
<td>βοΈ</td>
<td><code>flm bench</code> now defaults to 2 iterations per context length, configurable via <code>--bench-iterations</code></td>
</tr>
<tr>
<td>π</td>
<td>Fixed <code>qwen3vl-it</code> failing to recognize images when given two at once</td>
</tr>
</tbody>
</table>
<p>Thanks for your support β see you in the next one! π</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v0.9.46
2026-07-28T23:28:02Z
π FastFlowLM v0.9.46 β We're Moving!
<h2>π FastFlowLM is Now an Official AMD Project</h2>
<p>π FastFlowLM has joined the <strong>ROCm organization</strong> and is now officially maintained by AMD. This is the last release under <code>FastFlowLM/FastFlowLM</code> β starting from <strong>v1.0.0</strong>, everything moves to <strong><a href="https://github.com/ROCm/FastFlowLM"><code>ROCm/FastFlowLM</code></a></strong>. Please update your bookmarks, forks, and remotes accordingly. See you there!</p>
<hr>
<h2>π ModelScope Support</h2>
<p>Models are pulled from HuggingFace by default. You can now opt into <strong>ModelScope</strong> as an alternative source with a single flag.</p>
<p>Pull a model from ModelScope:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm pull llama3.2:1b --modelscope 1"><pre>flm pull llama3.2:1b --modelscope 1</pre></div>
<p>Auto-pull from ModelScope in CLI mode:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm run llama3.2:1b --modelscope 1"><pre>flm run llama3.2:1b --modelscope 1</pre></div>
<p>Auto-pull from ModelScope in server mode:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm serve --modelscope 1"><pre>flm serve --modelscope 1</pre></div>
<p>Check the compatibility of your local model with ModelScope:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm check llama3.2:1b --modelscope 1"><pre>flm check llama3.2:1b --modelscope 1</pre></div>
<hr>
<h2>πΌοΈ More Image Resize Levels for Qwen-VL Models</h2>
<p>Fine-grained control over input image resolution is now available via the <code>-r</code> flag:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm serve -r <img-pre-resize-level>"><pre>flm serve -r <span class="pl-k"><</span>img-pre-resize-level<span class="pl-k">></span></pre></div>
<table>
<thead>
<tr>
<th>Level</th>
<th>Resolution</th>
</tr>
</thead>
<tbody>
<tr>
<td>0</td>
<td>Original size</td>
</tr>
<tr>
<td>1</td>
<td>Height = 480</td>
</tr>
<tr>
<td>2</td>
<td>Height = 720</td>
</tr>
<tr>
<td>3</td>
<td>Height = 1080</td>
</tr>
<tr>
<td>4</td>
<td>Height = 1440</td>
</tr>
<tr>
<td>5</td>
<td>Height = 2160</td>
</tr>
<tr>
<td>6</td>
<td>Height = 2880</td>
</tr>
<tr>
<td>7</td>
<td>Height = 3240</td>
</tr>
<tr>
<td>8</td>
<td>Height = 4320</td>
</tr>
</tbody>
</table>
<hr>
<h2>β‘ Speed Improvements for Qwen3.6-MoE</h2>
<p>Both prefill and decoding throughput have been improved for Qwen3.6-MoE across all context lengths.</p>
<p><strong>Decoding throughput (tokens/s):</strong></p>
<table>
<thead>
<tr>
<th>Context</th>
<th>Old</th>
<th>New</th>
<th>Gain</th>
</tr>
</thead>
<tbody>
<tr>
<td>1k</td>
<td>12.41</td>
<td>13.65</td>
<td>+9.99%</td>
</tr>
<tr>
<td>2k</td>
<td>12.26</td>
<td>13.41</td>
<td>+9.38%</td>
</tr>
<tr>
<td>4k</td>
<td>11.96</td>
<td>13.09</td>
<td>+9.45%</td>
</tr>
<tr>
<td>8k</td>
<td>11.38</td>
<td>12.51</td>
<td>+9.93%</td>
</tr>
<tr>
<td>16k</td>
<td>10.40</td>
<td>11.24</td>
<td>+8.08%</td>
</tr>
<tr>
<td>32k</td>
<td>8.88</td>
<td>9.51</td>
<td>+7.09%</td>
</tr>
</tbody>
</table>
<p><strong>Prefill throughput (tokens/s):</strong></p>
<table>
<thead>
<tr>
<th>Context</th>
<th>Old</th>
<th>New</th>
<th>Gain</th>
</tr>
</thead>
<tbody>
<tr>
<td>1k</td>
<td>75.18</td>
<td>78.98</td>
<td>+5.05%</td>
</tr>
<tr>
<td>2k</td>
<td>109.85</td>
<td>118.04</td>
<td>+7.46%</td>
</tr>
<tr>
<td>4k</td>
<td>150.93</td>
<td>156.43</td>
<td>+3.64%</td>
</tr>
<tr>
<td>8k</td>
<td>181.56</td>
<td>197.93</td>
<td>+9.02%</td>
</tr>
<tr>
<td>16k</td>
<td>214.46</td>
<td>218.84</td>
<td>+2.04%</td>
</tr>
<tr>
<td>32k</td>
<td>219.72</td>
<td>221.96</td>
<td>+1.02%</td>
</tr>
</tbody>
</table>
<hr>
<h2>π Summary</h2>
<table>
<thead>
<tr>
<th></th>
<th>Highlight</th>
</tr>
</thead>
<tbody>
<tr>
<td>π </td>
<td>FastFlowLM is now an official AMD project β repo moved to <code>ROCm/FastFlowLM</code></td>
</tr>
<tr>
<td>π</td>
<td>ModelScope support: pull, serve, and check model compatibility</td>
</tr>
<tr>
<td>πΌοΈ</td>
<td>9-level image resize control for Qwen-VL models</td>
</tr>
<tr>
<td>β‘</td>
<td>Up to ~+10% decoding and ~+9% prefill speedup for Qwen3.6-MoE</td>
</tr>
</tbody>
</table>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v0.9.45
2026-07-21T21:16:18Z
π FastFlowLM v0.9.45 β Qwen3.6-35B-A3B & Smoother KV Cache Through Multi-Backend Support
<p>Here's what's new π</p>
<h2>π€ New Model: <code>Qwen3.6-35B-A3B</code></h2>
<p>Say hello to <code>Qwen3.6-35B-A3B</code> β the second MoE model in FLM, joining <code>GPT-OSS</code>. It packs 35B total parameters with only 3B activated per forward pass, so you get strong reasoning quality at a fraction of the compute cost.</p>
<p><strong>Tag:</strong> <code>qwen3.6-moe:35b-a3b</code></p>
<p>Run in CLI mode:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm run qwen3.6-moe:35b-a3b"><pre class="notranslate"><code>flm run qwen3.6-moe:35b-a3b
</code></pre></div>
<p>Run in server mode:</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm serve qwen3.6-moe:35b-a3b"><pre class="notranslate"><code>flm serve qwen3.6-moe:35b-a3b
</code></pre></div>
<p>Check out the <a href="https://fastflowlm.com/docs/models/qwen/#-model-card-qwen36-35b-a3b" rel="nofollow">model card</a> and <a href="https://fastflowlm.com/docs/benchmarks/qwen3.6_results/" rel="nofollow">benchmark results</a> for more details.</p>
<hr>
<h2>β‘ Smoother KV Cache Through Multi-Backend Support</h2>
<p>FLM now supports per-round KV cache checks, making context management more precise and robust when mixing inference backends.</p>
<p>Here's a real example β imagine a Lemonade user running <code>gemma4-it:e2b</code> on both NPU (via FLM) and GPU (via llama.cpp):</p>
<ol>
<li>They send an initial prompt to the NPU and get a response.</li>
<li>They continue on the GPU and get a second response.</li>
<li>They switch back to the NPU with the full conversation history.</li>
</ol>
<p>Previously, the NPU would see gaps from the GPU round, fail the KV cache check, and re-prefill everything from scratch. Now, with per-round KV cache checks, it knows the first round is already cached β so it only prefills the GPU round and the new prompt. Mixing backends is no longer a headache! π</p>
<hr>
<h2>π Summary</h2>
<ul>
<li><strong>New model: <code>Qwen3.6-35B-A3B</code></strong> β our second MoE model, with 35B total parameters and just 3B activated per token π§ </li>
<li><strong>Smarter KV cache</strong> β per-round checks let you mix FLM and other backends without losing cache or re-prefilling the whole context π</li>
</ul>
<p>Thanks for your support β more good stuff is on the way. See you in the next one! π</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v0.9.44
2026-07-06T02:45:58Z
π FastFlowLM v0.9.44 β Portable Linux Support
<h2>π§ Portable FLM for Linux</h2>
<p>FastFlowLM is now available as a portable build for Linux! This makes it easy to bundle and integrate the app into any higher-level application or custom environment β no installation required, just plug it in and go.</p>
<h2>β‘ Quick Start</h2>
<p>Download the portable package from the <a href="https://github.com/FastFlowLM/FastFlowLM/releases/latest/download">releases page</a> and place it in your home directory, then run:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="cd ~/portable
tar -xzf fastflowlm_v0.9.44_linux.tar.gz
cd fastflowlm_v0.9.44_linux
./flm run llama3.2:1b"><pre><span class="pl-c1">cd</span> <span class="pl-k">~</span>/portable
tar -xzf fastflowlm_v0.9.44_linux.tar.gz
<span class="pl-c1">cd</span> fastflowlm_v0.9.44_linux
./flm run llama3.2:1b</pre></div>
<h2>π Acknowledgements</h2>
<p>A huge thank you to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/superm1/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/superm1">@superm1</a> and the <a href="https://github.com/lemonade-sdk/lemonade">Lemonade team</a> β your invaluable insights and contributions made this portable Linux release possible. We truly appreciate your support! π</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v0.9.43
2026-05-26T19:17:02Z
π FastFlowLM v0.9.43 - FLM Benchmarking Tool, KV Cache, Chat Templates, and Tool Calling
<h2>π FLM Benchmarking Tool</h2>
<p>You can now use the FLM benchmarking tool to test models and compare performance across context lengths.</p>
<p>Each benchmark runs from <code>1k</code> to <code>32k</code> context length for <code>8</code> iterations.</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm bench <model_tag>"><pre>flm bench <span class="pl-k"><</span>model_tag<span class="pl-k">></span></pre></div>
<p>The benchmark reports:</p>
<ul>
<li><strong>TTFT:</strong> Time to first token in seconds.</li>
<li><strong>Prefill speed:</strong> Prompt processing speed in tokens per second.</li>
<li><strong>Decoding speed:</strong> Token generation speed in tokens per second.</li>
</ul>
<p>Results are printed as a table in the shell and saved as a CSV file in the current folder for later reference.</p>
<hr>
<h2>β‘ KV Cache Improvements</h2>
<p>FLM now uses a safer KV cache flow that preserves the full conversation history before applying the chat template.</p>
<h3>π What Changed</h3>
<p>Previously, FLM computed a checksum for incoming messages and handled cache state as follows:</p>
<ul>
<li><strong>Cache hit:</strong> Manually removed the old message from the message list.</li>
<li><strong>Cache miss:</strong> Cleared the context.</li>
</ul>
<p>This approach caused several issues: message-list manipulation happened before chat templating, <code>reasoning content</code> could remain in the KV cache indefinitely, and the <code>bos</code> token could be prefilled repeatedly on each conversation turn. Together, these issues could lead to fragmented or incorrectly formatted conversation histories.</p>
<p>The updated flow keeps the original message history intact:</p>
<ul>
<li><strong>Checksum tracking:</strong> FLM still computes a checksum for incoming messages.</li>
<li><strong>Cache hit:</strong> The existing context is retained while the complete original message history is processed.</li>
<li><strong>Cache miss:</strong> The context is cleared as before.</li>
<li><strong>Consistent templating:</strong> The chat template is always applied to the complete, unbroken message history sent by the client.</li>
<li><strong>Token-level diffing:</strong> FLM tokenizes the fully templated prompt, compares the token IDs against the cached token IDs from the previous turn, removes the overlapping cached portion, and keeps only the newly generated token IDs for prefill.</li>
</ul>
<p>The new implementation uses a <code>checkpoint</code> and <code>restore</code> mechanism to record the current KV cache state and restore it in the next round of conversation.</p>
<p>This keeps <code>reasoning content</code> out of the KV cache, making cached conversations more consistent and reliable, especially for complex prompts. It also helps ensure the final templated prompt remains correctly formatted.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/NVolcz/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/NVolcz">@NVolcz</a> for inspiring this improvement and providing feedback on the previous flow!</p>
<hr>
<h2>π§© Chat Template Updates</h2>
<p>The Gemma4 chat template has been updated to match the latest official release from Google. To support the improved KV cache logic, we also updated chat templates for several model families so they can safely handle complete message histories.</p>
<p>Updated model families:</p>
<ul>
<li><code>gemma4-it</code></li>
<li><code>qwen3.5</code></li>
<li><code>medgemma</code></li>
<li><code>deepseek-r1</code></li>
</ul>
<p>You do not need to redownload all model files manually. Use <code>flm list</code> to check whether your local model files are compatible with the latest chat templates.</p>
<p>There are two ways to update a model's chat template:</p>
<ol>
<li>Run <code>flm check <model_tag></code> and then <code>flm pull <model_tag></code> to update the model before your next use.</li>
<li>Run <code>flm run <model_tag></code> or <code>flm serve <model_tag></code>, and FLM will automatically check and update the chat template if needed.</li>
</ol>
<blockquote>
<p><g-emoji class="g-emoji" alias="warning">β οΈ</g-emoji> <strong>Note:</strong> The checking process may take about <code>20</code> seconds, depending on the model, but it only needs to run once per model version. After the chat template is updated, you can run or serve the model as usual without the extra delay.</p>
</blockquote>
<hr>
<h2>π Smarter Force Pulls</h2>
<p><code>flm pull <model_tag> --force</code> now checks existing model files first and only re-downloads missing or incorrect files, instead of removing and re-downloading the entire model.</p>
<hr>
<h2>π§° <code>Gemma4</code> Tool-Calling Message Formatting</h2>
<p><code>Gemma4</code> uses a different tool-calling message format from <code>OpenAI</code>-compatible tool-calling APIs. Previously, upstream <code>OpenAI</code>-formatted tool-call messages could be sent directly to <code>Gemma4</code> without conversion, which could lead to parsing issues.</p>
<p>FLM now detects <code>OpenAI</code>-formatted tool-call messages and converts them to the <code>Gemma4</code> tool-calling format before sending them to the model. This improves parsing reliability for <code>Gemma4</code> tool calls.</p>
<p>For more details about the format differences, see the <a href="https://ai.google.dev/gemma/docs/capabilities/text/function-calling-gemma4" rel="nofollow"><code>Gemma4</code> tool-calling documentation</a>.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/NVolcz/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/NVolcz">@NVolcz</a> for pointing this out!</p>
<hr>
<h2>π οΈ <code>Gemma4</code> Tool-Calling Reliability</h2>
<p>We received a lot of feedback about <code>Gemma4</code> tool-calling reliability, especially around tool-argument parsing.</p>
<p>Tool arguments are generated by the model as text and then parsed into JSON. Parsing can fail for several reasons, including malformed strings, complex nested arguments, or edge-case formatting.</p>
<p>This release improves <code>Gemma4</code> tool-call argument parsing across several edge cases, making tool calling more robust overall.</p>
<p>Some malformed outputs may still be impossible to recover automatically, such as tool calls that completely miss the expected chat-template format. We will continue improving coverage for these cases where possible.</p>
<p>We will also continue monitoring feedback and improving tool-calling reliability for <code>Qwen3</code> and <code>Qwen3.5</code>.</p>
<p>Thanks to the community members who shared their tool-calling experiences and edge cases. <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TatuLund/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TatuLund">@TatuLund</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jtmonroe/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jtmonroe">@jtmonroe</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/cqh963852/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/cqh963852">@cqh963852</a>, <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/antrv/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/antrv">@antrv</a>, and more than a dozen others contributed suggestions that helped improve tool-calling reliability.</p>
<hr>
<h2>π Tool Calling Disabled for Low-Reliability Models</h2>
<p>Tool-calling support has been disabled for the following models due to low usage and low reliability:</p>
<ul>
<li><code>lfm2.5-tk:1.2b</code></li>
<li><code>nanbeige4.1:3b</code></li>
</ul>
<hr>
<h2>π Summary</h2>
<p>FastFlowLM v0.9.43 adds benchmarking output for easier performance comparison, makes cached conversations more reliable by preserving complete message histories, keeping <code>reasoning content</code> out of the KV cache, and using token-level diffing after chat templating. This release also updates chat templates for <code>Gemma4</code>, <code>Qwen3.5</code>, <code>MedGemma</code>, and <code>DeepSeek-R1</code>, makes <code>flm pull <model_tag> --force</code> more efficient by re-downloading only missing or incorrect files, improves <code>Gemma4</code> tool-call formatting and argument parsing, and disables tool calling for low-reliability models.</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v0.9.42
2026-05-13T16:55:31Z
π FastFlowLM v0.9.42 - Tool Calling Reliability Updates
<h2>π οΈ Tool Calling Improvements</h2>
<h3>π§ Gemma4 JSON Validation</h3>
<p>Added a sanity check for Gemma4 tool-call JSON output to catch malformed responses more reliably.</p>
<p>This should help prevent failures caused by invalid JSON in tool calls.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/TatuLund/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/TatuLund">@TatuLund</a> for reporting this issue!</p>
<h3>π§° Qwen3.5 Non-Streaming Tool Calls</h3>
<p>Fixed incorrect parsing that could cause Qwen3.5 tool-calling support to fail in non-streaming mode.</p>
<h3>π§± Qwen3.5 Tool-Call Robustness</h3>
<p>Enhanced Qwen3.5 tool-call parsing to better handle cases where the model may miss closing tool tags.</p>
<p>This improves robustness for tool-calling workflows and reduces failures caused by incomplete tool-call markup.</p>
<hr>
<h2>π Summary</h2>
<p>FastFlowLM v0.9.42 focuses on more reliable tool calling. This release adds Gemma4 JSON sanity checks and improves Qwen3.5 tool-call behavior in both non-streaming and edge-case parsing scenarios.</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v0.9.41
2026-05-06T11:25:35Z
π FastFlowLM v0.9.41 - Tool Calling + Usability Updates
<h2>π Bug Fixes</h2>
<h3>π οΈ Tool Calling Cache Redundancy</h3>
<p>Fixed an issue where tool schemas could be injected into the KV cache on every cached turn.</p>
<p>This could fill the model context window with repeated copies of the same schema, wasting context tokens, increasing compute time, and making tool-calling behavior less reliable as instructions appeared repeatedly mid-conversation.</p>
<p>Thanks to <strong>nvolcz</strong> from Discord for reporting this issue!</p>
<hr>
<h2>β¨ Improvements</h2>
<h3>π Chat Completions Logging</h3>
<p>OpenAI-compatible chat completions logging now reports KV cache usage for each conversation round.</p>
<p>Example fields include:</p>
<div class="highlight highlight-source-json notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="{
"active_kv_tokens": 4096,
"max_kv_token_capacity": 32768,
"kv_token_occupancy_percentage": 12.5
}"><pre>{
<span class="pl-ent">"active_kv_tokens"</span>: <span class="pl-c1">4096</span>,
<span class="pl-ent">"max_kv_token_capacity"</span>: <span class="pl-c1">32768</span>,
<span class="pl-ent">"kv_token_occupancy_percentage"</span>: <span class="pl-c1">12.5</span>
}</pre></div>
<p><code>kv_token_occupancy_percentage = 4096 / 32768 Γ 100% = 12.5%</code></p>
<h3>π§ Gemma4 Audio Logging</h3>
<p>Improved Gemma4 audio logging to make long-audio handling easier to understand.</p>
<p>The previous wording could suggest that audio was clipped or partially dropped. The updated message now makes it clear when audio has been split into chunks for processing.</p>
<p>Example:</p>
<div class="highlight highlight-text-adblock notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="Audio in message is split into 33 chunks for processing."><pre>Audio in message is split into 33 chunks for processing.</pre></div>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/gdkrmr/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/gdkrmr">@gdkrmr</a> for the suggestion.</p>
<h3>π§ Arch Linux Guide</h3>
<p>Added an Arch Linux installation guide to the documentation.</p>
<p>For installation instructions, see the <a href="https://fastflowlm.com/docs/install_lin/#arch-linux" rel="nofollow">Arch Linux guide</a>.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/filipenf/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/filipenf">@filipenf</a> for the contribution!</p>
<h3>π Version Check Control</h3>
<p>FastFlowLM can now disable the automatic version check before <code>run</code> and <code>serve</code> modes.</p>
<p>To disable the startup version check, set system environment variable <code>FLM_DISABLE_UPDATE_CHECK</code> to <code>1</code>:</p>
<p>Linux:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="export FLM_DISABLE_UPDATE_CHECK=1"><pre><span class="pl-k">export</span> FLM_DISABLE_UPDATE_CHECK=1</pre></div>
<p>Windows:</p>
<div class="highlight highlight-source-powershell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="setx FLM_DISABLE_UPDATE_CHECK 1"><pre>setx FLM_DISABLE_UPDATE_CHECK <span class="pl-c1">1</span></pre></div>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/heliosran/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/heliosran">@heliosran</a> for the contribution.</p>
<h3>π§° Linux <code>make install</code></h3>
<p>Improved Linux installation behavior by avoiding accidental third-party submodule installs and placing bundled FLM shared libraries under <code>${CMAKE_INSTALL_LIBDIR}/flm</code>.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/J-Bu/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/J-Bu">@J-Bu</a> for the contribution.</p>
<hr>
<h2>π Summary</h2>
<p>FastFlowLM v0.9.41 fixes redundant tool schema injection in cached conversations and adds clearer observability for KV cache usage and Gemma4 audio processing. This release also improves Linux installation workflows with a new Arch Linux guide, better <code>make install</code> behavior, and an option to disable automatic version checks.</p>
github-actions[bot]
tag:github.com,2008:Repository/1003248064/v0.9.40
2026-04-28T20:28:26Z
π FastFlowLM v0.9.40 - Gemma4 E4B + Reliability Updates
<h2>π¦ New Model Support</h2>
<h3>π Gemma4-IT-E4B</h3>
<p>FastFlowLM now supports <code>gemma4-it:e4b</code> for language, vision, audio workloads, including concurrent multimodal input for omni-model use cases.</p>
<ul>
<li><strong>Tag:</strong> <code>gemma4-it:e4b</code></li>
</ul>
<p>Run in CLI mode:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm run gemma4-it:e4b"><pre>flm run gemma4-it:e4b</pre></div>
<p>Run in server mode:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm serve gemma4-it:e4b"><pre>flm serve gemma4-it:e4b</pre></div>
<p>For more details, see the <a href="https://fastflowlm.com/docs/models/gemma/#-model-card-gemma-4-e4b-it" rel="nofollow">model card</a> and <a href="https://fastflowlm.com/docs/benchmarks/gemma4_results/" rel="nofollow">benchmark results</a>.</p>
<hr>
<h2>β¨ Improvements</h2>
<h3>π₯ Performance Boosts for <code>gemma4-it:e2b</code></h3>
<p>This release brings meaningful speed improvements to the <code>gemma4-it:e2b</code> model:</p>
<ul>
<li><strong>Prefill:</strong> up to 11.4% faster</li>
<li><strong>Decoding:</strong> up to 10.2% faster</li>
</ul>
<h3>β‘ Chunk Prefill</h3>
<p>This release adds chunk prefill support, significantly reducing memory usage for long prompts and larger workloads.</p>
<p>You can configure the prefill chunk length with <code>--prefill-chunk-len</code> in both <code>CLI</code> and <code>server</code> modes. The default value is <code>4096</code>.</p>
<p>Run in CLI mode:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm run gemma4-it:e4b --prefill-chunk-len 8192"><pre>flm run gemma4-it:e4b --prefill-chunk-len 8192</pre></div>
<p>Run in server mode:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm serve gemma4-it:e4b --prefill-chunk-len 8192"><pre>flm serve gemma4-it:e4b --prefill-chunk-len 8192</pre></div>
<p>In server mode, you can now cancel a request even while it is still in the prefill stage. No more waiting around for a huge prompt to finish prefill: just hit the stop button in higher-level apps such as Open WebUI and move on.</p>
<h3>π Hash Checking</h3>
<p>A new hash checking command is now available to help verify downloaded model files.</p>
<p>If you have trouble running a model and suspect a corrupted download, run:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="flm check gemma4-it:e4b"><pre>flm check gemma4-it:e4b</pre></div>
<p>If corrupted files are detected, you will see output like this:</p>
<div class="highlight highlight-text-adblock notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="[FLM] Checking model: llama3.2:1b...
[FLM] Checking file: config.json...
[FLM] Fail!
[FLM] Removing corrupted file: config.json...
[FLM] Successfully removed config.json!
[FLM] Checking file: model.q4nx...
[FLM] Success!
[FLM] Checking file: tokenizer.json...
[FLM] Success!
[FLM] Checking file: tokenizer_config.json...
[FLM] Success!
[FLM] Model check completed with errors. Please use `flm pull llama3.2:1b` to re-download corrupted files."><pre>[FLM] Checking model: llama3.2:1b...
[FLM] Checking file: config.json...
[FLM] Fail!
[FLM] Removing corrupted file: config.json...
[FLM] Successfully removed config.json!
[FLM] Checking file: model.q4nx...
[FLM] Success!
[FLM] Checking file: tokenizer.json...
[FLM] Success!
[FLM] Checking file: tokenizer_config.json...
[FLM] Success!
[FLM] Model check completed with errors. Please use `flm pull llama3.2:1b` to re-download corrupted files.</pre></div>
<hr>
<h2>π Bug Fixes</h2>
<h3>π οΈ Tool Calling</h3>
<p>Fixed an issue where tool calls could return an incorrect finish reason.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/antrv/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/antrv">@antrv</a> for reporting this issue!</p>
<h3>βοΈβπ₯ Empty Multimodal Input Handling</h3>
<p>Fixed an issue where empty image or audio input in server mode could cause the server to break.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/antrv/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/antrv">@antrv</a> for reporting this issue!</p>
<h3>π§ Memory Limits</h3>
<p>Fixed a memlock limit issue that could affect loading ASR or embedding models standalone.</p>
<p>Thanks to <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/sofiageo/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/sofiageo">@sofiageo</a> for reporting this issue!</p>
<hr>
<h2>π Summary</h2>
<p>FastFlowLM v0.9.40 expands the Gemma4 lineup with <code>gemma4-it:e4b</code>. This release delivers meaningful speed improvements to <code>gemma4-it:e2b</code>, introduces chunk prefill for more efficient handling of long prompts, adds <code>check</code> command for verifying model files, and improves reliability across tool calling, multimodal input, and memory handling.</p>
github-actions[bot]