Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
-
Updated
Sep 3, 2026 - Swift
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
[Official] Prima.cpp: Scale Your Local AI Beyond One Device.
DeepSeek-V4-Flash-0731 284B inference in ~30 GB of RAM / Qwen3.8-Next-Flash-FP8 inference in ~20 GB of RAM on any M-series MacBook
Knolo turns documents into a verifiable Knowledge Image (.knolo): one portable file you can mount, query, and prove. Deterministic retrieval, cryptographic roots, fail-closed verification. Local-first. No vector DB required.
Open-source production-ready mobile server-less AI apps. Clone, customize, and ship on Android & iOS with MLange.
On-device shell command generator for macOS Tahoe. Uses Apple's 3B model with dynamic few-shot retrieval from 21k tldr examples.
Android 16 fork. AI as a platform primitive. Twelve capabilities, one shared runtime, every app. OEM-pluggable. Apache 2.0.
Agentic Android Open Source Project (AAOSP) — Android fork with native LLM system service, MCP-aware apps, and an agent-driven launcher. On-device Qwen 2.5 via llama.cpp. Apps declare tools in their manifest. The OS runs the model.
Declare the outcome, skip the logic. The developer-first infrastructure engine for structured, multi-turn AI conversations with built-in safety, async follow-ups, and native compliance
Rivo Agent is a private, local-first mobile AI assistant that runs compact GGUF language models directly on your phone. It lets users download a compatible model, chat offline, keep local memory, customize assistant behavior, and use AI privately without sending conversations to a remote inference server.
Apple FoundationModels API on iOS 18+. Same call site, native passthrough on iOS 26 (Apple Intelligence), CoreML / MLX backends on older OSes. Drop-in source compatible.
OnDevice Local AI Studio — run LLMs on iPhone offline. SwiftUI workbench with MLX, llama.cpp, whisper.cpp, RAG, and authenticated OpenAI/Anthropic/Ollama-compatible local APIs.
Run LLMs on Snapdragon NPU — including the 'unsupported' 8 Gen 1 (Hexagon v69). Verified at 31 tok/s on OnePlus 10 Pro.
Curated resource for mobile teams shipping on-device LLMs — runtime benchmarks, model picks, GDPR-friendly architecture, and real production use cases.
Kotlin Multiplatform engine for running Gemma LLMs on-device on Android via LiteRT-LM — stateful KV-cache chat sessions, resumable model management, function calling. Includes NativeLM, a private Local AI chat app. AGPL-3.0 / commercial.
High-performance Android SDK for on-device LLM inference (GGUF). Privacy-focused, offline-first, and powered by llama.cpp with a clean Kotlin Coroutines API.
Reverse-engineering notes on fm, Apple's Foundation Models CLI in macOS 27: on-device model catalog (9M/85M/300M/3B + code/vision/speech), Private Cloud Compute, Siri local<->cloud routing, and the OpenAI-compatible 'fm serve' API.
📱 手机端 AI 操作系统全景知识库 — 334+ 篇深度页面,覆盖端侧大模型、AI Agent、芯片适配、推理优化 | 自动更新
To associate your repository with the on-device-llm topic, visit your repo's landing page and select "manage topics."