Share to action, privately on your phone.
Understand screenshots, receipts, messages, and documents locally.
No account, no cloud, no telemetry — works in airplane mode.
Share something → understand it locally → propose the next action → approve and execute it on the phone. Marmot turns a screenshot, receipt, message, or document into an editable preview for a calendar event, reminder, draft reply, or saved note. Nothing is written or sent until you approve it.
share something → understand it locally → propose the next action → approve and execute it on the phone
The canonical repository is github.com/stancsz/marmot.
The older super-marmot/marmot parent is not a publishing target; releases,
issues, Discussions, and store metadata are maintained from this repository.
Start with the flagship path: share a screenshot or message, let the local model propose a calendar action, review the resolved time, and approve or undo the phone-side write. The Android proof run is recorded in the flagship verification log. See the sanitized before/after examples for verified calendar, offline travel, safe screenshot rejection, and the current draft reply quality gate.
Marmot is a lightweight iOS + Android app that does for phones what Ollama does for laptops: pick a model from a curated library, download it once, and chat with it locally. Inference runs on your device with llama.cpp — Metal-accelerated on iOS, ARM-optimized CPU on Android.
Instead of exposing thousands of models, Marmot ships one champion per weight class — the highest-rated open model at each size a phone can run, as of July 2026 — each with a RAM-fit badge computed from your device's actual memory so you know before downloading whether it will run comfortably.
| Home | Chat | Agent mode |
|---|---|---|
| Voice mode | Quick actions | Memory & RAG |
|---|---|---|
| Every launch | Model library | Export |
|---|---|---|
| Settings |
|---|
The mockups above show the design; these are actual captures from the app running on an Android 15 emulator (Pixel 7 profile) with Qwen3.5 0.8B generating on-device — 8.6 tok/s on emulated x86 CPU (real ARM phones with NEON are considerably faster).
| First launch | Model library | Real inference | Reasoning folded |
|---|---|---|---|
![]() |
![]() |
![]() |
![]() |
Chat
- 🐹 A proper hello — every launch, the marmot pops up with a random one-liner (“Zero servers were consulted in the making of this app.”) before getting out of your way. Tap to skip.
- ⚡ Streaming replies with tok/s stats, rendered as markdown (bold, lists, code blocks, tappable links) in assistant bubbles.
- 🤔 Reasoning-model aware —
<think>…</think>blocks fold into a "Thinking…" indicator instead of leaking raw tags. - 🎭 Personas — five built-ins (Concise, Coach, Writer, Tutor, Developer) plus save-your-own; the active persona shapes chat and agent answers alike.
- 🎛️ Tunable sampling — temperature, top-p, max tokens, context length, system prompt; 🌗 light by default, with dark and follow-system themes.
Agent
- 🤖 Agent Mode — flip the ⚙ chip and the model works step-by-step with local tools (calculator, clock, chat search, document search), showing a live thought/tool/observation timeline. Multi-step tasks are orchestrated: a planner decomposes, fresh-context executors run each step, live ☑ check-offs, then a synthesizer produces the answer (docs/AGENT.md).
- ⚖️ Verified answers (optional) — a reflection pass may revise, an independent judge scores; the verdict badge persists on the message.
- 🧠 Memory — user/project facts (editable in Settings → Memory) plus auto-captured episode summaries, recalled by meaning with on-device embeddings and injected into agent runs.
- 📚 Document & repo RAG — import text/markdown files or a public GitHub
repo (
owner/repo) and chat with them via semantic search, fully on-device. - 🌐 Web research (opt-in) —
web_search+fetch_pagetools with cited sources; with the switch off, Marmot is provably offline. - 🔌 MCP client — connect Model Context Protocol servers over HTTP (home server, Home Assistant, work tools); their tools appear in Agent Mode, namespaced per server, gated by the same policy layer.
Everyday
- ⚡ Share-to-Marmot — share a screenshot, receipt, message, page, or document from any app (or paste) and get one-tap local actions. Calendar previews show the resolved date/time, require approval, and support undo; transforms, drafts, and saves remain visibly reviewable.
- 🔬 Deep research — with web access on, a Research toggle steers the agent into multi-angle searches, source fetches, and a cited report.
- 🔗 Automation hooks —
marmot://ask?text=…deep link for iOS Shortcuts and Android automation.
Voice
- 🎙️ Conversation mode — a hands-free listen → think → speak loop with an animated, phase-aware voice stage.
- 📝 Meeting mode — continuous transcription with a recording timer; say "Marmot, …" for a suggested contribution (tap-to-speak card — it never talks into the room uninvited); transcripts save into searchable documents.
Models & data
-
Vision attachments — the SmolVLM 256M starter model downloads with its paired projector, so screenshots and receipts can be understood locally; PDFs and audio remain explicitly unsupported until their dedicated paths are shipped.
-
📦 Resumable downloads that continue in the background and survive restarts; a
.ggufon disk is always complete (atomic.partmoves). -
📱 RAM-fit badges ("Runs great / Should run / May be too big") from your device's actual memory; import any local
.ggufas a first-class model; experimental Android GPU toggle. -
📤 Export, import & share — Markdown chat sharing and JSON backups via the OS share sheet (Drive, OneDrive, Files — no cloud SDKs, no OAuth); restores merge without ever overwriting newer local history.
-
🔒 Private by architecture — no account, no backend, no telemetry; ~2 MB JS bundle; one model in memory at a time. Capability designs for what's next live in docs/CAPABILITIES.md.
| Class | Model | Download | Runs on | Why it's the pick |
|---|---|---|---|---|
| Vision starter | SmolVLM 256M Vision | 365 MB | ~1.5 GB+ RAM | Local screenshot and receipt understanding with a paired projector |
| 🪶 Featherweight | Qwen3.5 0.8B | 0.53 GB | any phone | Rated far above every other sub-1B; 200+ languages |
| 🥊 Lightweight | Qwen3.5 2B | 1.28 GB | 4–6 GB RAM | Beats Gemma 4 E2B on reasoning, GPQA, intelligence |
| 🥋 Middleweight | SmolLM3 3B | 1.92 GB | 6 GB RAM | Strongest fully-open 3B, dual-mode reasoning |
| 🏋️ Cruiserweight | Qwen3.5 4B | 2.74 GB | 8 GB RAM | Strongest dense 4B: knowledge, science, agentic wins |
| 👑 Heavyweight | Gemma 4 E4B | 4.06 GB | 12 GB+ RAM | 8B weights at a 4B footprint; closest to cloud quality |
All catalog entries use Apache 2.0 licenses. Text models are Q4_K_M GGUF builds
(Gemma 4 E4B ships as Q3_K_M to stay phone-sized) from
unsloth; the SmolVLM vision starter uses the
ggml-org build. URLs and
exact byte sizes verified against the hosted files. Rankings from
Artificial Analysis and per-tier benchmark
comparisons; re-evaluated as new models ship. Beyond the catalog, any local
.gguf can be imported from the Files app ("Import .gguf" in the model
library) and used as a first-class model.
Just want the app? Download the latest Android APK and sideload it. Browse release history for checksums and notes. The current public artifact is the sideload baseline; Google Play and TestFlight/App Store builds are prepared in the store runbook but are not claimed shipped until credentialed internal installs are verified.
Building from source: Marmot uses native modules, so it needs a development build (not Expo Go).
Prerequisites: Node 20+, and Android Studio (Android) or Xcode on macOS (iOS).
git clone https://github.com/stancsz/marmot.git
cd marmot
npm install
npx expo run:android # Android
npx expo run:ios # iOS (macOS only)expo run generates the native android/ and ios/ projects automatically
(continuous native generation) — they are disposable and not checked in.
Windows note: if llama.rn's postinstall fails with a tar error under Git Bash, rerun it with Windows' native tar:
PATH="/c/Windows/System32:$PATH" node node_modules/llama.rn/install/download-native-artifacts.js
src/
models/catalog.ts # curated model list (id, URL, exact size, license)
agent/ # pure-TS agent core — fully unit-tested
loop.ts # Observe→Decide→Act→Verify + JSON tool protocol
orchestrator.ts # per-step subagent executors + judge gate
planner.ts skills.ts # task decomposition; trigger→procedure registry
memory.ts documents.ts semantic.ts # memory + RAG, cosine retrieval
tools.ts web.ts # calculator/clock/search + opt-in web tools
reflection.ts verify.ts # self-critique + judge scoring
lib/
engine.ts # llama.cpp context lifecycle, embeddings, GPU opts
agentRuntime.ts # wires the agent core to the engine + stores
downloads.ts # resumable background downloads, atomic moves
chatStore.ts importParse.ts exportShare.ts # persistence + backup
customModels.ts repoCore.ts repoImport.ts # .gguf + GitHub imports
voiceSession.ts # voice-mode state machine (tested)
markdown.ts personas… thinking… deviceMemory…
screens/ # ChatList, Chat, Models, Memory, Voice, Settings
Stack: React Native (Expo SDK 57) · llama.rn · TypeScript.
One entry in src/models/catalog.ts:
{
id: 'my-model-1b',
name: 'My Model 1B',
family: 'Vendor',
params: '1B',
quant: 'Q4_K_M',
sizeBytes: 807_694_464, // exact — verify with: curl -sIL <url> | grep -i content-length
url: 'https://huggingface.co/…/resolve/main/….gguf',
description: 'One or two sentences on what it is good at.',
license: 'Apache 2.0',
thinking: false, // true if it emits <think> blocks
}Catalog PRs are welcome, with two constraints: the model must run acceptably on a phone (≤ ~4 GB quantized), and the size must be the exact byte count of the hosted file — the download manager and RAM-fit badges depend on it.
Image and screenshot attachments are now locally grounded through the SmolVLM starter path, and Flight mode now provides five bounded offline activities. The next product work is turning extracted facts into typed action previews, then adding explicit companion milestone saves; PDF/audio decoding remains a separate milestone.
The next product release is E4B-first: make a capable phone immediately feel like a private, multimodal assistant. The flagship loop is:
share something → understand it locally → propose the next action → approve and execute it on the phone.
Near-term priorities:
- Chat UI/UX foundation: professional icons, accessible touch targets, fluid state motion, clean Markdown history previews, model dropdown, and calm left history drawer
- E4B device-fit detection, model recommendation, and offline first-run demo
- Native Share-to-Marmot with action cards for summarize, extract, reply, save, and explain
- Local text/Markdown attachments with bounded grounding and an explicit untrusted-data boundary
- Screenshot/image grounding for E4B-capable devices; [ ] PDF/audio decoding
- Bounded Flight mode with local-only activities and no background work
- Calendar event preview, explicit approval, local-calendar fallback, and undo
- Reminders, contacts, and approval-gated compose actions
- Personal context with projects, grounded sources, and local retrieval
- Voice notes → transcript → decisions/action items → reminders
- Validate the core loop on real E4B hardware, then add one email provider
Web research, MCP, repo coding, file organization, live meeting participation, and deep research remain available as advanced Labs. The full product roadmap and quality bar live in docs/CAPABILITIES.md.
The five-minute contributor guide covers the Windows and macOS build path, focused lanes, issue templates, and privacy-safe evidence. Use the benchmark matrix for reproducible device reports and the model-catalog workflow for new local models. Discussions are for design questions; issues are for reproducible bugs and scoped tasks. Track installs separately from GitHub activity in the distribution ledger.
- llama.cpp — the inference engine
- llama.rn — llama.cpp bindings for React Native
- unsloth and bartowski — the quantized GGUF builds
- Meta, Google, Alibaba, and Hugging Face for the open models
MIT — the app. Each model has its own license (shown in the catalog and in-app) that you accept by downloading it.



