Turn any video into actionable insight — AI-powered subtitles, translation, summaries, and chat for any video source.
English | 简体中文
- Subtitle generation — Deepgram speech-to-text or Gemini/OpenAI vision for any language
- Translation — Translate subtitles to Simplified Chinese, Traditional Chinese, or English with natural phrasing
- AI Insights — One-click summary, key information extraction, topic analysis
- Chat with video — Ask anything about the video; AI uses transcript + insights as context
- Subtitle overlay — Draggable, customisable overlay on the video player with fullscreen support
- Picture-in-Picture — Float video over other apps with subtitle overlay (Document PiP, Chrome 116+)
- Recording — Capture screen audio, microphone, or both and auto-transcribe
- YouTube / Bilibili import — Paste a URL to import captions directly
- Cloud sync — Optional Supabase backend for multi-device sync
Prerequisites: Node.js 18+, pnpm (or npm)
# 1. Install dependencies
pnpm install
# 2. Copy the example env file
cp .env.example .env.local
# 3. Fill in at minimum:
# LLM_API_KEY=your_gemini_or_openai_key
# VITE_DEEPGRAM_API_KEY=your_deepgram_key
# 4. Start dev server
pnpm devOpen http://localhost:5173 and drag a video file onto the page.
All variables are set in .env.local (local dev) or Vercel → Project Settings → Environment Variables (production).
| Variable | Required | Description |
|---|---|---|
LLM_API_KEY |
Yes | API key for your LLM provider (Gemini, OpenAI, or compatible) |
LLM_BASE_URL |
No | Base URL of the LLM API. Defaults to Gemini if omitted. |
LLM_MODEL |
No | Model name. Defaults to gemini-2.5-flash. |
VITE_USE_PROXY |
Yes (Vercel) | Set to true to route all LLM calls through /api/proxy (keeps your key server-side) |
How to get an API key:
- Gemini (recommended): https://aistudio.google.com/app/apikey — free tier available, no credit card required
- OpenAI: https://platform.openai.com/api-keys — pay-as-you-go pricing
- OpenRouter (multi-model gateway): https://openrouter.ai/keys — set
LLM_BASE_URL=https://openrouter.ai/api/v1
Common LLM_BASE_URL values:
# Gemini (default — omit LLM_BASE_URL or set to):
https://generativelanguage.googleapis.com
# OpenAI:
https://api.openai.com/v1
# OpenRouter:
https://openrouter.ai/api/v1
# Local Ollama:
http://localhost:11434/v1
| Variable | Required | Description |
|---|---|---|
VITE_DEEPGRAM_API_KEY |
Yes | Deepgram API key for audio transcription |
How to get a Deepgram key:
- Sign up at https://deepgram.com — $200 free credit on sign-up, no credit card required initially
- Go to Console → API Keys → Create Key
- Copy the key and set
VITE_DEEPGRAM_API_KEY
Deepgram handles large audio files automatically. For very long videos (>60 min), audio is split into ≤3.5 MB chunks and transcribed in parallel.
Enables cross-device sync of subtitles, analyses, and chat history.
| Variable | Required | Description |
|---|---|---|
VITE_SUPABASE_URL |
No | Supabase project URL (https://xxxxx.supabase.co) |
VITE_SUPABASE_ANON_KEY |
No | Supabase anonymous/public key |
SUPABASE_SERVICE_ROLE_KEY |
No | Service role key — if set, enforces JWT auth on /api/proxy (recommended for shared deployments) |
How to set up Supabase:
- Create a free project at https://supabase.com
- Go to Project Settings → API and copy Project URL and anon/public key
- For auth gating, also copy the service_role key (keep this server-side only — never expose it to the browser)
- Run the database migrations in
/supabase/migrations/(if any)
When SUPABASE_SERVICE_ROLE_KEY is set, the proxy requires a valid Supabase JWT. Users must log in to use the system API key; otherwise they configure their own key in Settings.
| Variable | Description |
|---|---|
VITE_MODEL |
Model name shown in the Settings UI (cosmetic only) |
VITE_FFMPEG_BASE_URL |
CDN URL for FFmpeg WASM (used for video segmentation). Self-hosted at /ffmpeg by default via postinstall script. |
GEMINI_API_KEY |
Legacy — falls back from LLM_API_KEY |
OPENAI_API_KEY |
Legacy — falls back from LLM_API_KEY; also used for OpenAI Whisper transcription |
CUSTOM_API_KEY |
Legacy — falls back from LLM_API_KEY |
- Push the repo to GitHub
- Import into Vercel: https://vercel.com/new
- Set environment variables in Project Settings → Environment Variables:
LLM_API_KEY = your_gemini_or_openai_key
LLM_BASE_URL = https://generativelanguage.googleapis.com # or your provider
LLM_MODEL = gemini-2.5-flash
VITE_USE_PROXY = true
VITE_DEEPGRAM_API_KEY = your_deepgram_key
# Optional (cloud sync):
VITE_SUPABASE_URL = https://xxxxx.supabase.co
VITE_SUPABASE_ANON_KEY = eyJ...
SUPABASE_SERVICE_ROLE_KEY = eyJ... # enables auth gating
- Deploy — Vercel auto-detects Vite and deploys the serverless functions in
/api/
The app auto-routes requests based on LLM_BASE_URL:
| URL contains | Format used |
|---|---|
generativelanguage.googleapis.com |
Gemini API |
| anything else | OpenAI-compatible API |
This means you can use OpenRouter, Azure OpenAI, Ollama, Together AI, Groq, or any OpenAI-compatible API by setting LLM_BASE_URL accordingly.
Browser (React + Vite)
├── IndexedDB — videos, subtitles, analyses, notes, chat history
├── /api/proxy (Vercel serverless) — LLM calls (Gemini / OpenAI format)
├── /api/deepgram-proxy — Deepgram transcription
├── /api/youtube-captions — YouTube caption import (InnerTube API)
└── Supabase (optional) — auth + cloud sync
pnpm dev # start dev server with API proxy
pnpm build # production build
pnpm typecheck # TypeScript check
pnpm lint # ESLint
pnpm test # Vitest unit tests