Local speech-to-text with speaker diarization and known-speaker recognition. Everything runs on your own machine via faster-whisper and pyannote.audio - no cloud APIs, no audio ever leaves the host.
Built for meeting and conversation recordings, primarily in Ukrainian, with automatic language detection (Whisper is multilingual, so other languages work too).
- You add audio/video files with Add files... (or drop them into the
input/folder) and click Recognize speech. - faster-whisper transcribes the speech.
- If enabled (or later, via Identify speakers), pyannote
speaker-diarization-3.1splits the audio into speaker turns. - Each turn is turned into a voice embedding and matched against a local speaker database, so known people get their real names in the transcript.
- Unrecognized speakers are shown with sample lines; you can name them and optionally save their voiceprint for next time.
- The result is written as plain text to the
output/folder.
If no HuggingFace token is configured, diarization is unavailable and the tool falls back to a plain transcript with timestamps.
- Native desktop GUI (Windows) - no command line needed for everyday use.
- Built-in recording: microphone, system audio (what's playing through your speakers/headset), or both at once, mixed into one file - with live level meters, pause/resume, a 3-2-1 countdown so nothing gets clipped at the start, a system-wide start/stop shortcut, and a recording-in-progress badge on the taskbar icon.
- Fully local - runs offline once models are downloaded; no audio uploaded anywhere.
- Speaker diarization (who spoke when) via pyannote
speaker-diarization-3.1. - Recognition and speaker identification are separate steps: run both at once, or identify speakers later for an already-recognized file, without re-running speech recognition.
- Known-speaker recognition against a local, versioned voiceprint database.
- On-the-fly enrollment: name an unknown speaker and save their voiceprint
(
Name (v1).npy,Name (v2).npy, …), with overwrite / add-version / skip prompts. - No silent downloads - the app always asks before fetching a model, with the recommended choice pre-checked.
- In-app updates: MeetingScribe checks for a newer version on launch and can download and install it with one confirmation, no browser required.
- Light/dark theme (follows Windows automatically, or set manually) and an English/Ukrainian interface language switch.
- Automatic language detection, optimized for Ukrainian (other languages are supported via Whisper's multilingual model).
- CPU or NVIDIA CUDA, auto-detected.
- Skips files already transcribed in
output/(asks before overwriting), so re-running a batch does not redo finished work. - Plain-text output, easy to read and diff.
What you need yourself:
- Windows 10/11.
- An internet connection for installation and the first run (to fetch dependencies and models).
- Optional: a HuggingFace account and token to enable speaker diarization. Without it the tool still works as plain transcription. See Configuration.
Everything else is installed automatically by the installer or the app itself, you do not need to prepare any of it:
- Python 3.12 and ffmpeg (only if not already installed).
- Python dependencies (faster-whisper, pyannote.audio, torch, numpy, huggingface_hub, PySide6) - installed into a private environment the app keeps to itself, never your system Python.
- The CUDA build of torch if an NVIDIA GPU is detected (otherwise it runs on CPU). Auto-detected, nothing to configure.
- The models (see Models and licenses).
- Download the latest installer from the
Releases page
(
MeetingScribe-Setup-<version>.exe) and run it. - It installs Python and ffmpeg via winget if either is missing, sets up a private environment for MeetingScribe (never your system Python), and launches the app.
- On first launch, the app installs its remaining dependencies (with a
progress window) and, if no recognition model is installed yet, asks
before downloading one - nothing happens silently, and the recommended
choice is pre-checked. A small demo recording (
samples/demo.wav) is bundled so there's something to try right away, no recording of your own required.
MeetingScribe checks for updates on launch (and via Help > Check for updates); installing one is a single confirmation away.
Diarization uses gated pyannote models, which require a HuggingFace token with the model licenses accepted.
- Get a token at https://hf.co/settings/tokens.
- Paste it into the app's Settings > Models tab and click "Save token" -
this writes it to
config.envfor you. (You can also editconfig.envdirectly; seeconfig.env.example.HF_TOKENcan also be set as an environment variable, which takes precedence.) - Accept the license for each gated pyannote model while logged in to HuggingFace:
config.env is git-ignored and must never be committed with a real token.
- Launch MeetingScribe from the Start Menu or desktop shortcut (it also opens automatically right after installing).
- Use Add files... in the app (simplest), or put files in the
input/folder. - Select one or more files in the table, pick a recognition model and language.
- Click Recognize speech. With "Automatically identify speakers after recognition" off, it stops after producing a plain transcript - click Identify speakers whenever you're ready to add speaker labels to an already-recognized file, without re-running speech recognition.
- For unrecognized speakers, name them and optionally save their voiceprint for next time.
- Find the transcript in
output/(one.txtper input file).
Installed models, saved speakers, your HuggingFace token, theme, and interface language (English/Українська) are all managed from Settings.
.webm .mp4 .mkv .mov .avi .m4a .mp3 .wav
Use the built-in recorder (see Features) - check Microphone and/or System audio and click Record. It handles the audio format automatically, so there's nothing to configure. Always click Stop to finish, rather than closing the app mid-recording, so the file finalizes correctly.
Plain text, one block per speaker turn:
[Speaker name or Speaker N] mm:ss
spoken text for this turn...
[Another speaker] mm:ss
their spoken text...
If diarization is unavailable, the tool falls back to a continuous transcript with periodic timestamps.
No model weights are stored in this repository. They are downloaded through
the app itself (the first-run prompt for the mandatory model, or Settings >
Models for the rest) into the local models/ folder, and each is covered
by its own license - not by this project's MIT license. You are responsible
for accepting and complying with the terms of every model you download.
These require a HuggingFace account and accepting the conditions on each model page (see Configuration):
| Model | Page | Used for |
|---|---|---|
pyannote/speaker-diarization-3.1 |
https://hf.co/pyannote/speaker-diarization-3.1 | Diarization pipeline |
pyannote/segmentation-3.0 |
https://hf.co/pyannote/segmentation-3.0 | Speech segmentation |
pyannote/wespeaker-voxceleb-resnet34-LM |
https://hf.co/pyannote/wespeaker-voxceleb-resnet34-LM | Diarization embeddings |
pyannote/embedding |
https://hf.co/pyannote/embedding | Known-speaker enrollment & matching |
Public repositories, downloaded under their own terms:
| Model | Page | Notes |
|---|---|---|
Systran/faster-whisper-small |
https://hf.co/Systran/faster-whisper-small | Mandatory - light, fast, multilingual (~460 MB) |
mobiuslabsgmbh/faster-whisper-large-v3-turbo |
https://hf.co/mobiuslabsgmbh/faster-whisper-large-v3-turbo | Optional - higher accuracy (~1.6 GB) |
Systran/faster-whisper-large-v3 |
https://hf.co/Systran/faster-whisper-large-v3 | Optional - best accuracy (~3 GB, slower) |
The underlying Whisper model is by OpenAI (https://github.com/openai/whisper), released under the MIT license.
Mandatory set: the four pyannote models plus Whisper small.
Optional set: Whisper large-v3-turbo and large-v3.
This tool stands on these open-source projects (their own licenses apply):
| Library | Project | License |
|---|---|---|
| faster-whisper | https://github.com/SYSTRAN/faster-whisper | MIT |
| pyannote.audio | https://github.com/pyannote/pyannote-audio | MIT |
| PySide6 (Qt for Python) | https://www.qt.io/qt-for-python | LGPL-3.0 |
| PyTorch | https://github.com/pytorch/pytorch | BSD-3-Clause |
| NumPy | https://github.com/numpy/numpy | BSD-3-Clause |
| huggingface_hub | https://github.com/huggingface/huggingface_hub | Apache-2.0 |
| FFmpeg | https://ffmpeg.org/ | LGPL-2.1+/GPL |
This project's source code is released under the MIT License. Model weights are not covered by this license - see "Models and licenses" above.
