AI outbound calling automation, Google Voice browser control, realtime speech, CRM logging, and Windows app packaging.
Google Voice Dispatch Agent is a Windows-first live calling system that turns Google Voice into an AI-assisted outbound call console. It controls Google Voice in Chrome, dials contacts, waits for real answer evidence, speaks through a virtual audio cable, listens to call audio, transcribes speech, generates replies with Groq, and writes outcomes back to logs and CRM files.
It can run as a console app, a FastAPI operator console, or a packaged Windows EXE with Desktop and Startup shortcuts.
Compliance note: this project can place real calls. You are responsible for consent, caller-ID rules, Do Not Call compliance, recording disclosure, spam prevention, local law, Google Voice terms, and carrier/account limits.
AI calling agent Google Voice automation outbound call automation realtime voice AI FastAPI call console Selenium Google Voice bot Groq voice assistant Python call center automation Windows desktop AI app VB-CABLE voice routing AI dispatcher assistant CRM call logging voicemail detection speech to text calling agent text to speech phone agent
- Browser automation for Google Voice using Selenium and persistent Chrome profiles.
- Live call-state detection with conservative ringing, connected, ended, failed, and voicemail transitions.
- Realtime speech capture with VAD, STT, LLM response generation, and TTS playback.
- Voicemail detection through DOM cues plus early-call audio classifier diagnostics.
- Pre-generated opening line and voicemail audio assets before dialing.
- Call logs, transcripts, recordings, connected-call archive, failed-call archive, and CRM enrichment.
- FastAPI web console for settings, preflight, audio tools, live runs, logs, recordings, leads, and carrier CRM.
- Windows EXE build with PyInstaller.
- Desktop shortcut and Windows Startup shortcut installer.
- Automated tests for state handling, mock dialing, CRM, audio helpers, web routes, and realtime loop behavior.
Most outbound AI voice demos assume a telephony API. This project is different: it works around the practical reality of Google Voice running in a browser. That means the hard parts are not only AI responses. The hard parts are state detection, browser DOM drift, audio routing, ringing safety, voicemail timing, and long-running operational reliability.
The project is designed around those real-world constraints:
- Never speak until the call is actually connected.
- Never treat immediate DOM noise as a completed call.
- Respect a minimum ringing duration.
- Detect voicemail only after connected/greeting evidence.
- Keep full logs so every call can be audited.
- Provide preflight checks before live dialing.
The most important production fix is the conservative call-state gate in src/google_voice.py.
Current behavior:
- Ringing cannot instantly transition to ended.
min_ring_secondsmust pass before accepting connected, ended, or voicemail transitions.- Active call controls seen too early are logged and held in ringing.
- Stale ended banners are ignored during the minimum ringing window.
- Voicemail page/DOM cues are ignored while still ringing.
- Connected state waits for stable evidence such as answered controls or call timer signals.
- The AI opening line starts only after confirmed connection.
max_ring_secondsends the attempt as no answer only after the configured ring timeout.- Logs include current state, elapsed ringing time, DOM cues, answered controls, call-active status, voicemail cues, audio classifier result, and timeout reason.
GoogleVoiceAgent-Active/
src/
main.py CLI runner, safe test, batch calling
google_voice.py Selenium Google Voice automation and call-state detection
call_session.py Call session model and CallState transitions
conversation_loop.py Realtime capture, VAD, STT, LLM, TTS, voicemail handoff
voicemail_detector.py Audio classifier for greeting/beep/live-call signals
ai_groq.py Groq-backed AI response behavior
realtime_tts.py Edge TTS playback and cache integration
audio_capture.py Input capture and VAD framing
audio_routing.py Windows audio-device discovery and playback routing
crm.py Carrier CRM, connected call archive, transcripts, recordings
web_app.py FastAPI operator console and API routes
desktop_app.py PyInstaller app entrypoint
tests/ Automated test suite
scripts/ Build and shortcut installers
packaging/ PyInstaller spec
installer/ Optional Inno Setup installer config
High-level runtime flow:
- Load
.env,dialer_config.json, contacts, and CRM data. - Run preflight checks for Groq, Chrome profile, contacts, callback number, and audio devices.
- Launch Google Voice with the configured Chrome profile.
- Prepare opening line and voicemail audio before dialing.
- Dial one contact.
- Poll
detect_call_state()until connected, voicemail, no answer, ended, or failed. - Start the realtime conversation loop only after confirmed connection.
- Detect voicemail during the early connected-call window and hand off to voicemail playback when confirmed.
- Save call logs, transcript, recording, CRM update, and call archive.
- Apply cooldown and continue to the next contact.
| Layer | Tools |
|---|---|
| Language | Python |
| Web console | FastAPI, Jinja2, HTML, CSS, JavaScript |
| Browser automation | Selenium, webdriver-manager, Google Chrome |
| AI provider | Groq |
| Speech/TTS | Groq STT, Edge TTS, pyttsx3 fallback |
| Audio routing | VB-CABLE, sounddevice, soundcard, soundfile |
| Data | CSV/XLSX contacts, JSON config, local CRM artifacts |
| Desktop packaging | PyInstaller, PowerShell |
| CI/testing | pytest, GitHub Actions |
- Windows 10/11 recommended.
- Python 3.10+ for source runs.
- Google Chrome.
- Google Voice account logged in through the configured Chrome profile.
- VB-CABLE or similar virtual audio device.
- Groq API key.
- Contacts file or CRM data.
- A callback number configured in Google Voice when required by the account.
Install dependencies:
python -m pip install -r requirements.txtCreate local config:
Copy-Item .env.example .env
Copy-Item dialer_config.example.json dialer_config.jsonThen edit .env and dialer_config.json.
Use this when moving SPECTRUM CODEX from this laptop to another Windows PC.
Copy the full project folder to the new PC, for example:
C:\Users\YourName\Downloads\spectrum codex\spectrum codex
Do not copy only src/; keep requirements.txt, setup_project.bat, setup_project.ps1, data/, and the launcher files together.
Double-click:
setup_project.bat
If Windows blocks dependency install or audio packages, right-click setup_project.bat and choose Run as administrator.
The setup script will:
- Use
wingetwhen available to install/check Python 3.12, Google Chrome, Microsoft Visual C++ Redistributable, and FFmpeg. - Check Python and Python version.
- Create
.venvif missing. - Rebuild
.venvif it was copied from another PC and points to an old Python path. - Upgrade pip.
- Install
requirements.txt. - Create missing runtime folders.
- Create
.envonly if it does not already exist. - Create
.env.exampleand a starterdata/contacts.csvif missing. - Check for Chrome, Edge, or Chromium.
- Create launcher files for backend, tests, quickstart, and agent runs.
The installer is safe: it does not delete call data and does not overwrite an existing .env.
VB-CABLE still needs manual installation from https://vb-audio.com/Cable/ if Windows does not already have the virtual audio cable. The installer detects and warns about it, but it does not silently install audio drivers.
Open .env and set at least:
GROQ_API_KEY=your_real_groq_key
CALLBACK_NUMBER=your_google_voice_callback_number
CONTACTS_FILE=data/contacts.csv
LOOPBACK_DEVICE=CABLE Input
CAPTURE_DEVICE=defaultFor Google Voice calling, install/configure VB-CABLE and make sure Chrome has permission to use the correct microphone.
Start the web console:
run_backend.bat
Then open:
http://127.0.0.1:8000
Run tests:
run_tests.bat
Run the Spectrum agent directly:
run_agent.bat
Run the existing quickstart through the virtual environment:
run_quickstart.bat
Before live calling, open the Chrome profile, log into Google Voice, and run the app preflight checks.
PowerShell execution policy blocked:
powershell -NoProfile -ExecutionPolicy Bypass -File .\setup_project.ps1winget not available:
- Install App Installer from Microsoft Store, or install Python 3.12 and Google Chrome manually.
- Then run
setup_project.batagain.
Python 3.12 not found:
- Install Python 3.12 from
https://www.python.org/downloads/release/python-312/. - During install, check Add Python to PATH.
- Close and reopen Command Prompt or PowerShell.
- Run
setup_project.batagain. - Do not use Python 3.14 for this project; some dependencies may not support it yet.
Copied .venv from another computer fails with No Python at old path:
- This happens because virtual environments store absolute paths to the Python install that created them.
- The installer now checks
.venv\Scripts\python.exe; if it cannot run or is not Python 3.12, it removes and recreates.venvautomatically. - If automatic recreation fails, close all terminals using the project and run
setup_project.batagain. - Manual fix:
Remove-Item -Recurse -Force .\.venv
py -3.12 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade pip
.\.venv\Scripts\python.exe -m pip install -r requirements.txtpip install failed:
- Check internet connection.
- Run
setup_project.batas administrator. - Confirm Python 3.12 is installed:
py -3.12 --version- Try:
.\.venv\Scripts\python.exe -m pip install --upgrade pip
.\.venv\Scripts\python.exe -m pip install -r requirements.txt- If audio wheels fail, install Microsoft Visual C++ Redistributable and retry.
sounddevice or soundcard issue:
- Confirm the packages installed:
.\.venv\Scripts\python.exe -m pip show sounddevice soundcard soundfile- Reinstall if needed:
.\.venv\Scripts\python.exe -m pip install sounddevice soundcard soundfile- Run the app Audio page or:
.\.venv\Scripts\python.exe -m src.main --list-audio-devicesChrome driver or Selenium issue:
- Install/update Google Chrome.
- Close all Chrome windows using the same profile.
- Delete only stale
DevToolsActivePortfiles if Chrome crashed, not the whole profile. - Re-run setup so
webdriver-manageris installed. - Open Google Voice manually once and log in before automation.
Microphone or audio device issue:
- Install VB-CABLE.
- Set agent playback to
CABLE Input. - Set Chrome/Google Voice microphone to
CABLE Output. - Keep
CAPTURE_DEVICE=defaultfor simple testing, or use a dedicated second cable for cleaner caller capture. - Run:
.\.venv\Scripts\python.exe -m src.main --audio-route-testIf the app starts but Google Voice does not dial, first confirm the saved Chrome profile is logged into Google Voice on the new PC.
Common .env values:
GROQ_API_KEY=your_key_here
GOOGLE_VOICE_URL=https://voice.google.com/u/0/calls
CHROME_PROFILE_DIR=chrome_profiles/sales_profile
PLAYBACK_DEVICE=CABLE Input
CAPTURE_DEVICE=default
CALLBACK_NUMBER=your_callback_numberCommon dialer_config.json values:
{
"min_ring_seconds": 2,
"max_ring_seconds": 45,
"voicemail_detect_seconds": 15,
"silence_does_not_end_call": true,
"call_cooldown_seconds": 10,
"tts_warmup": true,
"stt_retry_count": 2,
"vad_silence_frames": 12,
"vad_speech_frames": 2
}For 24/7 calling, keep min_ring_seconds above zero, use a cooldown, and do not rapid-fire failed numbers.
Minimum VB-CABLE setup:
- Agent playback device:
CABLE Input. - Chrome microphone input:
CABLE Output. - Chrome speaker output: a device the agent can capture.
Recommended unattended setup:
- Use separate routes for agent TTS and inbound call capture.
- Avoid capturing the same speaker output that plays the agent voice.
- Keep Windows default audio devices stable.
- Re-run audio discovery after Windows updates or driver changes.
Useful commands:
python -m src.main --list-audio-devices
python -m src.main --audio-route-test
python -m src.main --preflightImportant: CAPTURE_DEVICE=default uses Windows speaker loopback. It works for testing, but for nonstop calling a dedicated capture route is more reliable because it reduces echo and self-transcription risk.
Preflight:
python -m src.main --preflightSafe one-number live test:
python -m src.main --safe-test +15551234567Start normal calling:
python -m src.mainStart the web console:
python -m src.web_appOpen:
http://127.0.0.1:8000/run
Build the EXE:
powershell -ExecutionPolicy Bypass -File .\scripts\build_windows_exe.ps1Fast rebuild after tests already passed:
powershell -ExecutionPolicy Bypass -File .\scripts\build_windows_exe.ps1 -SkipTestsBuild outputs:
dist\IndusDispatchConsole.exerelease\IndusDispatchConsole-portable.zip
Install Desktop and Windows Startup shortcuts:
powershell -ExecutionPolicy Bypass -File .\scripts\install_windows_shortcuts.ps1Manual app launch:
.\dist\IndusDispatchConsole.exeCustom port:
.\dist\IndusDispatchConsole.exe --port 8787Startup without opening the browser immediately:
powershell -ExecutionPolicy Bypass -File .\scripts\install_windows_shortcuts.ps1 -StartupNoBrowserThe FastAPI console provides operational pages for:
- Live run control.
- Preflight checks.
- Settings.
- Contacts upload/review.
- Audio diagnostics.
- Logs.
- Leads.
- Recordings.
- Connected calls.
- Carrier CRM.
Primary API and page routes live in src/web_app.py.
Run the full suite:
python -m pytest tests -qCompile check:
python -m compileall src testsRecommended before live calling:
python -m src.main --preflight
python -m src.main --safe-test +15551234567
python -m pytest tests -qCurrent local verification:
231 tests passed
Packaged EXE served /run with HTTP 200
Desktop and Startup shortcuts installed successfully
Runtime files are intentionally not committed:
.envdialer_config.jsonlogs/connected_calls/failed_calls/voicemail_calls/audio/voicemails/audio/scripts/chrome_profiles/dist/build/release/
These can contain secrets, phone numbers, recordings, transcripts, local browser sessions, and generated build artifacts.
For long-running calling:
- Keep the Chrome profile persistent.
- Confirm Google Voice login before dialing batches.
- Use conservative cooldowns between calls.
- Avoid repeated attempts to failed or busy numbers.
- Monitor Groq rate limits.
- Rotate logs and archive recordings.
- Use a dedicated audio route instead of default speaker loopback.
- Watch Google Voice account health.
- Keep a manual Chrome fallback for login challenges.
- Restart the app on a schedule if the browser session drifts.
Recommended future upgrades:
- Silero VAD for stronger speech detection.
- Deepgram for lower-latency realtime STT.
- Langfuse for AI prompt/response tracing.
- OpenTelemetry for runtime metrics.
Tools to delay unless the architecture changes:
- WhisperX is better for offline recordings than live low-latency calls.
- pyannote-audio is useful for diarization, but it adds compute cost.
- LiveKit is powerful, but it would be a larger move away from browser-based Google Voice automation.
Call ends while still ringing:
- Check
min_ring_seconds. - Review
CALL_STATElogs. - Confirm Google Voice DOM did not change.
- Confirm no stale ended banner is being treated as the current call.
AI speaks before pickup:
- Confirm connected evidence appears before
ConversationLoopstarts. - Run a safe test and inspect state transition logs.
Voicemail triggers too early:
- Confirm voicemail cues are ignored while ringing.
- Confirm voicemail is only accepted after connected/greeting/beep evidence.
No audio into Google Voice:
- Set Chrome microphone to
CABLE Output. - Set agent playback to
CABLE Input. - Run
--audio-route-test.
Agent hears itself:
- Avoid
CAPTURE_DEVICE=defaultfor unattended runs. - Use a dedicated inbound capture path.
Chrome login expires:
- Open the configured Chrome profile manually.
- Log into Google Voice.
- Run
python -m src.main --preflight.
Use a number you own or have permission to call:
- Run
python -m src.main --preflight. - Run
python -m src.main --safe-test +15551234567. - Let the phone ring for several seconds.
- Confirm the AI does not speak while ringing.
- Answer the call.
- Confirm the AI speaks only after connection.
- Say a short phrase and confirm it responds.
- Decline/send one call to voicemail.
- Confirm voicemail is detected after greeting or beep.
- Review logs, transcript, recording, and call archive.
This project can place real outbound calls through Google Voice. It is intended for lawful, permission-based business workflows only. Users are responsible for:
- Consent and permission-based outreach before calling
- Do Not Call lists and applicable telemarketing regulations
- Caller ID accuracy and business identification
- Recording disclosure where required by law
- Google Voice terms, carrier/account limits, and spam prevention
- Local privacy, employment, and telecommunications laws
This software must not be used for spam, harassment, deceptive robocalling, or unauthorized calling. The maintainers are not responsible for misuse.
Status: Client-ready · Windows desktop · use demo contacts for testing
The repository keeps source, tests, packaging scripts, workflows, and examples. It intentionally excludes local runtime data, browser profiles, build folders, logs, recordings, generated voicemails, secrets, and private contact data.
Local editor folders, generated test caches, temporary build folders, private runtime data, and one-off planning notes are not part of the product repository.