Voice Mode
الوضع الصوتي
Start with meaning, then move to detail.
This lesson explains Voice Mode as part of Hermes internals and extension points. You will learn what it does, when it matters, and the smallest safe test that proves it works.
If you are new, do not memorize names. Focus on three questions: what problem does this solve, what access does it need, and how can you verify the result?
For practice, inspect the first example, identify its effects, run it on test data, and compare the result with the source claim.
For advanced readers, inspect Prerequisites, Overview, Requirements, then verify failure modes and version compatibility.
Know Python, Git, and basic project structure before changing code.
A clear outcome before you read.
- Understand Voice Mode without assumed prior knowledge.
- Separate the source description from what still needs testing in your environment.
- Read the first command and identify its inputs and outputs before copying it.
Real-time voice conversations with Hermes Agent — CLI, Telegram, Discord (DMs, text channels, and voice channels)
What does the source say, and in what order?
- 01Prerequisites
Start here to understand the core idea or structure.
- 02Overview
Read this after the foundation, then connect it to the previous step.
- 03Requirements
Read this after the foundation, then connect it to the previous step.
- 04Python Packages
Read this after the foundation, then connect it to the previous step.
- 05System Dependencies
Read this after the foundation, then connect it to the previous step.
- 06API Keys
Read this after the foundation, then connect it to the previous step.
- 07CLI Voice Mode
Read this after the foundation, then connect it to the previous step.
- 08Quick Start
Read this after the foundation, then connect it to the previous step.
- 09How It Works
Read this after the foundation, then connect it to the previous step.
- 10Silence Detection
Finish here to verify the result and special cases.
Copy only after you understand the effect.
# CLI voice mode (microphone + audio playback)
cd ~/.hermes/hermes-agent && uv pip install -e ".[voice]"
# Discord + Telegram messaging (includes discord.py[voice] for VC support)
cd ~/.hermes/hermes-agent && uv pip install -e ".[messaging]"
# Premium TTS (ElevenLabs)
cd ~/.hermes/hermes-agent && uv pip install -e ".[tts-premium]"
# Local TTS (NeuTTS, optional)
python -m pip install -U neutts[all]
# Everything at once
cd ~/.hermes/hermes-agent && uv pip install -e ".[all]"# macOS
brew install portaudio ffmpeg opus
brew install espeak-ng # for NeuTTS
# Ubuntu/Debian
sudo apt install portaudio19-dev ffmpeg libopus0
sudo apt install espeak-ng # for NeuTTS# Speech-to-Text — local provider needs NO key at all
# pip install faster-whisper # Free, runs locally, recommended
GROQ_API_KEY=your-key # Groq Whisper — fast, free tier (cloud)
VOICE_TOOLS_OPENAI_KEY=your-key # OpenAI Whisper — paid (cloud)
# Text-to-Speech (optional — Edge TTS and NeuTTS work without any key)
ELEVENLABS_API_KEY=*** # ElevenLabs — premium quality
# VOICE_TOOLS_OPENAI_KEY above also enables OpenAI TTSRead the first command and identify its inputs and outputs before copying it.
Match every command to your installed Hermes version, review the files and accounts it can reach, and use non-sensitive data for the first test. If this explanation differs from the source, the official source wins.