استخدام الوضع الصوتي
Use Voice Mode with Hermes
ما هذه الصفحة، وماذا تحتوي.
الصوت: أن تتكلّم مع Hermes وتسمع ردّه بدل الكتابة والقراءة. مفيد وأنت تقود أو تمشي أو يداك مشغولتان. ستستعمل هنا hermes gateway وhermes setup tts، والقراءة نحو 8 دقائق. انتبه: الصوت يعني ميكروفونًا يستمع. اعرف متى يكون مفتوحًا، خاصة على جهاز مشترك.
A practical guide to setting up and using Hermes voice mode across CLI, Telegram, Discord, and Discord voice channels
نتائج مأخوذة من هذه الصفحة، لا من قالب.
- تعرف ما الصوت ولماذا قد تحتاجه.
- تنفّذ
hermes gatewayوhermes setup ttsوتفهم ما يحدث بعدها. - تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
- تضبط
GROQ_API_KEYفي المكان الصحيح.
كما تظهر تمامًا داخل Hermes.
hermes gatewayhermes setup tts
GROQ_API_KEYVOICE_TOOLS_OPENAI_KEYELEVENLABS_API_KEYDISCORD_ALLOWED_USERS
انتقل مباشرة إلى ما تحتاجه.
- 01What voice mode is good for
- 02Choose your voice mode setup
- 03Step 1: make sure normal Hermes works first
- 04Step 2: install the right extras
- 05Step 3: install system dependencies
- 06Step 4: choose STT and TTS providers
- 07Step 5: recommended config
- 08Use case 1: CLI voice mode
- 09Turn it on
- 10Tuning CLI behavior
- 11Use case 2: voice replies in Telegram or Discord
- 12Use case 3: Discord voice channels
- 13Required Discord permissions
- 14Join and leave
- 15Voice quality recommendations
- 16Common failure modes
- 17Suggested first-week setup
- 18Where to read next
بلا اختصار أو حذف.
النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.
This guide is the practical companion to the Voice Mode feature reference.
If the feature page explains what voice mode can do, this guide shows how to actually use it well.
What voice mode is good for
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: أن تتكلّم مع Hermes وتسمع ردّه بدل الكتابة والقراءة.
Voice mode is especially useful when:
- you want a hands-free CLI workflow
- you want spoken responses in Telegram or Discord
- you want Hermes sitting in a Discord voice channel for live conversation
- you want quick idea capture, debugging, or back-and-forth while walking around instead of typing
Choose your voice mode setup
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
There are really three different voice experiences in Hermes.
| Mode | Best for | Platform |
|---|---|---|
| Interactive microphone loop | Personal hands-free use while coding or researching | CLI |
| Voice replies in chat | Spoken responses alongside normal messaging | Telegram, Discord |
| Live voice channel bot | Group or personal live conversation in a VC | Discord voice channels |
A good path is:
- get text working first
- enable voice replies second
- move to Discord voice channels last if you want the full experience
Step 1: make sure normal Hermes works first
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
Before touching voice mode, verify that:
- Hermes starts
- your provider is configured
- the agent can answer text prompts normally
hermesAsk something simple:
What tools do you have available?If that is not solid yet, fix text mode first.
Step 2: install the right extras
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
CLI microphone + playback
cd ~/.hermes/hermes-agent && uv pip install -e ".[voice]"Messaging platforms
cd ~/.hermes/hermes-agent && uv pip install -e ".[messaging]"Premium ElevenLabs TTS
cd ~/.hermes/hermes-agent && uv pip install -e ".[tts-premium]"Local NeuTTS (optional)
python -m pip install -U neutts[all]Everything
cd ~/.hermes/hermes-agent && uv pip install -e ".[all]"Step 3: install system dependencies
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
macOS
brew install portaudio ffmpeg opus
brew install espeak-ngUbuntu / Debian
sudo apt install portaudio19-dev ffmpeg libopus0
sudo apt install espeak-ngWhy these matter:
portaudio→ microphone input / playback for CLI voice modeffmpeg→ audio conversion for TTS and messaging deliveryopus→ Discord voice codec supportespeak-ng→ phonemizer backend for NeuTTS
Step 4: choose STT and TTS providers
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
Hermes supports both local and cloud speech stacks.
Easiest / cheapest setup
Use local STT and free Edge TTS:
- STT provider:
local - TTS provider:
edge
This is usually the best place to start.
Environment file example
Add to ~/.hermes/.env:
# Cloud STT options (local needs no key)
GROQ_API_KEY=***
VOICE_TOOLS_OPENAI_KEY=***
# Premium TTS (optional)
ELEVENLABS_API_KEY=***Provider recommendations
Speech-to-text
local→ best default for privacy and zero-cost usegroq→ very fast cloud transcriptionopenai→ good paid fallback
Text-to-speech
edge→ free and good enough for most usersneutts→ free local/on-device TTSelevenlabs→ best qualityopenai→ good middle groundmistral→ multilingual, native Opus
If you use hermes setup
If you choose NeuTTS in the setup wizard, Hermes checks whether neutts is already installed. If it is missing, the wizard tells you NeuTTS needs the Python package neutts and the system package espeak-ng, offers to install them for you, installs espeak-ng with your platform package manager, and then runs:
python -m pip install -U neutts[all]If you skip that install or it fails, the wizard falls back to Edge TTS.
Step 5: recommended config
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
voice:
record_key: "ctrl+b"
submit_mode: "direct" # TUI: direct | draft
max_recording_seconds: 120
auto_tts: false
beep_enabled: true
silence_threshold: 200
silence_duration: 3.0
stt:
provider: "local"
local:
model: "base"
tts:
provider: "edge"
edge:
voice: "en-US-AriaNeural"This is a good conservative default for most people.
In the TUI, voice.submit_mode controls what happens after transcription:
direct(default) submits the transcript immediately.draftputs the transcript in the composer so you can edit or cancel it before pressing Enter.
For editable voice drafts, set:
voice:
submit_mode: "draft"If you want local TTS instead, switch the tts block to:
tts:
provider: "neutts"
neutts:
ref_audio: ''
ref_text: ''
model: neuphonic/neutts-air-q4-gguf
device: cpuUse case 1: CLI voice mode
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
Turn it on
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
Start Hermes:
hermesInside the CLI:
/voice onRecording flow
Default key:
Ctrl+B
Workflow:
- press
Ctrl+B - speak
- wait for silence detection to stop recording automatically
- Hermes transcribes and responds
- if TTS is on, it speaks the answer
- the loop can automatically restart for continuous use
Useful commands
/voice
/voice on
/voice off
/voice tts
/voice statusGood CLI workflows
Walk-up debugging
Say:
I keep getting a docker permission error. Help me debug it.Then continue hands-free:
- "Read the last error again"
- "Explain the root cause in simpler terms"
- "Now give me the exact fix"
Research / brainstorming
Great for:
- walking around while thinking
- dictating half-formed ideas
- asking Hermes to structure your thoughts in real time
Accessibility / low-typing sessions
If typing is inconvenient, voice mode is one of the fastest ways to stay in the full Hermes loop.
Tuning CLI behavior
إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير.
Silence threshold
If Hermes starts/stops too aggressively, tune:
voice:
silence_threshold: 250Higher threshold = less sensitive.
Silence duration
If you pause a lot between sentences, increase:
voice:
silence_duration: 4.0Record key
If Ctrl+B conflicts with your terminal or tmux habits:
voice:
record_key: "ctrl+space"Use case 2: voice replies in Telegram or Discord
أوامر تكتبها في الطرفية. افهم ما يفعله الأمر قبل نسخه. الأوامر هنا: hermes gateway.
This mode is simpler than full voice channels.
Hermes stays a normal chat bot, but can speak replies.
Start the gateway
hermes gatewayTurn on voice replies
Inside Telegram or Discord:
/voice onor
/voice ttsModes
| Mode | Meaning |
|---|---|
off | text only |
voice_only | speak only when the user sent voice |
all | speak every reply |
When to use which mode
/voice onif you want spoken replies only for voice-originating messages/voice ttsif you want a full spoken assistant all the time
Good messaging workflows
Telegram assistant on your phone
Use when:
- you are away from your machine
- you want to send voice notes and get quick spoken replies
- you want Hermes to function like a portable research or ops assistant
Discord DMs with spoken output
Useful when you want private interaction without server-channel mention behavior.
Use case 3: Discord voice channels
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
This is the most advanced mode.
Hermes joins a Discord VC, listens to user speech, transcribes it, runs the normal agent pipeline, and speaks replies back into the channel.
Required Discord permissions
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
In addition to the normal text-bot setup, make sure the bot has:
- Connect
- Speak
- preferably Use Voice Activity
Also enable privileged intents in the Developer Portal:
- Presence Intent
- Server Members Intent
- Message Content Intent
Join and leave
إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير. تضبط DISCORD_ALLOWED_USERS خارج المحادثة، في بيئة التشغيل.
In a Discord text channel where the bot is present:
/voice join
/voice leave
/voice statusWhat happens when joined
- users speak in the VC
- Hermes detects speech boundaries
- transcripts are posted in the associated text channel
- Hermes responds in text and audio
- the text channel is the one where
/voice joinwas issued
Best practices for Discord VC use
- keep
DISCORD_ALLOWED_USERStight - use a dedicated bot/testing channel at first
- verify STT and TTS work in ordinary text-chat voice mode before trying VC mode
Voice quality recommendations
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
Best quality setup
- STT: local
large-v3or Groqwhisper-large-v3 - TTS: ElevenLabs
Best speed / convenience setup
- STT: local
baseor Groq - TTS: Edge
Best zero-cost setup
- STT: local
- TTS: Edge
Common failure modes
إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير. تضبط DISCORD_ALLOWED_USERS خارج المحادثة، في بيئة التشغيل.
"No audio device found"
Install portaudio.
"Bot joins but hears nothing"
Check:
- your Discord user ID is in
DISCORD_ALLOWED_USERS - you are not muted
- privileged intents are enabled
- the bot has Connect/Speak permissions
"It transcribes but does not speak"
Check:
- TTS provider config
- API key / quota for ElevenLabs or OpenAI
ffmpeginstall for Edge conversion paths
"Whisper outputs garbage"
Try:
- quieter environment
- higher
silence_threshold - different STT provider/model
- shorter, clearer utterances
"It works in DMs but not in server channels"
That is often mention policy.
By default, the bot needs an @mention in Discord server text channels unless configured otherwise.
Suggested first-week setup
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية. الأوامر هنا: hermes setup tts.
If you want the shortest path to success:
- get text Hermes working
- run
hermes setup ttsto enable voice support - use CLI voice mode with local STT + Edge TTS
- then enable
/voice onin Telegram or Discord - only after that, try Discord VC mode
That progression keeps the debugging surface small.
Where to read next
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
4 أسئلة إجاباتها كلها في هذه الصفحة.
كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.