الأكاديمية ← أدلة تطبيقيةتوثيق رسمي · إرشاد عربي

استخدام الوضع الصوتي

Use Voice Mode with Hermes

متوسط8 دقائق قراءةالدرس 94 أسئلة✓ 2026-08-18
قبل أن تقرأ

ما هذه الصفحة، وماذا تحتوي.

الصوت: أن تتكلّم مع Hermes وتسمع ردّه بدل الكتابة والقراءة. مفيد وأنت تقود أو تمشي أو يداك مشغولتان. ستستعمل هنا hermes gateway وhermes setup tts، والقراءة نحو 8 دقائق. انتبه: الصوت يعني ميكروفونًا يستمع. اعرف متى يكون مفتوحًا، خاصة على جهاز مشترك.

18أقسام
25أمثلة برمجية
2جداول
2أوامر
1,346كلمة من المصدر
الوصف الرسمي في سطر

A practical guide to setting up and using Hermes voice mode across CLI, Telegram, Discord, and Discord voice channels

ماذا ستستطيع بعدها

نتائج مأخوذة من هذه الصفحة، لا من قالب.

  • تعرف ما الصوت ولماذا قد تحتاجه.
  • تنفّذ hermes gateway وhermes setup tts وتفهم ما يحدث بعدها.
  • تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
  • تضبط GROQ_API_KEY في المكان الصحيح.
ما ستقابله من أسماء

كما تظهر تمامًا داخل Hermes.

الأوامر
  • hermes gateway
  • hermes setup tts
متغيرات البيئة
  • GROQ_API_KEY
  • VOICE_TOOLS_OPENAI_KEY
  • ELEVENLABS_API_KEY
  • DISCORD_ALLOWED_USERS
خريطة الصفحة

انتقل مباشرة إلى ما تحتاجه.

  1. 01What voice mode is good for
  2. 02Choose your voice mode setup
  3. 03Step 1: make sure normal Hermes works first
  4. 04Step 2: install the right extras
  5. 05Step 3: install system dependencies
  6. 06Step 4: choose STT and TTS providers
  7. 07Step 5: recommended config
  8. 08Use case 1: CLI voice mode
  9. 09Turn it on
  10. 10Tuning CLI behavior
  11. 11Use case 2: voice replies in Telegram or Discord
  12. 12Use case 3: Discord voice channels
  13. 13Required Discord permissions
  14. 14Join and leave
  15. 15Voice quality recommendations
  16. 16Common failure modes
  17. 17Suggested first-week setup
  18. 18Where to read next
الصفحة الرسمية كاملة

بلا اختصار أو حذف.

النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.

This guide is the practical companion to the Voice Mode feature reference.

If the feature page explains what voice mode can do, this guide shows how to actually use it well.

What voice mode is good for

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: أن تتكلّم مع Hermes وتسمع ردّه بدل الكتابة والقراءة.

Voice mode is especially useful when:

  • you want a hands-free CLI workflow
  • you want spoken responses in Telegram or Discord
  • you want Hermes sitting in a Discord voice channel for live conversation
  • you want quick idea capture, debugging, or back-and-forth while walking around instead of typing

Choose your voice mode setup

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

There are really three different voice experiences in Hermes.

ModeBest forPlatform
Interactive microphone loopPersonal hands-free use while coding or researchingCLI
Voice replies in chatSpoken responses alongside normal messagingTelegram, Discord
Live voice channel botGroup or personal live conversation in a VCDiscord voice channels

A good path is:

  1. get text working first
  2. enable voice replies second
  3. move to Discord voice channels last if you want the full experience

Step 1: make sure normal Hermes works first

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Before touching voice mode, verify that:

  • Hermes starts
  • your provider is configured
  • the agent can answer text prompts normally
Shellسطر واحد
hermes

Ask something simple:

Textسطر واحد
What tools do you have available?

If that is not solid yet, fix text mode first.

Step 2: install the right extras

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

CLI microphone + playback

Shellسطر واحد
cd ~/.hermes/hermes-agent && uv pip install -e ".[voice]"

Messaging platforms

Shellسطر واحد
cd ~/.hermes/hermes-agent && uv pip install -e ".[messaging]"

Premium ElevenLabs TTS

Shellسطر واحد
cd ~/.hermes/hermes-agent && uv pip install -e ".[tts-premium]"

Local NeuTTS (optional)

Shellسطر واحد
python -m pip install -U neutts[all]

Everything

Shellسطر واحد
cd ~/.hermes/hermes-agent && uv pip install -e ".[all]"

Step 3: install system dependencies

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

macOS

Shellسطران
brew install portaudio ffmpeg opus
brew install espeak-ng

Ubuntu / Debian

Shellسطران
sudo apt install portaudio19-dev ffmpeg libopus0
sudo apt install espeak-ng

Why these matter:

  • portaudio → microphone input / playback for CLI voice mode
  • ffmpeg → audio conversion for TTS and messaging delivery
  • opus → Discord voice codec support
  • espeak-ng → phonemizer backend for NeuTTS

Step 4: choose STT and TTS providers

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Hermes supports both local and cloud speech stacks.

Easiest / cheapest setup

Use local STT and free Edge TTS:

  • STT provider: local
  • TTS provider: edge

This is usually the best place to start.

Environment file example

Add to ~/.hermes/.env:

Shell6 أسطر
# Cloud STT options (local needs no key)
GROQ_API_KEY=***
VOICE_TOOLS_OPENAI_KEY=***

# Premium TTS (optional)
ELEVENLABS_API_KEY=***

Provider recommendations

Speech-to-text
  • local → best default for privacy and zero-cost use
  • groq → very fast cloud transcription
  • openai → good paid fallback
Text-to-speech
  • edge → free and good enough for most users
  • neutts → free local/on-device TTS
  • elevenlabs → best quality
  • openai → good middle ground
  • mistral → multilingual, native Opus

If you use hermes setup

If you choose NeuTTS in the setup wizard, Hermes checks whether neutts is already installed. If it is missing, the wizard tells you NeuTTS needs the Python package neutts and the system package espeak-ng, offers to install them for you, installs espeak-ng with your platform package manager, and then runs:

Shellسطر واحد
python -m pip install -U neutts[all]

If you skip that install or it fails, the wizard falls back to Edge TTS.

Use case 1: CLI voice mode

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Turn it on

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Start Hermes:

Shellسطر واحد
hermes

Inside the CLI:

Textسطر واحد
/voice on

Recording flow

Default key:

  • Ctrl+B

Workflow:

  1. press Ctrl+B
  2. speak
  3. wait for silence detection to stop recording automatically
  4. Hermes transcribes and responds
  5. if TTS is on, it speaks the answer
  6. the loop can automatically restart for continuous use

Useful commands

Text5 أسطر
/voice
/voice on
/voice off
/voice tts
/voice status

Good CLI workflows

Walk-up debugging

Say:

Textسطر واحد
I keep getting a docker permission error. Help me debug it.

Then continue hands-free:

  • "Read the last error again"
  • "Explain the root cause in simpler terms"
  • "Now give me the exact fix"
Research / brainstorming

Great for:

  • walking around while thinking
  • dictating half-formed ideas
  • asking Hermes to structure your thoughts in real time
Accessibility / low-typing sessions

If typing is inconvenient, voice mode is one of the fastest ways to stay in the full Hermes loop.

Tuning CLI behavior

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير.

Silence threshold

If Hermes starts/stops too aggressively, tune:

YAMLسطران
voice:
  silence_threshold: 250

Higher threshold = less sensitive.

Silence duration

If you pause a lot between sentences, increase:

YAMLسطران
voice:
  silence_duration: 4.0

Record key

If Ctrl+B conflicts with your terminal or tmux habits:

YAMLسطران
voice:
  record_key: "ctrl+space"

Use case 2: voice replies in Telegram or Discord

أوامر تكتبها في الطرفية. افهم ما يفعله الأمر قبل نسخه. الأوامر هنا: hermes gateway.

This mode is simpler than full voice channels.

Hermes stays a normal chat bot, but can speak replies.

Start the gateway

Shellسطر واحد
hermes gateway

Turn on voice replies

Inside Telegram or Discord:

Textسطر واحد
/voice on

or

Textسطر واحد
/voice tts

Modes

ModeMeaning
offtext only
voice_onlyspeak only when the user sent voice
allspeak every reply

When to use which mode

  • /voice on if you want spoken replies only for voice-originating messages
  • /voice tts if you want a full spoken assistant all the time

Good messaging workflows

Telegram assistant on your phone

Use when:

  • you are away from your machine
  • you want to send voice notes and get quick spoken replies
  • you want Hermes to function like a portable research or ops assistant
Discord DMs with spoken output

Useful when you want private interaction without server-channel mention behavior.

Use case 3: Discord voice channels

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

This is the most advanced mode.

Hermes joins a Discord VC, listens to user speech, transcribes it, runs the normal agent pipeline, and speaks replies back into the channel.

Required Discord permissions

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

In addition to the normal text-bot setup, make sure the bot has:

  • Connect
  • Speak
  • preferably Use Voice Activity

Also enable privileged intents in the Developer Portal:

  • Presence Intent
  • Server Members Intent
  • Message Content Intent

Join and leave

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير. تضبط DISCORD_ALLOWED_USERS خارج المحادثة، في بيئة التشغيل.

In a Discord text channel where the bot is present:

Text3 أسطر
/voice join
/voice leave
/voice status

What happens when joined

  • users speak in the VC
  • Hermes detects speech boundaries
  • transcripts are posted in the associated text channel
  • Hermes responds in text and audio
  • the text channel is the one where /voice join was issued

Best practices for Discord VC use

  • keep DISCORD_ALLOWED_USERS tight
  • use a dedicated bot/testing channel at first
  • verify STT and TTS work in ordinary text-chat voice mode before trying VC mode

Voice quality recommendations

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Best quality setup

  • STT: local large-v3 or Groq whisper-large-v3
  • TTS: ElevenLabs

Best speed / convenience setup

  • STT: local base or Groq
  • TTS: Edge

Best zero-cost setup

  • STT: local
  • TTS: Edge

Common failure modes

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير. تضبط DISCORD_ALLOWED_USERS خارج المحادثة، في بيئة التشغيل.

"No audio device found"

Install portaudio.

"Bot joins but hears nothing"

Check:

  • your Discord user ID is in DISCORD_ALLOWED_USERS
  • you are not muted
  • privileged intents are enabled
  • the bot has Connect/Speak permissions

"It transcribes but does not speak"

Check:

  • TTS provider config
  • API key / quota for ElevenLabs or OpenAI
  • ffmpeg install for Edge conversion paths

"Whisper outputs garbage"

Try:

  • quieter environment
  • higher silence_threshold
  • different STT provider/model
  • shorter, clearer utterances

"It works in DMs but not in server channels"

That is often mention policy.

By default, the bot needs an @mention in Discord server text channels unless configured otherwise.

Suggested first-week setup

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية. الأوامر هنا: hermes setup tts.

If you want the shortest path to success:

  1. get text Hermes working
  2. run hermes setup tts to enable voice support
  3. use CLI voice mode with local STT + Edge TTS
  4. then enable /voice on in Telegram or Discord
  5. only after that, try Discord VC mode

That progression keeps the debugging surface small.

اختبار الفهم

4 أسئلة إجاباتها كلها في هذه الصفحة.

كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.

1. في جدول هذا الدرس، ما «Best for» المقابل لـ«Live voice channel bot»؟
2. أي متغير بيئة من التالي يظهر فعليًا في هذا الدرس؟
3. أي عنوان من التالي لا يظهر في هذا الدرس؟
4. أي مفتاح إعداد يظهر في أمثلة هذا الدرس؟