تشغيل Hermes محليًا عبر Ollama بلا تكلفة
Run Hermes Locally with Ollama — Zero API Cost
ما هذه الصفحة، وماذا تحتوي.
المزوّد والنموذج: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes. Hermes نفسه لا يفكّر؛ هو ينظّم العمل ويستدعي النموذج. لذلك اختيار النموذج يحدّد جودة النتيجة وتكلفتها. الصفحة فيها تحذير من المصدر، و10 دقائق قراءة. انتبه: الأغلى ليس دائمًا الأفضل لمهمتك. جرّب مهمة واحدة على نموذجين وقارن، وضع سقفًا للإنفاق من البداية.
Step-by-step guide to running Hermes Agent entirely on your own machine with Ollama and open-weight models like Gemma 4, no cloud API keys or paid subscriptions needed
نتائج مأخوذة من هذه الصفحة، لا من قالب.
- تعرف ما المزوّد والنموذج ولماذا قد تحتاجه.
- تنفّذ
hermes gatewayوhermes setupوتفهم ما يحدث بعدها. - تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
- تضبط
HERMES_API_TIMEOUTفي المكان الصحيح.
كما تظهر تمامًا داخل Hermes.
hermes gatewayhermes setuphermes toolshermes skillshermes prompt-size
HERMES_API_TIMEOUTOLLAMA_KEEP_ALIVEYOUR_TELEGRAM_BOT_TOKENYOUR_DISCORD_BOT_TOKENHERMES_STREAM_READ_TIMEOUT
انتقل مباشرة إلى ما تحتاجه.
- 01The Problem
- 02What This Guide Solves
- 03What You Need
- 04Step 1: Install Ollama
- 05Step 2: Pull a Model
- 06Step 3: Configure Hermes
- 07Step 4: Start Using Hermes
- 08Step 5: Pick the Right Model for Your Task
- 09Step 6: Optimize for Speed
- 10Step 7: Run as a Gateway Bot (Optional)
- 11Step 8: Set Up Fallbacks (Optional)
- 12Troubleshooting
- 13Cost Comparison
- 14What Works Well Locally
- 15What's Better with Cloud Models
بلا اختصار أو حذف.
النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.
The Problem
قسم لحل المشكلات. ابحث فيه عن العطل الذي يشبه حالتك بدل قراءته كاملًا.
Cloud LLM APIs charge per token. A heavy coding session can cost $5–20. For personal projects, learning, or privacy-sensitive work, that adds up — and you're sending every conversation to a third party.
What This Guide Solves
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes.
You'll set up Hermes Agent running entirely on your own hardware, using Ollama ↗ as the model backend. No API keys, no subscriptions, no data leaving your machine. Once configured, Hermes works exactly like it does with OpenRouter or Anthropic — terminal commands, file editing, web browsing, delegation — but the model runs locally.
By the end, you'll have:
- Ollama serving one or more open-weight models
- Hermes connected to Ollama as a custom endpoint
- A working local agent that can edit files, run commands, and browse the web
- Optional: a Telegram/Discord bot powered entirely by your own hardware
What You Need
جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.
| Component | Minimum | Recommended |
|---|---|---|
| RAM | 8 GB (for 3B models) | 32+ GB (for 27B+ models) |
| Storage | 5 GB free | 30+ GB (for multiple models) |
| CPU | 4 cores | 8+ cores (AMD EPYC, Ryzen, Intel Xeon) |
| GPU | Not required | NVIDIA GPU with 8+ GB VRAM speeds things up significantly |
Step 1: Install Ollama
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
curl -fsSL https://ollama.com/install.sh | shVerify it's running:
ollama --version
curl http://localhost:11434/api/tags # Should return {"models":[]}Step 2: Pull a Model
فيه تحذير مهم. اقرأه قبل أن تنفّذ أي شيء من هذا القسم. نصّ التحذير من المصدر مذكور أسفل هذا الشرح.
Choose based on your hardware:
| Model | Size on Disk | RAM Needed | Tool Calling | Best For |
|---|---|---|---|---|
gemma4:31b | ~20 GB | 24+ GB | Yes | Best quality — strong tool use and reasoning |
gemma2:27b | ~16 GB | 20+ GB | No | Conversational tasks, no tool use |
gemma2:9b | ~5 GB | 8+ GB | No | Fast chat, Q&A — cannot call tools |
llama3.2:3b | ~2 GB | 4+ GB | No | Lightweight quick answers only |
Pull your chosen model:
ollama pull gemma4:31bVerify the model works:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gemma4:31b",
"messages": [{"role": "user", "content": "Say hello"}],
"max_tokens": 50
}'You should see a JSON response with the model's reply.
Step 3: Configure Hermes
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية. الأوامر هنا: hermes setup.
Run the Hermes setup wizard:
hermes setupWhen prompted for a provider, select Custom Endpoint and enter:
- Base URL:
http://localhost:11434/v1 - API Key: Leave empty or type
no-key(Ollama doesn't need one) - Model:
gemma4:31b(or whichever model you pulled)
Alternatively, edit ~/.hermes/config.yaml directly:
model:
default: "gemma4:31b"
provider: "custom"
base_url: "http://localhost:11434/v1"Step 4: Start Using Hermes
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
hermesThat's it. You're now running a fully local agent. Try it out:
You: List all Python files in this directory and count the lines of code in each
You: Read the README.md and summarize what this project does
You: Create a Python script that fetches the weather for Ho Chi Minh CityHermes will use the terminal tool, file operations, and your local model — no cloud calls.
Step 5: Pick the Right Model for Your Task
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
Not every task needs the biggest model. Here's a practical guide:
| Task | Recommended Model | Why |
|---|---|---|
| File edits, code, terminal commands | gemma4:31b | Only model with reliable tool calling |
| Quick Q&A (no tool use needed) | gemma2:9b | Fast responses for conversational tasks |
| Lightweight chat | llama3.2:3b | Fastest, but very limited capabilities |
Switch models on the fly inside a session:
/model gemma2:9bStep 6: Optimize for Speed
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
Increase Ollama's Context Window
By default, Ollama uses a 2048-token context. Hermes requires at least 64,000 tokens for agentic work with tools:
# Create a Modelfile that extends context
cat > /tmp/Modelfile << 'EOF'
FROM gemma4:31b
PARAMETER num_ctx 64000
EOF
ollama create gemma4-64k -f /tmp/ModelfileThen update your Hermes config to use gemma4-64k as the model name.
Keep the Model Loaded
By default, Ollama unloads models after 5 minutes of inactivity. For a persistent gateway bot, keep it loaded:
# Set keep-alive to 24 hours
curl http://localhost:11434/api/generate \
-d '{"model": "gemma4:31b", "keep_alive": "24h"}'Or set it globally in Ollama's environment:
# /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_KEEP_ALIVE=24h"Use GPU Offloading (If Available)
If you have an NVIDIA GPU, Ollama automatically offloads layers to it. Check with:
ollama ps # Shows which model is loaded and how many GPU layersFor a 31B model on a 12 GB GPU, you'll get partial offload (~40 layers on GPU, rest on CPU), which still gives a significant speedup.
Step 7: Run as a Gateway Bot (Optional)
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية. الأوامر هنا: hermes gateway.
Once Hermes works locally in the CLI, you can expose it as a Telegram or Discord bot — still running entirely on your hardware.
Telegram
- Create a bot via @BotFather ↗ and get the token
- Add to your
~/.hermes/config.yaml:
model:
default: "gemma4:31b"
provider: "custom"
base_url: "http://localhost:11434/v1"
platforms:
telegram:
enabled: true
token: "YOUR_TELEGRAM_BOT_TOKEN"- Start the gateway:
hermes gatewayNow message your bot on Telegram — it responds using your local model.
Discord
- Create a Discord application at discord.com/developers ↗
- Add to config:
platforms:
discord:
enabled: true
token: "YOUR_DISCORD_BOT_TOKEN"- Start:
hermes gateway
Step 8: Set Up Fallbacks (Optional)
خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.
Local models can struggle with complex tasks. Set up a cloud fallback that only activates when the local model fails:
model:
default: "gemma4:31b"
provider: "custom"
base_url: "http://localhost:11434/v1"
fallback_providers:
- provider: openrouter
model: anthropic/claude-sonnet-4This way, 90% of your usage is free (local), and only the hard tasks hit the paid API.
Troubleshooting
قسم لحل المشكلات. ابحث فيه عن العطل الذي يشبه حالتك بدل قراءته كاملًا. الأوامر هنا: hermes tools، hermes skills.
"Connection refused" on startup
Ollama isn't running. Start it:
sudo systemctl start ollama
# or
ollama serveSlow responses
- Check model size vs RAM: If your model needs more RAM than available, it swaps to disk. Use a smaller model or add RAM.
- Check
ollama ps: If no GPU layers are offloaded, responses are CPU-bound. This is normal for CPU-only servers. - Reduce context: Large conversations slow down inference. Use
/compressregularly, or set a lower compression threshold in config.
Slow first response (prefill)
Hermes sends a fixed payload on every API call — the system prompt plus the tool schemas for all enabled tools — before any of your conversation content. On CPU-only or low-VRAM setups, processing that prompt (the prefill phase) dominates the first turn: the model can sit silent for minutes while it works through the prompt, then generate at its normal pace. This is expected behaviour, not a hang. The Mac local-LLM guide documents the same effect — during prefill on large contexts, local models may produce no output for minutes while processing the prompt — and Hermes automatically raises its stream read timeout from 120s to 1800s for local endpoints (HERMES_STREAM_READ_TIMEOUT).
What helps:
- Keep the model loaded — Ollama unloads idle models after 5 minutes, adding a full reload before the next prefill. Set
OLLAMA_KEEP_ALIVE=24h(see Step 6 ↗). - Widen the API timeout — set
HERMES_API_TIMEOUT=1800in~/.hermes/.env(see What You Need ↗). - Measure and trim the fixed prompt — run
hermes prompt-sizefor a byte breakdown of the system prompt and tool schemas, then disable unused toolsets withhermes toolsand uninstall skills you don't need withhermes skills. - Use GPU offloading — even a partial offload gives a significant speedup (see Step 6 ↗).
Model doesn't follow tool calls
Models without tool-call support produce plain text instead of structured function calls. Solutions:
- Use a model with tool-call support — of the models listed above, only
gemma4:31bhas reliable tool calling. - Hermes has auto-repair — it detects malformed tool calls and attempts to fix them automatically.
- Set up a fallback — if the local model fails 3 times, Hermes falls back to a cloud provider.
If the model prints raw JSON like {"name": "web_search", ...} in its reply instead of actually running the tool, that's usually the server, not the model — tool calling isn't enabled or the tool-call format isn't parsed. See the per-server fix table in Tool calls appear as text instead of executing (llama.cpp needs --jinja, vLLM needs --enable-auto-tool-choice --tool-call-parser hermes, and so on).
Context window errors
The default Ollama context (2048 tokens) is too small for agentic work. See Step 6 ↗ to increase it.
Cost Comparison
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
Here's what running locally saves compared to cloud APIs, based on a typical coding session (~100K tokens input, ~20K tokens output):
| Provider | Cost per Session | Monthly (daily use) |
|---|---|---|
| Anthropic Claude Sonnet | ~$0.80 | ~$24 |
| OpenRouter (GPT-4o) | ~$0.60 | ~$18 |
| Ollama (local) | $0.00 | $0.00 |
Your only cost is electricity — roughly $0.01–0.05 per session depending on hardware.
What Works Well Locally
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
- File editing and code generation — models 9B+ handle this well
- Terminal commands — Hermes wraps the command, runs it, reads output regardless of model
- Web browsing — the browser tool does the fetching; the model just interprets results
- Cron jobs and scheduled tasks — work identically to cloud setups
- Multi-platform gateway — Telegram, Discord, Slack all work with local models
What's Better with Cloud Models
شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.
- Very complex multi-step reasoning — 70B+ or cloud models like Claude Opus are noticeably better
- Long context windows — cloud models offer 100K–1M tokens; local runtimes often default below Hermes' 64K minimum unless you configure them
- Speed on large responses — cloud inference is faster than CPU-only local for long generations
The sweet spot: use local for everyday tasks, set up a cloud fallback for the hard stuff.
5 أسئلة إجاباتها كلها في هذه الصفحة.
كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.