الأكاديمية ← أدلة تطبيقيةتوثيق رسمي · إرشاد عربي

تشغيل Hermes محليًا عبر Ollama بلا تكلفة

Run Hermes Locally with Ollama — Zero API Cost

متوسط10 دقائق قراءةالدرس 105 أسئلة✓ 2026-08-18
قبل أن تقرأ

ما هذه الصفحة، وماذا تحتوي.

المزوّد والنموذج: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes. Hermes نفسه لا يفكّر؛ هو ينظّم العمل ويستدعي النموذج. لذلك اختيار النموذج يحدّد جودة النتيجة وتكلفتها. الصفحة فيها تحذير من المصدر، و10 دقائق قراءة. انتبه: الأغلى ليس دائمًا الأفضل لمهمتك. جرّب مهمة واحدة على نموذجين وقارن، وضع سقفًا للإنفاق من البداية.

15أقسام
19أمثلة برمجية
4جداول
5أوامر
1,641كلمة من المصدر
الوصف الرسمي في سطر

Step-by-step guide to running Hermes Agent entirely on your own machine with Ollama and open-weight models like Gemma 4, no cloud API keys or paid subscriptions needed

ماذا ستستطيع بعدها

نتائج مأخوذة من هذه الصفحة، لا من قالب.

  • تعرف ما المزوّد والنموذج ولماذا قد تحتاجه.
  • تنفّذ hermes gateway وhermes setup وتفهم ما يحدث بعدها.
  • تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
  • تضبط HERMES_API_TIMEOUT في المكان الصحيح.
ما ستقابله من أسماء

كما تظهر تمامًا داخل Hermes.

الأوامر
  • hermes gateway
  • hermes setup
  • hermes tools
  • hermes skills
  • hermes prompt-size
متغيرات البيئة
  • HERMES_API_TIMEOUT
  • OLLAMA_KEEP_ALIVE
  • YOUR_TELEGRAM_BOT_TOKEN
  • YOUR_DISCORD_BOT_TOKEN
  • HERMES_STREAM_READ_TIMEOUT
خريطة الصفحة

انتقل مباشرة إلى ما تحتاجه.

  1. 01The Problem
  2. 02What This Guide Solves
  3. 03What You Need
  4. 04Step 1: Install Ollama
  5. 05Step 2: Pull a Model
  6. 06Step 3: Configure Hermes
  7. 07Step 4: Start Using Hermes
  8. 08Step 5: Pick the Right Model for Your Task
  9. 09Step 6: Optimize for Speed
  10. 10Step 7: Run as a Gateway Bot (Optional)
  11. 11Step 8: Set Up Fallbacks (Optional)
  12. 12Troubleshooting
  13. 13Cost Comparison
  14. 14What Works Well Locally
  15. 15What's Better with Cloud Models
الصفحة الرسمية كاملة

بلا اختصار أو حذف.

النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.

The Problem

قسم لحل المشكلات. ابحث فيه عن العطل الذي يشبه حالتك بدل قراءته كاملًا.

Cloud LLM APIs charge per token. A heavy coding session can cost $5–20. For personal projects, learning, or privacy-sensitive work, that adds up — and you're sending every conversation to a third party.

What This Guide Solves

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes.

You'll set up Hermes Agent running entirely on your own hardware, using Ollama ↗ as the model backend. No API keys, no subscriptions, no data leaving your machine. Once configured, Hermes works exactly like it does with OpenRouter or Anthropic — terminal commands, file editing, web browsing, delegation — but the model runs locally.

By the end, you'll have:

  • Ollama serving one or more open-weight models
  • Hermes connected to Ollama as a custom endpoint
  • A working local agent that can edit files, run commands, and browse the web
  • Optional: a Telegram/Discord bot powered entirely by your own hardware

What You Need

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

ComponentMinimumRecommended
RAM8 GB (for 3B models)32+ GB (for 27B+ models)
Storage5 GB free30+ GB (for multiple models)
CPU4 cores8+ cores (AMD EPYC, Ryzen, Intel Xeon)
GPUNot requiredNVIDIA GPU with 8+ GB VRAM speeds things up significantly

Step 1: Install Ollama

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Shellسطر واحد
curl -fsSL https://ollama.com/install.sh | sh

Verify it's running:

Shellسطران
ollama --version
curl http://localhost:11434/api/tags   # Should return {"models":[]}

Step 2: Pull a Model

فيه تحذير مهم. اقرأه قبل أن تنفّذ أي شيء من هذا القسم. نصّ التحذير من المصدر مذكور أسفل هذا الشرح.

Choose based on your hardware:

ModelSize on DiskRAM NeededTool CallingBest For
gemma4:31b~20 GB24+ GBYesBest quality — strong tool use and reasoning
gemma2:27b~16 GB20+ GBNoConversational tasks, no tool use
gemma2:9b~5 GB8+ GBNoFast chat, Q&A — cannot call tools
llama3.2:3b~2 GB4+ GBNoLightweight quick answers only

Pull your chosen model:

Shellسطر واحد
ollama pull gemma4:31b

Verify the model works:

Shell7 أسطر
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma4:31b",
    "messages": [{"role": "user", "content": "Say hello"}],
    "max_tokens": 50
  }'

You should see a JSON response with the model's reply.

Step 3: Configure Hermes

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية. الأوامر هنا: hermes setup.

Run the Hermes setup wizard:

Shellسطر واحد
hermes setup

When prompted for a provider, select Custom Endpoint and enter:

  • Base URL: http://localhost:11434/v1
  • API Key: Leave empty or type no-key (Ollama doesn't need one)
  • Model: gemma4:31b (or whichever model you pulled)

Alternatively, edit ~/.hermes/config.yaml directly:

YAML4 أسطر
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

Step 4: Start Using Hermes

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Shellسطر واحد
hermes

That's it. You're now running a fully local agent. Try it out:

Text5 أسطر
You: List all Python files in this directory and count the lines of code in each

You: Read the README.md and summarize what this project does

You: Create a Python script that fetches the weather for Ho Chi Minh City

Hermes will use the terminal tool, file operations, and your local model — no cloud calls.

Step 5: Pick the Right Model for Your Task

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Not every task needs the biggest model. Here's a practical guide:

TaskRecommended ModelWhy
File edits, code, terminal commandsgemma4:31bOnly model with reliable tool calling
Quick Q&A (no tool use needed)gemma2:9bFast responses for conversational tasks
Lightweight chatllama3.2:3bFastest, but very limited capabilities

Switch models on the fly inside a session:

Textسطر واحد
/model gemma2:9b

Step 6: Optimize for Speed

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Increase Ollama's Context Window

By default, Ollama uses a 2048-token context. Hermes requires at least 64,000 tokens for agentic work with tools:

Shell7 أسطر
# Create a Modelfile that extends context
cat > /tmp/Modelfile << 'EOF'
FROM gemma4:31b
PARAMETER num_ctx 64000
EOF

ollama create gemma4-64k -f /tmp/Modelfile

Then update your Hermes config to use gemma4-64k as the model name.

Keep the Model Loaded

By default, Ollama unloads models after 5 minutes of inactivity. For a persistent gateway bot, keep it loaded:

Shell3 أسطر
# Set keep-alive to 24 hours
curl http://localhost:11434/api/generate \
  -d '{"model": "gemma4:31b", "keep_alive": "24h"}'

Or set it globally in Ollama's environment:

Shell3 أسطر
# /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_KEEP_ALIVE=24h"

Use GPU Offloading (If Available)

If you have an NVIDIA GPU, Ollama automatically offloads layers to it. Check with:

Shellسطر واحد
ollama ps   # Shows which model is loaded and how many GPU layers

For a 31B model on a 12 GB GPU, you'll get partial offload (~40 layers on GPU, rest on CPU), which still gives a significant speedup.

Step 7: Run as a Gateway Bot (Optional)

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية. الأوامر هنا: hermes gateway.

Once Hermes works locally in the CLI, you can expose it as a Telegram or Discord bot — still running entirely on your hardware.

Telegram

  1. Create a bot via @BotFather ↗ and get the token
  2. Add to your ~/.hermes/config.yaml:
YAML9 أسطر
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

platforms:
  telegram:
    enabled: true
    token: "YOUR_TELEGRAM_BOT_TOKEN"
  1. Start the gateway:
Shellسطر واحد
hermes gateway

Now message your bot on Telegram — it responds using your local model.

Discord

  1. Create a Discord application at discord.com/developers ↗
  2. Add to config:
YAML4 أسطر
platforms:
  discord:
    enabled: true
    token: "YOUR_DISCORD_BOT_TOKEN"
  1. Start: hermes gateway

Step 8: Set Up Fallbacks (Optional)

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Local models can struggle with complex tasks. Set up a cloud fallback that only activates when the local model fails:

YAML8 أسطر
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4

This way, 90% of your usage is free (local), and only the hard tasks hit the paid API.

Troubleshooting

قسم لحل المشكلات. ابحث فيه عن العطل الذي يشبه حالتك بدل قراءته كاملًا. الأوامر هنا: hermes tools، hermes skills.

"Connection refused" on startup

Ollama isn't running. Start it:

Shell3 أسطر
sudo systemctl start ollama
# or
ollama serve

Slow responses

  • Check model size vs RAM: If your model needs more RAM than available, it swaps to disk. Use a smaller model or add RAM.
  • Check ollama ps: If no GPU layers are offloaded, responses are CPU-bound. This is normal for CPU-only servers.
  • Reduce context: Large conversations slow down inference. Use /compress regularly, or set a lower compression threshold in config.

Slow first response (prefill)

Hermes sends a fixed payload on every API call — the system prompt plus the tool schemas for all enabled tools — before any of your conversation content. On CPU-only or low-VRAM setups, processing that prompt (the prefill phase) dominates the first turn: the model can sit silent for minutes while it works through the prompt, then generate at its normal pace. This is expected behaviour, not a hang. The Mac local-LLM guide documents the same effect — during prefill on large contexts, local models may produce no output for minutes while processing the prompt — and Hermes automatically raises its stream read timeout from 120s to 1800s for local endpoints (HERMES_STREAM_READ_TIMEOUT).

What helps:

  • Keep the model loaded — Ollama unloads idle models after 5 minutes, adding a full reload before the next prefill. Set OLLAMA_KEEP_ALIVE=24h (see Step 6 ↗).
  • Widen the API timeout — set HERMES_API_TIMEOUT=1800 in ~/.hermes/.env (see What You Need ↗).
  • Measure and trim the fixed prompt — run hermes prompt-size for a byte breakdown of the system prompt and tool schemas, then disable unused toolsets with hermes tools and uninstall skills you don't need with hermes skills.
  • Use GPU offloading — even a partial offload gives a significant speedup (see Step 6 ↗).

Model doesn't follow tool calls

Models without tool-call support produce plain text instead of structured function calls. Solutions:

  • Use a model with tool-call support — of the models listed above, only gemma4:31b has reliable tool calling.
  • Hermes has auto-repair — it detects malformed tool calls and attempts to fix them automatically.
  • Set up a fallback — if the local model fails 3 times, Hermes falls back to a cloud provider.

If the model prints raw JSON like {"name": "web_search", ...} in its reply instead of actually running the tool, that's usually the server, not the model — tool calling isn't enabled or the tool-call format isn't parsed. See the per-server fix table in Tool calls appear as text instead of executing (llama.cpp needs --jinja, vLLM needs --enable-auto-tool-choice --tool-call-parser hermes, and so on).

Context window errors

The default Ollama context (2048 tokens) is too small for agentic work. See Step 6 ↗ to increase it.

Cost Comparison

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Here's what running locally saves compared to cloud APIs, based on a typical coding session (~100K tokens input, ~20K tokens output):

ProviderCost per SessionMonthly (daily use)
Anthropic Claude Sonnet~$0.80~$24
OpenRouter (GPT-4o)~$0.60~$18
Ollama (local)$0.00$0.00

Your only cost is electricity — roughly $0.01–0.05 per session depending on hardware.

What Works Well Locally

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

  • File editing and code generation — models 9B+ handle this well
  • Terminal commands — Hermes wraps the command, runs it, reads output regardless of model
  • Web browsing — the browser tool does the fetching; the model just interprets results
  • Cron jobs and scheduled tasks — work identically to cloud setups
  • Multi-platform gateway — Telegram, Discord, Slack all work with local models

What's Better with Cloud Models

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

  • Very complex multi-step reasoning — 70B+ or cloud models like Claude Opus are noticeably better
  • Long context windows — cloud models offer 100K–1M tokens; local runtimes often default below Hermes' 64K minimum unless you configure them
  • Speed on large responses — cloud inference is faster than CPU-only local for long generations

The sweet spot: use local for everyday tasks, set up a cloud fallback for the hard stuff.

اختبار الفهم

5 أسئلة إجاباتها كلها في هذه الصفحة.

كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.

1. في جدول هذا الدرس، ما «Minimum» المقابل لـ«GPU»؟
2. أي متغير بيئة من التالي يظهر فعليًا في هذا الدرس؟
3. ما التحذير الذي يذكره المصدر في هذا الدرس؟
4. أي عنوان من التالي لا يظهر في هذا الدرس؟
5. أي مفتاح إعداد يظهر في أمثلة هذا الدرس؟