Run Hermes Locally with Ollama — Zero API Cost
تشغيل Hermes محليًا عبر Ollama بلا تكلفة
What this page is, and what it holds.
This page covers Run Hermes Locally with Ollama — Zero API Cost. It carries a source warning and takes about 10 minutes to read. The priciest model is not always best for your task. Compare on one task and set a spend cap.
Step-by-step guide to running Hermes Agent entirely on your own machine with Ollama and open-weight models like Gemma 4, no cloud API keys or paid subscriptions needed
Outcomes taken from this page, not a template.
- Understand what المزوّد والنموذج is and when you need it.
- Run
hermes gatewayandhermes setupand understand what happens next. - Read the table and take only the row that applies to you.
- Set
HERMES_API_TIMEOUTin the right place.
Exactly as they appear in Hermes.
hermes gatewayhermes setuphermes toolshermes skillshermes prompt-size
HERMES_API_TIMEOUTOLLAMA_KEEP_ALIVEYOUR_TELEGRAM_BOT_TOKENYOUR_DISCORD_BOT_TOKENHERMES_STREAM_READ_TIMEOUT
Jump to the part you need.
- 01The Problem
- 02What This Guide Solves
- 03What You Need
- 04Step 1: Install Ollama
- 05Step 2: Pull a Model
- 06Step 3: Configure Hermes
- 07Step 4: Start Using Hermes
- 08Step 5: Pick the Right Model for Your Task
- 09Step 6: Optimize for Speed
- 10Step 7: Run as a Gateway Bot (Optional)
- 11Step 8: Set Up Fallbacks (Optional)
- 12Troubleshooting
- 13Cost Comparison
- 14What Works Well Locally
- 15What's Better with Cloud Models
Nothing summarised away.
The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.
The Problem
A troubleshooting section. Find the symptom that matches yours rather than reading it end to end.
Cloud LLM APIs charge per token. A heavy coding session can cost $5–20. For personal projects, learning, or privacy-sensitive work, that adds up — and you're sending every conversation to a third party.
What This Guide Solves
Explains the idea itself. Read it slowly; the later sections build on it.
You'll set up Hermes Agent running entirely on your own hardware, using Ollama ↗ as the model backend. No API keys, no subscriptions, no data leaving your machine. Once configured, Hermes works exactly like it does with OpenRouter or Anthropic — terminal commands, file editing, web browsing, delegation — but the model runs locally.
By the end, you'll have:
- Ollama serving one or more open-weight models
- Hermes connected to Ollama as a custom endpoint
- A working local agent that can edit files, run commands, and browse the web
- Optional: a Telegram/Discord bot powered entirely by your own hardware
What You Need
A lookup table. Do not read it all; find the row that applies to you.
| Component | Minimum | Recommended |
|---|---|---|
| RAM | 8 GB (for 3B models) | 32+ GB (for 27B+ models) |
| Storage | 5 GB free | 30+ GB (for multiple models) |
| CPU | 4 cores | 8+ cores (AMD EPYC, Ryzen, Intel Xeon) |
| GPU | Not required | NVIDIA GPU with 8+ GB VRAM speeds things up significantly |
Step 1: Install Ollama
Ordered, practical steps. Run one and confirm it worked before moving on.
curl -fsSL https://ollama.com/install.sh | shVerify it's running:
ollama --version
curl http://localhost:11434/api/tags # Should return {"models":[]}Step 2: Pull a Model
Carries a warning. Read it before running anything here. The upstream warning appears below.
Choose based on your hardware:
| Model | Size on Disk | RAM Needed | Tool Calling | Best For |
|---|---|---|---|---|
gemma4:31b | ~20 GB | 24+ GB | Yes | Best quality — strong tool use and reasoning |
gemma2:27b | ~16 GB | 20+ GB | No | Conversational tasks, no tool use |
gemma2:9b | ~5 GB | 8+ GB | No | Fast chat, Q&A — cannot call tools |
llama3.2:3b | ~2 GB | 4+ GB | No | Lightweight quick answers only |
Pull your chosen model:
ollama pull gemma4:31bVerify the model works:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gemma4:31b",
"messages": [{"role": "user", "content": "Say hello"}],
"max_tokens": 50
}'You should see a JSON response with the model's reply.
Step 3: Configure Hermes
Ordered, practical steps. Run one and confirm it worked before moving on. Commands here: hermes setup.
Run the Hermes setup wizard:
hermes setupWhen prompted for a provider, select Custom Endpoint and enter:
- Base URL:
http://localhost:11434/v1 - API Key: Leave empty or type
no-key(Ollama doesn't need one) - Model:
gemma4:31b(or whichever model you pulled)
Alternatively, edit ~/.hermes/config.yaml directly:
model:
default: "gemma4:31b"
provider: "custom"
base_url: "http://localhost:11434/v1"Step 4: Start Using Hermes
Ordered, practical steps. Run one and confirm it worked before moving on.
hermesThat's it. You're now running a fully local agent. Try it out:
You: List all Python files in this directory and count the lines of code in each
You: Read the README.md and summarize what this project does
You: Create a Python script that fetches the weather for Ho Chi Minh CityHermes will use the terminal tool, file operations, and your local model — no cloud calls.
Step 5: Pick the Right Model for Your Task
Ordered, practical steps. Run one and confirm it worked before moving on.
Not every task needs the biggest model. Here's a practical guide:
| Task | Recommended Model | Why |
|---|---|---|
| File edits, code, terminal commands | gemma4:31b | Only model with reliable tool calling |
| Quick Q&A (no tool use needed) | gemma2:9b | Fast responses for conversational tasks |
| Lightweight chat | llama3.2:3b | Fastest, but very limited capabilities |
Switch models on the fly inside a session:
/model gemma2:9bStep 6: Optimize for Speed
Ordered, practical steps. Run one and confirm it worked before moving on.
Increase Ollama's Context Window
By default, Ollama uses a 2048-token context. Hermes requires at least 64,000 tokens for agentic work with tools:
# Create a Modelfile that extends context
cat > /tmp/Modelfile << 'EOF'
FROM gemma4:31b
PARAMETER num_ctx 64000
EOF
ollama create gemma4-64k -f /tmp/ModelfileThen update your Hermes config to use gemma4-64k as the model name.
Keep the Model Loaded
By default, Ollama unloads models after 5 minutes of inactivity. For a persistent gateway bot, keep it loaded:
# Set keep-alive to 24 hours
curl http://localhost:11434/api/generate \
-d '{"model": "gemma4:31b", "keep_alive": "24h"}'Or set it globally in Ollama's environment:
# /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_KEEP_ALIVE=24h"Use GPU Offloading (If Available)
If you have an NVIDIA GPU, Ollama automatically offloads layers to it. Check with:
ollama ps # Shows which model is loaded and how many GPU layersFor a 31B model on a 12 GB GPU, you'll get partial offload (~40 layers on GPU, rest on CPU), which still gives a significant speedup.
Step 7: Run as a Gateway Bot (Optional)
Ordered, practical steps. Run one and confirm it worked before moving on. Commands here: hermes gateway.
Once Hermes works locally in the CLI, you can expose it as a Telegram or Discord bot — still running entirely on your hardware.
Telegram
- Create a bot via @BotFather ↗ and get the token
- Add to your
~/.hermes/config.yaml:
model:
default: "gemma4:31b"
provider: "custom"
base_url: "http://localhost:11434/v1"
platforms:
telegram:
enabled: true
token: "YOUR_TELEGRAM_BOT_TOKEN"- Start the gateway:
hermes gatewayNow message your bot on Telegram — it responds using your local model.
Discord
- Create a Discord application at discord.com/developers ↗
- Add to config:
platforms:
discord:
enabled: true
token: "YOUR_DISCORD_BOT_TOKEN"- Start:
hermes gateway
Step 8: Set Up Fallbacks (Optional)
Ordered, practical steps. Run one and confirm it worked before moving on.
Local models can struggle with complex tasks. Set up a cloud fallback that only activates when the local model fails:
model:
default: "gemma4:31b"
provider: "custom"
base_url: "http://localhost:11434/v1"
fallback_providers:
- provider: openrouter
model: anthropic/claude-sonnet-4This way, 90% of your usage is free (local), and only the hard tasks hit the paid API.
Troubleshooting
A troubleshooting section. Find the symptom that matches yours rather than reading it end to end. Commands here: hermes tools, hermes skills.
"Connection refused" on startup
Ollama isn't running. Start it:
sudo systemctl start ollama
# or
ollama serveSlow responses
- Check model size vs RAM: If your model needs more RAM than available, it swaps to disk. Use a smaller model or add RAM.
- Check
ollama ps: If no GPU layers are offloaded, responses are CPU-bound. This is normal for CPU-only servers. - Reduce context: Large conversations slow down inference. Use
/compressregularly, or set a lower compression threshold in config.
Slow first response (prefill)
Hermes sends a fixed payload on every API call — the system prompt plus the tool schemas for all enabled tools — before any of your conversation content. On CPU-only or low-VRAM setups, processing that prompt (the prefill phase) dominates the first turn: the model can sit silent for minutes while it works through the prompt, then generate at its normal pace. This is expected behaviour, not a hang. The Mac local-LLM guide documents the same effect — during prefill on large contexts, local models may produce no output for minutes while processing the prompt — and Hermes automatically raises its stream read timeout from 120s to 1800s for local endpoints (HERMES_STREAM_READ_TIMEOUT).
What helps:
- Keep the model loaded — Ollama unloads idle models after 5 minutes, adding a full reload before the next prefill. Set
OLLAMA_KEEP_ALIVE=24h(see Step 6 ↗). - Widen the API timeout — set
HERMES_API_TIMEOUT=1800in~/.hermes/.env(see What You Need ↗). - Measure and trim the fixed prompt — run
hermes prompt-sizefor a byte breakdown of the system prompt and tool schemas, then disable unused toolsets withhermes toolsand uninstall skills you don't need withhermes skills. - Use GPU offloading — even a partial offload gives a significant speedup (see Step 6 ↗).
Model doesn't follow tool calls
Models without tool-call support produce plain text instead of structured function calls. Solutions:
- Use a model with tool-call support — of the models listed above, only
gemma4:31bhas reliable tool calling. - Hermes has auto-repair — it detects malformed tool calls and attempts to fix them automatically.
- Set up a fallback — if the local model fails 3 times, Hermes falls back to a cloud provider.
If the model prints raw JSON like {"name": "web_search", ...} in its reply instead of actually running the tool, that's usually the server, not the model — tool calling isn't enabled or the tool-call format isn't parsed. See the per-server fix table in Tool calls appear as text instead of executing (llama.cpp needs --jinja, vLLM needs --enable-auto-tool-choice --tool-call-parser hermes, and so on).
Context window errors
The default Ollama context (2048 tokens) is too small for agentic work. See Step 6 ↗ to increase it.
Cost Comparison
Explains the idea itself. Read it slowly; the later sections build on it.
Here's what running locally saves compared to cloud APIs, based on a typical coding session (~100K tokens input, ~20K tokens output):
| Provider | Cost per Session | Monthly (daily use) |
|---|---|---|
| Anthropic Claude Sonnet | ~$0.80 | ~$24 |
| OpenRouter (GPT-4o) | ~$0.60 | ~$18 |
| Ollama (local) | $0.00 | $0.00 |
Your only cost is electricity — roughly $0.01–0.05 per session depending on hardware.
What Works Well Locally
Explains the idea itself. Read it slowly; the later sections build on it.
- File editing and code generation — models 9B+ handle this well
- Terminal commands — Hermes wraps the command, runs it, reads output regardless of model
- Web browsing — the browser tool does the fetching; the model just interprets results
- Cron jobs and scheduled tasks — work identically to cloud setups
- Multi-platform gateway — Telegram, Discord, Slack all work with local models
What's Better with Cloud Models
Explains the idea itself. Read it slowly; the later sections build on it.
- Very complex multi-step reasoning — 70B+ or cloud models like Claude Opus are noticeably better
- Long context windows — cloud models offer 100K–1M tokens; local runtimes often default below Hermes' 64K minimum unless you configure them
- Speed on large responses — cloud inference is faster than CPU-only local for long generations
The sweet spot: use local for everyday tasks, set up a cloud fallback for the hard stuff.
5 questions answered by this page alone.
Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.