Academy → Practical GuidesOfficial documentation · Arabic guidance

Run Hermes Locally with Ollama — Zero API Cost

تشغيل Hermes محليًا عبر Ollama بلا تكلفة

Intermediate10 min readLesson 105 questions✓ 2026-08-18
Before you read

What this page is, and what it holds.

This page covers Run Hermes Locally with Ollama — Zero API Cost. It carries a source warning and takes about 10 minutes to read. The priciest model is not always best for your task. Compare on one task and set a spend cap.

15sections
19code examples
4tables
5commands
1,641source words
The official one-line description

Step-by-step guide to running Hermes Agent entirely on your own machine with Ollama and open-weight models like Gemma 4, no cloud API keys or paid subscriptions needed

What you will be able to do

Outcomes taken from this page, not a template.

  • Understand what المزوّد والنموذج is and when you need it.
  • Run hermes gateway and hermes setup and understand what happens next.
  • Read the table and take only the row that applies to you.
  • Set HERMES_API_TIMEOUT in the right place.
Identifiers you will meet

Exactly as they appear in Hermes.

Commands
  • hermes gateway
  • hermes setup
  • hermes tools
  • hermes skills
  • hermes prompt-size
Environment variables
  • HERMES_API_TIMEOUT
  • OLLAMA_KEEP_ALIVE
  • YOUR_TELEGRAM_BOT_TOKEN
  • YOUR_DISCORD_BOT_TOKEN
  • HERMES_STREAM_READ_TIMEOUT
Page map

Jump to the part you need.

  1. 01The Problem
  2. 02What This Guide Solves
  3. 03What You Need
  4. 04Step 1: Install Ollama
  5. 05Step 2: Pull a Model
  6. 06Step 3: Configure Hermes
  7. 07Step 4: Start Using Hermes
  8. 08Step 5: Pick the Right Model for Your Task
  9. 09Step 6: Optimize for Speed
  10. 10Step 7: Run as a Gateway Bot (Optional)
  11. 11Step 8: Set Up Fallbacks (Optional)
  12. 12Troubleshooting
  13. 13Cost Comparison
  14. 14What Works Well Locally
  15. 15What's Better with Cloud Models
The full official page

Nothing summarised away.

The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.

The Problem

A troubleshooting section. Find the symptom that matches yours rather than reading it end to end.

Cloud LLM APIs charge per token. A heavy coding session can cost $5–20. For personal projects, learning, or privacy-sensitive work, that adds up — and you're sending every conversation to a third party.

What This Guide Solves

Explains the idea itself. Read it slowly; the later sections build on it.

You'll set up Hermes Agent running entirely on your own hardware, using Ollama ↗ as the model backend. No API keys, no subscriptions, no data leaving your machine. Once configured, Hermes works exactly like it does with OpenRouter or Anthropic — terminal commands, file editing, web browsing, delegation — but the model runs locally.

By the end, you'll have:

  • Ollama serving one or more open-weight models
  • Hermes connected to Ollama as a custom endpoint
  • A working local agent that can edit files, run commands, and browse the web
  • Optional: a Telegram/Discord bot powered entirely by your own hardware

What You Need

A lookup table. Do not read it all; find the row that applies to you.

ComponentMinimumRecommended
RAM8 GB (for 3B models)32+ GB (for 27B+ models)
Storage5 GB free30+ GB (for multiple models)
CPU4 cores8+ cores (AMD EPYC, Ryzen, Intel Xeon)
GPUNot requiredNVIDIA GPU with 8+ GB VRAM speeds things up significantly

Step 1: Install Ollama

Ordered, practical steps. Run one and confirm it worked before moving on.

Shell1 line
curl -fsSL https://ollama.com/install.sh | sh

Verify it's running:

Shell2 lines
ollama --version
curl http://localhost:11434/api/tags   # Should return {"models":[]}

Step 2: Pull a Model

Carries a warning. Read it before running anything here. The upstream warning appears below.

Choose based on your hardware:

ModelSize on DiskRAM NeededTool CallingBest For
gemma4:31b~20 GB24+ GBYesBest quality — strong tool use and reasoning
gemma2:27b~16 GB20+ GBNoConversational tasks, no tool use
gemma2:9b~5 GB8+ GBNoFast chat, Q&A — cannot call tools
llama3.2:3b~2 GB4+ GBNoLightweight quick answers only

Pull your chosen model:

Shell1 line
ollama pull gemma4:31b

Verify the model works:

Shell7 lines
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma4:31b",
    "messages": [{"role": "user", "content": "Say hello"}],
    "max_tokens": 50
  }'

You should see a JSON response with the model's reply.

Step 3: Configure Hermes

Ordered, practical steps. Run one and confirm it worked before moving on. Commands here: hermes setup.

Run the Hermes setup wizard:

Shell1 line
hermes setup

When prompted for a provider, select Custom Endpoint and enter:

  • Base URL: http://localhost:11434/v1
  • API Key: Leave empty or type no-key (Ollama doesn't need one)
  • Model: gemma4:31b (or whichever model you pulled)

Alternatively, edit ~/.hermes/config.yaml directly:

YAML4 lines
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

Step 4: Start Using Hermes

Ordered, practical steps. Run one and confirm it worked before moving on.

Shell1 line
hermes

That's it. You're now running a fully local agent. Try it out:

Text5 lines
You: List all Python files in this directory and count the lines of code in each

You: Read the README.md and summarize what this project does

You: Create a Python script that fetches the weather for Ho Chi Minh City

Hermes will use the terminal tool, file operations, and your local model — no cloud calls.

Step 5: Pick the Right Model for Your Task

Ordered, practical steps. Run one and confirm it worked before moving on.

Not every task needs the biggest model. Here's a practical guide:

TaskRecommended ModelWhy
File edits, code, terminal commandsgemma4:31bOnly model with reliable tool calling
Quick Q&A (no tool use needed)gemma2:9bFast responses for conversational tasks
Lightweight chatllama3.2:3bFastest, but very limited capabilities

Switch models on the fly inside a session:

Text1 line
/model gemma2:9b

Step 6: Optimize for Speed

Ordered, practical steps. Run one and confirm it worked before moving on.

Increase Ollama's Context Window

By default, Ollama uses a 2048-token context. Hermes requires at least 64,000 tokens for agentic work with tools:

Shell7 lines
# Create a Modelfile that extends context
cat > /tmp/Modelfile << 'EOF'
FROM gemma4:31b
PARAMETER num_ctx 64000
EOF

ollama create gemma4-64k -f /tmp/Modelfile

Then update your Hermes config to use gemma4-64k as the model name.

Keep the Model Loaded

By default, Ollama unloads models after 5 minutes of inactivity. For a persistent gateway bot, keep it loaded:

Shell3 lines
# Set keep-alive to 24 hours
curl http://localhost:11434/api/generate \
  -d '{"model": "gemma4:31b", "keep_alive": "24h"}'

Or set it globally in Ollama's environment:

Shell3 lines
# /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_KEEP_ALIVE=24h"

Use GPU Offloading (If Available)

If you have an NVIDIA GPU, Ollama automatically offloads layers to it. Check with:

Shell1 line
ollama ps   # Shows which model is loaded and how many GPU layers

For a 31B model on a 12 GB GPU, you'll get partial offload (~40 layers on GPU, rest on CPU), which still gives a significant speedup.

Step 7: Run as a Gateway Bot (Optional)

Ordered, practical steps. Run one and confirm it worked before moving on. Commands here: hermes gateway.

Once Hermes works locally in the CLI, you can expose it as a Telegram or Discord bot — still running entirely on your hardware.

Telegram

  1. Create a bot via @BotFather ↗ and get the token
  2. Add to your ~/.hermes/config.yaml:
YAML9 lines
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

platforms:
  telegram:
    enabled: true
    token: "YOUR_TELEGRAM_BOT_TOKEN"
  1. Start the gateway:
Shell1 line
hermes gateway

Now message your bot on Telegram — it responds using your local model.

Discord

  1. Create a Discord application at discord.com/developers ↗
  2. Add to config:
YAML4 lines
platforms:
  discord:
    enabled: true
    token: "YOUR_DISCORD_BOT_TOKEN"
  1. Start: hermes gateway

Step 8: Set Up Fallbacks (Optional)

Ordered, practical steps. Run one and confirm it worked before moving on.

Local models can struggle with complex tasks. Set up a cloud fallback that only activates when the local model fails:

YAML8 lines
model:
  default: "gemma4:31b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4

This way, 90% of your usage is free (local), and only the hard tasks hit the paid API.

Troubleshooting

A troubleshooting section. Find the symptom that matches yours rather than reading it end to end. Commands here: hermes tools, hermes skills.

"Connection refused" on startup

Ollama isn't running. Start it:

Shell3 lines
sudo systemctl start ollama
# or
ollama serve

Slow responses

  • Check model size vs RAM: If your model needs more RAM than available, it swaps to disk. Use a smaller model or add RAM.
  • Check ollama ps: If no GPU layers are offloaded, responses are CPU-bound. This is normal for CPU-only servers.
  • Reduce context: Large conversations slow down inference. Use /compress regularly, or set a lower compression threshold in config.

Slow first response (prefill)

Hermes sends a fixed payload on every API call — the system prompt plus the tool schemas for all enabled tools — before any of your conversation content. On CPU-only or low-VRAM setups, processing that prompt (the prefill phase) dominates the first turn: the model can sit silent for minutes while it works through the prompt, then generate at its normal pace. This is expected behaviour, not a hang. The Mac local-LLM guide documents the same effect — during prefill on large contexts, local models may produce no output for minutes while processing the prompt — and Hermes automatically raises its stream read timeout from 120s to 1800s for local endpoints (HERMES_STREAM_READ_TIMEOUT).

What helps:

  • Keep the model loaded — Ollama unloads idle models after 5 minutes, adding a full reload before the next prefill. Set OLLAMA_KEEP_ALIVE=24h (see Step 6 ↗).
  • Widen the API timeout — set HERMES_API_TIMEOUT=1800 in ~/.hermes/.env (see What You Need ↗).
  • Measure and trim the fixed prompt — run hermes prompt-size for a byte breakdown of the system prompt and tool schemas, then disable unused toolsets with hermes tools and uninstall skills you don't need with hermes skills.
  • Use GPU offloading — even a partial offload gives a significant speedup (see Step 6 ↗).

Model doesn't follow tool calls

Models without tool-call support produce plain text instead of structured function calls. Solutions:

  • Use a model with tool-call support — of the models listed above, only gemma4:31b has reliable tool calling.
  • Hermes has auto-repair — it detects malformed tool calls and attempts to fix them automatically.
  • Set up a fallback — if the local model fails 3 times, Hermes falls back to a cloud provider.

If the model prints raw JSON like {"name": "web_search", ...} in its reply instead of actually running the tool, that's usually the server, not the model — tool calling isn't enabled or the tool-call format isn't parsed. See the per-server fix table in Tool calls appear as text instead of executing (llama.cpp needs --jinja, vLLM needs --enable-auto-tool-choice --tool-call-parser hermes, and so on).

Context window errors

The default Ollama context (2048 tokens) is too small for agentic work. See Step 6 ↗ to increase it.

Cost Comparison

Explains the idea itself. Read it slowly; the later sections build on it.

Here's what running locally saves compared to cloud APIs, based on a typical coding session (~100K tokens input, ~20K tokens output):

ProviderCost per SessionMonthly (daily use)
Anthropic Claude Sonnet~$0.80~$24
OpenRouter (GPT-4o)~$0.60~$18
Ollama (local)$0.00$0.00

Your only cost is electricity — roughly $0.01–0.05 per session depending on hardware.

What Works Well Locally

Explains the idea itself. Read it slowly; the later sections build on it.

  • File editing and code generation — models 9B+ handle this well
  • Terminal commands — Hermes wraps the command, runs it, reads output regardless of model
  • Web browsing — the browser tool does the fetching; the model just interprets results
  • Cron jobs and scheduled tasks — work identically to cloud setups
  • Multi-platform gateway — Telegram, Discord, Slack all work with local models

What's Better with Cloud Models

Explains the idea itself. Read it slowly; the later sections build on it.

  • Very complex multi-step reasoning — 70B+ or cloud models like Claude Opus are noticeably better
  • Long context windows — cloud models offer 100K–1M tokens; local runtimes often default below Hermes' 64K minimum unless you configure them
  • Speed on large responses — cloud inference is faster than CPU-only local for long generations

The sweet spot: use local for everyday tasks, set up a cloud fallback for the hard stuff.

Knowledge check

5 questions answered by this page alone.

Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.

1. In this lesson's table, what is the “Minimum” for “GPU”?
2. Which of these environment variables actually appears in this lesson?
3. Which warning does the source state in this lesson?
4. Which of these headings does not appear in this lesson?
5. Which configuration key appears in this lesson's examples?