الأكاديمية ← التكاملات والمزوّدونتوثيق رسمي · إرشاد عربي

مزوّدو نماذج الذكاء الاصطناعي

AI Providers

متوسط إلى متقدم60 دقيقة قراءةالدرس 25 أسئلة✓ 2026-08-18
قبل أن تقرأ

ما هذه الصفحة، وماذا تحتوي.

المزوّد والنموذج: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes. Hermes نفسه لا يفكّر؛ هو ينظّم العمل ويستدعي النموذج. لذلك اختيار النموذج يحدّد جودة النتيجة وتكلفتها. الصفحة فيها تحذير من المصدر، و60 دقيقة قراءة. انتبه: الأغلى ليس دائمًا الأفضل لمهمتك. جرّب مهمة واحدة على نموذجين وقارن، وضع سقفًا للإنفاق من البداية.

7أقسام
91أمثلة برمجية
13جداول
13أوامر
10,126كلمة من المصدر
ماذا ستستطيع بعدها

نتائج مأخوذة من هذه الصفحة، لا من قالب.

  • تعرف ما المزوّد والنموذج ولماذا قد تحتاجه.
  • تنفّذ hermes model وhermes chat وتفهم ما يحدث بعدها.
  • تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
  • تضبط ANTHROPIC_API_KEY في المكان الصحيح.
ما ستقابله من أسماء

كما تظهر تمامًا داخل Hermes.

الأوامر
  • hermes model
  • hermes chat
  • hermes setup
  • hermes migrate xai
  • hermes tools
  • hermes doctor
  • hermes update
  • hermes portal info
متغيرات البيئة
  • ANTHROPIC_API_KEY
  • XAI_API_KEY
  • ANTHROPIC_TOKEN
  • COPILOT_GITHUB_TOKEN
  • GITHUB_TOKEN
  • FIREWORKS_API_KEY
  • NOVITA_API_KEY
  • GLM_API_KEY
خريطة الصفحة

انتقل مباشرة إلى ما تحتاجه.

  1. 01Inference Providers
  2. 02Custom & Self-Hosted LLM Providers
  3. 03Optional API Keys
  4. 04OpenRouter Provider Routing
  5. 05OpenRouter Pareto Code Router
  6. 06Fallback Providers
  7. 07See Also
الصفحة الرسمية كاملة

بلا اختصار أو حذف.

النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.

This page covers setting up inference providers for Hermes Agent — from cloud APIs like OpenRouter and Anthropic, to self-hosted endpoints like Ollama and vLLM, to advanced routing and fallback configurations. You need at least one provider configured to use Hermes.

Inference Providers

فيه تحذير مهم. اقرأه قبل أن تنفّذ أي شيء من هذا القسم. الأوامر هنا: hermes model، hermes portal info. نصّ التحذير من المصدر مذكور أسفل هذا الشرح.

You need at least one way to connect to an LLM. Use hermes model to switch providers and models interactively, or configure directly:

ProviderSetup
Nous Portalhermes model (OAuth, subscription-based)
OpenAI Codexhermes model → ChatGPT or Codex Subscription (ChatGPT OAuth, uses Codex models)
GitHub Copilothermes model (OAuth device code flow, COPILOT_GITHUB_TOKEN, GH_TOKEN, or gh auth token)
GitHub Copilot ACPhermes model (spawns local copilot --acp --stdio)
Anthropichermes model (Claude Max + extra usage credits via OAuth; also supports Anthropic API key or manual setup-token — see note below)
OpenRouterOPENROUTER_API_KEY in ~/.hermes/.env
Fireworks AIFIREWORKS_API_KEY in ~/.hermes/.env (provider: fireworks; aliases: fireworks-ai, fw)
NovitaAINOVITA_API_KEY in ~/.hermes/.env (provider: novita, 200+ models, Model API, Agent Sandbox, GPU Cloud)
AI GatewayAI_GATEWAY_API_KEY in ~/.hermes/.env (provider: ai-gateway)
z.ai / GLMGLM_API_KEY in ~/.hermes/.env (provider: zai)
Kimi / MoonshotKIMI_API_KEY in ~/.hermes/.env (provider: kimi-coding)
Kimi / Moonshot (China)KIMI_CN_API_KEY in ~/.hermes/.env (provider: kimi-coding-cn; aliases: kimi-cn, moonshot-cn)
Arcee AIARCEEAI_API_KEY in ~/.hermes/.env (provider: arcee; aliases: arcee-ai, arceeai)
GMI CloudGMI_API_KEY in ~/.hermes/.env (provider: gmi; aliases: gmi-cloud, gmicloud)
Actual ComputerACTUAL_API_KEY in ~/.hermes/.env for the hosted relay, or ACTUAL_BASE_URL=http://127.0.0.1:8080 for the local daemon — no key needed on loopback (provider: actual; aliases: actual-computer, actualcomputer, aci)
MiniMaxMINIMAX_API_KEY in ~/.hermes/.env (provider: minimax)
MiniMax ChinaMINIMAX_CN_API_KEY in ~/.hermes/.env (provider: minimax-cn)
xAI (Grok) — Responses APIXAI_API_KEY in ~/.hermes/.env (provider: xai)
xAI Grok OAuth (SuperGrok)hermes model → "xAI Grok OAuth (SuperGrok / Premium+)" — browser login, no API key. See guide
Qwen Cloud (Alibaba DashScope)DASHSCOPE_API_KEY in ~/.hermes/.env (provider: alibaba)
Alibaba Cloud (Coding Plan)DASHSCOPE_API_KEY (provider: alibaba-coding-plan, alias: alibaba_coding) — separate billing SKU, different endpoint
Kilo CodeKILOCODE_API_KEY in ~/.hermes/.env (provider: kilocode)
Xiaomi MiMoXIAOMI_API_KEY in ~/.hermes/.env (provider: xiaomi, aliases: mimo, xiaomi-mimo)
Tencent TokenHubTOKENHUB_API_KEY in ~/.hermes/.env (provider: tencent-tokenhub, aliases: tencent, tokenhub, tencentmaas)
OpenCode ZenOPENCODE_ZEN_API_KEY in ~/.hermes/.env (provider: opencode-zen)
CommandCodeCOMMANDCODE_API_KEY in ~/.hermes/.env (provider: commandcode, alias: commandcode-chat; Claude models via commandcode-anthropic, alias: commandcode-claude). Works with GOAT/Pro/Max/Provider plans (not the $1 Go plan — no API access).
OpenCode GoOPENCODE_GO_API_KEY in ~/.hermes/.env (provider: opencode-go)
DeepSeekDEEPSEEK_API_KEY in ~/.hermes/.env (provider: deepseek)
Hugging FaceHF_TOKEN in ~/.hermes/.env (provider: huggingface, aliases: hf)
Google / GeminiGOOGLE_API_KEY (or GEMINI_API_KEY) in ~/.hermes/.env (provider: gemini)
Google Vertex AIhermes model → "Google Vertex AI" (provider: vertex; OAuth2 via service-account JSON or ADC, GCP billing)
OpenAI API (direct)OPENAI_API_KEY in ~/.hermes/.env (provider: openai-api, optional OPENAI_BASE_URL)
Azure AI Foundryhermes model → "Azure AI Foundry" (provider: azure-foundry; uses Azure OpenAI / Foundry endpoint and key)
AWS Bedrockhermes model → "AWS Bedrock" (provider: bedrock; standard AWS credentials chain via boto3)
NVIDIA BuildNVIDIA_API_KEY in ~/.hermes/.env (provider: nvidia; NIM-hosted models on build.nvidia.com)
Ollama Cloudhermes model → "Ollama Cloud" (provider: ollama-cloud; cloud-hosted Ollama API)
Qwen OAuthhermes model → "Qwen OAuth" (provider: qwen-oauth; browser PKCE login)
MiniMax OAuthhermes model → "MiniMax (OAuth)" (provider: minimax-oauth; browser PKCE login)
StepFunSTEPFUN_API_KEY in ~/.hermes/.env (provider: stepfun)
LM Studiohermes model → "LM Studio" (provider: lmstudio, optional LM_API_KEY)
Custom Endpointhermes model → choose "Custom endpoint" (saved in config.yaml)

For the official API-key path, see the dedicated Google Gemini guide.

Nous Portal

Nous Portal ↗ is Nous Research's unified subscription gateway and the recommended way to run Hermes Agent. One OAuth login covers 300+ frontier agentic models (Claude, GPT, Gemini, DeepSeek, Qwen, Kimi, GLM, MiniMax, Grok, ...) plus the Tool Gateway (web search, image generation, TTS, browser automation) — billed against your Nous subscription instead of separate per-provider accounts.

Shell3 أسطر
hermes setup --portal     # fresh install — OAuth + provider + gateway in one command
hermes model              # existing install — pick "Nous Portal" from the list
hermes portal info        # inspect login + routing at any time

Don't have a subscription yet? Get one at portal.nousresearch.com/manage-subscription ↗.

For full details: see the dedicated Nous Portal integration page (what's in the subscription, model catalog, troubleshooting) and the step-by-step Run Hermes Agent with Nous Portal guide.

Client identification. Every Portal request from Hermes Agent carries a client=hermes-client-v<version> tag (e.g. client=hermes-client-v0.13.0) auto-aligned to your installed release. This is sent on all Portal pathways — main chat loop, auxiliary calls, compression summarizer, web extraction — and lets Portal-side telemetry distinguish Hermes traffic from other clients. No config required; the tag updates automatically when you hermes update.

JWT auth (automatic). Hermes prefers scoped inference:invoke JWTs for Portal requests with the legacy opaque session-key path as a fallback. No configuration is required — credentials are managed by the OAuth flow and rotate transparently. Revoked refresh tokens are quarantined to avoid replay loops.

Two Commands for Model Management

Hermes has two model commands that serve different purposes:

CommandWhere to runWhat it does
hermes modelYour terminal (outside any session)Full setup wizard — add providers, run OAuth, enter API keys, configure endpoints
/modelInside a Hermes chat sessionQuick switch between already-configured providers and models

If you're trying to switch to a provider you haven't set up yet (e.g. you only have OpenRouter configured and want to use Anthropic), you need hermes model, not /model. Exit your session first (Ctrl+C or /quit), run hermes model, complete the provider setup, then start a new session.

Subscription plans: what your plan pays for

Several providers let you sign in to Hermes with a consumer subscription (Claude Max, ChatGPT, SuperGrok / X Premium+, …) instead of an API key. What that subscription actually pays for — and what it doesn't — differs per provider, and it's the single most common source of billing surprises. The table below is the short version; each provider's own section has the details.

Cells marked not currently documented mean exactly that: Hermes docs do not yet specify the behavior. Don't assume — check your provider's billing dashboard, and treat these as open questions.
Plan / pathCan Hermes use it?What gets consumedWhat does NOT get consumedCommon surprise
Anthropic — Claude Max + OAuth✅ Yes — hermes model → Anthropic OAuth. Requires Max and purchased extra usage creditsThe extra/overage credits you've added on top of the Max planThe base Max plan allowance (the usage included in Claude Code by default)All Hermes usage bills as "extra usage" even while your included Max allowance sits untouched
Anthropic — Claude Pro❌ No — Pro subscribers cannot use the OAuth pathNothing (path unavailable)Your Pro subscriptionPro looks like it should work; it doesn't. Use an ANTHROPIC_API_KEY instead (pay-per-token, independent of any Claude subscription)
OpenAI Codex — ChatGPT plan OAuth✅ Yes — hermes model → ChatGPT or Codex Subscription (ChatGPT OAuth device-code login, uses Codex models)Not currently documentedNot currently documentedDocs cover auth and token refresh only; plan-quota semantics are not yet documented
xAI — SuperGrok / X Premium+ OAuth✅ Yes — browser OAuth, no API key neededYour subscription quota (documented explicitly for X Search: OAuth is preferred over an API key and "uses your subscription quota instead of API spend"). Inference quota semantics beyond that: not currently documentedXAI_API_KEY / pay-per-token API spend, when OAuth credentials are configured and preferredHTTP 403 after a successful login — xAI has restricted OAuth API access to specific SuperGrok tiers despite an active in-app subscription
Google — Gemini consumer plan (Google AI Pro / Ultra)❌ No documented path — the gemini provider is API-key only (GOOGLE_API_KEY / GEMINI_API_KEY); Vertex AI uses GCP billingYour API key's quota (free tier or billing-enabled Google Cloud project) — consumer-plan consumption not currently documentedNot currently documentedFree-tier keys can be exhausted after a handful of agent turns, because Hermes may make several model calls per user turn

Anthropic. The OAuth path routes as Claude Code against your Anthropic account and only works on a Claude Max plan with purchased extra usage credits — the base Max allowance is never consumed by Hermes, only the extra/overage credits on top. Claude Pro subscribers cannot use this path; the supported alternative is an ANTHROPIC_API_KEY, billed pay-per-token against that key's organization at standard API pricing. See Anthropic (Native) ↗ below.

OpenAI Codex. Hermes authenticates via ChatGPT device-code OAuth, stores credentials in ~/.hermes/auth.json, and can import existing Codex CLI credentials from ~/.codex/auth.json. Which ChatGPT plan tiers are eligible, and how Hermes usage counts against your plan's Codex limits, are not currently documented — the Codex note under Nous Portal ↗ covers authentication and token-refresh behavior only.

xAI (SuperGrok / X Premium+). Browser OAuth works with either an active SuperGrok subscription or an X Premium+ subscription on the linked X account, and the same bearer token is reused by direct-to-xAI tools (TTS, image gen, video gen, transcription, X Search). If inference returns HTTP 403 after a successful login, that's a tier/entitlement restriction on xAI's side, not a stale token — the workaround is switching to an XAI_API_KEY. See xAI (Grok) ↗ below and the xAI Grok OAuth guide.

Google Gemini. There is currently no way to sign in to Hermes with a consumer Gemini subscription — the gemini provider takes an API key, and Google Vertex AI ↗ bills to your GCP project. A billing-enabled Google Cloud project is recommended for agent use; free-tier quotas are too small for long-running agent sessions. See the Google Gemini guide.

Anthropic (Native)

Use Claude models directly through the Anthropic API — no OpenRouter proxy needed. Supports three auth methods:

Shell14 سطرًا
# With an API key (pay-per-token)
export ANTHROPIC_API_KEY=***
hermes chat --provider anthropic --model claude-sonnet-4-6

# Preferred: authenticate through `hermes model`
# Hermes will use Claude Code's credential store directly when available
hermes model

# Manual override with a setup-token (fallback / legacy)
export ANTHROPIC_TOKEN=***  # setup-token or manual OAuth token
hermes chat --provider anthropic

# Auto-detect Claude Code credentials (if you already use Claude Code)
hermes chat --provider anthropic  # reads Claude Code credential files automatically

When you choose Anthropic OAuth through hermes model, Hermes prefers Claude Code's own credential store over copying the token into ~/.hermes/.env. That keeps refreshable Claude credentials refreshable.

Or set it permanently:

YAML3 أسطر
model:
  provider: "anthropic"
  default: "claude-sonnet-4-6"

GitHub Copilot

Hermes supports GitHub Copilot as a first-class provider with two modes:

copilot — Direct Copilot API (recommended). Uses your GitHub Copilot subscription to access GPT-5.x, Claude, Gemini, and other models through the Copilot API.

Shellسطر واحد
hermes chat --provider copilot --model gpt-5.4

Authentication options (checked in this order):

  1. COPILOT_GITHUB_TOKEN environment variable
  2. GH_TOKEN environment variable
  3. GITHUB_TOKEN environment variable
  4. gh auth token CLI fallback

If no token is found, hermes model offers an OAuth device code login — the same flow used by the Copilot CLI and opencode.

API routing: GPT-5+ models (except gpt-5-mini) automatically use the Responses API. All other models (GPT-4o, Claude, Gemini, etc.) use Chat Completions. Models are auto-detected from the live Copilot catalog.

copilot-acp — Copilot ACP agent backend. Spawns the local Copilot CLI as a subprocess:

Shellسطران
hermes chat --provider copilot-acp --model copilot-acp
# Requires the GitHub Copilot CLI in PATH and an existing `copilot login` session

Permanent config:

YAML3 أسطر
model:
  provider: "copilot"
  default: "gpt-5.4"
Environment variableDescription
COPILOT_GITHUB_TOKENGitHub token for Copilot API (first priority)
HERMES_COPILOT_ACP_COMMANDOverride the Copilot CLI binary path (default: copilot)
HERMES_COPILOT_ACP_ARGSOverride ACP args (default: --acp --stdio)

First-Class API-Key Providers

These providers have built-in support with dedicated provider IDs. Set the API key and use --provider to select:

Shell52 سطرًا
# Fireworks AI
hermes chat --provider fireworks --model accounts/fireworks/models/kimi-k2p6
# Requires: FIREWORKS_API_KEY in ~/.hermes/.env

# NovitaAI Model API
hermes chat --provider novita --model moonshotai/kimi-k2.5
# Requires: NOVITA_API_KEY in ~/.hermes/.env

# z.ai / ZhipuAI GLM
hermes chat --provider zai --model glm-5
# Requires: GLM_API_KEY in ~/.hermes/.env

# Kimi / Moonshot AI (international: api.moonshot.ai)
hermes chat --provider kimi-coding --model kimi-for-coding
# Requires: KIMI_API_KEY in ~/.hermes/.env

# Kimi / Moonshot AI (China: api.moonshot.cn)
hermes chat --provider kimi-coding-cn --model kimi-k2.5
# Requires: KIMI_CN_API_KEY in ~/.hermes/.env

# MiniMax (global endpoint)
hermes chat --provider minimax --model MiniMax-M2.7
# Requires: MINIMAX_API_KEY in ~/.hermes/.env

# MiniMax (China endpoint)
hermes chat --provider minimax-cn --model MiniMax-M2.7
# Requires: MINIMAX_CN_API_KEY in ~/.hermes/.env

# Qwen Cloud / DashScope (Qwen models)
hermes chat --provider alibaba --model qwen3.5-plus
# Requires: DASHSCOPE_API_KEY in ~/.hermes/.env

# Xiaomi MiMo
hermes chat --provider xiaomi --model mimo-v2-pro
# Requires: XIAOMI_API_KEY in ~/.hermes/.env

# Tencent TokenHub (Hy3 Preview)
hermes chat --provider tencent-tokenhub --model hy3-preview
# Requires: TOKENHUB_API_KEY in ~/.hermes/.env

# Arcee AI (Trinity models)
hermes chat --provider arcee --model trinity-large-thinking
# Requires: ARCEEAI_API_KEY in ~/.hermes/.env

# Meta Model API (Muse Spark family)
hermes chat --provider meta-ai --model muse-spark-1.2
# Requires: MODEL_API_KEY in ~/.hermes/.env

# GMI Cloud
# Use the exact model ID returned by GMI's /v1/models endpoint.
hermes chat --provider gmi --model zai-org/GLM-5.1-FP8
# Requires: GMI_API_KEY in ~/.hermes/.env

Fireworks uses its native slash-form catalog IDs, such as accounts/fireworks/models/kimi-k2p6. Run hermes model, choose Fireworks AI, and select from the live catalog or enter another Fireworks model ID. The default endpoint is https://api.fireworks.ai/inference/v1; configure a different endpoint through model.base_url in config.yaml, not .env.

Or set the provider permanently in config.yaml:

YAML3 أسطر
model:
  provider: "gmi"
  default: "zai-org/GLM-5.1-FP8"

Base URLs can be overridden with NOVITA_BASE_URL, GLM_BASE_URL, KIMI_BASE_URL, MINIMAX_BASE_URL, MINIMAX_CN_BASE_URL, DASHSCOPE_BASE_URL, XIAOMI_BASE_URL, GMI_BASE_URL, META_BASE_URL, or TOKENHUB_BASE_URL environment variables.

xAI (Grok) — Responses API + Prompt Caching

xAI is wired through the Responses API (codex_responses transport) for automatic reasoning support on Grok 4 models — no reasoning_effort parameter needed, the server reasons by default. Set XAI_API_KEY in ~/.hermes/.env and pick xAI in hermes model, or drop grok as a shortcut into /model grok-4-fast-reasoning.

SuperGrok and X Premium+ subscribers can sign in with browser OAuth instead of using an API key — pick xAI Grok OAuth (SuperGrok / Premium+) in hermes model, or run hermes auth add xai-oauth. The same OAuth bearer token is automatically reused by direct-to-xAI tools (TTS, image gen, video gen, transcription). See the xAI Grok OAuth guide for the full flow — and if Hermes runs on a remote host, also see OAuth over SSH / Remote Hosts for the required ssh -L tunnel.

When using xAI as a provider (any base URL containing x.ai), Hermes automatically enables prompt caching by sending the x-grok-conv-id header with every API request. This routes requests to the same server within a conversation session, allowing xAI's infrastructure to reuse cached system prompts and conversation history.

No configuration is needed — caching activates automatically when an xAI endpoint is detected and a session ID is available. This reduces latency and cost for multi-turn conversations.

xAI also ships a dedicated TTS endpoint (/v1/tts). Select xAI TTS in hermes tools → Voice & TTS, or see the Voice & TTS page for config.

Retired xAI model migration (May 15, 2026): xAI is retiring grok-4*, grok-3, grok-code-fast-1, and grok-imagine-image-pro on 2026-05-15. hermes doctor and hermes chat startup both detect any config still pointing at a retired ref and print the recommended replacement. Use hermes migrate xai for a one-shot config rewrite — dry-run by default, add --apply to write changes (a timestamped config.yaml.bak-pre-migrate-xai-* backup is created automatically).

Shellسطران
hermes migrate xai          # preview replacements
hermes migrate xai --apply  # rewrite ~/.hermes/config.yaml in place

xAI Web Search backend. When the Web Search toolset is enabled, web.backend: xai routes search through xAI's hosted search endpoint using the same XAI_API_KEY / OAuth credentials. No additional setup required if xAI is already configured as a provider.

NovitaAI

NovitaAI ↗ is the AI-native cloud for builders and agents. Its three product lines are Model API for 200+ models, Agent Sandbox for building and running AI agents, and GPU Cloud for scalable compute, all available from one platform.

Shell6 أسطر
# Use any available model
hermes chat --provider novita --model moonshotai/kimi-k2.5
# Requires: NOVITA_API_KEY in ~/.hermes/.env

# Short alias
hermes chat --provider novita-ai --model deepseek/deepseek-v3-0324

Or set it permanently in config.yaml:

YAML4 أسطر
model:
  provider: "novita"
  default: "moonshotai/kimi-k2.5"
  base_url: "https://api.novita.ai/openai/v1"

Get your API key at novita.ai/settings/key-management ↗. The base URL can be overridden with NOVITA_BASE_URL.

Ollama Cloud — Managed Ollama Models, OAuth + API Key

Ollama Cloud ↗ hosts the same open-weight catalog as local Ollama but without the GPU requirement. Pick it in hermes model as Ollama Cloud, paste your API key from ollama.com/settings/keys ↗, and Hermes auto-discovers the available models.

Shell4 أسطر
hermes model
# → pick "Ollama Cloud"
# → paste your OLLAMA_API_KEY
# → select from discovered models (gpt-oss:120b, glm-4.6:cloud, qwen3-coder:480b-cloud, etc.)

Or config.yaml directly:

YAML3 أسطر
model:
  provider: "ollama-cloud"
  default: "gpt-oss:120b"

The model catalog is fetched dynamically from ollama.com/v1/models and cached for one hour. model:tag notation (e.g. qwen3-coder:480b-cloud) is preserved through normalization — don't use dashes.

AWS Bedrock

Anthropic Claude, Amazon Nova, DeepSeek v3.2, Meta Llama 4, and other models via AWS Bedrock. Uses the AWS SDK (boto3) credential chain — no API key, just standard AWS auth.

Shell5 أسطر
# Simplest — named profile in ~/.aws/credentials
hermes chat --provider bedrock --model us.anthropic.claude-sonnet-4-6

# Or with explicit env vars
AWS_PROFILE=myprofile AWS_REGION=us-east-1 hermes chat --provider bedrock --model us.anthropic.claude-sonnet-4-6

Or permanently in config.yaml:

YAML10 أسطر
model:
  provider: "bedrock"
  default: "us.anthropic.claude-sonnet-4-6"
bedrock:
  region: "us-east-1"          # or set AWS_REGION
  # profile: "myprofile"       # or set AWS_PROFILE
  # discovery: true            # auto-discover region from IAM
  # guardrail:                 # optional Bedrock Guardrails
  #   guardrail_identifier: "your-guardrail-id"
  #   guardrail_version: "DRAFT"

Authentication uses the standard boto3 chain: explicit AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, AWS_PROFILE from ~/.aws/credentials, IAM role on EC2/ECS/Lambda, IMDS, or SSO. No env var is required if you're already authenticated with the AWS CLI.

Bedrock uses the Converse API under the hood — requests are translated to Bedrock's model-agnostic shape, so the same config works for Claude, Nova, DeepSeek, and Llama models. Set BEDROCK_BASE_URL only if you're calling a non-default regional endpoint.

See the AWS Bedrock guide for a walkthrough of IAM setup, region selection, and cross-region inference.

Google Vertex AI

Gemini models on Google Cloud Vertex AI via Vertex's OpenAI-compatible endpoint. Authentication is OAuth2 — a short-lived access token (~1 hour) minted from a service-account JSON or Application Default Credentials (ADC). There is no static API key; Hermes mints and auto-refreshes the token for you, including re-minting on a mid-session 401.

Shell6 أسطر
# Service account JSON (recommended for servers / gateways)
echo "VERTEX_CREDENTIALS_PATH=/path/to/service-account.json" >> ~/.hermes/.env
# or Application Default Credentials
gcloud auth application-default login

hermes model   # → "Google Vertex AI" → project → region → model

Or in config.yaml (project/region are non-secret and live here; the credential path stays in .env):

YAML6 أسطر
model:
  provider: "vertex"
  default: "google/gemini-3-flash-preview"   # Vertex requires the google/ prefix
vertex:
  project_id: "my-gcp-project"   # blank → use the project embedded in the credentials
  region: "global"               # required for the Gemini 3.x previews

VERTEX_PROJECT_ID / VERTEX_REGION env vars override the config.yaml values. Hermes lazy-installs google-auth on first use; run hermes setup if the managed install needs repair. See the Google Vertex AI guide for the full walkthrough, and the Google Gemini guide for the static-API-key AI Studio path instead.

Qwen Portal (OAuth)

Alibaba's Qwen Portal with browser-based OAuth login. Pick Qwen OAuth (Portal) in hermes model, sign in through the browser, and Hermes persists the refresh token.

Shell6 أسطر
hermes model
# → pick "Qwen OAuth (Portal)"
# → browser opens; sign in with your Alibaba account
# → confirm — credentials are saved to ~/.hermes/auth.json

hermes chat   # uses portal.qwen.ai/v1 endpoint

Or configure config.yaml:

YAML3 أسطر
model:
  provider: "qwen-oauth"
  default: "qwen3-coder-plus"

Set HERMES_QWEN_BASE_URL only if the portal endpoint relocates (default: https://portal.qwen.ai/v1).

Alibaba Cloud (Coding Plan)

If you're subscribed to Alibaba's Coding Plan (a pricing SKU separate from standard DashScope API access), Hermes exposes it as its own first-class provider: alibaba-coding-plan. Endpoint: https://coding-intl.dashscope.aliyuncs.com/v1. It's OpenAI-compatible like the regular alibaba provider but with a different base URL and billing surface.

YAML3 أسطر
model:
  provider: alibaba_coding     # alias for alibaba-coding-plan
  model: qwen3-coder-plus

Or from the CLI:

Shellسطر واحد
hermes chat --provider alibaba_coding --model qwen3-coder-plus

alibaba_coding uses the same DASHSCOPE_API_KEY your alibaba entry already uses — no separate key needed, just a different routing target. Before this provider was registered, users who set provider: alibaba_coding in config.yaml silently fell through to OpenRouter routing.

MiniMax (OAuth)

MiniMax-M2.7 via browser OAuth login — no API key needed. Pick MiniMax (OAuth) in hermes model, sign in through the browser, and Hermes persists the access + refresh tokens. Uses the Anthropic Messages-compatible endpoint (/anthropic) under the hood.

Shell6 أسطر
hermes model
# → pick "MiniMax (OAuth)"
# → browser opens; sign in with your MiniMax account (global or CN region)
# → confirm — credentials are saved to ~/.hermes/auth.json

hermes chat   # uses api.minimax.io/anthropic endpoint

Or configure config.yaml:

YAML3 أسطر
model:
  provider: "minimax-oauth"
  default: "MiniMax-M2.7"

Supported models: MiniMax-M2.7 (main) and MiniMax-M2.7-highspeed (wired as the default auxiliary model). The OAuth path ignores MINIMAX_API_KEY / MINIMAX_BASE_URL.

NVIDIA NIM

Nemotron and other open source models via build.nvidia.com ↗ (free API key) or a local NIM endpoint.

Shell6 أسطر
# Cloud (build.nvidia.com)
hermes chat --provider nvidia --model nvidia/nemotron-3-super-120b-a12b
# Requires: NVIDIA_API_KEY in ~/.hermes/.env

# Local NIM endpoint — override base URL
NVIDIA_BASE_URL=http://localhost:8000/v1 hermes chat --provider nvidia --model nvidia/nemotron-3-super-120b-a12b

Or set it permanently in config.yaml:

YAML3 أسطر
model:
  provider: "nvidia"
  default: "nvidia/nemotron-3-super-120b-a12b"

Hermes automatically attaches the NIM billing-origin header on every request to build.nvidia.com — no configuration needed. This routes consumption against the correct origin in NVIDIA's billing dashboard.

GMI Cloud

Open and reasoning models via GMI Cloud ↗ — OpenAI-compatible API, API key authentication.

Shell3 أسطر
# GMI Cloud
hermes chat --provider gmi --model deepseek-ai/DeepSeek-V3.2
# Requires: GMI_API_KEY in ~/.hermes/.env

Or set it permanently in config.yaml:

YAML3 أسطر
model:
  provider: "gmi"
  default: "deepseek-ai/DeepSeek-V3.2"

The base URL can be overridden with GMI_BASE_URL (default: https://api.gmi-serving.com/v1).

Actual Computer

Your own hardware as a private inference cluster via Actual Computer ↗. Two serving modes, both OpenAI-compatible (Hermes uses the Responses API transport):

  • Hosted relay — https://api.actual.inc, end-to-end encrypted, routes to your cluster. Authenticate with an ac_ inference key from actual.inc/user/keys ↗.
  • Local daemon — on-device at http://127.0.0.1:8080, fully offline. No API key needed: Hermes detects the loopback base URL and authenticates with an internal placeholder automatically.
Shell5 أسطر
# Hosted relay (ACTUAL_API_KEY in ~/.hermes/.env)
hermes chat --provider actual --model <model-id-from-your-cluster>

# Local daemon (ACTUAL_BASE_URL=http://127.0.0.1:8080 in ~/.hermes/.env, no key)
hermes chat --provider actual --model <installed-model-name>

Or set it permanently in config.yaml:

YAML3 أسطر
model:
  provider: "actual"
  default: "<model-id>"

Notes:

  • Model IDs come from your cluster's GET /v1/models — discover with hermes model or curl -s https://api.actual.inc/v1/models -H "Authorization: Bearer $ACTUAL_API_KEY".
  • Bare hosts are normalized: ACTUAL_BASE_URL=http://127.0.0.1:8080 becomes http://127.0.0.1:8080/v1 automatically.
  • Reasoning effort is clamped to Actual's supported range (none/low/medium/high/max) — a global xhigh/ultra setting will not 400 requests.
  • Small local models: Hermes' full default toolset plus the system prompt can exceed a 32k context window, producing an empty-stream error from llama.cpp-family servers. Restrict the toolset (-t file,web) or load the model with a larger context. The optional actual-setup skill (hermes skills install official/devops/actual-setup) covers setup and troubleshooting in detail.
  • Aliases: actual-computer, actualcomputer, aci.

StepFun

Step-series models via StepFun ↗ — OpenAI-compatible API, API key authentication.

Shell3 أسطر
# StepFun
hermes chat --provider stepfun --model step-3.5-flash
# Requires: STEPFUN_API_KEY in ~/.hermes/.env

Or set it permanently in config.yaml:

YAML3 أسطر
model:
  provider: "stepfun"
  default: "step-3.5-flash"

The base URL can be overridden with STEPFUN_BASE_URL (default: https://api.stepfun.com/v1).

Hugging Face Inference Providers

Hugging Face Inference Providers ↗ routes to 20+ open models through a unified OpenAI-compatible endpoint (router.huggingface.co/v1). Requests are automatically routed to the fastest available backend (Groq, Together, SambaNova, etc.) with automatic failover.

Shell6 أسطر
# Use any available model
hermes chat --provider huggingface --model Qwen/Qwen3.5-397B-A17B
# Requires: HF_TOKEN in ~/.hermes/.env

# Short alias
hermes chat --provider hf --model deepseek-ai/DeepSeek-V3.2

Or set it permanently in config.yaml:

YAML3 أسطر
model:
  provider: "huggingface"
  default: "Qwen/Qwen3.5-397B-A17B"

Get your token at huggingface.co/settings/tokens ↗ — make sure to enable the "Make calls to Inference Providers" permission. Free tier included ($0.10/month credit, no markup on provider rates).

You can append routing suffixes to model names: :fastest (default), :cheapest, or :provider_name to force a specific backend.

The base URL can be overridden with HF_BASE_URL.

Custom & Self-Hosted LLM Providers

فيه تحذير مهم. اقرأه قبل أن تنفّذ أي شيء من هذا القسم. الأوامر هنا: hermes model، hermes config set model. نصّ التحذير من المصدر مذكور أسفل هذا الشرح.

Hermes Agent works with any OpenAI-compatible API endpoint. If a server implements /v1/chat/completions, you can point Hermes at it. This means you can use local models, GPU inference servers, multi-provider routers, or any third-party API.

General Setup

Three ways to configure a custom endpoint:

Interactive setup (recommended):

Shell3 أسطر
hermes model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter: API base URL, API key, Model name

Manual config (config.yaml):

YAML6 أسطر
# In ~/.hermes/config.yaml
model:
  default: your-model-name
  provider: custom
  base_url: http://localhost:8000/v1
  api_key: your-key-or-leave-empty-for-local

Both approaches persist to config.yaml, which is the source of truth for model, provider, and base URL.

Switching Models with /model

Once you have at least one custom endpoint configured, you can switch models mid-session:

Text3 أسطر
/model custom:qwen-2.5          # Switch to a model on your custom endpoint
/model custom                    # Auto-detect the model from the endpoint
/model openrouter:claude-sonnet-4 # Switch back to a cloud provider

If you have named custom providers configured (see below), use the triple syntax:

Textسطران
/model custom:local:qwen-2.5    # Use the "local" custom provider with model qwen-2.5
/model custom:work:llama3       # Use the "work" custom provider with llama3

When switching providers, Hermes persists the base URL and provider to config so the change survives restarts. When switching away from a custom endpoint to a built-in provider, the stale base URL is automatically cleared.

Everything below follows this same pattern — just change the URL, key, and model name.

---

Ollama — Local Models, Zero Config

Ollama ↗ runs open-weight models locally with one command. Best for: quick local experimentation, privacy-sensitive work, offline use. Supports tool calling via the OpenAI-compatible API.

Shell3 أسطر
# Install and run a model
ollama pull qwen2.5-coder:32b
ollama serve   # Starts on port 11434

Then configure Hermes:

Shell5 أسطر
hermes model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:11434/v1
# Skip API key (Ollama doesn't need one)
# Enter model name (e.g. qwen2.5-coder:32b)

Or configure config.yaml directly:

YAML5 أسطر
model:
  default: qwen2.5-coder:32b
  provider: custom
  base_url: http://localhost:11434/v1
  context_length: 64000   # See warning below

Verify your context is set correctly:

Shellسطران
ollama ps
# Look at the CONTEXT column — it should show your configured value

---

vLLM — High-Performance GPU Inference

vLLM ↗ is the standard for production LLM serving. Best for: maximum throughput on GPU hardware, serving large models, continuous batching.

Shell7 أسطر
pip install vllm
vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --port 8000 \
  --max-model-len 65536 \
  --tensor-parallel-size 2 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

Then configure Hermes:

Shell5 أسطر
hermes model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:8000/v1
# Skip API key (or enter one if you configured vLLM with --api-key)
# Enter model name: meta-llama/Llama-3.1-70B-Instruct

Context length: vLLM reads the model's max_position_embeddings by default. If that exceeds your GPU memory, it errors and asks you to set --max-model-len lower. You can also use --max-model-len auto to automatically find the maximum that fits. Set --gpu-memory-utilization 0.95 (default 0.9) to squeeze more context into VRAM.

Tool calling requires explicit flags:

FlagPurpose
--enable-auto-tool-choiceRequired for tool_choice: "auto" (the default in Hermes)
--tool-call-parser <name>Parser for the model's tool call format

Supported parsers: hermes (Qwen 2.5, Hermes 2/3), llama3_json (Llama 3.x), mistral, deepseek_v3, deepseek_v31, xlam, pythonic. Without these flags, tool calls won't work — the model will output tool calls as text.

Qwen reasoning parsers: Hermes preserves structured reasoning metadata such as reasoning, reasoning_content, and streamed reasoning deltas when OpenAI-compatible servers return them. That metadata is treated as reasoning/thinking trace data, not as a replacement for the assistant's visible answer. For Qwen reasoning models served by vLLM, make sure the final user-visible response still appears in content. If --reasoning-parser qwen3 leaves content empty in your deployment, either disable that parser or pass a server-supported request option such as chat_template_kwargs.enable_thinking: false through extra_body.

---

SGLang — Fast Serving with RadixAttention

SGLang ↗ is an alternative to vLLM with RadixAttention for KV cache reuse. Best for: multi-turn conversations (prefix caching), constrained decoding, structured output.

Shell7 أسطر
pip install "sglang[all]"
python -m sglang.launch_server \
  --model meta-llama/Llama-3.1-70B-Instruct \
  --port 30000 \
  --context-length 65536 \
  --tp 2 \
  --tool-call-parser qwen

Then configure Hermes:

Shell4 أسطر
hermes model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:30000/v1
# Enter model name: meta-llama/Llama-3.1-70B-Instruct

Context length: SGLang reads from the model's config by default. Use --context-length to override. If you need to exceed the model's declared maximum, set SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1.

Tool calling: Use --tool-call-parser with the appropriate parser for your model family: qwen (Qwen 2.5), llama3, llama4, deepseekv3, mistral, glm. Without this flag, tool calls come back as plain text.

---

llama.cpp / llama-server — CPU & Metal Inference

llama.cpp ↗ runs quantized models on CPU, Apple Silicon (Metal), and consumer GPUs. Best for: running models without a datacenter GPU, Mac users, edge deployment.

Shell8 أسطر
# Build and start llama-server
cmake -B build && cmake --build build --config Release
./build/bin/llama-server \
  --jinja -fa \
  -c 64000 \
  -ngl 99 \
  -m models/qwen2.5-coder-32b-instruct-Q4_K_M.gguf \
  --port 8080 --host 0.0.0.0

Context length (-c): Recent builds default to 0 which reads the model's training context from the GGUF metadata. For models with 128k+ training context, this can OOM trying to allocate the full KV cache. Set -c explicitly to at least 64,000 tokens for Hermes. If using parallel slots (-np), the total context is divided among slots — with -c 64000 -np 4, each slot only gets 16k, which is below Hermes' minimum per active session.

Then configure Hermes to point at it:

Shell5 أسطر
hermes model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:8080/v1
# Skip API key (local servers don't need one)
# Enter model name — or leave blank to auto-detect if only one model is loaded

This saves the endpoint to config.yaml so it persists across sessions.

---

LM Studio — Desktop App with Local Models

LM Studio ↗ is a desktop app for running local models with a GUI. Best for: users who prefer a visual interface, quick model testing, developers on macOS/Windows/Linux.

Start the server from the LM Studio app (Developer tab → Start Server), or use the CLI:

Shellسطران
lms server start                        # Starts on port 1234
lms load qwen2.5-coder --context-length 64000

Then configure Hermes:

Shell5 أسطر
hermes model
# Select "LM Studio"
# Press Enter to use http://localhost:1234/v1
# Pick one of the discovered models
# If LM Studio server auth is enabled, enter LM_API_KEY when prompted

Hermes preserves the context of an already-loaded LM Studio instance. For an unloaded model in the default explicit mode, Hermes omits context_length unless you configured one in Hermes, so LM Studio can apply its own model setting. Hermes then uses only the context length LM Studio reports after loading.

To change context length in LM Studio:

  1. Click the gear icon next to the model picker
  2. Set "Context Length" to at least 64000 for a smooth experience
  3. Reload the model for the change to take effect
  4. If your machine cannot fit 64000, consider using a smaller model with larger context lengths.

Alternatively, use the CLI: lms load model-name --context-length 64000

You can use the CLI to estimate if the model will fit: lms load model-name --context-length 64000 --estimate-only

To set persistent per-model defaults: My Models tab → gear icon on the model → set context size. :::

If you use LM Studio's Just-In-Time loading / Auto-Evict feature and want LM Studio to manage model loading and eviction from normal chat requests, skip Hermes' explicit preload step:

Shellسطر واحد
hermes config set model.lmstudio_load_mode jit

Set it back to the default explicit preload behavior with:

Shellسطر واحد
hermes config set model.lmstudio_load_mode explicit

Tool calling: Supported since LM Studio 0.3.6. Models with native tool-calling training (Qwen 2.5, Llama 3.x, Mistral, Hermes) are auto-detected and shown with a tool badge. Other models use a generic fallback that may be less reliable.

---

WSL2 Networking (Windows Users)

Since Hermes Agent requires a Unix environment, Windows users run it inside WSL2. If your model server (Ollama, LM Studio, etc.) runs on the Windows host, you need to bridge the network gap — WSL2 uses a virtual network adapter with its own subnet, so localhost inside WSL2 refers to the Linux VM, not the Windows host.

Available on Windows 11 22H2+, mirrored mode makes localhost work bidirectionally between Windows and WSL2 — the simplest fix.

  1. Create or edit %USERPROFILE%\.wslconfig (e.g., C:\Users\YourName\.wslconfig):
INIسطران
   [wsl2]
   networkingMode=mirrored
  1. Restart WSL from PowerShell:
PowerShellسطر واحد
   wsl --shutdown
  1. Reopen your WSL2 terminal. localhost now reaches Windows services:
Shellسطر واحد
   curl http://localhost:11434/v1/models   # Ollama on Windows — works
Option 2: Use the Windows Host IP (Windows 10 / older builds)

If you can't use mirrored mode, find the Windows host IP from inside WSL2 and use that instead of localhost:

Shell3 أسطر
# Get the Windows host IP (the default gateway of WSL2's virtual network)
ip route show | grep -i default | awk '{ print $3 }'
# Example output: 172.29.192.1

Use that IP in your Hermes config:

YAML4 أسطر
model:
  default: qwen2.5-coder:32b
  provider: custom
  base_url: http://172.29.192.1:11434/v1   # Windows host IP, not localhost
Server Bind Address (Required for NAT Mode)

If you're using Option 2 (NAT mode with the host IP), the model server on Windows must accept connections from outside 127.0.0.1. By default, most servers only listen on localhost — WSL2 connections in NAT mode come from a different virtual subnet and will be refused. In mirrored mode, localhost maps directly so the default 127.0.0.1 binding works fine.

ServerDefault bindHow to fix
Ollama127.0.0.1Set OLLAMA_HOST=0.0.0.0 environment variable before starting Ollama (System Settings → Environment Variables on Windows, or edit the Ollama service)
LM Studio127.0.0.1Enable "Serve on Network" in the Developer tab → Server settings
llama-server127.0.0.1Add --host 0.0.0.0 to the startup command
vLLM0.0.0.0Already binds to all interfaces by default
SGLang127.0.0.1Add --host 0.0.0.0 to the startup command

Ollama on Windows (detailed): Ollama runs as a Windows service. To set OLLAMA_HOST:

  1. Open System Properties → Environment Variables
  2. Add a new System variable: OLLAMA_HOST = 0.0.0.0
  3. Restart the Ollama service (or reboot)
Windows Firewall

Windows Firewall treats WSL2 as a separate network (in both NAT and mirrored mode). If connections still fail after the steps above, add a firewall rule for your model server's port:

PowerShellسطران
# Run in Admin PowerShell — replace PORT with your server's port
New-NetFirewallRule -DisplayName "Allow WSL2 to Model Server" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 11434

Common ports: Ollama 11434, vLLM 8000, SGLang 30000, llama-server 8080, LM Studio 1234.

Quick Verification

From inside WSL2, test that you can reach your model server:

Shell3 أسطر
# Replace URL with your server's address and port
curl http://localhost:11434/v1/models          # Mirrored mode
curl http://172.29.192.1:11434/v1/models       # NAT mode (use your actual host IP)

If you get a JSON response listing your models, you're good. Use that same URL as the base_url in your Hermes config.

---

Troubleshooting Local Models

These issues affect all local inference servers when used with Hermes.

"Connection refused" from WSL2 to a Windows-hosted model server

If you're running Hermes inside WSL2 and your model server on the Windows host, http://localhost:<port> won't work in WSL2's default NAT networking mode. See WSL2 Networking ↗ above for the fix.

Tool calls appear as text instead of executing

The model outputs something like {"name": "web_search", "arguments": {...}} as a message instead of actually calling the tool.

Cause: Your server doesn't have tool calling enabled, or the model doesn't support it through the server's tool calling implementation.

ServerFix
llama.cppAdd --jinja to the startup command
vLLMAdd --enable-auto-tool-choice --tool-call-parser hermes
SGLangAdd --tool-call-parser qwen (or appropriate parser)
OllamaTool calling is enabled by default — make sure your model supports it (check with ollama show model-name)
LM StudioUpdate to 0.3.6+ and use a model with native tool support
Model seems to forget context or give incoherent responses

Cause: Context window is too small. When the conversation exceeds the context limit, most servers silently drop older messages. Hermes's system prompt + tool schemas alone can use 4k–8k tokens.

Diagnosis:

Shell7 أسطر
# Check what Hermes thinks the context is
# Look at startup line: "Context limit: X tokens"

# Check your server's actual context
# Ollama: ollama ps (CONTEXT column)
# llama.cpp: curl http://localhost:8080/props | jq '.default_generation_settings.n_ctx'
# vLLM: check --max-model-len in startup args

Fix: Set context to at least 64,000 tokens for agent use. See each server's section above for the specific flag.

"Context limit: 2048 tokens" at startup

Hermes auto-detects context length from your server's /v1/models endpoint. If the server reports a low value (or doesn't report one at all), Hermes uses the model's declared limit which may be wrong.

Fix: Set it explicitly in config.yaml:

YAML5 أسطر
model:
  default: your-model
  provider: custom
  base_url: http://localhost:11434/v1
  context_length: 64000
Responses get cut off mid-sentence

Possible causes:

  1. Low output cap (max_tokens) on the server — SGLang defaults to 128 tokens per response. Set --default-max-tokens on the server or configure Hermes with model.max_tokens in config.yaml. Note: max_tokens controls response length only — it is unrelated to how long your conversation history can be (that is context_length).
  2. Context exhaustion — The model filled its context window. Increase model.context_length or enable context compression in Hermes.

---

LiteLLM Proxy — Multi-Provider Gateway

LiteLLM ↗ is an OpenAI-compatible proxy that unifies 100+ LLM providers behind a single API. Best for: switching between providers without config changes, load balancing, fallback chains, budget controls.

Shell6 أسطر
# Install and start
pip install "litellm[proxy]"
litellm --model anthropic/claude-sonnet-4 --port 4000

# Or with a config file for multiple models:
litellm --config litellm_config.yaml --port 4000

Then configure Hermes with hermes model → Custom endpoint → http://localhost:4000/v1.

Example litellm_config.yaml with fallback:

YAML11 سطرًا
model_list:
  - model_name: "best"
    litellm_params:
      model: anthropic/claude-sonnet-4
      api_key: sk-ant-...
  - model_name: "best"
    litellm_params:
      model: openai/gpt-4o
      api_key: sk-...
router_settings:
  routing_strategy: "latency-based-routing"

---

ClawRouter — Cost-Optimized Routing

ClawRouter ↗ by BlockRunAI is a local routing proxy that auto-selects models based on query complexity. It classifies requests across 14 dimensions and routes to the cheapest model that can handle the task. Payment is via USDC cryptocurrency (no API keys).

Shellسطران
# Install and start
npx @blockrun/clawrouter    # Starts on port 8402

Then configure Hermes with hermes model → Custom endpoint → http://localhost:8402/v1 → model name blockrun/auto.

Routing profiles:

ProfileStrategySavings
blockrun/autoBalanced quality/cost74-100%
blockrun/ecoCheapest possible95-100%
blockrun/premiumBest quality models0%
blockrun/freeFree models only100%
blockrun/agenticOptimized for tool usevaries

---

Other Compatible Providers

Any service with an OpenAI-compatible API works. Some popular options:

ProviderBase URLNotes
Together AI ↗https://api.together.xyz/v1Cloud-hosted open models
Groq ↗https://api.groq.com/openai/v1Ultra-fast inference
DeepSeek ↗https://api.deepseek.com/v1DeepSeek models
Fireworks AI ↗https://api.fireworks.ai/inference/v1Fast open model hosting
GMI Cloud ↗https://api.gmi-serving.com/v1Managed OpenAI-compatible inference
Actual Computer ↗https://api.actual.inc/v1Private relay to your own cluster; local daemon at http://127.0.0.1:8080/v1
Cerebras ↗https://api.cerebras.ai/v1Wafer-scale chip inference
Mistral AI ↗https://api.mistral.ai/v1Mistral models
OpenAI ↗https://api.openai.com/v1Direct OpenAI access
Azure OpenAI ↗https://YOUR.openai.azure.com/Enterprise OpenAI
LocalAI ↗http://localhost:8080/v1Self-hosted, multi-model
Jan ↗http://localhost:1337/v1Desktop app with local models

Configure any of these with hermes model → Custom endpoint, or in config.yaml:

YAML5 أسطر
model:
  default: meta-llama/Llama-3.1-70B-Instruct-Turbo
  provider: custom
  base_url: https://api.together.xyz/v1
  api_key: your-together-key

---

Context Length Detection

Hermes uses a multi-source resolution chain to detect the correct context window for your model and provider:

  1. Config override — model.context_length in config.yaml (highest priority)
  2. Custom provider per-model — providers.<name>.models.<id>.context_length
  3. Persistent cache — previously discovered values (survives restarts)
  4. Endpoint /models — queries your server's API (local/custom endpoints)
  5. Anthropic /v1/models — queries Anthropic's API for max_input_tokens (API-key users only)
  6. OpenRouter API — live model metadata from OpenRouter
  7. Nous Portal — suffix-matches Nous model IDs against OpenRouter metadata
  8. models.dev ↗ — community-maintained registry with provider-specific context lengths for 3800+ models across 100+ providers
  9. Fallback defaults — broad model family patterns (128K default)

For most setups this works out of the box. The system is provider-aware — the same model can have different context limits depending on who serves it (e.g., claude-opus-4.6 is 1M on Anthropic direct but 128K on GitHub Copilot).

To set the context length explicitly, add context_length to your model config:

YAML4 أسطر
model:
  default: "qwen3.5:9b"
  base_url: "http://localhost:8080/v1"
  context_length: 131072  # tokens

For custom endpoints, you can also set context length per model:

YAML8 أسطر
providers:
  my-local-llm:
    api: "http://localhost:11434/v1"
    models:
      qwen3.5:27b:
        context_length: 64000
      deepseek-r1:70b:
        context_length: 65536

hermes model will prompt for context length when configuring a custom endpoint. Leave it blank for auto-detection.

---

Named Custom Providers

If you work with multiple custom endpoints (e.g., a local dev server and a remote GPU server), you can define them as named custom providers under the providers: dict in config.yaml, keyed by provider name:

YAML12 سطرًا
providers:
  local:
    api: http://localhost:8080/v1
    # api_key omitted — Hermes uses "no-key-required" for keyless local servers
  work:
    api: https://gpu-server.internal.corp/v1
    key_env: CORP_API_KEY
    transport: chat_completions   # set explicitly by `hermes model` → Custom Endpoint wizard; auto-detection still happens as a fallback
  anthropic-proxy:
    api: https://proxy.example.com/anthropic
    key_env: ANTHROPIC_PROXY_KEY
    transport: anthropic_messages  # for Anthropic-compatible proxies

Each entry accepts: api (the endpoint base URL — base_url/url are accepted aliases), name (optional display name; defaults to the dict key), key_env or inline api_key or key_cmd (see below), transport (chat_completions / anthropic_messages / codex_responses), default_model, models, context_length, discover_models, extra_body, extra_headers, ssl_ca_cert / ssl_verify, and enabled: false to hide an entry without deleting it.

Command-minted credentials (keycmd)

Enterprise gateways often issue short-lived bearer tokens (SSO/OIDC brokers, cloud IAM, internal auth proxies) rather than static API keys, so a token copied into .env goes stale mid-session and requests start returning 401. key_cmd names a command that prints a token; Hermes runs it and caches the result until shortly before expiry, so long sessions keep working with no restart:

YAML5 أسطر
providers:
  my-gateway:
    base_url: "https://gateway.internal.example.com/v1"
    api_mode: chat_completions
    key_cmd: "my-auth-cli print-token --profile prod"

Works with any helper that prints a token — databricks auth token, gcloud auth print-access-token, az account get-access-token, vault read, or Claude Code-style apiKeyHelper scripts.

The command must print only the token on stdout: either bare, or as JSON with an access_token field (expires_in is honored; absolute expiry/expiresOn ISO timestamps too). Multi-line output is rejected rather than guessed at. If no expiry is advertised, the token is re-minted on a bounded window.

Precedence: an explicit --api-key flag still wins; otherwise key_cmd beats a static api_key/key_env on the same entry. The minted credential applies to the main agent turn and to auxiliary tasks (title generation, compression, vision, embedding) alike.

Not to be confused with secrets.command, which runs a helper once at startup to populate env vars process-wide. Use that for a vault/keychain helper handing back many secrets; use key_cmd when one provider's credential must be re-minted during a session.

Some OpenAI-compatible endpoints need provider-specific request body fields. Add an extra_body map to the matching custom provider and Hermes will merge it into each chat-completions request for that endpoint:

YAML7 أسطر
providers:
  gemma-local:
    api: http://localhost:8080/v1
    default_model: google/gemma-4-31b-it
    extra_body:
      enable_thinking: true
      reasoning_effort: high

Use the shape your server documents. For example, vLLM Gemma deployments and some NVIDIA NIM endpoints expect enable_thinking under chat_template_kwargs instead of as a top-level extra_body field:

YAML3 أسطر
extra_body:
  chat_template_kwargs:
    enable_thinking: true

For Qwen reasoning models served by vLLM, this same shape can be used to disable thinking when a reasoning parser separates all generated text into reasoning fields and leaves the assistant content empty:

YAML3 أسطر
extra_body:
  chat_template_kwargs:
    enable_thinking: false

The hermes model → Custom Endpoint wizard now prompts for the API mode explicitly and persists your answer to config.yaml (as transport on the provider entry). URL-based auto-detection (e.g. /anthropic paths → anthropic_messages) still happens as a fallback when the field is left blank.

Native vision for custom-provider models. If your custom endpoint serves a vision-capable model that isn't in models.dev, set model.supports_vision: true so Hermes routes attached images natively (as image_url parts) instead of pre-processing them through vision_analyze. Single knob — no need to also set agent.image_input_mode: native.

YAML5 أسطر
model:
  provider: custom
  base_url: http://localhost:8080/v1
  default: qwen3.6-35b-a3b
  supports_vision: true   # send images natively; otherwise vision_analyze pre-describes them

The same key is honored on per-named-provider models (providers.<name>.models.<id>.supports_vision) and accepts standard YAML booleans (true/false/yes/no/on/off/1/0).

Switch between them mid-session with the triple syntax:

Text3 أسطر
/model custom:local:qwen-2.5       # Use the "local" endpoint with qwen-2.5
/model custom:work:llama3-70b      # Use the "work" endpoint with llama3-70b
/model custom:anthropic-proxy:claude-sonnet-4  # Use the proxy

You can also select named custom providers from the interactive hermes model menu.

---

Cookbook: Together AI, Groq, Perplexity

The cloud providers listed in Other Compatible Providers ↗ all speak OpenAI's REST dialect, so they wire up the same way under the providers: dict. Three worked recipes follow. Each drops into ~/.hermes/config.yaml and the matching API key goes in ~/.hermes/.env.

Together AI

Hosts open-weight models (Llama, MiniMax, Gemma, DeepSeek, Qwen) at prices significantly below first-party APIs. Good default for multi-model fleets.

YAML10 أسطر
# ~/.hermes/config.yaml
providers:
  together:
    api: https://api.together.xyz/v1
    key_env: TOGETHER_API_KEY
    # transport: chat_completions  # default — no need to set

model:
  default: MiniMaxAI/MiniMax-M2.7   # or any model from together.ai/models
  provider: custom:together
Shellسطران
# ~/.hermes/.env
TOGETHER_API_KEY=your-together-key

Switch models mid-session:

Text3 أسطر
/model custom:together:meta-llama/Llama-3.3-70B-Instruct-Turbo
/model custom:together:google/gemma-4-31b-it
/model custom:together:deepseek-ai/DeepSeek-V3

Together's /v1/models endpoint works, so hermes model can auto-discover available models.

Groq

Ultra-fast inference (~500 tok/s on Llama-3.3-70B). Small catalog but strong for latency-sensitive interactive use.

YAML9 أسطر
# ~/.hermes/config.yaml
providers:
  groq:
    api: https://api.groq.com/openai/v1
    key_env: GROQ_API_KEY

model:
  default: llama-3.3-70b-versatile
  provider: custom:groq
Shellسطران
# ~/.hermes/.env
GROQ_API_KEY=your-groq-key
Perplexity

Useful when you want a model that does live web search and citation automatically. Strict about which models are available — check perplexity.ai/settings/api ↗ for the current list.

YAML9 أسطر
# ~/.hermes/config.yaml
providers:
  perplexity:
    api: https://api.perplexity.ai
    key_env: PERPLEXITY_API_KEY

model:
  default: sonar
  provider: custom:perplexity
Shellسطران
# ~/.hermes/.env
PERPLEXITY_API_KEY=your-perplexity-key
Multiple providers in one config

The three recipes compose — use all of them together and switch per turn with /model custom:<name>:<model>:

YAML14 سطرًا
providers:
  together:
    api: https://api.together.xyz/v1
    key_env: TOGETHER_API_KEY
  groq:
    api: https://api.groq.com/openai/v1
    key_env: GROQ_API_KEY
  perplexity:
    api: https://api.perplexity.ai
    key_env: PERPLEXITY_API_KEY

model:
  default: MiniMaxAI/MiniMax-M2.7
  provider: custom:together      # boot to Together; switch freely after

---

Choosing the Right Setup

Use CaseRecommended
Just want it to workOpenRouter (default) or Nous Portal
Local models, easy setupOllama
Production GPU servingvLLM or SGLang
Mac / no GPUOllama or llama.cpp
Multi-provider routingLiteLLM Proxy or OpenRouter
Cost optimizationClawRouter or OpenRouter with sort: "price"
Maximum privacyOllama, vLLM, or llama.cpp (fully local)
Enterprise / AzureAzure OpenAI with custom endpoint
Chinese AI modelsz.ai (GLM), Kimi/Moonshot (kimi-coding or kimi-coding-cn), MiniMax, Xiaomi MiMo, or Tencent TokenHub (first-class providers)

Optional API Keys

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط. الأوامر هنا: hermes config set.

FeatureProviderEnv Variable
Web scrapingFirecrawl ↗FIRECRAWL_API_KEY, FIRECRAWL_API_URL
Browser automationBrowserbase ↗BROWSERBASE_API_KEY, BROWSERBASE_PROJECT_ID
Image generationFAL ↗FAL_KEY
Premium TTS voicesElevenLabs ↗ELEVENLABS_API_KEY
OpenAI TTS + voice transcriptionOpenAI ↗VOICE_TOOLS_OPENAI_KEY
Mistral TTS + voice transcriptionMistral ↗MISTRAL_API_KEY
Cross-session user modelingHoncho ↗HONCHO_API_KEY
Semantic long-term memorySupermemory ↗SUPERMEMORY_API_KEY

Self-Hosting Firecrawl

By default, Hermes uses the Firecrawl cloud API ↗ for web search and scraping. If you prefer to run Firecrawl locally, you can point Hermes at a self-hosted instance instead. See Firecrawl's SELF_HOST.md ↗ for complete setup instructions.

What you get: No API key required, no rate limits, no per-page costs, full data sovereignty.

What you lose: The cloud version uses Firecrawl's proprietary "Fire-engine" for advanced anti-bot bypassing (Cloudflare, CAPTCHAs, IP rotation). Self-hosted uses basic fetch + Playwright, so some protected sites may fail. Search uses DuckDuckGo instead of Google.

Setup:

  1. Clone and start the Firecrawl Docker stack (5 containers: API, Playwright, Redis, RabbitMQ, PostgreSQL — requires ~4-8 GB RAM):
Shell4 أسطر
   git clone https://github.com/firecrawl/firecrawl
   cd firecrawl
   # In .env, set: USE_DB_AUTHENTICATION=false, HOST=0.0.0.0, PORT=3002
   docker compose up -d
  1. Point Hermes at your instance (no API key needed):
Shellسطر واحد
   hermes config set FIRECRAWL_API_URL http://localhost:3002

You can also set both FIRECRAWL_API_KEY and FIRECRAWL_API_URL if your self-hosted instance has authentication enabled.

OpenRouter Provider Routing

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير.

When using OpenRouter, you can control how requests are routed across providers. Add a provider_routing section to ~/.hermes/config.yaml:

YAML7 أسطر
provider_routing:
  sort: "throughput"          # "price" (default), "throughput", or "latency"
  # only: ["anthropic"]      # Only use these providers
  # ignore: ["deepinfra"]    # Skip these providers
  # order: ["anthropic", "google"]  # Try providers in this order
  # require_parameters: true  # Only use providers that support all request params
  # data_collection: "deny"   # Exclude providers that may store/train on data

Shortcuts: Append :nitro to any model name for throughput sorting (e.g., anthropic/claude-sonnet-4:nitro), or :floor for price sorting.

OpenRouter Pareto Code Router

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير.

OpenRouter ships an experimental coding-model router at openrouter/pareto-code that auto-routes requests to the cheapest model meeting a coding-quality bar (ranked by Artificial Analysis ↗). Pick this model and tune the min_coding_score knob in ~/.hermes/config.yaml:

YAML6 أسطر
model:
  provider: openrouter
  model: openrouter/pareto-code

openrouter:
  min_coding_score: 0.65   # 0.0–1.0; higher = stronger (more expensive) coders. Default 0.65.

Notes:

  • min_coding_score is only sent when model.model is openrouter/pareto-code. On any other model the value is a no-op.
  • Set to empty string (or remove the line) to let OpenRouter pick the strongest available coder — its documented behavior when the plugins block is omitted.
  • Selection is deterministic per score on a given day, but the actual model chosen can shift as the Pareto frontier moves (new models, benchmark updates).
  • See OpenRouter's Pareto Router docs ↗ for the full router behavior.
  • To use the Pareto Code router for a specific auxiliary task (compression, vision, etc.) instead of the main agent, set extra_body.plugins under that task — see Auxiliary Models → OpenRouter routing & Pareto Code for auxiliary tasks.

Fallback Providers

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير. الأوامر هنا: hermes fallback.

Configure a chain of backup providers Hermes tries in order when the primary model fails (rate limits, server errors, auth failures). The canonical format is a top-level fallback_providers: list:

YAML7 أسطر
fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4
  - provider: anthropic
    model: claude-sonnet-4
    # base_url: http://localhost:8000/v1    # optional, for custom endpoints
    # api_mode: chat_completions           # optional override

The legacy single-pair fallback_model: dict is still accepted for back-compat:

YAML3 أسطر
fallback_model:
  provider: openrouter
  model: anthropic/claude-sonnet-4

When activated, the fallback swaps the model and provider mid-session without losing your conversation. The chain is tried entry-by-entry; activation is one-shot per session.

Supported providers: openrouter, nous, novita, openai-codex, copilot, copilot-acp, anthropic, gemini, qwen-oauth, huggingface, zai, kimi-coding, kimi-coding-cn, minimax, minimax-cn, minimax-oauth, deepseek, nvidia, xai, xai-oauth, ollama-cloud, bedrock, ai-gateway, azure-foundry, opencode-zen, opencode-go, commandcode, commandcode-anthropic, kilocode, xiaomi, arcee, gmi, actual, stepfun, lmstudio, alibaba, alibaba-coding-plan, tencent-tokenhub, custom.

---

See Also

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes.

  • Configuration — General configuration (directory structure, config precedence, terminal backends, memory, compression, and more)
  • Environment Variables — Complete reference of all environment variables
اختبار الفهم

5 أسئلة إجاباتها كلها في هذه الصفحة.

كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.

1. بحسب هذا الدرس، أي أمر يقوم بـ«existing install — pick "Nous Portal" from the list»؟
2. بحسب هذا الدرس، أي أمر يقوم بـ«inspect login + routing at any time»؟
3. في جدول هذا الدرس، ما «Setup» المقابل لـ«Actual Computer»؟
4. أي متغير بيئة من التالي يظهر فعليًا في هذا الدرس؟
5. ما التحذير الذي يذكره المصدر في هذا الدرس؟