الأكاديمية ← ميزات Hermesتوثيق رسمي · إرشاد عربي

مزوّدون احتياطيون عند الفشل

Fallback Providers

متوسط15 دقيقة قراءةالدرس 225 أسئلة✓ 2026-08-18
قبل أن تقرأ

ما هذه الصفحة، وماذا تحتوي.

المزوّد والنموذج: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes. Hermes نفسه لا يفكّر؛ هو ينظّم العمل ويستدعي النموذج. لذلك اختيار النموذج يحدّد جودة النتيجة وتكلفتها. الصفحة فيها تحذير من المصدر، و15 دقيقة قراءة. انتبه: الأغلى ليس دائمًا الأفضل لمهمتك. جرّب مهمة واحدة على نموذجين وقارن، وضع سقفًا للإنفاق من البداية.

7أقسام
18أمثلة برمجية
5جداول
2أوامر
2,497كلمة من المصدر
الوصف الرسمي في سطر

Configure automatic failover to backup LLM providers when your primary model is unavailable.

ماذا ستستطيع بعدها

نتائج مأخوذة من هذه الصفحة، لا من قالب.

  • تعرف ما المزوّد والنموذج ولماذا قد تحتاجه.
  • تنفّذ hermes fallback وhermes model وتفهم ما يحدث بعدها.
  • تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
  • تضبط LOCAL_API_KEY في المكان الصحيح.
ما ستقابله من أسماء

كما تظهر تمامًا داخل Hermes.

الأوامر
  • hermes fallback
  • hermes model
متغيرات البيئة
  • LOCAL_API_KEY
  • OPENAI_API_KEY
  • OPENROUTER_API_KEY
  • RESOURCE_EXHAUSTED
خريطة الصفحة

انتقل مباشرة إلى ما تحتاجه.

  1. 01Primary Model Fallback
  2. 02Auxiliary Task Fallback
  3. 03Auxiliary Capacity-Error Fallback
  4. 04Context Compression Fallback
  5. 05Delegation Provider Override
  6. 06Cron Job Providers
  7. 07Summary
الصفحة الرسمية كاملة

بلا اختصار أو حذف.

النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.

Hermes Agent has three layers of resilience that keep your sessions running when providers hit issues:

  1. Credential pools — rotate across multiple API keys for the same provider (tried first)
  2. Primary model fallback — automatically switches to a different provider:model when your main model fails
  3. Auxiliary task fallback — independent provider resolution for side tasks like vision, compression, and web extraction

Credential pools handle same-provider rotation (e.g., multiple OpenRouter keys). This page covers cross-provider fallback. Both are optional and work independently.

Primary Model Fallback

فيه تحذير مهم. اقرأه قبل أن تنفّذ أي شيء من هذا القسم. الأوامر هنا: hermes fallback، hermes model. نصّ التحذير من المصدر مذكور أسفل هذا الشرح.

When your main LLM provider encounters errors — rate limits, server overload, auth failures, connection drops — Hermes can automatically switch to a backup provider:model pair mid-session without losing your conversation.

Configuration

The easiest path is the interactive manager:

Shellسطر واحد
hermes fallback

hermes fallback reuses the provider picker from hermes model — same provider list, same credential prompts, same validation. Use the subcommands add, list (alias ls), remove (alias rm), and clear to manage the chain. Changes persist under the top-level fallback_providers: list in config.yaml.

If you'd rather edit the YAML directly, add a top-level fallback_providers list to ~/.hermes/config.yaml:

YAML3 أسطر
fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4

Each entry requires both provider and model. Entries missing either field are ignored.

Supported Providers

ProviderValueRequirements
AI Gatewayai-gatewayAI_GATEWAY_API_KEY
OpenRouteropenrouterOPENROUTER_API_KEY
Nous Portalnoushermes setup --portal (fresh) or hermes auth add nous (OAuth)
OpenAI Codexopenai-codexhermes model → ChatGPT or Codex Subscription (ChatGPT OAuth)
GitHub CopilotcopilotCOPILOT_GITHUB_TOKEN, GH_TOKEN, or GITHUB_TOKEN
GitHub Copilot ACPcopilot-acpExternal process (editor integration)
AnthropicanthropicANTHROPIC_API_KEY or Claude Code credentials
z.ai / GLMzaiGLM_API_KEY
Kimi / Moonshotkimi-codingKIMI_API_KEY
MiniMaxminimaxMINIMAX_API_KEY
MiniMax (China)minimax-cnMINIMAX_CN_API_KEY
DeepSeekdeepseekDEEPSEEK_API_KEY
NVIDIA NIMnvidiaNVIDIA_API_KEY (optional: NVIDIA_BASE_URL)
GMI CloudgmiGMI_API_KEY (optional: GMI_BASE_URL)
Upstage Solarupstage (alias solar)UPSTAGE_API_KEY (optional: UPSTAGE_BASE_URL)
StepFunstepfunSTEPFUN_API_KEY (optional: STEPFUN_BASE_URL)
Ollama Cloudollama-cloudOLLAMA_API_KEY
Google AI StudiogeminiGOOGLE_API_KEY (alias: GEMINI_API_KEY)
xAI (Grok)xai (alias grok)XAI_API_KEY (optional: XAI_BASE_URL)
xAI Grok OAuth (SuperGrok)xai-oauth (alias grok-oauth)hermes model → xAI Grok OAuth (browser login; SuperGrok subscription)
AWS BedrockbedrockStandard boto3 auth (AWS_REGION + AWS_PROFILE or AWS_ACCESS_KEY_ID)
Qwen Portal (OAuth)qwen-oauthhermes model (Qwen Portal OAuth; optional: HERMES_QWEN_BASE_URL)
MiniMax (OAuth)minimax-oauthhermes model (MiniMax portal OAuth)
OpenCode Zenopencode-zenOPENCODE_ZEN_API_KEY
CommandCodecommandcode (alias commandcode-chat; Claude via commandcode-anthropic)COMMANDCODE_API_KEY
OpenCode Goopencode-goOPENCODE_GO_API_KEY
Kilo CodekilocodeKILOCODE_API_KEY
Xiaomi MiMoxiaomiXIAOMI_API_KEY
Arcee AIarceeARCEEAI_API_KEY
GMI CloudgmiGMI_API_KEY
Alibaba / DashScopealibabaDASHSCOPE_API_KEY
Alibaba Coding Planalibaba-coding-planALIBABA_CODING_PLAN_API_KEY (falls back to DASHSCOPE_API_KEY)
Kimi / Moonshot (China)kimi-coding-cnKIMI_CN_API_KEY
StepFunstepfunSTEPFUN_API_KEY
Tencent TokenHubtencent-tokenhubTOKENHUB_API_KEY
Microsoft Foundryazure-foundryAZURE_FOUNDRY_API_KEY + AZURE_FOUNDRY_BASE_URL
LM Studio (local)lmstudioLM_API_KEY (or none for local) + LM_BASE_URL
Hugging FacehuggingfaceHF_TOKEN
Custom endpointcustombase_url + key_env (see below)

Custom Endpoint Fallback

For a custom OpenAI-compatible endpoint, add base_url and optionally key_env:

YAML5 أسطر
fallback_providers:
  - provider: custom
    model: my-local-model
    base_url: http://localhost:8000/v1
    key_env: MY_LOCAL_KEY            # env var name containing the API key

When Fallback Triggers

The fallback activates automatically when the primary model fails with:

  • Rate limits (HTTP 429) — after exhausting retry attempts
  • Server errors (HTTP 500, 502, 503) — after exhausting retry attempts
  • Auth failures (HTTP 401, 403) — immediately (no point retrying)
  • Not found (HTTP 404) — immediately
  • Invalid responses — when the API returns malformed or empty responses repeatedly

When triggered, Hermes:

  1. Resolves credentials for the fallback provider
  2. Builds a new API client
  3. Swaps the model, provider, and client in-place
  4. Resets the retry counter and continues the conversation

The switch is seamless — your conversation history, tool calls, and context are preserved. The agent continues from exactly where it left off, just using a different model.

Examples

OpenRouter as fallback for Anthropic native:

YAML7 أسطر
model:
  provider: anthropic
  default: claude-sonnet-4-6

fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4

Nous Portal as fallback for OpenRouter:

YAML7 أسطر
model:
  provider: openrouter
  default: anthropic/claude-opus-4

fallback_providers:
  - provider: nous
    model: nous-hermes-3

Local model as fallback for cloud:

YAML5 أسطر
fallback_providers:
  - provider: custom
    model: llama-3.1-70b
    base_url: http://localhost:8000/v1
    key_env: LOCAL_API_KEY

Codex OAuth as fallback:

YAML3 أسطر
fallback_providers:
  - provider: openai-codex
    model: gpt-5.3-codex

Where Fallback Works

ContextFallback Supported
CLI sessions✔
Messaging gateway (Telegram, Discord, etc.)✔
Subagent delegation✔ (subagents inherit the parent fallback chain)
Cron jobs✔ (cron agents inherit configured fallback providers)
Auxiliary tasks on provider: auto✔ (try per-task fallback, then the main fallback chain before built-in aux discovery)

---

Auxiliary Task Fallback

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

Hermes uses separate lightweight models for side tasks. Each task has its own provider resolution chain that acts as a built-in fallback system.

Tasks with Independent Provider Resolution

TaskWhat It DoesConfig Key
VisionImage analysis, browser screenshotsauxiliary.vision
Web ExtractWeb page summarizationauxiliary.web_extract
CompressionContext compression summariesauxiliary.compression
Skills HubSkill search and discoveryauxiliary.skills_hub
MCPMCP helper operationsauxiliary.mcp
ApprovalSmart command-approval classificationauxiliary.approval
Title GenerationSession title summariesauxiliary.title_generation
Triage Specifierhermes kanban specify / dashboard ✨ button — fleshes out a one-liner triage task into a real specauxiliary.triage_specifier

Auto-Detection Chain

When a task's provider is set to "auto" (the default), Hermes first tries the main provider + main model for that auxiliary task. If that route is unavailable or later fails with a capacity-style error, Hermes now honors user-configured fallback policy before using the built-in discovery chain:

Textسطران
Main provider + main model → auxiliary.<task>.fallback_chain →
fallback_providers / fallback_model → built-in auxiliary discovery chain

The task-specific chain is most precise and wins when present. The top-level fallback_providers chain is the same policy the main agent uses, so free-only or same-provider fallback rules apply to auxiliary tasks on auto as well.

Built-in text discovery chain (compression, web extract, title generation, etc.):

Textسطران
OpenRouter → Nous Portal → Custom endpoint → Codex OAuth →
API-key providers (z.ai, Kimi, MiniMax, Xiaomi MiMo, Hugging Face, Anthropic) → give up

Built-in vision discovery chain:

Textسطران
Main provider (if vision-capable) → OpenRouter → Nous Portal →
Codex OAuth → Anthropic → Custom endpoint → give up

Those built-in chains are a convenience fallback for users who have not declared a task-specific or main fallback policy.

Configuring Auxiliary Providers

Each task can be configured independently in config.yaml:

YAML25 سطرًا
auxiliary:
  vision:
    provider: "auto"              # auto | openrouter | nous | codex | main | anthropic
    model: ""                     # e.g. "openai/gpt-4o"
    base_url: ""                  # direct endpoint (takes precedence over provider)
    api_key: ""                   # API key for base_url

  web_extract:
    provider: "auto"
    model: ""

  compression:
    provider: "auto"
    model: ""
    fallback_chain:              # optional, task-specific fallback policy
      - provider: openrouter
        model: inclusionai/ring-2.6-1t:free

  skills_hub:
    provider: "auto"
    model: ""

  mcp:
    provider: "auto"
    model: ""

Every task above follows the same provider / model / base_url pattern. Each task can also declare its own fallback_chain; if omitted, provider: auto uses the top-level fallback_providers chain before Hermes' built-in auxiliary discovery chain.

Context compression is configured under auxiliary.compression:

YAML5 أسطر
auxiliary:
  compression:
    provider: main                                    # Same provider options as other auxiliary tasks
    model: google/gemini-3-flash-preview
    base_url: null                                    # Custom OpenAI-compatible endpoint

And the primary fallback chain uses:

YAML4 أسطر
fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4
    # base_url: http://localhost:8000/v1             # Optional custom endpoint

All three — auxiliary, compression, fallback — work the same way: set provider to pick who handles the request, model to pick which model, and base_url to point at a custom endpoint (overrides provider).

Provider Options for Auxiliary Tasks

These options apply to auxiliary:, compression:, and fallback_providers: entries only — "main" is not a valid value for your top-level model.provider. For custom endpoints, use provider: custom in your model: section (see AI Providers).

ProviderDescriptionRequirements
"auto"Try providers in order until one works (default)At least one provider configured
"openrouter"Force OpenRouterOPENROUTER_API_KEY
"nous"Force Nous Portalhermes auth
"codex"Force Codex OAuthhermes model → ChatGPT or Codex Subscription
"main"Use whatever provider the main agent uses (auxiliary tasks only)Active main provider configured
"anthropic"Force Anthropic nativeANTHROPIC_API_KEY or Claude Code credentials

Direct Endpoint Override

For any auxiliary task, setting base_url bypasses provider resolution entirely and sends requests directly to that endpoint:

YAML5 أسطر
auxiliary:
  vision:
    base_url: "http://localhost:1234/v1"
    api_key: "local-key"
    model: "qwen2.5-vl"

base_url takes precedence over provider. Hermes uses the configured api_key for authentication, falling back to OPENAI_API_KEY if not set. It does not reuse OPENROUTER_API_KEY for custom endpoints.

---

Auxiliary Capacity-Error Fallback

قسم لحل المشكلات. ابحث فيه عن العطل الذي يشبه حالتك بدل قراءته كاملًا.

When you set an explicit auxiliary provider (e.g. auxiliary.vision.provider: glm), Hermes treats that as your preferred choice — but if the provider literally cannot serve the request because of a capacity error (HTTP 402 payment required, HTTP 429 daily-quota exhaustion, connection failure), Hermes falls back through a layered chain instead of failing silently:

  1. Primary aux provider — the one you configured (tried first, always)
  2. auxiliary.<task>.fallback_chain — your per-task override list, if you wrote one
  3. Main agent provider + model — last-resort safety net (always tried, even if you didn't write a chain)
  4. Warn + re-raise — if every layer fails, Hermes logs Auxiliary <task>: ... all fallbacks exhausted at WARNING level and re-raises the original error

Transient HTTP 429 rate limits (Retry-After: ...) are treated as request constraints, not capacity problems — they respect your explicit provider choice and do not trigger the fallback ladder. Only daily/monthly quota exhaustion, payment errors, and connection failures bypass the explicit-provider gate.

For users on provider: auto (no explicit aux provider), the existing auto-detection chain runs in place of steps 2–3. Its first step is already the main agent model, so auto users get the same outcome with zero config.

Optional: per-task fallback chain

If you want a different fallback ordering than "main agent model first", configure fallback_chain explicitly. Each entry needs at least provider; model, base_url, and api_key are optional.

YAML16 سطرًا
auxiliary:
  vision:
    provider: glm
    model: glm-4v-flash
    fallback_chain:
      - provider: openrouter
        model: google/gemini-3-flash-preview
      - provider: nous
        model: anthropic/claude-sonnet-4

  compression:
    provider: openrouter
    fallback_chain:
      - provider: openai
        model: gpt-4o-mini
        timeout: 240            # optional — this candidate's own deadline (seconds)

You do not need to configure fallback_chain to get fallback — the main-agent safety net runs regardless. Use it only when you specifically want a different order than the default.

Each fallback_chain entry may also declare its own timeout (seconds). Without it, a fallback candidate inherits the task-level timeout — which may be tuned for the primary provider. Declaring a per-entry timeout lets a slower-but-reliable fallback (e.g. a large-context summarizer) get the budget it actually needs instead of dying on the primary's clock.

Provider quota errors that trigger fallback

Hermes recognizes these as capacity-equivalent to 402 credit exhaustion (not transient rate limits):

  • Bedrock / LiteLLM: Too many tokens per day, daily limit, tokens per day
  • Vertex AI / GCP: quota exceeded, resource exhausted, RESOURCE_EXHAUSTED
  • Generic: daily quota, quota_exceeded

If your provider returns a different phrase for daily-quota exhaustion and Hermes doesn't trigger fallback, that's a bug — open an issue with the exact error string.

---

Context Compression Fallback

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير.

Context compression uses the auxiliary.compression config block to control which model and provider handles summarization:

YAML4 أسطر
auxiliary:
  compression:
    provider: "auto"                              # auto | openrouter | nous | main
    model: "google/gemini-3-flash-preview"

If no provider is available for compression, Hermes drops middle conversation turns without generating a summary rather than failing the session.

---

Delegation Provider Override

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير.

Subagents spawned by delegate_task inherit the parent agent's primary fallback chain. You can still route subagents to a different primary provider:model pair for cost optimization:

YAML5 أسطر
delegation:
  provider: "openrouter"                      # override provider for all subagents
  model: "google/gemini-3-flash-preview"      # override model
  # base_url: "http://localhost:1234/v1"      # or use a direct endpoint
  # api_key: "local-key"

See Subagent Delegation for full configuration details.

---

Cron Job Providers

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: المزوّد هو الشركة التي تشغّل نموذج الذكاء الاصطناعي، والنموذج هو «العقل» الذي يفكّر لـHermes.

Cron jobs inherit your configured fallback_providers chain (or legacy fallback_model) when they create an agent. To use a different primary provider for a cron job, configure provider and model overrides on the cron job itself:

Python7 أسطر
cronjob(
    action="create",
    schedule="every 2h",
    prompt="Check server status",
    provider="openrouter",
    model="google/gemini-3-flash-preview"
)

See Scheduled Tasks (Cron) for full configuration details.

---

Summary

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

FeatureFallback MechanismConfig Location
Main agent modelfallback_providers in config.yaml — per-turn failover on errors (primary restored each turn)fallback_providers: (top-level list)
Auxiliary tasks (any) — auto usersFull auto-detection chain (main agent model first, then provider chain) on capacity errorsauxiliary.<task>.provider: auto
Auxiliary tasks (any) — explicit providerfallback_chain (if set) → main agent model → warn + raise, on capacity errors onlyauxiliary.<task>.fallback_chain
VisionLayered (see above) + internal OpenRouter retryauxiliary.vision
Web extractionLayered (see above) + internal OpenRouter retryauxiliary.web_extract
Context compressionLayered (see above); degrades to no-summary if all layers unavailableauxiliary.compression
Skills hubLayered (see above)auxiliary.skills_hub
MCP helpersLayered (see above)auxiliary.mcp
Approval classificationLayered (see above)auxiliary.approval
Title generationLayered (see above)auxiliary.title_generation
Triage specifierLayered (see above)auxiliary.triage_specifier
DelegationInherits the parent's fallback_providers chain; optional provider/model overridedelegation.provider / delegation.model
Cron jobsInherit the configured fallback_providers chain; optional per-job provider overridePer-job provider / model
اختبار الفهم

5 أسئلة إجاباتها كلها في هذه الصفحة.

كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.

1. في جدول هذا الدرس، ما «Value» المقابل لـ«Kimi / Moonshot»؟
2. أي متغير بيئة من التالي يظهر فعليًا في هذا الدرس؟
3. ما التحذير الذي يذكره المصدر في هذا الدرس؟
4. أي عنوان من التالي لا يظهر في هذا الدرس؟
5. أي مفتاح إعداد يظهر في أمثلة هذا الدرس؟