Academy → Developer GuideOfficial documentation · Arabic guidance

Agent Loop Internals

حلقة الوكيل من الداخل

Advanced7 min readLesson 42 questions✓ 2026-08-18
Before you read

What this page is, and what it holds.

This page covers Agent Loop Internals. About 7 minutes to read. A job that runs while you sleep also fails while you sleep. Make it report, and test it by hand first.

11sections
5code examples
4tables
0commands
1,185source words
The official one-line description

Detailed walkthrough of AIAgent execution, API modes, tools, callbacks, and fallback behavior

What you will be able to do

Outcomes taken from this page, not a template.

  • Understand what المهام المجدولة is and when you need it.
  • Read the table and take only the row that applies to you.
  • Know the common mistake before you hit it.
Page map

Jump to the part you need.

  1. 01Core Responsibilities
  2. 02Two Entry Points
  3. 03API Modes
  4. 04Turn Lifecycle
  5. 05Interruptible API Calls
  6. 06Tool Execution
  7. 07Callback Surfaces
  8. 08Budget and Fallback Behavior
  9. 09Compression and Persistence
  10. 10Key Source Files
  11. 11Related Docs
The full official page

Nothing summarised away.

The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.

The core orchestration engine is run_agent.py's AIAgent class — a large file that handles everything from prompt assembly to tool dispatch to provider failover.

Core Responsibilities

Explains the idea itself. Read it slowly; the later sections build on it.

AIAgent is responsible for:

  • Assembling the effective system prompt and tool schemas via prompt_builder.py
  • Selecting the correct provider/API mode (chat_completions, codex_responses, anthropic_messages)
  • Making interruptible model calls with cancellation support
  • Executing tool calls (sequentially or concurrently via thread pool)
  • Maintaining conversation history in OpenAI message format
  • Handling compression, retries, and fallback model switching
  • Tracking iteration budgets across parent and child agents
  • Flushing persistent memory before context is lost

Two Entry Points

Explains the idea itself. Read it slowly; the later sections build on it.

Python10 lines
# Simple interface — returns final response string
response = agent.chat("Fix the bug in main.py")

# Full interface — returns dict with messages, metadata, usage stats
result = agent.run_conversation(
    user_message="Fix the bug in main.py",
    system_message=None,           # auto-built if omitted
    conversation_history=None,      # auto-loaded from session if omitted
    task_id="task_abc123"
)

chat() is a thin wrapper around run_conversation() that extracts the final_response field from the result dict.

API Modes

Explains the idea itself. Read it slowly; the later sections build on it.

Hermes supports three API execution modes, resolved from provider selection, explicit args, and base URL heuristics:

API modeUsed forClient type
chat_completionsOpenAI-compatible endpoints (OpenRouter, custom, most providers)openai.OpenAI
codex_responsesOpenAI Codex / Responses APIopenai.OpenAI with Responses format
anthropic_messagesNative Anthropic Messages APIanthropic.Anthropic via adapter

The mode determines how messages are formatted, how tool calls are structured, how responses are parsed, and how caching/streaming works. All three converge on the same internal message format (OpenAI-style role/content/tool_calls dicts) before and after API calls.

Mode resolution order:

  1. Explicit api_mode constructor arg (highest priority)
  2. Provider-specific detection (e.g., anthropic provider → anthropic_messages)
  3. Base URL heuristics (e.g., api.anthropic.com → anthropic_messages)
  4. Default: chat_completions

Turn Lifecycle

Explains the idea itself. Read it slowly; the later sections build on it.

Each iteration of the agent loop follows this sequence:

Text15 lines
run_conversation()
  1. Generate task_id if not provided
  2. Append user message to conversation history
  3. Build or reuse cached system prompt (prompt_builder.py)
  4. Check if preflight compression is needed (>50% context)
  5. Build API messages from conversation history
     - chat_completions: OpenAI format as-is
     - codex_responses: convert to Responses API input items
     - anthropic_messages: convert via anthropic_adapter.py
  6. Inject ephemeral prompt layers (budget warnings, context pressure)
  7. Apply prompt caching markers if on Anthropic
  8. Make interruptible API call (_interruptible_api_call)
  9. Parse response:
     - If tool_calls: execute them, append results, loop back to step 5
     - If text response: persist session, flush memory if needed, return

Message Format

All messages use OpenAI-compatible format internally:

Python4 lines
{"role": "system", "content": "..."}
{"role": "user", "content": "..."}
{"role": "assistant", "content": "...", "tool_calls": [...]}
{"role": "tool", "tool_call_id": "...", "content": "..."}

Reasoning content (from models that support extended thinking) is stored in assistant_msg["reasoning"] and optionally displayed via the reasoning_callback.

Message Alternation Rules

The agent loop enforces strict message role alternation:

  • After the system message: User → Assistant → User → Assistant → ...
  • During tool calling: Assistant (with tool_calls) → Tool → Tool → ... → Assistant
  • Never two assistant messages in a row
  • Never two user messages in a row
  • Only tool role can have consecutive entries (parallel tool results)

Providers validate these sequences and will reject malformed histories.

Interruptible API Calls

Explains the idea itself. Read it slowly; the later sections build on it.

API requests are wrapped in _interruptible_api_call() which runs the actual HTTP call in a background thread while monitoring an interrupt event:

Text8 lines
┌────────────────────────────────────────────────────┐
│  Main thread                  API thread           │
│                                                    │
│   wait on:                     HTTP POST           │
│    - response ready     ───▶   to provider         │
│    - interrupt event                               │
│    - timeout                                       │
└────────────────────────────────────────────────────┘

When interrupted (user sends new message, /stop command, or signal):

  • The API thread is abandoned (response discarded)
  • The agent can process the new input or shut down cleanly
  • No partial response is injected into conversation history

Tool Execution

A lookup table. Do not read it all; find the row that applies to you.

Sequential vs Concurrent

When the model returns tool calls:

  • Single tool call → executed directly in the main thread
  • Multiple tool calls → executed concurrently via ThreadPoolExecutor
  • Exception: tools marked as interactive (e.g., clarify) force sequential execution
  • Results are reinserted in the original tool call order regardless of completion order

Execution Flow

Text8 lines
for each tool_call in response.tool_calls:
    1. Resolve handler from tools/registry.py
    2. Fire pre_tool_call plugin hook
    3. Check if dangerous command (tools/approval.py)
       - If dangerous: invoke approval_callback, wait for user
    4. Execute handler with args + task_id
    5. Fire post_tool_call plugin hook
    6. Append {"role": "tool", "content": result} to history

Agent-Level Tools

Some tools are intercepted by run_agent.py before reaching handle_function_call():

ToolWhy intercepted
todoReads/writes agent-local task state
memoryWrites to persistent memory files with character limits
session_searchQueries session history via the agent's session DB
delegate_taskSpawns subagent(s) with isolated context

These tools modify agent state directly and return synthetic tool results without going through the registry.

Callback Surfaces

A lookup table. Do not read it all; find the row that applies to you.

AIAgent supports platform-specific callbacks that enable real-time progress in the CLI, gateway, and ACP integrations:

CallbackWhen firedUsed by
tool_progress_callbackBefore/after each tool executionCLI spinner, gateway progress messages
thinking_callbackWhen model starts/stops thinkingCLI "thinking..." indicator
reasoning_callbackWhen model returns reasoning contentCLI reasoning display, gateway reasoning blocks
clarify_callbackWhen clarify tool is calledCLI input prompt, gateway interactive message
step_callbackAfter each complete agent turnGateway step tracking, ACP progress
stream_delta_callbackEach streaming token (when enabled)CLI streaming display
tool_gen_callbackWhen tool call is parsed from streamCLI tool preview in spinner
status_callbackState changes (thinking, executing, etc.)ACP status updates

Budget and Fallback Behavior

Explains the idea itself. Read it slowly; the later sections build on it.

Iteration Budget

The agent tracks iterations via IterationBudget:

  • Default: 500 iterations (configurable via agent.max_turns)
  • Each agent gets its own budget. Subagents get independent budgets capped at delegation.max_iterations (default 50) — total iterations across parent + subagents can exceed the parent's cap
  • At 100%, the agent stops and returns a summary of work done

Fallback Model

When the primary model fails (429 rate limit, 5xx server error, 401/403 auth error):

  1. Check fallback_providers list in config
  2. Try each fallback in order
  3. On success, continue the conversation with the new provider
  4. On 401/403, attempt credential refresh before failing over

The fallback system also covers auxiliary tasks independently — vision, compression, and web extraction each have their own fallback chain configurable via the auxiliary.* config section.

Compression and Persistence

Explains the idea itself. Read it slowly; the later sections build on it.

When Compression Triggers

  • Preflight (before API call): If conversation exceeds 50% of model's context window
  • Gateway auto-compression: If conversation exceeds 85% (more aggressive, runs between turns)

What Happens During Compression

  1. Memory is flushed to disk first (preventing data loss)
  2. Middle conversation turns are summarized into a compact summary
  3. The last N messages are preserved intact (compression.protect_last_n, default: 20)
  4. Tool call/result message pairs are kept together (never split)
  5. A new session lineage ID is generated (compression creates a "child" session)

Session Persistence

After each turn:

  • Messages are saved to the session store (SQLite via hermes_state.py)
  • Memory changes are flushed to MEMORY.md / USER.md
  • The session can be resumed later via /resume or hermes chat --resume

Key Source Files

A lookup table. Do not read it all; find the row that applies to you.

FilePurpose
run_agent.pyAIAgent class — the complete agent loop
agent/prompt_builder.pySystem prompt assembly from memory, skills, context files, personality
agent/context_engine.pyContextEngine ABC — pluggable context management
agent/context_compressor.pyDefault engine — lossy summarization algorithm
agent/prompt_caching.pyAnthropic prompt caching markers and cache metrics
agent/auxiliary_client.pyAuxiliary LLM client for side tasks (vision, summarization)
model_tools.pyTool schema collection, handle_function_call() dispatch
Knowledge check

2 questions answered by this page alone.

Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.

1. In this lesson's table, what is the “Used for” for “anthropicmessages”?
2. Which of these headings does not appear in this lesson?