Academy → Developer GuideOfficial documentation · Arabic guidance

Context Compression and Caching

ضغط السياق والتخزين المؤقت

Advanced19 min readLesson 294 questions✓ 2026-08-18
Before you read

What this page is, and what it holds.

This page covers Context Compression and Caching. It carries a source warning and takes about 19 minutes to read. Stuffing everything into context weakens the answer and raises the cost. Give it only what the task needs.

7sections
16code examples
2tables
2commands
3,284source words
What you will be able to do

Outcomes taken from this page, not a template.

  • Understand what السياق is and when you need it.
  • Run hermes config set compression and hermes plugins and understand what happens next.
  • Read the table and take only the row that applies to you.
  • Avoid the mistake the source warns about.
Identifiers you will meet

Exactly as they appear in Hermes.

Commands
  • hermes config set compression
  • hermes plugins
Page map

Jump to the part you need.

  1. 01Pluggable Context Engine
  2. 02Dual Compression System
  3. 03Configuration
  4. 04Compression Algorithm
  5. 05Before/After Example
  6. 06Prompt Caching (Anthropic)
  7. 07Context Pressure Warnings
The full official page

Nothing summarised away.

The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.

Hermes Agent uses a dual compression system and Anthropic prompt caching to manage context window usage efficiently across long conversations.

Source files: agent/context_engine.py (ABC), agent/context_compressor.py (default engine), agent/prompt_caching.py, gateway/run.py (session hygiene), run_agent.py (search for _compress_context)

Pluggable Context Engine

Settings you configure once. Change one at a time so you can see what each does. Commands here: hermes plugins.

Context management is built on the ContextEngine ABC (agent/context_engine.py). The built-in ContextCompressor is the default implementation, but plugins can replace it with alternative engines (e.g., Lossless Context Management).

YAML3 lines
context:
  engine: "compressor"    # default — built-in lossy summarization
  engine: "lcm"           # example — plugin providing lossless context

The engine is responsible for:

  • Deciding when compaction should fire (should_compress())
  • Performing compaction (compress())
  • Optionally exposing tools the agent can call (e.g., lcm_grep)
  • Tracking token usage from API responses

Selection is config-driven via context.engine in config.yaml. The resolution order:

  1. Check plugins/context_engine/<name>/ directory
  2. Check general plugin system (register_context_engine())
  3. Fall back to built-in ContextCompressor

Plugin engines are never auto-activated — the user must explicitly set context.engine to the plugin's name. The default "compressor" always uses the built-in.

Configure via hermes plugins → Provider Plugins → Context Engine, or edit config.yaml directly.

For building a context engine plugin, see Context Engine Plugins.

Dual Compression System

Explains the idea itself. Read it slowly; the later sections build on it.

Hermes has two separate compression layers that operate independently:

Text10 lines
                     ┌──────────────────────────┐
  Incoming message   │   Gateway Session Hygiene │  Fires at 85% of context
  ─────────────────► │   (pre-agent, rough est.) │  Safety net for large sessions
                     └─────────────┬────────────┘
                                   │
                                   ▼
                     ┌──────────────────────────┐
                     │   Agent ContextCompressor │  Fires at 50% of context (default)
                     │   (in-loop, real tokens)  │  Normal context management
                     └──────────────────────────┘

1. Gateway Session Hygiene (85% threshold)

Located in gateway/run.py (search for Session hygiene: auto-compress). This is a safety net that runs before the agent processes a message. It prevents API failures when sessions grow too large between turns (e.g., overnight accumulation in Telegram/Discord).

  • Threshold: Fixed at 85% of model context length
  • Token source: Prefers actual API-reported tokens from last turn; falls back to rough character-based estimate (estimate_messages_tokens_rough)
  • Fires: Only when len(history) >= 4 and compression is enabled
  • Purpose: Catch sessions that escaped the agent's own compressor

The gateway hygiene threshold is intentionally higher than the agent's compressor. Setting it at 50% (same as the agent) caused premature compression on every turn in long gateway sessions.

2. Agent ContextCompressor (50% threshold, configurable)

Located in agent/context_compressor.py. This is the **primary compression system** that runs inside the agent's tool loop with access to accurate, API-reported token counts.

Configuration

A lookup table. Do not read it all; find the row that applies to you. Commands here: hermes config set compression.

All compression settings are read from config.yaml under the compression key:

YAML23 lines
compression:
  enabled: true              # Enable/disable compression (default: true)
  threshold: 0.50            # Fraction of context window (default: 0.50 = 50%)
  # model_thresholds:        # Per-model threshold overrides (substring match,
  #   "glm-5.2": 0.40        # longest key wins). See "Per-model threshold
  #   "claude-sonnet": 0.35  # overrides" below.
  target_ratio: 0.20         # How much of threshold to keep as tail (default: 0.20)
  tail_mode: legacy          # Tail retention policy: legacy | lean (default: legacy)
  protect_last_n: 20         # Minimum protected tail messages (default: 20)
  min_tail_user_messages: 1  # Real user messages guaranteed in the tail (default: 1)
  codex_gpt55_autoraise: true  # gpt-5.5 on Codex OAuth: raise trigger to 85% (default: true)
  codex_gpt55_autoraise_notice: true  # Show the one-time autoraise notice (default: true)
  codex_app_server_auto: native  # native|hermes|off for Codex app-server thread compaction
  codex_responses_native: false  # gpt-5.6 on direct OpenAI/Codex: server-side compaction (opt-in)
  codex_responses_compact_threshold: 200000  # Server-side compaction trigger (input tokens)
  in_place: true             # Compact on the same session id, no rotation (default: true)

# Summarization model/provider configured under auxiliary:
auxiliary:
  compression:
    model: null              # Override model for summaries (default: auto-detect)
    provider: auto           # Provider: "auto", "openrouter", "nous", "main", etc.
    base_url: null           # Custom OpenAI-compatible endpoint

Parameter Details

ParameterDefaultRangeDescription
threshold0.500.0-1.0Compression triggers when prompt tokens ≥ threshold × context_length
model_thresholds{}mapPer-model overrides of threshold. Keys are substring-matched against the model name (longest match wins). The small-context floor still applies on top (see below)
target_ratio0.200.10-0.80Controls tail protection token budget: threshold_tokens × target_ratio (legacy mode only — lean uses its own clamp)
tail_modelegacylegacy, leanTail retention policy. legacy keeps a target_ratio-sized verbatim tail (~100K+ tokens on big-window models). lean keeps a clamped tail of 2.5% × context window (10K floor, 25K cap) and instead carries continuity in the summary: chunked identifier-preserving digests of the compacted region, a mechanically extracted anchor index (PR numbers, SHAs, paths, error strings — regex, never paraphrased), every real user message quoted verbatim (newest-first budget), and a session_search recovery pointer so the agent can re-access anything summarized away. Result on 500K-token real sessions: ~49K retained vs ~162K, with higher recall when paired with recovery (see evals/compaction/results/). Costs a few extra summarizer calls at the compaction boundary. Old tool results inside the lean tail are demoted to one-line stubs carrying a recovery pointer
protect_last_n20≥1Minimum number of recent messages always preserved
min_tail_user_messages1≥1Minimum number of REAL (actionable) user messages guaranteed to survive in the uncompressed tail. 1 = the existing single last-user anchor (behavior-preserving default). Raise to e.g. 3 to keep the last 3 real user turns verbatim even when bulky tool outputs fill the tail token budget. Blank platform echoes, compaction handoffs, and synthetic continuation rows never count toward N. The guarantee wins over the tail token budget — the tail may exceed the budget when the anchor pulls the cut back
protect_first_n3(hardcoded)System prompt + first exchange always preserved
idle_compact_after_seconds0≥0 secondsOpt-in: compact up front when a session resumes after this many seconds idle (0 = disabled). Skips when context ≤ threshold × target_ratio; honors cooldown/anti-thrash/lock guards
codex_gpt55_autoraisetrueboolRaise the trigger to 85% for gpt-5.5 on the ChatGPT Codex OAuth route (see below). Set false to keep the global threshold
codex_gpt55_autoraise_noticetrueboolShow the one-time Codex gpt-5.5 autoraise notice. Set false to keep the 85% autoraise but suppress the banner
codex_app_server_autonativenative, hermes, offThread-compaction mode for Codex app-server sessions (see below)
codex_responses_nativefalseboolOpt in to OpenAI's server-side compaction on the Responses API. Engages only for gpt-5.6-family models on the direct OpenAI API or a ChatGPT Codex subscription (see below)
codex_responses_compact_threshold200000≥1 tokensServer-side compaction trigger in input tokens. Clamped below the local compression threshold at request time so the server compacts first
in_placetrueboolCompact on the same session id instead of rotating to a new one (see below)

In-place compaction (single stable session id)

With compression.in_place: true (the default), a compaction rewrites the live message list on the same session id: the system prompt is rebuilt, the summarized middle is swapped in, and the pre-compaction turns are soft-archived under the same id (active=0, compacted=1 in the session store) — still searchable via session_search and recoverable, never deleted. There is no parent_session_id chain and no name #N renumbering; one conversation keeps one durable id for its whole life. This eliminated the session-rotation bug cluster (lost /goal state, orphaned sessions, search gaps across boundaries).

Consumers observe the mode rather than diffing session ids:

  • The session:compress event carries in_place: true/false and old_session_id (empty string in in-place mode, since there is no old id).
  • The gateway re-baselines transcript handling from the agent's rotation-independent _last_compaction_in_place flag, not from an id-change diff.

Set in_place: false to restore the legacy rotating path, where each compaction commits a new session id linked to the previous one via parent_session_id.

Per-model threshold overrides

compression.model_thresholds lets you trigger compaction at different points depending on the active model — useful when you swap between models with very different context windows (e.g. a 1M-context model can compress later while a 128K model should compress earlier):

YAML6 lines
compression:
  threshold: 0.50
  model_thresholds:
    "glm-5.2": 0.40
    "glm-5.2-1M": 0.25
    "claude-sonnet": 0.35

Resolution rules:

  • Keys are substring-matched against the model name; the longest matching key wins (glm-5.2-1M beats glm-5.2 for model glm-5.2-1M).
  • When no key matches (or the map is empty), the global threshold applies.
  • The override is re-resolved on every /model switch; switching to a model with no matching key falls back to the global threshold.
  • The small-context floor still applies on top of overrides (raise-only): models with context windows below 512K are floored at 0.75, so an override below the floor is raised to 0.75, while an override above it (e.g. 0.80) wins.

Plugin context engines can reuse the same resolution logic via from agent.context_compressor import resolve_model_threshold; engines that override update_model() own their own compaction policy and may ignore the map.

Codex gpt-5.5 threshold autoraise

The ChatGPT Codex OAuth backend hard-caps gpt-5.5 at a 272K context window (the same slug exposes 1.05M on OpenAI's direct API and OpenRouter, and 400K on GitHub Copilot). At the default 50% trigger, compaction would fire at ~136K — half the window the model can actually use. When the active route is Codex OAuth (provider: openai-codex) and the model is gpt-5.5, Hermes raises the trigger to 85% (~231K) and shows a notice with the opt-out command. The notice is shown once per profile — a marker under $HERMES_HOME (.codex_gpt55_autoraise_notice) records that it ran, so repeated agent/session inits (e.g. every inbound gateway message) don't re-emit it; if the raised threshold later changes it re-notifies once. Only this exact route is affected; gpt-5.5 on any other provider keeps your global threshold. To opt back down to the global value:

Shell1 line
hermes config set compression.codex_gpt55_autoraise false

To keep the 85% autoraise but hide only the one-time notice:

Shell1 line
hermes config set compression.codex_gpt55_autoraise_notice false

Codex app-server thread compaction

Codex app-server sessions (api_mode: codex_app_server — the codex CLI/agent runtime) are different from every other route: the codex agent owns the backing thread context, so Hermes' auxiliary summarizer cannot shrink it — rewriting the local transcript mirror leaves the real thread growing unbounded until a hard context reset. For this runtime, compaction goes through the app-server's own mechanism instead:

  • Manual compaction (/compress) asks the app-server to compact the thread (thread/compact/start) and waits for the compaction turn to complete.
  • Automatic compaction is controlled by compression.codex_app_server_auto: the default native lets the app-server decide when to compact and Hermes records the resulting compaction events (compression counters, session events). Set hermes to let Hermes' compression threshold initiate app-server compaction, or off to disable Hermes-initiated automatic compaction entirely (codex may still compact natively).

Hermes' local transcript is never rewritten on this runtime — state.db records the compaction boundary while the visible transcript stays intact. All other routes (including Codex OAuth chat sessions) keep Hermes' summary compressor.

Native Responses compaction (gpt-5.6 on direct OpenAI / Codex subscription)

OpenAI's Responses API supports server-side compaction: when a request includes context_management: [{type: "compaction", compact_threshold: N}] and the rendered input crosses N tokens, the server prunes older context into an opaque encrypted compaction output item. Hermes captures that item into the assistant message's existing replay sidecar and sends it back on subsequent turns, standing in for the pruned history — long-horizon recall without a client-side summary pass, and ZDR-friendly (store: false, no previous_response_id).

Opt in with compression.codex_responses_native: true. The gate is deliberately narrow, re-checked on every request:

  • Models: the gpt-5.6 family only. Other models fail server-side when the field is present (gpt-5.1/5.2 return HTTP 500 or stall the stream — there is no structured rejection to downgrade on, verified live Aug 2026).
  • Routes: api.openai.com (OpenAI API key) or the ChatGPT Codex backend (Codex subscription OAuth) only. xAI, GitHub/Copilot, OpenRouter, relays, and local servers never see the field.

Everything else about compression is unchanged: the local compressor stays armed as the fallback owner (the native threshold is clamped ~8K tokens below the local trigger so the server compacts first), and a structured provider rejection of the field disables native compaction for the session and retries the request without it. Switching the session to a non-eligible model or route simply stops the field from being sent — captured checkpoints are dropped from replay by the existing cross-issuer guard when the endpoint changes.

Computed Values (for a 200K context model at defaults)

Text4 lines
context_length       = 200,000
threshold_tokens     = 200,000 × 0.50 = 100,000
tail_token_budget    = 100,000 × 0.20 = 20,000
max_summary_tokens   = min(200,000 × 0.05, 12,000) = 10,000

Compression Algorithm

Carries a warning. Read it before running anything here. The upstream warning appears below.

The ContextCompressor.compress() method follows a 4-phase algorithm:

Phase 1: Prune Old Tool Results (cheap, no LLM call)

Old tool results (>200 chars) outside the protected tail are replaced with:

Text1 line
[Old tool output cleared to save context space]

This is a cheap pre-pass that saves significant tokens from verbose tool outputs (file contents, terminal output, search results).

Phase 2: Determine Boundaries

Text8 lines
┌─────────────────────────────────────────────────────────────┐
│  Message list                                               │
│                                                             │
│  [0..2]  ← protect_first_n (system + first exchange)        │
│  [3..N]  ← middle turns → SUMMARIZED                        │
│  [N..end] ← tail (by token budget OR protect_last_n)        │
│                                                             │
└─────────────────────────────────────────────────────────────┘

Tail protection is token-budget based: walks backward from the end, accumulating tokens until the budget is exhausted. Falls back to the fixed protect_last_n count if the budget would protect fewer messages.

Boundaries are aligned to avoid splitting tool_call/tool_result groups. The _align_boundary_backward() method walks past consecutive tool results to find the parent assistant message, keeping groups intact.

Phase 3: Generate Structured Summary

The middle turns are summarized using the auxiliary LLM with a structured template:

Text25 lines
## Goal
[What the user is trying to accomplish]

## Constraints & Preferences
[User preferences, coding style, constraints, important decisions]

## Progress
### Done
[Completed work — specific file paths, commands run, results]
### In Progress
[Work currently underway]
### Blocked
[Any blockers or issues encountered]

## Key Decisions
[Important technical decisions and why]

## Relevant Files
[Files read, modified, or created — with brief note on each]

## Next Steps
[What needs to happen next]

## Critical Context
[Specific values, error messages, configuration details]

Summary budget scales with the amount of content being compressed:

  • Formula: content_tokens × 0.20 (the _SUMMARY_RATIO constant)
  • Minimum: 2,000 tokens
  • Maximum: min(context_length × 0.05, 12,000) tokens

Phase 4: Assemble Compressed Messages

The compressed message list is:

  1. Head messages (with a note appended to system prompt on first compression)
  2. Summary message (role chosen to avoid consecutive same-role violations)
  3. Tail messages (unmodified)

Orphaned tool_call/tool_result pairs are cleaned up by _sanitize_tool_pairs():

  • Tool results referencing removed calls → removed
  • Tool calls whose results were removed → stub result injected

Iterative Re-compression

On subsequent compressions, the previous summary is passed to the LLM with instructions to update it rather than summarize from scratch. This preserves information across multiple compactions — items move from "In Progress" to "Done", new progress is added, and obsolete information is removed.

The _previous_summary field on the compressor instance stores the last summary text for this purpose.

Before/After Example

Explains the idea itself. Read it slowly; the later sections build on it.

Before Compression (45 messages, ~95K tokens)

Text14 lines
[0] system:    "You are a helpful assistant..." (system prompt)
[1] user:      "Help me set up a FastAPI project"
[2] assistant: <tool_call> terminal: mkdir project </tool_call>
[3] tool:      "directory created"
[4] assistant: <tool_call> write_file: main.py </tool_call>
[5] tool:      "file written (2.3KB)"
    ... 30 more turns of file editing, testing, debugging ...
[38] assistant: <tool_call> terminal: pytest </tool_call>
[39] tool:      "8 passed, 2 failed\n..."  (5KB output)
[40] user:      "Fix the failing tests"
[41] assistant: <tool_call> read_file: tests/test_api.py </tool_call>
[42] tool:      "import pytest\n..."  (3KB)
[43] assistant: "I see the issue with the test fixtures..."
[44] user:      "Great, also add error handling"

After Compression (25 messages, ~45K tokens)

Text31 lines
[0] system:    "You are a helpful assistant...
               [Note: Some earlier conversation turns have been compacted...]"
[1] user:      "Help me set up a FastAPI project"
[2] assistant: "[CONTEXT COMPACTION] Earlier turns were compacted...

               ## Goal
               Set up a FastAPI project with tests and error handling

               ## Progress
               ### Done
               - Created project structure: main.py, tests/, requirements.txt
               - Implemented 5 API endpoints in main.py
               - Wrote 10 test cases in tests/test_api.py
               - 8/10 tests passing

               ### In Progress
               - Fixing 2 failing tests (test_create_user, test_delete_user)

               ## Relevant Files
               - main.py — FastAPI app with 5 endpoints
               - tests/test_api.py — 10 test cases
               - requirements.txt — fastapi, pytest, httpx

               ## Next Steps
               - Fix failing test fixtures
               - Add error handling"
[3] user:      "Fix the failing tests"
[4] assistant: <tool_call> read_file: tests/test_api.py </tool_call>
[5] tool:      "import pytest\n..."
[6] assistant: "I see the issue with the test fixtures..."
[7] user:      "Great, also add error handling"

Prompt Caching (Anthropic)

A lookup table. Do not read it all; find the row that applies to you.

Source: agent/prompt_caching.py

Reduces input token costs by ~75% on multi-turn conversations by caching the conversation prefix. Uses Anthropic's cache_control breakpoints.

Strategy: systemand3

Anthropic allows a maximum of 4 cache_control breakpoints per request. Hermes uses the "system_and_3" strategy:

Text4 lines
Breakpoint 1: System prompt           (stable across all turns)
Breakpoint 2: 3rd-to-last non-system message  ─┐
Breakpoint 3: 2nd-to-last non-system message   ├─ Rolling window
Breakpoint 4: Last non-system message          ─┘

How It Works

apply_anthropic_cache_control() deep-copies the messages and injects cache_control markers:

Python4 lines
# Cache marker format
marker = {"type": "ephemeral"}
# Or for 1-hour TTL:
marker = {"type": "ephemeral", "ttl": "1h"}

The marker is applied differently based on content type:

Content TypeWhere Marker Goes
String contentConverted to [{"type": "text", "text": ..., "cache_control": ...}]
List contentAdded to the last element's dict
None/emptyAdded as msg["cache_control"]
Tool messagesAdded as msg["cache_control"] (native Anthropic only)

Cache-Aware Design Patterns

  1. Stable system prompt: The system prompt is breakpoint 1 and cached across all turns. Avoid mutating it mid-conversation (compression appends a note only on the first compaction).
  1. Message ordering matters: Cache hits require prefix matching. Adding or removing messages in the middle invalidates the cache for everything after.
  1. Compression cache interaction: After compression, the cache is invalidated for the compressed region but the system prompt cache survives. The rolling 3-message window re-establishes caching within 1-2 turns.
  1. TTL selection: Default is 5m (5 minutes). Use 1h for long-running sessions where the user takes breaks between turns.
  1. Model identity is part of the cache key: Provider-side caches are scoped to the model (and account/API key) serving the request. Any mid-conversation model change — an explicit /model switch, primary-model fallback, or a credential-pool rotation onto a different account — means the next request gets zero cache hits and re-reads the full conversation at undiscounted input price. This is inherent to how provider caches work, not something Hermes can avoid; user-facing docs for /model, fallback providers, and credential pools carry cost warnings for this reason. Don't add features that silently swap the model or credentials mid-session.

Enabling Prompt Caching

Prompt caching is automatically enabled when:

  • The model is an Anthropic Claude model (detected by model name)
  • The provider supports cache_control (native Anthropic API or OpenRouter)
YAML3 lines
# config.yaml — TTL is configurable (must be "5m" or "1h")
prompt_caching:
  cache_ttl: "5m"

The CLI shows caching status at startup:

Text1 line
💾 Prompt caching: ENABLED (Claude via OpenRouter, 5m TTL)

Context Pressure Warnings

Explains the idea itself. Read it slowly; the later sections build on it.

Intermediate context-pressure warnings have been removed (see the iteration-budget block in run_agent.py, which notes: "No intermediate pressure warnings — they caused models to 'give up' prematurely on complex tasks"). Compression fires when prompt tokens reach the configured compression.threshold (default 50%) with no prior warning step; gateway session hygiene fires as the secondary safety net at 85% of the model's context window.

Knowledge check

4 questions answered by this page alone.

Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.

1. In this lesson's table, what is the “Default” for “threshold”?
2. Which warning does the source state in this lesson?
3. Which of these headings does not appear in this lesson?
4. Which configuration key appears in this lesson's examples?