الأكاديمية ← ميزات Hermesتوثيق رسمي · إرشاد عربي

المعالجة الدفعية للمهام المتكررة

Batch Processing

متوسط إلى متقدم6 دقائق قراءةالدرس 333 أسئلة✓ 2026-08-18
قبل أن تقرأ

ما هذه الصفحة، وماذا تحتوي.

الجلسة: محادثة واحدة بكل ما دار فيها. تفتحها وتغلقها وتعود إليها لاحقًا. فصل العمل إلى جلسات يبقي كل موضوع نظيفًا، ويجعل الرجوع خطوة إلى الوراء ممكنًا عند الخطأ. القراءة نحو 6 دقائق. انتبه: الجلسة الطويلة جدًا تُنسي الوكيل بدايتها وتكلّف أكثر. ابدأ جلسة جديدة لكل مهمة مختلفة.

10أقسام
7أمثلة برمجية
4جداول
0أوامر
1,027كلمة من المصدر
الوصف الرسمي في سطر

Generate agent trajectories at scale — parallel processing, checkpointing, and toolset distributions

ماذا ستستطيع بعدها

نتائج مأخوذة من هذه الصفحة، لا من قالب.

  • تعرف ما الجلسة ولماذا قد تحتاجه.
  • تقرأ الجدول وتأخذ منه السطر الذي يخصّك فقط.
  • تعرف الخطأ الشائع: الجلسة الطويلة جدًا تُنسي الوكيل بدايتها وتكلّف أكثر. ابدأ جلسة جديدة لكل مهمة مختلفة.
خريطة الصفحة

انتقل مباشرة إلى ما تحتاجه.

  1. 01Overview
  2. 02Quick Start
  3. 03Dataset Format
  4. 04Configuration Options
  5. 05Toolset Distributions
  6. 06Output Format
  7. 07Checkpointing
  8. 08Quality Filtering
  9. 09Statistics
  10. 10Use Cases
الصفحة الرسمية كاملة

بلا اختصار أو حذف.

النص أدناه منقول من المصدر الرسمي بالإنجليزية حتى تبقى الأوامر والأسماء دقيقة كما هي. قبل كل قسم شرح عربي يوضّح ما بداخله.

Batch processing lets you run the Hermes agent across hundreds or thousands of prompts in parallel, generating structured trajectory data. This is primarily used for training data generation — producing ShareGPT-format trajectories with tool usage statistics that can be used for fine-tuning or evaluation.

Overview

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه. تذكير: محادثة واحدة بكل ما دار فيها. تفتحها وتغلقها وتعود إليها لاحقًا.

The batch runner (batch_runner.py) processes a JSONL dataset of prompts, running each through a full agent session with tool access. Each prompt gets its own isolated environment. The output is structured trajectory data with full conversation history, tool call statistics, and reasoning coverage metrics.

Quick Start

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية.

Shell17 سطرًا
# Basic batch run
python batch_runner.py \
    --dataset_file=data/prompts.jsonl \
    --batch_size=10 \
    --run_name=my_first_run \
    --model=anthropic/claude-sonnet-4.6 \
    --num_workers=4

# Resume an interrupted run
python batch_runner.py \
    --dataset_file=data/prompts.jsonl \
    --batch_size=10 \
    --run_name=my_first_run \
    --resume

# List available toolset distributions
python batch_runner.py --list_distributions

Dataset Format

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

The input dataset is a JSONL file (one JSON object per line). Each entry must have a prompt field:

JSONL3 أسطر
{"prompt": "Write a Python function that finds the longest palindromic substring"}
{"prompt": "Create a REST API endpoint for user authentication using Flask"}
{"prompt": "Debug this error: TypeError: cannot unpack non-iterable NoneType object"}

Entries can optionally include:

  • image or docker_image: A container image to use for this prompt's sandbox (works with Docker, Modal, and Singularity backends)
  • cwd: Working directory override for the task's terminal session

Configuration Options

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

ParameterDefaultDescription
--dataset_file(required)Path to JSONL dataset
--batch_size(required)Prompts per batch
--run_name(required)Name for this run (used for output dir and checkpointing)
--distribution"default"Toolset distribution to sample from
--modelclaude-sonnet-4.6Model to use
--base_urlhttps://openrouter.ai/api/v1API base URL
--api_key(env var)API key for model
--max_turns10Maximum tool-calling iterations per prompt
--num_workers4Parallel worker processes
--resumefalseResume from checkpoint
--verbosefalseEnable verbose logging
--max_samplesallOnly process first N samples from dataset
--max_tokensmodel defaultMaximum tokens per model response

Provider Routing (OpenRouter)

ParameterDescription
--providers_allowedComma-separated providers to allow (e.g., "anthropic,openai")
--providers_ignoredComma-separated providers to ignore (e.g., "together,deepinfra")
--providers_orderComma-separated preferred provider order
--provider_sortSort by "price", "throughput", or "latency"

Reasoning Control

ParameterDescription
--reasoning_effortReasoning effort: none, minimal, low, medium, high, xhigh, max, ultra
--reasoning_disabledCompletely disable reasoning/thinking tokens

Advanced Options

ParameterDescription
--ephemeral_system_promptSystem prompt used during execution but NOT saved to trajectories
--log_prefix_charsCharacters to show in log previews (default: 100)
--prefill_messages_filePath to JSON file with prefill messages for few-shot priming

Toolset Distributions

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Each prompt gets a randomly sampled set of toolsets from a distribution. This ensures training data covers diverse tool combinations. Use --list_distributions to see all available distributions.

In the current implementation, distributions assign a probability to each individual toolset. The sampler flips each toolset independently, then guarantees that at least one toolset is enabled. This is different from a hand-authored table of prebuilt combinations.

Output Format

إعدادات تضبطها مرة وتنساها. غيّر واحدًا في كل مرة حتى تعرف أثر كل تغيير.

All output goes to data/<run_name>/:

Text7 أسطر
data/my_run/
├── trajectories.jsonl    # Combined final output (all batches merged)
├── batch_0.jsonl         # Individual batch results
├── batch_1.jsonl
├── ...
├── checkpoint.json       # Resume checkpoint
└── statistics.json       # Aggregate tool usage stats

Trajectory Format

Each line in trajectories.jsonl is a JSON object:

JSON27 سطرًا
{
  "prompt_index": 42,
  "conversations": [
    {"from": "human", "value": "Write a function..."},
    {"from": "gpt", "value": "I'll create that function...",
     "tool_calls": [...]},
    {"from": "tool", "value": "..."},
    {"from": "gpt", "value": "Here's the completed function..."}
  ],
  "metadata": {
    "batch_num": 2,
    "timestamp": "2026-01-15T10:30:00",
    "model": "anthropic/claude-sonnet-4.6"
  },
  "completed": true,
  "partial": false,
  "api_calls": 3,
  "toolsets_used": ["terminal", "file"],
  "tool_stats": {
    "terminal": {"count": 2, "success": 2, "failure": 0},
    "read_file": {"count": 1, "success": 1, "failure": 0}
  },
  "tool_error_counts": {
    "terminal": 0,
    "read_file": 0
  }
}

The conversations field uses a ShareGPT-like format with from and value fields. Tool stats are normalized to include all possible tools with zero defaults, ensuring consistent schema across entries for HuggingFace datasets compatibility.

Checkpointing

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

The batch runner has robust checkpointing for fault tolerance:

  • Checkpoint file: Saved after each batch completes, tracking which prompt indices are done
  • Content-based resume: On --resume, the runner scans existing batch files and matches completed prompts by their actual text content (not just indices), enabling recovery even if the dataset order changes
  • Failed prompts: Only successfully completed prompts are marked as done — failed prompts will be retried on resume
  • Batch merging: On completion, all batch files (including from previous runs) are merged into a single trajectories.jsonl

How Resume Works

  1. Scan all batch_*.jsonl files for completed prompts (by content matching)
  2. Filter the dataset to exclude already-completed prompts
  3. Re-batch the remaining prompts
  4. Process only the remaining prompts
  5. Merge all batch files (old + new) into final output

Quality Filtering

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

The batch runner applies automatic quality filtering:

  • No-reasoning filter: Samples where zero assistant turns contain reasoning (no <REASONING_SCRATCHPAD> or native thinking tokens) are discarded
  • Corrupted entry filter: Entries with hallucinated tool names (not in the valid tool list) are filtered out during the final merge
  • Reasoning statistics: Tracks percentage of turns with/without reasoning across the entire run

Statistics

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

After completion, the runner prints comprehensive statistics:

  • Tool usage: Call counts, success/failure rates per tool
  • Reasoning coverage: Percentage of assistant turns with reasoning
  • Samples discarded: Count of samples filtered for lacking reasoning
  • Duration: Total processing time

Statistics are also saved to statistics.json for programmatic analysis.

Use Cases

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Training Data Generation

Generate diverse tool-use trajectories for fine-tuning:

Shell8 أسطر
python batch_runner.py \
    --dataset_file=data/coding_prompts.jsonl \
    --batch_size=20 \
    --run_name=coding_v1 \
    --model=anthropic/claude-sonnet-4.6 \
    --num_workers=8 \
    --distribution=default \
    --max_turns=15

Model Evaluation

Evaluate how well a model uses tools across standardized prompts:

Shell7 أسطر
python batch_runner.py \
    --dataset_file=data/eval_suite.jsonl \
    --batch_size=10 \
    --run_name=eval_gpt4 \
    --model=openai/gpt-4o \
    --num_workers=4 \
    --max_turns=10

Per-Prompt Container Images

For benchmarks requiring specific environments, each prompt can specify its own container image:

JSONL3 أسطر
{"prompt": "Install numpy and compute eigenvalues of a 3x3 matrix", "image": "python:3.11-slim"}
{"prompt": "Compile this Rust program and run it", "image": "rust:1.75"}
{"prompt": "Set up a Node.js Express server", "image": "node:20-alpine", "cwd": "/app"}

The batch runner verifies Docker images are accessible before running each prompt.

اختبار الفهم

3 أسئلة إجاباتها كلها في هذه الصفحة.

كل خيار اسم حقيقي من توثيق Hermes. حتى الخيارات الخاطئة حقيقية، لكنها من صفحات أخرى.

1. في جدول هذا الدرس، ما «Default» المقابل لـ«--resume»؟
2. أي عنوان من التالي لا يظهر في هذا الدرس؟
3. أي مفتاح إعداد يظهر في أمثلة هذا الدرس؟