Batch Processing
Batch Processing
Start with meaning, then move to detail.
This lesson explains Batch Processing as part of extending Hermes and connecting external tools. You will learn what it does, when it matters, and the smallest safe test that proves it works.
If you are new, do not memorize names. Focus on three questions: what problem does this solve, what access does it need, and how can you verify the result?
For practice, inspect the first example, identify its effects, run it on test data, and compare the result with the source claim.
For advanced readers, inspect Overview, Quick Start, Dataset Format, then verify failure modes and version compatibility.
Complete installation and one successful task before adding new capabilities.
A clear outcome before you read.
- Understand Batch Processing without assumed prior knowledge.
- Separate the source description from what still needs testing in your environment.
- Read the first command and identify its inputs and outputs before copying it.
Short definitions before the details.
- Provider
- The service that runs or provides access and authentication to a model.
- Tool
- A structured action the agent can call to read or change something.
Generate agent trajectories at scale — parallel processing, checkpointing, and toolset distributions
What does the source say, and in what order?
- 01Overview
Start here to understand the core idea or structure.
- 02Quick Start
Read this after the foundation, then connect it to the previous step.
- 03Dataset Format
Read this after the foundation, then connect it to the previous step.
- 04Configuration Options
Read this after the foundation, then connect it to the previous step.
- 05Provider Routing (OpenRouter)
Read this after the foundation, then connect it to the previous step.
- 06Reasoning Control
Read this after the foundation, then connect it to the previous step.
- 07Advanced Options
Read this after the foundation, then connect it to the previous step.
- 08Toolset Distributions
Read this after the foundation, then connect it to the previous step.
- 09Output Format
Read this after the foundation, then connect it to the previous step.
- 10Trajectory Format
Finish here to verify the result and special cases.
Copy only after you understand the effect.
# Basic batch run
python batch_runner.py \
--dataset_file=data/prompts.jsonl \
--batch_size=10 \
--run_name=my_first_run \
--model=anthropic/claude-sonnet-4.6 \
--num_workers=4
# Resume an interrupted run
python batch_runner.py \
--dataset_file=data/prompts.jsonl \
--batch_size=10 \
--run_name=my_first_run \
--resume
# List available toolset distributions
python batch_runner.py --list_distributionsEntries can optionally include:
- `image` or `docker_image`: A container image to use for this prompt's sandbox (works with Docker, Modal, and Singularity backends)
- `cwd`: Working directory override for the task's terminal session
## Configuration Options
| Parameter | Default | Description |
|-----------|---------|-------------|
| `--dataset_file` | (required) | Path to JSONL dataset |
| `--batch_size` | (required) | Prompts per batch |
| `--run_name` | (required) | Name for this run (used for output dir and checkpointing) |
| `--distribution` | `"default"` | Toolset distribution to sampl### Trajectory Format
Each line in `trajectories.jsonl` is a JSON object:Read the first command and identify its inputs and outputs before copying it.
Match every command to your installed Hermes version, review the files and accounts it can reach, and use non-sensitive data for the first test. If this explanation differs from the source, the official source wins.