Directory → SKILL
SKILLoptionalHermes Optional Skills

Slime

RL post-training for LLMs with Megatron and SGLang

Reinforcement LearningMegatron-LMSGLangGRPOPost-TrainingGLMOptionalHermes skill
Last registry verification2026-08-18v1.0.0Orchestra Research
Plain meaning

What does it add to Hermes?

RL post-training for LLMs with Megatron and SGLang

Slime is a skill related to extending the agent. It adds a capability or workflow to Hermes. The publisher description explains the intent, while granted permissions determine what it can actually do.

This plain-language explanation is based on the publisher description. The original text remains visible for verification.

Use it when

Use it when your goal in extending the agent is clear and you can limit it to the data and actions it actually needs.

Skip it when

Do not add it merely to experiment when Hermes already has a simpler path, or when you cannot review its source and permissions.

Who is it for?

Best for users who want a repeatable way of working inside Hermes.

Safe first test

Start with non-sensitive data and a small task whose result can be verified and reversed.

Original publisher description

RL post-training for LLMs with Megatron and SGLang

✓
Data source

This entry was indexed from Hermes Optional Skills. Our explanation interprets the type and domain without inventing a capability not present upstream.

!
Security review

The source is official or editorially reviewed, but you still need to review permissions and version compatibility.

Safe setup path

Inspect, install, then test.

  1. 01
    Open the source

    Match the publisher, license, and description to your need. Check the real update history.

  2. 02
    Review permissions and secrets

    Never paste a secret value into this site. Use environment-variable names and grant the smallest scope.

  3. 03
    Copy setup only after review

    The controls below copy text. They do not execute commands on your device.

  4. 04
    Test with a non-sensitive task

    Inspect the visible tools, then exclude write or delete tools you do not need.

Install command

Review the command, then copy it.

hermes skills install slime

Hermes Belarabi does not execute this command. Installation happens on your device and remains subject to Hermes scanning and your review.

The full skill definition

Exactly what Hermes loads when this skill runs.

Reproduced from the official documentation. Read it before enabling the skill: this text becomes the agent's instructions.

RL post-training for LLMs with Megatron and SGLang.

Skill metadata

A lookup table. Do not read it all; find the row that applies to you.

SourceOptional — install with hermes skills install official/mlops/slime
Pathoptional-skills/mlops/slime
Version1.0.0
AuthorOrchestra Research
LicenseMIT
Dependenciessglang-router>=0.2.3, ray, torch>=2.0.0, transformers>=4.40.0
Platformslinux, macos
TagsReinforcement Learning, Megatron-LM, SGLang, GRPO, Post-Training, GLM

Reference: full SKILL.md

Explains the idea itself. Read it slowly; the later sections build on it.

slime is an LLM post-training framework from Tsinghua's THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation.

When to Use slime

Explains the idea itself. Read it slowly; the later sections build on it.

Choose slime when you need:

  • Megatron-LM native training with SGLang inference
  • Custom data generation workflows with flexible data buffers
  • Training GLM, Qwen3, DeepSeek V3, or Llama 3 models
  • Research-grade framework with production backing (Z.ai)

Consider alternatives when:

  • You need enterprise-grade stability features → use miles
  • You want flexible backend swapping → use verl
  • You need PyTorch-native abstractions → use torchforge

Key Features

Explains the idea itself. Read it slowly; the later sections build on it.

  • Training: Megatron-LM with full parallelism support (TP, PP, DP, SP)
  • Rollout: SGLang-based high-throughput generation with router
  • Data Buffer: Flexible prompt management and sample storage
  • Models: GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3

Architecture Overview

Explains the idea itself. Read it slowly; the later sections build on it.

Text13 lines
┌─────────────────────────────────────────────────────────┐
│                    Data Buffer                          │
│ - Prompt initialization and management                  │
│ - Custom data generation and filtering                  │
│ - Rollout sample storage                                │
└─────────────┬───────────────────────────┬───────────────┘
              │                           │
┌─────────────▼───────────┐ ┌─────────────▼───────────────┐
│ Training (Megatron-LM)  │ │ Rollout (SGLang + Router)   │
│ - Actor model training  │ │ - Response generation       │
│ - Critic (optional)     │ │ - Reward/verifier output    │
│ - Weight sync to rollout│ │ - Multi-turn support        │
└─────────────────────────┘ └─────────────────────────────┘

Installation

Ordered, practical steps. Run one and confirm it worked before moving on.

Shell7 lines
# Recommended: Docker
docker pull slimerl/slime:latest
docker run --rm --gpus all --ipc=host --shm-size=16g \
  -it slimerl/slime:latest /bin/bash

# Inside container
cd /root/slime && pip install -e . --no-deps

From Source

Shell4 lines
git clone https://github.com/THUDM/slime.git
cd slime
pip install -r requirements.txt
pip install -e .

Quick Start: GRPO Training

Ordered, practical steps. Run one and confirm it worked before moving on.

Shell16 lines
# Source model configuration
source scripts/models/qwen3-4B.sh

# Launch training
python train.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 4 \
    --rollout-num-gpus 4 \
    --advantage-estimator grpo \
    --use-kl-loss --kl-loss-coef 0.001 \
    --rollout-batch-size 32 \
    --n-samples-per-prompt 8 \
    --global-batch-size 256 \
    --num-rollout 3000 \
    --prompt-data /path/to/data.jsonl \
    ${MODEL_ARGS[@]} ${CKPT_ARGS[@]}

---

Workflow 1: Standard GRPO Training

Settings you configure once. Change one at a time so you can see what each does. Set MODEL_ARGS in your environment, not in the chat.

Use this workflow for training reasoning models with group-relative advantages.

Prerequisites Checklist

  • [ ] Docker environment or Megatron-LM + SGLang installed
  • [ ] Model checkpoint (HuggingFace or Megatron format)
  • [ ] Training data in JSONL format

Step 1: Prepare Data

Python3 lines
# data.jsonl format
{"prompt": "What is 2 + 2?", "label": "4"}
{"prompt": "Solve: 3x = 12", "label": "x = 4"}

Or with chat format:

Python7 lines
{
    "prompt": [
        {"role": "system", "content": "You are a math tutor."},
        {"role": "user", "content": "What is 15 + 27?"}
    ],
    "label": "42"
}

Step 2: Configure Model

Choose a pre-configured model script:

Shell6 lines
# List available models
ls scripts/models/
# glm4-9B.sh, qwen3-4B.sh, qwen3-30B-A3B.sh, deepseek-v3.sh, llama3-8B.sh, ...

# Source your model
source scripts/models/qwen3-4B.sh

Step 3: Launch Training

Shell18 lines
python train.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 8 \
    --rollout-num-gpus 8 \
    --advantage-estimator grpo \
    --use-kl-loss \
    --kl-loss-coef 0.001 \
    --prompt-data /path/to/train.jsonl \
    --input-key prompt \
    --label-key label \
    --apply-chat-template \
    --rollout-batch-size 32 \
    --n-samples-per-prompt 8 \
    --global-batch-size 256 \
    --num-rollout 3000 \
    --save-interval 100 \
    --eval-interval 50 \
    ${MODEL_ARGS[@]}

Step 4: Monitor Training

  • [ ] Check TensorBoard: tensorboard --logdir outputs/
  • [ ] Verify reward curves are increasing
  • [ ] Monitor GPU utilization across nodes

---

Workflow 2: Asynchronous Training

Settings you configure once. Change one at a time so you can see what each does. Set MODEL_ARGS in your environment, not in the chat.

Use async mode for higher throughput by overlapping rollout and training.

When to Use Async

  • Large models with long generation times
  • High GPU idle time in synchronous mode
  • Sufficient memory for buffering

Launch Async Training

Shell8 lines
python train_async.py \
    --actor-num-nodes 1 \
    --actor-num-gpus-per-node 8 \
    --rollout-num-gpus 8 \
    --advantage-estimator grpo \
    --async-buffer-size 4 \
    --prompt-data /path/to/train.jsonl \
    ${MODEL_ARGS[@]}

Async-Specific Parameters

Shell2 lines
--async-buffer-size 4        # Number of rollouts to buffer
--update-weights-interval 2  # Sync weights every N rollouts

---

Workflow 3: Multi-Turn Agentic Training

Settings you configure once. Change one at a time so you can see what each does. Set MODEL_ARGS in your environment, not in the chat.

Use this workflow for training agents with tool use or multi-step reasoning.

Prerequisites

  • [ ] Custom generate function for multi-turn logic
  • [ ] Tool/environment interface

Step 1: Define Custom Generate Function

Python23 lines
# custom_generate.py
async def custom_generate(args, samples, evaluation=False):
    """Multi-turn generation with tool calling."""
    for sample in samples:
        conversation = sample.prompt

        for turn in range(args.max_turns):
            # Generate response
            response = await generate_single(conversation)

            # Check for tool call
            tool_call = extract_tool_call(response)
            if tool_call:
                tool_result = execute_tool(tool_call)
                conversation.append({"role": "assistant", "content": response})
                conversation.append({"role": "tool", "content": tool_result})
            else:
                break

        sample.response = response
        sample.reward = compute_reward(sample)

    return samples

Step 2: Launch with Custom Function

Shell5 lines
python train.py \
    --custom-generate-function-path custom_generate.py \
    --max-turns 5 \
    --prompt-data /path/to/agent_data.jsonl \
    ${MODEL_ARGS[@]}

See examples/search-r1/ for a complete multi-turn search example.

---

Configuration Reference

Explains the idea itself. Read it slowly; the later sections build on it.

Three Argument Categories

slime uses three types of arguments:

1. Megatron Arguments (passed directly):

Shell4 lines
--tensor-model-parallel-size 2
--pipeline-model-parallel-size 1
--num-layers 32
--hidden-size 4096

2. SGLang Arguments (prefixed with --sglang-):

Shell3 lines
--sglang-mem-fraction-static 0.8
--sglang-context-length 8192
--sglang-log-level INFO

3. slime Arguments:

Shell21 lines
# Resource allocation
--actor-num-nodes 1
--actor-num-gpus-per-node 8
--rollout-num-gpus 8
--colocate  # Share GPUs between training/inference

# Data
--prompt-data /path/to/data.jsonl
--input-key prompt
--label-key label

# Training loop
--num-rollout 3000
--rollout-batch-size 32
--n-samples-per-prompt 8
--global-batch-size 256

# Algorithm
--advantage-estimator grpo  # or: gspo, ppo, reinforce_plus_plus
--use-kl-loss
--kl-loss-coef 0.001

Key Constraints

Text1 line
rollout_batch_size × n_samples_per_prompt = global_batch_size × num_steps_per_rollout

Example: 32 × 8 = 256 × 1

---

Data Buffer System

Explains the idea itself. Read it slowly; the later sections build on it.

slime's data buffer enables flexible data management:

Basic Data Source

Python8 lines
class RolloutDataSource:
    def get_samples(self, num_samples):
        """Fetch prompts from dataset."""
        return self.dataset.sample(num_samples)

    def add_samples(self, samples):
        """Called after generation (no-op by default)."""
        pass

Buffered Data Source (Off-Policy)

Python11 lines
class RolloutDataSourceWithBuffer(RolloutDataSource):
    def __init__(self):
        self.buffer = []

    def add_samples(self, samples):
        """Store generated samples for reuse."""
        self.buffer.extend(samples)

    def buffer_filter(self, args, buffer, num_samples):
        """Custom selection logic (prioritized, stratified, etc.)."""
        return select_best(buffer, num_samples)

---

Common Issues and Solutions

Explains the idea itself. Read it slowly; the later sections build on it.

Issue: SGLang Engine Crash

Symptoms: Inference engine dies mid-training

Solutions:

Shell8 lines
# Enable fault tolerance
--use-fault-tolerance

# Increase memory allocation
--sglang-mem-fraction-static 0.85

# Reduce batch size
--rollout-batch-size 16

Issue: Weight Sync Timeout

Symptoms: Training hangs after rollout

Solutions:

Shell5 lines
# Increase sync interval
--update-weights-interval 5

# Use colocated mode (no network transfer)
--colocate

Issue: OOM During Training

Symptoms: CUDA OOM in backward pass

Solutions:

Shell8 lines
# Enable gradient checkpointing
--recompute-activations

# Reduce micro-batch size
--micro-batch-size 1

# Enable sequence parallelism
--sequence-parallel

Issue: Slow Data Loading

Symptoms: GPU idle during data fetch

Solutions:

Shell5 lines
# Increase data workers
--num-data-workers 4

# Use streaming dataset
--streaming-data

---

Supported Models

A lookup table. Do not read it all; find the row that applies to you.

Model FamilyConfigurations
GLMGLM-4.5, GLM-4.6, GLM-4.7, GLM-Z1-9B
QwenQwen3 (4B, 8B, 30B-A3B), Qwen3-MoE, Qwen2.5
DeepSeekV3, V3.1, R1
LlamaLlama 3 (8B, 70B)
OthersKimi K2, Moonlight-16B

Each model has pre-configured scripts in scripts/models/.

---

Advanced Topics

Settings you configure once. Change one at a time so you can see what each does. Set MODEL_ARGS in your environment, not in the chat.

Co-location Mode

Share GPUs between training and inference to reduce memory:

Shell5 lines
python train.py \
    --colocate \
    --actor-num-gpus-per-node 8 \
    --sglang-mem-fraction-static 0.4 \
    ${MODEL_ARGS[@]}

Custom Reward Model

Python9 lines
# custom_rm.py
class CustomRewardModel:
    def __init__(self, model_path):
        self.model = load_model(model_path)

    def compute_reward(self, prompts, responses):
        inputs = self.tokenize(prompts, responses)
        scores = self.model(inputs)
        return scores.tolist()
Shell1 line
--custom-rm-path custom_rm.py

Evaluation Multi-Task

Shell3 lines
--eval-prompt-data aime /path/to/aime.jsonl \
--eval-prompt-data gsm8k /path/to/gsm8k.jsonl \
--n-samples-per-eval-prompt 16

---

Resources

Explains the idea itself. Read it slowly; the later sections build on it.

  • Documentation: https://thudm.github.io/slime/
  • GitHub: https://github.com/THUDM/slime
  • Blog: https://lmsys.org/blog/2025-07-09-slime/
  • Examples: See examples/ directory for 14+ worked examples