Slime — RL post-training for LLMs with Megatron and SGLang
Slime — RL post-training for LLMs with Megatron and SGLang
Start with meaning, then move to detail.
This lesson explains Slime — RL post-training for LLMs with Megatron and SGLang as part of Hermes internals and extension points. You will learn what it does, when it matters, and the smallest safe test that proves it works.
If you are new, do not memorize names. Focus on three questions: what problem does this solve, what access does it need, and how can you verify the result?
For practice, inspect the first example, identify its effects, run it on test data, and compare the result with the source claim.
For advanced readers, inspect Skill metadata, Reference: full SKILL.md, When to Use slime, then verify failure modes and version compatibility.
Know Python, Git, and basic project structure before changing code.
A clear outcome before you read.
- Understand Slime — RL post-training for LLMs with Megatron and SGLang without assumed prior knowledge.
- Separate the source description from what still needs testing in your environment.
- Read the first command and identify its inputs and outputs before copying it.
Short definitions before the details.
- Skill
- An instruction bundle that teaches Hermes a repeatable workflow without necessarily adding an external service.
RL post-training for LLMs with Megatron and SGLang
What does the source say, and in what order?
- 01Skill metadata
Start here to understand the core idea or structure.
- 02Reference: full SKILL.md
Read this after the foundation, then connect it to the previous step.
- 03When to Use slime
Read this after the foundation, then connect it to the previous step.
- 04Key Features
Read this after the foundation, then connect it to the previous step.
- 05Architecture Overview
Read this after the foundation, then connect it to the previous step.
- 06Installation
Read this after the foundation, then connect it to the previous step.
- 07From Source
Read this after the foundation, then connect it to the previous step.
- 08Quick Start: GRPO Training
Read this after the foundation, then connect it to the previous step.
- 09Workflow 1: Standard GRPO Training
Read this after the foundation, then connect it to the previous step.
- 10Prerequisites Checklist
Finish here to verify the result and special cases.
Copy only after you understand the effect.
┌─────────────────────────────────────────────────────────┐
│ Data Buffer │
│ - Prompt initialization and management │
│ - Custom data generation and filtering │
│ - Rollout sample storage │
└─────────────┬───────────────────────────┬───────────────┘
│ │
┌─────────────▼───────────┐ ┌─────────────▼───────────────┐
│ Training (Megatron-LM) │ │ Rollout (SGLang + Router) │
│ - Actor model training │ │ - Response generation │
│ - Critic (opti# Recommended: Docker
docker pull slimerl/slime:latest
docker run --rm --gpus all --ipc=host --shm-size=16g \
-it slimerl/slime:latest /bin/bash
# Inside container
cd /root/slime && pip install -e . --no-depsgit clone https://github.com/THUDM/slime.git
cd slime
pip install -r requirements.txt
pip install -e .Read the first command and identify its inputs and outputs before copying it.
Match every command to your installed Hermes version, review the files and accounts it can reach, and use non-sensitive data for the first test. If this explanation differs from the source, the official source wins.