Directory → SKILL
SKILLoptionalHermes Optional Skills

Trl Fine Tuning

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF

Post-TrainingTRLReinforcement LearningFine-TuningSFTDPOGRPORLOO
Last registry verification2026-08-18v1.0.1Orchestra Research
Plain meaning

What does it add to Hermes?

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF

Trl Fine Tuning is a skill related to extending the agent. It adds a capability or workflow to Hermes. The publisher description explains the intent, while granted permissions determine what it can actually do.

This plain-language explanation is based on the publisher description. The original text remains visible for verification.

Use it when

Use it when your goal in extending the agent is clear and you can limit it to the data and actions it actually needs.

Skip it when

Do not add it merely to experiment when Hermes already has a simpler path, or when you cannot review its source and permissions.

Who is it for?

Best for users who want a repeatable way of working inside Hermes.

Safe first test

Start with non-sensitive data and a small task whose result can be verified and reversed.

Original publisher description

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF

✓
Data source

This entry was indexed from Hermes Optional Skills. Our explanation interprets the type and domain without inventing a capability not present upstream.

!
Security review

The source is official or editorially reviewed, but you still need to review permissions and version compatibility.

Safe setup path

Inspect, install, then test.

  1. 01
    Open the source

    Match the publisher, license, and description to your need. Check the real update history.

  2. 02
    Review permissions and secrets

    Never paste a secret value into this site. Use environment-variable names and grant the smallest scope.

  3. 03
    Copy setup only after review

    The controls below copy text. They do not execute commands on your device.

  4. 04
    Test with a non-sensitive task

    Inspect the visible tools, then exclude write or delete tools you do not need.

Install command

Review the command, then copy it.

hermes skills install trl-fine-tuning

Hermes Belarabi does not execute this command. Installation happens on your device and remains subject to Hermes scanning and your review.

The full skill definition

Exactly what Hermes loads when this skill runs.

Reproduced from the official documentation. Read it before enabling the skill: this text becomes the agent's instructions.

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

Skill metadata

A lookup table. Do not read it all; find the row that applies to you.

SourceOptional — install with hermes skills install official/mlops/trl-fine-tuning
Pathoptional-skills/mlops/training/trl-fine-tuning
Version1.0.1
AuthorOrchestra Research
LicenseMIT
Dependenciestrl, transformers, datasets, peft, accelerate, torch
Platformslinux, macos, windows
TagsPost-Training, TRL, Reinforcement Learning, Fine-Tuning, SFT, DPO, GRPO, RLOO, RLHF, Preference Alignment, HuggingFace

Reference: full SKILL.md

Explains the idea itself. Read it slowly; the later sections build on it.

Quick start

Ordered, practical steps. Run one and confirm it worked before moving on.

TRL provides post-training methods for aligning language models with human preferences.

Installation:

Shell1 line
pip install trl transformers datasets peft accelerate

Supervised Fine-Tuning (instruction tuning):

Python7 lines
from trl import SFTTrainer

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-0.5B",
    train_dataset=dataset,  # Prompt-completion pairs
)
trainer.train()

DPO (align with preferences):

Python10 lines
from trl import DPOTrainer, DPOConfig

config = DPOConfig(output_dir="model-dpo", beta=0.1)
trainer = DPOTrainer(
    model=model,
    args=config,
    train_dataset=preference_dataset,  # chosen/rejected pairs
    processing_class=tokenizer
)
trainer.train()

Common workflows

Settings you configure once. Change one at a time so you can see what each does.

Workflow 1: Full RLHF pipeline (SFT → Reward Model → RLOO)

Complete pipeline from base model to human-aligned model.

Note (TRL 1.x): PPO has been removed from TRL — PPOTrainer, PPOConfig, and python -m trl.scripts.ppo no longer exist. Use an online-RL trainer TRL still ships: RLOO (RLOOTrainer / trl rloo) is the closest drop-in for a reward-model-driven RLHF pipeline, and GRPO (GRPOTrainer / trl grpo, see Workflow 3) is the memory-efficient alternative. The step below uses RLOO.

Copy this checklist:

Text5 lines
RLHF Training:
- [ ] Step 1: Supervised fine-tuning (SFT)
- [ ] Step 2: Train reward model
- [ ] Step 3: RLOO reinforcement learning
- [ ] Step 4: Evaluate aligned model

Step 1: Supervised fine-tuning

Train base model on instruction-following data:

Python30 lines
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import SFTTrainer, SFTConfig
from datasets import load_dataset

# Load model
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B")

# Load instruction dataset
dataset = load_dataset("trl-lib/Capybara", split="train")

# Configure training
training_args = SFTConfig(
    output_dir="Qwen2.5-0.5B-SFT",
    per_device_train_batch_size=4,
    num_train_epochs=1,
    learning_rate=2e-5,
    logging_steps=10,
    save_strategy="epoch"
)

# Train
trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset,
    processing_class=tokenizer
)
trainer.train()
trainer.save_model()

Step 2: Train reward model

Train model to predict human preferences:

Python30 lines
from transformers import AutoModelForSequenceClassification
from trl import RewardTrainer, RewardConfig

# Load SFT model as base
model = AutoModelForSequenceClassification.from_pretrained(
    "Qwen2.5-0.5B-SFT",
    num_labels=1  # Single reward score
)
tokenizer = AutoTokenizer.from_pretrained("Qwen2.5-0.5B-SFT")

# Load preference data (chosen/rejected pairs)
dataset = load_dataset("trl-lib/ultrafeedback_binarized", split="train")

# Configure training
training_args = RewardConfig(
    output_dir="Qwen2.5-0.5B-Reward",
    per_device_train_batch_size=2,
    num_train_epochs=1,
    learning_rate=1e-5
)

# Train reward model
trainer = RewardTrainer(
    model=model,
    args=training_args,
    processing_class=tokenizer,
    train_dataset=dataset
)
trainer.train()
trainer.save_model()

Step 3: RLOO reinforcement learning

Optimize policy using the reward model. PPO was removed in TRL 1.x; use the RLOO CLI (trl rloo) with the trained reward model passed via --reward_model_name_or_path:

Shell8 lines
trl rloo \
    --model_name_or_path Qwen2.5-0.5B-SFT \
    --reward_model_name_or_path Qwen2.5-0.5B-Reward \
    --dataset_name trl-internal-testing/descriptiveness-sentiment-trl-style \
    --output_dir Qwen2.5-0.5B-RLOO \
    --learning_rate 3e-6 \
    --per_device_train_batch_size 64 \
    --num_generations 4

Equivalent Python (RLOOTrainer / RLOOConfig):

Python22 lines
from trl import RLOOTrainer, RLOOConfig
from transformers import AutoModelForSequenceClassification, AutoTokenizer

reward_model = AutoModelForSequenceClassification.from_pretrained(
    "Qwen2.5-0.5B-Reward", num_labels=1
)

config = RLOOConfig(
    output_dir="Qwen2.5-0.5B-RLOO",
    per_device_train_batch_size=64,
    learning_rate=3e-6,
    num_generations=4,
)

trainer = RLOOTrainer(
    model="Qwen2.5-0.5B-SFT",
    reward_funcs=reward_model,   # a reward model (or a callable reward function)
    args=config,
    train_dataset=dataset,       # prompt-only dataset
    processing_class=tokenizer,
)
trainer.train()

Step 4: Evaluate

Python9 lines
from transformers import pipeline

# Load aligned model
generator = pipeline("text-generation", model="Qwen2.5-0.5B-RLOO")

# Test
prompt = "Explain quantum computing to a 10-year-old"
output = generator(prompt, max_length=200)[0]["generated_text"]
print(output)

Workflow 2: Simple preference alignment with DPO

Align model with preferences without reward model.

Copy this checklist:

Text5 lines
DPO Training:
- [ ] Step 1: Prepare preference dataset
- [ ] Step 2: Configure DPO
- [ ] Step 3: Train with DPOTrainer
- [ ] Step 4: Evaluate alignment

Step 1: Prepare preference dataset

Dataset format:

JSON5 lines
{
  "prompt": "What is the capital of France?",
  "chosen": "The capital of France is Paris.",
  "rejected": "I don't know."
}

Load dataset:

Python5 lines
from datasets import load_dataset

dataset = load_dataset("trl-lib/ultrafeedback_binarized", split="train")
# Or load your own
# dataset = load_dataset("json", data_files="preferences.json")

Step 2: Configure DPO

Python12 lines
from trl import DPOConfig

config = DPOConfig(
    output_dir="Qwen2.5-0.5B-DPO",
    per_device_train_batch_size=4,
    num_train_epochs=1,
    learning_rate=5e-7,
    beta=0.1,  # KL penalty strength
    max_prompt_length=512,
    max_length=1024,
    logging_steps=10
)

Step 3: Train with DPOTrainer

Python15 lines
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import DPOTrainer

model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")

trainer = DPOTrainer(
    model=model,
    args=config,
    train_dataset=dataset,
    processing_class=tokenizer
)

trainer.train()
trainer.save_model()

CLI alternative:

Shell7 lines
trl dpo \
    --model_name_or_path Qwen/Qwen2.5-0.5B-Instruct \
    --dataset_name argilla/Capybara-Preferences \
    --output_dir Qwen2.5-0.5B-DPO \
    --per_device_train_batch_size 4 \
    --learning_rate 5e-7 \
    --beta 0.1

Workflow 3: Memory-efficient online RL with GRPO

Train with reinforcement learning using minimal memory.

For in-depth GRPO guidance — reward function design, critical training insights (loss behavior, mode collapse, tuning), and advanced multi-stage patterns — see references/grpo-training.md ↗. A production-ready training script is in templates/basic_grpo_training.py ↗.

Copy this checklist:

Text4 lines
GRPO Training:
- [ ] Step 1: Define reward function
- [ ] Step 2: Configure GRPO
- [ ] Step 3: Train with GRPOTrainer

Step 1: Define reward function

Python17 lines
def reward_function(completions, **kwargs):
    """
    Compute rewards for completions.

    Args:
        completions: List of generated texts

    Returns:
        List of reward scores (floats)
    """
    rewards = []
    for completion in completions:
        # Example: reward based on length and unique words
        score = len(completion.split())  # Favor longer responses
        score += len(set(completion.lower().split()))  # Reward unique words
        rewards.append(score)
    return rewards

Or use a reward model:

Python10 lines
from transformers import pipeline

reward_model = pipeline("text-classification", model="reward-model-path")

def reward_from_model(completions, prompts, **kwargs):
    # Combine prompt + completion
    full_texts = [p + c for p, c in zip(prompts, completions)]
    # Get reward scores
    results = reward_model(full_texts)
    return [r["score"] for r in results]

Step 2: Configure GRPO

Python10 lines
from trl import GRPOConfig

config = GRPOConfig(
    output_dir="Qwen2-GRPO",
    per_device_train_batch_size=4,
    num_train_epochs=1,
    learning_rate=1e-5,
    num_generations=4,  # Generate 4 completions per prompt
    max_new_tokens=128
)

Step 3: Train with GRPOTrainer

Python14 lines
from datasets import load_dataset
from trl import GRPOTrainer

# Load prompt-only dataset
dataset = load_dataset("trl-lib/tldr", split="train")

trainer = GRPOTrainer(
    model="Qwen/Qwen2-0.5B-Instruct",
    reward_funcs=reward_function,  # Your reward function
    args=config,
    train_dataset=dataset
)

trainer.train()

CLI:

Shell5 lines
trl grpo \
    --model_name_or_path Qwen/Qwen2-0.5B-Instruct \
    --dataset_name trl-lib/tldr \
    --output_dir Qwen2-GRPO \
    --num_generations 4

When to use vs alternatives

Explains the idea itself. Read it slowly; the later sections build on it.

Use TRL when:

  • Need to align model with human preferences
  • Have preference data (chosen/rejected pairs)
  • Want to use reinforcement learning (RLOO, GRPO)
  • Need reward model training
  • Doing RLHF (full pipeline)

Method selection:

  • SFT: Have prompt-completion pairs, want basic instruction following
  • DPO: Have preferences, want simple alignment (no reward model needed)
  • RLOO: Have a reward model, want online RL (the reward-model-driven RLHF path; PPO was removed in TRL 1.x)
  • GRPO: Memory-constrained, want online RL with reward functions
  • Reward Model: Building RLHF pipeline, need to score generations

Use alternatives instead:

  • HuggingFace Trainer: Basic fine-tuning without RL
  • Axolotl: YAML-based training configuration
  • LitGPT: Educational, minimal fine-tuning
  • Unsloth: Fast LoRA training

Common issues

Explains the idea itself. Read it slowly; the later sections build on it.

Issue: OOM during DPO training

Reduce batch size and sequence length:

Python5 lines
config = DPOConfig(
    per_device_train_batch_size=1,  # Reduce from 4
    max_length=512,  # Reduce from 1024
    gradient_accumulation_steps=8  # Maintain effective batch
)

Or use gradient checkpointing:

Python1 line
model.gradient_checkpointing_enable()

Issue: Poor alignment quality

Tune beta parameter:

Python5 lines
# Higher beta = more conservative (stays closer to reference)
config = DPOConfig(beta=0.5)  # Default 0.1

# Lower beta = more aggressive alignment
config = DPOConfig(beta=0.01)

Issue: Reward model not learning

Check loss type and learning rate:

Python4 lines
config = RewardConfig(
    learning_rate=1e-5,  # Try different LR
    num_train_epochs=3  # Train longer
)

Ensure preference dataset has clear winners:

Python3 lines
# Verify dataset
print(dataset[0])
# Should have clear chosen > rejected

Issue: Online RL (RLOO/GRPO) training unstable

Adjust the KL/beta regularization toward the reference policy:

Python6 lines
from trl import RLOOConfig

config = RLOOConfig(
    beta=0.05,          # KL coefficient toward the reference model (increase for stability)
    num_generations=4,  # more samples per prompt = lower-variance advantage estimates
)

Advanced topics

Explains the idea itself. Read it slowly; the later sections build on it.

SFT training guide: See references/sft-training.md ↗ for dataset formats, chat templates, packing strategies, and multi-GPU training.

DPO variants: See references/dpo-variants.md ↗ for IPO, cDPO, RPO, and other DPO loss functions with recommended hyperparameters.

Reward modeling: See references/reward-modeling.md ↗ for outcome vs process rewards, Bradley-Terry loss, and reward model evaluation.

Online RL methods: See references/online-rl.md ↗ for PPO, GRPO, RLOO, and OnlineDPO with detailed configurations.

GRPO deep dive: See references/grpo-training.md ↗ for expert-level GRPO patterns — reward function design philosophy, training insights (why loss increases, mode collapse detection), hyperparameter tuning, multi-stage training, and troubleshooting. Production-ready template in templates/basic_grpo_training.py ↗.

Hardware requirements

Explains the idea itself. Read it slowly; the later sections build on it.

  • GPU: NVIDIA (CUDA required)
  • VRAM: Depends on model and method
  • SFT 7B: 16GB (with LoRA)
  • DPO 7B: 24GB (stores reference model)
  • RLOO 7B: 40GB (policy + reward model)
  • GRPO 7B: 24GB (more memory efficient)
  • Multi-GPU: Supported via accelerate
  • Mixed precision: BF16 recommended (A100/H100)

Memory optimization:

  • Use LoRA/QLoRA for all methods
  • Enable gradient checkpointing
  • Use smaller batch sizes with gradient accumulation

Resources

Explains the idea itself. Read it slowly; the later sections build on it.

  • Docs: https://huggingface.co/docs/trl/
  • GitHub: https://github.com/huggingface/trl
  • Papers:
  • "Training language models to follow instructions with human feedback" (InstructGPT, 2022)
  • "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (DPO, 2023)
  • "Group Relative Policy Optimization" (GRPO, 2024)
  • Examples: https://github.com/huggingface/trl/tree/main/examples/scripts