Torchtitan — Pretrain LLMs at scale with PyTorch 4D parallelism
Torchtitan — Pretrain LLMs at scale with PyTorch 4D parallelism
Start with meaning, then move to detail.
This lesson explains Torchtitan — Pretrain LLMs at scale with PyTorch 4D parallelism as part of extending Hermes and connecting external tools. You will learn what it does, when it matters, and the smallest safe test that proves it works.
If you are new, do not memorize names. Focus on three questions: what problem does this solve, what access does it need, and how can you verify the result?
For practice, inspect the first example, identify its effects, run it on test data, and compare the result with the source claim.
For advanced readers, inspect Skill metadata, Reference: full SKILL.md, Quick start, then verify failure modes and version compatibility.
Complete installation and one successful task before adding new capabilities.
A clear outcome before you read.
- Understand Torchtitan — Pretrain LLMs at scale with PyTorch 4D parallelism without assumed prior knowledge.
- Separate the source description from what still needs testing in your environment.
- Read the first command and identify its inputs and outputs before copying it.
Short definitions before the details.
- Provider
- The service that runs or provides access and authentication to a model.
- Skill
- An instruction bundle that teaches Hermes a repeatable workflow without necessarily adding an external service.
Pretrain LLMs at scale with PyTorch 4D parallelism
What does the source say, and in what order?
- 01Skill metadata
Start here to understand the core idea or structure.
- 02Reference: full SKILL.md
Read this after the foundation, then connect it to the previous step.
- 03Quick start
Read this after the foundation, then connect it to the previous step.
- 04Common workflows
Read this after the foundation, then connect it to the previous step.
- 05Workflow 1: Pretrain Llama 3.1 8B on single node
Read this after the foundation, then connect it to the previous step.
- 06Workflow 2: Multi-node training with SLURM
Read this after the foundation, then connect it to the previous step.
- 07Workflow 3: Enable Float8 training for H100s
Read this after the foundation, then connect it to the previous step.
- 08Workflow 4: 4D parallelism for 405B models
Read this after the foundation, then connect it to the previous step.
- 09When to use vs alternatives
Read this after the foundation, then connect it to the previous step.
- 10Common issues
Finish here to verify the result and special cases.
Copy only after you understand the effect.
# From PyPI (stable)
pip install torchtitan
# From source (latest features, requires PyTorch nightly)
git clone https://github.com/pytorch/torchtitan
cd torchtitan
pip install -r requirements.txt# Get HF token from https://huggingface.co/settings/tokens
python scripts/download_hf_assets.py --repo_id meta-llama/Llama-3.1-8B --assets tokenizer --hf_token=...# Configs are selected by name from the Python config registry
# (torchtitan/models/llama3/config_registry.py), not by TOML path
MODULE=llama3 CONFIG=llama3_8b ./run_train.shRead the first command and identify its inputs and outputs before copying it.
Match every command to your installed Hermes version, review the files and accounts it can reach, and use non-sensitive data for the first test. If this explanation differs from the source, the official source wins.