Modal — Serverless GPU cloud for ML jobs and model APIs
Modal — Serverless GPU cloud for ML jobs and model APIs
Start with meaning, then move to detail.
This lesson explains Modal — Serverless GPU cloud for ML jobs and model APIs as part of Hermes internals and extension points. You will learn what it does, when it matters, and the smallest safe test that proves it works.
If you are new, do not memorize names. Focus on three questions: what problem does this solve, what access does it need, and how can you verify the result?
For practice, inspect the first example, identify its effects, run it on test data, and compare the result with the source claim.
For advanced readers, inspect Skill metadata, Reference: full SKILL.md, When to use Modal, then verify failure modes and version compatibility.
Know Python, Git, and basic project structure before changing code.
A clear outcome before you read.
- Understand Modal — Serverless GPU cloud for ML jobs and model APIs without assumed prior knowledge.
- Separate the source description from what still needs testing in your environment.
- Read the first command and identify its inputs and outputs before copying it.
Short definitions before the details.
- Provider
- The service that runs or provides access and authentication to a model.
- Skill
- An instruction bundle that teaches Hermes a repeatable workflow without necessarily adding an external service.
Serverless GPU cloud for ML jobs and model APIs
What does the source say, and in what order?
- 01Skill metadata
Start here to understand the core idea or structure.
- 02Reference: full SKILL.md
Read this after the foundation, then connect it to the previous step.
- 03When to use Modal
Read this after the foundation, then connect it to the previous step.
- 04Quick start
Read this after the foundation, then connect it to the previous step.
- 05Installation
Read this after the foundation, then connect it to the previous step.
- 06Hello World with GPU
Read this after the foundation, then connect it to the previous step.
- 07Basic inference endpoint
Read this after the foundation, then connect it to the previous step.
- 08Core concepts
Read this after the foundation, then connect it to the previous step.
- 09Key components
Read this after the foundation, then connect it to the previous step.
- 10Execution modes
Finish here to verify the result and special cases.
Copy only after you understand the effect.
pip install modal
modal setup # Opens browser for authenticationRun: `modal run hello_gpu.py`
### Basic inference endpoint## Core concepts
### Key components
| Component | Purpose |
|-----------|---------|
| `App` | Container for functions and resources |
| `Function` | Serverless function with compute specs |
| `Cls` | Class-based functions with lifecycle hooks |
| `Image` | Container image definition |
| `Volume` | Persistent storage for models/data |
| `Secret` | Secure credential storage |
### Execution modes
| Command | Description |
|---------|-------------|
| `modal run script.py` | Execute and exit |
| `modal serve script.py` | Development with live reload |
| `modal deploy script.py` | Persistent cloRead the first command and identify its inputs and outputs before copying it.
Match every command to your installed Hermes version, review the files and accounts it can reach, and use non-sensitive data for the first test. If this explanation differs from the source, the official source wins.