Directory → SKILL
SKILLoptionalHermes Optional Skills

Lambda Labs

On-demand GPU cloud instances for ML training

InfrastructureGPU CloudTrainingInferenceLambda LabsOptionalHermes skill
Last registry verification2026-08-18v1.0.0Orchestra Research
Plain meaning

What does it add to Hermes?

On-demand GPU cloud instances for ML training

Lambda Labs is a skill related to servers and infrastructure. It helps the agent understand or operate technical resources that can affect cost, availability, and security.

This plain-language explanation is based on the publisher description. The original text remains visible for verification.

Use it when

Use it when your goal in servers and infrastructure is clear and you can limit it to the data and actions it actually needs.

Skip it when

Do not add it merely to experiment when Hermes already has a simpler path, or when you cannot review its source and permissions.

Who is it for?

Best for users who want a repeatable way of working inside Hermes.

Safe first test

Start in a test environment with a limited account and inspect status before creating or deleting anything.

Original publisher description

On-demand GPU cloud instances for ML training

✓
Data source

This entry was indexed from Hermes Optional Skills. Our explanation interprets the type and domain without inventing a capability not present upstream.

!
Security review

The source is official or editorially reviewed, but you still need to review permissions and version compatibility.

Safe setup path

Inspect, install, then test.

  1. 01
    Open the source

    Match the publisher, license, and description to your need. Check the real update history.

  2. 02
    Review permissions and secrets

    Never paste a secret value into this site. Use environment-variable names and grant the smallest scope.

  3. 03
    Copy setup only after review

    The controls below copy text. They do not execute commands on your device.

  4. 04
    Test with a non-sensitive task

    Inspect the visible tools, then exclude write or delete tools you do not need.

Install command

Review the command, then copy it.

hermes skills install lambda-labs

Hermes Belarabi does not execute this command. Installation happens on your device and remains subject to Hermes scanning and your review.

The full skill definition

Exactly what Hermes loads when this skill runs.

Reproduced from the official documentation. Read it before enabling the skill: this text becomes the agent's instructions.

On-demand GPU cloud instances for ML training.

Skill metadata

A lookup table. Do not read it all; find the row that applies to you.

SourceOptional — install with hermes skills install official/mlops/lambda-labs
Pathoptional-skills/mlops/lambda-labs
Version1.0.0
AuthorOrchestra Research
LicenseMIT
Dependencieslambda-cloud-client>=1.0.0
Platformslinux, macos, windows
TagsInfrastructure, GPU Cloud, Training, Inference, Lambda Labs

Reference: full SKILL.md

Explains the idea itself. Read it slowly; the later sections build on it.

Guide to running ML workloads on Lambda Labs GPU cloud with on-demand instances and 1-Click Clusters.

When to use Lambda Labs

Explains the idea itself. Read it slowly; the later sections build on it.

Use Lambda Labs when:

  • Need dedicated GPU instances with full SSH access
  • Running long training jobs (hours to days)
  • Want simple pricing with no egress fees
  • Need persistent storage across sessions
  • Require high-performance multi-node clusters (16-512 GPUs)
  • Want pre-installed ML stack (Lambda Stack with PyTorch, CUDA, NCCL)

Key features:

  • GPU variety: B200, H100, GH200, A100, A10, A6000, V100
  • Lambda Stack: Pre-installed PyTorch, TensorFlow, CUDA, cuDNN, NCCL
  • Persistent filesystems: Keep data across instance restarts
  • 1-Click Clusters: 16-512 GPU Slurm clusters with InfiniBand
  • Simple pricing: Pay-per-minute, no egress fees
  • Global regions: 12+ regions worldwide

Use alternatives instead:

  • Modal: For serverless, auto-scaling workloads
  • SkyPilot: For multi-cloud orchestration and cost optimization
  • RunPod: For cheaper spot instances and serverless endpoints
  • Vast.ai: For GPU marketplace with lowest prices

Quick start

Ordered, practical steps. Run one and confirm it worked before moving on.

Account setup

  1. Create account at https://lambda.ai
  2. Add payment method
  3. Generate API key from dashboard
  4. Add SSH key (required before launching instances)

Launch via console

  1. Go to https://cloud.lambda.ai/instances
  2. Click "Launch instance"
  3. Select GPU type and region
  4. Choose SSH key
  5. Optionally attach filesystem
  6. Launch and wait 3-15 minutes

Connect via SSH

Shell5 lines
# Get instance IP from console
ssh ubuntu@<INSTANCE-IP>

# Or with specific key
ssh -i ~/.ssh/lambda_key ubuntu@<INSTANCE-IP>

GPU instances

A lookup table. Do not read it all; find the row that applies to you.

Available GPUs

GPUVRAMPrice/GPU/hrBest For
B200 SXM6180 GB$4.99Largest models, fastest training
H100 SXM80 GB$2.99-3.29Large model training
H100 PCIe80 GB$2.49Cost-effective H100
GH20096 GB$1.49Single-GPU large models
A100 80GB80 GB$1.79Production training
A100 40GB40 GB$1.29Standard training
A1024 GB$0.75Inference, fine-tuning
A600048 GB$0.80Good VRAM/price ratio
V10016 GB$0.55Budget training

Instance configurations

Text4 lines
8x GPU: Best for distributed training (DDP, FSDP)
4x GPU: Large models, multi-GPU training
2x GPU: Medium workloads
1x GPU: Fine-tuning, inference, development

Launch times

  • Single-GPU: 3-5 minutes
  • Multi-GPU: 10-15 minutes

Lambda Stack

Explains the idea itself. Read it slowly; the later sections build on it.

All instances come with Lambda Stack pre-installed:

Shell10 lines
# Included software
- Ubuntu 22.04 LTS
- NVIDIA drivers (latest)
- CUDA 12.x
- cuDNN 8.x
- NCCL (for multi-GPU)
- PyTorch (latest)
- TensorFlow (latest)
- JAX
- JupyterLab

Verify installation

Shell8 lines
# Check GPU
nvidia-smi

# Check PyTorch
python -c "import torch; print(torch.cuda.is_available())"

# Check CUDA version
nvcc --version

Python API

Settings you configure once. Change one at a time so you can see what each does. Set LAMBDA_API_KEY in your environment, not in the chat.

Installation

Shell1 line
pip install lambda-cloud-client

Authentication

Python8 lines



# Configure with API key
configuration = lambda_cloud_client.Configuration(
    host="https://cloud.lambdalabs.com/api/v1",
    access_token=os.environ["LAMBDA_API_KEY"]
)

List available instances

Python7 lines
with lambda_cloud_client.ApiClient(configuration) as api_client:
    api = lambda_cloud_client.DefaultApi(api_client)

    # Get available instance types
    types = api.instance_types()
    for name, info in types.data.items():
        print(f"{name}: {info.instance_type.description}")

Launch instance

Python13 lines
from lambda_cloud_client.models import LaunchInstanceRequest

request = LaunchInstanceRequest(
    region_name="us-west-1",
    instance_type_name="gpu_1x_h100_sxm5",
    ssh_key_names=["my-ssh-key"],
    file_system_names=["my-filesystem"],  # Optional
    name="training-job"
)

response = api.launch_instance(request)
instance_id = response.data.instance_ids[0]
print(f"Launched: {instance_id}")

List running instances

Python3 lines
instances = api.list_instances()
for instance in instances.data:
    print(f"{instance.name}: {instance.ip} ({instance.status})")

Terminate instance

Python6 lines
from lambda_cloud_client.models import TerminateInstanceRequest

request = TerminateInstanceRequest(
    instance_ids=[instance_id]
)
api.terminate_instance(request)

SSH key management

Python14 lines
from lambda_cloud_client.models import AddSshKeyRequest

# Add SSH key
request = AddSshKeyRequest(
    name="my-key",
    public_key="ssh-rsa AAAA..."
)
api.add_ssh_key(request)

# List keys
keys = api.list_ssh_keys()

# Delete key
api.delete_ssh_key(key_id)

CLI with curl

Settings you configure once. Change one at a time so you can see what each does. Set LAMBDA_API_KEY in your environment, not in the chat.

List instance types

Shell2 lines
curl -u $LAMBDA_API_KEY: \
  https://cloud.lambdalabs.com/api/v1/instance-types | jq

Launch instance

Shell8 lines
curl -u $LAMBDA_API_KEY: \
  -X POST https://cloud.lambdalabs.com/api/v1/instance-operations/launch \
  -H "Content-Type: application/json" \
  -d '{
    "region_name": "us-west-1",
    "instance_type_name": "gpu_1x_h100_sxm5",
    "ssh_key_names": ["my-key"]
  }' | jq

Terminate instance

Shell4 lines
curl -u $LAMBDA_API_KEY: \
  -X POST https://cloud.lambdalabs.com/api/v1/instance-operations/terminate \
  -H "Content-Type: application/json" \
  -d '{"instance_ids": ["<INSTANCE-ID>"]}' | jq

Persistent storage

Settings you configure once. Change one at a time so you can see what each does. Set FILESYSTEM_NAME in your environment, not in the chat.

Filesystems

Filesystems persist data across instance restarts:

Shell5 lines
# Mount location
/lambda/nfs/<FILESYSTEM_NAME>

# Example: save checkpoints
python train.py --checkpoint-dir /lambda/nfs/my-storage/checkpoints

Create filesystem

  1. Go to Storage in Lambda console
  2. Click "Create filesystem"
  3. Select region (must match instance region)
  4. Name and create

Attach to instance

Filesystems must be attached at instance launch time:

  • Via console: Select filesystem when launching
  • Via API: Include file_system_names in launch request

Best practices

Shell10 lines
# Store on filesystem (persists)
/lambda/nfs/storage/
  ├── datasets/
  ├── checkpoints/
  ├── models/
  └── outputs/

# Local SSD (faster, ephemeral)
~/ (instance home)
  └── working/  # Temporary files

SSH configuration

Explains the idea itself. Read it slowly; the later sections build on it.

Add SSH key

Shell5 lines
# Generate key locally
ssh-keygen -t ed25519 -f ~/.ssh/lambda_key

# Add public key to Lambda console
# Or via API

Multiple keys

Shell2 lines
# On instance, add more keys
echo 'ssh-rsa AAAA...' >> ~/.ssh/authorized_keys

Import from GitHub

Shell2 lines
# On instance
ssh-import-id gh:username

SSH tunneling

Shell8 lines
# Forward Jupyter
ssh -L 8888:localhost:8888 ubuntu@<IP>

# Forward TensorBoard
ssh -L 6006:localhost:6006 ubuntu@<IP>

# Multiple ports
ssh -L 8888:localhost:8888 -L 6006:localhost:6006 ubuntu@<IP>

JupyterLab

Explains the idea itself. Read it slowly; the later sections build on it.

Launch from console

  1. Go to Instances page
  2. Click "Launch" in Cloud IDE column
  3. JupyterLab opens in browser

Manual access

Shell6 lines
# On instance
jupyter lab --ip=0.0.0.0 --port=8888

# From local machine with tunnel
ssh -L 8888:localhost:8888 ubuntu@<IP>
# Open http://localhost:8888

Training workflows

Explains the idea itself. Read it slowly; the later sections build on it.

Single-GPU training

Shell12 lines
# SSH to instance
ssh ubuntu@<IP>

# Clone repo
git clone https://github.com/user/project
cd project

# Install dependencies
pip install -r requirements.txt

# Train
python train.py --epochs 100 --checkpoint-dir /lambda/nfs/storage/checkpoints

Multi-GPU training (single node)

Python17 lines
# train_ddp.py


from torch.nn.parallel import DistributedDataParallel as DDP

def main():
    dist.init_process_group("nccl")
    rank = dist.get_rank()
    device = rank % torch.cuda.device_count()

    model = MyModel().to(device)
    model = DDP(model, device_ids=[device])

    # Training loop...

if __name__ == "__main__":
    main()
Shell2 lines
# Launch with torchrun (8 GPUs)
torchrun --nproc_per_node=8 train_ddp.py

Checkpoint to filesystem

Python12 lines


checkpoint_dir = "/lambda/nfs/my-storage/checkpoints"
os.makedirs(checkpoint_dir, exist_ok=True)

# Save checkpoint
torch.save({
    'epoch': epoch,
    'model_state_dict': model.state_dict(),
    'optimizer_state_dict': optimizer.state_dict(),
    'loss': loss,
}, f"{checkpoint_dir}/checkpoint_{epoch}.pt")

1-Click Clusters

Settings you configure once. Change one at a time so you can see what each does. Set MASTER_ADDR in your environment, not in the chat.

Overview

High-performance Slurm clusters with:

  • 16-512 NVIDIA H100 or B200 GPUs
  • NVIDIA Quantum-2 400 Gb/s InfiniBand
  • GPUDirect RDMA at 3200 Gb/s
  • Pre-installed distributed ML stack

Included software

  • Ubuntu 22.04 LTS + Lambda Stack
  • NCCL, Open MPI
  • PyTorch with DDP and FSDP
  • TensorFlow
  • OFED drivers

Storage

  • 24 TB NVMe per compute node (ephemeral)
  • Lambda filesystems for persistent data

Multi-node training

Shell5 lines
# On Slurm cluster
srun --nodes=4 --ntasks-per-node=8 --gpus-per-node=8 \
  torchrun --nnodes=4 --nproc_per_node=8 \
  --rdzv_backend=c10d --rdzv_endpoint=$MASTER_ADDR:29500 \
  train.py

Networking

Explains the idea itself. Read it slowly; the later sections build on it.

Bandwidth

  • Inter-instance (same region): up to 200 Gbps
  • Internet outbound: 20 Gbps max

Firewall

  • Default: Only port 22 (SSH) open
  • Configure additional ports in Lambda console
  • ICMP traffic allowed by default

Private IPs

Shell2 lines
# Find private IP
ip addr show | grep 'inet '

Common workflows

Explains the idea itself. Read it slowly; the later sections build on it.

Workflow 1: Fine-tuning LLM

Shell18 lines
# 1. Launch 8x H100 instance with filesystem

# 2. SSH and setup
ssh ubuntu@<IP>
pip install transformers accelerate peft

# 3. Download model to filesystem
python -c "
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained('meta-llama/Llama-2-7b-hf')
model.save_pretrained('/lambda/nfs/storage/models/llama-2-7b')
"

# 4. Fine-tune with checkpoints on filesystem
accelerate launch --num_processes 8 train.py \
  --model_path /lambda/nfs/storage/models/llama-2-7b \
  --output_dir /lambda/nfs/storage/outputs \
  --checkpoint_dir /lambda/nfs/storage/checkpoints

Workflow 2: Batch inference

Shell7 lines
# 1. Launch A10 instance (cost-effective for inference)

# 2. Run inference
python inference.py \
  --model /lambda/nfs/storage/models/fine-tuned \
  --input /lambda/nfs/storage/data/inputs.jsonl \
  --output /lambda/nfs/storage/data/outputs.jsonl

Cost optimization

A lookup table. Do not read it all; find the row that applies to you.

Choose right GPU

TaskRecommended GPU
LLM fine-tuning (7B)A100 40GB
LLM fine-tuning (70B)8x H100
InferenceA10, A6000
DevelopmentV100, A10
Maximum performanceB200

Reduce costs

  1. Use filesystems: Avoid re-downloading data
  2. Checkpoint frequently: Resume interrupted training
  3. Right-size: Don't over-provision GPUs
  4. Terminate idle: No auto-stop, manually terminate

Monitor usage

  • Dashboard shows real-time GPU utilization
  • API for programmatic monitoring

Common issues

A lookup table. Do not read it all; find the row that applies to you.

IssueSolution
Instance won't launchCheck region availability, try different GPU
SSH connection refusedWait for instance to initialize (3-15 min)
Data lost after terminateUse persistent filesystems
Slow data transferUse filesystem in same region
GPU not detectedReboot instance, check drivers

References

Explains the idea itself. Read it slowly; the later sections build on it.

Resources

Explains the idea itself. Read it slowly; the later sections build on it.

  • Documentation: https://docs.lambda.ai
  • Console: https://cloud.lambda.ai
  • Pricing: https://lambda.ai/instances
  • Support: https://support.lambdalabs.com
  • Blog: https://lambda.ai/blog