الدليل ← Skill
SkilloptionalHermes Optional Skills

Stable Diffusion

Stable Diffusion: مهارة ينشرها Orchestra Research. مجالها التصميم والوسائط. يوفّر السجل أمر تثبيت جاهزًا يظهر في هذه الصفحة. الرخصة MIT. الإصدار الموثّق 1.0.0. تُفعَّل بأمر `hermes skills install stable-diffusion` بعد مراجعة ما تطلبه. تعتمد على: diffusers>=0.30.0, transformers>=4.41.0, accelerate>=0.31.0, torch>=2.0.0. تدعم: linux، macos، windows.

Image GenerationStable DiffusionDiffusersText-to-ImageMultimodalComputer VisionOptionalHermes skill
آخر تحقق من السجل2026-08-18v1.0.0Orchestra Research
افتح المصدر الأصلي ↗Read in English
المعنى ببساطة

ماذا يضيف إلى Hermes؟

Stable Diffusion: مهارة ينشرها Orchestra Research. مجالها التصميم والوسائط. يوفّر السجل أمر تثبيت جاهزًا يظهر في هذه الصفحة. الرخصة MIT. الإصدار الموثّق 1.0.0. تُفعَّل بأمر hermes skills install stable-diffusion بعد مراجعة ما تطلبه. تعتمد على: diffusers>=0.30.0, transformers>=4.41.0, accelerate>=0.31.0, torch>=2.0.0. تدعم: linux، macos، windows.

Stable Diffusion هي مهارة مرتبطة بمجال التصميم والوسائط. يضيف أدوات لإنشاء أو قراءة أو تعديل وسائط مثل التصميم والصورة والصوت والفيديو.

هذا تفسير مبسّط مبني على وصف الناشر. أبقينا الوصف الإنجليزي بجانبه حتى تستطيع مقارنة المعنى بالمصدر.

استخدمه عندما

استخدمه عندما يكون هدفك واضحًا في التصميم والوسائط وتستطيع تحديد البيانات والأفعال التي يحتاجها فقط.

لا تحتاجه عندما

لا تضفه لمجرد التجربة إذا كان لديك طريق أبسط داخل Hermes، أو إذا لم تستطع مراجعة المصدر والصلاحيات.

لمن يناسب؟

مناسب لمن يريد طريقة عمل قابلة للتكرار داخل Hermes.

أول اختبار آمن

ابدأ بأصل تجريبي، واطلب نسخة جديدة بدل تعديل الملف الأصلي حتى تتأكد من النتيجة.

الوصف الأصلي من الناشر، من دون ترجمة تغيّر المعنى

Text-to-image generation, inpainting, and img2img

✓
مصدر البيانات

فُهرس هذا الإدخال من Hermes Optional Skills. الشرح العربي يفسّر النوع والمجال ولا يضيف وظيفة غير مذكورة في المصدر.

!
مراجعة الأمان

المصدر رسمي أو خضع لمراجعة تحريرية، لكن ذلك لا يغني عن مراجعة الصلاحيات والإصدار.

مسار تثبيت آمن

افحص، ثبّت، ثم اختبر.

  1. 01
    افتح المصدر

    طابق اسم الناشر والرخصة والوصف مع حاجتك، وراجع آخر تحديث فعلي.

  2. 02
    راجع الصلاحيات والأسرار

    لا تلصق قيمة سر داخل الموقع. استخدم أسماء متغيرات البيئة وامنح أقل نطاق ممكن.

  3. 03
    انسخ الإعداد فقط بعد المراجعة

    الأزرار أدناه تنسخ نصًا إلى الحافظة ولا تشغّل أمرًا على جهازك.

  4. 04
    اختبر بمهمة غير حساسة

    تحقق من الأدوات الظاهرة، ثم استبعد أدوات الكتابة أو الحذف التي لا تحتاجها.

أمر التثبيت

راجع الأمر ثم انسخه.

hermes skills install stable-diffusion

لا ينفّذ Hermes بالعربي هذا الأمر. التثبيت يحدث داخل جهازك ويظل خاضعًا لفحص Hermes ومراجعتك.

تعريف المهارة كاملًا

ما الذي يحمّله Hermes بالضبط عند تشغيل هذه المهارة.

منقول من التوثيق الرسمي. اقرأه قبل تفعيل المهارة، فهذا النص يصبح تعليمات الوكيل نفسه.

Text-to-image generation, inpainting, and img2img.

Skill metadata

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

SourceOptional — install with hermes skills install official/mlops/stable-diffusion
Pathoptional-skills/mlops/stable-diffusion
Version1.0.0
AuthorOrchestra Research
LicenseMIT
Dependenciesdiffusers>=0.30.0, transformers>=4.41.0, accelerate>=0.31.0, torch>=2.0.0
Platformslinux, macos, windows
TagsImage Generation, Stable Diffusion, Diffusers, Text-to-Image, Multimodal, Computer Vision

Reference: full SKILL.md

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Guide to generating images with Stable Diffusion using the HuggingFace Diffusers library.

When to use Stable Diffusion

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Use Stable Diffusion when:

  • Generating images from text descriptions
  • Performing image-to-image translation (style transfer, enhancement)
  • Inpainting (filling in masked regions)
  • Outpainting (extending images beyond boundaries)
  • Creating variations of existing images
  • Building custom image generation workflows

Key features:

  • Text-to-Image: Generate images from natural language prompts
  • Image-to-Image: Transform existing images with text guidance
  • Inpainting: Fill masked regions with context-aware content
  • ControlNet: Add spatial conditioning (edges, poses, depth)
  • LoRA Support: Efficient fine-tuning and style adaptation
  • Multiple Models: SD 1.5, SDXL, SD 3.0, Flux support

Use alternatives instead:

  • DALL-E 3: For API-based generation without GPU
  • Midjourney: For artistic, stylized outputs
  • Imagen: For Google Cloud integration
  • Leonardo.ai: For web-based creative workflows

Quick start

خطوات عملية بالترتيب. نفّذ خطوة وتأكد أنها نجحت قبل الانتقال للتالية. الأوامر هنا: pip install xformers.

Installation

Shellسطران
pip install diffusers transformers accelerate torch
pip install xformers  # Optional: memory-efficient attention

Basic text-to-image

Python18 سطرًا
from diffusers import DiffusionPipeline


# Load pipeline (auto-detects model type)
pipe = DiffusionPipeline.from_pretrained(
    "stable-diffusion-v1-5/stable-diffusion-v1-5",
    torch_dtype=torch.float16
)
pipe.to("cuda")

# Generate image
image = pipe(
    "A serene mountain landscape at sunset, highly detailed",
    num_inference_steps=50,
    guidance_scale=7.5
).images[0]

image.save("output.png")

Using SDXL (higher quality)

Python19 سطرًا
from diffusers import AutoPipelineForText2Image


pipe = AutoPipelineForText2Image.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
    variant="fp16"
)
pipe.to("cuda")

# Enable memory optimization
pipe.enable_model_cpu_offload()

image = pipe(
    prompt="A futuristic city with flying cars, cinematic lighting",
    height=1024,
    width=1024,
    num_inference_steps=30
).images[0]

Architecture overview

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Three-pillar design

Diffusers is built around three core components:

Text6 أسطر
Pipeline (orchestration)
├── Model (neural networks)
│   ├── UNet / Transformer (noise prediction)
│   ├── VAE (latent encoding/decoding)
│   └── Text Encoder (CLIP/T5)
└── Scheduler (denoising algorithm)

Pipeline inference flow

Text7 أسطر
Text Prompt → Text Encoder → Text Embeddings
                                    ↓
Random Noise → [Denoising Loop] ← Scheduler
                      ↓
               Predicted Noise
                      ↓
              VAE Decoder → Final Image

Core concepts

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

Pipelines

Pipelines orchestrate complete workflows:

PipelinePurpose
StableDiffusionPipelineText-to-image (SD 1.x/2.x)
StableDiffusionXLPipelineText-to-image (SDXL)
StableDiffusion3PipelineText-to-image (SD 3.0)
FluxPipelineText-to-image (Flux models)
StableDiffusionImg2ImgPipelineImage-to-image
StableDiffusionInpaintPipelineInpainting

Schedulers

Schedulers control the denoising process:

SchedulerStepsQualityUse Case
EulerDiscreteScheduler20-50GoodDefault choice
EulerAncestralDiscreteScheduler20-50GoodMore variation
DPMSolverMultistepScheduler15-25ExcellentFast, high quality
DDIMScheduler50-100GoodDeterministic
LCMScheduler4-8GoodVery fast
UniPCMultistepScheduler15-25ExcellentFast convergence

Swapping schedulers

Python9 أسطر
from diffusers import DPMSolverMultistepScheduler

# Swap for faster generation
pipe.scheduler = DPMSolverMultistepScheduler.from_config(
    pipe.scheduler.config
)

# Now generate with fewer steps
image = pipe(prompt, num_inference_steps=20).images[0]

Generation parameters

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

Key parameters

ParameterDefaultDescription
promptRequiredText description of desired image
negative_promptNoneWhat to avoid in the image
num_inference_steps50Denoising steps (more = better quality)
guidance_scale7.5Prompt adherence (7-12 typical)
height, width512/1024Output dimensions (multiples of 8)
generatorNoneTorch generator for reproducibility
num_images_per_prompt1Batch size

Reproducible generation

Python9 أسطر


generator = torch.Generator(device="cuda").manual_seed(42)

image = pipe(
    prompt="A cat wearing a top hat",
    generator=generator,
    num_inference_steps=50
).images[0]

Negative prompts

Python5 أسطر
image = pipe(
    prompt="Professional photo of a dog in a garden",
    negative_prompt="blurry, low quality, distorted, ugly, bad anatomy",
    guidance_scale=7.5
).images[0]

Image-to-image

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Transform existing images with text guidance:

Python16 سطرًا
from diffusers import AutoPipelineForImage2Image
from PIL import Image

pipe = AutoPipelineForImage2Image.from_pretrained(
    "stable-diffusion-v1-5/stable-diffusion-v1-5",
    torch_dtype=torch.float16
).to("cuda")

init_image = Image.open("input.jpg").resize((512, 512))

image = pipe(
    prompt="A watercolor painting of the scene",
    image=init_image,
    strength=0.75,  # How much to transform (0-1)
    num_inference_steps=50
).images[0]

Inpainting

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Fill masked regions:

Python17 سطرًا
from diffusers import AutoPipelineForInpainting
from PIL import Image

pipe = AutoPipelineForInpainting.from_pretrained(
    "runwayml/stable-diffusion-inpainting",
    torch_dtype=torch.float16
).to("cuda")

image = Image.open("photo.jpg")
mask = Image.open("mask.png")  # White = inpaint region

result = pipe(
    prompt="A red car parked on the street",
    image=image,
    mask_image=mask,
    num_inference_steps=50
).images[0]

ControlNet

جدول مرجعي. لا تقرأه كله، ابحث عن السطر الذي يخصّك فقط.

Add spatial conditioning for precise control:

Python23 سطرًا
from diffusers import StableDiffusionControlNetPipeline, ControlNetModel


# Load ControlNet for edge conditioning
controlnet = ControlNetModel.from_pretrained(
    "lllyasviel/control_v11p_sd15_canny",
    torch_dtype=torch.float16
)

pipe = StableDiffusionControlNetPipeline.from_pretrained(
    "stable-diffusion-v1-5/stable-diffusion-v1-5",
    controlnet=controlnet,
    torch_dtype=torch.float16
).to("cuda")

# Use Canny edge image as control
control_image = get_canny_image(input_image)

image = pipe(
    prompt="A beautiful house in the style of Van Gogh",
    image=control_image,
    num_inference_steps=30
).images[0]

Available ControlNets

ControlNetInput TypeUse Case
cannyEdge mapsPreserve structure
openposePose skeletonsHuman poses
depthDepth maps3D-aware generation
normalNormal mapsSurface details
mlsdLine segmentsArchitectural lines
scribbleRough sketchesSketch-to-image

LoRA adapters

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Load fine-tuned style adapters:

Python18 سطرًا
from diffusers import DiffusionPipeline

pipe = DiffusionPipeline.from_pretrained(
    "stable-diffusion-v1-5/stable-diffusion-v1-5",
    torch_dtype=torch.float16
).to("cuda")

# Load LoRA weights
pipe.load_lora_weights("path/to/lora", weight_name="style.safetensors")

# Generate with LoRA style
image = pipe("A portrait in the trained style").images[0]

# Adjust LoRA strength
pipe.fuse_lora(lora_scale=0.8)

# Unload LoRA
pipe.unload_lora_weights()

Multiple LoRAs

Python8 أسطر
# Load multiple LoRAs
pipe.load_lora_weights("lora1", adapter_name="style")
pipe.load_lora_weights("lora2", adapter_name="character")

# Set weights for each
pipe.set_adapters(["style", "character"], adapter_weights=[0.7, 0.5])

image = pipe("A portrait").images[0]

Memory optimization

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Enable CPU offloading

Python5 أسطر
# Model CPU offload - moves models to CPU when not in use
pipe.enable_model_cpu_offload()

# Sequential CPU offload - more aggressive, slower
pipe.enable_sequential_cpu_offload()

Attention slicing

Python5 أسطر
# Reduce memory by computing attention in chunks
pipe.enable_attention_slicing()

# Or specific chunk size
pipe.enable_attention_slicing("max")

xFormers memory-efficient attention

Pythonسطران
# Requires xformers package
pipe.enable_xformers_memory_efficient_attention()

VAE slicing for large images

Python3 أسطر
# Decode latents in tiles for large images
pipe.enable_vae_slicing()
pipe.enable_vae_tiling()

Model variants

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Loading different precisions

Python12 سطرًا
# FP16 (recommended for GPU)
pipe = DiffusionPipeline.from_pretrained(
    "model-id",
    torch_dtype=torch.float16,
    variant="fp16"
)

# BF16 (better precision, requires Ampere+ GPU)
pipe = DiffusionPipeline.from_pretrained(
    "model-id",
    torch_dtype=torch.bfloat16
)

Loading specific components

Python11 سطرًا
from diffusers import UNet2DConditionModel, AutoencoderKL

# Load custom VAE
vae = AutoencoderKL.from_pretrained("stabilityai/sd-vae-ft-mse")

# Use with pipeline
pipe = DiffusionPipeline.from_pretrained(
    "stable-diffusion-v1-5/stable-diffusion-v1-5",
    vae=vae,
    torch_dtype=torch.float16
)

Batch generation

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Generate multiple images efficiently:

Python15 سطرًا
# Multiple prompts
prompts = [
    "A cat playing piano",
    "A dog reading a book",
    "A bird painting a picture"
]

images = pipe(prompts, num_inference_steps=30).images

# Multiple images per prompt
images = pipe(
    "A beautiful sunset",
    num_images_per_prompt=4,
    num_inference_steps=30
).images

Common workflows

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Workflow 1: High-quality generation

Python22 سطرًا
from diffusers import StableDiffusionXLPipeline, DPMSolverMultistepScheduler


# 1. Load SDXL with optimizations
pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
    variant="fp16"
)
pipe.to("cuda")
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)
pipe.enable_model_cpu_offload()

# 2. Generate with quality settings
image = pipe(
    prompt="A majestic lion in the savanna, golden hour lighting, 8k, detailed fur",
    negative_prompt="blurry, low quality, cartoon, anime, sketch",
    num_inference_steps=30,
    guidance_scale=7.5,
    height=1024,
    width=1024
).images[0]

Workflow 2: Fast prototyping

Python20 سطرًا
from diffusers import AutoPipelineForText2Image, LCMScheduler


# Use LCM for 4-8 step generation
pipe = AutoPipelineForText2Image.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16
).to("cuda")

# Load LCM LoRA for fast generation
pipe.load_lora_weights("latent-consistency/lcm-lora-sdxl")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
pipe.fuse_lora()

# Generate in ~1 second
image = pipe(
    "A beautiful landscape",
    num_inference_steps=4,
    guidance_scale=1.0
).images[0]

Common issues

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

CUDA out of memory:

Python7 أسطر
# Enable memory optimizations
pipe.enable_model_cpu_offload()
pipe.enable_attention_slicing()
pipe.enable_vae_slicing()

# Or use lower precision
pipe = DiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float16)

Black/noise images:

Python6 أسطر
# Check VAE configuration
# Use safety checker bypass if needed
pipe.safety_checker = None

# Ensure proper dtype consistency
pipe = pipe.to(dtype=torch.float16)

Slow generation:

Python6 أسطر
# Use faster scheduler
from diffusers import DPMSolverMultistepScheduler
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)

# Reduce steps
image = pipe(prompt, num_inference_steps=20).images[0]

References

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

Resources

شرح للفكرة نفسها. اقرأه ببطء، فبقية الأقسام تبني عليه.

  • Documentation: https://huggingface.co/docs/diffusers
  • Repository: https://github.com/huggingface/diffusers
  • Model Hub: https://huggingface.co/models?library=diffusers
  • Discord: https://discord.gg/diffusers