Understanding diffusion model theory and mechanics (forward/reverse process)
Scope: Diffusion model fundamentals, architectures, inference pipelines, and model selection Lines: ~340 Last Updated: 2025-10-18
Activate this skill when:
Forward Process (Noise Addition):
Reverse Process (Denoising):
Key Insight: Training predicts noise, inference removes it step-by-step
Purpose: Controls noise addition/removal over time
Common Schedulers:
Scheduler Selection:
Core Component: Modified UNet backbone for noise prediction
Architecture:
Attention Mechanisms:
Key Innovation: Operate in compressed latent space, not pixel space
Components:
Benefits:
from diffusers import StableDiffusionPipeline
import torch
# Load pipeline
pipe = StableDiffusionPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
torch_dtype=torch.float16,
safety_checker=None, # Optional: disable for speed
)
pipe = pipe.to("cuda")
# Generate image
prompt = "A serene mountain landscape at sunset, highly detailed, 4k"
negative_prompt = "blurry, low quality, distorted, ugly"
image = pipe(
prompt=prompt,
negative_prompt=negative_prompt,
num_inference_steps=50, # More steps = better quality
guidance_scale=7.5, # How strongly to follow prompt
height=512,
width=512,
generator=torch.Generator("cuda").manual_seed(42), # Reproducibility
).images[0]
image.save("output.png")
When to use:
from diffusers import StableDiffusionXLPipeline
import torch
# SDXL has better text understanding and higher resolution
pipe = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
variant="fp16",
use_safetensors=True,
)
pipe = pipe.to("cuda")
# Enable CPU offloading for lower VRAM (slower but uses ~6GB vs 12GB)
# pipe.enable_model_cpu_offload()
prompt = "Astronaut riding a horse on Mars, photorealistic, 8k uhd"
negative_prompt = "cartoon, painting, illustration"
image = pipe(
prompt=prompt,
negative_prompt=negative_prompt,
num_inference_steps=40, # SDXL converges faster
guidance_scale=7.0, # Lower than SD 1.5
height=1024, # Native resolution
width=1024,
).images[0]
image.save("sdxl_output.png")
Benefits:
from diffusers import (
StableDiffusionPipeline,
DDIMScheduler,
DPMSolverMultistepScheduler,
EulerAncestralDiscreteScheduler,
UniPCMultistepScheduler,
)
import torch
pipe = StableDiffusionPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
prompt = "A magical forest with glowing mushrooms"
# Fast and deterministic (recommended default)
pipe.scheduler = UniPCMultistepScheduler.from_config(pipe.scheduler.config)
image_unipc = pipe(prompt, num_inference_steps=20).images[0]
# High quality, slower
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)
image_dpm = pipe(prompt, num_inference_steps=25).images[0]
# Artistic/creative (non-deterministic)
pipe.scheduler = EulerAncestralDiscreteScheduler.from_config(pipe.scheduler.config)
image_euler = pipe(prompt, num_inference_steps=30).images[0]
# Original DDIM (fast, deterministic)
pipe.scheduler = DDIMScheduler.from_config(pipe.scheduler.config)
image_ddim = pipe(prompt, num_inference_steps=50).images[0]
When to use each:
import torch
from diffusers import StableDiffusionPipeline
pipe = StableDiffusionPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
torch_dtype=torch.float16,
).to("cuda")
prompt = "A red apple on a wooden table"
# Generate with different guidance scales
guidance_scales = [1.0, 3.0, 7.5, 12.0, 20.0]
images = []
for scale in guidance_scales:
image = pipe(
prompt=prompt,
guidance_scale=scale,
num_inference_steps=30,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
images.append(image)
image.save(f"guidance_{scale}.png")
# Guidance scale effects:
# 1.0-3.0: More creative, less prompt adherence, diverse
# 7.0-9.0: Balanced (recommended range)
# 10.0-15.0: Strong prompt adherence, less variation
# 15.0+: Over-saturated, artifacts, "burnt" look
Guidance Scale Guidelines:
from diffusers import StableDiffusionPipeline
import torch
pipe = StableDiffusionPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
torch_dtype=torch.float16,
).to("cuda")
prompt = "A futuristic cityscape at night"
# Generate multiple images in one call (efficient)
images = pipe(
prompt=prompt,
num_images_per_prompt=4, # Batch size (watch VRAM)
num_inference_steps=30,
guidance_scale=7.5,
).images
# Save all images
for i, img in enumerate(images):
img.save(f"batch_{i}.png")
VRAM considerations:
from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image
import torch
# Load img2img pipeline
pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
torch_dtype=torch.float16,
).to("cuda")
# Load input image
init_image = Image.open("input.png").convert("RGB")
init_image = init_image.resize((512, 512))
prompt = "A watercolor painting of the same scene"
# Generate variation
image = pipe(
prompt=prompt,
image=init_image,
strength=0.75, # 0.0 = no change, 1.0 = full generation
num_inference_steps=50,
guidance_scale=7.5,
).images[0]
image.save("img2img_output.png")
Strength parameter:
from diffusers import StableDiffusionInpaintPipeline
from PIL import Image
import torch
pipe = StableDiffusionInpaintPipeline.from_pretrained(
"runwayml/stable-diffusion-inpainting",
torch_dtype=torch.float16,
).to("cuda")
# Load image and mask (white = inpaint area)
image = Image.open("input.png").convert("RGB").resize((512, 512))
mask = Image.open("mask.png").convert("RGB").resize((512, 512))
prompt = "A red sports car"
result = pipe(
prompt=prompt,
image=image,
mask_image=mask,
num_inference_steps=50,
guidance_scale=7.5,
).images[0]
result.save("inpainted.png")
Use cases:
Model | Resolution | VRAM | Speed | Quality | Best For
---------------|------------|--------|-------|---------|------------------
SD 1.5 | 512x512 | 4GB | Fast | Good | General, fast iteration
SD 2.1 | 768x768 | 6GB | Mid | Better | Higher res, improved quality
SDXL 1.0 | 1024x1024 | 8-12GB | Slow | Best | Production, high quality
SD 3.0 | 1024x1024 | 10GB+ | Mid | Best | Latest, best text understanding
Parameter | Range | Default | Effect
----------------------|------------|---------|--------------------------------
num_inference_steps | 20-100 | 50 | Quality vs speed (diminishing returns >50)
guidance_scale | 1.0-20.0 | 7.5 | Prompt adherence (7-9 optimal)
height/width | 512-1024 | 512 | Output resolution (multiple of 64)
num_images_per_prompt | 1-8 | 1 | Batch generation (VRAM limited)
strength (img2img) | 0.0-1.0 | 0.8 | How much to change input
✅ DO: Use UniPC or DPM-Solver for best quality/speed balance
✅ DO: Use Euler Ancestral for artistic/creative outputs
✅ DO: Reduce steps with better schedulers (20-30 vs 50)
✅ DO: Match scheduler to use case (deterministic vs creative)
❌ DON'T: Use DDPM unless you need exact original algorithm
❌ DON'T: Use 100+ steps (waste of compute, minimal gain)
❌ DON'T: Ignore scheduler choice (huge impact on results)
✅ DO: Be specific and descriptive
✅ DO: Use quality tags ("highly detailed", "4k", "professional")
✅ DO: Use negative prompts to avoid common issues
✅ DO: Mention style/medium if important ("oil painting", "photograph")
❌ DON'T: Be vague or ambiguous
❌ DON'T: Forget negative prompts (prevents common artifacts)
❌ DON'T: Over-complicate (model has limits on prompt length)
❌ Using default scheduler without consideration: DDPM is slow, better options exist ✅ Use UniPC, DPM-Solver, or Euler Ancestral based on use case
❌ Ignoring guidance scale: Using default 7.5 for everything ✅ Tune guidance scale (7-9 for SD 1.5, 5-7 for SDXL)
❌ Too many inference steps: Using 100+ steps for marginal gains ✅ Use 20-30 steps with good scheduler, 40-50 max for quality
❌ Wrong resolution: Generating 1024x1024 with SD 1.5 ✅ Use native resolution (512 for SD 1.5, 1024 for SDXL)
❌ No negative prompts: Forgetting to specify what to avoid ✅ Always use negative prompts ("blurry, low quality, distorted")
❌ Not setting seed for debugging: Non-reproducible results ✅ Use fixed seed during development/debugging
❌ Loading full precision models: Using float32 on GPU ✅ Use float16/bfloat16 for 2x speed and 50% VRAM reduction
❌ Ignoring safety checker overhead: Keeping it on when not needed ✅ Disable safety checker for faster inference (if appropriate)
stable-diffusion-deployment.md - Production deployment, API setup, optimizationdiffusion-finetuning.md - Fine-tuning with DreamBooth, LoRA, textual inversionmodal-gpu-workloads.md - GPU selection and configuration for diffusion modelsmodal-web-endpoints.md - Creating API endpoints for diffusion inferencelora-peft-techniques.md - LoRA for parameter-efficient fine-tuningLast Updated: 2025-10-18 Format Version: 1.0 (Atomic)