PTENES
Skip to content
Level 2 ~120 min

ComfyUI — Open-Source Ecosystem Hub

Master ComfyUI, the node-based interface that has become the central hub for running open-source image and video models with full control over every step of the pipeline.

1

🔧 What Is ComfyUI?

O ComfyUI is an open-source, node-based graphical interface for running generative AI models. While tools like Midjourney or Runway offer a simplified experience (enter a prompt → get a result), ComfyUI exposes each step of the generation pipeline as a visual node that you can configure, connect, and customize.

Think of it as the "Blender of generative AI" — powerful, flexible, with a learning curve, but with virtually unlimited possibilities.

Why Did ComfyUI Dominate?

Three reasons: (1) it's 100% free and open-source, (2) it supports virtually any open-source model through custom nodes, and (3) the community has created more than 1000 custom node packages that extend its capabilities indefinitely.

2

📦 1000+ Custom Nodes

The custom node ecosystem is what makes ComfyUI truly powerful. Each node package adds new features:

Node Package Functionality
ComfyUI-Wan Integration with Wan 2.5/2.7 for video generation
ComfyUI-FLUX Support for FLUX.2 for photorealistic images
ComfyUI-SkyReels SkyReels V4 pipeline for video with audio
Rodin3D Nodes Generate 3D models from images
ControlNet Nodes Control over pose, depth, edges, and composition
LatentCut Node Cropping and compositing in latent space (without artifacts)
IP-Adapter Nodes Style transfer and character consistency
3

🧩 Subgraphs — Modular Workflows

The Subgraphs are an advanced feature that lets you encapsulate complex workflows in a single reusable node. This makes it possible to:

  • • Modularity — Create “blocks” that can be reused across different projects
  • • Organization — Keep complex workflows readable and manageable
  • • Sharing — Export subgraphs for the community or team
  • • Abstraction — Hide the complexity and expose only the relevant parameters
4

🤖 ComfyUI Copilot (Alibaba)

O ComfyUI Copilot, developed by Alibaba, is an AI assistant integrated with ComfyUI that dramatically speeds up the learning curve:

  • • Generate workflows with prompts — Describe what you want, and Copilot builds the workflow
  • • Node explanation — Hover over any node to get a detailed explanation
  • • Assisted debugging — Copilot identifies and suggests fixes for workflow errors
  • • 10x faster — According to Alibaba, Copilot speeds up learning ComfyUI by 10x

How to Enable Copilot

Install the ComfyUI-Copilot extension via ComfyUI Manager. After restarting, a chat icon appears in the interface. You can ask: "Create a workflow to generate an image with FLUX.2, apply pose ControlNet, and upscale 4x".

5

💾 VRAM optimization

Running large models locally requires careful VRAM management. ComfyUI offers several strategies:

Strategy Command/Configuration VRAM savings
FP16 Intermediates --fp16-intermediates ~30-40%
Low VRAM Mode --lowvram ~50%
CPU Offload --cpu Maximum (slower)
Tiled VAE VAE Decode (Tiled) node ~20% in decode

Hardware recommendation

For image generation with FLUX.2: at least 12GB VRAM (RTX 3060/4060). For video with Wan 2.7: at least 24GB VRAM (RTX 4090 or A100). With --fp16-intermediates, an RTX 3090 (24GB) runs most models comfortably.

6

🛠️ Setting Up ComfyUI Locally

Step 1 — Install prerequisites

Python 3.10+, Git, CUDA toolkit (NVIDIA) or ROCm (AMD). Check with nvidia-smi whether your GPU is recognized.

Step 2 — Clone the repository

git clone https://github.com/comfyanonymous/ComfyUI.git

Step 3 — Install dependencies

pip install -r requirements.txt

Step 4 — Download the FLUX.2 model

Download the FLUX.2-schnell model from Hugging Face and place it in ComfyUI/models/checkpoints/

Step 5 — Start ComfyUI

python main.py --fp16-intermediates

Access http://127.0.0.1:8188 in the browser.

7

🎨 First workflow: Text-to-Image with FLUX.2

Create your first workflow by connecting these nodes in order:

[Load Checkpoint: FLUX.2] → [CLIP Text Encode: your prompt] → [KSampler] → [VAE Decode] → [Save Image]

  • 1. Load Checkpoint — Select the FLUX.2 model you downloaded
  • 2. CLIP Text Encode — Add two: one for the positive prompt, another for the negative prompt
  • 3. KSampler — Configure steps (20), cfg (7.5), sampler (euler), scheduler (normal)
  • 4. VAE Decode — Converts the latent space into a visible image
  • 5. Save Image — Saves the result as a PNG
8

✅ Lesson Checklist