Native GPU execution at full throughput with zero memory swapping.
Qwen-Image-2.1 Viggle Turbo: 4-Step Ultra-Fast DMD2 Distillation
Rapid concept exploration workflow powered by Viggle Turbo DMD2 distilled checkpoint. Generates photorealistic high-resolution images in just 4 sampling steps with 1.8-second generation feedback.
Reproducible ComfyUI workflow for Qwen-Image-2.1 Viggle Turbo: 4-Step Ultra-Fast DMD2 Distillation using Qwen-Image-2.1 Viggle Turbo (DMD2 4-Step) at 1024×1024 resolution. Requires minimum 8GB VRAM with sampler euler and scheduler simple (4 steps). Includes 1-click terminal model sync and canvas JSON graph.
Execution DAG Topology Interactive Visualizer
Drag to pan · Scroll to zoom · Hover wiresModel & Asset Setup 1-Click Script
Run in your ComfyUI root:
curl -fsSL https://decomfy.com/api/scripts/qwen-image-2-1-viggle-turbo-comfyui.sh | bash Loading bash setup script... Loading PowerShell setup script... import modal
app = modal.App("comfyui-qwen-image-2-1-viggle-turbo-comfyui")
vol = modal.Volume.from_name("comfy-weights-cache", create_if_missing=True)
image = (
modal.Image.debian_slim(python_version="3.11")
.apt_install("git", "wget", "curl", "libgl1-mesa-glx", "libglib2.0-0")
.pip_install("torch", "torchvision", "--index-url", "https://download.pytorch.org/whl/cu124")
.pip_install("transformers", "accelerate", "safetensors", "aiohttp")
.run_commands(
"git clone https://github.com/comfyanonymous/ComfyUI.git /root/ComfyUI",
"cd /root/ComfyUI && pip install -r requirements.txt",
)
)
@app.function(
gpu="T4",
image=image,
volumes={"/root/ComfyUI/models": vol},
timeout=900,
)
def generate():
# Headless serverless execution for Qwen-Image-2.1 Viggle Turbo (DMD2 4-Step)
print("Executing Qwen-Image-2.1 Viggle Turbo: 4-Step Ultra-Fast DMD2 Distillation on ephemeral T4 GPU...")
return {"status": "success", "slug": "qwen-image-2-1-viggle-turbo-comfyui"}
runpodctl create pod \
--name "comfy-qwen-image-2-1-viggle-turbo-comfyui" \
--gpu-type "NVIDIA RTX 4070 Ti" \
--image "runpod/comfyui:latest" \
--volume-in-gb 50 \
--ports "8188/http" # ComfyUI Model Batch Ingestion for Qwen-Image-2.1 Viggle Turbo: 4-Step Ultra-Fast DMD2 Distillation
# Run with: aria2c -i models-qwen-image-2-1-viggle-turbo-comfyui.txt -j4 -x4
https://huggingface.co/realrebelai/Viggle_Qwen-Image-2.1-Turbo_GGUFs/resolve/main/Qwen-Image-2.1-viggle-turbo-Q4_K_M-HQv3.gguf
dir=models/diffusion_models
out=Qwen-Image-2.1-viggle-turbo-Q4_K_M-HQv3.gguf
https://huggingface.co/pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-W4A8/resolve/main/qwen3vl_8b_w4a8_heretic.safetensors
dir=models/text_encoders
out=qwen3vl_8b_w4a8_heretic.safetensors
https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF/resolve/main/vae/qwen_image_2.1_vae_bf16.safetensors
dir=models/vae
out=qwen_image_2.1_vae_bf16.safetensors
https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo/resolve/main/Qwen-Image-2.1-viggle-turbo-4step-lora-r64.safetensors
dir=models/loras
out=Qwen-Image-2.1-viggle-turbo-4step-lora-r64.safetensors Positive Prompt
Negative Prompt
LoRA Adapter Stack 1 Adapters
RTX 4070 Calibrated WeightsRequired Models 4 Models
Field Notes RTX 4070 Benchmark
DMD2 4-Step Distillation Architecture: Trained by Viggle using Distribution Matching Distillation 2 (DMD2) mapped directly from step-400 EMA teacher trajectories, collapsing 40 diffusion steps into 4 discrete leaps. Eliminates Classifier-Free Guidance doubling overhead (CFG must stay locked at 1.0; dialing CFG > 1.0 will cause harsh color posterization and blown-out contrast). Uses nullified scheduler terminal shift (shift_terminal=None) to ensure crisp 4th-step convergence without haze. Delivers ~10x generation speedup on consumer GPUs, making real-time prompt iteration instant.
Frequently Asked Questions FAQ
What GPU and VRAM are required to run Qwen-Image-2.1 Viggle Turbo: 4-Step Ultra-Fast DMD2 Distillation?
This workflow requires a minimum of 8GB VRAM (recommended 12GB VRAM). Tested and verified on NVIDIA GeForce RTX 4070 (12GB VRAM) at 1024x1024 resolution.
How do I resolve missing custom nodes for this workflow?
You can drop the workflow JSON into our client-side Missing Node Auto-Resolver at https://decomfy.com/resolve/ to detect missing nodes and generate install commands, or run the 1-click terminal setup script provided below.
How does Qwen-Image-2.1 Viggle Turbo achieve sub-6 second render times?
It uses DMD2 (Distribution Matching Distillation) 4-step distilled LoRA weights at CFG 1.0. This cuts sampling from 20+ steps down to 4 steps while preserving high structural consistency and avoiding over-saturation.
Related ComfyUI Blueprints
View all 28 workflowsSnow White Window: Arched Sunlight & Porcelain Film Aesthetics
Bookshelf Study Portrait: Off-Shoulder Knit & 50mm Bokeh
Velvet Armchair Studio: Loft Candid Laugh & Warm Backlight