Wan 2.1 1.3B: Fast Local Cinematic Video on Consumer GPUs (8GB-12GB)
Full Resolution
Generation Specs RTX 4070
Sampler euler
Scheduler simple
Steps 25
CFG Scale 6
Seed 10842918520
Target VRAM 12GB
VRAM 8GB - 12GB
GPU Hardware Compatibility
Optimal Ground Truth

Native GPU execution at full throughput with zero memory swapping.

Wan 2.1 1.3B T2V (Flow Matching DiT) Min 8GB VRAM 832×480 RTX 4070 Verified

Wan 2.1 1.3B: Fast Local Cinematic Video on Consumer GPUs (8GB-12GB)

Production-grade ComfyUI workflow for Wan 2.1 1.3B Text-to-Video. Generates smooth 16fps cinematic clips at 832x480 resolution on RTX 3060/4060Ti/4070 without VRAM overflow.

Blueprint Summary RTX 4070 Verified

Reproducible ComfyUI workflow for Wan 2.1 1.3B: Fast Local Cinematic Video on Consumer GPUs (8GB-12GB) using Wan 2.1 1.3B T2V (Flow Matching DiT) at 832×480 resolution. Requires minimum 8GB VRAM with sampler euler and scheduler simple (25 steps). Includes 1-click terminal model sync and canvas JSON graph.

Node Graph Pipeline 11 Total Nodes
Verify in Resolver
01 Load Models
UNETLoader
3 loaders (DiT, CLIP, VAE)
02 Conditioning
CLIPTextEncode
3 prompt encodings
03 KSampler
KSampler
DiT latent denoising
04 Decode & Save
VAEDecode
Latent to pixel space

Execution DAG Topology Interactive Visualizer

Drag to pan · Scroll to zoom · Hover wires
Topology DAG 0 Nodes
MODEL CLIP LATENT VAE IMAGE

Model & Asset Setup 1-Click Script

Run in your ComfyUI root:

curl -fsSL https://decomfy.com/api/scripts/wan-2-1-t2v-1-3b-cinematic.sh | bash

Positive Prompt

A breathtaking high-fashion photograph of an adorably cute and charming 22-year-old Korean girl with a sweet innocent smile and soft wavy dark hair. She has a pleasantly plump, delightfully chubby and curvy soft feminine body with soft rounded curves, full feminine figure, cuddly soft cheeks with cute aegyo-sal, and natural soft healthy fullness, absolutely not skinny. Translucent milky porcelain-white glowing skin with natural rosy flushed pink cheeks and blushing undertones, soft pink glossy lips, healthy radiant blooming complexion. She is wearing a cute pastel pink ruffled bikini, sitting gracefully on the shallow steps of a luxury resort infinity pool during warm golden afternoon light. Fine realistic skin micro-texture, visible subtle pores, delicate water droplets glistening on milky skin, shallow depth of field, 85mm f/1.4 portrait lens, soft creamy background bokeh, authentic realism, 8k resolution, masterpiece.

Negative Prompt

skinny, thin, bony, sharp bones, prominent clavicle, visible ribs, emaciated, anorexic, angular, harsh face, mature, old, ugly, deformed, noisy, blurry, cartoon, anime, 3d, illustration, extra limbs, bad hands, mutated anatomy, bad eyes, low quality

Required Models 3 Models

Disk Space Required: 8.3 GB (3 models · DiT/Base: 2.8 GB · Text Encoder: 5.2 GB · VAE: 254 MB)
models/diffusion_models/ 2.84 GB HuggingFace
wan2.1_t2v_1.3B_bf16.safetensors Official Wan 2.1 1.3B DiT Video Core (Flow Matching 3D Causal Attention, consumer GPU friendly)
models/text_encoders/ 5.2 GB HuggingFace
umt5_xxl_fp8_e4m3fn_scaled.safetensors UMT5-XXL FP8 Scaled Text Encoder (CLIPLoader type: wan)
models/vae/ 254 MB HuggingFace
wan_2.1_vae.safetensors 16-channel 3D Video VAE (eliminates contrast artifacts)

Field Notes RTX 4070 Benchmark

Benchmark: [object Object]

Wan 2.1 1.3B DiT represents the optimal sweet spot between VRAM efficiency and temporal coherence. While 14B models choke on consumer GPUs, the 1.3B core with UMT5-XXL FP8 text encoding renders 5-second cinematic 16fps clips comfortably within 8GB-12GB VRAM. Paired with Wan 2.1 16-channel 3D VAE to prevent latent color banding and contrast inversion.

Frequently Asked Questions FAQ

What GPU and VRAM are required to run Wan 2.1 1.3B: Fast Local Cinematic Video on Consumer GPUs (8GB-12GB)?

This workflow requires a minimum of 8GB VRAM (recommended 12GB VRAM). Tested and verified on NVIDIA GeForce RTX 4070 (12GB VRAM) at 832x480 resolution.

How do I resolve missing custom nodes for this workflow?

You can drop the workflow JSON into our client-side Missing Node Auto-Resolver at https://decomfy.com/resolve/ to detect missing nodes and generate install commands, or run the 1-click terminal setup script provided below.

What hardware and precision are required for Wan 2.1 1.3B Video DiT?

Wan 2.1 1.3B Text-to-Video uses a flow-matching 3D diffusion transformer with UMT5-XXL text encoder. On an RTX 4070 (12GB VRAM), load the 1.3B DiT in BF16 alongside FP8-scaled UMT5 text encoder for fluid 5-second 720p generations.

Related ComfyUI Blueprints

View all 28 workflows
Krea 2 DiT #144303076
DiT Blueprint
Krea 2 DiT #144290923
Photorealism Portrait
Krea 2 DiT #144303051
DiT Blueprint
Action completed