Wan2.2 I2V 14B fp8 scaledのワークフロー
以下の記事では、Wan2.2 I2V 14Bのワークフローを試しました。
(環境構築はこの記事を参考にしてください。)
上記の記事では、とにかく動かすことを意識してましたが、今回はある程度最適化しましたので、試したワークフローなどを共有します。
解像度:480×832、全体フレーム:81フレーム、5分程度、
解像度:720×1280、全体フレーム:81フレーム、15分程度で生成できます。
kijaiさんワークフローとRedditで共有されているワークフローを試しました。
動画でも解説しています。
共通で使うインプット
プロンプト
Anime schoolgirl gracefully swaying her upper body to guitar rhythm, fingers elegantly pressing frets and strumming strings with fluid wrist movements, gentle head tilting and nodding to the beat, shoulders subtly rising and falling with deep breaths, left foot tapping softly on ground keeping time, right leg slightly shifting weight, torso leaning forward during intense passages then relaxing back, facial expressions transitioning from concentration to pure joy, eyes occasionally closing in musical bliss then opening with sparkling delight, soft lip movements as if humming along, hair bouncing gently with each head movement, cardigan sleeves sliding naturally with arm motions, skirt swaying with body rhythm, entire pose flowing seamlessly from tense musical focus to relaxed emotional release画像

kijaiワークフロー(480×832)
ワークフロー

生成結果
みはるが投稿
— kongo jun (@jun_kongo) August 2, 2025
Wan2.2 I2V 14B fp8 scaled step1+4
・解像度:480×832
・全体フレーム:81フレーム
・生成時間:5分
・VRAM:16GB pic.twitter.com/OfxdCMS2Bg
生成時間:5分
VRAM:16GB
got prompt
T5Encoder: 100%|███████████████████████████████████████████████████████████████████████| 24/24 [00:00<00:00, 39.32it/s]
T5Encoder: 100%|██████████████████████████████████████████████████████████████████████| 24/24 [00:00<00:00, 200.10it/s]
C:\Users\XXXXX\comfy-ui\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-WanVideoWrapper\wanvideo\wan_video_vae.py:332: UserWarning: 1Torch was not compiled with flash attention. (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\builder\windows\pytorch\aten\src\ATen\native\transformers\cuda\sdp_utils.cpp:555.)
x = F.scaled_dot_product_attention(
Detected model in_channels: 36
Model type: t2v, num_heads: 40, num_layers: 40
Model variant detected: 14B
model_type FLOW
Using accelerate to load and assign model weights to device...
Loading transformer parameters to cpu: 100%|██████████████████████████████████████| 1095/1095 [00:01<00:00, 647.31it/s]
Using FP8 scaled linear quantization
Moving diffusion model from cuda:0 to cpu
Loading LoRA: lightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16 with strength: 1.0
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Using 1053 LoRA weight patches for WanVideo model
sigmas: tensor([1.0000, 0.0000])
Sampling until step 4, timestep: 999
timesteps: tensor([999], device='cuda:0')
image_cond shape: torch.Size([20, 21, 104, 60])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:07<00:00, 5.65it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6704.63MB
Transformer blocks on cuda:0: 6704.63MB
Total memory used by transformer blocks: 13409.26MB
Non-blocking memory transfer: False
----------------------
Seq len: 32760
Sampling 81 frames at 480x832 with 1 steps
100%|████████████████████████████████████████████████████████████████████████████████████| 1/1 [01:15<00:00, 75.45s/it]
Allocated memory: memory=0.096 GB
Max allocated memory: max_memory=13.021 GB
Max reserved memory: max_reserved=13.812 GB
Detected model in_channels: 36
Model type: t2v, num_heads: 40, num_layers: 40
Model variant detected: 14B
model_type FLOW
Using accelerate to load and assign model weights to device...
Loading transformer parameters to cpu: 100%|██████████████████████████████████████| 1095/1095 [00:01<00:00, 618.18it/s]
Using FP8 scaled linear quantization
Moving diffusion model from cuda:0 to cpu
Loading LoRA: lightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16 with strength: 1.0
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Using 1053 LoRA weight patches for WanVideo model
sigmas: tensor([1.0000, 0.9600, 0.8889, 0.7272, 0.0000])
Skipping first 1 steps, starting from timestep 959
timesteps: tensor([959, 888, 727], device='cuda:0')
image_cond shape: torch.Size([20, 21, 104, 60])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:08<00:00, 4.52it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6704.63MB
Transformer blocks on cuda:0: 6704.63MB
Total memory used by transformer blocks: 13409.26MB
Non-blocking memory transfer: False
----------------------
Seq len: 32760
Sampling 81 frames at 480x832 with 4 steps
100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [01:50<00:00, 36.93s/it]
Allocated memory: memory=0.143 GB
Max allocated memory: max_memory=13.084 GB
Max reserved memory: max_reserved=13.844 GB
Prompt executed in 275.52 secondskijaiワークフロー(720×1280)
ワークフロー

生成結果
みはるが投稿
— kongo jun (@jun_kongo) August 2, 2025
Wan2.2 I2V 14B fp8 scaled step1+5
・解像度:720×1280
・全体フレーム:81フレーム
・生成時間:15分
・VRAM:23.5GB pic.twitter.com/6S1TtwrLW6
生成時間:15分
VRAM:23.5GB
got prompt
T5Encoder: 100%|███████████████████████████████████████████████████████████████████████| 24/24 [00:00<00:00, 38.54it/s]
T5Encoder: 100%|██████████████████████████████████████████████████████████████████████| 24/24 [00:00<00:00, 308.11it/s]
C:\Users\XXXXX\comfy-ui\ComfyUI_windows_portable\ComfyUI\custom_nodes\ComfyUI-WanVideoWrapper\wanvideo\wan_video_vae.py:332: UserWarning: 1Torch was not compiled with flash attention. (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\builder\windows\pytorch\aten\src\ATen\native\transformers\cuda\sdp_utils.cpp:555.)
x = F.scaled_dot_product_attention(
Detected model in_channels: 36
Model type: t2v, num_heads: 40, num_layers: 40
Model variant detected: 14B
model_type FLOW
Using accelerate to load and assign model weights to device...
Loading transformer parameters to cpu: 100%|██████████████████████████████████████| 1095/1095 [00:01<00:00, 680.80it/s]
Using FP8 scaled linear quantization
Moving diffusion model from cuda:0 to cpu
Loading LoRA: lightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16 with strength: 1.0
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Using 1053 LoRA weight patches for WanVideo model
sigmas: tensor([1.0000, 0.0000])
Sampling until step 4, timestep: 999
timesteps: tensor([999], device='cuda:0')
image_cond shape: torch.Size([20, 21, 160, 88])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:07<00:00, 5.46it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6704.63MB
Transformer blocks on cuda:0: 6704.63MB
Total memory used by transformer blocks: 13409.26MB
Non-blocking memory transfer: False
----------------------
Seq len: 73920
Sampling 81 frames at 704x1280 with 1 steps
100%|███████████████████████████████████████████████████████████████████████████████████| 1/1 [04:12<00:00, 252.26s/it]
Allocated memory: memory=0.196 GB
Max allocated memory: max_memory=20.169 GB
Max reserved memory: max_reserved=21.281 GB
Detected model in_channels: 36
Model type: t2v, num_heads: 40, num_layers: 40
Model variant detected: 14B
model_type FLOW
Using accelerate to load and assign model weights to device...
Loading transformer parameters to cpu: 100%|██████████████████████████████████████| 1095/1095 [00:01<00:00, 633.42it/s]
Using FP8 scaled linear quantization
Moving diffusion model from cuda:0 to cpu
Loading LoRA: lightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16 with strength: 1.0
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Using 1053 LoRA weight patches for WanVideo model
sigmas: tensor([1.0000, 0.9697, 0.9231, 0.8421, 0.6666, 0.0000])
Skipping first 1 steps, starting from timestep 969
timesteps: tensor([969, 923, 842, 666], device='cuda:0')
image_cond shape: torch.Size([20, 21, 160, 88])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:08<00:00, 4.80it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6704.63MB
Transformer blocks on cuda:0: 6704.63MB
Total memory used by transformer blocks: 13409.26MB
Non-blocking memory transfer: False
----------------------
Seq len: 73920
Sampling 81 frames at 704x1280 with 5 steps
100%|███████████████████████████████████████████████████████████████████████████████████| 4/4 [08:21<00:00, 125.31s/it]
Allocated memory: memory=0.302 GB
Max allocated memory: max_memory=20.310 GB
Max reserved memory: max_reserved=21.438 GB
Prompt executed in 00:14:44Redditワークフロー(480×832)
ワークフロー

生成結果
みはるが投稿
— kongo jun (@jun_kongo) August 2, 2025
Wan2.2 I2V 14B fp8 scaled step8+8
・解像度:480×832
・全体フレーム:81フレーム
・生成時間:5分
・VRAM:21.5GB pic.twitter.com/IfokIxKhZV
生成時間:5.5分
VRAM:21.5GB
got prompt
Using xformers attention in VAE
Using xformers attention in VAE
VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
Using scaled fp8: fp8 matrix mult: False, scale input: False
CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
Requested to load WanTEModel
loaded completely 21474.8 6419.477203369141 True
C:\Users\XXXXX\comfy-ui\ComfyUI_windows_portable\ComfyUI\comfy\ldm\modules\attention.py:451: UserWarning: 1Torch was not compiled with flash attention. (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\builder\windows\pytorch\aten\src\ATen\native\transformers\cuda\sdp_utils.cpp:555.)
out = torch.nn.functional.scaled_dot_product_attention(q, k, v, attn_mask=mask, dropout_p=0.0, is_causal=False)
Requested to load WanVAE
loaded completely 11212.085292816162 242.02829551696777 True
Using scaled fp8: fp8 matrix mult: False, scale input: False
model weight dtype torch.float8_e4m3fn, manual cast: torch.float16
model_type FLOW
unet missing: ['text_embedding.0.scale_weight', 'text_embedding.2.scale_weight', 'time_embedding.0.scale_weight', 'time_embedding.2.scale_weight', 'time_projection.1.scale_weight', 'head.head.scale_weight']
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Requested to load WAN21
loaded completely 15484.626965667725 13627.902000427246 True
100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [01:42<00:00, 25.63s/it]
Using scaled fp8: fp8 matrix mult: False, scale input: False
model weight dtype torch.float8_e4m3fn, manual cast: torch.float16
model_type FLOW
unet missing: ['text_embedding.0.scale_weight', 'text_embedding.2.scale_weight', 'time_embedding.0.scale_weight', 'time_embedding.2.scale_weight', 'time_projection.1.scale_weight', 'head.head.scale_weight']
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Requested to load WAN21
loaded completely 15478.626965667725 13627.902000427246 True
100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [01:43<00:00, 25.98s/it]
Requested to load WanVAE
loaded completely 3226.7562942504883 242.02829551696777 True
Prompt executed in 323.52 secondsRedditワークフロー(720×1280)
ワークフロー

生成結果
みはるが投稿
— kongo jun (@jun_kongo) August 2, 2025
Wan2.2 I2V 14B fp8 scaled step8+8
・解像度:720×1280
・全体フレーム:81フレーム
・生成時間:15分
・VRAM:23.5GB pic.twitter.com/rIk0VFZSlu
生成時間:15分
VRAM:23.5GB
got prompt
Using xformers attention in VAE
Using xformers attention in VAE
FETCH ComfyRegistry Data [DONE]
[ComfyUI-Manager] default cache updated: https://api.comfy.org/nodes
FETCH DATA from: https://raw.githubusercontent.com/ltdrdata/ComfyUI-Manager/main/custom-node-list.jsonVAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
[DONE]
[ComfyUI-Manager] All startup tasks have been completed.
Using scaled fp8: fp8 matrix mult: False, scale input: False
CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
Requested to load WanTEModel
loaded completely 21474.8 6419.47723369141 True
C:\Users\XXXXX\comfy-ui\ComfyUI_windows_portable\ComfyUI\comfy\ldm\modules\attention.py:451: UserWarning: 1Torch was not compiled with flash attention. (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\builder\windows\pytorch\aten\src\ATen\native\transformers\cuda\sdp_utils.cpp:555.)
out = torch.nn.functional.scaled_dot_product_attention(q, k, v, attn_mask=mask, dropout_p=0.0, is_causal=False)
Requested to load WanVAE
loaded completely 5235.522792816162 242.02829551696777 True
Using scaled fp8: fp8 matrix mult: False, scale input: False
model weight dtype torch.float8_e4m3fn, manual cast: torch.float16
model_type FLOW
unet missing: ['text_embedding.0.scale_weight', 'text_embedding.2.scale_weight', 'time_embedding.0.scale_weight', 'time_embedding.2.scale_weight', 'time_projection.1.scale_weight', 'head.head.scale_weight']
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Requested to load WAN21
loaded partially 6710.994965667724 6706.5484046936035 554
100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [06:19<00:00, 94.89s/it]
Using scaled fp8: fp8 matrix mult: False, scale input: False
model weight dtype torch.float8_e4m3fn, manual cast: torch.float16
model_type FLOW
unet missing: ['text_embedding.0.scale_weight', 'text_embedding.2.scale_weight', 'time_embedding.0.scale_weight', 'time_embedding.2.scale_weight', 'time_projection.1.scale_weight', 'head.head.scale_weight']
lora key not loaded: diffusion_model.blocks.0.cross_attn.k_img.diff_b
...
lora key not loaded: diffusion_model.img_emb.proj.4.diff_b
Requested to load WAN21
loaded partially 6704.994965667724 6704.990783691406 618
100%|████████████████████████████████████████████████████████████████████████████████████| 4/4 [06:22<00:00, 95.60s/it]
Requested to load WanVAE
loaded completely 3178.1967163085938 242.02829551696777 True
Prompt executed in 00:14:53感想
検証の中で大量に生成しましたが、Wan2.2 14Bは非常に質がいいですね。
生成時間もかなり改善されているので、かなり実用的なレベルになってきています。
