Wan2.2 I2V 14B試してみた
遂にWan2.2がリリースされたので、試せるものから試していこうと思っています。
言わずと知れたkijaiさんのワークフローをローカルで動かします。
ライブラリとか分からないところがありながら、動かすことを重視したので、パラメータの調整とかは必要に応じてする必要があると思います。(sageattentionとかtritonとかは使ってません。)
環境
OS:Windows 11
GPU:GeForce RTX 4090
CPU:i9-13900KF
memory:128G
手順(480×832、81フレーム、step数4+12)
モデル
以下のリンクから、「Wan2_2-I2V-A14B-HIGH_fp8_e4m3fn_scaled_KJ.safetensors」と「Wan2_2-I2V-A14B-LOW_fp8_e4m3fn_scaled_KJ.safetensors」をダウンロードし、「ComfyUI/models/diffusion_models/」に格納する。
以下のリンクから、「lightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16.safetensors」をダウンロードし、「ComfyUI/models/loras/」に格納する。
以下のリンクから、「umt5-xxl-enc-bf16.safetensors」をダウンロードし、「ComfyUI/models/text_encoders/」に格納する。
以下のリンクから、「Wan2_1_VAE_bf16.safetensors」をダウンロードし、「ComfyUI/models/vae/」に格納する。
ワークフロー
ワークフローについては以下のリンクから取得できます。「wanvideo2_2_I2V_A14B_example_WIP.json」ただ、そのままだと動かなかったり、tritonをインストールしてなかったりするので、少し修正して動くようにしました。
試したワークフローは以下です。

プロンプト
Anime schoolgirl gracefully swaying her upper body to guitar rhythm, fingers elegantly pressing frets and strumming strings with fluid wrist movements, gentle head tilting and nodding to the beat, shoulders subtly rising and falling with deep breaths, left foot tapping softly on ground keeping time, right leg slightly shifting weight, torso leaning forward during intense passages then relaxing back, facial expressions transitioning from concentration to pure joy, eyes occasionally closing in musical bliss then opening with sparkling delight, soft lip movements as if humming along, hair bouncing gently with each head movement, cardigan sleeves sliding naturally with arm motions, skirt swaying with body rhythm, entire pose flowing seamlessly from tense musical focus to relaxed emotional release画像

結果(480×832、81フレーム、step数4+12)
結果は以下です。
みはるが投稿
— kongo jun (@jun_kongo) July 30, 2025
Wan2.2 14B fp8 scaled 4+12step
解像度:480×832
全体フレーム:81フレーム
生成時間:8分
VRAM:18GB pic.twitter.com/vK0rjhmYFg
生成時間は8分で、VRAMは18.5GBぐらいでした。
got prompt
Using 1053 LoRA weight patches for WanVideo model
timesteps: tensor([999, 988, 975, 959, 941, 918, 888, 851, 799, 727, 615, 421],
device='cuda:0')
sigmas: tensor([1.0000, 0.9887, 0.9756, 0.9600, 0.9412, 0.9180, 0.8889, 0.8510, 0.8000,
0.7272, 0.6153, 0.4210, 0.0000])
image_cond shape: torch.Size([20, 21, 104, 60])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:03<00:00, 11.17it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6709.26MB
Transformer blocks on cuda:0: 6709.26MB
Total memory used by transformer blocks: 13418.52MB
Non-blocking memory transfer: False
----------------------
Seq len: 32760
Sampling 81 frames at 480x832 with 12 steps
100%|██████████████████████████████████████████████████████████████████████████████████| 12/12 [07:14<00:00, 36.20s/it]
Allocated memory: memory=0.135 GB
Max allocated memory: max_memory=14.585 GB
Max reserved memory: max_reserved=16.094 GB
samples out stats: mean 0.21082305908203125 std 1.24538254737854 min -4.645794868469238 max 4.9678544998168945
Prompt executed in 458.76 seconds手順(630×1120、81フレーム、step数4+12)
モデル
「手順(480×832、81フレーム、step数4+12)」と同じ
ワークフロー
試したワークフローは以下です。

プロンプト
「手順(480×832、81フレーム、step数4+12)」と同じ
画像
「手順(480×832、81フレーム、step数4+12)」と同じ
結果(630×1120、81フレーム、step数4+12)
結果は以下です。
みはるが投稿
— kongo jun (@jun_kongo) July 30, 2025
Wan2.2 14B fp8 scaled 4+12step
解像度:630×1120
全体フレーム:81フレーム
生成時間:24分
VRAM:23.5GB pic.twitter.com/thCOc7t34b
生成時間は24分で、VRAMは23.5GBぐらいでした。
got prompt
Unloading all LoRAs
timesteps: tensor([999, 959, 888, 727], device='cuda:0')
sigmas: tensor([1.0000, 0.9600, 0.8889, 0.7272, 0.0000])
image_cond shape: torch.Size([20, 21, 140, 76])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:04<00:00, 8.98it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6709.26MB
Transformer blocks on cuda:0: 6709.26MB
Total memory used by transformer blocks: 13418.52MB
Non-blocking memory transfer: False
----------------------
Seq len: 55860
Sampling 81 frames at 608x1120 with 4 steps
100%|███████████████████████████████████████████████████████████████████████████████████| 4/4 [06:42<00:00, 100.56s/it]
Allocated memory: memory=0.197 GB
Max allocated memory: max_memory=19.800 GB
Max reserved memory: max_reserved=21.562 GB
samples out stats: mean 0.16913951933383942 std 1.3858981132507324 min -3.853821277618408 max 4.130667686462402
Using 1053 LoRA weight patches for WanVideo model
timesteps: tensor([999, 988, 975, 959, 941, 918, 888, 851, 799, 727, 615, 421],
device='cuda:0')
sigmas: tensor([1.0000, 0.9887, 0.9756, 0.9600, 0.9412, 0.9180, 0.8889, 0.8510, 0.8000,
0.7272, 0.6153, 0.4210, 0.0000])
image_cond shape: torch.Size([20, 21, 140, 76])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:05<00:00, 7.39it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6709.26MB
Transformer blocks on cuda:0: 6709.26MB
Total memory used by transformer blocks: 13418.52MB
Non-blocking memory transfer: False
----------------------
Seq len: 55860
Sampling 81 frames at 608x1120 with 12 steps
100%|██████████████████████████████████████████████████████████████████████████████████| 12/12 [16:11<00:00, 80.92s/it]
Allocated memory: memory=0.219 GB
Max allocated memory: max_memory=19.390 GB
Max reserved memory: max_reserved=20.969 GB
samples out stats: mean 0.19180047512054443 std 1.28163480758667 min -4.3833160400390625 max 4.80842399597168
Prompt executed in 00:23:54手順(630×1120、81フレーム、step数20+20)
モデル
「手順(480×832、81フレーム、step数4+12)」と同じ
ワークフロー
試したワークフローは以下です。

プロンプト
「手順(480×832、81フレーム、step数4+12)」と同じ
画像
「手順(480×832、81フレーム、step数4+12)」と同じ
結果(630×1120、81フレーム、step数20+20)
結果は以下です。
みはるが投稿
— kongo jun (@jun_kongo) July 30, 2025
Wan2.2 14B fp8 scaled 20+20step
解像度:630×1120
全体フレーム:81フレーム
生成時間:56分
VRAM:23.5GB pic.twitter.com/abSgNodk4U
生成時間は56分で、VRAMは23.5GBぐらいでした。
got prompt
Using 1053 LoRA weight patches for WanVideo model
timesteps: tensor([999, 993, 986, 978, 969, 959, 949, 936, 923, 907, 888, 867, 842, 811,
774, 727, 666, 585, 470, 296], device='cuda:0')
sigmas: tensor([1.0000, 0.9934, 0.9863, 0.9784, 0.9697, 0.9600, 0.9491, 0.9369, 0.9231,
0.9072, 0.8889, 0.8674, 0.8421, 0.8116, 0.7742, 0.7272, 0.6666, 0.5853,
0.4706, 0.2963, 0.0000])
image_cond shape: torch.Size([20, 21, 140, 76])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:04<00:00, 8.43it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6709.26MB
Transformer blocks on cuda:0: 6709.26MB
Total memory used by transformer blocks: 13418.52MB
Non-blocking memory transfer: False
----------------------
Seq len: 55860
Sampling 81 frames at 608x1120 with 20 steps
100%|██████████████████████████████████████████████████████████████████████████████████| 20/20 [28:09<00:00, 84.46s/it]
Allocated memory: memory=0.204 GB
Max allocated memory: max_memory=19.375 GB
Max reserved memory: max_reserved=21.406 GB
samples out stats: mean 0.48051169514656067 std 1.8571292161941528 min -5.107395172119141 max 5.415038108825684
Using 1053 LoRA weight patches for WanVideo model
timesteps: tensor([999, 993, 986, 978, 969, 959, 949, 936, 923, 907, 888, 867, 842, 811,
774, 727, 666, 585, 470, 296], device='cuda:0')
sigmas: tensor([1.0000, 0.9934, 0.9863, 0.9784, 0.9697, 0.9600, 0.9491, 0.9369, 0.9231,
0.9072, 0.8889, 0.8674, 0.8421, 0.8116, 0.7742, 0.7272, 0.6666, 0.5853,
0.4706, 0.2963, 0.0000])
image_cond shape: torch.Size([20, 21, 140, 76])
Swapping 20 transformer blocks
Initializing block swap: 100%|█████████████████████████████████████████████████████████| 40/40 [00:05<00:00, 7.60it/s]
----------------------
Block swap memory summary:
Transformer blocks on cpu: 6709.26MB
Transformer blocks on cuda:0: 6709.26MB
Total memory used by transformer blocks: 13418.52MB
Non-blocking memory transfer: False
----------------------
Seq len: 55860
Sampling 81 frames at 608x1120 with 20 steps
100%|██████████████████████████████████████████████████████████████████████████████████| 20/20 [26:46<00:00, 80.33s/it]
Allocated memory: memory=0.219 GB
Max allocated memory: max_memory=19.390 GB
Max reserved memory: max_reserved=21.156 GB
samples out stats: mean 0.1612669974565506 std 1.1109400987625122 min -4.131563663482666 max 4.928290367126465
Prompt executed in 00:55:54感想
正直使っているライブラリを理解できてないので、パラメータの調整とかに自信はないです。
動かすことを重視しました。
色々試した結果、step数が多ければ多いほど、動画の性能が上がるという、当然の結果となりました。lightx2vを使ったstep数の削減は、動画の質とのトレードオフですね。
