見出し画像

MiniMax H3がきましたね でもうちのは遅すぎる

Seedance2.0が出てから、こんなのがいずれはローカルで動くようになるのかなと思っていたら、それが実現するまであっという間でした。


ComfyUIも即日にネイティブサポートしました。テンプレートからワークフローが使えます。




プロンプトを書くのはちょっと難しそうだけど、プロンプトガイドがあるので、これをChatGPTのプロジェクトに置いておけば、うまいことやってくれます。

こっちがT2V、I2V用です。


こっちがR2V用です。


ChatGPTの新規プロジェクトを作って、情報源にこれらのファイルを置いておきます。

プロジェクト設定のプロンプトに、

参照画像を元に、MiniMax H3で動画を生成するためのプロンプトを生成してください。

とでもしておけば、準備完了です。今回はR2Vでストーリーボードからの動画生成をやってみます。




ChatGPT-Image2.0でストーリーボードを作ります。とりあえず5秒の動画をやってみようと思うので、4コマで。


キャラクターの登場シーンを書いてみました。


で、これをさっきのChatGPTのプロジェクトに渡して、

ハニービーの登場シーン。5秒で。

subject_definitions:
<Subject 1> is Honeybee, the female bee-themed combat android shown throughout <Picture 1>, with long blonde twin-tails, vivid violet eyes, black-and-gold segmented mechanical armor, glossy black limbs, luminous amber-gold honeycomb-pattern energy wings, and a long black-and-gold stinger lance. Her visual identity, costume, proportions, hairstyle, facial features, and wing design remain consistent across all shots.
<Subject 2> is Honeybee's Bee Swarm shown in <Picture 1>, consisting of multiple small black-and-gold flying combat drones with glowing violet central optics, moving in coordinated formation around Honeybee.
<Subject 3> is the nighttime futuristic megacity environment shown in <Picture 1>, with dense high-rise buildings, purple and blue neon signage, wet reflective surfaces, deep violet atmospheric haze, and a dark clouded sky illuminated by scattered city light.
<Subject 4> is Honeybee's golden flight-energy effect shown in <Picture 1>, consisting of intense amber-gold streaks, sparks, glowing particles, and luminous trails generated around her wings and lance during high-speed movement and landing.
<Picture 1> is the four-panel storyboard reference for [Shot 1] through [Shot 4], defining the shot order, framing progression, Honeybee's approach trajectory, landing pose, close-up composition, Bee Swarm placement, and the black-violet-gold cinematic visual style.

summary:
[reference generation] The target video is a short cinematic entrance sequence for <Subject 1>, Honeybee, following the four-panel structure of <Picture 1>. Above <Subject 3>, a distant golden light rapidly descends toward the camera, revealing Honeybee flying with <Subject 2>. The camera retreats as she aggressively closes the distance, she impacts the ground in a low combat landing, then rises into an intimidating close-up and confidently announces her arrival.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3], [Shot 4]): fully_preserved - Honeybee's blonde twin-tails, violet eyes, black-and-gold mechanical battle suit, glossy armored limbs, honeycomb energy wings, stinger lance, and confident predatory personality are retained throughout.
<Subject 2> (appears in [Shot 1], [Shot 2], [Shot 3], [Shot 4]): fully_preserved - the small black-and-gold Bee Swarm drones with violet glowing optics remain coordinated around Honeybee during her descent and landing.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3], [Shot 4]): fully_preserved - the dense neon megacity, purple-blue illumination, dark sky, atmospheric haze, and reflective urban surfaces remain visually consistent.
<Subject 4> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the amber-gold flight trails, sparks, wing glow, and impact energy intensify naturally with Honeybee's acceleration and landing.
<Picture 1> ([Shot 1] through [Shot 4] storyboard structure): fully_preserved - the distant aerial establishing view, frontal high-speed approach, low impact landing, and final dominant close-up follow the storyboard's composition and progression.

detailed_description:
The target video uses a high-end cinematic anime action style with glossy mechanical surfaces, dramatic black-and-gold contrast, saturated purple neon, volumetric night haze, intense motion blur, and bright amber energy effects. The sequence rapidly progresses from a vast aerial establishing shot to an aggressive frontal approach, a violent landing, and an intimidating close-up.

[Shot 1] A very wide aerial shot establishes <Subject 3>, a vast futuristic megacity at night beneath a dark, cloud-filled sky. Purple, magenta, and blue neon signs illuminate countless skyscrapers below. <Subject 2>, Honeybee's Bee Swarm drones, dart through the air at different depths. High above the city, a tiny point of amber-gold light suddenly appears and accelerates downward. The camera initially holds the enormous scale of the skyline, then begins a fast backward tracking movement as the approaching light expands. Golden streaks from <Subject 4> cut through the dark sky, revealing the descending silhouette of <Subject 1> at their center. A distant mechanical flight whine rapidly grows louder.

[Shot 2] At 00:01.200, the camera continues retreating at high speed as <Subject 1>, Honeybee, flies directly toward the lens. Her long blonde twin-tails stream violently backward in the airflow. Her violet eyes remain locked on the camera with a confident, predatory expression. Her translucent honeycomb-pattern energy wings flare brightly behind her while <Subject 4> produces branching amber-gold trails and sparks. Honeybee holds her long stinger lance forward and slightly downward, its golden blade glowing intensely. <Subject 2> races alongside her in a loose escort formation, some drones crossing close to the lens to emphasize speed and depth. Strong radial motion blur surrounds the city while Honeybee's face and upper body remain sharply readable.

[Shot 3] At 00:02.700, Honeybee suddenly drops below the camera and slams into a wet rooftop or elevated urban platform. The camera whips downward to follow her. She lands in a deep three-point crouch, one hand braced against the ground and the stinger lance angled diagonally beside her. A sharp golden impact pulse spreads across the surface as fragments, dust, sparks, and droplets burst outward. Her wings flex and remain brilliantly illuminated behind her. The Bee Swarm brakes and stabilizes in the air around the landing zone. Honeybee holds the low pose for a brief beat, then slowly raises her head and looks directly toward the camera.

[Shot 4] At 00:04.000, cut to a tight low-angle close-up of <Subject 1>. Honeybee straightens slightly while towering over the camera. Her luminous amber wings fill the background like a glowing honeycomb halo, framed by blurred purple neon towers and hovering Bee Swarm drones. Loose blonde strands drift across her face as the turbulence settles. Her violet eyes narrow and she gives a confident, mischievous smile. Honeybee (S1) looks directly down into the lens and declares in a youthful female voice with playful arrogance, <d>[Japanese] ハニービー、参上。女王の命により、お前を蜂の巣にしてあげる。</d> The camera performs a subtle slow push toward her eyes as the final words land, while golden particles drift through the frame.

overall_soundscape:
Continuous distant futuristic city ambience, high-altitude wind, rapidly intensifying mechanical flight noise, buzzing Bee Swarm propulsion, sharp energy-wing crackle, lance resonance, a heavy metallic landing impact, scattering debris, and fading golden electrical sparks. The drone buzz becomes quieter during the final close-up so Honeybee's voice remains dominant.

non_diegetic_music:
A fast cinematic electronic-orchestral action score begins with a low synthetic pulse and distant percussion, rapidly adding aggressive drums and rising metallic synth textures during Honeybee's descent. The score peaks at the landing impact, then drops into a tense sustained bass-and-string layer beneath her final declaration.

あっさりプロンプトが完成。


ワークフローを見てみると、参照画像が0番から始まっているけど、プロンプトには、1番からのつもりで書かれているので、1番にストーリーボードの画像を入れて、0番には参照用としてキャラの画像を配置した。



そして、Runを押して、のんびり待つ。



長い。



やっとできた。

[INFO] Prompt executed in 00:15:55

15分55秒。




うちの環境は、RTX4060ti16GBにRAM32GBです。


遅すぎん?3060(12GB)とRAM32GBでも10分かからないって記事を見たような気もするけど。


ちなみに、できた動画はこちら。



さてと、原因究明をしていこう。


ChatGPTに聞いてみた。

ComfyUIでMiniMax H3を動かしたんだけど、5秒の動画生成に15分くらいかかっていて、どうも世間一般より遅いらしい。色々調べてみると、CUDAが12.8なのがネックだという話もある。とりあえず少しでも速度を改善したいんだけど、何からやってみるのがいいかな。

その状況なら、CUDA 12.8をいきなり13.xへ上げるより、まずAttentionと実際のGPU稼働状況を確認するのがよさそうです。調べた限り、CUDA 12.8そのものがMiniMax H3で決定的に遅い、とはまだ言い切れません。むしろH3では SageAttentionが効いているかどうかでかなり差が出ています。

ふーん、そうなんだ。うちの環境はWindowsで動かしているときは、その辺りもいろいろ更新したけど、今のLinuxの環境ではあまり触ってないから一からやっていかなあかんか。

ChatGPTは、

現在の環境を正確に確認する。 ComfyUIの起動ログで torch、CUDA、xformers、SageAttention、Triton のバージョンを確認します。特に「CUDA Toolkit 12.8」より、ComfyUIのPythonから見える torch.version.cuda が重要です。Linuxなら次で一発です。


cd ~/ComfyUI source venv/bin/activate
python -c "import torch; print('Torch:',torch.__version__); print('CUDA:',torch.version.cuda); print('GPU:',torch.cuda.get_device_name(0)); print('Capability:',torch.cuda.get_device_capability(0))"

と言うので、やってみると、

Torch: 2.10.0+cu128
CUDA: 12.8
GPU: NVIDIA GeForce RTX 4060 Ti
Capability: (8, 9)

となっている。

次は TritonとSageAttentionが入っているかを確認しましょう。仮想環境に入った状態で、この2つを実行してください。

python -c "import triton; print('Triton:', triton.__version__)"

python -c "import sageattention; print('SageAttention:', sageattention.__file__)"

やってみると、

Triton: 3.6.0
ModuleNotFoundError: No module named 'sageattention'

SageAttentionが入っていない模様。

これでかなり絞れました。Triton 3.6.0 は入っているが、SageAttentionは未導入です。CUDA 12.8自体を触る前に、ここを改善するのが第一候補です。SageAttention公式はAda世代も最適化対象に含めており、SageAttention 2.2.0を利用できます。

そうか。じゃあSageAttentionを入れるか。

SageAttention 2.2はCUDA拡張を含むため、単純な pip install sageattention で必ず最適な2.2が入るとは限りません。公式リポジトリからのビルドが基本になります。

その前に一つだけ確認しましょう。
nvcc --version
さらに、
nvidia-smi
も実行してください。

ここで知りたいのは、
nvcc --version → CUDA Toolkit の実体 nvidia-smi → NVIDIA Driver と対応可能CUDA torch → 2.10.0+cu128

の3つです。

やってみると、nvccはコマンドなし。

なので、NVIDIA公式のCUDA 12.8 Toolkitをインストールすることにした。

wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb

sudo dpkg -i cuda-keyring_1.1-1_all.deb

sudo apt update

sudo apt install cuda-toolkit-12-8


これで、再度vnccの確認をすると、

$ /usr/local/cuda-12.8/bin/nvcc --version

nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2025 NVIDIA Corporation
Built on Fri_Feb_21_20:23:50_PST_2025
Cuda compilation tools, release 12.8, V12.8.93
Build cuda_12.8.r12.8/compiler.35583870_0

次に、SageAttentionを仮想環境に入れていきます。

source ~/comfy-env/bin/activate
export CUDA_HOME=/usr/local/cuda-12.8
export PATH=$CUDA_HOME/bin:$PATH
export LD_LIBRARY_PATH=$CUDA_HOME/lib64:$LD_LIBRARY_PATH
cd ~
git clone https://github.com/thu-ml/SageAttention.git
cd SageAttention
export EXT_PARALLEL=2
export NVCC_APPEND_FLAGS="--threads 4"
export MAX_JOBS=4
python setup.py install


これで、SageAttentionを有効にしてComfyUIを起動してみよう。

cd ~/comfy/ComfyUI
python main.py --use-sage-attention



これで、同じものを生成してみました。


[INFO] Prompt executed in 00:14:04

初回とはいえ、ほとんど変わってないじゃないか。いや、15分55秒からなら2分弱も短縮したとみるべきか。


2回めをやってみると、

[INFO] Prompt executed in 00:13:04

まぁちょっとは速くなってるけど、世間とかけ離れてない?



うーん。

今日はここまで。



ではまた。

いいなと思ったら応援しよう!