How to Get Started with MiniMax H3 in ComfyUI Part 3 | Reference to Video, Prompts, Speed Optimization, and Troubleshooting
Introduction | About this article

This article is a personal study and verification memo created with ChatGPT based on information as of August 11, 2026.
Since MiniMax H3 and ComfyUI are still being updated, the model names, node names, supported features, number of reference materials, recommended settings, speed optimization methods, and licenses mentioned in this article may change in the future.
Also, there is a possibility that this article contains errors or misunderstandings.
When actually installing or using these, be sure to check the latest information from the official MiniMax, official ComfyUI, and Hugging Face sites.

This time: Reference to Video
In the first part, we installed ComfyUI, and in the second part, we proceeded to basic video generation using Text to Video and Image to Video.
This time, I will organize the information focusing on **Reference to Video (R2V)**, which is an even more interesting feature of MiniMax H3.
Furthermore, I will summarize my personal notes on Reference Image, Reference Video, Reference Audio, how to create prompts, speed optimization, and common errors.
What is Reference to Video?
Simply put, Reference to Video is a method of instructing the AI by saying, **"Create a new video using this image, this video, and this audio as references."**
In T2V, you create a video from text. In I2V, you create a video from an image.
In R2V, you provide images, videos, audio, etc., to the AI as reference materials and generate a new video based on that information.

What can you do with R2V?
For example, you can think of it as using Image 1 as a reference for a character's face, Image 2 for their outfit, Video 1 for the character's movement, Video 2 for camera work, and Audio 1 for the character's voice.
In other words, it is possible to provide information such as **"who," "what they are wearing," "how they move," "what kind of camera is used to film them," and "what kind of voice they speak with"** from multiple materials.
This is a significantly different way of using it compared to simple T2V or I2V.
Using it for character consistency
One of the difficult aspects of AI video generation is character consistency.
Phenomena such as the face changing, hairstyle changing, outfit changing, or the character looking like a different person in the middle of the video can occur.
By using a Reference Image, you can provide the AI with reference material for a character's face, hairstyle, clothing, and more during generation.
While it cannot guarantee an identical character, it is expected to be more effective at conveying the desired visual direction than simply describing a woman in text using T2V alone.
Referencing Motion
By using a Reference Video, you can use a person's walking style, running, dancing, gestures, product rotations, or camera movements as a reference.
For example, you could use a 'character image + a video of a real person running' as a reference to create a video of the same character running in a similar way.
You could also consider using a separate video specifically for camera work reference.
Referencing Audio
In environments where Reference Audio is available, you can use audio information such as voice and speaking style for video generation.
If you can reference visuals, motion, and voice from separate sources, it opens up a very interesting production workflow.
However, when using someone else's voice, you must be careful about rights and consent issues.
In R2V, 'what you reference' is important
In R2V, it is not necessarily better to include as much material as possible.
I believe the key is to clearly define which parts of each material you want the AI to reference.I believe.
For example, you can assign roles such as: 'Picture 1 for the person's face and hairstyle,' 'Picture 2 for clothing,' 'Video 1 for the person's movement,' 'Video 2 for camera work,' and 'Audio 1 for the person's voice.'
Defining Roles in Prompts
In addition to loading reference materials, it is easier for the AI to understand if you also clarify the roles within the prompt.
For example, the approach would be: 'Maintain the face and hairstyle of the person in Picture 1. Wear the outfit from Picture 2. Reference Video 1 for the person's movement. Reference Video 2 for camera work. Reference Audio 1 for the person's voice.'
Since the actual tags and formatting may vary depending on the workflow you use, be sure to check the latest official templates for this as well.
How should you think about MiniMax H3 prompts?
From here on, I will organize my thoughts on prompts so that I don't get lost when using H3.
In my case, it seems easiest to think in the order of: "Overall scene → Person → Action → Camera → Time progression → Dialogue → Ambient sound → BGM."
1. Overall scene
First, write down where, who, and what is happening.
For example, content like: "A quiet cafe in Tokyo in the evening. A young woman is sitting by the window. Warm sunset light is streaming into the shop."
First, convey the situation of the entire video to the AI.
2. Person
Next, write down the person's appearance, clothing, expression, posture, etc.
For example: "A Japanese woman in her 20s. Shoulder-length black hair. Wearing a beige cardigan. Calm expression."
When using a Reference Image, it may be more about the mindset of "maintaining the person from the Reference Image" rather than describing the person in detail with text.
3. Action
Write down what the person does.
For example: "The woman lifts a coffee cup. She looks out the window. After a short pause, she slowly turns toward us."
In AI video generation, specifying too many actions at once can easily lead to failures, so simple movements are better for short videos.
4. Camera
In video generation, camera specifications are quite important.
There are fixed camera, pan, tilt, dolly in, dolly out, tracking shot, close-up, medium shot, wide shot, etc.
You can use English camera terminology, or you can write in Japanese, such as "the camera slowly approaches the person."
For example: "The camera captures the woman's profile in a medium shot. Then, it slowly approaches the woman."
5. Time progression
In the case of video, it is easier to understand if you organize what happens over time.
For example, for a 10-second video, you can structure it as: '0-3 seconds: A woman looks out the window,' '3-6 seconds: The camera slowly zooms in,' '6-10 seconds: The woman turns to face the camera and speaks.'
Doing this makes it easier for you to visualize the finished video.
⑥ Dialogue
If you want the character to speak, write down the dialogue.
For example, you can write: 'A woman speaks quietly in Japanese, saying, "Today was a good day, wasn't it?"'
If the dialogue is too long, it might not fit within the video duration or could sound unnatural, so it is better to start with shorter lines.
⑦ Ambient Sound
Write down ambient sounds appropriate for the setting, such as rain, traffic, city noise, wind, birds chirping, or the sound of dishes in a cafe.
By considering not just the visuals but also 'what would be heard in this location,' you can create videos that feel truly like H3.
⑧ BGM
Include BGM if necessary.
Examples include quiet jazz piano, ambient music, soft acoustic music, or electronic music.
By matching the BGM to the atmosphere of the visuals, you can control the overall impression of the video.
Example Prompt for a 10-Second Cafe Video
For example, it would look like this:
'A quiet cafe in Tokyo in the evening. A Japanese woman in her 20s is sitting by the window. Warm sunset light is streaming into the shop.
0-3 seconds: The camera captures the woman's profile in a medium shot. The woman is looking out the window.
3-6 seconds: The camera slowly zooms in on the woman. The woman slowly turns to look at the camera.
6-10 seconds: The woman smiles slightly and says, "Today was a good day, wasn't it?" in natural Japanese.
Quiet jazz piano is playing in the shop. Ambient city sounds can be heard in the distance. There is a small sound of the woman placing her cup on the table.'
It seems easiest to use this format as a base and make adjustments while checking the generation results.
Don't be too greedy with camera movements
For example, if you put a large number of camera specifications into a 10-second video, such as "drone shot → rapid zoom → pan → 360-degree rotation → close-up → overhead view," the footage may become unstable.
I think it is easier to judge the results if you limit yourself to one or two camera movements per video at first.
Start with one person
When there are multiple people, problems such as faces swapping, arms and bodies interfering with each other, and not knowing who is speaking are more likely to occur.
It is easier to isolate the cause of problems if you test with one person first, and then increase it to two or more people once that is stable.
R2V may require a different model
For R2V, you may need to use a different model than for T2V or I2V.
Therefore, if "Model not found" or similar is displayed when you open an R2V template, do not try to run it only with the model you were using for T2V; check the model required by the template.
Since specific model names may change in the future, it is safer to base your decisions on the model names displayed in the actual template and the latest distribution page.
About speed optimization
MiniMax H3 is a fairly heavy model, so generation speed is a concern.
Various methods such as quantized models, Attention optimization, Sage Attention, and VRAM optimization are sometimes used to increase speed.
However, if a beginner tries to speed things up from the start, it may lead to more problems, such as ComfyUI failing to launch, CUDA errors, PyTorch version mismatches, or Custom Nodes not working.
Create one video in a standard environment first
If it were me, I would first generate one video using the official or standard template with the default settings.
Once you have confirmed that "MiniMax H3 is working properly," you can proceed to speed optimization, quantization, or Custom Nodes.
By doing this, if you introduce speed optimizations and it stops working, it will be easier to isolate the cause.
ComfyUI Manager and Custom Nodes
ComfyUI has an extension feature called "Custom Nodes".
Think of them like plugins for Photoshop.
They allow you to add convenient features not found in standard ComfyUI or to support new AI models.
You can add and manage Custom Nodes using tools like ComfyUI Manager.
However, since Custom Nodes are external programs, it is safer to check the author, GitHub, update status, and user feedback before installing them, rather than just adding everything.
Common Trouble 1: Model not found
This is an error indicating that the model cannot be found.
First, check if the model file exists, if it is in the correct folder, if the filename has been changed, and if you have reloaded ComfyUI.
Mistakes in the model storage folder are particularly common.
Common Trouble 2: Nodes turning red
When loading a workflow, some nodes may appear red.
Possible causes include missing required Custom Nodes, an outdated ComfyUI version, missing models, or the workflow using newer features.
First, check the page where the template was distributed or the official ComfyUI information.
Common Trouble 3: Out of Memory
If you see "Out of Memory" or "OOM," it is likely that you are running out of GPU memory.
Solutions include lowering the resolution, shortening the video duration, closing other apps that use the GPU, exiting programs like Photoshop or Premiere Pro, and cleaning up browser tabs.
Since video generation uses more GPU memory than image generation, it is important to start with low resolution.
Common Trouble 4: Generation is abnormally slow
First, check whether you are actually using the GPU.
In versions like the Windows Portable version, there may be separate launch methods for NVIDIA GPUs and for CPUs.
If you are accidentally running it on the CPU, it will take a very long time.
It is also a good idea to check the GPU information displayed on the console screen.
Common Trouble 5: It stopped when I closed the black window
In the portable version, the ComfyUI server may be running in a black command window.
Even if ComfyUI is displayed in your browser, closing the black window may cause ComfyUI to stop.
Be careful not to close it by mistake, thinking it is an error screen.
Common Trouble 6: Model download stops midway
Since the MiniMax H3 model is very large, downloads may fail due to connection drops, insufficient disk space, or insufficient temporary file space.
After downloading, it is safer to check not only that the file exists, but also that it is the expected file size.
The biggest advantage of using ComfyUI
After doing all this, you might think, "Isn't it easier to make videos using a web service?"
In fact, if you are only making one video, a web-based service might be easier in some cases.
However, the interesting thing about ComfyUI is that you can create the production process itself, including video generation.
For example, you can create a single workflow that performs "image generation -> character adjustment -> background generation -> video conversion with MiniMax H3 -> upscaling -> frame interpolation -> color adjustment -> video saving."
This seems like it would be quite convenient for production tasks where you repeat the same steps every time.
Combining with design production
Personally, I would like to try combining ChatGPT, Adobe Firefly, Photoshop, ComfyUI, MiniMax H3, and Premiere Pro.
For example, the flow would be: "Create ideas and prompts with ChatGPT -> generate base images with Firefly or ChatGPT -> refine details in Photoshop -> convert to video with ComfyUI + MiniMax H3 -> adjust subtitles and length in Premiere Pro."
This could be highly applicable to video production in design departments, such as for advertising videos, social media videos, concept movies, and product introductions.
About licenses
When using open models locally, you must also check the model's license.
In particular, if you are using it for commercial purposes, client projects, advertisements, social media publication, redistribution, or integration into services, always check the full text of the latest license.
Just because it is an open-weight model does not mean it is free to use for anything.
Being 'AI-generated' does not mean it is free to use.
Apart from the model license, there are also rights regarding the input materials.
When using photographs of people, brand logos, products, characters, audio sources, videos, or illustrations as references, copyright, trademark rights, portrait rights, and publicity rights may be involved for each.
Therefore, for actual projects, it is necessary to check not only the AI model's terms of use but also both the input materials and the generated results.
Regarding the disclosure of AI generation
For AI-generated content, disclosure that it is AI-generated may be required depending on the service's terms of service, the model's license, the region of use, the publication medium, or internal company guidelines.
Since the rules in this area are highly likely to change in the future, it is best to check the latest information at the stage of actual publication.
For personal use | Basic rules for trying out MiniMax H3
Finally, I will summarize the basic rules I follow when trying out MiniMax H3.
Start with a short video of about 5 seconds, a lower resolution, one person, one movement, and about one type of camera work.
First, check the results and then refine the prompt.
Once the content improves, increase the resolution and duration, and finally adjust the audio and background music.
If necessary, take it to Premiere Pro to finish up with subtitles, editing, color grading, and volume adjustments.
Instead of aiming for a finished video from the start, it is better to test on a small scale and gradually increase the level of completion.
What has been achieved through these three parts
In Part 1, I covered the basics of ComfyUI and installation; in Part 2, the introduction of the MiniMax H3 model, Text to Video, and Image to Video; and in Part 3, I organized Reference to Video, reference images, reference videos, reference audio, prompts, speed optimization, and error countermeasures.
MiniMax H3 is not just an AI that creates videos from text; it is a video generation model with the potential to generate visuals, dialogue, sound effects, ambient sounds, and music by combining text, images, videos, and audio.
And by combining it with ComfyUI, I find it interesting that you can build the actual production workflow using generative AI rather than just one-off AI generation.
Finally
To reiterate, this article is a personal study and verification memo created with ChatGPT based on information as of August 11, 2026.
The content may contain errors, misunderstandings, outdated information, or specifications that may change in the future.
Therefore, if you are actually installing ComfyUI or MiniMax H3 while following this article, please be sure to check the latest information from the official MiniMax, official ComfyUI, and Hugging Face sources.
I also plan to build the environment myself based on this article and will update or correct it as I learn more or find any inaccuracies.
