見出し画像

ComfyUIでFP8/GGUF版のHunyuanImage-2.1 Distilledを試す(推奨VRAM 12~16GB)

Last update 9-15-2025
※ 公開されて間もないため、何度か記事の修正等を行う場合があります。
Distilled版ではない低速な通常版の記事もあります。
※ 関連記事は、ComfyUIの記事一覧にまとめてあります。




■ 0. 概要

▼ 0-0. はじめに

 本記事では、Windows上のComfyUIで「HunyuanImage-2.1 Distilled」を利用して、生成を高速化する手順を説明します。モデルはFP8版とGGUF版を利用します。

 執筆時点ではHunyuanImage-2.1 DistilledのLoRAは無く、Modelとして提供されています。

 それぞれの説明は下記の項目にあります。

  • 1. 留意事項と使用リソース量

  • 2.~3. FP8版の準備と生成

  • 4.~5. GGUF版の準備と生成

 HunyuanImage-2.1の概要や、Distilledモデルを利用しない通常の生成の手順は下記の記事を参照してください。

 ComfyUIのインストール方法は、下記の記事を参照してください。


▼ 0-1. 関連リンク

 公式の各種モデル(Model、Text Encoder、VAE)です。本記事では利用しません。

 本記事で利用する、FP8版のModel、Text EncoderとVAEです。

 FP8版のModelはCivitaiにも掲載されています。Civitaiの仕組み上、ファイル名が異なる可能性があります。

 本記事で利用する、GGUF版の各種モデルと拡張機能です。

 下記のGGUF版Modelも利用できると思います(未確認)。



■ 1. 留意事項と使用リソース量

▼ 1-1. 留意事項

 ComfyUIを更新していないと、必要なノードが存在しない場合があります。バグフィックスを含むこともあるので、なるべく最新版にしてください。

 次回の生成を高速化するため、VRAMに読み込んだモデルが自動的にメインRAMへ待避(オフロード)されることがあります。そのため、メインRAMが32GB以上であることが望ましいです。


▼ 1-2. VRAM使用量

 FP8版の場合、VRAM 16GBでは確実に利用できます。VRAM(占有GPUメモリ)の空きが十分にあれば12GBでも可能とみられますが、共有GPUメモリにはみ出すと生成速度がかなり低下します。

 GGUF版の場合、量子化タイプ次第となります。それでも、基本的にはVRAM 12~16GBの環境が適しています。

 GGUF版は、Modelに「hunyuanimage2.1-v2-q3_k_m.gguf」等を利用することで、VRAM 8GBでも共有GPUメモリにはみ出さずに生成できる可能性があります。


▼ 1-3. 生成時間

 筆者の環境(GeForce RTX 5060Ti 16GB)での生成時間は下記のとおりです。ディスクからのモデルの読み込みや、プロンプトの処理の時間は含まれていません。GGUF形式はModelのサイズ(量子化タイプ)によって多少変化します。

  • Distill FP8
    54秒程度(8Steps、euler/simple、CFG=2.0)
    40秒程度(変更:weight_dtype=fp8_e4m3fn_fast)
    21秒程度(変更:weight_dtype=fp8_e4m3fn_fast、CFG=1.0)

  • Distill GGUF Q4_K_M
    57秒程度(8Steps、euler/simple、CFG=2.0)
    30秒程度(変更:CFG=1.0)

 参考まで、通常の生成時では下記のとおりです。

  • FP8
    115秒程度(20Steps、euler/simple、CFG=3.5)
    80秒程度(変更:weight_dtype=fp8_e4m3fn_fast)

  • GGUF Q4_K_M
    135秒程度(20Steps、euler/simple、CFG=3.5)



■ 2. 準備(FP8版)

▼ 2-1. はじめに

 HunyuanImage Distilledと通常版の違いは、Modelと一部の設定のみです。また、Qwen-Imageと同じText Encoderを利用します。


▼ 2-2. モデルの設置

 筆者は種類ごとにディレクトリを分けているのでそのように記述しますが、お好みで決めていただいて構いません。

 下記のURLに掲載されたそれぞれのファイルをダウンロードして、適切なフォルダへ移動します。設置済みの場合はスキップしてください。なお、Qwen-Imageの記事で利用したものと同一のText Encoderが含まれています。

  • drbaph/HunyuanImage-2.1_fp8
    https://huggingface.co/drbaph/HunyuanImage-2.1_fp8 

    • hunyuanimage2.1-distilled_fp8_e4m3fn.safetensors (16.2GB)
      (※ 9-12-2025に再アップロードされたファイルが必要)
      → models\diffusion_models\hunyuan_image\

  • Comfy-Org/HunyuanImage_2.1_ComfyUI
    https://huggingface.co/Comfy-Org/HunyuanImage_2.1_ComfyUI 

    • split_files/text_encoders/
      qwen_2.5_vl_7b_fp8_scaled.safetensors (8.73GB)
      byt5_small_glyphxl_fp16.safetensors (418MB)
      → models\text_encoders\hunyuan_image\

    • split_files/vae/
      hunyuan_image_2.1_vae_fp16.safetensors (773MB)
      → models\vae\hunyuan_image\



■ 3. 生成(FP8版)

▼ 3-1. 概要

 生成用のワークフローを提供しますので、改造、再公開を含め自由にご利用ください。


▼ 3-2. 補足

 Distilledモデルは出力の品質がやや劣るので、パラメーターの調整が必要そうです。下記の知見をもとに、2種類のワークフローを提供します。

  • FP8版のModelを掲載した方は、CFGの値として1.5~2.5を推奨しています。

  • SamplerとSchedulerの組み合わせはeuler/simpleが標準で、dpmpp_sde/betaに変更すると改善されたように感じられました。ただし、dpmpp_sdeは生成速度が遅いです。

  • CFGの値を1.0にすると、生成速度がかなり上がる代わりにコントラストが若干低下するように見えます。

  • Modelのdtypeをdefaultからfp8_e4m3fn_fastに変更すると、品質を多少犠牲にして生成速度を上げられます。


▼ 3-3. ワークフロー(ノーマル)と生成例

 本来の設定値を用いたワークフローです。

 下記のファイルをダウンロードして、ComfyUIの画面にドラッグ&ドロップしてください。

 ワークフローの全体です。プロンプトは、既存のものをClaudeとGrokで変更しました。筆者の環境では1回あたり54秒程度かかります。

ワークフローの全体

 解像度は、公式には「2048 x 2048 (1:1)」「2304 x 1792 (4:3)」「1792 x 2304 (3:4)」「2560 x 1536 (16:9)」「1536 x 2560 (9:16)」の5種類が挙げられています。Stepsは8、CFGは2.0です。

Japanese anime-style illustration of a young girl with iconic anime features: large, expressive emerald green eyes, a soft, rounded face, and long, arranged golden blonde hair adorned with a delicate hair accessory - a small white ribbon bow with pearl details positioned on the side of her head. She displays a longing, wishful expression with slightly parted lips and dreamy eyes as she gazes at the mannequin. She is currently wearing a white three-quarter sleeve blouse with a charming pink ribbon bow at the collar, paired with dark blue jeans. She stands at the storefront entrance of a modern apparel store, turned toward a white fashion mannequin that displays a beautiful summer dress - a sleeveless midi-length dress with a soft lavender base color and very subtle, faded white flower patterns, featuring thin spaghetti straps and a fitted bodice with flowing A-line skirt. Her pose shows desire and admiration - one hand gently touching her cheek, her body slightly leaning forward toward the mannequin with a wistful, yearning posture. The store interior behind them features clean white walls, clothing racks along the walls, and display tables in the center showcasing folded garments, accessories, and other merchandise. The scene captures her moment of longing for the beautiful dress, rendered in a detailed and vibrant anime art style with warm, inviting lighting.

▼ 3-4. ワークフロー(カスタマイズ)と生成例

 筆者が独自に設定を変更したワークフローです。

  • Load Diffusion Modelのweight_dtypeを変更(default→fp8_e4m3fn_fast)

  • sampler/schedulerを変更(euler/simple→dpmpp_sde/beta)

  • cfgを1.0に変更(Negative Promptは無効)

 下記のファイルをダウンロードして、ComfyUIの画面にドラッグ&ドロップしてください。

 ワークフローの全体です。筆者の環境では1回あたり41秒程度かかります。

ワークフローの全体

 dpmpp_sde/betaに変更することで、品質が改善されていると感じられます。ただしdpmpp_sdeは速度が遅いので、weight_dtypeとCFGを変更して速度を上げています。

プロンプトは前出と同じ



■ 4. 準備(GGUF版)

▼ 4-1. はじめに

 FP8版と同じ流れで説明するので、内容が重複する部分があります。

 HunyuanImage Distilledと通常版の違いは、Modelと一部の設定のみです。また、Qwen-Imageと同じText Encoderを利用します。


▼ 4-2. モデルの設置

 筆者は種類ごとにディレクトリを分けているのでそのように記述しますが、お好みで決めていただいて構いません。

 下記のURLに掲載されたそれぞれのファイルをダウンロードして、適切なフォルダへ移動します。設置済みの場合はスキップしてください。ModelはQ4_K_Mでおよそ十分な品質となりますが、VRAM使用量と品質を見ながら上げ下げすることができます。なお、Qwen-Imageの記事で利用したものと同一のText Encoderが含まれています。


▼ 4-3. 必要な拡張機能について

 本記事に掲載したGGUF版のワークフローを読み込んだとき、拡張機能が不足していると「Some Nodes Are Missing」の表示が出ます。その場合は、ComfyUI Managerの機能を利用してインストールすることができます。

 拡張機能をインストールする方法は、下記の記事の「自動インストールの手順 」または「手動インストールの手順」を参照してください。

 本記事のワークフローでは下記の拡張機能を利用します。



■ 5. 生成(GGUF版)

▼ 5-1. 概要

 FP8版と同じ流れで説明するので、内容が重複する部分があります。

 生成用のワークフローを提供しますので、改造、再公開を含め自由にご利用ください。


▼ 5-2. 補足

 Distilledモデルは出力の品質がやや劣るので、パラメーターの調整が必要そうです。下記の知見をもとに、2種類のワークフローを提供します。

  • FP8版のModelを掲載した方は、CFGの値として1.5~2.5を推奨しています。

  • SamplerとSchedulerの組み合わせはeuler/simpleが標準で、dpmpp_sde/betaに変更すると改善されたように感じられました。ただし、dpmpp_sdeは生成速度が遅いです。

  • CFGの値を1.0にすると、生成速度がかなり上がる代わりにコントラストが若干低下するように見えます。

  • Modelのdtypeをdefaultからfp8_e4m3fn_fastに変更すると、品質を多少犠牲にして生成速度を上げられます。

 利用にあたり注意点があります。2回目以降でプロンプトを変更した場合は生成開始前に「Unload Models」を実行してください。専用GPUメモリ(VRAM)にデータが残った状態のままText Encoderのロードが行われて共有GPUメモリにはみ出し、プロンプトの処理時間がかなり長くなる場合があります。

適宜「Unload Models」を実行する

▼ 5-3. ワークフロー(ノーマル)と生成例

 本来の設定値を用いたワークフローです。

 下記のファイルをダウンロードして、ComfyUIの画面にドラッグ&ドロップしてください。

 ワークフローの全体です。プロンプトは、既存のものをClaudeとGrokで変更しました。筆者の環境では1回あたり57秒程度かかります。

ワークフローの全体

 解像度は、公式には「2048 x 2048 (1:1)」「2304 x 1792 (4:3)」「1792 x 2304 (3:4)」「2560 x 1536 (16:9)」「1536 x 2560 (9:16)」の5種類が挙げられています。Stepsは8、CFGは2.0です。

Japanese anime-style illustration of a young girl with iconic anime features: large, expressive emerald green eyes, a soft, rounded face, and long, arranged golden blonde hair adorned with a delicate hair accessory - a small white ribbon bow with pearl details positioned on the side of her head. She displays a longing, wishful expression with slightly parted lips and dreamy eyes as she gazes at the mannequin. She is currently wearing a white three-quarter sleeve blouse with a charming pink ribbon bow at the collar, paired with dark blue jeans. She stands at the storefront entrance of a modern apparel store, turned toward a white fashion mannequin that displays a beautiful summer dress - a sleeveless midi-length dress with a soft lavender base color and very subtle, faded white flower patterns, featuring thin spaghetti straps and a fitted bodice with flowing A-line skirt. Her pose shows desire and admiration - one hand gently touching her cheek, her body slightly leaning forward toward the mannequin with a wistful, yearning posture. The store interior behind them features clean white walls, clothing racks along the walls, and display tables in the center showcasing folded garments, accessories, and other merchandise. The scene captures her moment of longing for the beautiful dress, rendered in a detailed and vibrant anime art style with warm, inviting lighting.

▼ 5-4. ワークフロー(カスタム)と生成例

 筆者が独自に設定を変更したワークフローです。

  • sampler/schedulerを変更(euler/simple→dpmpp_sde/beta)

  • cfgを1.0に変更(Negative Promptは無効)

 下記のファイルをダウンロードして、ComfyUIの画面にドラッグ&ドロップしてください。

 ワークフローの全体です。筆者の環境では1回あたり55秒程度かかります。

ワークフローの全体

 dpmpp_sde/betaに変更することで、品質が改善されていると感じられます。ただしdpmpp_sdeは速度が遅いので、weight_dtypeとCFGを変更して速度を上げています。

プロンプトは前出と同じ



■ 6. おまけ

▼ 6-1. おまけ画像

 HunyuanImage-2.1で生成した画像を何枚か掲載します。ワークフローは3-4.のカスタム版を使用しています。プロンプトは通常版の記事に掲載されているものを、改めてHunyuan-PromptEnhancerを通した後のものです。

A young girl with characteristic anime features is seated at a rustic wooden table on a charming countryside cafe terrace. She possesses a softly rounded face, large and expressive brown eyes, and long, dark brown hair arranged neatly. A shy, bashful smile graces her lips, complemented by a subtle blush on her cheeks. She strikes a modest pose, with one hand resting gently against her cheek. Her attire consists of a white short-sleeved blouse with a delicate lace collar, a navy floral midi skirt, white socks, and brown loafers. On the table before her sits a coffee cup. A large speech bubble is positioned above her head, containing the Japanese text "一緒に休も?" with a line break. The background depicts a serene, remote mountain village, complete with quaint stone cottages, winding dirt paths, and vibrant wildflower meadows. Distant rolling hills are visible under a sky illuminated by the warm golden light of the hour, which casts dappled lighting across the entire scene. The image is presented as a Japanese anime-style illustration, characterized by a soft earth-tone color palette and a detailed, vibrant art style.
A joyful yet shy young Japanese woman is captured in a close-up portrait, performing elegantly while the night sky illuminates the scene. The central figure gazes directly forward, her expression a mix of delight and bashfulness. Her long, flowing silver-gray hair frames her face, meticulously styled with small double chignones, each tied with a vibrant red ribbon. She wears a flowing white dress, its fabric gently rippling from a subtle wind. She is gracefully playing a violin, with the instrument positioned near her shoulder. Ethereally floating around her are shimmering musical notes, conveying a sense of serene performance. The setting is at night, dominated by a large, luminous full moon in a deep blue sky dotted with scattered stars. In the far distance, a range of mountains is silhouetted against the sky. The soft moonlight casts a gentle glow on her figure and the immediate surroundings, which include Japanese silver grass (susuki) with their distinctive bushy seed heads, all of which are affected by the gentle wind creating soft movement. The overall image is rendered in a detailed and emotive anime art style.
An illustration captures a crisp autumn day in a park filled with maple trees, with a cheerful young girl serving as the central figure. In the foreground, the young girl stands with fair skin, large, expressive eyes, and a joyful expression. Her auburn hair is styled in intricate braids. She wears a cream-colored knit sweater dress and a bright red scarf around her neck. She holds a comically large wooden sign, on which the text "Welcome to HunyuanImage-2.1!" is written in an adorable, wobbly script. A cascade of colorful maple leaves falls from its sides, appearing as if caught in a gentle breeze. Behind her, a winding stone path is covered in a thick layer of fallen red and gold leaves. In the middle distance, other visitors are depicted taking photographs, their figures slightly blurred to create a sense of depth. The background is dominated by a dense grove of maple trees, their brilliant red and gold foliage forming a rich canopy. The entire scene is illuminated by warm, golden afternoon sunlight filtering through the leaves, casting long shadows and creating a magical, glowing atmosphere with visible light particles dancing in the air. The artwork is presented as a high-quality digital illustration in a classic anime style.
The scene is set within a grand, ornate concert hall, focusing on a young girl seated attentively at a grand piano. The central figure is a young girl with light gray hair neatly styled in two pigtails that frame her face. She wears a white ribbon blouse and a sky blue skirt that is decorated with a pattern of small white floral designs. Her delicate fingers are positioned on the ivory keys of the piano, poised as if she is about to play. Perched atop the polished, glossy black surface of the grand piano is an elegant orange tabby cat, its coat featuring distinct, darker orange stripes. In the background, the vast hall is defined by a high, vaulted ceiling, which is illuminated by the warm, golden glow cast by large, ornate crystal chandeliers. Far below, the audience occupies the rows of deep red velvet seats, whose forms are rendered as dim, indistinct silhouettes in the hall's recesses of light. The overall artistic style is a soft watercolor painting, influenced by anime character design, notable for its large, expressive eyes and delicate line work, which captures an intimate performance moment.
A young girl with smooth, fair skin is seated at a dessert buffet table within a bright and luxurious hotel restaurant. Her long, dark brown hair is styled in a high ponytail, accessorized with a simple red bow. Her facial features are characterized by large, warm brown eyes, a small nose, and a gentle, natural smile. She is wearing a pastel pink summer dress featuring delicate white lace trim along the neckline. In front of her, a white ceramic plate holds a modest assortment of desserts, including a slice of strawberry shortcake with visible layers, a small scoop of rich chocolate mousse, a couple of colorful fruit tarts, and a few pastel-colored macarons. Adjacent to her plate, a tall, clear glass contains iced tea, complete with visible ice cubes and a slice of lemon submerged within the liquid. The background is softly blurred, showing the indistinct forms of other dessert stations and the upscale interior of the restaurant, creating a shallow depth of field. The entire scene is bathed in soft, warm lighting that casts gentle highlights and creates a cozy, inviting atmosphere. This image presents a photorealistic photography style.
A slightly elevated, close-up shot captures a young Japanese woman in mid-stride while hiking on a sun-drenched mountain trail, presented from a dynamic, low-angle perspective. The central figure is the woman, her upper body and walking motion occupying the majority of the frame. She has joyful and energetic facial expressions, with her eyes looking forward, and her hair is pulled back into a high ponytail, partially covered by a simple cap or headband. She is dressed in light summer hiking gear, consisting of a colorful, short-sleeved tank top or a light, breathable hiking shirt, paired with shorts. In one hand, she grips a trekking pole, using it for balance as she moves along the path. A smaller daypack is visible, strapped to her side. The trail itself is narrow and unpaved, flanked by the lush, vibrant green vegetation of a mountainous region during summer. In the background, layers of forested slopes and distant mountain peaks are visible under a bright, clear sky. The entire scene is illuminated by bright, natural sunlight. This image presents a photography style, characterized by a shallow depth of field that keeps the subject sharp against a softly blurred background.
A professional businessman is captured in a slightly elevated, close-up shot, frozen in mid-stride as he walks with purpose across a bustling New York City street. The central figure is a man with a determined expression, his face showing confidence as he moves forward. He is dressed in a sharp, well-tailored business suit of a deep navy or charcoal gray color, worn over a crisp dress shirt and a complementary tie. In one hand, he carries a dark leather briefcase or a sleek laptop bag. The camera angle is a low diagonal, focusing on his upper body and capturing his dynamic walking motion, which creates a sense of forward momentum. He moves along a wide city sidewalk, where the pavement is lightly textured and reflective. In the background, the iconic urban scenery of New York City is rendered with a shallow depth of field, creating a soft, blurred bokeh effect. Recognizable shapes include towering skyscrapers, the vibrant yellow of a taxi cabs, and the general bustle of city life. Natural daytime lighting illuminates the scene, casting soft shadows and highlighting the textures of his suit and the surrounding environment. This image presents a photography style.
A lively scene depicts a group of diverse plush dolls engaged in conversation within a brightly lit school hallway. In the foreground, three central dolls form an animated group. The doll on the left is a girl with long brown hair styled in pigtails, wearing a classic navy blue and white sailor-style uniform with a red ribbon. She holds a miniature textbook in one hand and a tiny pencil case in the other, her fabric mouth open as if speaking. The central doll has short black hair and wears a light pink sailor uniform with a matching bow, holding a cute pink eraser and tilting her head curiously. To the right, a third doll with blonde hair in a neat bun is dressed in a blue-and-white uniform with a yellow ribbon, gesturing with one hand as she listens. The background features the hallway setting, including a row of colorful lockers lining one wall and large windows on the other, which allow warm, cheerful light to stream in. A bulletin board, scaled down to fit the doll's height, is mounted on the wall between the lockers. The entire scene is rendered in a soft-focus, 3D digital art style, emphasizing the plush texture and fuzzy details of the toys.
A surreal and cinematic landscape unfolds on an alien world, where realism and fantasy converge. In the foreground, unknown beings with elongated limbs and skin that has a glossy, liquid-like quality move gracefully through the dreamscape. They glide over shimmering metallic vegetation that shifts between organic, leaf-like plants and sharp, geometric crystal formations. The middle ground is dominated by a series of massive, spiraling towers that ascend towards the sky, their metallic surfaces twisting into complex, fantastical shapes. Above the towers, strange bioluminescent creatures, resembling living jellyfish with translucent, bell-shaped bodies and long, delicate tentacles, drift slowly through the air, emitting a soft, ethereal glow. The sky itself is a deep, cosmic expanse filled with multiple moons of varying sizes and colors, which are surrounded by vibrant aurora streams that swirl in shades of green, purple, and blue. The overall lighting shifts between warm, golden ethereal glows and cool, blue-toned illuminations, casting dramatic shadows and highlights across the entire scene. This image is rendered in a highly detailed, surreal cinematic style, reminiscent of a still from an epic science-fiction film.



■ 7. その他

 私が書いた他の記事は、メニューよりたどってください。

 ComfyUIに限定した記事の一覧もあります。

 記事に関することで何かありましたら、Xの@riddi0908までお願いします。

いいなと思ったら応援しよう!