見出し画像

ComfyUIでFP8版のHiDream-I1を試す(要VRAM 16GB)

Last update 7-9-2025
※ (7-9-2025) GGUF版の記事を書きました(12GB対応)
※ コメントにて、12GBでも動作可能との情報をいただきました。ただし本当にギリギリで、VRAMを(画面表示に使用せず)生成のみに割り当てた方が良さそうです。




■ 0. 概要

▼ 0-0. はじめに

 本記事では、ComfyUIで「HiDream-I1」を利用するための手順を説明します。FP8版のモデルでは基本的に要VRAM 16GBとなっています。なお、下記の記事は12GBに対応しています。

▼ 0-1. HiDream-I1について

 HiDream-I1はHiDream.aiの名義で公開された、香港のvivago.ai(Sparking Innovations Limited)による画像生成AIです。17Bの大きなパラメーター数を持ち(SDXLは 2.3B、SD 3.5 Largeは8B、FLUX.1は12B)、PCで利用できるものとしては最大クラスの規模です。リポジトリとモデルはMIT Licenseで4-7-2025に公開されました。

 その他の特徴として、本来の性能を持つモデルのFullのほか、少ないステップ数で生成できる蒸留モデルのDevFastがあります。また、Text EncoderはCLIP-L、CLIP-G、T5xxl、Llama-3.1-8B-Instructの4種類が採用されています。VAEはFLUX.1のものを利用しています。

 さらには4-29-2025に、編集ができるHiDream-E1 Fullも公開されました。本記事では紹介にとどめます。

▼ 0-2. 関連リンク1

 公式のリポジトリとモデルです。

▼ 0-3. 関連リンク2

 本項は、本記事で説明する内容に直接関係します。

▼ 0-4. 関連リンク3

 お試し向けのデモと有料サービスです。



■ 1. 使用リソース(FP8版)

▼ 1-1. VRAM使用量

 Comfy Orgが提供するFP8版の場合、VRAM 16GBを最大限に利用します。下記は注意事項です。

  • 筆者は起動するWebブラウザをMicrosoft Edgeに限定し、タブの数も減らすことで対策している。

  • 共有GPUメモリにはみ出すと生成時間がかなり増える。

  • アプリの起動などで、生成中にVRAM使用量が変化していっぱいになると生成処理が停止する。現在のComfyUIではそのような挙動を起こすことがあるので注意(Stepsが完全に止まり、GPUの温度が下がるので判別できる)。

16GBのVRAM(占有GPUメモリ)のほとんどを使用する

▼ 1-2. 生成時間

 下記は筆者の環境(RTX 5060 Ti)にて、標準の設定での1枚あたりの生成時間です。モデルの読み込み時間、プロンプトの処理時間は含まれていません(筆者の環境では両方合わせて20秒程度)。

  • Fast … 55秒(16 Steps)

  • Dev … 95秒(28 Steps)

  • Full … 320秒(50 Steps)



■ 2. 準備(FP8版)

▼ 2-1. 概要と補足

 情報は ComfyUI_examples/hidream に記載されていて、必要な作業はモデルの設置とワークフローの入手のみです。ComfyUIのバージョンが古いと生成できないので、その場合は更新を行ってください。

 モデルの設置場所が、過去の筆者の記事での説明と異なっているので、念のためお知らせします。なお、どちらの場所に設置しても同様に認識します。

  • models\unet\ → models\diffusion_models\

  • models\clip\ → models\text_encoders\

▼ 2-2. モデルの設置

 ModelはFullとDevとFastがあります。お好みで、いずれかまたはすべてを選んでください。筆者は、種類ごとにディレクトリを分けて設置する方針をとっています。

 Text Encoderです。通常は4種類すべてを使用します。非推奨ですが、「clip_gとllama_3.1_8b_instructのみ」や、「llama_3.1_8b_instructのみ」でも生成できるようです。

 VAEはFLUX.1と共通なので、入手済みであればダウンロードは不要です。



■ 3. 生成(FP8版)

▼ 3-1. 概要と補足

 3種類(Fast、Dev、Full)のワークフローと生成画像を掲載します。参考まで、2-2.で触れた「2種類または1種類のみのText Encoder」も試せるようにノードを用意しました。

 ワークフローは下記の画像のようになっています。

ワークフローの全体

▼ 3-2. HiDream-I1 Fast

 ワークフローを掲載します。ComfyUIの画面上にファイルをドラッグ&ドロップすると読み込まれます。

 Fastの場合、ModelSamplingは「3.0」、Stepsは「16」、CFGは「1.0」、Sampler+Schedulerは「lcm+normal」となっています。Negative Promptは使用できません。

An anime-style horizontal image of a young girl holding a soft-serve ice cream cone, viewed from a slight low-angle perspective to emphasize her gentle presence. She has large, sparkling eyes and short, light brown hair with a subtle inward curl. She wears a cozy, cream-colored knitted sweater with a ribbed collar, her expression soft and content as she looks at her ice cream. The vanilla soft-serve is perfectly swirled in a classic waffle cone, partially wrapped in a white paper sleeve with faint blue "Le Glace" lettering. Thin, golden-brown crispy snack pieces, resembling delicate potato chips, fan out elegantly atop the ice cream. Her hand, with a subtle tan, holds the cone delicately. The background is a clear, well-lit indoor shop with a warm, inviting atmosphere. Wooden shelves, sharply detailed, hold neatly arranged packaged goods and snacks in soft colors. A counter with clear signage and subtle outlines of staff or customers adds depth without focus. Soft, warm lighting enhances the gentle mood. The art style uses clean anime lines, smooth digital painting, and a soft, inviting color palette. Horizontal aspect ratio.

▼ 3-3. HiDream-I1 Dev

 ワークフローを掲載します。

 Devの場合、ModelSamplingは「6.0」、Stepsは「28」、CFGは「1.0」、Sampler+Schedulerは「lcm+normal」となっています。Fast同様、Negative Promptは使用できません。

A soft anime moment unfolds in a three-quarter profile view, capturing a bashful young woman from her temple to the top of her light summer blouse. Her long brown hair cascades in gentle waves, framing delicate features and warm, downcast eyes that shine with a hint of curiosity. A faint blush warms her cheeks as her lips curl into a shy, knowing smile. In her hands she cradles a vivid bubble tea, the cup’s pastel hues and tapioca pearls rendered with crisp clarity. On the table before her lie a neatly wrapped sandwich and a petite dessert plate. Behind her, the elegant restaurant reveals wooden tables with soft overhead lighting and relaxed diners, the depth of field kept subtle enough that details remain discernible rather than obscured. The entire scene balances her gentle presence against a refined summer restaurant atmosphere, blending serene charm with intimate warmth.

▼ 3-4. HiDream-I1 Full

 ワークフローを掲載します。

 Fullの場合、ModelSamplingは「3.0」、Stepsは「50」、CFGは「5.0」、Sampler+Schedulerは「uni_pc+simple」となっています。Fullのみ、Negative Promptが使用可となっています。

A massive curved LED display mounted on the corner of a historic Art Deco skyscraper overlooking Times Square in New York City at night, displaying a giant digital artwork of a short deform girl with medium-length gray hair reaching below her shoulder blades stands with a tiny high ponytail on the side, the rest of her hair falling naturally, and a few soft strands framing her cheerful face. She has large, expressive green eyes. She wears a white blouse with a black ribbon, a thin yellow cardigan, a light brown skirt with a white petticoat, white short socks, and pastel-colored sneakers. The displayed artwork shows her holding a cat amidst blooming spring flowers, with a traditional farm village and snow-capped mountains in the background. The artwork maintains its soft watercolor style with anime-influenced character design and large expressive eyes. The curved display glows against the Manhattan skyline, surrounded by Broadway theater marquees, yellow taxis, and the ceaseless flow of pedestrians below. The intersection buzzes with diverse crowds, street performers, and the iconic energy of the city that never sleeps.

 このプロンプトでは、車は無秩序に存在します(左右から同時に交差点へ進入する場合もある)。これは、現在の高度な画像生成AIですら難しい領域のようです。



■ 4. おまけ

▼ 4-1. i2iのワークフロー(Fast)

 HiDream-I1はFLUX.1のVAEを利用しているため、i2i(image to image)も可能です。

i2iのワークフロー

 ここでは例として、ImageFX(Imagen 3)で生成した画像を使用してみます。

Japanese anime style illustration. A girl with shoulder-length straight brown hair that cascades elegantly, with a small section tied on the right side with a cute hairpin. A delicate veil is draped over her head, extending downwards, and a sparkling tiara adorns her crown. She has large, expressive brown eyes with a lively spark, and a charming, endearing smile, with slightly rounded cheeks. She is posing in a chapel with her head tilted to the side in an artful, non-upright stance, wearing an elegant white wedding dress. The dress is a classic A-line silhouette with delicate lace detailing on the bodice and a flowing, floor-length skirt. She holds a small bouquet of white roses and baby's breath. The chapel background features soft, diffused lighting filtering through stained glass windows, creating a serene and ethereal atmosphere. Wooden pews are subtly visible in the foreground, and a gentle glow emanates from behind her, highlighting the lace of the dress.

 i2iは意外と奥が深く、入力画像、プロンプト、設定、モデルの能力のいずれも影響が大きく、どのような出力が得られるかは試してみないと分かりません。

i2iの結果(プロンプトは同じ)

 本例は比較的良好な出力にも見えますが、i2iを行う意味があまり無いように感じられます。HiDream-E1 Fullで画像を編集する方が実用的かもしれません。

 HiDream-I1でi2iを行うなら、自身の出力をインペイントで一部だけ変更するのはありかもしれません。



■ 5. その他

 私が書いた他の記事は、メニューよりたどってください。

 記事に関することで何かありましたら、Xの@riddi0908までお願いします。

いいなと思ったら応援しよう!